How often are the agents right?
Every day we save what the agents said, wait, and then check it against what actually happened. No cherry-picking — every call counts.
Live Track Record
5-day horizonEach daily signal is recorded and scored 5 trading days later: a bullish call is “correct” if the price is higher, a bearish call if lower. This is the live record — it grows every day and cannot be edited. Collecting since 7/13/2026.
These four numbers are the deterministic advisor’s alone — the engine behind today’s signals. The retired editorial baseline is kept in the record for honesty and reported separately below; the two are never added together. Across every stream, 568 calls have been evaluated and 80 are pending.
Which engine produced the calls
Every row in the record is labeled with the stack that produced it, and the two streams are never blended into one number. “Deterministic advisor” is the live quant engine (measured candles, ADX, volatility, calibrated confidence) that generates today’s signals; “editorial baseline” is the fallback it replaced, kept here because its history is part of the honest record.
| Signal engine | Recorded | Evaluated | Pending | Accuracy | 95% range | Avg move |
|---|---|---|---|---|---|---|
| Deterministic advisor (the quant engine) | 119 | 87 | 32 | 49% | 39–60% | -0.292% |
| Editorial baseline (pre-engine fallback) | 529 | 481 | 48 | 49% | 45–54% | +0.756% |
Accuracy by market
No accuracy claim is made below 30 evaluated signals — until then this page reports collection progress only. For the multi-year simulation of the same logic, see the admin Backtest Lab. Past accuracy does not guarantee future results.
Does the stated confidence mean anything?
collectingThe receipt behind every confidence number: signals are grouped by what we said (stated confidence) and scored by what happened (measured hit rate). Perfect honesty means the two columns match. Calibration is reported per engine — a calibrated stream blended with an uncalibrated one would describe neither. This table is machine-readable at /api/track-record/reliability and cannot be edited.
Showing only signals recorded when the market's ATR% sat in its middle third (33-67) of the trailing 250 days (atr_pct_rank_250 recorded at signal time (migration 034)). Signals recorded before this measurement existed are excluded, not guessed — the conditioned view fills in as new signals are recorded and evaluated.
Deterministic advisor (the quant engine)
No evaluated signals yet for this engine — 12 recorded and waiting for the 5-day horizon to pass. Results appear here automatically; nothing is claimed until they do.
Reliability = measured hit rate per stated-confidence bucket over the live, append-only signal record, reported per signal stream: 'advisor_live' is the deterministic quant engine, 'seed_baseline' the editorial fallback it replaced. Buckets under 30 evaluated signals are still collecting and make no claim. Past accuracy does not guarantee future results. 'brier' splits forecast error via isotonic recalibration (CORP): miscalibration is the fixable part, discrimination is genuine ordering skill, uncertainty is the irreducible base-rate variance; identity brier = miscalibration - discrimination + uncertainty.
Do the expected-move bands cover what they promise?
measured 2026-08-25 · 2 of 12 core cells under 93%A 95% band should contain the realised move 95% of the time. This is the receipt for that promise, per pair and horizon, measured on the 2015–2026 research corpus — not on the live record, which carries a forecast volatility on 0 signals so far (collecting). The split by the forecast volatility at issue time is where the band falls short: calm periods under-cover because a 30-day volatility estimate lags the next move.
| Pair | Horizon | Days | 95% band covered | 95% range | Calm-σ̂ third | Wild-σ̂ third |
|---|---|---|---|---|---|---|
| EURUSD | 1d | 2976 | 94.3% | 93.5%–95.1% | 93.3% | 95.6% |
| EURUSD | 5d | 2972 | 93.6% | 92.7%–94.5% | 92.6% | 94.3% |
| EURUSD | 20d | 2957 | 95.1% | 94.3%–95.9% | 93.4% | 97.5% |
| GBPUSD | 1d | 2980 | 93.9% | 93.0%–94.7% | 92.7% | 95.9% |
| GBPUSD | 5d | 2976 | 93.6% | 92.7%–94.4% | 92.0% | 95.6% |
| GBPUSD | 20d | 2961 | 94.9% | 94.0%–95.6% | 93.7% | 96.7% |
| USDJPY | 1d | 2993 | 93.5% | 92.5%–94.3% | 91.8% | 94.4% |
| USDJPY | 5d | 2989 | 92.6% | 91.6%–93.5% | 88.5% | 95.1% |
| USDJPY | 20d | 2974 | 92.7% | 91.7%–93.6% | 89.2% | 95.5% |
| AUDUSD | 1d | 2975 | 94.3% | 93.4%–95.1% | 93.4% | 94.8% |
| AUDUSD | 5d | 2971 | 95.0% | 94.2%–95.7% | 92.5% | 96.5% |
| AUDUSD | 20d | 2956 | 96.0% | 95.2%–96.7% | 91.1% | 98.4% |
Replication on the other 6 pairs
| Pair | Horizon | Days | 95% band covered | 95% range | Calm-σ̂ third | Wild-σ̂ third |
|---|---|---|---|---|---|---|
| USDCAD | 1d | 3001 | 94.4% | 93.6%–95.2% | 92.9% | 96.2% |
| USDCAD | 5d | 2997 | 94.6% | 93.7%–95.3% | 92.6% | 95.9% |
| USDCAD | 20d | 2982 | 94.9% | 94.1%–95.7% | 92.0% | 97.6% |
| USDCHF | 1d | 2984 | 93.7% | 92.7%–94.5% | 90.7% | 95.8% |
| USDCHF | 5d | 2980 | 92.7% | 91.7%–93.6% | 88.4% | 95.2% |
| USDCHF | 20d | 2965 | 95.2% | 94.3%–95.9% | 91.4% | 99.2% |
| NZDUSD | 1d | 2991 | 94.7% | 93.9%–95.5% | 93.4% | 95.7% |
| NZDUSD | 5d | 2987 | 94.5% | 93.6%–95.2% | 92.8% | 95.5% |
| NZDUSD | 20d | 2972 | 95.5% | 94.7%–96.2% | 92.1% | 98.4% |
| EURGBP | 1d | 2978 | 94.1% | 93.2%–94.8% | 93.3% | 95.9% |
| EURGBP | 5d | 2974 | 94.5% | 93.6%–95.3% | 94.7% | 94.7% |
| EURGBP | 20d | 2959 | 96.5% | 95.8%–97.2% | 97.1% | 96.5% |
| EURJPY | 1d | 2977 | 94.8% | 93.9%–95.5% | 94.0% | 96.5% |
| EURJPY | 5d | 2973 | 94.0% | 93.1%–94.8% | 93.1% | 96.6% |
| EURJPY | 20d | 2958 | 95.5% | 94.7%–96.2% | 96.7% | 98.6% |
| GBPJPY | 1d | 2981 | 93.6% | 92.7%–94.5% | 90.7% | 95.8% |
| GBPJPY | 5d | 2977 | 92.2% | 91.1%–93.1% | 88.1% | 94.9% |
| GBPJPY | 20d | 2962 | 92.0% | 91.0%–93.0% | 85.7% | 95.5% |
Band: ±1.96·σ̂_t·√k on ln close; σ̂ = compute_market_metrics realizedPct (median c2c/Parkinson/GK/EWMA over 30 daily NY-roll candles). Method: empirical coverage of the central 95% band with Wilson 95% CI (block-k bootstrap CI for k>1); σ̂-tercile split by the forecast σ̂ at issue time; split-half 2015–2020 vs 2021–2026. Source: docs/experiments/2026-08-25-interval-coverage-crps.md (+ -results.json); regenerate with 2026-08-25-interval-coverage-export.py. Coverage measured on historical data does not guarantee future coverage.
Can people actually read this?
measured, not assertedWe claim the product explains itself to beginners. Onboarding ends with three questions about how to read a card — what a model score is, what a stop means, and what “follow with paper money” does. These are the answers, including the wrong ones. A rate only counts as a claim at 50 answers per mode.
No one has taken the check yet. It appears at the end of onboarding.
Comprehension is a claim only at N >= 50 per mode; below that this is collection progress.