Detection Quality
The bottom line: on data the system had never seen, when it flags a market gap, it’s right ~78% of the time.
This page measures the prediction itself — not trading, not returns. A prediction counts as correct only if the market actually moved 2%+ in the direction it called, within the window. Scored against real historical prices, no AI in the scoring, reproducible to the digit.
The number that matters
The system’s rules were built on market history through 2022. So its record before 2023 is partly “graded on data it studied.” The only honest test of future performance is 2023 onward — signals it never saw:
| On unseen data (2023+) | ||
|---|---|---|
| Hit rate when it fires | 78% | of high-conviction gaps came true |
| Detection | 100% | every real gap was detected |
| High-conviction share | 62% | of correct gaps cleared the “act on it” bar |
| Sample | 92 high-conviction signals |
78% correct on data it never saw is the number to judge the project on.
It detected every real gap. The 62% is not a miss rate — it’s how many of the correct gaps were flagged at high conviction. The other 38% were still detected and still right, just scored below the “act on it” threshold. That threshold is a dial: lower it and more gaps become actionable, at the cost of a slightly lower hit rate.
We report performance on unseen data rather than a blended lifetime average, because only unseen data reflects how the system will do on the next call.
Two terms, plainly
- Hit rate (precision) — when the alarm rings, is there really a fire? A false alarm = a losing trade, so this is the one that costs money. This is what we optimise.
- Coverage (recall) — of all the real fires it detected, how many did it flag at high enough conviction to act on? (It detected every one; this is about how many cleared the confidence bar.) A low-conviction call only costs a missed opportunity, not capital.
Did it hold up out-of-sample? Yes.
| Built-on (pre-2023) | Unseen (2023+) | Verdict | |
|---|---|---|---|
| Hit rate | 86% | 78% | Holds — the gap is statistical noise, not real decay |
| High-conviction share | 78% | 62% | Fewer recent gaps cleared the conviction bar (all were still detected) |
If the system had merely memorised the past, its hit rate would collapse on new data. It didn’t. That’s the core result: it detects a real, repeatable pattern.
Where it’s strongest
| What it detects | Hit rate | Sample |
|---|---|---|
| Oil / gold / gas shocks | 100% | 64 |
| Currency-peg breaks | 76% | 29 |
| Central-bank policy gaps | 88% | 67 |
| Sovereign-debt stress | 80% | 105 |
| Real estate | 80% | 69 |
| Geopolitical shocks | 73% | 33 |
Commodities are flawless — 64 for 64, zero false alarms. Geopolitical is the hardest and the weakest.
Not a one-off era
Hit rate is stable across every market regime in the 12-year window — 77% to 92%, including the COVID shock and the rate-hike cycle. No single lucky period is carrying the result.
Every figure is reproducible from the committed signal record and price cache.