Quant Lab — What Eleven Experiments Found
The synthesis: forecastable volatility, regimes that assess rather than size, and the discipline that tells the two from noise
The Quant Lab set out to do one thing the libraries could not: test the proposal’s bets on real data, honestly, with a real chance of a null. Ten experiments later, the answers cohere into a single argument — and, because several came back as nulls, one that has already changed the proposal’s framing rather than merely decorated it. This page is the whole of it, on one screen.
The eleven verdicts
| Experiment | Tests | Verdict |
|---|---|---|
| Backtest overfitting | discipline | The best of N strategies looks great from noise alone; the max Sharpe grows like \sqrt{2\ln N} — hence the deflated Sharpe ratio. |
| Volatility-managed portfolios | Finding 2 | Forecastable volatility buys risk control, not alpha — drawdown roughly halved, but no statistically significant excess return. |
| Volatility-forecasting horizon | Finding 2 | That foresight is short: a volatility shock half-lives in ~1 month; GARCH ≈ HAR, both beating persistence more at longer horizons. |
| Regime-conditioning, out-of-sample | H5 | For sizing, a discrete regime adds nothing beyond volatility — no significant edge over a regime-blind vol scaler. |
| Regime-conditional calibration | H2 | ⚠️ Superseded. Reported that regime-conditioning sharpens risk probabilities (Brier 0.155 → 0.153) — a 0.002 edge, never significance-tested, against a benchmark cruder than the treatment. |
| Does regime-conditioning really sharpen tail probabilities? | H2 | Re-tested with a same-information benchmark and DM tests: 0 of 6 comparisons significant, and the sign reverses on the asset the claim was made on. H2 does not survive. |
| Benefit at transitions | H6 | The risk protection concentrates at transitions and the turbulence they open — 142% of it in 38% of days (p = 0.006), nothing in the calm middle. |
| Hard risk limits | H8 | Leverage caps and volatility/CVaR budgets make an untradeable strategy viable (−63% → −26% drawdown); reactive drawdown kill-switches backfire. |
| Calendar anomalies | discipline | Of 29 seasonal effects, none survive an honest multiple-testing correction; the largest |t| is exactly \sqrt{2\ln N}. |
| Is a factored market state better than one label? | RQ2 / H1 | At matched cardinality a 3×3 factored state beats a flat 9-state label 5 of 9 head-to-head and loses none — but is more stable on only 1 of 3 assets, and 18 of 18 comparisons say both lose to no state at all. |
| Is a market-state estimate worth anything? | RQ1 | Across 7 assets and 168 comparisons, state makes direction forecasts worse in 53 of 56 pairs and adds no risk-adjusted value; 18 comparisons are significant harms against 0.23 expected, and the one apparent gain clears the bar on 2 of 7 assets. |
The arc, in one paragraph
Volatility is the one thing this market lets you forecast — but only its risk, not its return, and only for about a month. A regime built on that volatility is therefore worth little as a position-sizing device (it merely repackages the volatility a continuous signal already captures). For a while it looked genuinely valuable as an assessment device — sharpening the probability of a dangerous day — but that claim did not survive a fair benchmark and a significance test, and has been retracted. And the payoff of all this adaptivity is not spread evenly: it is concentrated at the regime transitions, earned in the weeks after the market turns and nowhere in the calm between, which is why detecting the turn early matters far more than reacting to it. Around all of it, hard, forward-looking risk limits are what make an aggressive strategy actually survivable — while the two discipline experiments stand guard, confirming that the best backtest and the best calendar effect are exactly what pure noise produces.
The ninth experiment tests the premise all of that rests on — that a market state can be identified usefully in the first place — and returns the hardest verdict of the set. Put eight estimators, simple to sophisticated, against a volatility-only model with the same information, and every one degrades direction forecasting; the single calibration gain that survives correction belongs to the method that changes its mind every five days. Reliability decomposes into recovery, stability, timeliness and decision value, and these do not co-vary: the best estimator by one is the worst by another. That is not a reason to abandon state — it is the reason RQ1 must be answered by measurement across competing methods rather than by choosing one and defending it.
The tenth asks the structural version of the same question — RQ2/H1, whether overlapping dimensions beat one mutually exclusive label. At matched cardinality, where a 3×3 factored state and a flat 9-state label describe the same nine cells for 24 parameters against 108, factoring wins 5 of 9 head-to-head and loses none. It is not reliably more stable, and both descriptions lose to no state at all in 18 of 18 comparisons. The proposal’s preference for a multidimensional state is vindicated as a modelling choice and undercut as a source of edge — which is the same shape as every other finding here.
What it means for the proposal
The six tests that bear on the proposal directly (H1, H2, H5, H6, H8 and RQ1) point repeatedly at the same redesign, and it is a sharper — and thinner — proposal for it. The Regime Agent’s job is detection and assessment at the turning points — not sizing in the stable middle: a discrete regime earns its place by flagging when the market changes (change-point detection) and by conditioning relationships a scalar cannot (Markov-switching). The calibration leg of that argument is gone — H2 did not survive re-testing — so detection and relationship-conditioning are what is left to defend, and neither has yet been tested to this standard. Position sizing is left to volatility and to the Risk Agent’s hard, forward-looking limits, which the evidence shows are what deliver viability. RQ1 adds the constraint that keeps this honest: which estimator supplies that state is not a detail to be settled by preference, because the panel’s rankings on stability and on decision value are close to opposite. The proposal therefore commits to a benchmark protocol — every method scored on the same data, against the same volatility-only benchmark, with a deflated bar — rather than to a favoured model. None of this was assumed; it was tested, including where the test said no, and including when the test was pointed at this lab’s own best result.
The honest summary as it stands: every hypothesis that has faced the full protocol has returned a null or a near-null. H6 and H8 have not faced it. Until they do, the proposal’s positive case should be read as untested rather than established — which is a less comfortable position than this page held a day ago, and a more accurate one.
Every figure in the first eight experiments is computed and verified against the same multi_daily.csv Nasdaq-100 dataset; the ninth uses a seven-asset panel to 2026-08-24 and the tenth three indices from 2015. All are out-of-sample where a forecast or strategy is involved, and report nulls as readily as positives. Start from any row of the table above.