Quant Research Lab

Reproducible strategy experiments — explicit hypotheses, sound methods, and honestly reported results.

Self-contained research experiments on market data. The purpose is not to publish profitable strategies — it is to ask well-posed questions and answer them rigorously. A negative result, analysed carefully, is as valuable here as a positive one.

In a hurry? Read the synthesis → — what the experiments found, on one screen, and how they resolve the proposal’s central hypotheses.

Every experiment follows the same structure:

  1. Research question
  2. Hypothesis
  3. Dataset
  4. Methodology
  5. Code
  6. Results
  7. Limitations
  8. What I learned
  9. Next experiment

The tone is deliberately that of a researcher, not a signal seller. Not “this model predicts Tesla with 80% accuracy,” but: this experiment tests whether lagged return, volatility, and volume features contain statistically useful information for short-horizon directional classification, evaluated with walk-forward validation to reduce look-ahead bias.

Experiments

The original eight-experiment agenda is complete, and two further experiments extend it to RQ1 and RQ2. Each is computed and verified against the same real Nasdaq-100 data, reporting nulls as readily as positives.

The research agenda

These experiments are not a scattershot search for edges. Each one stress-tests a load-bearing claim of the PhD proposal on real data, and each is designed so that a null result is as publishable as a positive one. Together they are the empirical bridge between the libraries — which establish the methods — and the proposal, which bets on them: before proposing that regime and volatility awareness improve decisions, it is worth testing, honestly and out of sample, whether they actually do.

Each experiment maps to a proposal research question (RQ1) or hypothesis (H1, H2, H5, H6, H8), or to a capstone finding, and reports honestly whether it holds.
Experiment The question it settles Tests Status
How backtest selection inflates the Sharpe ratio How easily does searching over strategies manufacture a fake edge? discipline ✅ Done
Volatility-managed portfolios If volatility is forecastable, can you trade it? Finding 2 ✅ Done
Volatility-forecasting horizon (HAR vs GARCH vs random walk) How many days ahead is volatility actually predictable? Finding 2 ✅ Done
Regime-conditioning, out-of-sample Does conditioning on regime beat a regime-blind strategy, net of costs? H5 ✅ Done
Regime-conditional calibration Does regime-conditioning improve probabilistic forecasts, where it failed at sizing? H2 ⚠️ Superseded
Does regime-conditioning really sharpen tail probabilities? Does the H2 result survive a fair benchmark and a significance test? H2 ✅ Done
Benefit concentration at transitions Do the gains cluster around regime changes rather than in calm periods? H6 ✅ Done
Hard risk limits, net of their cost Do deterministic risk caps help after the return they sacrifice? H8 ✅ Done
Calendar anomalies after multiple-testing Do day-of-week / turn-of-month effects survive an honest correction? discipline ✅ Done
Is a market-state estimate worth anything? Do any of eight state estimators add decision value beyond a volatility-only model? RQ1 ✅ Done
Is a factored market state better than one label? At matched cardinality, does a 3×3 factored state beat a flat 9-state label? RQ2 / H1 ✅ Done

The verdicts so far sharpen the proposal rather than flatter it: forecastable volatility buys risk control, not alpha; that foresight is short — a roughly one-month half-life; and, most pointedly, conditioning on a discrete regime adds nothing beyond a volatility signal for position sizing, and no edge is statistically significant. A companion test appeared to show the same regime sharpening the probability of a dangerous day — but that result was measured against a weaker benchmark and never tested, and it has since been retracted. H5 and H2 both fail. And a test of H6 confirms it from a third angle — the risk protection of adaptivity is concentrated at regime transitions and the turbulence they open (142% of it in 38% of days, p = 0.006), earning nothing in the calm middle — so the payoff is dominated by how early the turning points are detected. Finally H8 closes the trilogy: hard risk limits turn an untradeable 2× strategy (−63% drawdown) into a viable one (−26%, higher Calmar) even as they cut returns — but only the right ones (leverage caps, volatility/CVaR budgets); a reactive drawdown kill-switch backfires, twice confirmed. And a discipline check completes the agenda: of 29 calendar anomalies, none survive an honest multiple-testing correction — the control that makes the careful positive results worth believing. Tested, not assumed.

The ninth experiment turns the same benchmark on RQ1 itself, and is the most uncomfortable of the set. Eight state estimators — moving-average crossover through Kalman filter, HMM, Markov-switching regression and two change-point detectors — were scored against a volatility-only model seeing the same features, across seven assets and 168 comparisons at a deflated bar of |t| > 3.20. State made direction forecasts worse in 53 of 56 asset-method pairs, and the risk-control claim dissolved once exposure was controlled for: the overlay improved raw drawdown in 44 of 56 pairs but Sharpe in only 4. Eighteen comparisons are significant harms against 0.23 expected by chance — adding state where it does not belong is not free. The one apparent gain points the right way on 7 of 7 assets but clears the bar on 2, and not on the asset it was discovered on. Reliability, it turns out, is not one property but several, and they do not point the same way. The entry also corrects itself in public, twice: a first version scored each method on its own span and reported harms that were really a property of 2015–2018, and a single-asset result that survived every correction at N = 24 did not survive being asked of six more assets.

The tenth turns to RQ2 and H1 — can overlapping dimensions be estimated more reliably than one mutually exclusive label? Tested at matched cardinality, so a 3×3 factored state and a flat 9-state label describe the same nine cells and differ only in price (24 parameters against 108), factoring wins 5 of 9 head-to-head comparisons and loses none. But it is more temporally stable on only one of three assets, and in 18 of 18 comparisons both descriptions are worse than using no state at all. RQ2’s answer is yes, and it barely matters: factoring is the better way to build a state representation, and still not a good enough reason to condition on one.

The eleventh is the one I least wanted to run. Every claim that had faced the full protocol had weakened or died, and every claim still standing had never faced it — including H2, this lab’s own surviving positive. Re-tested across three indices against a same-information benchmark rather than the volatility terciles it originally beat, none of six comparisons is significant, and on the asset the claim was made on the sign reverses. The published edge was 0.002 Brier with no significance test. H2 does not survive, which leaves only H6 and H8 untested and the proposal’s positive case thinner than this page claimed a day ago.

No matching items