PhD Research Proposal

Adaptive Market Intelligence — a probabilistic, multi-agent architecture for real-time market-state inference, strategy suitability and risk-governed decision-making. The full proposal, the evidence already built for it, and the road ahead.

Working title: Adaptive Market Intelligence: A Probabilistic Multi-Agent Architecture for Real-Time Market-State Inference, Strategy Selection and Risk-Governed Financial Decision-Making (academic alternative: Probabilistic Multi-Agent Market-State Inference and Adaptive Strategy Selection under Non-Stationarity and Uncertainty). Proposed award: PhD, three years. This is the updated working document (August 2026) developed in the open; the formatted version is available as a PDF for supervisors.

Central proposition. Financial-market systems should first infer what type of market is currently being observed, quantify the uncertainty in that assessment, and identify which strategy families are suitable — before moving to any directional decision. They should not move directly from raw data to trading actions.

Proposal at a glance

Field Updated proposal
Primary contribution A hierarchical architecture that converts heterogeneous real-time data into calibrated specialist evidence, competing market-state hypotheses, consensus estimates and regime-conditioned strategy suitability — under deterministic risk control.
Primary domain Quantitative finance, machine learning, financial econometrics and agentic decision systems.
Core methods Regime and change-point detection, probabilistic calibration, ensemble and expert-weighting methods, financial NLP, constrained agent coordination, strategy selection and formal risk control.
Initial empirical scope The NASDAQ-100 ecosystem — NQ futures, QQQ, volatility, rates, breadth, liquidity, cross-asset and macro-event information.
Commercial pathway An institutional market-intelligence and decision-support platform, with later extensions into allocation, risk overlays and execution support.

Executive summary

This research proposes a real-time Adaptive Market Intelligence architecture designed to infer the current market state before selecting a strategy or action. The central problem is not framed as direct price prediction. It is framed as point-in-time inference under non-stationarity: what kind of market is currently being observed, how certain is that assessment, which evidence supports or contradicts it, and which strategy families are appropriate under those conditions?

The architecture ingests synchronised market, microstructure, volatility, cross-asset, macroeconomic and textual data. Specialist quantitative models transform those inputs into bounded, calibrated evidence about trend, volatility, liquidity, structural change, macro drivers and cross-asset confirmation. A multi-agent reasoning layer then compares competing market-state hypotheses rather than averaging incompatible model scores. The resulting market-state representation is multidimensional and probabilistic — it can identify, for example, a macro-driven risk-off transition with expanding volatility, weakened but orderly liquidity and elevated uncertainty. A separate strategy council assesses the suitability of trend, mean-reversion, long-volatility, statistical-arbitrage, market-neutral and cash-preservation strategies, and a deterministic risk agent retains final authority to scale, reject or restrict any proposed action.

The academic contribution is the design and rigorous evaluation of a calibrated, hierarchical and auditable market-state inference system. The commercial contribution is a blueprint for an institutional market-intelligence platform providing real-time regime context, transition alerts, strategy suitability, risk recommendations, evidence traceability and decision history. Unusually for a proposal, much of its methodological foundation is already built and verified on this site, entry by entry — see the evidence tracker, which is also honest about the specialist methods still to add.

What changed from the previous outline

This is a deliberate reframing of an earlier, more autonomy-centred proposal. The revision sharpens the research question, demotes reinforcement learning from a pillar to an optional late-stage comparison, and makes the commercial thesis explicit.

Previous emphasis Updated emphasis
One principal Market Regime Agent Multiple specialist evidence models — trend, volatility, liquidity, macro narrative, cross-asset confirmation and structural change
Combined state passed directly to an RL trading agent An explicit hypothesis-generation and consensus layer before any strategy decision
Primary output was exposure or a trade action Primary output is probabilistic market-state intelligence and strategy suitability
RL central from an early stage RL is optional, introduced only after regime inference and deterministic allocation baselines are validated
Agents loosely described as functional modules A clear distinction between quantitative models, specialist agents, a consensus mechanism, strategy agents and deterministic risk control
Commercial value implied Commercial product, user outputs, proprietary data assets and a staged route to market made explicit

The research problem

Financial markets are non-stationary and only partially observable: relationships between returns, volatility, liquidity, order flow, rates, macro information and investor behaviour vary through time. A rule that works in a liquid directional trend can fail in an event-driven repricing, a volatility shock or a liquidity withdrawal. Yet many financial-ML studies frame the central task as predicting the next return or choosing buy/sell/hold — which produces systems that look strong in-sample and fail when the process changes. The market state is not a single label such as bull, bear or sideways; it is better represented through simultaneous dimensions — directionality, volatility, liquidity, dominant driver, transition risk and cross-asset confirmation.

Eight specific weaknesses motivate the work: static models apply historical relationships after the environment changes; single-label regimes oversimplify overlapping conditions; numerical and textual information are reconciled informally; model confidence is frequently uncalibrated; ensembles average outputs even when models answer different questions; systems move from observation to action without explicit hypothesis testing; RL is often tasked with trading before the state representation is shown to be stable; and commercial systems require auditability, latency control, deterministic risk limits and evidence traceability — not only backtest returns.

Central research problem. Can a hierarchical system of specialist quantitative models and bounded reasoning agents infer multidimensional financial-market states and transitions in real time, with calibrated uncertainty, and can those inferences improve out-of-sample strategy selection and risk management relative to static, regime-unaware and single-model alternatives?

Aim and objectives

The aim is to design, develop and evaluate a real-time Adaptive Market Intelligence framework that converts heterogeneous point-in-time data into calibrated specialist evidence, reconciles competing market-state hypotheses, estimates multidimensional market conditions, and supports regime-conditioned strategy selection under deterministic risk constraints.

# Objective
O1 Define an economically interpretable multidimensional taxonomy of market states and transitions.
O2 Construct a timestamp-safe data architecture combining market, microstructure, volatility, cross-asset, macro and textual information.
O3 Develop specialist model families for trend, volatility, liquidity, structural change, macro narrative and cross-asset confirmation.
O4 Calibrate specialist probabilities and estimate reliability conditional on data quality, horizon and previously observed conditions.
O5 Develop a bounded multi-agent consensus mechanism that compares competing market-state hypotheses and records supporting and contradictory evidence.
O6 Evaluate whether inferred market states improve strategy suitability, exposure management and risk outcomes versus regime-unaware alternatives.
O7 Test whether adaptive expert weighting provides incremental value over transparent static weighting and Bayesian or ensemble baselines.
O8 Design an auditable decision and review framework tracing each output from source data to models, consensus, strategy recommendation and realised outcome.
O9 Assess the architecture as the technical and intellectual-property foundation for an institutional market-intelligence platform.

Research questions and hypotheses

Principal research question. Can calibrated specialist models and a bounded multi-agent consensus architecture identify multidimensional market states and regime transitions in real time, and does this information improve out-of-sample strategy selection, risk control and decision robustness?

The sharp, testable core. Within that broad aim sits one identified, falsifiable claim the thesis is built to settle: does conditioning on a real-time, filtered market-state estimate improve the out-of-sample calibration and risk-adjusted robustness of equity decisions — concentrated at regime transitions — beyond a continuous volatility signal using the same information, after costs and multiple-testing correction? The load-bearing clause is “beyond a continuous volatility signal”: it fixes the benchmark and names the null the Quant Lab has already tested — a discrete regime added nothing beyond volatility for sizing, yet sharpened calibration and paid at transitions. Identification rests on a same-information benchmark, filtered point-in-time state, orthogonalisation against volatility and volume, a pre-committed metric, and a survivorship-bias-free equity cross-section (CRSP/Compustat) held to deflated-Sharpe and PBO standards. This falsifiable core gives the broader architecture an empirical spine without narrowing its ambition.

A dash means the question is proposed but not yet tested. Five of the ten already have out-of-sample evidence in the Quant Lab — and where it came back negative, the proposal was changed rather than the result.
Ref. Supporting research question Evidence so far
RQ1 Which statistical, econometric and ML methods most reliably identify trend, volatility, liquidity and structural-change states out of sample? Decision value of a state estimate
RQ2 Can overlapping market dimensions be estimated more reliably than a single mutually exclusive regime label? Factored vs flat state — yes at matched cardinality, but both lose to no state at all
RQ3 How should probabilities from heterogeneous models be calibrated, normalised and compared across different horizons? Partial — heterogeneous models compared across horizons
RQ4 Can structured information from announcements and financial text improve identification of market drivers beyond numerical data?
RQ5 Does explicit comparison of competing hypotheses outperform simple averaging, stacking or voting ensembles?
RQ6 Does market-state information improve the selection, activation and risk scaling of strategy families after realistic costs? Regime-conditioning OOS, hard risk limits
RQ7 Can expert reliability be estimated conditionally, so models receive different weights in different environments?
RQ8 Are adaptive-weighting or RL methods superior to transparent deterministic and Bayesian weighting methods?
RQ9 How stable, explainable and auditable are the outputs across assets, periods, latency assumptions and repeated trials? Backtest selection, calendar anomalies
RQ10 Which components create commercially useful decision intelligence even where direct trading performance is not consistently superior?

The testable hypotheses follow directly: H1 a multidimensional state is more temporally stable and useful than a single-label regime; H2 calibrated ensemble probabilities beat uncalibrated confidence out of sample; H3 macro-event and narrative features add value primarily during event-driven and transition periods; H4 explicit hypothesis reconciliation improves state classification and uncertainty over voting or averaging; H5 regime-conditioned strategy selection improves risk-adjusted performance, drawdown or tail behaviour versus static allocation; H6 the greatest incremental value arises at regime transitions and during specialist disagreement; H7 condition-specific expert weighting beats equal weighting, but complex adaptive methods do not necessarily beat well-designed transparent baselines; and H8 deterministic risk controls improve robustness and viability even where they reduce gross return.

These are not left as promises. The Quant Research Lab has already stress-tested the load-bearing ones on real data, and the results are shaping the design rather than being fitted to it. A deliberately fair out-of-sample test — a real-time, no-look-ahead regime-conditional strategy pitted against a regime-blind volatility scaler using the same information — found that, for position sizing, a discrete regime adds nothing statistically distinguishable beyond volatility (experiment). A companion test suggested regime conditioning did sharpen probability calibration — but that claim was measured against a benchmark cruder than the treatment and never significance-tested, and a re-test across three indices found nothing significant in six comparisons. H2 is retracted. What remains standing is the concentration of benefit at transitions and the value of hard risk limits — and neither has yet been tested to the standard that overturned the others.

RQ1 itself has since been put to the same test. Eight state estimators — from a moving-average crossover to a Kalman filter, a hidden Markov model, a Markov-switching regression and two change-point detectors — were scored against a volatility-only benchmark seeing the same features (experiment). Across seven assets and 168 comparisons, state made direction forecasts worse in 53 of 56 pairs, eighteen comparisons were significant harms against 0.23 expected by chance, and the single apparent calibration gain — from the least stable estimator in the panel — cleared the deflated bar on only two assets, not including the one it was found on. That is a result RQ1 has to explain rather than average away, and it is why the proposal treats reliability as a question to be measured across competing methods rather than a property to be assumed of any one of them. A proposal whose own experiments can, and do, return nulls is one built on evidence rather than hope.

Proposed conceptual architecture

The system is a hierarchical decision process: data → point-in-time platform → specialist evidence → standardised, calibrated evidence → competing hypotheses → consensus and uncertainty → strategy suitability → deterministic risk → outcome review. Crucially, the proposal distinguishes a model from an agent: a model estimates a defined statistical quantity; an agent owns a bounded task, may consult several models, applies validation rules, and communicates through a fixed schema. This distinction is deliberate — it prevents “agent” terminology from becoming a substitute for methodological precision.

Five-stage flowchart: load point-in-time data; apply specialist econometric and ML models in parallel; agents weigh each output into a confidence-weighted hypothesis; an agent council reconciles these into one market-state posterior; a risk-governed decision follows. A feedback loop returns realized outcomes to update the reliability weights.
Figure 1: The Adaptive Market Intelligence pipeline — five decision stages condensing the nine-layer stack below. Point-in-time data becomes calibrated specialist evidence, competing hypotheses, one reconciled market-state posterior, and a risk-governed action; realized outcomes feed back to re-weight the agents. Consistent with the evidence already built, market state is treated as a calibration- and risk-control signal, not a return-timing one — the LLM is a measured input, never an executor — and the online-learning loop is evaluated, not assumed to pay.
The nine-layer decision stack.
Layer Function Illustrative output
1 Data acquisition Real-time and historical feeds with source, event and available-at timestamps
2 Point-in-time data platform Raw, cleaned and research-ready layers; synchronisation, quality flags, feature lineage
3 Specialist quantitative models Trend, volatility, liquidity, structural-change, macro-narrative and cross-asset evidence
4 Evidence standardisation Calibrated probability, reliability, horizon, freshness, model version, support and anomaly warnings
5 Market-state hypothesis engine Construction and comparison of competing explanations of the current market
6 Consensus and uncertainty Posterior state estimate, disagreement, transition probability and confidence
7 Strategy council Suitability for trend, mean-reversion, volatility, stat-arb, market-neutral and cash preservation
8 Risk governance Deterministic exposure, leverage, liquidity, drawdown, concentration and execution constraints
9 Outcome review and learning State accuracy, strategy suitability, model contribution, overrides and realised outcomes

Specialist evidence domains

Each specialist owns one evidence domain, a defined data set, candidate methods and a structured, schema-constrained output — never a raw trading instruction.

Specialist Candidate methods Structured output
Trend & persistence HMM, Kalman filter & state-space, gradient boosting Direction, persistence, transition probability, model agreement
Volatility GARCH family, EGARCH & HAR-RV, boosting, sequence models Expansion probability, expected horizon, tail-risk condition
Liquidity & microstructure Anomaly detection, supervised stress models, order-book/impact features Liquidity condition, execution risk, stress probability
Structural change Bayesian online change-point, CUSUM & drift detection Break probability, detection delay, affected variables
Macro & narrative Schema-constrained LLM, event/surprise extraction Event class, direction, novelty, duration, affected assets, evidence
Cross-asset confirmation Dynamic correlation (DCC), factor & graphical methods Confirming/contradicting markets; risk-on vs risk-off evidence

An operational design note — how this architecture becomes a monitored, human-approved interface, and how it would be built — is set out separately in The AMI Console.

Data architecture and a worked example

The initial prototype focuses on the NASDAQ-100 ecosystem — rich enough to test macro sensitivity, tech concentration, volatility, rates and cross-asset interaction, yet narrow enough for a three-year programme. Inputs include NQ futures and QQQ prices/quotes, constituent returns and breadth, VIX/VXN and options-derived volatility, 2y/10y Treasury yields, the dollar and cross-asset risk indicators, scheduled US macro releases, and timestamped news where licensing permits. Data is organised in three layers — Bronze (unmodified vendor messages with source and receive timestamps and quality flags), Silver (deduplicated, normalised, time-aligned observations with corporate-action and session handling), and Gold (feature vectors, labels, model outputs, calibration data, state estimates and decision records) — so that every feature has an auditable, point-in-time lineage.

The value of this framing is easiest to see concretely. Suppose that one minute after a hawkish inflation surprise, the specialists report:

Specialist State Prob. Key evidence
Trend Downward 0.78 Negative returns, weak breadth, price below VWAP
Volatility Expansion 0.91 Realised-vol jump, volatility-index confirmation
Liquidity Deteriorating 0.74 Wider spreads, lower depth, negative imbalance
Macro narrative Hawkish repricing 0.84 Inflation surprise, rapid yield adjustment
Cross-asset Risk-off 0.69 Yields, dollar and volatility broadly confirm

The hypothesis engine then compares competing explanations rather than averaging these scores — an ordinary technical pullback (posterior 18%), a macro-driven risk-off transition (76%), or a full liquidity crisis (6%) — reconciling supporting and contradictory evidence. The consensus layer emits a final state: macro-driven risk-off transition (76%), volatility expansion (83%), transition probability 71%, liquidity weakened but orderly, overall confidence 73%. Only then does the strategy council assess suitability — directional trend 82%, long volatility 71%, market-neutral 54%, stat-arb 39%, mean-reversion 17% (inactive) — and the risk agent applies non-negotiable limits.

Strategy suitability and risk governance

The strategy layer is deliberately separated from market-state inference. Strategy agents receive the same consensus state and independently judge whether their approach should be active, inactive or risk-reduced. This makes it possible to test whether market-state intelligence is valuable without claiming any single model forecasts every environment. Above all sits a deterministic risk authority with final say: maximum gross/net exposure; volatility- and liquidity-conditional position-size limits; leverage, concentration and correlation caps; maximum-drawdown and Expected-Shortfall controls; minimum-liquidity and maximum-spread thresholds; scheduled-event and overnight restrictions; explicit transaction-cost and slippage assumptions; and a mandatory no-action state when evidence is stale, contradictory or outside the training distribution. The AI Trading Systems section develops why an LLM is never granted execution authority and what a layered control stack requires.

Methodology

The project follows a staged empirical design. Each component is evaluated independently first; integration proceeds only where ablation evidence demonstrates incremental value; and point-in-time integrity, transparent baselines and out-of-sample performance are prioritised over architectural complexity. The work is organised as six studies: (1) market-state taxonomy and detection; (2) calibrated specialist evidence; (3) macro and narrative contribution; (4) consensus and hypothesis reconciliation; (5) strategy suitability and risk; and (6) the integrated system, governance and commercial prototype.

Baselines are transparent and demanding: rule-based volatility/trend/liquidity states; HMM and Markov-switching models; K-means and Gaussian mixtures; Bayesian online change-point detection; logistic regression, random forests and gradient boosting; equal-weight, majority-vote, Bayesian model averaging and stacking ensembles; static versus regime-conditioned allocation; and single-agent, regime-unaware comparators. Validation enforces chronological splits, walk-forward and rolling windows, purged cross-validation with embargo, available-at timestamps and frozen information sets, transaction costs, spreads, slippage and delayed execution, regime- and transition-specific evaluation, repeated seeds, stress testing across monetary-policy and volatility environments, and out-of-distribution and staleness testing.

Reinforcement learning is not assumed to be necessary. It is tested only after the state representation, specialist evidence and deterministic strategy mappings are validated — as a candidate for expert weighting, strategy allocation, exposure scaling or conflict resolution — and compared directly against transparent rules, contextual bandits, Bayesian updating and supervised meta-models. A finding that simpler methods win is an academically valid result. Where a risk-sensitive objective is used, it penalises variance, drawdown, cost and turnover rather than maximising raw profit,

R_t \;=\; r_t \;-\; \lambda_1 \sigma_t^2 \;-\; \lambda_2\,\mathrm{DD}_t \;-\; \lambda_3\,\mathrm{TC}_t \;-\; \lambda_4\,\mathrm{TO}_t,

and, because the true state is latent, the formal setting is a partially observable process in which the system maintains a belief b_t(s) = P(S_t = s \mid o_{1:t}) over states — the object the MDP entry formalises and the specialists estimate.

Evaluation framework

Evaluation spans nine areas: state identification (accuracy where labels exist, adjusted Rand index, persistence, stability, detection delay, economic interpretability); probabilistic quality (log loss, Brier score, ECE, reliability diagrams, sharpness); transition detection (lead/delay, false alarms, missed transitions); consensus quality (hypothesis ranking, disagreement resolution, contradiction handling); strategy suitability (Sharpe, Sortino, Calmar, drawdown, Expected Shortfall, turnover, tail behaviour); execution realism (spread, slippage, impact, latency, implementation shortfall); agent contribution (ablation and Shapley-style attribution, overrides); robustness (block-bootstrap intervals, multiple-testing controls, Deflated Sharpe and PBO); and governance (traceability, reproducibility, schema compliance, source attribution, override frequency). Commercial usefulness — alert timeliness, interpretability, actionability, stability and value without direct execution — is evaluated as a first-class outcome, not an afterthought.

What is already built: the evidence tracker

This architecture’s methodological base now exists on this site as verified, tested entries. The tracker maps each layer and specialist domain to the evidence. The econometrics-heavier specialist set the updated proposal introduced — state-space, asymmetric and long-memory volatility, multivariate cross-asset dependence, and ensemble and expert-weighting methods — has been built and verified alongside the original foundation, so every row below is now green; the work that remains is depth and data, not coverage.

Legend: ✅ built and verified · ◐ partially built (needs extension) · ○ named in the proposal, not yet on the site.

Every row is now built and verified — the specialist and combination layer the updated architecture adds is complete, each mapped to its evidence. The three-year plan below rebuilds this prototype core on a proper cross-section.
Architecture component Status Evidence / what remains
Regime detection — clustering & mixtures K-means, GMM recover calm/sell-off/crisis states
Regime detection — HMM & Markov-switching 2-state HMM, point-in-time filter; beta doubling by regime
Structural change — online change-point BOCPD flags COVID and 2022 onset point-in-time
Structural change — CUSUM & drift detection CUSUM & concept drift — the cheap online tripwire; flags COVID 2020-02-25 (one day after BOCPD) and 2022, with the ARL false-alarm/delay trade-off made explicit
Trend specialist — state-space / Kalman filter Kalman filter & state-space — from scratch, verified vs statsmodels to 2.8e-10; real-time time-varying β (median 1.04, drifting 0.67–1.27) with honest ±2σ bands and the filtered-vs-smoothed look-ahead
Volatility — GARCH & dynamics GARCH, Hurst, stationarity/ADF
Volatility — EGARCH/GJR & HAR-RV EGARCH, GJR & HAR-RV — leverage effect significant (GJR γ t 4.2, EGARCH t −5.7; BIC 9202→9118), and HAR-RV on a realized measure forecasts best OOS (QLIKE 0.70 vs GARCH 0.87)
Liquidity specialist — stress / anomaly detection Liquidity-stress anomaly detection — Isolation Forest flags 5% of days, 5.4× into crises, preceding 41.6% vs 17.3% forward vol; catches “quiet illiquidity” a return threshold misses
Cross-asset confirmation — DCC, factor & graphical models Cross-asset dependence — from-scratch DCC (a+b=0.99) shows correlation tripling calm→crisis (0.23→0.76) and the defensive staple re-coupling (0.31→0.70): the risk-on/off signal
Macro & narrative specialist Schema-constrained extraction, narrative/surprise features, transformers/embeddings
Evidence standardisation — calibration Calibration & conformal prediction (ECE 0.105→0.062)
Conditional expert reliability Contextual bandits — LinUCB learns which expert to trust per regime, recovering the routing exactly when the structure exists (0.75 = oracle); honest null when it’s weak
Consensus — ensemble baselines (equal, vote, BMA, stacking) Ensemble methods & model combination — four combiners on a forecastable target; BMA edges the best single model, and the gain is small by design (base models 0.97-correlated) — the diversity lesson that motivates specialist decorrelation
Consensus — hypothesis reconciliation (posterior over states) Bayesian hypothesis reconciliation — beats averaging & voting on a known 3-state truth (acc 0.78 vs 0.73 vs 0.70; H4 confirmed), with a disagreement measure that spikes at transitions
Adaptive expert weighting (prediction with expert advice / Hedge) Online expert weighting — Hedge matches the best of four rotating volatility experts (regret 11.5 vs bound 44.3) and beats equal weighting; fixed-share an honest H7 null
Strategy council — regime-conditioned suitability Strategy-suitability council — tests H5: a hand-mapped council loses to diversification (0.39 vs 0.58), a learned-mapping council beats it (0.66); suitability must be learned, not assumed
Risk governance VaR & Expected Shortfall, Kelly, drawdown, hard risk budget
Optional RL layer (MDP → PPO, constrained) MDPsQ-learning, policy gradients, PPO, constrained RL
Optional RL — POMDP belief filtering & contextual bandits Belief filter (Kalman & state-space) and the transparent RL comparator (contextual bandits) both built and verified
Validation & anti-overfitting toolkit Walk-forward, purged CV, leakage, PBO/DSR
Explainability & audit Permutation importance & PD, exact SHAP

The site’s empirical findings also de-risk the positioning. The capstone shows every model from linear regression to a transformer agreeing that next-day direction is unpredictable (AUC ≈ 0.50) while volatility and regimes are structured and forecastable — which is exactly why this proposal targets calibrated market-state intelligence and risk-governed decision quality, not return prediction. The research question survives the efficient-market evidence because it does not depend on beating it.

Expected contribution

The methodological contribution is a hierarchical separation of data, quantitative models, specialist agents, consensus reasoning, strategy suitability and risk authority; a multidimensional probabilistic state representation; a common evidence schema carrying calibrated probability, conditional reliability, horizon, freshness, support, contradiction and distribution-shift information; a bounded hypothesis-reconciliation mechanism; and a rigorous comparison of adaptive expert weighting against transparent baselines. The empirical contribution is evidence on which regime dimensions are detectable and useful out of sample, whether state information improves strategy selection after costs, when text and macro add value, and whether complexity pays mainly during transitions and disagreement. The practical/governance contribution is a real-time architecture that produces useful intelligence without granting an LLM trading authority, an auditable path from observation to outcome, and a framework for deterministic risk controls, no-action states and human oversight.

Commercialisation blueprint

The first commercial product should be a decision-support platform, not an autonomous hedge fund or signal-selling service: institutional users receive a live interpretation of market conditions, evidence and uncertainty. Potential customers include asset managers and hedge funds; bank treasury, trading and risk functions; family offices and prop firms; execution desks; wealth-management platforms; and risk, model-validation and market-surveillance teams. Defensible IP accrues in the point-in-time data engineering and feature lineage, the market-state taxonomy and transition definitions, calibration and conditional-reliability histories, the hypothesis-reconciliation logic, the regime-to-strategy suitability mappings, a historical-analogue and outcome library, and the decision/override/attribution audit trail.

The route to market is staged: (1) a research dashboard on delayed/historical data validating state inference and explanation; (2) a live market-state API and alert service for a narrow universe; (3) an institutional strategy-suitability and risk-overlay module integrated with client workflows; (4) client-specific allocation, execution or portfolio-control extensions subject to regulation and validation.

Commercial discipline. The PhD should create reusable data assets, methods and evaluation evidence. It should not depend on the assumption that one proprietary strategy will generate persistent excess returns — the business case is stronger when the platform improves decision context, model governance and strategy selection across many clients and strategies.

Because commercialisation is explicit, one item belongs on the critical path before enrolment: agreeing the intellectual-property position with the host university, since ownership of PhD outputs varies by institution and standard routes exist (IP carve-outs, industry-partnered PhDs).

Chapter structure and three-year plan

The thesis is planned in ten chapters: introduction; literature review; research design and data architecture; multidimensional market-state detection; calibrated specialist evidence; macro-event and narrative intelligence; multi-agent hypothesis reconciliation; regime-conditioned strategy selection; integrated evaluation and commercial prototype; and governance, conclusions and future research.

Year 1 — complete the literature review; finalise the taxonomy; acquire and govern data; build the point-in-time platform; establish baselines; develop trend, volatility, liquidity and transition models (the specialist prototypes already built on this site, re-established on the full equity cross-section); first paper. Year 2 — probability calibration and reliability; macro-event agent; ensemble versus consensus methods; a live research dashboard; second paper. Year 3 — regime-conditioned strategies and risk controls; adaptive weighting; ablations and robustness; institutional prototype and commercialisation chapter; submit thesis.

Supervision and fit

The proposal draws on two distinct bodies of expertise, and is explicit about where each guides it. Its empirical and econometric coreregime detection, volatility modelling (GARCH, EGARCH, HAR-RV and implied measures), forecasting, probability calibration and the deflated-Sharpe / PBO evaluation discipline — sits in quantitative finance, financial econometrics and forecasting, where a primary supervisor in those areas would most directly lead, sharpen and examine the work. Much of that core is already built and verified on this site, which should make supervision concrete rather than speculative.

The AI and agentic-systems layerLLM narrative extraction, multi-agent coordination and the control and governance stack — is a genuinely separate competency, and the proposal treats it as one: kept deliberately subordinate to the econometric core (the LLM is a measured information source, never an execution authority) and well-suited to co-supervision with machine-learning-systems or applied-AI expertise. Positioning the AI component as a co-supervised extension, rather than the centrepiece, keeps the thesis anchored in defensible empirical finance while still pursuing the AI direction that motivates it.

Scope, limitations and risks

Scope is controlled deliberately: begin with one liquid market ecosystem; use one-minute decision intervals initially, adding tick-level components only where a defined question requires them; treat the LLM as a structured event-information processor, not a trader; evaluate market-state inference and strategy suitability before trade optimisation; introduce RL only after transparent baselines exist; and prioritise calibration, robustness and attribution over headline returns. The principal risks each have a named mitigation already demonstrated on this site — unstable or subjective regimes (multiple operational definitions, sensitivity reporting), look-ahead and revision bias (available-at timestamps, frozen vintages), agent complexity without value (mandatory ablations), uncalibrated confidence (formal calibration), LLM hallucination (schema constraints, retrieval, validation), RL instability (delayed RL, restricted actions, hard limits), and backtest overfitting (walk-forward, PBO, Deflated Sharpe).

Recommended positioning. Frame the doctorate as research into probabilistic market-state inference and regime-conditioned decision intelligence. Treat trading performance as one downstream evaluation, not the sole proof of value.

Concluding proposition

The revised proposal is not principally a claim that a collection of agents can predict market direction. It is a study of whether heterogeneous evidence can be transformed into a calibrated, multidimensional and auditable understanding of the current market environment. Its central innovation is the separation of perception, evidence, hypothesis reconciliation, strategy suitability and risk control: quantitative models describe specific aspects of the market; specialist agents validate and standardise their outputs; a consensus layer compares competing explanations; strategy agents judge suitability; and deterministic controls retain final authority. This is academically more defensible than a broad autonomous-trading proposal and commercially more valuable than a single strategy.

Appendix A — key formal definitions

Concept Formulation or interpretation
Market-state vector z_t = [\text{trend}_t,\ \text{vol}_t,\ \text{liq}_t,\ \text{driver}_t,\ \text{transition}_t,\ \text{xasset}_t]
Specialist evidence e_{i,t} = \{\text{state},\ p,\ \hat p_{\text{cal}},\ \text{reliability},\ \text{horizon},\ \text{freshness},\ \text{support},\ \text{contradiction}\}
Calibration condition P(Y = 1 \mid \hat p = p) \approx p
Competing hypotheses H_k are alternative explanations; posteriors are updated from specialist evidence
Belief state b_t(s) = P(S_t = s \mid o_{1:t})
Transition probability P(S_{t+1} = j \mid S_t = i), or an online probability the current state is changing
Strategy suitability u_{j,t} = expected utility of strategy j conditional on the inferred state and risk constraints
Expert reliability w_{i,t} depends on calibration history, similar-state performance, freshness and distribution shift
No-action condition Action withheld where confidence, data quality, liquidity or distributional validity falls below thresholds

Appendix B — minimum viable research prototype

A first prototype is intentionally small: NQ futures, QQQ, NASDAQ breadth, a volatility index, 2y/10y yields and a dollar proxy; one-minute decisions; a Python / Parquet / TimescaleDB / MLflow / FastAPI stack; six initial specialists (trend, volatility, liquidity, structural change, macro-event, cross-asset); consensus baselines (equal weighting, majority vote, Bayesian model averaging, supervised stacking); five strategy families (trend, mean-reversion, long volatility, market-neutral, cash preservation); and a first user output of live market state, confidence, disagreement, evidence, strategy suitability and risk recommendation. The first research milestone is concrete: demonstrate that market-state information improves at least one of calibration, drawdown, tail risk or strategy selection out of sample.

Literature review

A full thematic literature review surveys the literatures the proposal integrates — market efficiency and the limits of predictability, regime detection and non-stationarity, volatility, ensemble learning and prediction with expert advice, NLP and LLMs in finance, multi-agent systems, and backtest methodology — read critically and positioned against the 2024–2026 wave of multi-agent LLM trading systems (TradingAgents, FinRobot) whose own critics document the gap this proposal fills: control-first, calibrated, regime-aware, risk-governed and honestly evaluated. Structured notes on individual papers follow.

Date Title Categories
Jul 28, 2026 Li, Kim, Cucuringu & Ma (2025) — Can LLM-based Financial Investing Strategies Outperform the Market in the Long Run? LLM agents, evaluation, positioning
Jul 22, 2026 Xiao, Sun, Luo & Wang (2024) — TradingAgents: Multi-Agents LLM Financial Trading Framework LLM agents, multi-agent, positioning
Jul 14, 2026 Bailey & López de Prado (2014) — The Deflated Sharpe Ratio methodology, backtesting, statistics
Jul 8, 2026 Moreira & Muir (2017) — Volatility-Managed Portfolios volatility, portfolio, tradability
Jul 2, 2026 Hamilton (1989) — A New Approach to the Analysis of Nonstationary Time Series and the Business Cycle regimes, time series, foundational
Jun 28, 2026 Ang, Hodrick, Xing & Zhang (2006) — The Cross-Section of Volatility and Expected Returns asset pricing, volatility, cross-section
No matching items

The full formatted proposal — the ten-chapter thesis structure, the complete evaluation framework and the appendix of formal definitions — is available as a PDF. Institution-specific requirements, the final literature base, data-access arrangements and formal ethics sections will be completed during proposal development with supervisors.