Regime-detection models — hidden Markov models in their classic Hamilton 1989 form, and clustering approaches in the same family — summarize historical data into a small number of states with different statistical properties, and their documented strength and weakness are the same fact: states are inferred from the data that defined them, so the label of a regime is most reliable after the regime has run, with real-time identification lagging and state counts and boundaries depending on specification choices. UZU NEWS publishes information, not investment advice, and prices regime tools at their measured value.
"The market has entered a new regime" is a sentence used both as analysis and as soothsaying. The quantitative version is checkable: a model, a state definition, an inferred label and a timestamp — and the checkable version is considerably more modest about the present than the conversational one.
What are regime models, mechanically?
A hidden Markov model posits an unobserved state variable following a Markov chain, with observed returns or volatilities drawn from state-dependent distributions — for markets, typically a low-volatility and a high-volatility state, sometimes a third intermediate state, with transition probabilities estimated from history. Hamilton's 1989 application to U.S. output founded the macro literature; market applications followed through the 1990s. Clustering variants group periods by statistical similarity without explicit transition dynamics. The output is a state sequence: a label for every historical period and, at each moment, filtered probabilities of being in each state given data to date — the model's live guess, always conditional on the fitted parameters.
What are they genuinely good at?
Three uses, evidenced. Description: the two-state volatility characterization of equity markets is one of the most replicated statistical regularities in finance — volatility clusters, and HMMs capture the clustering compactly, which is why they appear throughout the academic risk literature. Smoothed historical analysis: with the full sample available, smoothed state probabilities give a clean retrospective segmentation — crisis periods light up, calm stretches group — useful for conditioning other research, such as evaluating how strategy performance varied across inferred states. Volatility forecasting: Markov-switching variance models rank among the competitive volatility forecasters in the comparisons literature, largely by admitting that today's variance regime predicts tomorrow's — persistence the single-state models force into their parameters.
Where do they fail?
Four documented failure modes — the mandatory limitations section. Lag: filtered — real-time — probabilities recognize a new regime only after enough data has accumulated in it; by construction the model needs observations of the new state to believe in it, so the sharpest regime shifts are identified latest, the exact periods when identification matters most. Specification dependence: the number of states, the distributions allowed, the inclusion of returns or volatility or both — each choice moves the boundaries, and different specifications of the same market disagree at the margins, which is why published regime counts vary from two to five with equal conviction. Non-stationarity: models fitted on one era inherit its regime structure; when the structure itself changes — new policy regimes, structural breaks the training set never saw — the state vocabulary no longer describes the world. In-sample seduction: regime-conditional performance looks impressive when states were fitted on the same data that performance is measured on — the honest regime analysis evaluates out-of-sample, which few marketing versions do.
How did 2020 and 2022 stress the tools?
Both episodes are documented case studies. March 2020: volatility moved to levels outside the training distribution of essentially every fitted model; state probabilities collapsed into the high-volatility state almost immediately — the one case where identification lag was short — but the state's parameters were being extrapolated beyond anything fitted, and transition probabilities estimated from tame years said nothing useful about how long it would last. 2022: an inflation-led drawdown with rising rates behaved statistically differently from demand-shock crises, and models calibrated on the 2000s-2010s disinflation era treated the episode with the same two states it always had — the mapping was available, and the vocabulary was stale. Both cases illustrate the same limit: regime models compress history, and history that has not happened yet is not in the compression.
| Use | Standing | Key condition |
|---|---|---|
| Descriptive state segmentation | Well-evidenced | Smoothed, retrospective |
| Volatility forecasting | Competitive | Refit discipline stated |
| Real-time regime labels | Lagging, noisy | Filtered probabilities only |
| Strategy conditioning | Mixed | Must be evaluated out-of-sample |
How should regime claims be evaluated?
Ask four questions. Specification: how many states, on what variables, fitted over what window — and how sensitive are the boundaries to those choices. Timing: is the quoted label filtered at the time or smoothed after — retrospective labels presented as timely calls are the genre's signature move. Evaluation: was regime-conditional performance measured out-of-sample, or on the same data that defined the regimes. And purpose: description and volatility forecasting have evidential support; regime-timed trading claims inherit every problem this site has catalogued for strategies, plus the identification lag. The primary literature is accessible — Hamilton 1989 and the Markov-switching literature that followed — and the supervisory conversation about model limitations in production systems is documented in the Federal Reserve's SR 11-7 guidance at federalreserve.gov. Regime models are honest instruments with a tense problem: they describe the past fluently and are taught to speak of the present only in probabilities.
For more context, read What does out-of-sample mean in a forecasting paper?.
For more context, read alternative data evaluation.
For more context, read forecast calibration.




