Skip to content
Saturday, August 29, 2026 · Global Edition
UZU News
MARKETS · INVESTING
Loading market quotes…
BTC · ETH · SOL · XRP · ADA · DOGE · AAPL · MSFT · NVDA · AMZN · GOOGL · TSLA
Market data by TradingView
Home / Analysis

What does it mean for a probability forecast to be calibrated?

Calibration is the match between stated probabilities and observed frequencies — measurable with Brier scores and reliability diagrams, and distinct from resolution.

Meteorologist style verification chart with diagonal reliability line
Probabilities are commitments to frequencies — counted after the fact.

A probability forecast is calibrated when stated probabilities match observed frequencies over many forecasts — a 40 percent event happens about 40 percent of the time among all statements of 40 percent — and the standard measurement is the Brier score, the mean squared difference between forecast probabilities and outcomes, with reliability diagrams binning forecasts by stated probability and plotting observed hit rates against them. UZU NEWS publishes information, not investment advice, and measures forecast quality without issuing forecasts.

Calibration is the least appreciated property of probabilistic prediction and the most checkable: unlike accuracy or skill, it requires no benchmark opinion, only a tally of stated probabilities against outcomes. A forecaster who says 40 percent repeatedly has made a claim that arithmetic can audit — and the audit, run on enough statements, is unforgiving in a way single predictions never are.

What is calibration, exactly?

The copula-form definition: calibration is the match between a model's stated probabilities and the observed frequencies of the events they describe, assessed across the distribution of statements. If every 30 percent forecast is followed by the event 30 percent of the time, every 70 percent forecast 70 percent of the time, and so on through the bins, the forecaster is calibrated. Calibration says nothing about whether the forecasts are useful: a forecaster who always says 50 percent for binary questions is perfectly calibrated and perfectly useless — the utility comes from resolution, the ability to sort events into bins that differ from the base rate. The two properties are orthogonal and both are measurable, which is why serious forecast evaluation reports them separately.

How is it measured?

Two standard instruments. The reliability diagram: bin forecasts by stated probability, plot the observed frequency against the stated probability for each bin, and read calibration off the diagonal — points on the diagonal are calibrated, points above indicate under-forecasting, below indicate over-forecasting; the hedgehog-with-45-degree-line caricatures make the diagram memorable. The Brier score: the mean of squared differences between probability and outcome indicator across forecasts — lower is better — and its decomposition, published by Allan Murphy in 1973, splits the score into reliability, resolution and uncertainty terms, formalizing the orthogonality: a score can improve by better calibration, better resolution, or both. Ranked probability scores extend the arithmetic to multi-category forecasts. The instruments require sample: single forecasts cannot be calibrated or miscalibrated, only sequences can, and bin-level assessments need enough statements per bin to make frequencies meaningful.

What do known applications show?

The classic documented comparisons: meteorological precipitation probabilities — where probabilistic forecasting matured earliest and calibration is maintained as an operational property, with published reliability curves near the diagonal for national weather services — and judgmental forecasting tournaments, where multi-year geopolitical forecasting competitions found that aggregation and training improved both calibration and resolution, with the results published by the research programs that ran them. In finance, calibration appears in the evaluation of probability outputs from models — default probabilities, event probabilities, direction probabilities — where the documented regularity is that miscalibration is common and direction varies: retail-facing optimism in some domains, excessive hedging in others. The election-forecasting record provides public case studies in both directions across recent cycles.

PropertyDefinitionMeasured byCan exist alone?
CalibrationStated probability matches observed frequencyReliability diagram; Brier reliability termYes — the always-50 forecaster
ResolutionBins differ from base rateBrier resolution termYes — sharply wrong bins
SkillBeats a stated reference forecastSkill scores vs baselineRequires a baseline

Where does calibration talk mislead?

Four recurring slips. Single-event calibration language: "that 30 percent call was well calibrated" is category error — calibration is a property of distributions of forecasts, not of individual ones; a 30 percent event occurring or not occurring is consistent with any calibration. Cherry-binned reliability: diagrams computed on few forecasts per bin are noise dressed as miscalibration; state the bin counts. Base-rate laundering: a forecaster calibrated on easy questions is not thereby calibrated on hard ones — calibration is domain- and regime-specific, and pooled diagrams hide the difference. And the calibration-equals-usefulness error: the always-50 forecaster returns calibrated and useless, so any evaluation quoting calibration alone has reported half the measurement. The Murphy decomposition exists precisely to keep the halves separate.

How should readers evaluate a probability source?

Ask for the track record in a form that admits arithmetic: the list of stated probabilities with outcomes, long enough to bin. Sources that publish their full forecast history — as the major tournament programs and some public forecast platforms do — can be audited by anyone; sources that quote only hits or only the latest surprise cannot. The National Weather Service documentation of probabilistic forecasting and the academic verification literature provide the reference standards, and the underlying decomposition mathematics traces to Murphy's 1973 paper. For a site that covers forecasting models, the editorial rule this publication follows is the one it recommends to readers: probabilities are commitments to distributions of outcomes, and the only respect they can be paid is in frequencies, counted after the fact.

Sofia Lindqvist

Sofia Lindqvist builds models for a living and is unusually honest about how often they are wrong.

More about Sofia Lindqvist

Frequently Asked Questions

What does it mean for forecasts to be calibrated?
That stated probabilities match observed frequencies across many forecasts: 40 percent events occur about 40 percent of the time among all 40 percent statements. Calibration is a property of distributions of forecasts, never of a single one.
What is the Brier score?
The mean squared difference between forecast probabilities and outcomes, lower is better. Murphy's 1973 decomposition splits it into reliability (calibration), resolution (usefulness beyond base rate) and uncertainty — keeping the two virtues separately measurable.
Can a forecaster be calibrated but useless?
Yes — always saying 50 percent is perfectly calibrated and carries no information. Usefulness requires resolution: sorting events into bins whose frequencies differ from the base rate. Serious evaluation reports both.
How do you check a probability source's calibration?
Get the full history of stated probabilities with outcomes, long enough to bin, and tally observed frequency per bin against stated probability. Sources publishing complete track records can be audited; sources quoting only hits cannot.