A high Sharpe ratio from a backtest is weak evidence, because it reports the outcome of a search rather than the performance of a model. Bailey and López de Prado's Deflated Sharpe Ratio, published in the Journal of Portfolio Management in 2014, shows that a ratio of 1.0 earned as the best of 100 trials on three years of data can deflate to near zero once the trial count is accounted for. The number is not wrong; it is answering a different question than the reader is asking.
Uzu News publishes information and analysis about forecasting methods, not investment advice, and this analysis makes no claim about any live strategy's future performance. Its subject is what a reported backtest can and cannot support as evidence.
Why do backtests overstate live performance?
Two mechanisms do most of the damage. The first is multiple testing: when a researcher tries many configurations on one history and publishes the winner, the winner's margin is partly luck, and the luck does not repeat. Halbert White's Reality Check, published in Econometrica in 2000, formalized the correction for this in econometrics.
The second is selection under scarcity. Financial history offers one realized path per asset. A configuration that fit the particular sequence of the 2010s is fitting noise that will not be regenerated, and the evaluator usually cannot tell the fit from the mechanism.
The 2014 Journal of Computational Finance paper by Bailey, Borwein, López de Prado and Zhu measured the combined effect: the probability of backtest overfitting, meaning the chance that the best in-sample configuration lands in the bottom half out-of-sample, rises toward and past 50 percent as the number of configurations tried on a fixed history grows. Overfit backtests are the expected output of an ordinary research process, not a rare failure of it.
What would a valid backtest report?
The answer is a set of conditions, not a single metric. Each one is checkable in principle, which is what makes the standard usable.
| Evidence item | Why it is required | Common omission |
|---|---|---|
| Full trial count | Every correction formula needs the number of configurations tried | Only the winner is published |
| Baseline comparison | A metric without a benchmark has no reference point | Baseline absent or trivial |
| Evaluation window and universe | Performance claims hold only where measured | Window chosen post hoc |
| Out-of-sample protocol | Separation of scoring data from fitting decisions | Test set used during tuning |
| Verification status | Claimed performance differs from independently verified performance | Distinction not stated |
The last row is the one readers most often have to supply themselves. A vendor's stated Sharpe ratio is a vendor claim, carried with its stated conditions; it becomes evidence only when an independent party recomputes it from data the vendor did not select.
How do regulators read the same evidence?
Banking supervisors reached a similar position a decade before the machine-learning literature made it explicit. The Federal Reserve's Supervisory Guidance on Model Risk Management, SR 11-7 of 2011, treats a model's developer as an interested party and requires independent validation: outcome analysis against data the model did not influence, benchmarking against alternatives, and documentation of known weaknesses and their compensating controls.
The parallel is exact enough to be worth stating. The literature's contribution was to quantify how much selection inflates a reported ratio; the guidance's contribution was to make independent checking a structural requirement rather than a courtesy. A backtest that would fail the guidance's independence test should not impress a reader simply because it appears in a paper instead of a bank memo.
What does this leave the reader to do?
Read the conditions before the number. A Sharpe ratio stated with its trial count, sample length, variance across trials, universe, and window is a report; the same number stated alone is a claim. When the conditions are missing, the deflation arithmetic cannot be run, and the honest default is to treat the figure as unverified.
This publication's own judgment, stated once: the persistent problem is not dishonesty but asymmetry of effort. Running a hundred backtests is cheap and quick, and verifying one is slow and expensive, so markets accumulate unverified winners faster than anyone retires them. That asymmetry, more than any single flawed paper, is why claimed backtest performance should be read as a hypothesis awaiting independent testing.
What remains unknown after this analysis is the base rate: no public dataset records how many published backtests have been independently replicated, so the share that survives verification cannot be stated from evidence. The corrections define the inflation. The audit counts do not yet exist at scale.
For more context, read How many signals were tested before this one worked?.

