Skip to content
Tuesday, August 25, 2026 · Global Edition
UZU News
MARKETS · INVESTING
Loading market quotes…
BTC · ETH · SOL · XRP · ADA · DOGE · AAPL · MSFT · NVDA · AMZN · GOOGL · TSLA
Market data by TradingView
Home / Investing

How many signals were tested before this one worked?

Multiple testing turns luck into apparent skill. What the correction procedures control, the hurdles the replication literature settled on, and what a corrected t-statistic still cannot tell you.

How many signals were tested before this one worked?

The significance threshold that most signal research reports was designed for a single test, and signal research is never a single test. Harvey, Liu and Zhu, writing in the Review of Financial Studies in 2016, documented 316 factors in the published literature and argued that a new one should clear a t-statistic above 3.0 rather than the conventional 2.0.

What is the multiple-testing problem in signal research?

Multiple testing is the inflation of false positives that occurs when many candidate signals are evaluated against the same data and only the survivors are reported. A single test at the 5% level accepts a one-in-twenty chance of a false discovery. Run the procedure across hundreds of candidates and false discoveries stop being an accident and become an expected yield.

The count in the academic record is the part that can be documented. Harvey, Liu and Zhu assembled 313 articles — 250 published papers and 63 working papers — covering 316 distinct factors, with empirical tests running from 1967 forward. That is the visible denominator. The authors' conclusion from it was blunt: they argue that most claimed research findings in financial economics are likely false.

Their remedy was arithmetic rather than rhetorical. If the literature is a single large family of tests, the significance hurdle for admission to it has to rise with the size of the family. The hurdle they recommend for a newly proposed factor is a t-statistic greater than 3.0.

How is the correction actually measured?

The corrections come from statistics, not from finance, and they differ in what they promise. Family-wise error rate, or FWER, is the probability of making at least one false discovery across the whole set of tests. False discovery rate, or FDR, is the expected proportion of the discoveries you report that are false. Controlling the first is stricter; controlling the second is more permissive by design.

Harvey, Liu and Zhu apply three established procedures to the factor literature: Bonferroni, Holm, and Benjamini, Hochberg and Yekutieli, usually shortened to BHY. They also build a model that accounts for correlation among the test statistics, because the standard procedures assume a level of independence that a library of overlapping accounting ratios does not have.

ProcedureWhat it controlsCharacter of the adjustment
BonferroniFamily-wise error rateSingle-step; the strictest of the three
HolmFamily-wise error rateSequential; uniformly less conservative than Bonferroni
BHYFalse discovery rateSequential; tolerates a controlled share of false positives

The practical output of all three is the same object: a higher bar. A t-statistic of 2.1 that would be reported as significant in isolation is, inside a family of several hundred tests, an unremarkable draw. Nothing about the signal changes. What changes is the number of competitors it was silently measured against.

What happened when the anomaly literature was re-tested?

The correction stopped being theoretical when someone rebuilt the anomalies from scratch. Hou, Xue and Zhang compiled 452 anomaly variables and replicated them over a January 1967 to December 2016 sample, published in the Review of Financial Studies in 2020. Their replication used NYSE breakpoints and value-weighted returns rather than the more permissive conventions common in the original papers.

The headline number is a failure rate, not a performance number. In the authors' own words, “65% of the 452 anomalies in our extensive data library, including 96% of the trading frictions category, cannot clear the single test hurdle of the absolute t-value of 1.96.” That is before any multiple-testing adjustment is applied at all.

Applying one makes it worse. The same abstract continues: “Imposing the higher multiple test hurdle of 2.78 at the 5% significance level raises the failure rate to 82%.” Read carefully, that is two separate findings stacked — methodological choices alone kill most of the sample, and the multiplicity correction kills most of what survives. The published replication reports both hurdles side by side for exactly that reason.

Neither paper claims the survivors are tradeable, and neither is a statement about what any signal will do next. They are statements about how much of a published record survives contact with a stricter test.

How do supervisors treat the same problem?

Bank supervisors arrived at the multiplicity problem from the governance side rather than the statistical one. The Federal Reserve's SR 11-7, issued April 4, 2011, defines model risk as “the potential for adverse consequences from decisions based on incorrect or misused model outputs and reports,” and locates it in two places: models with fundamental errors, and models used incorrectly or inappropriately.

Its definition of backtesting is the operative one for signal work. The guidance defines backtesting as “the comparison of actual outcomes with model forecasts during a sample time period not used in model development.” The clause after the comma is the whole test. A backtest run on the data that selected the signal is not out-of-sample, whatever the code returns.

The guidance also requires that validation carry “a degree of independence from model development and use,” and treats outcomes analysis — comparing model outputs to corresponding actual outcomes — as a core validation element rather than an optional one. The supervisory letter is addressed to banking organizations, not to researchers, but the structural point transfers: the party that built the signal is the worst-placed party to certify it.

On the disclosure side, the Securities and Exchange Commission's adviser marketing rule, Rule 206(4)-1, restricts hypothetical performance — the category that covers backtested results in advertisements. Under the agency's compliance guide, an adviser must adopt policies ensuring the hypothetical performance is “relevant to the likely financial situation and investment objectives of the intended audience” and must provide certain information underlying it. The compliance date was November 4, 2022. This is a description of a rule, not legal advice.

Where does the correction fail?

The first failure is the denominator. Every procedure described above needs to know how many tests were run, and the true count is unobservable. Harvey, Liu and Zhu could count 316 published factors; nobody can count the specifications that were tried and abandoned, which is where most of the multiplicity actually lives.

The second is correlation. Bonferroni, Holm and BHY were built for tests that are independent or nearly so. Financial signals constructed from overlapping accounting data are heavily correlated, which makes the naive corrections misstate the true hurdle — the reason Harvey, Liu and Zhu develop a separate framework that models the correlation structure instead of assuming it away.

The third is that a corrected t-statistic is still a statement about a historical sample. It says a pattern is unlikely to be a fluke of that window. It does not say the pattern has an economic cause, that the cause persists, or that the effect survives trading costs. Hou, Xue and Zhang's finding that 96% of the trading-frictions category fails even the uncorrected 1.96 hurdle is a reminder of how much apparent structure lives in the implementation details.

What these methods cannot establish

A multiple-testing correction cannot rescue a signal from a badly specified test. If the sample, the breakpoints, or the weighting scheme were chosen after seeing the result, the corrected statistic inherits the same contamination at a higher threshold. The Hou, Xue and Zhang exercise is instructive precisely because the convention change and the multiplicity adjustment are reported separately.

These procedures also cannot verify claims made outside the published record. A vendor result presented without its evaluation window, sample construction, and the number of variants tested cannot be corrected, because the inputs the correction requires are missing. A claim of that shape is a claim, not a measured result, and remains one until its conditions are stated.

None of the sources here measured live, post-publication returns of the signals in question, and none of them make a statement about future performance. Neither does this article. Uzu News covers forecasting methods; it does not forecast, and nothing above is a recommendation to act on any model output. The papers cited report what happened to a defined set of variables over a stated historical window, on stated conventions — and that, stated exactly, is the entire claim.

Karim Al-Rashid

Independent editorial contributor focused on market analysis, corporate earnings, property markets, economic indicators.

Karim Al-Rashid tracks company results and property markets, with a habit of looking past the loudest number in the room.

More about Karim Al-Rashid

Sources

  1. Factor count, hurdle of t>3.0, multiple-testing procedures (Bonferroni, Holm, BHY), correlation framework, 'most claimed research findings are likely false'Harvey, Liu and Zhu, '. . . and the Cross-Section of Expected Returns', Review of Financial Studies 29(1), 2016 (author copy, Duke University)
  2. 452 anomalies, 1967-2016 sample, NYSE breakpoints and value-weighted returns, 65% failure at |t|=1.96, 96% of trading frictions, 82% failure at t=2.78Hou, Xue and Zhang, 'Replicating Anomalies', Review of Financial Studies 33(5), 2020
  3. Definition of model risk, definition of backtesting, independence of validation, outcomes analysis; letter date April 4, 2011Board of Governors of the Federal Reserve System, SR 11-7: Guidance on Model Risk Management
  4. Rule 206(4)-1 restrictions on hypothetical performance, relevance and underlying-information conditions, November 4, 2022 compliance dateU.S. Securities and Exchange Commission, Investment Adviser Marketing compliance guide