MethodsUpdated 888 words, 4 minutes
Backtest overfitting: the deflated Sharpe ratio and a simulation
Why the best of many backtests looks good even when every strategy is noise: a simulation you can run, plus the deflated Sharpe ratio applied to it.
If you try enough strategies on the same data, one of them will have a good backtest. That sentence is obvious and it is also the most common way a research process produces a confident wrong answer, because the person who tried a thousand variants and reports the best one rarely reports the thousand. This article shows the effect with a simulation where every strategy is pure noise, then applies the correction Bailey and Lopez de Prado proposed in 2014, the deflated Sharpe ratio. The script is attached and uses only NumPy.
The simulation
Each "strategy" is a series of 1,260 daily returns (five trading years) drawn from a normal distribution with mean exactly zero and an annual volatility of 16 percent. There is no signal anywhere. For a given number of trials N, the script draws N such strategies, computes each one's in-sample Sharpe ratio, and keeps the best. It repeats that 200 times per N and averages.
rets = rng.normal(0.0, DAILY_VOL, size=(n, T)) # n strategies x T days, true mean 0
srs = rets.mean(axis=1) / rets.std(axis=1, ddof=1) # daily Sharpe of each
k = int(np.argmax(srs)) # the one you would have reported
best_annualised = srs[k] * math.sqrt(252)
Output of python code/backtest-overfitting-deflated-sharpe.py (seed 20260905):
| strategies tried N | mean of best in-sample Sharpe (annualised) | expected max under the null (formula) | share of runs where best Sharpe > 1.0 | DSR of the best strategy (mean) |
|---|---|---|---|---|
| 1 | 0.01 | 0.00 | 2% | 0.50 |
| 10 | 0.69 | 0.61 | 16% | 0.51 |
| 100 | 1.11 | 1.12 | 67% | 0.48 |
| 1000 | 1.45 | 1.47 | 100% | 0.49 |
| 10000 | 1.73 | 1.73 | 100% | 0.50 |
With one strategy and no selection, the average Sharpe is zero, as it should be. With a hundred candidates, the best one has an annualised Sharpe above 1.0 in two runs out of three. With a thousand, it always does, and the average "winner" reports 1.45. Nothing was learned about markets in any of these runs; the number is the price of looking.
The expected maximum
The third column is not from the simulation. It is the closed-form expectation the paper gives for the maximum of N Sharpe estimates when the true Sharpe of every candidate is zero:
SR0 = sqrt(V) × [ (1 − γ) Z⁻¹(1 − 1/N) + γ Z⁻¹(1 − 1/(N·e)) ]
where V is the variance of the Sharpe estimates across the trials, Z⁻¹ is the inverse of the standard normal CDF, and γ is the Euler-Mascheroni constant, about 0.5772. The formula and the simulation agree to within a few hundredths at every N above 10, which is a check that the script implements what the paper describes. The formula grows with the logarithm of N, which is why going from 1,000 trials to 10,000 buys only 0.26 of extra apparent Sharpe: the damage is done early.
The deflated Sharpe ratio
The paper's proposal is to test the reported Sharpe not against zero but against SR0, and to account for the non-normality of the returns while doing so. As we read the paper, the statistic is
DSR = Φ[ (SR̂ − SR0) × sqrt(T − 1) / sqrt(1 − γ₃ SR̂ + ((γ₄ − 1)/4) SR̂²) ]
with SR̂ the non-annualised (here daily) Sharpe estimate, T the number of returns, γ₃ the skewness and γ₄ the kurtosis of the returns, and Φ the standard normal CDF. It reads as a probability that the true Sharpe exceeds zero after the selection is accounted for.
def dsr(sr_hat, sr0, t, skew, kurt):
z = (sr_hat - sr0) * math.sqrt(t - 1) / math.sqrt(1 - skew * sr_hat + (kurt - 1) / 4 * sr_hat ** 2)
return norm_cdf(z)
The last column of the table is the mean DSR of the winning strategy, and it is about 0.5 at every N. That is the correct answer: the winner of a noise tournament is, on average, exactly as good as the expected winner of a noise tournament, so the evidence that it is better than nothing is a coin flip. The worked example the script prints for N = 1,000 makes the contrast explicit. The best strategy had a daily Sharpe of 0.0918 (1.46 annualised). Tested naively against zero, ignoring the 999 discarded candidates, its p-value would be 0.0006, a result most people would call highly significant. Its DSR is 0.462. The same Sharpe with N = 1, meaning it was the only strategy ever tried, would carry a DSR of 0.999.
What the correction needs, and what it cannot do
The formula needs N, the number of independent trials, and that is the number nobody records. Variants of a strategy that differ by one parameter are not independent trials, so N overstates; but a research team's collective history of discarded ideas is usually far larger than anyone's notebook shows, so N understates too. The paper discusses estimating an effective number of trials from the correlation of the candidates' returns, and that step is not implemented here. The V in the formula was estimated from the trials' own Sharpe estimates, which is available in a simulation and rarely in practice.
The companion paper in the Notices of the AMS makes the broader point with a different construction: the authors show that with enough parameter combinations one can always find a strategy that fits a given sample, and that the in-sample fit carries no information about the out-of-sample result. The deflated Sharpe ratio is a repair for one measurement; it does not repair a process that keeps searching until something passes.
None of this is a recommendation about any strategy, and the simulated strategies here are not tradable objects. The article shows one arithmetic fact, that the best of N noise draws has an expected Sharpe that grows with log N, and one published correction for it, with the code to check both.