Strategy research
Backtest Overfitting: Why Good Backtests Fail Live
Backtest overfitting explained: curve fitting, multiple testing, degrees of freedom, out-of-sample checks and cost assumptions, with the research behind them.
Strategy research
Backtest overfitting explained: curve fitting, multiple testing, degrees of freedom, out-of-sample checks and cost assumptions, with the research behind them.
A backtest shows how a set of rules would have performed on past prices. Overfitting happens when those rules have been shaped, deliberately or not, to match the noise in that particular history rather than any repeatable pattern. The result is a backtest that looks excellent and a live account that does not. The researchers cited below suspect it is a large part of why many systematic funds fail to live up to the expectations their managers create, and it is hard to detect from the backtest alone.
Curve fitting is the obvious form: adjusting parameters until the equity curve is smooth. A moving-average crossover tested at every length from 5 to 200, with a stop tested at every distance from 10 to 100 pips, will produce some combination with a beautiful history. That combination was selected because it fit the past, not because it captures anything about the future.
The less obvious form happens without any optimizer. Each time you look at a result, decide it is not good enough and change a rule, you have run another test. A trading idea refined over months of trial and error can be heavily overfitted even if the final version has only two parameters.
The core problem is statistical. If you test enough variations of a strategy with no real edge, some of them will look good by chance. The best of many random results is not a random result; it is the extreme of a distribution.
Bailey, Borwein, López de Prado and Zhu quantified this for trading strategies in "Pseudo-Mathematics and Financial Charlatanism", published in the Notices of the American Mathematical Society in 2014. They derived a minimum backtest length needed to avoid selecting a skill-less strategy given the number of independent configurations tried. One of their illustrations: with only five years of data, trying more than about 45 independent configurations makes it close to certain that the best one will show an annualized in-sample Sharpe ratio of 1 even when its expected out-of-sample Sharpe ratio is zero. With two years of data, seven configurations are enough.
Their conclusion is that a backtest which does not report how many trials were run cannot be assessed for overfitting at all. They also point out that a simple hold-out sample does not fix the problem, because it does not account for how many models were tried before one was selected.
The same issue appears across finance research. Harvey, Liu and Zhu argued in The Review of Financial Studies (2016) that, given the hundreds of factors already tested, a newly proposed return factor should clear a t-statistic above 3.0 rather than the conventional 2.0.
A natural assumption is that an overfit strategy will simply perform at zero out-of-sample. Bailey and co-authors show that this holds only when performance has no memory. When the process has memory, for example when performance is serially correlated or tends to revert toward a long-run level, they find that optimizing in-sample can produce negative expected performance out-of-sample, with a negative relationship between in-sample and out-of-sample results. The strategy that looked best can be selected precisely because it caught a temporary move that later reverses.
Every parameter, filter, time-of-day rule and special case adds a degree of freedom. The search space grows multiplicatively. Five parameters with ten possible values each give 10^5 = 100,000 combinations. Even if you only inspect a few hundred, the opportunity to find a spurious fit is large.
Some practical habits help:
The same research group proposed a method to estimate the probability of backtest overfitting directly, published in the Journal of Computational Finance in 2017 (DOI 10.21314/JCF.2016.322). Their combinatorially symmetric cross-validation splits the history into many blocks, forms every combination of half the blocks as in-sample and the rest as out-of-sample, picks the best configuration in-sample each time and checks where it ranks out-of-sample. The probability of overfitting is the share of combinations in which the in-sample winner ranks below the median out-of-sample. A high figure means the selection process is mostly picking noise.
Out-of-sample data is the first defense, and walk-forward analysis extends it into a repeated, structured process. Both only help if the out-of-sample period is truly untouched. Once you look at out-of-sample results and go back to adjust the rules, that period has become in-sample.
Paper trading or live trading at minimal size adds two things a backtest cannot: genuinely unseen data and real execution. It tests fills, slippage, platform behavior and your own ability to follow the rules. Its weakness is speed, since collecting enough trades to judge an edge can take months.
Many overfit strategies are also under-costed. A strategy that averages 0.8 pips of gross profit per trade over 2,000 trades shows 1,600 pips in the backtest. Add 1.0 pip of real round-trip cost and the same trades lose 400 pips. High-frequency, short-target strategies are the most sensitive.
Model spreads that vary by session and widen around news and the daily rollover, include commission and swap, and add slippage on stops. In MT5, testing on generated ticks applies a fixed spread within each minute bar, which hides spread spikes; the MT5 report guide covers the settings to check.
None of these steps guarantees a strategy will work. They reduce the chance of trading one that never had an edge.