Strategy research
Walk-Forward Analysis: Testing Strategies Out of Sample
Walk-forward analysis explained: in-sample and out-of-sample windows, rolling vs anchored designs, walk-forward efficiency and the common pitfalls.
Strategy research
Walk-forward analysis explained: in-sample and out-of-sample windows, rolling vs anchored designs, walk-forward efficiency and the common pitfalls.
Walk-forward analysis tests a strategy the way it would actually be run: optimize on past data, trade the next stretch with the chosen settings, then repeat. Instead of one backtest over the whole history, you get a chain of out-of-sample results, each produced by parameters that never saw the data they were tested on. It does not prove a strategy works. It is one of the better ways to find out that it does not.
In-sample data is the period used to design or optimize a strategy. Any result measured on it is contaminated by the choices made while looking at it: indicator settings, filters, stop distances, even which strategy idea survived to be tested.
Out-of-sample data is held back and used only to evaluate those choices. A strategy that performs well in-sample and collapses out-of-sample has most likely been fitted to noise.
A single split, such as optimizing on 2016 to 2022 and testing on 2023 to 2025, is a start. MT5's built-in forward test does exactly this, reserving a half, third or quarter of the period, or a custom date, for the forward segment. The weakness of a single split is that it gives one out-of-sample result from one market regime. Walk-forward analysis repeats the split many times across the history.
The stitched out-of-sample record is the result. The in-sample results are only the means of choosing parameters.
The two standard designs differ in what happens to the start of the in-sample window. Take ten years of data from January 2016 to December 2025, a three-year in-sample window, one-year out-of-sample windows and a one-year step:
| Out-of-sample year | Rolling in-sample window | Anchored in-sample window |
|---|---|---|
| 2019 | 2016 to 2018 | 2016 to 2018 |
| 2020 | 2017 to 2019 | 2016 to 2019 |
| 2021 | 2018 to 2020 | 2016 to 2020 |
| 2022 | 2019 to 2021 | 2016 to 2021 |
| 2023 | 2020 to 2022 | 2016 to 2022 |
| 2024 | 2021 to 2023 | 2016 to 2023 |
| 2025 | 2022 to 2024 | 2016 to 2024 |
Both designs produce seven out-of-sample years.
Rolling windows keep a fixed length and drop the oldest data. They adapt faster to changing conditions, at the cost of optimizing on less history each time.
Anchored windows keep the original start date and grow. Parameters tend to be more stable because each optimization sees more data, but the strategy reacts slowly if the market's behavior has shifted.
Neither is correct in general. If a strategy only survives under one design, that sensitivity is itself useful information.
Robert Pardo, whose book The Evaluation and Optimization of Trading Strategies (2nd edition, Wiley, 2008) popularized the method among traders, defines walk-forward efficiency as the annualized out-of-sample return divided by the annualized in-sample return.
If the in-sample optimizations returned 20% a year on average and the out-of-sample segments returned 9% a year, walk-forward efficiency is 9 ÷ 20 = 45%.
Annualizing matters because the in-sample and out-of-sample windows differ in length. Some decline from in-sample to out-of-sample is normal, because optimization picks the parameters that happened to fit the past best. A very low or negative efficiency says the optimization is mostly fitting noise. An efficiency near or above 100% is worth checking for errors before celebrating.
Look at the windows individually as well as on average. A strategy where one out-of-sample year produces all the profit and the rest lose money has a fragile record whatever the aggregate efficiency says.
Treat the joined out-of-sample trades as the strategy's real track record and run the same checks you would on live results: expectancy in R, profit factor, maximum drawdown, longest losing streak and trade count. The trade log analyzer does this from a CSV of the out-of-sample trades and flags small samples. The guide to expectancy and profit factor explains how wide the error bands are at different trade counts.
Also look at how the chosen parameters changed between windows. If the optimal moving-average length jumps from 15 to 80 to 30, the strategy has no stable optimum, and the out-of-sample results probably reflect luck in parameter selection.
The last point connects walk-forward analysis to the wider problem of multiple testing, which the guide to backtest overfitting covers in detail. Walk-forward results reduce the risk of fooling yourself; they do not remove it, and a strategy that passes can still lose money live.