Methodology · Deep dive
Walk-forward analysis, explained
Most backtests answer the question "how would this rule have done on the past I tuned it on?" — which is no question at all. Walk-forward analysis asks the more demanding question: how does the rule behave on data it has never seen, again and again, as time rolls forward? This page explains the method and its limitations. Numerical performance claims are withheld pending a complete reproducibility file and evidence review.
The method in four steps
- Split history into folds. Divide the sample into consecutive chronological windows. The current dataset size and fold count are withheld while the reproducibility file and evidence review remain incomplete.
- Fit on the past, judge on the future. Within each fold, parameters are chosen using only the training span, then performance is measured only on the unseen test span that follows it.
- Roll forward and repeat. Each fold's test span becomes later folds' history. No information travels backwards.
- Grade on retention and consistency. How much of the in-sample performance survives out-of-sample (the retention ratio), and does it survive in every fold — one negative fold means the average is luck-dependent.
The pitfall almost everyone hits: annualisation
Inside a fold, the test window is shorter than the training window. Comparing raw cumulative returns across windows of different lengths can make the shorter window appear weaker because of window arithmetic. Both legs must be annualised consistently before comparison.
Why the current assessment remains WEAK
A single train/test split can land on a friendly window and create false confidence. Rolling across multiple chronological windows is intended to make that instability visible.
The current internal assessment of the twelve-measurement composite is WEAK and not robust. It is not presented as a verified performance claim. Historical, backtested and walk-forward results do not guarantee future results. Full context is on the main methodology page.
Walk-forward vs. the alternatives
| Approach | What it catches | What it misses |
|---|---|---|
| In-sample backtest | Coding errors, roughly nothing else | Everything — the rule has seen the answers |
| Single train/test split | Gross overfitting | Window luck — one friendly test span flatters, as our RSI example shows |
| Walk-forward (n folds) | Window luck, regime dependence, parameter fragility | Regimes absent from the whole sample; structural breaks still to come |
Note the last cell: walk-forward is more demanding than the other approaches shown and still is not proof. Five years of history contains only the regimes it contains. That limit is why no Farlens output is a recommendation — the method quantifies the past's consistency, not the future's behaviour.
Checklist for reading anyone's backtest (including ours)
- Was any parameter chosen after seeing the test data? (If unstated, assume yes.)
- Is performance reported per-fold, or only as an average that can hide a losing fold?
- Are train and test legs annualised before comparison?
- Are transaction costs and missing-data handling stated?
- Is the sample size (observations, not years) disclosed?
Related reading
- The Farlens composite methodology
- Sharpe ratio calculator — general risk arithmetic with its limits stated