The Horizon You Fix First
A live record is the one test you cannot search, so its bar falls from lesson 64’s 4.499 to 1.645 and it needs 7.48 times fewer trades than a 15,000-cell backtest to support the same claim. That is the whole case for forward testing, and it holds only if you fix the number of trades before you start. Judge a 156-trade record once, at trade 156, and a system with no edge whatsoever passes 5.05 per cent of the time. Judge the same record after every trade from trade 20, stopping the first time it clears 1.645, and the same dead system passes 24.25 per cent of the time.
Prerequisites: Lesson 64, for the bar and the count it is computed from, lesson 19, for how a record accumulates evidence at all, and lesson 62, for the sentence the rule was written as before any of it was run.
What a forward test is actually for
The usual case for forward testing is realism: real fills, real latency, real costs, none of the conveniences a backtest quietly grants itself. That case is true and it is not the important one. Lesson 63 showed that a careful backtest can model costs and slippage honestly and still be worthless, and lesson 64 showed why: the number that discredits it is not any of the market assumptions but the count of configurations that were auditioned to produce it.
A forward test fixes that count at one. There is exactly one rule running, and no version of it that you might have preferred is being run alongside and quietly discarded. That, and not the realism, is the property worth paying for, and lesson 64 priced it: a single configuration clears at 1.645 where a 15,000-cell grid must clear 4.499, and since trades needed scale as the square of the bar, the live record needs 7.48 times fewer of them to say the same thing.
So a forward test is not a cheaper backtest or a more honest one. It is the same evidence bought in a different currency. A backtest spends configurations and gets its answer instantly; a forward test spends one configuration and pays in time. What time costs is arithmetic, and here it is in full. An edge of d in R per trade produces a t-statistic of d times the square root of the trade count, so clearing the bar takes 1.645 divided by d, all squared. Converting that into months needs a trading pace, and this page uses forty trades a month throughout: a rate that is not derived from anything, only stated, so a reader on a different pace can divide the trade count by their own.
| Edge, R per trade | Trades to clear 1.645 | Months at forty a month |
|---|---|---|
| 0.05 | 1,083 | 27.1 |
| 0.10 | 271 | 6.8 |
| 0.15 | 121 | 3.0 |
| 0.20 | 68 | 1.7 |
| 0.30 | 31 | 0.8 |
| 0.50 | 11 | 0.3 |
Read the first two rows and the shape of the problem is visible. The edges people actually find, once costs and the benchmark have been subtracted, live in the top half of that table, and the top half is measured in years. Halving the edge you are willing to accept quadruples the wait, because the bar is fixed and only the square root of the count grows to meet it. That is why forward testing gets abandoned: not because anyone disputes it, but because the honest horizon for a tenth of an R is six months and eight tenths, and almost nobody starts a test they have to leave alone that long.
The sheet this module began in lesson 62 gains a fourth column here, and it is the shortest one on the page: the number of trades the test will run for, written down before the first trade. Almost nobody you are trading against writes it down. It costs one integer and an act of self-restraint, and the reason it is missing is the same reason the configuration count was missing in lesson 64 — committing to it in advance is the only version that can cost you anything.
The same record, judged five ways
Take a system with no edge at all: 156 trades whose outcomes are drawn from a distribution centred exactly on zero. Do that twenty thousand times from seed 20260903, and each time compute the running t-statistic in the ordinary way, the mean trade over the standard error of the mean. Now vary one thing only, which is when you are allowed to look.
Judged once, at trade 156, and only there: 5.05 per cent of these dead systems clear 1.645. The test is calibrated, which is the point of running the check — a bar meant to pass five in a hundred passes 5.05 in a hundred, and everything below is measured against that.
| When you may look | Looks | Dead systems that pass | Median trade where they pass |
|---|---|---|---|
| Once, at trade 156 | 1 | 0.0505 | 156 |
| Halfway and at the end | 2 | 0.0795 | 78 |
| Four evenly spaced gates | 4 | 0.1187 | 78 |
| Monthly, thirteen trades apart | 12 | 0.1992 | 39 |
| After every trade from trade 20 | 137 | 0.2425 | 32 |
Nothing in that table changes the system, the bar, the trade count or the distribution the trades come from. The only thing that changes is how many times the record is allowed to be declared a success. Four gates — paper, then a quarter size, then half, then full, which is the deployment plan every systematic-trading guide recommends including the one this page replaces — passes a dead system 11.87 per cent of the time. That is more than twice the rate the bar was chosen to permit, and it is bought entirely with looking.
A record you are allowed to stop is a search over stopping times.
Which means it can be priced exactly as lesson 64 priced a grid. Solve for the threshold that restores five per cent under after-every-trade watching and it is 2.569. That is not a new kind of number: it is the bar lesson 64 gives for eleven configurations. Watching one live record as it accumulates is worth searching eleven, and it costs 2.44 times as many trades to reach the same conclusion, which turns the 6.8-month horizon for a tenth of an R into 660 trades and 16.5 months.
The version that should worry a reader running three candidate systems side by side, which is the ordinary way people forward test, is worse than the sum of its parts. Three records, each watched after every trade from trade 20, each with the same 1.645 bar: 55.65 per cent of the time at least one of the three declares itself alive when none of them is. A majority. The plan looks like diligence and it is a coin weighted in favour of adopting something.
So fix the horizon before the first trade, and read the record once, at the end.
What this does not settle
That the trades are independent and identically distributed. They are neither, and the assumption is doing real work here: consecutive trades from one system share the same weeks, the same regime and often the same open position risk, so the effective number of independent observations in a 156-trade record is smaller than 156. That pushes every horizon in the first table up rather than down, so the page is optimistic in the direction that matters most, and it does not say by how much because that depends on the system.
That the trade distribution is normal. It is not, and a system with rare large winners has a t-statistic that converges to normality slowly, which makes the 1.645 bar wrong at exactly the trade counts the first table calls short. On thirty-one trades at three tenths of an R the normal approximation is doing more work than the evidence is.
That stopping early is the only way looking hurts. The simulation models one rule — stop the first time the record clears the bar — and real traders also stop when a record looks bad, which is the mirror rule and is not modelled anywhere on this page. It removes systems that would have recovered rather than adopting systems that never worked, so its cost is invisible and uncountable, and nothing here prices it.
That a fixed horizon is compatible with running an account. It frequently is not. A system losing money for four consecutive months should be switched off, and switching it off is correct risk management even though it destroys the statistical property this whole page rests on. The honest position is that the two goals genuinely conflict, that risk management wins, and that a test abandoned for risk reasons has to be reported as abandoned rather than quietly recorded as a failure of the idea.
That the edge is the same size at the end as at the start. Everything above treats d as a constant over the horizon, and the entire reason a forward test can be worth doing is that live conditions differ from historical ones. If the edge is decaying while the test runs, the record is measuring an average of something that no longer exists by the time the average is significant, and the longer the honest horizon the worse that gets.
And the concession that costs most: the horizon has to be computed from the effect size, and the effect size is what the test is for. The first table asks for d before it will tell you how long, which is circular, and there is no way round it. The least dishonest repair is to fix the horizon from the smallest edge that would be worth trading after costs rather than from the one the backtest reported, because the second of those is the number lesson 63 spent a page showing you cannot use. Lesson 66 takes the record that does pass and asks what produced it, and finds that on this course’s own sixty closes 57.7 per cent of the winning rule’s return was the market rather than the rule.
Problems
- Price your own horizon. Take the smallest edge in R per trade that would be worth trading on your instrument after the round-trip cost you computed in lesson 63, and read the trade count off the first table, or compute 1.645 divided by that edge, squared. Ten minutes, and you end holding one number, the trades your forward test has to run for, which is the number you have to commit to before you place the first one.
- Count your looks. Write out your actual deployment plan — paper, small, larger, full, plus every review you have scheduled — and count the points at which you could decide the system is working. Half an hour, and you end holding one number, the count of gates, which the second table converts into the rate at which your plan adopts a system that has nothing in it.
- Run the null on your own plan. Simulate a system with no edge over your horizon, a thousand times from a seed you write down, and apply your own gates and your own bar to each simulated record. Count how often it would have been adopted. An evening, and you end holding one number, that fraction, which is the false-adoption rate of the process you were about to trust with money.
Sources. Peter Armitage, C. K. McPherson and B. C. Rowe, “Repeated Significance Tests on Accumulating Data” (Journal of the Royal Statistical Society, Series A, 1969), for the result the second table reproduces: that testing the same accumulating record repeatedly inflates the false-positive rate far above its nominal level. Stuart J. Pocock, “Group Sequential Methods in the Design and Analysis of Clinical Trials” (Biometrika, 1977), for the practice of fixing the number of interim looks in advance and raising the threshold to match, which is what the 2.569 on this page is. Peter C. O’Brien and Thomas R. Fleming, “A Multiple Testing Procedure for Clinical Trials” (Biometrics, 1979), for the alternative of spending the error rate unevenly across looks, which is the better design this page does not use. David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, “Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance” (Notices of the American Mathematical Society, 2014), for the bar this lesson carries in from lesson 64 and converts into eleven configurations.
The Price of Looking
The bar a search has to clear, and the count it is computed from.
Read Lesson →Backtesting as Evidence
Why the edge a backtest reports cannot be the one you fix a horizon from.
Read Lesson →Educational only. Trading involves substantial risk of loss. Not financial advice. Past performance does not guarantee future results.
💬 Discussion (0 comments)
Loading comments...
Ready to Trade with Signal Pilot?
Apply your trading education with professional indicators and real-time market analysis tools.
Back to Signal Pilot →