Signal Pilot
🟠 Advanced • Lesson 70 of 85

Machine Learning as a Filter

Reading time ~13 min • Module 8: Building a System
Signal Pilot
Professional Trading Education
0%
You’re making progress!
Keep reading to mark this lesson complete

Lesson 69 left this page a grid. Six hyperparameters at five values each is 15,625 configurations, a grid a beginner would call modest, and lesson 64’s arithmetic says a search that wide, judged over 156 trades, manufactures 0.318 of an R a trade before the market has done anything at all. That is the familiar objection to machine learning in trading and it is the answerable one: count the grid, raise the bar from 1.645 to 4.507, and see whether the model clears it. The objection that cannot be answered that way has no model in it at all. A filter does not create trades, it removes them, so it can only ever hand you back the losers it managed to catch, and a strategy that looks good has almost none to catch. Lesson 63’s winning rule has six winners and one loser. A filter equally good at keeping winners and rejecting losers would have to be right about 96.04 per cent of unseen trades to break even on that rule, and a perfect one, catching the single loser and touching nothing else, would improve it by 4.30 per cent.

Prerequisites: Lesson 64, for the expected-maximum formula and the bar a configuration count sets, lesson 63, for the seven trades this page takes apart, and lesson 17, for the two numbers a profit factor is made of.

What a grid costs before it has seen a price

A rule has parameters and lesson 64 counted them. A model has hyperparameters, which are the same thing under a different name, and it has one dimension a rule does not: which features go in. Both multiply out, and the arithmetic that prices a search does not care which is which. It counts configurations, and the count is known before anything runs.

Take the grid the tease left. Six hyperparameters at five values each is 15,625, and the substitution goes in full so that none of it has to be taken on trust. One divided by 15,625 is 0.0000640, and the standard normal value leaving that much in the upper tail is 3.83028. Then 15,625 times e is 42,473, one divided by that is 0.0000235, and its normal value is 4.06963. Weighting the two by 0.4228 and 0.5772 gives 3.96843, which is a t-statistic. Divide by the square root of 156, which is 12.49, and the manufactured edge is 0.318 of an R a trade. The bar a search that wide has to clear before five per cent of worthless searches would beat it is 4.507.

Now the dimension nobody counts. A menu of twenty features with an instruction to pick three to six of them is the usual shape of this exercise, and that instruction reads like a simplification. It is a search. Choosing exactly four from twenty is 4,845 distinct models, choosing three to six is 60,249, and the menu was written down before you started, so the count is available to anybody who wants it. Multiply the four-from-twenty figure by the hyperparameter grid and an afternoon’s work is 75,703,125 configurations.

What you variedConfigurationsBar at five per centR a trade manufactured over 156 trades
Three hyperparameters, five values each1253.3460.209
Four hyperparameters, five values each6253.7690.250
Six hyperparameters, five values each15,6254.5070.318
Four features chosen from a menu of twenty4,8454.2520.295
Both of the last two together75,703,1256.0610.453

The last row is the honest count for one afternoon, and 6.061 is a t-statistic that a few hundred trades will not produce however good the idea was. Read the third column down the table, though, and the shape is gentler than that sounds: going from 15,625 configurations to 75,703,125 is 4,845 times as much searching and moves the bar by 1.554. The penalty grows with the logarithm of the count, which cuts the way nobody mentions. You do not need a fashionable model to be badly exposed. Four hyperparameters and no feature selection at all is already at 3.769.

So suppose you do all of that properly. Suppose you count the grid, clear the bar, and hold in your hand a model that genuinely ranks your own setups better than chance. The rest of this page is about whether that is worth anything, and it needs no configuration count, no simulation and no assumption about markets.

Set it up the way the money actually works. A filter does not find new trades. It sits on the setups your method already produces and removes some of them, so the honest comparison is total profit over the same original opportunities, not profit per trade taken. Write your win rate as p and your payoff as b, both from lesson 17, so a winner pays b and a loser costs one. Let the filter keep a fraction s of the winners and correctly reject a fraction t of the losers. Per original trade it forgoes p times one minus s times b, and it saves one minus p times t. It is worth having when the second is at least the first:

(1 − p) t  ≥  p (1 − s) b

Divide it two ways and it says two different useful things. Divide by p times one minus s and the losers the filter drops must outnumber the winners it drops by at least b, your payoff ratio, which is a test you can run on a trade log with two counts and no statistics in it at all. Divide instead by one minus p and the right-hand side becomes p b over one minus p, which is gross profit over gross loss, the quantity usually called the profit factor and the only new name on this page. It is lesson 17’s two numbers written as one.

If the model is equally good in both directions, so that s and t are the same number a, the condition collapses to a threshold on accuracy alone: a must be at least the profit factor divided by one more than the profit factor. Turned round, a filter of accuracy a earns its place on any strategy whose profit factor is below a divided by one minus a.

At 52 per cent accuracy, on any strategy with a profit factor below 1.08.

At 55 per cent, below 1.22.

At 60 per cent, below 1.50.

At 65 per cent, below 1.86.

Fifty-two per cent is not a placeholder. Gu, Kelly and Xiu measured the whole family of these models out of sample on monthly returns across thousands of stocks, and their best network reached a monthly out-of-sample R-squared of about four tenths of one per cent. Convert that into the language of this page: under a bivariate normal the chance a prediction’s sign matches the outcome’s sign is one half plus the arcsine of the correlation, divided by pi; the correlation is the square root of the R-squared, which is 0.0632, and the accuracy is 52.01 per cent. The best published result in the field is a filter for strategies with a profit factor under 1.08.

Almost nobody who has fitted a model to their own trades has divided their gross profit by their gross loss and put the two numbers side by side. It is two sums off a trade log and one division, and it is the whole of the decision.

The sheet this module has been building since lesson 62 gains no ninth column. This is where it gets spent, all eight at once, on the one rule the course has followed since lesson 63.

The sheet, spent on one rule

Lesson 63 searched 253 moving-average configurations across sixty closes and kept a 2-bar average against a 5-bar one: seven trades, 10.70 a share gross, 9.84 after costs, against 5.68 for holding. Fill in the eight columns and the sheet reads like this.

The sentence, from lesson 62: go long when the 2-bar average of the closes crosses above the 5-bar, go flat when it crosses back, execute at the next close.

The pair, from lesson 63: 4.16 over the benchmark, against 52.5 per cent of structureless searches returning 4.16 or more.

The bar and the count, from lesson 64: 253 cells, a bar of 3.537, a t of 3.65, which clears until you set a minimum-trade floor, at which point the fraction runs from 0.68 per cent at a floor of seven to 50.48 per cent at a floor of two.

The trades fixed in advance, from lesson 65: blank. The seven are the ones that happened.

The exposure share, from lesson 66: 0.5085, and of the 10.70 gross the market supplied 2.95.

The quit depth, from lesson 67: blank.

The capacity, from lesson 68: an average trade of 0.9101R is 1.4055 a share, which on a price near 103 is 136 basis points, so in lesson 68’s small cap the trade holds 1,390,545 dollars and the best single position is 618,020.

The position cap and its check interval, from lesson 69: blank.

Three blanks, and a second column whose two figures are the wrong way round. That is the verdict, and it arrives before the sheet is half read. But run it all the way to the end anyway, because the question this page is asking is not whether the rule is any good. It is whether a model could make it better, and that question has an answer even for a rule that is worthless.

The seven trades net of costs are 2.577, 1.877, 2.377, 1.377, minus 0.423, 1.177 and 0.877. Six winners summing to 10.262 and one loser of 0.423, so the profit factor is 24.26, the average winner is 1.1075R, the average loser is 0.2739R and the payoff ratio is 4.04. Now put the sheet’s other columns to work and watch what happens as each one is applied. The last row of the table below is lesson 63’s own null, rerun for this page: the 59 bar-to-bar moves drawn with replacement, 4,000 series from seed 20260903, the identical 253-cell search on each at the seven-trade floor the winner itself met, and the profit factor of whichever cell won. Its median is 2.34, and 3.08 per cent of those structureless searches return a winner whose profit factor is 24.26 or better.

What has been subtractedProfit factorAccuracy a filter must beatThe most a perfect filter could add
Nothing36.6797.35%2.80%
Costs, at lesson 63’s convention24.2696.04%4.30%
Costs, and the market’s share from lesson 6614.2293.43%7.57%
What the same search returns on nothing, median2.3470.01%74.91%

The fourth column is the one worth stopping on, and it is a second pure number falling out of the same inequality. A perfect filter removes every loser and keeps every winner, so the most it can possibly add is the entire gross loss, and as a share of what the strategy already nets that is one divided by the profit factor minus one. On this rule it is 4.30 per cent. Not 4.30 per cent if the model is good: 4.30 per cent if the model is flawless, because there is only one loser in the record and it is worth 0.423 out of 9.839.

Read the two middle columns together, down the table, and they move in opposite directions for the same reason. As the profit factor is deflated toward the truth — costs first, then the market’s share, then what a search of the same width returns on data with nothing in it — the accuracy the filter needs falls from 97.35 per cent to 70.01, and the prize rises from 2.80 per cent to 74.91. The search manufactured the edge, and in the same motion it manufactured the reason no model could improve it.

A filter can only give you back your losers, and a good strategy has almost none.

Which is why the bottom row matters more than it looks. A profit factor of one is a strategy with exactly no expectancy, and there the threshold is fifty per cent and a 52 per cent filter is genuinely worth having. Count it out on three hundred trades at a payoff of two: a hundred winners paying two and two hundred losers costing one is zero. Keep 52 of the winners and 96 of the losers and you have 104 minus 96, which is 8R over the original three hundred, or 0.027 of an R a trade. The filter is not improving the strategy. It is the entire strategy, and it is a small one. Move up to a profit factor of two and the same 52 per cent filter takes 0.22R a trade away from a strategy making 0.50.

So add your winners, add your losers, divide, and read the answer as the accuracy any filter has to beat before you build it.

What this does not settle

That the filter is balanced. The accuracy threshold assumes the model is exactly as good at keeping winners as at rejecting losers, and the inequality it came from assumes nothing of the kind: it asks only that specificity divided by one minus sensitivity clears the profit factor. A model tuned hard toward rejecting losers, accepting that it will keep almost every winner it is shown, can clear that bar at an overall accuracy far below the threshold, and that is the direction worth building toward. This page prints the balanced number because it is the number a classifier reports by default, which is a statement about software rather than about markets.

That the trades the filter keeps pay what the ones it drops would have. If the model preferentially keeps the large winners, the arithmetic improves and the threshold falls, and a model that ranks trades by expected size rather than sorting them by sign is a different and better object than the one priced here. The page prices a classifier because a classifier is what people build, and because the accuracy figure the field quotes is only defined for one.

That a profit factor computed on seven trades is a number. It is a ratio of two sums, and one of those sums has a single term in it. Had lesson 63’s one losing trade cost 1.423 rather than 0.423, the profit factor would be 7.21, the threshold 87.82 per cent and the prize 16.10 per cent, from the same rule on the same closes. On a few hundred trades the ratio steadies, but it is never precise, and the second decimal on any threshold computed from it is decoration.

That 52 per cent is the right benchmark for your model. Gu, Kelly and Xiu forecast monthly returns across a cross-section of thousands of stocks using close to a hundred predictors, which is a different problem from grading setups your own method already produced: they have vastly more data and no access to your entry rule, and you have the reverse. The figure belongs on this page as a check on the order of magnitude that the field’s best work achieves, and it is not a prediction of what your own classifier will do.

That the configurations are countable. This is lesson 64’s concession and a model makes it worse rather than better. The hyperparameters you typed into a grid are countable. The feature you engineered, plotted, disliked and deleted is not, and neither is the target you redefined from next-bar direction to next-three-bar direction because the first one would not train. Every retraining on fresh data is another configuration, and a schedule that retrains monthly for two years has quietly run twenty-four more.

And the concession that costs most: this page has priced a model and said nothing about whether one might be right for a reason. A filter that refuses trades taken in the first fifteen minutes because your losers cluster there is a rule, it can be written in one line, and it counts as a single configuration with a bar of 1.645. What the arithmetic on this page penalises is not learning something from your own record. It is learning many things from it and keeping the best, which is what fitting a model is. Lesson 71 opens the portfolio module by asking what four rules buy that one does not, measures the correlation between rules instead of assuming it, and finds that the 14.7 months lesson 67 priced comes down to 14.1 and not to 3.7.

Problems

  1. Count your grid. Take the last model you fitted, or the last one you were shown, and count it properly: every hyperparameter value you tried, multiplied out, multiplied again by the number of feature subsets you considered. Read the bar off the first table above or compute it from the count. Ten minutes, and you end holding one number, the t your model had to beat, which is almost certainly not 1.645.
  2. Compute your own threshold. Take your last two hundred trades, add up everything the winners made, add up everything the losers cost, and divide the first by the second. Half an hour, and you end holding one number, your profit factor, from which the accuracy a filter has to beat is that number divided by one more than it, and the most a perfect filter could add is one divided by that number minus one.
  3. Measure the filter you could actually build. Split your log in time order. On the first half, pick one feature you can give a reason for in a sentence and find the single threshold on it that best separates your winners from your losers. Apply that one threshold, unchanged, to the second half, and count two things: how many losers it dropped and how many winners it dropped. An evening, and you end holding one number, the ratio of the first count to the second, which has to exceed your payoff ratio before the simplest filter you can build is worth anything, and which you should compare against what you expected before you split the log.

Sources. David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, “Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance” (Notices of the American Mathematical Society, 2014), for the expected-maximum formula the first table is computed from, carried over from lesson 64 and applied here to a hyperparameter grid rather than to a parameter one. Shihao Gu, Bryan Kelly and Dacheng Xiu, “Empirical Asset Pricing via Machine Learning” (The Review of Financial Studies, 2020), for the out-of-sample measurement this page converts into the 52 per cent that every threshold in it is read against. Marcos López de Prado, Advances in Financial Machine Learning (Wiley, 2018), for the argument that ordinary cross-validation leaks in finance and for the multiple-testing correction the first table is a small instance of. Robert D. Arnott, Campbell R. Harvey and Harry Markowitz, “A Backtesting Protocol in the Era of Machine Learning” (The Journal of Financial Data Science, 2019), for the protocol that requires the trial count to be recorded before the search, which is the sheet column lesson 64 added and this page spends.

Related Lessons
Lesson 64

The Price of Looking

The expected-maximum formula and the bar a configuration count sets.

Read Lesson →
Lesson 63

Backtesting as Evidence

The seven trades this page takes apart, and the search that produced them.

Read Lesson →
Lesson 17

Expectancy

The win rate and the payoff a profit factor is made of.

Read Lesson →
Educational only. Trading involves substantial risk of loss. Not financial advice. Past performance does not guarantee future results.

💬 Discussion (0 comments)

0/1000

Loading comments...

← Previous Lesson Next Lesson →

Ready to Trade with Signal Pilot?

Apply your trading education with professional indicators and real-time market analysis tools.

Back to Signal Pilot →