Signal Pilot
📚 Appendix • Measurement

How to Collect a Base Rate

Reading time ~11 min • Measurement
Signal Pilot
Professional Trading Education
0%
You’re making progress!
Keep reading to mark this lesson complete

Fourteen problems in this course end by telling you to go and count something, and not one of them says how. This is the how, and it is four decisions, all of which have to be made before you look. That timing is the whole appendix: every one of the four can also be made afterwards, in a way that feels reasonable at the time and hands you the answer you were hoping for. Below, noticing a level hold a little more often than you notice one break turns a true 30% into 66%, with nobody anywhere being dishonest.

Prerequisites: Lesson 19, which is what a proportion from a small sample is worth once you have one, and lesson 23, which is the discipline of writing things down at the time.

This page is not a lesson. It is the piece the lessons keep assuming, gathered in one place so they can point at it instead of repeating it. Lessons 25, 26, 28, 29, 30 and 31 all end by asking for a rate or a two-by-two, and between them that is fourteen problems, all of which are the same exercise wearing different clothes.

Four decisions, and the reason they come first

The first is what counts as one observation. It has to be defined so that another person, handed your rule and your data, would produce the same list you did. Two things break that. A rule that needs a judgement at the moment of counting — “a clear touch”, “a proper wall” — is not a rule, because you will be making that judgement with the outcome in front of you. And a rule that mentions the outcome is worse than useless: “levels that acted as support” has already counted the answer.

The second is the population you are drawing from. You have to be able to list the candidates before you know how they turned out. On an instrument and a timeframe you name in advance, over a period you name in advance, every case matching the rule is in the list — including the boring ones, the ones you were not watching, and the ones where nothing happened.

The third is the criterion, and it is the one people set afterwards without noticing. How far past the level counts as through. How long counts as holding. How much size counts as large. Every one of those is a dial, and if the dial is still free when you start counting you will end up at whatever setting makes the number look best. Write it down before the first observation, in units that do not depend on the instrument — a fraction of an ATR rather than a number of cents, so the same rule survives moving to a different market.

The fourth is that the cases have to be consecutive. Not the ones you remember, not the ones you screenshotted, not the ones that were interesting. Take every case in the window, in order, until you have your thirty or forty. A record of memorable cases measures your memory rather than the market, and what that costs is arithmetic rather than a warning.

What selection does to a rate

Suppose the truth on your instrument is this: across forty consecutive touches of a level, twelve held and twenty-eight broke. The rate is 30%, and nothing below changes it. What changes is how much of it reaches your notebook. A hold that is defended with size is an event; a level that quietly gives way is not. So suppose you notice and log a hold somewhat more often than you notice and log a break.

You notice a holdYou notice a breakTouches you logHolds among themYour answer
100% of the time100% of the time40.012.030.0%
80%60%26.49.636.4%
80%40%20.89.646.2%
80%30%18.09.653.3%
90%20%16.410.865.9%

Read the first row first, because it is the only one that is not an assumption. Log everything and you get the rate. Every row after it is the same forty touches, the same twelve holds, and a notebook that saw some of them and not others.

The last row is the one worth sitting with. A level that holds three times in ten reads as holding two times in three, and there is no dishonesty anywhere in it. Nobody made a case up. Nobody deleted a loss. Somebody noticed the dramatic ones nine times in ten and the quiet ones two times in ten, which is not a character flaw, it is what attention does.

And the structure of the damage is exactly the structure lesson 26 costed. In odds rather than percentages, the true odds here are 12 to 28, and your notebook multiplies them by the ratio of the two noticing rates: 0.9 divided by 0.2 is 4.5, and 12 to 28 multiplied by 4.5 is 1.93 to 1, which is 65.9%. Selection is a multiplier on your odds. The difference from lesson 26 is that this one is applied without your choosing it and does not appear anywhere in the record it produces.

Counting into four cells

Once the sample is honest, everything the lessons ask for comes off one table. Take the same forty touches and add the second column: what you said before each one resolved.

What you called itIt heldIt brokeRow total
You called it a hold8816
You called it a break42024
Column total122840

Every number the course has asked you for is in there. The base rate is the column totals: 12 of 40, which is 30%. Your read is the two columns read downward: of the twelve that held you called eight, which is 67%, and of the twenty-eight that broke you called twenty, which is 71%. In lesson 26’s notation you are a 67/71 read.

And the answer to the question lesson 26 spent a table computing — you called this one a hold, what is the chance it is — is the first row, read across: 8 of the 16 you called, which is 50%. No formula, no multiplier, no arithmetic beyond division. The two-by-two gives you the posterior by counting, which is why it is the shape to collect in.

One thing falls out of it that is worth keeping. When it says hold, a 67/71 read multiplies your odds by 2.33, which is exactly what a 70/70 read does — because the multiplier is a ratio and not a pair, so a read that is worse on one side and better on the other can be the same read. When it says break the two come apart, 0.47 against 0.43, which is lesson 31’s point about asymmetric reads showing up in miniature. Either way, the second decimal of your own accuracy is not where the leverage is. The sample it was measured on is.

What this does not settle

That thirty or forty observations are enough. They are not, for anything subtle. Thirty is enough to see a large effect and nowhere near enough to see a small one, and lesson 19 is where that gets quantified. This appendix is about collecting a sample that means what it says, which is a different question from how big it has to be, and the second question does not arise until the first is answered.

That a consecutive sample is an unbiased one. It removes the selection you were doing. It does not remove the selection the market did: your window covered particular hours, a particular volatility regime and a particular instrument, and the answer belongs to those. Two consecutive samples from two months can disagree honestly, and the useful response is to run it again rather than to average them.

That fixing the criterion in advance makes it the right criterion. It makes the count honest, not correct. A badly chosen threshold, fixed in advance and applied consistently, gives you a clean measurement of the wrong thing — and you will find that out from the count, which is the point of writing the criterion down where you can read it later.

That the counting is the hard part. The counting takes an hour. The hard part is that you already have a belief, the outcome is visible before you write the row, and every one of the four decisions can be quietly revisited in the direction of that belief. Which is why they go on paper before the first observation, and why problem 1 is about handing your definition to somebody else.

That a number produced this way is an edge. It is a description of what happened, in a form that can be argued with. Turning it into a reason to trade takes lesson 17, and a rate that comes back at the base rate is a result too — usually the most valuable one, because it is the one that stops you.

The count is easy and the honesty is expensive. Everything above is a way of spending it before you can see what it buys.

Problems
  1. Make your definition survive a second reader. Write the rule for one observation in three sentences, then apply it to twenty cases. Hand the three sentences and the same twenty cases to somebody else — or to yourself in a week, which is nearly as good and much easier to arrange — and count the disagreements. More than two or three and the rule still contains a judgement you have not written down. Find it and write it down. This costs an evening and it is the difference between a measurement and an opinion with a number attached.
  2. Price your own selection. Pick a level rule and a month you have already traded. Count it twice: once from memory, listing the cases you can recall, and once consecutively from the chart, taking every case in order. Compare the two rates. The gap is your own multiplier from the first table, measured rather than assumed, and it is the number that tells you how much of anything you believe from watching came from watching.
  3. Build one two-by-two, and read the posterior off it. Over thirty consecutive cases, record your call before each resolves and the outcome after. Fill in the four cells. Then read the three numbers straight off: the base rate from the column totals, your two rates down the columns, and the posterior across the row you actually act on. Do this once and every problem in this course that asks for a rate becomes the same twenty minutes of work.

Sources. Joseph Berkson, “Limitations of the Application of Fourfold Table Analysis to Hospital Data” (Biometrics Bulletin, 1946), which is where the second table’s failure mode was first set out: a fourfold table built from cases that were selected on something related to both columns gives an association that is not in the population it came from. Amos Tversky and Daniel Kahneman, “Availability: A Heuristic for Judging Frequency and Probability” (Cognitive Psychology, 1973), for the mechanism behind the first table: how easily a case comes to mind stands in for how often it happened, and the vivid cases come to mind more easily. Joseph Simmons, Leif Nelson and Uri Simonsohn, “False-Positive Psychology” (Psychological Science, 2011), for the third decision: the paper demonstrates how much a result can be moved by choices about criteria and stopping rules that are made after the data is in view, by researchers who are not trying to cheat.

None of this is new and none of it is difficult. It is here because the difficult part is doing it before the answer is visible, and every lesson that ends by asking for a rate is asking for exactly that. Go back to the problem that sent you here: whatever it was counting, the method is this one.

Related Lessons
Lesson 19

How Long Until You Know

How big the sample has to be, once you can collect an honest one.

Read Lesson →
Lesson 23

Keeping the Record

The ten fields, and why they are written at the time.

Read Lesson →
Lesson 26

The Order Book Is Theater

Where the two-by-two first appeared, and what it is worth.

Read Lesson →
Lesson 30

Volume Profile

The most recent lesson to end by asking for a rate nobody has.

Read Lesson →
Educational only. Trading involves substantial risk of loss. Not financial advice. Past performance does not guarantee future results.

💬 Discussion (0 comments)

0/1000

Loading comments...

← Curriculum Curriculum →

Ready to Trade with Signal Pilot?

Apply your trading education with professional indicators and real-time market analysis tools.

Back to Signal Pilot →