Multiple Testing: Why Enough Hypotheses Guarantee a False Positive
Here is the uncomfortable thing about that sentence. Every single test in the run is exactly as careful as it was when it was the only one. Nobody loosened anything, nobody cheated, nobody picked a friendlier thresholdThe bar a reading must get past before anybody treats the result as worth reporting. What a bar set at 5 per cent claims about one single test is settled earlier in these notes and used here without being reopened.. And by the twentieth of them, the chance that at least one has fired for no reason at all has gone from one in twenty to roughly two in three. The gap between one in twenty and two in three is not a flaw in any test. The gap is arithmetic with one moving part, and the arithmetic breaks on the very record it was computed from.
Independence, and a threshold of 5 per cent. Independence between two things, and the claim a single test held to 5 per cent makes, are both settled earlier in these notes. The subject here is what happens to that threshold the moment the test is run more than once.
One multiplication, repeated. If a single test comes back clean with chance 0.95, and the tests do not lean on one another, then all of them come back clean with chance 0.95 multiplied by itself once per test. The repeated multiplication is the entire mechanism. There is no second idea underneath it.
The Ashwin rule and its thirteen settings. The invented Ashwin rule reads the Nakshatra unit's change for a finished month and calls the next month up when that reading sits above a threshold. Its gridThe full list of settings a search will walk through, decided and written down before any walking starts rather than added to along the way. holds thirteen whole number settings. Thirteen settings is a real count of tests to put into the arithmetic instead of an imagined one.
What is the arithmetic, in one line?
Picture a stall on a street with a crate of tomatoes. The stallholder picks up one basket, turns it over, and looks for a bad one. Turning over one basket is a single check, and whatever the chance is of a bad one turning up by pure accident, it is what it is. Now picture the same stallholder going through twenty baskets and reporting the worst one he found. Nothing about his eyesight changed between the two stalls. The number of times he looked before choosing what to report is the only quantity that changed.
Put a number on it. A test held to a threshold of 5 per cent comes back clean 95 times in a hundred when there is genuinely nothing there. Write that as 0.95. Run two such tests that do not lean on one another and both come back clean 0.95 multiplied by 0.95 of the time. The product is 90.25 per cent. Everything else is at least one of the two firing, and one hundred less 90.25 leaves 9.75 per cent. The whole mechanism is that one multiplication repeated once per test, and for any count of tests the chance that at least one fires is 1 minus 0.95 raised to that count.
Notice the quantities missing from the picture. There is no adjustment for how important the question was, no allowance for the researcher being careful, and no term for how convincing the result looked. The only quantity that goes in is a count. The arithmetic is easy to state and easy to forget for the same reason: the quantity driving it is the least interesting fact about the work.
Where does the 0.95 in this arithmetic come from?
What does the ladder look like as the count of tests rises?
Work it out at a spread of counts and put them in order. Nothing in this table is estimated and nothing in it is sampled. Each row is 0.95 raised to the count in the first column, taken away from one.
| Count of tests | Chance that at least one fires by accident | What that test added |
|---|---|---|
| 1 | 5.00 per cent | 5.00 points |
| 2 | 9.75 per cent | 4.75 points |
| 3 | 14.26 per cent | 4.51 points |
| 5 | 22.62 per cent | 4.07 points |
| 10 | 40.13 per cent | 3.15 points |
| 13 | 48.67 per cent | 2.70 points |
| 14 | 51.23 per cent | 2.57 points |
| 20 | 64.15 per cent | 1.89 points |
| 45 | 90.06 per cent | 0.52 points |
| 100 | 99.41 per cent | 0.03 points |
Every extra test adds less than the test before it did. The climb looks gentle at the start and is not. The first test hands over 5.00 points. The thirteenth hands over 2.70. The hundredth hands over three hundredths of a point. A glance at the right hand column alone suggests that the count of tests stops mattering, and that has it backwards: the column shrinks precisely because the figure it is being added to is running out of room underneath one hundred.
Ten tests read 40.13 per cent and twenty read 64.15 per cent. The count of tests doubled. Did the figure?
Drawn as a curve from one test to a hundred, the shape is easier to hold on to than the table is. The curve leaves the floor quickly, bends over somewhere in the teens, and then spends the rest of its length creeping along under the top of the frame without ever arriving there.
At how many tests does at least one false positive become more likely than not?
Read the table again at the two rows nobody looks at. Thirteen tests read 48.67 per cent. Fourteen tests read 51.23 per cent. Somewhere between the thirteenth test and the fourteenth, the run stops being a place where a false positive would be unlucky and becomes a place where the absence of one would be.
At the fourteenth test, at least one false positive becomes more likely than not, and nothing whatsoever about any of the tests changed to make that happen. No threshold moved. No test got weaker. No record was edited. Not one of the fourteen is any less careful than the single test that one question on its own would have called for. The only quantity that moved between 48.67 and 51.23 is the count, and the count is the part of the work that usually goes unrecorded.
The heading above deserves an honest reading. Enough tests do take the chance of at least one false positive above one half, then above nine tenths, and then as close to certain as anybody would like. The arithmetic never arrives. Climbing towards one hundred per cent and reaching it are two different claims, and the heading sounds like the second.
At what count of tests does at least one false positive become more likely than not, and what changed about the tests to make it happen?
The control below drags the count of tests from one up to a hundred. Before it moves: does the figure rise in even steps?
Drag the count of tests and watch where the curve stops climbing
One control, one number: how many tests were run. Every test is held to the same threshold of 5 per cent and none of them changes. The curve draws itself from the first test up to the count the control is set to, the dashed line across the middle is one half and never moves, and the solid line at the top is one hundred per cent and never moves either. The short red bar on the right is what is still left between the curve and that top line. From a start at twenty, then down to one and back up to a hundred, watch which of the two lines the curve gets close to.
Educational illustration. Every test on this panel is held to a threshold of 5 per cent, and the tests are assumed not to lean on one another. The assumption carries the whole curve, and tests run on a single record break it. The count includes tests that were run and then dropped.
The heading above uses a strong word about what enough hypotheses do. Before the next block answers it: does the arithmetic ever actually reach one hundred per cent?
Does the figure ever actually reach one?
Forty five tests take it past nine tenths, at 90.06 per cent. Ninety tests take it past ninety nine hundredths, at 99.01 per cent. A hundred tests read 99.41 per cent. Keep going and the nines keep arriving: two hundred tests read 99.9965 per cent, three hundred read 99.99998, five hundred read 99.9999999993. Check every whole count out to five hundred and the figure is below one hundred per cent at every single one of them.
The arithmetic forces that result. The figure is one minus 0.95 raised to a count, and 0.95 raised to any whole count is a positive number, small but never nought. Something positive is always being taken away from one hundred, so something is always left. Enough tests make at least one false positive as close to certain as anyone cares to name, and no count of tests makes it certain, and those are two different sentences.
The shape is familiar from ordinary life. A walker covers half the remaining distance to a wall, then half of what is left, then half of that. The gap closes until it stops mattering to anybody watching, and the wall is never touched. The arithmetic here is doing the same thing with a different fraction.
Why does the distinction matter? Repeat a strong word as though the arithmetic supplied it and the one quantity the arithmetic settles has been overstated. The correct reading is the practical one and it is not weaker: at fourteen tests a false positive somewhere in the run is more likely than not, and at forty five it is a nine in ten proposition. Nobody needs certainty to act on that. The count is what they need, and the count is what nobody writes down.
How strict does each test have to be to hold the whole run at 5 per cent?
Turn the question round. Instead of fixing the threshold each test is held to and asking what the whole run comes to, fix the whole run at 5 per cent and ask what each test is then allowed. The equation is the same one with the other unknown solved for, and solving it directly beats reaching for somebody else's rule.
Forwards. Each test is allowed a threshold. The chance that a single test comes back clean is one minus that threshold. All the tests come back clean when that clean chance is multiplied by itself once per test, and the chance that at least one fires is one minus the result.
Backwards. Hold the whole run at 5 per cent. Then all the tests coming back clean has to be 95 per cent. So the clean chance of one single test, multiplied by itself once per test, must equal 0.95. The clean chance of one test is therefore the appropriate root of 0.95. The threshold that test is allowed is one minus that root.
In plain words. With thirteen tests, take the thirteenth root of 0.95 and subtract it from one. The result reads 0.3938 per cent. With twenty tests, the twentieth root gives 0.2561 per cent. With 4,096 tests it reads 0.0013 per cent.
Two things are worth noticing about those three figures. The first is that they are not 5 per cent divided by the count, though they land close to it. Five divided by thirteen is 0.3846, and the correct figure is 0.3938. Five divided by twenty is 0.2500, and the correct figure is 0.2561. The gap is small at these counts and it is real, and it exists because multiplying thirteen clean chances together is not the same operation as sharing out a threshold.
The second is the sharpness of the fall: going from one test to twenty divides the threshold each single test is allowed by 19.52, so a search of twenty needs every one of its twenty tests to clear a bar roughly a twentieth of the height of the bar a single test faced. Think of a household budgeting for a wedding. The total does not move, so every additional guest makes every other guest's share smaller, and past a certain count each share is too small to buy anything with. A search behaves the same way, and the total being shared out is the credibility of the whole run.
The cost of that strictness reads most clearly in months. Score the Ashwin rule across the earlier half of the six year record, the stretch holding 29 scoreable months. A reading has to reach 20 of those 29 calls before it clears 5 per cent on its own. Hold the whole run of thirteen settings at 5 per cent instead, so each setting faces 0.3938 per cent, and the bar moves up to 23 of 29. Three more months, and three months is what the count of tests costs once the cost leaves chances and enters the only currency this record has.
Twenty independent tests each need a threshold of 0.2561 per cent to hold the whole run at 5 per cent. A colleague proposes running forty tests instead. Roughly what happens to the per test bar, and why is roughly the honest word?
What does the arithmetic assume, and what does this record do to it?
Every figure above rests on one assumption, stated once at the top and easy to walk past: the tests do not lean on one another. Tests run on a single record always lean on one another. The tests read the same months, in the same order, with the same six years of the same invented unit underneath them. So the ladder is not wrong. The ladder describes a run of tests nobody on a single record ever has.
Two real searches on the six year record fall on opposite sides of what the ladder would suggest. The first takes one fixed question and answers it twenty defensible ways, differing only in the stretchA handful of months that sit next to each other in the record, pulled out and examined without the rest. of the record used, the conventionA rule about how to count something, agreed in advance so that anybody repeating the work counts it the same way. applied to a month that did not move, and the signalWhatever number a rule looks at just before it commits to a call, built only out of months that have already closed. the rule reads. Twelve of those twenty clear the 5 per cent threshold. The second searches 4,096 calendar rules on a version of the record with the calendar pattern taken out of it. Nothing is left in that version to find. Ninety four of the 4,096 clear the threshold, on a denominatorThe count that sits under a share and fixes what the share is measuring. Move it and the share moves too, even when the number on top holds still. of 62 months.
Independence at 5 per cent would suggest one test in twenty. Line the two counts up against that. Twelve out of twenty is sixty per cent, far above one in twenty. Ninety four out of 4,096 is 2.29 per cent, far below it. The ladder governs neither count, and the direction of the error is not predictable from the count of tests, so twenty tests reading 64.15 per cent is an anchor for how much a count of tests matters and is not a figure that may be attached to either of those two tables.
So what is the ladder still for? The ladder answers one question well, and the question is worth asking: does the count of tests matter enough to bother about? At thirteen tests the answer is that a run has roughly a coin flip's chance of throwing up at least one accident even if everything in it is empty, and no amount of relatedness between the tests makes that stop being worth knowing. The ladder cannot give the figure for a particular search on a particular record. Only the search itself gives that.
On the six year record, twelve of twenty related analyses clear 5 per cent while ninety four of 4,096 related calendar rules do. Why can the ladder not be used on either, and what is it still good for?
What is asked when somebody reports one result?
The arithmetic is worth carrying only if it turns into something that can be said out loud. Somebody sends over one figure: a reading, a rule, a finding, one line long. There is a minute available, and nobody is about to redo anybody's arithmetic. Five questions do the work, ordered from the cheapest to ask to the most awkward.
| The question | What it protects against |
|---|---|
| How many tests were run in total, including the ones that were dropped? | The count is the only quantity the ladder needs, and a dropped test counts exactly as much as a reported one. A run of one and the best of a run of thirteen print identically. |
| Were the tests independent of one another, and if not, how do they lean? | On one record they always lean. Once they do, the ladder bounds nothing, and the answer can land well above or well below what independence would suggest. |
| Which threshold was each test held to? | Without it there is no 0.95 to raise to anything, and half the arguments about a reported result are really arguments about this number. |
| Was that threshold fixed before the count of tests was known? | A threshold chosen after the search finished is a description of the winner rather than a test of it. |
| Is there a stretch of the record that was held back and never looked at? | A held back stretch settles the matter and needs none of the arithmetic above. |
A person at a lender asks the first two out of habit rather than distrust. Write a rule down before looking, or pick one out of thirteen after looking, and what lands on the desk reads the same either way. A count is the only thing that tells those two apart. An analyst on the other side of the table gets the same value out of volunteering the count unasked. A figure that travels with its count can be weighed by somebody who does not trust the analyst, and a figure that does not cannot be weighed even by somebody who does.
Everyone has done the household version of this. Ten shops are rung for a quote on the same repair, and the cheapest is reported to whoever is paying. The cheapest of ten is a real number and nobody made it up. The cheapest of ten is also not the price of the repair, and the only quantity separating the two readings is that ten shops were rung. The ladder is an argument about how hard somebody looked, and a stretch of the record held back before anything was chosen is a measurement, so where the second is available the first stops being the interesting conversation.
Somebody answers a question about multiple testing by quoting a figure off this ladder. Which pair should be checked first, and what settles the matter without the ladder at all?
The failure: quoting 48.67 per cent as though it were an all clear
An analyst walks the thirteen settings of the Ashwin rule across the earlier half of the six year record and reports whichever of them read best. Somebody asks the right question: what does running thirteen tests do to the result? The analyst goes to the ladder, reads that thirteen tests give 48.67 per cent, and answers that since 48.67 is under one half the result is still more likely than not to be real. The arithmetic in the reply is correct. Three things about it are wrong, and only one of them is the obvious one.
First, it answers a different question from the one asked. 48.67 per cent is the chance that at least one of thirteen tests fires when there is nothing at all in the record. The figure is not the chance that this one reported reading is false. The two quantities are different and not interchangeable, however similar the sentences describing them sound.
Second, the thirteen tests are not independent. The thirteen settings read the same 29 scoreable months, one whole number step apart on the same grid. The ladder is built on tests that do not lean on one another, and these lean on one another about as hard as tests can. The figure does not apply to them at all, in either direction.
Third, and worst, the reported reading was the best of the thirteen. The best of a set is exactly the object the arithmetic was warning about. Using a figure derived from the count of tests to reassure yourself about the winner of that same count of tests is using the alarm as evidence that nothing is burning.
The cost is not that a number was misused in a meeting. The cost is that a figure whose only job was to raise a question has been used to close it, and everyone in the room now believes the question has been dealt with. The fix needs none of this arithmetic: the rule is read on a stretch of the record that was held back before any setting was chosen, and what it called there is reported. The held back stretch answers the question the ladder can only bound.
Covered elsewhere. The claim a threshold of 5 per cent makes about a single test is settled earlier in these notes, and so is independence between two things. Holding a search down in practice is settled elsewhere in these notes. Choosing among several defensible ways of analysing one result, and choosing the question itself after looking at the record, are covered on their own, and they are two different faults rather than one. Placing an order, paying a cost, holding a position and judging whether any tested rule deserves anybody's attention all sit elsewhere.
Where do the figures above come from?
Every quantity above comes out of three things: a threshold of 5 per cent, a count of tests, and one invented record built earlier in these notes. Nothing else goes into any of them.
| Quantity used | Where it sits | How to rebuild it |
|---|---|---|
| The ladder, from one test to five hundred | Computed here from a threshold of 5 per cent and a count, and from nothing else at all | Raise 0.95 to the count and take it away from one. Twenty tests hand back 64.15 per cent. |
| The threshold one single test is allowed | The same equation with the other unknown solved for, derived above rather than borrowed | Take the thirteenth root of 0.95 and subtract it from one. The answer reads 0.3938 per cent. |
| The six year record of the Nakshatra unit | 72 invented monthly observations, settled earlier in these notes and reused here without one figure being altered | Rebuild it and count: 65 of the 72 months can be scored, 29 of them in the first three years. |
| The Ashwin rule and its thirteen settings | Invented, and written down in full before any count here was taken | Score the rule setting by setting on the scoreable months and the counts follow. |
| The two searches on this record | Run in full by the arithmetic check kept beside these notes, and that check refuses to finish if a count moves | Twelve of twenty analyses clear the threshold, and ninety four of 4,096 calendar rules do. |
The Nakshatra unit, the six year record and the Ashwin rule are invented.
Educational material. Not advice on any investment, tax, budget or market position.
