P-Hacking: How Analysis Choices Manufacture Significance
P-hacking is choosing among defensible ways to analyse one result until the reported figure clears a threshold. Here the question is fixed first: does the Ashwin rule call direction better than a rule holding no information at all? Twenty defensible analyses of one invented record answer that question anywhere from 50.00 per cent to 89.66 per cent, and twelve of the twenty clear 5 per cent.
The Nakshatra unit, an invented price, has a six year record of 72 monthly changes running from January 2019 to December 2024, built out of three stated parts. The Ashwin rule is one sentence, set down in advance so that it could be tested: read this month's change, call next month up when that change sits above a fixed level, and call it down otherwise. Every count below follows from the record and the rule by arithmetic alone. Anybody holding the three parts can rebuild all twenty answers to the same decimal place.
What is being chosen here, and what is being held still?
Nothing about p-hacking is visible until the fixed part of the exercise has been pinned down. The fixed part comes first. The question is written out before anything is run, and it stays exactly as written: does the Ashwin rule, reading at a level of minus 1.00 per cent, call the direction of the next month better than a rule that knows nothing and simply guesses? The rule does not move. The six year record does not move. The level of minus 1.00 per cent does not move.
Only the way the question gets answered moves, and every single one of the moves available is defensible on its own. That last clause is the whole difficulty. Nobody involved has to do anything they could not justify out loud, in a meeting, to somebody sceptical. Each choice has a straight faced argument behind it. The trouble is that there are twenty combinations of those arguments, they hand back twenty different answers, and only one of the twenty ends up in the write-up that gets read.
Now the version off the street. Somebody asks a shopkeeper whether Tuesdays are busy. There is no dishonest answer available and there are several honest ones. Tuesdays this month. Tuesdays this quarter. Tuesdays all year. Counted by people through the door, or counted by what the till took. Counted with the festival week left in, or with it taken out because it distorts things. Every one of those is a real answer to a real question. The shopkeeper tends to work through them in some order and stop at the first one that agrees with what they already believed about Tuesdays. Nobody lied. Nobody even noticed choosing.
A written up sheet beside the answers it left out is the shape of the whole subject. Look at the sheet on the left and try to find the lie. There is not one. The record is the record, the rule is the rule, the count of 26 out of 29 is arithmetic, and the studyA written up account of research that states a question, names the record it used, and prints the figure it arrived at. Reading and judging one is treated in its own right elsewhere. that prints it has described exactly what it did. Then look at the dark panel and count. Nineteen other answers existed, all to the same fixed question, all from the same record and the same rule, and the sheet has no room for any of them.
In the situation described here, what is being held still and what is being allowed to move?
What does the menu of choices actually look like?
Most writing on this subject describes the menu instead of printing it, and that is exactly what makes the idea feel like an accusation rather than an arithmetic fact. So here is the whole menu, in words, with nothing held back. There are three choices and no more.
The first is the stretchOne unbroken slice of the record, taken as though it were the whole thing. Months inside the slice get counted and months outside it are ignored entirely., which is the run of months the counting happens on. Five of them are on offer: all seventy two months, the first thirty six, the last thirty six, the first forty eight, and the last forty eight. Every one has an argument. The whole record uses everything available. The first three years are the earlier and quieter part. The last three years are the more recent part, and recency is a real argument. Forty eight months is two thirds of the record. Keeping two thirds and dropping either the oldest or the newest slice is defensible on its own.
The second is the counting conventionA counting decision, settled and recorded ahead of time. Two people working apart must not hand back two different answers to one question.. Six of the seventy two months ended on exactly the mark they opened on, and those six need a decision. Such months can either be excluded, on the argument that a call of up or down has nothing to be judged against when the month went nowhere, or counted as a rise, on the argument that not falling is a kind of rising and dropping months makes the record shorter. Both arguments are perfectly serious and people hold both.
The third is the signalWhatever a rule looks at in the instant before it decides. Assembled purely from months already behind it. Nothing from further ahead can slip in., which is what the rule reads before it calls. The signal is either the change in the month just gone, taken alone, or the three most recent months averaged together. Reading one month is simpler. Averaging three is steadier and less shaken by a single odd month. Nobody would blink at either.
Five stretches multiplied by two conventions multiplied by two signals is twenty paths, and every one of the twenty is a complete, defensible analysis of the same fixed question. Not twenty attempts, not twenty rounds of tinkering. Twenty finished answers, any one of which a careful person could have arrived at first and stopped at.
Suppose a fourth choice were added to the menu and it had two settings, like the last two. How many paths would the menu produce then, and which way could the highest reading on offer move?
How far apart do the twenty answers sit?
Run all twenty and print them. The best path reports 89.66 per cent, from 26 of 29 months, taking the first three years with unchanged months excluded and last month as the signal. The worst reports 50.00 per cent, from 18 of 36, taking the last three years with last month as the signal. The two are 39.6552 points apart, and there is one record and one rule underneath both of them.
Notice what that gap is not. There is only one rule, so the gap is not the difference between a good rule and a bad one. There is only one record, so the gap is not the difference between two records. Both readings come out of choices anybody could defend, so the gap is not the difference between a careful analyst and a careless one. The gap is what the freedom to choose is worth, measured in points, on this record.
Something quieter is worth catching before the table. Four pairs of paths finish on exactly the same count of matched months out of exactly the same count of scoreable months, so the twenty paths land on only sixteen distinct counts. Two more paths that do not share a count still share a reading, so turning the counts into percentages leaves only fifteen distinct readings.
| Rank | Matched calls | Reading | Own figure | The three choices, in order |
|---|---|---|---|---|
| 1 | 26 of 29 | 89.66 | 0.0008 | first thirty six, unchanged excluded, last month |
| 2 | 29 of 35 | 82.86 | 0.0058 | first thirty six, unchanged counted as a rise, last month |
| 3 | 32 of 41 | 78.05 | 0.0215 | first forty eight, unchanged excluded, last month |
| 4 | 35 of 47 | 74.47 | 0.0544 | first forty eight, unchanged counted as a rise, last month |
| 5 | 19 of 27 | 70.37 | 2.6119 | first thirty six, unchanged excluded, average of the last three |
| 6 | 44 of 65 | 67.69 | 0.2959 | all seventy two, unchanged excluded, last month |
| 7 | 22 of 33 | 66.67 | 4.0072 | first thirty six, unchanged counted as a rise, average of the last three |
| 8 | 26 of 39 | 66.67 | 2.6626 | first forty eight, unchanged excluded, average of the last three |
| 9 | 47 of 71 | 66.20 | 0.4277 | all seventy two, unchanged counted as a rise, last month |
| 10 | 29 of 45 | 64.44 | 3.6227 | first forty eight, unchanged counted as a rise, average of the last three |
| 11 | 39 of 63 | 61.90 | 3.8463 | all seventy two, unchanged excluded, average of the last three |
| 12 | 42 of 69 | 60.87 | 4.5593 | all seventy two, unchanged counted as a rise, average of the last three |
| 13 | 26 of 45 | 57.78 | 18.5649 | last forty eight, unchanged excluded, last month |
| 14 | 26 of 45 | 57.78 | 18.5649 | last forty eight, unchanged excluded, average of the last three |
| 15 | 27 of 48 | 56.25 | 23.5440 | last forty eight, unchanged counted as a rise, last month |
| 16 | 27 of 48 | 56.25 | 23.5440 | last forty eight, unchanged counted as a rise, average of the last three |
| 17 | 20 of 36 | 55.56 | 30.8860 | last thirty six, unchanged excluded, average of the last three |
| 18 | 20 of 36 | 55.56 | 30.8860 | last thirty six, unchanged counted as a rise, average of the last three |
| 19 | 18 of 36 | 50.00 | 56.6030 | last thirty six, unchanged excluded, last month |
| 20 | 18 of 36 | 50.00 | 56.6030 | last thirty six, unchanged counted as a rise, last month |
The fourth column, the own figure, is the chance that a rule holding no information at all would match that many months or more out of that many. The own figure is worked out for each path from that path's own two numbers and nothing else, and a smaller own figure means a result harder to put down to luck. Every percentage in the table is either a share of months called correctly or a chance of that kind, so none of them is a return and none of them is anybody's result.
The twenty paths give sixteen distinct counts but only fifteen distinct readings. Where did the extra collapse come from?
How many of the twenty clear the bar, and what does the middle one read?
A reading on its own says nothing until the count of months it was computed on is known, so each path carries its own figure beside it. Set that own figure against the thresholdThe mark a figure has to get past before anybody treats a result as worth reporting. The promise the 5 per cent mark makes for one lone test is settled elsewhere and assumed here. of 5 per cent and a verdict follows for each path, one at a time.
Twelve of the twenty land under 5 per cent, running from 0.0008 per cent on the best path down to 4.5593 per cent on the twelfth. The middle figure of the twenty reads 3.9267 per cent, and it clears too. The twelve that clear are exactly the top twelve by reading, and the eight that miss are exactly the bottom eight, so the line falls in one clean place on this record.
Now the consequence. Somebody who picked one of these twenty paths before looking at any of them, wrote it down, ran it once and reported it, would clear the threshold twelve times out of twenty. Twelve times in twenty is more often than not. Somebody doing that would not have to choose anything dishonestly, would not have to run the other nineteen, and would never know the other nineteen existed. The middle of the pack clears. On this record, a threshold of 5 per cent is not doing the work a reader assumes it is doing.
Twelve of the twenty paths clear the threshold, and the middle figure of the twenty clears it too. What does that mean for somebody who picked one path in advance, ran only that one, and reported it?
There is a second thing hiding in that table, and it is the one most readers walk past. A reading of 70.37 per cent sounds plainly stronger than a reading of 67.69 per cent. But 70.37 per cent is 19 months out of 27, and 67.69 per cent is 44 months out of 65. The second was computed on more than twice as many months, and that matters enormously to how surprising it is. The bar a reading has to clear depends on the denominatorWhatever number a share is being taken out of. Swap it and an identical top half says something else: twenty six months out of twenty nine is nowhere near twenty six out of forty five. the reading was computed on, so the bar is not one fixed level.
Worth settling before the control moves at all: stepping through the twenty paths in order of the reading each reports, from 89.66 per cent down to 50.00 per cent, does every path's own figure grow weaker as the walk goes on?
Step through the twenty paths one at a time
One control, and it does nothing except choose which of the twenty paths is the one being reported. The question, the six year record and the Ashwin rule are identical at every setting, so every change on screen is an analysis choice and nothing else. Three views redraw together. The menu lights up the three choices that make the current path. The bar shows the reading it reports, against faint marks for the other nineteen. The lower scale slides a marker against the fixed line at 5 per cent, and that scale is squeezed at the right hand end so that figures a thousand times apart can share it. The control starts on path one, the sheet drawn at the top of this guide.
Educational illustration. The question and the level of minus 1.00 per cent are pinned down before any path runs, so the panel moves analysis choices and nothing besides. The twenty paths read one record and most of them overlap heavily, so they are not independent of one another and the running count of paths that clear is not a chance of anything.
Why does the many tests figure not settle this table?
A reader who has met the arithmetic of many tests will reach for it here, and the reach is understandable. Twenty tests that are independentTwo tests stand independent when the result of the first leaves the expectation for the second exactly where it was. Covered separately and assumed here. of one another, each judged at a threshold of 5 per cent, give a 64.15 per cent chance that at least one of them clears by luck alone. Twenty paths sit in the table above. The two twenties look like the same twenty.
The two twenties are not the same, and the reason is one word: these twenty paths are not independent of one another. They all read one record. Ranks one and two differ only in whether six months are dropped, so they share twenty nine of their months exactly. Ranks thirteen and fourteen share all forty five of theirs and hand back the identical count. Ranks six and nine read every month in the record. When two tests overlap that heavily, one of them clearing says a great deal about whether the other will. Independence rules out exactly that.
So 64.15 per cent is worth carrying as a rough sense of how much a count of twenty matters, and it is not the chance attached to this table. An account that prints the two side by side has taken a figure computed for one quantity and stuck it on another. The honest thing to say about this table is the thing the table already shows: twelve of these twenty cleared and the middle one cleared. No single number summarises how likely that was, and working one out would need the overlap between every pair of the twenty measured first.
Twenty tests that are independent of one another give a 64.15 per cent chance that at least one clears by luck. Why can that figure not be attached to this table of twenty paths?
Which of the agreements between paths mean anything?
Four pairs of paths land on exactly the same count. A reader will notice all four, and leaving them unlabelled would teach a coincidence as though it were a finding. Here they are, sorted, with the reason written out beside each.
Two of the four pairs agree because of something in the record, and two agree for no reason at all. The something is this: every one of the six months where the Nakshatra unit ended exactly level falls inside the first three years. The last three years hold none. So on any stretch made up only of the last three years, the two counting conventions are being asked what to do about a kind of month that is not there. The two conventions cannot possibly differ there, and both pairs on that stretch agree by arithmetic. Nothing has been discovered.
The other two pairs sit on the last four years, and there the two signals happen to call 26 of 45 as each other and 27 of 48 as each other. Reading last month and reading the average of the last three are genuinely different instructions, they disagree elsewhere in the table, and here they happen to land together. The agreement is a coincidence, it means nothing, and putting the two kinds in one sentence without saying which is which is how a reader ends up believing the signal does not matter.
One more equality is worth labelling for the same reason. Ranks seven and eight both read 66.67 per cent, but they do not share a count: one is 22 of 33 and the other is 26 of 39. Both reduce to exactly two thirds. The match is arithmetic about fractions and carries no news whatever about the Ashwin rule. The two paths have quite different own figures, 4.0072 per cent against 2.6626 per cent, and the second was computed on more months.
Two pairs of paths agree because of something in the record and two agree for no reason. Without being told which is which, how can the two kinds be told apart?
The failure: ranking the paths by the wrong column
A reviewer is handed the table of twenty and does the sensible looking thing. Sorts it by the reading each path reports, works down from 89.66 per cent, and concludes that the reported figure falls steadily as the analysis gets more honest and less flattering. The sorted column reads like a clean story about a fitted rule losing its shine.
The two columns do not move together. Walk the twenty in order of the reading and the path's own figure moves the other way at three of the nineteen steps. A smaller share of a larger count of months can be harder to put down to luck than a larger share of a smaller one. The clearest case is the step from rank five to rank six. Rank five is 19 of 27, reading 70.37 per cent, own figure 2.6119 per cent. Rank six is 44 of 65, reading a lower 67.69 per cent, and its own figure of 0.2959 per cent is nearly nine times stronger. The reviewer who sorts by the reading prefers a figure computed on 27 months over one computed on 65.
Sort the same twenty by their own figures instead and six of the twenty change places. Rank seven moves four places down to eleventh. Rank nine climbs three places to sixth. The order is genuinely different, and which order is used decides which path gets handed back as the one to keep.
The fix is free and it has to become reflex rather than something recalled. The count of months is taken in at the same glance as the reading, every single time. A reading stripped of its denominator has hidden the one thing that decides what it is worth, so a bare percentage should never sit unaccompanied in a table under judgement.
What actually stops this, in practice?
Four things, and they are not equally hard. All four are open to a person reviewing somebody else's work, and open to somebody producing the work before there is anything to review.
Write the three choices down before running any of them. One line, before the first count is taken: this stretch, this convention, this signal, and this is the path being reported. Of the four preventions this is the only one that costs nothing, and it is the one nobody does. It costs nothing because writing down a choice that was going to be made anyway takes a sentence. At the moment of writing there is no reason to bother, and by the time there is a reason the line is too late to write.
Report the pre-committed path, and print the other nineteen beside it if they were run. Printing the nineteen costs a table and is worth the table. A reader looking at all twenty can see the spread of 39.6552 points and judge the reported figure against it. A reader given one line has nothing to judge it against.
Say how many paths were run altogether, whichever of them is printed. A count of attempts describes the search rather than any single result, and a reader has no other route to it, so the count earns its place even when the twenty are not shown.
And score the rule over months that were kept aside and never looked at. Months held back settle the question outright, with no argument about analysis choices at all. The fourth prevention is also the expensive one: those months have to exist in the first place, and the moment somebody looks at them they are spent.
Here is how that lands for somebody on the receiving end. Imagine a household deciding whether a neighbourhood tuition centre is any good, and being handed one number: nine of ten students improved. The right question is not whether nine of ten is impressive. The right question is how many groups of ten were counted before that group of ten was the one mentioned, and whether the group was named before the counting or after. Somebody at a lending desk going through a borrower's own account of how a rule performed, an analyst working through a research note, and a household weighing up a claim about a tuition centre are all standing in the same position, and the reported figure by itself answers none of them.
A colleague offers to list every analysis they ran alongside the one they reported. Does that fix the problem, and which of the four preventions still has to happen?
What is covered elsewhere
Picking the question itself once the record has already been seen is a different fault carrying a different tell, and it is taken up separately. Everything above holds the question nailed down and lets only the analysis move. The two subjects divide along that line. Running many tests and what that arithmetic gives, keeping a slice of months aside, and setting a run down so that a second person can repeat it were all handled earlier, each in its own place. Placing an order, paying a cost or moving money belongs elsewhere again.
Where do the twenty figures above come from?
Every one of the twenty was produced by arithmetic that can be repeated. An empty list of authorities is the honest answer here rather than a missing one. A recipe short enough to retype stands in place of one.
| Source | Document | Site |
|---|---|---|
| The six year record | Seventy two invented monthly changes of the Nakshatra unit, running from January 2019 to December 2024, assembled from three stated parts | None. It was built rather than downloaded |
| The Ashwin rule | One written sentence and one level of minus 1.00 per cent, both fixed before any of the twenty paths was run | None. It was written for these notes |
| The menu of twenty paths | Five stretches, two counting conventions and two signals, each choice listed in words rather than described in general terms | None. The list is the whole of it |
| The own figure beside each path | The share of coin tosses that would match that many months or more out of that many, worked out as an exact fraction rather than a decimal | None. It is a sum of whole numbers |
| Any outside authority | Nothing at all. Every figure above rests on the record, the rule and the menu of twenty paths | None. Nothing outside the three stated parts was used |
The Nakshatra unit, its six year record and the Ashwin rule are invented.
Educational material. Not advice on any investment, tax, budget or market position.
