Data Snooping: Finding Patterns Because You Looked Too Hard
Data snooping has an uncomfortable shape. Every step in a snooped result is arithmetic that checks out. A search runs, a rule comes back at the top of the pile, the figure printed beside that rule is computed correctly, and the whole thing is written up by somebody who never once shaded a number. And the result is worthless, in a way that no amount of care with the arithmetic could have prevented. The fault is not in any calculation anywhere. The question was picked out of the record instead of being brought to it. The argument runs twice over one invented record, once where there is genuinely something to find and once where there provably is not.
Two records, not one. The six year record of the Nakshatra unit is 72 dated monthly readings, invented and settled elsewhere in these notes: an average month of 1.00 per cent, an opening mark of Rs 100.00/- and a closing price of Rs 187.4539/-. The unadjusted record has a repeating calendar pattern built into it on purpose. The calendar adjusted record is the same 72 readings with each calendar month's own factor subtracted, so every calendar month of what is left averages exactly 1.00 per cent. The second record was built that way, so the absence of a calendar difference in it is certain rather than likely.
A bar of 5 per cent. Settled under inference and used here without being explained again. Beside every score below sits that score's own figure: the chance that calls carrying no information at all would have got that many months right or better, out of the same count of months. Under 5 per cent and the result is treated as worth reporting.
Counting, and one pairing argument. Every quantity below is months tallied, rules tallied, or one tally set over another.
What is data snooping, and how is it different from choosing among analyses?
Both faults produce a figure that is too good, and the two get muddled constantly, so the separation has to be made on something sharper than the flavour of the mistake. Make it on a single question: what was fixed before the record was opened?
In the fault covered separately, the question was fixed first. Somebody wants to know whether one particular rule calls direction better than a call with no information behind it. The question exists in writing before anybody opens anything. Only the route to the answer moves after that: which stretchSome months in a row, cut from a longer record and then read as though they were a record of their own. The first three years here are one; the last three are another. of the record to read, which conventionA counting decision that somebody made and wrote down. Nothing in a record forces one convention over another, and swapping one for another moves the figures without anything having happened. to use for a month that did not move, which signal to read. Twenty defensible routes, and the flattering one gets reported.
Data snooping does something else entirely. The record is opened first, and the question is read off it. Nobody arrives wondering whether September tends to rise; they look at six years of months, see that September rose in three years out of four, and then produce the question about September. The hypothesis is manufactured by the record and afterwards presented as though somebody had thought of it beforehand.
The difference is easy to feel on a street. A shopkeeper who wonders in March whether Tuesdays are her busiest day, writes that down, and then goes back through the takings, is doing the first thing. A shopkeeper who scans a year of takings, notices that Tuesdays look strong, and announces a Tuesday effect, is doing the second. The takings are identical. The arithmetic is identical. Only the order of the two acts changed, and the order changed the whole worth of the finding.
The order of the two acts is also why the two faults need different repairs, and this is the part most readers get wrong. Listing every analysis that was run fixes the first fault completely: the reader sees all twenty routes, sees which one was reported, and discounts accordingly. The single analysis that was run is entirely honest and the list has one entry on it, so applied to data snooping that same list does nothing at all. The one thing chosen out of many was never an analysis. A question was.
In data snooping, what gets chosen after the record has been seen, and what got chosen after it in the fault covered separately?
What turns up when every calendar rule on this record is searched?
Let a rule be as simple as it is possible to be. For each of the twelve calendar months it says one thing and never changes its mind: January always up, or January always down. February always up, or February always down. All the way to December. Such a rule reads nothing, computes nothing and adapts to nothing. Each rule is twelve fixed opinions in a row.
Two choices for each of twelve months means two multiplied by itself twelve times, and that comes to 4,096 rules. The 4,096 are the entire set, and the set is small enough to write out completely. Sixty six months can be scored. Six of the 72 closed exactly where they opened, and a rule saying up or down has no right answer available on a month that did neither, so the convention throughout these notes is to drop it. So 66 is the denominatorThe bottom of a fraction. In 51 out of 66, the 66 is the denominator. Two searches with different bottoms produce shares that cannot be laid straight against one another. under every reading in this section.
The best of the 4,096 rules calls 51 of the 66 months right, a reading of 77.27 per cent. The best rule calls January, February, March, September, October, November and December up, and May, June, July and August down. April splits two and two across the six years, so whichever way a rule calls April it collects two months. Exactly two of the 4,096 rules therefore tie for the top place rather than one.
Read that paragraph again and notice what is missing from it. Nobody had a theory about September. Nobody predicted that May would be weak. The twelve calls were not derived from anything; they were read off the record, month by month, by asking which way each month happened to go more often. The rule is not a finding about the record. The rule is a description of the record, written in the grammar of a rule.
Two calls for each of twelve months gives 4,096 rules. Suppose the same search ran over twenty four fortnights instead of twelve months. Roughly how many rules would there be, and what would that do to the best score to be expected?
Why is the average score across all 4,096 rules exactly half?
Something in that set of 4,096 is fixed before a single month is counted, and seeing it is what stops the number 51 from ever impressing anybody again.
Take any rule in the set. Now take the rule that reverses every one of its twelve calls: where the first says January up, the second says January down, and so on down the year. The set holds every combination, so the reversed rule is in it too. Every month that the first rule calls correctly, the second calls incorrectly, and the other way round. So the two scores of any such pair add to exactly 66, the count of scoreable months, and they do so without anybody looking at what the record contains.
The 4,096 rules split cleanly into 2,048 pairs of that kind. Each pair contributes 66, so the whole set contributes 2,048 multiplied by 66, and dividing that by 4,096 gives 33. The average score across every rule in the search is exactly 33 of 66, exactly half, and the same argument on the adjusted record gives exactly 31 of 62. The pairing is arithmetic. The average is not a property of the Nakshatra unit, it is not a coincidence about these six years, and it would come out the same on any record substituted for this one.
Now sit with the consequence. A search of this shape is certain in advance to contain a best and a worst that are mirror images of each other, sitting the same distance either side of the middle. Finding an impressive best is not news; it is the shape of the search. On this record the best is 51 and the worst is 15, and the two of them add to 66 as they must. The existence of a good looking rule at the top of a search shows that a search was run, and nothing else at all.
One detail stops the picture being read as a tidy hill. The middle of that shape is 33 and the average is 33, but 33 is not the commonest score. 260 rules land on 33, 270 land on 32 and another 270 on 34. Symmetry about the middle and a peak at the middle are two different claims, and only the first one is true here.
Why is the average score across all 4,096 rules exactly half of the scoreable months?
The panel below runs the same search on a record that has no calendar difference in it at all. How many of the 4,096 rules does it take before the best one found so far clears a 5 per cent bar?
Watch the best rule appear out of a record that has nothing in it
Drag the control to search more of the 4,096 rules, taken in a stated order, and keep the best score found so far. The panel opens on the whole search of the calendar adjusted record, the worked instance above. Then switch records and run it again.
Educational illustration. The calendar adjusted record has every calendar month averaging exactly 1.00 per cent by construction, so any best rule found in it was manufactured by the search rather than discovered in the record. The 4,096 rules are searched in a stated order and they are not independentTwo tests are independent when knowing the answer to one says nothing whatever about the answer to the other. The 4,096 rules here are not: they all read the same 72 months. of one another, so no count of tests governs the readings below. Each readout is a tally of months, a tally of rules, or one tally divided by another.
What does the same search return on a record with nothing in it?
Take the six year record and remove the calendar pattern: subtract from every reading the average of its own calendar month. Every calendar month of what is left averages exactly 1.00 per cent, so no calendar difference remains in it whatever. Not a small one. Not a hard to find one. None, by construction, and it is known the way the fairness of a coin is known to whoever machined it. Sixty two months can be scored on that record.
Run the identical search over the identical 4,096 rules. The best of them calls 41 of 62 months right, a reading of 66.13 per cent, and the figure printed beside it is 0.7567 per cent, comfortably under the 5 per cent bar. A researcher who ran only that search, on only that record, would put the result in writing without a moment of doubt, and every number in the write up would be correct.
Set the two side by side. On the record as it stands, 51 of 66 and 77.27 per cent. On the record with nothing in it, 41 of 62 and 66.13 per cent. Both clear the bar. Both come with correct arithmetic. One of them was found in a record that genuinely has a calendar pattern built into it, and one of them was found in a record that provably has not, and the best figure alone cannot say which is which.
A milder version of the claim is easy to hear, so the claim is worth stating exactly. The claim does not say the first search found nothing. The calendar pattern in the unadjusted record is real, and it was put there on purpose. The claim says something worse: the output of the search does not carry the information needed to tell the two situations apart. The number reads the same in both worlds.
The best rule on the record with nothing in it comes with an own figure of 0.7567 per cent. Is that figure wrong?
How many rules clear the bar on a record with nothing in it?
The best rule is one rule. The more useful count is how many of the whole set got past the bar, so here are both. On the calendar adjusted record, the record with nothing in it, 94 of the 4,096 rules clear 5 per cent. On the record as it stands, 430 do. Any rule scoring 38 or more of the 62 clears it on the adjusted record, and any rule scoring 41 or more of the 66 clears it on the record as it stands.
Now do the division that everybody reaches for and watch it fail. A thresholdA bar fixed in advance. A finding has to get past it before anybody treats the finding as worth writing up. It sits at 5 per cent throughout this guide. of 5 per cent invites the expectation that about one test in twenty clears it by luck. Across 4,096 tests that would be 204.8 of them. The adjusted record produced 94, well under that. The record as it stands produced 430, well over it. The count of tests does not even show which direction the error runs in, let alone its size.
The reason sits in the assumption underneath the one in twenty arithmetic. The assumption is covered separately and is worth naming again here. The one in twenty arithmetic is for tests that are independent of one another, meaning that knowing how one turned out says nothing about how the next will. The 4,096 calendar rules are nothing of the sort. The rules all read the same 72 months, and two rules that differ only in how they call October agree about everything else. So the count of rules searched is not a quantity that can be pushed through a formula. The count is a quantity to be reported, leaving the reader to do what they can with it.
94 of the 4,096 rules clear the bar on the record with nothing in it, and 430 clear it on the record as it stands. Why can neither count be worked out from the count of rules searched?
One rule was picked out of a record with a real calendar pattern in it, and one out of a record with none, and both look impressive on the stretch that chose them. Which single figure separates the two?
Which single check tells the two searches apart?
There is one, and it is not the best figure and it is not the own figure. Pick the calendar rule using the first three years only, then read that rule on the last three years. The last three years took no part in choosing it.
On the record as it stands, the rule picked out of the first three years gets every one of the 30 months it can be scored on there. A perfect 100.00 per cent, and it is not luck: over those three years every calendar month moved the same way each time it appeared, so a rule built by copying the record cannot miss. Read on the last three years, the same twelve calls land 21 of 36, or 58.33 per cent, against a baselineThe reading that comes from calling every month the same way and reading nothing at all. Its only job is to sit beside a result so the result stops looking impressive on its own. of 20 of 36 for calling every month up. The rule bought exactly one extra month out of 36, and it fell 41.6667 points doing it.
On the calendar adjusted record, the rule chosen from the first three years calls 22 of 27, or 81.48 per cent. Read on the last three years it calls 18 of 35, or 51.43 per cent. And here is the detail that finishes the argument: that rule calls all twelve months up, so on the last three years it is making exactly the always up call and its 18 of 35 is precisely the baseline. Not close to the baseline. Identical to it. The two rules are the same rule, so the match is arithmetic rather than luck.
Both rules fall, and the only place in this entire guide where the two searches look different is that held back reading: 58.33 per cent against 51.43 per cent. Six and nine tenths of a point, on months nobody chose. Everything else about the two searches, the best score, the own figure, the count of rules, the impressive reading on the stretch that chose the rule, reads the same sort of number in both cases. And notice what the perfect 100.00 per cent bought as a warning: nothing. The perfect reading was the higher of the two in sample readings and it belonged to the record with the real pattern. A record with nothing in it produced 81.48 per cent on the same test, so a reader holding only that figure had no way to know.
A rule called 30 of 30 months right on the stretch that chose it. What does that perfect reading predict about the next three years, and what did it actually turn into?
What should be asked when a pattern arrives already found?
Most of the time the search belongs to somebody else. A pattern walks in with somebody, an analyst sends a note, a colleague has noticed something in a spreadsheet, and there is about a minute to decide how much weight it can carry. Five questions, and the first one does most of the work.
How many patterns of this shape could have been reported, whether or not anybody ran them? Not how many were run, and not how many the person remembers running. The set existed whether or not anybody enumerated it, so the count that matters is the size of the set the reported pattern was picked from. Somebody who eyeballs six years of a record and picks out September has effectively searched twelve months, and probably several other shapes besides, without ever writing a loop.
Then four more. When was this pattern first described, before or after this record was opened? On a record known to contain nothing, what does the same search return? The search above answers exactly that question. Are there months nobody used while the pattern was being found, and what does the pattern read across those? And the last question catches the most people: what would have been reported if the search had turned up nothing at all? If the honest answer is that nothing would have been written, what is in hand is the survivor of a process whose failures never become visible.
A lender does the same thing without calling it this. Somebody arrives saying that shops of a certain kind repay better, and the first question is not about the arithmetic; it is how many kinds of shop were compared before that one was named. A household does it too. If one relative reports that a particular scheme worked, the useful question is how many relatives tried something and how many of those would have mentioned it if it had not worked. In every version, the question is the same: what is the size of the set this result was picked out of?
Every figure above is a hit rateThe count of calls a rule got right, divided by the count of calls it was able to make. It is a share of months and never an amount of money., a count of months or a count of rules. A share of correct calls says nothing about the size of the moves those calls were made on, so a hit rate and a gain are separate quantities and neither follows from the other.
Somebody presents a pattern noticed in a record. Which single question comes first, and why is asking how many analyses were run the wrong question here?
Where this goes wrong, and why it is the hardest fault in the sequence
A researcher runs the calendar search on the calendar adjusted record. The best of the 4,096 rules comes back at 41 of 62, the figure beside it is 0.7567 per cent, and that clears the 5 per cent bar by a wide margin. The result gets written up. The write up names the record, states the rule, gives the count of months, gives the reading and gives the figure. Every one of those lines is correct.
Every calendar month of that record averages exactly 1.00 per cent by construction, so the absence of a calendar difference is known with certainty rather than suspected. So the finding is empty. And the figure of 0.7567 per cent is still correct: it is correct arithmetic about one rule read once, and it says nothing whatever about a rule that was chosen as the best of 4,096. Nothing in the write up is false and there is no line in it a careful reader could object to. Data snooping is the hardest fault taught here for exactly that reason.
The cost is that the reader has no way to discount the result. The one number that would have warned them is the count of rules searched, and that is precisely the number the write up does not contain. Two repairs, and they work together. Report the count of patterns that could have been chosen beside the one that was, and 0.7567 per cent then arrives with 4,096 attached to it. And read the chosen rule over months that had no say at all in choosing it. On this record that moves the reading from 66.13 per cent to 51.43 per cent and puts it exactly on the always up call.
Covered elsewhere. Choosing among several defensible analyses of one result came earlier and is a different fault needing a different repair, and the two are kept apart here on purpose. The arithmetic of many tests, holding a stretch back, rolling a test forward through time and writing a run down so somebody else can repeat it all came earlier too, and none of them is restated here.
The cost of running a rule, the route by which an instruction ever reaches an account, and whether any rule named here deserves to be run at all, are covered separately. The Ashwin rule and the calendar rules appear here only as things being tested.
What sits behind these numbers, and how can each be rebuilt?
Four ingredients, and every one of them is checkable by hand rather than takeable on trust. Searched results deserve suspicion, and so does a figure that cannot be rebuilt.
| What was used | Where it was settled | How it can be rebuilt | Checked on |
|---|---|---|---|
| The six year record of the Nakshatra unit, 72 dated monthly readings | The ordered data notes of this library | Lay the twelve calendar factors over a steady 1.00 per cent, add the six years of leftovers, then confirm the mean of 1.00 per cent, the spread of 5.00 per cent and the closing price of Rs 187.4539/- | 23 August 2026 |
| The calendar adjusted record, the same 72 readings with the calendar pattern taken out | The same ordered data notes | Subtract each calendar month's own factor from every reading in that month, then check that all twelve calendar months of what is left average exactly 1.00 per cent | 23 August 2026 |
| The calendar search, 4,096 rules and every score they land on | Computed here rather than quoted | Write out every up and down pattern across twelve months, score each pattern against the record, and confirm the average lands on exactly half the scoreable months | 23 August 2026 |
| The own figure printed beside each score | Computed here rather than quoted | Count how many of the equally likely up and down outcomes across the scoreable months reach that score or better, and divide by the total count of outcomes | 23 August 2026 |
The Nakshatra unit, the six year record, the calendar adjusted record and the Ashwin rule are invented.
Educational material. Not advice on any investment, tax, budget or market position.
