How to Prevent Backtest Overfitting Before You Start
What will be tried is decided before it is tried. The grid goes in writing, every attempt including the abandoned ones is counted, a stretch of months is kept back, the bar moves to match how many settings were tried, and the baseline is printed next to every figure. Across the six year record, thirteen attempts added 31.0345 points to the reading on months the search saw and took 11.1111 off the reading on months it did not.
Five steps, in that order, and every one of them happens before the first attempt. Running them first is the only thing separating the list from a description of what already went wrong. Each step below is priced on the same invented material: the six year record of the Nakshatra unit, seventy two monthly readings, and the Ashwin rule. The rule holds one number and reads the month just gone before calling the month ahead.
The Ashwin rule is small enough to search completely, thirteen settings and no more, so every figure below can be recounted by hand. The count of attempts is a fact about the search, not about the record, and it is the one fact that can be decided in advance.
What is actually being held down here?
A count. Not a rule, not a record, not a person, and not carelessness. Backtest overfitting arrives when the number of attempts made against one stretch of history climbs while the number of quantities inside the rule stays exactly where it was. In this guide the rule holds one number the entire way through, from the first attempt to the thirteenth. Nothing about it grew. The only thing that grew was how many times the same twenty nine months were asked.
Here is the everyday version, and it is worth holding on to because the arithmetic later is just this in numbers. A stall tries one recipe on forty customers and writes down what they said. A second stall tries thirteen recipes on the same forty customers and writes down the reaction to the best one. Both hand over a sentence about a favourite. Only one of those sentences was picked out of thirteen, and the customers were the same forty people either way. The second stall did more work, learned less about the world, and produced a better looking sentence, all at the same time.
The news in that is unexpectedly good. If overfitting were driven by talent or honesty or care, there would be nothing procedural to do about it. It is driven by none of those. Overfitting is a count, and a count is something that can be written down and held fixed. Nobody has to be clever. Somebody has to keep score.
The measure used throughout is a hit rateThe share of the months a rule called correctly out of the months it was able to call at all. It is a plain count divided by a plain count and says nothing about money, size or consequence.. A hit rate is the count of months called correctly over the count of months that could be called at all. Six months did not move, and the very first month has nothing before it to read, so sixty five of the seventy two months on the six year record can be scored. Twenty nine of those sixty five sit in the first three years and thirty six sit in the last three.
Backtest overfitting shows up as one count rising while another count stands still. Which pair is it?
Prevention one: why does the grid go on paper before the search starts?
The gridEvery value a search intends to work through, named in advance. It records what will be tried and says nothing whatever about how any of it might turn out. for the Ashwin rule holds thirteen settings, every whole number from minus 6.00 per cent through to 6.00 per cent, searched from the lowest upward. Writing that sentence down costs one line and takes fifteen seconds. Prevention one is nothing more than putting it in a dated note before anything is run.
The reason this works has nothing to do with discipline, and everything to do with what can be checked afterwards. A grid written down first is a commitment: thirteen, and reporting the best of forty afterwards visibly breaks that note. A grid written down afterwards is a description, and a description can always be adjusted to fit whatever was actually done. The two sentences look identical in print. Only one of them could have been wrong.
Think about a shopkeeper who says at the start of the week that he will judge a new supplier on the next twenty deliveries. Compare him with one who, at the end of the month, says he was judging the supplier on the deliveries that mattered. The second man has not lied. He has simply left himself nothing anybody can hold him to, including himself. Which twenty were the ones that mattered was settled after he saw them.
Prevention two: which attempts count, including the ones thrown away?
All of them. Thirteen settings is thirteen attempts, and the eight that were abandoned the moment they were read count exactly as much as the five that improved anything. Counting the discards is the step people skip, and they skip it because the discarded attempts feel like nothing happened. Nothing was written up, nothing was kept, nobody remembers them by Friday.
The abandoned attempts are precisely what turned the reported figure into the best of thirteen rather than the best of five. The claim is worth restating with the arithmetic alongside. The setting eventually held is a thresholdA cut-off a reading must get past before a rule will commit to one call rather than the other. Somebody chooses it; nothing in the months chooses it for them. of minus 1.00 per cent, and that threshold calls twenty six of the twenty nine early months correctly. The count of twenty six did not change because of the discards. The sentence anybody is entitled to write about it did change. Best of five and best of thirteen are different claims about the same twenty six, and only the second one is true.
So prevention two is a counter, started at nought before the first attempt and raised by one every time a setting is scored, whether or not the result was kept, written down, mentioned to anybody or looked at twice. A setting scored and instantly discarded still consumed one of the thirteen. ConventionA rule for the counting itself, settled up front so that two people looking at identical material land on identical numbers instead of two defensible different ones. matters here more than accuracy: whatever counts as an attempt is settled first and applied to every attempt.
What do the two readings do while the search is running?
The two readings move in opposite directions, and that divergence is what a search costs. Searching the thirteen settings in the stated order, from minus 6.00 per cent upward, and keeping after each one whichever setting has the most correct calls on the first three years so far, gives one number per attempt. The setting being held is then read quietly on the thirty six months of the last three years. The search has never been shown those months and never uses them to choose anything.
On the twenty nine months the search can see, the running best climbs 58.62, 58.62, 65.52, 75.86, 82.76 and 89.66 per cent, and then stops. Every attempt from the sixth to the thirteenth leaves it exactly where it was. On the thirty six months the search cannot see, that same held setting reads 61.11, 61.11, 52.78, 47.22, 47.22 and 50.00 per cent, and then it too stops. The rule held one number the whole way, and the search still bought 31.0345 points on the months it could see and gave back 11.1111 points on the months it could not.
The second column does not fall smoothly, and any chart that draws it as a tidy slope is a lie. Across the twelve steps from the first attempt to the thirteenth, the held back reading falls at two steps, rises at one, and does not move at all at the other nine. The held back reading is not a decay. The reading lurches when the search happens to swap settings and sits perfectly still the rest of the time. A figure that is nobody's target should behave in exactly that way.
The very first attempt is worth taking before anything else. The shape of the whole argument is visible there. At attempt one the held back reading of 61.11 per cent is the taller of the two, and the reading on the months the search could see, at 58.62 per cent, is the shorter. The two lines cross at the third attempt. Everything before that crossing is a search that has not yet cost anything. Everything after it is the two lines pulling apart, ending 39.6552 points from each other.
At the first attempt the held back reading is the taller of the two. What does that show about a single untried setting, and what does it stop being once the search continues?
One line of arithmetic hides inside that first row, and it is worth pulling out because it looks like a fluke and is not. The first setting on the grid is a threshold of minus 6.00 per cent. In the first three years of the six year record, the smallest month is minus 4.00 per cent, so no month ever falls to minus 6.00 per cent or below. The signalWhatever a rule looks at before committing itself. Here that is the change in the month which has just finished, already known and unalterable by the time any call is made. the rule reads is therefore always above the threshold, so at that setting the Ashwin rule calls every single month up.
The first setting on the grid is therefore the always up baseline wearing a threshold, and its 17 of 29 is the baseline's 17 of 29 by arithmetic rather than by coincidence. Do not carry that equality anywhere else. On the whole sixty five month record, six months do have a signal at or below minus 6.00 per cent, so the same setting reads 39 of 65 while calling every month up reads 37 of 65. Same rule, same threshold, different months, and the equality is gone.
Across those first three years, the lowest setting on the grid produces precisely the count that calling every month up produces. Is that a coincidence of this record?
One more way of looking at the same thirteen settings, and this one shows the whole field at once rather than the search path through it. Plot each setting twice: its reading on the twenty nine months across the bottom, and its reading on the thirty six held back months up the side. If the two readings agreed, every point would sit on one diagonal line. They do not. The points sag away from that line, and they sag furthest exactly where the fitting reading is highest.
The same thirteen attempts, watched twice
| Attempt | Setting tried | Setting held | On the 29 months seen | On the 36 months held back | Gap, points |
|---|---|---|---|---|---|
| 1 | minus 6.00 | minus 6.00 | 17 of 29, 58.62 per cent | 22 of 36, 61.11 per cent | 2.4904 the other way |
| 2 | minus 5.00 | minus 6.00 | 17 of 29, 58.62 per cent | 22 of 36, 61.11 per cent | 2.4904 the other way |
| 3 | minus 4.00 | minus 4.00 | 19 of 29, 65.52 per cent | 19 of 36, 52.78 per cent | 12.7395 |
| 4 | minus 3.00 | minus 3.00 | 22 of 29, 75.86 per cent | 17 of 36, 47.22 per cent | 28.6398 |
| 5 | minus 2.00 | minus 2.00 | 24 of 29, 82.76 per cent | 17 of 36, 47.22 per cent | 35.5364 |
| 6 | minus 1.00 | minus 1.00 | 26 of 29, 89.66 per cent | 18 of 36, 50.00 per cent | 39.6552 |
| 7 | 0.00 | minus 1.00 | 26 of 29, 89.66 per cent | 18 of 36, 50.00 per cent | 39.6552 |
| 8 | 1.00 | minus 1.00 | 26 of 29, 89.66 per cent | 18 of 36, 50.00 per cent | 39.6552 |
| 9 | 2.00 | minus 1.00 | 26 of 29, 89.66 per cent | 18 of 36, 50.00 per cent | 39.6552 |
| 10 | 3.00 | minus 1.00 | 26 of 29, 89.66 per cent | 18 of 36, 50.00 per cent | 39.6552 |
| 11 | 4.00 | minus 1.00 | 26 of 29, 89.66 per cent | 18 of 36, 50.00 per cent | 39.6552 |
| 12 | 5.00 | minus 1.00 | 26 of 29, 89.66 per cent | 18 of 36, 50.00 per cent | 39.6552 |
| 13 | 6.00 | minus 1.00 | 26 of 29, 89.66 per cent | 18 of 36, 50.00 per cent | 39.6552 |
Every count above comes from the six year record of the Nakshatra unit and the Ashwin rule, neither of them carrying a maintained series behind it, so there is no as of date to quote. Rebuilding the record from its three stated parts returns these counts or it does not, and that is the whole of the checking.
The control below makes thirteen attempts one at a time. Before it runs: does the reading on the months the search cannot see fall a little at every attempt?
Walk the thirteen attempts one at a time and watch the two readings come apart.
One control moves: how many attempts have been made so far. The grid is always searched in the same order, from minus 6.00 per cent upward, and after each attempt the search keeps whichever setting has called the most of the twenty nine early months correctly. The rule itself never changes and holds one number throughout. The two bars are that held setting read two ways, and the trace underneath shows the whole run with the current attempt ringed. At the opening setting thirteen attempts have been made, the setting held is a threshold of minus 1.00 per cent, the reading on the twenty nine months is 89.66 per cent and the reading on the thirty six months held back is 50.00 per cent.
Educational illustration, built entirely on invented figures. The grid is searched in the stated order rather than in whatever order a loop happens to run in, so this ladder is reproducible by hand. The thirty six held back months are never used to choose anything on this panel; they are only ever read. A new setting replaces the one being held only when it calls strictly more months correctly. The second attempt therefore changes nothing: minus 5.00 per cent also calls 17 of 29, and ties do not displace. Every quantity on the panel is a count of months rather than money.
Prevention three: what does keeping months back do that counting attempts cannot?
Keeping months back measures instead of arguing. Counting attempts records how hard the search looked, and that is a fact about the searcher's own behaviour. Keeping a stretchA run of consecutive months lifted out of a longer record and treated on its own, so that whatever was decided elsewhere in the record was decided without it. back records what the looking was worth, and that is a fact about the world. On the six year record the held back stretch is the last three years, thirty six scoreable months, and the number it produces at the end is 50.00 per cent.
The two preventions answer different questions and neither one replaces the other. A count of thirteen attempts, on its own, cannot show whether the setting kept is any good. The count never touches a month the search did not use. A held back reading, on its own, cannot show how hard the search worked. A single lucky guess and an exhausting hunt produce the same one number. With both in hand, two separate things can be said. With only one, the temptation is to say both anyway.
There is one way to lose prevention three without noticing, and it takes about four seconds. If the held back stretch is looked at, even once, even out of curiosity, even without changing anything, it is no longer held back. The reason is not moral. Once those months have been seen, every later choice is a choice made by somebody who has seen them, and there is no procedure that can subtract that. A stretch glanced at is a stretch already spent.
Counting attempts and keeping months back do different jobs. Which of the two is a measurement rather than an argument about the search?
Prevention four: how much stricter does the bar get once thirteen settings have been tried?
The arithmetic of the sequence can also be run backwards. Suppose the chance of at least one false positive across the whole run is to stay at 5 per cent, and suppose for a moment that the thirteen tests were independent of one another. Thirteen goes at a 5 per cent bar give a 48.6658 per cent chance of at least one false positive across the run, so a single test cannot be allowed a bar of 5 per cent. Each test has to clear a much tighter bar instead. Taking 0.95, raising it to the power of one thirteenth and subtracting from one gives 0.3938 per cent. Twenty tests would need 0.2561 per cent.
On the twenty nine early months, moving from a bar of 5 per cent to a bar of 0.3938 per cent moves what a setting has to do from 20 correct calls out of 29 to 23 out of 29. The tighter bar takes the settings on this grid that still clear it from six down to four. Nothing about any setting changed; what changed is the number of settings that were tried before one was picked out.
0.3938 per cent printed alone reads as exact and it is not, so say the caveat in the same breath as the figure. The thirteen settings are not thirteen independent tests. The thirteen settings read one record through thirteen thresholds a single step apart, and thresholds one step apart mostly make identical calls on the same months, so the thirteen results are heavily tangled together. The tangle makes the correction an approximation aimed at a quantity nobody has computed. The correction is better than doing nothing and it is not a measurement.
Moving the bar to 0.3938 per cent takes what a setting must do from 20 of 29 to 23 of 29. Why is that correction not exact?
Calling every month up reads 58.62 per cent on the first three years and 56.92 per cent on the whole record. Which one goes beside a figure computed on the first three years?
Prevention five: why does a baseline sit beside every figure?
Because a figure printed alone has to be looked up before it can be judged, and nobody looks it up. Put 89.66 per cent in front of a reader with nothing beside it and they have to know, from somewhere, what a rule reading nothing at all would have scored on those same months. On the first three years of the six year record that baselineThe reading something with no information at all would produce on the same material, kept beside a result so the result has something to be judged against. is 58.62 per cent, from calling every month up and never reading a thing. On the whole sixty five month record it is 56.92 per cent.
Three things have to match before a baseline is worth printing, and they are easy to get wrong. A baseline has to read no signal at all. The baseline has to share the denominatorWhatever a share is being measured out of. Change it and the identical top number becomes a different percentage, so two shares built on different bottoms cannot be set beside each other., and here that means the same twenty nine months rather than some other twenty nine. And it has to be counted under the same convention, so a month that did not move is dropped from both or from neither. A baseline computed on different months is not a comparison, it is a second unrelated number placed nearby.
Consider a stall holder who says her new signboard brought ninety customers last month. Ninety sounds like a lot until somebody asks what an ordinary month brings, and she says eighty five. Nothing about ninety was wrong. The ninety simply could not be read until the eighty five arrived, and she was the only person in the conversation who already knew it. Printing the baseline is nothing more than putting the eighty five on the same line as the ninety, so the reader is not relying on the writer to volunteer it.
What do the five look like as a list run before the first attempt?
Here is the whole thing as a lender, an analyst or anybody handed somebody else's study would actually run it. Five items, four of which take a single line each, and all five sitting before the search rather than after it.
- Grid written down and dated. Thirteen whole number settings, minus 6.00 per cent to 6.00 per cent, searched from the lowest upward. One line.
- Attempt counter started at nought. Raised by one for every setting scored, including every setting abandoned on sight. One line.
- Stretch named and set aside. The last three years, thirty six scoreable months, not opened until the search has finished. One line, and then the discipline of leaving it shut.
- Bar moved for the count. Thirteen settings, so 0.3938 per cent each rather than 5 per cent each, with the note that the thirteen are not independent. One line and one caveat.
- Baseline computed on the same months. Calling every month up on the same twenty nine months under the same convention gives 17 of 29, or 58.62 per cent. One line.
All five happen before the search, and that is the whole difference between a prevention and a description of what already went wrong. Run afterwards, items one, two and three are no longer available: the grid is now a recollection, the count of abandoned attempts is now a recollection, and any stretch that was to be kept back has already been looked at. Items four and five still work afterwards. Both are arithmetic on numbers that already exist. So the honest answer to a colleague who asks whether the list can be applied to last week's study is short. Two of the five, and the two weakest.
A colleague offers to run the five item list over a study that finished last week. Which item is already impossible?
The correction that was applied while the protection was not obtained
A team does everything on the list except item three. The team counts thirteen attempts honestly, moves the bar from 5 per cent to 0.3938 per cent exactly as prevention four says, finds that its best setting reads 0.0008 per cent and clears the tighter bar with room to spare, and writes the study up as corrected for the search. Every number in that paragraph is arithmetically right.
The thirteen settings were never thirteen independent tests, so the correction was applied and nothing was measured. Those settings are thirteen thresholds one step apart reading the same twenty nine months, and thresholds one step apart mostly make identical calls. The correction is an argument about how hard the team looked. The argument is a reasonable one. An argument is not a reading on anything.
The cost is precise and it is not the arithmetic. The team now believes a stretch was kept back when none was, so they will treat 89.66 per cent as a figure that survived a test. The one number that would have settled the question is the 50.00 per cent the same setting reads on the thirty six months nobody opened, and it is not in the study because nobody kept those months back. A correction is an argument about the search and a held back reading is a measurement, so the fix is to run prevention three as well as prevention four, never instead of it.
What is covered elsewhere
Settled before this and left out on purpose: what a backtest is and what backtest overfitting means, how a figure produced on history differs from one produced afterwards, and how to write a run up so a second person can repeat it. All three are assumed here rather than rebuilt.
Covered separately: keeping a stretch back in full, rolling that held back stretch forward through time, the arithmetic of many tests in full rather than the single correction used here, and the pair of habits that pick a question or an analysis once the months are already in view. One figure is borrowed from each of those subjects here.
Orders, costs, positions and money of any kind are covered separately, in their own reading.
Where did these counts come from, and what stands behind them?
Every figure here rests on arithmetic a reader can redo. No regulator, exchange, maintained series or published study, paper, fund or record of results stands behind any of it.
| Material used | What it actually is | How it can be checked |
|---|---|---|
| The six year record of the Nakshatra unit | Seventy two invented monthly readings, the end of January 2019 to the end of December 2024 | Rebuild it from its three stated parts and count the months again |
| The Ashwin rule and its thirteen settings | One threshold, thirteen whole number values from minus 6.00 per cent to 6.00 per cent | Score the rule month by month and compare the counts printed above |
| The counting convention | A month that did not move is dropped, and the first month has nothing before it to read | Sixty five months can be scored, twenty nine early and thirty six late |
| The attempts ladder, both columns | The grid walked from the lowest setting upward, keeping the best so far on the early months only | Six of the thirteen attempts change the held setting; the other seven change nothing |
| The tighter bar for thirteen tests | One power, worked here rather than taken from a named correction | Raise 0.95 to the power of one thirteenth and take the answer from one |
The Nakshatra unit, the six year record and the Ashwin rule are invented.
Educational material. Not advice on any investment, tax, budget or market position.
