Backtesting: Testing a Rule on History, and How a Strategy Fits the Past
Somebody hands over a figure. A rule was tried on six years of a record, and it called the direction right about two times in three. The figure is real arithmetic. Nobody made it up, nobody rounded it kindly, and a second count of it would land on exactly the same number. And it still might mean nothing at all. Backtesting lives entirely in the gap between those two sentences. One invented record and one invented rule are enough to walk the gap end to end, and every number in it can be rebuilt.
The six year record of the Nakshatra unit. 72 dated monthly observationsOne dated entry in a record. The six year record holds 72 of them, one for each month., invented and settled elsewhere: an average month of 1.00 per cent, a spreadHow far the months of a record typically sit from their own average. A wide spread means months land further out in both directions. of 5.00 per cent, an opening mark of Rs 100.00/- and a closing price of Rs 187.4539/-. The record already carries a drift, a repeating calendar pattern and a scatter that more than doubles partway through. Not one figure in it is changed here.
Fitting a rule on one stretch of a record and reading it on another. Splitting a record that way is settled elsewhere in these notes, and it is used below without being explained again.
Counting, and nothing harder. Every quantity below is a count of months, a count of attempts, or one count divided by another. There is no other arithmetic here.
What is a backtest, and what does it actually count?
A backtest takes a rule that can be stated in one sentence, walks it through a record of months that have already finished, and counts how often the rule would have been right. The counting is the whole of the operation. There is no forecasting in it, no judgement in it and nothing clever in it. A backtest is a counting machine, and the only thing it can count is what the record already contains.
Think about a samosa stall outside one office building. The owner writes down a rule at the end of the year: fry early whenever yesterday evening was busy. Then she goes back through forty evenings of her own notebook and marks, evening by evening, whether frying early would have been the right call. At the end she has a count. Forty evenings, thirty one of them marked. The count is a backtest, and a backtest is exactly as good as the notebook, the rule and the marking behind it.
Notice what she cannot get out of the notebook however hard she counts. She cannot get what those evenings would have done under different weather. Different weather is not in the notebook. She cannot get whether the rule will still hold next winter. Next winter is not in the notebook either. And she cannot get whether frying early was worth doing. A count of right calls is not a count of anything else. Everything a backtest cannot answer, it cannot answer for the same reason: a record holds what happened and nothing else.
A backtest produces a count. What exactly is being counted, and what can no amount of counting get out of the record?
Which of the 72 months can be scored at all?
Before a single call is made, there is a duller question to settle: which months are even eligible? The six year record of the Nakshatra unit holds 72 months. The record does not hold 72 scoreable months, and the difference is not a rounding detail. The scoreable count is the denominatorThe count sitting underneath a share, which decides what the share is a share of. Change it and the share changes without any new information arriving. under every reading below.
Two things take months out. January 2019 has nothing before it to read, so the rule cannot produce a call for it at all. And six months did not move: April 2019, September 2019, August 2020, March 2021, April 2021 and September 2021 each closed exactly where they opened. A rule that calls up or down has no correct answer available on a month that did neither. One month with no signal and six months with no direction leaves 65 scoreable months out of 72.
The treatment of those six months is a conventionA choice about how something gets counted, written down in advance so that somebody else can repeat the count and land in the same place., and the honest version of a convention is one written down before anything is counted rather than settled afterwards when the counts are already on screen. The convention here is that an unmoved month is excluded. An unmoved month could just as defensibly have been counted as a rise. Counting the six as rises would give a different denominator and a different reading. The point is not which choice is better but that the choice was made in advance and is printed where a reader can see it.
All six unmoved months sit in the first three years, and none at all sits in the last three. The clustering is not a detail either. Those six months mean the record does not split into two equal halves of 36 scoreable months. The record splits into 29 and 36, and every later comparison carries those two uneven counts underneath it.
The six year record holds 72 months but every reading below sits over 65. Where did the other seven go?
Every one of the six unmoved months landed in the first three years and not one landed in the last three. Accident, or something about that stretch?
What does the Ashwin rule read, and how many settings does it have?
Only one object is still missing, and it is a small one. The Ashwin rule runs like this. At the end of each month, read the change the Nakshatra unit made that month. If that reading sits above a thresholdThe number a reading has to clear before a rule calls one way rather than the other., call the next month up. Otherwise call it down. Nothing else belongs to the rule. The rule fits in one sentence because it has to: a rule that cannot be stated in one sentence is a rule that cannot be checked.
The number the rule reads before it calls is the signalThe number a rule reads before it makes a call, computed only from months that have already finished.. Here the signal is just the month that has finished, taken as it stands. Nothing is averaged, nothing is smoothed and nothing from later is folded in. Every number the rule reads on the last day of a month was already sitting there on the last day of that month. Being able to say that sentence about a rule matters more than it sounds.
The threshold has to be some number, and there is no principle that supplies one. So the rule comes with a gridThe stated list of settings a search walks through, fixed before the search starts, so that the attempts can be counted afterwards.: thirteen whole number settings running from minus 6.00 per cent up to 6.00 per cent. The grid is stated in advance. A stated grid is a countable grid, and a count of attempts is the single most useful thing anybody can say about a tested figure.
Running all thirteen over the 65 scoreable months of the six year record gives a ladder. The best setting is minus 1.00 per cent, calling 44 of 65 months, or 67.69 per cent, and the worst is 6.00 per cent, calling 28 of 65, or 43.08 per cent. One more figure has to sit beside that ladder. Calling every single month up and reading nothing at all lands 37 of the same 65 months, or 56.92 per cent. The always up count is the baseline, and every reading on this record has to be measured against it.
| Threshold setting | Matched calls | Hit rate | Points against the line | Where it sits |
|---|---|---|---|---|
| minus 6.00 per cent | 39 of 65 | 60.00 per cent | 3.08 | above the line |
| minus 5.00 per cent | 38 of 65 | 58.46 per cent | 1.54 | above the line |
| minus 4.00 per cent | 38 of 65 | 58.46 per cent | 1.54 | above the line |
| minus 3.00 per cent | 39 of 65 | 60.00 per cent | 3.08 | above the line |
| minus 2.00 per cent | 41 of 65 | 63.08 per cent | 6.15 | above the line |
| minus 1.00 per cent | 44 of 65 | 67.69 per cent | 10.77 | above the line |
| 0.00 per cent | 43 of 65 | 66.15 per cent | 9.23 | above the line |
| 1.00 per cent | 42 of 65 | 64.62 per cent | 7.69 | above the line |
| 2.00 per cent | 40 of 65 | 61.54 per cent | 4.62 | above the line |
| 3.00 per cent | 38 of 65 | 58.46 per cent | 1.54 | above the line |
| 4.00 per cent | 33 of 65 | 50.77 per cent | minus 6.15 | below the line |
| 5.00 per cent | 30 of 65 | 46.15 per cent | minus 10.77 | below the line |
| 6.00 per cent | 28 of 65 | 43.08 per cent | minus 13.85 | below the line |
Read the last two columns rather than the middle one. Ten of the thirteen settings read above the always up line and three read below it, and not one of them lands exactly on it. The best setting beats reading nothing by 10.77 points, or four extra months out of 65. Four extra months is the honest size of what the rule appears to add on this record, and it is a very different sentence from the one that starts with 67.69 per cent.
The best setting reads 67.69 per cent and the always up line reads 56.92 per cent. How much is the rule actually adding, and what has to be known before treating that as real?
The calculator below walks the threshold from minus 6.00 per cent up to 6.00 per cent. Does the count climb steadily to a peak and fall away, or does it wander on the way?
Walk the grid, and watch how many attempts it takes to find the tallest bar
One control, one number inside the rule, thirteen places it can sit. The bar redraws to the count of matched calls at that setting, the dashed line across the middle is the always up baseline and never moves, and every setting landed on keeps a small tick above it. The green line marks the best found so far and never comes down. Walking upward from the first setting on the grid shows how many attempts pass before that green line stops moving.
Educational illustration. An unmoved month is excluded, so the denominator is 65 rather than 71. The grid was fixed at thirteen settings before any of them was tried. Every bar counts matched calls out of 65 months, never money.
What does the whole read look like, in order?
Four readings, taken in the order anybody would actually take them, on one record with one rule. Read down the table and watch the same rule produce two answers that are nowhere near each other.
| Step | What is being read | What it lands on |
|---|---|---|
| First | How many months can be scored at all | 65 of 72 |
| Second | What calling every month up and reading nothing would land | 37 of 65, 56.92 per cent |
| Third | The best of the thirteen settings on the whole record | 44 of 65, 67.69 per cent |
| Fourth | That threshold read on the first three years, then on the last three | 89.66 per cent, then 50.00 per cent |
One record, one rule, one threshold, and two readings 39.6552 points apart. Nothing was rerun, nothing was recounted a second way and no arithmetic changed between the third row and the fourth. The only thing that changed is which months the rule was standing on.
A threshold picked on the first three years calls 26 of the 29 scoreable months there. On the 36 months that arrived afterwards, does it hold up, slip a few points, or fall a long way?
Live Performance: what did the rule do on the months nobody chose?
Suppose the threshold had actually been settled at the close of December 2021, using the first three years and nothing else. The setting that calls the most months there is minus 1.00 per cent. On those 29 scoreable months it calls 26, or 89.66 per cent. If the work stopped there, that is the figure that would have been written down.
The 36 months that arrived afterwards are the live stretch. Every one of them turned up after the threshold was already fixed. A live stretch cannot be rerun, cannot be recounted a second way, and does not care in the slightest what was chosen before it began. The live stretch is worth something for exactly that reason. Every other reading in this guide can be taken again with a different convention until it comes out nicer. The choosing was over before those months arrived, so the live reading cannot be retaken.
On those 36 months the same threshold, unchanged, calls 18. Eighteen of 36 is 50.00 per cent. The reading falls 39.6552 points, from 89.66 per cent to 50.00 per cent, without one thing about the rule changing.
Two cautions about that second figure, and both matter more than the figure does. The first is that 50.00 per cent is not a law. The figure is 18 matched calls out of 36 on these particular months, and nothing made it land on a round number. A rule whose fitted reading collapses does not have to land there, and reading this one figure as though it were the natural resting place of a failed test would be inventing a pattern out of one observation. The second is that nobody arranged the live stretch to be hostile. Its months simply scatter more than twice as widely as the months the threshold was chosen on, and nobody chose that either.
Backtest Overfitting: what did searching thirteen settings buy, and what did it give back?
Go back to the moment before the threshold was chosen and watch the choosing happen. Walk the grid in its stated order, from minus 6.00 per cent upward, keeping the best setting found so far. On the first three years the first setting tried calls 17 of 29, or 58.62 per cent. By the sixth attempt the best so far calls 26 of 29, or 89.66 per cent. Six attempts bought 31.0345 points on the stretch the search could see, and the seventh attempt onward bought nothing more.
Now read the same six settings on the 36 months the search could not see, one at a time, in the same order. The first attempt reads 61.11 per cent there. The sixth attempt reads 50.00 per cent. The same six attempts that bought 31.0345 points on the visible stretch gave back 11.1111 points on the invisible one. The two lines walk in opposite directions, and they do it while the rule stays exactly the same size.
Backtest overfitting is the name for what just happened, and precision about it matters because the words invite a vaguer story. Nothing was tuned. Nothing was smoothed. No extra term was bolted onto the rule. Somebody looked at thirteen candidates and kept the one that looked best on the months in front of them. Keeping the best of thirteen is a completely reasonable thing to do and is what anybody would do. The searching itself is what moved the reading, and the searching leaves no mark on the rule for anybody to find afterwards.
Searching lifted the fitted reading by 31.0345 points. A colleague says that proves the search worked. What single figure answers them?
Backtest Overfitting vs Model Overfitting: what exactly gets added in each?
Overfitting a model and overfitting a backtest get treated as the same worry in a slightly different hat, and they are not. The difference is a count, and once the count is in view it stops being a matter of judgement altogether.
Overfitting a model adds numbers to one rule. Give a rule a second number, then a third, then a tenth, and it gains the flexibility to bend around whatever record it is shown. The rule visibly grows. Anybody looking at it can see it growing. Overfitting a model therefore has a recognisable smell: the rule that fits beautifully is the complicated one.
Backtest overfitting adds attempts at one record. The rule stays exactly as simple as it was. The Ashwin rule held exactly one number at every one of its thirteen attempts. The thirteen candidates differ in that one number and in nothing else. Not one of them is more flexible than any other, not one of them could bend around anything, and every single one of them is a rule anybody would be happy to explain in a sentence. The count of attempts rose from one to thirteen, and the rule that came out at the end is exactly as simple as the rule that went in.
Backtest overfitting is harder to catch for exactly that reason. There is nothing on the finished rule to inspect. The complicated rule announces itself; the well searched one does not. The only trace the search leaves is a count of how many candidates were looked at, and that count lives in somebody's head or in a script nobody reads, not in the rule and not in the figure.
Overfitting a model and overfitting a backtest are not the same thing. What gets added in each?
What should be asked of any backtested figure handed over?
A single habit survives everything else above. A tested figure arrives in a meeting, in a note, in a message at eleven at night, and there is about a minute to weigh it. Here are the seven questions, in the order they cost the least, and what each one is actually protecting against.
| The question | What it protects against |
|---|---|
| What record, and how many observations of it? | A hit rate off twelve months and a hit rate off seventy two months are not the same kind of claim, and nothing in the figure itself says which one it is. |
| How many of those could be scored, and where did the rest go? | Dropped months change the count sitting underneath the share. Here seven of the 72 went, and all six of the unmoved ones came out of one stretch rather than both. |
| What is the baseline, and does the figure beat it? | Without the line, every reading looks like an achievement. With it, three of these thirteen settings are worse than reading nothing at all. |
| How many settings were tried before this one was reported? | A reading that won a search of thirteen and a reading that was the only one taken are different objects, and they print identically. |
| Was any stretch held back, and was it looked at first? | A stretch that was peeked at before the threshold was chosen is not held back any more, whatever it is called in the write up. |
| Could every number the rule reads have been known on the day it read it? | A signal that quietly includes information from later is the one error that makes a rule look better rather than worse. |
| What would the same figure be under a different counting convention? | Counting the unmoved months as rises instead of dropping them moves the denominator and the count together, and the reading moves with them. |
Somebody who works at a lender reads a scoring rule this way as a matter of routine, and the habit is not suspicion. The reason is that a rule chosen from many candidates and a rule written down first are two completely different objects wearing the same clothes, and the only thing separating them is a count nobody volunteers. A backtested figure quoted without its count of attempts is not a small gap in the record; it is a figure nobody can weigh, including the person who quoted it.
The reverse habit is the useful one to build. A tested figure produced in house travels with those five fields every time: the record, the scoreable count, the baseline, the count of settings tried and whether anything was held back. The five fields cost one line. One line buys a figure that somebody who trusts nobody can check, and that is the only kind of checkable worth having.
A colleague sends a backtested hit rate with no other detail at all. Which three questions come first?
The failure: writing down 67.69 per cent and walking away with it
Somebody reads the ladder, sees 67.69 per cent at the top of it, and concludes that the Ashwin rule calls direction right about two times in three. The conclusion goes into a note. Nothing they did was dishonest and nothing they computed was wrong. The figure is exactly what it says it is: 44 matched calls out of 65 scoreable months of the six year record.
Two things were skipped, and both of them were sitting alongside the figure itself. The first is the baseline. Calling every month up and reading nothing at all lands 56.92 per cent on those identical 65 months, so the whole of what the rule appears to add is 10.77 points rather than 67.69 per cent. The second is the count of attempts. 67.69 per cent is the best of thirteen readings that run down to 43.08 per cent, and the best of thirteen is not the same kind of number as a single reading, even though it is written in exactly the same way.
The cost is not hypothetical on this record, and that is the uncomfortable part. Carry the same threshold across to the 36 months nobody chose and it calls 18 of them. The note was written, the figure travelled, and the months that would have contradicted it were already in the record when the note was written.
The habit that prevents it costs less than the quoting did. Before any backtested figure goes into a sentence, ask for the baseline and ask how many settings were tried. If neither answer is available, the figure is not ready to be written down yet, and saying so is a great deal cheaper than un-saying it later.
Subjects covered elsewhere. Holding a stretch of the record back before anything is chosen, and reading the rule on it afterwards, is set out separately, and so is rolling that hold out forward through time rather than doing it once. Running many tests against one threshold, the way choices made while analysing one result change what gets reported, and the effect of picking the question itself after looking at the record each have a treatment of their own. Placing an order, paying a cost, moving money, or deciding whether any tested rule is a strategy worth anybody's time, is covered separately.
Who stands behind the figures printed above?
No outside source stands behind the figures. The recipe is printed instead, and the table below says where each part of it lives.
| The input used | What it is |
|---|---|
| The arithmetic check kept with these notes | A script that rebuilds the six year record from its three parts and recomputes every count printed above, then refuses to finish if any of them moves |
| The six year record of the Nakshatra unit | 72 invented monthly observations, settled where the record was built and reused here without a single figure being changed |
| The Ashwin rule and its thirteen settings | One sentence and a stated list of thresholds, both fixed before any count was taken |
| Everything else | Nothing else. No outside body, no maintained series, no published finding and no record of anybody's results stands behind any figure here |
The Nakshatra unit, its six year record and the Ashwin rule are invented.
Educational material. Not advice on any investment, tax, budget or market position.
