Out-of-Sample Testing: The Only Test That Counts
An out of sample test reads a rule on months that were not looked at while the rule was being chosen. In sample, the Ashwin rule's chosen threshold calls 26 of 29 months, or 89.66 per cent. Out of sample, on 36 months held back, the same threshold calls 18, or 50.00 per cent. Nothing about the rule changed between the two.
Two figures, one rule, and a drop of nearly forty points between them. The two figures are not two attempts at measuring the same thing. One of them was produced by months that had a hand in picking the winner. The other was produced by months that had no say at all, and it is the second one that carries any weight. The record underneath is the six year record: 72 dated monthly readings of the Nakshatra unit, an invented thing built for exactly this kind of dismantling, running from January 2019 to December 2024. The rule laid over it is the Ashwin rule. The Ashwin rule reads last month's change and calls the coming month up or down depending on whether that reading sits above a thresholdA cut off number. The rule holds its reading up against this line and calls one way when it sits above, the other way when it does not. picked from thirteen whole number settings.
In-Sample vs Out-of-Sample Testing: which single question sorts a month into one or the other?
Ask one thing about each month, and it is not a question about the month at all. Did this month help choose the rule? If it did, it is in sample. If it did not, it is out of sample. The question is the entire test, and every distinction that follows comes out of it.
Notice what the question leaves out. The question does not ask when the month happened, whether it was early or late, whether the reading was large or small, or whether it looks representative of anything. The question asks what the month was used for. In sample and out of sample are therefore facts about the process rather than facts about the calendar. The very same December 2021 is an in sample month in a study that fitted on everything up to the end of 2021, and an out of sample month in a study that fitted on everything up to the end of 2020. December 2021 did not change. The study did.
Try the kitchen version, and it tends to stick where the technical wording slides off. A cook is making dal for a wedding, tasting and adjusting the salt as she goes. She takes a spoonful, decides it needs more salt, adds some, tastes again. The spoonful is what she has been adjusting towards, so by the time she is happy it cannot tell her whether the dal is well salted. To find out whether the dal actually works she has to hand a bowl to somebody who was not standing at the pot. A reading taken from the thing being adjusted against is not a measurement, it is a description of how hard the adjusting was.
What single question decides whether a month is in sample or out of sample?
Where is this record cut, and what does the cut cost in months?
The cut is the end of December 2021, and it was written down before anything was fitted. Writing the cut down first matters more than where it landed. The cut leaves three years on each side: 36 calendar months in the first stretchA block of months standing side by side in time, marked off from the full record and then examined on its own terms. and 36 in the second. Two halves, evenly divided, nothing clever about it.
Then the counting rules bite, and the halves stop being even. Two months cannot be scored at all. The opening month of the record, January 2019, is preceded by nothing, so the rule has no reading available and never makes a call about it at all. A month that finished exactly where it began cannot be scored either. There is no direction for the call to have matched or missed. Excluding it is the conventionAn agreed way of counting, set down in writing beforehand, so that anybody working from the same record arrives at the same figures. this record uses, and it was fixed before any fitting happened rather than chosen once the counts were visible.
On six occasions the record finished a month exactly where it started, and all six of those occasions land inside the first three years. April 2019, September 2019, August 2020, March 2021, April 2021 and September 2021. Add January 2019 and the first stretch loses seven of its 36 months. The second stretch loses none of its 36. So the two denominatorsThe bottom half of a fraction. It fixes what a share is being taken OUT OF, which is why two shares with different bottoms cannot be lined up against each other. are 29 and 36 rather than the 35 and 36 an even split would give. Assume the two halves are the same size and every comparison that follows comes out wrong.
The cut divides 36 calendar months from 36 calendar months, yet the scoreable counts come out at 29 and 36. Where did the seven missing months go?
What do the two readings actually come out at?
Fit first. On the 29 scoreable months of the first stretch, try all thirteen settings and keep whichever calls the most months correctly. The winner is a threshold of minus 1.00 per cent, calling 26 of the 29, or 89.66 per cent. Nine calls in ten. On its own that figure looks like something worth writing up.
Now carry that threshold, unchanged, to the 36 months on the other side of the cut. Same signalWhatever quantity a rule consults before committing to a call. Here it is worked out only from months that have already closed.: last month's reading. Same threshold of minus 1.00 per cent. Same convention about months that did not move. Not one character of the rule is different. The threshold calls 18 of the 36, or 50.00 per cent.
The rule did not change between those two lines, so the drop of 39.6552 points is not something that happened to the rule, it is something that was always true about the first figure. The 89.66 per cent was never a statement about how well the rule reads months. The figure stated how well thirteen attempts can be made to fit 29 particular months, a different claim wearing similar clothes.
One warning about that second figure before anybody builds anything on it. The reading lands on exactly 50.00 per cent because 18 is exactly half of 36, and nothing whatsoever made it land there. Move one month and the count is 17 or 19, and the tidy figure disappears. A failed reading does not have to come out at a coin, and this one did purely by accident. Read it as a count, 18 of 36, and the false neatness goes away.
The held back reading lands on exactly 50.00 per cent. Does that exactness mean anything?
Does the in sample reading always sit above the out of sample one?
Most treatments of this subject stop at the chosen setting, print the two figures, and let the reader draw the obvious conclusion: fitting months flatter, held back months tell the truth, the first bar is always taller. The story is satisfying, and this record refuses to support it.
The same comparison now runs at every one of the thirteen settings, not just the one the fit picked. Is the fitting bar the taller of the two at all thirteen?
Move the threshold across all thirteen settings and watch what each stretch says about it.
One control moves: the threshold, stepping through all thirteen whole number settings between minus 6.00 per cent and 6.00 per cent. Three things stay nailed down while it moves, namely the record, the cut and the way calls are counted. Two bars redraw at each step, and beneath them a track fills in the distance between those bars at every setting that has actually been opened up. The panel opens at a threshold of minus 1.00 per cent, where the rule calls 26 of the 29 fitting months for 89.66 per cent and 18 of the 36 held back months for 50.00 per cent, the pair everything above turns on.
Educational illustration. The cut stays pinned at the end of December 2021 and the control touches the threshold alone. Months that stood still are left out of the counting, and the awkward pair of totals comes from that exclusion. Every last one of those months lands on the fitting side. The lower track stays blank at settings nobody has opened yet. Blankness is the position a study is genuinely in before it has looked. Every measure on this panel is a share of months called correctly. None is a return.
Across the whole grid the fitting bar is the taller of the two at twelve of the thirteen settings. At one setting, a threshold of minus 6.00 per cent, it is the held back bar that is taller, by 2.4904 points. So a held back reading does not always come out lower, and any account claiming that it does has overclaimed on this very record.
The one exception has a plain explanation, and the explanation is arithmetic rather than luck. No month in the first three years fell as far as minus 6.00 per cent, so at that setting the rule finds nothing below its threshold and calls every single fitting month up. Its 17 of 29 is therefore not a reading of the rule at all; it is the baselineWhat something carrying nothing at all would produce. It is parked next to a result to show how much of that result is more than nothing. of calling everything up, arriving by a different route. Be careful, though: the fact holds for the earlier stretch and not for the record as a whole. Widen the view to all 72 months and six of the readings there do sink below that line, at which point the very same setting is genuinely sorting one month from another.
Now look at where the distance is widest. The widest distance is 39.6552 points, and it sits at a threshold of minus 1.00 per cent, precisely the setting the fit chose. Unlike the 50.00 per cent, the placement is no coincidence. The search went looking for whatever suited those 29 months best, and the setting that suits 29 particular months best is exactly the setting with the most fitting-to-those-months built into it. The gap is widest where the search worked hardest. At the other end of the range the two readings sit 0.2874 points apart, close enough that no chart at any sensible scale can separate them.
At a threshold of minus 6.00 per cent the held back reading is the higher of the two. Does that mean the rule reads unfamiliar months better than familiar ones?
How many calls does a reading need before it means anything?
A share on its own is not a result. Nine of ten and ninety of a hundred are the same share and they are not the same evidence, so every reading here has to be placed against the count it was computed on. The relevant question is how many correct calls a rule would need on a stretch of a given length before the reading clears the 5 per cent threshold that statistical significance sets.
On the 29 month fitting stretch the answer is 20 calls. On the 36 month held back stretch it is 24. The two bars are not the same share, and each stretch therefore needs a bar of its own: 20 of 29 is about 69 per cent, and 24 of 36 is exactly two thirds. A longer stretch asks for a smaller share and more calls at the same time.
Set the two actual results against them. The fitted rule made 26 of 29 on the fitting stretch, six calls clear of the bar. On the held back stretch the same rule made 18 of 36, six calls short of the 24 it needed, and 18 of 36 reads 56.60 per cent against that threshold rather than anywhere near the 5 per cent line. The 18 of 36 is not a near miss that better luck would have rescued. The result sits on the wrong side of the bar by a wide margin, produced by the same rule that cleared the other bar comfortably.
The rule made 18 of 36 held back calls and needed 24 to clear the threshold. How far off is that, and what does the 56.60 per cent mean here?
A researcher reads the held back stretch, dislikes the 18 of 36, decides the cut fell in an unlucky place, moves it to the end of December 2022 and reruns. What has just happened to the held back stretch?
Why can a held back stretch only be read once?
Back to the one question. A month is in sample if it helped choose the rule. Now consider what happens the moment a held back reading appears on the researcher's screen. The reading gets looked at, and then something follows: the study is written up, or the work goes back for another attempt. Both are decisions, and both were made in the light of what those months said.
A result that changes what happens next has helped choose the study, and that is true even when what it prompted was starting over. There is no version of looking at a number in which it also went unlooked at. So the second reading of the same 36 months is a fitting reading with an out of sample label stuck on the front, and the label is the only part that is out of sample.
The practical advice is correspondingly blunt. Write the cut date down before the first fit. Treat one reading of the held back stretch as the end of the study rather than a checkpoint in the middle of it. If the reading disappoints, the honest options are to publish it as it stands, or to wait for months that have not happened yet. Going back to the same months is not one of the options, however it is presented.
Why can a held back stretch be read only once?
What does the whole thing look like laid out step by step?
One cut, one threshold and thirteen comparisons. The run goes in order, and the step most write ups leave out comes at the end rather than nowhere at all.
| Step | What happens | Result |
|---|---|---|
| 1 | The cut date is written down: the end of December 2021 | 29 and 36 scoreable months |
| 2 | The grid is written down: every whole number threshold lying between minus 6.00 per cent and 6.00 per cent | 13 settings |
| 3 | Every setting is scored on the 29 fitting months, and the best is kept | minus 1.00 per cent |
| 4 | That setting's reading on the months that chose it | 26 of 29, 89.66 per cent |
| 5 | The same setting, unchanged, read once on the 36 held back months | 18 of 36, 50.00 per cent |
| 6 | The distance between the two readings at the chosen setting | 39.6552 points |
And then the step almost nobody runs. Score both sides at all thirteen settings rather than only at the one the fit landed on. The fitting counts come out at 17, 17, 19, 22, 24, 26, 25, 23, 21, 19, 17, 16 and 13 of 29. The held back counts come out at 22, 21, 19, 17, 17, 18, 18, 19, 19, 19, 16, 14 and 15 of 36. With that comparison run, the honest summary is not that held back months read lower, it is that held back months read lower at twelve of thirteen settings and higher at one, and the sentence a researcher is entitled to write changes because of it.
The failure: the held back stretch that vanished, and nothing announced it
A researcher works through the run above and reaches 18 of 36. The 18 of 36 is a disappointing number after a fitting reading of 89.66 per cent, and there is an explanation ready to hand: the cut fell in an awkward place, right where the record gets choppier. So the cut moves to the end of December 2022, the fit reruns on 41 months instead of 29, and the study proceeds.
The held back stretch is now gone, and no error message said so. The 36 months were read, the reading produced a decision, and the decision changed the study. Helping choose the study is precisely what an in sample month does. Every figure computed after that point is a fitted figure, whatever heading it appears under.
The rerun makes the loss concrete, so it is worth following through. On the moved cut the search picks the same threshold of minus 1.00 per cent, and the new held back stretch of 24 scoreable months returns 12 of 24. The 12 of 24 is 50.00 per cent again, from a stretch that is no longer held back at all. Two separate readings on this record land on exactly half, and they are unrelated to one another: the first is 18 of 36 and the second is 12 of 24, and neither number was pushed there by anything.
The cost is not that the study got a worse figure. The cost is that the study no longer has a measurement of any kind, and cannot get one back without waiting for months that have not happened yet. The fix costs nothing and has to be done first: write the cut date down before the first fit, and treat one reading of the held back stretch as the end of the work rather than a checkpoint inside it.
What has to be asked before calling something an out of sample test?
Somebody hands an analyst a figure and describes it as an out of sample result. Five questions decide whether the label survives, and none of them is about the figure itself.
When was the cut fixed, and was it fixed before or after the first fit? How many scoreable months sit on each side, and do those counts match the calendar months quoted? Has the held back stretch been read before, and how many times? Did anything at all about the rule change after the held back reading came in, including the grid it was picked from? And what does the always up baseline read on the held back stretch on its own? A rule that fails to beat calling everything up has not cleared a very low bar.
A second trip to the same months almost never feels like cheating to the person taking it, so the third question is the one that quietly undoes most held back readings. The second trip feels like diligence. The researcher is checking, refining, being careful. The arithmetic does not care what it felt like, and neither does the label.
The trap is not a specialist's concern either. A lender running a scoring rule past its first year, a household deciding whether the rule of thumb that worked for their last three big purchases is worth trusting for the fourth, an analyst who has tried nine variants of a screen and is about to report the tenth: all three are one question away from the same trap. Did the record about to do the measuring have any hand in choosing what is being measured? If it did, what is in hand is a description rather than a test.
Somebody describes a figure as an out of sample result. Which two questions most often turn out to undo the label?
What stands behind these figures?
Every count above was produced the same way: run the stated rule across the stated months, and add up how often the call matched. A count made like that rests on arithmetic rather than on an authority, so what stands behind it is a recipe rather than a citation. Rerunning the four ingredients below gives 26 of 29 and 18 of 36 every time, on any machine, in any year, and a figure is legitimate only if the recipe reproduces it.
Four ingredients regenerate every figure above, and all four are stated where the counts are.
| Ingredient | Where it is stated |
|---|---|
| The 72 monthly readings of the Nakshatra unit, January 2019 to December 2024 | Their shape is described in the block on where the record is split, and the whole set is what the panel walks across |
| The Ashwin rule, together with the thirteen settings its threshold is allowed to take | Named and restated wherever a count is quoted, so no count sits without its rule beside it |
| The counting convention that excludes a month that did not move | The block on the split, also the source of the denominators 29 and 36 |
| The split point, the end of December 2021 | Stated before any fitting is described: the order is the point |
The Nakshatra unit and the Ashwin rule are invented.
Educational material. Not advice on any investment, tax, budget or market position.
