Training, Validation and Test Sets, Compared in Full
Training months fit the model, validation months choose between models, and test months are opened once at the end so the figure reported was never tuned against. Cut the ten paired months of the Nakshatra unit and the Vasant unit six, two and two and one fitted line misses by 4.8774 on the months it learned from, 5.2169 on the validation pair and 5.8266 on the test pair.
Two words in that paragraph carry hidden weight, and both get settled before anything is built on them. A traded unitA label with a quoted price behind it and no story of any kind. Nobody issues it, nobody runs it, nothing is made or sold. All the arithmetic here needs is that the quote lands somewhere different each month. is a bare name with a quote against it. Its monthly change records the distance that quote covered across one month, stated as a share of where the month opened. When one reading sits at 2.00 per cent and the next at 18.50 per cent, what separates them is 16.50 percentage pointsWhat is left after one percentage reading is subtracted from a second, counted as plain units rather than as a share of anything. A reading that climbs from 2.00 to 18.50 has crossed 16.50 of them. Calling that a rise of 16.50 per cent would describe a completely different journey., and every miss below is measured in those.
Everything after this is one cut through one record, made three ways and then made every possible way. Where the cut falls decides more than most reported results admit.
What arrives already settled, and what does the split add?
Three things are taken as read here, and none of them gets explained a second time. The mechanics of laying a straight line through two columns of readings belong elsewhere. So does the idea of a miss, along with squaring the misses and totalling them. So does the range of regression forms on offer and what separates one from another. All three are treated in their own right elsewhere and are simply borrowed here.
The split adds a vocabulary and a warning: three named parts of a record with one job each, and the demonstration that a figure taken from any one cut is a single draw from a range wide enough to swallow the result. Somebody who already knows how to fit a line can still report a number that means nothing.
What are the three parts, and what is each one for?
The record comes first, and it is then cut into three named stretches. Each stretch gets one job. The contrast between the three is worthless until each side has been pinned down on its own, so all three are defined in full before any of them is put against the others.
The training months are where the numbers inside the model come from. A fitting routine reads those months and nothing else, and returns a set of coefficientsThe numbers a fitting routine settles on and writes into the model, such as the slope and the intercept of a straight line. They are chosen from the data rather than picked by a person, which is what makes the months they were chosen from a poor place to judge them.: on a straight line, a slope and an intercept. The coefficients are the model. Every month outside the training stretch is invisible to the routine while this happens, and that invisibility is the only thing that makes the rest of the exercise mean anything.
The validation months are where a choice between finished models is made. Nothing is fitted on them. Two or three or ten already-fitted candidates are each asked for their readings on these months, the misses are scored, and one candidate is kept. The validation months therefore do get used, repeatedly, and they leave their fingerprints on the model that survives. A candidate that wins on the validation months has been selected partly for suiting those particular months. A validation figure is therefore not a clean report either.
The test months are opened once, at the end, and report a figure that nothing was tunedAdjusted after seeing a result. Changing an input, a setting or a choice of model because a score came back disappointing is tuning, whether or not anybody calls it that, and whatever the score was measured on stops being an independent reading the moment it happens. against. Reporting is their whole job and their only job. The test months are not there to improve anything. The moment a test figure is looked at and something is changed because of it, those months have been tuned against, they have become validation months, and there is no test figure left on the desk at all.
Here is the everyday version, and it is worth holding on to because the three acts feel identical until they are named. A cook making a large pot for a wedding tastes it while cooking and adds salt, over and over. Tasting and salting is training: the pot is being changed by what the tasting says. Then two versions get made and one friend is asked which is better. Asking the friend is validation: nothing is being adjusted any more, one of two finished things is being picked. Then the pot goes out to the table. Serving the pot is the test, and it happens once. A cook who carries the pot back into the kitchen after the first guest pulls a face has not run a test, they have run a very expensive tasting.
What is the one job of the validation months?
Why can a model not be judged on the months it was fitted to?
The plainest point in the whole subject is regularly treated as a subtlety, and it needs stating without decoration. The coefficients were chosen to make the misses on the training months as small as they could be made. Making those misses small is not a side effect of fitting, it is the definition of fitting. Scoring the model on those same months therefore asks the arithmetic to mark its own work, using the very criterion it was told to optimise.
The circularity is not a delicate statistical caveat, and the trap is that the circular figure is the one a fitting routine prints without being asked for it. Run any fitting tool and a miss on the fitted rows appears in the output by default. Nobody requested it, nobody had to earn it, and it sits at the top of the report looking exactly like a result. A figure computed on months the model has never seen has to be asked for deliberately. The honest number is always the one somebody has to go out of their way to produce.
The household version: a candidate sets their own exam, writes the answer key, and then sits the exam. Scoring well tells the room that the candidate can copy their own key. A good score tells them nothing whatsoever about the paper somebody else will set next month. The fitting routine is doing exactly that, at speed, without malice, and it prints the score at the top.
A fitting routine prints a miss on the very months it was just fitted to. What is that number evidence of?
What does the cut look like on these ten months?
The record runs to ten months, made up for teaching, and each one holds two readings. The Nakshatra unit's monthly change goes in and the Vasant unit's monthly change comes out. Both columns sit in time orderRows left exactly where the calendar put them, oldest at the top, never reshuffled and never ranked by size. Reordering tidies up a chart and throws away the one thing that lets anybody ask what followed what. and they stay in it. Months 1 to 6 train, months 7 and 8 validate, months 9 and 10 test.
| Month | Nakshatra change | Vasant change | Part |
|---|---|---|---|
| 1 | 1.00 per cent | 3.00 per cent | Training |
| 2 | 6.00 per cent | 18.50 per cent | Training |
| 3 | minus 4.00 per cent | 2.50 per cent | Training |
| 4 | 11.00 per cent | 17.00 per cent | Training |
| 5 | 1.00 per cent | minus 3.00 per cent | Training |
| 6 | minus 9.00 per cent | minus 13.00 per cent | Training |
| 7 | 6.00 per cent | 6.50 per cent | Validation |
| 8 | 1.00 per cent | minus 1.00 per cent | Validation |
| 9 | minus 4.00 per cent | minus 7.50 per cent | Test |
| 10 | 1.00 per cent | minus 3.00 per cent | Test |
Fitted on the training six alone, the line reads 2.6467 plus 1.5200 times the Nakshatra unit's monthly change. The line fitted to all ten months is worth setting beside it. The fuller line reads 0.5000 plus 1.5000 times the same input and leaves squared misses of 218 across the record. Notice which half moved. The slope barely shifted, from 1.5000 to 1.5200. The intercept more than quintupled, from 0.5000 to 2.6467. Dropping four months out of ten left the tilt of the line almost untouched and moved its height substantially, so a model can be badly wrong about the level while looking steady about the relationship.
And the honest note that belongs here rather than in a footnote. Six months is a very short record to fit anything on. The shortness is part of what makes the demonstration below as violent as it is. A longer record would make every figure here calmer without changing a single one of the arguments.
The line fitted to the first six months misses those six by 4.8774. Before the panel below is read, by how much does that line miss months 9 and 10?
How much worse is the model on the months it has not seen?
One line, fitted once on the training six, then asked for its readings on all three stretches. The score used throughout is the root mean squared missSquare each miss so the sign stops mattering, average the squares across the readings, then take the square root to get back into the original units. It comes out in percentage points here, and how it compares with the other error measures is covered separately.. The score comes out in percentage points and can therefore be compared across stretches of different sizes.
| Stretch | Months | What it did there | Root mean squared miss |
|---|---|---|---|
| Training | 1 to 6 | The coefficients came from here | 4.8774 |
| Validation | 7 and 8 | Never fitted on, but a choice would be made here | 5.2169 |
| Test | 9 and 10 | Never seen by anything | 5.8266 |
The three figures rise in order, and the model is 19.46 per cent worse on the months it has never seen than on the months it learned from, both figures computed here rather than asserted. The rise is not dramatic and it is not supposed to be. The ordering matters, and so does the fact that it was measured at all. A report that quoted 4.8774 and stopped would have understated the miss by about a fifth without saying anything false.
Why must the cut follow the time order?
The ten months arrived in an order, and that order is part of the data rather than a presentational choice. Cut the record at random and month 10 can land in the training block while month 2 lands in the test block. The line is then fitted using something that happened in the tenth month and marked on something that happened in the second. No future record will ever be able to do that.
Shuffling the record before cutting it does not make the test fairer, it makes the test easier, and it does so silently. The output never records that the months were reordered. The score comes back looking like every other score. Think of a shopkeeper working out how much stock to hold on a Friday, using what the following Monday turned out to need. The rule that comes out will look wonderful on the record. Monday has not happened yet, so the rule cannot be run on a single future Friday at all.
Why does cutting these ten months at random make the test easier rather than fairer?
Can a held-out score be perfect purely by luck?
A held-out score can be perfect purely by luck, and this record contains the proof rather than an argument for it. Change the exercise slightly: instead of one six, two and two cut, hold out just two months, fit on the remaining eight, and score the two that were kept back. There are 45 ways to pick two months out of ten, and all 45 were worked through.
Exactly one of them returns a miss of exactly nothing. Hold out months 4 and 6, and the line refitted on the other eight comes back reading 0.5000 plus 1.5000 times the input. The refitted line is the same line, to the last decimal place, as the one fitted on all ten. The two match because months 4 and 6 are the two the full line misses by nothing at all. Dropping a month the line already passes exactly through does not tug the line anywhere, so the refitted line arrives back at those two months and hits both of them dead on.
A held-out score of exactly nothing was produced here by which two months were chosen and by nothing else, and no step in the procedure raised a hand about it. The output would read: model tested on data it had never seen, miss of nought. Every word of that sentence is true and the number is worthless.
One of the 45 two-month holdouts returns a miss of exactly nothing. Is that model perfect on unseen months?
Move the two months held out, and watch the refitted line and the reported miss move with them.
One control moves: which two of the ten months are kept back. The line is refitted on the other eight every time, so the model form never changes and the data never changes. The strip marks the two months held out, the panel redraws the refitted line against the pinned line fitted on all ten, and the track underneath rescales the reported miss against the widest figure any cut of this record produced. At the opening setting the two held-out months are 9 and 10, the refitted line reads 1.4609 plus 1.4471 times the input, and the miss on those two months is 4.7418. The miss of 4.7418 is the median of all 45, and months 9 and 10 are the pair used throughout.
Educational illustration, built entirely on invented figures. Nothing named on this panel trades. The two units, and the ten monthly readings written against them, exist only to be worked through, and correspond to no security, no fund, no index and no business anywhere. The ten months never change and never leave the order they arrived in. Months 5 and 10 hold identical readings, so one marker can end up sitting over the other. The control moves exactly one thing: which two months are kept back. The model stays a straight line throughout, so no other kind of model is being tested.
Does a model always do worse on months it has not seen?
No, and this is where a great many explanations quietly overstate the case. The claim that a model always scores worse on unseen data is comfortable, memorable and false, and the record here can settle it because every possibility was enumeratedEvery case worked through, one after another, rather than a handful tried and the rest assumed to follow. It is only available when the number of cases is small enough to finish, which is exactly why a ten month record is a good place to look. rather than sampled.
There are 1,260 ways to cut ten months into a six, a two and a two. All 1,260 were solved: the line refitted on each training six, then scored on that split's own training months and on its own test pair. The held-out months came out worse in 797 of the 1,260 and better in 463. The honest statement is that a model does worse on unseen months on average and 63.25 per cent of the time on this record, and that is a far weaker claim than always.
The spread is the part worth sitting with. Averaged across all 1,260 cuts the training miss is 4.1784 and the test miss is 5.7340. The general pattern shows up cleanly in those two averages. But the individual test figures run from 0.1865 all the way to 16.5680. Same ten months, same straight line, same procedure. The only thing that changed was which months landed in which block, and the reported quality of the model moved by a factor of about eighty-nine.
Across all 1,260 cuts the held-out months came out better in 463 of them. Does a model always do worse on data it has not seen?
How is the cut chosen, and what has to be said about it?
Three instructions, and a lender, an analyst or anybody reading somebody else's model result can apply all three from the outside without touching the arithmetic.
- Cut along the order the record arrived in, so nothing later ever helps fit anything earlier.
- Fix the three sizes before a single result has been looked at, and write them down where they cannot be quietly revised.
- Report which cut was used and why it was chosen. A cut selected after the answers were visible is not a test of anything.
The third instruction is the one that gets skipped, so it is kept here. The test pair used throughout is months 9 and 10. Months 9 and 10 were chosen because their miss of 4.7418 is exactly the medianThe middle value once everything is lined up in order, so half the readings sit below it and half above. It is used here rather than the average because a single extreme reading drags an average around and leaves the middle value where it was. of all 45 two-month holdouts. A median cut is a typical one rather than a flattering one. Had months 4 and 6 been quoted instead, the reported miss would have been exactly nothing. Had months 2 and 3 been quoted, it would have been 10.6419. All three of those figures come from the same model on the same ten months, so quoting any one of them without naming the cut is publishing a choice while calling it a measurement.
Anybody on the receiving end of a result is left with a single question. Before a held-out figure is taken seriously, the thing to establish is how far it would move if the cut moved. If nobody in the room knows, the figure is one draw from a range nobody has looked at, and it is worth precisely as much as the range is narrow.
A held-out miss of 4.7418 arrives with nothing else attached. What is the question to ask?
The split that got tried once, and the figure that came out of it
A team cuts the record once, scores the held-out months, and writes the figure down as the quality of the model. Nothing in that sentence is wrong. The arithmetic is right, the months really were held out, and the procedure was followed.
The cost sits entirely in what was never written beside the figure. On these ten months the reported number could have been 0.1865 or 16.5680 or anything in between, depending only on which months landed in which block. One particular pair of held-out months returns exactly nothing, not because the model is perfect but because those two months happen to sit on the line. Whatever cut was tried first became the answer, and nothing in the output records that a different cut was ever available.
Notice who pays. Not the person who ran it, who knows how long the record was and how tentative the whole thing felt. The person who pays is the second reader, who receives one number with no range attached and reasonably assumes that a figure quoted to four decimal places was pinned down to four decimal places. The habit that fixes this costs one line: before quoting a held-out figure, ask how far it would move if the cut moved, and if nobody has looked, say that instead of quoting the figure. Being able to say that the answer is unknown is not a weakness in the report. Saying so is the only honest response to a ten month record.
Somebody reports a test figure, adjusts the model to improve it, then reports the improved figure as a test result. What broke?
The problem set out above has remedies, and they are treated in their own right. Testing across many cuts at once is covered separately and in full, and no single lucky pair then decides the figure. So is what happens when a model is allowed to bend further and further toward the months in front of it, and so is where the leftover miss actually comes from and which parts of it more months would cure. Whether any fitted line is worth pointing at something that trades is a separate question, and a much harder one.
Where did every figure above come from?
Every figure above came out of arithmetic on ten made-up months, so each one can be recomputed and none can be looked up.
| What was used | Where it came from | What can be checked independently |
|---|---|---|
| The ten paired monthly changes | Written for teaching. Not observed, not sampled, not collected from anywhere | The sums, month by month, from the columns printed above |
| The line fitted on the training six | Worked out here from those made-up months and from nothing else | Refitting it on months 1 to 6 shows whether the reading lands where it is stated above |
| The 45 holdouts and the 1,260 cuts | Counted by solving every cut rather than by trying a few and assuming the rest | The counts close: 45 pairs, 1,260 cuts, 797 one way and 463 the other |
The Nakshatra unit and the Vasant unit, with their ten paired months, are invented.
Educational material. Not advice on any investment, tax, budget or market position.
