Residuals: What the Model Could Not Explain
A residual is how far the fitted line was out in one month: what actually happened, less what the line said would happen. Across ten invented months the ten residuals are 1, 9, 8, 0, minus 5, 0, minus 3, minus 3, minus 2 and minus 5, and they add to exactly zero. The largest is 9.00 percentage points, and the typical miss is 4.6690 per cent.
Three words in that paragraph carry weight, and none of them is assumed here. A traded unitA nameless thing with a price on it. Subtracting one figure from another never needs to know, so whether it stands for shares, a fund or a basket is left blank deliberately. The unit is worth something at the start of a month and something else at the end, and that is all it has to be. is such a thing, and its monthly change answers one question: by what percentage did its price end a month away from where it began? When the line names 9.50 for some month and the month delivers 18.50 instead, the distance separating those two readings measures 9.00 percentage pointsThe unit that results from taking one percentage away from another. The climb from a reading of 9.50 up to a reading of 18.50 measures 9.00 of them. Calling that a rise of 9.00 per cent describes an entirely separate quantity. The two labels are kept apart for that reason., not 9.00 per cent. The distance between the two readings, taken one month at a time, is the entire subject here.
The record is ten paired observationsA pair of readings taken off one case, worthless the moment they are separated. A month here supplies one reading to each of the two columns, and re-ordering either column while leaving the other alone destroys every figure below., all invented for teaching. Against every month sit two readings: how far the Nakshatra unit moved, the reading that does the explaining, and how far the Vasant unit moved, the reading being explained. The ten months are held in time orderOldest first, newest last, exactly as the months arrived. Any other arrangement, by size or by neatness, is a different record altogether. Questions about what follows what stop having answers in it. and they stay there. One of the checks below reads along that order, and a tidied record cannot answer it.
What arrives already settled, and what do the leftovers add?
Two things arrive settled and are used without being taught a second time. How a straight line gets drawn through ten pairs of readings, and the two conditions that drawing it forces on whatever is left over, are covered separately. What the intercept and the slope each claim is covered separately as well. The line itself is simply handed over: the fitted value for any month is 0.5000 plus 1.5000 times that month's input, and both of those figures are exact rather than rounded.
The leftovers are evidence in their own right, not merely the part of the data the model failed to use. A fit statistic collapses ten months into one figure and then that figure gets reported. The ten misses underneath it are ten separate facts, and every interesting thing that can go wrong with a fitted line shows up in them before it shows up anywhere else. Learning to look at them is a habit, not a technique, and it takes about a minute per model.
What exactly is a residual, and how is one worked out?
Month two on its own, done in the open. In that month the Nakshatra unit moved 6.00 per cent. Put that into the line: 0.5000 plus 1.5000 times 6.00 gives 9.50, so the line said the Vasant unit would move 9.50 per cent. The Vasant unit actually moved 18.50 per cent. Subtracting one from the other leaves 9.00.
Actual less fitted is the convention, worth fixing once and never varying. A positive answer means the month came in above what the line said. Reversed, every figure here keeps its size and flips its meaning, which is a fine way to publish a table that is exactly backwards. Because the direction matters more than the sign does, it is written here as a word. Month two came in 9.00 percentage points above the line. Month five came in 5.00 below it. Nobody has to remember which way round a minus sign was pointing.
Month three: the Nakshatra unit moved minus 4.00 per cent and the Vasant unit moved 2.50 per cent. The line is 0.5000 plus 1.5000 times the input. What is the residual?
Why do the ten misses always add up to nothing?
The same four steps run on all ten months give the table below. Every fitted value comes off the same line, every actual figure is the one recorded for that month, and every residual is the second column taken away from the third. Reading down the last column and adding as it goes: 1, then 10, then 18, then 18 still, then 13, 13, 10, 7, 5, and 0.
| Month | The Nakshatra unit | The line said | The Vasant unit did | The residual |
|---|---|---|---|---|
| 1 | 1.00 | 2.00 | 3.00 | 1.00 above |
| 2 | 6.00 | 9.50 | 18.50 | 9.00 above |
| 3 | minus 4.00 | minus 5.50 | 2.50 | 8.00 above |
| 4 | 11.00 | 17.00 | 17.00 | 0.00, exact |
| 5 | 1.00 | 2.00 | minus 3.00 | 5.00 below |
| 6 | minus 9.00 | minus 13.00 | minus 13.00 | 0.00, exact |
| 7 | 6.00 | 9.50 | 6.50 | 3.00 below |
| 8 | 1.00 | 2.00 | minus 1.00 | 3.00 below |
| 9 | minus 4.00 | minus 5.50 | minus 7.50 | 2.00 below |
| 10 | 1.00 | 2.00 | minus 3.00 | 5.00 below |
| All ten | 10.00 | 20.00 | 20.00 | 0.00 |
Figures: an invented ten month record, recomputed here from the twenty numbers in the two middle columns. No maintained series is used, so there is no as-of date to state. Every reading is a percentage; every residual is in percentage points.
Eighteen points of overshoot in the first three months, eighteen points of undershoot spread across the last six, and the two cancel to the last decimal. The closure at zero is satisfying to watch, and it is also the single most over-read fact about residuals.
The sum closing at zero confirms that the arithmetic ran, and says nothing whatever about the data. The method that produced this line is built to force it, and it forces it on any two columns of numbers whatever. Hand it ten months of pure noise and the misses will still sum to zero. Hand it the chair count at a community hall against the monthly change of a traded unit and they will sum to zero. A reader who checks this and feels reassured has been reassured by a tautology: the check can only ever come back clean, and a check that cannot fail carries no information.
There is a second condition on the pile, and it is the same kind of thing. The misses on this record are exactly uncorrelated with the input. Not nearly, not to four decimals, but exactly. The fitting method has no choice about that either. Both properties are worth knowing precisely so that they are not read as findings.
An analyst checks the ten residuals, finds they add to exactly zero, and writes clean beside it. What has that check established about the data?
What is one miss worth in familiar terms?
Percentage points are honest units. Percentage points are also slippery ones, and most people cannot hold nine of them in their head as a quantity. So put one miss into money. Suppose an invented holdingAn amount of something actually held, valued in rupees. Here it is a round Rs 2,00,000/- invented purely so that a percentage can be turned into a figure a reader already has a feel for. of Rs 2,00,000/- that moves exactly as the Vasant unit moves. Nine percentage points of that holding is Rs 18,000/-. Eighteen thousand rupees of movement on a two lakh holding, in one month, is what the worst month's miss is worth.
Two more figures follow the same way. Ignore which side of the line each one fell on, and the average size of the ten misses is 3.60 percentage points, or Rs 7,200/-. The root mean squared miss squares each one before averaging and then unsquares the answer, giving 4.6690 per cent, or Rs 9,338/-. The three of them together give a range: usually out by around seven thousand rupees, occasionally out by eighteen.
None of these three rupee figures is a gain, a loss, or an expectation about anything; each one is the size of a model's error, translated into a unit a reader can picture. Nobody made or lost Rs 18,000/- in month two. In month two a line drawn through ten invented months named a figure and the record named a different one, and the gap between the two was worth that much on a holding of that size. Kept separate, the two ideas make the money version useful. Merged, they quietly turn a diagnostic into a profit and loss statement.
The residual in month two is 9.00 percentage points. On an invented holding of Rs 2,00,000/-, what is that worth, and what is the figure not?
Which month do the two error measures each call the worst?
There are two ordinary ways to add ten misses into one score. Square each one and add, or take each one's size and add. Squaring settles in advance how heavily one bad month ought to weigh, so the two totals are both defensible and not the same instrument. Rather than argue about it, hand each month its share of each total and read the two columns side by side.
The squared total comes to 218, and month two takes 81 of it, working out at 37.16 per cent. The total by size comes to 36, and month two takes 9 of that, working out at 25.00 per cent. Do that for all ten months and both columns close at 100.00 per cent.
| Month | The residual | Share of the squared score, per cent | Share of the score by size, per cent |
|---|---|---|---|
| 1 | 1.00 above | 0.46 | 2.78 |
| 2 | 9.00 above | 37.16 | 25.00 |
| 3 | 8.00 above | 29.36 | 22.22 |
| 4 | 0.00, exact | 0.00 | 0.00 |
| 5 | 5.00 below | 11.47 | 13.89 |
| 6 | 0.00, exact | 0.00 | 0.00 |
| 7 | 3.00 below | 4.13 | 8.33 |
| 8 | 3.00 below | 4.13 | 8.33 |
| 9 | 2.00 below | 1.83 | 5.56 |
| 10 | 5.00 below | 11.47 | 13.89 |
| All ten | 0.00 | 100.00 | 100.00 |
Figures: an invented ten month record, recomputed here. Each share is printed to two places, so adding the ten printed figures in the squared column returns 100.01 while adding the unrounded ones returns exactly 100.00. The unrounded sum is the honest one, and the total row carries it.
Now read the two columns against each other. Both agree that month two is the worst and month three the next worst, so the ranking is not in dispute. The two measures disagree about the weight rather than the order, and squaring hands month two half again as much of the blame as counting by size does. Under squaring, month two alone is 37.16 per cent of everything the model got wrong. Under size, it is a quarter. Look at the other end and the disagreement reverses: month nine is worth 1.83 per cent squared and 5.56 per cent by size, three times as much.
The disagreement matters for a practical reason rather than a philosophical one. Attention is limited, and only one or two months will actually get looked at. The measure used for scoring decides which month gets the visit, and on most tools that measure was a default setting nobody chose. Squaring pulls attention towards the single worst month. Counting by size spreads it across the middling ones. Neither is wrong; being unaware of which one is steering is.
Before the panel below is touched. Month five and month ten both have residuals of 5.00 below the line. Will their two shares come out the same as each other?
Selecting a month shows the two measures disagreeing about how much it cost.
One control moves: which of the ten months is selected. The ten misses, the line and both totals stay exactly where they are, so selecting a month changes the view and never the arithmetic. The month can be set with the slider, with one of the jump buttons, or by clicking a bar in the chart itself. The dashed guide on each share bar marks where the other measure would have put that month, so the disagreement is visible as a distance rather than as two numbers. At the opening setting, month 2, the panel reads a miss of 9.00 percentage points above the line, a squared share of 37.16 per cent and a share by size of 25.00 per cent. Those are the figures printed in the table above.
Educational illustration, built entirely on invented figures. Both traded units on this panel were made up, along with every monthly change written against them, and neither corresponds to a security, a company, an index or a market anywhere. The ten months stay in the order they arrived and are never sorted. The one thing the control moves is which month gets picked out: the line stays pinned at 0.5000 plus 1.5000 times the input, so the misses themselves are untouched. Whichever month is selected, the two share columns still total 100.00 per cent across all ten. The rupee readings describe an invented holding of Rs 2,00,000/- and measure the size of an error, never a gain, a loss or an expectation.
Month two carries 37.16 per cent of one score and 25.00 per cent of the other. There is time to look at one month properly. Which measure points to month two hardest, and why?
Is the typical miss 4.6690 per cent or 5.2202 per cent?
Three summary figures come off the same ten misses and they are not interchangeable. The average size of a miss is 3.60 per cent, arrived at by ignoring direction and averaging. The root mean squared miss is 4.6690 per cent, arrived at by squaring, averaging over all ten months, and taking the square root back. The residual standard error is 5.2202 per cent, arrived at the same way except that it divides by eight instead of ten.
Why eight? Because the line was not handed down from anywhere; it was worked out from these same ten months, and working it out consumed two of them in the sense that two figures, the intercept and the slope, were fixed to make the misses as small as they could be. Ten readings minus two figures used up leaves eight degrees of freedomThe count of readings still free to vary once a calculation has already spent some of them. Fit a straight line and two are spent on the two figures in the line, so ten readings leave eight.. Dividing by eight rather than ten inflates the answer, and that inflation is the correction for having fitted the line on the very record it is now being scored against.
The figure 4.6690 per cent answers how far the line was out on these ten months, and 5.2202 per cent answers how far it is likely to be out on months not yet seen. The first describes the record in hand. The second treats those ten months as a sampleThe cases actually in hand, standing in for a larger set that is not. Ten months is a sample; every month that ever was or could be is what it stands in for. and estimates the wider picture. Ten months is a very small record on which to be estimating anything. The correction is not cosmetic.
For scale, hold either figure against the spreadHow far a set of readings strings out away from its own average. A wide spread means the readings are scattered. A narrow spread means they crowd near the middle. of the outcome itself, which on these ten months is 9.96 per cent. The line took a series that wanders by about ten percentage points and left misses that wander by about five. Halving the wander is a real improvement, and five percentage points is not a small remaining error. An honest write-up states both halves.
Somebody asks how far the line was typically out across the ten months actually in hand. Which figure answers that question as asked?
What should the misses be plotted against?
A list of ten numbers hides its own shape. Drawn, it gives it up immediately, and there are exactly three things worth drawing the misses against. Any more and the exercise is fishing.
- Against the input. This catches a wrong shape. If the misses curve, sitting above the line at both ends and below it in the middle, then the relationship was never straight and fitting a straight thing to it was the error.
- Against the fitted value. This catches misses that grow with the size of the prediction: small errors on quiet months and large ones on loud months, fanning out towards the right.
- Against the clock. This catches misses that follow one another, where a month above the line tends to be followed by another month above it.
One honest note about the first two on a record like this. Because there is a single input here, and the fitted value is that input multiplied by 1.5000 and shifted by 0.5000, plotting against the fitted value redraws exactly the same picture as plotting against the input, stretched. The two checks only pull apart once a model carries more than one input. On this record they are one check wearing two labels, and it passes.
The ten misses are clean on the first two plots and not clean on the third. Set the ten misses out again in the order they arrived, which is exactly how they are printed above, and the shape is impossible to overlook: the first three months all came in above the line, then two exact months and a middling one, and then the last four all came in below it. Misses that arrive in runs like that are not behaving like independent misses.
Putting a figure on how strongly one month's miss predicts the next one, and deciding what should be done about a model with that property, is covered separately. The habit is what belongs here. All three plots are run, and when one of them fails, that failure gets written down beside the two that passed.
Name the three things residuals are worth plotting against, and what each one is looking for.
This record passes two of those three plots and fails the third. Which one does it fail, and what follows from that?
What can a residual not tell?
A large residual is a loud fact and it says less than it sounds like it says. Three things it does not establish, and each one gets treated as though it did.
A large residual does not establish that the month was unusual. A miss of 9.00 in month two is equally consistent with two stories: something odd happened that month, or the line is the wrong shape and month two is where the wrongness shows up hardest. The residual is computed from the line and inherits whatever is wrong with it, so the residual cannot separate the two stories.
A large residual does not establish that the observation is bad. Bad means wrongly recorded, mistyped, a stale price, a corporate action nobody adjusted for. A residual has no view on any of that. A residual only reports disagreement with a line, and a perfectly correct figure can disagree loudly.
And a residual licenses no removal of anything. Removal is the one that costs money. An observation removed because it fits badly is an observation removed for disagreeing, and the fit reported afterwards is a fit on data selected to agree with it. Take month two out of this record and refit: the fit statistic climbs from 0.7559 to 0.7988 and the squared misses fall from 218 to 118.82. All of those figures survive an audit. The month removed was chosen by the very disagreement the score is measuring, so the improvement is manufactured anyway.
A residual is a question about a month, not a verdict on it. The question is worth asking every time. The verdict is not available from this column.
A colleague removes month two as an outlier, refits, and reports a fit statistic of 0.7988. What is wrong, and what should have been reported instead?
What happens to a miss that cannot be explained?
The habit earns its keep on a miss nobody can account for, and it is worth watching how a lender, an analyst or anyone running a household budget handles the same situation. A household on one salary budgets nine thousand rupees for the month's electricity and the bill comes in at fourteen. The useful move is not to change the budgeting rule on the spot. The useful move is to find out what happened: a visitor stayed three weeks, a meter was read late, a tariff slab shifted. The explanation is almost never inside the two columns under examination.
Four moves, in order, and every one of them ends in something written down.
- Write down what happened in that period before touching the model. The month itself is where to look. If nothing turns up, record that the search found nothing, which is a different and more useful statement than silence.
- Check whether the same month is extreme on the input too. A month that is unusual on both columns is a different case from one that is ordinary on the input and wild on the outcome, and the two lead somewhere different.
- If the model is refitted without it, report both fits. Not the better one. Both, side by side, with a note of which measure pointed to that month in the first place, because squaring and counting by size would have pointed to different months.
- If it is kept, the write-up says that it could not be explained. An unexplained miss that is disclosed is a known limitation. The same miss undisclosed is a claim of completeness that is not true.
Reporting both fits is the whole difference between a diagnostic and a decision taken quietly. A reader handed both can disagree and reach their own view. A reader handed only the better one cannot tell there was ever a choice. The second version is more persuasive and less useful at the same time.
Two clean checks, one improved fit, and a report nobody can audit
An analyst opens the ten misses and does two sensible things. She adds them, gets exactly zero, and writes centred beside it. She plots them against the Nakshatra unit, sees no shape at all, and writes clean beside that. Both statements are true. Neither carries any information: the sum was forced to close by the fitting method, and the absence of any relationship with the input was forced by the same method. She has run the two checks that cannot fail and skipped the one that does. The failure is expensive precisely because it looks like diligence, and it looks like diligence to a reviewer as well.
Then she notices month two. Residual 9.00, carrying 37.16 per cent of the squared error on its own. Concentration like that invites a decision. She takes it out as an outlier and refits. The fit statistic goes from 0.7559 to 0.7988 and the squared misses fall from 218 to 118.82. The report reads better, and not a single figure in it is arithmetically wrong. The month taken out was the single month that disagreed most with the line, and nothing in the write-up tells a reader that a month is missing at all.
The cost is not a wrong number. The cost is the ability of anybody outside the room where the work happened to push back on the result at all. The fix is three lines of discipline: all three plots are run and the one that fails is recorded beside the two that pass, a large residual is treated as a question about that month rather than a verdict on it, and where an observation is removed the fit is reported with and without it, along with the measure that led there. None of the three is clever and none of them costs anything but honesty.
What stands under these figures, and what does not?
No regulator, exchange or published record stands behind these figures. Subtracting a fitted figure from a figure that happened is arithmetic, and arithmetic behaves the same way in Mumbai, in Nairobi and on a kitchen table. There is no threshold to look up, no rate to confirm and no maintained series to cite.
The check that does exist can be run by anyone. Every figure here descends from twenty numbers printed in full in the build table above: ten monthly changes for the Nakshatra unit and ten for the Vasant unit. From those twenty, subtracting the two averages, the line, the ten misses, the two error totals and both share columns fall out with nothing left over. The standard here is not whose name sits at the bottom, but whether the arithmetic closes, and it closes.
| What a reference block usually carries | What sits there |
|---|---|
| An authority whose rule is being restated | None. Subtracting a fitted figure from a recorded one is arithmetic, and arithmetic answers to no market and no regulator. |
| A maintained record, with the date it was last opened | None, so there is no as-of date. Ten months written down for teaching were never current on any day. |
| A named author for the technique | None. Looking at what a fitted line left behind is ordinary practice, older and plainer than any single text, and attributing it to one would be a guess. |
| A figure carried in from an earlier note on trust | None. Every figure above was worked out again from the twenty numbers. |
| Something that would have to be fetched | Nothing at all. Two columns, ten rows, and a pencil. |
The Nakshatra unit and the Vasant unit are invented.
Educational material. Not advice on any investment, tax, budget or market position.
