Model Error: The Sources and Which Ones You Can Reduce
Model error has three sources, and they do not behave alike. The shape of the rule costs 72 out of a total 218 on the ten paired months of the Nakshatra unit and the Vasant unit. The estimating of the rule's own numbers from a short record shrinks as the record lengthens. A third part sits beyond the reach of any rule reading that input, 146 here, and never shrinks.
Where the figures below come from. Every one of them was worked out from ten paired monthly changes made up for teaching, using the ordinary least squaresThe fitting step that picks the line making the total of the squared misses as small as it will go. Settled in earlier reading and used here as it stands. arithmetic and nothing else. The Nakshatra unit is the input and the Vasant unit is the outcome, the fitted line is 1.5000 times the input with 0.5000 added, and the ten misses it leaves add to nothing and square up to 218. Every ladder printed below was produced by fitting each subset of the ten months in turn.
Where does model error actually come from?
This question earns a full treatment rather than a paragraph. A model gets something wrong. The size of the miss is unwelcome, and the instinct that arrives next is almost always the same one: get more data. The instinct is not silly, but it is aimed at one of three targets without checking which.
The household version comes first, and it puts the same three problems in a form that can be felt. Somebody living on one salary sits down to work out what next month will cost. The estimate comes out wrong by a fair margin. There are three separate reasons, and only one of them is fixed by keeping records for longer.
The first is that the rule they used was crude. The rule took last month and added a little. A rule that separated rent from food from travel would have missed by less, and no amount of record keeping improves a rule that has not been changed. The second is that the numbers inside their rule came from a short stretch of months. With three years of records, the typical grocery figure they are working from settles down. The third is that some bills genuinely arrive when they arrive. A relative visits, a tyre goes, a school asks for something. Three different problems, three different fixes, and only the middle one answers to a longer record.
A fitted model has exactly these three, and on the ten months of the Nakshatra unit and the Vasant unit each of them has a size.
Two of those three were priced on the earlier reading and are carried in here whole. The straight line misses these ten months by 218 when the misses are squared and added. Four of the ten months carry an identical Nakshatra reading of 1.00 per cent, and their Vasant outcomes are 3.00, then minus 3.00, then minus 1.00, then minus 3.00 per cent. Of the 218, 146 therefore sits past whatever any rule reading only the Nakshatra unit could get at. Any rule reading only that input has to hand those four months the same answer, so it must miss on at least three of them. The remaining 72 is what the straightness costs.
The third source is the one this guide is really about, and it is worth noticing that it does not appear in the 218 at all. It cannot. The 218 is what the line missed by on the very months it was fitted to, and on those months the line has already been placed as well as it can be placed. The estimating error shows up somewhere else: where the line would have sat had a different ten months been collected instead.
Name the three sources of model error, and say which one more record actually reduces.
Which of the three does a longer record actually shrink?
Which of the three shrinks is provable rather than arguable, and proving it needs no new data at all. The ten months already on the record are enough. Fitting the line on every four of them, then on every five, and so on up to all ten, leaves at each size a collection of fitted slopes rather than one, and two questions can be asked of that collection: how far apart do the slopes sit, and how much is left over once each line has done its best on its own months.
The two questions answer differently, and that difference is the whole of the matter. Settle what each of them does as the number of months rises before reading the panel below.
As the number of months in a subsetSome of the items from a larger set, taken without altering any of them. Six of these ten months, whichever six, is one subset of them. rises from four to ten, what happens to the spread of the fitted slopes, and what happens to the spread of what is left over?
Refit on every subset of a chosen size, and watch two spreads part company
One control, seven settings. Moving the number of months refits every subset of that size from scratch. The tick strip redraws with one tick per fitted slope, and the two lines below carry a marker at whichever size is selected.
Educational illustration. The ten paired months were made up for teaching and are not a record of anything; every reading the panel prints is illustrative. The shape of the rule is held fixed at a straight line throughout, so the only thing moving is how many months each fit is allowed to see. Every subset of the chosen size is fitted, not a sample of them, and that is why the counts read 209, 252, 210, 120, 45, 10 and 1. At four months one of the 210 possible subsets is the four that all read 1.00 per cent, and a fit needs the input to vary, so 209 of them return a slope.
The counts explain why the top of the ladder is so quiet. There are 210 ways to pick four months out of ten, 252 ways to pick five, and then the number falls away: 210, 120, 45, 10, and finally one single way to pick all ten. By the time the record is the whole record there is nothing left to disagree with, and that is exactly why the spread of the slopes reaches zero there and stays.
Across every subset of every size, the average fitted slope comes out at exactly 1.5000. What does that rule out?
And which one refuses to shrink, however long the record gets?
Put the two columns side by side and the answer is not a matter of interpretation. As the number of months climbs from four to ten, the spreadHow far apart a collection of figures sit from one another. A collection whose members all sit close together has a small spread; one with members flung wide has a large one. of the fitted slopes falls at every single step: 0.6718, then 0.4648, then 0.3107, then 0.2093, then 0.1382, then 0.0811, then exactly 0.0000. Over the same seven steps the spread of what is left over goes the other way: 4.7968, then 5.0106, then 5.1175, then 5.1716, then 5.2002, then 5.2146, then 5.2202.
One of those columns is cured by a longer record and the other is not, and reading them against each other is the whole of this guide. The falling column is uncertainty about where the line sits. The rising column is the part of the outcome that the line was never going to reach, and giving it more months does not shorten its reach. A longer record only measures that reach more honestly.
One more reading of the falling column says something the spread on its own does not. At every size on that ladder the average of all the fitted slopes comes out at exactly 1.5000. At four months the individual fits run from a slope below zero up to one above four, a wild spread, and yet they are scattered evenly around 1.5000 rather than piled to one side of it. So the estimating on this record is uncertain without being biasedLeaning the same way every time rather than merely being uncertain. A figure that comes out too high on average is biased; one that comes out too high and too low by turns, in equal measure, is not., and those are two different faults with two different consequences.
The direction of the rising column surprises people and deserves a sentence of its own. The leftover gets bigger with more months because a line fitted to four months has four months to satisfy and two numbers with which to satisfy them, so it can hug them. Given ten months, it cannot hug all ten. The rising column does not show the model getting worse. A flattering figure is losing its flattery and settling on 5.2202, the amount the record actually contains.
How much record does it take to halve the uncertainty?
Granting that a longer record does shrink the middle source, the next question is the practical one: by how much, for how much effort. The answer is unkind, and it is unkind in a way that is fixed by arithmetic rather than by circumstance.
The standard errorA figure saying how much a fitted number would jump about had a different record of the same size and spread been collected instead. Settled in earlier reading and used here as it stands. of the slope on these ten months is 0.3014. With ten more months carrying the same spread of Nakshatra readings it becomes 0.2131. With thirty more, taking the record to forty, it becomes 0.1507. Halving that again takes a hundred and sixty months, more than thirteen years of monthly readings, and it buys 0.0753.
Four times the record halves this figure, sixteen times the record quarters it, and no row of that table moves the 146 by so much as a hundredth. The reason is that the standard error carries a square root of the record length underneath it, so effort goes in linearly and improvement comes out at the square root. Doubling is barely felt. The doubling buys the difference between 0.3014 and 0.2131, a reduction of about three tenths, for twice the collecting.
Halving the standard error of the slope, starting from ten months: how much record does that take?
How to Run Sensitivity Analysis on a Financial Model: what does running one actually involve here?
The instruction to run a sensitivity analysis turns up constantly, and its two halves are worth separating. The building of a model of a business, and the valuing of anything, is covered separately and is not attempted here. The procedure travels across, and the procedure is short enough to state in one sentence: change one stated input by a stated amount, hold everything else where it was, and record what the answer does. That is a sensitivity analysis, whether the inputs are the numbers inside a fitted line or the numbers inside anything else.
So run one. The fitted line reads 0.5000 plus 1.5000 times the Nakshatra change. At a Nakshatra change of 6.00 per cent the reading is 9.50 per cent. Now the two coefficientsThe numbers a fitting step works out and then keeps. Here there are two of them, the slope and the intercept, and every reading the line produces is built from the pair. are not facts; each carries a standard error. Move the slope up by one standard error, from 1.5000 to 1.8014, and leave the intercept alone. The slope is multiplied by six on the way through, so the reading at 6.00 per cent moves by 1.8083. Put the slope back and move the intercept instead, from 0.5000 to 2.1780. The intercept is added once and not multiplied by anything, so the reading moves by 1.6780.
The procedure has given two clean what-if answers, each traceable to a stated change. The procedure has not given a probability, a confidence statement or a range that anything is likely to fall inside. A sensitivity analysis teaches only as much as the what-ifs somebody chose to run, so the list of inputs varied belongs beside the results and not in a footnote.
What exactly does a sensitivity analysis do?
Why does varying one input at a time overstate the range?
Here is where the procedure gets misused, and the misuse is so ordinary that it usually goes unremarked. There are two swings, 1.8083 and 1.6780. The obvious next move is to add them. Both could go wrong at once, so the total exposure must be 3.4864, and quoting that feels careful rather than careless.
The honest combined figure is 2.2351. Adding the swings one at a time overstates the range by 55.98 per cent. Two separate things are wrong with the addition, and it is worth having both, because the usual telling gives only the second.
The first and larger one is geometry. Two uncertainties that are not the same uncertainty do not lay end to end. The two combine through their squares, the way two sides of a right angle combine into the side opposite. Squaring 1.8083 and 1.6780, adding, and taking the root gives 2.4670, not 3.4864. Squaring alone accounts for a little over four fifths of the overstatement, and it would apply even if the two fitted numbers had nothing to do with each other.
The second is that they do have something to do with each other, and it works in the analyst's favour here. The line has to keep passing through the middle of what it saw, so a line fitted a little too steep, on this record, tends also to have been fitted a little too low. The two errors lean against each other rather than stacking, and that lean carries 2.4670 down the rest of the way to 2.2351. The remaining fifth of the overstatement is exactly that.
The lean is easier to believe once it is drawn. Moving both numbers of the line by one standard error in the directions they actually tend to move together gives a slope up to 1.8014 and an intercept down to minus 1.1780. The mirror image of that is a slope down to 1.1986 and an intercept up to 2.1780. Both sit a full standard error away on both numbers at once, a combination that sounds severe. At a Nakshatra change of 6.00 per cent they land at 9.6304 per cent and 9.3696 per cent, a total width of 0.2608, against the 6.9727 that adding the swings both ways would have promised.
Two one at a time swings of 1.8083 and 1.6780 add to 3.4864, and the honest combined figure is 2.2351. Why is the sum so much bigger?
Which months is the model actually sensitive to?
Everything so far varied the numbers inside the rule. There is a second question hiding behind it, and on this record it turns out to be the more interesting one: how much does the answer depend on the individual months that went in?
Drop each month in turn and refitRun the fitting step again on a changed set of records, so the numbers it works out are allowed to come out different from the ones before.. Ten drops, ten refitted slopes. Six of the ten come back at exactly 1.5000, unmoved, and that is not a rounding. Four of the ten months carry a Nakshatra reading of 1.00 per cent, the average reading, and dropping a month sitting at the average leaves the tilt of the line untouched. Two more months, the fourth and the sixth, sit exactly on the fitted line with nothing left over, and dropping a month the line already passes through gives the line no reason to move.
Of the four that do move, two barely do: dropping the seventh month takes the slope to 1.5612 and dropping the ninth takes it to 1.4592. The other two carry everything. Drop the second month and the slope falls to 1.3163; drop the third and it rises to 1.6633; the whole swing of 0.3469 between those two lives in two months out of ten.
A small note on that 0.3469. Each figure was rounded to four places before printing, so subtracting the two as printed gives 0.3470. The gap worked out from the unrounded slopes is 0.3469, and that is the one quoted, on the standing habit of rounding once on the way to the screen rather than on both sides of a comparison.
Dropping any one of six of the ten months leaves the fitted slope at exactly 1.5000. What does that show about a sensitivity analysis run over the inputs?
What does an honest statement of error contain?
Somebody has to read this. A lender sizing a facility, an analyst handed a model built by somebody who has left, a household deciding whether next month's figure is worth acting on. None of them needs a smaller number. Each needs a number whose parts are named. A named part can be argued with and a single figure cannot.
Four things make the statement honest. Split the miss into the shape of the rule, the estimating, and the part that is neither. Give the length of the record and the standard errors on it. Say whether the swings were combined properly or laid end to end. Name the individual records the answer leans on. On this material that means saying out loud that two months of ten carry the whole swing.
Then there is the figure that usually gets left out, and on this record it is the largest one in this guide. The fitted reading at a Nakshatra change of 6.00 per cent carries 2.2351. A single new month at that same reading carries 5.6785. The part that will never shrink is more than twice the part that will, so a statement about one month that quotes 2.2351 has left out the biggest source of error there is.
At a Nakshatra change of 6.00 per cent the fitted reading carries 2.2351 and a single new month carries 5.6785. Which of the two belongs in a statement about one particular month?
The failure: asking for more data without saying what it is meant to fix
A model is built. The answer comes out uncertain and somebody says the obvious thing: get more data. The record doubles from ten months to twenty, ten months of waiting or of digging.
Here is what that bought. The standard error of the slope went from 0.3014 to 0.2131, a reduction of about three tenths. Nobody changed the shape of the rule, so the 72 it costs is exactly where it was. Nothing about collecting more months of the same input can reach the 146 that no rule reading only that one input could ever remove, so it did not move by a hundredth. The expectation was that a longer record would make the answer right; what it actually did was make one of three sources about three tenths smaller and leave the other two untouched.
The habit that fixes this costs one sentence, written before the request goes out rather than after the data arrives. The sentence states which of the three sources the extra record is expected to reduce, and by how much. If the answer is the third source, more months is not the fix and no amount of it will become the fix. If the answer is the second, the multiple has to be said out loud: four times the record for a halving, sixteen times for a quartering, with somebody then free to decide that is not worth ten months of waiting.
Covered elsewhere. How a miss is turned into a measure, and how the four common measures differ from one another, is covered separately. The two opposite ways of getting a rule's flexibility wrong, one bending too little and one bending too much, are covered separately. So is the practice of testing a rule across many cuts of one record at once. The individual misses month by month, and what their pattern says, were settled in earlier reading and are used here only as a total. Building a model of a business, or valuing anything at all, is covered separately.
Which numbers above were looked up, and which were worked out here?
Every figure above was worked out from the ten made-up months rather than looked up. The table says which arithmetic produced each group, so checking means redoing rather than trusting.
| What this guide used | How it was produced |
|---|---|
| The ten paired monthly changes of the Nakshatra unit and the Vasant unit | Made up for teaching and carried in unaltered, so the arithmetic here lines up with the arithmetic beside it |
| The ladder of every subset of every size, from four months to ten | Each subset of each size fitted in turn and the results counted: 209, 252, 210, 120, 45, 10 and 1 |
| The standard errors, 0.3014 on the slope and 1.6780 on the intercept | Worked out from the same ten months by the ordinary fitting arithmetic |
| The reading of 9.50 per cent and every swing around it | Put through the fitted line and its standard errors above |
| Any rate, threshold, period or standard | None appears above, because refitting a made-up record needs none |
The Nakshatra unit and the Vasant unit are invented.
Educational material. Not advice on any investment, tax, budget or market position.
