Ordinary Least Squares: The Method and Its Assumptions
Ordinary least squares picks the one line, out of every line that could be drawn, whose squared misses add to the smallest total. On ten invented months that total is 218 and the line is 0.50 plus 1.50 times the input. The answer is not searched for but computed, and the computed line always leaves misses that add to zero and carry no relationship to the input.
A fitted line has already been established: 0.5000 plus 1.5000 times the input, ten paired months behind it, a sum of squared misses of 218, and an R squaredThe single headline figure for how much of the outcome's movement a fitted line accounts for, running from nothing to one. of 0.7559. Left unsaid was why that particular line and not one of the infinitely many others that could have been drawn through the same ten dots. Somebody made a choice. The choice buys something, and it carries six conditions under which the resulting number means what it looks like it means.
Everything below runs on two made up objects. The Nakshatra unit is a traded unitA stand in for any object with a price attached, so that a stretch of its life can be written down as a single figure. whose monthly changeHow far a price travelled across one month, set against where it stood when the month opened. Rs 200/- climbing to Rs 206/- across a month gets written down as 3.00 per cent. is the input. The Vasant unit is a second one, and its monthly change is the outcome the line is trying to land on. Ten months, ten paired observationsOne row holding two readings taken from the same occasion, so the two can be lined up against each other. Ten months here give ten rows, and nothing may be shuffled between rows., kept in the order the months happened.
| Month | The Nakshatra unit, the input | The Vasant unit, the outcome |
|---|---|---|
| 1 | 1.00 per cent | 3.00 per cent |
| 2 | 6.00 per cent | 18.50 per cent |
| 3 | minus 4.00 per cent | 2.50 per cent |
| 4 | 11.00 per cent | 17.00 per cent |
| 5 | 1.00 per cent | minus 3.00 per cent |
| 6 | minus 9.00 per cent | minus 13.00 per cent |
| 7 | 6.00 per cent | 6.50 per cent |
| 8 | 1.00 per cent | minus 1.00 per cent |
| 9 | minus 4.00 per cent | minus 7.50 per cent |
| 10 | 1.00 per cent | minus 3.00 per cent |
| Average | 1.00 per cent | 2.00 per cent |
What does ordinary least squares actually do?
Ordinary least squares does one thing, and the whole method fits inside a single sentence. Out of every straight line anybody could draw across those ten dots, ordinary least squares hands back the one for which the misses, squared and added up, come to the smallest total available. The specification stops there. Not the truest line, not the most sensible line, not the line an expert would draw. The line with the smallest total of squared misses, and nothing about the arithmetic argues that this is the total worth making small.
Here is the everyday version. A caterer works twelve weddings a season and wants a single rule turning the guest count into kilograms of rice to cook. Whatever rule gets written down, it will be wrong at every wedding by some amount. Two rules cannot be compared until each of them has a score, so somebody has to decide how those twelve wrongnesses combine into one. Suppose the decision is that being out by four kilograms counts sixteen times as badly as being out by one. The sixteen to one ruling is a decision. A person took it, it could have gone the other way, and once taken it crowns a different rule than a gentler decision would have crowned. Least squares takes exactly that decision, and then never mentions it again.
So the honest way to read a fitted line is as the winner of a competition whose scoring rules were fixed in advance. Change the scoring and a different line wins. To see that the competition is real, hold the line's shape and slide the slope. At a slope of 1.30, refitting the starting value to 0.70 so the line still passes through the middle of the record, the ten squared misses add to 230.00. At 1.40 with a starting value of 0.60 they add to 221.00. At 1.50 with 0.50 they add to 218.00. Then the total climbs back: 221.00 at a slope of 1.60 and 230.00 at 1.70.
| Slope | Starting value | Squared misses, added up | Cost of not being the winner |
|---|---|---|---|
| 1.30 | 0.70 | 230.00 | 12.00 |
| 1.40 | 0.60 | 221.00 | 3.00 |
| 1.50 | 0.50 | 218.00 | nothing, this is the winner |
| 1.60 | 0.40 | 221.00 | 3.00 |
| 1.70 | 0.30 | 230.00 | 12.00 |
A slope of 1.50 is the bottom of that list, and moving one tenth in either direction costs exactly 3.00 in squared misses, the same 3.00 going up as going down. The symmetry is not a coincidence and it is not rounding. The total is a perfect bowl in the slope, and its floor sits at 1.50. The walls of the bowl climb with the square of the distance from its floor, so moving two tenths costs 12.00, four times three.
State what ordinary least squares does, in one sentence, without using the word best. Which of these is the sentence?
Moving the slope from 1.50 to 1.60 raises the squared misses from 218.00 to 221.00. What does that 3.00 establish?
Why squared misses rather than plain ones?
The method never answers the question out loud, and it is worth asking directly. There are ten misses. Some are above the line and some below. The ten misses have to be turned into one number. Why square them?
There are three answers and they are not the same kind of answer. Two of them are judgements somebody made and only the third is mathematics.
The first reason is that squaring stops the misses cancelling. The ten misses as they stand add to zero, every single time, for reasons set out below. A total that is always zero cannot rank two lines. Squaring turns every miss into a positive contribution, so a line that is wrong in both directions is scored for being wrong in both directions. The cancelling reason, though, does not require squaring specifically. Dropping the sign and adding the sizes does the same job. So the first reason rules out the raw total and does not pick squaring out of the field.
The second reason is that squaring makes a big miss count more than proportionately, and that is a value judgement wearing a formula. A miss of 9.00 per cent is three times a miss of 3.00 per cent. Square them and the first contributes 81 while the second contributes 9, nine times as much. Whether one terrible month should count nine times a mediocre one, or three times, or twenty seven times, is a question about consequences in the world. The rent is due either way, so a household with one salary treats one month of no income as far worse than three months of income cut by a third, and it is right to. A vegetable seller with a stall outside one office building might feel the opposite. A series of thin weeks is what actually empties the cash box. Neither view is a fact about arithmetic. Squaring picks one of them and stops asking.
The third reason is the only one that is mathematics rather than opinion: squaring makes the answer come out in closed form, and the alternatives do not. Scored by squares, the winning line can be written down in two divisions, worked below. Scored by plain sizes, there is no such formula: the answer has to be hunted for, and on some records more than one line ties for the win. Closed form is why least squares became the method everybody learns first. Convenience of solution is a real reason and not a shameful one, though it is not a reason about the world.
The honest counterweight has to come next. On this particular record of ten months, the rule that makes the average size of a miss smallest lands on precisely the line least squares already chose, at an average miss size of 3.60 per cent. Two criteria that disagree in principle about which month matters have agreed here about which line wins. The two rules disagree about how to weigh the months and they happen to agree on this record about which line to pick, and both halves of that sentence have to be said. On a different ten months they can and do part company.
Of the three reasons for squaring, one is mathematics rather than a judgement about what matters. Which one?
The answer is worth predicting before the panel below is touched. As a miss is scored more harshly, what happens to the largest month's share of the total score?
Change how harshly a miss is scored and watch the same ten misses redistribute the blame.
One control, and it does not move the line. The fitted line stays at 0.5000 plus 1.5000 times the input at every setting, and the ten misses stay exactly as they are. The control moves the power each miss size is raised to before the shares are worked out. At a power of 1.00 each month is scored by the plain size of its miss. At 2.00 it is scored by the square. Squaring is what ordinary least squares does, and 2.00 is the default setting. The pale outlines behind the bars are the shares at a power of 1.00, left in place to show how far each month has travelled from there.
Educational illustration on an invented record. The line is held at 0.5000 plus 1.5000 times the input at every setting and cannot be moved from here; only the scoring changes. The ten shares add to 100.00 per cent at every setting. Raising every size to the same power cannot overtake one month with another, so the ranking of the months by size never changes.
On this record the line that minimises the average miss size turns out to be the same line least squares picked. Does that mean the two criteria always agree?
How does the method find the answer instead of hunting for it?
The bowl in the first figure invites a picture of somebody trying slopes one after another and keeping the best score. Trying slopes one after another is not what happens, and the difference matters more than it sounds. The winning line is computed in two divisions, from quantities read straight off the ten months, and no candidate line is ever tried.
Step one is the slope. Take how much the two units move together, measured as the paired departures from their own averages multiplied and added. On this record that comes to 450. Take how much the input moves on its own, measured as its departures squared and added. The input's own movement comes to 300. Divide: 450 over 300 is 1.5000, exactly. Step two is the starting value. The winning line always passes through the point made by the two averages, so put the input's average of 1.00 into the line and demand that it return the outcome's average of 2.00. The demand means 2.00 less 1.50 times 1.00, exactly 0.5000.
Both figures land on a clean decimal because the twenty numbers were composed to make them land there. On a measured record they would not, and the arithmetic would be identical. The exactness is worth taking one thing from: no estimateA figure worked out from the rows available that stands in for something that cannot be observed directly. What estimation costs, and how far an estimate can be trusted, are covered under sampling. on this record came out of a search that might have stopped in the wrong place. Every one came out of a division.
What two conditions does the answer always satisfy?
Take the fitted line back to the ten months and write down what it missed by each time. The fitted valueWhat the line says the outcome should have been at a given input, worked out before the actual reading is set beside it. for month one is 0.50 plus 1.50 times 1.00, or 2.00, and the outcome was 3.00, so the miss is 1.00. Do that ten times and the misses are 1.00, 9.00, 8.00, 0.00, minus 5.00, 0.00, minus 3.00, minus 3.00, minus 2.00 and minus 5.00.
Two things are now true about that list, and they are true of every ordinary least squares fit ever performed, on any data whatever.
First, the misses add to exactly zero. One plus nine plus eight plus nothing, less five, plus nothing, less three, less three, less two, less five. The positives come to eighteen and the negatives come to eighteen. Zero. Second, the misses carry no relationship to the input at all. Weight each miss by how far its month's input sat from the input's own average of 1.00 and add those up, and that also comes to zero. Neither of these is luck and neither is evidence. Both are consequences of how the line was picked: the two divisions in the previous section are precisely the arithmetic that forces them.
Now the trap that both conditions set. Both conditions are properties of the method rather than findings about the record, so both hold even when the straight line is completely the wrong shape for the data. Take seven readings whose scatterA plot with one dot per row, placed left to right by the first reading and up and down by the second, so the shape of a relationship can be looked at rather than summarised. forms an obvious U. Fit a straight line by least squares and it comes out perfectly flat, sitting at the average outcome of 4.00, a visibly wrong description of a U. Its misses are 5.00, 0.00, minus 3.00, minus 4.00, minus 3.00, 0.00 and 5.00. The seven misses add to zero. Weighted against the input they also add to zero. Both conditions pass with full marks on a fit that any eye can see is wrong.
A line is fitted, the misses are added, and the total comes to exactly zero. What has that established?
What does the method assume about the data it is handed?
Six conditions. Not one of the six is needed for the arithmetic to run, and that is the source of nearly all the trouble. Hand ordinary least squares any two columns of numbers at all and it will return a slope and a starting value without complaint, so the assumptions are conditions for the answer to mean what it appears to mean, never conditions for getting one.
| The condition | What it says | What it looks like when it fails |
|---|---|---|
| A straight shape | The real relationship between the two is a straight line, not a curve or a step | Misses plotted against the input form an arch or a bowl instead of a formless cloud |
| No pattern left over | Whatever the line failed to catch is not itself related to the input | The misses drift systematically as the input grows, so the line is too low at one end |
| Independent misses | Knowing one month's miss reveals nothing about the next month's | Long runs of misses on the same side of the line when the record is put back in order |
| Similar sized misses | The misses are about as large at one end of the input's range as at the other | A fan shape, tight on the left and spreading wide on the right |
| A cleanly measured input | The input column carries no measurement error of its own | A recorded input that is itself an approximation, so the slope is pulled toward zero |
| No near copy among the inputs | No input is very nearly a repeat of another input already in the fit | Two columns that move together almost perfectly, and coefficients that swing wildly |
The middle column and the right column read as a pair. Each condition is a claim about the world that the numbers alone cannot settle, and each has an artefact that can actually be looked at. The last one only bites when there is more than one input, and it is demonstrated separately. The first four are checkable immediately on any record, and checking them is looking rather than computing.
Which pairing of a condition with the thing to look at when checking it is right?
Which of those assumptions do these ten months break?
The third one, and there is no need to hedge about it. Put the ten misses back into the order the months happened, the order they have been in all along, and they do not look like ten independent surprises. The ten misses look like a story. The first three months miss above the line, in a row. Then a miss below, then an exact hit, then four misses below the line, in a row. In five of the nine consecutive pairs, both months land on the same side.
The claim deserves a figure, so here is one. The ten misses carry a lag one autocorrelationOne figure for how strongly each entry in an ordered list resembles the entry immediately before it. How it is worked out, and what to do when it is large, are covered under autocorrelation. of 0.4862, and R squared for the same fit on the same ten months is 0.7559, a perfectly respectable looking result. The 0.4862 is computed under autocorrelation, and is quoted here so that the two numbers stand next to each other.
Why can R squared not see any of it? Because computing R squared throws the order away. R squared adds up squared misses, and addition does not care which month came first. Shuffled into any arrangement at all and recomputed, the ten months give 0.7559 every time. The only evidence against the fit lives entirely in the ordering, and the headline number is computed in a way that destroys the ordering before it starts. The blindness is not a flaw in R squared. R squared was never asked about the order. The flaw is in reading R squared as though it had been.
The record breaks a stated condition, and the fit's own headline figure is structurally incapable of noticing. How big a problem the dependence is, and what anybody does about it, is set out under autocorrelation.
R squared is 0.7559 and the misses carry a lag one autocorrelation of 0.4862. Which condition has failed, and why is R squared silent about it?
What is ordinary least squares not?
Four denials, and the method is stronger for each of them.
Least squares is not a test of anything. Fitting a line returns two numbers and passes no verdict. Whether a slope of 1.5000 is distinguishable from no relationship at all is a separate question with separate machinery behind it, and the fit does not answer it by existing. Least squares is not a method for choosing which inputs to use. Hand it a column of anything and it will fit that column. The decision about which columns deserve to be in the room is made by a person, before the arithmetic runs, on grounds the arithmetic has no access to.
A fitted line is not a forecast. The line describes the ten months it was fitted on. Reading it as a statement about a month that has not happened adds a claim the fitting never made, and that claim needs its own defence. And it is not a method that shrinks a fit, penalises a fit, or resamples a record to pick among several fits. Shrinking, penalising and resampling all exist and are useful. All three are a different subject with their own conditions, covered under model building.
A method that shrinks the fitted coefficients toward zero to improve how a fit behaves is proposed. What is the right thing to say about where that belongs?
What should be checked before a fitted line is quoted?
Checking is the part a lender, an analyst or anybody with a spreadsheet actually does, and it takes about ten minutes. The cheapest checks are the least informative, and doing them first creates exactly the false comfort the two closing conditions invite, so the order matters.
The first check is that the misses add to zero, and what it establishes should be discounted immediately. The zero is a receipt that the arithmetic ran, in the way that a till receipt confirms the till worked rather than confirming the shopping was needed. Next comes plotting the misses against the input and looking at the shape. A formless cloud is what a satisfied straight line assumption looks like; an arch or a bowl means the shape is wrong and no amount of refitting a line will fix it. Then the misses are plotted against the clock, in the order the rows arrived, and examined for runs on one side. The clock plot is the check this record fails.
Then, if there is more than one input, check whether any two of them move together almost perfectly. A near copy destroys the individual coefficients while leaving the overall fit looking untouched. And last, the one people skip: record which condition went unchecked. The condition nobody checks is the one that fails. A sampleThe rows actually to hand, as against every row the world could have produced. The distinction, and what it costs, is covered under sampling. of ten months cannot be checked as thoroughly as a sample of ten thousand, and saying so out loud in the same paragraph as the slope is what separates a quoted figure from an oversold one.
The two checks that could not have come out any other way
Here is the mistake in the shape it actually takes. Somebody fits the line, adds the ten misses and gets zero. The same somebody weights the misses by the input and gets zero again, writes in the working note that the misses are unbiased and unrelated to the input, concludes that the conditions are satisfied, and quotes the slope of 1.5000 and the R squared of 0.7559 to three people who then build on both.
Both checks passed and neither could have failed. Ordinary least squares forces both of them to hold on any two columns of numbers put into it, including columns whose true relationship is a curve, a step or nothing at all. The flat line through the U in the fourth figure passes both, and it is wrong about everything. The checker did not verify a condition. The checker verified that the software had run.
Meanwhile the condition that does fail on this record, independence, was never looked at. Looking at it needs a different plot: the misses against the clock rather than against the input. And it fails at 0.4862, on the same ten months. The cost is not that one figure is wrong. The cost is that every figure downstream is quoted with more confidence than it earned, and the confidence rests on two checks that were incapable of coming out badly.
The fix is a distinction, not a technique. The two conditions are a receipt that the arithmetic ran. The six assumptions are checked by looking, at the misses against the input and at the misses against the clock, and that is work the arithmetic will never do on the analyst's behalf, however many decimal places it prints.
Covered elsewhere. The slope and the starting value each claim something on their own, and what each claims is set out separately, as is how the misses are read one at a time, and as is the reading that emerges when those misses are set back against the clock and the ordering is taken seriously. Methods that shrink a fit, penalise it, or resample a record to choose among several fits are a different subject with their own conditions, covered under model building.
Whose authority backs the figures here?
Nobody's. Ordinary least squares never asks where its ten pairs came from. The method takes whatever two columns are handed over and returns the line with the smallest squared misses. A citation would point somewhere else; these sums point at themselves.
The method is centuries old and more than one claimant is put forward for it in more than one country. A name given without a text to check it against is decoration rather than a credit, so the mechanism is described and no discoverer is named.
| Quantity | How it was produced | The check that would catch a mistake in it |
|---|---|---|
| The ten paired monthly changes | Written down in advance, in the order shown, chosen so the two closing conditions hold to the last decimal | None. The pairs were never measured, so nothing outside can contradict them |
| The slope 1.5000 and the starting value 0.5000 | 450 divided by 300, then the line pushed through the pair of averages | Redo both from the table at the top. A different answer means a pair was typed wrong |
| The five totals 230.00, 221.00, 218.00, 221.00 and 230.00 | Ten squared misses added at each of five slopes, the starting value refitted every time | Add the ten squares by hand at any one of the five settings |
| The shares of 25.00 and 37.16 per cent | One miss size raised to a power, divided by the same total across all ten months | Set the panel to a power of 1.00 and then 2.00 and read the largest bar |
| The lag one figure of 0.4862 | Quoted, not worked out. The figure is computed under autocorrelation, where reading the misses in order is the subject | Not available here; that working is covered separately |
| Least squares as a method | Described from its mechanism only | Not applicable. No discoverer is named and none is needed |
The Nakshatra unit and the Vasant unit are invented.
Educational material. Not advice on any investment, tax, budget or market position.
