Regression Coefficients: The Intercept, the Slope, and What Each One Claims
A regression coefficient is a claim with units attached to it. The slope coefficient of 1.5000 says the Vasant unit moves 1.50 percentage points for every one percentage point the Nakshatra unit moves. The intercept of 0.5000 says the fitted line reads 0.50 per cent at an input of zero. The two claims are separate, and neither of them is the correlation of 0.8694.
The two numbers come out of twenty figures that were written down rather than measured. Ten monthly changes for one invented object, ten for another, chosen so the arithmetic would land on round answers instead of ragged ones. Take the twenty figures as the practice material. The method is what survives contact with a measured record: the way a slope is stated, the conditions an intercept needs, and the reason neither of them equals a correlation.
The line has already appeared. Fitting one straight line through ten paired months of two invented objects gave 0.5000 plus 1.5000 times the input. The fitting produced those two numbers, and reading them is the next job. A fitted line is a sentence with two blanks in it, and the coefficients are what somebody wrote in the blanks. Once each coefficient can be said out loud as a full sentence, with the units in place, most of the ways a regression gets misread stop being available.
Two invented objects do all the carrying here, and neither of them gets built again from scratch. Call the first one Nakshatra and the second one Vasant. Each is a traded unitsA deliberately vague word for anything whose price is free to wander. Who issued the thing, and whether it is a share, a fund or a bar of metal, makes not the slightest difference to the divisions done further down., made up for teaching and labelled so wherever it turns up. Each carries ten readings of its own monthly changeThe distance a price covered inside one month, measured against wherever that month opened. Plus 6.00 describes a month ending six per cent up on its own starting point, and minus 4.00 a month ending four per cent down on it., covering the same ten months in time order, which is exactly what makes them ten paired observationsReadings that belong together because they describe the same stretch of time, letting one be lined up against the other. Shuffle one column against its partner and not a single figure below this line survives.. Nakshatra was picked as the input and Vasant as the outcome, and a person made that choice before any arithmetic ran, rather than the arithmetic making it.
What is a regression coefficient, and what makes it a claim?
The physical form of a coefficient comes first. Somewhere in a printout there is a column of names and a column of numbers, and beside the name of an input sits a figure. The figure beside the input is the coefficient. A coefficient is one number. The coefficient looks exactly like every other number in the printout, is printed to the same width, and arrives with no more ceremony than a row total.
And a coefficient stripped of its units is not a small omission, it is an unfinished sentence. The figure 1.50, sitting alone, means nothing whatsoever. The bare figure is not large, it is not small, it is not strong, it is not weak, and there is no honest thing to be said about it. Written out in full it becomes 1.50 percentage pointsA unit for the gap between one percentage and another. A reading that climbs from four to six has climbed two percentage points, and calling that a two per cent climb would be a different and much smaller claim. The distinction is what keeps a coefficient from being read at the wrong size. of the Vasant unit for every one percentage point of the Nakshatra unit, measured across ten paired months. The full version can be checked, argued with, or rejected outright. The bare figure has not said anything, so nobody can disagree with it.
The point is not a finance point, so take it out of finance for a second. Suppose a household takes a careful look at a year of its own record and finds that in the months where the salary was Rs 1,000/- higher than usual, the spending was Rs 1,500/- higher than usual. The household has found a coefficient of 1.50, and stated in that form everybody in the room understands it, can question the record, and can name at least three reasons it might not hold next year. Now delete the rupees and say only that the relationship between salary and spending is 1.50. The same room has no idea whether that is alarming, comforting or meaningless. Nothing was rounded away and nothing was estimated badly. The units were dropped, and the sentence collapsed.
Before anything moves. A fitted line reads 0.50 plus 1.50 times the input. As the input changes from one month to the next, which of the two contributions to the reading changes, and which one does not? Commit to an answer now.
What does the Slope Coefficient claim, and in what units?
The slope on these ten months is 1.5000, and it is exact rather than rounded. Here is the whole claim, written as one sentence, and it is worth reading aloud once rather than skimming.
For every one percentage point the Nakshatra unit moves in a month, the Vasant unit moves 1.50 percentage points in the same direction, on average, across these ten months.
Three qualifications are doing real work in that sentence and each of them can be dropped by accident. IN THE SAME DIRECTION is carried by the sign being positive: had the slope come out at minus 1.50 the sentence would read in the opposite direction, and nothing else about it would change. ON AVERAGE is carried by the line being a line: not one of the ten months moved by exactly 1.50 times its input, and the sentence never claimed one did. ACROSS THESE TEN MONTHS is the boundary of the record. The line was fitted to ten months and it describes ten months. The line has been told nothing about an eleventh month, and reading it as though it had is a separate and much larger claim than the arithmetic made.
The other thing to hold onto is that a slope is a rate and not a size. It says how much of one movement goes with one unit of another movement. The slope says nothing at all about how big either movement was. A slope of 1.50 is entirely consistent with a record in which nothing moved much: if the input crawled between minus 0.1 and plus 0.1 per cent all year, and the outcome crawled at 1.5 times that, the slope is still 1.50 and the whole record is a flat line to the eye. A record that swung violently fits a slope of 1.50 just as comfortably. The slope is silent on the question, and any sentence that reads it as a statement about size has borrowed something the number never carried.
Watching one reading come apart makes the rate concrete. Set the Nakshatra unit at 6.00 per cent and the fitted line reads 0.50 plus 1.50 times 6.00, which is 0.50 plus 9.00, which is 9.50 per cent. Two contributions, and only one of them had anything to do with the input. Push the input to 11.00 per cent and the slope contribution grows to 16.50 while the intercept contribution sits at 0.50 exactly where it was. Pull the input to zero and the slope contribution vanishes entirely, leaving 0.50 as the whole reading.
Say the slope of 1.5000 out loud as a full sentence, with the units of both sides in it. Which of these is that sentence?
What does the Intercept say, and what does it need to say anything?
The intercept on these ten months is 0.5000, also exact. Its claim is narrower than most readers expect, and stating it narrowly is the whole job.
When the Nakshatra unit reads exactly zero, the fitted line reads 0.50 per cent. The intercept says that much, and it is not a shorthand for anything else.
The average of the Vasant unit is 2.00 per cent, and the intercept is not that average. The intercept is not a baseline the Vasant unit returns to. Nothing happening is not the same as one particular input reading zero, so the intercept is not what the Vasant unit does when nothing is happening. The intercept is the fitted reading at one specific setting of the input, and its usefulness depends entirely on whether that setting means anything.
Two conditions have to hold before an intercept tells a reader something, and they are separate conditions that fail separately.
The first: zero has to be a value the input can actually take. Some inputs can be zero perfectly well. A monthly change of zero is an ordinary month in which the price finished where it started, and nothing about it is strange. Other inputs cannot be zero at all. A count of days in a month cannot be zero. A person's height cannot be zero. An input that has been measured on a scale starting somewhere else entirely cannot be zero either.
The second: zero has to sit inside the stretch of the input the record actually covers. An input can be capable of reading zero while every month in the record sits a long way from it, and the second condition therefore fails quietly. If ten months all had inputs between 40 and 60, the line at zero is being read forty units beyond anything anybody observed, and the record has no evidence about that region whatsoever.
On these ten months both conditions hold, and comfortably. The Nakshatra unit reads 1.00, 6.00, minus 4.00, 11.00, 1.00, minus 9.00, 6.00, 1.00, minus 4.00 and 1.00 per cent across the record, so it runs from minus 9.00 up to 11.00. Zero is not merely inside that spreadHow far apart the readings in a record sit from one another. Settled in earlier notes and used here without rebuilding; here it appears only as a size that the slope carries and the correlation has divided out., it sits close to the middle of it, with four months below it and six above. The intercept of 0.50 therefore describes a kind of month this record has genuinely seen: an unremarkable one where the Nakshatra unit barely moved.
And that is a lucky property of this particular record rather than a general property of intercepts. It happened because the input is a monthly change, and monthly changes wander either side of zero as a matter of course. Change the input to something that lives entirely on one side of zero and the same arithmetic produces an intercept that nobody in the room should quote.
The intercept is 0.5000. What exactly is it the reading for?
What two conditions does an intercept need before it says anything, and does this record meet them?
When is an intercept a number about nothing?
Change the input and watch the same arithmetic produce something unusable. Suppose the input is not a monthly change at all but a count of days in a month. Every observation in that record sits at 28, 29, 30 or 31. The arithmetic runs perfectly happily. The arithmetic finds a slope, finds an intercept, and the intercept is the fitted reading at a count of zero days.
There is no such month. There never has been and there never will be. The intercept in that fit is the model's answer to a question nobody can ask, and the answer will be printed to four decimal places with a standard error beside it, in exactly the same typeface as every honest number in the same table.
Notice how thoroughly both conditions fail at once here, and that they fail independently. Zero is not a value a count of days can take, so the first condition is gone. And the record covers 28 to 31, so reading the line at zero means extrapolatingReading a fitted line at an input outside the stretch the record covers. The arithmetic never refuses to do it. The arithmetic cannot show that the line was ever tested out there, and it was not. twenty eight units past the nearest observation. Even if calendars were different and zero day months existed somewhere, this record would still have nothing to say about them.
The arithmetic will never announce that the number is about nothing. This is the part worth carrying away from the block. No warning is printed. No flag is raised. The fit does not slow down, the fit statistics do not deteriorate, and the intercept does not come out looking odd. The intercept is a coefficient, so it comes out looking like one. The only thing that catches the problem is a person reading the units of the input and asking two plain questions: can this thing be zero, and did this record ever go near zero? The check takes thirty seconds and there is no automated substitute for it.
The counter-case is stated in words rather than in figures. The shape is what matters: an input whose zero is impossible, an intercept computed anyway, and nothing in the output that distinguishes it from a good one.
An input is a count of days in a month. What is wrong with reading its intercept?
Why is the slope not the correlation?
Set the two figures side by side. On these ten months the slope is 1.5000 and the correlation is 0.8694. Same ten months, same two objects, same pairing, two different numbers. Neither is a rounded version of the other and neither is an approximation to the other. The slope and the correlation answer different questions.
The correlation asks a question about tightness: how nearly did the pair rise and fall together, scored against a fixed ceiling of plus one in one direction and a floor of minus one in the other, once the sizes have been taken out of both columns. The slope asks a question about magnitude instead: how far did the outcome travel for one unit of travel in the input, with the sizes left precisely where they were found. THE CORRELATION IS THE RELATIONSHIP WITH THE SIZES DIVIDED OUT, AND THE SLOPE IS THE RELATIONSHIP WITH THE SIZES LEFT IN. That single line will keep the two apart for good.
The sizes in question are the two spreads. The Nakshatra unit spreads by 5.7735 per cent across the record and the Vasant unit by 9.9610 per cent. The Vasant unit is the more violent of the two. The ratio of those spreads, 9.9610 divided by 5.7735, comes to 1.7253. The spread ratio is the entire difference between the two figures, and the arithmetic closes exactly: the correlation of 0.869413, multiplied by the spread ratio of 1.725302, is 1.500000.
The decimals matter more here than usual. Multiply the two rounded four decimal displays, 0.8694 times 1.7253, and a calculator returns 1.49998 rather than 1.5000. Nothing is wrong with the identity. Both displayed figures were rounded before they were printed, so the product of the rounded pair lands a whisker short of the product of the true pair. The identity is exact and only the printing is approximate. A measured record never allows that check, so the distinction is worth learning here, where the true answer is known.
There is a second way to see that they are different objects, and it takes about three seconds. Rewrite the record with every Vasant reading multiplied by two. The outcome now travels twice as far as it used to for an unchanged input, so the slope reports 3.0000. The correlation does not budge by so much as a decimal place. Doubling the outcome column also doubles that column's own spread, and the two doublings cancel each other inside the arithmetic, leaving 0.8694 exactly where it was. One figure kept the measurement units of both columns. The other threw them away. Keeping the units or throwing them away is not a nuance, it is the whole difference between the two.
So there are two sentences that can be built from these ten months, and they answer different questions. Using the correlation: the pair tracked each other tightly, scoring 0.8694 where the ceiling is one. Using the slope: one point of movement in the Nakshatra unit went with 1.50 percentage points of movement in the Vasant unit. Both are true. Neither implies the other. And a reader handed only the figure 1.50 has no way at all to tell which of the two questions somebody answered.
The slope is 1.5000 and the correlation is 0.8694. Where does the difference between them come from?
How firmly does the record pin each coefficient down?
Ten months is not many, and the two coefficients are not pinned down to anything like the same degree. A printout hides the difference, giving both coefficients to four decimals, and four decimals look equally confident whatever is behind them.
Each coefficient comes with a standard errorA measure of how much an estimate would jump about if the same exercise were repeated on a different record of the same size. Built in earlier notes and used here as a ready-made tool rather than rebuilt. of its own, and the two are wildly different. The slope of 1.5000 has a standard error of 0.3014, so the slope stands at about five times its own standard error: 1.5000 divided by 0.3014 is 4.9770. The intercept of 0.5000 has a standard error of 1.6780, so the intercept stands at about three tenths of its own: 0.5000 divided by 1.6780 is 0.2980.
Both of those come off the same residual standard error of 5.2202, computed on the eight degrees of freedomThe count of observations left over after the fit has spent some of them. Ten months went in and two were used up fixing the two coefficients, so eight remain. Settled elsewhere and used here as a given. that remain once ten months have paid for two coefficients. The reason the two standard errors come out so far apart is not the data being worse in one place than another. The slope is pinned by how far apart the inputs are, and these inputs are nicely spread from minus 9.00 to 11.00. The intercept is pinned instead by how close the inputs sit to zero, and the intercept inherits the wobble of everything else besides.
THE RECORD HAS A GREAT DEAL TO SAY ABOUT THE SLOPE AND ALMOST NOTHING TO SAY ABOUT THE INTERCEPT. Both are printed to the same number of decimals, and the printing hides the gap completely. Read that as a statement about how much the record supports each figure, and stop there. The comparison is not a test result, it does not clear the slope for any purpose, and it is not grounds for doing anything with either unit.
Both coefficients are printed to four decimals. Which one does this record actually pin down, and how would anybody know?
The ten months, and every coefficient figure read off them
Everything above comes out of one table, and here it is in full so that no figure needs anybody to take it on faith. The months sit in time order and they stay in time order. Sorting them into a tidier picture would destroy the evidence the order itself carries.
| Month | Nakshatra change | Vasant change | 0.50 plus 1.50 times the input |
|---|---|---|---|
| 1 | 1.00 | 3.00 | 2.00 |
| 2 | 6.00 | 18.50 | 9.50 |
| 3 | minus 4.00 | 2.50 | minus 5.50 |
| 4 | 11.00 | 17.00 | 17.00 |
| 5 | 1.00 | minus 3.00 | 2.00 |
| 6 | minus 9.00 | minus 13.00 | minus 13.00 |
| 7 | 6.00 | 6.50 | 9.50 |
| 8 | 1.00 | minus 1.00 | 2.00 |
| 9 | minus 4.00 | minus 7.50 | minus 5.50 |
| 10 | 1.00 | minus 3.00 | 2.00 |
| Mean | 1.00 | 2.00 | 2.00 |
| Figure | Value | Worked from |
|---|---|---|
| Slope coefficient | 1.5000 | 450 over 300, and it divides exactly |
| Standard error of the slope | 0.3014 | 5.2202 over the square root of 300 |
| Slope against its own standard error | 4.9770 | 1.5000 over 0.3014 |
| Intercept | 0.5000 | 2.00 less 1.50 times 1.00 |
| Standard error of the intercept | 1.6780 | 5.2202 times the root of one tenth plus 1 over 300 |
| Intercept against its own standard error | 0.2980 | 0.5000 over 1.6780 |
| Range of the input | minus 9.00 to 11.00 | The lowest and highest of the ten inputs |
| Spread of the Nakshatra unit | 5.7735 | The root of 300 over nine |
| Spread of the Vasant unit | 9.9610 | The root of 893 over nine |
| Ratio of the two spreads | 1.7253 | 9.9610 over 5.7735 |
| Correlation | 0.8694 | 450 over the root of 300 times 893 |
| Residual standard error | 5.2202 | The root of 218 over eight |
| The identity, closing | 1.5000 | 0.869413 times 1.725302 |
Slide the input and watch one contribution stand still while the other sweeps.
One control moves: the Nakshatra monthly change, across the whole stretch the record covers, from minus 9.00 to 11.00 per cent. The two coefficients are held exactly where the fit put them, at 0.5000 and 1.5000, so everything that moves is bought entirely with the input. The intercept block is under four pixels tall at true scale, and the smallness is itself the lesson, so the left column is drawn at true scale and the right column magnifies the little window around zero. The opening setting is an input of 6.00 per cent, giving a reading of 9.50 per cent made of 0.50 from the intercept and 9.00 from the slope, which are the same three figures written out in the worked table above.
Educational illustration on invented data. The two coefficients are frozen at 0.5000 and 1.5000 at every setting and only the input moves, so the panel decomposes one fitted line rather than refitting anything. The slider is bounded at minus 9.00 and 11.00 per cent because those are the lowest and highest inputs the record contains, and the line has been given no evidence at all beyond them.
A printout gives a coefficient of 1.50 and no units at all. What are the two quite different things it could be, and how would they be told apart?
What must be said alongside a coefficient, every time?
An analyst reading somebody else, a lender reading a submitted model, a household reading a statement about its own spending: all three hold a number and have to decide whether they have been told enough to use it. Six things travel with a coefficient, and each one has a specific failure attached to leaving it out.
The units of both sides. Without them the figure is an unfinished sentence, as the first block laid out, and nobody in the room can tell whether it is alarming or trivial. The number of observations. Ten paired months is a small record and the reader is entitled to know that before deciding how much weight the figure carries. Whether zero sits inside the input's range. This is the one line that separates an intercept worth quoting from one about a month with no days, and it costs nothing to state. The standard error, or at least whether the coefficient is large or small against it. Four decimals is not evidence, and two coefficients printed identically can be supported completely differently, as they are here.
Whether any other input was in the fit alongside it. A coefficient means something quite different when it is the only input than when it is sharing the work with others, and that shift is covered separately because it deserves its own treatment. And the direction as a word rather than as a sign. This last one sounds trivial and it is the one that actually goes wrong. A minus sign is one character wide. The sign survives a copy and paste badly, disappears in a column that has been reformatted, and is easy to lose entirely when a figure is retyped into an email. Writing the two units moved up together and down together takes four extra words and cannot be lost.
Name three things that must be quoted alongside any coefficient. Which set below, and what goes wrong when each is missing?
How this goes wrong, in a note where every number was correct
Nobody makes an arithmetic mistake in this one. An analyst writes a short note and the sentence in it reads that the Nakshatra unit and the Vasant unit have a relationship of 1.50. Every figure behind that sentence is right. The slope really is 1.5000 and it was computed properly.
A reader picks it up and takes 1.50 as a correlation. A relationship of 1.50 is not a correlation. Across these ten months the correlation comes out at 0.8694, and the entire gap between that figure and 1.50 is the ratio of the two spreads, 1.7253. The damage is that 1.50 read as a correlation is off the top of the scale a correlation lives on, so the reader now believes the two objects move in lockstep beyond anything the record could show. And nothing in the sentence flags the error, because the sentence never said which of the two quantities the 1.50 was. The word relationship covered both, and it should not have.
The second half of the failure survives the first half being corrected, and it is quieter and costs more. The same note prints the intercept as 0.5000 to four decimals beside a slope of 1.5000 to four decimals. The reader treats them as two equally solid findings, and the formatting invites exactly that. In fact the record supports the slope at about five times its own standard error and the intercept at about three tenths of its own. One of those two figures is well pinned down by ten months and the other is barely distinguishable from nothing at all, and the printout has given the reader no way to see it.
The fix is two habits and neither takes any skill. The units go with every coefficient, so nobody has to guess which quantity was meant. And each coefficient is quoted against its own standard error rather than against the number of decimal places it happened to be printed to. Four decimals is a formatting choice made by software. Four decimals are not evidence, and never were.
Who is standing behind a coefficient of 1.5000?
Nobody at all. A coefficient is a claim, and a reader is right to ask whose claim it is. Here it belongs to twenty invented figures and to the division that turned them into 1.5000. There is no institution to appeal to, no published series to fetch again, and no document that would settle a disagreement about it. A division a reader can redo in ninety seconds stands in place of a citation: 450 over 300, with nothing between the reader and checking it.
Two consequences follow and both cut in a direction worth knowing. The first is comfortable: every figure above is checkable by hand from the ten paired months, so a reader who suspects an error can go and find one without asking anybody's permission. The second is not. Because these twenty figures were built to be tidy, the slope came out at exactly 1.5000 and the intercept at exactly 0.5000. Measured records do not behave that way, and a reader who has only ever seen worked figures close exactly will be startled the first time a slope arrives as 1.4873 with a standard error wider than any of the worked figures above. The tidiness is a teaching convenience, and it is the one property of this worked case that will not survive contact with a real record.
| Source | Document | Site | |
|---|---|---|---|
| No maintained record was opened to build any figure above, so there is no organisation to name, no document to cite, no site to link and no date on which anybody looked. | |||
The Nakshatra unit and the Vasant unit are invented.
Educational material. Not advice on any investment, tax, budget or market position.
