Multicollinearity: When Inputs Say the Same Thing Twice
Multicollinearity is what turns up when two inputs carry nearly the same information. On ten invented months, adding a near copy of the Nakshatra unit drags its coefficient from 1.5000 down to minus 1.0000 while the copy walks off with 2.5000. R squared meanwhile shifts from 0.7559 to 0.7560, so the quality of the fit cannot see any of it.
Earlier notes fitted one straight line through ten paired months and got 0.5000 plus 1.5000 times the input. The notes on the method behind that line listed the conditions it needs, and one of them was that no input may be a near copy of another. The near-copy condition has been sitting there like a smoke alarm nobody has tested. Testing it takes one step. A second input, almost the same thing as the first, is added, the line is refitted, and what breaks and, more importantly, what does not, comes straight out of the arithmetic.
Three invented objects do the carrying. Two of them have already appeared. The Nakshatra unit is a traded unitA deliberately loose word for anything with a price that is free to move. Whether it is a share, a fund, a bond or a bar of metal changes nothing about the divisions done below, so the notes decline to specify. whose ten monthly readings are 1.00, 6.00, minus 4.00, 11.00, 1.00, minus 9.00, 6.00, 1.00, minus 4.00 and 1.00 per cent. The Vasant unit is the outcome, at 3.00, 18.50, 2.50, 17.00, minus 3.00, minus 13.00, 6.50, minus 1.00, minus 7.50 and minus 3.00 per cent. Each reading is a monthly change, meaning the distance between where a price finished a month and where it began that month, put as a percentage of the beginning. Because the two columns describe the same ten stretches of time they line up as ten paired observationsTwo readings tied to each other by covering one and the same stretch of time, which is the only thing that permits either to be weighed against the other. Break the pairing and every quantity worked out below this line stops meaning anything., and they are held in time orderThe arrangement in which the readings actually happened, first month first. Sorting them by size would make a neater picture and would destroy the only evidence some of these notes have left to spend. throughout.
What is multicollinearity, and why define it by what it does?
Most definitions of this subject start from the cause: two or more inputs are strongly related to each other. The cause is true enough, and it is still the wrong end to hand a reader first. Naming the situation says nothing about why anyone should care. The damage is the better starting point.
Multicollinearity is the condition in which the coefficients of a fit become unstable and individually meaningless while the fit itself stays exactly as good as it was.
Read the second half of that sentence twice. A problem that leaves the quality of the fit untouched is a problem no measure of fit can ever report. That is the entire reason this has a name of its own. If it degraded the fit, nobody would need a word for it: the fit would degrade, somebody would notice, and the matter would end there. Multicollinearity does not degrade the fit. Multicollinearity sits underneath a perfectly healthy looking report and quietly rewrites what the report claims.
Here is the everyday version. Two people share a flat and both of them pay for the same electricity bill of Rs 4,000/- one month, one by transfer and one in cash to the landlord, and both write it into the shared accounts sheet. At the end of the month the sheet balances to the rupee. Total spending is right. Total contributed is right. There is no error in any total, so nothing anywhere reports an error. The broken question is a different one: how much did each person pay? The sheet can produce an answer to that, and it can produce several, and it has no way of showing which one is real. The totals are fine and the attribution is destroyed, and those two facts are compatible because they are different questions. Totals stand for the quality of a fit and attribution for the coefficients, and that is the subject here.
Define multicollinearity by what it does rather than by what causes it. Which of these is the working definition?
What does a near copy actually look like?
The third object is the Chandana unit, invented for this guide. Its ten monthly changes are 1.00, 5.95, minus 3.95, 11.05, 0.95, minus 9.00, 5.95, 1.05, minus 4.05 and 1.05 per cent. Its average is 1.00 per cent, and that figure is the Nakshatra unit's average as well. Set the two columns beside each other and the Chandana unit is the Nakshatra unit moved by five hundredths of a percentage pointThe unit for a gap between two percentages. A reading that goes from 6.00 to 5.95 has moved five hundredths of a percentage point, and calling that a five hundredths of a per cent move would be a different and much smaller statement. in eight of the ten months, and left exactly alone in months one and six.
The two correlate at 0.999967. On a measure that stops at one, that is about as close to being the same column as two different columns can manage.
The duplication is obvious on sight. Nothing in the argument depends on the duplication being hard to spot. The interesting question is not whether a person could catch this by eye. The interesting question is how the arithmetic behaves when nobody catches it, and the arithmetic behaves identically whether the duplication is this blatant or subtle enough that no reader would ever see it. Two inputs that are both, say, some measure of size, gathered by different departments in slightly different ways, will do exactly what the Chandana unit is about to do, and nobody will be looking at the two columns side by side when it happens.
What happens to the coefficients when the copy is added?
Now the one step everything so far has been leading to. The Vasant unit stays the outcome, the Nakshatra unit stays an input, the Chandana unit joins as a second input, and the line is refitted. Nothing else changes. Not one of the thirty figures is edited. The same method that produced 1.5000 produces the new answer.
The coefficient on the Nakshatra unit comes out at minus 1.0000. The coefficient on the Chandana unit comes out at 2.5000. The intercept does not move at all: it is 0.5000 before and 0.5000 after.
The sign flipped. Before the copy arrived, the fit said the Nakshatra unit moves the Vasant unit up. After, on the identical ten months, it says the Nakshatra unit moves the Vasant unit down. Nothing was measured differently, nobody made an arithmetic slip, and no assumption was quietly changed. A column that duplicated an input already present was handed to the method, and the method returned an answer in which the original input is now working in the opposite direction.
The arithmetic has not gone mad, and what stayed true shows why. Minus 1.0000 plus 2.5000 is exactly 1.5000, the very figure the single input was saying on its own. The pair has not lost the relationship. The pair has kept the relationship perfectly and then split it between two columns, and because the two columns are near copies there is no arithmetic reason to prefer one split over another. Push a large positive weight onto one and a matching negative weight onto the other and the two almost cancel; whatever survives the cancellation is the relationship. The fit found one of the enormously many ways of doing that, and the way it found happens to put a minus in front of the first column.
The household version again. Two people paid one Rs 4,000/- bill between them. Any split at all satisfies the total: 4,000 and nothing, 1,000 and 3,000, or minus 6,000 and 10,000 if one of them was reimbursed along the way. The sum is nailed down and each share is not, and the sheet holds no information that would settle it. Only the pair's claim is real; what either column says alone is an artefact of the split.
The coefficient on the Nakshatra unit goes from 1.5000 to minus 1.0000. What changed in the data to make that happen?
The two coefficients are minus 1.0000 and 2.5000. What do they add to, and what does that sum show?
What happened to the fit while all that was going on?
Almost nothing, and this is the half of the subject that people skip. Put the two fits side by side and read the bottom rows rather than the top ones.
| What the fit reports | One input | Both inputs |
|---|---|---|
| Intercept | 0.5000 | 0.5000 |
| Coefficient, the Nakshatra unit | 1.5000 | minus 1.0000 |
| Coefficient, the Chandana unit | not in the fit | 2.5000 |
| Sum of squared misses | 218.000 | 217.875 |
| R squared | 0.7559 | 0.7560 |
| Adjusted R squared | 0.7254 | 0.6863 |
| Largest move in any single fitted reading | 0.125 of a percentage point | |
The sum of squared misses drops from 218.000 to 217.875, a saving of an eighth. R squared rises from 0.7559 to 0.7560, a move in the fourth decimal place. And not one of the ten fitted readingsThe number the line produces for a given month, as opposed to the number that month actually recorded. Built in earlier notes and used here only as a thing that did or did not move. moves by more than 0.125 of a percentage point, on an outcome whose own spreadA single figure standing for how widely a set of readings is scattered, worked out in earlier notes and taken here as already built. Its whole job here is to be the yardstick a movement gets called large or small against. is 9.96 per cent. In fact every fitted reading moves by exactly 0.125 or by exactly nothing, because the change is 2.5000 multiplied by the five hundredths of a point separating the two inputs.
The coefficients moved by two and a half points. The predictions moved by an eighth of one.
The whole symptom fits in one line, and anyone who watches only the sign flip has seen half the subject and the less useful half. If the fit had fallen apart at the same time, everyone would find this on the first pass. The fit did not fall apart. The fit got a hair better.
One number on the report did move honestly, and it is the one that charges for inputs: adjusted R squared fell from 0.7254 to 0.6863. Adjusted R squared fell because the second column was billed for and delivered nothing worth the bill. Why the adjusted measure can fall while the plain one cannot is settled in the opening notes on measuring fit, where the comparison is worked through on this very same pair of columns. The fall of 0.0391 is 391 times the size of the rise of 0.0001, and it is the single signal on the whole report pointing at the trouble.
R squared moves from 0.7559 to 0.7560. Why could it not have moved much, whatever the data had been?
Adjusted R squared falls from 0.7254 to 0.6863 while plain R squared rises. What is the adjusted one charging for?
As the copy is allowed to differ more and more from the original, becoming less and less of a duplicate, what happens to R squared?
Widen the gap between the two inputs and watch two coefficients walk apart while one fit refuses to move.
One control moves: how far the Chandana unit is allowed to differ from the Nakshatra unit, from five hundredths of a percentage point up to a full point. The Nakshatra readings never change and the Vasant readings never change. Only the size of the wobble does. With the trail turned on, every setting visited leaves a mark, so after a minute of dragging the coefficients stand scattered across the scale and the fit marker piled up in one place. The opening setting is a wobble of 0.05, giving minus 1.0000 and 2.5000, the same two figures written out in the table above.
Educational illustration on invented data. The Nakshatra and Vasant readings are frozen at every setting and only the size of the wobble moves. The two coefficients add to 1.5000 at every setting and R squared reads 0.756019036954 at every setting, and both of those are consequences of how the third column is built rather than coincidences worth remarking on.
Why can R squared not see any of this?
State the reason rather than repeating the observation. R squared is worked out from the fitted readings and from nothing else: it compares how far the fitted readings sit from the actual ones against how far the average sits from the actual ones. The fitted readings barely moved. Therefore R squared barely moved. R squared could not have reported this trouble whatever the data had been. The trouble does not live in the fitted readings, and R squared has never looked anywhere else.
The panel above makes that concrete in a way a sentence cannot. Widen the wobble and the two coefficients walk right across the scale, from minus 1.0000 and 2.5000 to 1.3750 and 0.1250. The sum stays at 1.5000. And R squared reads 0.756019036954 at every single setting, identical to twelve decimal places. The fitted readings at every setting are the same ten numbers. The fit is not roughly the same. The fit is the same.
So a reader who checks the quality of the fit and stops has checked the one quantity that was never able to show the problem. The failure is one of process rather than of arithmetic. The arithmetic did what it was asked. The order in which somebody looked at things is what went wrong, and no amount of care with the divisions can fix an order-of-looking problem.
A caretaker checks a building by walking the corridors and looking for damage. The corridors are spotless every week. Meanwhile the labels on the two water tanks have been swapped, so every reading anyone takes about which tank is which is wrong. The corridors stay spotless throughout: corridor condition and tank labelling are not related. Checking harder does not help. Checking something else does.
How is it measured, if the fit cannot report it?
By turning away from the outcome entirely and looking only at the inputs. The variance inflation factor takes each input in turn, asks how well that input can be predicted from all the others, and converts the answer into how much wider the coefficient's own uncertainty has become because of the overlap. Take the share of an input the other inputs can already account for, subtract it from one, and divide one by what is left. When nothing overlaps, the factor is one and no widening has happened.
Here the Nakshatra unit can be predicted from the Chandana unit almost perfectly, so the leftover is almost nothing, so the division by almost nothing produces something enormous. The factor reads 15,001 on this record, and the rule of thumb people carry around is that anything above ten is worth a second look. The rule of thumb is a habit of practice rather than a limit anyone sets, and it is treated as one here; the point is not the ten, it is that the reading here is fifteen thousand and the fit still moved only in the fourth decimal.
The factor is worked out from the inputs alone and never once looks at the outcome, and that is exactly why it can see what R squared cannot. R squared asks how close the fitted readings came to the actual ones. The variance inflation factor asks a question about the input columns that would have the same answer if the outcome column were deleted altogether. Two different questions, and the trouble lives in the second one.
The widening can be watched directly. Fitted on its own, the Nakshatra coefficient carries a standard errorHow far a computed figure would be expected to land from where this particular record placed it, had a different record of the same size turned up instead. Built in earlier notes; here it is simply the width that gets inflated. of 0.3014. Fitted alongside the copy, the same coefficient carries a standard error of 39.4506. The standard error is about 122 times wider from the duplication itself, and 130.90 times wider once the small change in the leftover spread is counted as well. A coefficient of minus 1.0000 with a standard error of 39.4506 beside it is a number the record has essentially nothing to say about, and the fit prints it to four decimal places with a straight face.
The variance inflation factor reads 15,001. What is it worked out from, and why can it see something R squared cannot?
How unstable do the coefficients actually become?
Two and a half points of movement, from one added column, is already alarming. Here is the diagnostic that shows how little it would take to move them again. Drop one month from the record, refit on the remaining nine, and write down the coefficient. Do that ten times, once for each month, and look at the range.
With the Nakshatra unit as the only input, the slope stays between 1.3163 and 1.6633 whichever month is dropped, a swing of 0.3469. Six of the ten drops leave it at exactly 1.5000. The record is small, so the slope wobbles a little, and a third of a point is roughly the wobble to be expected from a sampleThe handful of readings actually to hand, as against the far larger set of readings that could have turned up. Built in earlier notes; the point here is only that a small one moves about when it is disturbed. of ten.
With both inputs in the fit, the coefficient on the Nakshatra unit runs from minus 34.2018 to 27.6536. The swing is 61.8554 on the same ten months, 178 times wider than the single input's swing. Dropping the second month alone sends it to minus 34.2018; dropping the third sends it to plus 27.6536. A coefficient that travels sixty two points depending on which single month is removed is not measuring anything, and printing it to four decimal places tells a reader the exact opposite of the truth about it.
Refitting on nine months is a diagnostic here and nothing more. Refitting pokes the record to see how firmly a number is held. The refit is not a method for choosing between models, not a way of scoring a model on data it has not seen, and not a smaller version of anything that does those jobs. Model choice and out-of-sample scoring belong to the notes on building models.
Dropping one month moves the paired coefficient anywhere between minus 34.2018 and 27.6536. What follows about that coefficient being reported to four decimal places?
What can actually be done about it?
Four responses, and every one of them costs something. The record does not contain the information being asked of it, so no arithmetic removes the problem while leaving everything else where it was. The choice is which loss to accept.
Keep one input and drop the other. Fitted on the Nakshatra unit alone, the Vasant unit gives a coefficient back at 1.5000 with a standard error of 0.3014, and the report is honest again. The cost is whatever the Chandana unit was carrying that the Nakshatra unit was not. On this invented record that is almost nothing, and on a real one it might be the more accurate of the two columns. Nothing in the arithmetic settles which to keep. The choice is a judgement about what the two columns mean, and it belongs to the analyst.
Combine the two into a single input. The two are averaged, or added, and the fit runs on the one column that results. Fitted on the average of the two, the coefficient is 1.500058 and the squared misses are 217.9362, so almost nothing is lost. The cost is the ability to say anything about either column separately, forever. The report will carry one number for the pair and no number for either member of it.
Keep both and report only what the pair says. The two together carry 1.5000 and neither minus 1.0000 nor 2.5000 is quoted on its own. Reporting the pair alone is the honest reading of the fit exactly as it stands, and it costs the individual claims that somebody probably wanted.
Collect months in which the two inputs come apart. Collecting such months is the only response that adds information rather than surrendering some. If a stretch of record exists where the Nakshatra unit rose while the Chandana unit fell, the arithmetic finally has something to separate them with, and the coefficients settle down on their own. The panel further up is that idea in miniature: widening the wobble is exactly what makes the two columns less alike, and the coefficients calm down as it widens.
No arithmetic fix exists, and anything offered as one is a method from a different subject. Techniques that shrink coefficients towards zero until they stop misbehaving, and techniques that pick among inputs automatically, are real, they are useful, and they belong to the notes on building models where they are treated properly with their own assumptions. Reaching for one of them here would swap an understood problem for an untaught method, and it would not answer the question the pair of columns is refusing to answer.
Somebody in the room proposes a technique that shrinks both coefficients towards zero until they stop misbehaving. What is the right thing to say about it?
How is this caught before it does damage?
Here is the checklist worth carrying away. Anyone handed a fitted model that somebody else built stands in one spot: holding a report, judging whether enough has been disclosed to act on it. A lender assessing a scoring model that arrived with an application, an analyst going through a colleague's work, an operator given a forecast by a supplier, a household opening the note its bank sent about where the money went. Five checks, and every one of them looks at the coefficients or at the inputs. Not one of them looks at the quality of the fit. The quality of the fit is the one thing established above as certain not to show the problem.
Every pair of inputs is correlated before anything is fitted. Not after a coefficient looks odd: by then the coefficient is already in a document. With two inputs there is one pair to check. With six there are fifteen pairs, one small table and about a minute. The table for the three units here would have carried 0.999967 in a cell and the whole exercise would have ended there.
Each coefficient is compared with the one that input gives on its own. Fit the Vasant unit on the Nakshatra unit alone and get 1.5000. Fit it on both and get minus 1.0000. A gap of two and a half points between the solo reading and the joint reading is the signal, and it costs one extra fit to find.
Treat a flipped sign as the loudest version of that signal. A coefficient that changes size when another input joins is ordinary and often correct. A coefficient that changes direction says the fit has rearranged something structural, and it deserves an explanation before anybody writes a sentence based on it.
Read the plain and the adjusted measures of fit side by side. Here the plain one rose by 0.0001 and the adjusted one fell by 0.0391. Only one of those two can fall, so only one of them was ever capable of saying anything.
And refit on nine tenths of the record, then see how far the coefficients travel. Sixty two points of travel from removing one month is not a subtle warning. Again: a diagnostic, not a way of picking a model.
How this goes wrong in a real room
A team fits the Vasant unit on both traded units. The report comes back with an R squared of 0.7560 and a coefficient of minus 1.0000 on the Nakshatra unit. Somebody writes the obvious sentence: the Nakshatra unit moves the Vasant unit down.
On the identical ten months, fitted on its own, the Nakshatra unit moves the Vasant unit up at 1.5000. Nothing was miscalculated anywhere. Two columns carrying the same information were handed to a method with no way to tell them apart, and the method split one relationship between them in a way that happened to put a minus in front of the first.
The cost is not a poor fit. The fit is fine, and every prediction moved by less than an eighth of a percentage point. The cost is a reported direction that is the reverse of the truth, carried into every sentence written downstream, with a healthy looking R squared sitting beside it and an adjusted figure of 0.6863 that nobody read because the plain one had gone up.
The fix is a habit, not a technique: correlate the inputs before fitting, compare every coefficient against the version from that input alone, and never report a direction out of a fit that has not been checked for duplicate columns.
Where this guide stops. A single coefficient's claim when it stands alone is covered separately. So is the case of an input that alters another input's slope instead of repeating what it says, a wholly different subject. Reading the misses in the order the months arrived is covered separately as well. And methods that shrink coefficients towards zero, or pick among inputs without being asked, are real and useful and sit in the notes on building models: refitting on nine months here is a diagnostic and not a small version of any of them.
Which outside document decided any of this?
None. Fitting a line through paired numbers and reading what the line leaves behind is arithmetic that can be redone with a pencil and a patient afternoon. The arithmetic carries no limit set by a supervisor, no quantity published by a trading venue and no record kept by any institution, so there is no document to name.
The single quantity here that comes from custom rather than from arithmetic is the working habit of treating a variance inflation factor above ten as worth a second look. Treating ten as the line is a convention, nobody enforces it, and it is not a limit. The habit is widespread and its origin is not settled.
Two consequences follow, and they cut in opposite directions. The comfortable one: every quantity above can be checked by hand from the thirty figures, so anyone who doubts a line of it can sit down and test that line, needing nothing from anybody. The uncomfortable one: because the columns were built to be tidy, the sum of the two coefficients lands on exactly 1.5000 and R squared repeats to twelve decimal places. Records that were measured rather than composed do not close like that, and somebody whose entire acquaintance with fitting is made of worked examples will get a shock when a pair of coefficients first arrives summing to 1.4983 with the fit shifting in the third decimal rather than the twelfth. Neatness here is a convenience of teaching, and of everything in this guide it is the first property a genuine record would take away.
| What it supplied | Where it can be read |
|---|---|
| Nothing at all | Nowhere |
| Every quantity above | Only here |
| A rule of thumb, named as one | Common practice, credited to nobody |
The Nakshatra unit, the Vasant unit and the Chandana unit are invented.
Educational material. Not advice on any investment, tax, budget or market position.
