Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quant Analyst · CoreTrack
1Quantitative Methods, Financial Data & Programming
iProbability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
iiStatistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
iiiCorrelation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
ivTime Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
vSimulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
viOptimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
viiModelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
viiiBacktesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
ixData Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
xProgramming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
xiQuantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity
2Stochastic Calculus & Derivative Pricing Theory
iProbability Foundations
The Probability SpaceRandom VectorsSigma-AlgebraExpectationSample Space and EventsDensity and Distribution FunctionsRisk-Neutral ProbabilityState Price Density vs…
iiStochastic Processes and Jumps
Properties of a Stochastic ProcessMartingaleBrownian Motion and Its PropertiesBrownian Motion vs Geometric…Stopping TimeThe Markov PropertyState VariablesTransition ProbabilityQuadratic VariationQuadratic Variation vs Ordinary…Submartingale and SupermartingaleMartingale RepresentationMarkov Process vs MartingaleOptional StoppingFiltrationJump ProcessesThe Poisson ProcessLevy ProcessesJump Diffusion
iiiIto Calculus
The Ito IntegralThe Ito Integral vs the Riemann IntegralInfinitesimals in Stochastic CalculusQuadratic CovariationIto's LemmaHow to Apply Ito's…The Infinitesimal GeneratorIto Calculus vs Ordinary Calculus
ivStochastic Differential Equations
Stochastic Differential EquationsStochastic Differential Equation vs…Drift and DiffusionStrong and Weak Solutions ComparedDiscretisationGeometric Brownian Motion
vPricing Theory and No-Arbitrage
No-ArbitrageGirsanov, Radon-Nikodym and Change…Physical and Risk-Neutral Measures…The Fundamental Theorems of…The Law of One PriceThe Pricing KernelDiscount Factors and Zero-Coupon PricesReplication vs HedgingComplete Market vs Incomplete MarketClearing Margin Architecture
viOption Pricing Theory
European and American OptionsMonte Carlo European OptionThe Black-Scholes PDEBlack Scholes and the GreeksThe Payoff FunctionThe Binomial ModelBinomial Option PricingDelta Hedging in TheoryBoundary, Initial and Terminal ConditionsThe Exercise BoundaryHow to Check Put-Call…
viiVolatility Models
Constant, Local and Stochastic…Vasicek Model vs CIR ModelThe Heston ModelThe SABR ModelThe Volatility ProcessImplied VolatilityVolatility Smile vs Skew vs Surface
viiiInterest Rate Models
Interest-Rate DerivativesMean ReversionThe Zero-Coupon BondThe Ornstein-Uhlenbeck ProcessThe Discount CurveZero RatesShort-Rate Model vs Market Model
ixNumerical Pricing
Closed Form and Numerical…Monte Carlo PricingEuler and Milstein Schemes ComparedTree MethodsFinite Difference MethodsNumerical Error and StabilityVariance Reduction
xCalibration and Model Risk
Model OverrideMarket Price and Model PriceCalibrationHow to Document a Pricing ModelThe Educational Illustration LabelMarket ConventionsModel Uncertainty and LimitationsBacktesting a Pricing ModelIdentifiabilityCalibrated ParametersThe Calibration Loss Function

Multicollinearity: When Inputs Say the Same Thing Twice

Multicollinearity is what turns up when two inputs carry nearly the same information. On ten invented months, adding a near copy of the Nakshatra unit drags its coefficient from 1.5000 down to minus 1.0000 while the copy walks off with 2.5000. R squared meanwhile shifts from 0.7559 to 0.7560, so the quality of the fit cannot see any of it.

Earlier notes fitted one straight line through ten paired months and got 0.5000 plus 1.5000 times the input. The notes on the method behind that line listed the conditions it needs, and one of them was that no input may be a near copy of another. The near-copy condition has been sitting there like a smoke alarm nobody has tested. Testing it takes one step. A second input, almost the same thing as the first, is added, the line is refitted, and what breaks and, more importantly, what does not, comes straight out of the arithmetic.

Three invented objects do the carrying. Two of them have already appeared. The Nakshatra unit is a traded unitA deliberately loose word for anything with a price that is free to move. Whether it is a share, a fund, a bond or a bar of metal changes nothing about the divisions done below, so the notes decline to specify. whose ten monthly readings are 1.00, 6.00, minus 4.00, 11.00, 1.00, minus 9.00, 6.00, 1.00, minus 4.00 and 1.00 per cent. The Vasant unit is the outcome, at 3.00, 18.50, 2.50, 17.00, minus 3.00, minus 13.00, 6.50, minus 1.00, minus 7.50 and minus 3.00 per cent. Each reading is a monthly change, meaning the distance between where a price finished a month and where it began that month, put as a percentage of the beginning. Because the two columns describe the same ten stretches of time they line up as ten paired observationsTwo readings tied to each other by covering one and the same stretch of time, which is the only thing that permits either to be weighed against the other. Break the pairing and every quantity worked out below this line stops meaning anything., and they are held in time orderThe arrangement in which the readings actually happened, first month first. Sorting them by size would make a neater picture and would destroy the only evidence some of these notes have left to spend. throughout.

What is multicollinearity, and why define it by what it does?

Most definitions of this subject start from the cause: two or more inputs are strongly related to each other. The cause is true enough, and it is still the wrong end to hand a reader first. Naming the situation says nothing about why anyone should care. The damage is the better starting point.

Multicollinearity is the condition in which the coefficients of a fit become unstable and individually meaningless while the fit itself stays exactly as good as it was.

Read the second half of that sentence twice. A problem that leaves the quality of the fit untouched is a problem no measure of fit can ever report. That is the entire reason this has a name of its own. If it degraded the fit, nobody would need a word for it: the fit would degrade, somebody would notice, and the matter would end there. Multicollinearity does not degrade the fit. Multicollinearity sits underneath a perfectly healthy looking report and quietly rewrites what the report claims.

Here is the everyday version. Two people share a flat and both of them pay for the same electricity bill of Rs 4,000/- one month, one by transfer and one in cash to the landlord, and both write it into the shared accounts sheet. At the end of the month the sheet balances to the rupee. Total spending is right. Total contributed is right. There is no error in any total, so nothing anywhere reports an error. The broken question is a different one: how much did each person pay? The sheet can produce an answer to that, and it can produce several, and it has no way of showing which one is real. The totals are fine and the attribution is destroyed, and those two facts are compatible because they are different questions. Totals stand for the quality of a fit and attribution for the coefficients, and that is the subject here.

Try it out

Define multicollinearity by what it does rather than by what causes it. Which of these is the working definition?

Breaking Into Quants Bootcamp — Fin Maverick

What does a near copy actually look like?

The third object is the Chandana unit, invented for this guide. Its ten monthly changes are 1.00, 5.95, minus 3.95, 11.05, 0.95, minus 9.00, 5.95, 1.05, minus 4.05 and 1.05 per cent. Its average is 1.00 per cent, and that figure is the Nakshatra unit's average as well. Set the two columns beside each other and the Chandana unit is the Nakshatra unit moved by five hundredths of a percentage pointThe unit for a gap between two percentages. A reading that goes from 6.00 to 5.95 has moved five hundredths of a percentage point, and calling that a five hundredths of a per cent move would be a different and much smaller statement. in eight of the ten months, and left exactly alone in months one and six.

The two correlate at 0.999967. On a measure that stops at one, that is about as close to being the same column as two different columns can manage.

The duplication is obvious on sight. Nothing in the argument depends on the duplication being hard to spot. The interesting question is not whether a person could catch this by eye. The interesting question is how the arithmetic behaves when nobody catches it, and the arithmetic behaves identically whether the duplication is this blatant or subtle enough that no reader would ever see it. Two inputs that are both, say, some measure of size, gathered by different departments in slightly different ways, will do exactly what the Chandana unit is about to do, and nobody will be looking at the two columns side by side when it happens.

TWO SERIES, ONE AXIS, AND NO VISIBLE DIFFERENCE ANYWHERE Ten months in time order. The largest gap between the two lines is five hundredths of a percentage point. 0 11 -9 1 4 6 10 along the bottom: the month, kept in time order the Nakshatra unit the Chandana unit MONTH FOUR ALONE, DRAWN AT ABOUT FIFTY TIMES THE SCALE ABOVE 11.00 the Nakshatra unit 11.05 the Chandana unit 0.05 In the panel above, that same gap is drawn under half a pixel high. It is not hidden and it is not small print: it is simply the correct size, and at the correct size it disappears. The two columns correlate at 0.999967.
The Chandana unit sits within five hundredths of a percentage point of the Nakshatra unit in every one of the ten months, the two lines are indistinguishable at reading scale, and the correlation between them is 0.999967.

What happens to the coefficients when the copy is added?

Now the one step everything so far has been leading to. The Vasant unit stays the outcome, the Nakshatra unit stays an input, the Chandana unit joins as a second input, and the line is refitted. Nothing else changes. Not one of the thirty figures is edited. The same method that produced 1.5000 produces the new answer.

The coefficient on the Nakshatra unit comes out at minus 1.0000. The coefficient on the Chandana unit comes out at 2.5000. The intercept does not move at all: it is 0.5000 before and 0.5000 after.

The sign flipped. Before the copy arrived, the fit said the Nakshatra unit moves the Vasant unit up. After, on the identical ten months, it says the Nakshatra unit moves the Vasant unit down. Nothing was measured differently, nobody made an arithmetic slip, and no assumption was quietly changed. A column that duplicated an input already present was handed to the method, and the method returned an answer in which the original input is now working in the opposite direction.

The arithmetic has not gone mad, and what stayed true shows why. Minus 1.0000 plus 2.5000 is exactly 1.5000, the very figure the single input was saying on its own. The pair has not lost the relationship. The pair has kept the relationship perfectly and then split it between two columns, and because the two columns are near copies there is no arithmetic reason to prefer one split over another. Push a large positive weight onto one and a matching negative weight onto the other and the two almost cancel; whatever survives the cancellation is the relationship. The fit found one of the enormously many ways of doing that, and the way it found happens to put a minus in front of the first column.

The household version again. Two people paid one Rs 4,000/- bill between them. Any split at all satisfies the total: 4,000 and nothing, 1,000 and 3,000, or minus 6,000 and 10,000 if one of them was reimbursed along the way. The sum is nailed down and each share is not, and the sheet holds no information that would settle it. Only the pair's claim is real; what either column says alone is an artefact of the split.

THE PAIR STILL ENDS WHERE THE SINGLE INPUT ENDED Each bar is laid end to end from zero. Where the last bar stops is what the fit is saying in total. 1.5000 zero ONE INPUT the Nakshatra unit, 1.5000 TWO INPUTS minus 1.0000 the Nakshatra unit the Chandana unit, 2.5000 starts where the first bar stopped Sum: minus 1.0000 plus 2.5000 is 1.5000, exactly. The intercept is 0.5000 in both fits. Neither of the two lower bars is saying anything on its own. Only where they finish is settled. -1.5 -1.0 -0.5 0 0.5 1.0 1.5 2.0 2.5 3.0 along the bottom: the coefficient, in percentage points of the Vasant unit per percentage point of an input
The two coefficients of minus 1.0000 and 2.5000 add to exactly 1.5000, which is what the single input claimed on its own, so the pair is settled while neither number alone is saying anything.
Try it out

The coefficient on the Nakshatra unit goes from 1.5000 to minus 1.0000. What changed in the data to make that happen?

Try it out

The two coefficients are minus 1.0000 and 2.5000. What do they add to, and what does that sum show?

AI For Finance Bootcamp — Fin Maverick

What happened to the fit while all that was going on?

Almost nothing, and this is the half of the subject that people skip. Put the two fits side by side and read the bottom rows rather than the top ones.

Both fits, on the identical ten paired months. Every quantity here was worked out afresh rather than carried over.
What the fit reportsOne inputBoth inputs
Intercept0.50000.5000
Coefficient, the Nakshatra unit1.5000minus 1.0000
Coefficient, the Chandana unitnot in the fit2.5000
Sum of squared misses218.000217.875
R squared0.75590.7560
Adjusted R squared0.72540.6863
Largest move in any single fitted reading0.125 of a percentage point

The sum of squared misses drops from 218.000 to 217.875, a saving of an eighth. R squared rises from 0.7559 to 0.7560, a move in the fourth decimal place. And not one of the ten fitted readingsThe number the line produces for a given month, as opposed to the number that month actually recorded. Built in earlier notes and used here only as a thing that did or did not move. moves by more than 0.125 of a percentage point, on an outcome whose own spreadA single figure standing for how widely a set of readings is scattered, worked out in earlier notes and taken here as already built. Its whole job here is to be the yardstick a movement gets called large or small against. is 9.96 per cent. In fact every fitted reading moves by exactly 0.125 or by exactly nothing, because the change is 2.5000 multiplied by the five hundredths of a point separating the two inputs.

The coefficients moved by two and a half points. The predictions moved by an eighth of one.

The whole symptom fits in one line, and anyone who watches only the sign flip has seen half the subject and the less useful half. If the fit had fallen apart at the same time, everyone would find this on the first pass. The fit did not fall apart. The fit got a hair better.

One number on the report did move honestly, and it is the one that charges for inputs: adjusted R squared fell from 0.7254 to 0.6863. Adjusted R squared fell because the second column was billed for and delivered nothing worth the bill. Why the adjusted measure can fall while the plain one cannot is settled in the opening notes on measuring fit, where the comparison is worked through on this very same pair of columns. The fall of 0.0391 is 391 times the size of the rise of 0.0001, and it is the single signal on the whole report pointing at the trouble.

THE SAME REPORT, TWICE, WITH THE ROWS SORTED BY WHAT MOVED Identical ten months, identical method. The only difference is one extra column of inputs. ONE INPUT BOTH INPUTS THE COEFFICIENTS Intercept 0.5000 0.5000 did not move The Nakshatra unit 1.5000 minus 1.0000 sign flipped The Chandana unit not in the fit 2.5000 took the relationship THE FIT Sum of squared misses 218.000 217.875 an eighth better R squared 0.7559 0.7560 fourth decimal Adjusted R squared 0.7254 0.6863 the one honest fall Largest move in a fitted reading 0.125 of a percentage point on a spread of 9.96 THE COEFFICIENTS MOVED BY TWO AND A HALF POINTS.THE PREDICTIONS MOVED BY AN EIGHTH OF ONE.
Adding a near copy moves the Nakshatra coefficient from 1.5000 to minus 1.0000 while R squared moves only from 0.7559 to 0.7560 and no fitted reading shifts by more than 0.125 of a percentage point.
Try it out

R squared moves from 0.7559 to 0.7560. Why could it not have moved much, whatever the data had been?

Try it out

Adjusted R squared falls from 0.7254 to 0.6863 while plain R squared rises. What is the adjusted one charging for?

Try it out

As the copy is allowed to differ more and more from the original, becoming less and less of a duplicate, what happens to R squared?

Play with it

Widen the gap between the two inputs and watch two coefficients walk apart while one fit refuses to move.

One control moves: how far the Chandana unit is allowed to differ from the Nakshatra unit, from five hundredths of a percentage point up to a full point. The Nakshatra readings never change and the Vasant readings never change. Only the size of the wobble does. With the trail turned on, every setting visited leaves a mark, so after a minute of dragging the coefficients stand scattered across the scale and the fit marker piled up in one place. The opening setting is a wobble of 0.05, giving minus 1.0000 and 2.5000, the same two figures written out in the table above.

Jump to a published setting:
The Nakshatra coefficient
minus 1.0000
The Chandana coefficient
2.5000
The two added together
1.5000
R squared, twelve decimals
0.756019036954
Correlation of the inputs
0.999967
Variance inflation factor
15,001
Loading the panel.

Educational illustration on invented data. The Nakshatra and Vasant readings are frozen at every setting and only the size of the wobble moves. The two coefficients add to 1.5000 at every setting and R squared reads 0.756019036954 at every setting, and both of those are consequences of how the third column is built rather than coincidences worth remarking on.

Why can R squared not see any of this?

State the reason rather than repeating the observation. R squared is worked out from the fitted readings and from nothing else: it compares how far the fitted readings sit from the actual ones against how far the average sits from the actual ones. The fitted readings barely moved. Therefore R squared barely moved. R squared could not have reported this trouble whatever the data had been. The trouble does not live in the fitted readings, and R squared has never looked anywhere else.

The panel above makes that concrete in a way a sentence cannot. Widen the wobble and the two coefficients walk right across the scale, from minus 1.0000 and 2.5000 to 1.3750 and 0.1250. The sum stays at 1.5000. And R squared reads 0.756019036954 at every single setting, identical to twelve decimal places. The fitted readings at every setting are the same ten numbers. The fit is not roughly the same. The fit is the same.

So a reader who checks the quality of the fit and stops has checked the one quantity that was never able to show the problem. The failure is one of process rather than of arithmetic. The arithmetic did what it was asked. The order in which somebody looked at things is what went wrong, and no amount of care with the divisions can fix an order-of-looking problem.

A caretaker checks a building by walking the corridors and looking for damage. The corridors are spotless every week. Meanwhile the labels on the two water tanks have been swapped, so every reading anyone takes about which tank is which is wrong. The corridors stay spotless throughout: corridor condition and tank labelling are not related. Checking harder does not help. Checking something else does.

FIVE SETTINGS. THE COEFFICIENTS WALK. THE FIT DOES NOT MOVE A DIGIT. Each row is the same ten months with the copy allowed to differ by a little more. THE WOBBLE THE TWO COEFFICIENTS R SQUARED 0.05 minus 1.0000 2.5000 0.756019036954 0.10 0.2500 1.2500 0.756019036954 0.25 1.0000 0.5000 0.756019036954 0.50 1.2500 0.2500 0.756019036954 1.00 1.3750 0.1250 0.756019036954 -1.5 -1.0 0 1.0 2.0 3.0 the Nakshatra coefficient the Chandana coefficient five identical rows
Widening the copy walks the Nakshatra coefficient from minus 1.0000 up to 1.3750 and the Chandana coefficient from 2.5000 down to 0.1250, while R squared reads 0.756019036954 on every row.
Risk Management Program Bootcamp — Fin Maverick

How is it measured, if the fit cannot report it?

By turning away from the outcome entirely and looking only at the inputs. The variance inflation factor takes each input in turn, asks how well that input can be predicted from all the others, and converts the answer into how much wider the coefficient's own uncertainty has become because of the overlap. Take the share of an input the other inputs can already account for, subtract it from one, and divide one by what is left. When nothing overlaps, the factor is one and no widening has happened.

Here the Nakshatra unit can be predicted from the Chandana unit almost perfectly, so the leftover is almost nothing, so the division by almost nothing produces something enormous. The factor reads 15,001 on this record, and the rule of thumb people carry around is that anything above ten is worth a second look. The rule of thumb is a habit of practice rather than a limit anyone sets, and it is treated as one here; the point is not the ten, it is that the reading here is fifteen thousand and the fit still moved only in the fourth decimal.

The factor is worked out from the inputs alone and never once looks at the outcome, and that is exactly why it can see what R squared cannot. R squared asks how close the fitted readings came to the actual ones. The variance inflation factor asks a question about the input columns that would have the same answer if the outcome column were deleted altogether. Two different questions, and the trouble lives in the second one.

The widening can be watched directly. Fitted on its own, the Nakshatra coefficient carries a standard errorHow far a computed figure would be expected to land from where this particular record placed it, had a different record of the same size turned up instead. Built in earlier notes; here it is simply the width that gets inflated. of 0.3014. Fitted alongside the copy, the same coefficient carries a standard error of 39.4506. The standard error is about 122 times wider from the duplication itself, and 130.90 times wider once the small change in the leftover spread is counted as well. A coefficient of minus 1.0000 with a standard error of 39.4506 beside it is a number the record has essentially nothing to say about, and the fit prints it to four decimal places with a straight face.

THE SCALE HAD TO BE MADE LOGARITHMIC OR ONE BAR WOULD VANISH Each step along the bottom multiplies by ten. The variance inflation factor for the Nakshatra unit, with the copy present. the usual rule of thumb 10 this record 15,001 1 10 100 1,000 10,000 100,000 Drawn on an ordinary evenly spaced scale, the upper bar would be under half a pixel wide beside the lower one. And through all of that, R squared moved from 0.7559 to 0.7560. THIS IS WORKED OUT FROM THE INPUT COLUMNS ALONE. THE OUTCOME IS NEVER CONSULTED.
The variance inflation factor here reads 15,001 against a rule of thumb that treats ten as worth a second look, and the scale had to be made logarithmic for both bars to be visible at once.
Try it out

The variance inflation factor reads 15,001. What is it worked out from, and why can it see something R squared cannot?

How unstable do the coefficients actually become?

Two and a half points of movement, from one added column, is already alarming. Here is the diagnostic that shows how little it would take to move them again. Drop one month from the record, refit on the remaining nine, and write down the coefficient. Do that ten times, once for each month, and look at the range.

With the Nakshatra unit as the only input, the slope stays between 1.3163 and 1.6633 whichever month is dropped, a swing of 0.3469. Six of the ten drops leave it at exactly 1.5000. The record is small, so the slope wobbles a little, and a third of a point is roughly the wobble to be expected from a sampleThe handful of readings actually to hand, as against the far larger set of readings that could have turned up. Built in earlier notes; the point here is only that a small one moves about when it is disturbed. of ten.

With both inputs in the fit, the coefficient on the Nakshatra unit runs from minus 34.2018 to 27.6536. The swing is 61.8554 on the same ten months, 178 times wider than the single input's swing. Dropping the second month alone sends it to minus 34.2018; dropping the third sends it to plus 27.6536. A coefficient that travels sixty two points depending on which single month is removed is not measuring anything, and printing it to four decimal places tells a reader the exact opposite of the truth about it.

Refitting on nine months is a diagnostic here and nothing more. Refitting pokes the record to see how firmly a number is held. The refit is not a method for choosing between models, not a way of scoring a model on data it has not seen, and not a smaller version of anything that does those jobs. Model choice and out-of-sample scoring belong to the notes on building models.

DROP ONE MONTH, REFIT, TEN TIMES OVER Both panels plot the coefficient on the Nakshatra unit. Only the scale differs. ON ONE SCALE WIDE ENOUGH TO HOLD THE PAIRED FIT one input all ten land inside three pixels here two inputs minus 34.2018 27.6536 -40 -20 0 20 35 THE SAME ONE INPUT REFITS, MAGNIFIED ONTO THEIR OWN SCALE 1.3163 1.4592 1.5000, where six of the ten land 1.5612 1.6633 1.25 1.35 1.45 1.55 1.65 1.75 One input: a swing of 0.3469. Two inputs: a swing of 61.8554, which is 178 times wider on the same ten months. THIS IS A DIAGNOSTIC. IT IS NOT A WAY OF CHOOSING A MODEL.
Dropping any one month moves the single input slope only between 1.3163 and 1.6633 while it moves the paired coefficient between minus 34.2018 and 27.6536, a swing 178 times wider.
Try it out

Dropping one month moves the paired coefficient anywhere between minus 34.2018 and 27.6536. What follows about that coefficient being reported to four decimal places?

What can actually be done about it?

Four responses, and every one of them costs something. The record does not contain the information being asked of it, so no arithmetic removes the problem while leaving everything else where it was. The choice is which loss to accept.

Keep one input and drop the other. Fitted on the Nakshatra unit alone, the Vasant unit gives a coefficient back at 1.5000 with a standard error of 0.3014, and the report is honest again. The cost is whatever the Chandana unit was carrying that the Nakshatra unit was not. On this invented record that is almost nothing, and on a real one it might be the more accurate of the two columns. Nothing in the arithmetic settles which to keep. The choice is a judgement about what the two columns mean, and it belongs to the analyst.

Combine the two into a single input. The two are averaged, or added, and the fit runs on the one column that results. Fitted on the average of the two, the coefficient is 1.500058 and the squared misses are 217.9362, so almost nothing is lost. The cost is the ability to say anything about either column separately, forever. The report will carry one number for the pair and no number for either member of it.

Keep both and report only what the pair says. The two together carry 1.5000 and neither minus 1.0000 nor 2.5000 is quoted on its own. Reporting the pair alone is the honest reading of the fit exactly as it stands, and it costs the individual claims that somebody probably wanted.

Collect months in which the two inputs come apart. Collecting such months is the only response that adds information rather than surrendering some. If a stretch of record exists where the Nakshatra unit rose while the Chandana unit fell, the arithmetic finally has something to separate them with, and the coefficients settle down on their own. The panel further up is that idea in miniature: widening the wobble is exactly what makes the two columns less alike, and the coefficients calm down as it widens.

No arithmetic fix exists, and anything offered as one is a method from a different subject. Techniques that shrink coefficients towards zero until they stop misbehaving, and techniques that pick among inputs automatically, are real, they are useful, and they belong to the notes on building models where they are treated properly with their own assumptions. Reaching for one of them here would swap an understood problem for an untaught method, and it would not answer the question the pair of columns is refusing to answer.

FOUR RESPONSES, AND WHAT EACH ONE COSTS None of these is a recommendation. Which loss is acceptable is a judgement about the columns, not about the arithmetic. THE RESPONSE WHAT IT GIVES UP Keep one input, drop the other the coefficient returns to 1.5000, error 0.3014 whatever the dropped column carried that the other did not Combine the two into one column on their average: 1.500058, misses 217.9362 any ability to speak about either one separately, for good Keep both, report only the pair the pair says 1.5000 and neither member speaks the individual claims that somebody probably wanted Collect months where the two come apart the only one that adds information nothing, beyond the time and access it takes to gather them THERE IS NO ARITHMETIC FIX. ANYTHING OFFERED AS ONE IS FROM A DIFFERENT SUBJECT.
Drop one, combine them, report only what the pair says, or collect months in which the two inputs come apart, and only the last of the four adds information rather than surrendering some.
Try it out

Somebody in the room proposes a technique that shrinks both coefficients towards zero until they stop misbehaving. What is the right thing to say about it?

Debt Capital Markets Bootcamp — Fin Maverick Regression for Finance — free micro-course from Fin Maverick

How is this caught before it does damage?

Here is the checklist worth carrying away. Anyone handed a fitted model that somebody else built stands in one spot: holding a report, judging whether enough has been disclosed to act on it. A lender assessing a scoring model that arrived with an application, an analyst going through a colleague's work, an operator given a forecast by a supplier, a household opening the note its bank sent about where the money went. Five checks, and every one of them looks at the coefficients or at the inputs. Not one of them looks at the quality of the fit. The quality of the fit is the one thing established above as certain not to show the problem.

Every pair of inputs is correlated before anything is fitted. Not after a coefficient looks odd: by then the coefficient is already in a document. With two inputs there is one pair to check. With six there are fifteen pairs, one small table and about a minute. The table for the three units here would have carried 0.999967 in a cell and the whole exercise would have ended there.

Each coefficient is compared with the one that input gives on its own. Fit the Vasant unit on the Nakshatra unit alone and get 1.5000. Fit it on both and get minus 1.0000. A gap of two and a half points between the solo reading and the joint reading is the signal, and it costs one extra fit to find.

Treat a flipped sign as the loudest version of that signal. A coefficient that changes size when another input joins is ordinary and often correct. A coefficient that changes direction says the fit has rearranged something structural, and it deserves an explanation before anybody writes a sentence based on it.

Read the plain and the adjusted measures of fit side by side. Here the plain one rose by 0.0001 and the adjusted one fell by 0.0391. Only one of those two can fall, so only one of them was ever capable of saying anything.

And refit on nine tenths of the record, then see how far the coefficients travel. Sixty two points of travel from removing one month is not a subtle warning. Again: a diagnostic, not a way of picking a model.

THE TABLE THAT SHOULD HAVE BEEN MADE FIRST Correlations between every pair of columns. One table, one minute, made before any fit is run. NAKSHATRA CHANDANA VASANT the Nakshatra unit the Chandana unit the Vasant unit 1.0000 0.999967 0.8694 0.999967 1.0000 0.8695 0.8694 0.8695 1.0000 two inputs, one column of information THE REPORT THAT GOT PRODUCED INSTEAD R squared 0.7560 adjusted R squared 0.6863 the Nakshatra unit minus 1.0000 error 39.4506 Nothing printed in this box names what the cell above would have shown. CORRELATING THE INPUTS FIRST CATCHES THIS. NO MEASURE OF FIT COMPUTED AFTERWARDS CAN.
Correlating every pair of inputs before fitting anything would have shown 0.999967 in one cell, and no measure of fit produced afterwards reports the trouble at all.

How this goes wrong in a real room

A team fits the Vasant unit on both traded units. The report comes back with an R squared of 0.7560 and a coefficient of minus 1.0000 on the Nakshatra unit. Somebody writes the obvious sentence: the Nakshatra unit moves the Vasant unit down.

On the identical ten months, fitted on its own, the Nakshatra unit moves the Vasant unit up at 1.5000. Nothing was miscalculated anywhere. Two columns carrying the same information were handed to a method with no way to tell them apart, and the method split one relationship between them in a way that happened to put a minus in front of the first.

The cost is not a poor fit. The fit is fine, and every prediction moved by less than an eighth of a percentage point. The cost is a reported direction that is the reverse of the truth, carried into every sentence written downstream, with a healthy looking R squared sitting beside it and an adjusted figure of 0.6863 that nobody read because the plain one had gone up.

The fix is a habit, not a technique: correlate the inputs before fitting, compare every coefficient against the version from that input alone, and never report a direction out of a fit that has not been checked for duplicate columns.

Where this guide stops. A single coefficient's claim when it stands alone is covered separately. So is the case of an input that alters another input's slope instead of repeating what it says, a wholly different subject. Reading the misses in the order the months arrived is covered separately as well. And methods that shrink coefficients towards zero, or pick among inputs without being asked, are real and useful and sit in the notes on building models: refitting on nine months here is a diagnostic and not a small version of any of them.

The sign flipped and the fit never caught it. See where a desk looks.

Which outside document decided any of this?

None. Fitting a line through paired numbers and reading what the line leaves behind is arithmetic that can be redone with a pencil and a patient afternoon. The arithmetic carries no limit set by a supervisor, no quantity published by a trading venue and no record kept by any institution, so there is no document to name.

The single quantity here that comes from custom rather than from arithmetic is the working habit of treating a variance inflation factor above ten as worth a second look. Treating ten as the line is a convention, nobody enforces it, and it is not a limit. The habit is widespread and its origin is not settled.

Two consequences follow, and they cut in opposite directions. The comfortable one: every quantity above can be checked by hand from the thirty figures, so anyone who doubts a line of it can sit down and test that line, needing nothing from anybody. The uncomfortable one: because the columns were built to be tidy, the sum of the two coefficients lands on exactly 1.5000 and R squared repeats to twelve decimal places. Records that were measured rather than composed do not close like that, and somebody whose entire acquaintance with fitting is made of worked examples will get a shock when a pair of coefficients first arrives summing to 1.4983 with the fit shifting in the third decimal rather than the twelfth. Neatness here is a convenience of teaching, and of everything in this guide it is the first property a genuine record would take away.

What was consulted
What it suppliedWhere it can be read
Nothing at allNowhere
Every quantity aboveOnly here
A rule of thumb, named as oneCommon practice, credited to nobody

The Nakshatra unit, the Vasant unit and the Chandana unit are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.