Case 057Regression and model reviewCore
A factor model of monthly returns shows a Durbin-Watson of 0.9, residual variance rising with market volatility (Breusch-Pagan p = 0.01) and variance inflation factors of 12 on two value factors. Which assumption does each break, what happens to the coefficients and t-statistics, and what is the fix?
1The situation
A junior analyst at Rajvantika Capital has regressed a strategy's monthly returns over 120 months on six factors, including two value factors: book-to-price and earnings yield. The headline result is a coefficient of 0.35 on book-to-price with a t-statistic of 2.4, which the analyst wants to report as significant.
The diagnostics printed under the regression show a Durbin-Watson statistic of 0.9, a Breusch-Pagan test for heteroskedasticity with a p-value of 0.01 (residual variance rises with market volatility), and variance inflation factors of 12 on both value factors. You are asked to review the model before it goes to the investment committee.
2Your task
For each diagnostic, name the assumption it breaks, say what it does to the coefficients and to the t-statistics, and propose the fix. Then say whether the t-statistic of 2.4 survives.
Quick check
What do these three problems mainly do to the regression?
Worked solution
Try it on paper, then open one step at a time.
30-second answerThe answer to give first
All three leave the coefficients roughly unbiased and all three make the t-statistic unreliable, so the 2.4 does not survive as reported. A Durbin-Watson of 0.9 means residual autocorrelation near 0.55, which understates standard errors; fix with Newey-West errors. The Breusch-Pagan result means non-constant variance; use robust errors. Variance inflation factors of 12 mean the value factors overlap by 92%; combine them. Corrected, the t-statistic falls to about 1.8.
Step 1What does a Durbin-Watson of 0.9 tell you?
A Durbin-Watson statistic near 2 means each month's residual is unrelated to the last; it is roughly 2 x (1 - rho), where rho is the residuals' first-order autocorrelation. A value of 0.9 implies rho of about 0.55: misses come in runs, which breaks the assumption that errors are independent. Think of a cricket batter's form. If a good innings makes the next one more likely good, ten innings tell you less about true ability than ten unrelated ones would. The coefficients are still roughly unbiased, but the usual standard errors assume every month is fresh information. If the factor itself is autocorrelated, say 0.5 month to month, the true standard error is about sqrt((1 + 0.55 x 0.5)/(1 - 0.55 x 0.5)) = 1.33 times the reported one. The fix is Newey-West standard errorsStandard errors that allow residuals to be correlated over several lags and to have changing variance, so t-statistics stay honest when the usual assumptions fail., or modelling the dynamics if a lagged term belongs in the model.
Step 2What does the Breusch-Pagan result change?
The test regresses squared residuals on candidate variables; a p-value of 0.01 says their variance moves with market volatility. That breaks the constant-variance assumption, which again leaves coefficients unbiased but makes the textbook standard errors wrong, usually too small when the volatile months are also the months when the factors move most. The fix is heteroskedasticity-robust standard errorsWhite standard errors: they estimate each coefficient’s uncertainty without assuming the residual variance is the same in every observation.; Newey-West already includes this, so one correction handles both problems. Weighted least squares, down-weighting the volatile months, is the alternative if you also want more precise estimates.
Step 3What does a variance inflation factor of 12 do to the value factors?
A variance inflation factor is 1/(1 - R-squared) from regressing one factor on all the others. A value of 12 means 91.7% of book-to-price is explained by the other factors, mostly earnings yield, so the regression cannot tell the two apart and each coefficient's standard error is sqrt 12 = 3.46 times what it would be with unrelated factors. This one is not a lie in the standard error; it is a true lack of information. The symptom is coefficients that swing, even flip sign, when a few months are added, and two value factors that are each insignificant though jointly strong. The fix is to combine them into one value composite, drop one, or test them jointly with an F-test rather than one at a time.
| Diagnostic | Assumption broken | Coefficients | t-statistics | Fix |
|---|---|---|---|---|
| Durbin-Watson 0.9 | Independent errors | Roughly unbiased | Too large | Newey-West errors, or model the lag |
| Breusch-Pagan p = 0.01 | Constant error variance | Roughly unbiased | Usually too large | Robust errors, or weighted least squares |
| VIF 12 on two factors | Regressors not too closely related | Unbiased but unstable | Small, honestly | Combine, drop one, or test jointly |
Step 4Does the t-statistic of 2.4 survive?
Apply the autocorrelation correction alone: 2.4 / 1.33 = 1.81, below the usual 1.96. Once the standard errors are made honest, book-to-price is no longer significant on its own, and the collinearity means the regression could not have separated it from earnings yield anyway. What you tell the committee is that value as a group may matter, which a joint test or a single composite factor can check, but the claim about book-to-price specifically does not stand. The limitation: the 0.5 autocorrelation of the factor is an assumption, so rerun with Newey-West errors on the actual data before quoting the corrected figure.
Where candidates lose it
Most candidates say each problem biases the coefficients. In this setting none of them does; the damage is to the standard errors, and the interviewer is listening for that distinction.
The second miss is treating multicollinearity like the other two. Autocorrelation and heteroskedasticity make the reported standard error wrong; collinearity makes the true standard error large. One is fixed with a better formula, the other only with different factors or more data.
What the interviewer asks next
- When would autocorrelated residuals bias the coefficients as well?
- How many lags would you choose for Newey-West standard errors on monthly data, and why?
- How would ridge regression change the two value coefficients?
Asked at Citadel, Quantitative Research, London, 2026 (Wall Street Oasis): assumptions of linear regressions how to detect them in practice
Company names and figures are illustrative.
