Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
038How would you fix violations of the OLS assumptions?AQR Capital ManagementInvestments · Greenwich · 2022
Say this
Depends which assumption breaks, and the fixes fall into two very different classes: violations that only break your standard errors, and violations that break the coefficients themselves. The first class you patch; the second class you have to re-specify the model.
Then walk it
- Heteroskedasticity and autocorrelated errors: coefficients stay unbiased, only inference is wrong. Fix with White or Newey-West robust standard errors, or clustered errors if the dependence is by group. Cheap fix, always worth doing on financial data.
- Endogeneity, meaning a regressor correlated with the error, whether from omitted variables, simultaneity or measurement error: this biases the coefficients and no standard error fix helps. You need an instrument, a control for the omitted factor, a fixed effect, or a different design.
- Multicollinearity: coefficients are still unbiased but the variances explode and the signs flip sample to sample. Drop or combine the collinear regressors, use ridge, or work with principal components. And check the variance inflation factors before you interpret anything.
- Non-normal or fat-tailed errors: inference is still fine asymptotically thanks to the CLT, but outliers dominate the fit because OLS minimises squares. Use robust regression, Huber loss or quantile regression, and always look at the influence diagnostics.
- Non-linearity: add the relevant transform or interaction rather than pretending it away. And I would say the order I actually work in on real data: plot residuals against fitted values and against time first, because most violations announce themselves visually before any test does.
Where candidates lose it
Listing fixes without separating what biases the coefficients from what only biases the standard errors. That distinction is the question. Slapping Newey-West errors on an endogenous regression is a common and useless move, and an interviewer at a research shop will push on exactly that.
Expect next
- Which of those actually biases your coefficients?
- How do you detect endogeneity if you have no instrument?
- What do you do when the residuals are fat-tailed and autocorrelated at the same time?
Reported by candidates at AQR Capital Management (Investments, Greenwich, 2022). Source: Wall Street Oasis.
046What are the assumptions behind ordinary least squares, and which of them actually matter?Quant researchRisk
Say this
Linearity in parameters, exogeneity meaning the error has zero mean conditional on the regressors, no perfect collinearity, homoskedasticity, and no autocorrelation. Only exogeneity is essential for unbiasedness. The last two affect efficiency and standard errors, not the coefficients.
Then walk it
- Exogeneity, E of error given X equals zero, is the load-bearing assumption. Break it and every coefficient is biased and inconsistent, and no amount of data or robust standard errors saves you.
- Homoskedasticity and no autocorrelation give you Gauss-Markov efficiency and the usual standard error formula. Break them and OLS is still unbiased, just no longer the minimum-variance linear estimator, and your t-statistics are wrong. Robust or Newey-West errors fix the inference.
- Normality of errors is not needed for unbiasedness or consistency at all. It only buys exact small-sample t and F distributions. Asymptotically the CLT handles it.
- No perfect collinearity is a requirement for the estimator to exist, since X'X must be invertible. Near-collinearity is not a violation, it just inflates variances.
- On financial data the realistic picture is: heteroskedasticity almost always, autocorrelation often, and exogeneity frequently violated because everything is jointly determined. So I default to robust standard errors, and I spend my thinking time on whether my regressor is endogenous, because that is the one that actually changes the answer.
Where candidates lose it
Listing normality as a core assumption, or treating all five as equally important. Rank them. The interviewer wants to hear which violations bias the coefficients and which only bias the standard errors, because that distinction determines whether you patch the model or rebuild it.
Expect next
- Give me a concrete example of endogeneity in a returns regression.
- Why is normality not needed?
- What does Gauss-Markov actually claim?
047Your regression has two highly correlated predictors. What happens, how do you detect it, and what do you do?Quant researchRisk
Say this
The coefficients stay unbiased but their variances blow up, so individual t-statistics collapse and signs flip from sample to sample while the overall fit and the joint prediction stay fine. Detect it with variance inflation factors or the condition number, then either combine the predictors or regularise.
Then walk it
- The mechanism: the variance of a coefficient is proportional to 1 over (1 minus R squared of that regressor on the others). At a pairwise correlation of 0.95 the variance inflation factor is about 10, so your standard error is roughly three times larger than it would otherwise be.
- The tell-tale symptom is a regression with a high overall R squared and an F test that rejects, but no individual coefficient significant. That combination is almost always collinearity.
- Detection: VIFs above 5 or 10 as a rough flag, or the condition number of the scaled X matrix above 30. Better still, look at the eigenvalues of the correlation matrix, since a near-zero eigenvalue is the direction that is unidentified.
- Fixes in order of preference: drop one if they are measuring the same thing, combine them into a single factor such as a sum or a principal component, or use ridge, which trades a little bias for a large variance reduction and is the textbook answer for exactly this problem.
- The thing to say before they ask: if you only care about prediction, collinearity is close to harmless, because the fitted values are stable even when the coefficients are not. It only matters if you want to interpret the individual coefficients or attribute risk to individual factors. That is why it is a bigger problem in a risk model than in a forecasting model.
Where candidates lose it
Claiming collinearity biases the coefficients. It does not. And do not automatically drop a variable, because if both belong in the model economically, dropping one creates omitted variable bias, which is a worse problem than inflated variances. Distinguish the prediction case from the interpretation case.
Expect next
- Why does ridge help here, mathematically?
- Is collinearity a problem if you only care about forecasting?
- How is this different from omitted variable bias?
050A colleague is excited about an R squared of 0.9 on a returns regression. What is your reaction?Quant researchQuant trading
Say this
Suspicion, not excitement. An R squared of 0.9 on returns almost always means a bug: a look-ahead leak, a regression of a price level on another price level, or the dependent variable included on the right-hand side. Real return predictability lives at an R squared of a fraction of a percent.
Then walk it
- Benchmark it. A genuinely good daily return predictor has an R squared around 0.001 to 0.01. A monthly cross-sectional factor model might reach a few percent. Anything above 0.1 on returns is a red flag rather than a result.
- Most likely causes in order: the target is in the features, the features are computed with future information, you regressed levels on levels where both are trending, or you regressed a variable on itself lagged by zero periods.
- The levels problem deserves a name. Two independent random walks regressed on each other will produce a high R squared and a significant t statistic almost every time, because the standard errors are wrong under non-stationarity. That is spurious regression, and it is Granger and Newbold's result.
- Also note what R squared does not tell you even when it is right: nothing about out-of-sample performance, nothing about economic significance, and it always rises when you add regressors, which is why adjusted R squared exists, penalising by (n-1)/(n-k-1).
- So what I would do: check for leakage first, difference the series and re-run, then look at out-of-sample R squared. And the thing worth knowing is that an out-of-sample R squared of 0.005 on daily returns, if it is real and tradeable, is a very good strategy. Small numbers are the norm and big numbers are bugs.
Where candidates lose it
Congratulating them. Knowing the realistic magnitude of return predictability is a strong signal that you have done real work, and not knowing it is a strong signal that you have not. Name look-ahead bias and spurious regression on levels as the two prime suspects.
Expect next
- What is a realistic R squared for a daily return forecast?
- Explain spurious regression between two random walks.
- What is out-of-sample R squared and how do you compute it honestly?
051What is omitted variable bias, and how would it show up in a factor regression?Quant researchPortfolio management
Say this
If you leave out a variable that belongs in the model and it is correlated with a regressor you kept, the kept coefficient absorbs part of its effect. The bias equals the true coefficient on the omitted variable times the regression coefficient of the omitted variable on the included one.
Then walk it
- Formula worth knowing: if the truth is y equals b1 x1 plus b2 x2 plus e and you regress y on x1 alone, you estimate b1 plus b2 times delta, where delta is from regressing x2 on x1.
- So the bias has a sign you can reason about. If the omitted factor has a positive premium and your included factor loads positively on it, you overstate your factor's premium.
- In factor work this is everywhere. Run a single-factor CAPM regression on a value portfolio and the alpha looks large, because you omitted the value factor. Add HML and the alpha collapses. That is not a bug, it is omitted variable bias doing exactly what the formula says.
- Momentum is the classic trap in the other direction. Omit momentum from a regression on a quality portfolio and quality's alpha inherits whatever momentum exposure quality happens to carry in your sample.
- How to handle it honestly: report alpha against a nested sequence of models, one factor then three then five plus momentum, and show what survives. And say the limitation out loud, because you can never rule out the factor nobody has published yet. The defence is out-of-sample and out-of-market evidence, not a longer regression.
Where candidates lose it
Defining it abstractly without giving the sign and magnitude formula, or without a concrete factor example. The interviewer wants to see you reason about the direction of the bias. Also do not confuse it with multicollinearity, which inflates variance without biasing anything.
Expect next
- Which direction does the bias go if the omitted factor is positively correlated with yours?
- How do you test whether your alpha survives the addition of a new factor?
- How is this different from multicollinearity?
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

