Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
041Explain the structure of a probabilistic graphical model you have worked with.Tower Research CapitalQuantitative Research · New York · 2015
Say this
Pick one model you actually built and describe it in four parts: the variables, the graph and what the missing edges assert, how you did inference, and how you checked it. The missing edges are the interesting part, because a graphical model is a set of conditional independence claims.
Then walk it
- Name the class first. A directed model, a Bayes net, factorises the joint as a product of each node given its parents and encodes causal or generative structure. An undirected model, a Markov random field, factorises into potentials over cliques and is better when the interactions have no natural direction.
- Then say what the graph buys you. Without structure, a joint over n binary variables needs 2 to the n minus 1 parameters. With a sparse graph it needs a handful per node. That reduction is the whole point, and the missing edges are the assumptions you are making.
- Inference: exact by belief propagation or the junction tree if the graph is a tree or has small treewidth, otherwise approximate by variational methods, loopy BP or MCMC. Say which you used and why, and say what the cost was.
- A concrete example is worth more than the taxonomy. A hidden Markov model is the simplest useful case: a latent state that evolves as a Markov chain with observations conditionally independent given the state. In markets people use it as a regime model, with the latent state as calm or stressed, fitted by Baum-Welch, and decoded with Viterbi.
- Then the honest part: on financial data the latent states are unstable, the number of regimes is not identified, and the fitted model will happily tell you the regime changed last week when it changed two months ago. So I used it as a descriptive overlay, never as a standalone signal.
Where candidates lose it
Reciting textbook definitions of Bayes nets and MRFs without ever describing a model you built. This question is a depth probe, and the interviewer will go three levels down on whichever model you name, so name the one you know cold. Be able to state the conditional independence your graph asserts and how you validated it.
Expect next
- What conditional independences does your graph assert, and did you test them?
- How did you do inference, and what was the complexity?
- How would you learn the graph structure from data?
Reported by candidates at Tower Research Capital (Quantitative Research, New York, 2015). Source: Wall Street Oasis.
042Derive the update rules for alternating least squares in a matrix factorisation.Tower Research CapitalQuantitative Research · New York · 2015
Say this
Fix one factor and the objective becomes an ordinary ridge regression in the other, so each update is a closed-form normal equation. With R approximated by U times V transpose and an L2 penalty, the update for a row of U is (V'V plus lambda I) inverse V'r.
Then walk it
- Objective: minimise the sum over observed entries of (r_ij minus u_i dot v_j) squared plus lambda times the sum of the squared norms of u and v. It is non-convex jointly in U and V, but convex in each one separately. That is the entire reason alternating minimisation works here.
- Differentiate with respect to u_i holding V fixed. The gradient is minus 2 times the sum over observed j of (r_ij minus u_i dot v_j) v_j plus 2 lambda u_i. Set it to zero.
- Rearranged: (sum over observed j of v_j v_j' plus lambda I) u_i equals the sum over observed j of r_ij v_j. So u_i equals that Gram matrix inverse times the weighted sum. Symmetric for v_j with U fixed.
- Cost per update is k cubed for the k by k solve plus k squared per observed entry, and it parallelises perfectly by row, which is exactly why ALS beat SGD for large recommender systems.
- Say the limitations. It converges to a local optimum only, so initialisation matters, usually small random or SVD-based. The lambda is essential because otherwise the Gram matrix is singular for users with fewer than k observations. And it monotonically decreases the objective every half-step, so if your loss ever goes up you have a bug in the derivation, which is a useful debugging fact.
Where candidates lose it
Writing down the gradient-descent update instead of the closed-form solve. ALS is defined by exploiting the per-block convexity to solve exactly, not by stepping. Also do not forget the lambda I, since without it the system is singular for sparse rows, and do not sum over all j when only observed entries enter the loss.
Expect next
- Why does ALS converge, and to what?
- When would you prefer SGD over ALS?
- How would you handle implicit feedback where you only see the ones?
Reported by candidates at Tower Research Capital (Quantitative Research, New York, 2015). Source: Wall Street Oasis.
043State the central limit theorem and tell me where it fails.Quant researchQuant trading
Say this
For independent identically distributed variables with finite mean and finite variance, the standardised sample mean converges in distribution to a standard normal. The key conditions are finite variance and enough independence, and both fail regularly in markets.
Then walk it
- Precisely: root n times (X bar minus mu) over sigma converges in distribution to N(0,1). Note it is the standardised mean that converges, and the rate is 1 over root n.
- Failure one, infinite variance. A Cauchy distribution has no variance and the sample mean of Cauchys is Cauchy again, no matter how large n is. Averaging buys you nothing. More generally, stable distributions with tail index alpha below 2 converge to a stable law, not a normal.
- Failure two, dependence. With strongly autocorrelated data the effective sample size is far below n, so you converge much more slowly and your standard errors are too small. Long-range dependence can break it entirely.
- Failure three, the rate in the tails. Even where the CLT holds, convergence is fastest in the middle and slowest in the tails, which is precisely where a risk manager needs accuracy. Berry-Esseen gives an error bound of order 1 over root n times the third absolute moment, so skewed data converges slowly.
- The practical version: daily equity returns have kurtosis of 5 to 10 and volatility clustering, so ten-day sums are much closer to normal than daily returns, but a 99.9 percent quantile computed from a normal assumption will still understate the tail badly. That is why value at risk models use empirical or extreme-value tails rather than leaning on the CLT.
Where candidates lose it
Stating the theorem without the finite variance condition, or claiming everything becomes normal for large n. Also do not confuse it with the law of large numbers, which is about convergence of the mean to a constant and needs only finite mean. Be ready to say what happens with infinite variance, because that is the follow-up.
Expect next
- What happens with a Cauchy distribution?
- How is that different from the law of large numbers?
- How large does n have to be in practice for returns data?
044What is a p-value, and what is it not?Quant researchRisk
Say this
It is the probability of seeing data at least as extreme as what you saw, assuming the null hypothesis is true. It is not the probability that the null is true, and it is not the probability you are wrong.
Then walk it
- The conditioning runs the wrong way from what people assume. A p-value is P(data given null), and what you actually want is P(null given data). Those are different objects and Bayes tells you the second depends on your prior.
- Concretely: if you test a thousand strategies of which fifty genuinely work, at a five percent significance level you get roughly 47 true discoveries and 47 false ones. A p-value of 0.05 in that setting means a coin flip on whether the finding is real.
- It also says nothing about effect size. With a million observations a completely useless one-basis-point edge will have a p-value of 0.0001. Significance is not importance, and in high-frequency data everything is significant.
- And it is only valid for a pre-specified test. Choosing the test after looking at the data, or stopping data collection when the p-value crosses 0.05, invalidates it completely.
- What I would report instead on a desk: the effect size with a confidence interval, out-of-sample performance, and how many specifications I tried. A p-value on its own is close to useless in a research process where hundreds of hypotheses get screened.
Where candidates lose it
Defining it as the probability the null is true. That is the single most common statistical error in finance interviews and it is disqualifying at a research shop. Also be ready with the multiple-testing consequence, because the interviewer's real target is whether you understand why published anomalies do not replicate.
Expect next
- So what significance level would you use if you screened a thousand signals?
- Explain the false discovery rate.
- What would you report instead of a p-value?
045You test two hundred signals and three come back significant at the five percent level. What do you conclude?Quant researchQuant trading
Say this
That you have found nothing. Under a pure null you would expect ten false positives from two hundred tests at five percent, so three is fewer than chance. If anything the result is evidence against there being any signal at all.
Then walk it
- Expected false positives are 200 times 0.05 equals 10. Getting three significant results is below what noise alone produces, so the finding is not just unimpressive, it is worse than random.
- The right frame is family-wise error or false discovery rate. Bonferroni sets the threshold at 0.05 over 200, which is 0.00025, brutal but valid. Benjamini-Hochberg controls the expected proportion of false discoveries among the rejections and is much less conservative, which is usually the better choice when you are screening.
- The subtlety with financial signals: they are heavily correlated with each other, so the effective number of independent tests is far below 200. Bonferroni is then too harsh. I would estimate the effective number of tests, for example from the eigenvalue spectrum of the signal correlation matrix, or use a permutation or block-bootstrap null that preserves the correlation structure.
- The right test of whether anything survived is not a p-value at all. It is out-of-sample: hold back a period, or better a different market, and see whether the three signals still work with the sign you predicted.
- And the disclosure discipline, which is the answer a research head wants to hear: I would report the number of specifications tried alongside the result. The deflated Sharpe ratio and Harvey and Liu's work on multiple testing in finance both exist because the profession spent decades not doing this.
Where candidates lose it
Getting excited about the three and building a strategy on them. The whole question is whether you compute the expected number of false positives before you get attached. Say ten out of two hundred immediately, then talk about correlated tests, because that is where the technical depth is.
Expect next
- How would you estimate the effective number of independent tests?
- What is the deflated Sharpe ratio?
- How would you set up the experiment properly from the start?
046What are the assumptions behind ordinary least squares, and which of them actually matter?Quant researchRisk
Say this
Linearity in parameters, exogeneity meaning the error has zero mean conditional on the regressors, no perfect collinearity, homoskedasticity, and no autocorrelation. Only exogeneity is essential for unbiasedness. The last two affect efficiency and standard errors, not the coefficients.
Then walk it
- Exogeneity, E of error given X equals zero, is the load-bearing assumption. Break it and every coefficient is biased and inconsistent, and no amount of data or robust standard errors saves you.
- Homoskedasticity and no autocorrelation give you Gauss-Markov efficiency and the usual standard error formula. Break them and OLS is still unbiased, just no longer the minimum-variance linear estimator, and your t-statistics are wrong. Robust or Newey-West errors fix the inference.
- Normality of errors is not needed for unbiasedness or consistency at all. It only buys exact small-sample t and F distributions. Asymptotically the CLT handles it.
- No perfect collinearity is a requirement for the estimator to exist, since X'X must be invertible. Near-collinearity is not a violation, it just inflates variances.
- On financial data the realistic picture is: heteroskedasticity almost always, autocorrelation often, and exogeneity frequently violated because everything is jointly determined. So I default to robust standard errors, and I spend my thinking time on whether my regressor is endogenous, because that is the one that actually changes the answer.
Where candidates lose it
Listing normality as a core assumption, or treating all five as equally important. Rank them. The interviewer wants to hear which violations bias the coefficients and which only bias the standard errors, because that distinction determines whether you patch the model or rebuild it.
Expect next
- Give me a concrete example of endogeneity in a returns regression.
- Why is normality not needed?
- What does Gauss-Markov actually claim?
047Your regression has two highly correlated predictors. What happens, how do you detect it, and what do you do?Quant researchRisk
Say this
The coefficients stay unbiased but their variances blow up, so individual t-statistics collapse and signs flip from sample to sample while the overall fit and the joint prediction stay fine. Detect it with variance inflation factors or the condition number, then either combine the predictors or regularise.
Then walk it
- The mechanism: the variance of a coefficient is proportional to 1 over (1 minus R squared of that regressor on the others). At a pairwise correlation of 0.95 the variance inflation factor is about 10, so your standard error is roughly three times larger than it would otherwise be.
- The tell-tale symptom is a regression with a high overall R squared and an F test that rejects, but no individual coefficient significant. That combination is almost always collinearity.
- Detection: VIFs above 5 or 10 as a rough flag, or the condition number of the scaled X matrix above 30. Better still, look at the eigenvalues of the correlation matrix, since a near-zero eigenvalue is the direction that is unidentified.
- Fixes in order of preference: drop one if they are measuring the same thing, combine them into a single factor such as a sum or a principal component, or use ridge, which trades a little bias for a large variance reduction and is the textbook answer for exactly this problem.
- The thing to say before they ask: if you only care about prediction, collinearity is close to harmless, because the fitted values are stable even when the coefficients are not. It only matters if you want to interpret the individual coefficients or attribute risk to individual factors. That is why it is a bigger problem in a risk model than in a forecasting model.
Where candidates lose it
Claiming collinearity biases the coefficients. It does not. And do not automatically drop a variable, because if both belong in the model economically, dropping one creates omitted variable bias, which is a worse problem than inflated variances. Distinguish the prediction case from the interpretation case.
Expect next
- Why does ridge help here, mathematically?
- Is collinearity a problem if you only care about forecasting?
- How is this different from omitted variable bias?
048Asset volatility comes in clusters. What does that break, and how do you model it?Quant researchRisk
Say this
It breaks the constant-variance assumption behind almost everything: OLS standard errors, iid return models and Black-Scholes. The standard answer is a GARCH model, where today's variance depends on yesterday's variance and yesterday's squared shock.
Then walk it
- The empirical fact first: returns are close to unpredictable in the mean but their squares and absolute values are strongly autocorrelated, with the autocorrelation of squared returns decaying over weeks. Big moves cluster.
- GARCH(1,1) is sigma squared at t equals omega plus alpha times the last squared return plus beta times the last variance. On daily equities alpha is typically around 0.05 to 0.1 and beta around 0.85 to 0.92, with alpha plus beta just under one, meaning very persistent but eventually mean reverting.
- Long-run variance is omega over (1 minus alpha minus beta). If alpha plus beta hits one you get integrated GARCH, which is essentially an exponentially weighted moving average with no mean reversion, and that is what RiskMetrics used.
- It matters for options because it generates both fat unconditional tails and a term structure of volatility, which is why implied vol curves upward or downward towards the long-run level depending on where spot vol sits.
- Variants worth naming and the honest limitation: GJR-GARCH or EGARCH add the leverage effect, since negative returns raise vol more than positive ones, which plain GARCH cannot capture. And for anything intraday I would prefer realised volatility from high-frequency data, because a HAR model on realised vol usually forecasts better than GARCH on daily closes.
Where candidates lose it
Describing GARCH mechanically without saying what it is for. The point is that conditional variance is forecastable even when the mean is not, which is why volatility trading exists and directional trading is hard. Also do not forget the leverage effect, since plain GARCH is symmetric in the sign of returns and equity vol is not.
Expect next
- Why does alpha plus beta sit so close to one?
- What is the leverage effect and which model captures it?
- Would you use GARCH or realised volatility to forecast tomorrow's vol?
049You need a covariance matrix for five hundred assets and you have two years of daily data. What is the problem and how do you fix it?Quant researchRisk
Say this
You have 500 assets and roughly 500 observations, so the sample covariance matrix is nearly singular and its smallest eigenvalues are garbage. Any optimiser will load up on exactly those directions, so you have to shrink or impose factor structure.
Then walk it
- Count the parameters: 500 times 501 over 2 is about 125,000 numbers estimated from 250,000 data points. The ratio of assets to observations, roughly one here, is what governs the damage, and the sample eigenvalue spectrum is badly biased even at a ratio of a quarter.
- Marchenko-Pastur describes exactly how the eigenvalues spread out. The largest are overstated and the smallest understated, and the smallest ones are the low-variance directions a mean-variance optimiser will concentrate in. That is why naive optimisers produce absurd leveraged long-short positions.
- Fix one, shrinkage. Ledoit-Wolf shrinks the sample matrix towards a structured target like a constant-correlation matrix, with an optimal intensity derived in closed form. Cheap, well-behaved and hard to beat as a default.
- Fix two, factor structure. Model returns as exposures to a few factors plus idiosyncratic noise, so the covariance is B times F times B transpose plus a diagonal. You have gone from 125,000 parameters to a few thousand. This is what every commercial risk model does.
- Fix three, random matrix filtering: keep the eigenvalues above the Marchenko-Pastur bulk edge as signal and replace the bulk with its average. Then state the practical check, which is out-of-sample portfolio variance rather than any in-sample fit statistic, because in-sample the sample matrix always wins and is always wrong.
Where candidates lose it
Saying you would just use the sample covariance matrix because two years is a lot of data. It is not, relative to 500 assets. The interviewer is testing whether you know that estimation error in the covariance matrix, not in the means, is what breaks portfolio optimisation in practice, and whether you can name shrinkage or factor models as the fix.
Expect next
- Why does the optimiser concentrate in the smallest eigenvalue directions?
- How do you choose the shrinkage intensity?
- How would you test whether your covariance matrix is any good?
050A colleague is excited about an R squared of 0.9 on a returns regression. What is your reaction?Quant researchQuant trading
Say this
Suspicion, not excitement. An R squared of 0.9 on returns almost always means a bug: a look-ahead leak, a regression of a price level on another price level, or the dependent variable included on the right-hand side. Real return predictability lives at an R squared of a fraction of a percent.
Then walk it
- Benchmark it. A genuinely good daily return predictor has an R squared around 0.001 to 0.01. A monthly cross-sectional factor model might reach a few percent. Anything above 0.1 on returns is a red flag rather than a result.
- Most likely causes in order: the target is in the features, the features are computed with future information, you regressed levels on levels where both are trending, or you regressed a variable on itself lagged by zero periods.
- The levels problem deserves a name. Two independent random walks regressed on each other will produce a high R squared and a significant t statistic almost every time, because the standard errors are wrong under non-stationarity. That is spurious regression, and it is Granger and Newbold's result.
- Also note what R squared does not tell you even when it is right: nothing about out-of-sample performance, nothing about economic significance, and it always rises when you add regressors, which is why adjusted R squared exists, penalising by (n-1)/(n-k-1).
- So what I would do: check for leakage first, difference the series and re-run, then look at out-of-sample R squared. And the thing worth knowing is that an out-of-sample R squared of 0.005 on daily returns, if it is real and tradeable, is a very good strategy. Small numbers are the norm and big numbers are bugs.
Where candidates lose it
Congratulating them. Knowing the realistic magnitude of return predictability is a strong signal that you have done real work, and not knowing it is a strong signal that you have not. Name look-ahead bias and spurious regression on levels as the two prime suspects.
Expect next
- What is a realistic R squared for a daily return forecast?
- Explain spurious regression between two random walks.
- What is out-of-sample R squared and how do you compute it honestly?
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

