Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
052Explain the bias-variance tradeoff, and where a quant strategy usually sits on it.Quant researchQuant development
Say this
Expected prediction error decomposes into squared bias, variance and irreducible noise. Bias is how wrong your model class is on average, variance is how much your fit moves with a different sample. In financial data the noise term dominates everything, so you sit far towards the high-bias, low-variance end.
Then walk it
- The decomposition: E of (y minus f hat) squared equals bias squared plus variance plus sigma squared. Only the first two are under your control.
- Flexible models cut bias and raise variance. A deep tree fits any shape and moves wildly with resampling. A linear model with three factors barely moves but cannot represent an interaction.
- The financial context is what makes the answer different from a generic machine learning answer. Signal-to-noise on returns is tiny, sigma squared swamps the other terms, and a flexible model spends all its capacity fitting noise. So simple, heavily regularised, few-parameter models win out of sample far more often than they should on pure machine learning intuition.
- How you find your place on the curve: cross-validation that respects the time ordering, learning curves, and watching the gap between in-sample and out-of-sample performance. If the gap is large you are on the variance side.
- One honest complication: the classic U-shaped curve is not the whole story. Very overparameterised models can show double descent, where test error falls again past the interpolation threshold. That is real in vision and language. I have not seen it be useful on noisy financial data, where the tiny signal means regularisation still dominates.
Where candidates lose it
Giving the textbook decomposition with no view on where financial data sits. Every candidate can recite bias plus variance. The differentiator is saying that low signal-to-noise pushes you towards simple models, and being able to say how you would diagnose which side you are on.
Expect next
- How would you diagnose which side of the tradeoff you are on?
- Why do simple models often win on financial data?
- What is double descent?
054What does PCA do, how do you choose the number of components, and what are its limitations on financial data?Quant researchRisk
Say this
It finds the orthogonal directions of maximum variance, which are the eigenvectors of the covariance matrix, and lets you describe the data with fewer numbers. Choose the number of components by explained variance, a scree elbow, or the Marchenko-Pastur bulk edge if you want a principled cutoff.
Then walk it
- Mechanically: eigendecompose the covariance or correlation matrix, or take the SVD of the centred data. Eigenvalues are the variance along each component, eigenvectors are the directions.
- Correlation versus covariance matters. On assets with wildly different volatilities, PCA on the covariance matrix is dominated by the most volatile names, so standardise first unless the scale is meaningful.
- Concrete example everyone in rates knows: PCA on the yield curve gives level, slope and curvature, explaining roughly 90, 8 and 2 percent of variance. On equities the first component is the market, explaining 25 to 40 percent depending on the regime, and it rises sharply in a crisis.
- Choosing k: cumulative explained variance at 90 or 95 percent, the scree elbow, or eigenvalues above the random matrix bulk edge, which is the statistically defensible version because it separates signal from estimation noise.
- Limitations, and these are the answer to the real question. PCA maximises variance, not predictive power, so the components need not have anything to do with your target. It is unstable: eigenvectors rotate sample to sample when eigenvalues are close, so your factor two and factor three swap places. It assumes linearity. And the components are usually uninterpretable outside a structured setting like the yield curve, which makes them awkward to risk-manage.
Where candidates lose it
Describing PCA as dimensionality reduction and stopping. Two things get graded: that it is unsupervised so high-variance directions are not necessarily predictive, and that you must standardise when scales differ. Also have a real example ready, because level-slope-curvature or the equity market factor proves you have used it rather than read about it.
Expect next
- Why is PCA not necessarily good for prediction?
- What does the first principal component of an equity universe represent, and what happens to it in a crisis?
- How is PCA related to a factor risk model?
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

