Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
030You draw n independent uniforms on zero to one. What are the expected values of the maximum and the minimum, and of the kth smallest?Quant researchQuant trading
Say this
The maximum has mean n/(n+1), the minimum 1/(n+1), and the kth smallest k/(n+1). The n points cut the interval into n plus 1 gaps that are exchangeable, so each gap averages 1/(n+1).
Then walk it
- Derive the max directly: P(max at most x) is x to the n, so the density is n x to the n minus 1, and the integral of x times that from 0 to 1 is n/(n+1).
- The gap argument is faster and generalises. The n order statistics plus the two endpoints create n plus 1 spacings, which are exchangeable with total length 1, so each has mean 1/(n+1). The kth order statistic is the sum of the first k spacings, hence k/(n+1).
- The kth order statistic is Beta(k, n minus k plus 1), which gives you the variance too: k(n-k+1) over ((n+1) squared (n+2)).
- Numbers: with 10 draws the max averages 0.909 and the min 0.091. With 100 draws the max averages 0.990. The max creeps to the boundary at rate 1/n, which is why extreme-value estimates converge slowly.
- Why a quant desk cares: the max of n draws is your model for the best of n signals, the worst drawdown of n periods, and the winning quote in an auction with n bidders. And it explains selection bias, because the best of a hundred backtests looks good even when none of them has any edge.
Where candidates lose it
Answering only for the max with a calculus derivation and then being stuck on the general kth. Learn the spacings argument, it gives all of them at once. And be ready to connect it to selection bias, because the practical follow-up is almost always about why the best of many strategies overstates its own quality.
Expect next
- What is the variance of the maximum?
- What is the expected range, max minus min?
- How does this explain the selection bias in picking the best of a hundred backtests?
056What is maximum likelihood estimation, and when would you prefer method of moments?Quant researchRisk
Say this
MLE picks the parameters that make the observed data most probable under your assumed distribution. It is asymptotically efficient if the model is right, which is exactly the condition that makes method of moments attractive when it is not.
Then walk it
- MLE: maximise the log likelihood, which is the sum of log densities. Under regularity conditions it is consistent, asymptotically normal, and attains the Cramer-Rao bound, with variance given by the inverse Fisher information.
- Method of moments: match sample moments to their theoretical expressions and solve. Generalised method of moments extends this to more moment conditions than parameters, weighting them optimally, and it needs no full distributional assumption.
- So the tradeoff is efficiency versus robustness. MLE uses the whole density, so it extracts every bit of information and pays for it with sensitivity to misspecification. GMM uses only the moments you trust.
- Concrete case: fitting a distribution to daily returns. MLE under a normal assumption gives you the sample mean and variance and will be badly misled by the tails. MLE under a Student t estimates the degrees of freedom and is much better behaved. GMM on a few robust moments avoids committing to either.
- Practical points worth raising: MLE can be biased in small samples even when consistent, the classic example being the variance estimator with n rather than n minus 1 in the denominator. And numerically you should always check the Hessian at the optimum, because a flat likelihood means your parameter is not identified, which is common in GARCH and regime models.
Where candidates lose it
Describing MLE as the best estimator without the qualifier if the model is correctly specified. That caveat is the entire content of the comparison. Also be ready for the small-sample bias point, since MLE being biased while still consistent catches people who have only memorised the asymptotic properties.
Expect next
- Give me an example where MLE is biased.
- What is the Cramer-Rao bound?
- How would you fit a Student t to returns, and what does the estimated degrees of freedom tell you?
058Daily equity returns are not normal. How are they different, and what do you do about it?Quant researchRisk
Say this
They have fat tails, negative skew and volatility clustering. Daily equity index kurtosis is typically 5 to 10 against 3 for a normal, so moves the normal says should happen once a century happen every few years.
Then walk it
- Put a number on it. A normal assigns a five standard deviation daily move a probability of about one in 3.5 million, roughly once in 14,000 years of trading. The S&P has had several since 1950. The tails are not slightly wrong, they are wrong by orders of magnitude.
- Negative skew: large down moves are bigger and faster than large up moves. That is why index option skew exists and why puts are persistently richer than calls in implied vol terms.
- Volatility clustering means part of the unconditional fat tail is a mixture effect. Returns standardised by a GARCH-type conditional volatility are much closer to normal, though still fat-tailed, which tells you some but not all of the kurtosis is time-varying vol rather than genuinely fat conditional tails.
- What I would do depends on the use. For risk: empirical quantiles, a Student t or a generalised Pareto fit to the tail via extreme value theory, and expected shortfall rather than value at risk, because expected shortfall is sensitive to how bad the tail is. For pricing: a model with jumps or stochastic volatility rather than plain Black-Scholes.
- And the aggregation point: monthly returns are considerably closer to normal than daily returns because of the CLT, so the right distributional assumption depends on your horizon. That is worth saying because it stops the conversation becoming a generic tails are fat sermon.
Where candidates lose it
Saying fat tails and stopping. Quantify it, because the five sigma comparison is what makes the point land. Also do not forget the skew, since symmetric fat tails would not explain the option skew, and be ready to distinguish unconditional fat tails from conditional heteroskedasticity.
Expect next
- How much of the kurtosis is explained by volatility clustering?
- What is expected shortfall and why prefer it to value at risk?
- How does this show up in the option surface?
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

