Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
011A random variable is uniform on the interval zero to ten. What are its expected value and variance?Old Mission CapitalFinance · New York · 2018
Say this
Mean 5, variance 100 over 12, which is 8.33, so standard deviation about 2.89. For a uniform on a to b the mean is the midpoint and the variance is (b minus a) squared over 12.
Then walk it
- Mean by symmetry: the midpoint of 0 and 10 is 5. No integration needed.
- Variance from the formula (b-a) squared over 12: 100 over 12 equals 8.33, standard deviation 2.887.
- If you want to derive it, E of X squared is the integral of x squared over 10 from 0 to 10, which is 1000/30 equals 33.33. Subtract 25 and you get 8.33. Good to be able to do it either way.
- The 1/12 is worth carrying in your head because it recurs: a fair n-sided die has variance (n squared minus 1)/12, and the rounding error of a value rounded to the nearest tick has variance tick squared over 12. That last one comes up in real microstructure work.
- Practical note: the uniform has thin support and no tails, so it is a bad default for anything financial. The moment somebody hands you a uniform in a trading context, ask what it is meant to represent.
Where candidates lose it
Reaching for integration under time pressure and fumbling the arithmetic. Know the (b-a) squared over 12 form cold. Also do not quote variance when they asked for standard deviation or the other way round, and say which one you are giving.
Expect next
- What is the expected value of the maximum of two independent draws?
- What is the distribution of the sum of two independent uniforms?
- What is the variance of the rounding error when you round to the nearest penny?
Reported by candidates at Old Mission Capital (Finance, New York, 2018). Source: Wall Street Oasis.
030You draw n independent uniforms on zero to one. What are the expected values of the maximum and the minimum, and of the kth smallest?Quant researchQuant trading
Say this
The maximum has mean n/(n+1), the minimum 1/(n+1), and the kth smallest k/(n+1). The n points cut the interval into n plus 1 gaps that are exchangeable, so each gap averages 1/(n+1).
Then walk it
- Derive the max directly: P(max at most x) is x to the n, so the density is n x to the n minus 1, and the integral of x times that from 0 to 1 is n/(n+1).
- The gap argument is faster and generalises. The n order statistics plus the two endpoints create n plus 1 spacings, which are exchangeable with total length 1, so each has mean 1/(n+1). The kth order statistic is the sum of the first k spacings, hence k/(n+1).
- The kth order statistic is Beta(k, n minus k plus 1), which gives you the variance too: k(n-k+1) over ((n+1) squared (n+2)).
- Numbers: with 10 draws the max averages 0.909 and the min 0.091. With 100 draws the max averages 0.990. The max creeps to the boundary at rate 1/n, which is why extreme-value estimates converge slowly.
- Why a quant desk cares: the max of n draws is your model for the best of n signals, the worst drawdown of n periods, and the winning quote in an auction with n bidders. And it explains selection bias, because the best of a hundred backtests looks good even when none of them has any edge.
Where candidates lose it
Answering only for the max with a calculus derivation and then being stuck on the general kth. Learn the spacings argument, it gives all of them at once. And be ready to connect it to selection bias, because the practical follow-up is almost always about why the best of many strategies overstates its own quality.
Expect next
- What is the variance of the maximum?
- What is the expected range, max minus min?
- How does this explain the selection bias in picking the best of a hundred backtests?
036The sample variance with the n minus one correction is unbiased. Is its square root an unbiased estimator of the standard deviation?Squarepoint CapitalQuantitative Research · London · 2026
Say this
No. The square root is concave, so by Jensen's inequality the expected square root is strictly less than the square root of the expected value. The sample standard deviation is biased downwards, always, for any distribution with positive variance.
Then walk it
- Jensen: for a strictly concave g, E of g(X) is less than g of E of X unless X is degenerate. With g the square root and X the unbiased sample variance, E of s is less than sigma.
- Size the bias for normal data. E of s equals c4(n) times sigma, where c4 is a known constant involving gamma functions. At n equal to 2, c4 is about 0.798, so you understate sigma by 20 percent. At n equal to 10 it is 0.9727, a 2.7 percent understatement. At n equal to 30 it is 0.9914.
- So the bias is order 1/(4n) and it vanishes as n grows. It is a real problem for short samples and irrelevant for long ones.
- Unbiasedness is also not preserved under any nonlinear transform, which is the general lesson. The unbiased estimator of sigma squared does not give you an unbiased estimator of sigma, or of 1/sigma, or of log sigma.
- Where this bites on a desk: annualised volatility estimated from a few weeks of data, and any Sharpe ratio, since the Sharpe divides by s. Understating s inflates the Sharpe, so short-sample Sharpes are biased upwards. That is worth saying because it connects a textbook Jensen question to a live problem in strategy evaluation.
Where candidates lose it
Saying yes because the variance estimator is unbiased. Unbiasedness does not survive a nonlinear function. Name Jensen explicitly, give the direction of the bias, and quantify it with c4 for at least one small n. The follow-up about Sharpe ratios is where the real conversation is, so get there yourself.
Expect next
- How would you correct it?
- What does that imply for a Sharpe ratio estimated on a short sample?
- Is the sample correlation coefficient unbiased?
Reported by candidates at Squarepoint Capital (Quantitative Research, London, 2026). Source: Wall Street Oasis.
043State the central limit theorem and tell me where it fails.Quant researchQuant trading
Say this
For independent identically distributed variables with finite mean and finite variance, the standardised sample mean converges in distribution to a standard normal. The key conditions are finite variance and enough independence, and both fail regularly in markets.
Then walk it
- Precisely: root n times (X bar minus mu) over sigma converges in distribution to N(0,1). Note it is the standardised mean that converges, and the rate is 1 over root n.
- Failure one, infinite variance. A Cauchy distribution has no variance and the sample mean of Cauchys is Cauchy again, no matter how large n is. Averaging buys you nothing. More generally, stable distributions with tail index alpha below 2 converge to a stable law, not a normal.
- Failure two, dependence. With strongly autocorrelated data the effective sample size is far below n, so you converge much more slowly and your standard errors are too small. Long-range dependence can break it entirely.
- Failure three, the rate in the tails. Even where the CLT holds, convergence is fastest in the middle and slowest in the tails, which is precisely where a risk manager needs accuracy. Berry-Esseen gives an error bound of order 1 over root n times the third absolute moment, so skewed data converges slowly.
- The practical version: daily equity returns have kurtosis of 5 to 10 and volatility clustering, so ten-day sums are much closer to normal than daily returns, but a 99.9 percent quantile computed from a normal assumption will still understate the tail badly. That is why value at risk models use empirical or extreme-value tails rather than leaning on the CLT.
Where candidates lose it
Stating the theorem without the finite variance condition, or claiming everything becomes normal for large n. Also do not confuse it with the law of large numbers, which is about convergence of the mean to a constant and needs only finite mean. Be ready to say what happens with infinite variance, because that is the follow-up.
Expect next
- What happens with a Cauchy distribution?
- How is that different from the law of large numbers?
- How large does n have to be in practice for returns data?
044What is a p-value, and what is it not?Quant researchRisk
Say this
It is the probability of seeing data at least as extreme as what you saw, assuming the null hypothesis is true. It is not the probability that the null is true, and it is not the probability you are wrong.
Then walk it
- The conditioning runs the wrong way from what people assume. A p-value is P(data given null), and what you actually want is P(null given data). Those are different objects and Bayes tells you the second depends on your prior.
- Concretely: if you test a thousand strategies of which fifty genuinely work, at a five percent significance level you get roughly 47 true discoveries and 47 false ones. A p-value of 0.05 in that setting means a coin flip on whether the finding is real.
- It also says nothing about effect size. With a million observations a completely useless one-basis-point edge will have a p-value of 0.0001. Significance is not importance, and in high-frequency data everything is significant.
- And it is only valid for a pre-specified test. Choosing the test after looking at the data, or stopping data collection when the p-value crosses 0.05, invalidates it completely.
- What I would report instead on a desk: the effect size with a confidence interval, out-of-sample performance, and how many specifications I tried. A p-value on its own is close to useless in a research process where hundreds of hypotheses get screened.
Where candidates lose it
Defining it as the probability the null is true. That is the single most common statistical error in finance interviews and it is disqualifying at a research shop. Also be ready with the multiple-testing consequence, because the interviewer's real target is whether you understand why published anomalies do not replicate.
Expect next
- So what significance level would you use if you screened a thousand signals?
- Explain the false discovery rate.
- What would you report instead of a p-value?
056What is maximum likelihood estimation, and when would you prefer method of moments?Quant researchRisk
Say this
MLE picks the parameters that make the observed data most probable under your assumed distribution. It is asymptotically efficient if the model is right, which is exactly the condition that makes method of moments attractive when it is not.
Then walk it
- MLE: maximise the log likelihood, which is the sum of log densities. Under regularity conditions it is consistent, asymptotically normal, and attains the Cramer-Rao bound, with variance given by the inverse Fisher information.
- Method of moments: match sample moments to their theoretical expressions and solve. Generalised method of moments extends this to more moment conditions than parameters, weighting them optimally, and it needs no full distributional assumption.
- So the tradeoff is efficiency versus robustness. MLE uses the whole density, so it extracts every bit of information and pays for it with sensitivity to misspecification. GMM uses only the moments you trust.
- Concrete case: fitting a distribution to daily returns. MLE under a normal assumption gives you the sample mean and variance and will be badly misled by the tails. MLE under a Student t estimates the degrees of freedom and is much better behaved. GMM on a few robust moments avoids committing to either.
- Practical points worth raising: MLE can be biased in small samples even when consistent, the classic example being the variance estimator with n rather than n minus 1 in the denominator. And numerically you should always check the Hessian at the optimum, because a flat likelihood means your parameter is not identified, which is common in GARCH and regime models.
Where candidates lose it
Describing MLE as the best estimator without the qualifier if the model is correctly specified. That caveat is the entire content of the comparison. Also be ready for the small-sample bias point, since MLE being biased while still consistent catches people who have only memorised the asymptotic properties.
Expect next
- Give me an example where MLE is biased.
- What is the Cramer-Rao bound?
- How would you fit a Student t to returns, and what does the estimated degrees of freedom tell you?
058Daily equity returns are not normal. How are they different, and what do you do about it?Quant researchRisk
Say this
They have fat tails, negative skew and volatility clustering. Daily equity index kurtosis is typically 5 to 10 against 3 for a normal, so moves the normal says should happen once a century happen every few years.
Then walk it
- Put a number on it. A normal assigns a five standard deviation daily move a probability of about one in 3.5 million, roughly once in 14,000 years of trading. The S&P has had several since 1950. The tails are not slightly wrong, they are wrong by orders of magnitude.
- Negative skew: large down moves are bigger and faster than large up moves. That is why index option skew exists and why puts are persistently richer than calls in implied vol terms.
- Volatility clustering means part of the unconditional fat tail is a mixture effect. Returns standardised by a GARCH-type conditional volatility are much closer to normal, though still fat-tailed, which tells you some but not all of the kurtosis is time-varying vol rather than genuinely fat conditional tails.
- What I would do depends on the use. For risk: empirical quantiles, a Student t or a generalised Pareto fit to the tail via extreme value theory, and expected shortfall rather than value at risk, because expected shortfall is sensitive to how bad the tail is. For pricing: a model with jumps or stochastic volatility rather than plain Black-Scholes.
- And the aggregation point: monthly returns are considerably closer to normal than daily returns because of the CLT, so the right distributional assumption depends on your horizon. That is worth saying because it stops the conversation becoming a generic tails are fat sermon.
Where candidates lose it
Saying fat tails and stopping. Quantify it, because the five sigma comparison is what makes the point land. Also do not forget the skew, since symmetric fat tails would not explain the option skew, and be ready to distinguish unconditional fat tails from conditional heteroskedasticity.
Expect next
- How much of the kurtosis is explained by volatility clustering?
- What is expected shortfall and why prefer it to value at risk?
- How does this show up in the option surface?
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

