Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
011A random variable is uniform on the interval zero to ten. What are its expected value and variance?Old Mission CapitalFinance · New York · 2018
Say this
Mean 5, variance 100 over 12, which is 8.33, so standard deviation about 2.89. For a uniform on a to b the mean is the midpoint and the variance is (b minus a) squared over 12.
Then walk it
- Mean by symmetry: the midpoint of 0 and 10 is 5. No integration needed.
- Variance from the formula (b-a) squared over 12: 100 over 12 equals 8.33, standard deviation 2.887.
- If you want to derive it, E of X squared is the integral of x squared over 10 from 0 to 10, which is 1000/30 equals 33.33. Subtract 25 and you get 8.33. Good to be able to do it either way.
- The 1/12 is worth carrying in your head because it recurs: a fair n-sided die has variance (n squared minus 1)/12, and the rounding error of a value rounded to the nearest tick has variance tick squared over 12. That last one comes up in real microstructure work.
- Practical note: the uniform has thin support and no tails, so it is a bad default for anything financial. The moment somebody hands you a uniform in a trading context, ask what it is meant to represent.
Where candidates lose it
Reaching for integration under time pressure and fumbling the arithmetic. Know the (b-a) squared over 12 form cold. Also do not quote variance when they asked for standard deviation or the other way round, and say which one you are giving.
Expect next
- What is the expected value of the maximum of two independent draws?
- What is the distribution of the sum of two independent uniforms?
- What is the variance of the rounding error when you round to the nearest penny?
Reported by candidates at Old Mission Capital (Finance, New York, 2018). Source: Wall Street Oasis.
013You have a feed of a hundred thousand data points and you know fifteen of them are missing, recorded as zeros at the end. If you pull a window, what is the probability of at least one missing value?Jump TradingProp Trading · Remote · 2022
Say this
Use the complement. For a sample of n points drawn without replacement from 100,000 of which 15 are bad, the probability of at least one bad is one minus the hypergeometric probability of none, which is one minus the product over i of (99,985 minus i)/(100,000 minus i). For small n that is well approximated by one minus (1 minus 0.00015) to the n.
Then walk it
- Always compute at least one as one minus none. Summing the cases is the slow road and it invites double counting.
- The exact object is hypergeometric: choose n from 99,985 good over choose n from 100,000. For n much smaller than 100,000 the with and without replacement answers agree to several decimals.
- Numbers give it life. p is 15 over 100,000, which is 0.00015. For a window of 1,000 points, one minus 0.99985 to the 1000 is about 13.9 percent. For a window of 100 it is about 1.5 percent. So this is a real problem, not a rounding issue.
- Useful shortcut: for small p and moderate n the answer is roughly n times p, capped by 1. A thousand times 0.00015 is 0.15, close to the exact 0.139, and the Poisson approximation 1 minus e to the minus 0.15 gives 0.1393, which is very close.
- The thing I would say next on a desk, because it is the real question: they are at the end of the series, which is not random at all. If they are the most recent 15 points, then any window containing the tail hits all 15 with certainty and every other window hits none. Position matters more than the count.
Where candidates lose it
Treating the missing points as randomly scattered when the question says they sit at the end. That is the detail being tested. Give the hypergeometric answer for the random case, then flag the structural point: trailing zeros are usually a feed-truncation artefact, so the right fix is to detect and drop the tail, not to price the probability.
Expect next
- How would you detect that the zeros are missing values rather than genuine zeros?
- What is the Poisson approximation and when does it break?
- How do you handle those points in a model without leaking future information?
Reported by candidates at Jump Trading (Prop Trading, Remote, 2022). Source: Wall Street Oasis.
022If X and Y are dependent, does that tell you anything about the relationship between X and Z?Tower Research CapitalProp Trading · New York · 2019
Say this
Nothing at all. Dependence is not transitive and it says nothing about a third variable you have not mentioned. X can be dependent on Y and completely independent of Z.
Then walk it
- Trivial counterexample: let X and Y be the same fair coin and let Z be a separate independent coin. X and Y are maximally dependent, X and Z are independent.
- The deeper point is that even if X depends on Y and Y depends on Z, X need not depend on Z. Let Y be X plus Z with X and Z independent. Y is dependent on both, and X and Z remain independent of each other.
- Correlation is a bit more constrained than dependence because the correlation matrix must be positive semi-definite. If corr(X,Y) is 0.9 and corr(Y,Z) is 0.9, then corr(X,Z) is bounded below by about 0.62. So high correlations do restrict the third pair, but only through that PSD constraint, and dependence in general carries no such bound.
- The formula for the bound: rho_xz is at least rho_xy times rho_yz minus the square root of (1 minus rho_xy squared)(1 minus rho_yz squared). Plug in 0.9 and 0.9 and you get 0.81 minus 0.19, which is 0.62.
- Why this matters on a desk: people assume that if two assets both correlate with a factor they must correlate with each other. If the loadings are moderate, say 0.5 and 0.5, the bound is minus 0.5, so they can be strongly negatively correlated. That mistake shows up in risk models constantly.
Where candidates lose it
Answering yes because it feels like dependence should chain. Give the counterexample in one breath, then earn the extra credit with the correlation bound, because the interviewer's follow-up is almost always the correlation version. And be precise that zero correlation does not mean independence, only the converse holds.
Expect next
- Now with correlations. If corr(X,Y) is 0.9 and corr(Y,Z) is 0.9, what do you know about corr(X,Z)?
- Give me an example of zero correlation with strong dependence.
- What is conditional independence and why does it matter for factor models?
Reported by candidates at Tower Research Capital (Prop Trading, New York, 2019). Source: Wall Street Oasis.
026You start with fifty dollars and bet a dollar on a fair coin each time. What is the probability you reach a hundred before going broke, and how does it change if the coin is slightly against you?Quant tradingQuant research
Say this
In a fair game it is exactly one half, because your wealth is a martingale and the stopping value must average back to fifty. Tilt the odds slightly against you and the probability collapses, not linearly but exponentially in the number of steps.
Then walk it
- Fair case: wealth is a martingale, so by optional stopping, 50 equals 100 times p plus 0 times (1 minus p), giving p equal to 0.5. In general starting at a with an upper barrier b, the probability is a over b.
- Biased case: with win probability q the hitting probability is (1 minus r to the a) over (1 minus r to the b) where r is (1-q)/q.
- Put a number on it. At q equal to 0.49, r is about 1.0408. With a equal to 50 and b equal to 100, the probability of reaching 100 drops to roughly 12 percent. A one percent edge against you turns a coin flip into a 1-in-8 shot.
- That sensitivity is the entire lesson. Expected value per bet is minus two cents, which sounds trivial, but over the hundreds of bets you need to walk the barrier it compounds into near certainty of ruin.
- And the practical version on a desk: expected time to absorption in the fair case is a times (b minus a), so 50 times 50 equals 2,500 bets. Casinos and market makers both live on this asymmetry. Small edge, high repetition, deep pockets.
Where candidates lose it
Giving a over b and stopping. The interesting content is how brutally the biased case differs, and candidates who cannot state the r to the power formula usually also guess that a one percent edge changes the answer by about one percent. It changes it from 50 percent to 12 percent. Put a number on it.
Expect next
- What is the expected number of bets until you stop?
- What happens if you bet your whole stack each time instead?
- How does this relate to a trader's drawdown limit?
027What is a martingale, and how would you use optional stopping to solve a problem?Quant researchQuant trading
Say this
A martingale is a process whose expected next value, given everything you know now, equals its current value. Optional stopping says that for a suitably bounded stopping time, the expected value at the stopping time equals the starting value, which is what turns a hard path-dependent question into one line of algebra.
Then walk it
- Formally: E of X_{n+1} given the filtration F_n equals X_n. No drift, conditional on history. It is not the same as independence, and increments need not be identically distributed.
- Optional stopping needs a condition, and you should name one: bounded stopping time, or bounded increments plus finite expected stopping time, or uniform integrability. Without it the theorem fails, and the classic failure is the doubling strategy, where a stopping time that is finite with probability one still produces E of X_tau equal to 1 rather than 0.
- How I use it: find a quantity that is conserved in expectation, then evaluate it at the stopping time. Gambler's ruin falls out immediately from wealth being a martingale.
- A second example, expected time in a symmetric random walk: W_n squared minus n is a martingale, so E of tau equals E of W_tau squared. With barriers at 0 and b starting from a, that gives E of tau equal to a(b minus a) in a line.
- And the reason it matters beyond puzzles: risk-neutral pricing is exactly the statement that the discounted price is a martingale under the pricing measure. Delta hedging is the construction of that martingale. If you can say that connection, the puzzle answer becomes a conversation about derivatives.
Where candidates lose it
Defining a martingale as a fair game and stopping there, or applying optional stopping without checking the integrability condition. Interviewers at the good shops will hand you the doubling strategy specifically to see whether you know why the theorem does not apply. Name the condition before you use the theorem.
Expect next
- Why does optional stopping fail for the doubling strategy?
- Is the square of a martingale a martingale?
- Connect this to risk-neutral pricing.
030You draw n independent uniforms on zero to one. What are the expected values of the maximum and the minimum, and of the kth smallest?Quant researchQuant trading
Say this
The maximum has mean n/(n+1), the minimum 1/(n+1), and the kth smallest k/(n+1). The n points cut the interval into n plus 1 gaps that are exchangeable, so each gap averages 1/(n+1).
Then walk it
- Derive the max directly: P(max at most x) is x to the n, so the density is n x to the n minus 1, and the integral of x times that from 0 to 1 is n/(n+1).
- The gap argument is faster and generalises. The n order statistics plus the two endpoints create n plus 1 spacings, which are exchangeable with total length 1, so each has mean 1/(n+1). The kth order statistic is the sum of the first k spacings, hence k/(n+1).
- The kth order statistic is Beta(k, n minus k plus 1), which gives you the variance too: k(n-k+1) over ((n+1) squared (n+2)).
- Numbers: with 10 draws the max averages 0.909 and the min 0.091. With 100 draws the max averages 0.990. The max creeps to the boundary at rate 1/n, which is why extreme-value estimates converge slowly.
- Why a quant desk cares: the max of n draws is your model for the best of n signals, the worst drawdown of n periods, and the winning quote in an auction with n bidders. And it explains selection bias, because the best of a hundred backtests looks good even when none of them has any edge.
Where candidates lose it
Answering only for the max with a calculus derivation and then being stuck on the general kth. Learn the spacings argument, it gives all of them at once. And be ready to connect it to selection bias, because the practical follow-up is almost always about why the best of many strategies overstates its own quality.
Expect next
- What is the variance of the maximum?
- What is the expected range, max minus min?
- How does this explain the selection bias in picking the best of a hundred backtests?
033A test for a disease is 99 percent accurate and the disease affects one in ten thousand people. You test positive. What is the probability you have it?Quant researchQuant trading
Say this
About one percent. Out of a million people, 100 are sick and about 99 of them test positive, while 999,900 are healthy and about 9,999 of them test positive falsely. So 99 out of roughly 10,098 positives are real, which is 0.98 percent.
Then walk it
- Do it in counts, not Bayes notation. A population of a million makes the arithmetic trivial and the answer intuitive.
- The formula check: P(sick given positive) equals 0.0001 times 0.99 divided by (0.0001 times 0.99 plus 0.9999 times 0.01), which is 0.000099 over 0.010098, about 0.0098.
- The driver is base rate. False positives from the huge healthy population swamp the true positives from the tiny sick population. At a prevalence of 1 in 10,000 and a 1 percent false positive rate, you get a hundred false positives for every true one before adjusting for sensitivity.
- So the useful quantity is the likelihood ratio: 0.99 over 0.01 equals 99. It multiplies your prior odds of 1 in 9,999 into posterior odds of about 99 in 9,999, which is 1 percent. Thinking in odds and likelihood ratios is far faster than the fraction form.
- Where this shows up in trading: any rare-event detector, from fraud flags to regime-change signals to strategy alerts. A signal with 99 percent accuracy on a one-in-ten-thousand event fires 99 false alarms per real one, which is why alert systems get ignored.
Where candidates lose it
Answering 99 percent. The second trap is being sloppy about what 99 percent accurate means, since sensitivity and specificity need not be equal. State your reading, do it in counts per million, and name base rate neglect as the reason the intuitive answer is wrong by two orders of magnitude.
Expect next
- What prevalence would make the positive predictive value fifty percent?
- You test positive twice. Now what?
- How does this apply to a trading signal that fires rarely?
036The sample variance with the n minus one correction is unbiased. Is its square root an unbiased estimator of the standard deviation?Squarepoint CapitalQuantitative Research · London · 2026
Say this
No. The square root is concave, so by Jensen's inequality the expected square root is strictly less than the square root of the expected value. The sample standard deviation is biased downwards, always, for any distribution with positive variance.
Then walk it
- Jensen: for a strictly concave g, E of g(X) is less than g of E of X unless X is degenerate. With g the square root and X the unbiased sample variance, E of s is less than sigma.
- Size the bias for normal data. E of s equals c4(n) times sigma, where c4 is a known constant involving gamma functions. At n equal to 2, c4 is about 0.798, so you understate sigma by 20 percent. At n equal to 10 it is 0.9727, a 2.7 percent understatement. At n equal to 30 it is 0.9914.
- So the bias is order 1/(4n) and it vanishes as n grows. It is a real problem for short samples and irrelevant for long ones.
- Unbiasedness is also not preserved under any nonlinear transform, which is the general lesson. The unbiased estimator of sigma squared does not give you an unbiased estimator of sigma, or of 1/sigma, or of log sigma.
- Where this bites on a desk: annualised volatility estimated from a few weeks of data, and any Sharpe ratio, since the Sharpe divides by s. Understating s inflates the Sharpe, so short-sample Sharpes are biased upwards. That is worth saying because it connects a textbook Jensen question to a live problem in strategy evaluation.
Where candidates lose it
Saying yes because the variance estimator is unbiased. Unbiasedness does not survive a nonlinear function. Name Jensen explicitly, give the direction of the bias, and quantify it with c4 for at least one small n. The follow-up about Sharpe ratios is where the real conversation is, so get there yourself.
Expect next
- How would you correct it?
- What does that imply for a Sharpe ratio estimated on a short sample?
- Is the sample correlation coefficient unbiased?
Reported by candidates at Squarepoint Capital (Quantitative Research, London, 2026). Source: Wall Street Oasis.
038How would you fix violations of the OLS assumptions?AQR Capital ManagementInvestments · Greenwich · 2022
Say this
Depends which assumption breaks, and the fixes fall into two very different classes: violations that only break your standard errors, and violations that break the coefficients themselves. The first class you patch; the second class you have to re-specify the model.
Then walk it
- Heteroskedasticity and autocorrelated errors: coefficients stay unbiased, only inference is wrong. Fix with White or Newey-West robust standard errors, or clustered errors if the dependence is by group. Cheap fix, always worth doing on financial data.
- Endogeneity, meaning a regressor correlated with the error, whether from omitted variables, simultaneity or measurement error: this biases the coefficients and no standard error fix helps. You need an instrument, a control for the omitted factor, a fixed effect, or a different design.
- Multicollinearity: coefficients are still unbiased but the variances explode and the signs flip sample to sample. Drop or combine the collinear regressors, use ridge, or work with principal components. And check the variance inflation factors before you interpret anything.
- Non-normal or fat-tailed errors: inference is still fine asymptotically thanks to the CLT, but outliers dominate the fit because OLS minimises squares. Use robust regression, Huber loss or quantile regression, and always look at the influence diagnostics.
- Non-linearity: add the relevant transform or interaction rather than pretending it away. And I would say the order I actually work in on real data: plot residuals against fitted values and against time first, because most violations announce themselves visually before any test does.
Where candidates lose it
Listing fixes without separating what biases the coefficients from what only biases the standard errors. That distinction is the question. Slapping Newey-West errors on an endogenous regression is a common and useless move, and an interviewer at a research shop will push on exactly that.
Expect next
- Which of those actually biases your coefficients?
- How do you detect endogeneity if you have no instrument?
- What do you do when the residuals are fat-tailed and autocorrelated at the same time?
Reported by candidates at AQR Capital Management (Investments, Greenwich, 2022). Source: Wall Street Oasis.
039What are the differences between Lasso and Ridge regression?Tower Research CapitalTrading · Princeton · 2018
Say this
Both add a penalty on coefficient size to trade variance for bias. Ridge penalises the sum of squares and shrinks everything smoothly towards zero without eliminating anything. Lasso penalises the sum of absolute values and sets coefficients exactly to zero, so it selects features.
Then walk it
- The geometry explains it. The L1 constraint region is a diamond with corners on the axes, so the solution tends to land on a corner, which means a zero coefficient. The L2 region is a ball with no corners, so solutions are interior and nothing is exactly zero.
- Ridge has a closed form, beta equals (X'X plus lambda I) inverse X'y, which is why it also fixes a singular X'X. Lasso has no closed form and needs coordinate descent or LARS.
- Correlated predictors behave very differently. Ridge splits the weight across a group of correlated features, which is stable. Lasso arbitrarily picks one and zeroes the rest, which is unstable across samples. Elastic net, which mixes both penalties, exists precisely to get sparsity without that instability.
- In a Bayesian reading, ridge is a Gaussian prior on the coefficients and lasso is a Laplace prior. The Laplace prior's spike at zero is what produces exact zeros.
- What I would say about which to use on financial data: predictors are usually highly correlated and the signal-to-noise ratio is awful, so ridge or elastic net typically beats pure lasso out of sample. Lasso is attractive when you need an interpretable short list of factors, but do not confuse the features it selected with the features that matter, because a slightly different sample gives you a different list.
Where candidates lose it
Stopping at L1 gives sparsity, L2 does not. Everyone says that. The differentiators are the diamond-versus-ball geometry, the behaviour under correlated predictors, and the Bayesian priors. Also always say that both require standardised features, because the penalty is scale-dependent and forgetting to standardise silently ruins the fit.
Expect next
- What is elastic net for?
- How do you choose lambda?
- Why do you have to standardise your features first?
Reported by candidates at Tower Research Capital (Trading, Princeton, 2018). Source: Wall Street Oasis.
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

