Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
014Here is a game. What is the expected value of winning under three different strategies, and which one would you choose?Jane StreetTrading · London · 2025OptiverGeneralist · Chicago · 2025
Say this
Set up the state and the decision rule before you compute anything, price each strategy with a clean conditional expectation, then choose on expected value first and on variance and ruin risk second. Say the comparison out loud as you go so the interviewer can follow your bookkeeping.
Then walk it
- Step one, define the state precisely: what you know when you decide, and what the payoff function is. Most errors in these problems are specification errors, not arithmetic.
- Step two, price each strategy by conditioning on the first move. E of payoff equals the sum over first outcomes of probability times conditional value. If the game is repeated or recursive, write V in terms of V and solve the fixed point.
- Step three, do the arithmetic in fractions, not decimals. Fractions let the interviewer audit you and they do not accumulate error.
- Step four, choose. If one strategy dominates on expected value, say so and stop. If they are close, break the tie on the second moment: I would take the lower-variance strategy at the same expected value, and I would pay a small amount of expected value to avoid a path that can lose more than my stake.
- Then state the assumption you are relying on, unprompted: whether you may stop adaptively, whether the game is repeated, and whether the payoff is linear in money. Those three change the answer more than the arithmetic does.
Where candidates lose it
Diving into arithmetic before defining the state, and then losing track of which branch you are on. The other failure is picking the highest expected value without a word about variance. A trading floor cares about the distribution of outcomes, so say which strategy you would actually run with real money and why.
Expect next
- Now suppose you can play the game a hundred times. Does your choice change?
- What if the payoff were doubled but the probability halved?
- What is the variance of your preferred strategy?
Reported by candidates at Jane Street (Trading, London, 2025); Optiver (Generalist, Chicago, 2025). Source: Wall Street Oasis.
015Two games have exactly the same expected value. Which one would you choose to play?Akuna CapitalSales and Trading · Chicago · 2025Belvedere TradingProp Trading · Chicago · 2022
Say this
If the expected values tie, I choose on variance, on how many times I get to play, and on whether any outcome can wipe me out. As a one-off with a fixed stake I take the lower-variance game. Repeated many times with the ability to size, I might prefer the higher-variance one.
Then walk it
- First, ask the question the interviewer wants you to ask: how many times do I get to play, and can I choose my size? Those two facts change the answer completely.
- One shot, fixed size: take low variance. Same mean, less dispersion, strictly better under any concave utility, and a trader's utility is concave because a bad first day costs them their limits.
- Repeated, and I can size: variance becomes something I can dial. Kelly says bet a fraction proportional to edge over variance, so the high-variance game just gets a smaller position. Per unit of risk they may be identical.
- Then the killer criterion, which is ruin. If one game has any probability of a loss larger than my capital, its long-run growth rate is minus infinity regardless of its expected value. Expected value is a bad objective when the bet is not repeatable.
- One more real consideration: correlation with everything else I have on. A game with the same mean and variance but zero correlation to my book is worth more than one that doubles my existing exposure. On a desk that is usually the deciding factor.
Where candidates lose it
Saying I am indifferent because the expected values are equal. That answers the arithmetic and fails the question, which is about risk preference. Also do not just say I prefer lower variance and stop, because the interesting answer depends on repetition, sizing and ruin. Ask the clarifying question first.
Expect next
- What if you could play one of them a thousand times?
- How would you size each one?
- Explain the Kelly criterion and why traders bet less than Kelly.
Reported by candidates at Akuna Capital (Sales and Trading, Chicago, 2025); Belvedere Trading (Prop Trading, Chicago, 2022). Source: Wall Street Oasis.
016You win a hundred dollars if you roll a ten with two dice. How much would you risk to play?Akuna CapitalTrading · Chicago · 2025
Say this
Fair value is eight dollars and a third. Three of the 36 outcomes make ten, so probability is 1/12 and the expected payoff is 100 over 12. I would pay up to about seven to leave edge, and if I am being asked to make a two-way price I would quote around 7 at 9.
Then walk it
- Count the outcomes: 6-4, 4-6, 5-5. Three ways out of 36, so 1/12, about 8.33 percent.
- Expected payoff 100 times 1/12 equals 8.33. That is fair value, and fair value is where you break even, not where you trade.
- So I need edge. I would bid 7 and offer 9 if I have to two-way it, which is about a dollar and a half of edge either side, roughly fifteen percent of fair value. That width reflects the fact that I cannot hedge a one-off die roll.
- Size matters as much as price. This bet has a standard deviation of about 28 dollars against a mean of 8.33, which is a terrible ratio. I would do it small even at a good price, and I would want to repeat it many times rather than do it once large.
- If the game is repeatable and I can do it a thousand times, I pay closer to 8. The edge I demand is compensation for variance I cannot diversify, and repetition diversifies it.
Where candidates lose it
Answering with the fair value of 8.33 as if that were your bid. A trader never pays fair value, and saying eight and a third is what I would risk tells the interviewer you do not understand where the money comes from. Quote a price below fair value, name your width, and say your size.
Expect next
- Now make me a two-way market on it and I will trade you.
- What if I could roll a hundred times?
- What is the standard deviation of your P&L on one play?
Reported by candidates at Akuna Capital (Trading, Chicago, 2025). Source: Wall Street Oasis.
037You stand on a road and watch cars drive past. How would you estimate the parameter of the underlying distribution?Jump TradingQuantitative Research · Chicago · 2018
Say this
First I would state the model: arrivals as a Poisson process with rate lambda, so inter-arrival times are exponential with mean 1 over lambda. Then the maximum likelihood estimate of lambda is just the count divided by the observation time, and its standard error is lambda over the square root of the count.
Then walk it
- Model choice first, and justify it: independent arrivals at a constant rate with no memory gives a Poisson process. That is reasonable on a quiet road, and clearly wrong near a traffic light where cars arrive in platoons.
- MLE: for n arrivals in time T, lambda hat is n over T. It is unbiased, and the variance is lambda over T, so the relative standard error is 1 over the square root of n. Twenty-five cars gives you a 20 percent standard error, a hundred cars gives 10 percent.
- That tells you the sample size you need before you open your mouth about precision. If someone wants the rate to five percent, you need 400 cars.
- Now the diagnostics, which are what a research interview is actually about. Plot the inter-arrival times and check whether they look exponential. Over-dispersion, meaning variance above the mean of the counts, tells you arrivals are clustered and Poisson is wrong. Then I would go to a Cox process or a Hawkes process with self-excitation.
- And I would flag the estimation trap: if instead I sampled by picking a random moment and measuring the gap I happened to land in, I would oversample long gaps. That is the inspection paradox, and it biases the mean gap upward by a factor of one plus the squared coefficient of variation. It is the same bias that makes waiting times feel longer than the timetable says.
Where candidates lose it
Jumping to a formula without stating the model or checking it. The interviewer wants model, estimator, standard error, then diagnostics. The specific failure mode they are hunting is the inspection paradox, so mention length-biased sampling unprompted. Hawkes processes are the right answer for clustered arrivals and they are also how trade arrivals actually behave in markets.
Expect next
- How would you test whether the Poisson assumption holds?
- What if the cars arrive in clusters?
- How long do you need to watch to get the rate within five percent?
Reported by candidates at Jump Trading (Quantitative Research, Chicago, 2018). Source: Wall Street Oasis.
040Here is a dataset. Analyse it using probability metrics and tell me what you find.Jane StreetCredit Risk · London · 2025
Say this
I would spend the first third of the time on the data itself before any modelling: shape, missingness, duplicates, timestamps, and the univariate distributions. Then state a hypothesis, test it, and report the effect size with an honest uncertainty. Narrate every step, because the interviewer is grading the process, not the punchline.
Then walk it
- Start with the boring checks, out loud. Row count, date range, obvious duplicates, missing values and whether they are missing at random, and whether any column is a leak of the outcome. Most real findings in interviews of this kind are data artefacts.
- Then univariates: mean, median, standard deviation, skew, kurtosis, and the tails. Plot histograms and the empirical CDF. If a column is heavy-tailed or bimodal, say so, because it changes every subsequent choice.
- Then the relationship you were asked about. Give a point estimate plus a confidence interval, and prefer a plot to a coefficient. If the data are time-ordered, check for autocorrelation and regime change before quoting any p-value, because serial dependence inflates significance badly.
- Then the discipline: state your null, say what result would change your mind, and count how many hypotheses you have looked at. If you tested twenty things, say so and adjust.
- Close with what the data cannot tell you. A credit dataset with survivors only cannot tell you about defaults. Ending on the limitation is what separates an analyst from someone producing numbers, and in a live exercise it is the cheapest way to sound senior.
Where candidates lose it
Going straight to a model. Almost every candidate opens a regression and never looks at a histogram, then reports a spurious result driven by three outliers or a broken timestamp. Talk through the data integrity checks first, and say your uncertainty on every number you quote.
Expect next
- What would you check before trusting that correlation?
- How many hypotheses did you test, and how does that change your p-value?
- What would you want that is not in this dataset?
Reported by candidates at Jane Street (Credit Risk, London, 2025). Source: Wall Street Oasis.
045You test two hundred signals and three come back significant at the five percent level. What do you conclude?Quant researchQuant trading
Say this
That you have found nothing. Under a pure null you would expect ten false positives from two hundred tests at five percent, so three is fewer than chance. If anything the result is evidence against there being any signal at all.
Then walk it
- Expected false positives are 200 times 0.05 equals 10. Getting three significant results is below what noise alone produces, so the finding is not just unimpressive, it is worse than random.
- The right frame is family-wise error or false discovery rate. Bonferroni sets the threshold at 0.05 over 200, which is 0.00025, brutal but valid. Benjamini-Hochberg controls the expected proportion of false discoveries among the rejections and is much less conservative, which is usually the better choice when you are screening.
- The subtlety with financial signals: they are heavily correlated with each other, so the effective number of independent tests is far below 200. Bonferroni is then too harsh. I would estimate the effective number of tests, for example from the eigenvalue spectrum of the signal correlation matrix, or use a permutation or block-bootstrap null that preserves the correlation structure.
- The right test of whether anything survived is not a p-value at all. It is out-of-sample: hold back a period, or better a different market, and see whether the three signals still work with the sign you predicted.
- And the disclosure discipline, which is the answer a research head wants to hear: I would report the number of specifications tried alongside the result. The deflated Sharpe ratio and Harvey and Liu's work on multiple testing in finance both exist because the profession spent decades not doing this.
Where candidates lose it
Getting excited about the three and building a strategy on them. The whole question is whether you compute the expected number of false positives before you get attached. Say ten out of two hundred immediately, then talk about correlated tests, because that is where the technical depth is.
Expect next
- How would you estimate the effective number of independent tests?
- What is the deflated Sharpe ratio?
- How would you set up the experiment properly from the start?
049You need a covariance matrix for five hundred assets and you have two years of daily data. What is the problem and how do you fix it?Quant researchRisk
Say this
You have 500 assets and roughly 500 observations, so the sample covariance matrix is nearly singular and its smallest eigenvalues are garbage. Any optimiser will load up on exactly those directions, so you have to shrink or impose factor structure.
Then walk it
- Count the parameters: 500 times 501 over 2 is about 125,000 numbers estimated from 250,000 data points. The ratio of assets to observations, roughly one here, is what governs the damage, and the sample eigenvalue spectrum is badly biased even at a ratio of a quarter.
- Marchenko-Pastur describes exactly how the eigenvalues spread out. The largest are overstated and the smallest understated, and the smallest ones are the low-variance directions a mean-variance optimiser will concentrate in. That is why naive optimisers produce absurd leveraged long-short positions.
- Fix one, shrinkage. Ledoit-Wolf shrinks the sample matrix towards a structured target like a constant-correlation matrix, with an optimal intensity derived in closed form. Cheap, well-behaved and hard to beat as a default.
- Fix two, factor structure. Model returns as exposures to a few factors plus idiosyncratic noise, so the covariance is B times F times B transpose plus a diagonal. You have gone from 125,000 parameters to a few thousand. This is what every commercial risk model does.
- Fix three, random matrix filtering: keep the eigenvalues above the Marchenko-Pastur bulk edge as signal and replace the bulk with its average. Then state the practical check, which is out-of-sample portfolio variance rather than any in-sample fit statistic, because in-sample the sample matrix always wins and is always wrong.
Where candidates lose it
Saying you would just use the sample covariance matrix because two years is a lot of data. It is not, relative to 500 assets. The interviewer is testing whether you know that estimation error in the covariance matrix, not in the means, is what breaks portfolio optimisation in practice, and whether you can name shrinkage or factor models as the fix.
Expect next
- Why does the optimiser concentrate in the smallest eigenvalue directions?
- How do you choose the shrinkage intensity?
- How would you test whether your covariance matrix is any good?
050A colleague is excited about an R squared of 0.9 on a returns regression. What is your reaction?Quant researchQuant trading
Say this
Suspicion, not excitement. An R squared of 0.9 on returns almost always means a bug: a look-ahead leak, a regression of a price level on another price level, or the dependent variable included on the right-hand side. Real return predictability lives at an R squared of a fraction of a percent.
Then walk it
- Benchmark it. A genuinely good daily return predictor has an R squared around 0.001 to 0.01. A monthly cross-sectional factor model might reach a few percent. Anything above 0.1 on returns is a red flag rather than a result.
- Most likely causes in order: the target is in the features, the features are computed with future information, you regressed levels on levels where both are trending, or you regressed a variable on itself lagged by zero periods.
- The levels problem deserves a name. Two independent random walks regressed on each other will produce a high R squared and a significant t statistic almost every time, because the standard errors are wrong under non-stationarity. That is spurious regression, and it is Granger and Newbold's result.
- Also note what R squared does not tell you even when it is right: nothing about out-of-sample performance, nothing about economic significance, and it always rises when you add regressors, which is why adjusted R squared exists, penalising by (n-1)/(n-k-1).
- So what I would do: check for leakage first, difference the series and re-run, then look at out-of-sample R squared. And the thing worth knowing is that an out-of-sample R squared of 0.005 on daily returns, if it is real and tradeable, is a very good strategy. Small numbers are the norm and big numbers are bugs.
Where candidates lose it
Congratulating them. Knowing the realistic magnitude of return predictability is a strong signal that you have done real work, and not knowing it is a strong signal that you have not. Name look-ahead bias and spurious regression on levels as the two prime suspects.
Expect next
- What is a realistic R squared for a daily return forecast?
- Explain spurious regression between two random walks.
- What is out-of-sample R squared and how do you compute it honestly?
055When would you use gradient boosting on market data, and when would you stick with a linear model?Quant researchQuant trading
Say this
Boosting earns its keep when you have a lot of data, genuine non-linearity and interactions, and a target with enough signal to support the extra capacity. For low-frequency return prediction with a few hundred monthly observations I would use a regularised linear model almost every time.
Then walk it
- The case for trees: they capture interactions and thresholds automatically, handle mixed feature types, are insensitive to monotone transforms, and do not care about outliers in the features. On microstructure problems with millions of observations they genuinely win.
- The case against on returns: the signal-to-noise is so low that a flexible learner mostly memorises noise, and the model cannot extrapolate beyond the range it saw, which is exactly where the interesting market states live. A boosted tree trained through 2019 has no representation of March 2020.
- Data volume is the deciding variable. Daily cross-sectional data with 3,000 names times 20 years is 15 million rows and trees are viable. Monthly aggregate time series with 300 observations is not, no matter how you tune it.
- If I use boosting, I use it with heavy constraints: shallow trees of depth three to five, low learning rate, strong subsampling, early stopping on a purged time-series split, and monotonic constraints where I have a prior on the sign.
- And I would always run the regularised linear baseline first and report both. In practice the boosted model often adds a modest amount of out-of-sample R squared over a good linear model on financial data, which is a real gain but far from the step change people expect. Knowing that the gain is modest rather than transformative is the useful piece of experience here.
Where candidates lose it
Defaulting to whatever is fashionable with no reference to data volume or signal-to-noise. The interviewer wants a judgement, not a preference. Also name the extrapolation limitation of trees, because that is the specific reason they fail in a regime the training set never saw, which is when you most need the model.
Expect next
- How would you stop a boosted model overfitting on financial data?
- Why can trees not extrapolate, and when does that hurt you?
- How much out-of-sample improvement would make you switch from the linear model?
059A strategy shows a Sharpe ratio of 2 over one year. How much do you believe it?Quant researchQuant trading
Say this
Not much. The standard error of an annualised Sharpe estimated over T years is roughly the square root of (1 plus half the Sharpe squared) divided by T, so with one year and a Sharpe of 2 the standard error is about 1.7. The 95 percent interval runs from roughly minus 1.4 to 5.4, which comfortably includes zero.
Then walk it
- The formula, for iid normal returns: standard error of the Sharpe estimate is root of ((1 plus SR squared over 2) divided by T), with T in years for an annualised Sharpe.
- With T equal to 1 and SR equal to 2, that is the square root of (1 plus 2) over 1, which is the square root of 3, about 1.73. Two standard errors either side of the point estimate spans minus 1.4 to 5.4, so one year of data cannot even establish that the strategy makes money.
- Turn it around into the useful statement: to establish statistical significance at two standard errors you need roughly T of at least 4 over SR squared years. A Sharpe of 2 needs about a year to be marginally significant, a Sharpe of 1 needs four years, and a Sharpe of 0.5 needs sixteen years. Most equity factors fall in that last bucket, which is why the factor literature is so contested.
- The estimation error is only half the problem. The other half is selection. If this strategy is the best of a hundred I tested, the honest benchmark is the expected maximum Sharpe under the null, which for a hundred trials is around 2.5 standard errors above zero. The deflated Sharpe ratio adjusts for exactly this.
- And the formula assumes iid normal returns. Autocorrelated returns, which is common in anything holding illiquid or smoothed positions, inflate the Sharpe substantially, and negative skew means the Sharpe misses the risk that actually matters. So I would also want the drawdown profile, the turnover, and the capacity before I believed anything.
Where candidates lose it
Treating a one-year Sharpe as a fact. This question separates people who have evaluated real strategies from people who have read about them. Give the standard error formula, invert it into how many years you need, and then raise selection bias yourself.
Expect next
- How many years would you need for a Sharpe of 0.5 to be significant?
- What if the returns are autocorrelated?
- What else would you want to see besides the Sharpe?
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

