Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
036The sample variance with the n minus one correction is unbiased. Is its square root an unbiased estimator of the standard deviation?Squarepoint CapitalQuantitative Research · London · 2026
Say this
No. The square root is concave, so by Jensen's inequality the expected square root is strictly less than the square root of the expected value. The sample standard deviation is biased downwards, always, for any distribution with positive variance.
Then walk it
- Jensen: for a strictly concave g, E of g(X) is less than g of E of X unless X is degenerate. With g the square root and X the unbiased sample variance, E of s is less than sigma.
- Size the bias for normal data. E of s equals c4(n) times sigma, where c4 is a known constant involving gamma functions. At n equal to 2, c4 is about 0.798, so you understate sigma by 20 percent. At n equal to 10 it is 0.9727, a 2.7 percent understatement. At n equal to 30 it is 0.9914.
- So the bias is order 1/(4n) and it vanishes as n grows. It is a real problem for short samples and irrelevant for long ones.
- Unbiasedness is also not preserved under any nonlinear transform, which is the general lesson. The unbiased estimator of sigma squared does not give you an unbiased estimator of sigma, or of 1/sigma, or of log sigma.
- Where this bites on a desk: annualised volatility estimated from a few weeks of data, and any Sharpe ratio, since the Sharpe divides by s. Understating s inflates the Sharpe, so short-sample Sharpes are biased upwards. That is worth saying because it connects a textbook Jensen question to a live problem in strategy evaluation.
Where candidates lose it
Saying yes because the variance estimator is unbiased. Unbiasedness does not survive a nonlinear function. Name Jensen explicitly, give the direction of the bias, and quantify it with c4 for at least one small n. The follow-up about Sharpe ratios is where the real conversation is, so get there yourself.
Expect next
- How would you correct it?
- What does that imply for a Sharpe ratio estimated on a short sample?
- Is the sample correlation coefficient unbiased?
Reported by candidates at Squarepoint Capital (Quantitative Research, London, 2026). Source: Wall Street Oasis.
037You stand on a road and watch cars drive past. How would you estimate the parameter of the underlying distribution?Jump TradingQuantitative Research · Chicago · 2018
Say this
First I would state the model: arrivals as a Poisson process with rate lambda, so inter-arrival times are exponential with mean 1 over lambda. Then the maximum likelihood estimate of lambda is just the count divided by the observation time, and its standard error is lambda over the square root of the count.
Then walk it
- Model choice first, and justify it: independent arrivals at a constant rate with no memory gives a Poisson process. That is reasonable on a quiet road, and clearly wrong near a traffic light where cars arrive in platoons.
- MLE: for n arrivals in time T, lambda hat is n over T. It is unbiased, and the variance is lambda over T, so the relative standard error is 1 over the square root of n. Twenty-five cars gives you a 20 percent standard error, a hundred cars gives 10 percent.
- That tells you the sample size you need before you open your mouth about precision. If someone wants the rate to five percent, you need 400 cars.
- Now the diagnostics, which are what a research interview is actually about. Plot the inter-arrival times and check whether they look exponential. Over-dispersion, meaning variance above the mean of the counts, tells you arrivals are clustered and Poisson is wrong. Then I would go to a Cox process or a Hawkes process with self-excitation.
- And I would flag the estimation trap: if instead I sampled by picking a random moment and measuring the gap I happened to land in, I would oversample long gaps. That is the inspection paradox, and it biases the mean gap upward by a factor of one plus the squared coefficient of variation. It is the same bias that makes waiting times feel longer than the timetable says.
Where candidates lose it
Jumping to a formula without stating the model or checking it. The interviewer wants model, estimator, standard error, then diagnostics. The specific failure mode they are hunting is the inspection paradox, so mention length-biased sampling unprompted. Hawkes processes are the right answer for clustered arrivals and they are also how trade arrivals actually behave in markets.
Expect next
- How would you test whether the Poisson assumption holds?
- What if the cars arrive in clusters?
- How long do you need to watch to get the rate within five percent?
Reported by candidates at Jump Trading (Quantitative Research, Chicago, 2018). Source: Wall Street Oasis.
040Here is a dataset. Analyse it using probability metrics and tell me what you find.Jane StreetCredit Risk · London · 2025
Say this
I would spend the first third of the time on the data itself before any modelling: shape, missingness, duplicates, timestamps, and the univariate distributions. Then state a hypothesis, test it, and report the effect size with an honest uncertainty. Narrate every step, because the interviewer is grading the process, not the punchline.
Then walk it
- Start with the boring checks, out loud. Row count, date range, obvious duplicates, missing values and whether they are missing at random, and whether any column is a leak of the outcome. Most real findings in interviews of this kind are data artefacts.
- Then univariates: mean, median, standard deviation, skew, kurtosis, and the tails. Plot histograms and the empirical CDF. If a column is heavy-tailed or bimodal, say so, because it changes every subsequent choice.
- Then the relationship you were asked about. Give a point estimate plus a confidence interval, and prefer a plot to a coefficient. If the data are time-ordered, check for autocorrelation and regime change before quoting any p-value, because serial dependence inflates significance badly.
- Then the discipline: state your null, say what result would change your mind, and count how many hypotheses you have looked at. If you tested twenty things, say so and adjust.
- Close with what the data cannot tell you. A credit dataset with survivors only cannot tell you about defaults. Ending on the limitation is what separates an analyst from someone producing numbers, and in a live exercise it is the cheapest way to sound senior.
Where candidates lose it
Going straight to a model. Almost every candidate opens a regression and never looks at a histogram, then reports a spurious result driven by three outliers or a broken timestamp. Talk through the data integrity checks first, and say your uncertainty on every number you quote.
Expect next
- What would you check before trusting that correlation?
- How many hypotheses did you test, and how does that change your p-value?
- What would you want that is not in this dataset?
Reported by candidates at Jane Street (Credit Risk, London, 2025). Source: Wall Street Oasis.
045You test two hundred signals and three come back significant at the five percent level. What do you conclude?Quant researchQuant trading
Say this
That you have found nothing. Under a pure null you would expect ten false positives from two hundred tests at five percent, so three is fewer than chance. If anything the result is evidence against there being any signal at all.
Then walk it
- Expected false positives are 200 times 0.05 equals 10. Getting three significant results is below what noise alone produces, so the finding is not just unimpressive, it is worse than random.
- The right frame is family-wise error or false discovery rate. Bonferroni sets the threshold at 0.05 over 200, which is 0.00025, brutal but valid. Benjamini-Hochberg controls the expected proportion of false discoveries among the rejections and is much less conservative, which is usually the better choice when you are screening.
- The subtlety with financial signals: they are heavily correlated with each other, so the effective number of independent tests is far below 200. Bonferroni is then too harsh. I would estimate the effective number of tests, for example from the eigenvalue spectrum of the signal correlation matrix, or use a permutation or block-bootstrap null that preserves the correlation structure.
- The right test of whether anything survived is not a p-value at all. It is out-of-sample: hold back a period, or better a different market, and see whether the three signals still work with the sign you predicted.
- And the disclosure discipline, which is the answer a research head wants to hear: I would report the number of specifications tried alongside the result. The deflated Sharpe ratio and Harvey and Liu's work on multiple testing in finance both exist because the profession spent decades not doing this.
Where candidates lose it
Getting excited about the three and building a strategy on them. The whole question is whether you compute the expected number of false positives before you get attached. Say ten out of two hundred immediately, then talk about correlated tests, because that is where the technical depth is.
Expect next
- How would you estimate the effective number of independent tests?
- What is the deflated Sharpe ratio?
- How would you set up the experiment properly from the start?
049You need a covariance matrix for five hundred assets and you have two years of daily data. What is the problem and how do you fix it?Quant researchRisk
Say this
You have 500 assets and roughly 500 observations, so the sample covariance matrix is nearly singular and its smallest eigenvalues are garbage. Any optimiser will load up on exactly those directions, so you have to shrink or impose factor structure.
Then walk it
- Count the parameters: 500 times 501 over 2 is about 125,000 numbers estimated from 250,000 data points. The ratio of assets to observations, roughly one here, is what governs the damage, and the sample eigenvalue spectrum is badly biased even at a ratio of a quarter.
- Marchenko-Pastur describes exactly how the eigenvalues spread out. The largest are overstated and the smallest understated, and the smallest ones are the low-variance directions a mean-variance optimiser will concentrate in. That is why naive optimisers produce absurd leveraged long-short positions.
- Fix one, shrinkage. Ledoit-Wolf shrinks the sample matrix towards a structured target like a constant-correlation matrix, with an optimal intensity derived in closed form. Cheap, well-behaved and hard to beat as a default.
- Fix two, factor structure. Model returns as exposures to a few factors plus idiosyncratic noise, so the covariance is B times F times B transpose plus a diagonal. You have gone from 125,000 parameters to a few thousand. This is what every commercial risk model does.
- Fix three, random matrix filtering: keep the eigenvalues above the Marchenko-Pastur bulk edge as signal and replace the bulk with its average. Then state the practical check, which is out-of-sample portfolio variance rather than any in-sample fit statistic, because in-sample the sample matrix always wins and is always wrong.
Where candidates lose it
Saying you would just use the sample covariance matrix because two years is a lot of data. It is not, relative to 500 assets. The interviewer is testing whether you know that estimation error in the covariance matrix, not in the means, is what breaks portfolio optimisation in practice, and whether you can name shrinkage or factor models as the fix.
Expect next
- Why does the optimiser concentrate in the smallest eigenvalue directions?
- How do you choose the shrinkage intensity?
- How would you test whether your covariance matrix is any good?
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

