Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
036The sample variance with the n minus one correction is unbiased. Is its square root an unbiased estimator of the standard deviation?Squarepoint CapitalQuantitative Research · London · 2026
Say this
No. The square root is concave, so by Jensen's inequality the expected square root is strictly less than the square root of the expected value. The sample standard deviation is biased downwards, always, for any distribution with positive variance.
Then walk it
- Jensen: for a strictly concave g, E of g(X) is less than g of E of X unless X is degenerate. With g the square root and X the unbiased sample variance, E of s is less than sigma.
- Size the bias for normal data. E of s equals c4(n) times sigma, where c4 is a known constant involving gamma functions. At n equal to 2, c4 is about 0.798, so you understate sigma by 20 percent. At n equal to 10 it is 0.9727, a 2.7 percent understatement. At n equal to 30 it is 0.9914.
- So the bias is order 1/(4n) and it vanishes as n grows. It is a real problem for short samples and irrelevant for long ones.
- Unbiasedness is also not preserved under any nonlinear transform, which is the general lesson. The unbiased estimator of sigma squared does not give you an unbiased estimator of sigma, or of 1/sigma, or of log sigma.
- Where this bites on a desk: annualised volatility estimated from a few weeks of data, and any Sharpe ratio, since the Sharpe divides by s. Understating s inflates the Sharpe, so short-sample Sharpes are biased upwards. That is worth saying because it connects a textbook Jensen question to a live problem in strategy evaluation.
Where candidates lose it
Saying yes because the variance estimator is unbiased. Unbiasedness does not survive a nonlinear function. Name Jensen explicitly, give the direction of the bias, and quantify it with c4 for at least one small n. The follow-up about Sharpe ratios is where the real conversation is, so get there yourself.
Expect next
- How would you correct it?
- What does that imply for a Sharpe ratio estimated on a short sample?
- Is the sample correlation coefficient unbiased?
Reported by candidates at Squarepoint Capital (Quantitative Research, London, 2026). Source: Wall Street Oasis.
037You stand on a road and watch cars drive past. How would you estimate the parameter of the underlying distribution?Jump TradingQuantitative Research · Chicago · 2018
Say this
First I would state the model: arrivals as a Poisson process with rate lambda, so inter-arrival times are exponential with mean 1 over lambda. Then the maximum likelihood estimate of lambda is just the count divided by the observation time, and its standard error is lambda over the square root of the count.
Then walk it
- Model choice first, and justify it: independent arrivals at a constant rate with no memory gives a Poisson process. That is reasonable on a quiet road, and clearly wrong near a traffic light where cars arrive in platoons.
- MLE: for n arrivals in time T, lambda hat is n over T. It is unbiased, and the variance is lambda over T, so the relative standard error is 1 over the square root of n. Twenty-five cars gives you a 20 percent standard error, a hundred cars gives 10 percent.
- That tells you the sample size you need before you open your mouth about precision. If someone wants the rate to five percent, you need 400 cars.
- Now the diagnostics, which are what a research interview is actually about. Plot the inter-arrival times and check whether they look exponential. Over-dispersion, meaning variance above the mean of the counts, tells you arrivals are clustered and Poisson is wrong. Then I would go to a Cox process or a Hawkes process with self-excitation.
- And I would flag the estimation trap: if instead I sampled by picking a random moment and measuring the gap I happened to land in, I would oversample long gaps. That is the inspection paradox, and it biases the mean gap upward by a factor of one plus the squared coefficient of variation. It is the same bias that makes waiting times feel longer than the timetable says.
Where candidates lose it
Jumping to a formula without stating the model or checking it. The interviewer wants model, estimator, standard error, then diagnostics. The specific failure mode they are hunting is the inspection paradox, so mention length-biased sampling unprompted. Hawkes processes are the right answer for clustered arrivals and they are also how trade arrivals actually behave in markets.
Expect next
- How would you test whether the Poisson assumption holds?
- What if the cars arrive in clusters?
- How long do you need to watch to get the rate within five percent?
Reported by candidates at Jump Trading (Quantitative Research, Chicago, 2018). Source: Wall Street Oasis.
040Here is a dataset. Analyse it using probability metrics and tell me what you find.Jane StreetCredit Risk · London · 2025
Say this
I would spend the first third of the time on the data itself before any modelling: shape, missingness, duplicates, timestamps, and the univariate distributions. Then state a hypothesis, test it, and report the effect size with an honest uncertainty. Narrate every step, because the interviewer is grading the process, not the punchline.
Then walk it
- Start with the boring checks, out loud. Row count, date range, obvious duplicates, missing values and whether they are missing at random, and whether any column is a leak of the outcome. Most real findings in interviews of this kind are data artefacts.
- Then univariates: mean, median, standard deviation, skew, kurtosis, and the tails. Plot histograms and the empirical CDF. If a column is heavy-tailed or bimodal, say so, because it changes every subsequent choice.
- Then the relationship you were asked about. Give a point estimate plus a confidence interval, and prefer a plot to a coefficient. If the data are time-ordered, check for autocorrelation and regime change before quoting any p-value, because serial dependence inflates significance badly.
- Then the discipline: state your null, say what result would change your mind, and count how many hypotheses you have looked at. If you tested twenty things, say so and adjust.
- Close with what the data cannot tell you. A credit dataset with survivors only cannot tell you about defaults. Ending on the limitation is what separates an analyst from someone producing numbers, and in a live exercise it is the cheapest way to sound senior.
Where candidates lose it
Going straight to a model. Almost every candidate opens a regression and never looks at a histogram, then reports a spurious result driven by three outliers or a broken timestamp. Talk through the data integrity checks first, and say your uncertainty on every number you quote.
Expect next
- What would you check before trusting that correlation?
- How many hypotheses did you test, and how does that change your p-value?
- What would you want that is not in this dataset?
Reported by candidates at Jane Street (Credit Risk, London, 2025). Source: Wall Street Oasis.
041Explain the structure of a probabilistic graphical model you have worked with.Tower Research CapitalQuantitative Research · New York · 2015
Say this
Pick one model you actually built and describe it in four parts: the variables, the graph and what the missing edges assert, how you did inference, and how you checked it. The missing edges are the interesting part, because a graphical model is a set of conditional independence claims.
Then walk it
- Name the class first. A directed model, a Bayes net, factorises the joint as a product of each node given its parents and encodes causal or generative structure. An undirected model, a Markov random field, factorises into potentials over cliques and is better when the interactions have no natural direction.
- Then say what the graph buys you. Without structure, a joint over n binary variables needs 2 to the n minus 1 parameters. With a sparse graph it needs a handful per node. That reduction is the whole point, and the missing edges are the assumptions you are making.
- Inference: exact by belief propagation or the junction tree if the graph is a tree or has small treewidth, otherwise approximate by variational methods, loopy BP or MCMC. Say which you used and why, and say what the cost was.
- A concrete example is worth more than the taxonomy. A hidden Markov model is the simplest useful case: a latent state that evolves as a Markov chain with observations conditionally independent given the state. In markets people use it as a regime model, with the latent state as calm or stressed, fitted by Baum-Welch, and decoded with Viterbi.
- Then the honest part: on financial data the latent states are unstable, the number of regimes is not identified, and the fitted model will happily tell you the regime changed last week when it changed two months ago. So I used it as a descriptive overlay, never as a standalone signal.
Where candidates lose it
Reciting textbook definitions of Bayes nets and MRFs without ever describing a model you built. This question is a depth probe, and the interviewer will go three levels down on whichever model you name, so name the one you know cold. Be able to state the conditional independence your graph asserts and how you validated it.
Expect next
- What conditional independences does your graph assert, and did you test them?
- How did you do inference, and what was the complexity?
- How would you learn the graph structure from data?
Reported by candidates at Tower Research Capital (Quantitative Research, New York, 2015). Source: Wall Street Oasis.
042Derive the update rules for alternating least squares in a matrix factorisation.Tower Research CapitalQuantitative Research · New York · 2015
Say this
Fix one factor and the objective becomes an ordinary ridge regression in the other, so each update is a closed-form normal equation. With R approximated by U times V transpose and an L2 penalty, the update for a row of U is (V'V plus lambda I) inverse V'r.
Then walk it
- Objective: minimise the sum over observed entries of (r_ij minus u_i dot v_j) squared plus lambda times the sum of the squared norms of u and v. It is non-convex jointly in U and V, but convex in each one separately. That is the entire reason alternating minimisation works here.
- Differentiate with respect to u_i holding V fixed. The gradient is minus 2 times the sum over observed j of (r_ij minus u_i dot v_j) v_j plus 2 lambda u_i. Set it to zero.
- Rearranged: (sum over observed j of v_j v_j' plus lambda I) u_i equals the sum over observed j of r_ij v_j. So u_i equals that Gram matrix inverse times the weighted sum. Symmetric for v_j with U fixed.
- Cost per update is k cubed for the k by k solve plus k squared per observed entry, and it parallelises perfectly by row, which is exactly why ALS beat SGD for large recommender systems.
- Say the limitations. It converges to a local optimum only, so initialisation matters, usually small random or SVD-based. The lambda is essential because otherwise the Gram matrix is singular for users with fewer than k observations. And it monotonically decreases the objective every half-step, so if your loss ever goes up you have a bug in the derivation, which is a useful debugging fact.
Where candidates lose it
Writing down the gradient-descent update instead of the closed-form solve. ALS is defined by exploiting the per-block convexity to solve exactly, not by stepping. Also do not forget the lambda I, since without it the system is singular for sparse rows, and do not sum over all j when only observed entries enter the loss.
Expect next
- Why does ALS converge, and to what?
- When would you prefer SGD over ALS?
- How would you handle implicit feedback where you only see the ones?
Reported by candidates at Tower Research Capital (Quantitative Research, New York, 2015). Source: Wall Street Oasis.
045You test two hundred signals and three come back significant at the five percent level. What do you conclude?Quant researchQuant trading
Say this
That you have found nothing. Under a pure null you would expect ten false positives from two hundred tests at five percent, so three is fewer than chance. If anything the result is evidence against there being any signal at all.
Then walk it
- Expected false positives are 200 times 0.05 equals 10. Getting three significant results is below what noise alone produces, so the finding is not just unimpressive, it is worse than random.
- The right frame is family-wise error or false discovery rate. Bonferroni sets the threshold at 0.05 over 200, which is 0.00025, brutal but valid. Benjamini-Hochberg controls the expected proportion of false discoveries among the rejections and is much less conservative, which is usually the better choice when you are screening.
- The subtlety with financial signals: they are heavily correlated with each other, so the effective number of independent tests is far below 200. Bonferroni is then too harsh. I would estimate the effective number of tests, for example from the eigenvalue spectrum of the signal correlation matrix, or use a permutation or block-bootstrap null that preserves the correlation structure.
- The right test of whether anything survived is not a p-value at all. It is out-of-sample: hold back a period, or better a different market, and see whether the three signals still work with the sign you predicted.
- And the disclosure discipline, which is the answer a research head wants to hear: I would report the number of specifications tried alongside the result. The deflated Sharpe ratio and Harvey and Liu's work on multiple testing in finance both exist because the profession spent decades not doing this.
Where candidates lose it
Getting excited about the three and building a strategy on them. The whole question is whether you compute the expected number of false positives before you get attached. Say ten out of two hundred immediately, then talk about correlated tests, because that is where the technical depth is.
Expect next
- How would you estimate the effective number of independent tests?
- What is the deflated Sharpe ratio?
- How would you set up the experiment properly from the start?
049You need a covariance matrix for five hundred assets and you have two years of daily data. What is the problem and how do you fix it?Quant researchRisk
Say this
You have 500 assets and roughly 500 observations, so the sample covariance matrix is nearly singular and its smallest eigenvalues are garbage. Any optimiser will load up on exactly those directions, so you have to shrink or impose factor structure.
Then walk it
- Count the parameters: 500 times 501 over 2 is about 125,000 numbers estimated from 250,000 data points. The ratio of assets to observations, roughly one here, is what governs the damage, and the sample eigenvalue spectrum is badly biased even at a ratio of a quarter.
- Marchenko-Pastur describes exactly how the eigenvalues spread out. The largest are overstated and the smallest understated, and the smallest ones are the low-variance directions a mean-variance optimiser will concentrate in. That is why naive optimisers produce absurd leveraged long-short positions.
- Fix one, shrinkage. Ledoit-Wolf shrinks the sample matrix towards a structured target like a constant-correlation matrix, with an optimal intensity derived in closed form. Cheap, well-behaved and hard to beat as a default.
- Fix two, factor structure. Model returns as exposures to a few factors plus idiosyncratic noise, so the covariance is B times F times B transpose plus a diagonal. You have gone from 125,000 parameters to a few thousand. This is what every commercial risk model does.
- Fix three, random matrix filtering: keep the eigenvalues above the Marchenko-Pastur bulk edge as signal and replace the bulk with its average. Then state the practical check, which is out-of-sample portfolio variance rather than any in-sample fit statistic, because in-sample the sample matrix always wins and is always wrong.
Where candidates lose it
Saying you would just use the sample covariance matrix because two years is a lot of data. It is not, relative to 500 assets. The interviewer is testing whether you know that estimation error in the covariance matrix, not in the means, is what breaks portfolio optimisation in practice, and whether you can name shrinkage or factor models as the fix.
Expect next
- Why does the optimiser concentrate in the smallest eigenvalue directions?
- How do you choose the shrinkage intensity?
- How would you test whether your covariance matrix is any good?
053How would you cross-validate a model on time series data, and why is standard k-fold wrong?Quant researchQuant trading
Say this
Standard k-fold trains on data that comes after your test set, which leaks the future. You need a forward-walking scheme: train on a window, test on the next block, roll forward, and put a gap between train and test so overlapping labels do not bleed across the boundary.
Then walk it
- Two distinct leaks. First, random folds put future observations in the training set, so the model learns things it could not have known. Second, features and labels are usually built from overlapping windows, so even adjacent-in-time observations share information across a fold boundary.
- The fix for the first is walk-forward or expanding-window validation: fit on 1 to t, test on t plus 1 to t plus h, roll. Expanding window mimics how you would actually retrain in production. A fixed rolling window is better if the process is non-stationary.
- The fix for the second is purging and embargoing, from Lopez de Prado. Remove training observations whose label window overlaps the test period, and embargo a short period immediately after the test block. On a 20-day forward return label you need at least a 20-day purge.
- Also beware the hidden leaks that sit outside the folds entirely: fitting a scaler, doing feature selection, or choosing hyperparameters on the full dataset before splitting. Every preprocessing step has to sit inside the fold.
- What I would actually report, and this is the part that matters: one final untouched hold-out period tested once, plus how many configurations I tried before I got there. Walk-forward validation run a hundred times is itself an overfitting device, and the number of trials is the honest measure of how much to discount the result.
Where candidates lose it
Saying you would use k-fold with shuffle turned off and stopping there. That fixes the ordering but not the overlapping-label leak, and interviewers at systematic shops probe exactly that. Mention purging and embargo, and mention that scalers and feature selection must live inside the fold.
Expect next
- How long should the embargo be?
- Expanding window or fixed rolling window, and why?
- How do you account for the number of configurations you tried?
059A strategy shows a Sharpe ratio of 2 over one year. How much do you believe it?Quant researchQuant trading
Say this
Not much. The standard error of an annualised Sharpe estimated over T years is roughly the square root of (1 plus half the Sharpe squared) divided by T, so with one year and a Sharpe of 2 the standard error is about 1.7. The 95 percent interval runs from roughly minus 1.4 to 5.4, which comfortably includes zero.
Then walk it
- The formula, for iid normal returns: standard error of the Sharpe estimate is root of ((1 plus SR squared over 2) divided by T), with T in years for an annualised Sharpe.
- With T equal to 1 and SR equal to 2, that is the square root of (1 plus 2) over 1, which is the square root of 3, about 1.73. Two standard errors either side of the point estimate spans minus 1.4 to 5.4, so one year of data cannot even establish that the strategy makes money.
- Turn it around into the useful statement: to establish statistical significance at two standard errors you need roughly T of at least 4 over SR squared years. A Sharpe of 2 needs about a year to be marginally significant, a Sharpe of 1 needs four years, and a Sharpe of 0.5 needs sixteen years. Most equity factors fall in that last bucket, which is why the factor literature is so contested.
- The estimation error is only half the problem. The other half is selection. If this strategy is the best of a hundred I tested, the honest benchmark is the expected maximum Sharpe under the null, which for a hundred trials is around 2.5 standard errors above zero. The deflated Sharpe ratio adjusts for exactly this.
- And the formula assumes iid normal returns. Autocorrelated returns, which is common in anything holding illiquid or smoothed positions, inflate the Sharpe substantially, and negative skew means the Sharpe misses the risk that actually matters. So I would also want the drawdown profile, the turnover, and the capacity before I believed anything.
Where candidates lose it
Treating a one-year Sharpe as a fact. This question separates people who have evaluated real strategies from people who have read about them. Give the standard error formula, invert it into how many years you need, and then raise selection bias yourself.
Expect next
- How many years would you need for a Sharpe of 0.5 to be significant?
- What if the returns are autocorrelated?
- What else would you want to see besides the Sharpe?
060You backtested a strategy and it performed brilliantly, but in live trading you keep losing money. What would you do?Jump TradingQuantitative Research · Chicago · 2018
Say this
First I would cut the size, because the priority is to stop bleeding while I diagnose. Then I would work through the causes in order of likelihood: costs and slippage, look-ahead or survivorship bias in the backtest, overfitting from too many trials, and only last the possibility that the edge was real and has decayed.
Then walk it
- Costs first, because it is the most common and the easiest to check. Compare realised fill prices against the prices the backtest assumed. If the backtest filled at mid and you are paying the spread plus impact, a strategy with a one basis point edge and a two basis point cost is a losing strategy that looked like a winner. Reconstruct the P&L attribution trade by trade against the simulated trades.
- Then look-ahead bias. Did any feature use data timestamped after the decision, including restated fundamentals, index membership known only later, or a corporate action applied on the announcement date rather than the effective date? Survivorship bias in the universe is the same family of error.
- Then overfitting. How many variants did I try before this one? If the answer is hundreds, the in-sample Sharpe is a maximum over many draws, and the deflated Sharpe is the honest number. Test on a market or a period I never touched.
- Then regime and decay. Plot the backtest P&L by year and see whether the edge was concentrated in one period. Check whether the alpha has been crowded out, which usually shows up as the signal still predicting but the entry price already moved.
- And the meta-answer, which is the one they want: I would write the diagnosis as a hypothesis with a test, not a list of possibilities. For example, if costs are the cause, the loss should scale with turnover, so I would compare the live P&L of the highest and lowest turnover sleeves. Then I would say what would make me shut it off permanently, and I would set that threshold before I looked at any more data.
Where candidates lose it
Jumping straight to the market regime changed. That is the excuse every losing strategy gets and it is almost never the first cause. The ordered list of costs, bias, overfitting, then decay is what a research head wants to hear, along with the instinct to reduce size before you finish diagnosing.
Expect next
- How exactly would you test whether costs are the cause?
- How many strategy variants did you try, and how should that change your prior?
- At what point do you shut it off for good?
Reported by candidates at Jump Trading (Quantitative Research, Chicago, 2018). Source: Wall Street Oasis.
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

