Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
054What does PCA do, how do you choose the number of components, and what are its limitations on financial data?Quant researchRisk
Say this
It finds the orthogonal directions of maximum variance, which are the eigenvectors of the covariance matrix, and lets you describe the data with fewer numbers. Choose the number of components by explained variance, a scree elbow, or the Marchenko-Pastur bulk edge if you want a principled cutoff.
Then walk it
- Mechanically: eigendecompose the covariance or correlation matrix, or take the SVD of the centred data. Eigenvalues are the variance along each component, eigenvectors are the directions.
- Correlation versus covariance matters. On assets with wildly different volatilities, PCA on the covariance matrix is dominated by the most volatile names, so standardise first unless the scale is meaningful.
- Concrete example everyone in rates knows: PCA on the yield curve gives level, slope and curvature, explaining roughly 90, 8 and 2 percent of variance. On equities the first component is the market, explaining 25 to 40 percent depending on the regime, and it rises sharply in a crisis.
- Choosing k: cumulative explained variance at 90 or 95 percent, the scree elbow, or eigenvalues above the random matrix bulk edge, which is the statistically defensible version because it separates signal from estimation noise.
- Limitations, and these are the answer to the real question. PCA maximises variance, not predictive power, so the components need not have anything to do with your target. It is unstable: eigenvectors rotate sample to sample when eigenvalues are close, so your factor two and factor three swap places. It assumes linearity. And the components are usually uninterpretable outside a structured setting like the yield curve, which makes them awkward to risk-manage.
Where candidates lose it
Describing PCA as dimensionality reduction and stopping. Two things get graded: that it is unsupervised so high-variance directions are not necessarily predictive, and that you must standardise when scales differ. Also have a real example ready, because level-slope-curvature or the equity market factor proves you have used it rather than read about it.
Expect next
- Why is PCA not necessarily good for prediction?
- What does the first principal component of an equity universe represent, and what happens to it in a crisis?
- How is PCA related to a factor risk model?
056What is maximum likelihood estimation, and when would you prefer method of moments?Quant researchRisk
Say this
MLE picks the parameters that make the observed data most probable under your assumed distribution. It is asymptotically efficient if the model is right, which is exactly the condition that makes method of moments attractive when it is not.
Then walk it
- MLE: maximise the log likelihood, which is the sum of log densities. Under regularity conditions it is consistent, asymptotically normal, and attains the Cramer-Rao bound, with variance given by the inverse Fisher information.
- Method of moments: match sample moments to their theoretical expressions and solve. Generalised method of moments extends this to more moment conditions than parameters, weighting them optimally, and it needs no full distributional assumption.
- So the tradeoff is efficiency versus robustness. MLE uses the whole density, so it extracts every bit of information and pays for it with sensitivity to misspecification. GMM uses only the moments you trust.
- Concrete case: fitting a distribution to daily returns. MLE under a normal assumption gives you the sample mean and variance and will be badly misled by the tails. MLE under a Student t estimates the degrees of freedom and is much better behaved. GMM on a few robust moments avoids committing to either.
- Practical points worth raising: MLE can be biased in small samples even when consistent, the classic example being the variance estimator with n rather than n minus 1 in the denominator. And numerically you should always check the Hessian at the optimum, because a flat likelihood means your parameter is not identified, which is common in GARCH and regime models.
Where candidates lose it
Describing MLE as the best estimator without the qualifier if the model is correctly specified. That caveat is the entire content of the comparison. Also be ready for the small-sample bias point, since MLE being biased while still consistent catches people who have only memorised the asymptotic properties.
Expect next
- Give me an example where MLE is biased.
- What is the Cramer-Rao bound?
- How would you fit a Student t to returns, and what does the estimated degrees of freedom tell you?
057What does stationarity mean, how do you test for it, and why do you care?Quant researchRisk
Say this
Weak stationarity means constant mean, constant variance and an autocovariance that depends only on the lag. You care because the standard inference machinery assumes it, and regressing non-stationary series on each other produces spurious relationships with impressive t statistics.
Then walk it
- Prices are not stationary, they are close to a random walk with a unit root. Returns are much closer to stationary, which is why every model works on returns and not on levels.
- Tests: augmented Dickey-Fuller and Phillips-Perron test the null of a unit root, KPSS tests the null of stationarity. Run both, because they have opposite nulls and agreeing tests are more convincing than either alone. And these tests have low power, so failing to reject is weak evidence.
- Spurious regression is the cost of getting it wrong. Regress one independent random walk on another and you reject the null of no relationship far more often than five percent of the time, with an R squared that looks respectable. Granger and Newbold showed this in 1974 and people still do it.
- The exception that matters for trading: cointegration. Two non-stationary series can have a stationary linear combination, which is precisely the statistical statement of a pair trade. Test it with Engle-Granger or Johansen, and then the correct specification is an error-correction model rather than a regression in levels.
- The practical honesty: financial series are not stationary even in returns, because volatility and correlation regimes shift. So I treat stationarity as a working approximation over a limited window, and I check parameter stability across subsamples rather than trusting one test on the full history.
Where candidates lose it
Answering just difference it until the test passes. Over-differencing destroys the signal, and a cointegrated pair loses its whole tradeable relationship if you difference both series. Say what stationarity buys you, name the spurious regression result, and bring up cointegration unprompted since it is where the money is.
Expect next
- What is cointegration and how does it differ from correlation?
- How do you test it, and what is an error-correction model?
- What if a series is stationary in one decade and not the next?
058Daily equity returns are not normal. How are they different, and what do you do about it?Quant researchRisk
Say this
They have fat tails, negative skew and volatility clustering. Daily equity index kurtosis is typically 5 to 10 against 3 for a normal, so moves the normal says should happen once a century happen every few years.
Then walk it
- Put a number on it. A normal assigns a five standard deviation daily move a probability of about one in 3.5 million, roughly once in 14,000 years of trading. The S&P has had several since 1950. The tails are not slightly wrong, they are wrong by orders of magnitude.
- Negative skew: large down moves are bigger and faster than large up moves. That is why index option skew exists and why puts are persistently richer than calls in implied vol terms.
- Volatility clustering means part of the unconditional fat tail is a mixture effect. Returns standardised by a GARCH-type conditional volatility are much closer to normal, though still fat-tailed, which tells you some but not all of the kurtosis is time-varying vol rather than genuinely fat conditional tails.
- What I would do depends on the use. For risk: empirical quantiles, a Student t or a generalised Pareto fit to the tail via extreme value theory, and expected shortfall rather than value at risk, because expected shortfall is sensitive to how bad the tail is. For pricing: a model with jumps or stochastic volatility rather than plain Black-Scholes.
- And the aggregation point: monthly returns are considerably closer to normal than daily returns because of the CLT, so the right distributional assumption depends on your horizon. That is worth saying because it stops the conversation becoming a generic tails are fat sermon.
Where candidates lose it
Saying fat tails and stopping. Quantify it, because the five sigma comparison is what makes the point land. Also do not forget the skew, since symmetric fat tails would not explain the option skew, and be ready to distinguish unconditional fat tails from conditional heteroskedasticity.
Expect next
- How much of the kurtosis is explained by volatility clustering?
- What is expected shortfall and why prefer it to value at risk?
- How does this show up in the option surface?
065What is adverse selection and why is it a market maker's real cost?Prop trading firmsQuant trading
Say this
Adverse selection is the fact that whoever trades with you chose to, and sometimes they chose because they know something you do not. Your quote gets hit disproportionately when it is wrong, so on average the trades you get are worse than the trades you wanted.
Then walk it
- The mechanism: you post a two-sided quote at your fair value. Uninformed flow hits both sides roughly equally and you earn the spread. Informed flow only takes the side that is mispriced, so those trades lose you money immediately.
- The measurement is simple and it is what every market making desk tracks: mark your fills against the mid price a few seconds or minutes later. If your buys are systematically below where the market goes, you are being adversely selected. The industry term is markout.
- This is why the spread must be wide enough that the profit from uninformed flow covers the loss to informed flow. Glosten and Milgrom's model makes the spread purely a function of the probability of informed trading, with zero inventory risk at all.
- It explains observable behaviour. Spreads widen before earnings and economic releases, when the probability of informed flow spikes. Market makers pay for retail order flow precisely because retail flow is less informed, so it is worth more.
- And the extreme version is why quotes get pulled. If adverse selection becomes severe enough that no spread compensates, the correct response is not to widen but to stop quoting. That is what a flash crash looks like from the inside, and saying that shows you understand the business rather than just the term.
Where candidates lose it
Confusing adverse selection with inventory risk. Inventory risk is the price moving while you hold a position you did not want. Adverse selection is getting the position in the first place precisely when it is wrong. Interviewers ask candidates to distinguish them, so have both definitions crisp and know that markout is how you measure it.
Expect next
- How is that different from inventory risk?
- How would you measure it on your own fills?
- Why is retail order flow worth paying for?
066Explain the Kelly criterion, and why do real traders bet less than it says?Quant tradingProp trading firms
Say this
Kelly maximises the expected growth rate of your capital by betting a fraction equal to your edge divided by the odds. For an even-money bet at probability p, that fraction is 2p minus 1. Real traders bet a fraction of it because Kelly assumes you know your edge exactly, and overbetting is far more damaging than underbetting.
Then walk it
- The derivation in one line: maximise the expected log of wealth, because log wealth is additive across repeated bets and its expectation governs the long-run growth rate. For a bet paying b to 1 with win probability p, the optimal fraction is (pb minus (1-p)) over b.
- Numbers: a 55 percent even-money bet gives f equal to 0.1, so ten percent of capital. A 60 percent bet gives 20 percent. That is a lot more than most people's intuition, which is the first surprise of Kelly.
- For continuous returns the analogue is mean over variance, which is why a Sharpe ratio maps directly to a leverage level. Full Kelly leverage equals the Sharpe divided by the volatility.
- The asymmetry is the key insight. Growth rate as a function of bet size is a concave parabola, so betting half Kelly gives you three quarters of the growth with half the volatility. Betting double Kelly gives you zero growth. Overestimating your edge by a factor of two therefore destroys the entire benefit.
- And full Kelly's drawdowns are intolerable in practice: the probability of at some point halving your capital under full Kelly is fifty percent. Nobody running other people's money survives that, and no risk manager permits it. So a quarter to a half Kelly is standard, and the honest reason is parameter uncertainty plus career risk, not mathematics.
Where candidates lose it
Reciting the formula without the asymmetry. The gradeable insight is that the growth curve is flat near the optimum and falls off a cliff past it, which is why uncertainty in your edge estimate pushes you to bet less. Also mention the fifty percent chance of a fifty percent drawdown, because it makes the practical argument concrete.
Expect next
- What is the probability of a fifty percent drawdown under full Kelly?
- How does Kelly relate to mean-variance optimisation?
- How would you size when your edge estimate itself has a standard error?
067What is the difference between a market order and a limit order, and who pays the spread?Prop trading firmsQuant trading
Say this
A market order takes whatever price is available and pays the spread for certainty of execution. A limit order posts a price and waits, earning the spread if it fills, but with no guarantee it fills at all. You are choosing between price risk and execution risk.
Then walk it
- The taker of liquidity pays. Buy with a market order and you pay the offer, which is above mid, so you start down by half the spread. The passive side on the other end of that trade collects it.
- On most exchanges the fee structure reinforces this: makers get a rebate, takers pay a fee. So the maker's economics are spread capture plus rebate minus adverse selection.
- The cost of a limit order is not zero, it is optionality you are giving away. A resting bid is a free put you have written to the market: it fills when the price is falling and does not fill when the price rises. That is adverse selection expressed as execution risk.
- So the choice is horizon-dependent. If I need to be done now because I have information or a hedge to put on, I pay the spread. If I am providing liquidity or my signal has a multi-day horizon, I post and wait.
- Worth adding the practical middle ground, since this is what execution desks actually do: split the order, post passively and cross only when the queue is not filling or when the signal decays. Implementation shortfall against the arrival price is how you measure whether you got that balance right.
Where candidates lose it
Getting the definitions right but not answering who pays the spread. The taker pays. The second miss is treating a limit order as free, when the real cost is the option you have written to anyone with better information. Say that and you are ahead of most candidates.
Expect next
- What is the hidden cost of a resting limit order?
- How would you decide between posting and crossing?
- What is implementation shortfall?
074How would you model market impact and slippage for a strategy you are sizing?Quant researchQuant trading
Say this
Split the cost into spread, temporary impact and permanent impact. The empirical regularity worth knowing is the square-root law: impact scales roughly with the square root of the order size as a fraction of daily volume, times the volatility.
Then walk it
- The square-root law: impact in volatility units is approximately a constant times the square root of order size over average daily volume, with the constant usually estimated around 0.5 to 1. So trading 1 percent of ADV in a 2 percent daily vol name costs roughly 0.1 times 2 percent, about 20 basis points.
- That non-linearity is what caps capacity. Doubling your size only increases impact by 41 percent per share, but total cost grows as size to the power 1.5, so cost eats your edge faster than your edge grows.
- Separate temporary from permanent. Temporary impact reverts after you stop trading and is a function of how fast you trade. Permanent impact is the information your trading revealed, and it does not come back. Almgren-Chriss style frameworks trade off the two against the risk of trading slowly.
- Estimating it honestly: use your own fills against arrival price, not a vendor model, and regress realised shortfall on participation rate, volatility and spread. You need a lot of trades, and you must control for the fact that you traded more aggressively when you had more signal, which biases the estimate.
- And the modelling discipline: be conservative, because impact is the parameter most likely to turn a profitable backtest into a losing strategy. I would rather assume twice the cost and discover I was pessimistic than the reverse. That preference is the answer they want to hear.
Where candidates lose it
Assuming linear impact or using the quoted spread as the whole cost. For any size that matters the spread is the small part. Know the square-root law and know that cost scaling as size to the power 1.5 is what determines capacity, because that is the link between a research result and a business decision.
Expect next
- Why does cost scale as size to the power one and a half?
- How do you separate permanent from temporary impact empirically?
- How does impact determine the capacity of a strategy?
076You have K sorted arrays on disk, too large to load at once. How do you merge them into one sorted output?CitadelEquity Capital Markets · New York · 2026
Say this
K-way merge with a min heap of size K. Push the first element of each array into the heap, repeatedly pop the minimum and write it out, then push the next element from whichever array the minimum came from. Time is N log K, memory is O(K) plus your buffers.
Then walk it
- The heap holds one candidate per array, each entry tagged with which array it came from and the index within it. Pop the smallest, emit it, and refill from that same array.
- Complexity: N total elements, each pushed and popped once, each operation log K. So N log K, which beats concatenate-and-sort at N log N whenever K is much smaller than N.
- The disk part is the real content of the question. You do not read element by element, you read blocks. Keep a buffer per array, say a few megabytes each, refill it when it drains, and write the output through a large buffer too. The heap operations are free compared with I/O, so the design goal is sequential reads and few of them.
- If K is very large, K times the buffer size exceeds memory, and then you merge in passes: merge groups of, say, 100 files at a time, then merge the results. That is exactly how external merge sort works, and total I/O is N times the number of passes.
- Practical notes I would raise: use a tournament tree or a loser tree instead of a binary heap if you want fewer comparisons per element, handle the tie-breaking rule explicitly if stability matters, and if this is a real system, check whether the operating system's readahead is already doing your buffering for you before you build it yourself.
Where candidates lose it
Answering merge them pairwise, which is K times N in the worst case, or ignoring the on-disk part entirely. The interviewer put the data on disk deliberately, so talk about block-sized buffered reads and what happens when K is too large to buffer. State the N log K complexity explicitly.
Expect next
- What if K is a million?
- How large would you make the buffers, and why?
- How would you parallelise it?
Reported by candidates at Citadel (Equity Capital Markets, New York, 2026). Source: Wall Street Oasis.
077Given an array and a window of size k, return the maximum in each window as it slides.Akuna CapitalQuant Development · Chicago · 2025
Say this
Monotonic deque, O(n) total. Keep a deque of indices whose values are strictly decreasing. Before pushing a new index, pop from the back everything smaller than the new value, and pop from the front anything that has fallen out of the window. The front is always the maximum.
Then walk it
- Why the deque is monotonic: if a new element is larger than something behind it, that older smaller element can never be the maximum of any future window, because the new one is both larger and more recent. So it is safe to discard permanently.
- Each index is pushed once and popped once, so the total work is O(n) even though a single step can pop many elements. That amortised argument is the thing to say out loud, because it is what distinguishes this from the naive O(n k).
- Store indices, not values, so you can test whether the front has expired by comparing front index against i minus k plus 1.
- Alternatives and why they are worse: a max heap gives O(n log k) and needs lazy deletion of expired entries. A balanced BST or a multiset gives O(n log k) too. Both are fine and both are beaten by the deque.
- Where this actually matters on a trading system, which is worth mentioning: rolling extremes over a tick window, running high and low for a breakout signal, and rolling maximum drawdown. The same structure with the comparison reversed gives you the rolling minimum, and the O(1) amortised cost per tick is what makes it usable in a hot path.
Where candidates lose it
Reaching for a heap and stopping there. The heap answer is acceptable but it is not the answer to this question, and the interviewer is specifically looking for the monotonic deque and the amortised O(n) argument. Also remember to expire the front by index, which is the bug that shows up most often in live coding.
Expect next
- Prove the amortised complexity.
- Now give me the rolling median instead.
- How would you handle a window defined by time rather than by count?
Reported by candidates at Akuna Capital (Quant Development, Chicago, 2025). Source: Wall Street Oasis.
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

