Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
038How would you fix violations of the OLS assumptions?AQR Capital ManagementInvestments · Greenwich · 2022
Say this
Depends which assumption breaks, and the fixes fall into two very different classes: violations that only break your standard errors, and violations that break the coefficients themselves. The first class you patch; the second class you have to re-specify the model.
Then walk it
- Heteroskedasticity and autocorrelated errors: coefficients stay unbiased, only inference is wrong. Fix with White or Newey-West robust standard errors, or clustered errors if the dependence is by group. Cheap fix, always worth doing on financial data.
- Endogeneity, meaning a regressor correlated with the error, whether from omitted variables, simultaneity or measurement error: this biases the coefficients and no standard error fix helps. You need an instrument, a control for the omitted factor, a fixed effect, or a different design.
- Multicollinearity: coefficients are still unbiased but the variances explode and the signs flip sample to sample. Drop or combine the collinear regressors, use ridge, or work with principal components. And check the variance inflation factors before you interpret anything.
- Non-normal or fat-tailed errors: inference is still fine asymptotically thanks to the CLT, but outliers dominate the fit because OLS minimises squares. Use robust regression, Huber loss or quantile regression, and always look at the influence diagnostics.
- Non-linearity: add the relevant transform or interaction rather than pretending it away. And I would say the order I actually work in on real data: plot residuals against fitted values and against time first, because most violations announce themselves visually before any test does.
Where candidates lose it
Listing fixes without separating what biases the coefficients from what only biases the standard errors. That distinction is the question. Slapping Newey-West errors on an endogenous regression is a common and useless move, and an interviewer at a research shop will push on exactly that.
Expect next
- Which of those actually biases your coefficients?
- How do you detect endogeneity if you have no instrument?
- What do you do when the residuals are fat-tailed and autocorrelated at the same time?
Reported by candidates at AQR Capital Management (Investments, Greenwich, 2022). Source: Wall Street Oasis.
083Write an algorithm to find all the primes from one to n, and then optimise it.AQR Capital ManagementResearch · Greenwich · 2015
Say this
Sieve of Eratosthenes. Mark every multiple of each prime as composite, and the unmarked survivors are the primes. Time is n log log n, which is essentially linear, and memory is n bits.
Then walk it
- The baseline to reject first: trial division on each number up to its square root is about n times root n over log n, far worse. Say why the sieve wins before you write it.
- The sieve itself: start at p equal to 2, mark 4, 6, 8 and so on, then advance to the next unmarked number. Two optimisations that come free: start marking at p squared rather than 2p, because smaller multiples are already marked, and stop the outer loop at root n.
- Memory optimisations: store only odd numbers, halving memory, use a bit array rather than bytes for an eightfold saving, and if n is large, sieve in cache-sized blocks. That last one matters more than anything else in practice, because a naive sieve over 10 to the 9 is dominated by cache misses, and segmenting it can be several times faster at identical complexity.
- Further refinements if pushed: a wheel sieve skipping multiples of 2, 3 and 5 removes about 77 percent of the candidates, and the sieve of Atkin is asymptotically better at n over log log n but is slower in practice and much harder to get right.
- And the answer to a different question they may be asking: if you want to test whether one large number is prime rather than enumerate a range, the sieve is the wrong tool entirely and you want Miller-Rabin, which is probabilistic and fast. Recognising that enumerate and test are different problems is worth saying.
Where candidates lose it
Giving trial division and calling it done, or giving the sieve with no optimisation when the question explicitly asks for one. The optimisations they want in order are: start at p squared, skip evens, use a bit array, then segment for cache. Naming cache blocking is what marks you out, because it is the one that matters at scale and it is not in the textbook answer.
Expect next
- What is the memory cost for n equal to a billion, and how would you reduce it?
- How would you parallelise the sieve?
- Now test whether one very large number is prime.
Reported by candidates at AQR Capital Management (Research, Greenwich, 2015). Source: Wall Street Oasis.
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

