Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
008What is the expected value of a roll of a fair six-sided die?Old Mission CapitalTrading · Chicago · 2020
Say this
Three and a half. Each face has probability one sixth, and one through six sums to 21, so 21 over 6 is 3.5. Or faster: for a uniform run of integers the mean is the midpoint, which is (1 plus 6) over 2.
Then walk it
- The sum one to n is n(n+1)/2, so the mean of a uniform die with n faces is (n+1)/2. For six faces that is 3.5.
- Variance is worth having ready too, because it is the immediate follow-up. E of X squared is (1+4+9+16+25+36)/6, which is 91/6, about 15.17. Subtract 3.5 squared, 12.25, and you get 35/12, about 2.917. Standard deviation is about 1.71.
- Also useful to know the general formula: the variance of a uniform die with n faces is (n squared minus 1) over 12. For n equal to 6 that is 35/12, matching.
- Two dice: expected sum is 7 by linearity, and variance is 35/6 because they are independent, so standard deviation is about 2.42.
- The reason a trading desk opens with this is speed and composure, not difficulty. Answer in under two seconds and have the variance ready before they ask, because the real question is the next one.
Where candidates lose it
Hesitating. This is a warm-up and any pause reads badly. The second trap is being caught flat-footed on variance, which follows more than half the time. Know 35/12 and know that the standard deviation of a die is about 1.71.
Expect next
- What is the variance?
- What about the sum of two dice?
- What is the expected value of the maximum of two dice?
Reported by candidates at Old Mission Capital (Trading, Chicago, 2020). Source: Wall Street Oasis.
011A random variable is uniform on the interval zero to ten. What are its expected value and variance?Old Mission CapitalFinance · New York · 2018
Say this
Mean 5, variance 100 over 12, which is 8.33, so standard deviation about 2.89. For a uniform on a to b the mean is the midpoint and the variance is (b minus a) squared over 12.
Then walk it
- Mean by symmetry: the midpoint of 0 and 10 is 5. No integration needed.
- Variance from the formula (b-a) squared over 12: 100 over 12 equals 8.33, standard deviation 2.887.
- If you want to derive it, E of X squared is the integral of x squared over 10 from 0 to 10, which is 1000/30 equals 33.33. Subtract 25 and you get 8.33. Good to be able to do it either way.
- The 1/12 is worth carrying in your head because it recurs: a fair n-sided die has variance (n squared minus 1)/12, and the rounding error of a value rounded to the nearest tick has variance tick squared over 12. That last one comes up in real microstructure work.
- Practical note: the uniform has thin support and no tails, so it is a bad default for anything financial. The moment somebody hands you a uniform in a trading context, ask what it is meant to represent.
Where candidates lose it
Reaching for integration under time pressure and fumbling the arithmetic. Know the (b-a) squared over 12 form cold. Also do not quote variance when they asked for standard deviation or the other way round, and say which one you are giving.
Expect next
- What is the expected value of the maximum of two independent draws?
- What is the distribution of the sum of two independent uniforms?
- What is the variance of the rounding error when you round to the nearest penny?
Reported by candidates at Old Mission Capital (Finance, New York, 2018). Source: Wall Street Oasis.
020How many zeros are at the end of a thousand factorial?Jump TradingTrading · Chicago · 2013
Say this
249. A trailing zero needs a factor of ten, which needs a two and a five, and fives are scarcer than twos, so just count the fives: 200 plus 40 plus 8 plus 1 equals 249.
Then walk it
- Trailing zeros equal the number of times 10 divides the number, which is the minimum of the exponent of 2 and the exponent of 5 in the prime factorisation. In a factorial, 5 always binds.
- Legendre's formula: sum of floor(1000 divided by 5 to the k). That is floor(1000/5) equals 200, floor(1000/25) equals 40, floor(1000/125) equals 8, floor(1000/625) equals 1, and floor(1000/3125) equals 0.
- 200 plus 40 plus 8 plus 1 gives 249.
- Why the higher powers: 25 contributes two fives, not one, so it must be counted again. Missing that is the single most common error and it costs you 49.
- Quick sanity check on the order of magnitude: roughly 1000/4 is 250, because each multiple of five contributes one and a bit. 249 sits right where it should.
Where candidates lose it
Answering 200 by counting only the multiples of five. Multiples of 25, 125 and 625 carry extra factors of five and each must be counted again. Say out loud why five binds rather than two, because that is the part of the reasoning being graded.
Expect next
- How many zeros in 100 factorial?
- How many digits does 1000 factorial have?
- What is the last non-zero digit of 100 factorial?
Reported by candidates at Jump Trading (Trading, Chicago, 2013). Source: Wall Street Oasis.
021If n(n+1)/2 is the sum of the integers from one to n, what is the formula for the sum of the squares, and can you derive it?Squarepoint CapitalTrading · London · 2025
Say this
n(n+1)(2n+1)/6. The fastest derivation is telescoping: expand (k+1) cubed minus k cubed as 3k squared plus 3k plus 1, sum both sides from 1 to n, and solve for the sum of squares.
Then walk it
- Left side telescopes to (n+1) cubed minus 1.
- Right side is 3S2 plus 3S1 plus n, where S2 is what you want and S1 is the known n(n+1)/2.
- So 3S2 equals (n+1) cubed minus 1 minus 3n(n+1)/2 minus n. Grind it out and you get S2 equals n(n+1)(2n+1)/6.
- Check at n equal to 3: 1 plus 4 plus 9 equals 14, and 3 times 4 times 7 over 6 equals 14. Always test a small case out loud, it costs three seconds and catches sign errors.
- The reason a desk asks this: it is the derivation that matters, not the formula. The same telescoping trick gives the sum of cubes, which is n squared (n+1) squared over 4, and it is the discrete analogue of integration by parts. And the practical use is immediate, because the variance of a uniform die and the variance of a linear time trend in a regression both fall straight out of this sum.
Where candidates lose it
Reciting the formula with no derivation. The interviewer already knows the formula, so the answer is worth nothing on its own. Show the telescoping and verify on n equals 3. Fumbling the algebra after setting it up correctly is forgivable; having no method is not.
Expect next
- Now the sum of cubes.
- Use it to get the variance of a fair n-sided die.
- What is the sum of 1 over k squared as n goes to infinity?
Reported by candidates at Squarepoint Capital (Trading, London, 2025). Source: Wall Street Oasis.
023Monty Hall. Three doors, one car, you pick one, I open a door with a goat, do you switch?Prop trading firmsQuant trading
Say this
Switch. Your original door wins one third of the time, so the other door wins two thirds. The host's choice is not random, and that is where the information comes from.
Then walk it
- Condition on your first pick. One third of the time you picked the car, and switching loses. Two thirds of the time you picked a goat, the host is forced to reveal the only other goat, and switching wins.
- So switching wins two thirds. The Bayes calculation agrees: the likelihood of the host opening door 3 is 1/2 if the car is behind your door 1, and 1 if the car is behind door 2, which is what tilts the posterior two to one.
- The intuition people find convincing: extend it to a hundred doors. You pick one, the host opens 98 goats, and switching wins 99 times out of 100. The host did all the work of avoiding the car.
- The critical assumption, and this is what a quant interview is really checking: the host knows where the car is and always opens a goat. If the host opens a door at random and happens to show a goat, the posterior is fifty-fifty and switching gains nothing.
- So the honest answer is: switch, and the reason it works is that the host's constraint leaks information. Change the host's rule and the answer changes.
Where candidates lose it
Getting the right answer for the wrong reason, or failing to state the host's rule. Everyone knows the answer is switch, so the only thing being graded is whether you can name the assumption that makes it true. Say explicitly that the host knows and is forced to reveal a goat.
Expect next
- What if the host does not know where the car is?
- What if the host only offers the switch when you picked the car?
- Do it with a hundred doors.
032How many people do you need in a room for a better than even chance that two share a birthday, and why is the answer so small?Prop trading firmsQuant trading
Say this
Twenty-three. The reason it feels small is that you are counting pairs, not people. Twenty-three people generate 253 pairs, and each pair matches with probability 1/365, so you expect about 0.69 matches.
Then walk it
- Compute the complement: the probability all birthdays differ is 365/365 times 364/365 times down to 343/365. At 23 people that product is about 0.493, so the match probability is about 0.507.
- The back-of-envelope version: the probability of no match is approximately exp of minus n(n-1)/(2 times 365). Set that to 0.5, so n squared over 730 equals ln 2, giving n about 22.5. Round up to 23.
- The pair-counting intuition is the answer to why. n choose 2 grows quadratically, so the number of chances grows fast while your intuition tracks n linearly.
- Contrast with the question people confuse it with: for someone to share your specific birthday you need about 253 people, because now you have only n pairs, not n squared over 2.
- Where this bites in real work: hash collisions and the birthday attack follow the same square-root law, and so does the chance that two of your supposedly independent signals are accidentally the same trade. You need about the square root of the space to get a collision, which is far fewer than people expect.
Where candidates lose it
Confusing it with the probability that someone shares your birthday, which needs 253 people. Also do not just recite 23. The gradeable part is the pair-counting argument and the exp of minus n squared over 730 approximation, which lets you answer variants like how many for a 99 percent chance without a calculator.
Expect next
- How many for a 99 percent chance?
- How many to share a birthday with you specifically?
- What is the connection to hash collisions?
034How many Starbucks are there in New York City?Tower Research CapitalProp Trading · New York · 2019
Say this
I would say roughly 250 to 350, and I would build it from demand rather than from geography. Eight million people, maybe one in ten buys a Starbucks on a given day, a store serves around a thousand cups a day, so 800,000 over 1,000 is about 800 store-days of demand, which I would then cut for the fact that Manhattan stores are much busier than a thousand cups.
Then walk it
- Build two independent estimates and reconcile them. That is the actual skill being tested, not the number.
- Demand side: 8 million residents plus commuters and tourists, call it 9 million daytime people. Ten percent buy coffee from Starbucks on a given day gives 900,000 cups. A busy Manhattan store does 1,500 to 3,000 cups a day, so 900,000 over 2,500 is about 360 stores.
- Supply side: Manhattan has roughly 200 avenue-blocks of dense commercial frontage and you see a Starbucks every few blocks in midtown, which suggests 150 to 200 in Manhattan alone, plus maybe the same again across the four outer boroughs. That lands around 300.
- Both routes land in the same band, 250 to 400, which is the useful output. I would quote 300 as my point estimate with a range.
- Then state your uncertainty honestly and where it sits: the biggest lever is cups per store, which I could be wrong on by a factor of two. The population number I am confident in to ten percent. Naming which assumption dominates the error is what separates an estimate from a guess.
Where candidates lose it
Producing one chain of assumptions and asserting the answer with false precision. Build two independent routes, reconcile them, give a range, and say which assumption carries the error. Also do not freeze because you do not know the answer. Nobody knows it, and the interviewer is grading the structure and your composure, not the number.
Expect next
- Now make me a market on it and I will trade you.
- How many coffee shops in total?
- How would you check your estimate if you had the internet for thirty seconds?
Reported by candidates at Tower Research Capital (Prop Trading, New York, 2019). Source: Wall Street Oasis.
035How many golf balls fit in the Empire State Building?Tower Research CapitalAssistant Trader · New York · 2013
Say this
Order of a hundred billion. The building is roughly a hundred million cubic feet, a golf ball plus its packing waste takes about 0.0015 cubic feet, so 100 million over 0.0015 is about 70 billion. I would quote 50 to 100 billion.
Then walk it
- Volume of the building: footprint about 200 by 400 feet, so 80,000 square feet, times 1,250 feet of height. That is 100 million cubic feet. Taper the tower and subtract structure and you might call it 80 million usable.
- Volume of a golf ball: diameter 1.68 inches, so radius 0.84 inches. Four thirds pi r cubed is about 2.5 cubic inches. There are 1,728 cubic inches in a cubic foot, so a ball is 0.00145 cubic feet.
- Packing efficiency: random close packing of spheres is about 64 percent, so effective volume per ball is 0.00145 over 0.64, about 0.00226 cubic feet.
- 80 million divided by 0.00226 gives about 35 billion. Using the full 100 million cubic feet gives 44 billion. So my range is tens of billions, call it 40 billion, and I would say 20 to 100 billion to be honest about the error bars.
- Say the two things you are least sure about: the usable fraction of the volume, and whether the question means the empty shell or the building with floors, furniture and lift shafts. Those swing the answer by a factor of two, and the packing fraction only matters at the 30 percent level.
Where candidates lose it
Forgetting the 1,728 cubic inches per cubic foot conversion, which throws you off by three orders of magnitude, or ignoring packing efficiency entirely. Also decide out loud whether you are filling the empty shell or the furnished building. And always sanity check the magnitude: if your answer is in millions or trillions, something went wrong by a factor of a thousand.
Expect next
- What is the packing efficiency of spheres and why?
- How much would they weigh?
- Now estimate the market value of that many golf balls.
Reported by candidates at Tower Research Capital (Assistant Trader, New York, 2013). Source: Wall Street Oasis.
043State the central limit theorem and tell me where it fails.Quant researchQuant trading
Say this
For independent identically distributed variables with finite mean and finite variance, the standardised sample mean converges in distribution to a standard normal. The key conditions are finite variance and enough independence, and both fail regularly in markets.
Then walk it
- Precisely: root n times (X bar minus mu) over sigma converges in distribution to N(0,1). Note it is the standardised mean that converges, and the rate is 1 over root n.
- Failure one, infinite variance. A Cauchy distribution has no variance and the sample mean of Cauchys is Cauchy again, no matter how large n is. Averaging buys you nothing. More generally, stable distributions with tail index alpha below 2 converge to a stable law, not a normal.
- Failure two, dependence. With strongly autocorrelated data the effective sample size is far below n, so you converge much more slowly and your standard errors are too small. Long-range dependence can break it entirely.
- Failure three, the rate in the tails. Even where the CLT holds, convergence is fastest in the middle and slowest in the tails, which is precisely where a risk manager needs accuracy. Berry-Esseen gives an error bound of order 1 over root n times the third absolute moment, so skewed data converges slowly.
- The practical version: daily equity returns have kurtosis of 5 to 10 and volatility clustering, so ten-day sums are much closer to normal than daily returns, but a 99.9 percent quantile computed from a normal assumption will still understate the tail badly. That is why value at risk models use empirical or extreme-value tails rather than leaning on the CLT.
Where candidates lose it
Stating the theorem without the finite variance condition, or claiming everything becomes normal for large n. Also do not confuse it with the law of large numbers, which is about convergence of the mean to a constant and needs only finite mean. Be ready to say what happens with infinite variance, because that is the follow-up.
Expect next
- What happens with a Cauchy distribution?
- How is that different from the law of large numbers?
- How large does n have to be in practice for returns data?
044What is a p-value, and what is it not?Quant researchRisk
Say this
It is the probability of seeing data at least as extreme as what you saw, assuming the null hypothesis is true. It is not the probability that the null is true, and it is not the probability you are wrong.
Then walk it
- The conditioning runs the wrong way from what people assume. A p-value is P(data given null), and what you actually want is P(null given data). Those are different objects and Bayes tells you the second depends on your prior.
- Concretely: if you test a thousand strategies of which fifty genuinely work, at a five percent significance level you get roughly 47 true discoveries and 47 false ones. A p-value of 0.05 in that setting means a coin flip on whether the finding is real.
- It also says nothing about effect size. With a million observations a completely useless one-basis-point edge will have a p-value of 0.0001. Significance is not importance, and in high-frequency data everything is significant.
- And it is only valid for a pre-specified test. Choosing the test after looking at the data, or stopping data collection when the p-value crosses 0.05, invalidates it completely.
- What I would report instead on a desk: the effect size with a confidence interval, out-of-sample performance, and how many specifications I tried. A p-value on its own is close to useless in a research process where hundreds of hypotheses get screened.
Where candidates lose it
Defining it as the probability the null is true. That is the single most common statistical error in finance interviews and it is disqualifying at a research shop. Also be ready with the multiple-testing consequence, because the interviewer's real target is whether you understand why published anomalies do not replicate.
Expect next
- So what significance level would you use if you screened a thousand signals?
- Explain the false discovery rate.
- What would you report instead of a p-value?
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

