Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
001There are two bags of stones and you do not know how many black or white are in each. You draw two stones and both are black. What is the probability the next one is black, and will you bet on it?CitadelQuantitative Trading · New York · 2025
Say this
Higher than a half, and yes I would bet on black. Because I do not know the composition, the two black draws are evidence about the composition itself, so I update towards bags that are black-heavy. The draws are not independent trials, they are a sample that teaches me about the urn.
Then walk it
- Set it up properly: put a prior over the unknown mixture, say the proportion of black p is uniform on zero to one, and the draws are conditionally independent given p.
- Then this is Laplace's rule of succession. With k blacks out of n draws the posterior predictive probability of another black is (k+1)/(n+2). Two blacks out of two gives 3/4.
- The intuition without algebra: seeing black twice shifts the posterior mass towards high p, and the predictive probability is the posterior mean of p, which is now above a half.
- Compare it with the alternative model. If I were told the bag was exactly 50/50 and I was drawing with replacement, the answer would be exactly a half and the history would be irrelevant. The whole question is which model you are in.
- On the betting half: I would take anything better than even money on black, and I would size it small because 3/4 is a function of my prior, not of data. Two draws is almost no information. If the prior were concentrated near a half the answer moves back towards a half.
Where candidates lose it
Saying one half because the draws are independent. They are only independent conditional on the unknown composition, and the composition is exactly what you are learning. The second failure is giving 3/4 with no mention of the prior, as if it were a fact rather than the output of a uniform prior you chose.
Expect next
- What if the prior were Beta(2,2) instead of uniform?
- Now make me a market on the probability and I will trade it.
- Same question but sampling without replacement from a bag of 10 stones. Does the answer move?
Reported by candidates at Citadel (Quantitative Trading, New York, 2025). Source: Wall Street Oasis.
076You have K sorted arrays on disk, too large to load at once. How do you merge them into one sorted output?CitadelEquity Capital Markets · New York · 2026
Say this
K-way merge with a min heap of size K. Push the first element of each array into the heap, repeatedly pop the minimum and write it out, then push the next element from whichever array the minimum came from. Time is N log K, memory is O(K) plus your buffers.
Then walk it
- The heap holds one candidate per array, each entry tagged with which array it came from and the index within it. Pop the smallest, emit it, and refill from that same array.
- Complexity: N total elements, each pushed and popped once, each operation log K. So N log K, which beats concatenate-and-sort at N log N whenever K is much smaller than N.
- The disk part is the real content of the question. You do not read element by element, you read blocks. Keep a buffer per array, say a few megabytes each, refill it when it drains, and write the output through a large buffer too. The heap operations are free compared with I/O, so the design goal is sequential reads and few of them.
- If K is very large, K times the buffer size exceeds memory, and then you merge in passes: merge groups of, say, 100 files at a time, then merge the results. That is exactly how external merge sort works, and total I/O is N times the number of passes.
- Practical notes I would raise: use a tournament tree or a loser tree instead of a binary heap if you want fewer comparisons per element, handle the tie-breaking rule explicitly if stability matters, and if this is a real system, check whether the operating system's readahead is already doing your buffering for you before you build it yourself.
Where candidates lose it
Answering merge them pairwise, which is K times N in the worst case, or ignoring the on-disk part entirely. The interviewer put the data on disk deliberately, so talk about block-sized buffered reads and what happens when K is too large to buffer. State the N log K complexity explicitly.
Expect next
- What if K is a million?
- How large would you make the buffers, and why?
- How would you parallelise it?
Reported by candidates at Citadel (Equity Capital Markets, New York, 2026). Source: Wall Street Oasis.
084How would you design a system to troubleshoot latency in a trading stack?CitadelProp Trading · New York · 2026
Say this
Timestamp at every hop with one clock, measure distributions not averages, and make the whole path attributable so you can say which segment consumed the microseconds. The design principle is that you cannot fix what you cannot decompose.
Then walk it
- Instrumentation: hardware timestamps at the network card for packet in and packet out, plus software timestamps at each stage, market data decode, book update, strategy decision, order encode, and kernel bypass send. Carry a correlation id through the whole chain so a single event can be reconstructed end to end.
- Clocks are the hard part. Use PTP with hardware timestamping across hosts, not NTP, and record clock offset and drift as first-class data. Two hosts disagreeing by fifty microseconds will invent latency that does not exist and hide latency that does.
- Statistics: report the median, the 99th, the 99.9th and the maximum. Averages are useless here because the distribution is heavily right-tailed and the tail is exactly what costs money. Track per-segment histograms, ideally with HDR histograms so the tail resolution survives.
- Storage and analysis: stream the records off the critical path into a time-series store, then build the two views that actually get used, a per-segment breakdown over time and a drill-down into the slowest individual events. Alert on percentile regressions against a rolling baseline rather than on fixed thresholds.
- Then the causes to design for, because the system exists to distinguish them: garbage collection or allocation pauses, page faults, context switches and CPU migration, interrupt coalescing settings, cache misses and false sharing, queueing at the exchange gateway, and simple network congestion. And I would say the measurement must not itself be on the hot path, so lock-free ring buffers with a separate reader thread, because an observability system that adds ten microseconds has destroyed what it measures.
Where candidates lose it
Describing logging and monitoring generically. This is a specific systems question and the differentiators are clock synchronisation, percentile rather than mean reporting, and keeping instrumentation off the critical path. Talk in microseconds, and be able to name concrete causes of a tail latency spike.
Expect next
- How do you synchronise clocks across hosts, and to what accuracy?
- Why report the 99.9th percentile rather than the average?
- Walk me through diagnosing a spike that happens once a day.
Reported by candidates at Citadel (Prop Trading, New York, 2026). Source: Wall Street Oasis.
098How would you describe the internet to someone who has just time travelled from thirty years ago?CitadelQuantitative Trading · Sydney · 2025
Say this
Anchor it in something they already have, then build outward. Thirty years ago they had a telephone network and a public library, so I would say: imagine the phone system, but instead of connecting two people it connects every machine, and instead of voices it carries copies of anything written down.
Then walk it
- Start from their reference frame. Someone from the mid-nineties knows telephones, fax machines, libraries, post and television. Every explanation should be a modification of one of those, not a new concept from scratch.
- Then the one structural idea that makes everything else follow: the cost of copying and sending information fell to approximately zero, and nobody owns the middle of the network. From that single fact you can derive email, search, shopping, streaming and social media without explaining any of them separately.
- Give one concrete, checkable example. The entire contents of that library, searchable in a quarter of a second from a device in your pocket, and it is free because someone sells advertising next to the answer.
- Then say what they would find strangest, because that is the interesting part of the answer. Not the technology but the consequences: that attention became the scarce resource, that distance stopped mattering for information but not for goods, and that privacy became something you trade rather than something you have.
- Check understanding rather than lecturing. I would ask them what they want to do with it and explain that, because the tailored answer lands and the general one does not. This question is a communication test, so the move that scores is structure plus adapting to the listener, not encyclopaedic completeness.
Where candidates lose it
Diving into TCP/IP, packets and DNS. The question tests whether you can teach, not whether you know how it works. Anchor in their existing mental model, lead with one organising idea, and use a concrete example. Rambling without structure is the common failure, so state your three-part plan before you start.
Expect next
- Now explain machine learning to the same person.
- What would they find most surprising about the consequences?
- Do it in thirty seconds instead.
Reported by candidates at Citadel (Quantitative Trading, Sydney, 2025). Source: Wall Street Oasis.
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

