Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
013You have a feed of a hundred thousand data points and you know fifteen of them are missing, recorded as zeros at the end. If you pull a window, what is the probability of at least one missing value?Jump TradingProp Trading · Remote · 2022
Say this
Use the complement. For a sample of n points drawn without replacement from 100,000 of which 15 are bad, the probability of at least one bad is one minus the hypergeometric probability of none, which is one minus the product over i of (99,985 minus i)/(100,000 minus i). For small n that is well approximated by one minus (1 minus 0.00015) to the n.
Then walk it
- Always compute at least one as one minus none. Summing the cases is the slow road and it invites double counting.
- The exact object is hypergeometric: choose n from 99,985 good over choose n from 100,000. For n much smaller than 100,000 the with and without replacement answers agree to several decimals.
- Numbers give it life. p is 15 over 100,000, which is 0.00015. For a window of 1,000 points, one minus 0.99985 to the 1000 is about 13.9 percent. For a window of 100 it is about 1.5 percent. So this is a real problem, not a rounding issue.
- Useful shortcut: for small p and moderate n the answer is roughly n times p, capped by 1. A thousand times 0.00015 is 0.15, close to the exact 0.139, and the Poisson approximation 1 minus e to the minus 0.15 gives 0.1393, which is very close.
- The thing I would say next on a desk, because it is the real question: they are at the end of the series, which is not random at all. If they are the most recent 15 points, then any window containing the tail hits all 15 with certainty and every other window hits none. Position matters more than the count.
Where candidates lose it
Treating the missing points as randomly scattered when the question says they sit at the end. That is the detail being tested. Give the hypergeometric answer for the random case, then flag the structural point: trailing zeros are usually a feed-truncation artefact, so the right fix is to detect and drop the tail, not to price the probability.
Expect next
- How would you detect that the zeros are missing values rather than genuine zeros?
- What is the Poisson approximation and when does it break?
- How do you handle those points in a model without leaking future information?
Reported by candidates at Jump Trading (Prop Trading, Remote, 2022). Source: Wall Street Oasis.
079How would you store key-value pairs, and what are the tradeoffs between the implementations?Jump TradingEngineering · Cambridge · 2019
Say this
Hash table for O(1) average lookup with no ordering, balanced tree for O(log n) with ordered iteration and range queries, and a flat sorted array if the data is static and you care about cache behaviour. The choice is driven by whether you need ordering and what your access pattern looks like in memory.
Then walk it
- Hash table: O(1) average, O(n) worst case on collisions, no ordering, and rehashing causes an occasional large latency spike. That spike is a real problem on a trading hot path and it is why people pre-size their maps.
- Balanced tree, red-black or B-tree: O(log n) guaranteed, ordered traversal, range queries, and predictable latency. Worse constants and worse cache locality because of pointer chasing.
- The tradeoff that matters most in practice is memory layout, not big-O. C++ unordered_map uses separate chaining with nodes scattered across the heap, so every lookup is potentially a cache miss. An open-addressing flat hash map keeps everything in one array and is commonly two to three times faster in real workloads at the same asymptotic complexity.
- For a mostly-static table, a sorted array with binary search beats both: contiguous memory, no pointers, and for small n a linear scan beats binary search because it is branch-predictable and prefetchable. Under about 16 to 32 entries, linear wins.
- And on disk the answer changes completely: B-trees for read-heavy workloads because of the branching factor against block size, LSM trees for write-heavy because they turn random writes into sequential ones. I would want to know the read-write ratio and whether the working set fits in cache before choosing anything.
Where candidates lose it
Answering hash map, O(1), done. The question says tradeoffs, so it is a systems question and the interviewer at a trading firm cares about tail latency and cache behaviour more than asymptotic complexity. Mention rehashing spikes and pointer chasing, and ask what the access pattern is.
Expect next
- Why is std::unordered_map often slow in practice?
- How would you avoid latency spikes from rehashing?
- What changes if the data lives on disk?
Reported by candidates at Jump Trading (Engineering, Cambridge, 2019). Source: Wall Street Oasis.
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

