Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
071Here is a scenario. Walk me through how you would analyse the trade.SchonfeldQuantitative Research · New York · 2021
Say this
I would structure it as five questions: what is the thesis and what would make it wrong, what is the expected value, how do I size it, how do I hedge what I am not trying to be exposed to, and what is my exit. Then say the number, because a trade analysis without a number is an opinion.
Then walk it
- Thesis first, stated as a falsifiable claim with a horizon. Not this looks cheap, but I think this spread compresses from 80 to 50 basis points over three months because of a specific mechanism, and if it is still at 80 in three months I am wrong.
- Expected value: probability times payoff on each branch. If there is a 60 percent chance of making 3 and a 40 percent chance of losing 2, that is 1.8 minus 0.8, so plus 1 with a 5-point range of outcomes. The range matters as much as the mean.
- Sizing: from the loss branch, not the win branch. I size so that the bad case is a loss I can carry, which in practice means a fraction of my risk budget, and I say what that fraction is.
- Hedging: separate the exposure I want from the ones that come attached. If the view is idiosyncratic, hedge out the market beta, the sector, and the rate duration, then check what basis risk remains after hedging, because that is the risk I did not choose.
- Exit and monitoring: the level or the date at which I am out, plus the two or three observables that would tell me the thesis is breaking before the P&L does. And I would name the thing I cannot hedge, because every trade has one and being explicit about it is what makes the analysis credible rather than promotional.
Where candidates lose it
Describing the thesis at length and never getting to sizing, hedging or the exit. Anyone can have a view. What a multi-manager platform is hiring for is the risk framework around it, so spend at least half your answer on size, hedge and exit, and name the unhedgeable residual yourself.
Expect next
- What is your stop, and why there?
- What would make you double the position?
- What risk are you left with after hedging?
Reported by candidates at Schonfeld (Quantitative Research, New York, 2021). Source: Wall Street Oasis.
072Where does the money come from in market making versus a systematic hedge fund strategy?Prop trading firmsQuant trading
Say this
A market maker gets paid a fee for providing immediacy and aims to be flat at the end of the day. A systematic fund takes a position because it forecasts a return and holds risk overnight. One sells a service, the other takes a view.
Then walk it
- Market making: high turnover, tiny edge per trade, thousands of trades a day, Sharpe ratios that can be very high because the law of large numbers works for you, and capacity limited by volume rather than capital. Risk is inventory and adverse selection, measured in seconds to minutes.
- Systematic trading: lower turnover, larger edge per position, Sharpe typically 0.5 to 2, capacity limited by market impact, and risk measured in days to months. You are exposed to being simply wrong about the forecast.
- The counterparty differs, which is the deepest version of the answer. A market maker's profit comes from other participants' demand for immediate execution. A systematic fund's profit comes from other participants' mispricing, behavioural bias, or need to shed risk.
- Which tells you what kills each one. Market makers die from a fast informed move against a large inventory, or from technology failure. Systematic funds die from crowding, regime change, and leverage in a deleveraging.
- And it explains the career difference, which is usually the real reason the question is asked. Market making gives you feedback in minutes and rewards fast reaction under pressure. Research gives you feedback in months and rewards patience and statistical honesty. Saying which one suits you, with a reason, is what they are listening for.
Where candidates lose it
Treating them as the same job with different time horizons. The economic source of the profit is different, and saying it plainly, a fee for liquidity versus a return for taking a view, is what demonstrates real understanding. Then connect it to which seat you want, because that is where the question is going.
Expect next
- Which of those do you want to do and why?
- Why can market makers run much higher Sharpe ratios?
- What kills each business?
073Why do alphas decay, and how would you detect that yours is dying?Quant researchQuant trading
Say this
Because a profitable pattern attracts capital until the price moves to where the profit was. Detect it by tracking realised versus expected performance, the signal's own predictive power separately from the P&L, and crowding measures, and set the decision rule before performance deteriorates.
Then walk it
- Mechanisms in order of frequency. Crowding, where other people trade the same signal and the entry price moves. Structural change, where the market feature the signal exploited is regulated or engineered away. Arbitrage by faster participants. And plain overfitting, where the alpha was never there.
- Separate the two things that can break. Is the signal still predicting, measured by information coefficient, the correlation between forecast and subsequent return? Or is it predicting but no longer profitable after costs? The first is decay, the second is crowding or impact, and the fixes differ.
- Concrete measures: rolling information coefficient, rolling Sharpe, realised transaction cost versus modelled, and the fraction of your expected edge captured on a typical fill. If the signal is intact and the capture rate is falling, other people are in front of you.
- Crowding proxies: short interest and borrow costs for the short leg, correlation of your P&L with published factor returns, and how your strategy behaves on days when leveraged players deleverage. A crowded trade has fat negative tails on those days.
- The discipline is the answer though. Set the decay threshold in advance, for example halve the allocation if the rolling one-year information coefficient falls below half its backtest level for two consecutive quarters. Deciding in the middle of a drawdown is how people turn a decayed alpha into a large loss, and having the rule written down before you need it is the part an interviewer is actually testing.
Where candidates lose it
Answering only markets get more efficient. Be specific about mechanisms and about measurement, and above all separate whether the signal stopped predicting from whether the trade stopped being profitable. A pre-committed decision rule is the piece most candidates never mention.
Expect next
- What is an information coefficient and what is a good value?
- How would you measure crowding in a trade?
- Would you turn it off, or reduce it, and who decides?
074How would you model market impact and slippage for a strategy you are sizing?Quant researchQuant trading
Say this
Split the cost into spread, temporary impact and permanent impact. The empirical regularity worth knowing is the square-root law: impact scales roughly with the square root of the order size as a fraction of daily volume, times the volatility.
Then walk it
- The square-root law: impact in volatility units is approximately a constant times the square root of order size over average daily volume, with the constant usually estimated around 0.5 to 1. So trading 1 percent of ADV in a 2 percent daily vol name costs roughly 0.1 times 2 percent, about 20 basis points.
- That non-linearity is what caps capacity. Doubling your size only increases impact by 41 percent per share, but total cost grows as size to the power 1.5, so cost eats your edge faster than your edge grows.
- Separate temporary from permanent. Temporary impact reverts after you stop trading and is a function of how fast you trade. Permanent impact is the information your trading revealed, and it does not come back. Almgren-Chriss style frameworks trade off the two against the risk of trading slowly.
- Estimating it honestly: use your own fills against arrival price, not a vendor model, and regress realised shortfall on participation rate, volatility and spread. You need a lot of trades, and you must control for the fact that you traded more aggressively when you had more signal, which biases the estimate.
- And the modelling discipline: be conservative, because impact is the parameter most likely to turn a profitable backtest into a losing strategy. I would rather assume twice the cost and discover I was pessimistic than the reverse. That preference is the answer they want to hear.
Where candidates lose it
Assuming linear impact or using the quoted spread as the whole cost. For any size that matters the spread is the small part. Know the square-root law and know that cost scaling as size to the power 1.5 is what determines capacity, because that is the link between a research result and a business decision.
Expect next
- Why does cost scale as size to the power one and a half?
- How do you separate permanent from temporary impact empirically?
- How does impact determine the capacity of a strategy?
075How has electronic market making changed over the last decade, and where do you think the edge is now?Prop trading firmsQuant trading
Say this
Spreads have compressed to a tick or less in liquid products, the pure speed race has largely been won and commoditised by a handful of firms, and the remaining edge has moved to breadth of product, quality of the pricing model, and access to less-contested flow.
Then walk it
- What changed: colocation and microwave or hollow-core fibre links turned latency into a fixed capital cost rather than an edge, exchange data got faster and cheaper, and the number of firms who can compete at the top tier is small.
- Where it went. First, breadth: applying the same infrastructure across equities, options, futures, crypto, ETFs and fixed income, since each new product is incremental revenue on a paid-for stack. Second, modelling: in options and ETFs the hard part is pricing thousands of related instruments consistently, which is a research problem, not a wire problem.
- Third, flow quality. Internalising or purchasing retail flow is valuable precisely because it is less informed. That is the economics behind payment for order flow, and it is the reason the regulatory debate about it matters commercially.
- The structural trend in fixed income and credit is worth naming: electronic market making has moved into products that were voice-traded a decade ago, and ETF creation and redemption is the mechanism that makes bond market making hedgeable at all.
- My honest view, offered as a view and not a fact: the marginal edge now sits in products where pricing is genuinely hard rather than where speed is hard, because speed has a ceiling that has been reached and modelling does not. And I would caveat that I am reading this from the outside, which is part of why I want to work somewhere that sees it from the inside.
Where candidates lose it
Reciting high-frequency trading is about speed as if it were still 2010. The interviewer works at one of these firms and will know instantly. Have a specific, current view, name the shift from latency to breadth and modelling, and flag that it is your view rather than asserting inside knowledge you do not have.
Expect next
- Is payment for order flow good or bad for the end investor?
- Why is options market making harder than equities?
- What do you think our firm's edge is?
076You have K sorted arrays on disk, too large to load at once. How do you merge them into one sorted output?CitadelEquity Capital Markets · New York · 2026
Say this
K-way merge with a min heap of size K. Push the first element of each array into the heap, repeatedly pop the minimum and write it out, then push the next element from whichever array the minimum came from. Time is N log K, memory is O(K) plus your buffers.
Then walk it
- The heap holds one candidate per array, each entry tagged with which array it came from and the index within it. Pop the smallest, emit it, and refill from that same array.
- Complexity: N total elements, each pushed and popped once, each operation log K. So N log K, which beats concatenate-and-sort at N log N whenever K is much smaller than N.
- The disk part is the real content of the question. You do not read element by element, you read blocks. Keep a buffer per array, say a few megabytes each, refill it when it drains, and write the output through a large buffer too. The heap operations are free compared with I/O, so the design goal is sequential reads and few of them.
- If K is very large, K times the buffer size exceeds memory, and then you merge in passes: merge groups of, say, 100 files at a time, then merge the results. That is exactly how external merge sort works, and total I/O is N times the number of passes.
- Practical notes I would raise: use a tournament tree or a loser tree instead of a binary heap if you want fewer comparisons per element, handle the tie-breaking rule explicitly if stability matters, and if this is a real system, check whether the operating system's readahead is already doing your buffering for you before you build it yourself.
Where candidates lose it
Answering merge them pairwise, which is K times N in the worst case, or ignoring the on-disk part entirely. The interviewer put the data on disk deliberately, so talk about block-sized buffered reads and what happens when K is too large to buffer. State the N log K complexity explicitly.
Expect next
- What if K is a million?
- How large would you make the buffers, and why?
- How would you parallelise it?
Reported by candidates at Citadel (Equity Capital Markets, New York, 2026). Source: Wall Street Oasis.
077Given an array and a window of size k, return the maximum in each window as it slides.Akuna CapitalQuant Development · Chicago · 2025
Say this
Monotonic deque, O(n) total. Keep a deque of indices whose values are strictly decreasing. Before pushing a new index, pop from the back everything smaller than the new value, and pop from the front anything that has fallen out of the window. The front is always the maximum.
Then walk it
- Why the deque is monotonic: if a new element is larger than something behind it, that older smaller element can never be the maximum of any future window, because the new one is both larger and more recent. So it is safe to discard permanently.
- Each index is pushed once and popped once, so the total work is O(n) even though a single step can pop many elements. That amortised argument is the thing to say out loud, because it is what distinguishes this from the naive O(n k).
- Store indices, not values, so you can test whether the front has expired by comparing front index against i minus k plus 1.
- Alternatives and why they are worse: a max heap gives O(n log k) and needs lazy deletion of expired entries. A balanced BST or a multiset gives O(n log k) too. Both are fine and both are beaten by the deque.
- Where this actually matters on a trading system, which is worth mentioning: rolling extremes over a tick window, running high and low for a breakout signal, and rolling maximum drawdown. The same structure with the comparison reversed gives you the rolling minimum, and the O(1) amortised cost per tick is what makes it usable in a hot path.
Where candidates lose it
Reaching for a heap and stopping there. The heap answer is acceptable but it is not the answer to this question, and the interviewer is specifically looking for the monotonic deque and the amortised O(n) argument. Also remember to expire the front by index, which is the bug that shows up most often in live coding.
Expect next
- Prove the amortised complexity.
- Now give me the rolling median instead.
- How would you handle a window defined by time rather than by count?
Reported by candidates at Akuna Capital (Quant Development, Chicago, 2025). Source: Wall Street Oasis.
078Given a stream of numbers, return the median after each element arrives.Old Mission CapitalEquities · Boston · 2024
Say this
Two heaps. A max heap for the lower half and a min heap for the upper half, kept balanced so their sizes differ by at most one. The median is the top of the larger heap, or the average of the two tops. Insert is O(log n), query is O(1).
Then walk it
- Insert rule: if the new value is at most the max of the lower heap, push it there, otherwise push to the upper heap. Then rebalance by moving one element across if the sizes differ by more than one.
- Query: if the sizes are equal, the median is the average of the two tops. Otherwise it is the top of the larger heap. Constant time either way.
- Total cost for n elements is n log n, and memory is O(n) because you must retain everything. That memory cost is the honest limitation, and it is the first thing an interviewer will probe.
- If the median must be over a sliding window rather than the whole prefix, the two-heap approach needs deletions from the middle. Use an indexed multiset or two heaps with lazy deletion and a hash of pending removals. That is the version that comes up in practice on a tick stream.
- And if approximate is acceptable, which on a trading system it usually is, the right answer is a streaming quantile sketch: t-digest or the Greenwald-Khanna algorithm, giving you any quantile in bounded memory rather than O(n). Naming that unprompted is what turns a correct interview answer into a practical one.
Where candidates lose it
Sorting on every element, which is O(n squared log n) overall, or maintaining a sorted list with insertion, which is O(n) per element because of the shifting even though the binary search is fast. Say two heaps immediately, then volunteer the sliding-window and bounded-memory variants, because that is where the conversation is heading.
Expect next
- Now do it over a sliding window of the last thousand values.
- What if you cannot store all the data?
- How would you get the 99th percentile instead of the median?
Reported by candidates at Old Mission Capital (Equities, Boston, 2024). Source: Wall Street Oasis.
079How would you store key-value pairs, and what are the tradeoffs between the implementations?Jump TradingEngineering · Cambridge · 2019
Say this
Hash table for O(1) average lookup with no ordering, balanced tree for O(log n) with ordered iteration and range queries, and a flat sorted array if the data is static and you care about cache behaviour. The choice is driven by whether you need ordering and what your access pattern looks like in memory.
Then walk it
- Hash table: O(1) average, O(n) worst case on collisions, no ordering, and rehashing causes an occasional large latency spike. That spike is a real problem on a trading hot path and it is why people pre-size their maps.
- Balanced tree, red-black or B-tree: O(log n) guaranteed, ordered traversal, range queries, and predictable latency. Worse constants and worse cache locality because of pointer chasing.
- The tradeoff that matters most in practice is memory layout, not big-O. C++ unordered_map uses separate chaining with nodes scattered across the heap, so every lookup is potentially a cache miss. An open-addressing flat hash map keeps everything in one array and is commonly two to three times faster in real workloads at the same asymptotic complexity.
- For a mostly-static table, a sorted array with binary search beats both: contiguous memory, no pointers, and for small n a linear scan beats binary search because it is branch-predictable and prefetchable. Under about 16 to 32 entries, linear wins.
- And on disk the answer changes completely: B-trees for read-heavy workloads because of the branching factor against block size, LSM trees for write-heavy because they turn random writes into sequential ones. I would want to know the read-write ratio and whether the working set fits in cache before choosing anything.
Where candidates lose it
Answering hash map, O(1), done. The question says tradeoffs, so it is a systems question and the interviewer at a trading firm cares about tail latency and cache behaviour more than asymptotic complexity. Mention rehashing spikes and pointer chasing, and ask what the access pattern is.
Expect next
- Why is std::unordered_map often slow in practice?
- How would you avoid latency spikes from rehashing?
- What changes if the data lives on disk?
Reported by candidates at Jump Trading (Engineering, Cambridge, 2019). Source: Wall Street Oasis.
080Can you implement a linked list, and when would you actually use one on a trading system?Jump TradingProp Trading · Remote · 2022
Say this
Yes, a node with a value and a next pointer, plus a head, and the usual care about the empty list and about updating head when you insert or delete at the front. But the honest answer to the second half is: rarely, because pointer chasing destroys cache performance.
Then walk it
- The implementation: struct with value and next, insert at head in O(1), search in O(n), delete given the previous node in O(1). Use a dummy head node and most of the edge cases disappear, which is the trick worth knowing for interviews.
- The standard edge cases they will check: empty list, single element, deleting the head, and not leaking the node you unlinked. In C++ that means being explicit about ownership, and in a real codebase it means a unique pointer or an arena.
- What a linked list genuinely buys you: O(1) splice of a node from the middle if you already hold a pointer to it, and stable addresses so a pointer stays valid across insertions. That is exactly the requirement in a limit order book, where you need to cancel an arbitrary resting order in constant time, so orders at a price level are typically an intrusive doubly linked list with a hash from order id to node.
- What it costs: every traversal is a potential cache miss, and a vector beats a list for iteration by an order of magnitude even when the asymptotics say otherwise.
- So the real-world answer is an intrusive list over a pre-allocated node pool, not a textbook list with individual heap allocations. Saying that is the difference between having done the exercise and having written low-latency code.
Where candidates lose it
Writing the code correctly and having nothing to say about why you would use one. At a trading firm the interesting half is the memory and cache discussion, and the order book cancel case is the one concrete example where a linked list is genuinely the right structure. Also do not forget the dummy head trick, it removes most of the bugs.
Expect next
- Reverse it in place.
- Detect a cycle in constant space.
- Why would a vector usually beat a list even when the complexity says otherwise?
Reported by candidates at Jump Trading (Prop Trading, Remote, 2022). Source: Wall Street Oasis.
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

