Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
031Take a random ordering of n distinct numbers and run exactly one left-to-right pass of bubble sort, swapping each adjacent pair that is out of order. What is the probability the list is fully sorted afterwards? Work it for n = 5.Jump TradingChicago · 2018
Try it first
For n = 5, how likely is the list sorted after one pass?
Show the worked solution
2 to the power (n - 1) divided by n factorial, which is 16/120 = 2/15 for n = 5. One pass moves every number that is not carried rightwards exactly one place left. So the list ends sorted only if no number starts more than one place right of its final spot. Placing 1, then 2, then 3 and so on, each has two allowed spots and the largest takes the last one, giving 2 to the power (n - 1) orderings.
What does one pass actually do to each number?
Picture a queue at a ticket window where the tallest person seen so far keeps stepping back past anyone shorter. That person travels a long way to the right; everyone they pass shifts one step forward. In one pass, the running maximum is carried right until it meets something larger, and every number it passes moves exactly one place left. Nothing moves left by two in a single pass. That limit is the whole problem.
In 3 1 2 5 4 every number starts at most one place right of its home, so one pass sorts it; in 2 3 1 4 5 the 1 starts two places right of home and ends one short, so only 16 of the 120 orderings of five numbers, 2 in 15, sort in one pass. Which orderings survive, and how do you count them?
Because a number can shift left by one at most, the list sorts only if each number starts no more than one place right of its home. The converse also holds: when every number meets that condition, the pass carries each big number to exactly where it belongs. For five numbers, a brute-force check of all 120 orderings finds exactly the 16 that meet the condition, and all 16 sort.
Now count them without listing. Place the numbers in increasing order. The 1 may sit in position 1 or 2. The 2 may sit anywhere in positions 1 to 3, one of which the 1 already took: two choices. The same holds for 3 and 4: each has k + 1 allowed spots, k - 1 of them already used by smaller numbers, so two choices each. The 5 fills the one position left. That is 2 x 2 x 2 x 2 x 1 = 16.
The relationship2^(n-1) orderings where no number starts more than one place right of its home n! all orderings of n distinct numbers, equally likely What it says in wordsTwo choices for each number except the largest, over all possible orderings.Check small cases out loud: for n = 2 both orderings sort, 2 of 2; for n = 3 it is 4 of 6. The probability collapses fast, because n factorial outruns 2 to the power n: about 4.4% for n = 6 and 1.3% for n = 7.
Where candidates lose it
The common wrong start is to think one pass only fixes the largest number, and answer that the other n - 1 must already be sorted, which gives 1/(n - 1)! and 1/24 for n = 5. It misses that every passed number also moves left one place, which rescues many orderings.
The other loss is guessing a rule from one example. State the one-step-left limit, derive the condition from it, then count by placing numbers in increasing order.
What the interviewer asks next
- What is the probability the list is sorted after two passes?
- How many passes does bubble sort need on average for a random list of n numbers, roughly?
- What if the pass runs right to left instead?
Asked at Jump Trading, Research, Chicago, 2018 (Wall Street Oasis):
one iteration of bubble sort, what's the probability that the array will be sorted
032n points are placed independently and uniformly on a circle of circumference 1, with n at least 3. Each point colours the arc between itself and its nearest neighbour. What is the expected total length that gets coloured?Susquehanna International GroupLondon · 2026
Try it first
Which is closest to the expected coloured length?
Show the worked solution
7/18, about 0.389, for every n from 3 upwards. A gap is left uncoloured only when it is longer than both gaps beside it, because then neither endpoint has it as its nearest. For three points that gap is simply the longest of three pieces, which averages 11/18, so 7/18 is coloured. For larger n the same 11/18 comes out, so the answer does not depend on n.
When is a gap left uncoloured?
Picture people standing round a circular table, each turning to talk to whoever is closer, left or right. A stretch of table between two people stays silent only if both of them turned away, which means each had a closer person on their other side. A gap is uncoloured exactly when it is longer than both of its neighbouring gaps. A gap coloured from both ends is still coloured once, so the question becomes: what is the expected total length of gaps that are local maxima?
Each gap is coloured if it is the shorter gap for at least one endpoint and left uncoloured if it is longer than both neighbours; this sample of ten points colours 0.57 of the circle, and the average over all placements is 7/18, about 0.389, for any n of 3 or more. How do you get 11/18 for the uncoloured part?
Start with n = 3, the case you can finish in the room. With three gaps, every gap's two neighbours are the other two gaps, so the only uncoloured gap is the longest one. Three random points cut the circle like a stick broken into three, and the longest of three pieces averages (1/3)(1 + 1/2 + 1/3) = 11/18. So the coloured length is 7/18.
For larger n, use the fact that the n gaps behave like n independent exponentialA random length whose chance of ending is the same at every instant; waiting times between random arrivals follow it. lengths rescaled to add up to 1, and that the rescaling is independent of the shape. For three unit exponentials X, Y and Z, the expected value of X counted only when X is the largest is 1 - 2/4 + 1/9 = 11/18. Each of the n gaps contributes that, divided by the expected total of n, and the n gaps sum to 11/18 again. The uncoloured share is 11/18 whatever n is, so the coloured share is always 7/18.
The relationshipx e^(-x) a gap's length times its density, in the exponential picture (1 - e^(-x))^2 the chance both neighbouring gaps are shorter 1/n rescaling so the n gaps add to a circle of length 1 What it says in wordsThe expected length of gaps longer than both neighbours is 11/18, and the rest of the circle is coloured.Say the check: a seeded simulation of 40,000 random circles gives 0.389 for n = 3, 0.388 for n = 5 and 0.389 for n = 10. The limitation is that the exponential step is a known result you should name, not derive, in an interview; the n = 3 case is the part you prove on the spot.
Where candidates lose it
The usual loss is counting gaps instead of measuring them. One gap in three is a local maximum, so candidates answer 2/3 coloured. The uncoloured gaps are selected for being long, which is why their share of length, 11/18, is far above one third.
The second is double counting a gap that both endpoints colour. It is coloured once. Frame the problem around uncoloured gaps and both mistakes disappear.
What the interviewer asks next
- What is the expected number of uncoloured gaps?
- What if each point colours the arc to its farther neighbour instead?
- Does the answer change for points on a line segment rather than a circle?
Asked at Susquehanna International Group, Quantitative Research, London, 2026 (Wall Street Oasis):
if n points are placed on a circle and each point colours in the arc to its nearest neighbour
033There are 21 matches on the table. Two players alternate taking 1, 2 or 3 matches, and whoever takes the last match loses. Would you rather go first or second, and what is your strategy?Quant tradingProp trading firms
Try it first
First or second?
Show the worked solution
Go second. The player facing 1, 5, 9, 13, 17 or 21 matches loses against good play, and 21 is on that list. Whatever your opponent takes, take 4 minus that, so every round removes exactly 4. They then face 17, 13, 9, 5 and finally 1, and are forced to take the last match.
How do you find the losing positions?
Start from the end of the game, the way you would plan the last few stops of a journey before the first. With 1 match in front of you, you must take it and lose. With 2, 3 or 4, you take enough to leave exactly 1 and win. With 5, every move leaves 2, 3 or 4, each a winning spot for the other player, so 5 loses. A position is losing when every move from it hands your opponent a winning position, and here that happens every 4 matches: 1, 5, 9, 13, 17, 21.
Counting down from 21, the positions 21, 17, 13, 9, 5 and 1 lose for the player about to move; moving second and answering each take of k with 4 minus k keeps the opponent on those positions until they must take the last match. Why does answering with 4 minus k always work?
Your opponent can take 1, 2 or 3; you can always take 3, 2 or 1 in reply. The pair of moves removes exactly 4 matches whatever they chose, so you control the count at the end of every round. From 21 the rounds end at 17, 13, 9 and 5, and then your opponent faces a single match. The number 4 is the maximum take plus one; that is where the period comes from.
A short check you can say aloud: the game has only 21 positions, and marking each as winning or losing from the bottom up, a position wins if any move reaches a losing one, reproduces the list 1, 5, 9, 13, 17, 21. If the pile had been 20, you would go first and take 3 to leave 17. This type of game has a backward inductionSolving a game by working out the best move at the last step first, then the step before, back to the start. solution, and the interviewer mainly wants to hear you build it from the end.
Where candidates lose it
The usual loss is playing forward: taking a few matches and hoping to spot the pattern mid-game. Under time pressure that becomes guessing. Work back from one match and the period of 4 appears in three steps.
The other slip is copying the rule for the version where taking the last match wins. There the losing spots are multiples of 4, and 21 means you should go first and take 1. Read which way the last match counts before you answer.
What the interviewer asks next
- What if taking the last match wins instead?
- What if each player may take 1 to 4 matches?
- What if there are two piles and you may take any number from one pile?
034Make me a two-way market on the number of heads in 100 flips of a fair coin, and justify the width.DRWNew York · 2026
Try it first
What is the standard deviation of the number of heads?
Show the worked solution
Centre it at 50 and quote around 46 at 54. The fair value is exactly 50. The standard deviation is √(100 x 0.5 x 0.5) = 5, so settlement lands between 45 and 55 about 73% of the time. A market 4 either side of fair earns 4 per lot on any trade, loses on a single sale at 54 only 18% of the time, and leaves room to move the quote if the other side seems to know something.
Where does the centre come from, and what sets the width?
A shopkeeper selling mangoes by the dozen knows the fair price; the margin he adds depends on how much the price of the next crate can swing and on whether the buyer knows something he does not. The centre of your market is the expected value, and the width is a choice about risk and information, scaled by how much the outcome can move. Here the expected value is 100 x 0.5 = 50, and nobody can know more than you about fresh flips of a fair coin, so the width is about risk alone.
The number of heads in 100 fair flips is centred at 50 with a standard deviation of 5, landing in 45 to 55 72.9% of the time and in 40 to 60 96.5% of the time, so a market of 46 at 54 sits inside one standard deviation and earns 4 per lot on each side. How do you justify 46 at 54 rather than 49 at 51?
Use the standard deviation as the ruler. The count has variance 100 x 0.5 x 0.5 = 25, so a standard deviation of 5. A quote 4 either side of fair earns 4 on each lot traded, against a settlement that typically moves 5, so every trade has an edge worth a large fraction of its risk. If someone buys at 54, you lose only if the count finishes at 55 or more, about 18% of the time. A tight 49 at 51 earns 1 per lot and a sale at 51 loses whenever the count reaches 52, about 38% of the time. Tighter wins more trades and earns less on each; in an interview game, start around one standard deviation wide and tighten as you learn.
The relationshipn = 100 number of flips p = 0.5 chance of heads on each flip sigma standard deviation of the number of heads What it says in wordsThe count of heads averages 50 and typically lands within 5 of it.Then say how you would react to trades, because that is the follow-up. If the interviewer lifts your 54 again and again, either they are testing your nerve or they know something, perhaps that the coin is not fair or that some flips are already done. Repeated one-way trading is information: move your market toward it and cut your size, rather than defending 50. The limitation of the simple answer is exactly that it assumes nobody knows more than you.
Where candidates lose it
The common loss is quoting 50 at 50, or 49.5 at 50.5, and calling it fair. A market maker earns the spread; a zero-width quote gives away every trade at no edge and leaves no room to adjust when the other side knows more.
The second is quoting a width with no reason. Name the standard deviation of 5, then choose a width against it. The number you say matters less than showing that width and risk are linked.
What the interviewer asks next
- I buy 10 lots at 54. Where is your new market?
- Now 60 flips have already happened and I have seen them. How does your market change?
- Make a market on the number of heads squared.
Asked at DRW, Quantitative Trading, New York, 2026 (Wall Street Oasis):
Make a market on the number of heads out of 100 coin flips.
035Regressing y on x gives a slope of 0.8; regressing x on y gives a slope of 0.45. What is the R-squared of either regression, and what is the correlation?Tower Research CapitalNew York · 2014
Try it first
What is the correlation between x and y?
Show the worked solution
R-squared is 0.36 for both regressions and the correlation is 0.6. The slope of y on x is r times sd(y)/sd(x); the slope of x on y is r times sd(x)/sd(y). Multiplying them cancels the standard deviations and leaves r squared: 0.8 x 0.45 = 0.36. The correlation is +0.6, positive because both slopes are positive, and the ratio sd(y)/sd(x) is √(0.8/0.45) = 4/3.
Why are the two slopes not reciprocals of each other?
Tall parents tend to have tall children, but a little less tall; and tall children tend to have tall parents, but a little less tall. Both statements are true at once. Each regression predicts toward the mean, so neither slope is the inverse of the other unless the fit is perfect. If the points lay exactly on a line, the slope of x on y would be 1/0.8 = 1.25. It is 0.45 instead, and the size of that shortfall is what measures how loose the relationship is.
Fitting y on x gives the shallower line with slope 0.8 and fitting x on y gives the steeper line, slope 0.45 in its own terms; their product, 0.36, is R-squared, so the correlation is 0.6 and the standard deviation of y is 4/3 that of x. How do the two slopes give R-squared?
Write each slope in terms of the correlation. The least squares slope of y on x is the covariance over the variance of x, which is r times sd(y)/sd(x). Swap the roles and the slope of x on y is r times sd(x)/sd(y). The standard deviation ratios are reciprocals, so the product of the two slopes is r squared, and in a one-variable regression r squared is exactly the R-squared. Here 0.8 x 0.45 = 0.36, so r = 0.6; the sign is positive because both slopes are positive, and the two slopes always share a sign.
The relationshipb_y|x slope from regressing y on x, 0.8 b_x|y slope from regressing x on y, 0.45 s_x, s_y standard deviations of x and y r the correlation of x and y What it says in wordsThe two slopes multiply to the squared correlation because the scale factors cancel.The figure uses 40 points built with standard deviations 3 and 4 and a correlation of exactly 0.6, and fitting both regressions returns slopes of 0.80 and 0.45. A quick sanity test comes free: the product of the two slopes can never exceed 1. If an interviewer quotes slopes of 0.8 and 1.5, the product 1.2 is impossible, and saying so is worth more than any calculation.
Where candidates lose it
The fast wrong answer is to say the slopes should be reciprocals and call the data inconsistent, or to answer 0.36 when asked for the correlation. 0.36 is R-squared; the correlation is its square root.
The second loss is dropping the sign. The square root of 0.36 could be plus or minus 0.6; both slopes are positive, so the correlation is positive, and saying why takes one sentence.
What the interviewer asks next
- What is the ratio of the standard deviation of y to that of x?
- If the slope of x on y were 1.5, what would you conclude?
- How does adding measurement noise to x change each slope?
Asked at Tower Research Capital, Quantitative Research, New York, 2014 (Wall Street Oasis):
Another detailed linear regression questions were asked, including problems about residual, variance and R^2
036A stock pays a growing dividend and is valued with the Gordon model at a discount rate of 10% and growth of 6%. What is its duration, and roughly how much does its price change if the discount rate rises by one point?BlackRockNew York · 2026
Try it first
What is the stock's duration, its percentage price sensitivity to the discount rate?
Show the worked solution
Duration is 1/(r - g) = 25 years, so a one-point rise cuts the value by about a fifth. The Gordon price is D1/(r - g), and its percentage sensitivity to r is 1/(r - g) = 1/0.04 = 25. With a Rs 4 dividend the price moves from Rs 100 at 10% to Rs 80 at 11%, a 20% fall. The 25% duration estimate overshoots because the price curve is convex.
Why does a stock have a duration at all?
A promise of money in one year hardly changes in value when rates move; a promise of money in twenty five years changes a lot, because the rate is compounded over every one of those years. A stock is a stream of dividends stretching forever, and when the dividends grow, most of its value sits in cash flows far in the future, so it behaves like a very long bond. Duration measures exactly that: the percentage price change for a change in the discount rate.
The relationshipD1 next year's dividend, Rs 4 in the illustration r discount rate, 10% g dividend growth rate, 6% 1/(r - g) percentage price change per unit change in r What it says in wordsDifferentiate the Gordon price and divide by price: the sensitivity is one over the gap between the discount rate and growth.At a 10% discount rate the Rs 4 dividend stock is worth Rs 100 and its duration is 25; at 11% the price is Rs 80, a 20% fall, while the tangent line predicts Rs 75, and at 12% the gap widens to Rs 66.7 against Rs 50 because the price curve is convex. Why does the estimate say 25% when the price falls 20%?
Duration is the slope at one point, and the price curve bends. For a one-point rise the tangent predicts a 25% fall, but the exact move from Rs 100 to Rs 80 is 20%, because the curve is convexCurving upward, so it always sits above any of its tangent lines. and flattens as r rises. The same bend makes a one-point fall worth more than 25%: at 9% the price is Rs 133.3, up 33.3%. For a big rate move, reprice exactly instead of trusting the slope. Duration here also equals price over dividend, 100/4, which is a quick way to say it: one over the dividend yield.
For precision, the Macaulay durationThe present-value weighted average time at which cash flows arrive. is (1 + r)/(r - g) = 27.5 years, and dividing by 1 + r gives the modified duration of 25. A 10-year bond paying 10% at a 10% yield has a modified duration of about 6.1. The gap r - g is what matters, which is why high-growth stocks carry the most duration: the same stock with 2% growth would have a duration of 12.5 years. The limitation is that the Gordon model holds growth fixed while rates move; in practice both shift together.
Where candidates lose it
The usual loss is saying a stock has no duration because it has no maturity, or that it is infinite because it pays forever. Both skip the one line of calculus that gives 1/(r - g).
The second is quoting 25% as the exact price change. It is the slope at 10%; the exact fall to 11% is 20%, and saying why, convexity, is what separates a strong answer.
What the interviewer asks next
- What happens to duration as growth approaches the discount rate?
- Why might a stock's measured sensitivity to bond yields be much lower than 25?
- What is the price change for a one-point fall in the discount rate?
Asked at BlackRock, Restructuring, New York, 2026 (Wall Street Oasis):
Which equities have duration ? multiple stocks vs value stocks MSE Forecasting equation
037Without paper: work out 56 x 56 and 73 x 74, and say the shortcut you used for each.Akuna CapitalChicago · 2025
Try it first
What is 56 x 56?
Show the worked solution
56 x 56 = 3,136 and 73 x 74 = 5,402. For 56 squared, split it as 50 + 6: 2,500, plus two strips of 300, plus 36. For 73 x 74, anchor both on 70: 4,900, plus 70 x 7 = 490, plus 3 x 4 = 12. A second route checks each: (60 - 4) squared = 3,136, and 73.5 squared minus a quarter = 5,402.
What is the shortcut for squaring a two-digit number?
Tiling a floor that is 56 tiles on each side, you would lay the big 50 by 50 block first, then two thin strips along the edges, then a small corner. Splitting a number into a round base plus a small part turns one hard product into one easy square and a few small ones: (a + b) squared = a squared + 2ab + b squared. For 56: 2,500 + 2 x 300 + 36 = 3,136. You can also go down from the next round number: (60 - 4) squared = 3,600 - 480 + 16, again 3,136.
56 squared splits into a 2,500 block, two 300 strips and a 36 corner, total 3,136; 73 x 74 splits into 4,900, 280, 210 and 12, total 5,402, the same answer as 73.5 squared minus a quarter. What changes when the two numbers differ, as in 73 x 74?
When two numbers share a tens digit, anchor both on it. For (70 + 3)(70 + 4), the product is 70 squared, plus 70 times the sum of the units, plus the product of the units: 4,900 + 490 + 12 = 5,402. The midpoint route gives the same: numbers equally spaced around 73.5 multiply to 73.5 squared minus the square of the half gap, 0.25, and 73.5 squared is 4,900 + 490 + 12.25. Another quick path: 73 x 74 = 73 squared + 73 = 5,329 + 73.
The relationshipa the round base, 70 b, c the small parts, 3 and 4 What it says in wordsMultiply the round parts, add the round part times the sum of the small parts, then add the small product.Timed tests reward a fixed routine more than cleverness. Pick one decomposition, say the partial products in order, and check with a second route only if time allows. The last digit is a free check: 6 x 6 ends in 6 and 3 x 4 ends in 2, so 3,136 and 5,402 pass. A good habit is to sanity check the size too: 56 squared must sit between 50 squared, 2,500, and 60 squared, 3,600.
Where candidates lose it
The usual slip in 56 squared is adding one strip of 300 instead of two, giving 2,836, or dropping the 36. The area picture makes both errors visible: a square has two strips and a corner.
On a timed screen the other loss is switching methods halfway. Commit to the split, say each partial product, then add. Checking the last digit costs a second and catches most slips.
What the interviewer asks next
- Work out 97 x 103 in your head.
- What is 35 squared, and what is the trick for squares ending in 5?
- Estimate 48 x 52 without multiplying directly.
Asked at Akuna Capital, Prop Trading, Chicago, 2025 (Wall Street Oasis):
The mental math problems which were timed, one example was the 56*56
038Walking up a moving escalator at one step per second you take 20 steps; walking at two steps per second you take 32 steps. How many steps are visible on the escalator?Susquehanna International GroupNew York · 2026
Try it first
How many steps are visible?
Show the worked solution
80 steps. At one step a second the climb takes 20 seconds; at two steps a second it takes 16. If the escalator moves v steps a second, the visible steps are 20 + 20v and also 32 + 16v. Setting them equal gives v = 3, so the escalator is 20 + 60 = 80 steps long, and the check 32 + 48 = 80 agrees.
What stays the same between the two walks?
On an airport moving walkway, walk slowly and the belt does most of the work; stride out and you do more of it yourself, but you reach the end sooner. The length of the walkway does not change. Every visible step is covered either by your legs or by the escalator, so your steps plus the escalator's movement during your climb always equal the same total. That fixed total is the unknown; the escalator's speed is the second unknown, and two walks give two equations.
Walking at one step a second you climb 20 steps in 20 seconds while the escalator carries 60; at two steps a second you climb 32 in 16 seconds while it carries 48; both add to the same 80 visible steps because the escalator moves 3 steps a second. How do you set up and solve the two equations?
Turn step counts into time first, because the escalator's contribution depends on time. The slow walk: 20 steps at one a second is 20 seconds. The fast walk: 32 steps at two a second is 16 seconds. The faster walk loses 4 seconds of escalator help and makes it up with 12 extra steps of its own, so the escalator moves 3 steps a second. Then the total is 20 + 20 x 3 = 80, and 32 + 16 x 3 = 80 confirms it.
The relationshipN visible steps on the escalator v escalator speed, in steps per second 20, 16 seconds taken on the slow and fast walks What it says in wordsThe same number of visible steps is covered on both walks, split differently between you and the machine.Say the check aloud, then the sense check: the escalator at 3 steps a second is faster than either walking pace, which is plausible for a long escalator. If the question had you walking down an up escalator, the escalator's steps would subtract instead of add, and the same method still works. The trap in variants is mixing up steps and seconds; keep one unit for each quantity.
Where candidates lose it
The usual loss is treating the step counts as if they were times, writing 20 + 20v = 32 + 32v, or averaging 20 and 32. The escalator helps for as long as you are on it, and the fast walk is shorter: 16 seconds, not 32.
The second is solving for the speed and stopping. The question asks for the visible steps; plug back in and check both walks give 80.
What the interviewer asks next
- How long does the climb take if you stand still?
- You now walk down the same escalator while it moves up, at 4 steps a second. How many steps do you take?
- A second escalator is twice as fast. How many steps does the slow walker take on it, for the same length?
Asked at Susquehanna International Group, Quantitative Trading, New York, 2026 (Wall Street Oasis):
A stairs question, ask for some physics m/s type of questions
039A surveillance screen flags suspicious trades. One order in 100 is genuinely manipulative. Alert A fires with a likelihood ratio of 9, and an independent alert B with a likelihood ratio of 4. Both fire on the same order: what is the probability it is manipulative?CitadelMiami · 2022
Try it first
Both alerts fire. Roughly how likely is the order manipulative?
Show the worked solution
About 26.7%. Work in odds. The prior odds are 1 to 99. Independent evidence multiplies the odds by each likelihood ratio: 1 x 9 x 4 = 36, so the posterior odds are 36 to 99. As a probability that is 36/135, about 26.7%. Even with both alerts, roughly three flagged orders in four are clean, because manipulation is rare to begin with.
Why is odds form the fast way to combine alerts?
Think of two smoke detectors in a kitchen where real fires are rare. Each beep makes a fire more likely, but toast sets both off far more often than fire does. In odds form, Bayes' rule is one multiplication per piece of independent evidence: posterior odds equal prior odds times each likelihood ratioHow much more often the evidence appears when the hypothesis is true than when it is false.. A ratio of 9 means alert A fires nine times as often on manipulative orders as on clean ones, for example on 90% of manipulative orders and 10% of clean ones.
Starting from odds of 1 to 99, alert A multiplies the odds by 9 to reach 9 to 99, an 8.3% chance, and alert B multiplies by 4 to reach 36 to 99, which is only 26.7% because the prior was so low. How do you check 26.7% by counting?
Take 10,000 orders: 100 manipulative and 9,900 clean. Suppose A fires on 90% of manipulative orders and 10% of clean ones, and B on 80% and 20%, which gives the stated ratios of 9 and 4. Both fire on 100 x 0.9 x 0.8 = 72 manipulative orders and on 9,900 x 0.1 x 0.2 = 198 clean ones. Of the 270 orders where both fire, 72 are manipulative: 26.7%, the same as the odds route.
The relationshipM the order is manipulative 1/99 prior odds: 1 manipulative order per 99 clean 9, 4 likelihood ratios of alerts A and B What it says in wordsMultiply the prior odds by each alert's likelihood ratio, then turn odds back into a probability.State the assumption that made multiplication legal: the alerts are independent given the truth. If both alerts key off the same feature, say order size, the second adds little new information and multiplying by 4 overstates the case. With one alert alone the chance is 8.3% for A and 3.9% for B, which is why a desk reviews orders on combined evidence rather than a single flag.
Where candidates lose it
The common loss is treating a likelihood ratio of 36 as odds of 36 to 1 and answering about 97%. That throws away the base rate: the evidence multiplies the prior odds of 1 to 99, not even odds.
The second is adding the ratios, 9 + 4 = 13, instead of multiplying. Independent evidence compounds, and odds form makes that one line of arithmetic.
What the interviewer asks next
- How many independent alerts with a ratio of 4 would you need to pass 50%?
- Alert B is triggered by the same feature as alert A. How does that change your answer?
- A third alert has a likelihood ratio of 0.5 and does not fire. What does that do?
Asked at Citadel, Sales and Trading, Miami, 2022 (Wall Street Oasis):
I got a question about Bayes' theorem applied to a practical scenario
040Game: roll a fair die and receive its face in rupees; whenever you roll a six you also roll again and add the next result, with no limit on repeats. What is the expected payout of the game?Quant tradingOptions market making
Try it first
What is the game worth?
Show the worked solution
Rs 4.20. Every roll pays its face, averaging 3.5, and with probability 1/6 the game then starts again, worth the same V. So V = 3.5 + V/6, which gives (5/6)V = 3.5 and V = 4.2. A second route agrees: the number of rolls averages 1/(5/6) = 1.2, and each averages 3.5, so 1.2 x 3.5 = 4.2.
Why write the game in terms of itself?
A pass that gets a free renewal each time you use it on a lucky day is worth its first use plus, on lucky days, another pass exactly like it. When a game can repeat with no memory, the value after the repeat is the value of the whole game, so one equation replaces an infinite sum. Here, after a six you are paid 6 and then face precisely the game you started with, worth V.
Faces 1 to 5 end the game, and a six pays 6 and restarts the same game, so V = (5/6) x 3 + (1/6) x (6 + V), which solves to V = 4.2; the number of rolls averages 1.2, and 1.2 x 3.5 gives the same 4.2. How do you set up the equation without slipping?
Condition on the first roll. With probability 5/6 it shows 1 to 5, averaging 3, and you stop. With probability 1/6 it shows 6: you collect 6 and then expect V more. So V = (5/6) x 3 + (1/6) x (6 + V), which simplifies to V = 3.5 + V/6, and V = 3.5 x 6/5 = 4.2. The form 3.5 + V/6 is worth saying: every roll pays 3.5 on average, and one time in six you get another go at the whole game.
The relationshipV expected payout of the game, in rupees 3 average of faces 1 to 5 6 + V payout after a six: the six itself plus a fresh game What it says in wordsThe game's value is one roll's average plus a one-in-six chance of the whole game again.Check by counting rolls. The chance a roll triggers another is 1/6, so the number of rolls averages 1/(1 - 1/6) = 1.2, and each roll averages 3.5 whatever came before, giving 1.2 x 3.5 = 4.2. Summing the series directly, k sixes and then a stop, also lands on 4.2, but the recursion gets there in two lines. The limitation to name if asked to price it: 4.2 is a fair value for one play; the payout has a long right tail, since two sixes in a row, one time in 36, already pay at least 13.
Where candidates lose it
The usual loss is capping the chain: adding one re-roll, 3.5 + 3.5/6 = 4.08, and stopping. The re-roll can be a six too, and the question says there is no limit.
The second is writing V = 3.5 + V without the 1/6, which has no solution, or forgetting the six itself is paid before the restart. Condition on the first roll and write each branch in full.
What the interviewer asks next
- What would you pay to play if a six pays nothing but gives a re-roll?
- What is the probability the payout exceeds 12?
- Now a one ends the game with zero payout. What is the game worth?
