Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
020A broad stock index has returned an average of 11% a year over the past 30 years, with an annual standard deviation of 16% (illustrative figures). Assuming yearly returns are independent, give a 95% confidence interval for its true expected annual return.Old Mission CapitalChicago · 2025
Try it first
Roughly how wide is the 95% interval?
Show the worked solution
About 5.3% to 16.7%. The standard error of a 30-year average is 16% divided by the square root of 30, 2.92 points. A 95% interval is 1.96 standard errors either side: 11% plus or minus 5.7. With a t value for 29 degrees of freedom it is a touch wider, plus or minus 6.0. Thirty years of data still leave the expected return very uncertain.
Why is the band so wide after thirty years?
Imagine judging a new cricketer's true batting average from a handful of innings. Scores swing wildly from one innings to the next, so a few innings tell you little. The precision of an average grows only with the square root of the number of observations, and yearly returns swing by 16 points, so 30 years give a standard error of about 2.9 points. That is large next to an average of 11%: the true figure could plausibly be 6% or 16%.
Around an 11% average with 16% annual volatility, the 95% interval is plus or minus 9.9 points with 10 years of data, 5.7 points with 30 years and still 2.9 points with 120 years, because precision grows only with the square root of the sample. How do you set it up in the room?
State the assumption, then compute. Treat the 30 annual returns as independent draws with a standard deviation of 16%, so the sample average has a standard error of 16 over root 30. Root 30 is about 5.48, so the standard error is 2.92. Multiply by 1.96: 5.73. The interval is 5.3% to 16.7%. If the interviewer pushes, note that with only 30 observations a t distributionThe distribution used for an average when the standard deviation is estimated from the same small sample; it has fatter tails than the normal. with 29 degrees of freedom gives a multiplier of about 2.045, widening the band slightly.
The relationship\bar r the average annual return, 11% \sigma the annual standard deviation, 16% n the number of yearly observations, 30 What it says in wordsThe interval is the average plus or minus 1.96 standard errors, and the standard error is volatility over the square root of years.What limitation should you add?
Halving the band needs four times the data: 120 years still only pins the average to within about 2.9 points, and markets change over such spans. Returns are also not quite independent from year to year, fat tails make the normal multiplier optimistic, and the arithmetic average overstates the compound growth rate. The practical lesson is that expected returns are the hardest input in finance to estimate, far harder than volatility, which is why many quant processes lean on risk estimates and treat return forecasts with caution.
Where candidates lose it
The common mistake is using 16% as the width, the spread of a single year, rather than the standard error of the average. That gives an interval from -20% to 42%, which describes one year's return, not the long-run mean.
The opposite slip is dividing by 30 instead of the square root of 30, which gives a band of about 1 point and badly overstates what the data can tell you.
What the interviewer asks next
- How many years of data would you need to pin the mean to within plus or minus 1 point?
- How would monthly data change the interval for the mean?
- Why is volatility much easier to estimate than the mean from the same data?
Asked at Old Mission Capital, Prop Trading, Chicago, 2025 (Wall Street Oasis):
Confidence interval on S&P 500 return past 30 years
024Differentiate f(x) = x to the power x, and find where it reaches its minimum for positive x.ScotiabankToronto · 2026
Try it first
What is the derivative of x to the x?
Show the worked solution
f'(x) = x to the x times (ln x + 1), and the minimum is at x = 1/e, about 0.368, where f is about 0.692. Take logs: ln f = x ln x. Differentiating, f'/f = ln x + 1, so f' = x to the x (ln x + 1). The derivative is zero when ln x = -1, that is x = 1/e, negative before it and positive after, so this is a minimum.
Why do both standard rules fail?
The power rule, n x to the (n - 1), treats the exponent as fixed; the exponential rule, a to the x times ln a, treats the base as fixed. In x to the x both the base and the exponent move, so neither rule applies on its own, and each one gives half of the right answer. Indeed the correct derivative is the sum of the two: x times x to the (x - 1), which is x to the x, plus x to the x ln x. That sum is a quick check on your final answer.
The curve x to the x falls from near 1 at zero to a minimum of 0.692 at x = 1/e and then rises to 4 at x = 2, and its logarithm x ln x has its minimum at the same point, which is why taking logs first is safe. How does taking logs make it routine?
Think of converting a messy multiplication into addition before doing it, the way a slide rule does. Write ln f = x ln x; the right-hand side is a product of two simple functions, and the product rule gives ln x + x times 1/x = ln x + 1. The left side differentiates to f'/f by the chain rule, so f' = f (ln x + 1). This is logarithmic differentiationDifferentiating the logarithm of a function instead of the function, then multiplying back; useful when the variable sits in an exponent., and it works for any function of the form g(x) to the h(x).
The relationshipe^{x ln x} x to the x rewritten with a fixed base ln x + 1 the derivative of x ln x e^{-1} where ln x = -1, the minimum What it says in wordsRewrite with base e, differentiate the exponent, and set it to zero.How do you confirm it is a minimum and state the value?
Check the sign of ln x + 1, since x to the x is always positive. For x below 1/e, ln x is below -1 and the slope is negative; above 1/e it is positive, so the function falls and then rises: a minimum. The value is (1/e) to the (1/e) = e to the (-1/e), about 0.692. Also say what happens at the edges: as x shrinks towards zero, x ln x tends to zero, so x to the x tends to 1, and at x = 1 it is exactly 1 again.
Where candidates lose it
The fast wrong answer applies the power rule, x times x to the (x - 1), which is just x to the x. It treats the exponent as a constant, and candidates who give it usually do so in the first three seconds.
The second loss is finding x = 1/e and stopping. The question asks for the minimum, so check the sign change and give the value, e to the (-1/e), about 0.692, together with the behaviour near zero.
What the interviewer asks next
- Differentiate x to the (x to the x).
- What is the limit of x to the x as x approaches 0 from above, and why?
- Which is larger, e to the pi or pi to the e, and how does x to the (1/x) settle it?
Asked at Scotiabank, Quant, Toronto, 2026 (Wall Street Oasis):
technical questions covering calculus (including derivatives of standard functions)
025Four people queue at a cash machine wanting 7, 3, 10 and 2 thousand rupees. Each visit allows at most 4 thousand, and anyone who has not got their full amount rejoins the back of the queue. In what order do they leave, and how would you compute the order quickly for a very long queue?Squarepoint CapitalLondon · 2026
Try it first
In what order do the four people leave?
Show the worked solution
They leave in the order 2, 4, 1, 3. Person 1 takes 4 and rejoins with 3; person 2 takes 3 and leaves; person 3 takes 4 and rejoins with 6; person 4 takes 2 and leaves. Then person 1 takes 3 and leaves, and person 3 needs two more visits. The shortcut: each person leaves in round amount divided by 4, rounded up, and ties go to whoever stood first.
How do you simulate it cleanly?
Model the line as a queueA first-in, first-out list: items join at the back and leave from the front, as in a real line. of pairs, person and amount still wanted. Pop the front, subtract the lesser of the cap and what they still want, and if anything is left push them onto the back; otherwise record them as leaving. The four people take seven visits in all. This is the answer most interviewers expect first, and it is correct, but its cost grows with the total number of visits, which is the sum of each amount over the cap.
Seven visits clear the queue: persons 2 and 4 leave on their first visit, person 1 on the second round and person 3 on the third, so the exit order is 2, 4, 1, 3, matching each person's amount divided by 4, rounded up. Is there a faster way than simulating?
Yes. Think of a canteen that serves one plate per person per pass: someone wanting three plates leaves on the third pass, whatever the others want. Person i leaves in round ceiling(a_i / k), and within a round the queue keeps its original order, so the exit order is simply the people sorted by their round number, ties broken by starting position. Here the rounds are 2, 1, 3 and 1, which sorts to 2, 4, 1, 3. That costs n log n, however large the amounts are, instead of the number of visits.
The relationshipa_i the amount person i wants k the cap per visit, 4 r_i the round in which person i leaves What it says in wordsSort people by how many rounds they need, and by queue position within a round.Why does queue order survive between rounds?
Everyone still waiting after a round rejoins in the same relative order they were served, because the queue is first in, first out. So round two serves the survivors of round one in their original order, and so on. That invariant is what lets you replace the simulation with a sort, and saying it out loud is what separates an answer that works from one you can defend. If a very large cap or tiny amounts made most people finish in round one, the sort still costs n log n, and a counting sort on round numbers can make it linear.
Where candidates lose it
The common wrong answer sorts by amount: 4, 2, 1, 3. It ignores that people who finish in the same round leave in queue order, and person 2 stands ahead of person 4.
The second loss is stopping at the simulation when the question asks how to do it quickly. With amounts in the crores and a small cap, simulating each visit could take billions of steps. The ceiling formula plus a stable sort is the answer to the second half.
What the interviewer asks next
- Return the time at which each person leaves if each visit takes one minute.
- What if the cap differs by visit, for example 4 thousand on odd visits and 2 on even ones?
- Implement the sort-based version and state its complexity.
Asked at Squarepoint Capital, Quant Research Intern Interview, London, 2026 (Wall Street Oasis):
returning the order in which people leave a queue given a list of amounts people want to withdraw from an ATM
027A bag holds three dice: one fair, one that shows six half the time with its other faces equally likely, and one that never shows six. You draw one at random and roll it twice, getting two sixes. What is the probability it is the loaded die?Belvedere TradingChicago · 2022
Try it first
Before you calculate: how likely is it now that you hold the loaded die?
Show the worked solution
90%. Each die starts at one in three. The chance of two sixes is 1/36 for the fair die, 1/4 for the loaded die and zero for the die with no six. Weight each by its prior: 1/108 for the fair die, 1/12 for the loaded die, nothing for the third. The loaded die's weight is nine times the fair die's, so its probability is 9/10.
What does the die that never shows six do to the answer?
A neighbour tells you a red car blocked the gate this morning. If one of your suspects owns only a blue scooter, that suspect is out, however likely they looked before. A hypothesis that cannot produce the evidence gets zero weight afterwards, no matter what its prior was. The no-six die could never give two sixes, so it drops out, and the question becomes a contest between the fair die and the loaded die, which started level at one third each.
Each die starts at one third; multiplying by the chance of two sixes gives weights of 1/108 for the fair die, 1/12 for the loaded die and zero for the no-six die, so the loaded die ends at 90% and the fair die at 10%. How much does each surviving die's likelihood count?
Now compare how easily each remaining die produces what you saw. The fair die gives two sixes 1 time in 36. The loaded die gives a six half the time, so two in a row 1 time in 4. With equal priors, the posterior odds are just the ratio of the likelihoodsThe probability of the observed evidence under each hypothesis, before any prior is applied.: 1/4 against 1/36, which is 9 to 1. Nine parts in ten is 90%.
The relationshipL the loaded die was drawn 66 the evidence: two sixes in two rolls 1/3 the prior for each die 1/36, 1/4, 0 the chance of two sixes from the fair, loaded and no-six dice What it says in wordsThe loaded die's share of all the ways two sixes can happen is nine tenths.Check it by counting, the safer habit under pressure. Imagine 108 rounds of drawing a die and rolling it twice, 36 rounds with each die. The fair die gives two sixes once, the loaded die 9 times, the no-six die never. Of the 10 double sixes, 9 came from the loaded die, and the prior of one third cancels because every die got the same number of rounds. The loaded die's other faces, 1 in 10 each, never enter, because only sixes were seen.
Where candidates lose it
The quick wrong answer is 1/2: two dice can roll a six, so it must be one or the other. That ignores how differently they produce two sixes in a row, a gap of nine to one.
The other loss is getting tangled in the loaded die's other faces, or leaving the no-six die in the denominator with some weight. Neither belongs: only the chance of the observed rolls counts, and for the no-six die that chance is zero.
What the interviewer asks next
- A third roll is also a six. What is the probability of the loaded die now? (It rises to 27/28.)
- The rolls were a six and then a two. Which die is most likely now?
- What is the chance the next roll is a six? (It is 7/15.)
Asked at Belvedere Trading, Capital Markets, Chicago, 2022 (Wall Street Oasis):
The technical portion of the interview consisted of probability questions including one questions relating to Bayes' theorem
028You roll two fair dice and are paid the larger of the two faces in rupees. What is the expected payout?Jane StreetNew York · 2026
Try it first
Pick the expected payout before you count.
Show the worked solution
161/36, about Rs 4.47. The larger face equals k in 2k minus 1 of the 36 equally likely outcomes: 1, 3, 5, 7, 9 and 11 cells for k from 1 to 6. Multiply each value by its count and add: 1 + 6 + 15 + 28 + 45 + 66 = 161. Divided by 36, that is 4.47, almost a full point above a single die's 3.5.
Why is the answer well above 3.5?
When two friends each suggest a restaurant and you always go with the better rated one, your average dinner beats either friend's average. Taking the larger of two draws pulls the result toward the top, because a low result survives only if both draws are low. A payout of 1 needs both dice on 1, one cell in 36. A payout of 6 needs just one six, and 11 cells in 36 contain at least one.
The larger face is 1 in one cell, 2 in three cells and so on up to 6 in eleven cells; face times count sums to 161, so the expected payout is 161/36, about 4.47, against 3.50 for one die. How do you count the cells without listing all 36?
Count the outcomes where the larger face is at most k: both dice must be at most k, which is k squared cells. The cells where the larger face is exactly k are k squared minus (k - 1) squared, which is 2k - 1. That is the L-shaped band in the grid: a new row and a new column, sharing one corner cell. The bands are 1, 3, 5, 7, 9 and 11, and they add to 36, which is the check that nothing was double counted.
The relationshipk the value of the larger face 2k - 1 the number of the 36 outcomes where the larger face is exactly k What it says in wordsWeight each possible payout by how many of the 36 outcomes produce it, then divide by 36.A second route helps when the interviewer changes the dice. Add up the chance that the payout reaches each level: the payout is at least k unless both dice are below k, so the sum of 1 minus (k - 1) squared over 36, for k from 1 to 6, is 6 minus 55/36, which is 161/36 again. Two methods landing on the same fraction is the check worth saying out loud. By symmetry the smaller face averages 7 minus 4.47, about 2.53, and with three dice the larger face rises to 4.96.
Where candidates lose it
The common slip is to treat the six payouts as equally likely and answer 3.5, or to say a bit more than 3.5 without a number. The grid shows how uneven the counts are: eleven ways to be paid 6 against one way to be paid 1.
The second slip is counting 12 cells for a payout of 6, which counts the double six twice. The row of sixes and the column of sixes share one cell.
What the interviewer asks next
- What is the expected value of the smaller face?
- What is the expected larger face with three dice?
- I pay you the larger face minus the smaller. What is that worth?
Asked at Jane Street, Investment Operations, New York, 2026 (Wall Street Oasis):
First interview was testing simple math brainteasers (e.g. expected value of dice throws, etc.)
035Regressing y on x gives a slope of 0.8; regressing x on y gives a slope of 0.45. What is the R-squared of either regression, and what is the correlation?Tower Research CapitalNew York · 2014
Try it first
What is the correlation between x and y?
Show the worked solution
R-squared is 0.36 for both regressions and the correlation is 0.6. The slope of y on x is r times sd(y)/sd(x); the slope of x on y is r times sd(x)/sd(y). Multiplying them cancels the standard deviations and leaves r squared: 0.8 x 0.45 = 0.36. The correlation is +0.6, positive because both slopes are positive, and the ratio sd(y)/sd(x) is √(0.8/0.45) = 4/3.
Why are the two slopes not reciprocals of each other?
Tall parents tend to have tall children, but a little less tall; and tall children tend to have tall parents, but a little less tall. Both statements are true at once. Each regression predicts toward the mean, so neither slope is the inverse of the other unless the fit is perfect. If the points lay exactly on a line, the slope of x on y would be 1/0.8 = 1.25. It is 0.45 instead, and the size of that shortfall is what measures how loose the relationship is.
Fitting y on x gives the shallower line with slope 0.8 and fitting x on y gives the steeper line, slope 0.45 in its own terms; their product, 0.36, is R-squared, so the correlation is 0.6 and the standard deviation of y is 4/3 that of x. How do the two slopes give R-squared?
Write each slope in terms of the correlation. The least squares slope of y on x is the covariance over the variance of x, which is r times sd(y)/sd(x). Swap the roles and the slope of x on y is r times sd(x)/sd(y). The standard deviation ratios are reciprocals, so the product of the two slopes is r squared, and in a one-variable regression r squared is exactly the R-squared. Here 0.8 x 0.45 = 0.36, so r = 0.6; the sign is positive because both slopes are positive, and the two slopes always share a sign.
The relationshipb_y|x slope from regressing y on x, 0.8 b_x|y slope from regressing x on y, 0.45 s_x, s_y standard deviations of x and y r the correlation of x and y What it says in wordsThe two slopes multiply to the squared correlation because the scale factors cancel.The figure uses 40 points built with standard deviations 3 and 4 and a correlation of exactly 0.6, and fitting both regressions returns slopes of 0.80 and 0.45. A quick sanity test comes free: the product of the two slopes can never exceed 1. If an interviewer quotes slopes of 0.8 and 1.5, the product 1.2 is impossible, and saying so is worth more than any calculation.
Where candidates lose it
The fast wrong answer is to say the slopes should be reciprocals and call the data inconsistent, or to answer 0.36 when asked for the correlation. 0.36 is R-squared; the correlation is its square root.
The second loss is dropping the sign. The square root of 0.36 could be plus or minus 0.6; both slopes are positive, so the correlation is positive, and saying why takes one sentence.
What the interviewer asks next
- What is the ratio of the standard deviation of y to that of x?
- If the slope of x on y were 1.5, what would you conclude?
- How does adding measurement noise to x change each slope?
Asked at Tower Research Capital, Quantitative Research, New York, 2014 (Wall Street Oasis):
Another detailed linear regression questions were asked, including problems about residual, variance and R^2
036A stock pays a growing dividend and is valued with the Gordon model at a discount rate of 10% and growth of 6%. What is its duration, and roughly how much does its price change if the discount rate rises by one point?BlackRockNew York · 2026
Try it first
What is the stock's duration, its percentage price sensitivity to the discount rate?
Show the worked solution
Duration is 1/(r - g) = 25 years, so a one-point rise cuts the value by about a fifth. The Gordon price is D1/(r - g), and its percentage sensitivity to r is 1/(r - g) = 1/0.04 = 25. With a Rs 4 dividend the price moves from Rs 100 at 10% to Rs 80 at 11%, a 20% fall. The 25% duration estimate overshoots because the price curve is convex.
Why does a stock have a duration at all?
A promise of money in one year hardly changes in value when rates move; a promise of money in twenty five years changes a lot, because the rate is compounded over every one of those years. A stock is a stream of dividends stretching forever, and when the dividends grow, most of its value sits in cash flows far in the future, so it behaves like a very long bond. Duration measures exactly that: the percentage price change for a change in the discount rate.
The relationshipD1 next year's dividend, Rs 4 in the illustration r discount rate, 10% g dividend growth rate, 6% 1/(r - g) percentage price change per unit change in r What it says in wordsDifferentiate the Gordon price and divide by price: the sensitivity is one over the gap between the discount rate and growth.At a 10% discount rate the Rs 4 dividend stock is worth Rs 100 and its duration is 25; at 11% the price is Rs 80, a 20% fall, while the tangent line predicts Rs 75, and at 12% the gap widens to Rs 66.7 against Rs 50 because the price curve is convex. Why does the estimate say 25% when the price falls 20%?
Duration is the slope at one point, and the price curve bends. For a one-point rise the tangent predicts a 25% fall, but the exact move from Rs 100 to Rs 80 is 20%, because the curve is convexCurving upward, so it always sits above any of its tangent lines. and flattens as r rises. The same bend makes a one-point fall worth more than 25%: at 9% the price is Rs 133.3, up 33.3%. For a big rate move, reprice exactly instead of trusting the slope. Duration here also equals price over dividend, 100/4, which is a quick way to say it: one over the dividend yield.
For precision, the Macaulay durationThe present-value weighted average time at which cash flows arrive. is (1 + r)/(r - g) = 27.5 years, and dividing by 1 + r gives the modified duration of 25. A 10-year bond paying 10% at a 10% yield has a modified duration of about 6.1. The gap r - g is what matters, which is why high-growth stocks carry the most duration: the same stock with 2% growth would have a duration of 12.5 years. The limitation is that the Gordon model holds growth fixed while rates move; in practice both shift together.
Where candidates lose it
The usual loss is saying a stock has no duration because it has no maturity, or that it is infinite because it pays forever. Both skip the one line of calculus that gives 1/(r - g).
The second is quoting 25% as the exact price change. It is the slope at 10%; the exact fall to 11% is 20%, and saying why, convexity, is what separates a strong answer.
What the interviewer asks next
- What happens to duration as growth approaches the discount rate?
- Why might a stock's measured sensitivity to bond yields be much lower than 25?
- What is the price change for a one-point fall in the discount rate?
Asked at BlackRock, Restructuring, New York, 2026 (Wall Street Oasis):
Which equities have duration ? multiple stocks vs value stocks MSE Forecasting equation
038Walking up a moving escalator at one step per second you take 20 steps; walking at two steps per second you take 32 steps. How many steps are visible on the escalator?Susquehanna International GroupNew York · 2026
Try it first
How many steps are visible?
Show the worked solution
80 steps. At one step a second the climb takes 20 seconds; at two steps a second it takes 16. If the escalator moves v steps a second, the visible steps are 20 + 20v and also 32 + 16v. Setting them equal gives v = 3, so the escalator is 20 + 60 = 80 steps long, and the check 32 + 48 = 80 agrees.
What stays the same between the two walks?
On an airport moving walkway, walk slowly and the belt does most of the work; stride out and you do more of it yourself, but you reach the end sooner. The length of the walkway does not change. Every visible step is covered either by your legs or by the escalator, so your steps plus the escalator's movement during your climb always equal the same total. That fixed total is the unknown; the escalator's speed is the second unknown, and two walks give two equations.
Walking at one step a second you climb 20 steps in 20 seconds while the escalator carries 60; at two steps a second you climb 32 in 16 seconds while it carries 48; both add to the same 80 visible steps because the escalator moves 3 steps a second. How do you set up and solve the two equations?
Turn step counts into time first, because the escalator's contribution depends on time. The slow walk: 20 steps at one a second is 20 seconds. The fast walk: 32 steps at two a second is 16 seconds. The faster walk loses 4 seconds of escalator help and makes it up with 12 extra steps of its own, so the escalator moves 3 steps a second. Then the total is 20 + 20 x 3 = 80, and 32 + 16 x 3 = 80 confirms it.
The relationshipN visible steps on the escalator v escalator speed, in steps per second 20, 16 seconds taken on the slow and fast walks What it says in wordsThe same number of visible steps is covered on both walks, split differently between you and the machine.Say the check aloud, then the sense check: the escalator at 3 steps a second is faster than either walking pace, which is plausible for a long escalator. If the question had you walking down an up escalator, the escalator's steps would subtract instead of add, and the same method still works. The trap in variants is mixing up steps and seconds; keep one unit for each quantity.
Where candidates lose it
The usual loss is treating the step counts as if they were times, writing 20 + 20v = 32 + 32v, or averaging 20 and 32. The escalator helps for as long as you are on it, and the fast walk is shorter: 16 seconds, not 32.
The second is solving for the speed and stopping. The question asks for the visible steps; plug back in and check both walks give 80.
What the interviewer asks next
- How long does the climb take if you stand still?
- You now walk down the same escalator while it moves up, at 4 steps a second. How many steps do you take?
- A second escalator is twice as fast. How many steps does the slow walker take on it, for the same length?
Asked at Susquehanna International Group, Quantitative Trading, New York, 2026 (Wall Street Oasis):
A stairs question, ask for some physics m/s type of questions
039A surveillance screen flags suspicious trades. One order in 100 is genuinely manipulative. Alert A fires with a likelihood ratio of 9, and an independent alert B with a likelihood ratio of 4. Both fire on the same order: what is the probability it is manipulative?CitadelMiami · 2022
Try it first
Both alerts fire. Roughly how likely is the order manipulative?
Show the worked solution
About 26.7%. Work in odds. The prior odds are 1 to 99. Independent evidence multiplies the odds by each likelihood ratio: 1 x 9 x 4 = 36, so the posterior odds are 36 to 99. As a probability that is 36/135, about 26.7%. Even with both alerts, roughly three flagged orders in four are clean, because manipulation is rare to begin with.
Why is odds form the fast way to combine alerts?
Think of two smoke detectors in a kitchen where real fires are rare. Each beep makes a fire more likely, but toast sets both off far more often than fire does. In odds form, Bayes' rule is one multiplication per piece of independent evidence: posterior odds equal prior odds times each likelihood ratioHow much more often the evidence appears when the hypothesis is true than when it is false.. A ratio of 9 means alert A fires nine times as often on manipulative orders as on clean ones, for example on 90% of manipulative orders and 10% of clean ones.
Starting from odds of 1 to 99, alert A multiplies the odds by 9 to reach 9 to 99, an 8.3% chance, and alert B multiplies by 4 to reach 36 to 99, which is only 26.7% because the prior was so low. How do you check 26.7% by counting?
Take 10,000 orders: 100 manipulative and 9,900 clean. Suppose A fires on 90% of manipulative orders and 10% of clean ones, and B on 80% and 20%, which gives the stated ratios of 9 and 4. Both fire on 100 x 0.9 x 0.8 = 72 manipulative orders and on 9,900 x 0.1 x 0.2 = 198 clean ones. Of the 270 orders where both fire, 72 are manipulative: 26.7%, the same as the odds route.
The relationshipM the order is manipulative 1/99 prior odds: 1 manipulative order per 99 clean 9, 4 likelihood ratios of alerts A and B What it says in wordsMultiply the prior odds by each alert's likelihood ratio, then turn odds back into a probability.State the assumption that made multiplication legal: the alerts are independent given the truth. If both alerts key off the same feature, say order size, the second adds little new information and multiplying by 4 overstates the case. With one alert alone the chance is 8.3% for A and 3.9% for B, which is why a desk reviews orders on combined evidence rather than a single flag.
Where candidates lose it
The common loss is treating a likelihood ratio of 36 as odds of 36 to 1 and answering about 97%. That throws away the base rate: the evidence multiplies the prior odds of 1 to 99, not even odds.
The second is adding the ratios, 9 + 4 = 13, instead of multiplying. Independent evidence compounds, and odds form makes that one line of arithmetic.
What the interviewer asks next
- How many independent alerts with a ratio of 4 would you need to pass 50%?
- Alert B is triggered by the same feature as alert A. How does that change your answer?
- A third alert has a likelihood ratio of 0.5 and does not fire. What does that do?
Asked at Citadel, Sales and Trading, Miami, 2022 (Wall Street Oasis):
I got a question about Bayes' theorem applied to a practical scenario
040Game: roll a fair die and receive its face in rupees; whenever you roll a six you also roll again and add the next result, with no limit on repeats. What is the expected payout of the game?Quant tradingOptions market making
Try it first
What is the game worth?
Show the worked solution
Rs 4.20. Every roll pays its face, averaging 3.5, and with probability 1/6 the game then starts again, worth the same V. So V = 3.5 + V/6, which gives (5/6)V = 3.5 and V = 4.2. A second route agrees: the number of rolls averages 1/(5/6) = 1.2, and each averages 3.5, so 1.2 x 3.5 = 4.2.
Why write the game in terms of itself?
A pass that gets a free renewal each time you use it on a lucky day is worth its first use plus, on lucky days, another pass exactly like it. When a game can repeat with no memory, the value after the repeat is the value of the whole game, so one equation replaces an infinite sum. Here, after a six you are paid 6 and then face precisely the game you started with, worth V.
Faces 1 to 5 end the game, and a six pays 6 and restarts the same game, so V = (5/6) x 3 + (1/6) x (6 + V), which solves to V = 4.2; the number of rolls averages 1.2, and 1.2 x 3.5 gives the same 4.2. How do you set up the equation without slipping?
Condition on the first roll. With probability 5/6 it shows 1 to 5, averaging 3, and you stop. With probability 1/6 it shows 6: you collect 6 and then expect V more. So V = (5/6) x 3 + (1/6) x (6 + V), which simplifies to V = 3.5 + V/6, and V = 3.5 x 6/5 = 4.2. The form 3.5 + V/6 is worth saying: every roll pays 3.5 on average, and one time in six you get another go at the whole game.
The relationshipV expected payout of the game, in rupees 3 average of faces 1 to 5 6 + V payout after a six: the six itself plus a fresh game What it says in wordsThe game's value is one roll's average plus a one-in-six chance of the whole game again.Check by counting rolls. The chance a roll triggers another is 1/6, so the number of rolls averages 1/(1 - 1/6) = 1.2, and each roll averages 3.5 whatever came before, giving 1.2 x 3.5 = 4.2. Summing the series directly, k sixes and then a stop, also lands on 4.2, but the recursion gets there in two lines. The limitation to name if asked to price it: 4.2 is a fair value for one play; the payout has a long right tail, since two sixes in a row, one time in 36, already pay at least 13.
Where candidates lose it
The usual loss is capping the chain: adding one re-roll, 3.5 + 3.5/6 = 4.08, and stopping. The re-roll can be a six too, and the question says there is no limit.
The second is writing V = 3.5 + V without the 1/6, which has no solution, or forgetting the six itself is paid before the restart. Condition on the first roll and write each branch in full.
What the interviewer asks next
- What would you pay to play if a six pays nothing but gives a re-roll?
- What is the probability the payout exceeds 12?
- Now a one ends the game with zero payout. What is the game worth?

