Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
056You roll a fair die until each of 2, 4 and 6 has appeared at least once. Given that the last even number to make its first appearance was 2, what is the probability that the very first roll was a 1? Why is it not 1/5?Squarepoint CapitalLondon · 2026
Try it first
Given that 2 was the last even to show up, what is the chance the first roll was a 1?
Show the worked solution
1/6, the same as with no information. An odd first roll says nothing about the order in which 2, 4 and 6 first appear, so it is independent of 2 finishing last. A first roll of 4 or 6 raises the chance 2 is last from 1/3 to 1/2, so conditioning on that ending shifts weight onto 4 and 6, which rise to 1/4 each. The odd faces keep 1/6 each; 1/5 wrongly spreads the weight evenly.
Why does the ending tell you anything about the start?
Suppose you hear that a friend reached a party last. That makes it a little more likely they left home late, because leaving late and arriving last go together. It says nothing about whether they wore a blue shirt, which has no bearing on arrival order. Conditioning on an outcome reweights every starting state by how likely that state makes the outcome, and a state that does not affect the outcome keeps its original probability. Here the outcome is 2 finishing last among the evens; the question is which first rolls make that more or less likely.
A first roll of 1, 3 or 5 leaves 2 a one in three chance of finishing last, a first roll of 4 or 6 raises it to one in two, and a first roll of 2 makes it impossible, so given that 2 finished last the odd faces are worth 1/6 each and 4 and 6 are worth 1/4 each. How do the numbers work out with Bayes?
Odd rolls never change which new even appears next, so only the order of first appearances matters, and without information it is a random ordering of three: 2 is last with chance 1/3. If the first roll is 4, then 2 and 6 are left to race, and each is equally likely to show first, so 2 ends last with chance 1/2. Now weigh: each face has prior 1/6. The joint chance of first roll 1 and 2 last is 1/6 x 1/3 = 1/18; of first roll 4 and 2 last, 1/6 x 1/2 = 1/12. The total is 1/3, so first roll 1 has posterior (1/18)/(1/3) = 1/6 and first roll 4 has (1/12)/(1/3) = 1/4.
The relationship1/6 the prior chance of any face on the first roll 1/3 the chance 2 is last when the first roll is odd, and also overall 1/2 the chance 2 is last when 4 or 6 is already seen What it says in wordsAn odd first roll is independent of the ending and keeps 1/6; the even faces 4 and 6 absorb the weight that 2 loses.Where does the 1/5 intuition go wrong?
It treats the information as simply ruling out one face and renormalising the rest. Ruling out an outcome and conditioning on an event are the same thing only when every remaining outcome makes the event equally likely, and here they do not. A check: the posteriors 1/6, 1/6, 1/6, 1/4, 1/4 and 0 add to 1, while five faces at 1/5 would give 4 and 6 the same weight as 1. On a desk this is the error of reading a trade's outcome as if it said nothing about which signal triggered it.
Where candidates lose it
Nearly everyone's first answer is 1/5. The interviewer is not testing the arithmetic; the question itself says it is not 1/5 and asks you to explain why, so an answer that only produces 1/6 without the reason loses most of the credit.
The second trap is getting lost in the odd rolls. They can be ignored completely, because they never change which even appears next. Say that early and the problem shrinks to the order of three numbers.
What the interviewer asks next
- Given that 2 finished last, what is the probability the first roll was a 4?
- What is the expected number of rolls until all three evens have appeared?
- Given that 2 finished last, what is the probability the first even to appear was 4?
Asked at Squarepoint Capital, Quant Research Intern Interview, London, 2026 (Wall Street Oasis):
why is the probability of seeing a 1 on our first roll, given that we end on a 2, not 1/5
087A new strategy has won on 9 of its first 10 trading days. With no prior knowledge, you treat its daily win rate as uniform between 0 and 1. What is the probability that it wins tomorrow, and why is the answer not 0.9?Wolverine TradingChicago · 2017
Try it first
What probability do you give to a win tomorrow?
Show the worked solution
10/12, about 0.833. A uniform prior on the win rate, updated with 9 wins and 1 loss, gives a Beta(10, 2) posterior. The chance of winning tomorrow is the posterior mean, (9 + 1)/(10 + 2). The answer is below 0.9 because ten days cannot rule out a lower true win rate, and averaging over that uncertainty pulls the estimate towards one half.
Why is 0.9 too confident?
A new restaurant with nine five-star reviews out of ten looks excellent, but you would not bet that the next diner rates it five stars with 90% certainty. Ten reviews is a small sample, and a restaurant that truly earns five stars 70% of the time could easily post nine out of ten. The right forecast averages over every win rate the evidence still allows, and with only ten days that includes plenty of rates below 0.9. The average of that spread of possibilities is what you should quote for tomorrow.
After 9 wins in 10 days from a flat prior, the win rate follows a Beta(10, 2) curve that peaks at 0.9 but has a long left tail, so its mean, the chance of winning tomorrow, is 10/12, about 0.833, and its middle 90% still spans 0.64 to 0.97. How does the Bayesian update give 10/12?
Start with every win rate p between 0 and 1 equally likely. The chance of the observed record is proportional to p^9 (1 - p), so the posterior is proportional to that, which is the Beta(10, 2) distribution. The chance of a win tomorrow is the average of p over the posterior, and the mean of a Beta(a, b) is a/(a + b), here 10/12. The shortcut is Laplace's rule of succession: add one imaginary win and one imaginary loss to the record, then divide.
The relationshipp the unknown daily win rate p^9 (1 - p) the likelihood of 9 wins and 1 loss Beta(10, 2) the posterior after a flat prior What it says in wordsMultiply the flat prior by the likelihood of the record, and the mean of the result is the chance of a win tomorrow.When does the pull towards one half stop mattering?
When the record is long. At 90 wins out of 100 the rule gives 91/102, about 0.892, almost exactly the raw 0.9, because a hundred days of data swamp the one imaginary win and loss. The limitation is the prior itself. A uniform prior says a 99% win rate was as plausible as a 50% one before you saw any data, which no trader believes about a new strategy. With a sceptical prior centred near one half, ten days would pull the estimate even further below 0.9. The posterior also tells you more than one number: the chance of winning both of the next two days is (10 x 11)/(12 x 13), about 0.705, not 0.833 squared.
Where candidates lose it
The trap is answering 0.9, the maximum likelihood estimate. It treats the observed win rate as the truth and ignores how little ten days can tell you.
The second loss is getting 10/12 by the rule of succession without being able to say where it comes from. Name the flat prior, the Beta(10, 2) posterior and its mean, and say that more data pushes the answer back towards 0.9.
What the interviewer asks next
- What is the probability the strategy's true win rate is above one half?
- With a Beta(5, 5) prior instead of a flat one, what is your forecast for tomorrow?
- How many consecutive wins would you need before your forecast exceeds 0.95?
Asked at Wolverine Trading, Prop Trading, Chicago, 2017 (Wall Street Oasis):
Phone interviews were pretty standard brainteasers and fit questions. There was a Bayes question
