Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
003Speed round, ninety seconds: you draw three cards from a well-shuffled 52-card deck without replacement. What is the probability that all three are of different suits?Quant tradingProp trading firms
Try it first
Closest answer, fast.
Show the worked solution
About 39.8%. The first card can be anything. The second must come from one of the three other suits: 39 of the 51 cards left. The third must avoid both suits already seen: 26 of the 50 left. Multiply: 39/51 x 26/50 = 1,014/2,550, just under 40%. Drawing with replacement would give 3/4 x 1/2 = 37.5%.
Why does the first card cost nothing?
Picture three guests arriving at a party with four dress colours in the wardrobe, and you want no two to match. The first guest cannot clash with anyone. A condition about the cards differing only bites from the second card on, so the first card contributes a factor of 1 and you start counting from card two. Candidates who write 13/52 for the first card have fixed a particular suit and then have to multiply by the number of suit orders to recover, which is where the slips happen.
The first card is free, the second must avoid one used suit with 39 of 51 cards still good, and the third must avoid two with 26 of 50 still good, so three different suits happen with probability 39.8%. How do you check it a second way in the time?
Count unordered hands. Choose which three suits appear, 4 ways, then one card from each, 13 cubed, and divide by all three-card hands, 52 choose 3. That is 4 x 2,197 = 8,788 over 22,100, which is the same 0.3976. In a speed round you do not have time for both, but knowing the counting route exists lets you sanity-check the product: 0.765 x 0.52 is a little under 0.40.
The relationship39/51 cards of a new suit among those left after one draw 26/50 cards of a third suit after two draws \binom{52}{3} the number of possible three-card hands What it says in wordsSequential dodging and direct counting give the same 39.8%.What is a speed round actually testing?
Thirty questions in forty-five minutes cannot all be worked in full. The skill being tested is choosing the shortest correct route and estimating the product well enough to pick from the options. Here, 39/51 is about 0.76 and 26/50 is 0.52; 0.76 x 0.52 is about 0.40, which eliminates every other option before you finish the exact fraction.
Where candidates lose it
The fast wrong answer is 37.5%, from treating the draws as if cards go back in the deck. It feels close enough, and in a multiple choice round it sits right next to the correct option on purpose.
The other slip is starting with 13/52 for the first card, which silently fixes that card's suit. You then need to multiply by 4 for the suit choice, and under time pressure most people forget.
What the interviewer asks next
- What is the probability that four cards are all of different suits?
- What is the probability that three cards share a suit?
- Draw until you have seen all four suits. What is the expected number of cards?
007You roll a fair six-sided die six times. What is the expected number of different faces that appear?Quant tradingProp trading firms
Try it first
Your estimate before any working?
Show the worked solution
About 3.99. Give each face an indicator that equals 1 if that face appears at least once. A given face is missed on all six rolls with probability (5/6) to the 6th, about 0.335, so it appears with probability 0.665. The expected count of distinct faces is the sum of the six indicators' expectations: 6 x 0.665 = 3.99. No case listing is needed.
Why not list the cases?
You could work out the chance of exactly one, two, up to six distinct faces and average them. It works, but it needs Stirling numbersCounts of the ways to split a set of items into a given number of non-empty groups; they appear when counting surjections. or a lot of careful counting, and it is easy to slip. Linearity of expectation lets you ignore how the faces interact: the expected total of several indicators is the sum of their expectations, whether or not they are independent. Here the six indicators are clearly dependent, since seeing many faces leaves fewer rolls for the others, and it does not matter at all.
Each of the six faces appears at least once with probability 66.5%, so the expected number of distinct faces is six times that, 3.99; the exact distribution peaks at four distinct faces and has all six only 1.5% of the time. How does the indicator trick work step by step?
Think of a teacher counting how many of six friends turn up to a party. Instead of listing every guest list, she asks for each friend separately how likely that friend is to come, then adds. Write the count as I1 + I2 + ... + I6, where I_k is 1 if face k appears; take expectations; each E[I_k] is just the probability face k appears. Face k is missed on one roll with probability 5/6, on all six with (5/6) to the 6th, 0.335. So each indicator averages 0.665 and the total averages 3.99.
The relationshipD the number of distinct faces seen in six rolls (5/6)^6 the chance a given face never appears in six rolls What it says in wordsThe expected number of distinct faces is six times the chance that any one face appears.Where does this pattern reappear?
The same shape answers how many distinct birthdays a group of n people has, how many of n hash buckets get used, and how many different stocks a random sample of trades touches. For n faces and n rolls, the expected share of faces seen is 1 - (1 - 1/n) to the n, which tends to 1 - 1/e, about 63.2%, as n grows. Six faces give 66.5%, already close. The exact enumeration of all 46,656 rolls gives a mean of 3.9906, the same number.
Where candidates lose it
The instinctive answer is 6, or something close to it, because six rolls over six faces feels like one of each. In reality repeats are the norm, and all six different faces happen in under 2% of runs.
The costlier trap is starting to enumerate cases under time pressure. Candidates who try to list exactly four distinct faces lose minutes. Say indicator variables and linearity in the first sentence.
What the interviewer asks next
- What is the expected number of faces that appear exactly once?
- How many rolls do you need, on average, to see all six faces?
- What is the variance of the number of distinct faces?
008Z1 and Z2 are independent standard normal random variables. What are the mean and variance of Z1 squared + Z2 squared, and what is the probability that it exceeds 2?Quant researchQuant trading
Try it first
What is P(Z1 squared + Z2 squared > 2)?
Show the worked solution
Mean 2, variance 4, and the probability of exceeding 2 is e to the -1, about 36.8%. Each Z squared has mean 1 and variance 2, so the sum has mean 2 and variance 4. The sum is a chi-squared with two degrees of freedom, which happens to be exactly an exponential with mean 2. Its tail beyond t is e to the -t/2, so beyond 2 it is e to the -1.
How do you get the mean and variance without the distribution?
E[Z squared] is the variance of Z, which is 1. For the variance of Z squared you need the fourth moment: E[Z to the 4] is 3 for a standard normal. So Var(Z squared) = 3 - 1 = 2, and because Z1 and Z2 are independent the variances add, giving a mean of 2 and a variance of 4. Say the fourth moment of 3 aloud; it is the number interviewers check you know, and it is why normal kurtosis is quoted as 3.
The sum of two squared standard normals has density one half times e to the minus x over 2, an exponential with mean 2 and variance 4, and the area beyond 2 is exactly e to the minus 1, about 36.8%. Why is this particular sum exponential?
Think of a dart thrown at a board where both the horizontal and vertical errors are independent standard normals. The joint density depends only on the distance from the centre, so the dart's direction is uniform and all the information is in the radiusThe distance of the point (Z1, Z2) from the origin, the square root of Z1 squared plus Z2 squared.. Switching to polar coordinates, the chance that the squared distance exceeds t is e to the -t/2, which is the tail of an exponential with mean 2. At t = 2 the answer is e to the -1.
The relationshipr the distance of (Z1, Z2) from the origin \frac{1}{2\pi} e^{-r^2/2} the joint density of two independent standard normals e^{-t/2} the tail of an exponential distribution with mean 2 What it says in wordsIn polar coordinates the angle integrates out and the radius gives an exponential tail.Why is 50% the tempting wrong answer?
Because 2 is the mean and people read the mean as the middle. For a right-skewed distribution the mean sits above the median, so less than half the mass lies beyond it; here the median is 2 ln 2, about 1.39. A simulation of 200,000 pairs gives a mean of 1.997, a variance of 3.96 and a tail share of 0.368, matching the exact results. The same polar trick is what powers the Box-Muller method for generating normal random numbers.
Where candidates lose it
The common wrong answer is 50%, from treating the mean as the median. Chi-squared variables are skewed to the right, and the skew is largest with few degrees of freedom.
The second trap is the variance. Candidates who say the variance of Z squared is 1 have confused it with the variance of Z. The fourth moment of 3 is the step, and missing it gives a variance of 2 for the sum instead of 4.
What the interviewer asks next
- What is the distribution of the square root of Z1 squared + Z2 squared?
- How would you use this to generate normal random numbers from uniforms?
- What are the mean and variance of a chi-squared with k degrees of freedom?
009A stock trades at 100 and in one period will be either 120 or 80. Interest rates are zero. Price a call option struck at 100 by building a portfolio of shares and borrowing that copies it, and explain why the real-world probability of the up move does not appear in the price.Options market makingQuant trading
Try it first
If you believe the stock goes up with probability 90%, what is the call worth?
Show the worked solution
The call is worth 10. It pays 20 if the stock goes to 120 and 0 at 80. Half a share pays 60 or 40, so half a share with a loan of 40 pays 20 or 0, exactly the call. That portfolio costs 50 - 40 = 10 today. If the call traded at any other price, you could buy the cheap one and sell the dear one for a riskless profit, so no probability is needed.
How do you build the copy?
Match the swing first. The call's payoff moves by 20 between the two states while the stock moves by 40, so the copy needs 20/40 = 0.5 of a share: that ratio is the option's deltaHow much an option's value changes for a one-unit change in the underlying price; here, the number of shares that copies the option.. Half a share is worth 60 or 40 at the end, which is 40 more than the call in both states. Borrow 40 today, repay 40 at the end with zero interest, and the copy pays exactly 20 or 0.
A call struck at 100 on a stock that moves to 120 or 80 is copied by half a share and a loan of 40, which pays 20 or 0 exactly as the call does and costs 10 today, so the call is worth 10 with no probability used. Why does the chance of the up move not matter?
Think of a shop selling a bundle of two items that you can also buy separately. The bundle's price is pinned by the parts, whatever you think about how useful the items are. The call is a bundle of half a share and a loan; the share price already reflects everyone's views about the up move, so the option inherits them and adds none of its own. A 90% view is a reason to hold the stock itself, not a reason to pay more for the call than its parts cost.
The relationship\Delta shares held in the copy B the amount borrowed, 0.5 x 80 - 0 = 40 q the risk-neutral weight on the up state, fixed by the prices, not by beliefs What it says in wordsThe copy's cost gives the price, and the same price is an average of the payoffs using weights set by today's stock price.What would you do if the call traded at 12?
Sell the dear thing and buy the cheap one. Sell the call for 12, buy half a share for 50 and borrow 40, a net cash inflow of 2 today; at the end the portfolio pays exactly what you owe on the call in either state. The 2 is kept whatever happens. The weight q = 0.5 that reproduces the price is called the risk-neutral probability, but it is a pricing weight backed out of the stock price, not a forecast. Say that distinction; interviewers listen for it.
Where candidates lose it
The trap is pricing the call as an expected payoff under your own view: 90% of 20 is 18. That price can be arbitraged against the stock, so nobody could trade it for long, and the interviewer wants to hear that the copying portfolio pins the price.
The second loss is getting 10 by assuming a 50% chance. The number is right by coincidence of the symmetric tree; ask yourself what happens with an up move to 130, and the risk-neutral weight changes to 1/2.5 = 0.4.
What the interviewer asks next
- Price the put struck at 100 and check put-call parity.
- What changes if interest rates are 5% for the period?
- The stock can go to 130 or 80 instead. Price the call again.
010A company's value to its current owner is equally likely to be anything from Rs 0 to Rs 100 crore, and only the owner knows the figure. In your hands the company would be worth 1.5 times that value. You may make one take-it-or-leave-it offer, which the owner accepts only if it is at least the company's value to them. What should you bid?Quant tradingQuant research
Try it first
Which bid maximises your expected profit?
Show the worked solution
Bid nothing. If a bid of b is accepted, the owner has told you the company is worth less than b to them, so its value is uniform on 0 to b and averages b/2. In your hands that is 1.5 x b/2 = 0.75b, a quarter less than you paid. Expected profit is (b/100) x (0.75b - b) = -b squared/400, negative for every positive bid. This is the winner's curse in its purest form.
Why does 75 look right and fail?
Picture buying a used car from someone who knows its history while you do not. If they agree to your price at once, that is itself news: sellers of good cars refuse low offers. Acceptance is not random; it happens exactly in the states where the company is worth less than you offered, so the average value you actually receive is the average below your bid, not the average overall. The naive 75 uses the unconditional average of 50 and forgets that you only trade when the owner is happy to sell.
The naive line values the company at its overall average and shows profit for any bid under 75, but conditioning on the owner accepting gives expected profit of minus b squared over 400, which is below zero for every positive bid, minus 6.25 crore at a bid of 50. How do you set up the expected profit?
Split it into the chance of a deal and the profit given a deal. A bid of b is accepted with probability b/100; given acceptance the owner's value is uniform on 0 to b, averaging b/2, so your value averages 0.75b and your profit averages minus 0.25b. Multiply: minus 0.25b x b/100, which is minus b squared over 400. At a bid of 50 that is minus 6.25 crore: you win half the time and lose 12.5 crore on average when you do.
The relationshipb your bid in Rs crore b/100 the chance the owner's value is below b b/2 the owner's average value, given that they accepted What it says in wordsThe chance of winning times the loss when you win is negative for every positive bid.When would bidding make sense, and where does this show up on a desk?
The multiplier is the lever. With a multiplier m, the profit given a deal is (m/2 - 1)b, so bidding pays only if you add more than double the owner's value; at exactly 2 you break even, and above 2 you should bid the full 100. On a trading desk the same logic is called adverse selectionThe tendency for the trades you actually get to come from counterparties with better information than you, so they are worse on average than a random trade.: the orders that fill against you are disproportionately the ones from people who know more. A quote that looks profitable against the average counterparty loses against the ones who choose to trade.
Where candidates lose it
Most candidates bid somewhere between 50 and 75, reasoning from the unconditional average value. That ignores the information in the owner's acceptance, which is the entire point of the question.
The second loss is a partial fix: realising acceptance is informative but then bidding a little lower, such as 60, to leave a margin. Any positive bid loses here. Write the expected profit as a function of b and let the algebra say zero.
What the interviewer asks next
- What multiplier would make you willing to bid, and how much would you then bid?
- What if the owner's value is uniform on 50 to 100 instead?
- How does this relate to a market maker who gets filled on their quotes?
011An urn starts with one red ball and one blue ball. Each turn you draw a ball at random, put it back, and add another ball of the same colour. After ten draws, what is the probability that exactly five of the draws were red?Quant researchQuant trading
Try it first
Your first instinct for P(exactly five reds)?
Show the worked solution
1/11, about 9.1%. Any particular sequence with five reds and five blues has probability 5! x 5! / 11!, whatever the order, because the numerators just count up the reds and blues separately. There are 10 choose 5 = 252 such sequences, and 252 x 5! x 5! / 11! = 1/11. The same algebra gives 1/11 for every count from 0 to 10.
Why is this not ten coin flips?
Think of a new café and its first customers. If the first few visitors like it and bring friends, it fills up; if they do not, it stays empty. Early luck compounds. In this urn each red draw adds a red ball, so it raises the chance of the next red: the draws reinforce each other, and runs that start lopsided tend to stay lopsided. That spreads the count of reds far more widely than independent coin flips, which bunch around five.
After ten draws from an urn that starts with one red and one blue ball and adds a ball of the drawn colour each time, every count of reds from 0 to 10 has probability 1/11, a flat line, while ten fair coin flips would pile up at five with 24.6%. Why does the order of the draws not matter?
Write out one sequence, say five reds then five blues. The chances are 1/2, 2/3, 3/4, 4/5, 5/6 for the reds, then 1/7, 2/8, 3/9, 4/10, 5/11 for the blues. Shuffle the order and the denominators are still 2 up to 11, while the red numerators still run 1 to 5 and the blue ones 1 to 5, so every arrangement of five reds and five blues has the same probability, 5! x 5! / 11!. Draws whose joint probability ignores order are called exchangeableA sequence of random variables whose joint distribution does not change when you reorder them, even though they need not be independent., and that is the property doing the work.
The relationship\binom{n}{k} the number of orders in which k reds can appear k!(n-k)!/(n+1)! the probability of any one such order What it says in wordsThe number of orders times the probability of each order is always one over n plus one.Is there a picture that makes 1/11 obvious?
Yes. This urn behaves exactly as if nature first picked a hidden red probability p uniformly between 0 and 1, and then flipped ten independent coins with that p. Averaging a binomial over a uniform p gives every count the same weight, 1/(n + 1). The urn is also a small model of momentum and of market share: an early lead makes further gains more likely, so the final split is highly uncertain, even though each step looks like a fair draw.
Where candidates lose it
The fast wrong answer is 252/1024, about 24.6%, from treating the draws as independent coin flips. The reinforcement is the whole question, and it spreads the outcomes out rather than pulling them to the middle.
The second trap is trying to add up paths through a tree of ten levels. Spot that every order of a given mix has the same probability, and the problem collapses to one line.
What the interviewer asks next
- What is the probability the eleventh draw is red, given that five of the first ten were red?
- What if the urn starts with two red and one blue ball?
- As the number of draws grows, what does the fraction of red balls converge to?
015You roll three fair dice. Is a total of 9 or a total of 10 more likely, given that each can be written as exactly six unordered combinations of three faces?Quant tradingProp trading firms
Try it first
Which total is more likely?
Show the worked solution
10 is more likely: 27 ways out of 216 against 25. The equally likely outcomes are the 216 ordered rolls, not the unordered combinations. A combination of three different faces covers six ordered rolls, one with a pair covers three, and a triple covers one. Total 9 includes 3 + 3 + 3, which counts once, and one fewer all-different combination, so it loses two orderings to 10.
Why are combinations the wrong thing to count?
Think of dealing two cards and asking whether a pair of kings or a king with a queen is more likely. There is one combination of each, but king and queen can arrive in either order while two kings are just two kings. Probability comes from counting outcomes that are equally likely, and with dice those are the ordered rolls: first die, second die, third die, 6 x 6 x 6 = 216 of them. Unordered combinations bundle different numbers of those outcomes, so counting combinations gives the wrong weights.
Totals of 9 and 10 each have six unordered combinations, but weighting each combination by its number of orderings gives 27 of 216 rolls for 10 against 25 for 9, mostly because 9 includes the triple 3, 3, 3, which happens only one way. How do you count the orderings fast?
Classify each combination by its repeats. Three different faces give 3! = 6 orders, a pair gives 3 orders, one for each position of the odd die, and a triple gives 1. For 9: 1 2 6, 1 3 5 and 2 3 4 are all different, 18; 1 4 4 and 2 2 5 are pairs, 6; 3 3 3 is a triple, 1; total 25. For 10: 1 3 6, 1 4 5 and 2 3 5 give 18; 2 2 6, 2 4 4 and 3 3 4 give 9; total 27. So 10 comes up 12.5% of the time and 9 only 11.6%.
The relationship6, 3, 1 the orderings of an all-different, a pair and a triple combination 216 the ordered outcomes of three dice What it says in wordsWeight each combination by its orderings and 10 beats 9 by two rolls in 216.Is there a shortcut that avoids listing?
Yes: symmetry. Replacing each face x by 7 - x maps a total of t to 21 - t, so the distribution of three dice is symmetric about 10.5, and 10 and 11 are the two most likely totals, each 27/216. Anything further from 10.5, including 9, must be less likely or equal; a quick count confirms it is 25. Historically this is the question gamblers put to Galileo, who answered it by counting ordered outcomes, which is still the method.
Where candidates lose it
The trap is the question's own framing: six combinations each invites the answer that the totals are equally likely. The interviewer wants you to reject the framing, not accept it.
The second loss is listing all 216 rolls, or writing out every ordering. Classifying by repeats, six, three or one, gets both totals in under a minute.
What the interviewer asks next
- What is the most likely total with four dice, and its probability?
- What is P(total is 9) with two dice, and why does 9 behave differently?
- How many ordered outcomes of three dice sum to 7?
017The true model is y = x1 + x2 + noise, where x1 and x2 are standardised and have correlation 0.5. You regress y on x1 alone, then regress the residuals on x2. What coefficient do you get on x2, and how would you recover the true value of 1 in two stages?Quant researchQuant trading
Try it first
What coefficient does the second stage give on x2?
Show the worked solution
You get 0.75, not 1. Regressing y on x1 alone gives a slope of 1 + 0.5 = 1.5, because x1 soaks up the half of x2 that moves with it. The residual is x2 - 0.5x1 + noise, whose slope on x2 is 1 - 0.5 squared = 0.75. To recover 1, residualise x2 on x1 as well and regress the residual of y on the residual of x2: the Frisch-Waugh-Lovell theorem.
Where does the missing quarter go?
Picture two salespeople who often work the same client. If you credit all joint sales to the first before looking at the second, the second looks worse than they are, because some of their work was already booked to the first. Stage one regresses y on x1 alone, and since x2 is correlated with x1, the coefficient on x1 rises to 1.5: it takes credit for 0.5 of x2. That piece has been removed from the residual, so stage two can only find what is left of x2's effect.
With a correlation of 0.5, x2 splits into 0.5 x1 plus an orthogonal part; stage one assigns the 0.5 x1 piece to x1, so regressing the residual on raw x2 gives 0.75, while regressing it on the orthogonal part of x2 recovers the true 1. How do you get 0.75 exactly?
Write the residual out. y - 1.5x1 = x2 - 0.5x1 + noise, and the slope of that on x2 is its covariance with x2 over the variance of x2: (1 - 0.5 x 0.5)/1 = 0.75. The formula generalises to 1 - rho squared times the true coefficient, so the bias gets worse as the regressors get more correlated: with rho = 0.9 you would find only 0.19. A simulation of 100,000 observations gives 1.506 for stage one and 0.752 for stage two.
The relationship\rho the correlation between x1 and x2, 0.5 x_2 - \rho x_1 the part of x2 left after regressing it on x1 What it says in wordsRegressing on raw x2 shrinks the answer by one minus rho squared; regressing on the part of x2 orthogonal to x1 gives the true coefficient.What does Frisch-Waugh-Lovell tell you to do?
To get a variable's coefficient from a multiple regression in stages, partial the other regressors out of both y and that variable, then regress residual on residual. Here that means regressing x2 on x1 as well, keeping the orthogonal part x2 - 0.5x1, and regressing the stage-one residual on it. The slope comes back as exactly 1; the simulation gives 1.003. This is why factor-neutralising a signal before testing it, rather than after, matters in quant research: the order of the stages changes the answer.
Where candidates lose it
The common answer is 1, on the belief that regressing residuals step by step is the same as a multiple regression. It is only the same when the regressors are uncorrelated, and the question gives you a correlation of 0.5 precisely to break that.
The second loss is saying the answer is biased without saying which way or by how much. Give 1.5 for stage one, 0.75 for stage two, the 1 - rho squared rule, and the fix.
What the interviewer asks next
- What would the stage-two coefficient be if the correlation were -0.5?
- In the two-stage FWL regression, how do the standard errors compare with the full multiple regression?
- You have a new signal correlated with a known factor. How do you test whether it adds anything?
022A three-way duel: you hit your target with probability 1/3, B with 2/3, and C never misses. You shoot first, then B, then C, repeating in that order until one person is left, and everyone aims to maximise their own survival. Where should you aim your first shot?Quant tradingQuant research
Try it first
Which first shot gives you the best chance of surviving?
Show the worked solution
Fire into the air. B and C each target the other, the bigger threat, so while both live nobody shoots at you. Aiming in the air gives survival of 2/3 x 3/7 + 1/3 x 1/3 = 25/63, about 39.7%. Aiming at C gives 31.2%, because a hit leaves you in a duel with B shooting first. Aiming at B gives 26.5%, because a hit leaves C, who never misses, to shoot you.
Who does everyone else aim at?
Start with the stronger players, because their choices fix yours. B aims at C, because if B shot you instead, C would kill B next turn for certain; C aims at B, the more dangerous of the two remaining threats. So while all three are alive, nobody is shooting at you. Think of two large firms in a price war while a small competitor stays out of it: the small firm's best move is often to let the giants weaken each other.
Firing into the air gives you 39.7% survival, against 31.2% for aiming at C and 26.5% for aiming at B, because hitting either rival makes you the survivor's only target while missing on purpose lets B and C shoot at each other first. How do you work out the two-player duels?
Against B with you shooting first, you win if you hit now, or if both miss and the same duel restarts. Call your survival x: x = 1/3 + (2/3)(1/3)x, so x = 3/7; if B shoots first, you must survive B's first shot, 1/3 of the time, giving 1/7. Against C you get exactly one shot, since C never misses: 1/3 if you shoot first, 0 if C does. Now combine. In the air: B hits C two times in three, giving you the 3/7 duel; otherwise C kills B and you get your one shot at C, 1/3. Total 25/63.
The relationship3/7 your survival in a duel with B when you shoot first 1/7 your survival in a duel with B when B shoots first 25/63 your survival after a deliberate miss What it says in wordsMissing on purpose beats both targeted shots: 75/189 against 59/189 and 50/189.What is the general lesson?
In a game with several players, weakening one rival can hurt you if it frees the strongest remaining player to turn on you. Your best shot is the one that keeps the others focused on each other. Say the limitation too: the answer depends on the hit rates and the order. Change the order of shooting, or let C aim at you, and the tree changes; the interviewer will often change a number or the order to see whether you rebuild the tree or repeat the slogan.
Where candidates lose it
The instinctive answer is to shoot at C, the most dangerous player. It ignores what happens after a hit: you have just made yourself B's only target, and B shoots first.
The second loss is assuming that firing into the air is allowed but not checking it is optimal. Candidates who have heard the answer before often cannot produce 25/63, 59/189 and 50/189 when asked. The numbers are the answer; the slogan is not.
What the interviewer asks next
- What if your hit rate were 1/2 instead of 1/3?
- What if C shot first and you shot last?
- What is B's overall survival probability when you fire into the air?
023You flip a fair coin until the pattern HTH appears. What is the expected number of flips? Why is it larger than the expected wait for HTT, when each pattern has the same probability of 1/8 at any given position?Quant tradingQuant research
Try it first
Expected flips to see HTH?
Show the worked solution
10 flips for HTH, against 8 for HTT. Track how much of the pattern you currently hold: nothing, H, or HT. For HTH, a tail after HT wrecks everything and you restart from nothing. For HTT, a head after HT breaks the pattern, but that head is itself a fresh start, so you keep an H. Solving the three expected-wait equations gives 10 and 8.
Why do equal probabilities give unequal waits?
Think of a combination lock where a wrong digit sometimes resets you to zero and sometimes lets you keep part of your progress. Two combinations can be equally likely to be dialled at random yet take different times to reach. Each three-flip window is HTH or HTT with the same 1/8 chance, but the windows overlap, and HTH occurrences tend to arrive in clusters, such as HTHTH, which spaces out the first appearance. The waiting time depends on where a near miss leaves you.
Waiting for HTH, a tail from state HT sends you back to the start and the average wait is 10 flips; waiting for HTT, a head from HT leaves you holding an H and the average wait is only 8. How do you set up the equations?
Let E0, E1 and E2 be the expected remaining flips when you hold nothing, H and HT. Each flip costs one and moves you to the next state with probability one half each way, so each state's wait is 1 plus the average of the two states it can move to. For HTH: E0 = 1 + (E1 + E0)/2, E1 = 1 + (E1 + E2)/2, E2 = 1 + (0 + E0)/2. Solving gives E2 = 6, E1 = 8, E0 = 10. For HTT only the last equation changes, to E2 = 1 + (0 + E1)/2, and the answers become 4, 6 and 8.
The relationship2^3 from the whole pattern matching itself 2^1 from HTH's last flip matching its first: the pattern overlaps itself What it says in wordsFor a fair coin, add 2 to the power k for every length k at which the pattern's start equals its end.Is there a shortcut an interviewer will accept?
Yes, the overlap rule, which comes from a fair-bet argument known as the ABRACADABRA methodA martingale argument in which gamblers arriving each flip bet on the pattern, used to compute expected waiting times for patterns.. For a fair coin, the expected wait is the sum of 2 to the k over every k where the first k flips of the pattern equal the last k. HTH matches itself at length 3 and at length 1, the single H, giving 8 + 2 = 10. HTT matches only at length 3, giving 8. HHH matches at 1, 2 and 3, giving 14. Derive the states first, then offer the rule as the check.
Where candidates lose it
The trap is answering 8 for every three-flip pattern, reasoning that each has probability 1/8 per window. That confuses frequency with first arrival: over a long run both patterns appear equally often, but HTH comes in overlapping clumps.
The second loss is getting the fall-back wrong in the state diagram. For HTT, after HT a head is not a return to nothing; it is a new H. Drawing that arrow to the start gives 10 for both and hides the whole point.
What the interviewer asks next
- What is the expected wait for HHH?
- Two players race, one waiting for HTH and one for HTT on the same flips. Who is more likely to win?
- With a biased coin that shows heads 60% of the time, what is the expected wait for HTH?
