Hedge Funds puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 38
- Topics
- 14
- Hard
- 30
009How would you price a digital option that pays Rs 100 if the index is above 11,000 at expiry, using only the prices of ordinary call options?Volatility and relative value fundsProp and quant trading firms
Try it first
A digital paying Rs 100 above 11,000 is closest to which position?
Show the worked solution
Replicate it with a tight call spread: buy 5 calls at 10,990 and sell 5 at 11,010. The position pays 0 below 10,990 and 100 above 11,010, a steep ramp standing in for the step. So the digital costs about 5 x (C at 10,990 minus C at 11,010). With calls at 212.40 and 203.60 that is 5 x 8.80 = Rs 44. In the limit, the price is minus 100 times the slope of call prices against strike.
Why does a call spread look like a step?
A steep enough ramp can stand in for a stair. A call spread's payoff is a ramp: nothing below the lower strike, rising point for point between the strikes, flat above the upper strike. Narrow the strikes and scale up the size, and the ramp tightens into the step a digital pays. Here the strikes are 20 points apart, so each spread pays at most 20, and five spreads pay at most 100, the digital's payout.
Five 10,990 / 11,010 call spreads pay 0 below 10,990 and 100 above 11,010, overpaying the digital just below 11,000 and underpaying just above it, and with illustrative calls at 212.40 and 203.60 the position costs Rs 44. What does the price of the spread tell you?
The spread costs the difference in call prices, so the digital costs five times that. As the strikes close in, the price becomes minus 100 times the slope of the call price against strike, and that slope is the discounted market-implied chance of finishing above the strike. With the illustrative quotes, 8.80 across 20 points is a slope of 0.44, a digital worth Rs 44 and an implied chance of about 44% before discounting.
The relationshipD the digital's price C(K) the price of a call struck at K h half the gap between the strikes, here 10 What it says in wordsA digital is a call spread scaled up as it narrows, so its price is the slope of call prices with strike.Which spread does a desk that sold the digital actually buy?
A desk that has sold the digital wants a hedge that pays at least 100 wherever the digital does. The centred spread overpays just below the strike and underpays just above it, so a seller hedges with five 10,980 / 11,000 spreads, which pay the full 100 by the strike and cost a little more. That difference is what the desk charges for an index that settles right at the strike. One more point marks a strong answer: the slope of call prices includes the change in implied volatility across strikes, so with the usual equity skew, where lower strikes carry higher volatility, the digital is worth more than a flat-volatility model says.
Where candidates lose it
Candidates reach for a pricing formula straight away. The question said using only call prices, and the interviewer wants the replication argument; the formula comes after, if at all.
The second loss is the size. A call spread 20 points wide pays at most 20, so it takes five of them to pay 100; a candidate who buys one spread prices the digital at a fifth of its value.
What the interviewer asks next
- How would you replicate a digital that pays 100 below 11,000?
- What do the digital call and the digital put at the same strike cost together?
- Why is a digital close to expiry, with the index at the strike, so hard to hedge?
017You roll a fair die until the first 6 appears and are paid the sum of every roll, including the final 6. What is the expected payout?Quant and systematic fundsProp and quant trading firms
Try it first
What is the expected payout?
Show the worked solution
21. The number of rolls until the first 6 averages 6, so there are on average 5 non-six rolls plus the 6. A roll known not to be a 6 is equally likely to be 1 to 5, so it averages 3, not 3.5, giving 5 x 3 + 6 = 21. Wald's identity confirms it: 6 expected rolls times 3.5 a roll is also 21, because a stop that looks only at past rolls does not bias the total.
Why is it tempting to get 23.5?
Suppose you keep buying scratch cards until one wins. Every card before the winner is, by definition, a loser, so those cards are worth less than an average card. The stopping rule changes the rolls before the stop: each one is known not to be a 6, so it averages 3, and valuing them at 3.5 overpays by 0.5 a roll, 2.5 in all. Five non-sixes at 3.5 plus a 6 is 23.5, which counts the high side twice: once in the 3.5 and again in the final 6.
In a typical run of 4, 1, 5, 2, 3 and then 6, the five rolls before the stop are non-sixes averaging 3 and the total is 21; 5 x 3 + 6 and 6 x 3.5 both give 21, while 5 x 3.5 + 6 = 23.5 wrongly treats the early rolls as ordinary rolls. How do the two routes agree?
Route one splits the sum: the expected number of non-six rolls times their average, plus the final 6. The count of rolls is a geometric wait with success chance 1/6, so it averages 6, of which 5 are non-sixes: 5 x 3 + 6 = 21. Route two is Wald's identityFor a stopping rule that uses only rolls already seen, the expected total equals the expected number of rolls times the average roll.: the expected total is the expected number of rolls times the average roll, 6 x 3.5 = 21. The low early rolls and the high final roll balance exactly.
The relationshipS the total paid N the number of rolls, including the 6 X a single roll, averaging 3.5 before any conditioning What it says in wordsCounted either as all rolls at 3.5 or as non-sixes at 3 plus a 6, the expected payout is 21.Why would a trading firm ask this?
Stopping rules are everywhere on a desk: exit at the first stop-loss hit, rebalance at the first breach of a band. The question checks whether you can tell when a stopping rule biases what you observe, as it does for the early rolls, and when it does not, as for the total. Say the condition too: Wald's identity needs the decision to stop to use only rolls already seen, and the expected number of rolls to be finite. A rule that could peek at the next roll would break it.
Where candidates lose it
The slip is 5 x 3.5 + 6 = 23.5. It treats the rolls before the 6 as ordinary rolls, when the stopping rule guarantees none of them is a 6, which pulls their average down to 3.
The opposite slip is to distrust 6 x 3.5 because stopping at a 6 seems to bias it. It does not: the total is unbiased for any stopping rule that looks only at the past. Give both routes and say why they agree.
What the interviewer asks next
- What is the expected payout if the final 6 is not paid?
- You stop at the first 5 or 6 instead. What is the expected payout?
- If you could choose to stop whenever you like, what would you pay to play?
020A book holds Rs 60 crore of a stock with 30% volatility and Rs 40 crore of another with 20% volatility, and the two have a correlation of 0.5. What is the book's volatility in rupees, and what share of the risk comes from each position?Man GroupBoston · 2022
Try it first
What share of the book's risk comes from the Rs 60 crore position?
Show the worked solution
The book's volatility is about Rs 23.1 crore a year, and the Rs 60 crore position carries about 74% of it on 60% of the capital. Stand-alone risks are Rs 18 crore and Rs 8 crore. Book variance is 18 squared plus 8 squared plus 2 x 0.5 x 18 x 8, which is 532, so volatility is Rs 23.07 crore. Each position's contribution is its covariance with the book over the book's volatility: Rs 17.17 crore and Rs 5.90 crore, which add back to the total.
Why is risk not shared out like capital?
Two friends share a taxi. One rides twice as far, straight through the traffic jam; the other gets off after a short hop. Splitting the fare by the number of bags each carries would be absurd. Risk belongs to a position in proportion to how much it moves and how much it moves with everything else, not to how much money sits in it. Here the first stock is larger, more volatile and positively correlated with the second, so it carries far more than its 60% of the capital.
Stand-alone risks of Rs 18 crore and Rs 8 crore add to Rs 26 crore, diversification at a 0.5 correlation removes Rs 2.93 crore, and the book's Rs 23.07 crore of volatility splits 74.4% to the first position and 25.6% to the second, against a 60 to 40 split of capital. The relationshipa, b stand-alone rupee volatilities: 60 x 30% = 18 and 40 x 20% = 8 \rho the correlation, 0.5 RC a position's contribution to book volatility What it says in wordsEach position owns its own variance plus half the shared term, and dividing by the book's volatility turns that into rupees of risk.Position Capital, Rs crore Volatility Stand-alone risk Risk contribution Share of risk A 60 30% 18.0 17.17 74.4% B 40 20% 8.0 5.90 25.6% Book 100 26.0 23.07 100.0% Rs crore of annual volatility. The stand-alone risks add to Rs 26.0 crore, but the book's volatility is Rs 23.07 crore, of which position A contributes 74.4% and position B 25.6%. Why do the contributions add up exactly to the total?
Split the variance. The cross term, 2 x 0.5 x 18 x 8 = 144, is shared equally, 72 to each position. So position A owns 324 + 72 = 396 of the 532 of variance and position B owns 64 + 72 = 136, and dividing each by the book's volatility of 23.07 gives rupee contributions that sum exactly to Rs 23.07 crore. The diversification benefit is the gap between the stand-alone total of Rs 26 crore and the book's Rs 23.07 crore.
What does a risk manager do with the split?
Cut where the risk is, not where the money is. Each rupee in position A carries 28.6 paise of marginal risk against 14.7 paise in position B, so trimming Rs 10 crore from A lowers book volatility by roughly Rs 2.9 crore. Recomputing exactly gives Rs 2.84 crore, close to the estimate. The limitation: the split is a snapshot at one correlation, and when correlations move, both the total and the split move with them.
Where candidates lose it
The quick answer shares risk like capital, 60 and 40, or like stand-alone risk, 18 and 8. The first ignores volatility and the second ignores correlation; neither sums to the book's actual Rs 23 crore of risk.
The second slip is adding the stand-alone risks to get the book's risk, Rs 26 crore. Volatilities do not add unless the correlation is exactly 1; variances do, with the cross term included.
What the interviewer asks next
- If the correlation fell to zero, how would the risk split between the two positions?
- How much of position B would you add to minimise the book's volatility, holding A fixed?
- How do transaction costs change which position you trim first?
Asked at Man Group, Investment Management, Boston, 2022 (Wall Street Oasis):
How do you understand portfolio risk and transaction cost?
021A stock with 20% annual volatility and no drift starts at Rs 100. Roughly what is the chance it ends the year more than 10% higher, and what is the chance it touches Rs 110 at some point during the year?Quant and systematic fundsProp and quant trading firms
Try it first
How does the chance of touching 110 compare with the chance of finishing above it?
Show the worked solution
About 31% to finish above Rs 110, and about 62% to touch it during the year. A 10% move is half of one year's 20% standard deviation, and a normal variable ends more than half a standard deviation up 30.9% of the time. By the reflection principle, every path that touches 110 and ends below has a mirror twin that ends above, so touching is twice as likely as finishing above: 61.7%.
Why is finishing above 110 about a one in three chance?
Scale the move by the volatility. Over one year the price spreads out with a standard deviation of about Rs 20, so Rs 110 is half a standard deviation above the start, and a normal variable finishes more than half a standard deviation up 30.9% of the time. Treating the price as an arithmetic random walk is close enough for a 10% move; a lognormal model, in which prices cannot go negative, gives a slightly lower figure, about 28%. Say you are approximating, and say which way the error runs.
A path that touches 110 and ends below it has a mirror twin, reflected in the barrier after the first touch, that ends above it, so the chance of touching 110 during the year, 61.7%, is twice the chance of finishing above it, 30.9%. Why is touching twice as likely as finishing above?
Think of a walker on a foggy path who is equally likely to step forward or back. Once she reaches a marker post, her remaining steps are a fair coin again: from the post she is as likely to end past it as short of it. So for every path that touches 110 and ends below, reflecting the part after the touch gives an equally likely path that ends above; touching paths split evenly between the two. Every path that ends above must have touched on the way, so the chance of touching is twice the chance of ending above.
The relationshipb the barrier, Rs 110 S_T the price at the end of the year 0.5 the barrier's distance in standard deviations: 10 / 20 What it says in wordsFor a driftless continuous walk, the chance of ever reaching a level is double the chance of finishing beyond it.Where does the factor of two matter on a desk?
Anything that triggers on a touch rather than on the finish. A stop-loss set 10% away is hit about twice as often as the price ends beyond it, and an option that pays on a touch is worth roughly twice one that pays only if the price finishes past the same level. The limitation: the factor of two holds for a driftless, continuously watched walk. Drift, jumps and checking the price only at the daily close all move it, and a checked-daily barrier is touched a little less often than a continuous one.
Where candidates lose it
The common error is answering the touch question with the finishing probability, 31%, as though the path does not matter. A price can visit 110 in March and be back at 100 by December, and the question asked about the visit.
The second is forgetting to scale by volatility. Ten per cent sounds small, but against 20% a year it is half a standard deviation, not a rare event. Say the scaling first, then the number.
What the interviewer asks next
- What is the chance the stock touches Rs 90 during the year?
- Roughly what is the chance it touches both 110 and 90?
- How does a positive drift change the ratio between touching and finishing above?
023A stock trades at 40 times forward earnings, pays out half its earnings as dividends, and investors want a 12% return. What long-run growth rate is the price implying?Long-short equity fundsGlobal macro funds
Try it first
What growth does the price imply?
Show the worked solution
About 10.75% a year, for ever. In a constant-growth model the forward P/E equals the payout ratio divided by (required return minus growth). With a P/E of 40 and a payout of 0.5, r minus g must be 0.5/40 = 1.25%, so g = 12% - 1.25% = 10.75%. Retaining half its earnings, the company would need a return on equity of 21.5% for ever to fund that growth.
How does a P/E hide a growth assumption?
A flat that rents for Rs 30,000 a month and sells for Rs 1.2 crore is priced at 400 months of rent; a buyer paying that is quietly assuming the rent will grow. A price multiple is a compressed forecast: fix the return investors want and the share of earnings paid out, and the multiple pins down the growth the price needs. The constant-growth model, price equals next year's dividend over (r minus g), divided through by earnings, gives P/E = payout/(r - g).
Rearranging P/E = payout/(r - g) with a P/E of 40, a 50% payout and a 12% required return leaves a dividend yield of 1.25% and implied growth of 10.75% a year for ever, which needs a return on equity of 21.5%. The relationshipE_1 next year's earnings, so the P/E is forward payout the share of earnings paid as dividends, 0.5 r the return investors require, 12% g the constant growth rate the price implies What it says in wordsThe required return is the dividend yield plus growth, so growth is whatever is left after the yield.Is 10.75% for ever plausible?
Test it against the business. Growth funded by retained earnings is return on equity times the share retained, so 10.75% growth with half the earnings kept needs a return on equity of 21.5%, held for ever. Few businesses hold returns like that for decades, and no company can outgrow the economy it sells into indefinitely, so compare the figure with the nominal growth you expect for that economy and say it as your assumption. The price is not wrong by arithmetic; it is demanding by assumption.
How sensitive is the answer?
Very. Because r - g is only 1.25%, every point on the required return moves the implied growth by a full point: at 11% the price implies 9.75%, at 13% it implies 11.75%. Using trailing rather than forward earnings shifts it too: 40 = 0.5(1 + g)/(0.12 - g) gives 10.62%. A high multiple rests on a thin gap between two large numbers, so small changes in either swing the value.
Where candidates lose it
Candidates treat the P/E as if it were price over dividend and forget the payout, which gives r - g = 2.5% and growth of 9.5%. The payout ratio is what turns earnings into the dividends the model actually discounts.
The other miss is stopping at 10.75% without judging it. The question asks what the price implies; the strong answer adds the return on equity it needs and whether that is believable.
What the interviewer asks next
- What P/E would 6% growth for ever justify at the same payout and required return?
- How does a rise in the required return to 13% change the implied growth?
- Why is a constant-growth model a poor fit for a young, fast-growing company?
024Five rational pirates, ranked A to E by seniority, must split 100 gold coins. The most senior proposes a split and all vote; it passes if at least half vote in favour, the proposer included. Otherwise the proposer is thrown overboard and the next most senior proposes. Each pirate wants first to survive, then to maximise coins, and votes against when indifferent. What does A propose?Prop and quant trading firmsLong-short equity funds
Try it first
What does A propose?
Show the worked solution
A proposes 98 for himself, 0 for B, 1 for C, 0 for D and 1 for E. Work backwards. With two pirates, D's own vote is half, so he keeps all 100. With three, C buys E with 1 coin. With four, B buys D with 1 coin. With five, A needs two votes beyond his own and buys the two pirates who get nothing in the four-pirate split, C and E, for one coin each.
Why start from the end?
Planning a train journey, you work back from when you must arrive, not forward from when you wake up. Each pirate votes by comparing the offer with what he would get if the proposal failed, so you can only price a vote once you know the next round's outcome, which means solving the smallest game first and working upwards. That method, backward induction, is the whole puzzle.
Solving from two pirates upwards gives splits of 100, 0 for two; 99, 0, 1 for three; 99, 0, 1, 0 for four; and 98, 0, 1, 0, 1 for five, because each proposer buys the pirates left with nothing in the next smaller game. How does each round play out?
Two pirates, D and E: D proposes 100 for himself, and his own vote is half, so it passes. Three pirates: C needs one more vote and buys E, who gets nothing in the two-pirate game, for 1 coin: 99, 0, 1. Four pirates: B needs one more vote and buys D, who gets nothing in the three-pirate game: 99, 0, 1, 0. Five pirates: A needs two more votes and buys C and E, both empty-handed in the four-pirate game: 98, 0, 1, 0, 1.
The relationshipn the number of pirates still aboard fallback coins what the voter gets if this proposal fails What it says in wordsA proposer needs half the votes and buys each one for a coin more than that pirate's next-round payoff.What is the interviewer really testing?
Whether you reason about the alternative each party faces rather than about fairness. A vote costs exactly one coin more than what the voter gets if the deal fails, so the cheapest supporters are the ones with the worst fallback. The same logic runs through any negotiation: a creditor backs a restructuring plan when it beats their recovery in liquidation, and support is cheapest from those whose alternative is worst. State the assumptions: perfect rationality, and a pirate who is indifferent votes against, which is why one coin, not zero, is needed.
Where candidates lose it
Candidates reach for a fair split, or reason forwards about who might be angry, and drown. Without the backward chain there is no way to know what any vote costs.
The second slip is offering coins to the wrong pirates: to B, or to D, who already does well in the four-pirate game. Buy the cheapest votes, from the pirates with nothing to lose, and say why.
What the interviewer asks next
- What happens with six pirates?
- What changes if a proposal needs a strict majority to pass?
- What if pirates vote yes when an offer merely equals their fallback?
025Four observations, 3.1, 7.4, 5.2 and 9.0, come from a uniform distribution on 0 to theta. What is the maximum likelihood estimate of theta, is it biased, and how would you correct it?AQR Capital ManagementTown of Greenwich · 2022
Try it first
Which statement is right?
Show the worked solution
The MLE is 9.0, the largest observation; it is biased low, and the unbiased correction is 5/4 x 9.0 = 11.25. The likelihood is 1 over theta to the fourth for any theta of at least 9.0 and zero below it, so it peaks at the sample maximum. But the maximum of n draws averages n/(n + 1) of theta, here 4/5, so scaling by (n + 1)/n removes the bias.
Why is the MLE the largest observation?
A friend draws four raffle tickets numbered from 1 up to some unknown top number, and the highest you see is 90. The top number is at least 90; guessing higher only spreads your belief over tickets nobody drew. The likelihood, 1 over theta to the n, is zero for any theta below the largest observation and falls as theta rises above it, so it is maximised exactly at the sample maximum, 9.0. This is a case where you do not differentiate: the maximum sits on a boundary, not where a slope is zero.
Four uniform draws on 0 to theta cut it into five gaps of equal expected size, so the largest draw averages four fifths of theta; the MLE of 9.0 therefore sits below theta, and scaling by 5/4 gives the unbiased 11.25, against 12.35 from doubling the sample mean. The relationshipL(\theta) the likelihood of the four observations n the number of observations, 4 \max the largest observation, 9.0 What it says in wordsThe likelihood peaks at the largest observation, which on average falls short of theta by a factor n/(n + 1).Why is it biased, and by how much?
The largest draw can never be above theta and is almost always below it. Four points dropped at random on 0 to theta cut it into five gaps of equal expected length, so the largest point sits on average four fifths of the way up, and the MLE underestimates theta by a fifth on average. Multiplying by 5/4 fixes it: 9.0 becomes 11.25. The bias shrinks as n grows, since n/(n + 1) tends to 1, but with four points it is large.
How does it compare with the obvious alternative?
The method of moments doubles the sample mean, since a uniform on 0 to theta averages theta/2: the mean here is 6.175, giving 12.35. Both 11.25 and 12.35 are unbiased, but the corrected maximum has a much smaller variance, theta squared over n(n + 2) against theta squared over 3n, because the largest draw carries the most information about the top of the range. With four points that is theta squared over 24 against theta squared over 12: half the variance.
Where candidates lose it
Candidates differentiate the log-likelihood, get minus n over theta, set it to zero and find no solution. The likelihood only falls on the allowed range, so the maximum sits at the boundary, the largest observation; say that before reaching for calculus.
The second miss is calling the MLE unbiased because maximum likelihood estimates are often well behaved. Here it is biased low by construction, and the interviewer expects the (n + 1)/n correction.
What the interviewer asks next
- What is the variance of the corrected estimator with four observations?
- What is the MLE if the distribution is uniform on theta to 2 theta?
- Derive the expected value of the maximum of n uniform draws.
Asked at AQR Capital Management, Trading, Town of Greenwich, 2022 (Wall Street Oasis):
Derive the mle for some given distribution. Explain linear regression intuitively and derive the ols estimate.
040A bag holds four stones, each black or white. Before you look, every count of black stones from 0 to 4 is equally likely. You draw two stones without replacement and both are black. What is the chance the next stone is black, and at what price would you bet on it?CitadelNew York · 2025
Try it first
Chance the third stone is black:
Show the worked solution
3/4, so fair odds are 3 to 1 on black. Two black draws rule out bags with 0 or 1 black. The ways to draw two blacks in order are 2, 6 and 12 for bags with 2, 3 and 4 black, so those bags now carry 10%, 30% and 60%. The next stone is black with chance 0, 1/2 and 1 in them, which averages to 3/4. A contract paying 100 if black is worth 75.
How do the two black draws change your view of the bag?
If a friend pulls two red sweets from a jar you have never seen, you start to suspect it is mostly red. Each possible bag is reweighted by how likely it was to produce what you saw: equal priors times the chance of two blacks. Drawing two blacks in order has 0 ways from bags with 0 or 1 black, 2 x 1 = 2 ways with 2 black, 3 x 2 = 6 with 3 black and 4 x 3 = 12 with 4 black. Out of 20 in total, that is 10%, 30% and 60%: the all-black bag is now the favourite.
Before the draw each bag is 20% likely; after two black stones the bags with 2, 3 and 4 black carry 10%, 30% and 60%, and since the next stone is black with probability 0, one half and 1 in those bags, the chance it is black is 3/4. How do you turn the posterior into a price?
Average the chance of black over the bags you still believe in. In the 2-black bag both remaining stones are white; in the 3-black bag one of two is black; in the 4-black bag both are, so the answer is 0.1 x 0 + 0.3 x 1/2 + 0.6 x 1 = 3/4. A contract paying 100 if the next stone is black is worth 75. You would buy it below 75 and sell it above; as a market maker you might quote 70 at 80. Offered even money on black, your expected profit per rupee staked is 0.75 minus 0.25, which is 50 paise.
The relationshipk the number of black stones in the bag B1 B2 the event that the first two draws are black 2, 6, 12 the ordered ways to draw two blacks from bags with 2, 3 and 4 black What it says in wordsThe chance of another black is the average of each bag's chance, weighted by how much the evidence now favours that bag.Check it with Laplace rule of successionWith a uniform prior, after s successes in n trials, the chance the next trial succeeds is (s + 1) / (n + 2).: after 2 blacks in 2 draws the next is black with chance (2 + 1)/(2 + 2) = 3/4, and the rule holds exactly for this finite bag. Two routes to 3/4 is what separates a solid answer from a lucky one. The limitation: everything rests on the flat prior. If you had reason to think mixed bags were more common, 3/4 would fall.
Where candidates lose it
The common loss is saying 1/2 because the remaining stones are unknown, which throws away the information in the two draws. The question is about updating, and the interviewer wants to see the reweighting.
The second loss is weighting the surviving bags equally, a third each, which gives 1/2. Each bag must be weighted by how likely it made two blacks: 2, 6 and 12. Then price it: a probability without a bet is half the answer at a trading firm.
What the interviewer asks next
- The third stone is black too. What is the chance the fourth is black?
- You quote 70 at 80 on a contract paying 100 if black and someone who has seen the bag lifts your offer. What now?
- How does the answer change if the prior is that each stone is black with probability one half, independently?
Asked at Citadel, Quantitative Trading, New York, 2025 (Wall Street Oasis):
Extended bayes derivative question about four stones in a bag (black and white stones).
041Two orders arrive one after the other. The first arrives after a wait that is exponential with a mean of one minute; the second arrives after a further, independent exponential wait with the same mean. What is the probability that both have arrived within one minute?CitadelChicago · 2025
Try it first
Your estimate:
Show the worked solution
1 minus 2/e, about 26.4%. The total wait is the sum of two independent exponential waits. Convolving the two densities gives t e^-t, a gamma shape that starts at zero because two steps cannot both be instant. Its area from 0 to 1 is 1 minus e^-1 (1 + 1), which is 1 minus 2/e. A second route: it is the chance that a Poisson process with rate 1 produces at least two arrivals in one minute.
Why is this not the chance of one wait, squared?
Squaring would be right if both orders were racing from the same start line. Here they queue: the second clock only starts when the first order lands. It is like two buses where you must take the first to reach the stop for the second. The event is that the sum of the two waits is under one minute, which is stricter than each wait being under one minute. The chance one wait is under a minute is 1 minus 1/e, about 63%; squaring gives 40.0%, which answers a different question.
How do you get the density of the sum?
Add up every way to split the total t between the two waits. The density of a sum of independent waits is the convolutionThe density of a sum of two independent variables, found by integrating one density against the other shifted over every possible split of the total. of their densities, and for two exponentials it is t e^-t. Every split of t into s and t minus s has density e^-s times e^-(t minus s), which is e^-t whatever s is, and there is a length t of possible splits. Integrate t e^-t from 0 to 1 by parts and you get 1 minus 2/e.
The total of two independent one-minute exponential waits has density t e^-t, which starts at zero and peaks at one minute, so only 26.4% of its area, 1 minus 2/e, lies below one minute. The relationshipX1, X2 the two independent exponential waits, each with mean one minute t e^-t the density of their sum, from convolving the two exponential densities What it says in wordsThe chance the total wait is under a minute is the area under the gamma density up to one minute.Check it with counting. Exponential waits are the gaps of a Poisson process, so both orders arrive within a minute exactly when the process makes at least two arrivals in that minute. With one arrival expected per minute, the chance of zero is e^-1 and of exactly one is e^-1, so at least two is 1 minus 2/e, the same number. Say both routes and the interviewer will usually skip ahead.
Where candidates lose it
The common loss is squaring the single-wait probability, which answers the question of two independent orders racing in parallel. Read the setup again: one after the other means the waits add.
The second loss is freezing on the convolution integral. If the integral will not come, switch to the Poisson count: at least two arrivals in one minute. Candidates who know one route and not the other are the ones interviewers push hardest.
What the interviewer asks next
- What is the probability that three orders in sequence all arrive within two minutes?
- Given that both orders arrived within one minute, what is the expected arrival time of the first?
- The two waits have means of one and two minutes. What is the density of their sum?
Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):
if i knew this was about convolutions, i would have answered better.
042A bowl holds 100 noodles. You repeatedly pick two free ends at random and tie them together, until no free ends remain. What is the expected number of loops?D.E. ShawNew York · 2026
Try it first
Roughly how many loops?
Show the worked solution
About 3.28 loops. With k strands in the bowl there are 2k free ends. Pick any end; the other end you pick is one of the remaining 2k minus 1, and exactly one of those belongs to the same strand, so this tie closes a loop with chance 1/(2k minus 1). Either way the number of strands falls by one. Adding 1/199 + 1/197 + ... + 1/3 + 1 gives about 3.284.
What does one tie do, whatever happens?
Start with the bookkeeping, because it makes the rest easy. Every tie reduces the number of loose strands by exactly one: either it closes a strand into a loop, or it joins two strands into one longer strand. So there are always exactly 100 ties, and with k strands left there are 2k free ends. The question becomes how many of those 100 ties happen to close a loop.
What is the chance a given tie closes a loop?
Think of a room of dancers holding hands in lines: grab one free hand, then pick a second free hand at random, and a circle forms only if the second hand is at the other end of the same line. With 2k free ends, the second end is one of 2k minus 1, and exactly one of them is the other end of the strand you picked, so the chance is 1/(2k minus 1). Give each tie an indicator that is 1 if it closes a loop; by linearity of expectation the expected number of loops is the sum of the chances, from 1/199 for the first tie up to 1 for the last.
The first tie closes a loop with chance 1 in 199 and the chances stay tiny until the last few ties, 1/5, 1/3 and 1, so the expected number of loops from 100 noodles is only 3.28, about a third of it from the last three ties. The relationshipk the number of strands in the bowl before a tie 1/(2k-1) the chance that tie joins the two ends of one strand What it says in wordsAdd each tie's chance of closing a loop to get the expected number of loops.For a sense check without a calculator: the sum of odd reciprocals up to 1/(2n minus 1) is about half of ln n plus ln 2 plus half of Euler's constant, which for n = 100 gives 3.28. The number of loops grows only like the logarithm of the number of noodles: a million noodles would give only about 7.9 loops.
Where candidates lose it
The first loss is trying to track the lengths of the strands, which quickly becomes impossible. The length of a strand never matters; only the count of strands does.
The second loss is getting 1/(2k minus 1) right but summing it wrong, for example as 100 x 1/199. Write the sum out from the last tie backwards, 1 + 1/3 + 1/5, and the size of the answer becomes obvious.
What the interviewer asks next
- What is the variance of the number of loops?
- What is the probability that you end with exactly one big loop?
- How does the answer grow with the number of noodles, roughly?
Asked at D.E. Shaw, Research, New York, 2026 (Wall Street Oasis):
What is the expected number of loops from tying 100 noodles' ends together randomly
