Hedge Funds puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 38
- Topics
- 14
- Hard
- 30
061How many flips of a fair coin do you expect to need before you see two heads in a row? Why is the answer different if you wait for a head followed by a tail?Squarepoint CapitalLondon · 2025
Try it first
What are the expected waits for HH and for HT?
Show the worked solution
Six flips on average for two heads in a row, and four for a head then a tail. Track how far along the pattern you are. For HH, a tail at any point sends you back to the start, including right after a head. For HT, a head after a head keeps you one step away, so progress is never lost. Solving the two small chains gives 6 and 4.
Why do two equally likely patterns take different times?
Think of two ladders where a slip costs you differently. On one, slipping from the first rung drops you to the ground; on the other, you can only ever slip back to the first rung. Both patterns are equally likely in any given pair of flips, but after one head the wrong next flip costs you everything for HH and nothing for HT. A head then a tail breaks HH and restarts it; a head then a head is still a perfect start for HT.
In the HH chain a tail from the one-head state falls back to the start, so the expected wait is 6 flips; in the HT chain a head from the one-head state stays where it is, so progress is never lost and the wait is 4 flips. How do you solve the chain?
Let E0 be the expected flips still needed from the start and E1 after one head. For HH, E0 = 1 + E1/2 + E0/2 and E1 = 1 + E0/2, because a tail from one head sends you back, and these solve to E1 = 4 and E0 = 6. For HT, the one-head state just waits for a tail, which takes 2 flips on average, and reaching the first head takes 2 more, giving 4. Say the states out loud before the algebra; the interviewer wants to hear them named.
The relationshipE0 expected flips still needed from the start E1 expected flips still needed after one head 1 the flip you are about to make What it says in wordsEach state's expected wait is one flip plus the average wait from wherever that flip sends you.What is the general pattern?
Patterns that can fail back to nothing take longer. The expected wait for n heads in a row is 2 to the power n + 1, minus 2: 2, 6 and 14 flips for one, two and three heads. This is a Markov chainA process whose next step depends only on the current state, not on how it got there, so it can be solved state by state. at heart, and the same state-by-state method handles anything that depends on a path: a streak, a barrier, a drawdown rule on a trading book.
Where candidates lose it
The usual wrong answer is 4 for both, reached by noting that each pattern has a one-in-four chance in a pair of flips. That treats the flips as separate pairs, which they are not: a pattern can start at any flip, and what happens after a failure depends on the pattern.
The second loss is setting up one equation instead of two. Name the states, start and one head, and write one equation for each.
What the interviewer asks next
- How many flips do you expect to need for three heads in a row?
- Two players race: one wins at the first HH, the other at the first HT. Who is more likely to win?
- With a coin that lands heads 60% of the time, how long do you expect to wait for HH?
Asked at Squarepoint Capital, Quantitative Research, London, 2025 (Wall Street Oasis):
statistical problems e.g. # of throws expected to get 2 heads in a row
063The sample variance computed with n minus 1 in the denominator is an unbiased estimator of the population variance. Is its square root an unbiased estimator of the standard deviation?Squarepoint CapitalLondon · 2026
Try it first
Is the square root of the unbiased sample variance unbiased for the standard deviation?
Show the worked solution
No. The square root of the unbiased variance underestimates the standard deviation on average. The square root is concave, so by Jensen's inequality the average of the square roots is below the square root of the average. For normal data with two observations the estimate averages about 0.80 sigma; the bias shrinks as the sample grows, to about 6% at five observations and under 1% at thirty.
Why does taking a square root break unbiasedness?
Two square rooms have floor areas of 4 and 16 square metres, so their sides are 2 and 4 metres. Average the areas, 10, and take the root: 3.16 metres. Average the sides instead: 3 metres. Averaging and then taking a square root gives a bigger answer than taking square roots and then averaging, because the square root bends downwards. The sample variance is right on average, so the average of its square roots must fall short of the true standard deviation.
Two equally likely variance estimates of 0.04 and 1.96 average to the true variance of 1.0, but their square roots, 0.2 and 1.4, average only 0.8, below the true standard deviation of 1.0, because the square-root curve bends downwards. How big is the bias?
It depends on the sample size and on the distribution. For normal data the expected sample standard deviation is c4 times sigma, with c4 about 0.80 at n = 2, 0.94 at n = 5, 0.97 at n = 10 and 0.99 at n = 30. At n = 2 you can check it directly: the sample standard deviation is the gap between the two draws divided by the square root of 2, and the average gap between two normal draws is 2 sigma over the square root of pi, which leaves the square root of 2/pi, about 0.798.
The relationships the square root of the unbiased sample variance sigma the true standard deviation c4(n) the correction factor for normal data, below 1 for every n What it says in wordsThe average sample standard deviation is a fixed fraction of the true one, and that fraction is below one.Does it matter in practice?
Sometimes. With a year of daily returns the bias is a rounding error; with a handful of monthly returns for a new fund it is not. A manager with five monthly returns has a volatility estimate that averages about 6% too low under normality, which flatters a Sharpe ratioAverage excess return divided by the standard deviation of returns, a measure of return per unit of risk. before anyone has looked at fat tails. Dividing by c4 removes the bias for normal data, but the fix depends on the distribution, so name the assumption. And unbiased is not the same as most accurate.
Where candidates lose it
The trap is assuming unbiasedness carries through any function of an estimate. It carries through straight-line transformations only; the square root is curved, so the property is lost.
The second loss is saying it is biased without the direction or the size. Say biased low, give the Jensen reason in one sentence, and quote about 0.80 at two observations, shrinking towards 1 as the sample grows.
What the interviewer asks next
- Is the square of an unbiased estimator of the standard deviation unbiased for the variance?
- Why does the sample variance divide by n minus 1 rather than n?
- Which estimator of sigma has the lowest mean squared error for normal data?
Asked at Squarepoint Capital, Quantitative Research, London, 2026 (Wall Street Oasis):
Is the square root of the unbiased estimator for sample variance unbiased for standard deviation?
077In a Monty Hall game you pick door 1. This host does not know where the car is: he opens one of the other two doors at random, and it happens to show a goat. Should you switch, and why does the usual two-thirds answer no longer hold?Squarepoint CapitalLondon · 2026
Try it first
The host opened a door at random and it happened to show a goat. What is your chance of winning if you switch?
Show the worked solution
It makes no difference: switching and sticking each win half the time. List the six equally likely cases of car position and the host's random pick. Two of them reveal the car, and you have seen that they did not happen. The four that remain split two and two. The knowing host gives two thirds only because he never risks the car, which pushes those two cases into the switch column.
Where does the usual two-thirds answer come from?
In the standard game the host knows where the car is and always opens a goat door. Your first pick is right one time in three, and nothing the knowing host does can change that, so the other two thirds sit on the remaining closed door. His choice carries information because it is forced: when the car is behind door 2, he must open door 3, and when it is behind door 3, he must open door 2.
What changes when the host picks at random?
Picture a friend who does not know the answer to a quiz question and strikes out one option on a whim. If that option happens to be wrong, you have learned less than if someone who knew had struck it. The random host is that friend. Write out six cases: the car behind door 1, 2 or 3, each with the host's coin choosing door 2 or door 3. In two of the six the random host opens the car door, and the goat you saw rules those two out, which removes switch wins rather than stick wins.
In the same six equally likely cases, a knowing host redirects the two where his coin points at the car, so switching wins 4 of 6; a random host shows the car in those two, they are ruled out, and switching wins 2 of the 4 that remain, one half. Count what is left. The car behind door 1 survives both host choices: two cases where sticking wins. The car behind door 2 survives only when the host opened door 3, and the car behind door 3 only when he opened door 2: two cases where switching wins. Two against two.
The relationship2/6 cases where the car is behind the other closed door and the host showed a goat 4/6 all cases where the host showed a goat, which is what you observed What it says in wordsCondition on what you saw: of the cases where a goat appears, half have the car behind the door you would switch to.Why would an interviewer want the intuitive answer broken rather than recited?
Because the lesson travels to every desk. The same observation carries different information depending on the process that produced it. A strong track record shown by a manager who launched ten funds and closed the nine that did badly is a host choosing which door to open for you. Before you update on evidence, ask whether the source could have shown you something else, and whether it chose what to show.
Where candidates lose it
Candidates who know the classic puzzle answer two thirds on reflex. The interviewer changed one fact, that the host knows, and is checking whether you notice that the host's knowledge is exactly what made switching better.
The other loss is saying one half without a reason, which sounds like the naive answer to the classic game. Name the two ruled-out cases, where the car would have been shown, and show that both come out of the switch column.
What the interviewer asks next
- With 100 doors and a knowing host who opens 98 goat doors, what is your chance if you switch?
- With 100 doors and a random host who happens to open 98 goat doors, what is it now?
- Where does the same logic show up when you read a fund family's track record?
Asked at Squarepoint Capital, Quant Research Intern Interview, London, 2026 (Wall Street Oasis):
notably I was asked why the 'intuitive answer' was not true rather than just what the correct answer was, related to the Monty Hall problem
087A fund charges 2 and 20 and earns a gross return of 12% on Rs 1,000 crore. Investors push the management fee down to 1%. What performance fee keeps the manager's total fee income unchanged at that return?Two SigmaNew York · 2026
Try it first
What performance fee keeps the manager whole?
Show the worked solution
About 27.3%. At 2 and 20 the manager earns Rs 20 crore of management fee and 20% of the remaining Rs 100 crore gain, Rs 20 crore: Rs 40 crore in all. At 1%, the management fee is Rs 10 crore and the gain after it is Rs 110 crore. To keep Rs 40 crore the performance fee must bring in Rs 30 crore, which is 30/110, or 27.3%. The two deals match only at a 12% gross return.
What does the manager earn today?
Think of a tailor who charges a fixed stitching fee plus a share of whatever the finished suit sells for above cost. Cut the fixed fee and the share must rise to keep the same income, but the share now applies to a slightly larger base. At 2 and 20 on Rs 1,000 crore earning 12%, the manager takes Rs 20 crore of management fee plus 20% of the Rs 100 crore gain left after it, Rs 40 crore in total. Investors keep Rs 80 crore, a net return of 8%.
Under 2 and 20 the manager's Rs 40 crore is Rs 20 crore fixed plus 20% of Rs 100 crore; under a 1% fixed fee it is Rs 10 crore plus 27.3% of Rs 110 crore, because the smaller fixed fee leaves a larger gain for the performance fee. How do you find the new performance fee?
Keep the total at Rs 40 crore. The management fee falls to Rs 10 crore, so the performance fee must bring in Rs 30 crore, and it is charged on a gain of Rs 110 crore, not Rs 100 crore. 30 divided by 110 is 27.3%. The base grows because less has been taken off the top before the performance fee is worked out. State the assumptions as you go: no hurdle rate, and no earlier losses to recover below a high-water markThe highest value an investor has paid a performance fee on; no new performance fee is charged until the fund climbs back above it..
The relationship40 the manager's total fee income under 2 and 20, Rs crore 10 the new 1% management fee, Rs crore 120 the gross gain, 12% of Rs 1,000 crore What it says in wordsThe new performance fee is the income still needed, divided by the gain left after the new management fee.Is the new deal really the same for investors?
Only at a 12% gross return. The new deal pays the manager less in poor years and more in good ones, so it moves risk from the investors to the manager. At a 4% gross return the old deal pays Rs 24 crore and the new one Rs 18.2 crore; at 20% the old pays Rs 56 crore and the new one Rs 61.8 crore. That is why allocators push for a lower fixed fee even at the price of a higher share: they would rather pay for performance than for size.
Gross return 2 and 20, Rs crore 1 and 27.3, Rs crore Who gains from the switch 4% 24.0 18.2 Investors 12% 40.0 40.0 Neither 20% 56.0 61.8 Manager Manager's fee income on Rs 1,000 crore under each deal: the two match at a 12% gross return, the new deal pays Rs 5.8 crore less at 4% and Rs 5.8 crore more at 20%. Where candidates lose it
The common slip is 30%: dividing the Rs 30 crore needed by the old Rs 100 crore base, forgetting that a smaller management fee leaves a larger gain for the performance fee to work on. The other is 25%, dividing by the gross Rs 120 crore.
The second loss is stopping at 27.3% as though the two deals were identical. They match only at a 12% return, and the interviewer wants to hear who comes out ahead in good years and in bad ones.
What the interviewer asks next
- At what gross return does the manager prefer the new deal?
- Add a 5% hurdle to the new deal. What performance fee keeps the manager whole now?
- How does a high-water mark change what the performance fee is worth to the manager?
Asked at Two Sigma, Equity Capital Markets, New York, 2026 (Wall Street Oasis):
the 2/20 rule, and if one part of this equation changed, how would the other variable make up for it
089The equity risk premium is 4.5%, and a market's fair multiple is 1 divided by (real yield + premium - real growth). Real yields rise from 1.5% to 2.5% while expected real growth rises from 2.0% to 2.5%. What happens to the fair multiple?CitadelNew York · 2026
Try it first
Where does the fair multiple go?
Show the worked solution
The fair multiple falls from 25x to about 22.2x, a compression of about 11%. The denominator is the real yield plus the premium minus growth: 1.5 + 4.5 - 2.0 = 4.0% before, and 2.5 + 4.5 - 2.5 = 4.5% after. Real yields rose by a full point and growth by only half a point, so the net rate rose by half a point, and the multiple, its inverse, fell.
Why is the multiple one over a spread?
Think of a shop that pays you rent forever, rising a little each year. What you would pay for it depends on the return you demand minus how fast the rent grows. For earnings paid out and growing forever, price over earnings is one divided by the required return minus growth, so the multiple depends only on the gap between the two. In real terms the required return is the real yield plus the equity risk premiumThe extra return investors demand for holding shares rather than government bonds.. The question treats all earnings as paid out, which is the assumption to name.
The net rate in the denominator rises from 4.0% to 4.5% because real yields climbed a full point while growth climbed half a point, so the fair multiple falls from 25x to 22.2x, about 11% lower. How do you work it out quickly?
Compute the denominator before and after. Before: 1.5 + 4.5 - 2.0 = 4.0%, a multiple of 25x. After: 2.5 + 4.5 - 2.5 = 4.5%, a multiple of 22.2x. Rates went up by 1.0 point and growth by 0.5, so the spread widened by 0.5 point. A 0.5-point rise on a 4.0% base is a 12.5% rise in the denominator, and the multiple falls by 1 minus 1/1.125, about 11.1%.
The relationshipr the real yield on government bonds ERP the equity risk premium, 4.5% g expected real growth of earnings What it says in wordsThe fair multiple is one over the net rate: what investors demand minus how fast the earnings grow.What does this teach beyond the arithmetic?
Higher yields do not hurt equities one for one if growth rises with them. What matters is whether real yields rise faster or slower than expected growth: faster compresses multiples, slower expands them. Had growth also risen a full point, to 3.0%, the net rate would be 4.0% again and the multiple 25x. The limitation to state is sensitivity: near a 4% net rate, a half-point move shifts the multiple by about 2.8 turns one way and 3.6 the other, so small errors in the premium or the growth guess swamp the answer.
Where candidates lose it
The quick wrong answer is that nothing happens because both rates went up. The question is built so that growth rises by only half as much as yields, and it is the spread, not the level, that sets the multiple.
The second loss is dropping the growth change and answering 20x. Write the denominator out in full, before and after; it takes ten seconds and removes both errors.
What the interviewer asks next
- Growth rises by a full point, to 3.0%. What is the multiple now?
- The equity risk premium also falls to 4.0%. What is the net effect?
- Why do shares whose value sits far in the future fall more than the market when real yields rise?
Asked at Citadel, Software, New York, 2026 (Wall Street Oasis):
real yields rising faster than growth expectations predicts equity multiple compression
