Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
061A Rs 1,000 crore fund charges a 2% management fee and 20% of gains, with the performance fee taken on the gain left after the management fee. In a year with a 10% gross return, the manager cuts the management fee to 1%. What performance fee keeps the manager's revenue unchanged?Two SigmaNew York · 2026
Try it first
Which performance fee keeps revenue at the old level?
Show the worked solution
About 28.9%, roughly 29%. Under 2 and 20 the manager earns Rs 20 crore of management fee plus 20% of the remaining Rs 80 crore gain, Rs 36 crore in all. At 1% the management fee is Rs 10 crore and the gain left is Rs 90 crore, so the performance fee must bring in Rs 26 crore: 26 / 90 = 28.9%. It only balances at a 10% return.
Why is the answer not simply 30%?
Think of a shopkeeper who takes a fixed rent from a stall plus a fifth of whatever the stall-holder clears after paying that rent. If the rent is halved, the stall-holder clears more, and the shopkeeper's fifth is a fifth of a larger amount. Cutting the management fee enlarges the gain on which the performance fee is charged, so part of the lost management revenue comes back automatically. Here the old fee of 20% on Rs 90 crore earns Rs 18 crore, up from Rs 16 crore, and the true shortfall is Rs 8 crore, not Rs 10 crore. Adding 10 points to get 30% repays the shortfall twice over in part.
At a 10% gross return the old terms earn Rs 36 crore, the fee cut alone earns Rs 28 crore, and a performance fee of 28.9% restores Rs 36 crore; at any other return the two sets of terms pay differently, crossing only at 10%. How do you set it up in one line?
Write revenue as management fee plus performance rate times the gain after the management fee, and hold it equal. The new rate is the old revenue less the new management fee, divided by the new fee base: (36 - 10) / 90 = 28.9%. The investor's position is the mirror image: gross gain Rs 100 crore less Rs 36 crore of fees leaves Rs 64 crore, a 6.4% net return under either set of terms, which is a good check that you have kept the total fixed.
The relationshipA assets under management, Rs 1,000 crore G the gross gain, 10% of A, Rs 100 crore m0, m1 the old and new management fee rates, 2% and 1% p0, p1 the old and new performance fee rates What it says in wordsHold total revenue fixed and solve for the performance rate on the new, larger fee base.What does the swap change, even when revenue matches?
The match holds at one return only. Swapping management fee for performance fee moves risk onto the manager: less is guaranteed, more depends on the year. At a 5% gross return the old terms pay Rs 26 crore and the new ones Rs 21.6 crore; at 20% the old terms pay Rs 56 crore and the new ones Rs 64.9 crore. State the convention too: if the performance fee were charged on the gross gain instead, the old revenue is Rs 40 crore and the new rate comes to exactly 30%. Say which convention you assume before giving a number.
Where candidates lose it
The fast wrong answers are 40%, from doubling the performance fee because the management fee halved, and 30%, from adding back ten points. Both forget that the performance fee is charged on a base the management fee itself shrinks.
The second loss is giving one number without the conditions. A fee trade-off is always at an assumed return and an assumed fee base; say both and offer the sensitivity, because the interviewer's next question is what happens in a bad year.
What the interviewer asks next
- At what gross return would the manager prefer the new terms to the old?
- How does a hurdle rate of 5% change the new performance fee?
- Why might investors prefer 1 and 29 to 2 and 20 even though revenue matches at 10%?
Asked at Two Sigma, Equity Capital Markets, New York, 2026 (Wall Street Oasis):
the question was regarding how hedge funds operate, the 2/20 rule, and if one part of this equation changed
062A symmetric random walk starts at 0 and moves up or down by one each step with equal probability. What is the expected number of steps until it first reaches +3 or -3?Quant researchQuant trading
Try it first
What is the expected number of steps to hit +3 or -3 from 0?
Show the worked solution
9 steps. Let E(k) be the expected steps left from state k. Each step costs one and lands on a neighbour with equal chance, so E(k) = 1 + (E(k - 1) + E(k + 1))/2, with E(3) = E(-3) = 0. The solution is E(k) = (3 - k)(3 + k): 9 from 0, 8 from plus or minus 1, 5 from plus or minus 2. A symmetric walk needs about the square of the distance to travel it.
Why is the answer not 3?
Think of someone pacing a narrow corridor, choosing left or right at random with every step. They do reach a wall eventually, but they spend most of their time doubling back. A random walk with no drift covers distance like the square root of time, so covering a distance of 3 takes about 3 squared steps, not 3. The answer 3 would only be right if the walk always moved in one direction. The structure to name in the room is a first-step recursion on the state, which turns the question into seven unknowns with simple links between them.
The expected number of steps left is 9 from the centre, 8 one step out, 5 two steps out and 0 at the walls, tracing the parabola 9 - k squared, which is the product of the distances to the two walls. How do you solve the recursion quickly?
Use symmetry first: E(1) = E(-1) and E(2) = E(-2). Then three equations remain. From 2: E(2) = 1 + (E(1) + 0)/2. From 1: E(1) = 1 + (E(0) + E(2))/2. From 0: E(0) = 1 + E(1), because both neighbours are worth E(1). Substituting gives E(1) = 8 and E(0) = 9, and every value matches the product of the distances to the two walls. The second differences of E are all minus 2, which is why the answer is a parabola: each step must remove exactly one unit of expected time.
The relationshipE(k) the expected number of steps left from state k 1 the step just taken E(plus or minus 3) = 0 the walls, where the walk stops What it says in wordsOne step plus the average of the two neighbouring states, solved, gives the product of the distances to the walls.How do you check 9 a second way?
Use a martingaleA process whose expected next value, given everything so far, equals its current value: a fair game.. Squaring the position and subtracting the number of steps gives a fair game, because each step adds exactly 1 to the expected square. At the moment the walk stops its square is exactly 9, so the expected number of steps must also be 9. The same argument handles uneven walls: starting between -a and +b, the expected time is a times b. The cousin result is the hitting probability: from +1, the walk reaches +3 before -3 with probability 2/3, the distance to the far wall over the total width.
Where candidates lose it
The trap answers are 3, from treating the walk as if it marched to a wall, and infinity, from a half-remembered fact that a symmetric walk on the whole line takes infinite expected time to reach a given level. With two walls, escape is certain and fast.
The second loss is setting up seven equations and grinding them. Use symmetry to cut to three, then state the product rule and check it with the squared-position argument.
What the interviewer asks next
- Starting at 0, what is the expected time to hit +2 or -4?
- From +1, what is the probability of reaching +3 before -3?
- With only one wall at +3 and no lower wall, what is the expected time to reach it?
063X and Y are independent random variables, each uniform on 0 to 1. What is the density of X + Y, and what is the probability that X + Y is less than 1.5?CitadelChicago · 2025
Try it first
What is P(X + Y < 1.5)?
Show the worked solution
The density is a triangle, f(s) = s for s up to 1 and 2 - s from 1 to 2, and P(X + Y < 1.5) = 7/8. Convolving two flat densities gives a tent peaking at 1. The part above 1.5 is a triangle with base 0.5 and height 0.5, area 1/8. In the unit square it is the same corner: the line x + y = 1.5 cuts off a triangle with legs of 0.5.
Why is the sum not uniform on 0 to 2?
Roll two dice: a total of 7 can be made six ways, a total of 12 only one way. Continuous uniforms behave the same. A sum near the middle can be made from many pairs, a sum near either end from very few, so the density of the sum rises to a peak and falls again. The mechanism that builds it is {term('convolution', 'The density of a sum of independent variables: for each possible total, add up the density of every pair of values that makes it.')}: the density at s is the length of the set of x values for which both x and s - x lie between 0 and 1.
In the unit square the line x + y = 1.5 cuts off a corner triangle of area 1/8, and in the triangular density of the sum the tail above 1.5 is the same 1/8, so X + Y is below 1.5 with probability 7/8. How does the convolution give the triangle?
Fix a total s. You need x between 0 and 1 and also s - x between 0 and 1, so x must lie between max(0, s - 1) and min(1, s). For s below 1 that interval has length s, and for s above 1 it has length 2 - s, so the density is a tent with its peak of 1 at s = 1. Check that the area is 1: a triangle with base 2 and height 1. The mean is 1 and the variance is 1/12 + 1/12 = 1/6, both of which you can read from symmetry and independence.
The relationshipf(s) the density of the sum at the value s the indicator 1 when s - x is a valid value of Y, otherwise 0 What it says in wordsThe density of the sum at s is how many ways of splitting s are allowed, which rises linearly to 1 and falls back.Why give the square picture as well?
It is a check that costs ten seconds. Because the pair is uniform on the unit square, any probability about X + Y is an area, and the event X + Y at least 1.5 is the corner triangle above the line x + y = 1.5. Its legs run from 0.5 to 1 on each axis, so its area is 1/8 and the answer is 7/8 again. The density gives you the whole distribution; the square gives you any single probability fast. Keep both, because the next question is usually three uniforms, where the density becomes piecewise quadratic and the square becomes a cube: P(X + Y + Z < 1) = 1/6.
Where candidates lose it
The common slip is 3/4, from assuming a sum of uniforms is uniform. Sums are never uniform unless one of the pieces is degenerate; they pile up in the middle.
The second loss is getting the convolution limits wrong and producing a density that does not integrate to 1. Write the two constraints on x out loud, and check the triangle's area before you use it.
What the interviewer asks next
- What is the density of X - Y?
- What is P(X + Y + Z < 1) for three independent uniforms?
- What is the expected value of max(X, Y), and of X + Y given that X + Y > 1?
Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):
He was asking some questions about the probability, especially on the convolution.
064You may roll a fair die up to three times. After each roll you either stop and are paid the face showing, or roll again; if you reach the third roll you must take it. What is your optimal stopping rule, and what is the game worth?RBC Capital MarketsToronto · 2025
Try it first
On the first roll you see a 4. What do you do?
Show the worked solution
Keep a 5 or 6 on the first roll, a 4, 5 or 6 on the second, and take whatever the third gives; the game is worth 14/3, about 4.67. Work backwards. The last roll is worth 3.5. With two rolls left, keep anything above 3.5: (4 + 5 + 6)/6 + (1/2)(3.5) = 4.25. With three, keep anything above 4.25: (5 + 6)/6 + (2/3)(4.25) = 14/3.
Why start from the last roll?
Think of house hunting with three viewings booked: whether to accept the first flat depends on what the remaining viewings are likely to offer, and you only know that once you know how you would behave at the last one. The value of continuing at any point is defined by what you would do later, so the only roll whose value you know outright is the last one, and every earlier decision is built on it. That is backward induction, and the interviewer wants to hear the words before any numbers.
Solving from the last roll upwards, the third roll is worth 3.5, so the second roll keeps 4, 5 or 6 and is worth 4.25, so the first roll keeps only 5 or 6 and the whole game is worth 14/3, about 4.67. How do the values build up?
On the last roll you take the face, worth 3.5. On the second roll, stop if the face beats 3.5, which means 4, 5 or 6; otherwise you get 3.5 from the last roll. The value is the average of the faces you keep plus the chance you continue times the value of continuing: (4 + 5 + 6)/6 + (3/6)(3.5) = 4.25. On the first roll, the bar to beat is now 4.25, so only 5 and 6 are kept: (5 + 6)/6 + (4/6)(4.25) = 11/6 + 17/6 = 14/3, about 4.667.
The relationshipV_n the value of the game with n rolls still available max(f, V_n) keep the face f if it beats rolling on, otherwise take the value of continuing What it says in wordsEach extra roll is worth the average of the better of the face and the value of carrying on.What does the common shortcut cost, and where does this lead?
The shortcut is to keep anything above the single-roll average of 3.5 at every stage. On the first roll that keeps a 4, which gives 4.625 instead of 4.667. The threshold rises with the number of rolls left, because each spare roll is an option, and an option is worth more the longer it lives. With six rolls the value is 5.27, and with many rolls it approaches 6, since you can wait for a six. The same structure prices an American option: exercise early only when the payoff beats the value of holding on.
Where candidates lose it
The trap is the fixed threshold: stopping on 4 at the first roll because 4 beats 3.5. It ignores that the comparison is with the value of continuing, which is 4.25 with two rolls left, not 3.5.
The second loss is computing forwards, trying to enumerate all paths from the first roll. Say backward induction, solve the last roll, and build up; three lines of arithmetic do the whole job.
What the interviewer asks next
- What is the game worth with four rolls?
- You now pay Rs 1 for each reroll. How does the rule change?
- If you are paid the square of the final face, what is the first-roll rule?
Asked at RBC Capital Markets, Quantitative Trading, Toronto, 2025 (Wall Street Oasis):
Best way to maximize EV across 3 chosen dice rolls (can choose to continue or not).
065The correlation between X and Y is 0.2, and the correlation between Y and Z is 0.5. What is the full range of possible values for the correlation between X and Z?Tower Research CapitalNew York · 2019
Try it first
Which statement about corr(X, Z) is right?
Show the worked solution
Anywhere from about -0.75 to 0.95. The correlation matrix must be positive semidefinite, which bounds the third correlation at 0.2 x 0.5 plus or minus sqrt((1 - 0.2^2)(1 - 0.5^2)), that is 0.1 plus or minus 0.849. Geometrically, correlations are cosines of angles: X sits 78.5 degrees from Y and Z sits 60 degrees from Y, so X and Z are between 18.5 and 138.5 degrees apart.
Why does knowing two correlations restrict the third at all?
Think of three towns on a map. If A is close to B and B is close to C, then A cannot be far from C; if B is only loosely near both, A and C could be almost anywhere. Correlations behave like distances in disguise: each is the cosine of the angle between two returns viewed as vectors, and angles obey a triangle rule. X makes an angle of {TH_A:.1f} degrees with Y, because its cosine is 0.2, and Z makes 60 degrees with Y. The angle between X and Z is therefore at least the difference and at most the sum.
With X at 78.5 degrees from Y and Z at 60 degrees from Y, the angle between X and Z ranges from 18.5 to 138.5 degrees, so their correlation can be anything from -0.75 to 0.95, centred on 0.1. How do you get the bound algebraically?
Any valid correlation matrix must be positive semidefiniteEvery portfolio built from the variables has a variance of zero or more; for a correlation matrix this means its determinant and all leading minors are non-negative., because a portfolio cannot have negative variance. For three variables with correlations a, b and c, the condition is 1 - a^2 - b^2 - c^2 + 2abc of at least zero. Treating that as a quadratic in c gives c = ab plus or minus sqrt((1 - a^2)(1 - b^2)), so the third correlation lies in an interval centred on the product of the other two. With a = 0.2 and b = 0.5 that is 0.1 plus or minus sqrt(0.72), or -0.7485 to 0.9485.
The relationshiprho_XY 0.2, the correlation of X and Y rho_YZ 0.5, the correlation of Y and Z the square root the room left over after the parts of X and Z explained by Y What it says in wordsThe third correlation is the product of the two given ones, plus or minus how much of X and Z is unexplained by Y.What does the centre value 0.1 mean, and when is the sign forced?
Split X and Z each into a part explained by Y and a leftover. The explained parts always contribute 0.2 x 0.5 = 0.1; the leftovers can be correlated however you like, and they can move the total by up to 0.849 either way. So 0.1 is the answer only if the leftovers are uncorrelated. The sign of corr(X, Z) is forced positive only when the two given correlations are strong, specifically when their squares add to more than 1; with 0.8 and 0.7 the range is about 0.13 to 0.99. On a desk this is why two hedges that each track an index only loosely say almost nothing about each other.
Where candidates lose it
The two fast wrong answers are 0.1, from multiplying, and must be positive, from assuming correlation is transitive. Both treat correlation like a chain of causes rather than a geometry.
The second loss is reaching the determinant condition and stalling on the algebra. Lead with the angle picture: arccos 0.2 is about 78.5 degrees, arccos 0.5 is 60, and the bounds are the cosines of their sum and difference.
What the interviewer asks next
- If corr(X, Y) = 0.8 and corr(Y, Z) = 0.7, can corr(X, Z) be negative?
- What is the most negative common correlation three variables can share?
- Given corr(X, Y) and corr(Y, Z), what value of corr(X, Z) makes X and Z uncorrelated once Y is controlled for?
Asked at Tower Research Capital, Prop Trading, New York, 2019 (Wall Street Oasis):
What if the correlation between X and Y is 0.2 and the correlation between Y and Z is 0.5.
066A 2-year bond pays a 6% annual coupon and trades at par. What are its Macaulay duration and modified duration?PIMCOLos Angeles · 2024
Try it first
Which is closest to the Macaulay duration?
Show the worked solution
Macaulay duration is 1.943 years and modified duration is 1.833. At par the yield equals the coupon, 6%. The year 1 coupon is worth 6 / 1.06 = 5.660 and the final payment 106 / 1.06^2 = 94.340, adding to 100. Weight each time by its share: 0.0566 x 1 + 0.9434 x 2 = 1.943. Divide by 1.06 for modified duration: a 1 point rise in yield cuts the price by about 1.83%.
What is Macaulay duration actually measuring?
Think of a seesaw with weights placed along it at the dates money arrives, each weight equal to what that payment is worth today. Macaulay duration is the point where that seesaw balances: the average time to your money, weighted by present value. A zero-coupon bond has all its weight at maturity, so its duration equals its maturity. Any coupon puts a little weight earlier and pulls the balance point forward. Here the coupon is small and the bond is short, so the pull is small, and that is the whole shape of the answer before any arithmetic.
The year 1 coupon is worth 5.66 today and the year 2 payment 94.34, so the present-value weighted average time is 1.943 years, just 0.057 years before maturity, and dividing by 1.06 gives a modified duration of 1.833. How do you get the numbers fast in your head?
At par, the yield is the coupon, so discount at 6%. The coupon's present value, 6 / 1.06, is about 5.66, and the rest of the price, 94.34, sits at year 2, so duration is 2 minus the coupon's weight: 2 - 0.0566 = 1.943. That shortcut works for any two-period bond: start at the maturity and subtract the early weight times the time it saves. For a par bond there is also a closed form, (1 + y)/y times (1 - 1/(1 + y)^n), which gives the same 1.9434.
The relationshipPV_t the present value of the cash flow at time t, discounted at the yield P the price, 100 at par y the yield per period, 6% What it says in wordsDuration is the average payment date weighted by value today; modified duration converts it into a price sensitivity.Why are there two durations, and how good is the estimate?
Macaulay duration is a time; modified duration is a slope. Modified duration is the percentage price change for a one point change in yield, so this bond loses about 1.83% if yields rise from 6% to 7%, about Rs 0.018 per Rs 100 for each basis point. Repricing exactly gives 98.192 at 7%, a fall of 1.808%, and 101.859 at 5%, a rise of 1.859%. The rise is bigger than the fall, which is convexityThe curvature of the price against yield: prices rise more when yields fall than they drop when yields rise by the same amount.; for a two-year bond it is a small correction, for a thirty-year bond it is not.
Where candidates lose it
The trap answers are 2 years, from forgetting the coupon, and 1.83 given as the Macaulay figure, from mixing the two definitions. Say which duration you are giving and in what units: years for Macaulay, percent per point of yield for modified.
The second slip is discounting at the wrong rate. A bond at par yields its coupon, so there is nothing to solve for; say that first and the arithmetic is two divisions.
What the interviewer asks next
- What is the duration of a 2-year zero-coupon bond, and of a 2-year bond with a 20% coupon?
- Estimate the price of this bond if yields rise 50 basis points.
- Why does duration fall as the coupon rises, holding maturity fixed?
Asked at PIMCO, Product & Strategy, Los Angeles, 2024 (Wall Street Oasis):
Lots of random bond math questions -- duration of this bond with x coupon sold at par
067Five traders drop their business cards in a bowl and each draws one at random. What is the probability that nobody draws their own card, and what does it approach as the number of traders grows?Quant tradingQuant research
Try it first
As the number of traders grows very large, the chance nobody gets their own card...
Show the worked solution
44/120, about 36.7%, and it tends to 1/e, about 36.8%. Count the orderings with no fixed point by inclusion and exclusion: 5! times (1 - 1 + 1/2! - 1/3! + 1/4! - 1/5!) = 44 of the 120 ways. The bracket is the start of the series for e^(-1), so the answer barely moves with the number of traders: it is within about 0.001 of 1/e at five people and closer still after that.
Why does the answer not go to 0 or to 1?
Think of a secret gift exchange at the office. With more people, each person is less likely to draw their own name, but there are more people who could. Each trader matches with chance 1/n and there are n traders, so the expected number of matches is exactly 1 at every size, and the chance of zero matches settles rather than vanishing. That is why the answer converges to a constant. It is also why the question is asked: it checks whether you can count an event defined by an absence, which needs inclusion and exclusion.
The chance that nobody draws their own card swings above and below 1/e for small groups and is 44/120 = 0.3667 for five traders, against 1/e = 0.3679, so the number of traders barely matters once there are more than four. How does inclusion and exclusion give 44?
Count the orderings where at least one trader gets their own card, then subtract. Fix any one trader: 4! orderings each, five traders, 120 in total, but that counts orderings with two fixed traders twice. Alternately subtract and add the counts with one, two, three, four and five traders fixed, and you get 120 - 120 + 60 - 20 + 5 - 1 = 44 orderings with no match. Dividing by 120 gives the probability as 1 - 1 + 1/2 - 1/6 + 1/24 - 1/120, and each term is the e^(-1) series truncated.
The relationshipD_n the number of orderings of n items with no item in its own place, a derangement count (-1)^k / k! the inclusion and exclusion term for k traders fixed What it says in wordsThe chance of no match is the alternating series for e to the minus 1, cut off after n terms.How do you check 44 another way?
Use the recursion. Trader 1 takes some other trader's card, say trader j's, in n - 1 ways. Either j takes trader 1's card back, leaving n - 2 traders to derange, or j does not, which is the same as deranging n - 1 traders. So D(n) = (n - 1)(D(n - 1) + D(n - 2)), and from D(1) = 0 and D(2) = 1 you get 2, 9 and then 4 x (9 + 2) = 44. The number of matches is close to Poisson with mean 1 for any decent n, so exactly one match has about the same chance as none: for five traders it is 3/8 = 45/120.
Where candidates lose it
The trap is answering (4/5)^5, about 33%, by treating each trader's miss as independent. The draws are without replacement, so the events are linked; the right answer is close but not equal, and the reasoning is wrong.
The second loss is getting 44/120 and stopping, when the follow-up about large groups is the point. Say the series is the start of e^(-1), so the answer is about 37% for any number of traders.
What the interviewer asks next
- What is the expected number of traders who draw their own card?
- What is the probability that exactly one trader draws their own card?
- With 100 traders, roughly what is the chance that at least two draw their own card?
068A box holds ten coins: one has heads on both sides and nine are fair. You pick a coin at random, flip it five times and see five heads. What is the probability you picked the double-headed coin?Jump TradingChicago · 2018
Try it first
After five heads, roughly how likely is the double-headed coin?
Show the worked solution
32/41, about 78%. Before flipping, the odds are 1 to 9 against the double-headed coin. Five heads happen for certain with it and with chance 1/32 with a fair coin, a likelihood ratio of 32. Multiply: posterior odds are 32 to 9, which is 32/41. Each extra head doubles the odds, so a sixth head would take it to 64/73, about 88%.
Why is the answer not close to certain?
Imagine a rare illness and a decent test. A positive result raises the chance you have it, but if the illness is rare enough, most positives still come from healthy people. Evidence is weighed against how common each explanation was to begin with, so five heads, which a fair coin produces only once in 32 tries, still has to overcome nine fair coins for every double-headed one. The 97% instinct takes 1 minus 1/32 and forgets the nine-to-one start.
Starting from odds of 1 to 9, five heads multiply the odds by 32 to give 32 to 9, so the chance of the double-headed coin rises from 10% to 78%, roughly doubling the odds with each head. How do you run Bayes in odds form?
Odds form is the fastest way to say it in the room. Posterior odds equal prior odds times the likelihood ratio: (1 to 9) times 32 gives 32 to 9. Converting back, 32 out of 32 + 9 is 32/41, about 78%. The long form gives the same thing: the joint chance of picking the special coin and seeing five heads is 1/10, the joint chance of a fair coin and five heads is 9/10 x 1/32 = 9/320, and the posterior is (32/320) / (41/320).
The relationshipD the event that the double-headed coin was picked 5H the observation of five heads in five flips 32 the likelihood ratio: 1 divided by 1/32 What it says in wordsPrior odds of one to nine, multiplied by a likelihood ratio of thirty-two, give odds of thirty-two to nine.What does the odds picture tell you about more flips?
Each head is twice as likely under the double-headed coin, so every head doubles the odds and every tail ends the question, since the special coin never shows tails. After 0 to 5 heads the chance runs 10%, 18%, 31%, 47%, 64% and 78%; it passes 50% only after the fourth head. The useful follow-up is the next flip: it lands heads with chance 32/41 + (9/41)(1/2) = 73/82, about 89%. On a desk, the same arithmetic tells you how many winning days it takes before a new strategy's record says anything about skill.
Where candidates lose it
The trap is answering 31/32, about 97%, by looking only at how unlikely five heads are from a fair coin. That ignores the prior; with nine fair coins in the box, the base rate matters as much as the evidence.
The second slip is the reverse: staying near 10% because the coin was chosen at random. Say the odds form, prior times likelihood ratio, and both errors disappear.
What the interviewer asks next
- What is the probability that the next flip is heads?
- How many heads in a row would you need to be 99% sure?
- If one of the ten coins were double-tailed instead, how would five heads change the answer?
Asked at Jump Trading, Quantitative Research, Chicago, 2018 (Wall Street Oasis):
Then he asked one question of probability which can be solved by Bayesian formula.
069Daily returns are drawn from a normal with mean zero and standard deviation 1% on 90% of days, and from a normal with mean zero and standard deviation 4% on the other 10%. What are the overall standard deviation and the kurtosis of daily returns?Two SigmaNew York · 2025
Try it first
What is the kurtosis of the mixture?
Show the worked solution
The standard deviation is sqrt(2.5), about 1.58%, and the kurtosis is 12.72, against 3 for a normal. Moments of a mixture are weighted averages of the pieces' moments. The variance is 0.9 x 1 + 0.1 x 16 = 2.5. A normal's fourth moment is 3 sigma^4, so the fourth moment is 3 x (0.9 x 1 + 0.1 x 256) = 79.5. Kurtosis is 79.5 / 6.25 = 12.72.
Why does mixing two normals create fat tails?
Think of a city's daily traffic: most days are ordinary, and a few days a year there is a festival or a strike and everything is wild. Averaged across the year, the typical day looks calmer than the average suggests, and the extreme days are far more extreme than a single bell curve would allow. Mixing a calm regime with a rare wild one concentrates most days near zero and puts the rest far out, which is exactly what kurtosis measures: a high peak with heavy tails. Every piece is normal; the mixture is not.
With the same 1.58% standard deviation, the mixture is more peaked than the normal and its tail is far heavier: a daily move beyond 6% comes about 1 day in 75 under the mixture against about 1 day in 6,766 under the normal. How do you compute the moments without integrating?
Condition on the regime. Any moment of a mixture is the weighted average of that moment in each regime, because the density itself is the weighted average of the two densities. The second moment is 0.9 x 1^2 + 0.1 x 4^2 = 2.5, so the standard deviation is about 1.58%. For the fourth, use the fact that a normal with standard deviation sigma has fourth moment 3 sigma^4: 3 x (0.9 x 1 + 0.1 x 256) = 79.5. Kurtosis is the fourth moment over the variance squared.
The relationshipw_i the regime weights, 0.9 and 0.1 sigma_i the regime standard deviations, 1% and 4% kappa kurtosis, the fourth moment divided by the variance squared; 3 for any normal What it says in wordsAverage the variances and the fourth moments across regimes, then compare the fourth moment with the squared variance.What does the number mean for risk?
Kurtosis is 3 times the ratio of the average of sigma^4 to the square of the average of sigma^2, so any variation in volatility pushes kurtosis above 3, and the more uneven the regimes, the further it goes. The practical cost shows in the tails. A risk model that fits one normal with a 1.58% standard deviation expects a move beyond 6% about once in 6,766 days, roughly once every 27 years of trading; the mixture produces one about once in 75 days, several times a year. This is the simplest model of volatility clustering, and the reason daily returns on real assets show kurtosis well above 3.
Where candidates lose it
The trap is answering 3 because each piece is normal, or averaging the two standard deviations to 1.3% and treating the mixture as one normal. Mixtures average densities and moments, not shapes or standard deviations.
The second slip is forgetting the factor of 3 in a normal's fourth moment, or reporting excess kurtosis without saying so. Say which you mean: kurtosis 12.72, excess kurtosis 9.72.
What the interviewer asks next
- What mix of the two regimes maximises the kurtosis for a fixed overall variance?
- If the two regimes had different means, what would happen to the skew?
- How would you estimate the two regime volatilities from a year of daily returns?
Asked at Two Sigma, Quantitative Research, New York, 2025 (Wall Street Oasis):
They asked a couple questions involving Mixture Gaussians (e.g., probability density and moments).
070I offer you a bet. I draw two cards from a well-shuffled 52-card deck. If they are the same colour you win Rs 100; if they differ you lose Rs 100. Do you take the bet, and what payout on a win would make it fair?Quant tradingOptions market making
Try it first
Should you take the bet?
Show the worked solution
Decline it: same colour comes up 25 times in 51, so the bet loses about Rs 1.96 per Rs 100. Whatever the first card is, 25 of the remaining 51 share its colour and 26 do not. The bet is fair only if a win pays Rs 26 for every Rs 25 risked, that is Rs 104 on a win against Rs 100 on a loss.
Why is same colour less likely than different colour?
Imagine a classroom with 26 girls and 26 boys. Pick one child, then pick a second: the second is slightly more likely to be of the other sex, because the first pick removed one of their own. Drawing without replacement makes the second card lean away from the first card's colour, since that colour is now short by one. The first card is irrelevant to the answer, whichever colour it is. Only what is left in the deck matters: 25 matching cards and 26 non-matching.
After any first card, 25 of the 51 cards left share its colour and 26 do not, so same colour has chance 25/51 = 49.0%, the Rs 100 even bet is worth Rs -1.96 on average, and it becomes fair at a payout of Rs 104. What is the bet worth, and what makes it fair?
Expected value is the win times its chance less the loss times its chance: 100 x 25/51 - 100 x 26/51 = -100/51, about Rs -1.96. A bet is fair when the payout ratio equals the odds against winning, here 26 to 25, so the win must pay Rs 104 for each Rs 100 at risk. Quoting it as odds rather than a probability is how a trader would answer: you would take the bet at 26 to 25 or better, and decline anything worse.
The relationship25/51 the chance the second card matches the first card's colour W* the win payout that makes the expected value zero What it says in wordsThe second card matches with chance 25 in 51, so an even-money bet loses a little under 2 rupees per 100.Why would an interviewer offer such a small edge?
To see whether you notice it at all, and then whether you size your response to it. An edge of about 2% per bet is small for one play but decisive over many, so the right answer is to decline at even money and quote the price at which you would play. It also tests whether you separate the first card, which is free, from the conditional chance of the second. With replacement, or with an infinite deck, the bet would be exactly fair; the whole edge comes from the deck being finite, and it shrinks as more decks are shuffled together.
Where candidates lose it
The trap answer is that the bet is fair because colours are fifty-fifty. That treats the two cards as independent draws with replacement, which a single deck is not.
The second loss is getting 25/51 and stopping there. The question asks whether you take the bet and at what price, so finish with the decision and the fair odds of 26 to 25.
What the interviewer asks next
- What if the cards come from two decks shuffled together?
- You win if the two cards are the same suit. What payout makes that fair?
- I draw three cards and you win if all three are the same colour. What is the chance?
