Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
060In how many ways can you place four queens on a 4 by 4 board so that no queen attacks another? How would you organise the search, and what does the same method give for five queens on a 5 by 5 board?Goldman SachsNew York · 2026
Try it first
How many non-attacking placements of four queens exist on a 4 by 4 board?
Show the worked solution
Two on a 4 by 4 board and 10 on a 5 by 5 board. Place one queen per row, trying columns left to right, and abandon a branch the moment the next row has no safe column. On 4 by 4 this visits 16 placements and finds columns 2, 4, 1, 3 and 3, 1, 4, 2. On 5 by 5 the same search visits 53 placements, against 3,125 boards for brute force.
How do you organise the search so it stays small?
Think of filling a seating plan for a wedding where some guests cannot sit near each other. You seat table by table, and the moment a table has no acceptable guest left you undo the previous choice instead of finishing a doomed plan. Backtracking builds the answer one decision at a time and abandons a partial answer as soon as it breaks a rule, so it never enumerates the boards that fail early. For queens, two rules come free from the structure: one queen per row, and one per column, which leaves only the diagonals to check at each step.
Row by row, the search tries 16 placements on the 4 by 4 board; four branches die when a row has no safe column, and two reach row 4, giving the solutions 2, 4, 1, 3 and 3, 1, 4, 2. How does the 4 by 4 search actually run?
Start with the corner. A queen in column 1 of row 1 leaves row 2 only columns 3 and 4, and both paths run out of safe squares by row 3 or row 4, so no solution uses a corner queen. A queen in column 2 forces column 4 in row 2, then column 1 in row 3 and column 3 in row 4, which works. Columns 3 and 4 are mirror images of 2 and 1. So there are exactly two solutions, and they are reflections of each other. Saying the symmetry out loud halves the work and is exactly what an interviewer building up from a base case wants to hear.
What changes on 5 by 5, and how does the method scale?
The larger board has more room, and every row 1 column leads somewhere. The same search finds 10 solutions after 53 placements, while brute force over one queen per row would test 3,125 boards. On 8 by 8 it finds all 92 solutions in 2,056 placements out of 16,777,216 one-per-row boards. In code, keep three sets, used columns, used down-diagonals (row minus column) and used up-diagonals (row plus column), so each safety check is constant time.
Where candidates lose it
Candidates start listing boards by eye and lose track, or they try all C(16, 4) = 1,820 ways to place four queens anywhere. The interviewer wants the structure: one per row, a column choice per row, and pruning.
The second loss is counting the two 4 by 4 solutions as four or eight by treating rotations as new. Say whether you count symmetric boards as distinct, and note that here the two solutions are each other's mirror image.
What the interviewer asks next
- Write the backtracking function and state its time complexity in the worst case.
- How would you count solutions up to rotation and reflection?
- Why do the 2 by 2 and 3 by 3 boards have no solution at all?
Asked at Goldman Sachs, Quantitative Research, New York, 2026 (Wall Street Oasis):
I was asked a backtracking question in 1 of the rounds in the superday.
061A Rs 1,000 crore fund charges a 2% management fee and 20% of gains, with the performance fee taken on the gain left after the management fee. In a year with a 10% gross return, the manager cuts the management fee to 1%. What performance fee keeps the manager's revenue unchanged?Two SigmaNew York · 2026
Try it first
Which performance fee keeps revenue at the old level?
Show the worked solution
About 28.9%, roughly 29%. Under 2 and 20 the manager earns Rs 20 crore of management fee plus 20% of the remaining Rs 80 crore gain, Rs 36 crore in all. At 1% the management fee is Rs 10 crore and the gain left is Rs 90 crore, so the performance fee must bring in Rs 26 crore: 26 / 90 = 28.9%. It only balances at a 10% return.
Why is the answer not simply 30%?
Think of a shopkeeper who takes a fixed rent from a stall plus a fifth of whatever the stall-holder clears after paying that rent. If the rent is halved, the stall-holder clears more, and the shopkeeper's fifth is a fifth of a larger amount. Cutting the management fee enlarges the gain on which the performance fee is charged, so part of the lost management revenue comes back automatically. Here the old fee of 20% on Rs 90 crore earns Rs 18 crore, up from Rs 16 crore, and the true shortfall is Rs 8 crore, not Rs 10 crore. Adding 10 points to get 30% repays the shortfall twice over in part.
At a 10% gross return the old terms earn Rs 36 crore, the fee cut alone earns Rs 28 crore, and a performance fee of 28.9% restores Rs 36 crore; at any other return the two sets of terms pay differently, crossing only at 10%. How do you set it up in one line?
Write revenue as management fee plus performance rate times the gain after the management fee, and hold it equal. The new rate is the old revenue less the new management fee, divided by the new fee base: (36 - 10) / 90 = 28.9%. The investor's position is the mirror image: gross gain Rs 100 crore less Rs 36 crore of fees leaves Rs 64 crore, a 6.4% net return under either set of terms, which is a good check that you have kept the total fixed.
The relationshipA assets under management, Rs 1,000 crore G the gross gain, 10% of A, Rs 100 crore m0, m1 the old and new management fee rates, 2% and 1% p0, p1 the old and new performance fee rates What it says in wordsHold total revenue fixed and solve for the performance rate on the new, larger fee base.What does the swap change, even when revenue matches?
The match holds at one return only. Swapping management fee for performance fee moves risk onto the manager: less is guaranteed, more depends on the year. At a 5% gross return the old terms pay Rs 26 crore and the new ones Rs 21.6 crore; at 20% the old terms pay Rs 56 crore and the new ones Rs 64.9 crore. State the convention too: if the performance fee were charged on the gross gain instead, the old revenue is Rs 40 crore and the new rate comes to exactly 30%. Say which convention you assume before giving a number.
Where candidates lose it
The fast wrong answers are 40%, from doubling the performance fee because the management fee halved, and 30%, from adding back ten points. Both forget that the performance fee is charged on a base the management fee itself shrinks.
The second loss is giving one number without the conditions. A fee trade-off is always at an assumed return and an assumed fee base; say both and offer the sensitivity, because the interviewer's next question is what happens in a bad year.
What the interviewer asks next
- At what gross return would the manager prefer the new terms to the old?
- How does a hurdle rate of 5% change the new performance fee?
- Why might investors prefer 1 and 29 to 2 and 20 even though revenue matches at 10%?
Asked at Two Sigma, Equity Capital Markets, New York, 2026 (Wall Street Oasis):
the question was regarding how hedge funds operate, the 2/20 rule, and if one part of this equation changed
063X and Y are independent random variables, each uniform on 0 to 1. What is the density of X + Y, and what is the probability that X + Y is less than 1.5?CitadelChicago · 2025
Try it first
What is P(X + Y < 1.5)?
Show the worked solution
The density is a triangle, f(s) = s for s up to 1 and 2 - s from 1 to 2, and P(X + Y < 1.5) = 7/8. Convolving two flat densities gives a tent peaking at 1. The part above 1.5 is a triangle with base 0.5 and height 0.5, area 1/8. In the unit square it is the same corner: the line x + y = 1.5 cuts off a triangle with legs of 0.5.
Why is the sum not uniform on 0 to 2?
Roll two dice: a total of 7 can be made six ways, a total of 12 only one way. Continuous uniforms behave the same. A sum near the middle can be made from many pairs, a sum near either end from very few, so the density of the sum rises to a peak and falls again. The mechanism that builds it is {term('convolution', 'The density of a sum of independent variables: for each possible total, add up the density of every pair of values that makes it.')}: the density at s is the length of the set of x values for which both x and s - x lie between 0 and 1.
In the unit square the line x + y = 1.5 cuts off a corner triangle of area 1/8, and in the triangular density of the sum the tail above 1.5 is the same 1/8, so X + Y is below 1.5 with probability 7/8. How does the convolution give the triangle?
Fix a total s. You need x between 0 and 1 and also s - x between 0 and 1, so x must lie between max(0, s - 1) and min(1, s). For s below 1 that interval has length s, and for s above 1 it has length 2 - s, so the density is a tent with its peak of 1 at s = 1. Check that the area is 1: a triangle with base 2 and height 1. The mean is 1 and the variance is 1/12 + 1/12 = 1/6, both of which you can read from symmetry and independence.
The relationshipf(s) the density of the sum at the value s the indicator 1 when s - x is a valid value of Y, otherwise 0 What it says in wordsThe density of the sum at s is how many ways of splitting s are allowed, which rises linearly to 1 and falls back.Why give the square picture as well?
It is a check that costs ten seconds. Because the pair is uniform on the unit square, any probability about X + Y is an area, and the event X + Y at least 1.5 is the corner triangle above the line x + y = 1.5. Its legs run from 0.5 to 1 on each axis, so its area is 1/8 and the answer is 7/8 again. The density gives you the whole distribution; the square gives you any single probability fast. Keep both, because the next question is usually three uniforms, where the density becomes piecewise quadratic and the square becomes a cube: P(X + Y + Z < 1) = 1/6.
Where candidates lose it
The common slip is 3/4, from assuming a sum of uniforms is uniform. Sums are never uniform unless one of the pieces is degenerate; they pile up in the middle.
The second loss is getting the convolution limits wrong and producing a density that does not integrate to 1. Write the two constraints on x out loud, and check the triangle's area before you use it.
What the interviewer asks next
- What is the density of X - Y?
- What is P(X + Y + Z < 1) for three independent uniforms?
- What is the expected value of max(X, Y), and of X + Y given that X + Y > 1?
Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):
He was asking some questions about the probability, especially on the convolution.
064You may roll a fair die up to three times. After each roll you either stop and are paid the face showing, or roll again; if you reach the third roll you must take it. What is your optimal stopping rule, and what is the game worth?RBC Capital MarketsToronto · 2025
Try it first
On the first roll you see a 4. What do you do?
Show the worked solution
Keep a 5 or 6 on the first roll, a 4, 5 or 6 on the second, and take whatever the third gives; the game is worth 14/3, about 4.67. Work backwards. The last roll is worth 3.5. With two rolls left, keep anything above 3.5: (4 + 5 + 6)/6 + (1/2)(3.5) = 4.25. With three, keep anything above 4.25: (5 + 6)/6 + (2/3)(4.25) = 14/3.
Why start from the last roll?
Think of house hunting with three viewings booked: whether to accept the first flat depends on what the remaining viewings are likely to offer, and you only know that once you know how you would behave at the last one. The value of continuing at any point is defined by what you would do later, so the only roll whose value you know outright is the last one, and every earlier decision is built on it. That is backward induction, and the interviewer wants to hear the words before any numbers.
Solving from the last roll upwards, the third roll is worth 3.5, so the second roll keeps 4, 5 or 6 and is worth 4.25, so the first roll keeps only 5 or 6 and the whole game is worth 14/3, about 4.67. How do the values build up?
On the last roll you take the face, worth 3.5. On the second roll, stop if the face beats 3.5, which means 4, 5 or 6; otherwise you get 3.5 from the last roll. The value is the average of the faces you keep plus the chance you continue times the value of continuing: (4 + 5 + 6)/6 + (3/6)(3.5) = 4.25. On the first roll, the bar to beat is now 4.25, so only 5 and 6 are kept: (5 + 6)/6 + (4/6)(4.25) = 11/6 + 17/6 = 14/3, about 4.667.
The relationshipV_n the value of the game with n rolls still available max(f, V_n) keep the face f if it beats rolling on, otherwise take the value of continuing What it says in wordsEach extra roll is worth the average of the better of the face and the value of carrying on.What does the common shortcut cost, and where does this lead?
The shortcut is to keep anything above the single-roll average of 3.5 at every stage. On the first roll that keeps a 4, which gives 4.625 instead of 4.667. The threshold rises with the number of rolls left, because each spare roll is an option, and an option is worth more the longer it lives. With six rolls the value is 5.27, and with many rolls it approaches 6, since you can wait for a six. The same structure prices an American option: exercise early only when the payoff beats the value of holding on.
Where candidates lose it
The trap is the fixed threshold: stopping on 4 at the first roll because 4 beats 3.5. It ignores that the comparison is with the value of continuing, which is 4.25 with two rolls left, not 3.5.
The second loss is computing forwards, trying to enumerate all paths from the first roll. Say backward induction, solve the last roll, and build up; three lines of arithmetic do the whole job.
What the interviewer asks next
- What is the game worth with four rolls?
- You now pay Rs 1 for each reroll. How does the rule change?
- If you are paid the square of the final face, what is the first-roll rule?
Asked at RBC Capital Markets, Quantitative Trading, Toronto, 2025 (Wall Street Oasis):
Best way to maximize EV across 3 chosen dice rolls (can choose to continue or not).
065The correlation between X and Y is 0.2, and the correlation between Y and Z is 0.5. What is the full range of possible values for the correlation between X and Z?Tower Research CapitalNew York · 2019
Try it first
Which statement about corr(X, Z) is right?
Show the worked solution
Anywhere from about -0.75 to 0.95. The correlation matrix must be positive semidefinite, which bounds the third correlation at 0.2 x 0.5 plus or minus sqrt((1 - 0.2^2)(1 - 0.5^2)), that is 0.1 plus or minus 0.849. Geometrically, correlations are cosines of angles: X sits 78.5 degrees from Y and Z sits 60 degrees from Y, so X and Z are between 18.5 and 138.5 degrees apart.
Why does knowing two correlations restrict the third at all?
Think of three towns on a map. If A is close to B and B is close to C, then A cannot be far from C; if B is only loosely near both, A and C could be almost anywhere. Correlations behave like distances in disguise: each is the cosine of the angle between two returns viewed as vectors, and angles obey a triangle rule. X makes an angle of {TH_A:.1f} degrees with Y, because its cosine is 0.2, and Z makes 60 degrees with Y. The angle between X and Z is therefore at least the difference and at most the sum.
With X at 78.5 degrees from Y and Z at 60 degrees from Y, the angle between X and Z ranges from 18.5 to 138.5 degrees, so their correlation can be anything from -0.75 to 0.95, centred on 0.1. How do you get the bound algebraically?
Any valid correlation matrix must be positive semidefiniteEvery portfolio built from the variables has a variance of zero or more; for a correlation matrix this means its determinant and all leading minors are non-negative., because a portfolio cannot have negative variance. For three variables with correlations a, b and c, the condition is 1 - a^2 - b^2 - c^2 + 2abc of at least zero. Treating that as a quadratic in c gives c = ab plus or minus sqrt((1 - a^2)(1 - b^2)), so the third correlation lies in an interval centred on the product of the other two. With a = 0.2 and b = 0.5 that is 0.1 plus or minus sqrt(0.72), or -0.7485 to 0.9485.
The relationshiprho_XY 0.2, the correlation of X and Y rho_YZ 0.5, the correlation of Y and Z the square root the room left over after the parts of X and Z explained by Y What it says in wordsThe third correlation is the product of the two given ones, plus or minus how much of X and Z is unexplained by Y.What does the centre value 0.1 mean, and when is the sign forced?
Split X and Z each into a part explained by Y and a leftover. The explained parts always contribute 0.2 x 0.5 = 0.1; the leftovers can be correlated however you like, and they can move the total by up to 0.849 either way. So 0.1 is the answer only if the leftovers are uncorrelated. The sign of corr(X, Z) is forced positive only when the two given correlations are strong, specifically when their squares add to more than 1; with 0.8 and 0.7 the range is about 0.13 to 0.99. On a desk this is why two hedges that each track an index only loosely say almost nothing about each other.
Where candidates lose it
The two fast wrong answers are 0.1, from multiplying, and must be positive, from assuming correlation is transitive. Both treat correlation like a chain of causes rather than a geometry.
The second loss is reaching the determinant condition and stalling on the algebra. Lead with the angle picture: arccos 0.2 is about 78.5 degrees, arccos 0.5 is 60, and the bounds are the cosines of their sum and difference.
What the interviewer asks next
- If corr(X, Y) = 0.8 and corr(Y, Z) = 0.7, can corr(X, Z) be negative?
- What is the most negative common correlation three variables can share?
- Given corr(X, Y) and corr(Y, Z), what value of corr(X, Z) makes X and Z uncorrelated once Y is controlled for?
Asked at Tower Research Capital, Prop Trading, New York, 2019 (Wall Street Oasis):
What if the correlation between X and Y is 0.2 and the correlation between Y and Z is 0.5.
066A 2-year bond pays a 6% annual coupon and trades at par. What are its Macaulay duration and modified duration?PIMCOLos Angeles · 2024
Try it first
Which is closest to the Macaulay duration?
Show the worked solution
Macaulay duration is 1.943 years and modified duration is 1.833. At par the yield equals the coupon, 6%. The year 1 coupon is worth 6 / 1.06 = 5.660 and the final payment 106 / 1.06^2 = 94.340, adding to 100. Weight each time by its share: 0.0566 x 1 + 0.9434 x 2 = 1.943. Divide by 1.06 for modified duration: a 1 point rise in yield cuts the price by about 1.83%.
What is Macaulay duration actually measuring?
Think of a seesaw with weights placed along it at the dates money arrives, each weight equal to what that payment is worth today. Macaulay duration is the point where that seesaw balances: the average time to your money, weighted by present value. A zero-coupon bond has all its weight at maturity, so its duration equals its maturity. Any coupon puts a little weight earlier and pulls the balance point forward. Here the coupon is small and the bond is short, so the pull is small, and that is the whole shape of the answer before any arithmetic.
The year 1 coupon is worth 5.66 today and the year 2 payment 94.34, so the present-value weighted average time is 1.943 years, just 0.057 years before maturity, and dividing by 1.06 gives a modified duration of 1.833. How do you get the numbers fast in your head?
At par, the yield is the coupon, so discount at 6%. The coupon's present value, 6 / 1.06, is about 5.66, and the rest of the price, 94.34, sits at year 2, so duration is 2 minus the coupon's weight: 2 - 0.0566 = 1.943. That shortcut works for any two-period bond: start at the maturity and subtract the early weight times the time it saves. For a par bond there is also a closed form, (1 + y)/y times (1 - 1/(1 + y)^n), which gives the same 1.9434.
The relationshipPV_t the present value of the cash flow at time t, discounted at the yield P the price, 100 at par y the yield per period, 6% What it says in wordsDuration is the average payment date weighted by value today; modified duration converts it into a price sensitivity.Why are there two durations, and how good is the estimate?
Macaulay duration is a time; modified duration is a slope. Modified duration is the percentage price change for a one point change in yield, so this bond loses about 1.83% if yields rise from 6% to 7%, about Rs 0.018 per Rs 100 for each basis point. Repricing exactly gives 98.192 at 7%, a fall of 1.808%, and 101.859 at 5%, a rise of 1.859%. The rise is bigger than the fall, which is convexityThe curvature of the price against yield: prices rise more when yields fall than they drop when yields rise by the same amount.; for a two-year bond it is a small correction, for a thirty-year bond it is not.
Where candidates lose it
The trap answers are 2 years, from forgetting the coupon, and 1.83 given as the Macaulay figure, from mixing the two definitions. Say which duration you are giving and in what units: years for Macaulay, percent per point of yield for modified.
The second slip is discounting at the wrong rate. A bond at par yields its coupon, so there is nothing to solve for; say that first and the arithmetic is two divisions.
What the interviewer asks next
- What is the duration of a 2-year zero-coupon bond, and of a 2-year bond with a 20% coupon?
- Estimate the price of this bond if yields rise 50 basis points.
- Why does duration fall as the coupon rises, holding maturity fixed?
Asked at PIMCO, Product & Strategy, Los Angeles, 2024 (Wall Street Oasis):
Lots of random bond math questions -- duration of this bond with x coupon sold at par
068A box holds ten coins: one has heads on both sides and nine are fair. You pick a coin at random, flip it five times and see five heads. What is the probability you picked the double-headed coin?Jump TradingChicago · 2018
Try it first
After five heads, roughly how likely is the double-headed coin?
Show the worked solution
32/41, about 78%. Before flipping, the odds are 1 to 9 against the double-headed coin. Five heads happen for certain with it and with chance 1/32 with a fair coin, a likelihood ratio of 32. Multiply: posterior odds are 32 to 9, which is 32/41. Each extra head doubles the odds, so a sixth head would take it to 64/73, about 88%.
Why is the answer not close to certain?
Imagine a rare illness and a decent test. A positive result raises the chance you have it, but if the illness is rare enough, most positives still come from healthy people. Evidence is weighed against how common each explanation was to begin with, so five heads, which a fair coin produces only once in 32 tries, still has to overcome nine fair coins for every double-headed one. The 97% instinct takes 1 minus 1/32 and forgets the nine-to-one start.
Starting from odds of 1 to 9, five heads multiply the odds by 32 to give 32 to 9, so the chance of the double-headed coin rises from 10% to 78%, roughly doubling the odds with each head. How do you run Bayes in odds form?
Odds form is the fastest way to say it in the room. Posterior odds equal prior odds times the likelihood ratio: (1 to 9) times 32 gives 32 to 9. Converting back, 32 out of 32 + 9 is 32/41, about 78%. The long form gives the same thing: the joint chance of picking the special coin and seeing five heads is 1/10, the joint chance of a fair coin and five heads is 9/10 x 1/32 = 9/320, and the posterior is (32/320) / (41/320).
The relationshipD the event that the double-headed coin was picked 5H the observation of five heads in five flips 32 the likelihood ratio: 1 divided by 1/32 What it says in wordsPrior odds of one to nine, multiplied by a likelihood ratio of thirty-two, give odds of thirty-two to nine.What does the odds picture tell you about more flips?
Each head is twice as likely under the double-headed coin, so every head doubles the odds and every tail ends the question, since the special coin never shows tails. After 0 to 5 heads the chance runs 10%, 18%, 31%, 47%, 64% and 78%; it passes 50% only after the fourth head. The useful follow-up is the next flip: it lands heads with chance 32/41 + (9/41)(1/2) = 73/82, about 89%. On a desk, the same arithmetic tells you how many winning days it takes before a new strategy's record says anything about skill.
Where candidates lose it
The trap is answering 31/32, about 97%, by looking only at how unlikely five heads are from a fair coin. That ignores the prior; with nine fair coins in the box, the base rate matters as much as the evidence.
The second slip is the reverse: staying near 10% because the coin was chosen at random. Say the odds form, prior times likelihood ratio, and both errors disappear.
What the interviewer asks next
- What is the probability that the next flip is heads?
- How many heads in a row would you need to be 99% sure?
- If one of the ten coins were double-tailed instead, how would five heads change the answer?
Asked at Jump Trading, Quantitative Research, Chicago, 2018 (Wall Street Oasis):
Then he asked one question of probability which can be solved by Bayesian formula.
069Daily returns are drawn from a normal with mean zero and standard deviation 1% on 90% of days, and from a normal with mean zero and standard deviation 4% on the other 10%. What are the overall standard deviation and the kurtosis of daily returns?Two SigmaNew York · 2025
Try it first
What is the kurtosis of the mixture?
Show the worked solution
The standard deviation is sqrt(2.5), about 1.58%, and the kurtosis is 12.72, against 3 for a normal. Moments of a mixture are weighted averages of the pieces' moments. The variance is 0.9 x 1 + 0.1 x 16 = 2.5. A normal's fourth moment is 3 sigma^4, so the fourth moment is 3 x (0.9 x 1 + 0.1 x 256) = 79.5. Kurtosis is 79.5 / 6.25 = 12.72.
Why does mixing two normals create fat tails?
Think of a city's daily traffic: most days are ordinary, and a few days a year there is a festival or a strike and everything is wild. Averaged across the year, the typical day looks calmer than the average suggests, and the extreme days are far more extreme than a single bell curve would allow. Mixing a calm regime with a rare wild one concentrates most days near zero and puts the rest far out, which is exactly what kurtosis measures: a high peak with heavy tails. Every piece is normal; the mixture is not.
With the same 1.58% standard deviation, the mixture is more peaked than the normal and its tail is far heavier: a daily move beyond 6% comes about 1 day in 75 under the mixture against about 1 day in 6,766 under the normal. How do you compute the moments without integrating?
Condition on the regime. Any moment of a mixture is the weighted average of that moment in each regime, because the density itself is the weighted average of the two densities. The second moment is 0.9 x 1^2 + 0.1 x 4^2 = 2.5, so the standard deviation is about 1.58%. For the fourth, use the fact that a normal with standard deviation sigma has fourth moment 3 sigma^4: 3 x (0.9 x 1 + 0.1 x 256) = 79.5. Kurtosis is the fourth moment over the variance squared.
The relationshipw_i the regime weights, 0.9 and 0.1 sigma_i the regime standard deviations, 1% and 4% kappa kurtosis, the fourth moment divided by the variance squared; 3 for any normal What it says in wordsAverage the variances and the fourth moments across regimes, then compare the fourth moment with the squared variance.What does the number mean for risk?
Kurtosis is 3 times the ratio of the average of sigma^4 to the square of the average of sigma^2, so any variation in volatility pushes kurtosis above 3, and the more uneven the regimes, the further it goes. The practical cost shows in the tails. A risk model that fits one normal with a 1.58% standard deviation expects a move beyond 6% about once in 6,766 days, roughly once every 27 years of trading; the mixture produces one about once in 75 days, several times a year. This is the simplest model of volatility clustering, and the reason daily returns on real assets show kurtosis well above 3.
Where candidates lose it
The trap is answering 3 because each piece is normal, or averaging the two standard deviations to 1.3% and treating the mixture as one normal. Mixtures average densities and moments, not shapes or standard deviations.
The second slip is forgetting the factor of 3 in a normal's fourth moment, or reporting excess kurtosis without saying so. Say which you mean: kurtosis 12.72, excess kurtosis 9.72.
What the interviewer asks next
- What mix of the two regimes maximises the kurtosis for a fixed overall variance?
- If the two regimes had different means, what would happen to the skew?
- How would you estimate the two regime volatilities from a year of daily returns?
Asked at Two Sigma, Quantitative Research, New York, 2025 (Wall Street Oasis):
They asked a couple questions involving Mixture Gaussians (e.g., probability density and moments).
072Towns A and B are 100 miles apart. A car leaves A for B at 50 mph. At the same moment a bird leaves B, flying towards the car at 100 mph; each time it meets the car it turns back to B, and each time it reaches B it turns towards the car again, until the car arrives at B. How far does the bird fly in total?BlackRockNew York · 2025
Try it first
How far does the bird fly?
Show the worked solution
200 miles. The car needs 100 / 50 = 2 hours to reach B, and the bird flies the whole time at 100 mph, so it covers 2 x 100 = 200 miles. Summing the zigzags gives the same answer: the first round trip is 133.3 miles, each later one is a third of the one before, and 133.3 / (1 - 1/3) = 200.
What is the question really asking you to count?
Think of a dog running back and forth between you and your front door while you walk home. You could trace every dash, or you could notice that the dog runs at a steady speed for exactly as long as your walk takes. Distance is speed times time, and the bird's flying time is fixed by the car, not by the zigzags, so the zigzag detail is a distraction. The car covers 100 miles at 50 mph in 2 hours; the bird flies at 100 mph for those same 2 hours. That is 200 miles, and it takes one sentence.
Plotted against time, the bird's zigzags shrink by a factor of three each round and all fit inside the car's 2-hour trip, so the bird flies for 2 hours at 100 mph, a total of 200 miles. How do you check it by summing the zigzags?
The bird and car close the first 100 miles at a combined 150 mph, so they meet after 40 minutes, 33.3 miles from A. The bird flies back to B, 66.7 miles, arriving at 80 minutes, by which time the car is at 66.7 miles. Each round trip starts with the gap to the car one third of the previous gap, so the round trips form a geometric series with ratio 1/3. The first is 133.3 miles; the sum is 133.3 / (1 - 1/3) = 200. It agrees, and it shows why infinitely many turns still add to a finite distance.
The relationshipv_bird the bird's speed, 100 mph t_car the car's travel time, 100 miles at 50 mph 133.3 the first round trip in miles, B to the first meeting and back What it says in wordsThe bird flies for exactly as long as the car drives; the zigzag series, summed, gives the same 200 miles.Why do interviewers still ask a puzzle this well known?
Because the way you answer tells them more than the answer. A candidate who starts summing legs has reached for the first method that fits; a candidate who asks what quantity is fixed has found the invariant, and that is the habit the interviewer is hiring. The story about von Neumann summing the series in his head is part of the folklore; you get more credit for the one-line method and the series as a check. The same move, looking for a quantity that does not depend on the messy path, solves many expected-value and stopping questions on this page.
Where candidates lose it
The trap is starting the series: solving for the first meeting, then the return, then the second meeting, and running out of time or making an arithmetic slip on the third leg. The infinite number of legs also tempts some candidates to answer infinity.
Lead with the time argument and give 200 within ten seconds; then offer the series with its ratio of one third as a check, which shows you could do it the long way.
What the interviewer asks next
- Where is the car when the bird reaches B for the second time?
- How many times does the bird turn around?
- If the bird started at A with the car, flying ahead to B and back, how far would it fly?
Asked at BlackRock, Quantitative Research, New York, 2025 (Wall Street Oasis):
A car starts at point A going 50 miles an hour towards point B
074A point is dropped uniformly at random in a unit square. What is the expected distance from the point to the nearest edge of the square?Hudson River TradingNew York · 2024
Try it first
What is the expected distance to the nearest edge?
Show the worked solution
1/6. Let D be the distance to the nearest edge. D exceeds d exactly when the point lies in the inner square of side 1 - 2d, so P(D > d) = (1 - 2d)^2 for d up to 1/2. The expected value of a non-negative variable is the integral of its tail, and the integral of (1 - 2d)^2 from 0 to 1/2 is 1/6.
Why work with the chance of being far rather than the distance itself?
Think of a sandpit where a child stands at a random spot and the question is how far they are from the nearest edge. Writing the distance as min(x, 1 - x, y, 1 - y) and integrating a minimum of four things means splitting the square into four triangles. Asking instead when the point is farther than d from every edge has a one-picture answer: the point must lie in a smaller square, shrunk by d on every side. That square has side 1 - 2d, so its area, (1 - 2d)^2, is the tail probability. One formula replaces four cases.
A point is more than d from every edge only inside the inner square of side 1 - 2d, so the chance of being farther than 0.1, 0.2, 0.3 and 0.4 is 0.64, 0.36, 0.16 and 0.04, and the area under that tail curve is the expected distance, 1/6. How does the tail give the expectation?
For any non-negative random variable, the expected value equals the integral of the chance that it exceeds each level, E[D] = integral of P(D > d). Here that is the integral of (1 - 2d)^2 from 0 to 1/2. Substitute u = 1 - 2d and it becomes half the integral of u^2 from 0 to 1, which is 1/6. A simulation with 200,000 random points gives 0.1669, against the exact 0.1667.
The relationshipD the distance from the random point to the nearest edge P(D > d) the area of the inner square of side 1 - 2d What it says in wordsAdd up the chance of being farther than each distance, and the total is the expected distance.How do you sanity-check 1/6 against simpler cases?
Build up the number of edges. The distance to one fixed edge averages 1/2, to the nearer of two opposite edges averages 1/4, and to the nearest of all four it falls to 1/6, so each added constraint pulls the minimum closer. The density of D is the slope of the tail, 4(1 - 2d), largest at the edge, which says most random points are near the boundary. That is the same reason most of the volume of a high-dimensional cube sits near its surface, a fact that matters when sampling scenarios in many risk factors at once.
Where candidates lose it
The trap answers are 1/2 and 1/4, from handling one edge or one axis and forgetting that the nearest of four edges is a minimum. A candidate who integrates min(x, 1 - x, y, 1 - y) directly often splits the square wrongly and lands on a different number.
Draw the inner square and say tail integral; the whole calculation is then one line.
What the interviewer asks next
- What is the expected distance to the nearest edge in a unit cube?
- What is the expected distance to the nearest corner of the square?
- What is the density of the distance to the nearest edge, and where is it highest?
Asked at Hudson River Trading, Campus Algo Dev Interview, New York, 2024 (Wall Street Oasis):
I was asked a expected value question involving the expected value among distance to an edge, with a randomly placed object.
