Derivatives Foundation puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 66
- Topics
- 12
- Hard
- 29
031A stock is at 1,000, volatility is 20% and rates are near zero. Estimate the three-month at-the-money call in your head, and the straddle.Market makingVolatility trading
Try it first
Say the call price before you reach for a formula.
Show the worked solution
Call about 40, straddle about 80. The rule is 0.4 x S x sigma x sqrt T. Three months is a quarter of a year, so sqrt T is 0.5 and the volatility over the period is 10%; 0.4 x 1,000 x 0.1 = 40. With rates at zero the at-the-money put is worth the same, so the straddle is 80. The full model gives 39.88 for the call, because the exact constant is 1 over sqrt(2 pi), 0.3989, not 0.4.
Where does the 0.4 come from?
A tailor who knows a customer's height is normally distributed around 170 cm with a spread of 10 cm can say how far above 170 the average tall customer stands: about 0.4 of the spread, which is 4 cm, because the mean of the positive half of a normal is sigma over sqrt(2 pi). An at-the-money call pays the positive half of the stock's move, and the average of the positive half of a normal is 0.4 of its standard deviation, so the call is worth 0.4 times the standard deviation of the move over its life. The standard deviation of the move is S x sigma x sqrt T, which is 1,000 x 0.2 x 0.5 = 100 here, and 0.4 of 100 is 40. The exact constant is 1 / sqrt(2 pi) = 0.3989, so the rule gives 40 where the precise version gives 39.89, and the model 39.88.
The relationshipS the stock price, 1,000 sigma sqrt T the volatility scaled to the option's life: 20% x 0.5 = 10% 0.4 the approximation to 1 over root 2 pi, which is 0.3989 N the standard normal distribution function What it says in wordsAn at-the-money call is about four tenths of one standard deviation of the stock's move over its life.Across volatilities from 0 to 100% the rule 0.4 x S x sigma x sqrt T sits almost on top of the full model price for a three-month at-the-money call, giving 40 against 39.88 at 20% volatility, and bends below it only at high volatility where the model's price curves. Why is the straddle just double, and when is it not?
With rates at zero and no dividends, the forward equals the spot, so an at-the-money call and put have the same value by put-call parity, and the straddle is simply two calls, about 80. The straddle costs 8% of the stock for a three-month bet, which is the whole quarter's one-standard-deviation move of 10% times 0.8: that is the number a volatility trader carries in their head. With rates or dividends the forward moves away from spot, at-the-money means at-the-forward, and the call and put split the straddle unevenly, though their sum barely changes. The rule also assumes the volatility is the right one for this strike, which on a real surface with skew it may not be.
How far can you push the rule?
The rule is linear in volatility and time, the model is not. For three months the gap is 0.12 on 40 at 20% volatility, and even at 60% volatility the rule gives 120 against 119.2. Over one year at 20% the rule gives 80 against the model's 79.66. Up to a total move of about 50% the rule is within a couple of percent, which is every interview case and most of the real book; beyond that the model price flattens because a call can never be worth more than the stock. Say the limitation and then use the rule anyway: on a desk the question is never whether 40 is exactly right, it is whether 44 on the screen is rich or cheap.
Where candidates lose it
The common loss is forgetting to scale the volatility to the horizon and quoting 0.4 x 1,000 x 0.2 = 80 for the call, which is the one-year number. Say sqrt T out loud: three months is a half.
The second is giving the call and then stalling on the straddle, or doubling the call without saying why. The put equals the call only because rates are zero and there is no dividend; name parity and the interviewer knows you understand what at the money means.
What the interviewer asks next
- Rates are now 8%. Which is worth more at the money, the call or the put, and by roughly how much?
- The stock pays a 2% dividend before expiry. What changes?
- Quote me the one-month straddle on the same stock, then the one-year.
- The market is paying 44 for the call. What volatility is it implying, roughly?
032One-month implied volatility is 20% and three-month implied volatility is 25%. What volatility is implied for the period from month one to month three?Volatility tradingEquity derivatives
Try it first
What is the forward volatility for months two and three?
Show the worked solution
About 27.2%. Variance is volatility squared times time, and variances add across periods. The three months carry 0.25 squared x 3 = 0.1875 of variance; the first month carries 0.20 squared x 1 = 0.04; so months two and three carry 0.1475 between them, 0.0737 per month on an annualised basis. The square root is 0.272. The forward volatility must sit above 25% because the quarter's average has to be pulled up from the 20% start.
Why do you add variances and not volatilities?
If you walk a random distance each hour, your spread after three hours is not three times the hourly spread, because some hours cancel others. What grows in a straight line is the variance, the spread squared. Implied volatility is quoted per year, so the variance an option carries is volatility squared times its life, and a longer option's variance is the sum of the variances of the pieces of time inside it. That is the whole mechanism. Twenty per cent for one month is 0.04 of variance; twenty-five per cent for three months is 0.1875; the difference belongs to the two months in between, and dividing by two months and taking the square root turns it back into a volatility.
Drawn as area, the first month's variance of 0.04 plus the two forward months' variance of 0.0737 each must fill the quarter's 0.1875, so the forward box has height 0.0737 and its square root is a forward volatility of 27.2%. The relationshipsigma 3, sigma 1 the three-month and one-month implied volatilities, 25% and 20% T the option life in months; the units cancel as long as both use the same one sigma 1,3 the forward volatility for the period between the two expiries What it says in wordsSubtract the short period's variance from the long period's and spread what is left over the time in between.What does the number tell a trader?
The 27.2% is the volatility you would lock in for months two and three by selling the one-month option and buying the three-month one in variance-weighted sizes, a calendar spread. If you believe realised volatility in those two months will be well below 27.2%, the three-month option is rich relative to the one-month, and the calendar is the trade that expresses it. Say the rounding: an upward sloping term structure, 20% then 25%, hides a forward that is steeper than either quote, 27.2%, and the trap of reading 25% as the forward is what the question is built to catch. The calculation uses calendar time in months; a desk would use trading days or variance-weighted business days, which shifts the number by a few tenths.
What is the limit of the calculation?
The subtraction has to leave something positive. If the three-month volatility were below 11.5% with the one-month still at 20%, the quarter would carry less variance than its first month alone, which is impossible without arbitrage: you could sell the one-month, buy the three-month and hold a position with negative forward variance. Forward variance can be small but never negative, so a term structure that inverts too sharply is a mispricing, not a forecast. The second limit is that both quotes must refer to the same strike in forward terms; mixing an at-the-money one-month with a three-month that has rolled away from the money adds skew to the comparison and the forward you compute is no longer clean.
Where candidates lose it
The common loss is averaging volatilities: 25% over three months with 20% in the first month gives 27.5% for the rest if you treat volatility as additive. It is close, which is why it survives, but it is the wrong quantity and on a steeper curve the gap is large.
The second is forgetting to weight by time: subtracting 0.04 from 0.0625 and taking the root. Write variance times time for each leg, every time.
What the interviewer asks next
- Six-month volatility is also 25%. What is the forward volatility for months four to six?
- The one-month is 30% and the three-month is 25%. What is the forward, and what does it say about the market?
- How would you actually lock in that forward volatility with listed options, and what would break the hedge?
- Why do desks compute this in trading days rather than calendar months?
033We play chess repeatedly. Half the games are draws; of the decisive games I win two thirds. The match ends when someone wins three games in a row, and a draw resets both streaks. What is the probability that I win the match?Old Mission CapitalChicago · 2018
Try it first
Before setting anything up: roughly how likely am I to win the match?
Show the worked solution
86/99, about 86.9%. Per game I win with probability 1/3, you win with 1/6 and we draw with 1/2. The match has five live states: no streak, my streak of one or two, your streak of one or two. Writing my chance of winning the match from each state as an unknown, each state's equation is a weighted average of its neighbours, and solving the five equations gives 86/99 from the start. The naive ratio of (1/3)^3 to (1/6)^3 gives 89% and is wrong.
Why does the match need states rather than a single formula?
A tennis game at deuce is the everyday version: whoever is a point ahead is in a different position from level, and the chance of winning the game from deuce is best found by naming the positions and linking them. What matters here is not the game count but the current streak, and only five positions are possible before the match ends: no streak, me on one, me on two, you on one, you on two. Every game moves the match from one of those positions to another, with the same three probabilities each time, so the match is a Markov chain and the answer is a small linear system rather than a series. Three in a row sounds like it needs a long sum over all the ways the match can go; the states collapse that sum into five unknowns.
From the no-streak start my chance of winning the match is 86/99, about 86.9%; on my streak of one or two it rises to 87.9% and 90.9%, on your streak of one or two it falls to 84.8% and 72.7%, and every draw returns the match to the start. How do you write and solve the equations?
Call my winning chance x from no streak, a1 and a2 from my streaks, b1 and b2 from yours. From any state a draw, probability 1/2, takes you to x. From no streak a win takes you to a1 and a loss to b1, so x = x/2 + a1/3 + b1/6. From a1 a win takes you to a2 and a loss to b1. From a2 a win ends the match in my favour, worth 1. From b1 a loss takes you to b2 and a win takes you to a1; from b2 a loss ends it, worth 0. Five equations in five unknowns, and the structure is friendly: substitute the draw term first, since x/2 appears everywhere, and the system reduces by hand in a few lines. The solution is x = 86/99, a1 = 29/33, a2 = 10/11, b1 = 28/33, b2 = 8/11. A simulation of 200,000 matches gives 0.869, which confirms the fraction.
The relationshipx my chance of winning the match with no streak live a1, a2 my chance when I have won one or two in a row b1, b2 my chance when you have won one or two in a row 1/2, 1/3, 1/6 the per-game chances of a draw, my win and your win What it says in wordsEach state's value is the average of the values of where the next game can send it, weighted by the chance of each result.Why is the naive ratio wrong, and in which direction?
The tempting shortcut compares the chance of three straight wins for me, (1/3)^3, with three straight for you, (1/6)^3, and takes my share: 8 over 9, 88.9%. That treats the match as a single race from scratch, but a broken streak is not a reset to equal footing: when you beat me on my streak of two, you start a streak of one, and the shortcut ignores every such hand-over. Those hand-overs favour the weaker player a little, which is why the true 86.9% sits below 88.9%. It is also worth saying that the draws change nothing about who wins: they only lengthen the match, which lasts about 33.9 games on average, because every draw sends both streaks back to zero.
Where candidates lose it
The common loss is the ratio shortcut, (1/3)^3 against (1/6)^3, which gives 8/9. It is close enough to sound right and the interviewer will ask you to defend it, at which point the missing hand-over of streaks becomes obvious.
The second is setting up too many states, tracking game counts or draw counts. Only the current streak matters. Five states, five equations, and the draw term is the same in every one.
What the interviewer asks next
- How long does the match last on average?
- The match now ends at two in a row. Does my chance go up or down, and why?
- Draws no longer reset the streaks, they are simply ignored. What is my chance now?
- Write the transition matrix and show which states are absorbing.
Asked at Old Mission Capital, Prop Trading, Chicago, 2018 (Wall Street Oasis):
You and I play chess. 1/2 games end in draws and in the other half I win with 2/3 probability and you win with 1/3
034A card is drawn from a 52-card deck. You may bet Rs 100 at even money on red or on black. How much would you pay to be told first whether the card is a heart?OptiverChicago · 2025
Try it first
What is the tip worth, at most?
Show the worked solution
Up to Rs 50. Without the tip the bet is a coin flip worth zero. With it: one time in four the card is a heart, you bet red and win Rs 100 for certain. Three times in four it is not a heart, which leaves 13 red and 26 black cards, so you bet black and win 100 with probability 2/3 and lose 100 with probability 1/3, an expected Rs 33.33. The average, 25 + 25, is Rs 50, and that is the most the information is worth.
Why is a tip about hearts worth anything for a bet on colour?
A friend who will tell you whether it is raining in one of the four districts of a city is not telling you the weather everywhere, but if you have to bet on whether it rains in the city at all, the tip changes your odds. The heart tip never names the colour, but every heart is red, so a yes makes red certain and a no tilts the remaining deck two to one towards black. Both answers leave you with a bet that has an edge, where before you had none. Information is worth the gap between what you can earn with it and what you could earn without it, and here that gap is the whole value, because without the tip the best you can do is zero.
Without the tip a bet on red or black has an expected value of zero, while with it a heart lets you win Rs 100 for certain and a non-heart lets you bet black at 26 to 13 for an expected Rs 33.33, so the tip is worth a quarter of 100 plus three quarters of 33.33, which is Rs 50. How do you set the calculation out so it cannot go wrong?
Price the decision in each branch, then weight the branches by how likely each answer is. If the answer is yes, probability 13/52, the card is red: bet red, expected gain 100. If the answer is no, probability 39/52, there are 39 cards left of which 13 are red diamonds and 26 are black: bet black, expected gain 100 x 26/39 minus 100 x 13/39, which is 33.33. The value of the tip is the probability-weighted average of the best you can do after each answer, minus the best you could do with no answer at all. That is 1/4 x 100 + 3/4 x 33.33 = 50, minus zero. Pay less than Rs 50 and the deal is in your favour; pay exactly 50 and you are indifferent.
The relationship13/52 the chance the card is a heart 100 a certain win on red once you know it is a heart 26/39, 13/39 the chance of black and of red among the 39 non-hearts 0 the value of the bet with no information What it says in wordsWeight the best decision after each possible answer by the chance of that answer, and subtract what the bet was worth before.What does this have to do with a trading desk?
The question is a small model of paying for data. Being told the colour outright would be worth Rs 100, the full value of perfect information; the heart tip, which answers a narrower question, is worth exactly half of that. A piece of information is worth what it changes in your best decision, not how interesting it sounds: a tip that the card is an ace would be worth nothing here, because it leaves red and black at even money. The limitation to say out loud is that the Rs 50 is an expected value: on any single hand you could pay 50 and lose 100, so a desk with a limited bankroll should pay less than the full value, by the same logic that sizes bets below the edge.
Where candidates lose it
The common loss is answering 25 by reasoning that hearts are a quarter of the deck, so the tip is right a quarter of the time. The tip is informative in both of its answers: a no still tilts the deck two to one.
The second is pricing only the yes branch and forgetting to re-count the deck after a no. Thirty-nine cards remain, thirteen of them red, and that recount is where the second Rs 25 lives.
What the interviewer asks next
- How much is it worth to be told whether the card is an ace?
- How much is it worth to be told whether the card is a face card or a heart?
- You can bet Rs 100 on the suit at 3 to 1 instead. What is the heart tip worth now?
- Why should a desk with a small bankroll pay less than Rs 50 for this tip?
Asked at Optiver, Quantitative Research, Chicago, 2025 (Wall Street Oasis):
Valuing information, taking directional bets when not plus EV.
035There are 100 coins on the table. Players take turns removing 1 to 10 coins, and whoever takes the last coin wins. Do you want to go first, and what is your first move?Quant trading
Try it first
Go first or second, and what is the opening?
Show the worked solution
Go first and take 1, leaving 99. Work backwards: whoever faces 11 coins loses, because any take of 1 to 10 leaves 1 to 10 for the other player to finish. The same holds for 22, 33 and every multiple of 11. From 100, taking 1 leaves 99, a multiple of 11; after that, whatever the opponent takes, you take 11 minus it, stepping down 88, 77, 66 and so on to 0, where you take the last coin.
Why work backwards from the last coin?
If you are climbing stairs with a friend and the rule is that the person who steps onto the top stair wins, you do not plan from the bottom; you ask which stair you must leave your friend on so that they cannot reach the top in one go. Games with a fixed last move are solved from the end: find the positions where the player to move loses, then find the positions from which you can push your opponent onto one of them. With 1 to 10 coins allowed, facing 1 to 10 coins is a win, you take them all. Facing 11 is a loss, because every move leaves between 1 and 10. Facing 12 to 21 is a win, since you can reduce to 11. Facing 22 is a loss again. The losing positions repeat every 11.
Every multiple of 11 from 0 to 99 is a losing position for the player who must move, so the first player takes 1 to leave 99 and then answers every take of t with 11 minus t, stepping down through 88, 77 and 66 until the last coin. How do you find the period without listing every position?
The period is the largest take plus one, 11, because that is the one total a pair of moves can always be made to add up to: whatever your opponent takes between 1 and 10, you can take the balance of 11. The losing positions are the multiples of the largest take plus one, and the winning opening move is the remainder when the pile is divided by that number. 100 divided by 11 is 9 remainder 1, so take 1. If the rule allowed 1 to 7 coins, the period would be 8 and the opening would be 100 mod 8, which is 4. If the pile had been 99 to start with, you would want to go second, because the first player cannot leave a multiple of 11.
The relationship11 the largest allowed take plus one, the amount you can always complete in a pair of moves 100 mod 11 the remainder when 100 is divided by 11; take exactly this many What it says in wordsTake the remainder on your first move, then keep each pair of moves summing to 11.What changes if the last coin loses instead of wins?
Then you want to hand your opponent the last coin, so the position you avoid facing is 1 coin, and the losing positions shift up by one: 1, 12, 23 and so on up to 100. Facing 100 in that version you are already lost, so you would want to go second, which shows the interviewer that you re-derive the pattern rather than remember it. The method is the same in every variant: name the terminal position, step back one move at a time to find the first losing position, then find the period. The limitation of the trick is that it needs a game with perfect information and no chance; add a die that sets each turn's maximum and the clean period disappears.
Where candidates lose it
The common loss is taking 10, because a bigger move feels like a stronger start. It leaves 90, which is not a multiple of 11, and a prepared opponent takes 2 to leave 88 and wins from there.
The second is knowing the answer and not the reason. Say why 11 is the period: any take of 1 to 10 can be completed to 11. Without that sentence the interviewer will change the numbers and watch you stall.
What the interviewer asks next
- Players may take 1 to 7 coins instead. Do you go first, and what is the opening?
- The player who takes the last coin loses. Do you go first?
- There are two piles, 100 and 60, and you may take from either pile. Who wins?
- Each turn a die sets the maximum take. Is there still a strategy, and what is it?
036Three assets all have the same pairwise correlation rho. What are the eigenvalues of the correlation matrix, and how low can rho go?Jump TradingPudong Xinqu · 2023
Try it first
How negative can the common correlation be?
Show the worked solution
The eigenvalues are 1 + 2 rho, once, and 1 - rho, twice; rho can go no lower than -1/2. The vector (1, 1, 1) is an eigenvector with eigenvalue 1 + 2 rho, and any vector whose entries sum to zero is an eigenvector with eigenvalue 1 - rho, which gives a two-dimensional space and hence a double root. A correlation matrix must have no negative eigenvalue, since each eigenvalue is a portfolio variance, so 1 + 2 rho is at least 0 and rho is at least -1/2.
Why can three assets not all be strongly negatively correlated?
Three friends cannot all sit opposite each other at a table: if A faces B and B faces C, then A and C are on the same side. Correlation has the same constraint. If asset A moves against B and B moves against C, A and C are pushed towards moving together, so there is a floor on how negative a common correlation can be, and for three assets that floor is -1/2. You can see the floor without any algebra by holding an equal-weight portfolio of the three, each with unit variance: its variance is (3 + 6 rho) / 9, which is (1 + 2 rho) / 3, and a variance cannot be negative. At rho = -1/2 the portfolio has zero variance; it is perfectly hedged, and nothing below that is possible.
Plotted against rho, the eigenvalue 1 + 2 rho rises steeply and crosses zero at rho = -1/2, while the double eigenvalue 1 - rho falls gently to zero at rho = 1, so the matrix is a valid correlation matrix only between those two points. How do you find the eigenvalues without expanding a determinant?
Write the matrix as (1 - rho) times the identity plus rho times the all-ones matrix J. The identity leaves every vector alone, so you only need the eigenvalues of J, and J is easy: it maps (1, 1, 1) to (3, 3, 3), eigenvalue 3, and it maps any vector whose entries sum to zero to the zero vector, eigenvalue 0, with a two-dimensional space of such vectors. Shifting and scaling by (1 - rho) turns those into 1 - rho + 3 rho = 1 + 2 rho for the market direction and 1 - rho for the two spread directions. Check with the trace: the eigenvalues add to 1 + 2 rho + 2(1 - rho) = 3, the sum of the diagonal, as they must. At rho = 0.3 they are 1.6, 0.7 and 0.7.
The relationshipI the identity matrix J the matrix of all ones, whose eigenvalues are 3 (once) and 0 (twice) lambda 1 the eigenvalue of the common or market direction lambda 2, 3 the double eigenvalue of the two directions that net to zero, the spread trades What it says in wordsThe matrix is a stretch of the all-ones matrix, so the market direction gets 1 + 2 rho and every spread direction gets 1 - rho.What does the structure tell a risk or trading desk?
The eigenvectors are the principal components. The (1, 1, 1) direction is the market factor, and its eigenvalue over the trace, (1 + 2 rho) / 3, is the share of total variance it explains: 53% at rho = 0.3 and 80% at rho = 0.7. The two spread directions carry the rest, equally. A long-short book that nets to zero across the three assets lives entirely in the 1 - rho directions, which is why pairs trades get calmer as correlation rises and why a correlation of 1 collapses them to nothing. For n assets the same argument gives eigenvalues 1 + (n - 1) rho and 1 - rho, so the floor is -1 / (n - 1): -1/3 for four assets and -1/9 for ten. Say the limitation as well: a historical correlation matrix estimated from more assets than observations is only barely positive semi-definite, and a hand-edited one, where a trader overrides a few pairs, can fail the test entirely, which is exactly the fault a risk system is built to catch.
Where candidates lose it
The common loss is answering -1, because a correlation can be -1. For a pair it can; for three assets pairwise, it cannot, and the interviewer wants the reason: a negative eigenvalue is a negative portfolio variance.
The second is expanding the characteristic polynomial by hand and getting lost. Spot the all-ones structure, name the eigenvector (1, 1, 1), and the rest is one line.
What the interviewer asks next
- What is the floor on rho for n equally correlated assets?
- The three assets have correlations 0.9, 0.9 and -0.9. Is that a valid correlation matrix?
- What is the variance of the equal-weight portfolio at rho = -1/2, and what does that portfolio look like?
- How would you repair an estimated correlation matrix that has a small negative eigenvalue?
Asked at Jump Trading, Prop Trading, Pudong Xinqu, 2023 (Wall Street Oasis):
Some very difficult linear algebra questions about PCA and eigenvalues
037X and Y are independent and uniform on 0 to 1. What is the density of X + Y, and what is the probability that X + Y is less than 1.5?CitadelChicago · 2025Citadel SecuritiesChicago · 2025
Try it first
What is P(X + Y < 1.5)?
Show the worked solution
The density is a triangle: f(s) = s for s between 0 and 1, and 2 - s for s between 1 and 2; P(X + Y < 1.5) = 7/8. The sum lands in a thin band near s only if X and Y fit inside the unit square along the line X + Y = s, and the length of that line segment grows until s = 1 and then shrinks. The sum exceeds 1.5 only in the top-right corner of the square, a triangle with legs 1/2 and area 1/8, so the probability of less than 1.5 is 0.875.
Why is the sum of two flat distributions a triangle?
Roll two ordinary dice: a total of 7 can be made six ways, a total of 2 or 12 only one way, so the totals pile up in the middle even though each die is flat. The sum of two independent uniforms is the continuous version: the density at s is the length of the diagonal line X + Y = s inside the unit square, which grows from 0 up to s = 1 and shrinks back to 0 at s = 2. Formally the density is the convolution, the integral over x of f(x) times g(s - x), and with both densities equal to 1 on the unit interval the integrand is 1 exactly where both x and s - x lie between 0 and 1. That interval has length s for s below 1 and length 2 - s above, which is the triangle.
The relationshipf X+Y the density of the sum the indicator equals 1 where both x and s - x lie in the unit interval, which is where both densities are non-zero the integral from 1.5 to 2 the area of the small triangle in the tail, 1/8 What it says in wordsConvolve two flat densities and the overlap of the two intervals gives a triangle; the tail beyond 1.5 is one eighth of it.On the unit square the event X + Y > 1.5 is the corner triangle with legs one half and area one eighth, and on the density plot the same one eighth is the tail of the triangle beyond s = 1.5, so P(X + Y < 1.5) is seven eighths. How do you get 7/8 in your head?
Use the square, not the integral. The pair (X, Y) is a uniformly random point in the unit square, so any probability is an area. X + Y > 1.5 is the region above the line from (0.5, 1) to (1, 0.5), a right triangle with both legs equal to 1/2, whose area is half of a quarter, which is 1/8. Everything else is 7/8. The same picture answers the mirror question at once: X + Y < 0.5 is the bottom-left corner, also 1/8. And it checks the density: integrating 2 - s from 1.5 to 2 gives (0.5)^2 / 2 = 1/8, the same number from the other direction, which is what you say to show the two views agree.
What does the interviewer ask next, and why?
The follow-ups test whether you understood the mechanism or memorised the triangle. Add a third uniform and the density becomes three parabolic pieces, smooth at the joins, with P(X + Y + Z < 1) = 1/6, the volume of the corner tetrahedron, and in general the sum of n uniforms heads towards a normal with mean n/2 and variance n/12, which is the central limit theorem arriving. The mean of X + Y is 1 and the variance is 1/12 + 1/12 = 1/6, by independence. Each extra convolution smooths the density by one degree: flat to triangular to parabolic, which is the general fact that sums of independent variables are smoother than their parts. The limitation worth saying: this all relies on independence; with X and Y correlated, the point is no longer uniform over the square and the areas no longer give probabilities.
Where candidates lose it
The common loss is treating the sum as uniform on 0 to 2 and answering 0.75. The sum is a triangle, and the mass in the top quarter of its range is only 1/8, not 1/4.
The second is setting up the convolution integral correctly and then mishandling the limits, integrating over the wrong interval of x. Draw the square; the limits are the overlap of two unit intervals, and the picture makes them obvious.
What the interviewer asks next
- What is P(X + Y + Z < 1) for three independent uniforms?
- What is the density of X - Y, and of max(X, Y)?
- What are the mean and variance of X + Y, and how do you get them without the density?
- X and Y are now correlated with rho = 0.5. Can you still find P(X + Y < 1.5) by area?
Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):
He was asking some questions about the probability, especially on the convolution.
Asked at Citadel Securities, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):He was asking some questions about the probability, especially on the convolution.
038Box A holds 3 red and 1 black ball, box B holds 1 red and 3 black. You pick a box at random and draw a red ball. You then draw again from the same box without replacement. What is the probability the second ball is red?Prop trading firmsMarket making
Try it first
Pick the probability of a second red before you compute.
Show the worked solution
Exactly 1/2. The first red is evidence: box A produces a red with probability 3/4, box B with 1/4, so after seeing red the odds are 3 to 1 for A, a posterior of 3/4. Then update the contents. Box A has 2 red of 3 left, so a second red has probability 2/3; box B has 0 red of 3, probability 0. Average over the updated boxes: 3/4 x 2/3 + 1/4 x 0 = 1/2. Forgetting the update gives 1/3, which is wrong.
Why does the first red change which box you think you hold?
If a friend's cooking is spicy three times out of four and yours one time in four, and a dish picked at random from one of you turns out spicy, you would guess it was your friend's, and you would be right three times in four. The first red ball is the same clue: it comes from box A three times as easily as from box B, so the box you are holding is A with probability 3/4, not 1/2. The arithmetic is Bayes. The weight for A is 1/2 x 3/4 = 3/8, for B it is 1/2 x 1/4 = 1/8, and A's share of the total 1/2 is 3/4. The question is built so that candidates who skip this update get a clean wrong answer.
After a red first draw the box is A with posterior 3/4 and B with 1/4, box A has 2 red of 3 left and box B none, so the chance of a second red is 3/4 x 2/3 plus 1/4 x 0, which is exactly one half. What has to be updated besides the belief about the box?
The box itself. One red ball has left it. Box A now holds 2 red and 1 black, so a red is 2/3; box B now holds 0 red and 3 black, so a red is impossible. Both updates matter and they pull in opposite directions: believing more in A pushes the answer up, removing a red from A pushes it down, and here they land on exactly 1/2. Say that out loud, then check it by counting. Imagine 8 rounds, 4 with each box. Box A yields a first red in 3 of its 4 rounds and box B in 1 of its 4: four red first draws in all. In the three from A the second draw is red 2/3 of the time, 2 rounds; in the one from B, never. Two of four is 1/2.
The relationshipR1, R2 the first and second draws are red P(A | R1) the posterior that the box is A after one red, 3/4 P(R2 | A, R1) a second red from box A with one red already gone, 2/3 What it says in wordsWeight each box's chance of a second red by how likely that box is after the first red.How do the variants separate candidates who understand from those who memorised?
Put the first ball back and the box contents do not change but the belief still does: 3/4 x 3/4 + 1/4 x 1/4 = 5/8, which is above 1/2 because the posterior leans to A and A stays rich in red. Forget the belief update and keep the contents update, and you get 1/2 x 2/3 + 1/2 x 0 = 1/3. The gap between 1/3, 1/2 and 5/8 is the whole lesson: the first draw tells you about the box and changes the box, and you need both. If a second red does appear, the posterior for A becomes 1, since B has no reds left, and a third red is then 1/2. The limitation is the usual one for Bayes: the 1/2 prior for each box is given here, and in a trading setting the prior is the thing you are least sure of.
Where candidates lose it
The common loss is 1/3: keeping the 1/2 prior on each box and only updating the contents. The first red is evidence about the box, and leaving it out throws away half the information in the question.
The second is 2/3: updating to box A and then forgetting that B is still possible with probability 1/4, or forgetting that A has lost a red. Carry both boxes through to the end, then average.
What the interviewer asks next
- The first ball is put back before the second draw. What is the probability now?
- The second ball is also red. What is the probability the box is A, and that a third draw is red?
- Box B held 0 red and 4 black instead. What changes, and what is the answer?
- A trader sees one winning trade from a new signal. What is the analogue of the box update?
039You start with Rs 2 and bet Rs 1 at a time on a coin that falls your way 60% of the time. You stop when you reach Rs 5 or go broke. What is the probability you reach Rs 5?Two SigmaNew York · 2023
Try it first
Roughly how likely are you to reach Rs 5 before going broke?
Show the worked solution
135/211, about 64.0%. Let r be q over p, which is 0.4 / 0.6 = 2/3. The probability of reaching N from a stake of i is (1 - r^i) / (1 - r^N). With i = 2 and N = 5 that is (1 - 4/9) / (1 - 32/243) = (5/9) x (243/211) = 135/211. A fair coin would give 2/5 = 40%; the 60% edge lifts it to 64.0%. The game lasts about 6.0 bets on average.
Why is the answer not simply 2 out of 5?
With a fair coin the answer is 2/5, because a fair game cannot create or destroy expected money: you start with Rs 2, you finish with Rs 5 or Rs 0, so the chance of Rs 5 must be 2/5 to keep the average at 2. With a 60% coin each bet gains you Rs 0.20 on average, so the walk drifts upward and the chance of hitting the top is higher than the fair-coin fraction; what you need is a quantity that is still conserved under the biased coin. That quantity is (q/p) to the power of your stake. A win multiplies it by q/p, a loss by p/q, and weighted by their probabilities the two moves cancel: p x (q/p) + q x (p/q) = q + p = 1. Because that quantity is conserved, its starting value must equal its average finishing value, and that one line gives the formula.
The relationshipr the loss probability over the win probability; below 1 when the coin favours you i the starting stake, Rs 2 N the target, Rs 5 P i the chance of reaching the target before going broke What it says in wordsSet the conserved quantity r to the stake equal to its average at the end, and solve for the chance of reaching the target.The chance of reaching Rs 5 before Rs 0 rises along a curve above the fair-coin straight line, reaching 0.640 from a starting stake of Rs 2 against 0.4 for a fair coin, because each rupee of stake multiplies the odds of ruin by q over p, two thirds. How do you derive it from the states if you forget the formula?
Write P_i for the chance of reaching 5 from a stake of i. Then P_0 = 0, P_5 = 1, and in between P_i = 0.6 P_(i+1) + 0.4 P_(i-1), one equation per state. That is a second-order linear recurrence whose solutions are of the form A + B r^i with r = q/p, and the two boundary conditions fix A and B. Solving the five equations directly gives P_1 = 81/211, P_2 = 135/211, P_3 = 171/211 and P_4 = 195/211, and the recurrence is the thing to write on the whiteboard first, because it works for any rule change. A simulation of 200,000 games gives 0.639, agreeing with the fraction to three places. The same system with a 1 on the right-hand side of each interior equation gives the expected duration, about 6.0 bets from Rs 2.
What does the biased formula tell you about trading with an edge?
Let the target go to infinity. With a fair coin the chance of never going broke is zero: any finite stake is eventually lost. With the 60% coin it is 1 minus r to the stake, which from Rs 2 is 1 - 4/9 = 5/9, about 56%, and from Rs 10 it is above 98%. An edge does not protect a thin stake: with Rs 2 behind a 60% coin you still go broke 44% of the time, and the cure is not a better coin but a bigger stake relative to the bet. That is why a desk with a genuine edge still caps position size, and why the question sits next to the Kelly one. The limitation is that the bets here are of fixed size; once you can resize the bet with your capital, the ruin arithmetic changes completely.
Where candidates lose it
The common loss is answering 2/5, the fair-coin answer, or guessing that 60% means roughly 60%. The edge changes the structure, and the interviewer wants to hear q over p.
The second is writing the formula with p/q instead of q/p, which gives a number below 40% for a coin that favours you. Sanity check the direction: an edge in your favour must raise the chance above the fair-coin 2/5.
What the interviewer asks next
- What is the chance of reaching Rs 5 from Rs 2 with a fair coin, and why is it exactly 2/5?
- The target is Rs 10 instead of Rs 5. What is the chance now?
- There is no target: you play until you go broke or forever. What is the chance you never go broke?
- How long does the game last on average from Rs 2?
Asked at Two Sigma, Research, New York, 2023 (Wall Street Oasis):
Biased gamblers ruin problems; Markov Chain problems; sampling uniformly from triangle
040Five rational pirates, A the most senior down to E, split 100 gold coins. The most senior proposes a split; it passes if at least half of the pirates, his own vote included, agree, otherwise he is thrown overboard and the next most senior proposes. Each pirate wants to survive, then to get the most coins, then to see others thrown overboard. How should A split the coins to survive and keep the most?Old Mission CapitalNew York · 2014
Try it first
How many coins can A keep?
Show the worked solution
A proposes 98 for himself, 0 for B, 1 for C, 0 for D and 1 for E. Work backwards. With two pirates, D keeps 100, since his own vote is half. With three, C keeps 99 and gives E 1, because E gets nothing if C dies. With four, B keeps 99 and gives D 1. With five, A needs two other votes and buys the cheapest: C and E, who get nothing under B's plan, accept one coin each. A's vote plus two makes three of five.
Why does the answer come from the end of the game, not the start?
When you negotiate a price you think about what the other side does if you walk away; what they do next determines what they will accept now. Each pirate compares the offer in front of him with what he would get under the next proposal, so to know what A must offer you first need to know what B would do, which needs C, which needs D with two pirates left. With two pirates, D proposes 100 for himself; his own vote is one of two, which is half, so it passes and E gets nothing. That is the anchor. Every earlier step is a question of who gets zero in the next round, because those are the pirates whose votes are cheapest to buy.
Working up from two pirates, D keeps 100, C keeps 99 by paying E one coin, B keeps 99 by paying D one coin, and A keeps 98 by paying C and E one coin each, the two pirates who would receive nothing under B's proposal. How does each proposer decide whose vote to buy?
He buys exactly the number of extra votes he needs, from the pirates who are cheapest, and a pirate is cheap if the next proposal gives him nothing. With three pirates, C needs one more vote; under D's plan E gets zero, so C offers E one coin and keeps 99, and E accepts because one is better than zero. With four, B needs one more; under C's plan D gets zero, so B pays D one coin. With five, A needs two more votes, and under B's plan the pirates with nothing are C and E, so A pays them one coin each and keeps 98; offering anything to B or D is wasted, because B would get 99 and D one coin without A, and neither can be bought for less. The alternation is the pattern to say out loud: each round, the pirates who were paid last time are the ones left out this time.
The relationshipeach tuple coins to the proposer first, then to the others in order of seniority the 1s votes bought for one coin from pirates who would get zero in the next round votes needed half of the pirates, rounded up, including the proposer: 1 of 2, 2 of 3, 2 of 4, 3 of 5 What it says in wordsEach proposer keeps everything except one coin for each extra vote he needs, bought from whoever the next round would leave empty-handed.What assumptions is the answer resting on, and what happens when they change?
Three assumptions, and the interviewer will test at least one. Pirates are perfectly rational and know the others are. A pirate who is indifferent between two outcomes prefers the one where a rival is thrown overboard, which is why a bought pirate must be paid one coin and not zero. And the voting rule is at least half, with the proposer voting. Change the rule to a strict majority of all votes and the two-pirate case flips, D cannot pass anything and dies, so E would reject everything in the three-pirate round, and the whole ladder shifts. With a strict majority of the other pirates' votes the ladder shifts again, and a candidate who re-derives it from the two-pirate anchor is the one who gets hired. The limitation is that real negotiators are not this rational; the puzzle is a model of backward induction, not of pirates.
Where candidates lose it
The common loss is reasoning forward from fairness, offering 20 each or 50 for A, and being unable to say why those numbers rather than others. Fairness is not in the rules; survival and coins are.
The second is paying the wrong pirates. B and D would do well without A and cannot be bought cheaply; C and E would get nothing. Candidates who pay B and D have not worked out the four-pirate round.
What the interviewer asks next
- The proposal now needs a strict majority of all votes. What does A propose?
- The proposal needs a majority of the other pirates' votes, excluding the proposer. Does A survive?
- There are 6 pirates. How does the pattern continue?
- What if a pirate who is indifferent prefers to keep the proposer alive?
Asked at Old Mission Capital, Prop Trading, New York, 2014 (Wall Street Oasis):
There are 5 pirates and they are trying to split 100 gold coins in a rational way.
