Derivatives Foundation puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 66
- Topics
- 12
- Hard
- 29
019You roll a fair die again and again until the first six appears. What is the expected total of all the numbers rolled, including the final six?Jane StreetNew York · 2026
Try it first
Commit to a number first: the expected total, six included, is
Show the worked solution
21. A six comes up one time in six, so you expect six rolls, and each roll averages 3.5; by Wald's identity the expected total is 6 x 3.5 = 21. From the other side: the five expected rolls before the six can only be 1 to 5, so each averages 3, giving 15, and the closing six makes 21. The tempting 5 x 3.5 + 6 = 23.5 forgets that the earlier rolls cannot be sixes.
Why is 23.5 wrong when every roll averages 3.5?
Count people walking through a door until the first one in a red shirt arrives. The people before that one are, by the way you counted, not in red; nobody would estimate their shirts from the whole crowd. Conditioning on a roll not being a six changes its average from 3.5 to 3, and every roll before the stopping roll carries that condition. So the honest split is five expected rolls at 3 each, which is 15, plus the six that ends the game. The slip to 23.5 counts the sixes twice: once in the 3.5 of the earlier rolls, and again as the final roll.
The five expected rolls before the first six can only show 1 to 5, so they average 3 and total 15, and with the closing six the expected total is 21, the same as six expected rolls at 3.5 each, while the tempting 23.5 wrongly lets the earlier rolls average 3.5; a seeded simulation of 200,000 games gives 20.92. Why does 6 x 3.5 still give the right answer?
Because of Wald's identity, and it is worth naming in the room. If the decision to stop depends only on rolls you have already seen, the expected total equals the expected number of rolls times the average roll. Before any roll is made, each one is a fresh die with average 3.5; the stopping rule only decides how many of them you take, and it cannot peek ahead. The two routes agree because the missing sixes in the early rolls are exactly balanced by the guaranteed six at the end. A one-line check also works: the first roll averages 3.5, and five times in six the game restarts.
The relationshipS the total of all rolls, six included N the number of rolls until the first six, with E[N] = 6 X one roll of the die, with E[X] = 3.5 5/6 the chance the first roll is not a six and the game starts afresh What it says in wordsConditioning on the first roll gives the same 21 as multiplying the expected number of rolls by the average roll.Where does this reasoning show up on a desk?
Any rule of the form trade until something happens. A rule such as take profit once up 2% cannot change the expected P&L of a fair bet, for the same reason: a stopping rule that only looks backwards changes when you stop, not the average of what you collect per step. The limitation to state is that Wald needs the expected number of steps to be finite; a rule like keep going until you are ahead can break it. A sharp follow-up the interviewer may use: given that every roll was even, what is the expected number of rolls? The answer is 1.5, not 3, and the same conditioning idea explains it.
Where candidates lose it
The common answer is 23.5: five rolls at the familiar 3.5, plus the six. It sounds careful because it treats the last roll separately, and it is wrong because the earlier rolls are conditioned on not being sixes.
The second loss is getting 21 from 6 x 3.5 and not being able to say why it is allowed. The interviewer will push: the number of rolls depends on the rolls, so why can you multiply? The answer is that the stop depends only on the past, which is Wald's identity.
What the interviewer asks next
- Given that every roll in the game was even, what is the expected number of rolls?
- What is the expected total if you stop at the first roll of 5 or 6, counting that roll?
- You are paid the total but must pay Rs 4 per roll. Would you play, and at what cost per roll is the game fair?
Asked at Jane Street, Investment Operations, New York, 2026 (Wall Street Oasis):
First interview was testing simple math brainteasers (e.g. expected value of dice throws, etc.)
022Make me a market on the total number of Olympic gold medals the United States has ever won, Summer and Winter Games together. You cannot look anything up.Old Mission CapitalBoston · 2024
Try it first
Once the structure is written down, which input deserves most of your quote's width?
Show the worked solution
About 1,096 in the middle, so quote 1,000 bid, 1,200 offered, and say the share of golds is where you are least sure. Summer: about 29 Games attended, times about 190 events, times about 18% won by the United States, gives about 992. Winter: 24 Games, about 62 events, about 7%, adds about 104. The honest range is 756 to 1,491; the quote sits around the middle and moves as soon as someone trades on it.
What is the interviewer actually marking?
If a friend asks how much your monthly grocery bill is, you do not recall a number; you think of trips a week, spend a trip, four weeks, and say a figure with a range. A market on an unknown quantity is marked on the structure behind the middle, the honesty of the range, and what you do when someone trades with you, not on knowing the answer. The interviewer knows the real count. Treat every input below as an assumption to be checked against the published medal table, and say so; inventing precision is worse than a wide range.
Summer Games contribute about 29 x 190 x 18% = 992 golds and Winter Games about 24 x 62 x 7% = 104, a total of about 1,096 with an honest range of 756 to 1,491, and the share of golds is the widest link, so the quote of 1,000 at 1,200 sits around the middle. How do you build the middle?
Split it into Summer and Winter, because they differ by an order of magnitude. Modern Summer Games run every four years from 1896, less three wartime cancellations, about 30 Games, and take one off for the boycotted 1980 Games, 29. Events grew from a few dozen to over three hundred, so call the average about 190, with 170 to 210 as the bracket. The share of golds is the input you own: the United States wins more than any other country, so something like a sixth to a fifth of all golds, 18% in the middle, is defensible but soft. Winter adds 24 Games at about 62 events and a smaller share, about 7%.
The relationship29, 24 Summer and Winter Games the United States took part in, as estimated 190, 62 average events per Games, one gold each 0.18, 0.07 the assumed United States share of golds in each What it says in wordsGolds are the number of Games times events per Games times the share the United States wins, added across Summer and Winter.How wide should the quote be, and what happens when someone trades?
Running all lows together and all highs together gives 756 to 1,491, but every input is unlikely to sit at its extreme at once, so the quote can be much tighter than that range. Quote about 10% either side of the middle, 1,000 bid and 1,200 offered, and treat the first trade as information from someone who may know the answer. If the interviewer pays 1,200, move both sides up, perhaps to 1,150 at 1,350, and ask yourself which link you underestimated; the share is the first suspect. The limitation to admit is that a quote built from memory of rough facts can be confidently off-centre, which is exactly why the market moves when it is traded.
Where candidates lose it
The common failure is to blurt one number, or a market as wide as 500 at 2,000, with no structure behind it. A number with no structure cannot be defended, and a market that wide is a refusal to quote dressed up as caution.
The second loss is holding the quote after the interviewer lifts it twice. A counterparty who keeps buying is telling you the answer is higher. Move up, widen if you must, and say which assumption you are revising.
What the interviewer asks next
- I buy 10 at your offer. Where is your new market?
- Now make a market on the total golds won by all countries ever. Which of your inputs carries over?
- Your structure gave about 1,100. If the true answer were 1,500, which link would you suspect, and why?
Asked at Old Mission Capital, Equities, Boston, 2024 (Wall Street Oasis):
One was on making a market on the total number of Olympic gold medals the US has won.
025A stock trades at 500. The one-year 500-strike call is priced at 40 and the 500-strike put at 45, and interest rates are zero. What dividend is the options market pricing in, and what would you check before trading on it?Equity derivativesMarket making
Try it first
Before using any formula: the put costs more than the call at the same at-the-money strike. With zero rates, that tells you
Show the worked solution
A dividend of 5 per share. With zero rates, put-call parity says call minus put equals spot minus the dividend minus the strike. Here 40 minus 45 = 500 minus D minus 500, so D = 5, and the implied forward is 495. Before trading against it, check that the options are European or that early exercise is worth nothing, what it costs to borrow the stock, whether the ex-date falls before expiry, and whether all four prices are live and tradable.
Why does a put costing more than a call point to a dividend?
Suppose you agree today to buy a friend's scooter in a year for its fair price, but the friend will keep using it to deliver parcels and pocket the fees until then. You would pay less than today's price, by the fees they collect. Holding a call and selling a put at the same strike is an agreement to buy the stock at the strike, so call minus put must equal the forward price minus the strike, and the forward is spot minus whatever the holder collects before expiry. With zero rates there is no interest to account for, so a put dearer than the call by 5 says the forward is 5 below spot.
The relationshipC, P the call and put prices at the same strike and expiry S the spot price, 500 K the strike, 500 D the cash the stock pays out before expiry, here the implied dividend What it says in wordsWith zero rates, the call minus the put equals the spot less the dividend less the strike, so the dividend is what balances the equation.Call minus put is minus 5, so the options imply a forward of 495, and since the stock's forward is 500 minus the dividend, the market is pricing a dividend of 5; if the company pays only 3, the remaining 2 has to be explained by something else, such as the cost of borrowing the stock. What would you check before trading on it?
Say your own estimate of the dividend is 3. Then a reversal looks attractive: short the stock, buy the call, sell the put, collecting 500 + 45 minus 40 = 505; at expiry the options deliver the stock back at 500 and you owe the 3 dividend, 503 in all, a locked-in 2 per share. That 2 is only yours if you can borrow the stock for less than 0.4% of its price over the year; a hard-to-borrow stock shows up in parity exactly as an extra dividend. Then check the rest: American calls can be exercised just before a dividend, the ex-date must fall before expiry, and four bid-offer spreads can eat a gap of 2 on their own.
What is the general lesson the interviewer wants?
That parity is an accounting identity between three things you can trade, and any gap between what it implies and what you believe is a claim about something you have not yet priced. The options market does not quote a dividend; it quotes a forward, and the dividend, the borrow cost and the interest rate are the pieces you split it into. A good answer gives the 5, names the implied forward of 495, and then lists what could make the 5 something other than a dividend. The limitation to say plainly: with non-zero rates, the strike is discounted and the arithmetic shifts, so state the zero-rate assumption before quoting the number.
Where candidates lose it
The common answer is that the put is dearer because the market expects the stock to fall. At one strike and one expiry, direction cannot make the put dearer than parity allows; anyone could sell the put, buy the call and short the stock against it. The gap is a forward, not a view.
The second loss is stopping at 5. The question asks what you would check, and the borrow cost is the one interviewers wait for: an implied dividend above the announced one is often a stock that is expensive to short.
What the interviewer asks next
- Interest rates are now 6% a year. Redo the implied dividend, and say which way it moves.
- The company announces a dividend of 8. Which trade would you put on with these four prices, and what is the risk?
- Why might an American call on this stock be worth more than its European twin, and how would that distort the implied dividend?
028How many zeros are at the end of 1000! (1000 factorial)?Jump TradingChicago · 2013
Try it first
Pick the count before you work it out.
Show the worked solution
249 trailing zeros. A zero at the end is a factor of 10, and a 10 is a 2 times a 5. Among 1 to 1,000 there are 994 factors of 2 but only 249 factors of 5, so the 5s decide. Multiples of 5 contribute 200, multiples of 25 add a second 5 each for 40 more, multiples of 125 add 8, and 625 adds 1: 200 + 40 + 8 + 1 = 249.
Why count 5s and not 10s?
Imagine packing gift boxes that each need one lid and one base, from a pile of 994 lids and 249 bases. You can make 249 boxes, and the extra lids are useless. A trailing zero is a box: it needs one factor of 2 and one factor of 5, and in the product 1 x 2 x 3 x ... x 1,000 the 5s run out long before the 2s. Counting multiples of 10 misses every zero that comes from pairing a 5 in one number with a 2 in another; 4 x 5 is 20, which ends in a zero though neither factor is a multiple of 10. So the question is simply: how many times does 5 divide into 1,000 factorial?
Multiples of 5 up to 1,000 contribute 200 factors of 5, multiples of 25 a further 40, multiples of 125 another 8 and 625 one more, which add to 249, while the 994 factors of 2 are never the constraint. How do you count the 5s without missing the doubled ones?
Count in layers. Every multiple of 5 carries at least one 5: 1,000 divided by 5 is 200. Every multiple of 25 carries a second 5 that the first layer did not see: 1,000 over 25 is 40. Multiples of 125 carry a third, 8 of them, and 625 carries a fourth, once. Each layer counts only the extra 5 that the layer before it missed, which is why you add the plain quotients and never multiply. The sum is 249. The same method gives 24 zeros for 100 factorial, 20 plus 4, and the layering stops as soon as the power of 5 exceeds the number.
The relationshipZ(n) the number of trailing zeros of n factorial floor of n over 5^k how many numbers up to n carry at least k factors of 5 What it says in wordsAdd the whole-number quotients of n by each power of 5 until the power is bigger than n.What does the interviewer learn from how you say it?
The answer is a test of whether you decompose a question into its prime factors rather than reach for the obvious count. Say the pairing argument first, then the layered count, then check the shape: the answer is a little under n over 4, because 200 + 40 + 8 + 1 is a geometric series that sums towards 1,000 over 4. That check is useful on the follow-ups. For the number of factors of 2 the same layering gives 994; for 3s it gives 498, so in base 6, where a trailing zero needs a 2 and a 3, 1,000 factorial ends in 498 zeros. The limitation is worth a sentence: this counts zeros at the end, not zeros anywhere in the 2,568-digit number, which is a different and much harder question.
Where candidates lose it
The fast wrong answer is 100, one per multiple of 10, or 200, one per multiple of 5. The first misses every zero made by pairing a 5 with a 2 from a different number; the second misses the extra 5s inside 25, 125 and 625.
The other slip is counting 625 as four and 125 as three in the final sum, which double counts. Each layer adds only one new 5 per number, so the layers are added as plain quotients.
What the interviewer asks next
- How many trailing zeros does 100! have?
- How many trailing zeros does 1000! have when written in base 6? In base 12?
- What is the smallest n for which n! ends in exactly 100 zeros, and is there an n that ends in exactly 5 zeros?
Asked at Jump Trading, Trading, Chicago, 2013 (Wall Street Oasis):
How many 0's are in 1000! (factorial)?
029A point is dropped uniformly at random inside a unit square. What is the expected distance from the point to the nearest edge?Hudson River TradingNew York · 2024
Try it first
Pick the expected distance before you integrate.
Show the worked solution
One sixth. The distance to the nearest edge exceeds t only when the point lands inside the inner square of side 1 minus 2t, which has area (1 - 2t)^2. The expected value of a non-negative quantity is the integral of its tail probability, so E[D] is the integral of (1 - 2t)^2 from 0 to 1/2, which is 1/6, about 0.167. One axis alone would give 1/4; the second axis trims it.
Why work with the chance of being further than t, rather than the distance itself?
Suppose you want the average waiting time at a counter, and the only thing you can observe is, for each t, the fraction of people still waiting after t minutes. That is enough: add up the survival fractions over all t and you have the average wait. For any non-negative quantity, the expected value equals the area under its tail curve, and here the tail curve is easy to see: the point is further than t from every edge exactly when it lands inside the inner square of side 1 minus 2t. That inner square has area (1 - 2t)^2, and it shrinks to nothing at t equal one half, the centre of the square. Writing the density of the minimum of four dependent distances directly is far messier.
A random point is further than t from every edge only when it lands in the inner square of side 1 minus 2t, so the tail probability is (1 - 2t)^2, and the area under that curve from 0 to 1/2 is the expected distance of one sixth, confirmed by a 0.1667 Monte Carlo estimate. How does the integral come out to exactly 1/6?
Substitute u = 1 - 2t. As t runs from 0 to 1/2, u runs from 1 down to 0, and dt is minus du over 2. The integral becomes one half of the integral of u squared from 0 to 1, which is one half times one third. The answer is 1/6 because a squared tail integrates to a third and the half-width of the square halves it again. A sanity check: the median distance is where (1 - 2t)^2 equals a half, which is t = 0.146, a little below the mean, as you expect for a distribution with a long right tail reaching 0.5.
The relationshipD the distance from the point to the nearest edge (1 - 2t)^2 the area of the inner square where every edge is further than t away u the side of that inner square, 1 - 2t What it says in wordsThe expected distance is the area under the tail curve, and the tail curve is the area of a shrinking inner square.What is the fast cross-check that shows you understand the structure?
Do one axis first. The distance to the nearer of the left and right edges is min(x, 1 - x), a triangle-shaped quantity with mean 0.25. The nearest edge of the square is the smaller of two such independent quantities, one per axis, and taking the smaller of two pulls the mean down from 1/4 to 1/6. Say 1/4 for one axis, then 1/6 for two, and the interviewer hears that you see the minimum of independent pieces rather than a formula. The same structure gives 1/8 for a unit cube, where the tail is (1 - 2t)^3, and it shows the limitation of the method: it works because the inner region stays the same shape as it shrinks, which fails for a disc, where the nearest point on the boundary is not along an axis.
Where candidates lose it
The common loss is answering 1/4, the one-axis answer, forgetting that the nearest of four edges is smaller on average than the nearest of two. The second axis is independent and matters.
The other is trying to write the density of the minimum directly and running out of time. The tail curve is one line, (1 - 2t)^2, and the expected value is its integral. Learn that identity; it solves half the expected value questions a desk asks.
What the interviewer asks next
- What is the expected distance to the nearest face for a point dropped in a unit cube?
- What is the expected distance to the nearest edge in a 2 by 1 rectangle?
- What is the expected distance from the point to the centre of the square? Why is that a different kind of integral?
- Two points are dropped. What is the chance that both are further than 0.1 from every edge?
Asked at Hudson River Trading, Campus Algo Dev Interview, New York, 2024 (Wall Street Oasis):
expected value question involving the expected value among distance to an edge, with a randomly placed object
030A coin falls your way 60% of the time and pays even money. What fraction of your capital should you bet each time, and what growth rate does that give?OptiverAustin · 2025
Try it first
What fraction of your capital goes on each flip?
Show the worked solution
Bet 20% of your capital each time, for a growth rate of about 2.0% per bet. The Kelly fraction at even money is the edge, p minus q, which is 0.6 minus 0.4. The expected log growth is 0.6 x ln(1.2) + 0.4 x ln(0.8) = 0.0201, so capital typically grows by a factor of about 7.5 over 100 bets. Bet twice that, 40%, and the growth turns slightly negative, -0.24% a bet: anything beyond about 39% loses money despite the edge.
Why not bet everything on a coin that favours you?
A street vendor with a stall that makes money six days in seven does not spend the whole float on stock each morning, because the seventh day would end the business. Expected value is the right guide for a single bet with money you can replace, but when you must survive to play again, what matters is the growth of your capital over many bets, and that is governed by the average of the logarithm, not the average of the rupees. Betting all of it has the highest expected value and a certainty of ruin: one loss and there is nothing left to compound. The logarithm punishes that loss infinitely, which is the mathematical way of saying you cannot come back from zero.
The relationshipf the fraction of capital staked on each bet p, q the chance of winning and of losing, 0.6 and 0.4 f star the fraction that maximises long-run growth; at even money it is the edge g the expected log growth per bet, about 2% here What it says in wordsMaximise the average log of your wealth after one bet, and at even money the best stake is simply the edge.Expected log growth per bet rises from zero to a peak of 2.01% at a 20% stake and falls back through zero at about 39%, so half Kelly keeps three quarters of the growth while double Kelly, just past the crossing, loses a little on every bet. Where does the 20% come from, and what if the odds are not even?
Differentiate g with respect to f and set it to zero: p over (1 + f) equals q over (1 - f), which gives f = p - q. At even money that is the edge, 0.2. For a bet that pays b to 1 the formula becomes (bp - q) over b, edge over odds. Say the general form, because the interviewer will often move the payout: at 2 to 1 on the same coin the stake becomes (1.2 - 0.4) / 2 = 40%, and at odds where bp is below q the right stake is zero. The growth at 20% is 0.0201 per bet, which compounds to a factor of e to the 2.01, about 7.5 times, over 100 bets on the typical path. That is the median outcome, not the average: a few lucky paths do far better and pull the mean up.
Why do desks bet less than Kelly?
The curve is flat on the left and steep on the right. Half Kelly, 10%, still delivers 1.50% a bet, three quarters of the peak, with far smaller swings. Double Kelly, 40%, delivers -0.24%, already below zero, with wild swings: the crossing point is about 39%, and anything beyond it shrinks capital in the long run even though every single bet has positive expected value. Over-betting is punished much harder than under-betting, and in real trading the 60% is an estimate rather than a fact, so sizing at a fraction of Kelly is the standard discipline. The limitation to state: Kelly assumes you know p, that bets are independent and repeated, and that you can resize freely; on a desk, position limits and the uncertainty in your edge usually bind before the formula does.
Where candidates lose it
The common loss is answering with the expected value: 60% of the time you win, so bet big, or even bet everything. The interviewer wants to hear that growth is about the log, that ruin ends the game, and that the stake is the edge, 20%.
The second is giving 20% but no growth rate, or quoting the 20% expected return per rupee as the growth rate. The growth per bet is about 2%, far below the 20% edge, because the losses bite on a shrinking base.
What the interviewer asks next
- The coin pays 2 to 1 instead of even money. What is the stake now?
- Why does half Kelly give three quarters of the growth? What is the general relationship?
- You are only 60% sure the coin is 60% biased. How does that change the stake?
- What is the probability that a full-Kelly bettor halves their capital at some point before doubling it?
Asked at Optiver, Quantitative Research, Austin, 2025 (Wall Street Oasis):
Use Kelly Criterion
032One-month implied volatility is 20% and three-month implied volatility is 25%. What volatility is implied for the period from month one to month three?Volatility tradingEquity derivatives
Try it first
What is the forward volatility for months two and three?
Show the worked solution
About 27.2%. Variance is volatility squared times time, and variances add across periods. The three months carry 0.25 squared x 3 = 0.1875 of variance; the first month carries 0.20 squared x 1 = 0.04; so months two and three carry 0.1475 between them, 0.0737 per month on an annualised basis. The square root is 0.272. The forward volatility must sit above 25% because the quarter's average has to be pulled up from the 20% start.
Why do you add variances and not volatilities?
If you walk a random distance each hour, your spread after three hours is not three times the hourly spread, because some hours cancel others. What grows in a straight line is the variance, the spread squared. Implied volatility is quoted per year, so the variance an option carries is volatility squared times its life, and a longer option's variance is the sum of the variances of the pieces of time inside it. That is the whole mechanism. Twenty per cent for one month is 0.04 of variance; twenty-five per cent for three months is 0.1875; the difference belongs to the two months in between, and dividing by two months and taking the square root turns it back into a volatility.
Drawn as area, the first month's variance of 0.04 plus the two forward months' variance of 0.0737 each must fill the quarter's 0.1875, so the forward box has height 0.0737 and its square root is a forward volatility of 27.2%. The relationshipsigma 3, sigma 1 the three-month and one-month implied volatilities, 25% and 20% T the option life in months; the units cancel as long as both use the same one sigma 1,3 the forward volatility for the period between the two expiries What it says in wordsSubtract the short period's variance from the long period's and spread what is left over the time in between.What does the number tell a trader?
The 27.2% is the volatility you would lock in for months two and three by selling the one-month option and buying the three-month one in variance-weighted sizes, a calendar spread. If you believe realised volatility in those two months will be well below 27.2%, the three-month option is rich relative to the one-month, and the calendar is the trade that expresses it. Say the rounding: an upward sloping term structure, 20% then 25%, hides a forward that is steeper than either quote, 27.2%, and the trap of reading 25% as the forward is what the question is built to catch. The calculation uses calendar time in months; a desk would use trading days or variance-weighted business days, which shifts the number by a few tenths.
What is the limit of the calculation?
The subtraction has to leave something positive. If the three-month volatility were below 11.5% with the one-month still at 20%, the quarter would carry less variance than its first month alone, which is impossible without arbitrage: you could sell the one-month, buy the three-month and hold a position with negative forward variance. Forward variance can be small but never negative, so a term structure that inverts too sharply is a mispricing, not a forecast. The second limit is that both quotes must refer to the same strike in forward terms; mixing an at-the-money one-month with a three-month that has rolled away from the money adds skew to the comparison and the forward you compute is no longer clean.
Where candidates lose it
The common loss is averaging volatilities: 25% over three months with 20% in the first month gives 27.5% for the rest if you treat volatility as additive. It is close, which is why it survives, but it is the wrong quantity and on a steeper curve the gap is large.
The second is forgetting to weight by time: subtracting 0.04 from 0.0625 and taking the root. Write variance times time for each leg, every time.
What the interviewer asks next
- Six-month volatility is also 25%. What is the forward volatility for months four to six?
- The one-month is 30% and the three-month is 25%. What is the forward, and what does it say about the market?
- How would you actually lock in that forward volatility with listed options, and what would break the hedge?
- Why do desks compute this in trading days rather than calendar months?
034A card is drawn from a 52-card deck. You may bet Rs 100 at even money on red or on black. How much would you pay to be told first whether the card is a heart?OptiverChicago · 2025
Try it first
What is the tip worth, at most?
Show the worked solution
Up to Rs 50. Without the tip the bet is a coin flip worth zero. With it: one time in four the card is a heart, you bet red and win Rs 100 for certain. Three times in four it is not a heart, which leaves 13 red and 26 black cards, so you bet black and win 100 with probability 2/3 and lose 100 with probability 1/3, an expected Rs 33.33. The average, 25 + 25, is Rs 50, and that is the most the information is worth.
Why is a tip about hearts worth anything for a bet on colour?
A friend who will tell you whether it is raining in one of the four districts of a city is not telling you the weather everywhere, but if you have to bet on whether it rains in the city at all, the tip changes your odds. The heart tip never names the colour, but every heart is red, so a yes makes red certain and a no tilts the remaining deck two to one towards black. Both answers leave you with a bet that has an edge, where before you had none. Information is worth the gap between what you can earn with it and what you could earn without it, and here that gap is the whole value, because without the tip the best you can do is zero.
Without the tip a bet on red or black has an expected value of zero, while with it a heart lets you win Rs 100 for certain and a non-heart lets you bet black at 26 to 13 for an expected Rs 33.33, so the tip is worth a quarter of 100 plus three quarters of 33.33, which is Rs 50. How do you set the calculation out so it cannot go wrong?
Price the decision in each branch, then weight the branches by how likely each answer is. If the answer is yes, probability 13/52, the card is red: bet red, expected gain 100. If the answer is no, probability 39/52, there are 39 cards left of which 13 are red diamonds and 26 are black: bet black, expected gain 100 x 26/39 minus 100 x 13/39, which is 33.33. The value of the tip is the probability-weighted average of the best you can do after each answer, minus the best you could do with no answer at all. That is 1/4 x 100 + 3/4 x 33.33 = 50, minus zero. Pay less than Rs 50 and the deal is in your favour; pay exactly 50 and you are indifferent.
The relationship13/52 the chance the card is a heart 100 a certain win on red once you know it is a heart 26/39, 13/39 the chance of black and of red among the 39 non-hearts 0 the value of the bet with no information What it says in wordsWeight the best decision after each possible answer by the chance of that answer, and subtract what the bet was worth before.What does this have to do with a trading desk?
The question is a small model of paying for data. Being told the colour outright would be worth Rs 100, the full value of perfect information; the heart tip, which answers a narrower question, is worth exactly half of that. A piece of information is worth what it changes in your best decision, not how interesting it sounds: a tip that the card is an ace would be worth nothing here, because it leaves red and black at even money. The limitation to say out loud is that the Rs 50 is an expected value: on any single hand you could pay 50 and lose 100, so a desk with a limited bankroll should pay less than the full value, by the same logic that sizes bets below the edge.
Where candidates lose it
The common loss is answering 25 by reasoning that hearts are a quarter of the deck, so the tip is right a quarter of the time. The tip is informative in both of its answers: a no still tilts the deck two to one.
The second is pricing only the yes branch and forgetting to re-count the deck after a no. Thirty-nine cards remain, thirteen of them red, and that recount is where the second Rs 25 lives.
What the interviewer asks next
- How much is it worth to be told whether the card is an ace?
- How much is it worth to be told whether the card is a face card or a heart?
- You can bet Rs 100 on the suit at 3 to 1 instead. What is the heart tip worth now?
- Why should a desk with a small bankroll pay less than Rs 50 for this tip?
Asked at Optiver, Quantitative Research, Chicago, 2025 (Wall Street Oasis):
Valuing information, taking directional bets when not plus EV.
037X and Y are independent and uniform on 0 to 1. What is the density of X + Y, and what is the probability that X + Y is less than 1.5?CitadelChicago · 2025Citadel SecuritiesChicago · 2025
Try it first
What is P(X + Y < 1.5)?
Show the worked solution
The density is a triangle: f(s) = s for s between 0 and 1, and 2 - s for s between 1 and 2; P(X + Y < 1.5) = 7/8. The sum lands in a thin band near s only if X and Y fit inside the unit square along the line X + Y = s, and the length of that line segment grows until s = 1 and then shrinks. The sum exceeds 1.5 only in the top-right corner of the square, a triangle with legs 1/2 and area 1/8, so the probability of less than 1.5 is 0.875.
Why is the sum of two flat distributions a triangle?
Roll two ordinary dice: a total of 7 can be made six ways, a total of 2 or 12 only one way, so the totals pile up in the middle even though each die is flat. The sum of two independent uniforms is the continuous version: the density at s is the length of the diagonal line X + Y = s inside the unit square, which grows from 0 up to s = 1 and shrinks back to 0 at s = 2. Formally the density is the convolution, the integral over x of f(x) times g(s - x), and with both densities equal to 1 on the unit interval the integrand is 1 exactly where both x and s - x lie between 0 and 1. That interval has length s for s below 1 and length 2 - s above, which is the triangle.
The relationshipf X+Y the density of the sum the indicator equals 1 where both x and s - x lie in the unit interval, which is where both densities are non-zero the integral from 1.5 to 2 the area of the small triangle in the tail, 1/8 What it says in wordsConvolve two flat densities and the overlap of the two intervals gives a triangle; the tail beyond 1.5 is one eighth of it.On the unit square the event X + Y > 1.5 is the corner triangle with legs one half and area one eighth, and on the density plot the same one eighth is the tail of the triangle beyond s = 1.5, so P(X + Y < 1.5) is seven eighths. How do you get 7/8 in your head?
Use the square, not the integral. The pair (X, Y) is a uniformly random point in the unit square, so any probability is an area. X + Y > 1.5 is the region above the line from (0.5, 1) to (1, 0.5), a right triangle with both legs equal to 1/2, whose area is half of a quarter, which is 1/8. Everything else is 7/8. The same picture answers the mirror question at once: X + Y < 0.5 is the bottom-left corner, also 1/8. And it checks the density: integrating 2 - s from 1.5 to 2 gives (0.5)^2 / 2 = 1/8, the same number from the other direction, which is what you say to show the two views agree.
What does the interviewer ask next, and why?
The follow-ups test whether you understood the mechanism or memorised the triangle. Add a third uniform and the density becomes three parabolic pieces, smooth at the joins, with P(X + Y + Z < 1) = 1/6, the volume of the corner tetrahedron, and in general the sum of n uniforms heads towards a normal with mean n/2 and variance n/12, which is the central limit theorem arriving. The mean of X + Y is 1 and the variance is 1/12 + 1/12 = 1/6, by independence. Each extra convolution smooths the density by one degree: flat to triangular to parabolic, which is the general fact that sums of independent variables are smoother than their parts. The limitation worth saying: this all relies on independence; with X and Y correlated, the point is no longer uniform over the square and the areas no longer give probabilities.
Where candidates lose it
The common loss is treating the sum as uniform on 0 to 2 and answering 0.75. The sum is a triangle, and the mass in the top quarter of its range is only 1/8, not 1/4.
The second is setting up the convolution integral correctly and then mishandling the limits, integrating over the wrong interval of x. Draw the square; the limits are the overlap of two unit intervals, and the picture makes them obvious.
What the interviewer asks next
- What is P(X + Y + Z < 1) for three independent uniforms?
- What is the density of X - Y, and of max(X, Y)?
- What are the mean and variance of X + Y, and how do you get them without the density?
- X and Y are now correlated with rho = 0.5. Can you still find P(X + Y < 1.5) by area?
Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):
He was asking some questions about the probability, especially on the convolution.
Asked at Citadel Securities, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):He was asking some questions about the probability, especially on the convolution.
038Box A holds 3 red and 1 black ball, box B holds 1 red and 3 black. You pick a box at random and draw a red ball. You then draw again from the same box without replacement. What is the probability the second ball is red?Prop trading firmsMarket making
Try it first
Pick the probability of a second red before you compute.
Show the worked solution
Exactly 1/2. The first red is evidence: box A produces a red with probability 3/4, box B with 1/4, so after seeing red the odds are 3 to 1 for A, a posterior of 3/4. Then update the contents. Box A has 2 red of 3 left, so a second red has probability 2/3; box B has 0 red of 3, probability 0. Average over the updated boxes: 3/4 x 2/3 + 1/4 x 0 = 1/2. Forgetting the update gives 1/3, which is wrong.
Why does the first red change which box you think you hold?
If a friend's cooking is spicy three times out of four and yours one time in four, and a dish picked at random from one of you turns out spicy, you would guess it was your friend's, and you would be right three times in four. The first red ball is the same clue: it comes from box A three times as easily as from box B, so the box you are holding is A with probability 3/4, not 1/2. The arithmetic is Bayes. The weight for A is 1/2 x 3/4 = 3/8, for B it is 1/2 x 1/4 = 1/8, and A's share of the total 1/2 is 3/4. The question is built so that candidates who skip this update get a clean wrong answer.
After a red first draw the box is A with posterior 3/4 and B with 1/4, box A has 2 red of 3 left and box B none, so the chance of a second red is 3/4 x 2/3 plus 1/4 x 0, which is exactly one half. What has to be updated besides the belief about the box?
The box itself. One red ball has left it. Box A now holds 2 red and 1 black, so a red is 2/3; box B now holds 0 red and 3 black, so a red is impossible. Both updates matter and they pull in opposite directions: believing more in A pushes the answer up, removing a red from A pushes it down, and here they land on exactly 1/2. Say that out loud, then check it by counting. Imagine 8 rounds, 4 with each box. Box A yields a first red in 3 of its 4 rounds and box B in 1 of its 4: four red first draws in all. In the three from A the second draw is red 2/3 of the time, 2 rounds; in the one from B, never. Two of four is 1/2.
The relationshipR1, R2 the first and second draws are red P(A | R1) the posterior that the box is A after one red, 3/4 P(R2 | A, R1) a second red from box A with one red already gone, 2/3 What it says in wordsWeight each box's chance of a second red by how likely that box is after the first red.How do the variants separate candidates who understand from those who memorised?
Put the first ball back and the box contents do not change but the belief still does: 3/4 x 3/4 + 1/4 x 1/4 = 5/8, which is above 1/2 because the posterior leans to A and A stays rich in red. Forget the belief update and keep the contents update, and you get 1/2 x 2/3 + 1/2 x 0 = 1/3. The gap between 1/3, 1/2 and 5/8 is the whole lesson: the first draw tells you about the box and changes the box, and you need both. If a second red does appear, the posterior for A becomes 1, since B has no reds left, and a third red is then 1/2. The limitation is the usual one for Bayes: the 1/2 prior for each box is given here, and in a trading setting the prior is the thing you are least sure of.
Where candidates lose it
The common loss is 1/3: keeping the 1/2 prior on each box and only updating the contents. The first red is evidence about the box, and leaving it out throws away half the information in the question.
The second is 2/3: updating to box A and then forgetting that B is still possible with probability 1/4, or forgetting that A has lost a red. Carry both boxes through to the end, then average.
What the interviewer asks next
- The first ball is put back before the second draw. What is the probability now?
- The second ball is also red. What is the probability the box is A, and that a third draw is red?
- Box B held 0 red and 4 black instead. What changes, and what is the answer?
- A trader sees one winning trade from a new signal. What is the analogue of the box update?

