Hedge Funds puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 38
- Topics
- 14
- Hard
- 30
031You roll two fair dice and are paid the higher of the two faces, in rupees. What is the expected payout?Wolverine TradingChicago · 2025
Try it first
Your quick estimate:
Show the worked solution
161/36, about Rs 4.47. The higher face equals k in 2k minus 1 of the 36 equally likely rolls: 1, 3, 5, 7, 9 and 11 rolls for k from 1 to 6. Multiply each value by its count, add to 161, and divide by 36. The lower face averages 91/36, about 2.53, and the two add to 7, the average total of two dice, which is a quick check.
How many of the 36 rolls give each maximum?
Think of two runners and a prize for the faster time: the winning time is better than a typical single runner's because you always keep the better of two. The higher face is at most k in k x k of the 36 rolls, so it equals exactly k in k squared minus (k minus 1) squared, which is 2k minus 1 rolls. That gives 1 roll with a maximum of 1, 3 with a maximum of 2, and on up to 11 with a maximum of 6. On the grid those cells form L shapes that grow as you move towards the corner.
Of the 36 equally likely rolls, the higher face is 1 in just one roll and 6 in eleven rolls, so the expected maximum is 161/36, about 4.47, well above the 3.5 of a single die. The relationshipk the value of the higher face 2k - 1 the number of rolls, out of 36, whose higher face is exactly k What it says in wordsWeight each possible maximum by how many of the 36 rolls produce it.How do you check 4.47 in ten seconds?
Use the pair. The higher face plus the lower face always equals the total of the two dice, so their averages must add to 7. The lower face is at least k in (7 minus k) squared rolls, which gives an average of 91/36, about 2.53. 4.47 plus 2.53 is 7.00. A second method that lands exactly is what makes an interviewer stop checking your arithmetic and move on to the follow-up.
The follow-up is usually a game. If you could pay to roll one die or to roll two and keep the higher, the second is worth about Rs 0.97 more. That gap, the value of a free second look, is the same idea as an option: the right to choose after seeing the outcome is worth paying for.
Where candidates lose it
The common loss is answering 3.5 plus something vague, or 5, from instinct. Both skip the count of how often each maximum occurs, which is the whole question.
The second is listing all 36 rolls one by one under time pressure. Say the 2k minus 1 rule, give 161 over 36, and use the lower face check to show the number is right.
What the interviewer asks next
- What is the expected higher face with three dice?
- What would you pay to roll two dice and keep the higher, if you could reroll both once?
- What is the expected value of the lower face, and why do the two add to 7?
Asked at Wolverine Trading, Equity Hedge, Chicago, 2025 (Wall Street Oasis):
Typical dice questions that you can find in most probability textbooks
032Two ice cream sellers each choose a spot on a straight 1 km beach. Sunbathers are spread evenly along it and each walks to the nearer seller. Where do the sellers end up, and is that the best outcome for the customers?Prop and quant trading firmsLong-short equity funds
Try it first
Where do two self-interested sellers settle?
Show the worked solution
Both end up side by side in the middle, and customers are worse off. From the quarter points, a seller who steps inward keeps everyone behind them and wins beach from the rival, so both drift to the centre. There they still split customers half and half, but the average walk doubles from 125 m to 250 m. Competition moves the sellers to the spot that maximises share, not the one that serves customers best.
Why can neither seller stay at the quarter points?
Picture two petrol pumps on a highway. Each wants the drivers on its own side plus as many from the middle as it can reach first. A seller keeps every customer on the far side of them wherever they stand, so moving towards the rival only ever adds customers. From 250 m, A steps to 400 m against B at 750 m: the dividing line moves to the midpoint, 575 m, and A's share rises from 50% to 57.5%. B then responds the same way, and the dance ends only when both stand at 500 m.
Sellers at the quarter points split the beach evenly with an average walk of 125 m; when A edges inward to 400 m it wins 57.5% of the beach, and the process ends with both at the middle, still splitting 50/50 but with the average walk doubled to 250 m. What does the middle cost the customers?
With both sellers in the centre, a sunbather is on average a quarter of the beach away, 250 m. With sellers at 250 m and 750 m, nobody is more than 250 m away and the average is 125 m. The shares are identical in both setups; the only thing that changed is how far customers walk, so the stable outcome is strictly worse for them and no better for the sellers. This is the {term('Hotelling model', 'A model of competition on a line, set out by Harold Hotelling in 1929, in which rivals crowd towards the centre to win the middle ground.')}, and it is why rival shops cluster and why two parties often converge on the middle voter.
Say the equilibrium idea in one line: a pair of positions is stable when neither player can do better by moving alone, and the middle is the only such pair here. Then say where the model breaks: if customers stop buying when the walk is too long, or if prices can differ, the sellers have a reason to spread out again. Interviewers ask this to see whether you can reason about another player's best response, which is most of trading.
Where candidates lose it
The common loss is answering the quarter points, because that is the sensible arrangement. The question asks where self-interested sellers end up, and the quarter points are not stable.
The second loss is getting to the middle but not saying what it costs. The interviewer wants the contrast: same shares, twice the walking. Name the stable point, then name the welfare cost.
What the interviewer asks next
- What happens with three sellers?
- If customers refuse to walk more than 300 m, where do the sellers stand?
- Where do you see the same pattern in markets or in fund positioning?
033Make a two-way market on the sum of three fair dice. Then one die is revealed to be a 6. Where do you move your market, and should it get wider or narrower?CitadelLondon · 2026
Try it first
After the 6 is shown, what happens to your market?
Show the worked solution
Move the mid from 10.5 to 13 and tighten the market by about a fifth. Each die averages 3.5, so three dice average 10.5. Once one die shows 6, the sum is 6 plus two unknown dice averaging 7, which is 13. The variance falls from 3 x 35/12 to 2 x 35/12, so the standard deviation drops from 2.96 to 2.42. If the first market was 9.5 at 11.5, the new one is about 12.2 at 13.8.
Where do you put the first market, and how wide?
Start from the fair value and then decide the width from how uncertain the outcome is. The mid is the expected sum, 3 x 3.5 = 10.5, and the width should scale with the standard deviation of the sum, because that is how far the answer typically lands from the mid. One die has variance 35/12, so three independent dice have 35/4 = 8.75, a standard deviation of 2.96. A market of 9.5 bid, 11.5 offered is a reasonable opening: tight enough to trade, with room for your edge.
What does revealing one die change?
A weather forecast for tomorrow is more precise than one for next week, because fewer things can still change. Once one die is known, it contributes a certain 6 and no uncertainty, so the mid rises by 2.5 and only two dice of variance remain. The mid becomes 6 + 7 = 13. The variance becomes 35/6, a standard deviation of 2.42, down from 2.96. Scale the width by the same ratio, about 0.82, and a 2.0 wide market becomes about 1.6 wide: 12.2 at 13.8.
Before the reveal the sum is centred on 10.5 with a standard deviation of 2.96; after one die shows 6 it is centred on 13 with a standard deviation of 2.42, so the market moves up by 2.5 and tightens from 2.0 wide to about 1.6. Say what would make you widen instead. If the person revealing the die can choose which die to show, or picks the moment, the reveal itself carries information and you should be more careful, not less. A 6 chosen as the highest of three tells you the other two are 6 or lower, and they no longer average 7. Interviewers like it when you ask who chose what to reveal before you requote.
Where candidates lose it
The common loss is widening after the 6 because it feels like a shock. A shock that is fully known removes uncertainty. The mid jumps, but the range of outcomes shrinks.
The second loss is moving the mid by the full 6, or to 16.5 as if all dice were sixes. Only the revealed die is known; the other two still average 3.5 each.
What the interviewer asks next
- A second die is revealed as a 1. Where is your market now?
- The revealer chose to show the highest of the three dice. Where do you quote?
- Someone lifts your 13.8 offer straight away. What do you do next?
Asked at Citadel, Quantitative Research, London, 2026 (Wall Street Oasis):
3rd I got rejected it was different brainteasers and trading game
034A stock trades at Rs 1,000 and its one-month options are priced at 30% implied volatility. Without a calculator, roughly what does a one-month at-the-money straddle cost, and how far must the stock move for the buyer to break even at expiry?Volatility and relative value fundsProp and quant trading firms
Try it first
Closest estimate for the straddle:
Show the worked solution
About Rs 69, so the stock must move about 6.9% either way. Scale the 30% annual volatility to one month by the square root of time: 30% x root(1/12) is about 8.7%. An at-the-money straddle is worth about 0.8 times that move times the price: 0.8 x 8.7% x 1,000 is about Rs 69. At expiry the buyer profits only below about Rs 931 or above Rs 1069.
Where does the 0.8 x sigma x root T rule come from?
A straddle pays the size of the move, whichever way it goes, so its value is the expected absolute move. For a normal distribution, the average absolute move is the standard deviation times root(2/pi), which is about 0.8. Over one month the standard deviation of the price is S x sigma x root T = 1,000 x 0.30 x root(1/12), about Rs 87. Times 0.8 gives about Rs 69. Black-Scholes with zero rates gives Rs 69.1, so the shortcut is good to within a rupee for short-dated at-the-money options.
A one month straddle struck at Rs 1,000 costs about Rs 69, so at expiry it loses money anywhere between about Rs 931 and Rs 1069 and profits only on a move of more than about 6.9% either way. Why is the breakeven move bigger than the typical move traders expect?
Buying a straddle is like buying insurance against a big move in either direction: you pay the average claim up front. The breakeven at expiry is the premium, about 6.9%, while the average absolute move is the same 6.9% by construction, so a buyer at fair implied volatility breaks even only on average. The buyer makes money when realised volatility turns out higher than the 30% priced in, and the seller makes money when it turns out lower. That is why desks talk about straddles as a bet on volatility, not on direction.
The relationshipsigma implied volatility, 30% a year T time to expiry in years, 1/12 S the stock price, Rs 1,000 What it says in wordsAn at-the-money straddle costs about 0.8 times the one standard deviation move over its life.Say the limits. The rule assumes the option is at the money and short dated, and rates are small; for long-dated or out-of-the-money options it drifts. A single call or put is half the straddle, about 0.4 x sigma x root T x S, which is a useful second number to have ready.
Where candidates lose it
The most common loss is forgetting to scale volatility to one month, which gives an answer near Rs 240 or Rs 300. Implied volatility is quoted per year; divide by root 12 for a month.
The second loss is forgetting the 0.8 and pricing the straddle at the full one standard deviation move, Rs 87. Say where the 0.8 comes from, the average absolute value of a normal, and the answer sounds reasoned rather than memorised.
What the interviewer asks next
- What does a one-week straddle cost on the same stock?
- Roughly what is the one month at-the-money call worth on its own?
- The stock moves 5% by expiry. Did the straddle buyer or seller win, and what does that say about realised volatility?
035Three assets all have the same pairwise correlation, rho. What is the lowest value rho can take? What is the answer for n assets?Multi-manager platformsQuant and systematic funds
Try it first
Lowest possible common correlation for three assets:
Show the worked solution
Minus one half for three assets, and minus 1/(n minus 1) for n. Give each asset unit variance and add them up. The variance of the sum is 3 plus 6 rho, because there are three variances and six pairwise covariance terms. A variance cannot be negative, so rho is at least -1/2. With n assets the sum's variance is n plus n(n minus 1) rho, which gives rho at least -1/(n minus 1): -0.33 for four, -0.11 for ten.
Why can three assets not all be perfectly opposed?
Three friends cannot all disagree with each other on a yes-or-no question: if Ravi says yes and Meena says no, Arjun agrees with one of them. Perfect opposition is a relationship between two things; with three, two of them must lean the same way. The same limit holds for correlations. A set of numbers in a correlation matrix has to be internally consistent, and equal pairwise correlations become impossible well before -1 once there are three or more assets.
How do you find the exact floor?
Use the one fact that can never fail: a variance is zero or more. Add the three assets with unit variance; the sum has variance 3 from the three diagonal terms plus 6 rho from the six covariance terms, and 3 + 6 rho must be at least zero. That gives rho of at least -1/2. The picture is three arrows 120 degrees apart: each pair has a cosine of -1/2, and the three add to exactly zero, which is the boundary case. For n assets, n + n(n minus 1) rho at least zero gives -1/(n minus 1).
Three unit arrows 120 degrees apart add to zero, the arrangement where every pair has correlation minus one half, and the variance of the sum, 3 + 6 rho, turns negative below that value, which is impossible, so rho cannot fall below minus one half. The relationshipn the number of assets, each with variance 1 rho the common pairwise correlation n(n-1) the number of ordered pairs, each contributing one covariance of rho What it says in wordsThe variance of the equal-weighted basket must be non-negative, which caps how negative a shared correlation can be.Say why a risk desk cares. A correlation matrix that breaks this rule is not a valid risk model: it implies some portfolio has negative variance, and an optimiser will pile into it. This happens in practice when correlations are estimated pair by pair from different data windows or overridden by hand in a stress test. The check is that the matrix is positive semi-definiteA matrix for which every weighted combination of the assets has a variance of zero or more., and the equal-correlation case is the cleanest example of the rule.
Where candidates lose it
The usual loss is answering -1, because that is the floor for any single pair. The interviewer is testing whether you see that the pairs constrain each other.
The second loss is getting -1/2 by intuition but having no proof. Say the variance of the sum in one line. It takes ten seconds, it generalises to n immediately, and it is the answer the follow-ups build on.
What the interviewer asks next
- Two assets have correlation 0.9 with a third. What is the lowest possible correlation between the first two?
- A stress test sets every pairwise correlation in a 20-asset book to -0.1. Is that a valid matrix?
- How would you repair a correlation matrix that is not positive semi-definite?
036A regime model says a bull month is followed by another bull month 90% of the time, and a bear month by another bear month 80% of the time. In the long run, what share of months are bull months?Quant and systematic fundsProp and quant trading firms
Try it first
Long-run share of bull months:
Show the worked solution
Two thirds of months are bull months, whatever the starting state. In the long run the number of months switching from bull to bear must equal the number switching back. 10% of bull months switch out and 20% of bear months switch in, so 0.1 x bull = 0.2 x bear, which makes bull twice as common as bear: 2/3 against 1/3. A second route: bull spells last 10 months on average and bear spells 5.
What has to balance in the long run?
Think of a shop with people walking in and out all day. Once the crowd inside stops growing or shrinking, the number walking in each minute must equal the number walking out. In the long run, the flow from bull to bear must equal the flow from bear to bull, because otherwise one state would keep filling up. The flow out of bull is 10% of bull months; the flow out of bear is 20% of bear months. Setting 0.1 x bull equal to 0.2 x bear, with bull plus bear equal to 1, gives bull = 2/3.
Bull months turn bear 10% of the time and bear months turn bull 20% of the time, so the long-run share of bull months is two thirds, and a chain started in either state is within a few points of two thirds after about a year. How do you check two thirds another way?
Use the length of each spell. A state you leave with probability p each month lasts 1/p months on average, so bull spells last 10 months and bear spells 5. Spells alternate, so over a long stretch the market spends 10 months bull for every 5 bear: 10 out of 15 is two thirds. Two methods that agree is what the interviewer is listening for.
The relationshippi_bull the long-run share of bull months 0.1, 0.2 the chances of leaving bull and leaving bear each month 0.7 how much of any starting gap survives each month What it says in wordsThe long-run share of a state is the chance of entering it divided by the total chance of switching, and the start is forgotten at a rate of 0.7 a month.Answer the part of the question people skip: why the starting state does not matter. The gap between today's odds and two thirds shrinks by a factor of 0.7 every month, so after 12 months only 1.4% of it is left. A model this sticky still forgets its starting point within about a year, which is why regime forecasts beyond a few months mostly return the long-run average.
Where candidates lose it
The first loss is answering 90%, the one-step persistence, as if it were the long-run share. The second is saying it depends on today's state, which is true for next month and false for the long run.
Set up the flow balance in one line, give two thirds, then check it with spell lengths of 10 and 5 months. If you have time, say how fast the start is forgotten: 0.7 a month.
What the interviewer asks next
- Today is a bear month. What is the chance that the month after next is a bull month?
- How long does the average bear spell last, and what is the chance one lasts more than a year?
- How would you estimate the two transition probabilities from 20 years of monthly data, and how wide would the error be?
037A strategy has an average annual return of 10% and annual volatility of 20%. Roughly what compound annual growth rate should an investor expect over many years?Fund of funds and allocatorsMulti-manager platforms
Try it first
Your estimate of the compound growth rate:
Show the worked solution
About 8% a year, two points below the 10% average. Compound growth is roughly the average return minus half the variance: 10% minus 0.5 x 0.20 squared, which is 10% minus 2%. Check with two years of +30% and -10%: the average is 10% and the volatility 20%, but 1.30 x 0.90 is 1.17, which compounds at 8.17% a year. The drag grows with the square of volatility.
Why does the average return overstate what you end up with?
A shopkeeper whose sales rise 50% one month and fall 50% the next has not broken even: 100 becomes 150 and then 75. A loss is applied to a bigger base after a gain, and a gain to a smaller base after a loss, so swings always drag compound growth below the simple average. The bigger the swings, the bigger the drag. The simple average of yearly returns is called the {term('arithmetic mean', 'The plain average of the yearly returns, adding them up and dividing by the number of years.')}; the rate your money actually grows at is the geometric mean, and it is always the lower of the two when returns vary.
How big is the drag, and where does half the variance come from?
For returns that are not too large, the geometric mean is close to the arithmetic mean minus half the variance. With 20% volatility the variance is 0.04, half of that is 0.02, so a 10% average compounds at about 8%. The two-year example makes it concrete: +30% and -10% average 10% with a standard deviation of 20%, and 1.30 x 0.90 = 1.17, a compound rate of 8.17% a year. The rule of thumb says 8.00%, close enough to trust in an interview.
Holding the average return at 10%, compound growth falls to 8% at 20% volatility and to 2% at 40%, because the drag is about half the variance; two years of +30% and -10% average 10% but compound at 8.17% a year. The relationshipg the compound annual growth rate mu the average annual return, 10% sigma annual volatility, 20% What it says in wordsCompound growth equals the average return less half the variance.Add why an allocator asks this. Two funds with the same average return and different volatility do not leave investors with the same money. Cutting volatility from 20% to 10% raises compound growth by 1.5 points with no change in the average, which is part of why lower-volatility strategies can be worth more than their averages suggest. The limitation: the half-variance rule is an approximation that weakens for very volatile or fat-tailed returns.
Where candidates lose it
The common loss is answering 10%, treating the average return as the growth rate. The interviewer wants to hear the word compounding and a number for the drag.
The second loss is subtracting the full variance or the volatility itself, giving 6% or -10%. It is half the variance, and 20% squared is 4%, not 40%. Say the rule, give 8%, and check it with a two-year example.
What the interviewer asks next
- At what volatility does a 10% average return compound to zero?
- Fund A averages 12% with 30% volatility; fund B averages 10% with 15%. Which grows money faster?
- How does leverage change the answer, and what leverage maximises compound growth here?
038X and Y are independent random variables with the same variance. What is the correlation between X and X + Y?Squarepoint CapitalMontreal · 2026
Try it first
Pick one:
Show the worked solution
1 over root 2, about 0.71. The covariance of X with X + Y is Var(X) plus Cov(X, Y), which is sigma squared plus zero. The standard deviation of X + Y is root 2 times sigma because the variances add. So the correlation is sigma squared over (sigma x root 2 sigma), which is 1/root 2. X explains half the variance of the sum, and the correlation is the square root of that half.
What is the fastest way to set it up?
A two-member team's score is the sum of both players' scores. If the players are equally good and play independently, knowing one player's score tells you something about the team total, but only half the story. Split the covariance: Cov(X, X + Y) = Cov(X, X) + Cov(X, Y) = sigma squared + 0. The variance of the sum is sigma squared + sigma squared = 2 sigma squared, because independent variances add. Correlation is covariance over the product of standard deviations: sigma squared over (sigma x root 2 sigma) = 1/root 2.
Drawn as arrows, independent X and Y sit at right angles and their sum lies at 45 degrees to X, so the correlation is cos 45, about 0.707; equivalently, X supplies half of the variance of X + Y, and the correlation is the square root of one half. Why is the answer not 0.5?
Because 0.5 is the R squaredThe share of one variable variance explained by another; for a simple regression it is the correlation squared., not the correlation. X explains exactly half of the variance of X + Y, and correlation is the square root of the share of variance explained, so it is root 0.5, about 0.707. The geometric picture makes it stick: treat independent variables as arrows at right angles, and correlation as the cosine of the angle between arrows. X + Y sits at 45 degrees to X, and cos 45 is 0.707.
Give the general version to show you own it. If Y has variance k times X's, the correlation is 1/root(1 + k): the more noise you add, the lower it falls. That is the logic behind a noisy signal: a forecast that is half signal and half independent noise, by variance, correlates about 0.71 with the signal, not 0.5.
Where candidates lose it
The common loss is answering 0.5 because X is half of the sum. That is the share of variance, and correlation is its square root.
The other loss is saying zero because X and Y are independent. The sum contains X, so it cannot be independent of X. Split the covariance in one line and the answer falls out.
What the interviewer asks next
- What is the correlation between X + Y and X - Y?
- Y has four times the variance of X. What is corr(X, X + Y) now?
- What is the correlation between the sum of the first 10 and the sum of the first 20 of a series of independent returns?
Asked at Squarepoint Capital, Desk Quant Analyst Interview, Montreal, 2026 (Wall Street Oasis):
There were also 3-4 basic math/stats questions about mean, covariance, correlation, etc.
039Depreciation rises by Rs 10 and the tax rate is 25%. Walk the change through net income, the cash flow statement and the balance sheet.Millennium ManagementNew York · 2024
Try it first
What happens to cash?
Show the worked solution
Net income falls Rs 7.5, cash rises Rs 2.5 and the balance sheet shrinks by Rs 7.5 on both sides. Pre-tax profit falls 10, tax falls 2.5, so net income falls 7.5. The cash flow statement starts at -7.5 and adds back the non-cash 10: cash up 2.5. On the balance sheet, cash is up 2.5 and fixed assets are down 10, so assets fall 7.5, matched by retained earnings down 7.5.
Why does cash go up when an expense goes up?
Imagine your employer lets you deduct the wear on your car from taxable income. No money leaves your pocket for the wear itself, but your tax bill falls. Depreciation is an expense that costs no cash but reduces tax, so the only cash effect is the tax saved: 25% of Rs 10, Rs 2.5. That is the {term('depreciation tax shield', 'The tax saved because depreciation is deductible even though it uses no cash; equal to depreciation times the tax rate.')}, and it is the one number the question is testing.
A Rs 10 rise in depreciation at a 25% tax rate cuts net income by Rs 7.5, raises cash by Rs 2.5 through the tax saved, and lowers fixed assets by Rs 10, so total assets and retained earnings both fall by Rs 7.5 and the balance sheet balances. What order do you walk it in so nothing gets lost?
Income statement first, then cash flow, then balance sheet, one line each. Net income is the bridge: it closes the income statement, opens the cash flow statement, and lands in retained earnings on the balance sheet. Income statement: depreciation +10, pre-tax -10, tax -2.5, net income -7.5. Cash flow: -7.5 plus 10 added back, cash +2.5. Balance sheet: cash +2.5, fixed assets -10, so assets -7.5; retained earnings -7.5, so the two sides move together.
Add one sentence on why a hedge fund analyst cares. Two companies with identical operations can report different earnings because of depreciation choices, while their cash generation differs only by the tax effect. That is one reason investors look at cash flow alongside earnings before trusting a P/E.
Where candidates lose it
The common loss is saying cash is unchanged because depreciation is non-cash. That forgets the tax: depreciation is deductible, so the tax bill falls and cash rises by Rs 2.5.
The second loss is saying cash falls 7.5 by reading net income as cash. Walk the add-back out loud and check that assets and equity both fall by 7.5 before you stop.
What the interviewer asks next
- Now the depreciation rise comes from a Rs 10 write-down of an asset that is not tax deductible. What changes?
- What if the company is loss-making and pays no tax this year?
- Walk a Rs 10 rise in inventory, bought with cash, through the three statements.
Asked at Millennium Management, Investment Research, New York, 2024 (Wall Street Oasis):
Nothing as much, technical questions were super basic like $10 depreciation
040A bag holds four stones, each black or white. Before you look, every count of black stones from 0 to 4 is equally likely. You draw two stones without replacement and both are black. What is the chance the next stone is black, and at what price would you bet on it?CitadelNew York · 2025
Try it first
Chance the third stone is black:
Show the worked solution
3/4, so fair odds are 3 to 1 on black. Two black draws rule out bags with 0 or 1 black. The ways to draw two blacks in order are 2, 6 and 12 for bags with 2, 3 and 4 black, so those bags now carry 10%, 30% and 60%. The next stone is black with chance 0, 1/2 and 1 in them, which averages to 3/4. A contract paying 100 if black is worth 75.
How do the two black draws change your view of the bag?
If a friend pulls two red sweets from a jar you have never seen, you start to suspect it is mostly red. Each possible bag is reweighted by how likely it was to produce what you saw: equal priors times the chance of two blacks. Drawing two blacks in order has 0 ways from bags with 0 or 1 black, 2 x 1 = 2 ways with 2 black, 3 x 2 = 6 with 3 black and 4 x 3 = 12 with 4 black. Out of 20 in total, that is 10%, 30% and 60%: the all-black bag is now the favourite.
Before the draw each bag is 20% likely; after two black stones the bags with 2, 3 and 4 black carry 10%, 30% and 60%, and since the next stone is black with probability 0, one half and 1 in those bags, the chance it is black is 3/4. How do you turn the posterior into a price?
Average the chance of black over the bags you still believe in. In the 2-black bag both remaining stones are white; in the 3-black bag one of two is black; in the 4-black bag both are, so the answer is 0.1 x 0 + 0.3 x 1/2 + 0.6 x 1 = 3/4. A contract paying 100 if the next stone is black is worth 75. You would buy it below 75 and sell it above; as a market maker you might quote 70 at 80. Offered even money on black, your expected profit per rupee staked is 0.75 minus 0.25, which is 50 paise.
The relationshipk the number of black stones in the bag B1 B2 the event that the first two draws are black 2, 6, 12 the ordered ways to draw two blacks from bags with 2, 3 and 4 black What it says in wordsThe chance of another black is the average of each bag's chance, weighted by how much the evidence now favours that bag.Check it with Laplace rule of successionWith a uniform prior, after s successes in n trials, the chance the next trial succeeds is (s + 1) / (n + 2).: after 2 blacks in 2 draws the next is black with chance (2 + 1)/(2 + 2) = 3/4, and the rule holds exactly for this finite bag. Two routes to 3/4 is what separates a solid answer from a lucky one. The limitation: everything rests on the flat prior. If you had reason to think mixed bags were more common, 3/4 would fall.
Where candidates lose it
The common loss is saying 1/2 because the remaining stones are unknown, which throws away the information in the two draws. The question is about updating, and the interviewer wants to see the reweighting.
The second loss is weighting the surviving bags equally, a third each, which gives 1/2. Each bag must be weighted by how likely it made two blacks: 2, 6 and 12. Then price it: a probability without a bet is half the answer at a trading firm.
What the interviewer asks next
- The third stone is black too. What is the chance the fourth is black?
- You quote 70 at 80 on a contract paying 100 if black and someone who has seen the bag lifts your offer. What now?
- How does the answer change if the prior is that each stone is black with probability one half, independently?
Asked at Citadel, Quantitative Trading, New York, 2025 (Wall Street Oasis):
Extended bayes derivative question about four stones in a bag (black and white stones).
