Hedge Funds puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 38
- Topics
- 14
- Hard
- 30
002The classic Russian roulette puzzle: a six-chamber revolver has two bullets in adjacent chambers. The cylinder is spun once, the trigger is pulled and it clicks empty. You must pull again. Is it safer to spin the cylinder again first, or not?SchonfeldCentral · 2022
Try it first
Which gives the better chance of surviving the second pull?
Show the worked solution
Do not spin: you survive 75% of the time, against 66.7% if you spin. The empty click puts the cylinder on one of the four empty chambers, each equally likely. Because the two bullets sit together, three of those four empties are followed by another empty and only one is followed by a bullet. A fresh spin throws that information away and gives four chances in six.
What does the empty click tell you?
Picture six people in a queue where two friends always stand together. Pick someone at random who is not one of the friends and ask whether the person behind them is a friend. Only one of the four has a friend behind them: the one standing just in front of the pair. The empty click is information: it tells you which chambers you could be on, and using it is the whole puzzle. The cylinder has no memory, but you do.
After an empty click the cylinder sits on chamber 3, 4, 5 or 6; three of those are followed by an empty chamber and only chamber 6 is followed by a bullet, so not spinning survives 75% of the time against 66.7% for a fresh spin. The relationship3/4 three of the four empty chambers are followed by another empty 4/6 four of six chambers are empty after a fresh spin What it says in wordsConditioning on the empty click beats resetting to the base rate when the bullets sit together.Why does the answer flip if the bullets are not adjacent?
Separate the bullets, say into chambers 1 and 4. The four empties are now 2, 3, 5 and 6, and the chambers after them are 3, 4, 6 and 1: two empties and two bullets. Not spinning now survives only 2 times in 4, 50%, so spinning, at 66.7%, becomes the better choice. Adjacency is what bunches both bullets behind a single empty chamber. Ask where the bullets sit before you answer, and say that the answer depends on it.
Where does this reasoning show up on a desk?
The same move, updating on what you have just observed instead of resetting to the base rate, is how a trader reads a fill. Getting filled on your bid tells you something about who was selling, just as the empty click tells you which chamber you are on. Ignoring it is the equivalent of spinning the cylinder: it feels neutral, but it throws away an edge you were handed for free.
Where candidates lose it
Candidates say it makes no difference, because a spin feels like a clean reset and the cylinder has no memory. The trap is treating no memory in the device as no information for you: the click has ruled out the two loaded chambers as your position.
The second loss is answering without checking the layout. The case for not spinning rests entirely on the bullets being adjacent; with the bullets apart, the answer reverses. Name that condition in your answer.
What the interviewer asks next
- You survive the second pull without spinning. Should you spin before a third?
- Three bullets in adjacent chambers: spin or not?
- What if the two bullets are in chambers 1 and 4?
Asked at Schonfeld, Quantitative Research, Central, 2022 (Wall Street Oasis):
Coding, requires to know DP and divde and conquer., Russian Roulette
003X and Y are independent and each uniform on 0 to 1. What is the probability that X + Y is less than 1.5, and what shape is the density of X + Y?CitadelChicago · 2025
Try it first
Pick before you draw anything.
Show the worked solution
The probability is 7/8, and the density of X + Y is a triangle, a tent peaking at 1. Because X and Y are independent and uniform, every point of the unit square is equally likely, so probability is area. The line x + y = 1.5 slices off a corner triangle with legs of 0.5, area 1/8. The sum's density rises in a straight line from 0 to 1 and falls back to 0 at 2.
Why does probability become area here?
Throw a dart at a square board so that every point is equally likely to be hit. The chance it lands in a region is that region's share of the board. Two independent uniforms are exactly that dart: the pair (X, Y) lands evenly on the unit square, so any question about X and Y becomes a question about an area. The condition X + Y below 1.5 is everything under the line x + y = 1.5, which is the whole square except one corner.
The line x + y = 1.5 removes a corner triangle of area 1/8 from the unit square, so X + Y is below 1.5 with probability 7/8, and the density of X + Y is a tent on 0 to 2 whose tail beyond 1.5 also has area 1/8. How do you get the shape of the sum's density?
Slide the line x + y = s across the square and watch how long it is inside. Near s = 0 it barely clips the corner; at s = 1 it runs corner to corner, the longest it gets; past 1 it shortens again. The density of the sum at s is proportional to the length of that line inside the square, which gives a triangle rising from 0 to a peak at 1 and falling to 2. This is the convolutionThe density of a sum of independent variables, found by adding up every way the two parts can combine to the same total. of two flat densities, and the same reason two dice most often total 7.
The relationshipf_{X+Y}(s) the density of the sum at the value s f_Y(s - x) equal to 1 when s - x lies between 0 and 1, otherwise 0 What it says in wordsAdd up every split of s into an x and a y that both lie in 0 to 1; the count of splits rises to s = 1 and then falls.Check the first answer with the tent. The area beyond 1.5 is a triangle with base 0.5 and height 0.5, which is 1/8 again. Two routes that agree is the check an interviewer wants to hear before you commit. Add a third uniform and the density becomes three joined curved pieces; add many and the sum looks normal, which is the central limit theorem arriving in slow motion.
Where candidates lose it
Candidates reach for a double integral before drawing, set the limits wrongly, and spend two minutes on what is a one-line area argument. Draw the square first; the corner triangle is visible at a glance.
The second loss is saying the sum of two uniforms is uniform on 0 to 2. It is not: there is only one way to get a sum near 0 and many ways to get a sum near 1, which is why the density is a tent and not a flat line.
What the interviewer asks next
- What is the probability that X + Y is less than 0.5?
- What is the probability that the larger of X and Y is below 0.5, and how does the picture change?
- What does the density of X + Y + Z look like?
Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):
He was asking some questions about the probability, especially on the convolution.
005Without paper, what are 997 x 1,003 and 67 squared?Akuna Capitalchicago · 2024
Try it first
What is 997 x 1,003?
Show the worked solution
997 x 1,003 = 999,991 and 67 squared = 4,489. The first is a difference of squares: the numbers sit 3 either side of 1,000, so the product is 1,000,000 minus 9. The second uses the same identity the other way round: 67 squared is 64 x 70 plus 3 squared, which is 4,480 plus 9. Both take one line once you spot the round number nearby.
Why does multiplying around a round number work?
Take a square garden 10 metres a side and reshape it to 13 by 7. The fence is the same length, but the plot shrinks from 100 square metres to 91. It always shrinks by the square of how far you moved each side, here 3 squared, 9. Two numbers spaced equally around a midpoint multiply to the midpoint squared minus the gap squared. For 997 x 1,003 the midpoint is 1,000 and the gap is 3, so the answer is 1,000,000 minus 9.
Removing a b by b corner from an a by a square and standing the leftover strip on its end makes a rectangle a + b wide and a - b tall, which is why 997 x 1,003 is 1,000,000 - 9 = 999,991 and 67 squared is 64 x 70 + 9 = 4,489. The relationshipa the round midpoint, or the number being squared b the gap you choose to make a factor round What it says in wordsA product around a midpoint is the midpoint squared less the gap squared; run it backwards to square any number.How do you square a number like 67 in your head?
Push it to a round neighbour and repair the difference. Move 3 down to 64 and 3 up to 70, multiply those, then add back the 3 squared you took away: 64 x 70 = 4,480, plus 9 is 4,489. You choose the gap so one factor is round. The expansion route agrees: 67 is 70 minus 3, so 67 squared is 4,900 - 420 + 9, the same 4,489. Two methods landing on one number is your check.
Why would a fund ask arithmetic at all?
A trader checks prices, spreads and position sizes in their head all day, and a slip costs money before any spreadsheet catches it. Interviewers use speed on arithmetic like this as a proxy for how quickly you would catch a quote that does not add up. Say the identity as you use it, so that if you slip, the interviewer can see where and you can recover out loud.
Where candidates lose it
The trap is grinding 997 x 1,003 column by column under time pressure, dropping a carry and landing on 1,000,009 or 999,909. Two numbers either side of 1,000 is the whole question, and it is there to see whether you notice.
For 67 squared, candidates who remember 65 squared is 4,225 try to count up from there and lose track of the cross term. Pick the route that gives you one round multiplication, and say it out loud as you go.
What the interviewer asks next
- What is 48 x 52?
- What is 95 squared, and what is the fastest route?
- Estimate 1.03 to the tenth power in your head.
Asked at Akuna Capital, Hedge Fund, chicago, 2024 (Wall Street Oasis):
focusing on quick probability puzzles, mental math, and some data structure questions
006You owe exactly Rs pi, that is Rs 3.14159..., and can only pay in whole paise. How do you pay a fair amount on average, and what is the chance you end up paying Rs 3.15?Millennium ManagementSheung Wan · 2025
Try it first
Under the fair scheme, what is the chance you pay Rs 3.15?
Show the worked solution
Randomise: pay Rs 3.15 with probability 0.159 and Rs 3.14 otherwise. Pi is 3.14159..., which sits 0.159 of the way from 3.14 to 3.15. Paying the higher amount with exactly that probability makes the expected payment 3.14 + 0.01 x 0.1593..., which is pi. So the chance you pay Rs 3.15 is about 15.9%, and over many meals nobody is short-changed.
Why can no fixed amount be fair?
Always round to Rs 3.14 and the restaurant loses 0.159 paise every time; always pay 3.15 and you overpay 0.841 paise. Any fixed amount is unfair to one side, so the only way to be exactly fair is to be fair on average. Two friends who split a Rs 101 bill by taking turns to pay the odd rupee are doing the same thing: neither is exact on any one night, both are exact over time.
Pi sits 0.159 paise above Rs 3.14 and 0.841 paise below Rs 3.15, so paying Rs 3.15 with probability 15.9% and Rs 3.14 otherwise balances exactly at pi, which makes the expected payment fair. The relationshipp the probability of paying Rs 3.15 \pi - 3.14 how far pi sits above the lower whole-paisa amount What it says in wordsThe chance of paying the higher amount equals how far along the gap pi lies.How do you actually draw a probability of 0.159?
Use any randomness you can split finely. Draw a uniform number between 0 and 1 and pay Rs 3.15 if it falls below 0.1593. With only a die, paying 3.15 on a six gives 1/6, an expected payment of Rs 3.141667: close, not exact. With only a fair coin you can be exact: toss it to generate the binary digits of a uniform number one at a time and stop as soon as the digits so far settle which side of 0.1593 it falls. Each toss settles it with probability one half, so on average it takes two tosses.
Where does randomised rounding show up in a fund?
Whenever a quantity has to be split in whole units. A fund allocating 1,003 shares across three accounts cannot give each 334.33; handing the odd share out by lottery, or in rotation, keeps each account fair on average. The principle is the same: when the exact amount is impossible, make the expected amount exact and keep the error unbiased. The limitation is that fair on average is not fair every time, which is why allocation policies also cap how far any account can drift.
Where candidates lose it
Candidates round to Rs 3.14 and argue the gap is too small to matter. The interviewer is not asking about a sixth of a paisa; the question is whether you see that a fair expected value can be built from amounts that are each individually wrong.
The second loss is the coin flip. Fifty-fifty between 3.14 and 3.15 feels even-handed but averages 3.145, overpaying by nearly half a paisa every time. The probability has to match where pi sits in the gap.
What the interviewer asks next
- How would you hit the probability exactly using only a fair coin?
- How many coin tosses does that take on average?
- What if you owe Rs e, 2.71828...?
Asked at Millennium Management, Quantitative Research, Sheung Wan, 2025 (Wall Street Oasis):
How to pay the restaurant fairly if I owe pi dollars. Need to pay with usual dollars and cents.
008A stock closes at 100, 96, 104, 99, 110, 105 and 112 on seven days, and short selling is not allowed. What is the maximum profit from one buy and one sell, and from any number of round trips?Man GroupLondon · 2019
Try it first
What is the most you can make with any number of round trips?
Show the worked solution
One round trip makes at most 16; unlimited round trips make 26. For one trade, walk the prices once, carrying the lowest price so far and the best sale against it: buy at 96, sell at 112. For many trades, add every day-on-day rise and skip every fall: 8 + 11 + 7 = 26. Without short selling the falls are simply sat out, never profited from.
How do you find the best single trade without checking every pair?
Imagine walking down a street of shops that all sell the same phone, planning to buy once and sell once further along. You do not need to compare every pair of shops: carry the cheapest price seen so far in your head, and at each shop ask what selling here would make against it. One pass, keeping the running minimum and the best gap found so far, gives the best single trade. Here the running minimum drops to 96 on day 2 and the best gap appears on day 7: 112 - 96 = 16.
The best single trade buys at 96 on day 2 and sells at 112 on day 7 for 16, while trading every rising leg, 96 to 104, 99 to 110 and 105 to 112, collects 8 + 11 + 7 = 26. Day Price Move Lowest so far Best single trade so far Sum of rises so far 1 100 100 0 0 2 96 -4 96 0 0 3 104 +8 96 8 8 4 99 -5 96 8 8 5 110 +11 96 14 19 6 105 -5 96 14 19 7 112 +7 96 16 26 One pass through the prices tracks both answers at once: the running minimum gives the best single trade, 16, and the running sum of positive moves gives the many-trade maximum, 26. Why is the many-trade answer just the sum of the rises?
Any rise from a low to a later high is the sum of the daily steps inside it, and some of those steps may be falls. With no short selling and no costs, the most you can make is the total of every positive day-on-day move, 26 here, because trading only the up steps collects everything a longer trade would and skips its falls. In practice that is three round trips: buy 96, sell 104; buy 99, sell 110; buy 105, sell 112.
What does the interviewer add next?
Costs. Once each round trip costs something, the sum of rises overstates the profit, because small moves stop being worth trading. With a cost of 6 per round trip, the three separate trades net 2 + 5 + 1 = 8, the best two-trade split nets 9, and the single trade from 96 to 112 nets 10, so the single trade now wins. The general version is a short dynamic programme that tracks the best profit on each day while holding and while flat.
Where candidates lose it
For the first part, candidates take the lowest and highest prices without checking the order. Here they happen to line up, 96 before 112, but an interviewer who swaps two prices will catch anyone who never checked that the low comes first.
For the second, the loss is counting falls as profit, which needs a short sale the question forbids, or stopping at 16 because it is the best single trade. Say the rule plainly: bank every rise, sit out every fall.
What the interviewer asks next
- What if each round trip costs 6?
- What if you may make at most two round trips?
- How does the answer change if short selling is allowed?
Asked at Man Group, Alternative Investments, London, 2019 (Wall Street Oasis):
Given a series of prices, find the one buy/sell trade pair which gives the maximum profit
014A corporate bond has a spread duration of 6 and convexity of 50. Its credit spread widens by 50 basis points. Roughly what happens to its price?AQR Capital ManagementGreenwich · 2021
Try it first
Which is closest?
Show the worked solution
The price falls by about 2.94%. Spread duration of 6 says a 0.50 percentage point widening costs 6 x 0.50% = 3.00%. Convexity of 50 adds back one half x 50 x 0.005 squared, about 0.06%, because the price curve bends upwards. On a bond priced at 100 that is a move to about 97.06. At 50 basis points the convexity term is small; at 300 or 500 it is not.
What do duration and convexity each measure?
Picture a playground slide that curves and flattens towards the bottom. Judge the drop from the steepness at the top and you overstate it, because the slide levels off as you go. Spread durationThe percentage change in a bond price for a one percentage point change in its credit spread, holding the risk-free rate fixed. is the steepness at today's spread; convexity is the flattening, so the straight-line estimate always overstates the loss when spreads widen. Duration gives the first-order move, 6 x 0.50% = 3.00% down; convexity corrects it by a term that depends on the square of the move.
For a 50 basis point widening, duration of 6 gives minus 3.00% and convexity of 50 adds back 0.06%, a fall of 2.94%; the convexity cushion grows with the square of the move, to 2.25 points at 300 basis points and 6.25 at 500. The relationshipD_s spread duration, 6 C convexity, 50 \Delta s the change in spread as a decimal, 50 basis points = 0.005 What it says in wordsThe price moves by the duration term plus a smaller correction that grows with the square of the spread change.When does the convexity term start to matter?
It grows with the square of the move. At 50 basis points convexity is worth 0.06% against a 3.00% duration loss; at 300 basis points duration says -18% and convexity adds back 2.25%, which is no longer small. That is why a credit desk can run duration-only risk for everyday moves but needs convexity for stress scenarios. One more distinction marks a strong answer: for a fixed-coupon bond spread duration and rate duration are close, but a floating-rate note has almost no rate duration and still carries several years of spread duration.
Say the limitation plainly. Both numbers are local, measured at today's spread, and a distressed bond stops behaving like this long before default, when its price starts tracking the expected recovery instead. For a bond trading near par, as here, the two-term estimate is good to a few hundredths of a per cent for moves of this size.
Where candidates lose it
Candidates give minus 3% and stop, which is fine as a first line but ignores the second number the question handed you. Worse is using convexity with the wrong sign, making the loss bigger: for a plain bond convexity always cushions a spread widening.
The other slip is units. Fifty basis points is 0.005 in the formula; squaring 0.50 instead turns a 0.06% correction into 6.25% and produces a price that rises when spreads widen.
What the interviewer asks next
- What if the spread tightens by 50 basis points instead?
- Why can a callable bond have negative convexity?
- How would you hedge the spread risk of this bond?
Asked at AQR Capital Management, Investment Research, Greenwich, 2021 (Wall Street Oasis):
Discussion on credit spreads on fixed income products and duration.
020A book holds Rs 60 crore of a stock with 30% volatility and Rs 40 crore of another with 20% volatility, and the two have a correlation of 0.5. What is the book's volatility in rupees, and what share of the risk comes from each position?Man GroupBoston · 2022
Try it first
What share of the book's risk comes from the Rs 60 crore position?
Show the worked solution
The book's volatility is about Rs 23.1 crore a year, and the Rs 60 crore position carries about 74% of it on 60% of the capital. Stand-alone risks are Rs 18 crore and Rs 8 crore. Book variance is 18 squared plus 8 squared plus 2 x 0.5 x 18 x 8, which is 532, so volatility is Rs 23.07 crore. Each position's contribution is its covariance with the book over the book's volatility: Rs 17.17 crore and Rs 5.90 crore, which add back to the total.
Why is risk not shared out like capital?
Two friends share a taxi. One rides twice as far, straight through the traffic jam; the other gets off after a short hop. Splitting the fare by the number of bags each carries would be absurd. Risk belongs to a position in proportion to how much it moves and how much it moves with everything else, not to how much money sits in it. Here the first stock is larger, more volatile and positively correlated with the second, so it carries far more than its 60% of the capital.
Stand-alone risks of Rs 18 crore and Rs 8 crore add to Rs 26 crore, diversification at a 0.5 correlation removes Rs 2.93 crore, and the book's Rs 23.07 crore of volatility splits 74.4% to the first position and 25.6% to the second, against a 60 to 40 split of capital. The relationshipa, b stand-alone rupee volatilities: 60 x 30% = 18 and 40 x 20% = 8 \rho the correlation, 0.5 RC a position's contribution to book volatility What it says in wordsEach position owns its own variance plus half the shared term, and dividing by the book's volatility turns that into rupees of risk.Position Capital, Rs crore Volatility Stand-alone risk Risk contribution Share of risk A 60 30% 18.0 17.17 74.4% B 40 20% 8.0 5.90 25.6% Book 100 26.0 23.07 100.0% Rs crore of annual volatility. The stand-alone risks add to Rs 26.0 crore, but the book's volatility is Rs 23.07 crore, of which position A contributes 74.4% and position B 25.6%. Why do the contributions add up exactly to the total?
Split the variance. The cross term, 2 x 0.5 x 18 x 8 = 144, is shared equally, 72 to each position. So position A owns 324 + 72 = 396 of the 532 of variance and position B owns 64 + 72 = 136, and dividing each by the book's volatility of 23.07 gives rupee contributions that sum exactly to Rs 23.07 crore. The diversification benefit is the gap between the stand-alone total of Rs 26 crore and the book's Rs 23.07 crore.
What does a risk manager do with the split?
Cut where the risk is, not where the money is. Each rupee in position A carries 28.6 paise of marginal risk against 14.7 paise in position B, so trimming Rs 10 crore from A lowers book volatility by roughly Rs 2.9 crore. Recomputing exactly gives Rs 2.84 crore, close to the estimate. The limitation: the split is a snapshot at one correlation, and when correlations move, both the total and the split move with them.
Where candidates lose it
The quick answer shares risk like capital, 60 and 40, or like stand-alone risk, 18 and 8. The first ignores volatility and the second ignores correlation; neither sums to the book's actual Rs 23 crore of risk.
The second slip is adding the stand-alone risks to get the book's risk, Rs 26 crore. Volatilities do not add unless the correlation is exactly 1; variances do, with the cross term included.
What the interviewer asks next
- If the correlation fell to zero, how would the risk split between the two positions?
- How much of position B would you add to minimise the book's volatility, holding A fixed?
- How do transaction costs change which position you trim first?
Asked at Man Group, Investment Management, Boston, 2022 (Wall Street Oasis):
How do you understand portfolio risk and transaction cost?
025Four observations, 3.1, 7.4, 5.2 and 9.0, come from a uniform distribution on 0 to theta. What is the maximum likelihood estimate of theta, is it biased, and how would you correct it?AQR Capital ManagementTown of Greenwich · 2022
Try it first
Which statement is right?
Show the worked solution
The MLE is 9.0, the largest observation; it is biased low, and the unbiased correction is 5/4 x 9.0 = 11.25. The likelihood is 1 over theta to the fourth for any theta of at least 9.0 and zero below it, so it peaks at the sample maximum. But the maximum of n draws averages n/(n + 1) of theta, here 4/5, so scaling by (n + 1)/n removes the bias.
Why is the MLE the largest observation?
A friend draws four raffle tickets numbered from 1 up to some unknown top number, and the highest you see is 90. The top number is at least 90; guessing higher only spreads your belief over tickets nobody drew. The likelihood, 1 over theta to the n, is zero for any theta below the largest observation and falls as theta rises above it, so it is maximised exactly at the sample maximum, 9.0. This is a case where you do not differentiate: the maximum sits on a boundary, not where a slope is zero.
Four uniform draws on 0 to theta cut it into five gaps of equal expected size, so the largest draw averages four fifths of theta; the MLE of 9.0 therefore sits below theta, and scaling by 5/4 gives the unbiased 11.25, against 12.35 from doubling the sample mean. The relationshipL(\theta) the likelihood of the four observations n the number of observations, 4 \max the largest observation, 9.0 What it says in wordsThe likelihood peaks at the largest observation, which on average falls short of theta by a factor n/(n + 1).Why is it biased, and by how much?
The largest draw can never be above theta and is almost always below it. Four points dropped at random on 0 to theta cut it into five gaps of equal expected length, so the largest point sits on average four fifths of the way up, and the MLE underestimates theta by a fifth on average. Multiplying by 5/4 fixes it: 9.0 becomes 11.25. The bias shrinks as n grows, since n/(n + 1) tends to 1, but with four points it is large.
How does it compare with the obvious alternative?
The method of moments doubles the sample mean, since a uniform on 0 to theta averages theta/2: the mean here is 6.175, giving 12.35. Both 11.25 and 12.35 are unbiased, but the corrected maximum has a much smaller variance, theta squared over n(n + 2) against theta squared over 3n, because the largest draw carries the most information about the top of the range. With four points that is theta squared over 24 against theta squared over 12: half the variance.
Where candidates lose it
Candidates differentiate the log-likelihood, get minus n over theta, set it to zero and find no solution. The likelihood only falls on the allowed range, so the maximum sits at the boundary, the largest observation; say that before reaching for calculus.
The second miss is calling the MLE unbiased because maximum likelihood estimates are often well behaved. Here it is biased low by construction, and the interviewer expects the (n + 1)/n correction.
What the interviewer asks next
- What is the variance of the corrected estimator with four observations?
- What is the MLE if the distribution is uniform on theta to 2 theta?
- Derive the expected value of the maximum of n uniform draws.
Asked at AQR Capital Management, Trading, Town of Greenwich, 2022 (Wall Street Oasis):
Derive the mle for some given distribution. Explain linear regression intuitively and derive the ols estimate.
026You have five trade ideas. Each needs some units of risk budget and carries an expected profit: A needs 4 units for Rs 9 crore, B 3 units for Rs 7 crore, C 5 units for Rs 10 crore, D 2 units for Rs 4 crore and E 6 units for Rs 11 crore. Your budget is 10 units and you cannot take part of an idea. Which ideas do you take?Bridgewater AssociatesNew York · 2025
Try it first
Before you enumerate: which set do you expect to win?
Show the worked solution
Take B, C and D: 10 units for Rs 21 crore. Ranking by profit per unit of risk picks B (2.33), then A (2.25), then D (2.00), which uses 9 units for Rs 20 crore and leaves one unit idle. Swapping A for C costs a little ratio but fills the budget and adds Rs 1 crore. With whole ideas, you check the few sets that fill the budget rather than trust the ratio alone.
Why is profit per unit of risk the right place to start?
Packing a suitcase for a flight with a weight limit, you favour the things that give the most use per kilo. When the budget is the scarce thing, the useful measure is profit per unit of budget, not profit alone. E earns the most in rupees but only 1.83 per unit, the worst of the five. B earns 2.33 per unit, A 2.25, and C and D 2.00 each. If ideas could be split, you would simply fill from the top of that list and the answer would be exact.
Why does the ratio ranking fail here?
Ideas come whole. Take B and A and you have used 7 units; D fits for 9 units, and nothing else fits in the last one. A unit left idle earns nothing, so a set with a slightly lower average ratio that uses the full budget can beat the greedy pick. Replace A (4 units) with C (5 units) and the budget is full: B, C and D earn 7 + 10 + 4 = Rs 21 crore against Rs 20 crore. This is the knapsack problemChoosing whole items, each with a size and a value, to maximise total value without exceeding a capacity., and the fix is to check the handful of sets that nearly fill the budget.
B, C and D use all 10 risk units for Rs 21 crore, the best set that fits; ranking by profit per unit picks B, A and D, which uses 9 units for Rs 20 crore, and A, B and C would earn Rs 26 crore but needs 12 units. With five ideas there are only 31 possible sets, so enumerate the full ones out loud: B, C, D for 21; A, B, D for 20; A, E for 20; A, C for 19. On a real book with hundreds of positions a desk uses an optimiser, but the logic is the same, and interviewers want to hear that you know the ratio rule is a starting point and not a proof.
Say the limitation too. Expected profit is an estimate, and the gap here is Rs 1 crore on Rs 21 crore. If C's estimate is shakier than A's, a portfolio manager could reasonably prefer the greedy set. The arithmetic tells you the best set on the stated numbers; confidence in each number decides whether that edge is real.
Where candidates lose it
The common loss is ranking by profit per unit, taking the top three, and stopping. That is right for divisible positions and wrong here, because the tenth unit of budget sits unused and the interviewer built the numbers to punish exactly that.
The opposite loss is ranking by rupee profit and grabbing E first. E is the least efficient idea on the list. Say the ratio rule, then check the sets that fill the budget, then say how sure you are of each estimate.
What the interviewer asks next
- The budget rises to 11 units. What changes?
- Ideas B and C are highly correlated, so together they use 9 units instead of 8. Does your answer move?
- You can take half of any idea for half its units and half its profit. What do you take now?
Asked at Bridgewater Associates, Generalist, New York, 2025 (Wall Street Oasis):
Given a list of items and their utilitites, give an algorithm to maximise utility
027A trade surveillance screen flags 90% of genuinely suspicious trades and wrongly flags 5% of clean ones. One trade in a hundred is suspicious. A trade has just been flagged. What is the chance it is actually suspicious?CitadelMiami · 2022
Try it first
Gut answer first: a flagged trade is suspicious with probability about
Show the worked solution
About 15%, not 90%. Picture 10,000 trades. 100 are suspicious and the screen flags 90 of them. 9,900 are clean and it wrongly flags 5%, which is 495. So 585 trades are flagged and only 90 are suspicious: 90 over 585 is 15.4%. The false flags swamp the true ones because clean trades are so common.
Why is 90% the wrong number?
A smoke alarm that goes off whenever there is a fire is good. But if it also goes off every time someone makes toast, most of its alarms are toast. The 90% tells you how often a suspicious trade gets flagged; the question asks how often a flag is suspicious, and those run in opposite directions. Mixing them up is called the {term('base rate', 'How common something is before any test is run. Here, one trade in a hundred is suspicious.')} fallacy, and it is exactly what this question is built to catch.
How do you get the number without the formula?
Use counts, not percentages. Start with 10,000 trades because it makes every number whole. Split by the truth first, then by what the screen says, and then read only the flagged column. 1% of 10,000 is 100 suspicious trades; 90% of those, 90, are flagged. 9,900 are clean; 5% of those, 495, are flagged anyway. The flagged column holds 585 trades, and 90 of them are the real thing.
Of 10,000 trades, 100 are suspicious and 90 of those are flagged, while 495 of the 9,900 clean trades are flagged by mistake, so only 90 of the 585 flags, 15.4%, point at a suspicious trade. The relationshipS the trade is suspicious F the screen flags the trade 0.01 the base rate: one trade in a hundred is suspicious What it says in wordsThe chance a flag is right equals true flags divided by all flags.Add the desk point after the number. A second, independent check changes things fast: run the 585 flagged trades through a second screen with the same error rates and the 15.4% prior becomes about 77%. A weak test is still useful as a first filter; it is only misleading when its hit rate is read as its accuracy.
Where candidates lose it
Most candidates say 90% within a second, because the question hands them that number. It is the chance of a flag given a suspicious trade, and the interviewer asked the reverse.
The second loss is starting on Bayes' formula with decimals and getting tangled. Say you will use 10,000 trades, draw the two splits, and the answer reads straight off the flagged column.
What the interviewer asks next
- What false flag rate would make a flag right half the time?
- The flagged trades go through a second independent screen and are flagged again. What is the chance now?
- Compliance wants to catch 99% of suspicious trades. What does that usually do to the false flag rate?
Asked at Citadel, Sales and Trading, Miami, 2022 (Wall Street Oasis):
I got a question about Bayes' theorem applied to a practical scenario
