Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
081A six-chamber revolver holds two bullets in adjacent chambers. The cylinder is spun, the interviewer pulls the trigger on himself and it clicks on an empty chamber. It is now your turn. Do you want him to spin the cylinder again first, or pull straight away?SchonfeldCentral · 2022
Try it first
Which gives you the better chance of surviving?
Show the worked solution
Do not spin: pulling straight away survives with probability 3/4, against 2/3 with a spin. The click tells you the hammer sat on one of the four empty chambers. Because the two bullets are adjacent, the four empties form a run, and only the last empty in that run is followed by a bullet. A spin throws that information away and resets you to 4 empties out of 6.
What does the click actually tell you?
Think of a row of six houses where two neighbours keep dogs. You knocked at a random house and no dog barked. If you now try the next house along, you are only in trouble if you had knocked on the one house sitting just before the dogs. The click narrows the hammer's position to the four empty chambers, and the question becomes how many of those four have a bullet immediately after them. With the bullets side by side, the empties run 3, 4, 5, 6 in firing order, and only chamber 6 hands over to a bullet.
With two adjacent bullets, only one of the four empty chambers is followed by a bullet, so pulling straight away after a click survives 75% of the time, while a fresh spin survives only 66.7%, four empties out of six. The relationshipempties followed by an empty chambers 3, 4 and 5 in firing order 4/6 the survival chance of a fresh, random chamber What it says in wordsWithout a spin you are conditioning on the click, which helps; a spin forgets it.Does the answer depend on the bullets being adjacent?
Completely, and that is the follow-up most interviewers ask. If the two bullets are not next to each other, the empties split into two runs, two empties now sit in front of a bullet, and pulling straight away survives only 1/2, worse than the 2/3 of a spin. The same count works for bullets one apart or directly opposite: in both cases two of the four empties are followed by a bullet. So the rule is not spin or do not spin; it is count the empties that border a bullet and compare with a fresh spin.
What is the general lesson for a trading interview?
A random reset destroys information, and information has value only if the structure of the problem lets you use it. Here the structure is clustering: the bullets sit together, so a safe chamber is likely followed by another safe one. Markets have the same feature in volatility: a calm day tends to be followed by a calm day, so conditioning on what just happened beats assuming each day is a fresh draw. Say that link in one line after the arithmetic.
Where candidates lose it
The common error is to treat both options as a fresh draw and say it makes no difference, or to say a spin is safer because it resets the odds. Both ignore the click, which is the one piece of information you were given.
The second loss is getting 3/4 without seeing that it hinges on adjacency. Say that with the bullets apart the answer flips to spin, and give the count, two bordering empties out of four.
What the interviewer asks next
- The two bullets are placed in random chambers, not necessarily adjacent. Spin or not?
- Three adjacent bullets and a click. Spin or not?
- After two clicks in a row without spins, what is your survival chance on the third pull?
Asked at Schonfeld, Quantitative Research, Central, 2022 (Wall Street Oasis):
Coding, requires to know DP and divde and conquer., Russian Roulette
082Two independent waiting times are each exponentially distributed with a mean of one minute. What is the probability that their total is less than one minute?CitadelChicago · 2025
Try it first
Pick the closest value.
Show the worked solution
1 - 2/e, about 26.4%. Convolving two exponential densities gives the total the density x e^-x, which starts at zero and peaks at one minute. Its area below one minute is 1 - 2/e. The same number drops out of the Poisson view: the total is under a minute exactly when at least two arrivals land in the first minute of a rate-one Poisson process.
Why does adding two waits change the shape so much?
Suppose you need two buses, one after the other, and each arrives on average a minute after you reach its stop. Catching the first bus quickly is common; catching both quickly is rare, because both have to cooperate. A single exponential wait is most likely near zero, but a sum of two is almost never near zero, so its density starts at zero and rises into a hump. That shift of mass away from zero is why the answer is much smaller than the 63.2% chance that one wait is under a minute.
The single exponential wait puts most of its mass near zero, but the total of two waits has density x e^-x, which starts at zero and peaks at one minute, so only 26.4% of its area, shaded, falls below one minute. The relationshipS the total of the two waits x the first wait, which can be anything from 0 to s e^{-x} the exponential density with mean 1 What it says in wordsTo land on a total of s, the first wait takes any value x and the second makes up the rest; adding over all x gives s e^-s.Is there a way to get 1 - 2/e without integrating?
Yes, and it is the cleaner answer to give aloud. Exponential waits with mean one are the gaps between arrivals of a Poisson process with rate one per minute. The second arrival comes before one minute exactly when at least two arrivals land in the first minute, and the Poisson chance of zero or one arrival is e^-1 + e^-1 = 2/e. So the answer is 1 - 2/e, about 26.4%, with no calculus at all.
Sanity-check the size. Both waits being under a minute has probability (1 - 1/e)^2, about 40.0%, and the total being under a minute is a stricter event, so the answer must be smaller: 26.4% is. The limitation is the independence assumption; if the two waits were driven by the same traffic, they would move together and the total would be more spread out.
Where candidates lose it
The frequent wrong answer is (1 - 1/e)^2, about 40%, which is the chance that each wait is under a minute. The question asks about the total, and two waits of 0.7 minutes each pass that test while failing this one.
The second loss is starting a convolution integral and getting lost in the limits. Say the Poisson route first: at least two arrivals in the first minute, one minus the chance of zero or one.
What the interviewer asks next
- What is the probability that the sum of three such waits is under one minute?
- Given the total is exactly 2 minutes, what is the distribution of the first wait?
- What is the probability that the first wait is shorter than the second?
Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):
if i knew this was about convolutions, i would have answered better
083A market is either calm or stressed each day. A calm day is followed by another calm day with probability 0.8, and a stressed day is followed by another stressed day with probability 0.6. In the long run, what fraction of days are calm?DRWLondon · 2026
Try it first
What share of days are calm in the long run?
Show the worked solution
Two thirds of days are calm. In the long run the fraction of days moving from calm to stressed must equal the fraction moving back, so calm share x 0.2 = stressed share x 0.4. Calm days are therefore twice as common as stressed days, 2/3 against 1/3. The same answer comes from spell lengths: calm spells average 5 days and stressed spells 2.5, and 5 / 7.5 = 2/3.
Why can you balance flows instead of solving equations?
Think of two rooms at a party. Every few minutes, one in five people in the kitchen wanders to the lounge, and two in five people in the lounge wander back. Once the crowd settles, the numbers crossing each way must match, or one room would keep filling up. In a two-state chain the long-run shares are fixed by one equation: the share of days leaving calm must equal the share of days leaving stressed. Calm leaves at rate 0.2 and stressed at rate 0.4, so calm must hold twice as many days.
Calm days leave at rate 0.2 and stressed days at rate 0.4, so in the long run calm must hold twice as many days as stressed for the flows to balance: two thirds calm and one third stressed, with each flow equal to 2/15 of all days. The relationshippi_C the long-run share of calm days pi_S the long-run share of stressed days 1 - 0.8 the chance a calm day is followed by a stressed one What it says in wordsEach state's long-run share is the rate of leaving the other state, divided by the two leaving rates added together.How do you check it a second way?
Use spell lengths. A calm spell ends each day with probability 0.2, so it lasts 1/0.2 = 5 days on average; a stressed spell ends with probability 0.4, so it lasts 2.5 days. The chain alternates calm spell, stressed spell, calm spell, so the calm share is 5 out of every 7.5 days, which is 2/3. Two methods, one answer, in under a minute.
How fast does the chain forget where it started?
The transition matrix has a second eigenvalue of 0.8 + 0.6 - 1 = 0.4, and any gap between today's probabilities and the long-run split shrinks by that factor each day. After five days the starting state explains only 0.4^5, about 1%, of the gap, so the answer does not depend on how the week began. A trader would add the limitation: real regimes are not memoryless, and a stress spell that has already lasted a month is not as likely to end tomorrow as one that started yesterday.
Where candidates lose it
The trap is answering 80%, reading the chance of staying calm as the share of calm days. The 0.8 describes one step, and the long-run share depends on how quickly both states are left, not just one of them.
The second slip is setting up a full eigenvector calculation and running out of time. Say the flow balance in one line, then use the spell lengths as the check.
What the interviewer asks next
- Today is stressed. What is the probability that the day after tomorrow is calm?
- What is the expected number of days until the first stressed day, starting calm?
- If a desk loses Rs 2 lakh on stressed days and makes Rs 1 lakh on calm days, what is its long-run average daily P&L?
Asked at DRW, Trader Intern Interview, London, 2026 (Wall Street Oasis):
consisted of math, statistics, and probability theory (eg. one question was about markov chains
084A price path runs 100, 120, 90, 130, 80, 110, 140, 112. What are the two largest drawdowns, and how do you compute the maximum drawdown in a single pass through the data?Balyasny Asset ManagementLondon · 2025
Try it first
What is the maximum drawdown of this path?
Show the worked solution
The two largest drawdowns are 130 to 80, 38.5%, and 120 to 90, 25.0%. Walk through the prices once, keeping the highest price seen so far. At each step the drawdown is one minus price over that running peak; the maximum drawdown is the largest value seen. For the n largest, close an episode each time a new high is set, record its trough, and sort the episodes.
What exactly is a drawdown measured from?
Think of a hiker measuring how far below the highest point reached so far she has dropped. A drop only counts from a summit already climbed, not from a peak further along the trail. A drawdown is the fall from the running maximum to a later price, so the order of the prices matters and the overall high and low cannot simply be paired. Here the 140 comes after the 80, so the tempting 42.9% never happened.
Tracking the running peak shows three separate drawdowns: 130 to 80 at 38.5%, 120 to 90 at 25.0%, and 140 to 112 at 20.0%, which is still open at the end of the data. The relationshipP_t the price on day t M_t the running peak, the highest price up to day t DD_t the drawdown on day t What it says in wordsKeep the highest price so far, measure today's fall from it, and remember the worst fall.How do you get the n largest drawdowns rather than just the worst?
Split the path into episodes. An episode opens at a running peak and closes when the price makes a new high; its size is the fall from that peak to the lowest price inside it. Here the 120 episode closes when the price reaches 130, with a trough of 90, so it is 25%. The 130 episode closes at 140 with a trough of 80, 38.5%. The 140 episode never closes, so report it as open at 20.0%. Sorting the episodes gives the n largest in one pass plus a sort, O(N log N) at worst, and a heap of size n keeps it at O(N log n).
Say the edge cases, because the interviewer is testing code judgement as much as arithmetic. Two drawdowns from the same peak must not be counted twice: the fall to 80 and the later level of 110 belong to one episode. An episode still open at the end of the data is real risk and should be reported with a flag. And drawdown on a price series is not the same as drawdown on a strategy's cumulative P&L, where you would use the equity curve, not the price.
Where candidates lose it
The instinctive error pairs the overall high with the overall low: 140 and 80, 42.9%. That ignores time order; a peak must come before its trough.
The coding version of the same mistake is returning the n largest daily drawdown values, which for n = 4 would add 15.4%, the day the price sat at 110 below its 130 peak, as a separate event when it is part of the 130 to 80 fall. Group by episode first, then rank.
What the interviewer asks next
- How long did the 130 episode last from peak to recovery?
- Write the one-pass code and state its time and memory cost.
- Why is maximum drawdown a noisy statistic for comparing two strategies with short track records?
Asked at Balyasny Asset Management, Quantitative Trading, London, 2025 (Wall Street Oasis):
There was an OA with a programming and data science problem. Programming asked to return the n largest drawdowns
085What are the last two digits of 4 raised to the power 3000?Belvedere TradingNew york · 2021
Try it first
Which ending is right?
Show the worked solution
76. Split 100 into 4 x 25. Any power of 4 is 0 mod 4. Mod 25, Euler's theorem applies because 4 and 25 share no factor, and 3000 is a multiple of phi(25) = 20, so 4^3000 is 1 mod 25. The numbers below 100 that are 1 mod 25 are 1, 26, 51 and 76, and only 76 is divisible by 4.
Why does the obvious Euler shortcut fail?
Euler's theorem says a to the power phi(n) is 1 mod n, and phi(100) = 40, so it is tempting to say 4^3000 = (4^40)^75 ends in 01. The theorem needs the base and the modulus to share no factor, and 4 and 100 share a factor of 4, so it does not apply. A quick sense check kills 01 anyway: every power of 4 is divisible by 4, and a number is divisible by 4 exactly when its last two digits are, which 01 is not.
How do you split the problem so the theorem does apply?
Think of a clock with 100 hours as two smaller clocks running together, one with 4 hours and one with 25. Knowing where both small clocks point fixes the big one exactly. Mod 4 the answer is 0, since 4^3000 is a multiple of 4; mod 25 the answer is 1, since 4 and 25 share no factor and 3000 is a multiple of phi(25) = 20. Now list the numbers below 100 that are 1 mod 25: 1, 26, 51, 76. Only 76 is a multiple of 4. This step is the Chinese remainder theorem, and naming it earns credit.
The last two digits of 4^n cycle through ten values and the tenth power ends in 76, so every multiple of 10 as an exponent, including 3000, ends in 76; splitting 100 into 4 and 25 confirms it, because 76 is the only number below 100 that is 0 mod 4 and 1 mod 25. The relationshipmod 4, mod 25 the remainders on division by 4 and by 25 phi(25) = 20 how many numbers below 25 share no factor with 25 What it says in wordsFind the remainder on each small clock, then find the one number below 100 that matches both.What is the fastest check if you have a pencil?
Just list the endings. Multiply each ending by 4 and keep the last two digits: 04, 16, 64, 56, 24, 96, 84, 36, 44, 76, and then 76 x 4 = 304, which ends in 04, so the cycle has length 10. Because 76 x 76 = 5,776 also ends in 76, every power of 76 ends in 76, and 4^3000 = (4^10)^300 must end in 76. On a multiple-choice test, this listing takes about thirty seconds and needs no theorem at all.
Where candidates lose it
The trap is applying Euler's theorem with phi(100) = 40 and answering 01. The theorem requires the base and modulus to share no factor, and 4 and 100 do share one.
The second loss is listing powers without noticing the cycle and running out of time. Say early that the endings must repeat, find the period of 10, and read the answer from 3000 being a multiple of 10.
What the interviewer asks next
- What are the last two digits of 7^2026?
- What are the last three digits of 4^3000?
- What is the remainder when 2^100 is divided by 7?
Asked at Belvedere Trading, Equities, New york, 2021 (Wall Street Oasis):
It was a 14 question multiple choice test. Some basic number theory (4^3000 modulo 100)
086A price-weighted index holds three stocks priced 50, 100 and 150, with a divisor of 3. The 150 stock splits 3 for 1. What is the new divisor, and how does a market-cap-weighted index handle the same split?MizuhoHong Kong · 2024
Try it first
What must the new divisor be?
Show the worked solution
The new divisor is 2. Before the split the prices sum to 300, and 300 / 3 is 100. After a 3 for 1 split the 150 stock trades at 50, the sum is 200, and only a divisor of 2 keeps the index at 100. A cap-weighted index needs no adjustment at all, because a split triples the share count as it cuts the price to a third, leaving market value unchanged.
Why must the divisor change when nothing about the company changed?
Cut a pizza into twelve slices instead of four and you have not made more pizza. A stock split does the same to a company: three times the shares, each worth a third. A price-weighted index adds up share prices, so a split drops the sum even though no value was lost, and the divisor must be cut to stop a fake fall in the index. Solve for it by keeping the index level fixed: 200 divided by the new divisor must equal 100, so the divisor is 2.
Before the split the prices sum to 300 and the index is 300 / 3 = 100; after C splits 3 for 1 the sum is 200, so the divisor falls to 2 to hold the index at 100, and C's weight falls from 50% to 25%. The relationshipP_i the price of stock i d the divisor before the split, 3 d' the divisor after the split What it says in wordsChoose the new divisor so the index is the same the moment after the split as the moment before.What else changes in a price-weighted index after the split?
The weights. In a price-weighted index a stock's weight is its price over the sum of prices, so the expensive stock dominates whatever the size of the company. Before the split C carried 50% of the index; after it, C carries only 25% and B, untouched, jumps to 50%. A 10% rise in C used to add 5 index points; now it adds 2.5. Nothing about C's business changed; the index simply started caring less about it, which is the main criticism of price weighting.
How do the other common methods treat the split?
A market-cap-weighted index sums price times shares, and a 3 for 1 split multiplies shares by 3 while dividing price by 3, so the stock's market value, its weight and the index are all unchanged; no divisor adjustment is needed for a split. Cap-weighted divisors still change for events that alter total market value without a price move, such as share issuance, buybacks or a constituent being replaced. An equal-weighted index is also untouched by a split, since weights are reset to equal at each rebalance, but it has to trade at every rebalance to get back to equal, which costs money.
Where candidates lose it
The common slip is dividing the old divisor by the split ratio and answering 1. The divisor is fixed by keeping the index level unchanged, and only one stock split, so the adjustment is smaller than the ratio.
The second loss is saying a cap-weighted index needs the same adjustment. Market value does not change in a split, so a cap-weighted index does nothing; say that, then name the events that do change its divisor.
What the interviewer asks next
- Stock B now pays a special dividend of 20. How does each index type handle it?
- Replace stock A with a new stock priced 200. What is the new divisor?
- Which stock has the most influence on a price-weighted index, and why is that a flaw?
Asked at Mizuho, Sales and Trading, Hong Kong, 2024 (Wall Street Oasis):
Different index methodology - need to know all of them with examples.
087A new strategy has won on 9 of its first 10 trading days. With no prior knowledge, you treat its daily win rate as uniform between 0 and 1. What is the probability that it wins tomorrow, and why is the answer not 0.9?Wolverine TradingChicago · 2017
Try it first
What probability do you give to a win tomorrow?
Show the worked solution
10/12, about 0.833. A uniform prior on the win rate, updated with 9 wins and 1 loss, gives a Beta(10, 2) posterior. The chance of winning tomorrow is the posterior mean, (9 + 1)/(10 + 2). The answer is below 0.9 because ten days cannot rule out a lower true win rate, and averaging over that uncertainty pulls the estimate towards one half.
Why is 0.9 too confident?
A new restaurant with nine five-star reviews out of ten looks excellent, but you would not bet that the next diner rates it five stars with 90% certainty. Ten reviews is a small sample, and a restaurant that truly earns five stars 70% of the time could easily post nine out of ten. The right forecast averages over every win rate the evidence still allows, and with only ten days that includes plenty of rates below 0.9. The average of that spread of possibilities is what you should quote for tomorrow.
After 9 wins in 10 days from a flat prior, the win rate follows a Beta(10, 2) curve that peaks at 0.9 but has a long left tail, so its mean, the chance of winning tomorrow, is 10/12, about 0.833, and its middle 90% still spans 0.64 to 0.97. How does the Bayesian update give 10/12?
Start with every win rate p between 0 and 1 equally likely. The chance of the observed record is proportional to p^9 (1 - p), so the posterior is proportional to that, which is the Beta(10, 2) distribution. The chance of a win tomorrow is the average of p over the posterior, and the mean of a Beta(a, b) is a/(a + b), here 10/12. The shortcut is Laplace's rule of succession: add one imaginary win and one imaginary loss to the record, then divide.
The relationshipp the unknown daily win rate p^9 (1 - p) the likelihood of 9 wins and 1 loss Beta(10, 2) the posterior after a flat prior What it says in wordsMultiply the flat prior by the likelihood of the record, and the mean of the result is the chance of a win tomorrow.When does the pull towards one half stop mattering?
When the record is long. At 90 wins out of 100 the rule gives 91/102, about 0.892, almost exactly the raw 0.9, because a hundred days of data swamp the one imaginary win and loss. The limitation is the prior itself. A uniform prior says a 99% win rate was as plausible as a 50% one before you saw any data, which no trader believes about a new strategy. With a sceptical prior centred near one half, ten days would pull the estimate even further below 0.9. The posterior also tells you more than one number: the chance of winning both of the next two days is (10 x 11)/(12 x 13), about 0.705, not 0.833 squared.
Where candidates lose it
The trap is answering 0.9, the maximum likelihood estimate. It treats the observed win rate as the truth and ignores how little ten days can tell you.
The second loss is getting 10/12 by the rule of succession without being able to say where it comes from. Name the flat prior, the Beta(10, 2) posterior and its mean, and say that more data pushes the answer back towards 0.9.
What the interviewer asks next
- What is the probability the strategy's true win rate is above one half?
- With a Beta(5, 5) prior instead of a flat one, what is your forecast for tomorrow?
- How many consecutive wins would you need before your forecast exceeds 0.95?
Asked at Wolverine Trading, Prop Trading, Chicago, 2017 (Wall Street Oasis):
Phone interviews were pretty standard brainteasers and fit questions. There was a Bayes question
088A strategy's true annualised Sharpe ratio is 1.0. Roughly how many years of returns do you need before a t-test rejects a zero mean at about the 5% level? What if the Sharpe ratio is 0.5?Viking Global InvestorsNew York · 2014
Try it first
How many years does a Sharpe 0.5 strategy need?
Show the worked solution
About 4 years at a Sharpe of 1, and about 16 years at 0.5. The t-statistic for a mean return is the Sharpe ratio times the square root of the number of years, so reaching t = 2 needs (2 / SR)^2 years. Halving the Sharpe ratio quadruples the evidence you need, and sampling daily instead of yearly does not shorten it.
Why does the t-statistic come out as Sharpe times root years?
A t-test on a mean divides the average return by its standard error, which is the volatility over the square root of the number of observations. With annual observations that ratio is (mean / volatility) x sqrt(years), and mean over volatility is exactly the annual Sharpe ratio. So a Sharpe of 1 gives t = sqrt(years): 2 after 4 years. A Sharpe of 0.5 gives t = 0.5 x sqrt(years): 2 only after 16 years. With the textbook 1.96 in place of 2 the numbers are 3.8 and 15.4 years; the round figures are what you say in the room.
The t-statistic equals the Sharpe ratio times the square root of years, so a Sharpe of 1 reaches t = 2 after 4 years and a Sharpe of 0.5 only after 16: halving the Sharpe quadruples the track record you need. The relationshipr bar the average annual return in excess of cash sigma the annual volatility N the number of years observed SR the annual Sharpe ratio, r bar over sigma What it says in wordsThe evidence for a real edge grows with the square root of time, scaled by the Sharpe ratio.Can you shortcut it with daily data?
This is the follow-up that separates candidates. Sampling daily gives about 252 times as many observations a year, but the daily Sharpe ratio is smaller by the square root of 252, because daily mean scales with time and daily volatility with its square root. The two effects cancel exactly, so the t-statistic depends on calendar time, not on how finely you slice it. Think of estimating a river's average level: measuring every minute instead of every day does not help if the river's slow swings are the uncertainty.
What makes the real requirement even longer?
Three things, each worth one sentence. Returns are not independent from year to year, and positive autocorrelation inflates the true standard error. If you tested twenty strategies and kept the best, a t of 2 is easy to get by luck, so the bar has to rise with the number of ideas tried. And the Sharpe ratio itself drifts as markets change, so a sixteen-year record may be measuring two different strategies. A desk that says a Sharpe 0.5 strategy is proven after three years is reading noise.
Where candidates lose it
The first loss is scaling linearly: a Sharpe half as big needs twice as long, so 8 years. The t-statistic grows with the square root of time, so the years scale with the square of 1/Sharpe, and 16 is right.
The second loss is proposing daily data as the fix. The number of observations goes up but the per-observation Sharpe goes down by the square root of that factor, and the two cancel.
What the interviewer asks next
- How many years does a Sharpe of 2 need, and why do high-frequency desks care?
- If you tested 50 strategies, roughly what t-statistic would you demand of the best one?
- How does positive autocorrelation in monthly returns change the answer?
Asked at Viking Global Investors, Quantitative Research, New York, 2014 (Wall Street Oasis):
how to reject a hypothesis test, what's your structure of your code, what's the sample size
089Make me a market on the number of disposable nappies used in the UK in one day. Build the estimate from stated assumptions and choose a width you would actually trade on.DRWLondon · 2025
Try it first
If each of four inputs could be about 10 to 25% off in either direction, how uncertain is the product?
Show the worked solution
About 9.4 million a day, and I would open at 8 bid, 11 offered, in millions. Assume about 700,000 births a year, 2.5 years in nappies, six changes a day and 90% disposable: 9.45 million. Multiplying the low and high ends of each input gives 5.5 to 15.0 million, so a quote of 8 at 11 is tight enough to trade and still honest about the uncertainty.
How do you build the estimate so the interviewer can follow it?
Chain it through things you can reason about. Children in nappies are roughly births a year times the years each child spends in them. Assume about 700,000 births a year, a round number worth checking against the latest official statistics, and 2.5 years in nappies: about 1.75 million children. Each child uses about six a day on average, more as a newborn and fewer as a toddler, and assume 90% of families use disposables: 1.75 million x 6 x 0.9 = 9.45 million a day. Say each assumption out loud and give it a range as you go.
Multiplying the four central assumptions gives 9.45 million nappies a day, but multiplying the four lows and the four highs gives 5.5 to 15.0 million, so the honest uncertainty is roughly a factor of 1.6 either side, and a quote of 8 at 11 million sits inside it. The relationshipB births a year, an assumption Y years a child spends in nappies c changes a day d share of families using disposables What it says in wordsBuild the count from quantities you can defend one at a time, and multiply.Where should the width of the market come from?
From the ranges, multiplied. A shopkeeper who is unsure of both price and quantity is more unsure of revenue than of either. Put a low and a high on every input and multiply the lows together and the highs together: here 5.5 million to 15.0 million, about a factor of 1.6 either side of the centre. Centre the quote near the middle on a multiplicative scale, the geometric mean of the ends, 9.1 million, which sits close to the central estimate.
Then choose the width you will actually trade. A market as wide as the whole range, 5.5 at 15, is useless: nobody trades against it and it tells the interviewer you have no view. Quote tighter, 8 at 11, and move it as they trade: if they keep buying at 11, raise both sides, because their trades carry information. Say the scope questions too: does the count include adult incontinence products, and a school-age child in night-time pants? Those can move the answer more than any of the four inputs.
Where candidates lose it
The first loss is giving one number, or a market whose width is a round guess such as plus or minus a million, with no link to the assumptions. The interviewer wants to see where the width came from.
The second loss is the opposite: a market so wide it is safe and worthless. Show the full range, then quote a tighter two-way price and explain how you would move it when they trade.
What the interviewer asks next
- I buy 5 lots at your offer. Where is your market now?
- What single piece of data would you buy to narrow the range most, and why?
- How would you size the market if the settlement were a count of nappies sold rather than used?
Asked at DRW, Trading, London, 2025 (Wall Street Oasis):
Make me a market on the amount of diapers used in the UK daily
090X and Y are independent random variables with the same variance. What is the correlation between X and X + Y?Squarepoint CapitalMontreal · 2026
Try it first
Pick the correlation.
Show the worked solution
1/sqrt(2), about 0.71. The covariance of X with X + Y is Var X plus Cov(X, Y), and the second term is zero, so it is sigma squared. The variance of X + Y is 2 sigma squared, because independent variances add. Dividing sigma squared by sigma times sqrt(2) sigma leaves 1/sqrt(2). Squared, that is 0.5: X explains half of the sum's variance.
Why isn't the answer one half?
Picture two people each tossing a coin for a rupee, and a pot holding their combined winnings. One player's result explains exactly half of the pot's variability, and the other half comes from the other player. Half is the share of variance explained, R squared, and correlation is its square root, so the correlation is 1/sqrt(2), not 1/2. This is the most common slip on the question, and it comes from mixing up the two measures.
In the covariance box, X's own variance fills one of the two non-zero cells, so X accounts for half of Var(X + Y), and the correlation between X and X + Y is sigma squared divided by sigma times sqrt(2) sigma, which is 1/sqrt(2), about 0.71. The relationshipCov(X, X + Y) Var X plus Cov(X, Y), which is sigma squared plus zero sigma_{X+Y} the standard deviation of the sum, sqrt(2) sigma What it says in wordsCovariance is linear, so split it into pieces; the only surviving piece is X's own variance.How does it change if the variances differ?
Let Var Y be k times Var X. The covariance is still Var X, and Var(X + Y) becomes (1 + k) Var X. The correlation is 1/sqrt(1 + k): the noisier Y is, the less the sum tracks X. At k = 1 you get 0.707; at k = 4, 0.447; at k = 0.25, 0.894. This is exactly the signal-plus-noise model: if a price move is a true signal plus independent noise of equal size, the best-case correlation between your signal and the move is about 0.71.
Where does this show up on a desk?
Any time one piece is part of a total. A stock's return is market return plus its own specific return; if the two had equal variance, the stock would correlate 0.71 with the market. The same arithmetic tells you the ceiling on a predictor: if half of tomorrow's move is unpredictable noise, no model can correlate more than 0.71 with it. The limitation is independence; if X and Y are correlated, add 2 Cov(X, Y) to the variance of the sum and Cov(X, Y) to the covariance.
Where candidates lose it
The frequent slip is answering one half, confusing the share of variance with the correlation. Correlation is the square root of that share.
The second loss is writing the standard deviation of X + Y as 2 sigma, adding standard deviations instead of variances, which gives one half again by a different road. Independent variances add; standard deviations do not.
What the interviewer asks next
- What is the correlation between X + Y and X - Y?
- If X and Y have correlation 0.5 and equal variance, what is corr(X, X + Y)?
- What is the correlation between the first die and the total of two dice?
Asked at Squarepoint Capital, Desk Quant Analyst Interview, Montreal, 2026 (Wall Street Oasis):
There were also 3-4 basic math/stats questions about mean, covariance, correlation, etc.
