Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
081A six-chamber revolver holds two bullets in adjacent chambers. The cylinder is spun, the interviewer pulls the trigger on himself and it clicks on an empty chamber. It is now your turn. Do you want him to spin the cylinder again first, or pull straight away?SchonfeldCentral · 2022
Try it first
Which gives you the better chance of surviving?
Show the worked solution
Do not spin: pulling straight away survives with probability 3/4, against 2/3 with a spin. The click tells you the hammer sat on one of the four empty chambers. Because the two bullets are adjacent, the four empties form a run, and only the last empty in that run is followed by a bullet. A spin throws that information away and resets you to 4 empties out of 6.
What does the click actually tell you?
Think of a row of six houses where two neighbours keep dogs. You knocked at a random house and no dog barked. If you now try the next house along, you are only in trouble if you had knocked on the one house sitting just before the dogs. The click narrows the hammer's position to the four empty chambers, and the question becomes how many of those four have a bullet immediately after them. With the bullets side by side, the empties run 3, 4, 5, 6 in firing order, and only chamber 6 hands over to a bullet.
With two adjacent bullets, only one of the four empty chambers is followed by a bullet, so pulling straight away after a click survives 75% of the time, while a fresh spin survives only 66.7%, four empties out of six. The relationshipempties followed by an empty chambers 3, 4 and 5 in firing order 4/6 the survival chance of a fresh, random chamber What it says in wordsWithout a spin you are conditioning on the click, which helps; a spin forgets it.Does the answer depend on the bullets being adjacent?
Completely, and that is the follow-up most interviewers ask. If the two bullets are not next to each other, the empties split into two runs, two empties now sit in front of a bullet, and pulling straight away survives only 1/2, worse than the 2/3 of a spin. The same count works for bullets one apart or directly opposite: in both cases two of the four empties are followed by a bullet. So the rule is not spin or do not spin; it is count the empties that border a bullet and compare with a fresh spin.
What is the general lesson for a trading interview?
A random reset destroys information, and information has value only if the structure of the problem lets you use it. Here the structure is clustering: the bullets sit together, so a safe chamber is likely followed by another safe one. Markets have the same feature in volatility: a calm day tends to be followed by a calm day, so conditioning on what just happened beats assuming each day is a fresh draw. Say that link in one line after the arithmetic.
Where candidates lose it
The common error is to treat both options as a fresh draw and say it makes no difference, or to say a spin is safer because it resets the odds. Both ignore the click, which is the one piece of information you were given.
The second loss is getting 3/4 without seeing that it hinges on adjacency. Say that with the bullets apart the answer flips to spin, and give the count, two bordering empties out of four.
What the interviewer asks next
- The two bullets are placed in random chambers, not necessarily adjacent. Spin or not?
- Three adjacent bullets and a click. Spin or not?
- After two clicks in a row without spins, what is your survival chance on the third pull?
Asked at Schonfeld, Quantitative Research, Central, 2022 (Wall Street Oasis):
Coding, requires to know DP and divde and conquer., Russian Roulette
082Two independent waiting times are each exponentially distributed with a mean of one minute. What is the probability that their total is less than one minute?CitadelChicago · 2025
Try it first
Pick the closest value.
Show the worked solution
1 - 2/e, about 26.4%. Convolving two exponential densities gives the total the density x e^-x, which starts at zero and peaks at one minute. Its area below one minute is 1 - 2/e. The same number drops out of the Poisson view: the total is under a minute exactly when at least two arrivals land in the first minute of a rate-one Poisson process.
Why does adding two waits change the shape so much?
Suppose you need two buses, one after the other, and each arrives on average a minute after you reach its stop. Catching the first bus quickly is common; catching both quickly is rare, because both have to cooperate. A single exponential wait is most likely near zero, but a sum of two is almost never near zero, so its density starts at zero and rises into a hump. That shift of mass away from zero is why the answer is much smaller than the 63.2% chance that one wait is under a minute.
The single exponential wait puts most of its mass near zero, but the total of two waits has density x e^-x, which starts at zero and peaks at one minute, so only 26.4% of its area, shaded, falls below one minute. The relationshipS the total of the two waits x the first wait, which can be anything from 0 to s e^{-x} the exponential density with mean 1 What it says in wordsTo land on a total of s, the first wait takes any value x and the second makes up the rest; adding over all x gives s e^-s.Is there a way to get 1 - 2/e without integrating?
Yes, and it is the cleaner answer to give aloud. Exponential waits with mean one are the gaps between arrivals of a Poisson process with rate one per minute. The second arrival comes before one minute exactly when at least two arrivals land in the first minute, and the Poisson chance of zero or one arrival is e^-1 + e^-1 = 2/e. So the answer is 1 - 2/e, about 26.4%, with no calculus at all.
Sanity-check the size. Both waits being under a minute has probability (1 - 1/e)^2, about 40.0%, and the total being under a minute is a stricter event, so the answer must be smaller: 26.4% is. The limitation is the independence assumption; if the two waits were driven by the same traffic, they would move together and the total would be more spread out.
Where candidates lose it
The frequent wrong answer is (1 - 1/e)^2, about 40%, which is the chance that each wait is under a minute. The question asks about the total, and two waits of 0.7 minutes each pass that test while failing this one.
The second loss is starting a convolution integral and getting lost in the limits. Say the Poisson route first: at least two arrivals in the first minute, one minus the chance of zero or one.
What the interviewer asks next
- What is the probability that the sum of three such waits is under one minute?
- Given the total is exactly 2 minutes, what is the distribution of the first wait?
- What is the probability that the first wait is shorter than the second?
Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):
if i knew this was about convolutions, i would have answered better
084A price path runs 100, 120, 90, 130, 80, 110, 140, 112. What are the two largest drawdowns, and how do you compute the maximum drawdown in a single pass through the data?Balyasny Asset ManagementLondon · 2025
Try it first
What is the maximum drawdown of this path?
Show the worked solution
The two largest drawdowns are 130 to 80, 38.5%, and 120 to 90, 25.0%. Walk through the prices once, keeping the highest price seen so far. At each step the drawdown is one minus price over that running peak; the maximum drawdown is the largest value seen. For the n largest, close an episode each time a new high is set, record its trough, and sort the episodes.
What exactly is a drawdown measured from?
Think of a hiker measuring how far below the highest point reached so far she has dropped. A drop only counts from a summit already climbed, not from a peak further along the trail. A drawdown is the fall from the running maximum to a later price, so the order of the prices matters and the overall high and low cannot simply be paired. Here the 140 comes after the 80, so the tempting 42.9% never happened.
Tracking the running peak shows three separate drawdowns: 130 to 80 at 38.5%, 120 to 90 at 25.0%, and 140 to 112 at 20.0%, which is still open at the end of the data. The relationshipP_t the price on day t M_t the running peak, the highest price up to day t DD_t the drawdown on day t What it says in wordsKeep the highest price so far, measure today's fall from it, and remember the worst fall.How do you get the n largest drawdowns rather than just the worst?
Split the path into episodes. An episode opens at a running peak and closes when the price makes a new high; its size is the fall from that peak to the lowest price inside it. Here the 120 episode closes when the price reaches 130, with a trough of 90, so it is 25%. The 130 episode closes at 140 with a trough of 80, 38.5%. The 140 episode never closes, so report it as open at 20.0%. Sorting the episodes gives the n largest in one pass plus a sort, O(N log N) at worst, and a heap of size n keeps it at O(N log n).
Say the edge cases, because the interviewer is testing code judgement as much as arithmetic. Two drawdowns from the same peak must not be counted twice: the fall to 80 and the later level of 110 belong to one episode. An episode still open at the end of the data is real risk and should be reported with a flag. And drawdown on a price series is not the same as drawdown on a strategy's cumulative P&L, where you would use the equity curve, not the price.
Where candidates lose it
The instinctive error pairs the overall high with the overall low: 140 and 80, 42.9%. That ignores time order; a peak must come before its trough.
The coding version of the same mistake is returning the n largest daily drawdown values, which for n = 4 would add 15.4%, the day the price sat at 110 below its 130 peak, as a separate event when it is part of the 130 to 80 fall. Group by episode first, then rank.
What the interviewer asks next
- How long did the 130 episode last from peak to recovery?
- Write the one-pass code and state its time and memory cost.
- Why is maximum drawdown a noisy statistic for comparing two strategies with short track records?
Asked at Balyasny Asset Management, Quantitative Trading, London, 2025 (Wall Street Oasis):
There was an OA with a programming and data science problem. Programming asked to return the n largest drawdowns
088A strategy's true annualised Sharpe ratio is 1.0. Roughly how many years of returns do you need before a t-test rejects a zero mean at about the 5% level? What if the Sharpe ratio is 0.5?Viking Global InvestorsNew York · 2014
Try it first
How many years does a Sharpe 0.5 strategy need?
Show the worked solution
About 4 years at a Sharpe of 1, and about 16 years at 0.5. The t-statistic for a mean return is the Sharpe ratio times the square root of the number of years, so reaching t = 2 needs (2 / SR)^2 years. Halving the Sharpe ratio quadruples the evidence you need, and sampling daily instead of yearly does not shorten it.
Why does the t-statistic come out as Sharpe times root years?
A t-test on a mean divides the average return by its standard error, which is the volatility over the square root of the number of observations. With annual observations that ratio is (mean / volatility) x sqrt(years), and mean over volatility is exactly the annual Sharpe ratio. So a Sharpe of 1 gives t = sqrt(years): 2 after 4 years. A Sharpe of 0.5 gives t = 0.5 x sqrt(years): 2 only after 16 years. With the textbook 1.96 in place of 2 the numbers are 3.8 and 15.4 years; the round figures are what you say in the room.
The t-statistic equals the Sharpe ratio times the square root of years, so a Sharpe of 1 reaches t = 2 after 4 years and a Sharpe of 0.5 only after 16: halving the Sharpe quadruples the track record you need. The relationshipr bar the average annual return in excess of cash sigma the annual volatility N the number of years observed SR the annual Sharpe ratio, r bar over sigma What it says in wordsThe evidence for a real edge grows with the square root of time, scaled by the Sharpe ratio.Can you shortcut it with daily data?
This is the follow-up that separates candidates. Sampling daily gives about 252 times as many observations a year, but the daily Sharpe ratio is smaller by the square root of 252, because daily mean scales with time and daily volatility with its square root. The two effects cancel exactly, so the t-statistic depends on calendar time, not on how finely you slice it. Think of estimating a river's average level: measuring every minute instead of every day does not help if the river's slow swings are the uncertainty.
What makes the real requirement even longer?
Three things, each worth one sentence. Returns are not independent from year to year, and positive autocorrelation inflates the true standard error. If you tested twenty strategies and kept the best, a t of 2 is easy to get by luck, so the bar has to rise with the number of ideas tried. And the Sharpe ratio itself drifts as markets change, so a sixteen-year record may be measuring two different strategies. A desk that says a Sharpe 0.5 strategy is proven after three years is reading noise.
Where candidates lose it
The first loss is scaling linearly: a Sharpe half as big needs twice as long, so 8 years. The t-statistic grows with the square root of time, so the years scale with the square of 1/Sharpe, and 16 is right.
The second loss is proposing daily data as the fix. The number of observations goes up but the per-observation Sharpe goes down by the square root of that factor, and the two cancel.
What the interviewer asks next
- How many years does a Sharpe of 2 need, and why do high-frequency desks care?
- If you tested 50 strategies, roughly what t-statistic would you demand of the best one?
- How does positive autocorrelation in monthly returns change the answer?
Asked at Viking Global Investors, Quantitative Research, New York, 2014 (Wall Street Oasis):
how to reject a hypothesis test, what's your structure of your code, what's the sample size
089Make me a market on the number of disposable nappies used in the UK in one day. Build the estimate from stated assumptions and choose a width you would actually trade on.DRWLondon · 2025
Try it first
If each of four inputs could be about 10 to 25% off in either direction, how uncertain is the product?
Show the worked solution
About 9.4 million a day, and I would open at 8 bid, 11 offered, in millions. Assume about 700,000 births a year, 2.5 years in nappies, six changes a day and 90% disposable: 9.45 million. Multiplying the low and high ends of each input gives 5.5 to 15.0 million, so a quote of 8 at 11 is tight enough to trade and still honest about the uncertainty.
How do you build the estimate so the interviewer can follow it?
Chain it through things you can reason about. Children in nappies are roughly births a year times the years each child spends in them. Assume about 700,000 births a year, a round number worth checking against the latest official statistics, and 2.5 years in nappies: about 1.75 million children. Each child uses about six a day on average, more as a newborn and fewer as a toddler, and assume 90% of families use disposables: 1.75 million x 6 x 0.9 = 9.45 million a day. Say each assumption out loud and give it a range as you go.
Multiplying the four central assumptions gives 9.45 million nappies a day, but multiplying the four lows and the four highs gives 5.5 to 15.0 million, so the honest uncertainty is roughly a factor of 1.6 either side, and a quote of 8 at 11 million sits inside it. The relationshipB births a year, an assumption Y years a child spends in nappies c changes a day d share of families using disposables What it says in wordsBuild the count from quantities you can defend one at a time, and multiply.Where should the width of the market come from?
From the ranges, multiplied. A shopkeeper who is unsure of both price and quantity is more unsure of revenue than of either. Put a low and a high on every input and multiply the lows together and the highs together: here 5.5 million to 15.0 million, about a factor of 1.6 either side of the centre. Centre the quote near the middle on a multiplicative scale, the geometric mean of the ends, 9.1 million, which sits close to the central estimate.
Then choose the width you will actually trade. A market as wide as the whole range, 5.5 at 15, is useless: nobody trades against it and it tells the interviewer you have no view. Quote tighter, 8 at 11, and move it as they trade: if they keep buying at 11, raise both sides, because their trades carry information. Say the scope questions too: does the count include adult incontinence products, and a school-age child in night-time pants? Those can move the answer more than any of the four inputs.
Where candidates lose it
The first loss is giving one number, or a market whose width is a round guess such as plus or minus a million, with no link to the assumptions. The interviewer wants to see where the width came from.
The second loss is the opposite: a market so wide it is safe and worthless. Show the full range, then quote a tighter two-way price and explain how you would move it when they trade.
What the interviewer asks next
- I buy 5 lots at your offer. Where is your market now?
- What single piece of data would you buy to narrow the range most, and why?
- How would you size the market if the settlement were a count of nappies sold rather than used?
Asked at DRW, Trading, London, 2025 (Wall Street Oasis):
Make me a market on the amount of diapers used in the UK daily
093Three dice: red has faces 2, 6 and 7; green has 1, 5 and 12; blue has 3, 4 and 8, each face appearing twice. You and I each pick a die and roll once, and the higher number wins. Which die do you want, and does it matter who picks first?Belvedere TradingChicago · 2022
Try it first
Which die is best against the other two?
Show the worked solution
No die is best: red beats green, green beats blue and blue beats red, each with probability 5/9. So who picks first matters a great deal. Let me choose, then take the die that beats mine and win 5/9 of the time. If you are forced to pick first, every choice loses 5/9 of the time against an opponent who knows the cycle.
How do you work out who beats whom?
Each matchup has only nine equally likely pairs of faces, so write the 3 by 3 grid and count. Red against green: red's 2 beats only the 1, while its 6 and 7 each beat the 1 and the 5, for 1 + 2 + 2 = 5 wins out of 9. Do the same for the other two pairs and every matchup comes out 5 to 4: red over green, green over blue, blue over red. It is rock, paper, scissors built out of dice, and in rock, paper, scissors nobody asks which hand shape is best.
Counting the nine face pairs in each matchup shows red beats green, green beats blue and blue beats red, each in 5 of 9 cases, so the three dice form a cycle and the second player can always pick a die that wins 5/9 of the time. The relationshipP(R > G) the chance red's roll beats green's 1 + 2 + 2 the wins for red's faces 2, 6 and 7 in turn What it says in wordsCount, face by face, how many of the opponent's three faces each face beats, and divide by nine.Why does green lose to red when green has the higher average?
The averages are red 5, green 6 and blue 5. Winning is about how often, not by how much. Green's 12 wins every time it shows, but it shows only a third of the time, and green's other two faces, 1 and 5, lose to both of red's high faces. A higher mean and a higher chance of winning are different things, and the gap between them is the whole puzzle. Change the rules so the winner collects the difference between the two numbers, and green's expected margin against red is 6 - 5 = +1: now you want green against red, and blue against red is a dead heat at 0.
Where does a trader meet the same thing?
Head-to-head comparisons need not line up into a ranking. Strategy A can beat strategy B on more days than not, B can beat C, and C can beat A, whenever one of them earns its money in rare large wins, as green does. Before choosing between strategies, decide whether you care about how often you win or how much you make, because under the first a cycle like this one means there may be no best choice at all. The limitation of the puzzle is that it is one roll; over many rolls with the total score counted, the mean matters more and green's 12 starts to pay.
Where candidates lose it
The common slip is choosing green because its average, 6, is highest. The game pays for winning, not for margin, and green loses to red five times in nine.
The second loss is answering the first question and missing the second. Because the dice form a cycle, the real answer is strategic: insist that your opponent picks first. Saying that unprompted is what the interviewer is listening for.
What the interviewer asks next
- If each player rolls their die twice and adds the results, does the cycle still hold?
- Design three dice whose faces sum to the same total and still form a cycle.
- With three players each taking one die, can any die be favoured against both others?
Asked at Belvedere Trading, Prop Trading, Chicago, 2022 (Wall Street Oasis):
You have 3 dice: red has 2, 6, 7; green has 1, 5, 12; blue has 3, 4, 8.
096A desk's daily P&L in Rs lakh over seven days is -1, 2, 4, -9, 8, -2, 3. Which run of consecutive days has the largest total, and how do you find it in one pass through the data?Wolverine TradingChicago · 2014
Try it first
Which run has the largest total?
Show the worked solution
Days 5 to 7, the run 8, -2, 3, which totals Rs 9 lakh. Walk through the days keeping the best total of a run ending today: either today alone or today added to yesterday's best run, whichever is larger. Record the largest value you see. The run 2, 4 looks attractive but totals only 6, and the -9 day makes it pointless to carry anything before it.
How do you find the best run without checking every start and end day?
Picture walking along a road with toll booths that either pay you or charge you. You may choose where to start and stop collecting. If the purse you have carried from earlier booths is in the red, the sensible move is to drop it and start fresh at the next booth. A run ending today is worth extending from yesterday only if the best run ending yesterday is positive; if it is negative it can only drag today down, so today starts a new run. That rule looks at each day once. Checking every pair of start and end days means 28 runs for seven days and 31,375 for a trading year of 250 days.
Carrying the best run forward only while it is positive, the running total drops to -3 after the loss of 9 and restarts at 8 on day 5, so the best run is 8, -2, 3 with a total of 9, ahead of the tempting run 2, 4 at 6. The relationshipx_t the P&L on day t c_t the best total of a run that ends on day t best the largest c_t seen so far What it says in wordsThe best run ending today either starts today or extends the best run ending yesterday; keep whichever is bigger, and remember the biggest.Day P&L Best run ending today Best so far 1 -1 -1 -1 2 2 2 2 3 4 6 6 4 -9 -3 6 5 8 8 8 6 -2 6 8 7 3 9 9 Running the rule day by day, the best run ending today drops to -3 after day 4, restarts at 8 on day 5 and reaches 9 on day 7, which is the answer. Why does the tempting run 2, 4 lose?
Because one later day beats it on its own. After the -9, the best run ending on day 4 is 6 - 9 = -3, so the rule drops the past and day 5 starts fresh at 8. The -2 on day 6 dips the run to 6, but the 3 on day 7 lifts it to 9. The best run can contain a losing day: 8, -2, 3 beats 8 alone because the day after the loss more than repays it. A candidate who stops a run at the first red day misses this, and the brute-force check over all 28 runs confirms 9 is the maximum.
What edge cases does the interviewer probe?
Three. If every day is a loss, the answer should be the least bad single day, so start the best at the first day's value, not at zero, or you will report an empty run worth 0. To report which days, store the start index whenever you restart and copy it when you record a new best. The same pass with the signs flipped finds the worst run, here -9, the single -9 day. The limitation on a desk is that the best run in hindsight is a selected statistic: a strategy that is judged by its best stretch will always look better than it trades.
Where candidates lose it
The quick wrong answer is the run 2, 4, because it is the first good stretch. The 8 on day 5 beats it alone, and carrying 8 through -2 and 3 beats 8.
The second loss is in the code: starting the best total at zero, which reports 0 for a week of all losses, or restarting at every losing day instead of only when the running total itself turns negative. State the rule exactly: carry yesterday's run only while it is positive.
What the interviewer asks next
- Return the start and end days of the best run, not just its total.
- What does your code return if every day in the series is a loss?
- Find the best run if you may skip at most one day inside it.
Asked at Wolverine Trading, Quantitative Research, Chicago, 2014 (Wall Street Oasis):
Develop an algorithm to find out the section that contains the maximum sum.
097How many integers from 1 to 1,000 share no common factor with 1,000 other than 1?Quant researchQuant trading
Try it first
Pick the count.
Show the worked solution
400. Since 1,000 = 2^3 x 5^3, a number shares a factor with 1,000 exactly when it is divisible by 2 or by 5. There are 500 multiples of 2 and 200 of 5, but the 100 multiples of 10 sit in both lists, so 600 numbers share a factor and 400 do not. Euler's formula agrees: 1,000 x 1/2 x 4/5 = 400.
Which numbers share a factor with 1,000?
Picture a hall of 1,000 people where everyone wearing a red badge or a blue badge is asked to leave. To count who stays, you need the red-badge count, the blue-badge count, and how many wear both, because they would otherwise be counted out twice. Write 1,000 as 2^3 x 5^3: a number shares a factor with it exactly when it is divisible by 2 or by 5, so only two badges matter, and the powers 3 do not add any new conditions. Every multiple of 4 or 8 is already a multiple of 2, and every multiple of 25 or 125 is already a multiple of 5.
Of the numbers 1 to 1,000, 500 are multiples of 2 and 200 are multiples of 5, with 100 multiples of 10 in both, so 600 share a factor with 1,000 and 400 lie outside both circles; Euler's product 1,000 x 1/2 x 4/5 gives the same 400. The relationshipphi(1000) Euler's totient: how many of 1 to 1,000 share no factor with 1,000 1000/2, 1000/5 the counts of multiples of 2 and of 5 1000/10 the multiples of both, added back once What it says in wordsRemove the multiples of each prime, add back the multiples of both, and you get the same answer as multiplying by the share that survives each prime.Why does the quick product formula work here?
Half of all numbers are odd, and among those, four in five are not multiples of 5. Because 1,000 is a multiple of 10, the numbers 1 to 1,000 contain exactly 100 full blocks of ten, and in each block exactly 4 numbers, 1, 3, 7 and 9, survive both tests, so 100 x 4 = 400. The strip of 1 to 20 in the figure shows the pattern repeating, 8 survivors in 20. The product is exact only when the range is a whole number of such blocks: for 1 to 1,234 it gives 493.6, while a direct count gives 494.
Where does a question like this lead in an interview?
Usually to powers and remainders. Euler's theorem says a number coprime to n, raised to the power phi(n), leaves remainder 1 when divided by n, and that is the engine behind last-digit puzzles: phi(100) = 40, so 3^40 ends in 01 and so does 3^400. It also leads to probability: the chance that two large random integers share no factor tends to 6/pi^2, about 0.608. The habit the question tests is factorising first: once you see only the primes 2 and 5 matter, a counting question becomes a two-circle Venn diagram.
Where candidates lose it
The usual slip is 1,000 - 500 - 200 = 300, subtracting both lists and forgetting that the multiples of 10 were removed twice. Add them back once and the answer is 400.
The second loss is treating each prime power as a new condition, subtracting multiples of 4, 8, 25 and 125 as well. Every multiple of 4 is already a multiple of 2; only the distinct primes matter.
What the interviewer asks next
- How many integers from 1 to 1,000 share no factor with 360?
- What are the last two digits of 3^400?
- What is the probability that two randomly chosen integers share no common factor?
098Two traders' monthly P&L are independent and normal. A has mean Rs 10 lakh and standard deviation Rs 3 lakh; B has mean Rs 8 lakh and standard deviation Rs 4 lakh. What is the probability that A out-earns B in a given month?DRWLondon · 2025
Try it first
Pick the probability that A earns more than B in a month.
Show the worked solution
About 65.5%. The gap A - B is normal with mean 10 - 8 = Rs 2 lakh and variance 3^2 + 4^2 = 25, so its standard deviation is Rs 5 lakh. A out-earns B when the gap is positive, and zero sits 2/5 = 0.4 standard deviations below the mean, so the probability is Phi(0.4), about 65.5%. The better trader loses about one month in three.
Why do the variances add when you subtract?
You and a colleague set off for the same meeting from different places, and each journey is uncertain by a few minutes. The gap between your two arrival times is more uncertain than either journey, not less, because either of you can be the late one. Subtracting an independent random amount adds its noise, so Var(A - B) = Var A + Var B = 9 + 16 = 25, and the gap's standard deviation is 5, not 1. The mean subtracts as you would expect, 10 - 8 = 2. The gap is normal because a difference of independent normals is normal.
The two traders' monthly P&L overlap heavily, and the gap A - B has mean 2 and standard deviation 5, so the area above zero where A wins is only 65.5%, leaving B ahead in 34.5% of months. The relationshipmu_A, mu_B the mean monthly P&L, 10 and 8 sigma_A, sigma_B the standard deviations, 3 and 4 Phi the standard normal cumulative distribution What it says in wordsThe gap's mean is the difference of the means, its variance the sum of the variances, and the answer is how many standard deviations zero sits below that mean.How much does a longer comparison window help?
A lot, and at a predictable rate. Over a quarter of independent months the total gap has mean 6 and standard deviation 5 x sqrt(3), about 8.7, so A comes out ahead with probability 75.6%. Over a year the mean is 24 and the standard deviation 5 x sqrt(12), about 17.3, so the probability is 91.7%. The edge grows with the number of months and the noise with its square root, so the z-score grows with the square root of time. A risk manager who ranks traders on one month of P&L is ranking mostly noise.
What if the two traders' P&L are correlated?
Then the shared part cancels in the gap. With correlation 0.5, the variance is 9 + 16 - 2 x 0.5 x 3 x 4 = 13, a standard deviation of 3.61, and A wins with probability 71.0%. Positive correlation makes the comparison sharper because common market moves drop out of the difference; negative correlation does the opposite. The limitation is the normal assumption: real P&L has fat tails and skew, and if one trader earns through rare large months, the month-by-month win rate can disagree with the mean, so check the shape before trusting the 65.5%.
Where candidates lose it
The commonest slip is subtracting the standard deviations, 4 - 3 = 1, which makes A look almost certain to win at 97.7%. Noise does not cancel when you subtract independent variables; it adds.
The second loss is subtracting the variances, 16 - 9, or adding the standard deviations, 3 + 4. Square, add, then take the root: sqrt(9 + 16) = 5. The answer is then a z-score of 0.4, and Phi(0.4) is about 0.655.
What the interviewer asks next
- What is the probability that A out-earns B over a full year of independent months?
- If their monthly P&L has correlation 0.5, what is the answer?
- What is the probability that A out-earns B by more than Rs 5 lakh in a month?
Asked at DRW, Trading, London, 2025 (Wall Street Oasis):
technical interview based on normal distribution and market making
100You regress a centred target y on one standardised feature x with no intercept. The sum of x squared is 100 and the sum of x times y is 80. What is the OLS slope, and what is the ridge slope with penalty lambda = 25?Citadel SecuritiesLondon · 2026
Try it first
Pick the pair.
Show the worked solution
OLS gives 0.8 and ridge gives 0.64. OLS minimises squared error and its slope is the sum of xy over the sum of x squared, 80/100. Ridge adds lambda times the slope squared to the loss, which puts lambda into the denominator: 80/(100 + 25) = 0.64. That is the OLS slope times 100/125 = 0.8, so ridge shrinks the slope towards zero but never to zero.
Where does lambda end up in the formula?
Think of a new analyst's forecast that you half trust: you do not discard it, you shade it towards zero, and the less data behind it the more you shade. Ridge does that mechanically. It minimises the squared errors plus lambda times the slope squared; setting the derivative to zero gives b = Sxy/(Sxx + lambda). Penalising the size of the slope acts exactly like adding observations whose x squared totals lambda and whose y is zero, data that say the slope is zero. With 100 of real evidence and 25 of make-believe evidence, the slope is 80/125 = 0.64.
The ridge slope 80/(100 + lambda) falls from the OLS value 0.8 to 0.64 at lambda 25 and 0.40 at lambda 100 without ever reaching zero, while the lasso slope falls in a straight line and hits exactly zero at lambda 160. The relationshipsum x_i y_i the cross-product of feature and target, 80 sum x_i^2 the sum of squares of the feature, 100 lambda the ridge penalty, 25 What it says in wordsRidge is OLS with lambda added to the sum of squares, so every slope is multiplied by Sxx/(Sxx + lambda).Why would you want a slope that is biased towards zero?
Because a smaller, steadier estimate can be closer to the truth on average. Suppose the true slope is 0.5 and the noise variance is 25. OLS is unbiased but its variance is 25/100 = 0.25. Ridge at lambda 25 has variance 0.16 and a bias of -0.1, so its mean squared error is 0.17. Ridge trades a little bias for a larger cut in variance, and when the signal is weak relative to the noise that trade wins. In this one-feature case the best lambda is noise variance over slope squared, 100, which halves the slope and cuts the error to 0.125. In practice the truth is unknown, so lambda is chosen by cross-validation.
lambda Slope on this data Variance Bias squared Mean squared error 0 0.80 0.2500 0.0000 0.2500 25 0.64 0.1600 0.0100 0.1700 100 0.40 0.0625 0.0625 0.1250 Assuming a true slope of 0.5 and noise variance 25, ridge at lambda 25 and 100 has a lower mean squared error than OLS because the drop in variance outweighs the bias it adds. How does lasso differ?
Lasso penalises lambda times the absolute slope instead. In one dimension that subtracts lambda/2 from the cross-product rather than adding to the denominator: (80 - 12.5)/100 = 0.675 at lambda 25, and exactly zero once lambda reaches 160. Ridge scales coefficients down; lasso shifts them down and can set them to exactly zero, which is why lasso selects features and ridge does not. Both penalties depend on the scale of x, which is why the feature must be standardised first; with correlated features, ridge spreads the weight across them while lasso tends to keep one.
Where candidates lose it
The fast wrong answer subtracts the penalty from the slope or from the numerator, which is lasso's mechanics, not ridge's. Ridge adds lambda to the sum of squares in the denominator, so the slope is scaled, not shifted.
The second loss is saying ridge is always better because it has lower variance. It trades variance for bias; if the true slope is large and the data plentiful, shrinking costs more in bias than it saves. Say that lambda is chosen by cross-validation, not by taste.
What the interviewer asks next
- What value of lambda halves the OLS slope?
- With two highly correlated features, how do ridge and lasso split the weight between them?
- Why must features be standardised before applying a ridge penalty?
Asked at Citadel Securities, Quantitative Research, London, 2026 (Wall Street Oasis):
very detailed and difficult questions about regularisation ridge and lasso
