Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
087A new strategy has won on 9 of its first 10 trading days. With no prior knowledge, you treat its daily win rate as uniform between 0 and 1. What is the probability that it wins tomorrow, and why is the answer not 0.9?Wolverine TradingChicago · 2017
Try it first
What probability do you give to a win tomorrow?
Show the worked solution
10/12, about 0.833. A uniform prior on the win rate, updated with 9 wins and 1 loss, gives a Beta(10, 2) posterior. The chance of winning tomorrow is the posterior mean, (9 + 1)/(10 + 2). The answer is below 0.9 because ten days cannot rule out a lower true win rate, and averaging over that uncertainty pulls the estimate towards one half.
Why is 0.9 too confident?
A new restaurant with nine five-star reviews out of ten looks excellent, but you would not bet that the next diner rates it five stars with 90% certainty. Ten reviews is a small sample, and a restaurant that truly earns five stars 70% of the time could easily post nine out of ten. The right forecast averages over every win rate the evidence still allows, and with only ten days that includes plenty of rates below 0.9. The average of that spread of possibilities is what you should quote for tomorrow.
After 9 wins in 10 days from a flat prior, the win rate follows a Beta(10, 2) curve that peaks at 0.9 but has a long left tail, so its mean, the chance of winning tomorrow, is 10/12, about 0.833, and its middle 90% still spans 0.64 to 0.97. How does the Bayesian update give 10/12?
Start with every win rate p between 0 and 1 equally likely. The chance of the observed record is proportional to p^9 (1 - p), so the posterior is proportional to that, which is the Beta(10, 2) distribution. The chance of a win tomorrow is the average of p over the posterior, and the mean of a Beta(a, b) is a/(a + b), here 10/12. The shortcut is Laplace's rule of succession: add one imaginary win and one imaginary loss to the record, then divide.
The relationshipp the unknown daily win rate p^9 (1 - p) the likelihood of 9 wins and 1 loss Beta(10, 2) the posterior after a flat prior What it says in wordsMultiply the flat prior by the likelihood of the record, and the mean of the result is the chance of a win tomorrow.When does the pull towards one half stop mattering?
When the record is long. At 90 wins out of 100 the rule gives 91/102, about 0.892, almost exactly the raw 0.9, because a hundred days of data swamp the one imaginary win and loss. The limitation is the prior itself. A uniform prior says a 99% win rate was as plausible as a 50% one before you saw any data, which no trader believes about a new strategy. With a sceptical prior centred near one half, ten days would pull the estimate even further below 0.9. The posterior also tells you more than one number: the chance of winning both of the next two days is (10 x 11)/(12 x 13), about 0.705, not 0.833 squared.
Where candidates lose it
The trap is answering 0.9, the maximum likelihood estimate. It treats the observed win rate as the truth and ignores how little ten days can tell you.
The second loss is getting 10/12 by the rule of succession without being able to say where it comes from. Name the flat prior, the Beta(10, 2) posterior and its mean, and say that more data pushes the answer back towards 0.9.
What the interviewer asks next
- What is the probability the strategy's true win rate is above one half?
- With a Beta(5, 5) prior instead of a flat one, what is your forecast for tomorrow?
- How many consecutive wins would you need before your forecast exceeds 0.95?
Asked at Wolverine Trading, Prop Trading, Chicago, 2017 (Wall Street Oasis):
Phone interviews were pretty standard brainteasers and fit questions. There was a Bayes question
088A strategy's true annualised Sharpe ratio is 1.0. Roughly how many years of returns do you need before a t-test rejects a zero mean at about the 5% level? What if the Sharpe ratio is 0.5?Viking Global InvestorsNew York · 2014
Try it first
How many years does a Sharpe 0.5 strategy need?
Show the worked solution
About 4 years at a Sharpe of 1, and about 16 years at 0.5. The t-statistic for a mean return is the Sharpe ratio times the square root of the number of years, so reaching t = 2 needs (2 / SR)^2 years. Halving the Sharpe ratio quadruples the evidence you need, and sampling daily instead of yearly does not shorten it.
Why does the t-statistic come out as Sharpe times root years?
A t-test on a mean divides the average return by its standard error, which is the volatility over the square root of the number of observations. With annual observations that ratio is (mean / volatility) x sqrt(years), and mean over volatility is exactly the annual Sharpe ratio. So a Sharpe of 1 gives t = sqrt(years): 2 after 4 years. A Sharpe of 0.5 gives t = 0.5 x sqrt(years): 2 only after 16 years. With the textbook 1.96 in place of 2 the numbers are 3.8 and 15.4 years; the round figures are what you say in the room.
The t-statistic equals the Sharpe ratio times the square root of years, so a Sharpe of 1 reaches t = 2 after 4 years and a Sharpe of 0.5 only after 16: halving the Sharpe quadruples the track record you need. The relationshipr bar the average annual return in excess of cash sigma the annual volatility N the number of years observed SR the annual Sharpe ratio, r bar over sigma What it says in wordsThe evidence for a real edge grows with the square root of time, scaled by the Sharpe ratio.Can you shortcut it with daily data?
This is the follow-up that separates candidates. Sampling daily gives about 252 times as many observations a year, but the daily Sharpe ratio is smaller by the square root of 252, because daily mean scales with time and daily volatility with its square root. The two effects cancel exactly, so the t-statistic depends on calendar time, not on how finely you slice it. Think of estimating a river's average level: measuring every minute instead of every day does not help if the river's slow swings are the uncertainty.
What makes the real requirement even longer?
Three things, each worth one sentence. Returns are not independent from year to year, and positive autocorrelation inflates the true standard error. If you tested twenty strategies and kept the best, a t of 2 is easy to get by luck, so the bar has to rise with the number of ideas tried. And the Sharpe ratio itself drifts as markets change, so a sixteen-year record may be measuring two different strategies. A desk that says a Sharpe 0.5 strategy is proven after three years is reading noise.
Where candidates lose it
The first loss is scaling linearly: a Sharpe half as big needs twice as long, so 8 years. The t-statistic grows with the square root of time, so the years scale with the square of 1/Sharpe, and 16 is right.
The second loss is proposing daily data as the fix. The number of observations goes up but the per-observation Sharpe goes down by the square root of that factor, and the two cancel.
What the interviewer asks next
- How many years does a Sharpe of 2 need, and why do high-frequency desks care?
- If you tested 50 strategies, roughly what t-statistic would you demand of the best one?
- How does positive autocorrelation in monthly returns change the answer?
Asked at Viking Global Investors, Quantitative Research, New York, 2014 (Wall Street Oasis):
how to reject a hypothesis test, what's your structure of your code, what's the sample size
089Make me a market on the number of disposable nappies used in the UK in one day. Build the estimate from stated assumptions and choose a width you would actually trade on.DRWLondon · 2025
Try it first
If each of four inputs could be about 10 to 25% off in either direction, how uncertain is the product?
Show the worked solution
About 9.4 million a day, and I would open at 8 bid, 11 offered, in millions. Assume about 700,000 births a year, 2.5 years in nappies, six changes a day and 90% disposable: 9.45 million. Multiplying the low and high ends of each input gives 5.5 to 15.0 million, so a quote of 8 at 11 is tight enough to trade and still honest about the uncertainty.
How do you build the estimate so the interviewer can follow it?
Chain it through things you can reason about. Children in nappies are roughly births a year times the years each child spends in them. Assume about 700,000 births a year, a round number worth checking against the latest official statistics, and 2.5 years in nappies: about 1.75 million children. Each child uses about six a day on average, more as a newborn and fewer as a toddler, and assume 90% of families use disposables: 1.75 million x 6 x 0.9 = 9.45 million a day. Say each assumption out loud and give it a range as you go.
Multiplying the four central assumptions gives 9.45 million nappies a day, but multiplying the four lows and the four highs gives 5.5 to 15.0 million, so the honest uncertainty is roughly a factor of 1.6 either side, and a quote of 8 at 11 million sits inside it. The relationshipB births a year, an assumption Y years a child spends in nappies c changes a day d share of families using disposables What it says in wordsBuild the count from quantities you can defend one at a time, and multiply.Where should the width of the market come from?
From the ranges, multiplied. A shopkeeper who is unsure of both price and quantity is more unsure of revenue than of either. Put a low and a high on every input and multiply the lows together and the highs together: here 5.5 million to 15.0 million, about a factor of 1.6 either side of the centre. Centre the quote near the middle on a multiplicative scale, the geometric mean of the ends, 9.1 million, which sits close to the central estimate.
Then choose the width you will actually trade. A market as wide as the whole range, 5.5 at 15, is useless: nobody trades against it and it tells the interviewer you have no view. Quote tighter, 8 at 11, and move it as they trade: if they keep buying at 11, raise both sides, because their trades carry information. Say the scope questions too: does the count include adult incontinence products, and a school-age child in night-time pants? Those can move the answer more than any of the four inputs.
Where candidates lose it
The first loss is giving one number, or a market whose width is a round guess such as plus or minus a million, with no link to the assumptions. The interviewer wants to see where the width came from.
The second loss is the opposite: a market so wide it is safe and worthless. Show the full range, then quote a tighter two-way price and explain how you would move it when they trade.
What the interviewer asks next
- I buy 5 lots at your offer. Where is your market now?
- What single piece of data would you buy to narrow the range most, and why?
- How would you size the market if the settlement were a count of nappies sold rather than used?
Asked at DRW, Trading, London, 2025 (Wall Street Oasis):
Make me a market on the amount of diapers used in the UK daily
090X and Y are independent random variables with the same variance. What is the correlation between X and X + Y?Squarepoint CapitalMontreal · 2026
Try it first
Pick the correlation.
Show the worked solution
1/sqrt(2), about 0.71. The covariance of X with X + Y is Var X plus Cov(X, Y), and the second term is zero, so it is sigma squared. The variance of X + Y is 2 sigma squared, because independent variances add. Dividing sigma squared by sigma times sqrt(2) sigma leaves 1/sqrt(2). Squared, that is 0.5: X explains half of the sum's variance.
Why isn't the answer one half?
Picture two people each tossing a coin for a rupee, and a pot holding their combined winnings. One player's result explains exactly half of the pot's variability, and the other half comes from the other player. Half is the share of variance explained, R squared, and correlation is its square root, so the correlation is 1/sqrt(2), not 1/2. This is the most common slip on the question, and it comes from mixing up the two measures.
In the covariance box, X's own variance fills one of the two non-zero cells, so X accounts for half of Var(X + Y), and the correlation between X and X + Y is sigma squared divided by sigma times sqrt(2) sigma, which is 1/sqrt(2), about 0.71. The relationshipCov(X, X + Y) Var X plus Cov(X, Y), which is sigma squared plus zero sigma_{X+Y} the standard deviation of the sum, sqrt(2) sigma What it says in wordsCovariance is linear, so split it into pieces; the only surviving piece is X's own variance.How does it change if the variances differ?
Let Var Y be k times Var X. The covariance is still Var X, and Var(X + Y) becomes (1 + k) Var X. The correlation is 1/sqrt(1 + k): the noisier Y is, the less the sum tracks X. At k = 1 you get 0.707; at k = 4, 0.447; at k = 0.25, 0.894. This is exactly the signal-plus-noise model: if a price move is a true signal plus independent noise of equal size, the best-case correlation between your signal and the move is about 0.71.
Where does this show up on a desk?
Any time one piece is part of a total. A stock's return is market return plus its own specific return; if the two had equal variance, the stock would correlate 0.71 with the market. The same arithmetic tells you the ceiling on a predictor: if half of tomorrow's move is unpredictable noise, no model can correlate more than 0.71 with it. The limitation is independence; if X and Y are correlated, add 2 Cov(X, Y) to the variance of the sum and Cov(X, Y) to the covariance.
Where candidates lose it
The frequent slip is answering one half, confusing the share of variance with the correlation. Correlation is the square root of that share.
The second loss is writing the standard deviation of X + Y as 2 sigma, adding standard deviations instead of variances, which gives one half again by a different road. Independent variances add; standard deviations do not.
What the interviewer asks next
- What is the correlation between X + Y and X - Y?
- If X and Y have correlation 0.5 and equal variance, what is corr(X, X + Y)?
- What is the correlation between the first die and the total of two dice?
Asked at Squarepoint Capital, Desk Quant Analyst Interview, Montreal, 2026 (Wall Street Oasis):
There were also 3-4 basic math/stats questions about mean, covariance, correlation, etc.
092You climb a staircase of 10 steps, taking either one step or two steps at a time. In how many different ways can you reach the top?Tower Research CapitalNew York · 2012
Try it first
Pick the number of ways.
Show the worked solution
89 ways. Split by the last move: you arrive at step 10 with a single from step 9 or a double from step 8, so ways(10) = ways(9) + ways(8). With 1 way to reach step 1 and 2 ways to reach step 2, the counts run 1, 2, 3, 5, 8, 13, 21, 34, 55, 89: the Fibonacci numbers, with 89 at the top.
How do you count without listing every route?
Imagine a friend at the top of the stairs asks how you got there. There is one thing you can say for certain about your last move: it was either a single from step 9 or a double from step 8, and never both. Every route to step 10 is a route to step 9 followed by a single, or a route to step 8 followed by a double, so the count at step 10 is the sum of the counts at steps 9 and 8. The same holds at every step, which turns the puzzle into a running sum.
Writing the number of ways on each step, each step is the sum of the two below it, so the counts follow the Fibonacci sequence and step 10 collects 55 routes ending with a single and 34 ending with a double, 89 in all. The relationshipw(n) the number of ways to reach step n w(n-1) routes whose last move is a single step w(n-2) routes whose last move is a double step What it says in wordsSort every route by its last move; the two groups do not overlap and together cover everything.Can you check 89 a second way?
Count by how many double steps you take. With k doubles, you make 10 - 2k singles, so 10 - k moves in all, and you only have to choose which k of those moves are the doubles. Summing the binomial counts over k = 0 to 5 gives 1 + 9 + 28 + 35 + 15 + 1 = 89, the same answer by a completely different road. Saying a second check aloud is worth more than the answer itself in a first round, because it shows you do not trust a pattern you have not tested.
Double steps k Moves in total Ways to place the doubles 0 10 C(10, 0) = 1 1 9 C(9, 1) = 9 2 8 C(8, 2) = 28 3 7 C(7, 3) = 35 4 6 C(6, 4) = 15 5 5 C(5, 5) = 1 total 89 Counting routes by the number of double steps gives 89 again, which confirms the Fibonacci running sum. What does the interviewer usually ask next?
Two things. First, allow steps of one, two or three: the same last-move argument gives w(n) = w(n - 1) + w(n - 2) + w(n - 3), and the count for ten steps becomes 274. Second, the coding version. A recursive function that calls itself for n - 1 and n - 2 recomputes the same steps again and again: for 30 steps it makes 1,664,079 calls to return 1,346,269. Storing each step's count once, or just keeping the last two numbers in a loop, does the job in 30 additions. The counts grow by about 1.618, the golden ratio, per step, which is the limitation of any approach that lists routes rather than counting them.
Where candidates lose it
The fast wrong answer is 2^10 = 1,024, treating each of ten stairs as a binary choice. A double step consumes two stairs, so routes have different numbers of moves and the choices are not ten independent coin flips.
The second loss is an off-by-one in the starting values, which lands on 55 or 144. Write the first three steps out by hand: 1 way to step 1, 2 ways to step 2, 3 ways to step 3. Anchor the sequence there and the tenth term is 89.
What the interviewer asks next
- What if you can also take three steps at a time?
- How many ways are there if step 5 is broken and cannot be stood on?
- Write code that counts the ways for 1,000 steps without the recursion blowing up.
Asked at Tower Research Capital, Intern Interview -, New York, 2012 (Wall Street Oasis):
How many ways can you jump up stairs if you can only jump either 1 or 2 steps?
093Three dice: red has faces 2, 6 and 7; green has 1, 5 and 12; blue has 3, 4 and 8, each face appearing twice. You and I each pick a die and roll once, and the higher number wins. Which die do you want, and does it matter who picks first?Belvedere TradingChicago · 2022
Try it first
Which die is best against the other two?
Show the worked solution
No die is best: red beats green, green beats blue and blue beats red, each with probability 5/9. So who picks first matters a great deal. Let me choose, then take the die that beats mine and win 5/9 of the time. If you are forced to pick first, every choice loses 5/9 of the time against an opponent who knows the cycle.
How do you work out who beats whom?
Each matchup has only nine equally likely pairs of faces, so write the 3 by 3 grid and count. Red against green: red's 2 beats only the 1, while its 6 and 7 each beat the 1 and the 5, for 1 + 2 + 2 = 5 wins out of 9. Do the same for the other two pairs and every matchup comes out 5 to 4: red over green, green over blue, blue over red. It is rock, paper, scissors built out of dice, and in rock, paper, scissors nobody asks which hand shape is best.
Counting the nine face pairs in each matchup shows red beats green, green beats blue and blue beats red, each in 5 of 9 cases, so the three dice form a cycle and the second player can always pick a die that wins 5/9 of the time. The relationshipP(R > G) the chance red's roll beats green's 1 + 2 + 2 the wins for red's faces 2, 6 and 7 in turn What it says in wordsCount, face by face, how many of the opponent's three faces each face beats, and divide by nine.Why does green lose to red when green has the higher average?
The averages are red 5, green 6 and blue 5. Winning is about how often, not by how much. Green's 12 wins every time it shows, but it shows only a third of the time, and green's other two faces, 1 and 5, lose to both of red's high faces. A higher mean and a higher chance of winning are different things, and the gap between them is the whole puzzle. Change the rules so the winner collects the difference between the two numbers, and green's expected margin against red is 6 - 5 = +1: now you want green against red, and blue against red is a dead heat at 0.
Where does a trader meet the same thing?
Head-to-head comparisons need not line up into a ranking. Strategy A can beat strategy B on more days than not, B can beat C, and C can beat A, whenever one of them earns its money in rare large wins, as green does. Before choosing between strategies, decide whether you care about how often you win or how much you make, because under the first a cycle like this one means there may be no best choice at all. The limitation of the puzzle is that it is one roll; over many rolls with the total score counted, the mean matters more and green's 12 starts to pay.
Where candidates lose it
The common slip is choosing green because its average, 6, is highest. The game pays for winning, not for margin, and green loses to red five times in nine.
The second loss is answering the first question and missing the second. Because the dice form a cycle, the real answer is strategic: insist that your opponent picks first. Saying that unprompted is what the interviewer is listening for.
What the interviewer asks next
- If each player rolls their die twice and adds the results, does the cycle still hold?
- Design three dice whose faces sum to the same total and still form a cycle.
- With three players each taking one die, can any die be favoured against both others?
Asked at Belvedere Trading, Prop Trading, Chicago, 2022 (Wall Street Oasis):
You have 3 dice: red has 2, 6, 7; green has 1, 5, 12; blue has 3, 4, 8.
094Two points are chosen independently and uniformly on the surface of a unit sphere. What is the expected distance between them measured along the surface, that is, the great-circle distance?Tower Research CapitalNew York · 2019
Try it first
What is the expected great-circle distance?
Show the worked solution
pi/2, about 1.571. Rotate the sphere so the first point sits at the north pole; nothing changes, because the second point is uniform. On a unit sphere the surface distance is the polar angle theta of the second point. The northern and southern hemispheres are mirror images, so theta is as likely to be pi/2 - t as pi/2 + t, and its mean is pi/2.
Why can you put the first point at the pole?
Ask how far apart two random towns are on a perfectly round planet, and you can simply stand in one of them: the globe looks the same from every spot on it. Symmetry lets you fix one point anywhere, so the problem shrinks to one random point and its angle from the pole. On a sphere of radius 1, the distance along the surface between the pole and a point at polar angle theta is theta itself, measured in radians, so the question becomes: what is the average polar angle of a uniform point?
With the first point at the pole, the surface distance is the polar angle theta of the second point, whose density (1/2) sin theta is symmetric about pi/2, so the expected distance is pi/2; a uniform angle, the dashed line, puts too many points near the poles. The relationshiptheta the polar angle of the second point, equal to the surface distance on a unit sphere f(theta) the density of that angle sin theta the relative size of the band of latitude at angle theta What it says in wordsThere is more surface near the equator than near the poles, in proportion to sin theta, and that density is symmetric about pi/2.Where does sin theta come from? The circle of latitude at angle theta from the pole has circumference 2 pi sin theta, so a thin band there holds surface in proportion to sin theta: almost none near the poles, the most at the equator. Archimedes put it more neatly: the area of a band is proportional to its height along the axis, so cos theta is uniform between -1 and 1. The density (1/2) sin theta is a mirror image about pi/2, so the mean is pi/2 without doing the integral. Integration by parts confirms it, a numerical integral gives 1.5708, and a seeded simulation of 100,000 pairs gives 1.570.
If a uniform angle gives the same mean, why does the shape matter?
Because the mean survives by luck of symmetry and almost nothing else does. Choosing theta uniformly on 0 to pi crowds points near the poles. Ask for the chance the two points are within 60 degrees of each other and the correct answer is (1 - cos 60 degrees)/2 = 0.25, while the uniform angle says 0.33. Ask for the expected straight-line chord, 2 sin(theta/2), and the correct density gives 4/3, about 1.333, while the uniform angle gives 4/pi, about 1.273. The simulation gives 1.333 for the chord. This is the limitation of the shortcut: it answers this one question and must not be reused for the next.
Where candidates lose it
The commonest wrong answer is 4/3, the expected straight-line chord, which some candidates remember from a related puzzle. The question asks for distance along the surface, which on a unit sphere is the angle itself.
The second loss is the right answer for the wrong reason: picking the angle uniformly between 0 and pi. The mean comes out right by symmetry, but any follow-up on the chord or on the chance of being close gives the wrong number. Say that the band of latitude grows like sin theta.
What the interviewer asks next
- What is the expected straight-line distance between the two points?
- What is the probability that the two points are within 60 degrees of each other?
- Four points are chosen uniformly on a sphere. What is the chance they all lie in one hemisphere?
Asked at Tower Research Capital, Quantitative Research, New York, 2019 (Wall Street Oasis):
a 3d geometry question about the surface distance between points chosen randomly on the surface of a sphere
096A desk's daily P&L in Rs lakh over seven days is -1, 2, 4, -9, 8, -2, 3. Which run of consecutive days has the largest total, and how do you find it in one pass through the data?Wolverine TradingChicago · 2014
Try it first
Which run has the largest total?
Show the worked solution
Days 5 to 7, the run 8, -2, 3, which totals Rs 9 lakh. Walk through the days keeping the best total of a run ending today: either today alone or today added to yesterday's best run, whichever is larger. Record the largest value you see. The run 2, 4 looks attractive but totals only 6, and the -9 day makes it pointless to carry anything before it.
How do you find the best run without checking every start and end day?
Picture walking along a road with toll booths that either pay you or charge you. You may choose where to start and stop collecting. If the purse you have carried from earlier booths is in the red, the sensible move is to drop it and start fresh at the next booth. A run ending today is worth extending from yesterday only if the best run ending yesterday is positive; if it is negative it can only drag today down, so today starts a new run. That rule looks at each day once. Checking every pair of start and end days means 28 runs for seven days and 31,375 for a trading year of 250 days.
Carrying the best run forward only while it is positive, the running total drops to -3 after the loss of 9 and restarts at 8 on day 5, so the best run is 8, -2, 3 with a total of 9, ahead of the tempting run 2, 4 at 6. The relationshipx_t the P&L on day t c_t the best total of a run that ends on day t best the largest c_t seen so far What it says in wordsThe best run ending today either starts today or extends the best run ending yesterday; keep whichever is bigger, and remember the biggest.Day P&L Best run ending today Best so far 1 -1 -1 -1 2 2 2 2 3 4 6 6 4 -9 -3 6 5 8 8 8 6 -2 6 8 7 3 9 9 Running the rule day by day, the best run ending today drops to -3 after day 4, restarts at 8 on day 5 and reaches 9 on day 7, which is the answer. Why does the tempting run 2, 4 lose?
Because one later day beats it on its own. After the -9, the best run ending on day 4 is 6 - 9 = -3, so the rule drops the past and day 5 starts fresh at 8. The -2 on day 6 dips the run to 6, but the 3 on day 7 lifts it to 9. The best run can contain a losing day: 8, -2, 3 beats 8 alone because the day after the loss more than repays it. A candidate who stops a run at the first red day misses this, and the brute-force check over all 28 runs confirms 9 is the maximum.
What edge cases does the interviewer probe?
Three. If every day is a loss, the answer should be the least bad single day, so start the best at the first day's value, not at zero, or you will report an empty run worth 0. To report which days, store the start index whenever you restart and copy it when you record a new best. The same pass with the signs flipped finds the worst run, here -9, the single -9 day. The limitation on a desk is that the best run in hindsight is a selected statistic: a strategy that is judged by its best stretch will always look better than it trades.
Where candidates lose it
The quick wrong answer is the run 2, 4, because it is the first good stretch. The 8 on day 5 beats it alone, and carrying 8 through -2 and 3 beats 8.
The second loss is in the code: starting the best total at zero, which reports 0 for a week of all losses, or restarting at every losing day instead of only when the running total itself turns negative. State the rule exactly: carry yesterday's run only while it is positive.
What the interviewer asks next
- Return the start and end days of the best run, not just its total.
- What does your code return if every day in the series is a loss?
- Find the best run if you may skip at most one day inside it.
Asked at Wolverine Trading, Quantitative Research, Chicago, 2014 (Wall Street Oasis):
Develop an algorithm to find out the section that contains the maximum sum.
098Two traders' monthly P&L are independent and normal. A has mean Rs 10 lakh and standard deviation Rs 3 lakh; B has mean Rs 8 lakh and standard deviation Rs 4 lakh. What is the probability that A out-earns B in a given month?DRWLondon · 2025
Try it first
Pick the probability that A earns more than B in a month.
Show the worked solution
About 65.5%. The gap A - B is normal with mean 10 - 8 = Rs 2 lakh and variance 3^2 + 4^2 = 25, so its standard deviation is Rs 5 lakh. A out-earns B when the gap is positive, and zero sits 2/5 = 0.4 standard deviations below the mean, so the probability is Phi(0.4), about 65.5%. The better trader loses about one month in three.
Why do the variances add when you subtract?
You and a colleague set off for the same meeting from different places, and each journey is uncertain by a few minutes. The gap between your two arrival times is more uncertain than either journey, not less, because either of you can be the late one. Subtracting an independent random amount adds its noise, so Var(A - B) = Var A + Var B = 9 + 16 = 25, and the gap's standard deviation is 5, not 1. The mean subtracts as you would expect, 10 - 8 = 2. The gap is normal because a difference of independent normals is normal.
The two traders' monthly P&L overlap heavily, and the gap A - B has mean 2 and standard deviation 5, so the area above zero where A wins is only 65.5%, leaving B ahead in 34.5% of months. The relationshipmu_A, mu_B the mean monthly P&L, 10 and 8 sigma_A, sigma_B the standard deviations, 3 and 4 Phi the standard normal cumulative distribution What it says in wordsThe gap's mean is the difference of the means, its variance the sum of the variances, and the answer is how many standard deviations zero sits below that mean.How much does a longer comparison window help?
A lot, and at a predictable rate. Over a quarter of independent months the total gap has mean 6 and standard deviation 5 x sqrt(3), about 8.7, so A comes out ahead with probability 75.6%. Over a year the mean is 24 and the standard deviation 5 x sqrt(12), about 17.3, so the probability is 91.7%. The edge grows with the number of months and the noise with its square root, so the z-score grows with the square root of time. A risk manager who ranks traders on one month of P&L is ranking mostly noise.
What if the two traders' P&L are correlated?
Then the shared part cancels in the gap. With correlation 0.5, the variance is 9 + 16 - 2 x 0.5 x 3 x 4 = 13, a standard deviation of 3.61, and A wins with probability 71.0%. Positive correlation makes the comparison sharper because common market moves drop out of the difference; negative correlation does the opposite. The limitation is the normal assumption: real P&L has fat tails and skew, and if one trader earns through rare large months, the month-by-month win rate can disagree with the mean, so check the shape before trusting the 65.5%.
Where candidates lose it
The commonest slip is subtracting the standard deviations, 4 - 3 = 1, which makes A look almost certain to win at 97.7%. Noise does not cancel when you subtract independent variables; it adds.
The second loss is subtracting the variances, 16 - 9, or adding the standard deviations, 3 + 4. Square, add, then take the root: sqrt(9 + 16) = 5. The answer is then a z-score of 0.4, and Phi(0.4) is about 0.655.
What the interviewer asks next
- What is the probability that A out-earns B over a full year of independent months?
- If their monthly P&L has correlation 0.5, what is the answer?
- What is the probability that A out-earns B by more than Rs 5 lakh in a month?
Asked at DRW, Trading, London, 2025 (Wall Street Oasis):
technical interview based on normal distribution and market making
099A stock is worth either 100 or 110, with equal probability. 20% of the traders who arrive know the true value: they buy if it is 110 and sell if it is 100. The other 80% buy or sell at random, half and half. Where should a market maker set its ask so that it breaks even, on average, when someone buys from it?Jane StreetNew York · 2025
Try it first
Where should the ask be?
Show the worked solution
Set the ask at 106, and by the same logic the bid at 104. If the stock is worth 110, a buy arrives with probability 0.2 + 0.8 x 0.5 = 0.6; if it is worth 100, with probability 0.4. Given a buy, Bayes puts the chance of 110 at 0.6, so the stock is worth 106 to the market maker selling it. The spread of 2 is the price of trading against informed flow.
Why can't the market maker just quote the expected value of 105?
A second-hand car dealer who pays the average price for every car will find that the owners of good cars go elsewhere and the owners of bad ones queue up. Who chooses to trade with you is information. A market maker does not care what the stock is worth on average; it cares what the stock is worth given that someone has just chosen to buy from it. At an ask of 105, noise buyers are harmless, a loss of 5 when the stock is worth 110 and a gain of 5 when it is worth 100. Informed buyers only appear in the 110 world, and they cost 0.5 per arriving trader on average, so 105 is a losing quote.
Tracing who sends a buy order in each world, buys come with probability 0.30 from the 110 world and 0.20 from the 100 world, so a buy lifts the chance of 110 from 0.5 to 0.6 and the break-even ask is 106. The relationship0.6 the chance of a buy when the stock is worth 110: 0.2 informed plus 0.8 x 0.5 noise 0.4 the chance of a buy when the stock is worth 100: noise only V the stock's true value What it says in wordsSet the ask at the value of the stock conditional on being bought from, which Bayes' rule gives directly.What sets the width of the spread?
The share of informed traders and the size of what they know. With a share alpha informed, a buy is alpha + (1 - alpha)/2 likely in the high world and (1 - alpha)/2 in the low world, and the ask works out to 105 + 5 alpha. The spread is a fee for adverse selection: it is zero when nobody is informed and widens to the full 100 to 110 range when everyone is. The table runs the formula for a few shares. Order processing and inventory costs add to this in real markets, but the information component is what makes spreads jump around earnings and news.
Informed share Ask Bid Spread 0% 105 105 0 10% 105.5 104.5 1 20% 106 104 2 50% 107.5 102.5 5 100% 110 100 10 The break-even spread equals the informed share times the 10-point value gap, so it is 2 at 20% informed and 5 at 50% informed. What happens after the first trade?
The market maker updates. After one buy, the chance of 110 is 0.6, and if a second buy arrives the same Bayes step lifts it to 0.692, so the next ask is about 106.92. Each order moves the quotes towards the true value, which is how prices come to reflect what the informed traders know. The limitation is that the model has one share size, no inventory risk and no competition between market makers; real desks also skew quotes to manage position, which this puzzle leaves out.
Where candidates lose it
The fast wrong answer is 105, the unconditional expected value. It ignores that the act of buying is evidence: informed traders buy only when the stock is worth 110, so a market maker at 105 loses on every informed buyer and only breaks even on noise traders.
The second loss is overreacting and quoting 110 because some buyers are informed. Most buyers are noise traders, and a quote at 110 drives them away. Bayes gives the exact weight, 0.6 on the high value, and the ask of 106.
What the interviewer asks next
- Where should the bid be, and why is the spread symmetric here?
- After one buy at 106, where is the next ask?
- How does the spread change if half of the traders are informed?
Asked at Jane Street, Generalist, New York, 2025 (Wall Street Oasis):
It was a probability theory based quant trading style market making questions which were intense

