Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
020A broad stock index has returned an average of 11% a year over the past 30 years, with an annual standard deviation of 16% (illustrative figures). Assuming yearly returns are independent, give a 95% confidence interval for its true expected annual return.Old Mission CapitalChicago · 2025
Try it first
Roughly how wide is the 95% interval?
Show the worked solution
About 5.3% to 16.7%. The standard error of a 30-year average is 16% divided by the square root of 30, 2.92 points. A 95% interval is 1.96 standard errors either side: 11% plus or minus 5.7. With a t value for 29 degrees of freedom it is a touch wider, plus or minus 6.0. Thirty years of data still leave the expected return very uncertain.
Why is the band so wide after thirty years?
Imagine judging a new cricketer's true batting average from a handful of innings. Scores swing wildly from one innings to the next, so a few innings tell you little. The precision of an average grows only with the square root of the number of observations, and yearly returns swing by 16 points, so 30 years give a standard error of about 2.9 points. That is large next to an average of 11%: the true figure could plausibly be 6% or 16%.
Around an 11% average with 16% annual volatility, the 95% interval is plus or minus 9.9 points with 10 years of data, 5.7 points with 30 years and still 2.9 points with 120 years, because precision grows only with the square root of the sample. How do you set it up in the room?
State the assumption, then compute. Treat the 30 annual returns as independent draws with a standard deviation of 16%, so the sample average has a standard error of 16 over root 30. Root 30 is about 5.48, so the standard error is 2.92. Multiply by 1.96: 5.73. The interval is 5.3% to 16.7%. If the interviewer pushes, note that with only 30 observations a t distributionThe distribution used for an average when the standard deviation is estimated from the same small sample; it has fatter tails than the normal. with 29 degrees of freedom gives a multiplier of about 2.045, widening the band slightly.
The relationship\bar r the average annual return, 11% \sigma the annual standard deviation, 16% n the number of yearly observations, 30 What it says in wordsThe interval is the average plus or minus 1.96 standard errors, and the standard error is volatility over the square root of years.What limitation should you add?
Halving the band needs four times the data: 120 years still only pins the average to within about 2.9 points, and markets change over such spans. Returns are also not quite independent from year to year, fat tails make the normal multiplier optimistic, and the arithmetic average overstates the compound growth rate. The practical lesson is that expected returns are the hardest input in finance to estimate, far harder than volatility, which is why many quant processes lean on risk estimates and treat return forecasts with caution.
Where candidates lose it
The common mistake is using 16% as the width, the spread of a single year, rather than the standard error of the average. That gives an interval from -20% to 42%, which describes one year's return, not the long-run mean.
The opposite slip is dividing by 30 instead of the square root of 30, which gives a band of about 1 point and badly overstates what the data can tell you.
What the interviewer asks next
- How many years of data would you need to pin the mean to within plus or minus 1 point?
- How would monthly data change the interval for the mean?
- Why is volatility much easier to estimate than the mean from the same data?
Asked at Old Mission Capital, Prop Trading, Chicago, 2025 (Wall Street Oasis):
Confidence interval on S&P 500 return past 30 years
029Rs 1 was invested in a broad stock index 30 years ago. If yearly log returns are independent with mean 7% and standard deviation 16% (illustrative inputs), give a median and a 95% interval for what it is worth today.Old Mission CapitalChicago · 2025
Try it first
Which is the best central 95% range for the Rs 1 today?
Show the worked solution
Median about Rs 8.2; 95% interval roughly Rs 1.5 to Rs 45. Log returns add, so after 30 years the log of wealth has mean 30 x 0.07 = 2.1 and standard deviation 0.16 x √30 = 0.88. The median is e to the 2.1, about 8.2. The band is e to the power 2.1 plus or minus 1.96 x 0.88. In rupees it is lopsided, and the mean, about Rs 12, sits above the median.
Why work in log returns rather than percentage returns?
Pay rises compound: 10% and then another 10% is 21%, not 20%. Logs turn that multiplication into addition. Log returns add across years, so the 30-year log return is a sum of 30 yearly pieces, and a sum of independent pieces is close to normal. With mean 0.07 and standard deviation 0.16 a year, the sum has mean 2.1 and variance 30 x 0.16 squared, so a standard deviation of 0.16 x √30, about 0.876. The spread grows with the square root of time, not with time.
The relationshipW_30 value of the Rs 1 after 30 years mu = 0.07 mean yearly log return, an illustrative input sigma = 0.16 standard deviation of the yearly log return 1.96 the number of standard deviations that cuts off 2.5% in each tail of a normal What it says in wordsBuild the interval for the log of wealth, where it is symmetric, then exponentiate the two ends.On a log scale the 95% band is symmetric around the median of Rs 8.2, running from Rs 1.5 to Rs 45.5; on an ordinary rupee scale the same band reaches Rs 6.7 below the median and Rs 37.3 above it, and the mean of Rs 12.0 sits right of the median. Why is the band so lopsided, and where does the mean sit?
Symmetric in the exponent means lopsided in rupees. Going 1.96 standard deviations down divides the median by e to the 1.72, a factor of 5.6; going the same distance up multiplies by 5.6. Dividing and multiplying by the same factor leaves Rs 6.7 of room below the median and Rs 37.3 above it. The same skew separates mean from median. The mean of a lognormalA variable whose logarithm is normally distributed; it is always positive and skewed to the right. variable is e to the power (mean plus half the variance), about Rs 12.0 here, because a few very good paths pull the average up while most paths finish below it.
Close with the limits. The 7% and 16% are illustrative inputs, not a claim about any real index. The calculation assumes independent years and constant volatility; real markets have fat tails and calm and stormy regimes, so treat the band as a floor on the true uncertainty. What the interviewer is testing is whether you scale the mean with t and the volatility with √t, and exponentiate only at the end. One useful extra: the chance the Rs 1 is worth less than Rs 1 is the chance the log falls below zero, about 0.8%.
Where candidates lose it
The most common slip is building the interval in rupees: take 8.2 and add and subtract a symmetric amount, which can even run below zero. Build it in logs and exponentiate the two ends.
The second is scaling the 16% by 30 instead of √30, which gives a log standard deviation of 4.8 and a band from paise to crores. Variance adds across years; standard deviation grows with the square root.
What the interviewer asks next
- What is the probability the Rs 1 is worth less than Rs 1 today?
- If you are given the average percentage return rather than the average log return, how do you convert?
- How does the band change over a 10-year horizon?
Asked at Old Mission Capital, Prop Trading, Chicago, 2025 (Wall Street Oasis):
Confidence interval of portfolio value if you invested $1 in S&P 500 30 years ago
041Five observations come from a uniform distribution on 0 to theta: 3.1, 7.4, 5.2, 9.0 and 1.8. What is the maximum likelihood estimate of theta, is it biased, and how would you correct it?AQR Capital ManagementTown of Greenwich · 2022
Try it first
What is the maximum likelihood estimate of theta?
Show the worked solution
The MLE is 9.0, the largest observation; it is biased low, and multiplying by (n + 1)/n = 6/5 gives an unbiased 10.8. Each observation has density 1/theta when theta covers it, so the likelihood is theta to the minus 5 for theta at least 9.0 and zero below. That peaks at 9.0. But the sample maximum averages 5/6 of theta, never above it, so scale it up by 6/5.
Why does the likelihood peak at the largest observation?
Suppose raffle tickets are numbered 1 to N and you see five, the highest being 90. N cannot be below 90, and the smaller N is, the more likely it was to produce those particular five tickets. For a uniform on 0 to theta, each observation has density 1/theta, so the likelihood is theta to the minus 5, which only falls as theta grows, but it is zero for any theta below an observation. The best allowed value is the smallest theta that covers all the data: the maximum, 9.0. Calculus does not help here, because the peak sits at the edge where the likelihood jumps from zero.
The likelihood is zero for theta below 9.0, peaks at 9.0 and then falls as theta to the minus 5, down to 40% of the peak at 10.8; on the data line, the MLE of 9.0 sits at the largest observation, the method of moments gives 10.6 and the bias-corrected estimate is 10.8. Why is 9.0 biased, and what is the right correction?
The sample maximum can never exceed theta, so it can only err on the low side. Five points drop into 0 to theta and cut it into six gaps of the same average size, so the largest point sits on average one gap short of theta: at 5/6 of theta. Scaling the maximum by (n + 1)/n removes that bias: 9.0 x 6/5 = 10.8. The same logic underlies the classic serial-number estimation problem from wartime production counts.
The relationshiptheta the unknown upper end of the uniform n = 5 number of observations max the largest observation, 9.0 What it says in wordsThe likelihood peaks at the sample maximum, which on average falls short of theta by a factor n/(n + 1), so scale it up.An interviewer may ask why not use twice the mean, 2 x 5.3 = 10.6, which is also unbiased. The corrected maximum is far more precise: its variance is theta squared over n(n + 2), against theta squared over 3n for twice the mean, so twice the mean is 2.3 times as variable with five points. Twice the mean can even land below the largest observation, an estimate the data have already ruled out. The limitation of the correction is that unbiased is not the only goal: the multiple of the maximum with the smallest mean squared error is (n + 2)/(n + 1), which gives 10.5 here, and saying you would choose by the loss that matters shows you know the trade.
Where candidates lose it
The common loss is setting the derivative of the log-likelihood to zero, getting -5/theta = 0, and concluding there is no maximum. The maximum is at a boundary, where the indicator switches on, and that is the point of the question.
The second is answering 9.0 and stopping. The follow-up is always the bias; say that the maximum sits below theta on average and give the (n + 1)/n correction with its one-line reason.
What the interviewer asks next
- What is the MLE if the distribution is uniform on theta to 2 theta?
- Derive the variance of the corrected estimator.
- The observations come from a uniform on theta minus 1 to theta plus 1. What is the MLE now?
Asked at AQR Capital Management, Trading, Town of Greenwich, 2022 (Wall Street Oasis):
Derive the mle for some given distribution. Explain linear regression intuitively and derive the ols estimate.
069Daily returns are drawn from a normal with mean zero and standard deviation 1% on 90% of days, and from a normal with mean zero and standard deviation 4% on the other 10%. What are the overall standard deviation and the kurtosis of daily returns?Two SigmaNew York · 2025
Try it first
What is the kurtosis of the mixture?
Show the worked solution
The standard deviation is sqrt(2.5), about 1.58%, and the kurtosis is 12.72, against 3 for a normal. Moments of a mixture are weighted averages of the pieces' moments. The variance is 0.9 x 1 + 0.1 x 16 = 2.5. A normal's fourth moment is 3 sigma^4, so the fourth moment is 3 x (0.9 x 1 + 0.1 x 256) = 79.5. Kurtosis is 79.5 / 6.25 = 12.72.
Why does mixing two normals create fat tails?
Think of a city's daily traffic: most days are ordinary, and a few days a year there is a festival or a strike and everything is wild. Averaged across the year, the typical day looks calmer than the average suggests, and the extreme days are far more extreme than a single bell curve would allow. Mixing a calm regime with a rare wild one concentrates most days near zero and puts the rest far out, which is exactly what kurtosis measures: a high peak with heavy tails. Every piece is normal; the mixture is not.
With the same 1.58% standard deviation, the mixture is more peaked than the normal and its tail is far heavier: a daily move beyond 6% comes about 1 day in 75 under the mixture against about 1 day in 6,766 under the normal. How do you compute the moments without integrating?
Condition on the regime. Any moment of a mixture is the weighted average of that moment in each regime, because the density itself is the weighted average of the two densities. The second moment is 0.9 x 1^2 + 0.1 x 4^2 = 2.5, so the standard deviation is about 1.58%. For the fourth, use the fact that a normal with standard deviation sigma has fourth moment 3 sigma^4: 3 x (0.9 x 1 + 0.1 x 256) = 79.5. Kurtosis is the fourth moment over the variance squared.
The relationshipw_i the regime weights, 0.9 and 0.1 sigma_i the regime standard deviations, 1% and 4% kappa kurtosis, the fourth moment divided by the variance squared; 3 for any normal What it says in wordsAverage the variances and the fourth moments across regimes, then compare the fourth moment with the squared variance.What does the number mean for risk?
Kurtosis is 3 times the ratio of the average of sigma^4 to the square of the average of sigma^2, so any variation in volatility pushes kurtosis above 3, and the more uneven the regimes, the further it goes. The practical cost shows in the tails. A risk model that fits one normal with a 1.58% standard deviation expects a move beyond 6% about once in 6,766 days, roughly once every 27 years of trading; the mixture produces one about once in 75 days, several times a year. This is the simplest model of volatility clustering, and the reason daily returns on real assets show kurtosis well above 3.
Where candidates lose it
The trap is answering 3 because each piece is normal, or averaging the two standard deviations to 1.3% and treating the mixture as one normal. Mixtures average densities and moments, not shapes or standard deviations.
The second slip is forgetting the factor of 3 in a normal's fourth moment, or reporting excess kurtosis without saying so. Say which you mean: kurtosis 12.72, excess kurtosis 9.72.
What the interviewer asks next
- What mix of the two regimes maximises the kurtosis for a fixed overall variance?
- If the two regimes had different means, what would happen to the skew?
- How would you estimate the two regime volatilities from a year of daily returns?
Asked at Two Sigma, Quantitative Research, New York, 2025 (Wall Street Oasis):
They asked a couple questions involving Mixture Gaussians (e.g., probability density and moments).
076You must predict a quantity with a single constant, and your data are 1, 2, 3, 4 and 40. Which constant minimises the mean squared error, which minimises the mean absolute error, and what does the difference tell you?Tower Research CapitalPrinceton · 2018
Try it first
Before you calculate: which pair of constants is right?
Show the worked solution
The mean, 10, minimises squared error; the median, 3, minimises absolute error. Squared error grows with the square of a miss, so the single 40 pulls the best constant towards it. Absolute error charges every unit of miss equally, so the best constant sits in the middle of the pack. Choosing the loss is choosing which of the two statistics you estimate.
Why does each loss land on a different constant?
Picture five friends choosing a meeting point on a straight road: four live at the 1, 2, 3 and 4 km marks and one at the 40 km mark. If the goal is the smallest total travel, you meet at 3 km: moving towards the far friend saves them one km per km moved but costs the other four one km each. If the goal is the smallest total of squared travel, the far friend's 37 km squared, 1,369, dominates everything and the meeting point slides to 10 km. Absolute error balances the count of points on each side, which is the median; squared error balances the total distance on each side, which is the mean.
Squared error is lowest at the mean, 226 at c = 10, and the median scores 275 on it; absolute error is lowest at the median, 8.2 at c = 3, and the mean scores 12 on it, so each constant is poor under the other loss. The relationshipx_i the five data points c the constant you predict #{x_i < c} how many points sit below c What it says in wordsSquared error is flat where the total distance above equals the total below, which is the mean; absolute error is flat where the count above equals the count below, which is the median.So which constant should you actually use?
It depends on what the 40 is, and that is the answer the interviewer wants to hear. If the 40 is a typing error or a one-off glitch in a price feed, the median is the honest summary: remove the 40 and the mean falls to 2.5, while the median barely moves. If the 40 is real, say the one big winning day in a strategy's P&L, the mean is the number that matters, because your total profit is the sum of the days and the sum is five times the mean. A robust estimator that ignores the big day would tell you the strategy earns 3 a day when it actually earns 10.
Where does this show up in a quant job?
Every regression makes this choice silently. Ordinary least squares minimises squared error and so fits conditional means; one extreme observation can swing the line. Least absolute deviation, or quantile regression at the 50% level, fits conditional medians and shrugs off the same extreme point. In practice, desks winsorise or clip returns before fitting squared-error models, or switch to a Huber loss that is squared near zero and linear in the tails. The limitation is that none of these fixes is free: every one of them throws away some of the information in genuine large moves.
Where candidates lose it
The common slip is to answer 10 for both, because the mean feels like the default best guess. The two losses answer different questions, and the interviewer is checking that you know squared error chases the outlier and absolute error does not.
The second loss is stopping at the arithmetic. Say what the 40 might be, a data error or a real big day, and which loss fits each case. That judgement is the point of the question.
What the interviewer asks next
- Which constant minimises the maximum absolute error, and what is that loss called?
- Add a sixth point at 5. What happens to the median, and to the set of absolute-error minimisers?
- Why does Lasso use an absolute-value penalty, and how is that related to this puzzle?
Asked at Tower Research Capital, Trading, Princeton, 2018 (Wall Street Oasis):
What if instead of minimizing mean squared error we look at mean absolute error?
088A strategy's true annualised Sharpe ratio is 1.0. Roughly how many years of returns do you need before a t-test rejects a zero mean at about the 5% level? What if the Sharpe ratio is 0.5?Viking Global InvestorsNew York · 2014
Try it first
How many years does a Sharpe 0.5 strategy need?
Show the worked solution
About 4 years at a Sharpe of 1, and about 16 years at 0.5. The t-statistic for a mean return is the Sharpe ratio times the square root of the number of years, so reaching t = 2 needs (2 / SR)^2 years. Halving the Sharpe ratio quadruples the evidence you need, and sampling daily instead of yearly does not shorten it.
Why does the t-statistic come out as Sharpe times root years?
A t-test on a mean divides the average return by its standard error, which is the volatility over the square root of the number of observations. With annual observations that ratio is (mean / volatility) x sqrt(years), and mean over volatility is exactly the annual Sharpe ratio. So a Sharpe of 1 gives t = sqrt(years): 2 after 4 years. A Sharpe of 0.5 gives t = 0.5 x sqrt(years): 2 only after 16 years. With the textbook 1.96 in place of 2 the numbers are 3.8 and 15.4 years; the round figures are what you say in the room.
The t-statistic equals the Sharpe ratio times the square root of years, so a Sharpe of 1 reaches t = 2 after 4 years and a Sharpe of 0.5 only after 16: halving the Sharpe quadruples the track record you need. The relationshipr bar the average annual return in excess of cash sigma the annual volatility N the number of years observed SR the annual Sharpe ratio, r bar over sigma What it says in wordsThe evidence for a real edge grows with the square root of time, scaled by the Sharpe ratio.Can you shortcut it with daily data?
This is the follow-up that separates candidates. Sampling daily gives about 252 times as many observations a year, but the daily Sharpe ratio is smaller by the square root of 252, because daily mean scales with time and daily volatility with its square root. The two effects cancel exactly, so the t-statistic depends on calendar time, not on how finely you slice it. Think of estimating a river's average level: measuring every minute instead of every day does not help if the river's slow swings are the uncertainty.
What makes the real requirement even longer?
Three things, each worth one sentence. Returns are not independent from year to year, and positive autocorrelation inflates the true standard error. If you tested twenty strategies and kept the best, a t of 2 is easy to get by luck, so the bar has to rise with the number of ideas tried. And the Sharpe ratio itself drifts as markets change, so a sixteen-year record may be measuring two different strategies. A desk that says a Sharpe 0.5 strategy is proven after three years is reading noise.
Where candidates lose it
The first loss is scaling linearly: a Sharpe half as big needs twice as long, so 8 years. The t-statistic grows with the square root of time, so the years scale with the square of 1/Sharpe, and 16 is right.
The second loss is proposing daily data as the fix. The number of observations goes up but the per-observation Sharpe goes down by the square root of that factor, and the two cancel.
What the interviewer asks next
- How many years does a Sharpe of 2 need, and why do high-frequency desks care?
- If you tested 50 strategies, roughly what t-statistic would you demand of the best one?
- How does positive autocorrelation in monthly returns change the answer?
Asked at Viking Global Investors, Quantitative Research, New York, 2014 (Wall Street Oasis):
how to reject a hypothesis test, what's your structure of your code, what's the sample size
098Two traders' monthly P&L are independent and normal. A has mean Rs 10 lakh and standard deviation Rs 3 lakh; B has mean Rs 8 lakh and standard deviation Rs 4 lakh. What is the probability that A out-earns B in a given month?DRWLondon · 2025
Try it first
Pick the probability that A earns more than B in a month.
Show the worked solution
About 65.5%. The gap A - B is normal with mean 10 - 8 = Rs 2 lakh and variance 3^2 + 4^2 = 25, so its standard deviation is Rs 5 lakh. A out-earns B when the gap is positive, and zero sits 2/5 = 0.4 standard deviations below the mean, so the probability is Phi(0.4), about 65.5%. The better trader loses about one month in three.
Why do the variances add when you subtract?
You and a colleague set off for the same meeting from different places, and each journey is uncertain by a few minutes. The gap between your two arrival times is more uncertain than either journey, not less, because either of you can be the late one. Subtracting an independent random amount adds its noise, so Var(A - B) = Var A + Var B = 9 + 16 = 25, and the gap's standard deviation is 5, not 1. The mean subtracts as you would expect, 10 - 8 = 2. The gap is normal because a difference of independent normals is normal.
The two traders' monthly P&L overlap heavily, and the gap A - B has mean 2 and standard deviation 5, so the area above zero where A wins is only 65.5%, leaving B ahead in 34.5% of months. The relationshipmu_A, mu_B the mean monthly P&L, 10 and 8 sigma_A, sigma_B the standard deviations, 3 and 4 Phi the standard normal cumulative distribution What it says in wordsThe gap's mean is the difference of the means, its variance the sum of the variances, and the answer is how many standard deviations zero sits below that mean.How much does a longer comparison window help?
A lot, and at a predictable rate. Over a quarter of independent months the total gap has mean 6 and standard deviation 5 x sqrt(3), about 8.7, so A comes out ahead with probability 75.6%. Over a year the mean is 24 and the standard deviation 5 x sqrt(12), about 17.3, so the probability is 91.7%. The edge grows with the number of months and the noise with its square root, so the z-score grows with the square root of time. A risk manager who ranks traders on one month of P&L is ranking mostly noise.
What if the two traders' P&L are correlated?
Then the shared part cancels in the gap. With correlation 0.5, the variance is 9 + 16 - 2 x 0.5 x 3 x 4 = 13, a standard deviation of 3.61, and A wins with probability 71.0%. Positive correlation makes the comparison sharper because common market moves drop out of the difference; negative correlation does the opposite. The limitation is the normal assumption: real P&L has fat tails and skew, and if one trader earns through rare large months, the month-by-month win rate can disagree with the mean, so check the shape before trusting the 65.5%.
Where candidates lose it
The commonest slip is subtracting the standard deviations, 4 - 3 = 1, which makes A look almost certain to win at 97.7%. Noise does not cancel when you subtract independent variables; it adds.
The second loss is subtracting the variances, 16 - 9, or adding the standard deviations, 3 + 4. Square, add, then take the root: sqrt(9 + 16) = 5. The answer is then a z-score of 0.4, and Phi(0.4) is about 0.655.
What the interviewer asks next
- What is the probability that A out-earns B over a full year of independent months?
- If their monthly P&L has correlation 0.5, what is the answer?
- What is the probability that A out-earns B by more than Rs 5 lakh in a month?
Asked at DRW, Trading, London, 2025 (Wall Street Oasis):
technical interview based on normal distribution and market making
