Derivatives Foundation puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 66
- Topics
- 12
- Hard
- 29
014The sample variance with n minus 1 in the denominator is an unbiased estimate of the true variance. Is its square root an unbiased estimate of the standard deviation? If not, which way is it off, and does it matter when you estimate volatility from 20 daily returns?Squarepoint CapitalLondon · 2026
Try it first
Pick before you reason: the sample standard deviation on average is
Show the worked solution
No. The sample standard deviation is biased low, by about 1.3% at n = 20, and for a volatility estimate that is small next to the sampling noise. The square root is concave, so by Jensen's inequality the expected value of the root is below the root of the expected value. For normal returns the exact factor is c4(n): the expected sample standard deviation is 0.9869 times the true one at n = 20. On the same 20 points the estimate's own standard error is around 16%, so the bias is not what you should worry about.
Why does an unbiased variance give a biased standard deviation?
If you average the areas of several square plots and then take the square root, you do not get the average of their side lengths; the big plots pull the area average up more than they pull the side average. Unbiasedness is a statement about averages, and averages do not pass through a curved function: for a concave function like the square root, the average of the outputs is below the output of the average. That is Jensen's inequality, and it is the whole answer. The variance is unbiased; its root is not, and the direction is down.
Two sample variances of 0.1 and 1.9 times the truth average to 1.0, whose root is 1.00, but their roots average only 0.85, and for 20 normal returns the exact expected shortfall is c4(20) = 0.9869, so the sample standard deviation runs about 1.3% low on average. The relationships the sample standard deviation, the root of the n minus 1 sample variance sigma the true standard deviation c4(n) the exact correction factor for normal data, always below 1 and rising to 1 as n grows What it says in wordsFor normal data the sample standard deviation underestimates the true one by a known factor that depends only on the sample size.n c4(n) shortfall 5 0.9400 6.0% 10 0.9727 2.7% 20 0.9869 1.3% 50 0.9949 0.5% 100 0.9975 0.2% The shortfall of the sample standard deviation falls quickly with the sample size, from about 6% at five observations to about 1.3% at twenty and half a percent at fifty. Does it matter for a 20-day volatility estimate?
Not much, and saying why is the second half of the marks. The bias is 1.3%, but the standard error of a standard deviation from 20 observations is roughly 1 over the root of 2 times 19, about 16% of the true value, so the noise is more than ten times the bias. You can multiply by 1/c4 to remove the bias if you like, but you cannot remove the noise without more data, and 20 daily returns is simply a short window. On a volatility desk the honest answer is that a 20-day estimate of 16% could easily have been 13% or 19%.
What assumption does the exact factor need?
Normal returns. The direction of the bias, downward, holds for any distribution, because it comes from concavity alone; the size of the factor c4 depends on the distribution, and fat-tailed daily returns change it. Say the general result first, Jensen, then the normal-case number, then the limitation. That order shows you know what is a theorem and what is a model.
Where candidates lose it
The fast wrong answer is yes, on the grounds that the variance is unbiased and the root is just a relabelling. The root is a curved function, and expectations do not commute with curved functions; the sign of the curvature gives the direction.
The second loss is stopping at biased low. The interviewer asked whether it matters, and the answer is a comparison of 1.3% of bias against roughly 16% of sampling noise. Numbers, not adjectives.
What the interviewer asks next
- Derive the direction of the bias from Jensen's inequality in one line.
- What is the standard error of the sample standard deviation for normal data with n observations, roughly?
- Would you use the n or the n minus 1 version for a volatility estimate, and does the choice matter at n = 20?
Asked at Squarepoint Capital, Quantitative Research, London, 2026 (Wall Street Oasis):
Is the square root of the unbiased estimator for sample variance unbiased for standard deviation?
026I will pay you if at least 60% of the flips of a fair coin come up heads. Do you want 10 flips or 100 flips?Hudson River TradingNew York · 2020
Try it first
Before you count anything: which do you take?
Show the worked solution
Take 10 flips. You are betting on luck, and luck averages out as the trials pile up. With 10 flips, 6 or more heads happens 386 times in 1,024, about 37.7%. With 100 flips, 60 or more heads happens about 2.8% of the time. The share of heads settles towards 50% at the rate of one over the square root of n, so the 60% line gets harder to reach with every extra flip.
Why does the number of flips change the odds at all?
Think of a school with two cricket teams, one of eleven and one of a hundred and ten. If someone offers a prize for a team whose average height is 10 cm above the national average, the small team is the one that can win: one or two tall players move its average, while the big team's average is pinned down by sheer numbers. The same coin, flipped more often, produces a share of heads that sits ever closer to a half, so a payout that needs an unusual share wants the fewest flips you can get. This is the law of large numbers working against you, and the question is testing whether you know which side of it you are on.
With 10 flips the bars at 6 heads and above hold 37.7% of the probability, but with 100 flips the bars at 60 heads and above hold only 2.8%, because the share of heads tightens around a half as the flips increase. How do you put a number on it without a table?
Count the small case exactly: 6 or more heads in 10 flips means adding the ways to get 6, 7, 8, 9 and 10 heads, which are 210, 120, 45, 10 and 1, a total of 386 out of 1,024, so 37.7%. For 100 flips use the normal approximation. The standard deviation of the share of heads is 0.5 over the square root of n: 15.8% for 10 flips, 5% for 100. The 60% line is 0.6 standard deviations out in the first case and 2.0 in the second, and two standard deviations in one tail is about 2.3%. The exact binomial answer is 2.8%; the approximation gets you to the right decision in one breath.
The relationshipsigma of X bar the standard deviation of the share of heads n the number of flips z how many standard deviations the 60% line sits from the mean of 50% What it says in wordsThe 60% line gets further from the centre, measured in standard deviations, as the flips increase, so it becomes rarer to cross.What is the interviewer listening for after the answer?
Say the general rule and then the exception. The rule: whenever a payout needs the sample to look unlike the population, choose the smallest sample. The exception: if the payout were for landing between 40% and 60%, you would want the most flips, for exactly the same reason. Read the sign of the bet before you choose the sample size: a bet on luck wants few trials and a bet on the average wants many. If the interviewer changes the wording to more than 60%, 7 or more heads in 10 is 17.2%, still far above the 100-flip figure. And if the choice is between 100 and 1,000 flips, 600 or more heads happens about 1.36e-10 of the time, which is as close to never as a desk needs.
Where candidates lose it
The fast wrong answer is 100 flips, because more flips feel like more chances. They are more chances for the average to assert itself, not for luck. Candidates who say it have the law of large numbers backwards, and the interviewer hears it immediately.
The second loss is saying 10 without a number. Have the 386 out of 1,024 ready, then the standard deviation argument for 100, so the answer sounds reasoned rather than remembered.
What the interviewer asks next
- I pay you if the share of heads is between 45% and 55%. Now which do you want?
- What if the coin has a 60% bias towards heads? Does the answer flip?
- Roughly how many flips make the 60% line a three standard deviation event?
Asked at Hudson River Trading, Prop Trading, New York, 2020 (Wall Street Oasis):
Questions on EV for coin tosses, law of large numbers, Bayes theorem
037X and Y are independent and uniform on 0 to 1. What is the density of X + Y, and what is the probability that X + Y is less than 1.5?CitadelChicago · 2025Citadel SecuritiesChicago · 2025
Try it first
What is P(X + Y < 1.5)?
Show the worked solution
The density is a triangle: f(s) = s for s between 0 and 1, and 2 - s for s between 1 and 2; P(X + Y < 1.5) = 7/8. The sum lands in a thin band near s only if X and Y fit inside the unit square along the line X + Y = s, and the length of that line segment grows until s = 1 and then shrinks. The sum exceeds 1.5 only in the top-right corner of the square, a triangle with legs 1/2 and area 1/8, so the probability of less than 1.5 is 0.875.
Why is the sum of two flat distributions a triangle?
Roll two ordinary dice: a total of 7 can be made six ways, a total of 2 or 12 only one way, so the totals pile up in the middle even though each die is flat. The sum of two independent uniforms is the continuous version: the density at s is the length of the diagonal line X + Y = s inside the unit square, which grows from 0 up to s = 1 and shrinks back to 0 at s = 2. Formally the density is the convolution, the integral over x of f(x) times g(s - x), and with both densities equal to 1 on the unit interval the integrand is 1 exactly where both x and s - x lie between 0 and 1. That interval has length s for s below 1 and length 2 - s above, which is the triangle.
The relationshipf X+Y the density of the sum the indicator equals 1 where both x and s - x lie in the unit interval, which is where both densities are non-zero the integral from 1.5 to 2 the area of the small triangle in the tail, 1/8 What it says in wordsConvolve two flat densities and the overlap of the two intervals gives a triangle; the tail beyond 1.5 is one eighth of it.On the unit square the event X + Y > 1.5 is the corner triangle with legs one half and area one eighth, and on the density plot the same one eighth is the tail of the triangle beyond s = 1.5, so P(X + Y < 1.5) is seven eighths. How do you get 7/8 in your head?
Use the square, not the integral. The pair (X, Y) is a uniformly random point in the unit square, so any probability is an area. X + Y > 1.5 is the region above the line from (0.5, 1) to (1, 0.5), a right triangle with both legs equal to 1/2, whose area is half of a quarter, which is 1/8. Everything else is 7/8. The same picture answers the mirror question at once: X + Y < 0.5 is the bottom-left corner, also 1/8. And it checks the density: integrating 2 - s from 1.5 to 2 gives (0.5)^2 / 2 = 1/8, the same number from the other direction, which is what you say to show the two views agree.
What does the interviewer ask next, and why?
The follow-ups test whether you understood the mechanism or memorised the triangle. Add a third uniform and the density becomes three parabolic pieces, smooth at the joins, with P(X + Y + Z < 1) = 1/6, the volume of the corner tetrahedron, and in general the sum of n uniforms heads towards a normal with mean n/2 and variance n/12, which is the central limit theorem arriving. The mean of X + Y is 1 and the variance is 1/12 + 1/12 = 1/6, by independence. Each extra convolution smooths the density by one degree: flat to triangular to parabolic, which is the general fact that sums of independent variables are smoother than their parts. The limitation worth saying: this all relies on independence; with X and Y correlated, the point is no longer uniform over the square and the areas no longer give probabilities.
Where candidates lose it
The common loss is treating the sum as uniform on 0 to 2 and answering 0.75. The sum is a triangle, and the mass in the top quarter of its range is only 1/8, not 1/4.
The second is setting up the convolution integral correctly and then mishandling the limits, integrating over the wrong interval of x. Draw the square; the limits are the overlap of two unit intervals, and the picture makes them obvious.
What the interviewer asks next
- What is P(X + Y + Z < 1) for three independent uniforms?
- What is the density of X - Y, and of max(X, Y)?
- What are the mean and variance of X + Y, and how do you get them without the density?
- X and Y are now correlated with rho = 0.5. Can you still find P(X + Y < 1.5) by area?
Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):
He was asking some questions about the probability, especially on the convolution.
Asked at Citadel Securities, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):He was asking some questions about the probability, especially on the convolution.
054Give me two random variables that are uncorrelated but clearly dependent, and show me that the correlation is zero.Two SigmaNew York · 2025Tower Research CapitalNew York · 2014
Try it first
X is a standard normal and Y = X squared. What is the correlation of X and Y?
Show the worked solution
Take X standard normal and Y = X squared. Y is completely determined by X, the strongest dependence there is, yet Cov(X, Y) = E[X cubed] - E[X] E[X squared] = 0 - 0 x 1 = 0, because a symmetric distribution has zero third moment. A discrete version is X uniform on minus 1, 0 and 1 with Y = X squared, where E[XY] = (minus 1 + 0 + 1)/3 = 0 and E[X] = 0.
What does correlation actually measure?
Think of a shop whose sales rise when it is very cold and when it is very hot, and sag in mild weather. Plot sales against temperature and you get a U. Ask a straight line to summarise the U and it comes out flat: the hot days pull it up on the right exactly as much as the cold days pull it up on the left. Correlation is the slope of the best straight line through the cloud, scaled to lie between minus 1 and 1, so it can only detect the part of a relationship that runs in one direction. A relationship that turns around is invisible to it.
Forty-one points on the parabola Y = X squared, with X spread like a standard normal, have a best-fit line that is exactly flat at height 0.97, so their correlation is zero while every Y is fixed by its X. How do you prove the zero in two lines?
Write the covariance as E[XY] minus E[X] E[Y]. With Y = X squared, E[XY] is E[X cubed], and every distribution symmetric about zero has E[X cubed] = 0, because each positive value x cubed is cancelled by the equally likely minus x cubed. E[X] is also zero, so both terms vanish and the correlation, which is the covariance divided by two positive standard deviations, is zero. The same proof works for any even function of a symmetric X, such as the absolute value, and for any symmetric X, not just the normal.
The relationshipE[X^3] the third moment, zero for any distribution symmetric about zero E[X] the mean of X, zero E[X^2] the variance of a standard normal, one What it says in wordsThe covariance of a symmetric variable with its own square is zero, so correlation is zero even though the square is a function of the variable.Why does a derivatives desk care about this distinction?
Because a hedge built on correlation only neutralises straight-line exposure. A stock's return and its squared return are close to uncorrelated, yet the squared return is exactly what a long option position pays for, so a delta hedge that zeroes the linear exposure leaves the gamma exposure untouched. The limitation runs the other way too: independence does imply zero correlation, so a measured correlation near zero rules nothing out, and a desk that reads zero correlation as no relationship will be surprised in both tails.
Where candidates lose it
Candidates reach for a definition, saying independence implies zero correlation but not the reverse, and then cannot produce an example. The interviewer wants the example first: X and X squared, said inside ten seconds.
The second loss is giving the example without the proof. Write the covariance as E[X cubed] minus E[X] E[X squared] and say why the third moment of a symmetric variable is zero.
What the interviewer asks next
- Give a second example where neither variable is a function of the other.
- X is uniform on 0 to 1 rather than symmetric. Is X still uncorrelated with X squared?
- What single number would catch the dependence between X and X squared that correlation misses?
Asked at Two Sigma, Generalist, New York, 2025 (Wall Street Oasis):
Come up with two uncorrelated but dependent variables. Lots of questions regarding regression
Asked at Tower Research Capital, Quantitative Research, New York, 2014 (Wall Street Oasis):He also asked a question about independence and correlation. I did not respond fast enough
063The gaps between trades in an illiquid option are exponentially distributed with an unknown rate. You observe gaps of 2, 3 and 7 seconds. Derive the maximum likelihood estimate of the rate, and tell me why you would not trust it much.AQR Capital ManagementTown of Greenwich · 2022
Try it first
Gaps of 2, 3 and 7 seconds. What is the maximum likelihood rate?
Show the worked solution
The maximum likelihood estimate is 3/12 = 0.25 trades per second, the count divided by the total time, and with three observations it is both noisy and biased high. The likelihood of gaps t_1, t_2, t_3 is lambda cubed times exp(minus lambda times 12). Its log, 3 ln lambda minus 12 lambda, has derivative 3/lambda minus 12, which is zero at lambda = 0.25. On average this estimator reads 1.5 times the true rate when n = 3.
What does the likelihood actually say, in words?
A shopkeeper who saw customers arrive 2, 3 and 7 minutes apart would say roughly one every four minutes, and the likelihood is the formal version of that. For each candidate rate, the likelihood is how probable the observed gaps would be under that rate; the maximum likelihood estimate is the rate that makes what you saw least surprising. With exponential gaps the density of each gap t is lambda times exp(minus lambda t), so three independent gaps multiply to lambda cubed times exp(minus lambda times their sum). The sum, 12 seconds, is all the data you need; the individual values 2, 3 and 7 drop out.
The likelihood of the gaps 2, 3 and 7 seconds, plotted against the trade rate, peaks at 3 divided by 12 = 0.25 trades per second, but stays above half its peak from about 0.11 to 0.47, so three observations pin the rate down only loosely, and the bias-corrected estimate of 0.167 sits well to the left. How do you derive the peak in three lines?
Take logs first, because a product of exponentials becomes a sum. The log likelihood is n ln lambda minus lambda times the sum of the gaps, its derivative is n over lambda minus the sum, and setting that to zero gives lambda equal to n over the sum, the number of events divided by the time they took. Here that is 3 over 12. The second derivative, minus n over lambda squared, is negative everywhere, so the stationary point is a maximum. The same derivation gives the familiar result that the maximum likelihood rate is one over the sample mean gap, 1 over 4 seconds.
The relationshipl(lambda) the log likelihood of the observed gaps as a function of the rate n the number of gaps observed, three sum of t_i the total time covered by the gaps, 12 seconds What it says in wordsThe maximum likelihood rate is the number of trades divided by the total time between them, which is one over the average gap.Why would you not trust 0.25, and what would you say instead?
Three reasons, in the order a desk cares about them. The estimate is built on three numbers, so the likelihood hill is wide and any rate from roughly 0.11 to 0.47 fits almost as well; one over a sample mean is biased upward, with an expected value of n over n minus 1 times the true rate, 1.5 times here, so the unbiased version is (n - 1) over the sum, 0.167; and nothing in three gaps tests the exponential assumption itself. Real trade arrivals cluster, with bursts after news and dead stretches overnight, so a single constant rate is a model you chose, not a fact you found. The honest statement is a rate near 0.25 with a wide interval and a flag that the model may be wrong.
Where candidates lose it
The common loss is a formula without a derivation: candidates say one over the mean and stop. The interviewer asked you to derive it, so write the likelihood, take the log, differentiate and check the sign of the second derivative.
The second loss is answering the trust question with only the sample size. Mention the bias, give the corrected estimate, and question the exponential assumption, because that is the part a desk actually gets wrong.
What the interviewer asks next
- What is the maximum likelihood estimate of the mean gap, and is it biased?
- Give an approximate 95% interval for the rate from these three gaps.
- How would you test whether trade gaps are really exponential?
- The fourth gap is 60 seconds. What happens to the estimate, and does that worry you?
Asked at AQR Capital Management, Trading, Town of Greenwich, 2022 (Wall Street Oasis):
Derive the mle for some given distribution. Explain linear regression intuitively and derive the ols estimate.
088Rs 1 was invested in a broad equity index 30 years ago. If its log returns average 9% a year with 16% volatility, give a 95% range for what the Rs 1 is worth today, and explain why the range is lopsided.Old Mission CapitalChicago · 2025
Try it first
Which quantity do you build the 95% range on first?
Show the worked solution
Roughly Rs 2.7 to Rs 82.9, around a median of Rs 14.9. Thirty years of log returns sum to a mean of 2.7 with a standard deviation of 0.16 x the square root of 30, about 0.876. The 95% band in logs is 2.7 plus or minus 1.72, from 0.98 to 4.42. Exponentiating gives Rs 2.7 to Rs 82.9. It is lopsided because e to the power stretches gains and compresses losses: 12.2 below the median, 68.0 above.
Why work in logs and not in rupees?
A savings balance grows by multiplying: up 10% then down 10% leaves you at 0.99, not 1. Multiplications are awkward to average, but their logarithms add, and sums of many independent pieces tend towards a bell curve. Log returns add across years, so their total over 30 years has a mean that grows with time and a spread that grows with the square root of time, and the 95% band is symmetric in logs. Here the mean is 30 x 0.09 = 2.7, the standard deviation is 0.16 x 5.477 = 0.876, and 1.96 standard deviations is 1.72. The log of the final value lies between 0.98 and 4.42 with 95% confidence.
Thirty years of log returns averaging 9% with 16% volatility give a final value of Rs 1 that peaks near Rs 6.9, has a median of Rs 14.9 and a mean of Rs 21.8, and lies between Rs 2.7 and Rs 82.9 with 95% confidence, a band that is symmetric in logs but runs 12.2 below the median and 68.0 above it in rupees. Why does the rupee range lean so far to the right?
Converting back means raising e to each end. e to the 0.98 is 2.67 and e to the 4.42 is 82.9, around a median of e to the 2.7, 14.9. The same 1.72 step in logs is a factor of 5.6 either way, and dividing by 5.6 moves you 12.2 rupees while multiplying by 5.6 moves you 68.0. That is why the distribution of wealth has a long right tail: the typical outcome, the median, sits well below the average, Rs 21.8, which is pulled up by the rare decades that compound very well.
The relationshipW_30 the value after 30 years of Rs 1 invested at the start mu the average yearly log return, 9% sigma the yearly volatility of log returns, 16% 1.96 the number of standard deviations that covers 95% of a normal What it says in wordsBuild a symmetric 95% band for the total log return, then raise e to each end to get a lopsided band in rupees.Two cautions the interviewer will reward. First, the 9% and 16% here are assumptions for the exercise; confirm the index's actual history before quoting any real figure, and note that a 30-year average return is itself estimated with a wide error. Second, the band assumes yearly log returns are independent and normal. Real markets have fat tails and runs of bad years, which widen the lower end in particular. The model gets the shape right, a long right tail with the median below the mean, even if the exact ends are only approximate.
Where candidates lose it
The common loss is building the band in rupees with a symmetric plus or minus, which gives a lower bound below zero or a band centred on the wrong number. Wealth cannot go below zero, and any method that says it can has skipped the logs.
The second loss is scaling the volatility by 30 instead of the square root of 30. That gives a standard deviation of 4.8 in logs, a band from almost nothing to tens of thousands, and tells the interviewer you have not met the square-root-of-time rule.
What the interviewer asks next
- What is the probability the Rs 1 is worth less than Rs 1 today?
- Why is the mean of the final value higher than the median?
- If volatility were 25% instead of 16%, what happens to the median and to the mean?
Asked at Old Mission Capital, Prop Trading, Chicago, 2025 (Wall Street Oasis):
Confidence interval of portfolio value if you invested $1 in S&P 500 30 years ago
094A broad equity index has returned an average of 11% a year over the past 30 years, with annual volatility of 16%. Give a 95% confidence interval for its true expected annual return.Old Mission CapitalChicago · 2025
Try it first
How wide is the 95% interval for the true expected return, either side of 11%?
Show the worked solution
About 5.3% to 16.7%. The standard error of a 30-year average is the yearly volatility over the square root of 30: 16/5.48 = 2.92 points. A 95% interval is 1.96 standard errors either side, about 5.7 points, so 11% plus or minus 5.7%. Thirty years of data leave the true expected return anywhere from modest to spectacular, and narrowing it to plus or minus one point would take nearly a thousand years.
Why is the interval so wide after 30 years?
Weigh yourself on a scale that is off by up to two kilos each time. One reading tells you little; the average of four readings is better, but only twice as good, not four times, because errors cancel in proportion to the square root of the count. The uncertainty in an average shrinks with the square root of the number of observations, so 30 noisy years cut a 16-point yearly spread only to about 2.9 points. That is the standard error. Multiply by 1.96 for 95% and you get roughly 5.7 points either side of 11%.
The 95% interval for the true expected return narrows only with the square root of the years of data, so with an 11% average and 16% volatility it runs from about -3% to 25% after 5 years, from 5.3% to 16.7% after 30 years, and would need about 983 years to shrink to plus or minus one point. What does the interviewer do with the answer?
Usually they push on what it means. A 30-year history cannot tell a 6% market from a 16% market with any confidence, which is why long-run return assumptions are judgements, not measurements. To get the interval down to plus or minus 2 points needs (1.96 x 16/2)^2 = 246 years, and to plus or minus 1 point, 983 years. Volatility, by contrast, is estimated far better from the same data, because daily or monthly returns give thousands of observations of spread, while there is only one 30-year path for the mean. Sampling more often does not help the mean: the average depends only on the first and last levels.
The relationshipr bar the sample average yearly return, 11% sigma the volatility of one year's return, 16% n the number of years, 30 1.96 the multiplier for a 95% normal interval What it says in wordsThe average of 30 noisy years is itself noisy, with an error of the volatility over the square root of 30, so the interval is about 5.7 points either side.Say the assumptions, because the interviewer will. The interval treats the 30 yearly returns as independent draws from one unchanging distribution, which real markets are not; regime shifts make the true uncertainty wider. The 11% and 16% are the inputs given for this exercise, so confirm any real index's figures before using them. And an arithmetic average of yearly returns is not the compound growth rate; with 16% volatility the compound rate is roughly 1.3 points lower, which is a separate question worth flagging rather than answering.
Where candidates lose it
The common loss is using 16% as the error, giving 11% plus or minus 31 points, which confuses the spread of one year with the uncertainty of an average. The opposite loss is a tiny interval from forgetting that only 30 independent years exist.
The second loss is reaching for daily data to shrink the interval. More frequent sampling sharpens the volatility estimate, not the mean: the total return over 30 years depends only on where the index started and where it ended.
What the interviewer asks next
- How many years of data would you need to tell a 6% expected return from an 8% one?
- Why does more frequent data help estimate volatility but not the expected return?
- If returns were autocorrelated, would the interval be wider or narrower?
Asked at Old Mission Capital, Prop Trading, Chicago, 2025 (Wall Street Oasis):
Confidence interval on S&P 500 return past 30 years
099Daily returns come from a normal with 1% volatility on 90% of days and a normal with 3% volatility on the other 10%, both with mean zero. What are the overall volatility and kurtosis, and what does this do to out-of-the-money option prices?Two SigmaNew York · 2025
Try it first
What is the overall daily volatility?
Show the worked solution
Volatility 1.34% a day and kurtosis 8.33, against 3 for a normal, which makes far out-of-the-money options worth much more than a single-volatility model says. Variance is 0.9 x 1 + 0.1 x 9 = 1.8. The fourth moment is 3 x (0.9 x 1 + 0.1 x 81) = 27, so kurtosis is 27/1.8^2 = 8.33. With the same volatility, a 4% move is 6.4 times as likely as the normal says, and a one-day put 4% out of the money is worth about 23 times as much.
How do the moments of a mixture combine?
Picture a road with quiet days and occasional storms. Most days the traffic varies a little; on storm days it varies a lot. Averaged over a year, the variability you see is dominated by the storm days, out of proportion to how rare they are. For a mixture with mean zero, each even moment is the weighted average of the components' moments: variances add up by weight, and fourth moments, which grow with the fourth power of volatility, are dominated by the rare wild days. Variance: 0.9 x 1 + 0.1 x 9 = 1.8, a volatility of 1.34%. Fourth moment: each normal's is 3 sigma^4, so 3 x (0.9 x 1 + 0.1 x 81) = 27. The 3% days are a tenth of the time but supply 90% of the fourth moment.
Against a normal with the same 1.34% daily volatility, the mixture of 1% and 3% days is taller in the middle and lies above the normal in the tails beyond about 3.7%, so its kurtosis is 8.33 instead of 3 and a move beyond 4% either way is 6.4 times as likely. What does the fat tail do to option prices?
Kurtosis is the fourth moment over the square of the variance: 27/3.24 = 8.33. A single normal fitted to the same data has the same volatility, 1.34%, but kurtosis 3, so it puts far too little weight on big moves. An option far out of the money pays only on a big move, so its value is driven by the tail, and a mixture that keeps the volatility but fattens the tail makes it worth many times what the matched normal says. A one-day put 4% out of the money is worth 0.0127% of spot under the mixture against 0.00055% under the normal, about 23 times as much. To match the mixture's price, a single-volatility model needs 1.91% for that put but only 1.20% at the money: the smile.
The relationshipsigma^2 the variance of the daily return, in square percentage points kappa the kurtosis, 3 for any single normal 3 sigma^4 the fourth moment of a normal with mean zero What it says in wordsVariances average by weight, fourth moments average by weight, and the ratio comes out near 8.3 because the rare 3% days dominate the fourth power.Say what it means on a desk. This mixture is the simplest model of regime-switching volatility, and it reproduces two things a single normal cannot: a peak that is too tall and tails that are too fat, the shape seen in daily returns of most liquid assets. It is also why at-the-money options can look rich and far wings cheap if you price both off one historical volatility. The limitation: the mixture here is symmetric and draws each day independently. Real markets cluster their wild days and fall harder than they rise, which adds skew to the smile, not just curvature.
Where candidates lose it
The common loss is averaging the volatilities, 0.9 x 1 + 0.1 x 3 = 1.2%, and then getting the kurtosis wrong as a result. Volatilities never average in a mixture; variances do.
The second loss is computing the moments correctly and stopping, without connecting kurtosis to option prices. The question asks what the fat tail does to the wings, and the answer is a smile: more implied volatility the further out of the money you go.
What the interviewer asks next
- What is the kurtosis if the wild days are 5% volatility but only 4% of days?
- Which is more mispriced by a single-volatility model, an at-the-money option or a far out-of-the-money one, and in which direction?
- How would you add skew to this model?
Asked at Two Sigma, Quantitative Research, New York, 2025 (Wall Street Oasis):
They asked a couple questions involving Mixture Gaussians (e.g., probability density and moments).
