Hedge Funds puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 38
- Topics
- 14
- Hard
- 30
025Four observations, 3.1, 7.4, 5.2 and 9.0, come from a uniform distribution on 0 to theta. What is the maximum likelihood estimate of theta, is it biased, and how would you correct it?AQR Capital ManagementTown of Greenwich · 2022
Try it first
Which statement is right?
Show the worked solution
The MLE is 9.0, the largest observation; it is biased low, and the unbiased correction is 5/4 x 9.0 = 11.25. The likelihood is 1 over theta to the fourth for any theta of at least 9.0 and zero below it, so it peaks at the sample maximum. But the maximum of n draws averages n/(n + 1) of theta, here 4/5, so scaling by (n + 1)/n removes the bias.
Why is the MLE the largest observation?
A friend draws four raffle tickets numbered from 1 up to some unknown top number, and the highest you see is 90. The top number is at least 90; guessing higher only spreads your belief over tickets nobody drew. The likelihood, 1 over theta to the n, is zero for any theta below the largest observation and falls as theta rises above it, so it is maximised exactly at the sample maximum, 9.0. This is a case where you do not differentiate: the maximum sits on a boundary, not where a slope is zero.
Four uniform draws on 0 to theta cut it into five gaps of equal expected size, so the largest draw averages four fifths of theta; the MLE of 9.0 therefore sits below theta, and scaling by 5/4 gives the unbiased 11.25, against 12.35 from doubling the sample mean. The relationshipL(\theta) the likelihood of the four observations n the number of observations, 4 \max the largest observation, 9.0 What it says in wordsThe likelihood peaks at the largest observation, which on average falls short of theta by a factor n/(n + 1).Why is it biased, and by how much?
The largest draw can never be above theta and is almost always below it. Four points dropped at random on 0 to theta cut it into five gaps of equal expected length, so the largest point sits on average four fifths of the way up, and the MLE underestimates theta by a fifth on average. Multiplying by 5/4 fixes it: 9.0 becomes 11.25. The bias shrinks as n grows, since n/(n + 1) tends to 1, but with four points it is large.
How does it compare with the obvious alternative?
The method of moments doubles the sample mean, since a uniform on 0 to theta averages theta/2: the mean here is 6.175, giving 12.35. Both 11.25 and 12.35 are unbiased, but the corrected maximum has a much smaller variance, theta squared over n(n + 2) against theta squared over 3n, because the largest draw carries the most information about the top of the range. With four points that is theta squared over 24 against theta squared over 12: half the variance.
Where candidates lose it
Candidates differentiate the log-likelihood, get minus n over theta, set it to zero and find no solution. The likelihood only falls on the allowed range, so the maximum sits at the boundary, the largest observation; say that before reaching for calculus.
The second miss is calling the MLE unbiased because maximum likelihood estimates are often well behaved. Here it is biased low by construction, and the interviewer expects the (n + 1)/n correction.
What the interviewer asks next
- What is the variance of the corrected estimator with four observations?
- What is the MLE if the distribution is uniform on theta to 2 theta?
- Derive the expected value of the maximum of n uniform draws.
Asked at AQR Capital Management, Trading, Town of Greenwich, 2022 (Wall Street Oasis):
Derive the mle for some given distribution. Explain linear regression intuitively and derive the ols estimate.
038X and Y are independent random variables with the same variance. What is the correlation between X and X + Y?Squarepoint CapitalMontreal · 2026
Try it first
Pick one:
Show the worked solution
1 over root 2, about 0.71. The covariance of X with X + Y is Var(X) plus Cov(X, Y), which is sigma squared plus zero. The standard deviation of X + Y is root 2 times sigma because the variances add. So the correlation is sigma squared over (sigma x root 2 sigma), which is 1/root 2. X explains half the variance of the sum, and the correlation is the square root of that half.
What is the fastest way to set it up?
A two-member team's score is the sum of both players' scores. If the players are equally good and play independently, knowing one player's score tells you something about the team total, but only half the story. Split the covariance: Cov(X, X + Y) = Cov(X, X) + Cov(X, Y) = sigma squared + 0. The variance of the sum is sigma squared + sigma squared = 2 sigma squared, because independent variances add. Correlation is covariance over the product of standard deviations: sigma squared over (sigma x root 2 sigma) = 1/root 2.
Drawn as arrows, independent X and Y sit at right angles and their sum lies at 45 degrees to X, so the correlation is cos 45, about 0.707; equivalently, X supplies half of the variance of X + Y, and the correlation is the square root of one half. Why is the answer not 0.5?
Because 0.5 is the R squaredThe share of one variable variance explained by another; for a simple regression it is the correlation squared., not the correlation. X explains exactly half of the variance of X + Y, and correlation is the square root of the share of variance explained, so it is root 0.5, about 0.707. The geometric picture makes it stick: treat independent variables as arrows at right angles, and correlation as the cosine of the angle between arrows. X + Y sits at 45 degrees to X, and cos 45 is 0.707.
Give the general version to show you own it. If Y has variance k times X's, the correlation is 1/root(1 + k): the more noise you add, the lower it falls. That is the logic behind a noisy signal: a forecast that is half signal and half independent noise, by variance, correlates about 0.71 with the signal, not 0.5.
Where candidates lose it
The common loss is answering 0.5 because X is half of the sum. That is the share of variance, and correlation is its square root.
The other loss is saying zero because X and Y are independent. The sum contains X, so it cannot be independent of X. Split the covariance in one line and the answer falls out.
What the interviewer asks next
- What is the correlation between X + Y and X - Y?
- Y has four times the variance of X. What is corr(X, X + Y) now?
- What is the correlation between the sum of the first 10 and the sum of the first 20 of a series of independent returns?
Asked at Squarepoint Capital, Desk Quant Analyst Interview, Montreal, 2026 (Wall Street Oasis):
There were also 3-4 basic math/stats questions about mean, covariance, correlation, etc.
048A strategy's true annual Sharpe ratio is 1.0. How many years of monthly returns do you need before its average return shows a t-statistic of 2? What if the true Sharpe is 0.5?Viking Global InvestorsNew York · 2014
Try it first
Years needed for a Sharpe of 0.5:
Show the worked solution
About 4 years for a Sharpe of 1.0 and about 16 years for a Sharpe of 0.5. The t-statistic of a mean return is the mean over its standard error, which works out to the annual Sharpe ratio times the square root of the number of years, whatever the data frequency. Setting Sharpe x root(years) = 2 gives years = (2 / Sharpe) squared: 4 for 1.0, 16 for 0.5 and just 1 for 2.0.
Why does the t-statistic grow with the square root of time?
A coin that lands heads 55% of the time looks fair after 20 tosses; you need hundreds before the bias shows through the noise. The average return grows in proportion to time, but the noise around it grows only with the square root of time, so the signal-to-noise ratio, the t-statistic, grows with root time. With monthly data, the t-statistic is the monthly Sharpe times root(12 x years), and the monthly Sharpe is the annual Sharpe divided by root 12, so the twelves cancel: t = annual Sharpe x root(years).
Because the t-statistic equals the Sharpe ratio times the square root of years, a Sharpe of 2.0 clears t = 2 after 1 year, a Sharpe of 1.0 after 4 years and a Sharpe of 0.5 only after 16 years. Why does monthly data not shorten the wait?
More frequent data gives more observations but each is noisier relative to its mean. Sampling the same years more often does not add information about the mean return; only more years do. This is why a {term('t-statistic', 'An estimate divided by its standard error; a value around 2 is the usual threshold for saying an effect is unlikely to be pure noise.')} on the average return depends on the span of the data, not the number of rows. Frequency helps you estimate volatility, not the mean.
The relationshipSR the true annual Sharpe ratio Y years of data 2 the target t-statistic What it says in wordsThe years needed to prove a strategy grow with the inverse square of its Sharpe ratio.Say the practical point. Most real strategies have Sharpe ratios well below 1, so their track records are too short to separate skill from luck with any confidence. An allocator looking at a three-year record with a Sharpe of 0.8 sees a t-statistic of about 1.4. The limitation of the rule: it assumes returns are independent and stable over the whole sample, and fat tails or regime changes make the real uncertainty larger.
Where candidates lose it
The common loss is thinking monthly data gives twelve times the evidence, which leads to answers like four months. The twelve cancels, because the monthly Sharpe is smaller by root 12.
The second loss is saying a Sharpe of 0.5 needs twice as long as 1.0. The dependence is on the square: half the Sharpe, four times the data.
What the interviewer asks next
- How many years for a Sharpe of 0.3?
- You test 20 strategies and pick the best one with t = 2.2. How much do you trust it?
- Would daily data change the answer for estimating the Sharpe ratio itself rather than the mean?
Asked at Viking Global Investors, Quantitative Research, New York, 2014 (Wall Street Oasis):
how to reject a hypothesis test, what's your structure of your code, what's the sample size
063The sample variance computed with n minus 1 in the denominator is an unbiased estimator of the population variance. Is its square root an unbiased estimator of the standard deviation?Squarepoint CapitalLondon · 2026
Try it first
Is the square root of the unbiased sample variance unbiased for the standard deviation?
Show the worked solution
No. The square root of the unbiased variance underestimates the standard deviation on average. The square root is concave, so by Jensen's inequality the average of the square roots is below the square root of the average. For normal data with two observations the estimate averages about 0.80 sigma; the bias shrinks as the sample grows, to about 6% at five observations and under 1% at thirty.
Why does taking a square root break unbiasedness?
Two square rooms have floor areas of 4 and 16 square metres, so their sides are 2 and 4 metres. Average the areas, 10, and take the root: 3.16 metres. Average the sides instead: 3 metres. Averaging and then taking a square root gives a bigger answer than taking square roots and then averaging, because the square root bends downwards. The sample variance is right on average, so the average of its square roots must fall short of the true standard deviation.
Two equally likely variance estimates of 0.04 and 1.96 average to the true variance of 1.0, but their square roots, 0.2 and 1.4, average only 0.8, below the true standard deviation of 1.0, because the square-root curve bends downwards. How big is the bias?
It depends on the sample size and on the distribution. For normal data the expected sample standard deviation is c4 times sigma, with c4 about 0.80 at n = 2, 0.94 at n = 5, 0.97 at n = 10 and 0.99 at n = 30. At n = 2 you can check it directly: the sample standard deviation is the gap between the two draws divided by the square root of 2, and the average gap between two normal draws is 2 sigma over the square root of pi, which leaves the square root of 2/pi, about 0.798.
The relationships the square root of the unbiased sample variance sigma the true standard deviation c4(n) the correction factor for normal data, below 1 for every n What it says in wordsThe average sample standard deviation is a fixed fraction of the true one, and that fraction is below one.Does it matter in practice?
Sometimes. With a year of daily returns the bias is a rounding error; with a handful of monthly returns for a new fund it is not. A manager with five monthly returns has a volatility estimate that averages about 6% too low under normality, which flatters a Sharpe ratioAverage excess return divided by the standard deviation of returns, a measure of return per unit of risk. before anyone has looked at fat tails. Dividing by c4 removes the bias for normal data, but the fix depends on the distribution, so name the assumption. And unbiased is not the same as most accurate.
Where candidates lose it
The trap is assuming unbiasedness carries through any function of an estimate. It carries through straight-line transformations only; the square root is curved, so the property is lost.
The second loss is saying it is biased without the direction or the size. Say biased low, give the Jensen reason in one sentence, and quote about 0.80 at two observations, shrinking towards 1 as the sample grows.
What the interviewer asks next
- Is the square of an unbiased estimator of the standard deviation unbiased for the variance?
- Why does the sample variance divide by n minus 1 rather than n?
- Which estimator of sigma has the lowest mean squared error for normal data?
Asked at Squarepoint Capital, Quantitative Research, London, 2026 (Wall Street Oasis):
Is the square root of the unbiased estimator for sample variance unbiased for standard deviation?
098With orthonormal regressors, ordinary least squares gives coefficients of 0.6 and 0.15. What do ridge and lasso with a penalty of 0.2 give for each, and why does only lasso set a coefficient to zero?CitadelLondon · 2026
Try it first
Using half the residual sum of squares plus the penalty, what do the two methods give?
Show the worked solution
Ridge gives 0.5 and 0.125; lasso gives 0.4 and exactly 0. With orthonormal regressors each coefficient is shrunk on its own. Ridge divides each by 1 + 0.2 = 1.2, so it scales both down and never reaches zero. Lasso subtracts 0.2 from each size and stops at zero, so the small coefficient, 0.15, is removed. This uses the scaling of half the residual sum of squares plus 0.2 times the penalty.
Why do orthonormal regressors make this a one-line problem?
When the regressors are uncorrelated and scaled to unit length, the fit for each coefficient does not depend on the others, so the penalised problem splits into separate one-variable problems. Think of adjusting the volume on two speakers that are not wired together: turning one down does not change the other. For each coefficient you minimise half of (beta minus b) squared plus the penalty, where b is its least squares value, and the answer depends only on that one number. Ridge's penalty is half of beta squared times 0.2; lasso's is the size of beta times 0.2.
Ridge scales every least squares coefficient by 1/1.2, so 0.6 becomes 0.5 and 0.15 becomes 0.125, while lasso subtracts 0.2 and stops at zero, so 0.6 becomes 0.4 and 0.15, inside the dead zone below 0.2, becomes exactly 0. What does each penalty do to a coefficient?
Ridge's squared penalty pulls hard on big coefficients and gently on small ones. Setting the slope to zero gives beta x (1 + 0.2) = b, so ridge divides by 1.2: 0.6 becomes 0.5 and 0.15 becomes 0.125, shrunk but never zero. Lasso's absolute-value penalty pulls with the same force, 0.2, whatever the size. Its answer is the sign of b times the larger of (the size of b minus 0.2) and zero: 0.6 becomes 0.4, and 0.15, weaker than the pull of 0.2, lands exactly on 0.
The relationshipb the least squares coefficient, 0.6 or 0.15 lambda the penalty weight, 0.2 What it says in wordsRidge divides every coefficient by the same factor; lasso takes the same amount off every coefficient and never goes past zero.Why does only lasso select variables?
At zero, the squared penalty is flat: its slope is zero, so any small coefficient still earns its place by improving the fit a little. The absolute-value penalty has a corner at zero with a slope of 0.2 on each side, so a coefficient stays at zero unless the fit improves by more than 0.2 per unit, and 0.15 does not. That is why lasso gives sparseHaving many coefficients exactly equal to zero, so the model uses only a few of the available variables. models and ridge does not. The limitation to state: with correlated regressors lasso tends to keep one of a group arbitrarily and drop the rest, which is why desks often blend the two penalties in an elastic net.
Where candidates lose it
The common slip is to swap the two, saying ridge sets small coefficients to zero because it penalises harder. Ridge's penalty is heavy on big coefficients and almost nothing on small ones, which is exactly why it never zeroes them.
The second loss is getting the lasso numbers off by a factor of two. Written as the full residual sum of squares plus 0.2 times the absolute values, without the half, the threshold is 0.1, giving 0.5 and 0.05. State your scaling before you give numbers.
What the interviewer asks next
- At what penalty does lasso set the 0.6 coefficient to zero as well?
- The two regressors now have a correlation of 0.9. How do ridge and lasso behave differently?
- How would you choose the penalty in practice without fitting it to noise?
Asked at Citadel, Quantitative Research, London, 2026 (Wall Street Oasis):
very detailed and difficult questions about regularisation ridge and lasso
