Hedge Funds puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 38
- Topics
- 14
- Hard
- 30
025Four observations, 3.1, 7.4, 5.2 and 9.0, come from a uniform distribution on 0 to theta. What is the maximum likelihood estimate of theta, is it biased, and how would you correct it?AQR Capital ManagementTown of Greenwich · 2022
Try it first
Which statement is right?
Show the worked solution
The MLE is 9.0, the largest observation; it is biased low, and the unbiased correction is 5/4 x 9.0 = 11.25. The likelihood is 1 over theta to the fourth for any theta of at least 9.0 and zero below it, so it peaks at the sample maximum. But the maximum of n draws averages n/(n + 1) of theta, here 4/5, so scaling by (n + 1)/n removes the bias.
Why is the MLE the largest observation?
A friend draws four raffle tickets numbered from 1 up to some unknown top number, and the highest you see is 90. The top number is at least 90; guessing higher only spreads your belief over tickets nobody drew. The likelihood, 1 over theta to the n, is zero for any theta below the largest observation and falls as theta rises above it, so it is maximised exactly at the sample maximum, 9.0. This is a case where you do not differentiate: the maximum sits on a boundary, not where a slope is zero.
Four uniform draws on 0 to theta cut it into five gaps of equal expected size, so the largest draw averages four fifths of theta; the MLE of 9.0 therefore sits below theta, and scaling by 5/4 gives the unbiased 11.25, against 12.35 from doubling the sample mean. The relationshipL(\theta) the likelihood of the four observations n the number of observations, 4 \max the largest observation, 9.0 What it says in wordsThe likelihood peaks at the largest observation, which on average falls short of theta by a factor n/(n + 1).Why is it biased, and by how much?
The largest draw can never be above theta and is almost always below it. Four points dropped at random on 0 to theta cut it into five gaps of equal expected length, so the largest point sits on average four fifths of the way up, and the MLE underestimates theta by a fifth on average. Multiplying by 5/4 fixes it: 9.0 becomes 11.25. The bias shrinks as n grows, since n/(n + 1) tends to 1, but with four points it is large.
How does it compare with the obvious alternative?
The method of moments doubles the sample mean, since a uniform on 0 to theta averages theta/2: the mean here is 6.175, giving 12.35. Both 11.25 and 12.35 are unbiased, but the corrected maximum has a much smaller variance, theta squared over n(n + 2) against theta squared over 3n, because the largest draw carries the most information about the top of the range. With four points that is theta squared over 24 against theta squared over 12: half the variance.
Where candidates lose it
Candidates differentiate the log-likelihood, get minus n over theta, set it to zero and find no solution. The likelihood only falls on the allowed range, so the maximum sits at the boundary, the largest observation; say that before reaching for calculus.
The second miss is calling the MLE unbiased because maximum likelihood estimates are often well behaved. Here it is biased low by construction, and the interviewer expects the (n + 1)/n correction.
What the interviewer asks next
- What is the variance of the corrected estimator with four observations?
- What is the MLE if the distribution is uniform on theta to 2 theta?
- Derive the expected value of the maximum of n uniform draws.
Asked at AQR Capital Management, Trading, Town of Greenwich, 2022 (Wall Street Oasis):
Derive the mle for some given distribution. Explain linear regression intuitively and derive the ols estimate.
098With orthonormal regressors, ordinary least squares gives coefficients of 0.6 and 0.15. What do ridge and lasso with a penalty of 0.2 give for each, and why does only lasso set a coefficient to zero?CitadelLondon · 2026
Try it first
Using half the residual sum of squares plus the penalty, what do the two methods give?
Show the worked solution
Ridge gives 0.5 and 0.125; lasso gives 0.4 and exactly 0. With orthonormal regressors each coefficient is shrunk on its own. Ridge divides each by 1 + 0.2 = 1.2, so it scales both down and never reaches zero. Lasso subtracts 0.2 from each size and stops at zero, so the small coefficient, 0.15, is removed. This uses the scaling of half the residual sum of squares plus 0.2 times the penalty.
Why do orthonormal regressors make this a one-line problem?
When the regressors are uncorrelated and scaled to unit length, the fit for each coefficient does not depend on the others, so the penalised problem splits into separate one-variable problems. Think of adjusting the volume on two speakers that are not wired together: turning one down does not change the other. For each coefficient you minimise half of (beta minus b) squared plus the penalty, where b is its least squares value, and the answer depends only on that one number. Ridge's penalty is half of beta squared times 0.2; lasso's is the size of beta times 0.2.
Ridge scales every least squares coefficient by 1/1.2, so 0.6 becomes 0.5 and 0.15 becomes 0.125, while lasso subtracts 0.2 and stops at zero, so 0.6 becomes 0.4 and 0.15, inside the dead zone below 0.2, becomes exactly 0. What does each penalty do to a coefficient?
Ridge's squared penalty pulls hard on big coefficients and gently on small ones. Setting the slope to zero gives beta x (1 + 0.2) = b, so ridge divides by 1.2: 0.6 becomes 0.5 and 0.15 becomes 0.125, shrunk but never zero. Lasso's absolute-value penalty pulls with the same force, 0.2, whatever the size. Its answer is the sign of b times the larger of (the size of b minus 0.2) and zero: 0.6 becomes 0.4, and 0.15, weaker than the pull of 0.2, lands exactly on 0.
The relationshipb the least squares coefficient, 0.6 or 0.15 lambda the penalty weight, 0.2 What it says in wordsRidge divides every coefficient by the same factor; lasso takes the same amount off every coefficient and never goes past zero.Why does only lasso select variables?
At zero, the squared penalty is flat: its slope is zero, so any small coefficient still earns its place by improving the fit a little. The absolute-value penalty has a corner at zero with a slope of 0.2 on each side, so a coefficient stays at zero unless the fit improves by more than 0.2 per unit, and 0.15 does not. That is why lasso gives sparseHaving many coefficients exactly equal to zero, so the model uses only a few of the available variables. models and ridge does not. The limitation to state: with correlated regressors lasso tends to keep one of a group arbitrarily and drop the rest, which is why desks often blend the two penalties in an elastic net.
Where candidates lose it
The common slip is to swap the two, saying ridge sets small coefficients to zero because it penalises harder. Ridge's penalty is heavy on big coefficients and almost nothing on small ones, which is exactly why it never zeroes them.
The second loss is getting the lasso numbers off by a factor of two. Written as the full residual sum of squares plus 0.2 times the absolute values, without the half, the threshold is 0.1, giving 0.5 and 0.05. State your scaling before you give numbers.
What the interviewer asks next
- At what penalty does lasso set the 0.6 coefficient to zero as well?
- The two regressors now have a correlation of 0.9. How do ridge and lasso behave differently?
- How would you choose the penalty in practice without fitting it to noise?
Asked at Citadel, Quantitative Research, London, 2026 (Wall Street Oasis):
very detailed and difficult questions about regularisation ridge and lasso
