Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
100You regress a centred target y on one standardised feature x with no intercept. The sum of x squared is 100 and the sum of x times y is 80. What is the OLS slope, and what is the ridge slope with penalty lambda = 25?Citadel SecuritiesLondon · 2026
Try it first
Pick the pair.
Show the worked solution
OLS gives 0.8 and ridge gives 0.64. OLS minimises squared error and its slope is the sum of xy over the sum of x squared, 80/100. Ridge adds lambda times the slope squared to the loss, which puts lambda into the denominator: 80/(100 + 25) = 0.64. That is the OLS slope times 100/125 = 0.8, so ridge shrinks the slope towards zero but never to zero.
Where does lambda end up in the formula?
Think of a new analyst's forecast that you half trust: you do not discard it, you shade it towards zero, and the less data behind it the more you shade. Ridge does that mechanically. It minimises the squared errors plus lambda times the slope squared; setting the derivative to zero gives b = Sxy/(Sxx + lambda). Penalising the size of the slope acts exactly like adding observations whose x squared totals lambda and whose y is zero, data that say the slope is zero. With 100 of real evidence and 25 of make-believe evidence, the slope is 80/125 = 0.64.
The ridge slope 80/(100 + lambda) falls from the OLS value 0.8 to 0.64 at lambda 25 and 0.40 at lambda 100 without ever reaching zero, while the lasso slope falls in a straight line and hits exactly zero at lambda 160. The relationshipsum x_i y_i the cross-product of feature and target, 80 sum x_i^2 the sum of squares of the feature, 100 lambda the ridge penalty, 25 What it says in wordsRidge is OLS with lambda added to the sum of squares, so every slope is multiplied by Sxx/(Sxx + lambda).Why would you want a slope that is biased towards zero?
Because a smaller, steadier estimate can be closer to the truth on average. Suppose the true slope is 0.5 and the noise variance is 25. OLS is unbiased but its variance is 25/100 = 0.25. Ridge at lambda 25 has variance 0.16 and a bias of -0.1, so its mean squared error is 0.17. Ridge trades a little bias for a larger cut in variance, and when the signal is weak relative to the noise that trade wins. In this one-feature case the best lambda is noise variance over slope squared, 100, which halves the slope and cuts the error to 0.125. In practice the truth is unknown, so lambda is chosen by cross-validation.
lambda Slope on this data Variance Bias squared Mean squared error 0 0.80 0.2500 0.0000 0.2500 25 0.64 0.1600 0.0100 0.1700 100 0.40 0.0625 0.0625 0.1250 Assuming a true slope of 0.5 and noise variance 25, ridge at lambda 25 and 100 has a lower mean squared error than OLS because the drop in variance outweighs the bias it adds. How does lasso differ?
Lasso penalises lambda times the absolute slope instead. In one dimension that subtracts lambda/2 from the cross-product rather than adding to the denominator: (80 - 12.5)/100 = 0.675 at lambda 25, and exactly zero once lambda reaches 160. Ridge scales coefficients down; lasso shifts them down and can set them to exactly zero, which is why lasso selects features and ridge does not. Both penalties depend on the scale of x, which is why the feature must be standardised first; with correlated features, ridge spreads the weight across them while lasso tends to keep one.
Where candidates lose it
The fast wrong answer subtracts the penalty from the slope or from the numerator, which is lasso's mechanics, not ridge's. Ridge adds lambda to the sum of squares in the denominator, so the slope is scaled, not shifted.
The second loss is saying ridge is always better because it has lower variance. It trades variance for bias; if the true slope is large and the data plentiful, shrinking costs more in bias than it saves. Say that lambda is chosen by cross-validation, not by taste.
What the interviewer asks next
- What value of lambda halves the OLS slope?
- With two highly correlated features, how do ridge and lasso split the weight between them?
- Why must features be standardised before applying a ridge penalty?
Asked at Citadel Securities, Quantitative Research, London, 2026 (Wall Street Oasis):
very detailed and difficult questions about regularisation ridge and lasso
