Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
035Regressing y on x gives a slope of 0.8; regressing x on y gives a slope of 0.45. What is the R-squared of either regression, and what is the correlation?Tower Research CapitalNew York · 2014
Try it first
What is the correlation between x and y?
Show the worked solution
R-squared is 0.36 for both regressions and the correlation is 0.6. The slope of y on x is r times sd(y)/sd(x); the slope of x on y is r times sd(x)/sd(y). Multiplying them cancels the standard deviations and leaves r squared: 0.8 x 0.45 = 0.36. The correlation is +0.6, positive because both slopes are positive, and the ratio sd(y)/sd(x) is √(0.8/0.45) = 4/3.
Why are the two slopes not reciprocals of each other?
Tall parents tend to have tall children, but a little less tall; and tall children tend to have tall parents, but a little less tall. Both statements are true at once. Each regression predicts toward the mean, so neither slope is the inverse of the other unless the fit is perfect. If the points lay exactly on a line, the slope of x on y would be 1/0.8 = 1.25. It is 0.45 instead, and the size of that shortfall is what measures how loose the relationship is.
Fitting y on x gives the shallower line with slope 0.8 and fitting x on y gives the steeper line, slope 0.45 in its own terms; their product, 0.36, is R-squared, so the correlation is 0.6 and the standard deviation of y is 4/3 that of x. How do the two slopes give R-squared?
Write each slope in terms of the correlation. The least squares slope of y on x is the covariance over the variance of x, which is r times sd(y)/sd(x). Swap the roles and the slope of x on y is r times sd(x)/sd(y). The standard deviation ratios are reciprocals, so the product of the two slopes is r squared, and in a one-variable regression r squared is exactly the R-squared. Here 0.8 x 0.45 = 0.36, so r = 0.6; the sign is positive because both slopes are positive, and the two slopes always share a sign.
The relationshipb_y|x slope from regressing y on x, 0.8 b_x|y slope from regressing x on y, 0.45 s_x, s_y standard deviations of x and y r the correlation of x and y What it says in wordsThe two slopes multiply to the squared correlation because the scale factors cancel.The figure uses 40 points built with standard deviations 3 and 4 and a correlation of exactly 0.6, and fitting both regressions returns slopes of 0.80 and 0.45. A quick sanity test comes free: the product of the two slopes can never exceed 1. If an interviewer quotes slopes of 0.8 and 1.5, the product 1.2 is impossible, and saying so is worth more than any calculation.
Where candidates lose it
The fast wrong answer is to say the slopes should be reciprocals and call the data inconsistent, or to answer 0.36 when asked for the correlation. 0.36 is R-squared; the correlation is its square root.
The second loss is dropping the sign. The square root of 0.36 could be plus or minus 0.6; both slopes are positive, so the correlation is positive, and saying why takes one sentence.
What the interviewer asks next
- What is the ratio of the standard deviation of y to that of x?
- If the slope of x on y were 1.5, what would you conclude?
- How does adding measurement noise to x change each slope?
Asked at Tower Research Capital, Quantitative Research, New York, 2014 (Wall Street Oasis):
Another detailed linear regression questions were asked, including problems about residual, variance and R^2
078Let A be the 2 by 2 matrix with 2 on the diagonal and 1 off the diagonal. Compute A to the power 10 without multiplying it out ten times.Quant researchQuant trading
Try it first
What is the top-left entry of A^10?
Show the worked solution
A^10 has 29,525 on the diagonal and 29,524 off it. A has eigenvalue 3 along (1, 1) and eigenvalue 1 along (1, -1). Writing A = Q D Q^T with D = diag(3, 1), the tenth power is Q D^10 Q^T, and only the numbers 3 and 1 get raised to the tenth. The entries are (3^10 + 1)/2 and (3^10 - 1)/2.
Why look for eigenvectors at all?
Think of a photocopier set to 300% on one axis and 100% on the other. Copy a copy ten times and you do not need to simulate every pass: that axis is 3 to the tenth times longer and the other is unchanged. An eigenvector is a direction the matrix only stretches, so applying the matrix ten times along it is just multiplying by the eigenvalue ten times. Symmetric matrices always have a full set of such directions at right angles, which is what makes this matrix easy.
Find them by inspection. Adding the two rows of A gives 3 in each, so A(1, 1) = (3, 3): eigenvalue 3. Subtracting gives 1, so A(1, -1) = (1, -1): eigenvalue 1. The trace is 4 and the determinant is 3, and 3 + 1 = 4 and 3 x 1 = 3, which confirms both in one line.
The matrix stretches the direction (1, 1) by a factor of 3 and leaves (1, -1) unchanged, so A to the tenth stretches them by 59,049 and 1, and converting back to ordinary coordinates gives 29,525 on the diagonal and 29,524 off it. The relationshipQ the matrix whose columns are the unit eigenvectors D the diagonal matrix of eigenvalues, 3 and 1 Q^T the transpose of Q, which is also its inverse What it says in wordsRotate into the eigenvector directions, raise each eigenvalue to the tenth, and rotate back.Is there an even faster route for this particular matrix?
Yes. Write A = I + J, where J is the all-ones matrix. J squared is 2J, so every power of J is a multiple of J, and (I + J)^n collapses to I + ((3^n - 1)/2) J. For n = 10 that is I + 29,524 J, which gives 29,525 on the diagonal and 29,524 off it: the same answer, and a good cross-check to say aloud. A brute-force multiplication in code agrees exactly.
Say why this matters on a desk. A covariance matrix with equal variances and one common correlation has exactly this shape, and its eigenvectors are the market direction and the spread directions. Powers of transition matrices in Markov chains are computed the same way, and the eigenvalue closest to 1 tells you how fast the chain forgets where it started.
Where candidates lose it
The fast wrong answer raises each entry to the tenth, giving 1,024 on the diagonal and 1 off it. Matrix multiplication mixes rows and columns, so entries do not power separately; A squared already has 5 on the diagonal, not 4.
The second loss is diagonalising correctly and then fumbling the conversion back. The Q matrix carries a 1/sqrt(2) on each side, which becomes the factor of one half in the final answer. Check with the trace: the diagonal entries of A^10 must sum to 3^10 + 1.
What the interviewer asks next
- What is A^n as n grows large, after dividing by 3^n?
- Compute the square root of A, a symmetric matrix B with B squared equal to A.
- Generalise: an n by n matrix with a on the diagonal and b everywhere else. What are its eigenvalues?
100You regress a centred target y on one standardised feature x with no intercept. The sum of x squared is 100 and the sum of x times y is 80. What is the OLS slope, and what is the ridge slope with penalty lambda = 25?Citadel SecuritiesLondon · 2026
Try it first
Pick the pair.
Show the worked solution
OLS gives 0.8 and ridge gives 0.64. OLS minimises squared error and its slope is the sum of xy over the sum of x squared, 80/100. Ridge adds lambda times the slope squared to the loss, which puts lambda into the denominator: 80/(100 + 25) = 0.64. That is the OLS slope times 100/125 = 0.8, so ridge shrinks the slope towards zero but never to zero.
Where does lambda end up in the formula?
Think of a new analyst's forecast that you half trust: you do not discard it, you shade it towards zero, and the less data behind it the more you shade. Ridge does that mechanically. It minimises the squared errors plus lambda times the slope squared; setting the derivative to zero gives b = Sxy/(Sxx + lambda). Penalising the size of the slope acts exactly like adding observations whose x squared totals lambda and whose y is zero, data that say the slope is zero. With 100 of real evidence and 25 of make-believe evidence, the slope is 80/125 = 0.64.
The ridge slope 80/(100 + lambda) falls from the OLS value 0.8 to 0.64 at lambda 25 and 0.40 at lambda 100 without ever reaching zero, while the lasso slope falls in a straight line and hits exactly zero at lambda 160. The relationshipsum x_i y_i the cross-product of feature and target, 80 sum x_i^2 the sum of squares of the feature, 100 lambda the ridge penalty, 25 What it says in wordsRidge is OLS with lambda added to the sum of squares, so every slope is multiplied by Sxx/(Sxx + lambda).Why would you want a slope that is biased towards zero?
Because a smaller, steadier estimate can be closer to the truth on average. Suppose the true slope is 0.5 and the noise variance is 25. OLS is unbiased but its variance is 25/100 = 0.25. Ridge at lambda 25 has variance 0.16 and a bias of -0.1, so its mean squared error is 0.17. Ridge trades a little bias for a larger cut in variance, and when the signal is weak relative to the noise that trade wins. In this one-feature case the best lambda is noise variance over slope squared, 100, which halves the slope and cuts the error to 0.125. In practice the truth is unknown, so lambda is chosen by cross-validation.
lambda Slope on this data Variance Bias squared Mean squared error 0 0.80 0.2500 0.0000 0.2500 25 0.64 0.1600 0.0100 0.1700 100 0.40 0.0625 0.0625 0.1250 Assuming a true slope of 0.5 and noise variance 25, ridge at lambda 25 and 100 has a lower mean squared error than OLS because the drop in variance outweighs the bias it adds. How does lasso differ?
Lasso penalises lambda times the absolute slope instead. In one dimension that subtracts lambda/2 from the cross-product rather than adding to the denominator: (80 - 12.5)/100 = 0.675 at lambda 25, and exactly zero once lambda reaches 160. Ridge scales coefficients down; lasso shifts them down and can set them to exactly zero, which is why lasso selects features and ridge does not. Both penalties depend on the scale of x, which is why the feature must be standardised first; with correlated features, ridge spreads the weight across them while lasso tends to keep one.
Where candidates lose it
The fast wrong answer subtracts the penalty from the slope or from the numerator, which is lasso's mechanics, not ridge's. Ridge adds lambda to the sum of squares in the denominator, so the slope is scaled, not shifted.
The second loss is saying ridge is always better because it has lower variance. It trades variance for bias; if the true slope is large and the data plentiful, shrinking costs more in bias than it saves. Say that lambda is chosen by cross-validation, not by taste.
What the interviewer asks next
- What value of lambda halves the OLS slope?
- With two highly correlated features, how do ridge and lasso split the weight between them?
- Why must features be standardised before applying a ridge penalty?
Asked at Citadel Securities, Quantitative Research, London, 2026 (Wall Street Oasis):
very detailed and difficult questions about regularisation ridge and lasso
