Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
017The true model is y = x1 + x2 + noise, where x1 and x2 are standardised and have correlation 0.5. You regress y on x1 alone, then regress the residuals on x2. What coefficient do you get on x2, and how would you recover the true value of 1 in two stages?Quant researchQuant trading
Try it first
What coefficient does the second stage give on x2?
Show the worked solution
You get 0.75, not 1. Regressing y on x1 alone gives a slope of 1 + 0.5 = 1.5, because x1 soaks up the half of x2 that moves with it. The residual is x2 - 0.5x1 + noise, whose slope on x2 is 1 - 0.5 squared = 0.75. To recover 1, residualise x2 on x1 as well and regress the residual of y on the residual of x2: the Frisch-Waugh-Lovell theorem.
Where does the missing quarter go?
Picture two salespeople who often work the same client. If you credit all joint sales to the first before looking at the second, the second looks worse than they are, because some of their work was already booked to the first. Stage one regresses y on x1 alone, and since x2 is correlated with x1, the coefficient on x1 rises to 1.5: it takes credit for 0.5 of x2. That piece has been removed from the residual, so stage two can only find what is left of x2's effect.
With a correlation of 0.5, x2 splits into 0.5 x1 plus an orthogonal part; stage one assigns the 0.5 x1 piece to x1, so regressing the residual on raw x2 gives 0.75, while regressing it on the orthogonal part of x2 recovers the true 1. How do you get 0.75 exactly?
Write the residual out. y - 1.5x1 = x2 - 0.5x1 + noise, and the slope of that on x2 is its covariance with x2 over the variance of x2: (1 - 0.5 x 0.5)/1 = 0.75. The formula generalises to 1 - rho squared times the true coefficient, so the bias gets worse as the regressors get more correlated: with rho = 0.9 you would find only 0.19. A simulation of 100,000 observations gives 1.506 for stage one and 0.752 for stage two.
The relationship\rho the correlation between x1 and x2, 0.5 x_2 - \rho x_1 the part of x2 left after regressing it on x1 What it says in wordsRegressing on raw x2 shrinks the answer by one minus rho squared; regressing on the part of x2 orthogonal to x1 gives the true coefficient.What does Frisch-Waugh-Lovell tell you to do?
To get a variable's coefficient from a multiple regression in stages, partial the other regressors out of both y and that variable, then regress residual on residual. Here that means regressing x2 on x1 as well, keeping the orthogonal part x2 - 0.5x1, and regressing the stage-one residual on it. The slope comes back as exactly 1; the simulation gives 1.003. This is why factor-neutralising a signal before testing it, rather than after, matters in quant research: the order of the stages changes the answer.
Where candidates lose it
The common answer is 1, on the belief that regressing residuals step by step is the same as a multiple regression. It is only the same when the regressors are uncorrelated, and the question gives you a correlation of 0.5 precisely to break that.
The second loss is saying the answer is biased without saying which way or by how much. Give 1.5 for stage one, 0.75 for stage two, the 1 - rho squared rule, and the fix.
What the interviewer asks next
- What would the stage-two coefficient be if the correlation were -0.5?
- In the two-stage FWL regression, how do the standard errors compare with the full multiple regression?
- You have a new signal correlated with a known factor. How do you test whether it adds anything?
047Two assets' daily returns are negatively correlated within every month, yet their monthly returns are positively correlated across the year. How can that happen? Build a small numerical example.Squarepoint CapitalMontreal · 2024
Try it first
Both assets share a drift that changes from month to month, plus daily noise that is negatively correlated. What happens to the correlation as you sum more days into one return?
Show the worked solution
A drift shared by both assets for the whole month can outweigh daily noise that moves them in opposite directions. Within a month the drift is constant, so only the noise shows and the correlation is negative. Summed over 21 days the drift's covariance grows with 21 squared but the noise's only with 21. With noise correlation -0.5 and drift standard deviation 0.3% a day, monthly correlation is +0.48.
What does a three-month example look like?
Picture two shops in the same market street. On any one day, a customer who buys from one did not buy from the other, so their daily takings move against each other. But in festival months the whole street is busy and in the rains the whole street is quiet, so their monthly takings rise and fall together. Correlation at one horizon says nothing on its own about another, because a slow common factor and fast opposing noise can sit in the same data.
Make it numerical with five-day months. In month 1 both assets drift at -0.8% a day, in month 2 at +0.2%, in month 3 at +1.2%. On top, asset A gets daily noise of +0.6, -0.3, 0, +0.3, -0.6 and asset B gets -0.3, +0.3, 0, -0.3, +0.3, which move in opposite directions. Within each month the correlation is -0.95. Summed over each month, the noise nets to zero, so both assets return -4%, +1% and +6%: identical, a monthly correlation of +1. Even all 15 days pooled show +0.71, because the month-to-month swing in drift is larger than the noise.
Within each five-day month the daily returns of the two assets slope downward with a correlation of -0.95, but the monthly drifts of -0.8%, +0.2% and +1.2% a day are shared, so the three cluster centres rise together and the monthly returns correlate at +1. Why does summing more days push the correlation positive?
Write each daily return as the month's drift m plus noise. Over n days the drift adds up to n times m, while the noise adds up to a sum of n separate shocks. The drift's contribution to covariance scales with n squared, the noise's only with n, so the longer the horizon the more the shared drift wins. Take noise with standard deviation 1% a day and a within-month correlation of -0.5, and a drift whose standard deviation across months is 0.3% a day. For a 21-day month the drift adds 39.69 to the covariance and the noise subtracts 10.5, for a monthly correlation of 29.19/60.69 = 0.48.
The relationshipn number of days summed into one return \sigma_m standard deviation of the shared daily drift across months, 0.3% c daily noise covariance within a month, -0.5 \sigma daily noise standard deviation, 1% What it says in wordsShared drift covariance grows with the square of the horizon, independent noise covariance only in proportion to it.For a 21-day month the shared drift adds 39.69 to the covariance and the opposing daily noise takes away 10.5, so monthly returns correlate at +0.48, and the correlation of summed returns crosses from negative to positive at about 5.6 days. The crossover sits where n times 0.09 equals 0.5, about 5.6 days, so even weekly returns of five days would still show a slightly negative correlation, -0.03. A hedge sized on daily correlation can therefore fail at a monthly horizon, which is why a desk measures correlation at the frequency it actually holds risk. The other mechanisms worth naming are mean reversion in the spread between the two assets and stale prices that lag by a day; both also make correlation depend on frequency. The limitation of the example is the assumption that drift is constant inside a month and noise is independent from day to day.
Where candidates lose it
The common loss is saying it is impossible, or that it must be a data error, because correlation feels like a fixed property of two assets. It is a property of two assets at a horizon, and the interviewer wants the decomposition into a slow shared part and a fast opposing part.
The second is a hand-waved answer with no numbers. Build the five-day example in a minute, state that drift covariance scales with n squared and noise with n, and the explanation becomes checkable.
What the interviewer asks next
- What would make daily correlation positive but monthly correlation negative?
- How would you estimate the shared monthly drift from daily data?
- A pairs trader hedges at the daily beta and holds for a month. What goes wrong?
Asked at Squarepoint Capital, Hedge Fund, Montreal, 2024 (Wall Street Oasis):
correlation can be negative intra-month but positive across a year, how?
053Five assets each have unit variance, and every pair has correlation 0.4. What are the eigenvalues of the correlation matrix, and what share of total variance does the first principal component explain?Jump TradingPudong Xinqu · 2023
Try it first
Before any algebra: what share of variance does the first component explain?
Show the worked solution
One eigenvalue of 2.6 and four of 0.6, so the first principal component explains 52%. Write the matrix as 0.6 times the identity plus 0.4 times a matrix of ones. The all-ones vector is an eigenvector with eigenvalue 0.6 + 5 x 0.4 = 2.6; any vector whose weights sum to zero is killed by the ones matrix and has eigenvalue 0.6. The trace check: 2.6 + 4 x 0.6 = 5.
What structure should you spot before touching a determinant?
Think of five students whose marks all move together when the paper is hard, plus their own good and bad days. There is one shared shock and five private ones. An equicorrelation matrix is exactly that: R = (1 - rho) I + rho J, where J is the matrix of all ones, so its eigenvectors are those of J and you never need a characteristic polynomial. J sends the all-ones vector to 5 times itself and sends any vector whose entries sum to zero to zero. Those two facts give every eigenvalue.
The 5 by 5 matrix with 0.4 off the diagonal has one eigenvalue of 2.6, carried by the equal weight portfolio and explaining 52% of the variance, and four eigenvalues of 0.6, carried by long short combinations and explaining 12% each. How do the eigenvalues fall out, and how do you check them?
Apply R to the all-ones vector: each row sums to 1 + 4 x 0.4, so the equal weight portfolio has eigenvalue 1 + (n - 1) rho = 2.6. Apply R to any vector with weights summing to zero, such as long asset 1 and short asset 2: the rho J part vanishes and only (1 - rho) = 0.6 is left, and there are four independent such vectors. The eigenvalues must add to the trace, the sum of the diagonal, which is 5: 2.6 + 2.4 = 5.
The relationshiprho the common pairwise correlation, 0.4 n the number of assets, 5 1 1^T the all-ones matrix J What it says in wordsA common correlation creates one large factor for the average and leaves every long short combination with the same small variance.What does the answer say about a real portfolio?
The first component is the market: equal weights, and its share rises towards rho as you add assets. With 50 assets at the same correlation the first eigenvalue is 1 + 49 x 0.4 = 20.6, 41.2% of the total, while each of the other 49 stays at 0.6. Diversification removes the private shocks but never the common one. The same formula gives a limit: the smallest eigenvalue 1 - rho is always fine, but 1 + (n - 1) rho must stay positive, so five assets cannot all share a correlation below -0.25.
Where candidates lose it
The loss is trying to expand a 5 by 5 determinant by hand. It is slow, error prone and signals that you did not see the structure. The interviewer is waiting for identity plus ones matrix.
The second trap is reading 40% as the explained share because the correlation is 0.4. The share is (1 + (n - 1) rho)/n, which is 52% here and only approaches rho as n grows.
What the interviewer asks next
- What is the most negative common correlation five assets can have?
- What are the eigenvectors of the four 0.6 eigenvalues, and why are they not unique?
- If one asset is removed, what share does the first component explain?
- How would you spot a second factor, such as a sector, in the eigenvalues?
Asked at Jump Trading, Prop Trading, Pudong Xinqu, 2023 (Wall Street Oasis):
Some very difficult linear algebra questions about PCA and eigenvalues
065The correlation between X and Y is 0.2, and the correlation between Y and Z is 0.5. What is the full range of possible values for the correlation between X and Z?Tower Research CapitalNew York · 2019
Try it first
Which statement about corr(X, Z) is right?
Show the worked solution
Anywhere from about -0.75 to 0.95. The correlation matrix must be positive semidefinite, which bounds the third correlation at 0.2 x 0.5 plus or minus sqrt((1 - 0.2^2)(1 - 0.5^2)), that is 0.1 plus or minus 0.849. Geometrically, correlations are cosines of angles: X sits 78.5 degrees from Y and Z sits 60 degrees from Y, so X and Z are between 18.5 and 138.5 degrees apart.
Why does knowing two correlations restrict the third at all?
Think of three towns on a map. If A is close to B and B is close to C, then A cannot be far from C; if B is only loosely near both, A and C could be almost anywhere. Correlations behave like distances in disguise: each is the cosine of the angle between two returns viewed as vectors, and angles obey a triangle rule. X makes an angle of {TH_A:.1f} degrees with Y, because its cosine is 0.2, and Z makes 60 degrees with Y. The angle between X and Z is therefore at least the difference and at most the sum.
With X at 78.5 degrees from Y and Z at 60 degrees from Y, the angle between X and Z ranges from 18.5 to 138.5 degrees, so their correlation can be anything from -0.75 to 0.95, centred on 0.1. How do you get the bound algebraically?
Any valid correlation matrix must be positive semidefiniteEvery portfolio built from the variables has a variance of zero or more; for a correlation matrix this means its determinant and all leading minors are non-negative., because a portfolio cannot have negative variance. For three variables with correlations a, b and c, the condition is 1 - a^2 - b^2 - c^2 + 2abc of at least zero. Treating that as a quadratic in c gives c = ab plus or minus sqrt((1 - a^2)(1 - b^2)), so the third correlation lies in an interval centred on the product of the other two. With a = 0.2 and b = 0.5 that is 0.1 plus or minus sqrt(0.72), or -0.7485 to 0.9485.
The relationshiprho_XY 0.2, the correlation of X and Y rho_YZ 0.5, the correlation of Y and Z the square root the room left over after the parts of X and Z explained by Y What it says in wordsThe third correlation is the product of the two given ones, plus or minus how much of X and Z is unexplained by Y.What does the centre value 0.1 mean, and when is the sign forced?
Split X and Z each into a part explained by Y and a leftover. The explained parts always contribute 0.2 x 0.5 = 0.1; the leftovers can be correlated however you like, and they can move the total by up to 0.849 either way. So 0.1 is the answer only if the leftovers are uncorrelated. The sign of corr(X, Z) is forced positive only when the two given correlations are strong, specifically when their squares add to more than 1; with 0.8 and 0.7 the range is about 0.13 to 0.99. On a desk this is why two hedges that each track an index only loosely say almost nothing about each other.
Where candidates lose it
The two fast wrong answers are 0.1, from multiplying, and must be positive, from assuming correlation is transitive. Both treat correlation like a chain of causes rather than a geometry.
The second loss is reaching the determinant condition and stalling on the algebra. Lead with the angle picture: arccos 0.2 is about 78.5 degrees, arccos 0.5 is 60, and the bounds are the cosines of their sum and difference.
What the interviewer asks next
- If corr(X, Y) = 0.8 and corr(Y, Z) = 0.7, can corr(X, Z) be negative?
- What is the most negative common correlation three variables can share?
- Given corr(X, Y) and corr(Y, Z), what value of corr(X, Z) makes X and Z uncorrelated once Y is controlled for?
Asked at Tower Research Capital, Prop Trading, New York, 2019 (Wall Street Oasis):
What if the correlation between X and Y is 0.2 and the correlation between Y and Z is 0.5.
