Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant puzzles, solved step by step

Puzzles
100
Traced to a firm
71
Topics
12
Hard
30
Topic
All topicsLogic and algorithmic reasoning10Conditional probability and Bayes7Counting and combinatorics8Continuous and geometric probability9Correlation, regression and linear algebra9Market making, betting and sizing9Expected value and optimal stopping9Statistics and estimation9Pricing, options and index maths7Games and strategic reasoning8Markov chains and random walks7Mental maths and number sense8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 1–9 of 9 · filtered from 100Clear filters
  1. 005Construct two random variables that are uncorrelated but clearly dependent, and show that their covariance is zero.Correlation, regression and linear algebraWarm upTwo SigmaNew York · 2025

    Try it first

    Which pair works?

    Show the worked solution

    Take X equal to -1, 0 or 1 with probability 1/3 each, and Y = X squared. Y is fixed by X, so they are as dependent as variables can be. But E[X] = 0 and E[XY] = E[X cubed] = (-1 + 0 + 1)/3 = 0, so the covariance E[XY] - E[X]E[Y] is zero. Correlation only measures straight-line association, and this relationship is a V.

    What does correlation actually measure?

    Think of a thermostat that runs the air conditioner hard on very hot days and the heater hard on very cold days. Energy use is clearly driven by temperature, but a straight line through the data is flat: high use at both ends, low in the middle. Correlation measures only how well a straight line summarises the relationship, so any symmetric U or V shape can have zero correlation while being completely determined. Independence is the stronger claim that knowing X tells you nothing about Y at all.

    Y is fixed by X, yet the best straight line is flatfit: Y = 2/3(-1, 1)(1, 1)(0, 0)X = -1, 0, 1 equally likely, Y = X squaredE[XY] = (-1 + 0 + 1)/3 = 0 = E[X] E[Y]fit: Y = 1/3Y = X squaredX uniform on -1 to 1, Y = X squaredCov(X, Y) = E[X cubed] = 0 by symmetry
    With X symmetric about zero and Y equal to X squared, Y is fixed exactly by X, yet the best straight line through the points is flat, so the covariance and the correlation are both zero.

    How do you show the covariance is zero in one line?

    Write the definition and let symmetry do the work. Cov(X, Y) = E[XY] - E[X]E[Y], and with Y = X squared the first term is E[X cubed], which is zero for any X symmetric about zero; the second term has E[X] = 0 in it. With the three-point version you can even list the products: -1 x 1, 0 x 0 and 1 x 1 add to zero. Yet P(Y = 0 given X = 0) is 1 while P(Y = 0) is 1/3, which is dependence in plain sight.

    The relationship
    Cov⁡(X,X2)=E[X3]−E[X] E[X2]=0−0⋅23=0P(Y=0∣X=0)=1≠P(Y=0)=13\operatorname{Cov}(X, X^2) = E[X^3] - E[X]\,E[X^2] = 0 - 0\cdot\tfrac{2}{3} = 0 \qquad P(Y=0\mid X=0) = 1 \ne P(Y=0) = \tfrac{1}{3}
    E[X^3]zero because the values of X are symmetric about zero
    P(Y = 0 | X = 0)knowing X changes the odds on Y, so they are dependent
    What it says in wordsThe covariance cancels by symmetry, while a single conditional probability proves the dependence.

    Why does a quant interviewer care?

    Because models quietly substitute zero correlation for no relationship. A delta-hedged option book gains or loses roughly with the square of the underlying's move, so its daily P and L can show near-zero correlation with the market while being entirely driven by it. The same holds for a volatility strategy or any payoff with a kink. The one case where zero correlation does mean independence is when the pair is jointly normal, which is worth adding before the interviewer asks.

    Where candidates lose it

    Candidates reach for two independent variables, which are uncorrelated but not dependent, or for X and -X, which are dependent but perfectly correlated. Both show the definitions are fuzzy.

    The quieter trap is choosing X uniform on 0 to 1 and Y = X squared. Without symmetry about zero the covariance is positive, 1/12, and the example fails. Centre X first.

    What the interviewer asks next

    • When does zero correlation imply independence?
    • Give an example with zero correlation where Y is not a function of X.
    • If you regress Y on X in the example, what do the fitted line and R squared look like?

    Asked at Two Sigma, Generalist, New York, 2025 (Wall Street Oasis): Come up with two uncorrelated but dependent variables.

  2. 017The true model is y = x1 + x2 + noise, where x1 and x2 are standardised and have correlation 0.5. You regress y on x1 alone, then regress the residuals on x2. What coefficient do you get on x2, and how would you recover the true value of 1 in two stages?Correlation, regression and linear algebraHardQuant researchQuant trading

    Try it first

    What coefficient does the second stage give on x2?

    Show the worked solution

    You get 0.75, not 1. Regressing y on x1 alone gives a slope of 1 + 0.5 = 1.5, because x1 soaks up the half of x2 that moves with it. The residual is x2 - 0.5x1 + noise, whose slope on x2 is 1 - 0.5 squared = 0.75. To recover 1, residualise x2 on x1 as well and regress the residual of y on the residual of x2: the Frisch-Waugh-Lovell theorem.

    Where does the missing quarter go?

    Picture two salespeople who often work the same client. If you credit all joint sales to the first before looking at the second, the second looks worse than they are, because some of their work was already booked to the first. Stage one regresses y on x1 alone, and since x2 is correlated with x1, the coefficient on x1 rises to 1.5: it takes credit for 0.5 of x2. That piece has been removed from the residual, so stage two can only find what is left of x2's effect.

    x1 has already eaten the part of x2 that points its way0.5 x1: already in x1x1x2part of x2 orthogonalto x1: length 0.87cos 0.5CoefficientsStage 1: y on x11.50True effect of x21.00Residual on raw x20.75Residual on x2 orthogonal1.0010
    With a correlation of 0.5, x2 splits into 0.5 x1 plus an orthogonal part; stage one assigns the 0.5 x1 piece to x1, so regressing the residual on raw x2 gives 0.75, while regressing it on the orthogonal part of x2 recovers the true 1.

    How do you get 0.75 exactly?

    Write the residual out. y - 1.5x1 = x2 - 0.5x1 + noise, and the slope of that on x2 is its covariance with x2 over the variance of x2: (1 - 0.5 x 0.5)/1 = 0.75. The formula generalises to 1 - rho squared times the true coefficient, so the bias gets worse as the regressors get more correlated: with rho = 0.9 you would find only 0.19. A simulation of 100,000 observations gives 1.506 for stage one and 0.752 for stage two.

    The relationship
    β^2,seq=Cov⁡(x2−ρx1, x2)Var⁡(x2)=1−ρ2=0.75β^2,FWL=Cov⁡(x2−ρx1, x2−ρx1)Var⁡(x2−ρx1)=1\hat\beta_{2,\text{seq}} = \frac{\operatorname{Cov}(x_2 - \rho x_1,\ x_2)}{\operatorname{Var}(x_2)} = 1 - \rho^2 = 0.75 \qquad \hat\beta_{2,\text{FWL}} = \frac{\operatorname{Cov}(x_2 - \rho x_1,\ x_2 - \rho x_1)}{\operatorname{Var}(x_2 - \rho x_1)} = 1
    \rhothe correlation between x1 and x2, 0.5
    x_2 - \rho x_1the part of x2 left after regressing it on x1
    What it says in wordsRegressing on raw x2 shrinks the answer by one minus rho squared; regressing on the part of x2 orthogonal to x1 gives the true coefficient.

    What does Frisch-Waugh-Lovell tell you to do?

    To get a variable's coefficient from a multiple regression in stages, partial the other regressors out of both y and that variable, then regress residual on residual. Here that means regressing x2 on x1 as well, keeping the orthogonal part x2 - 0.5x1, and regressing the stage-one residual on it. The slope comes back as exactly 1; the simulation gives 1.003. This is why factor-neutralising a signal before testing it, rather than after, matters in quant research: the order of the stages changes the answer.

    Where candidates lose it

    The common answer is 1, on the belief that regressing residuals step by step is the same as a multiple regression. It is only the same when the regressors are uncorrelated, and the question gives you a correlation of 0.5 precisely to break that.

    The second loss is saying the answer is biased without saying which way or by how much. Give 1.5 for stage one, 0.75 for stage two, the 1 - rho squared rule, and the fix.

    What the interviewer asks next

    • What would the stage-two coefficient be if the correlation were -0.5?
    • In the two-stage FWL regression, how do the standard errors compare with the full multiple regression?
    • You have a new signal correlated with a known factor. How do you test whether it adds anything?
  3. 035Regressing y on x gives a slope of 0.8; regressing x on y gives a slope of 0.45. What is the R-squared of either regression, and what is the correlation?Correlation, regression and linear algebraCoreTower Research CapitalNew York · 2014

    Try it first

    What is the correlation between x and y?

    Show the worked solution

    R-squared is 0.36 for both regressions and the correlation is 0.6. The slope of y on x is r times sd(y)/sd(x); the slope of x on y is r times sd(x)/sd(y). Multiplying them cancels the standard deviations and leaves r squared: 0.8 x 0.45 = 0.36. The correlation is +0.6, positive because both slopes are positive, and the ratio sd(y)/sd(x) is √(0.8/0.45) = 4/3.

    Why are the two slopes not reciprocals of each other?

    Tall parents tend to have tall children, but a little less tall; and tall children tend to have tall parents, but a little less tall. Both statements are true at once. Each regression predicts toward the mean, so neither slope is the inverse of the other unless the fit is perfect. If the points lay exactly on a line, the slope of x on y would be 1/0.8 = 1.25. It is 0.45 instead, and the size of that shortfall is what measures how loose the relationship is.

    Two regressions, two lines: the slopes multiply to R-squaredxyy on x: slope 0.8x on y: slope 0.45, drawn as 1/0.45 = 2.22Slope of y on x = r x (sd of y / sd of x)Slope of x on y = r x (sd of x / sd of y)Multiply the two: the sd ratios cancel0.8 x 0.45 = 0.36 = R-squaredr = +0.6Divide instead: sd of y / sd of x= √(0.8 / 0.45) = 1.333Neither slope is the reciprocal of the otherbecause |r| < 1: both regress toward the mean
    Fitting y on x gives the shallower line with slope 0.8 and fitting x on y gives the steeper line, slope 0.45 in its own terms; their product, 0.36, is R-squared, so the correlation is 0.6 and the standard deviation of y is 4/3 that of x.

    How do the two slopes give R-squared?

    Write each slope in terms of the correlation. The least squares slope of y on x is the covariance over the variance of x, which is r times sd(y)/sd(x). Swap the roles and the slope of x on y is r times sd(x)/sd(y). The standard deviation ratios are reciprocals, so the product of the two slopes is r squared, and in a one-variable regression r squared is exactly the R-squared. Here 0.8 x 0.45 = 0.36, so r = 0.6; the sign is positive because both slopes are positive, and the two slopes always share a sign.

    The relationship
    by∣x bx∣y=rsysx⋅rsxsy=r2=0.8×0.45=0.36b_{y|x}\, b_{x|y} = r\frac{s_y}{s_x}\cdot r\frac{s_x}{s_y} = r^2 = 0.8 \times 0.45 = 0.36
    b_y|xslope from regressing y on x, 0.8
    b_x|yslope from regressing x on y, 0.45
    s_x, s_ystandard deviations of x and y
    rthe correlation of x and y
    What it says in wordsThe two slopes multiply to the squared correlation because the scale factors cancel.

    The figure uses 40 points built with standard deviations 3 and 4 and a correlation of exactly 0.6, and fitting both regressions returns slopes of 0.80 and 0.45. A quick sanity test comes free: the product of the two slopes can never exceed 1. If an interviewer quotes slopes of 0.8 and 1.5, the product 1.2 is impossible, and saying so is worth more than any calculation.

    Where candidates lose it

    The fast wrong answer is to say the slopes should be reciprocals and call the data inconsistent, or to answer 0.36 when asked for the correlation. 0.36 is R-squared; the correlation is its square root.

    The second loss is dropping the sign. The square root of 0.36 could be plus or minus 0.6; both slopes are positive, so the correlation is positive, and saying why takes one sentence.

    What the interviewer asks next

    • What is the ratio of the standard deviation of y to that of x?
    • If the slope of x on y were 1.5, what would you conclude?
    • How does adding measurement noise to x change each slope?

    Asked at Tower Research Capital, Quantitative Research, New York, 2014 (Wall Street Oasis): Another detailed linear regression questions were asked, including problems about residual, variance and R^2

  4. 047Two assets' daily returns are negatively correlated within every month, yet their monthly returns are positively correlated across the year. How can that happen? Build a small numerical example.Correlation, regression and linear algebraHardSCSquarepoint CapitalMontreal · 2024

    Try it first

    Both assets share a drift that changes from month to month, plus daily noise that is negatively correlated. What happens to the correlation as you sum more days into one return?

    Show the worked solution

    A drift shared by both assets for the whole month can outweigh daily noise that moves them in opposite directions. Within a month the drift is constant, so only the noise shows and the correlation is negative. Summed over 21 days the drift's covariance grows with 21 squared but the noise's only with 21. With noise correlation -0.5 and drift standard deviation 0.3% a day, monthly correlation is +0.48.

    What does a three-month example look like?

    Picture two shops in the same market street. On any one day, a customer who buys from one did not buy from the other, so their daily takings move against each other. But in festival months the whole street is busy and in the rains the whole street is quiet, so their monthly takings rise and fall together. Correlation at one horizon says nothing on its own about another, because a slow common factor and fast opposing noise can sit in the same data.

    Make it numerical with five-day months. In month 1 both assets drift at -0.8% a day, in month 2 at +0.2%, in month 3 at +1.2%. On top, asset A gets daily noise of +0.6, -0.3, 0, +0.3, -0.6 and asset B gets -0.3, +0.3, 0, -0.3, +0.3, which move in opposite directions. Within each month the correlation is -0.95. Summed over each month, the noise nets to zero, so both assets return -4%, +1% and +6%: identical, a monthly correlation of +1. Even all 15 days pooled show +0.71, because the month-to-month swing in drift is larger than the noise.

    Inside each month the points fall; across months the clusters climb-1%0+1%+2%-1%0+1%asset A daily returnasset B daily returnMonth 1Month 2Month 3Five-day monthsDriftWithin rMonth A, B-0.8%-0.95-4%, -4%+0.2%-0.95+1%, +1%+1.2%-0.95+6%, +6%Noise nets to zero inside each monthCorrelation by frequencyWithin each month: -0.95All 15 days pooled: +0.71Monthly returns: +1.00
    Within each five-day month the daily returns of the two assets slope downward with a correlation of -0.95, but the monthly drifts of -0.8%, +0.2% and +1.2% a day are shared, so the three cluster centres rise together and the monthly returns correlate at +1.

    Why does summing more days push the correlation positive?

    Write each daily return as the month's drift m plus noise. Over n days the drift adds up to n times m, while the noise adds up to a sum of n separate shocks. The drift's contribution to covariance scales with n squared, the noise's only with n, so the longer the horizon the more the shared drift wins. Take noise with standard deviation 1% a day and a within-month correlation of -0.5, and a drift whose standard deviation across months is 0.3% a day. For a 21-day month the drift adds 39.69 to the covariance and the noise subtracts 10.5, for a monthly correlation of 29.19/60.69 = 0.48.

    The relationship
    ρ(n)=n2σm2+n cn2σm2+n σ2ρ(21)=39.69−10.539.69+21=0.48\rho(n) = \frac{n^2\sigma_m^2 + n\,c}{n^2\sigma_m^2 + n\,\sigma^2} \qquad \rho(21) = \frac{39.69 - 10.5}{39.69 + 21} = 0.48
    nnumber of days summed into one return
    \sigma_mstandard deviation of the shared daily drift across months, 0.3%
    cdaily noise covariance within a month, -0.5
    \sigmadaily noise standard deviation, 1%
    What it says in wordsShared drift covariance grows with the square of the horizon, independent noise covariance only in proportion to it.
    Drift covariance grows with days squared, noise covariance only with days+39.69Drift-10.5Noise+29.19Monthly cov21-day month, in % squared21 x 21 x 0.3^2 = 39.69; 21 x (-0.5) = -10.5Variance 60.69, so monthly r = 29.19/60.69 = 0.48-0.4+0.401 day: -0.385 days: -0.0321 days: +0.48flips at 5.6 daysdays summed into one return11121Correlation of summed returnsr(n) = (0.09 n - 0.5) / (0.09 n + 1)
    For a 21-day month the shared drift adds 39.69 to the covariance and the opposing daily noise takes away 10.5, so monthly returns correlate at +0.48, and the correlation of summed returns crosses from negative to positive at about 5.6 days.

    The crossover sits where n times 0.09 equals 0.5, about 5.6 days, so even weekly returns of five days would still show a slightly negative correlation, -0.03. A hedge sized on daily correlation can therefore fail at a monthly horizon, which is why a desk measures correlation at the frequency it actually holds risk. The other mechanisms worth naming are mean reversion in the spread between the two assets and stale prices that lag by a day; both also make correlation depend on frequency. The limitation of the example is the assumption that drift is constant inside a month and noise is independent from day to day.

    Where candidates lose it

    The common loss is saying it is impossible, or that it must be a data error, because correlation feels like a fixed property of two assets. It is a property of two assets at a horizon, and the interviewer wants the decomposition into a slow shared part and a fast opposing part.

    The second is a hand-waved answer with no numbers. Build the five-day example in a minute, state that drift covariance scales with n squared and noise with n, and the explanation becomes checkable.

    What the interviewer asks next

    • What would make daily correlation positive but monthly correlation negative?
    • How would you estimate the shared monthly drift from daily data?
    • A pairs trader hedges at the daily beta and holds for a month. What goes wrong?

    Asked at Squarepoint Capital, Hedge Fund, Montreal, 2024 (Wall Street Oasis): correlation can be negative intra-month but positive across a year, how?

  5. 053Five assets each have unit variance, and every pair has correlation 0.4. What are the eigenvalues of the correlation matrix, and what share of total variance does the first principal component explain?Correlation, regression and linear algebraHardJump TradingPudong Xinqu · 2023

    Try it first

    Before any algebra: what share of variance does the first component explain?

    Show the worked solution

    One eigenvalue of 2.6 and four of 0.6, so the first principal component explains 52%. Write the matrix as 0.6 times the identity plus 0.4 times a matrix of ones. The all-ones vector is an eigenvector with eigenvalue 0.6 + 5 x 0.4 = 2.6; any vector whose weights sum to zero is killed by the ones matrix and has eigenvalue 0.6. The trace check: 2.6 + 4 x 0.6 = 5.

    What structure should you spot before touching a determinant?

    Think of five students whose marks all move together when the paper is hard, plus their own good and bad days. There is one shared shock and five private ones. An equicorrelation matrix is exactly that: R = (1 - rho) I + rho J, where J is the matrix of all ones, so its eigenvectors are those of J and you never need a characteristic polynomial. J sends the all-ones vector to 5 times itself and sends any vector whose entries sum to zero to zero. Those two facts give every eigenvalue.

    One market factor and four equal leftovers: the eigenvalues of R10.40.40.40.40.410.40.40.40.40.410.40.40.40.40.410.40.40.40.40.41Correlation matrix Rtrace = 5 = sum of eigenvaluesaverage eigenvalue 12.6PC152%0.6PC212%0.6PC312%0.6PC412%0.6PC512%PC1: equal weights, the market1 + (n - 1) x 0.4 = 2.6, 52% of variancePC2 to PC5: weights summing to zero1 - 0.4 = 0.6 each, 12% each
    The 5 by 5 matrix with 0.4 off the diagonal has one eigenvalue of 2.6, carried by the equal weight portfolio and explaining 52% of the variance, and four eigenvalues of 0.6, carried by long short combinations and explaining 12% each.

    How do the eigenvalues fall out, and how do you check them?

    Apply R to the all-ones vector: each row sums to 1 + 4 x 0.4, so the equal weight portfolio has eigenvalue 1 + (n - 1) rho = 2.6. Apply R to any vector with weights summing to zero, such as long asset 1 and short asset 2: the rho J part vanishes and only (1 - rho) = 0.6 is left, and there are four independent such vectors. The eigenvalues must add to the trace, the sum of the diagonal, which is 5: 2.6 + 2.4 = 5.

    The relationship
    R=(1−ρ)I+ρ 11⊤λ1=1+(n−1)ρ=2.6,λ2..5=1−ρ=0.6,λ1n=52%R = (1-\rho)I + \rho\,\mathbf{1}\mathbf{1}^{\top} \qquad \lambda_1 = 1+(n-1)\rho = 2.6,\quad \lambda_{2..5} = 1-\rho = 0.6,\quad \frac{\lambda_1}{n} = 52\%
    rhothe common pairwise correlation, 0.4
    nthe number of assets, 5
    1 1^Tthe all-ones matrix J
    What it says in wordsA common correlation creates one large factor for the average and leaves every long short combination with the same small variance.

    What does the answer say about a real portfolio?

    The first component is the market: equal weights, and its share rises towards rho as you add assets. With 50 assets at the same correlation the first eigenvalue is 1 + 49 x 0.4 = 20.6, 41.2% of the total, while each of the other 49 stays at 0.6. Diversification removes the private shocks but never the common one. The same formula gives a limit: the smallest eigenvalue 1 - rho is always fine, but 1 + (n - 1) rho must stay positive, so five assets cannot all share a correlation below -0.25.

    Where candidates lose it

    The loss is trying to expand a 5 by 5 determinant by hand. It is slow, error prone and signals that you did not see the structure. The interviewer is waiting for identity plus ones matrix.

    The second trap is reading 40% as the explained share because the correlation is 0.4. The share is (1 + (n - 1) rho)/n, which is 52% here and only approaches rho as n grows.

    What the interviewer asks next

    • What is the most negative common correlation five assets can have?
    • What are the eigenvectors of the four 0.6 eigenvalues, and why are they not unique?
    • If one asset is removed, what share does the first component explain?
    • How would you spot a second factor, such as a sector, in the eigenvalues?

    Asked at Jump Trading, Prop Trading, Pudong Xinqu, 2023 (Wall Street Oasis): Some very difficult linear algebra questions about PCA and eigenvalues

  6. 065The correlation between X and Y is 0.2, and the correlation between Y and Z is 0.5. What is the full range of possible values for the correlation between X and Z?Correlation, regression and linear algebraHardTower Research CapitalNew York · 2019

    Try it first

    Which statement about corr(X, Z) is right?

    Show the worked solution

    Anywhere from about -0.75 to 0.95. The correlation matrix must be positive semidefinite, which bounds the third correlation at 0.2 x 0.5 plus or minus sqrt((1 - 0.2^2)(1 - 0.5^2)), that is 0.1 plus or minus 0.849. Geometrically, correlations are cosines of angles: X sits 78.5 degrees from Y and Z sits 60 degrees from Y, so X and Z are between 18.5 and 138.5 degrees apart.

    Why does knowing two correlations restrict the third at all?

    Think of three towns on a map. If A is close to B and B is close to C, then A cannot be far from C; if B is only loosely near both, A and C could be almost anywhere. Correlations behave like distances in disguise: each is the cosine of the angle between two returns viewed as vectors, and angles obey a triangle rule. X makes an angle of {TH_A:.1f} degrees with Y, because its cosine is 0.2, and Z makes 60 degrees with Y. The angle between X and Z is therefore at least the difference and at most the sum.

    Correlations are cosines of angles, so the third one is boxed inYX: 78.5 deg from YZ: 60 deg, same sideZ: 60 deg, other sidecos 78.5 deg = 0.2cos 60 deg = 0.5Same side: 78.5 - 60 = 18.5 deg, corr 0.95Other side: 78.5 + 60 = 138.5 deg, corr -0.75-101centre 0.2 x 0.5 = 0.1-0.750.95Possible values of corr(X, Z)0.1 plus or minus sqrt(0.96 x 0.75)= 0.1 plus or minus 0.849Anything outside makes the correlationmatrix impossible: a negative varianceEven a negative correlation is allowed
    With X at 78.5 degrees from Y and Z at 60 degrees from Y, the angle between X and Z ranges from 18.5 to 138.5 degrees, so their correlation can be anything from -0.75 to 0.95, centred on 0.1.

    How do you get the bound algebraically?

    Any valid correlation matrix must be positive semidefiniteEvery portfolio built from the variables has a variance of zero or more; for a correlation matrix this means its determinant and all leading minors are non-negative., because a portfolio cannot have negative variance. For three variables with correlations a, b and c, the condition is 1 - a^2 - b^2 - c^2 + 2abc of at least zero. Treating that as a quadratic in c gives c = ab plus or minus sqrt((1 - a^2)(1 - b^2)), so the third correlation lies in an interval centred on the product of the other two. With a = 0.2 and b = 0.5 that is 0.1 plus or minus sqrt(0.72), or -0.7485 to 0.9485.

    The relationship
    ρXZ∈[ρXYρYZ−(1−ρXY2)(1−ρYZ2), ρXYρYZ+(1−ρXY2)(1−ρYZ2)]=[−0.749, 0.949]\rho_{XZ} \in \Big[\rho_{XY}\rho_{YZ} - \sqrt{(1-\rho_{XY}^2)(1-\rho_{YZ}^2)},\ \rho_{XY}\rho_{YZ} + \sqrt{(1-\rho_{XY}^2)(1-\rho_{YZ}^2)}\Big] = [-0.749,\ 0.949]
    rho_XY0.2, the correlation of X and Y
    rho_YZ0.5, the correlation of Y and Z
    the square rootthe room left over after the parts of X and Z explained by Y
    What it says in wordsThe third correlation is the product of the two given ones, plus or minus how much of X and Z is unexplained by Y.

    What does the centre value 0.1 mean, and when is the sign forced?

    Split X and Z each into a part explained by Y and a leftover. The explained parts always contribute 0.2 x 0.5 = 0.1; the leftovers can be correlated however you like, and they can move the total by up to 0.849 either way. So 0.1 is the answer only if the leftovers are uncorrelated. The sign of corr(X, Z) is forced positive only when the two given correlations are strong, specifically when their squares add to more than 1; with 0.8 and 0.7 the range is about 0.13 to 0.99. On a desk this is why two hedges that each track an index only loosely say almost nothing about each other.

    Where candidates lose it

    The two fast wrong answers are 0.1, from multiplying, and must be positive, from assuming correlation is transitive. Both treat correlation like a chain of causes rather than a geometry.

    The second loss is reaching the determinant condition and stalling on the algebra. Lead with the angle picture: arccos 0.2 is about 78.5 degrees, arccos 0.5 is 60, and the bounds are the cosines of their sum and difference.

    What the interviewer asks next

    • If corr(X, Y) = 0.8 and corr(Y, Z) = 0.7, can corr(X, Z) be negative?
    • What is the most negative common correlation three variables can share?
    • Given corr(X, Y) and corr(Y, Z), what value of corr(X, Z) makes X and Z uncorrelated once Y is controlled for?

    Asked at Tower Research Capital, Prop Trading, New York, 2019 (Wall Street Oasis): What if the correlation between X and Y is 0.2 and the correlation between Y and Z is 0.5.

  7. 078Let A be the 2 by 2 matrix with 2 on the diagonal and 1 off the diagonal. Compute A to the power 10 without multiplying it out ten times.Correlation, regression and linear algebraCoreQuant researchQuant trading

    Try it first

    What is the top-left entry of A^10?

    Show the worked solution

    A^10 has 29,525 on the diagonal and 29,524 off it. A has eigenvalue 3 along (1, 1) and eigenvalue 1 along (1, -1). Writing A = Q D Q^T with D = diag(3, 1), the tenth power is Q D^10 Q^T, and only the numbers 3 and 1 get raised to the tenth. The entries are (3^10 + 1)/2 and (3^10 - 1)/2.

    Why look for eigenvectors at all?

    Think of a photocopier set to 300% on one axis and 100% on the other. Copy a copy ten times and you do not need to simulate every pass: that axis is 3 to the tenth times longer and the other is unchanged. An eigenvector is a direction the matrix only stretches, so applying the matrix ten times along it is just multiplying by the eigenvalue ten times. Symmetric matrices always have a full set of such directions at right angles, which is what makes this matrix easy.

    Find them by inspection. Adding the two rows of A gives 3 in each, so A(1, 1) = (3, 3): eigenvalue 3. Subtracting gives 1, so A(1, -1) = (1, -1): eigenvalue 1. The trace is 4 and the determinant is 3, and 3 + 1 = 4 and 3 x 1 = 3, which confirms both in one line.

    Two directions the matrix only stretches: powers become powers of numbersxy(1, 1)A(1, 1) = (3, 3)(1, -1) = A(1, -1)eigenvalue 3 along (1, 1); eigenvalue 1 along (1, -1)A = Q D Q^TQ holds the unit eigenvectors, D = diag(3, 1)A^10 = Q D^10 Q^Tthe Q^T Q pairs in the middle cancelD^10 = diag(59,049, 1)only two numbers get raised to the 10thA^10: diagonal (3^10 + 1)/2, off it (3^10 - 1)/229,52529,52429,52429,525
    The matrix stretches the direction (1, 1) by a factor of 3 and leaves (1, -1) unchanged, so A to the tenth stretches them by 59,049 and 1, and converting back to ordinary coordinates gives 29,525 on the diagonal and 29,524 off it.
    The relationship
    A=Q(3001)QT,  Q=12(111−1)⇒A10=12(310+1310−1310−1310+1)A = Q\begin{pmatrix}3&0\\0&1\end{pmatrix}Q^{T},\; Q=\tfrac{1}{\sqrt2}\begin{pmatrix}1&1\\1&-1\end{pmatrix} \Rightarrow A^{10} = \tfrac12\begin{pmatrix}3^{10}+1 & 3^{10}-1\\ 3^{10}-1 & 3^{10}+1\end{pmatrix}
    Qthe matrix whose columns are the unit eigenvectors
    Dthe diagonal matrix of eigenvalues, 3 and 1
    Q^Tthe transpose of Q, which is also its inverse
    What it says in wordsRotate into the eigenvector directions, raise each eigenvalue to the tenth, and rotate back.

    Is there an even faster route for this particular matrix?

    Yes. Write A = I + J, where J is the all-ones matrix. J squared is 2J, so every power of J is a multiple of J, and (I + J)^n collapses to I + ((3^n - 1)/2) J. For n = 10 that is I + 29,524 J, which gives 29,525 on the diagonal and 29,524 off it: the same answer, and a good cross-check to say aloud. A brute-force multiplication in code agrees exactly.

    Say why this matters on a desk. A covariance matrix with equal variances and one common correlation has exactly this shape, and its eigenvectors are the market direction and the spread directions. Powers of transition matrices in Markov chains are computed the same way, and the eigenvalue closest to 1 tells you how fast the chain forgets where it started.

    Where candidates lose it

    The fast wrong answer raises each entry to the tenth, giving 1,024 on the diagonal and 1 off it. Matrix multiplication mixes rows and columns, so entries do not power separately; A squared already has 5 on the diagonal, not 4.

    The second loss is diagonalising correctly and then fumbling the conversion back. The Q matrix carries a 1/sqrt(2) on each side, which becomes the factor of one half in the final answer. Check with the trace: the diagonal entries of A^10 must sum to 3^10 + 1.

    What the interviewer asks next

    • What is A^n as n grows large, after dividing by 3^n?
    • Compute the square root of A, a symmetric matrix B with B squared equal to A.
    • Generalise: an n by n matrix with a on the diagonal and b everywhere else. What are its eigenvalues?
  8. 090X and Y are independent random variables with the same variance. What is the correlation between X and X + Y?Correlation, regression and linear algebraWarm upSCSquarepoint CapitalMontreal · 2026

    Try it first

    Pick the correlation.

    Show the worked solution

    1/sqrt(2), about 0.71. The covariance of X with X + Y is Var X plus Cov(X, Y), and the second term is zero, so it is sigma squared. The variance of X + Y is 2 sigma squared, because independent variances add. Dividing sigma squared by sigma times sqrt(2) sigma leaves 1/sqrt(2). Squared, that is 0.5: X explains half of the sum's variance.

    Why isn't the answer one half?

    Picture two people each tossing a coin for a rupee, and a pot holding their combined winnings. One player's result explains exactly half of the pot's variability, and the other half comes from the other player. Half is the share of variance explained, R squared, and correlation is its square root, so the correlation is 1/sqrt(2), not 1/2. This is the most common slip on the question, and it comes from mixing up the two measures.

    X supplies half of the variance of X + Y, so the correlation is 1/sqrt(2)XYXYVar Xsigma^2Cov(X, Y)0Cov(Y, X)0Var Ysigma^2X row:Cov(X, X+Y)= sigma^2all four cells: Var(X + Y) = 2 sigma^2corr = Cov / (sd X x sd(X+Y))= sigma^2 / (sigma x sqrt(2) sigma)= 1/sqrt(2) = 0.707Share of Var(X + Y) explained by XR^2 = 0.5Unequal variances: corr = 1/sqrt(1 + k),k = Var Y / Var X; k = 4 gives 0.447
    In the covariance box, X's own variance fills one of the two non-zero cells, so X accounts for half of Var(X + Y), and the correlation between X and X + Y is sigma squared divided by sigma times sqrt(2) sigma, which is 1/sqrt(2), about 0.71.
    The relationship
    ρ=Cov⁡(X,X+Y)σX σX+Y=σ2+0σ⋅2 σ=12≈0.707\rho = \frac{\operatorname{Cov}(X, X+Y)}{\sigma_X\,\sigma_{X+Y}} = \frac{\sigma^2 + 0}{\sigma\cdot\sqrt{2}\,\sigma} = \frac{1}{\sqrt 2} \approx 0.707
    Cov(X, X + Y)Var X plus Cov(X, Y), which is sigma squared plus zero
    sigma_{X+Y}the standard deviation of the sum, sqrt(2) sigma
    What it says in wordsCovariance is linear, so split it into pieces; the only surviving piece is X's own variance.

    How does it change if the variances differ?

    Let Var Y be k times Var X. The covariance is still Var X, and Var(X + Y) becomes (1 + k) Var X. The correlation is 1/sqrt(1 + k): the noisier Y is, the less the sum tracks X. At k = 1 you get 0.707; at k = 4, 0.447; at k = 0.25, 0.894. This is exactly the signal-plus-noise model: if a price move is a true signal plus independent noise of equal size, the best-case correlation between your signal and the move is about 0.71.

    Where does this show up on a desk?

    Any time one piece is part of a total. A stock's return is market return plus its own specific return; if the two had equal variance, the stock would correlate 0.71 with the market. The same arithmetic tells you the ceiling on a predictor: if half of tomorrow's move is unpredictable noise, no model can correlate more than 0.71 with it. The limitation is independence; if X and Y are correlated, add 2 Cov(X, Y) to the variance of the sum and Cov(X, Y) to the covariance.

    Where candidates lose it

    The frequent slip is answering one half, confusing the share of variance with the correlation. Correlation is the square root of that share.

    The second loss is writing the standard deviation of X + Y as 2 sigma, adding standard deviations instead of variances, which gives one half again by a different road. Independent variances add; standard deviations do not.

    What the interviewer asks next

    • What is the correlation between X + Y and X - Y?
    • If X and Y have correlation 0.5 and equal variance, what is corr(X, X + Y)?
    • What is the correlation between the first die and the total of two dice?

    Asked at Squarepoint Capital, Desk Quant Analyst Interview, Montreal, 2026 (Wall Street Oasis): There were also 3-4 basic math/stats questions about mean, covariance, correlation, etc.

  9. 100You regress a centred target y on one standardised feature x with no intercept. The sum of x squared is 100 and the sum of x times y is 80. What is the OLS slope, and what is the ridge slope with penalty lambda = 25?Correlation, regression and linear algebraCoreCSCitadel SecuritiesLondon · 2026

    Try it first

    Pick the pair.

    Show the worked solution

    OLS gives 0.8 and ridge gives 0.64. OLS minimises squared error and its slope is the sum of xy over the sum of x squared, 80/100. Ridge adds lambda times the slope squared to the loss, which puts lambda into the denominator: 80/(100 + 25) = 0.64. That is the OLS slope times 100/125 = 0.8, so ridge shrinks the slope towards zero but never to zero.

    Where does lambda end up in the formula?

    Think of a new analyst's forecast that you half trust: you do not discard it, you shade it towards zero, and the less data behind it the more you shade. Ridge does that mechanically. It minimises the squared errors plus lambda times the slope squared; setting the derivative to zero gives b = Sxy/(Sxx + lambda). Penalising the size of the slope acts exactly like adding observations whose x squared totals lambda and whose y is zero, data that say the slope is zero. With 100 of real evidence and 25 of make-believe evidence, the slope is 80/125 = 0.64.

    Ridge divides the slope down; lasso subtracts it to exactly zero0501001502002503000.20.40.60.8penalty lambdaOLS 0.8lambda 25:ridge 0.64lambda 100: 0.40, half of OLSlasso hits 0 at 160ridge: 80 / (100 + lambda)lasso: (80 - lambda/2) / 100
    The ridge slope 80/(100 + lambda) falls from the OLS value 0.8 to 0.64 at lambda 25 and 0.40 at lambda 100 without ever reaching zero, while the lasso slope falls in a straight line and hits exactly zero at lambda 160.
    The relationship
    b^ridge=arg⁡min⁡b∑i(yi−bxi)2+λb2=∑xiyi∑xi2+λ=80125=0.64=0.8×100125\hat b_{\text{ridge}} = \arg\min_b \sum_i (y_i - b x_i)^2 + \lambda b^2 = \frac{\sum x_i y_i}{\sum x_i^2 + \lambda} = \frac{80}{125} = 0.64 = 0.8 \times \frac{100}{125}
    sum x_i y_ithe cross-product of feature and target, 80
    sum x_i^2the sum of squares of the feature, 100
    lambdathe ridge penalty, 25
    What it says in wordsRidge is OLS with lambda added to the sum of squares, so every slope is multiplied by Sxx/(Sxx + lambda).

    Why would you want a slope that is biased towards zero?

    Because a smaller, steadier estimate can be closer to the truth on average. Suppose the true slope is 0.5 and the noise variance is 25. OLS is unbiased but its variance is 25/100 = 0.25. Ridge at lambda 25 has variance 0.16 and a bias of -0.1, so its mean squared error is 0.17. Ridge trades a little bias for a larger cut in variance, and when the signal is weak relative to the noise that trade wins. In this one-feature case the best lambda is noise variance over slope squared, 100, which halves the slope and cuts the error to 0.125. In practice the truth is unknown, so lambda is chosen by cross-validation.

    lambdaSlope on this dataVarianceBias squaredMean squared error
    00.800.25000.00000.2500
    250.640.16000.01000.1700
    1000.400.06250.06250.1250
    Assuming a true slope of 0.5 and noise variance 25, ridge at lambda 25 and 100 has a lower mean squared error than OLS because the drop in variance outweighs the bias it adds.

    How does lasso differ?

    Lasso penalises lambda times the absolute slope instead. In one dimension that subtracts lambda/2 from the cross-product rather than adding to the denominator: (80 - 12.5)/100 = 0.675 at lambda 25, and exactly zero once lambda reaches 160. Ridge scales coefficients down; lasso shifts them down and can set them to exactly zero, which is why lasso selects features and ridge does not. Both penalties depend on the scale of x, which is why the feature must be standardised first; with correlated features, ridge spreads the weight across them while lasso tends to keep one.

    Where candidates lose it

    The fast wrong answer subtracts the penalty from the slope or from the numerator, which is lasso's mechanics, not ridge's. Ridge adds lambda to the sum of squares in the denominator, so the slope is scaled, not shifted.

    The second loss is saying ridge is always better because it has lower variance. It trades variance for bias; if the true slope is large and the data plentiful, shrinking costs more in bias than it saves. Say that lambda is chosen by cross-validation, not by taste.

    What the interviewer asks next

    • What value of lambda halves the OLS slope?
    • With two highly correlated features, how do ridge and lasso split the weight between them?
    • Why must features be standardised before applying a ridge penalty?

    Asked at Citadel Securities, Quantitative Research, London, 2026 (Wall Street Oasis): very detailed and difficult questions about regularisation ridge and lasso

Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.