Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant puzzles, solved step by step

Puzzles
100
Traced to a firm
71
Topics
12
Hard
30
Topic
All topicsLogic and algorithmic reasoning10Conditional probability and Bayes7Counting and combinatorics8Continuous and geometric probability9Correlation, regression and linear algebra9Market making, betting and sizing9Expected value and optimal stopping9Statistics and estimation9Pricing, options and index maths7Games and strategic reasoning8Markov chains and random walks7Mental maths and number sense8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 11–20 of 30 · filtered from 100Clear filters
  1. 032n points are placed independently and uniformly on a circle of circumference 1, with n at least 3. Each point colours the arc between itself and its nearest neighbour. What is the expected total length that gets coloured?Continuous and geometric probabilityHardSusquehanna International GroupLondon · 2026

    Try it first

    Which is closest to the expected coloured length?

    Show the worked solution

    7/18, about 0.389, for every n from 3 upwards. A gap is left uncoloured only when it is longer than both gaps beside it, because then neither endpoint has it as its nearest. For three points that gap is simply the longest of three pieces, which averages 11/18, so 7/18 is coloured. For larger n the same 11/18 comes out, so the answer does not depend on n.

    When is a gap left uncoloured?

    Picture people standing round a circular table, each turning to talk to whoever is closer, left or right. A stretch of table between two people stays silent only if both of them turned away, which means each had a closer person on their other side. A gap is uncoloured exactly when it is longer than both of its neighbouring gaps. A gap coloured from both ends is still coloured once, so the question becomes: what is the expected total length of gaps that are local maxima?

    A gap stays uncoloured only if it is longer than both of its neighbours0.060.040.150.080.100.120.070.080.160.14coloured: 0.57uncoloured: 0.43shorter gap for at least one endpointlonger than both neighboursThree points: the only uncoloured gap is thelongest of three pieces, which averages(1/3)(1 + 1/2 + 1/3) = 11/18Any n: each gap is uncoloured with the sameexpected length, and n of them add to 11/18Expected coloured length7/18 = 0.389Simulated, 40,000 circles: n = 3 gives 0.389,n = 5 gives 0.388, n = 10 gives 0.389
    Each gap is coloured if it is the shorter gap for at least one endpoint and left uncoloured if it is longer than both neighbours; this sample of ten points colours 0.57 of the circle, and the average over all placements is 7/18, about 0.389, for any n of 3 or more.

    How do you get 11/18 for the uncoloured part?

    Start with n = 3, the case you can finish in the room. With three gaps, every gap's two neighbours are the other two gaps, so the only uncoloured gap is the longest one. Three random points cut the circle like a stick broken into three, and the longest of three pieces averages (1/3)(1 + 1/2 + 1/3) = 11/18. So the coloured length is 7/18.

    For larger n, use the fact that the n gaps behave like n independent exponentialA random length whose chance of ending is the same at every instant; waiting times between random arrivals follow it. lengths rescaled to add up to 1, and that the rescaling is independent of the shape. For three unit exponentials X, Y and Z, the expected value of X counted only when X is the largest is 1 - 2/4 + 1/9 = 11/18. Each of the n gaps contributes that, divided by the expected total of n, and the n gaps sum to 11/18 again. The uncoloured share is 11/18 whatever n is, so the coloured share is always 7/18.

    The relationship
    E[coloured]=1−n⋅1n∫0∞xe−x(1−e−x)2 dx=1−1118=718E[\text{coloured}] = 1 - n\cdot\frac{1}{n}\int_0^\infty x e^{-x}(1-e^{-x})^2\,dx = 1 - \frac{11}{18} = \frac{7}{18}
    x e^(-x)a gap's length times its density, in the exponential picture
    (1 - e^(-x))^2the chance both neighbouring gaps are shorter
    1/nrescaling so the n gaps add to a circle of length 1
    What it says in wordsThe expected length of gaps longer than both neighbours is 11/18, and the rest of the circle is coloured.

    Say the check: a seeded simulation of 40,000 random circles gives 0.389 for n = 3, 0.388 for n = 5 and 0.389 for n = 10. The limitation is that the exponential step is a known result you should name, not derive, in an interview; the n = 3 case is the part you prove on the spot.

    Where candidates lose it

    The usual loss is counting gaps instead of measuring them. One gap in three is a local maximum, so candidates answer 2/3 coloured. The uncoloured gaps are selected for being long, which is why their share of length, 11/18, is far above one third.

    The second is double counting a gap that both endpoints colour. It is coloured once. Frame the problem around uncoloured gaps and both mistakes disappear.

    What the interviewer asks next

    • What is the expected number of uncoloured gaps?
    • What if each point colours the arc to its farther neighbour instead?
    • Does the answer change for points on a line segment rather than a circle?

    Asked at Susquehanna International Group, Quantitative Research, London, 2026 (Wall Street Oasis): if n points are placed on a circle and each point colours in the arc to its nearest neighbour

  2. 044A stick of length 1 is broken at three independent uniform points into four pieces. What is the expected length of the longest piece?Continuous and geometric probabilityHardHRHudson River TradingNew York · 2024

    Try it first

    What is the expected length of the longest piece?

    Show the worked solution

    25/48, about 0.521. For a stick broken into n pieces, the expected k-th smallest piece is (1/n)(1/n + 1/(n - 1) + ... ) with k terms. For n = 4 the sorted pieces average 3/48, 7/48, 13/48 and 25/48, which add to 1. The longest is (1/4)(1 + 1/2 + 1/3 + 1/4) = 25/48, more than twice the average piece of 1/4.

    Why is the longest piece so much longer than a quarter?

    Cut a sheet of dough at three random spots and the pieces are rarely even: one is usually a big slab and one a sliver. Random breaks produce uneven pieces, and the longest piece collects the unevenness, so its average sits far above the average piece. The average piece is always 1/4; the question asks about the largest of four correlated lengths, which is an {term('order statistic', 'The k-th smallest value in a sample, for example the minimum, the median or the maximum.')}.

    Four pieces sorted by length: the longest averages 25/48 of the stick3/487/4813/4825/48naive guess: 1/4longest: 25/48 = 0.521Each rank adds one harmonic step, scaled by 1/4shortest1/4 x (1/4)= 3/48 = 0.0625simulated 0.0625second1/4 x (1/4 + 1/3)= 7/48 = 0.1458simulated 0.1457third1/4 x (1/4 + 1/3 + 1/2)= 13/48 = 0.2708simulated 0.2707longest1/4 x (1/4 + 1/3 + 1/2 + 1)= 25/48 = 0.5208simulated 0.5210Simulation: 100,000 sticks, three uniform breaks each.
    Sorted by length, the four pieces of a randomly broken stick average 3/48, 7/48, 13/48 and 25/48 of its length, so the longest piece averages about 0.52, more than twice the naive quarter, and a 100,000-stick simulation agrees to three decimals.

    Where does the harmonic formula come from?

    Start with the shortest piece. The chance that all four pieces exceed x is (1 - 4x) cubed: take x off every piece and the three breaks must fit into the remaining length 1 - 4x. Integrating that from 0 to 1/4 gives an expected shortest piece of 1/16. Then the key fact: the step from each sorted piece to the next adds on average (1/n) times 1 over the number of pieces still longer. After the shortest, three pieces remain longer, so the next piece averages 1/16 + (1/4)(1/3); then add (1/4)(1/2), then (1/4)(1). The longest piece is (1/4)(1/4 + 1/3 + 1/2 + 1) = 25/48.

    The relationship
    E[L(4)]=14(1+12+13+14)=2548≈0.521E[L_{(4)}] = \frac{1}{4}\left(1 + \frac12 + \frac13 + \frac14\right) = \frac{25}{48} \approx 0.521
    L_(4)the longest of the four pieces
    1/4one over the number of pieces
    1 + 1/2 + 1/3 + 1/4the harmonic sum up to the number of pieces
    What it says in wordsThe longest piece averages one quarter of the fourth harmonic number.

    The step rule comes from the fact that the pieces behave like independent exponential lengths scaled to total 1, and the gap between successive minima of exponentials is memoryless. You can name that in the room rather than prove it. The check that the formula is right: the four sorted averages add to exactly 1, and a simulation of 100,000 sticks gives 0.521 for the longest. For n pieces in general, the longest averages (1/n) times the n-th harmonic number, which grows like (ln n)/n.

    Where candidates lose it

    The common loss is answering 1/4, the average piece. The question asks for the average of the largest piece, and the largest of four uneven pieces is usually more than half the stick.

    The second is trying to integrate the maximum directly over the three break points, which gets messy fast. Start from the minimum, use the step rule, and check that the four averages add to 1.

    What the interviewer asks next

    • What is the expected length of the shortest piece for n pieces?
    • What is the probability the four pieces can form a quadrilateral?
    • Break the stick at two points instead. What is the expected longest piece?

    Asked at Hudson River Trading, Campus Algo Dev Interview, New York, 2024 (Wall Street Oasis): I was asked an expected value question involving order statistics.

  3. 045Simplified poker with three cards, A, K and Q: each player antes 1, you are dealt one card and I am dealt another. You may bet 1 or check; if you bet, I call or fold, and if you check the higher card wins the antes. How often should you bluff with the Q, and how often should I call with the K, in equilibrium?Games and strategic reasoningHardOld Mission CapitalNew York · 2022

    Try it first

    How often should I call a bet when I hold the K?

    Show the worked solution

    Bluff with the Q one time in three, and call with the K one time in three. You always bet the A and always check the K; I always call with the A and fold the Q. A Q bluff risks 1 more to win the 2 antes, so I must defend two thirds of hands facing it: the A gives half, the K calling one time in three gives the rest. The game is worth 1/18 per hand to you.

    Which hands are easy, and where is the real decision?

    Clear away the obvious hands first. With the A you always bet, because you win whether I call or fold. With the K, betting only gets called by the A and folds out the Q, which you beat anyway, so you check. As the caller, I always call with the A and always fold the Q. The whole game comes down to two mixed choices: how often you bluff with the Q, and how often I call with the K.

    Equilibrium: bluff the Q one time in three, call with the K one time in threeYour cardeach 1/3AKQBet, alwaysvalue betCheck, alwaysshowdown for 1Bluff 1/3worth -1Check 2/3worth -1Caller facing a betA: call alwaysK: call 1/3, fold 2/3Q: fold alwaysDefends 1/2 + 1/2 x 1/3 = 2/3 vs a bluffWhy 1/3 eachQ bluff: 1/2(-2) + 1/2(c(-2) + (1-c)(+1))equals checking, -1, when c = 1/3K call: (-2 + 2b) / (1 + b)equals folding, -1, when b = 1/3Game value to the bettor: +1/18 per hand
    In equilibrium the bettor always bets the A, always checks the K and bluffs the Q one time in three, while the caller always calls the A, folds the Q and calls with the K one time in three; together the A and the K calls defend two thirds of hands facing a bluff.

    How do the indifference conditions fix both frequencies?

    A teacher who spot-checks homework faces the same logic: check every paper and time is wasted, never check and everyone copies, check at the right rate and copying stops paying. In equilibrium each player mixes at the rate that makes the other indifferent between their two options. Your Q loses 1 by checking. Bluffing loses 2 against my A, and against my K loses 2 if I call and wins 1 if I fold. The two are equal only when I call with the K one time in three. My K loses 1 by folding; calling loses 2 against your A and wins 2 against a bluff, which is worth -1 only when you bluff one time in three.

    The relationship
    12(−2)+12(c(−2)+(1−c)(1))⏟Q bluffs=−1⇒c=13−2+2b1+b⏟K calls=−1⇒b=13\underbrace{\tfrac12(-2) + \tfrac12\big(c(-2) + (1-c)(1)\big)}_{\text{Q bluffs}} = -1 \Rightarrow c = \tfrac13 \qquad \underbrace{\frac{-2 + 2b}{1 + b}}_{\text{K calls}} = -1 \Rightarrow b = \tfrac13
    chow often the caller calls with the K
    bhow often the bettor bluffs with the Q
    -1the payoff of the alternative: checking the Q or folding the K, losing the ante
    What it says in wordsEach frequency is set so the opponent's two choices are worth the same.

    Check against the pot-odds rule. A bluff risks 1 extra to win the 2 antes, so the caller must defend 2/(2 + 1) = 2/3 of the hands facing it; the A already covers half, and the K calling one time in three covers the other sixth. That two thirds is the number people misremember as the K's calling rate. Put the strategies together and the bettor, who acts first with more information about their own hand, earns 1/18 of a chip per hand. The limitation: with a bigger bet or more cards the ratios change, but the method, indifference on both sides, carries over.

    Where candidates lose it

    The common loss is setting the K's calling rate to two thirds. Two thirds is the total defence the caller needs against a bluff, and the A already provides half of it, so the K calls only one time in three.

    The second is never bluffing because the Q cannot win a showdown. A player who never bluffs lets the caller fold every K to a bet, and the A's bets stop earning. The bluff is what gets the A paid.

    What the interviewer asks next

    • What is the value of the game to each player?
    • The bet size doubles to 2. How do the bluffing and calling frequencies change?
    • Now the caller may also bet after a check. What changes?

    Asked at Old Mission Capital, Quantitative Research, New York, 2022 (Wall Street Oasis): Asking to find the game theory optimal strategy in a simplified poker game

  4. 047Two assets' daily returns are negatively correlated within every month, yet their monthly returns are positively correlated across the year. How can that happen? Build a small numerical example.Correlation, regression and linear algebraHardSCSquarepoint CapitalMontreal · 2024

    Try it first

    Both assets share a drift that changes from month to month, plus daily noise that is negatively correlated. What happens to the correlation as you sum more days into one return?

    Show the worked solution

    A drift shared by both assets for the whole month can outweigh daily noise that moves them in opposite directions. Within a month the drift is constant, so only the noise shows and the correlation is negative. Summed over 21 days the drift's covariance grows with 21 squared but the noise's only with 21. With noise correlation -0.5 and drift standard deviation 0.3% a day, monthly correlation is +0.48.

    What does a three-month example look like?

    Picture two shops in the same market street. On any one day, a customer who buys from one did not buy from the other, so their daily takings move against each other. But in festival months the whole street is busy and in the rains the whole street is quiet, so their monthly takings rise and fall together. Correlation at one horizon says nothing on its own about another, because a slow common factor and fast opposing noise can sit in the same data.

    Make it numerical with five-day months. In month 1 both assets drift at -0.8% a day, in month 2 at +0.2%, in month 3 at +1.2%. On top, asset A gets daily noise of +0.6, -0.3, 0, +0.3, -0.6 and asset B gets -0.3, +0.3, 0, -0.3, +0.3, which move in opposite directions. Within each month the correlation is -0.95. Summed over each month, the noise nets to zero, so both assets return -4%, +1% and +6%: identical, a monthly correlation of +1. Even all 15 days pooled show +0.71, because the month-to-month swing in drift is larger than the noise.

    Inside each month the points fall; across months the clusters climb-1%0+1%+2%-1%0+1%asset A daily returnasset B daily returnMonth 1Month 2Month 3Five-day monthsDriftWithin rMonth A, B-0.8%-0.95-4%, -4%+0.2%-0.95+1%, +1%+1.2%-0.95+6%, +6%Noise nets to zero inside each monthCorrelation by frequencyWithin each month: -0.95All 15 days pooled: +0.71Monthly returns: +1.00
    Within each five-day month the daily returns of the two assets slope downward with a correlation of -0.95, but the monthly drifts of -0.8%, +0.2% and +1.2% a day are shared, so the three cluster centres rise together and the monthly returns correlate at +1.

    Why does summing more days push the correlation positive?

    Write each daily return as the month's drift m plus noise. Over n days the drift adds up to n times m, while the noise adds up to a sum of n separate shocks. The drift's contribution to covariance scales with n squared, the noise's only with n, so the longer the horizon the more the shared drift wins. Take noise with standard deviation 1% a day and a within-month correlation of -0.5, and a drift whose standard deviation across months is 0.3% a day. For a 21-day month the drift adds 39.69 to the covariance and the noise subtracts 10.5, for a monthly correlation of 29.19/60.69 = 0.48.

    The relationship
    ρ(n)=n2σm2+n cn2σm2+n σ2ρ(21)=39.69−10.539.69+21=0.48\rho(n) = \frac{n^2\sigma_m^2 + n\,c}{n^2\sigma_m^2 + n\,\sigma^2} \qquad \rho(21) = \frac{39.69 - 10.5}{39.69 + 21} = 0.48
    nnumber of days summed into one return
    \sigma_mstandard deviation of the shared daily drift across months, 0.3%
    cdaily noise covariance within a month, -0.5
    \sigmadaily noise standard deviation, 1%
    What it says in wordsShared drift covariance grows with the square of the horizon, independent noise covariance only in proportion to it.
    Drift covariance grows with days squared, noise covariance only with days+39.69Drift-10.5Noise+29.19Monthly cov21-day month, in % squared21 x 21 x 0.3^2 = 39.69; 21 x (-0.5) = -10.5Variance 60.69, so monthly r = 29.19/60.69 = 0.48-0.4+0.401 day: -0.385 days: -0.0321 days: +0.48flips at 5.6 daysdays summed into one return11121Correlation of summed returnsr(n) = (0.09 n - 0.5) / (0.09 n + 1)
    For a 21-day month the shared drift adds 39.69 to the covariance and the opposing daily noise takes away 10.5, so monthly returns correlate at +0.48, and the correlation of summed returns crosses from negative to positive at about 5.6 days.

    The crossover sits where n times 0.09 equals 0.5, about 5.6 days, so even weekly returns of five days would still show a slightly negative correlation, -0.03. A hedge sized on daily correlation can therefore fail at a monthly horizon, which is why a desk measures correlation at the frequency it actually holds risk. The other mechanisms worth naming are mean reversion in the spread between the two assets and stale prices that lag by a day; both also make correlation depend on frequency. The limitation of the example is the assumption that drift is constant inside a month and noise is independent from day to day.

    Where candidates lose it

    The common loss is saying it is impossible, or that it must be a data error, because correlation feels like a fixed property of two assets. It is a property of two assets at a horizon, and the interviewer wants the decomposition into a slow shared part and a fast opposing part.

    The second is a hand-waved answer with no numbers. Build the five-day example in a minute, state that drift covariance scales with n squared and noise with n, and the explanation becomes checkable.

    What the interviewer asks next

    • What would make daily correlation positive but monthly correlation negative?
    • How would you estimate the shared monthly drift from daily data?
    • A pairs trader hedges at the daily beta and holds for a month. What goes wrong?

    Asked at Squarepoint Capital, Hedge Fund, Montreal, 2024 (Wall Street Oasis): correlation can be negative intra-month but positive across a year, how?

  5. 053Five assets each have unit variance, and every pair has correlation 0.4. What are the eigenvalues of the correlation matrix, and what share of total variance does the first principal component explain?Correlation, regression and linear algebraHardJump TradingPudong Xinqu · 2023

    Try it first

    Before any algebra: what share of variance does the first component explain?

    Show the worked solution

    One eigenvalue of 2.6 and four of 0.6, so the first principal component explains 52%. Write the matrix as 0.6 times the identity plus 0.4 times a matrix of ones. The all-ones vector is an eigenvector with eigenvalue 0.6 + 5 x 0.4 = 2.6; any vector whose weights sum to zero is killed by the ones matrix and has eigenvalue 0.6. The trace check: 2.6 + 4 x 0.6 = 5.

    What structure should you spot before touching a determinant?

    Think of five students whose marks all move together when the paper is hard, plus their own good and bad days. There is one shared shock and five private ones. An equicorrelation matrix is exactly that: R = (1 - rho) I + rho J, where J is the matrix of all ones, so its eigenvectors are those of J and you never need a characteristic polynomial. J sends the all-ones vector to 5 times itself and sends any vector whose entries sum to zero to zero. Those two facts give every eigenvalue.

    One market factor and four equal leftovers: the eigenvalues of R10.40.40.40.40.410.40.40.40.40.410.40.40.40.40.410.40.40.40.40.41Correlation matrix Rtrace = 5 = sum of eigenvaluesaverage eigenvalue 12.6PC152%0.6PC212%0.6PC312%0.6PC412%0.6PC512%PC1: equal weights, the market1 + (n - 1) x 0.4 = 2.6, 52% of variancePC2 to PC5: weights summing to zero1 - 0.4 = 0.6 each, 12% each
    The 5 by 5 matrix with 0.4 off the diagonal has one eigenvalue of 2.6, carried by the equal weight portfolio and explaining 52% of the variance, and four eigenvalues of 0.6, carried by long short combinations and explaining 12% each.

    How do the eigenvalues fall out, and how do you check them?

    Apply R to the all-ones vector: each row sums to 1 + 4 x 0.4, so the equal weight portfolio has eigenvalue 1 + (n - 1) rho = 2.6. Apply R to any vector with weights summing to zero, such as long asset 1 and short asset 2: the rho J part vanishes and only (1 - rho) = 0.6 is left, and there are four independent such vectors. The eigenvalues must add to the trace, the sum of the diagonal, which is 5: 2.6 + 2.4 = 5.

    The relationship
    R=(1−ρ)I+ρ 11⊤λ1=1+(n−1)ρ=2.6,λ2..5=1−ρ=0.6,λ1n=52%R = (1-\rho)I + \rho\,\mathbf{1}\mathbf{1}^{\top} \qquad \lambda_1 = 1+(n-1)\rho = 2.6,\quad \lambda_{2..5} = 1-\rho = 0.6,\quad \frac{\lambda_1}{n} = 52\%
    rhothe common pairwise correlation, 0.4
    nthe number of assets, 5
    1 1^Tthe all-ones matrix J
    What it says in wordsA common correlation creates one large factor for the average and leaves every long short combination with the same small variance.

    What does the answer say about a real portfolio?

    The first component is the market: equal weights, and its share rises towards rho as you add assets. With 50 assets at the same correlation the first eigenvalue is 1 + 49 x 0.4 = 20.6, 41.2% of the total, while each of the other 49 stays at 0.6. Diversification removes the private shocks but never the common one. The same formula gives a limit: the smallest eigenvalue 1 - rho is always fine, but 1 + (n - 1) rho must stay positive, so five assets cannot all share a correlation below -0.25.

    Where candidates lose it

    The loss is trying to expand a 5 by 5 determinant by hand. It is slow, error prone and signals that you did not see the structure. The interviewer is waiting for identity plus ones matrix.

    The second trap is reading 40% as the explained share because the correlation is 0.4. The share is (1 + (n - 1) rho)/n, which is 52% here and only approaches rho as n grows.

    What the interviewer asks next

    • What is the most negative common correlation five assets can have?
    • What are the eigenvectors of the four 0.6 eigenvalues, and why are they not unique?
    • If one asset is removed, what share does the first component explain?
    • How would you spot a second factor, such as a sector, in the eigenvalues?

    Asked at Jump Trading, Prop Trading, Pudong Xinqu, 2023 (Wall Street Oasis): Some very difficult linear algebra questions about PCA and eigenvalues

  6. 054With interest rate r and volatility sigma, check which of these satisfy the Black-Scholes equation: V = S, V = K e^(-r(T-t)), and V = S squared. Explain what the ones that pass are as trades, and fix the one that fails.Pricing, options and index mathsHardQuant researchOptions market making

    Try it first

    Which candidates pass?

    Show the worked solution

    V = S and V = K e^(-r(T-t)) satisfy it; V = S squared does not. The first is the stock itself and the second is a zero-coupon bond paying K at T, both traded assets that must earn r. S squared has gamma 2 and leaves (r + sigma squared) S squared unbalanced. Multiplying by e^((r + sigma squared)(T-t)) fixes it, which is the price of a claim paying S squared at expiry.

    What is the equation actually saying?

    Think of a household budget rule that any fair arrangement must obey: over one day, what you hold must earn the same as the same money in a savings account, once the risk has been hedged away. The Black-Scholes equation says that for a delta-hedged position, time decay plus the gamma term plus the financing of the hedge equals r times the value. Written in {term('greeks', 'Theta is the change in value with time, delta with the stock price, and gamma is the change in delta with the stock price.')}, it is theta + half sigma squared S squared gamma + r S delta = r V. A candidate price passes only if its greeks balance that line.

    The relationship
    ∂V∂t+12σ2S2∂2V∂S2+rS∂V∂S−rV=0\frac{\partial V}{\partial t} + \tfrac12\sigma^2 S^2 \frac{\partial^2 V}{\partial S^2} + rS\frac{\partial V}{\partial S} - rV = 0
    dV/dttheta, the change in value as time passes
    d2V/dS2gamma, how fast delta changes
    dV/dSdelta, the hedge ratio
    rthe interest rate
    What it says in wordsA hedged position's decay, convexity and financing must add up to exactly the interest the money would earn.
    Substitute each candidate into theta + half sigma^2 S^2 gamma + r S delta - r VCandidate VThetaDeltaGammaLeft side of equationPasses?Sthe stock010rS - rS = 0yesK e^(-r tau)a zero-coupon bondrV00rV - rV = 0yesS^2not a price02S2(r + sigma^2) S^2noS^2 ff = e^((r + sigma^2) tau)-(r + sigma^2) V2S f2 f0yestau = T - t. The failing S^2 leaves a surplus; the time factor f is exactly what cancels it, pricing a claim paying S_T^2.At S = 100, r = 5%, sigma = 20%, one year: the S^2 claim is worth 10,942, not 10,000.
    Substituting each candidate's theta, delta and gamma, the stock and the zero-coupon bond balance the equation exactly, S squared leaves a surplus of (r + sigma squared) S squared, and S squared times e^((r + sigma squared)(T - t)) balances it again.

    Why do the two that pass make sense as trades?

    Anything that is itself a traded, self-financing asset must satisfy the equation, because the equation is only the statement that no hedged position earns more than r. V = S is just holding the stock: delta 1, no gamma, no decay, and the financing term rS matches rV. V = K e^(-r(T-t)) is a zero-coupon bond: it does not depend on S at all, and its value grows at exactly r as it approaches T. The stock and the bond are also the two pieces of the call price formula, which is why the check is worth a minute.

    Why does S squared fail, and how do you repair it?

    S squared has gamma 2, so the half sigma squared S squared gamma term adds sigma squared S squared, the delta term adds 2rS squared, and subtracting rV leaves (r + sigma squared) S squared with nothing to cancel it. A convex payoff gains from every move, so a fair price for it has to decay over time to pay for that gain, and S squared on its own has no decay. Try V = S squared times f(t): the equation forces f' = -(r + sigma squared) f, so the price of a claim paying S squared at T is S squared e^((r + sigma squared)(T - t)). At S = 100, r = 5%, sigma = 20% and one year, that is about 10,942, not 10,000, and the extra is the value of volatility.

    Where candidates lose it

    Candidates often say every function of S and t is a solution, or differentiate correctly and then fail to say what the passing solutions are. The question asks for the trades: the stock and a bond. Naming them turns a calculus check into finance.

    The second trap is the sign of theta for the bond. Its value rises as t approaches T, so theta is +rV; getting that sign wrong makes the bond appear to fail.

    What the interviewer asks next

    • Which power of S, S to the a, satisfies the equation with no time factor?
    • What is the price today of a claim paying log S at expiry?
    • Why does the drift of the stock not appear anywhere in the equation?
  7. 056You roll a fair die until each of 2, 4 and 6 has appeared at least once. Given that the last even number to make its first appearance was 2, what is the probability that the very first roll was a 1? Why is it not 1/5?Conditional probability and BayesHardSCSquarepoint CapitalLondon · 2026

    Try it first

    Given that 2 was the last even to show up, what is the chance the first roll was a 1?

    Show the worked solution

    1/6, the same as with no information. An odd first roll says nothing about the order in which 2, 4 and 6 first appear, so it is independent of 2 finishing last. A first roll of 4 or 6 raises the chance 2 is last from 1/3 to 1/2, so conditioning on that ending shifts weight onto 4 and 6, which rise to 1/4 each. The odd faces keep 1/6 each; 1/5 wrongly spreads the weight evenly.

    Why does the ending tell you anything about the start?

    Suppose you hear that a friend reached a party last. That makes it a little more likely they left home late, because leaving late and arriving last go together. It says nothing about whether they wore a blue shirt, which has no bearing on arrival order. Conditioning on an outcome reweights every starting state by how likely that state makes the outcome, and a state that does not affect the outcome keeps its original probability. Here the outcome is 2 finishing last among the evens; the question is which first rolls make that more or less likely.

    The first roll changes how likely 2 is to finish lastFirst roll1/21, 3 or 5then 2 last: 1/3joint 1/2 x 1/3 = 1/61/34 or 6then 2 last: 1/2joint 1/3 x 1/2 = 1/61/62then 2 last: 0joint 0P(2 last) = 1/6 + 1/6 + 0 = 1/3Given 2 finished last, the first roll wasnaive 1/51/611/631/651/441/4602face on the first roll
    A first roll of 1, 3 or 5 leaves 2 a one in three chance of finishing last, a first roll of 4 or 6 raises it to one in two, and a first roll of 2 makes it impossible, so given that 2 finished last the odd faces are worth 1/6 each and 4 and 6 are worth 1/4 each.

    How do the numbers work out with Bayes?

    Odd rolls never change which new even appears next, so only the order of first appearances matters, and without information it is a random ordering of three: 2 is last with chance 1/3. If the first roll is 4, then 2 and 6 are left to race, and each is equally likely to show first, so 2 ends last with chance 1/2. Now weigh: each face has prior 1/6. The joint chance of first roll 1 and 2 last is 1/6 x 1/3 = 1/18; of first roll 4 and 2 last, 1/6 x 1/2 = 1/12. The total is 1/3, so first roll 1 has posterior (1/18)/(1/3) = 1/6 and first roll 4 has (1/12)/(1/3) = 1/4.

    The relationship
    P(first=1∣2 last)=16⋅1313=16P(first=4∣2 last)=16⋅1213=14P(\text{first}=1 \mid \text{2 last}) = \frac{\tfrac16\cdot\tfrac13}{\tfrac13} = \frac16 \qquad P(\text{first}=4 \mid \text{2 last}) = \frac{\tfrac16\cdot\tfrac12}{\tfrac13} = \frac14
    1/6the prior chance of any face on the first roll
    1/3the chance 2 is last when the first roll is odd, and also overall
    1/2the chance 2 is last when 4 or 6 is already seen
    What it says in wordsAn odd first roll is independent of the ending and keeps 1/6; the even faces 4 and 6 absorb the weight that 2 loses.

    Where does the 1/5 intuition go wrong?

    It treats the information as simply ruling out one face and renormalising the rest. Ruling out an outcome and conditioning on an event are the same thing only when every remaining outcome makes the event equally likely, and here they do not. A check: the posteriors 1/6, 1/6, 1/6, 1/4, 1/4 and 0 add to 1, while five faces at 1/5 would give 4 and 6 the same weight as 1. On a desk this is the error of reading a trade's outcome as if it said nothing about which signal triggered it.

    Where candidates lose it

    Nearly everyone's first answer is 1/5. The interviewer is not testing the arithmetic; the question itself says it is not 1/5 and asks you to explain why, so an answer that only produces 1/6 without the reason loses most of the credit.

    The second trap is getting lost in the odd rolls. They can be ignored completely, because they never change which even appears next. Say that early and the problem shrinks to the order of three numbers.

    What the interviewer asks next

    • Given that 2 finished last, what is the probability the first roll was a 4?
    • What is the expected number of rolls until all three evens have appeared?
    • Given that 2 finished last, what is the probability the first even to appear was 4?

    Asked at Squarepoint Capital, Quant Research Intern Interview, London, 2026 (Wall Street Oasis): why is the probability of seeing a 1 on our first roll, given that we end on a 2, not 1/5

  8. 059We play a coin game. I pick a sequence of three heads or tails, you then pick a different sequence after seeing mine, and we flip a fair coin until one of the two sequences appears; whoever's comes first wins. I pick HHH. What do you pick, and how often do you win?Games and strategic reasoningHardQuant tradingProp trading firms

    Try it first

    Which reply to HHH is best?

    Show the worked solution

    Pick THH; you win 7 times in 8. HHH can only win if the first three flips are all heads, which has chance 1/8. In any other run, the first HHH is preceded by a tail, and that tail with the next two heads spells THH, which is completed one flip before HHH. So THH wins every game except the one that opens with three heads.

    Why is this not a fair race between two 1/8 sequences?

    Think of two runners on the same track where one always starts one step ahead of the other on the only route to the finish. Their speeds are identical but the race is not even. In a race between patterns, what matters is not how often each appears but which one tends to appear first, and that depends on how the patterns overlap. THH is built from HHH's own first two heads with a tail placed in front, so it ambushes HHH whenever HHH has not already won at the start.

    Every HHH that is not at the very start has THH inside it, one flip earlierHHH wins only here:HHHflips 1 to 3 all headschance 1/2 x 1/2 x 1/2 = 1/8Any other run, e.g.T1H2T3T4H5H6H7THH done on flip 6HHH only on flip 7: too lateUnless the game opens H H H, the first HHH has a T right before it.That T and the next two heads are THH, completed one flip before HHH.So THH wins 1 - 1/8 = 7/8 of the time.
    HHH wins only when the first three flips are heads, a 1/8 chance; in every other run the first HHH is preceded by a tail, so THH is completed one flip earlier and wins the remaining 7/8 of games.

    How do you prove 7/8 without a Markov chain?

    Look at the first time HHH appears. If it does not start at flip 1, the flip immediately before it must be a tail, because otherwise an earlier HHH would already have appeared. That tail plus the first two heads of the HHH is THH, finished one flip before HHH. So HHH wins only if flips 1 to 3 are heads, chance 1/8, and THH wins otherwise: 7/8. The proof is one sentence, and interviewers want to hear it rather than a transition matrix.

    What is the general lesson for the second mover?

    This is Penney's gameA coin sequence race in which the second player, choosing after seeing the first, can always pick a sequence that wins more than half the time., and the second player always has an edge because the winning relation among three-flip sequences is not transitive: every sequence has another that beats it. The recipe: take the opponent's first two flips, and put in front of them the opposite of the opponent's second flip. Against HHH that gives THH at 7/8; against HTH it gives HHT at 2/3. In trading terms, a strategy that looks as good as any other in isolation can still lose systematically to one designed around it.

    Where candidates lose it

    The trap is answering that every sequence has probability 1/8, so the game is fair, or picking TTT because it has nothing in common with HHH. Both treat the race as independent draws of three flips rather than a stream where patterns overlap.

    The second trap is reaching for a four-state Markov chain and running out of time. The tail-before-the-run argument settles it in one sentence; set up the chain only if asked about a harder pair.

    What the interviewer asks next

    • I pick HTH. What do you pick, and how often do you win?
    • What is the expected number of flips to see HHH, and to see THH?
    • Why can no three-flip sequence be the best first choice?
  9. 064You may roll a fair die up to three times. After each roll you either stop and are paid the face showing, or roll again; if you reach the third roll you must take it. What is your optimal stopping rule, and what is the game worth?Expected value and optimal stoppingHardRCRBC Capital MarketsToronto · 2025

    Try it first

    On the first roll you see a 4. What do you do?

    Show the worked solution

    Keep a 5 or 6 on the first roll, a 4, 5 or 6 on the second, and take whatever the third gives; the game is worth 14/3, about 4.67. Work backwards. The last roll is worth 3.5. With two rolls left, keep anything above 3.5: (4 + 5 + 6)/6 + (1/2)(3.5) = 4.25. With three, keep anything above 4.25: (5 + 6)/6 + (2/3)(4.25) = 14/3.

    Why start from the last roll?

    Think of house hunting with three viewings booked: whether to accept the first flat depends on what the remaining viewings are likely to offer, and you only know that once you know how you would behave at the last one. The value of continuing at any point is defined by what you would do later, so the only roll whose value you know outright is the last one, and every earlier decision is built on it. That is backward induction, and the interviewer wants to hear the words before any numbers.

    Solve from the last roll up: keep a face only if it beats rolling onRoll 12 rolls left after it123456rerollkeepbeats 4.25?14/3 = 4.67Roll 21 roll left after it123456rerollkeepbeats 3.50?17/4 = 4.25Roll 3the last roll123456no choice7/2 = 3.50feedsfeedsRoll 2 value: (4 + 5 + 6)/6 + (3/6) x 3.5 = 4.25Roll 1 value: (5 + 6)/6 + (4/6) x 4.25 = 14/3
    Solving from the last roll upwards, the third roll is worth 3.5, so the second roll keeps 4, 5 or 6 and is worth 4.25, so the first roll keeps only 5 or 6 and the whole game is worth 14/3, about 4.67.

    How do the values build up?

    On the last roll you take the face, worth 3.5. On the second roll, stop if the face beats 3.5, which means 4, 5 or 6; otherwise you get 3.5 from the last roll. The value is the average of the faces you keep plus the chance you continue times the value of continuing: (4 + 5 + 6)/6 + (3/6)(3.5) = 4.25. On the first roll, the bar to beat is now 4.25, so only 5 and 6 are kept: (5 + 6)/6 + (4/6)(4.25) = 11/6 + 17/6 = 14/3, about 4.667.

    The relationship
    V1=3.5,Vn+1=16∑f=16max⁡(f,Vn)  ⇒  V2=174,V3=143V_1 = 3.5,\qquad V_{n+1} = \frac16\sum_{f=1}^{6}\max(f, V_n) \;\Rightarrow\; V_2 = \tfrac{17}{4},\quad V_3 = \tfrac{14}{3}
    V_nthe value of the game with n rolls still available
    max(f, V_n)keep the face f if it beats rolling on, otherwise take the value of continuing
    What it says in wordsEach extra roll is worth the average of the better of the face and the value of carrying on.

    What does the common shortcut cost, and where does this lead?

    The shortcut is to keep anything above the single-roll average of 3.5 at every stage. On the first roll that keeps a 4, which gives 4.625 instead of 4.667. The threshold rises with the number of rolls left, because each spare roll is an option, and an option is worth more the longer it lives. With six rolls the value is 5.27, and with many rolls it approaches 6, since you can wait for a six. The same structure prices an American option: exercise early only when the payoff beats the value of holding on.

    Where candidates lose it

    The trap is the fixed threshold: stopping on 4 at the first roll because 4 beats 3.5. It ignores that the comparison is with the value of continuing, which is 4.25 with two rolls left, not 3.5.

    The second loss is computing forwards, trying to enumerate all paths from the first roll. Say backward induction, solve the last roll, and build up; three lines of arithmetic do the whole job.

    What the interviewer asks next

    • What is the game worth with four rolls?
    • You now pay Rs 1 for each reroll. How does the rule change?
    • If you are paid the square of the final face, what is the first-roll rule?

    Asked at RBC Capital Markets, Quantitative Trading, Toronto, 2025 (Wall Street Oasis): Best way to maximize EV across 3 chosen dice rolls (can choose to continue or not).

  10. 065The correlation between X and Y is 0.2, and the correlation between Y and Z is 0.5. What is the full range of possible values for the correlation between X and Z?Correlation, regression and linear algebraHardTower Research CapitalNew York · 2019

    Try it first

    Which statement about corr(X, Z) is right?

    Show the worked solution

    Anywhere from about -0.75 to 0.95. The correlation matrix must be positive semidefinite, which bounds the third correlation at 0.2 x 0.5 plus or minus sqrt((1 - 0.2^2)(1 - 0.5^2)), that is 0.1 plus or minus 0.849. Geometrically, correlations are cosines of angles: X sits 78.5 degrees from Y and Z sits 60 degrees from Y, so X and Z are between 18.5 and 138.5 degrees apart.

    Why does knowing two correlations restrict the third at all?

    Think of three towns on a map. If A is close to B and B is close to C, then A cannot be far from C; if B is only loosely near both, A and C could be almost anywhere. Correlations behave like distances in disguise: each is the cosine of the angle between two returns viewed as vectors, and angles obey a triangle rule. X makes an angle of {TH_A:.1f} degrees with Y, because its cosine is 0.2, and Z makes 60 degrees with Y. The angle between X and Z is therefore at least the difference and at most the sum.

    Correlations are cosines of angles, so the third one is boxed inYX: 78.5 deg from YZ: 60 deg, same sideZ: 60 deg, other sidecos 78.5 deg = 0.2cos 60 deg = 0.5Same side: 78.5 - 60 = 18.5 deg, corr 0.95Other side: 78.5 + 60 = 138.5 deg, corr -0.75-101centre 0.2 x 0.5 = 0.1-0.750.95Possible values of corr(X, Z)0.1 plus or minus sqrt(0.96 x 0.75)= 0.1 plus or minus 0.849Anything outside makes the correlationmatrix impossible: a negative varianceEven a negative correlation is allowed
    With X at 78.5 degrees from Y and Z at 60 degrees from Y, the angle between X and Z ranges from 18.5 to 138.5 degrees, so their correlation can be anything from -0.75 to 0.95, centred on 0.1.

    How do you get the bound algebraically?

    Any valid correlation matrix must be positive semidefiniteEvery portfolio built from the variables has a variance of zero or more; for a correlation matrix this means its determinant and all leading minors are non-negative., because a portfolio cannot have negative variance. For three variables with correlations a, b and c, the condition is 1 - a^2 - b^2 - c^2 + 2abc of at least zero. Treating that as a quadratic in c gives c = ab plus or minus sqrt((1 - a^2)(1 - b^2)), so the third correlation lies in an interval centred on the product of the other two. With a = 0.2 and b = 0.5 that is 0.1 plus or minus sqrt(0.72), or -0.7485 to 0.9485.

    The relationship
    ρXZ∈[ρXYρYZ−(1−ρXY2)(1−ρYZ2), ρXYρYZ+(1−ρXY2)(1−ρYZ2)]=[−0.749, 0.949]\rho_{XZ} \in \Big[\rho_{XY}\rho_{YZ} - \sqrt{(1-\rho_{XY}^2)(1-\rho_{YZ}^2)},\ \rho_{XY}\rho_{YZ} + \sqrt{(1-\rho_{XY}^2)(1-\rho_{YZ}^2)}\Big] = [-0.749,\ 0.949]
    rho_XY0.2, the correlation of X and Y
    rho_YZ0.5, the correlation of Y and Z
    the square rootthe room left over after the parts of X and Z explained by Y
    What it says in wordsThe third correlation is the product of the two given ones, plus or minus how much of X and Z is unexplained by Y.

    What does the centre value 0.1 mean, and when is the sign forced?

    Split X and Z each into a part explained by Y and a leftover. The explained parts always contribute 0.2 x 0.5 = 0.1; the leftovers can be correlated however you like, and they can move the total by up to 0.849 either way. So 0.1 is the answer only if the leftovers are uncorrelated. The sign of corr(X, Z) is forced positive only when the two given correlations are strong, specifically when their squares add to more than 1; with 0.8 and 0.7 the range is about 0.13 to 0.99. On a desk this is why two hedges that each track an index only loosely say almost nothing about each other.

    Where candidates lose it

    The two fast wrong answers are 0.1, from multiplying, and must be positive, from assuming correlation is transitive. Both treat correlation like a chain of causes rather than a geometry.

    The second loss is reaching the determinant condition and stalling on the algebra. Lead with the angle picture: arccos 0.2 is about 78.5 degrees, arccos 0.5 is 60, and the bounds are the cosines of their sum and difference.

    What the interviewer asks next

    • If corr(X, Y) = 0.8 and corr(Y, Z) = 0.7, can corr(X, Z) be negative?
    • What is the most negative common correlation three variables can share?
    • Given corr(X, Y) and corr(Y, Z), what value of corr(X, Z) makes X and Z uncorrelated once Y is controlled for?

    Asked at Tower Research Capital, Prop Trading, New York, 2019 (Wall Street Oasis): What if the correlation between X and Y is 0.2 and the correlation between Y and Z is 0.5.

← PreviousPage 2 of 3
  1. 1
  2. 2
  3. 3
Next →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.