Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant puzzles, solved step by step

Puzzles
100
Traced to a firm
71
Topics
12
Hard
30
Topic
All topicsLogic and algorithmic reasoning10Conditional probability and Bayes7Counting and combinatorics8Continuous and geometric probability9Correlation, regression and linear algebra9Market making, betting and sizing9Expected value and optimal stopping9Statistics and estimation9Pricing, options and index maths7Games and strategic reasoning8Markov chains and random walks7Mental maths and number sense8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 11–20 of 20 · filtered from 100Clear filters
  1. 055In how many ways can three positive integers, in order, sum to 10? To 11? Give the general formula for any total n of at least 3.Counting and combinatoricsWarm upOld Mission CapitalNew York · 2018

    Try it first

    How many ordered triples of positive integers sum to 10?

    Show the worked solution

    36 for 10, 45 for 11, and (n - 1)(n - 2)/2 in general. Write n as a row of n stars. Splitting it into three positive parts means placing two bars in two different gaps among the n - 1 gaps between stars. Each choice gives exactly one ordered triple, so the count is n - 1 choose 2: 9 choose 2 = 36 and 10 choose 2 = 45.

    Why turn the sum into a row of stars?

    Think of ten sweets in a line to be shared among three children in order, each getting at least one. You do not need to decide amounts; you only need to decide where to cut the line. Every ordered split of n into three positive parts is exactly one choice of two cut points among the n - 1 gaps between items, and every choice of two gaps gives a valid split. That one-to-one match is the whole argument, and it is what the interviewer wants to hear you state.

    Ten stars, nine gaps: pick 2 gaps for the bars123456789gap3433 + 4 + 3 = 10Two bars, two different gaps out of n - 1:C(9, 2) = 36 and C(10, 2) = 45Count for each total n:1n=33n=46n=510n=615n=721n=828n=936n=1045n=11
    Ten stars have nine gaps between them; putting bars in two different gaps, here gaps 3 and 7, splits the stars into 3, 4 and 3, so the ordered triples summing to 10 number 9 choose 2, which is 36, and those summing to 11 number 45.

    What if the interviewer meant something slightly different?

    Ask two quick questions before you answer: are zeros allowed, and does order matter. The word numbers hides three different questions, and each has a different count. If zeros are allowed, add one to each part first so they become positive and sum to n + 3: for 10 that is 12 choose 2 = 66. If order does not matter, the count for 10 drops to 8 unordered triples, and for 11 to 10, which you would list rather than compute. Asking which one is wanted takes five seconds and is part of the answer.

    The relationship
    #{(a,b,c)≥1: a+b+c=n}=(n−12)=(n−1)(n−2)2(92)=36,  (102)=45\#\{(a,b,c)\ge 1:\ a+b+c=n\} = \binom{n-1}{2} = \frac{(n-1)(n-2)}{2} \qquad \binom{9}{2}=36,\ \ \binom{10}{2}=45
    n - 1the number of gaps between n stars
    2the number of bars needed to make three parts
    What it says in wordsChoose two of the gaps between the stars; each choice is one ordered triple.

    How do you check 36 without the formula?

    Fix the first number and count the rest. If the first part is a, the other two must sum to 10 - a, which can be done in 9 - a ordered ways. For a from 1 to 8 that is 8 + 7 + ... + 1 = 36. The counts for each total are the triangular numbers 1, 3, 6, 10 and so on, which is the same formula read another way.

    Where candidates lose it

    The fast wrong answer comes from listing unordered triples such as 1, 1, 8 and 2, 3, 5, getting 8, and not noticing the question counts order. The opposite slip is counting zeros and getting 66.

    Both are avoided by one clarifying question at the start. Then give the stars and bars picture in a sentence, because the interviewer's next question is usually four or five parts.

    What the interviewer asks next

    • How many ways can four positive integers sum to 10?
    • How many ordered triples of non-negative integers sum to 10?
    • How many ordered triples of positive integers sum to 10 with every part at most 5?

    Asked at Old Mission Capital, Finance, New York, 2018 (Wall Street Oasis): In how many ways can you have three numbers that sum to 10? What about 11?

  2. 063X and Y are independent random variables, each uniform on 0 to 1. What is the density of X + Y, and what is the probability that X + Y is less than 1.5?Continuous and geometric probabilityWarm upCitadelChicago · 2025

    Try it first

    What is P(X + Y < 1.5)?

    Show the worked solution

    The density is a triangle, f(s) = s for s up to 1 and 2 - s from 1 to 2, and P(X + Y < 1.5) = 7/8. Convolving two flat densities gives a tent peaking at 1. The part above 1.5 is a triangle with base 0.5 and height 0.5, area 1/8. In the unit square it is the same corner: the line x + y = 1.5 cuts off a triangle with legs of 0.5.

    Why is the sum not uniform on 0 to 2?

    Roll two dice: a total of 7 can be made six ways, a total of 12 only one way. Continuous uniforms behave the same. A sum near the middle can be made from many pairs, a sum near either end from very few, so the density of the sum rises to a peak and falls again. The mechanism that builds it is {term('convolution', 'The density of a sum of independent variables: for each possible total, add up the density of every pair of values that makes it.')}: the density at s is the length of the set of x values for which both x and s - x lie between 0 and 1.

    Two flat densities convolve into a triangle; the tail above 1.5 is 1/81/8X + Y < 1.5area 7/8x + y = 1.5000.50.511XY00.511.521value of X + Ypeak 1 at sum 1rises: f(s) = stail above 1.5:(1/2)(0.5)(0.5) = 1/87/8
    In the unit square the line x + y = 1.5 cuts off a corner triangle of area 1/8, and in the triangular density of the sum the tail above 1.5 is the same 1/8, so X + Y is below 1.5 with probability 7/8.

    How does the convolution give the triangle?

    Fix a total s. You need x between 0 and 1 and also s - x between 0 and 1, so x must lie between max(0, s - 1) and min(1, s). For s below 1 that interval has length s, and for s above 1 it has length 2 - s, so the density is a tent with its peak of 1 at s = 1. Check that the area is 1: a triangle with base 2 and height 1. The mean is 1 and the variance is 1/12 + 1/12 = 1/6, both of which you can read from symmetry and independence.

    The relationship
    fX+Y(s)=∫011{0≤s−x≤1} dx={s0≤s≤12−s1≤s≤2P(X+Y<1.5)=1−12(0.5)2=78f_{X+Y}(s) = \int_0^1 \mathbf{1}\{0 \le s-x \le 1\}\,dx = \begin{cases} s & 0\le s\le 1\\ 2-s & 1\le s\le 2\end{cases} \qquad P(X+Y<1.5) = 1 - \tfrac12(0.5)^2 = \tfrac78
    f(s)the density of the sum at the value s
    the indicator1 when s - x is a valid value of Y, otherwise 0
    What it says in wordsThe density of the sum at s is how many ways of splitting s are allowed, which rises linearly to 1 and falls back.

    Why give the square picture as well?

    It is a check that costs ten seconds. Because the pair is uniform on the unit square, any probability about X + Y is an area, and the event X + Y at least 1.5 is the corner triangle above the line x + y = 1.5. Its legs run from 0.5 to 1 on each axis, so its area is 1/8 and the answer is 7/8 again. The density gives you the whole distribution; the square gives you any single probability fast. Keep both, because the next question is usually three uniforms, where the density becomes piecewise quadratic and the square becomes a cube: P(X + Y + Z < 1) = 1/6.

    Where candidates lose it

    The common slip is 3/4, from assuming a sum of uniforms is uniform. Sums are never uniform unless one of the pieces is degenerate; they pile up in the middle.

    The second loss is getting the convolution limits wrong and producing a density that does not integrate to 1. Write the two constraints on x out loud, and check the triangle's area before you use it.

    What the interviewer asks next

    • What is the density of X - Y?
    • What is P(X + Y + Z < 1) for three independent uniforms?
    • What is the expected value of max(X, Y), and of X + Y given that X + Y > 1?

    Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis): He was asking some questions about the probability, especially on the convolution.

  3. 068A box holds ten coins: one has heads on both sides and nine are fair. You pick a coin at random, flip it five times and see five heads. What is the probability you picked the double-headed coin?Conditional probability and BayesWarm upJump TradingChicago · 2018

    Try it first

    After five heads, roughly how likely is the double-headed coin?

    Show the worked solution

    32/41, about 78%. Before flipping, the odds are 1 to 9 against the double-headed coin. Five heads happen for certain with it and with chance 1/32 with a fair coin, a likelihood ratio of 32. Multiply: posterior odds are 32 to 9, which is 32/41. Each extra head doubles the odds, so a sixth head would take it to 64/73, about 88%.

    Why is the answer not close to certain?

    Imagine a rare illness and a decent test. A positive result raises the chance you have it, but if the illness is rare enough, most positives still come from healthy people. Evidence is weighed against how common each explanation was to begin with, so five heads, which a fair coin produces only once in 32 tries, still has to overcome nine fair coins for every double-headed one. The 97% instinct takes 1 minus 1/32 and forgets the nine-to-one start.

    Each head doubles the odds on the double-headed coinPick1/10Double-headed5 heads: 1joint 1/10 = 32/3209/10Fair5 heads: 1/32joint 9/320Odds 32 : 9 for the double-headed coinP = 32/41 = 78.0%Chance it is the double-headed coin10%018%131%247%364%478%5heads seen in a rowodds 1:9, 2:9, 4:9 ... 32:9
    Starting from odds of 1 to 9, five heads multiply the odds by 32 to give 32 to 9, so the chance of the double-headed coin rises from 10% to 78%, roughly doubling the odds with each head.

    How do you run Bayes in odds form?

    Odds form is the fastest way to say it in the room. Posterior odds equal prior odds times the likelihood ratio: (1 to 9) times 32 gives 32 to 9. Converting back, 32 out of 32 + 9 is 32/41, about 78%. The long form gives the same thing: the joint chance of picking the special coin and seeing five heads is 1/10, the joint chance of a fair coin and five heads is 9/10 x 1/32 = 9/320, and the posterior is (32/320) / (41/320).

    The relationship
    P(D∣5H)=110⋅1110⋅1+910⋅132=3241≈0.78odds=19×32=329P(D \mid 5H) = \frac{\tfrac1{10}\cdot 1}{\tfrac1{10}\cdot 1 + \tfrac9{10}\cdot\tfrac1{32}} = \frac{32}{41} \approx 0.78 \qquad \text{odds} = \frac19 \times 32 = \frac{32}{9}
    Dthe event that the double-headed coin was picked
    5Hthe observation of five heads in five flips
    32the likelihood ratio: 1 divided by 1/32
    What it says in wordsPrior odds of one to nine, multiplied by a likelihood ratio of thirty-two, give odds of thirty-two to nine.

    What does the odds picture tell you about more flips?

    Each head is twice as likely under the double-headed coin, so every head doubles the odds and every tail ends the question, since the special coin never shows tails. After 0 to 5 heads the chance runs 10%, 18%, 31%, 47%, 64% and 78%; it passes 50% only after the fourth head. The useful follow-up is the next flip: it lands heads with chance 32/41 + (9/41)(1/2) = 73/82, about 89%. On a desk, the same arithmetic tells you how many winning days it takes before a new strategy's record says anything about skill.

    Where candidates lose it

    The trap is answering 31/32, about 97%, by looking only at how unlikely five heads are from a fair coin. That ignores the prior; with nine fair coins in the box, the base rate matters as much as the evidence.

    The second slip is the reverse: staying near 10% because the coin was chosen at random. Say the odds form, prior times likelihood ratio, and both errors disappear.

    What the interviewer asks next

    • What is the probability that the next flip is heads?
    • How many heads in a row would you need to be 99% sure?
    • If one of the ten coins were double-tailed instead, how would five heads change the answer?

    Asked at Jump Trading, Quantitative Research, Chicago, 2018 (Wall Street Oasis): Then he asked one question of probability which can be solved by Bayesian formula.

  4. 072Towns A and B are 100 miles apart. A car leaves A for B at 50 mph. At the same moment a bird leaves B, flying towards the car at 100 mph; each time it meets the car it turns back to B, and each time it reaches B it turns towards the car again, until the car arrives at B. How far does the bird fly in total?Logic and algorithmic reasoningWarm upBLBlackRockNew York · 2025

    Try it first

    How far does the bird fly?

    Show the worked solution

    200 miles. The car needs 100 / 50 = 2 hours to reach B, and the bird flies the whole time at 100 mph, so it covers 2 x 100 = 200 miles. Summing the zigzags gives the same answer: the first round trip is 133.3 miles, each later one is a third of the one before, and 133.3 / (1 - 1/3) = 200.

    What is the question really asking you to count?

    Think of a dog running back and forth between you and your front door while you walk home. You could trace every dash, or you could notice that the dog runs at a steady speed for exactly as long as your walk takes. Distance is speed times time, and the bird's flying time is fixed by the car, not by the zigzags, so the zigzag detail is a distraction. The car covers 100 miles at 50 mph in 2 hours; the bird flies at 100 mph for those same 2 hours. That is 200 miles, and it takes one sentence.

    The bird flies exactly as long as the car drives: 2 hoursAB500 h0.5 h1 h1.5 h2 htime since the car left Acar, 50 mphmeet at 40 min, 33 milesbird, 100 mph, starts at BHard way:sum the zigzags133.3 + 44.4 + ...each 1/3 of the lastEasy way:2 h x 100 mph= 200 miles
    Plotted against time, the bird's zigzags shrink by a factor of three each round and all fit inside the car's 2-hour trip, so the bird flies for 2 hours at 100 mph, a total of 200 miles.

    How do you check it by summing the zigzags?

    The bird and car close the first 100 miles at a combined 150 mph, so they meet after 40 minutes, 33.3 miles from A. The bird flies back to B, 66.7 miles, arriving at 80 minutes, by which time the car is at 66.7 miles. Each round trip starts with the gap to the car one third of the previous gap, so the round trips form a geometric series with ratio 1/3. The first is 133.3 miles; the sum is 133.3 / (1 - 1/3) = 200. It agrees, and it shows why infinitely many turns still add to a finite distance.

    The relationship
    distance=vbird×tcar=100×10050=200∑k≥0133.3(13)k=133.31−13=200\text{distance} = v_{bird}\times t_{car} = 100 \times \frac{100}{50} = 200 \qquad \sum_{k\ge 0} 133.3\left(\tfrac13\right)^k = \frac{133.3}{1-\tfrac13} = 200
    v_birdthe bird's speed, 100 mph
    t_carthe car's travel time, 100 miles at 50 mph
    133.3the first round trip in miles, B to the first meeting and back
    What it says in wordsThe bird flies for exactly as long as the car drives; the zigzag series, summed, gives the same 200 miles.

    Why do interviewers still ask a puzzle this well known?

    Because the way you answer tells them more than the answer. A candidate who starts summing legs has reached for the first method that fits; a candidate who asks what quantity is fixed has found the invariant, and that is the habit the interviewer is hiring. The story about von Neumann summing the series in his head is part of the folklore; you get more credit for the one-line method and the series as a check. The same move, looking for a quantity that does not depend on the messy path, solves many expected-value and stopping questions on this page.

    Where candidates lose it

    The trap is starting the series: solving for the first meeting, then the return, then the second meeting, and running out of time or making an arithmetic slip on the third leg. The infinite number of legs also tempts some candidates to answer infinity.

    Lead with the time argument and give 200 within ten seconds; then offer the series with its ratio of one third as a check, which shows you could do it the long way.

    What the interviewer asks next

    • Where is the car when the bird reaches B for the second time?
    • How many times does the bird turn around?
    • If the bird started at A with the car, flying ahead to B and back, how far would it fly?

    Asked at BlackRock, Quantitative Research, New York, 2025 (Wall Street Oasis): A car starts at point A going 50 miles an hour towards point B

  5. 075You roll two fair dice and are paid the product of the two faces in rupees. What is the expected payout?Expected value and optimal stoppingWarm upWolverine TradingChicago · 2024

    Try it first

    What is the expected product?

    Show the worked solution

    Rs 12.25. The two dice are independent, so the expected product equals the product of the expected faces: 3.5 x 3.5 = 12.25. A check on the 6 by 6 grid: row i averages 3.5i, and the six row averages, 3.5 to 21, average 12.25. If both numbers came from one die, the answer would be E[X squared] = 91/6, about 15.17, which is not the question.

    Why can you multiply the averages?

    Picture a shop whose daily takings are footfall times average spend, where the two have nothing to do with each other. Over a year, average takings are average footfall times average spend. When two quantities are independent, the average of their product is the product of their averages, because knowing one tells you nothing about how big the other will be. That fails the moment they move together: a crowded day with bigger spending lifts the product above the product of averages. Two dice thrown separately are the textbook independent pair.

    Average of the grid = row average x column average = 3.5 x 3.5123456second die1123456avg 3.5224681012avg 73369121518avg 10.544812162024avg 14551015202530avg 17.5661218243036avg 21first diedark cells: above 12.25 (13 of 36)Two independent diceE[XY] = E[X] E[Y]3.5 x 3.5 = 12.25Not the same die squaredE[X^2] = 91/6= 15.17, higher by the varianceThe gap, 35/12 = 2.92, is one die's variance
    Across the 36 equally likely cells of the multiplication grid each row averages its row number times 3.5, so the grand average is 3.5 x 3.5 = 12.25, even though 23 of the 36 products are below it; squaring a single die would give 15.17 instead.

    How do you check 12.25 by brute force quickly?

    Sum the grid by rows. Row i of the multiplication table sums to i x 21, so the whole grid sums to 21 x 21 = 441, and 441 / 36 = 12.25. That is the product rule made visible: the grid's total factors into the first die's total times the second die's total. Note too that the product is skewed: only 18 different values appear, most cells are small, and the big ones, 25, 30 and 36, pull the mean up, so {BELOW75} of the 36 outcomes pay less than the average.

    The relationship
    E[XY]=136∑i=16∑j=16ij=21×2136=44136=12.25=E[X] E[Y]E[XY] = \frac{1}{36}\sum_{i=1}^{6}\sum_{j=1}^{6} ij = \frac{21 \times 21}{36} = \frac{441}{36} = 12.25 = E[X]\,E[Y]
    X, Ythe two independent die faces
    21the sum of the faces 1 to 6
    441the sum of all 36 products
    What it says in wordsThe grid's total is the product of the two dice's totals, so the average product is the product of the averages.

    What is the interviewer likely to ask next?

    The usual next step changes the dependence. If the same die is used for both numbers, you are paid the square of one roll, and its expectation is 91/6, about 15.17, higher by exactly the variance of one die, 35/12. That gap is the covariance at work: E[XY] = E[X]E[Y] + Cov(X, Y). Then comes the price: if you would pay to play, you should quote around 12.25 and note that the payout's standard deviation is about 8.94, so a single play is very noisy relative to its mean.

    Where candidates lose it

    The trap is computing E[X squared] instead of E[XY], getting 15.17, usually because the candidate thinks of rolling one die and squaring. The other slip is answering 18 by taking the midpoint of 1 to 36.

    Say independent and give 3.5 x 3.5 inside five seconds; then offer the row-sum check, 21 x 21 / 36, so the interviewer sees you can verify it.

    What the interviewer asks next

    • What is the expected payout if you are paid the square of a single roll?
    • What is the variance of the product of two dice?
    • You may reroll one of the two dice once after seeing both. What is the game worth?

    Asked at Wolverine Trading, Sales and Trading, Chicago, 2024 (Wall Street Oasis): don't think I was what they were looking for. Questions on EV & Dice.

  6. 076You must predict a quantity with a single constant, and your data are 1, 2, 3, 4 and 40. Which constant minimises the mean squared error, which minimises the mean absolute error, and what does the difference tell you?Statistics and estimationWarm upTower Research CapitalPrinceton · 2018

    Try it first

    Before you calculate: which pair of constants is right?

    Show the worked solution

    The mean, 10, minimises squared error; the median, 3, minimises absolute error. Squared error grows with the square of a miss, so the single 40 pulls the best constant towards it. Absolute error charges every unit of miss equally, so the best constant sits in the middle of the pack. Choosing the loss is choosing which of the two statistics you estimate.

    Why does each loss land on a different constant?

    Picture five friends choosing a meeting point on a straight road: four live at the 1, 2, 3 and 4 km marks and one at the 40 km mark. If the goal is the smallest total travel, you meet at 3 km: moving towards the far friend saves them one km per km moved but costs the other four one km each. If the goal is the smallest total of squared travel, the far friend's 37 km squared, 1,369, dominates everything and the meeting point slides to 10 km. Absolute error balances the count of points on each side, which is the median; squared error balances the total distance on each side, which is the mean.

    One outlier, two losses, two different best constants010203040the outlier, 40median 3mean 10Mean squared errorlime dot, the lowest: 226 at c = 1005101520constant c you predictc = 3 scores 275Mean absolute errorlime dot, the lowest: 8.2 at c = 305101520constant c you predictc = 10 scores 12
    Squared error is lowest at the mean, 226 at c = 10, and the median scores 275 on it; absolute error is lowest at the median, 8.2 at c = 3, and the mean scores 12 on it, so each constant is poor under the other loss.
    The relationship
    ddc∑i(xi−c)2=−2∑i(xi−c)=0⇒c=xˉddc∑i∣xi−c∣=#{xi<c}−#{xi>c}=0⇒c=median\frac{d}{dc}\sum_i (x_i-c)^2 = -2\sum_i (x_i - c) = 0 \Rightarrow c=\bar x \qquad \frac{d}{dc}\sum_i |x_i-c| = \#\{x_i<c\} - \#\{x_i>c\} = 0 \Rightarrow c = \text{median}
    x_ithe five data points
    cthe constant you predict
    #{x_i < c}how many points sit below c
    What it says in wordsSquared error is flat where the total distance above equals the total below, which is the mean; absolute error is flat where the count above equals the count below, which is the median.

    So which constant should you actually use?

    It depends on what the 40 is, and that is the answer the interviewer wants to hear. If the 40 is a typing error or a one-off glitch in a price feed, the median is the honest summary: remove the 40 and the mean falls to 2.5, while the median barely moves. If the 40 is real, say the one big winning day in a strategy's P&L, the mean is the number that matters, because your total profit is the sum of the days and the sum is five times the mean. A robust estimator that ignores the big day would tell you the strategy earns 3 a day when it actually earns 10.

    Where does this show up in a quant job?

    Every regression makes this choice silently. Ordinary least squares minimises squared error and so fits conditional means; one extreme observation can swing the line. Least absolute deviation, or quantile regression at the 50% level, fits conditional medians and shrugs off the same extreme point. In practice, desks winsorise or clip returns before fitting squared-error models, or switch to a Huber loss that is squared near zero and linear in the tails. The limitation is that none of these fixes is free: every one of them throws away some of the information in genuine large moves.

    Where candidates lose it

    The common slip is to answer 10 for both, because the mean feels like the default best guess. The two losses answer different questions, and the interviewer is checking that you know squared error chases the outlier and absolute error does not.

    The second loss is stopping at the arithmetic. Say what the 40 might be, a data error or a real big day, and which loss fits each case. That judgement is the point of the question.

    What the interviewer asks next

    • Which constant minimises the maximum absolute error, and what is that loss called?
    • Add a sixth point at 5. What happens to the median, and to the set of absolute-error minimisers?
    • Why does Lasso use an absolute-value penalty, and how is that related to this puzzle?

    Asked at Tower Research Capital, Trading, Princeton, 2018 (Wall Street Oasis): What if instead of minimizing mean squared error we look at mean absolute error?

  7. 083A market is either calm or stressed each day. A calm day is followed by another calm day with probability 0.8, and a stressed day is followed by another stressed day with probability 0.6. In the long run, what fraction of days are calm?Markov chains and random walksWarm upDRWLondon · 2026

    Try it first

    What share of days are calm in the long run?

    Show the worked solution

    Two thirds of days are calm. In the long run the fraction of days moving from calm to stressed must equal the fraction moving back, so calm share x 0.2 = stressed share x 0.4. Calm days are therefore twice as common as stressed days, 2/3 against 1/3. The same answer comes from spell lengths: calm spells average 5 days and stressed spells 2.5, and 5 / 7.5 = 2/3.

    Why can you balance flows instead of solving equations?

    Think of two rooms at a party. Every few minutes, one in five people in the kitchen wanders to the lounge, and two in five people in the lounge wander back. Once the crowd settles, the numbers crossing each way must match, or one room would keep filling up. In a two-state chain the long-run shares are fixed by one equation: the share of days leaving calm must equal the share of days leaving stressed. Calm leaves at rate 0.2 and stressed at rate 0.4, so calm must hold twice as many days.

    In the long run the flow out of calm equals the flow back inCalmStressed0.80.60.2 calm to stressed0.4 stressed to calmBalance: (2/3) x 0.2 = (1/3) x 0.4 = 2/15 of all dayscalm 2/3 of daysstressed 1/3average calm spell 1/0.2 = 5 days; average stressed spell 1/0.4 = 2.5 days
    Calm days leave at rate 0.2 and stressed days at rate 0.4, so in the long run calm must hold twice as many days as stressed for the flows to balance: two thirds calm and one third stressed, with each flow equal to 2/15 of all days.
    The relationship
    πC (1−0.8)=πS (1−0.6),πC+πS=1  ⇒  πC=0.40.2+0.4=23\pi_C\,(1-0.8) = \pi_S\,(1-0.6), \quad \pi_C + \pi_S = 1 \;\Rightarrow\; \pi_C = \frac{0.4}{0.2+0.4} = \frac23
    pi_Cthe long-run share of calm days
    pi_Sthe long-run share of stressed days
    1 - 0.8the chance a calm day is followed by a stressed one
    What it says in wordsEach state's long-run share is the rate of leaving the other state, divided by the two leaving rates added together.

    How do you check it a second way?

    Use spell lengths. A calm spell ends each day with probability 0.2, so it lasts 1/0.2 = 5 days on average; a stressed spell ends with probability 0.4, so it lasts 2.5 days. The chain alternates calm spell, stressed spell, calm spell, so the calm share is 5 out of every 7.5 days, which is 2/3. Two methods, one answer, in under a minute.

    How fast does the chain forget where it started?

    The transition matrix has a second eigenvalue of 0.8 + 0.6 - 1 = 0.4, and any gap between today's probabilities and the long-run split shrinks by that factor each day. After five days the starting state explains only 0.4^5, about 1%, of the gap, so the answer does not depend on how the week began. A trader would add the limitation: real regimes are not memoryless, and a stress spell that has already lasted a month is not as likely to end tomorrow as one that started yesterday.

    Where candidates lose it

    The trap is answering 80%, reading the chance of staying calm as the share of calm days. The 0.8 describes one step, and the long-run share depends on how quickly both states are left, not just one of them.

    The second slip is setting up a full eigenvector calculation and running out of time. Say the flow balance in one line, then use the spell lengths as the check.

    What the interviewer asks next

    • Today is stressed. What is the probability that the day after tomorrow is calm?
    • What is the expected number of days until the first stressed day, starting calm?
    • If a desk loses Rs 2 lakh on stressed days and makes Rs 1 lakh on calm days, what is its long-run average daily P&amp;L?

    Asked at DRW, Trader Intern Interview, London, 2026 (Wall Street Oasis): consisted of math, statistics, and probability theory (eg. one question was about markov chains

  8. 086A price-weighted index holds three stocks priced 50, 100 and 150, with a divisor of 3. The 150 stock splits 3 for 1. What is the new divisor, and how does a market-cap-weighted index handle the same split?Pricing, options and index mathsWarm upMizuhoHong Kong · 2024

    Try it first

    What must the new divisor be?

    Show the worked solution

    The new divisor is 2. Before the split the prices sum to 300, and 300 / 3 is 100. After a 3 for 1 split the 150 stock trades at 50, the sum is 200, and only a divisor of 2 keeps the index at 100. A cap-weighted index needs no adjustment at all, because a split triples the share count as it cuts the price to a third, leaving market value unchanged.

    Why must the divisor change when nothing about the company changed?

    Cut a pizza into twelve slices instead of four and you have not made more pizza. A stock split does the same to a company: three times the shares, each worth a third. A price-weighted index adds up share prices, so a split drops the sum even though no value was lost, and the divisor must be cut to stop a fake fall in the index. Solve for it by keeping the index level fixed: 200 divided by the new divisor must equal 100, so the divisor is 2.

    A split changes the sum of prices, so the divisor must change to hold the indexBefore: C at 15050stock Aweight 17%100stock Bweight 33%150stock Cweight 50%sum 300 / divisor 3 = index 100After a 3 for 1 split: C at 5050stock Aweight 25%100stock Bweight 50%50stock Cweight 25%sum 200 / divisor 2 = index 100split
    Before the split the prices sum to 300 and the index is 300 / 3 = 100; after C splits 3 for 1 the sum is 200, so the divisor falls to 2 to hold the index at 100, and C's weight falls from 50% to 25%.
    The relationship
    I=∑iPid3003=200d′⇒d′=2I = \frac{\sum_i P_i}{d} \qquad \frac{300}{3} = \frac{200}{d'} \Rightarrow d' = 2
    P_ithe price of stock i
    dthe divisor before the split, 3
    d'the divisor after the split
    What it says in wordsChoose the new divisor so the index is the same the moment after the split as the moment before.

    What else changes in a price-weighted index after the split?

    The weights. In a price-weighted index a stock's weight is its price over the sum of prices, so the expensive stock dominates whatever the size of the company. Before the split C carried 50% of the index; after it, C carries only 25% and B, untouched, jumps to 50%. A 10% rise in C used to add 5 index points; now it adds 2.5. Nothing about C's business changed; the index simply started caring less about it, which is the main criticism of price weighting.

    How do the other common methods treat the split?

    A market-cap-weighted index sums price times shares, and a 3 for 1 split multiplies shares by 3 while dividing price by 3, so the stock's market value, its weight and the index are all unchanged; no divisor adjustment is needed for a split. Cap-weighted divisors still change for events that alter total market value without a price move, such as share issuance, buybacks or a constituent being replaced. An equal-weighted index is also untouched by a split, since weights are reset to equal at each rebalance, but it has to trade at every rebalance to get back to equal, which costs money.

    Where candidates lose it

    The common slip is dividing the old divisor by the split ratio and answering 1. The divisor is fixed by keeping the index level unchanged, and only one stock split, so the adjustment is smaller than the ratio.

    The second loss is saying a cap-weighted index needs the same adjustment. Market value does not change in a split, so a cap-weighted index does nothing; say that, then name the events that do change its divisor.

    What the interviewer asks next

    • Stock B now pays a special dividend of 20. How does each index type handle it?
    • Replace stock A with a new stock priced 200. What is the new divisor?
    • Which stock has the most influence on a price-weighted index, and why is that a flaw?

    Asked at Mizuho, Sales and Trading, Hong Kong, 2024 (Wall Street Oasis): Different index methodology - need to know all of them with examples.

  9. 090X and Y are independent random variables with the same variance. What is the correlation between X and X + Y?Correlation, regression and linear algebraWarm upSCSquarepoint CapitalMontreal · 2026

    Try it first

    Pick the correlation.

    Show the worked solution

    1/sqrt(2), about 0.71. The covariance of X with X + Y is Var X plus Cov(X, Y), and the second term is zero, so it is sigma squared. The variance of X + Y is 2 sigma squared, because independent variances add. Dividing sigma squared by sigma times sqrt(2) sigma leaves 1/sqrt(2). Squared, that is 0.5: X explains half of the sum's variance.

    Why isn't the answer one half?

    Picture two people each tossing a coin for a rupee, and a pot holding their combined winnings. One player's result explains exactly half of the pot's variability, and the other half comes from the other player. Half is the share of variance explained, R squared, and correlation is its square root, so the correlation is 1/sqrt(2), not 1/2. This is the most common slip on the question, and it comes from mixing up the two measures.

    X supplies half of the variance of X + Y, so the correlation is 1/sqrt(2)XYXYVar Xsigma^2Cov(X, Y)0Cov(Y, X)0Var Ysigma^2X row:Cov(X, X+Y)= sigma^2all four cells: Var(X + Y) = 2 sigma^2corr = Cov / (sd X x sd(X+Y))= sigma^2 / (sigma x sqrt(2) sigma)= 1/sqrt(2) = 0.707Share of Var(X + Y) explained by XR^2 = 0.5Unequal variances: corr = 1/sqrt(1 + k),k = Var Y / Var X; k = 4 gives 0.447
    In the covariance box, X's own variance fills one of the two non-zero cells, so X accounts for half of Var(X + Y), and the correlation between X and X + Y is sigma squared divided by sigma times sqrt(2) sigma, which is 1/sqrt(2), about 0.71.
    The relationship
    ρ=Cov⁡(X,X+Y)σX σX+Y=σ2+0σ⋅2 σ=12≈0.707\rho = \frac{\operatorname{Cov}(X, X+Y)}{\sigma_X\,\sigma_{X+Y}} = \frac{\sigma^2 + 0}{\sigma\cdot\sqrt{2}\,\sigma} = \frac{1}{\sqrt 2} \approx 0.707
    Cov(X, X + Y)Var X plus Cov(X, Y), which is sigma squared plus zero
    sigma_{X+Y}the standard deviation of the sum, sqrt(2) sigma
    What it says in wordsCovariance is linear, so split it into pieces; the only surviving piece is X's own variance.

    How does it change if the variances differ?

    Let Var Y be k times Var X. The covariance is still Var X, and Var(X + Y) becomes (1 + k) Var X. The correlation is 1/sqrt(1 + k): the noisier Y is, the less the sum tracks X. At k = 1 you get 0.707; at k = 4, 0.447; at k = 0.25, 0.894. This is exactly the signal-plus-noise model: if a price move is a true signal plus independent noise of equal size, the best-case correlation between your signal and the move is about 0.71.

    Where does this show up on a desk?

    Any time one piece is part of a total. A stock's return is market return plus its own specific return; if the two had equal variance, the stock would correlate 0.71 with the market. The same arithmetic tells you the ceiling on a predictor: if half of tomorrow's move is unpredictable noise, no model can correlate more than 0.71 with it. The limitation is independence; if X and Y are correlated, add 2 Cov(X, Y) to the variance of the sum and Cov(X, Y) to the covariance.

    Where candidates lose it

    The frequent slip is answering one half, confusing the share of variance with the correlation. Correlation is the square root of that share.

    The second loss is writing the standard deviation of X + Y as 2 sigma, adding standard deviations instead of variances, which gives one half again by a different road. Independent variances add; standard deviations do not.

    What the interviewer asks next

    • What is the correlation between X + Y and X - Y?
    • If X and Y have correlation 0.5 and equal variance, what is corr(X, X + Y)?
    • What is the correlation between the first die and the total of two dice?

    Asked at Squarepoint Capital, Desk Quant Analyst Interview, Montreal, 2026 (Wall Street Oasis): There were also 3-4 basic math/stats questions about mean, covariance, correlation, etc.

  10. 092You climb a staircase of 10 steps, taking either one step or two steps at a time. In how many different ways can you reach the top?Counting and combinatoricsWarm upTower Research CapitalNew York · 2012

    Try it first

    Pick the number of ways.

    Show the worked solution

    89 ways. Split by the last move: you arrive at step 10 with a single from step 9 or a double from step 8, so ways(10) = ways(9) + ways(8). With 1 way to reach step 1 and 2 ways to reach step 2, the counts run 1, 2, 3, 5, 8, 13, 21, 34, 55, 89: the Fibonacci numbers, with 89 at the top.

    How do you count without listing every route?

    Imagine a friend at the top of the stairs asks how you got there. There is one thing you can say for certain about your last move: it was either a single from step 9 or a double from step 8, and never both. Every route to step 10 is a route to step 9 followed by a single, or a route to step 8 followed by a double, so the count at step 10 is the sum of the counts at steps 9 and 8. The same holds at every step, which turns the puzzle into a running sum.

    Ways to reach each step: the sum of the two steps below1step 12step 23step 35step 48step 513step 621step 734step 855step 989step 10dashed: 34 ways end with a double from step 8solid: 55 ways end with a single from step 9step 10: 34 + 55 = 89start: 1 way to step 1, 2 ways to step 2 (1+1 or 2)
    Writing the number of ways on each step, each step is the sum of the two below it, so the counts follow the Fibonacci sequence and step 10 collects 55 routes ending with a single and 34 ending with a double, 89 in all.
    The relationship
    w(n)=w(n−1)+w(n−2),w(1)=1,  w(2)=2  ⇒  w(10)=89w(n) = w(n-1) + w(n-2), \qquad w(1) = 1,\; w(2) = 2 \;\Rightarrow\; w(10) = 89
    w(n)the number of ways to reach step n
    w(n-1)routes whose last move is a single step
    w(n-2)routes whose last move is a double step
    What it says in wordsSort every route by its last move; the two groups do not overlap and together cover everything.

    Can you check 89 a second way?

    Count by how many double steps you take. With k doubles, you make 10 - 2k singles, so 10 - k moves in all, and you only have to choose which k of those moves are the doubles. Summing the binomial counts over k = 0 to 5 gives 1 + 9 + 28 + 35 + 15 + 1 = 89, the same answer by a completely different road. Saying a second check aloud is worth more than the answer itself in a first round, because it shows you do not trust a pattern you have not tested.

    Double steps kMoves in totalWays to place the doubles
    010C(10, 0) = 1
    19C(9, 1) = 9
    28C(8, 2) = 28
    37C(7, 3) = 35
    46C(6, 4) = 15
    55C(5, 5) = 1
    total 89
    Counting routes by the number of double steps gives 89 again, which confirms the Fibonacci running sum.

    What does the interviewer usually ask next?

    Two things. First, allow steps of one, two or three: the same last-move argument gives w(n) = w(n - 1) + w(n - 2) + w(n - 3), and the count for ten steps becomes 274. Second, the coding version. A recursive function that calls itself for n - 1 and n - 2 recomputes the same steps again and again: for 30 steps it makes 1,664,079 calls to return 1,346,269. Storing each step's count once, or just keeping the last two numbers in a loop, does the job in 30 additions. The counts grow by about 1.618, the golden ratio, per step, which is the limitation of any approach that lists routes rather than counting them.

    Where candidates lose it

    The fast wrong answer is 2^10 = 1,024, treating each of ten stairs as a binary choice. A double step consumes two stairs, so routes have different numbers of moves and the choices are not ten independent coin flips.

    The second loss is an off-by-one in the starting values, which lands on 55 or 144. Write the first three steps out by hand: 1 way to step 1, 2 ways to step 2, 3 ways to step 3. Anchor the sequence there and the tenth term is 89.

    What the interviewer asks next

    • What if you can also take three steps at a time?
    • How many ways are there if step 5 is broken and cannot be stood on?
    • Write code that counts the ways for 1,000 steps without the recursion blowing up.

    Asked at Tower Research Capital, Intern Interview -, New York, 2012 (Wall Street Oasis): How many ways can you jump up stairs if you can only jump either 1 or 2 steps?

← PreviousPage 2 of 2
  1. 1
  2. 2
Next →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.