Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant puzzles, solved step by step

Puzzles
100
Traced to a firm
71
Topics
12
Hard
30
Topic
All topicsLogic and algorithmic reasoning10Conditional probability and Bayes7Counting and combinatorics8Continuous and geometric probability9Correlation, regression and linear algebra9Market making, betting and sizing9Expected value and optimal stopping9Statistics and estimation9Pricing, options and index maths7Games and strategic reasoning8Markov chains and random walks7Mental maths and number sense8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 71–80 of 100
  1. 071We play a matching game: each of us shows heads or tails at the same moment. If both show heads I pay you 3; if both show tails I pay you 1; if we mismatch you pay me 2. What mix should I use to make you indifferent, what mix should you use, and what is the game worth to you per round?Games and strategic reasoningCoreQuant tradingQuant research

    Try it first

    What is the game worth to you per round?

    Show the worked solution

    I show heads 3/8 of the time; so should you; and the game is worth -1/8 to you per round. At my mix q, your heads pays 3q - 2(1 - q) and your tails pays -2q + (1 - q). Setting them equal gives q = 3/8, where both pay -1/8. By the same algebra your 3/8 mix makes my choices equal, so neither of us can improve: the table looks fair and is not.

    Why do I randomise to make you indifferent, not to help myself?

    Think of a penalty taker and a goalkeeper. If the taker shoots left more often than he should, the keeper dives left and gains; any pattern is exploited. In a zero-sum game, the mix that protects you is the one that leaves your opponent with nothing to exploit, which means making every one of their choices pay the same. So I do not pick my heads probability by looking at my own payoffs in isolation; I pick it so that your heads and your tails earn you equal amounts. Then it no longer matters what you do.

    My mix of 3/8 heads makes your two choices pay the same: -1/8Payoff to youI showHTH+3-2T-2+1You showWins 3 + 1 = 4, losses 2 + 2 = 4:looks fair, is not-2-1+1+2+30you show Hyou show Tcross at q = 3/8,both pay -1/800.250.50.751q = my probability of showing heads
    Your payoff from heads rises and your payoff from tails falls as my heads probability grows; they cross at 3/8, where each pays -1/8, so that is my equalising mix and the most you can secure per round.

    How do the two equations give 3/8 and -1/8?

    Let q be my chance of heads. Your heads earns 3q - 2(1 - q) = 5q - 2 and your tails earns -2q + 1(1 - q) = 1 - 3q, and they are equal when 8q = 3, at q = 3/8. Plug back in: 5(3/8) - 2 = -1/8. For your side, let p be your chance of heads; my payoffs are the negatives of yours, and because the table is symmetric the same algebra gives p = 3/8. At those mixes, each cell's frequency is HH 9/64, TT 25/64 and the two mismatches 15/64 each: 3(9) + 1(25) - 2(30) = -8, over 64, is -1/8.

    The relationship
    5q−2=1−3q  ⇒  q∗=38,V=5⋅38−2=−185q - 2 = 1 - 3q \;\Rightarrow\; q^* = \tfrac38,\qquad V = 5\cdot\tfrac38 - 2 = -\tfrac18
    qmy probability of showing heads
    5q - 2your expected payoff from showing heads
    1 - 3qyour expected payoff from showing tails
    Vthe value of the game to you per round
    What it says in wordsMy heads probability is set where your two choices earn the same, and that common amount is what the game is worth to you.

    Why does a table with equal wins and losses favour me?

    Because the cells are not played equally often. Your big win needs both of us on heads, and I can make that cell rare; the mismatches, where I win, happen 30 times in 64 at equilibrium. If I played heads half the time you could earn +1/2 by always showing heads, so a naive 50/50 from me would be a gift. In trading the same idea sets a market maker's quotes: they are placed so that the informed side has no choice that beats the others, not so that the market maker profits on any one trade.

    Where candidates lose it

    The trap is adding up the table, 3 + 1 against 2 + 2, and calling the game fair. Equilibrium frequencies are not a quarter each, so a table's totals tell you nothing about its value.

    The second slip is solving for the mix that maximises your own payoff against a fixed opponent. In a mixed equilibrium each player's mix is pinned down by the other player's payoffs; say that sentence before the algebra.

    What the interviewer asks next

    • If I play heads half the time, what should you do and what do you earn?
    • What payoff for both tails would make the game fair?
    • How does the answer change if the mismatch payment is 3 instead of 2?
  2. 072Towns A and B are 100 miles apart. A car leaves A for B at 50 mph. At the same moment a bird leaves B, flying towards the car at 100 mph; each time it meets the car it turns back to B, and each time it reaches B it turns towards the car again, until the car arrives at B. How far does the bird fly in total?Logic and algorithmic reasoningWarm upBLBlackRockNew York · 2025

    Try it first

    How far does the bird fly?

    Show the worked solution

    200 miles. The car needs 100 / 50 = 2 hours to reach B, and the bird flies the whole time at 100 mph, so it covers 2 x 100 = 200 miles. Summing the zigzags gives the same answer: the first round trip is 133.3 miles, each later one is a third of the one before, and 133.3 / (1 - 1/3) = 200.

    What is the question really asking you to count?

    Think of a dog running back and forth between you and your front door while you walk home. You could trace every dash, or you could notice that the dog runs at a steady speed for exactly as long as your walk takes. Distance is speed times time, and the bird's flying time is fixed by the car, not by the zigzags, so the zigzag detail is a distraction. The car covers 100 miles at 50 mph in 2 hours; the bird flies at 100 mph for those same 2 hours. That is 200 miles, and it takes one sentence.

    The bird flies exactly as long as the car drives: 2 hoursAB500 h0.5 h1 h1.5 h2 htime since the car left Acar, 50 mphmeet at 40 min, 33 milesbird, 100 mph, starts at BHard way:sum the zigzags133.3 + 44.4 + ...each 1/3 of the lastEasy way:2 h x 100 mph= 200 miles
    Plotted against time, the bird's zigzags shrink by a factor of three each round and all fit inside the car's 2-hour trip, so the bird flies for 2 hours at 100 mph, a total of 200 miles.

    How do you check it by summing the zigzags?

    The bird and car close the first 100 miles at a combined 150 mph, so they meet after 40 minutes, 33.3 miles from A. The bird flies back to B, 66.7 miles, arriving at 80 minutes, by which time the car is at 66.7 miles. Each round trip starts with the gap to the car one third of the previous gap, so the round trips form a geometric series with ratio 1/3. The first is 133.3 miles; the sum is 133.3 / (1 - 1/3) = 200. It agrees, and it shows why infinitely many turns still add to a finite distance.

    The relationship
    distance=vbird×tcar=100×10050=200∑k≥0133.3(13)k=133.31−13=200\text{distance} = v_{bird}\times t_{car} = 100 \times \frac{100}{50} = 200 \qquad \sum_{k\ge 0} 133.3\left(\tfrac13\right)^k = \frac{133.3}{1-\tfrac13} = 200
    v_birdthe bird's speed, 100 mph
    t_carthe car's travel time, 100 miles at 50 mph
    133.3the first round trip in miles, B to the first meeting and back
    What it says in wordsThe bird flies for exactly as long as the car drives; the zigzag series, summed, gives the same 200 miles.

    Why do interviewers still ask a puzzle this well known?

    Because the way you answer tells them more than the answer. A candidate who starts summing legs has reached for the first method that fits; a candidate who asks what quantity is fixed has found the invariant, and that is the habit the interviewer is hiring. The story about von Neumann summing the series in his head is part of the folklore; you get more credit for the one-line method and the series as a check. The same move, looking for a quantity that does not depend on the messy path, solves many expected-value and stopping questions on this page.

    Where candidates lose it

    The trap is starting the series: solving for the first meeting, then the return, then the second meeting, and running out of time or making an arithmetic slip on the third leg. The infinite number of legs also tempts some candidates to answer infinity.

    Lead with the time argument and give 200 within ten seconds; then offer the series with its ratio of one third as a check, which shows you could do it the long way.

    What the interviewer asks next

    • Where is the car when the bird reaches B for the second time?
    • How many times does the bird turn around?
    • If the bird started at A with the car, flying ahead to B and back, how far would it fly?

    Asked at BlackRock, Quantitative Research, New York, 2025 (Wall Street Oasis): A car starts at point A going 50 miles an hour towards point B

  3. 073Use Newton's method to find the square root of 2, starting from 1.5. How many correct digits do you have after each step, and why?Mental maths and number senseCoreQuant researchDesk quant

    Try it first

    Starting from 1.5, about how many correct digits after three Newton steps?

    Show the worked solution

    About 1, 3, 6 and 12 correct digits: the count roughly doubles each step. The update is x(next) = (x + 2/x)/2: 1.5 gives 17/12 = 1.41667, then 577/408 = 1.4142157, then 665857/470832 = 1.41421356237469. Each new error is about the old error squared divided by 2x, so if the error is 10^-k, the next is about 10^-2k. That is quadratic convergence.

    Where does the update rule come from?

    Think of guessing a side of a square room whose area is 2. If your guess is too big, 2 divided by your guess is too small, and the truth sits between the two. Newton's method for x squared minus 2 is exactly that: replace x with the average of x and 2/x. Formally, Newton follows the tangent of f(x) = x^2 - 2 down to zero, x - f(x)/f'(x) = x - (x^2 - 2)/(2x), which simplifies to (x + 2/x)/2. The averaging form is the one to use in your head.

    Each Newton step roughly doubles the correct digitsStepFractionDecimal (correct digits in green)Correct digits03/21.51117/121.416666666666632577/4081.414215686274563665857/4708321.4142135623746124(digits run past the table)1.414213562373024Rule: x(next) = (x + 2/x) / 2. The new error is about the old error squared over 2x.
    Starting from 1.5, Newton's iterates for the square root of 2 have 1, 3, 6 and then 12 correct digits, doubling at each step because each new error is roughly the square of the old one.

    How do you get the iterates without a calculator?

    Keep fractions. From 3/2, the next value is (3/2 + 4/3)/2 = 17/12, then (17/12 + 24/17)/2 = 577/408, and the pattern continues: if x = p/q, the next is (p^2 + 2q^2)/(2pq). 17/12 is 1.41667, already right to 1.41. 577/408 is 1.4142157 against 1.4142136, right to 1.41421. The third step's fraction, 665857/470832, is too big for mental division, but you can predict its accuracy without doing it, which is the point of the question.

    The relationship
    xn+1=12(xn+2xn)xn+1−2=(xn−2)22xnx_{n+1} = \tfrac12\Big(x_n + \frac{2}{x_n}\Big) \qquad x_{n+1} - \sqrt2 = \frac{(x_n - \sqrt2)^2}{2x_n}
    x_nthe current estimate of the square root of 2
    x_n - sqrt 2the error of the current estimate
    2 x_nabout 2.8 near the root, so the new error is about a third of the old error squared
    What it says in wordsEach new error is the old error squared, divided by about 2.8, so the correct digits roughly double.

    Why is it the digits that double, and when does that fail?

    Subtract the root from the update and the algebra collapses to (x - root 2)^2 / 2x. Squaring an error of 10^-3 gives 10^-6, so each step doubles the number of correct digits once you are close. The errors here run about 0.09, 0.0025, 2 x 10^-6 and 1.6 x 10^-12. The doubling needs a good start and a simple root: far from the root, or where the slope is zero, Newton can creep or jump away. Bisection, by contrast, gains one binary digit per step whatever happens, which is why desk code often brackets with bisection and finishes with Newton when solving for implied volatility.

    Where candidates lose it

    The trap is guessing linear progress, one or two digits a step, because that is how most iterative methods feel. Newton is special near a simple root, and the interviewer wants the word quadratic and the reason for it.

    The second loss is getting lost in decimals. Work in fractions, 3/2, 17/12, 577/408, and state the error-squared rule instead of computing the third step.

    What the interviewer asks next

    • Write Newton's update for the cube root of 10, and start it from 2.
    • Why does Newton converge only linearly at a double root?
    • How would you use Newton's method to find an implied volatility, and what can go wrong?
  4. 074A point is dropped uniformly at random in a unit square. What is the expected distance from the point to the nearest edge of the square?Continuous and geometric probabilityCoreHRHudson River TradingNew York · 2024

    Try it first

    What is the expected distance to the nearest edge?

    Show the worked solution

    1/6. Let D be the distance to the nearest edge. D exceeds d exactly when the point lies in the inner square of side 1 - 2d, so P(D > d) = (1 - 2d)^2 for d up to 1/2. The expected value of a non-negative variable is the integral of its tail, and the integral of (1 - 2d)^2 from 0 to 1/2 is 1/6.

    Why work with the chance of being far rather than the distance itself?

    Think of a sandpit where a child stands at a random spot and the question is how far they are from the nearest edge. Writing the distance as min(x, 1 - x, y, 1 - y) and integrating a minimum of four things means splitting the square into four triangles. Asking instead when the point is farther than d from every edge has a one-picture answer: the point must lie in a smaller square, shrunk by d on every side. That square has side 1 - 2d, so its area, (1 - 2d)^2, is the tail probability. One formula replaces four cases.

    Farther than d from every edge means inside a square of side 1 - 2d0.640.360.160.04nearest edgenumbers: P(distance > d) for d = 0.1 to 0.400.10.20.30.40.50.51P(D > d) = (1 - 2d)^2area = 1/6d, distance to the nearest edgesimulated: 0.1669exact: 1/6 = 0.1667
    A point is more than d from every edge only inside the inner square of side 1 - 2d, so the chance of being farther than 0.1, 0.2, 0.3 and 0.4 is 0.64, 0.36, 0.16 and 0.04, and the area under that tail curve is the expected distance, 1/6.

    How does the tail give the expectation?

    For any non-negative random variable, the expected value equals the integral of the chance that it exceeds each level, E[D] = integral of P(D > d). Here that is the integral of (1 - 2d)^2 from 0 to 1/2. Substitute u = 1 - 2d and it becomes half the integral of u^2 from 0 to 1, which is 1/6. A simulation with 200,000 random points gives 0.1669, against the exact 0.1667.

    The relationship
    E[D]=∫01/2P(D>d) dd=∫01/2(1−2d)2 dd=[−(1−2d)36]01/2=16E[D] = \int_0^{1/2} P(D>d)\,dd = \int_0^{1/2} (1-2d)^2\,dd = \Big[-\tfrac{(1-2d)^3}{6}\Big]_0^{1/2} = \tfrac16
    Dthe distance from the random point to the nearest edge
    P(D > d)the area of the inner square of side 1 - 2d
    What it says in wordsAdd up the chance of being farther than each distance, and the total is the expected distance.

    How do you sanity-check 1/6 against simpler cases?

    Build up the number of edges. The distance to one fixed edge averages 1/2, to the nearer of two opposite edges averages 1/4, and to the nearest of all four it falls to 1/6, so each added constraint pulls the minimum closer. The density of D is the slope of the tail, 4(1 - 2d), largest at the edge, which says most random points are near the boundary. That is the same reason most of the volume of a high-dimensional cube sits near its surface, a fact that matters when sampling scenarios in many risk factors at once.

    Where candidates lose it

    The trap answers are 1/2 and 1/4, from handling one edge or one axis and forgetting that the nearest of four edges is a minimum. A candidate who integrates min(x, 1 - x, y, 1 - y) directly often splits the square wrongly and lands on a different number.

    Draw the inner square and say tail integral; the whole calculation is then one line.

    What the interviewer asks next

    • What is the expected distance to the nearest edge in a unit cube?
    • What is the expected distance to the nearest corner of the square?
    • What is the density of the distance to the nearest edge, and where is it highest?

    Asked at Hudson River Trading, Campus Algo Dev Interview, New York, 2024 (Wall Street Oasis): I was asked a expected value question involving the expected value among distance to an edge, with a randomly placed object.

  5. 075You roll two fair dice and are paid the product of the two faces in rupees. What is the expected payout?Expected value and optimal stoppingWarm upWolverine TradingChicago · 2024

    Try it first

    What is the expected product?

    Show the worked solution

    Rs 12.25. The two dice are independent, so the expected product equals the product of the expected faces: 3.5 x 3.5 = 12.25. A check on the 6 by 6 grid: row i averages 3.5i, and the six row averages, 3.5 to 21, average 12.25. If both numbers came from one die, the answer would be E[X squared] = 91/6, about 15.17, which is not the question.

    Why can you multiply the averages?

    Picture a shop whose daily takings are footfall times average spend, where the two have nothing to do with each other. Over a year, average takings are average footfall times average spend. When two quantities are independent, the average of their product is the product of their averages, because knowing one tells you nothing about how big the other will be. That fails the moment they move together: a crowded day with bigger spending lifts the product above the product of averages. Two dice thrown separately are the textbook independent pair.

    Average of the grid = row average x column average = 3.5 x 3.5123456second die1123456avg 3.5224681012avg 73369121518avg 10.544812162024avg 14551015202530avg 17.5661218243036avg 21first diedark cells: above 12.25 (13 of 36)Two independent diceE[XY] = E[X] E[Y]3.5 x 3.5 = 12.25Not the same die squaredE[X^2] = 91/6= 15.17, higher by the varianceThe gap, 35/12 = 2.92, is one die's variance
    Across the 36 equally likely cells of the multiplication grid each row averages its row number times 3.5, so the grand average is 3.5 x 3.5 = 12.25, even though 23 of the 36 products are below it; squaring a single die would give 15.17 instead.

    How do you check 12.25 by brute force quickly?

    Sum the grid by rows. Row i of the multiplication table sums to i x 21, so the whole grid sums to 21 x 21 = 441, and 441 / 36 = 12.25. That is the product rule made visible: the grid's total factors into the first die's total times the second die's total. Note too that the product is skewed: only 18 different values appear, most cells are small, and the big ones, 25, 30 and 36, pull the mean up, so {BELOW75} of the 36 outcomes pay less than the average.

    The relationship
    E[XY]=136∑i=16∑j=16ij=21×2136=44136=12.25=E[X] E[Y]E[XY] = \frac{1}{36}\sum_{i=1}^{6}\sum_{j=1}^{6} ij = \frac{21 \times 21}{36} = \frac{441}{36} = 12.25 = E[X]\,E[Y]
    X, Ythe two independent die faces
    21the sum of the faces 1 to 6
    441the sum of all 36 products
    What it says in wordsThe grid's total is the product of the two dice's totals, so the average product is the product of the averages.

    What is the interviewer likely to ask next?

    The usual next step changes the dependence. If the same die is used for both numbers, you are paid the square of one roll, and its expectation is 91/6, about 15.17, higher by exactly the variance of one die, 35/12. That gap is the covariance at work: E[XY] = E[X]E[Y] + Cov(X, Y). Then comes the price: if you would pay to play, you should quote around 12.25 and note that the payout's standard deviation is about 8.94, so a single play is very noisy relative to its mean.

    Where candidates lose it

    The trap is computing E[X squared] instead of E[XY], getting 15.17, usually because the candidate thinks of rolling one die and squaring. The other slip is answering 18 by taking the midpoint of 1 to 36.

    Say independent and give 3.5 x 3.5 inside five seconds; then offer the row-sum check, 21 x 21 / 36, so the interviewer sees you can verify it.

    What the interviewer asks next

    • What is the expected payout if you are paid the square of a single roll?
    • What is the variance of the product of two dice?
    • You may reroll one of the two dice once after seeing both. What is the game worth?

    Asked at Wolverine Trading, Sales and Trading, Chicago, 2024 (Wall Street Oasis): don't think I was what they were looking for. Questions on EV & Dice.

  6. 076You must predict a quantity with a single constant, and your data are 1, 2, 3, 4 and 40. Which constant minimises the mean squared error, which minimises the mean absolute error, and what does the difference tell you?Statistics and estimationWarm upTower Research CapitalPrinceton · 2018

    Try it first

    Before you calculate: which pair of constants is right?

    Show the worked solution

    The mean, 10, minimises squared error; the median, 3, minimises absolute error. Squared error grows with the square of a miss, so the single 40 pulls the best constant towards it. Absolute error charges every unit of miss equally, so the best constant sits in the middle of the pack. Choosing the loss is choosing which of the two statistics you estimate.

    Why does each loss land on a different constant?

    Picture five friends choosing a meeting point on a straight road: four live at the 1, 2, 3 and 4 km marks and one at the 40 km mark. If the goal is the smallest total travel, you meet at 3 km: moving towards the far friend saves them one km per km moved but costs the other four one km each. If the goal is the smallest total of squared travel, the far friend's 37 km squared, 1,369, dominates everything and the meeting point slides to 10 km. Absolute error balances the count of points on each side, which is the median; squared error balances the total distance on each side, which is the mean.

    One outlier, two losses, two different best constants010203040the outlier, 40median 3mean 10Mean squared errorlime dot, the lowest: 226 at c = 1005101520constant c you predictc = 3 scores 275Mean absolute errorlime dot, the lowest: 8.2 at c = 305101520constant c you predictc = 10 scores 12
    Squared error is lowest at the mean, 226 at c = 10, and the median scores 275 on it; absolute error is lowest at the median, 8.2 at c = 3, and the mean scores 12 on it, so each constant is poor under the other loss.
    The relationship
    ddc∑i(xi−c)2=−2∑i(xi−c)=0⇒c=xˉddc∑i∣xi−c∣=#{xi<c}−#{xi>c}=0⇒c=median\frac{d}{dc}\sum_i (x_i-c)^2 = -2\sum_i (x_i - c) = 0 \Rightarrow c=\bar x \qquad \frac{d}{dc}\sum_i |x_i-c| = \#\{x_i<c\} - \#\{x_i>c\} = 0 \Rightarrow c = \text{median}
    x_ithe five data points
    cthe constant you predict
    #{x_i < c}how many points sit below c
    What it says in wordsSquared error is flat where the total distance above equals the total below, which is the mean; absolute error is flat where the count above equals the count below, which is the median.

    So which constant should you actually use?

    It depends on what the 40 is, and that is the answer the interviewer wants to hear. If the 40 is a typing error or a one-off glitch in a price feed, the median is the honest summary: remove the 40 and the mean falls to 2.5, while the median barely moves. If the 40 is real, say the one big winning day in a strategy's P&L, the mean is the number that matters, because your total profit is the sum of the days and the sum is five times the mean. A robust estimator that ignores the big day would tell you the strategy earns 3 a day when it actually earns 10.

    Where does this show up in a quant job?

    Every regression makes this choice silently. Ordinary least squares minimises squared error and so fits conditional means; one extreme observation can swing the line. Least absolute deviation, or quantile regression at the 50% level, fits conditional medians and shrugs off the same extreme point. In practice, desks winsorise or clip returns before fitting squared-error models, or switch to a Huber loss that is squared near zero and linear in the tails. The limitation is that none of these fixes is free: every one of them throws away some of the information in genuine large moves.

    Where candidates lose it

    The common slip is to answer 10 for both, because the mean feels like the default best guess. The two losses answer different questions, and the interviewer is checking that you know squared error chases the outlier and absolute error does not.

    The second loss is stopping at the arithmetic. Say what the 40 might be, a data error or a real big day, and which loss fits each case. That judgement is the point of the question.

    What the interviewer asks next

    • Which constant minimises the maximum absolute error, and what is that loss called?
    • Add a sixth point at 5. What happens to the median, and to the set of absolute-error minimisers?
    • Why does Lasso use an absolute-value penalty, and how is that related to this puzzle?

    Asked at Tower Research Capital, Trading, Princeton, 2018 (Wall Street Oasis): What if instead of minimizing mean squared error we look at mean absolute error?

  7. 077An equity index stands at 20,000 and its implied volatility is 18% a year. Where do you think it closes in four months? Give a central value and a 90% range you would be willing to make a market around.Market making, betting and sizingHardMSMorgan StanleyTokyo · 2025

    Try it first

    Roughly how wide is a 90% range for the index four months out?

    Show the worked solution

    Centre on today's level, about 20,000, with a 90% range of roughly 16,800 to 23,600. Four months is a third of a year, so one standard deviation is 18% x sqrt(1/3), about 10.4%. In log terms the 90% band is 1.645 of those either side, which gives 16,767 and 23,601. The median sits a little below 20,000 and the upside tail is longer than the downside.

    Why is a single number the wrong answer?

    Ask a cab driver how long the airport run takes and a good one says forty minutes, maybe an hour in traffic. The range is the useful part, because you plan your flight around it. A trading interviewer asking where an index closes wants a distribution, because a market maker quotes against the spread of outcomes, not against a guess. Your central value should not be a view on the economy either: with no edge, the best central estimate of a traded index is roughly its forward, which for four months is close to today's 20,000 once financing and dividends roughly offset.

    The width comes from the implied volatility the market already quotes. Volatility grows with the square root of time, because independent daily moves add their variances, not their standard deviations. Four months is a third of a year, so one standard deviation is 18% x sqrt(1/3) = 10.4%, about 2,078 index points.

    The honest forecast is a distribution, 10.4% wide per standard deviation14,00016,00018,00020,00022,00024,00026,0005th pct 16,76795th pct 23,601median 19,892, mean 20,000middle 90%-3,233 points+3,601 pointsone sd: 18% x sqrt(1/3) = 10.4%
    With 18% volatility over four months, the index's middle 90% runs from about 16,767 to 23,601, which is 3,233 points below today's level and 3,601 points above, because a lognormal distribution stretches further up than down.
    The relationship
    ST=S0 e−σ2T/2+σTZ5th, 95th pct=S0 e−σ2T/2∓1.645 σTS_T = S_0\, e^{-\sigma^2 T/2 + \sigma\sqrt{T} Z} \qquad \text{5th, 95th pct} = S_0\, e^{-\sigma^2T/2 \mp 1.645\,\sigma\sqrt{T}}
    S_0today's level, 20,000
    sigmaimplied volatility, 0.18 a year
    Ttime in years, 1/3
    Za standard normal draw
    What it says in wordsLog returns are normal with a standard deviation of sigma times root T, and a small drift correction keeps the mean at today's level.

    Why is the range lopsided, and where does the median sit?

    A fall of 10% and a rise of 10% are not mirror images in log space. Normal log returns make the upside tail longer: the 90% band stretches 3,601 points up but only 3,233 points down. The same convexity pushes the median below the mean: if the mean is 20,000, the median is 20,000 x exp(-sigma squared T / 2), about 19,892. That gap of about 108 points is small here, but it grows with volatility and time, and a candidate who names it shows they know the difference between the most central outcome and the average one.

    What would you add before quoting a market on it?

    Two honest caveats. Implied volatility is a price, not a forecast: it tends to sit above the volatility that is later realised, because option sellers charge for bearing crash risk, so the band built from it is usually a little wide. Against that, real index returns have fatter tails than the lognormal, so the 5% tails are more likely to hold a larger move than the curve suggests. Say both, then give your market: a tight two-way price around 20,000 if asked for the level, and the 90% band as the range you would sell outside of.

    Where candidates lose it

    The first loss is scaling volatility linearly with time: a third of 18% is 6%, which gives a band far too narrow. Volatility scales with the square root of time, so four months is about 10.4%, not 6%.

    The second loss is answering with a macro story and a point forecast. The interviewer wants you to use the price the market already gives you, implied volatility, and to say that the honest answer is a distribution with a lopsided shape.

    What the interviewer asks next

    • What 90% range would you give for one week out?
    • How would the range change if implied volatility jumped to 30%?
    • If you had to bet on the index finishing above 22,000, what fair probability would you quote?

    Asked at Morgan Stanley, Sales and Trading, Tokyo, 2025 (Wall Street Oasis): What do you think this index will close at by the end of the year (4 months from now)

  8. 078Let A be the 2 by 2 matrix with 2 on the diagonal and 1 off the diagonal. Compute A to the power 10 without multiplying it out ten times.Correlation, regression and linear algebraCoreQuant researchQuant trading

    Try it first

    What is the top-left entry of A^10?

    Show the worked solution

    A^10 has 29,525 on the diagonal and 29,524 off it. A has eigenvalue 3 along (1, 1) and eigenvalue 1 along (1, -1). Writing A = Q D Q^T with D = diag(3, 1), the tenth power is Q D^10 Q^T, and only the numbers 3 and 1 get raised to the tenth. The entries are (3^10 + 1)/2 and (3^10 - 1)/2.

    Why look for eigenvectors at all?

    Think of a photocopier set to 300% on one axis and 100% on the other. Copy a copy ten times and you do not need to simulate every pass: that axis is 3 to the tenth times longer and the other is unchanged. An eigenvector is a direction the matrix only stretches, so applying the matrix ten times along it is just multiplying by the eigenvalue ten times. Symmetric matrices always have a full set of such directions at right angles, which is what makes this matrix easy.

    Find them by inspection. Adding the two rows of A gives 3 in each, so A(1, 1) = (3, 3): eigenvalue 3. Subtracting gives 1, so A(1, -1) = (1, -1): eigenvalue 1. The trace is 4 and the determinant is 3, and 3 + 1 = 4 and 3 x 1 = 3, which confirms both in one line.

    Two directions the matrix only stretches: powers become powers of numbersxy(1, 1)A(1, 1) = (3, 3)(1, -1) = A(1, -1)eigenvalue 3 along (1, 1); eigenvalue 1 along (1, -1)A = Q D Q^TQ holds the unit eigenvectors, D = diag(3, 1)A^10 = Q D^10 Q^Tthe Q^T Q pairs in the middle cancelD^10 = diag(59,049, 1)only two numbers get raised to the 10thA^10: diagonal (3^10 + 1)/2, off it (3^10 - 1)/229,52529,52429,52429,525
    The matrix stretches the direction (1, 1) by a factor of 3 and leaves (1, -1) unchanged, so A to the tenth stretches them by 59,049 and 1, and converting back to ordinary coordinates gives 29,525 on the diagonal and 29,524 off it.
    The relationship
    A=Q(3001)QT,  Q=12(111−1)⇒A10=12(310+1310−1310−1310+1)A = Q\begin{pmatrix}3&0\\0&1\end{pmatrix}Q^{T},\; Q=\tfrac{1}{\sqrt2}\begin{pmatrix}1&1\\1&-1\end{pmatrix} \Rightarrow A^{10} = \tfrac12\begin{pmatrix}3^{10}+1 & 3^{10}-1\\ 3^{10}-1 & 3^{10}+1\end{pmatrix}
    Qthe matrix whose columns are the unit eigenvectors
    Dthe diagonal matrix of eigenvalues, 3 and 1
    Q^Tthe transpose of Q, which is also its inverse
    What it says in wordsRotate into the eigenvector directions, raise each eigenvalue to the tenth, and rotate back.

    Is there an even faster route for this particular matrix?

    Yes. Write A = I + J, where J is the all-ones matrix. J squared is 2J, so every power of J is a multiple of J, and (I + J)^n collapses to I + ((3^n - 1)/2) J. For n = 10 that is I + 29,524 J, which gives 29,525 on the diagonal and 29,524 off it: the same answer, and a good cross-check to say aloud. A brute-force multiplication in code agrees exactly.

    Say why this matters on a desk. A covariance matrix with equal variances and one common correlation has exactly this shape, and its eigenvectors are the market direction and the spread directions. Powers of transition matrices in Markov chains are computed the same way, and the eigenvalue closest to 1 tells you how fast the chain forgets where it started.

    Where candidates lose it

    The fast wrong answer raises each entry to the tenth, giving 1,024 on the diagonal and 1 off it. Matrix multiplication mixes rows and columns, so entries do not power separately; A squared already has 5 on the diagonal, not 4.

    The second loss is diagonalising correctly and then fumbling the conversion back. The Q matrix carries a 1/sqrt(2) on each side, which becomes the factor of one half in the final answer. Check with the trace: the diagonal entries of A^10 must sum to 3^10 + 1.

    What the interviewer asks next

    • What is A^n as n grows large, after dividing by 3^n?
    • Compute the square root of A, a symmetric matrix B with B squared equal to A.
    • Generalise: an n by n matrix with a on the diagonal and b everywhere else. What are its eigenvalues?
  9. 079You may draw numbers uniform on 0 to 1, one after another. Each draw costs 0.02, and when you stop you keep the last number drawn. What is your optimal stopping threshold, and what is the game worth?Expected value and optimal stoppingHardQuant tradingQuant research

    Try it first

    What threshold should you stop at?

    Show the worked solution

    Stop at the first draw of 0.8 or more; the game is worth 0.8 after all fees. Holding x, one more draw improves you by (1 - x)^2 / 2 on average, which equals the 0.02 fee at x = 0.8. Check: you expect 5 draws costing 0.10 in total, and the draw you keep averages 0.90, so the net value is 0.80.

    How do you decide whether one more draw is worth it?

    Think of hunting for a flat. Each viewing costs you an evening. If the flat in hand is already good, another viewing rarely beats it, and when it does it only beats it by a little. The value of one more look is the chance of beating what you hold times the average margin when you do, and you stop when that falls below the price of looking. With a uniform draw and a current value x, the chance of beating x is 1 - x and the average margin is (1 - x)/2, so the expected gain is (1 - x)^2 / 2.

    Draw again only while the expected improvement beats the 0.02 feegain above cost:draw againgain below cost:stop and keep xcost of a draw, 0.02gain from one more draw = (1 - x)^2 / 2threshold x* = 0.80.50.60.70.80.91.00.050.100current value x
    The expected gain from one more draw, (1 - x)^2 / 2, falls below the 0.02 cost exactly at x = 0.8, so you keep drawing while you hold less than 0.8 and stop at the first draw above it.
    The relationship
    V=E[max⁡(U,V)]−c=V+(1−V)22−c  ⇒  (1−V)22=c  ⇒  V=1−2c=0.8V = \mathbb{E}[\max(U, V)] - c = V + \frac{(1-V)^2}{2} - c \;\Rightarrow\; \frac{(1-V)^2}{2} = c \;\Rightarrow\; V = 1 - \sqrt{2c} = 0.8
    Vthe value of the game before paying for the next draw
    Uthe next uniform draw
    cthe cost per draw, 0.02
    What it says in wordsThe game is worth one draw plus the option to walk away with it, less the fee, and solving that gives both the threshold and the value, 0.8.

    Why are the threshold and the value the same number?

    Because the game has no memory. After a disappointing draw you are back where you started, facing the same game worth V. You should accept a draw exactly when it beats what the fresh game is worth, so the threshold equals the value. This is the same logic as a reservation price: you walk away from any offer below what the next round is worth to you. The arithmetic check is worth saying: with threshold 0.8 a draw succeeds with probability 0.2, so you expect 5 draws and 0.10 of fees, and an accepted draw is uniform on 0.8 to 1 with mean 0.9. The net is 0.9 - 0.10 = 0.8. A seeded simulation of 200,000 games gives 0.800.

    What happens as the cost changes?

    The threshold is 1 - sqrt(2c), so it responds to the square root of the cost. A fee of 0.005, four times cheaper, only moves the threshold from 0.8 to 0.9. At a fee of 0.5 or more, the threshold hits zero and you take the first draw, because even the first draw's average of 0.5 barely covers what you paid. The limitation of the model is that draws are independent and the distribution is known; if you were learning the distribution as you drew, the first few draws would be worth more than this rule says.

    Where candidates lose it

    The common wrong threshold is 0.98, from reasoning that one more draw is worth it whenever the cost is less than what you could gain at best. That compares the fee with the best case, not with the average improvement, and it draws far too many times.

    The second loss is getting 0.8 as a threshold and then quoting the value as the average of the kept draw, 0.9. The fees paid on the way, 0.10 on average, come off. Say the check out loud: 0.9 minus 0.1 is 0.8.

    What the interviewer asks next

    • What if you are only allowed at most two draws in total?
    • What if each draw is uniform on 0 to 100 and costs 1?
    • Now the fee is charged only on draws after the first. What changes?
  10. 080A path moves one unit right or one unit up at a time, from (0,0) to (6,4). Every shortest path is equally likely. The point (3,2) is blocked. How many valid paths remain, and what is the probability that a random shortest path avoids the blocked point?Counting and combinatoricsCoreSusquehanna International GroupLondon · 2026

    Try it first

    How many of the shortest paths pass through (3,2)?

    Show the worked solution

    110 paths avoid the block, so the probability is 110/210 = 11/21, about 52.4%. All shortest paths use 6 rights and 4 ups in some order: C(10,4) = 210. Paths through (3,2) combine 10 ways in with 10 ways out, 100 in all. Subtract, and 110 survive.

    How do you count all the shortest paths?

    A shortest path is a string of 10 moves with exactly 6 rights and 4 ups, like a ten-letter word made of R and U. Choosing which 4 of the 10 slots are ups fixes the path completely, so there are C(10,4) = 210 shortest paths. That is the whole sample space, and every one of those strings is equally likely by the question's rule.

    Why multiply for the paths through the blocked point?

    Think of a trip from home to the office with a stop at a coffee shop. If there are 10 routes to the shop and 10 routes from the shop to the office, there are 10 x 10 = 100 full trips, because every first half pairs with every second half. Paths through (3,2) split the same way: C(5,2) = 10 in and C(5,2) = 10 out, so 100 of the 210 pass through the block. Subtract and 110 remain.

    Write the count at every point: left plus below, with the block set to zero11111123451361015140blocked1025155154016112666171844110start (0,0)end (6,4)All shortest pathsC(10,4) = 210Through (3,2)C(5,2) x C(5,2) = 100Avoiding the block210 - 100 = 110Probability of avoiding110/210 = 11/21 = 52.4%
    Adding the count from the left and the count from below at every point, with the blocked point set to zero, gives 110 paths at (6,4), which matches 210 total paths minus the 100 that pass through (3,2).
    The relationship
    (104)−(52)(52)=210−100=110,P=110210=1121\binom{10}{4} - \binom{5}{2}\binom{5}{2} = 210 - 100 = 110, \qquad P = \frac{110}{210} = \frac{11}{21}
    C(10,4)ways to place 4 ups among 10 moves
    C(5,2)ways to place 2 ups among the 5 moves on each side of the block
    What it says in wordsCount everything, subtract the paths forced through the block, and divide by everything.

    The grid method in the figure is the check, and it is also what you would code. Each point's count is the count from the left plus the count from below, because the last step into any point came from one of those two neighbours. Setting the block to zero removes every path through it automatically, and the same method handles several blocks, where the subtraction formula needs inclusion and exclusion.

    Does the answer change if the walker flips a coin at each step?

    Yes, and interviewers often ask this next. If the walker flips a fair coin for right or up at each step, any visit to (3,2) happens on move five, and a walker still able to get there has not yet touched the top or right edge, so all five of those moves were free coin flips. The chance of standing on (3,2) after five flips is C(5,2)/2^5 = 10/32 = 5/16, so the coin-flip walker avoids the block with probability 11/16, about 68.8%, well above 11/21. Choosing uniformly among complete paths is not the same as flipping coins: conditioning on the end point (6,4) pulls paths towards the diagonal that leads there, and (3,2) sits on it. Say which model the question means before you answer.

    Where candidates lose it

    The usual slip is adding the ways in and out, 10 + 10 = 20, instead of multiplying. Paths through a point are pairs of half-paths, and pairs multiply.

    The second loss is quietly switching models, treating each step as a coin flip while using the uniform-path count, or the reverse. State that the question picks among all 210 shortest paths with equal chance, and the answer 11/21 follows.

    What the interviewer asks next

    • What if both (3,2) and (2,3) are blocked?
    • How many shortest paths pass through (3,2) or (4,3), counting each path once?
    • How many paths from (0,0) to (6,4) never go above the line y = x?

    Asked at Susquehanna International Group, Quantitative Research, London, 2026 (Wall Street Oasis): Probability about crossing from (0,0) to (6,4). Some point in the middle cannot pass through

← PreviousPage 8 of 10
  1. 1
  2. …
  3. 7
  4. 8
  5. 9
  6. 10
Next →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.