Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
Explore NISM prep
Series-VIII · Equity DerivativesSeries-XII · Securities Markets FoundationSeries-V-A · Mutual Fund DistributorsSeries-XV · Research AnalystSeries-XIX-E · Category III AIF ManagersSeries-XIX-D · Category I & II AIF ManagersSeries-XIX-C · Alternative Investment Fund ManagersSeries-XVI · Commodity DerivativesSeries-VI · Depository OperationsSeries-II-A · Registrars & Transfer AgentsSeries-I · Currency DerivativesSeries-VII · Securities Operations & Risk Management
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant puzzles, solved step by step

Puzzles
100
Traced to a firm
71
Topics
12
Hard
30
Topic
All topicsLogic and algorithmic reasoning10Conditional probability and Bayes7Counting and combinatorics8Continuous and geometric probability9Correlation, regression and linear algebra9Market making, betting and sizing9Expected value and optimal stopping9Statistics and estimation9Pricing, options and index maths7Games and strategic reasoning8Markov chains and random walks7Mental maths and number sense8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 51–60 of 71 · filtered from 100Clear filters
  1. 075You roll two fair dice and are paid the product of the two faces in rupees. What is the expected payout?Expected value and optimal stoppingWarm upWolverine TradingChicago · 2024

    Try it first

    What is the expected product?

    Show the worked solution

    Rs 12.25. The two dice are independent, so the expected product equals the product of the expected faces: 3.5 x 3.5 = 12.25. A check on the 6 by 6 grid: row i averages 3.5i, and the six row averages, 3.5 to 21, average 12.25. If both numbers came from one die, the answer would be E[X squared] = 91/6, about 15.17, which is not the question.

    Why can you multiply the averages?

    Picture a shop whose daily takings are footfall times average spend, where the two have nothing to do with each other. Over a year, average takings are average footfall times average spend. When two quantities are independent, the average of their product is the product of their averages, because knowing one tells you nothing about how big the other will be. That fails the moment they move together: a crowded day with bigger spending lifts the product above the product of averages. Two dice thrown separately are the textbook independent pair.

    Average of the grid = row average x column average = 3.5 x 3.5123456second die1123456avg 3.5224681012avg 73369121518avg 10.544812162024avg 14551015202530avg 17.5661218243036avg 21first diedark cells: above 12.25 (13 of 36)Two independent diceE[XY] = E[X] E[Y]3.5 x 3.5 = 12.25Not the same die squaredE[X^2] = 91/6= 15.17, higher by the varianceThe gap, 35/12 = 2.92, is one die's variance
    Across the 36 equally likely cells of the multiplication grid each row averages its row number times 3.5, so the grand average is 3.5 x 3.5 = 12.25, even though 23 of the 36 products are below it; squaring a single die would give 15.17 instead.

    How do you check 12.25 by brute force quickly?

    Sum the grid by rows. Row i of the multiplication table sums to i x 21, so the whole grid sums to 21 x 21 = 441, and 441 / 36 = 12.25. That is the product rule made visible: the grid's total factors into the first die's total times the second die's total. Note too that the product is skewed: only 18 different values appear, most cells are small, and the big ones, 25, 30 and 36, pull the mean up, so {BELOW75} of the 36 outcomes pay less than the average.

    The relationship
    E[XY]=136∑i=16∑j=16ij=21×2136=44136=12.25=E[X] E[Y]E[XY] = \frac{1}{36}\sum_{i=1}^{6}\sum_{j=1}^{6} ij = \frac{21 \times 21}{36} = \frac{441}{36} = 12.25 = E[X]\,E[Y]
    X, Ythe two independent die faces
    21the sum of the faces 1 to 6
    441the sum of all 36 products
    What it says in wordsThe grid's total is the product of the two dice's totals, so the average product is the product of the averages.

    What is the interviewer likely to ask next?

    The usual next step changes the dependence. If the same die is used for both numbers, you are paid the square of one roll, and its expectation is 91/6, about 15.17, higher by exactly the variance of one die, 35/12. That gap is the covariance at work: E[XY] = E[X]E[Y] + Cov(X, Y). Then comes the price: if you would pay to play, you should quote around 12.25 and note that the payout's standard deviation is about 8.94, so a single play is very noisy relative to its mean.

    Where candidates lose it

    The trap is computing E[X squared] instead of E[XY], getting 15.17, usually because the candidate thinks of rolling one die and squaring. The other slip is answering 18 by taking the midpoint of 1 to 36.

    Say independent and give 3.5 x 3.5 inside five seconds; then offer the row-sum check, 21 x 21 / 36, so the interviewer sees you can verify it.

    What the interviewer asks next

    • What is the expected payout if you are paid the square of a single roll?
    • What is the variance of the product of two dice?
    • You may reroll one of the two dice once after seeing both. What is the game worth?

    Asked at Wolverine Trading, Sales and Trading, Chicago, 2024 (Wall Street Oasis): don't think I was what they were looking for. Questions on EV & Dice.

  2. 076You must predict a quantity with a single constant, and your data are 1, 2, 3, 4 and 40. Which constant minimises the mean squared error, which minimises the mean absolute error, and what does the difference tell you?Statistics and estimationWarm upTower Research CapitalPrinceton · 2018

    Try it first

    Before you calculate: which pair of constants is right?

    Show the worked solution

    The mean, 10, minimises squared error; the median, 3, minimises absolute error. Squared error grows with the square of a miss, so the single 40 pulls the best constant towards it. Absolute error charges every unit of miss equally, so the best constant sits in the middle of the pack. Choosing the loss is choosing which of the two statistics you estimate.

    Why does each loss land on a different constant?

    Picture five friends choosing a meeting point on a straight road: four live at the 1, 2, 3 and 4 km marks and one at the 40 km mark. If the goal is the smallest total travel, you meet at 3 km: moving towards the far friend saves them one km per km moved but costs the other four one km each. If the goal is the smallest total of squared travel, the far friend's 37 km squared, 1,369, dominates everything and the meeting point slides to 10 km. Absolute error balances the count of points on each side, which is the median; squared error balances the total distance on each side, which is the mean.

    One outlier, two losses, two different best constants010203040the outlier, 40median 3mean 10Mean squared errorlime dot, the lowest: 226 at c = 1005101520constant c you predictc = 3 scores 275Mean absolute errorlime dot, the lowest: 8.2 at c = 305101520constant c you predictc = 10 scores 12
    Squared error is lowest at the mean, 226 at c = 10, and the median scores 275 on it; absolute error is lowest at the median, 8.2 at c = 3, and the mean scores 12 on it, so each constant is poor under the other loss.
    The relationship
    ddc∑i(xi−c)2=−2∑i(xi−c)=0⇒c=xˉddc∑i∣xi−c∣=#{xi<c}−#{xi>c}=0⇒c=median\frac{d}{dc}\sum_i (x_i-c)^2 = -2\sum_i (x_i - c) = 0 \Rightarrow c=\bar x \qquad \frac{d}{dc}\sum_i |x_i-c| = \#\{x_i<c\} - \#\{x_i>c\} = 0 \Rightarrow c = \text{median}
    x_ithe five data points
    cthe constant you predict
    #{x_i < c}how many points sit below c
    What it says in wordsSquared error is flat where the total distance above equals the total below, which is the mean; absolute error is flat where the count above equals the count below, which is the median.

    So which constant should you actually use?

    It depends on what the 40 is, and that is the answer the interviewer wants to hear. If the 40 is a typing error or a one-off glitch in a price feed, the median is the honest summary: remove the 40 and the mean falls to 2.5, while the median barely moves. If the 40 is real, say the one big winning day in a strategy's P&L, the mean is the number that matters, because your total profit is the sum of the days and the sum is five times the mean. A robust estimator that ignores the big day would tell you the strategy earns 3 a day when it actually earns 10.

    Where does this show up in a quant job?

    Every regression makes this choice silently. Ordinary least squares minimises squared error and so fits conditional means; one extreme observation can swing the line. Least absolute deviation, or quantile regression at the 50% level, fits conditional medians and shrugs off the same extreme point. In practice, desks winsorise or clip returns before fitting squared-error models, or switch to a Huber loss that is squared near zero and linear in the tails. The limitation is that none of these fixes is free: every one of them throws away some of the information in genuine large moves.

    Where candidates lose it

    The common slip is to answer 10 for both, because the mean feels like the default best guess. The two losses answer different questions, and the interviewer is checking that you know squared error chases the outlier and absolute error does not.

    The second loss is stopping at the arithmetic. Say what the 40 might be, a data error or a real big day, and which loss fits each case. That judgement is the point of the question.

    What the interviewer asks next

    • Which constant minimises the maximum absolute error, and what is that loss called?
    • Add a sixth point at 5. What happens to the median, and to the set of absolute-error minimisers?
    • Why does Lasso use an absolute-value penalty, and how is that related to this puzzle?

    Asked at Tower Research Capital, Trading, Princeton, 2018 (Wall Street Oasis): What if instead of minimizing mean squared error we look at mean absolute error?

  3. 077An equity index stands at 20,000 and its implied volatility is 18% a year. Where do you think it closes in four months? Give a central value and a 90% range you would be willing to make a market around.Market making, betting and sizingHardMSMorgan StanleyTokyo · 2025

    Try it first

    Roughly how wide is a 90% range for the index four months out?

    Show the worked solution

    Centre on today's level, about 20,000, with a 90% range of roughly 16,800 to 23,600. Four months is a third of a year, so one standard deviation is 18% x sqrt(1/3), about 10.4%. In log terms the 90% band is 1.645 of those either side, which gives 16,767 and 23,601. The median sits a little below 20,000 and the upside tail is longer than the downside.

    Why is a single number the wrong answer?

    Ask a cab driver how long the airport run takes and a good one says forty minutes, maybe an hour in traffic. The range is the useful part, because you plan your flight around it. A trading interviewer asking where an index closes wants a distribution, because a market maker quotes against the spread of outcomes, not against a guess. Your central value should not be a view on the economy either: with no edge, the best central estimate of a traded index is roughly its forward, which for four months is close to today's 20,000 once financing and dividends roughly offset.

    The width comes from the implied volatility the market already quotes. Volatility grows with the square root of time, because independent daily moves add their variances, not their standard deviations. Four months is a third of a year, so one standard deviation is 18% x sqrt(1/3) = 10.4%, about 2,078 index points.

    The honest forecast is a distribution, 10.4% wide per standard deviation14,00016,00018,00020,00022,00024,00026,0005th pct 16,76795th pct 23,601median 19,892, mean 20,000middle 90%-3,233 points+3,601 pointsone sd: 18% x sqrt(1/3) = 10.4%
    With 18% volatility over four months, the index's middle 90% runs from about 16,767 to 23,601, which is 3,233 points below today's level and 3,601 points above, because a lognormal distribution stretches further up than down.
    The relationship
    ST=S0 e−σ2T/2+σTZ5th, 95th pct=S0 e−σ2T/2∓1.645 σTS_T = S_0\, e^{-\sigma^2 T/2 + \sigma\sqrt{T} Z} \qquad \text{5th, 95th pct} = S_0\, e^{-\sigma^2T/2 \mp 1.645\,\sigma\sqrt{T}}
    S_0today's level, 20,000
    sigmaimplied volatility, 0.18 a year
    Ttime in years, 1/3
    Za standard normal draw
    What it says in wordsLog returns are normal with a standard deviation of sigma times root T, and a small drift correction keeps the mean at today's level.

    Why is the range lopsided, and where does the median sit?

    A fall of 10% and a rise of 10% are not mirror images in log space. Normal log returns make the upside tail longer: the 90% band stretches 3,601 points up but only 3,233 points down. The same convexity pushes the median below the mean: if the mean is 20,000, the median is 20,000 x exp(-sigma squared T / 2), about 19,892. That gap of about 108 points is small here, but it grows with volatility and time, and a candidate who names it shows they know the difference between the most central outcome and the average one.

    What would you add before quoting a market on it?

    Two honest caveats. Implied volatility is a price, not a forecast: it tends to sit above the volatility that is later realised, because option sellers charge for bearing crash risk, so the band built from it is usually a little wide. Against that, real index returns have fatter tails than the lognormal, so the 5% tails are more likely to hold a larger move than the curve suggests. Say both, then give your market: a tight two-way price around 20,000 if asked for the level, and the 90% band as the range you would sell outside of.

    Where candidates lose it

    The first loss is scaling volatility linearly with time: a third of 18% is 6%, which gives a band far too narrow. Volatility scales with the square root of time, so four months is about 10.4%, not 6%.

    The second loss is answering with a macro story and a point forecast. The interviewer wants you to use the price the market already gives you, implied volatility, and to say that the honest answer is a distribution with a lopsided shape.

    What the interviewer asks next

    • What 90% range would you give for one week out?
    • How would the range change if implied volatility jumped to 30%?
    • If you had to bet on the index finishing above 22,000, what fair probability would you quote?

    Asked at Morgan Stanley, Sales and Trading, Tokyo, 2025 (Wall Street Oasis): What do you think this index will close at by the end of the year (4 months from now)

  4. 080A path moves one unit right or one unit up at a time, from (0,0) to (6,4). Every shortest path is equally likely. The point (3,2) is blocked. How many valid paths remain, and what is the probability that a random shortest path avoids the blocked point?Counting and combinatoricsCoreSusquehanna International GroupLondon · 2026

    Try it first

    How many of the shortest paths pass through (3,2)?

    Show the worked solution

    110 paths avoid the block, so the probability is 110/210 = 11/21, about 52.4%. All shortest paths use 6 rights and 4 ups in some order: C(10,4) = 210. Paths through (3,2) combine 10 ways in with 10 ways out, 100 in all. Subtract, and 110 survive.

    How do you count all the shortest paths?

    A shortest path is a string of 10 moves with exactly 6 rights and 4 ups, like a ten-letter word made of R and U. Choosing which 4 of the 10 slots are ups fixes the path completely, so there are C(10,4) = 210 shortest paths. That is the whole sample space, and every one of those strings is equally likely by the question's rule.

    Why multiply for the paths through the blocked point?

    Think of a trip from home to the office with a stop at a coffee shop. If there are 10 routes to the shop and 10 routes from the shop to the office, there are 10 x 10 = 100 full trips, because every first half pairs with every second half. Paths through (3,2) split the same way: C(5,2) = 10 in and C(5,2) = 10 out, so 100 of the 210 pass through the block. Subtract and 110 remain.

    Write the count at every point: left plus below, with the block set to zero11111123451361015140blocked1025155154016112666171844110start (0,0)end (6,4)All shortest pathsC(10,4) = 210Through (3,2)C(5,2) x C(5,2) = 100Avoiding the block210 - 100 = 110Probability of avoiding110/210 = 11/21 = 52.4%
    Adding the count from the left and the count from below at every point, with the blocked point set to zero, gives 110 paths at (6,4), which matches 210 total paths minus the 100 that pass through (3,2).
    The relationship
    (104)−(52)(52)=210−100=110,P=110210=1121\binom{10}{4} - \binom{5}{2}\binom{5}{2} = 210 - 100 = 110, \qquad P = \frac{110}{210} = \frac{11}{21}
    C(10,4)ways to place 4 ups among 10 moves
    C(5,2)ways to place 2 ups among the 5 moves on each side of the block
    What it says in wordsCount everything, subtract the paths forced through the block, and divide by everything.

    The grid method in the figure is the check, and it is also what you would code. Each point's count is the count from the left plus the count from below, because the last step into any point came from one of those two neighbours. Setting the block to zero removes every path through it automatically, and the same method handles several blocks, where the subtraction formula needs inclusion and exclusion.

    Does the answer change if the walker flips a coin at each step?

    Yes, and interviewers often ask this next. If the walker flips a fair coin for right or up at each step, any visit to (3,2) happens on move five, and a walker still able to get there has not yet touched the top or right edge, so all five of those moves were free coin flips. The chance of standing on (3,2) after five flips is C(5,2)/2^5 = 10/32 = 5/16, so the coin-flip walker avoids the block with probability 11/16, about 68.8%, well above 11/21. Choosing uniformly among complete paths is not the same as flipping coins: conditioning on the end point (6,4) pulls paths towards the diagonal that leads there, and (3,2) sits on it. Say which model the question means before you answer.

    Where candidates lose it

    The usual slip is adding the ways in and out, 10 + 10 = 20, instead of multiplying. Paths through a point are pairs of half-paths, and pairs multiply.

    The second loss is quietly switching models, treating each step as a coin flip while using the uniform-path count, or the reverse. State that the question picks among all 210 shortest paths with equal chance, and the answer 11/21 follows.

    What the interviewer asks next

    • What if both (3,2) and (2,3) are blocked?
    • How many shortest paths pass through (3,2) or (4,3), counting each path once?
    • How many paths from (0,0) to (6,4) never go above the line y = x?

    Asked at Susquehanna International Group, Quantitative Research, London, 2026 (Wall Street Oasis): Probability about crossing from (0,0) to (6,4). Some point in the middle cannot pass through

  5. 081A six-chamber revolver holds two bullets in adjacent chambers. The cylinder is spun, the interviewer pulls the trigger on himself and it clicks on an empty chamber. It is now your turn. Do you want him to spin the cylinder again first, or pull straight away?Games and strategic reasoningCoreSchonfeldCentral · 2022

    Try it first

    Which gives you the better chance of surviving?

    Show the worked solution

    Do not spin: pulling straight away survives with probability 3/4, against 2/3 with a spin. The click tells you the hammer sat on one of the four empty chambers. Because the two bullets are adjacent, the four empties form a run, and only the last empty in that run is followed by a bullet. A spin throws that information away and resets you to 4 empties out of 6.

    What does the click actually tell you?

    Think of a row of six houses where two neighbours keep dogs. You knocked at a random house and no dog barked. If you now try the next house along, you are only in trouble if you had knocked on the one house sitting just before the dogs. The click narrows the hammer's position to the four empty chambers, and the question becomes how many of those four have a bullet immediately after them. With the bullets side by side, the empties run 3, 4, 5, 6 in firing order, and only chamber 6 hands over to a bullet.

    After a click, only one of the four empty chambers sits in front of a bullet123456firing orderruns clockwiseloadedempty, bullet nextempty, empty nextGiven the click, the hammer now sits onone of the four empties, each equally likely.Pull straight away75.0%3 of the 4 empties are followed by an emptySpin first66.7%4 of the 6 chambers are emptybar length: the chance you survive your pull
    With two adjacent bullets, only one of the four empty chambers is followed by a bullet, so pulling straight away after a click survives 75% of the time, while a fresh spin survives only 66.7%, four empties out of six.
    The relationship
    P(survive∣click, no spin)=#{empties followed by an empty}#{empties}=34  >  P(survive∣spin)=46=23P(\text{survive}\mid\text{click, no spin}) = \frac{\#\{\text{empties followed by an empty}\}}{\#\{\text{empties}\}} = \frac34 \;>\; P(\text{survive}\mid\text{spin}) = \frac46 = \frac23
    empties followed by an emptychambers 3, 4 and 5 in firing order
    4/6the survival chance of a fresh, random chamber
    What it says in wordsWithout a spin you are conditioning on the click, which helps; a spin forgets it.

    Does the answer depend on the bullets being adjacent?

    Completely, and that is the follow-up most interviewers ask. If the two bullets are not next to each other, the empties split into two runs, two empties now sit in front of a bullet, and pulling straight away survives only 1/2, worse than the 2/3 of a spin. The same count works for bullets one apart or directly opposite: in both cases two of the four empties are followed by a bullet. So the rule is not spin or do not spin; it is count the empties that border a bullet and compare with a fresh spin.

    What is the general lesson for a trading interview?

    A random reset destroys information, and information has value only if the structure of the problem lets you use it. Here the structure is clustering: the bullets sit together, so a safe chamber is likely followed by another safe one. Markets have the same feature in volatility: a calm day tends to be followed by a calm day, so conditioning on what just happened beats assuming each day is a fresh draw. Say that link in one line after the arithmetic.

    Where candidates lose it

    The common error is to treat both options as a fresh draw and say it makes no difference, or to say a spin is safer because it resets the odds. Both ignore the click, which is the one piece of information you were given.

    The second loss is getting 3/4 without seeing that it hinges on adjacency. Say that with the bullets apart the answer flips to spin, and give the count, two bordering empties out of four.

    What the interviewer asks next

    • The two bullets are placed in random chambers, not necessarily adjacent. Spin or not?
    • Three adjacent bullets and a click. Spin or not?
    • After two clicks in a row without spins, what is your survival chance on the third pull?

    Asked at Schonfeld, Quantitative Research, Central, 2022 (Wall Street Oasis): Coding, requires to know DP and divde and conquer., Russian Roulette

  6. 082Two independent waiting times are each exponentially distributed with a mean of one minute. What is the probability that their total is less than one minute?Continuous and geometric probabilityCoreCitadelChicago · 2025

    Try it first

    Pick the closest value.

    Show the worked solution

    1 - 2/e, about 26.4%. Convolving two exponential densities gives the total the density x e^-x, which starts at zero and peaks at one minute. Its area below one minute is 1 - 2/e. The same number drops out of the Poisson view: the total is under a minute exactly when at least two arrivals land in the first minute of a rate-one Poisson process.

    Why does adding two waits change the shape so much?

    Suppose you need two buses, one after the other, and each arrives on average a minute after you reach its stop. Catching the first bus quickly is common; catching both quickly is rare, because both have to cooperate. A single exponential wait is most likely near zero, but a sum of two is almost never near zero, so its density starts at zero and rises into a hump. That shift of mass away from zero is why the answer is much smaller than the 63.2% chance that one wait is under a minute.

    Two memoryless waits add up to a hump: little mass near zero012345minutes0.51.0one wait: e^-xsum of two waits: x e^-x26.4%One wait under 11 - 1/e = 63.2%both: 40.0%Total under 11 - 2/e = 26.4%the shaded area
    The single exponential wait puts most of its mass near zero, but the total of two waits has density x e^-x, which starts at zero and peaks at one minute, so only 26.4% of its area, shaded, falls below one minute.
    The relationship
    fS(s)=∫0se−xe−(s−x) dx=s e−s,P(S<1)=∫01s e−s ds=1−2e≈0.264f_{S}(s) = \int_0^s e^{-x}e^{-(s-x)}\,dx = s\,e^{-s}, \qquad P(S<1) = \int_0^1 s\,e^{-s}\,ds = 1 - \frac{2}{e} \approx 0.264
    Sthe total of the two waits
    xthe first wait, which can be anything from 0 to s
    e^{-x}the exponential density with mean 1
    What it says in wordsTo land on a total of s, the first wait takes any value x and the second makes up the rest; adding over all x gives s e^-s.

    Is there a way to get 1 - 2/e without integrating?

    Yes, and it is the cleaner answer to give aloud. Exponential waits with mean one are the gaps between arrivals of a Poisson process with rate one per minute. The second arrival comes before one minute exactly when at least two arrivals land in the first minute, and the Poisson chance of zero or one arrival is e^-1 + e^-1 = 2/e. So the answer is 1 - 2/e, about 26.4%, with no calculus at all.

    Sanity-check the size. Both waits being under a minute has probability (1 - 1/e)^2, about 40.0%, and the total being under a minute is a stricter event, so the answer must be smaller: 26.4% is. The limitation is the independence assumption; if the two waits were driven by the same traffic, they would move together and the total would be more spread out.

    Where candidates lose it

    The frequent wrong answer is (1 - 1/e)^2, about 40%, which is the chance that each wait is under a minute. The question asks about the total, and two waits of 0.7 minutes each pass that test while failing this one.

    The second loss is starting a convolution integral and getting lost in the limits. Say the Poisson route first: at least two arrivals in the first minute, one minus the chance of zero or one.

    What the interviewer asks next

    • What is the probability that the sum of three such waits is under one minute?
    • Given the total is exactly 2 minutes, what is the distribution of the first wait?
    • What is the probability that the first wait is shorter than the second?

    Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis): if i knew this was about convolutions, i would have answered better

  7. 083A market is either calm or stressed each day. A calm day is followed by another calm day with probability 0.8, and a stressed day is followed by another stressed day with probability 0.6. In the long run, what fraction of days are calm?Markov chains and random walksWarm upDRWLondon · 2026

    Try it first

    What share of days are calm in the long run?

    Show the worked solution

    Two thirds of days are calm. In the long run the fraction of days moving from calm to stressed must equal the fraction moving back, so calm share x 0.2 = stressed share x 0.4. Calm days are therefore twice as common as stressed days, 2/3 against 1/3. The same answer comes from spell lengths: calm spells average 5 days and stressed spells 2.5, and 5 / 7.5 = 2/3.

    Why can you balance flows instead of solving equations?

    Think of two rooms at a party. Every few minutes, one in five people in the kitchen wanders to the lounge, and two in five people in the lounge wander back. Once the crowd settles, the numbers crossing each way must match, or one room would keep filling up. In a two-state chain the long-run shares are fixed by one equation: the share of days leaving calm must equal the share of days leaving stressed. Calm leaves at rate 0.2 and stressed at rate 0.4, so calm must hold twice as many days.

    In the long run the flow out of calm equals the flow back inCalmStressed0.80.60.2 calm to stressed0.4 stressed to calmBalance: (2/3) x 0.2 = (1/3) x 0.4 = 2/15 of all dayscalm 2/3 of daysstressed 1/3average calm spell 1/0.2 = 5 days; average stressed spell 1/0.4 = 2.5 days
    Calm days leave at rate 0.2 and stressed days at rate 0.4, so in the long run calm must hold twice as many days as stressed for the flows to balance: two thirds calm and one third stressed, with each flow equal to 2/15 of all days.
    The relationship
    πC (1−0.8)=πS (1−0.6),πC+πS=1  ⇒  πC=0.40.2+0.4=23\pi_C\,(1-0.8) = \pi_S\,(1-0.6), \quad \pi_C + \pi_S = 1 \;\Rightarrow\; \pi_C = \frac{0.4}{0.2+0.4} = \frac23
    pi_Cthe long-run share of calm days
    pi_Sthe long-run share of stressed days
    1 - 0.8the chance a calm day is followed by a stressed one
    What it says in wordsEach state's long-run share is the rate of leaving the other state, divided by the two leaving rates added together.

    How do you check it a second way?

    Use spell lengths. A calm spell ends each day with probability 0.2, so it lasts 1/0.2 = 5 days on average; a stressed spell ends with probability 0.4, so it lasts 2.5 days. The chain alternates calm spell, stressed spell, calm spell, so the calm share is 5 out of every 7.5 days, which is 2/3. Two methods, one answer, in under a minute.

    How fast does the chain forget where it started?

    The transition matrix has a second eigenvalue of 0.8 + 0.6 - 1 = 0.4, and any gap between today's probabilities and the long-run split shrinks by that factor each day. After five days the starting state explains only 0.4^5, about 1%, of the gap, so the answer does not depend on how the week began. A trader would add the limitation: real regimes are not memoryless, and a stress spell that has already lasted a month is not as likely to end tomorrow as one that started yesterday.

    Where candidates lose it

    The trap is answering 80%, reading the chance of staying calm as the share of calm days. The 0.8 describes one step, and the long-run share depends on how quickly both states are left, not just one of them.

    The second slip is setting up a full eigenvector calculation and running out of time. Say the flow balance in one line, then use the spell lengths as the check.

    What the interviewer asks next

    • Today is stressed. What is the probability that the day after tomorrow is calm?
    • What is the expected number of days until the first stressed day, starting calm?
    • If a desk loses Rs 2 lakh on stressed days and makes Rs 1 lakh on calm days, what is its long-run average daily P&amp;L?

    Asked at DRW, Trader Intern Interview, London, 2026 (Wall Street Oasis): consisted of math, statistics, and probability theory (eg. one question was about markov chains

  8. 084A price path runs 100, 120, 90, 130, 80, 110, 140, 112. What are the two largest drawdowns, and how do you compute the maximum drawdown in a single pass through the data?Logic and algorithmic reasoningCoreBalyasny Asset ManagementLondon · 2025

    Try it first

    What is the maximum drawdown of this path?

    Show the worked solution

    The two largest drawdowns are 130 to 80, 38.5%, and 120 to 90, 25.0%. Walk through the prices once, keeping the highest price seen so far. At each step the drawdown is one minus price over that running peak; the maximum drawdown is the largest value seen. For the n largest, close an episode each time a new high is set, record its trough, and sort the episodes.

    What exactly is a drawdown measured from?

    Think of a hiker measuring how far below the highest point reached so far she has dropped. A drop only counts from a summit already climbed, not from a peak further along the trail. A drawdown is the fall from the running maximum to a later price, so the order of the prices matters and the overall high and low cannot simply be paired. Here the 140 comes after the 80, so the tempting 42.9% never happened.

    Track the highest price so far; the drawdown is the fall from it100day 0120day 190day 2130day 380day 4110day 5140day 6112day 7running peak130 to 80-38.5%120 to 90-25.0%140 to 112-20.0%openlargest first
    Tracking the running peak shows three separate drawdowns: 130 to 80 at 38.5%, 120 to 90 at 25.0%, and 140 to 112 at 20.0%, which is still open at the end of the data.
    The relationship
    Mt=max⁡s≤tPs,DDt=1−PtMt,MDD=max⁡tDDtM_t = \max_{s\le t} P_s, \qquad DD_t = 1 - \frac{P_t}{M_t}, \qquad \text{MDD} = \max_t DD_t
    P_tthe price on day t
    M_tthe running peak, the highest price up to day t
    DD_tthe drawdown on day t
    What it says in wordsKeep the highest price so far, measure today's fall from it, and remember the worst fall.

    How do you get the n largest drawdowns rather than just the worst?

    Split the path into episodes. An episode opens at a running peak and closes when the price makes a new high; its size is the fall from that peak to the lowest price inside it. Here the 120 episode closes when the price reaches 130, with a trough of 90, so it is 25%. The 130 episode closes at 140 with a trough of 80, 38.5%. The 140 episode never closes, so report it as open at 20.0%. Sorting the episodes gives the n largest in one pass plus a sort, O(N log N) at worst, and a heap of size n keeps it at O(N log n).

    Say the edge cases, because the interviewer is testing code judgement as much as arithmetic. Two drawdowns from the same peak must not be counted twice: the fall to 80 and the later level of 110 belong to one episode. An episode still open at the end of the data is real risk and should be reported with a flag. And drawdown on a price series is not the same as drawdown on a strategy's cumulative P&L, where you would use the equity curve, not the price.

    Where candidates lose it

    The instinctive error pairs the overall high with the overall low: 140 and 80, 42.9%. That ignores time order; a peak must come before its trough.

    The coding version of the same mistake is returning the n largest daily drawdown values, which for n = 4 would add 15.4%, the day the price sat at 110 below its 130 peak, as a separate event when it is part of the 130 to 80 fall. Group by episode first, then rank.

    What the interviewer asks next

    • How long did the 130 episode last from peak to recovery?
    • Write the one-pass code and state its time and memory cost.
    • Why is maximum drawdown a noisy statistic for comparing two strategies with short track records?

    Asked at Balyasny Asset Management, Quantitative Trading, London, 2025 (Wall Street Oasis): There was an OA with a programming and data science problem. Programming asked to return the n largest drawdowns

  9. 085What are the last two digits of 4 raised to the power 3000?Mental maths and number senseHardBelvedere TradingNew york · 2021

    Try it first

    Which ending is right?

    Show the worked solution

    76. Split 100 into 4 x 25. Any power of 4 is 0 mod 4. Mod 25, Euler's theorem applies because 4 and 25 share no factor, and 3000 is a multiple of phi(25) = 20, so 4^3000 is 1 mod 25. The numbers below 100 that are 1 mod 25 are 1, 26, 51 and 76, and only 76 is divisible by 4.

    Why does the obvious Euler shortcut fail?

    Euler's theorem says a to the power phi(n) is 1 mod n, and phi(100) = 40, so it is tempting to say 4^3000 = (4^40)^75 ends in 01. The theorem needs the base and the modulus to share no factor, and 4 and 100 share a factor of 4, so it does not apply. A quick sense check kills 01 anyway: every power of 4 is divisible by 4, and a number is divisible by 4 exactly when its last two digits are, which 01 is not.

    How do you split the problem so the theorem does apply?

    Think of a clock with 100 hours as two smaller clocks running together, one with 4 hours and one with 25. Knowing where both small clocks point fixes the big one exactly. Mod 4 the answer is 0, since 4^3000 is a multiple of 4; mod 25 the answer is 1, since 4 and 25 share no factor and 3000 is a multiple of phi(25) = 20. Now list the numbers below 100 that are 1 mod 25: 1, 26, 51, 76. Only 76 is a multiple of 4. This step is the Chinese remainder theorem, and naming it earns credit.

    Powers of 4 cycle every 10 steps, and 3000 lands on 7604n=116n=264n=356n=424n=596n=684n=736n=844n=976n=104^n mod 100repeats every 103000 = 10 x 300, and 76 x 76 = 5,776so 4^3000 = (4^10)^300 ends in 76Check with 100 = 4 x 25:mod 44^3000 is 0mod 25 (phi = 20)4^3000 is 11265176numbers below 100 that are 1 mod 25only 76 is also divisible by 4
    The last two digits of 4^n cycle through ten values and the tenth power ends in 76, so every multiple of 10 as an exponent, including 3000, ends in 76; splitting 100 into 4 and 25 confirms it, because 76 is the only number below 100 that is 0 mod 4 and 1 mod 25.
    The relationship
    43000≡0(mod4),43000=(420)150≡1(mod25)  ⇒  43000≡76(mod100)4^{3000} \equiv 0 \pmod 4, \quad 4^{3000} = (4^{20})^{150} \equiv 1 \pmod{25} \;\Rightarrow\; 4^{3000} \equiv 76 \pmod{100}
    mod 4, mod 25the remainders on division by 4 and by 25
    phi(25) = 20how many numbers below 25 share no factor with 25
    What it says in wordsFind the remainder on each small clock, then find the one number below 100 that matches both.

    What is the fastest check if you have a pencil?

    Just list the endings. Multiply each ending by 4 and keep the last two digits: 04, 16, 64, 56, 24, 96, 84, 36, 44, 76, and then 76 x 4 = 304, which ends in 04, so the cycle has length 10. Because 76 x 76 = 5,776 also ends in 76, every power of 76 ends in 76, and 4^3000 = (4^10)^300 must end in 76. On a multiple-choice test, this listing takes about thirty seconds and needs no theorem at all.

    Where candidates lose it

    The trap is applying Euler's theorem with phi(100) = 40 and answering 01. The theorem requires the base and modulus to share no factor, and 4 and 100 do share one.

    The second loss is listing powers without noticing the cycle and running out of time. Say early that the endings must repeat, find the period of 10, and read the answer from 3000 being a multiple of 10.

    What the interviewer asks next

    • What are the last two digits of 7^2026?
    • What are the last three digits of 4^3000?
    • What is the remainder when 2^100 is divided by 7?

    Asked at Belvedere Trading, Equities, New york, 2021 (Wall Street Oasis): It was a 14 question multiple choice test. Some basic number theory (4^3000 modulo 100)

  10. 086A price-weighted index holds three stocks priced 50, 100 and 150, with a divisor of 3. The 150 stock splits 3 for 1. What is the new divisor, and how does a market-cap-weighted index handle the same split?Pricing, options and index mathsWarm upMizuhoHong Kong · 2024

    Try it first

    What must the new divisor be?

    Show the worked solution

    The new divisor is 2. Before the split the prices sum to 300, and 300 / 3 is 100. After a 3 for 1 split the 150 stock trades at 50, the sum is 200, and only a divisor of 2 keeps the index at 100. A cap-weighted index needs no adjustment at all, because a split triples the share count as it cuts the price to a third, leaving market value unchanged.

    Why must the divisor change when nothing about the company changed?

    Cut a pizza into twelve slices instead of four and you have not made more pizza. A stock split does the same to a company: three times the shares, each worth a third. A price-weighted index adds up share prices, so a split drops the sum even though no value was lost, and the divisor must be cut to stop a fake fall in the index. Solve for it by keeping the index level fixed: 200 divided by the new divisor must equal 100, so the divisor is 2.

    A split changes the sum of prices, so the divisor must change to hold the indexBefore: C at 15050stock Aweight 17%100stock Bweight 33%150stock Cweight 50%sum 300 / divisor 3 = index 100After a 3 for 1 split: C at 5050stock Aweight 25%100stock Bweight 50%50stock Cweight 25%sum 200 / divisor 2 = index 100split
    Before the split the prices sum to 300 and the index is 300 / 3 = 100; after C splits 3 for 1 the sum is 200, so the divisor falls to 2 to hold the index at 100, and C's weight falls from 50% to 25%.
    The relationship
    I=∑iPid3003=200d′⇒d′=2I = \frac{\sum_i P_i}{d} \qquad \frac{300}{3} = \frac{200}{d'} \Rightarrow d' = 2
    P_ithe price of stock i
    dthe divisor before the split, 3
    d'the divisor after the split
    What it says in wordsChoose the new divisor so the index is the same the moment after the split as the moment before.

    What else changes in a price-weighted index after the split?

    The weights. In a price-weighted index a stock's weight is its price over the sum of prices, so the expensive stock dominates whatever the size of the company. Before the split C carried 50% of the index; after it, C carries only 25% and B, untouched, jumps to 50%. A 10% rise in C used to add 5 index points; now it adds 2.5. Nothing about C's business changed; the index simply started caring less about it, which is the main criticism of price weighting.

    How do the other common methods treat the split?

    A market-cap-weighted index sums price times shares, and a 3 for 1 split multiplies shares by 3 while dividing price by 3, so the stock's market value, its weight and the index are all unchanged; no divisor adjustment is needed for a split. Cap-weighted divisors still change for events that alter total market value without a price move, such as share issuance, buybacks or a constituent being replaced. An equal-weighted index is also untouched by a split, since weights are reset to equal at each rebalance, but it has to trade at every rebalance to get back to equal, which costs money.

    Where candidates lose it

    The common slip is dividing the old divisor by the split ratio and answering 1. The divisor is fixed by keeping the index level unchanged, and only one stock split, so the adjustment is smaller than the ratio.

    The second loss is saying a cap-weighted index needs the same adjustment. Market value does not change in a split, so a cap-weighted index does nothing; say that, then name the events that do change its divisor.

    What the interviewer asks next

    • Stock B now pays a special dividend of 20. How does each index type handle it?
    • Replace stock A with a new stock priced 200. What is the new divisor?
    • Which stock has the most influence on a price-weighted index, and why is that a flaw?

    Asked at Mizuho, Sales and Trading, Hong Kong, 2024 (Wall Street Oasis): Different index methodology - need to know all of them with examples.

← PreviousPage 6 of 8
  1. 1
  2. …
  3. 5
  4. 6
  5. 7
  6. 8
Next →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.