Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant puzzles, solved step by step

Puzzles
100
Traced to a firm
71
Topics
12
Hard
30
Topic
All topicsLogic and algorithmic reasoning10Conditional probability and Bayes7Counting and combinatorics8Continuous and geometric probability9Correlation, regression and linear algebra9Market making, betting and sizing9Expected value and optimal stopping9Statistics and estimation9Pricing, options and index maths7Games and strategic reasoning8Markov chains and random walks7Mental maths and number sense8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 11–20 of 100
  1. 011An urn starts with one red ball and one blue ball. Each turn you draw a ball at random, put it back, and add another ball of the same colour. After ten draws, what is the probability that exactly five of the draws were red?Markov chains and random walksCoreQuant researchQuant trading

    Try it first

    Your first instinct for P(exactly five reds)?

    Show the worked solution

    1/11, about 9.1%. Any particular sequence with five reds and five blues has probability 5! x 5! / 11!, whatever the order, because the numerators just count up the reds and blues separately. There are 10 choose 5 = 252 such sequences, and 252 x 5! x 5! / 11! = 1/11. The same algebra gives 1/11 for every count from 0 to 10.

    Why is this not ten coin flips?

    Think of a new café and its first customers. If the first few visitors like it and bring friends, it fills up; if they do not, it stays empty. Early luck compounds. In this urn each red draw adds a red ball, so it raises the chance of the next red: the draws reinforce each other, and runs that start lopsided tend to stay lopsided. That spreads the count of reds far more widely than independent coin flips, which bunch around five.

    Ten draws from the urn: every count of reds from 0 to 10 is equally likely012345678910coin flips: 24.6% at five1/11eachnumber of red draws out of tenStart: 1 red, 1 blueDraw, return, add one of that colour
    After ten draws from an urn that starts with one red and one blue ball and adds a ball of the drawn colour each time, every count of reds from 0 to 10 has probability 1/11, a flat line, while ten fair coin flips would pile up at five with 24.6%.

    Why does the order of the draws not matter?

    Write out one sequence, say five reds then five blues. The chances are 1/2, 2/3, 3/4, 4/5, 5/6 for the reds, then 1/7, 2/8, 3/9, 4/10, 5/11 for the blues. Shuffle the order and the denominators are still 2 up to 11, while the red numerators still run 1 to 5 and the blue ones 1 to 5, so every arrangement of five reds and five blues has the same probability, 5! x 5! / 11!. Draws whose joint probability ignores order are called exchangeableA sequence of random variables whose joint distribution does not change when you reorder them, even though they need not be independent., and that is the property doing the work.

    The relationship
    P(k reds in n)=(nk)k! (n−k)!(n+1)!=1n+1n=10:  252⋅120⋅12039 916 800=111P(k \text{ reds in } n) = \binom{n}{k}\frac{k!\,(n-k)!}{(n+1)!} = \frac{1}{n+1} \qquad n = 10:\; \frac{252 \cdot 120 \cdot 120}{39\,916\,800} = \frac{1}{11}
    \binom{n}{k}the number of orders in which k reds can appear
    k!(n-k)!/(n+1)!the probability of any one such order
    What it says in wordsThe number of orders times the probability of each order is always one over n plus one.

    Is there a picture that makes 1/11 obvious?

    Yes. This urn behaves exactly as if nature first picked a hidden red probability p uniformly between 0 and 1, and then flipped ten independent coins with that p. Averaging a binomial over a uniform p gives every count the same weight, 1/(n + 1). The urn is also a small model of momentum and of market share: an early lead makes further gains more likely, so the final split is highly uncertain, even though each step looks like a fair draw.

    Where candidates lose it

    The fast wrong answer is 252/1024, about 24.6%, from treating the draws as independent coin flips. The reinforcement is the whole question, and it spreads the outcomes out rather than pulling them to the middle.

    The second trap is trying to add up paths through a tree of ten levels. Spot that every order of a given mix has the same probability, and the problem collapses to one line.

    What the interviewer asks next

    • What is the probability the eleventh draw is red, given that five of the first ten were red?
    • What if the urn starts with two red and one blue ball?
    • As the number of draws grows, what does the fraction of red balls converge to?
  2. 012Speed round: convert 3/32, 7/16 and 11/64 to decimals in your head, and explain the pattern you used.Mental maths and number senseWarm upBelvedere TradingChicago · 2021

    Try it first

    What is 3/32 as a decimal?

    Show the worked solution

    3/32 = 0.09375, 7/16 = 0.4375 and 11/64 = 0.171875. Every denominator here is a power of two, so the unit fraction is a chain of halvings: 1/2 = 0.5, 1/4 = 0.25, 1/8 = 0.125, 1/16 = 0.0625, 1/32 = 0.03125, 1/64 = 0.015625. Find the rung, then multiply by the numerator. Each decimal ends exactly, because 2 divides a power of 10.

    Why do powers of two give clean decimals?

    Think of cutting a one-metre ribbon in half again and again: 50 cm, 25 cm, 12.5 cm, 6.25 cm. Each cut adds at most one digit to the length. A fraction terminates in decimal exactly when its denominator has no prime factors other than 2 and 5, so every power-of-two fraction ends, and 1 over 2 to the n has exactly n decimal places. That tells you before you start that 11/64 will have six digits after the point.

    Every power-of-two fraction is a chain of halvings1/20.51/40.251/80.1251/160.06251/320.031251/640.015625halve the one above, rung by rungScale the matching rung3/32= 3 x 0.031250.093757/16= 7 x 0.06250.437511/64= 11 x 0.0156250.171875
    Each rung of the halving ladder is half the one above, from 0.5 down to 0.015625 for 1/64, so 3/32 is three of the 0.03125 rung, 0.09375, and 11/64 is eleven of the 0.015625 rung, 0.171875.

    How do you do 11/64 without losing a digit?

    Split the numerator into pieces you already know. 11/64 is 8/64 + 2/64 + 1/64, which is 1/8 + 1/32 + 1/64: 0.125 + 0.03125 + 0.015625 = 0.171875. Or take 11 x 0.015625 as 10 x 0.015625 plus one more, 0.15625 + 0.015625. Either way you add numbers you have memorised instead of dividing. 7/16 works the same way as 1/2 - 1/16, 0.5 - 0.0625 = 0.4375.

    The relationship
    1164=18+132+164=0.125+0.03125+0.015625=0.171875\frac{11}{64} = \frac{1}{8} + \frac{1}{32} + \frac{1}{64} = 0.125 + 0.03125 + 0.015625 = 0.171875
    1/8, 1/32, 1/64rungs of the halving ladder
    11 = 8 + 2 + 1the numerator written in binary
    What it says in wordsWrite the numerator as a sum of powers of two and add the matching rungs.

    Why do trading firms test this?

    Some bond and futures markets have long quoted prices in 32nds and 64ths of a point, and option deltas and odds come up as fractions all day. A trader who converts 3/32 at the speed of reading reacts to a price while a slower colleague is still dividing. The same round usually mixes in products such as 38 x 42, which is 40 squared minus 2 squared, 1,596: the test is spotting structure that turns long arithmetic into one step.

    Where candidates lose it

    Candidates try long division under pressure and drop or add a zero: 0.9375 for 3/32 is a common slip, and it is actually 15/16. Knowing the ladder by heart removes the division entirely.

    The second loss is rounding. The question asks for the decimal, and 0.094 or 0.17 sounds careless when the exact answer is short and available. Give all the digits, then the rounded figure if asked.

    What the interviewer asks next

    • What is 13/128 as a decimal?
    • Now 38 x 42 in your head, and say the trick you used.
    • Convert 0.859375 back to a fraction.

    Asked at Belvedere Trading, Trading, Chicago, 2021 (Wall Street Oasis): 3/32 mental math, 38*42, crossing the bridge in the shortest amount of time

  3. 013Users join a server at times 1, 2, 4, 5 and 7 and leave at times 7, 3, 8, 9 and 10 respectively. A leave at the same moment as a join is processed first. What is the maximum number of users online at once, and how would you compute it efficiently for a million users?Logic and algorithmic reasoningCoreTwo SigmaNew York · 2025

    Try it first

    What is the peak number of users online together?

    Show the worked solution

    The peak is 3 users. Turn every join into a +1 event and every leave into a -1 event, sort all ten events by time with leaves before joins at equal times, and keep a running total. It goes 1, 2, 1, 2, 3, then at time 7 down to 2 and back to 3, then 2, 1, 0. Sorting costs n log n, and the sweep itself is linear.

    Why not check every moment in time?

    A shopkeeper who wants to know the busiest moment of the day does not count heads every second; they note each time the door opens in or out and keep a tally. The count of users can change only at a join or a leave, so the maximum must occur just after some join, and you only need to look at the 2n event times. Checking every time step costs time proportional to the length of the day, and comparing every pair of users costs n squared; both are far too slow at a million users.

    Sort the joins and leaves, then keep a running totaluser 1+-user 2+-user 3+-user 4+-user 5+-t = 7: user 1 leaves,user 5 joins1234join first: false peak of 4true peak 301234567891011timeusers
    Each join adds one user and each leave removes one, so a running total over the sorted events finds the peak of 3; processing the join at time 7 before the leave would produce a false peak of 4.

    Why does the tie rule matter so much?

    At time 7 user 1 leaves and user 5 joins. If a user's session is taken to end just before the moment they leave, then a leave and a join at the same instant never overlap, and the leave must be sorted first; sorting the other way invents a user who was never there. In code this is one comparison in the sort key, and it is exactly the detail interviewers use to separate a working answer from a nearly working one. Ask which convention applies before writing any code.

    The relationship
    peak=max⁡j  ∑i≤jdi,di∈{+1,−1}, events sorted by (ti, di)cost O(nlog⁡n)\text{peak} = \max_{j}\;\sum_{i \le j} d_i,\quad d_i \in \{+1,-1\},\ \text{events sorted by } (t_i,\ d_i) \qquad \text{cost } O(n \log n)
    d_i+1 for a join and -1 for a leave
    (t_i, d_i)the sort key: time first, then leaves (-1) before joins (+1)
    What it says in wordsSort the events so leaves come first at equal times, then the peak is the largest running total.

    Is there a version that avoids building the event list?

    Sort the join times and the leave times separately and walk two pointers through them. At each step take the earlier of the next join and the next leave, taking the leave on a tie, and adjust the count; this is the same sweep without allocating 2n tuples. If times are small integers you can go further: add +1 and -1 into an array indexed by time and take a running sum, which is linear. Mention both and say which you would use for a million users with timestamps in milliseconds.

    Where candidates lose it

    The trap is the tie. Many candidates write a correct sweep, sort by time alone and get 4, because the join at time 7 is processed before the leave. The question states the convention precisely to see whether you use it.

    The second loss is proposing a double loop that checks every pair of sessions. It gives the right answer on five users and fails the question, which asked how you would do it efficiently.

    What the interviewer asks next

    • Return the time interval during which the peak occurs, not just the count.
    • Users arrive as a stream and you must report the current count at any moment. What data structure do you use?
    • How many servers are needed if each can hold at most two users at once?

    Asked at Two Sigma, Equity Hedge, New York, 2025 (Wall Street Oasis): Given arrays (start & end) of the times users join and leave a server, find the max number of concurrent users on the server

  4. 014Five per cent of fund managers are skilled and beat the market in any given year with probability 60%; the rest are unskilled and beat it with probability 50%. Years are independent. A manager has beaten the market in exactly 8 of the last 10 years. What is the probability the manager is skilled?Conditional probability and BayesCoreCSCitadel SecuritiesMiami · 2022

    Try it first

    Roughly how likely is it that this manager is skilled?

    Show the worked solution

    About 12.7%. A skilled manager wins exactly 8 of 10 with probability 0.1209; an unskilled one with 0.0439, a likelihood ratio of about 2.75. Prior odds of skill are 5 to 95, 1 to 19. Posterior odds are 2.75 to 19, so the probability is 0.05 x 0.1209 / (0.05 x 0.1209 + 0.95 x 0.0439) = 0.127. The record helps, but luck has far more players.

    Why does an impressive record move the needle so little?

    Imagine a thousand people each tossing a coin ten times. About 55 of them will get eight heads or better with a fair coin. If a handful of the thousand had slightly biased coins, you still could not pick them out from the lucky crowd by one run of ten. Evidence moves a belief in proportion to how much more likely it is under one explanation than the other, and 8 wins in 10 is not much more likely from a 60% manager than from a 50% one. The ratio is about 2.75.

    An eight-win record is 2.75 times likelier from skill, but luck has 19 times the playersSkilled: wins 60% of years01234567812.1%910winning years out of 10Unskilled: wins 50% of years0123456784.4%910winning years out of 10Weighted by how common each is: 5% x 0.1209 against 95% x 0.043912.7%unskilled but lucky: 87.3%P(skilled | 8 wins in 10) = 12.7%
    Eight wins in ten years has probability 12.1% for a 60% manager and 4.4% for a 50% manager, but after weighting by how common each type is, 5% against 95%, the chance that an eight-win manager is skilled is only 12.7%.

    How do you set it up quickly?

    Use odds. Prior odds of skill are 1 to 19; the likelihood ratioHow many times more likely the evidence is under one hypothesis than under the other. of the record is (0.6/0.5) to the 8 times (0.4/0.5) squared, which is 1.2 to the 8 times 0.64, about 2.75; multiply to get posterior odds of about 0.145. Converting, 2.75 over 2.75 + 19 is 12.7%. The binomial coefficient, 45, is the same in both likelihoods and cancels, so you never need it.

    The relationship
    P(S∣8)P(U∣8)=0.050.95⋅0.68 0.420.510≈119×2.75  ⇒  P(S∣8)≈0.127\frac{P(S\mid 8)}{P(U\mid 8)} = \frac{0.05}{0.95}\cdot\frac{0.6^8\,0.4^2}{0.5^{10}} \approx \frac{1}{19}\times 2.75 \;\Rightarrow\; P(S\mid 8) \approx 0.127
    S, Uskilled and unskilled
    0.6^8 0.4^2the chance of one particular sequence of 8 wins and 2 losses for a skilled manager
    0.5^{10}the same for an unskilled manager
    What it says in wordsPrior odds of 1 to 19, times a likelihood ratio of 2.75, give a posterior of about 12.7%.

    What does this say about picking managers?

    When skill is rare and its edge is small, even a long, strong track record leaves luck as the likelier explanation. Using 8 or more wins instead of exactly 8 barely changes things: the answer becomes 13.9%. The honest limitation is that the model is stylised: real skill is not a fixed 60%, and survivorship means the managers you hear about were already filtered for good records, which pushes the true figure lower still.

    Where candidates lose it

    The common answer is around 80%, reading the record's win rate as the chance of skill. That skips the prior entirely, and with only 5% of managers skilled, the prior dominates.

    The quieter trap is computing the full binomial probabilities, 45 x 0.6 to the 8 x 0.4 squared and so on, and getting lost in decimals. The coefficient cancels. Say odds and likelihood ratio and the arithmetic stays on one line.

    What the interviewer asks next

    • How many years of 80% wins would you need before the manager is more likely skilled than not?
    • What if 20% of managers were skilled?
    • How does survivorship bias change the answer if you only ever see managers with good records?

    Asked at Citadel Securities, Sales and Trading, Miami, 2022 (Wall Street Oasis): I got a question about Bayes' theorem applied to a practical scenario, which I handled decently

  5. 015You roll three fair dice. Is a total of 9 or a total of 10 more likely, given that each can be written as exactly six unordered combinations of three faces?Counting and combinatoricsCoreQuant tradingProp trading firms

    Try it first

    Which total is more likely?

    Show the worked solution

    10 is more likely: 27 ways out of 216 against 25. The equally likely outcomes are the 216 ordered rolls, not the unordered combinations. A combination of three different faces covers six ordered rolls, one with a pair covers three, and a triple covers one. Total 9 includes 3 + 3 + 3, which counts once, and one fewer all-different combination, so it loses two orderings to 10.

    Why are combinations the wrong thing to count?

    Think of dealing two cards and asking whether a pair of kings or a king with a queen is more likely. There is one combination of each, but king and queen can arrive in either order while two kings are just two kings. Probability comes from counting outcomes that are equally likely, and with dice those are the ordered rolls: first die, second die, third die, 6 x 6 x 6 = 216 of them. Unordered combinations bundle different numbers of those outcomes, so counting combinations gives the wrong weights.

    Six combinations each, but the orderings differ: 25 against 27Total of 9orderings1 + 2 + 66 all different1 + 3 + 56 all different1 + 4 + 43 one pair2 + 2 + 53 one pair2 + 3 + 46 all different3 + 3 + 31 tripleTotal25/216 = 11.6%Total of 10orderings1 + 3 + 66 all different1 + 4 + 56 all different2 + 2 + 63 one pair2 + 3 + 56 all different2 + 4 + 43 one pair3 + 3 + 43 one pairTotal27/216 = 12.5%The triple 3 + 3 + 3 counts once; it is what costs 9 the two orderings
    Totals of 9 and 10 each have six unordered combinations, but weighting each combination by its number of orderings gives 27 of 216 rolls for 10 against 25 for 9, mostly because 9 includes the triple 3, 3, 3, which happens only one way.

    How do you count the orderings fast?

    Classify each combination by its repeats. Three different faces give 3! = 6 orders, a pair gives 3 orders, one for each position of the odd die, and a triple gives 1. For 9: 1 2 6, 1 3 5 and 2 3 4 are all different, 18; 1 4 4 and 2 2 5 are pairs, 6; 3 3 3 is a triple, 1; total 25. For 10: 1 3 6, 1 4 5 and 2 3 5 give 18; 2 2 6, 2 4 4 and 3 3 4 give 9; total 27. So 10 comes up 12.5% of the time and 9 only 11.6%.

    The relationship
    P(9)=3⋅6+2⋅3+1216=25216P(10)=3⋅6+3⋅3216=27216P(9) = \frac{3\cdot 6 + 2\cdot 3 + 1}{216} = \frac{25}{216} \qquad P(10) = \frac{3\cdot 6 + 3\cdot 3}{216} = \frac{27}{216}
    6, 3, 1the orderings of an all-different, a pair and a triple combination
    216the ordered outcomes of three dice
    What it says in wordsWeight each combination by its orderings and 10 beats 9 by two rolls in 216.

    Is there a shortcut that avoids listing?

    Yes: symmetry. Replacing each face x by 7 - x maps a total of t to 21 - t, so the distribution of three dice is symmetric about 10.5, and 10 and 11 are the two most likely totals, each 27/216. Anything further from 10.5, including 9, must be less likely or equal; a quick count confirms it is 25. Historically this is the question gamblers put to Galileo, who answered it by counting ordered outcomes, which is still the method.

    Where candidates lose it

    The trap is the question's own framing: six combinations each invites the answer that the totals are equally likely. The interviewer wants you to reject the framing, not accept it.

    The second loss is listing all 216 rolls, or writing out every ordering. Classifying by repeats, six, three or one, gets both totals in under a minute.

    What the interviewer asks next

    • What is the most likely total with four dice, and its probability?
    • What is P(total is 9) with two dice, and why does 9 behave differently?
    • How many ordered outcomes of three dice sum to 7?
  6. 016You need to sample a point uniformly at random from a triangle with vertices A, B and C, using two independent uniform numbers u and v on 0 to 1. How do you do it, and why does the formula A + u(B - A) + v(C - A) fail on its own?Continuous and geometric probabilityHardTwo SigmaNew York · 2023

    Try it first

    What goes wrong with A + u(B - A) + v(C - A) for u, v uniform on 0 to 1?

    Show the worked solution

    Draw u and v; if u + v is above 1, replace them with 1 - u and 1 - v; then return A + u(B - A) + v(C - A). The plain formula maps the unit square onto a parallelogram twice the size of the triangle, so half the draws, 50.0% in a simulation, land outside. Reflecting through the square's centre folds that half exactly onto the other, keeping the density flat and wasting no draws.

    Why does the plain formula give a parallelogram?

    Think of a tiled floor where each tile is a parallelogram and you want to pick a spot on one triangular half of a tile. Pick any spot on the tile and half the time you are on the wrong half. A + u(B - A) + v(C - A) with u and v each free on 0 to 1 walks up to one full step along AB and one full step along AC, which covers the parallelogram with corners A, B, C and D = B + C - A, not the triangle. The triangle is exactly the part where u + v is at most 1.

    Half the square lands outside the triangle: reflect it back inu + v > 1u + v below 1(0.8, 0.6)(0.2, 0.4)0011uvreflect: (u, v) to (1 - u, 1 - v)ABCD, never wantednaive point,outside ABCreflected, insideA + u(B - A) + v(C - A)
    Two uniforms fill a unit square that the linear map turns into a parallelogram twice the size of triangle ABC, so draws with u + v above 1 land outside; reflecting such a draw from (0.8, 0.6) to (0.2, 0.4) brings it back inside at a uniformly distributed spot.

    Why does reflecting keep the distribution uniform?

    Two facts. The map (u, v) to (1 - u, 1 - v) is a half turn about the square's centre, so it carries the upper triangle onto the lower one without stretching any area; and an affine mapA linear map followed by a shift, such as A + u(B - A) + v(C - A); it scales every area by the same factor. scales every area by the same factor, so a flat density stays flat. Put together, each small patch of the triangle receives draws from exactly two equal patches of the square. A simulation that splits the triangle into four equal pieces finds 24.9%, 25.1%, 25.0%, 25.0% of the points in them, each a quarter.

    The relationship
    (u,v)↦{(u,v)u+v≤1(1−u, 1−v)u+v>1P=A+u(B−A)+v(C−A)(u,v) \mapsto \begin{cases}(u,v) & u+v \le 1\\ (1-u,\,1-v) & u+v>1\end{cases} \qquad P = A + u(B-A) + v(C-A)
    u, vindependent uniforms on 0 to 1
    (1-u, 1-v)the reflection of a draw through the square's centre
    Pthe sampled point, uniform on triangle ABC
    What it says in wordsFold the unwanted half of the square onto the wanted half, then map it linearly onto the triangle.

    What other methods would an interviewer accept, and which fail?

    Rejection works: throw away draws with u + v above 1. It is correct but wastes half the random numbers. A popular wrong method draws three uniforms and divides each by their sum to get weights on A, B and C; the weights add to 1, but the points pile up near the centre, so the result is not uniform. A correct closed form uses a square root: with r1 and r2 uniform, take (1 - root r1)A + root r1 (1 - r2)B + root r1 r2 C. Name one fast method, prove it, then name the tempting wrong one.

    Where candidates lose it

    The trap is writing the linear formula and stopping, because it looks like a weighted average of the vertices. Half the points leave the triangle, and the candidate who does not draw the square never sees it.

    The second loss is fixing the problem in a way that breaks uniformity, such as normalising random weights to sum to 1 or clamping u + v to 1. Both keep points inside but crowd them into part of the triangle. Say why your fix preserves area.

    What the interviewer asks next

    • Prove the square-root method gives a uniform point.
    • How would you sample uniformly from a convex polygon with n vertices?
    • How would you sample uniformly from the surface of a sphere?

    Asked at Two Sigma, Research, New York, 2023 (Wall Street Oasis): Biased gamblers ruin problems; Markov Chain problems; sampling uniformly from triangle

  7. 017The true model is y = x1 + x2 + noise, where x1 and x2 are standardised and have correlation 0.5. You regress y on x1 alone, then regress the residuals on x2. What coefficient do you get on x2, and how would you recover the true value of 1 in two stages?Correlation, regression and linear algebraHardQuant researchQuant trading

    Try it first

    What coefficient does the second stage give on x2?

    Show the worked solution

    You get 0.75, not 1. Regressing y on x1 alone gives a slope of 1 + 0.5 = 1.5, because x1 soaks up the half of x2 that moves with it. The residual is x2 - 0.5x1 + noise, whose slope on x2 is 1 - 0.5 squared = 0.75. To recover 1, residualise x2 on x1 as well and regress the residual of y on the residual of x2: the Frisch-Waugh-Lovell theorem.

    Where does the missing quarter go?

    Picture two salespeople who often work the same client. If you credit all joint sales to the first before looking at the second, the second looks worse than they are, because some of their work was already booked to the first. Stage one regresses y on x1 alone, and since x2 is correlated with x1, the coefficient on x1 rises to 1.5: it takes credit for 0.5 of x2. That piece has been removed from the residual, so stage two can only find what is left of x2's effect.

    x1 has already eaten the part of x2 that points its way0.5 x1: already in x1x1x2part of x2 orthogonalto x1: length 0.87cos 0.5CoefficientsStage 1: y on x11.50True effect of x21.00Residual on raw x20.75Residual on x2 orthogonal1.0010
    With a correlation of 0.5, x2 splits into 0.5 x1 plus an orthogonal part; stage one assigns the 0.5 x1 piece to x1, so regressing the residual on raw x2 gives 0.75, while regressing it on the orthogonal part of x2 recovers the true 1.

    How do you get 0.75 exactly?

    Write the residual out. y - 1.5x1 = x2 - 0.5x1 + noise, and the slope of that on x2 is its covariance with x2 over the variance of x2: (1 - 0.5 x 0.5)/1 = 0.75. The formula generalises to 1 - rho squared times the true coefficient, so the bias gets worse as the regressors get more correlated: with rho = 0.9 you would find only 0.19. A simulation of 100,000 observations gives 1.506 for stage one and 0.752 for stage two.

    The relationship
    β^2,seq=Cov⁡(x2−ρx1, x2)Var⁡(x2)=1−ρ2=0.75β^2,FWL=Cov⁡(x2−ρx1, x2−ρx1)Var⁡(x2−ρx1)=1\hat\beta_{2,\text{seq}} = \frac{\operatorname{Cov}(x_2 - \rho x_1,\ x_2)}{\operatorname{Var}(x_2)} = 1 - \rho^2 = 0.75 \qquad \hat\beta_{2,\text{FWL}} = \frac{\operatorname{Cov}(x_2 - \rho x_1,\ x_2 - \rho x_1)}{\operatorname{Var}(x_2 - \rho x_1)} = 1
    \rhothe correlation between x1 and x2, 0.5
    x_2 - \rho x_1the part of x2 left after regressing it on x1
    What it says in wordsRegressing on raw x2 shrinks the answer by one minus rho squared; regressing on the part of x2 orthogonal to x1 gives the true coefficient.

    What does Frisch-Waugh-Lovell tell you to do?

    To get a variable's coefficient from a multiple regression in stages, partial the other regressors out of both y and that variable, then regress residual on residual. Here that means regressing x2 on x1 as well, keeping the orthogonal part x2 - 0.5x1, and regressing the stage-one residual on it. The slope comes back as exactly 1; the simulation gives 1.003. This is why factor-neutralising a signal before testing it, rather than after, matters in quant research: the order of the stages changes the answer.

    Where candidates lose it

    The common answer is 1, on the belief that regressing residuals step by step is the same as a multiple regression. It is only the same when the regressors are uncorrelated, and the question gives you a correlation of 0.5 precisely to break that.

    The second loss is saying the answer is biased without saying which way or by how much. Give 1.5 for stage one, 0.75 for stage two, the 1 - rho squared rule, and the fix.

    What the interviewer asks next

    • What would the stage-two coefficient be if the correlation were -0.5?
    • In the two-stage FWL regression, how do the standard errors compare with the full multiple regression?
    • You have a new signal correlated with a known factor. How do you test whether it adds anything?
  8. 018I will draw a card from a shuffled deck. You may pay Rs 6 to play a bet that pays Rs 10 if the card is red. Before deciding, you may pay to be told the card's colour. What is the most you should pay for that information?Market making, betting and sizingCoreOptiverChicago · 2025

    Try it first

    What is the information worth?

    Show the worked solution

    Rs 2. Without information the bet is worth 0.5 x 10 - 6 = -1, so you decline and your value is 0. With the colour known, you play on red and make 4, and skip black and make 0, which averages 2. Information is worth the improvement in your best decision: 2 - 0 = 2. If it would not change what you do, it is worth nothing.

    How do you value a piece of information?

    Suppose a weather forecast costs money and you are deciding whether to carry an umbrella. If you would carry it anyway, the forecast is worthless to you; it is valuable only if some answer would change what you do. The value of information is the expected value of your best decision with it, minus the expected value of your best decision without it. Work out both decision trees separately and subtract. Never value information by the size of the payout it relates to.

    Information is worth what it changes in your decisionWithout informationYou decidePlay: pay 6half chance of 10EV = 5 - 6 = -1DeclineEV = 0Best choice: decline. Value = 0With information firstColour?red, 1/2black, 1/2Red: playwin 10 - 6 = +4Black: decline0Value = 1/2 x 4 + 1/2 x 0 = 2Worth paying for the information: up to 2 - 0 = Rs 2
    Blind, the bet has an expected value of minus 1 so you decline and get 0; told the colour first, you play only on red and make 4 half the time, an average of 2, so the information is worth Rs 2.

    Why is it not worth Rs 4 or Rs 5?

    Rs 4 is what you make when the card is red, but it is red only half the time. Rs 5 is half the payout, which ignores the Rs 6 you pay to play. The information saves you from the losing half of the bet and lets you keep the winning half, and that is worth half of Rs 4, which is Rs 2. Pay more than Rs 2 and you would do better declining the offer of information and declining the bet.

    The relationship
    VOI=E[max⁡(payoff,0)]−max⁡(E[payoff],0)=12max⁡(4,0)+12max⁡(−6,0)−max⁡(−1,0)=2\text{VOI} = E\big[\max(\text{payoff}, 0)\big] - \max\big(E[\text{payoff}], 0\big) = \tfrac12\max(4,0) + \tfrac12\max(-6,0) - \max(-1,0) = 2
    payoff10 - 6 = 4 on red, -6 on black
    E[max(payoff, 0)]your value when you can choose after seeing the colour
    max(E[payoff], 0)your value when you must choose blind
    What it says in wordsInformation is worth the gap between deciding after you know and deciding before.

    When is information worth the most?

    Vary the price of the bet. At a price of 5 you are exactly indifferent blind, and the information is worth 2.50, its maximum; at a price of 0 you would always play, and it is worth 0. Information is valuable when you are close to indifferent and the decision could go either way. The formula also has the shape of an option payoff: knowing first lets you exercise only when it pays, which is why traders talk about paying for optionality and paying for information in the same breath.

    Where candidates lose it

    The trap is answering with the size of the win, Rs 4, or half the payout, Rs 5. Both value the information by the bet it is about, not by the decision it improves.

    The second loss is forgetting that without information you would decline. Candidates who compare with playing blind, at -1, get 3. The comparison is always with your best action without the information, which here is to walk away.

    What the interviewer asks next

    • What is the information worth if the bet costs Rs 3?
    • What if the information is only 80% reliable?
    • You can pay to see one card of a two-card hand before betting. How do you decide what that is worth?

    Asked at Optiver, Quantitative Research, Chicago, 2025 (Wall Street Oasis): Valuing information, taking directional bets when not plus EV.

  9. 019A casino flips a fair coin until the first head appears and pays 2 to the power n rupees if that happens on flip n. Its bank holds only 2 to the power 20 rupees, so it pays at most that. What is the fair price of the game?Expected value and optimal stoppingHardHRHudson River TradingNew York · 2020

    Try it first

    What is the fair price with the bank capped at 2 to the 20 rupees?

    Show the worked solution

    Rs 21. Flip n pays 2 to the n with probability 1/2 to the n, so each of flips 1 to 20 contributes exactly 1 rupee: 20 rupees. Beyond flip 20 the casino pays its whole bank, 2 to the 20, and the chance of getting that far is 1/2 to the 20, which adds 1 more. The famous infinite value collapses to 21 as soon as the payer's bank is finite.

    Why is the uncapped game worth infinity?

    Each extra flip halves the chance and doubles the prize, so each flip adds the same 1 rupee to the average, forever. An expected value is a sum over outcomes of prize times chance, and when every term is 1 and there are infinitely many terms, the sum has no limit. That is the St Petersburg paradox, discussed by Daniel Bernoulli in the eighteenth century: the mathematics says pay anything, and nobody would pay more than a modest sum.

    Each flip adds 1 rupee until the bank runs out; the whole tail adds 1 more15101520212223242526bank cap: 2 to the 20= Rs 10,48,576payout 2 to the n x chance 1/2 to the n = 1 rupee per flipcapped: 1/2, 1/4, ...adds up to 1flip20 flips x 1 rupee+ capped tailFair price = 20 + 1 = Rs 21
    Every flip up to the twentieth contributes exactly 1 rupee to the expected value, and the capped payouts beyond it add up to just 1 more, so a casino with a bank of 2 to the 20 rupees offers a game worth Rs 21.

    What exactly does the cap remove?

    Picture a lottery that promises to double your prize every day for ever, run by a shop with a small safe. The later promises are worth nothing because the shop cannot keep them. The cap turns every term after flip 20 from 1 rupee into 2 to the 20 divided by 2 to the n, which is 1/2, 1/4, 1/8 and so on, and those add up to exactly 1. So the tail that made the value infinite is now worth a single rupee. The bank is Rs 10,48,576 and the game is worth Rs 21.

    The relationship
    E=∑n=1202n⋅2−n+∑n=21∞220⋅2−n=20+1=21E = \sum_{n=1}^{20} 2^n \cdot 2^{-n} + \sum_{n=21}^{\infty} 2^{20} \cdot 2^{-n} = 20 + 1 = 21
    2^nthe payout if the first head arrives on flip n
    2^{-n}the chance the first head arrives on flip n
    2^{20}the bank, which caps every later payout
    What it says in wordsTwenty flips worth a rupee each, plus a capped tail worth one rupee.

    What does this teach about pricing a payoff?

    Doubling the bank adds only one rupee to the fair price: a bank of 2 to the 30, about Rs 107 crore, makes the game worth Rs 31. The value grows with the logarithm of what the counterparty can pay, so the realistic price of a lottery-like payoff depends on who stands behind it. The same idea appears in trading as counterparty risk: a contract's promised payout in extreme states is only worth what the other side can deliver in those states.

    Where candidates lose it

    Candidates recite the St Petersburg paradox and answer infinity, missing the cap in the question. The interviewer has changed one word to see whether you hear it.

    The second loss is getting 20 by stopping at the cap and forgetting the tail. Every sequence of 20 tails still pays the full bank, and that last piece is worth exactly one more rupee.

    What the interviewer asks next

    • How big must the bank be for the game to be worth Rs 50?
    • How much would a player with logarithmic utility pay for the uncapped game?
    • If you could play the capped game a million times, how would the average payout behave?

    Asked at Hudson River Trading, Prop Trading, New York, 2020 (Wall Street Oasis): Questions on EV for coin tosses, law of large numbers, Bayes theorem

  10. 020A broad stock index has returned an average of 11% a year over the past 30 years, with an annual standard deviation of 16% (illustrative figures). Assuming yearly returns are independent, give a 95% confidence interval for its true expected annual return.Statistics and estimationCoreOld Mission CapitalChicago · 2025

    Try it first

    Roughly how wide is the 95% interval?

    Show the worked solution

    About 5.3% to 16.7%. The standard error of a 30-year average is 16% divided by the square root of 30, 2.92 points. A 95% interval is 1.96 standard errors either side: 11% plus or minus 5.7. With a t value for 29 degrees of freedom it is a touch wider, plus or minus 6.0. Thirty years of data still leave the expected return very uncertain.

    Why is the band so wide after thirty years?

    Imagine judging a new cricketer's true batting average from a handful of innings. Scores swing wildly from one innings to the next, so a few innings tell you little. The precision of an average grows only with the square root of the number of observations, and yearly returns swing by 16 points, so 30 years give a standard error of about 2.9 points. That is large next to an average of 11%: the true figure could plausibly be 6% or 16%.

    Thirty years pins the average return only to within about 6 points11% sample average10 years1.1%20.9%+/- 9.9 points30 years5.3%16.7%+/- 5.7 points120 years8.1%13.9%+/- 2.9 points0%5%10%15%20%Half width = 1.96 x 16% / square root of years
    Around an 11% average with 16% annual volatility, the 95% interval is plus or minus 9.9 points with 10 years of data, 5.7 points with 30 years and still 2.9 points with 120 years, because precision grows only with the square root of the sample.

    How do you set it up in the room?

    State the assumption, then compute. Treat the 30 annual returns as independent draws with a standard deviation of 16%, so the sample average has a standard error of 16 over root 30. Root 30 is about 5.48, so the standard error is 2.92. Multiply by 1.96: 5.73. The interval is 5.3% to 16.7%. If the interviewer pushes, note that with only 30 observations a t distributionThe distribution used for an average when the standard deviation is estimated from the same small sample; it has fatter tails than the normal. with 29 degrees of freedom gives a multiplier of about 2.045, widening the band slightly.

    The relationship
    rˉ±1.96 σn=11%±1.96×16%30=11%±5.7%  ⇒  [5.3%, 16.7%]\bar r \pm 1.96\,\frac{\sigma}{\sqrt n} = 11\% \pm 1.96 \times \frac{16\%}{\sqrt{30}} = 11\% \pm 5.7\% \;\Rightarrow\; [5.3\%,\ 16.7\%]
    \bar rthe average annual return, 11%
    \sigmathe annual standard deviation, 16%
    nthe number of yearly observations, 30
    What it says in wordsThe interval is the average plus or minus 1.96 standard errors, and the standard error is volatility over the square root of years.

    What limitation should you add?

    Halving the band needs four times the data: 120 years still only pins the average to within about 2.9 points, and markets change over such spans. Returns are also not quite independent from year to year, fat tails make the normal multiplier optimistic, and the arithmetic average overstates the compound growth rate. The practical lesson is that expected returns are the hardest input in finance to estimate, far harder than volatility, which is why many quant processes lean on risk estimates and treat return forecasts with caution.

    Where candidates lose it

    The common mistake is using 16% as the width, the spread of a single year, rather than the standard error of the average. That gives an interval from -20% to 42%, which describes one year's return, not the long-run mean.

    The opposite slip is dividing by 30 instead of the square root of 30, which gives a band of about 1 point and badly overstates what the data can tell you.

    What the interviewer asks next

    • How many years of data would you need to pin the mean to within plus or minus 1 point?
    • How would monthly data change the interval for the mean?
    • Why is volatility much easier to estimate than the mean from the same data?

    Asked at Old Mission Capital, Prop Trading, Chicago, 2025 (Wall Street Oasis): Confidence interval on S&P 500 return past 30 years

← PreviousPage 2 of 10
  1. 1
  2. 2
  3. 3
  4. …
  5. 10
Next →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.