Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Hedge Funds puzzles, solved step by step

Puzzles
100
Traced to a firm
38
Topics
14
Hard
30
Topic
All topicsBetting and sizing5Conditional probability and Bayes7Continuous probability and distributions7Counting and combinatorics7Estimation and mental maths4Expected value and dice games8Logic and brainteasers10Market making and trading games6Options and payoffs5Portfolio and risk maths8Random walks and Markov chains7Returns, compounding and fees7Statistics and estimation11Valuation, accounting and macro riddles8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 1–10 of 14 · filtered from 100Clear filters
  1. 020A book holds Rs 60 crore of a stock with 30% volatility and Rs 40 crore of another with 20% volatility, and the two have a correlation of 0.5. What is the book's volatility in rupees, and what share of the risk comes from each position?Portfolio and risk mathsHardMan GroupBoston · 2022

    Try it first

    What share of the book's risk comes from the Rs 60 crore position?

    Show the worked solution

    The book's volatility is about Rs 23.1 crore a year, and the Rs 60 crore position carries about 74% of it on 60% of the capital. Stand-alone risks are Rs 18 crore and Rs 8 crore. Book variance is 18 squared plus 8 squared plus 2 x 0.5 x 18 x 8, which is 532, so volatility is Rs 23.07 crore. Each position's contribution is its covariance with the book over the book's volatility: Rs 17.17 crore and Rs 5.90 crore, which add back to the total.

    Why is risk not shared out like capital?

    Two friends share a taxi. One rides twice as far, straight through the traffic jam; the other gets off after a short hop. Splitting the fare by the number of bags each carries would be absurd. Risk belongs to a position in proportion to how much it moves and how much it moves with everything else, not to how much money sits in it. Here the first stock is larger, more volatile and positively correlated with the second, so it carries far more than its 60% of the capital.

    The bigger, more volatile name carries 74% of the risk on 60% of the capital18.0A alone60 x 30%+8.0B alone40 x 20%-2.93Diversifiedrho = 0.523.07BookRs crore60%40%Capital74.4%25.6%RiskPosition APosition B
    Stand-alone risks of Rs 18 crore and Rs 8 crore add to Rs 26 crore, diversification at a 0.5 correlation removes Rs 2.93 crore, and the book's Rs 23.07 crore of volatility splits 74.4% to the first position and 25.6% to the second, against a 60 to 40 split of capital.
    The relationship
    σbook2=a2+b2+2ρabRCA=a2+ρabσbookRCB=b2+ρabσbook\sigma_{book}^2 = a^2 + b^2 + 2\rho ab \qquad RC_A = \frac{a^2 + \rho ab}{\sigma_{book}} \qquad RC_B = \frac{b^2 + \rho ab}{\sigma_{book}}
    a, bstand-alone rupee volatilities: 60 x 30% = 18 and 40 x 20% = 8
    \rhothe correlation, 0.5
    RCa position's contribution to book volatility
    What it says in wordsEach position owns its own variance plus half the shared term, and dividing by the book's volatility turns that into rupees of risk.
    PositionCapital, Rs croreVolatilityStand-alone riskRisk contributionShare of risk
    A6030%18.017.1774.4%
    B4020%8.05.9025.6%
    Book10026.023.07100.0%
    Rs crore of annual volatility. The stand-alone risks add to Rs 26.0 crore, but the book's volatility is Rs 23.07 crore, of which position A contributes 74.4% and position B 25.6%.

    Why do the contributions add up exactly to the total?

    Split the variance. The cross term, 2 x 0.5 x 18 x 8 = 144, is shared equally, 72 to each position. So position A owns 324 + 72 = 396 of the 532 of variance and position B owns 64 + 72 = 136, and dividing each by the book's volatility of 23.07 gives rupee contributions that sum exactly to Rs 23.07 crore. The diversification benefit is the gap between the stand-alone total of Rs 26 crore and the book's Rs 23.07 crore.

    What does a risk manager do with the split?

    Cut where the risk is, not where the money is. Each rupee in position A carries 28.6 paise of marginal risk against 14.7 paise in position B, so trimming Rs 10 crore from A lowers book volatility by roughly Rs 2.9 crore. Recomputing exactly gives Rs 2.84 crore, close to the estimate. The limitation: the split is a snapshot at one correlation, and when correlations move, both the total and the split move with them.

    Where candidates lose it

    The quick answer shares risk like capital, 60 and 40, or like stand-alone risk, 18 and 8. The first ignores volatility and the second ignores correlation; neither sums to the book's actual Rs 23 crore of risk.

    The second slip is adding the stand-alone risks to get the book's risk, Rs 26 crore. Volatilities do not add unless the correlation is exactly 1; variances do, with the cross term included.

    What the interviewer asks next

    • If the correlation fell to zero, how would the risk split between the two positions?
    • How much of position B would you add to minimise the book's volatility, holding A fixed?
    • How do transaction costs change which position you trim first?

    Asked at Man Group, Investment Management, Boston, 2022 (Wall Street Oasis): How do you understand portfolio risk and transaction cost?

  2. 025Four observations, 3.1, 7.4, 5.2 and 9.0, come from a uniform distribution on 0 to theta. What is the maximum likelihood estimate of theta, is it biased, and how would you correct it?Statistics and estimationHardACAQR Capital ManagementTown of Greenwich · 2022

    Try it first

    Which statement is right?

    Show the worked solution

    The MLE is 9.0, the largest observation; it is biased low, and the unbiased correction is 5/4 x 9.0 = 11.25. The likelihood is 1 over theta to the fourth for any theta of at least 9.0 and zero below it, so it peaks at the sample maximum. But the maximum of n draws averages n/(n + 1) of theta, here 4/5, so scaling by (n + 1)/n removes the bias.

    Why is the MLE the largest observation?

    A friend draws four raffle tickets numbered from 1 up to some unknown top number, and the highest you see is 90. The top number is at least 90; guessing higher only spreads your belief over tickets nobody drew. The likelihood, 1 over theta to the n, is zero for any theta below the largest observation and falls as theta rises above it, so it is maximised exactly at the sample maximum, 9.0. This is a case where you do not differentiate: the maximum sits on a boundary, not where a slope is zero.

    The largest draw always sits below theta; scaling by 5/4 corrects it0246810123.17.45.29.0The data, and three estimates of thetaMLE 9.0corrected 11.252 x mean 12.35On average, four draws cut 0 to theta into five equal gapsgapgapgapgapgap04/5 theta = 9.0theta = 11.25So E[max] = 4/5 theta, and theta = 5/4 x 9.0 = 11.25
    Four uniform draws on 0 to theta cut it into five gaps of equal expected size, so the largest draw averages four fifths of theta; the MLE of 9.0 therefore sits below theta, and scaling by 5/4 gives the unbiased 11.25, against 12.35 from doubling the sample mean.
    The relationship
    L(θ)=∏i=141θ=θ−4    (θ≥9.0)E[max⁡]=nn+1θ  ⇒  θ^=n+1nmax⁡=54×9.0=11.25L(\theta) = \prod_{i=1}^{4}\frac{1}{\theta} = \theta^{-4} \;\;(\theta \ge 9.0) \qquad E[\max] = \frac{n}{n+1}\theta \;\Rightarrow\; \hat\theta = \frac{n+1}{n}\max = \frac{5}{4}\times 9.0 = 11.25
    L(\theta)the likelihood of the four observations
    nthe number of observations, 4
    \maxthe largest observation, 9.0
    What it says in wordsThe likelihood peaks at the largest observation, which on average falls short of theta by a factor n/(n + 1).

    Why is it biased, and by how much?

    The largest draw can never be above theta and is almost always below it. Four points dropped at random on 0 to theta cut it into five gaps of equal expected length, so the largest point sits on average four fifths of the way up, and the MLE underestimates theta by a fifth on average. Multiplying by 5/4 fixes it: 9.0 becomes 11.25. The bias shrinks as n grows, since n/(n + 1) tends to 1, but with four points it is large.

    How does it compare with the obvious alternative?

    The method of moments doubles the sample mean, since a uniform on 0 to theta averages theta/2: the mean here is 6.175, giving 12.35. Both 11.25 and 12.35 are unbiased, but the corrected maximum has a much smaller variance, theta squared over n(n + 2) against theta squared over 3n, because the largest draw carries the most information about the top of the range. With four points that is theta squared over 24 against theta squared over 12: half the variance.

    Where candidates lose it

    Candidates differentiate the log-likelihood, get minus n over theta, set it to zero and find no solution. The likelihood only falls on the allowed range, so the maximum sits at the boundary, the largest observation; say that before reaching for calculus.

    The second miss is calling the MLE unbiased because maximum likelihood estimates are often well behaved. Here it is biased low by construction, and the interviewer expects the (n + 1)/n correction.

    What the interviewer asks next

    • What is the variance of the corrected estimator with four observations?
    • What is the MLE if the distribution is uniform on theta to 2 theta?
    • Derive the expected value of the maximum of n uniform draws.

    Asked at AQR Capital Management, Trading, Town of Greenwich, 2022 (Wall Street Oasis): Derive the mle for some given distribution. Explain linear regression intuitively and derive the ols estimate.

  3. 040A bag holds four stones, each black or white. Before you look, every count of black stones from 0 to 4 is equally likely. You draw two stones without replacement and both are black. What is the chance the next stone is black, and at what price would you bet on it?Conditional probability and BayesHardCitadelNew York · 2025

    Try it first

    Chance the third stone is black:

    Show the worked solution

    3/4, so fair odds are 3 to 1 on black. Two black draws rule out bags with 0 or 1 black. The ways to draw two blacks in order are 2, 6 and 12 for bags with 2, 3 and 4 black, so those bags now carry 10%, 30% and 60%. The next stone is black with chance 0, 1/2 and 1 in them, which averages to 3/4. A contract paying 100 if black is worth 75.

    How do the two black draws change your view of the bag?

    If a friend pulls two red sweets from a jar you have never seen, you start to suspect it is mostly red. Each possible bag is reweighted by how likely it was to produce what you saw: equal priors times the chance of two blacks. Drawing two blacks in order has 0 ways from bags with 0 or 1 black, 2 x 1 = 2 ways with 2 black, 3 x 2 = 6 with 3 black and 4 x 3 = 12 with 4 black. Out of 20 in total, that is 10%, 30% and 60%: the all-black bag is now the favourite.

    Two blacks drawn: the evidence shifts weight toward the all-black bag0%0 black0%1 black10%2 blackweight 2next black 030%3 blackweight 6next black 1/260%4 blackweight 12next black 1prior 20% eachbefore the drawafter two blacksChance the next is black10% x 0+ 30% x 1/2+ 60% x 13/4Fair odds: 3 to 1 onFair price: 75 per 100
    Before the draw each bag is 20% likely; after two black stones the bags with 2, 3 and 4 black carry 10%, 30% and 60%, and since the next stone is black with probability 0, one half and 1 in those bags, the chance it is black is 3/4.

    How do you turn the posterior into a price?

    Average the chance of black over the bags you still believe in. In the 2-black bag both remaining stones are white; in the 3-black bag one of two is black; in the 4-black bag both are, so the answer is 0.1 x 0 + 0.3 x 1/2 + 0.6 x 1 = 3/4. A contract paying 100 if the next stone is black is worth 75. You would buy it below 75 and sell it above; as a market maker you might quote 70 at 80. Offered even money on black, your expected profit per rupee staked is 0.75 minus 0.25, which is 50 paise.

    The relationship
    P(B3∣B1B2)=∑kP(k∣B1B2) P(B3∣k,B1B2)=220⋅0+620⋅12+1220⋅1=34P(B_3 \mid B_1 B_2) = \sum_{k} P(k \mid B_1 B_2)\,P(B_3 \mid k, B_1 B_2) = \tfrac{2}{20}\cdot 0 + \tfrac{6}{20}\cdot\tfrac{1}{2} + \tfrac{12}{20}\cdot 1 = \tfrac{3}{4}
    kthe number of black stones in the bag
    B1 B2the event that the first two draws are black
    2, 6, 12the ordered ways to draw two blacks from bags with 2, 3 and 4 black
    What it says in wordsThe chance of another black is the average of each bag's chance, weighted by how much the evidence now favours that bag.

    Check it with Laplace rule of successionWith a uniform prior, after s successes in n trials, the chance the next trial succeeds is (s + 1) / (n + 2).: after 2 blacks in 2 draws the next is black with chance (2 + 1)/(2 + 2) = 3/4, and the rule holds exactly for this finite bag. Two routes to 3/4 is what separates a solid answer from a lucky one. The limitation: everything rests on the flat prior. If you had reason to think mixed bags were more common, 3/4 would fall.

    Where candidates lose it

    The common loss is saying 1/2 because the remaining stones are unknown, which throws away the information in the two draws. The question is about updating, and the interviewer wants to see the reweighting.

    The second loss is weighting the surviving bags equally, a third each, which gives 1/2. Each bag must be weighted by how likely it made two blacks: 2, 6 and 12. Then price it: a probability without a bet is half the answer at a trading firm.

    What the interviewer asks next

    • The third stone is black too. What is the chance the fourth is black?
    • You quote 70 at 80 on a contract paying 100 if black and someone who has seen the bag lifts your offer. What now?
    • How does the answer change if the prior is that each stone is black with probability one half, independently?

    Asked at Citadel, Quantitative Trading, New York, 2025 (Wall Street Oasis): Extended bayes derivative question about four stones in a bag (black and white stones).

  4. 041Two orders arrive one after the other. The first arrives after a wait that is exponential with a mean of one minute; the second arrives after a further, independent exponential wait with the same mean. What is the probability that both have arrived within one minute?Continuous probability and distributionsHardCitadelChicago · 2025

    Try it first

    Your estimate:

    Show the worked solution

    1 minus 2/e, about 26.4%. The total wait is the sum of two independent exponential waits. Convolving the two densities gives t e^-t, a gamma shape that starts at zero because two steps cannot both be instant. Its area from 0 to 1 is 1 minus e^-1 (1 + 1), which is 1 minus 2/e. A second route: it is the chance that a Poisson process with rate 1 produces at least two arrivals in one minute.

    Why is this not the chance of one wait, squared?

    Squaring would be right if both orders were racing from the same start line. Here they queue: the second clock only starts when the first order lands. It is like two buses where you must take the first to reach the stop for the second. The event is that the sum of the two waits is under one minute, which is stricter than each wait being under one minute. The chance one wait is under a minute is 1 minus 1/e, about 63%; squaring gives 40.0%, which answers a different question.

    How do you get the density of the sum?

    Add up every way to split the total t between the two waits. The density of a sum of independent waits is the convolutionThe density of a sum of two independent variables, found by integrating one density against the other shifted over every possible split of the total. of their densities, and for two exponentials it is t e^-t. Every split of t into s and t minus s has density e^-s times e^-(t minus s), which is e^-t whatever s is, and there is a length t of possible splits. Integrate t e^-t from 0 to 1 by parts and you get 1 minus 2/e.

    Total wait of two exponential steps: the shaded area under 1 minute is 26.4%one exponential wait, e^-tsum of two: t e^-t, peak at 1 minute26.4%012345Total wait, minutes0.00.51.0Area under 1 minute1 - e^-1 (1 + 1)= 1 - 2/e26.4%Not the same as eachwait under 1 minute:(1 - 1/e)^2 = 40.0%
    The total of two independent one-minute exponential waits has density t e^-t, which starts at zero and peaks at one minute, so only 26.4% of its area, 1 minus 2/e, lies below one minute.
    The relationship
    P(X1+X2≤1)=∫01te−t dt=1−2e−1≈0.264P(X_1 + X_2 \le 1) = \int_0^1 t e^{-t}\,dt = 1 - 2e^{-1} \approx 0.264
    X1, X2the two independent exponential waits, each with mean one minute
    t e^-tthe density of their sum, from convolving the two exponential densities
    What it says in wordsThe chance the total wait is under a minute is the area under the gamma density up to one minute.

    Check it with counting. Exponential waits are the gaps of a Poisson process, so both orders arrive within a minute exactly when the process makes at least two arrivals in that minute. With one arrival expected per minute, the chance of zero is e^-1 and of exactly one is e^-1, so at least two is 1 minus 2/e, the same number. Say both routes and the interviewer will usually skip ahead.

    Where candidates lose it

    The common loss is squaring the single-wait probability, which answers the question of two independent orders racing in parallel. Read the setup again: one after the other means the waits add.

    The second loss is freezing on the convolution integral. If the integral will not come, switch to the Poisson count: at least two arrivals in one minute. Candidates who know one route and not the other are the ones interviewers push hardest.

    What the interviewer asks next

    • What is the probability that three orders in sequence all arrive within two minutes?
    • Given that both orders arrived within one minute, what is the expected arrival time of the first?
    • The two waits have means of one and two minutes. What is the density of their sum?

    Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis): if i knew this was about convolutions, i would have answered better.

  5. 042A bowl holds 100 noodles. You repeatedly pick two free ends at random and tie them together, until no free ends remain. What is the expected number of loops?Counting and combinatoricsHardDED.E. ShawNew York · 2026

    Try it first

    Roughly how many loops?

    Show the worked solution

    About 3.28 loops. With k strands in the bowl there are 2k free ends. Pick any end; the other end you pick is one of the remaining 2k minus 1, and exactly one of those belongs to the same strand, so this tie closes a loop with chance 1/(2k minus 1). Either way the number of strands falls by one. Adding 1/199 + 1/197 + ... + 1/3 + 1 gives about 3.284.

    What does one tie do, whatever happens?

    Start with the bookkeeping, because it makes the rest easy. Every tie reduces the number of loose strands by exactly one: either it closes a strand into a loop, or it joins two strands into one longer strand. So there are always exactly 100 ties, and with k strands left there are 2k free ends. The question becomes how many of those 100 ties happen to close a loop.

    What is the chance a given tie closes a loop?

    Think of a room of dancers holding hands in lines: grab one free hand, then pick a second free hand at random, and a circle forms only if the second hand is at the other end of the same line. With 2k free ends, the second end is one of 2k minus 1, and exactly one of them is the other end of the strand you picked, so the chance is 1/(2k minus 1). Give each tie an indicator that is 1 if it closes a loop; by linearity of expectation the expected number of loops is the sum of the chances, from 1/199 for the first tie up to 1 for the last.

    Each tie's chance of closing a loop: tiny for 90 ties, large only at the end00.51first tie: 1/199, 0.5%1/51/3last tie: 1Tie number 1 to 100 (chance this tie closes a loop)0123Running total of expected loopsafter 90 ties: 1.15all 100 ties3.28
    The first tie closes a loop with chance 1 in 199 and the chances stay tiny until the last few ties, 1/5, 1/3 and 1, so the expected number of loops from 100 noodles is only 3.28, about a third of it from the last three ties.
    The relationship
    E[loops]=∑k=110012k−1=1+13+15+⋯+1199≈3.28E[\text{loops}] = \sum_{k=1}^{100} \frac{1}{2k-1} = 1 + \tfrac13 + \tfrac15 + \dots + \tfrac{1}{199} \approx 3.28
    kthe number of strands in the bowl before a tie
    1/(2k-1)the chance that tie joins the two ends of one strand
    What it says in wordsAdd each tie's chance of closing a loop to get the expected number of loops.

    For a sense check without a calculator: the sum of odd reciprocals up to 1/(2n minus 1) is about half of ln n plus ln 2 plus half of Euler's constant, which for n = 100 gives 3.28. The number of loops grows only like the logarithm of the number of noodles: a million noodles would give only about 7.9 loops.

    Where candidates lose it

    The first loss is trying to track the lengths of the strands, which quickly becomes impossible. The length of a strand never matters; only the count of strands does.

    The second loss is getting 1/(2k minus 1) right but summing it wrong, for example as 100 x 1/199. Write the sum out from the last tie backwards, 1 + 1/3 + 1/5, and the size of the answer becomes obvious.

    What the interviewer asks next

    • What is the variance of the number of loops?
    • What is the probability that you end with exactly one big loop?
    • How does the answer grow with the number of noodles, roughly?

    Asked at D.E. Shaw, Research, New York, 2026 (Wall Street Oasis): What is the expected number of loops from tying 100 noodles' ends together randomly

  6. 043You may roll a fair die up to three times. After each roll you either stop and are paid the face in rupees, or throw that roll away and roll again; if you reach the third roll you must take it. What is the best stopping rule, and what is the game worth?Expected value and dice gamesHardSCSquarepoint CapitalLondon · 2026

    Try it first

    On the first roll you get a 4. What do you do?

    Show the worked solution

    Stop on the first roll only with a 5 or 6, on the second with a 4 or more, and the game is worth 14/3, about Rs 4.67. Work backwards. The last roll is worth 3.5. With two rolls left, keep 4, 5 or 6 and reroll otherwise: worth (4 + 5 + 6)/6 + 3/6 x 3.5 = 4.25. With three rolls left, keep only what beats 4.25, a 5 or 6: worth (5 + 6)/6 + 4/6 x 4.25 = 14/3.

    Why start from the last roll?

    Deciding whether to take a job offer is easier if you know what your fallback is worth. Each keep-or-reroll decision compares the roll in hand with the value of the rolls still to come, so you need the value of the future first, and the only stage with no future is the last one. On the last roll you must take whatever comes, which is worth 3.5 on average. That single number lets you solve the stage before it, and so on backwards. This is {term('backward induction', 'Solving a sequence of decisions from the last one to the first, using the value of each later stage to make the earlier decision.')}, the core of dynamic programming.

    Solve from the last roll backwards: each value sets the next thresholdFirst roll3 rolls left123456keep 5 or 6reroll the restWorth4.67= 14/3Second roll2 rolls left123456keep 4, 5 or 6reroll the restWorth4.25= 17/4Last roll1 roll left123456must keep itWorth3.50= 7/2Keep a roll only if it beats what the remaining rolls are worth: 3.5, then 4.25Arrows run right to left: each stage uses the value of the stage after it
    With one roll left the game is worth 3.5, so with two left you keep 4 or more and the game is worth 4.25; with three left you keep only 5 or 6, and the whole game is worth 14/3, about 4.67.

    How do the thresholds come out?

    With two rolls left, a roll of 4, 5 or 6 beats the 3.5 you expect from rerolling, and 1, 2 or 3 does not. So the two-roll game is worth the average of the kept faces times their chance, plus the chance of rerolling times 3.5: 15/6 + 1.75 = 4.25. On the first roll the fallback is now 4.25, so a 4 is no longer good enough: only 5 or 6 is kept. That gives 11/6 plus 4/6 x 4.25, which is 1.833 plus 2.833, or 14/3.

    The relationship
    Vn=16∑f=16max⁡(f, Vn−1),V1=3.5,  V2=4.25,  V3=143V_n = \frac{1}{6}\sum_{f=1}^{6} \max\left(f,\, V_{n-1}\right), \quad V_1 = 3.5,\; V_2 = 4.25,\; V_3 = \tfrac{14}{3}
    V_nthe value of the game with n rolls left
    fthe face you just rolled
    max(f, V_{n-1})keep the roll or throw it away, whichever is worth more
    What it says in wordsEach stage is worth the average, over the six faces, of the better of keeping the face or playing on.

    Add the pattern. More rolls always raise the value, but by less each time: 3.5, 4.25, 4.67, then about 4.94 with four rolls. An extra option is always worth something and never worth more than what it can still improve. That is the same logic as valuing a trade you can exit early: the right to wait is priced by what the future is worth, not by the average outcome.

    Where candidates lose it

    The common loss is using 3.5 as the threshold at every stage, which keeps a 4 on the first roll. The fallback on the first roll is the two-roll game, 4.25, not a single roll.

    The second loss is solving forwards and getting lost. Say you will start from the last roll, compute 3.5, 4.25 and 14/3 in that order, and the thresholds fall out.

    What the interviewer asks next

    • What is the game worth with four rolls, and what is the first-roll threshold?
    • You must pay Rs 1 for every reroll. How do the thresholds change?
    • You are paid the square of the face instead. What is the optimal rule?

    Asked at Squarepoint Capital, Quantitative Research, London, 2026 (Wall Street Oasis): Dynamic programming questions with focus on probability at the end.

  7. 045You make a market on the number of heads in 10 fair coin flips: 4.5 bid, 5.5 offered. A counterparty who has already seen the first three flips lifts your offer. What does the trade tell you, and where do you requote?Market making and trading gamesHardCitadelNew York · 2025

    Try it first

    Given that they bought at 5.5, the fair value is about

    Show the worked solution

    The lift says they saw at least two heads, so the fair value is now at least 5.75, not 5; requote around 5.75 bid, 6.5 offered. The other seven flips are worth 3.5 heads, so the buyer's value is heads seen plus 3.5. Paying 5.5 only makes sense with two heads (5.5) or three (6.5). Those are 3 to 1 likely, giving 5.75; a buyer who needs a strict edge saw three heads, worth 6.5.

    What is the trader's view before they trade?

    A friend offers to buy your raffle ticket after the first few numbers are drawn. The offer itself is the warning. The informed trader values the contract at heads already seen plus 3.5, the expected heads in the seven unseen flips, so their value is 3.5, 4.5, 5.5 or 6.5 with chances 1, 3, 3 and 1 in 8. Against your market of 4.5 bid and 5.5 offered, they buy only if their value is at least 5.5, and sell to you at 4.5 only if it is 4.5 or less. The flat 5 you quoted around is right only for someone who has seen nothing.

    What the buyer saw decides whether they lift: a lift means 2 or 3 heads34567your offer 5.5your bid 4.5value given a lift = 5.750 headsvalue 3.5chance 1/81 headvalue 4.5chance 3/82 headsvalue 5.5chance 3/8may lift3 headsvalue 6.5chance 1/8liftshits bidTrader's fair value after seeing the first 3 flips
    The informed trader's value is 3.5, 4.5, 5.5 or 6.5 depending on how many heads they saw, so a lift at 5.5 means two or three heads, and weighting those 3 to 1 puts the fair value given the trade at 5.75, above your 5.5 offer.

    How do you turn the trade into a new fair value?

    Condition on the fact that they traded. Only the two-head and three-head worlds produce a buy at 5.5, and they are 3/8 and 1/8 likely, so given a lift the value is (3 x 5.5 + 1 x 6.5) / 4 = 5.75. If you assume they would not bother trading at zero edge, only the three-head world is left and the value is 6.50. Either way you sold too cheaply: this is adverse selectionThe tendency of a market maker to trade most with the people who know more, so the trades that happen are the ones that lose money for the market maker., and it is the cost every market maker prices into the spread.

    Now requote. Your bid should not be below what you now believe the floor is, and your offer should sit where even the best informed buyer has no edge. Something like 5.75 bid, 6.5 offered does both: a buyer who saw three heads is indifferent at 6.5, and you are no longer selling below value. Cut your size too, because you know someone is trading with more information than you, and say you would ask whether they could see the flips before quoting again.

    Where candidates lose it

    The common loss is staying at 5 because the coin is fair. The coin is fair; the counterparty is not uninformed. The trade itself carries information and you must update on it.

    The other loss is overreacting and moving to 8 or 9, as if the trader knew all ten flips. They saw three. Condition on what could have made them trade, weight those worlds, and move by exactly that much.

    What the interviewer asks next

    • The same trader then hits your new bid. What do you conclude?
    • How wide should your first market have been if you knew one counterparty could see three flips?
    • What if the trader had seen the first three flips but traded a small size and then a large size?

    Asked at Citadel, Quantitative Trading, New York, 2025 (Wall Street Oasis): Superday was more market-making but requires very sold foundation in math and statistics.

  8. 066You roll a fair die repeatedly until you have seen every even number, 2, 4 and 6. Given that the last new even number to appear is a 2, what is the probability that your first roll was a 1? Why is the intuitive answer of 1/5 wrong?Conditional probability and BayesHardSCSquarepoint CapitalLondon · 2026

    Try it first

    Given the game ends on a 2, what is the chance the first roll was a 1?

    Show the worked solution

    The probability is 1/6, not 1/5. Ending on a 2 rules out a first roll of 2, but it does not leave the other five faces equally likely. A first roll of 4 or 6 has already cleared one rival, so 2 then finishes last half the time; after an odd first roll it finishes last only a third of the time. Weighting the faces by those chances gives 1/6 for a 1.

    Why does ruling out one face not spread its weight evenly?

    Suppose you hear that a friend arrived late to a meeting. Before, the bus, the train and the car were equally likely ways she travelled. Learning how things ended shifts weight towards the starts that make that ending more likely, in proportion to how strongly each one leads to it. If the bus is late twice as often as the train, the bus now carries twice the train's weight. The 1/5 answer treats the ending as if it only ruled a face out and said nothing else.

    How likely is a 2 to finish last after each first roll?

    Odd rolls never change the order in which the even numbers first appear, so after an odd first roll the three evens are still symmetric and 2 is last with chance 1/3; after a first roll of 4 or 6 only two evens remain and 2 is last with chance 1/2; after a first roll of 2 it can never be the last new even. Multiply each by the 1/6 chance of that first roll: the joint chances are 1/18 for each odd face, 0 for a 2 and 1/12 for each of 4 and 6. They add to 1/3.

    Knowing how it ends reweights how it beganFirst rolleach face 1/6First roll 1, 3 or 5chance 1/22 last: 1/3joint 1/6First roll 2chance 1/62 last: 0joint 0First roll 4 or 6chance 1/32 last: 1/2joint 1/6Total chance 2 ends it: 1/6 + 0 + 1/6 = 1/3Given the game ends on a 21/61021/631/441/651/46naive 1/5 each (dashed) against the truechances: 1/6 for each odd face
    An odd first roll leaves 2 last among the evens a third of the time, a 4 or 6 leaves it last half the time and a 2 never does, so given the game ends on a 2 each odd face has chance 1/6 and each of 4 and 6 has chance 1/4, not 1/5 each.
    The relationship
    P(first=1∣2 last)=P(first=1) P(2 last∣first=1)P(2 last)=16⋅1313=16P(\text{first}=1 \mid 2 \text{ last}) = \frac{P(\text{first}=1)\,P(2 \text{ last} \mid \text{first}=1)}{P(2 \text{ last})} = \frac{\tfrac16 \cdot \tfrac13}{\tfrac13} = \frac16
    P(first = 1)the chance of rolling a 1 first, 1/6
    P(2 last | first = 1)the chance 2 is the last even to appear after an odd first roll, 1/3
    P(2 last)the overall chance the game ends on a 2, 1/3 by symmetry
    What it says in wordsBayes' rule: the prior chance of a 1, times how strongly a 1 leads to ending on a 2, divided by the overall chance of ending on a 2.

    How do you check the answer?

    Make the six posterior chances add up. Three odd faces at 1/6 each and two even faces, 4 and 6, at 1/4 each give 1/2 plus 1/2, which is 1, with nothing left for a 2. The odd faces keep exactly their starting weight because an odd roll tells you nothing about the evens; all the weight removed from the 2 goes to 4 and 6. Saying that sentence shows the interviewer you understand Bayes ruleThe rule for updating a probability after new information: the prior times the likelihood of the information, divided by the overall chance of the information. rather than recite it.

    Where candidates lose it

    The trap is to condition only by elimination: the game ends on a 2, so the first roll was not a 2, so the five other faces share the weight at 1/5 each. That treats the ending as a filter when it is also evidence about the start.

    The second loss is getting 1/6 and being unable to say where the missing weight went. Name it: 4 and 6 each rise to 1/4, because they make ending on a 2 more likely.

    What the interviewer asks next

    • Given the game ends on a 2, what is the chance the first roll was a 4?
    • What is the expected number of rolls in this game?
    • Given the game ends on a 2, what is the chance the second roll was a 1?

    Asked at Squarepoint Capital, Quant Research Intern Interview, London, 2026 (Wall Street Oasis): why is the probability of seeing a 1 on our first roll, given that we end on a 2, not 1/5

  9. 067Daily returns are normal with 1% volatility on 80% of days and normal with 4% volatility on the other 20%, both with zero mean. What are the overall daily volatility and the kurtosis of this mixture?Continuous probability and distributionsHardTwo SigmaNew York · 2025

    Try it first

    What is the overall daily volatility of the mixture?

    Show the worked solution

    Overall volatility is 2% a day and the kurtosis is 9.75, against 3 for a normal distribution. Variances mix in proportion: 0.8 x 1 + 0.2 x 16 = 4, so volatility is 2%. Fourth moments mix the same way, and each normal contributes 3 times its volatility to the fourth: 0.8 x 3 + 0.2 x 768 = 156. Dividing by the variance squared, 16, gives 9.75.

    Why is a mixture of normals not normal?

    Think of a commute that takes 30 minutes on most days and two hours on strike days. The average trip hides the shape: most days cluster tightly and a few days sit far out. Mixing a calm regime with a wild one gives more small moves than a normal with the same overall spread, fewer medium ones, and far more big ones. The peak is taller, the shoulders thinner and the tails fatter, which is what {term('kurtosis', 'The fourth moment of a distribution divided by the variance squared; 3 for a normal distribution, higher when tails are fatter.')} measures.

    How do you get the two numbers?

    Work with moments, because they mix in proportion to the weights. The variance is 0.8 x 1 + 0.2 x 16 = 4, so volatility is 2%; the fourth moment is 0.8 x 3 x 1 + 0.2 x 3 x 256 = 2.4 + 153.6 = 156, and kurtosis is 156 / 4 squared = 9.75. Look at where the 156 comes from: 153.6 of it is the wild days, which occur only one day in five. The fourth power makes rare large moves dominate.

    The relationship
    σ2=∑iwiσi2=4,κ=∑iwi⋅3σi4(∑iwiσi2)2=15616=9.75\sigma^2 = \sum_i w_i\sigma_i^2 = 4, \qquad \kappa = \frac{\sum_i w_i \cdot 3\sigma_i^4}{\left(\sum_i w_i\sigma_i^2\right)^2} = \frac{156}{16} = 9.75
    w_ithe share of days in each regime, 0.8 and 0.2
    sigma_ithe volatility in each regime, 1% and 4%
    3 sigma_i^4the fourth moment of a zero-mean normal with volatility sigma_i
    What it says in wordsAverage the second and fourth moments across regimes, then divide the fourth moment by the variance squared.
    Calm days plus rare wild days: a taller peak and fatter tails-8%-4%0+4%+8%Daily returnmixture, peak 0.34normal, same 2% volshaded: beyond 5.0%,where the mixture is higherRight tail, heights magnified 21 times+5%+6%+7%+8%mixturenormalBeyond 6% either way:2.7% of days vs 0.27%
    Against a normal with the same 2% volatility, the mixture has a taller peak, thinner shoulders and fatter tails: a daily move beyond 6% either way happens on 2.7% of days under the mixture but only 0.27% under the normal, about 10 times as often.

    What does this mean for a risk model?

    A model that fits a normal to the 2% volatility is right about the average day and wrong about the days that matter. It says a move beyond 6% happens on about 0.27% of days, roughly once in 370 trading days; the mixture says 2.7%, roughly once in 37. That is the usual story of market returns: calm stretches and volatile stretches, each close to normal, adding up to fat tails. A model that lets volatility change over time captures much of it.

    Where candidates lose it

    The first slip is averaging the volatilities, 0.8 x 1% + 0.2 x 4% = 1.6%. Variances average in a mixture, not standard deviations, so the answer is 2%.

    The second is guessing that a mixture of normals has kurtosis 3 because each piece does. Mixing different variances always pushes kurtosis above 3, and here the fourth-power weight on the wild days takes it to 9.75.

    What the interviewer asks next

    • What mixture weight on the 4% regime maximises the kurtosis?
    • What is the probability density of the mixture at zero, compared with the normal?
    • If the two regimes had different means but the same volatility, what would happen to skew and kurtosis?

    Asked at Two Sigma, Quantitative Research, New York, 2025 (Wall Street Oasis): They asked a couple questions involving Mixture Gaussians (e.g., probability density and moments).

  10. 068You and an opponent each secretly show heads or tails. If both show heads you win Rs 3, if both show tails you win Rs 1, and if they differ you pay Rs 2. The payoffs look balanced. What mix should each player use, and what is the game worth to you?Expected value and dice gamesHardCitadelsydney · 2025

    Try it first

    If both of you play well, what is the game worth to you per round?

    Show the worked solution

    Both players should show heads 3/8 of the time, and the game is worth minus Rs 0.125 a round to you. Choose your mix so the opponent gains nothing by switching: 3p - 2(1 - p) = -2p + (1 - p) gives p = 3/8. The opponent's mix solves the same balance from your side, also 3/8. At those mixes you lose an eighth of a rupee a round, although the payoffs look even.

    Why is a fair coin the wrong strategy?

    Think of a penalty taker and a goalkeeper. If the taker always shoots left, the keeper dives left; the taker's only defence is to mix so the keeper cannot profit from guessing. A mix is right only if it leaves the other side indifferent, and a fair coin here does not: against it the opponent's tails pays you 0.5 x (-2) + 0.5 x 1 = -0.50 a round. So the opponent always shows tails, and you lose Rs 0.50 a round rather than breaking even.

    How do you find the equilibrium mix?

    Set up your payoff for each of your choices as the opponent's chance of heads, q, varies. Showing heads pays 3q - 2(1 - q) = 5q - 2; showing tails pays -2q + (1 - q) = 1 - 3q; they are equal at q = 3/8, where both pay -1/8. The opponent plays 3/8 heads so that nothing you do beats -1/8. By the same balance from the opponent's side, you play 3/8 heads so that nothing the opponent does pushes you below -1/8. That pair is the Nash equilibriumA pair of strategies in which neither player can do better by changing only their own choice..

    The relationship
    5q−2=1−3q  ⇒  q=38,V=5⋅38−2=−185q - 2 = 1 - 3q \;\Rightarrow\; q = \tfrac{3}{8}, \qquad V = 5 \cdot \tfrac38 - 2 = -\tfrac18
    qthe opponent's chance of showing heads
    5q - 2your expected payoff if you show heads
    1 - 3qyour expected payoff if you show tails
    Vthe value of the game to you
    What it says in wordsThe opponent's mix makes your two choices pay the same, and that common payoff is what the game is worth.
    The opponent picks the mix that makes your two choices pay the same-2-1+1+2+3000.250.50.751Opponent's chance of showing heads, qyou show heads: 5q - 2you show tails: 1 - 3qq = 3/8: you get-1/8 either wayYour payoff, Rsopp. Hopp. Tyou H+3-2you T-2+1Both show heads 3/8Value to you: -Rs 0.125Fair coin vs best reply:you get -0.50 a round
    Your payoff from showing heads rises with the opponent's chance of heads and your payoff from tails falls; the lines cross at 3/8, where either choice pays minus Rs 0.125, so the opponent plays 3/8 heads and the game is worth minus Rs 0.125 a round to you.

    Why is the game negative when the payoffs look even?

    Because the balance is in the totals, not in the play. Your two winning cells need coordination the opponent will not give you, while their winning cells, the mismatches, pay the same 2 either way. The opponent can lean towards tails, starving your big Rs 3 cell, and the only price is feeding your small Rs 1 cell. Say what a desk would do with it: ask to be paid about 13 paise a round to play, or ask to swap sides.

    Where candidates lose it

    The trap is adding up the payoffs, 3 and 1 against 2 and 2, calling the game fair and playing a fair coin. That ignores that the opponent chooses too, and against a fair coin their best reply costs you Rs 0.50 a round.

    The second slip is solving for your own mix by making yourself indifferent. Each player's mix is chosen to make the other player indifferent; set up the equation from the opponent's payoffs.

    What the interviewer asks next

    • Change the tails-tails payoff to Rs 2. What are the mixes and the value now?
    • What would you pay per round to play this game from the opponent's side?
    • The opponent is known to show heads half the time. What do you do, and what do you earn?

    Asked at Citadel, Quantitative Trading, sydney, 2025 (Wall Street Oasis): many probability questions for OA. mix of prob and game theory for technical

← PreviousPage 1 of 2
  1. 1
  2. 2
Next →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.