Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
Explore NISM prep
Series-VIII · Equity DerivativesSeries-XII · Securities Markets FoundationSeries-V-A · Mutual Fund DistributorsSeries-XV · Research AnalystSeries-XIX-E · Category III AIF ManagersSeries-XIX-D · Category I & II AIF ManagersSeries-XIX-C · Alternative Investment Fund ManagersSeries-XVI · Commodity DerivativesSeries-VI · Depository OperationsSeries-II-A · Registrars & Transfer AgentsSeries-I · Currency DerivativesSeries-VII · Securities Operations & Risk Management
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Hedge Funds puzzles, solved step by step

Puzzles
100
Traced to a firm
38
Topics
14
Hard
30
Topic
All topicsBetting and sizing5Conditional probability and Bayes7Continuous probability and distributions7Counting and combinatorics7Estimation and mental maths4Expected value and dice games8Logic and brainteasers10Market making and trading games6Options and payoffs5Portfolio and risk maths8Random walks and Markov chains7Returns, compounding and fees7Statistics and estimation11Valuation, accounting and macro riddles8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 1–10 of 15 · filtered from 100Clear filters
  1. 002The classic Russian roulette puzzle: a six-chamber revolver has two bullets in adjacent chambers. The cylinder is spun once, the trigger is pulled and it clicks empty. You must pull again. Is it safer to spin the cylinder again first, or not?Conditional probability and BayesCoreSchonfeldCentral · 2022

    Try it first

    Which gives the better chance of surviving the second pull?

    Show the worked solution

    Do not spin: you survive 75% of the time, against 66.7% if you spin. The empty click puts the cylinder on one of the four empty chambers, each equally likely. Because the two bullets sit together, three of those four empties are followed by another empty and only one is followed by a bullet. A fresh spin throws that information away and gives four chances in six.

    What does the empty click tell you?

    Picture six people in a queue where two friends always stand together. Pick someone at random who is not one of the friends and ask whether the person behind them is a friend. Only one of the four has a friend behind them: the one standing just in front of the pair. The empty click is information: it tells you which chambers you could be on, and using it is the whole puzzle. The cylinder has no memory, but you do.

    After an empty click, only one of four empty chambers leads into a bullet123456fires 1, 2, 3 ...loadedemptyYou just clicked on ...next pull fires ...Chamber 34: emptyChamber 45: emptyChamber 56: emptyChamber 61: LOADEDDo not spin3/4 = 75.0%Spin again4/6 = 66.7%
    After an empty click the cylinder sits on chamber 3, 4, 5 or 6; three of those are followed by an empty chamber and only chamber 6 is followed by a bullet, so not spinning survives 75% of the time against 66.7% for a fresh spin.
    The relationship
    P(safe∣no spin)=34=75%P(safe∣spin)=46≈66.7%P(\text{safe} \mid \text{no spin}) = \frac{3}{4} = 75\% \qquad P(\text{safe} \mid \text{spin}) = \frac{4}{6} \approx 66.7\%
    3/4three of the four empty chambers are followed by another empty
    4/6four of six chambers are empty after a fresh spin
    What it says in wordsConditioning on the empty click beats resetting to the base rate when the bullets sit together.

    Why does the answer flip if the bullets are not adjacent?

    Separate the bullets, say into chambers 1 and 4. The four empties are now 2, 3, 5 and 6, and the chambers after them are 3, 4, 6 and 1: two empties and two bullets. Not spinning now survives only 2 times in 4, 50%, so spinning, at 66.7%, becomes the better choice. Adjacency is what bunches both bullets behind a single empty chamber. Ask where the bullets sit before you answer, and say that the answer depends on it.

    Where does this reasoning show up on a desk?

    The same move, updating on what you have just observed instead of resetting to the base rate, is how a trader reads a fill. Getting filled on your bid tells you something about who was selling, just as the empty click tells you which chamber you are on. Ignoring it is the equivalent of spinning the cylinder: it feels neutral, but it throws away an edge you were handed for free.

    Where candidates lose it

    Candidates say it makes no difference, because a spin feels like a clean reset and the cylinder has no memory. The trap is treating no memory in the device as no information for you: the click has ruled out the two loaded chambers as your position.

    The second loss is answering without checking the layout. The case for not spinning rests entirely on the bullets being adjacent; with the bullets apart, the answer reverses. Name that condition in your answer.

    What the interviewer asks next

    • You survive the second pull without spinning. Should you spin before a third?
    • Three bullets in adjacent chambers: spin or not?
    • What if the two bullets are in chambers 1 and 4?

    Asked at Schonfeld, Quantitative Research, Central, 2022 (Wall Street Oasis): Coding, requires to know DP and divde and conquer., Russian Roulette

  2. 003X and Y are independent and each uniform on 0 to 1. What is the probability that X + Y is less than 1.5, and what shape is the density of X + Y?Continuous probability and distributionsCoreCitadelChicago · 2025

    Try it first

    Pick before you draw anything.

    Show the worked solution

    The probability is 7/8, and the density of X + Y is a triangle, a tent peaking at 1. Because X and Y are independent and uniform, every point of the unit square is equally likely, so probability is area. The line x + y = 1.5 slices off a corner triangle with legs of 0.5, area 1/8. The sum's density rises in a straight line from 0 to 1 and falls back to 0 at 2.

    Why does probability become area here?

    Throw a dart at a square board so that every point is equally likely to be hit. The chance it lands in a region is that region's share of the board. Two independent uniforms are exactly that dart: the pair (X, Y) lands evenly on the unit square, so any question about X and Y becomes a question about an area. The condition X + Y below 1.5 is everything under the line x + y = 1.5, which is the whole square except one corner.

    Probability is area: the missing corner is 1/8, so the answer is 7/8x + y = 1.51/8X + Y below 1.5area 1 - 1/8 = 7/8000.50.511XY011.521peak at 1: the tenttail beyond 1.5area 1/8Density of X + Y
    The line x + y = 1.5 removes a corner triangle of area 1/8 from the unit square, so X + Y is below 1.5 with probability 7/8, and the density of X + Y is a tent on 0 to 2 whose tail beyond 1.5 also has area 1/8.

    How do you get the shape of the sum's density?

    Slide the line x + y = s across the square and watch how long it is inside. Near s = 0 it barely clips the corner; at s = 1 it runs corner to corner, the longest it gets; past 1 it shortens again. The density of the sum at s is proportional to the length of that line inside the square, which gives a triangle rising from 0 to a peak at 1 and falling to 2. This is the convolutionThe density of a sum of independent variables, found by adding up every way the two parts can combine to the same total. of two flat densities, and the same reason two dice most often total 7.

    The relationship
    fX+Y(s)=∫01fX(x) fY(s−x) dx={s0≤s≤12−s1≤s≤2f_{X+Y}(s) = \int_0^1 f_X(x)\, f_Y(s-x)\, dx = \begin{cases} s & 0 \le s \le 1 \\ 2 - s & 1 \le s \le 2 \end{cases}
    f_{X+Y}(s)the density of the sum at the value s
    f_Y(s - x)equal to 1 when s - x lies between 0 and 1, otherwise 0
    What it says in wordsAdd up every split of s into an x and a y that both lie in 0 to 1; the count of splits rises to s = 1 and then falls.

    Check the first answer with the tent. The area beyond 1.5 is a triangle with base 0.5 and height 0.5, which is 1/8 again. Two routes that agree is the check an interviewer wants to hear before you commit. Add a third uniform and the density becomes three joined curved pieces; add many and the sum looks normal, which is the central limit theorem arriving in slow motion.

    Where candidates lose it

    Candidates reach for a double integral before drawing, set the limits wrongly, and spend two minutes on what is a one-line area argument. Draw the square first; the corner triangle is visible at a glance.

    The second loss is saying the sum of two uniforms is uniform on 0 to 2. It is not: there is only one way to get a sum near 0 and many ways to get a sum near 1, which is why the density is a tent and not a flat line.

    What the interviewer asks next

    • What is the probability that X + Y is less than 0.5?
    • What is the probability that the larger of X and Y is below 0.5, and how does the picture change?
    • What does the density of X + Y + Z look like?

    Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis): He was asking some questions about the probability, especially on the convolution.

  3. 006You owe exactly Rs pi, that is Rs 3.14159..., and can only pay in whole paise. How do you pay a fair amount on average, and what is the chance you end up paying Rs 3.15?Expected value and dice gamesCoreMillennium ManagementSheung Wan · 2025

    Try it first

    Under the fair scheme, what is the chance you pay Rs 3.15?

    Show the worked solution

    Randomise: pay Rs 3.15 with probability 0.159 and Rs 3.14 otherwise. Pi is 3.14159..., which sits 0.159 of the way from 3.14 to 3.15. Paying the higher amount with exactly that probability makes the expected payment 3.14 + 0.01 x 0.1593..., which is pi. So the chance you pay Rs 3.15 is about 15.9%, and over many meals nobody is short-changed.

    Why can no fixed amount be fair?

    Always round to Rs 3.14 and the restaurant loses 0.159 paise every time; always pay 3.15 and you overpay 0.841 paise. Any fixed amount is unfair to one side, so the only way to be exactly fair is to be fair on average. Two friends who split a Rs 101 bill by taking turns to pay the odd rupee are doing the same thing: neither is exact on any one night, both are exact over time.

    Weight each paisa by how close pi is to it, and the beam balances at pi84.1%15.9%pay Rs 3.14pay Rs 3.15Rs 3.14Rs 3.15pi = 3.14159...0.159 paise0.841 paise: the gap to 3.15Expected payment = 3.14 x 0.8407 + 3.15 x 0.1593= 3.14159265..., exactly pi
    Pi sits 0.159 paise above Rs 3.14 and 0.841 paise below Rs 3.15, so paying Rs 3.15 with probability 15.9% and Rs 3.14 otherwise balances exactly at pi, which makes the expected payment fair.
    The relationship
    E[pay]=3.14 (1−p)+3.15 p=π  ⟺  p=π−3.140.01≈0.1593E[\text{pay}] = 3.14\,(1-p) + 3.15\,p = \pi \iff p = \frac{\pi - 3.14}{0.01} \approx 0.1593
    pthe probability of paying Rs 3.15
    \pi - 3.14how far pi sits above the lower whole-paisa amount
    What it says in wordsThe chance of paying the higher amount equals how far along the gap pi lies.

    How do you actually draw a probability of 0.159?

    Use any randomness you can split finely. Draw a uniform number between 0 and 1 and pay Rs 3.15 if it falls below 0.1593. With only a die, paying 3.15 on a six gives 1/6, an expected payment of Rs 3.141667: close, not exact. With only a fair coin you can be exact: toss it to generate the binary digits of a uniform number one at a time and stop as soon as the digits so far settle which side of 0.1593 it falls. Each toss settles it with probability one half, so on average it takes two tosses.

    Where does randomised rounding show up in a fund?

    Whenever a quantity has to be split in whole units. A fund allocating 1,003 shares across three accounts cannot give each 334.33; handing the odd share out by lottery, or in rotation, keeps each account fair on average. The principle is the same: when the exact amount is impossible, make the expected amount exact and keep the error unbiased. The limitation is that fair on average is not fair every time, which is why allocation policies also cap how far any account can drift.

    Where candidates lose it

    Candidates round to Rs 3.14 and argue the gap is too small to matter. The interviewer is not asking about a sixth of a paisa; the question is whether you see that a fair expected value can be built from amounts that are each individually wrong.

    The second loss is the coin flip. Fifty-fifty between 3.14 and 3.15 feels even-handed but averages 3.145, overpaying by nearly half a paisa every time. The probability has to match where pi sits in the gap.

    What the interviewer asks next

    • How would you hit the probability exactly using only a fair coin?
    • How many coin tosses does that take on average?
    • What if you owe Rs e, 2.71828...?

    Asked at Millennium Management, Quantitative Research, Sheung Wan, 2025 (Wall Street Oasis): How to pay the restaurant fairly if I owe pi dollars. Need to pay with usual dollars and cents.

  4. 014A corporate bond has a spread duration of 6 and convexity of 50. Its credit spread widens by 50 basis points. Roughly what happens to its price?Valuation, accounting and macro riddlesCoreACAQR Capital ManagementGreenwich · 2021

    Try it first

    Which is closest?

    Show the worked solution

    The price falls by about 2.94%. Spread duration of 6 says a 0.50 percentage point widening costs 6 x 0.50% = 3.00%. Convexity of 50 adds back one half x 50 x 0.005 squared, about 0.06%, because the price curve bends upwards. On a bond priced at 100 that is a move to about 97.06. At 50 basis points the convexity term is small; at 300 or 500 it is not.

    What do duration and convexity each measure?

    Picture a playground slide that curves and flattens towards the bottom. Judge the drop from the steepness at the top and you overstate it, because the slide levels off as you go. Spread durationThe percentage change in a bond price for a one percentage point change in its credit spread, holding the risk-free rate fixed. is the steepness at today's spread; convexity is the flattening, so the straight-line estimate always overstates the loss when spreads widen. Duration gives the first-order move, 6 x 0.50% = 3.00% down; convexity corrects it by a term that depends on the square of the move.

    Duration is the straight line; convexity is how the curve bends away7080901000100200300400500Spread widening, basis points+2.25+6.25+50 bp: -2.94%curve: duration plus convexityduration onlyMoveDurationConvexityTotal+50 bp-3.00%+0.06%-2.94%+300 bp-18.00%+2.25%-15.75%+500 bp-30.00%+6.25%-23.75%Convexity grows with the squareof the move: tiny at 50 bp,a fifth of the gross loss at 500 bp
    For a 50 basis point widening, duration of 6 gives minus 3.00% and convexity of 50 adds back 0.06%, a fall of 2.94%; the convexity cushion grows with the square of the move, to 2.25 points at 300 basis points and 6.25 at 500.
    The relationship
    ΔPP≈−Ds Δs+12C (Δs)2=−6(0.005)+12(50)(0.005)2=−3.00%+0.0625%≈−2.94%\frac{\Delta P}{P} \approx -D_s\,\Delta s + \tfrac{1}{2}C\,(\Delta s)^2 = -6(0.005) + \tfrac{1}{2}(50)(0.005)^2 = -3.00\% + 0.0625\% \approx -2.94\%
    D_sspread duration, 6
    Cconvexity, 50
    \Delta sthe change in spread as a decimal, 50 basis points = 0.005
    What it says in wordsThe price moves by the duration term plus a smaller correction that grows with the square of the spread change.

    When does the convexity term start to matter?

    It grows with the square of the move. At 50 basis points convexity is worth 0.06% against a 3.00% duration loss; at 300 basis points duration says -18% and convexity adds back 2.25%, which is no longer small. That is why a credit desk can run duration-only risk for everyday moves but needs convexity for stress scenarios. One more distinction marks a strong answer: for a fixed-coupon bond spread duration and rate duration are close, but a floating-rate note has almost no rate duration and still carries several years of spread duration.

    Say the limitation plainly. Both numbers are local, measured at today's spread, and a distressed bond stops behaving like this long before default, when its price starts tracking the expected recovery instead. For a bond trading near par, as here, the two-term estimate is good to a few hundredths of a per cent for moves of this size.

    Where candidates lose it

    Candidates give minus 3% and stop, which is fine as a first line but ignores the second number the question handed you. Worse is using convexity with the wrong sign, making the loss bigger: for a plain bond convexity always cushions a spread widening.

    The other slip is units. Fifty basis points is 0.005 in the formula; squaring 0.50 instead turns a 0.06% correction into 6.25% and produces a price that rises when spreads widen.

    What the interviewer asks next

    • What if the spread tightens by 50 basis points instead?
    • Why can a callable bond have negative convexity?
    • How would you hedge the spread risk of this bond?

    Asked at AQR Capital Management, Investment Research, Greenwich, 2021 (Wall Street Oasis): Discussion on credit spreads on fixed income products and duration.

  5. 026You have five trade ideas. Each needs some units of risk budget and carries an expected profit: A needs 4 units for Rs 9 crore, B 3 units for Rs 7 crore, C 5 units for Rs 10 crore, D 2 units for Rs 4 crore and E 6 units for Rs 11 crore. Your budget is 10 units and you cannot take part of an idea. Which ideas do you take?Betting and sizingCoreBridgewater AssociatesNew York · 2025

    Try it first

    Before you enumerate: which set do you expect to win?

    Show the worked solution

    Take B, C and D: 10 units for Rs 21 crore. Ranking by profit per unit of risk picks B (2.33), then A (2.25), then D (2.00), which uses 9 units for Rs 20 crore and leaves one unit idle. Swapping A for C costs a little ratio but fills the budget and adds Rs 1 crore. With whole ideas, you check the few sets that fill the budget rather than trust the ratio alone.

    Why is profit per unit of risk the right place to start?

    Packing a suitcase for a flight with a weight limit, you favour the things that give the most use per kilo. When the budget is the scarce thing, the useful measure is profit per unit of budget, not profit alone. E earns the most in rupees but only 1.83 per unit, the worst of the five. B earns 2.33 per unit, A 2.25, and C and D 2.00 each. If ideas could be split, you would simply fill from the top of that list and the answer would be exact.

    Why does the ratio ranking fail here?

    Ideas come whole. Take B and A and you have used 7 units; D fits for 9 units, and nothing else fits in the last one. A unit left idle earns nothing, so a set with a slightly lower average ratio that uses the full budget can beat the greedy pick. Replace A (4 units) with C (5 units) and the budget is full: B, C and D earn 7 + 10 + 4 = Rs 21 crore against Rs 20 crore. This is the knapsack problemChoosing whole items, each with a size and a value, to maximise total value without exceeding a capacity., and the fix is to check the handful of sets that nearly fill the budget.

    Every full set that fits 10 risk units, ranked by expected P&LIdea A4u, Rs 9 cr, 2.25/uIdea B3u, Rs 7 cr, 2.33/uIdea C5u, Rs 10 cr, 2.00/uIdea D2u, Rs 4 cr, 2.00/uIdea E6u, Rs 11 cr, 1.83/uRisk units usedExpected P&LB + C + DBCDRs 21 crbest: all 10 units usedA + B + DABDRs 20 crper-unit ranking picks thisA + EAERs 20 crA + CACRs 19 crB + EBERs 18 crA + B + CABCRs 26 crover budgetbudget: 10 units05
    B, C and D use all 10 risk units for Rs 21 crore, the best set that fits; ranking by profit per unit picks B, A and D, which uses 9 units for Rs 20 crore, and A, B and C would earn Rs 26 crore but needs 12 units.

    With five ideas there are only 31 possible sets, so enumerate the full ones out loud: B, C, D for 21; A, B, D for 20; A, E for 20; A, C for 19. On a real book with hundreds of positions a desk uses an optimiser, but the logic is the same, and interviewers want to hear that you know the ratio rule is a starting point and not a proof.

    Say the limitation too. Expected profit is an estimate, and the gap here is Rs 1 crore on Rs 21 crore. If C's estimate is shakier than A's, a portfolio manager could reasonably prefer the greedy set. The arithmetic tells you the best set on the stated numbers; confidence in each number decides whether that edge is real.

    Where candidates lose it

    The common loss is ranking by profit per unit, taking the top three, and stopping. That is right for divisible positions and wrong here, because the tenth unit of budget sits unused and the interviewer built the numbers to punish exactly that.

    The opposite loss is ranking by rupee profit and grabbing E first. E is the least efficient idea on the list. Say the ratio rule, then check the sets that fill the budget, then say how sure you are of each estimate.

    What the interviewer asks next

    • The budget rises to 11 units. What changes?
    • Ideas B and C are highly correlated, so together they use 9 units instead of 8. Does your answer move?
    • You can take half of any idea for half its units and half its profit. What do you take now?

    Asked at Bridgewater Associates, Generalist, New York, 2025 (Wall Street Oasis): Given a list of items and their utilitites, give an algorithm to maximise utility

  6. 027A trade surveillance screen flags 90% of genuinely suspicious trades and wrongly flags 5% of clean ones. One trade in a hundred is suspicious. A trade has just been flagged. What is the chance it is actually suspicious?Conditional probability and BayesCoreCitadelMiami · 2022

    Try it first

    Gut answer first: a flagged trade is suspicious with probability about

    Show the worked solution

    About 15%, not 90%. Picture 10,000 trades. 100 are suspicious and the screen flags 90 of them. 9,900 are clean and it wrongly flags 5%, which is 495. So 585 trades are flagged and only 90 are suspicious: 90 over 585 is 15.4%. The false flags swamp the true ones because clean trades are so common.

    Why is 90% the wrong number?

    A smoke alarm that goes off whenever there is a fire is good. But if it also goes off every time someone makes toast, most of its alarms are toast. The 90% tells you how often a suspicious trade gets flagged; the question asks how often a flag is suspicious, and those run in opposite directions. Mixing them up is called the {term('base rate', 'How common something is before any test is run. Here, one trade in a hundred is suspicious.')} fallacy, and it is exactly what this question is built to catch.

    How do you get the number without the formula?

    Use counts, not percentages. Start with 10,000 trades because it makes every number whole. Split by the truth first, then by what the screen says, and then read only the flagged column. 1% of 10,000 is 100 suspicious trades; 90% of those, 90, are flagged. 9,900 are clean; 5% of those, 495, are flagged anyway. The flagged column holds 585 trades, and 90 of them are the real thing.

    Follow 10,000 trades through the screen: false flags outnumber true ones10,000trades1%99%100suspicious9,900clean90%10%5%95%90true flags10missed495false flags9,405clearedAll flags: 585495 false9090 / 58515.4%A flag is right about1 time in 6.5
    Of 10,000 trades, 100 are suspicious and 90 of those are flagged, while 495 of the 9,900 clean trades are flagged by mistake, so only 90 of the 585 flags, 15.4%, point at a suspicious trade.
    The relationship
    P(S∣F)=0.90×0.010.90×0.01+0.05×0.99=0.0090.0585≈15.4%P(S\mid F) = \frac{0.90 \times 0.01}{0.90 \times 0.01 + 0.05 \times 0.99} = \frac{0.009}{0.0585} \approx 15.4\%
    Sthe trade is suspicious
    Fthe screen flags the trade
    0.01the base rate: one trade in a hundred is suspicious
    What it says in wordsThe chance a flag is right equals true flags divided by all flags.

    Add the desk point after the number. A second, independent check changes things fast: run the 585 flagged trades through a second screen with the same error rates and the 15.4% prior becomes about 77%. A weak test is still useful as a first filter; it is only misleading when its hit rate is read as its accuracy.

    Where candidates lose it

    Most candidates say 90% within a second, because the question hands them that number. It is the chance of a flag given a suspicious trade, and the interviewer asked the reverse.

    The second loss is starting on Bayes' formula with decimals and getting tangled. Say you will use 10,000 trades, draw the two splits, and the answer reads straight off the flagged column.

    What the interviewer asks next

    • What false flag rate would make a flag right half the time?
    • The flagged trades go through a second independent screen and are flagged again. What is the chance now?
    • Compliance wants to catch 99% of suspicious trades. What does that usually do to the false flag rate?

    Asked at Citadel, Sales and Trading, Miami, 2022 (Wall Street Oasis): I got a question about Bayes' theorem applied to a practical scenario

  7. 033Make a two-way market on the sum of three fair dice. Then one die is revealed to be a 6. Where do you move your market, and should it get wider or narrower?Market making and trading gamesCoreCitadelLondon · 2026

    Try it first

    After the 6 is shown, what happens to your market?

    Show the worked solution

    Move the mid from 10.5 to 13 and tighten the market by about a fifth. Each die averages 3.5, so three dice average 10.5. Once one die shows 6, the sum is 6 plus two unknown dice averaging 7, which is 13. The variance falls from 3 x 35/12 to 2 x 35/12, so the standard deviation drops from 2.96 to 2.42. If the first market was 9.5 at 11.5, the new one is about 12.2 at 13.8.

    Where do you put the first market, and how wide?

    Start from the fair value and then decide the width from how uncertain the outcome is. The mid is the expected sum, 3 x 3.5 = 10.5, and the width should scale with the standard deviation of the sum, because that is how far the answer typically lands from the mid. One die has variance 35/12, so three independent dice have 35/4 = 8.75, a standard deviation of 2.96. A market of 9.5 bid, 11.5 offered is a reasonable opening: tight enough to trade, with room for your edge.

    What does revealing one die change?

    A weather forecast for tomorrow is more precise than one for next week, because fewer things can still change. Once one die is known, it contributes a certain 6 and no uncertainty, so the mid rises by 2.5 and only two dice of variance remain. The mid becomes 6 + 7 = 13. The variance becomes 35/6, a standard deviation of 2.42, down from 2.96. Scale the width by the same ratio, about 0.82, and a 2.0 wide market becomes about 1.6 wide: 12.2 at 13.8.

    The reveal shifts the centre up 2.5 and narrows the spread of outcomes5%10%15%3456789101112131415161718Sum of the three diceBefore: 9.5 / 11.5After a 6: 12.2 / 13.8Before: mean 10.5, sd 2.96After a 6: mean 13, sd 2.42
    Before the reveal the sum is centred on 10.5 with a standard deviation of 2.96; after one die shows 6 it is centred on 13 with a standard deviation of 2.42, so the market moves up by 2.5 and tightens from 2.0 wide to about 1.6.

    Say what would make you widen instead. If the person revealing the die can choose which die to show, or picks the moment, the reveal itself carries information and you should be more careful, not less. A 6 chosen as the highest of three tells you the other two are 6 or lower, and they no longer average 7. Interviewers like it when you ask who chose what to reveal before you requote.

    Where candidates lose it

    The common loss is widening after the 6 because it feels like a shock. A shock that is fully known removes uncertainty. The mid jumps, but the range of outcomes shrinks.

    The second loss is moving the mid by the full 6, or to 16.5 as if all dice were sixes. Only the revealed die is known; the other two still average 3.5 each.

    What the interviewer asks next

    • A second die is revealed as a 1. Where is your market now?
    • The revealer chose to show the highest of the three dice. Where do you quote?
    • Someone lifts your 13.8 offer straight away. What do you do next?

    Asked at Citadel, Quantitative Research, London, 2026 (Wall Street Oasis): 3rd I got rejected it was different brainteasers and trading game

  8. 048A strategy's true annual Sharpe ratio is 1.0. How many years of monthly returns do you need before its average return shows a t-statistic of 2? What if the true Sharpe is 0.5?Statistics and estimationCoreViking Global InvestorsNew York · 2014

    Try it first

    Years needed for a Sharpe of 0.5:

    Show the worked solution

    About 4 years for a Sharpe of 1.0 and about 16 years for a Sharpe of 0.5. The t-statistic of a mean return is the mean over its standard error, which works out to the annual Sharpe ratio times the square root of the number of years, whatever the data frequency. Setting Sharpe x root(years) = 2 gives years = (2 / Sharpe) squared: 4 for 1.0, 16 for 0.5 and just 1 for 2.0.

    Why does the t-statistic grow with the square root of time?

    A coin that lands heads 55% of the time looks fair after 20 tosses; you need hundreds before the bias shows through the noise. The average return grows in proportion to time, but the noise around it grows only with the square root of time, so the signal-to-noise ratio, the t-statistic, grows with root time. With monthly data, the t-statistic is the monthly Sharpe times root(12 x years), and the monthly Sharpe is the annual Sharpe divided by root 12, so the twelves cancel: t = annual Sharpe x root(years).

    t = Sharpe x root(years): halving the Sharpe quadruples the wait012345048121620Years of monthly returnst = 21 yr4 yrs16 yrsSharpe 2.0Sharpe 1.0Sharpe 0.5
    Because the t-statistic equals the Sharpe ratio times the square root of years, a Sharpe of 2.0 clears t = 2 after 1 year, a Sharpe of 1.0 after 4 years and a Sharpe of 0.5 only after 16 years.

    Why does monthly data not shorten the wait?

    More frequent data gives more observations but each is noisier relative to its mean. Sampling the same years more often does not add information about the mean return; only more years do. This is why a {term('t-statistic', 'An estimate divided by its standard error; a value around 2 is the usual threshold for saying an effect is unlikely to be pure noise.')} on the average return depends on the span of the data, not the number of rows. Frequency helps you estimate volatility, not the mean.

    The relationship
    t≈SRannualY  ⇒  Y=(2SR)2:SR=1→4,SR=0.5→16t \approx SR_{\text{annual}}\sqrt{Y} \;\Rightarrow\; Y = \left(\frac{2}{SR}\right)^2: \quad SR = 1 \to 4, \quad SR = 0.5 \to 16
    SRthe true annual Sharpe ratio
    Yyears of data
    2the target t-statistic
    What it says in wordsThe years needed to prove a strategy grow with the inverse square of its Sharpe ratio.

    Say the practical point. Most real strategies have Sharpe ratios well below 1, so their track records are too short to separate skill from luck with any confidence. An allocator looking at a three-year record with a Sharpe of 0.8 sees a t-statistic of about 1.4. The limitation of the rule: it assumes returns are independent and stable over the whole sample, and fat tails or regime changes make the real uncertainty larger.

    Where candidates lose it

    The common loss is thinking monthly data gives twelve times the evidence, which leads to answers like four months. The twelve cancels, because the monthly Sharpe is smaller by root 12.

    The second loss is saying a Sharpe of 0.5 needs twice as long as 1.0. The dependence is on the square: half the Sharpe, four times the data.

    What the interviewer asks next

    • How many years for a Sharpe of 0.3?
    • You test 20 strategies and pick the best one with t = 2.2. How much do you trust it?
    • Would daily data change the answer for estimating the Sharpe ratio itself rather than the mean?

    Asked at Viking Global Investors, Quantitative Research, New York, 2014 (Wall Street Oasis): how to reject a hypothesis test, what's your structure of your code, what's the sample size

  9. 051You start with Rs 1,000 and can bet any amount on a coin that wins 60% of the time at even money, as many times as you like. What fraction of your money do you bet each time, and what happens if you bet twice that?Betting and sizingCoreMillennium ManagementAtlanta · 2025

    Try it first

    Which fraction of your money gives the fastest long-run growth?

    Show the worked solution

    Bet 20% of your current money each time; bet 40% and the long-run growth is gone. For an even-money bet the Kelly fraction is the win probability minus the loss probability, 0.6 minus 0.4. At 20% the typical path grows about 2.0% a bet. At double that, the average log growth is about -0.2% a bet, so the same edge now earns nothing over time.

    Why not bet everything when the odds are in your favour?

    Think of a shopkeeper who puts all of today's takings into tomorrow's stock. Six days in ten the stock sells and she doubles her money; on the other four she is back to zero, and one zero ends the shop. Wealth compounds, so a single wipe-out is permanent, and the stake that maximises the average payoff of one round is not the stake that maximises what you end with after many. Betting 100% has the highest expected value per round and a certainty of eventual ruin.

    How do you find the best fraction?

    Maximise the growth rate, which is the average of the log of what each bet does to your money. Bet a fraction f and a win multiplies your wealth by 1 + f, a loss by 1 - f. The growth rate 0.6 ln(1 + f) + 0.4 ln(1 - f) peaks where f equals the edge, 0.6 - 0.4 = 20%, at about 2.01% a bet. Over 100 bets that turns Rs 1,000 into about Rs 7,490 on the typical path. This rule is the Kelly criterionA sizing rule that picks the bet size maximising the long-run growth rate of wealth, the expected log return per bet..

    The relationship
    g(f)=pln⁡(1+f)+qln⁡(1−f),f∗=p−q=0.6−0.4=0.2g(f) = p\ln(1+f) + q\ln(1-f), \qquad f^{*} = p - q = 0.6 - 0.4 = 0.2
    pthe chance of winning, 0.6
    qthe chance of losing, 0.4
    fthe fraction of current money bet
    g(f)the expected log growth per bet
    What it says in wordsGrowth per bet is the probability-weighted average of the log gain and the log loss, and it is highest when you bet the edge.
    Growth per bet against the fraction you bet: the peak sits at the edge-4%-2%+2%0%0%10%20%30%40%50%Fraction of current money bet each timeKelly: bet 20%, growth +2.0% a betDouble Kelly, 40%: -0.2% a beta real edge, and no growth leftshaded: negative long-run growth
    Expected log growth per bet rises to about 2.0% at a 20% stake, falls back to slightly below zero, -0.2%, at a 40% stake and turns sharply negative beyond it, so over-betting a real edge can remove all long-run growth.

    What happens at twice the Kelly bet?

    Near the peak the growth curve is close to a parabola centred on 20%. Double the Kelly fraction and you sit as far down the far side as not betting at all sits on the near side: growth of about zero, -0.24% a bet. Over 100 bets the typical path ends near Rs 783, below where you started, although every single bet had a positive expected value. The extra swings cost more in compounding than the extra stake earns.

    Say the limitation too. Kelly assumes you know the 60% exactly. On a desk the edge is an estimate, and overestimating it pushes you to the right of the peak, where the curve falls fastest. That is why many traders bet half Kelly: 10% here keeps about 75% of the growth with half the swings.

    Where candidates lose it

    The fast wrong answer is to bet big because the odds favour you, often 60% because that is the chance of winning. A candidate who only computes expected value per round will always pick the largest stake, and the interviewer is waiting to see whether you notice that money compounds.

    The second loss is reaching 20% and being unable to say what happens beyond it. Have the shape ready: growth peaks at the edge, is roughly zero at twice the edge, and is negative after that.

    What the interviewer asks next

    • The coin now pays 2 to 1 when you win and still wins 60% of the time. What is the Kelly fraction?
    • You are only 80% sure the coin is 60/40 rather than fair. How does that change your bet?
    • What would you pay to play 100 rounds of this game starting with Rs 1,000?

    Asked at Millennium Management, Software Engineering Intern Interview, Atlanta, 2025 (Wall Street Oasis): Start with 1000 and bet each round. Kelly criterion would be very useful for this step.

  10. 056Game A: roll two fair dice and be paid the product of the faces in rupees. Game B: roll one fair die and be paid the square of its face. Which game is worth more, and by exactly how much?Expected value and dice gamesCoreCitadelMiami · 2022

    Try it first

    Which game would you rather play, and by roughly how much?

    Show the worked solution

    The one-die game is worth more: Rs 15.17 against Rs 12.25, a gap of 35/12, about Rs 2.92. With two independent dice the average product is the product of the averages, 3.5 x 3.5 = 12.25. With one die the average square is (1 + 4 + 9 + 16 + 25 + 36)/6 = 91/6. The gap between the two is exactly the variance of one die.

    Why is averaging a square not the same as squaring an average?

    Take two students who score 2 and 8 in a test. Their average is 5, and 5 squared is 25. Square each score first and then average, and you get (4 + 64)/2 = 34. Squaring rewards the high value more than it penalises the low one, so the average of the squares is always at least the square of the average, and the gap is the spread. Here the gap, 9, is exactly the variance of the two scores.

    How do the two games come out?

    In game A the dice are independent, so the average of the product is the product of the averages. Two independent dice give 3.5 x 3.5 = Rs 12.25, because a high roll on one die is as likely to meet a low roll on the other as a high one. In game B one number is multiplied by itself, so a 6 always meets a 6 and a 1 always meets a 1. The average of the six squares is 91/6, about Rs 15.17.

    The relationship
    E[X2]=Var(X)+E[X]2=3512+12.25=916≈15.17,E[XY]=E[X] E[Y]=12.25E[X^2] = \mathrm{Var}(X) + E[X]^2 = \tfrac{35}{12} + 12.25 = \tfrac{91}{6} \approx 15.17, \qquad E[XY] = E[X]\,E[Y] = 12.25
    X, Ythe faces of two independent dice
    Var(X)the variance of one die, 35/12
    E[X]the average face, 3.5
    What it says in wordsThe average square is the square of the average plus the variance; the average product of independent dice has no variance term.
    Squaring one die pays for its spread; independent dice do notOne die squared: the six equally likely payouts1face 14face 29face 316face 425face 536face 6mean of squares 15.173.5 x 3.5 = 12.25Game A: two dice, paid the product12.25Game B: one die, paid its square15.17+2.92The gap is the variance of one die:15.17 - 12.25 = 35/12 = 2.92
    The six squares from 1 to 36 average 15.17, above the 12.25 that squaring the average roll gives; the two-dice product game is worth 12.25, so the one-die square game is worth 2.92 more, exactly the variance of one die.

    What is the desk lesson?

    A payoff that is the square of a move gains from dispersion; one built from two independent pieces does not. Correlation is what turns a product into a square: if the second die always copied the first, game A would be game B. Flip it and make the second die show 7 minus the first, and the average product falls to 9.33. Say that range out loud and the interviewer knows you see the payoff as a bet on varianceThe average squared distance of an outcome from its mean; for one fair die it is 35/12, about 2.92. as well as on the average.

    Where candidates lose it

    The quick wrong answer is that the games are worth the same, because both feel like three and a half times three and a half. That is true only for independent dice. Squaring one die ties the factors together.

    The second loss is getting 15.17 and 12.25 and not naming the gap. Say that 2.92 is the variance of a die; that one sentence is what the interviewer is waiting for.

    What the interviewer asks next

    • What is game A worth if you are paid the sum of the two dice instead of the product?
    • The second die always shows 7 minus the first. What is the average product now?
    • What is the expected value of the larger of two dice?

    Asked at Citadel, Sales and Trading, Miami, 2022 (Wall Street Oasis): No, pretty typical interview questions, dice questions, etc.

← PreviousPage 1 of 2
  1. 1
  2. 2
Next →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.