Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Derivatives Foundation puzzles, solved step by step

Puzzles
100
Traced to a firm
66
Topics
12
Hard
29
Topic
All topicsMental maths and estimation9Random walks and Markov chains7Conditional probability and Bayes7Volatility and correlation7Option pricing intuition7Expected value and optimal stopping10Market making11Option payoffs and no-arbitrage10Probability and counting11Distributions and statistics8Games and logic8Betting and sizing5
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 1–10 of 100
  1. 001How many golf balls fit inside a 100-storey office tower? Talk me through your thought process; I care more about the structure than the final number.Mental maths and estimationCoreTower Research CapitalNew York · 2013

    Try it first

    Before any arithmetic: which single assumption will move your answer the most, once the building's size is fixed?

    Show the worked solution

    About 12 billion, with an honest range of 10 to 14 billion. A floor plate of 4,000 square metres at 3.8 metres a floor gives 1.52 million cubic metres over 100 floors; take half as usable, 0.76 million. A golf ball is 4.27 centimetres across, about 41 cubic centimetres, so 24,500 would fit in a solid cubic metre. Spheres leave gaps: at a packing factor of 0.64 that is 15,700 per cubic metre, and 0.76 million times 15,700 is 11.9 billion.

    Why build a chain instead of guessing a number?

    If someone asks how much rice a storeroom holds, you do not guess a tonnage. You measure the room, decide how much of it is shelving, and know how much a sack holds. An estimation question is marked on whether each link in the chain is stated, sized and defensible, not on the final figure. Here the chain is floors, floor plate, floor height, usable fraction, ball volume, packing factor. Say the six links before you fill any of them in, so the interviewer can follow and correct a single link rather than the whole answer.

    Golf balls in a 100-storey tower: the chain, then the packing factorFloors100Floor plate4,000 sq mFloor height3.8 mGross volume1.52 million cu mUsable, 50%0.76 million cu mOne ball: 4.27 cm across40.8 cu cmIf balls were solid24,531 per cu mSpheres leave gaps: multiply bya packing factorPacking factor and the answer it givesLoose random, 0.5610.4 billionDense random, 0.6411.9 billionClosest packing, 0.7413.8 billionClosest over loose = 1.32: the packing factor alone moves the answer by about a third
    The chain runs from 100 floors, a 4,000 square metre plate and 3.8 metre floors to 1.52 million gross cubic metres and 0.76 million usable, then from a 40.8 cubic centimetre ball to 24,531 balls per solid cubic metre, and the packing factor turns that into 10.4, 11.9 or 13.8 billion balls, a spread of about a third from the packing assumption alone.

    Where does the packing factor come from, and why does it matter so much?

    Pour marbles into a jar and shake it: they settle at roughly 64% of the jar's volume filled, with the rest air. Stack them by hand in the tightest possible pattern and you reach about 74%. Tip them in gently without shaking and you can be as low as 56%. The floor plate and height are things you can look up or pace out, but the packing factor is a physical assumption you must own, and it moves the answer from 10.4 to 13.8 billion on its own. Say which one you are using and why: balls poured into a building are a shaken random pile, so 0.64 is the defensible middle.

    The relationship
    N=100×4,000×3.8×0.540.8×10−6×0.64≈1.19×1010N = \frac{100 \times 4{,}000 \times 3.8 \times 0.5}{40.8 \times 10^{-6}} \times 0.64 \approx 1.19 \times 10^{10}
    100 x 4,000 x 3.8floors times floor plate times floor height, the gross volume in cubic metres
    0.5the usable fraction after the core, structure and services
    40.8 x 10^-6one golf ball in cubic metres
    0.64the packing factor for a shaken random pile of spheres
    What it says in wordsUsable volume divided by the volume of one ball, scaled down for the gaps between balls.

    What do you say when the interviewer pushes on the usable fraction?

    Lifts, stairs, structural columns, ducts and ceiling voids take a large slice of a tower, and nobody knows it to the percent. Give a range and show which way it pushes the answer: 40% usable takes you to about 9.5 billion, 60% to about 14 billion, so the final figure is somewhere between 10 and 15 billion whatever you assume. A candidate who reports a single number to three figures has missed the point of the question. A candidate who says 12 billion, give or take a third, with the reasons, has answered it.

    Where candidates lose it

    The common failure is to fix on one number early, such as the number of balls in a room, and then multiply by a guessed room count without ever stating the building's volume. The interviewer cannot follow it, cannot correct it, and marks it as a guess.

    The second loss is forgetting that spheres do not fill space. Dividing the volume by the ball's volume overstates the answer by about half. Saying the words packing factor, and a number for it, is what separates a trader's estimate from a schoolchild's.

    What the interviewer asks next

    • Now the balls are tennis balls, 6.7 centimetres across. Roughly how does your answer change?
    • If I told you the real answer was 20 billion, which of your assumptions would you revisit first?
    • Make me a market on it: give me a bid and an offer in billions, and tell me how wide and why.

    Asked at Tower Research Capital, Assistant Trader, New York, 2013 (Wall Street Oasis): How many golf balls fit in the empire state building? Explain thought process and detailed solution

  2. 002Every market day is either trending or choppy. A trending day is followed by another trending day 70% of the time, and a choppy day by another choppy day 60% of the time. In the long run, what fraction of days are trending? And when a trend starts, how many days does it last on average?Random walks and Markov chainsCoreQuant tradingHedge funds

    Try it first

    Gut call before the algebra: which is larger, the share of trending days or the share of choppy days?

    Show the worked solution

    4/7 of days trend, about 57.1%, and a trend lasts 3.3 days on average. In the long run the flow out of trending must equal the flow in: 0.3 times the trending share equals 0.4 times the choppy share, so the shares stand 4 to 3. A trend ends on any given day with probability 0.3, so its expected length is 1/0.3, and a choppy spell lasts 1/0.4 = 2.5 days.

    Why must the two flows balance?

    Picture two rooms at a party with a door between them. Each minute, 30% of the people in room A wander into B and 40% of those in B wander into A. The crowd settles when the two queues through the door carry the same number of people; otherwise one room keeps filling. A stationary split is one where the number of days leaving each state equals the number entering it, and that single equation fixes the split. Days leaving trending: 0.3 times the trending share. Days entering it from choppy: 0.4 times the choppy share. Set them equal and the ratio is 4 to 3.

    Two states, four arrows: in the long run the two crossing flows must balanceTrendingdayChoppydaystays trending 0.7stays choppy 0.60.3 of trending days flip0.4 of choppy days flipBalance: 0.3 x (trending share) = 0.4 x (choppy share)so trending : choppy = 4 : 3Long-run share of daysTrending: 57.1% of days4/7average streak 1/0.3 = 3.3 daysChoppy: 42.9% of days3/7average streak 1/0.4 = 2.5 daysCheck: 4/7 x 3.3 against 3/7 x 2.5share = streak x how often a streak starts
    Trending days keep 0.7 of their successors and lose 0.3 to choppy, while choppy days keep 0.6 and lose 0.4 back, so the long-run shares stand 4 to 3, 57.1% trending and 42.9% choppy, with trends lasting 3.3 days and choppy spells 2.5 days on average.
    The relationship
    0.3 πT=0.4 πC,πT+πC=1  ⇒  πT=0.40.3+0.4=47E[streak]=10.3=3.330.3\,\pi_T = 0.4\,\pi_C,\quad \pi_T + \pi_C = 1 \;\Rightarrow\; \pi_T = \frac{0.4}{0.3 + 0.4} = \frac{4}{7} \qquad E[\text{streak}] = \frac{1}{0.3} = 3.33
    pi_T, pi_Cthe long-run shares of trending and choppy days
    0.3, 0.4the chance a trending day flips to choppy, and a choppy day flips to trending
    1/0.3expected length of a run that ends with probability 0.3 each day
    What it says in wordsEach state's share is the other state's flip rate divided by the sum of the two flip rates, and a run's length is one over its own flip rate.

    Why is the average streak 1/0.3 and not something longer?

    A trend that has lasted five days is no more likely to end tomorrow than one that started today: the chain has no memory beyond yesterday. Each trending day ends the run with probability 0.3, independent of its age, so the run length is a geometric count with mean 1/0.3 = 3.33 days. That is the same reason the expected number of rolls to a six is 6. The two answers also check each other: the share of a state equals how often a run of it starts, times how long it lasts, and 4/7 against 3/7 is exactly 3.33 against 2.5 scaled by the same start rate.

    What does the chain say about tomorrow, given today?

    This is where the puzzle connects to trading. Today's state carries real information: after a trending day, tomorrow trends with probability 0.7, well above the unconditional 57%. After a choppy day it is only 0.4. The long-run split tells you nothing about tomorrow; the transition row for today's state does. Say that distinction out loud, because an interviewer who hears 57% quoted as a one-day forecast knows you have confused the stationary distribution with a conditional one. The limitation to add: a two-state chain with fixed probabilities is a toy, and real regime persistence drifts over time.

    Where candidates lose it

    The first wrong answer is 50%, on the grounds that each state has one way in and one way out. The flows are not equal in rate: trending leaks at 0.3, choppy at 0.4, and the slower leak wins more of the time.

    The second loss is writing out eigenvectors of a two-by-two matrix under time pressure. The balance equation, flow out equals flow in, takes one line and is what the interviewer wants to hear.

    What the interviewer asks next

    • Starting from a choppy day, what is the chance that the day after tomorrow is trending?
    • Add a third state, a crash day, that follows a choppy day 5% of the time. How does the method change?
    • How would you estimate these transition probabilities from a year of daily data, and how noisy would they be?
  3. 003A logger stamps every event to the nanosecond, nine decimal places, and whenever the timestamp is missing it writes nine zeros instead. In 100,000 records you find 15 whose fractional part is exactly nine zeros. What is the probability that at least one of those 15 is a filled-in missing value?Conditional probability and BayesHardJump TradingAnonymous interview candidate in · 2022

    Try it first

    First instinct: roughly how many genuine timestamps, out of 100,000, should end in nine zeros by chance?

    Show the worked solution

    Essentially 1; the 15 are gaps. A genuine stamp ends in nine zeros with probability 10^-9, so in 100,000 records you expect 0.0001 such endings. The chance that 15 or more arise genuinely is about 10^-72, which no reasonable prior on missing data can overcome: even a missing rate of one in ten thousand would produce about 10 filled-in endings. So the probability that at least one of the 15 is a filled-in value is 1 to every decimal place you could print.

    What is the question really asking you to compare?

    A shopkeeper who finds 15 notes with the same serial number does not ask what the chance of a coincidence is; she asks which explanation makes 15 identical notes likely. This is a Bayes question in disguise: compare how likely 15 all-zero endings are if nothing is missing against how likely they are if some values are missing, then weight by a prior. The first likelihood is astronomically small; the second is ordinary. The prior would need to be more extreme than anything a real system justifies to change the answer.

    Expected against observed, on a log scale: five orders of magnitude apart0.0000010.00010.011100count of records ending in nine zeros (log scale)expected if genuine: 100,000 x 10^-9 = 0.0001observed: 15genuineseenChance of 15 or more genuine all-zero endings: about 10^-72Even if only one record in 10,000 were missing you would expect 10 filled-in endings.So the 15 are almost all gaps, and at least one of them certainly is.
    Genuine nanosecond stamps should produce 0.0001 all-zero endings in 100,000 records, while 15 were observed, five orders of magnitude more, and the chance of 15 or more genuine ones is about 10^-72, so the observation is explained only by filled-in missing values.

    How do you put a number on the genuine case?

    Each of the 100,000 stamps ends in a specific nine-digit string with probability one in a billion, so the count of genuine all-zero endings is Poisson with mean 0.0001. The probability of exactly 15 is e^(-0.0001) times 0.0001^15 over 15 factorial, which is about 10^-72. You do not need the exact figure in the room; say that 0.0001 to the fifteenth power is 10^-60 before dividing by 15 factorial, and the interviewer has what they need. The point is to show you can set up the count, not to print 72 zeros.

    The relationship
    P(≥1 missing∣15)=1−P(15∣none) P(none)P(15)≈1−10−72 P(none)P(15)≈1P(\geq 1 \text{ missing} \mid 15) = 1 - \frac{P(15 \mid \text{none}) \, P(\text{none})}{P(15)} \approx 1 - \frac{10^{-72} \, P(\text{none})}{P(15)} \approx 1
    P(15 | none)the chance of 15 genuine all-zero endings when no value is missing, Poisson with mean 0.0001
    P(none)your prior that the data set has no missing values at all
    P(15)the overall chance of seeing 15, dominated by the missing-value explanation
    What it says in wordsThe probability that none are missing is the genuine likelihood times its prior, divided by the total, and the genuine likelihood is so small that the result rounds to 1 whatever prior you hold.

    What does the interviewer want to hear about the prior?

    The honest answer is that the question is underspecified: without a prior on how often values go missing, you cannot write a single number. Say that, then show it does not matter: for the posterior to drop even to 99.9% you would need a prior of no missing values more than 10^69 times stronger than the alternative, and no logging system earns that confidence. For contrast, a modest missing rate of one record in ten thousand would give an expected 10 filled-in endings, right where the observed 15 sits. A limitation worth adding: the argument assumes the genuine fractional digits are uniform, which breaks if the clock quantises to microseconds and pads with zeros itself.

    Where candidates lose it

    Candidates reach for the binomial probability of 15 genuine zeros and stop, reporting a tiny number as if it were the answer. The question asks for the probability of a missing value given the data, which needs the comparison with the alternative, not a single likelihood.

    The second loss is freezing because no prior is given. The strong move is to name the missing input, then show that the likelihood ratio is so lopsided that the prior cannot matter. That is what a desk wants: a conclusion that survives the unknown.

    What the interviewer asks next

    • Now the logger stamps to the microsecond, six digits, and pads with three zeros. Does the argument survive?
    • Suppose only 1 record ends in nine zeros. What would you conclude then, and what would you need to know?
    • How would you check the data itself rather than reason about it?

    Asked at Jump Trading, Prop Trading, Anonymous interview candidate in, 2022 (Wall Street Oasis): What is the probability of at least 1 missing value given that we see 15 data points with 0's in the end

  4. 004Two stocks each have 30% annual volatility and a correlation of 0.5. What is the volatility of a basket that holds half of each? What if the correlation were zero?Volatility and correlationWarm upEquity derivativesRisk management

    Try it first

    Before the formula: can a 50/50 basket of two 30% stocks ever be more volatile than 30%?

    Show the worked solution

    25.98% at a correlation of 0.5, and 21.21% at zero. Basket variance is the sum of the two weighted variances plus twice the weighted covariance: 0.25 x 0.09 + 0.25 x 0.09 + 2 x 0.25 x 0.5 x 0.09 = 0.0675, whose square root is 25.98%. At zero correlation the cross term vanishes, leaving 0.045, whose root is 21.21%. The basket is less volatile than either stock because they do not move in step.

    Why is the basket calmer than the stocks inside it?

    Two commuters who each arrive late by a random ten minutes rarely arrive late together; the average of their lateness swings less than either does alone. Volatilities do not add; variances do, and the cross term that joins them is scaled by the correlation, so anything below a correlation of 1 cuts the basket's swing below its parts. At a correlation of 1 the two stocks are one stock and you get 30% back. At minus 1 they cancel exactly and the basket is flat.

    Basket volatility against correlation: it reaches the parts' 30% only at rho = 1-1.0-0.50+0.5+1.010%20%30%correlation between the two stockseach stock alone: 30%rho 0: 21.21%rho 0.5: 25.98%rho 1: 30.00%rho -1: 0%, a perfect hedgeThe gap below 30% is the diversificationit exists only because rho is below 1
    Basket volatility rises with correlation from 0% at minus 1 through 21.21% at zero and 25.98% at 0.5 to the parts' 30% only at a correlation of 1, so the gap below 30% is the diversification and it exists only because the stocks are imperfectly correlated.
    The relationship
    σB2=w2σ2+w2σ2+2w2ρ σ2=2×0.25×0.09 (1+ρ)⇒σB=0.301+ρ2\sigma_B^2 = w^2\sigma^2 + w^2\sigma^2 + 2w^2\rho\,\sigma^2 = 2 \times 0.25 \times 0.09\,(1+\rho) \quad\Rightarrow\quad \sigma_B = 0.30\sqrt{\tfrac{1+\rho}{2}}
    wthe weight of each stock, 0.5
    sigmaeach stock's volatility, 0.30
    rhothe correlation between the two stocks
    sigma_Bthe basket's volatility
    What it says in wordsWith equal weights and equal volatilities, the basket's volatility is the single-stock volatility times the square root of (1 plus rho) over 2.

    How do you do it in your head?

    Use the shortcut in the formula: with two equal stocks the basket volatility is 30% times the square root of (1 plus rho) over 2. At rho 0.5 that is 30% times the root of 0.75, about 0.866, giving 26.0%; at rho 0 it is 30% times the root of 0.5, about 0.707, giving 21.2%. Say the structure first, then the number, so a slip in the arithmetic does not look like a slip in the thinking.

    What is the limitation you should name?

    The formula treats correlation as a fixed number, and it is not. Correlations between stocks tend to rise in a sell-off, which is exactly when a basket holder wants the diversification, so the 25.98% is a fair-weather figure. On a derivatives desk that is why basket options and dispersion trades are priced with a correlation assumption that is marked, stressed and hedged rather than looked up once. Say that the answer depends on the correlation you assume, and that the assumption is the risk.

    Where candidates lose it

    The fast wrong answer is 30%, from averaging the two volatilities. Volatility is a square root, and square roots do not average. Add the variances and the covariance, then take the root.

    The second loss is forgetting the factor of 2 on the cross term. With it, the correlation 0.5 answer is 25.98%; without it, you get 23.72% and an interviewer who knows the number immediately.

    What the interviewer asks next

    • Three stocks at 30% volatility, all pairwise correlations 0.5, equal weights. What is the basket volatility?
    • As the number of equally correlated stocks grows large, where does the basket volatility settle, and why?
    • The basket option is quoted at 24% implied volatility. What correlation is the market pricing?
  5. 005A stock trades at 100. Each day for three days it moves up 10 or down 10, each with probability one half, and interest rates are zero. What is a call struck at 100 and expiring after the third day worth?Option pricing intuitionCore

    Try it first

    Before drawing the tree: how many distinct end prices are there, and how many equally likely paths?

    Show the worked solution

    7.50. Three moves of plus or minus 10 end at 130, 110, 90 or 70, reached by 1, 3, 3 and 1 of the eight equally likely paths. The 100 call pays 30 at 130, 10 at 110 and nothing below. Its value is the average payoff: (1 x 30 + 3 x 10) over 8, which is 60 over 8, or 7.50. With zero rates and symmetric moves, the real-world probabilities are already the pricing probabilities, so no discounting and no adjustment is needed.

    Why is counting paths all the tree needs?

    Think of three coin tosses where you get a sweet for each head. The chance of exactly two heads is not one in four; it is three in eight, because there are three orders in which two heads can arrive. The tree recombines, so the value of an end node is its payoff weighted by how many of the eight paths reach it, and the path counts are the binomial coefficients 1, 3, 3, 1. Nothing else in the problem carries information: the step size fixes the end prices and the counts fix the weights.

    Three up-or-down days: count the paths to each end price, then average the payoffs100901108010012070901101301 path of 8call pays 303 paths of 8call pays 103 paths of 8call pays 01 path of 8call pays 0day 1day 2day 3todayup +10down -10(1 x 30 + 3 x 10) / 8call = 7.50
    From 100, three moves of plus or minus 10 reach 130, 110, 90 or 70 by 1, 3, 3 and 1 of the eight equally likely paths, the call pays 30 and 10 at the two upper nodes and nothing below, and the average payoff (1 x 30 + 3 x 10) over 8 gives a price of 7.50.
    The relationship
    C=18∑k=03(3k) max⁡(100+10(2k−3)−100, 0)=1×30+3×10+3×0+1×08=7.50C = \frac{1}{8}\sum_{k=0}^{3}\binom{3}{k}\,\max(100 + 10(2k-3) - 100,\,0) = \frac{1 \times 30 + 3 \times 10 + 3 \times 0 + 1 \times 0}{8} = 7.50
    kthe number of up days out of three
    C(3, k)the number of paths with k up days: 1, 3, 3, 1
    100 + 10(2k - 3)the end price after k ups and 3 minus k downs
    What it says in wordsThe call is the payoff at each end price, weighted by the share of paths that reach it.

    Where does the risk-neutral machinery go?

    In a general tree you would replace the real probabilities with the risk-neutral ones, chosen so that the stock's expected growth equals the interest rate. Here rates are zero and the moves are symmetric, so the stock already has zero expected drift and the risk-neutral probability is the same one half you were given. Say that out loud: it shows you know the shortcut is a coincidence of the setup, not a rule. If the up move were 10 and the down move 5, one half would no longer price the stock and you would have to solve for the probability that does.

    What sanity checks do you say before the number?

    Two quick ones. The call cannot be worth more than the expected value of the stock above the strike ignoring the max, which is zero here, so the call is worth exactly the expected positive part, and that is what 7.50 is. And put-call parity with zero rates says the 100 put must also be 7.50, which you can confirm from the lower nodes: (3 x 10 + 1 x 30) over 8. Giving the put price unprompted, and showing it matches, is the cheapest way to prove the tree was right.

    Where candidates lose it

    The common error is to treat the four end prices as equally likely, which gives (30 + 10) over 4 = 10. The outer nodes are reached by one path each and the inner ones by three; the weights are 1, 3, 3, 1, not 1, 1, 1, 1.

    The second loss is reaching for a risk-neutral formula and getting lost in it. With zero rates and symmetric moves, the given probabilities already price the stock. Say why, then count.

    What the interviewer asks next

    • Now the up move is 10 and the down move 5. What probability prices the stock, and what is the call worth?
    • Price the 110 call and the 90 put on the same tree.
    • Four days instead of three: what is the 100 call worth, and why does it rise?
  6. 006A trade surveillance system raises an alert on 95% of genuinely suspicious trades and, wrongly, on 2% of normal trades. One trade in a thousand is genuinely suspicious. An alert has just fired on a trade. What is the probability the trade is suspicious?Conditional probability and BayesWarm upCitadelMiami · 2022

    Try it first

    Gut answer before you count anything.

    Show the worked solution

    About 4.5%. Count 100,000 trades. One in a thousand is suspicious, so 100 are, and 95 of those alert. The other 99,900 are normal, and 2% of them, 1,998, alert anyway. Alerts total 2,093, of which 95 are genuine, so the probability that an alerted trade is suspicious is 95 over 2,093, about 4.5%. The 95% hit rate is not the answer; the base rate is what decides it.

    Why does a 95% accurate system give a 4.5% answer?

    A smoke alarm that goes off for 2% of toast is a fine alarm in a house that is never on fire; nearly every ring will be toast. When the thing you are looking for is rare, even a small false-alarm rate applied to the huge normal pile produces more alerts than the true cases produce. Here 2% of 99,900 normal trades is 1,998, twenty times the 95 genuine alerts. The system is not bad; the base rate is low, and that is what the question is testing.

    Count 100,000 trades: the false alarms from the normal pile swamp the true onesAll trades100,0001 in 1,000999 in 1,000Genuinely suspicious100Normal99,90095% alert5% missed2% alert98% quietTrue alerts95Missed5False alerts1,998Quiet97,902Alerts in all: 95 + 1,998 = 2,093. Suspicious given an alert = 95 / 2,093 = 4.5%
    Of 100,000 trades, 100 are suspicious and raise 95 true alerts, while the 99,900 normal trades raise 1,998 false ones, so alerts total 2,093 and a trade that alerts is genuinely suspicious only 4.5% of the time.
    The relationship
    P(S∣A)=P(A∣S) P(S)P(A∣S) P(S)+P(A∣N) P(N)=0.95×0.0010.95×0.001+0.02×0.999=952,093≈4.5%P(S \mid A) = \frac{P(A \mid S)\,P(S)}{P(A \mid S)\,P(S) + P(A \mid N)\,P(N)} = \frac{0.95 \times 0.001}{0.95 \times 0.001 + 0.02 \times 0.999} = \frac{95}{2{,}093} \approx 4.5\%
    S, Na suspicious trade, a normal trade
    Aan alert fires
    P(A | S) = 0.95the hit rate
    P(A | N) = 0.02the false-alarm rate
    P(S) = 0.001the base rate
    What it says in wordsTrue alerts divided by all alerts, where all alerts are the true ones plus the false ones from the normal pile.

    What is the fastest way to say it in the room?

    Do not write Bayes' formula; count a round number of trades. Say: in 100,000 trades, 100 are suspicious and 95 alert; 99,900 are normal and 1,998 alert; 95 over 2,093 is about 4.5%. Three sentences, no algebra, and every number is checkable by the person listening. The odds form is just as quick: prior odds 1 to 999, likelihood ratio 0.95 over 0.02, about 47.5, so posterior odds 47.5 to 999, roughly 1 to 21.

    What does the desk do with a 4.5% answer?

    It decides what the alert is for. A 4.5% hit rate is fine for a filter that sends trades to a human for a second look, and useless for an automatic block, because 95% of blocked trades would be legitimate business. That is the trade-off every surveillance, fraud and risk-limit system lives with: a lower threshold catches more of the 100 but drags in more of the 99,900. Say the limitation too: the 2% and 95% are themselves estimates from past data, and a system tuned on last year's patterns can drift.

    Where candidates lose it

    The whole trap is answering 95%, confusing the probability of an alert given a suspicious trade with the probability of a suspicious trade given an alert. Interviewers ask this precisely because the two sound the same and are twenty times apart.

    The second loss is reaching for the formula and tangling the denominator. Count 100,000 trades and the denominator builds itself: 95 plus 1,998.

    What the interviewer asks next

    • The false-alarm rate is cut to 0.5%. What is the probability now?
    • Two independent systems both alert on the same trade. What is the probability it is suspicious?
    • What base rate would make an alert a coin flip, and what does that tell you about where surveillance is worth running?

    Asked at Citadel, Sales and Trading, Miami, 2022 (Wall Street Oasis): I got a question about Bayes' theorem applied to a practical scenario, which I handled decently

  7. 007A stock goes up 10% one day and down 10% the next, and keeps alternating for 250 trading days. Where does it end relative to its start? And where does a fund that delivers three times the stock's daily move end up?Volatility and correlationCoreVolatility tradingWealth management

    Try it first

    Before multiplying: after one up day and one down day, is the stock back where it started?

    Show the worked solution

    The stock ends at about 28% of its start; the three-times fund at roughly 8 millionths of its start, effectively zero. Each up-and-down pair multiplies the stock by 1.1 x 0.9 = 0.99, and 125 pairs give 0.99^125 = 0.285. The leveraged fund moves 30% each way, so each pair is 1.3 x 0.7 = 0.91, and 0.91^125 is about 7.6e-06. The arithmetic average return is zero in both cases; the compounded return is not.

    Why does a zero average return lose money?

    Take a 100 rupee note to a shop that marks everything up 10% in the morning and discounts 10% in the afternoon. The afternoon discount is taken off a bigger number, so the price ends at 99, not 100. A gain and a loss of the same percentage do not cancel, because the loss acts on the larger base; the pair costs the square of the move, 1% for a 10% swing. Repeat that 125 times and the 1% losses compound to a 72% fall. This is volatility drag: the gap between the average return and the compounded return.

    Alternating moves on a log scale: both drift down, and leverage multiplies the drift10.10.010.00110^-410^-510^-6050100150200250trading daystart = 1.0stock: 0.99^125 = 28% of start3x fund: 0.91^125 = 8 millionthseach up-down pair: 1.1 x 0.9 = 0.99, a loss of 1%with 3x: 1.3 x 0.7 = 0.91, a loss of 9%, nine times worse
    On a log scale both paths step down in straight lines: the stock loses 1% per up-and-down pair and ends at 28% of its start after 250 days, while the three-times fund loses 9% per pair and ends at about 8 millionths of its start, so leverage multiplies the drag by far more than three.
    The relationship
    (1+m)(1−m)=1−m20.99125=0.2850.91125≈7.8×10−6(1+m)(1-m) = 1 - m^2 \qquad 0.99^{125} = 0.285 \qquad 0.91^{125} \approx 7.8 \times 10^{-6}
    mthe daily move, 0.10 for the stock and 0.30 for the three-times fund
    1 - m^2what one up-and-down pair leaves of the value
    125the number of pairs in 250 days
    What it says in wordsEach pair loses the square of the move, and the leveraged fund's loss per pair is nine times the stock's because 0.3 squared is nine times 0.1 squared.

    Why is three times the move so much worse than three times the loss?

    The drag per pair is the square of the move. Tripling the move multiplies the drag by nine, not three: the stock loses 1% per pair, the fund loses 9%. That is why a daily-rebalanced leveraged fund in a choppy, sideways market bleeds even when the underlying ends flat. The general rule you can quote: over many periods the compounded growth rate is roughly the average return minus half the variance, and leverage multiplies the variance by the square of the leverage.

    What would you say to a client who holds the three-times fund?

    That the product tracks three times the daily move, exactly as promised, and that this is not the same as three times the return over a year. A leveraged fund is a tool for a view on the next day or week; held through a sideways year it loses to its own rebalancing. The limitation to state: the alternating path is the worst case for drag, and a strongly trending market can make a leveraged fund return more than three times the underlying. The drag is about path, not just direction.

    Where candidates lose it

    The common answer is that the stock ends flat, because plus 10 and minus 10 seem to cancel. They cancel in arithmetic and not in compounding; the second move acts on a different base.

    The second loss is saying the leveraged fund ends at three times the stock's loss, or at 28% cubed. The right route is per pair: 1.3 x 0.7 = 0.91, then raise to the 125th power. The drag scales with the square of the leverage.

    What the interviewer asks next

    • Make the daily move 1% instead of 10%. Where does the stock end after 250 days?
    • Over a year with 16% annual volatility and zero average daily return, roughly what is the compounded return?
    • Why do leveraged funds rebalance daily, and what would change if they rebalanced monthly?
  8. 008You roll a fair die again and again, adding each face to a pot. But if you roll a 1, the whole pot is wiped out and the game ends. You may stop and bank the pot at any time. When should you stop, and why?Expected value and optimal stoppingHardQuant tradingProp trading firms

    Try it first

    First instinct: with 15 in the pot, should you roll once more?

    Show the worked solution

    Roll while the pot is below 20 and stop once it reaches 20 or more. One more roll loses the pot with probability 1/6 and otherwise adds a face of 2, 3, 4, 5 or 6, which sum to 20. The expected change is (20 minus pot) over 6: +3.33 from an empty pot, +0.83 at 15, zero at 20 and negative beyond. The number of rolls so far is irrelevant; only the pot matters. Played this way, the expected bank from an empty pot is about 8.14.

    What does one more roll actually buy you?

    Think of a street game where you can keep picking envelopes that each add a few rupees to your winnings, but one envelope in six says lose everything. Whether to pick again depends on how much you already hold, not on how many envelopes you have opened. The expected change from one more roll is the chance of adding, five sixths, times the average addition, 4, minus the chance of ruin, one sixth, times the pot you would lose. That is (20 minus pot) over 6, and it is positive exactly while the pot is under 20.

    One more roll is worth (20 minus pot) / 6: positive below a pot of 20, negative above0102030pot already held-2+2+4expected gain from one more rollpot = 20: gain is exactly zeroempty pot: +3.33pot 30: -1.67ROLL AGAINBANK ITone more roll adds valueone more roll loses value
    The expected gain from one more roll falls in a straight line from +3.33 with an empty pot to zero at a pot of 20 and to -1.67 at 30, so rolling adds value on the left of 20 and destroys it on the right, and the stopping rule is simply to bank at 20 or above.
    The relationship
    E[Δ∣pot=p]=56×4−16×p=20−p6⇒stop when p≥20E[\Delta \mid \text{pot}=p] = \frac{5}{6}\times 4 - \frac{1}{6}\times p = \frac{20 - p}{6} \quad\Rightarrow\quad \text{stop when } p \ge 20
    pthe pot already held
    5/6 x 4the chance of surviving the roll times the average of the faces 2 to 6
    p/6the pot at risk, times the one-in-six chance of a 1
    What it says in wordsOne more roll is worth the expected addition minus the expected loss, and the two balance when the pot is 20.

    Why is the one-step rule the whole answer here?

    In many stopping problems you cannot trust a one-step look: a roll that loses value today might open a bigger gain tomorrow. Here the gain from one more roll only falls as the pot grows, so once rolling stops being worth it, it never becomes worth it again, and the one-step rule is optimal. You can check it by valuing the whole game for each stopping threshold: stopping at 20 (or 21, which is equivalent because the roll at exactly 20 is worth zero) gives an expected bank of about 8.14, stopping at 15 gives 7.85 and at 25 gives 8.00. Say the word monotone if you know it; say the reason either way.

    What would a desk add to the textbook answer?

    That the rule maximises expected value and nothing else. A player who cannot afford to lose the pot, or who is paid on a target rather than on the average, would stop earlier, and a trader who sizes positions knows that expected value is only the first thing to check. The limitation is the same in the puzzle and on the desk: expected value is right for a game you can play many times, and this game is played once.

    Where candidates lose it

    Candidates who answer by feel either stop far too early, because a one-in-six wipe-out sounds frightening, or never stop, because the pot keeps growing. The question is asking for the point where the two forces balance, and that point is a number: 20.

    The second loss is using 3.5 as the average face. The faces that keep you in the game are 2 to 6, averaging 4, and their total is 20. Using 3.5 gives a threshold of 21 and shows the ruin branch has been counted twice.

    What the interviewer asks next

    • Now rolling a 1 costs you only half the pot. What is the new stopping rule?
    • What is the expected bank from an empty pot under the optimal rule, and how would you compute it?
    • Suppose you must pay 1 for every roll. Does the threshold go up or down, and by how much?
  9. 009Three players each hold one hidden card from a standard deck, ace counting 1 up to king counting 13. You can see only your own card, a 10. You must make a two-way market on the total of all three cards. The player on your left lifts your offer, you requote, and he lifts again. What do you do now?Market makingHardOptiverChicago · 2025

    Try it first

    Before the second lift: what is a fair value for the total when all you know is your own 10?

    Show the worked solution

    Raise the market and widen it; do not sell a third time near the old level. With only your 10 known, the fair total is 24 and the two unknowns give a standard deviation of about 5.3. A player who sees his own card and lifts your offer is telling you his card is high: if it is 8 or more the fair total is 27.5; after a second lift, 11 or more, it is 29. You are short 2 at an average near 27 against a value near 29. Quote something like 28 at 32.

    What does a lift tell you that the cards do not?

    If you are selling a second-hand bike and the first viewer pays your asking price without haggling, you have probably priced it low; if he immediately asks to buy a second one, you certainly have. A counterparty who can see something you cannot and keeps buying is telling you the value is higher than your offer, and every trade you do with him before repricing is a trade you will regret. This is adverse selection, and a market-making game exists to see whether you notice it in time.

    Each lift is information: the fair value climbs and the quote climbs and widens with itQuote 1two hidden cards, 7 eachfair 2423 at 25they LIFT at 25position: short 1 at 25Quote 2their card likely 8+, avg 10.5fair 27.526 at 29they LIFT at 29position: short 2, avg 27Quote 3their card likely 11+, avg 12fair 2928 at 32wider: 4 not 2do not keep selling at 25Your card: 10. Each hidden card: 1 to 13, mean 7, so the prior total is 10 + 7 + 7 = 24, give or take 5.3.A buyer who keeps lifting is telling you their card is high. Reprice before you sell a third time.Short 2 at an average of 27 against a fair value near 29: expected loss about 4 so far. Stop the bleed, do not chase it.
    With your 10 the prior fair total is 24; one lift suggests the buyer's card is 8 or more and moves the fair value to 27.5; a second lift suggests 11 or more and moves it to 29, so the quote climbs from 23 at 25 to 28 at 32 and widens from 2 to 4 while you sit short 2.
    The relationship
    E[total]=10+E[L]+7,E[L∣L≥8]=10.5,E[L∣L≥11]=12  ⇒  24→27.5→29E[\text{total}] = 10 + E[L] + 7,\qquad E[L \mid L \ge 8] = 10.5,\quad E[L \mid L \ge 11] = 12 \;\Rightarrow\; 24 \to 27.5 \to 29
    Lthe card held by the player who keeps lifting
    7the mean of the third player's card, still unknown and uniform on 1 to 13
    L >= 8, L >= 11the rough information in a first and a second lift: his card is above what your offer implied
    What it says in wordsEach lift raises your estimate of the lifter's card, and the fair total moves by exactly that amount.

    How much do you move, and how much do you widen?

    Move at least as far as the information says and widen because your uncertainty about his behaviour has grown. After one lift a quote of 26 at 29 sits around the new estimate of 27.5; after two lifts 28 at 32 sits around 29 and is twice as wide, because the next lift would mean his card is 12 or 13 and the total near 30 or more. The width is not a penalty on him; it is the price of your own blindness. A standard deviation of 5.3 on the total from the two unknown cards is the natural scale for the width before any lifts.

    What about the position you already have?

    You are short 2 at an average near 27 and the value is near 29, so you are losing about 4 on paper. Do not try to earn it back by selling more at a worse price, and do not flip to buying from the other player at any cost; skew your quote up so the next trade is more likely to reduce the short than add to it. Say your position out loud when the interviewer asks; the game checks whether you can hold the fair value, the quote and the inventory in your head at once. The limitation to state: the thresholds 8 and 11 are a rough model of his behaviour, and a player who bluffs changes the inference.

    Where candidates lose it

    The common failure is to keep quoting 23 at 25 after the first lift, and to sell a third unit at 25 after the second, because the cards have not changed. Your information has changed. Two lifts from a player who sees his own card are worth more than the deck statistics.

    The second loss is overreacting: moving the quote to 35 at 40 after one lift. He may hold a 9. Move to what the evidence supports, widen for what it does not, and keep your position in mind.

    What the interviewer asks next

    • Now the third player hits your bid at 28. How do you reprice?
    • What if the lifter can see your card as well as his own?
    • Make a market on the product of the three cards instead of the sum. What changes about the width?

    Asked at Optiver, Quantitative Research, Chicago, 2025 (Wall Street Oasis): The next round was a poker style market making game as well as a separate behavioural interview

  10. 010In some stock, the 99-strike call trades at 5.60 and the 101-strike call at 4.70, same expiry. Estimate the price of a digital option that pays 1 if the stock finishes above 100 at that expiry.Option payoffs and no-arbitrageCoreExotics tradingStructured products

    Try it first

    Before any arithmetic: which combination of the two calls has a payoff that looks most like a step at 100?

    Show the worked solution

    About 0.45. Buying the 99 call and selling the 101 call pays 0 below 99, 2 above 101 and a straight ramp between. Divide that by the width of 2 and the payoff is 0 below 99, 1 above 101 and a ramp through 100: a digital with its edge smoothed over two points. Its cost is (5.60 minus 4.70) over 2, which is 0.45. The narrower the spread, the closer the ramp sits to the step, and the price converges to the digital.

    Why does a call spread stand in for a digital?

    A light switch is a step: off or on. A dimmer that goes from fully off to fully on over a tiny turn of the knob is, for every practical purpose, the same switch. A call spread over its width is a dimmer: it ramps from 0 to 1 across the two strikes, and as the strikes close in on 100 the ramp becomes the step. So the digital is the limit of a scaled call spread, and a traded call spread gives you a price for it without any model.

    A call spread over its width is a ramp; the digital is a step; shrink the width and they meetCall spread 99 / 101, scaled by 1/2959910010110510(call 99 - call 101) / 2rampstock price at expiryDigital struck at 100959910010110510pays 1 if S > 100stepstock price at expiryprice of the ramp = (5.60 - 4.70) / 2 = 0.45, so the digital is worth about 0.45
    The 99 to 101 call spread divided by its width of 2 pays 0 below 99, ramps to 1 at 101 and crosses the digital's step exactly at 100, so the two payoffs differ only inside the narrow band between the strikes and the spread's price, (5.60 minus 4.70) over 2, gives a digital value of 0.45.
    The relationship
    D(100)≈C(99)−C(101)101−99=5.60−4.702=0.45D(K)=−∂C∂KD(100) \approx \frac{C(99) - C(101)}{101 - 99} = \frac{5.60 - 4.70}{2} = 0.45 \qquad D(K) = -\frac{\partial C}{\partial K}
    C(K)the price of a call struck at K
    D(K)the price of a digital paying 1 above K
    (C(99) - C(101)) / 2the slope of the call price in strike, estimated across 100
    What it says in wordsThe digital is minus the slope of the call price with respect to strike, and a centred call spread measures that slope.

    Is 0.45 the digital's price or an approximation, and which way is it off?

    It is an approximation to the slope at 100 taken from two points either side. Because the spread is centred on 100, the first-order error cancels and what remains is small, of the order of the curvature of the call price between 99 and 101. If the digital were struck at 99 instead, the same spread would overstate it, because the call price is convex in strike and the ramp sits above the step on that side. On a desk you would quote the digital from the tightest spread the market will show you, and hedge it with that spread, so the approximation is also the hedge.

    What does 0.45 say about the market, and what is the limitation?

    A digital paying 1 above 100 at 0.45, with rates near zero, means the pricing probability of finishing above 100 is about 45%, slightly below one half. That is a risk-neutral probability, not a forecast, and it is pulled down by the skew: with a steeper put skew, out-of-the-money calls are cheaper in volatility terms and the slope in strike is steeper, which moves the digital. Say that a flat-volatility formula would miss this, and that the call spread picks the skew up automatically because it uses the two traded prices.

    Where candidates lose it

    The common error is to take the difference of the two call prices, 0.90, and present it as the digital. That is the price of a spread that pays 2 above 101, not 1. Divide by the width.

    The second loss is reaching for a lognormal formula with a guessed volatility. The question gives you two traded prices precisely so you can price the digital without a model; use them.

    What the interviewer asks next

    • The 99.5 and 100.5 calls are 5.37 and 4.93. What does that pair say about the digital, and why might it differ from 0.45?
    • How would you hedge a short digital you sold at 0.45, and what goes wrong near expiry?
    • Price a digital that pays 1 if the stock finishes below 100.
← PreviousPage 1 of 10
  1. 1
  2. 2
  3. …
  4. 10
Next →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.