Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant puzzles, solved step by step

Puzzles
100
Traced to a firm
71
Topics
12
Hard
30
Topic
All topicsLogic and algorithmic reasoning10Conditional probability and Bayes7Counting and combinatorics8Continuous and geometric probability9Correlation, regression and linear algebra9Market making, betting and sizing9Expected value and optimal stopping9Statistics and estimation9Pricing, options and index maths7Games and strategic reasoning8Markov chains and random walks7Mental maths and number sense8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 1–10 of 50 · filtered from 100Clear filters
  1. 001Four people must cross a narrow bridge at night with one torch. At most two can be on the bridge at once, anyone crossing must carry the torch, and a pair walks at the slower person's pace. They take 1, 2, 5 and 10 minutes. What is the shortest total time to get everyone across?Logic and algorithmic reasoningCoreBelvedere TradingChicago · 2021

    Try it first

    Before you plan it: what is the fastest time?

    Show the worked solution

    17 minutes. Send 1 and 2 over (2 minutes), 1 comes back (1), 5 and 10 cross together (10), 2 comes back (2), and 1 and 2 cross again (2). The obvious plan, where the fastest person escorts each of the others, takes 19. The saving comes from putting the two slowest walkers on the bridge at the same time.

    Why does the obvious plan lose two minutes?

    Think of two slow parcels going by the same courier. If each travels on its own trip you pay for both trips; if they share a van, you pay for the slower one only. Every crossing costs the slower walker's time, so a slow person paired with a fast one wastes the fast one, and two slow people paired together waste nothing. Escorting with the fastest person pays 10 and then 5 as separate crossings. Pairing 5 with 10 pays 10 once and the 5 minutes are free.

    Send the two slowest together and the 5 minutes hide inside the 10Pair the slowest21 and 2 over11 back105 and 10 over22 back21 and 2 over17 minFastest escorts all101 and 10 over11 back51 and 5 over11 back21 and 2 over19 min0510151719minutes
    Pairing the 5 and 10 minute walkers on one crossing finishes in 17 minutes, while letting the 1 minute walker escort everyone pays for the 10 and the 5 separately and finishes in 19.

    What is the price of pairing the slow two?

    Somebody has to bring the torch back after the slow pair crosses, and it must not be one of them. So the plan first ferries two fast people over, leaves one on the far side to carry the torch back later, and spends the 2 minute walker's return trip to buy the 5 minute saving. The trade is 1 + 2 extra minutes of shuttling against 5 minutes saved on the slow side, a net gain of 2. With different speeds the trade can flip, which is the real content of the puzzle.

    The relationship
    escort: 2a+b+c+dpair: a+3b+dpair wins when 2b<a+c\text{escort: } 2a + b + c + d \qquad \text{pair: } a + 3b + d \qquad \text{pair wins when } 2b < a + c
    a, bthe two fastest times, here 1 and 2
    c, dthe two slowest times, here 5 and 10
    What it says in wordsPairing the slow two is better exactly when twice the second fastest time is less than the fastest plus the second slowest.

    How do you convince the interviewer 17 cannot be beaten?

    There must be at least five crossings, three over and two back, because each trip over moves at most two people and someone must return the torch. The 10 minute walker costs 10 on whatever crossing carries them. If 5 and 10 cross separately you already spend 15 on those two trips, and the three remaining crossings cost at least 1 + 1 + 2, which is 19; if they cross together, the best you can do with the other four crossings is 2 + 1 + 2 + 2. A brute force over every schedule gives the same minimum, 17 minutes.

    Where candidates lose it

    The strong candidate's trap is a fast answer of 19. Letting the quickest person run every errand feels efficient, and it is the right instinct for returning the torch, but it is the wrong instinct for the slow walkers.

    The second loss is getting 17 by trial and error and then being unable to say why. State the principle, that the slow pair shares one crossing, and give the rule for when it wins: when twice the second fastest time is below the fastest plus the second slowest.

    What the interviewer asks next

    • What if the times are 1, 4, 5 and 10?
    • Six people with times 1, 2, 5, 10, 20 and 25: what is the plan?
    • Write the general algorithm for n people and say its running time.

    Asked at Belvedere Trading, Trading, Chicago, 2021 (Wall Street Oasis): crossing the bridge in the shortest amount of time with one flashlight brainteaser

  2. 002You are flying to a city where it rains on 25% of days. You phone three friends who live there. Each tells the truth with probability 2/3, independently of the others, and all three say it is raining. What is the probability that it is actually raining?Conditional probability and BayesCoreJane StreetNew York · 2025

    Try it first

    Pick your answer before working it.

    Show the worked solution

    8/11, about 72.7%. If it is raining, all three say yes with probability (2/3)^3 = 8/27. If it is dry, all three must be lying, (1/3)^3 = 1/27. Weight each by how often it happens: 1/4 x 8/27 against 3/4 x 1/27, which is 8 parts to 3. Three agreeing witnesses move a 25% prior a long way, but not to certainty.

    Why is the answer not simply 8/9?

    Picture a clinic where a test is quite reliable but the illness is uncommon. A positive result makes the illness more likely, but how much more depends on how rare it was to begin with. The friends' agreement tells you how much more likely rain makes their answer than dry does, eight times, but it does not erase the fact that dry days are three times as common. 8/9 is the answer you get if rain and dry start level. Here they do not.

    Three yeses: compare the two shaded areas, not the two stripsall 3 say yes8/27not all yes19/27all 3 lie and say yes: 1/27not all yes26/27Rain, 1/4Dry, 3/4Shaded areasRain: 1/4 x 8/27= 8/108Dry: 3/4 x 1/27= 3/108P(rain | 3 yes)8/11 = 72.7%rain: 8 partsdry: 3 partsInside the shaded region only
    Rain covers a quarter of days and all three friends say yes on 8/27 of those, while dry days cover three quarters and all three lie on only 1/27 of them, so the shaded areas stand 8 to 3 and the chance of rain given three yeses is 8/11, about 72.7%.

    How do you set it up so the arithmetic stays small?

    Use odds rather than probabilities. Posterior odds are prior odds times the likelihood ratioHow many times more likely the evidence is if the hypothesis is true than if it is false., and both are easy numbers here. Prior odds of rain are 1 to 3. The likelihood ratio of three yeses is (2/3)^3 over (1/3)^3, which is 2 cubed, 8. So the posterior odds are 8 to 3, and the probability is 8 over 8 plus 3, 8/11. Each additional agreeing friend would double the odds again.

    The relationship
    P(R∣YYY)P(D∣YYY)=P(R)P(D)⋅(2/3)3(1/3)3=13⋅8=83  ⇒  P(R∣YYY)=811\frac{P(R\mid YYY)}{P(D\mid YYY)} = \frac{P(R)}{P(D)}\cdot\frac{(2/3)^3}{(1/3)^3} = \frac{1}{3}\cdot 8 = \frac{8}{3} \;\Rightarrow\; P(R\mid YYY) = \frac{8}{11}
    R, Drain and dry
    YYYall three friends say yes
    (2/3)^3 and (1/3)^3the chance of three yeses when it rains, and when it is dry
    What it says in wordsMultiply the prior odds by how much more likely the evidence is under rain, then turn the odds back into a probability.

    What assumption is doing the work, and should you say it?

    The calculation needs the friends to lie independently. If they could be coordinating a joke, three yeses are really one piece of evidence, and the answer falls back towards the one-friend figure of 2/5. Say the independence assumption out loud, then give 8/11. Interviewers often follow up by making one friend unreliable or by letting them talk to each other.

    Where candidates lose it

    The most common wrong answer is 8/9: the candidate compares the chance of three truths with the chance of three lies and forgets the weather's own odds. The rain prior is a quarter, and leaving it out quietly assumes it is a coin flip.

    The second loss is writing out a full Bayes formula with 27ths and 108ths and losing the thread under time pressure. Odds times likelihood ratio gets 8 to 3 in two lines and is easier to check out loud.

    What the interviewer asks next

    • What if only two of the three friends say yes?
    • How many agreeing friends would you need before you were 95% sure it is raining?
    • What changes if the friends can talk to each other before answering?

    Asked at Jane Street, Generalist, New York, 2025 (Wall Street Oasis): There was a question about the probability of rain the next day that relied on a very in depth understanding of bayes theorem

  3. 007You roll a fair six-sided die six times. What is the expected number of different faces that appear?Expected value and optimal stoppingCoreQuant tradingProp trading firms

    Try it first

    Your estimate before any working?

    Show the worked solution

    About 3.99. Give each face an indicator that equals 1 if that face appears at least once. A given face is missed on all six rolls with probability (5/6) to the 6th, about 0.335, so it appears with probability 0.665. The expected count of distinct faces is the sum of the six indicators' expectations: 6 x 0.665 = 3.99. No case listing is needed.

    Why not list the cases?

    You could work out the chance of exactly one, two, up to six distinct faces and average them. It works, but it needs Stirling numbersCounts of the ways to split a set of items into a given number of non-empty groups; they appear when counting surjections. or a lot of careful counting, and it is easy to slip. Linearity of expectation lets you ignore how the faces interact: the expected total of several indicators is the sum of their expectations, whether or not they are independent. Here the six indicators are clearly dependent, since seeing many faces leaves fewer rolls for the others, and it does not matter at all.

    Six indicators, one per face, each switched on 66.5% of the time66.5%66.5%66.5%66.5%66.5%66.5%Each bar: 1 - (5/6) to the 6th, the chance that face shows up at least onceSum of the six bars = 6 x 0.6651 = 3.99Exact spread of distinct faces122%323%450%523%62%mean 3.99distinct faces in six rolls
    Each of the six faces appears at least once with probability 66.5%, so the expected number of distinct faces is six times that, 3.99; the exact distribution peaks at four distinct faces and has all six only 1.5% of the time.

    How does the indicator trick work step by step?

    Think of a teacher counting how many of six friends turn up to a party. Instead of listing every guest list, she asks for each friend separately how likely that friend is to come, then adds. Write the count as I1 + I2 + ... + I6, where I_k is 1 if face k appears; take expectations; each E[I_k] is just the probability face k appears. Face k is missed on one roll with probability 5/6, on all six with (5/6) to the 6th, 0.335. So each indicator averages 0.665 and the total averages 3.99.

    The relationship
    E[D]=∑k=16P(face k appears)=6(1−(56)6)≈6×0.665=3.99E[D] = \sum_{k=1}^{6} P(\text{face } k \text{ appears}) = 6\left(1 - \left(\tfrac{5}{6}\right)^6\right) \approx 6 \times 0.665 = 3.99
    Dthe number of distinct faces seen in six rolls
    (5/6)^6the chance a given face never appears in six rolls
    What it says in wordsThe expected number of distinct faces is six times the chance that any one face appears.

    Where does this pattern reappear?

    The same shape answers how many distinct birthdays a group of n people has, how many of n hash buckets get used, and how many different stocks a random sample of trades touches. For n faces and n rolls, the expected share of faces seen is 1 - (1 - 1/n) to the n, which tends to 1 - 1/e, about 63.2%, as n grows. Six faces give 66.5%, already close. The exact enumeration of all 46,656 rolls gives a mean of 3.9906, the same number.

    Where candidates lose it

    The instinctive answer is 6, or something close to it, because six rolls over six faces feels like one of each. In reality repeats are the norm, and all six different faces happen in under 2% of runs.

    The costlier trap is starting to enumerate cases under time pressure. Candidates who try to list exactly four distinct faces lose minutes. Say indicator variables and linearity in the first sentence.

    What the interviewer asks next

    • What is the expected number of faces that appear exactly once?
    • How many rolls do you need, on average, to see all six faces?
    • What is the variance of the number of distinct faces?
  4. 008Z1 and Z2 are independent standard normal random variables. What are the mean and variance of Z1 squared + Z2 squared, and what is the probability that it exceeds 2?Statistics and estimationCoreQuant researchQuant trading

    Try it first

    What is P(Z1 squared + Z2 squared > 2)?

    Show the worked solution

    Mean 2, variance 4, and the probability of exceeding 2 is e to the -1, about 36.8%. Each Z squared has mean 1 and variance 2, so the sum has mean 2 and variance 4. The sum is a chi-squared with two degrees of freedom, which happens to be exactly an exponential with mean 2. Its tail beyond t is e to the -t/2, so beyond 2 it is e to the -1.

    How do you get the mean and variance without the distribution?

    E[Z squared] is the variance of Z, which is 1. For the variance of Z squared you need the fourth moment: E[Z to the 4] is 3 for a standard normal. So Var(Z squared) = 3 - 1 = 2, and because Z1 and Z2 are independent the variances add, giving a mean of 2 and a variance of 4. Say the fourth moment of 3 aloud; it is the number interviewers check you know, and it is why normal kurtosis is quoted as 3.

    Two squared normals add up to an exponential with mean 2024680.250.5mean = 263.2%P(above 2) = e to the -1 = 36.8%value of Z1 squared + Z2 squaredDensity: (1/2) e to the -x/2Mean 2, variance 4Tail: P(X > t) = e to the -t/2
    The sum of two squared standard normals has density one half times e to the minus x over 2, an exponential with mean 2 and variance 4, and the area beyond 2 is exactly e to the minus 1, about 36.8%.

    Why is this particular sum exponential?

    Think of a dart thrown at a board where both the horizontal and vertical errors are independent standard normals. The joint density depends only on the distance from the centre, so the dart's direction is uniform and all the information is in the radiusThe distance of the point (Z1, Z2) from the origin, the square root of Z1 squared plus Z2 squared.. Switching to polar coordinates, the chance that the squared distance exceeds t is e to the -t/2, which is the tail of an exponential with mean 2. At t = 2 the answer is e to the -1.

    The relationship
    P(Z12+Z22>t)=∫02π ⁣ ⁣∫t∞12πe−r2/2 r dr dθ=e−t/2P( ⋅>2)=e−1P(Z_1^2+Z_2^2 > t) = \int_0^{2\pi}\!\!\int_{\sqrt t}^{\infty} \frac{1}{2\pi} e^{-r^2/2}\, r\,dr\,d\theta = e^{-t/2} \qquad P(\,\cdot > 2) = e^{-1}
    rthe distance of (Z1, Z2) from the origin
    \frac{1}{2\pi} e^{-r^2/2}the joint density of two independent standard normals
    e^{-t/2}the tail of an exponential distribution with mean 2
    What it says in wordsIn polar coordinates the angle integrates out and the radius gives an exponential tail.

    Why is 50% the tempting wrong answer?

    Because 2 is the mean and people read the mean as the middle. For a right-skewed distribution the mean sits above the median, so less than half the mass lies beyond it; here the median is 2 ln 2, about 1.39. A simulation of 200,000 pairs gives a mean of 1.997, a variance of 3.96 and a tail share of 0.368, matching the exact results. The same polar trick is what powers the Box-Muller method for generating normal random numbers.

    Where candidates lose it

    The common wrong answer is 50%, from treating the mean as the median. Chi-squared variables are skewed to the right, and the skew is largest with few degrees of freedom.

    The second trap is the variance. Candidates who say the variance of Z squared is 1 have confused it with the variance of Z. The fourth moment of 3 is the step, and missing it gives a variance of 2 for the sum instead of 4.

    What the interviewer asks next

    • What is the distribution of the square root of Z1 squared + Z2 squared?
    • How would you use this to generate normal random numbers from uniforms?
    • What are the mean and variance of a chi-squared with k degrees of freedom?
  5. 009A stock trades at 100 and in one period will be either 120 or 80. Interest rates are zero. Price a call option struck at 100 by building a portfolio of shares and borrowing that copies it, and explain why the real-world probability of the up move does not appear in the price.Pricing, options and index mathsCoreOptions market makingQuant trading

    Try it first

    If you believe the stock goes up with probability 90%, what is the call worth?

    Show the worked solution

    The call is worth 10. It pays 20 if the stock goes to 120 and 0 at 80. Half a share pays 60 or 40, so half a share with a loan of 40 pays 20 or 0, exactly the call. That portfolio costs 50 - 40 = 10 today. If the call traded at any other price, you could buy the cheap one and sell the dear one for a riskless profit, so no probability is needed.

    How do you build the copy?

    Match the swing first. The call's payoff moves by 20 between the two states while the stock moves by 40, so the copy needs 20/40 = 0.5 of a share: that ratio is the option's deltaHow much an option's value changes for a one-unit change in the underlying price; here, the number of shares that copies the option.. Half a share is worth 60 or 40 at the end, which is 40 more than the call in both states. Borrow 40 today, repay 40 at the end with zero interest, and the copy pays exactly 20 or 0.

    Copy the payoff with shares and borrowing, and price the copyStock 100Call ?Stock 120Call pays 20Stock 80Call pays 0updownDelta = (20 - 0) / (120 - 80) = 0.5 shareThe copy: 0.5 share, borrow 40TodayUp (120)Down (80)0.5 share506040Loan-40-40-40Total10200Matches the call in both statesCall = cost of the copy = 10No probability of up or down was used
    A call struck at 100 on a stock that moves to 120 or 80 is copied by half a share and a loan of 40, which pays 20 or 0 exactly as the call does and costs 10 today, so the call is worth 10 with no probability used.

    Why does the chance of the up move not matter?

    Think of a shop selling a bundle of two items that you can also buy separately. The bundle's price is pinned by the parts, whatever you think about how useful the items are. The call is a bundle of half a share and a loan; the share price already reflects everyone's views about the up move, so the option inherits them and adds none of its own. A 90% view is a reason to hold the stock itself, not a reason to pay more for the call than its parts cost.

    The relationship
    Δ=Cu−CdSu−Sd=20−0120−80=0.5C0=ΔS0−B=50−40=10=q Cu+(1−q) Cd,  q=S0−SdSu−Sd=0.5\Delta = \frac{C_u - C_d}{S_u - S_d} = \frac{20-0}{120-80} = 0.5 \qquad C_0 = \Delta S_0 - B = 50 - 40 = 10 = q\,C_u + (1-q)\,C_d,\; q = \frac{S_0 - S_d}{S_u - S_d} = 0.5
    \Deltashares held in the copy
    Bthe amount borrowed, 0.5 x 80 - 0 = 40
    qthe risk-neutral weight on the up state, fixed by the prices, not by beliefs
    What it says in wordsThe copy's cost gives the price, and the same price is an average of the payoffs using weights set by today's stock price.

    What would you do if the call traded at 12?

    Sell the dear thing and buy the cheap one. Sell the call for 12, buy half a share for 50 and borrow 40, a net cash inflow of 2 today; at the end the portfolio pays exactly what you owe on the call in either state. The 2 is kept whatever happens. The weight q = 0.5 that reproduces the price is called the risk-neutral probability, but it is a pricing weight backed out of the stock price, not a forecast. Say that distinction; interviewers listen for it.

    Where candidates lose it

    The trap is pricing the call as an expected payoff under your own view: 90% of 20 is 18. That price can be arbitraged against the stock, so nobody could trade it for long, and the interviewer wants to hear that the copying portfolio pins the price.

    The second loss is getting 10 by assuming a 50% chance. The number is right by coincidence of the symmetric tree; ask yourself what happens with an up move to 130, and the risk-neutral weight changes to 1/2.5 = 0.4.

    What the interviewer asks next

    • Price the put struck at 100 and check put-call parity.
    • What changes if interest rates are 5% for the period?
    • The stock can go to 130 or 80 instead. Price the call again.
  6. 011An urn starts with one red ball and one blue ball. Each turn you draw a ball at random, put it back, and add another ball of the same colour. After ten draws, what is the probability that exactly five of the draws were red?Markov chains and random walksCoreQuant researchQuant trading

    Try it first

    Your first instinct for P(exactly five reds)?

    Show the worked solution

    1/11, about 9.1%. Any particular sequence with five reds and five blues has probability 5! x 5! / 11!, whatever the order, because the numerators just count up the reds and blues separately. There are 10 choose 5 = 252 such sequences, and 252 x 5! x 5! / 11! = 1/11. The same algebra gives 1/11 for every count from 0 to 10.

    Why is this not ten coin flips?

    Think of a new café and its first customers. If the first few visitors like it and bring friends, it fills up; if they do not, it stays empty. Early luck compounds. In this urn each red draw adds a red ball, so it raises the chance of the next red: the draws reinforce each other, and runs that start lopsided tend to stay lopsided. That spreads the count of reds far more widely than independent coin flips, which bunch around five.

    Ten draws from the urn: every count of reds from 0 to 10 is equally likely012345678910coin flips: 24.6% at five1/11eachnumber of red draws out of tenStart: 1 red, 1 blueDraw, return, add one of that colour
    After ten draws from an urn that starts with one red and one blue ball and adds a ball of the drawn colour each time, every count of reds from 0 to 10 has probability 1/11, a flat line, while ten fair coin flips would pile up at five with 24.6%.

    Why does the order of the draws not matter?

    Write out one sequence, say five reds then five blues. The chances are 1/2, 2/3, 3/4, 4/5, 5/6 for the reds, then 1/7, 2/8, 3/9, 4/10, 5/11 for the blues. Shuffle the order and the denominators are still 2 up to 11, while the red numerators still run 1 to 5 and the blue ones 1 to 5, so every arrangement of five reds and five blues has the same probability, 5! x 5! / 11!. Draws whose joint probability ignores order are called exchangeableA sequence of random variables whose joint distribution does not change when you reorder them, even though they need not be independent., and that is the property doing the work.

    The relationship
    P(k reds in n)=(nk)k! (n−k)!(n+1)!=1n+1n=10:  252⋅120⋅12039 916 800=111P(k \text{ reds in } n) = \binom{n}{k}\frac{k!\,(n-k)!}{(n+1)!} = \frac{1}{n+1} \qquad n = 10:\; \frac{252 \cdot 120 \cdot 120}{39\,916\,800} = \frac{1}{11}
    \binom{n}{k}the number of orders in which k reds can appear
    k!(n-k)!/(n+1)!the probability of any one such order
    What it says in wordsThe number of orders times the probability of each order is always one over n plus one.

    Is there a picture that makes 1/11 obvious?

    Yes. This urn behaves exactly as if nature first picked a hidden red probability p uniformly between 0 and 1, and then flipped ten independent coins with that p. Averaging a binomial over a uniform p gives every count the same weight, 1/(n + 1). The urn is also a small model of momentum and of market share: an early lead makes further gains more likely, so the final split is highly uncertain, even though each step looks like a fair draw.

    Where candidates lose it

    The fast wrong answer is 252/1024, about 24.6%, from treating the draws as independent coin flips. The reinforcement is the whole question, and it spreads the outcomes out rather than pulling them to the middle.

    The second trap is trying to add up paths through a tree of ten levels. Spot that every order of a given mix has the same probability, and the problem collapses to one line.

    What the interviewer asks next

    • What is the probability the eleventh draw is red, given that five of the first ten were red?
    • What if the urn starts with two red and one blue ball?
    • As the number of draws grows, what does the fraction of red balls converge to?
  7. 013Users join a server at times 1, 2, 4, 5 and 7 and leave at times 7, 3, 8, 9 and 10 respectively. A leave at the same moment as a join is processed first. What is the maximum number of users online at once, and how would you compute it efficiently for a million users?Logic and algorithmic reasoningCoreTwo SigmaNew York · 2025

    Try it first

    What is the peak number of users online together?

    Show the worked solution

    The peak is 3 users. Turn every join into a +1 event and every leave into a -1 event, sort all ten events by time with leaves before joins at equal times, and keep a running total. It goes 1, 2, 1, 2, 3, then at time 7 down to 2 and back to 3, then 2, 1, 0. Sorting costs n log n, and the sweep itself is linear.

    Why not check every moment in time?

    A shopkeeper who wants to know the busiest moment of the day does not count heads every second; they note each time the door opens in or out and keep a tally. The count of users can change only at a join or a leave, so the maximum must occur just after some join, and you only need to look at the 2n event times. Checking every time step costs time proportional to the length of the day, and comparing every pair of users costs n squared; both are far too slow at a million users.

    Sort the joins and leaves, then keep a running totaluser 1+-user 2+-user 3+-user 4+-user 5+-t = 7: user 1 leaves,user 5 joins1234join first: false peak of 4true peak 301234567891011timeusers
    Each join adds one user and each leave removes one, so a running total over the sorted events finds the peak of 3; processing the join at time 7 before the leave would produce a false peak of 4.

    Why does the tie rule matter so much?

    At time 7 user 1 leaves and user 5 joins. If a user's session is taken to end just before the moment they leave, then a leave and a join at the same instant never overlap, and the leave must be sorted first; sorting the other way invents a user who was never there. In code this is one comparison in the sort key, and it is exactly the detail interviewers use to separate a working answer from a nearly working one. Ask which convention applies before writing any code.

    The relationship
    peak=max⁡j  ∑i≤jdi,di∈{+1,−1}, events sorted by (ti, di)cost O(nlog⁡n)\text{peak} = \max_{j}\;\sum_{i \le j} d_i,\quad d_i \in \{+1,-1\},\ \text{events sorted by } (t_i,\ d_i) \qquad \text{cost } O(n \log n)
    d_i+1 for a join and -1 for a leave
    (t_i, d_i)the sort key: time first, then leaves (-1) before joins (+1)
    What it says in wordsSort the events so leaves come first at equal times, then the peak is the largest running total.

    Is there a version that avoids building the event list?

    Sort the join times and the leave times separately and walk two pointers through them. At each step take the earlier of the next join and the next leave, taking the leave on a tie, and adjust the count; this is the same sweep without allocating 2n tuples. If times are small integers you can go further: add +1 and -1 into an array indexed by time and take a running sum, which is linear. Mention both and say which you would use for a million users with timestamps in milliseconds.

    Where candidates lose it

    The trap is the tie. Many candidates write a correct sweep, sort by time alone and get 4, because the join at time 7 is processed before the leave. The question states the convention precisely to see whether you use it.

    The second loss is proposing a double loop that checks every pair of sessions. It gives the right answer on five users and fails the question, which asked how you would do it efficiently.

    What the interviewer asks next

    • Return the time interval during which the peak occurs, not just the count.
    • Users arrive as a stream and you must report the current count at any moment. What data structure do you use?
    • How many servers are needed if each can hold at most two users at once?

    Asked at Two Sigma, Equity Hedge, New York, 2025 (Wall Street Oasis): Given arrays (start & end) of the times users join and leave a server, find the max number of concurrent users on the server

  8. 014Five per cent of fund managers are skilled and beat the market in any given year with probability 60%; the rest are unskilled and beat it with probability 50%. Years are independent. A manager has beaten the market in exactly 8 of the last 10 years. What is the probability the manager is skilled?Conditional probability and BayesCoreCSCitadel SecuritiesMiami · 2022

    Try it first

    Roughly how likely is it that this manager is skilled?

    Show the worked solution

    About 12.7%. A skilled manager wins exactly 8 of 10 with probability 0.1209; an unskilled one with 0.0439, a likelihood ratio of about 2.75. Prior odds of skill are 5 to 95, 1 to 19. Posterior odds are 2.75 to 19, so the probability is 0.05 x 0.1209 / (0.05 x 0.1209 + 0.95 x 0.0439) = 0.127. The record helps, but luck has far more players.

    Why does an impressive record move the needle so little?

    Imagine a thousand people each tossing a coin ten times. About 55 of them will get eight heads or better with a fair coin. If a handful of the thousand had slightly biased coins, you still could not pick them out from the lucky crowd by one run of ten. Evidence moves a belief in proportion to how much more likely it is under one explanation than the other, and 8 wins in 10 is not much more likely from a 60% manager than from a 50% one. The ratio is about 2.75.

    An eight-win record is 2.75 times likelier from skill, but luck has 19 times the playersSkilled: wins 60% of years01234567812.1%910winning years out of 10Unskilled: wins 50% of years0123456784.4%910winning years out of 10Weighted by how common each is: 5% x 0.1209 against 95% x 0.043912.7%unskilled but lucky: 87.3%P(skilled | 8 wins in 10) = 12.7%
    Eight wins in ten years has probability 12.1% for a 60% manager and 4.4% for a 50% manager, but after weighting by how common each type is, 5% against 95%, the chance that an eight-win manager is skilled is only 12.7%.

    How do you set it up quickly?

    Use odds. Prior odds of skill are 1 to 19; the likelihood ratioHow many times more likely the evidence is under one hypothesis than under the other. of the record is (0.6/0.5) to the 8 times (0.4/0.5) squared, which is 1.2 to the 8 times 0.64, about 2.75; multiply to get posterior odds of about 0.145. Converting, 2.75 over 2.75 + 19 is 12.7%. The binomial coefficient, 45, is the same in both likelihoods and cancels, so you never need it.

    The relationship
    P(S∣8)P(U∣8)=0.050.95⋅0.68 0.420.510≈119×2.75  ⇒  P(S∣8)≈0.127\frac{P(S\mid 8)}{P(U\mid 8)} = \frac{0.05}{0.95}\cdot\frac{0.6^8\,0.4^2}{0.5^{10}} \approx \frac{1}{19}\times 2.75 \;\Rightarrow\; P(S\mid 8) \approx 0.127
    S, Uskilled and unskilled
    0.6^8 0.4^2the chance of one particular sequence of 8 wins and 2 losses for a skilled manager
    0.5^{10}the same for an unskilled manager
    What it says in wordsPrior odds of 1 to 19, times a likelihood ratio of 2.75, give a posterior of about 12.7%.

    What does this say about picking managers?

    When skill is rare and its edge is small, even a long, strong track record leaves luck as the likelier explanation. Using 8 or more wins instead of exactly 8 barely changes things: the answer becomes 13.9%. The honest limitation is that the model is stylised: real skill is not a fixed 60%, and survivorship means the managers you hear about were already filtered for good records, which pushes the true figure lower still.

    Where candidates lose it

    The common answer is around 80%, reading the record's win rate as the chance of skill. That skips the prior entirely, and with only 5% of managers skilled, the prior dominates.

    The quieter trap is computing the full binomial probabilities, 45 x 0.6 to the 8 x 0.4 squared and so on, and getting lost in decimals. The coefficient cancels. Say odds and likelihood ratio and the arithmetic stays on one line.

    What the interviewer asks next

    • How many years of 80% wins would you need before the manager is more likely skilled than not?
    • What if 20% of managers were skilled?
    • How does survivorship bias change the answer if you only ever see managers with good records?

    Asked at Citadel Securities, Sales and Trading, Miami, 2022 (Wall Street Oasis): I got a question about Bayes' theorem applied to a practical scenario, which I handled decently

  9. 015You roll three fair dice. Is a total of 9 or a total of 10 more likely, given that each can be written as exactly six unordered combinations of three faces?Counting and combinatoricsCoreQuant tradingProp trading firms

    Try it first

    Which total is more likely?

    Show the worked solution

    10 is more likely: 27 ways out of 216 against 25. The equally likely outcomes are the 216 ordered rolls, not the unordered combinations. A combination of three different faces covers six ordered rolls, one with a pair covers three, and a triple covers one. Total 9 includes 3 + 3 + 3, which counts once, and one fewer all-different combination, so it loses two orderings to 10.

    Why are combinations the wrong thing to count?

    Think of dealing two cards and asking whether a pair of kings or a king with a queen is more likely. There is one combination of each, but king and queen can arrive in either order while two kings are just two kings. Probability comes from counting outcomes that are equally likely, and with dice those are the ordered rolls: first die, second die, third die, 6 x 6 x 6 = 216 of them. Unordered combinations bundle different numbers of those outcomes, so counting combinations gives the wrong weights.

    Six combinations each, but the orderings differ: 25 against 27Total of 9orderings1 + 2 + 66 all different1 + 3 + 56 all different1 + 4 + 43 one pair2 + 2 + 53 one pair2 + 3 + 46 all different3 + 3 + 31 tripleTotal25/216 = 11.6%Total of 10orderings1 + 3 + 66 all different1 + 4 + 56 all different2 + 2 + 63 one pair2 + 3 + 56 all different2 + 4 + 43 one pair3 + 3 + 43 one pairTotal27/216 = 12.5%The triple 3 + 3 + 3 counts once; it is what costs 9 the two orderings
    Totals of 9 and 10 each have six unordered combinations, but weighting each combination by its number of orderings gives 27 of 216 rolls for 10 against 25 for 9, mostly because 9 includes the triple 3, 3, 3, which happens only one way.

    How do you count the orderings fast?

    Classify each combination by its repeats. Three different faces give 3! = 6 orders, a pair gives 3 orders, one for each position of the odd die, and a triple gives 1. For 9: 1 2 6, 1 3 5 and 2 3 4 are all different, 18; 1 4 4 and 2 2 5 are pairs, 6; 3 3 3 is a triple, 1; total 25. For 10: 1 3 6, 1 4 5 and 2 3 5 give 18; 2 2 6, 2 4 4 and 3 3 4 give 9; total 27. So 10 comes up 12.5% of the time and 9 only 11.6%.

    The relationship
    P(9)=3⋅6+2⋅3+1216=25216P(10)=3⋅6+3⋅3216=27216P(9) = \frac{3\cdot 6 + 2\cdot 3 + 1}{216} = \frac{25}{216} \qquad P(10) = \frac{3\cdot 6 + 3\cdot 3}{216} = \frac{27}{216}
    6, 3, 1the orderings of an all-different, a pair and a triple combination
    216the ordered outcomes of three dice
    What it says in wordsWeight each combination by its orderings and 10 beats 9 by two rolls in 216.

    Is there a shortcut that avoids listing?

    Yes: symmetry. Replacing each face x by 7 - x maps a total of t to 21 - t, so the distribution of three dice is symmetric about 10.5, and 10 and 11 are the two most likely totals, each 27/216. Anything further from 10.5, including 9, must be less likely or equal; a quick count confirms it is 25. Historically this is the question gamblers put to Galileo, who answered it by counting ordered outcomes, which is still the method.

    Where candidates lose it

    The trap is the question's own framing: six combinations each invites the answer that the totals are equally likely. The interviewer wants you to reject the framing, not accept it.

    The second loss is listing all 216 rolls, or writing out every ordering. Classifying by repeats, six, three or one, gets both totals in under a minute.

    What the interviewer asks next

    • What is the most likely total with four dice, and its probability?
    • What is P(total is 9) with two dice, and why does 9 behave differently?
    • How many ordered outcomes of three dice sum to 7?
  10. 018I will draw a card from a shuffled deck. You may pay Rs 6 to play a bet that pays Rs 10 if the card is red. Before deciding, you may pay to be told the card's colour. What is the most you should pay for that information?Market making, betting and sizingCoreOptiverChicago · 2025

    Try it first

    What is the information worth?

    Show the worked solution

    Rs 2. Without information the bet is worth 0.5 x 10 - 6 = -1, so you decline and your value is 0. With the colour known, you play on red and make 4, and skip black and make 0, which averages 2. Information is worth the improvement in your best decision: 2 - 0 = 2. If it would not change what you do, it is worth nothing.

    How do you value a piece of information?

    Suppose a weather forecast costs money and you are deciding whether to carry an umbrella. If you would carry it anyway, the forecast is worthless to you; it is valuable only if some answer would change what you do. The value of information is the expected value of your best decision with it, minus the expected value of your best decision without it. Work out both decision trees separately and subtract. Never value information by the size of the payout it relates to.

    Information is worth what it changes in your decisionWithout informationYou decidePlay: pay 6half chance of 10EV = 5 - 6 = -1DeclineEV = 0Best choice: decline. Value = 0With information firstColour?red, 1/2black, 1/2Red: playwin 10 - 6 = +4Black: decline0Value = 1/2 x 4 + 1/2 x 0 = 2Worth paying for the information: up to 2 - 0 = Rs 2
    Blind, the bet has an expected value of minus 1 so you decline and get 0; told the colour first, you play only on red and make 4 half the time, an average of 2, so the information is worth Rs 2.

    Why is it not worth Rs 4 or Rs 5?

    Rs 4 is what you make when the card is red, but it is red only half the time. Rs 5 is half the payout, which ignores the Rs 6 you pay to play. The information saves you from the losing half of the bet and lets you keep the winning half, and that is worth half of Rs 4, which is Rs 2. Pay more than Rs 2 and you would do better declining the offer of information and declining the bet.

    The relationship
    VOI=E[max⁡(payoff,0)]−max⁡(E[payoff],0)=12max⁡(4,0)+12max⁡(−6,0)−max⁡(−1,0)=2\text{VOI} = E\big[\max(\text{payoff}, 0)\big] - \max\big(E[\text{payoff}], 0\big) = \tfrac12\max(4,0) + \tfrac12\max(-6,0) - \max(-1,0) = 2
    payoff10 - 6 = 4 on red, -6 on black
    E[max(payoff, 0)]your value when you can choose after seeing the colour
    max(E[payoff], 0)your value when you must choose blind
    What it says in wordsInformation is worth the gap between deciding after you know and deciding before.

    When is information worth the most?

    Vary the price of the bet. At a price of 5 you are exactly indifferent blind, and the information is worth 2.50, its maximum; at a price of 0 you would always play, and it is worth 0. Information is valuable when you are close to indifferent and the decision could go either way. The formula also has the shape of an option payoff: knowing first lets you exercise only when it pays, which is why traders talk about paying for optionality and paying for information in the same breath.

    Where candidates lose it

    The trap is answering with the size of the win, Rs 4, or half the payout, Rs 5. Both value the information by the bet it is about, not by the decision it improves.

    The second loss is forgetting that without information you would decline. Candidates who compare with playing blind, at -1, get 3. The comparison is always with your best action without the information, which here is to walk away.

    What the interviewer asks next

    • What is the information worth if the bet costs Rs 3?
    • What if the information is only 80% reliable?
    • You can pay to see one card of a two-card hand before betting. How do you decide what that is worth?

    Asked at Optiver, Quantitative Research, Chicago, 2025 (Wall Street Oasis): Valuing information, taking directional bets when not plus EV.

← PreviousPage 1 of 5
  1. 1
  2. 2
  3. …
  4. 5
Next →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.