Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant puzzles, solved step by step

Puzzles
100
Traced to a firm
71
Topics
12
Hard
30
Topic
All topicsLogic and algorithmic reasoning10Conditional probability and Bayes7Counting and combinatorics8Continuous and geometric probability9Correlation, regression and linear algebra9Market making, betting and sizing9Expected value and optimal stopping9Statistics and estimation9Pricing, options and index maths7Games and strategic reasoning8Markov chains and random walks7Mental maths and number sense8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 1–10 of 33 · filtered from 100Clear filters
  1. 001Four people must cross a narrow bridge at night with one torch. At most two can be on the bridge at once, anyone crossing must carry the torch, and a pair walks at the slower person's pace. They take 1, 2, 5 and 10 minutes. What is the shortest total time to get everyone across?Logic and algorithmic reasoningCoreBelvedere TradingChicago · 2021

    Try it first

    Before you plan it: what is the fastest time?

    Show the worked solution

    17 minutes. Send 1 and 2 over (2 minutes), 1 comes back (1), 5 and 10 cross together (10), 2 comes back (2), and 1 and 2 cross again (2). The obvious plan, where the fastest person escorts each of the others, takes 19. The saving comes from putting the two slowest walkers on the bridge at the same time.

    Why does the obvious plan lose two minutes?

    Think of two slow parcels going by the same courier. If each travels on its own trip you pay for both trips; if they share a van, you pay for the slower one only. Every crossing costs the slower walker's time, so a slow person paired with a fast one wastes the fast one, and two slow people paired together waste nothing. Escorting with the fastest person pays 10 and then 5 as separate crossings. Pairing 5 with 10 pays 10 once and the 5 minutes are free.

    Send the two slowest together and the 5 minutes hide inside the 10Pair the slowest21 and 2 over11 back105 and 10 over22 back21 and 2 over17 minFastest escorts all101 and 10 over11 back51 and 5 over11 back21 and 2 over19 min0510151719minutes
    Pairing the 5 and 10 minute walkers on one crossing finishes in 17 minutes, while letting the 1 minute walker escort everyone pays for the 10 and the 5 separately and finishes in 19.

    What is the price of pairing the slow two?

    Somebody has to bring the torch back after the slow pair crosses, and it must not be one of them. So the plan first ferries two fast people over, leaves one on the far side to carry the torch back later, and spends the 2 minute walker's return trip to buy the 5 minute saving. The trade is 1 + 2 extra minutes of shuttling against 5 minutes saved on the slow side, a net gain of 2. With different speeds the trade can flip, which is the real content of the puzzle.

    The relationship
    escort: 2a+b+c+dpair: a+3b+dpair wins when 2b<a+c\text{escort: } 2a + b + c + d \qquad \text{pair: } a + 3b + d \qquad \text{pair wins when } 2b < a + c
    a, bthe two fastest times, here 1 and 2
    c, dthe two slowest times, here 5 and 10
    What it says in wordsPairing the slow two is better exactly when twice the second fastest time is less than the fastest plus the second slowest.

    How do you convince the interviewer 17 cannot be beaten?

    There must be at least five crossings, three over and two back, because each trip over moves at most two people and someone must return the torch. The 10 minute walker costs 10 on whatever crossing carries them. If 5 and 10 cross separately you already spend 15 on those two trips, and the three remaining crossings cost at least 1 + 1 + 2, which is 19; if they cross together, the best you can do with the other four crossings is 2 + 1 + 2 + 2. A brute force over every schedule gives the same minimum, 17 minutes.

    Where candidates lose it

    The strong candidate's trap is a fast answer of 19. Letting the quickest person run every errand feels efficient, and it is the right instinct for returning the torch, but it is the wrong instinct for the slow walkers.

    The second loss is getting 17 by trial and error and then being unable to say why. State the principle, that the slow pair shares one crossing, and give the rule for when it wins: when twice the second fastest time is below the fastest plus the second slowest.

    What the interviewer asks next

    • What if the times are 1, 4, 5 and 10?
    • Six people with times 1, 2, 5, 10, 20 and 25: what is the plan?
    • Write the general algorithm for n people and say its running time.

    Asked at Belvedere Trading, Trading, Chicago, 2021 (Wall Street Oasis): crossing the bridge in the shortest amount of time with one flashlight brainteaser

  2. 002You are flying to a city where it rains on 25% of days. You phone three friends who live there. Each tells the truth with probability 2/3, independently of the others, and all three say it is raining. What is the probability that it is actually raining?Conditional probability and BayesCoreJane StreetNew York · 2025

    Try it first

    Pick your answer before working it.

    Show the worked solution

    8/11, about 72.7%. If it is raining, all three say yes with probability (2/3)^3 = 8/27. If it is dry, all three must be lying, (1/3)^3 = 1/27. Weight each by how often it happens: 1/4 x 8/27 against 3/4 x 1/27, which is 8 parts to 3. Three agreeing witnesses move a 25% prior a long way, but not to certainty.

    Why is the answer not simply 8/9?

    Picture a clinic where a test is quite reliable but the illness is uncommon. A positive result makes the illness more likely, but how much more depends on how rare it was to begin with. The friends' agreement tells you how much more likely rain makes their answer than dry does, eight times, but it does not erase the fact that dry days are three times as common. 8/9 is the answer you get if rain and dry start level. Here they do not.

    Three yeses: compare the two shaded areas, not the two stripsall 3 say yes8/27not all yes19/27all 3 lie and say yes: 1/27not all yes26/27Rain, 1/4Dry, 3/4Shaded areasRain: 1/4 x 8/27= 8/108Dry: 3/4 x 1/27= 3/108P(rain | 3 yes)8/11 = 72.7%rain: 8 partsdry: 3 partsInside the shaded region only
    Rain covers a quarter of days and all three friends say yes on 8/27 of those, while dry days cover three quarters and all three lie on only 1/27 of them, so the shaded areas stand 8 to 3 and the chance of rain given three yeses is 8/11, about 72.7%.

    How do you set it up so the arithmetic stays small?

    Use odds rather than probabilities. Posterior odds are prior odds times the likelihood ratioHow many times more likely the evidence is if the hypothesis is true than if it is false., and both are easy numbers here. Prior odds of rain are 1 to 3. The likelihood ratio of three yeses is (2/3)^3 over (1/3)^3, which is 2 cubed, 8. So the posterior odds are 8 to 3, and the probability is 8 over 8 plus 3, 8/11. Each additional agreeing friend would double the odds again.

    The relationship
    P(R∣YYY)P(D∣YYY)=P(R)P(D)⋅(2/3)3(1/3)3=13⋅8=83  ⇒  P(R∣YYY)=811\frac{P(R\mid YYY)}{P(D\mid YYY)} = \frac{P(R)}{P(D)}\cdot\frac{(2/3)^3}{(1/3)^3} = \frac{1}{3}\cdot 8 = \frac{8}{3} \;\Rightarrow\; P(R\mid YYY) = \frac{8}{11}
    R, Drain and dry
    YYYall three friends say yes
    (2/3)^3 and (1/3)^3the chance of three yeses when it rains, and when it is dry
    What it says in wordsMultiply the prior odds by how much more likely the evidence is under rain, then turn the odds back into a probability.

    What assumption is doing the work, and should you say it?

    The calculation needs the friends to lie independently. If they could be coordinating a joke, three yeses are really one piece of evidence, and the answer falls back towards the one-friend figure of 2/5. Say the independence assumption out loud, then give 8/11. Interviewers often follow up by making one friend unreliable or by letting them talk to each other.

    Where candidates lose it

    The most common wrong answer is 8/9: the candidate compares the chance of three truths with the chance of three lies and forgets the weather's own odds. The rain prior is a quarter, and leaving it out quietly assumes it is a coin flip.

    The second loss is writing out a full Bayes formula with 27ths and 108ths and losing the thread under time pressure. Odds times likelihood ratio gets 8 to 3 in two lines and is easier to check out loud.

    What the interviewer asks next

    • What if only two of the three friends say yes?
    • How many agreeing friends would you need before you were 95% sure it is raining?
    • What changes if the friends can talk to each other before answering?

    Asked at Jane Street, Generalist, New York, 2025 (Wall Street Oasis): There was a question about the probability of rain the next day that relied on a very in depth understanding of bayes theorem

  3. 013Users join a server at times 1, 2, 4, 5 and 7 and leave at times 7, 3, 8, 9 and 10 respectively. A leave at the same moment as a join is processed first. What is the maximum number of users online at once, and how would you compute it efficiently for a million users?Logic and algorithmic reasoningCoreTwo SigmaNew York · 2025

    Try it first

    What is the peak number of users online together?

    Show the worked solution

    The peak is 3 users. Turn every join into a +1 event and every leave into a -1 event, sort all ten events by time with leaves before joins at equal times, and keep a running total. It goes 1, 2, 1, 2, 3, then at time 7 down to 2 and back to 3, then 2, 1, 0. Sorting costs n log n, and the sweep itself is linear.

    Why not check every moment in time?

    A shopkeeper who wants to know the busiest moment of the day does not count heads every second; they note each time the door opens in or out and keep a tally. The count of users can change only at a join or a leave, so the maximum must occur just after some join, and you only need to look at the 2n event times. Checking every time step costs time proportional to the length of the day, and comparing every pair of users costs n squared; both are far too slow at a million users.

    Sort the joins and leaves, then keep a running totaluser 1+-user 2+-user 3+-user 4+-user 5+-t = 7: user 1 leaves,user 5 joins1234join first: false peak of 4true peak 301234567891011timeusers
    Each join adds one user and each leave removes one, so a running total over the sorted events finds the peak of 3; processing the join at time 7 before the leave would produce a false peak of 4.

    Why does the tie rule matter so much?

    At time 7 user 1 leaves and user 5 joins. If a user's session is taken to end just before the moment they leave, then a leave and a join at the same instant never overlap, and the leave must be sorted first; sorting the other way invents a user who was never there. In code this is one comparison in the sort key, and it is exactly the detail interviewers use to separate a working answer from a nearly working one. Ask which convention applies before writing any code.

    The relationship
    peak=max⁡j  ∑i≤jdi,di∈{+1,−1}, events sorted by (ti, di)cost O(nlog⁡n)\text{peak} = \max_{j}\;\sum_{i \le j} d_i,\quad d_i \in \{+1,-1\},\ \text{events sorted by } (t_i,\ d_i) \qquad \text{cost } O(n \log n)
    d_i+1 for a join and -1 for a leave
    (t_i, d_i)the sort key: time first, then leaves (-1) before joins (+1)
    What it says in wordsSort the events so leaves come first at equal times, then the peak is the largest running total.

    Is there a version that avoids building the event list?

    Sort the join times and the leave times separately and walk two pointers through them. At each step take the earlier of the next join and the next leave, taking the leave on a tie, and adjust the count; this is the same sweep without allocating 2n tuples. If times are small integers you can go further: add +1 and -1 into an array indexed by time and take a running sum, which is linear. Mention both and say which you would use for a million users with timestamps in milliseconds.

    Where candidates lose it

    The trap is the tie. Many candidates write a correct sweep, sort by time alone and get 4, because the join at time 7 is processed before the leave. The question states the convention precisely to see whether you use it.

    The second loss is proposing a double loop that checks every pair of sessions. It gives the right answer on five users and fails the question, which asked how you would do it efficiently.

    What the interviewer asks next

    • Return the time interval during which the peak occurs, not just the count.
    • Users arrive as a stream and you must report the current count at any moment. What data structure do you use?
    • How many servers are needed if each can hold at most two users at once?

    Asked at Two Sigma, Equity Hedge, New York, 2025 (Wall Street Oasis): Given arrays (start & end) of the times users join and leave a server, find the max number of concurrent users on the server

  4. 014Five per cent of fund managers are skilled and beat the market in any given year with probability 60%; the rest are unskilled and beat it with probability 50%. Years are independent. A manager has beaten the market in exactly 8 of the last 10 years. What is the probability the manager is skilled?Conditional probability and BayesCoreCSCitadel SecuritiesMiami · 2022

    Try it first

    Roughly how likely is it that this manager is skilled?

    Show the worked solution

    About 12.7%. A skilled manager wins exactly 8 of 10 with probability 0.1209; an unskilled one with 0.0439, a likelihood ratio of about 2.75. Prior odds of skill are 5 to 95, 1 to 19. Posterior odds are 2.75 to 19, so the probability is 0.05 x 0.1209 / (0.05 x 0.1209 + 0.95 x 0.0439) = 0.127. The record helps, but luck has far more players.

    Why does an impressive record move the needle so little?

    Imagine a thousand people each tossing a coin ten times. About 55 of them will get eight heads or better with a fair coin. If a handful of the thousand had slightly biased coins, you still could not pick them out from the lucky crowd by one run of ten. Evidence moves a belief in proportion to how much more likely it is under one explanation than the other, and 8 wins in 10 is not much more likely from a 60% manager than from a 50% one. The ratio is about 2.75.

    An eight-win record is 2.75 times likelier from skill, but luck has 19 times the playersSkilled: wins 60% of years01234567812.1%910winning years out of 10Unskilled: wins 50% of years0123456784.4%910winning years out of 10Weighted by how common each is: 5% x 0.1209 against 95% x 0.043912.7%unskilled but lucky: 87.3%P(skilled | 8 wins in 10) = 12.7%
    Eight wins in ten years has probability 12.1% for a 60% manager and 4.4% for a 50% manager, but after weighting by how common each type is, 5% against 95%, the chance that an eight-win manager is skilled is only 12.7%.

    How do you set it up quickly?

    Use odds. Prior odds of skill are 1 to 19; the likelihood ratioHow many times more likely the evidence is under one hypothesis than under the other. of the record is (0.6/0.5) to the 8 times (0.4/0.5) squared, which is 1.2 to the 8 times 0.64, about 2.75; multiply to get posterior odds of about 0.145. Converting, 2.75 over 2.75 + 19 is 12.7%. The binomial coefficient, 45, is the same in both likelihoods and cancels, so you never need it.

    The relationship
    P(S∣8)P(U∣8)=0.050.95⋅0.68 0.420.510≈119×2.75  ⇒  P(S∣8)≈0.127\frac{P(S\mid 8)}{P(U\mid 8)} = \frac{0.05}{0.95}\cdot\frac{0.6^8\,0.4^2}{0.5^{10}} \approx \frac{1}{19}\times 2.75 \;\Rightarrow\; P(S\mid 8) \approx 0.127
    S, Uskilled and unskilled
    0.6^8 0.4^2the chance of one particular sequence of 8 wins and 2 losses for a skilled manager
    0.5^{10}the same for an unskilled manager
    What it says in wordsPrior odds of 1 to 19, times a likelihood ratio of 2.75, give a posterior of about 12.7%.

    What does this say about picking managers?

    When skill is rare and its edge is small, even a long, strong track record leaves luck as the likelier explanation. Using 8 or more wins instead of exactly 8 barely changes things: the answer becomes 13.9%. The honest limitation is that the model is stylised: real skill is not a fixed 60%, and survivorship means the managers you hear about were already filtered for good records, which pushes the true figure lower still.

    Where candidates lose it

    The common answer is around 80%, reading the record's win rate as the chance of skill. That skips the prior entirely, and with only 5% of managers skilled, the prior dominates.

    The quieter trap is computing the full binomial probabilities, 45 x 0.6 to the 8 x 0.4 squared and so on, and getting lost in decimals. The coefficient cancels. Say odds and likelihood ratio and the arithmetic stays on one line.

    What the interviewer asks next

    • How many years of 80% wins would you need before the manager is more likely skilled than not?
    • What if 20% of managers were skilled?
    • How does survivorship bias change the answer if you only ever see managers with good records?

    Asked at Citadel Securities, Sales and Trading, Miami, 2022 (Wall Street Oasis): I got a question about Bayes' theorem applied to a practical scenario, which I handled decently

  5. 018I will draw a card from a shuffled deck. You may pay Rs 6 to play a bet that pays Rs 10 if the card is red. Before deciding, you may pay to be told the card's colour. What is the most you should pay for that information?Market making, betting and sizingCoreOptiverChicago · 2025

    Try it first

    What is the information worth?

    Show the worked solution

    Rs 2. Without information the bet is worth 0.5 x 10 - 6 = -1, so you decline and your value is 0. With the colour known, you play on red and make 4, and skip black and make 0, which averages 2. Information is worth the improvement in your best decision: 2 - 0 = 2. If it would not change what you do, it is worth nothing.

    How do you value a piece of information?

    Suppose a weather forecast costs money and you are deciding whether to carry an umbrella. If you would carry it anyway, the forecast is worthless to you; it is valuable only if some answer would change what you do. The value of information is the expected value of your best decision with it, minus the expected value of your best decision without it. Work out both decision trees separately and subtract. Never value information by the size of the payout it relates to.

    Information is worth what it changes in your decisionWithout informationYou decidePlay: pay 6half chance of 10EV = 5 - 6 = -1DeclineEV = 0Best choice: decline. Value = 0With information firstColour?red, 1/2black, 1/2Red: playwin 10 - 6 = +4Black: decline0Value = 1/2 x 4 + 1/2 x 0 = 2Worth paying for the information: up to 2 - 0 = Rs 2
    Blind, the bet has an expected value of minus 1 so you decline and get 0; told the colour first, you play only on red and make 4 half the time, an average of 2, so the information is worth Rs 2.

    Why is it not worth Rs 4 or Rs 5?

    Rs 4 is what you make when the card is red, but it is red only half the time. Rs 5 is half the payout, which ignores the Rs 6 you pay to play. The information saves you from the losing half of the bet and lets you keep the winning half, and that is worth half of Rs 4, which is Rs 2. Pay more than Rs 2 and you would do better declining the offer of information and declining the bet.

    The relationship
    VOI=E[max⁡(payoff,0)]−max⁡(E[payoff],0)=12max⁡(4,0)+12max⁡(−6,0)−max⁡(−1,0)=2\text{VOI} = E\big[\max(\text{payoff}, 0)\big] - \max\big(E[\text{payoff}], 0\big) = \tfrac12\max(4,0) + \tfrac12\max(-6,0) - \max(-1,0) = 2
    payoff10 - 6 = 4 on red, -6 on black
    E[max(payoff, 0)]your value when you can choose after seeing the colour
    max(E[payoff], 0)your value when you must choose blind
    What it says in wordsInformation is worth the gap between deciding after you know and deciding before.

    When is information worth the most?

    Vary the price of the bet. At a price of 5 you are exactly indifferent blind, and the information is worth 2.50, its maximum; at a price of 0 you would always play, and it is worth 0. Information is valuable when you are close to indifferent and the decision could go either way. The formula also has the shape of an option payoff: knowing first lets you exercise only when it pays, which is why traders talk about paying for optionality and paying for information in the same breath.

    Where candidates lose it

    The trap is answering with the size of the win, Rs 4, or half the payout, Rs 5. Both value the information by the bet it is about, not by the decision it improves.

    The second loss is forgetting that without information you would decline. Candidates who compare with playing blind, at -1, get 3. The comparison is always with your best action without the information, which here is to walk away.

    What the interviewer asks next

    • What is the information worth if the bet costs Rs 3?
    • What if the information is only 80% reliable?
    • You can pay to see one card of a two-card hand before betting. How do you decide what that is worth?

    Asked at Optiver, Quantitative Research, Chicago, 2025 (Wall Street Oasis): Valuing information, taking directional bets when not plus EV.

  6. 020A broad stock index has returned an average of 11% a year over the past 30 years, with an annual standard deviation of 16% (illustrative figures). Assuming yearly returns are independent, give a 95% confidence interval for its true expected annual return.Statistics and estimationCoreOld Mission CapitalChicago · 2025

    Try it first

    Roughly how wide is the 95% interval?

    Show the worked solution

    About 5.3% to 16.7%. The standard error of a 30-year average is 16% divided by the square root of 30, 2.92 points. A 95% interval is 1.96 standard errors either side: 11% plus or minus 5.7. With a t value for 29 degrees of freedom it is a touch wider, plus or minus 6.0. Thirty years of data still leave the expected return very uncertain.

    Why is the band so wide after thirty years?

    Imagine judging a new cricketer's true batting average from a handful of innings. Scores swing wildly from one innings to the next, so a few innings tell you little. The precision of an average grows only with the square root of the number of observations, and yearly returns swing by 16 points, so 30 years give a standard error of about 2.9 points. That is large next to an average of 11%: the true figure could plausibly be 6% or 16%.

    Thirty years pins the average return only to within about 6 points11% sample average10 years1.1%20.9%+/- 9.9 points30 years5.3%16.7%+/- 5.7 points120 years8.1%13.9%+/- 2.9 points0%5%10%15%20%Half width = 1.96 x 16% / square root of years
    Around an 11% average with 16% annual volatility, the 95% interval is plus or minus 9.9 points with 10 years of data, 5.7 points with 30 years and still 2.9 points with 120 years, because precision grows only with the square root of the sample.

    How do you set it up in the room?

    State the assumption, then compute. Treat the 30 annual returns as independent draws with a standard deviation of 16%, so the sample average has a standard error of 16 over root 30. Root 30 is about 5.48, so the standard error is 2.92. Multiply by 1.96: 5.73. The interval is 5.3% to 16.7%. If the interviewer pushes, note that with only 30 observations a t distributionThe distribution used for an average when the standard deviation is estimated from the same small sample; it has fatter tails than the normal. with 29 degrees of freedom gives a multiplier of about 2.045, widening the band slightly.

    The relationship
    rˉ±1.96 σn=11%±1.96×16%30=11%±5.7%  ⇒  [5.3%, 16.7%]\bar r \pm 1.96\,\frac{\sigma}{\sqrt n} = 11\% \pm 1.96 \times \frac{16\%}{\sqrt{30}} = 11\% \pm 5.7\% \;\Rightarrow\; [5.3\%,\ 16.7\%]
    \bar rthe average annual return, 11%
    \sigmathe annual standard deviation, 16%
    nthe number of yearly observations, 30
    What it says in wordsThe interval is the average plus or minus 1.96 standard errors, and the standard error is volatility over the square root of years.

    What limitation should you add?

    Halving the band needs four times the data: 120 years still only pins the average to within about 2.9 points, and markets change over such spans. Returns are also not quite independent from year to year, fat tails make the normal multiplier optimistic, and the arithmetic average overstates the compound growth rate. The practical lesson is that expected returns are the hardest input in finance to estimate, far harder than volatility, which is why many quant processes lean on risk estimates and treat return forecasts with caution.

    Where candidates lose it

    The common mistake is using 16% as the width, the spread of a single year, rather than the standard error of the average. That gives an interval from -20% to 42%, which describes one year's return, not the long-run mean.

    The opposite slip is dividing by 30 instead of the square root of 30, which gives a band of about 1 point and badly overstates what the data can tell you.

    What the interviewer asks next

    • How many years of data would you need to pin the mean to within plus or minus 1 point?
    • How would monthly data change the interval for the mean?
    • Why is volatility much easier to estimate than the mean from the same data?

    Asked at Old Mission Capital, Prop Trading, Chicago, 2025 (Wall Street Oasis): Confidence interval on S&P 500 return past 30 years

  7. 024Differentiate f(x) = x to the power x, and find where it reaches its minimum for positive x.Mental maths and number senseCoreScotiabankToronto · 2026

    Try it first

    What is the derivative of x to the x?

    Show the worked solution

    f'(x) = x to the x times (ln x + 1), and the minimum is at x = 1/e, about 0.368, where f is about 0.692. Take logs: ln f = x ln x. Differentiating, f'/f = ln x + 1, so f' = x to the x (ln x + 1). The derivative is zero when ln x = -1, that is x = 1/e, negative before it and positive after, so this is a minimum.

    Why do both standard rules fail?

    The power rule, n x to the (n - 1), treats the exponent as fixed; the exponential rule, a to the x times ln a, treats the base as fixed. In x to the x both the base and the exponent move, so neither rule applies on its own, and each one gives half of the right answer. Indeed the correct derivative is the sum of the two: x times x to the (x - 1), which is x to the x, plus x to the x ln x. That sum is a quick check on your final answer.

    Take logs first: the minimum sits at x = 1/e either waymin 0.692 at x = 1/e = 0.3682 to the 2 = 40121234f(x) = x to the xmin -1/e = -0.36812ln f(x) = x ln xd/dx of x ln x = ln x + 1f'(x) = x to the x (ln x + 1)
    The curve x to the x falls from near 1 at zero to a minimum of 0.692 at x = 1/e and then rises to 4 at x = 2, and its logarithm x ln x has its minimum at the same point, which is why taking logs first is safe.

    How does taking logs make it routine?

    Think of converting a messy multiplication into addition before doing it, the way a slide rule does. Write ln f = x ln x; the right-hand side is a product of two simple functions, and the product rule gives ln x + x times 1/x = ln x + 1. The left side differentiates to f'/f by the chain rule, so f' = f (ln x + 1). This is logarithmic differentiationDifferentiating the logarithm of a function instead of the function, then multiplying back; useful when the variable sits in an exponent., and it works for any function of the form g(x) to the h(x).

    The relationship
    f(x)=exln⁡x  ⇒  f′(x)=exln⁡x(ln⁡x+1)=xx(ln⁡x+1)=0  ⟺  x=e−1≈0.368f(x) = e^{x\ln x} \;\Rightarrow\; f'(x) = e^{x\ln x}(\ln x + 1) = x^x(\ln x + 1) = 0 \iff x = e^{-1} \approx 0.368
    e^{x ln x}x to the x rewritten with a fixed base
    ln x + 1the derivative of x ln x
    e^{-1}where ln x = -1, the minimum
    What it says in wordsRewrite with base e, differentiate the exponent, and set it to zero.

    How do you confirm it is a minimum and state the value?

    Check the sign of ln x + 1, since x to the x is always positive. For x below 1/e, ln x is below -1 and the slope is negative; above 1/e it is positive, so the function falls and then rises: a minimum. The value is (1/e) to the (1/e) = e to the (-1/e), about 0.692. Also say what happens at the edges: as x shrinks towards zero, x ln x tends to zero, so x to the x tends to 1, and at x = 1 it is exactly 1 again.

    Where candidates lose it

    The fast wrong answer applies the power rule, x times x to the (x - 1), which is just x to the x. It treats the exponent as a constant, and candidates who give it usually do so in the first three seconds.

    The second loss is finding x = 1/e and stopping. The question asks for the minimum, so check the sign change and give the value, e to the (-1/e), about 0.692, together with the behaviour near zero.

    What the interviewer asks next

    • Differentiate x to the (x to the x).
    • What is the limit of x to the x as x approaches 0 from above, and why?
    • Which is larger, e to the pi or pi to the e, and how does x to the (1/x) settle it?

    Asked at Scotiabank, Quant, Toronto, 2026 (Wall Street Oasis): technical questions covering calculus (including derivatives of standard functions)

  8. 025Four people queue at a cash machine wanting 7, 3, 10 and 2 thousand rupees. Each visit allows at most 4 thousand, and anyone who has not got their full amount rejoins the back of the queue. In what order do they leave, and how would you compute the order quickly for a very long queue?Logic and algorithmic reasoningCoreSCSquarepoint CapitalLondon · 2026

    Try it first

    In what order do the four people leave?

    Show the worked solution

    They leave in the order 2, 4, 1, 3. Person 1 takes 4 and rejoins with 3; person 2 takes 3 and leaves; person 3 takes 4 and rejoins with 6; person 4 takes 2 and leaves. Then person 1 takes 3 and leaves, and person 3 needs two more visits. The shortcut: each person leaves in round amount divided by 4, rounded up, and ties go to whoever stood first.

    How do you simulate it cleanly?

    Model the line as a queueA first-in, first-out list: items join at the back and leave from the front, as in a real line. of pairs, person and amount still wanted. Pop the front, subtract the lesser of the cap and what they still want, and if anything is left push them onto the back; otherwise record them as leaving. The four people take seven visits in all. This is the answer most interviewers expect first, and it is correct, but its cost grows with the total number of visits, which is the sum of each amount over the cap.

    Each visit takes at most 4; the rest goes to the back of the lineVisitroundamount still wanted (block = 1 thousand)1. person 1r1takes 4, back of queue with 32. person 2r1takes 3, leaves: exit 13. person 3r1takes 4, back of queue with 64. person 4r1takes 2, leaves: exit 25. person 1r2takes 3, leaves: exit 36. person 3r2takes 4, back of queue with 27. person 3r3takes 2, leaves: exit 4Exit order2, 4, 1, 3Rounds needed = amount / 4, rounded up: 2, 1, 3, 1
    Seven visits clear the queue: persons 2 and 4 leave on their first visit, person 1 on the second round and person 3 on the third, so the exit order is 2, 4, 1, 3, matching each person's amount divided by 4, rounded up.

    Is there a faster way than simulating?

    Yes. Think of a canteen that serves one plate per person per pass: someone wanting three plates leaves on the third pass, whatever the others want. Person i leaves in round ceiling(a_i / k), and within a round the queue keeps its original order, so the exit order is simply the people sorted by their round number, ties broken by starting position. Here the rounds are 2, 1, 3 and 1, which sorts to 2, 4, 1, 3. That costs n log n, however large the amounts are, instead of the number of visits.

    The relationship
    ri=⌈aik⌉order=sort⁡i (ri, i)r=(2,1,3,1)⇒2,4,1,3r_i = \left\lceil \frac{a_i}{k} \right\rceil \qquad \text{order} = \operatorname{sort}_i\,(r_i,\ i) \qquad r = (2,1,3,1) \Rightarrow 2, 4, 1, 3
    a_ithe amount person i wants
    kthe cap per visit, 4
    r_ithe round in which person i leaves
    What it says in wordsSort people by how many rounds they need, and by queue position within a round.

    Why does queue order survive between rounds?

    Everyone still waiting after a round rejoins in the same relative order they were served, because the queue is first in, first out. So round two serves the survivors of round one in their original order, and so on. That invariant is what lets you replace the simulation with a sort, and saying it out loud is what separates an answer that works from one you can defend. If a very large cap or tiny amounts made most people finish in round one, the sort still costs n log n, and a counting sort on round numbers can make it linear.

    Where candidates lose it

    The common wrong answer sorts by amount: 4, 2, 1, 3. It ignores that people who finish in the same round leave in queue order, and person 2 stands ahead of person 4.

    The second loss is stopping at the simulation when the question asks how to do it quickly. With amounts in the crores and a small cap, simulating each visit could take billions of steps. The ceiling formula plus a stable sort is the answer to the second half.

    What the interviewer asks next

    • Return the time at which each person leaves if each visit takes one minute.
    • What if the cap differs by visit, for example 4 thousand on odd visits and 2 on even ones?
    • Implement the sort-based version and state its complexity.

    Asked at Squarepoint Capital, Quant Research Intern Interview, London, 2026 (Wall Street Oasis): returning the order in which people leave a queue given a list of amounts people want to withdraw from an ATM

  9. 027A bag holds three dice: one fair, one that shows six half the time with its other faces equally likely, and one that never shows six. You draw one at random and roll it twice, getting two sixes. What is the probability it is the loaded die?Conditional probability and BayesCoreBelvedere TradingChicago · 2022

    Try it first

    Before you calculate: how likely is it now that you hold the loaded die?

    Show the worked solution

    90%. Each die starts at one in three. The chance of two sixes is 1/36 for the fair die, 1/4 for the loaded die and zero for the die with no six. Weight each by its prior: 1/108 for the fair die, 1/12 for the loaded die, nothing for the third. The loaded die's weight is nine times the fair die's, so its probability is 9/10.

    What does the die that never shows six do to the answer?

    A neighbour tells you a red car blocked the gate this morning. If one of your suspects owns only a blue scooter, that suspect is out, however likely they looked before. A hypothesis that cannot produce the evidence gets zero weight afterwards, no matter what its prior was. The no-six die could never give two sixes, so it drops out, and the question becomes a contest between the fair die and the loaded die, which started level at one third each.

    Two sixes: each die's prior times its chance of producing themDieP(six)P(two sixes)Prior x likelihoodPosteriorFair die1/61/361/3 x 1/36 = 1/10810%Loaded die1/21/41/3 x 1/4 = 1/1290%No-six die001/3 x 0 = 0eliminated: cannot roll a sixWeights 1/12 against 1/108: odds of 9 to 1 for the loaded dieThe prior of 1/3 is common to both, so it cancels9/10 = 90%
    Each die starts at one third; multiplying by the chance of two sixes gives weights of 1/108 for the fair die, 1/12 for the loaded die and zero for the no-six die, so the loaded die ends at 90% and the fair die at 10%.

    How much does each surviving die's likelihood count?

    Now compare how easily each remaining die produces what you saw. The fair die gives two sixes 1 time in 36. The loaded die gives a six half the time, so two in a row 1 time in 4. With equal priors, the posterior odds are just the ratio of the likelihoodsThe probability of the observed evidence under each hypothesis, before any prior is applied.: 1/4 against 1/36, which is 9 to 1. Nine parts in ten is 90%.

    The relationship
    P(L∣66)=13⋅1413⋅136+13⋅14+13⋅0=910P(L \mid 66) = \frac{\tfrac13\cdot\tfrac14}{\tfrac13\cdot\tfrac1{36} + \tfrac13\cdot\tfrac14 + \tfrac13\cdot 0} = \frac{9}{10}
    Lthe loaded die was drawn
    66the evidence: two sixes in two rolls
    1/3the prior for each die
    1/36, 1/4, 0the chance of two sixes from the fair, loaded and no-six dice
    What it says in wordsThe loaded die's share of all the ways two sixes can happen is nine tenths.

    Check it by counting, the safer habit under pressure. Imagine 108 rounds of drawing a die and rolling it twice, 36 rounds with each die. The fair die gives two sixes once, the loaded die 9 times, the no-six die never. Of the 10 double sixes, 9 came from the loaded die, and the prior of one third cancels because every die got the same number of rounds. The loaded die's other faces, 1 in 10 each, never enter, because only sixes were seen.

    Where candidates lose it

    The quick wrong answer is 1/2: two dice can roll a six, so it must be one or the other. That ignores how differently they produce two sixes in a row, a gap of nine to one.

    The other loss is getting tangled in the loaded die's other faces, or leaving the no-six die in the denominator with some weight. Neither belongs: only the chance of the observed rolls counts, and for the no-six die that chance is zero.

    What the interviewer asks next

    • A third roll is also a six. What is the probability of the loaded die now? (It rises to 27/28.)
    • The rolls were a six and then a two. Which die is most likely now?
    • What is the chance the next roll is a six? (It is 7/15.)

    Asked at Belvedere Trading, Capital Markets, Chicago, 2022 (Wall Street Oasis): The technical portion of the interview consisted of probability questions including one questions relating to Bayes' theorem

  10. 028You roll two fair dice and are paid the larger of the two faces in rupees. What is the expected payout?Expected value and optimal stoppingCoreJane StreetNew York · 2026

    Try it first

    Pick the expected payout before you count.

    Show the worked solution

    161/36, about Rs 4.47. The larger face equals k in 2k minus 1 of the 36 equally likely outcomes: 1, 3, 5, 7, 9 and 11 cells for k from 1 to 6. Multiply each value by its count and add: 1 + 6 + 15 + 28 + 45 + 66 = 161. Divided by 36, that is 4.47, almost a full point above a single die's 3.5.

    Why is the answer well above 3.5?

    When two friends each suggest a restaurant and you always go with the better rated one, your average dinner beats either friend's average. Taking the larger of two draws pulls the result toward the top, because a low result survives only if both draws are low. A payout of 1 needs both dice on 1, one cell in 36. A payout of 6 needs just one six, and 11 cells in 36 contain at least one.

    The larger face is k in 2k - 1 of the 36 cells: L-shaped bands123456223456333456444456555556666666112233445566Die 1Die 2Larger faceCellsFace x cellsk = 111k = 236k = 3515k = 4728k = 5945k = 61166Total36161161 / 36 = 4.47one die alone: 3.50
    The larger face is 1 in one cell, 2 in three cells and so on up to 6 in eleven cells; face times count sums to 161, so the expected payout is 161/36, about 4.47, against 3.50 for one die.

    How do you count the cells without listing all 36?

    Count the outcomes where the larger face is at most k: both dice must be at most k, which is k squared cells. The cells where the larger face is exactly k are k squared minus (k - 1) squared, which is 2k - 1. That is the L-shaped band in the grid: a new row and a new column, sharing one corner cell. The bands are 1, 3, 5, 7, 9 and 11, and they add to 36, which is the check that nothing was double counted.

    The relationship
    E[max⁡]=∑k=16k 2k−136=16136≈4.47E[\max] = \sum_{k=1}^{6} k\,\frac{2k-1}{36} = \frac{161}{36} \approx 4.47
    kthe value of the larger face
    2k - 1the number of the 36 outcomes where the larger face is exactly k
    What it says in wordsWeight each possible payout by how many of the 36 outcomes produce it, then divide by 36.

    A second route helps when the interviewer changes the dice. Add up the chance that the payout reaches each level: the payout is at least k unless both dice are below k, so the sum of 1 minus (k - 1) squared over 36, for k from 1 to 6, is 6 minus 55/36, which is 161/36 again. Two methods landing on the same fraction is the check worth saying out loud. By symmetry the smaller face averages 7 minus 4.47, about 2.53, and with three dice the larger face rises to 4.96.

    Where candidates lose it

    The common slip is to treat the six payouts as equally likely and answer 3.5, or to say a bit more than 3.5 without a number. The grid shows how uneven the counts are: eleven ways to be paid 6 against one way to be paid 1.

    The second slip is counting 12 cells for a payout of 6, which counts the double six twice. The row of sixes and the column of sixes share one cell.

    What the interviewer asks next

    • What is the expected value of the smaller face?
    • What is the expected larger face with three dice?
    • I pay you the larger face minus the smaller. What is that worth?

    Asked at Jane Street, Investment Operations, New York, 2026 (Wall Street Oasis): First interview was testing simple math brainteasers (e.g. expected value of dice throws, etc.)

← PreviousPage 1 of 4
  1. 1
  2. 2
  3. …
  4. 4
Next →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.