Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant puzzles, solved step by step

Puzzles
100
Traced to a firm
71
Topics
12
Hard
30
Topic
All topicsLogic and algorithmic reasoning10Conditional probability and Bayes7Counting and combinatorics8Continuous and geometric probability9Correlation, regression and linear algebra9Market making, betting and sizing9Expected value and optimal stopping9Statistics and estimation9Pricing, options and index maths7Games and strategic reasoning8Markov chains and random walks7Mental maths and number sense8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 61–70 of 71 · filtered from 100Clear filters
  1. 087A new strategy has won on 9 of its first 10 trading days. With no prior knowledge, you treat its daily win rate as uniform between 0 and 1. What is the probability that it wins tomorrow, and why is the answer not 0.9?Conditional probability and BayesHardWolverine TradingChicago · 2017

    Try it first

    What probability do you give to a win tomorrow?

    Show the worked solution

    10/12, about 0.833. A uniform prior on the win rate, updated with 9 wins and 1 loss, gives a Beta(10, 2) posterior. The chance of winning tomorrow is the posterior mean, (9 + 1)/(10 + 2). The answer is below 0.9 because ten days cannot rule out a lower true win rate, and averaging over that uncertainty pulls the estimate towards one half.

    Why is 0.9 too confident?

    A new restaurant with nine five-star reviews out of ten looks excellent, but you would not bet that the next diner rates it five stars with 90% certainty. Ten reviews is a small sample, and a restaurant that truly earns five stars 70% of the time could easily post nine out of ten. The right forecast averages over every win rate the evidence still allows, and with only ten days that includes plenty of rates below 0.9. The average of that spread of possibilities is what you should quote for tomorrow.

    After 9 wins in 10, the win rate is still a wide curve, centred at 0.833flat prior before any data0.00.20.40.60.81.0daily win rate pmean 10/12 = 0.833peak 0.9middle 90%Naive estimate9/10 = 0.900the peak, not the meanRule of succession(9 + 1)/(10 + 2)= 0.833
    After 9 wins in 10 days from a flat prior, the win rate follows a Beta(10, 2) curve that peaks at 0.9 but has a long left tail, so its mean, the chance of winning tomorrow, is 10/12, about 0.833, and its middle 90% still spans 0.64 to 0.97.

    How does the Bayesian update give 10/12?

    Start with every win rate p between 0 and 1 equally likely. The chance of the observed record is proportional to p^9 (1 - p), so the posterior is proportional to that, which is the Beta(10, 2) distribution. The chance of a win tomorrow is the average of p over the posterior, and the mean of a Beta(a, b) is a/(a + b), here 10/12. The shortcut is Laplace's rule of succession: add one imaginary win and one imaginary loss to the record, then divide.

    The relationship
    π(p∣data)∝p9(1−p)  =  Beta(10,2),P(win tomorrow)=E[p]=9+110+2=56\pi(p\mid\text{data}) \propto p^{9}(1-p) \;=\; \text{Beta}(10,2), \qquad P(\text{win tomorrow}) = \mathbb{E}[p] = \frac{9+1}{10+2} = \frac{5}{6}
    pthe unknown daily win rate
    p^9 (1 - p)the likelihood of 9 wins and 1 loss
    Beta(10, 2)the posterior after a flat prior
    What it says in wordsMultiply the flat prior by the likelihood of the record, and the mean of the result is the chance of a win tomorrow.

    When does the pull towards one half stop mattering?

    When the record is long. At 90 wins out of 100 the rule gives 91/102, about 0.892, almost exactly the raw 0.9, because a hundred days of data swamp the one imaginary win and loss. The limitation is the prior itself. A uniform prior says a 99% win rate was as plausible as a 50% one before you saw any data, which no trader believes about a new strategy. With a sceptical prior centred near one half, ten days would pull the estimate even further below 0.9. The posterior also tells you more than one number: the chance of winning both of the next two days is (10 x 11)/(12 x 13), about 0.705, not 0.833 squared.

    Where candidates lose it

    The trap is answering 0.9, the maximum likelihood estimate. It treats the observed win rate as the truth and ignores how little ten days can tell you.

    The second loss is getting 10/12 by the rule of succession without being able to say where it comes from. Name the flat prior, the Beta(10, 2) posterior and its mean, and say that more data pushes the answer back towards 0.9.

    What the interviewer asks next

    • What is the probability the strategy's true win rate is above one half?
    • With a Beta(5, 5) prior instead of a flat one, what is your forecast for tomorrow?
    • How many consecutive wins would you need before your forecast exceeds 0.95?

    Asked at Wolverine Trading, Prop Trading, Chicago, 2017 (Wall Street Oasis): Phone interviews were pretty standard brainteasers and fit questions. There was a Bayes question

  2. 088A strategy's true annualised Sharpe ratio is 1.0. Roughly how many years of returns do you need before a t-test rejects a zero mean at about the 5% level? What if the Sharpe ratio is 0.5?Statistics and estimationCoreViking Global InvestorsNew York · 2014

    Try it first

    How many years does a Sharpe 0.5 strategy need?

    Show the worked solution

    About 4 years at a Sharpe of 1, and about 16 years at 0.5. The t-statistic for a mean return is the Sharpe ratio times the square root of the number of years, so reaching t = 2 needs (2 / SR)^2 years. Halving the Sharpe ratio quadruples the evidence you need, and sampling daily instead of yearly does not shorten it.

    Why does the t-statistic come out as Sharpe times root years?

    A t-test on a mean divides the average return by its standard error, which is the volatility over the square root of the number of observations. With annual observations that ratio is (mean / volatility) x sqrt(years), and mean over volatility is exactly the annual Sharpe ratio. So a Sharpe of 1 gives t = sqrt(years): 2 after 4 years. A Sharpe of 0.5 gives t = 0.5 x sqrt(years): 2 only after 16 years. With the textbook 1.96 in place of 2 the numbers are 3.8 and 15.4 years; the round figures are what you say in the room.

    Halve the Sharpe ratio and you need four times as many yearst = 2, roughly the 5% bar4 years16 yearsSharpe 1.0Sharpe 0.5051015201234t-statisticyears of returnst = SR x sqrt(years)years =(2 / SR)^2SR 1: 4SR 0.5: 16
    The t-statistic equals the Sharpe ratio times the square root of years, so a Sharpe of 1 reaches t = 2 after 4 years and a Sharpe of 0.5 only after 16: halving the Sharpe quadruples the track record you need.
    The relationship
    t=rˉσ/N=SR⋅Nyears    ⇒    Nyears=(2SR)2t = \frac{\bar r}{\sigma/\sqrt{N}} = \text{SR}\cdot\sqrt{N_{\text{years}}} \;\;\Rightarrow\;\; N_{\text{years}} = \left(\frac{2}{\text{SR}}\right)^2
    r barthe average annual return in excess of cash
    sigmathe annual volatility
    Nthe number of years observed
    SRthe annual Sharpe ratio, r bar over sigma
    What it says in wordsThe evidence for a real edge grows with the square root of time, scaled by the Sharpe ratio.

    Can you shortcut it with daily data?

    This is the follow-up that separates candidates. Sampling daily gives about 252 times as many observations a year, but the daily Sharpe ratio is smaller by the square root of 252, because daily mean scales with time and daily volatility with its square root. The two effects cancel exactly, so the t-statistic depends on calendar time, not on how finely you slice it. Think of estimating a river's average level: measuring every minute instead of every day does not help if the river's slow swings are the uncertainty.

    What makes the real requirement even longer?

    Three things, each worth one sentence. Returns are not independent from year to year, and positive autocorrelation inflates the true standard error. If you tested twenty strategies and kept the best, a t of 2 is easy to get by luck, so the bar has to rise with the number of ideas tried. And the Sharpe ratio itself drifts as markets change, so a sixteen-year record may be measuring two different strategies. A desk that says a Sharpe 0.5 strategy is proven after three years is reading noise.

    Where candidates lose it

    The first loss is scaling linearly: a Sharpe half as big needs twice as long, so 8 years. The t-statistic grows with the square root of time, so the years scale with the square of 1/Sharpe, and 16 is right.

    The second loss is proposing daily data as the fix. The number of observations goes up but the per-observation Sharpe goes down by the square root of that factor, and the two cancel.

    What the interviewer asks next

    • How many years does a Sharpe of 2 need, and why do high-frequency desks care?
    • If you tested 50 strategies, roughly what t-statistic would you demand of the best one?
    • How does positive autocorrelation in monthly returns change the answer?

    Asked at Viking Global Investors, Quantitative Research, New York, 2014 (Wall Street Oasis): how to reject a hypothesis test, what's your structure of your code, what's the sample size

  3. 089Make me a market on the number of disposable nappies used in the UK in one day. Build the estimate from stated assumptions and choose a width you would actually trade on.Market making, betting and sizingCoreDRWLondon · 2025

    Try it first

    If each of four inputs could be about 10 to 25% off in either direction, how uncertain is the product?

    Show the worked solution

    About 9.4 million a day, and I would open at 8 bid, 11 offered, in millions. Assume about 700,000 births a year, 2.5 years in nappies, six changes a day and 90% disposable: 9.45 million. Multiplying the low and high ends of each input gives 5.5 to 15.0 million, so a quote of 8 at 11 is tight enough to trade and still honest about the uncertainty.

    How do you build the estimate so the interviewer can follow it?

    Chain it through things you can reason about. Children in nappies are roughly births a year times the years each child spends in them. Assume about 700,000 births a year, a round number worth checking against the latest official statistics, and 2.5 years in nappies: about 1.75 million children. Each child uses about six a day on average, more as a newborn and fewer as a toddler, and assume 90% of families use disposables: 1.75 million x 6 x 0.9 = 9.45 million a day. Say each assumption out loud and give it a range as you go.

    Multiply the ranges, not just the central guessesBirths a yearcentral 700krange 650k to 750kassumedxYears in nappiescentral 2.5range 2 to 3birth to toilet trainingxChanges a daycentral 6range 5 to 7newborns use morexDisposable sharecentral 90%range 85% to 95%the rest use clothNappies a day: low 5.5mcentral 9.45mhigh 15.0m5.5m15.0mbid 8moffer 11m9.45mlog scale: the multiplied range runs about 1.6 times either side of the centre; the quote sits inside it
    Multiplying the four central assumptions gives 9.45 million nappies a day, but multiplying the four lows and the four highs gives 5.5 to 15.0 million, so the honest uncertainty is roughly a factor of 1.6 either side, and a quote of 8 at 11 million sits inside it.
    The relationship
    N=B×Y×c×d=700,000×2.5×6×0.9≈9.45 millionN = B \times Y \times c \times d = 700{,}000 \times 2.5 \times 6 \times 0.9 \approx 9.45\text{ million}
    Bbirths a year, an assumption
    Yyears a child spends in nappies
    cchanges a day
    dshare of families using disposables
    What it says in wordsBuild the count from quantities you can defend one at a time, and multiply.

    Where should the width of the market come from?

    From the ranges, multiplied. A shopkeeper who is unsure of both price and quantity is more unsure of revenue than of either. Put a low and a high on every input and multiply the lows together and the highs together: here 5.5 million to 15.0 million, about a factor of 1.6 either side of the centre. Centre the quote near the middle on a multiplicative scale, the geometric mean of the ends, 9.1 million, which sits close to the central estimate.

    Then choose the width you will actually trade. A market as wide as the whole range, 5.5 at 15, is useless: nobody trades against it and it tells the interviewer you have no view. Quote tighter, 8 at 11, and move it as they trade: if they keep buying at 11, raise both sides, because their trades carry information. Say the scope questions too: does the count include adult incontinence products, and a school-age child in night-time pants? Those can move the answer more than any of the four inputs.

    Where candidates lose it

    The first loss is giving one number, or a market whose width is a round guess such as plus or minus a million, with no link to the assumptions. The interviewer wants to see where the width came from.

    The second loss is the opposite: a market so wide it is safe and worthless. Show the full range, then quote a tighter two-way price and explain how you would move it when they trade.

    What the interviewer asks next

    • I buy 5 lots at your offer. Where is your market now?
    • What single piece of data would you buy to narrow the range most, and why?
    • How would you size the market if the settlement were a count of nappies sold rather than used?

    Asked at DRW, Trading, London, 2025 (Wall Street Oasis): Make me a market on the amount of diapers used in the UK daily

  4. 090X and Y are independent random variables with the same variance. What is the correlation between X and X + Y?Correlation, regression and linear algebraWarm upSCSquarepoint CapitalMontreal · 2026

    Try it first

    Pick the correlation.

    Show the worked solution

    1/sqrt(2), about 0.71. The covariance of X with X + Y is Var X plus Cov(X, Y), and the second term is zero, so it is sigma squared. The variance of X + Y is 2 sigma squared, because independent variances add. Dividing sigma squared by sigma times sqrt(2) sigma leaves 1/sqrt(2). Squared, that is 0.5: X explains half of the sum's variance.

    Why isn't the answer one half?

    Picture two people each tossing a coin for a rupee, and a pot holding their combined winnings. One player's result explains exactly half of the pot's variability, and the other half comes from the other player. Half is the share of variance explained, R squared, and correlation is its square root, so the correlation is 1/sqrt(2), not 1/2. This is the most common slip on the question, and it comes from mixing up the two measures.

    X supplies half of the variance of X + Y, so the correlation is 1/sqrt(2)XYXYVar Xsigma^2Cov(X, Y)0Cov(Y, X)0Var Ysigma^2X row:Cov(X, X+Y)= sigma^2all four cells: Var(X + Y) = 2 sigma^2corr = Cov / (sd X x sd(X+Y))= sigma^2 / (sigma x sqrt(2) sigma)= 1/sqrt(2) = 0.707Share of Var(X + Y) explained by XR^2 = 0.5Unequal variances: corr = 1/sqrt(1 + k),k = Var Y / Var X; k = 4 gives 0.447
    In the covariance box, X's own variance fills one of the two non-zero cells, so X accounts for half of Var(X + Y), and the correlation between X and X + Y is sigma squared divided by sigma times sqrt(2) sigma, which is 1/sqrt(2), about 0.71.
    The relationship
    ρ=Cov⁡(X,X+Y)σX σX+Y=σ2+0σ⋅2 σ=12≈0.707\rho = \frac{\operatorname{Cov}(X, X+Y)}{\sigma_X\,\sigma_{X+Y}} = \frac{\sigma^2 + 0}{\sigma\cdot\sqrt{2}\,\sigma} = \frac{1}{\sqrt 2} \approx 0.707
    Cov(X, X + Y)Var X plus Cov(X, Y), which is sigma squared plus zero
    sigma_{X+Y}the standard deviation of the sum, sqrt(2) sigma
    What it says in wordsCovariance is linear, so split it into pieces; the only surviving piece is X's own variance.

    How does it change if the variances differ?

    Let Var Y be k times Var X. The covariance is still Var X, and Var(X + Y) becomes (1 + k) Var X. The correlation is 1/sqrt(1 + k): the noisier Y is, the less the sum tracks X. At k = 1 you get 0.707; at k = 4, 0.447; at k = 0.25, 0.894. This is exactly the signal-plus-noise model: if a price move is a true signal plus independent noise of equal size, the best-case correlation between your signal and the move is about 0.71.

    Where does this show up on a desk?

    Any time one piece is part of a total. A stock's return is market return plus its own specific return; if the two had equal variance, the stock would correlate 0.71 with the market. The same arithmetic tells you the ceiling on a predictor: if half of tomorrow's move is unpredictable noise, no model can correlate more than 0.71 with it. The limitation is independence; if X and Y are correlated, add 2 Cov(X, Y) to the variance of the sum and Cov(X, Y) to the covariance.

    Where candidates lose it

    The frequent slip is answering one half, confusing the share of variance with the correlation. Correlation is the square root of that share.

    The second loss is writing the standard deviation of X + Y as 2 sigma, adding standard deviations instead of variances, which gives one half again by a different road. Independent variances add; standard deviations do not.

    What the interviewer asks next

    • What is the correlation between X + Y and X - Y?
    • If X and Y have correlation 0.5 and equal variance, what is corr(X, X + Y)?
    • What is the correlation between the first die and the total of two dice?

    Asked at Squarepoint Capital, Desk Quant Analyst Interview, Montreal, 2026 (Wall Street Oasis): There were also 3-4 basic math/stats questions about mean, covariance, correlation, etc.

  5. 092You climb a staircase of 10 steps, taking either one step or two steps at a time. In how many different ways can you reach the top?Counting and combinatoricsWarm upTower Research CapitalNew York · 2012

    Try it first

    Pick the number of ways.

    Show the worked solution

    89 ways. Split by the last move: you arrive at step 10 with a single from step 9 or a double from step 8, so ways(10) = ways(9) + ways(8). With 1 way to reach step 1 and 2 ways to reach step 2, the counts run 1, 2, 3, 5, 8, 13, 21, 34, 55, 89: the Fibonacci numbers, with 89 at the top.

    How do you count without listing every route?

    Imagine a friend at the top of the stairs asks how you got there. There is one thing you can say for certain about your last move: it was either a single from step 9 or a double from step 8, and never both. Every route to step 10 is a route to step 9 followed by a single, or a route to step 8 followed by a double, so the count at step 10 is the sum of the counts at steps 9 and 8. The same holds at every step, which turns the puzzle into a running sum.

    Ways to reach each step: the sum of the two steps below1step 12step 23step 35step 48step 513step 621step 734step 855step 989step 10dashed: 34 ways end with a double from step 8solid: 55 ways end with a single from step 9step 10: 34 + 55 = 89start: 1 way to step 1, 2 ways to step 2 (1+1 or 2)
    Writing the number of ways on each step, each step is the sum of the two below it, so the counts follow the Fibonacci sequence and step 10 collects 55 routes ending with a single and 34 ending with a double, 89 in all.
    The relationship
    w(n)=w(n−1)+w(n−2),w(1)=1,  w(2)=2  ⇒  w(10)=89w(n) = w(n-1) + w(n-2), \qquad w(1) = 1,\; w(2) = 2 \;\Rightarrow\; w(10) = 89
    w(n)the number of ways to reach step n
    w(n-1)routes whose last move is a single step
    w(n-2)routes whose last move is a double step
    What it says in wordsSort every route by its last move; the two groups do not overlap and together cover everything.

    Can you check 89 a second way?

    Count by how many double steps you take. With k doubles, you make 10 - 2k singles, so 10 - k moves in all, and you only have to choose which k of those moves are the doubles. Summing the binomial counts over k = 0 to 5 gives 1 + 9 + 28 + 35 + 15 + 1 = 89, the same answer by a completely different road. Saying a second check aloud is worth more than the answer itself in a first round, because it shows you do not trust a pattern you have not tested.

    Double steps kMoves in totalWays to place the doubles
    010C(10, 0) = 1
    19C(9, 1) = 9
    28C(8, 2) = 28
    37C(7, 3) = 35
    46C(6, 4) = 15
    55C(5, 5) = 1
    total 89
    Counting routes by the number of double steps gives 89 again, which confirms the Fibonacci running sum.

    What does the interviewer usually ask next?

    Two things. First, allow steps of one, two or three: the same last-move argument gives w(n) = w(n - 1) + w(n - 2) + w(n - 3), and the count for ten steps becomes 274. Second, the coding version. A recursive function that calls itself for n - 1 and n - 2 recomputes the same steps again and again: for 30 steps it makes 1,664,079 calls to return 1,346,269. Storing each step's count once, or just keeping the last two numbers in a loop, does the job in 30 additions. The counts grow by about 1.618, the golden ratio, per step, which is the limitation of any approach that lists routes rather than counting them.

    Where candidates lose it

    The fast wrong answer is 2^10 = 1,024, treating each of ten stairs as a binary choice. A double step consumes two stairs, so routes have different numbers of moves and the choices are not ten independent coin flips.

    The second loss is an off-by-one in the starting values, which lands on 55 or 144. Write the first three steps out by hand: 1 way to step 1, 2 ways to step 2, 3 ways to step 3. Anchor the sequence there and the tenth term is 89.

    What the interviewer asks next

    • What if you can also take three steps at a time?
    • How many ways are there if step 5 is broken and cannot be stood on?
    • Write code that counts the ways for 1,000 steps without the recursion blowing up.

    Asked at Tower Research Capital, Intern Interview -, New York, 2012 (Wall Street Oasis): How many ways can you jump up stairs if you can only jump either 1 or 2 steps?

  6. 093Three dice: red has faces 2, 6 and 7; green has 1, 5 and 12; blue has 3, 4 and 8, each face appearing twice. You and I each pick a die and roll once, and the higher number wins. Which die do you want, and does it matter who picks first?Games and strategic reasoningCoreBelvedere TradingChicago · 2022

    Try it first

    Which die is best against the other two?

    Show the worked solution

    No die is best: red beats green, green beats blue and blue beats red, each with probability 5/9. So who picks first matters a great deal. Let me choose, then take the die that beats mine and win 5/9 of the time. If you are forced to pick first, every choice loses 5/9 of the time against an opponent who knows the cycle.

    How do you work out who beats whom?

    Each matchup has only nine equally likely pairs of faces, so write the 3 by 3 grid and count. Red against green: red's 2 beats only the 1, while its 6 and 7 each beat the 1 and the 5, for 1 + 2 + 2 = 5 wins out of 9. Do the same for the other two pairs and every matchup comes out 5 to 4: red over green, green over blue, blue over red. It is rock, paper, scissors built out of dice, and in rock, paper, scissors nobody asks which hand shape is best.

    Each die wins 5 of 9 face pairs against the next: a cycle, not a rankingRed (rows) vs Green (columns)15122winloselose6winwinlose7winwinloseRed wins 5 of 9Green (rows) vs Blue (columns)3481loseloselose5winwinlose12winwinwinGreen wins 5 of 9Blue (rows) vs Red (columns)2673winloselose4winloselose8winwinwinBlue wins 5 of 9Red, mean 5Green, mean 6Blue, mean 55/95/9Blue beats Red 5/9: back to the start, so let your opponent choose first
    Counting the nine face pairs in each matchup shows red beats green, green beats blue and blue beats red, each in 5 of 9 cases, so the three dice form a cycle and the second player can always pick a die that wins 5/9 of the time.
    The relationship
    P(R>G)=1+2+29,P(G>B)=0+2+39,P(B>R)=1+1+39, each =59P(R>G) = \tfrac{1+2+2}{9},\quad P(G>B) = \tfrac{0+2+3}{9},\quad P(B>R) = \tfrac{1+1+3}{9}, \text{ each } = \tfrac59
    P(R > G)the chance red's roll beats green's
    1 + 2 + 2the wins for red's faces 2, 6 and 7 in turn
    What it says in wordsCount, face by face, how many of the opponent's three faces each face beats, and divide by nine.

    Why does green lose to red when green has the higher average?

    The averages are red 5, green 6 and blue 5. Winning is about how often, not by how much. Green's 12 wins every time it shows, but it shows only a third of the time, and green's other two faces, 1 and 5, lose to both of red's high faces. A higher mean and a higher chance of winning are different things, and the gap between them is the whole puzzle. Change the rules so the winner collects the difference between the two numbers, and green's expected margin against red is 6 - 5 = +1: now you want green against red, and blue against red is a dead heat at 0.

    Where does a trader meet the same thing?

    Head-to-head comparisons need not line up into a ranking. Strategy A can beat strategy B on more days than not, B can beat C, and C can beat A, whenever one of them earns its money in rare large wins, as green does. Before choosing between strategies, decide whether you care about how often you win or how much you make, because under the first a cycle like this one means there may be no best choice at all. The limitation of the puzzle is that it is one roll; over many rolls with the total score counted, the mean matters more and green's 12 starts to pay.

    Where candidates lose it

    The common slip is choosing green because its average, 6, is highest. The game pays for winning, not for margin, and green loses to red five times in nine.

    The second loss is answering the first question and missing the second. Because the dice form a cycle, the real answer is strategic: insist that your opponent picks first. Saying that unprompted is what the interviewer is listening for.

    What the interviewer asks next

    • If each player rolls their die twice and adds the results, does the cycle still hold?
    • Design three dice whose faces sum to the same total and still form a cycle.
    • With three players each taking one die, can any die be favoured against both others?

    Asked at Belvedere Trading, Prop Trading, Chicago, 2022 (Wall Street Oasis): You have 3 dice: red has 2, 6, 7; green has 1, 5, 12; blue has 3, 4, 8.

  7. 094Two points are chosen independently and uniformly on the surface of a unit sphere. What is the expected distance between them measured along the surface, that is, the great-circle distance?Continuous and geometric probabilityHardTower Research CapitalNew York · 2019

    Try it first

    What is the expected great-circle distance?

    Show the worked solution

    pi/2, about 1.571. Rotate the sphere so the first point sits at the north pole; nothing changes, because the second point is uniform. On a unit sphere the surface distance is the polar angle theta of the second point. The northern and southern hemispheres are mirror images, so theta is as likely to be pi/2 - t as pi/2 + t, and its mean is pi/2.

    Why can you put the first point at the pole?

    Ask how far apart two random towns are on a perfectly round planet, and you can simply stand in one of them: the globe looks the same from every spot on it. Symmetry lets you fix one point anywhere, so the problem shrinks to one random point and its angle from the pole. On a sphere of radius 1, the distance along the surface between the pole and a point at polar angle theta is theta itself, measured in radians, so the question becomes: what is the average polar angle of a uniform point?

    Fix one point at the pole: the angle to the other is symmetric about pi/2thetafirst point, moved to the polesecond pointsurface arc = thetachord2 sin(theta/2)dashed red: the wrong uniform angle, density 1/pimean = pi/2,by symmetry0pi/2piangle theta from the poledensity of theta: (1/2) sin theta0.5
    With the first point at the pole, the surface distance is the polar angle theta of the second point, whose density (1/2) sin theta is symmetric about pi/2, so the expected distance is pi/2; a uniform angle, the dashed line, puts too many points near the poles.
    The relationship
    f(θ)=12sin⁡θ    (0≤θ≤π),E[θ]=∫0πθ 12sin⁡θ dθ=π2f(\theta) = \tfrac12 \sin\theta \;\; (0 \le \theta \le \pi), \qquad \mathbb{E}[\theta] = \int_0^{\pi} \theta\, \tfrac12 \sin\theta \, d\theta = \frac{\pi}{2}
    thetathe polar angle of the second point, equal to the surface distance on a unit sphere
    f(theta)the density of that angle
    sin thetathe relative size of the band of latitude at angle theta
    What it says in wordsThere is more surface near the equator than near the poles, in proportion to sin theta, and that density is symmetric about pi/2.

    Where does sin theta come from? The circle of latitude at angle theta from the pole has circumference 2 pi sin theta, so a thin band there holds surface in proportion to sin theta: almost none near the poles, the most at the equator. Archimedes put it more neatly: the area of a band is proportional to its height along the axis, so cos theta is uniform between -1 and 1. The density (1/2) sin theta is a mirror image about pi/2, so the mean is pi/2 without doing the integral. Integration by parts confirms it, a numerical integral gives 1.5708, and a seeded simulation of 100,000 pairs gives 1.570.

    If a uniform angle gives the same mean, why does the shape matter?

    Because the mean survives by luck of symmetry and almost nothing else does. Choosing theta uniformly on 0 to pi crowds points near the poles. Ask for the chance the two points are within 60 degrees of each other and the correct answer is (1 - cos 60 degrees)/2 = 0.25, while the uniform angle says 0.33. Ask for the expected straight-line chord, 2 sin(theta/2), and the correct density gives 4/3, about 1.333, while the uniform angle gives 4/pi, about 1.273. The simulation gives 1.333 for the chord. This is the limitation of the shortcut: it answers this one question and must not be reused for the next.

    Where candidates lose it

    The commonest wrong answer is 4/3, the expected straight-line chord, which some candidates remember from a related puzzle. The question asks for distance along the surface, which on a unit sphere is the angle itself.

    The second loss is the right answer for the wrong reason: picking the angle uniformly between 0 and pi. The mean comes out right by symmetry, but any follow-up on the chord or on the chance of being close gives the wrong number. Say that the band of latitude grows like sin theta.

    What the interviewer asks next

    • What is the expected straight-line distance between the two points?
    • What is the probability that the two points are within 60 degrees of each other?
    • Four points are chosen uniformly on a sphere. What is the chance they all lie in one hemisphere?

    Asked at Tower Research Capital, Quantitative Research, New York, 2019 (Wall Street Oasis): a 3d geometry question about the surface distance between points chosen randomly on the surface of a sphere

  8. 096A desk's daily P&L in Rs lakh over seven days is -1, 2, 4, -9, 8, -2, 3. Which run of consecutive days has the largest total, and how do you find it in one pass through the data?Logic and algorithmic reasoningCoreWolverine TradingChicago · 2014

    Try it first

    Which run has the largest total?

    Show the worked solution

    Days 5 to 7, the run 8, -2, 3, which totals Rs 9 lakh. Walk through the days keeping the best total of a run ending today: either today alone or today added to yesterday's best run, whichever is larger. Record the largest value you see. The run 2, 4 looks attractive but totals only 6, and the -9 day makes it pointless to carry anything before it.

    How do you find the best run without checking every start and end day?

    Picture walking along a road with toll booths that either pay you or charge you. You may choose where to start and stop collecting. If the purse you have carried from earlier booths is in the red, the sensible move is to drop it and start fresh at the next booth. A run ending today is worth extending from yesterday only if the best run ending yesterday is positive; if it is negative it can only drag today down, so today starts a new run. That rule looks at each day once. Checking every pair of start and end days means 28 runs for seven days and 31,375 for a trading year of 250 days.

    Carry yesterday's run only while it is positive: the best run totals 9-1+2+4-9+8-2+3P&L, Rs lakhbest run: 8 - 2 + 3 = 9tempting: 2 + 4 = 6best ending today-126-3869best so far-1266889day1234567restart: yesterday's -3 is dropped
    Carrying the best run forward only while it is positive, the running total drops to -3 after the loss of 9 and restarts at 8 on day 5, so the best run is 8, -2, 3 with a total of 9, ahead of the tempting run 2, 4 at 6.
    The relationship
    ct=max⁡(xt,  ct−1+xt),best=max⁡tctc_t = \max(x_t,\; c_{t-1} + x_t), \qquad \text{best} = \max_t c_t
    x_tthe P&L on day t
    c_tthe best total of a run that ends on day t
    bestthe largest c_t seen so far
    What it says in wordsThe best run ending today either starts today or extends the best run ending yesterday; keep whichever is bigger, and remember the biggest.
    DayP&LBest run ending todayBest so far
    1-1-1-1
    2222
    3466
    4-9-36
    5888
    6-268
    7399
    Running the rule day by day, the best run ending today drops to -3 after day 4, restarts at 8 on day 5 and reaches 9 on day 7, which is the answer.

    Why does the tempting run 2, 4 lose?

    Because one later day beats it on its own. After the -9, the best run ending on day 4 is 6 - 9 = -3, so the rule drops the past and day 5 starts fresh at 8. The -2 on day 6 dips the run to 6, but the 3 on day 7 lifts it to 9. The best run can contain a losing day: 8, -2, 3 beats 8 alone because the day after the loss more than repays it. A candidate who stops a run at the first red day misses this, and the brute-force check over all 28 runs confirms 9 is the maximum.

    What edge cases does the interviewer probe?

    Three. If every day is a loss, the answer should be the least bad single day, so start the best at the first day's value, not at zero, or you will report an empty run worth 0. To report which days, store the start index whenever you restart and copy it when you record a new best. The same pass with the signs flipped finds the worst run, here -9, the single -9 day. The limitation on a desk is that the best run in hindsight is a selected statistic: a strategy that is judged by its best stretch will always look better than it trades.

    Where candidates lose it

    The quick wrong answer is the run 2, 4, because it is the first good stretch. The 8 on day 5 beats it alone, and carrying 8 through -2 and 3 beats 8.

    The second loss is in the code: starting the best total at zero, which reports 0 for a week of all losses, or restarting at every losing day instead of only when the running total itself turns negative. State the rule exactly: carry yesterday's run only while it is positive.

    What the interviewer asks next

    • Return the start and end days of the best run, not just its total.
    • What does your code return if every day in the series is a loss?
    • Find the best run if you may skip at most one day inside it.

    Asked at Wolverine Trading, Quantitative Research, Chicago, 2014 (Wall Street Oasis): Develop an algorithm to find out the section that contains the maximum sum.

  9. 098Two traders' monthly P&L are independent and normal. A has mean Rs 10 lakh and standard deviation Rs 3 lakh; B has mean Rs 8 lakh and standard deviation Rs 4 lakh. What is the probability that A out-earns B in a given month?Statistics and estimationCoreDRWLondon · 2025

    Try it first

    Pick the probability that A earns more than B in a month.

    Show the worked solution

    About 65.5%. The gap A - B is normal with mean 10 - 8 = Rs 2 lakh and variance 3^2 + 4^2 = 25, so its standard deviation is Rs 5 lakh. A out-earns B when the gap is positive, and zero sits 2/5 = 0.4 standard deviations below the mean, so the probability is Phi(0.4), about 65.5%. The better trader loses about one month in three.

    Why do the variances add when you subtract?

    You and a colleague set off for the same meeting from different places, and each journey is uncertain by a few minutes. The gap between your two arrival times is more uncertain than either journey, not less, because either of you can be the late one. Subtracting an independent random amount adds its noise, so Var(A - B) = Var A + Var B = 9 + 16 = 25, and the gap's standard deviation is 5, not 1. The mean subtracts as you would expect, 10 - 8 = 2. The gap is normal because a difference of independent normals is normal.

    Subtracting adds noise: the gap A - B has mean 2 and sd 5-505101520A: 10, sd 3B: 8, sd 4monthly P&L, Rs lakh-10-5051015A wins:65.5%B wins: 34.5%mean 2, sd sqrt(9 + 16) = 5gap A - B, Rs lakh; zero is 0.4 sd below the mean
    The two traders' monthly P&L overlap heavily, and the gap A - B has mean 2 and standard deviation 5, so the area above zero where A wins is only 65.5%, leaving B ahead in 34.5% of months.
    The relationship
    A−B∼N ⁣(μA−μB,  σA2+σB2)=N(2, 25),P(A>B)=Φ ⁣(25)=Φ(0.4)≈0.655A - B \sim N\!\left(\mu_A - \mu_B,\; \sigma_A^2 + \sigma_B^2\right) = N(2,\, 25), \qquad P(A > B) = \Phi\!\left(\frac{2}{5}\right) = \Phi(0.4) \approx 0.655
    mu_A, mu_Bthe mean monthly P&L, 10 and 8
    sigma_A, sigma_Bthe standard deviations, 3 and 4
    Phithe standard normal cumulative distribution
    What it says in wordsThe gap's mean is the difference of the means, its variance the sum of the variances, and the answer is how many standard deviations zero sits below that mean.

    How much does a longer comparison window help?

    A lot, and at a predictable rate. Over a quarter of independent months the total gap has mean 6 and standard deviation 5 x sqrt(3), about 8.7, so A comes out ahead with probability 75.6%. Over a year the mean is 24 and the standard deviation 5 x sqrt(12), about 17.3, so the probability is 91.7%. The edge grows with the number of months and the noise with its square root, so the z-score grows with the square root of time. A risk manager who ranks traders on one month of P&L is ranking mostly noise.

    What if the two traders' P&L are correlated?

    Then the shared part cancels in the gap. With correlation 0.5, the variance is 9 + 16 - 2 x 0.5 x 3 x 4 = 13, a standard deviation of 3.61, and A wins with probability 71.0%. Positive correlation makes the comparison sharper because common market moves drop out of the difference; negative correlation does the opposite. The limitation is the normal assumption: real P&L has fat tails and skew, and if one trader earns through rare large months, the month-by-month win rate can disagree with the mean, so check the shape before trusting the 65.5%.

    Where candidates lose it

    The commonest slip is subtracting the standard deviations, 4 - 3 = 1, which makes A look almost certain to win at 97.7%. Noise does not cancel when you subtract independent variables; it adds.

    The second loss is subtracting the variances, 16 - 9, or adding the standard deviations, 3 + 4. Square, add, then take the root: sqrt(9 + 16) = 5. The answer is then a z-score of 0.4, and Phi(0.4) is about 0.655.

    What the interviewer asks next

    • What is the probability that A out-earns B over a full year of independent months?
    • If their monthly P&L has correlation 0.5, what is the answer?
    • What is the probability that A out-earns B by more than Rs 5 lakh in a month?

    Asked at DRW, Trading, London, 2025 (Wall Street Oasis): technical interview based on normal distribution and market making

  10. 099A stock is worth either 100 or 110, with equal probability. 20% of the traders who arrive know the true value: they buy if it is 110 and sell if it is 100. The other 80% buy or sell at random, half and half. Where should a market maker set its ask so that it breaks even, on average, when someone buys from it?Market making, betting and sizingHardJane StreetNew York · 2025

    Try it first

    Where should the ask be?

    Show the worked solution

    Set the ask at 106, and by the same logic the bid at 104. If the stock is worth 110, a buy arrives with probability 0.2 + 0.8 x 0.5 = 0.6; if it is worth 100, with probability 0.4. Given a buy, Bayes puts the chance of 110 at 0.6, so the stock is worth 106 to the market maker selling it. The spread of 2 is the price of trading against informed flow.

    Why can't the market maker just quote the expected value of 105?

    A second-hand car dealer who pays the average price for every car will find that the owners of good cars go elsewhere and the owners of bad ones queue up. Who chooses to trade with you is information. A market maker does not care what the stock is worth on average; it cares what the stock is worth given that someone has just chosen to buy from it. At an ask of 105, noise buyers are harmless, a loss of 5 when the stock is worth 110 and a gain of 5 when it is worth 100. Informed buyers only appear in the 110 world, and they cost 0.5 per arriving trader on average, so 105 is a losing quote.

    A buy is evidence: given a buy, the stock is worth 106, so the ask is 106start0.5worth 110informed 0.2noise 0.80.5worth 100informed 0.2noise 0.8buy0.10buy0.20sell0.20sell0.10buy0.20sell0.20buys from 110:0.10 + 0.20 = 0.30buys from 100: 0.20P(110 | buy) = 0.6ask = 106
    Tracing who sends a buy order in each world, buys come with probability 0.30 from the 110 world and 0.20 from the 100 world, so a buy lifts the chance of 110 from 0.5 to 0.6 and the break-even ask is 106.
    The relationship
    P(110∣buy)=0.5×0.60.5×0.6+0.5×0.4=0.6,ask=E[V∣buy]=100+0.6×10=106P(110 \mid \text{buy}) = \frac{0.5 \times 0.6}{0.5 \times 0.6 + 0.5 \times 0.4} = 0.6, \qquad \text{ask} = \mathbb{E}[V \mid \text{buy}] = 100 + 0.6 \times 10 = 106
    0.6the chance of a buy when the stock is worth 110: 0.2 informed plus 0.8 x 0.5 noise
    0.4the chance of a buy when the stock is worth 100: noise only
    Vthe stock's true value
    What it says in wordsSet the ask at the value of the stock conditional on being bought from, which Bayes' rule gives directly.

    What sets the width of the spread?

    The share of informed traders and the size of what they know. With a share alpha informed, a buy is alpha + (1 - alpha)/2 likely in the high world and (1 - alpha)/2 in the low world, and the ask works out to 105 + 5 alpha. The spread is a fee for adverse selection: it is zero when nobody is informed and widens to the full 100 to 110 range when everyone is. The table runs the formula for a few shares. Order processing and inventory costs add to this in real markets, but the information component is what makes spreads jump around earnings and news.

    Informed shareAskBidSpread
    0%1051050
    10%105.5104.51
    20%1061042
    50%107.5102.55
    100%11010010
    The break-even spread equals the informed share times the 10-point value gap, so it is 2 at 20% informed and 5 at 50% informed.

    What happens after the first trade?

    The market maker updates. After one buy, the chance of 110 is 0.6, and if a second buy arrives the same Bayes step lifts it to 0.692, so the next ask is about 106.92. Each order moves the quotes towards the true value, which is how prices come to reflect what the informed traders know. The limitation is that the model has one share size, no inventory risk and no competition between market makers; real desks also skew quotes to manage position, which this puzzle leaves out.

    Where candidates lose it

    The fast wrong answer is 105, the unconditional expected value. It ignores that the act of buying is evidence: informed traders buy only when the stock is worth 110, so a market maker at 105 loses on every informed buyer and only breaks even on noise traders.

    The second loss is overreacting and quoting 110 because some buyers are informed. Most buyers are noise traders, and a quote at 110 drives them away. Bayes gives the exact weight, 0.6 on the high value, and the ask of 106.

    What the interviewer asks next

    • Where should the bid be, and why is the spread symmetric here?
    • After one buy at 106, where is the next ask?
    • How does the spread change if half of the traders are informed?

    Asked at Jane Street, Generalist, New York, 2025 (Wall Street Oasis): It was a probability theory based quant trading style market making questions which were intense

← PreviousPage 7 of 8
  1. 1
  2. …
  3. 6
  4. 7
  5. 8
Next →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.