Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant puzzles, solved step by step

Puzzles
100
Traced to a firm
71
Topics
12
Hard
30
Topic
All topicsLogic and algorithmic reasoning10Conditional probability and Bayes7Counting and combinatorics8Continuous and geometric probability9Correlation, regression and linear algebra9Market making, betting and sizing9Expected value and optimal stopping9Statistics and estimation9Pricing, options and index maths7Games and strategic reasoning8Markov chains and random walks7Mental maths and number sense8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 11–20 of 33 · filtered from 100Clear filters
  1. 035Regressing y on x gives a slope of 0.8; regressing x on y gives a slope of 0.45. What is the R-squared of either regression, and what is the correlation?Correlation, regression and linear algebraCoreTower Research CapitalNew York · 2014

    Try it first

    What is the correlation between x and y?

    Show the worked solution

    R-squared is 0.36 for both regressions and the correlation is 0.6. The slope of y on x is r times sd(y)/sd(x); the slope of x on y is r times sd(x)/sd(y). Multiplying them cancels the standard deviations and leaves r squared: 0.8 x 0.45 = 0.36. The correlation is +0.6, positive because both slopes are positive, and the ratio sd(y)/sd(x) is √(0.8/0.45) = 4/3.

    Why are the two slopes not reciprocals of each other?

    Tall parents tend to have tall children, but a little less tall; and tall children tend to have tall parents, but a little less tall. Both statements are true at once. Each regression predicts toward the mean, so neither slope is the inverse of the other unless the fit is perfect. If the points lay exactly on a line, the slope of x on y would be 1/0.8 = 1.25. It is 0.45 instead, and the size of that shortfall is what measures how loose the relationship is.

    Two regressions, two lines: the slopes multiply to R-squaredxyy on x: slope 0.8x on y: slope 0.45, drawn as 1/0.45 = 2.22Slope of y on x = r x (sd of y / sd of x)Slope of x on y = r x (sd of x / sd of y)Multiply the two: the sd ratios cancel0.8 x 0.45 = 0.36 = R-squaredr = +0.6Divide instead: sd of y / sd of x= √(0.8 / 0.45) = 1.333Neither slope is the reciprocal of the otherbecause |r| < 1: both regress toward the mean
    Fitting y on x gives the shallower line with slope 0.8 and fitting x on y gives the steeper line, slope 0.45 in its own terms; their product, 0.36, is R-squared, so the correlation is 0.6 and the standard deviation of y is 4/3 that of x.

    How do the two slopes give R-squared?

    Write each slope in terms of the correlation. The least squares slope of y on x is the covariance over the variance of x, which is r times sd(y)/sd(x). Swap the roles and the slope of x on y is r times sd(x)/sd(y). The standard deviation ratios are reciprocals, so the product of the two slopes is r squared, and in a one-variable regression r squared is exactly the R-squared. Here 0.8 x 0.45 = 0.36, so r = 0.6; the sign is positive because both slopes are positive, and the two slopes always share a sign.

    The relationship
    by∣x bx∣y=rsysx⋅rsxsy=r2=0.8×0.45=0.36b_{y|x}\, b_{x|y} = r\frac{s_y}{s_x}\cdot r\frac{s_x}{s_y} = r^2 = 0.8 \times 0.45 = 0.36
    b_y|xslope from regressing y on x, 0.8
    b_x|yslope from regressing x on y, 0.45
    s_x, s_ystandard deviations of x and y
    rthe correlation of x and y
    What it says in wordsThe two slopes multiply to the squared correlation because the scale factors cancel.

    The figure uses 40 points built with standard deviations 3 and 4 and a correlation of exactly 0.6, and fitting both regressions returns slopes of 0.80 and 0.45. A quick sanity test comes free: the product of the two slopes can never exceed 1. If an interviewer quotes slopes of 0.8 and 1.5, the product 1.2 is impossible, and saying so is worth more than any calculation.

    Where candidates lose it

    The fast wrong answer is to say the slopes should be reciprocals and call the data inconsistent, or to answer 0.36 when asked for the correlation. 0.36 is R-squared; the correlation is its square root.

    The second loss is dropping the sign. The square root of 0.36 could be plus or minus 0.6; both slopes are positive, so the correlation is positive, and saying why takes one sentence.

    What the interviewer asks next

    • What is the ratio of the standard deviation of y to that of x?
    • If the slope of x on y were 1.5, what would you conclude?
    • How does adding measurement noise to x change each slope?

    Asked at Tower Research Capital, Quantitative Research, New York, 2014 (Wall Street Oasis): Another detailed linear regression questions were asked, including problems about residual, variance and R^2

  2. 036A stock pays a growing dividend and is valued with the Gordon model at a discount rate of 10% and growth of 6%. What is its duration, and roughly how much does its price change if the discount rate rises by one point?Pricing, options and index mathsCoreBLBlackRockNew York · 2026

    Try it first

    What is the stock's duration, its percentage price sensitivity to the discount rate?

    Show the worked solution

    Duration is 1/(r - g) = 25 years, so a one-point rise cuts the value by about a fifth. The Gordon price is D1/(r - g), and its percentage sensitivity to r is 1/(r - g) = 1/0.04 = 25. With a Rs 4 dividend the price moves from Rs 100 at 10% to Rs 80 at 11%, a 20% fall. The 25% duration estimate overshoots because the price curve is convex.

    Why does a stock have a duration at all?

    A promise of money in one year hardly changes in value when rates move; a promise of money in twenty five years changes a lot, because the rate is compounded over every one of those years. A stock is a stream of dividends stretching forever, and when the dividends grow, most of its value sits in cash flows far in the future, so it behaves like a very long bond. Duration measures exactly that: the percentage price change for a change in the discount rate.

    The relationship
    P=D1r−g−1PdPdr=1r−g=10.10−0.06=25P = \frac{D_1}{r-g} \qquad -\frac{1}{P}\frac{dP}{dr} = \frac{1}{r-g} = \frac{1}{0.10-0.06} = 25
    D1next year's dividend, Rs 4 in the illustration
    rdiscount rate, 10%
    gdividend growth rate, 6%
    1/(r - g)percentage price change per unit change in r
    What it says in wordsDifferentiate the Gordon price and divide by price: the sensitivity is one over the gap between the discount rate and growth.
    Gordon price against the discount rate, with the tangent at 10%9%10%11%12%13%5075100125150175Discount rate rPrice, RsRs 100 at 10%Rs 80: actual, -20%Rs 75: tangent, -25%Rs 66.7 at 12%tangent Rs 50P = D1 / (r - g)Rs 4 / (0.10 - 0.06) = 100Duration = 1 / (r - g)25yearsOne point on r:duration says -25%the curve says -20%A 10-year 10% bond at 10%:duration about 6.1 years
    At a 10% discount rate the Rs 4 dividend stock is worth Rs 100 and its duration is 25; at 11% the price is Rs 80, a 20% fall, while the tangent line predicts Rs 75, and at 12% the gap widens to Rs 66.7 against Rs 50 because the price curve is convex.

    Why does the estimate say 25% when the price falls 20%?

    Duration is the slope at one point, and the price curve bends. For a one-point rise the tangent predicts a 25% fall, but the exact move from Rs 100 to Rs 80 is 20%, because the curve is convexCurving upward, so it always sits above any of its tangent lines. and flattens as r rises. The same bend makes a one-point fall worth more than 25%: at 9% the price is Rs 133.3, up 33.3%. For a big rate move, reprice exactly instead of trusting the slope. Duration here also equals price over dividend, 100/4, which is a quick way to say it: one over the dividend yield.

    For precision, the Macaulay durationThe present-value weighted average time at which cash flows arrive. is (1 + r)/(r - g) = 27.5 years, and dividing by 1 + r gives the modified duration of 25. A 10-year bond paying 10% at a 10% yield has a modified duration of about 6.1. The gap r - g is what matters, which is why high-growth stocks carry the most duration: the same stock with 2% growth would have a duration of 12.5 years. The limitation is that the Gordon model holds growth fixed while rates move; in practice both shift together.

    Where candidates lose it

    The usual loss is saying a stock has no duration because it has no maturity, or that it is infinite because it pays forever. Both skip the one line of calculus that gives 1/(r - g).

    The second is quoting 25% as the exact price change. It is the slope at 10%; the exact fall to 11% is 20%, and saying why, convexity, is what separates a strong answer.

    What the interviewer asks next

    • What happens to duration as growth approaches the discount rate?
    • Why might a stock's measured sensitivity to bond yields be much lower than 25?
    • What is the price change for a one-point fall in the discount rate?

    Asked at BlackRock, Restructuring, New York, 2026 (Wall Street Oasis): Which equities have duration ? multiple stocks vs value stocks MSE Forecasting equation

  3. 038Walking up a moving escalator at one step per second you take 20 steps; walking at two steps per second you take 32 steps. How many steps are visible on the escalator?Logic and algorithmic reasoningCoreSusquehanna International GroupNew York · 2026

    Try it first

    How many steps are visible?

    Show the worked solution

    80 steps. At one step a second the climb takes 20 seconds; at two steps a second it takes 16. If the escalator moves v steps a second, the visible steps are 20 + 20v and also 32 + 16v. Setting them equal gives v = 3, so the escalator is 20 + 60 = 80 steps long, and the check 32 + 48 = 80 agrees.

    What stays the same between the two walks?

    On an airport moving walkway, walk slowly and the belt does most of the work; stride out and you do more of it yourself, but you reach the end sooner. The length of the walkway does not change. Every visible step is covered either by your legs or by the escalator, so your steps plus the escalator's movement during your climb always equal the same total. That fixed total is the unknown; the escalator's speed is the second unknown, and two walks give two equations.

    Two walks, same escalator: your steps plus the escalator's always make 801 step/s, 20 syou: 20escalator: 20 s x 3 = 602 steps/s, 16 syou: 32escalator: 16 s x 3 = 4880 visible steps20 + 20v = 32 + 16v, so 4v = 12The faster walk saves 4 seconds of escalator help and pays for them with 12 extra stepsv = 3, N = 80
    Walking at one step a second you climb 20 steps in 20 seconds while the escalator carries 60; at two steps a second you climb 32 in 16 seconds while it carries 48; both add to the same 80 visible steps because the escalator moves 3 steps a second.

    How do you set up and solve the two equations?

    Turn step counts into time first, because the escalator's contribution depends on time. The slow walk: 20 steps at one a second is 20 seconds. The fast walk: 32 steps at two a second is 16 seconds. The faster walk loses 4 seconds of escalator help and makes it up with 12 extra steps of its own, so the escalator moves 3 steps a second. Then the total is 20 + 20 x 3 = 80, and 32 + 16 x 3 = 80 confirms it.

    The relationship
    N=20+20v=32+16v  ⇒  v=3, N=80N = 20 + 20v = 32 + 16v \;\Rightarrow\; v = 3,\ N = 80
    Nvisible steps on the escalator
    vescalator speed, in steps per second
    20, 16seconds taken on the slow and fast walks
    What it says in wordsThe same number of visible steps is covered on both walks, split differently between you and the machine.

    Say the check aloud, then the sense check: the escalator at 3 steps a second is faster than either walking pace, which is plausible for a long escalator. If the question had you walking down an up escalator, the escalator's steps would subtract instead of add, and the same method still works. The trap in variants is mixing up steps and seconds; keep one unit for each quantity.

    Where candidates lose it

    The usual loss is treating the step counts as if they were times, writing 20 + 20v = 32 + 32v, or averaging 20 and 32. The escalator helps for as long as you are on it, and the fast walk is shorter: 16 seconds, not 32.

    The second is solving for the speed and stopping. The question asks for the visible steps; plug back in and check both walks give 80.

    What the interviewer asks next

    • How long does the climb take if you stand still?
    • You now walk down the same escalator while it moves up, at 4 steps a second. How many steps do you take?
    • A second escalator is twice as fast. How many steps does the slow walker take on it, for the same length?

    Asked at Susquehanna International Group, Quantitative Trading, New York, 2026 (Wall Street Oasis): A stairs question, ask for some physics m/s type of questions

  4. 039A surveillance screen flags suspicious trades. One order in 100 is genuinely manipulative. Alert A fires with a likelihood ratio of 9, and an independent alert B with a likelihood ratio of 4. Both fire on the same order: what is the probability it is manipulative?Conditional probability and BayesCoreCitadelMiami · 2022

    Try it first

    Both alerts fire. Roughly how likely is the order manipulative?

    Show the worked solution

    About 26.7%. Work in odds. The prior odds are 1 to 99. Independent evidence multiplies the odds by each likelihood ratio: 1 x 9 x 4 = 36, so the posterior odds are 36 to 99. As a probability that is 36/135, about 26.7%. Even with both alerts, roughly three flagged orders in four are clean, because manipulation is rare to begin with.

    Why is odds form the fast way to combine alerts?

    Think of two smoke detectors in a kitchen where real fires are rare. Each beep makes a fire more likely, but toast sets both off far more often than fire does. In odds form, Bayes' rule is one multiplication per piece of independent evidence: posterior odds equal prior odds times each likelihood ratioHow much more often the evidence appears when the hypothesis is true than when it is false.. A ratio of 9 means alert A fires nine times as often on manipulative orders as on clean ones, for example on 90% of manipulative orders and 10% of clean ones.

    In odds form, each independent alert multiplies: 1:99, then 9:99, then 36:99Before any alertodds 1 : 991.0%Alert A firesodds 9 : 998.3%x 9Alert B also firesodds 36 : 9926.7%x 4red share: manipulative; light share: clean36 / (36 + 99) = 26.7%
    Starting from odds of 1 to 99, alert A multiplies the odds by 9 to reach 9 to 99, an 8.3% chance, and alert B multiplies by 4 to reach 36 to 99, which is only 26.7% because the prior was so low.

    How do you check 26.7% by counting?

    Take 10,000 orders: 100 manipulative and 9,900 clean. Suppose A fires on 90% of manipulative orders and 10% of clean ones, and B on 80% and 20%, which gives the stated ratios of 9 and 4. Both fire on 100 x 0.9 x 0.8 = 72 manipulative orders and on 9,900 x 0.1 x 0.2 = 198 clean ones. Of the 270 orders where both fire, 72 are manipulative: 26.7%, the same as the odds route.

    The relationship
    P(M∣A,B)P(Mˉ∣A,B)=199×9×4=3699  ⇒  P=36135≈26.7%\frac{P(M\mid A,B)}{P(\bar M\mid A,B)} = \frac{1}{99}\times 9\times 4 = \frac{36}{99} \;\Rightarrow\; P = \frac{36}{135} \approx 26.7\%
    Mthe order is manipulative
    1/99prior odds: 1 manipulative order per 99 clean
    9, 4likelihood ratios of alerts A and B
    What it says in wordsMultiply the prior odds by each alert's likelihood ratio, then turn odds back into a probability.

    State the assumption that made multiplication legal: the alerts are independent given the truth. If both alerts key off the same feature, say order size, the second adds little new information and multiplying by 4 overstates the case. With one alert alone the chance is 8.3% for A and 3.9% for B, which is why a desk reviews orders on combined evidence rather than a single flag.

    Where candidates lose it

    The common loss is treating a likelihood ratio of 36 as odds of 36 to 1 and answering about 97%. That throws away the base rate: the evidence multiplies the prior odds of 1 to 99, not even odds.

    The second is adding the ratios, 9 + 4 = 13, instead of multiplying. Independent evidence compounds, and odds form makes that one line of arithmetic.

    What the interviewer asks next

    • How many independent alerts with a ratio of 4 would you need to pass 50%?
    • Alert B is triggered by the same feature as alert A. How does that change your answer?
    • A third alert has a likelihood ratio of 0.5 and does not fire. What does that do?

    Asked at Citadel, Sales and Trading, Miami, 2022 (Wall Street Oasis): I got a question about Bayes' theorem applied to a practical scenario

  5. 041Five observations come from a uniform distribution on 0 to theta: 3.1, 7.4, 5.2, 9.0 and 1.8. What is the maximum likelihood estimate of theta, is it biased, and how would you correct it?Statistics and estimationCoreACAQR Capital ManagementTown of Greenwich · 2022

    Try it first

    What is the maximum likelihood estimate of theta?

    Show the worked solution

    The MLE is 9.0, the largest observation; it is biased low, and multiplying by (n + 1)/n = 6/5 gives an unbiased 10.8. Each observation has density 1/theta when theta covers it, so the likelihood is theta to the minus 5 for theta at least 9.0 and zero below. That peaks at 9.0. But the sample maximum averages 5/6 of theta, never above it, so scale it up by 6/5.

    Why does the likelihood peak at the largest observation?

    Suppose raffle tickets are numbered 1 to N and you see five, the highest being 90. N cannot be below 90, and the smaller N is, the more likely it was to produce those particular five tickets. For a uniform on 0 to theta, each observation has density 1/theta, so the likelihood is theta to the minus 5, which only falls as theta grows, but it is zero for any theta below an observation. The best allowed value is the smallest theta that covers all the data: the maximum, 9.0. Calculus does not help here, because the peak sits at the edge where the likelihood jumps from zero.

    The likelihood is zero below the largest observation and falls after itzero: some observation would exceed thetapeak at theta = 9.0: the MLE40% of peak at 10.824% of peak at 12thetaL(theta) = theta to the power -5, for theta at least 9.0The data and three estimates on one line0246810121416MLE 9.02 x mean = 10.66/5 x 9.0 = 10.8unbiased, and much tighter
    The likelihood is zero for theta below 9.0, peaks at 9.0 and then falls as theta to the minus 5, down to 40% of the peak at 10.8; on the data line, the MLE of 9.0 sits at the largest observation, the method of moments gives 10.6 and the bias-corrected estimate is 10.8.

    Why is 9.0 biased, and what is the right correction?

    The sample maximum can never exceed theta, so it can only err on the low side. Five points drop into 0 to theta and cut it into six gaps of the same average size, so the largest point sits on average one gap short of theta: at 5/6 of theta. Scaling the maximum by (n + 1)/n removes that bias: 9.0 x 6/5 = 10.8. The same logic underlies the classic serial-number estimation problem from wartime production counts.

    The relationship
    L(θ)=θ−5 1{θ≥9.0}E[max⁡]=nn+1θ  ⇒  θ^=65×9.0=10.8L(\theta) = \theta^{-5}\,\mathbf{1}\{\theta \ge 9.0\} \qquad E[\max] = \frac{n}{n+1}\theta \;\Rightarrow\; \hat\theta = \frac{6}{5}\times 9.0 = 10.8
    thetathe unknown upper end of the uniform
    n = 5number of observations
    maxthe largest observation, 9.0
    What it says in wordsThe likelihood peaks at the sample maximum, which on average falls short of theta by a factor n/(n + 1), so scale it up.

    An interviewer may ask why not use twice the mean, 2 x 5.3 = 10.6, which is also unbiased. The corrected maximum is far more precise: its variance is theta squared over n(n + 2), against theta squared over 3n for twice the mean, so twice the mean is 2.3 times as variable with five points. Twice the mean can even land below the largest observation, an estimate the data have already ruled out. The limitation of the correction is that unbiased is not the only goal: the multiple of the maximum with the smallest mean squared error is (n + 2)/(n + 1), which gives 10.5 here, and saying you would choose by the loss that matters shows you know the trade.

    Where candidates lose it

    The common loss is setting the derivative of the log-likelihood to zero, getting -5/theta = 0, and concluding there is no maximum. The maximum is at a boundary, where the indicator switches on, and that is the point of the question.

    The second is answering 9.0 and stopping. The follow-up is always the bias; say that the maximum sits below theta on average and give the (n + 1)/n correction with its one-line reason.

    What the interviewer asks next

    • What is the MLE if the distribution is uniform on theta to 2 theta?
    • Derive the variance of the corrected estimator.
    • The observations come from a uniform on theta minus 1 to theta plus 1. What is the MLE now?

    Asked at AQR Capital Management, Trading, Town of Greenwich, 2022 (Wall Street Oasis): Derive the mle for some given distribution. Explain linear regression intuitively and derive the ols estimate.

  6. 042A queue holds between 0 and 3 orders. Each tick at most one thing happens: with probability 0.3 a new order arrives (if there is room), with probability 0.5 one order is filled (if the queue is not empty), and otherwise nothing changes. In the long run, what fraction of ticks is the queue full?Markov chains and random walksCoreDRWNew York · 2026

    Try it first

    Roughly what share of ticks is the queue full?

    Show the worked solution

    27/272, about 9.9% of ticks. In a birth-death chain the long-run flow up across each cut equals the flow down, so share(k) x 0.3 = share(k + 1) x 0.5. Each state's share is 0.6 times the one below: weights 1, 0.6, 0.36 and 0.216, summing to 2.176. The full state gets 0.216/2.176, about 9.9%, and the queue is empty about 46% of the time.

    Why can you skip solving the full set of equations?

    Stand at a doorway between two rooms at a party that has settled down. Over an evening, the number of people walking through one way must match the number walking back, or one room would keep filling. In a chain that only steps up or down by one, the long-run flow across the boundary between neighbouring states must balance, which gives one simple equation per cut. Here flow up from state k is its share times 0.3, and flow down from state k + 1 is its share times 0.5.

    Birth-death chain: across each cut, flow up equals flow down0 ordersempty1 order2 orders3 ordersfull0.30.50.30.50.30.5Cut balance: share(k) x 0.3 = share(k + 1) x 0.5, so each share is 0.6 times the last46.0%weight 127.6%weight 0.616.5%weight 0.369.9%weight 0.216
    Arrivals push the queue up with probability 0.3 and fills pull it down with probability 0.5, so each state's long-run share is 0.6 times the one below: 46.0% empty, 27.6% with one order, 16.5% with two and 9.9% full.

    How do the cut equations give the answer?

    Write each share relative to the empty state. Each cut gives share(k + 1) = share(k) x 0.3/0.5 = 0.6 x share(k), so the weights are 1, 0.6, 0.36 and 0.216. They sum to 2.176, so the full queue holds 0.216/2.176 = 27/272 of the time, about 9.9%. The staying probabilities, 0.2 in the middle states and 0.5 when full, never enter; a chain that pauses on a state does not change the balance across cuts.

    The relationship
    πk+1=πk⋅0.30.5π3=0.631+0.6+0.62+0.63=27272≈9.9%\pi_{k+1} = \pi_k\cdot\frac{0.3}{0.5} \qquad \pi_3 = \frac{0.6^3}{1 + 0.6 + 0.6^2 + 0.6^3} = \frac{27}{272} \approx 9.9\%
    pi_klong-run share of ticks with k orders in the queue
    0.3chance of an arrival when there is room
    0.5chance of a fill when the queue is not empty
    What it says in wordsEach state is visited 0.6 times as often as the one below it; normalise the four weights to add to one.

    Check with conservation. Orders accepted per tick are 0.3 x (1 - 0.099) = 0.2702, and orders filled per tick are 0.5 x (1 - 0.460) = 0.2702: the same, as they must be. That gives a useful business number: arrivals turned away because the queue is full run at 0.3 x 0.099, about 0.030 per tick, or one arrival in ten. The model assumes one event per tick; if an arrival and a fill could happen in the same tick, the chain changes and so do the numbers.

    Where candidates lose it

    The common loss is assuming the four states are equally likely, or writing out all four balance equations with the self-loops and solving a 4 by 4 system under time pressure. The cut method needs three one-line ratios.

    The second is inverting the ratio, using 0.5/0.3, which makes the full state the most common. Fills are faster than arrivals, so the queue must lean toward empty; check the direction before you normalise.

    What the interviewer asks next

    • What is the average queue length?
    • What arrival probability would make the queue full 25% of the time?
    • How does the answer change if the queue can hold unlimited orders?

    Asked at DRW, Quantitative Research, New York, 2026 (Wall Street Oasis): There was a problem on Chi-squared distributions which was difficult and also one on birth death chains.

  7. 043A 3 x 3 x 3 cube is painted on the outside and cut into 27 small cubes. How many small cubes have 3, 2, 1 and 0 painted faces? You pick a small cube at random and roll it like a die: what is the probability the top face is painted?Counting and combinatoricsCoreJane StreetNew York · 2026

    Try it first

    What is the probability the top face is painted?

    Show the worked solution

    8 cubes have 3 painted faces, 12 have 2, 6 have 1 and 1 has none; the chance the top face is painted is exactly 1/3. Corners carry three, edge middles two, face centres one, and the core none. Picking a random cube and rolling it picks a random small face out of 27 x 6 = 162. The painted ones are the big cube's surface, 6 x 9 = 54, so the probability is 54/162 = 1/3.

    Where do the 8, 12, 6 and 1 come from?

    Think of a Rubik's cube: its pieces are corners, edges and centres, plus a hidden core. A small cube's painted faces equal the number of outer walls it touches: a corner touches three, an edge middle two, a face centre one, the core none. A cube has 8 corners, 12 edges with one middle piece each, and 6 faces with one centre each. That accounts for 8 + 12 + 6 = 26 cubes; the 27th is the core.

    Three layers of the cube, each small cube labelled by painted facesTop layer323212323Middle layer212101212Bottom layer323212323Numbers are painted faces on that small cube; the 0 is the hidden core.CubesPainted faces8 corners x 32412 edge middles x 2246 face centres x 161 core x 00Total painted faces54Every small cube is equally likely and every face of it is equally likely,so the top face is a uniform pick from all 27 x 6 = 162 small faces.Painted small faces = the big cube's surface: 6 faces x 9 = 5454/162 = 1/3
    Slicing the cube into three layers shows 8 corner cubes with three painted faces, 12 edge cubes with two, 6 face centres with one and a single unpainted core, which together carry 54 painted faces out of 162, exactly one third.

    Why is the roll probability exactly one third?

    Do it the long way first: weight each cube type by its share of cubes and its share of painted faces. 8/27 x 3/6 + 12/27 x 2/6 + 6/27 x 1/6 + 1/27 x 0 = (24 + 24 + 6)/162 = 54/162. Then notice the shortcut: a random cube with a random face up is a uniform pick from all 162 small faces, and the painted small faces are exactly the big cube's surface, 6 x 9 = 54. That gives 1/3 without any case split, and it works for any size: an n x n x n cube gives 6n squared over 6n cubed, which is 1/n.

    The relationship
    P(painted top)=8⋅3+12⋅2+6⋅1+1⋅027⋅6=54162=13P(\text{painted top}) = \frac{8\cdot 3 + 12\cdot 2 + 6\cdot 1 + 1\cdot 0}{27\cdot 6} = \frac{54}{162} = \frac13
    8, 12, 6, 1numbers of corner, edge, face-centre and core cubes
    3, 2, 1, 0painted faces on each type
    27 x 6all small faces, each equally likely to land on top
    What it says in wordsCount painted small faces over all small faces, because the roll makes every small face equally likely.

    Interviewers use the second part to see whether you look for the structure before the arithmetic. Counting faces instead of cubes turns a four-case weighted average into one division. Say both routes: the case split proves you can count, the face count proves you can see.

    Where candidates lose it

    The common loss is answering about cubes when the question is about faces: 26 of 27 cubes have paint, so candidates say 26/27, forgetting that a painted cube still shows an unpainted face most of the time.

    The second is miscounting edges, using 8 or 24 instead of 12. Say the cube's shape out loud, 8 corners, 12 edges, 6 faces, and check 8 + 12 + 6 + 1 = 27.

    What the interviewer asks next

    • For a 4 x 4 x 4 cube, how many small cubes have exactly two painted faces?
    • You roll a random small cube and see a painted top. What is the chance it is a corner cube?
    • For which n does an n x n x n cube have more unpainted small cubes than painted ones?

    Asked at Jane Street, Engineering, New York, 2026 (Wall Street Oasis): How you got to the answer matters even if you got the question right. Strawberry question + 3x3 cube question

  8. 048Calls on the same stock and expiry are quoted: the 95 strike at 9.80 bid, 10.20 offered, and the 100 strike at 4.30 bid, 4.50 offered. Is there an arbitrage, and exactly how would you trade it?Pricing, options and index mathsCoreWTWolverine Trading, Chicago, ILUSA · 2019

    Try it first

    What can you lock in, per spread, at these quotes?

    Show the worked solution

    Yes: sell the 95 call at 9.80 and buy the 100 call at 4.50, collecting 5.30 for a position that can never cost more than 5. A 95/100 call spread pays between 0 and the strike gap of 5 at expiry, so its price must sit between 0 and 5. The market lets you sell it for 5.30, which locks in at least 0.30 per spread, more if the stock ends below 100.

    What is a 95/100 call spread worth at most?

    Think of two coupons for the same shirt: one lets you buy it for Rs 950, the other for Rs 1,000. The first is worth more, but never by more than Rs 50, because the most it can save you over the second is the Rs 50 difference in price. Long the 95 call and short the 100 call pays the stock's rise above 95, capped once it reaches 100, so at expiry it is worth between 0 and the strike gap of 5. Anything that is certain to pay no more than 5 cannot be worth more than 5 today; with interest it is worth at most 5 discounted, slightly less.

    A 95/100 call spread can never pay more than 5, and the market bids 5.30 for itCall strikeBidOffer959.8010.201004.304.50Lime: the price you actually trade atThe tradeSell the 95 call at its bid+9.80Buy the 100 call at its offer-4.50Credit now; most owed at expiry is 5.00+5.30Worst case: 5.30 - 5.00 = +0.30012345859095100105110stock price at expiryvalue of the spreadyou collected 5.30payoff capped at 50 below 95Zoom: 4.75 to 5.50collected 5.30most you owe 5.00+0.30 locked
    Selling the 95 call at its 9.80 bid and buying the 100 call at its 4.50 offer collects 5.30 for a spread whose payoff is zero below 95 and capped at 5 above 100, so at least 0.30 is kept whatever the stock does at expiry.

    Which side of each quote do you trade at?

    This is where the question is really won or lost. You sell at the bid and buy at the offer, so the spread you can sell is worth 9.80 - 4.50 = 5.30 to you, not the mid of 5.60. 5.30 is still above 5, so the bound is broken at prices you can actually deal at. Buying the spread would cost 10.20 - 4.30 = 5.90 for something worth at most 5, a certain loss, so only one direction works. Check the stock price cases: below 95 both calls expire worthless and you keep 5.30; at 97 you owe 2 on the short call and keep 3.30; at 100 or above you owe exactly 5 net and keep 0.30.

    Stock at expiryShort 95 call paysLong 100 call receivesNet owedYou keep
    900.000.000.005.30
    950.000.000.005.30
    97-2.000.002.003.30
    100-5.000.005.000.30
    110-15.00+10.005.000.30
    The 5.30 collected less what the spread owes at expiry is never below 0.30, because the short 95 call and the long 100 call together never owe more than 5.
    The relationship
    0≤C(95)−C(100)≤(100−95) e−rT9.80−4.50=5.30>50 \le C(95) - C(100) \le (100 - 95)\,e^{-rT} \qquad 9.80 - 4.50 = 5.30 > 5
    C(K)price of the call with strike K, same stock and expiry
    e^{-rT}discount factor to expiry; it makes the upper bound slightly below 5
    What it says in wordsA call spread is worth between zero and the discounted strike gap; selling it for more than the gap is free money.

    What could stop the arbitrage from paying?

    If the calls are American and the short 95 is exercised early, exercise the 100 call too: you pay the stock price minus 95 and receive the stock price minus 100, a net 5, and you already hold 5.30. The real frictions are fees, the margin the short call ties up, and the risk that the quote vanishes after you trade one leg. Trade both legs together as a spread order. On a real screen a 0.30 bound violation lasts seconds, which is why the interviewer is testing whether you can see it fast and name the side, not whether such quotes are common.

    Where candidates lose it

    The common loss is reasoning with mid prices: 10.00 - 4.40 = 5.60 and a claimed profit of 0.60. Nobody deals at mids; you sell at the bid and buy at the offer, and the honest edge is 0.30.

    The second is getting the direction backwards and buying the spread because the 95 call looks cheap next to its payoff. Say the bound first, the spread is worth at most 5, and the direction follows: sell it.

    What the interviewer asks next

    • The 105 call is quoted 1.10 bid, 1.30 offered. Is there a butterfly arbitrage across 95, 100 and 105?
    • What is the lower bound on the 95/100 call spread, and what quotes would break it?
    • How does a dividend before expiry change the early-exercise argument?

    Asked at Wolverine Trading, Prop Trading, Chicago, IL, USA, 2019 (Wall Street Oasis): pricing options given an ask and a bid price for options with different strikes if you were to short one and long another

  9. 051You and a friend agree to meet at a spot some time between 5 and 6 pm. Each of you arrives at an independent, uniformly random time in that hour and waits 20 minutes for the other before leaving (or until 6 pm, whichever comes first). What is the probability you meet?Continuous and geometric probabilityCoreJane StreetNew York · 2026

    Try it first

    Before you draw anything: what is the chance you meet?

    Show the worked solution

    5/9, about 55.6%. Put your arrival time on one axis and your friend's on the other, so every outcome is a point in a unit square. You meet when the two times differ by at most a third of an hour, a band along the diagonal. The two corner triangles outside the band each have legs of 2/3, area 2/9, so the band is 1 - 4/9 = 5/9.

    Why turn two arrival times into a square?

    Think of two people trying to catch each other at a tea stall with no phones. Nothing about the answer depends on who is you and who is the friend; it depends only on the pair of times. With two independent uniform times, every pair is equally likely, so the pair is a point spread evenly over a square and any probability is simply an area. The event you meet becomes a region: the set of points where the two times are within 20 minutes of each other. That region is a diagonal band, because the line x = y is where you arrive together.

    Every pair of arrival times is a point; you meet inside the bandMiss2/9Miss2/9Meet: 5/900202040406060Your arrival, minutes past the hourFriend's arrivalCount the white, not the greenWhole square1One miss triangle: legs of 40 minutes(2/3) x (2/3) / 2 = 2/9Both miss triangles4/9Meeting band1 - 4/9 = 5/9P(meet) = 5/9 = 55.6%
    Plotting your arrival against your friend's, the meeting region is the diagonal band where the times differ by 20 minutes or less; the two white corner triangles, each 2/9 of the square, are the misses, so you meet with probability 5/9.

    How do you get the area without any integration?

    Count the region you do not want. The two corners where one person arrives more than 20 minutes after the other are right triangles with both legs 40 minutes long, which is 2/3 of the side. Each has area (2/3) x (2/3) / 2 = 2/9, together 4/9, so the band is 5/9. Complements are the fastest route here, as they are for most geometric probability questions, because the leftover pieces are usually triangles.

    The relationship
    P(∣X−Y∣≤w)=1−(1−w)2=2w−w2w=13: 1−(23)2=59P(|X-Y|\le w) = 1-(1-w)^2 = 2w - w^2 \qquad w=\tfrac13:\ 1-\left(\tfrac23\right)^2 = \tfrac59
    X, Ythe two arrival times as fractions of the hour, independent and uniform on 0 to 1
    wthe waiting time as a fraction of the hour, here 20 of 60 minutes
    What it says in wordsThe chance of meeting is one minus the two corner triangles, whose legs are each one minus the waiting time.

    What does the general formula tell you that the number does not?

    Read 2w - w squared term by term. The 2w is the naive answer of either person waiting, and the minus w squared removes the double count and the clipping at the edges of the hour. It also tells you how waiting time buys certainty: to meet half the time each person must wait about 17.6 minutes, and to be sure each must wait the whole hour. The same picture prices any tolerance between two random arrivals, such as two orders landing in the same matching window of an auction.

    Where candidates lose it

    The common answer is 1/3, from reading 20 minutes as a third of the hour. It ignores that either person can be the one who waits, and it has no way to handle the edges of the hour, where a person arriving at 5:55 can only wait five minutes.

    The second trap is trying to integrate over one person's arrival time case by case near the edges. It works but wastes three minutes. Draw the square first and subtract the two triangles out loud.

    What the interviewer asks next

    • How long would each person need to wait for a 50% chance of meeting?
    • You wait 10 minutes and your friend waits 30. What is the chance now?
    • Three people arrive at random in the hour and each waits 20 minutes. What is the chance all three are together at some moment?

    Asked at Jane Street, Technology, New York, 2026 (Wall Street Oasis): 1v1 math problems. bus stop. two people meeting probelm

  10. 060In how many ways can you place four queens on a 4 by 4 board so that no queen attacks another? How would you organise the search, and what does the same method give for five queens on a 5 by 5 board?Logic and algorithmic reasoningCoreGoldman SachsNew York · 2026

    Try it first

    How many non-attacking placements of four queens exist on a 4 by 4 board?

    Show the worked solution

    Two on a 4 by 4 board and 10 on a 5 by 5 board. Place one queen per row, trying columns left to right, and abandon a branch the moment the next row has no safe column. On 4 by 4 this visits 16 placements and finds columns 2, 4, 1, 3 and 3, 1, 4, 2. On 5 by 5 the same search visits 53 placements, against 3,125 boards for brute force.

    How do you organise the search so it stays small?

    Think of filling a seating plan for a wedding where some guests cannot sit near each other. You seat table by table, and the moment a table has no acceptable guest left you undo the previous choice instead of finishing a doomed plan. Backtracking builds the answer one decision at a time and abandons a partial answer as soon as it breaks a rule, so it never enumerates the boards that fail early. For queens, two rules come free from the structure: one queen per row, and one per column, which leaves only the diagonals to check at each step.

    Place row by row, abandon a branch the moment it has no safe squarerow 1row 2row 3row 4start13dead end42dead end2413solution3142solution413dead end2dead endNumber in each circle = the column chosen for that row2 4 1 33 1 4 24 x 4: 16 placements tried,2 solutions5 x 5: 53 placements tried,10 solutionsBrute force: 4^4 = 256 and5^5 = 3,125 full boards
    Row by row, the search tries 16 placements on the 4 by 4 board; four branches die when a row has no safe column, and two reach row 4, giving the solutions 2, 4, 1, 3 and 3, 1, 4, 2.

    How does the 4 by 4 search actually run?

    Start with the corner. A queen in column 1 of row 1 leaves row 2 only columns 3 and 4, and both paths run out of safe squares by row 3 or row 4, so no solution uses a corner queen. A queen in column 2 forces column 4 in row 2, then column 1 in row 3 and column 3 in row 4, which works. Columns 3 and 4 are mirror images of 2 and 1. So there are exactly two solutions, and they are reflections of each other. Saying the symmetry out loud halves the work and is exactly what an interviewer building up from a base case wants to hear.

    What changes on 5 by 5, and how does the method scale?

    The larger board has more room, and every row 1 column leads somewhere. The same search finds 10 solutions after 53 placements, while brute force over one queen per row would test 3,125 boards. On 8 by 8 it finds all 92 solutions in 2,056 placements out of 16,777,216 one-per-row boards. In code, keep three sets, used columns, used down-diagonals (row minus column) and used up-diagonals (row plus column), so each safety check is constant time.

    Where candidates lose it

    Candidates start listing boards by eye and lose track, or they try all C(16, 4) = 1,820 ways to place four queens anywhere. The interviewer wants the structure: one per row, a column choice per row, and pruning.

    The second loss is counting the two 4 by 4 solutions as four or eight by treating rotations as new. Say whether you count symmetric boards as distinct, and note that here the two solutions are each other's mirror image.

    What the interviewer asks next

    • Write the backtracking function and state its time complexity in the worst case.
    • How would you count solutions up to rotation and reflection?
    • Why do the 2 by 2 and 3 by 3 boards have no solution at all?

    Asked at Goldman Sachs, Quantitative Research, New York, 2026 (Wall Street Oasis): I was asked a backtracking question in 1 of the rounds in the superday.

← PreviousPage 2 of 4
  1. 1
  2. 2
  3. 3
  4. 4
Next →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.