Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
Explore NISM prep
Series-VIII · Equity DerivativesSeries-XII · Securities Markets FoundationSeries-V-A · Mutual Fund DistributorsSeries-XV · Research AnalystSeries-XIX-E · Category III AIF ManagersSeries-XIX-D · Category I & II AIF ManagersSeries-XIX-C · Alternative Investment Fund ManagersSeries-XVI · Commodity DerivativesSeries-VI · Depository OperationsSeries-II-A · Registrars & Transfer AgentsSeries-I · Currency DerivativesSeries-VII · Securities Operations & Risk Management
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant puzzles, solved step by step

Puzzles
100
Traced to a firm
71
Topics
12
Hard
30
Topic
All topicsLogic and algorithmic reasoning10Conditional probability and Bayes7Counting and combinatorics8Continuous and geometric probability9Correlation, regression and linear algebra9Market making, betting and sizing9Expected value and optimal stopping9Statistics and estimation9Pricing, options and index maths7Games and strategic reasoning8Markov chains and random walks7Mental maths and number sense8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 21–30 of 71 · filtered from 100Clear filters
  1. 031Take a random ordering of n distinct numbers and run exactly one left-to-right pass of bubble sort, swapping each adjacent pair that is out of order. What is the probability the list is fully sorted afterwards? Work it for n = 5.Counting and combinatoricsHardJump TradingChicago · 2018

    Try it first

    For n = 5, how likely is the list sorted after one pass?

    Show the worked solution

    2 to the power (n - 1) divided by n factorial, which is 16/120 = 2/15 for n = 5. One pass moves every number that is not carried rightwards exactly one place left. So the list ends sorted only if no number starts more than one place right of its final spot. Placing 1, then 2, then 3 and so on, each has two allowed spots and the largest takes the last one, giving 2 to the power (n - 1) orderings.

    What does one pass actually do to each number?

    Picture a queue at a ticket window where the tallest person seen so far keeps stepping back past anyone shorter. That person travels a long way to the right; everyone they pass shifts one step forward. In one pass, the running maximum is carried right until it meets something larger, and every number it passes moves exactly one place left. Nothing moves left by two in a single pass. That limit is the whole problem.

    One pass moves a number left by one place at most3 1 2 5 4: sorts in one pass3125412345beforeaftersorted: every number is home2 3 1 4 5: does not2314521345beforeafternot sorted: the 1 needed two stepsRule: each number may start at most one place right of its home.Place 1, 2, 3, 4 in turn: 2 allowed spots each. The 5 takes the last spot.2 x 2 x 2 x 2 x 1 = 16 orderings out of 5! = 12016/120 = 2/15
    In 3 1 2 5 4 every number starts at most one place right of its home, so one pass sorts it; in 2 3 1 4 5 the 1 starts two places right of home and ends one short, so only 16 of the 120 orderings of five numbers, 2 in 15, sort in one pass.

    Which orderings survive, and how do you count them?

    Because a number can shift left by one at most, the list sorts only if each number starts no more than one place right of its home. The converse also holds: when every number meets that condition, the pass carries each big number to exactly where it belongs. For five numbers, a brute-force check of all 120 orderings finds exactly the 16 that meet the condition, and all 16 sort.

    Now count them without listing. Place the numbers in increasing order. The 1 may sit in position 1 or 2. The 2 may sit anywhere in positions 1 to 3, one of which the 1 already took: two choices. The same holds for 3 and 4: each has k + 1 allowed spots, k - 1 of them already used by smaller numbers, so two choices each. The 5 fills the one position left. That is 2 x 2 x 2 x 2 x 1 = 16.

    The relationship
    P(sorted after one pass)=2 n−1n!n=5: 16120=215P(\text{sorted after one pass}) = \frac{2^{\,n-1}}{n!} \qquad n = 5:\ \frac{16}{120} = \frac{2}{15}
    2^(n-1)orderings where no number starts more than one place right of its home
    n!all orderings of n distinct numbers, equally likely
    What it says in wordsTwo choices for each number except the largest, over all possible orderings.

    Check small cases out loud: for n = 2 both orderings sort, 2 of 2; for n = 3 it is 4 of 6. The probability collapses fast, because n factorial outruns 2 to the power n: about 4.4% for n = 6 and 1.3% for n = 7.

    Where candidates lose it

    The common wrong start is to think one pass only fixes the largest number, and answer that the other n - 1 must already be sorted, which gives 1/(n - 1)! and 1/24 for n = 5. It misses that every passed number also moves left one place, which rescues many orderings.

    The other loss is guessing a rule from one example. State the one-step-left limit, derive the condition from it, then count by placing numbers in increasing order.

    What the interviewer asks next

    • What is the probability the list is sorted after two passes?
    • How many passes does bubble sort need on average for a random list of n numbers, roughly?
    • What if the pass runs right to left instead?

    Asked at Jump Trading, Research, Chicago, 2018 (Wall Street Oasis): one iteration of bubble sort, what's the probability that the array will be sorted

  2. 032n points are placed independently and uniformly on a circle of circumference 1, with n at least 3. Each point colours the arc between itself and its nearest neighbour. What is the expected total length that gets coloured?Continuous and geometric probabilityHardSusquehanna International GroupLondon · 2026

    Try it first

    Which is closest to the expected coloured length?

    Show the worked solution

    7/18, about 0.389, for every n from 3 upwards. A gap is left uncoloured only when it is longer than both gaps beside it, because then neither endpoint has it as its nearest. For three points that gap is simply the longest of three pieces, which averages 11/18, so 7/18 is coloured. For larger n the same 11/18 comes out, so the answer does not depend on n.

    When is a gap left uncoloured?

    Picture people standing round a circular table, each turning to talk to whoever is closer, left or right. A stretch of table between two people stays silent only if both of them turned away, which means each had a closer person on their other side. A gap is uncoloured exactly when it is longer than both of its neighbouring gaps. A gap coloured from both ends is still coloured once, so the question becomes: what is the expected total length of gaps that are local maxima?

    A gap stays uncoloured only if it is longer than both of its neighbours0.060.040.150.080.100.120.070.080.160.14coloured: 0.57uncoloured: 0.43shorter gap for at least one endpointlonger than both neighboursThree points: the only uncoloured gap is thelongest of three pieces, which averages(1/3)(1 + 1/2 + 1/3) = 11/18Any n: each gap is uncoloured with the sameexpected length, and n of them add to 11/18Expected coloured length7/18 = 0.389Simulated, 40,000 circles: n = 3 gives 0.389,n = 5 gives 0.388, n = 10 gives 0.389
    Each gap is coloured if it is the shorter gap for at least one endpoint and left uncoloured if it is longer than both neighbours; this sample of ten points colours 0.57 of the circle, and the average over all placements is 7/18, about 0.389, for any n of 3 or more.

    How do you get 11/18 for the uncoloured part?

    Start with n = 3, the case you can finish in the room. With three gaps, every gap's two neighbours are the other two gaps, so the only uncoloured gap is the longest one. Three random points cut the circle like a stick broken into three, and the longest of three pieces averages (1/3)(1 + 1/2 + 1/3) = 11/18. So the coloured length is 7/18.

    For larger n, use the fact that the n gaps behave like n independent exponentialA random length whose chance of ending is the same at every instant; waiting times between random arrivals follow it. lengths rescaled to add up to 1, and that the rescaling is independent of the shape. For three unit exponentials X, Y and Z, the expected value of X counted only when X is the largest is 1 - 2/4 + 1/9 = 11/18. Each of the n gaps contributes that, divided by the expected total of n, and the n gaps sum to 11/18 again. The uncoloured share is 11/18 whatever n is, so the coloured share is always 7/18.

    The relationship
    E[coloured]=1−n⋅1n∫0∞xe−x(1−e−x)2 dx=1−1118=718E[\text{coloured}] = 1 - n\cdot\frac{1}{n}\int_0^\infty x e^{-x}(1-e^{-x})^2\,dx = 1 - \frac{11}{18} = \frac{7}{18}
    x e^(-x)a gap's length times its density, in the exponential picture
    (1 - e^(-x))^2the chance both neighbouring gaps are shorter
    1/nrescaling so the n gaps add to a circle of length 1
    What it says in wordsThe expected length of gaps longer than both neighbours is 11/18, and the rest of the circle is coloured.

    Say the check: a seeded simulation of 40,000 random circles gives 0.389 for n = 3, 0.388 for n = 5 and 0.389 for n = 10. The limitation is that the exponential step is a known result you should name, not derive, in an interview; the n = 3 case is the part you prove on the spot.

    Where candidates lose it

    The usual loss is counting gaps instead of measuring them. One gap in three is a local maximum, so candidates answer 2/3 coloured. The uncoloured gaps are selected for being long, which is why their share of length, 11/18, is far above one third.

    The second is double counting a gap that both endpoints colour. It is coloured once. Frame the problem around uncoloured gaps and both mistakes disappear.

    What the interviewer asks next

    • What is the expected number of uncoloured gaps?
    • What if each point colours the arc to its farther neighbour instead?
    • Does the answer change for points on a line segment rather than a circle?

    Asked at Susquehanna International Group, Quantitative Research, London, 2026 (Wall Street Oasis): if n points are placed on a circle and each point colours in the arc to its nearest neighbour

  3. 034Make me a two-way market on the number of heads in 100 flips of a fair coin, and justify the width.Market making, betting and sizingWarm upDRWNew York · 2026

    Try it first

    What is the standard deviation of the number of heads?

    Show the worked solution

    Centre it at 50 and quote around 46 at 54. The fair value is exactly 50. The standard deviation is √(100 x 0.5 x 0.5) = 5, so settlement lands between 45 and 55 about 73% of the time. A market 4 either side of fair earns 4 per lot on any trade, loses on a single sale at 54 only 18% of the time, and leaves room to move the quote if the other side seems to know something.

    Where does the centre come from, and what sets the width?

    A shopkeeper selling mangoes by the dozen knows the fair price; the margin he adds depends on how much the price of the next crate can swing and on whether the buyer knows something he does not. The centre of your market is the expected value, and the width is a choice about risk and information, scaled by how much the outcome can move. Here the expected value is 100 x 0.5 = 50, and nobody can know more than you about fresh flips of a fair coin, so the width is about risk alone.

    Heads in 100 flips: centred at 50, standard deviation 5303540455055606570bid 46offer 5445 to 5540 to 60Within 45 to 55 (one standard deviation): 72.9% of outcomesWithin 40 to 60 (two standard deviations): 96.5%Edge 4 per lot either sidea sale at 54 loses 18% of the time
    The number of heads in 100 fair flips is centred at 50 with a standard deviation of 5, landing in 45 to 55 72.9% of the time and in 40 to 60 96.5% of the time, so a market of 46 at 54 sits inside one standard deviation and earns 4 per lot on each side.

    How do you justify 46 at 54 rather than 49 at 51?

    Use the standard deviation as the ruler. The count has variance 100 x 0.5 x 0.5 = 25, so a standard deviation of 5. A quote 4 either side of fair earns 4 on each lot traded, against a settlement that typically moves 5, so every trade has an edge worth a large fraction of its risk. If someone buys at 54, you lose only if the count finishes at 55 or more, about 18% of the time. A tight 49 at 51 earns 1 per lot and a sale at 51 loses whenever the count reaches 52, about 38% of the time. Tighter wins more trades and earns less on each; in an interview game, start around one standard deviation wide and tighten as you learn.

    The relationship
    μ=np=50σ=np(1−p)=25=5\mu = np = 50 \qquad \sigma = \sqrt{np(1-p)} = \sqrt{25} = 5
    n = 100number of flips
    p = 0.5chance of heads on each flip
    sigmastandard deviation of the number of heads
    What it says in wordsThe count of heads averages 50 and typically lands within 5 of it.

    Then say how you would react to trades, because that is the follow-up. If the interviewer lifts your 54 again and again, either they are testing your nerve or they know something, perhaps that the coin is not fair or that some flips are already done. Repeated one-way trading is information: move your market toward it and cut your size, rather than defending 50. The limitation of the simple answer is exactly that it assumes nobody knows more than you.

    Where candidates lose it

    The common loss is quoting 50 at 50, or 49.5 at 50.5, and calling it fair. A market maker earns the spread; a zero-width quote gives away every trade at no edge and leaves no room to adjust when the other side knows more.

    The second is quoting a width with no reason. Name the standard deviation of 5, then choose a width against it. The number you say matters less than showing that width and risk are linked.

    What the interviewer asks next

    • I buy 10 lots at 54. Where is your new market?
    • Now 60 flips have already happened and I have seen them. How does your market change?
    • Make a market on the number of heads squared.

    Asked at DRW, Quantitative Trading, New York, 2026 (Wall Street Oasis): Make a market on the number of heads out of 100 coin flips.

  4. 035Regressing y on x gives a slope of 0.8; regressing x on y gives a slope of 0.45. What is the R-squared of either regression, and what is the correlation?Correlation, regression and linear algebraCoreTower Research CapitalNew York · 2014

    Try it first

    What is the correlation between x and y?

    Show the worked solution

    R-squared is 0.36 for both regressions and the correlation is 0.6. The slope of y on x is r times sd(y)/sd(x); the slope of x on y is r times sd(x)/sd(y). Multiplying them cancels the standard deviations and leaves r squared: 0.8 x 0.45 = 0.36. The correlation is +0.6, positive because both slopes are positive, and the ratio sd(y)/sd(x) is √(0.8/0.45) = 4/3.

    Why are the two slopes not reciprocals of each other?

    Tall parents tend to have tall children, but a little less tall; and tall children tend to have tall parents, but a little less tall. Both statements are true at once. Each regression predicts toward the mean, so neither slope is the inverse of the other unless the fit is perfect. If the points lay exactly on a line, the slope of x on y would be 1/0.8 = 1.25. It is 0.45 instead, and the size of that shortfall is what measures how loose the relationship is.

    Two regressions, two lines: the slopes multiply to R-squaredxyy on x: slope 0.8x on y: slope 0.45, drawn as 1/0.45 = 2.22Slope of y on x = r x (sd of y / sd of x)Slope of x on y = r x (sd of x / sd of y)Multiply the two: the sd ratios cancel0.8 x 0.45 = 0.36 = R-squaredr = +0.6Divide instead: sd of y / sd of x= √(0.8 / 0.45) = 1.333Neither slope is the reciprocal of the otherbecause |r| < 1: both regress toward the mean
    Fitting y on x gives the shallower line with slope 0.8 and fitting x on y gives the steeper line, slope 0.45 in its own terms; their product, 0.36, is R-squared, so the correlation is 0.6 and the standard deviation of y is 4/3 that of x.

    How do the two slopes give R-squared?

    Write each slope in terms of the correlation. The least squares slope of y on x is the covariance over the variance of x, which is r times sd(y)/sd(x). Swap the roles and the slope of x on y is r times sd(x)/sd(y). The standard deviation ratios are reciprocals, so the product of the two slopes is r squared, and in a one-variable regression r squared is exactly the R-squared. Here 0.8 x 0.45 = 0.36, so r = 0.6; the sign is positive because both slopes are positive, and the two slopes always share a sign.

    The relationship
    by∣x bx∣y=rsysx⋅rsxsy=r2=0.8×0.45=0.36b_{y|x}\, b_{x|y} = r\frac{s_y}{s_x}\cdot r\frac{s_x}{s_y} = r^2 = 0.8 \times 0.45 = 0.36
    b_y|xslope from regressing y on x, 0.8
    b_x|yslope from regressing x on y, 0.45
    s_x, s_ystandard deviations of x and y
    rthe correlation of x and y
    What it says in wordsThe two slopes multiply to the squared correlation because the scale factors cancel.

    The figure uses 40 points built with standard deviations 3 and 4 and a correlation of exactly 0.6, and fitting both regressions returns slopes of 0.80 and 0.45. A quick sanity test comes free: the product of the two slopes can never exceed 1. If an interviewer quotes slopes of 0.8 and 1.5, the product 1.2 is impossible, and saying so is worth more than any calculation.

    Where candidates lose it

    The fast wrong answer is to say the slopes should be reciprocals and call the data inconsistent, or to answer 0.36 when asked for the correlation. 0.36 is R-squared; the correlation is its square root.

    The second loss is dropping the sign. The square root of 0.36 could be plus or minus 0.6; both slopes are positive, so the correlation is positive, and saying why takes one sentence.

    What the interviewer asks next

    • What is the ratio of the standard deviation of y to that of x?
    • If the slope of x on y were 1.5, what would you conclude?
    • How does adding measurement noise to x change each slope?

    Asked at Tower Research Capital, Quantitative Research, New York, 2014 (Wall Street Oasis): Another detailed linear regression questions were asked, including problems about residual, variance and R^2

  5. 036A stock pays a growing dividend and is valued with the Gordon model at a discount rate of 10% and growth of 6%. What is its duration, and roughly how much does its price change if the discount rate rises by one point?Pricing, options and index mathsCoreBLBlackRockNew York · 2026

    Try it first

    What is the stock's duration, its percentage price sensitivity to the discount rate?

    Show the worked solution

    Duration is 1/(r - g) = 25 years, so a one-point rise cuts the value by about a fifth. The Gordon price is D1/(r - g), and its percentage sensitivity to r is 1/(r - g) = 1/0.04 = 25. With a Rs 4 dividend the price moves from Rs 100 at 10% to Rs 80 at 11%, a 20% fall. The 25% duration estimate overshoots because the price curve is convex.

    Why does a stock have a duration at all?

    A promise of money in one year hardly changes in value when rates move; a promise of money in twenty five years changes a lot, because the rate is compounded over every one of those years. A stock is a stream of dividends stretching forever, and when the dividends grow, most of its value sits in cash flows far in the future, so it behaves like a very long bond. Duration measures exactly that: the percentage price change for a change in the discount rate.

    The relationship
    P=D1r−g−1PdPdr=1r−g=10.10−0.06=25P = \frac{D_1}{r-g} \qquad -\frac{1}{P}\frac{dP}{dr} = \frac{1}{r-g} = \frac{1}{0.10-0.06} = 25
    D1next year's dividend, Rs 4 in the illustration
    rdiscount rate, 10%
    gdividend growth rate, 6%
    1/(r - g)percentage price change per unit change in r
    What it says in wordsDifferentiate the Gordon price and divide by price: the sensitivity is one over the gap between the discount rate and growth.
    Gordon price against the discount rate, with the tangent at 10%9%10%11%12%13%5075100125150175Discount rate rPrice, RsRs 100 at 10%Rs 80: actual, -20%Rs 75: tangent, -25%Rs 66.7 at 12%tangent Rs 50P = D1 / (r - g)Rs 4 / (0.10 - 0.06) = 100Duration = 1 / (r - g)25yearsOne point on r:duration says -25%the curve says -20%A 10-year 10% bond at 10%:duration about 6.1 years
    At a 10% discount rate the Rs 4 dividend stock is worth Rs 100 and its duration is 25; at 11% the price is Rs 80, a 20% fall, while the tangent line predicts Rs 75, and at 12% the gap widens to Rs 66.7 against Rs 50 because the price curve is convex.

    Why does the estimate say 25% when the price falls 20%?

    Duration is the slope at one point, and the price curve bends. For a one-point rise the tangent predicts a 25% fall, but the exact move from Rs 100 to Rs 80 is 20%, because the curve is convexCurving upward, so it always sits above any of its tangent lines. and flattens as r rises. The same bend makes a one-point fall worth more than 25%: at 9% the price is Rs 133.3, up 33.3%. For a big rate move, reprice exactly instead of trusting the slope. Duration here also equals price over dividend, 100/4, which is a quick way to say it: one over the dividend yield.

    For precision, the Macaulay durationThe present-value weighted average time at which cash flows arrive. is (1 + r)/(r - g) = 27.5 years, and dividing by 1 + r gives the modified duration of 25. A 10-year bond paying 10% at a 10% yield has a modified duration of about 6.1. The gap r - g is what matters, which is why high-growth stocks carry the most duration: the same stock with 2% growth would have a duration of 12.5 years. The limitation is that the Gordon model holds growth fixed while rates move; in practice both shift together.

    Where candidates lose it

    The usual loss is saying a stock has no duration because it has no maturity, or that it is infinite because it pays forever. Both skip the one line of calculus that gives 1/(r - g).

    The second is quoting 25% as the exact price change. It is the slope at 10%; the exact fall to 11% is 20%, and saying why, convexity, is what separates a strong answer.

    What the interviewer asks next

    • What happens to duration as growth approaches the discount rate?
    • Why might a stock's measured sensitivity to bond yields be much lower than 25?
    • What is the price change for a one-point fall in the discount rate?

    Asked at BlackRock, Restructuring, New York, 2026 (Wall Street Oasis): Which equities have duration ? multiple stocks vs value stocks MSE Forecasting equation

  6. 037Without paper: work out 56 x 56 and 73 x 74, and say the shortcut you used for each.Mental maths and number senseWarm upAkuna CapitalChicago · 2025

    Try it first

    What is 56 x 56?

    Show the worked solution

    56 x 56 = 3,136 and 73 x 74 = 5,402. For 56 squared, split it as 50 + 6: 2,500, plus two strips of 300, plus 36. For 73 x 74, anchor both on 70: 4,900, plus 70 x 7 = 490, plus 3 x 4 = 12. A second route checks each: (60 - 4) squared = 3,136, and 73.5 squared minus a quarter = 5,402.

    What is the shortcut for squaring a two-digit number?

    Tiling a floor that is 56 tiles on each side, you would lay the big 50 by 50 block first, then two thin strips along the edges, then a small corner. Splitting a number into a round base plus a small part turns one hard product into one easy square and a few small ones: (a + b) squared = a squared + 2ab + b squared. For 56: 2,500 + 2 x 300 + 36 = 3,136. You can also go down from the next round number: (60 - 4) squared = 3,600 - 480 + 16, again 3,136.

    Anchor on a round base: one big block plus small strips2,5003003003650656 x 56 = 2,500 + 300 + 300 + 36= 3,136check: (60 - 4) squared = 3,600 - 480 + 164,90070 x 4 = 2803 x 70 = 2103 x 4 = 1270473 x 74 = 4,900 + 280 + 210 + 12= 5,402check: 73.5 squared - 0.25 = 5,402.25 - 0.25
    56 squared splits into a 2,500 block, two 300 strips and a 36 corner, total 3,136; 73 x 74 splits into 4,900, 280, 210 and 12, total 5,402, the same answer as 73.5 squared minus a quarter.

    What changes when the two numbers differ, as in 73 x 74?

    When two numbers share a tens digit, anchor both on it. For (70 + 3)(70 + 4), the product is 70 squared, plus 70 times the sum of the units, plus the product of the units: 4,900 + 490 + 12 = 5,402. The midpoint route gives the same: numbers equally spaced around 73.5 multiply to 73.5 squared minus the square of the half gap, 0.25, and 73.5 squared is 4,900 + 490 + 12.25. Another quick path: 73 x 74 = 73 squared + 73 = 5,329 + 73.

    The relationship
    (a+b)(a+c)=a2+a(b+c)+bc73×74=4900+490+12=5402(a+b)(a+c) = a^2 + a(b+c) + bc \qquad 73 \times 74 = 4900 + 490 + 12 = 5402
    athe round base, 70
    b, cthe small parts, 3 and 4
    What it says in wordsMultiply the round parts, add the round part times the sum of the small parts, then add the small product.

    Timed tests reward a fixed routine more than cleverness. Pick one decomposition, say the partial products in order, and check with a second route only if time allows. The last digit is a free check: 6 x 6 ends in 6 and 3 x 4 ends in 2, so 3,136 and 5,402 pass. A good habit is to sanity check the size too: 56 squared must sit between 50 squared, 2,500, and 60 squared, 3,600.

    Where candidates lose it

    The usual slip in 56 squared is adding one strip of 300 instead of two, giving 2,836, or dropping the 36. The area picture makes both errors visible: a square has two strips and a corner.

    On a timed screen the other loss is switching methods halfway. Commit to the split, say each partial product, then add. Checking the last digit costs a second and catches most slips.

    What the interviewer asks next

    • Work out 97 x 103 in your head.
    • What is 35 squared, and what is the trick for squares ending in 5?
    • Estimate 48 x 52 without multiplying directly.

    Asked at Akuna Capital, Prop Trading, Chicago, 2025 (Wall Street Oasis): The mental math problems which were timed, one example was the 56*56

  7. 038Walking up a moving escalator at one step per second you take 20 steps; walking at two steps per second you take 32 steps. How many steps are visible on the escalator?Logic and algorithmic reasoningCoreSusquehanna International GroupNew York · 2026

    Try it first

    How many steps are visible?

    Show the worked solution

    80 steps. At one step a second the climb takes 20 seconds; at two steps a second it takes 16. If the escalator moves v steps a second, the visible steps are 20 + 20v and also 32 + 16v. Setting them equal gives v = 3, so the escalator is 20 + 60 = 80 steps long, and the check 32 + 48 = 80 agrees.

    What stays the same between the two walks?

    On an airport moving walkway, walk slowly and the belt does most of the work; stride out and you do more of it yourself, but you reach the end sooner. The length of the walkway does not change. Every visible step is covered either by your legs or by the escalator, so your steps plus the escalator's movement during your climb always equal the same total. That fixed total is the unknown; the escalator's speed is the second unknown, and two walks give two equations.

    Two walks, same escalator: your steps plus the escalator's always make 801 step/s, 20 syou: 20escalator: 20 s x 3 = 602 steps/s, 16 syou: 32escalator: 16 s x 3 = 4880 visible steps20 + 20v = 32 + 16v, so 4v = 12The faster walk saves 4 seconds of escalator help and pays for them with 12 extra stepsv = 3, N = 80
    Walking at one step a second you climb 20 steps in 20 seconds while the escalator carries 60; at two steps a second you climb 32 in 16 seconds while it carries 48; both add to the same 80 visible steps because the escalator moves 3 steps a second.

    How do you set up and solve the two equations?

    Turn step counts into time first, because the escalator's contribution depends on time. The slow walk: 20 steps at one a second is 20 seconds. The fast walk: 32 steps at two a second is 16 seconds. The faster walk loses 4 seconds of escalator help and makes it up with 12 extra steps of its own, so the escalator moves 3 steps a second. Then the total is 20 + 20 x 3 = 80, and 32 + 16 x 3 = 80 confirms it.

    The relationship
    N=20+20v=32+16v  ⇒  v=3, N=80N = 20 + 20v = 32 + 16v \;\Rightarrow\; v = 3,\ N = 80
    Nvisible steps on the escalator
    vescalator speed, in steps per second
    20, 16seconds taken on the slow and fast walks
    What it says in wordsThe same number of visible steps is covered on both walks, split differently between you and the machine.

    Say the check aloud, then the sense check: the escalator at 3 steps a second is faster than either walking pace, which is plausible for a long escalator. If the question had you walking down an up escalator, the escalator's steps would subtract instead of add, and the same method still works. The trap in variants is mixing up steps and seconds; keep one unit for each quantity.

    Where candidates lose it

    The usual loss is treating the step counts as if they were times, writing 20 + 20v = 32 + 32v, or averaging 20 and 32. The escalator helps for as long as you are on it, and the fast walk is shorter: 16 seconds, not 32.

    The second is solving for the speed and stopping. The question asks for the visible steps; plug back in and check both walks give 80.

    What the interviewer asks next

    • How long does the climb take if you stand still?
    • You now walk down the same escalator while it moves up, at 4 steps a second. How many steps do you take?
    • A second escalator is twice as fast. How many steps does the slow walker take on it, for the same length?

    Asked at Susquehanna International Group, Quantitative Trading, New York, 2026 (Wall Street Oasis): A stairs question, ask for some physics m/s type of questions

  8. 039A surveillance screen flags suspicious trades. One order in 100 is genuinely manipulative. Alert A fires with a likelihood ratio of 9, and an independent alert B with a likelihood ratio of 4. Both fire on the same order: what is the probability it is manipulative?Conditional probability and BayesCoreCitadelMiami · 2022

    Try it first

    Both alerts fire. Roughly how likely is the order manipulative?

    Show the worked solution

    About 26.7%. Work in odds. The prior odds are 1 to 99. Independent evidence multiplies the odds by each likelihood ratio: 1 x 9 x 4 = 36, so the posterior odds are 36 to 99. As a probability that is 36/135, about 26.7%. Even with both alerts, roughly three flagged orders in four are clean, because manipulation is rare to begin with.

    Why is odds form the fast way to combine alerts?

    Think of two smoke detectors in a kitchen where real fires are rare. Each beep makes a fire more likely, but toast sets both off far more often than fire does. In odds form, Bayes' rule is one multiplication per piece of independent evidence: posterior odds equal prior odds times each likelihood ratioHow much more often the evidence appears when the hypothesis is true than when it is false.. A ratio of 9 means alert A fires nine times as often on manipulative orders as on clean ones, for example on 90% of manipulative orders and 10% of clean ones.

    In odds form, each independent alert multiplies: 1:99, then 9:99, then 36:99Before any alertodds 1 : 991.0%Alert A firesodds 9 : 998.3%x 9Alert B also firesodds 36 : 9926.7%x 4red share: manipulative; light share: clean36 / (36 + 99) = 26.7%
    Starting from odds of 1 to 99, alert A multiplies the odds by 9 to reach 9 to 99, an 8.3% chance, and alert B multiplies by 4 to reach 36 to 99, which is only 26.7% because the prior was so low.

    How do you check 26.7% by counting?

    Take 10,000 orders: 100 manipulative and 9,900 clean. Suppose A fires on 90% of manipulative orders and 10% of clean ones, and B on 80% and 20%, which gives the stated ratios of 9 and 4. Both fire on 100 x 0.9 x 0.8 = 72 manipulative orders and on 9,900 x 0.1 x 0.2 = 198 clean ones. Of the 270 orders where both fire, 72 are manipulative: 26.7%, the same as the odds route.

    The relationship
    P(M∣A,B)P(Mˉ∣A,B)=199×9×4=3699  ⇒  P=36135≈26.7%\frac{P(M\mid A,B)}{P(\bar M\mid A,B)} = \frac{1}{99}\times 9\times 4 = \frac{36}{99} \;\Rightarrow\; P = \frac{36}{135} \approx 26.7\%
    Mthe order is manipulative
    1/99prior odds: 1 manipulative order per 99 clean
    9, 4likelihood ratios of alerts A and B
    What it says in wordsMultiply the prior odds by each alert's likelihood ratio, then turn odds back into a probability.

    State the assumption that made multiplication legal: the alerts are independent given the truth. If both alerts key off the same feature, say order size, the second adds little new information and multiplying by 4 overstates the case. With one alert alone the chance is 8.3% for A and 3.9% for B, which is why a desk reviews orders on combined evidence rather than a single flag.

    Where candidates lose it

    The common loss is treating a likelihood ratio of 36 as odds of 36 to 1 and answering about 97%. That throws away the base rate: the evidence multiplies the prior odds of 1 to 99, not even odds.

    The second is adding the ratios, 9 + 4 = 13, instead of multiplying. Independent evidence compounds, and odds form makes that one line of arithmetic.

    What the interviewer asks next

    • How many independent alerts with a ratio of 4 would you need to pass 50%?
    • Alert B is triggered by the same feature as alert A. How does that change your answer?
    • A third alert has a likelihood ratio of 0.5 and does not fire. What does that do?

    Asked at Citadel, Sales and Trading, Miami, 2022 (Wall Street Oasis): I got a question about Bayes' theorem applied to a practical scenario

  9. 041Five observations come from a uniform distribution on 0 to theta: 3.1, 7.4, 5.2, 9.0 and 1.8. What is the maximum likelihood estimate of theta, is it biased, and how would you correct it?Statistics and estimationCoreACAQR Capital ManagementTown of Greenwich · 2022

    Try it first

    What is the maximum likelihood estimate of theta?

    Show the worked solution

    The MLE is 9.0, the largest observation; it is biased low, and multiplying by (n + 1)/n = 6/5 gives an unbiased 10.8. Each observation has density 1/theta when theta covers it, so the likelihood is theta to the minus 5 for theta at least 9.0 and zero below. That peaks at 9.0. But the sample maximum averages 5/6 of theta, never above it, so scale it up by 6/5.

    Why does the likelihood peak at the largest observation?

    Suppose raffle tickets are numbered 1 to N and you see five, the highest being 90. N cannot be below 90, and the smaller N is, the more likely it was to produce those particular five tickets. For a uniform on 0 to theta, each observation has density 1/theta, so the likelihood is theta to the minus 5, which only falls as theta grows, but it is zero for any theta below an observation. The best allowed value is the smallest theta that covers all the data: the maximum, 9.0. Calculus does not help here, because the peak sits at the edge where the likelihood jumps from zero.

    The likelihood is zero below the largest observation and falls after itzero: some observation would exceed thetapeak at theta = 9.0: the MLE40% of peak at 10.824% of peak at 12thetaL(theta) = theta to the power -5, for theta at least 9.0The data and three estimates on one line0246810121416MLE 9.02 x mean = 10.66/5 x 9.0 = 10.8unbiased, and much tighter
    The likelihood is zero for theta below 9.0, peaks at 9.0 and then falls as theta to the minus 5, down to 40% of the peak at 10.8; on the data line, the MLE of 9.0 sits at the largest observation, the method of moments gives 10.6 and the bias-corrected estimate is 10.8.

    Why is 9.0 biased, and what is the right correction?

    The sample maximum can never exceed theta, so it can only err on the low side. Five points drop into 0 to theta and cut it into six gaps of the same average size, so the largest point sits on average one gap short of theta: at 5/6 of theta. Scaling the maximum by (n + 1)/n removes that bias: 9.0 x 6/5 = 10.8. The same logic underlies the classic serial-number estimation problem from wartime production counts.

    The relationship
    L(θ)=θ−5 1{θ≥9.0}E[max⁡]=nn+1θ  ⇒  θ^=65×9.0=10.8L(\theta) = \theta^{-5}\,\mathbf{1}\{\theta \ge 9.0\} \qquad E[\max] = \frac{n}{n+1}\theta \;\Rightarrow\; \hat\theta = \frac{6}{5}\times 9.0 = 10.8
    thetathe unknown upper end of the uniform
    n = 5number of observations
    maxthe largest observation, 9.0
    What it says in wordsThe likelihood peaks at the sample maximum, which on average falls short of theta by a factor n/(n + 1), so scale it up.

    An interviewer may ask why not use twice the mean, 2 x 5.3 = 10.6, which is also unbiased. The corrected maximum is far more precise: its variance is theta squared over n(n + 2), against theta squared over 3n for twice the mean, so twice the mean is 2.3 times as variable with five points. Twice the mean can even land below the largest observation, an estimate the data have already ruled out. The limitation of the correction is that unbiased is not the only goal: the multiple of the maximum with the smallest mean squared error is (n + 2)/(n + 1), which gives 10.5 here, and saying you would choose by the loss that matters shows you know the trade.

    Where candidates lose it

    The common loss is setting the derivative of the log-likelihood to zero, getting -5/theta = 0, and concluding there is no maximum. The maximum is at a boundary, where the indicator switches on, and that is the point of the question.

    The second is answering 9.0 and stopping. The follow-up is always the bias; say that the maximum sits below theta on average and give the (n + 1)/n correction with its one-line reason.

    What the interviewer asks next

    • What is the MLE if the distribution is uniform on theta to 2 theta?
    • Derive the variance of the corrected estimator.
    • The observations come from a uniform on theta minus 1 to theta plus 1. What is the MLE now?

    Asked at AQR Capital Management, Trading, Town of Greenwich, 2022 (Wall Street Oasis): Derive the mle for some given distribution. Explain linear regression intuitively and derive the ols estimate.

  10. 042A queue holds between 0 and 3 orders. Each tick at most one thing happens: with probability 0.3 a new order arrives (if there is room), with probability 0.5 one order is filled (if the queue is not empty), and otherwise nothing changes. In the long run, what fraction of ticks is the queue full?Markov chains and random walksCoreDRWNew York · 2026

    Try it first

    Roughly what share of ticks is the queue full?

    Show the worked solution

    27/272, about 9.9% of ticks. In a birth-death chain the long-run flow up across each cut equals the flow down, so share(k) x 0.3 = share(k + 1) x 0.5. Each state's share is 0.6 times the one below: weights 1, 0.6, 0.36 and 0.216, summing to 2.176. The full state gets 0.216/2.176, about 9.9%, and the queue is empty about 46% of the time.

    Why can you skip solving the full set of equations?

    Stand at a doorway between two rooms at a party that has settled down. Over an evening, the number of people walking through one way must match the number walking back, or one room would keep filling. In a chain that only steps up or down by one, the long-run flow across the boundary between neighbouring states must balance, which gives one simple equation per cut. Here flow up from state k is its share times 0.3, and flow down from state k + 1 is its share times 0.5.

    Birth-death chain: across each cut, flow up equals flow down0 ordersempty1 order2 orders3 ordersfull0.30.50.30.50.30.5Cut balance: share(k) x 0.3 = share(k + 1) x 0.5, so each share is 0.6 times the last46.0%weight 127.6%weight 0.616.5%weight 0.369.9%weight 0.216
    Arrivals push the queue up with probability 0.3 and fills pull it down with probability 0.5, so each state's long-run share is 0.6 times the one below: 46.0% empty, 27.6% with one order, 16.5% with two and 9.9% full.

    How do the cut equations give the answer?

    Write each share relative to the empty state. Each cut gives share(k + 1) = share(k) x 0.3/0.5 = 0.6 x share(k), so the weights are 1, 0.6, 0.36 and 0.216. They sum to 2.176, so the full queue holds 0.216/2.176 = 27/272 of the time, about 9.9%. The staying probabilities, 0.2 in the middle states and 0.5 when full, never enter; a chain that pauses on a state does not change the balance across cuts.

    The relationship
    πk+1=πk⋅0.30.5π3=0.631+0.6+0.62+0.63=27272≈9.9%\pi_{k+1} = \pi_k\cdot\frac{0.3}{0.5} \qquad \pi_3 = \frac{0.6^3}{1 + 0.6 + 0.6^2 + 0.6^3} = \frac{27}{272} \approx 9.9\%
    pi_klong-run share of ticks with k orders in the queue
    0.3chance of an arrival when there is room
    0.5chance of a fill when the queue is not empty
    What it says in wordsEach state is visited 0.6 times as often as the one below it; normalise the four weights to add to one.

    Check with conservation. Orders accepted per tick are 0.3 x (1 - 0.099) = 0.2702, and orders filled per tick are 0.5 x (1 - 0.460) = 0.2702: the same, as they must be. That gives a useful business number: arrivals turned away because the queue is full run at 0.3 x 0.099, about 0.030 per tick, or one arrival in ten. The model assumes one event per tick; if an arrival and a fill could happen in the same tick, the chain changes and so do the numbers.

    Where candidates lose it

    The common loss is assuming the four states are equally likely, or writing out all four balance equations with the self-loops and solving a 4 by 4 system under time pressure. The cut method needs three one-line ratios.

    The second is inverting the ratio, using 0.5/0.3, which makes the full state the most common. Fills are faster than arrivals, so the queue must lean toward empty; check the direction before you normalise.

    What the interviewer asks next

    • What is the average queue length?
    • What arrival probability would make the queue full 25% of the time?
    • How does the answer change if the queue can hold unlimited orders?

    Asked at DRW, Quantitative Research, New York, 2026 (Wall Street Oasis): There was a problem on Chi-squared distributions which was difficult and also one on birth death chains.

← PreviousPage 3 of 8
  1. 1
  2. 2
  3. 3
  4. 4
  5. …
  6. 8
Next →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.