Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant interview preparation

Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.

Jump to the question bank
Go deeper

Quant & Hedge Fund Analyst Bootcamp

Question banks tell you what gets asked. This course gives you the work behind an answer that survives a follow-up.

Explore the course →
Question bank

100 questions, mapped to the firms that asked them

Questions
100
Traced to a firm
53
Firms
15
Updated
September 2026
Asked at
All firmsOld Mission Capital12Tower Research Capital10Jump Trading7Akuna Capital5Citadel4DED.E. Shaw3Jane Street3ACAQR Capital Management2DRW2Millennium Management2Schonfeld2SCSquarepoint Capital2Susquehanna International Group2Belvedere Trading1Optiver1
Topic
All topicsProbability10Coins, cards and games6Expected value8Statistics11Market making15Estimation and mental maths4Stochastic processes4Regression5Machine learning6Time series6Programming10Options and derivatives8Fit and motivation7
Level
AnyCoreIntermediateHard
Type
AnyBrainteaserTechnicalCaseMarket viewFit
Showing 11–20 of 47 · filtered from 100Clear filters
  1. 024I have two children and at least one is a boy. What is the probability both are boys?ProbabilityIntermediatetechnicalProp trading firmsQuant trading

    Say this

    One third, if the information came from a statement about the pair. The sample space is BB, BG, GB, GG, the condition kills GG, and one of the three survivors is BB. But the answer becomes a half if you learned it by meeting one specific child.

    Then walk it

    1. Equally likely and independent births give four ordered outcomes, each one quarter. Conditioning on at least one boy leaves three, of which one is BB. So one third.
    2. Now the version that makes it a real question. Suppose instead I introduce you to my elder child and he is a boy. Now you have conditioned on the elder being a boy, which leaves BB and BG, so the answer is one half.
    3. Same words in English, different conditioning event, different answer. The phrase at least one is a boy is a statement about the pair; this is my son is a statement about a position.
    4. The famous extension is the Tuesday boy: at least one is a boy born on a Tuesday. Now the answer is 13/27, because the extra detail changes how many pairs satisfy the condition and it breaks the symmetry between the two children.
    5. What I would actually say in an interview: the answer is one third under the standard reading, and then immediately name the ambiguity, because the entire point of the question is whether you notice that the conditioning event is underspecified.

    Where candidates lose it

    Answering one half on instinct, or answering one third and stopping. Both are half answers. Give one third with the sample space, then say precisely which conditioning event gives a half, because a quant interviewer is testing whether you can spot an ill-posed conditioning statement, which is a daily hazard in real data work.

    Expect next

    • Now: at least one is a boy born on a Tuesday.
    • What if I tell you my eldest is a boy?
    • How does this relate to survivorship bias in a dataset?
  2. 025What is the expected number of fair coin flips to see two heads in a row, and how does it compare to heads followed by tails?Stochastic processesIntermediatetechnicalQuant tradingQuant research

    Say this

    Six flips for HH and four for HT. They differ because HH can destroy its own progress: a tail after a single head sends you back to nothing, while for HT a head after a head keeps you one step from done.

    Then walk it

    1. Set up states for HH. Let A be the expected flips from scratch and B from having one head. A equals 1 plus half A plus half B. B equals 1 plus half times 0 plus half A.
    2. Substitute: B equals 1 plus A/2, so A equals 1 plus A/2 plus (1 plus A/2)/2, which gives A equals 1.5 plus 0.75A, so 0.25A equals 1.5 and A equals 6.
    3. Now HT. Let A be from scratch, B from having a head. A equals 1 plus half A plus half B. But B equals 1 plus half times 0 plus half B, because another head leaves you still in state B rather than resetting. So B equals 2.
    4. Then A equals 1 plus A/2 plus 1, so A/2 equals 2 and A equals 4.
    5. The lesson worth saying out loud: patterns with self-overlap take longer. Both patterns have probability 1/4 per pair of positions, yet the waiting times differ, and that is purely about overlap structure. It generalises: the expected wait for a pattern equals the sum of 2 to the power of the length of each of its self-overlapping prefixes. HH gives 4 plus 2 equals 6, HT gives 4 plus 0 equals 4.

    Where candidates lose it

    Assuming both answers are 4 because each two-flip pattern has probability a quarter. That is the intuition the question is designed to break. Set up the state equations explicitly and pay attention to where a failed attempt lands you, because that is the only difference between the two problems.

    Expect next

    • Now do HHH.
    • In a race between HH and HT, which appears first and with what probability?
    • Derive it with the martingale approach instead.
  3. 026You start with fifty dollars and bet a dollar on a fair coin each time. What is the probability you reach a hundred before going broke, and how does it change if the coin is slightly against you?Stochastic processesIntermediatetechnicalQuant tradingQuant research

    Say this

    In a fair game it is exactly one half, because your wealth is a martingale and the stopping value must average back to fifty. Tilt the odds slightly against you and the probability collapses, not linearly but exponentially in the number of steps.

    Then walk it

    1. Fair case: wealth is a martingale, so by optional stopping, 50 equals 100 times p plus 0 times (1 minus p), giving p equal to 0.5. In general starting at a with an upper barrier b, the probability is a over b.
    2. Biased case: with win probability q the hitting probability is (1 minus r to the a) over (1 minus r to the b) where r is (1-q)/q.
    3. Put a number on it. At q equal to 0.49, r is about 1.0408. With a equal to 50 and b equal to 100, the probability of reaching 100 drops to roughly 12 percent. A one percent edge against you turns a coin flip into a 1-in-8 shot.
    4. That sensitivity is the entire lesson. Expected value per bet is minus two cents, which sounds trivial, but over the hundreds of bets you need to walk the barrier it compounds into near certainty of ruin.
    5. And the practical version on a desk: expected time to absorption in the fair case is a times (b minus a), so 50 times 50 equals 2,500 bets. Casinos and market makers both live on this asymmetry. Small edge, high repetition, deep pockets.

    Where candidates lose it

    Giving a over b and stopping. The interesting content is how brutally the biased case differs, and candidates who cannot state the r to the power formula usually also guess that a one percent edge changes the answer by about one percent. It changes it from 50 percent to 12 percent. Put a number on it.

    Expect next

    • What is the expected number of bets until you stop?
    • What happens if you bet your whole stack each time instead?
    • How does this relate to a trader's drawdown limit?
  4. 029There are n distinct types of card in cereal boxes, uniformly at random. How many boxes do you expect to buy to collect all n?Expected valueIntermediatetechnicalQuant tradingQuant research

    Say this

    n times the harmonic number H_n, which is roughly n times (ln n plus 0.577). For 50 cards that is about 225 boxes, so four and a half times the number of cards.

    Then walk it

    1. Decompose by waiting times. Once you hold k distinct cards, the chance the next box is new is (n minus k)/n, so the wait for the next new card is geometric with mean n/(n minus k).
    2. Sum over k from 0 to n minus 1: n times (1/n plus 1/(n-1) up to 1/1), which is n H_n.
    3. Numbers: n equal to 6 gives 14.7 boxes, n equal to 50 gives 224.9, n equal to 365 gives about 2,364. The last one is the expected days to see every birthday.
    4. The tail is where the cost is. Getting the first half of the set takes about 0.69n boxes; the last single card alone takes n boxes in expectation. Most of the pain is the final few.
    5. Variance is worth flagging: it is about n squared times pi squared over 6, so the standard deviation is roughly 1.28n. For n equal to 50 that is 64 boxes, which is enormous relative to the mean of 225. Quoting the mean without the spread would be misleading if you were budgeting for it.

    Where candidates lose it

    Trying to compute it by inclusion-exclusion over the whole collection. The decomposition into independent geometric waits plus linearity of expectation is the intended route and it takes twenty seconds. Also note the harmonic sum by name, because the log growth is the insight the interviewer wants.

    Expect next

    • What is the variance?
    • What if the cards are not equally likely?
    • How many boxes for a 90 percent chance of completing the set?
  5. 030You draw n independent uniforms on zero to one. What are the expected values of the maximum and the minimum, and of the kth smallest?StatisticsIntermediatetechnicalQuant researchQuant trading

    Say this

    The maximum has mean n/(n+1), the minimum 1/(n+1), and the kth smallest k/(n+1). The n points cut the interval into n plus 1 gaps that are exchangeable, so each gap averages 1/(n+1).

    Then walk it

    1. Derive the max directly: P(max at most x) is x to the n, so the density is n x to the n minus 1, and the integral of x times that from 0 to 1 is n/(n+1).
    2. The gap argument is faster and generalises. The n order statistics plus the two endpoints create n plus 1 spacings, which are exchangeable with total length 1, so each has mean 1/(n+1). The kth order statistic is the sum of the first k spacings, hence k/(n+1).
    3. The kth order statistic is Beta(k, n minus k plus 1), which gives you the variance too: k(n-k+1) over ((n+1) squared (n+2)).
    4. Numbers: with 10 draws the max averages 0.909 and the min 0.091. With 100 draws the max averages 0.990. The max creeps to the boundary at rate 1/n, which is why extreme-value estimates converge slowly.
    5. Why a quant desk cares: the max of n draws is your model for the best of n signals, the worst drawdown of n periods, and the winning quote in an auction with n bidders. And it explains selection bias, because the best of a hundred backtests looks good even when none of them has any edge.

    Where candidates lose it

    Answering only for the max with a calculus derivation and then being stuck on the general kth. Learn the spacings argument, it gives all of them at once. And be ready to connect it to selection bias, because the practical follow-up is almost always about why the best of many strategies overstates its own quality.

    Expect next

    • What is the variance of the maximum?
    • What is the expected range, max minus min?
    • How does this explain the selection bias in picking the best of a hundred backtests?
  6. 031Break a stick at two uniformly random points. What is the probability the three pieces form a triangle?ProbabilityIntermediatetechnicalProp trading firmsQuant trading

    Say this

    One quarter. Let the cuts be x and y on a stick of length one. The triangle condition is that no piece exceeds one half, and that region is a quarter of the unit square.

    Then walk it

    1. The triangle inequality for three pieces summing to 1 reduces to a single condition: every piece must be strictly less than 1/2. If any piece is at least a half it is at least as long as the other two together.
    2. Draw the unit square in x and y. Take x less than y without loss of generality, which is the lower triangle of area 1/2. The three pieces are x, y minus x, and 1 minus y.
    3. The three conditions x less than 1/2, y minus x less than 1/2, and 1 minus y less than 1/2 carve out the middle triangle with vertices at (0, 1/2), (1/2, 1/2) and (1/2, 1). That has area 1/8.
    4. Double it for the other ordering and divide by the total area 1, giving 1/4.
    5. Different setup, different answer, and this is the part worth saying: if instead you break the stick once and then break the longer piece, the probability drops to 2 ln 2 minus 1, about 0.386. The phrase break at two random points must mean both cuts on the original stick, and you should confirm that reading before you compute.

    Where candidates lose it

    Not reducing the three triangle inequalities to the single condition no piece over a half. Candidates who try to handle three inequalities geometrically in one pass usually get 1/2 or 1/8. Also state the sampling scheme, because the sequential-break version has a completely different answer and interviewers use the ambiguity deliberately.

    Expect next

    • Now break the stick once and then break the longer piece.
    • What is the expected length of the longest piece?
    • What is the probability the triangle is obtuse?
  7. 033A test for a disease is 99 percent accurate and the disease affects one in ten thousand people. You test positive. What is the probability you have it?ProbabilityIntermediatephone / first roundQuant researchQuant trading

    Say this

    About one percent. Out of a million people, 100 are sick and about 99 of them test positive, while 999,900 are healthy and about 9,999 of them test positive falsely. So 99 out of roughly 10,098 positives are real, which is 0.98 percent.

    Then walk it

    1. Do it in counts, not Bayes notation. A population of a million makes the arithmetic trivial and the answer intuitive.
    2. The formula check: P(sick given positive) equals 0.0001 times 0.99 divided by (0.0001 times 0.99 plus 0.9999 times 0.01), which is 0.000099 over 0.010098, about 0.0098.
    3. The driver is base rate. False positives from the huge healthy population swamp the true positives from the tiny sick population. At a prevalence of 1 in 10,000 and a 1 percent false positive rate, you get a hundred false positives for every true one before adjusting for sensitivity.
    4. So the useful quantity is the likelihood ratio: 0.99 over 0.01 equals 99. It multiplies your prior odds of 1 in 9,999 into posterior odds of about 99 in 9,999, which is 1 percent. Thinking in odds and likelihood ratios is far faster than the fraction form.
    5. Where this shows up in trading: any rare-event detector, from fraud flags to regime-change signals to strategy alerts. A signal with 99 percent accuracy on a one-in-ten-thousand event fires 99 false alarms per real one, which is why alert systems get ignored.

    Where candidates lose it

    Answering 99 percent. The second trap is being sloppy about what 99 percent accurate means, since sensitivity and specificity need not be equal. State your reading, do it in counts per million, and name base rate neglect as the reason the intuitive answer is wrong by two orders of magnitude.

    Expect next

    • What prevalence would make the positive predictive value fifty percent?
    • You test positive twice. Now what?
    • How does this apply to a trading signal that fires rarely?
  8. 038How would you fix violations of the OLS assumptions?RegressionIntermediatetechnicalACAQR Capital ManagementInvestments · Greenwich · 2022

    Say this

    Depends which assumption breaks, and the fixes fall into two very different classes: violations that only break your standard errors, and violations that break the coefficients themselves. The first class you patch; the second class you have to re-specify the model.

    Then walk it

    1. Heteroskedasticity and autocorrelated errors: coefficients stay unbiased, only inference is wrong. Fix with White or Newey-West robust standard errors, or clustered errors if the dependence is by group. Cheap fix, always worth doing on financial data.
    2. Endogeneity, meaning a regressor correlated with the error, whether from omitted variables, simultaneity or measurement error: this biases the coefficients and no standard error fix helps. You need an instrument, a control for the omitted factor, a fixed effect, or a different design.
    3. Multicollinearity: coefficients are still unbiased but the variances explode and the signs flip sample to sample. Drop or combine the collinear regressors, use ridge, or work with principal components. And check the variance inflation factors before you interpret anything.
    4. Non-normal or fat-tailed errors: inference is still fine asymptotically thanks to the CLT, but outliers dominate the fit because OLS minimises squares. Use robust regression, Huber loss or quantile regression, and always look at the influence diagnostics.
    5. Non-linearity: add the relevant transform or interaction rather than pretending it away. And I would say the order I actually work in on real data: plot residuals against fitted values and against time first, because most violations announce themselves visually before any test does.

    Where candidates lose it

    Listing fixes without separating what biases the coefficients from what only biases the standard errors. That distinction is the question. Slapping Newey-West errors on an endogenous regression is a common and useless move, and an interviewer at a research shop will push on exactly that.

    Expect next

    • Which of those actually biases your coefficients?
    • How do you detect endogeneity if you have no instrument?
    • What do you do when the residuals are fat-tailed and autocorrelated at the same time?

    Reported by candidates at AQR Capital Management (Investments, Greenwich, 2022). Source: Wall Street Oasis.

  9. 039What are the differences between Lasso and Ridge regression?Machine learningIntermediatetechnicalTower Research CapitalTrading · Princeton · 2018

    Say this

    Both add a penalty on coefficient size to trade variance for bias. Ridge penalises the sum of squares and shrinks everything smoothly towards zero without eliminating anything. Lasso penalises the sum of absolute values and sets coefficients exactly to zero, so it selects features.

    Then walk it

    1. The geometry explains it. The L1 constraint region is a diamond with corners on the axes, so the solution tends to land on a corner, which means a zero coefficient. The L2 region is a ball with no corners, so solutions are interior and nothing is exactly zero.
    2. Ridge has a closed form, beta equals (X'X plus lambda I) inverse X'y, which is why it also fixes a singular X'X. Lasso has no closed form and needs coordinate descent or LARS.
    3. Correlated predictors behave very differently. Ridge splits the weight across a group of correlated features, which is stable. Lasso arbitrarily picks one and zeroes the rest, which is unstable across samples. Elastic net, which mixes both penalties, exists precisely to get sparsity without that instability.
    4. In a Bayesian reading, ridge is a Gaussian prior on the coefficients and lasso is a Laplace prior. The Laplace prior's spike at zero is what produces exact zeros.
    5. What I would say about which to use on financial data: predictors are usually highly correlated and the signal-to-noise ratio is awful, so ridge or elastic net typically beats pure lasso out of sample. Lasso is attractive when you need an interpretable short list of factors, but do not confuse the features it selected with the features that matter, because a slightly different sample gives you a different list.

    Where candidates lose it

    Stopping at L1 gives sparsity, L2 does not. Everyone says that. The differentiators are the diamond-versus-ball geometry, the behaviour under correlated predictors, and the Bayesian priors. Also always say that both require standardised features, because the penalty is scale-dependent and forgetting to standardise silently ruins the fit.

    Expect next

    • What is elastic net for?
    • How do you choose lambda?
    • Why do you have to standardise your features first?

    Reported by candidates at Tower Research Capital (Trading, Princeton, 2018). Source: Wall Street Oasis.

  10. 047Your regression has two highly correlated predictors. What happens, how do you detect it, and what do you do?RegressionIntermediatetechnicalQuant researchRisk

    Say this

    The coefficients stay unbiased but their variances blow up, so individual t-statistics collapse and signs flip from sample to sample while the overall fit and the joint prediction stay fine. Detect it with variance inflation factors or the condition number, then either combine the predictors or regularise.

    Then walk it

    1. The mechanism: the variance of a coefficient is proportional to 1 over (1 minus R squared of that regressor on the others). At a pairwise correlation of 0.95 the variance inflation factor is about 10, so your standard error is roughly three times larger than it would otherwise be.
    2. The tell-tale symptom is a regression with a high overall R squared and an F test that rejects, but no individual coefficient significant. That combination is almost always collinearity.
    3. Detection: VIFs above 5 or 10 as a rough flag, or the condition number of the scaled X matrix above 30. Better still, look at the eigenvalues of the correlation matrix, since a near-zero eigenvalue is the direction that is unidentified.
    4. Fixes in order of preference: drop one if they are measuring the same thing, combine them into a single factor such as a sum or a principal component, or use ridge, which trades a little bias for a large variance reduction and is the textbook answer for exactly this problem.
    5. The thing to say before they ask: if you only care about prediction, collinearity is close to harmless, because the fitted values are stable even when the coefficients are not. It only matters if you want to interpret the individual coefficients or attribute risk to individual factors. That is why it is a bigger problem in a risk model than in a forecasting model.

    Where candidates lose it

    Claiming collinearity biases the coefficients. It does not. And do not automatically drop a variable, because if both belong in the model economically, dropping one creates omitted variable bias, which is a worse problem than inflated variances. Distinguish the prediction case from the interpretation case.

    Expect next

    • Why does ridge help here, mathematically?
    • Is collinearity a problem if you only care about forecasting?
    • How is this different from omitted variable bias?
← PreviousPage 2 of 5
  1. 1
  2. 2
  3. 3
  4. …
  5. 5
Next →

Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

Puzzles

100 Quant puzzles, solved step by step

Try each one before you read the answer: probability, mental maths and the brainteasers interviewers use to watch you think.

Solve the puzzles →
Case studies

100 Quant case studies, worked step by step

A business, its numbers and a task, as in an assessment day or a case round. Work it on paper, then open the solution one step at a time.

Work the cases →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.