Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant interview preparation

Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.

Jump to the question bank
Go deeper

Quant & Hedge Fund Analyst Bootcamp

Question banks tell you what gets asked. This course gives you the work behind an answer that survives a follow-up.

Explore the course →
Question bank

100 questions, mapped to the firms that asked them

Questions
100
Traced to a firm
53
Firms
15
Updated
September 2026
Asked at
All firmsOld Mission Capital12Tower Research Capital10Jump Trading7Akuna Capital5Citadel4DED.E. Shaw3Jane Street3ACAQR Capital Management2DRW2Millennium Management2Schonfeld2SCSquarepoint Capital2Susquehanna International Group2Belvedere Trading1Optiver1
Topic
All topicsProbability10Coins, cards and games6Expected value8Statistics11Market making15Estimation and mental maths4Stochastic processes4Regression5Machine learning6Time series6Programming10Options and derivatives8Fit and motivation7
Level
AnyCoreIntermediateHard
Type
AnyBrainteaserTechnicalCaseMarket viewFit
Showing 11–20 of 42 · filtered from 100Clear filters
  1. 041Explain the structure of a probabilistic graphical model you have worked with.Machine learningHardtechnicalTower Research CapitalQuantitative Research · New York · 2015

    Say this

    Pick one model you actually built and describe it in four parts: the variables, the graph and what the missing edges assert, how you did inference, and how you checked it. The missing edges are the interesting part, because a graphical model is a set of conditional independence claims.

    Then walk it

    1. Name the class first. A directed model, a Bayes net, factorises the joint as a product of each node given its parents and encodes causal or generative structure. An undirected model, a Markov random field, factorises into potentials over cliques and is better when the interactions have no natural direction.
    2. Then say what the graph buys you. Without structure, a joint over n binary variables needs 2 to the n minus 1 parameters. With a sparse graph it needs a handful per node. That reduction is the whole point, and the missing edges are the assumptions you are making.
    3. Inference: exact by belief propagation or the junction tree if the graph is a tree or has small treewidth, otherwise approximate by variational methods, loopy BP or MCMC. Say which you used and why, and say what the cost was.
    4. A concrete example is worth more than the taxonomy. A hidden Markov model is the simplest useful case: a latent state that evolves as a Markov chain with observations conditionally independent given the state. In markets people use it as a regime model, with the latent state as calm or stressed, fitted by Baum-Welch, and decoded with Viterbi.
    5. Then the honest part: on financial data the latent states are unstable, the number of regimes is not identified, and the fitted model will happily tell you the regime changed last week when it changed two months ago. So I used it as a descriptive overlay, never as a standalone signal.

    Where candidates lose it

    Reciting textbook definitions of Bayes nets and MRFs without ever describing a model you built. This question is a depth probe, and the interviewer will go three levels down on whichever model you name, so name the one you know cold. Be able to state the conditional independence your graph asserts and how you validated it.

    Expect next

    • What conditional independences does your graph assert, and did you test them?
    • How did you do inference, and what was the complexity?
    • How would you learn the graph structure from data?

    Reported by candidates at Tower Research Capital (Quantitative Research, New York, 2015). Source: Wall Street Oasis.

  2. 042Derive the update rules for alternating least squares in a matrix factorisation.Machine learningHardsuperdayTower Research CapitalQuantitative Research · New York · 2015

    Say this

    Fix one factor and the objective becomes an ordinary ridge regression in the other, so each update is a closed-form normal equation. With R approximated by U times V transpose and an L2 penalty, the update for a row of U is (V'V plus lambda I) inverse V'r.

    Then walk it

    1. Objective: minimise the sum over observed entries of (r_ij minus u_i dot v_j) squared plus lambda times the sum of the squared norms of u and v. It is non-convex jointly in U and V, but convex in each one separately. That is the entire reason alternating minimisation works here.
    2. Differentiate with respect to u_i holding V fixed. The gradient is minus 2 times the sum over observed j of (r_ij minus u_i dot v_j) v_j plus 2 lambda u_i. Set it to zero.
    3. Rearranged: (sum over observed j of v_j v_j' plus lambda I) u_i equals the sum over observed j of r_ij v_j. So u_i equals that Gram matrix inverse times the weighted sum. Symmetric for v_j with U fixed.
    4. Cost per update is k cubed for the k by k solve plus k squared per observed entry, and it parallelises perfectly by row, which is exactly why ALS beat SGD for large recommender systems.
    5. Say the limitations. It converges to a local optimum only, so initialisation matters, usually small random or SVD-based. The lambda is essential because otherwise the Gram matrix is singular for users with fewer than k observations. And it monotonically decreases the objective every half-step, so if your loss ever goes up you have a bug in the derivation, which is a useful debugging fact.

    Where candidates lose it

    Writing down the gradient-descent update instead of the closed-form solve. ALS is defined by exploiting the per-block convexity to solve exactly, not by stepping. Also do not forget the lambda I, since without it the system is singular for sparse rows, and do not sum over all j when only observed entries enter the loss.

    Expect next

    • Why does ALS converge, and to what?
    • When would you prefer SGD over ALS?
    • How would you handle implicit feedback where you only see the ones?

    Reported by candidates at Tower Research Capital (Quantitative Research, New York, 2015). Source: Wall Street Oasis.

  3. 043State the central limit theorem and tell me where it fails.StatisticsCoretechnicalQuant researchQuant trading

    Say this

    For independent identically distributed variables with finite mean and finite variance, the standardised sample mean converges in distribution to a standard normal. The key conditions are finite variance and enough independence, and both fail regularly in markets.

    Then walk it

    1. Precisely: root n times (X bar minus mu) over sigma converges in distribution to N(0,1). Note it is the standardised mean that converges, and the rate is 1 over root n.
    2. Failure one, infinite variance. A Cauchy distribution has no variance and the sample mean of Cauchys is Cauchy again, no matter how large n is. Averaging buys you nothing. More generally, stable distributions with tail index alpha below 2 converge to a stable law, not a normal.
    3. Failure two, dependence. With strongly autocorrelated data the effective sample size is far below n, so you converge much more slowly and your standard errors are too small. Long-range dependence can break it entirely.
    4. Failure three, the rate in the tails. Even where the CLT holds, convergence is fastest in the middle and slowest in the tails, which is precisely where a risk manager needs accuracy. Berry-Esseen gives an error bound of order 1 over root n times the third absolute moment, so skewed data converges slowly.
    5. The practical version: daily equity returns have kurtosis of 5 to 10 and volatility clustering, so ten-day sums are much closer to normal than daily returns, but a 99.9 percent quantile computed from a normal assumption will still understate the tail badly. That is why value at risk models use empirical or extreme-value tails rather than leaning on the CLT.

    Where candidates lose it

    Stating the theorem without the finite variance condition, or claiming everything becomes normal for large n. Also do not confuse it with the law of large numbers, which is about convergence of the mean to a constant and needs only finite mean. Be ready to say what happens with infinite variance, because that is the follow-up.

    Expect next

    • What happens with a Cauchy distribution?
    • How is that different from the law of large numbers?
    • How large does n have to be in practice for returns data?
  4. 044What is a p-value, and what is it not?StatisticsCoretechnicalQuant researchRisk

    Say this

    It is the probability of seeing data at least as extreme as what you saw, assuming the null hypothesis is true. It is not the probability that the null is true, and it is not the probability you are wrong.

    Then walk it

    1. The conditioning runs the wrong way from what people assume. A p-value is P(data given null), and what you actually want is P(null given data). Those are different objects and Bayes tells you the second depends on your prior.
    2. Concretely: if you test a thousand strategies of which fifty genuinely work, at a five percent significance level you get roughly 47 true discoveries and 47 false ones. A p-value of 0.05 in that setting means a coin flip on whether the finding is real.
    3. It also says nothing about effect size. With a million observations a completely useless one-basis-point edge will have a p-value of 0.0001. Significance is not importance, and in high-frequency data everything is significant.
    4. And it is only valid for a pre-specified test. Choosing the test after looking at the data, or stopping data collection when the p-value crosses 0.05, invalidates it completely.
    5. What I would report instead on a desk: the effect size with a confidence interval, out-of-sample performance, and how many specifications I tried. A p-value on its own is close to useless in a research process where hundreds of hypotheses get screened.

    Where candidates lose it

    Defining it as the probability the null is true. That is the single most common statistical error in finance interviews and it is disqualifying at a research shop. Also be ready with the multiple-testing consequence, because the interviewer's real target is whether you understand why published anomalies do not replicate.

    Expect next

    • So what significance level would you use if you screened a thousand signals?
    • Explain the false discovery rate.
    • What would you report instead of a p-value?
  5. 046What are the assumptions behind ordinary least squares, and which of them actually matter?RegressionCoretechnicalQuant researchRisk

    Say this

    Linearity in parameters, exogeneity meaning the error has zero mean conditional on the regressors, no perfect collinearity, homoskedasticity, and no autocorrelation. Only exogeneity is essential for unbiasedness. The last two affect efficiency and standard errors, not the coefficients.

    Then walk it

    1. Exogeneity, E of error given X equals zero, is the load-bearing assumption. Break it and every coefficient is biased and inconsistent, and no amount of data or robust standard errors saves you.
    2. Homoskedasticity and no autocorrelation give you Gauss-Markov efficiency and the usual standard error formula. Break them and OLS is still unbiased, just no longer the minimum-variance linear estimator, and your t-statistics are wrong. Robust or Newey-West errors fix the inference.
    3. Normality of errors is not needed for unbiasedness or consistency at all. It only buys exact small-sample t and F distributions. Asymptotically the CLT handles it.
    4. No perfect collinearity is a requirement for the estimator to exist, since X'X must be invertible. Near-collinearity is not a violation, it just inflates variances.
    5. On financial data the realistic picture is: heteroskedasticity almost always, autocorrelation often, and exogeneity frequently violated because everything is jointly determined. So I default to robust standard errors, and I spend my thinking time on whether my regressor is endogenous, because that is the one that actually changes the answer.

    Where candidates lose it

    Listing normality as a core assumption, or treating all five as equally important. Rank them. The interviewer wants to hear which violations bias the coefficients and which only bias the standard errors, because that distinction determines whether you patch the model or rebuild it.

    Expect next

    • Give me a concrete example of endogeneity in a returns regression.
    • Why is normality not needed?
    • What does Gauss-Markov actually claim?
  6. 047Your regression has two highly correlated predictors. What happens, how do you detect it, and what do you do?RegressionIntermediatetechnicalQuant researchRisk

    Say this

    The coefficients stay unbiased but their variances blow up, so individual t-statistics collapse and signs flip from sample to sample while the overall fit and the joint prediction stay fine. Detect it with variance inflation factors or the condition number, then either combine the predictors or regularise.

    Then walk it

    1. The mechanism: the variance of a coefficient is proportional to 1 over (1 minus R squared of that regressor on the others). At a pairwise correlation of 0.95 the variance inflation factor is about 10, so your standard error is roughly three times larger than it would otherwise be.
    2. The tell-tale symptom is a regression with a high overall R squared and an F test that rejects, but no individual coefficient significant. That combination is almost always collinearity.
    3. Detection: VIFs above 5 or 10 as a rough flag, or the condition number of the scaled X matrix above 30. Better still, look at the eigenvalues of the correlation matrix, since a near-zero eigenvalue is the direction that is unidentified.
    4. Fixes in order of preference: drop one if they are measuring the same thing, combine them into a single factor such as a sum or a principal component, or use ridge, which trades a little bias for a large variance reduction and is the textbook answer for exactly this problem.
    5. The thing to say before they ask: if you only care about prediction, collinearity is close to harmless, because the fitted values are stable even when the coefficients are not. It only matters if you want to interpret the individual coefficients or attribute risk to individual factors. That is why it is a bigger problem in a risk model than in a forecasting model.

    Where candidates lose it

    Claiming collinearity biases the coefficients. It does not. And do not automatically drop a variable, because if both belong in the model economically, dropping one creates omitted variable bias, which is a worse problem than inflated variances. Distinguish the prediction case from the interpretation case.

    Expect next

    • Why does ridge help here, mathematically?
    • Is collinearity a problem if you only care about forecasting?
    • How is this different from omitted variable bias?
  7. 048Asset volatility comes in clusters. What does that break, and how do you model it?Time seriesIntermediatetechnicalQuant researchRisk

    Say this

    It breaks the constant-variance assumption behind almost everything: OLS standard errors, iid return models and Black-Scholes. The standard answer is a GARCH model, where today's variance depends on yesterday's variance and yesterday's squared shock.

    Then walk it

    1. The empirical fact first: returns are close to unpredictable in the mean but their squares and absolute values are strongly autocorrelated, with the autocorrelation of squared returns decaying over weeks. Big moves cluster.
    2. GARCH(1,1) is sigma squared at t equals omega plus alpha times the last squared return plus beta times the last variance. On daily equities alpha is typically around 0.05 to 0.1 and beta around 0.85 to 0.92, with alpha plus beta just under one, meaning very persistent but eventually mean reverting.
    3. Long-run variance is omega over (1 minus alpha minus beta). If alpha plus beta hits one you get integrated GARCH, which is essentially an exponentially weighted moving average with no mean reversion, and that is what RiskMetrics used.
    4. It matters for options because it generates both fat unconditional tails and a term structure of volatility, which is why implied vol curves upward or downward towards the long-run level depending on where spot vol sits.
    5. Variants worth naming and the honest limitation: GJR-GARCH or EGARCH add the leverage effect, since negative returns raise vol more than positive ones, which plain GARCH cannot capture. And for anything intraday I would prefer realised volatility from high-frequency data, because a HAR model on realised vol usually forecasts better than GARCH on daily closes.

    Where candidates lose it

    Describing GARCH mechanically without saying what it is for. The point is that conditional variance is forecastable even when the mean is not, which is why volatility trading exists and directional trading is hard. Also do not forget the leverage effect, since plain GARCH is symmetric in the sign of returns and equity vol is not.

    Expect next

    • Why does alpha plus beta sit so close to one?
    • What is the leverage effect and which model captures it?
    • Would you use GARCH or realised volatility to forecast tomorrow's vol?
  8. 051What is omitted variable bias, and how would it show up in a factor regression?RegressionIntermediatetechnicalQuant researchPortfolio management

    Say this

    If you leave out a variable that belongs in the model and it is correlated with a regressor you kept, the kept coefficient absorbs part of its effect. The bias equals the true coefficient on the omitted variable times the regression coefficient of the omitted variable on the included one.

    Then walk it

    1. Formula worth knowing: if the truth is y equals b1 x1 plus b2 x2 plus e and you regress y on x1 alone, you estimate b1 plus b2 times delta, where delta is from regressing x2 on x1.
    2. So the bias has a sign you can reason about. If the omitted factor has a positive premium and your included factor loads positively on it, you overstate your factor's premium.
    3. In factor work this is everywhere. Run a single-factor CAPM regression on a value portfolio and the alpha looks large, because you omitted the value factor. Add HML and the alpha collapses. That is not a bug, it is omitted variable bias doing exactly what the formula says.
    4. Momentum is the classic trap in the other direction. Omit momentum from a regression on a quality portfolio and quality's alpha inherits whatever momentum exposure quality happens to carry in your sample.
    5. How to handle it honestly: report alpha against a nested sequence of models, one factor then three then five plus momentum, and show what survives. And say the limitation out loud, because you can never rule out the factor nobody has published yet. The defence is out-of-sample and out-of-market evidence, not a longer regression.

    Where candidates lose it

    Defining it abstractly without giving the sign and magnitude formula, or without a concrete factor example. The interviewer wants to see you reason about the direction of the bias. Also do not confuse it with multicollinearity, which inflates variance without biasing anything.

    Expect next

    • Which direction does the bias go if the omitted factor is positively correlated with yours?
    • How do you test whether your alpha survives the addition of a new factor?
    • How is this different from multicollinearity?
  9. 052Explain the bias-variance tradeoff, and where a quant strategy usually sits on it.Machine learningCoretechnicalQuant researchQuant development

    Say this

    Expected prediction error decomposes into squared bias, variance and irreducible noise. Bias is how wrong your model class is on average, variance is how much your fit moves with a different sample. In financial data the noise term dominates everything, so you sit far towards the high-bias, low-variance end.

    Then walk it

    1. The decomposition: E of (y minus f hat) squared equals bias squared plus variance plus sigma squared. Only the first two are under your control.
    2. Flexible models cut bias and raise variance. A deep tree fits any shape and moves wildly with resampling. A linear model with three factors barely moves but cannot represent an interaction.
    3. The financial context is what makes the answer different from a generic machine learning answer. Signal-to-noise on returns is tiny, sigma squared swamps the other terms, and a flexible model spends all its capacity fitting noise. So simple, heavily regularised, few-parameter models win out of sample far more often than they should on pure machine learning intuition.
    4. How you find your place on the curve: cross-validation that respects the time ordering, learning curves, and watching the gap between in-sample and out-of-sample performance. If the gap is large you are on the variance side.
    5. One honest complication: the classic U-shaped curve is not the whole story. Very overparameterised models can show double descent, where test error falls again past the interpolation threshold. That is real in vision and language. I have not seen it be useful on noisy financial data, where the tiny signal means regularisation still dominates.

    Where candidates lose it

    Giving the textbook decomposition with no view on where financial data sits. Every candidate can recite bias plus variance. The differentiator is saying that low signal-to-noise pushes you towards simple models, and being able to say how you would diagnose which side you are on.

    Expect next

    • How would you diagnose which side of the tradeoff you are on?
    • Why do simple models often win on financial data?
    • What is double descent?
  10. 053How would you cross-validate a model on time series data, and why is standard k-fold wrong?Time seriesHardtechnicalQuant researchQuant trading

    Say this

    Standard k-fold trains on data that comes after your test set, which leaks the future. You need a forward-walking scheme: train on a window, test on the next block, roll forward, and put a gap between train and test so overlapping labels do not bleed across the boundary.

    Then walk it

    1. Two distinct leaks. First, random folds put future observations in the training set, so the model learns things it could not have known. Second, features and labels are usually built from overlapping windows, so even adjacent-in-time observations share information across a fold boundary.
    2. The fix for the first is walk-forward or expanding-window validation: fit on 1 to t, test on t plus 1 to t plus h, roll. Expanding window mimics how you would actually retrain in production. A fixed rolling window is better if the process is non-stationary.
    3. The fix for the second is purging and embargoing, from Lopez de Prado. Remove training observations whose label window overlaps the test period, and embargo a short period immediately after the test block. On a 20-day forward return label you need at least a 20-day purge.
    4. Also beware the hidden leaks that sit outside the folds entirely: fitting a scaler, doing feature selection, or choosing hyperparameters on the full dataset before splitting. Every preprocessing step has to sit inside the fold.
    5. What I would actually report, and this is the part that matters: one final untouched hold-out period tested once, plus how many configurations I tried before I got there. Walk-forward validation run a hundred times is itself an overfitting device, and the number of trials is the honest measure of how much to discount the result.

    Where candidates lose it

    Saying you would use k-fold with shuffle turned off and stopping there. That fixes the ordering but not the overlapping-label leak, and interviewers at systematic shops probe exactly that. Mention purging and embargo, and mention that scalers and feature selection must live inside the fold.

    Expect next

    • How long should the embargo be?
    • Expanding window or fixed rolling window, and why?
    • How do you account for the number of configurations you tried?
← PreviousPage 2 of 5
  1. 1
  2. 2
  3. 3
  4. …
  5. 5
Next →

Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

Puzzles

100 Quant puzzles, solved step by step

Try each one before you read the answer: probability, mental maths and the brainteasers interviewers use to watch you think.

Solve the puzzles →
Case studies

100 Quant case studies, worked step by step

A business, its numbers and a task, as in an assessment day or a case round. Work it on paper, then open the solution one step at a time.

Work the cases →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.