Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant interview preparation

Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.

Jump to the question bank
Go deeper

Quant & Hedge Fund Analyst Bootcamp

Question banks tell you what gets asked. This course gives you the work behind an answer that survives a follow-up.

Explore the course →
Question bank

100 questions, mapped to the firms that asked them

Questions
100
Traced to a firm
53
Firms
15
Updated
September 2026
Asked at
All firmsOld Mission Capital12Tower Research Capital10Jump Trading7Akuna Capital5Citadel4DED.E. Shaw3Jane Street3ACAQR Capital Management2DRW2Millennium Management2Schonfeld2SCSquarepoint Capital2Susquehanna International Group2Belvedere Trading1Optiver1
Topic
All topicsProbability10Coins, cards and games6Expected value8Statistics11Market making15Estimation and mental maths4Stochastic processes4Regression5Machine learning6Time series6Programming10Options and derivatives8Fit and motivation7
Level
AnyCoreIntermediateHard
Type
AnyBrainteaserTechnicalCaseMarket viewFit
Showing 1–5 of 5 · filtered from 100Clear filters
  1. 039What are the differences between Lasso and Ridge regression?Machine learningIntermediatetechnicalTower Research CapitalTrading · Princeton · 2018

    Say this

    Both add a penalty on coefficient size to trade variance for bias. Ridge penalises the sum of squares and shrinks everything smoothly towards zero without eliminating anything. Lasso penalises the sum of absolute values and sets coefficients exactly to zero, so it selects features.

    Then walk it

    1. The geometry explains it. The L1 constraint region is a diamond with corners on the axes, so the solution tends to land on a corner, which means a zero coefficient. The L2 region is a ball with no corners, so solutions are interior and nothing is exactly zero.
    2. Ridge has a closed form, beta equals (X'X plus lambda I) inverse X'y, which is why it also fixes a singular X'X. Lasso has no closed form and needs coordinate descent or LARS.
    3. Correlated predictors behave very differently. Ridge splits the weight across a group of correlated features, which is stable. Lasso arbitrarily picks one and zeroes the rest, which is unstable across samples. Elastic net, which mixes both penalties, exists precisely to get sparsity without that instability.
    4. In a Bayesian reading, ridge is a Gaussian prior on the coefficients and lasso is a Laplace prior. The Laplace prior's spike at zero is what produces exact zeros.
    5. What I would say about which to use on financial data: predictors are usually highly correlated and the signal-to-noise ratio is awful, so ridge or elastic net typically beats pure lasso out of sample. Lasso is attractive when you need an interpretable short list of factors, but do not confuse the features it selected with the features that matter, because a slightly different sample gives you a different list.

    Where candidates lose it

    Stopping at L1 gives sparsity, L2 does not. Everyone says that. The differentiators are the diamond-versus-ball geometry, the behaviour under correlated predictors, and the Bayesian priors. Also always say that both require standardised features, because the penalty is scale-dependent and forgetting to standardise silently ruins the fit.

    Expect next

    • What is elastic net for?
    • How do you choose lambda?
    • Why do you have to standardise your features first?

    Reported by candidates at Tower Research Capital (Trading, Princeton, 2018). Source: Wall Street Oasis.

  2. 041Explain the structure of a probabilistic graphical model you have worked with.Machine learningHardtechnicalTower Research CapitalQuantitative Research · New York · 2015

    Say this

    Pick one model you actually built and describe it in four parts: the variables, the graph and what the missing edges assert, how you did inference, and how you checked it. The missing edges are the interesting part, because a graphical model is a set of conditional independence claims.

    Then walk it

    1. Name the class first. A directed model, a Bayes net, factorises the joint as a product of each node given its parents and encodes causal or generative structure. An undirected model, a Markov random field, factorises into potentials over cliques and is better when the interactions have no natural direction.
    2. Then say what the graph buys you. Without structure, a joint over n binary variables needs 2 to the n minus 1 parameters. With a sparse graph it needs a handful per node. That reduction is the whole point, and the missing edges are the assumptions you are making.
    3. Inference: exact by belief propagation or the junction tree if the graph is a tree or has small treewidth, otherwise approximate by variational methods, loopy BP or MCMC. Say which you used and why, and say what the cost was.
    4. A concrete example is worth more than the taxonomy. A hidden Markov model is the simplest useful case: a latent state that evolves as a Markov chain with observations conditionally independent given the state. In markets people use it as a regime model, with the latent state as calm or stressed, fitted by Baum-Welch, and decoded with Viterbi.
    5. Then the honest part: on financial data the latent states are unstable, the number of regimes is not identified, and the fitted model will happily tell you the regime changed last week when it changed two months ago. So I used it as a descriptive overlay, never as a standalone signal.

    Where candidates lose it

    Reciting textbook definitions of Bayes nets and MRFs without ever describing a model you built. This question is a depth probe, and the interviewer will go three levels down on whichever model you name, so name the one you know cold. Be able to state the conditional independence your graph asserts and how you validated it.

    Expect next

    • What conditional independences does your graph assert, and did you test them?
    • How did you do inference, and what was the complexity?
    • How would you learn the graph structure from data?

    Reported by candidates at Tower Research Capital (Quantitative Research, New York, 2015). Source: Wall Street Oasis.

  3. 042Derive the update rules for alternating least squares in a matrix factorisation.Machine learningHardsuperdayTower Research CapitalQuantitative Research · New York · 2015

    Say this

    Fix one factor and the objective becomes an ordinary ridge regression in the other, so each update is a closed-form normal equation. With R approximated by U times V transpose and an L2 penalty, the update for a row of U is (V'V plus lambda I) inverse V'r.

    Then walk it

    1. Objective: minimise the sum over observed entries of (r_ij minus u_i dot v_j) squared plus lambda times the sum of the squared norms of u and v. It is non-convex jointly in U and V, but convex in each one separately. That is the entire reason alternating minimisation works here.
    2. Differentiate with respect to u_i holding V fixed. The gradient is minus 2 times the sum over observed j of (r_ij minus u_i dot v_j) v_j plus 2 lambda u_i. Set it to zero.
    3. Rearranged: (sum over observed j of v_j v_j' plus lambda I) u_i equals the sum over observed j of r_ij v_j. So u_i equals that Gram matrix inverse times the weighted sum. Symmetric for v_j with U fixed.
    4. Cost per update is k cubed for the k by k solve plus k squared per observed entry, and it parallelises perfectly by row, which is exactly why ALS beat SGD for large recommender systems.
    5. Say the limitations. It converges to a local optimum only, so initialisation matters, usually small random or SVD-based. The lambda is essential because otherwise the Gram matrix is singular for users with fewer than k observations. And it monotonically decreases the objective every half-step, so if your loss ever goes up you have a bug in the derivation, which is a useful debugging fact.

    Where candidates lose it

    Writing down the gradient-descent update instead of the closed-form solve. ALS is defined by exploiting the per-block convexity to solve exactly, not by stepping. Also do not forget the lambda I, since without it the system is singular for sparse rows, and do not sum over all j when only observed entries enter the loss.

    Expect next

    • Why does ALS converge, and to what?
    • When would you prefer SGD over ALS?
    • How would you handle implicit feedback where you only see the ones?

    Reported by candidates at Tower Research Capital (Quantitative Research, New York, 2015). Source: Wall Street Oasis.

  4. 052Explain the bias-variance tradeoff, and where a quant strategy usually sits on it.Machine learningCoretechnicalQuant researchQuant development

    Say this

    Expected prediction error decomposes into squared bias, variance and irreducible noise. Bias is how wrong your model class is on average, variance is how much your fit moves with a different sample. In financial data the noise term dominates everything, so you sit far towards the high-bias, low-variance end.

    Then walk it

    1. The decomposition: E of (y minus f hat) squared equals bias squared plus variance plus sigma squared. Only the first two are under your control.
    2. Flexible models cut bias and raise variance. A deep tree fits any shape and moves wildly with resampling. A linear model with three factors barely moves but cannot represent an interaction.
    3. The financial context is what makes the answer different from a generic machine learning answer. Signal-to-noise on returns is tiny, sigma squared swamps the other terms, and a flexible model spends all its capacity fitting noise. So simple, heavily regularised, few-parameter models win out of sample far more often than they should on pure machine learning intuition.
    4. How you find your place on the curve: cross-validation that respects the time ordering, learning curves, and watching the gap between in-sample and out-of-sample performance. If the gap is large you are on the variance side.
    5. One honest complication: the classic U-shaped curve is not the whole story. Very overparameterised models can show double descent, where test error falls again past the interpolation threshold. That is real in vision and language. I have not seen it be useful on noisy financial data, where the tiny signal means regularisation still dominates.

    Where candidates lose it

    Giving the textbook decomposition with no view on where financial data sits. Every candidate can recite bias plus variance. The differentiator is saying that low signal-to-noise pushes you towards simple models, and being able to say how you would diagnose which side you are on.

    Expect next

    • How would you diagnose which side of the tradeoff you are on?
    • Why do simple models often win on financial data?
    • What is double descent?
  5. 054What does PCA do, how do you choose the number of components, and what are its limitations on financial data?Machine learningCoretechnicalQuant researchRisk

    Say this

    It finds the orthogonal directions of maximum variance, which are the eigenvectors of the covariance matrix, and lets you describe the data with fewer numbers. Choose the number of components by explained variance, a scree elbow, or the Marchenko-Pastur bulk edge if you want a principled cutoff.

    Then walk it

    1. Mechanically: eigendecompose the covariance or correlation matrix, or take the SVD of the centred data. Eigenvalues are the variance along each component, eigenvectors are the directions.
    2. Correlation versus covariance matters. On assets with wildly different volatilities, PCA on the covariance matrix is dominated by the most volatile names, so standardise first unless the scale is meaningful.
    3. Concrete example everyone in rates knows: PCA on the yield curve gives level, slope and curvature, explaining roughly 90, 8 and 2 percent of variance. On equities the first component is the market, explaining 25 to 40 percent depending on the regime, and it rises sharply in a crisis.
    4. Choosing k: cumulative explained variance at 90 or 95 percent, the scree elbow, or eigenvalues above the random matrix bulk edge, which is the statistically defensible version because it separates signal from estimation noise.
    5. Limitations, and these are the answer to the real question. PCA maximises variance, not predictive power, so the components need not have anything to do with your target. It is unstable: eigenvectors rotate sample to sample when eigenvalues are close, so your factor two and factor three swap places. It assumes linearity. And the components are usually uninterpretable outside a structured setting like the yield curve, which makes them awkward to risk-manage.

    Where candidates lose it

    Describing PCA as dimensionality reduction and stopping. Two things get graded: that it is unsupervised so high-variance directions are not necessarily predictive, and that you must standardise when scales differ. Also have a real example ready, because level-slope-curvature or the equity market factor proves you have used it rather than read about it.

    Expect next

    • Why is PCA not necessarily good for prediction?
    • What does the first principal component of an equity universe represent, and what happens to it in a crisis?
    • How is PCA related to a factor risk model?

Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

Puzzles

100 Quant puzzles, solved step by step

Try each one before you read the answer: probability, mental maths and the brainteasers interviewers use to watch you think.

Solve the puzzles →
Case studies

100 Quant case studies, worked step by step

A business, its numbers and a task, as in an assessment day or a case round. Work it on paper, then open the solution one step at a time.

Work the cases →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.