Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Quant interview preparation

Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.

Jump to the question bank
Go deeper

Quant & Hedge Fund Analyst Bootcamp

Question banks tell you what gets asked. This course gives you the work behind an answer that survives a follow-up.

Explore the course →
Question bank

100 questions, mapped to the firms that asked them

Questions
100
Traced to a firm
53
Firms
15
Updated
September 2026
Asked at
All firmsOld Mission Capital12Tower Research Capital10Jump Trading7Akuna Capital5Citadel4DED.E. Shaw3Jane Street3ACAQR Capital Management2DRW2Millennium Management2Schonfeld2SCSquarepoint Capital2Susquehanna International Group2Belvedere Trading1Optiver1
Topic
All topicsProbability10Coins, cards and games6Expected value8Statistics11Market making15Estimation and mental maths4Stochastic processes4Regression5Machine learning6Time series6Programming10Options and derivatives8Fit and motivation7
Level
AnyCoreIntermediateHard
Type
AnyBrainteaserTechnicalCaseMarket viewFit
Showing 1–10 of 11 · filtered from 100Clear filters
  1. 027What is a martingale, and how would you use optional stopping to solve a problem?Stochastic processesHardtechnicalQuant researchQuant trading

    Say this

    A martingale is a process whose expected next value, given everything you know now, equals its current value. Optional stopping says that for a suitably bounded stopping time, the expected value at the stopping time equals the starting value, which is what turns a hard path-dependent question into one line of algebra.

    Then walk it

    1. Formally: E of X_{n+1} given the filtration F_n equals X_n. No drift, conditional on history. It is not the same as independence, and increments need not be identically distributed.
    2. Optional stopping needs a condition, and you should name one: bounded stopping time, or bounded increments plus finite expected stopping time, or uniform integrability. Without it the theorem fails, and the classic failure is the doubling strategy, where a stopping time that is finite with probability one still produces E of X_tau equal to 1 rather than 0.
    3. How I use it: find a quantity that is conserved in expectation, then evaluate it at the stopping time. Gambler's ruin falls out immediately from wealth being a martingale.
    4. A second example, expected time in a symmetric random walk: W_n squared minus n is a martingale, so E of tau equals E of W_tau squared. With barriers at 0 and b starting from a, that gives E of tau equal to a(b minus a) in a line.
    5. And the reason it matters beyond puzzles: risk-neutral pricing is exactly the statement that the discounted price is a martingale under the pricing measure. Delta hedging is the construction of that martingale. If you can say that connection, the puzzle answer becomes a conversation about derivatives.

    Where candidates lose it

    Defining a martingale as a fair game and stopping there, or applying optional stopping without checking the integrability condition. Interviewers at the good shops will hand you the doubling strategy specifically to see whether you know why the theorem does not apply. Name the condition before you use the theorem.

    Expect next

    • Why does optional stopping fail for the doubling strategy?
    • Is the square of a martingale a martingale?
    • Connect this to risk-neutral pricing.
  2. 036The sample variance with the n minus one correction is unbiased. Is its square root an unbiased estimator of the standard deviation?StatisticsHardtechnicalSCSquarepoint CapitalQuantitative Research · London · 2026

    Say this

    No. The square root is concave, so by Jensen's inequality the expected square root is strictly less than the square root of the expected value. The sample standard deviation is biased downwards, always, for any distribution with positive variance.

    Then walk it

    1. Jensen: for a strictly concave g, E of g(X) is less than g of E of X unless X is degenerate. With g the square root and X the unbiased sample variance, E of s is less than sigma.
    2. Size the bias for normal data. E of s equals c4(n) times sigma, where c4 is a known constant involving gamma functions. At n equal to 2, c4 is about 0.798, so you understate sigma by 20 percent. At n equal to 10 it is 0.9727, a 2.7 percent understatement. At n equal to 30 it is 0.9914.
    3. So the bias is order 1/(4n) and it vanishes as n grows. It is a real problem for short samples and irrelevant for long ones.
    4. Unbiasedness is also not preserved under any nonlinear transform, which is the general lesson. The unbiased estimator of sigma squared does not give you an unbiased estimator of sigma, or of 1/sigma, or of log sigma.
    5. Where this bites on a desk: annualised volatility estimated from a few weeks of data, and any Sharpe ratio, since the Sharpe divides by s. Understating s inflates the Sharpe, so short-sample Sharpes are biased upwards. That is worth saying because it connects a textbook Jensen question to a live problem in strategy evaluation.

    Where candidates lose it

    Saying yes because the variance estimator is unbiased. Unbiasedness does not survive a nonlinear function. Name Jensen explicitly, give the direction of the bias, and quantify it with c4 for at least one small n. The follow-up about Sharpe ratios is where the real conversation is, so get there yourself.

    Expect next

    • How would you correct it?
    • What does that imply for a Sharpe ratio estimated on a short sample?
    • Is the sample correlation coefficient unbiased?

    Reported by candidates at Squarepoint Capital (Quantitative Research, London, 2026). Source: Wall Street Oasis.

  3. 041Explain the structure of a probabilistic graphical model you have worked with.Machine learningHardtechnicalTower Research CapitalQuantitative Research · New York · 2015

    Say this

    Pick one model you actually built and describe it in four parts: the variables, the graph and what the missing edges assert, how you did inference, and how you checked it. The missing edges are the interesting part, because a graphical model is a set of conditional independence claims.

    Then walk it

    1. Name the class first. A directed model, a Bayes net, factorises the joint as a product of each node given its parents and encodes causal or generative structure. An undirected model, a Markov random field, factorises into potentials over cliques and is better when the interactions have no natural direction.
    2. Then say what the graph buys you. Without structure, a joint over n binary variables needs 2 to the n minus 1 parameters. With a sparse graph it needs a handful per node. That reduction is the whole point, and the missing edges are the assumptions you are making.
    3. Inference: exact by belief propagation or the junction tree if the graph is a tree or has small treewidth, otherwise approximate by variational methods, loopy BP or MCMC. Say which you used and why, and say what the cost was.
    4. A concrete example is worth more than the taxonomy. A hidden Markov model is the simplest useful case: a latent state that evolves as a Markov chain with observations conditionally independent given the state. In markets people use it as a regime model, with the latent state as calm or stressed, fitted by Baum-Welch, and decoded with Viterbi.
    5. Then the honest part: on financial data the latent states are unstable, the number of regimes is not identified, and the fitted model will happily tell you the regime changed last week when it changed two months ago. So I used it as a descriptive overlay, never as a standalone signal.

    Where candidates lose it

    Reciting textbook definitions of Bayes nets and MRFs without ever describing a model you built. This question is a depth probe, and the interviewer will go three levels down on whichever model you name, so name the one you know cold. Be able to state the conditional independence your graph asserts and how you validated it.

    Expect next

    • What conditional independences does your graph assert, and did you test them?
    • How did you do inference, and what was the complexity?
    • How would you learn the graph structure from data?

    Reported by candidates at Tower Research Capital (Quantitative Research, New York, 2015). Source: Wall Street Oasis.

  4. 042Derive the update rules for alternating least squares in a matrix factorisation.Machine learningHardsuperdayTower Research CapitalQuantitative Research · New York · 2015

    Say this

    Fix one factor and the objective becomes an ordinary ridge regression in the other, so each update is a closed-form normal equation. With R approximated by U times V transpose and an L2 penalty, the update for a row of U is (V'V plus lambda I) inverse V'r.

    Then walk it

    1. Objective: minimise the sum over observed entries of (r_ij minus u_i dot v_j) squared plus lambda times the sum of the squared norms of u and v. It is non-convex jointly in U and V, but convex in each one separately. That is the entire reason alternating minimisation works here.
    2. Differentiate with respect to u_i holding V fixed. The gradient is minus 2 times the sum over observed j of (r_ij minus u_i dot v_j) v_j plus 2 lambda u_i. Set it to zero.
    3. Rearranged: (sum over observed j of v_j v_j' plus lambda I) u_i equals the sum over observed j of r_ij v_j. So u_i equals that Gram matrix inverse times the weighted sum. Symmetric for v_j with U fixed.
    4. Cost per update is k cubed for the k by k solve plus k squared per observed entry, and it parallelises perfectly by row, which is exactly why ALS beat SGD for large recommender systems.
    5. Say the limitations. It converges to a local optimum only, so initialisation matters, usually small random or SVD-based. The lambda is essential because otherwise the Gram matrix is singular for users with fewer than k observations. And it monotonically decreases the objective every half-step, so if your loss ever goes up you have a bug in the derivation, which is a useful debugging fact.

    Where candidates lose it

    Writing down the gradient-descent update instead of the closed-form solve. ALS is defined by exploiting the per-block convexity to solve exactly, not by stepping. Also do not forget the lambda I, since without it the system is singular for sparse rows, and do not sum over all j when only observed entries enter the loss.

    Expect next

    • Why does ALS converge, and to what?
    • When would you prefer SGD over ALS?
    • How would you handle implicit feedback where you only see the ones?

    Reported by candidates at Tower Research Capital (Quantitative Research, New York, 2015). Source: Wall Street Oasis.

  5. 053How would you cross-validate a model on time series data, and why is standard k-fold wrong?Time seriesHardtechnicalQuant researchQuant trading

    Say this

    Standard k-fold trains on data that comes after your test set, which leaks the future. You need a forward-walking scheme: train on a window, test on the next block, roll forward, and put a gap between train and test so overlapping labels do not bleed across the boundary.

    Then walk it

    1. Two distinct leaks. First, random folds put future observations in the training set, so the model learns things it could not have known. Second, features and labels are usually built from overlapping windows, so even adjacent-in-time observations share information across a fold boundary.
    2. The fix for the first is walk-forward or expanding-window validation: fit on 1 to t, test on t plus 1 to t plus h, roll. Expanding window mimics how you would actually retrain in production. A fixed rolling window is better if the process is non-stationary.
    3. The fix for the second is purging and embargoing, from Lopez de Prado. Remove training observations whose label window overlaps the test period, and embargo a short period immediately after the test block. On a 20-day forward return label you need at least a 20-day purge.
    4. Also beware the hidden leaks that sit outside the folds entirely: fitting a scaler, doing feature selection, or choosing hyperparameters on the full dataset before splitting. Every preprocessing step has to sit inside the fold.
    5. What I would actually report, and this is the part that matters: one final untouched hold-out period tested once, plus how many configurations I tried before I got there. Walk-forward validation run a hundred times is itself an overfitting device, and the number of trials is the honest measure of how much to discount the result.

    Where candidates lose it

    Saying you would use k-fold with shuffle turned off and stopping there. That fixes the ordering but not the overlapping-label leak, and interviewers at systematic shops probe exactly that. Mention purging and embargo, and mention that scalers and feature selection must live inside the fold.

    Expect next

    • How long should the embargo be?
    • Expanding window or fixed rolling window, and why?
    • How do you account for the number of configurations you tried?
  6. 066Explain the Kelly criterion, and why do real traders bet less than it says?Market makingHardtechnicalQuant tradingProp trading firms

    Say this

    Kelly maximises the expected growth rate of your capital by betting a fraction equal to your edge divided by the odds. For an even-money bet at probability p, that fraction is 2p minus 1. Real traders bet a fraction of it because Kelly assumes you know your edge exactly, and overbetting is far more damaging than underbetting.

    Then walk it

    1. The derivation in one line: maximise the expected log of wealth, because log wealth is additive across repeated bets and its expectation governs the long-run growth rate. For a bet paying b to 1 with win probability p, the optimal fraction is (pb minus (1-p)) over b.
    2. Numbers: a 55 percent even-money bet gives f equal to 0.1, so ten percent of capital. A 60 percent bet gives 20 percent. That is a lot more than most people's intuition, which is the first surprise of Kelly.
    3. For continuous returns the analogue is mean over variance, which is why a Sharpe ratio maps directly to a leverage level. Full Kelly leverage equals the Sharpe divided by the volatility.
    4. The asymmetry is the key insight. Growth rate as a function of bet size is a concave parabola, so betting half Kelly gives you three quarters of the growth with half the volatility. Betting double Kelly gives you zero growth. Overestimating your edge by a factor of two therefore destroys the entire benefit.
    5. And full Kelly's drawdowns are intolerable in practice: the probability of at some point halving your capital under full Kelly is fifty percent. Nobody running other people's money survives that, and no risk manager permits it. So a quarter to a half Kelly is standard, and the honest reason is parameter uncertainty plus career risk, not mathematics.

    Where candidates lose it

    Reciting the formula without the asymmetry. The gradeable insight is that the growth curve is flat near the optimum and falls off a cliff past it, which is why uncertainty in your edge estimate pushes you to bet less. Also mention the fifty percent chance of a fifty percent drawdown, because it makes the practical argument concrete.

    Expect next

    • What is the probability of a fifty percent drawdown under full Kelly?
    • How does Kelly relate to mean-variance optimisation?
    • How would you size when your edge estimate itself has a standard error?
  7. 074How would you model market impact and slippage for a strategy you are sizing?Market makingHardtechnicalQuant researchQuant trading

    Say this

    Split the cost into spread, temporary impact and permanent impact. The empirical regularity worth knowing is the square-root law: impact scales roughly with the square root of the order size as a fraction of daily volume, times the volatility.

    Then walk it

    1. The square-root law: impact in volatility units is approximately a constant times the square root of order size over average daily volume, with the constant usually estimated around 0.5 to 1. So trading 1 percent of ADV in a 2 percent daily vol name costs roughly 0.1 times 2 percent, about 20 basis points.
    2. That non-linearity is what caps capacity. Doubling your size only increases impact by 41 percent per share, but total cost grows as size to the power 1.5, so cost eats your edge faster than your edge grows.
    3. Separate temporary from permanent. Temporary impact reverts after you stop trading and is a function of how fast you trade. Permanent impact is the information your trading revealed, and it does not come back. Almgren-Chriss style frameworks trade off the two against the risk of trading slowly.
    4. Estimating it honestly: use your own fills against arrival price, not a vendor model, and regress realised shortfall on participation rate, volatility and spread. You need a lot of trades, and you must control for the fact that you traded more aggressively when you had more signal, which biases the estimate.
    5. And the modelling discipline: be conservative, because impact is the parameter most likely to turn a profitable backtest into a losing strategy. I would rather assume twice the cost and discover I was pessimistic than the reverse. That preference is the answer they want to hear.

    Where candidates lose it

    Assuming linear impact or using the quoted spread as the whole cost. For any size that matters the spread is the small part. Know the square-root law and know that cost scaling as size to the power 1.5 is what determines capacity, because that is the link between a research result and a business decision.

    Expect next

    • Why does cost scale as size to the power one and a half?
    • How do you separate permanent from temporary impact empirically?
    • How does impact determine the capacity of a strategy?
  8. 078Given a stream of numbers, return the median after each element arrives.ProgrammingHardtechnicalOld Mission CapitalEquities · Boston · 2024

    Say this

    Two heaps. A max heap for the lower half and a min heap for the upper half, kept balanced so their sizes differ by at most one. The median is the top of the larger heap, or the average of the two tops. Insert is O(log n), query is O(1).

    Then walk it

    1. Insert rule: if the new value is at most the max of the lower heap, push it there, otherwise push to the upper heap. Then rebalance by moving one element across if the sizes differ by more than one.
    2. Query: if the sizes are equal, the median is the average of the two tops. Otherwise it is the top of the larger heap. Constant time either way.
    3. Total cost for n elements is n log n, and memory is O(n) because you must retain everything. That memory cost is the honest limitation, and it is the first thing an interviewer will probe.
    4. If the median must be over a sliding window rather than the whole prefix, the two-heap approach needs deletions from the middle. Use an indexed multiset or two heaps with lazy deletion and a hash of pending removals. That is the version that comes up in practice on a tick stream.
    5. And if approximate is acceptable, which on a trading system it usually is, the right answer is a streaming quantile sketch: t-digest or the Greenwald-Khanna algorithm, giving you any quantile in bounded memory rather than O(n). Naming that unprompted is what turns a correct interview answer into a practical one.

    Where candidates lose it

    Sorting on every element, which is O(n squared log n) overall, or maintaining a sorted list with insertion, which is O(n) per element because of the shifting even though the binary search is fast. Say two heaps immediately, then volunteer the sliding-window and bounded-memory variants, because that is where the conversation is heading.

    Expect next

    • Now do it over a sliding window of the last thousand values.
    • What if you cannot store all the data?
    • How would you get the 99th percentile instead of the median?

    Reported by candidates at Old Mission Capital (Equities, Boston, 2024). Source: Wall Street Oasis.

  9. 081Write me an unordered_map class. What is actually inside a hash map?ProgrammingHardsuperdayOld Mission CapitalTrading · Chicago · 2021

    Say this

    An array of buckets, a hash function mapping keys to bucket indices, a collision resolution strategy, and a resize policy driven by load factor. The three decisions that define the implementation are the hash, the collision handling, and when you grow.

    Then walk it

    1. Core operations: index equals hash of key modulo bucket count, then search within that bucket comparing keys for equality. Insert, find and erase all follow that pattern, and all are O(1) expected under a good hash.
    2. Collision resolution, and this is the main design choice. Separate chaining stores a list per bucket, which is simple and is what the C++ standard effectively mandates for unordered_map because of its iterator and reference stability guarantees. Open addressing stores entries inline and probes forward, which is far more cache-friendly but complicates erase, since you need tombstones or backward shifting.
    3. Resize: track load factor as elements over buckets, and when it exceeds a threshold, typically 0.75 for chaining or 0.5 to 0.7 for open addressing, allocate a bigger array and rehash everything. Use a power-of-two bucket count so the modulo is a bitmask, but then your hash must mix the high bits or a weak hash collides badly.
    4. The correctness details an interviewer will probe: key equality is separate from the hash, two equal keys must hash the same, iterator invalidation on rehash, and what happens when the key type has a bad hash. A hash that is the identity on integers plus power-of-two buckets means sequential keys with a stride collide catastrophically.
    5. If I were writing this for a trading system I would use open addressing with linear probing over a pre-allocated power-of-two array, reserve capacity up front so no rehash ever happens in the hot path, and store keys and values in separate arrays if the values are large. The reason is tail latency: one rehash mid-session is a millisecond spike, and a millisecond is forever.

    Where candidates lose it

    Describing the interface rather than the internals. The question is about buckets, hashing, collisions and resizing. Also be ready for why is std::unordered_map slow, whose answer is node-per-element allocation and the standard's stability guarantees forcing chaining. And never forget that erase under open addressing needs tombstones, which is the bug candidates ship.

    Expect next

    • How does erase work under open addressing?
    • What load factor would you choose and why?
    • What makes a good hash function, and what happens with a bad one?

    Reported by candidates at Old Mission Capital (Trading, Chicago, 2021). Source: Wall Street Oasis.

  10. 092What are the assumptions behind Black-Scholes, and which one fails hardest?Options and derivativesHardtechnicalDerivativesProp trading firms

    Say this

    Constant known volatility, geometric Brownian motion with no jumps, continuous frictionless hedging, constant rates, no dividends and European exercise. The one that fails hardest is constant volatility, and the proof that it fails is the volatility smile.

    Then walk it

    1. If the model were right, every strike and expiry on the same underlying would have the same implied vol. They do not. Equity index options show a pronounced skew, with out-of-the-money puts trading at much higher implied vol than calls, and the smile steepens for shorter expiries.
    2. Two economic reasons for the skew: returns are negatively skewed with crash risk, which a lognormal cannot represent, and there is genuine demand for downside protection that pushes puts rich. Both are real and they reinforce each other.
    3. No jumps is the second failure, and it is the same failure in a different form. Under continuous paths a delta hedge is riskless in the limit; with jumps it is not, and that unhedgeable jump risk is precisely what the skew prices.
    4. Continuous costless hedging fails too, which matters practically. You hedge discretely and pay the spread, so your realised hedging error has a variance proportional to the hedge interval, and a short-gamma book pays that cost repeatedly.
    5. But here is the thing worth saying: the model is still used everywhere despite being false, because it is a lossless translator between price and implied vol. Traders quote in vol, not in price, and Black-Scholes is the shared language. The correct summary is that it is a wrong model used as a coordinate system, with local vol, stochastic vol models like Heston, and jump models layered on top for anything path-dependent.

    Where candidates lose it

    Listing the assumptions without naming the smile as the empirical refutation. That link is the whole point. And do not conclude the model is useless, because that misses why every desk still quotes in Black-Scholes implied vol. Wrong but indispensable as a change of variables is the answer.

    Expect next

    • If you know the model is wrong, why still use it?
    • What is local volatility, and what does it fix?
    • How would you price a barrier option given a smile?
← PreviousPage 1 of 2
  1. 1
  2. 2
Next →

Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

Puzzles

100 Quant puzzles, solved step by step

Try each one before you read the answer: probability, mental maths and the brainteasers interviewers use to watch you think.

Solve the puzzles →
Case studies

100 Quant case studies, worked step by step

A business, its numbers and a task, as in an assessment day or a case round. Work it on paper, then open the solution one step at a time.

Work the cases →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.