Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Risk Management puzzles, solved step by step

Puzzles
100
Traced to a firm
17
Topics
13
Hard
30
Topic
All topicsCapital and leverage6Compounding and drawdowns8Correlation and diversification8Counterparty exposure and collateral7Credit risk arithmetic10Duration and rates7Liquidity and balance sheet7Logic, estimation and brainteasers7Operational loss and fraud7Options and Greeks7Probability and base rates8Statistics and estimation10VaR and expected shortfall8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 1–5 of 5 · filtered from 100Clear filters
  1. 037Forecaster A has a bias of 1 point and a forecast error standard deviation of 2. Forecaster B is unbiased with a standard deviation of 2.5. Using the mean squared error decomposition, which forecaster is better?Statistics and estimationCoreBLBlackRockNew York · 2026

    Try it first

    Which forecaster has the lower mean squared error?

    Show the worked solution

    Forecaster A, with a mean squared error of 5 against 6.25. Mean squared error splits into bias squared plus variance. A pays 1 squared for its bias and 2 squared for its spread, 5 in all. B pays nothing for bias but 2.5 squared, 6.25, for its spread. A's typical error, the square root, is 2.24 against 2.50.

    How can a biased forecaster beat an unbiased one?

    Two archers. One groups every arrow tightly but slightly left of centre; the other is centred on average but scatters arrows all over the target. Ask which one lands closer to the bullseye on a typical shot, and the tight grouping wins. Mean squared error charges for two things, how far off you are on average and how much you scatter, and a small, steady bias can cost far less than a large scatter. Being unbiased only removes the first charge.

    The relationship
    MSE=bias2+variance:A=12+22=5,B=02+2.52=6.25\text{MSE} = \text{bias}^2 + \text{variance}: \quad A = 1^2 + 2^2 = 5, \qquad B = 0^2 + 2.5^2 = 6.25
    biasthe average forecast error, forecast minus actual
    variancethe spread of the errors around their own average, the standard deviation squared
    What it says in wordsSquared error on average equals the squared average error plus the spread of errors around it.
    Mean squared error = bias squared + variancevariance 4.00bias sq 1MSE 5.00Forecaster Avariance 6.25MSE 6.25Forecaster BTypical error (root MSE)A: 2.24B: 2.50A wins while its biasstays under 1.50
    Forecaster A's mean squared error is 1 of bias squared plus 4 of variance, 5 in total, while unbiased Forecaster B carries 6.25 of pure variance, so A's small bias buys a larger cut in variance and gives the lower error.

    When would your answer flip, and what would you do with A?

    Solve for the tie: A matches B when bias squared plus 4 equals 6.25, so a bias of 1.50. Below that, A wins. More useful still, a bias that is stable can be measured and subtracted: correct A by one point and its MSE falls to 4, better than either original. That is the practical lesson for a risk team: a model that is consistently off in one direction is fixable, while a noisy model is not. The trade-off is also why risk teams use shrinkagePulling a noisy estimate towards a simpler, steadier target, accepting a little bias in return for much lower variance. on covariance matrices built from short histories.

    The limit: MSE punishes large errors heavily because it squares them, and it treats over-forecasts and under-forecasts alike. A risk manager forecasting losses may care more about under-forecasting than over-forecasting, in which case a symmetric score is the wrong yardstick and the ranking could change.

    Where candidates lose it

    Candidates pick B on reflex because unbiased sounds like correct. The question is built to see whether you know that MSE has two parts and can do the two-line arithmetic.

    The quieter miss is stopping at 5 against 6.25. Add that A's bias can be corrected, taking its MSE to 4, and you have turned a statistics answer into a model-risk judgement.

    What the interviewer asks next

    • What bias would make the two forecasters exactly equal?
    • Why might a regulator prefer the unbiased forecaster even with a higher MSE?
    • How would you test whether A's bias is stable over time?

    Asked at BlackRock, Restructuring, New York, 2026 (Wall Street Oasis): Which equities have duration ? multiple stocks vs value stocks MSE Forecasting equation

  2. 048You estimate a desk's daily P&L variance from five observations, once dividing the sum of squared deviations by 5 and once by 4. Which estimator is unbiased, which is consistent, and how large is the bias?Statistics and estimationCoreUBSZurich · 2021

    Try it first

    Which statement is true?

    Show the worked solution

    Dividing by 4 is unbiased; both estimators are consistent; dividing by 5 is low by one fifth. The sample mean is estimated from the same five points, which uses up one degree of freedom, so the divide-by-n estimator averages (n - 1)/n of the true variance, 80% here. If the true variance is 4, it centres on 3.2, a bias of -0.8. As n grows that factor tends to 1, so both estimators converge on the truth.

    Why does dividing by n come out too low?

    Measure how spread out five friends' heights are by comparing each to the group's own average, and you will understate the spread, because that average was pulled towards those five people. Deviations measured from the sample mean are smaller on average than deviations from the true mean, so their sum of squares understates the spread by exactly one observation's worth. Dividing by n minus 1, the degrees of freedomThe number of independent pieces of information left after estimating something from the same data; estimating the mean uses up one., corrects it exactly.

    Five observations: divide by 4 and you are centred; divide by 5 and you land low0481216Estimated daily variance, Rs crore squaredtrue variance 4divide by 5: mean 3.2divide by 4: centred at 4Mean squared errordivide by 4: 8.00divide by 5: 5.76biased, yet less error
    With five observations and a true variance of 4, the divide-by-4 estimator is centred on 4 while the divide-by-5 estimator is centred on 3.2, 80% of the truth, yet the biased version is narrower and has a lower mean squared error, 5.76 against 8.00.

    What is the difference between unbiased and consistent?

    Unbiased is about the average over many repeated samples of the same size; consistent is about what happens to one estimate as the sample grows. Divide-by-5 fails the first: repeat the five-day exercise many times and the estimates average 3.2, not 4. It passes the second: with 250 days the bias is only -0.016, and it shrinks to zero with more data. An estimator can be unbiased but inconsistent too, such as using only the first observation to estimate a mean: right on average, never improving.

    The relationship
    E ⁣[1n∑(xi−xˉ)2]=n−1n σ2=0.8×4=3.2E\!\left[\tfrac{1}{n}\textstyle\sum (x_i - \bar{x})^2\right] = \frac{n-1}{n}\,\sigma^2 = 0.8 \times 4 = 3.2
    nthe number of observations, 5
    \bar{x}the sample mean, estimated from the same five points
    \sigma^2the true variance, 4 in the illustration
    What it says in wordsDividing by n recovers only (n minus 1) over n of the true variance on average.

    Now the twist a model validator should add. For normal data the unbiased estimator has variance 8.00 here, while the divide-by-5 version has 5.12 plus a squared bias of 0.64, a mean squared error of 5.76. The biased estimator is closer to the truth on a typical sample. Which you prefer depends on the use: unbiasedness matters when estimates are averaged across many desks; a smaller typical error matters for a single desk's limit. With five data points, neither is reliable, and that is the more important thing to say.

    Where candidates lose it

    The usual slip is to treat unbiased and consistent as the same thing, and so to call the divide-by-5 estimator inconsistent. The interviewer asked both words together precisely to hear you separate them.

    The second miss is answering from memory without the reason. One sentence on the sample mean using up a degree of freedom shows you know why n minus 1 exists, not just that it does.

    What the interviewer asks next

    • Give an example of an estimator that is unbiased but not consistent.
    • Is the sample standard deviation, the square root of the unbiased variance, itself unbiased?
    • With 250 days of P&L, does the choice between n and n - 1 matter for VaR?

    Asked at UBS, Risk Management, Zurich, 2021 (Wall Street Oasis): And several other questions on econometrics - what is an unbiased estimator vs consistent estimator?

  3. 050An EWMA volatility model with a decay factor of 0.94 had yesterday's daily volatility estimate at 1.2%. Today's return is minus 3%. What is the updated volatility estimate?Statistics and estimationCoreBank market riskQuant risk

    Try it first

    Roughly where does the new daily volatility land?

    Show the worked solution

    About 1.38%. EWMA updates the variance, not the volatility: new variance is 0.94 x 1.2 squared plus 0.06 x 3 squared, which is 1.3536 plus 0.54, or 1.8936. The square root is 1.376%. The shock has only a 6% weight but enters squared, so it supplies 29% of the new variance.

    Why blend variances rather than volatilities?

    A household tracking how much its grocery bill swings would be misled if it averaged the size of the swings in rupees but ignored that one big swing matters far more than several small ones. Variance is the average squared move, so the model updates squared returns, and a move 2.5 times normal size counts 6.25 times as much. An EWMAExponentially weighted moving average: each day the estimate keeps a fixed share of yesterday and adds the rest from today, so older days fade geometrically. of variances is what makes one large shock move the estimate quickly. Blending the volatilities directly gives 1.31%, understating the jump.

    Six per cent of the weight, but more than a quarter of the resultWeights94% on yesterday's variance6%Share ofnew variance0.94 x 1.2 sq = 1.35471.5%0.06 x 3 sq = 0.5428.5%New variance 1.3536 + 0.54 = 1.8936Volatility: square root = 1.38% a day, up from 1.20%
    Today's minus 3% return carries only a 6% weight in the EWMA update, but because it enters squared it supplies 28.5% of the new variance of 1.8936, lifting daily volatility from 1.20% to 1.38%.
    The relationship
    σt2=λ σt−12+(1−λ) rt2=0.94(1.44)+0.06(9)=1.8936,σt=1.376%\sigma_t^2 = \lambda\,\sigma_{t-1}^2 + (1-\lambda)\,r_t^2 = 0.94(1.44) + 0.06(9) = 1.8936, \quad \sigma_t = 1.376\%
    \lambdathe decay factor, 0.94
    \sigma_{t-1}yesterday's volatility estimate, 1.2%
    r_ttoday's return, minus 3%
    What it says in wordsKeep 94% of yesterday's variance, add 6% of today's squared return, then take the square root.

    What happens next, and what are the model's limits?

    If tomorrow is flat, the estimate decays to the square root of 0.94 x 1.8936, about 1.33%. Each day of calm keeps 94% of the variance, so a shock's influence halves in about 11 trading days: ln 0.5 over ln 0.94. That is the design choice behind 0.94: fast enough to react to a new regime within days, slow enough not to swing on every move.

    The limits: EWMA has no pull towards a long-run average, so after a calm spell it can drift very low and understate risk just before volatility returns, which GARCH-type models address with a mean-reversion term. It also treats up and down moves alike, while falling markets often raise volatility more than rising ones. And the choice of 0.94 is a convention for daily data, not a law, so the estimate should be backtested against realised moves.

    Where candidates lose it

    The usual slip is to blend the two volatilities directly, 0.94 x 1.2 plus 0.06 x 3, and answer 1.31%. It looks like the formula but applies it to the wrong quantity and understates the effect of the shock.

    The second is forgetting the square root at the end and quoting 1.89% as a volatility. Say the units at each step: variance in, variance out, then volatility.

    What the interviewer asks next

    • What would the estimate be with lambda of 0.97 instead?
    • How many quiet days until the estimate is back below 1.25%?
    • Why might a risk manager prefer GARCH to EWMA for a ten-day VaR?
  4. 062Your prior estimate of a hidden fair price is 10 with variance 4. A noisy measurement comes in at 12 with measurement variance 1. After one Kalman filter update, what is your new estimate and its variance?Statistics and estimationCoreUBSLondon · 2022

    Try it first

    Where does the new estimate land?

    Show the worked solution

    The new estimate is 11.6 with variance 0.8. The Kalman gain is the prior variance over the total, 4 over 5, or 0.8. The estimate moves 80% of the way from 10 towards 12, landing at 11.6. The variance becomes (1 - 0.8) x 4 = 0.8, smaller than either the prior's 4 or the measurement's 1.

    How does the filter decide how far to move?

    Two friends guess your commute time. One has ridden with you a hundred times, the other once. You would average their guesses, but lean heavily on the first. A Kalman filterA method that updates an estimate of something you cannot see directly each time a noisy measurement arrives, weighting old estimate and new data by how precise each is. does exactly that: it weights the prior and the measurement by their precision, one over variance. Precision 0.25 against 1 gives the measurement 80% of the weight.

    The update leans towards whichever source is more precise, and ends up sharper than both46810121416Estimate of the hidden fair priceprior: 10, variance 4measurement: 12, variance 1update: 11.6, variance 0.8gain = 4 / (4 + 1) = 0.810 + 0.8 x (12 - 10) = 11.6
    The prior centred at 10 is wide, with variance 4, and the measurement at 12 is narrow, with variance 1. The update lands at 11.6, four fifths of the way to the measurement, and its variance of 0.8 makes it narrower than either source.

    Why is the new variance smaller than both inputs?

    Because two independent pieces of evidence together know more than either alone. Precisions add: 1 over 4 plus 1 over 1 is 1.25, and one over 1.25 is a variance of 0.8. That is the part candidates skip. The filter does not just move the estimate; it becomes more confident with every measurement, until new data carry little weight and the estimate settles.

    The relationship
    K=PP+R=0.8x^=10+0.8(12−10)=11.6P′=(1−K)P=0.8K = \frac{P}{P+R} = 0.8 \qquad \hat{x} = 10 + 0.8(12-10) = 11.6 \qquad P' = (1-K)P = 0.8
    Pthe prior variance, 4
    Rthe measurement variance, 1
    Kthe Kalman gain, the weight on the new measurement
    \hat{x}the updated estimate
    What it says in wordsMove from the prior towards the measurement by the gain, and shrink the variance by the same share.

    Say the limitation: this single step assumes both errors are normal and the hidden price did not move between the prior and the measurement. A full filter adds a prediction step that lets the price drift and widens the variance before each update. Risk teams use the idea to track hidden quantities such as a hedge ratio that changes over time.

    Where candidates lose it

    The trap is averaging the two numbers and answering 11. That treats the measurement and the prior as equally trustworthy, which the variances say they are not.

    The second trap is getting 11.6 and then saying the variance is somewhere between 1 and 4. Combining evidence always reduces uncertainty, so the new variance must be below both: 0.8.

    What the interviewer asks next

    • A second measurement of 11 arrives with variance 1. What is the estimate now?
    • What happens to the gain as the number of measurements grows?
    • How would you use a Kalman filter to estimate a time-varying hedge ratio?

    Asked at UBS, Risk, London, 2022 (Wall Street Oasis): Explain what kalman filter is.

  5. 087A desk's daily returns have a volatility of 1%. Over 250 days its average daily return is 0.05%. How precisely is that mean estimated, and can you say the desk has skill?Statistics and estimationCoreAsset manager riskQuant risk

    Try it first

    Is 0.05% a day, measured over one year, clearly different from zero?

    Show the worked solution

    Not precisely enough to claim skill. The standard error of the mean is the daily volatility over the square root of the number of days: 1% over the square root of 250, about 0.063%. The estimate of 0.05% is only 0.79 standard errors from zero, and a two standard error band runs from -0.076% to 0.176%. You would need about 1,600 days, 6.4 years, to clear zero.

    Why is a year of daily data not enough?

    Weigh yourself on a bathroom scale that jumps by two kilos each time you step on it. If you lost 100 grams last month, a week of readings will not show it; the jumps drown the signal. The precision of an average improves only with the square root of the number of observations, so a small edge buried in large daily noise takes years to show. Here the daily noise is 1% and the edge is 0.05%, twenty times smaller.

    The relationship
    SE=σn=1%250=0.063%t=0.050.063=0.79SE = \frac{\sigma}{\sqrt{n}} = \frac{1\%}{\sqrt{250}} = 0.063\% \qquad t = \frac{0.05}{0.063} = 0.79
    sigmadaily return volatility, 1%
    nnumber of daily observations, 250
    thow many standard errors the mean is from zero
    What it says in wordsThe mean's uncertainty is the daily volatility divided by the square root of the number of days, and the edge is less than one of those units from zero.
    One year of data cannot tell a 0.05% edge from zerozero: no skill250 daysone year-0.076%+0.176%1,600 daysabout 6.4 years0.000%+0.100%estimate 0.05%-0.10%0+0.10%+0.20%Average daily return, with a two standard error band
    After 250 days the two standard error band around the 0.05% daily mean runs from -0.076% to 0.176% and straddles zero; after about 1,600 days, 6.4 years, the band narrows to 0.000% to 0.100% and only just clears zero.

    How long would it take, and what does that mean for judging desks?

    Set the t-statistic to 2 and solve for n: n equals (2 times 1% over 0.05%) squared, which is 1,600 days. An annual Sharpe ratio of about 0.79 needs more than six years of data before it is statistically distinguishable from zero. That is longer than most desks keep the same strategy, so a risk team cannot rely on the P&L record alone; it looks at whether the edge has a reason, whether it survives out of sample, and how much the result depends on a few days.

    State the assumptions. The calculation treats daily returns as independent with constant volatility. Fat tails and volatility clustering make the true uncertainty larger, so six years is a floor, not a promise.

    Where candidates lose it

    The common error is annualising the mean to 12.5% and declaring skill, as if a big annual number were proof. The annual volatility grows too, to about 16%, and the ratio of the two is what matters.

    The other slip is dividing by 250 instead of its square root, which makes the mean look fifty times more precise than it is.

    What the interviewer asks next

    • The desk's volatility is 0.5% instead. How many days now?
    • Why is the mean so much harder to estimate than the volatility?
    • How would you judge a new desk that has only six months of history?
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.