Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Risk Management puzzles, solved step by step

Puzzles
100
Traced to a firm
17
Topics
13
Hard
30
Topic
All topicsCapital and leverage6Compounding and drawdowns8Correlation and diversification8Counterparty exposure and collateral7Credit risk arithmetic10Duration and rates7Liquidity and balance sheet7Logic, estimation and brainteasers7Operational loss and fraud7Options and Greeks7Probability and base rates8Statistics and estimation10VaR and expected shortfall8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 1–10 of 17 · filtered from 100Clear filters
  1. 006Write NPV as the product of two vectors. The cash flows are minus 100, 30, 40, 50 and 20 in years 0 to 4, and the discount rate is 10%. What is the NPV?Duration and ratesWarm upMoody'sNew York · 2018

    Try it first

    Which operation turns the two vectors into the NPV?

    Show the worked solution

    NPV is the dot product of the cash flow vector and the discount factor vector, and here it is about 11.56. The discount factors at 10% are 1, 0.909, 0.826, 0.751 and 0.683. Multiplying term by term gives minus 100, 27.27, 33.06, 37.57 and 13.66, which sum to 11.56. A positive NPV means the project earns more than 10%.

    Why is NPV a dot product at all?

    Think of a grocery bill. One list holds the quantity of each item, another holds each item's price, and the bill is quantity times price for each line, added up. NPV has the same shape: one vector holds the cash flows, the other holds what one rupee in each year is worth today, and the NPV is the sum of their products. Writing it this way separates the project, which is the cash flow vector, from the market, which is the discount factor vector. Change the rate and only the second vector changes.

    The relationship
    NPV=c⋅d=∑t=04ct dt,dt=1(1+r)t\text{NPV} = \mathbf{c} \cdot \mathbf{d} = \sum_{t=0}^{4} c_t\, d_t, \qquad d_t = \frac{1}{(1+r)^t}
    cthe cash flow vector, year 0 to year 4
    dthe discount factor vector, one entry per year
    rthe discount rate, 10%
    What it says in wordsMultiply each cash flow by the value today of one rupee in that year, and add the results.
    Multiply the two vectors term by term, then add the productsYear 0Year 1Year 2Year 3Year 4Cash flow-10030405020xDiscount factor1.0000.9090.8260.7510.683=Present value-100.0027.2733.0637.5713.660-100+27.27+33.06+37.57+13.66+11.56Year 0Year 1Year 2Year 3Year 4NPVRunning sum
    Multiplying the cash flows by the discount factors year by year gives present values of minus 100, 27.27, 33.06, 37.57 and 13.66, and adding them from minus 100 upward reaches an NPV of 11.56.

    What does the interviewer want to hear beyond the number?

    The thought process was part of the question, so say it in order. First build the discount factor vector from the rate, then take the dot product, then sanity check the sign and size. The undiscounted inflows are 140 against 100 out, so a positive but much smaller NPV is expected once four years of 10% are taken out. And in a spreadsheet the same idea is one SUMPRODUCT of two ranges, which is why the vector form is how a model is usually built.

    Then give the extension that shows range. With a term structure of rates, only the discount factor vector changes: each entry uses its own year's rate. With several scenarios, stack the cash flow vectors into a matrix and one matrix multiplication gives every scenario's NPV at once. The limitation is that the vector form assumes the cash flows are known; uncertain cash flows need expected values or scenarios first.

    Where candidates lose it

    Candidates reach for the NPV formula and start adding fractions, which gets the number but misses the question. The interviewer asked for two vectors precisely to see whether you can separate what the project pays from what time is worth.

    The other slip is discounting year 0. The first discount factor is 1; the minus 100 is already in today's money.

    What the interviewer asks next

    • How would you write the IRR condition using the same two vectors?
    • The rate for year 1 is 8% and for later years 10%. What changes in the vector form?
    • How would you compute the NPV for 1,000 cash flow scenarios in one operation?

    Asked at Moody's, Analytics, New York, 2018 (Wall Street Oasis): Construct an NPV formula using 2 vectors and show me your thought process.

  2. 012A Monte Carlo VaR at 99% uses 10,000 simulated paths. How many scenarios sit in the tail that defines the VaR, and how many paths do you need to halve the standard error of the estimate?Statistics and estimationHardMSCISan Francisco · 2018

    Try it first

    How many paths halve the standard error of the VaR estimate?

    Show the worked solution

    100 scenarios define the tail, and you need 40,000 paths to halve the error. At 99%, 1% of 10,000 paths is 100 scenarios, and the VaR is read at the edge of that group. Monte Carlo error falls as one over the square root of the number of paths, so halving it means four times the paths. For a book with Rs 10 crore of daily volatility, the error is about Rs 37 lakh at 10,000 paths.

    Why does a 99% VaR rest on so few scenarios?

    Picture an exit poll that interviews 10,000 voters but reports only on the one voter in a hundred who picked a small party. Your estimate for that party rests on about 100 people, not 10,000. A 99% VaR is read at the boundary of the worst 1% of outcomes, so only about 100 of 10,000 paths carry information about where that boundary is. For a normal book with Rs 10 crore of daily volatility the true VaR is Rs 23.26 crore, and the estimate wobbles around it by about Rs 37 lakh from one run of 10,000 paths to the next.

    The relationship
    SE(VaR^)≈1f(q)p(1−p)n\text{SE}(\widehat{\text{VaR}}) \approx \frac{1}{f(q)}\sqrt{\frac{p(1-p)}{n}}
    pthe tail probability, 1%
    nthe number of paths
    f(q)the height of the P&L density at the VaR point
    What it says in wordsThe error of a quantile estimate shrinks with the square root of the paths and grows where the tail is thin.
    Precision costs the square: half the error needs four times the paths02040608010k20k30k40k50kMonte Carlo pathsStandard error, Rs lakh10,000 paths: Rs 37.3 lakh40,000: Rs 18.7 lakhTail at 10,000 paths100 scenarios1% of the pathsVaR estimateRs 23.26 crore+/- 1.6% at 10k
    For a book with Rs 10 crore of daily volatility, the standard error of a 99% Monte Carlo VaR is about Rs 37.3 lakh at 10,000 paths and Rs 18.7 lakh at 40,000, because the error falls with the square root of the number of paths.

    Why does halving the error cost four times the paths?

    Because the paths sit under a square root. To halve the error, the square root of the path count has to double, which means the path count itself must quadruple. That is why precision in Monte Carlo is expensive: going from about 1.6% error to 0.8% of the VaR costs four times the computing time. It also explains why desks use variance reductionSimulation techniques, such as antithetic paths or importance sampling, that give a more precise estimate from the same number of paths., especially importance sampling, which pushes more paths into the tail where the VaR is decided.

    Say the limit plainly. More paths shrink sampling error only; they do nothing about model error. If the simulated distribution has the wrong tail, a million paths give a very precise estimate of the wrong number. And expected shortfall at 97.5% averages the worst 250 paths, so it is usually more stable than the 99% VaR read from a single boundary.

    Where candidates lose it

    The trap is answering 20,000 paths, assuming error falls in proportion to the path count. It falls with the square root, and the difference is a factor of two in cost that a quant interviewer will not let pass.

    The other miss is thinking all 10,000 paths inform the VaR. Say the number 100 out loud; it shows you know why tail estimates are noisy.

    What the interviewer asks next

    • How many paths would you need for the same precision on a 99.9% VaR?
    • How does importance sampling reduce the error without adding paths?
    • Why is a Monte Carlo VaR often less stable from day to day than a historical VaR on 500 days of data?

    Asked at MSCI, Risk Management, San Francisco, 2018 (Wall Street Oasis): the next round was paper test, materials were similar to CFA level 1, and it mostly focused on monte carlo, dividend, and risk

  3. 013For a normally distributed P&L, compare the 99% VaR with the 97.5% expected shortfall. Why would switching from one to the other barely change capital for a normal book but raise it for a fat-tailed one?VaR and expected shortfallHardUBSZurich · 2021

    Try it first

    For a normal P&L, how do the 99% VaR and the 97.5% expected shortfall compare?

    Show the worked solution

    For a normal P&L they are almost identical: 2.33 against 2.34 standard deviations. On Rs 10 crore of daily volatility that is Rs 23.3 crore against Rs 23.4 crore. The 97.5% level was chosen so the switch would be neutral for a normal book. For a fat-tailed book with the same volatility, VaR is Rs 26.2 crore and expected shortfall Rs 29.1 crore, about 11% higher.

    What does each number actually measure?

    Think of a river's flood level. VaR is the line on the wall that the water passes one year in a hundred. Expected shortfallThe average loss on the days when losses exceed the VaR at a chosen confidence level. is how deep the water gets, on average, in the years it passes a lower line. VaR reads a single point in the tail; expected shortfall averages everything beyond its cut-off, so it responds to how far the tail stretches. For a normal distribution the 97.5% cut-off is 1.96 standard deviations, and the average of the losses beyond it is 2.338 standard deviations.

    The relationship
    ES97.5%=σ φ(1.96)0.025=2.338 σVaR99%=2.326 σ\text{ES}_{97.5\%} = \sigma\,\frac{\varphi(1.96)}{0.025} = 2.338\,\sigma \qquad \text{VaR}_{99\%} = 2.326\,\sigma
    sigmathe standard deviation of daily P&L
    phi(1.96)the height of the standard normal curve at the 97.5% cut-off
    0.025the probability of being beyond that cut-off
    What it says in wordsFor a normal book, the average loss beyond 1.96 standard deviations lands almost exactly on the 99% VaR.
    Same volatility, two tails: where VaR and expected shortfall landNormal tailVaR 99%: 23.3ES 97.5%: 23.412223242Daily loss, Rs croreFat tail (Student t, 3 degrees of freedom)VaR 99%: 26.2ES 97.5%: 29.112223242Daily loss, Rs croreGap between ES and VaR: 0.5%Gap between ES and VaR: 11.0%
    With the same Rs 10 crore volatility, a normal P&L puts the 99% VaR at Rs 23.3 crore and the 97.5% expected shortfall at Rs 23.4 crore, almost the same point, while a fat-tailed P&L puts them at Rs 26.2 crore and Rs 29.1 crore, 11% apart.

    Why does the fat-tailed book pay more under expected shortfall?

    Because its extreme losses are larger even though its everyday volatility is the same. Expected shortfall averages the tail, so a book that sells protection against crashes, whose losses are rare but very large, shows a much higher number under ES than under VaR. Using a Student t with three degrees of freedom and the same Rs 10 crore volatility, the 99% VaR rises only to Rs 26.2 crore, but the 97.5% expected shortfall reaches Rs 29.1 crore. That gap is exactly the risk VaR was criticised for ignoring.

    Give both sides, because the question asks for advantages and disadvantages. Expected shortfall sees the tail's depth and adds up sensibly across desks, since it is subadditive. But it is harder to backtest, because you are checking an average of rare events rather than a count of breaches, and it needs more data to estimate. VaR is easy to backtest and explain, and blind beyond its own line.

    Where candidates lose it

    The trap is assuming that a 97.5% measure must be smaller than a 99% measure. It compares the confidence levels and forgets that expected shortfall averages beyond its line while VaR stops at it.

    The second miss is stating that ES is always much larger than VaR. For a normal book it is not; say the 2.33 and 2.34 and explain that the gap only opens when the tail is fat.

    What the interviewer asks next

    • Why is VaR not subadditive, and can you build an example with two bonds?
    • How would you backtest an expected shortfall model?
    • A desk sells deep out-of-the-money puts. Which measure shows its risk better, and why?

    Asked at UBS, Risk Management, Zurich, 2021 (Wall Street Oasis): what are the advantages and disadvantages of ES compared to VaR?

  4. 017A Rs 1,000 crore loan pool is tranched into equity from 0 to 5%, mezzanine from 5 to 15% and senior from 15 to 100%. The pool loses 12%. How much does each tranche lose as a share of its size, and what pool loss wipes out the mezzanine?Credit risk arithmeticCoreMoody'sNew York · 2024

    Try it first

    What share of the mezzanine tranche is lost when the pool loses 12%?

    Show the worked solution

    Equity loses 100%, mezzanine 70% and senior nothing; the mezzanine is wiped out at a 15% pool loss. The Rs 120 crore loss fills the tranches from the bottom. Equity absorbs its full Rs 50 crore. The remaining Rs 70 crore falls on the Rs 100 crore mezzanine. The senior tranche starts losing only once pool losses pass 15%, the point where the mezzanine is gone.

    How do losses move through a tranche stack?

    Picture a building flooding from the ground up. The ground floor is soaked before a drop reaches the first floor, and the top floors stay dry until the water climbs to them. Losses fill the tranches from the bottom: each tranche loses nothing until the pool loss passes its attachment pointThe level of pool loss at which a tranche starts to lose money., and everything once the loss passes its detachment point. The equity attaches at 0% and detaches at 5%; the mezzanine attaches at 5% and detaches at 15%.

    Losses fill the stack from the bottom: a 12% pool loss//Equity 0-5%Mezzanine 5-15%Senior 15-100%Pool loss 12%at 15%: mezzanine gone0%5%10%15%20%100%TrancheSize, Rs croreLoss, Rs croreShare of tranche lostEquity5050100%Mezzanine1007070%Senior85000%
    A 12% loss on the Rs 1,000 crore pool wipes out the Rs 50 crore equity tranche, takes Rs 70 crore, or 70%, of the Rs 100 crore mezzanine, and leaves the senior tranche untouched until pool losses pass 15%.
    The relationship
    tranche loss share=min⁡(L,D)−min⁡(L,A)D−A=12%−5%15%−5%=70%\text{tranche loss share} = \frac{\min(L, D) - \min(L, A)}{D - A} = \frac{12\% - 5\%}{15\% - 5\%} = 70\%
    Lthe pool loss, 12%
    Athe attachment point, 5% for the mezzanine
    Dthe detachment point, 15% for the mezzanine
    What it says in wordsThe part of the pool loss that falls between a tranche's lower and upper edges, divided by the tranche's thickness.

    Why does thickness decide how risky a tranche is?

    Because a thin tranche goes from untouched to wiped out over a small range of pool losses. The mezzanine is only 10 points thick, so a pool loss moving from 5% to 15% takes it from zero to total loss, while the same move barely registers on the pool as a whole. That is the leverage inside structured finance: the mezzanine's loss share moved 7 times as far as the pool's 12% average suggests from 5% onwards. A rating analyst evaluating the deal asks how likely the pool loss is to cross each attachment point, which depends heavily on how correlated the loans are.

    Name the risks the structure does not remove. Correlation among the loans decides whether pool losses cluster at a few percent or occasionally jump past 15%. The collateral data may be weak. And the waterfall rules in the documents, such as when cash is diverted to protect senior holders, can shift losses between tranches in ways this simple loss-only picture does not show.

    Where candidates lose it

    The trap is answering 12% for every tranche, as if losses were shared in proportion. The whole point of tranching is that they are not.

    The second miss is saying the mezzanine loses 7%, the points above its attachment, and forgetting to divide by its 10 point thickness. Loss share is always relative to the tranche's own size.

    What the interviewer asks next

    • What pool loss would cost the senior tranche 10% of its value?
    • How does rising correlation among the loans change the risk of the equity versus the senior tranche?
    • Why might a mezzanine tranche be rated well below the pool's average credit quality?

    Asked at Moody's, Credit Risk, New York, 2024 (Wall Street Oasis): What is structured finance, how would you evaluate it, and what are the credit risks?

  5. 022An institutional investor asks for an 8% expected annual return with 10% volatility. Assuming returns are normal, what is the chance of a losing year, and what volatility would keep that chance below 10%?Probability and base ratesCoreMSCIAnonymous interview candidate in · 2013

    Try it first

    Roughly how often does this portfolio lose money in a year?

    Show the worked solution

    About 21%, and volatility would need to fall to about 6.2%. A loss means a return below zero, which is 8 points, or 0.8 standard deviations, under the mean. About 21.2% of a normal distribution lies below that, roughly one year in five. For a 10% chance, zero must sit 1.28 standard deviations below the mean, so volatility must be 8 divided by 1.28, about 6.2%.

    How do a return target and a volatility target fix the chance of loss?

    Think of a commute that takes 40 minutes on average but varies from day to day. Whether you are ever late for a 50 minute deadline depends on how much it varies, not only on the average. The chance of a losing year depends on how many standard deviations the expected return sits above zero: here 8 divided by 10, which is 0.8. Look up 0.8 in the normal table and about 21.2% of years fall below zero. The investor who hears 8% and thinks losses are rare is wrong one year in five.

    Same 8% target, two volatilities: the area below zero is the chance of a losing year-20%-10%0%8%20%30%Annual returnloss | gaintarget 8%Volatility 10%P(loss) = 21.2%Volatility 6.2%P(loss) = 10.0%Loss chance = N(-mean / vol)
    With an 8% expected return and 10% volatility, 21.2% of the return distribution falls below zero, while cutting volatility to 6.2% narrows the curve until exactly 10% of years show a loss.
    The relationship
    P(R<0)=N ⁣(−μσ)=N(−0.8)≈21.2%σmax⁡=μ1.2816≈6.2%P(R < 0) = N\!\left(-\frac{\mu}{\sigma}\right) = N(-0.8) \approx 21.2\% \qquad \sigma_{\max} = \frac{\mu}{1.2816} \approx 6.2\%
    muthe expected annual return, 8%
    sigmathe annual volatility
    Nthe standard normal cumulative distribution
    1.2816the number of standard deviations that leaves 10% in the lower tail
    What it says in wordsDivide the expected return by the volatility, and the normal table tells you how often returns fall below zero.

    What would you actually set as targets, and what is wrong with this model?

    Set the targets as a pair, and state the trade-off. If the investor cannot tolerate losing more than one year in ten, then either volatility must come down to about 6.2%, which usually lowers the expected return too, or the loss tolerance must be stated over a longer horizon. Over five years the mean grows five times but the volatility only by the square root of five, so the chance of a losing five-year stretch is much lower. Asking about the horizon is the question a good risk manager raises first.

    Then name the model's limits. Real returns have fatter left tails than a normal curve, so the chance of a large loss is understated; returns are not independent from year to year; and the 8% expected return is an assumption, not a promise. A drawdown limit, such as no more than a 15% fall from peak, is often more useful to an institution than a probability of a losing year.

    Where candidates lose it

    The trap is assuming that a positive expected return makes losing years rare. At 0.8 standard deviations above zero, they happen about one year in five.

    The second miss is solving for volatility with the wrong number from the normal table. For a 10% tail you need 1.28 standard deviations, not 1.645, which is the 5% tail.

    What the interviewer asks next

    • What is the chance of a negative return over five years with the same targets, assuming independent years?
    • The investor adds a limit of no more than a 15% loss in any year. What volatility does that imply at 99% confidence?
    • Why might a pension fund care more about a drawdown limit than a volatility target?

    Asked at MSCI, Risk Management, Anonymous interview candidate in, 2013 (Wall Street Oasis): What risk-return targets would you set for an institutional investor?

  6. 025A Kalman filter tracks a random walk with process variance 1 and measurement variance 4. What gain does it settle at, and which simple smoother is it then equivalent to?Statistics and estimationHardUBSLondon · 2022

    Try it first

    Where does the gain settle?

    Show the worked solution

    The gain settles at about 0.39, and the filter becomes an exponentially weighted moving average. In steady state the prior variance P solves P squared minus P minus 4 equals zero, so P is about 2.56. The gain is P over P plus 4, about 0.39. Each new estimate is then 0.39 times the new reading plus 0.61 times the old estimate, which is exactly an EWMA.

    What is a Kalman filter doing, in one picture?

    Think of estimating how many people are in a stadium from a noisy turnstile count that you update every few minutes. Your last estimate is useful but the crowd keeps changing, and each new count is useful but noisy. A Kalman filter blends the old estimate and the new reading, weighting each by how much you trust it; the weight on the new reading is the Kalman gainThe share of the gap between a new measurement and the prior estimate that the filter accepts as news.. Process variance of 1 says the true value drifts by about one unit each step; measurement variance of 4 says each reading is off by about two units.

    The relationship
    P−=P+Q,K=P−P−+R,P=(1−K)P−  ⇒  (P−)2−QP−−QR=0P^- = P + Q, \quad K = \frac{P^-}{P^- + R}, \quad P = (1-K)P^- \;\Rightarrow\; (P^-)^2 - Q P^- - QR = 0
    Qthe process variance, 1: how much the true value moves each step
    Rthe measurement variance, 4: how noisy each reading is
    P^-the variance of the estimate just before a reading arrives
    Kthe Kalman gain
    What it says in wordsIn steady state the uncertainty added by the drift each step exactly balances the uncertainty removed by each reading.
    The gain settles fast, and then the filter is just a moving average0.000.250.500.751.0012345678910Update numberKalman gain Ksteady state 0.3900.960.550.44Steady statenew = 0.39 x reading+ 0.61 x oldan EWMAProcess variance Q = 1Measurement R = 4
    Starting from a vague prior the Kalman gain begins at 0.96, drops to 0.55 and 0.44, and settles at 0.390 by about the sixth update, after which the filter is an exponentially weighted moving average with weight 0.39 on each new reading.

    Why does the gain settle rather than fall to zero?

    Because the thing being tracked keeps moving. If the true value were fixed, every reading would add certainty and the gain would shrink towards zero, like a running average; with a random walk, each step adds one unit of variance back, so certainty stops improving at a balance point. Plugging in Q of 1 and R of 4, the prior variance solves P squared minus P minus 4, giving (1 plus the square root of 17) over 2, about 2.56, and a gain of 0.390. Only the ratio of R to Q matters: a noisier measurement lowers the gain, a faster-moving state raises it.

    That is the link worth saying in a risk interview. An EWMA volatility or correlation estimate, the kind many risk systems use, is a Kalman filter in steady state with a particular noise ratio, whether or not anyone calls it that. The filter's advantage is that it chooses the weight from stated assumptions about noise and drift, and adapts it early on; the limit is that those assumptions, a linear model with normal noise, have to be right for the weight to be the best one.

    Where candidates lose it

    The trap is describing a Kalman filter in general terms and never producing a number. The interviewer wants to see you set up the variance recursion and solve the steady state.

    The second miss is saying the gain goes to zero. That is true only for a constant state; for a random walk it settles at a positive value, and that is why the filter reduces to an EWMA.

    What the interviewer asks next

    • What steady-state gain do you get with measurement variance 1 instead of 4?
    • What EWMA decay factor corresponds to this filter, and what is its half-life in updates?
    • How would you estimate Q and R from data?

    Asked at UBS, Risk, London, 2022 (Wall Street Oasis): Explain what a Kalman filter is

  7. 028A fund's annual volatility is 18%, its benchmark's is 16%, and the correlation between their returns is 0.95. What is the fund's tracking error?Correlation and diversificationCoreMSCIMonterrey · 2013

    Try it first

    Quick instinct: roughly how big is the tracking error?

    Show the worked solution

    About 5.73%. Tracking error is the volatility of the fund's return minus the benchmark's. Its variance is 18 squared plus 16 squared minus 2 x 0.95 x 18 x 16, which is 324 plus 256 minus 547.2, or 32.8. The square root is 5.73%, nearly three times the 2-point gap in volatilities.

    What exactly is tracking error measuring?

    Two friends walk to the same office. How far apart they are at any moment depends less on how fast each walks than on whether they take the same streets. Tracking error is the volatility of the return difference, fund minus benchmark, so it depends on how much the two move apart, not on how much each moves. The fund and index can both swing wildly and still track closely if they swing together. That is why the formula needs the correlationA number from minus 1 to 1 describing how closely two returns move together; 1 means perfect lockstep., not just the two volatilities.

    The relationship
    TE=σF2+σB2−2ρ σFσB=324+256−547.2=32.8=5.73%TE = \sqrt{\sigma_F^2 + \sigma_B^2 - 2\rho\,\sigma_F\sigma_B} = \sqrt{324 + 256 - 547.2} = \sqrt{32.8} = 5.73\%
    \sigma_F, \sigma_Bfund and benchmark volatility, 18% and 16%
    \rhocorrelation of their returns, 0.95
    What it says in wordsThe variance of a difference is the two variances added, less twice the part they share.
    Tracking error is the short side of the volatility triangleangle 18.2 degrees, cos = 0.95Benchmark volatility 16%Fund volatility 18%TE 5.73%Same 18 and 16, change onlythe correlationcorrelation 0.993.12%correlation 0.955.73%correlation 0.907.85%Vols differ by only 2 pointsVol gap alone: 18 - 16 = 2Correlation adds the rest
    Drawing the two volatilities as sides 18 and 16 at the angle whose cosine is 0.95 makes tracking error the short third side, 5.73%; nudging correlation from 0.95 to 0.99 cuts it to 3.12%, and dropping it to 0.90 raises it to 7.85%.

    Why does the correlation matter more than the volatilities?

    Look at the table in the figure. Keeping 18 and 16 fixed, moving correlation from 0.99 to 0.90 takes tracking error from 3.12% to 7.85%, more than doubling it. At high correlations each hundredth of correlation moves tracking error a lot, because the large shared term 2 x rho x 18 x 16 almost cancels the two variances. Now hold correlation at 0.95 and give both sides 16% volatility: tracking error is still 5.06%. The volatility gap contributes a little; the imperfect correlation contributes most.

    Close with the limit. The formula uses a correlation estimated from history, and correlations drift, often falling in stressed markets. A fund reporting 5.7% tracking error in calm years can run well above it in a sell-off, so a risk team watches realised tracking error alongside the model figure.

    Where candidates lose it

    The fast wrong answer is 2%, subtracting the volatilities. It silently assumes correlation of exactly 1, which the question has just told you is false. Candidates who say it have treated volatility as if it were a return.

    The second trap is fumbling the formula under pressure. Anchor it to one line you already know: the variance of A minus B is var A plus var B minus twice the covariance. Everything else follows.

    What the interviewer asks next

    • What correlation would give a tracking error of exactly 2%?
    • The fund's beta to the benchmark is 1.07. Split the tracking error into a beta part and a residual part.
    • Why might a fund with low tracking error still underperform its benchmark every year?

    Asked at MSCI, Financial Tools, Monterrey, 2013 (Wall Street Oasis): What's the tracking error formula?

  8. 037Forecaster A has a bias of 1 point and a forecast error standard deviation of 2. Forecaster B is unbiased with a standard deviation of 2.5. Using the mean squared error decomposition, which forecaster is better?Statistics and estimationCoreBLBlackRockNew York · 2026

    Try it first

    Which forecaster has the lower mean squared error?

    Show the worked solution

    Forecaster A, with a mean squared error of 5 against 6.25. Mean squared error splits into bias squared plus variance. A pays 1 squared for its bias and 2 squared for its spread, 5 in all. B pays nothing for bias but 2.5 squared, 6.25, for its spread. A's typical error, the square root, is 2.24 against 2.50.

    How can a biased forecaster beat an unbiased one?

    Two archers. One groups every arrow tightly but slightly left of centre; the other is centred on average but scatters arrows all over the target. Ask which one lands closer to the bullseye on a typical shot, and the tight grouping wins. Mean squared error charges for two things, how far off you are on average and how much you scatter, and a small, steady bias can cost far less than a large scatter. Being unbiased only removes the first charge.

    The relationship
    MSE=bias2+variance:A=12+22=5,B=02+2.52=6.25\text{MSE} = \text{bias}^2 + \text{variance}: \quad A = 1^2 + 2^2 = 5, \qquad B = 0^2 + 2.5^2 = 6.25
    biasthe average forecast error, forecast minus actual
    variancethe spread of the errors around their own average, the standard deviation squared
    What it says in wordsSquared error on average equals the squared average error plus the spread of errors around it.
    Mean squared error = bias squared + variancevariance 4.00bias sq 1MSE 5.00Forecaster Avariance 6.25MSE 6.25Forecaster BTypical error (root MSE)A: 2.24B: 2.50A wins while its biasstays under 1.50
    Forecaster A's mean squared error is 1 of bias squared plus 4 of variance, 5 in total, while unbiased Forecaster B carries 6.25 of pure variance, so A's small bias buys a larger cut in variance and gives the lower error.

    When would your answer flip, and what would you do with A?

    Solve for the tie: A matches B when bias squared plus 4 equals 6.25, so a bias of 1.50. Below that, A wins. More useful still, a bias that is stable can be measured and subtracted: correct A by one point and its MSE falls to 4, better than either original. That is the practical lesson for a risk team: a model that is consistently off in one direction is fixable, while a noisy model is not. The trade-off is also why risk teams use shrinkagePulling a noisy estimate towards a simpler, steadier target, accepting a little bias in return for much lower variance. on covariance matrices built from short histories.

    The limit: MSE punishes large errors heavily because it squares them, and it treats over-forecasts and under-forecasts alike. A risk manager forecasting losses may care more about under-forecasting than over-forecasting, in which case a symmetric score is the wrong yardstick and the ranking could change.

    Where candidates lose it

    Candidates pick B on reflex because unbiased sounds like correct. The question is built to see whether you know that MSE has two parts and can do the two-line arithmetic.

    The quieter miss is stopping at 5 against 6.25. Add that A's bias can be corrected, taking its MSE to 4, and you have turned a statistics answer into a model-risk judgement.

    What the interviewer asks next

    • What bias would make the two forecasters exactly equal?
    • Why might a regulator prefer the unbiased forecaster even with a higher MSE?
    • How would you test whether A's bias is stable over time?

    Asked at BlackRock, Restructuring, New York, 2026 (Wall Street Oasis): Which equities have duration ? multiple stocks vs value stocks MSE Forecasting equation

  9. 048You estimate a desk's daily P&L variance from five observations, once dividing the sum of squared deviations by 5 and once by 4. Which estimator is unbiased, which is consistent, and how large is the bias?Statistics and estimationCoreUBSZurich · 2021

    Try it first

    Which statement is true?

    Show the worked solution

    Dividing by 4 is unbiased; both estimators are consistent; dividing by 5 is low by one fifth. The sample mean is estimated from the same five points, which uses up one degree of freedom, so the divide-by-n estimator averages (n - 1)/n of the true variance, 80% here. If the true variance is 4, it centres on 3.2, a bias of -0.8. As n grows that factor tends to 1, so both estimators converge on the truth.

    Why does dividing by n come out too low?

    Measure how spread out five friends' heights are by comparing each to the group's own average, and you will understate the spread, because that average was pulled towards those five people. Deviations measured from the sample mean are smaller on average than deviations from the true mean, so their sum of squares understates the spread by exactly one observation's worth. Dividing by n minus 1, the degrees of freedomThe number of independent pieces of information left after estimating something from the same data; estimating the mean uses up one., corrects it exactly.

    Five observations: divide by 4 and you are centred; divide by 5 and you land low0481216Estimated daily variance, Rs crore squaredtrue variance 4divide by 5: mean 3.2divide by 4: centred at 4Mean squared errordivide by 4: 8.00divide by 5: 5.76biased, yet less error
    With five observations and a true variance of 4, the divide-by-4 estimator is centred on 4 while the divide-by-5 estimator is centred on 3.2, 80% of the truth, yet the biased version is narrower and has a lower mean squared error, 5.76 against 8.00.

    What is the difference between unbiased and consistent?

    Unbiased is about the average over many repeated samples of the same size; consistent is about what happens to one estimate as the sample grows. Divide-by-5 fails the first: repeat the five-day exercise many times and the estimates average 3.2, not 4. It passes the second: with 250 days the bias is only -0.016, and it shrinks to zero with more data. An estimator can be unbiased but inconsistent too, such as using only the first observation to estimate a mean: right on average, never improving.

    The relationship
    E ⁣[1n∑(xi−xˉ)2]=n−1n σ2=0.8×4=3.2E\!\left[\tfrac{1}{n}\textstyle\sum (x_i - \bar{x})^2\right] = \frac{n-1}{n}\,\sigma^2 = 0.8 \times 4 = 3.2
    nthe number of observations, 5
    \bar{x}the sample mean, estimated from the same five points
    \sigma^2the true variance, 4 in the illustration
    What it says in wordsDividing by n recovers only (n minus 1) over n of the true variance on average.

    Now the twist a model validator should add. For normal data the unbiased estimator has variance 8.00 here, while the divide-by-5 version has 5.12 plus a squared bias of 0.64, a mean squared error of 5.76. The biased estimator is closer to the truth on a typical sample. Which you prefer depends on the use: unbiasedness matters when estimates are averaged across many desks; a smaller typical error matters for a single desk's limit. With five data points, neither is reliable, and that is the more important thing to say.

    Where candidates lose it

    The usual slip is to treat unbiased and consistent as the same thing, and so to call the divide-by-5 estimator inconsistent. The interviewer asked both words together precisely to hear you separate them.

    The second miss is answering from memory without the reason. One sentence on the sample mean using up a degree of freedom shows you know why n minus 1 exists, not just that it does.

    What the interviewer asks next

    • Give an example of an estimator that is unbiased but not consistent.
    • Is the sample standard deviation, the square root of the unbiased variance, itself unbiased?
    • With 250 days of P&amp;L, does the choice between n and n - 1 matter for VaR?

    Asked at UBS, Risk Management, Zurich, 2021 (Wall Street Oasis): And several other questions on econometrics - what is an unbiased estimator vs consistent estimator?

  10. 056Two stocks both have a 10% cost of equity. One grows its dividends at 8% a year, the other at 2%. Using the Gordon growth model, how much does each price fall if the discount rate rises by 50 basis points?Duration and ratesCoreBLBlackRockNew York · 2026

    Try it first

    Which stock falls more when the discount rate rises half a point?

    Show the worked solution

    The 8% grower falls 20%; the 2% grower falls about 5.9%. Under Gordon growth, price is next year's dividend over r minus g. For the fast grower that gap widens from 2% to 2.5%, so the price falls to 0.02 over 0.025, or 80% of what it was. For the slow grower the gap goes from 8% to 8.5%, and the price keeps 0.08 over 0.085 of its value.

    Why does the fast grower react so much more?

    Think of two ways to be paid Rs 10 lakh: most of it next year, or a trickle that grows for decades. If someone doubles the rate at which you discount the future, the trickle loses far more, because most of its money is far away. A fast-growing dividend is that trickle: its value sits in cash flows many years out. A stock whose value rests on distant cash flows behaves like a long bond, so the same rise in the discount rate cuts its price far more.

    Same 50 basis point rise, very different price falls4060801001201409.5%10.0%10.5%11.0%11.5%12.0%Discount rate (cost of equity)both 100 at 10%g = 2%: 94.1, down 5.9%g = 8%: 80.0, down 20.0%Equity duration = 1 / (r - g)50 years against 12.5 years
    Both stocks are priced at 100 with a 10% cost of equity. A rise to 10.5% takes the 8% grower to 80, a fall of 20%, and the 2% grower to 94.1, a fall of 5.9%, because the fast grower's price rests on a gap of only 2 points between r and g.

    How do you turn this into a duration number?

    Differentiate the price with respect to r and divide by price: the answer is 1 over (r minus g). That gives the fast grower an equity duration of 50 years and the slow grower 12.5 years. Duration times 0.5% predicts falls of 25% and 6.25%; the exact falls are a little smaller, 20% and 5.9%, because the price curve bends, the same convexity a bond has.

    The relationship
    P=D1r−gDeq=−1PdPdr=1r−gP = \frac{D_1}{r-g} \qquad D_{\text{eq}} = -\frac{1}{P}\frac{dP}{dr} = \frac{1}{r-g}
    D_1next year's dividend
    rthe cost of equity, 10%
    gthe constant dividend growth rate, 8% or 2%
    What it says in wordsAn equity's sensitivity to the discount rate is one over the gap between the discount rate and growth.

    The limitation is that Gordon growth assumes growth never changes and runs for ever, which exaggerates duration for a fast grower that will slow. The direction survives any sensible model: growth stocks carry more rate risk than stocks priced on today's cash.

    Where candidates lose it

    The trap is answering that both fall by about the same amount because the rate change is the same. The rate change is the same; the base it lands on is not. The fast grower's r minus g is a quarter of the slow grower's, so the same half point is four times as large relative to it.

    The second slip is quoting the duration answer, 25%, as exact. Give 20% and say duration overstates it because the price curve bends.

    What the interviewer asks next

    • What happens to each price if growth expectations for the fast grower fall to 7% at the same time?
    • Why might a portfolio of growth stocks behave like a long-duration bond fund?
    • What does equity duration mean for a pension fund that holds equities against long liabilities?

    Asked at BlackRock, Risk and Quantitative Analysis, New York, 2026 (Wall Street Oasis): Which equities have duration? Technical and behavioural on VaR, market views and stock valuation.

← PreviousPage 1 of 2
  1. 1
  2. 2
Next →
Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.