Risk Management puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 17
- Topics
- 13
- Hard
- 30
012A Monte Carlo VaR at 99% uses 10,000 simulated paths. How many scenarios sit in the tail that defines the VaR, and how many paths do you need to halve the standard error of the estimate?MSCISan Francisco · 2018
Try it first
How many paths halve the standard error of the VaR estimate?
Show the worked solution
100 scenarios define the tail, and you need 40,000 paths to halve the error. At 99%, 1% of 10,000 paths is 100 scenarios, and the VaR is read at the edge of that group. Monte Carlo error falls as one over the square root of the number of paths, so halving it means four times the paths. For a book with Rs 10 crore of daily volatility, the error is about Rs 37 lakh at 10,000 paths.
Why does a 99% VaR rest on so few scenarios?
Picture an exit poll that interviews 10,000 voters but reports only on the one voter in a hundred who picked a small party. Your estimate for that party rests on about 100 people, not 10,000. A 99% VaR is read at the boundary of the worst 1% of outcomes, so only about 100 of 10,000 paths carry information about where that boundary is. For a normal book with Rs 10 crore of daily volatility the true VaR is Rs 23.26 crore, and the estimate wobbles around it by about Rs 37 lakh from one run of 10,000 paths to the next.
The relationshipp the tail probability, 1% n the number of paths f(q) the height of the P&L density at the VaR point What it says in wordsThe error of a quantile estimate shrinks with the square root of the paths and grows where the tail is thin.For a book with Rs 10 crore of daily volatility, the standard error of a 99% Monte Carlo VaR is about Rs 37.3 lakh at 10,000 paths and Rs 18.7 lakh at 40,000, because the error falls with the square root of the number of paths. Why does halving the error cost four times the paths?
Because the paths sit under a square root. To halve the error, the square root of the path count has to double, which means the path count itself must quadruple. That is why precision in Monte Carlo is expensive: going from about 1.6% error to 0.8% of the VaR costs four times the computing time. It also explains why desks use variance reductionSimulation techniques, such as antithetic paths or importance sampling, that give a more precise estimate from the same number of paths., especially importance sampling, which pushes more paths into the tail where the VaR is decided.
Say the limit plainly. More paths shrink sampling error only; they do nothing about model error. If the simulated distribution has the wrong tail, a million paths give a very precise estimate of the wrong number. And expected shortfall at 97.5% averages the worst 250 paths, so it is usually more stable than the 99% VaR read from a single boundary.
Where candidates lose it
The trap is answering 20,000 paths, assuming error falls in proportion to the path count. It falls with the square root, and the difference is a factor of two in cost that a quant interviewer will not let pass.
The other miss is thinking all 10,000 paths inform the VaR. Say the number 100 out loud; it shows you know why tail estimates are noisy.
What the interviewer asks next
- How many paths would you need for the same precision on a 99.9% VaR?
- How does importance sampling reduce the error without adding paths?
- Why is a Monte Carlo VaR often less stable from day to day than a historical VaR on 500 days of data?
Asked at MSCI, Risk Management, San Francisco, 2018 (Wall Street Oasis):
the next round was paper test, materials were similar to CFA level 1, and it mostly focused on monte carlo, dividend, and risk
023You estimate a correlation of 0.5 between two assets from 36 monthly observations. Roughly what is its standard error, and is a later estimate of 0.3 from the next 36 months evidence that the relationship has changed?Quant riskAsset manager risk
Try it first
Is the drop from 0.5 to 0.3 statistically meaningful?
Show the worked solution
The standard error is about 0.125, and 0.3 is not evidence of a change. A rough rule is one minus the correlation squared, over the square root of n: 0.75 over 6. Using Fisher's transformation, the 95% interval for 0.5 runs from about 0.21 to 0.71, and 0.3 sits inside it. Testing the difference between two 36 month estimates gives z of about 0.97, well below 1.96.
Why is a correlation from 36 months so imprecise?
Think of judging whether two friends like the same films after watching three dozen films together. They might agree on 20 by chance even if their tastes barely overlap. A correlation is estimated from how often two series move together, and with 36 points random coincidences move the estimate a lot: the standard error of 0.5 is about 0.125, a quarter of the estimate itself. The rough rule is one minus the correlation squared over the square root of the sample size, 0.75 over 6.
The relationshiprho the estimated correlation, 0.5 n the number of observations, 36 tanh^{-1} Fisher's transformation, which makes the sampling error close to normal What it says in wordsThe error of a correlation shrinks only with the square root of the sample size, and Fisher's transformation gives a fair interval around it.The 95% interval for a correlation of 0.5 from 36 months runs from 0.21 to 0.71, and the interval for 0.3 from another 36 months runs from -0.03 to 0.57, so the two overlap heavily and the drop is within sampling noise. How do you test whether the relationship changed?
Transform both estimates with Fisher's z, which turns 0.5 into 0.549 and 0.3 into 0.310. The difference, 0.240, divided by its standard error of the square root of 2 over 33, about 0.246, gives a z of 0.97: well short of the 1.96 needed for 95% confidence. A gap that size would appear by chance about 33% of the time even if nothing had changed. The honest statement is that the data cannot tell 0.5 from 0.3.
Then say why a risk manager cares. Hedge ratios, portfolio volatility and diversification benefits are all built on correlations like this one, and treating 0.5 as exact makes a model look more certain than its data allow. The practical responses are longer or higher-frequency data, shrinking noisy estimates towards a sensible prior, and stress testing the portfolio at correlations across the whole interval rather than at the point estimate.
Where candidates lose it
The trap is reading the drop from 0.5 to 0.3 as a regime change. It is a 40% fall in the number but well inside the noise that three years of monthly data carry.
The second miss is quoting a standard error without saying it is rough. The simple formula is an approximation; Fisher's transformation is what you use for an interval or a test.
What the interviewer asks next
- How many monthly observations would you need to bring the standard error down to 0.05?
- Would weekly data over the same three years help, and what could go wrong?
- How would you stress test a portfolio whose risk depends on this correlation?
025A Kalman filter tracks a random walk with process variance 1 and measurement variance 4. What gain does it settle at, and which simple smoother is it then equivalent to?UBSLondon · 2022
Try it first
Where does the gain settle?
Show the worked solution
The gain settles at about 0.39, and the filter becomes an exponentially weighted moving average. In steady state the prior variance P solves P squared minus P minus 4 equals zero, so P is about 2.56. The gain is P over P plus 4, about 0.39. Each new estimate is then 0.39 times the new reading plus 0.61 times the old estimate, which is exactly an EWMA.
What is a Kalman filter doing, in one picture?
Think of estimating how many people are in a stadium from a noisy turnstile count that you update every few minutes. Your last estimate is useful but the crowd keeps changing, and each new count is useful but noisy. A Kalman filter blends the old estimate and the new reading, weighting each by how much you trust it; the weight on the new reading is the Kalman gainThe share of the gap between a new measurement and the prior estimate that the filter accepts as news.. Process variance of 1 says the true value drifts by about one unit each step; measurement variance of 4 says each reading is off by about two units.
The relationshipQ the process variance, 1: how much the true value moves each step R the measurement variance, 4: how noisy each reading is P^- the variance of the estimate just before a reading arrives K the Kalman gain What it says in wordsIn steady state the uncertainty added by the drift each step exactly balances the uncertainty removed by each reading.Starting from a vague prior the Kalman gain begins at 0.96, drops to 0.55 and 0.44, and settles at 0.390 by about the sixth update, after which the filter is an exponentially weighted moving average with weight 0.39 on each new reading. Why does the gain settle rather than fall to zero?
Because the thing being tracked keeps moving. If the true value were fixed, every reading would add certainty and the gain would shrink towards zero, like a running average; with a random walk, each step adds one unit of variance back, so certainty stops improving at a balance point. Plugging in Q of 1 and R of 4, the prior variance solves P squared minus P minus 4, giving (1 plus the square root of 17) over 2, about 2.56, and a gain of 0.390. Only the ratio of R to Q matters: a noisier measurement lowers the gain, a faster-moving state raises it.
That is the link worth saying in a risk interview. An EWMA volatility or correlation estimate, the kind many risk systems use, is a Kalman filter in steady state with a particular noise ratio, whether or not anyone calls it that. The filter's advantage is that it chooses the weight from stated assumptions about noise and drift, and adapts it early on; the limit is that those assumptions, a linear model with normal noise, have to be right for the weight to be the best one.
Where candidates lose it
The trap is describing a Kalman filter in general terms and never producing a number. The interviewer wants to see you set up the variance recursion and solve the steady state.
The second miss is saying the gain goes to zero. That is true only for a constant state; for a random walk it settles at a positive value, and that is why the filter reduces to an EWMA.
What the interviewer asks next
- What steady-state gain do you get with measurement variance 1 instead of 4?
- What EWMA decay factor corresponds to this filter, and what is its half-life in updates?
- How would you estimate Q and R from data?
Asked at UBS, Risk, London, 2022 (Wall Street Oasis):
Explain what a Kalman filter is
073A rating grade carries a predicted default rate of 1%. It has 500 obligors, and 9 defaulted last year. At 95% confidence, is the grade miscalibrated?Rating agencyModel validation
Try it first
Nine defaults against five expected. What does an exact one-sided test say?
Show the worked solution
No, not at 95% confidence, though it is close. With 500 obligors at 1%, you expect 5 defaults. The exact binomial chance of 9 or more is about 6.7%, above the 5% cut-off; the smallest count that would reject is 10. A normal approximation without a continuity correction gives z = 1.80 and wrongly rejects. One year of data rarely proves a grade wrong.
How surprising are nine defaults if the 1% is right?
If a coin lands heads 7 times in 10, you would not call it biased; 70 in 100 you would. Small samples swing. With 500 obligors at 1%, the expected count is only 5, and counts that small are lumpy. The question is not whether 9 is bigger than 5, but how often a correct 1% grade produces 9 or more by chance, and the exact answer is about one year in fifteen.
At a true 1% default rate across 500 obligors the count centres on 5, and 9 or more defaults happens 6.7% of the time. The one-sided 95% critical value is 10 defaults, so the observed 9 sits just inside the acceptance region. Why does the normal approximation get it wrong here?
The standard deviation of the count is the square root of 500 x 0.01 x 0.99, about 2.22. Nine is 1.80 standard deviations above five, past 1.645, so a quick z-test rejects. But the binomial is skewed and discrete at small counts; with a continuity correction z falls to 1.57 and the conclusion flips back. For counts under about ten, use the exact binomial.
The relationshipX the number of defaults in the grade 500 obligors in the grade 0.01 the predicted default rate under test What it says in wordsAdd up the chance of every outcome at least as bad as nine; if that total is above 5%, you do not reject.Two further limitations make the case for caution stronger. Defaults are correlated through the economy, so a single bad year moves many obligors together and the true spread of counts is wider than the binomial says. And a validation team would look at several years and at neighbouring grades before recalibrating. The right answer is to flag the grade for watching, not to declare it broken.
Where candidates lose it
The trap is the quick z-test: (9 - 5) / 2.22 = 1.8, above 1.645, reject. It is the answer most candidates give, and it is wrong because a normal curve is a poor stand-in for a skewed count with a mean of five.
The second miss is failing to mention correlation. Independent defaults are the most favourable assumption for rejecting; real defaults cluster, which makes nine even less conclusive.
What the interviewer asks next
- How many defaults would reject at 99% confidence?
- The grade shows 8, 9 and 7 defaults in three consecutive years. What now?
- How does default correlation change the test?
098A delinquency model was trained on a sample oversampled to 50% bad accounts, while the true bad rate is 2%. It scores an applicant at 30%. What is that applicant's probability of going bad in the real population?Neuberger BermanChicago · 2024
Try it first
Roughly what is the applicant's real-world probability of going bad?
Show the worked solution
About 0.87%. Oversampling inflates every score by the same factor on the odds. The sample's odds of bad are 50 to 50; the population's are 2 to 98, one forty-ninth as high. The applicant's sample odds are 30 to 70, 0.429; divide by 49 to get 0.00875, which is a probability of 0.87%. The model ranks correctly but must be recalibrated.
Why does oversampling change the score?
Suppose a doctor learns to spot a rare illness from a teaching ward where half the patients have it. In a village clinic, where only one patient in fifty has it, the same symptoms mean far less. A model trained on a 50% bad sample has learned a base rate 25 times too high, so every score it produces is inflated, even though the ordering from safest to riskiest is still right. Oversampling is done on purpose, to give the model enough bad accounts to learn from, so the correction is a routine step, not a sign of a broken model.
On a log-odds scale, oversampling shifts every score by the same amount, the log of 49, so a sample score of 30% maps to 0.87% in the real population and a sample score of 50% maps back to exactly the 2% base rate. The relationshippi true bad rate in the population, 2% s bad rate in the training sample, 50% odds probability of bad divided by probability of good What it says in wordsMultiply the model's odds by the ratio of the population's odds of bad to the sample's; then convert odds back to a probability.How do you sanity-check the answer?
Take an applicant the model scores at exactly 50%, the sample average. After correction that applicant should sit at the population average, and the formula gives exactly 2%, which confirms the factor. Then note that 30% is below the sample average, so the corrected figure should be below 2%, and 0.87% is. The shortcut of scaling the probability by 2% over 50% gives 1.2%, which is close for low scores but breaks down badly for high ones: a sample score of 90% would scale to 3.6%, while the correct answer is about 15.5%.
Say what a validator would do next. The correction assumes the good and bad accounts were each sampled at random within their class. If the bad accounts were drawn from a different period or channel, the model's ranking may also be off, and the fix is to check calibration on a recent, unsampled holdout, comparing predicted and actual bad rates by score band.
Where candidates lose it
The first trap is reporting 30% as the applicant's risk, which overstates it about thirty-fivefold and would, in a pricing or provisioning model, charge far too much for the loan.
The second is correcting the probability instead of the odds. It is close at low scores and wrong at high ones; say you adjust the odds, and check with the 50% applicant.
What the interviewer asks next
- What does the corrected probability become for an applicant the model scores at 90%?
- Does oversampling change the model's Gini or only its calibration?
- How would you recalibrate if the true bad rate itself shifts in a downturn?
Asked at Neuberger Berman, Risk, Chicago, 2024 (Wall Street Oasis):
How would you approach building a delinquency model?
