Risk Management puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 17
- Topics
- 13
- Hard
- 30
012A Monte Carlo VaR at 99% uses 10,000 simulated paths. How many scenarios sit in the tail that defines the VaR, and how many paths do you need to halve the standard error of the estimate?MSCISan Francisco · 2018
Try it first
How many paths halve the standard error of the VaR estimate?
Show the worked solution
100 scenarios define the tail, and you need 40,000 paths to halve the error. At 99%, 1% of 10,000 paths is 100 scenarios, and the VaR is read at the edge of that group. Monte Carlo error falls as one over the square root of the number of paths, so halving it means four times the paths. For a book with Rs 10 crore of daily volatility, the error is about Rs 37 lakh at 10,000 paths.
Why does a 99% VaR rest on so few scenarios?
Picture an exit poll that interviews 10,000 voters but reports only on the one voter in a hundred who picked a small party. Your estimate for that party rests on about 100 people, not 10,000. A 99% VaR is read at the boundary of the worst 1% of outcomes, so only about 100 of 10,000 paths carry information about where that boundary is. For a normal book with Rs 10 crore of daily volatility the true VaR is Rs 23.26 crore, and the estimate wobbles around it by about Rs 37 lakh from one run of 10,000 paths to the next.
The relationshipp the tail probability, 1% n the number of paths f(q) the height of the P&L density at the VaR point What it says in wordsThe error of a quantile estimate shrinks with the square root of the paths and grows where the tail is thin.For a book with Rs 10 crore of daily volatility, the standard error of a 99% Monte Carlo VaR is about Rs 37.3 lakh at 10,000 paths and Rs 18.7 lakh at 40,000, because the error falls with the square root of the number of paths. Why does halving the error cost four times the paths?
Because the paths sit under a square root. To halve the error, the square root of the path count has to double, which means the path count itself must quadruple. That is why precision in Monte Carlo is expensive: going from about 1.6% error to 0.8% of the VaR costs four times the computing time. It also explains why desks use variance reductionSimulation techniques, such as antithetic paths or importance sampling, that give a more precise estimate from the same number of paths., especially importance sampling, which pushes more paths into the tail where the VaR is decided.
Say the limit plainly. More paths shrink sampling error only; they do nothing about model error. If the simulated distribution has the wrong tail, a million paths give a very precise estimate of the wrong number. And expected shortfall at 97.5% averages the worst 250 paths, so it is usually more stable than the 99% VaR read from a single boundary.
Where candidates lose it
The trap is answering 20,000 paths, assuming error falls in proportion to the path count. It falls with the square root, and the difference is a factor of two in cost that a quant interviewer will not let pass.
The other miss is thinking all 10,000 paths inform the VaR. Say the number 100 out loud; it shows you know why tail estimates are noisy.
What the interviewer asks next
- How many paths would you need for the same precision on a 99.9% VaR?
- How does importance sampling reduce the error without adding paths?
- Why is a Monte Carlo VaR often less stable from day to day than a historical VaR on 500 days of data?
Asked at MSCI, Risk Management, San Francisco, 2018 (Wall Street Oasis):
the next round was paper test, materials were similar to CFA level 1, and it mostly focused on monte carlo, dividend, and risk
025A Kalman filter tracks a random walk with process variance 1 and measurement variance 4. What gain does it settle at, and which simple smoother is it then equivalent to?UBSLondon · 2022
Try it first
Where does the gain settle?
Show the worked solution
The gain settles at about 0.39, and the filter becomes an exponentially weighted moving average. In steady state the prior variance P solves P squared minus P minus 4 equals zero, so P is about 2.56. The gain is P over P plus 4, about 0.39. Each new estimate is then 0.39 times the new reading plus 0.61 times the old estimate, which is exactly an EWMA.
What is a Kalman filter doing, in one picture?
Think of estimating how many people are in a stadium from a noisy turnstile count that you update every few minutes. Your last estimate is useful but the crowd keeps changing, and each new count is useful but noisy. A Kalman filter blends the old estimate and the new reading, weighting each by how much you trust it; the weight on the new reading is the Kalman gainThe share of the gap between a new measurement and the prior estimate that the filter accepts as news.. Process variance of 1 says the true value drifts by about one unit each step; measurement variance of 4 says each reading is off by about two units.
The relationshipQ the process variance, 1: how much the true value moves each step R the measurement variance, 4: how noisy each reading is P^- the variance of the estimate just before a reading arrives K the Kalman gain What it says in wordsIn steady state the uncertainty added by the drift each step exactly balances the uncertainty removed by each reading.Starting from a vague prior the Kalman gain begins at 0.96, drops to 0.55 and 0.44, and settles at 0.390 by about the sixth update, after which the filter is an exponentially weighted moving average with weight 0.39 on each new reading. Why does the gain settle rather than fall to zero?
Because the thing being tracked keeps moving. If the true value were fixed, every reading would add certainty and the gain would shrink towards zero, like a running average; with a random walk, each step adds one unit of variance back, so certainty stops improving at a balance point. Plugging in Q of 1 and R of 4, the prior variance solves P squared minus P minus 4, giving (1 plus the square root of 17) over 2, about 2.56, and a gain of 0.390. Only the ratio of R to Q matters: a noisier measurement lowers the gain, a faster-moving state raises it.
That is the link worth saying in a risk interview. An EWMA volatility or correlation estimate, the kind many risk systems use, is a Kalman filter in steady state with a particular noise ratio, whether or not anyone calls it that. The filter's advantage is that it chooses the weight from stated assumptions about noise and drift, and adapts it early on; the limit is that those assumptions, a linear model with normal noise, have to be right for the weight to be the best one.
Where candidates lose it
The trap is describing a Kalman filter in general terms and never producing a number. The interviewer wants to see you set up the variance recursion and solve the steady state.
The second miss is saying the gain goes to zero. That is true only for a constant state; for a random walk it settles at a positive value, and that is why the filter reduces to an EWMA.
What the interviewer asks next
- What steady-state gain do you get with measurement variance 1 instead of 4?
- What EWMA decay factor corresponds to this filter, and what is its half-life in updates?
- How would you estimate Q and R from data?
Asked at UBS, Risk, London, 2022 (Wall Street Oasis):
Explain what a Kalman filter is
037Forecaster A has a bias of 1 point and a forecast error standard deviation of 2. Forecaster B is unbiased with a standard deviation of 2.5. Using the mean squared error decomposition, which forecaster is better?BlackRockNew York · 2026
Try it first
Which forecaster has the lower mean squared error?
Show the worked solution
Forecaster A, with a mean squared error of 5 against 6.25. Mean squared error splits into bias squared plus variance. A pays 1 squared for its bias and 2 squared for its spread, 5 in all. B pays nothing for bias but 2.5 squared, 6.25, for its spread. A's typical error, the square root, is 2.24 against 2.50.
How can a biased forecaster beat an unbiased one?
Two archers. One groups every arrow tightly but slightly left of centre; the other is centred on average but scatters arrows all over the target. Ask which one lands closer to the bullseye on a typical shot, and the tight grouping wins. Mean squared error charges for two things, how far off you are on average and how much you scatter, and a small, steady bias can cost far less than a large scatter. Being unbiased only removes the first charge.
The relationshipbias the average forecast error, forecast minus actual variance the spread of the errors around their own average, the standard deviation squared What it says in wordsSquared error on average equals the squared average error plus the spread of errors around it.Forecaster A's mean squared error is 1 of bias squared plus 4 of variance, 5 in total, while unbiased Forecaster B carries 6.25 of pure variance, so A's small bias buys a larger cut in variance and gives the lower error. When would your answer flip, and what would you do with A?
Solve for the tie: A matches B when bias squared plus 4 equals 6.25, so a bias of 1.50. Below that, A wins. More useful still, a bias that is stable can be measured and subtracted: correct A by one point and its MSE falls to 4, better than either original. That is the practical lesson for a risk team: a model that is consistently off in one direction is fixable, while a noisy model is not. The trade-off is also why risk teams use shrinkagePulling a noisy estimate towards a simpler, steadier target, accepting a little bias in return for much lower variance. on covariance matrices built from short histories.
The limit: MSE punishes large errors heavily because it squares them, and it treats over-forecasts and under-forecasts alike. A risk manager forecasting losses may care more about under-forecasting than over-forecasting, in which case a symmetric score is the wrong yardstick and the ranking could change.
Where candidates lose it
Candidates pick B on reflex because unbiased sounds like correct. The question is built to see whether you know that MSE has two parts and can do the two-line arithmetic.
The quieter miss is stopping at 5 against 6.25. Add that A's bias can be corrected, taking its MSE to 4, and you have turned a statistics answer into a model-risk judgement.
What the interviewer asks next
- What bias would make the two forecasters exactly equal?
- Why might a regulator prefer the unbiased forecaster even with a higher MSE?
- How would you test whether A's bias is stable over time?
Asked at BlackRock, Restructuring, New York, 2026 (Wall Street Oasis):
Which equities have duration ? multiple stocks vs value stocks MSE Forecasting equation
048You estimate a desk's daily P&L variance from five observations, once dividing the sum of squared deviations by 5 and once by 4. Which estimator is unbiased, which is consistent, and how large is the bias?UBSZurich · 2021
Try it first
Which statement is true?
Show the worked solution
Dividing by 4 is unbiased; both estimators are consistent; dividing by 5 is low by one fifth. The sample mean is estimated from the same five points, which uses up one degree of freedom, so the divide-by-n estimator averages (n - 1)/n of the true variance, 80% here. If the true variance is 4, it centres on 3.2, a bias of -0.8. As n grows that factor tends to 1, so both estimators converge on the truth.
Why does dividing by n come out too low?
Measure how spread out five friends' heights are by comparing each to the group's own average, and you will understate the spread, because that average was pulled towards those five people. Deviations measured from the sample mean are smaller on average than deviations from the true mean, so their sum of squares understates the spread by exactly one observation's worth. Dividing by n minus 1, the degrees of freedomThe number of independent pieces of information left after estimating something from the same data; estimating the mean uses up one., corrects it exactly.
With five observations and a true variance of 4, the divide-by-4 estimator is centred on 4 while the divide-by-5 estimator is centred on 3.2, 80% of the truth, yet the biased version is narrower and has a lower mean squared error, 5.76 against 8.00. What is the difference between unbiased and consistent?
Unbiased is about the average over many repeated samples of the same size; consistent is about what happens to one estimate as the sample grows. Divide-by-5 fails the first: repeat the five-day exercise many times and the estimates average 3.2, not 4. It passes the second: with 250 days the bias is only -0.016, and it shrinks to zero with more data. An estimator can be unbiased but inconsistent too, such as using only the first observation to estimate a mean: right on average, never improving.
The relationshipn the number of observations, 5 \bar{x} the sample mean, estimated from the same five points \sigma^2 the true variance, 4 in the illustration What it says in wordsDividing by n recovers only (n minus 1) over n of the true variance on average.Now the twist a model validator should add. For normal data the unbiased estimator has variance 8.00 here, while the divide-by-5 version has 5.12 plus a squared bias of 0.64, a mean squared error of 5.76. The biased estimator is closer to the truth on a typical sample. Which you prefer depends on the use: unbiasedness matters when estimates are averaged across many desks; a smaller typical error matters for a single desk's limit. With five data points, neither is reliable, and that is the more important thing to say.
Where candidates lose it
The usual slip is to treat unbiased and consistent as the same thing, and so to call the divide-by-5 estimator inconsistent. The interviewer asked both words together precisely to hear you separate them.
The second miss is answering from memory without the reason. One sentence on the sample mean using up a degree of freedom shows you know why n minus 1 exists, not just that it does.
What the interviewer asks next
- Give an example of an estimator that is unbiased but not consistent.
- Is the sample standard deviation, the square root of the unbiased variance, itself unbiased?
- With 250 days of P&L, does the choice between n and n - 1 matter for VaR?
Asked at UBS, Risk Management, Zurich, 2021 (Wall Street Oasis):
And several other questions on econometrics - what is an unbiased estimator vs consistent estimator?
062Your prior estimate of a hidden fair price is 10 with variance 4. A noisy measurement comes in at 12 with measurement variance 1. After one Kalman filter update, what is your new estimate and its variance?UBSLondon · 2022
Try it first
Where does the new estimate land?
Show the worked solution
The new estimate is 11.6 with variance 0.8. The Kalman gain is the prior variance over the total, 4 over 5, or 0.8. The estimate moves 80% of the way from 10 towards 12, landing at 11.6. The variance becomes (1 - 0.8) x 4 = 0.8, smaller than either the prior's 4 or the measurement's 1.
How does the filter decide how far to move?
Two friends guess your commute time. One has ridden with you a hundred times, the other once. You would average their guesses, but lean heavily on the first. A Kalman filterA method that updates an estimate of something you cannot see directly each time a noisy measurement arrives, weighting old estimate and new data by how precise each is. does exactly that: it weights the prior and the measurement by their precision, one over variance. Precision 0.25 against 1 gives the measurement 80% of the weight.
The prior centred at 10 is wide, with variance 4, and the measurement at 12 is narrow, with variance 1. The update lands at 11.6, four fifths of the way to the measurement, and its variance of 0.8 makes it narrower than either source. Why is the new variance smaller than both inputs?
Because two independent pieces of evidence together know more than either alone. Precisions add: 1 over 4 plus 1 over 1 is 1.25, and one over 1.25 is a variance of 0.8. That is the part candidates skip. The filter does not just move the estimate; it becomes more confident with every measurement, until new data carry little weight and the estimate settles.
The relationshipP the prior variance, 4 R the measurement variance, 1 K the Kalman gain, the weight on the new measurement \hat{x} the updated estimate What it says in wordsMove from the prior towards the measurement by the gain, and shrink the variance by the same share.Say the limitation: this single step assumes both errors are normal and the hidden price did not move between the prior and the measurement. A full filter adds a prediction step that lets the price drift and widens the variance before each update. Risk teams use the idea to track hidden quantities such as a hedge ratio that changes over time.
Where candidates lose it
The trap is averaging the two numbers and answering 11. That treats the measurement and the prior as equally trustworthy, which the variances say they are not.
The second trap is getting 11.6 and then saying the variance is somewhere between 1 and 4. Combining evidence always reduces uncertainty, so the new variance must be below both: 0.8.
What the interviewer asks next
- A second measurement of 11 arrives with variance 1. What is the estimate now?
- What happens to the gain as the number of measurements grows?
- How would you use a Kalman filter to estimate a time-varying hedge ratio?
Asked at UBS, Risk, London, 2022 (Wall Street Oasis):
Explain what kalman filter is.
098A delinquency model was trained on a sample oversampled to 50% bad accounts, while the true bad rate is 2%. It scores an applicant at 30%. What is that applicant's probability of going bad in the real population?Neuberger BermanChicago · 2024
Try it first
Roughly what is the applicant's real-world probability of going bad?
Show the worked solution
About 0.87%. Oversampling inflates every score by the same factor on the odds. The sample's odds of bad are 50 to 50; the population's are 2 to 98, one forty-ninth as high. The applicant's sample odds are 30 to 70, 0.429; divide by 49 to get 0.00875, which is a probability of 0.87%. The model ranks correctly but must be recalibrated.
Why does oversampling change the score?
Suppose a doctor learns to spot a rare illness from a teaching ward where half the patients have it. In a village clinic, where only one patient in fifty has it, the same symptoms mean far less. A model trained on a 50% bad sample has learned a base rate 25 times too high, so every score it produces is inflated, even though the ordering from safest to riskiest is still right. Oversampling is done on purpose, to give the model enough bad accounts to learn from, so the correction is a routine step, not a sign of a broken model.
On a log-odds scale, oversampling shifts every score by the same amount, the log of 49, so a sample score of 30% maps to 0.87% in the real population and a sample score of 50% maps back to exactly the 2% base rate. The relationshippi true bad rate in the population, 2% s bad rate in the training sample, 50% odds probability of bad divided by probability of good What it says in wordsMultiply the model's odds by the ratio of the population's odds of bad to the sample's; then convert odds back to a probability.How do you sanity-check the answer?
Take an applicant the model scores at exactly 50%, the sample average. After correction that applicant should sit at the population average, and the formula gives exactly 2%, which confirms the factor. Then note that 30% is below the sample average, so the corrected figure should be below 2%, and 0.87% is. The shortcut of scaling the probability by 2% over 50% gives 1.2%, which is close for low scores but breaks down badly for high ones: a sample score of 90% would scale to 3.6%, while the correct answer is about 15.5%.
Say what a validator would do next. The correction assumes the good and bad accounts were each sampled at random within their class. If the bad accounts were drawn from a different period or channel, the model's ranking may also be off, and the fix is to check calibration on a recent, unsampled holdout, comparing predicted and actual bad rates by score band.
Where candidates lose it
The first trap is reporting 30% as the applicant's risk, which overstates it about thirty-fivefold and would, in a pricing or provisioning model, charge far too much for the loan.
The second is correcting the probability instead of the odds. It is close at low scores and wrong at high ones; say you adjust the odds, and check with the 50% applicant.
What the interviewer asks next
- What does the corrected probability become for an applicant the model scores at 90%?
- Does oversampling change the model's Gini or only its calibration?
- How would you recalibrate if the true bad rate itself shifts in a downturn?
Asked at Neuberger Berman, Risk, Chicago, 2024 (Wall Street Oasis):
How would you approach building a delinquency model?
