Risk Management puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 17
- Topics
- 13
- Hard
- 30
023You estimate a correlation of 0.5 between two assets from 36 monthly observations. Roughly what is its standard error, and is a later estimate of 0.3 from the next 36 months evidence that the relationship has changed?Quant riskAsset manager risk
Try it first
Is the drop from 0.5 to 0.3 statistically meaningful?
Show the worked solution
The standard error is about 0.125, and 0.3 is not evidence of a change. A rough rule is one minus the correlation squared, over the square root of n: 0.75 over 6. Using Fisher's transformation, the 95% interval for 0.5 runs from about 0.21 to 0.71, and 0.3 sits inside it. Testing the difference between two 36 month estimates gives z of about 0.97, well below 1.96.
Why is a correlation from 36 months so imprecise?
Think of judging whether two friends like the same films after watching three dozen films together. They might agree on 20 by chance even if their tastes barely overlap. A correlation is estimated from how often two series move together, and with 36 points random coincidences move the estimate a lot: the standard error of 0.5 is about 0.125, a quarter of the estimate itself. The rough rule is one minus the correlation squared over the square root of the sample size, 0.75 over 6.
The relationshiprho the estimated correlation, 0.5 n the number of observations, 36 tanh^{-1} Fisher's transformation, which makes the sampling error close to normal What it says in wordsThe error of a correlation shrinks only with the square root of the sample size, and Fisher's transformation gives a fair interval around it.The 95% interval for a correlation of 0.5 from 36 months runs from 0.21 to 0.71, and the interval for 0.3 from another 36 months runs from -0.03 to 0.57, so the two overlap heavily and the drop is within sampling noise. How do you test whether the relationship changed?
Transform both estimates with Fisher's z, which turns 0.5 into 0.549 and 0.3 into 0.310. The difference, 0.240, divided by its standard error of the square root of 2 over 33, about 0.246, gives a z of 0.97: well short of the 1.96 needed for 95% confidence. A gap that size would appear by chance about 33% of the time even if nothing had changed. The honest statement is that the data cannot tell 0.5 from 0.3.
Then say why a risk manager cares. Hedge ratios, portfolio volatility and diversification benefits are all built on correlations like this one, and treating 0.5 as exact makes a model look more certain than its data allow. The practical responses are longer or higher-frequency data, shrinking noisy estimates towards a sensible prior, and stress testing the portfolio at correlations across the whole interval rather than at the point estimate.
Where candidates lose it
The trap is reading the drop from 0.5 to 0.3 as a regime change. It is a 40% fall in the number but well inside the noise that three years of monthly data carry.
The second miss is quoting a standard error without saying it is rough. The simple formula is an approximation; Fisher's transformation is what you use for an interval or a test.
What the interviewer asks next
- How many monthly observations would you need to bring the standard error down to 0.05?
- Would weekly data over the same three years help, and what could go wrong?
- How would you stress test a portfolio whose risk depends on this correlation?
050An EWMA volatility model with a decay factor of 0.94 had yesterday's daily volatility estimate at 1.2%. Today's return is minus 3%. What is the updated volatility estimate?Bank market riskQuant risk
Try it first
Roughly where does the new daily volatility land?
Show the worked solution
About 1.38%. EWMA updates the variance, not the volatility: new variance is 0.94 x 1.2 squared plus 0.06 x 3 squared, which is 1.3536 plus 0.54, or 1.8936. The square root is 1.376%. The shock has only a 6% weight but enters squared, so it supplies 29% of the new variance.
Why blend variances rather than volatilities?
A household tracking how much its grocery bill swings would be misled if it averaged the size of the swings in rupees but ignored that one big swing matters far more than several small ones. Variance is the average squared move, so the model updates squared returns, and a move 2.5 times normal size counts 6.25 times as much. An EWMAExponentially weighted moving average: each day the estimate keeps a fixed share of yesterday and adds the rest from today, so older days fade geometrically. of variances is what makes one large shock move the estimate quickly. Blending the volatilities directly gives 1.31%, understating the jump.
Today's minus 3% return carries only a 6% weight in the EWMA update, but because it enters squared it supplies 28.5% of the new variance of 1.8936, lifting daily volatility from 1.20% to 1.38%. The relationship\lambda the decay factor, 0.94 \sigma_{t-1} yesterday's volatility estimate, 1.2% r_t today's return, minus 3% What it says in wordsKeep 94% of yesterday's variance, add 6% of today's squared return, then take the square root.What happens next, and what are the model's limits?
If tomorrow is flat, the estimate decays to the square root of 0.94 x 1.8936, about 1.33%. Each day of calm keeps 94% of the variance, so a shock's influence halves in about 11 trading days: ln 0.5 over ln 0.94. That is the design choice behind 0.94: fast enough to react to a new regime within days, slow enough not to swing on every move.
The limits: EWMA has no pull towards a long-run average, so after a calm spell it can drift very low and understate risk just before volatility returns, which GARCH-type models address with a mean-reversion term. It also treats up and down moves alike, while falling markets often raise volatility more than rising ones. And the choice of 0.94 is a convention for daily data, not a law, so the estimate should be backtested against realised moves.
Where candidates lose it
The usual slip is to blend the two volatilities directly, 0.94 x 1.2 plus 0.06 x 3, and answer 1.31%. It looks like the formula but applies it to the wrong quantity and understates the effect of the shock.
The second is forgetting the square root at the end and quoting 1.89% as a volatility. Say the units at each step: variance in, variance out, then volatility.
What the interviewer asks next
- What would the estimate be with lambda of 0.97 instead?
- How many quiet days until the estimate is back below 1.25%?
- Why might a risk manager prefer GARCH to EWMA for a ten-day VaR?
073A rating grade carries a predicted default rate of 1%. It has 500 obligors, and 9 defaulted last year. At 95% confidence, is the grade miscalibrated?Rating agencyModel validation
Try it first
Nine defaults against five expected. What does an exact one-sided test say?
Show the worked solution
No, not at 95% confidence, though it is close. With 500 obligors at 1%, you expect 5 defaults. The exact binomial chance of 9 or more is about 6.7%, above the 5% cut-off; the smallest count that would reject is 10. A normal approximation without a continuity correction gives z = 1.80 and wrongly rejects. One year of data rarely proves a grade wrong.
How surprising are nine defaults if the 1% is right?
If a coin lands heads 7 times in 10, you would not call it biased; 70 in 100 you would. Small samples swing. With 500 obligors at 1%, the expected count is only 5, and counts that small are lumpy. The question is not whether 9 is bigger than 5, but how often a correct 1% grade produces 9 or more by chance, and the exact answer is about one year in fifteen.
At a true 1% default rate across 500 obligors the count centres on 5, and 9 or more defaults happens 6.7% of the time. The one-sided 95% critical value is 10 defaults, so the observed 9 sits just inside the acceptance region. Why does the normal approximation get it wrong here?
The standard deviation of the count is the square root of 500 x 0.01 x 0.99, about 2.22. Nine is 1.80 standard deviations above five, past 1.645, so a quick z-test rejects. But the binomial is skewed and discrete at small counts; with a continuity correction z falls to 1.57 and the conclusion flips back. For counts under about ten, use the exact binomial.
The relationshipX the number of defaults in the grade 500 obligors in the grade 0.01 the predicted default rate under test What it says in wordsAdd up the chance of every outcome at least as bad as nine; if that total is above 5%, you do not reject.Two further limitations make the case for caution stronger. Defaults are correlated through the economy, so a single bad year moves many obligors together and the true spread of counts is wider than the binomial says. And a validation team would look at several years and at neighbouring grades before recalibrating. The right answer is to flag the grade for watching, not to declare it broken.
Where candidates lose it
The trap is the quick z-test: (9 - 5) / 2.22 = 1.8, above 1.645, reject. It is the answer most candidates give, and it is wrong because a normal curve is a poor stand-in for a skewed count with a mean of five.
The second miss is failing to mention correlation. Independent defaults are the most favourable assumption for rejecting; real defaults cluster, which makes nine even less conclusive.
What the interviewer asks next
- How many defaults would reject at 99% confidence?
- The grade shows 8, 9 and 7 defaults in three consecutive years. What now?
- How does default correlation change the test?
087A desk's daily returns have a volatility of 1%. Over 250 days its average daily return is 0.05%. How precisely is that mean estimated, and can you say the desk has skill?Asset manager riskQuant risk
Try it first
Is 0.05% a day, measured over one year, clearly different from zero?
Show the worked solution
Not precisely enough to claim skill. The standard error of the mean is the daily volatility over the square root of the number of days: 1% over the square root of 250, about 0.063%. The estimate of 0.05% is only 0.79 standard errors from zero, and a two standard error band runs from -0.076% to 0.176%. You would need about 1,600 days, 6.4 years, to clear zero.
Why is a year of daily data not enough?
Weigh yourself on a bathroom scale that jumps by two kilos each time you step on it. If you lost 100 grams last month, a week of readings will not show it; the jumps drown the signal. The precision of an average improves only with the square root of the number of observations, so a small edge buried in large daily noise takes years to show. Here the daily noise is 1% and the edge is 0.05%, twenty times smaller.
The relationshipsigma daily return volatility, 1% n number of daily observations, 250 t how many standard errors the mean is from zero What it says in wordsThe mean's uncertainty is the daily volatility divided by the square root of the number of days, and the edge is less than one of those units from zero.After 250 days the two standard error band around the 0.05% daily mean runs from -0.076% to 0.176% and straddles zero; after about 1,600 days, 6.4 years, the band narrows to 0.000% to 0.100% and only just clears zero. How long would it take, and what does that mean for judging desks?
Set the t-statistic to 2 and solve for n: n equals (2 times 1% over 0.05%) squared, which is 1,600 days. An annual Sharpe ratio of about 0.79 needs more than six years of data before it is statistically distinguishable from zero. That is longer than most desks keep the same strategy, so a risk team cannot rely on the P&L record alone; it looks at whether the edge has a reason, whether it survives out of sample, and how much the result depends on a few days.
State the assumptions. The calculation treats daily returns as independent with constant volatility. Fat tails and volatility clustering make the true uncertainty larger, so six years is a floor, not a promise.
Where candidates lose it
The common error is annualising the mean to 12.5% and declaring skill, as if a big annual number were proof. The annual volatility grows too, to about 16%, and the ratio of the two is what matters.
The other slip is dividing by 250 instead of its square root, which makes the mean look fifty times more precise than it is.
What the interviewer asks next
- The desk's volatility is 0.5% instead. How many days now?
- Why is the mean so much harder to estimate than the volatility?
- How would you judge a new desk that has only six months of history?
