Case 097Risk measurement and limitsHard
Vairagi's 99% one-day VaR was exceeded 8 times in 250 days, against 2.5 expected. Is the model wrong? Use the binomial distribution and the traffic-light idea to decide.
1The situation
Vairagi Capital's risk team backtests the fund's one-day 99% value at risk over the past 250 trading days. An exception is a day on which the realised loss exceeded that morning's VaR. A correct model should produce one exception in a hundred days, 2.5 over the year. The log shows 8, on days 37, 112, 118, 119, 121, 122, 205, 236.
The portfolio manager's view is that the year was unusually turbulent and eight is nothing to act on. The chief risk officer wants a number, not a view.
2Your task
Decide whether eight exceptions are consistent with a correct model, using the binomial distribution and the traffic-light zones, and say what else in the log matters.
Quick check
Under a correct 99% model, how likely are 8 or more exceptions in 250 days?
Worked solution
Try it on paper, then open one step at a time.
30-second answerThe answer to give first
Under a correct model the chance of 8 or more exceptions in 250 days is 0.40%, about one year in 248, so the model is presumed wrong until shown otherwise. Eight sits in the amber zone of the traffic-light framework, where a correct model lands only 10.8% of the time. The observed rate of 3.2% says the number reported as 99% VaR behaves like a 96.8% VaR, and the clustering of five exceptions in eleven days points at a volatility input that lagged the market.
Step 1What is the chance of eight under a correct model?
A coin that is meant to land heads one time in a hundred lands heads eight times in 250 throws. You do not need a theory of coins to be suspicious; you need the binomial distribution. If each day has an independent 1% chance of an exception, the count over 250 days is binomial with mean 2.5 and standard deviation 1.57, and the chance of 8 or more is 0.40%. That is one year in about 248. The portfolio manager's turbulent-year point is not an answer, because the VaR is supposed to update to the turbulence: a model that produces exceptions whenever markets are volatile is a model that does not measure volatility. The formal version is the Kupiec test: the likelihood ratio comparing p = 1% with the observed 3.2% is 7.7, against a chi-square critical value of 3.841, with a p-value of about 0.54%.
| X | the number of exceptions in 250 days, binomial under a correct model |
| LR | Kupiec's likelihood ratio, chi-square with one degree of freedom under the null |
| 0.032 | the observed exception rate, 8 over 250 |
Step 2What do the traffic-light zones add?
The zones turn the probability into a rule a committee can apply without arguing. The Basel framework for 250-day backtests puts 0 to 4 exceptions in green, 5 to 9 in amber and 10 or more in red, chosen so that a correct model is green about 89% of the time and red almost never. Amber carries a capital multiplier that rises with the count, and the burden of proof shifts to the risk team to explain why the model is not wrong. Confirm the current thresholds and multipliers before quoting them; the numbers here are the framework's design logic, not a live regulatory figure. The zones also admit what the test cannot do: with 250 days it is weak. If the true exception rate were 2%, double the stated risk, the model would still land in green 44% of the time. A green light is not a clean bill of health; a red or deep amber is a strong signal.
| Zone | Exceptions in 250 days | Chance under a correct model | Response |
|---|---|---|---|
| Green | 0 to 4 | 89.2% | Accept; no capital add-on |
| Amber | 5 to 9 | 10.8% | Investigate; add-on scales with the count |
| Red | 10 or more | 0.03% | Model presumed wrong; maximum add-on |
Step 3How wrong is the model, and where?
Eight exceptions estimate the true tail probability at 3.2%, and a 95% lower bound on it is about 1.6%, so the data rule out 1% but not 1.6%. Under a normal model, a quantile that is breached 3.2% of the time instead of 1% is at 1.85 standard deviations instead of 2.33, so the model understates volatility by roughly 26%, if the shape is right and the scale is wrong. The log tells you where to look: five of the eight exceptions fall between day 112 and day 122, and two pairs are on consecutive days. Under independence at a 3.2% rate, the chance that two or more of eight exceptions are followed next day by another is about 2.5%. Clustering is the signature of a volatility input that lagged the market for one stretch, a long lookback window or a slow decay, rather than a tail that is permanently fat. The Christoffersen test formalises this: it tests the count and the independence separately, and this log fails both.
Step 4What do you tell the chief risk officer?
Lead with the number, then the diagnosis, then the fix. The count is inconsistent with a correct model at any reasonable confidence, the realised rate is about three times the stated one, and the exceptions cluster, so the first suspect is the volatility estimate's speed of response, not the distribution's shape. Rerun the year with a faster decay or a shorter window and count again; if the cluster disappears and the count falls to the green zone, the fix is the window. If exceptions remain spread through calm periods, the tail is fat and the model needs a heavier-tailed distribution or a historical simulation. Say the limitation: 250 days is a short record, the test cannot see a model that is 50% too small, and a single cluster can be one bad fortnight. The response is to fix and re-test, not to raise the VaR by 26% and move on.
Where candidates lose it
The common loss is accepting the turbulent-year defence: eight exceptions because markets were wild. The VaR's job is to track that wildness. Exceptions that arrive together are evidence of a model that lagged, not of bad luck.
The second is treating the mean of 2.5 as a target and 8 as merely above it, without computing the tail. The binomial puts 8 at 3.5 standard deviations; the probability, under half a percent, is the answer.
What the interviewer asks next
- How many exceptions would a 97.5% expected-shortfall backtest expect, and how would you test it?
- Rerun the count for 500 days with 12 exceptions: which zone, and how does the test's power change?
- How would you test the independence of exceptions formally rather than by eye?
- If the eight exceptions were all small breaches of VaR, does that change your conclusion?
Company names and figures are illustrative.
