VaR Backtesting: Counting Exceptions and What They Mean
Backtesting lays a risk measure against what actually happened. An exception is a day on which the realised loss went past the measure taken that morning. Vindhya Commercial Bank Limited, an invented bank, recorded seven exceptions in 250 days against about 2.5 expected, and four of the seven fell in consecutive pairs. The pattern matters more than the count.
Backtesting is what happens when one claim is taken seriously. A measure at 99 per cent is a claim about a share of days: on roughly ninety nine days in a hundred, the loss will stay inside the figure. A claim about a share of days is a testable claim, and counting the days is the test. No statistics beyond counting are needed. The method counts days, subtracts one column from another, and divides. The whole method fits in three arithmetic operations, and almost everything interesting about it lies in what is done with the answer rather than in getting it.
What is a backtest actually laying against what?
BacktestingLaying a risk measure against realised outcomes over a stated window and counting the occasions it was wrong. puts two numbers side by side that were produced at different ends of the same day, by different processes, for different purposes. The first is a forecast. The second is a fact. At Vindhya Commercial Bank Limited the forecast is produced every morning before the market opens: a one day figure at 99 per cent, computed on the Rs 3,600 crore held for trading book by historical simulation. The fact is produced every evening after the book is revalued: the rupee amount the trading book actually lost or made that day. A backtest is the discipline of writing both numbers on the same line and doing that again tomorrow.
Consider a household running a monthly budget. The shape is the same and it is easier to feel. On the first of each month the household sets aside Rs 8,000/- for the things that go wrong: a scooter service, a doctor, a school demand nobody warned it about. At the end of the month it adds up what actually went wrong. Some months the spending is Rs 3,000/-, some months Rs 7,500/-. Neither is interesting on its own. The interesting month is the month the household spent Rs 11,000/-. In that month the set-aside was not enough. More interesting still is how many such months there were in a year. The set-aside is not judged by how close it came each month; it is judged by how often it was not enough. A household running that comparison is running a backtest, in the same shape a bank runs it.
Two things about that shape are worth pinning down before anything else. Both are places a reader can quietly go wrong. The first is that the morning figure was never a promise about the size of the loss. The morning figure is a threshold, and the test asks only whether the day landed on one side of it or the other. The second is that the two numbers are produced independently. Nobody who values the book in the evening is looking at what the model said that morning. If the two processes ever start informing each other, the test stops being a test and becomes a conversation.
What exactly makes a day an exception?
An exceptionA day on which the realised loss went past the measure taken that morning, which is why the same rupee loss can be an exception on one day and not on another. is a day on which the realised loss went past the measure taken that morning. Read that sentence twice. The phrase that morning is doing almost all of the work, and it is the phrase people drop. An exception is not a loss above some fixed rupee amount, not a loss above the worst loss of last year, and not a loss that upset anybody. An exception is a loss that went past a specific number, computed a few hours earlier, that has to be looked up before the day can be classified at all.
The consequence of that is not obvious until it is seen in figures, so two days from this invented bank's year are worth setting side by side. On month 11 day 22, the trading book lost Rs 18.6 crore and the measure taken that morning read Rs 15.8 crore. The loss went past the measure, so the day is exception X7. The identical loss on a morning when the measure had read Rs 19.2 crore instead would be a different matter: same rupees out of the door, same book, same desk. The day with the higher morning measure would not be an exception at all, and no row would be written. The same rupee loss can be an exception on one day and an ordinary day on another, and nothing about the loss changed.
The dependence on the morning figure is not a defect in the test. The test exists to measure exactly that. The test is not asking whether the trading book had a bad day; the trading book is allowed to have bad days and it is measured every morning precisely because it will. The test is asking whether the model that produces the morning figure knew that the day ahead could be bad. A day on which the model said Rs 19.2 crore and the book lost Rs 18.6 crore is a day the model got approximately right. A day on which the model said Rs 15.8 crore and the book lost Rs 18.6 crore is a day the model got wrong, and it got wrong by Rs 2.8 crore.
Recording an exception is therefore a three column act and not a one column act. The record carries the day. The record carries the realised loss. And the record carries the measure that was standing that morning. Without the third column what has been recorded is a bad day rather than a statement about a model, and the difference between those two things is the entire subject. Vindhya Commercial Bank Limited keeps all three columns, and the seven rows it wrote in the year are set out in full below.
What exactly makes a day an exception?
How many exceptions should 250 days at 99 per cent produce?
About 2.5, and the arithmetic that gets there needs nothing beyond a multiplication. A measure at 99 per cent claims that the loss stays inside it on 99 per cent of days. Saying that is the same as saying it goes past on 1 per cent of them. Over an observation windowThe number of days the test runs over, being 250 in this invented bank, and part of what any exception count means. of 250 days, 1 per cent of 250 is 2.5. The expected count is not a property of the model, it is a property of the confidence level multiplied by the length of the window, and moving either one moves it.
Vindhya Commercial Bank Limited recorded seven. Seven against 2.5 is 2.8 times the expectation. Put another way, the measure went past on 7 of 250 days, being 2.8 per cent of them, where the level it was set at claimed 1.0 per cent. A miss of that size is large by any reading, and large enough that nobody at this invented bank argued about whether it mattered.
Two cautions belong right here, before the number gets carried anywhere. The first is that the expected countThe number of exceptions a measure at a stated confidence level implies over the window, being about 2.5 here, which is an expectation rather than a permission. is an expectation and not a permission. Nothing entitles a bank to two exceptions and no more. A year with two exceptions is unremarkable, and a year with four is not evidence of anything much on its own. 250 days is a short window, and counts of small numbers bounce around. The second caution is the reverse of the first and matters more: a count that is a long way from the expectation is worth taking seriously precisely because the window is short, and seven where 2.5 was expected is a long way.
The shape of the problem has just changed. The problem began with a model producing a number every morning. A morning number is an object nobody outside a risk function can inspect. The problem ends with a single integer, seven. Anybody can inspect an integer. The compression from a model to an integer is what makes backtesting useful and also what makes it dangerous, and both halves of that are developed below.
Seven exceptions arrived where about 2.5 were expected. What does that say about how bad those seven days were?
What do the seven days look like with all three columns printed?
Here is the whole year of Vindhya Commercial Bank Limited in seven rows. The bank labels them X1 to X7 and each row carries the day, the realised loss on the held for trading book, and the measure that had been taken that morning. Every figure is in Rs crore and every one of them belongs to that invented bank.
| Exception | Day | Realised loss | Measure that morning | Excess |
|---|---|---|---|---|
| X1 | month 2 day 9 | Rs 19.4 crore | Rs 14.8 crore | Rs 4.6 crore |
| X2 | month 3 day 14 | Rs 24.6 crore | Rs 15.2 crore | Rs 9.4 crore |
| X3 | month 3 day 15 | Rs 21.0 crore | Rs 15.4 crore | Rs 5.6 crore |
| X4 | month 6 day 3 | Rs 17.8 crore | Rs 15.0 crore | Rs 2.8 crore |
| X5 | month 9 day 2 | Rs 28.2 crore | Rs 15.6 crore | Rs 12.6 crore |
| X6 | month 9 day 3 | Rs 20.4 crore | Rs 16.2 crore | Rs 4.2 crore |
| X7 | month 11 day 22 | Rs 18.6 crore | Rs 15.8 crore | Rs 2.8 crore |
| Seven days | five months of twelve | Rs 150.0 crore | Rs 108.0 crore | Rs 42.0 crore |
Two of those figures need naming carefully. Each of them wears a second meaning elsewhere in this invented case, and a reader who carries the wrong one away will get an argument wrong later. The Rs 16.2 crore in the X6 row is the measure taken on the morning of month 9 day 3, and it is not the Monte Carlo figure of the same size that appears when this bank compares its three ways of computing the measure. And the Rs 18.6 crore in the X7 row is a realised loss on month 11 day 22. The Rs 18.6 crore realised loss is a different object entirely from the Rs 18.6 crore measured figure that ran against the Rs 18.0 crore cap of limit L5 on month 3 day 22. Two objects, one number, and the only protection is to name which object is meant every time the figure is printed.
The third column deserves attention before anything else. The morning measure ranges from Rs 14.8 crore to Rs 16.2 crore, a spread of Rs 1.4 crore across an entire year. A spread that narrow is a model producing a stable, unexciting figure morning after morning. The second column ranges from Rs 17.8 crore to Rs 28.2 crore, a spread of Rs 10.4 crore. The loss being measured moved more than seven times as much as the measure of it did, and that observation is available from the table alone before any counting starts.
How big were the misses, and what does the gap between the two columns sum to?
The excessThe realised loss less the measure taken that morning, which is the only part of an exception that carries its size. on an exception is the realised loss less the measure taken that morning, and it is the only part of the row that carries size. Worked down the seven rows of Vindhya Commercial Bank Limited, it gives Rs 4.6 crore, Rs 9.4 crore, Rs 5.6 crore, Rs 2.8 crore, Rs 12.6 crore, Rs 4.2 crore and Rs 2.8 crore. The seven excesses add to Rs 42.0 crore, and Rs 42.0 crore over seven days is exactly Rs 6.0 crore a day.
The whole of that can be checked in one line without adding seven numbers. The one-line check is worth running for a second reason: it proves the two columns tie. The seven realised losses sum to Rs 150.0 crore. The seven morning measures sum to Rs 108.0 crore. Rs 150.0 crore less Rs 108.0 crore is Rs 42.0 crore. The sum of the excesses is one column total less the other. The excess column is therefore not extra information so much as the same information written as a gap.
The Rs 42.0 crore needs naming carefully wherever it is carried. Here it is the sum of the seven backtesting excesses at Vindhya Commercial Bank Limited. The same invented case carries a Rs 42.0 crore figure that is the gross loss on one operational incident, and another that is a count of 42 control findings, and a third that is the number of risk data elements carrying all eight of their attributes. None of those is this one. A number without its object attached is a number waiting to be misread.
Two figures held next to each other change what the count appeared to say. The average realised loss across the seven days was Rs 21.43 crore. The average morning measure across the same seven days was Rs 15.43 crore. On the days the measure was wrong, it was wrong by about forty per cent of itself, every time, on average. A miss of forty per cent every time is a very different sentence from seven exceptions, and it came out of the same seven rows.
The seven realised losses sum to Rs 150.0 crore. What do the seven morning measures sum to, and what is the difference?
Do the seven rank the same way in rupees as they do in ratio?
Almost. There are two honest ways to say which of the seven days was worst, and it is worth running both because they answer slightly different questions. Ranking by rupee excess asks how much money the model failed to see. Ranking by the ratio of the realised loss to the morning measure asks how badly the model was wrong in proportion to what it had claimed. A Rs 4.2 crore miss against a measure of Rs 16.2 crore and a Rs 4.6 crore miss against a measure of Rs 14.8 crore are close in rupees and further apart in proportion.
By excess in rupees at Vindhya Commercial Bank Limited, the order runs X5 at Rs 12.6 crore, X2 at Rs 9.4 crore, X3 at Rs 5.6 crore, X1 at Rs 4.6 crore, X6 at Rs 4.2 crore, and then X4 and X7 both at Rs 2.8 crore with nothing to separate them. By the ratio of loss to measure, the order runs X5 at 1.81, X2 at 1.62, X3 at 1.36, X1 at 1.31, X6 at 1.26, X4 at 1.19 and X7 at 1.18. The two rankings agree everywhere except at the tie, where the rupee measure cannot see a difference and the ratio can.
The agreement between the two rankings is not an accident here and not a rule anywhere. The agreement holds at this invented bank because the seven morning measures sit in a narrow band of Rs 14.8 crore to Rs 16.2 crore, so dividing by them barely reorders anything. On a book whose morning measure genuinely moved with market conditions, a small rupee miss against a small measure could easily outrank a larger rupee miss against a much larger one. The lesson to hold is not that the two rankings agree. The lesson is that both rankings should be run and the answer found out.
Why is the pattern of the seven worth more than the count of them?
Because of where they fell. Sorted by day rather than by size, the seven show something that no amount of staring at the integer seven would ever reveal. X2 fell on month 3 day 14 and X3 fell on month 3 day 15. X5 fell on month 9 day 2 and X6 fell on month 9 day 3. Four of the seven exceptions at Vindhya Commercial Bank Limited fell in two consecutive pairs.
The position of the pairs matters because of what sits underneath a confidence level. When a measure claims that the loss stays inside it on 99 per cent of days, there is an independence assumptionThe assumption behind a confidence level that one day's outcome says nothing about the next, and the assumption pairs of exceptions call into question. riding along with the claim: that today's outcome says nothing about tomorrow's. Days that genuinely behave that way scatter their exceptions across the window. Exceptions on independent days do not arrive next to each other. ClusteringExceptions falling next to each other in time, which a model assuming each day independent of the last does not produce. is what appears when a model is not reacting to what has just happened.
Think about a weather forecast that keeps missing on the second day of every storm. The forecast is not wrong about how often it rains. The forecast is wrong about the fact that a wet day makes the next day likelier to be wet, and its miss on Tuesday is a direct consequence of its miss on Monday. A miss on the second day of a storm is precisely the shape of the pairs here. A model computing its morning figure from a rolling window of past days does not fully register a violent day until that day is already inside the window. Registering it then is a day late, and a day late is exactly when the second half of a pair arrives.
Now read the same seven along the calendar. The seven fall in months 2, 3, 6, 9 and 11. Five months of twelve carry all seven exceptions and seven months carry none at all. Months 3 and 9, the two months holding the pairs, carry Rs 45.6 crore and Rs 48.6 crore of realised loss between them, being Rs 94.2 crore of the Rs 150.0 crore lost across all seven days. Four of the seven by count, and 62.8 per cent by value, sit inside two months of the year. A backtest that reports a rate has described a year in which nothing much happened seven times. The calendar says the year had two episodes.
Four of the seven exceptions fell in two consecutive pairs. Why is that worse news than the count of seven?
How unusual is two pairs, roughly?
A rough sizing follows, and the word rough is doing real work in that sentence. Across 250 days there are 249 consecutive day pairs, being day 1 with day 2, day 2 with day 3, and so on to the end. If days really were independent of each other, the chance that both days of any given pair are exceptions would be the exception rate multiplied by itself. At the rate that was actually observed, 2.8 per cent, that is 249 multiplied by 0.028 multiplied by 0.028. The product is about 0.2 pairs across the whole year. Two occurred, about ten times as many.
The pair arithmetic is an illustration and not a test. The arithmetic falls short of a test for three plain reasons. The arithmetic uses the observed rate rather than the claimed one, and using the observed rate is already circular. The arithmetic treats the 249 pairs as if they were separate from each other. The pairs overlap: day 2 appears in two of them. And it produces no statement at all about how likely two pairs would be under any honest assumption. Formal tests of coverage and of independence exist, they are the right instruments for exactly this question, and they sit with model validation rather than here.
The sizing is good for one thing: deciding whether to look. Two against about 0.2 is far enough apart that nobody at Vindhya Commercial Bank Limited needed a formal result to justify opening the model up. The honest use of a back-of-an-envelope number is exactly that: a reason to look rather than a conclusion to publish, and a paper that presents it as a finding has claimed something the arithmetic never said.
Seven exceptions arrived where about 2.5 were expected. Before the control below is moved: how much bigger would the measure have to be each morning to leave only about two of the seven?
How much bigger would the morning measure have to be?
Multiply every one of the seven morning measures by the same amount and watch which of the seven realised losses is still beyond it. Bars that stay red are still exceptions; bars that turn grey have been cleared. The count gauge underneath carries a fixed marker at the roughly 2.5 exceptions that 250 days at the 99 per cent level implied.
The eight solved points are worth having as text as well as in the picture. The eight points are the whole of the behaviour the control has. A multiplier of 1.00 leaves seven exceptions. At 1.20 five remain, being X1, X2, X3, X5 and X6, and 1.25 leaves the same five. At 1.30 four remain, being X1, X2, X3 and X5. At 1.35 three remain, being X2, X3 and X5. At 1.40 exactly two remain, being X2 and X5, and two is the closest whole number to the roughly 2.5 that was expected. At 1.65 only X5 remains, and at 1.81 none of the seven is left. The seven ratios of realised loss to morning measure that produce all of that are X1 at 1.31, X2 at 1.62, X3 at 1.36, X4 at 1.19, X5 at 1.81, X6 at 1.26 and X7 at 1.18.
| Multiplier | Exceptions remaining | Which of the seven |
|---|---|---|
| 1.00 times | 7 | every one of X1 to X7, as recorded |
| 1.20 times | 5 | X1, X2, X3, X5 and X6 |
| 1.25 times | 5 | X1, X2, X3, X5 and X6 |
| 1.30 times | 4 | X1, X2, X3 and X5 |
| 1.35 times | 3 | X2, X3 and X5 |
| 1.40 times | 2 | X2 and X5 |
| 1.65 times | 1 | X5 alone |
| 1.81 times | 0 | none of the seven |
One thing about the 1.40 is worth noticing and not worth over-reading. The bank's ratio of its expected shortfall estimate to its value at risk figure is also 1.40, and 1.4038 multiplied by the average morning measure of Rs 15.43 crore gives Rs 21.7 crore against a tail estimate of Rs 21.9 crore. Two arithmetics landing near each other on one year of data at one invented bank is a coincidence worth a sentence and not a rule worth carrying anywhere. Neither of those two numbers is derived from the other.
One obvious thing about the control is worth saying plainly. No bank fixes a backtest by multiplying its measure by a constant. Doing that would raise the figure on all 250 days, including the 243 on which it was already comfortably enough. A flat increase turns a model that misses too often into a model that is useless too often. A real remedy changes the model itself: the window it looks back over, how it weights recent days, how it behaves when the market has just moved violently. The control shows, faster than anything else, how far the seven days sit from the model that measured them.
Is a backtesting exception the same thing as a limit breach?
No, and the two even fell in the same month at Vindhya Commercial Bank Limited. Falling inside one month makes the confusion very easy to commit. A limit breachA measured number going over a cap, which is a different event from an exception and in this bank fell on different days. is a measured number going over a cap. A backtesting exception is a realised loss going over the measured number. A breach and an exception are different objects, compared against different things, discovered by different people at different times of day.
Watch month 3 of this invented bank's year. On days 14 and 15 the trading book lost Rs 24.6 crore and Rs 21.0 crore against morning measures of Rs 15.2 crore and Rs 15.4 crore. The two days are exceptions X2 and X3, they were found in the evening, and nothing about them touched a cap. Then on days 21 and 22 the measured figure itself read Rs 19.2 crore and Rs 18.6 crore against the Rs 18.0 crore cap of limit L5. The reading over the cap is breach B5, it was found in the morning, and no realised loss was involved in finding it. The cause was a widening between two government bond maturities that the model had been treating as offsetting each other, and positions were cut on day 23.
Two consequences follow and both are worth having. The first is that the two worst backtesting days of the year were not the two days the limit was breached, so a report of the year that lists only breaches has not listed the days the model was most wrong. A limit breach asks whether the desk has taken more risk than it was allowed; a backtesting exception asks whether the number measuring that risk was any good. The second is about the vocabulary a draft paper uses. The sentence the bank breached its value at risk limit seven times this year merges two events, and it is wrong twice over: seven is the count of exceptions, and the limit was breached on two days.
A draft paper says the bank breached its value at risk limit seven times this year. What is wrong with that sentence?
How can one day carry two different failures at once?
Because a bad day in a market is bad in more than one place, and the controls that catch each one are looking at different things. Month 9 day 2 at Vindhya Commercial Bank Limited is the single worst backtesting day of the year: the held for trading book lost Rs 28.2 crore against a morning measure of Rs 15.6 crore, an excess of Rs 12.6 crore and a ratio of 1.81 times. The day is exception X5. The very same day, the bank's net overnight open currency position reached Rs 276 crore against the Rs 240 crore cap of limit L7, an excess of Rs 36 crore being 15.0 per cent over. The currency excess is breach B2, and it was caused by a customer deal booked after the cut-off and squared the following morning.
Drawing an arrow between those two is extremely tempting. Do not. The trading book loss and the currency position excess were found by different reports, concern different exposures, and have different causes on the record. Two failures on one bad day is a coincidence of the calendar and it becomes an invention the moment somebody writes that one caused the other.
There is a practical reason this matters beyond tidiness. Recording X5 as having been caused by the currency deal writes down a cause for the worst backtesting day of the year that is not the cause, and the model review that follows will go looking in the wrong place. A wrong cause is worse than a missing one. A missing cause keeps the question open and a wrong one closes it.
Month 9 day 2 carries both the worst backtesting exception and a currency limit breach. Did one cause the other?
What did seven exceptions actually trigger, and who decided?
Two things fired at Vindhya Commercial Bank Limited, and both of them are that bank's own policy rather than anybody's requirement. Its market risk policy escalates the backtest to the market risk committee, committee G7, at five exceptions in 250 days. The policy requires a model reviewWhat this bank's own policy requires at seven exceptions in 250 days, and an activity that belongs to model validation rather than to backtesting. at seven. Seven arrived, so both thresholds were crossed and both actions followed.
Notice the shape of that policy. The shape is a good one, and it is copyable. The policy has two steps rather than one. The first step is telling somebody, and it is cheap: a committee that meets monthly gets the backtest in front of it earlier than it otherwise would. The second step is doing something expensive. The expensive step is opening the model. A single threshold forces a bank to choose between reacting too early and reacting too late, and two thresholds let it do the cheap thing first.
A model review has its own contents, its own performer, its own way of preserving independence and its own entry in the model inventory, and all four belong to model risk rather than to backtesting. The one thing worth stating at this boundary is that a backtest and a validation are not the same activity even though both are about a model being right. A backtest tests one model's output against realised outcomes on 250 days. A validation asks whether the model is fit for the use it is being put to, reaching its data, its assumptions, its implementation and its stated limitations. A validation can find a model unfit that backtests perfectly well.
What did seven exceptions trigger at this invented bank?
Where does the backtesting approach come from, and what applies in India?
The mechanism is jurisdiction-free and can be run anywhere with a pencil. A measure is taken in the morning, an outcome is observed in the evening, the days on which the second went past the first are counted, and the count, the sizes and the pattern are read. Nothing in that requires a rule. A rule is required for the next question, namely what anybody official does about a particular count, and that question belongs to the authorities that set the rule.
Where the supervisory figures are to be confirmed
The Bank for International Settlements at bis.org is the origin of the Basel market risk framework and of the traffic lightA supervisory device banding an exception count, published by the Bank for International Settlements, with no band or boundary stated here. approach that bands an exception count into zones. The Reserve Bank of India at rbi.org.in sets what actually binds an Indian bank: which measurement approach may be used on the trading book, what must be computed and reported, how a supervisor treats an exception count and from what date. The Indian Banks Association at iba.org.in carries banking operational convention in India. Five exceptions and seven are the invented bank's own escalation policy and bind nobody else. Bands, zone boundaries, multipliers, capital consequences and effective dates come from those two bodies. Confirm every one of the rest at source before it is used for anything.
A supervisory band is a live number attached to a live rule, and any reproduction of one is wrong the day the rule moves and gives no warning that it has become wrong. Naming the mechanism, naming the body that publishes it and going to that body for the number is the only handling that stays true. The seven rows here stand on exactly the opposite footing: they belong to an invented bank and nothing in the world can move them.
Is the supervisory band that seven exceptions in 250 days falls into stated above?
Who actually reads a backtest, and what does each reader take from it?
Four people read the same seven rows at and around Vindhya Commercial Bank Limited, and no two of them are reaching for the same thing. Start with the one closest to the position. Devendra Achar, the invented bank's head of treasury, reads the backtest as a statement about the instrument he uses every day. His whole management of the trading book runs against a measure and a cap, and the backtest is the only evidence he has about whether that measure is worth managing against. Seven exceptions with two pairs tells him the number he steers by is late on exactly the days it matters. For the person using the measure, a backtest is a report card on the steering wheel rather than on the driving.
Second, an analyst outside the bank looking at published market risk disclosure. An outside analyst cannot see the model and would not understand it if shown. An analyst can compare a count against a window and a level. The comparison is arithmetic anybody can check, and that is precisely why an exception count is disclosed at all. The honest reading for an outside analyst is narrow but real: a count far above its expectation says the disclosed measure understates the tail of that book, and it says nothing whatever about the size of the misses unless the excesses are disclosed too.
Third, a lender to this bank or a large wholesale depositor placing money with it. Their question is not about the model, it is about how much of the institution the model covers. The measure being backtested here is computed on Rs 3,600 crore of held for trading positions out of a Rs 96,000 crore balance sheet. A very good backtest on that book would still say nothing at all about the rest, and a very bad one, of the kind this year produced, still concerns only that slice. A lender who takes an exception count as a verdict on the whole institution has taken a verdict on 3.75 per cent of it.
Fourth, a household with a term deposit at this bank, who will never see any of these figures. The reason this reader still matters is that the entire apparatus, the morning measure, the evening comparison, the count, the escalation at five and the model review at seven, exists so that somebody looks at a book with money in it on the days it behaves badly. Seven exceptions in a year with two adjacent pairs is not a scandal and nobody lost a deposit. The year produced a model that was late twice, found by a control that worked, escalated by a policy that fired. A functioning backtest looks like exactly that from outside: not a clean result, but a bad result that reached somebody.
What is the commonest way a backtest goes wrong?
Stopping at the count, which is the commonest thing that happens to a backtest
Seven against an expectation of about 2.5 is a big enough miss to fill a committee paper on its own. The comfort of that single number is the trap. Once the paper says seven, the meeting becomes an argument about whether seven is acceptable, and an argument about whether a number is acceptable is a comfortable meeting to have. The argument has two sides, it has people who feel strongly, and it reaches a conclusion. The argument also throws away both of the findings the seven rows actually contain.
The first thing thrown away is the pattern. A measure at 99 per cent carries an assumption underneath it that one day's outcome says nothing about the next. Days that behave that way do not produce pairs. At this invented bank's observed rate, the rough arithmetic gives about 0.2 consecutive pairs across the year, and two occurred. Clustering is a different finding from the count and a worse one. A count that is too high can be answered by raising the measure. Pairs mean the measure is not reacting to the market on the day it most needs to. Seven exceptions spread evenly across twelve months and seven falling in five months with two adjacent pairs are the same count and completely different diagnoses.
The second thing thrown away is the size. The count is blind to how bad each day was, and deliberately so: a day that goes past the measure by Rs 2.8 crore and a day that goes past by Rs 12.6 crore are both exactly one. At Vindhya Commercial Bank Limited the seven realised losses ranged from Rs 17.8 crore to Rs 28.2 crore, and the integer seven cannot see a single rupee of that range.
And there is a third failure, a merge rather than an omission. Writing that the bank breached its value at risk limit seven times this year is very easy, and the sentence is wrong in two ways at once. Limit breach B5 fell on month 3 days 21 and 22, when the measured figure went over the Rs 18.0 crore cap of limit L5. Exceptions X2 and X3 fell on month 3 days 14 and 15, when a realised loss went past the measure. Different days, different faults, and the two worst backtesting days of the year were not the two days the limit was breached. The whole mistake costs a year of evidence about one model, traded for a single integer that a committee can argue about.
What do these seven rows actually give?
Backtesting began as a word and ends as a method and a diagnosis. Backtesting lays a measure against realised outcomes over a stated window and counts the days the measure was wrong. An exception is a realised loss going past the measure taken that morning. The same Rs 18.6 crore loss is therefore an exception on one morning and an ordinary day on another. At the 99 per cent level over 250 days the expectation is about 2.5, and Vindhya Commercial Bank Limited recorded seven, being 2.8 per cent of days against a claimed 1.0 per cent.
There are also three readings of the same seven rows that a count on its own cannot give. The sizes: seven realised losses of Rs 150.0 crore against seven morning measures of Rs 108.0 crore, leaving backtesting excesses of Rs 42.0 crore that average exactly Rs 6.0 crore. The rankings: by rupee excess and by ratio to the measure, the same order except that the ratio separates X4 and X7 where both sat at Rs 2.8 crore. The pattern: seven exceptions in five months of twelve, with two adjacent pairs, and the two paired months alone carrying Rs 94.2 crore of the Rs 150.0 crore lost across all seven. Three readings, one table, and only the first of the three is arithmetic anybody would call difficult.
And there are the two separations that stop a reader merging events that are not the same event. A limit breach is a measured number going over a cap and fell on month 3 days 21 and 22; a backtesting exception is a realised loss going over the measured number and fell, in that same month, on days 14 and 15. Month 9 day 2 carries exception X5 and currency limit breach B2, two failures on one bad day, and neither caused the other. A backtest produces a table and not a number, and almost everything worth knowing is in the columns nobody prints.
Sources
| Source | Document | Site |
|---|---|---|
| Bank for International Settlements | The Basel market risk framework and the backtesting approach that bands an exception count | bis.org |
| Reserve Bank of India | What actually binds a bank in India: which measurement approach may be used on the trading book, what must be computed and reported, how a supervisor treats an exception count and from what date | rbi.org.in |
| Indian Banks Association | Banking operational convention in India | iba.org.in |
Vindhya Commercial Bank Limited and Devendra Achar are invented.
Educational material. Not advice on any investment, tax, budget or market position.
