Value at Risk vs Expected Shortfall: The Cut-Off and the Tail
Value at risk and expected shortfall are not one question asked at two strengths. Value at risk locates a cut-off and says nothing about the days beyond it. Expected shortfall averages exactly those days. At Vindhya Commercial Bank Limited, invented, the two read Rs 15.6 crore and Rs 21.9 crore, and in its own recorded year the averaging measure was out by 2.2 per cent while the counting one was out by 180 per cent.
Almost everybody meets these two measures in the wrong order. Readers see the smaller figure first, meet the larger one later, and quietly file the larger one as the same measure set more conservatively. The misreading is expensive. The two measures do not disagree about how bad a bad day is. Their disagreement is about what a bad day even is a question about.
Are these two measures the same question asked at two strengths?
The shape of the difference is older than any trading book, and it shows most plainly outside finance. Suppose a household keeps two rain records for the roof over a workshop. The first record answers one question: on how many days did more than 40 millimetres fall? The second answers a different one: on the days when more than 40 millimetres fell, how much fell on average? A day is one day however hard it rains. A monsoon that drops 300 millimetres in a morning changes the second record and leaves the first exactly where it was. The first record can count catastrophes and never measure one. The information was never collected, so no amount of care in keeping the record fixes that.
The difference is that simple, and here it is in the language of a trading book. Value at riskA loss a stated book is not expected to exceed over a stated horizon on a stated share of outcomes, and nothing at all about the outcomes where it is exceeded. states a loss that a stated book is not expected to exceed over a stated horizon on a stated share of outcomes. Three things have to be named for the figure to mean anything: which book, over how long, and on what share of days. At Vindhya Commercial Bank Limited the book is the held for trading book of Rs 3,600 crore, the horizon is one day, the share is 99 per cent, and the figure that comes out of the bank's own historical simulation over 500 days is Rs 15.6 crore.
Expected shortfallThe average of the losses beyond a stated cut-off, on the same book over the same horizon, which is a statement about the size of bad days rather than about where they begin. states the average of the losses beyond a stated cut-off, on the same book over the same horizon. Expected shortfall needs the same three things named and one more besides: what is being averaged. At the same invented bank, on the same Rs 3,600 crore book, over the same one day, beyond a 97.5 per cent cut-off, the figure is Rs 21.9 crore. Read those two definitions again and notice that neither sentence contains the other: one marks where the bad days begin and the other says what they cost once they have begun. Every difference in the rest of this guide falls out of those two sentences and nothing else.
Which single sentence states what expected shortfall claims about a book?
Why does one of them never move when the worst day gets worse?
Draw the range of daily outcomes on a single axis, gains to the left and losses to the right, and the two measures become two completely different objects on the same picture. The threshold measureA measure that marks a point in the range of outcomes, which is what the first of these two is and why it carries no information about anything past that point. is one vertical mark. A mark is all it is. The mark sits at the 99 per cent point, it has a rupee value of Rs 15.6 crore at this invented bank, and everything to the right of it is outside what the mark describes. The tail measureA measure that averages what lies beyond a point, which is what the second of these two is and why it moves whenever those outcomes get worse. is not a mark at all. The tail measure is the centre of mass of an area: take everything beyond the 97.5 per cent point and average it.
Now stretch the far right of that picture. Make the worst day worse, and the next worst worse again, and keep going. The mark does not move. A mark answers one question, where does the worst one per cent begin, and stretching the outcomes past it does not change where it begins. The centre of mass of the shaded area moves every single time. The stretched outcomes are the only thing it is made of. The contrast is not a defect in the first measure and not a virtue in the second: it is what each one was built to answer, and reading a threshold as though it described the region past it is reading a signpost as though it were a map of the road ahead.
A cut-off can be known as a position without being known as an amount. Vindhya Commercial Bank Limited records a rupee figure for its 99 per cent mark and for the average beyond its 97.5 per cent cut-off, and none for the 97.5 per cent mark itself. The 97.5 per cent cut-off therefore has a place in the range of outcomes and no rupee value to print beside it.
What happens when two books are put together?
Here the argument stops being about tails and becomes about arithmetic that a risk committee has to live with. Artzner, Delbaen, Eber and Heath set out, in Coherent Measures of Risk, published in 1999, the properties any sensible measure of risk should satisfy. There are four of them, and three are uncontroversial: a book that is worse in every state should measure worse, doubling a book should double the measure, and adding cash should reduce the measure by the cash. The fourth is subadditivityThe property that combining two books never makes the measured risk larger than the two measured separately, which is the one that separates these two measures., and it says that combining two books should never make the measured risk larger than the two books measured apart. Together those four are what the paper calls coherenceA set of properties a risk measure can be tested against, set out by Artzner, Delbaen, Eber and Heath in 1999..
The averaging measure satisfies all four; the threshold measure can fail the fourth, and that is the property on which the two genuinely part company. The reason is easy to see once the picture from the last section is in mind. A threshold is a point, and where two books have their bad days on different days, the combined book can push a lump of outcomes just past a point that neither book alone reached. An average of everything past a point cannot do that. The average is built from the same rupees whichever way the books are cut.
Now the part that matters more than the mathematics. Ask what a desk head does with a measure that can report a combined book as riskier than its parts. The answer is that the desk head has a perfectly respectable technical argument for not being aggregated at all: put my book with theirs and the number gets worse, so leave the numbers apart. A measure that is not subadditive hands every desk in the institution an argument against being added up. Subadditivity is therefore a governance problem long before it is a mathematical one.
Who set out the properties a risk measure should satisfy, and which of those properties separates these two measures?
Which of the two got this bank's own year right?
Definitions settle what each measure claims. Definitions do not settle which one told this bank anything useful, and the invented case happens to contain a clean test of exactly that. Over 250 observation days Vindhya Commercial Bank Limited recorded seven exceptions, an exception being a day on which the realised loss went past the measure taken that same morning. The seven are X1 at Rs 19.4 crore, X2 at Rs 24.6 crore, X3 at Rs 21.0 crore, X4 at Rs 17.8 crore, X5 at Rs 28.2 crore, X6 at Rs 20.4 crore and X7 at Rs 18.6 crore.
The counting measure goes first. At 99 per cent over 250 days about 2.5 exceptions are expected. Seven arrived. Seven over 2.5 is 2.8 times the expectation, an error of 180 per cent, and as a share of days it is 2.8 per cent against an expected 1.0 per cent. The averaging measure comes second. The seven losses sum to exactly Rs 150.0 crore and average Rs 21.43 crore. The bank's tail estimate was Rs 21.9 crore. The estimate stood Rs 0.47 crore above what actually happened. On the realised average of Rs 21.43 crore that gap is an error of 2.2 per cent.
Same book, same 250 mornings, same seven days: the measure that estimates the size of the bad days was out by 2.2 per cent and the measure that counts them was out by 180 per cent. The comparison is the cleanest argument for the averaging measure anywhere in this invented case, and it is worth saying plainly that it is one year on one book and settles nothing on its own. Settling nothing is not the same as showing nothing. One of the two measures was nearly right about this year and the other was not close, so these are not two strengths of one reading.
In the recorded year, by how much was the count wrong and by how much was the tail average wrong?
Why does the fair version of that test use 6.25 days?
The comparison above is stacked, and that has to be said before somebody else says it. The seven exceptions are the days that went past a 99 per cent measure. The tail estimate of Rs 21.9 crore sits at 97.5 per cent, a cut-off covering the worst 2.5 per cent of days. The two sets of days are not the same, so averaging seven of them and comparing the result to a 97.5 per cent estimate is comparing two slightly different things and calling the closeness a triumph.
Run it properly. Two and a half per cent of 250 days is 6.25 days, and that is a fractional cut-offThe arithmetic of averaging a non-whole number of days, needed here because 2.5 per cent of 250 days is 6.25 days rather than a whole one.: a quarter of a day cannot be averaged by ignoring the quarter. Sorted from worst down, the worst six of the seven recorded losses are Rs 28.2, 24.6, 21.0, 20.4, 19.4 and 18.6 crore, summing to Rs 132.2 crore. The seventh, Rs 17.8 crore, is counted at a weight of 0.25, adding Rs 4.45 crore. Total Rs 136.65 crore over 6.25 days is Rs 21.86 crore against an estimate of Rs 21.9 crore, a difference of Rs 0.04 crore, being 0.2 per cent.
Notice what the honest version did to the argument. The honest version made the argument stronger, not weaker. And notice equally that Rs 21.86 crore against Rs 21.9 crore is still not an identity. The seven recorded days are not the same object as the worst 6.25 days of the year. The distinction is not pedantry, and the reason lies in what the seven recorded days actually are.
Why does the fair version of that comparison use 6.25 days rather than seven?
Are those seven the year's seven largest losses?
No, and nothing in the case proves they are. The question matters because every figure above rests on the seven, and the seven are not a list of the year's worst days. The seven are a list of the days on which the realised loss went past that morning's measure. The two lists are different, and the difference is the morning measure. The morning measure moves.
Watch it move. Across the seven recorded exceptions the morning measure ranged from Rs 14.8 crore to Rs 16.2 crore. Now turn to breach B5, a different event in the same year. On month 3 days 21 and 22 the measured value at risk stood at Rs 19.2 crore and Rs 18.6 crore, so it broke the bank's Rs 18.0 crore cap on those two days. The measure that morning stood at Rs 19.2 crore. A realised loss of, say, Rs 18.5 crore on month 3 day 21 would have been larger than five of the seven recorded exceptions and would not have been recorded as an exception at all.
The case records no realised loss for any of the other 243 days, so nobody may claim the seven were the year's seven largest, and any figure built on them inherits that limit. The exception countThe number of days a realised loss went past the measure, which this guide shows to be completely insensitive to how far past it went. is a count of a relationship between two numbers, not a ranking of losses, and reading it as a ranking is a mistake that survives every improvement that can be made to either measure.
Are the seven recorded losses the seven largest losses of the year?
What does a cap do if the measure under it is switched?
Now put the distinction where it costs money. The bank's limit L5 caps its traded measure at Rs 18.0 crore. The reported figure is Rs 15.6 crore, so utilisation is 15.6 over 18.0, being 86.7 per cent, and the limit is within. Suppose somebody proposes moving to the tail measure, on the entirely reasonable ground that it describes something the threshold does not. Put Rs 21.9 crore under the same Rs 18.0 crore cap and utilisation reads 121.7 per cent. A breach appears on a book where not one position has moved and not one price has changed.
The reason is that a limit was never only a cap. A limit is a cap plus a measurement basisWhat a cap is measured on, being the second half of every limit and the half nobody quotes when they quote the limit., and the basis is the half that never gets quoted in the corridor. People say the limit is Rs 18.0 crore. Nobody says the limit is Rs 18.0 crore of one day 99 per cent historical simulation value at risk on the held for trading book. Yet the full description is what the limit actually is. Changing the second half changes the control, whatever the first half still says.
The bank's cap is Rs 18.0 crore against a value at risk of Rs 15.6 crore. What does the same cap read against its expected shortfall of Rs 21.9 crore?
Is the ratio between the two one number?
The obvious repair to the last problem is to scale the cap. Rs 21.9 crore over Rs 15.6 crore is 1.4038, so Rs 18.0 crore multiplied by 1.4038 is Rs 25.27 crore, and utilisation drops back to 86.7 per cent. Rescaling the cap looks like a unit conversion, the way feet become metres. The ratio is not a conversion factor, and this bank contains the counter-example in its own records.
The Rs 15.6 crore comes from historical simulation, the method the bank's policy records. Its two other method results on the same book on the same day are Rs 13.2 crore on variance covariance and Rs 16.2 crore on Monte Carlo. Divide the same Rs 21.9 crore by each and the ratio reads 1.66, 1.40 and 1.35. The three ratios spread 22.7 per cent between the highest and the lowest, produced by nothing whatever except the method underneath the denominator. A cap moved by a ratio observed once is a cap set by yesterday's tail on yesterday's method, and the number that was supposed to describe the shape of the tail moves by nearly a quarter without the tail changing at all.
A paper quotes a tail ratio of 1.40 times for this bank. What has to be printed beside it?
What happens to each measure if every bad day gets worse?
Everything above has been argued in words and in one year's arithmetic. The control below makes the year worse by degrees and shows how the two measures react. No table can do that. Its value lies entirely in whether its answer comes as a surprise, so an answer committed to in advance is worth more than one read off the screen.
This bank's seven exceptions average Rs 21.43 crore. Before the control is moved: if every one of those seven days had been twice as far past its morning measure, what happens to the count of exceptions?
Deepen every recorded loss at once and watch only one of the two measures move
One control: a tail depth multiplier applied to the excess of each of the seven recorded losses over the measure taken that morning, from 0 to 5.00 in steps of 0.05. The seven morning measures are Rs 14.8, 15.2, 15.4, 15.0, 15.6, 16.2 and 15.8 crore, summing to Rs 108.0 crore and averaging Rs 15.43 crore. The seven excesses are Rs 4.6, 9.4, 5.6, 2.8, 12.6, 4.2 and 2.8 crore, summing to exactly Rs 42.0 crore of excess and averaging exactly Rs 6.0 crore. So the average of the seven is Rs 15.43 crore plus Rs 6.00 crore for every whole unit of depth, exactly. The count of exceptions is seven at every setting above zero. Solved: a depth of 0 gives Rs 15.43 crore and a count of zero, 0.50 gives Rs 18.43 crore, 1.00 gives Rs 21.43 crore, 1.50 gives Rs 24.43 crore, 2.00 gives Rs 27.43 crore, 3.00 gives Rs 33.43 crore and 5.00 gives Rs 45.43 crore, with a count of seven at every one of those. The point at which the average equals the bank's own Rs 21.9 crore estimate is a depth of 1.0786, and no step lands on it: at 1.05 the average reads Rs 21.73 crore and at 1.10 it reads Rs 22.03 crore.
At a tail depth of 1.00, the seven recorded days average Rs 21.43 crore and the count of exceptions is 7, against an estimate of Rs 21.9 crore and an expectation of about 2.5 exceptions.
What does neither of them say?
Somebody who has followed the argument this far will reach the reasonable conclusion that moving to the tail measure has solved the problem of the days beyond the cut-off. The move has not solved it, and three things are still missing. The first is the worst single day. An average of Rs 21.43 crore is exactly as consistent with seven days all sitting near Rs 21 crore as it is with six modest days and one catastrophe. The invented bank's own seven range from Rs 17.8 crore to Rs 28.2 crore, a spread of Rs 10.4 crore hidden inside one number.
The second is the shape past the average. An average is a single point in a region that has a shape, and nothing about that shape survives the averaging. The third is the largest of the three, and it applies to both measures equally: both are built from a stated window of past days, so tail depthHow far the losses beyond the cut-off run, which is the only thing in this guide that one of the two measures can see and the other cannot. that the window never contained is a depth neither measure can reach. Lowenstein, in When Genius Failed, published in 2000, is the standing account of what a modelled tail does when the world produces a day outside its own history, and no average over the days a window contains can reach a day it does not. Moving from a threshold to a tail average buys the size of the recorded bad days and buys nothing whatever about the ones that were never in the record.
Somebody argues that moving to the tail measure has solved the problem of the days beyond the cut-off. What is still missing?
Reading the larger figure as the smaller one with the dial turned up
The failure has three layers and each one survives the repair applied to the one before it. The first layer is the switch itself. Somebody moves the reported measure from Rs 15.6 crore to Rs 21.9 crore and leaves limit L5 at Rs 18.0 crore because the cap was set by the board and the measure is a technical matter. Utilisation goes from 86.7 per cent to 121.7 per cent overnight and the breach log gains an entry describing a book on which nothing happened.
The second layer is the fix. Scale the cap by 1.4038 to Rs 25.27 crore, utilisation returns to 86.7 per cent, and everybody agrees the translation has been done properly. The translation has not been done: 1.4038 is this bank's ratio today on the method its policy records, and the same Rs 21.9 crore over its other two method results gives 1.66 and 1.35. The rescaled cap is not a conversion, it is a photograph of one morning.
The third layer survives both fixes and is the one worth carrying away. Neither measure describes the worst day. One is a point with nothing behind it and the other is an average. Half the story therefore sits in what the average was taken over. The bank's Rs 21.43 crore is built from seven days between Rs 17.8 crore and Rs 28.2 crore, and the same Rs 21.43 crore would come out of a far uglier set. A committee that has replaced one number with a larger number has not replaced a question with a better one unless somebody also changed what gets shown alongside it.
Who actually picks up this distinction, and what do they do with it?
Four different people meet these two figures and use them for four different things, and the distinction is worth exactly what it saves each of them.
The independent member of a board risk committee reads a limit paper that says the traded measure is Rs 15.6 crore against a cap of Rs 18.0 crore, and the useful question is not whether 86.7 per cent is comfortable. The useful question is what the paper says about the days past the cut-off. The number in front of that member is silent about every one of them. If the paper carries only a threshold, the committee has been shown where the bad days start and nothing about what they cost.
The analyst at another institution, looking at this bank as a counterparty rather than as an employer, reads whatever the bank discloses and has to work out what the disclosed number is a statement about. A threshold and an average beyond a threshold are not comparable across two banks that chose different cut-offs, and the arithmetic in this guide gives the analyst the exact reason: the same tail figure over three method results moved by 22.7 per cent at one bank on one day, so a difference between two banks' ratios may be a difference of method and nothing else.
The person running the model validation function reads the same seven days and asks a further question: whether seven exceptions in 250 days is more than chance would produce, and whether two consecutive pairs are more clustering than chance would produce. Both questions are formal tests, and formal tests belong to model validation.
The mechanism does not care about scale, so here is the household version. A person who keeps a note of every month the household spent more than it earned has a count. A person who also writes down how much more has a size. The count gives how often the budget failed; only the size gives whether the failures were the price of a wedding or the price of a hospital. Anybody who has ever budgeted has met both of these measures without knowing their names, and the one that hurts is always the second.
Which of the two should carry the cap?
The honest answer is that this is a decision about what an institution wants controlled, not a contest that one measure wins. Capping the threshold measure controls how often the book is expected to go past a line. Frequency is easy to test after the fact by counting. Capping the tail average controls how expensive the bad days are allowed to be on average. A year gives very few observations of severity, so severity is much harder to test.
Notice what that costs each way. The bank's counting measure was out by 180 per cent on the year and its tail average was out by 2.2 per cent, and the record argues one way. Its counting measure can be checked against a year of days by anybody with the record, and its tail average essentially cannot. Testability argues the other way. An institution that has thought about this can say which of frequency and severity it is controlling and why; an institution that has not simply has whichever measure its system happened to produce first, and calls it a policy.
Where these standards come from, and what actually binds an Indian bank
A confidence level, a window length and a cap are choices somebody made, and two banks can make them differently without either bank being wrong. The 99 per cent, the 97.5 per cent, the 250 day window, the 500 day historical simulation, the Rs 18.0 crore cap in limit L5 and every rupee figure above are Vindhya Commercial Bank Limited's own choices.
The market risk framework in which both of these measures sit is published by the Basel Committee on Banking Supervision through the Bank for International Settlements at bis.org, and that is where the origin of the standard, including the treatment of a tail measure inside a capital framework, should be read.
The Reserve Bank of India at rbi.org.in sets what an Indian bank must actually compute, at what cut-off, on which book, how often, what it must report and what it must hold against the result, and it is the only place to confirm any of it. Any confidence level, multiplier, cap, band or effective date appearing above is the invented bank's own working number rather than a requirement.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | What actually binds an Indian bank on market risk measurement, which book a position sits in, what must be computed and reported and what must be held against it | rbi.org.in |
| Bank for International Settlements | The Basel Committee market risk framework in which both a threshold measure and a tail measure sit, cited as the origin of the standard | bis.org |
| Artzner, Delbaen, Eber and Heath | Coherent Measures of Risk, 1999, for the four properties a risk measure is tested against and for subadditivity in particular | ssrn.com |
| Lowenstein | When Genius Failed, 2000, for what a modelled tail does when the world produces a day outside its own history | published book, no site |
Vindhya Commercial Bank Limited is invented.
Educational material. Not advice on any investment, tax, budget or market position.
