Risk Score: Compressing Two Judgements Into One Number
A risk score is impact multiplied by likelihood on an institution's own scale, produced so that a register can be sorted and a dashboard can show a colour. The score is a reporting device and not an assessment. On a five by five grid the 25 cells produce only 14 distinct scores, and only 4 of the 25 can be recovered from the number they produce.
A risk score is arithmetic anybody can check with a pencil, and the check is worth making. The multiplication does something to the information that nobody announces when the report is handed over. Two judgements go in. One number comes out. The number is easier to work with in every way that matters to a meeting, and it has quietly stopped being able to say what to do.
What is a risk score, and what is it actually for?
A risk scoreA single number formed from two judgements, usually impact multiplied by likelihood, so that a list can be ordered. is a single number formed from two separate judgements so that a list of risks can be put in an order. Nothing else is going on. The score is not a measurement of anything, it does not come from a model, and it carries no unit. The score exists because a list has to be sortable and two numbers cannot sort a list.
Take it out of a bank for a moment. A household has four jobs waiting and one free Sunday. The roof has a slow leak that would ruin a ceiling if the monsoon caught it, and it has held for two years. A bulb on the stairs fuses about once a month and somebody replaces it in four minutes. The gate latch sticks. The geyser makes an unwelcome sound. All four cannot be done, so they have to be ranked, and ranking them means squashing two thoughts into one: how bad would it be, and how often is it going to happen. Squashing two thoughts into one is a risk score. The leaking roof and the fusing bulb can very easily come out of it with the same rank, and that is the whole problem.
Vindhya Commercial Bank Limited, invented, does the same thing at a balance sheet of Rs 96,000 crore. The bank keeps a record of the risks it has identified, rates each one for how much it would hurt and how probable it is, multiplies the two, and reports the result.
There are exactly three things the resulting number does well, and they are all things a single value can do that a pair cannot. The number can order a list, and 46 entries acquire a top. A score can drive a colour, turning a band of scores into red, amber or green on a dashboard. And a score can carry an escalation levelA score at or above which something is reported upward, which is one of the few genuine uses of the compressed number.. Anything at or above a chosen value goes upward to a committee without anybody having to argue the case again. There is a fourth thing people use it for, deciding what to actually do about an entry, and it is the one thing the number cannot support.
A colleague reads a score off the register and uses it to decide what to do about that entry. What is wrong with that?
How is the number computed, and who decides the scale?
The computation is one multiplication. ImpactHow much it would hurt if the thing happened, rated on the institution's own scale and estimated elsewhere. is rated from 1 to 5. LikelihoodHow probable the thing is over a stated horizon, rated on the institution's own scale and estimated elsewhere. is rated from 1 to 5. Multiplying the two gives a score somewhere between 1 and 25. Both ratings are taken here as given inputs: how each one is arrived at, what a five point scale means and how an entry gets onto a register in the first place are covered separately.
Now the question people almost never ask. Where does five by five come from? Nowhere. No external body prescribes a risk scoring scale, so every scale in use anywhere is the institution's own choice. Five by five is common because an odd number gives a middle option and five rows stay readable at a glance. Three by three, four by four and six by six all exist and all work. Nothing in the arithmetic requires five, and an institution that chose four by four would get a different set of possible scores, a different amount of crowding and a different answer to every question below. Crowding, holes and lost pairs are properties of the scale somebody picked, not properties of risk.
There is an honest technical caution to put beside the multiplication. A rating of 4 for impact is not twice a rating of 2 in any measurable sense. A rating is a rank chosen from a ladder of descriptions, and multiplying two ranks produces a number that looks arithmetical and is really an ordering convention. Frank Knight, in Risk, Uncertainty and Profit, published in 1921, separated the risk that can be measured from the uncertainty that cannot, and a five by five rating sits closer to the second than most reports admit. The score is still useful. The score is useful as a sorting key and not as a quantity.
Where does a five by five scale come from?
A five by five grid has 25 cells. How many different scores can it produce?
How many different scores can twenty five cells actually produce?
Write the grid out and put the score inside every cell. Almost nobody does that, and nothing else anybody can do with a scoring scheme is half as useful. Twenty five cells go in. Fourteen distinct numbers come out: 1, 2, 3, 4, 5, 6, 8, 9, 10, 12, 15, 16, 20 and 25. 7, 11, 13, 14, 17, 18, 19, 21, 22, 23 and 24 are simply not products of two whole numbers from 1 to 5, so eleven of the values between 1 and 25 cannot be produced at all. The grid produces 14 values out of 25, being 56.0 per cent, and 11 that never can, being 44.0 per cent.
Look at what that does to a scale nobody thought to check. A report that speaks of a risk scoring range of 1 to 25 sounds continuous, sounds fine grained, and sounds like it has twenty five settings on the dial. The dial has fourteen settings, with holes scattered all through the upper half. A committee member who asks whether anything on the register scores 13 is asking a question the scheme can never answer yes to, and nothing in the scheme tells them so.
The same finding is set out below as a table, to be looked up rather than taken on trust. Read the middle column as the pairs that produce each score, impact first and likelihood second.
| Score | The pairs that produce it, impact first | Cells | Recoverable from the score alone |
|---|---|---|---|
| 1 | 1 with 1 | 1 | Yes |
| 2 | 1 with 2, 2 with 1 | 2 | No |
| 3 | 1 with 3, 3 with 1 | 2 | No |
| 4 | 1 with 4, 2 with 2, 4 with 1 | 3 | No |
| 5 | 1 with 5, 5 with 1 | 2 | No |
| 6 | 2 with 3, 3 with 2 | 2 | No |
| 8 | 2 with 4, 4 with 2 | 2 | No |
| 9 | 3 with 3 | 1 | Yes |
| 10 | 2 with 5, 5 with 2 | 2 | No |
| 12 | 3 with 4, 4 with 3 | 2 | No |
| 15 | 3 with 5, 5 with 3 | 2 | No |
| 16 | 4 with 4 | 1 | Yes |
| 20 | 4 with 5, 5 with 4 | 2 | No |
| 25 | 5 with 5 | 1 | Yes |
| 14 | Every pair on the grid | 25 | 4 cells, being 16.0 per cent |
Where on the scale does the crowding sit?
Counting the cells behind each score is where the shape of the thing appears. Four scores come from a single cell. Nine scores come from two cells each. One score, and only one, comes from three cells, and that score is 4. Check the total: 4 single cells plus 18 in pairs plus 3 at the score of 4 gives 4 plus 18 plus 3, and that is 25, so every cell is accounted for exactly once.
Where do the four single cell scores sit? At 1, at 9, at 16 and at 25. Two of those are the absolute extremes of the scale, the corner where nothing much is at stake and the corner where everything is. The other two sit high. Every score in the busy part of the range, the 4s and 5s and 6s and 8s and 10s and 12s where the ordinary business of a register actually lives, comes from two or three cells. The compression bites hardest exactly where the population is, and the parts of the scale where the number is trustworthy are the parts almost nothing lands on.
One more number for shape. The average of all 25 scores is 225 divided by 25, being exactly 9.0, and the middle value when all 25 are lined up in order is 8. So the arithmetic centre of the scheme sits at 8 and 9. The value 9 is one of the four that identify their cell, and 8 comes from two cells. Nothing designed that. The centre is what falls out of multiplying two ladders together.
Where on the scale does the compression do the most damage?
Impact 2 with likelihood 2 scores 4. Which other cells give exactly the same score?
Can impact and likelihood be recovered from a score?
The question has an exact answer, and the calculator below is built around it. Where two cells produce a score there are two candidates, and the number offers no way to choose between them. So the pair can be recovered only where exactly one cell produces the score. A recoverable cellA combination of impact and likelihood that can be identified from its score alone, because no other combination produces it. is therefore a cell whose score is 1, 9, 16 or 25 and nothing else.
Four of the 25 cells are recoverable, being 16.0 per cent, and 21 are not, being 84.0 per cent. That second figure needs its object named every time it is used. The 84.0 per cent is the share of grid cells that cannot be identified from their score and nothing else: this same invented bank carries an 84.8 per cent somewhere else in its papers that means a limit utilisation, and another 84.8 per cent that means how complete a register of models is, and the three have no relationship whatsoever beyond looking alike in a report.
Now the part that is easy to get backwards, and worth being careful about. All four recoverable cells sit on the diagonalThe cells where impact equals likelihood, which is where almost all the recoverable scores sit.. Impact equals likelihood there: 1 with 1, 3 with 3, 4 with 4 and 5 with 5. But the diagonal has five cells, not four. The fifth is impact 2 with likelihood 2. 1 times 4 and 4 times 1 also make 4, so that cell is not recoverable. So the rule runs one way and not the other: every recoverable cell is on the diagonal, and one cell on the diagonal is not recoverable.
There is a smaller way to say the same thing that is worth holding on to. Of the 14 distinct scores, 4 identify their cell. Printed as a percentage that is 28.6 per cent, and it is the share of distinct scores that identify a cell and not any of the other things 28.6 per cent stands for elsewhere in this invented bank's papers. Same fraction, different object, and naming the object is the whole discipline.
An entry scores 12. What does that establish about its impact and its likelihood?
Set the pair and watch every other cell that makes the same number
Two controls, and any cell in the grid can also be clicked directly. The slider moves impact, the buttons pick likelihood, and the grid shows the chosen cell in dark and every other cell producing the identical score in green. The default is impact 2 with likelihood 2, scoring 4. The two green cells that appear beside it are the fastest demonstration of what the compression costs. CompressionTurning two values into one, which is what makes sorting possible and what makes the pair unrecoverable. is not a metaphor here; it can be watched happening.
Impact 2 with likelihood 2 scores 4, which is also produced by impact 1 with likelihood 4 and by impact 4 with likelihood 1, so this cell is not recoverable from a score of 4 alone.
What does a score of 4 actually hide?
A score of 4 puts the whole argument on one screen. Three cells make it, and they are three completely different situations that a register orders as equals.
Three entries score 4, and they need three different responses
Impact 4 with likelihood 1 is a rare severe event. The event will probably not happen, and if it does it hurts a great deal. In this bank's own loss record that shape is incident I13, the trade finance fraud in which a member of staff and an outside party issued nine letters of credit against forged shipping documents over fourteen months ending in month 8. The fraud happened once and cost Rs 15.4 crore net, the largest net loss of the year.
Impact 1 with likelihood 4 is a frequent trivial one. Such an event happens all the time and each occurrence is small. The frequent trivial shape is incident I1, card-not-present fraud on the debit card portfolio, many small events across the year adding to Rs 4.8 crore net. This bank's record fixes no impact or likelihood rating for any incident, so the two incidents illustrate the two shapes and carry no rating of their own.
Impact 2 with likelihood 2 is neither, and it is the one nobody argues about at a meeting.
The response each one calls for differs. A rare severe event is an argument for a control that stops it happening at all, or for a reserve, or for insurance. The purchase is protection against a tail that may never be seen. Stopping every instance costs more than the instances do, so a frequent trivial one is an argument for reducing the cost of each occurrence or for accepting it as an ordinary operating expense. The middling one argues for neither in particular. Three responses, one number. A register sorted by score puts all three at the same rank, and a dashboard paints all three the same colour. The multiplication did precisely what it was asked to do. The limit lies in what a sorted list can carry.
Three entries on the register all score 4. What might the three of them be?
What does this bank's own register look like?
The registerThe record of risks an institution has identified, each carrying a rating, an owner and a status. at month 12 carries 46 entries, each rated on the bank's own five by five scale, and the result is 4 red, 12 amber and 30 green. The check: 4 plus 12 plus 30 is 46. As shares that is 8.7 per cent red, 26.1 per cent amber and 65.2 per cent green, and those three rounded figures do sum to 100.0. One honest limitation before anything is read into the shape: this bank's record fixes the colour of each entry and does not fix the score of each entry, so there is no distribution of the 46 across the fourteen scores to show.
The four reds are worth looking at, and not because of the number 4. Every one of the four is already written down somewhere else in this bank's own records. A register that is actually working looks exactly like that. RR1 is sector concentration, and it is also open breach B1. RR2 is the dependence on wholesale funding, and it is also open breach B3. RR3 is the collateral valuation control, and it is also incident I10 and the year's single material weakness. RR4 is the behavioural deposit assumption, and it is also model V1, one of the three models in this bank that have never been validated. A register that agrees with the breach log, the loss log and the model inventory is doing its job; one that disagrees with them is itself the finding.
This bank's register has 46 entries with 4 red. What is notable about those four?
What goes wrong when a register is sorted by score?
Sorting is why a score exists, and sorting is also the moment the information leaves the report. The loss is a design problem and not a mistake anybody made. A committee cannot read 46 entries with equal attention. Members read from the top and stop when the meeting runs out of time, and the sorting was built to support exactly that behaviour. So what rises is anything with a big product, meaning anything rated high on both judgements. Anything rated high on one and low on the other falls, no matter how much of a problem it is. And what becomes invisible is the difference between entries that share a rank.
Think about what that does to the two shapes above. The rare severe event, impact 4 with likelihood 1, sorts at 4 and sits in the lower half of any register. The frequent trivial one, impact 1 with likelihood 4, sorts at exactly the same place. If the committee reads down to the score of 8 and stops, both of them are below the line together, and the only fact that distinguishes them, that one of them is the shape that produced the largest single net loss in this bank's year, never reaches the room. A ranking is trusted and a pile of unsorted entries is not, so a ranking that hides the one thing deciding the response is worse than no ranking at all.
How can the sorting be kept while none of the information is lost?
The fix is so small it is faintly embarrassing, and that is the best possible property for a fix to have. Print the two judgements next to the score. A reported pairShowing impact and likelihood beside the score, which is the cheapest way to keep the sorting and lose none of the information. costs one extra column on a report line and recovers everything the multiplication removed.
Compare the two lines directly. A line reading score 4 tells a reader nothing they can act on. A line reading score 4, impact 4, likelihood 1 tells them it is the rare severe one, and that points at a preventive control or a reserve. The second line still sorts by its first figure exactly as well as the first line does, so the trade between sorting and knowing was never a real trade at all. The trade only looks real if somebody has decided in advance that a report line can carry one number and not three. Nothing about the arithmetic forces that choice, and nothing about a report layout does either.
How can the sorting be kept while none of the information is lost?
What does a committee member actually do with a scored register?
Three habits are worth having, and none of them needs any authority to adopt. The first is to ask, of any entry being discussed, what the pair was. Rating is somebody else's craft and is covered separately, so the question is not a challenge to it. The pair shows whether the conversation should be about prevention or about unit cost. If the answer is not on the paper, that is the finding, and it is a cheap one to fix.
The second is to distrust a comparison between two entries with the same score and to make the meeting say which is which. Two entries at 12 are 3 with 4 and 4 with 3, and those are genuinely different arguments even though the number refuses to distinguish them.
The third is for whoever maintains the record. When a score changes between one meeting and the next, say which of the two judgements moved. A score falling from 12 to 8 could be impact 4 with likelihood 3 becoming impact 4 with likelihood 2, meaning the thing became less probable, or it could be impact 3 with likelihood 4 becoming impact 2 with likelihood 4, meaning the thing became less damaging. The two are opposite stories about the same fall, and a movement column carrying only the product cannot tell them apart. Every score change with no stated cause is a claim nobody has to defend, so an internal auditor reading a register can find real work in the movement column alone. The same discipline serves a lender reading a borrower's own risk record, or anybody reading a supplier's: the score shows where somebody put the entry, and the pair shows what they were worried about.
What is named here, and where the binding version lives
No external body prescribes a risk scoring scale. There is no standard five by five, no standard escalation level and no standard colour boundary. Every scale, rating, colour and count here belongs to the invented Vindhya Commercial Bank Limited.
Where a score feeds a report that goes to a board, what an Indian bank must actually compute, report and place before its board comes from the Reserve Bank of India at rbi.org.in. The principles on risk data aggregation and risk reporting that sit behind how a bank assembles such a report were published by the Basel Committee on Banking Supervision at the Bank for International Settlements at bis.org, and what binds in India is the Reserve Bank of India's version and not the international text.
A threshold, ratio, escalation level, colour boundary or effective date that binds is published by the issuing body itself, and the wording on that body's own site is the wording that binds.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | What an Indian bank must actually compute, report and place before its board on risk management arrangements | rbi.org.in |
| Bank for International Settlements | The Basel Committee principles on risk data aggregation and risk reporting that sit behind how a report is assembled | bis.org |
| Ministry of Corporate Affairs | The Companies Act duty on a board in respect of the risk management policy it reports on, and the form of that report | mca.gov.in |
| Frank Knight | Risk, Uncertainty and Profit, 1921, where measurable risk is separated from the uncertainty that cannot be measured | Houghton Mifflin |
Vindhya Commercial Bank Limited is invented.
Educational material. Not advice on any investment, tax, budget or market position.
