Risk Maturity: How Developed a Risk Function Actually Is
Risk maturity is a claim about how developed a whole risk function is, and it is tested against that function's own records rather than a questionnaire. At Vindhya Commercial Bank Limited, invented, seven readings numbered M1 to M7 run from a register that reconciles at 100.0 per cent down to a risk data dictionary complete on only 28.6 per cent. One function, seven different stages.
One awkwardness runs through this subject, and it is worth carrying from the first line. A maturity claim is made about a whole function. A risk function is not one object that was built on one day. A risk function is a register somebody rebuilt last year, a limit set that grew one limit at a time as each new worry arrived, a model list that began life as a spreadsheet, a data dictionary nobody has ever been given a quarter to finish, and a dashboard assembled out of whatever the bank was already counting. The parts of a risk function were built at different times by different people for different reasons, and no mechanism exists anywhere that would bring them to the same stage together. So the honest output of a maturity assessment is a set of readings, and the tempting output is a score.
What is risk maturity a claim about?
The shape is clearer away from banking. Consider a household that has its money in reasonable order. The health cover is renewed on the day it falls due, every year, without a reminder. The rent agreement is filed in a folder with the electricity connection papers. No will exists, none ever has, and what happens to the flat has been raised twice at dinner and dropped twice. Now a question somebody might reasonably ask about that household: on a scale of one to five, how organised are they? Every answer is wrong. Two of those three parts are finished and one has not been started, and no number in the middle describes either of them.
Risk maturityA claim about how developed a whole risk function is, as against how developed one part of it is. is that question asked about a risk function instead of a household. The claim is about the whole. The parts sit at different stages, and the whole has no stage of its own. Somebody genuinely needs to know whether the function that produces the numbers is capable of producing them reliably, so the claim is worth making anyway. A board is entitled to ask it. A supervisor is entitled to ask it. Somebody joining the function is entitled to ask it. None of them is entitled to a single digit that answers it, and what to send instead is set out below.
Where does the evidence come from, if not from a questionnaire?
The usual instrument is a questionnaire. Somebody in the risk function is asked to rate the function's practices against a set of descriptions, and the ratings are averaged. There is a structural problem with that instrument, and it is not that people lie. The problem is that the questions are answered by the people whose work is being rated, from memory, about practices they designed, using words like embedded and consistent that nobody has defined. The answer is an opinion about a process, when the thing actually worth knowing was whether the process leaves a trace.
So take the evidence from somewhere else. Every reading in this guide is a count the bank already holds in a record it already keeps, computed at month 12, with nobody asked to rate anything. A maturity readingOne measurable statement about one part of a risk function, computed from a record rather than rated by a person. is one measurable statement about one part of the function: how many of a thing exist, and how many of them are in the state the bank says they should be in. Two counts, one ratio, and a record that already exists to be checked against. The method stops there, and it is the same method the culture readings used, applied to the machinery rather than to behaviour.
Vindhya Commercial Bank Limited is a mid-sized Indian commercial bank with a balance sheet of Rs 96,000 crore, and the case runs over twelve numbered months with month 12 as the reporting date. Seven readings come out of its records, numbered M1 to M7. Three of them describe things the function does well. Four describe things it does badly. All seven describe the same function on the same day.
What in this bank's records says the risk function is developed?
Three readings say something good, and the first of them is genuinely rare. M1 asks whether the risk register agrees with the other records the bank keeps. The register at month 12 carries 46 entries scored on the bank's own five by five matrix: 4 red, 12 amber and 30 green, and 4 plus 12 plus 30 is 46. The test is applied to the four reds, numbered RR1 to RR4, and it is a simple one. A risk that has actually crystallised leaves a mark in a log, so a red entry should be visible somewhere else in the bank.
All four tie out. RR1, sector concentration, is breach B1 in the breach log. RR2, dependence on wholesale funding, is breach B3 in the same log. RR3, the collateral valuation control, is incident I10 in the operational loss log and the single most serious control finding of the year. RR4, the behavioural deposit assumption, is model V1 in the model inventory. Four of four, being 100.0 per cent. A register that reconciles with the breach log, the loss log and the model inventory is the one reading in this guide where the bank is genuinely finished.
Why is that rare? Because a register drifts by standing still. The breach log, the loss log and the model inventory are all updated by the events that hit them, week by week, by the people those events land on. The register is updated when somebody sits down to update it. So the ordinary finding, in most functions, is a register that is six months behind three other records, still showing a risk as amber that the loss log priced at Rs 15.4 crore in month 8. The register at Vindhya agrees with the other records for one reason: somebody rebuilt it in month 6, entry by entry, against those records. The 100.0 per cent is not the resting state of a register, it is the state of a register six months after somebody rebuilt it.
Vindhya's register agrees with its breach log, its loss log and its model inventory. Is that the normal condition of a register?
M2, the model inventory: 28 registered against 33 found running
The second good reading is about the model inventoryThe register of models an institution says it runs, listing each one with its purpose, its tier and its validation status.. At month 12 this bank's inventory carries 28 models, tiered by its own materiality test: 6 in tier 1, 13 in tier 2 and 9 in tier 3, and 6 plus 13 plus 9 is 28. An inventory is a list somebody typed, so the number 28 by itself says nothing at all. The instrument that says something is the sweepA search for what is actually running in an institution, as against what somebody wrote down as running.: a search of the bank at month 12 for anything that takes inputs, applies a method and produces a number somebody relies on. The sweep found 33. So 28 over 33 is 84.8 per cent, and 5 models were running with nobody having registered them.
An inventory is a claim about what is running, and only a sweep tests the claim. Nobody hid the five, and the absence of concealment is the point worth sitting with. Somebody built a scorecard in a spreadsheet because a decision needed making on a Tuesday, it worked, it stayed, and the registration form was never the next thing anybody did. The gap between 28 and 33 is not evidence of concealment, it is evidence of what happens to any list that is maintained by people volunteering to be on it. At 84.8 per cent the inventory is one of the better readings in this function, and the four still to come are worse.
The model inventory lists 28 models. A sweep found 33 in use. What kind of statement is an inventory, then?
M3, validation currency: 19 of 28, and three that have never been done
The third reading looks inside the 28 that are registered and asks a different question. Registration is administrative. Model validationAn independent check that a model is fit for the use it is put to, covering its data, its assumptions, its build and its stated limits. is the substantive check: somebody who did not build the model examines whether it is fit for the use it is being put to. Vindhya runs a twelve month validation cycle of its own choosing, and against that cycle 19 of the 28 are validated and current, 6 are overdue, and 3 have never been validated at all. 19 plus 6 plus 3 is 28, and 19 over 28 is 67.9 per cent.
Notice that 6 and 3 are different failures wearing the same colour. An overdue validation is a queue problem: the work has been done before, it is late, and a name and a date will clear it. A model that has never been validated is not in the queue at all. The model predates the validation cycle itself. Three of the twenty eight, being 10.7 per cent, are in that second state, and they are numbered V1, V2 and V3.
V3 is a spreadsheet scorecard maintained by one person with no written specification of what it does. V2 sets the haircut applied to collateral. V1 sets the behavioural life the bank assumes for the Rs 36,000 crore sitting in current and savings accounts. Those accounts are contractually repayable on demand, and somebody has to slot them somewhere. V1 decides the sign of this bank's headline interest rate risk number: at the assumption in force the answer is minus Rs 840 crore, and at the deposit life the bank's own repricing ladder implies it is plus Rs 240 crore, and that model has never been validated. How the interest rate arithmetic works is covered separately. One fact about maturity matters here: the least examined model in the inventory is the one the largest number in the risk report turns on.
What in the same records says it is not?
Four readings say something less comfortable, and the first two are the same structure examined from opposite ends. The board at Vindhya has set eight appetite clausesOne sentence of what a board says it is willing to accept, written in the risk appetite statement it approves., numbered A1 to A8: capital, earnings volatility, counterparty concentration, sector concentration, asset quality, liquidity, operational loss and conduct. Underneath them sit twelve limits, numbered L1 to L12, each with a number and a utilisation somebody reports every month. The arrangement between the two is a cascadeAn appetite clause turning into a limit somebody reports against, so a sentence the board approved becomes a number a person watches.: what the board said it was willing to accept becomes a number somebody watches between board meetings. How a cascade is supposed to be built is covered separately, and only two counts from it are used here.
M4 looks downward. Of the eight clauses, four reach a limit: A2 into L8, A3 into L2, A4 into L3 and A7 into L11. Four reach nothing at all: A1 on capital, A5 on asset quality, A6 on liquidity and A8 on conduct. So 4 over 8, being 50.0 per cent. M5 looks upward from the other end. Of the twelve limits, eight have no clause above them: L1, L4, L5, L6, L7, L9, L10 and L12 came from somewhere other than the appetite statement, mostly from an incident, a committee decision or a previous supervisor's question. So 4 over 12, being 33.3 per cent. The framework is broken at both ends, and checking either direction alone would have found only half of it.
The everyday version is a house with a rule and a house with a habit. The rule is what the household agreed at the start of the year: no more than a fixed sum on eating out each month. The habit is the four things somebody actually watches. One of them is the electricity bill, and nobody ever agreed anything about that bill. Both are real. The rule and the habit are simply not connected, and a stranger asked whether that household controls its spending would answer differently depending on which end they were shown.
Four of the eight appetite clauses reach a limit. Is checking that enough to say whether the framework works?
M6, the indicators: 5 of the 16 move before the loss
The dashboard that goes to committee carries 16 key risk indicators, at month 12 showing 9 green, 5 amber and 2 red, and 9 plus 5 plus 2 is 16. M6 asks a question about them that has nothing to do with their colour. The question is when each one moves. Eleven of the sixteen are lagging counts: losses booked, breaches recorded, issues overdue, complaints received. Every one of those changes after the thing has already happened. Five are genuinely leading: staff attrition in the dealing room, the age profile of open access rights, the share of manual journal entries at close, the proportion of exceptions approved by the same person who raised them, and the certificate of deposit roll rate. So 5 over 16, being 31.3 per cent.
A leading indicatorOne that moves before the loss rather than counting it afterwards, so there is still time to act when it changes. and a lagging count are two different objects doing two different jobs, and a dashboard that mixes them without saying which is which lets the reader believe the whole dashboard is a warning. It is not. Two thirds of this dashboard is a history of the quarter that has already closed, and the test of that is in the loss log: of the 13 incidents in the year, 4 were preceded by an amber or red indicator, being 30.8 per cent, so 9 arrived with the dashboard showing nothing at all.
What separates an indicator that moves before a loss from one that counts losses?
M7, the risk data dictionary: 42 of the 147 elements are complete
The last reading is the lowest and it is the one that has already cost money. The monthly risk report is fed by 147 data elements. The bank's own dictionary says each element should carry eight things, numbered T1 to T8: the definition, the source system, the named owner, the named steward, the permitted values and format, the refresh frequency, what happens when the value is absent, and the lineage from source to the report line it feeds. All eight are present for 42 of the 147, being 28.6 per cent.
The interesting part is which data attributeOne of the eight things a risk data element should carry, from its definition through to what happens when it does not arrive. is missing. T3, a named owner, is present for 103 of the 147. T8, lineage, is missing for 71. T7, what happens when the value is absent, is missing for 105, being 71.4 per cent, and 147 less 105 is 42. The complete elements are exactly the ones carrying T7. The attribute this bank most often leaves out is the only one of the eight that describes what to do on a day when something goes wrong.
The single score, and the two readings it buries
Here is what goes wrong, and it goes wrong at the reporting step rather than the measuring step. The seven readings are computed honestly and then compressed into one number, and the function is reported as 3 out of 5, or as 56.6 per cent, or as developing. Everybody files it. Nobody reads further. Both ends of the range disappear into the middle, and the two ends are the only parts of this assessment that anybody could act on.
At the top end, M1 at 100.0 per cent is a register that reconciles with three other records. Reconciliation like that is uncommon, and a rebuild in month 6 earned it. Compressed into a middling score, that work is invisible, and the next person to propose a rebuild has no evidence that the last one paid for itself.
At the bottom end, M7 at 28.6 per cent is the risk data dictionary, and it is the reading that produced a loss. The collateral valuation feed carried T1 to T6 and T8 and did not carry T7. In month 10 the feed stopped arriving. Nothing was defined to happen, so nothing happened, for 11 working days, and 340 loans stayed marked at values that were no longer current. The stale marking is incident I10 at Rs 1.4 crore net, and it is the bank's single material weakness, sitting on the valuation of Rs 8,640 crore of secured advances. A single score would have reported both of those readings as one middling number and pointed nobody at the feed.
A data element carries seven of its eight attributes. The one missing is what happens when the value is absent. What went wrong on the day the feed stopped?
What do the seven readings look like side by side?
Placed on one axis, the argument stops being an argument. The seven readings are 100.0, 84.8, 67.9, 50.0, 33.3, 31.3 and 28.6 per cent. The seven describe one risk function, at one bank, on one date, and each of them was computed from a record that function keeps. The distance from the top reading to the bottom one is 71.4 percentage points, and no summary of the seven can be honest and short at the same time.
| Reading | What it counts | Count | Per cent |
|---|---|---|---|
| M1 register reconciles | reds RR1 to RR4 that tie to another record | 4 of 4 | 100.0 |
| M2 inventory coverage | models registered against models found running | 28 of 33 | 84.8 |
| M3 validation currency | registered models validated and current | 19 of 28 | 67.9 |
| M4 cascade downward | appetite clauses that reach a limit | 4 of 8 | 50.0 |
| M5 cascade upward | limits with an appetite clause above them | 4 of 12 | 33.3 |
| M6 leading indicators | indicators that move before the loss | 5 of 16 | 31.3 |
| M7 data dictionary | elements carrying all eight attributes T1 to T8 | 42 of 147 | 28.6 |
| Seven readings | sum and average of the column | sum 395.9 | 56.6 |
Can one score describe a whole risk function?
Take the average and see what it does. The seven readings sum to 395.9, and 395.9 divided by 7 is 56.6 per cent. The 56.6 per cent now exists, and it is arithmetically correct, and it describes almost nothing. Work out how far each reading sits from it: 43.4, 28.2, 11.3, 6.6, 23.3, 25.3 and 28.0 percentage points. Only one of the seven readings, M4 at 50.0 per cent, is within 10 percentage points of the average of the seven. The other six are all further away than that, and two of them are more than 28 points away in opposite directions.
The seven readings average 56.6 per cent. How many of them are within 10 percentage points of that average?
The average might simply be the wrong summary, with some cleverer single number fitting better. Testing for a better single number is the natural next thought, and the test turns out to be answerable exactly. The question is which single score, anywhere from 0 to 100, sits closest to all seven readings at once, measuring closeness as the plain average of the seven gaps. The answer is the middle reading of the seven, M4 at 50.0 per cent, and at that score the average gap is 22.8 percentage points. No score anywhere on the scale gets the average gap below 22.8 points, so the best single number available is still, on average, more than a fifth of the whole scale away from each reading it claims to describe.
One thing about this function has to be reported to somebody who will read nothing else. What should it be?
Seven readings run from 100.0 per cent down to 28.6 per cent. Before the control below is moved: how close can the best single score get to all seven at once, on average?
Pick one score for the whole function, and watch the seven distances refuse to close
One control: a single maturity score for the whole risk function, anywhere from 0 to 100 per cent. One consequence: how far that score sits from each of the seven readings, drawn as seven lines whose lengths change, and averaged into one number underneath. The default is 50.0 per cent, the best single score available, and at that setting the average distance is 22.8 percentage points. The control moves anywhere on the scale. The static readings are these: at 0 per cent the average distance is 56.6 points, at 50.0 per cent it is 22.8 points, at 56.6 per cent, the average of the seven readings, it is 23.7 points, and at 100 per cent it is 43.4 points. 22.8 points is the floor and nothing on the scale beats it.
a single score of 50.0 per cent
At a single score of 50.0 per cent the average distance from each of the seven readings is 22.8 percentage points, and that is as close as any single number gets.
What would it take to move one of the readings?
Anybody who reads a maturity assessment asks next what it would take to move a reading, and the answer is unwelcome. Take M3, validation currency, at 67.9 per cent. The obvious lever is a policy, and this bank already pulled it: PL6, the model risk policy, was approved in month 4, and it sets out what has to be validated, by whom and how often. So why are three models still sitting outside it at month 12?
Because V1, V2 and V3 all existed before month 4. A policy binds forward and inherits everything that came before it, so approving a policy is the beginning of the work rather than the end of it. From month 4 onward, every new model gets registered and enters the cycle. The three that predate it move only when a person sits down and validates them. V1 needs validating most and is the hardest of the three: the outcome it predicts, how long a deposit actually stays, is only observable over years and cannot be tested against last year's outcomes at all. The reading moves at the speed of the backlog, not at the speed of the approval.
The model risk policy was approved in month 4 and three unvalidated models already existed. What does that say about how fast a maturity reading moves?
Who actually uses a reading like this, and for what
Three people use these readings, and none of them uses the score. Every one of them is deciding where to spend a scarce week, and a set of readings tells them that while a single number cannot.
The head of internal audit is building next year's coverage plan and has a fixed number of weeks to allocate. M1 at 100.0 per cent tells him the register can be relied on as a starting point for scoping. Trusting it saves him a fortnight of reconciliation. M7 at 28.6 per cent tells him where the next incident is likely to come from. He plans towards the low reading and away from the high one, and a score of 56.6 per cent would have told him neither thing.
A supervisor reading the bank's own submissions is deciding a different question: how much weight to put on the numbers the bank sends in. M3 at 67.9 per cent, with three models never validated and one of them setting an assumption a headline number turns on, is directly relevant to that question. So is M2. A number produced by an unregistered model is a number nobody has reviewed.
Somebody joining the risk function in month 13 is deciding what to do first. The readings answer that in about four minutes. The score answers it not at all, and would have cost her a quarter of finding out by hand what M5 and M7 already state side by side.
Which bodies set the rules, and what has to be confirmed at source
No regulator sets a maturity ratio, a minimum, a threshold or an effective date for any of these seven readings. Every reading is the bank's own arithmetic against its own internal expectations, and the twelve month validation cycle is one the bank chose for itself.
Where the underlying subjects are regulated, the bodies are these and they are named rather than quoted. The international standards on model use, data aggregation and risk reporting come from the Basel Committee at the Bank for International Settlements, bis.org. The rules that actually bind an Indian bank come from the Reserve Bank of India at rbi.org.in, and naming only the global standard is the error worth avoiding. The duty on a board and an auditor in respect of internal financial controls sits with the Ministry of Corporate Affairs at mca.gov.in, with the assurance standards from the Institute of Chartered Accountants of India at icai.org. Confirm the text, the applicability and the dates at those sources.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | What actually binds an Indian bank on model use, risk data and internal reporting | rbi.org.in |
| Bank for International Settlements | The Basel Committee standards behind risk data aggregation and risk reporting | bis.org |
| Ministry of Corporate Affairs | The Companies Act duty on the board and the auditor in respect of internal financial controls | mca.gov.in |
| Institute of Chartered Accountants of India | The assurance standards and guidance behind an independent check | icai.org |
Vindhya Commercial Bank Limited is invented.
Educational material. Not advice on any investment, tax, budget or market position.
