Model Validation: Independent Challenge Before Approval
Validation is somebody who did not build the thing checking the choices behind it before it is approved again. Validation is not testing and it is not monitoring: it runs once, it looks at decisions rather than behaviour, and it cannot detect anything. At one invented bank eleven working days produced seven findings, and the interval arithmetic shows why funding it as a detector fails.
One word, two meanings, and the confusion between them costs firms real money. In statistics, validating a model means checking a fitted thing against data it was not fitted on, and that craft is covered separately. In model risk work the word means something completely different: a review, by a person who did not build the arrangement, of the choices sitting underneath it, carried out before a firm decides again whether the arrangement may keep running. Everything in the second meaning follows from one property that has nothing to do with statistics: the reviewer did not build the thing.
What is an independent review actually challenging?
Start with a school. A teacher sets a question paper, teaches the course, marks the answers and reports the results, and every one of those steps can be done carefully. Now ask a different question. Was the paper a reasonable test of the course? The teacher cannot answer that, and not because of any dishonesty. The teacher chose the questions, so whether those questions were the right ones is something the teacher settled already, without noticing it was being settled. Only a second examiner, who did not set the paper, can ask it at all.
The second examiner is the whole idea. Independent validationA review by somebody who did not build the thing, of the choices behind it, before it is approved again. is the second examiner arriving before a firm decides whether an arrangement may keep running for another year. The decision point has a name in most firms, re-approvalThe point at which a firm decides again whether a deployed arrangement may keep running., and the review exists to put something in front of it other than the building team's own account of itself.
The invented bank used throughout, Sumeru Bank Limited, decides retail loan applications along a chain of nine numbered components, five of them fitted to data and four written out as rules by a person. At month 12, before the annual re-approval, Neelima Rao, in the risk function, reviewed the scoring model. The scoring model is component 6, and it determined the outcome of 5,981 of one month's 8,600 files. The part worth slowing down on is this. When she reviewed the four written rule sets, there was a procedure to read: component 5, the income corroboration rule, is 34 lines of written instruction and she read all 34 end to end in 25 minutes. When she reviewed the fitted component there was no procedure at all. A fitted component does not contain one. So a review of a fitted thing reads four things instead.
The four are the fitting populationThe set of past cases a model was fitted on, and the cases that were left out of it., meaning which past cases the thing was built from and which were left out; the labelThe definition of the outcome a model was fitted to predict, which is a set of choices rather than a fact., meaning what somebody decided to call a bad outcome; the inputs, meaning what it is fed and who is accountable for each field where it comes from; and the behaviour across recent months, read afterwards from the record. Every one of the four is a decision a person made before anything was built. A review of a fitted component is therefore a review of decisions and not a review of a procedure. When behaviour is derived from data rather than written down, readability does not survive at all, and repeatability survives only as something tested rather than something read.
What does this kind of review look at that a test does not?
What does independent mean here, and who qualifies?
Independence sounds like a compliment paid to a person, and it is not. Independence is a structural property: a fact about where somebody sits rather than a fact about how honest or how rigorous they are. A scrupulous reviewer in the wrong position is not independent, and a mediocre one in the right position is. The distinction is worth holding on to. Firms constantly try to buy the scrupulous reviewer when the position is what does the work.
The idea is already familiar from buying a flat. The seller offers a survey. The surveyor may be excellent. But the survey was commissioned by the person who wants the sale to happen, so if the surveyor writes something inconvenient, the cost of writing it falls on the surveyor. The buyer wants a surveyor whose fee does not depend on the answer. IndependenceNot having built it, not being accountable for its performance, and not reporting to whoever is. is three separate conditions, and a reviewer who meets two of them is not independent. They are: the reviewer did not build the thing; the reviewer is not accountable for how it performs; and the reviewer does not report to whoever is.
Watch what each condition removes. Not having built it removes the second examiner problem, and it is the easiest of the three to arrange and the least useful on its own. A person whose numbers depend on the thing working has a stake in the answer, so not being accountable for performance removes the incentive to find nothing. Not reporting to whoever is accountable removes the quiet cost of writing an unwelcome sentence. Neelima Rao met all three: she is in the risk function, she built no part of the chain, and the person accountable for component 6 is Revathi Balan in retail credit, who is not in her reporting line at all.
A reviewer did not build the component but reports to the person accountable for how it performs. Independent?
Model Validation vs Model Monitoring: which of the two can notice a fault while it is happening?
The distinction that everything following rests on is clearest in a building. The electrician who wires a flat tests each circuit before anybody moves in. A smoke alarm sits in the ceiling afterwards and watches continuously. And once a year somebody who did not do the wiring inspects the installation and asks whether the design was reasonable for the load the building actually carries. An annual inspection will not catch a fire and a smoke alarm will never show that the wiring plan was wrong, and no amount of spending on either one turns it into the other.
Testing, monitoring and independent challenge sit in exactly that relationship. Testing runs before anybody is affected and asks whether the thing behaves as intended. Monitoring runs continuously afterwards and asks whether it still does. Independent challenge runs at a point in the calendar and asks whether what the thing was built to do was reasonably defined in the first place. Three questions, three positions in time. The first two both measure against the definition, so only the third can put the definition itself in doubt.
Now put the bank's own record against those three columns. Its monitoring watched 3 of the 9 available signals, and all three were downstream of anything that could go wrong at the front. When an upstream income field changed format in month 8, the change ran for 30 working days before anybody noticed, across about 12,900 files decided in that window, of which 176 moved out of accept and into the referral band. The independent review at month 12 read that entire episode, wrote two findings about the monitoring arrangement, and could not have detected it, for the plain reason that the review happened afterwards. The review was the right arrangement for saying the watching was thin and the wrong arrangement entirely for doing the watching.
Which of the three runs continuously?
What did eleven working days actually consist of?
Numbers like eleven working days are usually reported as a total and never opened. The shape inside the total is the interesting part. At this bank the working day is 7 hours, so eleven working days is 77 hours, and those 77 hours split into six stretches of work. Two days on the label and its four choices. Two on the fitting population and the channel mix. One on the outcome distribution across the groups the bank can construct from what it records. Two on the four written rule sets, line by line. Two on the monitoring arrangement and the six week episode. Two on writing the finding. Two plus two plus one plus two plus two plus two is 11.
The shape of that spread is itself a finding: more time went on what the thing was fitted on and what it was fitted to predict than on anything the thing actually did. Set the two days of writing up aside as overhead and 4 of the 9 remaining days, being 44.4 per cent, went on the fitting population and the label. Close to half the reviewing effort went on two decisions that were made before a single line of the arrangement existed, and neither of which a test could have questioned.
Why did thirty four written lines take twenty five minutes and one fitted component take eleven days?
The same reviewer did both. She read component 5, the income corroboration rule, all 34 lines of it, end to end, in 25 minutes, and formed a view on whether the rule said something sensible. Reviewing component 6 took eleven working days. Eleven working days is 77 hours, or 4,620 minutes. Dividing 4,620 by 25 gives about 185, so the second reading took roughly 185 times as long as the first.
The ratio of 185 to one is a property of what was being reviewed, not of who was doing the reviewing. A written rule hands over its reasoning in order, in words, and the review is an argument about the words. A fitted component hands over nothing to argue with, so the reviewer has to go and find the four things instead: pull the fitting population, reconstruct the label from whatever definitions were recorded, trace each input back to whoever is accountable for it at source, and read months of behaviour after the fact. Every one of those is a separate job of work, and none of them existed as a document waiting to be read. A firm choosing a fitted component over a written rule is buying that ratio, and the price shows up in the review long after the build is finished.
Why did reviewing the fitted component take about 185 times as long as reading the 34 line written rule?
What did the eleven days produce, and where did the seven findings land?
Eleven working days produced seven findings. Two on the label's four choices. One on the channel mix in the fitting population. One on the 9 written lines, out of 126 across the four rule sets, that contradicted another line or could never be reached at all. Two on the monitoring arrangement. One on the absence of any procedure for files that had already been decided inside a suspect window. Two plus one plus one plus two plus one is 7. Seven findings from a competent review of a working arrangement is an ordinary yield, not a scandal, and a review that produces none is usually a review that was not allowed to look.
Notice what is absent from that list: not one of the seven says the chain got a file wrong. That absence is the clearest evidence available that this is a different arrangement from a test. A test produces statements of the form it behaved this way on these cases. The independent review produced statements of the form this choice cannot be supported as it stands. The two kinds of sentence go to different people, and only one of them can be answered by fixing something in the build.
Two of the seven findings landed on the label's four choices. Why would a test never have produced those?
Why did two of the seven land on the label's four choices?
Because the label is where the most consequential decisions hide, and because nothing downstream can question it. Component 6 was fitted on a past window of 3,00,000 applications, of which 2,40,000 were accepted and therefore observable, being 80.0 per cent, and 60,000 were declined and never observed at all, being 20.0 per cent. On top of that population sit four numbered choices: what counts as bad, set at 90 days past due; the observation window, set at 12 months; the population, being accepted applications only; and accounts closed early, excluded. Excluding accounts closed early removes 4,320 cases, being 1.8 per cent. On those definitions 8,160 of the 2,40,000 carry a bad label, being 3.4 per cent, 4,320 carry no label either way, and 2,27,520 are good, being 94.8 per cent. The three add to 2,40,000 and the shares add to 100.0.
Now change one choice and watch the ground move. The bank's own readings for the observation window are these, and every one of them is the same population read differently.
| Observation window chosen | Cases carrying a bad label | As a share of 2,40,000 |
|---|---|---|
| 6 months | 5,040 | 2.1 per cent |
| 12 months, which this bank chose | 8,160 | 3.4 per cent |
| 18 months | 10,080 | 4.2 per cent |
| 24 months | 11,520 | 4.8 per cent |
| Difference between 12 months and 18 | 1,920 more | a rise of 23.5 per cent |
Nothing in that table is a fact about borrowers. The table is four readings of one population under four choices, and the bank picked one of them. A chain can pass every test ever written and still be predicting an outcome that somebody defined in a meeting nobody wrote down, and the 90 days and the 12 months are that bank's own choices rather than any standard. Every test downstream measures against the chosen label, so the label is invisible to all of them by construction. The invisibility of the label is why two of the eleven days and two of the seven findings landed there, and it is the single strongest argument for the whole arrangement existing.
What is a finding, and what happens to it next?
A findingA written statement that a choice cannot be supported as it stands, which somebody then has to answer. is a written object, and its parts are worth being exact about. A firm that produces the object without the parts has produced a document rather than a control. Each one names the choice it is challenging, states why that choice cannot be supported as it stands, goes to the person accountable for the component, and carries a date by which an answer is due.
The third kind of answer, the accepted position set out above, deserves defending. People new to control work read it as a loophole. A firm may look at a challenged choice, decide the choice stands, and write down why, signed by the person accountable. Recording the reason is a complete and honest answer. The choice has now been made deliberately, by a named person, on the record. Before the review it was none of those three things. The value of the review sits in the answering, not in the finding, so a firm that funds eleven working days of reading and names nobody to answer has bought a document and called it a control.
A review produces seven findings and nobody is named to answer them. What has the firm bought?
Why can a review like this never be a detector of faults?
Because of when it happens, and this is arithmetic rather than an opinion about how good anybody is at reviewing. A detectorSomething whose purpose is to notice a fault while it is happening, which a periodic review is not. has to be present while the fault is present. A periodic review is present on the days it runs and absent on every other day, so whether it sits inside a given episode is a question about how the episode falls in the calendar and has nothing whatever to do with the quality of the review.
Take the bank's own episode as the thing to be caught. The episode ran 30 working days, from the format change in month 8 to the flag in month 9. The bank's working year is 240 working days, being 12 months of 20. One review a year means a gap of 240 working days between reviews, and a 30 working day episode dropped at random into that gap sits across a review point 30 times out of 240, being 12.5 per cent of the ways it could fall. Twelve and a half per cent is the whole of what once-a-year coverage is worth against an episode of that length, and no reviewer, however good, moves that figure by a single point.
What does each frequency cost a year, and what does it buy?
State the cost first. A firm commits to the cost before it learns what the cost buys. Each review costs 11 working days of somebody's time whether it finds anything or not, and a cost paid on that footing is a standing costWhat an arrangement costs every year whether or not it finds anything. rather than a cost of finding something. Then read what the money buys in the last column. The last column reports coverage of an episodeThe share of the ways a fault of a given length could fall in time on which a periodic review sits inside it., meaning the share of the ways a 30 working day episode could fall on which a review sits inside it at all.
| Reviews a year | Standing cost | Share of a 240 day working year | At the assumed rate | What it buys: coverage |
|---|---|---|---|---|
| 1, which this bank does | 11 working days | 4.6 per cent | Rs 41,250/- | 12.5 per cent |
| 2 | 22 working days | 9.2 per cent | Rs 82,500/- | 25.0 per cent |
| 4 | 44 working days | 18.3 per cent | Rs 1,65,000/- | 50.0 per cent |
| 8 | 88 working days | 36.7 per cent | Rs 3,30,000/- | all of it |
| 12 | 132 working days | 55.0 per cent | Rs 4,95,000/- | all of it, and no more |
Two rows in that table are the argument. Two reviews a year put a 120 working day gap between them, so 30 of 120 is 25.0 per cent of coverage bought at 22 working days of effort. Monthly reviewing costs 132 of 240 working days, being 55.0 per cent of one person's working year, and at the bank's assumed fully loaded Rs 9,00,000/- a year a post that is about Rs 4,95,000/- a year. Both figures are arithmetic on the bank's own assumed rates rather than a costing the bank performed. And the eighth row is where the argument ends: at eight reviews a year the gap is 30 working days, exactly the length of the episode, so coverage reaches all of it and everything bought after that is pure cost. Coverage rises 12.5 points for each review added and then stops dead; cost rises 4.6 points of a working year for each review added and never stops.
Before the control below is moved: one review a year, against an episode lasting six weeks. What share of the ways that episode could fall does the review sit inside?
Buy more reviews, and watch where the coverage stops
One control: how many reviews a year the firm pays for. Two consequences on one scale, and the cost is stated before the coverage. The default is this bank's actual arrangement, one review a year, costing 11 working days and covering 12.5 per cent of the ways a 30 working day episode could fall. Monthly reviewing costs 132 working days, being 55.0 per cent of one person's working year, and at the assumed Rs 9,00,000/- a post that is about Rs 4,95,000/- a year. Then drop the episode into the year yourself and see whether this particular fall is one of the ones a review sits inside.
Educational illustration. Figures are the invented bank's own and describe one deployment. One review costs 11 working days, the working year is 240 working days and the episode is 30, all taken from one measured case. The assumption the picture turns on: a review counts as covering an episode whenever it begins inside it. The assumption is generous. A review looking at the choices behind a component is not looking for a live fault and may well sit inside an episode without noticing it.
A firm proposes quarterly reviews to catch drift faster. What is wrong with the proposal?
What should a review like this be funded for instead?
Fund it for the thing only it can do. Of the nine monitoring signals available on this chain, the bank watched 3, and signal 1, being the share of each input field arriving in the expected format, would have caught the month 8 change on the day it happened. Signal 1 is available the same day and costs a fraction of 11 working days a year to watch. So the sensible arrangement is not one arrangement watching two grounds badly, it is monitoring watching the behaviour continuously and an independent review reading the choices once, and the review's most valuable output at this bank was two findings saying the monitoring was thin. That is the review doing exactly what it is for: it could not detect the episode, and it could say why nobody detected it.
What can a review of this kind not settle at all?
Four things, and each of them is a limit worth stating plainly. A review cannot say a choice is wrong. When the review says the 90 day definition or the 12 month window cannot be supported as it stands, it is making a statement about the support, not about the choice, and the person accountable may answer by producing support or by recording a reason. Treating findings as verdicts is how firms learn to argue with their reviewers instead of answering them.
A review cannot make anybody answer. The power to compel an answer sits with whoever governs the re-approval, and where nobody has been named, the finding simply sits. A review cannot see a day it did not look at, and the coverage arithmetic above is the whole of that limit. And a review cannot make a fitted component readable. After eleven working days the reviewer knew a great deal about the four things surrounding component 6 and still could not read a procedure inside it. There is not one. A review turns an unexamined choice into an examined one, and that is the entire product; it does not turn a fitted component into a written rule.
The review says the label's choice of ninety days past due cannot be supported as it stands. Has it said the choice is wrong?
How does a lender, an analyst or a household read a claim of independent review?
The claim travels well outside a control function. A firm tells a counterparty that its models are independently validated annually. A lender assessing another lender reads the same sentence. Somebody comparing two service providers is handed it in a sales pack. Four questions with checkable answers sit behind that sentence, and the sentence on its own answers none of them.
Ask who did it, against the three conditions. A reviewer inside the building team meets one condition and reads as independent in a sales pack. Ask what it produced. A review producing no findings is either a review that was not allowed to look or one whose findings were not written down, and seven from eleven working days is what a real one looks like. Ask who answered them and by when. The answering is where the control lives. And ask what is watching between reviews. The phrase annual review is most often used to avoid exactly that question. A household choosing between two lenders cannot audit either of them, but it can notice which one can say what its review found and which one can only say that a review happened.
The error that gets made, and what it costs
The error is substitution, and it is made in good faith by capable people. A firm has a real, independent, annual review that produces real findings, and a thin monitoring arrangement that watches 3 signals of 9. Asked whether it has covered its automated decisioning, it says yes, and points at the review. Two different grounds have each been covered once, and the ground that carries live faults got covered by three downstream signals that could not see the front of the chain at all.
Follow it through on these figures. The episode ran 30 working days and about 12,900 files were decided inside it, of which 176 moved out of accept and into the referral band. The review at month 12 read that episode and wrote two findings about it, roughly three months after the last file in the window had been decided. Buying more of the review would not have changed that by a day. A review at any frequency the firm can afford sits inside such an episode on a share of the ways it can fall and notices nothing on the rest. The input signal that would have caught it was available on the day it happened.
The reverse error costs less but is quieter. The reverse error is to conclude that the review is overhead, and to run monitoring alone. A firm running monitoring alone watches its behaviour beautifully against a bad outcome definition that nobody beyond the people who built it has ever read, and it will go on doing so for years. Every signal it watches measures against that same definition. Neither arrangement is optional and neither is a substitute.
What the reader has to confirm at source
Where a regulated lender uses a model in a lending decision, the expectations about independent review before approval sit with the Reserve Bank of India and are published at rbi.org.in, and where the deployer is a market intermediary the Securities and Exchange Board of India publishes its own at sebi.gov.in. The standing international discipline from which model risk practice descends originates with the Bank for International Settlements at bis.org, an origin rather than the position in India. A firm in India is bound by what the Indian supervisor states, and the current position has to be read at the source.
The annual cycle described here is Sumeru Bank Limited's own arrangement, the eleven working days and the seven findings are that invented bank's own record, and the 90 day and 12 month choices in the label are that bank's own choices rather than any standard.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Published expectations on a regulated lender covering digital lending, outsourcing, data and consent, and what a firm should be able to show about arrangements that decide customer outcomes | rbi.org.in |
| Securities and Exchange Board of India | Published expectations where the deployer of such an arrangement is a market intermediary rather than a lender | sebi.gov.in |
| Bank for International Settlements | International supervisory material from which the standing discipline of model risk work originates, and not the position in India | bis.org |
Sumeru Bank Limited, its intake chain, Revathi Balan and Neelima Rao are invented.
Educational material. Not advice on any investment, tax, budget or market position.
