AI Governance vs Model Risk Management: Which Layer Holds What
Model risk work is the older discipline, and its test is what a thing produces: anything supplying a value a decision uses gets an inventory entry, a named accountable person, an approval and a periodic re-approval. The newer layer adds three tests, about how the behaviour was set, where the output lands and whether it was bought. At one invented bank the two lists met on exactly one component of nine.
Two lists exist in most large firms that have deployed anything, and they were built by different people worried about different things. Neither list is a longer version of the other, and neither is a stricter version of the other. The two lists ask different opening questions, so they catch different items, and the place where they meet is usually much smaller than anybody expects. The size of that meeting place settles most of the argument about which of the two a firm needs, and it settles it with counting rather than with opinion.
What is the older discipline of model risk work actually for?
Model risk work is a large discipline on its own. The part of it that draws the boundary fits in one paragraph. A firm that runs numbers into decisions keeps a model inventoryThe list a firm keeps of everything producing a value a decision uses, with an accountable person named against each entry.: a list of everything producing a value a decision uses. Against each entry sits a named accountable personThe named individual who answers for a thing, as distinct from whoever operates it day to day., an approval given before the thing was allowed to run, a record of how it was built and checked, and a re-approvalThe point at which a firm decides again whether something may keep running, rather than letting an old approval stand forever. at a stated interval. The arrangement grew up around pricing, provisioning and capital numbers, and it has one opening question that decides everything else: does a decision use what this thing produces.
The question of what a thing produces is the entire boundary, and everything else in the older discipline comes from taking that question literally. The older discipline is not indifferent to how a thing was built; it asks about that in enormous detail once an item is inside. Construction is simply not the test for getting in. A thing enters because of what it produces, and a thing that produces nothing a decision uses never reaches the part where construction is examined at all.
What did the newer layer add, and why was it added?
Think about a household for a moment. The shape is exactly the same, and nobody needs a bank to see it. A household keeps, without ever calling it a list, a mental record of everything with a warranty card. A warranty answers who pays when something breaks. The same household keeps a second record of everything left running when the house is empty. The second record is about what could burn the place down. The refrigerator is on both. The table lamp is on neither. The water heater is on the second only, and the wristwatch is on the first only. Two honest lists, two different worries, and a small overlap that nobody planned.
The newer layer is the second list. The newer layer was built because components began arriving whose behaviour was not written by anybody, whose outputs reached customers without passing a desk, and which increasingly were not built in the firm at all. A use case registerThe list a firm keeps of where behaviour derived from data is in use, built on different tests from an inventory of things feeding decisions. is the artefact, and its opening questions are about construction and about where an output lands rather than about whether a decision consumes a number. The newer layer exists because three worries had no test that would catch them, and a firm cannot control what its scoping question never picks up.
Which four questions decide where a thing belongs?
A scoping testThe written question that decides whether a thing is inside a policy or outside it, applied before anything else happens. is the written question that decides whether a thing is inside a policy or outside it, and the whole of the boundary between these two layers is four such questions asked in order. Ask them of any component, any spreadsheet, any bought service, and the answers place it. Question 1: does it produce a value a decision uses. Question 2: was its behaviour fitted to data rather than written by somebody. Question 3: does its output reach a customer or a reported figure with no person deciding in between. Question 4: was it bought as a service rather than built.
Running them in that order produces one of four placements. A yes on question 1 puts a thing inside the older discipline. A yes on any of questions 2, 3 or 4 puts it inside the newer layer. A yes on question 1 and a yes on one of the other three puts it inside both layers at once. The double placement needs the most care and gets it least often. And a no on all four means neither layer holds it. The fourth placement is a legitimate answer and worth writing down as one. An item nobody can find later was never really scoped at all.
Which of the four questions is the older discipline's own test?
Why is the first question the older test and the other three the newer layer?
Because of what each question is worried about. Question 1 is worried about a number being wrong in a way that moves money, and that worry does not care how the number was produced. A written formula in a spreadsheet and a component fitted to three hundred thousand past applications are equally capable of feeding a wrong number into a credit decision, so the older discipline treats them identically and asks about construction only once they are inside.
Questions 2, 3 and 4 are worried about three things that the first question cannot reach. Question 2 is worried that nobody can read the behaviour off the thing: a written rule can be read line by line and a fitted component cannot, so an inspection that works on one does not work on the other. Question 3 is worried about consequence without a person: an output that lands on a customer with nobody in between has no step at which a mistake gets caught by somebody noticing. Agrawal, Gans and Goldfarb, in Prediction Machines, 2018, frame a fitted component as supplying a prediction that a person still has to act on, and question 3 asks precisely whether that person exists in the arrangement at all. Question 4 is worried about inspection rights: a component running on somebody else's machines under an agreement can only be examined as far as the agreement allows. None of those three worries is about whether a decision consumes a number, and the older question therefore never picks any of them up.
Where do the nine components of one deployed chain land on the first two questions?
Sumeru Bank Limited, invented, runs a retail loan intake chain built from nine numbered components. Five of them have behaviour fitted to data: the liveness check on the selfie image, the document classifier, the field reading step, the scoring model and the drafting assistant. Four are rule sets somebody wrote: the identity match, the income corroboration rule, the fraud rules on the servicing book and the workflow router. Place all nine against question 1 on one axis and question 2 on the other and the picture is unusually lopsided. At this bank exactly one component produces a value a decision uses.
The empty corner is the surprise. The older discipline is built to hold a written calculation feeding a decision, and this chain does not contain one. Every decision the four written components determine is a routing decision: the income corroboration rule sent 602 files to a person in the month, the fraud rules 452, the identity match 189 and the workflow router 112, a total of 1,355 files, and not one of those 1,355 was a decision on a loan. The four written components decide where a file goes. In this chain the written components never decide what happens to the applicant, and the fitted side does both. The older discipline's opening question therefore caught very little of this chain.
The four at the bottom left are not stranded, and that matters. Question 3 asks whether the output reaches a customer or a reported figure with no person deciding, and on that test all four of them are inside the newer layer. A file routed away from a straight-through answer changes what the applicant experiences whether or not a person eventually looks at it. Only the document classifier, whose output never leaves the system, and the drafting assistant, whose every output is signed by a person, fall outside. Seven of the nine are caught by that third question and two are not.
Which two of the nine components fall outside the question about an output reaching somebody with no person deciding?
Nine components went live in one chain at Sumeru Bank Limited. How many of them were already in the bank's model inventory?
What did this bank's model inventory hold when the chain arrived?
The model inventory at Sumeru Bank Limited predates every part of the intake chain, and it held exactly one of the nine components: number 6, the scoring model. The scoring model was there because it produces a value a credit decision uses, and the risk function had always held things of that kind. There was no oversight, no backlog and no neglect involved. The inventory held precisely what its own test told it to hold, and it applied that test correctly to all nine components, admitting the one that passed and declining the eight that did not.
Sit with that for a moment. The instinct on first reading is that somebody was asleep. Nobody was. An inventory built on what a thing produces will hold one component out of nine in a chain like this one, and it will do so while working perfectly. Perfect operation producing that result is a far more uncomfortable finding than negligence would have been. A control that fails through neglect is fixed by paying attention. A control that misses eight of nine while working exactly as designed can only be fixed by adding a different question, and adding a different question is what the newer layer is.
What did the newer layer's own list hold, and how complete was it on day one?
The bank built a use case register for the newer layer and signed it off with 9 entries. A sweep run by Ashok Pillai in technology risk later found 14 uses actually running, so the register was 64.3 per cent complete on the day it was signed, and 9 against 14 is the bank's own count from its own sweep. Of the 14, only 6 carried a named accountable person, being 42.9 per cent, and 4 of the 14 sat inside a service bought as a serviceA component running on somebody else's machines under an agreement, which changes what the buying firm can inspect and how much. rather than built inside the bank.
Do not read 64.3 per cent as a criticism. A list of things feeding decisions can largely be inherited. The numbers feeding decisions have names and appear in reports. A list of where behaviour derived from data is in use has to be found rather than inherited. Much of it arrives inside something that was bought, and finding it is the work rather than a preliminary to the work. The four bought entries are the clearest case: nothing about them appeared on any technology plan the bank wrote, and the sweep found them by asking what each part of a process was actually doing rather than by reading a list of projects.
How complete was the use case register at Sumeru Bank Limited on the day it was signed off?
How big is the overlap between the two layers, really?
The overlapThe set of things both layers hold, which needs one arrangement rather than two running side by side. at Sumeru Bank Limited is one component. Component 6, the scoring model, sits in the model inventory because a credit decision uses the value it produces, and it sits inside the newer layer because its behaviour was fitted to data and its output reaches an applicant with nobody deciding in between. Nothing else in the chain is on both lists. One of nine components is 11.1 per cent of the chain, and against the wider count of 19 items that the bank's chosen scoping test caught across the whole institution, that same single component is 5.3 per cent.
The overlap is one component of nine. Is that an argument for dropping one of the two layers?
Why is a small overlap the finding rather than a curiosity?
Because the same overlap reads two completely different ways depending on what is counted, and both readings are true at once. Counted in components, the older discipline held 1 of 9, being 11.1 per cent of the chain. Counted in decisions, that one component determined the outcome of 5,981 of the month's 8,600 decided files, being 69.5 per cent. The eight components it did not hold determined the other 2,619, being 30.5 per cent, and 1,264 plus 602 plus 452 plus 189 plus 112 is 2,619 exactly. All of these are Sumeru Bank Limited's own invented figures for one steady month.
A small overlap is evidence that the two layers are doing different work, and a large overlap would have been the argument for merging them. If the inventory and the register had returned the same eight or nine items, one of them would be a second copy of the other and a firm could reasonably run only one. In fact they returned one item in common out of nine. Each list is therefore almost entirely made of things the other never picks up. Drop either list and nearly everything it was holding is lost, and lost silently. A scoping question that does not catch a thing also does not report that it failed to catch it.
The one overlapping component is 11.1 per cent of the components and 69.5 per cent of the month's decisions. Which figure should a board hear?
What does each layer demand that the other does not?
The four questions are not just a placement device; each one carries a demand that follows from it, and the demands do not substitute for one another. Set the four questions below against the layer each belongs to and against what each one actually caught at Sumeru Bank Limited. Every count is the bank's own, drawn from its own scoping exercise and its own sweep, and none of it is a threshold or a requirement of any authority.
| No. | The question | Layer | What it caught at this bank |
|---|---|---|---|
| 1 | Does it produce a value a decision uses? | The older discipline | 1 of the 9 components |
| 2 | Was its behaviour fitted to data rather than written by somebody? | The newer layer | 5 of the 9 components |
| 3 | Does its output reach a customer or a reported figure with no person deciding? | The newer layer | 7 of the 9, and 19 items across the bank |
| 4 | Was it bought as a service rather than built? | The newer layer | 4 of the 14 register entries |
| Components the older discipline's question admitted, of nine | 1 | ||
Read the third row carefully. The third row is where the two counting bases meet, and they are not the same base. Applied across the whole institution, the consequence testScoping by what an output does to somebody rather than by how the thing was built or by what number it produces. the bank chose caught 19 items: 7 of the nine intake chain components plus 12 of the 38 other routing and calculation rules sitting elsewhere in the bank, and 19 less 7 is 12. The register of 14 entries is a different exercise: a sweep of where behaviour derived from data was in use. The consequence test catches written rules as well as fitted ones. Its count of 19 therefore includes a dozen routing and calculation rules that no register of fitted behaviour would ever have listed.
So the demand of the older discipline is that anything feeding a decision has a named person, an approval and a re-approval, whatever it is made of. The demand of the newer layer is that anything whose behaviour was fitted, or whose output lands on somebody with no step in between, or which runs on machines the firm does not control, is written down with the same weight even when the number it produces feeds nothing anybody would call a decision. Neither demand covers the other, and a firm that has one has genuinely not got the other.
A firm says its existing model risk work already covers its AI. What is the specific gap?
What happens where both layers apply: two accountable people, two approvals, two records?
The overlap is the part firms discover late, and it shows up in paperwork rather than in principle. Component 6 at Sumeru Bank Limited sits in the model inventory with Revathi Balan, head of retail credit, named against it, an approval that was given before it ran and an independent challenge carried out by Neelima Rao in the risk function over eleven working days at month 12, before the annual re-approval. The same component sits in the use case register, entered under the use case that was defined and approved at month 0 and reviewed again in the sweep at month 10. Two records, two approval events, two review cycles, one component.
Two named people for one component sounds like belt and braces and is worse than one. When two people are each accountable for a thing, each can reasonably believe the other is the one answering, and the belief is honest on both sides. The fix is not to merge the two layers but to say plainly, in writing and in advance: one approval governs, one review cycle governs, and one record is the one anybody reads for the facts about the component. The other layer's record then points at it rather than restating it, and stops being a second version that can quietly go out of date.
A component sits in both an inventory and a register. What has to be decided?
How can both layers run without running everything twice?
The duplication people fear does not come from running two sets of questions. The duplication comes from keeping two lists. Every item then has to be found twice, named twice and reconciled afterwards, and the reconciliation is where most of the effort goes and most of the errors arrive. The sequence that avoids it is short. Build one list of everything in use. Run question 1 across all of it. Run questions 2, 3 and 4 across all of it. Then attach one accountable person, one review cycle and one record to every item, whichever question caught it.
How can the whole scoping exercise be done only once?
What does somebody joining a control function do with this in the first week?
Something quite concrete, and it takes an afternoon. Ask for both lists. Count the items on each. Count the items on both. Then ask the only question that turns those three numbers into information: what share of the outcomes customers actually experience is determined by items that appear on neither list. At Sumeru Bank Limited that answer, for the intake chain alone, is that the eight components outside the model inventory determined 2,619 of 8,600 decided files in a steady month, being 30.5 per cent, and that a further 5 uses were running in the bank with no register entry at all on the day the register was signed.
The same three counts tell a supervisor, an internal auditor or a non-executive something none of the policy documents will. A firm that cannot say how many items sit on both lists has not run either scoping test recently enough to know. The overlap is the one number that can only be produced by holding the two lists side by side. The three counts are also the cheapest diagnostic available: they need no access to any component, no technical skill and no cooperation from anybody who built anything. Two lists and the patience to count are the whole of it.
The mistake: deciding which of the two layers is the real one
The obvious response to two overlapping arrangements is to declare one of them the real one and let the other become a form somebody fills in. Both versions of that fail here, and they fail on the invented bank's own counts rather than on principle.
Declare the older discipline sufficient and eight of the nine components have no entry, no named accountable person, no approval and no re-approval. The eight include the reading step that routed 1,264 files a month to a person, and the liveness check that ended 620 applications a month before anybody was scored at all, of which a hand review of 200 rejections found 31 genuine applicants, being 15.5 per cent, and all 31 shared one condition, a low-light image taken on a low-specification handset. The concentration of those errors on one group is the pattern O'Neil describes in Weapons of Math Destruction, 2016: the errors of a deployed component landing on one part of a population rather than spreading evenly across it. Not one of those eight components produces a value a credit decision consumes, so the older question will never reach them however diligently it is asked.
Declaring the newer layer sufficient rests the component that determined 69.5 per cent of the month's outcomes on a list that was 64.3 per cent complete on the day it was signed and that carried a named person against only 6 of its 14 entries. The failure is not choosing wrongly, it is choosing at all: the older discipline held the one component that mattered most and missed the other eight, the newer layer caught fourteen uses and was a third incomplete on day one, and a firm running one of them is not running a stricter version of the other.
There is a quieter version of the same mistake, and it is more common. A firm keeps both layers, gives each its own list, and then reports coverage as the larger of the two numbers. Nobody has lied. The two lists were built by different tests, so neither number describes the institution. The only figure that would describe it is the one nobody computed: the count of what sits on neither list.
What the reader has to confirm at source
The standing discipline of model risk work, including the idea of an inventory, a named accountable person and an independent challenge, originates in international supervisory material published by the Bank for International Settlements at bis.org. The material there is where the discipline began rather than the binding position in any one country. The Reserve Bank of India states what actually applies to a regulated lender in India and publishes it at rbi.org.in, including its expectations on outsourcing, digital lending, data and consent. Where the deployer is a market intermediary rather than a lender, the position is stated by the Securities and Exchange Board of India at sebi.gov.in. The accountability of a board for what its institution runs sits under company law administered by the Ministry of Corporate Affairs at mca.gov.in.
The four scoping questions, the register, the inventory and every count attached to them are Sumeru Bank Limited's own invented arrangements and are not a standard, a norm or a requirement of any authority. The current position is best read at the source before any of it is relied upon.
Sources
| Source | Document | Site |
|---|---|---|
| Bank for International Settlements | International supervisory material from which the standing discipline of model risk work originates: the inventory, the named accountable person, independent challenge and periodic re-approval | bis.org |
| Reserve Bank of India | Published expectations on a regulated lender covering outsourcing, digital lending, data and consent, and the oversight of arrangements that decide customer outcomes | rbi.org.in |
| Securities and Exchange Board of India | Published expectations where the deployer is a market intermediary rather than a lender | sebi.gov.in |
| Ministry of Corporate Affairs | Company law material on the accountability of a board for what its institution runs | mca.gov.in |
| Agrawal, Gans and Goldfarb | Prediction Machines, 2018, on a fitted component as a supplier of predictions that a person still has to act on | Harvard Business Review Press |
| O'Neil | Weapons of Math Destruction, 2016, on the errors of a deployed component falling unevenly across a population | Crown |
Sumeru Bank Limited, its intake chain, Revathi Balan, Neelima Rao and Ashok Pillai are invented.
Educational material. Not advice on any investment, tax, budget or market position.
