How to Create an AI Governance Framework for Finance
Build it in this order: write the scoping test, sweep for what is already running, write the register, name a person against every entry, then write the approval, documentation, monitoring and route back rules around what the sweep found. The order matters because every rule after the first one applies only to whatever the first one catches, and a rule written against an imagined system governs an imagined system.
Almost every arrangement of this kind gets built backwards. Somebody is asked to put governance in place, and the first artefact produced is a policy document. A document can be finished in a fortnight and shown to somebody. The document is careful, it is internally consistent, and nobody who wrote it has yet looked at a single thing the firm is running. A document written in that order is not wrong so much as empty: it is a set of rules about a set of things nobody has counted. The eight steps below run the other way round, and the first two of them are the whole reason the rest of it lands on anything real.
Why does the order matter more than any single rule in it?
Think about a household that decides to get its money under control. One household writes a rulebook first: nothing over a certain amount without agreeing it, everything on one card, a monthly sit down. The rulebook is a good one. Six months later the electricity bill has doubled. The rulebook covered spending decisions, and nobody had ever listed what the house was already paying for. The other household starts by sitting down and writing out every standing payment on the account, gets a list of fourteen things, discovers four of them nobody remembered agreeing to, and only then writes rules. The second household wrote worse rules and got a better result. The rules landed on a list.
A firm is the same and larger. Every rule in an AI governance arrangement is a rule about a set of things, and until somebody has written the sentence that defines the set and then gone and found the things, each rule governs nothing in particular. The order therefore matters more than the wording of any rule inside it. Approval rules apply to whatever is in scope. Register entries exist for whatever is in scope. Monitoring, records, a named person and a route back all attach to entries. Pull the first step out and the whole structure is standing on a definition nobody has tested against reality.
At Sumeru Bank Limited, an invented lender, the policy was written in month 0, before a single component of its retail loan intake chain existed. Three of its ten sections turned out to describe a system nobody had built, and the wording of the first section had quietly set the size of every job that followed. None of that was carelessness. Building in the other order produces exactly that, every time, and the bank is a fair example rather than a bad one.
Here is the sequence in full, before any of it is explained. Every line of it ends in an artefact somebody can hold. Read it as a list of things to produce rather than as a list of things to believe.
- Write the scoping testOne sentence deciding what the arrangement covers.Produces: a wording, and the number of things it catches.
- Sweep for what is already runningSearch for uses before writing any further rule.Produces: a list. At this bank, 14 uses where 9 were known.
- Write the registerOne line for each thing the sweep found.Produces: 14 lines, each readable by somebody outside the team.
- Name one person against every lineA person, by name, not a team and not a committee.Produces: 6 names against 14 lines, and a gap of 8.
- Write the approval ruleWho permits a new use, and what they are shown.Produces: six items an approver reads before signing.
- Write what must exist before deployment and afterDocumentation before, a record of each decision after.Produces: a six field record on every decision the thing takes.
- Write the route back, then price itThe way of working without the component, costed in people.Produces: a rehearsal date and a headcount of 11.26 posts.
- Set the review date and the two measuresWhen the whole arrangement is looked at again, and against what.Produces: two dated ratios, 64.3 per cent and 42.9 per cent.
What is the first step, and why is it first?
Step 1: how the scoping test is written, and which wordings are on the table
The scoping testThe written wording that decides which things a policy applies to, and therefore which things must be approved, recorded and monitored. is one sentence, and it is the only sentence in the whole arrangement that can be rewritten without touching anything else while changing what everything else applies to. Three wordings come up in practice, and all three are worth writing out before choosing. The three read as though they differ by shades of emphasis, and they do not.
The first says the arrangement covers anything that learns from data. The second says it covers anything the firm calls artificial intelligence. The third, the consequence wordingScoping by what an output does to somebody, rather than by how the thing producing it was built., says it covers anything whose output reaches a customer or a reported figure without a person deciding. The first two describe how a thing was built and only the third describes what it does, and that single difference is what makes the third one a governance test rather than a definition of a technique.
Written out as a procedure, the third one is two questions in a row, and both have to be answered yes. Does the output reach a customer or a reported figure? Then, does it get there without a person deciding? Two yeses put the thing in scopeCovered by the arrangement, so its approvals, register entry, records and monitoring all apply to it. whatever it is made of, and a single no leaves it out. The consequence wording is short enough that a head of a business area can apply it to their own work without ringing anybody. No definition of artificial intelligence has ever had that property.
What does each wording catch, counted over one bank?
Now run all three over the same population and count. At Sumeru Bank Limited the population is the nine components of the retail loan intake chain plus 38 other routing and calculation rules sitting elsewhere in the bank: the calculation that sets an overdraft limit, the recalculation on a floating rate loan, the rule that decides which call goes to which queue, and so on. Forty seven candidate things in total, and not one of them moves while the wording changes.
Scoped to anything that learns from data, 5 things are in: the five components of the chain that were fitted rather than written, and nothing at all from the other 38. A rule somebody typed does not learn. Scoped to anything called artificial intelligence, all 9 components come in and so do all 38 of the other rules. Once the test is a label rather than a property, everybody applies the label generously, and 9 plus 38 is 47. Scoped by consequence, 7 of the 9 components are in and 12 of the 38 survive, giving 19. The 12 is what 19 less 7 leaves. The same bank, on the same day, is a bank with 5 things to govern, 47 things to govern or 19 things to govern, and the only thing that changed was a sentence.
The two components the consequence wording drops are worth naming. Both show the test doing its job rather than being lenient. Component 3, the document classifier, is dropped because its output never leaves the system: it hands its answer to the reading step next door and no customer or reported figure ever sees it. Component 8, the drafting assistant, is dropped because a person signs every output it produces. Both are learned models, both would be in scope under the first wording, and neither reaches anybody without a human being in between.
What does each wording cost to run, in working days?
Here is the part that is almost never on the table when the sentence is being argued over. Bringing one thing into scope is work: somebody has to write down what it does, what it consumes, what its output does and to whom, and who answers for it. For anything already running, that is first-time documentationThe one-off work of writing down what an existing thing already does, before it can be governed. The writing down builds nothing new., and it builds nothing. First-time documentation produces a description of something that was already working perfectly well without one.
Sumeru Bank Limited costed that work at about 3 working days an item, its own rate and nobody else's. The rate of 3 working days an item turns each wording into a budget. Five items is 15 working days. Nineteen items is 57 working days, being under three months of one person at the bank's 20 working days a month. Forty seven items is 141 working days, being about seven months of one person doing nothing else at all, before a single control exists that did not exist the day before. A scoping wording is a budget decision wearing the clothes of a definition, and the person choosing the wording is almost never the person who will spend the days.
| Wording of the scoping test | Items | Working days | Months of one person |
|---|---|---|---|
| Anything that learns from data | 5 | 15 | 0.75 |
| Anything the firm calls artificial intelligence | 47 | 141 | 7.05 |
| Anything whose output reaches a customer or a reported figure with nobody deciding | 19 | 57 | 2.85 |
Look at the shape of that rather than the rows. Cost runs in a straight line through the item count, at three days a step, so the wording sets the bill before anybody has looked at a single item. The argument about the sentence is therefore worth having slowly, and a firm that adopts a broad wording in a morning meeting and then quietly does not apply it has not been dishonest so much as unpriced.
A scoping wording produces 47 items at about 3 working days each. What does that commit the bank to?
Before the control below is moved: the same nine components, scoped to artificial intelligence rather than to machine learning. How many items?
Change the sentence, and watch the bank change size
One control with three settings, and the only thing it changes is the wording of the scoping test. Nothing in the bank moves. Three views redraw together: the population of 47 candidate things with the caught ones filled in, a bar splitting what was caught into components of the chain and other written rules, and a marker sliding along the cost line. The default is the wording Sumeru Bank Limited chose, the consequence one: 7 components plus 12 other rules is 19 items, being 57 working days of first-time documentation. Set to the machine learning wording, the reading falls to 5 items and 15 working days, and the panel underneath names what walked out of scope when it did.
Scoped by consequence, Sumeru Bank Limited has 19 items in scope, being 7 of the 9 components of its intake chain and 12 of the 38 other written rules, and that is 57 working days of first-time documentation before a single new control exists.
Educational illustration. One invented bank, nine components in one chain and 38 other routing and calculation rules elsewhere in it. About 3 working days of first-time documentation an item is that bank's own rate and covers writing down what an existing thing already does, not building anything new. A working month is taken as 20 working days. Figures are the invented bank's own and describe one deployment.
Why is the cheapest wording the worst one?
Fifteen working days against fifty seven looks like an easy saving, and on any reasonable view of effort it is. The trouble is what the cheap wording buys, and at this bank it can be stated as a single component. Component 9 is the workflow router: a written rule of 22 lines that decides where every file goes. The router touched every one of the month's 8,600 files. The router is the highest volume thing in the entire chain. Somebody typed it, and no wording built on whether a thing learns from data can ever catch a thing somebody typed.
Under the machine learning wording the router was outside the policy, so when its waiting time before escalation was changed in month 7, no approval was needed and no record was required, and neither omission broke a single rule. There was nothing to break. The change was found in the month 10 sweep, three months later, by somebody looking for something else. Fifteen working days of documentation had bought a policy that governed five fitted components and left the thing deciding the path of every file to run on whatever anybody last typed into it.
One point carries the whole sequence, and it is worth saying without any softening. A scoping test is not a definition of artificial intelligence and it was never trying to be one. A scoping test is a decision about which things get the controls. Asking how a thing was built catches some things that decide almost nothing and misses some things that decide everything. Asking what its output does catches by consequence, and consequence is the only property the person on the other end of the output can feel.
The saving that was real, and the thing it bought
Nobody at Sumeru Bank Limited chose badly on purpose. The machine learning wording was the one that sounded most precise in the room, it was defensible in a sentence, and it cost 42 working days less than the wording eventually chosen. The saving was real.
The saving bought a policy that governed the five components deciding 84.2 per cent of the month's outcomes and left the four written ones deciding the other 15.8 per cent outside it entirely. One of those four was the router that touched every file. The cost did not show up as a breach. There was no rule to breach. The cost showed up three months late, as a change nobody could describe, in a sweep that was looking for something else.
The general form of the mistake: any scoping test written about how a thing was built will be argued over by people who know how things are built, and none of them is the person the output lands on.
Why does scoping by how a thing was built fail on a workflow router?
Step 2: what does a sweep find that no rule ever would?
The sentence exists. Before a second sentence is written, the things themselves have to be found. A sweepA deliberate search for uses already running, done before the rest of the rules are written rather than after. is a search for what is already in use, and it is the step that turns a set of rules into a set of rules about particular things. At Sumeru Bank Limited, Ashok Pillai in technology risk ran one in month 10 and found 14 uses in place where the register held 9.
He searched four routes and no single route found more than three of the five missing ones. The payments ledger, looked at for a recurring charge to a supplier of software services, found 3, and every one of them was something bought. A calculation somebody built inside a spreadsheet raises no invoice. A question put to the head of every business area found 3, a different 3. The network record of which external services staff systems had connected to found the same 3 the ledger did. The change record of systems the bank already had found 1. One route catches what is bought and the other catches what somebody built, so routes 1 and 2 together find all five, and a sweep running only the first route reports a clean bank.
The finding that matters most is the one nobody expects. Not one of the five was concealed. Every one had been set up by somebody solving a real problem in their own area, and four of the five had told somebody about it. The policy had simply never said who to tell. The absence of concealment changes what a sweep is for. A sweep is not an enforcement exercise. A sweep is the discovery step that the approval rule structurally cannot perform.
The scoping test has been written. What comes before writing anything else?
Step 3: what does the register turn the sweep into?
The sweep produces a pile of findings in somebody's notebook. Step 3 turns it into a list that survives that person leaving. One line per use, and the unit is the use rather than the component, so the whole intake chain with its nine components inside it is a single line, exactly the same size in the list as a scoring calculation somebody built in a spreadsheet in small business lending.
The output of this step is a fixed count that can be stated out loud: 14 lines, of which 9 existed before the sweep and 5 did not. Four of the 14 sit inside a service bought from a supplier rather than built by the bank, being 28.6 per cent, and three of those four are among the five the sweep added. Buying something is easier to do quietly than building it. The register is the first artefact in the sequence that anybody outside the team can read, and that readability is the whole of its value: a list only a builder can interpret has not left the builder's head.
One discipline is worth insisting on at this step. Each line is written for somebody in another part of the bank to read and say yes, that is the work this team actually does. Written in the language of the people who built the thing, the list is technically perfect and socially useless: nobody recognises their own work in it, so nobody corrects it, and an uncorrected list decays without anybody noticing that it has.
Step 4: why does a name against every line change what the other steps are worth?
Now write a person's name against each of the 14 lines. Not a team, not a committee, not a function. A person, spelled the way they spell it, who can be asked a question and will answer it. At Sumeru Bank Limited this step produced 6 names against 14 lines, being 42.9 per cent, and that gap is the second measured finding the sequence has generated so far.
Look at where the names were and were not. All 6 sit among the 9 registered entries, so two thirds of what was registered had somebody accountable, and not one of the 5 the sweep found had anybody at all. The pattern is not a coincidence: a use that arrived through an approval door was approved by somebody who had to write a name down, and a use that arrived because somebody in an area needed it solved never passed a moment where a name was required. The naming step is where the arrangement stops being a description of the firm and starts being a set of obligations somebody has accepted, and every step after it delivers its output to a name.
A name is worth pinning down by what it is not. A name is not authorship: Revathi Balan, the named accountable person for the scoring model, built no part of it and could not have. A name is not blame allocation either. A name says who a supervisor, an auditor or a board member asks, and that has to have a single answer: a question addressed to a team is a question addressed to nobody.
Step 5: who approves, and what has to be in front of them?
Only now is the approval rule worth writing. The set of things it will be applied to is finally known. The rule names the approverThe person or committee that permits a use case to be built and deployed, and who signs against a fixed set of evidence. and, more importantly, fixes what they are shown. Six things, and they fit on one side of a sheet: what it does, what it consumes, what its output does and to whom, whether a person stands in the path of that output, who will be accountable for it, and what the route back is if it has to stop.
Notice what is not on that list. Performance is not on it. Approving is a decision about consequence and accountability rather than about accuracy, so an approval paper that leads with how well the thing scored against a test population is answering a question the approver cannot act on. A signature taken without those six items is not an approval, it is a signature, and the difference shows up two years later when somebody asks who agreed to this and on what basis.
The sixth item turns the rule from paperwork into a cost, and that is why it belongs on the approval paper itself. Across the nine components of this bank's chain, a written route back to a manual process existed for 3 and had been rehearsed for 1. If the approver has to see the route back before signing, the absence of one becomes visible at the only moment when it is cheap to fix.
An approver is shown a use case's performance against a test population and nothing else. Enough?
Step 6: what must exist before deployment, and what after?
Step 6 writes two lists, and the boundary between them is the go-live date. Before deployment: what a component does, what it consumes and on what basis, what its output does, who is accountable and what the route back is. The first list is the approval paper, filed. After deployment: a record of what it actually did, decision by decision, giving a question asked in eight months somewhere to go.
Sumeru Bank Limited's record of a decision carries six numbered fields: which version of which component acted, what it read, what it produced, what the previous version would have produced where that is known, who could have intervened and did not, and the time. Four of the six were recorded at go-live. Fields 4 and 5 were added in month 9, after an upstream income field changed format in month 8 and monitoring did not flag it for six weeks. The two fields nobody recorded were precisely the two that answer what would have happened otherwise, and that is the only question anybody asks after a period of quiet trouble.
The pattern is the general shape of record keeping rather than one bank's oversight, and it is worth sitting with. The fields that get built first are the ones that describe the thing working. The fields that get built after the first incident are the ones that describe the difference between what happened and what would have happened. About 12,900 files were decided during that six week window at this bank, and 176 of them moved out of straightforward acceptance into a referral, and for those six weeks nobody could say which files those were.
Who sets expectations on a firm building an arrangement like this?
The wordings, the rate of 3 working days an item, the six record fields and the eight steps are one invented bank's own drafting. None of them is a requirement and none of them is a standard. Where a regulated lender has expectations placed on it about outsourcing, digital lending, the data it consumes, consent and the records it keeps, those are stated by the Reserve Bank of India at rbi.org.in. Where the deployer is a market intermediary rather than a bank, the equivalent position is stated by the Securities and Exchange Board of India at sebi.gov.in. The accountability of a board and its officers for the records a firm keeps sits with the Ministry of Corporate Affairs at mca.gov.in. The international standard on governance of deployed systems at a bank is set out by the Bank for International Settlements at bis.org, named as the origin of the expectation rather than as the position in India.
Step 7: what does a route back cost before the rule requiring one is signed?
Every arrangement of this kind contains a rule saying that each use case must have a route backThe written and rehearsed way of doing the work without the component, for the period when the component is unavailable or wrong. to a manual process. The route back rule is the easiest rule in the whole document to write and the only one priced in people rather than in paper. So write it, then price it, then sign it, in that order.
The route was exercised once at this bank, so the price is known exactly. In month 9 one channel fell back to manual decisioning for 4 working days, being 1,720 files at 430 a working day. At the pre-chain hands-on time of 11 minutes a file, 430 files is 4,730 minutes of desk time a day, and at the assumed 420 minutes a person that is 11.26 posts. The fallback needed 11.26 posts, exactly the 11.26 posts the chain had removed, so the route back costs the whole of the saving for as long as it runs. The exception desk had 7 people, so the bank borrowed 5 from elsewhere in retail operations for four days.
Two numbers follow from that and both belong on the paper before anybody signs. The desk's own 2,940 minutes covers 2,940 divided by 11, being 267.3 files, or 62.2 per cent of a day's book: a fallback covering more than that needs people the firm no longer employs. And the 310 minutes left over after the borrowed staff covered the fallback absorbed only 16.3 of the 150.5 exceptions still arriving each day, so the desk fell behind by about 134 cases a day and the standing queue of about 285 cases reached about 822, being 2.9 times its normal size. A route back that has never been priced is an aspiration written in the voice of a control.
A rule requiring a route back for every use case is about to be signed. What comes first?
Step 8: what two measures say whether any of this is real?
The last step sets a review dateThe stated date on which the whole arrangement is looked at again, set when it is written rather than when somebody remembers. and the two things that get reported on it. Both are ratios rather than counts. A count only ever goes up and says nothing about what has been missed.
The first is completeness: how many entries the register holds against how many a sweep finds. At Sumeru Bank Limited that was 9 against 14 on the day the register was signed off, being 64.3 per cent, and all five it was missing were already running that day, so the shortfall is a day one number rather than growth. The second is accountability: how many entries carry a named person. The named-person count was 6 of 14, being 42.9 per cent. Both ratios are cheap to produce and neither can be produced at all without the sweep, and that is why step 2 sits where it does rather than somewhere more convenient.
Without one refinement written into the measure, the reading is unfair to the arrangement. Measured against everything running, the register was 64.3 per cent complete. Measured against the 11 things its own chosen scoping test actually catches, it was 9 of 11, being 81.8 per cent. Both readings are true and they answer different questions: the first counts everything the firm runs, the second counts only what the firm said it would govern. Report the denominator alongside the number, every time, or the measure becomes an argument about arithmetic in a room where nobody has the sweep in front of them.
Name the two measures that say whether the arrangement is real rather than written.
How does somebody actually start this on a Monday morning with nothing?
For a small lender with three or four things running and no arrangement at all, the first week costs almost nothing and produces two artefacts. The Monday goes on writing the two questions of the consequence test on a single sheet. The Tuesday and Wednesday go on the two sweep routes that between them find everything: the payments ledger is read down for recurring charges to suppliers of software services, and one plain question is put to the head of each business area, phrased as what does this area use that produces an answer it acts on rather than as what artificial intelligence does it use. Nobody answers the second question truthfully, and almost everybody answers the first one. By Friday there is a list, and it is longer than the one anybody expected.
The week after that, the list becomes lines and the lines get names. Only then is it worth writing anything that reads like a policy. By that point the policy writes itself: every section has a real thing to point at. The whole sequence is designed so that the expensive artefact, the document, is written last against evidence rather than first against imagination.
The same eight steps are what somebody reading a firm from the outside works backwards through. A lender assessing a borrower that runs a chain like this, an internal audit team arriving in week one, or a diligence team looking at an acquisition all ask for the same three artefacts in the same order: the scoping sentence, the register, and the two ratios with their denominators. If the scoping sentence is about how things are built, they know before opening the register what kind of thing will be missing from it. If the register exists but no sweep has ever been run, the completeness figure is unknown rather than high. And if the two ratios have never been produced, the honest reading is that the arrangement is a document rather than a practice, whatever the document says about itself. The cost of the whole exercise is small and knowable: at this invented bank, the chain it governs took Rs 2,40,00,000/- to build and Rs 65,00,000/- a year to run, and the governance around it was 57 working days of one person's time.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Published expectations on a regulated lender covering outsourcing, digital lending, the data a system consumes, consent and record keeping. Any obligation on a lender building an arrangement of this kind is stated there | rbi.org.in |
| Securities and Exchange Board of India | The equivalent published position where the deployer is a market intermediary rather than a bank | sebi.gov.in |
| Ministry of Corporate Affairs | The accountability of a board and of its officers for the records a firm keeps, being the layer a named accountable person inside a firm reports into | mca.gov.in |
| Bank for International Settlements | The international standard on governance of deployed systems at a bank, named as the origin of the expectation rather than as the position in India | bis.org |
Sumeru Bank Limited, its retail loan intake chain, Revathi Balan and Ashok Pillai are invented.
Educational material. Not advice on any investment, tax, budget or market position.
