Artificial Intelligence in Finance: What the Term Actually Covers
Artificial intelligence (AI) is a label for a group of methods rather than a property a system either has or lacks. Inside one deployed financial system the label usually covers several separate components, some deriving their behaviour from past data and some following steps a person wrote down. Which of those two a component is decides who can explain its output and what a reviewer has to read.
The mismatch is between two attachments. The label attaches to a system. Accountability attaches to a part. A system is approved once, in one line, by people who mean something quite specific by the phrase, and it then runs as nine separate things that fail separately, change separately and have to be explained separately. Everything awkward about the words artificial intelligence inside a bank comes out of that gap between how the thing was approved and how it actually runs.
What does the term cover, and what is it not?
Think about the word vehicle for a moment. Vehicle is a useful word at a toll gate and a useless one in a workshop. Nobody repairs a vehicle; they repair a clutch, a radiator, a brake line. Artificial intelligenceA label for a group of methods that produce an output without a person working it out each time. sits in the same place. The label is a collecting word for a set of methods that produce an output without a person working the answer out each time, and the label is genuinely useful in a conversation about a budget, a strategy or a market. The label stops being useful the moment somebody has to fix something, explain something, or answer for something.
The term is not a switch, a threshold or a certificate. No line was crossed on the day a system became one. No measurement separates a system that has it from a system that does not, and no supplier can hand over a document proving the presence of it. Presence is not the sort of property the term names. The term describes how a particular behaviour was arrived at. Descriptions of that kind attach to parts, not to products, and certainly not to whole platforms.
Finance is where a missing description hurts most. Somebody always has to say why. Why was this application declined. Why did this file wait two working days. Why did the number in the monthly pack move. A label that covers a hundred things at once cannot answer any of those questions, and the person holding the label is usually the person who gets asked.
Is artificial intelligence a property a system has, or a label for a group of methods?
Why is the useful question never whether this is AI?
Because whichever answer comes back, nothing follows from it. Say yes: what to read, who to call and what to put in front of a customer are all still unknown. Say no: the system is still deciding things about people, and the decisions still need explaining. A question is worth asking only when the answer changes what somebody does next, and this one changes nothing.
Swap it for a question that does work. Which parts of this learn from data, and which parts did a person write down. The second question has consequences immediately. If a person wrote it, there is a document, and the document can be read this afternoon. If the behaviour came out of past examples, there is no document to read, and what gets reviewed instead is the examples and the measured behaviour. The first is an afternoon. The second is a project. Sizing that difference correctly is most of what a reviewer is paid for.
A household version, if it helps. Asking whether a kitchen is modern produces an opinion. Asking which appliances have a manual with a wiring diagram in it and which ones need somebody called out produces a plan for the evening. The second question sorts the room. The first just describes it.
What are the two kinds of component, and how do they differ?
A deployed system is made of componentsOne separately built and separately changeable part of a deployed system., and each one is either written or learned. A rule setA procedure a person wrote down, which can be read from the first line to the last. is a procedure somebody wrote down: thirty-four lines, or four hundred, but readable from the first line to the last by anyone who is allowed to open the file. A learned modelA component whose behaviour was derived from past examples rather than written by a person. derived its behaviour from past examples instead. Nothing was ever written down, so there is no line anywhere that says what it does.
The two kinds differ on four things a reviewer actually cares about, and not one of the four is about how modern the component is. Where the behaviour came from, what has to be read to check it, what has to happen for it to change, and what can be put in front of the customer. The four differences are the entire practical content of the distinction, and they are why the distinction is worth drawing at all.
A component compares declared income against a bank statement and routes the file if the gap exceeds a set tolerance. Written or learned?
One line in a bank register read: AI system, retail lending. How many separately changeable components sat underneath it?
What actually sits inside one deployed lending chain?
Sumeru Bank Limited, an invented mid-sized Indian bank, built one thing: a retail personal loan intake chainIn the case here, one bank's retail loan system running from an application on a handset through to a decision. running from a customer starting an application on a handset through to a decision, a disbursal and the monitoring after it. The board approved it as an AI system. The chain has nine components, in a fixed order, and a file passes through all of them.
Five components learn and four are rules, and the split does not follow the order or the glamour of the work. The learned five are the liveness check, the document classifier, the field reading step, the scoring model and the drafting assistant. The written four are the identity match, the income corroboration step of thirty-four lines, the fraud rules on the servicing book and the workflow router. Numbered as above: two, three, four, six and eight learn; one, five, seven and nine were written by a person.
Now the month. In one steady month, 10,000 applications started on a handset. 8,600 of them completed digital onboarding and reached the decision engineThe point in the chain where a file receives an accept, a decline or a referral., being 86.0 per cent. The other 1,400 did not: 620 were rejected by the liveness check, 480 were abandoned at document upload, and 300 could not complete the consent step. Of the 8,600 that arrived, 5,590 were decided straight throughA file that reaches its outcome with no person touching it at any point. with nobody touching the file, being 65.0 per cent, and 3,010 were routed to a person.
Which component determined the outcome, and is it the impressive one?
Take the 8,600 files that reached the engine and ask a narrow question of each: which single component settled what happened to this file. Not which component touched the file, but which one settled it. Attribute every one of the 8,600 to exactly one component and the counts come out as follows. The scoring model determined 5,981. The field reading step determined 1,264. A field it could not read with enough confidence sent the file to a person, and the person settled the outcome. Income corroboration determined 602, the fraud rules 452, the identity match 189 and the router 112. The six counts add to 8,600 exactly.
Here is the number all of this is built towards. The five components the bank called AI determined 7,245 of the month's decisions, being 84.2 per cent, and the four it called rules determined the other 1,355, being 15.8 per cent. One file in six a month, settled by something nobody in the approval conversation had thought of as part of the system at all. Notice also that the biggest single number, 5,981, belongs to a component that only ever produces a prediction on which a written policy then acts. Agrawal, Gans and Goldfarb frame it this way in Prediction Machines: the machine supplies the prediction, and a person still has to decide what to do about it.
Three of the nine components determined none of the 8,600 outcomes. Does that make them safe to leave out of a review?
Which component touched every file, and why did nobody look at it?
Component 9, the workflow router, is a set of rules somebody wrote that decides where each file goes next: onward to the engine, into a queue, out to the exceptionA file the system sends to a person instead of deciding it automatically. desk, back to the customer for another document. Every single one of the 8,600 files passed through it, and so did the 1,400 that never reached the engine. The component with the widest reach in the chain is a rule set, and the reason nobody examined it is that nobody found it interesting enough to call AI.
The router determined only 112 outcomes outright, a number small enough to keep it invisible. But determining an outcome and shaping an experience are different things. The router decided which queue a file joined, and therefore whether a customer waited about four minutes or two working days. When somebody eventually asked why a particular file sat for two days, the answer lived in the router, and the router had never been opened by anyone reviewing the system.
Everyday version: in a hospital, the machine everyone talks about is the scanner. The desk that puts the form in one tray rather than another decides whether a patient is seen in twenty minutes or four hours. Nobody writes a paper about the tray.
Which component in the chain touched every single one of the 8,600 files?
What does accountability look like for each kind?
Accountability for a written component and accountability for a learned one are two different jobs, and treating them as one is where most of the trouble starts. For a written component, accountability is document work. Somebody can be handed the thirty-four lines of the income corroboration rule and asked whether the tolerance in line nineteen is the tolerance the credit policy intended. Neelima Rao, in the risk function at Sumeru, read all thirty-four lines in twenty-five minutes. There is a version history, a person who made the last change, and a sentence that can be lifted straight out and put in front of a customer.
For a learned component, none of that exists. There is no line nineteen. Accountability for a learned component is accountability for a population rather than for a sentence: what examples it was fitted on, how its behaviour has been measured since, and whether that measurement is still being taken. The independent review of the scoring model took eleven working days against twenty-five minutes for the rule, and the difference is not effort or seniority. One of them can be read, and the other can only be measured.
The customer-facing consequence is sharper still. When the written income rule routes a file, the bank can show the applicant the procedure that did it. When the scoring model declines a file, there is no procedure to show, only an attributed reason. Two applicants can get what feels like the same refusal, and only one of them can be handed the reason in its original form.
Which four questions tell a reviewer what is in front of them?
A reviewer will often be looking at a system somebody else built, described by people using words that carry no information. Four questions, in order, settle what kind of thing each part is, and the first three can be answered without opening anything technical at all.
Question two catches a case neither of the others does, and that is what earns it a place. A written rule with unchanged inputs returns the same answer every time, always. So if the answer moved, either the component is not a plain written procedure, or something feeding it changed underneath. Both are worth knowing, and in practice the second is the more common finding by a wide margin.
A component gives a different answer to the same file on two different days. What does that establish?
Where does the label get attached to something that neither learns nor decides?
The label travels downward as well as upward, and this direction wastes real effort. At Sumeru, a monthly report assembled by a database query was listed as an AI use case. Nothing in it was fitted to anything. The report does not decide; it counts rows and prints them. The report determined none of the 8,600 outcomes and reaches nobody outside the reporting pack. Meanwhile the workflow router touched every file in the month and was not on the list at all.
A monthly report built by a database query was listed as an AI use case. What does that listing actually cost?
What does the label cost a register when it covers a whole system at once?
A firm keeps a list of the systems it has approved, and each register entryOne line in the list a firm keeps of the systems it has approved, naming what was approved and who answers for it. names what was approved and who answers for it. Sumeru wrote one row. The row is worth reading closely for what it cannot say.
A register row that names a system rather than its parts cannot tell anybody which part was approved. The missing component list is not a documentation quibble. The row is what triggers reviews, sets change control and tells an accountable person what she is accountable for. Nine components under one line means nine components sharing one approval date, one review cycle and one name. Each of them can change on its own on any Tuesday.
Before the control below moves: of nine components, five learn and four are rules. Opening the five learned ones covers what share of the month's 8,600 decisions?
Open components, and see how much of the month is covered
One control: how many of the nine components a reviewer opens, from one to nine. One consequence: the share of the month's 8,600 decisions whose determining component has now been opened. Two orders are available. By decisions determined works down the list from the largest. The order the bank used opens the five learned components first. Anybody at Sumeru who said AI meant those five. The default below is the second order at five components opened. Five opened covers 7,245 decisions, being 84.2 per cent, and leaves 1,355 files a month determined by a component nobody opened.
5 of 9 components opened
Five components opened, taken in the order the bank actually used, learned components first. That covers 7,245 of the month's 8,600 decisions, being 84.2 per cent, and leaves 1,355 files a month whose determining component nobody opened.
The error that gets made, and what it costs
The board approved the intake chain as one register entry reading AI system, retail lending, with one accountable name against it. Nine components sat under that line. When the review pass ran, it opened the five that had been described as the AI: the liveness check, the classifier, the reading step, the scoring model and the drafting assistant. Nobody had called the four rule sets AI, so nobody opened them.
The four rule sets determined 1,355 outcomes a month, being 15.8 per cent, and one of them decided where all 8,600 files went.
The cost was not an error in a number. The cost was that when a customer asked why a file had waited two working days, the only people who could answer had never been asked to look.
How does a lender, an analyst or a board actually use this?
What each of them does with the count
A board member gets one row to approve and a cost to sign off. Sumeru spent Rs 2,40,00,000/- to build the intake chain once and Rs 65,00,000/- a year to run it, and both numbers sat behind that single line. The useful question in the room is not whether the spend is justified. The useful count is how many separately changeable parts the row covers, and how many of them she is being asked to hold one name against.
A reviewer or an internal auditor uses the count to size the work before agreeing to it. Four written components at roughly the reading time of the income rule is days. Five learned components at roughly the effort of the scoring model review is weeks. Agreeing to review an AI system without first counting its parts is agreeing to an unknown quantity of work.
A customer-facing manager uses it to know which complaints she can answer and which she cannot. Where a written component acted, she can show the procedure. Where a learned one acted, she can give a reason but not the rule. Knowing which of the two produced a given outcome, before the complaint arrives, is the difference between a straight answer and a fortnight of internal enquiry.
Who sets expectations on a deployer here
A bank deploying a chain like this in India sits under the Reserve Bank of India. The Reserve Bank publishes its expectations on outsourcing, digital lending, customer data and consent at rbi.org.in. Where the deployer is a market intermediary rather than a lender, the Securities and Exchange Board of India sets the equivalent expectations at sebi.gov.in. Requirements, thresholds and effective dates move, and the issuing body's own site carries the current position.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering outsourcing, digital lending, customer data and consent | rbi.org.in |
| Securities and Exchange Board of India | Expectations where the deployer of such a system is a market intermediary | sebi.gov.in |
| Ministry of Corporate Affairs | Material on the accountability of a board for what it approves | mca.gov.in |
| Bank for International Settlements | International supervisory material on the deployment of such systems by banks | bis.org |
| Agrawal, Gans and Goldfarb | Prediction Machines, on a learned component supplying a prediction that a person must still act on | Harvard Business Review Press |
Sumeru Bank Limited, Revathi Balan and Neelima Rao are invented.
Educational material. Not advice on any investment, tax, budget or market position.
