Credit Decisioning Systems: From Application to Answer
Credit decisioning turns a loan application into an answer. At Sumeru Bank Limited, invented, a fitted component scored 5,981 of one month's 8,600 files and two cut-offs the bank set for itself sorted them: 4,902 accepted, 688 declined, 391 sent to a person. The other 2,619 files reached an answer without that component scoring them at all.
A decisioning arrangement is a chain of parts, and only one of them is fitted to data. The fitted part produces a number. The outcome, the speed, the words in the letter and whether anybody can explain it afterwards are the whole of what an applicant experiences, and every one of them comes out of the bank's use of that number. The outcome, the speed, the wording and the record are choices people made, and whether an arrangement like this can be defended turns on whether anybody wrote the choices down.
What is a credit decisioning system, and which of its parts is fitted?
Picture a school sports trial. A teacher stands at the finish line with a stopwatch and times every child over sixty metres. The stopwatch produces a number and does nothing else with it. Somebody else, usually the coach, has already decided that anybody under nine seconds goes straight into the team, anybody over eleven does not, and everyone in between gets watched again next week. Two schools with identical stopwatches and identical children can field completely different teams, and the stopwatch is not the reason. The lines are.
Credit decisioningThe whole arrangement that turns a loan application into an answer, of which a fitted component is only one part. is that arrangement running on loan applications. At Sumeru Bank Limited, invented, an application starts on a handset, passes a set of written checks, reaches a component that returns a number, meets two lines somebody chose, gets routed, gets answered, and leaves a record behind. Six parts, and exactly one of them was fitted to data. Every part except the fitted one is a decision a person made and could have written down. The imbalance is not a criticism of the design. Every deployed decisioning arrangement has that shape, and losing sight of it is how a bank ends up describing a whole workflow by naming the one component in the middle of it.
What is the credit model at this bank, and what does it produce?
The bank calls component 6 its scoring modelA component fitted to past data that returns a number about an application, rather than an outcome or an explanation., and the phrase credit model does a lot of work in conversation that it does not do in the arrangement itself. Here is what it is at Sumeru Bank Limited, invented. Component 6 was fitted on a past window of 3,00,000 applications, of which 2,40,000 were accepted and therefore observable and 60,000 were declined and carry no observable outcome. The fitting ran against a labelThe outcome a fitted component was trained to recognise, defined by choices somebody made about what counts and over what period. carrying four numbered choices: what counts as bad, set at 90 days past due; the observation window, set at 12 months; the population, being accepted applications only; and accounts closed early, excluded. All four are that bank's own choices and none of them is a standard. How a component is fitted, and how it is evaluated afterwards, are settled elsewhere and used here as they stand.
Component 6 produces a single value between 0 and 1000 on a scale the bank invented for itself, where a higher number means the component placed the application further from the outcome it was fitted to recognise. The number carries no decision, no reason, no letter, no priority and no view about any person. Agrawal, Gans and Goldfarb, in Prediction Machines, 2018, separate the prediction from the deciding that follows it, and this component sits entirely on the prediction side of that line. Everything after the number is somebody else's work.
Component 6 is the odd one out among the fitted parts of this chain, and one more thing about it is worth holding. The other fitted components read images, documents and free text, where nobody could ever write down in advance the set of things that might arrive. Component 6 reads a set of fields on an application form, and that set is completely enumerable. The part of component 6 that could not be stated in advance is not the input but the outcome, meaning whether an account repays. That is why it was fitted at all, and it is also why the arrangement around it has to carry so much of the weight.
What does the scoring component at this bank produce?
Which files did the credit model score, and where did the other 2,619 go?
Anybody describing this arrangement would say the scoring model decides the applications. Put a month's counts against that sentence and it turns out to be true of about two thirds of them. 8,600 files reached the decision engine in the month. Component 6 scored 5,981 of them. The other 2,619 files, being 30.5 per cent of the month, reached an answer without that component ever scoring them. They had already left the automatic path earlier, each for a reason that had nothing to do with creditworthiness.
Think of the queue at a passport or municipal counter. Most people who leave without being seen by the officer were not turned down by the officer. The door sent them back for a missing photocopy, an address that did not match, a form filled in the wrong colour of ink. The officer never formed a view about them at all, and an officer asked why so few people were rejected today would give an answer about the door and not about the desk.
| Why the file left before the score | Files | What was acting |
|---|---|---|
| A document field the reading step could not read with enough confidence | 1,264 | a fitted component |
| The declared income could not be corroborated from the statement | 602 | a written line |
| A fraud line fired on the file | 452 | a written line |
| An identity record did not match across two sources | 189 | a written line |
| The consent record was incomplete | 112 | a written line |
| Files that reached an answer with no score at all | 2,619 | 30.5% of 8,600 |
Two readings sit inside that table and both are worth taking away. First, every one of those 2,619 files went to the exception desk, so a person and not the scoring component decided what happened to them. Second, 1,355 of the 2,619 were sent there by written lines rather than by anything fitted, and not one of those 1,355 was a decision about a loan. The written lines in this chain decide where a file goes and never what happens to the applicant. The fitted side does both. Add the 391 files the scoring component sent to a person to the 2,619 that never reached it, and the total is the 3,010 exceptions the desk worked in the month.
8,600 files reached the decision engine and the scoring component scored 5,981. Where did the other 2,619 go?
Who chose the two cut-offs, and what are they a choice about?
Every June, colleges publish cut-off lists. The board exam produced a mark for each student and the college chose the line, and two colleges looking at exactly the same marks publish different lines on the same afternoon. Nobody confuses the examiner with the admissions committee. In a decisioning arrangement the same two jobs exist, and both of them happen inside one system in about four minutes, so the two are much easier to confuse.
At Sumeru Bank Limited the cut-offThe value at which an arrangement stops asking and acts. The deployer picks it; the fitted component has no say in it. pair is 720 and 580 on the bank's own 0 to 1000 scale. At or above 720 the file is accepted automatically. Below 580 it is declined automatically. Everything between goes to a person. Revathi Balan, head of retail credit, set both numbers and is the named accountable person for the component. Neither number came out of the component, neither is anybody's standard, and both are Sumeru Bank Limited's own choices. The component would return exactly the same 5,981 numbers if the pair were 740 and 560, and a different set of people would be accepted, declined and referred.
Where the expectations on an automatic credit decision come from
A regulated lender running an automatic decision on a retail borrower sits under the Reserve Bank of India at rbi.org.in. The Reserve Bank publishes its expectations on digital lending, fair practice, the treatment of a borrower and the use of data and consent. Where the deployer is a market intermediary, the Securities and Exchange Board of India at sebi.gov.in is the equivalent. The accountability of a board for what a company does is a matter for the Ministry of Corporate Affairs at mca.gov.in.
The 0 to 1000 scale, the accept cut-off of 720, the decline cut-off of 580, the 90 days, the 12 months and the nine record items are all Sumeru Bank Limited's own choices, and not one of them is a standard, a norm or anybody's requirement.
The referral range runs from 580 to 719, so it is 140 of the 1,000 points wide. What share of the month's 5,981 scored files sat inside it?
How much of the scale does each outcome cover, and how many files sit in it?
Take the two questions apart. The answers are nothing like each other. On the scale itself, the declined region runs from 0 to 579 and is 58.0 per cent of the whole thing, the referral range is 14.0 per cent and the accepted region is 28.0 per cent. Now count the files. 688 of the scored files were declined, being 11.5 per cent, 391 were referred, being 6.5 per cent, and 4,902 were accepted, being 82.0 per cent. The narrowest region on the scale holds by far the most files, and more than half the scale holds about a tenth of them. 3,012 files, being 50.4 per cent of everything scored, sat at 800 or above on their own.
The lopsidedness is the fact underneath every argument anybody will ever have about this arrangement. Move a cut-off twenty points near the top of the scale and it walks through crowded ground. Move it twenty points down at the bottom and it passes through almost nobody. Where the two lines should sit, and what a move costs in files and in expected trouble, is set out under the decision threshold. The shape of the population comes first. A line cannot be swept through a population nobody has counted.
What happens to a file inside the referral range?
The referral rangeThe values between the two cut-offs, where the arrangement has been told not to act automatically and a person decides instead. is the part of the design people skip when they describe it, and it is the part that costs money. A file landing between 580 and 719 has not been declined and has not been accepted. The component has returned its number and the arrangement has been told, by the two chosen lines, that this number is not enough on its own. So a person now makes a credit decision on that file, using the score along with everything else in it. From month 5 each of those files also carries a second reviewer.
391 files a month took that route. The exception desk works to a locked average of 19 minutes a file, an average across all exception kinds rather than a measurement of these particular files. 391 files at 19 minutes is 7,429 minutes a month, being about 0.88 of a full post. The second reviewer repeats the re-reading and the deciding, 7 of those 19 minutes, so 391 times 7 is 2,737 minutes, being another 0.33 of a post. The middle range is 6.5 per cent of the scored files and about 1.2 posts of standing capacity, and a design that forgets to fund it has quietly decided that those files will wait.
Two very different arrangements sit inside this one month, and both get described with the same phrase about a person overseeing the system. On the 391 referrals, a person is inside each file. On the 4,902 accepts and 688 declines, being 5,590 files nobody touched, a person watches the component's behaviour across files and has the standing to stop it. The first acted 391 times in one month. The second was exercised once in six months from go-live, in month 9 week 3, when one channel fell back to manual decisioning for four working days, covering 1,720 files.
The accept cut-off moves from 720 to 740 and nothing else changes. What has changed about the component?
Why is the component that decides most files still a minority of the system?
Both of these sentences are true of component 6 in the same month, and they have to be held together. The first: it determined the outcome of 5,981 of the month's 8,600 files, being 69.5 per cent, more than every other component in the chain put together. The second: on any one of those files it was one step out of six, and the five steps around it were written down by people. A component can determine most of the decisions and still be a minority of the system, and an arrangement is defended or lost in the other five steps.
An everyday version helps here. A household deciding whether to lend money to a cousin might look at what he earns. The number matters. But the amount, the date it has to come back, whether anybody says it aloud at the next wedding, and what happens if the first instalment is late are all decided by the household, and if the arrangement goes wrong it will go wrong in one of those and not in the earnings figure. O'Neil, in Weapons of Math Destruction, 2016, makes the related point that a component's errors do not fall evenly across the people it is applied to, and where they fall is settled by the parts around it rather than by the fitting.
What has to be written down before an automated credit decision runs?
Ask a household that has ever had building work done what they wish they had written down. Almost nobody says the specification of the cement. The answers are all about what happens to them: what the final amount will be, who decides when the work is finished, what happens if the tiles do not arrive, and who they can call when something cracks. The builder needs the specification in order to work, so the specification gets written down anyway. The rest is the part that only matters later, to the person on the receiving end.
A decision recordWhat has to be written down about an automated decision before it runs, so that somebody who was not there can examine it afterwards. for an automated credit workflow has exactly that structure. Sumeru Bank Limited settled on nine numbered items, and the list below is that invented bank's own. The list is not a standard, not a checklist issued by anybody and not a requirement. The shape of the list is worth more study than the list itself. Five of the nine items are about how the component was made, and four are about what it does to a file and who can stop it.
How to Document an Automated Credit-Decision Workflow
| Item | What it has to answer | What it is about |
|---|---|---|
| 1 | The population the component was fitted on, and what was excluded from it | how it was made |
| 2 | The label, with all four of its choices written out | how it was made |
| 3 | Every input field, where it comes from and who is accountable for it there | how it was made |
| 4 | The cut-offs, who set them and on what date | what it does to a file |
| 5 | What happens to a file the component declines to decide | what it does to a file |
| 6 | The reason attributed to each declined file, and how that reason is produced | what it does to a file |
| 7 | The named person accountable for the component | how it was made |
| 8 | What would have to be true for the component to be stopped, and who may stop it | what it does to a file |
| 9 | The record of every change to any of the above | how it was made |
Read the list once as a builder and once as an applicant, and it splits cleanly in two. A builder needs items 1, 2, 3, 7 and 9 in order to finish the work and hand it over. An applicant, or anybody answering an applicant, needs items 4, 5, 6 and 8, and needs none of the first five. The split between what a builder needs and what an applicant needs is the whole result arriving early, and the month 12 validation turned it from an observation into a count.
Which of these is not one of the nine items of the credit decision record?
Of those nine items, how many did the independent validation at month 12 find documented?
What did one month and one validation actually look like?
Here is the whole worked instance in one place. Every figure belongs to Sumeru Bank Limited, invented, and describes one month at steady state in one deployment. Month 6 is the month; month 12 is when Neelima Rao, who sits in the risk function and built no part of the chain, ran the independent validation.
| The month, and then the record | Count | Check |
|---|---|---|
| Files reaching the decision engine | 8,600 | the month at steady state |
| Files that left before the score, for a person | 2,619 | 1,264 + 602 + 452 + 189 + 112 |
| Files scored by component 6 | 5,981 | 5,981 + 2,619 = 8,600 |
| Accepted automatically at 720 and above | 4,902 | 82.0% of the scored files |
| Sent to a person, scoring between 580 and 719 | 391 | 6.5% of the scored files |
| Declined automatically below 580 | 688 | 11.5% of the scored files |
| The three outcomes together | 5,981 | 4,902 + 391 + 688 |
| Record items documented at the month 12 validation | 5 | items 1, 2, 3, 7 and 9 |
| Record items not documented | 4 | items 4, 5, 6 and 8 |
| The record as a share | 55.6% | 5 of 9 |
One nuance in that table matters more than the count. Item 4 asks for the cut-offs, who set them and on what date. Files were being sorted by the two numbers every four minutes, so the numbers were obviously in the running arrangement. The pair of numbers written down with a name and a date beside them did not exist, so who set them had to be established by asking rather than by reading. Item 5 has the same shape: the referral route was working, 391 files a month were going down it and a second reviewer had been added in month 5, and none of that had been written anywhere as the thing the arrangement does when the component declines to decide. A practice that everybody follows and nobody has written down is not a record. Such a practice is a habit, and habits leave with the people who hold them.
Why were the four missing items the ones about the applicant?
The four missing items are not a story about carelessness, and reading them that way loses the lesson. A different question can be asked of each of the nine items: did somebody have to settle this in order to finish building the thing? A component cannot be fitted without deciding the population. Nor can a component be fitted without the label and its four choices, or run without the input fields and their sources. Approval demanded a named accountable person, and the build process kept a change record because engineering work produces one. Every item that was documented is an item somebody had to answer before the work could be called done.
Now ask the same question of the other four. Nobody has to write down who chose 720 in order for 720 to work. The routing rule sends a referred file to a person whether anybody has described the arrangement or not, so nobody has to write down what happens to it. Nobody has to write down how a reason is produced in order for a reason to appear on a file. And stopping is exactly the situation nobody is in on the day they go live, so nobody has to write down what would stop the component in order for it to run. Documentation follows the work, and the four items about what happens to a person are the four that no step of the work demanded.
Why were those four items the ones missing?
The error that gets made, and what it costs
The mistake is to read 5 of 9 as a score and to raise it to 9 of 9 by writing the four missing items as a description of what the arrangement already does. The mistake is made by somebody senior, in good faith, usually within a fortnight of the finding, and it produces four new paragraphs saying the cut-offs are 720 and 580, referred files go to a person, a reason is attributed automatically, and the component would be stopped if something went seriously wrong. Every sentence there is true. A record item is not a description of current behaviour, so not one of the four adds anything. A record item is a commitment made in advance that somebody can later be held to.
The cost at Sumeru Bank Limited was not a bad decision. The cost was that four months of decisions could be described and not defended. Ask why this file was declined and the answer is a description. Ask who chose 720, on what evidence and on what date, or what the arrangement had committed to do when the component declined to decide, or what would have had to happen for anybody to stop it, and there is nothing to read, so the answer becomes whatever somebody remembers today.
The reverse mistake is just as available, and it is to read this as evidence that the people who built the component were careless. The people who built it were not. The four items went missing because no step of the work required them. A record therefore has to be settled before go-live rather than assembled from whatever the build happened to leave behind.
How long before any of this can be argued with?
Here is the fact that decides how much the record matters. Under this bank's own label, an account is bad if it reaches 90 days past due inside a 12 month observation window. So the outcome of a decision taken in month 6 is not knowable for 15 months, and the month 6 book cannot be scored against reality until month 21. Nobody can check whether these decisions were good ones for well over a year, whatever anybody says in the meantime.
Set that beside the independent validation at month 12. No outcomes existed yet, so Neelima Rao could not possibly have been checking them. She was checking the record, and the record is the only thing there is to check for nine more months. Every tuition class knows this feeling: nobody can tell whether this year's teaching worked until the results come out, so what can be examined in the meantime is whether the class was run the way it said it would be.
How long after a month's decisions can they first be scored against real outcomes at this bank?
Where does accountability for one decision sit, and what does that person need?
An accountable personThe named individual who answers for a component and its outputs, as distinct from whoever built it or operates it. was named for component 6 and that is item 7, one of the five that existed. Revathi Balan, head of retail credit, is that person. Naming her is the easy half. The hard half is whether she has, in front of her, the things she would need on the day one applicant asks the only question that matters to them. The question is always the same. Why was this application declined?
Walk it through with the record as it stood. She can see the score the file received. She can see the route it took. She can see that a reason was attributed on the file. Three things are missing from what she can read. Item 6 is how that reason is produced. Item 4 is the pair of cut-offs written down with the name and date of whoever set them. Item 8 is what would have to be true for anybody to stop the component and who is allowed to. Naming somebody accountable without giving them the record that answers the question makes them the person who carries the blame rather than the person who can respond.
The applicant at the other end is not a problem being handled. A declined application is a person who wanted something and did not get it, and it does not follow that any of the 688 deserved that outcome, brought it on themselves or should have known better. Whether any individual decision was right is not something the record can settle. The record settles something narrower. A bank running an arrangement like this should be able to answer the question, and whether it can is settled months earlier, in a document, by people who have never met the applicant.
One applicant asks why their application was declined. Who has to be able to answer, and what do they need in front of them?
What is worth looking at in one hour with an arrangement like this?
An arrangement like this gets read by more people than might be expected: a lender's own second line of credit review, an internal auditor, a member of a credit committee who has to approve the thing, a risk officer being asked to sign, and an analyst reading what a lender discloses about how much of its book is decided automatically. A borrower's adviser can use the same four questions to ask a lender something answerable.
First, ask which parts are fitted and which are written, and get the list by name. A chain described as a scoring model usually turns out to have one fitted component and several written ones, and the written ones are the parts anybody can read this afternoon. Second, ask how many of last month's files the fitted component never scored. At this invented bank the answer was 2,619 of 8,600, and an arrangement described by its model was deciding nearly a third of its volume somewhere else entirely. Third, ask whether the cut-offs carry a name and a date. A threshold that nobody has signed is the single most common gap in an arrangement like this, and it is also the cheapest one to close. Fourth, ask what happens in the middle range and who pays for it. The middle range is where the design either funds a person or quietly makes people wait.
None of those four questions requires an understanding of how a component was fitted. Three of them can be answered from a document. The four questions are worth more than any single number, and the arithmetic of moving a cut-off belongs with the cut-off itself.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Published expectations on a regulated lender covering digital lending, fair practice, the treatment of a borrower and the use of data and consent | rbi.org.in |
| Securities and Exchange Board of India | The equivalent expectations where the deployer of an automatic decision arrangement is a market intermediary | sebi.gov.in |
| Ministry of Corporate Affairs | The accountability of a board for what a company does, into which a named accountable person for a component ultimately reports | mca.gov.in |
| Agrawal, Gans and Goldfarb | Prediction Machines, 2018, on the separation between a prediction and the deciding that follows it | Harvard Business Review Press |
| O'Neil | Weapons of Math Destruction, 2016, on errors falling unevenly across the people a component is applied to | Crown |
Sumeru Bank Limited, its intake chain, Revathi Balan and Neelima Rao are invented.
Educational material. Not advice on any investment, tax, budget or market position.
