Machine Learning vs Rule-Based Automation
A rule set needs somebody who can state the correct outcome before the case arrives. A learned component needs past cases with the outcomes already attached. Neither requirement is about capability, and both are settled before anybody builds anything. Where both can be met, the choice turns on what a reviewer, a complaints desk and the person affected need to be shown, and rarely on which one performs better.
Almost every argument between these two is held on the wrong ground. Somebody says the learned option is more accurate, somebody else says the written one is safer, and both sentences are opinions about performance. The four things that actually separate them are not opinions. The four separators are facts about what exists: what has to exist before either can be built at all, what exists afterwards for somebody to look at, how each one behaves when it is wrong, and what a year of keeping it honest costs in days of somebody's time. All four are settled before a line of either one is written. The choice is usually made long before anybody in the room notices that a choice is being made.
What does each one need before it can exist at all?
The gate of a housing society holds both of these, familiar to everybody and rarely called by these names. One society writes a rule at the gate: no vehicle enters without a sticker, and a visitor waits while the resident is called. Anybody can read that rule, including the visitor being turned away. The second society has had the same watchman on the gate for nine years, and he simply knows who belongs. He is right most of the time. Nothing about what he knows is written down anywhere, and it never was.
The first society needed one thing before the gate could work: somebody able to say, in advance, what the correct outcome is for every vehicle that will arrive. The second needed something completely different: nine years of vehicles arriving, and of being right or wrong about them, with somebody noticing which. The two requirements are the whole comparison, and everything below follows from them.
Rule-based automationAutomation whose behaviour is a written procedure a person can read line by line. needs the first thing. Somebody has to be able to write down the correct answer for the cases the step will meet, before it meets them. Machine learningA component whose behaviour was derived from a pile of past cases rather than written down by a person. needs the second. A pile of past cases has to exist, with the outcome attached to each one, and the outcome has to be something that was actually observed rather than something somebody would like to have observed.
At Sumeru Bank Limited both requirements were met in different places in the same loan intake chain. Component 5, the income corroboration rule, exists because somebody could write the correct outcome down: route the file when the declared monthly income runs ahead of the median monthly salary credit over three months of statement by more than the tolerance the bank chose for itself. The worked case declares Rs 45,000/- a month against a corroborated Rs 38,000/-, a gap of Rs 7,000/- which is 15.6 per cent of the declared figure, so the file is routed to a person. Component 6, the scoring model, exists because a pile existed: a past window holding 300,000 applications, of which 240,000 were accepted and therefore observable, and 8,160 of those carry a bad label on the bank's own definitions of what counts as bad and how long the outcome is watched for.
Notice what neither of those sentences says. Neither is a claim about how clever anything is. The requirement is a fact about the world before the project starts. A step can therefore rule one of the two options out years before anybody proposes it. A step nobody has ever performed has no pile of past cases, so the learned option is not available at any price. A step where the right answer genuinely depends on something nobody has written down has no stated in advanceThe correct outcome being known before the case arrives, rather than only after it plays out. answer, so the written option is not available either.
What does each one need before it can exist? Pick the pair that is right for both.
What does each one give a reviewer to look at?
The gate makes the difference concrete. To learn why a particular car was turned away by the first society, a reader consults the rule. To learn why a particular car was turned away by the watchman, there is nothing to consult. He can be asked, and he will say something, but what he says is a story about a decision he already made. The only thing that can actually be examined is what he has been doing: a month of the gate log, sorted.
The second difference is what a reviewer gets to look at, and it decides most real arguments. A written rule set hands a reviewer the lines it was given; a learned component hands a reviewer nothing to read and only its behaviour to examine. Neelima Rao, in the bank's risk function, read all 34 lines of the income rule in 25 minutes, and at the end of it she could point at the line that routed any one of the 602 files it routed that month. The behaviour reviewExamining what a learned component has been doing across many cases, in place of reading what it was told to do. of the scoring model took her eleven working days and produced nothing she could point at in the same way.
The loose version of the claim is wrong, so it is worth being precise about what survives on the learned side. Repeatability survives: the same file fed in twice produces the same answer, for as long as the fitted numbers stay where they are. But it survives as something to be tested rather than something to be read, and it ends the moment the component is fitted again. Readability does not survive at all. A finding on the written side names a line. A finding on the learned side names a pattern.
A reviewer sits down in front of a deployed component with a week to spend. Which of these can be read directly?
Who sets the expectations when the person affected asks why
The reason a reviewer can be shown matters most when the person being refused asks for one, and a lender making retail credit decisions in India sits under the Reserve Bank of India, at rbi.org.in, for its digital lending, outsourcing, customer data and record-keeping expectations. The expectations fall on the lender whichever of the two options sits behind the decision.
How does each one fail, and how would the failure be noticed?
A written rule set fails by saying something wrong, and then saying exactly the same wrong thing to every file that meets it. The sameness is uncomfortable and it is also a gift. The fault behind one complaint is sitting on a line, and it produced the same fault every previous time, so one complaint is a complete audit of a written rule. The bank found nine such faults across all four of its rule sets at the month 12 validation: nine lines out of 126, being 7.1 per cent, that either contradicted another line or could never be reached at all.
A learned component fails in the opposite direction, and the difference is not a matter of degree. It drifts. In this chain an upstream income field changed format on one channel in month 8, and monitoring did not flag it until month 9. Across those six weeks about 12,900 files were decided and 176 of them moved out of accept and into the referral band. Stated as a month, the approval rate slipped from 57.0 per cent to 55.6, and referrals rose from 391 a month to 508, a rise of 117 which is the same 176 files spread over the six weeks. Nothing about the component changed. The data arriving at it changed.
A deployer has to ask how the fault would have been noticed. Not from any single file. No single file looked wrong. A file that scored into the referral band was reviewed by a person and got a sensible answer. The fault existed only in the aggregate. A complaints desk could never see it, and only somebody comparing the shape of one month against the shape of the month before could. The asymmetry is why the two need different kinds of watching, and why a bank that has only a complaints desk has cover for one of them and none for the other.
A second failure shape on the learned side is worth naming. Cathy O'Neil built her 2016 book Weapons of Math Destruction around it: the errors do not fall evenly. When the bank reviewed 200 of the liveness rejections by hand, 31 of them turned out to be genuine applicants, being 15.5 per cent, and all 31 shared one condition. Every one was a low-light image taken on a low-specification handset. No individual rejection carried that information on its face. The shared condition appeared only when 200 of the rejections were put in a row.
A single customer complaint reveals a fault in one component of the chain. Which of the two was it, and why is that identifiable from the complaint alone?
What does each one cost to keep working, once it is running?
Two descriptions can both sound reasonable, so descriptions are useless. Put both costs in the same unit and the argument becomes something a person can lose. The unit the bank used is a working day of one reviewer, assumed at 420 minutes.
On the written side the cost is a pair checkComparing every line of a rule set against every other line, to find pairs that contradict each other or that make one unreachable.. The bank's own assumed rate is about 400 pairs a working day, a little over a minute a pair. Its four rule sets hold 126 lines between them, and here is what that costs when each set is checked against itself.
| Rule set | Lines | Pairs to check | Working days |
|---|---|---|---|
| Component 5, income corroboration | 34 | 561 | 1.4 |
| Component 7, the fraud rules | 61 | 1,830 | 4.6 |
| Component 9, the workflow router | 22 | 231 | 0.6 |
| Component 1, the identity match | 9 | 36 | 0.1 |
| All four, checked set by set | 126 | 2,658 | 6.6 |
The day column is rounded to one decimal, so it does not add exactly to the total; the total is worked from the 2,658 pairs. Now the number that surprises people. Four sets that interact have to be checked as one. Treating all 126 lines as a single set and checking every line against every other line takes the pairs to 7,875 and the cost to 19.7 working days. The extra 5,217 pairs are nothing but lines in one set being checked against lines in another. Where the boundary is drawn around a rule set changes its review cost by a factor of three without a single line being written or deleted.
On the learned side the cost was eleven working days, being 4,620 minutes on the same assumed day, spent on the fitting population, the label, the inputs and the recent behaviour. The behaviour review is one number with no lines column, for the plain reason that there are no lines.
One distinction is worth holding on to, and getting it wrong is what makes this comparison come out backwards. Reading a rule set is not checking it. Neelima Rao read the 34 lines in 25 minutes. Checking those same 34 lines against each other is 561 pairs and about 1.4 working days. The honest comparison is 1.4 days against eleven, not 25 minutes against eleven days, and people who quote the second version are comparing a read to a review.
Why bother converting both review costs into working days instead of describing each one in words?
At what size does the cheaper option to review change?
Here is the interesting part, and it is the reason this comparison is worth doing with numbers rather than instincts. The two costs behave completely differently as the problem gets bigger. Pairs of n lines are n times n less one, over two, so the pair check grows with the lines. The behaviour review does not grow at all. Eleven working days covered a component fitted on 240,000 applications, and about eleven working days would cover one fitted on a tenth of that. The work is examining the population, the label, the inputs and the behaviour, and none of those gets smaller in an interesting way.
Two costs, one growing and one flat, cross exactly once. Everything below that crossing pointThe size of problem at which the cheaper of the two options to review changes over. is cheaper to review as a written rule set, and everything above it is cheaper to review as a fitted component.
At roughly how many distinct casesOne situation the step has to handle differently from all the others, which needs about one line of its own in a written rule set. does a learned component become the cheaper of the two to review?
Work the two sides of it exactly. At 94 distinct cases a rule set of 94 lines makes 4,371 pairs. At 400 pairs a working day that comes to 10.9 working days, just under the eleven. At 95 lines it makes 4,465 pairs and 11.2 working days, just over. The crossing sits between 94 and 95 distinct cases, and every rule set in this chain is on the cheap side of it: 61 lines is the largest, at 4.6 working days against eleven. Move the control below and watch the two bars swap over.
Where does the cheaper review switch over?
One thing moves: how many distinct cases the step has to handle. Everything else is held: one line per distinct case, every pair checked, 400 pairs a working day, and a behaviour review of eleven working days whatever the component holds.
Static readings to check the control against: at 34 distinct cases the pair check is 561 pairs and about 1.4 working days against 11 for a behaviour review. The crossing is at about 94. At 126 lines the pair check is 7,875 pairs and 19.7 working days against the same 11.
At 34 distinct cases a rule set of 34 lines makes 561 pairs to check, about 1.4 working days, against 11 working days for a behaviour review, so the written option is 7.8 times cheaper to review at this size.
Why does the crossing sit so much higher than teams expect?
Ask a room to guess and most of the guesses land somewhere between fifteen and forty. The crossing is at 94. Three things produce that gap, and each is worth catching in yourself.
The first is the one from the previous block. People compare a read against a review. Twenty five minutes of reading against eleven working days flatters the written option and makes it look nearly free. Then somebody who has been burned by a badly written rule set overcorrects in the other direction. The comparison that matters is a check against a check.
The second is that people compare a large rule set against a small learned component, as though the flat cost shrinks for a modest one. It does not. Examining the fitting population, the label, the inputs and the recent behaviour is the same work whether the component was fitted on 240,000 applications or 24,000, and eleven working days is what the bank spent on it.
The third is that the flat eleven days is the shakiest number in the whole comparison, and it deserves saying plainly rather than in a footnote. The eleven days is one measurement, taken once, at one bank, on one component. The crossing is not a constant to memorise, it is an arithmetic to be redone with an institution's own two numbers. If a behaviour review at a given institution genuinely costs six working days, the crossing falls to about 70 distinct cases. If it costs twenty, the crossing moves out to about 127. The shape of the answer holds in every version: the crossing is much higher than instinct suggests, and almost every rule set anybody actually writes sits below it.
Where did one deployed chain actually put each of them?
Sumeru Bank Limited runs nine components in the intake chain. Five of them learn from data and four are rules somebody wrote. Counted that way it sounds like a balanced estate, four against five, and that is exactly how it was described inside the bank.
Count it as decisions instead. In one steady month at month six, 8,600 applications reached the decision engine. The four written components determined the outcome of 1,355 of them, being 15.8 per cent. The five learned components determined the other 7,245, being 84.2 per cent. Four against five becomes one file in six against five files in six, and which of those two counts a governance conversation uses decides how serious the conversation gets.
Then look at what kind of decision each side was making. Nobody in the bank had noticed this part. Component 5 routed 602 files, the fraud rules routed 452, the identity match routed 189 and the router itself held back 112 where the consent record was incomplete. The four counts add to the whole 1,355. Every one of the decisions the written components determined was a decision about where the file goes, and not one of them was a decision about the loan. On the other side, component 6 determined 5,981 outcomes, of which 4,902 were accepted and 688 were declined outright, and component 4 sent 1,264 to a person because a field could not be read. In this chain the written components decide where a file goes and the learned components decide what happens to the applicant.
Four written components and five learned ones. Does that mean the chain is roughly evenly split between the two?
What do the four written components have in common?
Line them up and one property is shared by all four. Every one of them sits at a place where somebody could state the correct outcome in advance. Two identity records either agree or they do not, and somebody can write which disagreements matter. A declared income either runs ahead of the corroborated one by more than the tolerance the bank chose or it does not. A fraud pattern is a pattern somebody decided to act on, written into 61 lines by people who could say why each line was there. And where a file goes next is a question with a correct answer given what is already known about the file.
A second property follows from the first without anybody planning it. All four are small. A place where the right answer can be stated tends not to need hundreds of distinct cases to describe it. The largest holds 61 lines and the smallest holds 9, and every one of them sits far below the crossing.
What do the five learned components have in common?
Four of the five share the mirror-image property. The input is a document, an image or a pattern nobody could sit down and enumerate. Nobody can write the rule for whether a selfie is a photograph of a live person, or for which of four kinds a document is, or for where the income field sits in a bank statement when every lender formats its statements differently, or for how to phrase the first draft of an explanation. In each of those the outcome could be checked once it existed, but the input could never be written out case by case.
The fifth is component 6, the scoring model, and it does not belong with the other four. Its input is not an image and not a document. Its input is a set of fields on an application form, all of them enumerable, all of them written down. The part of component 6 that cannot be stated in advance is not the input at all. The unstatable part is the outcome: whether this applicant will repay. Component 6 sits on the learned side for the opposite reason to the other four, and that difference is the whole of what went wrong in this chain.
What do the four written components in this chain have in common with each other?
What is the wrong way to choose between the two?
Two grounds decide most of these arguments in most institutions, and neither one survives contact with the four criteria set out above.
The first is choosing on performance when the two options are not both available. If nobody can state the correct outcome for a step in advance, comparing accuracy is comparing something that exists against something that cannot be built, and the written option loses an argument it was never in. If no pile of past cases with outcomes exists, the same thing happens the other way round. Availability comes first and performance comes second, and a room that starts with performance never gets back to availability.
The second is more subtle and it is the expensive one. A reason against one option gets read as a reason for the other, and it is not one. Somebody says nobody can write this down, everyone nods, and the learned option is adopted by default. But nobody being able to write it down is a statement about the step, not about either option. The sentence says the correct outcome cannot be stated in advance. Saying so rules the written option out. Saying so rules nothing in. The same sentence should have made everybody more careful about the learned option too. A component that cannot be checked against a stated correct outcome cannot be checked against a stated correct outcome, no matter how it was built.
The tell for this move is a sentence like it will produce an answer regardless. Producing an answer regardless is true of anything, including a coin, and it is not a reason.
The error that gets made, and what it costs
The bank scored all eleven steps of its old loan process on a five-part readiness scoreA scored check on whether a step can be handed to a machine safely, covering volume, stability, data, decidability and consequence. before it automated anything, out of ten. Eight steps scored seven or above and were automated. Step 8, the decision itself, scored four, and it was automated anyway because the business case needed it. The argument in the room was between a rule set and a learned component, and the learned option won on being able to handle cases nobody had thought of.
Read the score again. The score was four because of one test: nobody could state the correct outcome for that step in advance. The failed test ruled the rule set out, and was then treated as the case for the learned component. The failed test is a reason against both. The honest answer available in that room was that step 8 belonged with a person until somebody could say what right looked like.
Nobody in that room was incompetent, and overriding a readiness score is sometimes exactly right: a low score on a step with a small consequence, overridden knowingly and written down with what would make it reversible, is a normal and defensible decision. The step 8 override was none of those things. The override was taken for a stated reason that did not support the conclusion drawn from it.
The cost is still running. In one steady month 1,290 files reach an outcome the applicant would call a refusal: 602 routed by the written income rule, each with a printed procedure sitting behind it, and 688 declined by the scoring model, with no procedure behind them and only an attributed reason. Answering all 1,290 of those people by quoting the written procedure is wrong for 53.3 per cent of them, and every governance difficulty that follows in this chain traces back to that one override.
Nobody can state the correct outcome for a step in advance. What follows from that, on its own?
Is there a step where neither one is the answer?
Yes, and this chain has two of them. For anybody holding a list of steps and a budget, those two steps matter more than anything else in the comparison. Three questions, asked in order, place a step without an argument.
Can the correct outcome be stated in advance? If no, the written option is out. Do past cases with the outcome attached exist? If no, the learned option is out. Two noes and the step stays with a person, and no amount of enthusiasm changes that. Must one outcome be explainable to the person it lands on? If yes and only the learned option survived the first two questions, the honest placement is a person deciding with the component's output in front of them, rather than the component deciding.
The third arrangement has a name in the literature. Ajay Agrawal, Joshua Gans and Avi Goldfarb built their 2018 book Prediction Machines around exactly it: what a fitted component produces is a prediction, and somebody still has to decide what to do about it. Step 10 of the bank's old process, drafting and sending the letter, scored six and was placed exactly there. A draft is produced and a person signs it. Step 7, writing the assessment, scored three and was left with people untouched. Eight steps automated, one augmented, one left alone and one overridden makes eleven.
How does this apply to a step handed over tomorrow?
Take the household version first. The shape is identical and the stakes are low enough to see it clearly. A housing society is choosing how to run its visitor gate. Writing the rule down means every resident can read it, every turned-away visitor gets a reason, and the committee can be held to it. Leaving it to the watchman's judgement means better decisions on the odd cases and no way to answer the resident who asks why her mother was stopped. Neither is obviously right. The obviously wrong move is choosing without noticing that this is the trade being made.
The person being asked to approve a step puts the three questions in order and writes the answers down, then adds two more that usually get skipped. The first extra question is what a reviewer will be able to look at in eighteen months, and how many days that will take. The second is who finds out when this is wrong, and from what. A step whose honest answers are nothing much, eleven days, nobody, and a complaint that never comes is a step being approved blind, whatever the business case says.
For somebody reading another party's claim instead, as a lender's reviewer looking at an outsourced service, or as an analyst reading a company that says it has automated its operations, the useful question is never how much was automated. The useful question is which of the two is behind the decisions that reach a customer, and what the complaints desk can be shown when one of them is challenged. A company that can answer the second question has thought about this; a company that answers with a percentage has not. The percentage is a count of components, and that number moves once decisions are counted instead.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering digital lending, outsourcing, customer data, consent and record keeping | rbi.org.in |
| Securities and Exchange Board of India | Expectations where the deployer of such a system is a market intermediary | sebi.gov.in |
| Bank for International Settlements | International supervisory material on the deployment of such systems by banks | bis.org |
| Ajay Agrawal, Joshua Gans and Avi Goldfarb | Prediction Machines, 2018, on a fitted component producing a prediction that somebody still has to decide what to do about | Harvard Business Review Press |
| Cathy O'Neil | Weapons of Math Destruction, 2016, on the errors of a fitted component falling unevenly across the people it decides about | Crown |
Sumeru Bank Limited and Neelima Rao are invented.
Educational material. Not advice on any investment, tax, budget or market position.
