Automation vs Augmentation: Replacing Work and Assisting It
Automation replaces a step: the system produces the outcome and nobody sees it before it takes effect. Augmentation assists a step: the system produces something a person then acts on, changes or rejects. The difference is not how capable the component is. The difference is whether a person stands between the output and its effect, and that one question decides who answers for the result.
Two arrangements only need separating when something turns on the separation, and here one large thing does. Accountability for an outcome follows the path the outcome took, not the technology that produced it. Drawn on a whiteboard, the difference between the two paths is obvious and slightly boring. Watched over six months in a working process, that same difference becomes the hardest distinction in the subject to hold on to. A person who accepts ninety nine outputs in a row is not reviewing the hundredth in any way a control could rely on. Assistance turns into replacement without anybody deciding that it should, and the record still says a person signed.
What does automation actually mean, before it is set against anything?
A step is automatedThe system produces the outcome and nobody sees it before it takes effect. when the outcome it produces takes effect without a person seeing it first. The definition ends there, and notice what is not in it. Nothing about how the outcome was produced. Nothing about whether the thing producing it was written by somebody or fitted to past examples. Nothing about how clever it looks in a demonstration. Completeness is about the path and not about the machinery. A written rule of nine lines automates a step exactly as completely as a component fitted on three hundred thousand records.
Automation is already part of ordinary life and is not exotic. A standing instruction on a savings account moves money on the fifth of the month. Nobody approves it that month. Nobody reads it before it goes. Somebody set it up once, and the setting up was the last human decision in the sequence. A prepaid electricity meter cuts supply at a balance of zero without a supervisor confirming that this particular household should be cut off today. In both cases a person made one decision at the start and then stepped out of the path entirely.
At Sumeru Bank Limited, an invented bank running a retail loan intake chain, 8,600 applications reached the decision engine in the month measured. Of those, 5,590 were decided with no person touching the file at all, or 65.0 per cent. The 5,590 outcomes are automated in the exact sense above. An applicant received an answer in about four minutes. No member of staff read that file before the answer went out, and none was meant to. Automation is defined by the absence of a person in the path, never by the sophistication of what fills the path.
What does augmentation mean, on exactly the same terms?
A step is augmentedThe system produces something a person acts on, changes or rejects before it takes effect. when the component produces something, a person acts on that something, and only then does anything take effect. The person can pass it through unchanged, change it, or throw it away and start again. Three outcomes, and the arrangement is only doing its job if all three are genuinely available. Suppose there is no time, no alternative and no procedure for rejecting. Then the person cannot in practice reject the output, the third outcome is decorative, and so is the arrangement.
The everyday version sits in the hand. Predictive text on a handset produces a word, and the sender presses send. If the word was read first, that is assistance and the sender is accountable for the message. If send is pressed without reading, the picture on the screen is identical, the arrangement has quietly become something else, and the message that goes out is whatever the component produced. Nothing visible changed between those two cases. The whole difficulty of the subject sits in that gap, and it is worth holding on to before any of the finance arrives.
In the intake chain, step 10 of the old eleven step process, the decision letter drafted and sent, is arranged this way. A drafting component produces a first draft of the letter and a person signs it before it leaves the bank. Agrawal, Gans and Goldfarb, in Prediction Machines, 2018, put the general shape of this well: a fitted component produces something, and a person still has to decide what to do about it. An augmented step therefore contains two outputs, what the component produced and what the person let out, and only the second one ever reaches anybody.
What is the one question that separates the two arrangements?
Does a person stand between the output and its effect? That is it. Ask it of any step in any process and the answer sorts the step into one of the two arrangements, and no second question is needed. A rule set of nine lines with a person signing every result is augmentation. A component fitted to three hundred thousand past applications with nobody looking is automation. The question is about the path, and the path does not care what produced the thing travelling along it.
The trap is that most people sort by capability instead. Sorting by capability puts simple, mechanical work in the automation column and difficult, judgement-carrying work in the augmentation column, as though the two words described how hard the task was. Capability feels like the natural axis, and it is the wrong one. The capability axis produces the sentence a review committee hears most often. On that account a step is only assisted because the component is not clever enough yet, and the person will drop out once the component improves. The person does not drop out because the component improved; the person drops out because somebody decided the output could take effect unseen, and that is a decision with a name on it.
What single question separates automation from augmentation?
Who answers for the output, in each of the two arrangements?
In an automated step, nobody is in the path at the moment the output takes effect, so accountability sits entirely upstream. Accountability sits with whoever approved the arrangement, whoever maintains it, and whoever watches it. At Sumeru Bank Limited, Revathi Balan, head of retail credit, is the named accountable person for the scoring component. Being named means she answers for outcomes she did not see and could not have seen. The design she approved is what decided that nobody would see them. Answering for the unseen is not a loophole. Upstream is the correct place for the accountability to sit, and that is the reason approving an automated step is a heavier act than approving an assisted one.
In an augmented step, accountability splits in two and only the first half usually gets recorded. The person in the path answers for the specific output that went out under their name. The design answers for whether that person could possibly have reviewed it. If the design put a person in the path and then gave that person the same minute they had before, with the same queue and the same target, the design has arranged for a sign-offThe act of a person accepting an output, which is evidence of a review only where the rate of change or rejection is recorded. and called it a review. The person in the path answers for the output only to the extent the design paid for the review that would have caught the problem.
A component produces a letter, a person signs it, and the letter turns out to be wrong. The reviewer had 40 letters and 40 minutes. Who answers for it?
What does each arrangement need to have in its record?
The intake chain writes an audit trail with six fields on it: which version of which component acted, what it read, what it produced, what the previous version would have produced where that is known, who could have intervened and did not, and the time. Four of the six were recorded from go-live. Fields four and five were added in month nine, after an episode in which nobody could say afterwards which files would have been decided differently. The six fields are a good record of an automated step, and they were designed for one.
An assisted step needs something none of the six holds. Field five records who could have intervened and did not. For an automated step where a person had the standing to stop it, that is exactly the right field. Field five says nothing about a person who did intervene, and nothing about what they changed when they did. The record therefore shows a signature on every output and holds no way of telling a reviewer who changed one output in ten from a reviewer who changed none at all in a year. An assisted step needs one field the six do not carry. The missing field is what the person did with what the component produced, recorded as accepted unchanged, changed, or rejected.
| What the record has to hold | Automated step | Assisted step |
|---|---|---|
| Which version of which component acted | Required | Required |
| What it read and what it produced | Required | Required, and the draft has to be kept, not overwritten by the signed version |
| The time | Required | Required, including how long the review took |
| Who could have intervened and did not | Required | Somebody did intervene, so this is not the field that matters |
| What the person did with the output | Not applicable | The field that makes the arrangement checkable, and the one most often absent |
An assisted step keeps only the signed final version of every output and overwrites the draft. What has that record lost?
How does each of the two fail, and why do the failures look nothing alike?
An automated step fails in bulk and it fails identically. Whatever is wrong is wrong for every file that meets the same condition. The failure therefore has a population, a start date and a count. In the intake chain, an upstream income field changed format on one channel in month eight, monitoring flagged it in month nine, and about 12,900 files were decided in that window with 176 of them moving out of acceptance into the referral band. Every one of those 176 failed in the same way for the same reason. Bulk failure is unpleasant, and it is also the most findable kind of failure there is. Anything that happens to thousands of files at once shows up in a rate.
An assisted step fails one output at a time, in whatever particular way the review happened not to catch on that occasion. There is no population. The two outputs that went out wrong this quarter have nothing in common with each other except that both went out. A monitoring rate cannot see them: two in two hundred does not move any average anybody watches. The only thing that finds them is going back and pulling a sample of the outputs the reviewer accepted and checking those against the underlying file. An automated failure is a population that can be counted, and an assisted failure is a sample that somebody has to go out and draw.
How does an automated step fail differently from an assisted one?
Which arrangement is actually in force, once it is running?
The answer comes from counting how often the person disagrees. The disagreement rateHow often the reviewing person changes or rejects what the component produced, counted against everything they saw. is the number of outputs the reviewer changed or rejected, over the number of outputs they saw, and it is the only measure that tells the two arrangements apart in a running process. Everything else available describes the arrangement on paper. A procedure describes it. A signature block describes it. An organisation chart with a review step in it describes it. The disagreement rate describes what happened.
A reviewer who changes one output in five is reviewing, obviously and unarguably. A reviewer who changes one output in two hundred is signing, and the point on which everything else turns is this: the second is not a stricter, better version of the first. Signing is a different arrangement wearing the paperwork of reviewing. Somewhere between the two the step stopped being assisted, and nothing in the process announced the crossing. From the inside every day looked the same as the last one. The measure that separates the two arrangements is not in the design document, it is in what the reviewer did last month, and if nobody counted it then which arrangement is in force is simply unknown.
A reviewer signs 199 of 200 outputs unchanged. Automated or assisted?
Why is good assistance the thing that turns assistance into replacement?
Because nothing goes wrong. Nothing going wrong is the entire mechanism, and the mechanism is worth stating slowly. Every instinct pulls the other way. If the component produced rubbish, the reviewer would read every line, the disagreement rate would sit high, and the arrangement would stay exactly what it was designed to be. The component being right, output after output after output, is what teaches everyone involved that reading the next one closely is time spent for no return. Review decayThe drift from a genuine review to a signature, caused by the output usually being right rather than by anybody being careless. is the rational response to a hundred correct outputs in a row, not a lapse.
Anybody who has done this knows it did not feel like a failure of character. The first three months of a mobile bill get checked line by line. In the eleventh year the total gets a glance. Nothing about the person changed. Ten years of correct bills taught that person where to spend attention, and attention is finite and has somewhere else to be. A desk under a queue target does the same arithmetic, faster, and the arithmetic is not wrong. Review decay is caused by the assistance being good. Good assistance cannot be corrected by asking people to be more careful, so any design relying on a review has to measure the review rather than exhort it.
What causes review decay?
Where does a step carrying judgement belong, and why?
One question comes before any design work: will somebody eventually have to explain one of these outputs to the person it was about? If the answer is yes, the step needs a person between the output and its effect, and the reason is not sentiment about human oversight. The reason is that an explanation given afterwards has to come from somewhere, and in an automated step the only thing available afterwards is what the record kept. If nobody looked at the file, nobody can say more about that file than the record already says.
The intake chain scored each of the eleven steps of the old process on five tests, each marked nought to two for a maximum of ten: how often the step happens, how stable the rule behind it is, whether the input arrives in a form a machine can read, whether the correct outcome can be stated in advance, and what happens when it is wrong. Step 10, the letter, scored six and was assisted. Step 8, the decision itself, scored four and was automated anyway. The business case needed it. Every governance difficulty that shows up later in this bank's story traces back to that one override, and an override is what happens when a score is treated as something to be argued with rather than something to be believed.
| What happened to the step | Steps | The reason recorded at the time |
|---|---|---|
| Automated outright | 8 | Scored seven or above on the five tests |
| Automated anyway, against the score | 1 | Step 8, the decision, scored four, and the business case needed it |
| Assisted: a draft produced, a person signs | 1 | Step 10, the letter, scored six |
| Left with people entirely | 1 | Step 7, the written assessment, scored three |
| The eleven steps of the old process | 11 | 8 plus 1 plus 1 plus 1 is 11 |
A step will produce outcomes somebody has to explain to the person affected. Which arrangement does it need?
Before the next block. Of 200 drafted notes, 23 carried a statement that was not in the file. How many of the 23 did the signing person catch?
What happened to the signing at one bank, measured?
Sumeru Bank Limited checked 200 drafted exception notes against the files they were written from. The bank did the check for another purpose entirely, and the finding that matters most fell out of it by accident. Of the 200 drafts, 23 carried a statement that was not in the file, or 11.5 per cent. The signing person caught 21 of them. Two did not get caught and reached a customer. 21 plus 2 is 23, and both halves of that sentence belong in the same breath. The arrangement neither failed nor was vindicated.
Read the 21 first. Each of the 21 was wrong in its own particular way. Twenty one times, a person in the path stopped something from reaching a customer that no monitoring rate would have found. The 21 are the arrangement doing exactly the work it was designed to do, and they are the reason the step was assisted rather than automated in the first place. Then read the 2. Two outputs went out carrying a statement that was not in the file, under a signature, through a control everybody in the building believed in.
The 21 changes in 200 also produced this bank's only measurement of a disagreement rate, at 10.5 per cent. The rate is a floor rather than the whole of it, counting only the changes made for that one reason and not whatever else the reviewer altered. On the evidence, the arrangement at this desk was on the reviewing side of the line, and the uncomfortable part is that nobody knew that at the time they relied on it.
The same arrangement had already been used to settle a scoping question. The bank scoped its controls by a consequence testScoping by what an output does to somebody rather than by how the thing producing it was built., asking whether an output reaches a customer or a reported figure without a person deciding. Seven of the nine components in the chain came in scopeCovered by a policy, so its controls, approvals and records apply to the thing named. on that test. The drafting component was one of the two left out, and the reason recorded was that a person signs every output. A signature on every output is a control. At the moment that reason was written down, nothing had ever measured the control.
A scope leaves a component out on the ground that a person signs every output. What has to exist for that reasoning to hold?
The error that gets made, and what it costs
The record at Sumeru Bank Limited said a person signed every decision letter, and it was true. Every row of it was true. The record held no field at all for how often that person changed anything, and without that one number a reviewer who rewrites one letter in ten and a reviewer who has changed nothing since go-live produce exactly the same evidence. The bank had no measure of the disagreement rate at all until the note review produced one by accident, and by then the scoping decision that relied on the signing had already been taken and written down.
The fault is easy to put in the wrong place, so note carefully where it sits. The fault is not with the person signing. The signing person did the work the design asked for, caught 21 of the 23 unsupported statements in the sample, and had never been given either the minutes or the instruction that a fuller review would require. The failure belongs to a design that funded a signature and then described it, in a document that other decisions rested on, as a review.
The cost was not the two letters, unwelcome as those were for the two people who received them. The cost was that a control the whole building believed in had no evidence behind it, and nobody had ever asked for any. When the evidence finally arrived it happened to be favourable. Finding out by accident is the part worth losing sleep over. A bank that learns by accident that its control was working would have learned the same way if the control had not been.
How does an assisted step stay assisted?
Three things, and a signature is none of them. Record the disagreement rate. Count how many outputs the reviewer changed or rejected out of how many they saw, and put that count somewhere a person outside the process reads it. Sample what was accepted. The accepted pile is the only place an assisted failure can hide, so pull a handful of the outputs that went out unchanged and check those against the underlying file. And fund the review time explicitly, in the handling time the desk is measured on. Reading the output then becomes work the process has paid for rather than work squeezed between two other things.
The intake chain does one version of this well and it is worth naming. The example shows the shape of a design paying for review where the judgement actually sits. From month five, a second reviewer is given only to the files sitting in the scoring component's referral bandThe range of scores where a component returns neither an acceptance nor a decline, so a person takes the decision instead., being 391 of the 3,010 files routed to a person that month, or 13.0 per cent. The reasoning behind that choice is exact: those 391 are the only exceptions where a person is making the decision rather than supplying information the chain was missing. Review effort was spent where a person holds the judgement, rather than spread evenly across every file a person happens to touch.
Who sets the expectations behind a control like this one
Where a regulated lender relies on a person's sign-off as a control, the expectations that apply sit with the Reserve Bank of India, whose material on outsourcing, digital lending and record keeping is published at rbi.org.in and must be read there. Where the deployer is a market intermediary rather than a lender, the Securities and Exchange Board of India at sebi.gov.in is the body to read. The current position should be confirmed at the source before any of it is acted on.
Why is this a fork rather than a spectrum?
How this gets used, by three people who never meet
Whoever signs off a business case reads it as a cost question before it is a control question. An assisted step is not a cheaper automated step. An assisted step carries the build cost, the running cost and the review minutes, permanently, and a case that shows a step as assisted while showing no review time in the headcount has priced something it has not budgeted for. At Sumeru Bank Limited, the chain runs at Rs 65,00,000 a year against Rs 45,00,000 a year of posts saved, a shortfall of Rs 20,00,000 a year before the Rs 2,40,00,000 build is counted at all, and unfunded review time is exactly the sort of thing that hides inside a gap like that.
Whoever runs the desk reads it as a queue question. Ismail Sheikh, who heads the exception desk in this bank's story, has seven people where there used to be twelve. Each remaining file costs 19 minutes against the old 11, and the reason is that the easy files stopped arriving. Any review minutes added to an assisted step land on that same 19, and a design that adds a review without adding a minute has moved the cost onto the queue rather than removing it.
And anybody handed a process diagram can test it in five minutes without knowing anything about the technology. Find every box where an output leaves the system. For each one, ask who saw it before it left, then ask what the record holds about what that person did. If the answer to the second question is a signature and nothing else, the diagram is describing an arrangement the process cannot demonstrate it has, and the honest repair is either to measure the review or to redraw the box as automated and put the controls an automated box needs around it.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering outsourcing, digital lending, record keeping and reliance on internal controls | rbi.org.in |
| Securities and Exchange Board of India | Expectations where the deployer of such an arrangement is a market intermediary | sebi.gov.in |
| Bank for International Settlements | International supervisory material on the deployment of automated decision processes by banks | bis.org |
| Agrawal, Gans and Goldfarb | Prediction Machines, 2018, on a fitted component producing something a person still has to decide what to do about | Harvard Business Review Press |
Sumeru Bank Limited, Revathi Balan and Ismail Sheikh are invented.
Educational material. Not advice on any investment, tax, budget or market position.
