Human in the Loop and Human on the Loop Compared
Human in the loop means a person decides each case before anything takes effect. Human on the loop means nobody decides any case, and somebody watches the whole run of files and holds the authority to halt it. At Sumeru Bank Limited, invented, the first arrangement acted on 391 files in one month and the second was exercised once in six.
Both of those get called human oversight, and in most descriptions of a deployed system they arrive in the same sentence. Human in the loop and human on the loop are not two settings of one dial. The two are different jobs, and once it is clear which job the person has been given, everything else about the arrangement stops being a matter of opinion: how many people it needs, how quickly it can notice anything, what it costs in minutes, and which of the two things that can go wrong it is even capable of catching. The two arrangements are separated not by how closely anybody is watching but by whether a person decides the individual case.
What does human in the loop mean, and what does human on the loop mean?
Start with a wedding. At the gate stands a cousin with a printed guest list, and nobody walks in until the cousin has looked at the invitation card and looked at the list. The cousin is a person in the path. If the cousin steps away for ten minutes, the queue stops. Inside the hall, the caterer has a supervisor walking the length of the buffet all evening. The supervisor serves nobody and stops nobody at the entrance, and if a tray comes out cold or the paneer runs low, the supervisor can pull the whole tray and send it back. The supervisor is a person over the whole thing. Two people, two entirely different jobs, and only one of them can be told about a single guest.
Human in the loopA person decides each case before the outcome takes effect, so the file cannot move until they act. is the cousin at the gate. The system prepares something, a person decides it, and only then does the outcome exist. Take the person away and the work halts. The halt is the test: if pulling the person out stops nothing, they were never in the loop.
Human on the loopNobody decides any individual case, and a person watches the system across many cases with the authority to halt it. is the supervisor. Nobody is in the path. Files are decided and outcomes take effect at whatever rate they arrive, and somewhere off to the side a person is looking at what the whole run of them is doing, holding the authority to stop it. Take that person away and nothing halts at all. Nothing halting is the second half of the same test, and it explains why this arrangement can be quietly absent for a long time without anybody noticing. In the loop, the person is a step in the work; on the loop, the person is a condition on the work, and the difference shows up in whether anything stops when they do.
What separates human in the loop from human on the loop?
What is the person actually doing in each arrangement?
Take the two jobs apart by what lands in front of the person and what they are expected to produce. In the loop, one file lands. The person has the application, the documents, the reading the component produced and the reason it stopped, and what they produce is an outcome for that one applicant. Their unit is a file, their evidence is everything about that file, and their output is a decision somebody could later be asked to explain.
On the loop, no file lands at all. A patternWhat a person on the loop watches: many files read together as one shape, rather than any single case. lands instead, meaning many files read together: an approval share, a count of referrals, a spread of scores, a comparison against last month. The person is not deciding anything about anybody. The person is judging whether the shape of the whole has moved, and their output is not an outcome but a judgement about the arrangement itself. Two moves are available to them: let it run, or stop it.
The smallest thing a person can see is the smallest thing they can act on, and that single difference in unit sets everything else. A person holding one file can see that this applicant declared a salary of Rs 45,000 against a corroborated Rs 38,000 and think about what that means for this application. A person holding a monthly reading cannot see that applicant at all, and no amount of attention will make them visible. A reading built from thousands of files does not contain any one file in a form a human eye can find.
Which files at this bank got each arrangement?
At Sumeru Bank Limited, invented, the split is not a description. The split is a count, and it is exact. In one steady month the scoring component returned a value on 5,981 files. Of those, 391 came back inside the referral bandThe range of scores where the component declines to decide, so a person has to., the range where the component declines to decide, and every one of them went to a person who made the credit decision. The 391 files are the in-the-loop arrangement, entire. In share terms the 391 are 6.5 per cent of the scored files and 13.0 per cent of everything the exception desk saw that month.
The other 5,590 files were decided and answered without a person reading any of them: 4,902 accepted and 688 declined, and 4,902 plus 688 is 5,590. The 5,590 files are the on-the-loop arrangement, entire. The on-the-loop files are 65.0 per cent of the 8,600 that reached the decision engine, and 391 plus 5,590 is 5,981, so every scored file in the month sits in one arrangement or the other and none sits in both. The month's remaining 2,619 files left the automatic route before the component ever scored them, and where they went is set out under credit decisioning systems. Write the two counts side by side and the phrase human oversight stops being one idea: the arrangement most people describe covers 6.5 per cent of the month, and the other 93.5 per cent has a different arrangement with a different name doing a different job.
391 files a month go to a person and 5,590 do not. Which arrangement covers which?
What does the in-the-loop arrangement cost in minutes and posts?
Only one of the two arrangements has a price, and the price can be built out of numbers somebody measured for another purpose entirely. The comparison becomes checkable at exactly that point. A file at the exception desk at Sumeru Bank Limited takes 19 minutes on average, and that 19 is not a lump. The 19 breaks into 4 minutes re-reading what the chain already produced, 6 minutes finding whatever is missing, 5 minutes obtaining it from the applicant or another system, 3 minutes deciding, and 1 minute recording, and 4 plus 6 plus 5 plus 3 plus 1 is 19.
From month 5, every referral band file at this bank carries a second reviewerA further person put on the same case, used where the first person is deciding rather than fetching information., a design choice of the bank's own and nobody's requirement. Ask what that second person actually does and the cost falls out of the breakdown rather than out of a new measurement. The information is already sitting there by the time the second reviewer sees it, so the finding and the obtaining are not repeated. The second reviewer repeats the re-reading, being 4 minutes, and the deciding, being 3, and 4 plus 3 is 7. So 391 files times 7 minutes is 2,737 minutes a month, and at the assumed working month of 8,400 minutes a person that is 0.33 of a post.
Put the whole arrangement on one line. The first pass is 391 times 19, being 7,429 minutes and 0.88 of a post. The second reviewer is 2,737 minutes and 0.33. Together the in-the-loop arrangement is 391 times 26, being 10,166 minutes a month and 1.21 posts, and 7,429 plus 2,737 is 10,166 either way it is built. Every one of those minutes is arithmetic on figures somebody already measured. A number built that way is arguable instead of merely asserted. If the second reviewer really repeats the finding as well, that can be said and the total recomputed; no such recomputation is possible against a cost nobody decomposed.
Where do the second reviewer's seven minutes come from?
What does the on-the-loop arrangement cost, and why does it look free?
Now do the same exercise on the other side and watch it fail. The exception desk at this bank is 7 people, and that 7 is built from 3,010 exceptions a month at 19 minutes each, being 57,190 minutes, over the assumed working month of 8,400. Every one of those minutes belongs to a file a person handled. Not one of them is the arrangement standing over the 5,590 files nobody handled. Search the desk sizing for the watching and it is not a small line; it is not a line.
The missing minutes are not a scandal and they are not sloppiness. The missing minutes are what happens when a cost has no unit. The in-the-loop cost has a natural unit, being minutes a file, so it multiplies straight into a headcount and lands in a budget. The person on the loop does not touch files, so the on-the-loop cost has no per-file unit at all. The on-the-loop cost consumes somebody's attention on a periodic reading, the work of producing that reading, and the standing capacity to act on it, and none of those three divide neatly by 5,590.
An arrangement whose cost has no unit does not get costed at zero by anybody's decision; it simply never reaches the sheet where costs are written down. Think of a housing society. The lift attendant is on the muster roll at a wage everybody can name. The person who is supposed to notice that the lift has been jerking for a fortnight is nobody in particular, costs nothing anybody has written down, and the first time the arrangement is tested is the day the lift stops between floors.
Why does the on-the-loop arrangement carry no minutes in this bank's desk sizing?
How often was each arrangement actually exercised?
An arrangement that exists on a diagram and an arrangement that has been used are two different things, and the way to tell them apart is to count occasions. An interventionAn occasion on which a person actually changed or halted what the system would otherwise have done. is an occasion on which a person actually changed or stopped what would otherwise have happened, and it leaves a trace that can be counted.
In the loop, the count for one steady month is 391. Every one of those files reached an outcome because a person read it and decided it, and every one of them left a record with a name on it. In that same month, the on-the-loop arrangement standing over 5,590 files was exercised zero times. Widen the window to the whole six months from go-live in month 4 to month 9 and the count becomes one. The single occasion came in month 9 week 3, when one channel fell back to manual decisioning for 4 working days, covering 1,720 files at the bank's rate of 430 files a working day.
The two windows are not the same length, and the case does not lock what the earlier months looked like, so be careful setting one against the other. Month 6 is the steady month; month 4 was go-live on a single channel and month 5 was the extension to the rest, so there is no honest six-month total to set against the one. The safe comparison is 391 occasions in one month against a single occasion in six, and within the one month where both counts are known it is 391 against nothing at all.
The in-the-loop arrangement acted on 391 files in one month. How many times was the on-the-loop arrangement exercised in six?
What does a person on the loop need before they can act at all?
Naming somebody is the cheapest part of this arrangement and the part firms do first. Four things have to be true underneath the name before that person can do anything at all, and a firm can have the name in place and none of the four.
The first is standing to stopThe authority to halt an automated system, resting with a named person, settled in advance rather than argued at the time. it, settled in advance and written down. An authority that has to be argued for on the day is not an authority. The second is a reading that moves before harm has finished accumulating. A person told about last quarter cannot act during this one. The third is a fallbackThe route work takes when the automated one is halted, which has to exist on the day it is needed. that exists on the day. Stopping a system with nowhere for the work to go is not a decision anybody will make twice. The fourth is a record showing what the component would have done differently. A person who cannot compare cannot judge.
Hold this bank's arrangement against those four. The standing to stop is precisely item 8 of the nine item credit decision record, being what would have to be true for the component to be stopped and who may stop it, and at the month 12 validation item 8 was one of the four items not documented. The reading took six weeks, a length worked through below. The fallback did exist and was used, running one channel manually for 4 working days over 1,720 files, and that is the one part of the arrangement this bank can show worked. The fourth is not among the nine items the validation examined at all, so nothing is on record either way. Three of the four are things a firm can only demonstrate before it needs them, and the only one this bank demonstrated is the one it happened to use.
A firm names a person on the loop and sends them a monthly report. What is still missing?
How does each arrangement fail, and why do the two failures look nothing alike?
The in-the-loop failure is well known and it is a failure of decay. The person stays in the path, the files keep passing through them, and over time the reading becomes a glance and the decision becomes a signature. The arrangement is intact on the diagram and hollow in practice, and how that happens is set out under automation and augmentation. The shape of that failure matters here: the trace keeps being produced every single day, and the failure hides inside a record that looks healthy.
The on-the-loop failure has the opposite shape, and it is not decay at all. The on-the-loop failure is a limit built into the job. A person on the loop cannot see one file going wrong, ever, and no degree of diligence changes that. Their unit is the pattern, and one file does not move a pattern. Put a number on it with this bank's own volumes. The month's accept share is 4,902 of 8,600, being 57.0 per cent. One file moving from accept into referral changes that share by 0.012 of a percentage point. The move that eventually did get flagged was a fall from 57.0 per cent to 55.6, being 1.4 points, and 1.4 points at this volume is 117 files going the same way. So roughly a hundred and twenty files have to move together before the reading a person on the loop is watching shifts as far as the shift that was actually noticed.
Calling the on-the-loop failure inattention therefore gets it exactly backwards. The person did not miss a signal. There was no signal of the size they were looking at until enough files had already been decided to make one. Asking them to catch a single wrong file is asking for something the arrangement is not built to deliver, and the fault sits with whoever designed the arrangement to be relied on that way, never with the person standing in it.
One file is decided wrongly. Which arrangement could have caught it?
Why did the on-the-loop failure take six weeks to appear?
Follow the one episode this bank has. In month 8 week 2, an upstream income field changed format on one channel. Nothing about the component changed: the same fitted numbers, the same cut-offs, the same code. The change was in what arrived in one field, and the effect was that files began scoring lower and sliding out of accept into the referral range. Monitoring flagged it in month 9 week 3, six weeks later, and the channel was put on the manual route for 4 working days while the reading step was corrected.
Six weeks at this bank's rate is 30 working days at 430 files a day, which is about 12,900 files decided inside the window, and 176 of them moved from accept into the referral band. Check that against the steady state and it holds: at 8,600 files a month the accept share fell from 57.0 per cent to 55.6, being a fall of 117 accepts a month, and referrals rose from 391 to 508, which is the same 117. Six weeks is a month and a half, and 117 times 1.5 is 175.5, which rounds to the 176 the case records. The window is not an anecdote, it is a length, and everything about the on-the-loop arrangement that matters is a property of that length rather than of anybody's effort inside it.
Why did it take six weeks to notice that the income field had changed format?
The mistake that gets made, and what it costs
The error is answering the question about human oversight with one arrangement and letting it stand for both. The error is made in good faith and usually by somebody who can describe the referral route accurately, in detail, with the second reviewer and the 19 minutes and the name of the person who signs. Every word of that is true. The trouble is that it describes 6.5 per cent of the month and is offered as an answer about all of it.
Follow the cost through. The 391 files get a person, a second person, 26 minutes and 1.21 posts of funded capacity, and every one of them leaves a record with a name attached. The 5,590 files get an arrangement with no minutes in any costing line, no per-file trace, and one exercised occasion in six months. Both are real arrangements. Only one of them has ever produced evidence about itself, and the one that has is the smaller one by a factor of fourteen.
Then there is what the single occasion means, said as arithmetic rather than as alarm. One exercise in six months is one observation. One observation shows that the fallback worked on that day, and that is worth knowing, more than many firms can show. One observation does not show how often the arrangement would work: a rate needs occasions, and there is one. At that rate, ten occasions would take five years to accumulate, so nobody at this bank is ever going to hold a reliability figure for it. An arrangement exercised once in six months is not an arrangement with a known reliability, it is an arrangement nobody has yet had the chance to test, and the honest thing to write next to it is that sentence rather than a number.
Why is neither arrangement an optional extra?
Reading both as additions is tempting: build the system, get it working, then decide how much human oversight to bolt on according to appetite and budget. Reading them that way is what produces an arrangement nobody funded, and it fails on a test that has nothing to do with appetite. The test is whether somebody will one day have to explain one outcome to the person it was about.
Where the answer is yes, a person has to stand in the path of that outcome, because an explanation of a single case can only be produced by somebody who looked at that single case. Where the answer is no, and only the shape of the whole matters, a person stands over the whole. The choice is not a preference; it follows from what each arrangement is capable of producing.
Now apply the test honestly to this bank, and it does not flatter anybody. Of the 5,590 files on the loop, 688 were declines, and a decline is exactly the outcome somebody may later have to explain to the person it was about. The 688 sit in the arrangement where nobody read the file. The placement is a fact about one invented bank's own design rather than a requirement of anybody, and what the applicant was actually told is set out under adverse action. Structurally the two arrangements are not ranked by strength but selected by what the outcome will later have to be able to do, and getting that selection wrong is not fixed by watching harder.
Is a person in the loop an extra that a deployer adds if it wants to?
What does each arrangement need in the record before it runs?
Both arrangements were running at Sumeru Bank Limited and both were working in the ordinary sense: files went to people, people decided them, a fallback existed and was used once. The month 12 validation, done by Neelima Rao, who built no part of the chain, checked the nine item credit decision record and found 5 of the 9 documented, being 55.6 per cent. Two of the four that were missing are precisely the two arrangements set out in this guide.
| Record item | Which arrangement it governs | At month 12 |
|---|---|---|
| Item 5. What happens to a file the component declines to decide | The in-the-loop arrangement, being the 391 | Not documented |
| Item 8. What would have to be true for the component to be stopped, and who may stop it | The on-the-loop arrangement, being the 5,590 | Not documented |
| Both arrangements were running anyway | 391 files a month and 5,590 files a month | Neither written down |
Read that table once and the symmetry is obvious. Read it twice and the asymmetry underneath it is the thing worth taking away. Both arrangements were unwritten, and only one of them had six months of daily evidence standing behind it anyway. The referral route could be reconstructed from the record of 391 decided files a month even with item 5 blank, because the work itself left a trail with names and times on it. The on-the-loop arrangement had item 8 blank and one occasion in six months behind it, so with the record silent there was almost nothing else to read it off.
The practical reason a record item matters more for one arrangement than the other has nothing to do with which is more important. A commitment written in advance is the only evidence that exists for an arrangement which, when it is working perfectly, produces no events at all. The in-the-loop arrangement generates its own evidence as a by-product of doing the work. The on-the-loop arrangement generates evidence only when something goes wrong, so on a quiet day the writing is all the evidence there is that it exists.
Why the difference is not a quantity
The difference between the two arrangements is not a quantity: a person either decides the individual case or nobody does, and there is no setting between the two. No continuum runs from one arrangement to the other, and the detection lag that makes the on-the-loop arrangement expensive is set out under monitoring, where it can be held properly rather than half held in two places. The month at its real widths replaces any continuum: 391 files in one arrangement and 5,590 in the other, with nothing between them.
How would somebody reviewing an arrangement like this tell the two apart in an hour?
The two arrangements can be told apart from outside, from a risk function, from a credit committee, or from across the table from a lender explaining its process. Neither the model documentation nor an understanding of how anything was fitted is needed. Four questions are enough, and each one has a number as its answer rather than a paragraph.
Ask first: on how many files does a person decide before the outcome exists? Not oversee, not review, decide. At this bank the answer is 391 a month out of 5,981 scored, so 6.5 per cent, and any answer given as a percentage of a percentage or as a description of a governance forum is not an answer to this question. Ask second: how many minutes a month are funded for it, and where do those minutes come from? Here it is 10,166 minutes and they come from a measured breakdown, so the figure can be argued with. Ask third: how many times in the last six months did anybody exercise the arrangement standing over the automatic files, and what did they do? One occasion, month 9 week 3, four working days, 1,720 files. Ask fourth: how long did the reading take to move? Six weeks.
Four numbers, and between them they say more about an arrangement than any amount of description, because each one is either a count somebody can produce from a system or an admission that nobody counted. A firm that answers all four crisply may still have a badly designed arrangement, and the argument will at least be about the design. A firm that cannot answer the third question has said something important without meaning to, which is that whatever is standing over the automatic files has never left a trace, and nobody has yet found out whether it works.
Where the expectation on human oversight of an automated decision sits
Where a regulated lender relies on a person deciding, or on a person standing over an automated decision, the expectations covering digital lending, fair practice, outsourcing, data, consent and the treatment of a borrower are set by the Reserve Bank of India and published at rbi.org.in. Where the deployer is a market intermediary rather than a bank, the equivalent expectations sit with the Securities and Exchange Board of India at sebi.gov.in, and the accountability of a board for what its systems do sits under company law administered by the Ministry of Corporate Affairs at mca.gov.in. Read all three at source and confirm them there. The referral range, the second reviewer added in month 5, the nine item record and every count here are one invented bank's own arrangements rather than anybody's standard.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Published expectations on a regulated lender covering digital lending, fair practice, outsourcing, the treatment of a borrower and the use of data and consent, the body of expectations under which accountability for a person standing over an automated decision ultimately sits | rbi.org.in |
| Securities and Exchange Board of India | The equivalent expectations where the deployer of an automatic decision arrangement is a market intermediary rather than a bank | sebi.gov.in |
| Ministry of Corporate Affairs | The accountability of a board for what a company does, under which the standing to halt an automated arrangement ultimately reports | mca.gov.in |
| Ajay Agrawal, Joshua Gans and Avi Goldfarb | Prediction Machines, 2018, for the separation between what a fitted component produces and the deciding a person still has to do with it | Harvard Business Review Press |
Sumeru Bank Limited and Neelima Rao are invented.
Educational material. Not advice on any investment, tax, budget or market position.
