Stopping an Automated System: The Kill Switch and the Fallback
Stopping happens at three levels: one file overridden, one population moved to people, or the whole arrangement halted. A stop is six actions rather than a switch, and the expensive one decides what happens to the files already decided. A fallback is a capacity claim: at one invented bank the desk could carry 62.2 per cent of the volume handed to it.
Picture a toll plaza at seven in the morning. The automatic lane reads the tag, lifts the barrier and takes about four seconds a car. One morning the reader stops working and somebody switches that lane off. Notice what the switch did and did not do. The switch stopped the reader. It did not stop the cars. The cars keep arriving at the rate they were always arriving at, and now a person with a receipt book has to serve them at whatever rate a person with a receipt book can manage. Whether the plaza survives the morning is not a question about the switch at all. The question is how many people are standing there, how many minutes they have, and how many cars a minute keep coming.
A switch is a metaphor for something that is actually a transfer of work, and work that has been transferred has to be done by somebody else or not done at all. Every stopping arrangement, whatever the policy calls it, is therefore a claim about capacity. Either that capacity exists in posts and minutes, or the arrangement is a sentence in a document. The finance version carries a particular sting, and one invented bank below shows it: the capacity that makes a stop survivable is very often the first thing the deployment took away.
What does stopping an automated system actually mean, and at what levels?
People say stop as though it were one thing. Stopping is three things, and the three are not degrees of severity. The three levels differ in what they cover, and coverage is a different axis from severity. At Sumeru Bank Limited, invented, the retail loan intake chain runs nine numbered components and decides about 8,600 files a month, and its arrangements for stopping sit at three numbered levels.
Level 1 is a manual overrideA person changing the outcome on one file while the arrangement carries on running normally for everybody else.. A person changes the outcome on one file. The chain keeps running, unaltered, for every other file in the building. Level 2 is a fallbackTaking one population or one component out of the arrangement and having people carry that work instead.. One population, or one component, is taken out and people carry that work instead. Nobody's individual outcome is being argued about; a whole stream is being routed somewhere else. Level 3 is the kill switchA full stop of the whole arrangement, and the level almost nobody ever uses., a full stop of the chain, and at this bank it was never used at all.
The three are not three settings of one dial: each covers a different set of files, and a firm that holds one of them does not thereby hold the other two. This matters because policies get written as though holding an override meant being able to stop. Level 1 at this bank ran at 34 files a month against the 5,590 the chain decided with nobody touching them, being 0.6 per cent. Level 2 was used exactly once, in month 9, for 4 working days, being 1,720 files at 430 a working day and exactly 20.0 per cent of a month's volume. Level 3 has a line in the procedure and no history behind it.
What are the three levels of stopping?
Human Override vs Manual Fallback: which one covers a fault nobody noticed?
Levels 1 and 2 get talked about as though the second were simply a bigger version of the first. Override and fallback are not on the same line. An override is retail and reactive: it changes one outcome, for one person, usually because that person asked. A fallback is wholesale and planned: it moves an entire population out of the arrangement whether or not a single person in that population has said a word.
The everyday version is a housing society water pump. If one flat has no water and a plumber goes up and clears that flat's line, that is an override. If the pump itself is suspect and the society switches the whole building on to tanker supply for four days, that is a fallback. The second is not a bigger plumber. The tanker is a different decision, taken by a different person, on different evidence, and it reaches the flats that never rang anybody.
The reason this distinction earns its place is that a fault whose victims feel almost nothing produces almost no complaints, so a firm holding only the override has no route to act on it. At this bank the six week episode moved 176 files out of the accepted population and into the referral band. Of those 176, 152 were accepted by a person anyway. Twenty four were not, being 13.6 per cent of the 176. So the overwhelming majority of the people affected experienced a two day wait instead of a four minute answer, and then got their loan. Almost nobody had anything to complain about, so almost nothing would ever have reached an override.
What does a stop actually require, action by action?
A firm asked whether it can stop its arrangement will show a switch. A firm actually stopping one does six separate things, only two of which the switch does. At this bank the six are numbered, and the numbering is the bank's own rather than anybody's standard.
Action 1, stop new files entering. Action 2, hold the files already inside the arrangement. Action 3, route those held files to people. Action 4, tell those people what the chain would have done with each file and why that answer is now suspect. Action 5, tell the customers whose files are being held. Action 6, decide what happens to the files that were already decided while the fault was running.
The switch is actions 1 and 2. The other four are procedures, capacity and records rather than software, and none of them can be bought in the week they are needed. A firm that has a switch still does not have a stop. Look at what each of the four actually asks for. Action 3 asks for a desk with room in it. Action 4 asks for a record of what the arrangement produced and why, kept file by file, before anybody knew it would be needed. Action 5 asks for wording somebody has already drafted and cleared. Action 6 asks for a decision nobody has ever made before, under time pressure, about people who have already gone home.
A firm has a switch that stops new files entering and holds the ones inside. How much of a stop does it have?
Which of those six actions could this bank do inside an hour?
In month 9, week 3, the monitoring pack flagged that an income field had been arriving in a changed format from one channel since month 8, week 2. Ismail Sheikh, who runs the exception desk, and Revathi Balan, the named accountable person for the scoring model, took the decision to fall that channel back. Here is what then happened to each of the six actions, and the pattern in it is worth more than the story.
Actions 1, 2 and 3 were done within the hour. New files from that channel stopped entering the chain, the files already inside were held, and the held files were routed to the exception desk. The switch worked exactly as designed, on the day, at speed. Nobody at this bank has ever had cause to complain about the switch.
Action 5 took 2 working days. There was no cleared wording for telling a customer that a decision was being held while the bank checked something, so wording had to be written, read by the people who read such things, and sent. Two working days is not a scandal, but it is two days in which a person who expected an answer in about four minutes heard nothing.
Actions 4 and 6 could not be done at all, and the two of them are exactly the two that need a record the bank had never kept. Action 4 asks what the chain would have produced on each held file. The trail did not carry it. Two fields that would have held it, numbered 4 and 5 in the bank's own trail specification, did not exist until that month. So the desk received files with no statement of what the arrangement had been about to do. Every one had to be worked from the beginning as though the chain had never seen it. Action 6, the files already decided, had no procedure at all. Not a bad procedure. None.
What happens to the files that were already decided?
One question separates a firm that has thought about stopping from one that has bought a switch. The fault started in month 8, week 2. Somebody found out in month 9, week 3. Everything discussed so far concerns files that were inside the arrangement at the moment of the stop. So what happened to the files that went through while nobody knew?
The stretch between a fault starting and a fault being confirmed has a name worth learning: the already-decided windowThe stretch of time between a fault starting and being confirmed, and the files that were decided inside it.. At this bank it ran 30 working days and about 12,900 files were decided inside it. Of those, 176 were moved from the accepted population into the referral band by the fault, being 1.4 per cent of the window. The 176 did not have their files held. Each had already been told something. Some had been sent away to wait two days; some had gone elsewhere; all of them had gone home.
Consider a chemist whose weighing scale reads five per cent light and who discovers it on a Thursday. Stopping the scale is trivial. Deciding what is owed to the people who were served on Monday, Tuesday and Wednesday is the whole of the problem, and it does not get easier by being left. Action 6 is the only action that costs money after the fault is fixed, and the only one whose subject has already walked out of the building. Both of those make it the expensive action. This bank had to find the 176 by re-scoring all 12,900 files against the corrected reading step, because nothing in the trail said what the previous version would have produced. Every one of the 176 was then re-decided by a person.
The fault ran for six weeks before anybody knew. Which of the six actions covers that period?
What does a Fallback Procedure need in order to exist at all?
Now to the level that actually gets used. A Fallback Procedure is written as though it were a routing instruction: if the chain is suspect, send the work to the exception desk. Read it again and count what that sentence is assuming. The sentence assumes there is a desk. It assumes the desk has room. It assumes the people on it still know how to do the work by hand. And it assumes somebody has worked out how many files a day they can pass, and compared that number to how many files a day are arriving.
None of those is a software question and none of them is free. A fallback is a capacity claimA statement that somebody can absorb the work, either true in posts and minutes or not true at all.: a statement that somebody can absorb the work. A capacity claim is either true in posts and minutes or it is not true at all, and there is no third state in which it is true in principle. A fallback that cannot carry the volume is not a fallback; it is a slower way of failing, and the slowness is the only thing it adds.
Four things have to exist before the sentence in the policy means anything. First, somewhere for the work to go, named. Second, people there, counted. Third, minutes in their day that are not already spoken for. Fourth, a rate: how long one file takes a person, from picking it up to putting it down. Miss any one of the four and what has been written is an intention. This bank's own four numbers go into that arithmetic below.
What does the capacity arithmetic say on this bank's own desk?
The exception desk at Sumeru Bank Limited is 7 people. The bank assumes a working day of 420 minutes a person, the same 7 hour day it uses everywhere. So the desk holds 7 times 420, being 2,940 minutes a day, and that is the entire resource on the fallback side of the equation.
The rate is the bank's own locked figure for what a file used to take before the chain existed: 11 minutes of hands-on timeMinutes a person actually spends on a file, as distinct from how long the file sits waiting for somebody to pick it up. a file. Dividing gives the desk's throughput: 2,940 over 11 is 267.3 files a working day. On the arriving side, the channel that fell back handed over 430 files a working day, being the chain's locked daily volume. Two rates, and everything else in this guide is what happens when one is subtracted from the other.
| The fallback capacity build, month 9 | Figure |
|---|---|
| Posts on the exception desk | 7 |
| Assumed minutes a post a working day | 420 |
| Desk minutes available a working day | 2,940 |
| Hands-on minutes a file, pre-chain rate | 11 |
| Files the desk can pass a working day | 267.3 |
| Files handed over a working day | 430 |
| Files accumulating a working day | 162.7 |
Now run it over the 4 working days the fallback actually lasted. The channel handed the desk 430 times 4, being 1,720 files. The desk could pass 267.3 times 4, being 1,069. So 1,720 less 1,069 is 651 files that could not pass the desk inside the window. The fallback carried 62.2 per cent of what it was handed and left 37.8 per cent behind, and 62.2 plus 37.8 is 100.0, so nothing is hiding in a rounding.
One caution before the picture, and it makes the reading worse rather than better. The 1,069 is a ceiling on throughput and not a count of decisions. The 11 minutes covers checking, keying, filing and recording; the assessment and the credit decision itself sat with credit officers elsewhere in the building, who are not in this arithmetic at all. So 1,069 is the number of files the desk could move, and the number it could actually see decided is smaller by however much of the credit officers' time was free that week. A capacity claim that counts only the desk is already the optimistic version.
Is the 1,069 files a count of decisions?
How long can a fallback run before the backlog decides the matter?
Look again at the two rates. Neither of them is a risk. Both are constants. 430 files arrive a working day. 267.3 can pass. The difference, 162.7 files, is not something that might happen; it is what happens, every day, from the first day, for as long as the arrangement runs. A backlogWork handed to a desk faster than the desk can pass it, accumulating every day the arrangement runs. under these conditions is not a hazard. A backlog here is a line with a slope.
A constant difference changes the question a firm ought to be asking. The policy answers whether the firm can fall back, and answers yes. The policy does not answer for how long. Running the same two rates out gives a schedule rather than a warning. After 2.6 working days the accumulated backlog is already larger than a full day of arrivals. After 4 working days it is 651 files. After 20 working days, being a whole month, the desk has passed 5,346 and 3,254 could not pass. Two constants make a constant ratio, so the share never moves: the desk carries 62.2 per cent at every duration, and only the count of files left behind grows.
So here is the honest reading of this bank's month 9, stated as arithmetic rather than as alarm. The fallback worked. Four working days was inside what the desk could absorb, the backlog it left was recoverable, and nobody outside the building noticed. The same fallback, run for twenty working days, would not have worked at all, and nothing about the desk, the people or the policy would have been any different. A four day fallback and a twenty day fallback are not the same control used for longer; the first is a control and the second is an unfunded promise, and the only thing separating them is a duration nobody had written down.
One more calculation, and it is arithmetic on locked figures rather than a measured reading. The 651 files carry 651 times 11, being 7,161 minutes of hands-on work. In ordinary running the desk holds a margin of about 4.2 cases a day at 19 minutes a case, being roughly 80 spare minutes a working day. Dividing shows that the standing margin would take about 90 working days to absorb what four days of fallback left behind. Which says something blunt about every recovery plan that promises the desk will catch up: a backlog of this shape is cleared by putting more minutes on the desk, never by waiting for the margin to eat it.
A fallback is written into a policy with no stated duration. What is missing?
Where did the posts that make a fallback work actually go?
The arithmetic also runs the other way. Instead of asking what 7 people can carry, the question is what carrying the whole arriving volume would take. 430 files a working day at 11 minutes a file is 4,730 minutes. Divided by the 420 minute day, that is 11.26 posts, so 12. Twelve posts is the size of desk at which the arriving line and the passing line are the same line, and nothing accumulates at any duration whatever.
Now the part that is uncomfortable and has to be said carefully. The exception desk at this bank was 12 people before go-live and is 7 people now. The deployment removed five posts. So the desk size that makes the fallback work is exactly the desk size the business case deleted, to the post, and neither of those two decisions was wrong when it was taken. Five posts at the bank's assumed fully loaded cost of Rs 9,00,000/- a year is Rs 45,00,000/- a year, against a running cost of Rs 65,00,000/- a year for the chain itself.
The arithmetic argues for costing the fallback separately, and it does not argue against the deployment. The chain gives about 65 per cent of applicants an answer in roughly four minutes instead of two working days, and that is worth having. The mistake was never removing the posts. The mistake was writing a fallback into a policy afterwards without going back to the arithmetic and noticing that the sentence now described a desk that no longer existed. A promise made in one document and priced in another is how firms end up holding neither.
A business case removes five posts and a policy promises a manual fallback. What has been promised?
Before the control below is moved: the fallback runs 4 working days and hands the desk 1,720 files. How many can the desk pass?
How long can this desk hold a fallback open?
One variable: the number of working days the fallback runs. Everything else is held at this invented bank's own figures. Watch the two lines separate, and notice that they separate from the very first day.
Working days the fallback runs: 4
One invented bank, one deployment. 430 files a working day, 7 posts at an assumed 420 minutes each, and 11 minutes of hands-on time a file as it stood before the chain existed. A desk out of practice is slower than a desk that never stopped, so holding the 11 minutes constant across every duration flatters the fallback. Second assumption, stated rather than buried: the credit assessment sat with credit officers elsewhere, so the 267.3 a day is desk throughput and not decisions.
Is Manual Override a control, or a complaints route wearing its name?
Back to level 1, the level every firm is sure it has. At this bank the override ran at 34 files a month. Set against the 5,590 files a month the chain decided with nobody touching them, that is 0.6 per cent. A small number is not automatically a bad number; a well designed arrangement should be right most of the time, and an override that fired constantly would say something worse. The count is not the interesting number. Who raised them is.
Every one of the 34 arrived because a customer complained. Not one was found by anybody watching the chain. Once that is known, the name on the control has to change. The bank holds a complaints routeAn arrangement that only ever reaches the people who complain, whatever the policy calls it., a perfectly respectable thing to have and a completely different thing from a control.
The difference is not about quality, it is about coverage, and coverage is the axis this whole guide runs on. A complaints route reaches exactly the people who noticed something and had the time, the confidence and the phone number to say so. The route never reaches anybody else. Set that against the six week episode: 176 files moved, 152 of them accepted by a person anyway. The overwhelming majority of the people the fault touched waited two days and were then approved. A control that depends on the affected person noticing cannot catch a fault whose entire experience, for almost everybody it touches, is a two day wait. Cathy O'Neil, in Weapons of Math Destruction, 2016, makes the related point that the people an arrangement gets wrong are frequently the people least placed to contest it, and a route that waits to be contested inherits that problem whole.
The one question that separated the two at this bank
Ask it exactly this way, of last year's records: of the overrides used, how many were raised by somebody other than the customer? Not how many overrides there were, not what share of decisions they touched, not whether the authority is documented. Just who raised them.
At Sumeru Bank Limited the answer was none. Thirty four a month, twelve months, every single one arriving through a complaint. The control existed, the authority was real, and the people using it were doing their jobs properly. The fault that actually happened produced almost nobody with a reason to ring anybody, so the control could never have found it. A firm that discovers this about itself has not found a scandal. Such a firm has found that it holds one of the three levels and has been reporting it as though it held the arrangement.
Every override last year was raised by a customer. What kind of control is it?
How does a lender, an analyst or a household read a promise to fall back?
Take this off the invented bank for a moment. The shape turns up everywhere, and being able to see it fast is worth the detour.
A household version first. A spare key left with the neighbour on the second floor is a fallback, and a good one: the capacity claim behind it is tiny, being one key, one door, one person, thirty seconds. Now suppose the arrangement is that the neighbour will look after the household's parents for a week while it is travelling. Same word, same warm feeling, completely different claim, and the honest question is not whether the neighbour is willing. The question is how many hours a day the job actually takes and how many hours a day that neighbour has. Every fallback anybody has ever promised is a rate against a rate, and the ones that fail are the ones where nobody did the division.
A credit analyst reading a lender's operational resilience section is looking for exactly two numbers, and will usually find neither. First, at what rate can the manual route carry the work, in files a day or claims a day or payments an hour. Second, for how many days can it carry it before the queue is unmanageable. A disclosure that says a manual process is available and stops there has said nothing checkable. A disclosure that says the manual route carries about 62 per cent of the daily flow and is designed to hold for up to five working days has said something an analyst can test, argue with and hold the firm to.
A person sitting on a board or a risk committee has a cheaper question available, and it is the one from the failure block above: of last year's overrides, how many were raised by somebody other than the customer. One query against the records answers it, and the answer separates a control from a complaints route without any argument about definitions. The same two questions apply to a bought-in service rather than a built one, and there they get harder. The capacity that has to absorb the fallback usually sits in the buyer's own building. The failed arrangement sits in somebody else's.
How to Design an AI Fallback and Override Procedure: what has to be written first?
Everything above points the same way. Every part of a stop that matters has to exist before the day it is needed, so stopping an automated arrangement is a design decision taken in advance and never a reaction. The switch can be built in an afternoon. The desk, the trail, the wording and the policy for files already decided cannot.
The document that has to be written is short. Not a plan, not a framework, a numbered list of six answers, each of which is a fact about the firm rather than an intention.
The six answers have something in common. Every one of them is answerable on a quiet Tuesday and unanswerable on the day. Item 4 is the one the arithmetic above turns on, and it is the item most often left as an adjective: adequate capacity, appropriate resourcing, sufficient cover. Adequate for how many files a day, and for how many days. If the procedure cannot answer that in two numbers, the fallback in it has been mentioned rather than designed.
Where does the accountability for a stop sit?
Where a lender suspends an automated decisioning arrangement, falls a channel back to people, or tells customers that their applications are being held, the expectations that apply to a regulated lender in India are stated by the Reserve Bank of India at rbi.org.in, covering matters such as outsourcing, digital lending, data and consent and the oversight of arrangements that decide customer outcomes. The standing discipline of model risk work, from which much of this vocabulary descends, originates in international supervisory material published by the Bank for International Settlements at bis.org, an origin rather than a statement of the position in India. The three levels, the six actions and every figure attached to them are one invented bank's own arrangement rather than anybody's standard, and the position that binds a given lender is stated at the source.
Covered elsewhere. How a fault is detected in the first place, and which signals speak and how fast, is set out under model drift and monitoring; stopping starts after a fault has been confirmed. The whole response once a fault is confirmed, run against a clock across eight stages, is set out under how to create an AI incident-response playbook, and it uses these same six actions. How the exception desk is staffed and what one exception costs to work is established elsewhere and used here.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Published expectations on a regulated lender covering digital lending, outsourcing, data and consent, and the oversight of arrangements that decide customer outcomes, including the suspension of such an arrangement | rbi.org.in |
| Bank for International Settlements | International supervisory material from which the standing discipline of model risk work originates | bis.org |
| Cathy O'Neil | Weapons of Math Destruction, 2016, on the errors of an arrangement falling on the people least placed to contest them | named in the text |
Sumeru Bank Limited, Revathi Balan and Ismail Sheikh are invented.
Educational material. Not advice on any investment, tax, budget or market position.
