Workflow Automation: Moving Work Without Moving People
Workflow automation moves the carrying out of a step from a person to a system. Automation does not move the judgement in the step, the accountability for the outcome or the work of handling what the system cannot finish. Most automation in finance is a set of written rules rather than anything that learns, and describing it as a rule engine is both more accurate and easier to govern.
Every automated step rests on a single claim: that the correct outcome could have been stated in advance. Where that claim holds, the machine is carrying out a decision somebody has already made and written down, so the automation is unremarkable and dependable. Where the claim does not hold, automating the step does not remove the judgement from the process. Automating the step only removes the record of who exercised the judgement.
What does automation actually move, and what stays exactly where it was?
Start somewhere ordinary. A household buys a washing machine. The machine takes over the scrubbing, and that is a real change: two hours on a Sunday come back. But look at what did not move. The machine will happily ruin a silk saree at sixty degrees, so somebody still has to decide which clothes can take a hot wash. Somebody is still the person who has to answer for the ruined saree. And the hand-wash pile is still sitting in the bucket, and it is now the whole of what is left to do by hand.
The washing machine shows the entire shape of workflow automationMoving the carrying out of a step from a person to a system, while the deciding and the answering for it stay where they were. in a finance process, and the shape does not get more complicated at scale. One thing moves: the carrying out of the step. Three things stay. A machine executing a rule is not making a choice, it is applying one somebody else made, so the judgement inside the step stays. No supervisor, board or customer has ever accepted a system as the answer to who decided this, so the accountability for the outcome stays. And the handling of everything the system could not finish stays, and that last one is the expensive surprise. The leftover work does not merely stay the same size. Its shape changes.
Automation moves one thing and leaves three behind, and the three it leaves are precisely the ones a business case forgets to put a number against. A case that costs the minutes of the step and stops there is not slightly optimistic. A case like that is optimistic in the same direction every single time, and that consistency is what makes the error worth learning as a shape rather than as an anecdote.
A bank moves the carrying out of a step to a system. Which three things stay exactly where they were?
What is most finance automation actually made of?
Rule-Based Automation
Rule-based automationAutomation whose behaviour is a written procedure that a person composed, rather than behaviour derived from past data. is automation whose behaviour is a document. Somebody sat down, wrote a condition and an action, and the system now performs that condition and that action several thousand times a month without getting bored. There is no mystery in it and no cleverness in it, and both of those absences are advantages rather than apologies.
Take the retail intake chain at Sumeru Bank Limited, an invented lender. Four of its parts are rule sets somebody wrote: the income corroboration rule, the fraud rules on the servicing book, the workflow router, and the identity match. Between them they hold 126 lines. The income rule is 34 of those lines, the fraud rules 61, the router 22 and the identity match 9. Neelima Rao, in the risk function, read all 34 lines of the income rule in twenty-five minutes and could then state what it would do with any file handed to her. The scoring model derives its behaviour from data rather than from a document. Reviewing that model took eleven working days and still ended in a description of tendencies rather than a rule.
Most of what a bank calls automation is a set of written lines that one competent person can read in an afternoon. The claim is not a small one, and it decides the whole governance question. A rule can be printed, disagreed with, dated and changed by a named individual. A rule can also be found wrong in a way that has an address: at the month twelve validation, 9 of the 126 lines, being 7.1 per cent, either contradicted another line or could never be reached by any file at all. Nobody needed a statistical method to find those nine. Finding those nine needed a reader with a printout.
Most automation deployed in a finance process is which kind of thing?
Why does calling it something cleverer make it harder to govern?
Because the description sets the question people ask next. Tell a reviewer that a step is carried out by a rule and the next question is the obvious one: show me the rule. Tell the same reviewer that the step is handled by an intelligent system and the next question quietly becomes how well it performs. How well it performs is a different question with a different answer and no document at the end of it. The grander word does not make the step better. The grander word makes the step unreadable, and it does that by suggesting there was never anything to read.
The inflation runs in the other direction too, and that is the half people miss. At Sumeru the workflow router is 22 lines and it decided where every one of the month's 8,600 files went. Nobody had ever described those 22 lines as intelligent, so nobody put them on any list of things that needed opening. The part of the chain with the widest reach went unexamined precisely because it had been described accurately and modestly. A label that flatters one part of a system does not only overstate that part. A flattering label also starves everything it declined to cover.
How can a step be judged ready to be moved?
How to Assess Whether a Finance Process Is Ready for Automation
The tea seller outside a large office building illustrates the point. Between eight and eleven he makes the same drink about four hundred times. The recipe has not changed in six years. He holds exactly three sizes of glass, so his inputs arrive in exactly three forms. The correct answer for any order can be stated before the order arrives. And when he gets one wrong, the cost is one glass of tea and he pours another. A machine could do that work, and the tea machine in an office lobby is there for that reason.
Now watch the same man do something else. A regular is short of cash and asks to pay on Friday. He looks at the person, thinks about the last three times, and says yes or no. The credit decision happens perhaps four times a day, the rule is in his head and moves with his mood, the input is a face rather than a field, the correct outcome cannot be stated in advance, and getting it wrong costs him money and a customer. Same man, same stall, same morning. Utterly different step.
The readiness testA scored check on whether a step can be automated safely, run before anything is built rather than after it has gone wrong. Sumeru used is that comparison written down. Five tests, each scored zero, one or two, for a maximum of ten.
| Test | What it asks | A score of two looks like |
|---|---|---|
| 1 Volume | Does this step happen often enough to be worth building for? | Thousands of times a month, every month |
| 2 Stability | Does the rule change less often than the volume arrives? | The rule has not moved in years |
| 3 Data | Does the input arrive in a form a machine can read? | A field, already structured, already there |
| 4 Decidability | Can the correct outcome be stated in advance? | The answer can be written before the case arrives |
| 5 Consequence | What happens when it is wrong, and who bears it? | The error is cheap, visible, and lands on the bank |
The first four tests ask whether a step can be moved. Only the fifth asks what it costs when the answer is wrong, and who pays for that. The fifth test is the one that gets skipped, and it gets skipped because by the time anybody is scoring a step, somebody senior has usually already decided to automate it and is looking for a number that agrees. Where a wrong answer lands on a customer who has no way to see it coming and no procedure to appeal to, a step can pass volume, stability, data and decidability handsomely and still be a terrible thing to automate.
Sumeru set its own bar at seven. Seven or above, automate the step outright. Five or six, assist the step rather than replace it. Below five, leave it with people. The bar of seven is the bank's own choice, arrived at in a room, and it is not anybody's standard. A different bank could reasonably set it at six or at eight, and the useful thing about writing it down is not that the number is right but that it exists before the scoring starts rather than after.
A step happens 8,600 times a month, its rule has not changed in two years, and its input arrives as a clean structured field. Ready to automate?
What happened when the test was run over one real process?
Before the intake chain existed, a retail personal loan at Sumeru went through eleven steps, in this order: the application received, the identity documents checked, the documents sorted and filed, the fields keyed into the system, the income corroborated against the statement, the credit record pulled, the assessment written, the decision taken, the decision recorded, the letter drafted and sent, and the disbursal instruction raised.
Two of those eleven decide what every later number means, so they are worth pausing on before the scores. The intake and processing desk handled steps 1, 2, 3, 4, 9, 10 and 11 and nothing else. At 1.0, 2.0, 1.5, 3.0, 1.0, 1.5 and 1.0 minutes of hands-on timeMinutes a person actually spends working on a file, as distinct from how long the file sits waiting for somebody to pick it up. those seven steps sum to exactly 11.0 minutes a file. The 11.0 minutes is the desk's handling time before anything was built, and every later comparison of desk cost rests on it. Steps 5, 6, 7 and 8 sat with credit officers somewhere else in the building and are not inside the desk's twelve posts, or the seven that came after.
So the 11.0 minutes is hands-on desk time on checking, keying, filing and recording. The figure is not the whole assessment, and it is emphatically not the two working days a file took to come back, most of which was a file waiting in a queue rather than anybody working on it. Both numbers describe the same file, so sliding between the two is easy to do by accident, and it is the commonest way an automation saving gets counted twice.
Now look at the scores as a shape rather than as a list. The shape is the finding. The scores are not spread evenly and they are not spread randomly. Steps 7 and 8, writing the assessment and taking the decision, scored three and four. Every single step around them scored seven or more. The dip sits exactly where the judgement sits, in the middle of the process rather than at either end, and the middle is where a lending process puts its thinking.
The trough is not a coincidence about one bank; it is what the fourth test measures. Decidability asks whether the correct outcome can be stated in advance, and the steps where it cannot are, by definition, the steps a person was hired for. A readiness scoring exercise across almost any finance process draws this same silhouette: high at the receiving end, high at the recording end, and a trough in the middle where somebody has to weigh something.
The desk's handling time before the chain was 11.0 minutes a file. What does that 11.0 actually cover?
Which step failed the test and was automated anyway?
Step 8, the decision itself, scored four out of ten. Step 8 was automated regardless, and the reason was neither stupidity nor stealth. The reason was arithmetic. A chain that carries a file all the way to the edge of a decision and then stops cannot return an answer in about four minutes, and the four minutes was the entire point of building it. The business case needed step 8, so step 8 went in.
Overriding a readiness score is sometimes the right call. The test is an input to a decision, not the decision, and a bank that never overrode its own scoring would never build anything at all. The difference between a defensible override and a quiet one is small and entirely procedural. Somebody writes down that the decision was an override, what the score was, the test it failed out of the five, why the exposure was accepted anyway, and whose name sits against that acceptance. None of that is difficult. The writing down is simply easy to skip on a week when the go-live date is close.
The consequence at Sumeru is the general lesson in a specific shape. The judgement did not go anywhere; only the record of it left. In one steady month, 1,290 files reach an outcome the applicant would describe as a refusal. Of those, 602 were routed by the written income rule, and behind each of them sits a printed procedure a manager can read out. The other 688 were declined by the scoring model, and behind those sits no procedure at all, only an attributed reason. 602 plus 688 is 1,290. A complaints manager who answers every refusal by quoting the written procedure is giving the wrong kind of answer to 53.3 per cent of them, and that is a direct consequence of one score of four being waved through.
Step 8, the decision, scored four out of ten on the readiness test and was automated anyway. What follows from that?
What does it mean to assist a step rather than replace it?
Step 10, drafting and sending the decision letter, scored six. Six is the interesting number on the whole sheet, sitting between the bank's two bars: too low to hand over outright, too high to leave alone. Sumeru's answer was neither yes nor no. A draft of the letter is produced by the chain, and a person reads it and signs it before it goes anywhere. The keystrokes moved. The signature did not.
Assisting a step moves the typing and leaves the signature, and that is a different decision from automating it, not a softer version of the same one. The difference shows up the moment something is wrong with a letter. On an assisted step there is a named person at the end of it who can be asked what they read before they signed. On an automated step the equivalent question exists too, but it points somewhere else and much further back: not who signed this letter, but who approved the rule that wrote it, and when. Both are answerable. The two questions are not the same, and a process design that leaves it unclear which one applies will discover the ambiguity on a bad day rather than a good one. Where assistance pays for itself and where it quietly costs more than it saves is worked through separately.
Which four figures flatter an automation programme?
Four numbers turn up in nearly every report on a deployment like this one, and all four of them are true. Their truth is exactly what makes them slippery: nobody has to lie to mislead with them, they only have to stop the sentence halfway through.
The first is the straight throughA file that reaches its outcome with no person touching it at any stage. rate quoted on its own. At Sumeru, 5,590 of the month's 8,600 files were decided with nobody touching them, a rate of 65.0 per cent. Nothing sits beside the rate to read it against, so read alone it sounds like a strong result. The second is the volume processed. The chain handled 8,600 files in the month. The number is large, but volume moves with demand and most of it would have arrived whether or not anything had been built. The third is the cost a file, averaged across the whole month. The average quietly blends a group that now costs the bank almost nothing with a group that costs more than it used to. The fourth is the posts removed, presented as a saving with no cost line anywhere near it.
Each of the four is a true figure with its second half missing, so the fix is never to find a cleverer measure but to finish the sentence. An exceptionA file the system sends to a person instead of completing, because something in it could not be resolved by rule. desk, a board and a regulator all read a half sentence the same way, generously.
How can the result be reported without flattering it?
How to Measure Automation Outcomes Responsibly
Restoring to each flattering figure the half that was removed gives four honest measures, and they are the same four numbers already to hand.
One, the straight-through rate with the business case beside it: 65.0 per cent against a business case of 85. Two, the time to an answer for both groups rather than the fast one: about four minutes for the 65.0 per cent who go straight through, and two working days, entirely unchanged from before the chain existed, for the other 35.0 per cent. Three, the handling time on what is left, up from 11 minutes a file to 19. Four, the total cost including the cost of running it: Rs 65,00,000/- a year to operate the chain, against five posts saved at an assumed fully loaded Rs 9,00,000/- a year each, being Rs 45,00,000/-. The running cost exceeds the headcount saving by Rs 20,00,000/- a year, and that is before the Rs 2,40,00,000/- spent to build it is counted at all.
Written that way, the chain does not pay back on headcount. Saying so plainly is not an admission of failure and it is not a criticism of the build; it is the sentence that makes everything else in the report believable. The chain does pay back on something else, stated in the same breath: 65.0 per cent of applicants now get an answer in about four minutes instead of two working days, and the same desk can absorb a great deal more volume than it could before without a single extra post. Both gains are real and both are worth money. Neither is the gain the business case was written around.
A report states that the straight-through rate is 65.0 per cent and moves on. What is the missing half?
The error that gets made, and what it costs
The business case for the intake chain promised ten posts off the exception desk. Five came off. Nobody in this case was careless and nobody was dishonest. The gap of five has an entirely arithmetic explanation, and the arithmetic is worth walking slowly.
The case made two assumptions. The first was that 85 per cent of files would run straight through. The deployment produced 65.0 per cent, a real shortfall but not the expensive one. The second assumption was the costly one, and it was so natural that it was never written down as an assumption at all: that a file still needing a person would take the same 11 minutes it had always taken. On that arithmetic, 1,290 exceptions a month at 11 minutes is 14,190 minutes, and at an assumed working month of 8,400 minutes a person that is 1.69 posts, so two. Twelve before, two after, ten saved.
The chain produced 3,010 exceptions a month at 19 minutes each, being 57,190 minutes and 6.81 posts, so seven. Twelve before, seven after, five saved. The 19 minutes is not a sign that anything got worse. The rise is what happens when every file that used to be quick stops arriving at the desk.
The cost was not the five posts. Sumeru could carry five posts. The cost was that a number presented to a board as a saving turned out to have a hidden assumption inside it, and once that is discovered, every other figure in the same paper has to be re-opened by somebody who now has a reason to doubt it.
| The reading | Exceptions a month | Minutes each | Total minutes | Posts |
|---|---|---|---|---|
| The desk before the chain | 8,600 | 11 | 94,600 | 12 |
| The business case assumed | 1,290 | 11 | 14,190 | 2 |
| What the deployment produced | 3,010 | 19 | 57,190 | 7 |
Read the table downward and the whole gap is visible in two cells. The promised saving is 12 less 2, being ten. The actual saving is 12 less 7, being five. Every other number in the calculation, including the working month of 8,400 minutes and the rounding up of a part post to a whole one, is identical in both rows. One assumption, that the work left behind resembles the work that left, produced the entire difference between ten posts and five.
Before the control below is moved: the business case assumed 85 per cent straight through and 11 minutes an exception, and on those two assumptions it promised ten posts. How many were actually saved?
Move the straight-through rate, and watch the two desks separate
One control: the share of the month's 8,600 files that reach an outcome with nobody touching them, from 40 to 95 per cent. One consequence: the posts the exception desk needs, costed twice over, once at the old handling time of 11 minutes a file and once at the 19 minutes an exception actually takes. The default below is the real reading of 65.0 per cent, producing 3,010 exceptions, 4 posts on the old handling time and 7 posts on the real one, against the 12 posts the desk had before the chain existed. The business case sat at 85 per cent and 11 minutes together, giving 1,290 exceptions and 2 posts, and that corner is the only one where ten posts come off.
65.0 per cent of files straight through
At 65.0 per cent straight through, 3,010 files a month reach a person. Costed at the old handling time of 11 minutes each, that is 33,110 minutes and 4 posts. At the 19 minutes an exception actually takes, it is 57,190 minutes and 7 posts, against the 12 the desk had before. The 3 posts between the two bars are what the harder residue costs, and the business case had no line for them.
Why does the work that is left take longer per file?
Walk past a vegetable seller at eight in the morning and again at eight at night. The evening pile looks worse, and the obvious conclusion is that the vegetables got worse during the day. The vegetables did not. Every tomato in the evening pile was sitting in the morning pile too. Everybody who came in between simply reached for the good ones, and what is left is what nobody chose. Nothing about any single tomato changed. The composition of the pile changed.
The evening pile is the residueThe work left after the easy cases stop arriving, which is harder than the old average was even though no individual case became harder., and the residue is the single most reliable surprise in automation. The chain at Sumeru took the files it could finish, overwhelmingly the straightforward ones, and handed back the rest. The desk did not get slower and the staff did not get worse. The desk simply stopped receiving the two-minute files. Every file arriving there now is one the chain could not resolve, a completely different population from the one the old 11 minute average described.
The 19 minutes an exception now costs breaks down like this: 4 minutes to re-read the file, 6 to find what the chain could not, 5 to obtain the missing information, 3 to decide or refer, and 1 to record. The five parts add to 19. Look at what that composition says. Only 3 of the 19 minutes are the deciding. Ten of them, the 4 to re-read and the 6 to find what the chain could not, are work that barely existed when the desk saw every file and had already touched most of them once.
Total desk work fell by 39.5 per cent while the time each remaining file takes rose by 72.7 per cent, and an honest report holds both of those in the same paragraph. Desk minutes a month went from 94,600 to 57,190. Hands-on minutes a file went from 11 to 19. Presenting only the first is selling; presenting only the second is complaining. Presenting both describes what happened.
One design choice followed directly from this, and it is worth copying. At go-live every exception was given one reviewer. From month 5 a second reviewer is given only to the 391 files a month sitting in the scoring model's referral band, being 13.0 per cent of the 3,010. The reasoning is precise rather than budgetary: those 391 are the only exceptions where a person is actually making the decision. On the other 2,619 the person supplies information the chain was missing and then lets the chain finish, a different task that does not need a second pair of eyes.
What does an automated decision owe an auditor?
How to Create an Audit Trail for Automated Financial Work
Once a step runs without a person, the only evidence that anything happened is whatever the system chose to write down. The audit trailThe record of what acted, on what, producing what, and who had the standing to stop it and did not. is therefore a design decision rather than a technical by-product, and the design has to be settled before go-live. A field nobody recorded in month 4 cannot be recovered in month 9.
Sumeru's trail carries six fields for every action by every component. One, the component and version that acted. Two, what it read. Three, what it produced. Four, what the previous version would have produced, where that is known. Five, who could have intervened and did not. Six, the time.
Four of those six were recorded from go-live: which version acted, what it read, what it produced, and when. Fields four and five were added in month 9, and the reason is the whole argument for them. In month 8 an upstream income field quietly changed format on one channel, and roughly six weeks passed before monitoring flagged it. When it surfaced, one question was asked in every room: which files would have been decided differently? Nobody could answer it. Nothing in the record said what the earlier version of the reading step would have produced on those files, and nothing said which named person had the standing to stop the run and had not used it.
Four of the six fields record what happened, and only the two that were missing record what would have happened otherwise, the only question anybody actually asks after something goes wrong. The asymmetry is a general point rather than a Sumeru one. A trail designed to prove that a system ran is easy to build and answers nothing; a trail designed to answer a counterfactual has to be planned while the system is still being built, by somebody imagining the enquiry.
Field one earns its place for a reason worth naming. The core system offers no other way in, so two of the eleven steps, recording the decision and raising the disbursal instruction, are carried out by an automation that operates that system through its own screens. In the period it broke 14 times, every single break following a change to a screen, at roughly 4 hours to restore each time, being 56 hours in total. Without a record of which version of that automation acted on a given file, there is no way to tell an entry made cleanly from one made by a component reading the wrong box on a changed screen. How that kind of automation is built and why it breaks the way it does is covered separately.
Which two of the six trail fields were missing at go-live, and why did that matter?
How does an operations head, an auditor or a board member actually use this?
What each of them asks for, and when
An operations head sizing a desk after a deployment has exactly one question that matters, and it is not the straight-through rate. The question is what the remaining files now cost to handle. Costing the residue at the old average is the single mistake that produced the whole gap at Sumeru, and it is avoidable with one line in a plan: measure handling time on the files that actually remain, three months after go-live, and do not reuse the pre-deployment average for anything.
An internal auditor arriving at a chain like this can save weeks by asking two questions on day one. Which steps were automated with a readiness score below the bank's own bar, and who signed the override. And which of the six trail fields are actually populated. The answer decides whether a counterfactual enquiry is even possible before anybody starts one. At Sumeru the first question surfaces step 8 immediately and the second surfaces fields four and five.
A board member is usually shown a saving and a cost of build, and is rarely shown the cost of running. The one question that reliably repays asking in that room is what the chain costs to operate against what it saves, stated as one sentence with both halves in it. At Sumeru that sentence is Rs 65,00,000/- a year to run against Rs 45,00,000/- a year of posts saved, a shortfall of Rs 20,00,000/- a year before the Rs 2,40,00,000/- of build cost is counted at all, and the case is made instead on the 65.0 per cent of applicants who now get an answer in about four minutes. A case built on the four-minute answer is a perfectly good one. The four-minute case is simply a different one from the case that was originally approved, and a board is entitled to be told which of the two it is being asked to back.
Who sets expectations on a lender running a chain like this
A bank in India operating an automated intake and decision process sits under the Reserve Bank of India, whose expectations on outsourcing, digital lending, record-keeping and customer consent are published at rbi.org.in. Where the deployer is a market intermediary rather than a lender, the equivalent expectations come from the Securities and Exchange Board of India at sebi.gov.in. The five readiness tests, the bar of seven, the working month of 8,400 minutes and the six trail fields are all the invented bank's own design and are nobody's rule. The current position is stated at the issuing body's own site.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering outsourcing, digital lending, record-keeping and customer consent | rbi.org.in |
| Securities and Exchange Board of India | Equivalent expectations where the deployer of such a process is a market intermediary | sebi.gov.in |
| Ministry of Corporate Affairs | Material on the accountability of a board for what it approves | mca.gov.in |
| Bank for International Settlements | International supervisory material on the deployment of automated processes by banks | bis.org |
| Agrawal, Gans and Goldfarb | Prediction Machines, on a component that derives its behaviour from data as producing a prediction a person still has to act on | Harvard Business Review Press |
Sumeru Bank Limited and Neelima Rao are invented.
Educational material. Not advice on any investment, tax, budget or market position.
