Straight-Through Processing and Exception Handling
Straight-through processing means a file reaching its outcome with nobody touching it, and the rate is the share of files in a period that do. Exception handling is everything that happens to the rest. The two are not opposites in difficulty. As the rate rises, the files still stopping are the ones the automation found hardest, so the work left behind gets slower for each file even while it shrinks in total.
A rate of this kind is a statement about the easy cases. The rate says nothing at all about the hard ones, and it says nothing about whether any of the outcomes was right. Every design decision that follows rests on one fact: what is left over after automation is not a smaller, random sample of what used to arrive. The leftover is the part the machinery could not finish, and a part that could not be finished is a different population with different costs. Costing it as though it were the old work in miniature is the standard error, and it is made in almost every business case written for this kind of programme.
What is straight-through processing, and what is the rate counting?
Straight-through processingA file reaching its outcome with no person touching it at any point along the way. is a plain idea with a precise boundary. A file arrives, passes through every step of the machinery, and comes out the other end with an outcome recorded, and at no point did a person open it, key anything into it, telephone anybody about it or sign it. If a person touched it once, for one minute, it is not straight through. There is no partial credit and no half mark. The absence of a half mark is what makes the count usable at all.
The straight-through rateThe share of files in a period that reached their outcome untouched, expressed as a percentage. is then just division: the files that reached an outcome untouched in a period, divided by all the files that entered the machinery in that period, with the base of that division stated. A rate quoted without its denominator can be moved several points simply by choosing a different starting line, so the base is the part people get wrong. Counting from the moment an application is submitted gives one number. Counting from the moment it reaches the decision engine gives another. Everything that fell away before that point has quietly left the calculation.
Here is the household version. The share of a household's monthly bills that paid themselves without anyone doing anything depends on which bills are being counted. All of them, including the one that arrives by post and always needs a phone call? Or only the ones set up on a standing instruction in the first place? The second question flatters the household enormously and describes nothing.
Notice the three things the rate leaves out. The rate does not measure whether the untouched outcomes were correct. The rate does not measure how long anything took. The rate does not measure what the machinery cost. A component that accepted every application without looking would produce a rate of one hundred per cent, and a bank reading that number as a quality score would be reading it exactly backwards. The rate counts work avoided, not work done well, and those are two separate questions that need two separate measures.
What does the straight-through rate count?
What did the rate come out at, and what had the business case assumed?
Sumeru Bank Limited, an invented lender, ran its retail loan intake machinery through one steady month and received 10,000 applications, of which 8,600 completed digital onboarding and reached the decision engine. Of those 8,600, exactly 5,590 came out with an outcome and no fingerprints on them, a straight-through rate of 65.0 per cent. The remaining 3,010 stopped somewhere and waited for a person.
The paper that had authorised the spend assumed 85 per cent. Twenty points of straight-through rate is not a rounding difference in a business case; it is the difference between a desk of two people and a desk of seven. The rate is the headline everybody quotes, so it is worth being exact about what each version implies before going anywhere near the reasons.
| What was assumed, and what happened | Business case | Month six |
|---|---|---|
| Files reaching the decision engine | 8,600 | 8,600 |
| Straight-through rate | 85 per cent | 65.0 per cent |
| Files stopping for a person | 1,290 | 3,010 |
| Hands-on minutes on one such file | 11 | 19 |
| Desk minutes a month | 14,190 | 57,190 |
| Posts needed, at an assumed 8,400 working minutes a person a month | 2 | 7 |
The desk held 12 posts before any of this was built. The paper promised it would fall to 2, a saving of 10. The desk fell to 7, a saving of 5. Half the promised benefit arrived. Both of the numbers that produced the shortfall sit in the table above, and only one of them is the rate.
Why do the exceptions get harder as the rate rises?
An exceptionA file the machinery sends to a person instead of completing, because some step could not be finished as written. is a file the machinery could not finish. The whole mechanism sits inside that one sentence, so read it slowly. The machinery is a set of written rules and fitted components arranged in a fixed order. The machinery does not survey a file, form a view of its own difficulty and elect to hand over the tricky ones. Each step either completes on the inputs it was given or it does not, and the router sends the ones that did not into a queue. The router is itself a written rule of twenty-two lines. Nothing in that arrangement chooses. The arrangement only fails to finish, and failing to finish is not spread at random.
So think about which files stop. A file with four clean, well-lit document images, an income that matches the statement to the rupee and an identity record agreeing across two sources will sail through, and that same file used to be the one a clerk finished in four minutes. A file with a smudged payslip photographed at night, a declared income fifteen per cent above what the statement shows and a name spelt two ways is the one that stops, and that same file always took a clerk twenty minutes. Automation removes work from the quick end of the spread first, so the average difficulty of what remains rises without any individual file having changed at all.
The leftover behaves as a group, so it deserves a name. Call it the residueThe files still stopping once the easy ones stop arriving. They are harder than the old average because the easy ones have gone.. The residue is what the desk actually handles now, and it is harder than the old average by construction rather than by misfortune.
The household version is a shelf of books being sorted into keep and give away. The obvious keeps go first and the obvious discards go next, and both are quick. The pile left in hand after twenty minutes is the one for which there is no rule, and every book in it takes a minute of real thought. Nobody made the books harder. The easy ones have been taken out.
Hands-on time rose from 11 minutes to 19 minutes a file after the machinery went live. Did the work get harder?
What does the gap between two posts and seven actually break into?
Two numbers moved against the business case. Each has a different owner and a different fix, so the two are worth separating. Hold everything in desk minutes a month, the only unit in which the two effects add up cleanly.
The business case sat at 14,190 minutes: 1,290 exceptions at 11 minutes each. Now take the rate shortfall on its own. At 65.0 per cent instead of 85, 3,010 files stop rather than 1,290, and if each still took 11 minutes that would be 33,110 minutes. The rate alone therefore added 18,920 minutes. Now add the residue. The same 3,010 files take 19 minutes rather than 11, giving 57,190 minutes, so the residue added a further 24,080. The total gap of 43,000 minutes splits 44.0 per cent to the rate that was missed and 56.0 per cent to the handling time that was assumed constant, so the larger half of the shortfall was never about the rate at all.
The split is worth sitting with. Every conversation about this programme was about the rate, and the rate is the smaller half. The bigger half is a single line in a spreadsheet where somebody typed 11 into the row for after and did not think about it again.
Of the 43,000 extra desk minutes a month, which half is larger: the missed straight-through rate, or the rise in handling time?
What causes an exception, and in what proportions?
A queue of 3,010 files a month is not one thing. The queue is six things, arriving in very unequal quantities, and until somebody counts them the desk is a single undifferentiated pile that every improvement idea gets aimed at in general. Counting them takes an afternoon and changes what the next year of work is spent on.
At Sumeru Bank the six causes were counted for one steady month, and they sum to the 3,010. A document field the reading step could not read with enough confidence: 1,264 files, being 42.0 per cent. The declared income could not be corroborated against the statement: 602, being 20.0 per cent. A fraud rule fired on the servicing book: 452, being 15.0 per cent. The scoring model returned a value inside its referral bandThe range of output values where the component declines to decide either way and the file has to go to a person.: 391, being 13.0 per cent. An identity record did not match across two sources: 189, being 6.3 per cent. The consent record was incomplete: 112, being 3.7 per cent.
The largest three causes carry 2,318 files between them, being 77.0 per cent of the queue, and the smallest carries 112, so effort spread evenly across all six is spent mostly where it cannot help. Halving the smallest cause saves the desk 56 files a month. Halving the largest saves 632. The two projects will take about the same amount of somebody's time to scope, and one of them is worth eleven times the other.
Which single cause carries the largest share of the 3,010 exceptions?
Which causes come from a written rule, and which from a learned component?
Sort the same six causes a second way and the desk turns into two different jobs. Two of them, the unreadable document field and the score in the referral band, come from components whose behaviour was fitted to past examples. Together those are 1,655 files, being 55.0 per cent. The other four, income corroboration, the fraud rules, the identity match and the consent record, come from rules a person wrote down. Together those are 1,355 files, being 45.0 per cent.
Why does that matter to somebody sitting at the desk? Because for the second group there is a document. The four written rule sets run to 126 lines between them: 34 for income corroboration, 61 for the fraud rules, 22 for the router and 9 for the identity match. A reviewer can print those, read them and say precisely which line stopped this file and what would have to be different for it not to. For the learned group there is no line to point at, so the person is told what the component produced rather than why, and the work becomes supplying the missing input rather than reading the reason.
Agrawal, Gans and Goldfarb make the useful framing here: a fitted component supplies a prediction, and somebody still has to act on it. The prediction arrives as a number and a confidence, not as an argument. An exception raised that way reads differently from one raised by line 17 of a written procedure.
What does handling one exception actually consist of, minute by minute?
Handling timeThe hands-on minutes a person spends on one exception, which is not the same as the elapsed time the applicant waits. of 19 minutes sounds like a slower version of the old 11 minutes. It is not. The 19 minutes is a different set of activities on a different set of files, and reading it as the old job plus eight minutes is how a desk gets staffed wrongly.
The old 11 minutes was routine handling on every file that came in: checking identity documents, sorting and filing, keying fields into the system, recording the decision, drafting and sending the letter, raising the disbursal instruction. Seven small tasks, each of them the same on every file, at 1.0, 2.0, 1.5, 3.0, 1.0, 1.5 and 1.0 minutes, summing to exactly 11.0.
The machinery now does all seven, so the new 19 minutes contains none of them. The 19 minutes breaks into 4 minutes to re-read the file from the beginning, 6 to find what the machinery could not, 5 to obtain the missing information, 3 to decide or refer, and 1 to record. The five slices sum to 19. The largest single slice is finding what the machinery could not find. Finding is investigation rather than decision, and it is the slice a badly built exception queue makes longer.
How to design an exception-handling workflow: what goes in front of the person?
Six minutes of every nineteen goes on finding what the machinery could not. The design of the queue controls that slice, and it is the only one of the five a queue can make shorter or longer without changing anything about the file. So the design question is narrow and answerable: what does the person need in front of them at the moment the file opens?
Five things, and the list does not grow. The file itself, the reason it stopped, what the component read, what the component produced, and what the person must now supply or decide. Any queue that omits one of those five sends the person hunting through another system for it, and hunting through another system is a minute and a half, every time, on every file.
| What the queue must show | What it prevents |
|---|---|
| The file, complete, including the images as uploaded | Opening a second system to see what the applicant actually sent |
| The reason it stopped, named as one of the six causes | Re-deriving the cause by reading the whole file from the top |
| What the component read, as the value it received | Guessing whether the input was wrong or the reading of it was |
| What the component produced, as the value and the confidence | Arguing about the outcome without knowing what was returned |
| What is being asked: supply this, or decide this | A person deciding a file where nobody wanted a decision |
Notice that the fifth line is the one that turns a queue into a workflow. A file that says supply the missing payslip is a fetch task, and it can be handed to whoever is free. A file that says decide whether this applicant is acceptable is a judgement task, and it cannot. Marking every file with which of the two it is takes one field, and it is the field most exception queues do not have.
There is a household version of the bad queue. Somebody hands over a bank envelope and says there is a problem with it. The envelope then has to be read from the top to find the problem, and half that reading time is spent arriving at a fact the person who handed it over already knew. Reading it from the top is the six minutes.
An exception queue shows the file reference and a status code and nothing else. Which of the nineteen minutes does that lengthen?
How to build human review into an automated finance workflow: who gets a second pair of eyes?
Every exception already has one person on it. The question is which of them warrants a second reviewerA further person checking an exception, given only where a person is making the decision rather than fetching something. as well, and the answer that most banks reach for is the wrong one: the important files, or the large ones, or the ones above some amount.
Sumeru Bank used a different test, and it is the one worth learning. A second pair of eyes only adds anything where the first pair is exercising judgement, so ask what the first person is actually doing on this file. On five of the six causes the person is supplying something the machinery could not obtain: a legible payslip figure, a matching identity record, a completed consent, a corroborating statement line, a fraud alert cleared against the servicing history. All five are fetching, and the answer is either found or it is not. A second person fetching the same fact again finds the same fact and has added cost without adding anything else.
On one cause the person is deciding. When the scoring model returns a value inside its referral band it has, in effect, declined to decide, and a person now makes the credit decision on that file. The referral band is 391 files a month, being 13.0 per cent of the 3,010, and from month 5 those are the files that get a second reviewer. The other 2,619, being 87.0 per cent, get one.
Which of the six causes warrants a second reviewer, and on what ground?
Why is a second reviewer on everything worse than a second reviewer on some?
The tempting answer to a control question is always more control everywhere. Putting a second person on all 3,010 leaves nobody able to call the bank light touch. A second person on all 3,010 costs two things, and the second one is the one that actually matters.
The first cost is arithmetic. The fetched fact is either right or it is not, and it will be checked downstream anyway when the file completes, so a second pass on the 2,619 fetching files buys nothing. A second pass is a large amount of somebody's month spent producing agreement.
The second cost is that a control which almost never finds anything stops being a control. The reviewer learns from several hundred empty reviews that there is nothing to find, and starts signing. The signature is still there on every file. The looking is not. And the files where the looking mattered, the 391 where a person actually decided something, are now reviewed by somebody who has been trained by the other 2,619 to expect nothing. Spreading a control evenly across a queue is one of the more reliable ways to switch it off.
O'Neil is useful here, in Weapons of Math Destruction, on the general shape of the problem: the errors of an automated system do not fall evenly across the people it processes, so a control spread evenly across its outputs is not aimed at where the harm concentrates. The lesson for an exception desk is narrower but the same. Aim the review at the files where a person is exercising judgement, and leave the fetch tasks to one person and the downstream check.
Why is a second reviewer on every exception worse than a second reviewer on some of them?
What does a rising straight-through rate cost, and who pays it?
The desk fell from 12 posts to 7, so five posts were saved. At an assumed fully loaded Rs 9,00,000/- a year each, that is Rs 45,00,000/- a year. The machinery costs Rs 65,00,000/- a year to run, and it cost Rs 2,40,00,000/- to build once. The running cost alone exceeds the headcount saving by Rs 20,00,000/- a year, before a single rupee of the build cost is counted, so this does not pay back on headcount and no honest paper should say it does.
A programme that does not pay back on headcount is not thereby a bad thing to build. The machinery pays back somewhere else. 5,590 applicants a month now get an answer in about 4 minutes instead of waiting 2 working days, and the per-file work now lands on the machinery rather than on a person, so the same desk of 7 can absorb a great deal more volume than the old desk of 12 could. Both of those are real. Neither of them is a headcount line, and a paper that promises a headcount line and delivers a service line has still misled the people who approved it.
There is a further cost that arrives later and does not appear in any of these numbers. The core system has no other way in, so two of the eleven original steps are now carried out by an automation that operates it through its screens. The automation broke 14 times in the period, every time after a screen changed, and each break took about 4 hours to restore. The breakages cost 56 hours of somebody's year that nobody costed.
How a lender, an analyst or a board actually reads this
Four measures get quoted about a programme like this, and each of them is half a measure. The straight-through rate on its own, the volume processed, the cost a file, and the posts removed. Each becomes honest the moment its missing half is put beside it, and a reader whose only skill is asking for the second half will catch most of what goes wrong here.
The rate, with the business case beside it: 65.0 per cent against 85. The time to an answer, for both groups rather than on average: about 4 minutes for 65.0 per cent and 2 working days for the other 35.0 per cent. The handling time on what is left: 11 minutes rising to 19. The total cost including running it: Rs 65,00,000/- a year to run against Rs 45,00,000/- a year of posts saved. Where a paper gives the first of each pair and not the second, the second is the one worth asking for.
Whose experience of this did not improve at all?
Time to answerElapsed time from application to decision, as the applicant experiences it, rather than hands-on minutes at a desk. is the measure an applicant would recognise, and it is the one where the arithmetic goes furthest wrong. On the bank's own assumption of an eight hour working day, 65.0 per cent of the files that reached the decision engine now come back in about 4 minutes and 35.0 per cent still take 2 working days. Weight those and the average time to answer is about 339 minutes.
Nobody waits 339 minutes. The two groups sit at about 4 minutes and at 2 working days with nothing in between, so not one applicant in the month had that experience and no applicant ever will. An average across a distribution with two clumps and nothing in the middle describes a bank no applicant ever met, and a report quoting only that average has described a bank in which nobody waits.
Be exact about who the second group is. The second group is 3,010 people a month: 35.0 per cent of the 8,600 files that reached the decision engine, and 30.1 per cent of the 10,000 applications started in the month. The 3,010 are not a rounding error and not a residual category. The 3,010 are more than one applicant in three, they waited exactly as long as they would have waited if none of this had been built, and several hundred of them are waiting because a photograph taken in a dim room could not be read. The programme was worth doing for the 5,590. The programme did nothing whatever for these 3,010, and a bank that stops measuring the second group because the average looks good has stopped being able to see them.
The average time to an answer fell dramatically. How many applicants experienced that average?
Before the control below is moved: the rate rises from 65.0 to 90 per cent. Does the unreadable document field stay at 42.0 per cent of what is left?
Raise the rate, and watch what the desk is left holding
One control: the straight-through rate, from 50 to 90 per cent, against a fixed 8,600 files a month. One consequence: the six causes redrawn. The six do not fall together. Two views are available. Counts shows every cause shrinking on a fixed scale. Share of the queue shows the mix the desk is left holding, and the mix is what changes. The default below is the bank's real month: 65.0 per cent, giving 3,010 exceptions of which 1,264 are unreadable document fields, being 42.0 per cent and the dominant cause. At 90 per cent there are 860 exceptions and the largest cause is down to 29.1 per cent, with three other causes within striking distance of it.
Straight-through rate: 65.0 per cent
At a straight-through rate of 65.0 per cent, 3,010 files a month stop for a person. The largest cause is an unreadable document field at 1,264 files, being 42.0 per cent of the queue and 22.0 points ahead of the second largest.
Move the control to 90 and look at the share view. At 65.0 per cent the desk has one problem: unreadable document fields, 42.0 per cent of everything, 22.0 points clear of anything else, and an improvement programme aimed at document capture is obviously the right call. At 90 per cent that same cause is 29.1 per cent and only 7.0 points clear. The document field cause fell 80.2 per cent and the referral band cause fell only 51.4 per cent. The causes automation removes first are the ones caused by bad inputs, and what survives a high straight-through rate is the work that needed judgement in the first place, so the desk at 90 per cent is doing four different jobs rather than a smaller version of one. A desk doing four jobs is a different desk, with different skills on it, and a plan that assumes the same seven people doing less of the same thing has planned for a queue that will not arrive.
The error that gets made, and what it costs
The paper that authorised the intake machinery held two numbers in one row. The paper assumed 85 per cent of files would go straight through, and it assumed the files that did not would take 11 minutes each. Eleven minutes was what a file took on the day the paper was written. The first assumption was optimistic and everybody argued about it. The second was not an assumption anybody noticed making, and it was the larger error.
The desk was told it would fall from 12 posts to 2. The desk fell to 7. Ismail Sheikh, who runs the exception desk, spent the first two quarters after go-live explaining a shortfall that was arithmetic rather than performance: 18,920 extra minutes a month from the rate and 24,080 from handling time nobody had re-estimated, being 43,000 minutes and 5.12 posts.
The cost was not the five posts. The cost was that the desk spent two quarters defending itself over a line in a spreadsheet. Meanwhile the 3,010 applicants a month still waiting 2 working days for an answer were not on anybody's report at all. The average time to an answer had improved, and the average was what the report carried.
Who sets expectations on a lender running a chain like this?
A bank deciding retail loan applications in India sits under the Reserve Bank of India. The Reserve Bank publishes its expectations on outsourcing, digital lending, customer data and consent, and on how an applicant waiting on a decision is to be treated, at rbi.org.in. Where the deployer is a market intermediary rather than a lender, the Securities and Exchange Board of India sets the equivalent expectations at sebi.gov.in. Requirements, thresholds and effective dates move, and the current position at the issuing body's own site governs.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering outsourcing, digital lending, customer data and consent, and the treatment of an applicant awaiting a decision | rbi.org.in |
| Securities and Exchange Board of India | Equivalent expectations where the deployer of such a chain is a market intermediary | sebi.gov.in |
| Bank for International Settlements | International supervisory material on the deployment of automated decision chains by banks | bis.org |
| Agrawal, Gans and Goldfarb | Prediction Machines, on a fitted component supplying a prediction that a person must still act on | Harvard Business Review Press |
| O'Neil | Weapons of Math Destruction, on an automated system distributing its errors unevenly across the people it processes | Crown |
Sumeru Bank Limited and Ismail Sheikh are invented.
Educational material. Not advice on any investment, tax, budget or market position.
