How to Build Human Escalation Into a Fraud-Alert Workflow
Building escalation into an alert workflow means settling three things before anything runs: which cases cannot wait for the queue, who they go to and what happens to the payment meanwhile. The build takes seven steps, each producing something the next one needs. At Sumeru Bank Limited, an invented bank, five written criteria sent 63 of one month's 540 kept cases to an investigator the same working day, carrying 19 of the month's 27 confirmed cases.
Underneath all seven steps sits one small idea. A queue puts cases in the order they arrived, and some cases lose their value while they wait in it. Everything below is an attempt to write down, in advance, what waiting would cost, and to arrange matters so that somebody can check afterwards whether the writing was any good. An escalation route is not a faster desk but a different order of service, and an order of service is a design decision that has to be made on paper before a single case turns up.
Why escalate at all, rather than working the queue in order?
Stand in the queue at a passport office for a moment. Everybody has a token, the tokens are called in order, and nobody can jump. The token system is doing something valuable: it is being fair, and it is being fair in a way that any person standing in it can see and check. A queueCases waiting for a person, worked in the order they arrived. is a fairness machine, and it is very good at that one job.
Now put one more person in that hall. Her flight leaves in two hours. Waiting costs her the whole trip. For the person ahead of her it costs forty minutes of a free afternoon. A queue only knows arrival time, and arrival time carries no information at all about what waiting costs. So the queue has no way to know any of that. Taking her out of the line and dealing with her first is escalationSending a case out of the queue because waiting would cost something., and it is the one thing a queue cannot do for itself. A queue orders cases by when they turned up and an escalation route orders them by what waiting would cost, and the two agree only by accident.
At Sumeru Bank Limited, triage keeps 540 cases in a month for a fuller review, and each of those 540 gets the same forty minutes of somebody's attention whatever order they are worked in. Worked purely in arrival order, a case where the money left the bank yesterday and cannot be pulled back sits patiently behind a case where the payment is still sitting inside the bank and can be stopped with an internal entry. Nothing is broken. Nobody has been careless. The arrangement is simply sorting on the one thing that does not matter here.
What question is an escalation criterion actually asking about a case?
What are the seven steps, and what does each one hand to the next?
The order matters more than any single step in it, so here is the whole build before any of it is unpacked. Name what waiting costs. Decide who receives an escalated case. Write the criteria. Decide what happens to the payment. Record which criterion fired. Measure whether the criteria select anything. Review the weakest one out loud. Each step produces a concrete thing, and the next step cannot be attempted without it. Skipping one does not slow the build down so much as leave the next question unanswerable.
Taken out of order, the steps fail in predictable ways. Writing the criteria first, before naming what waiting costs, produces the criterion that is easiest to agree in a meeting rather than the one that catches the case that cannot wait. Skipping the record of which criterion fired leaves step six with nothing to measure, so the criteria stay in force forever with nobody able to say what any one of them contributes. The shape of the build is seven steps, seven outputs, and one honest measurement at the end that only exists because of a field somebody added at step five.
Step one, what does waiting actually cost, case by case?
Start with a household example. The step is easier to feel than to define. A wedding is being planned. Two things are outstanding on Monday morning: the hall is not yet confirmed, and the tent order is not yet placed. Leave the tent order until Wednesday and absolutely nothing happens. The tent people have stock and will take the order whenever it comes. Leave the hall until Wednesday and it may be gone. Somebody else is asking about the same date. Both jobs matter. Only one of them is damaged by waiting.
Step one produces a written list of the situations in which waiting one working day changes the outcome, and the whole of the rest of the build is downstream of that list. At Sumeru Bank Limited, that conversation was held between Ismail Sheikh, who runs the exception desk, and Revathi Balan, who heads retail credit, and it was held with no reference to how much money any case involved. The question in the room was narrow on purpose: name a situation where tomorrow morning is too late.
Four kinds of answer came back. The money is already out of the bank, so tomorrow there is nothing left to stop. The account holder is already on the phone, so tomorrow the bank is answering a person who has been waiting a day for a reply it could have given today. The account is new enough that no pattern of ordinary use has been established at all. And the same beneficiary is showing up across several unrelated accounts at once, a pattern that stops being visible once the week moves on. A fifth answer, the size of the amount, came in later and is where this walkthrough ends up.
Triage has already judged which cases are most likely to be fraud, and kept the 540. Step one does not ask that question again, and does not ask which cases are biggest. The step asks one question, in one shape, and throws away every suggestion that cannot answer it. Most of what people first propose in that meeting will not survive the filter, and that is the filter doing its job rather than the people being wrong.
Step two, who does an escalated case go to, and when are they there?
An escalation with no named receiver is a piece of paper moving sideways. Step two produces two things and neither is optional: the name of the person an escalated case lands on, and the hours during which that person is actually at a desk. A same-day route that delivers a case to an empty chair at half past five is a queue wearing a different name, and it will look identical in every report the arrangement produces.
Think of a hospital that has an emergency room and a doctor on call for it. The route only means something because somebody is rostered onto the other end of it. Take the roster away and the sign above the door still says emergency, the corridor still exists, and every patient still waits until morning.
At Sumeru Bank Limited, the fraud desk is five people. The number of posts comes from the month's own work: 6,120 alerts read for ninety seconds each is 9,180 minutes, 540 kept cases at forty minutes each is 21,600, and 27 investigations at about three hours each is 4,860. The three together are 35,640 minutes a month. At an assumed working month of 8,400 minutes a person, that is 4.24 posts, so the desk is staffed at five. The investigatorThe person an escalated case goes to. At this bank, the person who carries out the three hour investigation rather than the forty minute review. at the end of the same-day routeThe path a case takes when a criterion is met, reaching a person on the working day the case was kept rather than the next one. is drawn from those same five. Saying so plainly settles an argument that otherwise runs forever.
Here is the part that surprises people. Sending 63 of the 540 kept cases down the same-day route adds no minutes at all to that 35,640. Every one of the 540 was always going to get its forty minutes, and 63 of them at forty minutes is 2,520 of the 21,600 already inside the total. Escalation at this bank changes when the work happens and not how much of it there is. The design can therefore be argued for without asking anybody for a budget. The reading is arithmetic on the desk figures rather than a separate measurement, and it holds only while the route reorders cases inside the same 540 rather than pulling extra cases in.
Step three, how is a criterion written that a person can apply inside the first read?
Step three produces the written criteria themselves. The constraint that shapes every word of them is not correctness, it is speed, and that constraint comes from the desk rather than from anybody's preference. The person applying a criterionA written test that sends a case to the same-day route when it is met. at this bank is working at the pace of a first readThe 90 second look at an alert, inside which a criterion has to be applicable., ninety seconds an alert, and has just decided to keep this one. The criterion has to be answerable in the seconds that follow, on the screen already open, or it will not be answered at all.
So the criteria at Sumeru Bank Limited are written as flat statements of fact about the case in front of the person, and every one of them can be answered by looking. Has the money left. Has the customer called. When was the account opened. Has this beneficiary appeared elsewhere this week. Is the amount at or above the value the bank set for itself. A criterion is a sentence somebody reads under time pressure, so the test of a good one is whether it can be answered from what is already on the screen, and every other virtue comes second to that.
What makes a criterion applicable in ninety seconds, and what makes one useless?
Three properties, and a criterion missing any one of them will quietly stop being applied. A criterion has to be answerable from what is already on the case, so nobody has to open a second system. It has to be answerable yes or no, so nobody has to form a view. And it has to be answerable without asking anybody, so nobody has to wait for a reply from a person who is themselves in a queue. A criterion that is perfectly correct and takes four minutes to apply produces an escalation route that looks designed on paper and in practice fires on whatever happened to be easy that afternoon.
Good intentions go wrong at exactly this point, so it is worth being blunt. Somebody will propose a criterion of the form: escalate where the transaction is inconsistent with the customer's usual behaviour. The sentence is a reasonable one. It is also unanswerable in ninety seconds, unanswerable as a yes or no, and unanswerable without opening something else, so it fails all three properties at once. The result is not that cases get escalated badly. The criterion stops being read, nobody records that it stopped being read, and it survives in the document long after it stopped existing on the desk.
A proposed criterion needs the person reading the case to open a second system and check something there. Will it be applied?
Step four, what happens to the payment while the case is open?
This is the step that gets left until last and belongs fourth, because it is the only one where the design reaches outside the bank and touches somebody who has done nothing wrong. Step four produces two sentences: whether the payment on an escalated case is heldStopping a payment while the case on it is open. while the case is open, and who may release it, in what time. The hold rule is where the whole arrangement stops being an internal workflow question and becomes something a person outside the bank experiences directly.
Picture a small trader paying a supplier on the last day before a festival closing. The payment stops. Nobody has accused her of anything, nobody has told her anything, and the goods do not move. The stopped payment is not a hypothetical cost, it is an afternoon of her business, and it lands whether or not the case turns out to be anything at all.
At Sumeru Bank Limited, the numbers for one month are these. Of the 540 cases triage kept, the bank held the transaction on 218 while the case was open. After review it released 191 of them. 218 less 27 is 191, and the 27 are the cases an investigation confirmed. So in one month 191 people had a lawful payment stopped and were doing nothing wrong, being 87.6 per cent of the holds, and that figure is a cost of this design rather than an accident that happened alongside it. Nobody at this bank has ever put a rupee figure on that cost, so it appears in no budget line.
Two things follow for the design, and both belong in step four rather than in a later apology. The first is that whoever may release a held payment has to be named, and the time within which they will look at it has to be written down. An unbounded hold does more damage than any other form this design takes. The second is that the hold rule and the criteria are separate decisions. A case can be escalated without the payment being held, and at this bank most kept cases were not held at all: 322 of the 540. How the 218 holds distribute across the same-day route and the queue is not something this bank recorded.
Where the rules on monitoring and on a stopped payment come from
The international standard under which a regulated institution monitors transactions originates with the Financial Action Task Force, at fatf-gafi.org. The Task Force sets the standard rather than the rules any one country applies. The rules that apply to a bank in India, including how a customer whose payment is stopped is to be treated, are stated by the Reserve Bank of India at rbi.org.in. Where the institution running an arrangement like this is a market intermediary rather than a bank, the equivalent expectations are stated by the Securities and Exchange Board of India at sebi.gov.in.
The five criteria, the thirty day account age and the money value are the bank's own design choices, set by its own people, and none of them is a standard.
What has to be decided about the payment before a single criterion is written?
Step five, what does the record have to hold on every escalated case?
Step five produces one field. Not a report, not a dashboard, one field on the case: which criterion fired. The field is written at the moment the case is escalated, by the person escalating it, and it is never reconstructed afterwards. One field is the difference between an arrangement that can be reviewed and an arrangement that can only be defended, and it costs the person about two seconds.
Reconstruction does not work, and the reason is concrete rather than a matter of principle. Six months later somebody asks whether the criterion about new accounts is earning its place. Without the field, the only way to answer is to go back over the escalated cases and work out which criteria each one would have met. Reconstruction answers a different question: a case can meet three criteria, and the question is which one actually sent it. Any one case might have been escalated on the money having gone, and also happens to sit on an account opened three weeks ago. Reconstructing it credits the new-account criterion with a case it never sent.
Why record which criterion fired, rather than simply recording that the case was escalated?
Step six, how can anyone tell whether the criteria select anything at all?
Step six produces a comparison, and it is the first moment in the build where the design can be wrong in a way anybody can see. The cases that went down the same-day route and the cases that waited in the queue have their confirmation rates set side by side. If the two rates are the same, the criteria are sorting cases and nothing more. A route that picks out no difference has picked out nothing. Picking out a difference is what selectivityWhether the criteria pick out cases that differ from the ones they leave behind. means here, and it is a property of the criteria rather than of the people working them.
At Sumeru Bank Limited, the month reads as follows. Of the 540 kept cases, 63 met at least one criterion and went same day, being 11.7 per cent, and 477 waited for the next working day. 63 plus 477 is 540. Of the month's 27 confirmed cases, 19 came out of the 63 and 8 came out of the 477. 19 plus 8 is 27. So 11.7 per cent of the cases carried 70.4 per cent of the confirmations, and the same-day route confirmed at 30.2 per cent against 1.7 per cent for the queue, about eighteen times the rate. The two rates are the output of step six, and neither of them exists without the field added at step five.
Two cautions travel with that comparison and neither is optional. The first is that a difference this size does not establish that the criteria caused anything. Cases that meet a criterion are different cases to begin with, and making them different is the whole point of writing the criteria, so this measurement cannot separate the selection from the speed. The second is that both numbers describe one month at a single bank, and 27 confirmations is a small number to build any rate on. Step six gives a comparison that could not previously be made at all, not a proof.
How would anyone tell whether a set of escalation criteria is selecting anything?
Which five criteria did one bank arrive at, and what did they catch?
Here is the output of steps one to three at Sumeru Bank Limited, in the order the bank added them. Any one of the five being met is enough to send a case down the same-day route, and that matters for how the cumulative reading below works. Each row shows what that criterion added on top of everything above it.
| In the order added | The criterion, as written | Cases | Confirmed |
|---|---|---|---|
| 1 | The money has already left the bank and cannot be reversed by an internal entry | 21 | 8 |
| 2 | The account holder has already contacted the bank about the same transaction | 33 | 11 |
| 3 | The account was opened within the last 30 days | 44 | 14 |
| 4 | The same beneficiary appears on alerts across three or more unrelated accounts within one week | 55 | 18 |
| 5 | The amount is at or above a value the bank set for itself | 63 | 19 |
The cumulative figures carry the finding, so read down the two number columns rather than across the rows. The cases run 21, 33, 44, 55 and 63. The confirmations run 8, 11, 14, 18 and 19. The first criterion on its own accounts for a third of the route's cases and 42 per cent of everything the route confirmed. The last one adds 8 cases and 1 confirmation. The five criteria contribute very unevenly, and that shape only appears because each escalated case carries the field from step five.
One honesty note before the shape is drawn, and it belongs in the open rather than in a footnote. A case can meet more than one criterion, so adding them in a different order would hand the overlaps to different criteria and change every row above except the last. The readings above are cumulative in the order this bank added them and describe that order, and no reading of any criterion standing entirely alone exists for criteria 2 to 5 at this bank.
The same unevenness reads more sharply as a price. Dividing the cases each criterion added by the confirmations it added gives what that criterion cost in cases for each confirmed case it brought in.
| Criterion added | Cases added | Confirmations added | Cases per confirmation |
|---|---|---|---|
| 1, the money has already gone | 21 | 8 | 2.63 |
| 2, the account holder has called | 12 | 3 | 4.00 |
| 3, the account is under 30 days old | 11 | 3 | 3.67 |
| 4, the beneficiary is on three accounts | 11 | 4 | 2.75 |
| 5, the money value | 8 | 1 | 8.00 |
The four columns reconcile: 21 plus 12 plus 11 plus 11 plus 8 is 63, and 8 plus 3 plus 3 plus 4 plus 1 is 19. The fifth criterion costs 8 cases for each confirmed case it brings, against 2.63 for the first, so it is roughly three times the price of the criterion it was added after. The per-step prices are arithmetic on the cumulative readings rather than separate measurements, and like the readings they belong to this order of adding.
Before the control below is moved: which of the five criteria adds the fewest confirmed cases?
Switch the criteria on one at a time and watch the second bar stop moving
One thing moves: how many of the five criteria are in force, added cumulatively in the order this bank added them. Everything else is held where the month put it. Triage still keeps 540 cases, the month still confirms 27, and every kept case still gets its forty minutes.
With all five criteria in force, 63 of the 540 kept cases go to an investigator the same working day and carry 19 of the month's 27 confirmed cases, so the route confirms at 30.2 per cent while the 477 left in the queue confirm at 1.7 per cent.
Which criterion turned out to be the weakest, and why is that surprising?
Look at the last row of the control again. The money value criterion adds 8 cases to the same-day route and 1 confirmed case, and it is the criterion that almost every design writes first. The money value criterion is the easiest of the five to state, the easiest to agree in a room full of people who disagree about everything else, and the only one that sounds like risk when it is read out. The one criterion everybody writes first is the weakest of the five here, and the reason is structural rather than careless.
The money value criterion, and what it is actually sorting
The value of a transaction states what is at stake. The value says nothing whatsoever about whether waiting until tomorrow changes the outcome, and escalation is a question about waiting. A large payment that can still be stopped on Monday morning needs no same-day route at all, and a small payment whose money left the bank an hour ago cannot wait until Monday for anybody.
So an arrangement built on the value criterion alone is not a weaker escalation. The value criterion alone gives a different thing wearing the same name: a queue re-sorted by size. A queue re-sorted by size will look completely correct in every report, will fire reliably, will be easy to explain to anybody who asks, and will put the reversible cases at the front on the morning it matters.
At Sumeru Bank Limited, the cost of finding this out was one field on a case record and one month of counting. The cost of not finding it out is that nothing about the criterion ever looks wrong, so it stays first in the document forever.
Why does the money value criterion contribute so little at this bank?
Step seven, which criterion is reviewed first, and what is said out loud?
Step seven produces a decision, taken by a named person and written down. Not a conclusion, a decision, and the difference matters because the obvious conclusion here is wrong. Finding the weakest criterion is not an instruction to remove it; it establishes that keeping it is now something somebody is choosing rather than something nobody knows they are doing.
There are real arguments on both sides at Sumeru Bank Limited, and a review that does not put both on paper has not done step seven. Against keeping the money value criterion: it added 8 cases and 1 confirmation, it costs the desk the same forty minutes on each of those 8, and where a payment is held on one of them it costs somebody outside the bank an afternoon. For keeping it: 8 cases a month is a small cost, and this is the criterion most likely to fire on the large single loss that has not happened yet and that no reading of one month can see. Neelima Rao, in the risk function, is not the person who built any of this, and reviews like this one land with her precisely because she did not.
Step seven forbids a third option, the common one. The third option is to keep the criterion, say nothing, and let the next person to read the document assume the order of the five carries some meaning. The criterion that stays in force because nobody looked is indistinguishable, on paper, from the criterion that stays in force because somebody looked hard and decided to keep it, and only one of those is a design.
The criteria are measured and the weakest one turns out to contribute almost nothing. What follows?
What does the escalation itself owe the record, and who reads it?
Everything above produces documents, and the documents outlive the people. At Sumeru Bank Limited, the parts the record has to carry are exactly the outputs of the seven steps: the list of what waiting costs, the name of the receiver and the hours, the criteria as written, the hold rule with who may release and how quickly, the field on each escalated case, the two confirmation rates for the month, and the dated decision on the weakest criterion with its reasons.
Read that list again and notice that only two of the seven items are about catching fraud. The other five are about who did what, when, and on whose authority. The balance is not an accident of this bank. An escalation route is a standing arrangement to treat some people's payments differently and faster than others, and an arrangement like that has to be able to say why, in writing, to somebody who was not in the room. Ashok Pillai, in technology risk, ran the register sweep at this bank that found arrangements in use which nobody had written down, and the reason a sweep finds things is that undocumented arrangements do not stop working, they only stop being visible.
How would somebody reviewing an arrangement like this read it in an hour?
Suppose an analyst is handed a fraud alert workflow they did not build and given an hour. The seven steps make a serviceable order to ask questions in, and each question has a document behind it or does not.
The first thing to ask for is the written list of what waiting costs, and if there is no such list, everything after it was written from somebody's memory of a meeting. Next comes who receives an escalated case and when they are at a desk, with the two answers checked against each other. Then the criteria themselves, read for the three properties rather than for correctness: answerable from the case, answerable yes or no, answerable without asking anybody. Then what happens to the payment, who may release it and how fast, followed by how many payments were held last month and how many were released. The pair of numbers on holds is where this design touches people who did nothing wrong. Then whether the case record carries which criterion fired. Last come the two confirmation rates. The final question is when the weakest criterion was last reviewed, and if nobody can produce those rates the answer is already settled.
One sentence governs how every number produced in that hour should be read. The arrangement at Sumeru Bank Limited is right 0.15 per cent of the times it speaks, being 27 confirmed cases out of 18,000 alerts, and that is the design working rather than failing. The two errors do not cost the same thing. A missed case at this bank carries an average amount at risk of Rs 84,000/-, an amount the bank can name and budget for. An alert on a lawful payment costs the desk ninety seconds and costs somebody outside the bank an afternoon, and 191 people paid that second cost in one month while doing nothing wrong. Both halves of that go together every time. The first half alone condemns a design that is working, and the second half alone will excuse any amount of wrongness at all.
Sources
| Source | Document | Site |
|---|---|---|
| Financial Action Task Force | The origin of the international standard under which a regulated institution monitors transactions and escalates what it finds. The Task Force sets the standard and not the rules a bank in India is held to. | fatf-gafi.org |
| Reserve Bank of India | Published expectations on a regulated lender covering fraud monitoring, outsourcing, digital lending, data and consent, and the treatment of a customer whose payment is stopped. Together they are the rules a bank in India is held to. | rbi.org.in |
| Securities and Exchange Board of India | The equivalent expectations where the institution running an arrangement like this is a market intermediary rather than a bank | sebi.gov.in |
| Agrawal, Gans and Goldfarb | Prediction Machines, 2018, on the split between what a system produces and what a person still has to decide | Harvard Business Review Press |
| Cathy O'Neil | Weapons of Math Destruction, 2016, for a system's errors falling unevenly across the people they land on | Crown |
Sumeru Bank Limited, Revathi Balan, Ismail Sheikh, Neelima Rao and Ashok Pillai are invented.
Educational material. Not advice on any investment, tax, budget or market position.
