Alert Triage and Escalation: Sorting Signal From Volume
Alert triage is the sorting that stands between a monitoring system and an investigator. At Sumeru Bank Limited, invented, four stages handle one month: 11,880 of 18,000 alerts close automatically as repeats, 6,120 get a 90 second read, 540 get 40 minutes and 27 become investigations. Escalation is the separate question of which of those 540 cannot wait.
A monitoring arrangement that raises eighteen thousand messages a month and puts five people behind them is not solving a detection problem but an attention problem. Detection already happened, in the sense that something matched a written line and a message exists; what remains is the far harder question of how much of a working life each of those messages is worth. Every stage is a bet that a cheaper look can safely close a large share of what arrives, and the whole design stands or falls on whether each of those bets can be checked afterwards rather than asserted.
What is alert triage, and what is it actually sorting?
Think about a small shop that takes orders on a phone. By evening there are two hundred messages. The shopkeeper does not read them in order and does not read them equally. She thumbs down the list in about a second each, and what she is deciding in that second is not whether a message is true. She is deciding whether this one gets answered now, gets answered tonight, or gets nothing at all. She will be wrong sometimes. She will still be right that reading two hundred messages carefully is not available to her, and that pretending otherwise means the urgent ones wait behind the routine ones.
Alert triageSorting alerts by how much attention each one is worth before anybody investigates any of them. is that thumb moving down the list, written down as a procedure so that it happens the same way every day and can be argued with. Triage sorts attention, and never truth. The distinction does real work: deciding what can be said about a closed alert. An alert that closed at the first read has not been declared lawful. Nobody established anything about it. The decision was that the ninety seconds already spent was the right amount, and that a second look was worth less than the same time spent on the next alert in the list.
Get that backwards and two things go wrong at once. The desk starts recording closures as findings. A month's report then says the arrangement examined 18,000 transactions and found 27 problems. No such examination happened. And the people working the list start feeling that closing an alert is a judgement about a customer rather than a judgement about a queue. Carried sixty times a day, that feeling is heavier than the work needs it to be.
What is triage sorting?
What are the four stages, and what closes at each one?
Sumeru Bank Limited runs four numbered stages on the month's 18,000 alerts. The figures are that bank's own and describe one month of one deployment. Stage 1 is automatic suppressionClosing an alert automatically, with a record, because it repeats a pattern already cleared on that same account recently.: an alert repeating a pattern the desk already cleared on that same account within the last 30 days closes with a record and nobody reads it, and that is 11,880 alerts, being 66.0 per cent of the month. Stage 2 is a first readA short look at an alert, long enough to close most of them and no longer. of 90 seconds on the 6,120 that survive, of which 5,580 close there. Stage 3 is a fuller review of 40 minutes on the 540 that are kept. Stage 4 is an investigationThe work that establishes whether a case is actually fraud, rather than whether it is worth more looking at. of about 3 hours on each of the 27 that are confirmed.
The two closing identities are what make a stage list a description rather than a story, and they are worth reading slowly. Every alert closes exactly once: 11,880 plus 5,580 plus 540 is 18,000. And the 540 are the survivors of the first read: 6,120 less 5,580 is 540. A stage list that does not close in both directions is a diagram, not a design, and the difference is visible in about a minute.
| Stage | What happens | Who does it | Arrives | Closes here | Goes on |
|---|---|---|---|---|---|
| 1 | Automatic suppression of a repeat already cleared on that account within 30 days | Nobody. A record is written | 18,000 | 11,880 | 6,120 |
| 2 | A first read of 90 seconds | A person on the desk | 6,120 | 5,580 | 540 |
| 3 | A fuller review of 40 minutes on a kept caseAn alert that survives triage and is given a fuller review rather than being closed. | A person on the desk | 540 | 513 | 27 |
| 4 | An investigation of about 3 hours | An investigator | 27 | 27 | 0 |
| One month on the servicing book | 18,000 | 18,000 | 0 | ||
Look down the last three columns and notice what is happening to the price. Stage 1 costs nothing per alert and handles the most. Stage 2 costs a minute and a half and handles a third of the month. Stage 3 costs forty minutes and handles three per cent. Stage 4 costs about three hours and handles fifteen alerts in every ten thousand. The cost per alert rises by a factor of roughly a hundred and twenty across the four stages, and the volume falls by a factor of nearly seven hundred, and those two movements are the whole engineering idea. The desk buys the right to spend three hours on twenty seven things by refusing to spend three minutes on eighteen thousand.
What does each stage cost, and how big a desk does that make?
The desk falls straight out of the volumes and the rates, with nothing else added. Multiply each stage by its own price and add. The 6,120 first reads at 1.5 minutes are 9,180 minutes. The 540 fuller reviews at 40 minutes are 21,600. The 27 investigations at 180 minutes are 4,860. The three totals sum to 35,640 minutes a month. At the bank's assumed working month of 8,400 minutes a person, that is 4.24 posts, so the fraud desk at Sumeru Bank Limited is 5 people, and the fifth exists entirely because 4.24 does not round down.
| Stage | Items | Minutes each | Minutes a month | Share of the desk |
|---|---|---|---|---|
| 1 Suppression | 11,880 | 0 | 0 | 0.0% |
| 2 First read | 6,120 | 1.5 | 9,180 | 25.8% |
| 3 Fuller review | 540 | 40 | 21,600 | 60.6% |
| 4 Investigation | 27 | 180 | 4,860 | 13.6% |
| The month | 35,640 | 100.0% |
There is a second way to cut the same 35,640 minutes and it is more useful than the first. Instead of asking what each stage cost, ask what each alert cost by the time it finally closed. An alert that closes at stage 1 costs nothing. One that closes at the first read costs 1.5 minutes. One kept and then taken no further costs 1.5 plus 40, being 41.5. One that becomes a confirmed case costs 1.5 plus 40 plus 180, being 221.5. Now count the alerts by where they ended: 11,880 at stage 1, 5,580 at the first read, 513 at the fuller review and 27 at investigation. The four counts sum to 18,000. The minutes come out at 0, 8,370, 21,289.5 and 5,980.5, and those sum to exactly 35,640. The second cut is arithmetic on the bank's own locked figures rather than a new measurement, and landing on the same total from a different direction is the check that it is right.
Two readings fall out of that picture and both are useful. The average alert costs this bank 1.98 minutes from arrival to closure. A figure that small is easy to hold in the head and multiply. And the expensive half of the desk is not where the volume is: the 540 kept cases, three per cent of the month, take about sixty per cent of everything. A cheaper desk comes not from attacking the 18,000, but from attacking the 540, and the only honest way to do that is to make the first read better at deciding which cases deserve forty minutes.
The desk needs 35,640 minutes a month. Which stage eats most of them?
What does stage 1 do, given that nobody reads those alerts?
Everybody has a version of stage 1 at home. The smoke alarm outside the kitchen goes off whenever fish is fried. After the fourth time, nobody in the house runs to the kitchen; they hear it, register that it is seven in the evening on a frying day, and carry on. Running to the kitchen four times a week costs something, and finding fish there four times a week teaches nothing. The household has built a suppression rule, and a sensible one. The rule also carries the only risk a suppression rule ever carries: the evening the alarm means something else.
Stage 1 at Sumeru Bank Limited is exactly that written down. If an alert repeats a pattern the desk has already cleared on that same account within the last 30 days, it closes automatically with a record and nobody reads it. The 30 days is that bank's own choice and is not anybody's requirement. The one written line disposes of 11,880 alerts a month, being 66.0 per cent of everything the monitoring raises. The stage that does two thirds of the work is the one nobody would call detection, and it consumes no minutes at all.
Take it away and the arithmetic moves in a way that surprises people. Without stage 1, every alert reaches the first read: 18,000 at 1.5 minutes is 27,000 minutes. Suppression only ever removes repeats of patterns already cleared on that account, so stages 3 and 4 do not move. The desk then needs 27,000 plus 21,600 plus 4,860, being 53,460 minutes, or 6.36 posts and therefore 7 people. Suppression saves 17,820 minutes a month, being 2.12 posts, and in whole heads it is the difference between a desk of 5 and a desk of 7. At the bank's assumed fully loaded Rs 9,00,000/- a post, two posts is Rs 18,00,000/- a year.
Before the control below is moved: suppression closes 66.0 per cent of alerts with nobody reading them. With it switched off, how many posts does the desk need?
Move the suppression share and watch the desk resize
One control: the share of the month's 18,000 alerts closed automatically by stage 1, from 0 to 80 per cent. One consequence: the alerts reaching a person redraw on the upper bar, the desk minutes restack on the lower one, and the post count is read off the grid line the bar rounds up to. Stages 3 and 4 are held fixed at 26,460 minutes throughout, on the assumption that suppression removes only repeats of patterns already cleared on that account.
At the deployed suppression share of 66 per cent, 11,880 alerts close with nobody reading them and 6,120 reach a person, so the desk needs 35,640 minutes a month, being 4.24 posts, and therefore 5 people at about Rs 45,00,000/- a year.
How would anybody know the suppression is safe?
Here is the uncomfortable property of stage 1. Every other stage leaves behind a person who looked. A closure at the first read has somebody's name on it, and if that person was wrong somebody can go back and ask them what they saw. A suppressed alert has a record that it was suppressed and nothing else, so being wrong at stage 1 leaves exactly the same trace as being right. The only way to learn anything about it is to go back and read some of the suppressed alerts by hand. Reading them is precisely the work the stage exists to avoid.
Sumeru Bank Limited did that once. The bank pulled 400 of the month's 11,880 suppressed alerts and had them read as though they had arrived at the first read. None of them would have been kept by triage. A count of zero is worth something only because there is a number to compare it against. Among the alerts a person actually reads, 540 of 6,120 are kept, a keep rate of 8.8 per cent. Had the 400 behaved like the alerts people read, about 35 of them would have been kept. The check found zero against an expectation of thirty five, and that gap is the whole of the evidence.
| The check on stage 1 | Reading | Where it comes from |
|---|---|---|
| Suppressed alerts in the month | 11,880 | Stage 1, being 66.0 per cent of 18,000 |
| Read by hand | 400 | 3.4 per cent of the suppressed alerts |
| Keep rate among alerts a person reads | 8.8% | 540 kept of 6,120 read at stage 2 |
| Expected keeps in the 400, at that rate | 35 | 400 times 8.8 per cent |
| Keeps actually found | 0 | The hand check |
| What the same rate would imply for all 11,880 | 1,048 | Arithmetic on the locked keep rate, not an observation |
The 1,048 in the last row is the reason anybody bothered. If suppressed alerts were just ordinary alerts that nobody happened to read, stage 1 would be discarding roughly 1,048 keepable cases a month, nearly twice the 540 the desk actually keeps. The hand check argues hard against that. Zero out of four hundred is not the result a stage quietly throwing away a fifth of the desk's real work would produce.
400 suppressed alerts were read by hand and none would have been kept. What does that establish?
Why is that check smaller than it sounds?
Because 400 is 3.4 per cent of one month of suppressed alerts, and one month is one month. A hand check on a sample can show that a stage is not making a large, evenly spread mistake. A hand check cannot show that the stage is not making a narrow one. If there is one kind of alert that suppression removes systematically, and that kind is rare enough that a sample of four hundred is unlikely to contain any of it, the check comes back clean and the kind stays invisible.
There is also a cost reason nobody escapes. Reading all 400 by hand at the first read rate of 1.5 minutes is 600 minutes, about a working day and a half. Reading every suppressed alert in the month would be 11,880 times 1.5, or 17,820 minutes. The 17,820 minutes are exactly, to the minute, what stage 1 saves, so a complete audit of the suppression costs precisely as much as the suppression is worth. The identity is not a coincidence and it is not a trick; it is the definition of the stage restated. The same identity also means the only checks available are partial ones, and the honest way to describe stage 1 forever is: a stage that saves two posts, checked once on 3.4 per cent of one month, with a result that argues in its favour and cannot settle it.
What is case escalation, and how is it a different question from triage?
Everything so far has answered one question: how much attention is this worth. Case escalationSending a case out of the ordinary queue because waiting for its turn would itself cost something. answers a different one: does waiting cost anything. The two questions have different answers often enough that a design carrying only the first will reliably work the wrong cases first.
The household version is a leaking tap against a gas smell. The tap will cost more to fix if it is left, but it will cost the same amount tomorrow as today, so it belongs in the ordinary list. The gas smell may turn out to be nothing. The cost of being wrong about it grows by the hour, so it is still the one dealt with in the next four minutes. Nobody thinks the gas smell is more important than the tap. The gas smell is less patient.
At Sumeru Bank Limited the impatient cases are concrete. Of the 540 kept cases in the month, the bank held the transaction pending on 218 of them and released 191 after review, so 191 people had a lawful payment stopped in one month. Every day a kept case sits in the queueThe cases waiting for a person, worked in the order they arrived. is a day of somebody's held payment or a day in which money that has already left the bank gets further away, and neither of those costs is visible in the attention ranking triage produces. Escalation exists to close that gap.
A case is small in value and the money has already left the bank. Triage or escalation?
Which five criteria send a case the same working day?
Sumeru Bank Limited wrote down five. Any one of them, met by a case kept at stage 3, sends it to an investigator the same working day instead of into the queue. All five are that bank's own choices, and none of them is a standard, a norm or anybody's requirement.
| Criterion | What it says | Why waiting costs something |
|---|---|---|
| 1 | The money has already left the bank and cannot be reversed by an internal entry | Recovery gets harder by the hour once funds are outside |
| 2 | The account holder has already contacted the bank about the same transaction | A person is already waiting, and the bank is already late |
| 3 | The account was opened within the last 30 days | A new account has no ordinary behaviour to compare against |
| 4 | The same beneficiary appears on alerts across three or more unrelated accounts within one week | A pattern across accounts is still running while the case waits |
| 5 | The amount is at or above a value the bank set for itself | The single loss is larger if it completes |
One criterion is enough, and that choice is what makes the route usable at all. A rule requiring two or three would be more selective, and it would also take longer to apply than the ninety seconds of the read it sits inside. Taking longer than the read defeats the purpose of having it. The five are written as things somebody can check by looking at the case in front of them, not as things somebody has to work out.
How many of the five criteria have to be met to send a case the same day?
Where the expectation behind any of this comes from
The international standard on monitoring transactions originates with the Financial Action Task Force at fatf-gafi.org, and what actually applies to a bank operating in India is stated by the Reserve Bank of India at rbi.org.in, with the Securities and Exchange Board of India at sebi.gov.in where the deployer is a market intermediary. Read the position at those sites and confirm the current version there.
The four stages, the 30 day suppression window, the 90 seconds, the 40 minutes, the five escalation criteria and the money value criterion 5 compares against are all Sumeru Bank Limited's own choices, not requirements, thresholds or effective dates.
What did the same-day route catch that the queue did not?
A route is only worth building if it can be shown to be selecting something, and that demands a count almost nobody produces by default: not how many cases went each way, but how many confirmations came out of each. In the month, 63 of the 540 kept cases went same day, being 11.7 per cent, and 19 of the month's 27 confirmed cases came from those 63. Nineteen of twenty seven is 70.4 per cent of all the confirmations, arriving through 11.7 per cent of the cases. The other 477 produced 8.
Put it as hit rates and the gap sharpens. The same-day routeThe path a case takes out of the ordinary queue when an escalation criterion is met. confirmed 19 of 63, being 30.2 per cent. The queue confirmed 8 of 477, being 1.7 per cent. A case on the same-day route was about eighteen times as likely to end in a confirmation as a case in the queue, and that ratio is the only evidence anybody has that the five criteria are doing more than sorting. Without it there is a route, a policy and no way to tell whether either was worth having.
One caution before that is taken as proof the criteria are well chosen. The five are not equal. Criterion 1, whether the money has already gone, accounts on its own for 21 of the 63 cases and 8 of the 19 confirmations. Criterion 5, the money value, is the one written first in most places, and it adds 8 cases and 1 confirmation. The value threshold is the weakest of the five and it is usually the only one anybody writes down.
63 of 540 kept cases went same day and carried 19 of the month's 27 confirmations. What does that say about the criteria?
What does a case need in front of the person working it?
Ninety seconds is the entire budget for the first read, and 6,120 of them is 9,180 minutes of somebody's working life. Whether ninety seconds is enough is decided by one thing only: whether the case arrives assembled. The contents of an alert record when it is raised are set out under the fraud alert. A case has an age and an alert does not, so a triaged case needs that list plus one more item.
The eight items are the transaction, the account, the line that fired, what that line compared, what this account usually looks like, whether the payment is held, the age of the case, and who holds it. Two of those are the ones that quietly break the budget. If the case does not carry what the line compared, or what this account usually looks like, the person has to go and assemble the comparison instead of reading it. A first read that drifts from 90 seconds to four minutes adds 2.5 minutes across 6,120 alerts, being 15,300 minutes a month, and that takes the desk from 35,640 minutes to 50,940 and from 5 posts to 7. Seven people is the identical answer produced by switching suppression off altogether. A presentation fault and a missing stage cost the same two people.
Note how this list sits beside the one the exception desk on the lending side uses. The lending case carries seven numbered fields, and Ismail Sheikh's desk works them at 19 minutes each. The fraud case carries eight of its own, and only the last two are shared: the age of the case and who holds it. Everything else differs because the two desks are answering different questions, and a bank that gives both desks the same case layout has decided that consistency matters more than either of them working.
A first read is meant to take 90 seconds and takes four minutes. What is most likely missing?
How does a triage arrangement fail, and what does the failure look like?
The suppression rule that quietly stopped being right
Suppose one of the patterns stage 1 treats as an already cleared repeat starts being used by somebody who worked out that it is treated that way. Or suppose an ordinary change in how a payment type is recorded makes a new kind of alert look like an old cleared one. Nothing announces either event. The rule keeps matching, the alerts keep closing, and the records keep being written.
Now ask what moves in the numbers the desk reports. The alert count does not move. The alerts are still being raised. The suppression share does not move. The 6,120 read does not move. The keep rate does not move. The cases that would have been kept never reach the read. The confirmation count does not move. The cases that would have been confirmed were never opened. Every count the desk produces looks exactly as it did before. A suppression fault cannot be found by watching the counts, however carefully anybody watches them.
The shape of the failure is worth naming precisely: not one alert missed, but one kind of alert never seen, for as long as the rule that suppresses it stands. A missed instance shows up eventually because the loss arrives. A missed kind never does. There is nothing to compare the month against.
A suppression rule is quietly wrong. What does the failure look like in the numbers?
How does somebody running this desk read the four stages?
A triage arrangement is a capacity arrangement wearing a detection costume. Start with the margin. Five people at the assumed working month of 8,400 minutes is 42,000 minutes of capacity against 35,640 of work, so the margin is 6,360 minutes a month, being 15.1 per cent, or about 318 minutes a working day. In cases that is room for roughly eight extra fuller reviews a day. The margin exists only because 4.24 posts rounded up to 5, and reading it that way makes it obvious how fragile it is: the same desk with 4 people would be 6,240 minutes short every month.
Now put it beside the other desk in the same bank. The exception desk on the lending side runs at a margin of 4.2 cases a day, being 2.7 per cent of its capacity. A week carrying three per cent more volume turns its stable queue into a growing one. The fraud desk sits at 15.1 per cent. Two desks in one bank, one with five times the other's headroom, and neither number appears in any report either desk produces. Working it out takes about four minutes and a locked working month, and it is the first thing anybody reviewing either arrangement should compute.
Then ask for four things, in this order. First, the two closing identities. A stage list that does not close is not a description of anything. Second, the date and result of the last hand check on suppressed alerts, with the expected count beside the found count. A check reported without its expectation is a number with nothing to lean on. Third, the confirmation counts by route, not just the case counts. The 11.7 per cent of cases carrying 70.4 per cent of the findings is the only evidence the criteria work. Fourth, the measured first read time against the 90 seconds the desk was sized on. The measured time moves the desk by two posts before anybody notices it has moved at all.
Covered elsewhere. The six numbered lines that raise the month's 18,000 alerts in the first place are set out under alert sources. How to build the escalation into a workflow step by step, and how the five criteria behave when they are added one at a time, are covered under escalation design. Scoring a transaction by how far it sits from an account's own pattern is covered under anomaly scoring, and how false positives are counted and reported as a measure is covered separately.
Sources
| Source | Document | Site |
|---|---|---|
| Financial Action Task Force | Origin of the international standard on monitoring transactions, named as the origin only. The position for a bank in India is stated by the Reserve Bank of India instead | fatf-gafi.org |
| Reserve Bank of India | Published expectations on a regulated lender covering monitoring, outsourcing, digital lending, data and consent, and the treatment of a customer where a decision is automated. Must be read at source | rbi.org.in |
| Securities and Exchange Board of India | Equivalent expectations where the deployer of an automated monitoring arrangement is a market intermediary rather than a bank | sebi.gov.in |
| Agrawal, Gans and Goldfarb | Prediction Machines, 2018, for the split between a prediction and the deciding that follows it | Harvard Business Review Press |
Sumeru Bank Limited and Ismail Sheikh are invented.
Educational material. Not advice on any investment, tax, budget or market position.
