The Fraud Alert: What Triggers It and What Happens Next
A fraud alert is a message a monitoring system raises when a transaction matches something the bank wrote down as worth looking at. At Sumeru Bank Limited, invented, 18,000 were raised in one month, 540 survived triage and 27 were confirmed. A set of 61 written lines raises them rather than anything fitted to data, and the bank still calls it a model.
An alert is not an accusation and it is not a finding. An alert is a request that somebody look, and every design question underneath it is a question about how much looking each request is worth. The question has an answer only because the two ways of being wrong cost different things. A case nobody looks at costs the bank money. An alert raised on a payment that was perfectly lawful costs the desk ninety seconds, and sometimes it costs a person a great deal more than ninety seconds. Everything below follows from those two costs being different sizes.
What does a fraud alert actually claim about a transaction?
Start at the gate of a housing society. The guard has a list of residents and an instruction that anybody not on it gets stopped and asked who they are here to see. Forty people a day get stopped. Thirty nine are couriers, plumbers and somebody's cousin. The guard is not accusing any of them of anything, and nobody in the building thinks he is. He is doing the one thing the arrangement asks of him. An unfamiliar face becomes a question, and somebody else answers it.
Transaction monitoringWatching every transaction on a book of accounts against a set of written lines, and raising a message wherever one of them matches. at a bank is that same arrangement running on payments instead of visitors. A payment goes out, the system compares it against the lines the bank has written down, and where it matches one of them a message is raised. The message asserts one thing only, that this transaction matched something written down as worth a look. The message asserts nothing whatever about the person who made it. Read it as a finding and everything downstream bends: the desk starts working from a presumption instead of a question, and the bank starts making an accusation it has no basis for.
Two of the hardest numbers at this bank depend on that distinction. At Sumeru Bank Limited the system spoke 18,000 times in one month and 27 of those turned out to be fraud. If an alert were a finding, that reads as a system wrong 17,973 times. An alert is not a finding, so what actually happened is that the arrangement asked 18,000 questions and got 27 answers worth having, and whether that is a good arrangement is a question about what the questions cost and what the answers are worth. Nothing else.
What does a fraud alert assert about the transaction it names?
What is the fraud model at this bank, and is it fitted to anything?
Sumeru Bank Limited calls whatever raises its alerts a fraud model. The word model covers two very different kinds of machinery, and which kind a bank has decides what anybody can do about it. The fraud model at this bank is 61 lines somebody wrote down, and not one of them was derived from data. A person sat with what the bank had lost money to before, wrote each line as a sentence, and put it into the monitoring that reads every transaction on the book. The word model is doing no work at all in that sentence. Lines somebody wrote can be read, argued with, and changed next week.
A shopkeeper keeps a strip of paper taped beside the till. Do not accept a torn note. Do not hand over goods against a transfer until the message arrives. Do not deliver to a new address on a first order. Somebody wrote each of those after something went wrong, any assistant can read the strip, and anybody can disagree with line three and cross it out. Another shopkeeper has stood at that counter for twenty years and simply knows when something is off. She is right more often. She also cannot hand over the strip of paper. A rule setBehaviour written down by a person as lines that can be read, rather than derived from data by fitting. is the strip of paper, and a component fitted to data is the twenty years.
There is a real advantage in the strip of paper and it is not accuracy. The advantage is that the reasoning is a document. When somebody rings up angry about a stopped payment, the desk can say which line acted and what that line compared, and the person on the phone can disagree with the line rather than with the institution. Agrawal, Gans and Goldfarb, in Prediction Machines, 2018, observe that a fitted component produces a prediction that a person still has to act on. A written line skips the prediction entirely and hands over the action. The argument about a written line is therefore much simpler, and its blindness is much larger.
What raises the alerts at this bank?
Which six lines raised the month's 18,000 alerts?
The servicing bookThe accounts a bank already has and is running, as distinct from the applications arriving at the front door. at Sumeru Bank Limited carried 2,000,000 transactions in the month, and the 61 lines raised 18,000 alerts across them, or 0.9 per cent of everything that moved. The bank groups its 61 lines into six numbered sources. A source is what anybody at the desk actually argues about, so the grouping matters more than the line count does. Every figure below is Sumeru Bank Limited's own and describes one month of one deployment.
| Source | What the line looks for | Alerts | Share | Kept | Confirmed |
|---|---|---|---|---|---|
| 1 | A payment to a beneficiary not seen before on that account | 7,020 | 39.0% | 141 | 3 |
| 2 | A transaction outside the account's usual hours | 4,140 | 23.0% | 62 | 1 |
| 3 | A change of registered handset followed by a transfer within 24 hours | 2,880 | 16.0% | 173 | 14 |
| 4 | A single transfer at or above a value the bank set for itself | 1,980 | 11.0% | 89 | 4 |
| 5 | Several transfers in quick succession summing above that same value | 1,260 | 7.0% | 61 | 4 |
| 6 | The remaining lines together | 720 | 4.0% | 14 | 1 |
| One month on the servicing book | 18,000 | 100.0% | 540 | 27 | |
The first two number columns describe the lines; the last column describes what was actually there. The two descriptions do not rank the same way, and the gap between them comes back twice below. The behaviour each line watches also matters. Not one of the six looks at the person. The lines look at a beneficiary this account has not paid before, at an hour this account does not usually transact in, at a handset that changed. Every line is a statement about how this account normally behaves, and an alert is what happens when today does not look like the last three months.
Where the expectation on transaction monitoring comes from
Watching transactions against written lines is not a habit any single bank invented. The international standard on it originates with the Financial Action Task Force at fatf-gafi.org, and what actually applies to a bank operating in India is stated by the Reserve Bank of India at rbi.org.in, with the Securities and Exchange Board of India at sebi.gov.in where the deployer is a market intermediary. Both authorities restate their position from time to time, so the current version is the one on the site.
The 61 lines, the six groupings, the 30 day suppression window and the money value that lines 4 and 5 compare against are all Sumeru Bank Limited's own choices, and none of them is a standard, a norm or anybody's requirement.
What happens to an alert after it is raised?
An alert is only worth raising if the arrangement behind it can absorb the volume. Eighteen thousand messages a month arriving at five people is 3,600 each, or about 180 a working day each, and no desk reads 180 alerts a day carefully. So the design does not try to read them all, and the four stages exist to decide how much attention each alert is worth before anybody investigates anything. The sorting is called triageSorting alerts by how much attention each one is worth, before anybody investigates any of them., and triage is where most of the engineering in a monitoring arrangement actually sits.
Stage 1 is the one worth staring at. SuppressionClosing an alert automatically because it repeats something already cleared on that same account. closes an alert automatically where it repeats a pattern the desk already cleared on that same account within the last 30 days. The bank chose that window itself. Suppression writes a record and no person ever reads it. The single stage closes 11,880 of the 18,000, being 66.0 per cent, and it is the only one of the four that nobody would describe as detection. Suppression is also the reason the desk is five people instead of seven, and the arithmetic for that follows below.
Because stage 1 closes two thirds of everything unread, it is the stage that has to be checked by hand, and Sumeru Bank Limited checked it. Four hundred of the month's 11,880 suppressed alerts, being 3.4 per cent of them, were pulled out and read by a person, and none of the 400 would have been kept by triage. Had those 400 behaved like the alerts a person actually read, where the keep rate is 540 of 6,120 or 8.8 per cent, about 35 of them would have been kept. The check is reassuring and it is small: 400 alerts from one month, and nothing in it makes the suppression safe next month.
Which triage stage removes the most alerts, and does anybody read them?
How do 18,000 alerts become 27 confirmed cases?
The desk arithmetic never appears in a description of a monitoring arrangement, and it decides whether the arrangement exists at all. Put it on the table. Every minute below is Sumeru Bank Limited's own assumption, and the working month of 8,400 minutes a person is the same assumption used across the bank.
| Stage | Cases | Minutes each | Minutes a month |
|---|---|---|---|
| 2, a first read | 6,120 | 1.5 | 9,180 |
| 3, a fuller review | 540 | 40 | 21,600 |
| 4, an investigation | 27 | 180 | 4,860 |
| The month, at 8,400 minutes a person | 4.24 posts | 35,640 |
Four point two four posts is why the fraud desk at this bank is 5 people. Now take stage 1 away and give every alert a first read: 18,000 at 1.5 minutes is 27,000, plus the same 21,600 and 4,860. The total is 53,460 minutes and 6.36 posts, so 7 people. The stage nobody would call detection saves 17,820 minutes a month, being 2.12 posts, and it saves them by closing alerts that no person ever sees. The trade sits inside every monitoring arrangement, stated in posts rather than in adjectives.
The three headline numbers measure three different things and are easy to blur. Notice what each one is measuring. The 18,000 measures the lines: it went up or down because somebody wrote a line or edited one. The 540 measures the triage: it went up or down because of how the desk sorted, and it would move if the same 18,000 arrived at a differently trained desk. The 27 is the two of them together plus whatever an investigation was able to establish. An investigation establishes less than everything that was actually fraud. A confirmed caseAn alert that an investigation established was fraud. Cases that were fraud but could not be established are not in this count. is a finding about what could be shown, and the count of things that were fraud and were never shown is in none of the bank's columns.
18,000 alerts, 540 kept, 27 confirmed. Which of those is a measure of the lines, and which is a measure of the desk?
How often is an alert right, and is 0.15 per cent a failure?
Twenty seven confirmations out of 18,000 alerts is 0.15 per cent. Say it out loud in the shape it usually gets said in: the system is right on 0.15 per cent of the times it speaks. The two errors do not cost the same thing, so a rate of 0.15 per cent is the design working rather than failing. Both halves of that sentence have to travel together. Anybody who states the first half without the second has described a system that looks broken, and anybody who states the second half without the first has described a system with no cost at all. Neither of those banks exists.
The reason a rate on its own settles nothing is that it has no units of consequence in it. A smoke alarm in a kitchen goes off at toast far more often than at fire, and nobody calculates its accuracy. Everybody already knows what the two errors cost: a minute of waving a newspaper at the ceiling against a burnt house. The alarm is deliberately tuned to speak too often, and that is not a defect in the alarm, it is the whole specification. A confirmation rate answers how often the system spoke usefully and says nothing about what it would have cost to stay quiet.
A monitoring arrangement confirms fraud on 0.15 per cent of the alerts it raises. Is it working?
Why do the two errors not cost the same thing?
Here is the arithmetic, and every rupee in it is the bank's own. Sumeru Bank Limited's average amount at risk on a confirmed case is Rs 84,000/-. Twenty seven cases in a month carry Rs 22,68,000/- between them, and 324 cases in a year carry Rs 2,72,16,000/-. The fraud desk is 5 posts at an assumed fully loaded Rs 9,00,000/- a year, or Rs 45,00,000/-. So an arrangement right on 0.15 per cent of the times it speaks pays for itself about six times over, 2,72,16,000 over 45,00,000 being 6.05. The break-even sits at Rs 13,889/- an average case. The desk would still be worth running if the average case were a sixth of the size it is.
The six times cover is one half of the answer, and it is the half a bank finds easy to say. The other half is that 17,973 of the month's 18,000 alerts were not fraud, being 99.85 per cent, and that 6,120 alerts were opened by a person of which 6,093 were not fraud, being 99.56 per cent. And on 191 of those, the payment was stopped while somebody reviewed it and then released. The 191 never travels with the six times cover and always should. Neither half cancels the other. The arrangement is worth about six times its cost and it is also wrong nearly every time it speaks, and the only way to hold both is to notice that being wrong in the two directions buys and costs entirely different things.
What does an alert cost the person on the other side of it?
Of the 540 cases triage kept, Sumeru Bank Limited held the payment pending review on 218 and released 191 of them afterwards. Twenty seven were confirmed and 191 were not, so 87.6 per cent of the holds came off. One hundred and ninety one people had a lawful payment stopped in one month, and every one of them was doing exactly what they thought they were doing. A holdStopping a payment while the alert raised on it is reviewed, so the money does not move until somebody decides. is not a letter that arrives later. A hold stops the payment now, and the person is standing at a counter or on the phone or watching a screen when it stops.
The cost has a shape. A payment to a hall for a wedding booking, made on a Saturday because that was the day the amount was arranged, to a beneficiary that account has never paid before. Source 1 fires. The payment stops. Somebody at the other end is waiting for a confirmation that does not arrive, and the person who sent it spends the afternoon on a helpline instead of at the thing they were paying for. The review clears it, correctly. There was never anything to find. The afternoon is a real cost, it was borne by somebody who did nothing wrong, and no arithmetic anywhere in the bank's figures nets it off against the Rs 2,72,16,000/-.
O'Neil, in Weapons of Math Destruction, 2016, shows that the errors of an automated arrangement do not fall evenly across the people it acts on. Nothing in these 61 lines looks at who anybody is, but the lines look at behaviour, and behaviour is not evenly distributed either. An account that pays new beneficiaries often, transacts late, or has just changed handsets meets these lines more often than an account that pays the same four people at the same hour every month. Sumeru Bank Limited has not measured who its 191 are. A cost nobody counts is a cost nobody manages, and 191 a month is 2,292 a year on the same steady volumes. The absence is worth saying out loud.
The desk costs Rs 45,00,000/- a year and the cases it finds carry Rs 2,72,16,000/-. What is missing from that comparison?
Which lines earn their volume, and which do not?
Go back to the six sources and read the alert column against the confirmation column. Source 1 raises 39.0 per cent of the month's alerts and carries 3 of the 27 confirmations, being 11.1 per cent. Source 3 raises 16.0 per cent and carries 14, being 51.9 per cent. Source 3 speaks 41.0 per cent as often as source 1 and is right about eleven times as often per alert, 0.49 per cent of its own alerts against 0.04 per cent. The loudest line is not the most productive one, and there is nothing whatever in an alert count that could have told anybody so.
Before moving the control below. Source 1 raises 39.0 per cent of the month's alerts. What share of the confirmed cases does it carry?
Pick one of the six lines, and watch the three counts stop agreeing
One control: which of the six numbered sources is in view, with an all six setting. One consequence: three tracks redraw as that source's share of the month's alerts, of the month's kept cases and of the month's confirmations, a marker slides along a scale showing how often that line is right about its own alerts, and the 27 confirmed cases light up one square at a time. The default is all six together, giving 18,000 alerts, 540 kept and 27 confirmed, or 0.15 per cent of everything raised. Source 1 gives 7,020 alerts and 3 confirmations. Source 3 gives 2,880 alerts and 14.
Across all six sources together, the 61 written lines raised 18,000 alerts in the month, triage kept 540 of them and 27 were confirmed, which is 0.15 per cent of everything raised.
What can a written line never see, however many lines there are?
Forty four of Sumeru Bank Limited's 61 lines were written after a loss the bank had already taken, and 17 were written from an expectation of one. Sources 3, 4 and 5 carry 22 of the 27 confirmations between them, being 81.5 per cent, and all three sit among the 44. A rule set is a record of what has already happened to the bank, and that is both the whole of its strength and the exact shape of what it cannot see. The lines that find the most were bought with losses the bank had already taken. A sentence like that is worth sitting with.
So what happens when a way of taking money appears that nobody has written a line for? Nothing happens. Not a low alert, not a weak signal, not a message somewhere down a list. The transaction passes through the 61 comparisons, matches none of them, and is never mentioned. The only way it gets seen is if it happens to trip a line written for something entirely different, and that is luck rather than design. The blindness is not a defect in this bank's lines; it is what a line is. Ask a guard who has a list of residents to spot somebody who is on the list and should not be, and he will wave them through every single time. Spotting that person was never the question he was given.
A new way of taking money appears that nobody has written a line for. What does the monitoring arrangement do?
What does the record have to hold on every alert?
Ninety seconds is the whole budget for a first read, and 6,120 of them a month is 9,180 minutes of somebody's working life. An alert that makes the reader go and look something up has already spent its budget, so whether ninety seconds is enough is decided entirely by what the alert arrives carrying. The record is the least glamorous part of a monitoring arrangement, and it decides whether the desk is five people or eight.
Field 6 deserves its own sentence. An alert with a hold on it is a different object from one without: somebody is waiting. The record has to say whether a payment is held. Ninety seconds of queue time on an alert with no hold is ninety seconds of desk cost, and ninety seconds on an alert with a hold is ninety seconds of somebody's afternoon. A record that does not distinguish the two makes it impossible to work the queue in an order that reflects the difference, and 191 people a month are on the wrong side of that.
How does a person reviewing this arrangement read the three numbers?
Ismail Sheikh runs the exception desk at Sumeru Bank Limited and Neelima Rao sits in its risk function, and neither of them reads 18,000, 540 and 27 as a scoreboard. The two of them read the numbers as three separate instruments pointing at three separate things, in a fixed order.
The first reading is a ratio, not a level: 540 out of 18,000 is 3.0 per cent, and it says how much of what the lines produced was worth a second look. If that share moves without anybody editing a line, the sorting has changed rather than the world. The second reading is 27 out of 540, being 5.0 per cent, and it says how much of what the desk kept turned into something an investigation could establish. The two ratios separate the lines from the desk, and a firm that only tracks the confirmation rate against all alerts has fused them into one number that cannot tell it which end moved. Rewriting lines when the sorting changed, or retraining a desk when a line changed, is the ordinary consequence of reading only the fused number.
The third reading is the one that is not in any of the usual reports: 218 payments held and 191 released. The third reading is not a performance measure and it does not go up when things go well. The third reading is a count of people, and a reviewer asks for it separately because it is the only number that describes a cost the bank does not pay. Next to the household running on one salary, a held payment on a Saturday is not an inconvenience of the same size for everybody, and nothing in a monthly summary of 18,000 alerts will tell anybody that.
The error that gets made, and what it costs
The mistake is to read 0.15 per cent as a hit rate and conclude the arrangement is broken. The mistake is made most often by somebody senior seeing a monitoring summary for the first time, and the move it produces is to tighten the lines so they speak less often. Follow that through on this bank's own figures. Fewer alerts means fewer of the 6,120 first reads and real desk time saved, and it also means fewer of the 27, at Rs 84,000/- an average confirmed case. Sources 3, 4 and 5 carry 22 of the 27 and are also among the quieter lines, so a tightening aimed at volume lands on source 1 and source 2. Source 1 and source 2 raise 62.0 per cent of the alerts between them and carry 4 of the confirmations.
The reverse mistake costs more and is made less loudly: reading the 6.05 times cover as the whole answer, and letting the lines speak more and hold more because the bank's side of the arithmetic keeps working. Every extra hold is another person waiting. Both mistakes have the same root: pricing one of the two errors and leaving the other one out.
And there is a third figure neither mistake touches. Every count above is of what an investigation could establish. Cases that were fraud and were never established appear nowhere, and no amount of care with the 18,000, the 540 and the 27 will make them appear.
Sources
| Source | Document | Site |
|---|---|---|
| Financial Action Task Force | The international standard on transaction monitoring at financial institutions, named as the origin of the expectation and not as the position in India | fatf-gafi.org |
| Reserve Bank of India | Published expectations on a regulated lender covering fraud risk management, customer service, digital lending, outsourcing and record keeping. The position for a bank in India is stated here | rbi.org.in |
| Securities and Exchange Board of India | Equivalent expectations where the deployer of a monitoring arrangement is a market intermediary | sebi.gov.in |
Sumeru Bank Limited, Ismail Sheikh and Neelima Rao are invented.
Educational material. Not advice on any investment, tax, budget or market position.
