Fraud Detection and Transaction Monitoring: Rules and Models
Transaction monitoring watches every transaction on a book against written lines and raises an alert when one matches. Fraud detection is the whole arrangement that turns those alerts into confirmed cases. At Sumeru Bank Limited, invented, the arrangement is right 0.15 per cent of the times it speaks, and that is the design working rather than failing, because the two errors do not cost the same thing.
The whole arrangement rests on one asymmetry, and the asymmetry is worth stating before any number arrives. A case that gets missed costs this bank an amount of money it can name and put in a budget. An alert on a lawful payment costs the desk ninety seconds, and it costs somebody outside the bank an afternoon of not being able to pay for something. Because those two errors are different sizes, an arrangement that is wrong almost every single time it opens its mouth can still be the right arrangement, and the moment that is said without the second half it becomes a sentence that defends anything. So the second half has to travel with the first, in the same breath, every time the claim is made.
What is Transaction Monitoring, and what does it actually produce?
Start at the gate of a residential society. The guard has a printed sheet on a clipboard, and the sheet says things like: stop anybody carrying a television out after eight in the evening, and stop any vehicle whose number is not on the resident list. The guard is not deciding who is a thief. The guard is comparing what walks past against a sheet somebody wrote, and when something matches, a name goes in a register and a supervisor gets a call. The sheet does not know what theft is. The sheet knows what a previous theft looked like.
Transaction monitoringWatching every transaction on a book against a set of written lines, and raising an alert whenever one of them matches. is that clipboard, running over money instead of televisions. At Sumeru Bank Limited the servicing book carried 2,000,000 transactions in one month. Component 7 of the intake chain sat across all of them. Component 7 is not fitted to anything: it is 61 written linesA rule somebody stated in advance. A written line fires only on the thing it names, and the person who wrote it can read it out loud., each one a comparison somebody stated in advance. The 61 lines raised 18,000 alerts in the month, being 0.9 per cent of the transactions.
The entire output of transaction monitoring is a list, and the list is not a finding, an accusation or a decision. Nothing on it has been read. Nothing on it has been stopped. Nobody outside the bank knows any of it exists. An alert says only that a transaction matched a comparison, a much smaller claim than the word fraud makes it sound. Almost every argument about whether a monitoring arrangement is any good turns out to be an argument about what happens after the list, not about the list itself.
Fraud Detection vs Transaction Monitoring: where exactly do the two come apart?
The society gate makes the same point. The clipboard is the monitoring. But nobody in that society would say the building is protected by a clipboard. Protection is the clipboard plus the guard who reads it, plus the supervisor who is called, plus the register the entry goes into, plus the decision somebody makes about whether to stop the person at the gate or wave them through and ring the resident instead. Strip all of that away and what remains is a clipboard, and a clipboard is not protection.
Fraud detectionThe whole arrangement that turns raised alerts into confirmed cases, including every person, stage and record between the two. is the whole arrangement, and monitoring is one component inside it. At Sumeru Bank Limited the arrangement runs from those 18,000 alerts through four stages, set out under alert triage and escalation, down to 27 confirmed casesAn alert that an investigation established was fraud. Nothing else the arrangement produces would be called a finding.. Monitoring produces alerts. Detection produces confirmed cases. Alerts and confirmed cases are two different things, and no amount of good monitoring turns one into the other on its own.
The split matters commercially, and it is where a firm most often buys half an arrangement and thinks it has bought the whole one. The monitoring half can be bought. Monitoring arrives as a running service, it can be switched on in a quarter, and its cost sits in one line of a budget. The detection half arrives as salaries: at this bank it is 5 posts on a fraud desk, and those posts appear in a headcount request that a completely different person has to approve, in a completely different meeting, often in a different year. So the two halves are approved separately, and it is entirely possible to end up with excellent monitoring, 18,000 alerts a month, and nobody with time to read them. The gap is not a hypothetical failure mode. The arithmetic below makes it almost inevitable whenever the second approval does not happen.
What does transaction monitoring produce, and what does fraud detection produce?
What does a written line give the person who has to answer for it?
Component 7 is a rule set, and that was a choice. The bank could have fitted something to its own past fraud cases and let it rank every transaction by how unusual it looked. The bank did not, and the reason has nothing to do with which approach finds more fraud. The reason is what happens on a Tuesday afternoon when an account holder rings up and asks why their payment was stopped.
Two security guards stand at the same gate. The first can state the rule: nobody carries a television out after eight without a gate pass, here is the sheet, that is why the person was stopped. The rule may well be a stupid one, and it can be argued with, and if enough people argue the society can change the sheet at the next meeting. The second guard says the person looked wrong to him. He may well be a better guard, with sharper instincts and a better record. But there is nothing there to argue with, and nothing the society can change on Friday.
Agrawal, Gans and Goldfarb, in Prediction Machines, published in 2018, draw the line between the prediction a component produces and the deciding that a person still has to do afterwards. The 61 written lines buy this bank something other than accuracy: a sentence that can be read out to the person the alert was about, and the ability to change the rule inside a week when the pattern moves. A fitted component would very likely rank better. A fitted component would also leave the person on the phone with nothing to say.
An account holder rings and asks why their payment was held. Which arrangement leaves somebody able to answer?
How does a line get written, and what does that say about what it cannot see?
Of component 7's 61 lines, 44 were written after a loss the bank had already taken and 17 were written from an expectation of one. The split between the 44 and the 17 explains almost everything about how a rule set behaves. A line exists because somebody lost money, somebody looked at how it was done, and somebody wrote down the comparison that would have caught it. A rule set is therefore a written record of what has already happened to the bank. A record of the past is accurate about the past and blind in a shape that cannot be seen from inside it.
The productive lines confirm this. Alert sources 3, 4 and 5 together carry 22 of the month's 27 confirmations, being 81.5 per cent, and all three sit among the 44 that were written after a loss. The bank knows those patterns because it paid for them once each. Nothing in the rule set knows about the pattern that has not yet cost anybody anything.
There is a second consequence, and it is a comfort rather than a warning. A fraud case is confirmed or not within days, so somebody can look at last month's lines against last month's outcomes and argue about them at a meeting this month. The scoring model on the credit side cannot be looked at that way at all: under this bank's own definition an outcome is not known until twelve months of observation have run, so the earliest a month of credit decisions can be scored is fifteen months after them. The fraud lines can be argued with monthly. The speed at which an outcome arrives, more than anything about the technology, is why one part of this chain is 61 sentences and the other is a fitted component.
Why are so many alerts not fraud, and is that a fault in the lines?
One number decides everything downstream, and it is not a number about the lines at all. In the month, 2,000,000 transactions produced 27 confirmed cases. The rate is about one in every 74,074. Whatever is being looked for is that rare, and the base rateHow rare the thing being looked for actually is. The base rate sets how many alerts will not be that thing, whatever the lines say. of a thing sets how often anything looking for it will speak about something else.
A wedding at scale, the kind with 2,000 guests over three days, has one person present who is there to steal from the gift table. Instructions for the staff good enough to catch that one person now have to be written. Every instruction available, stop anybody who walks towards the gift table alone, ask about anybody nobody at the top table recognises, will stop dozens of perfectly ordinary guests for every single time it is right. Not because the instruction is badly written. Because there is one of them and 2,000 of everybody else.
So the 17,973 alerts a month that were not fraud are arithmetic before they are anything else, and rewriting the lines to speak less often moves the confirmations before it moves much else. This is the part that gets argued about wrongly in almost every review. Somebody looks at 18,000 alerts and 27 cases and concludes the lines are badly written. Sometimes they are. But a line sensitive enough to catch a one in 74,074 event will speak far more often than that event happens. The only way to make it speak much less is to make it less sensitive, and a less sensitive line takes cases off the bottom of the list first.
27 confirmed cases in 2,000,000 transactions. Why does that produce 18,000 alerts rather than a few hundred?
What does a missed case cost this bank, in rupees?
Now price the first of the two errors. Sumeru Bank Limited's own amount at riskWhat a confirmed case would have cost the bank if it had not been stopped, on the bank's own average across its own cases. on a confirmed case is Rs 84,000/-, an average across its own cases rather than any published figure. So the 27 cases in the month carry Rs 22,68,000/- between them, and on the same steady volumes 324 cases a year carry Rs 2,72,16,000/-.
| The first error, priced | Working | Amount |
|---|---|---|
| Average amount at risk on one confirmed case | the bank's own average | Rs 84,000/- |
| Confirmed cases in the month | from 18,000 alerts | 27 |
| Carried by the month's cases | 27 times Rs 84,000/- | Rs 22,68,000/- |
| Confirmed cases in a year | 27 times 12 | 324 |
| Carried by the year's cases | 324 times Rs 84,000/- | Rs 2,72,16,000/- |
Every rupee in that column is the bank's own money, and somebody in the building is accountable for it. Every bank has already written this column for exactly that reason. Notice what it assumes: that every confirmed case would have completed if nobody had stopped it, and that a stopped case saves the whole amount. A real month delivers neither cleanly. So the total prices the shape of the exposure rather than money anybody got back.
What does an alert cost the desk, and what does a hold cost somebody else?
Now price the second error. The second error lands in two different places on two different people, so pricing it takes two units rather than one.
Inside the bank it is minutes. Of the month's 18,000 alerts, 11,880 were closed by automatic suppression and nobody read them, so they cost almost nothing. The remaining 6,120 were read. Take out the 27 that turned out to be fraud and the arithmetic on this bank's own locked handling times runs like this: 5,580 first reads at 90 seconds is 8,370 minutes, the 513 kept cases that were investigated and found nothing carry another 769.5 minutes of first reading and 20,520 minutes of fuller review, and the three together are 29,659.5 minutes a month. At the assumed working month of 8,400 minutes that is 3.53 posts. So of the 4.24 posts of work the desk actually does, 3.53 of them go on alerts that were not fraud, being 83.2 per cent of every minute the desk spends. That is not a scandal. The arithmetic of the base rate makes it unavoidable.
Outside the bank it is not minutes at all. Of the 540 cases kept, Sumeru Bank Limited stopped the payment while the case was reviewed on 218 of them. Twenty seven were confirmed. HoldingStopping a payment from completing while the alert on it is being reviewed. The money does not move until somebody releases it. a payment is the only step in this whole arrangement that anybody outside the bank can feel, and it happened to 218 people in the month.
191 of those payments were released after review. 191 people were stopped from paying for something and had done nothing whatever wrong. They are not a rounding error and they are not a cost the design absorbs on their behalf. Somebody's rent went late. Somebody's supplier was not paid on the day. Nothing in the right-hand column above is denominated in a unit that can be compared with the left-hand one, and the comparison everybody makes therefore uses only the left.
218 payments were held and 191 of them were released after review. Who paid for those 191?
Where is the break-even, and what makes the trade arguable at all?
Put the two priced sides together and the argument becomes arithmetic. The year's 324 confirmed cases carry Rs 2,72,16,000/-. The desk that produces them costs Rs 45,00,000/-. The ratio is 6.05 times, and it is where the sentence about paying for itself six times over comes from.
But the far more useful number is the one underneath it. The desk divided by the year's cases gives the break-evenThe average case value at which the desk costs exactly what it prevents. Above it the arrangement pays; below it, it does not.: Rs 45,00,000/- over 324 cases is Rs 13,889/- an average case. Stating the break-even converts a question nobody can settle, is this arrangement worth it, into a question anybody can settle, is the average case bigger than Rs 13,889/-. At this bank it is six times bigger, and that is the whole defence, resting on one figure that most firms have never written down.
Suppose this bank's average amount at risk on a confirmed case fell to Rs 10,000/-. What has changed about the design?
Before the control below is touched: at roughly what average case value does this desk stop paying for itself?
Move the one number the whole defence rests on
Everything else is held where the month put it and does not move: 27 confirmed cases a month, being 324 a year, and a fraud desk of 5 posts at an assumed fully loaded Rs 9,00,000/- each, being Rs 45,00,000/- a year. Only the average amount at risk on a confirmed case moves.
At Rs 84,000/- an average case, the year's 324 confirmed cases carry Rs 2,72,16,000/- against a desk costing Rs 45,00,000/-, being 6.05 times over.
Why do both halves of that sentence have to be said together?
One sentence carries the whole case, and it only works whole. The arrangement at Sumeru Bank Limited is right 0.15 per cent of the times it speaks, and that is the design working rather than failing, because the two errors do not cost the same thing. Said with only the first half, it condemns a design that pays for itself six times over on the strength of a rate that was never the right measure. Said with only the second half, the only party left in the comparison is the bank, and the sentence will defend absolutely any amount of wrongness.
Watch how the second failure works. The second failure is the more comfortable one and therefore the more common. The claim is that the desk carries six times its own cost. Now imagine the alerts double to 36,000 while the confirmations stay at 27. Every rupee in the comparison is unchanged, so the sentence still reads six times over. Twice as many people have been looked at, and plausibly twice as many lawful payments have been held. An argument that does not get worse as the wrongness rises is not an argument about wrongness at all, and that is the precise defect in offering the six times figure alone.
Somebody defends this monitoring arrangement by saying it pays for itself six times over. What is missing from that defence?
The failure that hides behind a correct slide, and what it costs
A monitoring arrangement is reviewed. Somebody produces one slide: Rs 2,72,16,000/- carried a year against a desk costing Rs 45,00,000/-, six times over, approved. Every figure on the slide is correct and the slide is not wrong about anything it says. The slide is wrong about what it leaves out. The same month held 17,973 alerts that were not fraud and 191 people who were stopped from paying for something and had done nothing.
The cost is not that the arrangement gets approved. On these numbers it should be. The cost is that the slide sets the standard of proof for every future review, so when the alerts double and the confirmations do not, the same slide will be produced, it will still read six times over, and nothing in the room will have got worse. A measure that cannot deteriorate cannot govern anything.
What does this arrangement cost the people it is wrong about?
The 17,973 alerts that were not fraud are usually written as one number. Three completely different experiences are stacked inside that one number, so writing them as one is itself part of the problem.
Follow the third block down. 513 alerts were kept, investigated by a person for about 40 minutes each, and nothing was found. The 513 are people about whom a bank formed a question and answered it in their favour, and almost all of them will never know it happened. But 191 of those 513 also had a lawful payment stopped while the question was being answered. The 191 did nothing wrong at all, and they are not a cost the design quietly absorbs on their behalf: the cost was moved onto them, and it was moved without being counted anywhere.
O'Neil, in Weapons of Math Destruction, published in 2016, makes the general point that a system's errors rarely fall evenly across the people they land on. Sumeru Bank Limited has not measured who its 191 are, and that absence is worth stating plainly rather than filling with a guess. A payment held for two days is a small inconvenience to somebody with a balance and a serious event for somebody paying a hospital. Sumeru Bank Limited does not know which of those it did 191 times last month, and no arrangement that has not asked the question can claim to know the answer.
Where does the judgement in all of this actually sit?
Looking at 61 lines running over 2,000,000 transactions, somebody could easily conclude that the arrangement makes the decisions. The arrangement makes none of them. Every judgement was made in advance by a person, and it is worth being able to point at all three.
The monitoring compares; the people chose what was worth comparing, how much attention each stage was worth, and whether to stop somebody's money while the question was open. That is why the interesting review questions are never about the lines. Ask who set the 90 seconds and what evidence they had. Ask who decided that 218 payments a month is an acceptable number to stop. Both are answerable questions with names attached, and neither is a technical question.
Name one of the three places the judgement in this arrangement sits.
How would somebody reviewing this arrangement test it in one hour?
A risk reviewer, an internal auditor or a board member can run this test without any access to the monitoring itself. Neelima Rao, in the risk function at Sumeru Bank Limited, built no part of the chain, and building no part of it is exactly why she can ask these four things. Every one of them is answerable from records the bank already keeps.
| What to ask for | What a usable answer looks like |
|---|---|
| The break-even, written down | The desk's annual cost divided by the year's confirmed cases. At this bank, Rs 45,00,000/- over 324, being Rs 13,889/-. If nobody has ever computed it, the arrangement has never been argued about, only asserted. |
| The second column of the ledger | A count of alerts that were not fraud, split by how much attention each one consumed, and a count of payments held and released. At this bank: 11,880, 5,580, 513, and 191. |
| Which lines earn their volume | Alerts and confirmations by source, side by side, so a line that speaks constantly and is almost never right is visible as such rather than hidden inside a total. |
| Who chose the holds | A name and a date against the decision to stop a payment pending review, and the standing arrangement for releasing one quickly. This is a conduct question long before it is a fraud question. |
Notice what is not on that list. Nothing about how the monitoring works internally, nothing about statistics, and nothing that requires the reviewer to have built anything. Every one of those four questions is answerable from records the bank already keeps, and the ones a firm cannot answer are the finding.
Who states what applies here
The international standard on transaction monitoring originates with the Financial Action Task Force at fatf-gafi.org. The expectation that a regulated institution watches transactions comes from there in the first place. The Reserve Bank of India at rbi.org.in states what actually applies to a bank in India, and where the deployer is a market intermediary rather than a bank the Securities and Exchange Board of India at sebi.gov.in states it.
The 61 lines, the 90 seconds, the two suppression and review windows and the Rs 84,000/- are Sumeru Bank Limited's own choices, not a standard, a norm or a requirement of any authority. Anything resting on them should be checked against the position at source.
Sources
| Source | Document | Site |
|---|---|---|
| Financial Action Task Force | The origin of the international standard under which a regulated institution monitors transactions. Named here as the origin only. What applies to a bank in India is stated by the Reserve Bank of India rather than here | fatf-gafi.org |
| Reserve Bank of India | Published expectations on a regulated lender covering fraud monitoring, outsourcing, digital lending, data and consent, and the treatment of a customer whose transaction is stopped. What applies to a bank in India is stated here and must be read at source | rbi.org.in |
| Securities and Exchange Board of India | Equivalent expectations where the deployer of a monitoring arrangement is a market intermediary rather than a bank | sebi.gov.in |
| Agrawal, Gans and Goldfarb | Prediction Machines, 2018, for the split between the prediction a component produces and the deciding a person still has to do afterwards | Harvard Business Review Press |
| O'Neil | Weapons of Math Destruction, 2016, for a system's errors falling unevenly across the people they land on | Crown |
Sumeru Bank Limited and Neelima Rao are invented.
Educational material. Not advice on any investment, tax, budget or market position.
