Anomaly Detection: Finding the Unusual Without a Rule for It
Anomaly detection scores how far something sits from a pattern instead of checking it against a written rule, so it can flag what nobody thought to write down. At Sumeru Bank Limited, invented, a trial scored 2,000,000 transactions in one month and flagged the 2,000 furthest from their own account's recent behaviour. Nine were fraud. The bank did not deploy it.
Everything distinctive about this approach follows from what it does not need. Nobody has to state in advance what is being looked for. The freedom from stating it is the whole gain. Unusual and unlawful are different words and the scoring only knows the first one, so the whole cost arrives with the same freedom. A distance score can establish that something is unlike the rest and it can never establish that something is wrong, and every design decision below is a consequence of that one sentence. One month at one invented bank, where the approach ran beside 61 written lines and was measured, shows what it found, what it would have cost and why the bank left it on the shelf.
What is anomaly detection, and how is it different from a written line?
Picture a night watchman on a street he has walked for eleven years. Nobody has given him a list of what a thief looks like. He could not write that list if asked for it. Eleven years of what this street sounds like at two in the morning is what he has instead, and when a shutter rattles in a way shutters here do not rattle, he turns around. He is not identifying a crime. He is noticing a difference, and then a person still has to walk over and look.
Anomaly detectionScoring how far something sits from a pattern rather than matching it to a written rule that somebody stated in advance. is that turning around, written down as arithmetic. The arithmetic takes an item, compares it against a pattern built from items like it, and returns a number saying how far apart the two are. The number answers one question and only one: how unlike the rest is this? Nothing in how the score was built ever showed it an example of a lawful, a harmful, a deliberate or an accidental item, so the number carries no opinion about which of those the item is.
Set that beside a written lineA rule somebody stated in advance, which fires only on the thing it names and stays silent on everything else.. Most monitoring is actually made of written lines. A written line names something: a payment to a beneficiary this account has not paid before, or a change of registered handset followed by a transfer within a day. A line fires when the thing it names happens and stays silent otherwise. Somebody wrote it, somebody can read it, and somebody can be asked why it says what it says. Anomaly detection offers none of that and buys something else in exchange. A distance score does not have to be told what to look for, so it is not silent on the thing nobody has named yet.
The two are not competing versions of the same answer. A written line and a distance score answer different questions, and everything below rests on that difference. A written line answers a question about a description: does this transaction match something a person named? A distance score answers a question about a comparison: how unlike this account's own recent behaviour is this transaction? Neither answer can be converted into the other, and an arrangement that treats a high score as though a line had fired has quietly changed the question without telling anybody.
What does a distance score measure?
What is the score measuring distance from, and who decides what normal means?
A distance is meaningless until somebody says distance from what. The choice of what to compare against is the whole design, and a person makes it in advance, out of sight of everybody who later reads a flag. At Sumeru Bank Limited, invented, the trial measured every transaction against that same account's own preceding three months. Not against other customers, not against a picture of a typical account, and not against anything the bank had learned about fraud. Against that one account's own recent life.
Measuring an account against itself is a real decision with real consequences, and it is worth naming both directions. Judging an account against itself means a household living on a modest salary and a household moving ten times as much every week are each judged on their own terms, so the larger one is not permanently unusual for being larger. The same choice gives an account with three months of quiet behind it a very tight idea of usual, so almost anything new looks far away. An account that already moves money in every direction has an idea of usual so loose that very little can fall outside it. The same arrangement is therefore more sensitive on a steady account and less sensitive on a busy one, and nothing in a flag tells the person reading it which kind of account they are looking at.
The reference periodThe stretch of history a distance score treats as usual. Everything the score says is a comparison against this window and nothing else. is the second half of that choice, and it is the half people forget. Three months is the invented bank's own selection. Three months is not a standard, not a requirement and not a finding. Change it and the same transaction moves. Take a shorter window and the account's most recent change becomes the new normal, so a person whose salary rose two months ago stops looking odd. Take a longer one and that same rise stays odd for the better part of a year. Neither answer is more correct than the other; they answer different questions, and only one of them was asked.
Sumeru Bank Limited, invented, chose three months, and the honest way to record that is as a choice rather than as a fact about the world. Nothing establishes three months as right. The number sat in the design, it moved the answer, and a flag arriving on somebody's desk carried no trace of it. The period is where the meaning of the word normal was actually decided, and it was decided months earlier by people who were not looking at this transaction.
What was normal defined as in this trial?
What can a distance score find that a written line never will?
Here is the sharpest way to put it, and it is worth sitting with. The two approaches are built in opposite orders, and the order is the whole difference.
A written line is built backwards from an event. Something happened, somebody suffered a loss, somebody in the bank sat down and described the shape of what happened, and a line went into the rule set. Every line is therefore a sentence about the past, and its coverage is exactly the set of things somebody has already lived through and bothered to write down. At Sumeru Bank Limited, invented, component 7 is 61 written lines, of which 44 were written after a loss the bank had already taken and 17 were written from an expectation of one. The ratio is not an embarrassment. A rule set is nothing but a written record of what has already gone wrong.
A distance score is built forwards from behaviour. Nothing has to have happened yet. The scoring takes whatever an account did over the reference period and computes a comparison. Nobody had to recognise the trouble first, so the comparison is available on the first day a new kind of trouble appears. A written line can only see what somebody has already named, and a distance score can see something nobody has ever named. The unnamed thing is precisely and only what a rule set is structurally blind to.
Notice what that does and does not promise. The claim is not that the scoring will find the new thing. The claim is only that the way the scoring was built does not disqualify it from finding one, which is much weaker and much more honest. The written line is disqualified. The gap between the two is real whether or not anything ever walks into it.
Why is unusual not the same word as unlawful?
The difference between unusual and unlawful is the sentence the whole approach turns on, and it is easy to nod at and then forget within a paragraph, so here it is with the arithmetic attached. At Sumeru Bank Limited, invented, the trial flagged 2,000 transactions in one month. Nine of them were later established as fraud. The other 1,991, being 99.55 per cent, were not.
The 1,991 are people. Somebody paid a hospital. Somebody paid a caterer three weeks before a wedding and then paid a tailor and then paid a hall. Somebody moved cities and put down a deposit on a room, and then bought a bed and a cylinder and a fan in four days. Somebody's mother needed a bus ticket home at two in the morning. Every single one of those months is unusual against its own last three months, and not one of them is any of the bank's business.
The arithmetic invites the error rather than the other way round. When an arrangement puts a small number of items in front of a reader and calls them the furthest from the pattern, the natural thing to feel is that they are the most suspicious. A short list usually does mean that. The trial's short list does not. The list means these were the least like the rest, and on any month of any book the least like the rest is dominated by ordinary people whose month was unlike their last three. Reading unusual as suspicious is not a lapse of care; it is what the shape of the output suggests, and a design that does not push back on it will produce that reading every time.
1,991 of the 2,000 flagged transactions were not fraud. What were they?
What did one bank actually run, and on how many transactions?
Sumeru Bank Limited, invented, ran the approach for one month on its servicing bookThe accounts a bank is already running, as against the applications coming in the front door. Monitoring happens here, after a loan exists., beside the written lines it already had. Every figure in this section is that bank's own, describes one month of one deployment, and is an illustration rather than a measurement of anything outside it.
The book carried 2,000,000 transactions in the month. Every one of them was scored by how far it sat from that account's own preceding three months. The scores were then ranked, and the furthest 2,000 were taken, being the top 0.1 per cent. The top 0.1 per cent is the flag shareThe proportion of the items scored that get looked at. It is chosen by whoever runs the arrangement, not produced by the scoring., and it is worth stopping on. A reader meeting the flag share for the first time often assumes it came out of the arithmetic. It did not. The scoring produces an ordering of 2,000,000 items and nothing else. Somebody then has to say how far down that ordering anybody is going to read, and that somebody is choosing a budget.
Neither the scale nor the month makes this a trialA run set up beside the live arrangement, with its outputs recorded and deliberately not acted on, so it can be measured without touching a customer. rather than a deployment. Nothing was acted on, and that is the whole difference. Each of the 2,000 flags was written down with its score, with whether a written line had also fired on the same transaction, and later with whatever an investigation established. No payment was held because of a flag. No customer was contacted because of a flag. Nobody on the desk had a queue lengthened by a flag. Recording without acting is the only arrangement under which a bank can find out what an approach does before deciding whether to let it do anything, and it is the reason these numbers can be read as a measurement rather than as a defence of a decision already taken.
Not acting also fixes what the trial can and cannot answer. The trial can say how many flags there were, how many overlapped the lines and how many turned out to be fraud. Nobody worked a flag on the day, so the trial cannot say what would have happened if somebody had.
Where the rules for monitoring transactions are actually stated
Monitoring transactions on a customer book is not an optional exercise a bank invents for itself, and the international standard behind it originates with the Financial Action Task Force at fatf-gafi.org. The Reserve Bank of India at rbi.org.in states what applies to a bank in India, and the Securities and Exchange Board of India at sebi.gov.in states what applies where the deployer is a market intermediary rather than a bank. The three month reference period, the 0.1 per cent flag share and the 61 written lines all belong to the invented bank.
Which of the 2,000 flags had already raised a written line, and which had not?
The overlap decides whether an approach is worth anything on top of what a bank already has, and the overlap is the number most often skipped. A flag count looks like a finding on its own and is not one. A flag on a transaction the existing lines already raised has told the bank nothing it did not know ninety seconds earlier.
Of the trial's 2,000 flags, 612 had also raised one of the 61 written lines, being 30.6 per cent of the flags. The same 612 are 3.4 per cent of the month's 18,000 alerts. The other 1,388 had raised no line at all. And of the month's 27 confirmed cases, 9 sat inside the 2,000, with every one of the nine sitting inside the 612. The 612 are the overlapThe flags a written line had already raised, so nothing about them is new work for anybody. and the 1,388 are the new work, and confusing the two is how a trial gets reported as a success it did not have.
The trial flagged 2,000 transactions and 612 had also raised a written line. Which group is the new work?
What did the 1,388 with no line behind them turn out to be?
Not one of the 1,388 turned out to be fraud in the trial month. The sentence is the whole result and it is very easy to over-read, so it needs saying slowly, in both directions.
Read one way it is the strongest thing the trial produced. The part of the output that was genuinely new, the part that would have cost the bank something, contained nothing the bank needed. All nine of the confirmed casesAn alert or a flag that an investigation established was actually fraud, as against one that was merely worth looking at. the trial found were in the overlap. The written lines had already raised every one of the nine. On this month's evidence the approach added no confirmations at all.
Read the other way it is one month with 27 confirmed cases in it, and 27 is a small number to draw a conclusion from. The trial cannot tell the difference between an approach that finds nothing and an approach that would have found something in a month when somebody tried something new. Both look exactly like this. A negative result on a month with 27 events is evidence about a month, and treating it as evidence about an approach is a bigger claim than the arithmetic will carry.
Recording both readings, rather than pretending the first was the only one, is the part of the trial worth copying. The finding was written as a statement about cost against yield in one month on one book, and not as a verdict.
What would it have cost to look at everything the scoring flagged?
Now the arithmetic that decided the matter, and it has nothing to do with how good the scoring is. Working a flag properly at this bank means the fuller review the fraud desk already runs on a kept case, and that review costs 40 minutes. Applying that rate to the 1,388 flags with no written line behind them gives 55,520 minutes a month. At the bank's assumed working month of 8,400 minutes a person, that is 6.61 posts.
The fraud desk at Sumeru Bank Limited is 5 people. The extra looking the trial would have created is 6.61 posts against a whole desk of 5, so deploying it at this flag share meant more than doubling the fraud function to work a group of flags that produced nothing that month. Put in the desk's own units, the 55,520 minutes are 1.56 times the 35,640 minutes the desk spends on everything it currently does, and 8.7 times the 6,360 minutes of spare capacity it carries. At the bank's assumed fully loaded Rs 9,00,000/- a post, 6.61 posts is about Rs 59,49,000/- a year.
| What is being worked | How many | At what rate | Minutes a month | Posts |
|---|---|---|---|---|
| Trial flags with no written line behind them | 1,388 | 40 min | 55,520 | 6.61 |
| Everything the fraud desk already does across its four stages | 18,000 | mixed | 35,640 | 4.24, staffed at 5 |
| Spare capacity the desk carries at 5 posts | 6,360 | 0.76 | ||
| The new work as a multiple of the desk that would absorb it | 1.56 times | 1.32 times | ||
Notice which half of the arrangement is expensive. The scoring itself is close to free: it runs on transactions the bank already stores and produces an ordering nobody has to read. A person looking is what costs money, and a person looking costs the same 40 minutes whether the flag came from a written line or from a distance. In every detection arrangement the looking is the expensive half, and it is almost always the half nobody costs before the trial starts.
Working the 1,388 flags with no line behind them takes 40 minutes each. How does that compare with the fraud desk?
What happens to the arithmetic if the share is widened instead?
The obvious objection to a disappointing trial is that the flag share was too tight. Looking at more of the ordering would surely find more. The bank swept it, and the sweep is the most useful thing the trial produced. A sweep settles the shape of the trade rather than one point on it.
Every row below is the invented bank's own reading from the same month and the same ordering, and the sweep runs from a flag share of 0.02 per cent to 1.00 per cent. Read the second and third columns together and the whole argument is there. Across the sweep the flags grow fifty times, from 400 to 20,000, and the confirmed cases inside them grow seven times, from 3 to 21, and at no setting does the count reach the month's 27.
| Flag share | Flags | Confirmed cases inside | Worked per confirmed case | Desk minutes | Posts |
|---|---|---|---|---|---|
| 0.02 per cent | 400 | 3 | 133 | 16,000 | 1.90 |
| 0.05 per cent | 1,000 | 6 | 167 | 40,000 | 4.76 |
| 0.10 per cent, the trial as run | 2,000 | 9 | 222 | 80,000 | 9.52 |
| 0.25 per cent | 5,000 | 14 | 357 | 200,000 | 23.81 |
| 0.50 per cent | 10,000 | 18 | 556 | 400,000 | 47.62 |
| 1.00 per cent | 20,000 | 21 | 952 | 800,000 | 95.24 |
The fourth column is the one to hold on to, because it puts both sides on the same scale. At the tightest setting the bank works 133 transactions for each confirmed case it finds. At the widest it works 952. The price of a confirmed case rises about seven times across the sweep while the case itself does not get any more valuable.
Read it as steps rather than as levels and it is starker still. Going from 400 flags to 1,000 buys 3 more confirmed cases, so each one costs 200 extra flags. Going from 10,000 to 20,000 buys 3 more confirmed cases as well, and each one costs 3,333 extra flags. The last step of the sweep is about seventeen times the price of the first, so there is no single sentence about what widening the share costs, only a sentence about widening it from where to where. Ten thousand extra flags at 40 minutes each is 400,000 minutes, being 47.62 posts, for those last 3 confirmed cases.
And the top of the sweep is worth one more look. At 1.00 per cent the arrangement produces 20,000 flags in a month, against the 18,000 alerts the 61 written lines produced across the whole book. Costed at the desk rate that is 95.24 posts, or about Rs 8,57,16,000/- a year at the bank's assumed Rs 9,00,000/- a post, against the Rs 2,72,16,000/- a year of amount at risk that the bank's 324 confirmed cases carry between them at its own average of Rs 84,000/- a case. Looking at one per cent of everything would cost this invented bank roughly three times its whole annual exposure to confirmed fraud. At that point the detection question has become an arithmetic question and stopped being a technical one.
Before the control is moved: widening the flag share from 0.10 to 1.00 per cent multiplies the flags by ten. What happens to the confirmed cases found?
Move the flag share and watch the squares stop filling
One control: the share of the month's 2,000,000 transactions flagged as furthest from their own account's pattern, across the six settings the invented bank swept. Three consequences redraw together. The first bar is how many transactions somebody would have to look at, against a scale to 20,000 with the month's 18,000 rule alerts marked for comparison. The middle row is the month's 27 confirmed cases, one square each, filling in as the flags widen. The third bar is the posts the looking would need, against the dashed line marking the whole fraud desk of 5.
Educational illustration. Figures are the invented bank's own and describe one deployment. One bank, one month, 2,000,000 transactions and 27 confirmed cases in total. The readings at shares other than 0.10 per cent are that bank's own estimates from the same trial rather than separate months. A confirmed case is one an investigation established. The count of transactions worked assumes every flag is looked at once at the desk's own 40 minutes a case, so it includes the flags a written line had already raised; the bank's own stated reason for not deploying costed only the 1,388 that had not.
Why did the bank not deploy it, and did the bank say why?
Sumeru Bank Limited, invented, did not deploy the arrangement, and the stated reason was the cost of the looking. Working the 1,388 flags that had raised no written line, at the desk's own 40 minutes a case, is 55,520 minutes a month and 6.61 posts, against a whole fraud desk of 5. The cost of the looking is the reason as recorded, and it is worth noticing what is not in it.
The reason says nothing about the scoring being poor. The reason says nothing about distance from a pattern being the wrong idea. The approach did find nothing that month, and even that is not the recorded reason. The decision was taken on the arithmetic of what a person would have had to do with the output. The output is not what costs money, so that arithmetic is the only ground on which a deployment decision can honestly be taken.
Recording the reason that precisely is itself the practice worth taking away. Nothing in a decision written as the approach did not work says what would have to change, so a decision written that way cannot be revisited. A decision written as at 40 minutes a flag and 1,388 flags a month this costs more than the desk it sits beside can be revisited the day either of those two numbers moves, and one of them is a design choice rather than a fact.
The trial found nothing the written lines had not raised. Does that show the approach does not work?
The failure: reading a cost result as a capability result
The trial found nothing the written lines had not also raised, at every flag share the bank tried. The tempting reading, and the one that gets written into a slide within a week, is that the approach does not work. The tempting reading goes past the evidence in two directions at once, and both matter.
One direction is the evidence about the approach. One month at one bank with 27 confirmed cases in it cannot separate an approach that finds nothing from one that would have found something in a month when somebody tried something the bank had never seen. And 44 of the 61 written lines were written after a loss Sumeru Bank Limited had already taken. A loss nobody had named yet is exactly the situation an approach that needs nothing named exists for. The month simply did not contain a new kind of trouble, and that is a fact about the month.
The reading also goes past the evidence in the other direction, and that second mistake has the bigger bill attached. Concluding that the written lines are therefore sufficient reads a month with no new kind of trouble in it as proof that no new kind will come. The honest finding is narrow and it is about cost: at this bank, at this flag share, in this month, the extra looking cost more than a whole desk and returned nothing, and every one of those four qualifiers is doing work.
Who makes the error and what it costs: the first version is made by whoever presents the trial, and it closes off a rebuild that might have been cheap. The second is made by whoever reads the presentation, and it costs whatever the first unnamed thing costs when it arrives.
What would have to be true for this to be worth deploying here?
Everything above points at one answer, and it is not the answer people expect. Nothing about the scoring has to improve. A sharper ordering narrows the flags and helps, but it does not touch what made the arithmetic fail. The change has to come in what happens to a flag: either the flag has to cost far less than 40 minutes to work, or it has to arrive carrying a reason a person can act on beyond the distance itself.
The 40 minutes is not a law of nature. The figure is what a fuller review costs at this desk, and what makes it cost that is reconstruction. The person opening a flag has a transaction, an account and a number saying far. Everything else, what this account usually looks like, what this transaction did that was different, whether the money has already left, has to be assembled by hand. A flag that arrived with the comparison already drawn would take a fraction of the time, and at a fraction of the time the whole sweep above changes shape.
There is a second reading of the trial hiding inside the same numbers, and it points at a different deployment altogether. Take the 612 flags that overlapped. The 612 are 3.4 per cent of the month's 18,000 alerts, and they held 9 of the month's 27 confirmed cases, being a third of them. An alert in that group was confirmed 1.47 per cent of the time against 0.15 per cent across the whole month, or about ten times as often. Used as a way of raising new alerts the scoring cost more than a desk and returned nothing, and used as a way of ordering alerts the written lines had already raised it concentrated the month's confirmations by roughly ten times at no extra looking whatsoever.
A claim like that is easy to over-sell, so two cautions belong in the same breath. The first is that the 612 are not a random slice of the alerts; they are alerts that were unusual for their account as well as matching a line, so the comparison mixes selection with ranking and cannot separate the two. The second is that it is still one month with 27 events in it. The reading supports a different trial, not a deployment, and the difference between those two sentences is the whole of the argument above.
What would have to change for this approach to be worth deploying at this bank?
How would somebody assessing a trial like this read it in an hour?
The arithmetic becomes useful to somebody who is not building anything at all. A risk reviewer, an internal auditor, a lender's credit committee or an analyst reading a claim about a detection arrangement all face the same problem: a trial arrives as a headline, and the headline is almost always the flag count and the case count. Neither of those is the finding. Six questions get to the finding, and none of them requires knowing how the scoring works.
One, what is the reference period, and who chose it? Everything the arrangement says is a comparison against that window. If nobody in the room can name it, nobody in the room knows what normal means here. At Sumeru Bank Limited, invented, it was that account's own preceding three months, chosen by the bank.
Two, what is the flag share, and is it a budget or a finding? The scoring produces an ordering. Somebody chose how far down it anybody reads, and that choice is where the volume comes from. At this bank it was 0.1 per cent of 2,000,000 transactions, being 2,000 flags.
Three, how much of the output overlaps what already exists? The overlap separates a trial with a result from a trial with a number. 612 of the 2,000 overlapped, so 1,388 were new, and all of the cost of the arrangement sits in the 1,388.
Four, what does one flag cost to work, and who works it? The output is close to free and the looking is not. 40 minutes a flag at this desk is what turned a trial into a decision.
Five, what was the yield in the part that did not overlap? At this bank it was nothing in one month, which is a real result and a narrow one.
Six, how many events did the month contain? Twenty seven. A month with twenty seven events cannot settle much, and a presentation that does not say the number is hoping nobody asks.
Two named ideas sit behind all six, and both are worth naming. Agrawal, Gans and Goldfarb, in Prediction Machines, 2018, separate the prediction a component produces from the deciding a person still has to do, and the whole argument above is that separation with an invoice attached: the prediction got cheaper and the deciding did not. O'Neil, in Weapons of Math Destruction, 2016, is the reason the 1,991 are accounted for at all. Errors do not fall evenly, and the people who carry them are rarely the people reading the report. Neither idea is a caution to add at the end; each is a question to ask before the flag count is read out.
A household meets this idea outside a bank, and the same arithmetic explains something familiar there. A payment stopped for looking unusual has not been judged. The payment has been noticed by an arrangement that compares this month against recent months and cannot tell a hospital from anything else. Knowing that the comparison is against the household's own recent behaviour, and not against any judgement of the household, is the difference between an unpleasant afternoon and an insulting one.
Sources
| Source | Document | Site |
|---|---|---|
| Financial Action Task Force | Origin of the international standard on monitoring transactions. The standard binding a bank in India is the Reserve Bank of India's, not this one | fatf-gafi.org |
| Reserve Bank of India | Published expectations on a regulated lender covering monitoring, outsourcing, data and consent, and the treatment of a customer where a decision or a hold is automated | rbi.org.in |
| Securities and Exchange Board of India | Equivalent expectations where the deployer of an automated monitoring arrangement is a market intermediary rather than a bank | sebi.gov.in |
| Agrawal, Gans and Goldfarb | Prediction Machines, 2018, for the split between the prediction a component produces and the deciding a person still has to do with it | Harvard Business Review Press |
| Cathy O'Neil | Weapons of Math Destruction, 2016, for errors falling unevenly across the people they land on | Crown |
Sumeru Bank Limited is invented.
Educational material. Not advice on any investment, tax, budget or market position.
