False Positives and False Negatives: The Cost of Each
A yes or no system fails in two directions, and the two failures do not cost the same thing or land on the same person. Describing one takes four counts, not two. One invented bank held two of those counts exactly, held a third only as a floor it could never close, and never took the fourth at all.
Start with the smoke alarm in a kitchen. The alarm has exactly two ways of being wrong. One is to shriek at burnt toast, an annoyance that wakes the house and costs somebody two minutes on a chair with a tea towel. The other is to stay silent while something is actually burning. The silence costs something else entirely. Nobody sensible argues about which of those two is the error. Both are errors, both are borne by different people in different units, and no way of adjusting the alarm makes both of them go away.
Every count that follows comes from one month of one deployment at Sumeru Bank Limited, an invented bank.
What are the two ways a system that says yes or no can be wrong?
Sumeru Bank Limited runs fraud monitoring on its servicing book. The monitoring is component 7 of the bank's intake chain, and it is a rule set somebody wrote rather than a fitted model. The two errors below belong to any arrangement that answers yes or no, whether the answer came out of a rule a person typed or out of numbers fitted to data. In one month that monitoring raised 18,000 alerts. The bank kept 540 of them after triageThe step that decides how much looking each flagged item is worth, before anybody investigates it properly., being 3.0 per cent, and 27 were confirmed as fraud, being 5.0 per cent of the ones kept and 0.15 per cent of everything it flagged.
The first way to be wrong is to speak when there was nothing to say. An alert with nothing behind it is a false positiveThe system said yes and the answer was no. It raised a flag and there was nothing behind it.: the alert fired, somebody looked, and the transaction was exactly what it appeared to be. The second is to stay quiet when there was something to say. A silence with something behind it is a false negativeThe system said no and the answer was yes. It stayed silent and something was there.: the fraud went through, the monitoring never spoke, and nobody knew anything had happened.
One property shapes everything that follows, and it is not a property of the fraud rules. The property belongs to the world. An error where the system speaks announces itself and gets counted; an error where the system stays silent produces no event at all, and only turns into a record much later, if it ever does. A wrong alert leaves an alert line, a triage decision and sometimes an investigation file. A missed fraud leaves nothing on the day. The miss becomes visible only when the account is eventually written offAn account the bank has accepted it will not recover. It is the only record a missed fraud reliably leaves behind., and only if it is.
How many counts does it take to describe how a yes or no system is performing?
How many numbers does it take to count them, and what does each hold?
Being wrong in two ways is only half of what a system does, so two counts are not enough. A system is also right in two ways, and the two kinds of rightness are as different from each other as the two kinds of wrongness. So there are four countsThe four numbers any yes or no system produces: right when it spoke, wrong when it spoke, right when it stayed silent, wrong when it stayed silent. in total: right when it spoke, wrong when it spoke, right when it stayed silent, wrong when it stayed silent. Nothing else can happen. Any file the monitoring saw in that month sits in exactly one of the four.
Each of the four counts takes a number. Count 1 is an alert that turned out to be fraud, and at this bank that is 27. Count 2 is an alert that turned out to be nothing, and that is 17,973, made of 17,460 dropped at triage and 513 kept and then found to be nothing. Count 3 is silence that was correct, every transaction the rules let through that was exactly what it looked like. Count 4 is silence that was wrong, the frauds that got past.
The unevenness between an error that announces itself and one that leaves no event is a fact about what this firm writes down rather than about how its rules behave, and it travels into everything computed from the four. Two of those counts are already remarkable before anybody asks what the errors cost. Of the 18,000 things this system flagged, 17,973 were not fraud, or 99.85 per cent of everything it said. Read cold, that sounds like an arrangement that does not work. The 99.85 per cent is in fact the design working as intended, and the sentence is not a defence of anything until the cost of each error, and who it lands on, is known.
| Count | What it holds | This bank, one month |
|---|---|---|
| 1 | An alert that was fraud | 27 |
| 2 | An alert that was not fraud, being 17,460 dropped at triage and 513 investigated finding nothing | 17,973 |
| 3 | Silence that was correct | never counted |
| 4 | Silence that was wrong | at least 11 |
| Alerts raised, being counts 1 and 2 together | 18,000 | |
Why is the count of missed frauds only ever a floor?
Count 4 at this bank is at least 11. The 11 are accounts written off in the month that carried a pattern the fraud rules had been written to catch, and which raised no alert at all. Read the sentence again and notice where the number came from. The monitoring cannot have produced it. By definition the monitoring never saw those accounts. The count came from the write-off ledger, months downstream, and only for the accounts that went bad in a way that forced somebody to look.
So 11 is a floorA count that is at least this and cannot be shown to be exactly this, because the thing being counted does not always leave a record.. The floor counts only the misses that eventually surfaced. A fraud that got through, was never written off, and left the book quietly is in count 4 as surely as any of the 11, and no arrangement anywhere in the bank contains a line for it. The number is not uncertain in the ordinary way, where a better measurement would tighten it; it is a floor by construction, and it stays a floor no matter how hard anybody works.
Put the two known sides together and the reading is uncomfortable. Sumeru Bank Limited confirmed 27 frauds and knows of at least 11 it missed, so it knows about 38 frauds in total, and at least 28.9 per cent of the frauds it knows about were ones the monitoring never mentioned. The 28.9 per cent is not the miss rate. The figure is a floor under the miss rate, computed on the only frauds anybody at this bank can name.
A firm says it missed 11 frauds last month. What kind of number is that?
Why does nobody ever miss the count that nobody takes?
Count 3 is the times the rules correctly said nothing. On a servicing book that is almost every transaction that happened all month, and it is by a very long distance the largest number in the whole arrangement. Count 3 appears in no pack, no report and no committee paper at this bank. There is no line for it, no person accountable for it and no budget attached to it.
Ask why nobody has ever complained about that and the answer is embarrassing in its simplicity. Nothing happens in count 3. No minute is spent. No customer is inconvenienced. No money moves anywhere it should not. Count 3 produces no event, and firms count events. Every number a firm keeps about a control is a count of something going wrong. So the one count that would make the arrangement look reasonable is the one that was never taken.
The missing count matters more than it sounds. The 99.85 per cent figure from earlier looks damning precisely because it is computed inside the alerts, where all the wrongness lives, and never against the silence, where all the rightness lives. Neither reading is dishonest and both use real numbers. The two readings simply answer different questions, and a firm that holds only one of them will only ever have one of the answers.
Why does nobody notice that the largest of the four counts is never taken?
What does a wrong flag actually cost, counted in minutes?
Count 2 is the one this bank can price, so the costs start there. On its own flat planning rates, triage takes 2 minutes an alert and an investigation takes 30 minutes. Every one of the 18,000 alerts gets triaged whatever happens to it afterwards, so triage is 36,000 minutes a month. The 540 kept then get investigated, another 16,200 minutes. Together that is 52,200 minutes, being 870 hours.
Hold those 870 hours against what came out of them. Twenty seven confirmed cases. The 870 hours work out at about 1,933 minutes for each confirmed case, or roughly 32 hours. Four full working days of somebody's attention go into every fraud actually caught. The overwhelming majority of the work a fraud desk does is work on count 2, and that is true of every fraud desk that has ever existed rather than being a fault of this one.
One caution before that figure travels anywhere. The 2 minutes and the 30 minutes are Sumeru Bank Limited's own flat planning rates, and 52,200 is what it would cost to look at every alert at those rates. The 52,200 is not the desk's measured monthly workload. The measured workload is a different and smaller number, recorded where the triage arrangement itself is set out. The distinction matters because a planning figure is for comparing one arrangement with another, and a measured figure is for paying for the arrangement actually in place.
What kind of number is the 52,200 minutes?
What does that same wrong flag cost the person it lands on?
The minutes are the bank's own cost, and a bank can decide to bear them. There is a second cost inside the same count 2 that the bank does not bear at all. Of the 540 alerts kept, 513 were investigated and nothing was found. The 513 are people, in one month, who were looked into for having done nothing whatever. On 191 of them a lawful payment was stopped while somebody looked, and then released.
Sit with those 191 for a moment. The 191 are the part that gets written up as a percentage and forgotten. Somebody paid a school fee, or a hospital bill, or a supplier who does not extend credit, and the payment did not go. The customer rang the bank. The answer was that the payment was under review. The payment was released, and no apology was owed because no rule was broken by anybody. The people inside count 2 were doing nothing wrong, and the cost they bore is not recorded as an error anywhere in this bank. From the arrangement's point of view nothing went wrong at all.
The mirror image sits in count 4. There the harm is caused by an absence rather than by an action: money left an account and nothing stopped it. Neither group is a statistic. One was inconvenienced by a control doing its job imperfectly and the other was harmed by the same control not doing it at all, and a firm that describes either group only as a percentage has stopped being able to see the argument it is actually having.
What does a missed one cost, in money and in accounts?
Count 4 is borne in a different unit again. The at least 11 accounts were written off, and a write-off means the bank accepted it would not get the money back. A count of accounts is the whole of the recorded cost, and what is absent from it matters. Sumeru Bank Limited never put a rupee total in front of anybody, so none stands against those 11. The unit in which the miss was recorded is accounts, not money, and the two are not the same thing.
Notice also the direction that pulls every argument at this bank. The wrong flags are known exactly, at 17,973, priced in minutes at a planning rate, and reported every month. The misses are known only as a floor, at 11, in an entirely different unit, and arrive months later from a different part of the bank. The side of the trade that is easy to count is the side that gets managed, and it gets managed whether or not it is the side that matters more.
Cathy O'Neil, in Weapons of Math Destruction, 2016, shows a system's errors falling unevenly across the people subject to it, and the ones with the least recourse being usually the least visible in the numbers. The same shape appears at this bank in a small and undramatic way. The 191 stopped payments and the missing count 3 are both invisible for one structural reason: neither of them generated a line anybody had a reason to open.
Which of the two errors at this fraud desk is measured precisely, and which is not?
Who bears each of the four, and is it ever the same person?
The numbering is the point, so here is the whole of it, numbered. Party 1 is the bank's own people, who bear count 2 in time, at 52,200 planning minutes a month. Party 2 is the customer, who bears the same count 2 in an investigation and sometimes a stopped payment, at 513 investigated and 191 stopped and released. Party 3 is the bank itself, bearing count 4 in money, at least 11 accounts written off. Party 4 is nobody at all, bearing count 3. Count 3 costs nothing and nobody counts it.
Look at what that list does to the idea of a single cost of errors. Party 1 pays in minutes. Party 2 pays in a day of not being able to move money and in the experience of being investigated. Party 3 pays in written-off accounts. There is no exchange rate between a minute of a bank employee's time, a stopped school fee and a written-off account, so any single money figure covering all three is somebody's judgement wearing a number's clothes.
None of that is an argument against putting a number on it. Firms have to choose, and choosing requires comparison. The argument is for saying out loud whose judgement the exchange rate is. The moment the four parties are collapsed into one figure, the person who chose the weights has quietly decided which party the arrangement is going to lean on, and that decision is much harder to reopen once it is a single line in a paper.
A report shows 27 confirmed frauds against the desk's cost. How many of the four parties has it counted?
Why does one movement look trivial against one count and alarming against another?
Leave the fraud desk for a moment and take the other yes or no system in the same bank, the accept decision on the loan intake chain. In one month an upstream income field started arriving in a different format on one channel, and 117 files a month moved out of the accepted group and into the referral group. Nothing about the model changed. The data changed.
Now watch what happens to those identical 117 files when they are set against two different basesThe count a movement is expressed against, which decides whether the same movement looks large or small.. Against the 4,902 files the chain accepts in a month, 117 is 2.4 per cent, and 2.4 per cent of anything reads as noise. Against the 391 files it refers in a month, the same 117 files are 29.9 per cent, and nearly a third more work arriving at a desk reads as an event. Same files. Same month. Same 117.
The base decides which party can see the movement at all, so the choice of base is not presentation. Whoever is watching the accept count sees nothing worth a meeting. Whoever is sitting at the exception desk feels the whole of it. The reason the base belongs in an account of costs is simple. The argument about whether an arrangement is set correctly is nearly always conducted against the large base, where every movement is small. The cost is borne against the small one.
117 files move. Against which count does the movement look small, and why does that matter?
Across the whole episode 176 files moved from accepted into the referral group. How many did a person then accept anyway?
What did that movement actually cost the people inside it?
Across the six weeks the format change ran, about 12,900 files were decided and 176 of them moved from accepted into the referral group. Every one of the 176 went to a person, and 152 of them were then accepted by that person anyway, being 86.4 per cent. The remaining 24 were not accepted. Not one file was declined by the chain that would not otherwise have been declined.
So what did the error cost? Almost never a wrong outcome. The cost was time, and the applicant paid it rather than the bank. The chain answers in about four minutes when nobody touches the file, and the exception path takes two working days. On this bank's own seven hour working day, two working days is about 210 times as long as four minutes, and 152 people waited all of that to be told what the chain had already worked out. The cost of a wrong flag is usually a delay, and a delay is paid by whoever is waiting rather than by whoever caused it. So an error can be almost entirely harmless in its outcomes and still be expensive.
Can moving the setting reduce both kinds of error at once?
Every committee eventually asks that question, and the answer is no. The settingThe point at which a system changes its answer. It is a choice made after the model rather than a part of the model. is the point at which the system changes its answer, and it is a choice made after the model rather than a part of it. Moved one way, more files get flagged. Flagging more finds more of what is being looked for and also flags more of what is not. Moved the other way, both counts fall together.
Sumeru Bank Limited swept its own accept setting on the intake chain and the readings show the trade exactly. Each equal step tightens it by about 59 files: the accepted count falls 4,902, 4,843, 4,785, 4,727, 4,668, and the referred count rises 391, 450, 508, 566, 625 in perfect step. A different cut-off decides the declined count, so it sits at 688 and does not move at any setting. Every file the setting removes from one group appears in the other. So no position on that sweep reduces both kinds of error, and asking for both is asking for a different system rather than a different setting.
Can moving the setting reduce both kinds of error at once?
Where should a firm sit on that trade, and who decides?
If no position reduces both, then choosing a position is choosing who bears what. Draw it with the load on the bank's people and its customers running along one direction and the money the arrangement lets through running along the other, and every real setting is a point that has values on both at once. Sumeru Bank Limited sits at one such point: 52,200 planning minutes, 513 people investigated, 191 payments stopped and released, and at least 11 accounts written off.
Drawing the choice this way makes it a question about people rather than a question about tuning, and that is the only form in which it can honestly be argued. A firm that says its setting is correct has said that it prefers 513 investigated customers to some smaller number of missed frauds, and that preference is a policy rather than a calculation. The measured sweep of a fraud triage setting against what it finds is covered with the triage arrangement itself.
How does a lender, an analyst or a household actually read this?
Start at home. The habit is easier to build there. Everybody has met a household smoke alarm that went off at breakfast until somebody took the battery out. Taking the battery out is a real decision about the four counts, made in about four seconds, and it moved the setting all the way to one end. The household did not decide the risk was acceptable. The household decided that the annoyance was being borne by people who were present and the risk by people who were asleep.
A lender reading a fraud or credit report should ask which of the four counts is actually in front of it. If the paper carries confirmed cases and a desk cost, that is party 1 and part of count 1, and the other three parties are missing. The questions that follow are the count of investigations that found nothing, whether the missed count is a floor or a total, and what base sits under every percentage in the report. An analyst reading a firm's claim about its controls should do the same, and should treat a very high catch rate with more suspicion than a low one. A catch rate says nothing at all about what was never seen.
And whoever is being asked to approve the setting should ask the only question that survives all of this. Not is the setting correct, but which party is this setting choosing to load, and has anybody asked them. The question can be answered by a person with no statistical training whatever, and that is precisely why it is the one worth asking in the room.
The error that gets made, and what it costs
The failure at this bank was not the setting and was not the rules. The failure was the adding up. A monthly report showing 27 confirmed frauds against the desk's cost is arithmetically fine and every figure in it is correct. The report has also counted exactly one of the four parties, and it happens to be the party that sits inside the bank and turns up to the meeting. The 513 people investigated for nothing are absent. The 191 stopped payments are absent. The at least 11 written off accounts arrive months later through a different door and land in somebody else's paper. Count 3 is absent for a different reason. Nobody has ever taken it.
So every argument this bank ever had about whether its fraud monitoring was set correctly was conducted on a quarter of the evidence, and it was the quarter that makes tightening look expensive and loosening look cheap. Nobody has to be careless for this to happen: the report is accurate, the reader is diligent, and the arrangement still produces a systematically one sided argument because three of the four costs were never routed to the paper on which the decision gets made.
The reverse error is just as costly and reads as rigour. The reverse error takes 99.85 per cent of alerts being wrong as proof that the arrangement does not work, and demands a system that raises far fewer alerts and misses nothing. The demand is for a different system rather than a different setting, and answering it by moving the setting is how a fraud desk ends up with a smaller workload, a proud report and a floor under its miss count that nobody has looked at in a year.
What the reader has to confirm at source
Where a customer's lawful payment is stopped while an account is investigated, and where a bank must be able to show how it decided to do that, the expectations on a regulated lender are set by the Reserve Bank of India and published at rbi.org.in. The standing international discipline for the oversight of models and the arrangements built on them originates in supervisory material published by the Bank for International Settlements at bis.org, and what actually binds a bank in India is what the Reserve Bank of India states rather than what the international material says.
The 2 minute triage rate, the 30 minute investigation rate, the setting on the fraud rules and the accept setting on the intake chain are all Sumeru Bank Limited's own inventions and are not a standard, a norm or a requirement of any authority. The current position should be read at the source before any of it is relied on.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Published expectations on a regulated lender covering conduct towards customers, record keeping, and the oversight of arrangements that decide or interrupt customer transactions. The position that applies to a bank in India is stated in this material | rbi.org.in |
| Bank for International Settlements | International supervisory material from which the standing discipline of model oversight originates rather than the position that binds a bank in India | bis.org |
| Cathy O'Neil | Weapons of Math Destruction, 2016. The book shows a system's errors falling unevenly across the people subject to it, and the ones with least recourse being least visible in the numbers | Crown |
Sumeru Bank Limited and its intake chain are invented.
Educational material. Not advice on any investment, tax, budget or market position.
