Classification Metrics: Which Question Each One Answers
A measure of a system that says yes or no is a question with counts behind it, so the records a firm keeps settle what it can ask. At one invented bank two of the four counts were exact, one was a floor and one was never taken, so one question could be answered outright, one only as a ceiling and one not at all.
Every measure of a yes or no system is built from the same four counts: the times it spoke and was right, the times it spoke and was wrong, the times it stayed silent and was right, and the times it stayed silent and was wrong. Nothing else goes in. The choice of measure therefore sits downstream of the record keeping and never upstream of it. Where a count is missing the measure is not approximate but unavailable, and reporting it anyway quietly puts an assumption where a number should be.
What is a measure of a classifier actually made of?
A housing society gate on a busy evening is the place to start. One watchman stops some visitors and waves the rest through. Somebody upstairs wants to know whether he is any good at it. Four counts come to hand immediately: the visitors he stopped who really were up to something, the visitors he stopped who were simply late for dinner at a flat on the third floor, the visitors he waved through who were exactly who they said they were, and the visitors he waved through who should have been stopped. Four counts, and every sensible question about that watchman is made out of some pair of them.
A measure of a classifierA question about how a yes or no system performed, with counts behind it. It is never a property the system carries around on its own. works the same way and has no other ingredients. A measure is a question, and behind the question sit counts of what actually happened. The question and the counts are the whole construction. A measure is not a property the system carries around with it. A measure is an answer to one question, worked out from records somebody either kept or did not keep. The distinction matters more than it sounds. A firm can hold a perfectly good system and still be unable to say a single true thing about how it performed, purely because the fourth count was never written down.
Now the trap opens. Asked how good the watchman is, the person upstairs will offer a number. Asked which of the four counts stands behind that number, the conversation changes completely. At most gates in the world nobody has ever counted the visitors who were waved through and were fine, and nobody has ever counted the ones who were waved through and were not. The counts that exist are the ones somebody had a reason to write down, and nobody has ever had a reason to write down a quiet evening.
What is a measure of a classifier built from?
Which questions can a system that says yes or no be asked?
There are three, and they are worth holding in the plainest possible words before anybody attaches a name to them. The first: of the times the system spoke, how often was it right? The second: of the things that were there to be caught, how many did it catch? The third: across everything it decided about, how often was it right overall?
The three questions are not variations on one theme. Each looks at a different slice of what happened, and one system can look magnificent on one of them and dreadful on another with nothing about it changing in between. The first question is about the noise a system makes, so it asks what it costs to listen. The second is about coverage, so it asks what got past. The third tries to swallow everything at once, and for exactly that reason it is the one that goes wrong most quietly.
Each of the three has a common name. The name is the part people repeat and the question is the part they need, so the plain question matters more than the name. The first is usually called precision, the second is usually called recall or sensitivity, and the third is usually called accuracy. Naming runs differently from one write-up to the next, and how each one is worked out from its counts is a separate subject. Which question each one puts to a firm's records is what settles whether it can be asked at all.
A monitoring pack quotes one number for a yes or no system and calls it the score. What is the first thing worth asking about it?
Which counts does each question need?
Set the three questions against the four counts and the picture resolves in one look. The question about what the system flagged needs the two counts of the times it spoke, and nothing else. The question about what the system found needs the times it spoke and was right together with the times it stayed silent and was wrong. The question about how often it was right overall needs all four, including the one that hardly anybody anywhere counts.
Sumeru Bank Limited, invented, runs fraud monitoring on its servicing book, and one steady month gives the four counts drawn above. The times it spoke and was right: 27 confirmed cases. The times it spoke and was wrong: 17,973, being the 17,460 alerts discarded at triage plus the 513 that were kept, investigated and came to nothing. The two spoken counts together make up all 18,000 alerts the month produced. The times it stayed silent and was wrong: at least 11, being accounts written off in the month that carried a pattern the fraud rules were written to catch and that raised no alert at all. The times it stayed silent and was right: never counted, not once, not by anybody.
| The count | This month | What kind of number it is |
|---|---|---|
| It spoke and was right | 27 | Exact. Each one was investigated and confirmed |
| It spoke and was wrong | 17,973 | Exact. 17,460 discarded at triage plus 513 investigated and not confirmed |
| It stayed silent and was wrong | at least 11 | A floor. Only the misses that later surfaced can appear in it |
| It stayed silent and was right | no count | Never taken. Nobody has a reason to record a quiet evening |
| Total alerts the month produced | 18,000 | The two spoken counts, and only those two, add to this |
Read the third column of that table rather than the second. Two of these are numbers, one is a lower limit wearing the clothes of a number, and one is a blank. Any measure that reaches into the third row inherits a lower limit, and any measure that reaches into the fourth inherits a blank, and no amount of care in the arithmetic afterwards repairs either.
Which question could this bank answer exactly?
One of the three, and only one. The question about what the system flagged asks how often it was right on the occasions it spoke, and both counts it stands on are exact numbers that somebody wrote down deliberately. Sumeru Bank Limited knows it spoke 18,000 times in the month, and it knows 27 of those occasions turned into a confirmed caseA flagged item that turned out on investigation to be the thing the system was looking for.. There is no estimate anywhere in that pair. The bank records the reading as 0.15 per cent of everything it said, and among the 540 alerts triage kept for a proper look, 27 came back confirmed, being 5.0 per cent of those kept.
An exact answer is not the same as a comfortable one, and 0.15 per cent is the sound of a system built so that the cheap error happens constantly and the expensive one does not. What each of the two errors costs and who ends up bearing it is set out under false positives and false negatives. The short version is that the wrong speakings are paid for in two currencies at once. The bank's own people pay in time. Customers pay in inconvenience: of the 540 alerts investigated, a payment was held on 218 and released on 191 after review, and 513 of the 540 investigations found nothing at all.
Hold on to what made this question answerable. The exactness has exactly one source. Both counts it needs sit on the spoken side of the ledger, and a firm that speaks has to record what it said in order to act on it. The counts a firm holds exactly are the counts its own work forced it to create. Every measure built only from what a system said is therefore available, and every measure built from what it did not say is in trouble.
Which question could it answer only as a ceiling, and why?
The question about what the system found asks what got past, and it is the one everybody actually wants answered. Answering that question needs two counts. Sumeru Bank Limited has the times the system spoke and was right at 27, and it puts the times the system stayed silent and was wrong at 11. Each of the 11 is an account written off during the month carrying a pattern the fraud rules were written to catch, on which no alert ever fired. So there were at least 38 frauds in the month, of which 27 were caught, and the reading comes out at 71.1 per cent.
Everything turns on that 11, so look hard at it. The 11 is not a count of the misses. The 11 counts the misses that surfaced, and the only way one surfaced was by becoming bad enough to be written off inside the same month. A ceilingA reading that cannot be exceeded, produced when the base of a measure is itself a floor rather than a total. is what follows from that. A fraud that was quietly successful, or one that will not surface for another year, leaves no trace in the 11 and no trace in the 38. Every fraud missing from that count would enlarge the base of known frauds and push the share caught downwards, so 71.1 per cent is not an estimate that might be too high or too low: it is a limit the true figure sits at or below.
Reading a floor as a floor is the most useful habit in the whole subject, and it costs nothing to acquire. When a count in a measure is a floor, the value of the reading is lost and the direction of it is kept. Saying that the share caught is at most 71.1 per cent is honest and usable. Saying the share caught is 71.1 per cent is neither, and the two sentences differ by two words.
The count of missed cases is a floor. What does that make any share built on it?
Which question could the bank not answer at all, and why did nobody notice?
The question about how often the system was right overall is the one a board is most likely to ask for and the one Sumeru Bank Limited cannot answer at any level of care. The overall question needs all four counts, and the fourth is the times the system stayed silent and was right. A non-event on an account that nobody looked at is not the sort of thing an operation generates a record of. Nobody at the bank has ever counted those, and nobody at any comparable bank has either. There is no file, no queue, no timestamp and no name attached to it.
An unavailable measureA question whose counts the firm does not hold, which is a different thing from a question with a poor or wide answer. is not a measure with a wide margin around it. The firm is simply unable to put the question. The moment a report quotes an overall share of times the system was right, somebody has quietly supplied the missing count, and the usual way of supplying it is to treat every account that raised no alert as a correct silence. Treating every quiet account as a correct silence is not a measurement. The treatment is a choice, made silently, and the choice hands the system credit for every second of every day on which it did nothing at all.
Why does nobody notice? Because of the one property this count has that the other three do not: nobody bears its cost. The wrong speakings cost the bank time and cost customers inconvenience, so somebody complains and somebody counts. The misses cost the bank money on written-off accounts, so somebody eventually writes them down. Correct silences cost nobody anything. No process anywhere in the bank therefore produces a record of them, and a reading that leans on them arrives looking as ordinary as any other line in a pack.
A report gives an overall share of times the system was right. What must it have assumed?
What does this tool do with the counts it is given?
The tool below is not a calculator and it produces no reading at all. The tool takes which of the four counts an arrangement holds, and for each one whether it is an exact number or a floor, and returns which of the three questions may be asked. The sequence a firm actually has to work through runs in the opposite direction to the one most people expect: not picking a measure and then hunting for the counts, but taking the counts genuinely held and reading off what they will support.
The four counts, set to what an arrangement actually holds
Four inputs, one consequence: the three verdicts redraw as the holdings change. The defaults are the settings of Sumeru Bank Limited: the two spoken counts exact, the count of correct silences never taken and the count of misses a floor. The three holdings together make the first question answerable, the second a ceiling and the third unavailable.
How do the same counts read at six settings of the same system?
Nothing about the system has to change for the counts to move. Triage ranks the month's 18,000 alerts and passes the top slice on for a proper look, and the size of that slice is a dial somebody set. The keep countHow many of the flagged items a triage step passes on for a proper look, out of everything the system flagged. at Sumeru Bank Limited is 540, being 3.0 per cent of the alerts, and a reading taken at one keep count belongs to that setting rather than to the system in the abstract.
Two costings of the same desk exist and they are not the same thing, so the reconciliation has to be stated before the second number appears. The 2 minutes an alert and the 30 minutes an investigation used below are the bank's own flat planning rateA flat per-item time a firm uses to compare one setting against another, distinct from what the work actually took. figures, and they serve one purpose only: comparing one setting of the keep count against another. The fraud desk's measured monthly minutes are counted differently and come out at a different total. The planning rates are not the desk's measured cost, and the measured costing is a separate subject. Every minute in the sweep is a planning minute.
| What is counted | How many | The bank's own planning rate | Planning minutes |
|---|---|---|---|
| Every alert is triaged, whatever happens to it afterwards | 18,000 | 2 minutes | 36,000 |
| Only a kept alert is investigated | 540 | 30 minutes | 16,200 |
| The month at the bank's own setting | 52,200 |
52,200 planning minutes is 870 hours, and it stands behind 27 confirmed cases. Now move the dial and watch what happens to both sides at once. The bank has swept its own ranking across six settings, and it is honest about where those numbers come from: the confirmed counts at the settings it does not run are its own estimate from ranking the same month's alerts, not six months of actually running them.
| Alerts kept | Cases confirmed | Planning minutes | Minutes for each confirmed case |
|---|---|---|---|
| 180 | 16 | 41,400 | 2,588 |
| 360 | 23 | 46,800 | 2,035 |
| 540, the deployed setting | 27 | 52,200 | 1,933 |
| 900 | 31 | 63,000 | 2,032 |
| 1,800 | 35 | 90,000 | 2,571 |
| 18,000, everything | 38 | 576,000 | 15,158 |
Read the two middle columns as shapes rather than as numbers. The confirmed cases climb steeply and then flatten while the planning minutes climb without any flattening at all, and the whole of the decision sits in the gap between those two shapes. Keeping 180 alerts confirms 16 cases. Keeping all 18,000 is a hundredfold rise in looking, and it confirms 38. The ranking is doing most of the work before a single person opens a file.
Before the control is moved: the bank keeps 540 alerts and confirms 27. If it kept all 18,000, how many would it confirm?
Move the keep count and watch the two bars come apart
One control: how many of the month's 18,000 alerts triage passes on for investigation, at the six settings the bank swept. One consequence: two bars redrawing side by side, the confirmed cases and the planning minutes, each carrying a dashed line at the deployed reading. Any move then shows both of its sides at once, what it buys and what it costs. The default is the bank's actual setting of 540 kept, giving 27 confirmed cases at 52,200 planning minutes, being 1,933 minutes for each confirmed case. Move one notch up and 900 kept gives 31 confirmed at 63,000 minutes, being 2,032 minutes each, and the 4 extra cases that move buys cost 2,700 minutes apiece.
Where is the cheapest point for each confirmed case, and why is it a trap?
Somebody now takes the sweep and does the obvious thing with it. The obvious move is to ask what each confirmed case cost at every setting and pick the lowest. The average readingThe cost divided across the cases found, taken over the whole arrangement rather than over the move being considered. runs 2,588 minutes for each confirmed case at 180 kept, 2,035 at 360, 1,933 at 540, 2,032 at 900, 2,571 at 1,800 and 15,158 if everything is kept. The lowest of those six is 1,933, and 1,933 is the setting Sumeru Bank Limited already runs.
Sit with how that lands in a meeting. A reading has been produced, it is arithmetically correct, it was not chosen to flatter anybody, and it says the current setting is the best of the six. The trouble is that the cost for each case found counts only the cases found, so it prices every case missed at nothing, and a measure that prices misses at nothing will always prefer looking less. The trough is not evidence about the setting. The trough is a property of the measure. The bank tuned the setting until the work felt worth it, and the measure rewards exactly that tuning, so the trough would appear at whatever setting the bank happened to be running.
The reading that was chosen instead of the arrangement
A monitoring pack at Sumeru Bank Limited carries one line: at the current setting each confirmed case costs 1,933 planning minutes, the lowest figure on the bank's own six point sweep. Nobody has fabricated anything. The arithmetic is right, the sweep is the bank's own, and the conclusion drawn from it is that the setting is correct and needs no discussion this quarter.
Neither of two facts appears anywhere in a cost for each case found: the same sweep shows 4 more confirmed cases available at the next setting up, and the 11 known misses are a floor rather than a total. The line cannot say either. The people who bear the consequence of that line are the customers behind the misses and the fraud desk. The customers are not counted in it. The desk now has a number that closes the conversation. The reading did not lie. The reading answered a question nobody in the room had asked.
The current setting gives the lowest cost for each confirmed case on the whole sweep. Is that a finding?
What does the marginal reading say that the average hides?
Somebody at the fraud desk asks a real question: should the desk investigate 360 more alerts a month? The average is about the whole arrangement and the question is about one move, so the average reading cannot answer it. The marginal readingWhat the next block of cases found actually cost, which is the figure a decision about the next step needs. is the one that answers it. Moving from 540 kept to 900 kept adds 10,800 planning minutes and buys 4 more confirmed cases, so those 4 cost 2,700 minutes each.
Put the three figures side by side and the disagreement is plain. The average at the current setting is 1,933 minutes for each case, the average at the new setting is 2,032, and the next 4 cases cost 2,700 apiece. The ranking put the easiest cases first, so the next cases always cost more than the average, and that is exactly the fact an average reading is built to hide. Whoever quotes 1,933 into that conversation is answering a question about the past, and nobody in the room is deciding anything about the past.
One more feature of those three numbers explains why the trough exists at all. The 4 cases bought on the way into the current setting cost 1,350 minutes each and sit below the average. The 4 cases bought on the way out cost 2,700 each and sit above it. An average bottoms out exactly where the next block of cases starts costing more than everything found so far, so the minimum marks the crossing point of two readings rather than a good place to stop. Nothing in that arithmetic knows what a missed fraud does to the person it happens to.
Somebody asks whether to investigate 360 more alerts a month. Which reading answers them?
Why can the sweep never find more than the bank already knows about?
Look at the top of the sweep once more. Keeping every one of the 18,000 alerts confirms 38 cases, and 38 is not a discovery. The 38 is the 27 the bank already caught plus the floor of 11 it already knew it had missed. The sweep is bounded by the bank's own records rather than by its effort, so the most exhaustive setting available finds only the misses that had already surfaced somewhere else in the bank.
A fraud that was never written off, never complained about and never noticed contributes nothing to any of the four counts. Such a fraud is not in the 27, not in the 11, and therefore cannot appear at the top of any sweep however much time is spent. The bound is a limit of the record, not a limit of the work, and no amount of investigating changes it. Think of a household counting how much food gets wasted by weighing what goes into the bin. Weighing more carefully improves nothing about the portion that never reached the bin.
Why can the sweep never confirm more than 38 cases?
How does somebody reading a monitoring pack actually use all this?
Here is the working habit, and it takes about a minute. A line arrives in a pack quoting a measure of a yes or no system. Before the value is reacted to, the question it answers is named, then the two or four counts it stands on are named, then each count is tested for whether it is exact, a floor or absent. The sequence is the whole tool, and it turns a number that cannot be argued with into a claim that can be inspected.
A household example makes the shape obvious. A shopkeeper says he catches most of the bad notes that come across his counter. Asked which counts he holds, it turns out he has the notes he stopped and checked, and nothing whatsoever about the notes he accepted without a second glance. He can say exactly how often he was right when he stopped somebody. He cannot say what got past him, and he should not answer that question at all. The reading a firm most wants is usually the one its records support least, and the discipline is to say so rather than to fill the gap.
Analysts and validators use the same sequence in the other direction. A model validator reading a pack line at Sumeru Bank Limited would ask Revathi Balan, as the named accountable person, which of the four counts the bank actually holds, and would expect the answer to be two exact, one floor and one absent before any reading is discussed at all. Neelima Rao, in the risk function, treats a reading whose counts nobody can name as an unresolved item rather than as a poor result, and the difference between an unresolved item and a poor result is what keeps a monitoring pack honest. O'Neil, in Weapons of Math Destruction, 2016, makes the related point that a model's errors fall unevenly across the people they land on, and the counts nobody keeps tend to be the errors borne by people with the least ability to complain about them.
One last point of practice explains why every cost in the sweep is stated in minutes. Sumeru Bank Limited holds an assumed fully loaded cost of Rs 9,00,000/- a year for a post, and its own build and running figures for the intake chain, Rs 2,40,00,000/- once and Rs 65,00,000/- a year. None of the three enters any reading of the sweep. Converting minutes into money adds another assumption on top of the two planning rates, and a reading that stands on three assumptions is harder to defend than one that stands on two, so the bank's own sweep stays in minutes.
Who expects a lender to be able to answer this?
Where a measure of a deployed system is reported to a board or to a supervisor, the expectations on a regulated lender in India are set by the Reserve Bank of India and published at rbi.org.in, covering the oversight of arrangements that decide customer outcomes. The standing discipline of model risk work originates in international supervisory material from the Bank for International Settlements at bis.org. The material is an origin rather than the position in India, and the two are quoted interchangeably far more often than the difference allows.
What do the three questions leave out?
Which question each measure of a yes or no system answers, and which counts stand behind it, is settled by the three questions and the four counts above. How each of those measures is computed from its counts is a separate subject, and how a model is checked before it goes live is set out under model testing. So is the naming itself, where one write-up calls a measure something another does not. The cost of each of the two errors, and who bears it, is set out under false positives and false negatives. How a fraud desk is staffed, how alerts are suppressed before anybody reads them, and what a bad account costs the bank are separate subjects again.
Is the calculation of each measure set out above?
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Published expectations on a regulated lender covering the oversight of arrangements that decide customer outcomes, and what is reported about them | rbi.org.in |
| Bank for International Settlements | International supervisory material from which the standing discipline of model risk work originates. An origin rather than the position in India | bis.org |
| O'Neil | Weapons of Math Destruction, 2016. Argues that the errors of a model fall unevenly across the people they land on | Crown |
Sumeru Bank Limited, its intake chain, Revathi Balan and Neelima Rao are invented.
Educational material. Not advice on any investment, tax, budget or market position.
