The ROC Curve and AUC: Trading False Positives Against False Negatives, in One Number
A receiver operating characteristic (ROC) curve plots right calls up and wrong calls across as the calling threshold sweeps from high to low, and the area under it compresses that whole sweep into one number. On the ten paired months of the Nakshatra unit and the Vasant unit relabelled up or down, that area is 0.8000, computed twice: by measuring under the curve, and by counting all 25 pairs.
Every quantity below falls out of ten paired months made up for teaching, put through arithmetic short enough to check by hand: five distinct scores, six points, twenty five pairings. Counting the pairings on a sheet of paper lands on the same twenty out of twenty five printed below. Arithmetic on made up months needs no market, no maintained record and no published return. A pen settles all of it.
Picture a shopkeeper with a crate of mangoes. He handles each one, judges how ripe it feels, and lays the crate out in a line from softest to firmest. Then, separately, he decides where along that line to put the divider between sell today and keep for tomorrow. Two completely different acts. How well he judged ripeness is settled the moment the line is laid out. Where the divider goes is a decision he makes afterwards, and he can move it a dozen times without touching a single mango.
Grading the crate and placing the divider come apart completely, and one number measures the grading while taking no account of the divider. The number is the area under a curve called the ROC curve. What that area counts, and what it refuses to count, is the whole subject.
Three things carry in from earlier reading and none of them is rebuilt here. The ten paired months of the Nakshatra unit and the Vasant unit, both invented and standing for nothing. The relabelling of the Vasant unit as simply up or down, five months up and five months down. And a fitted straight line on that label which hands every month a scoreA number a fitted rule attaches to one row, used only for putting the rows in an arrangement. Where the number came from is settled in earlier notes and taken as given here. between nought and one: 0.4500 plus 0.0500 times that month's Nakshatra reading. The Nakshatra readings run 1.00, 6.00, minus 4.00, 11.00, 1.00, minus 9.00, 6.00, 1.00, minus 4.00 and 1.00 per cent, so the scores run 0.50, 0.75, 0.25, 1.00, 0.50, 0.00, 0.75, 0.50, 0.25 and 0.50.
The last thing carried in is a rule for turning a score into a call. Pick a thresholdA cut off value chosen by hand. Any row scoring at or above it gets called one way and every row below it gets called the other. Choosing one sensibly is a judgement covered separately., then call a month up when its score reaches that threshold and down when it does not. Move the threshold and the calls move. Nothing else about the model changes at all: not the fitted line, not one of the ten scores, not the arrangement they sit in. Everything that follows grows out of that single sentence.
What is actually being traded here, and against what?
Start with a tea cart outside an office gate. The man running it has to decide how many cups to brew before the five o'clock rush arrives. Brew a lot and he serves everyone who wants tea, and he also tips a good quantity down the drain. Brew a little and he wastes nothing, and he turns away people who came for a cup. There is no quantity that avoids both. He is not choosing between a good outcome and a bad one; he is choosing which of two bad ones he would rather absorb.
Calling months up and down has exactly that shape. Dropping the threshold calls more months up. More of the five months that really did go up are caught, and more of the five months that really went down are wrongly called up. Both counts move the same way, always. Raising the threshold shrinks both together. There is no setting of the threshold that improves both counts at once, so picking one is a decision about which of the two mistakes costs more.
| Threshold | Months called up | Of the five up months, caught | Of the five down months, wrongly called up |
|---|---|---|---|
| 1.00 | 1 | 1 | 0 |
| 0.75 | 3 | 3 | 0 |
| 0.50 | 7 | 4 | 3 |
| 0.25 | 9 | 5 | 4 |
| 0.00 | 10 | 5 | 5 |
The threshold comes down from 0.75 to 0.50. Which two things move in the same direction as it falls?
What do the two axes measure?
The curve lives inside a square whose sides both run from nought to one. Across the bottom goes the share of the five down months that were wrongly called up. Up the side goes the share of the five up months that were rightly called up. Both are shares, and here is the part that matters more than it first looks: both divide by a group whose size was fixed before anyone touched the threshold, so both are trapped between nought and one whatever the threshold does, and the entire sweep fits inside one square.
Five down months. Five up months. The two fives were settled the moment the outcome was relabelled, and no threshold can add a month to either group or take one away. All a threshold does is decide how many of each five end up on the called-up side. At a threshold of 0.50 the answer is three of the five down months and four of the five up months, so the point sits at 0.60 across and 0.80 up.
Had the axes been raw counts instead of shares, the square would have stretched or shrunk with the record, and two records holding different numbers of up and down months could never have been drawn on the same picture. Dividing by the fixed group is what makes the drawing portable. The same instinct quotes a shop's wastage as a share of what it stocked rather than as a number of unsold items. A stall and a supermarket can then be set beside each other at all.
Both axes are shares of a group whose size is fixed in advance. Why is that worth insisting on?
How is the curve built from one set of scores?
Look again at the ten scores. Written out they are 0.50, 0.75, 0.25, 1.00, 0.50, 0.00, 0.75, 0.50, 0.25 and 0.50, and although there are ten of them there are only five distinct values in the list: 1.00, 0.75, 0.50, 0.25 and 0.00. A threshold sitting anywhere between two neighbouring scores produces exactly the same ten calls as a threshold sitting on the lower of the two. Only the moments when the threshold crosses a score change anything.
So the sweepRunning one setting across its whole range from one end to the other and recording what happens at every stop, rather than trying a single value and reporting that. has a natural set of stopping places. Start with the threshold above everything, higher than 1.00. Nothing is called up. No up month is caught and no down month is wrongly called. The point sits at nought across and nought up, the originThe corner of a chart where both axes read nought. Here it is the bottom left of the square.. Now drop the threshold past each distinct score in turn and mark where the two shares land each time. Five drops, five more points, six in all.
- Threshold above 1.00: nothing called up. 0.00 across, 0.00 up.
- Threshold 1.00: only the month scoring 1.00 is called up, and it was an up month. 0.00 across, 0.20 up.
- Threshold 0.75: the two months scoring 0.75 join it, and both were up months. 0.00 across, 0.60 up.
- Threshold 0.50: four months scoring 0.50 join in, one up and three down. 0.60 across, 0.80 up.
- Threshold 0.25: two months scoring 0.25 join in, one up and one down. 0.80 across, 1.00 up.
- Threshold 0.00: the last month joins in, and it was a down month. 1.00 across, 1.00 up.
The curve is not a shape fitted to anything: it is the sweep itself, six readings joined in the order the threshold visited them. Notice how the record announces itself in the shape. The three highest scores all belong to months that really did go up, so the first three points climb straight up the left edge and the model is spending them for free. The value 0.50 is shared by one up month and three down months, and they all cross together, so the fourth point lurches sideways. The steepness at the start is the model being right; the sideways lurch is the price it pays for that.
Set the threshold at 0.60, a value between two of the distinct scores rather than on one of them. Every month scoring 0.60 or more is called up. Which point on the curve does that setting land on?
What is the AUC, and what does an area of one half mean?
The area under the curve (AUC) is exactly that: the region beneath the joined points, down to the bottom of the square, measured as a share of the square. Since the square has sides of one, its whole area is one, and the shaded part is a number between nought and one. Nothing more elaborate is going on.
Measuring it here needs no calculus, only the area of a trapeziumA four sided shape with one pair of parallel sides. Its area is the distance between those two sides multiplied by their average length.. The area of one is its width multiplied by the average of its two heights. Walk along the six points from left to right. The first two steps go straight up with no sideways movement at all, so they are nought wide and contribute nothing. Three steps remain.
| The step across | Width | Average height | Area of that piece |
|---|---|---|---|
| 0.00 to 0.60 | 0.60 | 0.70 | 0.4200 |
| 0.60 to 0.80 | 0.20 | 0.90 | 0.1800 |
| 0.80 to 1.00 | 0.20 | 1.00 | 0.2000 |
| The whole square beneath the curve | 0.8000 |
Now the two ends of the scale. A curve that runs hard up the left edge and then straight along the top misses almost none of the square and has an area close to one: that is a rule whose scores separate the up months from the down months cleanly. A curve that lies along the diagonal cuts the square in half and has an area of 0.5000. A coin flipA rule that decides by chance alone and carries no information about the thing it is deciding, used here as the do-nothing case to measure against. puts the months in no arrangement worth having, so a coin flip scores an area of 0.5000. At every threshold it drags up months and down months across the line at the same rate, and that traces the diagonal exactly.
One warning before going further. Two numbers below land on identical digits while bearing no relationship of any kind. A coin flip's area of 0.5000 is not the same creature as a rule that calls every single month up. Such a rule gets 50.00 per cent of its calls right on this record, and it does so because five of the ten months really did go up. One is an area between nought and one and the other is a share of calls; they land on the same digits by arithmetic accident on this particular record, and the resemblance carries no meaning whatsoever.
Why does counting pairs land on the same number?
Here is a second road to the identical figure, and it never mentions a threshold, a curve or an area. Take the five up months and the five down months. Pair each up month with each down month, for 25 pairings in all. For each pairing ask one question: did the up month carry the higher score? Count one where it did, nothing where it did not, and a half where the two tieTwo rows carrying exactly the same score, so neither can be placed above the other. Splitting the credit down the middle is the only even handed way to settle one..
Arrange the ten months by score, highest first, and the answer is almost visible. The arrangement runs: up, up, up, up, down, down, down, up, down, down. The top four are all up months. Every one of them beats every down month underneath, and the wins pile up straight away. Then three down months sit at 0.50, level with one up month. Then one up month at 0.25 sits below those three down months. The arrangement goes wrong in that one place and nowhere else.
Counted out in full: 18 pairings the up month wins outright, 4 pairings the two scores tie, and 3 pairings the down month scores higher. All 25 pairings are accounted for. Eighteen wins count one each and four ties count a half each, so the total is 18 plus 2, or 20. And 20 out of 25 is 0.8000.
The two roads agree exactly rather than to four decimal places, and the second road is the one that says what the area actually means: it is the chance that a month picked at random from the up group carries a higher score than a month picked at random from the down group. Keep that sentence. The quantity genuinely does not involve a threshold, a call or an accuracy, so the sentence mentions none of them. The area is a statement about arrangement, and nothing else.
The area on this record is 0.8000. Said as a plain sentence about the ten months, what does that number claim?
Why does the area not move when the threshold moves?
The calls change a great deal as the threshold sweeps: at 0.00 every month is called up and at 1.00 only one is. Does the area move with them?
Across the five thresholds on this record the share of calls that come out right runs from 50.00 per cent up to 80.00 per cent. Over that same sweep, what does the area under the curve do?
The answer is that the area does not move, and it is worth being precise about why rather than just noting it. Look at how the curve was built. The picture is the whole sweep, so every single one of the six points was already on it before any threshold was chosen. Choosing a threshold does not build a different curve; it picks out one of the six points that were sitting there all along and says this is the one I am working at. The shaded region beneath the curve is untouched by that choice, in the same way that circling a name on a printed list does not change the list.
The area is a property of the arrangement the months are put in by their scores, and a threshold does not change that arrangement, only where a line is drawn through it. The shopkeeper again. Sliding the divider along the row of mangoes says nothing new about how well he graded them; it says only what he is doing with the grading today. Slid far left, he sells nearly the whole crate, including some hard ones. Slid far right, he sells only the softest few and holds back some that were ready. The row itself has not moved a centimetre.
| Threshold | Called up and up | Called up and down | Called down and up | Called down and down | Right calls | Area |
|---|---|---|---|---|---|---|
| 1.00 | 1 | 0 | 4 | 5 | 60.00 per cent | 0.8000 |
| 0.75 | 3 | 0 | 2 | 5 | 80.00 per cent | 0.8000 |
| 0.50 | 4 | 3 | 1 | 2 | 60.00 per cent | 0.8000 |
| 0.25 | 5 | 4 | 0 | 1 | 60.00 per cent | 0.8000 |
| 0.00 | 5 | 5 | 0 | 0 | 50.00 per cent | 0.8000 |
Read the last two columns against each other. One of them swings by thirty percentage points on a model nobody has retrained, refitted or altered in any way. The other prints the same six characters five times. The disagreement is not a defect in either measure. The two columns answer different questions, and a reader who wants both has to ask for both.
Drag the threshold and watch three things move while a fourth refuses to.
One control moves: the calling threshold, from 0.00 to 1.00 in hundredths, held as a whole number of hundredths so nothing can drift. The top strip shows the ten months arranged by score, highest on the left, with the cut sliding through them. The left panel puts a marker on the fixed curve. The middle panel redraws the four calls. The bottom panel plots the share of right calls against the threshold as a staircase, with the area drawn flat across it. The opening setting is a threshold of 0.50, where the four calls read 4, 3, 1 and 2, the share of right calls is 60.00 per cent and the area is 0.8000, and every one of those readings is also sitting in the table above as ordinary text, for anyone who never touches the control.
Educational illustration on an invented record. The Nakshatra unit and the Vasant unit exist only in these notes, and the up and down labels describe made up outcomes rather than anything that happened. The ten scores are held fixed while the threshold moves: the only thing changing anywhere on the panel is where the line falls. Arranging the months by score here is a display choice for this panel alone, and the order the months arrived in is untouched by it.
Strip everything else away. What is the area actually a property of?
Can two different forms share one area?
Earlier reading established a second fitted form on this same up or down label, a logistic one, and it hands out visibly different numbers. Where the straight line gives 0.00, 0.25, 0.50, 0.75 and 1.00 across the five distinct Nakshatra readings, the logistic form gives 0.0553, 0.1947, 0.5000, 0.8053 and 0.9447. Different at every reading except the middle one. Feed those scores through everything above and the area comes out at 0.8000, identical to four decimal places and beyond.
Writing that down as a finding is tempting. Resist it. The logistic form is a rising transformA rule that turns each number into another number and never turns a bigger one into a smaller one. Feeding a list through one leaves every row in the same relative position. of the same underlying reading, and a rising transform cannot make any month overtake any other. The two areas had to match, and no other outcome was arithmetically available. Month four had the highest score under the straight line and it still has the highest under the logistic form. Month six had the lowest and it still does. Every tie under one form is a tie under the other. Since the area counts nothing but which of two months sits higher, and nothing has changed about which of two months sits higher, the count cannot change.
Presenting that as evidence about the two forms would be inventing a result out of a definition. The match is closer to observing that a queue is in the same order whether the people are numbered from the front or measured by their distance from the door. The match does say something about the measure rather than about the forms: the area is deaf to how far apart the scores are and hears only which is bigger. Reading it as a statement about how confident any score was is therefore ruled out.
A reader tends to assume this cannot happen, so one demonstration of that deafness is worth having. The straight line score is not a chance and was never constrained to behave like one. Push its input far enough and it walks straight past one: at a Nakshatra reading of 15.00 per cent, four points past the largest month anywhere on this record, the straight line reads 1.20, and no chance can read 1.20. The logistic form at that same reading of 15.00 per cent gives 0.9816 and stays under one, as it always will. On the ten months actually present the area sees only the arrangement, and the arrangement is identical, so the area cannot tell those two behaviours apart.
The logistic form gives quite different scores and exactly the same area. Is that evidence that the two fitted forms are equally good?
Ten months is a very small record, and it was built to cooperate. Five up and five down keeps both denominators equal. Records met in the wild are almost never that neat, and the neatness is a property of the lesson rather than of the subject. Carried onto a lopsided record, what still holds is the argument's shape: a sorting can be measured on its own terms, a cut through that sorting is a separate decision made afterwards, and one number cannot report on both at once.
What does the area not say?
Somebody supplies a report carrying one line: the area is 0.8000. Before that line is any use, three questions need answering, and the area answers none of them itself.
The area does not say which threshold to use. The same number appears at every threshold, so the area cannot possibly prefer one. Choosing a threshold means deciding which of the two mistakes hurts more, and that is a judgement about consequences rather than a calculation, covered separately.
The area does not say how many calls will come out right. The count depends entirely on the threshold, and on this record the range is wide: 50.00 per cent at one end of the sweep and 80.00 per cent at the other, on a model that never changed. Taking one of those figures as the answer means quietly picking a threshold without saying so.
The area does not say whether the scores mean anything as chances. A score of 0.75 under the straight line and a score of 0.8053 under the logistic form produce the identical area, so the area cannot be sensitive to what the numbers claim about themselves. Where it matters whether a score of 0.75 corresponds to anything, the area is silent and a different check is required.
A model with an area of 0.8000 can be right half the time or four fifths of the time on the very same ten months, depending entirely on a choice the area is built not to see. So the working habit is simple and worth adopting permanently: an area is never quoted on its own. An area is quoted together with the threshold actually in use and the four calls that go with it. The three together describe a decision somebody made, and the area alone describes none.
Think of it as a reference on a job applicant that says only ranks well against others seen so far. Genuinely useful information, and completely silent on what the person should be hired to do on Monday morning. The reference is not thrown away. The reference is also not treated as the job description.
An area of 0.8000 arrives with a question attached: how many of the model's calls will come out right? What is the correct response?
The failure: reading an area of 0.8000 as an accuracy of 80.00 per cent
The mistake is the commonest failure with this measure, and it goes wrong quietly. Somebody reads that the area is 0.8000, writes in a summary that the model is right about 80 per cent of the time, and nobody catches it because the sentence sounds entirely reasonable.
The sentence is wrong twice over. First, the two are different quantities: one counts pairings and involves no threshold at all, the other counts calls and is meaningless without one. Second, on this record the figure is simply not true at the threshold most people would reach for. At 0.50 the model gets 60.00 per cent of its calls right, not 80.00 per cent. The value 80.00 per cent does appear on this record, at a threshold of 0.75, and that coincidence is more dangerous than a clean error would be. The mistaken sentence survives a spot check.
Three figures in this guide land on the digits eight and nought and no two of them have anything to do with each other: the area of 0.8000, the share of right calls at a threshold of 0.75 which reads 80.00 per cent, and the height of one point on the curve which reads 0.80. The resemblance is arithmetic accident on this particular record. Read nothing into it.
One habit removes the risk completely, and it asks for no extra effort. Never write an area down on its own line. Write the area, then the threshold, then the four calls, in that order, every time. The accuracy stands right there beside it in a form nobody can mistake, so a summary that says area 0.8000, threshold 0.50, calls 4, 3, 1 and 2 cannot be misread as a claim about accuracy.
Why is reading an area of 0.8000 as an accuracy of 80.00 per cent such an easy mistake to make?
Covered elsewhere. The two named views of a classifier that are built from the same four calls are covered separately, as is how a model is tested across many different cuts of one record. How a threshold should be chosen when one mistake genuinely costs more than the other is a judgement about consequences rather than a calculation and is treated on its own. The error measures used when the outcome is a size rather than a label are also covered separately. Whether calling a month up or down would be worth anything to anybody in a market is a separate question again.
How can every number here be checked?
Arithmetic carries its own proof: how many of twenty five comparisons went one way is settled by comparing them, not by citing anybody. Every figure below can be rebuilt with a pen, and each recipe is short enough to run in a minute.
| What is printed | What it falls out of | What rebuilding it takes |
|---|---|---|
| The ten scores, 0.50 through 0.00 | 0.4500 plus 0.0500 times each Nakshatra reading in turn | Ten multiplications and ten additions |
| The six points on the curve | At each distinct score, how many of the five up months and how many of the five down months sit at or above it | Five pairs of counts, then a division by five |
| The area 0.8000, measured | Three trapezium pieces, each a width times an average of two heights | Three multiplications and one addition |
| The area 0.8000, counted | Twenty five comparisons, a win counting one and a tie counting a half | Twenty five glances at a five by five grid |
| The share of right calls, 50.00 to 80.00 per cent | The two agreeing calls added and divided by ten, at each of the five settings | One addition and one division per setting |
| The logistic scores, 0.0553 to 0.9447 | One over one plus the exponential of minus 0.2839 times the reading less 1.00 | A calculator carrying an exponential key |
The Nakshatra unit and the Vasant unit are invented.
Educational material. Not advice on any investment, tax, budget or market position.
