Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quantitative Methods, Financial Data & Programming
1Probability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
2Statistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
3Correlation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
4Time Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
5Simulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
6Optimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
7Modelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
8Backtesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
9Data Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
10Programming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
11Quantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity

The ROC Curve and AUC: Trading False Positives Against False Negatives, in One Number

A receiver operating characteristic (ROC) curve plots right calls up and wrong calls across as the calling threshold sweeps from high to low, and the area under it compresses that whole sweep into one number. On the ten paired months of the Nakshatra unit and the Vasant unit relabelled up or down, that area is 0.8000, computed twice: by measuring under the curve, and by counting all 25 pairs.

Every quantity below falls out of ten paired months made up for teaching, put through arithmetic short enough to check by hand: five distinct scores, six points, twenty five pairings. Counting the pairings on a sheet of paper lands on the same twenty out of twenty five printed below. Arithmetic on made up months needs no market, no maintained record and no published return. A pen settles all of it.

Picture a shopkeeper with a crate of mangoes. He handles each one, judges how ripe it feels, and lays the crate out in a line from softest to firmest. Then, separately, he decides where along that line to put the divider between sell today and keep for tomorrow. Two completely different acts. How well he judged ripeness is settled the moment the line is laid out. Where the divider goes is a decision he makes afterwards, and he can move it a dozen times without touching a single mango.

Grading the crate and placing the divider come apart completely, and one number measures the grading while taking no account of the divider. The number is the area under a curve called the ROC curve. What that area counts, and what it refuses to count, is the whole subject.

Three things carry in from earlier reading and none of them is rebuilt here. The ten paired months of the Nakshatra unit and the Vasant unit, both invented and standing for nothing. The relabelling of the Vasant unit as simply up or down, five months up and five months down. And a fitted straight line on that label which hands every month a scoreA number a fitted rule attaches to one row, used only for putting the rows in an arrangement. Where the number came from is settled in earlier notes and taken as given here. between nought and one: 0.4500 plus 0.0500 times that month's Nakshatra reading. The Nakshatra readings run 1.00, 6.00, minus 4.00, 11.00, 1.00, minus 9.00, 6.00, 1.00, minus 4.00 and 1.00 per cent, so the scores run 0.50, 0.75, 0.25, 1.00, 0.50, 0.00, 0.75, 0.50, 0.25 and 0.50.

The last thing carried in is a rule for turning a score into a call. Pick a thresholdA cut off value chosen by hand. Any row scoring at or above it gets called one way and every row below it gets called the other. Choosing one sensibly is a judgement covered separately., then call a month up when its score reaches that threshold and down when it does not. Move the threshold and the calls move. Nothing else about the model changes at all: not the fitted line, not one of the ten scores, not the arrangement they sit in. Everything that follows grows out of that single sentence.

What is actually being traded here, and against what?

Start with a tea cart outside an office gate. The man running it has to decide how many cups to brew before the five o'clock rush arrives. Brew a lot and he serves everyone who wants tea, and he also tips a good quantity down the drain. Brew a little and he wastes nothing, and he turns away people who came for a cup. There is no quantity that avoids both. He is not choosing between a good outcome and a bad one; he is choosing which of two bad ones he would rather absorb.

Calling months up and down has exactly that shape. Dropping the threshold calls more months up. More of the five months that really did go up are caught, and more of the five months that really went down are wrongly called up. Both counts move the same way, always. Raising the threshold shrinks both together. There is no setting of the threshold that improves both counts at once, so picking one is a decision about which of the two mistakes costs more.

The five thresholds worth trying on this record, and what each one does to the two counts. Every figure computed from the ten scores above.
ThresholdMonths called upOf the five up months, caughtOf the five down months, wrongly called up
1.00110
0.75330
0.50743
0.25954
0.001055
AS THE THRESHOLD FALLS, BOTH COUNTS CLIMBof the five up months, caughtof the five down months, wrongly called upthreshold 1.001 of 50 of 5threshold 0.753 of 50 of 5threshold 0.504 of 53 of 5threshold 0.255 of 54 of 5threshold 0.005 of 55 of 5Neither count ever falls as the line comes down. There is no setting that trims one without swelling the other.
Both counts climb together as the threshold falls, from one caught and none wrongly called at a threshold of 1.00 to five and five at a threshold of nought, so no setting trims one without swelling the other.
Try it out

The threshold comes down from 0.75 to 0.50. Which two things move in the same direction as it falls?

What do the two axes measure?

The curve lives inside a square whose sides both run from nought to one. Across the bottom goes the share of the five down months that were wrongly called up. Up the side goes the share of the five up months that were rightly called up. Both are shares, and here is the part that matters more than it first looks: both divide by a group whose size was fixed before anyone touched the threshold, so both are trapped between nought and one whatever the threshold does, and the entire sweep fits inside one square.

Five down months. Five up months. The two fives were settled the moment the outcome was relabelled, and no threshold can add a month to either group or take one away. All a threshold does is decide how many of each five end up on the called-up side. At a threshold of 0.50 the answer is three of the five down months and four of the five up months, so the point sits at 0.60 across and 0.80 up.

Had the axes been raw counts instead of shares, the square would have stretched or shrunk with the record, and two records holding different numbers of up and down months could never have been drawn on the same picture. Dividing by the fixed group is what makes the drawing portable. The same instinct quotes a shop's wastage as a share of what it stocked rather than as a number of unsold items. A stall and a supermarket can then be set beside each other at all.

ONE SQUARE, TWO SHARES, BOTH OUT OF FIVE0.00.00.20.20.40.40.60.60.80.81.01.0the perfect corner: every upmonth caught, no down monthwrongly called upeverything called upnothing called up at alla rule that sorts nothingruns along this lineacross: the share of the FIVE DOWN MONTHS wrongly called upup: the share of the FIVE UP MONTHS rightly called upThe two groups arecounted once, before thethreshold is touched:five up monthsfive down monthsMoving the threshold cannotchange either of those fives,so both shares stay insidenought and one.
The square is fixed before any threshold is chosen: the across axis divides by the five down months and the up axis divides by the five up months, so every possible setting lands somewhere inside it.
Try it out

Both axes are shares of a group whose size is fixed in advance. Why is that worth insisting on?

How is the curve built from one set of scores?

Look again at the ten scores. Written out they are 0.50, 0.75, 0.25, 1.00, 0.50, 0.00, 0.75, 0.50, 0.25 and 0.50, and although there are ten of them there are only five distinct values in the list: 1.00, 0.75, 0.50, 0.25 and 0.00. A threshold sitting anywhere between two neighbouring scores produces exactly the same ten calls as a threshold sitting on the lower of the two. Only the moments when the threshold crosses a score change anything.

So the sweepRunning one setting across its whole range from one end to the other and recording what happens at every stop, rather than trying a single value and reporting that. has a natural set of stopping places. Start with the threshold above everything, higher than 1.00. Nothing is called up. No up month is caught and no down month is wrongly called. The point sits at nought across and nought up, the originThe corner of a chart where both axes read nought. Here it is the bottom left of the square.. Now drop the threshold past each distinct score in turn and mark where the two shares land each time. Five drops, five more points, six in all.

  • Threshold above 1.00: nothing called up. 0.00 across, 0.00 up.
  • Threshold 1.00: only the month scoring 1.00 is called up, and it was an up month. 0.00 across, 0.20 up.
  • Threshold 0.75: the two months scoring 0.75 join it, and both were up months. 0.00 across, 0.60 up.
  • Threshold 0.50: four months scoring 0.50 join in, one up and three down. 0.60 across, 0.80 up.
  • Threshold 0.25: two months scoring 0.25 join in, one up and one down. 0.80 across, 1.00 up.
  • Threshold 0.00: the last month joins in, and it was a down month. 1.00 across, 1.00 up.

The curve is not a shape fitted to anything: it is the sweep itself, six readings joined in the order the threshold visited them. Notice how the record announces itself in the shape. The three highest scores all belong to months that really did go up, so the first three points climb straight up the left edge and the model is spending them for free. The value 0.50 is shared by one up month and three down months, and they all cross together, so the fourth point lurches sideways. The steepness at the start is the model being right; the sideways lurch is the price it pays for that.

THE THRESHOLD DROPPED PAST EACH DISTINCT SCORE, AND WHERE IT LANDSthresholdthe ten months in the order they arrivedthe point that landsabove 1.00UUUUDDUDDD0.00 across0.00 up1.00UUUUDDUDDD0.00 across0.20 up0.75UUUUDDUDDD0.00 across0.60 up0.50UUUUDDUDDD0.60 across0.80 up0.25UUUUDDUDDD0.80 across1.00 up0.00UUUUDDUDDD1.00 across1.00 upU is a month whose outcome was up, D a month whose outcome was down. A shaded box is a month called up at that threshold.
Each row drops the threshold past one more distinct score, shades the months called up at that setting, and prints the point that lands, which is how the six readings on the curve are produced.
THE SIX POINTS, JOINED, AND THE AREA THEY ENCLOSE0.00.00.20.20.40.40.60.60.80.81.01.00.60 across, 0.80 up0.00 across, 0.60 uparea 0.8000a line thatsorts nothing:area 0.5000share of the five down months wrongly called upshare of the five up months rightly called upThe three pieces thatcarry any width at all:0.00 to 0.600.60 wide, 0.70 tall on average0.42000.60 to 0.800.20 wide, 0.90 tall on average0.18000.80 to 1.000.20 wide, 1.00 tall on average0.2000added: 0.8000The first two stepsare nought wide, sothey add nothing.
The six points joined and the area beneath them shaded, giving 0.8000 against 0.5000 for the dashed diagonal that a rule sorting nothing would trace.
Try it out

Set the threshold at 0.60, a value between two of the distinct scores rather than on one of them. Every month scoring 0.60 or more is called up. Which point on the curve does that setting land on?

What is the AUC, and what does an area of one half mean?

The area under the curve (AUC) is exactly that: the region beneath the joined points, down to the bottom of the square, measured as a share of the square. Since the square has sides of one, its whole area is one, and the shaded part is a number between nought and one. Nothing more elaborate is going on.

Measuring it here needs no calculus, only the area of a trapeziumA four sided shape with one pair of parallel sides. Its area is the distance between those two sides multiplied by their average length.. The area of one is its width multiplied by the average of its two heights. Walk along the six points from left to right. The first two steps go straight up with no sideways movement at all, so they are nought wide and contribute nothing. Three steps remain.

The area under the six points, worked out as three trapezium pieces. Widths and heights are read straight off the point list above.
The step acrossWidthAverage heightArea of that piece
0.00 to 0.600.600.700.4200
0.60 to 0.800.200.900.1800
0.80 to 1.000.201.000.2000
The whole square beneath the curve0.8000

Now the two ends of the scale. A curve that runs hard up the left edge and then straight along the top misses almost none of the square and has an area close to one: that is a rule whose scores separate the up months from the down months cleanly. A curve that lies along the diagonal cuts the square in half and has an area of 0.5000. A coin flipA rule that decides by chance alone and carries no information about the thing it is deciding, used here as the do-nothing case to measure against. puts the months in no arrangement worth having, so a coin flip scores an area of 0.5000. At every threshold it drags up months and down months across the line at the same rate, and that traces the diagonal exactly.

One warning before going further. Two numbers below land on identical digits while bearing no relationship of any kind. A coin flip's area of 0.5000 is not the same creature as a rule that calls every single month up. Such a rule gets 50.00 per cent of its calls right on this record, and it does so because five of the ten months really did go up. One is an area between nought and one and the other is a share of calls; they land on the same digits by arithmetic accident on this particular record, and the resemblance carries no meaning whatsoever.

Why does counting pairs land on the same number?

Here is a second road to the identical figure, and it never mentions a threshold, a curve or an area. Take the five up months and the five down months. Pair each up month with each down month, for 25 pairings in all. For each pairing ask one question: did the up month carry the higher score? Count one where it did, nothing where it did not, and a half where the two tieTwo rows carrying exactly the same score, so neither can be placed above the other. Splitting the credit down the middle is the only even handed way to settle one..

Arrange the ten months by score, highest first, and the answer is almost visible. The arrangement runs: up, up, up, up, down, down, down, up, down, down. The top four are all up months. Every one of them beats every down month underneath, and the wins pile up straight away. Then three down months sit at 0.50, level with one up month. Then one up month at 0.25 sits below those three down months. The arrangement goes wrong in that one place and nowhere else.

Counted out in full: 18 pairings the up month wins outright, 4 pairings the two scores tie, and 3 pairings the down month scores higher. All 25 pairings are accounted for. Eighteen wins count one each and four ties count a half each, so the total is 18 plus 2, or 20. And 20 out of 25 is 0.8000.

The two roads agree exactly rather than to four decimal places, and the second road is the one that says what the area actually means: it is the chance that a month picked at random from the up group carries a higher score than a month picked at random from the down group. Keep that sentence. The quantity genuinely does not involve a threshold, a call or an accuracy, so the sentence mentions none of them. The area is a statement about arrangement, and nothing else.

ALL 25 PAIRINGS OF ONE UP MONTH AGAINST ONE DOWN MONTHdown months, and the score each carriesup months0.500.500.500.250.001.00winwinwinwinwin0.75winwinwinwinwin0.75winwinwinwinwin0.50tietietiewinwin0.25losslosslosstiewin18 wins count one each4 ties count a half each3 losses count nothing18 plus 4 halves is 20, and 20 out of 25 is 0.8000
All 25 pairings laid out as a grid, with eighteen wins, four ties and three losses, so eighteen plus four halves is twenty and twenty out of twenty five is 0.8000.
TWO ROADS TO THE SAME NUMBER, AND THEY MEET EXACTLYROAD ONE: MEASURESweep the threshold, marksix points, join them, andtake the area beneath.0.4200 plus 0.1800 plus 0.2000ROAD TWO: COUNTPair every up month withevery down month, 25 pairs,and count the higher score.18 wins plus 4 halves is 200.8000one number, arrived at twiceNot equal to four decimal places by luck: the trapezium sum and the pair count are the same arithmetic written two ways.
Measuring under the six joined points and counting the twenty five pairings are two different pieces of arithmetic that land on the same 0.8000 exactly, not approximately.
Try it out

The area on this record is 0.8000. Said as a plain sentence about the ten months, what does that number claim?

Breaking Into Quants Bootcamp — Fin Maverick

Why does the area not move when the threshold moves?

The calls change a great deal as the threshold sweeps: at 0.00 every month is called up and at 1.00 only one is. Does the area move with them?

Try it out

Across the five thresholds on this record the share of calls that come out right runs from 50.00 per cent up to 80.00 per cent. Over that same sweep, what does the area under the curve do?

The answer is that the area does not move, and it is worth being precise about why rather than just noting it. Look at how the curve was built. The picture is the whole sweep, so every single one of the six points was already on it before any threshold was chosen. Choosing a threshold does not build a different curve; it picks out one of the six points that were sitting there all along and says this is the one I am working at. The shaded region beneath the curve is untouched by that choice, in the same way that circling a name on a printed list does not change the list.

The area is a property of the arrangement the months are put in by their scores, and a threshold does not change that arrangement, only where a line is drawn through it. The shopkeeper again. Sliding the divider along the row of mangoes says nothing new about how well he graded them; it says only what he is doing with the grading today. Slid far left, he sells nearly the whole crate, including some hard ones. Slid far right, he sells only the softest few and holds back some that were ready. The row itself has not moved a centimetre.

The five thresholds, the four calls at each, the share of calls that come out right, and the area. Every row computed from the same ten scores.
ThresholdCalled up and upCalled up and downCalled down and upCalled down and downRight callsArea
1.00104560.00 per cent0.8000
0.75302580.00 per cent0.8000
0.50431260.00 per cent0.8000
0.25540160.00 per cent0.8000
0.00550050.00 per cent0.8000

Read the last two columns against each other. One of them swings by thirty percentage points on a model nobody has retrained, refitted or altered in any way. The other prints the same six characters five times. The disagreement is not a defect in either measure. The two columns answer different questions, and a reader who wants both has to ask for both.

ONE COLUMN MOVES THIRTY POINTS, THE OTHER DOES NOT MOVE AT ALLthresholdthe four callsaccuracythe area1.00called up: 1 up and 0 downcalled down: 4 up and 5 down60.00 per cent0.80000.75called up: 3 up and 0 downcalled down: 2 up and 5 down80.00 per cent0.80000.50called up: 4 up and 3 downcalled down: 1 up and 2 down60.00 per cent0.80000.25called up: 5 up and 4 downcalled down: 0 up and 1 down60.00 per cent0.80000.00called up: 5 up and 5 downcalled down: 0 up and 0 down50.00 per cent0.8000Same model, same scores, same order. Only the place the line is drawn has changed between one row and the next.
Across the five thresholds the share of right calls moves between 50.00 and 80.00 per cent while the area beside it repeats 0.8000 unchanged in every row.
Play with it

Drag the threshold and watch three things move while a fourth refuses to.

One control moves: the calling threshold, from 0.00 to 1.00 in hundredths, held as a whole number of hundredths so nothing can drift. The top strip shows the ten months arranged by score, highest on the left, with the cut sliding through them. The left panel puts a marker on the fixed curve. The middle panel redraws the four calls. The bottom panel plots the share of right calls against the threshold as a staircase, with the area drawn flat across it. The opening setting is a threshold of 0.50, where the four calls read 4, 3, 1 and 2, the share of right calls is 60.00 per cent and the area is 0.8000, and every one of those readings is also sitting in the table above as ordinary text, for anyone who never touches the control.

Jump to a setting, or nudge it one hundredth at a time:
Called up and it was up
4
Called up and it was down
3
Called down and it was up
1
Called down and it was down
2
Share of calls right
60.00 per cent
The point on the curve
0.60 across, 0.80 up
The area under the curve
0.8000
Loading the panel.

Educational illustration on an invented record. The Nakshatra unit and the Vasant unit exist only in these notes, and the up and down labels describe made up outcomes rather than anything that happened. The ten scores are held fixed while the threshold moves: the only thing changing anywhere on the panel is where the line falls. Arranging the months by score here is a display choice for this panel alone, and the order the months arrived in is untouched by it.

Try it out

Strip everything else away. What is the area actually a property of?

Can two different forms share one area?

Earlier reading established a second fitted form on this same up or down label, a logistic one, and it hands out visibly different numbers. Where the straight line gives 0.00, 0.25, 0.50, 0.75 and 1.00 across the five distinct Nakshatra readings, the logistic form gives 0.0553, 0.1947, 0.5000, 0.8053 and 0.9447. Different at every reading except the middle one. Feed those scores through everything above and the area comes out at 0.8000, identical to four decimal places and beyond.

Writing that down as a finding is tempting. Resist it. The logistic form is a rising transformA rule that turns each number into another number and never turns a bigger one into a smaller one. Feeding a list through one leaves every row in the same relative position. of the same underlying reading, and a rising transform cannot make any month overtake any other. The two areas had to match, and no other outcome was arithmetically available. Month four had the highest score under the straight line and it still has the highest under the logistic form. Month six had the lowest and it still does. Every tie under one form is a tie under the other. Since the area counts nothing but which of two months sits higher, and nothing has changed about which of two months sits higher, the count cannot change.

Presenting that as evidence about the two forms would be inventing a result out of a definition. The match is closer to observing that a queue is in the same order whether the people are numbered from the front or measured by their distance from the door. The match does say something about the measure rather than about the forms: the area is deaf to how far apart the scores are and hears only which is bigger. Reading it as a statement about how confident any score was is therefore ruled out.

A reader tends to assume this cannot happen, so one demonstration of that deafness is worth having. The straight line score is not a chance and was never constrained to behave like one. Push its input far enough and it walks straight past one: at a Nakshatra reading of 15.00 per cent, four points past the largest month anywhere on this record, the straight line reads 1.20, and no chance can read 1.20. The logistic form at that same reading of 15.00 per cent gives 0.9816 and stays under one, as it always will. On the ten months actually present the area sees only the arrangement, and the arrangement is identical, so the area cannot tell those two behaviours apart.

DIFFERENT SCORES, IDENTICAL ORDER, THEREFORE IDENTICAL AREAthe straight line on the labelthe logistic form on the same labelmonth 41.000.9447outcome upmonth 20.750.8053outcome upmonth 70.750.8053outcome upmonth 10.500.5000outcome upmonth 50.500.5000outcome downmonth 80.500.5000outcome downmonth 100.500.5000outcome downmonth 30.250.1947outcome upmonth 90.250.1947outcome downmonth 60.000.0553outcome downhighestlowestEvery joining line runs flat. No month overtakes another, so the twenty five pairings come out the same and so does the area.
The straight line scores and the logistic scores differ at nearly every month, yet no month overtakes another between the two columns, which is why both return an area of 0.8000.
Try it out

The logistic form gives quite different scores and exactly the same area. Is that evidence that the two fitted forms are equally good?

Ten months is a very small record, and it was built to cooperate. Five up and five down keeps both denominators equal. Records met in the wild are almost never that neat, and the neatness is a property of the lesson rather than of the subject. Carried onto a lopsided record, what still holds is the argument's shape: a sorting can be measured on its own terms, a cut through that sorting is a separate decision made afterwards, and one number cannot report on both at once.

AI For Finance Bootcamp — Fin Maverick Rebalancing: When, Why and What It Costs — free micro-course from Fin Maverick

What does the area not say?

Somebody supplies a report carrying one line: the area is 0.8000. Before that line is any use, three questions need answering, and the area answers none of them itself.

The area does not say which threshold to use. The same number appears at every threshold, so the area cannot possibly prefer one. Choosing a threshold means deciding which of the two mistakes hurts more, and that is a judgement about consequences rather than a calculation, covered separately.

The area does not say how many calls will come out right. The count depends entirely on the threshold, and on this record the range is wide: 50.00 per cent at one end of the sweep and 80.00 per cent at the other, on a model that never changed. Taking one of those figures as the answer means quietly picking a threshold without saying so.

The area does not say whether the scores mean anything as chances. A score of 0.75 under the straight line and a score of 0.8053 under the logistic form produce the identical area, so the area cannot be sensitive to what the numbers claim about themselves. Where it matters whether a score of 0.75 corresponds to anything, the area is silent and a different check is required.

A model with an area of 0.8000 can be right half the time or four fifths of the time on the very same ten months, depending entirely on a choice the area is built not to see. So the working habit is simple and worth adopting permanently: an area is never quoted on its own. An area is quoted together with the threshold actually in use and the four calls that go with it. The three together describe a decision somebody made, and the area alone describes none.

Think of it as a reference on a job applicant that says only ranks well against others seen so far. Genuinely useful information, and completely silent on what the person should be hired to do on Monday morning. The reference is not thrown away. The reference is also not treated as the job description.

Try it out

An area of 0.8000 arrives with a question attached: how many of the model's calls will come out right? What is the correct response?

The failure: reading an area of 0.8000 as an accuracy of 80.00 per cent

The mistake is the commonest failure with this measure, and it goes wrong quietly. Somebody reads that the area is 0.8000, writes in a summary that the model is right about 80 per cent of the time, and nobody catches it because the sentence sounds entirely reasonable.

The sentence is wrong twice over. First, the two are different quantities: one counts pairings and involves no threshold at all, the other counts calls and is meaningless without one. Second, on this record the figure is simply not true at the threshold most people would reach for. At 0.50 the model gets 60.00 per cent of its calls right, not 80.00 per cent. The value 80.00 per cent does appear on this record, at a threshold of 0.75, and that coincidence is more dangerous than a clean error would be. The mistaken sentence survives a spot check.

Three figures in this guide land on the digits eight and nought and no two of them have anything to do with each other: the area of 0.8000, the share of right calls at a threshold of 0.75 which reads 80.00 per cent, and the height of one point on the curve which reads 0.80. The resemblance is arithmetic accident on this particular record. Read nothing into it.

One habit removes the risk completely, and it asks for no extra effort. Never write an area down on its own line. Write the area, then the threshold, then the four calls, in that order, every time. The accuracy stands right there beside it in a form nobody can mistake, so a summary that says area 0.8000, threshold 0.50, calls 4, 3, 1 and 2 cannot be misread as a claim about accuracy.

THE SAME DIGITS, THREE DIFFERENT QUANTITIES0.8000the areaa fact about the order80.00 per centthe misreadingnot a quantity at all60.00 per centthe accuracy at 0.50a fact about one cutThe left panel and the right panel are both true of this same model at the same moment.The middle panel is what somebody writes down when they read the left panel as though it werea share of calls. It is struck through because nothing on this record carries that value: at athreshold of 0.75 the accuracy does read 80.00 per cent, and that is a different setting entirely.The digits agreeing is arithmetic coincidence and carries no meaning.
The area of 0.8000 and the share of right calls of 60.00 per cent are both true of this model at the same moment, and the 80.00 per cent struck through between them is neither of those things.
Try it out

Why is reading an area of 0.8000 as an accuracy of 80.00 per cent such an easy mistake to make?

Covered elsewhere. The two named views of a classifier that are built from the same four calls are covered separately, as is how a model is tested across many different cuts of one record. How a threshold should be chosen when one mistake genuinely costs more than the other is a judgement about consequences rather than a calculation and is treated on its own. The error measures used when the outcome is a size rather than a label are also covered separately. Whether calling a month up or down would be worth anything to anybody in a market is a separate question again.

One area across every threshold, preferring none. See what the area cannot say.

How can every number here be checked?

Arithmetic carries its own proof: how many of twenty five comparisons went one way is settled by comparing them, not by citing anybody. Every figure below can be rebuilt with a pen, and each recipe is short enough to run in a minute.

Every figure in this guide, what it falls out of, and the arithmetic needed to rebuild it from scratch.
What is printedWhat it falls out ofWhat rebuilding it takes
The ten scores, 0.50 through 0.000.4500 plus 0.0500 times each Nakshatra reading in turnTen multiplications and ten additions
The six points on the curveAt each distinct score, how many of the five up months and how many of the five down months sit at or above itFive pairs of counts, then a division by five
The area 0.8000, measuredThree trapezium pieces, each a width times an average of two heightsThree multiplications and one addition
The area 0.8000, countedTwenty five comparisons, a win counting one and a tie counting a halfTwenty five glances at a five by five grid
The share of right calls, 50.00 to 80.00 per centThe two agreeing calls added and divided by ten, at each of the five settingsOne addition and one division per setting
The logistic scores, 0.0553 to 0.9447One over one plus the exponential of minus 0.2839 times the reading less 1.00A calculator carrying an exponential key

The Nakshatra unit and the Vasant unit are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.