Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quantitative Methods, Financial Data & Programming
1Probability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
2Statistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
3Correlation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
4Time Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
5Simulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
6Optimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
7Modelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
8Backtesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
9Data Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
10Programming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
11Quantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity

Multiple Testing: Why Enough Hypotheses Guarantee a False Positive

Here is the uncomfortable thing about that sentence. Every single test in the run is exactly as careful as it was when it was the only one. Nobody loosened anything, nobody cheated, nobody picked a friendlier thresholdThe bar a reading must get past before anybody treats the result as worth reporting. What a bar set at 5 per cent claims about one single test is settled earlier in these notes and used here without being reopened.. And by the twentieth of them, the chance that at least one has fired for no reason at all has gone from one in twenty to roughly two in three. The gap between one in twenty and two in three is not a flaw in any test. The gap is arithmetic with one moving part, and the arithmetic breaks on the very record it was computed from.

Three things the arithmetic stands on

Independence, and a threshold of 5 per cent. Independence between two things, and the claim a single test held to 5 per cent makes, are both settled earlier in these notes. The subject here is what happens to that threshold the moment the test is run more than once.

One multiplication, repeated. If a single test comes back clean with chance 0.95, and the tests do not lean on one another, then all of them come back clean with chance 0.95 multiplied by itself once per test. The repeated multiplication is the entire mechanism. There is no second idea underneath it.

The Ashwin rule and its thirteen settings. The invented Ashwin rule reads the Nakshatra unit's change for a finished month and calls the next month up when that reading sits above a threshold. Its gridThe full list of settings a search will walk through, decided and written down before any walking starts rather than added to along the way. holds thirteen whole number settings. Thirteen settings is a real count of tests to put into the arithmetic instead of an imagined one.

What is the arithmetic, in one line?

Picture a stall on a street with a crate of tomatoes. The stallholder picks up one basket, turns it over, and looks for a bad one. Turning over one basket is a single check, and whatever the chance is of a bad one turning up by pure accident, it is what it is. Now picture the same stallholder going through twenty baskets and reporting the worst one he found. Nothing about his eyesight changed between the two stalls. The number of times he looked before choosing what to report is the only quantity that changed.

Put a number on it. A test held to a threshold of 5 per cent comes back clean 95 times in a hundred when there is genuinely nothing there. Write that as 0.95. Run two such tests that do not lean on one another and both come back clean 0.95 multiplied by 0.95 of the time. The product is 90.25 per cent. Everything else is at least one of the two firing, and one hundred less 90.25 leaves 9.75 per cent. The whole mechanism is that one multiplication repeated once per test, and for any count of tests the chance that at least one fires is 1 minus 0.95 raised to that count.

TWO TESTS, FOUR OUTCOMES, AND ONLY ONE OF THEM IS QUIETThe big block is both tests coming back clean. Everything outside it is at least one false positive.90.25per centboth tests come back cleansecond test, left to rightfirst testfiresboth come back clean90.25 per centonly the second one fires4.75 per centonly the first one fires4.75 per centboth of them fire0.25 per centat least one of the two fires9.75 per centwhich is one hundred less 90.25,and nothing else went into it
Two independent tests both come back clean 90.25 per cent of the time, so at least one of the two fires 9.75 per cent of the time, and that 9.75 is the whole of the square outside the plain block.

Notice the quantities missing from the picture. There is no adjustment for how important the question was, no allowance for the researcher being careful, and no term for how convincing the result looked. The only quantity that goes in is a count. The arithmetic is easy to state and easy to forget for the same reason: the quantity driving it is the least interesting fact about the work.

Try it out

Where does the 0.95 in this arithmetic come from?

What does the ladder look like as the count of tests rises?

Work it out at a spread of counts and put them in order. Nothing in this table is estimated and nothing in it is sampled. Each row is 0.95 raised to the count in the first column, taken away from one.

Count of testsChance that at least one fires by accidentWhat that test added
15.00 per cent5.00 points
29.75 per cent4.75 points
314.26 per cent4.51 points
522.62 per cent4.07 points
1040.13 per cent3.15 points
1348.67 per cent2.70 points
1451.23 per cent2.57 points
2064.15 per cent1.89 points
4590.06 per cent0.52 points
10099.41 per cent0.03 points

Every extra test adds less than the test before it did. The climb looks gentle at the start and is not. The first test hands over 5.00 points. The thirteenth hands over 2.70. The hundredth hands over three hundredths of a point. A glance at the right hand column alone suggests that the count of tests stops mattering, and that has it backwards: the column shrinks precisely because the figure it is being added to is running out of room underneath one hundred.

EACH TEST ADDS LESS THAN THE TEST BEFORE ITThirteen tests, one bar each. The bright cap is what that single extra test added to the running figure.01020304050per cent5.001+5.009.752+4.7514.263+4.5118.554+4.2922.625+4.0726.49630.17733.66836.98940.131043.121145.961248.6713count of teststhe thirteenth test adds 2.70 points where the first added 5.00
Each extra test adds less to the running figure than the test before it, which is why the climb from one test to thirteen looks gentle and still ends at 48.67 per cent.
Try it out

Ten tests read 40.13 per cent and twenty read 64.15 per cent. The count of tests doubled. Did the figure?

Drawn as a curve from one test to a hundred, the shape is easier to hold on to than the table is. The curve leaves the floor quickly, bends over somewhere in the teens, and then spends the rest of its length creeping along under the top of the frame without ever arriving there.

THE LADDER FROM ONE TEST TO A HUNDREDThe chance of at least one false positive, plotted at every whole count of tests.0255075100per cent120406080100count of testsone hundred per cent, and the curve does not touch itone half13 tests read 48.67 per cent14 tests read 51.23 per cent20 tests: 64.15 per cent100 tests: 99.41 per cent
The chance rises at every one of its ninety nine steps and every step is smaller than the one before it, crossing one half between the thirteenth test and the fourteenth and reaching 99.41 per cent at a hundred.
Breaking Into Quants Bootcamp — Fin Maverick

At how many tests does at least one false positive become more likely than not?

Read the table again at the two rows nobody looks at. Thirteen tests read 48.67 per cent. Fourteen tests read 51.23 per cent. Somewhere between the thirteenth test and the fourteenth, the run stops being a place where a false positive would be unlucky and becomes a place where the absence of one would be.

At the fourteenth test, at least one false positive becomes more likely than not, and nothing whatsoever about any of the tests changed to make that happen. No threshold moved. No test got weaker. No record was edited. Not one of the fourteen is any less careful than the single test that one question on its own would have called for. The only quantity that moved between 48.67 and 51.23 is the count, and the count is the part of the work that usually goes unrecorded.

The heading above deserves an honest reading. Enough tests do take the chance of at least one false positive above one half, then above nine tenths, and then as close to certain as anybody would like. The arithmetic never arrives. Climbing towards one hundred per cent and reaching it are two different claims, and the heading sounds like the second.

Try it out

At what count of tests does at least one false positive become more likely than not, and what changed about the tests to make it happen?

Try it out

The control below drags the count of tests from one up to a hundred. Before it moves: does the figure rise in even steps?

Play with it

Drag the count of tests and watch where the curve stops climbing

One control, one number: how many tests were run. Every test is held to the same threshold of 5 per cent and none of them changes. The curve draws itself from the first test up to the count the control is set to, the dashed line across the middle is one half and never moves, and the solid line at the top is one hundred per cent and never moves either. The short red bar on the right is what is still left between the curve and that top line. From a start at twenty, then down to one and back up to a hundred, watch which of the two lines the curve gets close to.

DRAG THE COUNT OF TESTS AND WATCH THE CURVE BUILDOne control. The dashed one half line and the solid ceiling never move.0255075100per cent120406080100count of testsone halfone hundred per centThe threshold each test is held to never moves. Only the count of tests does.
20 tests
Tests run
20
At least one fires, per cent
64.15
The last test added, points
1.89
Short of one hundred, points
35.85

Educational illustration. Every test on this panel is held to a threshold of 5 per cent, and the tests are assumed not to lean on one another. The assumption carries the whole curve, and tests run on a single record break it. The count includes tests that were run and then dropped.

Try it out

The heading above uses a strong word about what enough hypotheses do. Before the next block answers it: does the arithmetic ever actually reach one hundred per cent?

Does the figure ever actually reach one?

Forty five tests take it past nine tenths, at 90.06 per cent. Ninety tests take it past ninety nine hundredths, at 99.01 per cent. A hundred tests read 99.41 per cent. Keep going and the nines keep arriving: two hundred tests read 99.9965 per cent, three hundred read 99.99998, five hundred read 99.9999999993. Check every whole count out to five hundred and the figure is below one hundred per cent at every single one of them.

The arithmetic forces that result. The figure is one minus 0.95 raised to a count, and 0.95 raised to any whole count is a positive number, small but never nought. Something positive is always being taken away from one hundred, so something is always left. Enough tests make at least one false positive as close to certain as anyone cares to name, and no count of tests makes it certain, and those are two different sentences.

The shape is familiar from ordinary life. A walker covers half the remaining distance to a wall, then half of what is left, then half of that. The gap closes until it stops mattering to anybody watching, and the wall is never touched. The arithmetic here is doing the same thing with a different fraction.

ZOOM IN THREE TIMES AND THE GAP IS STILL THEREEach strip magnifies the right hand end of the strip above it a hundred times. The ceiling never gets touched.the whole scale, nought to one hundred per cent010013 tests45 teststhe last hundredth of that strip, magnified a hundred times99.0010090 tests100 tests150 testsand the last hundredth of THAT strip, magnified again99.99100200 tests300 and 500 tests both land inside this last stroke, and neither sits on the linestill short of one hundred by, in points:45 tests9.9440100 tests0.5921200 tests0.0035300 tests0.0000208500 tests0.0000000007
The figure sits below one hundred per cent at every whole count checked out to five hundred, so enough tests make at least one false positive as close to certain as anyone likes and never certain.

Why does the distinction matter? Repeat a strong word as though the arithmetic supplied it and the one quantity the arithmetic settles has been overstated. The correct reading is the practical one and it is not weaker: at fourteen tests a false positive somewhere in the run is more likely than not, and at forty five it is a nine in ten proposition. Nobody needs certainty to act on that. The count is what they need, and the count is what nobody writes down.

AI For Finance Bootcamp — Fin Maverick

How strict does each test have to be to hold the whole run at 5 per cent?

Turn the question round. Instead of fixing the threshold each test is held to and asking what the whole run comes to, fix the whole run at 5 per cent and ask what each test is then allowed. The equation is the same one with the other unknown solved for, and solving it directly beats reaching for somebody else's rule.

The same equation, solved the other way

Forwards. Each test is allowed a threshold. The chance that a single test comes back clean is one minus that threshold. All the tests come back clean when that clean chance is multiplied by itself once per test, and the chance that at least one fires is one minus the result.

Backwards. Hold the whole run at 5 per cent. Then all the tests coming back clean has to be 95 per cent. So the clean chance of one single test, multiplied by itself once per test, must equal 0.95. The clean chance of one test is therefore the appropriate root of 0.95. The threshold that test is allowed is one minus that root.

In plain words. With thirteen tests, take the thirteenth root of 0.95 and subtract it from one. The result reads 0.3938 per cent. With twenty tests, the twentieth root gives 0.2561 per cent. With 4,096 tests it reads 0.0013 per cent.

Two things are worth noticing about those three figures. The first is that they are not 5 per cent divided by the count, though they land close to it. Five divided by thirteen is 0.3846, and the correct figure is 0.3938. Five divided by twenty is 0.2500, and the correct figure is 0.2561. The gap is small at these counts and it is real, and it exists because multiplying thirteen clean chances together is not the same operation as sharing out a threshold.

The second is the sharpness of the fall: going from one test to twenty divides the threshold each single test is allowed by 19.52, so a search of twenty needs every one of its twenty tests to clear a bar roughly a twentieth of the height of the bar a single test faced. Think of a household budgeting for a wedding. The total does not move, so every additional guest makes every other guest's share smaller, and past a certain count each share is too small to buy anything with. A search behaves the same way, and the total being shared out is the credibility of the whole run.

WHAT IT COSTS EACH TEST TO HOLD THE WHOLE RUN AT 5 PER CENTThe bar is the threshold one single test is allowed, once the count of tests is fixed.on a scale running to 5.00 per cent1 test5.0000 per cent13 tests0.3938 per cent20 tests0.2561 per cent4,096 tests0.0013 per centthe last three, magnified, on a scale running to 0.45 per cent13 tests0.3938 per cent20 tests0.2561 per cent4,096 tests0.0013 per centGoing from one test to twenty divides the per test bar by 19.52, and the bar for 4,096 tests is a hairlineat both scales, which is the honest picture of it.
Holding a whole run at 5 per cent takes the per test threshold from 5.0000 per cent at one test to 0.2561 per cent at twenty, which is nearly a twentieth of it, and to a hairline at 4,096.

The cost of that strictness reads most clearly in months. Score the Ashwin rule across the earlier half of the six year record, the stretch holding 29 scoreable months. A reading has to reach 20 of those 29 calls before it clears 5 per cent on its own. Hold the whole run of thirteen settings at 5 per cent instead, so each setting faces 0.3938 per cent, and the bar moves up to 23 of 29. Three more months, and three months is what the count of tests costs once the cost leaves chances and enters the only currency this record has.

Try it out

Twenty independent tests each need a threshold of 0.2561 per cent to hold the whole run at 5 per cent. A colleague proposes running forty tests instead. Roughly what happens to the per test bar, and why is roughly the honest word?

Rebalancing: When, Why and What It Costs — free micro-course from Fin Maverick

What does the arithmetic assume, and what does this record do to it?

Every figure above rests on one assumption, stated once at the top and easy to walk past: the tests do not lean on one another. Tests run on a single record always lean on one another. The tests read the same months, in the same order, with the same six years of the same invented unit underneath them. So the ladder is not wrong. The ladder describes a run of tests nobody on a single record ever has.

Two real searches on the six year record fall on opposite sides of what the ladder would suggest. The first takes one fixed question and answers it twenty defensible ways, differing only in the stretchA handful of months that sit next to each other in the record, pulled out and examined without the rest. of the record used, the conventionA rule about how to count something, agreed in advance so that anybody repeating the work counts it the same way. applied to a month that did not move, and the signalWhatever number a rule looks at just before it commits to a call, built only out of months that have already closed. the rule reads. Twelve of those twenty clear the 5 per cent threshold. The second searches 4,096 calendar rules on a version of the record with the calendar pattern taken out of it. Nothing is left in that version to find. Ninety four of the 4,096 clear the threshold, on a denominatorThe count that sits under a share and fixes what the share is measuring. Move it and the share moves too, even when the number on top holds still. of 62 months.

Independence at 5 per cent would suggest one test in twenty. Line the two counts up against that. Twelve out of twenty is sixty per cent, far above one in twenty. Ninety four out of 4,096 is 2.29 per cent, far below it. The ladder governs neither count, and the direction of the error is not predictable from the count of tests, so twenty tests reading 64.15 per cent is an anchor for how much a count of tests matters and is not a figure that may be attached to either of those two tables.

WHAT ACTUALLY CLEARED, AGAINST WHAT ONE TEST IN TWENTY WOULD SUGGESTTwo searches run on the six year record. Both sets of tests read one record, so neither is independent.twenty analyses of oneresult12 of 20 cleared, 60.00 per centone test in twenty would be 5.00 per cent, or 1 of them4,096 calendar rules onthe adjusted record94 of 4,096 cleared, 2.29 per centone test in twenty would be 5.00 per cent, or 204.8 of themshare of the tests run that cleared the 5 per cent thresholdThe first search cleared far more than independence would suggest and the second far less.The ladder bounds neither, because in both searches the tests read the same record. It is ananchor for how much a count of tests matters, and it is not a figure to attach to eitherrow.
Twelve of twenty related analyses clear the threshold and ninety four of 4,096 related calendar rules do, one far above what independence would suggest and one far below it.

So what is the ladder still for? The ladder answers one question well, and the question is worth asking: does the count of tests matter enough to bother about? At thirteen tests the answer is that a run has roughly a coin flip's chance of throwing up at least one accident even if everything in it is empty, and no amount of relatedness between the tests makes that stop being worth knowing. The ladder cannot give the figure for a particular search on a particular record. Only the search itself gives that.

Try it out

On the six year record, twelve of twenty related analyses clear 5 per cent while ninety four of 4,096 related calendar rules do. Why can the ladder not be used on either, and what is it still good for?

Cleaning Financial Data teaches you to find the errors that survive every check and break every model.

What is asked when somebody reports one result?

The arithmetic is worth carrying only if it turns into something that can be said out loud. Somebody sends over one figure: a reading, a rule, a finding, one line long. There is a minute available, and nobody is about to redo anybody's arithmetic. Five questions do the work, ordered from the cheapest to ask to the most awkward.

The questionWhat it protects against
How many tests were run in total, including the ones that were dropped?The count is the only quantity the ladder needs, and a dropped test counts exactly as much as a reported one. A run of one and the best of a run of thirteen print identically.
Were the tests independent of one another, and if not, how do they lean?On one record they always lean. Once they do, the ladder bounds nothing, and the answer can land well above or well below what independence would suggest.
Which threshold was each test held to?Without it there is no 0.95 to raise to anything, and half the arguments about a reported result are really arguments about this number.
Was that threshold fixed before the count of tests was known?A threshold chosen after the search finished is a description of the winner rather than a test of it.
Is there a stretch of the record that was held back and never looked at?A held back stretch settles the matter and needs none of the arithmetic above.

A person at a lender asks the first two out of habit rather than distrust. Write a rule down before looking, or pick one out of thirteen after looking, and what lands on the desk reads the same either way. A count is the only thing that tells those two apart. An analyst on the other side of the table gets the same value out of volunteering the count unasked. A figure that travels with its count can be weighed by somebody who does not trust the analyst, and a figure that does not cannot be weighed even by somebody who does.

Everyone has done the household version of this. Ten shops are rung for a quote on the same repair, and the cheapest is reported to whoever is paying. The cheapest of ten is a real number and nobody made it up. The cheapest of ten is also not the price of the repair, and the only quantity separating the two readings is that ten shops were rung. The ladder is an argument about how hard somebody looked, and a stretch of the record held back before anything was chosen is a measurement, so where the second is available the first stops being the interesting conversation.

TWO ROUTES OUT OF THE SAME QUESTION, AND ONE OF THEM IS FOUR STEPS SHORTEREverything in this guide is the long route. The short one is a measurement.Was a stretch of the record held back?yesnoRead the rule on it once andreport what it called. One step,and it is a measurement ratherthan an argument.Count every test that was run, includingthe ones dropped.1Ask whether the tests lean on oneanother. On one record they do.2Work out the threshold each test wouldneed for the run to hold at 5 per cent.3Recheck the reported reading againstthat stricter bar.4Accept that the answer is still a boundand not a measurement.5done
The ladder is an argument about how hard somebody looked and a held back stretch is a measurement, so the second answers the question that the first can only put a bound on.
Try it out

Somebody answers a question about multiple testing by quoting a figure off this ladder. Which pair should be checked first, and what settles the matter without the ladder at all?

ONE FIGURE, QUOTED AS AN ALL CLEAR, WITH THREE THINGS WRONG WITH ITThe arithmetic in the reply is correct. What it is being asked to do is not.THE REPLY48.67per centthirteen tests, so the result isstill more likely than not to berealwrong question48.67 per cent is the chance that at least oneof thirteen tests fires when there is nothingthere. It is not the chance that this onereported reading is false.the tests are not independentThe thirteen settings read the same 29 monthsone step apart, so the ladder does not governthem at all.this was the best of the thirteenThe reported reading won a search. The best of aset is exactly the object the arithmetic waswarning about.The fix needs none of thisarithmetic. Read the rule on astretch that was held back beforeanything was chosen, and count whatit calls there.
How often at least one of thirteen tests fires by accident is a different quantity from how likely the one reported reading is to be false, and the reported one had already won a search.

The failure: quoting 48.67 per cent as though it were an all clear

An analyst walks the thirteen settings of the Ashwin rule across the earlier half of the six year record and reports whichever of them read best. Somebody asks the right question: what does running thirteen tests do to the result? The analyst goes to the ladder, reads that thirteen tests give 48.67 per cent, and answers that since 48.67 is under one half the result is still more likely than not to be real. The arithmetic in the reply is correct. Three things about it are wrong, and only one of them is the obvious one.

First, it answers a different question from the one asked. 48.67 per cent is the chance that at least one of thirteen tests fires when there is nothing at all in the record. The figure is not the chance that this one reported reading is false. The two quantities are different and not interchangeable, however similar the sentences describing them sound.

Second, the thirteen tests are not independent. The thirteen settings read the same 29 scoreable months, one whole number step apart on the same grid. The ladder is built on tests that do not lean on one another, and these lean on one another about as hard as tests can. The figure does not apply to them at all, in either direction.

Third, and worst, the reported reading was the best of the thirteen. The best of a set is exactly the object the arithmetic was warning about. Using a figure derived from the count of tests to reassure yourself about the winner of that same count of tests is using the alarm as evidence that nothing is burning.

The cost is not that a number was misused in a meeting. The cost is that a figure whose only job was to raise a question has been used to close it, and everyone in the room now believes the question has been dealt with. The fix needs none of this arithmetic: the rule is read on a stretch of the record that was held back before any setting was chosen, and what it called there is reported. The held back stretch answers the question the ladder can only bound.

Covered elsewhere. The claim a threshold of 5 per cent makes about a single test is settled earlier in these notes, and so is independence between two things. Holding a search down in practice is settled elsewhere in these notes. Choosing among several defensible ways of analysing one result, and choosing the question itself after looking at the record, are covered on their own, and they are two different faults rather than one. Placing an order, paying a cost, holding a position and judging whether any tested rule deserves anybody's attention all sit elsewhere.

Debt Capital Markets Bootcamp — Fin Maverick

Where do the figures above come from?

Every quantity above comes out of three things: a threshold of 5 per cent, a count of tests, and one invented record built earlier in these notes. Nothing else goes into any of them.

Quantity usedWhere it sitsHow to rebuild it
The ladder, from one test to five hundredComputed here from a threshold of 5 per cent and a count, and from nothing else at allRaise 0.95 to the count and take it away from one. Twenty tests hand back 64.15 per cent.
The threshold one single test is allowedThe same equation with the other unknown solved for, derived above rather than borrowedTake the thirteenth root of 0.95 and subtract it from one. The answer reads 0.3938 per cent.
The six year record of the Nakshatra unit72 invented monthly observations, settled earlier in these notes and reused here without one figure being alteredRebuild it and count: 65 of the 72 months can be scored, 29 of them in the first three years.
The Ashwin rule and its thirteen settingsInvented, and written down in full before any count here was takenScore the rule setting by setting on the scoreable months and the counts follow.
The two searches on this recordRun in full by the arithmetic check kept beside these notes, and that check refuses to finish if a count movesTwelve of twenty analyses clear the threshold, and ninety four of 4,096 calendar rules do.

The Nakshatra unit, the six year record and the Ashwin rule are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.