Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quantitative Methods, Financial Data & Programming
1Probability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
2Statistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
3Correlation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
4Time Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
5Simulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
6Optimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
7Modelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
8Backtesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
9Data Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
10Programming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
11Quantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity

Statistical Bias: The Types and Where Each Creeps In

Bias is an error that points the same way every time, so averaging more cases never removes it. Selection bias enters when cases are chosen. Survivorship bias enters when cases leave. Look ahead bias enters when information from later is used earlier. Measurement bias enters in the arithmetic itself. Drop the five worst months from the fifty month record and its mean climbs from 0.50 to 1.56 per cent, past a true 1.00.

Reading a record for bias takes two things and nothing else: the ability to average a list of numbers, and the willingness to ask where the list came from. The object being measured here is the Nakshatra unit, an invented traded unitSomething whose price is quoted often enough that what it did in each period can be written down. Used purely as something to have readings about. whose monthly changeThe percentage by which a price moved over one month. One reading per month, and the only kind of reading any of the four biases needs. takes one of five values. The five values and the five weights beside them were written down before a single month was drawn, so the true average monthly change of the Nakshatra unit is 1.00 per cent and its true spread is 5.00 per cent. Not estimated. Stated, the way the number of chairs in a room is stated.

A known truth is a strange luxury. In almost every record handed to an analyst the truth is the missing quantity, so arguments about a one way error can run for years without anyone being able to settle them. Here the truth is printed above. So every kind of one way error named below can be let into a record on purpose and then measured, to two decimal places, against a number already in view. A description of bias that never measures one leaves the thing itself unshown.

What makes an error bias rather than simply being wrong?

Picture two weighing scales at a grocery shop. The first reads fifty grams heavy on every packet it ever weighs. The second wobbles, reading fifty grams heavy on one packet and fifty grams light on the next, with no pattern to it. Both scales are wrong by about the same amount on any single packet. Now weigh a hundred packets on each and average the readings. The wobbly scale's errors cancel and the average lands close to the truth. The heavy scale's average lands fifty grams heavy. A hundred readings agree with each other, so it lands there with more confidence than before.

Bias is not a bigger error than ordinary error, it is an error with a direction, and a direction is precisely the property that averaging cannot remove. A direction is therefore a different fault from a wobble, and it needs a different remedy. Every estimatorAny rule that turns a set of readings into a single answer. Adding up fifty monthly changes and dividing by fifty is an estimator. So is a rule that ignores half of them. run on a sample is wrong by something. The useful test is whether running an estimator again and again would eventually land on the right answer, or circle a wrong one forever.

Here is the same idea with the Nakshatra unit, where the truth is available. Five separate records of fifty months each were drawn from the same generatorThe written down set of possible values and their weights that the readings come out of. Because it was written first here, it is knowable; in an ordinary record it never is.. Honestly averaged, their means come to 0.50, 1.30, 0.90, 0.40 and 1.60 per cent. Three sit below the true 1.00 and two sit above it, and the five of them average 0.94, a bare 0.06 below the truth. Now apply one rule to all five: throw out every month at the worst value, on the grounds that those months were unusual. The five means become 1.56, 1.96, 1.76, 1.68 and 2.04 per cent. Every single one is above the truth, and the five now average 1.80, a full 0.80 above it.

ERRORS THAT SCATTER, AND ERRORS THAT POINT Ten records of the same unit. The same truth behind all ten. TRUE MEAN 1.00 PER CENT FIVE RECORDS AVERAGED HONESTLY THEIR AVERAGE 0.94 0.50, 1.30, 0.90, 0.40 and 1.60 per cent. Three fall short of the truth, two overshoot it. THE SAME FIVE WITH THEIR WORST MONTHS DROPPED THEIR AVERAGE 1.80 1.56, 1.96, 1.76, 1.68 and 2.04 per cent. Not one of the five falls short. All five overshoot. 0.00 0.50 1.00 1.50 2.00 Monthly change, per cent. The truth is known here only because the five values and their weights were written first.
Averaged honestly the five records straddle the true 1.00 per cent and land within 0.06 of it; cleaned by one shared rule they all sit above it and land 0.80 away, which is the difference between an error that scatters and an error that points.
Try it out

In one sentence, what makes bias different from ordinary error?

Try it out

A method is run on twenty times the data and gives a tighter range around the same wrong answer. What has improved and what has not?

Where does selection bias get into a record?

Selection bias enters at the moment cases are chosen for the record, before anybody has written down a single number. The timing is what makes selection bias so easy to miss: by the time the file reaches the analyst, the choosing is invisible, and everything inside the file is correct.

A tea stall outside one office building asks its regulars whether they liked the new blend, and four in five say yes. The four in five is a true count, and it is a true count of the people who came back. Everybody who tried the new blend once and never returned is not in the room to be asked, and they are not a small correction to the answer, they are the answer. Selection bias does not require anybody to cheat, only for somebody to decide which cases are worth writing down.

Now the same thing with numbers that can be scored. Suppose a record of the Nakshatra unit is assembled not from all fifty months, but from the months somebody thought were worth noting at the time. The dramatic months get written down and most of the flat ones do not: all five of the worst months and all three of the best survive, but only five of the twenty five ordinary months at 1.00 per cent make it in. The result is a thirty month record. Its tallyThe count of how many times each possible value turned up. A tally holds everything a mean or a spread needs, without keeping the months in order. is 5, 9, 5, 8 and 3, its total is 5.00, and its mean is 0.17 per cent against a true 1.00. The selected record is wrong by 0.83 per cent, and it did not lose a single reading to error. The selected record lost twenty perfectly ordinary months to a judgement about what was interesting.

TWO RECORDS OF THE SAME FIFTY MONTHS Every reading in both records is correct. One of them was chosen. EVERY MONTH WRITTEN DOWN, 50 MONTHS 5 9 25 8 3 minus 9.00 minus 4.00 1.00 6.00 11.00 MONTHLY CHANGE, PER CENT RECORD MEAN 0.50 per cent, 50 months ONLY THE MONTHS SOMEBODY THOUGHT WORTH WRITING DOWN, 30 MONTHS 5 9 5 8 3 20 ORDINARY MONTHS NEVER WRITTEN DOWN minus 9.00 minus 4.00 1.00 6.00 11.00 RECORD MEAN 0.17 per cent, 30 months Nobody altered a reading. Twenty ordinary months were simply never thought worth the ink.
Choosing which months are interesting enough to record drops the mean of the Nakshatra unit from 0.50 per cent to 0.17 against a true 1.00, without any reading in the shorter record being wrong.

What does survivorship bias do to a mean that can be checked?

Survivorship bias enters at the opposite end of a record's life from selection: not when cases are chosen, but when cases leave. Something was in the record, and then it was not, and the reason it left was related to how bad it was. Because survivorship happens one row at a time, it is the easiest of the four to watch.

Take the fifty month record exactly as it stands, tallied 5, 9, 25, 8 and 3 against its five possible values in order from worst to best. Value times count adds to 25.00 across fifty months, so the mean is 0.50 per cent. Now remove the worst months one at a time, on the entirely reasonable sounding ground that they were unusual and should not distort the picture. Each removal takes one month at minus 9.00 out of the record. A total that had 9.00 subtracted from it no longer does, so the count of months drops by one and the total rises by nine.

So after removing a number of months, the mean is 25 plus 9 times that number, over 50 minus that number. Work the six settings and the picture is not subtle. The true mean is 1.00 per cent, the honest estimate was 0.50 per cent, and dropping five months out of fifty carries the answer from half the truth to half again above it, without one figure in the record being falsified.

THE RECORD'S MEAN AS THE WORST MONTHS COME OUT Six settings of one rule. Nothing else about the record changes. BETWEEN 2 AND 3 MONTHS REMOVED IT PASSES THE TRUTH TRUE MEAN 1.00 PER CENT 0.50 0.69 0.90 1.11 1.33 1.56 0 removed 1 removed 2 removed 3 removed 4 removed 5 removed 50 months 49 months 48 months 47 months 46 months 45 months Each mean is 25 plus 9 times the number removed, divided by 50 minus the number removed. Per cent.
Removing the worst months one at a time lifts the record's mean from 0.50 per cent through 0.69, 0.90, 1.11 and 1.33 to 1.56, crossing the true 1.00 per cent somewhere between the second and third month removed.
Try it out

Drop the five worst months and the record's mean goes from 0.50 to 1.56 per cent. Where is the truth in relation to those two figures?

What does the whole build look like, line by line?

Every setting below is worked from the tally rather than quoted, so each row can be checked with a pen. The count of months is 50 minus the number removed. Each removed month was carrying a minus 9.00 that is no longer in the sum, so the total is 25.00 plus 9.00 for each month removed. The mean is the second column divided by the first.

Worst months removedTally that remainsMonthsTotal of value times countRecord mean, per centAgainst the true 1.00
None5, 9, 25, 8, 35025.000.500.50 below
One4, 9, 25, 8, 34934.000.690.31 below
Two3, 9, 25, 8, 34843.000.900.10 below
Three2, 9, 25, 8, 34752.001.110.11 above
Four1, 9, 25, 8, 34661.001.330.33 above
Five0, 9, 25, 8, 34570.001.560.56 above

The crossing point is not a matter of opinion either. Setting 25 plus 9 times the number removed equal to 50 minus that number gives ten times the number equal to twenty five, so the record's mean equals the true 1.00 per cent at exactly two and a half months removed. Since half a month cannot be removed, the third removal is the one that carries the record from understating the truth to overstating it. Three months out of fifty is six per cent of the record, and the shaded rows above are the ones on the wrong side.

The other thing worth noticing about the cleaned record is what it looks like from the inside. The five most disagreeable months are gone, so the cleaned record's months agree with each other far better than the honest record's months did. Anybody who judges a record by how tidy it is will prefer the wrong one.

THE CLEANING LOG, AND THE SAME RECORD WITHOUT ONE Five lines removed. Not one reading altered. CLEANING LOG, NAKSHATRA UNIT, 50 MONTHS LINE READING REASON GIVEN STATUS 07 minus 9.00 per cent unusual, set aside REMOVED 13 minus 9.00 per cent unusual, set aside REMOVED 21 minus 9.00 per cent unusual, set aside REMOVED 34 minus 9.00 per cent unusual, set aside REMOVED 48 minus 9.00 per cent unusual, set aside REMOVED TOTAL REMOVED, WHICH IS THE ONE NUMBER THAT CAN BE AUDITED 5 lines of 50 WHAT THE LOG BUYS With the count in view, the mean can be reported both ways: 0.50 per cent across all fifty months, and 1.56 per cent across the forty five that were kept. WITH NO LOG Nothing was falsified and nothing is missing from view, so no total finds the five. The cleaning rule shown here is invented for teaching. No record kept by anybody is being described.
A cleaning log turns five removals into a number somebody can audit, and a cleaning rule with no count beside it cannot be checked by anyone, including the person who wrote the rule.
Try it out

How many of the fifty months have to be removed before the record's mean passes the true 1.00 per cent?

Play with it

Take out the worst months, one at a time, and watch the answer walk past the truth

Fifty months, one square each, grouped by what the month did. The control removes the worst months and nothing else: no reading is edited, no month is added, and the true mean of 1.00 per cent stays pinned exactly where it is at every setting. The second control switches between five separate records of the same unit. The walk then shows up as either a quirk of one record or a property of the rule. Start at the default, the published fifty month record with nothing removed and a mean reading 0.50 per cent to the decimal, then drag right.

PINNED AND NOT ADJUSTABLE: the true mean of the Nakshatra unit is 1.00 per cent and its true spread is 5.00 per cent. Both come from the five values and their weights, written down before any month was drawn, so no control on this panel can move either one.
Jump to a setting:
ONE SQUARE PER MONTH, GROUPED BY WHAT THE MONTH DID Removed months are drawn hollow and struck through. minus 9.00 minus 4.00 1.00 6.00 11.00 MONTHLY CHANGE, PER CENT TRUE MEAN 1.00, PINNED 0.00 0.50 1.50 2.00 0.50 per cent
Loading the panel.
The five counts
5 + 9 + 25 + 8 + 3
Months in the record
50
Value times count, added
25.00
The record's mean
0.50 per cent
Distance from the true 1.00
0.50 per cent below
Where this record crosses
between 2 and 3 removed
Educational illustration. The Nakshatra unit and its five fifty month records are teaching objects, not readings taken from any market. Only the worst months are removed and nothing else about a record changes at any setting. The five values and their weights were written down before any month was drawn, and that alone is why the true mean of 1.00 per cent is knowable. No record handed to an analyst arrives with its own truth attached. Once every worst month a record holds has gone there is nothing further to remove, and the control stops. A rule that pushes a figure one way is a fact about the rule, not a reason to act on the figure.

Two things are worth watching as the control moves. The months being removed are the ones furthest from everything else, so the count of months falls very slowly while the mean moves very fast. The second is that switching records changes where the crossing sits but not the direction of travel. Record four lands exactly on the true 1.00 per cent at three months removed and goes past it at four. Record five is already above the truth before anything is removed at all. On all five records every removal pushes the answer the same way. A rule that only ever removes bad cases can only ever push an average up, and a push is a direction rather than a wobble.

Why is look ahead bias the hardest of the four to see?

Look ahead bias enters when information that only became available later is used as though it had been available at the time. Nothing leaves the record and nothing joins it. Every figure in it is correct. The rows all add up.

Look ahead bias is the hardest of the four to find for exactly that reason: nothing is wrong with the record as a set of numbers, so no count, no total and no cross check on the record itself will reveal it. Only a question about timing will, and timing is not a property of any column. The question to put to each figure is when it became knowable.

Take it home first. A household sits down to score last year's monthly budget and works out that it managed groceries on Rs 9,000/- a month. Today's prices are the ones the household has in its head, so it does the sums with those. Today's prices are higher, so last year's grocery basket now looks like it cost more than the household actually paid, and the household concludes that it was better at shopping than it was. Every figure it used is a real price. None of them was a price it faced. The mistake is not in the arithmetic, it is in the calendar.

The same shape appears in a record of monthly changes. Suppose somebody scores each month of the Nakshatra unit as good or bad, and does the sorting using the full fifty month record in front of them. The scorer now knows which months turned out badly. If any part of the scoring leans on that knowledge, even by deciding which months were worth a second look, then the score is not something anybody could have produced at the time. The record has not been damaged; the claim made about it has.

Try it out

A record is complete, every figure in it is correct, and it is still biased. Which of the four is that, and what question finds it?

Breaking Into Quants Bootcamp — Fin Maverick

What has a denominator got to do with any of this?

The fourth kind hides in the arithmetic rather than in the data, and so survives every check aimed at the rows. Working out a spread from a sample means adding up the squared gaps between each reading and the sample's own mean, then dividing. The choice of what to divide by is the whole story. Dividing by the count of readings gives one answer; dividing by the count less one gives a slightly larger one.

On the fifty month record, the squared gaps add to 1,212.50. Dividing by 50 and taking the square root gives 4.92 per cent. Dividing by 49 and taking the square root gives 4.97 per cent. Both figures sit under the true 5.00 on this particular record, and that part is luck. The relationship between them is not luck. On every record, without exception, the count denominatorThe divisor. Either the count of readings, or that count less one, and picking between the two settles the answer. gives a smaller answer than the corrected one. The two differ by a fixed multiplier that never changes sign.

Run it on all five records and the multiplier is identical to five decimal places each time. An identical multiplier every time is the signature of a one way error. A gap that is the same size and the same direction in every record examined is not coming from the sample at all, it is coming from the formula, and a bigger sample cannot average away a formula. The size of it depends only on how many readings there are: with fifty readings the count denominator returns about 0.99 of the corrected figure, and with five readings it returns about 0.89.

With five readings the same arithmetic gets loud. A five month stretch of the same unit reading minus 4.00, minus 4.00, 1.00, 6.00 and 6.00 per cent has a mean of 1.00 and squared gaps adding to 100.00. Dividing by four and taking the root gives 5.00 per cent; dividing by five and taking the root gives 4.47. The shortfall is 0.53 rather than 0.05, ten times as wide on a stretch a tenth as long. The corrected figure landing exactly on the true 5.00 here is luck and nothing more, exactly as its landing on 4.97 on the fifty month record was luck. The part that is not luck is which of the two figures is the smaller one, and it is the same one both times.

RecordSquared gaps addedSpread on the count denominatorSpread on the corrected denominatorCount over corrected
Record one, the fifty month record1,212.504.924.970.98995
Record two1,170.504.844.890.98995
Record three1,274.505.055.100.98995
Record four1,282.005.065.120.98995
Record five1,032.004.544.590.98995

Look at the last column. Five different records, five different spreads, five different distances from the true 5.00 per cent, and one multiplier that does not budge. Two of the five records put both figures above the truth and three put both below. The scatter is the wobble. The last column is the push.

THE COUNT DENOMINATOR AGAINST THE CORRECTED ONE The same scale in both panels. Only the number of readings changes. PANEL A, FIFTY READINGS 0.05 4.30 4.50 4.70 4.90 5.10 TRUE SPREAD 5.00 COUNT DENOMINATOR, 4.92 CORRECTED DENOMINATOR, 4.97 PANEL B, FIVE READINGS OF THE SAME UNIT 0.53 4.30 4.50 4.70 4.90 5.10 TRUE SPREAD 5.00 COUNT DENOMINATOR, 4.47 CORRECTED DENOMINATOR, 5.00 The count denominator returns about 0.99 of the corrected figure on fifty readings and about 0.89 on five. The same direction both times. That the corrected figure lands on 5.00 in Panel B is luck, not a rule.
The count denominator falls short of the corrected one by 0.05 on fifty readings and by 0.53 on five, always in the same direction, which makes the shortfall a property of the formula rather than of the readings.
Try it out

The count denominator gives 4.92 per cent and the corrected one gives 4.97, against a true 5.00. Is that gap noise or bias?

AI For Finance Bootcamp — Fin Maverick

Why does adding more data not fix any of it?

The instinct to gather more data is a good one, and a good instinct is exactly what makes this trap catch careful people. More data really does help with the other problem. If an estimate wanders because fifty months is a small window, then two hundred months narrows the wander, and the figure that summarises how far a single estimate typically lands from what it is estimating, the standard errorA number describing how far one estimate typically falls from the value it is estimating. How it is built, and how it shrinks as the window grows, is covered separately., gets smaller as the window grows. The narrowing is real and it is worth having.

But bias is not wander, and the arithmetic of more data does nothing to it. Go back to the cleaning rule that drops the worst months. If a hundred records are built and every one of them is cleaned the same way, then all hundred report high. Their average is high. And here is the sting: their spread is small. They agree with each other, and they agree because they were all pushed in the same direction by the same rule.

The one thing that feels most like confirmation, many separate records agreeing, is exactly what a shared method produces, so agreement between records is evidence about the method they have in common and not evidence about the truth. Watch it with the five records. Averaged honestly, their running average moves 0.50, 0.90, 0.90, 0.78, 0.94, settling near the true 1.00. Cleaned the same way, their running average moves 1.56, 1.76, 1.76, 1.74, 1.80, settling near 1.80. Both settle down. Only one settles on the truth.

WHAT ADDING RECORDS ACTUALLY DOES Both lines settle. Settling and being right are two different things. 0.00 0.50 1.00 1.50 2.00 TRUE MEAN 1.00 PER CENT CLEANED THE SAME WAY, SETTLES NEAR 1.80 AVERAGED HONESTLY, SETTLES NEAR THE TRUTH 1 record 2 3 4 5 records Each point is the average of every record to its left. Running average of the record mean, per cent.
Adding records makes both running averages steadier, but only the honestly averaged one steadies on the true 1.00 per cent while the cleaned one steadies near 1.80, so steadiness is not the same as being right.
Try it out

A hundred records built the same way all report the same high figure. What does the agreement between them establish?

Can one record carry all four at once?

One record can, and there is nothing unusual about it. The four are not four names for one thing, they are four different moments in the life of a record, and a record passes through all four moments on its way to becoming a figure somebody quotes.

Cases are chosen, and selection can enter. Readings are written down. Cases leave, and survivorship can enter. Readings get dated, or fail to, and look ahead can enter. The arithmetic runs, and measurement can enter. Then a figure is reported. By that stage everything that could go one way has already gone, so nothing further can get in. A record can carry all four at once, and because each one enters at a different moment, each one needs a different question to find it and no single check finds more than one.

WHERE EACH ONE GETS IN, ALONG A RECORD'S LIFE Six stages. Four of them let a one way error in, and each needs its own question. SELECTION BIAS SURVIVORSHIP BIAS LOOK AHEAD BIAS MEASUREMENT BIAS CASES ARE CHOSEN READINGS ARE WRITTEN DOWN CASES LEAVE THE RECORD READINGS ARE DATED THE ARITHMETIC RUNS THE FIGURE IS REPORTED ASK WHO PICKED THEM NOTHING ENTERS HERE ASK WHAT LEFT ASK WHEN IT WAS KNOWN ASK WHICH FORMULA NOTHING ENTERS HERE THREE OF THE FOUR ARE ALREADY IN THE RECORD BEFORE ANYBODY OPENS IT. The fourth enters after collection is over, inside the arithmetic, where no count of rows can reach it.
Selection enters when cases are picked, survivorship when they drop out, look ahead when the clock is ignored and measurement when the arithmetic runs, so one record can carry all four and each needs its own question.
Try it out

Which of the four enters after the collecting of data is finished, and where does it live instead?

Cleaning Financial Data — free micro-course from Fin Maverick

How is bias hunted in a record built by somebody else?

Most records that reach anybody arrive finished. A lender is handed a summary of a borrower's monthly receipts. An analyst is handed a spreadsheet of monthly figures somebody else assembled. A household is handed a statement of what a shop's takings have been. In none of those cases can the record be rebuilt, and in none of them is the truth known. Five questions can still be put to a finished record, one aimed at each moment in the timeline and one aimed at the answer itself.

Where did these cases come from, and who decided which ones went in. Which cases used to be in the record and are not in it now. Could anything here have been known only afterwards. Which formula was used, and on what denominator. And then the fifth question, the one that changes the conversation. If the answer is wrong, which direction is it wrong in? Asking which way rather than merely whether turns a vague worry into something checkable.

The fifth question is worth dwelling on. A person reading a record can always ask it and can usually answer it. A lender looking at a summary of monthly receipts does not need to know the true average to reason that a summary prepared by the borrower, from months the borrower chose, with the poor months set aside for review, can only be wrong in one direction. Reasoning about direction is not a calculation and does not need one. The reasoning is a statement about which way a rule pushes, and a rule that only ever removes bad cases pushes up. The same reasoning runs the other way when the person assembling the record had a reason to look cautious.

FIVE QUESTIONS TO PUT TO A RECORD BUILT BY SOMEBODY ELSE Four of them are aimed at a moment. The fifth is aimed at the answer. NO. THE QUESTION WHAT IT IS AIMED AT 1 Where did these cases come from, and who decided? SELECTION 2 What used to be in this record and is not in it now? SURVIVORSHIP 3 Could anything here have been known only afterwards? LOOK AHEAD 4 Which formula was used, and on what denominator? MEASUREMENT 5 If the answer is wrong, which direction is it wrong in? ALL FOUR AT ONCE The fifth question is the one that converts a worry into a check, because it names a direction rather than a doubt.
Ask where the cases came from, what left the record, what could only have been known afterwards, which formula ran, and which direction the error points, because each question is aimed at a different moment in the record's life.
Try it out

A cleaned record arrives with no cleaning log attached. Which single number should be asked for first?

The tidy record that everybody trusted more

An invented illustration. A set of monthly changes is cleaned before anyone analyses it, and the cleaning rule is a sensible sounding one: months that look unusual are set aside for review. Nobody ever reviews them and they never come back. Unusual and bad are the same months in this record, so five months of fifty go, all of them at minus 9.00 per cent.

The figure reported is 1.56 per cent against a true 1.00. The expensive part is not the 0.56 of error. The person reporting 1.56 is more confident than the person who reported 0.50, and has better looking evidence. The cleaned record's months agree with one another far better. Its spread is narrower. Its range is tighter. Every one of those improvements is a direct consequence of the fault, so the record that is further from the truth now looks like the more careful job, and there is nothing inside it that says otherwise.

No cleaning rule exists that cannot go one way, so the fix is not a better cleaning rule. The fix is a count. Keep a tally of everything removed and report the mean both ways whenever anything was dropped, 0.50 per cent across all fifty months and 1.56 across the forty five kept. A cleaning rule with no count beside it cannot be audited by anybody, and that includes the person who wrote the rule, who six months later will not remember what it took out.

The four moments in a record's life at which a one way error gets in are named above, and three of them are measured against a true mean of 1.00 per cent that was written down before any record existed. How tightly a method repeats itself, as against whether it lands on the truth, is built separately, where the same five records are used to make exactly that distinction. Where the wandering in an honest estimate comes from, and how many months are needed to shrink it, is covered separately, where the standard error named above is built rather than merely named. Deciding whether a gap between two figures is large enough to be worth acting on is also covered separately. A one way error inside a fitted relationship, or inside a test of a rule run over past readings, is covered separately and much later. Every point estimateA single number offered as the answer, with no range attached to it. Every figure compared above is one of these, and a missing range is part of why the differences between them are so easy to miss. compared above is a single number, and comparing single numbers is exactly where a one way error hides best.
Cleaning Financial Data teaches you to find the errors that survive every check and break every model.

What the figures are built on

What a figure rests onWhere it came fromHow to redo it
The five values of the Nakshatra unit and the five weights beside themWritten down before a single month was drawn, so they are stated rather than measuredMultiply each value by its weight and add the five products. They come to 1.00 per cent.
The fifty month record, tallied 5, 9, 25, 8 and 3Invented for teaching and then held fixed everywhere it appearsThe five counts add to 50, and value times count adds to 25.00, so the mean is 0.50 per cent.
The six means in the survivorship build, 0.50 through 1.56 per centRecomputed line by line from that tallyDivide 25 plus 9 times the number removed by 50 minus the number removed.
The two spread figures, 4.92 and 4.97 per centRecomputed from the same tally, both from a total of squared gaps of 1,212.50Divide 1,212.50 by 50 and by 49 in turn, then take the square root of each.
The named ways a one way error gets inCommon teaching property, described here by the mechanism that produces each oneNamed by the mechanism that produces each one rather than by attribution to a text.

The Nakshatra unit, the fifty month record, the four further records of fifty months, the five month stretch, the record of interesting months and the cleaning log are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.