Statistical Bias: The Types and Where Each Creeps In
Bias is an error that points the same way every time, so averaging more cases never removes it. Selection bias enters when cases are chosen. Survivorship bias enters when cases leave. Look ahead bias enters when information from later is used earlier. Measurement bias enters in the arithmetic itself. Drop the five worst months from the fifty month record and its mean climbs from 0.50 to 1.56 per cent, past a true 1.00.
Reading a record for bias takes two things and nothing else: the ability to average a list of numbers, and the willingness to ask where the list came from. The object being measured here is the Nakshatra unit, an invented traded unitSomething whose price is quoted often enough that what it did in each period can be written down. Used purely as something to have readings about. whose monthly changeThe percentage by which a price moved over one month. One reading per month, and the only kind of reading any of the four biases needs. takes one of five values. The five values and the five weights beside them were written down before a single month was drawn, so the true average monthly change of the Nakshatra unit is 1.00 per cent and its true spread is 5.00 per cent. Not estimated. Stated, the way the number of chairs in a room is stated.
A known truth is a strange luxury. In almost every record handed to an analyst the truth is the missing quantity, so arguments about a one way error can run for years without anyone being able to settle them. Here the truth is printed above. So every kind of one way error named below can be let into a record on purpose and then measured, to two decimal places, against a number already in view. A description of bias that never measures one leaves the thing itself unshown.
What makes an error bias rather than simply being wrong?
Picture two weighing scales at a grocery shop. The first reads fifty grams heavy on every packet it ever weighs. The second wobbles, reading fifty grams heavy on one packet and fifty grams light on the next, with no pattern to it. Both scales are wrong by about the same amount on any single packet. Now weigh a hundred packets on each and average the readings. The wobbly scale's errors cancel and the average lands close to the truth. The heavy scale's average lands fifty grams heavy. A hundred readings agree with each other, so it lands there with more confidence than before.
Bias is not a bigger error than ordinary error, it is an error with a direction, and a direction is precisely the property that averaging cannot remove. A direction is therefore a different fault from a wobble, and it needs a different remedy. Every estimatorAny rule that turns a set of readings into a single answer. Adding up fifty monthly changes and dividing by fifty is an estimator. So is a rule that ignores half of them. run on a sample is wrong by something. The useful test is whether running an estimator again and again would eventually land on the right answer, or circle a wrong one forever.
Here is the same idea with the Nakshatra unit, where the truth is available. Five separate records of fifty months each were drawn from the same generatorThe written down set of possible values and their weights that the readings come out of. Because it was written first here, it is knowable; in an ordinary record it never is.. Honestly averaged, their means come to 0.50, 1.30, 0.90, 0.40 and 1.60 per cent. Three sit below the true 1.00 and two sit above it, and the five of them average 0.94, a bare 0.06 below the truth. Now apply one rule to all five: throw out every month at the worst value, on the grounds that those months were unusual. The five means become 1.56, 1.96, 1.76, 1.68 and 2.04 per cent. Every single one is above the truth, and the five now average 1.80, a full 0.80 above it.
In one sentence, what makes bias different from ordinary error?
A method is run on twenty times the data and gives a tighter range around the same wrong answer. What has improved and what has not?
Where does selection bias get into a record?
Selection bias enters at the moment cases are chosen for the record, before anybody has written down a single number. The timing is what makes selection bias so easy to miss: by the time the file reaches the analyst, the choosing is invisible, and everything inside the file is correct.
A tea stall outside one office building asks its regulars whether they liked the new blend, and four in five say yes. The four in five is a true count, and it is a true count of the people who came back. Everybody who tried the new blend once and never returned is not in the room to be asked, and they are not a small correction to the answer, they are the answer. Selection bias does not require anybody to cheat, only for somebody to decide which cases are worth writing down.
Now the same thing with numbers that can be scored. Suppose a record of the Nakshatra unit is assembled not from all fifty months, but from the months somebody thought were worth noting at the time. The dramatic months get written down and most of the flat ones do not: all five of the worst months and all three of the best survive, but only five of the twenty five ordinary months at 1.00 per cent make it in. The result is a thirty month record. Its tallyThe count of how many times each possible value turned up. A tally holds everything a mean or a spread needs, without keeping the months in order. is 5, 9, 5, 8 and 3, its total is 5.00, and its mean is 0.17 per cent against a true 1.00. The selected record is wrong by 0.83 per cent, and it did not lose a single reading to error. The selected record lost twenty perfectly ordinary months to a judgement about what was interesting.
What does survivorship bias do to a mean that can be checked?
Survivorship bias enters at the opposite end of a record's life from selection: not when cases are chosen, but when cases leave. Something was in the record, and then it was not, and the reason it left was related to how bad it was. Because survivorship happens one row at a time, it is the easiest of the four to watch.
Take the fifty month record exactly as it stands, tallied 5, 9, 25, 8 and 3 against its five possible values in order from worst to best. Value times count adds to 25.00 across fifty months, so the mean is 0.50 per cent. Now remove the worst months one at a time, on the entirely reasonable sounding ground that they were unusual and should not distort the picture. Each removal takes one month at minus 9.00 out of the record. A total that had 9.00 subtracted from it no longer does, so the count of months drops by one and the total rises by nine.
So after removing a number of months, the mean is 25 plus 9 times that number, over 50 minus that number. Work the six settings and the picture is not subtle. The true mean is 1.00 per cent, the honest estimate was 0.50 per cent, and dropping five months out of fifty carries the answer from half the truth to half again above it, without one figure in the record being falsified.
Drop the five worst months and the record's mean goes from 0.50 to 1.56 per cent. Where is the truth in relation to those two figures?
What does the whole build look like, line by line?
Every setting below is worked from the tally rather than quoted, so each row can be checked with a pen. The count of months is 50 minus the number removed. Each removed month was carrying a minus 9.00 that is no longer in the sum, so the total is 25.00 plus 9.00 for each month removed. The mean is the second column divided by the first.
| Worst months removed | Tally that remains | Months | Total of value times count | Record mean, per cent | Against the true 1.00 |
|---|---|---|---|---|---|
| None | 5, 9, 25, 8, 3 | 50 | 25.00 | 0.50 | 0.50 below |
| One | 4, 9, 25, 8, 3 | 49 | 34.00 | 0.69 | 0.31 below |
| Two | 3, 9, 25, 8, 3 | 48 | 43.00 | 0.90 | 0.10 below |
| Three | 2, 9, 25, 8, 3 | 47 | 52.00 | 1.11 | 0.11 above |
| Four | 1, 9, 25, 8, 3 | 46 | 61.00 | 1.33 | 0.33 above |
| Five | 0, 9, 25, 8, 3 | 45 | 70.00 | 1.56 | 0.56 above |
The crossing point is not a matter of opinion either. Setting 25 plus 9 times the number removed equal to 50 minus that number gives ten times the number equal to twenty five, so the record's mean equals the true 1.00 per cent at exactly two and a half months removed. Since half a month cannot be removed, the third removal is the one that carries the record from understating the truth to overstating it. Three months out of fifty is six per cent of the record, and the shaded rows above are the ones on the wrong side.
The other thing worth noticing about the cleaned record is what it looks like from the inside. The five most disagreeable months are gone, so the cleaned record's months agree with each other far better than the honest record's months did. Anybody who judges a record by how tidy it is will prefer the wrong one.
How many of the fifty months have to be removed before the record's mean passes the true 1.00 per cent?
Take out the worst months, one at a time, and watch the answer walk past the truth
Fifty months, one square each, grouped by what the month did. The control removes the worst months and nothing else: no reading is edited, no month is added, and the true mean of 1.00 per cent stays pinned exactly where it is at every setting. The second control switches between five separate records of the same unit. The walk then shows up as either a quirk of one record or a property of the rule. Start at the default, the published fifty month record with nothing removed and a mean reading 0.50 per cent to the decimal, then drag right.
Two things are worth watching as the control moves. The months being removed are the ones furthest from everything else, so the count of months falls very slowly while the mean moves very fast. The second is that switching records changes where the crossing sits but not the direction of travel. Record four lands exactly on the true 1.00 per cent at three months removed and goes past it at four. Record five is already above the truth before anything is removed at all. On all five records every removal pushes the answer the same way. A rule that only ever removes bad cases can only ever push an average up, and a push is a direction rather than a wobble.
Why is look ahead bias the hardest of the four to see?
Look ahead bias enters when information that only became available later is used as though it had been available at the time. Nothing leaves the record and nothing joins it. Every figure in it is correct. The rows all add up.
Look ahead bias is the hardest of the four to find for exactly that reason: nothing is wrong with the record as a set of numbers, so no count, no total and no cross check on the record itself will reveal it. Only a question about timing will, and timing is not a property of any column. The question to put to each figure is when it became knowable.
Take it home first. A household sits down to score last year's monthly budget and works out that it managed groceries on Rs 9,000/- a month. Today's prices are the ones the household has in its head, so it does the sums with those. Today's prices are higher, so last year's grocery basket now looks like it cost more than the household actually paid, and the household concludes that it was better at shopping than it was. Every figure it used is a real price. None of them was a price it faced. The mistake is not in the arithmetic, it is in the calendar.
The same shape appears in a record of monthly changes. Suppose somebody scores each month of the Nakshatra unit as good or bad, and does the sorting using the full fifty month record in front of them. The scorer now knows which months turned out badly. If any part of the scoring leans on that knowledge, even by deciding which months were worth a second look, then the score is not something anybody could have produced at the time. The record has not been damaged; the claim made about it has.
A record is complete, every figure in it is correct, and it is still biased. Which of the four is that, and what question finds it?
What has a denominator got to do with any of this?
The fourth kind hides in the arithmetic rather than in the data, and so survives every check aimed at the rows. Working out a spread from a sample means adding up the squared gaps between each reading and the sample's own mean, then dividing. The choice of what to divide by is the whole story. Dividing by the count of readings gives one answer; dividing by the count less one gives a slightly larger one.
On the fifty month record, the squared gaps add to 1,212.50. Dividing by 50 and taking the square root gives 4.92 per cent. Dividing by 49 and taking the square root gives 4.97 per cent. Both figures sit under the true 5.00 on this particular record, and that part is luck. The relationship between them is not luck. On every record, without exception, the count denominatorThe divisor. Either the count of readings, or that count less one, and picking between the two settles the answer. gives a smaller answer than the corrected one. The two differ by a fixed multiplier that never changes sign.
Run it on all five records and the multiplier is identical to five decimal places each time. An identical multiplier every time is the signature of a one way error. A gap that is the same size and the same direction in every record examined is not coming from the sample at all, it is coming from the formula, and a bigger sample cannot average away a formula. The size of it depends only on how many readings there are: with fifty readings the count denominator returns about 0.99 of the corrected figure, and with five readings it returns about 0.89.
With five readings the same arithmetic gets loud. A five month stretch of the same unit reading minus 4.00, minus 4.00, 1.00, 6.00 and 6.00 per cent has a mean of 1.00 and squared gaps adding to 100.00. Dividing by four and taking the root gives 5.00 per cent; dividing by five and taking the root gives 4.47. The shortfall is 0.53 rather than 0.05, ten times as wide on a stretch a tenth as long. The corrected figure landing exactly on the true 5.00 here is luck and nothing more, exactly as its landing on 4.97 on the fifty month record was luck. The part that is not luck is which of the two figures is the smaller one, and it is the same one both times.
| Record | Squared gaps added | Spread on the count denominator | Spread on the corrected denominator | Count over corrected |
|---|---|---|---|---|
| Record one, the fifty month record | 1,212.50 | 4.92 | 4.97 | 0.98995 |
| Record two | 1,170.50 | 4.84 | 4.89 | 0.98995 |
| Record three | 1,274.50 | 5.05 | 5.10 | 0.98995 |
| Record four | 1,282.00 | 5.06 | 5.12 | 0.98995 |
| Record five | 1,032.00 | 4.54 | 4.59 | 0.98995 |
Look at the last column. Five different records, five different spreads, five different distances from the true 5.00 per cent, and one multiplier that does not budge. Two of the five records put both figures above the truth and three put both below. The scatter is the wobble. The last column is the push.
The count denominator gives 4.92 per cent and the corrected one gives 4.97, against a true 5.00. Is that gap noise or bias?
Why does adding more data not fix any of it?
The instinct to gather more data is a good one, and a good instinct is exactly what makes this trap catch careful people. More data really does help with the other problem. If an estimate wanders because fifty months is a small window, then two hundred months narrows the wander, and the figure that summarises how far a single estimate typically lands from what it is estimating, the standard errorA number describing how far one estimate typically falls from the value it is estimating. How it is built, and how it shrinks as the window grows, is covered separately., gets smaller as the window grows. The narrowing is real and it is worth having.
But bias is not wander, and the arithmetic of more data does nothing to it. Go back to the cleaning rule that drops the worst months. If a hundred records are built and every one of them is cleaned the same way, then all hundred report high. Their average is high. And here is the sting: their spread is small. They agree with each other, and they agree because they were all pushed in the same direction by the same rule.
The one thing that feels most like confirmation, many separate records agreeing, is exactly what a shared method produces, so agreement between records is evidence about the method they have in common and not evidence about the truth. Watch it with the five records. Averaged honestly, their running average moves 0.50, 0.90, 0.90, 0.78, 0.94, settling near the true 1.00. Cleaned the same way, their running average moves 1.56, 1.76, 1.76, 1.74, 1.80, settling near 1.80. Both settle down. Only one settles on the truth.
A hundred records built the same way all report the same high figure. What does the agreement between them establish?
Can one record carry all four at once?
One record can, and there is nothing unusual about it. The four are not four names for one thing, they are four different moments in the life of a record, and a record passes through all four moments on its way to becoming a figure somebody quotes.
Cases are chosen, and selection can enter. Readings are written down. Cases leave, and survivorship can enter. Readings get dated, or fail to, and look ahead can enter. The arithmetic runs, and measurement can enter. Then a figure is reported. By that stage everything that could go one way has already gone, so nothing further can get in. A record can carry all four at once, and because each one enters at a different moment, each one needs a different question to find it and no single check finds more than one.
Which of the four enters after the collecting of data is finished, and where does it live instead?
How is bias hunted in a record built by somebody else?
Most records that reach anybody arrive finished. A lender is handed a summary of a borrower's monthly receipts. An analyst is handed a spreadsheet of monthly figures somebody else assembled. A household is handed a statement of what a shop's takings have been. In none of those cases can the record be rebuilt, and in none of them is the truth known. Five questions can still be put to a finished record, one aimed at each moment in the timeline and one aimed at the answer itself.
Where did these cases come from, and who decided which ones went in. Which cases used to be in the record and are not in it now. Could anything here have been known only afterwards. Which formula was used, and on what denominator. And then the fifth question, the one that changes the conversation. If the answer is wrong, which direction is it wrong in? Asking which way rather than merely whether turns a vague worry into something checkable.
The fifth question is worth dwelling on. A person reading a record can always ask it and can usually answer it. A lender looking at a summary of monthly receipts does not need to know the true average to reason that a summary prepared by the borrower, from months the borrower chose, with the poor months set aside for review, can only be wrong in one direction. Reasoning about direction is not a calculation and does not need one. The reasoning is a statement about which way a rule pushes, and a rule that only ever removes bad cases pushes up. The same reasoning runs the other way when the person assembling the record had a reason to look cautious.
A cleaned record arrives with no cleaning log attached. Which single number should be asked for first?
The tidy record that everybody trusted more
An invented illustration. A set of monthly changes is cleaned before anyone analyses it, and the cleaning rule is a sensible sounding one: months that look unusual are set aside for review. Nobody ever reviews them and they never come back. Unusual and bad are the same months in this record, so five months of fifty go, all of them at minus 9.00 per cent.
The figure reported is 1.56 per cent against a true 1.00. The expensive part is not the 0.56 of error. The person reporting 1.56 is more confident than the person who reported 0.50, and has better looking evidence. The cleaned record's months agree with one another far better. Its spread is narrower. Its range is tighter. Every one of those improvements is a direct consequence of the fault, so the record that is further from the truth now looks like the more careful job, and there is nothing inside it that says otherwise.
No cleaning rule exists that cannot go one way, so the fix is not a better cleaning rule. The fix is a count. Keep a tally of everything removed and report the mean both ways whenever anything was dropped, 0.50 per cent across all fifty months and 1.56 across the forty five kept. A cleaning rule with no count beside it cannot be audited by anybody, and that includes the person who wrote the rule, who six months later will not remember what it took out.
What the figures are built on
| What a figure rests on | Where it came from | How to redo it |
|---|---|---|
| The five values of the Nakshatra unit and the five weights beside them | Written down before a single month was drawn, so they are stated rather than measured | Multiply each value by its weight and add the five products. They come to 1.00 per cent. |
| The fifty month record, tallied 5, 9, 25, 8 and 3 | Invented for teaching and then held fixed everywhere it appears | The five counts add to 50, and value times count adds to 25.00, so the mean is 0.50 per cent. |
| The six means in the survivorship build, 0.50 through 1.56 per cent | Recomputed line by line from that tally | Divide 25 plus 9 times the number removed by 50 minus the number removed. |
| The two spread figures, 4.92 and 4.97 per cent | Recomputed from the same tally, both from a total of squared gaps of 1,212.50 | Divide 1,212.50 by 50 and by 49 in turn, then take the square root of each. |
| The named ways a one way error gets in | Common teaching property, described here by the mechanism that produces each one | Named by the mechanism that produces each one rather than by attribution to a text. |
The Nakshatra unit, the fifty month record, the four further records of fifty months, the five month stretch, the record of interesting months and the cleaning log are invented.
Educational material. Not advice on any investment, tax, budget or market position.
