Selection and Survivorship Bias: When the Sample Was Never Representative
A record is unrepresentative when the things inside it were chosen in a way that moves the answer. The Neelbagh market office, invented, keeps only stalls that are still trading, so two that closed were deleted along with every month they traded. Month 1 then averages Rs 38,500/- across the eight kept and Rs 35,000/- across the ten that really traded, 10.00 per cent apart.
Two things are already in hand, and both get used below without being rebuilt. The Neelbagh stall record, invented for these notes, has already been printed and walked through: four months numbered 1 to 4, one row for each stall in each month, and eight columns whose meanings are written down. And the arithmetic averageAdd up the numbers that are usable, then divide by how many were used. Nothing more than that, and it is settled in full separately. has been defined along with the convention this record uses for it. The convention is repeated below every time a figure appears.
The convention decides every average below, so here it is in one line. The average takingsThe money that came across the counter during a month, counted before rent, wages, stock or anything else is paid out of it. A gross figure, not a profit. per stall month is the sum of the takings cells that carry a usable number, divided by how many carry one. A blank is not a number. The office code for a return that never arrived is not a number. A real zero is a number. Where a division below uses some other divisor, the divisor is named in the same sentence as the figure.
Nothing else is assumed. The whole argument is arithmetic on ten whole numbers, and it turns on which ten numbers were handed over and which ten were not. The word bias is used here in one narrow sense only: the record was built out of the wrong set of things. Bias in that narrow sense says nothing about a procedure, a formula or a long run.
What does it mean for a record to be unrepresentative?
A record is a set of things somebody decided to write down. Somebody drew a line around what would be included, somebody kept it up over time, and somebody decided what happened to a row when the thing it described stopped existing. Every one of those is a decision made by a person, before the file is ever opened.
A record is unrepresentative when the way things got into it is connected to the thing being measured. That connection is the whole fault. If entries arrive and leave for reasons that have nothing to do with the quantity of interest, the record can be small, scrappy and still honest. If entries leave precisely because of that quantity, the record can be enormous, tidy and completely wrong.
Here is the everyday version, and it costs nothing to picture. Somebody stands outside a cinema at the end of a film and asks ten people coming out whether the film was any good. Ten answers come back, and they can be averaged. Walking out is exactly what stops somebody being available to ask, so the four people who walked out at the interval are never met. The people missing from those ten are missing for the very reason the question was being asked, and no amount of politeness in the questioning repairs that.
Now notice what does not fix it. Asking a hundred people instead of ten does not fix it. Asking on three different evenings does not fix it. A bigger record gathered the same way is wrong by the same amount, and the only thing that grows is the confidence with which the wrong figure gets quoted. Size and correctness are separate properties, and a fault in how things were chosen is untouched by how many were chosen.
Does collecting more rows fix an unrepresentative record?
Who decided what this record would ever cover?
Before a single row exists, somebody settles which things the record will cover. The Neelbagh market has fourteen pitchesA marked out patch of floor inside a covered market, let by the month. It has no walls of its own, and a trader who takes one brings the stall.. In month 1, ten of them were trading and four stood empty. All four stood empty across all four months, so they produce no rows and never could. The market office keeps a record of trading stalls, and an empty pitch is not a trading stall.
The set of ten trading stalls is the frame, and the frame answers one question: what was this record ever going to cover. A frame is a decision rather than a fact. The office could have written the record differently, keeping one row per pitch per month with the empty ones showing nothing, and then fourteen would have been the frame. The office chose not to, and that choice was made before the first row was written and is nowhere recorded inside the file.
Four pitches standing empty is not a fault in the record. The record was built to cover trading stalls, and an empty pitch was never going to produce a row. The empty pitches are a frame, not a loss. But a frame decides which questions a record can answer, and the moment the frame is forgotten, the record gets used to answer a question it was never built for.
Four of the Neelbagh market's fourteen pitches stood empty across all four months. Is that a fault in the record?
Why do three different averages of the same month's money all come out right?
Month 1 is worth dwelling on because it has no faults in it at all. No blank, no repeated row, no odd code, no misspelling. Ten stalls traded, ten rows were written, and the takings add to Rs 3,50,000/-. Everything that follows comes out of that one clean month.
Rs 3,50,000/- divided by the ten stalls that traded gives Rs 35,000/-. The same Rs 3,50,000/- divided by all fourteen pitches in the market gives Rs 25,000/-, lower by 28.57 per cent. Both divisions are correct. The two divisions are not two attempts at one figure, one of them sloppier than the other. Each answers a different question. The first says what a trading stall took in month 1. The second says what a pitch in this market earned in month 1, counting the empty ones as the nothing they earned.
Choosing between those two divisors is selection, and it is worth being exact about what moved. The money did not move. Nobody deleted a row, nobody lost a return, nobody hid anything. The person doing the dividing changed their mind about which set they were describing, so the denominatorThe number that is divided by. It sits underneath the line in a division, and it is what decides which set the answer is an average over. changed. Selection changes the denominator and leaves the money alone, and the only harm it does is when the denominator goes unstated.
Month 1's takings across the market are Rs 3,50,000/-. Which is the right average: Rs 25,000/- or Rs 35,000/-?
What is survivorship bias, and why is it not the same fault as selection?
Now the second fault, and it arrives quietly. The market office does not keep a historical archive. For rent, for licences and for the notice board, the office needs only the stalls currently trading, and a working file of those is what it keeps. NB-09 Roshni Juice and NB-10 Amber Rolls both closed after month 2. When they closed, the office removed them from the file, and it removed them completely: their month 1 rows and their month 2 rows went with them.
Look at what that does to month 1 inside the file. The takings no longer add to Rs 3,50,000/-. Rs 20,000/- and Rs 22,000/- have left, so the takings inside the file add to Rs 3,08,000/-. The count is no longer ten. It is eight. Survivorship changes the money and the count together. Selection moves only the count, and that is precisely why the two faults must never be given one name.
Feel the difference with something ordinary. A shopkeeper who counts every till in a street, including the two that took nothing, and one who counts only the tills that rang, are doing selection: same street, same money, different denominator, and either answer is fine as long as they say which they meant. A shopkeeper who quietly throws away the takings slips of every shop that shut during the year is doing something else entirely. Both the total and the count have shifted, and the slips are gone, so nobody looking at the drawer afterwards can even tell.
The worked instance: month 1, printed whole
Here is every stall that traded in month 1 of the Neelbagh stall record, with what it took and whether it is in the file the office hands over.
| Stall | Name | Took in month 1 | In the office file? |
|---|---|---|---|
| NB-01 | Kadamba Idli | Rs 42,000/- | yes |
| NB-02 | Chandan Tea | Rs 31,000/- | yes |
| NB-03 | Harit Greens | Rs 38,000/- | yes |
| NB-04 | Peetal Utensils | Rs 26,000/- | yes |
| NB-05 | Bansi Flour | Rs 55,000/- | yes |
| NB-06 | Ilaka Fruit | Rs 34,000/- | yes |
| NB-07 | Sundari Chaat | Rs 40,000/- | yes |
| NB-08 | Peeli Mithai | Rs 42,000/- | yes |
| NB-09 | Roshni Juice | Rs 20,000/- | no, closed after month 2 and deleted |
| NB-10 | Amber Rolls | Rs 22,000/- | no, closed after month 2 and deleted |
| Ten stalls that traded | Rs 3,50,000/- | eight of them survive in the file | |
And the three averages, each printed with the divisor that made it, under the convention named at the top.
| Money divided | Divided by | Average takings per stall month |
|---|---|---|
| Rs 3,50,000/- | 14 pitches in the market | Rs 25,000/- |
| Rs 3,50,000/- | 10 stalls that traded | Rs 35,000/- |
| Rs 3,08,000/- | 8 stalls the office kept | Rs 38,500/- |
Read the left column of that second table before anything else. The first two rows divide the same money. The third does not, and the Rs 42,000/- difference in the numerator is the two deleted stalls. Printing the three figures side by side without saying that the third one moved the money too draws a picture in which survivorship looks like one more choice of divisor, and it is not.
What does survivorship cost here, in rupees?
Month 1 averages Rs 38,500/- across the eight the office kept and Rs 35,000/- across the ten that actually traded. The gap is Rs 3,500/- per stall monthOne stall for one month, which is the thing being counted here. Ten stalls across four months would make forty of them at most., and that gap is exactly 10.00 per cent above the honest figure. The fault is worth Rs 3,500/- a stall a month, stated as money rather than as a caution.
Ask why it is that large and the answer is the mechanism itself. The two stalls that closed were the two smallest takers in month 1, at Rs 20,000/- and Rs 22,000/-. The two averaged Rs 21,000/-, a full 40.00 per cent below what the ten averaged. The office deleted 20.00 per cent of the market, and it deleted the weakest fifth of it. A record never loses a random selection of what was in it. Here the reason a stall left is the very quantity being measured, its takings.
The whole of survivorship sits in that mechanism, and the mechanism is worth stating twice. Closing is not an accident that strikes stalls at random. A stall closes because it was not taking enough. Takings is also the column being averaged. So the reason for removal and the quantity being measured are the same quantity, and every removal pushes the average one way only.
When month 1's weakest stalls are dropped from the record one at a time, does the average move a lot or a little?
Drop month 1's weakest stalls one at a time, then look at the same record the way its file would show it.
One control moves: how many of month 1's ten stalls have been dropped, always from the weakest upward, from none to four. The ten takings figures never move; the record wrote them once and they stay written. As the control is dragged, the dropped bars grey out, the average line slides upward and the live sentence restates the reading. The second button below redraws the same setting as a complete file, with the dropped stalls gone rather than greyed, and a file handed over looks exactly like that. The panel opens at two dropped, the number the market office deleted.
The eight surviving stalls averaged Rs 38,500/- in month 1 and all ten averaged Rs 35,000/-. What is the overstatement?
Why does a file with this fault in it look completely clean?
Take the six counts that get run against any record of this shape and run them against the file the office hands over. Rows in the file, thirty two. Distinct pairs of stall and month, thirty one, one short of the rows because one stall month was filed twice. Takings cells carrying a number, thirty one, one short of the rows because one cell is blank. Distinct stall identifiers, eight. Distinct stall names, nine, one more than the identifiers because one stall is spelled two ways. Stall months the eight by four grid expects, thirty two, one more than the pairs present because one row never arrived.
Four faults show up as a gap of one between two counts. Now put the four deleted rows back, run all six counts again, and watch what happens. Every count rises. Not one gap changes size. A check reads what is there, and the deleted rows are not there to be read, so the fault that costs most in this record leaves no trace in it at all.
One more column shows how thorough a clean profile can look while still saying nothing. The filed on dayA column in this record holding the day of the following month on which a stall's monthly return reached the office. It is used here only as one more column that checks out. column is complete on every row in the file, in range on every row, and perfectly consistent. The two deleted stalls filed on time in both months they traded, so the column was complete for them too. With those rows gone, nobody will ever know it.
Every count on the office's file comes out clean. What does that establish about survivorship in it?
Is one stall month missing from this file, or nine?
Two correct answers make this question worth asking. Lay the file against the grid it implies: eight stalls, four months, thirty two stall months expected. Thirty one distinct pairs are present, so one is missing. The missing pair is NB-04 Peetal Utensils in month 3, whose return never arrived.
Now lay the same file against the grid the market really had: ten stalls, four months, forty stall months. Thirty one are present, so nine are missing. One of those nine is the absent return. Four are the rows the office deleted, being NB-09 and NB-10 in months 1 and 2. The remaining four are months 3 and 4 for those two stalls, when they no longer existed at all.
The grid was chosen before the file was written, and nothing inside the file records the choice, so both counts are correct arithmetic and the file cannot show which grid is right. A class register with the students who left already rubbed out adds up perfectly, every time it is checked, forever.
Is one stall month missing from this file, or nine?
What should be asked of any record before it is averaged?
Everything above is one market and ten numbers, but the habit it teaches is portable, and it is four questions long. A lender sizing a loan against a trade whose failures quietly leave the books, an analyst reading a set of results supplied by whoever is still around to supply them, and a household comparing what neighbours say they spend on school fees are all standing in the same place. The four questions are put to any record, in this order, before anything is divided by anything.
Question four is the one that does the work. A yes to it is the whole of survivorship, and here is the uncomfortable part: the answers are facts about how the file is kept rather than facts inside it, so not one of the four questions can be answered by reading the file. The answers come from asking the keeper. In the Neelbagh market the answer to question two sits in the office day bookA paper notebook the office writes sentences into, one line for each thing that happened, with no columns at all. What a notebook like that can and cannot answer is covered separately., in a line about two stalls handing back their pitches. A notebook of sentences has its own treatment. The narrow claim here is that the sentence exists, it is not in the file, and nobody who only opened the file will ever go looking for it.
What went wrong: a leaflet whose arithmetic was perfect
The Neelbagh market office prints a short leaflet for traders asking about a vacant pitch. The leaflet quotes Rs 38,500/- a month as what a stall in this market takes. The figure is computed correctly, from the file the office keeps, under the convention repeated throughout. Nobody lied and nobody miscalculated.
Every stall in that file is still trading, and being still trading is exactly why it is in the file. The two stalls that took Rs 20,000/- and Rs 22,000/- and then shut are not in it. Neither are the four pitches nobody ever took. So a trader reading the leaflet is shown Rs 38,500/- for a market whose trading stalls averaged Rs 35,000/- and whose pitches averaged Rs 25,000/-, and the two figures they most needed are the two the file cannot produce.
The habit that fixes it is to ask what left the record before averaging anything, and to take the answer from the person who keeps it rather than from the record itself. Had the leaflet carried one extra sentence naming what the figure was divided by and what had been removed, the arithmetic would have been unchanged and the reader would have been in a completely different position.
The office leaflet quotes Rs 38,500/- a month, computed correctly. What is the trader not being told?
Where each of these threads is picked up. The counts run against a file, and what each of them can and cannot expose, are covered separately. Missing values inside a record, why a cell is empty and what that decides about the fix, are covered separately, as are rows that repeat and what they distort. Whether a large figure is a real month or a recording error is a question about provenance rather than about size, and it has its own treatment. How research gets designed so that a result is able to fail honestly, and so that anybody can retrace which data produced which answer, is covered separately again and sits outside these notes on data.
What was read to produce these figures?
Nothing was read. Every figure below was invented for teaching and then checked by working the arithmetic twice, once forward from the ten takings figures and once backward from each average.
| The material used | What it gave | Standing |
|---|---|---|
| The Neelbagh stall record, month 1 | Ten takings figures totalling Rs 3,50,000/- | Invented for teaching, and faulted on purpose |
| The market office file | The eight stalls kept and the two deleted | Invented for teaching |
| The Neelbagh pitch list | Fourteen pitches, ten of them trading | Invented for teaching |
| The arithmetic, worked twice | Every average, gap and share printed above | Recomputed here, never transcribed |
The Neelbagh market, the Neelbagh stall record, the market office and the ten stalls from NB-01 Kadamba Idli to NB-10 Amber Rolls are invented.
Educational material. Not advice on any investment, tax, budget or market position.
