Base-Rate Neglect: Ignoring How Common Something Actually Is
A base rate is how common something is before any particular evidence is looked at. Neglecting it means answering with the force of the evidence alone, as though the question asked how convincing the signal is rather than how likely the conclusion is. The size of the resulting error depends entirely on how rare the thing was to begin with.
Base-rate neglect is one count done properly, and the count is short. The whole error is that two numbers have to be combined and only one of them feels like an input. Kahneman and Tversky set the effect out in On the Psychology of Prediction, in Psychological Review in 1973. Fifty years on, the effect is still live, and not because people are bad at arithmetic. The second number does not feel like arithmetic at all.
What is a base rate, and where does one come from?
A base rateHow common something is before any particular evidence is considered. Also called the prior, or the prevalence. is a count taken over a group, taken before the case in front of the analyst is looked at. How many of the pens in this box write. How many parcels ring the bell twice. How many of the decisions logged last quarter were later reversed. A base rate is a fact about a population, and says nothing whatever about the particular case in hand. Saying nothing about the case is exactly why a base rate feels irrelevant, and exactly why it cannot be dropped.
Start away from money. A watchman at a housing gate hears a bell ring twice, sharply. He has been told that couriers ring twice. Should he expect a courier? The answer depends on something the bell cannot tell him: how many people ring that gate in a morning, and how many of them are couriers. Suppose the gate takes forty rings a day and two are couriers. So many more people get the chance to ring twice out of habit that a doubled ring is a weak clue, no matter how reliably couriers ring twice. The strength of a clue and the frequency of what it points to are separate facts, and the answer needs both.
Base rates come from three places and it is worth knowing which one is in hand. The first is a record that can be counted directly, which is the strongest and the rarest. The second is a published count somebody else took, usable once what was counted and over what period are known. The third is an honest estimate with a range attached. An estimate is much weaker, and still enormously better than nothing: even a rough count of how common something is will move an answer further than an extra opinion about the case. The Palash decision log, an invented record of 240 decisions taken by 60 investors over eight quarters, is an example of the first kind. Of the 240, 84 carried a written reason, or 35.0 per cent, and 71 were taken within 48 hours of a news item, or 29.6 per cent. Both of those counts are base rates.
What is a base rate?
Which two numbers does an honest answer need?
The whole apparatus is three numbers and a division, set up once. There is a thing that is either true or not true of the case at hand. There is a signalEvidence that appears more often when something is true than when it is false. A test result, a flag, a symptom, a tell.. A signal is evidence that appears more often when the thing is true than when it is false. The signal has two properties, and both matter. Its hit rateHow often the signal appears when the thing really is true. Also called sensitivity. is how often it appears when the thing is true. Its false alarm rateHow often the signal appears when the thing is false. Every signal has one, and it is the property that usually goes unmentioned. is how often it appears when the thing is false.
Now the question. The signal has appeared. How likely is it that the thing is true? The number has a name, the posteriorThe correct likelihood after the base rate and the evidence have been combined. Before is the base rate; after is the posterior., and it needs three inputs: the base rate, the hit rate and the false alarm rate. Take any one of the three away and the question has no answer at all. Having no answer is a stronger statement than getting a less accurate one. The signal usually arrives with its hit rate attached, so people almost never leave that one out. The base rate is what gets left out, and left out so completely that nobody notices a number is missing.
The failure belongs to one person forming one judgement from evidence that has reached them, not to a price and not to anything a crowd does to a market. A watchman at a gate, an adviser reading a review flag and a person deciding alone at a kitchen table are running the identical machinery, and it fails the identical way.
Why does the specific crowd out the general?
The mechanism is what makes base-rate neglect a bias rather than a mistake in long division. The specific evidence is about this case. Specific evidence arrives recently, has detail, can be pictured, and a story can be told with it. The base rate is about cases in general. The base rate has no detail, cannot be pictured, and no story can be told with it at all. So the mind files one of them as evidence and the other as background, and background does not get added to anything. The filing decision is the whole of the bias. The name for the sorting is crowding outSpecific evidence displacing general information rather than combining with it. The general number is not disputed, it is simply not used.: the general number is not argued with, it is not used.
The substitution underneath is covered under representativeness, the stated prerequisite. The hard question, how likely is this, gets swapped for an easy one, how well does this case match the picture I have of that kind of case. The swap is not lazy and it is often excellent. The swap cannot carry a frequency. Resemblance has no room in it for how many of a thing there are. Representativeness establishes the swap. Base-rate neglect is the count the swap skipped, left undone.
Watch it happen in the log. On 19 February a television segment names Suvarna Chemicals Limited and Meera Sundaram adds Rs 1,00,000/- to that holding the same evening, taking its cost to Rs 4,00,000/-. The segment is specific, vivid and about this holding. The general number sits beside it in the same log. In a given week 11.0 per cent of the eligible list gets mentioned at all, and 41 of the 96 logged buys, or 42.7 per cent, followed a mention within three days. The mention was treated as information about the holding when it was mostly information about what a television segment does in a week. Whether the purchase was wrong is a separate question, and so is what happened to the holding afterwards. The point is narrower and harder: one of the two numbers was used and the other was in the same record, unread.
Why does specific evidence crowd out the base rate?
What does the arithmetic look like when it is counted out?
Counting a worked case out rather than reaching for a formula shows what a formula would only hide. Devika Rao, the adviser at the invented Palash Advisory Services Private Limited, runs a review flag over each logged decision. The flag is meant to catch decisions that get reversed within a quarter. The flag fires on 80 per cent of the decisions that really do get reversed, and on 30 per cent of the ones that do not. Suppose that 20 per cent of decisions get reversed within a quarter.
Now count out 100 decisions. Twenty of them are reversals and eighty are not. Of the twenty reversals, the flag fires on 80 per cent, or 16, and stays quiet on the other 4. Of the eighty non-reversals, the flag fires on 30 per cent, or 24, and stays quiet on the other 56. Check the four groups add up: 16 plus 4 plus 24 plus 56 is 100. The flag has now fired 16 plus 24 times, or 40 times, and it was right on 16 of them. Sixteen out of forty is 40.0 per cent, exactly half of the 80 per cent the flag arrived advertising.
| Group | The working | Out of 100 |
|---|---|---|
| Reversed, and the flag fired | 20 reversals, the flag fires on 80 per cent of them | 16 |
| Reversed, and the flag stayed quiet | the other 20 per cent of the 20 | 4 |
| Not reversed, and the flag fired | 80 non-reversals, the flag fires on 30 per cent of them | 24 |
| Not reversed, and the flag stayed quiet | the other 70 per cent of the 80 | 56 |
| Every decision, counted once | 16 plus 4 plus 24 plus 56 | 100 |
| Times the flag fired | 16 fired on a reversal, 24 fired on a non-reversal | 40 |
| How often a fired flag is right | 16 of the 40 | 40.0 per cent |
Nothing in that count is clever. Four multiplications and one division are the whole of it, and a reader who has never seen the updating rule written in notation has just applied it correctly. The four groups do the work that notation would do, and a count can be checked by pointing at it. The reason the answer surprises people is not the arithmetic. The surprise is that the 24 false alarms come from a much bigger pile than the 16 true detections do, and the size of that pile is the base rate, arriving from off to one side where nobody was looking.
Of 100 decisions at a 20 per cent base rate, the flag fires 40 times. How many of those 40 are really reversals?
How far is the correct answer from the number the signal arrived with?
Put the two numbers on one line and the size of the error stops being an abstraction. The flag announces 80 per cent. The number is printed on the flag, and it is a true number about the flag. The correct likelihood, once the base rate has been counted in, is 40.0 per cent. The distance between them is 40.0 points, which on this setting is as large as the answer itself. Somebody acting on the printed number is not being slightly optimistic. The belief being held is twice as strong as the evidence supports.
And here is the part that keeps this effect alive: the flag was not lying. Eighty per cent of reversals really do get flagged. The statement is true, and it is a statement about the flag. The question the reader actually asked, though, was about a decision that has been flagged. Flagged decisions are a different group entirely, made up mostly of decisions that were never going to be reversed. The signal was not wrong, it was answering a different question from the one that was asked. Almost every real instance of this error has that shape: a true number, correctly reported, quietly answering the wrong question.
A flag fires on 80 per cent of the cases where the thing is true and on 30 per cent where it is false, and the thing happens 20 per cent of the time. Before the control below is moved: how likely is the conclusion once the flag has fired?
What happens when the signal is held fixed and only the base rate moves?
One demonstration settles the argument, and what is held still in it matters as much as what moves. Both properties of the flag are frozen at every setting: it fires on 80 per cent of the cases where the thing is true, and on 30 per cent where it is false, from one end of the control to the other. Nothing about the quality of the evidence changes. The only thing that moves is how common the thing was before the flag was ever consulted, and the correct answer moves across almost the whole range with it.
Freeze the signal, move only how common the thing is
One variable moves: the base rate, from 1 to 50 per cent. The flag's two properties are held fixed at 80 per cent and 30 per cent throughout. Every movement shown is therefore caused by the base rate alone. The flat line is what the flag advertises about itself. Nothing about the flag changes, so the line never moves. The default setting of 20 per cent reproduces the worked count above exactly: 16, 4, 24 and 56 out of a hundred, and 16 of 40 firings, or 40.0 per cent. At a base rate of 5 per cent the correct answer is 12.3 per cent; at 10 it is 22.9; at 20 it is 40.0; at 35 it is 58.9; and at 50 it is 72.7.
At a base rate of 20 per cent, 1,000 cases split into 200 where the thing is true and 800 where it is not. The flag fires on 160 of the true and 240 of the false, so of 400 firings only 160 are right, and the correct likelihood is 40.0 per cent against the 80 per cent the flag advertises.
Moving the base rate from 20 per cent to 50 per cent makes the answer climb from 40.0 to 72.7. What has changed about the signal?
Why is the error worst exactly when the thing is rarest?
Now the finding that earns the effect its place. With the control dragged down towards the left the answer does not fall gently, it collapses. At a base rate of 5 per cent the correct likelihood is 12.3 per cent. Counted out of 1,000 cases, the reason is plain. Fifty of the thousand are true and 950 are not. The flag fires on 40 of the 50 true cases, and on 285 of the 950 false ones. So it fires 325 times and is right 40 of them. Of every eight times the flag fires, just over seven of them are firing on something that is not there.
Read that again with the flag's reputation in mind. Nobody would call this a weak flag. The flag catches four out of five of the things it is looking for, and stays quiet on seven out of ten of the rest. Both are respectable properties, and a person told only those two numbers would trust the flag. The flag has not failed at all: rarity alone produced the result, and rarity is a fact about the population rather than a fault in the evidence. The effect is therefore at its most dangerous exactly where it does the most harm. Rare things are the ones worth detecting. Rarity is also what makes the correct answer collapse. Both statements are about the same number.
At a base rate of 5 per cent, how often does this flag fire on something that is not there?
Why does a perfectly respectable signal arrive nearly useless?
Look at how a signal actually reaches somebody. The shape of the delivery explains why the error is structural rather than personal. A signal arrives with its own accuracy attached. Whoever built it measured how often it fires when the thing is true, and that number gets printed alongside. Printing it is what makes the signal worth having. The number that never gets printed is how common the thing is in the group being tested. The base rate belongs to the population, not to the signal, so the person who built the signal had no reason to carry it and often did not know it.
The result is evidence that is complete in its own terms and unusable in the terms that matter to the person holding it. The one number needed to turn a signal into a likelihood is the one number nobody prints beside it. Stating it plainly changes what the repair has to be. The repair cannot be to read the signal more carefully. The missing number is not there to be read. The repair has to be to go and get the number from somewhere else, before the signal is interpreted at all.
What is the repair, and where does it fit in?
The repair is a single question asked in a particular place. Before interpreting any signal, ask: how common is this thing anyway, in the group this case came from? Not afterwards, as a sanity check on a conclusion already formed. By then the conclusion has a story attached, and the number will be argued with rather than used. Before, when the signal is still just a signal.
Three things then happen, and all three are useful. If the number can be counted, it combines with the signal to give an answer that can be defended, as the four groups did above. If it can only be estimated with a range, the range combines to give an answer with a range. An honest range beats the printed figure by a long way. And if it cannot be obtained at all, something genuinely valuable has been learned: this signal cannot be turned into a likelihood by anybody, and any confident number stated from it, by the reader or by whoever sent it, is not a measurement. Not being able to find the base rate is a finding, not a dead end.
Devika Rao's version of this at the practice is a line in the checklist that 20 of the 60 logged investors adopted on 4 November: before acting on any flag, write down how often this thing happens at all. Notice how modest that is. The line is not general scepticism about evidence. General scepticism would be worse than useless: a person who distrusts every signal equally throws away the 80 per cent along with the error. The line is one number, fetched once, before reading.
What is the repair, stated as one question?
When can a base rate be ignored without doing damage?
Sometimes, and it is worth being precise about when. A rule applied everywhere stops being used anywhere. A base rate can be set aside when the evidence is strong enough that no plausible value for it would change the conclusion. If a signal almost never fires when the thing is false, the false alarm pile stays small even when the population is enormous, and the answer stays high across every base rate that might reasonably be assumed. Direct observation is the extreme case. Nothing is being estimated about how many pens in the box write: this pen is in hand and writing.
The test is worth running as a sentence rather than a feeling. The highest and lowest base rates that could be defended are taken, the answer is worked at both, and the two answers are compared to see whether they point the same way. If they do, the base rate was not the binding input and the conclusion stands. If they do not, the conclusion is resting on an assumption that has not been made explicit. The exemption is real and it is much rarer than the confidence people place in ordinary signals suggests. A flag that fires on 30 per cent of the cases where nothing is happening is nowhere near it.
When can a base rate legitimately be ignored?
The error that gets made, and what it costs
The error is answering the question the signal came with instead of the question actually asked. The flag reports 80 per cent, the reader repeats 80 per cent, and nobody involved has said anything false. A fact about the flag has been used as though it were a fact about this case, and the two are only the same number when the thing being looked for is about as common as not.
The cost is everything downstream of the belief. At a base rate of 20 per cent, a reader acting on the printed figure holds a belief twice as strong as the count supports. At 5 per cent they hold one nearly seven times too strong, and they will act on it repeatedly. Of the firings they respond to, 87.7 per cent are firing on nothing, and each of those looks exactly like the ones that were right. The cost is not one bad call. A steady stream of effort, attention and follow-up goes on cases that were never there, with no feedback that would ever reveal it.
The second cost is subtler and lands on whoever built the check. A signal judged by how often it catches the thing looks excellent and gets kept. Judged by how often its firings are right, the same signal at a low base rate looks poor, and the difference between those two verdicts is not a disagreement about the evidence. One group of cases has been counted two different ways.
How does somebody actually use this, at a desk or at a kitchen table?
For a professional deciding on behalf of other people, this is a filing habit rather than a technique. Any check, screen, flag or alert that a practice runs is a signal with two properties and an invisible third input. The practical move is to write the base rate onto the same sheet as the check, once. The person reading the output is then not required to remember that a number is missing. Devika Rao can do this for the Palash log because she holds the rows: she can count how often the flagged thing actually happens across the 240 logged decisions and print it beside the flag. Where the count cannot be taken, the honest output says so and gives a range, and that is better work than a confident figure nobody can support.
For a person deciding alone, with no adviser and no committee, the same habit is one written line before acting on anything that arrived looking convincing: out of a hundred cases like this, how many turn out this way? Meera Sundaram can ask it of a television segment as easily as of a review flag, and the answer that stopped the 19 February decision from being automatic was sitting in the same record all along, at 11.0 per cent of the eligible list being mentioned in a week. The habit that survives contact with a busy week is writing the number down, not remembering to think about it.
One caution about the evidence itself. A single log of 240 decisions illustrates a mechanism and could not show that any procedure works. Showing that would need far more than one record.
Sources
| Source | Document | Site |
|---|---|---|
| Daniel Kahneman and Amos Tversky | On the Psychology of Prediction, Psychological Review, 1973, the paper in which base-rate neglect is set out | ssrn.com |
| Daniel Kahneman and Amos Tversky | Subjective Probability: A Judgment of Representativeness, Cognitive Psychology, 1972, the paper establishing judgement by resemblance | ssrn.com |
| Amos Tversky and Daniel Kahneman | Availability: A Heuristic for Judging Frequency and Probability, Cognitive Psychology, 1973 | ssrn.com |
| Amos Tversky and Daniel Kahneman | Judgment under Uncertainty: Heuristics and Biases, Science, 1974 | ssrn.com |
| Social Science Research Network and the National Bureau of Economic Research | open-access repositories of working papers in economics and the social sciences | ssrn.com and nber.org |
Meera Sundaram, Devika Rao, Palash Advisory Services Private Limited, the Palash decision log and Suvarna Chemicals Limited are invented.
Educational material. Not advice on any investment, tax, budget or market position.
