Heuristics and Biases: Shortcuts That Work Until They Do Not
A heuristic is a shortcut that answers a hard question by substituting an easy one. Three were named originally: judging by resemblance, judging by what comes to mind, and starting from a number and adjusting. Each usually works, which is why it survives. A bias is the error left behind when it fails in the same direction every time.
One idea carries the whole subject. The mind is not answering the hard question badly. The mind is answering a different, easier question rather well, and then presenting that answer as though it had been about the first one. Nothing feels wrong while this happens. There is no moment of substitution that can be caught in the act, no sense of having taken a lesser route. The swap is invisible from the inside, which is why a shortcut cannot be spotted by paying closer attention to how confident an answer feels.
Bounded rationality makes the case that deciding costs something, that capacity for it is limited, and that stopping early is the sensible response rather than a failure of will. Herbert Simon set that out in the Quarterly Journal of Economics in 1955. The obvious next question follows. If a person stops early, what exactly are they doing instead of the full calculation? The answer is that they run a shortcut, and shortcuts have a shape that can be drawn.
What is a Mental Shortcut, and what does it substitute?
A heuristicA shortcut that answers a hard question by substituting an easier one. is a rule that produces an answer quickly, without the work the question actually calls for. The definition worth carrying is narrower and more useful than the everyday one. A shortcut is not simply a quick answer or a rough estimate. A shortcut is a substitutionAnswering a different question from the one asked without noticing the swap.: one question is put in place of another, the second one is answered properly, and the answer is then reported under the first question's name.
Take it out of money first. Consider a stranger who has just been introduced as somebody's new business partner. The hard question is whether this person will do what they say they will. Future conduct is the subject, and there is no evidence about it at all. A completely different question gets answered instead, in about a second: does this person seem like the sort who does what they say? A lifetime of impressions is available to compare against, so the resemblance question is answerable. The answer arrives feeling like a judgement about the future, and it never announces itself as a judgement about resemblance.
Now put it in the setting this subject cares about. Meera Sundaram, an invented investor, is looking at a holding on a Tuesday evening. The hard question is whether that holding will do better than the alternatives over the years she has left before her stated goal. Nobody can answer that. The easy question, the one that actually gets answered, is whether this holding looks like the sort of thing that has done well recently. The two questions have different answers, and the mind returns the second one wearing the label of the first.
What does a heuristic substitute?
Which three shortcuts were named first, and by whom?
Amos Tversky and Daniel Kahneman set out three of them in a paper called Judgment under Uncertainty: Heuristics and Biases, published in Science in 1974. The 1974 paper is the origin of all three, and everything written about them afterwards is a restatement of it. Naming the paper matters here for a reason that is not politeness. The three were not collected from folklore or assembled from observation of investors. The three were proposed as specific substitutions, each one saying which easy question stands in for which hard one, and each one predicting a particular error in advance rather than explaining one afterwards.
The first is judging by resemblance. The hard question is how likely something is to belong to a category; the easy question is how much it looks like a typical member of that category. The second is judging by what comes to mind. The hard question is how common something is; the easy question is how readily examples of it arrive when somebody goes looking. The third is starting from a number and adjusting. The hard question is what a quantity actually is; the easy question is how far to move from whatever number is already sitting in front of the person deciding.
Naming which of the three is running says in advance which error to expect, and that is the entire practical value of having three names rather than one. Resemblance ignores how common the category is to begin with. Judging by what comes to mind mistakes vividness for frequency. Starting from a number leaves the estimate too close to wherever it started. Three substitutions, three characteristic errors, and each one repairable by a different move.
Somebody estimates how common a kind of failure is by how easily examples come to mind. Which shortcut is running?
Why does the substitution stay invisible to the person making it?
Here is the part that decides whether the rest of this subject makes sense. A shortcut does not feel like a shortcut. Because the easy question really has been answered well, a shortcut feels like an answer, and often a rather confident one. Confidence tracks how smoothly the answer arrived, not how good the evidence was, and a substituted question produces a much smoother arrival than the hard question ever could. The hard question would have prompted hesitation. The easy one does not.
The usual advice to think more carefully so often does nothing for exactly this reason. Thinking harder about a question that is not actually being asked produces a more careful answer to the wrong question. On 19 February, in the invented Palash decision log, a television segment names Suvarna Chemicals Limited and Meera adds Rs 1,00,000/- to that position the same evening, taking its cost to Rs 4,00,000/-. Asked why, she will give reasons, and the reasons will be about the business. The reasons were assembled after the easy question had already been answered, so the reasons are real and are also not what produced the decision. Dishonesty has nothing to do with it. The order of events is simply the ordinary one.
There is a household version of this that everybody has lived. Somebody asks whether a particular route home will be quick tonight. The hard question is about traffic at this hour on this day. Memory answers a different one instead: was that route quick the last two or three times it was taken? The answer comes with real confidence, and it is not experienced as an answer to a different question.
Why do shortcuts survive if they produce errors?
Because they are usually right, and because being usually right quickly is worth more than being always right slowly in nearly every setting a person has ever lived in. A great many summaries of the subject go wrong at exactly this point. Summaries present the shortcuts as defects, as bugs that a better designed mind would not have. Shortcuts are not defects. Shortcuts are answers to a design problem: how to decide well enough, fast enough, on evidence that is not available and will not become available.
Consider what is actually being optimised. A shortcut trades a small amount of accuracy for a very large amount of speed and effort, and it makes that trade in a setting where the cost of a wrong answer is usually small and recoverable. Judging a stranger by resemblance is wrong reasonably often, and it is wrong cheaply: the judgement adjusts as more is learned. The test that matters is ecological validityWhether a shortcut works in the setting it actually gets used in.. Ecological validity asks whether a shortcut works in the environment it evolved to work in, not whether it survives a laboratory problem designed to break it.
A shortcut is a bargain. The bargain is good in most settings and bad in a few, and the few are exactly the ones that need naming. Financial decisions happen to sit in the bad set unusually often, for three reasons that have nothing to do with anybody's intelligence. The feedback is slow and noisy, so which shortcut failed is rarely learned. The costs of being wrong are not small. And the setting is full of things that resemble each other closely while behaving completely differently. Judging by resemblance breaks down under precisely that condition.
How well can a shortcut do, at its best and at its worst?
The question deserves arithmetic rather than adjectives, and the arithmetic is simple enough to do in the head once it has been seen done once. Set up the cleanest possible version. Out of 100 cases, 20 genuinely belong to the category in question, so the base rateHow common something is before any particular evidence is considered. is 20.0 per cent. The shortcut looks at all 100 and flags 30 of them as belonging. The question is how many of those 30 flags are right.
Start at the bottom. Suppose the easy question tracks the hard one not at all, so the flags land without any relationship to the truth. Then the 30 flagged cases contain the same share of true cases as any other 30 would. The share is 6. Six right out of thirty is 20.0 per cent, exactly the base rate that held before the shortcut said a word. A shortcut that tracks nothing is not worse than nothing; it is precisely nothing, dressed up as information.
Now the top. Suppose the easy question tracks the hard one perfectly, so every one of the 20 true cases is among the 30 flagged. Perfect tracking gives 20 right out of 30, or 66.7 per cent. And notice what that ceiling is made of. The ceiling is not a limit on how good the shortcut is. The ceiling is arithmetic: the shortcut flagged 30 cases when only 20 could possibly qualify, so 10 of the flags are wrong no matter how well it works. Even a shortcut that tracks the truth perfectly is wrong a third of the time here, purely because it flagged more cases than the world contains.
| How well the easy question tracks the hard one | The working | Right, of 30 flagged | Share |
|---|---|---|---|
| Not at all | the 30 flagged hold the base rate share, 6 | 6 | 20.0 per cent |
| Halfway | 6 plus half of the remaining 14 | 13 | 43.3 per cent |
| Perfectly | all 20 true cases sit inside the 30 flagged | 20 | 66.7 per cent |
A shortcut flags 30 of 100 cases and 20 are truly in the category. Before anything moves: what is the best it can ever do?
Move the tracking strength and watch both halves stay true at once
One control moves: how strongly the easy question tracks the hard one, from 0 to 1. The base rate of 20 in 100 and the 30 cases the shortcut flags are both held fixed. One control teaching one relationship is the only way to see the relationship. Watch the two readings that matter. The share of flags that are right never falls below the 20.0 per cent base rate, and it never rises above the 66.7 per cent ceiling.
At a tracking strength of 0.50 the shortcut flags 30 of the 100 cases and 13 of them truly belong, which is 43.3 per cent against a base rate of 20.0 per cent. It is better than knowing nothing, and it is still wrong more often than it is right.
At zero tracking the shortcut returns 20.0 per cent. Why is that number familiar?
What is a Cognitive Bias, as against an ordinary mistake?
A biasAn error that repeats in the same direction and therefore survives averaging. is not the shortcut. A bias is what the shortcut leaves behind when it fails, and it earns the name only when the failure has a particular shape. An ordinary mistake is a mistake. A bias is an error that runs in one direction, repeats across occasions, and has a mechanism behind it that says when to expect it again. Three tests, and all three have to clear.
Run those tests on something measurable. In the Palash decision log, 96 of the 240 recorded decisions were buys. Of those 96 buys, 41 followed a media mention of the holding within three days, or 42.7 per cent. Against what? Against the fact that in any given week, 11.0 per cent of the eligible list gets mentioned at all. So mentions are attached to buys nearly four times as often as they are attached to the list a buyer was choosing from. The gap is 31.7 percentage points.
Stating a bias as the distance between an observed rate and the rate that would have been expected is what turns it from a complaint into a measurement. The distance gives a number that can be checked, that can move, and that somebody else could disagree with by producing a different expected rate. Without the 11.0 per cent, the 42.7 per cent is just a figure that sounds high. With it, there is something to argue about, and something to argue about is the condition for there being something to learn.
How does a bias differ from noise?
Bias and noise differ in direction, and the difference decides which cure will work. Daniel Kahneman, Olivier Sibony and Cass Sunstein separated the two carefully in Noise, published in 2021, and the separation is worth more than most of what gets written about either one alone. A bias is systematicHappening in a consistent direction rather than at random. error: it leans. NoiseScatter in judgements that points in no particular direction and averages away. is scatter: it does not lean, it just sprays.
The log gives one of each, side by side. The 41 of 96 buys following a media mention leans hard in one direction, repeats across quarters, and has a mechanism that can be stated in a sentence. All three tests clear, so the pattern is bias. The 24 pauses of a standing instruction, being 10.0 per cent of the 240 decisions, are the other case. Pauses happen after falls and after rises. The pauses cluster in no particular quarter. Different investors pause for different reasons and no shared trigger shows up. The pauses are noise. There is no direction to find, so no amount of staring at them will produce one.
Two failures that look identical in a summary need opposite responses, and telling them apart is the whole reason the distinction is worth learning. Noise is a spread problem: it is fixed with a process that makes judgements more alike, such as everybody using the same checklist in the same order. Bias is a location problem: it is fixed, if it can be fixed at all, by moving where the judgements sit. A checklist that makes everybody equally consistent and equally wrong has cured the noise and left the bias untouched.
41 of 96 buys followed a media mention, against an 11.0 per cent rate of the eligible list being mentioned at all. Bias or noise?
Why does averaging remove one and not the other?
Averaging is the test that makes the distinction operational instead of decorative, and the test follows from what the two words mean. Noise points in no particular direction, so when many noisy judgements are added up the departures cancel each other out, and the more of them added the more completely they cancel. The scatter shrinks with the square root of the number of judgements, so gathering four times as many decisions halves it. Bias points in one direction, so adding up many biased judgements adds up the lean along with them. Average a thousand of them and the lean is exactly where it started.
Meera lives this without knowing the words for it. If she reviews four of her own past decisions, the picture is mostly scatter: one was rushed, one was thought about for a fortnight, one happened while she was travelling. If she reviews all sixty of her decisions in the log, the scatter flattens and something else appears from underneath it: a steady lean toward whatever she had recently read about. Reviewing more of one's own decisions does not shrink a bias; it removes the noise that was hiding it. More data makes a bias look bigger while changing nothing about its size.
There is a household version. One person's estimate of what a wedding will cost is scattered by mood, by which supplier they last spoke to and by what they had for lunch. Ask twenty people from the same circle and the mood effects cancel out. The shared lean that everybody in that circle has survives: weddings always cost more than anybody plans for. Twenty estimates removed the noise and left the bias standing exactly where it was.
Somebody reviews far more of their own past decisions than they used to. What happens to their bias?
The error that gets made, and what it costs
The error is treating every departure as a bias. A bias is a satisfying thing to find: it has a name, it has a paper behind it, and it makes a decision feel explained. The temptation is enormous. So a run of scattered outcomes gets a label attached to it, and the label sticks.
Put the two entries from the log side by side and the difference is plain. The 41 of 96 buys that followed a mention lean one way, repeat across quarters, and have a mechanism. The 24 pauses of a standing instruction lean nowhere, cluster nowhere, and share no trigger. Calling the second one a bias is not a small imprecision. The label sends the search after a cause that is not there, and it recommends a cure that cannot work. The cure for scatter is more consistency, and the cure for a lean is a different position.
The error costs the ability to check anything. A named bias with no expected rate behind it cannot be measured, cannot be argued with, and cannot be shown to have improved. A named bias becomes a way of describing whatever already happened, and an explanation that does that has stopped being worth having.
When is a shortcut simply the right tool?
Often, and it needs saying without hedging. Everything above could be misread as an argument that shortcuts are a defect to be trained out. Shortcuts are not a defect, and anybody who has ever used one has not thereby decided badly. Two conditions decide it, and where both hold, the shortcut is the correct choice and the careful method is waste.
The first condition is that the easy question tracks the hard one in this particular setting. The whole line in the figure above was about that condition. Where the resemblance really does correlate with the outcome, the shortcut carries genuine information, and it carries it in a second rather than an afternoon. The second condition is that being wrong is cheap and reversible. A wrong answer that can be noticed and undone tomorrow costs almost nothing; a wrong answer that locks in for eleven years costs a great deal.
Where the easy question tracks the hard one and the error is cheap to undo, running the careful method instead is not diligence, it is waste of the only resource available. Choosing a vegetable stall by how busy it looks is a shortcut, and it is the right one: busyness genuinely tracks freshness, and being wrong costs one poor dinner. Choosing where the education money for eleven years goes on the same kind of evidence is the same shortcut in a setting where neither condition holds. The shortcut did not change. The setting did.
Is there a setting where using a shortcut is straightforwardly the correct choice?
How does a person deciding alone, or one deciding for others, use any of this?
Devika Rao, the adviser at the invented Palash Advisory Services Private Limited, does not use this material to work out which biases her clients hold. She has no measurement that would let her do that, and neither does anybody else in a first meeting. Her use for it is narrower and far more useful: sorting which of her own review questions are worth asking. When a client explains a decision, the reasons offered are the reasons assembled afterwards, so she asks instead what the client was looking at in the minutes before deciding. Asking what the client was looking at goes at the easy question rather than at the justification.
She also uses the bias and noise split to decide what a process change can possibly achieve. Of the 60 investors in the log, 20 adopted a written checklist on 4 November. Across the four quarters that followed, those 20 recorded a written reason on 34 of 41 decisions, being 82.9 per cent, against 19 of 63 decisions, being 30.2 per cent, for the other 40. A checklist makes decisions more alike, so it is a noise instrument first. A checklist is not evidence that any return improved, and the log makes no such claim: 60 people over eight quarters cannot carry one.
For a person deciding alone with no adviser and no committee, one habit carries everything above: writing down the question apparently being answered, before answering it. Writing the question down is not a debiasing technique, and correction is set out under debiasing. The habit simply catches a swap in the only moment when catching it is possible, before the answer arrives and starts to feel obvious. An analyst reading a research note gets the same value from the same habit, and so does a lender reading a file: naming the hard question, then checking whether the evidence in front of them is about that question or about something that merely resembles it.
Fourteen named effects follow. What should each of them be expected to be?
What should be expected of the fourteen effects that follow?
Fourteen named effects make up the largest run in behavioural finance, and every one of them rests on the shape set out above. Each is set out separately under cognitive biases. Each has a mechanism worth a full treatment of its own, an original paper that named it, and a worked case in the log. Summarising fourteen mechanisms in a paragraph each would produce a list of labels, and one more list of labels is the last thing the subject needs.
Every one of the fourteen has the shape just established, so the shape carries into all of them. Each is a shortcut before it is an error. Each substitutes some easy question for some hard one. Each is right often enough that it survived, and each fails in one direction under conditions that can be stated in advance. A shortcut that failed every time would never have become common enough to name, so reading the fourteen as a catalogue of human failure is the single easiest way to misunderstand all of them at once.
So the test to apply to each entry is not whether somebody recognises themselves in it. Self-recognition feels satisfying and teaches nothing. The test that works is the one already run twice: what is the hard question, what easy question got put in its place, and in which direction does the error go when the two come apart.
Sources
| Source | Document | Site |
|---|---|---|
| Amos Tversky and Daniel Kahneman | Judgment under Uncertainty: Heuristics and Biases, Science, 1974, in which all three original shortcuts were first set out | ssrn.com |
| Daniel Kahneman, Olivier Sibony and Cass Sunstein | Noise, 2021, in which systematic error and scatter are separated as two different failures needing two different cures | cited to the book itself |
| Herbert Simon | the paper setting out limited capacity as the starting point for the shortcuts, Quarterly Journal of Economics, 1955 | nber.org |
Meera Sundaram, Devika Rao, Palash Advisory Services Private Limited, the Palash decision log and Suvarna Chemicals Limited are invented.
Educational material. Not advice on any investment, tax, budget or market position.
