Representativeness: Judging by Resemblance Rather Than Probability
Representativeness judges how likely something is by how much it resembles the type, rather than by how common the type actually is. The shortcut is fast, it is usually roughly right, and it fails in two specific ways: it ignores how common the thing is to begin with, and it treats a small sample as though it were a large one.
The shortcut rests on a swap nobody notices making. The situation poses a question about probability, and probability is something a mind computes slowly and badly. Similarity is something a mind computes instantly, unasked, with no sense of effort at all. When a hard question has an easy question sitting beside it that feels like the same question, the easy one gets answered and its answer is handed over under the hard question's name. The substitution of an easy question for a hard one is the shape of every heuristic. RepresentativenessJudging likelihood by how closely something resembles the type. is the first of the three original ones, set out by Daniel Kahneman and Amos Tversky in Subjective Probability, in Cognitive Psychology, in 1972.
What does representativeness substitute, exactly?
The whole subject is contained in the gap between the two questions, so start with them side by side. The hard question is: how likely is it that this process is a good one? Answering that properly needs a starting proportion, a count of observations, and some arithmetic about how far chance alone can carry a result. The easy question is: how much does this look like a good process? Answering that needs nothing but a mental picture and a glance.
Representativeness answers the second and reports it as an answer to the first. The substitution has a cost. Similarity has no units and no scale. Because a strong resemblance to the type looks exactly the same either way, similarity cannot distinguish four observations from four hundred. Similarity is silent about quantity. Quantity is precisely the information the probability question needed most. A run of four decisions that all went well resembles the record of somebody skilled just as strongly as a run of four hundred does, and the mind returns the same confident feeling for both.
The everyday version runs like this. At a wedding, somebody points out a guest in an expensive sherwani, speaking softly, carrying himself well. Asked whether he is more likely to be a surgeon or a shopkeeper, the picture answers instantly: surgeon. The picture is answering resemblance. The picture has not asked how many surgeons and how many shopkeepers were invited, and at a wedding of four hundred people, that count decides the answer. The stereotypeThe mental picture of a type that resemblance is measured against. is doing the work, and the counting never happens.
Representativeness answers an easy question in place of a hard one. Which easy question does it answer?
Why does judging by resemblance work at all?
No mind would carry a shortcut that was simply wrong. Representativeness survives because most of the time it is right, and it is worth being exact about why. Resemblance is a sound answer to a probability question in the situations where two conditions hold together. The type has to be reasonably common. Meeting one must not be a surprise in the first place. The features being matched on have to genuinely track the type. Things that look like it usually are it.
Both conditions hold constantly in ordinary life. Smoke in a stairwell means something is burning. Dark cloud gathering at four in the afternoon in June means rain within the hour. Hungry office workers are picky and there are other stalls, so a queue outside a food stall at one o'clock means the food is worth queueing for. In each case the resemblance is doing real work, and stopping to compute a probability would cost more than the shortcut ever could.
Representativeness is not a defect bolted onto an otherwise sound mind; it is a good rule applied outside the situations it was built for. The distinction matters for how representativeness should be treated. A rule that was always wrong could be switched off. A rule that is usually right cannot be, and never should be, so the work is learning to recognise the narrow situations where it stops holding.
What is the first failure, and how is its shape recognised?
The first failure is the one that follows straight from the wedding example. Representativeness ignores how common the thing was to begin with. The proper name for that is base-rate neglect, and its full arithmetic is set out under base-rate neglect. The two failures look alike from a distance and need completely different repairs, so recognising the shape of the first one matters.
The shape is this. Somebody has a mental picture of a type, meets a case that matches the picture, and concludes the case is probably of that type, without ever asking how many cases of that type exist to be met. The correction for that first failure is a count taken before the evidence arrives, not a count taken after it. The count of how many surgeons were invited has to be in hand before the guest in the sherwani walks past. Once he has walked past, the picture is already answering. Kahneman and Tversky took that failure apart in On the Psychology of Prediction, in Psychological Review, in 1973.
What is the law of small numbers?
The second failure is the one worked through in full below, and it is the more interesting of the two because it survives even after the first has been fixed. The law of small numbersWrongly expecting a short run to look like the long-run average. is the mistaken expectation that a short stretch of results will look like the long stretch it came from. The name is deliberately ironic. A real law of large numbers says that averages settle down as observations pile up, and the real law is true. The law of small numbers is the belief that the settling down has already happened after five or ten observations, and that belief is false.
The feeling from the inside is familiar. A coin tossed four times comes up heads four times, and something says the next one is somehow more likely to be tails. Four heads does not look like what a coin does. The feeling is representativeness reporting on the runA short consecutive sequence of results from the same process.: four heads does not resemble the mental picture of coin tossing, so it is registered as unusual. The coin, of course, has no memory. Four heads is an ordinary product of an ordinary coin, and the mental picture of what a coin does was never a picture of four tosses. It was a picture of a thousand.
The mistake is not about coins; it is about expecting the small sample to inherit the properties of the process that produced it. The same expectation set out by Kahneman and Tversky in Subjective Probability, in Cognitive Psychology, in 1972, is what makes a short record of decisions feel like a description of the decider. Four good decisions resemble the record of somebody good, and resemblance is exactly what the shortcut answers with. The shortcut never asks how likely that same run is from somebody entirely ordinary.
What does the law of small numbers describe?
Are the two failures really the same failure?
No, and this is worth pinning down before any arithmetic, because the two are constantly bundled together as though one correction covered both. The two failures differ in what they ignore, in what they look like, and in what fixes them. The first ignores the count that existed before any evidence arrived. The second ignores the count of the evidence itself. One is about prior frequency, the other about sample sizeHow many observations a conclusion rests on., and they are logically independent of one another.
Taking the two together in one situation shows the independence. Suppose how common careful deciders are is known exactly, so the first failure is fully corrected. The record in front of the reader is still a four decision record, and it still gives no way of telling a careful decider from a lucky one. Correcting for how common something is leaves the sample size problem exactly where it was, and a large sample leaves the starting proportion exactly where it was. Each needs its own repair, and a reader who has only made one of them has made half the journey.
Somebody corrects carefully for how common a type actually is before judging a case. Does that also fix the sample size problem?
How far can a run of results drift on chance alone?
Now the arithmetic, which is the part that turns an intuition into something checkable. Suppose a decision process is genuinely a coin: right half the time, no better, no worse, with each decision independent of the one before it. Over a stretch of it, the proportion that came out right is recorded. How far can that observed proportion sit from 50 per cent purely by chance?
The spreadHow far an observed rate can drift from the true one by chance alone. of an observed rate from an even process is the square root of 0.25 divided by the number of observations, and about 95 runs in every 100 land within twice that spread. Twice the square root of 0.25 divided by n works out, conveniently, to 100 divided by the square root of n, in percentage points. The band expression carries the whole argument, and its values are worth working through.
| Decisions watched | Spread of the observed rate | Band, plus or minus | An ordinary result runs from |
|---|---|---|---|
| 4 | 0.2500 | 50.0 points | 0.0 to 100.0 per cent |
| 5 | 0.2236 | 44.7 points | 5.3 to 94.7 per cent |
| 10 | 0.1581 | 31.6 points | 18.4 to 81.6 per cent |
| 25 | 0.1000 | 20.0 points | 30.0 to 70.0 per cent |
| 100 | 0.0500 | 10.0 points | 40.0 to 60.0 per cent |
| 500 | 0.0224 | 4.5 points | 45.5 to 54.5 per cent |
Read the first row again, because it is the one that stings. At four observations a process with no skill in it at all can produce any result from nothing right to everything right, and every one of those results is ordinary. At five observations the band is 44.7 points, so anything between about 5 and 95 per cent is what an even process routinely does. Nothing in that range is evidence of anything. The coin is an assumption rather than a finding. A process with exactly no skill in it is the benchmark an observed rate has to be measured against.
A decision process is genuinely no better than a coin. Over five decisions, how far can its observed success rate drift from 50 per cent by chance alone?
Lengthen the run and watch the band refuse to close
One variable moves: the number of decisions observed, from 5 to 500. Everything else is held still. The green funnel is the range an observed success rate can fall inside by chance alone, drawn around the 50 per cent line for a process with no skill in it whatever. The default sits at 25 decisions, where the band is plus or minus 20.0 points, so anything from 30.0 to 70.0 per cent is an entirely ordinary result. Moved to the right, the funnel narrows grudgingly: 5 decisions give plus or minus 44.7 points, 10 give 31.6, 25 give 20.0, 100 give 10.0, and 500 still leave 4.5. Quadrupling the number of decisions only halves the band, and that is why the curve falls away so slowly.
Watch 25 decisions from a process that is genuinely even and about 95 runs in 100 land between 30.0 and 70.0 per cent, so an observed rate anywhere inside that band is an ordinary result rather than a sign of skill.
Going from 25 decisions to 100 decisions quadruples the evidence. What happens to the band?
The quadruple-to-halve relationship is the thing most worth carrying away, because it sets how much patience the work requires. Uncertainty falls with the square root of the number of observations, not with the number itself. Four times as much evidence buys half as much uncertainty, and that exchange rate never improves.
Followed forward, that exchange rate gets uncomfortable. Pulling the band down to plus or minus 5.0 points takes 400 decisions. Reaching plus or minus 2.5 points takes 1,600. A band tight enough to separate a genuinely good process from an even one takes a number of decisions almost nobody ever accumulates. The Palash decision log, an invented record, holds 240 decisions in total across everybody in it, so even pooled it does not reach 400.
What does the Palash decision log actually settle?
The arithmetic applies directly to the case. The Palash decision log, kept by Palash Advisory Services Private Limited, holds 240 logged decisions taken by 60 investors across eight quarters. Dividing 240 by 60 gives 4 decisions each. Meera Sundaram, one of the 60, is a salaried investor of 41 with four positions running, and her own stretch of decisions is exactly that long.
Put 4 into the expression and the band comes out at plus or minus 50.0 points. At four decisions the range an even process routinely produces covers every possible result from 0 to 100 per cent, so the invented log cannot separate skill from chance for any single person in it, and no honest reading of it should try. None of that is a criticism of the record keeping. The limit is a fact about four, and it would be equally true of a record kept perfectly.
Meera Sundaram's own four positions on 30 September show it. The Vindhya index scheme cost Rs 3,00,000/- and stood at Rs 3,36,000/-, up 12.0 per cent. The Nilgiri mid-cap scheme cost Rs 3,00,000/- and stood at Rs 2,55,000/-, down 15.0 per cent. Suvarna Chemicals Limited cost Rs 4,00,000/- and stood at Rs 4,60,000/-, up 15.0 per cent. Kesari Logistics Limited cost Rs 3,00,000/- and stood at Rs 1,95,000/-, down 35.0 per cent. Two up, two down, on a cost of Rs 13,00,000/- now worth Rs 12,46,000/-. Two right out of four is 50.0 per cent, dead centre of a band running from 0 to 100, and it says precisely nothing about how Meera Sundaram decides.
Bring that same count down to one person's own stretch of decisions and it does not improve.
The Palash decision log holds 240 decisions across 60 investors. What can it establish about any single investor's skill?
How can a real pattern be distinguished from a short one?
The correction is not scepticism and it is not a habit of doubting records. Scepticism has no dial on it, so it talks a reader out of long records as easily as short ones. The correction is a computation, and it takes about ten seconds. The observations are counted. The band an even process would produce over that many follows from 100 divided by the square root of the count. The remaining question is whether the observed rate sits inside that band or outside it.
A result inside the band is not evidence of anything, and a result outside it is the first thing worth taking seriously. The order in which the steps run is the point. The count comes first and the impression comes last. The shortcut works the other way round, and that reversal is the whole repair. The reversal also explains why the repair feels effortful: a slow question is being asked before a fast one has finished answering.
One more thing belongs in the same check. Regression to the averageThe tendency for an extreme result to be followed by a less extreme one. means an extreme stretch tends to be followed by a less extreme one, purely because the extreme stretch was partly chance and chance does not repeat to order. So a record that sat at the top of a band last year will usually sit closer to the middle next year, with nothing at all having changed about the decider. Expecting that shift to continue in the same direction, or reading its reversal as somebody losing their touch, is representativeness again in a different coat.
The check itself is short enough to hold in mind, and it runs as three steps and two branches.
Where is judging by resemblance the correct move?
Stated without hedging: judging by resemblance is the right move whenever the type being matched against is common and the features being matched on genuinely track it. Resemblance covers most of ordinary recognition, most of the time, and treating the shortcut as a flaw to be suppressed would leave a person slower and no more accurate at almost everything done in a day.
Resemblance is also the right move when the alternative is not available. If nobody has counted how common the type is, and nobody can, then a resemblance judgement openly labelled as a resemblance judgement is a reasonable thing to work with. The error is not using a sound shortcut; the error is forgetting that it was used and reporting the output as though a probability had been computed. Devika Rao, the adviser at Palash Advisory Services Private Limited, can look at a decision and think it resembles a careful one. The moment that becomes a statement about how likely the decider is to be careful, an unstated count has been smuggled in.
When is judging by resemblance the right move?
How does anybody actually use this in practice?
Take the professional case first, because it is the one where the count is usually available. Somebody assessing a decision record for another person, an analyst reading a manager's history or an adviser reviewing a client's own log, has a first question that costs nothing: how many decisions is this, exactly? Not how many years, not how many holdings, but how many separate decisions. Eleven years of a record with three decisions in it is a three decision record. The number of observations, and never the length of the story built on them, is what a record can support.
Then the same person computes the band and states it out loud beside the record. A record of 62 per cent from 25 decisions sits inside a band running from 30.0 to 70.0 per cent, so the honest sentence is that the record is consistent with skill and equally consistent with an even process, and 25 decisions cannot tell the two apart. The sentence is neither flattering nor hostile; it is a statement about the count.
The household version needs no professional at all. A person deciding alone, with no adviser and no committee, still keeps a mental record and still draws conclusions from it. The useful habit is to write down how many decisions the conclusion rests on before writing the conclusion. Meera Sundaram, deciding on her own account, has four. In the log, a written reason was recorded on 84 of the 240 decisions, or 35.0 per cent, so most decisions in it were never even written down to be counted later. Counting first is cheap, it needs no arithmetic beyond a square root, and it is the only step that reliably stops the resemblance answer from being reported as a probability.
The error that gets made, and what it costs
The failure is treating a short stretch of results as a description of a process. Four decisions that went well resemble the record of somebody good at this, and resemblance is exactly what the shortcut answers with. The shortcut never asks how likely that run is from somebody entirely ordinary, and at four decisions such a run is very likely indeed.
The part almost everybody misses is that the same arithmetic runs in both directions. A short poor stretch is exactly as uninformative as a short good one. Somebody who concludes from four disappointing decisions that they are simply not built for this has made precisely the error of somebody who concludes from four pleasing ones that they have a gift. The shortcut does not have an optimistic setting and a pessimistic setting. The shortcut has one setting, resemblance, and it is silent about how many observations there were in either direction.
The cost is not usually money in a single stroke. The cost is a conclusion about a person, held with a confidence the evidence cannot support, that then shapes every decision after it. A person who has decided at four observations that they are bad at this stops deciding, and a person who has decided at four that they are gifted stops checking. Both conclusions were drawn from a band running from 0 to 100 per cent.
Somebody concludes from four poor decisions that they are simply bad at this. What error is that?
Sources
| Source | Document | Site |
|---|---|---|
| Daniel Kahneman and Amos Tversky | Subjective Probability, the 1972 paper in Cognitive Psychology in which representativeness is set out and small samples are expected to resemble the process behind them | ssrn.com |
| Amos Tversky and Daniel Kahneman | Judgment under Uncertainty: Heuristics and Biases, Science, 1974, the paper naming the three original heuristics together | ssrn.com |
| Daniel Kahneman and Amos Tversky | On the Psychology of Prediction, Psychological Review, 1973, the paper taking apart the neglect of how common a type is | ssrn.com |
| Amos Tversky and Daniel Kahneman | the 1973 paper in Cognitive Psychology setting out the availability shortcut | ssrn.com |
| Working paper repository | where the papers above are findable in preprint or reprint form | nber.org |
Meera Sundaram, Devika Rao, Palash Advisory Services Private Limited, the Palash decision log, the Vindhya index scheme, the Nilgiri mid-cap scheme, Suvarna Chemicals Limited and Kesari Logistics Limited are invented.
Educational material. Not advice on any investment, tax, budget or market position.
