Debiasing: Checklists, Outside View, Red Teams and Base Rates
Debiasing is the attempt to reduce a known reasoning error. The uncomfortable finding is that most attempts fail: telling somebody about a bias, including telling them accurately and at length, usually leaves the bias intact. Structure works better. A procedure forces a different question, where an intention only asks for a better one.
The reading so far has supplied the names of things. Anchoring, the disposition effect, herding, the pull of a recent quarter, the way a loss feels heavier than the same gain. Naming looks like the cure. A person who can define anchoring should, on that expectation, be somebody who no longer anchors. The expectation that naming is the cure is the single most common mistake made about this whole subject, and the evidence has been against it for forty years. What to do instead is narrower and less satisfying than a reader would like.
Everything below rests on a distinction the earlier subjects never had to make. There is a difference between knowing that an error exists and being routed away from it at the moment it happens. Knowledge is a thing a person holds. Routing is a thing that happens to that person whether or not anybody is paying attention. The earlier reading supplied the first. Everything below is the second, and the reason the two have to be separated is that the second does not follow from the first, however carefully the first was taught.
What is debiasing, and how much of it actually works?
DebiasingAn attempt to reduce a known reasoning error. covers every deliberate attempt to make a known reasoning error smaller. Debiasing includes explaining the error, warning about it in advance, paying people to avoid it, training them on worked cases, giving them feedback afterwards, and rearranging the task so the error has nowhere to occur. Six different attempts sit in that list, and lumping them together as awareness is what produces the disappointment.
Baruch Fischhoff surveyed the attempts in a chapter titled Debiasing, published in Judgment under Uncertainty in 1982, and the survey is not encouraging reading for anybody who has just learned a list of biases. Most of the interventions tested were ineffective. Warning people that a number they had been shown would pull their estimate did not stop the estimate being pulled. Somebody could be told exactly how a known outcome makes the past look more predictable than it was. They still reconstructed the past as more predictable than it was. The knowledge was present and the correction was absent, in the same person, at the same time.
The failure is not a failure of effort, and the reason is worth seeing. A bias is not a belief a person holds and could therefore drop. A bias is closer to the way a judgement gets assembled: what comes to mind first, what feels representative, what the starting number was. By the time the conclusion arrives it already has the shape the error gave it, and the person inspecting it is inspecting the finished product rather than the assembly. Nobody can audit a process they did not watch, and nobody watches their own.
Take it out of finance for a moment. Somebody is told that the first price quoted in a negotiation will pull whatever is countered with, and accepts this completely. The same person then walks into a shop where the seller opens at Rs 40,000/- for a table they had loosely valued at Rs 25,000/-. The counter comes out at Rs 30,000/- and feels reasonable while it is being said. The warning was not forgotten. The warning simply had nothing to attach to. At no point did a signal arrive saying the pull is happening now. The missing signal is the whole problem, and every method below is an attempt to work around it.
Does telling somebody about a bias usually remove it?
Why does knowing about a bias correct so little?
Three reasons, and each one suggests a different repair. Separating them is better than shrugging at the finding as a whole.
The first is timing. A bias operates while the judgement is being made, and awareness is a thing a person has between judgements. Asked the day after whether the television segment influenced the purchase, somebody will consider the question seriously. Asked at nine in the evening with the segment still running, the question does not arise at all. Nothing in the situation raises it. The moment a correction is most needed is the moment it is least likely to be summoned, and no amount of prior explanation changes that ordering.
The second is that the output of a biased process feels exactly like the output of an unbiased one. There is no signal. A conclusion reached by anchoring arrives with the same confidence as a conclusion reached by working through the evidence, and it comes with reasons attached. Reasons are generated after the conclusion at least as often as before it. So a person told to check their reasoning checks the reasons that were produced to support the conclusion, finds them sound, and stops. The check ran and it could not have failed.
The third is that the reader has been given a list of roughly twenty named errors and no way of knowing which one is live. Even a person who has genuinely learned all of them faces a search problem at the moment of decision: which of the twenty applies here? Searching a list of twenty under time pressure, using the same reasoning that is currently compromised, is not a promising design. A method that requires the decider to correctly diagnose their own error before correcting it has already asked for the hardest part.
Which technique has the best record, and what does it ask of the decider?
One technique comes out of the testing better than the rest, and it is almost aggressively simple. Considering the oppositeActively generating reasons a conclusion might be wrong. means, having reached a conclusion, actively generating the reasons it might be wrong before doing anything else. Not an internal query about whether the conclusion might be wrong. Such a query produces a brief nod and no content. Generating the reasons, out loud or on paper, until there are some.
Charles Lord, Mark Lepper and Elizabeth Preston tested this in the Journal of Personality and Social Psychology in 1984. The instruction that worked was not a request to be fair or balanced. A request for balance changed little. The direction that worked was specific: consider the opposite, ask what the evidence would look like if it pointed the other way, and produce that content rather than merely acknowledge its possibility.
The mechanism is worth stating because it explains why the vaguer instruction fails. Once a conclusion is reached, the material that comes to mind afterwards is the material that fits it. An instruction to be balanced does not change what comes to mind; it only changes how the material feels once it has arrived. Asking for reasons the conclusion is wrong changes the retrieval instruction, so different material arrives. The technique works on what gets fetched from memory, not on how carefully the material already fetched is weighed.
Considering the opposite has something in common with everything that follows. The instruction does not ask for less bias. It asks for a specific action whose output is visible. Whether three contrary reasons were generated is knowable, in the way that whether a search was fair is not. Every method worth the time replaces a state of mind with an action that leaves a trace.
What do the techniques with the better record have in common?
How to Build a Behavioural-Bias Checklist that gets used rather than filed?
A checklist is the plainest way to move a correction into the path of the decision. Atul Gawande set out the discipline in The Checklist Manifesto in 2009. Everybody already believed that lists are useful. The important part of his argument is that the good ones are short, sit at a defined pause, and ask for observable facts rather than for judgement.
Most behavioural checklists fail because they are written as a list of biases. A line reading am I anchoring cannot be answered. The line asks the decider to diagnose their own reasoning, and that capacity is the one in question. A line reading what number did I see first, and write it down, can be answered by anybody, and once the number is written down the anchoring question becomes visible without anyone having to name it.
So the construction rule is this. Each line names an observable, asks for something written, and can be answered wrongly. If a line cannot be answered wrongly it is decoration. Here is a six line version, and every line satisfies that test.
| Line | What it asks for | What it catches |
|---|---|---|
| 1 | Write the first number seen about this, and where it came from | a starting point doing work nobody has looked at |
| 2 | Write what is expected to happen, in a form that could turn out false | a conclusion too vague to be checked later |
| 3 | Write the group of comparable past cases and what happened to them | a judgement made only from the details of the case in hand |
| 4 | Write three reasons this could be wrong, generated after the conclusion | the search that never went looking for contrary material |
| 5 | Write what would have to be seen to change the conclusion, before acting | a position that cannot be disturbed by any future evidence |
| 6 | Write the date, and the earliest date on which action may be taken | the decision taken inside the hour it was first considered |
Now the part that matters more than the wording. A checklist has to sit at a pause that already exists, or one has to be created for it. In a practice, that pause is the moment before an instruction is sent. For somebody deciding alone, the natural pause is the moment the conclusion is reached and before anything is acted on. The sixth line puts it exactly there. A list that lives in a folder and is consulted when somebody remembers has been converted back into an intention, whose record is the one already described.
A procedure that only works on a holding has been written too narrowly, so run the same six lines on a decision with no money in it. A household is choosing a school for a nine year old. Line one: the first number seen was a fee of Rs 1,80,000/- a year, quoted by the first school visited, and every later school has been read against it. Line two: the expectation is that the child will settle within one term, stated so it could turn out false. Line three: the comparable cases are the four children on the same street who changed schools in the last three years, and two of them took longer than a year. Line four: three reasons this could be wrong, including that the visit was on a sports day. Line five: what would change the mind, namely the child being unhappy at the end of a second term. Line six: the decision date, and the earliest action date a week later. Nothing in those six lines is financial, and passing that test is what makes a procedure worth carrying into a decision that is.
Did the checklist change anything measurable in the invented Palash decision log? One thing, and only one. Twenty of the sixty investors adopted a written checklist on 4 November. Across quarters five to eight those twenty recorded a written reason on 34 of 41 decisions, or 82.9 per cent, against 19 of 63 decisions, or 30.2 per cent, for the other forty. Their realisation ratioHow readily gains were closed compared with how readily losses were, after allowing for how many of each were available to close. fell from 3.2 to 1.6, meaning gains were still closed more readily than losses but by half as much.
The log records no difference in return, and nobody should read one into those figures. There is one stretch of time, no comparison group assembled in advance, and no measurement of returns before or after. The measured effect is that reasoning became inspectable. Inspectable reasoning is a smaller claim than most readers want, and it is the only one the record supports.
What is the outside view, and where do its numbers come from?
Daniel Kahneman and Amos Tversky set out the correction in Intuitive Prediction: Biases and Corrective Procedures in 1979, and it is the most transferable idea in the whole subject. A case can be judged from its own details, which is the inside view, or from the record of cases like it, which is the outside viewJudging a case by the record of similar cases rather than its own details.. Almost everybody, almost always, does the first, and the second gives the better forecast in nearly every setting where both have been compared.
The reason is not that details are useless. Details are compelling out of proportion to what they predict. A plan that has been thought through carefully feels different from a plan that has not, and that feeling is real, and it is a poor guide to how long the plan will take. The record of comparable plans is a duller input and a better one.
The same thing happens in a household. Somebody says the kitchen work will take three weeks, and the estimate is built from the tasks: two days for the plumbing, four for the tiling, and so on, summed carefully. The outside view asks a different question. How long did the last four kitchens on this street take? Six weeks, nine weeks, seven weeks and five weeks. The inside view produced a number from the plan. The outside view produced a number from the record, and the record is not impressed by how carefully the plan was made.
The two inputs the outside view needs are a reference classThe set of comparable cases a base rate is drawn from., meaning the set of comparable past cases, and a base rateHow often something happens across a reference class., meaning how often the thing in question happened across that set. Everything hard about the technique is in getting those two, and the next section is entirely about that difficulty.
Take the outside-view question that matters most to somebody deciding about holdings: does acting more improve what somebody picks? The invented Palash decision log has a reference class for it. Sixty investors sit in five turnover groups of twelve each, with annual turnover of 9, 34, 71, 128 and 210 per cent. Gross returns across those five groups ran 11.2, 11.0, 11.1, 10.9 and 11.0 per cent. Costs ran 0.3, 0.6, 1.5, 2.5 and 4.1 points, and net returns therefore ran 10.9, 10.4, 9.6, 8.4 and 6.9 per cent.
Gross returns sit inside 0.3 points of each other across a range of activity that varies more than twenty-threefold. Net returns run 4.0 points apart. The base-rate answer to the question is that in this cohort, acting more did not improve what was picked. Acting more changed only what the acting cost.
The figure above carries the whole argument for taking an outside view at all. Any one of those sixty investors, asked why they traded, would give an inside-view answer built from the details of each decision. The record says the details did not matter to what was picked. The details mattered a great deal to what was kept.
How to Use Base Rates in a Forecasting Exercise, and which question comes second?
The procedure is two questions in a fixed order, and the second one is the one that does the work.
Question one: what is the reference class? The set of comparable past cases this one belongs to has to be named, and the boundary has to be specific. A class drawn too wide stops being comparable. A class drawn too narrow has one member, the case in hand.
Question two: were the outcomes of that class actually recorded? Not could they have been, not does somebody probably know. Were they written down somewhere a reader can find them. If the answer is no, the forecast has no base rate, and the honest output of the procedure is to say so rather than to estimate one from impression.
Most people run the first question, feel that they have taken an outside view, and never run the second. The result is worse than the inside view it replaced. A number produced from a class whose outcomes nobody recorded carries the authority of a statistic and the content of a guess.
What is the second question in the base-rate procedure?
What do an available base rate and an unavailable one look like side by side?
A method taught only on the easy case does not survive practice, so here are both halves worked on the same invented record. The Palash decision log holds 240 decisions taken by 60 investors over eight quarters, four decisions each: 96 buys, 84 sells, 36 switches and 24 pauses of a standing instruction, and those four counts sum to 240.
The base rate that is available
The question is the one already asked: does acting more improve what somebody picks? The reference class is the five turnover groups, twelve investors each. Gross returns were computed for every group, so the outcomes are recorded. Reading the rate off the record gives 11.2, 11.0, 11.1, 10.9 and 11.0 per cent, a span of 0.3 points across turnover running from 9 to 210 per cent. The second question passes, so the forecast has a base rate, and the answer is that in this cohort more activity did not go with better picking.
The base rate that is not available
Now a question that sounds just as answerable. Do names mentioned in the media do well? The log has something on the subject. Of the 96 buys, 41 followed a media mention within three days, or 42.7 per cent, against 11.0 per cent of the eligible list being mentioned at all in a given week. The gap is large and it is real.
But it is a base rate for attention, not a base rate for outcomes. The figure records how often a mention preceded a purchase. The log recorded no result for mentioned names as a group, so the figure says nothing whatever about what happened to them afterwards. Question one passes: the class is nameable, being the eligible list of holdings that were mentioned in a given week. Question two fails: nobody wrote down how those holdings did. So the honest output is that this forecast has no base rate in this record, and any number produced for it would be manufactured.
The attention figure is worth drawing on its own. A careless reader will lift the 42.7 per cent out of this record and treat it as a finding about performance.
The takeaway a reader can carry into any subject is short. Name the reference class first. Then check whether its outcomes were recorded. If they were not, stop, and say the forecast has no base rate rather than producing one anyway. A procedure whose most valuable output is sometimes the word no is doing something a general instruction to think harder cannot do.
Why can the invented log not settle whether media-mentioned names do well?
Does widening the reference class make the individual case clearer?
The control below is built around this question, and it is worth settling before anything moves.
Before the control moves: does widening the reference class make the individual case clearer?
Widen the reference class and watch what actually gets firmer
One variable moves: how many investors sit in the reference class, from a single person up to all 60. Everything else is held. The invented Palash decision log records 240 decisions by 60 investors over eight quarters, four decisions each, so a class of one person contributes 4 observations, one turnover group of twelve contributes 48, and the whole cohort contributes 240. The rate being estimated is held at 35.0 per cent, the share of the 240 that carried a written reason, being 84 of 240. The five turnover groups ran annual turnover of 9, 34, 71, 128 and 210 per cent with gross returns of 11.2, 11.0, 11.1, 10.9 and 11.0 per cent, a span of 0.3 points, and those group figures only become worth quoting at the full 240. Widening the class answers a different question rather than the same question better: 240 decisions say something about the sixty and nothing about any member of them.
With all 60 investors in the class the record contributes 240 decisions, so the recorded-reason rate of 35.0 per cent carries a plausible band of 28.8 to 41.2 per cent, a width of 12.3 points. That is a statement about the sixty and not about any one of them.
Two things are worth noticing as the control moves. The band narrows steeply at first and then slowly, so the gain from going from four decisions to forty-eight is enormous and the gain from forty-eight to two hundred and forty is modest. And at every setting, the thing getting firmer is the group figure. Nothing that happens to the band says anything more about the single investor the case began with, and that is exactly what makes the outside view a different question rather than a sharper answer.
Red Team: what does one do that a reviewer does not?
A red teamA group tasked with defeating a case rather than improving it. is defined by its instruction, not by its expertise. Hand the same document to two people. Ask the first to review it, and they will look for weaknesses in order to strengthen it. Reviewing means exactly that. Ask the second to defeat it, and they will look for the one weakness that ends it. Defeating means exactly that. Same document, opposite instruction, and the two come back with different things.
The wording of the instruction matters, because a reviewer is, structurally, on the side of the case. Improving a case presumes the case survives improvement. A reviewer who found that the whole thing was wrong would have done something outside the job as given. Nobody has to be timid or political for this to happen; it follows from the wording of the task.
Three conditions make the difference real rather than theatrical. The instruction is written down, so nobody has to guess how adversarial to be. The output is a written objection rather than a conversation, so it survives the meeting. And whoever holds the case has to answer the objection in writing. Most arrangements skip that part, and an objection nobody answered is indistinguishable from an objection nobody made.
What separates a red team from a reviewer?
How to Use a Devil’s-Advocate Review in Research, and who argues what?
The devil-advocate review is the small, repeatable version of the same idea, sized for a person or a two-person practice rather than for a committee. The review has four parts and takes about twenty minutes.
First, the seat is named before the work starts, not after. Somebody is the objector for this note, and they know it in advance. A person asked to object on the spot will object politely and about nothing important. Second, the objector gets a written instruction with a single target: produce the strongest reason this conclusion is wrong, and produce it as a claim that could be checked. Third, the objection goes on paper. Fourth, the author answers it on paper, and the answer stays with the note.
The fourth part is the one that decides whether any of this was worth doing. An objection that was raised, nodded at and forgotten has left no trace, and six months later nobody can tell whether it was answered or ignored. An objection with a written answer beside it lets a later reader see the reasoning that was actually applied, and that is the entire purpose of the exercise.
For somebody working alone with nobody to hand the objector seat to, the substitute is time rather than a person. The conclusion is written and then left. Later the strongest available objection is written out, as though somebody else had produced it, before the original reasons are reread. The substitute is weaker than a second person and much better than nothing. The retrieval instruction has still been changed.
How to Check for Confirmation Bias in a Research Note, and in what order?
Confirmation bias is not usually visible in a conclusion. Confirmation bias is visible in the sources, and specifically in what is missing from them. So the check is a sweep of the source list, run in a fixed order. Run in any other order, the sweep collapses into an impression about whether the note felt fair.
Step one is to list every source the note used, without judging any of them. Step two is to mark each one as supporting the conclusion, cutting against it, or neutral. Step three, the sharp one, is to describe what a source arguing the other way would look like, concretely enough that it would be recognised. Step four is to ask whether such a source was looked for, and whether the search that would have found it was ever run.
A note with nine supporting sources and a note with two can be equally sound or equally selective, so the tally of supporting sources settles almost nothing. What settles something is step four. A note that names the contrary source it went looking for and could not find is in a different condition from a note whose search was never pointed that way, even where both end with the same nine supporting citations.
What is the sharpest question to ask of a research note's sources?
How to Build a Personal Research Pause Protocol that actually gets kept?
A pause protocolA fixed delay between reaching a conclusion and acting on it. is a fixed delay between reaching a conclusion and acting on it, with a fixed thing done inside the delay. A pause protocol is not a general resolution to slow down, an intention wearing a procedure’s coat, and not a rule about being calm.
Three settings define it, and all three are set once, in advance, when nothing is happening. The length of the wait, stated in hours or days rather than in feelings. The trigger, meaning which decisions the wait applies to. A protocol that applies to everything gets dropped in the first busy week. And the reread, meaning the specific thing done before acting, normally rereading the written reasons and checking whether they still say what they were thought to say.
A reasonable starting shape for a person deciding alone is this: any decision arriving within forty-eight hours of a news item waits until the following day, and before acting the written reason is reread and confirmed not to consist mainly of the news item. The forty-eight hour trigger is not arbitrary. In the invented Palash log, 71 of the 240 decisions were taken within 48 hours of a news item, or 29.6 per cent, so almost a third of the record would pass through that gate.
Look at what such a gate would have touched in the invented record, and be careful about what that means. On 19 February a television segment named Suvarna Chemicals Limited and Meera Sundaram added Rs 1,00,000/- to that holding the same evening, taking its cost to Rs 4,00,000/- and the total cost to Rs 13,00,000/-. A pause protocol would have moved that decision to the following day. The wait would not have told her whether the decision was right, and nothing in this record shows that waiting would have produced a better result. What the wait produces is a written reason read twice, once while the segment was running and once when it was not, and a reader who can see both.
What does none of this fix?
The procedures leave four substantial things untouched.
First, none of it fixes a bad reference class. The procedures force a class to be named; they cannot say that the class named was wrong. A person who compares this decision to the four most memorable past cases has run the outside view correctly on a class assembled by memory, and memory is a biased sampler. The written class helps a later reader spot the problem. Spotting is not preventing.
Second, none of it fixes the thing nobody has thought of. A checklist covers the errors somebody wrote down. A red team objects to the case as presented. Both operate inside the frame the work already has. Genuine surprises live outside it.
Third, none of it fixes incentives. If somebody is paid for the conclusion, an objector seat and a written answer produce a well-documented version of the conclusion they were paid to reach. Procedure and interest are separate things, and procedure is the weaker of the two.
Fourth, and most importantly for a reader who has come this far, none of it has been shown to improve returns, on this record or anywhere in this sequence. What the twenty who adopted the checklist demonstrably produced was written reasons on 82.9 per cent of their decisions instead of 30.2 per cent, and a realisation ratio that fell from 3.2 to 1.6. Both figures are changes in how decisions were recorded and in which positions were closed. Whether it made them better off is not measured, and cannot be measured from one stretch of eight quarters with no comparison group.
Do these procedures improve returns?
The error that gets made, and what it costs
The error is treating debiasing as a matter of vigilance. Vigilance sounds responsible, it flatters the reader who has just learned a long list of mechanisms, and it is what the evidence contradicts. Fischhoff's survey did not find that people were not trying hard enough. The survey found that trying, in the form of knowing and intending, mostly did not move the outcome.
The reason is a timing problem that no amount of resolve fixes. An intention is available in the calm hour when it is formed and unavailable in the loud minute when it is needed. The same conditions that produce the error also suppress the search for the correction. A written step does not have this property. A written step sits in the path, fires whether or not anybody thought of it, and leaves an output that somebody who was not in the room can read afterwards.
The error costs the wrong repair. A person who believes debiasing is vigilance responds to a bad decision by resolving to be more careful, and a resolution changes nothing measurable. They do not build the six-line list or set the wait, and both of those change what gets recorded. The whole distinction is between a resolution, whose only evidence is a feeling, and a procedure, whose evidence is an artefact.
How does somebody actually run this, alone or on behalf of others?
Devika Rao, the adviser at the invented Palash Advisory Services Private Limited, does not run six procedures on every decision. A practice that did would stop functioning inside a month. She runs a trigger. Any instruction arriving within two days of a news item, or any instruction that closes a position at a round number matching its cost, goes through the six-line list before it is sent, and the objector seat is used on the three or four research notes a quarter that carry the most weight.
The lone decider runs the same shape with the seats collapsed. The trigger is the same, the list is the same six lines, the objector seat becomes a stated wait plus the strongest objection written out the next morning, and the source sweep runs on whatever was read before deciding, even where that is three articles rather than a research note.
The common design in both is that the procedure is attached to a trigger rather than to a good intention, and it produces an artefact rather than a feeling. A lender reading a credit file, an analyst signing a note and a household deciding on a school are all doing the same thing: making the reasoning visible to somebody who was not there, including their own later self.
One last figure, on the only measured movement in the invented record, drawn so that nobody can read more into it than it holds. The realisation ratio for the twenty who adopted the list fell from 3.2 to 1.6, meaning gains were still closed more readily than losses but by half as much. The fall is a change in which positions were closed. The ratio is not a return.
Sources
| Source | Document | Site |
|---|---|---|
| Baruch Fischhoff | the survey chapter titled Debiasing, in Judgment under Uncertainty, 1982 | ssrn.com |
| Charles Lord, Mark Lepper and Elizabeth Preston | the paper testing the instruction to consider the opposite, Journal of Personality and Social Psychology, 1984 | ssrn.com |
| Daniel Kahneman and Amos Tversky | Intuitive Prediction: Biases and Corrective Procedures, 1979, where the outside view is set out | nber.org |
| Amos Tversky and Daniel Kahneman | Judgment under Uncertainty: Heuristics and Biases, Science, 1974 | ssrn.com |
| Atul Gawande | The Checklist Manifesto, 2009, on how a short list is built and where it sits | cited to the book itself |
| Karl Popper | The Logic of Scientific Discovery, 1934, for the test a claim has to be able to fail | cited to the book itself |
| J Edward Russo and Paul Schoemaker | Decision Traps, 1989, on keeping the written record of a decision | cited to the book itself |
| Securities and Exchange Board of India | conduct and suitability duties applying to registered intermediaries | sebi.gov.in |
| Association of Mutual Funds in India | investor-facing practice material for distributors and advisers | amfiindia.com |
Meera Sundaram, Devika Rao, Palash Advisory Services Private Limited, the Palash decision log, the Palash 100 index, Suvarna Chemicals Limited and Kesari Logistics Limited are invented.
Educational material. Not advice on any investment, tax, budget or market position.
