Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Behavioural Finance & Investor Decision-Making
1Foundations
The Rational InvestorJudgment Under UncertaintyPreferencesBehavioural FinanceInvestor and Market BehaviourFinancial Well-BeingBounded RationalityHeuristics and Biases
2Cognitive Biases, Emotion and Attention
Limited AttentionRepresentativenessThe Affect HeuristicAnchoring and AdjustmentEmotion and Decision QualityOverconfidence and OptimismAmbiguity and Complexity AversionAvailability and SalienceHome Bias, Local Bias…FramingThe Halo EffectHindsight BiasThe Narrative FallacyPresent Bias and Hyperbolic DiscountingBase-Rate NeglectStatus Quo Bias and the Default Effect
3Preferences and Prospect Theory
Prospect TheoryRegretThe Endowment EffectMental AccountingThe Sunk Cost FallacyLoss AversionRisk Seeking in Losses
4Social Behaviour
HerdingNarrative EconomicsFear of Missing OutGroupthinkSocial Proof
5Investment and Trading Behaviour
Excess TradingNaive DiversificationThe Disposition EffectLottery PreferencesNoise TradersPortfolio InertiaRecency Bias
6Markets and Anomalies
Mania, Panic and CapitulationMarket EfficiencyEfficient Market Hypothesis vs…Speculative BubblesReflexivityInvestor SentimentMarket AnomaliesShort-Sale ConstraintsPrice DiscoveryLimits to Arbitrage
7Decision, Research and Debiasing
The Decision JournalDebiasingChoice Architecture, Defaults and…The Pre-Mortem and Process QualityDecision Quality
8Advice, Conduct and Communication
Communication ConductSuitability and AppropriatenessChoice OverloadComplaint BehaviourRisk DisclosureVulnerable Investors

Debiasing: Checklists, Outside View, Red Teams and Base Rates

Debiasing is the attempt to reduce a known reasoning error. The uncomfortable finding is that most attempts fail: telling somebody about a bias, including telling them accurately and at length, usually leaves the bias intact. Structure works better. A procedure forces a different question, where an intention only asks for a better one.

The reading so far has supplied the names of things. Anchoring, the disposition effect, herding, the pull of a recent quarter, the way a loss feels heavier than the same gain. Naming looks like the cure. A person who can define anchoring should, on that expectation, be somebody who no longer anchors. The expectation that naming is the cure is the single most common mistake made about this whole subject, and the evidence has been against it for forty years. What to do instead is narrower and less satisfying than a reader would like.

Everything below rests on a distinction the earlier subjects never had to make. There is a difference between knowing that an error exists and being routed away from it at the moment it happens. Knowledge is a thing a person holds. Routing is a thing that happens to that person whether or not anybody is paying attention. The earlier reading supplied the first. Everything below is the second, and the reason the two have to be separated is that the second does not follow from the first, however carefully the first was taught.

What is debiasing, and how much of it actually works?

DebiasingAn attempt to reduce a known reasoning error. covers every deliberate attempt to make a known reasoning error smaller. Debiasing includes explaining the error, warning about it in advance, paying people to avoid it, training them on worked cases, giving them feedback afterwards, and rearranging the task so the error has nowhere to occur. Six different attempts sit in that list, and lumping them together as awareness is what produces the disappointment.

Baruch Fischhoff surveyed the attempts in a chapter titled Debiasing, published in Judgment under Uncertainty in 1982, and the survey is not encouraging reading for anybody who has just learned a list of biases. Most of the interventions tested were ineffective. Warning people that a number they had been shown would pull their estimate did not stop the estimate being pulled. Somebody could be told exactly how a known outcome makes the past look more predictable than it was. They still reconstructed the past as more predictable than it was. The knowledge was present and the correction was absent, in the same person, at the same time.

The warning arrives, and the estimate lands in the same place. STAGE ONE The error is explained, carefully and at length. The listener follows it. STAGE TWO A number appears in the room shortly before a judgement is asked for. STAGE THREE The judgement settles near that number, with the warning still fresh. HOLDING THE NAME OF AN ERROR IS NOT THE SAME EVENT AS BEING ROUTED PAST IT Stage one is what the earlier reading gave. Stage three is what the survey work kept finding anyway. None of this is measured on the invented cohort used later; it is the shape of the finding, not a figure.
Explaining an error and then watching the same error occur is the ordinary result rather than a curiosity, and that is the reason a correction belongs in the task rather than in the person.

The failure is not a failure of effort, and the reason is worth seeing. A bias is not a belief a person holds and could therefore drop. A bias is closer to the way a judgement gets assembled: what comes to mind first, what feels representative, what the starting number was. By the time the conclusion arrives it already has the shape the error gave it, and the person inspecting it is inspecting the finished product rather than the assembly. Nobody can audit a process they did not watch, and nobody watches their own.

Take it out of finance for a moment. Somebody is told that the first price quoted in a negotiation will pull whatever is countered with, and accepts this completely. The same person then walks into a shop where the seller opens at Rs 40,000/- for a table they had loosely valued at Rs 25,000/-. The counter comes out at Rs 30,000/- and feels reasonable while it is being said. The warning was not forgotten. The warning simply had nothing to attach to. At no point did a signal arrive saying the pull is happening now. The missing signal is the whole problem, and every method below is an attempt to work around it.

Six different things get called debiasing. They are not equally hopeful. THE ATTEMPT WHAT IT ASKS OF THE DECIDER 1. Explain the error to them notice it happening, unaided 2. Warn them just before the task notice it happening, unaided 3. Pay them for a better answer try harder at the same method 4. Train them on worked cases carry the training across 5. Make them argue the other side follow an instruction, once 6. Rebuild the task around it nothing at the moment itself The two shaded rows are the ones with a defensible record. Read down the right column to see why.
Ranking the six attempts by what each one demands of the person at the moment of decision explains the pattern in the evidence better than ranking them by how clever they sound.
Try it out

Does telling somebody about a bias usually remove it?

Financial Literacy Bootcamp — Fin Maverick

Why does knowing about a bias correct so little?

Three reasons, and each one suggests a different repair. Separating them is better than shrugging at the finding as a whole.

The first is timing. A bias operates while the judgement is being made, and awareness is a thing a person has between judgements. Asked the day after whether the television segment influenced the purchase, somebody will consider the question seriously. Asked at nine in the evening with the segment still running, the question does not arise at all. Nothing in the situation raises it. The moment a correction is most needed is the moment it is least likely to be summoned, and no amount of prior explanation changes that ordering.

The second is that the output of a biased process feels exactly like the output of an unbiased one. There is no signal. A conclusion reached by anchoring arrives with the same confidence as a conclusion reached by working through the evidence, and it comes with reasons attached. Reasons are generated after the conclusion at least as often as before it. So a person told to check their reasoning checks the reasons that were produced to support the conclusion, finds them sound, and stops. The check ran and it could not have failed.

The third is that the reader has been given a list of roughly twenty named errors and no way of knowing which one is live. Even a person who has genuinely learned all of them faces a search problem at the moment of decision: which of the twenty applies here? Searching a list of twenty under time pressure, using the same reasoning that is currently compromised, is not a promising design. A method that requires the decider to correctly diagnose their own error before correcting it has already asked for the hardest part.

Twenty of these have been named. Which one is happening right now? Anchoring Availability Confirmation Hindsight Overconfidence Recency Framing Loss aversion Mental accounts Status quo Regret Endowment Sunk cost Herding Social proof Narrow framing Home bias Disposition Present bias Ambiguity NONE OF THEM IS FLAGGED AT THE MOMENT IT MATTERS. THE RIGHT ONE MUST BE FOUND And it must be found using the same reasoning that is currently the thing under suspicion.
Twenty names with nothing marking which one is live turns every correction into a search problem run by the faculty that is already compromised.
Two ways to place a correction. Only one of them is present when it is needed. A JUDGEMENT IS ABOUT TO BE MADE ROUTE ONE, THE INTENTION it fires only if noticed, and noticing is what already failed ROUTE TWO, THE WRITTEN STEP it fires because it sits in the path, noticed or not Every method that follows is an attempt to move a correction from the upper route to the lower one.
The difference between a resolution and a procedure is whether the correction depends on the decider noticing, which is precisely the capacity the error has already taken away.
Portfolio Management Bootcamp — Fin Maverick

Which technique has the best record, and what does it ask of the decider?

One technique comes out of the testing better than the rest, and it is almost aggressively simple. Considering the oppositeActively generating reasons a conclusion might be wrong. means, having reached a conclusion, actively generating the reasons it might be wrong before doing anything else. Not an internal query about whether the conclusion might be wrong. Such a query produces a brief nod and no content. Generating the reasons, out loud or on paper, until there are some.

Charles Lord, Mark Lepper and Elizabeth Preston tested this in the Journal of Personality and Social Psychology in 1984. The instruction that worked was not a request to be fair or balanced. A request for balance changed little. The direction that worked was specific: consider the opposite, ask what the evidence would look like if it pointed the other way, and produce that content rather than merely acknowledge its possibility.

The mechanism is worth stating because it explains why the vaguer instruction fails. Once a conclusion is reached, the material that comes to mind afterwards is the material that fits it. An instruction to be balanced does not change what comes to mind; it only changes how the material feels once it has arrived. Asking for reasons the conclusion is wrong changes the retrieval instruction, so different material arrives. The technique works on what gets fetched from memory, not on how carefully the material already fetched is weighed.

Same person, same conclusion, two instructions. What comes back differs. INSTRUCTION: BE BALANCED What arrives in the next minute: the reason already held a second reason of the same kind a feeling of having been fair Nothing contrary was requested, so nothing contrary was fetched. INSTRUCTION: WHY MIGHT THIS BE WRONG What arrives in the next minute: the fact that was set aside the case that ended differently the assumption nobody stated The retrieval instruction changed, so the returned material changed. Illustrative contents. The point is the instruction, not the three items listed under each one.
Being told to be balanced leaves the search unchanged, while being told to produce reasons against the conclusion changes what memory returns, and that is the whole of the effect.

Considering the opposite has something in common with everything that follows. The instruction does not ask for less bias. It asks for a specific action whose output is visible. Whether three contrary reasons were generated is knowable, in the way that whether a search was fair is not. Every method worth the time replaces a state of mind with an action that leaves a trace.

Try it out

What do the techniques with the better record have in common?

How to Build a Behavioural-Bias Checklist that gets used rather than filed?

A checklist is the plainest way to move a correction into the path of the decision. Atul Gawande set out the discipline in The Checklist Manifesto in 2009. Everybody already believed that lists are useful. The important part of his argument is that the good ones are short, sit at a defined pause, and ask for observable facts rather than for judgement.

Most behavioural checklists fail because they are written as a list of biases. A line reading am I anchoring cannot be answered. The line asks the decider to diagnose their own reasoning, and that capacity is the one in question. A line reading what number did I see first, and write it down, can be answered by anybody, and once the number is written down the anchoring question becomes visible without anyone having to name it.

So the construction rule is this. Each line names an observable, asks for something written, and can be answered wrongly. If a line cannot be answered wrongly it is decoration. Here is a six line version, and every line satisfies that test.

LineWhat it asks forWhat it catches
1Write the first number seen about this, and where it came froma starting point doing work nobody has looked at
2Write what is expected to happen, in a form that could turn out falsea conclusion too vague to be checked later
3Write the group of comparable past cases and what happened to thema judgement made only from the details of the case in hand
4Write three reasons this could be wrong, generated after the conclusionthe search that never went looking for contrary material
5Write what would have to be seen to change the conclusion, before actinga position that cannot be disturbed by any future evidence
6Write the date, and the earliest date on which action may be takenthe decision taken inside the hour it was first considered
The same intention, written two ways. One of them can be answered. THE VERSION THAT GETS FILED Am I anchoring on anything? Answered by inspecting one's own reasoning, which is the thing that has already gone wrong. Cannot be answered incorrectly, so ticking it costs nothing and proves nothing. THE VERSION THAT GETS USED Write the first number seen, and where it came from. Answered by looking outward at a fact, which anybody can do under pressure. Leaves a written trace, so a later reader can see what was in the room. Both lines are aimed at the same error. Only the lower one produces anything a reviewer could read.
A checklist line that asks for an observable fact leaves a written trace and can be answered under pressure, while a line asking for self diagnosis produces a tick and nothing else.

Now the part that matters more than the wording. A checklist has to sit at a pause that already exists, or one has to be created for it. In a practice, that pause is the moment before an instruction is sent. For somebody deciding alone, the natural pause is the moment the conclusion is reached and before anything is acted on. The sixth line puts it exactly there. A list that lives in a folder and is consulted when somebody remembers has been converted back into an intention, whose record is the one already described.

A procedure that only works on a holding has been written too narrowly, so run the same six lines on a decision with no money in it. A household is choosing a school for a nine year old. Line one: the first number seen was a fee of Rs 1,80,000/- a year, quoted by the first school visited, and every later school has been read against it. Line two: the expectation is that the child will settle within one term, stated so it could turn out false. Line three: the comparable cases are the four children on the same street who changed schools in the last three years, and two of them took longer than a year. Line four: three reasons this could be wrong, including that the visit was on a sports day. Line five: what would change the mind, namely the child being unhappy at the end of a second term. Line six: the decision date, and the earliest action date a week later. Nothing in those six lines is financial, and passing that test is what makes a procedure worth carrying into a decision that is.

One list, two decisions. The lines do not move; only the answers do. THE LINE CHOOSING A SCHOOL ADDING TO A HOLDING 1. First number seen a fee of Rs 1,80,000/- a cost of Rs 3,00,000/- 2. Falsifiable expectation settled within one term the reason still true in a year 3. Comparable past cases four children on the street the logged decisions of others 4. Three reasons against the visit fell on a sports day the segment ran that evening 5. What would change it unhappy after two terms the stated reason turns false 6. Earliest action date one week later one week later Amounts are invented. A list that stops working when the subject changes was a list about the subject.
Running the identical six lines on a school choice and on a holding shows that the procedure belongs to deciding rather than to investing, which is the test it has to pass.

Did the checklist change anything measurable in the invented Palash decision log? One thing, and only one. Twenty of the sixty investors adopted a written checklist on 4 November. Across quarters five to eight those twenty recorded a written reason on 34 of 41 decisions, or 82.9 per cent, against 19 of 63 decisions, or 30.2 per cent, for the other forty. Their realisation ratioHow readily gains were closed compared with how readily losses were, after allowing for how many of each were available to close. fell from 3.2 to 1.6, meaning gains were still closed more readily than losses but by half as much.

The log records no difference in return, and nobody should read one into those figures. There is one stretch of time, no comparison group assembled in advance, and no measurement of returns before or after. The measured effect is that reasoning became inspectable. Inspectable reasoning is a smaller claim than most readers want, and it is the only one the record supports.

Share of decisions carrying a written reason, quarters five to eight. 82.9 THE 20 WHO ADOPTED A WRITTEN CHECKLIST 34 of 41 decisions 30.2 THE OTHER 40 19 of 63 decisions 0 25 50 75 100 per cent of that group's logged decisions No return difference is recorded, measured or implied here. The invented log holds none. Illustrative figures.
What the invented log measures after the checklist is whether a reason was written down, and the two rates differ sharply, while no return comparison exists in the record at all.

What is the outside view, and where do its numbers come from?

Daniel Kahneman and Amos Tversky set out the correction in Intuitive Prediction: Biases and Corrective Procedures in 1979, and it is the most transferable idea in the whole subject. A case can be judged from its own details, which is the inside view, or from the record of cases like it, which is the outside viewJudging a case by the record of similar cases rather than its own details.. Almost everybody, almost always, does the first, and the second gives the better forecast in nearly every setting where both have been compared.

The reason is not that details are useless. Details are compelling out of proportion to what they predict. A plan that has been thought through carefully feels different from a plan that has not, and that feeling is real, and it is a poor guide to how long the plan will take. The record of comparable plans is a duller input and a better one.

The same thing happens in a household. Somebody says the kitchen work will take three weeks, and the estimate is built from the tasks: two days for the plumbing, four for the tiling, and so on, summed carefully. The outside view asks a different question. How long did the last four kitchens on this street take? Six weeks, nine weeks, seven weeks and five weeks. The inside view produced a number from the plan. The outside view produced a number from the record, and the record is not impressed by how carefully the plan was made.

Same job, two ways of answering. The record is not impressed by the plan. THE INSIDE VIEW built from the tasks in this plan plumbing, two days tiling, four days cabinets, five days paint and finishing, four days THE ANSWER: THREE WEEKS THE OUTSIDE VIEW the four comparable jobs on the street the plan said 3 6 9 7 5 0 5 10 weeks taken All five durations are invented. Every recorded job took longer than the careful estimate for this one.
Setting a carefully built three week estimate against four recorded jobs of six, nine, seven and five weeks shows what the outside view adds that no amount of care inside the plan supplies.

The two inputs the outside view needs are a reference classThe set of comparable cases a base rate is drawn from., meaning the set of comparable past cases, and a base rateHow often something happens across a reference class., meaning how often the thing in question happened across that set. Everything hard about the technique is in getting those two, and the next section is entirely about that difficulty.

Take the outside-view question that matters most to somebody deciding about holdings: does acting more improve what somebody picks? The invented Palash decision log has a reference class for it. Sixty investors sit in five turnover groups of twelve each, with annual turnover of 9, 34, 71, 128 and 210 per cent. Gross returns across those five groups ran 11.2, 11.0, 11.1, 10.9 and 11.0 per cent. Costs ran 0.3, 0.6, 1.5, 2.5 and 4.1 points, and net returns therefore ran 10.9, 10.4, 9.6, 8.4 and 6.9 per cent.

Gross returns sit inside 0.3 points of each other across a range of activity that varies more than twenty-threefold. Net returns run 4.0 points apart. The base-rate answer to the question is that in this cohort, acting more did not improve what was picked. Acting more changed only what the acting cost.

Five turnover groups, twelve investors each. Net on the left dot, gross on the right. 10.9 11.2 turnover 9 per cent 10.4 11.0 turnover 34 per cent 9.6 11.1 turnover 71 per cent 8.4 10.9 turnover 128 per cent 6.9 11.0 turnover 210 per cent 6.0 7.0 8.0 9.0 10.0 11.0 annual return, per cent. Axis starts at 6.0. Invented cohort, illustrative throughout.
The right-hand dots barely move down the five rows while the left-hand dots march leftward, so the widening red bar is the cost of activity rather than any difference in what was chosen.

The figure above carries the whole argument for taking an outside view at all. Any one of those sixty investors, asked why they traded, would give an inside-view answer built from the details of each decision. The record says the details did not matter to what was picked. The details mattered a great deal to what was kept.

The two spreads, drawn on one scale. This is the finding in a single picture. 0.3 points, gross returns GROSS 11.2 to 10.9 4.0 points, net returns NET 10.9 to 6.9 0 1 2 3 4 spread from highest group to lowest group, in percentage points Both bars start at zero and share one scale, so the shorter bar is one thirteenth of the longer one.
Setting the gross spread of 0.3 points beside the net spread of 4.0 points on one scale shows that almost all the difference between these groups was made after the picking, not during it.

How to Use Base Rates in a Forecasting Exercise, and which question comes second?

The procedure is two questions in a fixed order, and the second one is the one that does the work.

Question one: what is the reference class? The set of comparable past cases this one belongs to has to be named, and the boundary has to be specific. A class drawn too wide stops being comparable. A class drawn too narrow has one member, the case in hand.

Question two: were the outcomes of that class actually recorded? Not could they have been, not does somebody probably know. Were they written down somewhere a reader can find them. If the answer is no, the forecast has no base rate, and the honest output of the procedure is to say so rather than to estimate one from impression.

Most people run the first question, feel that they have taken an outside view, and never run the second. The result is worse than the inside view it replaced. A number produced from a class whose outcomes nobody recorded carries the authority of a statistic and the content of a guess.

Three candidate classes for the same decision. Only the middle one is a class. TOO NARROW 4 observations the decisions taken by this one investor the class has one member, which is the case in hand USABLE 48 observations the decisions taken by this turnover group of twelve comparable to each other, and their results were written down TOO WIDE any decision at all every decision taken by anybody, anywhere large, and the members stopped being comparable on the way Counts are from the invented log, where 240 decisions across 60 investors is four each.
Widening a class buys observations and spends comparability, so the boundary has to be argued for rather than pushed outward until the count looks respectable.
Two questions, in this order. The second is where most attempts stop being honest. 1. WHAT IS THE REFERENCE CLASS? 2. WERE ITS OUTCOMES ACTUALLY RECORDED? NO YES STOP, AND SAY SO This question has no base rate here. Estimating one anyway invents a number. READ THE RATE OFF THE RECORD Start the forecast from that rate, then adjust for what this case does differently. Naming a class is quick. Finding one whose results were written down is the half that fails.
The base-rate procedure is two questions in sequence, and the second one has a stopping branch, which is the feature that separates it from a general instruction to think about comparable cases.
Try it out

What is the second question in the base-rate procedure?

What do an available base rate and an unavailable one look like side by side?

A method taught only on the easy case does not survive practice, so here are both halves worked on the same invented record. The Palash decision log holds 240 decisions taken by 60 investors over eight quarters, four decisions each: 96 buys, 84 sells, 36 switches and 24 pauses of a standing instruction, and those four counts sum to 240.

The base rate that is available

The question is the one already asked: does acting more improve what somebody picks? The reference class is the five turnover groups, twelve investors each. Gross returns were computed for every group, so the outcomes are recorded. Reading the rate off the record gives 11.2, 11.0, 11.1, 10.9 and 11.0 per cent, a span of 0.3 points across turnover running from 9 to 210 per cent. The second question passes, so the forecast has a base rate, and the answer is that in this cohort more activity did not go with better picking.

The base rate that is not available

Now a question that sounds just as answerable. Do names mentioned in the media do well? The log has something on the subject. Of the 96 buys, 41 followed a media mention within three days, or 42.7 per cent, against 11.0 per cent of the eligible list being mentioned at all in a given week. The gap is large and it is real.

But it is a base rate for attention, not a base rate for outcomes. The figure records how often a mention preceded a purchase. The log recorded no result for mentioned names as a group, so the figure says nothing whatever about what happened to them afterwards. Question one passes: the class is nameable, being the eligible list of holdings that were mentioned in a given week. Question two fails: nobody wrote down how those holdings did. So the honest output is that this forecast has no base rate in this record, and any number produced for it would be manufactured.

Same procedure, two questions. One clears the second gate and one does not. DOES ACTING MORE IMPROVE PICKING? THE CLASS five turnover groups, twelve investors in each of them THE RECORDED OUTCOMES gross returns 11.2, 11.0, 11.1, 10.9 and 11.0 per cent, a span of 0.3 points THE QUESTION CAN BE ASKED DO MENTIONED NAMES DO WELL? THE CLASS holdings on the eligible list mentioned in a given week THE RECORDED OUTCOMES NOTHING WAS WRITTEN DOWN HERE THE QUESTION CANNOT BE ASKED Invented log throughout. The difference between the panels is a fact about the record, not about the questions.
What decides whether a base rate exists is whether somebody wrote the outcomes down, and the right-hand panel shows that a nameable class with no recorded results yields nothing at all.

The attention figure is worth drawing on its own. A careless reader will lift the 42.7 per cent out of this record and treat it as a finding about performance.

What the log measured about mentions, and what it did not. 42.7 buys following a mention, 41 of 96 11.0 share of the list mentioned in a week NO OUTCOME WAS RECORDED FOR MENTIONED NAMES AS A GROUP how those names then performed 0 10 20 30 40 50 per cent The dashed row is the point of the figure. It is empty because the invented record is empty there, not because the bar happens to be short.
Two solid bars measure how often attention preceded a purchase while the dashed row stays empty, and that empty row is the reason the performance question cannot be answered from this log.

The takeaway a reader can carry into any subject is short. Name the reference class first. Then check whether its outcomes were recorded. If they were not, stop, and say the forecast has no base rate rather than producing one anyway. A procedure whose most valuable output is sometimes the word no is doing something a general instruction to think harder cannot do.

Try it out

Why can the invented log not settle whether media-mentioned names do well?

Private Wealth Management Bootcamp — Fin Maverick

Does widening the reference class make the individual case clearer?

The control below is built around this question, and it is worth settling before anything moves.

Try it out

Before the control moves: does widening the reference class make the individual case clearer?

Play with it

Widen the reference class and watch what actually gets firmer

One variable moves: how many investors sit in the reference class, from a single person up to all 60. Everything else is held. The invented Palash decision log records 240 decisions by 60 investors over eight quarters, four decisions each, so a class of one person contributes 4 observations, one turnover group of twelve contributes 48, and the whole cohort contributes 240. The rate being estimated is held at 35.0 per cent, the share of the 240 that carried a written reason, being 84 of 240. The five turnover groups ran annual turnover of 9, 34, 71, 128 and 210 per cent with gross returns of 11.2, 11.0, 11.1, 10.9 and 11.0 per cent, a span of 0.3 points, and those group figures only become worth quoting at the full 240. Widening the class answers a different question rather than the same question better: 240 decisions say something about the sixty and nothing about any member of them.

1 investor, 4 decisions60 investors60 investors, 240 decisions
How firm is the group answer, at this width of class? The band is the range the group rate could plausibly sit in, given how many decisions the class contributes. PLAUSIBLE BAND 35.0 28.8 to 41.2 per cent 0 20 40 60 80 100 per cent of decisions carrying a written reason DECISIONS THE CLASS CONTRIBUTES 240 decisions the bar is full at 240, the whole logged record
Investors in the class, what moves
60
Decisions observed
240
Held constant, the recorded rate
35.0
Width of the band, points
12.3

With all 60 investors in the class the record contributes 240 decisions, so the recorded-reason rate of 35.0 per cent carries a plausible band of 28.8 to 41.2 per cent, a width of 12.3 points. That is a statement about the sixty and not about any one of them.

Educational illustration. The band is two standard errors either side of 35.0 per cent for the number of decisions in the class, clipped at nought and a hundred, which is a rough width rather than a formal interval. The cohort was written for teaching. Gross returns are before costs. A base rate describes a group and never an individual inside it.

Two things are worth noticing as the control moves. The band narrows steeply at first and then slowly, so the gain from going from four decisions to forty-eight is enormous and the gain from forty-eight to two hundred and forty is modest. And at every setting, the thing getting firmer is the group figure. Nothing that happens to the band says anything more about the single investor the case began with, and that is exactly what makes the outside view a different question rather than a sharper answer.

Red Team: what does one do that a reviewer does not?

A red teamA group tasked with defeating a case rather than improving it. is defined by its instruction, not by its expertise. Hand the same document to two people. Ask the first to review it, and they will look for weaknesses in order to strengthen it. Reviewing means exactly that. Ask the second to defeat it, and they will look for the one weakness that ends it. Defeating means exactly that. Same document, opposite instruction, and the two come back with different things.

The wording of the instruction matters, because a reviewer is, structurally, on the side of the case. Improving a case presumes the case survives improvement. A reviewer who found that the whole thing was wrong would have done something outside the job as given. Nobody has to be timid or political for this to happen; it follows from the wording of the task.

One document. Two instructions. Two completely different outputs. THE SAME DOCUMENT INSTRUCTION: IMPROVE THIS CASE the case is presumed to survive the exercise INSTRUCTION: DEFEAT THIS CASE the case is presumed to be defeatable WHAT COMES BACK eleven improvements, a tighter argument, the same conclusion WHAT COMES BACK one objection that, if true, ends the case entirely Neither output is better in general. They answer different questions, and only one of them can say stop.
Because reviewing presumes the case survives and defeating presumes it can be ended, the instruction alone decides which of the two outputs a reader will ever see.

Three conditions make the difference real rather than theatrical. The instruction is written down, so nobody has to guess how adversarial to be. The output is a written objection rather than a conversation, so it survives the meeting. And whoever holds the case has to answer the objection in writing. Most arrangements skip that part, and an objection nobody answered is indistinguishable from an objection nobody made.

Try it out

What separates a red team from a reviewer?

Investment Banking Analyst Bootcamp — Fin Maverick Writing an Investment Thesis — free micro-course from Fin Maverick

How to Use a Devil’s-Advocate Review in Research, and who argues what?

The devil-advocate review is the small, repeatable version of the same idea, sized for a person or a two-person practice rather than for a committee. The review has four parts and takes about twenty minutes.

First, the seat is named before the work starts, not after. Somebody is the objector for this note, and they know it in advance. A person asked to object on the spot will object politely and about nothing important. Second, the objector gets a written instruction with a single target: produce the strongest reason this conclusion is wrong, and produce it as a claim that could be checked. Third, the objection goes on paper. Fourth, the author answers it on paper, and the answer stays with the note.

The fourth part is the one that decides whether any of this was worth doing. An objection that was raised, nodded at and forgotten has left no trace, and six months later nobody can tell whether it was answered or ignored. An objection with a written answer beside it lets a later reader see the reasoning that was actually applied, and that is the entire purpose of the exercise.

For somebody working alone with nobody to hand the objector seat to, the substitute is time rather than a person. The conclusion is written and then left. Later the strongest available objection is written out, as though somebody else had produced it, before the original reasons are reread. The substitute is weaker than a second person and much better than nothing. The retrieval instruction has still been changed.

Four steps. The last one is the step almost everybody skips. 1. NAME THE SEAT before the work starts, never at the meeting itself 2. ONE TARGET the strongest reason this is wrong, checkable 3. ON PAPER a written claim, not a remark in a conversation 4. ANSWERED in writing, filed beside the note, readable later An objection that was raised and nodded at leaves no trace at all, so six months later nobody can tell it apart from an objection nobody made. Working alone, the fourth step still applies. Time replaces the second person. The instruction and the written answer are what make this different from asking a colleague what they think.
Naming the objector in advance and filing a written answer beside the objection are what turn a conversation into a record a later reader can inspect.
Backtesting a Strategy teaches you to build a backtest, name how it flatters itself, and state what the result establishes.

How to Check for Confirmation Bias in a Research Note, and in what order?

Confirmation bias is not usually visible in a conclusion. Confirmation bias is visible in the sources, and specifically in what is missing from them. So the check is a sweep of the source list, run in a fixed order. Run in any other order, the sweep collapses into an impression about whether the note felt fair.

Step one is to list every source the note used, without judging any of them. Step two is to mark each one as supporting the conclusion, cutting against it, or neutral. Step three, the sharp one, is to describe what a source arguing the other way would look like, concretely enough that it would be recognised. Step four is to ask whether such a source was looked for, and whether the search that would have found it was ever run.

A note with nine supporting sources and a note with two can be equally sound or equally selective, so the tally of supporting sources settles almost nothing. What settles something is step four. A note that names the contrary source it went looking for and could not find is in a different condition from a note whose search was never pointed that way, even where both end with the same nine supporting citations.

The sweep, in order. Running it out of order turns it into an impression. 1 LIST EVERY SOURCE THE NOTE USED no judgement yet, just the list, including the ones that were only glanced at 2 MARK EACH ONE: SUPPORTS, CUTS AGAINST, NEUTRAL a mechanical pass, so the marking cannot be argued with afterwards 3 DESCRIBE WHAT A CONTRARY SOURCE WOULD LOOK LIKE concretely enough that one would be recognised if it appeared 4 ASK WHETHER ONE WAS ACTUALLY LOOKED FOR this is the step that finds something, and the only one that can fail Counting supporting sources cannot separate a sound note from a selective one. Step four can.
Because a supporting tally looks identical for a sound note and a selective one, the sweep only earns its time at the fourth step, which asks what was searched for.
Two notes with identical tallies. One of them has done the work. NOTE A sources supporting the view 9 sources cutting against it 0 No contrary search is recorded. Nobody knows what was skipped. NOTE B sources supporting the view 9 sources cutting against it 0 The contrary search is described, was run, and returned nothing. Counts are illustrative. The bars are identical on purpose, because the tally is not what separates the two.
Identical source tallies can sit on either side of the only distinction that matters, which is whether the search that would have found contrary material was ever pointed that way.
Try it out

What is the sharpest question to ask of a research note's sources?

How to Build a Personal Research Pause Protocol that actually gets kept?

A pause protocolA fixed delay between reaching a conclusion and acting on it. is a fixed delay between reaching a conclusion and acting on it, with a fixed thing done inside the delay. A pause protocol is not a general resolution to slow down, an intention wearing a procedure’s coat, and not a rule about being calm.

Three settings define it, and all three are set once, in advance, when nothing is happening. The length of the wait, stated in hours or days rather than in feelings. The trigger, meaning which decisions the wait applies to. A protocol that applies to everything gets dropped in the first busy week. And the reread, meaning the specific thing done before acting, normally rereading the written reasons and checking whether they still say what they were thought to say.

A reasonable starting shape for a person deciding alone is this: any decision arriving within forty-eight hours of a news item waits until the following day, and before acting the written reason is reread and confirmed not to consist mainly of the news item. The forty-eight hour trigger is not arbitrary. In the invented Palash log, 71 of the 240 decisions were taken within 48 hours of a news item, or 29.6 per cent, so almost a third of the record would pass through that gate.

Set once, in advance. Then it runs without anybody deciding to run it. CONCLUSION reached, written THE STATED WAIT a length in hours or days, fixed beforehand nothing is acted on here THE REREAD of the written reason ACT, OR NOT either is a result SETTING ONE, THE LENGTH stated in hours or days, never in how it feels SETTING TWO, THE TRIGGER which decisions it covers, narrow enough to survive SETTING THREE, THE REREAD the exact thing done before anything is sent Choosing not to act at the end of the wait is a result, not a failure of the protocol.
Fixing the length, the trigger and the reread in advance is what stops a pause protocol collapsing back into a general resolution to be more careful.

Look at what such a gate would have touched in the invented record, and be careful about what that means. On 19 February a television segment named Suvarna Chemicals Limited and Meera Sundaram added Rs 1,00,000/- to that holding the same evening, taking its cost to Rs 4,00,000/- and the total cost to Rs 13,00,000/-. A pause protocol would have moved that decision to the following day. The wait would not have told her whether the decision was right, and nothing in this record shows that waiting would have produced a better result. What the wait produces is a written reason read twice, once while the segment was running and once when it was not, and a reader who can see both.

Where a stated wait would sit in the logged sequence. 4 January holding opens 19 FEBRUARY segment runs, Rs 1,00,000/- added the same evening the wait would sit here 6 June questionnaire done 30 September valuation struck 12 OCTOBER the decision the case turns on 4 November checklist adopted A wait changes when a decision is taken and what is on paper when it is. It does not decide whether the decision was right, and this invented record contains no measurement of that. Dates sit inside a single unnamed year.
Marking the two same-day decisions on the logged sequence shows exactly where a stated wait would insert itself, and the panel below states what such a wait does not settle.
Ratio Analysis That Says Something — free micro-course from Fin Maverick

What does none of this fix?

The procedures leave four substantial things untouched.

First, none of it fixes a bad reference class. The procedures force a class to be named; they cannot say that the class named was wrong. A person who compares this decision to the four most memorable past cases has run the outside view correctly on a class assembled by memory, and memory is a biased sampler. The written class helps a later reader spot the problem. Spotting is not preventing.

Second, none of it fixes the thing nobody has thought of. A checklist covers the errors somebody wrote down. A red team objects to the case as presented. Both operate inside the frame the work already has. Genuine surprises live outside it.

Third, none of it fixes incentives. If somebody is paid for the conclusion, an objector seat and a written answer produce a well-documented version of the conclusion they were paid to reach. Procedure and interest are separate things, and procedure is the weaker of the two.

Fourth, and most importantly for a reader who has come this far, none of it has been shown to improve returns, on this record or anywhere in this sequence. What the twenty who adopted the checklist demonstrably produced was written reasons on 82.9 per cent of their decisions instead of 30.2 per cent, and a realisation ratio that fell from 3.2 to 1.6. Both figures are changes in how decisions were recorded and in which positions were closed. Whether it made them better off is not measured, and cannot be measured from one stretch of eight quarters with no comparison group.

Four things still standing after every one of these procedures has been run. A BAD REFERENCE CLASS Naming one is required. Nothing reveals that the class named was assembled out of memorable cases. WHAT NOBODY THOUGHT OF A list covers written errors. An objector argues the case as put. Surprises sit outside both frames. AN INTEREST IN THE ANSWER Where somebody is paid for the conclusion, procedure produces a well-documented version of it. ANY CLAIM ABOUT RETURNS Not measured here and not claimed. One stretch, no comparison group, no before-and-after measurement. The fourth panel is the one every claim made elsewhere has to stay consistent with.
Setting the four unfixed things beside each other keeps the six procedures inside the claim the record supports, which is about inspectability rather than about outcomes.
Try it out

Do these procedures improve returns?

Each procedure aims at one failure and leaves one thing on paper. THE PROCEDURE THE FAILURE IT TARGETS WHAT IT LEAVES ON PAPER Considering the opposite a one-sided search three written reasons against The checklist the step nobody remembers six answered lines Base rates a forecast built from details a named class and its record The objector seat review that cannot say stop an objection and its answer The source sweep evidence gathered one way a marked source list The stated wait acting inside the first hour two datestamps, and a reread Read the right column on its own. Every entry is something a later reader could pick up and inspect.
Laying the six procedures against the failure each targets shows that all of them are judged by an artefact somebody else could read rather than by how the decider felt.

The error that gets made, and what it costs

The error is treating debiasing as a matter of vigilance. Vigilance sounds responsible, it flatters the reader who has just learned a long list of mechanisms, and it is what the evidence contradicts. Fischhoff's survey did not find that people were not trying hard enough. The survey found that trying, in the form of knowing and intending, mostly did not move the outcome.

The reason is a timing problem that no amount of resolve fixes. An intention is available in the calm hour when it is formed and unavailable in the loud minute when it is needed. The same conditions that produce the error also suppress the search for the correction. A written step does not have this property. A written step sits in the path, fires whether or not anybody thought of it, and leaves an output that somebody who was not in the room can read afterwards.

The error costs the wrong repair. A person who believes debiasing is vigilance responds to a bad decision by resolving to be more careful, and a resolution changes nothing measurable. They do not build the six-line list or set the wait, and both of those change what gets recorded. The whole distinction is between a resolution, whose only evidence is a feeling, and a procedure, whose evidence is an artefact.

Availability at the two moments that matter. THE CALM HOUR when the rule is made THE LOUD MINUTE when it is needed A RESOLUTION PRESENT GONE A PROCEDURE PRESENT STILL PRESENT The upper row is why awareness performs so poorly. The lower row is the only design that avoids it. None of this is measured on the invented cohort; it is the shape of the timing problem.
Because a resolution is present when it is formed and absent when it is needed, the timing gap alone explains most of the disappointing record without anybody having tried less hard.

How does somebody actually run this, alone or on behalf of others?

Devika Rao, the adviser at the invented Palash Advisory Services Private Limited, does not run six procedures on every decision. A practice that did would stop functioning inside a month. She runs a trigger. Any instruction arriving within two days of a news item, or any instruction that closes a position at a round number matching its cost, goes through the six-line list before it is sent, and the objector seat is used on the three or four research notes a quarter that carry the most weight.

The lone decider runs the same shape with the seats collapsed. The trigger is the same, the list is the same six lines, the objector seat becomes a stated wait plus the strongest objection written out the next morning, and the source sweep runs on whatever was read before deciding, even where that is three articles rather than a research note.

The common design in both is that the procedure is attached to a trigger rather than to a good intention, and it produces an artefact rather than a feeling. A lender reading a credit file, an analyst signing a note and a household deciding on a school are all doing the same thing: making the reasoning visible to somebody who was not there, including their own later self.

Same design, two settings. Only who fills the seats changes. THE TRIGGER, IDENTICAL IN BOTH: A DECISION WITHIN TWO DAYS OF A NEWS ITEM IN A PRACTICE the six-line list, before sending a named objector, three or four times a quarter the source sweep on the note the wait, held by the process DECIDING ALONE the six-line list, before acting one's own objection, written the next morning the source sweep on what was read the wait, held by the calendar Neither column contains anything specific to money, which is why the same shape runs on a school or a job.
Running the identical trigger and list in a practice and alone shows that what changes between the two settings is who fills each seat, not the procedure itself.
One decision, passing through all six in order. CONCLUSION REACHED the six-line list reasons the opposite way the reference class and gate the objector and the answer the source sweep the stated wait ACT, OR DO NOT What leaves the path is a written record of the reasoning. What does not leave it is any statement about whether the decision will turn out well, which no procedure here is able to supply. Nobody runs all six on every decision. The trigger decides which decisions enter the path at all.
Following one decision through the whole sequence shows what comes out at the end, which is a readable record of reasoning rather than any statement about the outcome.

One last figure, on the only measured movement in the invented record, drawn so that nobody can read more into it than it holds. The realisation ratio for the twenty who adopted the list fell from 3.2 to 1.6, meaning gains were still closed more readily than losses but by half as much. The fall is a change in which positions were closed. The ratio is not a return.

The one measured movement, and the boundary around it. 3.2 before 1.6 after 0 1 2 3 4 gains closed for every one loss closed, at equal opportunity THE 20 WHO ADOPTED THE LIST The ratio moved. No return figure was recorded before or after, and no comparison group was set up in advance, so the question stays open.
Drawing the ratio moving from 3.2 to 1.6 beside an explicit statement of what was not measured keeps the only recorded movement from being read as evidence about outcomes.
The pre-mortem is a separate technique, set out under the pre-mortem, and how to judge a decision once its outcome is known is a different problem, set out under decision quality. The decision journal and its falsifiability test are assumed here rather than retaught. Reasoning becomes inspectable afterwards, and no claim beyond that survives this record. Where a duty of an adviser is at stake, the Securities and Exchange Board of India states the requirement.
Debiasing forces a reference class to be named, never checked. See what stays untouched.

Sources

SourceDocumentSite
Baruch Fischhoffthe survey chapter titled Debiasing, in Judgment under Uncertainty, 1982ssrn.com
Charles Lord, Mark Lepper and Elizabeth Prestonthe paper testing the instruction to consider the opposite, Journal of Personality and Social Psychology, 1984ssrn.com
Daniel Kahneman and Amos TverskyIntuitive Prediction: Biases and Corrective Procedures, 1979, where the outside view is set outnber.org
Amos Tversky and Daniel KahnemanJudgment under Uncertainty: Heuristics and Biases, Science, 1974ssrn.com
Atul GawandeThe Checklist Manifesto, 2009, on how a short list is built and where it sitscited to the book itself
Karl PopperThe Logic of Scientific Discovery, 1934, for the test a claim has to be able to failcited to the book itself
J Edward Russo and Paul SchoemakerDecision Traps, 1989, on keeping the written record of a decisioncited to the book itself
Securities and Exchange Board of Indiaconduct and suitability duties applying to registered intermediariessebi.gov.in
Association of Mutual Funds in Indiainvestor-facing practice material for distributors and advisersamfiindia.com

Meera Sundaram, Devika Rao, Palash Advisory Services Private Limited, the Palash decision log, the Palash 100 index, Suvarna Chemicals Limited and Kesari Logistics Limited are invented.
Educational material. Not advice on any investment, tax, budget or market position.

Covered in this topic

Subtopics

Red TeamHow to Build a Behavioural-Bias ChecklistHow to Use Base Rates in a Forecasting ExerciseHow to Use a Devil's-Advocate Review in ResearchHow to Check for Confirmation Bias in a Research NoteHow to Build a Personal Research Pause Protocol
← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.