Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Behavioural Finance & Investor Decision-Making
1Foundations
The Rational InvestorJudgment Under UncertaintyPreferencesBehavioural FinanceInvestor and Market BehaviourFinancial Well-BeingBounded RationalityHeuristics and Biases
2Cognitive Biases, Emotion and Attention
Limited AttentionRepresentativenessThe Affect HeuristicAnchoring and AdjustmentEmotion and Decision QualityOverconfidence and OptimismAmbiguity and Complexity AversionAvailability and SalienceHome Bias, Local Bias…FramingThe Halo EffectHindsight BiasThe Narrative FallacyPresent Bias and Hyperbolic DiscountingBase-Rate NeglectStatus Quo Bias and the Default Effect
3Preferences and Prospect Theory
Prospect TheoryRegretThe Endowment EffectMental AccountingThe Sunk Cost FallacyLoss AversionRisk Seeking in Losses
4Social Behaviour
HerdingNarrative EconomicsFear of Missing OutGroupthinkSocial Proof
5Investment and Trading Behaviour
Excess TradingNaive DiversificationThe Disposition EffectLottery PreferencesNoise TradersPortfolio InertiaRecency Bias
6Markets and Anomalies
Mania, Panic and CapitulationMarket EfficiencyEfficient Market Hypothesis vs…Speculative BubblesReflexivityInvestor SentimentMarket AnomaliesShort-Sale ConstraintsPrice DiscoveryLimits to Arbitrage
7Decision, Research and Debiasing
The Decision JournalDebiasingChoice Architecture, Defaults and…The Pre-Mortem and Process QualityDecision Quality
8Advice, Conduct and Communication
Communication ConductSuitability and AppropriatenessChoice OverloadComplaint BehaviourRisk DisclosureVulnerable Investors

Decision Quality: Judging the Decision, Not the Outcome

Decision quality asks one thing only: was the reasoning sound on the information that existed when the choice was made. An outcome is one draw from a range nobody can see, so judging by it means judging on a sample of one. The two come apart in both directions, and the direction that does the damage is a poor process that happened to work.

There is a sentence almost everybody agrees with and almost nobody lives by: a good decision and a good result are two different things. The agreement lasts until a result actually arrives. Then the result is sitting there, large and specific and impossible to ignore, and the reasoning that produced it has gone quiet. The distance between agreeing with that sentence and behaving as though it were true is where the whole difficulty lives. Closing that distance takes something more definite than good intentions. A short set of questions, all of them answerable before anything has happened at all, is one way of doing it.

What is decision quality, and how is it different from a good result?

Decision qualityWhether the thinking behind a choice holds up on the information that existed when it was made. asks one question and only one: given what could be known when the choice was made, was the reasoning sound? Not was it clever, not was it bold, not did it feel right. Was it sound, and could somebody else pick it up afterwards and check it. Outcome qualityHow the decision actually turned out. asks something completely different: how did it turn out. Both are real questions. Only one of them is about the decider.

The pattern is clearer away from money, where nothing is at stake for the reader. A guest leaves a wedding at one in the morning, having slept two hours the night before, and drives ninety minutes home. The drive ends safely. Because the ending was fine, almost nobody telling that story afterwards describes it as a bad decision. Yet nothing about the ending changed the state of the driver at the moment the key turned. The decision was taken with a tired person, a dark road and no plan for what happens if concentration goes. Arriving home safely was one of the things that could have happened, and it happened. Arriving home safely is not a report on the driving.

Turn it round and it is just as clear. A household saves for eleven years towards a stated goal, chooses a boring monthly instruction they will not be tempted to interrupt, keeps two months of outgo aside so a bad month never forces a sale, and writes down why. Then something outside all of it goes wrong and the goal is missed. The reasoning did not become poor when the result arrived. The reasoning was the same reasoning the whole time.

So a decision, at the moment it is taken, is a choice made across a set of futures nobody can see. Some of those futures are kind to it and some are not. One of those futures arrives. The reasoning is the thing the decider was in charge of; the draw is the thing the decider was not. A judgement that scores the decider on the draw is scoring the decider on somebody else's work.

One decision. Many futures. One draw. THE DECISION the reasoning the information held the reason written the check set in advance ALL OF IT THE DECIDER'S and all of it knowable on the day THE RANGE OF FUTURES none visible in advance THE ONE THAT ARRIVED loud, specific, and easy to mistake for a verdict not the decider's to control The middle panel is the part a review never sees, which is why the right panel gets asked to do its work.
A decision is taken across a spread of futures that stays permanently invisible, so the single future that turns up is asked to stand in for the whole spread, and it cannot.

The phrase repays being precise about. Decision quality is not a softer, kinder version of outcome quality for people who did badly. Decision quality is a separate measurement of a separate thing, taken at a separate time, and it is the only one of the two that exists on the day it is needed.

The question a review can askWhen it can be answeredWho supplies the answer
Was the reasoning sound given what was knowable?on the day of the decision, before anything happensthe person deciding
Did it turn out well?only later, and sometimes much laterthe world
Which one can be used to improve the next decision?the first, because it is available while there is still timethe person deciding
Try it out

What is decision quality measured against?

Financial Literacy Bootcamp — Fin Maverick

What is resulting, and why does explaining it not stop it?

ResultingJudging a decision by its outcome rather than by its reasoning. is the name for judging a decision by how it came out. The word is Annie Duke's, from Thinking in Bets, 2018, and it is a useful word precisely because the thing it names is so ordinary that without a word for it nobody notices it happening. A trade worked, so it was a good call. The dinner was excellent, so the booking was inspired. The candidate turned out badly, so hiring them was a mistake. Each of those sentences has skipped a step so smoothly that the tape has to be slowed down to see it.

Resulting is hard to shift because resulting is not a piece of ignorance, and being told about it does not cure it. Jonathan Baron and John Hershey showed as much in Outcome Bias in Decision Evaluation, published in the Journal of Personality and Social Psychology in 1988. Baron and Hershey gave people descriptions of decisions, stated plainly the information that had been available when each decision was taken, and then attached different results to the same description. The ratings people gave separated according to the result. The information at the time was identical and it was written down in front of them, and the ratings moved anyway.

Read the finding again, slowly. The result is stronger than it first sounds. The experiment is not about people forgetting what was known at the time. The people rating those decisions had the information in their hands and were told outright that it was identical in both versions, and they still rated the decision differently once a result was attached. Resulting survives being explained, so an explanation of it is not a cure and was never going to be one. What works is not knowing about it; what works is scoring the decision on something the outcome cannot reach.

Same decision, same stated information, two different results attached. VERSION ONE VERSION TWO the reasoning, word for word the same the information held, stated in full the reasoning, word for word the same the information held, stated in full RESULT ATTACHED: IT WENT BADLY RESULT ATTACHED: IT WENT WELL HOW THE SAME DECISION THEN GETS RATED rated worse rated better the reasoning is identical in both
Baron and Hershey held the stated information constant and moved only the result, and the ratings of the very same reasoning separated, so each arrow carries a direction and no size.

The gap between the two sets of ratings has a direction and a size, and only the direction carries over from one experiment to the next. How far apart the ratings went depends on the wording of the particular scenarios, so a number put on the size would be an invented measurement rather than a finding.

Why does the result get to be so loud? Partly because it is definite. The reasoning is a paragraph of qualified sentences about what might happen; the result is a single number that has already happened. Partly because a result arrives with feelings attached, and a paragraph does not. And partly because reviewing a result costs nothing while reconstructing what somebody knew four months ago is work, and the work is skipped unless it was written down at the time.

What a review looks at, drawn at the size it actually gets looked at. THE ENTRY, AS IT IS READ FOUR MONTHS LATER COST Rs 75,800/- the reason recorded on the day, in the same entry sell when it gets back to Rs 3,00,000/- The type sizes are the argument. Nothing else in the entry has to be hidden for the result to win.
The outcome and the reasoning sit in the same record, but one of them arrives at four times the size of the other and takes the review with it.
Try it out

Baron and Hershey gave people identical decisions with different results attached. What happened?

Why is one outcome a sample of one?

The arithmetic underneath all of this is not difficult arithmetic. When a decision is taken, there is a spread of things that could happen. One item from that spread arrives. To judge the decision by what arrived is to estimate the whole spread from a sample of oneA single observation, from which no general conclusion follows.. How wide the band of pure chance results grows when the number of observations is small is set out under representativeness. One is the smallest number there is, so the band is at its widest, and a judgement resting on it has no evidential content at all. Not weak content. None.

An everyday version. A street vendor outside an office building takes one afternoon off and misses the busiest day of the month. Was taking the afternoon off a poor decision? One afternoon cannot settle it. Thirty afternoons might, and even then only if something is written down about which afternoons were chosen and why. The single afternoon is loud because it is the one that happened, and it is silent because it is one.

The invented Palash 100 index shows the same thing in a shape that is easy to look at. The index opens at 100.0, and its eight quarter ends run 118.0, 131.0, 112.0, 104.0, 116.0, 124.0, 121.0 and 127.0. The peak is 131.0 at the second quarter end and the low is 104.0 at the fourth. Somebody who happened to read the number at the second quarter end saw a rise of 31.0 per cent from the open. Somebody reading at the fourth saw a rise of only 4.0 per cent, and from the peak a fall of 20.6 per cent. The path is one path. The reading depends entirely on which single moment supplied it.

One path. Two single readings. Two completely different stories. 96 108 120 132 131.0, the peak 104.0, the low open Q1 Q2 Q3 Q4 Q5 Q6 Q7 Q8 Invented and illustrative throughout. The index is not a description of any real market.
Reading the invented index at the second quarter end gives a rise of 31.0 per cent and reading it at the fourth gives 4.0 per cent, from exactly the same path.

Take four consecutive readings of the same invented index and the point stops being abstract. From the open to the first quarter end the Palash 100 rose 18.0 per cent, from there to the second a further 11.0 per cent, then it fell 14.5 per cent to the third and a further 7.1 per cent to the fourth. Four moves, one path, and any single one of them read on its own describes a rise or a fall nobody actually lived through.

Four quarters of one path, each read on its own. open to Q1 Q1 to Q2 Q2 to Q3 Q3 to Q4 up 18.0 per cent up 11.0 per cent down 14.5 per cent down 7.1 per cent Eight pixels to a percentage point on every row. The invented index is illustrative and describes no real market.
Four consecutive moves of one invented path run from up 18.0 per cent to down 14.5 per cent, so any single quarter read alone reports an index nobody held.

The same trap sits inside the logged record. Across the eight quarters, 84 of the 240 logged decisions were sales, and 61 of those sales, or 72.6 per cent, were of positions standing in gain. Left there, the figure sounds like a finding about a preference for winners. The figure is nothing of the sort, because it ignores what was there to be sold. Of the 240 positions open across the period, 108 stood in gain and 132 in loss. Gains were realised at 61 of 108, or 56.5 per cent, and losses at 23 of 132, or 17.4 per cent, a ratio of 3.2 to 1. The ratio is the measurement and the 72.6 per cent is the trap, and the reason the trap works is the same reason resulting works: a single loud number, read on its own, with nothing underneath it.

The loud number, and the two numbers underneath it. READ ON ITS OWN 72.6 per cent of sales were positions in gain, 61 of 84 and it says nothing, because it ignores what there was to sell READ AGAINST WHAT WAS AVAILABLE 56.5 pc gains, 61 of 108 17.4 per cent losses, 23 of 132 a ratio of 3.2 to 1 which is the measurement All three bars share one scale, three pixels to a percentage point, so they can be compared by eye. Invented figures from the logged record of 240 decisions.
Set against the positions actually available to sell, the realisation rates come out at 56.5 and 17.4 per cent, and the headline share of 72.6 per cent turns out to have measured nothing.

What does Process Quality vs Outcome Quality look like, set side by side?

Most of the confusion between the two comes from only ever meeting them as a pair, so each one has to be defined on its own before they can be contrasted. So take them one at a time.

Process quality is a property of how a decision was made. Process quality is made of four things that could be photographed on the day: whether a reason was recorded, whether the reason committed to something that could later be found false, whether the standard being used was a fact about the thing being decided or a fact about the decider, and whether anybody was given the job of arguing the other side. None of those four needs the future. All four are finished by the time the decision is taken. Process quality is therefore complete on the day, and it is complete whether the decider is a committee of nine or a person at a kitchen table with nobody to consult.

Process quality, defined on its own: four things, all of them finished on the day. 1 2 3 4 A REASON, WRITTEN not remembered later, written at the time SOMETHING FALSIFIABLE a claim that could turn out to be wrong AN HONEST STANDARD about the thing, not about the decider's own past A CASE AGAINST somebody, or the decider, arguing the other way ALL FOUR ARE COMPLETE BEFORE THE DECISION LEAVES THE ROOM and none of them needs to know what happens next
Process quality has four components and every one of them is settled at the moment of deciding, which is why it can be scored before anything at all has happened.

Outcome quality is a property of what turned up. Outcome quality is made of the result itself, the costs that were taken out of it on the way through, and the stretch of time it happens to be measured over. Two of those three have nothing to do with the decider either. A result measured to one date and the same result measured to a date three months later are two different outcome qualities from one decision, and nobody decided the date. Outcome quality is real, it matters, and it is what the money actually does. Outcome quality is simply not a report on anybody’s thinking.

Outcome quality, defined on its own: three things, and two of them are not the decider's. THE DRAW THAT ARRIVED THE COSTS TAKEN OUT THE STRETCH MEASURED one item out of a spread nobody could see dealing charges, spread and tax, all real a date somebody picked for reasons of their own NOT THE DECIDER'S PARTLY THE DECIDER'S NOT THE DECIDER'S NONE OF THE THREE CAN BE READ ON THE DAY OF THE DECISION which is the practical difference, and it is the whole of it
Outcome quality is built from a draw nobody chose, the costs of getting there and a measuring date somebody else picked, so most of it is settled outside the decider.
One sale. Two measuring dates. Two different verdicts. MEASURED ON 12 OCTOBER MEASURED TO 31 MARCH FOLLOWING Rs 60,000/- booked Rs 36,800/- forgone 15.0 per cent on a cost of Rs 4,00,000/-, and it looks well timed it rose 8.0 per cent after the sale, and now it looks early THE DECISION DID NOT CHANGE BETWEEN THE TWO PANELS The only thing that changed is the date somebody chose to measure on, and nobody decided that.
The same sale reads as Rs 60,000/- booked on one date and as Rs 36,800/- forgone on another, which is two outcome qualities from one unchanged decision.

Now the two can be put side by side, and the contrast is sharper than most readers expect. Process quality is knowable at the time and outcome quality is not, and that single difference is what decides which of the two can be steered by. Everything else in the comparison follows from it.

Asked ofProcess qualityOutcome quality
What it measureshow the decision was madehow it turned out
When it can be readon the day, before anything happenslater, and the date is arbitrary
Who settles itthe decider, entirelythe world, mostly
What moves itchanging how a decision is madechanging what the world does
How many observations arriveone per decision, every timeone per decision, and it is a sample of one
Which one can be improved on purposethis onenot this one

A short warning about the table above. Reading it, it is tempting to conclude that outcome quality does not matter. Outcome quality matters enormously; it is the only thing that ever pays for anything. The narrower claim is worth stating exactly. Everything that happened after the reasoning stopped is mixed into the outcome, so the outcome cannot serve as a measurement of the reasoning.

What each one actually responds to. PROCESS QUALITY OUTCOME QUALITY HOW THE DECISION IS MADE THE SCORE MOVES Write the reason down and the score rises the same afternoon. The lever and the dial are in the same room. STEERABLE THE WORLD MOVES THE RESULT MOVES Decide exactly as well and the result still lands wherever it lands. The lever is not in the room at all. NOT STEERABLE
Only one of the two responds to anything the decider can do, which is why the other one makes such a poor thing to aim at directly.
Portfolio Management Bootcamp — Fin Maverick

Which of the four combinations does the damage?

Sound and poor reasoning set against good and poor results give four combinations. The four combinations are not equally common and, far more importantly, they are not equally dangerous. Most treatments stop at drawing the grid. The grid is not the teaching; the asymmetry inside it is.

Sound reasoning with a good result is the pleasant cell. There is nothing to correct, so it teaches almost nothing. Poor reasoning with a poor result is the obvious cell: the result hurts, somebody goes back to the entry, the fault is visible and it gets fixed. Sound reasoning with a poor result is the painful cell, and its risk is over-correction, a working method thrown away because one draw went against it. All three of those cells get looked at.

Poor reasoning with a good result is the cell that does the damage, and it does the damage precisely because it is never examined. The cell arrives looking like a success and is described as a good call. Nobody changes something that worked, so the reasoning gets repeated, and each repetition collects more evidence for a method that has none. Nobody in the history of reviews has opened a file and said, that one made money, so find out what went wrong.

Four combinations, and only one of them never gets opened. RESULT: GOOD RESULT: POOR REASONING: SOUND REASONING: POOR EARNED the reasoning and the result agree, and there is nothing much left to correct GETS LOOKED AT: SOMETIMES PAINFUL examined hard, often too hard, and a method that works can get dropped GETS LOOKED AT: ALWAYS THE DANGEROUS ONE rewarded, repeated, and collecting more evidence for a method that has none GETS LOOKED AT: NEVER OBVIOUS the result forces somebody back to the entry, and the fault is sitting there GETS LOOKED AT: ALWAYS
Three of the four combinations end up being examined by somebody, and the fourth is rewarded instead, which is what lets a method with nothing behind it survive.

There is a second, quieter cost. Because the rewarded cell is never opened, the person in it learns the wrong lesson with total confidence. The person now believes something that is not true about how the world works, and believes it on the strength of evidence lived through personally. Lived evidence is the most persuasive kind there is and the hardest to argue with.

The error that gets made, and what it costs

The error is running the review on the result. The habit looks responsible. Somebody sorts the decisions by how they turned out, opens the worst ones, and asks what went wrong. Every step of that is reasonable and the whole thing is upside down. Sorting by result guarantees that the poorly reasoned decisions that happened to work are never in the pile that gets opened.

The cost is precise, and the cost is not the money on any single entry. The review can only ever find faults in the cell it opens. Run that way for two years, a review systematically corrects sound methods that got unlucky and leaves untouched every scrap of loose reasoning that got lucky. The record of the logged decisions shows how much material there is to work with. A written reason exists on only 84 of 240 entries, or 35.0 per cent. On nearly two thirds of them there is no reasoning left to examine at all, whichever way the pile is sorted.

The fix is not to review harder. The fix is to sort the pile by the process score instead of by the result, and sorting that way is only possible if something was written down at the time. Keeping a record is what the whole argument rests on, which is why scoring a decision belongs alongside the decision journal rather than apart from it.

The loop a result-sorted review runs, and it has no exit. SORT THE PILE BY RESULT OPEN THE POOR RESULTS CORRECT WHAT IS IN THAT PILE THE LUCKY LOOSE ENTRIES STAY nothing in this loop ever opens the entries that worked by luck Each turn corrects a sound method that got unlucky and leaves the loose reasoning exactly where it was.
Sorting a review by result produces a closed loop in which the poorly reasoned successes are never reached, however many times the review is run.
Try it out

Which of the four combinations gets reviewed least?

Private Wealth Management Bootcamp — Fin Maverick

How is a decision scored before the outcome exists?

Everything so far has been diagnosis. Here is the method, and the test it has to pass is severe: every question in it must be answerable on the day, by somebody who has no idea what is going to happen. A scoring method that needs the outcome is not a scoring method. Such a method is the outcome wearing a different hat.

Four questions, offered as one workable set rather than as a standard issued by anybody, and short enough to be answered in the time it takes to make a cup of tea.

Four questions, in this order, all of them answerable before anything happens. 1 2 3 4 WAS A REASON WRITTEN DOWN? in words, at the time, not reconstructed afterwards WAS A DATE SET? by which the expected thing should have happened WAS AN ENDING OBSERVATION NAMED? something observable that would close the position early IS THE STANDARD A FACT ABOUT THE THING? or is it a fact about the decider's own history, such as the price paid One point each, four available. One workable set, offered as an illustration and not as anybody's standard.
Each of the four questions is settled by what exists on the day, so the total can be written into the record before there is any result to be influenced by.

Question one, the written reason, is the one everything else depends on. Russo and Schoemaker made the case for the written record in Decision Traps, 1989, and the argument is simply that memory of a reason is reconstructed after the fact and reconstructed in the direction of what happened. If it was not written, there is nothing to score. Atul Gawande, in The Checklist Manifesto, 2009, makes the related point about why a fixed short list beats a good intention: the list runs the same way on a tired day as on a sharp one.

The same reason, written down and remembered. WRITTEN ON 12 OCTOBER RECALLED AFTER THE RESULT sell it when it gets back to Rs 3,00,000/- I was uneasy about it and had meant to look at it again CHECKABLE, AND FIXED MOVES WITH WHAT HAPPENED A recalled reason drifts towards the result, and it drifts sincerely, which is what spoils it as evidence. ONLY THE LEFT PANEL CAN BE SCORED, AND ONLY BECAUSE IT WAS WRITTEN
A remembered reason shifts towards whatever happened while a written one cannot, which is why the first question of the four carries all the others.

Question two, the date, converts a hope into something that can expire. Without a date, a position can be held indefinitely on a reason that stopped applying months ago, and nobody notices because nothing ever falls due.

What a date does, drawn as the moment it creates. NO DATE SET nothing ever falls due, so nobody ever reads the reason back and the position can be held for years on a reason that expired A DATE SET THE DATE ARRIVES one moment exists at which the reason has to be looked at again and that single moment is the whole of what the question buys
A date creates one moment at which somebody has to read the reason back, and without it no such moment ever arrives at all.

Question three, the ending observation, is falsification wearing working clothes. Karl Popper argued in The Logic of Scientific Discovery, 1934, that what separates a claim worth holding from one that is not is whether anything could show it to be false. Applied to a decision, the question is: what would have to be seen to conclude that the reasoning was wrong, and would it actually be seen? A reason with no such observation attached cannot be wrong. Being unable to be wrong sounds like a strength and is the opposite of one.

Question four, the reference pointThe level a change is measured against, which may be a fact about the decider., catches the most common single fault. Daniel Kahneman and Amos Tversky showed in Econometrica in 1979 that people evaluate outcomes as gains and losses against a reference level rather than as final positions, and that the reference level is often just where they happened to start. A purchase cost is exactly that. A purchase cost is a fact about the decider’s own history and carries no information whatsoever about the thing being held. If the standard being applied is a number out of the decider's own past, question four scores zero.

What each question is actually looking for. THE QUESTION WHAT SCORES A POINT WHAT SCORES NOTHING A written reason a sentence dated that day a clear memory of it A date a day it falls due in due course, eventually An ending observation a thing that could be seen if it stops looking good An honest standard a fact about the thing held the price paid for it
The right hand column is what most entries actually contain, and every item in it fails for the same reason, which is that none of them can be checked.

Two things stand out about the set. The four questions carry no opinion about what should have been decided: the moment a scoring method starts marking the conclusion, it has become the reviewer’s own view of the world with a number attached. The four questions also say nothing about money. The same four questions run on a decision to change jobs. Was a reason written down? Was a date set by which the new work should have felt like the right move? Was an observation named that would say it was not working? Is the standard a fact about the job, or is it the salary held before, a number out of the decider’s own past? The method does not know it is being applied to a career, and that is the test it has to pass to be worth anything.

The same four questions, on a decision with no money in it. THE DECISION: TAKE THE OTHER JOB 1. A REASON, WRITTEN the work is closer to what I want to do 2. A DATE nine months, then read this back 3. AN ENDING OBSERVATION still no work of that kind by month six 4. AN HONEST STANDARD the work itself, not the old salary SCORE 4 OF 4, AND THE JOB MAY STILL TURN OUT BADLY The two statements are not in conflict, and holding both at once is the skill.
Scored on a job move rather than a holding, the same four questions still work, which is the test any method of this kind has to survive.
Try it out

Why is the reference point question in the scoring set?

What does the scoring do to the decision of 12 October?

Now run the four questions on one real entry, using nothing that was not on the table that day. The invented Palash decision log records that on 12 October, Meera Sundaram sold Suvarna Chemicals Limited whole at Rs 4,60,000/- against a cost of Rs 4,00,000/-, booking Rs 60,000/-, or 15.0 per cent. In the same sitting she kept Kesari Logistics Limited, which had cost Rs 3,00,000/- and stood at Rs 1,95,000/-, down Rs 1,05,000/-. She recorded that she would sell it when it got back to Rs 3,00,000/-.

The rest of the holding shows where that decision sat. On 30 September the holding stood at four positions against a cost of Rs 13,00,000/-. The Vindhya index scheme had gone from Rs 3,00,000/- to Rs 3,36,000/-, up 12.0 per cent, and Suvarna Chemicals Limited from Rs 4,00,000/- to Rs 4,60,000/-, up 15.0 per cent. The Nilgiri mid-cap scheme had fallen from Rs 3,00,000/- to Rs 2,55,000/-, down 15.0 per cent, and Kesari Logistics Limited from Rs 3,00,000/- to Rs 1,95,000/-, down 35.0 per cent. Two positions stood in gain, worth Rs 7,96,000/- between them, and two stood in loss, worth Rs 4,50,000/-. The four values total Rs 12,46,000/- against the cost of Rs 13,00,000/-, so the holding was down Rs 54,000/-, or 4.2 per cent.

The four positions at 30 September, value drawn against cost. THE POSITION VALUE, WITH THE COST MARKED the Vindhya index scheme the Nilgiri mid-cap scheme Suvarna Chemicals Limited Kesari Logistics Limited Rs 3,36,000/- Rs 2,55,000/- Rs 4,60,000/- Rs 1,95,000/- The upright mark on each row is what that position cost, drawn on the same scale as the bar. Cost Rs 13,00,000/- against value Rs 12,46,000/-, so down Rs 54,000/- and 4.2 per cent. All figures invented.
Two positions stood above their cost mark and two below it, and the holding as a whole was down Rs 54,000/- on a cost of Rs 13,00,000/-.

That pairing has two named mechanisms in it. Realising the position standing in gain while holding the position standing in loss is the disposition effect, set out under the disposition effect and measured by Terrance Odean in the Journal of Finance in 1998. Using the purchase cost as the level everything is measured against is the reference point at work, set out under prospect theory and drawn from Kahneman and Tversky in Econometrica in 1979. Neither of those is what is being scored here. The scoring does not care which holding she picked; it cares whether the reasoning behind the pick could be checked.

The questionWhat the entry of 12 October actually showsScore
Was a reason written down?yes, in her own words: sell when it gets back to Rs 3,00,000/-1
Was a date set?no date at all, so the reason can never fall due0
Was an ending observation named?nothing named that would close the position early0
Is the standard a fact about the holding?no, it is Rs 3,00,000/-, which is what she paid0
Totalone of the four questions met, which is 25.0 per cent1 of 4
One logged entry, scored on the day it was written. 12 OCTOBER: KEEP KESARI LOGISTICS, SELL WHEN IT REACHES Rs 3,00,000/- 1 0 0 0 A REASON, WRITTEN it is there, in her own words, dated that day A DATE none set, so the reason can never fall due AN ENDING OBSERVATION nothing named that would close the position early AN HONEST STANDARD the standard is Rs 3,00,000/-, which is what she paid MET NOT MET NOT MET NOT MET 1 of 4
Scored question by question, the entry of 12 October meets one of the four and misses three, and all four verdicts were available on the day it was written.

One of four, or 25.0 per cent. Now the part that matters more than the number. Every one of those four verdicts could have been reached on 12 October, before a single thing had happened. Availability on the day is the entire practical difference between the two measurements, and it is why the scoring is worth the four minutes it takes.

Everything the score needs sits on the left. Everything the result needs sits on the right. ALL FOUR QUESTIONS ANSWERABLE the reason, the date, the ending observation and the standard THE RESULT ARRIVES Rs 36,800/- forgone and Rs 39,000/- further lost, Rs 75,800/- in all 12 OCTOBER 31 MARCH FOLLOWING the score is complete here, and could have been written into the entry itself nothing on this side is allowed to reach back and change the score Dates fall inside one unnamed year. Every figure here is invented.
The four answers are all settled on the left hand date while the result only exists on the right hand one, and no arrow runs backwards between them.

A scoring method gets misused at exactly this point. State the limit of the score plainly. A score of 1 of 4 does not say the decision was wrong. The score says the reasoning was not in a condition to be checked. Those are separate claims and only the second one is being made. Keeping Kesari Logistics Limited on 12 October might have been entirely sensible. On this record there is no way to find out, because the entry contains no date at which anybody would look again and no observation that would have said the reason had failed. A 1 of 4 decision can be followed by an excellent result, and often is. The mismatch is not a defect in the scoring. The mismatch is the thing the scoring is built to survive.

The record does show what happened next, and it is worth stating so long as it is stated with its limits. By 31 March following, Suvarna Chemicals Limited had risen 8.0 per cent after the sale, so the Rs 4,60,000/- released would have been Rs 4,96,800/-, a forgone Rs 36,800/-. Kesari Logistics Limited had fallen a further 20.0 per cent, from Rs 1,95,000/- to Rs 1,56,000/-, a further Rs 39,000/- lost. The pair costs Rs 75,800/-. One case is not evidence that any rule works, and this one is offered as an illustration of the arithmetic and nothing more.

Change the result. Look at what happens to the score. IF THE RESULT HAD BEEN POOR 1 OF 4 as it in fact was, costing the pair Rs 75,800/- by 31 March following IF THE RESULT HAD BEEN GOOD 1 OF 4 had the kept holding recovered, as it might just as easily have done = The reasoning was the same reasoning either way, and the reasoning is what is being scored.
Attach a poor result or a good one and the score reads one out of four in both places, which is the property that makes it worth recording at all.

Notice what would have happened in the second panel. Had Kesari Logistics Limited recovered, the entry would have been remembered as patience rewarded, filed as a success, and used as the model for the next one. Same missing date. Same missing observation. Same purchase cost doing the work of a standard. Poor reasoning with a good result is the dangerous cell from earlier, arriving in a specific record on a specific day.

Try it out

The decision of 12 October scores 1 of 4. Does that mean it was wrong?

Try it out

Before the control below is moved: what will happen to the process score as the result changes?

Play with it

Change the result and watch the score refuse to move

One variable moves: the result attached to the decision of 12 October, from a cost of Rs 1,50,000/- through to a gain of Rs 1,50,000/-. The default is the result actually recorded, a cost of Rs 75,800/-, being Rs 36,800/- forgone on the sale and Rs 39,000/- further lost on the holding. The process score is the four questions from above: a written reason met, a date not set, no ending observation named, and a standard taken from the purchase cost, so 1 of 4, or 25.0 per cent. Every one of the four questions is answered from what existed on 12 October, so the score reads 1 of 4 at every setting of the control. A measurement that refuses to move when the result moves is doing exactly what it was built to do.

a cost of Rs 1,50,000/-a cost of Rs 75,800/-a gain of Rs 1,50,000/-
One control. Two readings. Only one of them moves. THE RESULT ATTACHED, WHICH MOVES: A COST TO THE LEFT, A GAIN TO THE RIGHT a cost of Rs 75,800/- 1,50,000 1,00,000 50,000 0 50,000 1,00,000 1,50,000 the result actually recorded, a cost of Rs 75,800/- THE TWO SCORES, ON THE SAME RESULT SCALE AS THE PANEL ABOVE the process score, 1 of 4 how it gets described SCORE 0 1 2 3 4 The nought to four scale of approval is an illustrative one, drawn to show a shape and not to measure anything.
The result attached, what moves
Rs 75,800/- cost
The process score
1 of 4
Held constant, as a percentage
25.0 per cent
How it then gets described
a bad decision

With the decision of 12 October costing Rs 75,800/-, the reasoning scores 1 of 4, which is 25.0 per cent, and a decision costing that much tends to get described afterwards as a bad decision.

Educational illustration. The four questions are an illustrative set rather than anybody’s published standard. Every result other than the recorded cost of Rs 75,800/- is hypothetical. Figures held in whole rupees throughout.

What is a Devil’s Advocate, and what makes the role fail?

A devil's advocateSomebody assigned to argue against a case within the group. is somebody given the job of arguing against a case from inside the group considering it. The phrase comes from a formal office that existed to put the case against a candidate before a decision was confirmed, and the idea it carries is a good one: a case that has never been argued against has not been tested, it has only been agreed with. Applied to the four questions above, the role is what tends to produce an answer to question three. The person arguing the other way is exactly the person who will say out loud what would have to be seen for the case to be wrong.

The role is one of the standard corrections, and Baruch Fischhoff set out the general shape of such corrections in his chapter on debiasing in Judgment under Uncertainty, 1982: a correction has to target a specific failure, and it has to run at the moment the failure happens rather than afterwards. A devil’s advocate targets one specific failure, a case that goes unchallenged because everybody in the room already agrees.

The role fails when everybody in the room knows it is a role. The objection arrives already labelled. The objection is heard not as somebody’s view but as somebody performing an assigned duty, and an argument heard that way is discounted before it has been weighed on its content. The person is thanked, the point is noted, and the case proceeds exactly as it would have done. Worse, the group now feels it has tested the case, so the appearance of a challenge has replaced the challenge. Nobody has measured how often an announced role fails this way. Treat the pattern as a regularity people report from inside such rooms rather than as a finding.

What decides whether an objection does any work. AN OBJECTION IS RAISED IS IT HEARD AS ASSIGNED? YES NO DISCOUNTED IN ADVANCE the argument arrives already labelled as not sincerely meant WEIGHED ON ITS CONTENT the case is actually tested, which is all the role was for the case survives untested the case is stronger, or it is dropped
Whether an objection changes anything turns on how it is heard rather than on how good it is, which is why an announced role can defeat itself.
What a real objection leaves behind in the record. THE OBJECTION somebody says why this might be wrong WHAT IT NAMES the thing there would be to see to be wrong WHAT IT BECOMES a line in the entry, and a point on question three AN OBJECTION THAT LEAVES NOTHING IN THE ENTRY HAS NOT HAPPENED which is a usefully hard test to apply to a review that felt thorough at the time, and it is the same test whether the room held nine people or one.
An objection is only worth the meeting it took if it ends as a written observation somebody could later go and look for.

So what does work? Four adjustments, and each one attacks the labelling rather than the argument. Objections are asked for in writing before any discussion, so nobody has heard the room first. The job rotates, so it is never one person's known position. The request is for the opposite conclusion rather than for an attack. Charles Lord, Mark Lepper and Elizabeth Preston showed in the Journal of Personality and Social Psychology in 1984 that instructing people to consider the opposite reduced the pull of what they already believed, and considering the opposite is a different instruction from telling somebody to object. And for a person deciding alone, with no group and nobody to assign, the case against is written out by hand before deciding, since a sheet of paper does not know it is playing a part.

Four adjustments, all of them aimed at the label rather than the argument. WRITTEN BEFORE THE DISCUSSION nobody has heard the room yet, so nothing has been adjusted to it THE JOB ROTATES so objecting never becomes one person's known standing position ASK FOR THE OPPOSITE reasons the other conclusion might hold, which is not an instruction to attack anything DECIDING ALONE the case against written in one's own hand, before the decision, on paper
Each adjustment removes the label that lets an objection be discounted, which is a different repair from asking for better objections.
Try it out

What makes the devil's advocate role fail?

Investment Banking Analyst Bootcamp — Fin Maverick Cleaning Financial Data — free micro-course from Fin Maverick

If outcomes cannot be used, what is left to judge by?

The heading is the fair objection to everything above, and it deserves a concrete answer rather than a principle. Take away the result and a surprising amount is left, but only if somebody wrote things down at the time. Everything on the list below is a fact about the record rather than about the world.

Was there an entry at all. Did it contain a reason in words. Did the reason commit to something that could later be found false. Was a date set. Was the standard a fact about the thing decided, or about the decider. Was anybody, including yourself, given the job of arguing the other way. And did the review actually happen on the date it was supposed to. Seven checkable things, and the result is not among them.

What a review can actually check, and what it only pretends to. CHECKABLE, AND ALL OF IT ABOUT THE RECORD NOT CHECKABLE 1. an entry exists at all 2. it contains a reason, in words 3. the reason could be found false 4. a date was set 5. the standard is about the thing held 6. somebody argued the other way 7. the review happened on its date the result, as a report on the reasoning what would have happened had she decided otherwise how confident anybody felt at the time, recalled later ALL THREE ARRIVE LATE
Seven things about a decision can be checked from the record alone, and every one of them is settled before any result exists.

The logged record shows how thin most entries are on exactly these points. A written reason exists on 84 of the 240 decisions, or 35.0 per cent. And 71 of the 240, or 29.6 per cent, were taken within 48 hours of a news item. Of the 96 buys, 41 followed a media mention within three days, or 42.7 per cent, against 11.0 per cent of the eligible list being mentioned at all in a given week. Every one of those counts measures the record and the timing, and not one of them measures a return.

Something did move after 4 November, when 20 of the 60 adopted a written checklist. Across the fifth to eighth quarters those 20 recorded a written reason on 34 of 41 decisions, or 82.9 per cent, against 19 of 63, or 30.2 per cent, for the other 40. Their realisation ratio, gains realised against losses realised, fell from 3.2 to 1.6. Both of those are changes in behaviour and in what the record contains, and neither of them is a change in return. The distinction is not a technicality; it is the difference between what the record can support and what somebody would like it to say.

How often a written reason exists, in three groups. all 240 decisions the 20 who adopted the list the other 40 35.0 per cent 82.9 per cent 30.2 per cent 84 of 240, then 34 of 41 and 19 of 63 across the fifth to eighth quarters. Four and a half pixels to a point. Every one of these is a measurement of the record, and none of them is a measurement of a return.
Written reasons ran at 82.9 per cent among those who adopted the checklist against 30.2 per cent among the rest, which is a difference in what the record holds.
What moved, and what was never measured at all. WHAT MOVED WHAT WAS NOT MEASURED before 4 November 3.2 to 1 after 4 November 1.6 to 1 gains realised against losses realised, which is a change in behaviour NO RETURN FIGURE none recorded before adoption and none recorded after it
The realisation ratio halved from 3.2 to 1.6 while the empty panel beside it stays empty, because no return was recorded on either side of the change.
How this gets used in practice

Sorting the pile by score instead of by result

Devika Rao, the adviser at the invented Palash Advisory Services Private Limited, runs a quarterly review, and the sort order of the pile is the one structural change to it that matters. Sorted by result, the review opens the losers and never opens the poorly reasoned entries that happened to work. Sorted by the process score, it opens every entry that could not be checked, whichever way each one turned out. The second sort takes the same amount of time and looks at a completely different set of decisions.

The same sort works for one person with nobody to answer to, which is the harder case and the more common one. Once a quarter, the entries from three months ago are opened, each one is scored out of four without looking at what it is worth now, and the entries scoring zero on the date question are counted. The count of failed date questions is the only number in the exercise, and it is entirely within the decider’s control. A lender reading a credit file does a version of the same thing when it asks whether the assumptions were written down and dated rather than whether the borrower ended up repaying; an analyst does it when the note is judged on whether it named what would prove it wrong.

One caution for anybody deciding on behalf of somebody else. A process score is not a substitute for the conduct and suitability duties that apply to a registered intermediary. The Securities and Exchange Board of India sets those duties and publishes them at sebi.gov.in, and they bind the adviser whatever any internal scoring says.

The same pile, sorted two ways, opening two different sets of entries. SORTED BY RESULT SORTED BY PROCESS SCORE poor result, poor reasoning: opened poor result, sound reasoning: opened good result, sound reasoning: skipped good result, poor reasoning: skipped THE DANGEROUS CELL IS NEVER OPENED scored 0 of 4: opened scored 1 of 4: opened scored 2 of 4: opened scored 3 or 4 of 4: skipped THE RESULT NEVER ENTERS THE SORT
Sorting by result leaves the poorly reasoned successes permanently unopened, while sorting by score reaches them without ever consulting how anything turned out.
Cleaning Financial Data teaches you to find the errors that survive every check and break every model.

What is not claimed about results?

One limit belongs at the end, stated plainly. A record of this shape cannot show that a better way of deciding produces a better return. The claim is not measured by such a record, not implied by it, and not quietly assemblable from it either.

The reasons are specific rather than a general shrug about evidence. Look at each of them. The checklist was adopted on 4 November and the log was read back on 31 March following, one stretch of time with no comparison group running alongside it. The 240 decisions are spread across 60 people, four decisions each. There is no measurement of return before adoption and none after, so there is nothing on either side of a comparison that has not been made. And the 20 who adopted the checklist chose to adopt it, so they were already the kind of people who adopt a checklist. That willingness is a difference between the two groups that has nothing to do with the checklist.

What a claim about returns would need, and what is actually here. WHAT WOULD BE NEEDED WHAT THIS RECORD HAS VERDICT a comparison group one stretch, nothing alongside MISSING returns measured before and after no return measured at all MISSING enough decisions per person 240 across 60 people, so four each MISSING groups that differ only in the method the 20 chose to adopt it themselves MISSING Four requirements, none of them met, which is why the claim is refused rather than merely left unstated.
A claim about returns would need four things this record does not have, so the refusal is a matter of arithmetic rather than of caution.

So what is the process scoreA fixed set of questions answerable before any outcome exists. actually for? One thing only. The score makes the reasoning inspectable afterwards. Inspectable reasoning is worth having because a decision whose reasoning cannot be examined cannot be improved, and improving the next decision is the only use anybody has ever had for reviewing the last one. Whether examinable reasoning also earns more is a separate question needing a completely different kind of evidence, and this record is not that evidence. Concluding otherwise would be the exact error the scoring exists to correct: treating a result as though it settled a question about reasoning.

Try it out

Does a higher process score produce better returns?

The pre-mortem and the post-mortem are procedures with their own steps and are covered separately, as are the journal and the debiasing methods that supply the record this scoring reads from. The four questions are one workable set rather than a standard issued by anybody, and somebody who prefers five different questions has lost nothing so long as every one of them can be answered on the day.

Sources

SourceDocumentSite
Jonathan Baron and John Hersheythe paper demonstrating outcome bias in the evaluation of decisions, Journal of Personality and Social Psychology, 1988ssrn.com
Annie DukeThinking in Bets, 2018, in which the word resulting is set outcited to the book itself
Daniel Kahneman and Amos Tverskythe 1979 paper in Econometrica in which outcomes are evaluated against a reference levelssrn.com
Karl PopperThe Logic of Scientific Discovery, 1934, on what it takes for a claim to be testablecited to the book itself
J Edward Russo and Paul SchoemakerDecision Traps, 1989, on keeping a written record of reasoningcited to the book itself
Charles Lord, Mark Lepper and Elizabeth Prestonthe paper on considering the opposite, Journal of Personality and Social Psychology, 1984ssrn.com
Baruch Fischhoffthe chapter on debiasing in Judgment under Uncertainty, 1982cited to the book itself
Atul GawandeThe Checklist Manifesto, 2009, on why a fixed short list beats an intentioncited to the book itself
Terrance Odeanthe paper measuring the realising of positions in gain against the holding of positions in loss, Journal of Finance, 1998ssrn.com
Securities and Exchange Board of Indiathe conduct and suitability requirements applying to registered intermediariessebi.gov.in

Meera Sundaram, Devika Rao, Palash Advisory Services Private Limited, the Palash decision log, the Palash 100 index, the Vindhya index scheme, the Nilgiri mid-cap scheme, Suvarna Chemicals Limited and Kesari Logistics Limited are invented.
Educational material. Not advice on any investment, tax, budget or market position.

Covered in this topic

Subtopics

Devil's AdvocateProcess Quality vs Outcome Quality
← Previous
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.