The Pre-Mortem and Process Quality: Imagining Failure
A pre-mortem assumes the failure has already happened and asks what caused it. The single change of tense is the technique. Asking what might go wrong produces a short polite list; asking why it did go wrong produces a longer and far more specific one. Klein named the practice; the effect underneath it is prospective hindsight.
Ask a group what could go wrong with a plan and watch what happens to the room. Somebody says the market could turn. Somebody says it might take longer than expected. Somebody who does not want to be difficult says nothing at all. Three minutes later the list is closed and the plan proceeds, and every objection on the list is so general that none of them can be acted on. The problem is not that the group lacks doubts; it is that the question asked for speculation, and speculation is the one mental task people are worst at.
Now one word changes. The question is no longer what could go wrong. Instead: it is six months from now, this decision has gone badly, and the room is sitting there explaining it. Write down why. The room fills. The causes get longer, more concrete and more uncomfortable, and several of them name things that were sitting in plain view the whole time. Nothing about the plan changed between the two questions. The only thing that changed was the tense.
The change of tense is the whole of the technique, and the rest is what to do with it. The sequence covers the pre-mortemA review that assumes the failure has happened and asks what caused it. itself and the order its steps run in, the post-mortemA review conducted after the outcome is known. that comes afterwards and what it can see that the pre-mortem could not, and the difference between an objector placed inside a group and a team placed outside it. The idea that ties all four together is that the reasoning is the only part of a decision anybody actually controls.
What is a pre-mortem, and why is the tense the whole technique?
A post-mortem is run after something died. A pre-mortem is run before, on a death that has not happened, and the trick is that it is described as though it already had. The facilitator does not ask for worries. The facilitator states a fact: the date is such and such, this decision has failed, and the failure is now common knowledge. Everyone in the room is then asked to write the explanation. Gary Klein set the practice out in Performing a Project Premortem in the Harvard Business Review in 2007, and the whole instruction fits in two sentences.
The reason the wording matters so much is that it swaps one mental task for a different one. Asking what could go wrong asks a person to generate possibilities out of nothing. There is no anchor, no target, and no way to know when the list is finished, so people produce two or three safe items and stop. Asking why it went wrong hands the person a fact and asks for an account of it, and explaining a fact is something people do fluently, constantly, and at length. The pre-mortem does not make anybody more imaginative; it moves the same imagination onto a task it is already good at.
The effect the technique leans on has a name and a paper behind it. Deborah Mitchell, J. Edward Russo and Nancy Pennington called it prospective hindsightExplaining a hypothetical outcome as though it were settled fact. in Back to the Future, published in the Journal of Behavioral Decision Making in 1989, and what they reported is that describing an event as certain rather than possible increases the number of reasons people can produce for it. Klein's contribution was to turn that into a meeting procedure rather than a finding. The two citations are separate on purpose: one is the mechanism, the other is the practice, and quoting the practice without the mechanism is how the technique gets softened back into an ordinary risk discussion.
What is the pre-mortem instruction, stated exactly?
How to Run a Pre-Mortem for a Financial Decision: what order do the steps run in?
The procedure is short enough to memorise, and the order is not decorative. Each step exists because of a specific way the previous one gets undone. The pre-mortem runs on any decision about to be taken that would be expensive to reverse, whether or not money is involved, and twenty minutes is the right length rather than an afternoon.
- Fix the decision in writing first. One paragraph: what is being done, by when, and on what reasoning. A pre-mortem on a decision nobody has written down becomes a pre-mortem on whatever each person privately assumed the decision was.
- State the failure as a settled fact, with a date on it. Not it might fail. It is the fifteenth of the month six months out, this has gone badly, and everybody here knows it. The date does the work, because a failure without a date cannot be explained, only imagined.
- Everybody writes causes alone, in silence, for two minutes. No discussion, no going round the table first. Silent writing is the step that gets dropped, and the step the whole procedure rests on.
- Read every cause out, one at a time, without debate. Go round until the causes run out. Argument is allowed later; it is not allowed while the list is still being built, because the first challenge ends the collection.
- Keep the causes that could actually be checked. A cause that can be tested, watched, priced or diarised now is worth carrying. A cause that amounts to conditions might change is not, and cutting it is not pessimism about the cause, it is honesty about what could be done with it.
- Change the decision, or write down why not. Either the plan moves or the record states which cause was accepted and on what grounds. A pre-mortem that ends with everybody nodding and nothing written has produced nothing.
Step three is the one that decides whether the exercise worked, and it is the one everybody softens. Writing alone before anybody speaks is not a courtesy to shy participants. Silent writing is there because the first cause spoken aloud reshapes every cause after it, so a room that discusses before it writes produces one cause with five endorsements rather than six causes. Anybody who has watched a group settle on a view within ninety seconds of the loudest person speaking has seen the effect the step exists to block.
Why does everybody write their causes alone, in silence, before anybody speaks?
Does any of this work on a decision that has nothing to do with money?
The procedure has to work there, and portability is a test rather than a nicety. Not one of the six steps mentions a price, a market, an instrument or a portfolio. The six steps mention a decision, a date, a failure and a written record. If a procedure only works when there is something to buy, it was written too narrowly, and it will quietly stop working the moment the decision at hand is a job offer, a house, a course of treatment or a move to another city.
Take a household deciding whether one earner should leave a salaried job to run a stall. The pre-mortem instruction is the same: it is eighteen months from now, the stall has closed, write down why. The causes that come back are specific in a way that a list of worries never is. The pitch was next to one office building and the building emptied. The takings covered the stock but never the rent. Nobody set a date to check whether it was working, so the decision to stop was made by the bank balance rather than by the household. Every one of those is a cause somebody could have watched for from the first week, which is the only kind of cause worth keeping.
How to Run an Investment Decision Post-Mortem: what changes once the outcome is known?
The post-mortem is the same shape of exercise run from the other end, and it looks deceptively easy because the answer is already sitting on the table. The ease is exactly the danger. Once the outcome is known, the account of what was being thought beforehand quietly rearranges itself to fit, and it does so without any sensation of rearranging. So the post-mortem procedure spends its first step on protecting the record rather than on analysing the result.
- Reread the written record before any discussion. The reasoning as it was written on the day comes first and memory comes second, never the other way round. J. Edward Russo and Paul Schoemaker made the written record the centre of Decision Traps in 1989 for this reason: without it there is nothing to check the recollection against.
- Write down what actually happened, in figures, on its own. Keep the result physically separate from any story about why. Mixing them on the same line is how the story starts borrowing authority from the numbers.
- Mark which of the stated causes actually occurred. Go back to the pre-mortem list and tick the ones that came true. Then add, in a different column, the causes that nobody listed at all, because those are the ones that show where the procedure is blind.
- Judge the reasoning only on what was knowable on the day. Anything that arrived afterwards is excluded from the verdict on the reasoning, however tempting it is. Judging on what was knowable is the step that takes discipline, and it is the reason the written record exists.
- Write two verdicts, never one. One on the reasoning, one on the result. A single combined verdict is not a shorter version of two, it is a different and much worse thing, and the next section is entirely about why.
- Change the procedure, not the memory. The output of a post-mortem is an amended checklist for next time. The output is not a revised account of what was believed then.
What can the post-mortem see that the pre-mortem could not?
Three things, and they are worth naming separately. First, the post-mortem can see which of the imagined causes actually turned up, and this is the only feedback the pre-mortem ever gets. Second, the post-mortem can see the causes nobody imagined at all. The blank column of unlisted causes is the most useful part of the record. The blank column shows what kind of cause the procedure does not reach. Third, the post-mortem can see the result. The result is precisely the information that has to be kept out of the verdict on the reasoning.
So the post-mortem holds one extra fact and spends half its procedure making sure that fact does not contaminate the other half. The arrangement sounds paradoxical until the timeline is laid out. The pre-mortem sits before the decision and can only use what was knowable then. The post-mortem sits after the outcome and can use everything. The verdict on the reasoning, though, has to be reached using only the left hand part of that timeline, even though the person reaching it is standing at the right hand end.
Process Quality: what is it, and why is it the only part of a decision anybody controls?
Any decision splits in two. One half is the reasoning: which question was asked, which information was gathered, which alternatives were weighed, what was written down and what was said would change the decider’s mind. And there is the result: what actually happened afterwards. Process qualityWhether the reasoning was sound given what was knowable at the time. is a verdict on the first of those. Outcome qualityHow the decision actually turned out. is a verdict on the second.
Only one of those two was ever open to influence. The decider chose the question. The decider chose how hard to look. The decider chose whether to write anything down. The decider did not choose what a few thousand strangers would do afterwards, or what news would arrive in February, or which of two roughly equal candidates would take the role that was turned down. The reasoning is the only part of a decision anybody controls, and so the only part it makes sense to grade. A student can control how they revise and not whether the paper suits them; a surgeon can control the preparation and not the patient's response; and the logic is identical in all three places.
Judging the reasoning has one hard requirement, and it is the reason the written record from earlier in this sequence keeps reappearing. Reasoning that was never written down does not sit still. Once the result is known, recollection of what was being thought shifts to accommodate it, smoothly and invisibly, so that a decision that turned out well is remembered as having been better reasoned than it was, and one that turned out badly is remembered as having been full of doubts that were never actually voiced. A post-mortem without a record does not judge the reasoning that ran; it judges a reconstruction built to be consistent with an answer that is already known.
So what is actually being graded when a process is graded? Four things, and it helps to have them as a fixed list rather than as a general sense of whether somebody was being careful. The fourth is the one people skip, and it comes from an idea far older than behavioural work: Karl Popper argued in The Logic of Scientific Discovery in 1934 that a claim earns its standing from what would refute it, not from what agrees with it. Applied to a decision, that means the case has to name in advance the observation that would show it was wrong.
Why is grading a process after the outcome is known unreliable, unless the reasoning was written down?
How to Separate Process Quality From Outcome Quality: why must the separation be procedural?
Everybody agrees with the separation in the abstract. Ask any group whether a decision should be judged by the reasoning rather than by the result and every hand goes up. Then hand the same group two cases, one where careful reasoning lost money and one where careless reasoning made money, and watch the verdicts. The careful one gets picked apart. The careless one gets a shrug and a compliment. Agreement is not the problem. The separation fails not because people reject it but because nothing in an ordinary week ever forces them to apply it.
Which is why the fix is a piece of scheduling rather than a resolution to try harder. There are exactly two ways a review can come into existence. Either somebody arranges it in advance, at the moment the decision is taken, with a date written next to it. Or it happens because somebody became unhappy. Reviewing only on disappointment sounds efficient: why review what did not go wrong? The triggered arrangement is also the more common one in both households and firms, and it is structurally incapable of ever reaching the one cell that most needs reaching.
Follow that through. A review triggered by disappointment fires when the result is poor. A triggered review therefore examines careful reasoning that lost and careless reasoning that lost. Both are worth examining. A triggered review never once fires on a good result, so the reasoning behind everything that worked goes uninspected for ever, including all the reasoning that was terrible and worked anyway. A review scheduled in advance fires regardless of the result, costs some time spent on decisions nobody is unhappy about, and is the only arrangement that sees the whole picture.
Why must a post-mortem be scheduled in advance rather than triggered by a poor result?
Put the two verdicts on two axes and four cells appear, each one worth sitting with before touching the control below. Sound reasoning with a good result is earned, and the correct response is to repeat the reasoning, not to celebrate the result. Sound reasoning with a poor result is the one people find hardest. The honest response is to keep doing exactly what was done. Poor reasoning with a poor result is uncomfortable but harmless in the long run. The discomfort itself prompts a look. And the cell that quietly destroys people is poor reasoning with a good result, rewarded and repeated and never examined by anybody.
Before the control below is used: which of the four cells is the most dangerous over a working lifetime?
Move one decision around the grid and watch what the reviewer owes it
One variable moves: which of the four combinations is in view. Everything else is held still. The four cells are sound reasoning with a good result, sound reasoning with a poor result, poor reasoning with a good result, and poor reasoning with a poor result. The 12 October entry in the invented Palash decision log sits in the last of those: by 31 March the sold holding had risen 8.0 per cent, so Rs 36,800/- was forgone, and the kept holding had fallen a further 20.0 per cent, so Rs 39,000/- more was lost, making Rs 75,800/- on the pair. Had the kept holding instead returned to the Rs 3,00,000/- she was waiting for, the reasoning would have been word for word identical and the same entry would have shown Rs 68,200/- in her favour, and would have been repeated untouched. The comparison between the two cells is why the control exists, and one invented case is not evidence about anything.
Poor reasoning, poor result. This is where the 12 October entry sits: Rs 36,800/- forgone on the sale and Rs 39,000/- further lost on the holding, Rs 75,800/- on the pair. One invented case says nothing about how often any cell happens.
What does a pre-mortem produce when it is run on the logged decision?
The invented Palash decision log carries an entry for 12 October. Meera Sundaram intends to sell Suvarna Chemicals Limited whole at Rs 4,60,000/- against a cost of Rs 4,00,000/-, booking Rs 60,000/-, or 15.0 per cent on what she paid. She intends to keep Kesari Logistics Limited, then standing at Rs 1,95,000/- against a cost of Rs 3,00,000/-, until it gets back to the Rs 3,00,000/-. Run the pre-mortem on 11 October, the day before, and remember the exact wording. The instruction is not what could go wrong. The instruction is this: it is 31 March, this decision has gone badly, write down why.
Four causes come back, and they are the sort a room produces in two silent minutes. The holding that was sold carries on rising, so the gain was booked early for no reason connected to the holding. The holding that was kept falls further, so the loss grows while nobody is watching. The recovery target has no date attached, so there is nothing in the calendar that ever forces a look at it. And the target is a purchase cost. A purchase cost is a fact about what was paid rather than a fact about the worth of the holding.
Now put the record beside the list. By 31 March, Suvarna Chemicals had risen 8.0 per cent after the sale, so the Rs 4,60,000/- would have been Rs 4,96,800/- and Rs 36,800/- was forgone. Kesari Logistics had fallen a further 20.0 per cent, from Rs 1,95,000/- to Rs 1,56,000/-, so Rs 39,000/- more was lost. The pair costs Rs 75,800/-, or 5.8 per cent of the Rs 13,00,000/- put in, and two of the four written causes are exactly what happened; and one case is not evidence that a procedure works, which has to be said in the same breath as the number rather than in a paragraph further down. That is a single illustration of what the technique produces, not a demonstration that it produces good outcomes. Four causes were written and two matched. Had none matched, the procedure would have been no worse and no better: a match rate over one case is not a rate at all.
| The leg | What the record shows | Money |
|---|---|---|
| Sold holding, after the sale | Rs 4,60,000/- rising 8.0 per cent to Rs 4,96,800/- | Rs 36,800/- forgone |
| Kept holding, after 12 October | Rs 1,95,000/- falling 20.0 per cent to Rs 1,56,000/- | Rs 39,000/- lost |
| The pair, by 31 March | the two legs added, on an invented log and one case only | Rs 75,800/- |
The fourth bar is doing the real work. Suppose one thing about the world changes and nothing at all about Meera Sundaram: Kesari Logistics drifts back up to the Rs 3,00,000/- she was waiting for. The drift back is Rs 1,05,000/- recovered on the kept holding, against Rs 36,800/- still forgone on the sold one, so the same entry would have shown Rs 68,200/- in her favour. The purchase-cost target would still be a purchase-cost target, the review would still have had no date, and every one of those defects would have been rewarded rather than exposed. Same reasoning, different cell, opposite lesson learned. The pair of outcomes is the entire argument for grading the reasoning on its own.
The pre-mortem produced four causes and two of them turned up in the record. What does the match establish?
Devil’s Advocate vs Red Team: what is the structural difference between them?
Both exist to attack a case before reality does, and they are constantly used as though they were the same arrangement. The two arrangements are not the same. A devil’s advocateOne person assigned to argue against a case within the group. is one member of the deciding group, appointed to argue against whatever the group is minded to do. A red teamA separate group tasked with defeating the case from outside it. is a separate group, outside the deciding one, asked to build the strongest case against and handed nothing except the written case and the evidence behind it.
The difference is position, and position decides what each one is willing and able to say. The devil’s advocate has to sit in the same room tomorrow. The objections are made to colleagues, in front of whoever will be writing the next appraisal, and the role is understood by everybody present to be a role rather than a belief. The assigned quality of the role is what defuses it: an objection everybody knows is assigned carries less weight than the same objection made sincerely, so the group can absorb it, thank the advocate, and proceed unchanged. The advantage is that this person knows the case in detail and can attack it precisely.
The red team pays the mirror-image price. A red team has no stake in the case surviving, no appraisal riding on the mood in the room, and no habit of thought shared with the people who built the case. The absence of all three is exactly why a red team can find objections that nobody inside would have voiced. A red team also does not know the case as well, has less context, and will spend some of its effort attacking things the group had already dealt with quietly. Charles Lord, Lee Ross and Mark Lepper reported in the Journal of Personality and Social Psychology in 1984 that deliberately considering the opposite reduces the pull of the view a person already holds, and both arrangements are attempts to build that instruction into a structure rather than leave it to willpower.
What is the structural difference between a devil’s advocate and a red team?
How to Run a Red-Team Review of an Investment Case: what does the outside group actually do?
A red-team review has three moving parts, and getting them wrong is what turns the exercise into a formality. The brief handed over, the output demanded, and the person the team reports to. Give it the written case and the evidence that was used, nothing else and nothing extra, so that it is attacking the case that actually exists rather than a summary of it. Ask it for the strongest case against, not for a balanced assessment. A balanced assessment is what the deciding group already produced, and duplicating that is the one output with no value. And have it report to somebody other than the people whose case it is attacking.
The output should be short enough to be read: one printed side of case against, one named test that would settle the question, and a list of what the team could not check and why. The list of unchecked claims is the part people leave out, and it is the part that shows how much weight the review can carry. A red team that reports it could not check three of the five load-bearing claims has said more than one that quietly reviewed the two it could reach. The same discipline works on a household decision, where the red team is simply the friend who is asked to argue the other side properly and told not to be kind about it.
Two things break a red team, and both of them look like cooperation while they are happening. The first is reporting lines. If the team's findings are graded by the people whose case is under attack, the findings get softer every round until the exercise produces agreement, and nobody involved will experience that as dishonesty. The second is the request for balance. A red team asked for a fair assessment has been asked to become a second copy of the deciding group, and a second copy was already available for free.
The error this whole sequence exists to correct
The error is grading a process by the result that happened to follow it. Jonathan Baron and John Hershey named it in Outcome Bias in Decision Evaluation, in the Journal of Personality and Social Psychology in 1988, and what they reported is that people rate the identical decision differently once they are told how it turned out. Annie Duke later gave the everyday version of it a one-word name, resulting, in Thinking in Bets in 2018. The finding is not that people are careless. The finding is that knowing the result changes the grade even when the person grading has been asked not to let it.
The asymmetry is what makes outcome bias expensive. The two directions are not equally dangerous. Careful reasoning followed by a poor result gets examined to death: everybody is unhappy, somebody wants an explanation, and the reasoning is picked over line by line, sometimes until a perfectly sound method is abandoned for having had a bad quarter. Careless reasoning followed by a good result gets nothing. Careless reasoning followed by a good result is rewarded and repeated, and no post-mortem is ever run on it. Nobody in any household or any firm calls a meeting to ask why something worked.
So the damage accumulates on the side nobody looks at. Every unexamined success leaves the same defect in place for next time. The defect only ever surfaces on the occasion it finally meets a bad result, perhaps years later and much larger. The separation therefore cannot be left to good intentions. Intention fails in exactly the cell where it is needed: in that cell nothing has gone wrong yet, so nothing prompts anybody to look.
How is any of this used by somebody doing the job?
Devika Rao, the adviser at the invented Palash Advisory Services Private Limited, does not run a twenty-minute pre-mortem on every instruction that arrives. She runs one on decisions that would be expensive to reverse, and she keeps the artefact rather than the meeting. The artefact is a single printed side written before the decision: what is being done, why in one sentence, what would show the reasoning was wrong, the date the review will happen, and the causes the pre-mortem produced that were worth keeping. Six months later that written record is what the post-mortem reads first, and there is nothing else that could serve.
Look at what the invented log actually recorded when 20 of the 60 investors adopted a written checklist on 4 November. Across quarters five to eight they wrote a reason on 34 of their 41 decisions, being 82.9 per cent, against 19 of 63, being 30.2 per cent, for the other 40. Their ratio of realising gains to realising losses fell from 3.2 to 1.6. Eight quarters, 60 people and no comparison group cannot carry a claim that they did better, and no return was measured before or after in any case. The log shows that more decisions became inspectable, and a written procedure is for exactly that and for nothing more.
Atul Gawande made the general version of this argument in The Checklist Manifesto in 2009: the value of a checklist is not that it makes anybody cleverer but that it removes the discretion to skip a step on a busy afternoon. For a person deciding alone, with no adviser and no committee, the same record works unchanged, and the red team is one honest friend. Where a professional is required to record a review as part of a conduct duty, that requirement is set by the Securities and Exchange Board of India at sebi.gov.in.
What can none of these procedures promise?
Results. None of them, not one. The worked instance above is the most persuasive and the most misleading part of the whole sequence. A pre-mortem generated four causes and two of them turned up. Two matches out of four is a single observation from an invented log. The count is not a measurement, it has no comparison group, nobody ran the same decision without the pre-mortem to see what happened, and no version of that arithmetic turns into evidence about how the technique performs.
The same restraint applies to the cohort figures. The 20 investors who adopted a written checklist recorded far more written reasons afterwards and their realisation ratio moved. Nothing about their returns was measured before or after, so no statement about returns can be made, and any sentence that invites such an inference has committed the exact error these four procedures exist to correct. The four procedures do something narrower and more honest than a promise: they make reasoning inspectable after the fact, and they generate causes that a list of worries would not have produced. Whether that pays is a separate question that this case cannot answer.
Do the pre-mortem, the post-mortem and the red-team review improve results?
Sources
| Source | Document | Site |
|---|---|---|
| Gary Klein | Performing a Project Premortem, Harvard Business Review, 2007 | ssrn.com |
| Deborah Mitchell, J. Edward Russo and Nancy Pennington | Back to the Future, the prospective hindsight paper, Journal of Behavioral Decision Making, 1989 | ssrn.com |
| Jonathan Baron and John Hershey | Outcome Bias in Decision Evaluation, Journal of Personality and Social Psychology, 1988 | ssrn.com |
| Charles Lord, Lee Ross and Mark Lepper | the considering-the-opposite paper, Journal of Personality and Social Psychology, 1984 | ssrn.com |
| Karl Popper | The Logic of Scientific Discovery, 1934 | cited to the book itself |
| J. Edward Russo and Paul Schoemaker | Decision Traps, 1989 | cited to the book itself |
| Atul Gawande | The Checklist Manifesto, 2009 | cited to the book itself |
| Annie Duke | Thinking in Bets, 2018, for the word resulting | cited to the book itself |
| Securities and Exchange Board of India | conduct and record-keeping requirements applying to registered intermediaries | sebi.gov.in |
Meera Sundaram, Devika Rao, Palash Advisory Services Private Limited, the Palash decision log, the Palash 100 index, the Vindhya index scheme, the Nilgiri mid-cap scheme, Suvarna Chemicals Limited and Kesari Logistics Limited are invented.
Educational material. Not advice on any investment, tax, budget or market position.
