How Research Post-Mortems Improve Decision Discipline
A post mortem improves research discipline by attaching a verdict to the work instead of to the result. Done properly it separates a wrong mechanism from a wrong magnitude, and those two failures take opposite repairs. Done as a hunt for blame, or run only on calls that turned out badly, it manufactures rules that would have caught the last error and will cause the next one.
Running an examination and using one are separate skills. The ten step procedure and its four way classification are covered separately. The harder question comes one step later: a finished examination sits on the desk, and what happens to it decides whether the hour was worth spending.
What does the exercise actually improve?
One assumption is worth ruling out at the start. A post mortem does not make the next forecast more accurate. If accuracy is the only thing that counts as improvement, the exercise will disappoint.
What the exercise improves is the ability to tell sound work from unsound work without waiting for the market to rule on it. The claim sounds modest and it is not. Think about a cook who has been running a stall outside one office building for four years. On a Thursday the queue is long and on the Friday it is empty. The office had a training day, so neither day tells him anything about whether his sambar is good. If the only signal he trusts is the queue, he will change the recipe after every quiet Friday and change it back after every busy Thursday, and after two years he will be cooking something nobody chose. The information he actually needs sits in the pot, not in the queue.
Research is the same shape with a longer lag. A view about a company can be beautifully argued and lose money for a year because a completely separate thing happened, and a view can be carelessly argued and make money because the argument was never what moved the price. If the desk only ever grades work by the result, it keeps whatever was lucky and drops whatever was unlucky, and its method drifts towards noise. Grading a call by its outcome instead of by the thinking inside it has a name of its own, resulting, and the name belongs to Annie Duke, Thinking in Bets, 2018. The word is common enough to be worth knowing; the idea behind it is covered separately.
So the improvement is a governance one rather than a predictive one. A desk that can grade its work independently of the outcome can keep a good mechanism after a bad year and drop a bad one after a good year. An outcome-graded desk finds exactly that pair of moves impossible.
Why does a wrong mechanism need a different repair from a wrong size?
The difference between the two kinds of error is the practical centre of the argument, and it is the part that gets fumbled most often in practice. Two calls can both be wrong and be wrong in ways that have almost nothing in common.
A mechanism error means the thing named as the driver of the outcome does not drive it. The written claim was that the gross marginSales less what the materials inside the product cost, carried as a share of sales. Wages, freight, advertising and interest all sit below this line, not in it. would move because of a variable, and that variable turns out to be irrelevant, or to work in the opposite direction, or to be a consequence of the thing it was thought to cause. The repair is structural. A variable leaves the work, or a different one enters it, and the chain of reasoning is rebuilt.
A magnitude error means the chain was right and one number in it was wrong. The variable belongs there, it moves the outcome in the direction claimed, and the mistake was about how far it would move. The repair is completely different. The chain stays, the offending assumption is restated as a number rather than as an adjective, a range is attached to it, and the level at which the conclusion would flip is written down. Nothing is removed.
Applying the mechanism repair to a magnitude error throws away a chain that was working in order to look decisive, and that is the commonest damage a review process does to itself. An analyst whose reasoning was sound and whose sizing was loose walks out of the room having lost the reasoning, and the next piece of work is built on a weaker frame than the one that failed.
An examination finds that the chain of reasoning held, and that exactly one size in it was badly judged. Should that variable be dropped from future work?
Which calls actually get examined, and how many kinds are there?
Almost every review process on earth is triggered by a poor outcome. Something went wrong, somebody notices, and the examination is scheduled. The trigger sounds like common sense, and it is a trap. Work quality and the outcome that work met are two separate things, and they vary independently.
Laid out, they give four cases, not two. Sound work with a poor outcome. Unsound work with a poor outcome. Sound work with a good outcome. Unsound work with a good outcome. A review process triggered by poor outcomes opens the first two of those and never opens the other two. Half the grid goes unexamined by construction rather than by accident.
Of the two that never get opened, one is harmless. Sound work that turned out well is the case where nothing needs saying. The dangerous one is unsound work with a good outcome: the call that was right for a reason that was not true. No complaint is attached to it, so nobody opens it, and it enters the process as a method somebody has apparently validated. Every future piece of work is then allowed to lean on it. The failure is not that the method is wrong; the failure is that the process has certified it.
Of the four cases in that grid, which one is most dangerous to the process itself?
Two analysts wrote opposite calls on the same company in the same week. One got the direction wrong. One got it right. Which of them is worth sitting down and examining?
Two calls on the same company, side by side
Both calls below were written on Sarvani Coatings Limited, an invented paints and coatings maker, and every rupee in them is illustrative.
The first call is the one already sitting in the record. The call was written at the close of year two, when the freshly published statements put gross margin at 44.0 per cent, and it said margin would compress. The reasoning ran that input cost per unitWhat one unit of output costs in materials. It separates a price move from a volume move, because a bigger total bill can simply mean more units of the same thing. was climbing while the company had little room to push prices. Over the year that followed, margin went to 46.0 per cent instead, a one year gain of 2.0 points.
Now read the work rather than the result. Input cost per unit did rise, by about 3.6 per cent, so the mechanism was intact. She got the size of the other term wrong: realisationThe revenue a maker actually collects for each unit sold. It moves when prices move and it also moves when the blend of what is sold changes, so it is not the same thing as a price list. per unit rose about 7.5 per cent, far more than she had allowed for. The miss is a magnitude error, and a narrow one. Verdict: keep the mechanism, and state the realisation assumption as a number next time rather than as the phrase limited room.
The second call was written in the same week by a different analyst. Input prices were falling, she said, so margin would expand. She pointed at the cost of materials dropping as a share of revenue: 57.0 per cent in year one, 56.0 per cent in year two, and still falling. She was right. Margin expanded to 46.0 per cent, exactly as she said it would.
Now do the work anyway, and the second call falls apart. The cost of materials rose from Rs 1,187 crore to Rs 1,304 crore, an increase of about 9.9 per cent, while volume rose 6.0 per cent, so materials cost per unit of output ROSE about 3.6 per cent. Input prices did not fall at any point in the published record. The share of revenue fell because revenue rose about 13.9 per cent, giving realisation per unit up about 7.5 per cent. Compared as factors rather than subtracted as rates, realisation per unit outran input cost per unit by about 3.7 per cent. A share is a ratio of two things, and a ratio moving says nothing about which of its two terms moved.
Two things about that arithmetic are worth stating rather than glossing. Both are one year figures, year two to year three, and they sit against a one year margin gain of 2.0 points; the larger multi-year move in the record belongs to a different period and is not paired with them here. And the second analyst pointed at the year one to year two fall in the share. The fall is a genuine observation. But the record carries no volume growth for year one, so that earlier move cannot be split into a price part and a volume part at all. She had a real observation and no way to know what was behind it. She then wrote down a reason anyway.
So the second analyst was right in direction and wrong in mechanism. Her method will keep working for exactly as long as realisation keeps outrunning input cost, and will fail the first year it does not. She has no way of watching for that year. She does not believe realisation is what carries the margin, so she is not looking at it. The call that gets reviewed is the one with the recoverable error, and the call that walks into the process as a validated method is the one nobody opens.
How many examinations does it take before a finding means anything?
A desk's very first post mortem produces a striking finding, and it feels important. Should the process change now?
One post mortem is a story about one call. The story can be a very good one, and a story is exactly what a process should not be steered by. Nothing yet says whether the failure found is a habit or an accident. What turns a finding into something worth acting on is a base rateHow often something happens across many cases, rather than in the one case immediately at hand.: the same failure showing up across several independent calls, on different companies, with different people, over different periods. Until then the correct action is to write the finding down, date it, and leave it alone.
Waiting is the hardest of these instructions to follow. A fresh finding feels most urgent at exactly the moment it has the least support behind it. The room is still in it, everyone agrees, the evidence is vivid, and the one thing actually known is that it has been seen once. Independence matters as much as the count does. Three post mortems on three calls that all rested on the same forecast, made by the same person in the same quarter, are close to one post mortem repeated, and treating them as three is how a desk convinces itself it has a pattern.
What can an examination never recover?
There is a hard ceiling on every post mortem, and it was set before the examination started. If an assumption was never written down as a number, the examination cannot retrieve it. The examination can only reconstruct the assumption, and a reconstruction is a reading of the work by somebody who now knows the answer, not a record of what the writer believed at the time.
The first analyst wrote limited room to raise prices. The phrase carries a real thought and it is not a number. What did she think realisation would do? Hold flat? Rise two per cent? Rise five? Every one of those is compatible with the phrase, and only one of them makes her call a narrow miss rather than a large one. Reconstructing it after the fact means choosing, and a reviewer who knows realisation came in at about 7.5 per cent will choose differently from one who does not. If instead the assumption setEverything a piece of research quietly takes as given, set out in one place. Put down as numbers it can be tested later; put down as adjectives it never can be. had carried a number, the verdict would have been arithmetic rather than interpretation.
The quality of every post mortem is fixed by how the original work was written. The investment therefore sits upstream of the review rather than inside it. A desk that wants better examinations does not buy them by reviewing harder. The desk buys them by writing assumptions as numbers, and by naming its disconfirming testAn observation named in advance that would show the view is wrong. Naming it before the fact is what stops a view being quietly rescued afterwards. before anything can be observed.
The original work never stated its key assumption as a number. What is the ceiling on what the review can produce?
How does the exercise itself go bad?
Three ways, and they are not equally visible.
It becomes blame. Once the room understands that the examination is about who, people stop writing their assumptions where a reviewer can find them. The reasoning goes into somebody's head instead of into the document, and the process loses the raw material it runs on. Blame is loud and everybody notices it.
It becomes hindsight. Every failed assumption looks obviously wrong once the answer is known, so the review produces the finding that somebody should have seen it. A finding of that shape teaches nobody anything and cannot be acted on. Hindsight is quieter and at least detectable. Reviews that all reach the same conclusion are visibly not doing work.
And it becomes rule accumulation, the quietest of the three and by a wide margin the most expensive. Each error yields a new mandatory check. Each check is individually sensible, cheap and defensible. Nobody ever proposes removing one. Removing a check means arguing that the error it prevents is acceptable. So the list only grows, and it grows in the direction of the past.
A review checklist has reached forty mandatory items. Every one of them is reasonable. What is it costing?
The desk that reviewed itself into paralysis
A desk runs post mortems conscientiously. Every call that went badly gets examined, and every examination produces a new mandatory check. After three years the checklist runs to forty items, each of which would have prevented one specific past error, and an analyst now spends a large part of every piece of work confirming that things which went wrong once have not gone wrong again. On the suppositions above that is 40 units of a 100 unit capacity, or 40.0 per cent of the work, spent looking backwards.
Nothing on the list is unreasonable, and the list as a whole is unaffordable. Capacity spent on checks is capacity not spent on the reading that produces a view in the first place, so the desk is buying protection against repeats with the raw material of its next original thought. Meanwhile the calls that came out right for untrue reasons were never opened at all, so the methods most likely to fail next year are precisely the ones the process has certified. The result is a review system that consumes more every year and detects less.
The repair has three parts and none of them is more review. A finding waits for a second appearance on an unrelated call before it becomes a check. Any check that fires on every piece of work has to justify itself against the reading it displaces, and one that cannot is removed. And calls that went right are examined on a schedule rather than never, so the one quadrant nobody opens stops being the only place the process is blind.
Is a review that finds nothing a failed review?
No, and treating it as one is how the whole exercise quietly turns into theatre. A review that reads the work, checks the assumptions against what happened, finds the reasoning sound and recommends no change has produced a result. A no-change verdict is a finding with a date on it, and it belongs in the file exactly as a discovered error would.
A process that has never recorded a no-change verdict is not reviewing work; it is confirming that reviews produce changes. Once everybody understands that walking out of the room empty-handed reads as a wasted afternoon, somebody will find something, and what they find will be whatever is easiest to phrase as a check. The forty item list gets built exactly that way, out of examinations that should have produced eleven entries and twenty-nine confirmations.
There is a second reason, and it is the one people miss. A base rate needs the cases where nothing happened. If the file only holds the examinations that found something, the denominator was never written down. Nothing then tells the desk whether a failure appears in one call out of thirty or in one out of three. The no-change entries are the denominator.
An examination reads the work carefully and finds no error at all. Is that a failed review?
What does a run of examinations actually leave behind?
Not a stack of resolutions. What survives is a register: for each judgement the desk keeps having to make, the assumption that has to be stated as a number, and the level at which the conclusion would reverse. Two columns, one row per recurring judgement, and nothing else.
Look at what the two calls above would put into it. Both analysts were arguing about the same thing from opposite sides, and both of them left out the same number. One said limited room to raise prices; one said input prices are falling. Neither wrote down what realisation per unit would do, and realisation per unit is the term that decided the outcome. So the register gains one row, and that row would have made either call checkable in advance.
| The judgement that keeps recurring | The assumption that must be a number | Where the conclusion reverses |
|---|---|---|
| Will the gross margin hold? | Realisation per unit growth, stated as a rate for the year | When it falls below input cost per unit growth |
| Is a falling cost share good news? | Both terms separately, not the ratio | When volume growth is unavailable and the split cannot be made |
| Is this a narrow miss or a broken view? | The number the writer expected, recorded before the outcome | When the chain itself, not one input, turns out not to hold |
That register is small, it grows slowly, and its size is a fairer measure of what a desk has actually learned than the length of any checklist. A checklist grows by one item per past error, mechanically, so its length measures how many things have gone wrong. A register grows only when a genuine recurring judgement is identified, so its length measures how many judgements the desk has managed to make explicit. Twelve rows after five years would be a serious body of work. Forty checks after three years is a symptom.
After three years of doing this properly, what should a desk be able to produce?
Who actually does this, and what it looks like on a Wednesday
On a sell-sideResearch written by a house that also arranges trades or raises capital, and published to clients rather than kept inside the firm that wrote it. desk this is usually a supervisory routine rather than a ceremony: a scheduled hour, one call, the original document open on the left and the outturn on the right, and a verdict of a single sheet filed whether or not anything was found. The register is meant to be read before the next piece of work rather than after the next problem, so it lives with the templates and not with the compliance folder.
On the buy side the same routine is run against positions rather than published views, and the classification matters more. A magnitude error and a mechanism error have different implications for whether the position is resized or exited. An investor running personal money can do a cut-down version with nothing more than a notebook: before the act, the one number the reasoning depends on and the level at which the view would change are written down, and at review time that line is read first.
The household version is smaller still and the discipline is identical. A household that switched its recurring investment because one year was poor is grading the queue outside the stall, not the pot. What happened is already known. The useful question after a disappointing year is which of the things believed at the start have turned out not to be true. Usually the answer is none of them, and writing that down is the point.
Put this case in the grid. An analyst wrote that a mix shiftA change in the blend of what a business sells, which moves an average even when no individual price or cost has moved. towards industrial would lift the blended margin, the margin lifted, and it later turns out the mix moved by under one percentage point, far too little to have done it. Which cell?
A note on the arithmetic, for a reader who checks it. Every rate above was worked from the case record's published rupee absolutes held in whole rupees, and never by dividing one already rounded figure by a second. Rs 2,415 crore of revenue set beside the prior Rs 2,120 crore is a move of 13.9 per cent; materials at Rs 1,304 crore beside Rs 1,187 crore move 9.9 per cent; and on volume up 6.0 per cent those two become realisation per unit up about 7.5 per cent and input cost per unit up about 3.6 per cent. Behind the printed shares sit 55.9906 and 53.9959 for the cost of materials, and 44.0094 and 46.0041 for gross margin. The printed 2.0 point gain is 1.9947 unrounded. The record's own rounded figures are the ones printed, so that the sequence stays in step. Carrying the year two share forward by the two per-unit factors lands on the year three share exactly. The agreement is compelled by the algebra and confirms nothing. Volume cancels from both ends of the ratio, so the two routes are a single route written out twice. Every figure above describes a single year, year two to year three, set against a one year margin gain of 2.0 points. The larger multi-year move in the record belongs to a different period and is not paired with any of them. The record carries no year one volume growth, so the year one to year two share fall of 0.9659 of a point cannot be split per unit at all and is quoted only as a share move. The record carries no count of examinations and no cost for a check, so three quantities in the drawing on rule accumulation are suppositions and are labelled where they appear.
Who decides each of these, and where the current text lives
| The question | Who settles it | Where to read it |
|---|---|---|
| Conduct expected of a research analyst, and what a published document must disclose about the person who wrote it | The Securities and Exchange Board of India (SEBI) | sebi.gov.in |
| The primary venue record of a filing, including the timestamp the venue itself puts against it | National Stock Exchange of India | nseindia.com |
| The second venue copy, worth opening whenever a date or a figure reads oddly on the first | BSE Limited, the Bombay Stock Exchange | bseindia.com |
| Judging a decision by how it turned out rather than by the reasoning inside it, which is called resulting | Annie Duke, Thinking in Bets | published 2018 |
Sarvani Coatings Limited is invented.
Educational material. Not advice on any investment, tax, budget or market position.
