Judgment Under Uncertainty: The Step Between Information and Decision
Judgment is the step between information arriving and a decision being taken. Judgment turns what is known into an estimate of what is likely, and that estimate is separate from what is wanted. Most of what looks like a preference problem is a judgment problem, and the two fail in different ways and are corrected by different means.
Judgment sits inside a separation almost nobody has ever had drawn for them. Information, judgment, preference, choice and outcome are five distinct steps, and nearly every argument about a financial decision is really an argument about which of those five went wrong. Until the step is named, the argument has no subject, and naming the step is most of the work. Two people can spend an hour disagreeing about a sale, one of them talking about the evidence and the other about how much a loss would hurt, and never notice that they are discussing different steps.
What are the five steps that run from information to outcome?
The clearest illustration has nothing to do with money. A neighbour tells a commuter the road to the station is flooded. The neighbour's sentence is information. It rained hard last night and this neighbour has been right before, so the commuter decides the report is probably true without being certain of it. Settling on probably true is judgmentThe step that turns information into an estimate of how likely something is.. A missed train costs the whole morning, so being forty minutes early is much better than being five minutes late. Ranking the two mornings that way is preference. The commuter leaves by the long route. Taking the long route is choice. The road turns out to have been clear and the commuter arrives twenty minutes early with nothing to do. Arriving early with nothing to do is outcome.
Every one of those five is a separate thing that can go wrong on its own. Information fails when a person is told something false. Judgment fails when a person is told something true and misreads how likely it is. Preference fails when a person reads the odds correctly and is wrong about what actually matters to them. Choice fails when a person gets all three right and still acts badly. And getting all four right and still having a bad morning is no failure at all.
Which of the five steps turns information into an estimate of what is likely?
What is judgment, and why is it not the same as the decision?
Judgment produces an estimate, not an action. Its output is a number, or something that could be written as a number if pressed: probably, unlikely, about one in five, almost certainly not. Doing something requires knowing how much each outcome would matter to the person deciding, and judgment never asks that question. Nothing in the output says what to do. Judgment answers how likely, and it stops there.
The separation sounds academic until it collapses in an ordinary conversation. A household is deciding whether to keep a fixed deposit or move the money. One person says the scheme is too risky. The sentence has two things packed inside it: an estimate of how often the scheme falls badly, and a statement about how much a bad fall would hurt this household. Pulled apart, the conversation becomes tractable. Left packed together, the sentence keeps the two of them arguing past each other for an hour. One keeps producing evidence and the other keeps producing feelings, and both are correct about their own half.
There is a second reason for the separation, and it is the practical one. A judgment problem has a repair, and the repair is more information, a better method, or a check on how the estimate was formed. A preference problem has nothing to be right or wrong about, and no such repair exists. If a person would genuinely rather hold a smaller sum with no surprises than a larger sum with several, no evidence in the world corrects that. Spending three weeks gathering data to fix something that was never a judgment problem is one of the most common wastes of effort in the whole of investing.
What does Decision-Making Under Uncertainty mean, seen from the other end?
Decision-Making Under Uncertainty is the same step named from the far side. Where the language of judgment starts from the person and asks what estimate they formed, the language of Decision-Making Under Uncertainty starts from the situation and asks what the situation makes possible. Decision-Making Under Uncertainty is the older phrase and the one a textbook contents list uses. The outcome is not settled at the moment of deciding, so something has to stand in for knowledge, and that something is a judgment.
Framed that way, three things follow at once. First, a decision whose outcome is settled in advance is not a decision but an arithmetic exercise. Every financial decision is therefore Decision-Making Under Uncertainty, without exception. Second, the outcome was never fully in the decider's hands, and the quality of a decision cannot be read off it. Third, if judgments are estimates, then the way to grade them is the way any estimate is graded: over a run, not one at a time.
The third point has a name. CalibrationWhether things a person calls 70 per cent likely happen about 70 per cent of the time. asks whether the numbers a person attaches to their judgments match how often those things happen. Take everything somebody called seventy per cent likely over two years, count how many came true, and see whether the answer is close to seven in ten. Calibration is a property of a run of judgments and never of a single one. One decision can never establish that somebody judges well or badly. A single seventy per cent judgment that came out wrong is not evidence of anything; the same person being wrong on eight of ten such calls is.
What does calibration mean?
What is the difference between risk and uncertainty?
Frank Knight drew this line in Risk, Uncertainty and Profit in 1921, and it has held up for a century because it is a distinction about what is available rather than about how nervous anybody feels. RiskA situation where the possible outcomes and their odds can both be stated. is a situation where the possible outcomes can be listed and the odds attached to them. UncertaintyA situation where the odds themselves are not known, as against merely unknown outcomes. is a situation where the odds themselves are not available, and often the list of outcomes is not complete either.
The everyday version is two jars. The first jar is stated to hold ten balls, seven dark and three light, and looking inside is not allowed. Which ball comes out is unknown, and yet three in ten is a fact that can be built on. The first jar is risk. The second jar holds an unknown number of balls in unknown proportions, and nobody will say anything about it. Which ball comes out is still unknown, and now three in ten cannot even be said. There is no fraction to say. The second jar is uncertainty, and no amount of care turns it into the first.
The reason this matters is that almost everything financial is sold as if it were the first jar when it is really the second. A written number has a way of looking like a fact. Somebody writes a one in twenty chance of a bad quarter and the sentence reads exactly like the seven dark and three light, when the one in twenty may have been produced by counting eight quarters of an invented index or by a person forming an impression on a Thursday. Writing a number down does not turn uncertainty into risk. The written number only makes the uncertainty harder to see.
A holding is described as having a 1 in 20 chance of a bad quarter. Is that risk or uncertainty?
How can a judgment be wrong when every fact used was correct?
The Palash decision log, an invented record kept by Palash Advisory Services Private Limited, holds 240 decisions taken by 60 investors over eight quarters, and Meera Sundaram is one of the 60. Of those 240 decisions, 71 were taken within 48 hours of a news item appearing, or 29.6 per cent. Almost three in ten decisions followed hard on the heels of something arriving. The count shows information reaching the decision quickly. The count shows nothing at all about whether the step in between was done well, and the step in between mostly leaves no trace.
So do one of them properly. A signal appears. The signal might be a segment on a television channel, a line in a statement, a broker note, anything. Three numbers are needed and no others. First, how common is the thing the signal points to, before the signal arrived: suppose it fits one holding in five, so the starting rate is 20 per cent. The starting rate has a name, the priorWhat was believed before the new information arrived.. Second, the hit rateHow often a signal appears when the thing it points to is true.: how often the signal shows up when the thing really is true, say 80 per cent. Third, the false alarm rateHow often a signal appears when the thing it points to is not true.: how often the signal shows up anyway when the thing is not true, say 30 per cent.
In whole holdings the arithmetic becomes something visible, and the count is easier that way than in fractions. Suppose there are 100 holdings. Twenty of them are the thing. Eighty of them are not. Of the twenty that are, the signal fires on 80 per cent, or 16 holdings. Of the eighty that are not, the signal fires on 30 per cent, or 24 holdings. So the signal fires on 40 holdings in total, and only 16 of those 40 are the real thing. The correct judgment is 16 out of 40, or exactly 40.0 per cent, and not the 80 per cent the signal came advertising.
| The step | The working | Value |
|---|---|---|
| Holdings that are the thing | 20 per cent of 100 holdings, the starting rate | 20 |
| Holdings that are not | the other 80 of the 100 | 80 |
| True firings | 20 times the 80.0 per cent hit rate | 16 |
| False firings | 80 times the 30.0 per cent false alarm rate | 24 |
| All firings | 16 true plus 24 false | 40 |
| The correct judgment | 16 of the 40 holdings the signal fired on | 40.0 per cent |
Every fact in that table was correct and freely given. Nobody was lied to and nothing was hidden. The 80 per cent was true; it is simply an answer to a different question. The hit rate answers how often the signal fires when the thing is true. The reverse matters more: how often the thing is true when the signal has fired. The two numbers are not the same and not even close. The whole gap between them is the false alarm rate, and no seller quotes it.
The arithmetic has one trap in it, and the trap catches people who are otherwise following. The 80 that are not the thing and the 80 per cent hit rate are two completely different eighties that happen to collide in this example. Changing the starting rate to one holding in four makes the collision disappear while the method stays identical. Writing the working out in whole holdings earns its keep for exactly this reason: 16 true firings out of 40 total firings is a sentence that can be checked on the fingers, and 0.16 over 0.40 is a sentence that can only be checked by an act of trust.
What makes a signal worth anything, the hit rate or the false alarm rate?
The example above used one false alarm rate, 30 per cent, and got one answer, 40.0 per cent. The false alarm rate governs everything, and nobody publishes it. The interesting question is what happens as that number moves. Hold the starting rate at 20 per cent and the hit rate at 80 per cent, and turn the false alarm rate up from nothing.
At a false alarm rate near zero, almost the only holdings the signal fires on are the ones that really are the thing, and the signal is close to decisive. At 5 per cent the correct judgment is 80.0 per cent, the one place where the eighty everybody names is actually the right answer. At 30 per cent the correct judgment is 40.0 per cent, the worked case above. And at 80 per cent the correct judgment is 20.0 per cent, exactly the starting rate that held before the signal ever arrived. When a signal fires just as readily whether or not the thing is true, seeing it cannot move the estimate at all, however impressive its hit rate sounds.
Notice the shape of that. The collapse is fast at the start and slow later, so most of the damage is done while the false alarm rate is still in the range a person would describe as small. Going from 5 per cent to 30 per cent sounds like a modest deterioration, and it takes the judgment from 80.0 per cent to 40.0 per cent and halves it. The false alarm rate therefore cannot be eyeballed and has to be asked for. The everyday version is a car alarm in a crowded street. The alarm is excellent at going off when somebody breaks into a car, and it goes off just as readily when a lorry passes. Nobody in the street looks up.
The false alarm rate rises to 80 per cent while the hit rate stays at 80 per cent. What is the signal now worth?
A signal is right 80 per cent of the time when the thing is true, and the starting rate is 20 per cent. Before the control below is moved: what is the correct judgment?
Turn up the false alarm rate and watch the signal stop saying anything
One variable moves: the false alarm rate, from 0 to 100 per cent. One control can only teach one relationship, and two things are held fixed. The starting rate stays at 20 per cent and the hit rate stays at 80 per cent. The control opens at 30 per cent, reproducing the worked case above exactly and giving a correct judgment of 40.0 per cent.
At a false alarm rate of 30.0 per cent the signal moves a 20.0 per cent starting rate to 40.0 per cent, so it is worth 20.0 points of new belief and nothing like the 80.0 per cent it arrived carrying.
What happens when a judgment gets graded by its outcome?
On 12 October, Meera Sundaram sold Suvarna Chemicals Limited whole at Rs 4,60,000/- against a cost of Rs 4,00,000/-, booking Rs 60,000/- or 15.0 per cent on what she paid. By the time the log was read back on 31 March following, that holding had risen a further 8.0 per cent, so the same Rs 4,60,000/- would have been Rs 4,96,800/-. Rs 36,800/- was forgone.
Looking at Rs 36,800/- and concluding the judgment on 12 October was poor is very tempting. The conclusion does not follow, and the reason fits in one line. The 8.0 per cent had not happened yet on 12 October, so it was not available to any judgment made that day, and a judgment can only be assessed against what was knowable at the time it was made. Graded against Rs 36,800/-, it is not the judgment that is being graded but the world.
The error that gets made, and what it costs
The error is reading backwards from an outcome to a process. The error runs in both directions, and the second direction is the more dangerous one. Nobody investigates a decision that paid.
Run it forwards to see why it fails. On 12 October, what could Meera have known? The cost, Rs 4,00,000/-. The value, Rs 4,60,000/-. The gain of Rs 60,000/-. The statements in front of her. Her own sense of what was likely, if she formed one. The list of inputs available that day ends there. Rs 36,800/- was not on it and could not have been, so it cannot be evidence about the step she took.
Now run it the other way. Suppose the holding had fallen 8.0 per cent instead. The same sale would have looked shrewd, and nobody would have asked a single question about how the decision was reached. If a bad outcome cannot condemn a judgment, a good one cannot acquit it either, and a rule that only works when the money went the wrong way is not a rule at all.
The error costs the ability to learn anything. Judgments graded by outcomes preserve whatever process happened to precede the wins, including the parts of it that were nonsense, and discard whatever preceded the losses, including the parts that were sound. Over eight quarters that is a machine for getting steadily worse while feeling as though things are getting better. And one case, this one included, is not evidence that any rule works.
Rs 36,800/- was forgone by selling on 12 October. Does that show the judgment was poor?
How can a judgment problem be told from a preference problem?
There is a two line test and it works on almost every disagreement about money. The test applies to two people, or to the two halves of one person on two different afternoons. The first question is whether they agree on how likely the thing is. If they do not, that is a judgment problem, and it has a repair: the missing number, a check on where the estimate came from, a count of how the same call went the last ten times. If they do agree on how likely it is and still choose differently, that is a preference problem. They are not disagreeing about the world, and no amount of information will move it.
The test scales down to a kitchen table. Two people in one household look at the same statement. One says the scheme falls by a fifth about one year in seven. The other says the same. So far they agree completely, and the judgment step is settled. Then one of them wants out and the other does not. A fifth off in the year the school fees are due is a very different event in the two heads. Nothing about the likelihood separates them. Producing another chart is a waste of an evening. What would help is a conversation about which year the money is needed, which is a preference question with a factual anchor.
Two people see the same statement. Both agree a fall is unlikely. One sells anyway. Judgment problem or preference problem?
Why is judgment the step most worth writing down?
Of the five steps, four leave evidence behind them on their own. Information arrives with a date on it. Choice produces a transaction. Outcome produces a valuation. Preference is at least visible in what somebody keeps saying. Judgment happens silently, takes a few seconds, and feels at the time like something too obvious to be worth recording. Judgment is the only one of the five that leaves nothing behind unless somebody deliberately writes it down.
Of the 240 logged decisions in the Palash record, how many carried a written reason?
84 of the 240 carried a written reason, or 35.0 per cent. The other 156 record what was done and not what was thought. Set that beside the 71 of 240 taken within 48 hours of a news item, or 29.6 per cent. The two figures describe the same gap from opposite sides: information reaches the decision fast, and the step in between very often leaves no trace at all. An entry with the judgment left blank holds nothing to review. The record shows that the holding was sold. The record does not show what the seller thought was likely, so whether the seller was right can never be found out, and calibration over a run of judgments becomes impossible to compute.
The blank line is therefore worth looking at rather than the filled ones. A blank reasoning field is not a lapse and not carelessness. A blank reasoning field is the ordinary case. The judgment felt so obvious while it was being made that recording it seemed like paperwork. Six months later the obvious thing has gone, and what is left is a transaction with no reasoning attached to it. Nothing whatever can be learned from that artefact.
How does an adviser actually use any of this on a Tuesday afternoon?
What a practitioner does with the judgment step
Devika Rao, the adviser at Palash Advisory Services Private Limited, does not open a meeting by naming a psychological effect, and nor need anybody else. She asks three ordinary questions and writes the answers in the record, and the three questions are the five steps compressed into something that can be said out loud.
The first is what changed, and it pins the information step to something with a date rather than to a general mood. The second is how likely the client thinks it is now, and how likely the client thought it was last month. The second question is the judgment step and the only one that can be checked afterwards. The third is what it would mean for the client if it happened, and the third question is the preference step, with no right answer at all. Three lines in a record turn an unreviewable transaction into something that can be marked six months later.
A lender does the same thing under a different name. When a credit file is reviewed, the estimate of how likely the borrower is to fall behind is recorded separately from the decision about whether to lend, precisely so that the two can be criticised separately when the loan sours. An analyst does it by writing the estimate before the result rather than after. And a person deciding alone, with no adviser and no committee anywhere in sight, does it with a notebook and one sentence per decision. The sentence costs about forty seconds and is the only way to ever find out whether their seventy per cents come in at seven in ten.
Two cautions. Recording the judgment step is a described practice, not a practice shown to pay. Across quarters five to eight, 20 of the 60 investors in the invented Palash record kept a written checklist and recorded a reason on 34 of 41 decisions, being 82.9 per cent, against 19 of 63, being 30.2 per cent, for the other 40. The two shares measure what got written down and nothing else. Eight quarters and 60 people are far too few to separate a real difference in returns from ordinary variation.
What does a good judgment look like when the outcome is bad?
A good judgment looks exactly like a bad one from the outside, and that is the uncomfortable part. If the outcome cannot settle it, something else has to, and there are only three things available. The first is the inputs: were the numbers that were used actually the right numbers, and was any of them the false alarm rate that nobody quotes. The second is the method: was the estimate reached by counting something, or by an impression formed in the ninety seconds after a screen changed. The third is the run: over ten or twenty judgments of the same kind, did the sevens in ten come in at about seven in ten.
All three of those are available on the day the judgment is made, and none of them requires waiting for the outcome. Separating the step out has exactly that practical value. A judgment can be assessed before its outcome exists, and that is the only kind of assessment that can improve the next one. Waiting for the outcome gives a verdict that arrives too late to change anything and is contaminated by everything that happened in between.
Two more things follow, and both are worth carrying forward. A good judgment badly rewarded is still a good judgment and should be repeated. And a poor judgment handsomely rewarded is still a poor judgment and should not be, however pleasant the statement looks. Neither sentence is comfortable, and no single case, the one worked through above included, can establish either. The step exists, it can be done well or badly on its own terms, and the specific errors set out under heuristics and biases are all statements about how this one step goes wrong.
Sources
| Source | Document | Site |
|---|---|---|
| Frank Knight | Risk, Uncertainty and Profit, 1921, where the distinction between stateable odds and unstateable odds is drawn | cited to the book itself |
| Social Science Research Network (SSRN) | the repository where working papers on judgment and decision-making are findable | ssrn.com |
| National Bureau of Economic Research | the working paper series where research in this area is findable | nber.org |
Meera Sundaram, Devika Rao, Palash Advisory Services Private Limited, the Palash decision log, the Palash 100 index and Suvarna Chemicals Limited are invented.
Educational material. Not advice on any investment, tax, budget or market position.
