Forecasting Honestly: Consensus, Range and Probability
A consensus forecast is the set of what a collection of forecasters submitted, averaged into one number. The average is the least informative part. The range across the submissions is what shows whether the thing is genuinely uncertain. A probability-weighted forecast multiplies each case by its stated probability and sums, and the figure it produces usually matches none of the cases it came from.
Two ideas built earlier sit underneath the subject. The first is a collected central figure used as the benchmark a data surprise is measured against. A released number gets called higher or lower against that benchmark. The second is the practice of writing a base case, a bull case and a bear case and attaching a stated probability to each. Neither idea on its own asks what the collection and the weighting actually produce, or what a reader is entitled to conclude from either. The whole subject turns out to rest on a single distinction, between a number that somebody holds and a number that arithmetic produced.
What exactly is a consensus forecast, and who holds it?
A consensus forecast is a collection, not a view. Twelve desks each submit a figure, somebody adds the twelve and divides by twelve, and the result is published as the consensus. Averaging twelve forecasts produces a number no forecaster submitted, so nobody holds a consensus. That is not an occasional accident of the arithmetic. An average of twelve distinct figures coincides with one of them only when the numbers happen to line up around it exactly. The mismatch is close to the normal case.
Think of twelve neighbours guessing the weight of a sack of rice at a village fair. The guesses come in at nine kilograms, eleven, fourteen and so on, and the organiser announces that the average guess was 11.7 kilograms. No neighbour said 11.7 kilograms. If a thirteenth person arrived and asked who thinks it weighs 11.7 kilograms, the honest answer is nobody. The average is a thing the organiser made, using everyone's guesses as raw material, and it belongs to the organiser rather than to any of the twelve.
Take Nirvi Engineering, an invented manufacturer inside the invented Republic of Sankhya. Twelve research desks each submit a figure for what Nirvi Engineering's operating profitRevenue less operating costs, before interest and tax. Operating profit is what the business earned from running itself, and how to build it is covered under company accounts. would come to over the coming year. The twelve submissions, in Rs crore, are 84, 88, 90, 96, 98, 102, 108, 110, 114, 116, 118 and 124. The twelve submissions add to Rs 1,248 crore, and Rs 1,248 crore divided by twelve is Rs 104 crore. Read back down the list of twelve and look for Rs 104 crore. Rs 104 crore is not on the list. The closest submissions on either side are Rs 102 crore and Rs 108 crore, so the published consensus sits in a gap where nobody put anything.
The mismatch matters for a plain reason. The moment a reader treats Rs 104 crore as somebody's opinion, that reader has misattributed a view to twelve people who did not hold it and to no one who did. A sentence such as forecasters think Nirvi Engineering will earn Rs 104 crore is wrong twice over: the desks did not think that, and the number is not a thought at all. The consensus figure is the output of a division. Naming it correctly costs one word. The consensus average of twelve submissions is Rs 104 crore, and the twelve submissions themselves run from Rs 84 crore to Rs 124 crore.
Twelve desks submit figures and the published consensus for Nirvi Engineering is Rs 104 crore. Who holds that view?
Why is the range more informative than the central figure?
The average locates the middle of a collection. The range shows how much the collection disagrees, and disagreement is information about the world rather than about the desks. A central figure without its range has discarded the useful half of the collection. Two collections can publish the identical central figure while saying entirely different things, and only the range separates them.
Here is the test that makes it concrete. Panel one is the twelve submissions already listed: Rs 84 crore to Rs 124 crore, averaging Rs 104 crore. Panel two is a different collection of twelve desks, labelled the same way for comparison, whose submissions are 94, 96, 97, 100, 101, 103, 106, 107, 109, 110, 111 and 114, all in Rs crore. The panel two figures add to Rs 1,248 crore as well, so panel two also averages Rs 104 crore. The two panels publish the same headline. Panel one spans Rs 40 crore from lowest to highest. Panel two spans Rs 20 crore, exactly half as wide.
Each panel is saying something different. Panel two says twelve desks looked at Nirvi Engineering and landed within Rs 20 crore of one another. The inputs are reasonably agreed and the arithmetic between them is reasonably shared. Panel one says twelve desks looked at the same business and could not get within Rs 40 crore of one another. A gap that wide usually means at least one input is genuinely contested. Perhaps some desks assume the rupee cost of an imported input holds and others do not. The wide panel is not a worse panel. The width says something true and uncomfortable. The answer is not settled.
A range can be driven by a single stray submission, so one more measure is worth reading alongside it. DispersionAny measure of how spread out a set of numbers is, rather than where its middle lies. Standard deviation is one such measure; the average distance from the mean is another. here is measured as the average distance of a submission from the average, and it uses all twelve rather than only the two extremes. For panel one, the twelve distances are 20, 16, 14, 8, 6, 2, 4, 6, 10, 12, 14 and 20. The twelve distances add to Rs 132 crore, so the average distance is Rs 11 crore. For panel two the same calculation gives Rs 5.50 crore, again exactly half. Range and dispersion agree here that panel one is twice as spread as panel two.
A note on which central figure gets published. Some collections publish the average and some publish the medianThe middle value once a set of numbers is sorted, so half the set sits below it and half above. With twelve numbers the median is the average of the sixth and seventh. instead, which is less pulled about by one extreme submission. Either can be the published central figure, and neither carries the range on its own. Whichever one is published, the reader's job is the same: find the spread and read it first.
Two collections of twelve submissions both average Rs 104 crore. Panel one spans Rs 40 crore, panel two spans Rs 20 crore. What has panel one shown that panel two has not?
A reader is handed a single line: consensus Rs 104 crore. Which part of the collection has been thrown away?
| Submission, listed lowest to highest | Panel one, Rs crore | Panel two, Rs crore |
|---|---|---|
| Desk 01 | 84 | 94 |
| Desk 02 | 88 | 96 |
| Desk 03 | 90 | 97 |
| Desk 04 | 96 | 100 |
| Desk 05 | 98 | 101 |
| Desk 06 | 102 | 103 |
| Desk 07 | 108 | 106 |
| Desk 08 | 110 | 107 |
| Desk 09 | 114 | 109 |
| Desk 10 | 116 | 110 |
| Desk 11 | 118 | 111 |
| Desk 12 | 124 | 114 |
| Total of the twelve | 1,248 | 1,248 |
| Average of the twelve | 104.00 | 104.00 |
| Range, highest less lowest | 40.00 | 20.00 |
| Average distance from the average | 11.00 | 5.50 |
Every figure in both collections is printed above, so the arithmetic can be redone rather than believed. Adding the twelve figures in either column gives Rs 1,248 crore; dividing by twelve gives Rs 104 crore; subtracting the lowest from the highest gives the range. Panel two is exactly panel one with every submission moved half the distance towards Rs 104 crore. Every measure of spread therefore comes out at exactly half.
How is a Probability-Weighted Forecast built?
A probability-weighted forecast is a different construction from a consensus, and mixing the two is a common muddle. A consensus averages what several people said. A probability-weighted forecast takes what one person said across several stated cases and combines them using stated probabilities. The arithmetic is simple: multiply each case by its probability, add the products, and the sum is the probability-weighted forecast. The care is entirely in what the result is then called.
The three cases come from the scenario work built earlier for Nirvi Engineering. The base case puts operating profit at Rs 100 crore and carries a stated probability of 50.00 per cent. The bull case puts it at Rs 130 crore at 25.00 per cent. The bear caseThe unfavourable case in a set of stated scenarios, written down deliberately alongside a base and a favourable case. How the three are constructed is covered under scenario building. puts it at Rs 60 crore, also at 25.00 per cent. The three probabilities add to 100.00 per cent. Checking that total is the first thing to do on any weighting, and a set that does not add to 100.00 per cent is not a weighting at all.
Work it through. The base case contributes Rs 100 crore multiplied by 50.00 per cent, or Rs 50.00 crore. The bull case contributes Rs 130 crore multiplied by 25.00 per cent, or Rs 32.50 crore. The bear case contributes Rs 60 crore multiplied by 25.00 per cent, or Rs 15.00 crore. Add the three contributions: Rs 50.00 crore plus Rs 32.50 crore plus Rs 15.00 crore comes to Rs 97.50 crore. The probability-weighted figure of Rs 97.50 crore sits below the base case of Rs 100 crore and equals none of the three cases.
| Case | Operating profit, Rs crore | Stated probability | Contribution, Rs crore |
|---|---|---|---|
| Base case | 100.00 | 50.00 per cent | 50.00 |
| Bull case | 130.00 | 25.00 per cent | 32.50 |
| Bear case | 60.00 | 25.00 per cent | 15.00 |
| Probability-weighted figure, a summary of assumptions | 100.00 per cent | 97.50 |
Notice what the last row is called. The last row is not the forecast, and not what anybody thinks will happen. The last row is a summary of assumptions: three stated values and three stated probabilities, compressed into one figure. Change any of the six inputs and the summary changes. A summary is supposed to do exactly that. Notice also what the weighted figure is not. Rs 100 crore plus Rs 130 crore plus Rs 60 crore divided by three is Rs 96.67 crore, the plain average of the three cases. The plain average implies the three cases are equally probable, and the writer said they were not.
Rs 100 crore at 50.00 per cent, Rs 130 crore at 25.00 per cent and Rs 60 crore at 25.00 per cent. What does the weighting come to?
Which of the three cases for Nirvi Engineering is equal to Rs 97.50 crore?
Why does the weighted figure land on none of the cases?
The mismatch looks like a paradox on first reading and it is not one. A weighted average of distinct outcomes lands somewhere between them, and the only way it can land exactly on one of them is if the others happen to cancel around it. Unless the cases sit symmetricBalanced about a centre, so that what sits on one side is mirrored on the other. A set of numbers can be symmetric about its middle value without being equally probable. around the base case with matching probabilities, the weighting lands off the base case. Here they do not, and the arithmetic of why is worth doing slowly.
Measure the two side cases as gaps from the base case rather than as levels. The bull case at Rs 130 crore sits Rs 30 crore above the base case of Rs 100 crore. The bear case at Rs 60 crore sits Rs 40 crore below it. Both carry the same stated probability of 25.00 per cent. So the bull case pulls the weighted figure up by 25.00 per cent of Rs 30 crore, or Rs 7.50 crore, and the bear case pulls it down by 25.00 per cent of Rs 40 crore, or Rs 10.00 crore. The two pulls do not cancel. The net pull is Rs 2.50 crore downwards. Rs 100 crore less Rs 2.50 crore is Rs 97.50 crore, the figure the full weighting produced.
Change one input and watch the whole thing move. If the bull case had been written at Rs 140 crore instead of Rs 130 crore, it would sit Rs 40 crore above the base case, matching the bear case's Rs 40 crore below. At 25.00 per cent each, the upward pull of Rs 10.00 crore and the downward pull of Rs 10.00 crore cancel exactly, and the weighted figure lands on Rs 100.00 crore, precisely on the base case. So the weighted figure landing off the base is not a quirk of weighting. The gap is a direct readout of how unbalanced the writer's own cases were.
The figure is a summary of assumptions and not an outcome anyone described. That distinction has a practical edge. A reader who takes Rs 97.50 crore and builds a plan around it has built a plan for a state of the world that appears nowhere in the scenario work. The writer described three futures. In one of them Nirvi Engineering earns Rs 100 crore, in one Rs 130 crore and in one Rs 60 crore. There is no described future in which it earns Rs 97.50 crore, and a plan that only works at Rs 97.50 crore is a plan tested against nothing anybody wrote down.
The bull case sits Rs 30 crore above the base case and the bear case sits Rs 40 crore below it, both at 25.00 per cent. Why does the weighting land below the base case?
What makes a forecast honest rather than merely careful?
Honesty here is not a temperament and not a tone. Honesty is a property that can be tested after the fact, and testable honesty is the only kind a reader can act on. A forecast is honest when it states its inputs, states its range and states in advance what would falsify it. Somebody who was not in the room can check all three afterwards. A forecast that states none of the three may still be careful, thoughtful and produced by serious people, and there is no way for a reader to find out.
Take the three properties one at a time. Stating the inputs means writing down which assumptions were used and at what values, so a reader who disagrees with one can change it and redo the arithmetic. For Nirvi Engineering, the inputs behind a profit figure would include the assumed revenue, the assumed cost base, and the assumed rupee cost of the imported inputs. A reader who thinks one of them is wrong can substitute a different value and see the answer move. Substituting an input and redoing the arithmetic is a conversation. A figure with no stated inputs allows no such conversation, only agreement or disagreement in the abstract.
Stating the range means publishing how wide the possible set was, not only where its middle sat. Publishing the range is the same discipline as reading a consensus range, applied to a single forecaster instead of a collection. Rs 104 crore, with a plausible span of Rs 84 crore to Rs 124 crore, is a statement. Rs 104 crore on its own is a number with its uncertainty deleted, and deleting the uncertainty makes the figure look more solid than the person who wrote it believed it to be.
Stating what would falsifyTo show a claim is wrong by pointing at a specific observation that contradicts it. A claim no observation could contradict cannot be tested at all. the forecast is the hardest of the three and the most revealing. Stating a falsifier means naming, before the fact, the observation that would show the forecast was built on a wrong reading. Something of the form: if the imported input cost rises by more than the assumed amount, this figure is wrong and here is roughly how wrong. A forecast that cannot be contradicted by any observation is not cautious. The forecast is untestable instead, and an untestable figure gives a reader nothing to check when the outturnThe value a quantity actually turned out to be, once it is measured and published, as against what anyone said beforehand. arrives.
Where a reader in India can go for collected forecasts
The Reserve Bank of India is the central bank, and among its published material is a survey that collects submissions from professional forecasters. The survey is a real example of a collected set of submissions. The Ministry of Finance publishes the Economic Survey, and the National Statistical Office compiles the national accounts against which any forecast is eventually measured.
Every number, span and probability worked through above belongs to the Republic of Sankhya and to Nirvi Engineering inside it, so a reading of India's own figures has to be taken from each institution's own site.
Name one of the three properties that make a forecast testable rather than merely careful.
What does no forecast deliver, however carefully it is built?
No forecast delivers certainty. Everybody already knows that, and nobody is misled by it. The more interesting gap is the second one. No forecast delivers its own error rate, so a forecaster without a record of past errors is untested rather than accurate, and from the outside the two look exactly alike. The record is the only thing that tells them apart, and the record has to be kept before it is needed.
Two invented desks stand side by side. Desk 03 and Desk 07 both publish the same thing today: a central figure of Rs 104.00 crore for Nirvi Engineering with a span of Rs 84.00 crore to Rs 124.00 crore. As published, the two are indistinguishable. The difference is that Desk 03 has kept a log of what it submitted in four earlier years and what the outturn turned out to be, and Desk 07 has kept nothing.
| Period | Desk 03 submitted, Rs crore | Outturn, Rs crore | Error, submitted less outturn |
|---|---|---|---|
| Year one | 96 | 101 | minus 5 |
| Year two | 108 | 103 | plus 5 |
| Year three | 112 | 99 | plus 13 |
| Year four | 94 | 97 | minus 3 |
| Average error, signs kept | plus 2.50 | ||
| Average size of error, signs ignored | 6.50 |
Either row alone misleads, so read the last two together. The four errors are minus 5, plus 5, plus 13 and minus 3, adding to plus 10 for an average of plus 2.50. The signed average looks reassuring: Desk 03 has run only slightly high. But the sizes of the four errors are 5, 5, 13 and 3, adding to 26 for an average of 6.50. The offsetting signs hid more than half the typical miss. A desk that reports only the signed average error is reporting the smaller of two true numbers, and a reader who wants to know how far off this desk usually lands needs the unsigned one.
Desk 07 has neither figure. Having no record is not the same as being wrong, and the difference is worth sitting with. Desk 07 may be the better forecaster of the two. Desk 07 cannot be a forecaster whose accuracy a reader has any way of checking. Treating an unrecorded desk as reliable because its publications look tidy is a judgement made on typography. The household version is the same shape: a neighbour who has been right about the price of onions four times and can say when they were wrong has given the household something a confident stranger has not.
What should a lender, an analyst or a household take from a consensus?
The reading order is the opposite of the one most people use, and reversing it is most of the practical value here. The spread comes first, the dispersion second and the central figure last. Every other reader already has the central figure, and only the spread shows how much weight it deserves. A number everyone shares carries no advantage; the shape of the disagreement around it sometimes does.
A lender assessing Nirvi Engineering for a working capital facility does not need to know that twelve desks average Rs 104 crore. The facility has to survive the low end, so the lender needs to know that the twelve run from Rs 84 crore to Rs 124 crore. A covenant sized against Rs 104 crore is sized against a figure nobody submitted; a covenant sized against Rs 84 crore is sized against the worst thing anybody in the collection actually wrote down. The lender's question is not where the middle is, it is how bad the bottom of the published set gets.
An analyst uses the spread differently, as a map of where the disagreement is concentrated. If twelve desks agree on revenue and disagree only on the imported input cost, the entire Rs 40 crore span is one assumption wearing twelve values, and the analyst's work is to form a view on that single assumption rather than on the company as a whole. Forming a view on one assumption is a much smaller job than it first looks, and finding that assumption requires the submissions rather than their average.
An investor and a household face the same structure with different stakes. A household deciding whether one earner's variable income can support a fixed monthly commitment is doing exactly this arithmetic without the vocabulary: the average month is not the month that breaks the plan, the worst plausible month is. The span across forecasters says how settled the question is, and the same reading applies to any collected forecast of a macro variable, inflation or the policy rateThe rate at which a central bank lends to banks against collateral, which it sets as an instrument of policy. The policy rate and its movements are covered under monetary policy. alike.
The panel that reports a range beside every central figure
The slider compresses or stretches the twelve submissions around a fixed average of Rs 104.00 crore, leaving the headline exactly where it is while the information underneath it changes completely. The selector changes only the three stated probabilities. The three case values are held constant at Rs 100 crore, Rs 130 crore and Rs 60 crore throughout.
The reader who quotes the middle and plans against it
The mistake is quick and it reads as competence. A reader takes the published line, consensus Rs 104 crore, drops the twelve submissions it came from, and writes that forecasters expect Nirvi Engineering to earn Rs 104 crore. Panel one and panel two both produce that same sentence, and the sentence is equally uninformative in both cases even though one collection agreed within Rs 20 crore and the other could not agree within Rs 40 crore.
The second half of the mistake costs more. The same reader takes the probability-weighted Rs 97.50 crore and builds a plan sized to it. Rs 97.50 crore happens in none of the three cases. The plan is therefore sized to a state of the world that nobody described, that no case tested, and that the writer of the scenarios never claimed was possible. The plan will be too loose for the bear case at Rs 60 crore and too tight for the bull case at Rs 130 crore, the worst of both.
The fix is one habit, applied every time. Quote the range with the central figure, always. Two identical averages can sit on close agreement or on complete disagreement, and nothing else distinguishes them. Then, when a probability-weighted figure appears, call it a summary of assumptions out loud, and plan against the cases themselves rather than against the summary of them.
A reader sizes a plan for Nirvi Engineering against the probability-weighted Rs 97.50 crore. What has the plan been sized against?
Where a reader can go and check this
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Survey of Professional Forecasters, a collected set of submissions from professional forecasters | rbi.org.in |
| Ministry of Finance | Economic Survey, the annual review of the Indian economy | finmin.nic.in |
| National Statistical Office | Methodology notes accompanying the national accounts | mospi.gov.in |
| International Monetary Fund | World Economic Outlook, where ranges are published beside central figures | imf.org |
| National Bureau of Economic Research | Working paper series on the evaluation of forecast errors | nber.org |
| Bank for International Settlements | Annual Economic Report, a yearly survey of global macro conditions | bis.org |
The Republic of Sankhya, its currency the Marut and Nirvi Engineering are invented.
Educational material. Not advice on any investment, tax, budget or market position.
