Likelihood: Estimating Probability Without False Precision
Likelihood is a statement about a chance within a stated period, and the period is half of it. The future has not happened, so the chance is always an estimate, and the honest question is what the estimate rests on: a count of the past, a model, a graded scale built from an institution's own history, or somebody's judgement. Each of those is defensible and none of them is a measurement.
Almost everything in risk work that carries a number carries an estimated chance somewhere underneath it, and almost nobody looks at what that chance rested on. A register rating, a stress scenario, an expected loss, a capital figure: every one of them is built on somebody having said how likely something is. The honesty of the whole structure is bounded by the honesty of that one step, and that one step is usually the least examined thing in the building. Estimating that chance goes wrong in one particular way, and the way is not that people get the arithmetic wrong. People almost never do. The fault is that the answer gets written down more confidently than the evidence allows.
What is a likelihood actually a statement about?
A likelihoodThe chance of a stated event within a stated period, which is always an estimate. has three parts, and until it has all three it is not a statement at all. There is an event, defined precisely enough that somebody could recognise it if it happened. There is a chance. And there is a period over which that chance applies. Take away the event and there is nothing to count. Take away the chance and there is no estimate. Take away the period and the number becomes meaningless. A 5 per cent chance a day and a 5 per cent chance a decade are not near each other in any sense at all.
The institutional version is the same shape wearing better clothes, so the everyday version is worth holding on to. A neighbour says the scooter is likely to break down. Likely is not a statement. Break down how, and by when? Refuse to start on a cold morning, once this winter, is a statement: an event anybody would recognise, a chance that can be argued about, a period. Somebody can now disagree usefully. Nobody can usefully disagree with likely.
Moving the period changes the number without changing a single thing about the world, so the period does more work than people expect. Take one record. Vindhya Commercial Bank Limited, an invented bank used as the working example throughout, recorded two operational incidents in one event category over its twelve month reporting period. The same single record can be written three ways.
| The period stated | The same two events written as a rate | Is it a different claim? |
|---|---|---|
| One month | 0.167 events | No |
| One quarter | 0.500 events | No |
| One year | 2.000 events | No |
Three numbers, one record, and none of the three is more true than the other two. A likelihood without a period attached is not a weak estimate, it is not an estimate at all, and the reader who sees 0.167 and the reader who sees 2.000 have been told exactly the same thing. The period is therefore not a detail at the end of the sentence. The period is half the sentence.
What three things does a likelihood statement need before it is a statement at all?
Where does an estimate come from, if not from a formula?
There is no formula that produces a likelihood out of nothing. Every such figure comes from one of exactly four places, and knowing which one is most of what makes the figure defensible. A count of what has already happened. A model that predicts what will. A graded scale built from an institution's own history. Or a person's judgement placed on an ordered scale. Each of the four can be argued with, and each in a completely different way. A figure whose source is not stated cannot be argued with at all.
Notice that the four are not ranked. A count is not automatically better than a judgement. A count of two events is worse than the judgement of somebody who has watched the process for fifteen years, and both of them are honest as long as they say which they are. The mistake is not using a weak source. The mistake is using a weak source and presenting the output in the format of a strong one.
The fourth source is the one people apologise for, so it deserves a word. A person's judgement placed on an ordered scale is not a failure of rigour; it is often the only source available, and a large part of what a risk function does is collect it in a form that can be compared across people. A judgement is made defensible by the same thing that makes the other three defensible: saying out loud that a judgement is what was done.
A likelihood of 12 per cent arrives from somebody. What is the first question?
When does a count of the past stop being a count?
Start with the cleanest of the four sources. The moment it stops being clean is the moment worth watching. A frequencyHow often something has actually happened, which is a count of the past and not a statement about the future. is a count. Thirteen incidents happened. A count is a fact about the past, it is checkable, and nobody can argue with it except by going back to the record. Divided by the period, it becomes a rateA frequency divided by a period, which becomes a likelihood estimate only when somebody claims the past will continue., which is still a fact about the past: thirteen incidents in twelve months.
The count becomes an estimate at the exact moment somebody uses it to say something about a year that has not happened yet, and that step is silent, unwritten and almost never noticed. Nothing written down changes. The number does not change. The claim changes: from here is what was recorded to here is what is expected. The second claim carries an assumption the first one does not, that whatever produced the past will keep producing it.
Here is the invented bank's own record, sorted into the seven operational risk event categories it uses. The seven categories are the Basel Committee's, published by the Bank for International Settlements at bis.org. Their meaning does not matter here. The number of events sitting in each one does.
| Event category | Incidents in the year | Which ones | One further event moves the rate by |
|---|---|---|---|
| 1 internal fraud | 2 | I5, I13 | 50.0 per cent |
| 2 external fraud | 2 | I1, I7 | 50.0 per cent |
| 3 employment practices and workplace safety | 1 | I11 | 100.0 per cent |
| 4 clients, products and business practices | 2 | I4, I12 | 50.0 per cent |
| 5 damage to physical assets | 1 | I8 | 100.0 per cent |
| 6 business disruption and system failures | 2 | I3, I9 | 50.0 per cent |
| 7 execution, delivery and process management | 3 | I2, I6, I10 | 33.3 per cent |
| All thirteen incidents together | 13 | I1 to I13 | 7.7 per cent |
Read the count column before the rate column. The count column is the part everybody skips. Two of the seven event categories rest on a single recorded event. Six of the seven rest on two events or fewer, being 85.7 per cent of the categories. Only one category has as many as three. The record is not badly kept; it is an ordinary one, and any institution that has had a quiet year will have a record shaped exactly like it. The thirteen incidents sum correctly, 2 + 2 + 1 + 2 + 1 + 2 + 3 = 13, and the whole set is the only figure on the table with enough behind it to be worth two decimal places.
What does one more event do to a rate built on two?
There is one piece of arithmetic that ought to sit at the top of every report that carries an estimated chance, and it takes ten seconds. A rate built on a count of events moves by one over that count if a single further event turns up. The arithmetic is that short. At one observation, one more event doubles the estimate. At two it moves it by half. At three by a third. At thirteen by under a thirteenth.
Say it as a household would. One meal burnt this year, then another, and the burning rate has doubled. Fifty burnt already, then another, and nothing has happened. Nobody who has ever cooked finds that surprising, and yet the same person will read 8.33 per cent a month off a spreadsheet built on one event and treat it as a measurement.
The shape is the lesson. The curve is not a gentle decline; it is a cliff followed by a plain. Almost all of the instability lives between one observation and five, and after about twenty the curve is so flat that a further event barely registers. The crossings are worth memorising for that reason: the sensitivity falls below 10.0 per cent at ten observations and below 5.0 per cent at twenty, and both figures are exact rather than approximate.
A category has exactly one recorded event this year. What is the honest thing to say about its likelihood?
The question underneath the control below can be answered from the curve just read, before the control is moved. The arithmetic is easy and the intuition is not, and the gap between the two is worth finding out about.
A category has two recorded events in twelve months. How much would one further event move the estimated rate?
Move the number of events and watch the estimate stop wobbling
One control: how many events have been observed over the same twelve month window, from 1 to 30. Two consequences, and both of them redraw: the estimated rate in events a month, and how much one further event would move that estimate. The default is 2 events, where four of the invented bank's seven event categories sit: a rate of 0.167 events a month and a sensitivity of 50.0 per cent.
On 2 events in twelve months, the estimated rate is 0.167 a month, and one further event would move it by 50.0 per cent.
Where exactly is the fault in 8.33 per cent a month?
Event category 3, employment practices and workplace safety, is the thinnest row on the table above, with exactly one event in the year: incident I11, an employment tribunal settlement of Rs 1.2 crore in month 11. One event, twelve months. Its likelihood can be written four different ways without a single arithmetical error.
| How the same single observation is written | Is the division correct? | What it quietly claims |
|---|---|---|
| One event in twelve months | Yes | Nothing beyond what happened |
| One event a year | Yes | That next year looks like this one |
| 0.0833 events a month | Yes | That the year can be cut into twelve equal slices |
| 8.33 per cent a month | Yes | That the evidence can tell 8.33 from 8.31 |
Every one of those is arithmetically correct, and the arithmetic is not where the fault is. False precisionStating an estimate to more decimal places than the evidence can distinguish between. is a property of the presentation and not of the calculation. Arguing about the division therefore never gets anywhere. Nobody divided wrongly. The problem is that a single observation cannot distinguish a rate of one a year from a rate of one every three years, and the second decimal place says it can.
The claim can be sized rather than asserted, and sizing it is worth doing once. Assume, as a working assumption here rather than the invented bank's, that events of this kind arrive independently at a constant rate. Each of three quite different underlying rates then has some chance of producing exactly one event in a twelve month window.
The picture holds the whole of false precision. The estimate is not wrong. The estimate is stated in a format that carries a claim the evidence never made. And the fix is not a better calculation; it is a shorter sentence. About one a year, on one observation, is a defensible thing to write. 8.33 per cent a month is not, and no amount of recomputing will make it so.
Where does the fault lie in a figure of 8.33 per cent built on one observation?
What does a grading scale look like read as multiples?
The invented bank has a second way of saying how likely something is, and it looks far more solid than a count of thirteen. The bank grades borrowers 1 to 10 and attaches its own one year probability of defaultAn institution's own estimate of the chance a borrower fails within a year, here built from its own history. to each grade: 0.03 per cent at grade 1, then 0.08, 0.20, 0.45, 0.90, 1.80, 3.60, 7.20, and 15.00 per cent at grade 9. Grade 10 is default itself, a state and not a chance, and the Rs 1,764 crore sitting in it is the same Rs 1,764 crore as the gross non-performing book rather than a second number beside it. Every one of those figures is this invented bank's own estimate from its own default history, and none of them is a rating agency or market figure.
Set down as ten rows of equal height, that scale looks like a staircase with equal steps. The scale is nothing of the kind. The scale is a ladder of multiples, each grade roughly twice the one above it, and 15.00 divided by 0.03 is exactly 500, so end to end it spans 500 times. Read the step multiples out loud and something else appears: 2.667, then 2.500, then 2.250, then 2.000 four times over, then 2.083. Grades 4 to 8 double exactly, four consecutive times, and the two ends of the scale are wider than the middle.
Why does this matter for a reader rather than for a modeller? Because the even rows train the eye to think in differences and the scale is built in ratios. Moving a borrower from grade 6 to grade 7 does not add 1.80 percentage points in any meaningful sense; it doubles the estimated chance. On the printed table both moves look like one row. A one row move at the top of the scale and a one row move at the bottom are not the same event, and nothing in the layout of the table says so.
The scale runs from 0.03 per cent to 15.00 per cent across nine grades. What does printing them as nine even rows conceal?
The same two decimals at the two ends of one scale
Now look at the format rather than the figures. The scale is printed to two decimal places at both ends, in the same column, in the same typeface: 0.03 at the top and 15.00 at the bottom. The consistency looks like discipline. The consistency is the source of the trouble.
Work the two numbers yourself. Both are short. The gap between grade 1 and grade 2 is 0.05 percentage points, and 0.05 divided by 0.01 is 5, so the second decimal place cuts that gap into five pieces. Applied to grade 1's Rs 2,940 crore of advances at the invented bank's own 40.0 per cent loss given default assumption, the whole gap is worth Rs 0.588 crore. The gap between grade 8 and grade 9 is 7.80 percentage points, or 780 steps of the same decimal, and applied to grade 8's Rs 2,352 crore it is worth Rs 73.38 crore. The wider gap holds 156 times as many steps and 124.8 times as much money. A consistent format is not a neutral choice: printing every grade to two decimals tells the reader that the last digit means the same thing everywhere on the scale, and it does not.
What happens when ten grades are pushed onto a five point axis?
Sooner or later a carefully graded estimate has to be reported alongside everything else, and everything else lives on a coarse scale. CompressionMapping a finer scale onto a coarser one, which loses detail and can also hide how thin the detail was. happens next: the fine scale is mapped onto the coarse one, and different things can then sit in one view. Suppose, purely as an illustration, that ten grades are mapped two to a point on a five point axis. The two to a point mapping is not the invented bank's practice, and its record does not say how it maps its grades onto its own axis.
Something is lost, and everybody expects that. Grade 3 and grade 4 are a factor of 2.25 apart and after compression they are one point. Losing that distinction is the ordinary price of putting different things in one view, and the price is usually worth paying. The serious cost is not what compression loses, it is what it hides: after the mapping, a figure that rested on a long default history and a figure that rested on a single observation are the same point on the same axis, and nothing downstream can tell them apart. The coarse scale did not introduce the error. The coarse scale removed the last place where the thinness of the fine estimate was still visible.
A ten grade estimate is placed on a five point axis. What has been lost, and what has been hidden?
The estimate that travelled, and where the reader lost sight of it
Go back to event category 3 and its single event, incident I11. Somebody writes 8.33 per cent a month. Nothing has gone wrong yet; the number is correct and the person who wrote it knows exactly what it rests on. Now watch it move. The figure is placed on a five point likelihood axis, where it becomes a 2 rather than a 1. A five point axis is an ordinal scaleA scale whose points are ordered but not evenly spaced, so a 4 is not twice a 2., so the 2 is a position and not a quantity. The 2 is combined with a consequence rating and the entry lands in one cell. The cell carries a colour. The colour decides whether the entry is escalated or left where it is.
Five objects, and by the third one the count is gone. Nobody removed it and nobody hid it. An axis point simply does not carry that field. The failure is not that somebody stated an unearned estimate; it is that the format the estimate travelled in had nowhere to record how thin it was, so by the time a decision was taken the provenance had evaporated without anybody noticing.
And this is why arguing about the arithmetic at the far end never works. Somebody at the escalation meeting says the rating looks high, somebody else recomputes 1 divided by 12 and confirms it is 8.33, and everybody moves on. The recomputation is correct. The recomputation answers a question nobody should have been asking. Only one question mattered, how many events sit behind the figure, and it stopped being askable four objects ago.
What is the difference between something measurable and something merely uncertain?
Frank Knight, in Risk, Uncertainty and Profit, published in 1921, drew a line that this guide has been walking along without naming. He separated the case where the odds can be established from the case where they cannot be known at all, and he reserved the word risk for the first and uncertainty for the second. Knight's line sounds like a philosopher's distinction and it is nothing of the kind. The line is the most practical idea in this guide, and it gives an analyst permission to say that a situation does not support a number.
Look at the invented bank's two records side by side and the line draws itself. Grade 5 of its rating scale carries Rs 14,112 crore of advances, the largest book on the scale, and behind the 0.90 per cent it states is that institution's own accumulated default history. Grade 5 can count. Event category 3 carries one recorded event. Event category 3 cannot count anything at all: there is a thing that happened, and there is no population to put it in. The honest response to the second situation is not a smaller number or a wider range; it is a different kind of sentence altogether.
The practical use of the distinction is defensive. Every reporting form ever designed has a box for a likelihood, and a box exerts pressure: it wants filling. Knight's line is what an analyst holds up when somebody says just put a best guess in. Unmeasurable uncertaintyA situation in which the odds are not merely unknown but cannot be established, which Frank Knight separated from measurable risk in 1921. is not an excuse for laziness. Unmeasurable uncertainty describes the situation actually at hand, and saying so out loud is more useful than a number that will be believed.
Who separated measurable risk from unmeasurable uncertainty, and where?
How is a likelihood stated honestly when the evidence is thin?
All of this arrives at something practical, and it is a habit rather than a technique. An honest likelihood statement has four parts, and the first three are the ones from the start of this guide: the event as defined, the chance, and the period. The fourth is the evidence the chance rests on, with the number of observations behind it. The fourth part is the one everybody drops, and it is the only one that lets a reader who was not in the room disagree with it.
There is a harder version of the same habit, and it is worth naming because people find it difficult to do in a meeting. Sometimes the honest output is a refusal. Reporting the observation and declining to convert it into a rate is a legitimate answer, it is usually the right one on a single event, and it is far more defensible than a figure everybody will quietly round in their heads anyway. What makes the refusal work is offering something in its place: here is what happened, here is what would be needed before a rate meant anything, and here is what is being done to collect it.
What can a likelihood estimate never tell?
Three things, and each of them is misread often enough to be worth stating separately. A likelihood cannot say when. An event estimated at a one in five chance this year says nothing whatever about where inside the year it falls; it is not due in the tenth month and it is not overdue in the eleventh. A likelihood cannot say which one. A grade 5 estimate applies to the grade, and it does not nominate the borrower. And a chance is a statement about a population of possible years, not a property of the one under way, so a likelihood cannot say anything at all about a single case.
A likelihood feels like a description of the thing in front of the reader, and it is a description of the set that thing belongs to. The single case catches experienced people for exactly that reason. The household version makes it obvious. Knowing that one house in a thousand on a street floods in a given year says a great deal about the street and nothing whatever about any one house on it. Both statements are true at once, and the second is not a weaker version of the first. The second asks a different question, one the first was never asked.
An event is estimated at a one in five chance this year. Does that say anything about when in the year it happens?
Who actually has to state a likelihood, and what do they do with it?
The estimate turns straight into money for a lender, so a lender is the clearest case. When a credit officer at the invented bank puts a borrower in grade 5 rather than grade 6, the estimated chance halves from 1.80 per cent to 0.90 per cent, and every downstream figure built on that grade halves with it. The officer is not producing a measurement and everybody in the chain knows it. Two things make the grade defensible: it comes from a stated place, the institution's own recorded history, and anybody can ask which grades have enough behind them to be worth two decimals and which do not.
An analyst reading that institution from outside does something different with the same figure. The outside analyst cannot check the history, and so checks the shape instead: does the distribution of exposure across the grades look like a real book or like a book where everything drifted into the comfortable middle, and does the scale span enough to distinguish anything at all? A scale spanning 500 times from end to end can separate a strong borrower from a weak one. A scale where every grade sits within a factor of two of its neighbour is barely a scale at all, and the analyst reads the ladder of multiples for exactly that.
And a household does the same arithmetic without the vocabulary. A scooter that has failed to start twice this winter supplies two observations, and one more failure would move the sense of the thing by half. Nobody says 16.67 per cent a month about it. Nobody would state a rate to two decimals on evidence that could be counted on one hand in their own life, and the professional version does exactly that. The institutional setting does not make it more rigorous. The institutional setting makes the habit harder to see.
Whose figures are these, and who binds an institution in India?
Every probability, grade, count, rupee figure and rate belongs to Vindhya Commercial Bank Limited and to its own twelve month record. The ten grade scale and the probabilities attached to it are that institution's own estimates from its own default history, and they are not a rating agency scale, a market figure or a supervisory classification.
Where an international structure sits behind something named here, it comes from the Basel Committee on Banking Supervision, published by the Bank for International Settlements at bis.org. The seven operational risk event categories used to sort the incident record are theirs. The Reserve Bank of India at rbi.org.in sets what an institution in India must actually estimate, how it must be documented, and what a supervisor makes of it.
The mechanism itself is free of any jurisdiction. An event, a chance, a period and the evidence beneath them behave the same way in every country and in every institution, and so does the arithmetic that says a rate built on two observations moves by half when a third turns up.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | What actually binds an institution in India on risk estimation, documentation and the treatment of its own internal ratings | rbi.org.in |
| Bank for International Settlements | The Basel Committee material behind the seven operational risk event categories used to sort the incident record | bis.org |
| Frank Knight | Risk, Uncertainty and Profit, 1921, where measurable risk is separated from unmeasurable uncertainty | In print since 1921 |
| SSRN | The Social Science Research Network (SSRN) carries academic literature on estimating rates from small counts, where a statistical treatment rather than a standard is the origin | ssrn.com |
Vindhya Commercial Bank Limited is invented.
Educational material. Not advice on any investment, tax, budget or market position.
