Probability in Finance: What a Number Between Zero and One Claims
A probability is a number between zero and one attached to a stated outcome inside a stated set of cases. The set is half the claim, and it is the half that gets dropped. Conditioning does not change the world, it changes which cases are counted over. One invented record holds one thousand companies. Forty of them missed a payment, four in every hundred.
Probability asks almost nothing of anybody. Counting, division, and the patience to ask a second question before accepting the first answer. The three together are the whole toolkit. Every figure below comes off one invented record of one thousand companies, whose four counts are stated the way the number of chairs in a room is stated rather than measured or sampled from anything. A word from finance gets defined the moment it turns up, so no earlier reading about lending is needed.
What does a number between zero and one actually claim?
A number between zero and one claims that across a stated collection of cases, a stated share of them show a stated feature. Three parts, and all three have to be present or the number claims nothing at all. The share is the number itself. The feature is the outcomeThe specific thing being asked about, described tightly enough that any single case can be sorted into yes or no without argument. Nobody can count a vague outcome., meaning whatever either happened or did not. The collection is the set of cases counted over, and it is the part almost everyone leaves out.
A probability quoted without the set it was counted over is not a slightly incomplete statement, it is a statement with its subject missing. Try it on something everyday. A tea stall outside one office building says two in five of its customers buy a second cup. The sentence sounds like a fact about customers, and it is not one yet. Two in five of which customers? The morning queue between nine and ten, when people are settling in and have time? Or every customer all day, including the four o clock rush where nobody has a spare minute? The morning queue and the all-day crowd are two different collections of people carrying two different numbers, and the stall owner who confuses them will order the wrong quantity of milk. The number two in five did not lie. The number simply never finished its sentence.
A probability arrives with no set of cases named. What is missing, and does it matter?
Where does a probability come from, and do the two routes make the same claim?
There are two routes to a number between zero and one, and once written down they look identical. The first route is counting. Given a record of cases, count how many carry the outcome, divide by how many cases there are, and the number falls out. Anybody can recount, so nobody has to take it on trust. The second route is judgement. Somebody considers a one off event that has never been in any record, a specific building being finished by a specific month, and puts a number on how strongly they believe it. The judged number is also between zero and one, and it obeys the same arithmetic. There is nothing to count, so nobody can recount it.
Both routes produce a number on the same scale, both are written the same way, and a reader who is not told which one they have been handed cannot tell them apart. The warning is not an argument against the second route. A considered belief is often the only thing available, and refusing to put a number on it does not make the uncertainty go away, it just hides it inside a word like probably where nobody can check it either. The point is narrower and it is about honesty of labelling. Counting calls for saying what was counted. Judging calls for saying that it is judgement, and saying what would change it. The two routes fail in completely different ways, so the route matters as much as the number.
Every single number below came off the first route. Not one of them was estimatedWorking out a number for a population that cannot be seen, by measuring a smaller group that can be seen and reasoning across the gap. The width of that gap is covered separately. from a sample. The record described below is fully in view, all thousand cases of it, so counting is the whole method.
A colleague says there is a 30 per cent chance a particular new building is finished by March. Which route produced that number?
Which three rules turn a probability into arithmetic rather than a feeling?
The rest of the working runs on one made-up record. Call it the Ketaki record. The record holds one thousand companies watched over one year. Two things were noted about each of them. First, whether the company missed a payment, meaning it owed money on a fixed date and did not pay on that date. Forty of the thousand did. Second, whether an invented screening ruleA fixed procedure that reads whatever data a business already holds about a case and sorts it into flagged or not flagged, without a human deciding case by case. called the Ketaki alert fired on the company during the year. The alert fired one hundred and fifty times. The two facts cross, so every company sits in exactly one of four boxes, and that is a two-way tableA grid that sorts every case by two facts at once, so each cell holds the cases carrying one particular combination of the two, and the rows and columns still add up to the whole..
| The Ketaki record, one year, invented | The alert fired | No alert | Row total |
|---|---|---|---|
| Missed a payment | 30 | 10 | 40 |
| Paid on time | 120 | 840 | 960 |
| Column total | 150 | 850 | 1,000 |
The four inside numbers are the only facts on the table, so they come first. Thirty companies missed a payment and had the alert fire on them. One hundred and twenty had the alert fire and paid everything on time. Ten missed a payment with no alert at all. Eight hundred and forty had neither. Each of those four is a joint countThe number of cases carrying two stated features at the same time, rather than either one on its own. A joint count sits inside a grid, never on its edge., and 30 plus 120 plus 10 plus 840 comes to 1,000 exactly. Every other number below is one of those four divided by something.
Now the rules, and there are only three. Every probability sits between zero and one, the whole record adds to one, and outcomes that cannot both happen are added together. Take them one at a time on the record above. First, a share of a set can never be less than none of it or more than all of it, so 0 and 1 are hard walls and a number outside them is an arithmetic mistake rather than a bold opinion. Second, each of the four counts divided by one thousand gives 3.00, 12.00, 1.00 and 84.00 per cent. Every company is in one box and no company is in two, so the four add to exactly 100.00 per cent. Third, missing a payment with an alert and missing a payment without one cannot both happen to the same company. The chance of missing a payment at all is therefore 3.00 plus 1.00, or 4.00 per cent, matching the forty companies in the row total. Three rules are the entire arithmetic. Everything harder in this subject is these three rules applied more times.
Forty of the one thousand companies in the Ketaki record missed a payment. What is the chance a company picked from the whole record missed one?
What is Conditional Probability, and what does conditioning actually change?
Conditional Probability is what appears when counting stops running over everything and runs over one named part of the record instead. Nothing else changes. The cases are the same cases, the outcomes are the same outcomes, the record has not been rewritten. All that moves is the denominatorThe number underneath a division, meaning the total being divided by. Change it and the answer changes even when the number on top does not move at all., and the answer moves with it.
Work it on the Ketaki record. Across the whole thousand, forty missed a payment. The chance of a missed payment is 40 divided by 1,000, or 4.00 per cent. Now condition on the alert. Conditioning on the alert means throwing away every company the alert did not fire on and counting only inside the one hundred and fifty that it did. Inside those one hundred and fifty, thirty missed a payment. So the chance of a missed payment given that the alert fired is 30 divided by 150, or 20.00 per cent.
The only thing that moved was the denominator. The same one thousand companies produced both numbers, and not one company changed in any way between them. Four in a hundred became twenty in a hundred without a single fact about a single company being different. Conditioning is not an event that happens to the world. Conditioning is a decision about what is being counted over, taken by whoever is doing the counting.
The other side of the same cut is the half people forget, so look at it now. Condition instead on no alert. Eight hundred and fifty companies had no alert during the year, and ten of them missed a payment anyway. So the chance of a missed payment given no alert is 10 divided by 850, or 1.18 per cent. Conditioning on the alert pushed the number up from 4.00 to 20.00 per cent. Conditioning on the absence of the alert pushed it down from 4.00 to 1.18 per cent. Both directions came out of the same four counts. The two conditions between them use up every company in the record, 150 plus 850 making 1,000 and 30 plus 10 making 40.
One more way to see the same cut, for anybody who finds bars unconvincing. The record can be split as a branching count. Take the thousand and split it first by whether a payment was missed, then split each of those two groups by whether the alert fired. Four numbers sit at the ends of the branches. Splitting the thousand the other way round, first by whether the alert fired and then by whether a payment was missed, reaches the same four numbers. The order of the questions changed everything about the route and nothing about the destination. The two conditionals below can therefore disagree while both being correct.
The Ketaki alert fired on 150 companies, and 30 of those missed a payment. What is the chance of a missed payment given that the alert fired?
The chance of a missed payment given no alert is 1.18 per cent, against 4.00 per cent across the whole record. What has conditioning done here?
Why is the chance of an alert given trouble not the chance of trouble given an alert?
One failure in this subject costs more than all the others, and it is worth slowing down for. Two sentences that sound almost identical in English are two completely different divisions.
Sentence one. Of the companies that missed a payment, how many did the alert fire on? Forty missed a payment and the alert fired on thirty of them, so the answer is 30 divided by 40, or 75.00 per cent. The 75.00 per cent describes the alert as a catcher. Whoever built the alert quotes that figure, and the figure flatters the work.
Sentence two. Of the companies the alert fired on, how many missed a payment? One hundred and fifty were alerted and thirty of them missed a payment, so the answer is 30 divided by 150, or 20.00 per cent. The 20.00 per cent describes the alert as a signal just received. A person sitting in front of an alert needs that second figure. The person is not looking at the forty. The person is looking at one of the one hundred and fifty.
Both sentences are counting exactly the same thirty companies, and they disagree because they put those thirty over two different denominators, forty in the first and one hundred and fifty in the second. Say it in the plainest possible way. Of every hundred alerts, twenty came from a company that missed a payment and eighty came from a company that paid everything on time. Four out of every five alerts point at a company that was fine. And the alert is still the same alert that catches three quarters of all the trouble. Neither statement cancels the other, and a reader who only ever hears the first will read the record backwards for as long as they hold it.
The Ketaki alert catches 75.00 per cent of the companies that go on to miss a payment. Does that mean 75.00 per cent of alerts are right?
How does the answer move when trouble becomes more common?
One question trips up almost everybody, and it is worth committing to an answer before reading on. Suppose missed payments become twice as common in some other year, and the alert itself is completely unchanged: same rule, same catch rate of 75.00 per cent, same rate of firing on companies that turn out fine. Does the share of alerts that turn out to be right go up, go down, or stay put? The quiz below takes an answer, and the control in the panel underneath shows what actually happens.
A prediction first. If missed payments became twice as common and the alert did not change at all, what happens to the share of alerts that are right?
Move how common trouble is. The alert never changes, and the answer moves anyway.
One control, and it moves one thing only: how many of the thousand companies missed a payment. The alert is welded shut. The alert catches 75.00 per cent of the companies that miss a payment at every setting, and fires on 12.50 per cent of the companies that pay on time at every setting. The thousand squares resort themselves, the four counts update, and the bar underneath redraws to show whichever question is selected. Start at the default of 40 companies, the published Ketaki record, reading 20.00 per cent. Then drag right and watch a number nobody touched climb.
The second readout does not move as the control moves. The catch rate is the definition of the alert, and the control does not touch the alert, so the second readout cannot move. The first readout climbs from 5.71 per cent when one company in a hundred misses a payment, through 20.00 per cent at the published record, to 60.00 per cent when one company in five does. An alert that has not changed in any respect becomes far better news simply because trouble became more common, and that is not a property of the alert at all, it is a property of the record it is fired into. A rule that worked beautifully in a bad year can therefore look like noise in a calm one, without anybody having touched a line of it.
What does it mean for two outcomes to be independent?
Independence is one of those words that sounds like a judgement and is actually a division. Two outcomes are independent when the chance of both happening together equals the chance of one multiplied by the chance of the other. The multiplication is all the word says. Independence does not ask whether the two things feel related, whether one seems to cause the other, or whether anybody expects a link. Independence asks whether one particular multiplication comes out right.
Run it on the record. The alert fired on 150 of the 1,000, or 15.00 per cent. A payment was missed by 40 of the 1,000, or 4.00 per cent. If those two outcomes were independent, the share of companies with both would be 15.00 per cent multiplied by 4.00 per cent, or 0.60 per cent. And 0.60 per cent of a thousand companies is six companies. The record holds thirty. Not six. The alert and the missed payments are not independent. The alert-and-missed cell holds five times the count independence would put there, and carrying real information means exactly that.
Two things hold at once, and this is where the argument ends up. The alert carries genuine information, five times what chance alone would produce. And four out of every five alerts still point at a company that paid everything on time. Both are true, both come off the same four counts, and neither weakens the other. A signal can be strongly informative and still be wrong most of the times it appears, and the reason is simply that there were an enormous number of sound companies available to be wrongly flagged and only forty troubled ones available to be rightly flagged.
Independence would put 6 companies in the alerted and missed cell, and the Ketaki record holds 30. What does that show, and what does it not show?
Can every figure come off one four-cell table?
Yes. Showing the division beside every figure is the whole method. Below is every number stated above, in the order it was reached, with the division that produced it. A figure absent from that list was never stated at all.
| What it is called | The division | The answer |
|---|---|---|
| Base rate, the chance of a missed payment across the whole record | 40 out of 1,000 | 4.00 per cent |
| Share of the record the alert fired on | 150 out of 1,000 | 15.00 per cent |
| Catch rate, the chance the alert fires given a missed payment | 30 out of 40 | 75.00 per cent |
| False alarmA case a rule flagged where the thing being looked for was not present. It is counted against the cases where nothing was wrong, not against the flags raised. rate, the chance the alert fires on a company that pays on time | 120 out of 960 | 12.50 per cent |
| The chance of a missed payment given that the alert fired | 30 out of 150 | 20.00 per cent |
| The chance a company paid on time given that the alert fired | 120 out of 150 | 80.00 per cent |
| The chance of a missed payment given that no alert fired | 10 out of 850 | 1.18 per cent |
| Independence check on the alerted and missed cell | 15.00 per cent of 4.00 per cent of 1,000 gives 6 against the 30 held | 5 times |
Two habits are worth stealing from that table. The first is that every row states its denominator out loud. Stating the denominator is what stops the fifth row being mixed with the third. The second is that the rows are in the order the reasoning went, so the last row visibly could not have been reached without the first two. A treatment that shows the division beside every figure has made itself checkable, and one that shows only the answers is asking to be believed.
The failure a swapped denominator causes
A reviewer at a lenderA business whose work is handing over money now and being paid back later, usually with a charge for the wait. What a lender does when repayment stops is covered separately. is told the Ketaki alert catches three quarters of the companies that go on to miss a payment. Nothing about that sentence is false. The reviewer hears it as three quarters of alerts being right, a different sentence entirely, and starts the year treating each of the 150 alerts as a problem to be worked. One hundred and twenty of those 150 are companies that pay everything on time. So the reviewer opens 120 files that did not need opening, and each one takes real hours from a week that was already full.
The expensive part is not the wasted work. The expensive part is what the reviewer learns from the wasted work. Within a month the pattern is unmistakable from the inside: most of these alerts come to nothing. The reviewer begins skimming them, then deprioritising them, then clearing them in batches. By then the alert has been reclassified in one human head as noise, and the thirty that were genuine get cleared in exactly the same sweep. A rule that carried five times the information chance alone would produce has been switched off by a misreading of one number.
The fix costs one question, asked before any of that work starts. Which denominator is this number sitting on? Not how often the alert fires when there is trouble, 30 over 40. How often there is trouble when the alert fires, 30 over 150. The second is the number the person holding an alert is actually standing in. Ask for it by name and the whole misreading is impossible.
Four in five of the companies the Ketaki alert flags turn out to have paid everything on time, and a reviewer calls the alert useless. What is right about that, and what is wrong?
What should be asked before accepting any probability handed over by somebody else?
The method becomes usable rather than merely correct at the moment somebody else hands over a number. Anybody who works with numbers, whether they sit in a lending office, read reports for a living, or run a small business and are being shown a chart by a supplier, gets handed probabilities constantly. Most of them arrive stripped. Seven questions get the missing parts back. Each one has a one line answer on the Ketaki record, and an answerable question is not a rhetorical one.
The swap between the two conditionals is the failure that costs the most and shows the least, so question four does most of the work. A wrong denominator does not announce itself. The arithmetic is right, the sentence is grammatical, and the number is plausible. The only way it surfaces is when somebody says the phrase out of what out loud. Get into the habit of writing conditionals as a fraction with words on both parts, thirty of the one hundred and fifty alerted rather than twenty per cent, and the swap becomes physically difficult to make. Householders do a version of this every day without calling it anything: told that most burst pipes happen in old buildings, the useful follow up is not whether that is true, it is how many old buildings never had a burst pipe at all.
What sits behind the figures here, if no source is named?
Four invented counts sit behind every figure here, and every percentage is one of those counts divided by a total, so there is no outside document to consult and none is named. A lending rate or a tax slab is set somewhere by somebody and moves, so it would owe the reader a document and a date. The numbers here are not set anywhere. The percentages are arithmetic performed on counts invented for teaching, and the only instrument that can check them is a calculator.
| The statement made | How it was produced | What checking it needs |
|---|---|---|
| The four counts of the Ketaki record: 30, 120, 10 and 840 | Made up for teaching. The four counts are the starting point, not a measurement of anything | Nothing. There is nothing underneath them |
| The base rate, the share alerted, the catch rate and the false alarm rate | One of the four counts divided by a row total or a column total | A calculator |
| The chance of a missed payment given an alert, and given no alert | The same four counts, each read against a different total | A calculator |
| The independence check at six against thirty | Two of those percentages multiplied together and applied to one thousand cases | A calculator |
The Ketaki record and the Ketaki alert are invented.
Educational material. Not advice on any investment, tax, budget or market position.
