Randomness vs Uncertainty: A Known Shape or an Unknown One
Randomness is not knowing which outcome arrives when the weight on each outcome is known. Uncertainty is not knowing the weights. Under randomness a probability can be computed and defended. Under uncertainty a probability can still be computed, to any number of decimals, and those decimals describe the assumption rather than the world. One threshold here reads 8.00, 2.28 or 1.78 per cent depending on which shape was believed.
Underneath that answer sits one idea worth making solid. Every probability anybody quotes has two parts, and only one of them is arithmetic. The arithmetic sits on top of a set of weights, and those weights arrived from somewhere. Sometimes they were handed over, settled and countable, in which case the arithmetic inherits their solidity. Sometimes they were picked by whoever was doing the sum, in which case the arithmetic inherits that instead. The two cases produce numbers that look exactly alike in print, carried to the same number of decimals, and the difference between them is not visible anywhere in the working. The difference therefore has to be said in words. The distinction is about where an input came from, and no amount of care downstream will recover it if it was not stated upstream.
Two things built earlier carry the whole comparison, and no new machinery is needed. One is an invented one thousand company record called the Ketaki record, with an invented alert on it, where every count is settled and every conditional answer is exact. The other is a set of three invented shapes whose centre lands on 1.00 per cent and whose spread lands on 5.00 per cent in every case, and which still disagree about how often a bad month turns up. Both are re-read rather than rebuilt, and neither argument depends on remembering how either one was constructed.
What is randomness, stated on its own?
Randomness is the situation where the list of things that can happen is settled, the weight sitting on each of them is settled, and the only thing missing is which one is about to arrive. The fastest way to lose a comparison is to define each side as the absence of the other, so the other word is best set aside completely for now. Randomness is a state of knowing a very great deal. The claim that randomness is a state of knowing sounds wrong on first reading, and it is the most useful sentence here. The whole menu is known. How heavily each item on the menu is favoured is known. Exactly one fact is missing, and it is a fact about a single draw rather than about the world.
A shuffled deck is the clean case, and it is clean because nothing about it is in dispute. The deck holds fifty two cards. Every one of them is in there, once. Nobody has removed the queens or slipped in a sixth ace. After a proper shuffle, the card on top is equally favoured across all fifty two, and which one it is remains unknown. Everything worth knowing about the deck is known; the single thing missing is the order. Turning the top card over supplies the missing fact, and nothing it reveals revises what was believed about the deck itself.
The deck can feel like a toy. Take the household version instead. A shopkeeper runs a lucky draw at a wedding: two hundred slips in a drum, twenty of them marked. Anybody can count the slips, anybody can count the marks, and everybody in the hall agrees that a pull is favoured one in ten. Nobody in the hall knows which pull will land. The disagreement about the outcome is total and the agreement about the setup is total, at the same time, and those two facts sit together comfortably. Randomness is exactly that arrangement, and it is the condition under which a probability is a statement everybody can check. Randomness is also the condition under which a probability is a claim about the long runWhat would happen if the same setup were repeated a very large number of times. A probability of one in ten says twenty marked slips in two hundred, not that every tenth pull is the marked one. rather than a forecast of the next pull.
Before reading on, and without using the other word: what is randomness?
What is uncertainty, stated on its own?
Uncertainty is the situation where the weights themselves are not to hand. There may be a list of things that could happen, or only a partial one. There is often nothing to count, so a settled statement of how heavily each outcome is favoured is missing, and in many cases nobody has one. Uncertainty is not a larger amount of the same shortfall; it is a shortfall of a different kind, sitting one level further up. The gap is not in the draw. The gap is in the description of what is doing the drawing.
Here is the household version, and it is worth sitting with because it is so ordinary. A household has two years of electricity bills on a shelf. The household can count them, add them and see the spread: the low months, the high months, the months that surprised nobody. Now ask it what the next two years look like. The honest answer is not a wider version of the same spread. The honest answer is that nobody in that household knows whether the next two years are being produced by the same arrangement at all. A new appliance arrives. A room gets rented out. The tariff structure is redrawn. None of those is a bad draw from the old pattern; each is a different pattern. The bills on the shelf are countable and whatever will produce the next bills is not.
Uncertainty does something to the vocabulary. A number can still be written down. The chance of a bill above some level can be stated as such and such, carried to two decimals, and the arithmetic will be flawless. But the object the arithmetic was performed on was not handed over by anybody: whoever did the sum supplied it. The number is then a true statement about the thing supplied and an open question about the household. Under uncertainty the answer is exactly as good as the description it was computed from, and no better, however carefully it was computed. The sentence is a hard one to like, and it is the whole of this side of the distinction.
Now the other one, and again without using the first word: what is uncertainty?
Where exactly do the two part?
With both sides standing on their own, the work is already done and the contrast is short. The parting is not about how confident anybody feels, and it is not about how big the numbers are or how wide the spread is. The parting is about where the weights came from: whether they were an input handed over or an assumption somebody made. Everything else follows from that single fact, and nothing else needs to be added to it.
Here is the uncomfortable half. Once the weights are in place, the arithmetic downstream is identical in the two cases. Exactly the same steps, in exactly the same order, producing exactly the same kind of answer to exactly the same number of decimals. The finished figure carries no record of where its inputs came from, so a reader handed it cannot inspect it and work out which situation produced it. The lucky draw at the wedding and the household electricity bill both end in a number that reads as a percentage. The distinction has to be made in words instead, in the sentence carrying the answer, rather than looked for in the numbers. There is nothing in the numbers to find.
The second half is what each side permits as an honest statement, and this is more useful than any adjective. With settled weights a probability can be stated and defended, a disagreement about it can be called an error that somebody can be shown, and what would happen over a long run of identical setups can be stated as well. With unsettled weights three things can be stated: a range across the descriptions considered, the direction the answer moves as the description moves, and the name of the description in the same breath as the number. The second side does not permit presenting the number as though it came from the first side. The prohibition is not a small stylistic matter; it is the entire difference, expressed as a habit.
A properly shuffled deck of fifty two cards, face down. Which of the two is that, and what exactly is known?
What happens when three shapes meet one threshold?
Now the first of two worked cases, and both are re-read rather than rebuilt. An invented traded unitA made-up thing with a price that gets recorded each month. It stands in for anything whose value is recorded on a schedule, and it exists only so the arithmetic has something to be about. called the Nakshatra unit has its monthly change described in three different ways in the earlier notes. The first is a five value generator that puts weight on exactly five outcomes. The second is a smooth symmetric shape. The third is a two-part mixtureA shape built by holding two shapes at once, each with a share of the weight. Ninety five months in a hundred come from one and five in a hundred from the other. How it is built is covered separately. that holds a quiet month and a stressed month together. All three report a centre of 1.00 per cent and a spread of 5.00 per cent, and they agree on those two numbers exactly rather than approximately.
Now ask all three one question: how often does a month land at minus 9.00 per cent or below? The five value generator reads it straight off its running totalThe weight accumulated from the worst outcome upward, so that reading the total at any point gives the weight at or below that point. Built in the earlier note on shapes. and answers 8.00 per cent. The smooth symmetric shape answers 2.28 per cent. The two-part mixture answers 1.78 per cent. Push the question further out, to minus 14.00 per cent or below, and the generator answers 0.00 per cent because it cannot produce a month that bad at all, the symmetric shape answers 0.13 per cent and the mixture answers 0.72 per cent, or 5.32 times the symmetric reading. Look at what has happened to the ordering: at the nearer threshold the mixture gave the smallest of the three answers, and at the further one it gives the largest.
| One question, three descriptions, all invented | At minus 9.00 per cent or below | At minus 14.00 per cent or below |
|---|---|---|
| Five value generator | 8.00 per cent | 0.00 per cent |
| Smooth symmetric shape | 2.28 per cent | 0.13 per cent |
| Two-part mixture | 1.78 per cent | 0.72 per cent |
| Centre and spread reported by each | 1.00 and 5.00 per cent | 1.00 and 5.00 per cent |
The three agree on everything the arithmetic is allowed to look at, so no arithmetic anywhere chooses among them. The centre matches. The spread matches. All three summary figures can be held at once and the fork still has no signpost on it. And every single number computed after that fork inherits the branch that was taken, silently, with nothing in the output marking where the branch was. A reader who is handed 1.78 per cent has been handed the mixture's answer, and unless somebody wrote down that it is the mixture's answer, that fact has left the building.
A prediction, before the panel below. As the threshold is dragged further out into the bad tail, do the three readings converge on each other or separate?
Move the one threshold, and watch three descriptions that agree on everything give three answers.
One control, one variable: the threshold. The three shapes are drawn together and cannot be changed. Holding them fixed is the argument rather than a limitation of the panel. Choosing among them is not something the panel can do. The scale under the bars rescales as the threshold is dragged, so a gap that has become invisible in absolute terms is still drawn at full size, and the ratio beside it gives how many times over the widest reading covers the narrowest. The panel opens on the worked reading above, a threshold of minus 9.00 per cent.
Readings worth writing down so they survive without the panel. At minus 9.00 per cent the three answers are 8.00, 2.28 and 1.78 per cent, and the widest is 4.50 times the narrowest. Drag to minus 14.00 per cent and the generator drops to 0.00 per cent because it has no outcome that bad. The other two read 0.13 and 0.72 per cent, and the wider is 5.32 times the narrower. The absolute figures shrink and the disagreement between them grows. Most readers predict the opposite, and the far tail is exactly where a chosen description does the most damage. Pull the threshold back to 0.00 per cent instead and the three read 26.00, 42.07 and 40.48 per cent, still disagreeing, but now over ground where the disagreement is a matter of a few points rather than a matter of multiples.
Three descriptions share a centre of 1.00 per cent and a spread of 5.00 per cent. Which part of the arithmetic settles which one to use?
What happens when the alert itself does not change?
The second case is sharper still: this time not one thing about the machinery moves. The Ketaki record is an invented set of one thousand companies watched for a year, with an invented screening ruleA made-up test that looks at a company and either fires or stays quiet. It stands in for any rule that flags some cases for attention, and it exists only so the counting has something to count. called the Ketaki alert running over them. Forty of the thousand missed a payment, and of those forty it fired on thirty, a catch rateOut of the cases that really did have the problem, the share the rule fired on. Thirty of forty is seventy five in a hundred. Built in the earlier note on conditioning. of 75.00 per cent, and it fired on one hundred and twenty of the nine hundred and sixty that did not, which is a false alarm rateOut of the cases that were fine, the share the rule fired on anyway. One hundred and twenty of nine hundred and sixty is twelve and a half in a hundred. of 12.50 per cent. So it fired one hundred and fifty times, and of those one hundred and fifty, thirty had really missed a payment. The share works out at 20.00 per cent, exactly, and anybody can redo the division.
Now change one thing and one thing only. Leave the alert exactly as it is, with the same catch rate of 75.00 per cent and the same false alarm rate of 12.50 per cent, and move the base rateThe share of the whole record that had the problem before any rule looked at it. Four in a hundred here, and it is a property of the population rather than of the rule. from 4.00 per cent to 8.00 per cent. Eighty companies now miss a payment and nine hundred and twenty do not. The alert catches sixty of the eighty and falsely flags one hundred and fifteen of the nine hundred and twenty, so it fires one hundred and seventy five times, and sixty of those had really missed. Sixty divided by one hundred and seventy five is 34.29 per cent. The alert did not change, the arithmetic did not change, and the answer moved by more than fourteen points.
| The Ketaki record, invented, one thousand companies | Base rate 4.00 per cent | Base rate 8.00 per cent |
|---|---|---|
| Companies that missed a payment | 40 | 80 |
| Catch rate, held fixed | 75.00 per cent | 75.00 per cent |
| Alerts on companies that missed | 30 | 60 |
| False alarm rate, held fixed | 12.50 per cent | 12.50 per cent |
| Alerts on companies that were sound | 120 | 115 |
| Alerts in total | 150 | 175 |
| Of the alerts, the share that really missed | 20.00 per cent | 34.29 per cent |
Where the shortfall lives in that example is unusually clear. The shortfall is not in the rule; the rule was fully described and never touched. Nor is it in the division, a sum a school child can check. The shortfall is in one input: how many of the thousand were in trouble before anything looked at them. Handed over as a settled count, that number puts the whole calculation on the first side of the distinction and the answer is exact. Supplied as a judgement about the year ahead, it puts the calculation on the second side, and the exact answer is exact about that judgement. Same rule, same sum, and a fourteen point swing sitting entirely in an assumption that nobody was asked to defend.
The Ketaki alert is untouched, the division is untouched, and the answer goes from 20.00 per cent to 34.29 per cent. What moved?
Which of the two is a computed probability actually answering?
Put together, the two cases yield a plain answer. Every figure produced anywhere in these notes is an answer to a question asked of a stated description. How often does a month land at minus 9.00 per cent or below, given the two-part mixture? Of the alerts, what share really missed, given a base rate of 4.00 per cent? The description is not decoration in front of the question; it is part of the question. Strip it away and what remains is not a smaller claim, it is an incomplete one. A computed probability answers a question about the draw, and it does so conditional on a description that somebody supplied.
None of that is a criticism of the arithmetic, and it would be a bad reading to take it as one. The arithmetic is exact, it is checkable, and it is the only part of the whole business that cannot quietly go wrong. Refusing to compute because a description had to be chosen leaves nothing at all, and that is strictly worse than a number whose provenance is written down beside it. The honest position is neither of the two extremes. The honest position is that a probability is a conditional statement, that the condition is a real part of it, and that dropping the condition from the sentence is what turns a defensible number into a misleading one.
What can actually be done when the shape is unknown?
The list of what can be done is modest. There is a large and useful body of work on how to decide which description fits a record already in hand, and on how cases get built and stressed when no description is trusted at all. Both are covered separately. Four habits are on offer instead, smaller and available to anybody immediately, and they keep the shortfall visible instead of letting it disappear into a decimal.
The first is to put the assumption inside the sentence carrying the answer, rather than parking it somewhere further down. The sentence reads the chance is 0.72 per cent under the two-part description, not the chance is 0.72 per cent. The second is to compute under more than one description and report the range: 0.13 to 0.72 per cent says something true that either figure alone conceals. The third is to say which way the answer moves as the description moves. A reader who knows that a heavier far tail pushes this figure up can reason about descriptions that were never tried. The fourth is about the decision rather than the number: a course of action that survives every description considered is preferable to one that is best under the favoured description alone.
None of these four turns uncertainty into randomness, and anything claiming to has quietly put an assumption in somewhere. The claim is a working test rather than a slogan. When a method is offered that converts an unknown description into a known one, the honest version of that method always has a supplied input sitting inside it, and the useful question is simply where. Which input, what happens to the answer when it moves, and what it would take to be wrong about it. A method that cannot answer those three has not removed the assumption; it has only stopped showing it.
The right description genuinely cannot be told. Name two things that can still honestly be done.
Somebody offers a method that turns an unknown description into a known one. What should be looked for?
Why does a decimal place read like knowledge?
A false air of precision is the part a working reader meets most often, and it is worth stating as a portable rule. The figure 0.72 per cent has two decimals. The figure looks like something somebody measured, in the way a weight or a distance is measured, and every instinct a person has about numbers says that two decimals means somebody was being careful. Two decimals does mean somebody was being careful, and the care was entirely in the arithmetic. Precision is a property of the working and accuracy is a property of the assumption, and a figure can carry all of the first and none of the second.
Drawn out, the imbalance becomes silly. The last decimal place on that figure is one hundredth of a percentage point wide. The range that same figure takes across just the two descriptions used here runs from 0.13 to 0.72 per cent, fifty nine times wider than the decimal place being reported so carefully. Somebody has, in effect, measured a room to the nearest millimetre without being sure which room. Nobody involved is being dishonest, and that is what makes it durable: the arithmetic really was careful, and the carefulness really is in the wrong place.
Here is how a working reader uses this. When a figure with several decimals arrives, the first question is not whether the sum is right. Assume the sum is right; it usually is. The first question is which of its inputs was supplied by a person and how far the answer travels when that input moves. On the Ketaki alert, moving one input by four points of base rate moved the answer by more than fourteen points. On the Nakshatra unit, swapping one description moved a tail reading by more than five times over. A reader who asks that first is doing something a very careful arithmetic checker is not doing, and asking it decides whether the number can be leaned on.
A figure is quoted as 0.72 per cent. What is precise about it, and what is not?
Where this goes wrong, and it goes wrong in a sentence with no error in it
A reviewer writes one line into a note: the chance of a month at minus 14.00 per cent or below is 0.13 per cent. Checking it turns up nothing. The arithmetic is exact. The inputs are the ones they were handed. The reviewer did not round badly, drop a term or misread a table. Every objection that can be raised about the working comes back clean, and the sentence is still wrong.
The line is wrong because the description it stands on was chosen rather than counted, and a second description agreeing with every summary figure the reviewer holds, its centre landing on 1.00 per cent and its spread on 5.00 per cent just as theirs does, answers 0.72 per cent instead. The second description puts more than five times the weight in the same stretch of the tail. The reviewer never claimed the description was the only one. The reviewer simply did not mention that there was a choice, and the sentence they wrote carries no signal that anything in it is fragile. Nobody downstream can tell by looking which part was computed and which part was supplied.
The repair is a routine rather than a caution: put the assumption inside the sentence that carries the number, every time. The fragility then travels with the figure instead of being left behind at the desk where the sum was done. The reviewer's line becomes: under a smooth symmetric description the chance is 0.13 per cent, and under a two-part description carrying the same centre and spread it is 0.72 per cent. The rewritten line is longer by one clause and it cannot mislead anybody. The choice is now visible to every reader who gets it.
Why is no outside source named?
Because there is none to name. Every figure quoted here was produced by arithmetic on three invented objects, the Ketaki record, the Ketaki alert and the Nakshatra unit, and the arithmetic is shown on the notes those objects were built on. There is no outside document to consult, no maintained record to date, and so there is no reference table below. The absence is deliberate and it earns a sentence. A reader who has met reference tables on other notes will look for one here and should know why the space is empty.
Method notes of this kind take their warrant from the working rather than from a body that publishes the working. No supervisor, no trading venue, no index maintainer and no data vendor decides what a weighted count comes to, or what a count restricted to one column of a table comes to. The steps are the whole authority, and any number here is checked by redoing the count rather than by finding the institution that issued it.
The split between a shortfall that can carry weights and a shortfall that cannot has been written about in economics for a long time.
The Ketaki record, the Ketaki alert and the Nakshatra unit are invented.
Educational material. Not advice on any investment, tax, budget or market position.
