Random Variable: Turning an Uncertain Outcome Into a Number
A random variable is a rule that hands every outcome a number. The rule exists first and the number arrives only when an outcome does. Numbers can be added, averaged and ranked in a way that outcomes cannot, and the whole step exists for that reason. The expected value is those numbers averaged with each one weighted by its chance, and it is a balance point rather than a forecast.
Underneath the answer sits one habit of thought that costs nothing to acquire and pays for itself immediately. Uncertainty in ordinary life arrives as a description: the month went badly, the alert fired, the delivery was late. A description cannot be halved, doubled or averaged. So before any arithmetic can start, somebody has to decide what number each description gets. Choosing that number is a decision, not a discovery, and every later step inherits the choice.
A chance between zero and one, and the reference setThe stated collection of cases a chance is counted over. Two hundred deliveries this month, or one thousand recorded companies. it is counted over, are both set out under probability. Everything else is multiplication and addition.
What is a random variable, in plain words?
A random variable is a rule that says which number goes with which outcome. The definition is complete at that, and the word carrying it is rule. A random variable is not the number, and it is not the outcome either.
Before the uncertain thing resolves, the rule already exists and the number does not, and holding those two apart is the single distinction most readers slide past. The rule can be written down today, printed, argued about and changed, all without knowing a single outcome. When the outcome finally arrives, the rule does no work at all: it simply reads off the number it had always promised to that outcome.
Take a household running on one salary, about as far from a trading floor as anything gets. Nobody in that household knows what this month's electricity bill will be. But everybody knows the rule: so many units at one rate, so many more at the next rate up, plus fixed charges. The rule is printed on the back of last month's bill. The tariff exists in full, unambiguously, before the meter has been read. The bill is the number and the tariff is the random variable. Notice that if the tariff changes, the rule has changed even though the month has not.
The word random is doing less work here than it looks. Random does not mean chaotic, unknowable or lawless. Random means only that which outcome turns up is uncertain. The rule attached to those outcomes is often completely fixed and completely known, so a great deal can be worked out in advance.
Which of these three is the random variable: the number, the outcome, or the rule joining them?
Why bother turning an outcome into a number?
Because outcomes will not do arithmetic and numbers will. A bad month cannot be added to a good month. There is no average of the alert fired and the alert did not fire. A late delivery cannot be ranked against a wet monsoon. The moment each of those gets a number, all three operations become available at once.
The mapping from outcomes to numbers is the step that makes every later method possible, and every method that follows inherits whatever the mapping decided. Averages, spreads, comparisons and models all sit downstream of it. If the mapping is careless, nothing downstream can repair it. By then the descriptions are gone and only the numbers remain.
Which is why the second half of the point matters more than the first. Choosing the mapping is a decision somebody makes, not a fact about the world. The world produces months; it does not produce percentages. A percentage is what a person decided to call a month.
Think of a vegetable seller keeping a rough note of each day. The day could be recorded as takings in rupees, or as kilos sold, or as a mark out of ten for how the day felt. All three are legitimate rules over the same days, and they will disagree about which was the best day of the week. A day of high takings on cheap stock is a different day when measured in kilos. Nobody is wrong. The three rules simply differ, and the choice between them happened before any arithmetic did.
How does a yes-or-no outcome become a number?
The simplest mapping there is produces the neatest result of the lot. An outcome that either happens or does not is sent to 1 when it happens and to 0 when it does not. The rule is complete at that. Two outcomes, two numbers.
Use the Ketaki alert for it. The Ketaki alert, an invented screening ruleA test applied to every case in a record, firing on some cases and staying quiet on others. What such a rule is worth, and how often it is right, is set out under probability., was run across one thousand invented companies over one year. It fired on 150 of them and stayed quiet on the other 850. Send fired to 1 and quiet to 0.
Now average that variable across the record. Add up the numbers: 150 companies contribute 1 each and 850 contribute 0 each, so the total is 150. Divide by the 1,000 companies in the denominatorThe number sitting underneath in a division. The denominator fixes what the answer is being measured against, and changing it makes the same total give a different answer. and the result is 0.15.
The average of a one-or-zero variable is exactly the chance of the thing it marks, and that is not a coincidence of these particular numbers. It falls out of the arithmetic every time. The zeros contribute nothing to the total, so the total is just a count of how often the thing happened, and dividing a count by the size of the record is how the chance is obtained in the first place. Averaging and counting turn out to be the same act wearing two names.
The identity is the hinge between counting and averaging, and it is worth sitting with. Averaging works on numbers and counting works on cases, and the one-or-zero variable shows the two were never separate subjects.
An outcome is mapped to 1 when it happens and 0 when it does not. What is the average of that variable across the record?
What does the monthly change of the Nakshatra unit look like written out?
The Nakshatra unit is an invented traded unitSomething that can be bought and sold at a price, so the price is whatever the last buyer and seller settled on. Nothing about how such a price gets set is needed here. priced once a month and starting at Rs 100/-. Its monthly change is the change in the price over the month, written as a change per hundred rupees of the starting price. If the price goes from Rs 100/- to Rs 106/-, the monthly change is 6.00 per cent. If it goes to Rs 91/-, the monthly change is minus 9.00 per cent.
The Nakshatra monthly change takes five values and no others: minus 9.00, minus 4.00, 1.00, 6.00 and 11.00 per cent. The weight on each is stated, not measured: 0.08, 0.18, 0.48, 0.18 and 0.08.
The five weights are a stated property of a made-up generator, not an estimate of anything. Nobody watched a hundred months and counted. The generator was written down with those weights already attached, the way the tariff on the back of the electricity bill is written down. How anybody would work out a weight from months somebody had actually watched is a separate subject.
Five values is a simplification chosen so the arithmetic stays visible on one screen. The simplification is not a claim that a price can only do five things. A real price can land almost anywhere, and what changes when it can is the last question below.
| Monthly change | What that means at a starting price of Rs 100/- | Weight |
|---|---|---|
| minus 9.00 per cent | The price ends the month at Rs 91/- | 0.08 |
| minus 4.00 per cent | The price ends the month at Rs 96/- | 0.18 |
| 1.00 per cent | The price ends the month at Rs 101/- | 0.48 |
| 6.00 per cent | The price ends the month at Rs 106/- | 0.18 |
| 11.00 per cent | The price ends the month at Rs 111/- | 0.08 |
| Five values, nothing else possible | Every month lands on exactly one row | 1.00 |
The weights add to 1.00 and they have to. The five rows cover every month the generator can produce and no month can be on two rows at once, so the weights are a complete account of where a month can go. A set of weights adding to less than one has left something out, and a set adding to more than one has counted something twice.
The five weights are 0.08, 0.18, 0.48, 0.18 and 0.08. What must they add to, and why?
How is the expected value actually computed?
Multiply each value by its weight, then add the five products. The method is complete at that, and no second step hides behind it.
The shape of the arithmetic is easier to trust when every term is visible, so the five products are worth writing out one at a time. Minus 9.00 times 0.08 gives minus 0.72. Minus 4.00 times 0.18 gives minus 0.72 as well. Then 1.00 times 0.48 gives 0.48, 6.00 times 0.18 gives 1.08, and 11.00 times 0.08 gives 0.88.
| Value | Weight | Value times weight | Running total |
|---|---|---|---|
| minus 9.00 per cent | 0.08 | minus 0.72 | minus 0.72 |
| minus 4.00 per cent | 0.18 | minus 0.72 | minus 1.44 |
| 1.00 per cent | 0.48 | 0.48 | minus 0.96 |
| 6.00 per cent | 0.18 | 1.08 | 0.12 |
| 11.00 per cent | 0.08 | 0.88 | 1.00 |
| The expected monthly change | 1.00 | 1.00 | 1.00 per cent |
The answer of 1.00 per cent is a property of the stated generator rather than something anybody observed. The five products add to 1.00 in any order. Changing one weight moves the answer. The panel below does exactly that.
The running total takes a route of its own on the way. The total dives to minus 1.44 after the two falls, climbs back through minus 0.96, crosses zero on the fourth term and finishes at 1.00. The order of the terms is arbitrary and the path would be a different shape if they were added the other way round, but the destination would not move. Only the total is meaningful.
Add the five products shown above: minus 0.72, minus 0.72, 0.48, 1.08 and 0.88. What is the expected monthly change?
Worth settling before the panel below is touched: suppose the weight on the worst value, minus 9.00 per cent, doubles from 0.08 to 0.16, the middle value absorbs the difference, and the five values themselves do not change at all. What happens to the expected value?
Move the weight on the worst value and watch the balance point slide.
The panel opens on the generator exactly as it stands above: a weight of 0.08 on minus 9.00 per cent, 0.48 left on the middle value, an expected value of 1.00 per cent and a 26.00 per cent chance of a fall. Drag the slider and only one thing changes: how much weight sits on the worst value. The middle value gives up or takes back whatever is needed so the five always add to 1.00, and the five values themselves never move a millimetre. Watch the red marker slide while the tick marks stay exactly where they were.
Why is the expected value not a prediction?
Because nothing obliges a single month to land on it. The expected value of the Nakshatra monthly change is 1.00 per cent. The same generator produces a fall in 26.00 per cent of months, and one month in twelve and a half is a fall of 9.00 per cent. Both statements are true of the same object at the same time and neither softens the other.
An expected value is where the weighted values balance, not a value the next observation is under any obligation to take. The clearest way to feel that is a fair six-sided die. Its expected value is 3.5. No face on the die shows 3.5, and no throw of it ever will. The average is a real and useful number that describes the whole object, and it names an outcome that cannot happen.
For the Nakshatra unit, 1.00 per cent happens to be one of the five values it can produce. The coincidence belongs to this particular generator and is not a rule about expected values. Shift a weight in the panel above and the balance point slides off the tick mark it started on, at which point it names a monthly change the unit is incapable of producing, exactly like the die.
The expected value is for the long runWhat happens across very many repeats rather than in one instance. The phrase is used loosely here, as an intuition. The arithmetic that makes it precise is a separate subject. and for comparison. A second invented generator with an expected value of minus 1.00 per cent would say something real about how the two differ on average, without either statement saying anything about next month. How far apart the five values sit is the spreadHow widely spaced the values are around the centre. Two quantities can share a centre while one of them is far tighter than the other., and the expectation says nothing at all about it. Spread is a separate question.
The expected monthly change is 1.00 per cent, a rise. What is the chance the month is actually a fall?
When does averaging without the weights give the wrong answer?
Here is the trap, and it is sharper than it looks. On this generator the wrong method gives the right answer.
Average the five values with no weights at all. Add minus 9, minus 4, 1, 6 and 11 to get 5, then divide by the 5 values. The answer is 1.00 per cent. The careful weighted arithmetic produced exactly that answer, after five multiplications and a running total. A reader who spots this concludes, quite reasonably, that the weights were decoration and the shortcut works.
The agreement is a coincidence of these particular weights and it proves nothing whatsoever about the method. It happens because the five values are evenly spaced, five points apart, and the five weights are symmetricArranged as a mirror image about the middle. Here the first weight matches the fifth and the second matches the fourth, so the two sides balance each other out. about the middle: 0.08 matches 0.08 at the far end, and 0.18 matches 0.18. Every pull to the left has an identical pull to the right, so both methods land on the middle value. Change either of those two conditions and the agreement evaporates.
So change one. Keep the same five values, minus 9.00, minus 4.00, 1.00, 6.00 and 11.00 per cent, and put different weights on them: 0.30, 0.30, 0.20, 0.10 and 0.10. The new weights still add to 1.00, so they are a legitimate set. The weight has been piled onto the bad end, so they are simply not symmetric.
| Value | Alternative weight | Value times weight |
|---|---|---|
| minus 9.00 per cent | 0.30 | minus 2.70 |
| minus 4.00 per cent | 0.30 | minus 1.20 |
| 1.00 per cent | 0.20 | 0.20 |
| 6.00 per cent | 0.10 | 0.60 |
| 11.00 per cent | 0.10 | 1.10 |
| Weighted average, the right method | 1.00 | minus 2.00 per cent |
| Unweighted average, the shortcut | not used | 1.00 per cent |
The shortcut never looked at the weights, and the weights are the only thing that changed. So the right method now says minus 2.00 per cent while the shortcut still says 1.00 per cent. The two answers are 3.00 points apart, and they disagree about the direction: one says the unit tends to lose value and the other says it tends to gain.
A method that agrees with the right answer on one case has not been tested, it has been lucky. That is the general rule worth carrying away, and it is not really about averages. Any shortcut checked against exactly one worked example has been checked against the easiest possible evidence. The useful question is never does it agree here, it is what would have to be true for it to disagree, and does that thing hold.
Averaging the five values with no weights also gives 1.00 per cent. Does that show the shortcut works?
The same five values under weights of 0.30, 0.30, 0.20, 0.10 and 0.10 give a weighted average of minus 2.00 per cent. What does the unweighted average give on those same values, and what does the gap show?
The error: reading a balance point as a forecast, and a lucky agreement as a verified method
An analyst reports that the expected monthly change of the Nakshatra unit is 1.00 per cent. A reader hears that as what next month will do, and builds a plan on it. But the generator produces a fall in 26.00 per cent of months and can hand over a fall of 9.00 per cent. A fall that size takes a Rs 100/- unit to Rs 91/- inside one month. The plan was built on a centre of gravity that no single month is obliged to visit, and the expected value never carried any information about how far from that centre a month can land.
The quieter version of the same failure catches more careful readers. Somebody averages the five values without their weights, gets 1.00 per cent, sees it match the careful arithmetic, and files the shortcut away as verified. The shortcut was verified against exactly one case, chosen by nobody, that happened to be symmetric. The first time they meet a set of weights piled onto one end, the shortcut hands them 1.00 per cent when the answer is minus 2.00 per cent, and it does so silently and with the wrong sign.
The cost in both versions is the same shape: a number that looks fully worked and is being asked a question it cannot answer. The fix is two questions. Ask what else the object can produce, not just where it balances. And when a shortcut agrees with a careful method, ask what would have to change for it to disagree, before deciding the shortcut is safe.
What changes when the quantity can land anywhere in a range?
Everything so far has leaned on one convenience: there are five values and they can be listed. A quantity whose values can be listed is discrete, and listing is what makes the weights writable one by one.
Now suppose the monthly change can be any number at all between two limits, not just five. The change could be 1.00 per cent, or 1.0004 per cent, or 1.00041 per cent. Between any two values named there is always another, so there is no next value after 1.00 per cent. The list would have to be endless, so the weights cannot be written out one by one, and that is the whole of the difference.
A curve replaces the list. The weight sitting on exactly 1.0004 per cent turns out to be nothing at all, so the question becomes what weight sits on the stretch between 1.00 and 2.00 per cent, and the curve answers that. The honest question about a continuous quantity is always about a range, never about a point. Why a single point carries no weight, how the area under such a curve works, and what the whole picture is called are all part of the distributionThe complete account of a quantity: every value it can take, with the weight on each. Defining one properly, and what a centre and a spread fail to say about it, is a subject of its own., set out under probability distributions.
Everything established so far is unchanged. The mapping from outcomes to numbers is the same act. The expected value is still a weighted average and still a balance point rather than a forecast. The unweighted-average trap is still a trap. Only the bookkeeping of the weights changes, from a list to a curve.
Why can a quantity that lands anywhere in a range not have its weights listed one by one?
How does anybody put an expected value to work?
Four situations, all of them ordinary, none of them requiring a market.
A household deciding whether to take an annual maintenance contract is doing this arithmetic whether or not it writes anything down. The contract costs a fixed amount. Without it, some years cost nothing and one year in several costs a large repair. Listing the repair outcomes, putting a weight on each and adding the products gives a number that can be set beside the contract price. The value of the arithmetic is not that it settles the decision, it is that it forces the weights into the open where somebody can argue with them.
A lender pricing a small loan does the same thing with two outcomes and a one-or-zero variable. Repaid gets 0, not repaid gets 1, and the average of that variable across a book of loans is the chance of not being repaid. The one-or-zero variable is doing paid work there.
Expected values are good at comparison rather than prediction, and an analyst comparing two invented projects uses them for exactly that. Two projects with expected values of 1.00 per cent and minus 2.00 per cent differ in a way worth knowing, and neither number says anything about how either project will actually turn out.
And a reader handed any expected value at all has three questions to ask, in this order. Which outcomes were listed, and is anything missing from the list? Which weights were used, and where did they come from, given that the weights and not the values are what move the answer? And how far can a single case sit from this number? The expected value on its own is silent about that. A number offered without answers to those three is not wrong. The number is just less informative than it looks.
What sits behind the numbers?
Every figure below is made up and every step is arithmetic. The weights are stated rather than measured, so there is no series to look up and no institution whose judgement carries any part of the answer.
| What the arithmetic uses | Where it comes from | Site |
|---|---|---|
| The five monthly changes and the five weights on them | Stated properties of a made-up generator, not measured from anything | None. No site was read |
| Every product, total and percentage in the build | Worked step by step, so each line can be checked with a pen | None. Arithmetic only |
| The fired-or-quiet example that becomes a one and a zero | The made-up Ketaki alert, which belongs to probability | None. No site was read |
The Nakshatra unit, the Ketaki record and the Ketaki alert are invented.
Educational material. Not advice on any investment, tax, budget or market position.
