The Probability Space: The Formal Setting Behind Every Model
A probability space is three objects taken together: the set of everything that could happen, the collection of subsets that may be assigned a number, and the rule that gives each of those subsets a number between zero and one. Pricing theory changes the third object and leaves the first two alone. One model can therefore carry two probabilities for the same event.
Probability is not a property of a single outcome. Probability is a property of a set of outcomes, handed out by a rule that has to stay consistent across every set at once. The shift from outcomes to sets is the whole purchase of the formal setting, and it is what makes it possible to replace the rule without disturbing anything else in the model.
Why does a model need a formal setting before it needs a number?
Most people use probability for years without once writing down a probability spaceThe three objects together: the outcomes, the allowed subsets, and the rule that assigns each subset a number.. Working without one is fine, right up to the moment somebody changes the probabilities. Then the question no informal account can answer is the only question that matters: what else changed when they did that?
The formal setting exists so that it becomes possible to say exactly what is being held fixed while something else moves. Told only that a probability moved from 0.617911 to 0.559618, nobody can tell whether the world got worse, whether somebody added outcomes nobody had thought of, or whether the same world is simply being weighed on a different scale. A worse world, a wider outcome set and a different scale are not the same situation, and they have completely different consequences. The three part construction is what separates them.
Here is the everyday version. A bag of rice weighed on the kitchen scale and then on the scale at the shop gives two different readings. Nothing happened to the rice. The instrument differs, and no amount of staring at the bag will settle which reading to trust. A model that keeps the bag and the scale in separate boxes can talk about changing scales. A model that keeps only the reading has thrown away the thing the reading was about, so it cannot.
| \(\Omega\) | the set of everything that could happen, with each possibility appearing exactly once |
| \(\mathcal{F}\) | the collection of subsets of \(\Omega\) that may have a number attached |
| \(\mathbb{P}\) | the rule, a measure, sending each subset in \(\mathcal{F}\) to a number between zero and one |
Of the three objects, which one has to be settled before either of the other two means anything at all?
What exactly is the set of everything that could happen?
The outcome setEvery single thing that could happen, listed once, with nothing left out and nothing counted twice. is written with the capital Greek letter omega, and a single member of it is written with the small omega. For the standard process, an invented traded quantity this whole subject runs on, one outcome is not a price. One outcome is an entire year of prices: the value at every instant between the start and the horizon, taken as one indivisible thing.
The whole-path outcome is the first genuinely unfamiliar idea in this subject, and it repays sitting with. An outcome is a whole path, not a finishing value, and everything the model can ever be asked has to be a statement about paths. The locked path published for this subject is one member of the outcome set: Rs 100/- at the start, Rs 97.64/- after a month, up to Rs 111.08/- at month six, down to Rs 93.74/- at month nine, and Rs 106.18/- at the horizon. The entire sequence of readings is one omega. Change the reading in month three and the result is a different omega, even if the finish is identical.
Why insist on that? Because it decides what the model can be asked. Pacing out a coastline with a ruler is the same phenomenon. A long ruler gives one length; a shorter ruler finds inlets the long ruler stepped straight over, and the length grows. The ruler is not measuring the coast wrongly, it is measuring a coarser version of it. An outcome set of finishing values only is the long ruler: it can answer where the year ended and can never answer whether the process ever fell below Rs 95/-. Whether the process fell below Rs 95/- is not a statement about anything in that set. Choosing the outcome set is choosing, in advance, the entire list of questions the model is capable of answering.
Which subsets may be given a number?
An eventA subset of the outcome set, so a statement that is either true or false for each individual outcome. is a subset of the outcome set. An event is therefore a statement that is either true or false for each path taken one at a time. The process finishes above Rs 100/- is an event. The process never falls below Rs 95/- is an event. The process is at Rs 106.18/- at the horizon and was above Rs 110/- at month six is an event, and the locked path belongs to it.
The collection of events is written with a script capital F. The collection is not simply every subset that can be imagined; it is a collection with structure, called a sigma-algebra, and the structure is what makes the rule consistent. A sigma-algebra has to contain the whole outcome set, it has to contain the opposite of anything it contains, and it has to contain the union of any endless list of its members. The three closure requirements exist so that any question that can be asked has an askable negation, and so that any run of questions joined by the word or can be asked as well.
When the collection carries a time subscript it becomes the information available at that time, and everything in it is a question that could already have been settled by then. Whether the process rose in the first month sits inside the collection at month one. Whether it finishes above Rs 100/- does not. At month one the year has not happened yet. The subscripted collection is the object the rest of this subject leans on hardest, and it is set out under filtration.
Why the collection is not simply all subsets, and what breaks if it is insisted that it is, is a genuine result rather than a technicality, and it is set out under sigma-algebras. Here it is enough that the collection is chosen first and the rule only ever speaks about its members.
What must the rule satisfy to count as a probability at all?
Probability Measure
A probability measureThe rule assigning each allowed subset a number between zero and one, consistently across all of them at once. is a function whose input is a set and whose output is a number. A measure is not a formula about values and it is not a shape on a chart. Hand it an event and it returns that event's weight. Three conditions decide whether a candidate rule earns the name, and a rule that misses any one of them is not a slightly imperfect probability, it is not a probability.
| \(A\) | any event, meaning any member of the collection \(\mathcal{F}\) |
| \(\Omega\) | the whole outcome set, the event that is true for every path |
| \(A_i\) | an endless list of events, no two of which can both happen |
| \(\mathbb{P}\) | the candidate rule being tested against the three conditions |
The third condition is the one that does the work, and it has a name: countable additivityProbabilities of sets that cannot overlap add together, and this holds even for an endless list of such sets.. Finite additivity, where any two separate events add, sounds like the same thing and is strictly weaker. Extending it to an endless list is what permits limits to be taken, and taking limits is the only way to say anything at all about a set of continuous paths. Almost every interesting event about the standard process is built as a limit of simpler ones.
Notice what the conditions do not say: nothing in them mentions frequency, belief, betting or the world. They are consistency requirements on a bookkeeping system. Consistency requirements are precisely why a second, differently motivated rule can sit on the same outcome set without contradiction. The second rule satisfies the same three conditions, so it is just as much a probability as the first, whatever anybody intends it to mean.
Somebody proposes a rule that gives every event a number between zero and one and gives the whole outcome set the number one. The rule then gives two events that cannot both happen the numbers 0.6 and 0.7, and gives the event that either of them happens the number 0.9. Is it a probability measure?
How is a probability measure different from a probability distribution?
Probability Measure vs Probability Distribution
Measure and distribution get used as though they were interchangeable, and they are different kinds of object. A measure eats a set. A distribution eats a value. The confusion is harmless in an applied course and fatal in this one. The entire manoeuvre this subject is built around is replacing one of them while the other side of the model holds still.
| \(\mathbb{P}\) | the measure, whose input is a set drawn from \(\mathcal{F}\) |
| \(F(x)\) | the distribution function, whose input is a number on the rupee scale |
| \(\omega\) | one outcome, meaning one complete path of the process over the horizon |
| \(S_T\) | the value of the standard process at the horizon \(T\), read off a path |
The distribution is downstream: it is what remains after the measure has already spoken. The right hand side above shows it. The braces build a set of paths out of a number, and then the measure weighs that set. With the measure removed, the distribution has nothing to compute with. With the distribution removed, the measure is unharmed. The measure never needed the rupee scale in the first place.
Back to the scales in the shop. The bag of rice is the outcome set. A scale is a measure. The printed table of what the scale reads for each bag size is a distribution: useful, real, and entirely a consequence of which scale was used. Swapping the scale changes the table. Nothing done to the table will swap the scale.
Which of these two is handed a set of outcomes and returns a number: the measure, or the distribution function?
What does arriving information change, and what does it leave alone?
Conditional Probability
ConditioningRestricting attention to the outcomes still possible, then rescaling the rule so the remaining weights total one again. is usually taught as a formula with a fraction in it, and the fraction hides the construction underneath. Conditioning builds a new probability measure on a smaller outcome set. Not a fraction of the old one, and not a temporary adjustment. A full rule in its own right, satisfying all three conditions.
| \(A\) | the event being asked about |
| \(B\) | the event the arriving information has established as true |
| \(A\cap B\) | the outcomes on which both are true, meaning the part of \(A\) that survives |
Read as three steps it stops being a formula and becomes a procedure. Discard the outcomes the information has ruled out. Keep the relative weights of everything that survives untouched. The numerator preserves ratios exactly. Then divide through by the total that is left, so the survivors add to one again. The relative weights never move; only the total is restored, and that is the whole content of conditioning.
The everyday version is a lift that only knows which floor it is on. Standing on the fourth floor, the set of floors reachable next has shrunk, but nothing has changed about how the building is laid out. The lift did not become a different lift when the doors opened. The list of floors that can still be asked about changed, and the weights on what remains were stretched to fill the space the discarded floors used to occupy.
Conditioning on arriving information rescales the probabilities of whatever outcomes are still possible. What do those rescaled numbers add up to?
The standard process has an average finish of Rs 108.33/-. Is the chance of finishing at or above that average more than half, or less?
What do the three objects look like on the standard process?
Time to build one. The standard process starts at Rs 100/-, drifts at 8 per cent a year, carries a volatility of 20 per cent a year, pays nothing out, and is watched for one year. Four numbers describe the rule and nothing else. A rule alone does not describe a probability space, and the gap between the two is where the work sits.
| The object | What it is here | What it decides |
|---|---|---|
| The outcome set | Every continuous path the process could trace from Rs 100/- over the year | Which questions can be asked at all |
| The allowed subsets | Statements built from where the path sits at stated times, closed under negation and countable union | Which questions can carry a number |
| The rule, called P | Fixed by a drift of 8 per cent and a volatility of 20 per cent a year | What each of those numbers is |
Now ask the rule a question: what weight does it give the event that the process finishes above Rs 100/-? The event is a set of paths, so the answer is a number the measure returns, and it comes out of the parameters rather than out of any observation.
| \(S_0,\ S_T\) | the standard process at the start and at the horizon, in rupees |
| \(K\) | the level the event is about, here Rs 100/- |
| \(\mu\) | the drift the rule carries, as a decimal |
| \(\sigma\) | the volatility, as a decimal, the same under either rule |
| \(T\) | the horizon in years, here 1.0 |
| \(N\) | the standard normal distribution function |
Put the numbers in. The variance rate is 0.04, so half of it is 0.02 and the drift of the logarithm is 0.06 exactly. The level is the starting value, so the logarithm term is zero. The calculation leaves 0.06 divided by 0.20, or 0.30, and the standard normal distribution function at 0.30 is 0.617911. The rule called P gives the event that the standard process finishes above Rs 100/- a weight of 0.617911, and every digit of that came from the four invented parameters rather than from anything observed.
The same rule gives an average finish of Rs 108.33/- and a middle finish of Rs 106.18/-. Average and middle are not equal, and the gap of Rs 2.15/- is entirely the half variance correction showing up in rupees. The gap is covered separately, under the half variance correction. Here it stands as a fact about the rule rather than as something to unpick.
Why can one outcome set carry two different rules at once?
Risk-Neutral Probability
Everything above was built with a single rule. Now put a second one on the same outcome set. The first is the physical measureThe rule describing what the model itself says is likely, written with the letter P., written P, and it carries the drift of 8 per cent. The second is the risk-neutral measureThe second rule placed on the same outcome set, written with the letter Q, used for pricing rather than for describing what is likely., written Q, and it carries 5 per cent, the continuously compounded rate this subject holds fixed.
Where Q comes from, why anybody would want it, and how it is constructed are all set out under physical and risk-neutral measures. Here it is an object and nothing more: a second rule, satisfying the same three conditions, sitting on the same outcome set and the same collection of allowed subsets. A second rule of that kind is all the argument needs.
Run the formula again with 0.05 in place of 0.08. The drift of the logarithm becomes 0.03, dividing by 0.20 gives 0.15, and the distribution function at 0.15 is 0.559618. The identical event now carries 0.559618 instead of 0.617911, a difference of 0.058293, and not one path was added to the outcome set or removed from it. The set of things that could happen is byte for byte what it was. Only the weights moved.
The move is easier to believe seen as a shape rather than as two numbers. Each rule spreads a total weight of one across the same range of finishing values, and the second spreads it slightly further to the left. The whole difference between the two pictures is a shift of the drift of the logarithm, from 0.06 down to 0.03, exactly the difference between the drift and the rate. Nothing about the width changes. The volatility is the same under both.
The same event carries 0.617911 under one rule and 0.559618 under another. Which of the two, if either, is the chance the event happens?
The rule is about to change from a drift of 8 per cent to one of 5 per cent. What happens to the set of paths the process could take?
Move the rule. Watch the outcome set refuse to move.
One control: the drift the rule carries, from 0 to 12 per cent a year. The curve over the finishing value redraws, the shaded weight above Rs 100/- moves with it, and the panel underneath, the outcome set itself, is drawn once and never touched again. The volatility stays at 20 per cent, the horizon stays at one year, and the starting value stays at Rs 100/-.
How is the space written out for a teaching example?
How to Define a Probability Space for a Financial Teaching Example
Here is the practical residue of everything above, and it fits on a card. Five lines, in this order, and a teaching example that cannot fill all five has a hole in it rather than a shorthand.
Most incomplete teaching examples are missing line 1 and line 4, and they are missing them together. A worked example that states a starting value, a drift and a volatility has described the rule and nothing else. Ask it whether the process ever touched Rs 95/- and it cannot answer, not because the arithmetic is hard but because it never said what a single outcome is or how long the year runs. Parameters describe the scale. Parameters do not say what is being weighed.
Line 5 is not decoration either. A computed illustration and an observed figure look identical once written down, and the label on the card is the only thing that separates them.
A teaching example states a starting value, a drift and a volatility, and nothing else. What is missing from its probability space?
What is actually done when somebody hands over a probability?
The three part construction is abstract, and the use it gets put to is not. Any time a number arrives describing how likely something is, whether it comes out of a model, a report or a product illustration, the three part construction supplies three questions to ask in order, and they take about ten seconds.
First, what is the outcome set? If nobody can say what a single outcome is, the number is decorative. A weight has to be a weight on something, and there is a real difference between a statement about a year of readings and a statement about one reading at the end of it.
Second, which rule produced it? Here two rules on one outcome set produced 0.617911 and 0.559618 for the identical event. Both are correct arithmetic. The two numbers answer different questions, and only the number naming its own rule is usable. A figure handed over without its rule attached is a reading without a scale.
Third, is this a likelihood or a weight? A number can satisfy all three conditions of a probability and still make no claim whatever about how often anything happens. That is not a defect. The gap between a weight and a likelihood is what makes the second rule useful in the first place, and it is the single most common thing practitioners get wrong about this subject.
The household version costs nothing. When a shopkeeper quotes the weight of a bag, the number is accepted because the scale it came off is visible and the bag it was on is known. When neither is on show, the bag gets weighed again. Numbers about likelihood deserve the same reflex and rarely get it.
The error that gets made, and what it costs
Reading the second rule as a forecast. The number 0.559618 attaches to exactly the same event as 0.617911, and a reader who meets it first will naturally take it as the chance the standard process finishes above Rs 100/-. The number is not a forecast. The number 0.559618 is the weight under which discounted values behave consistently, and the model's own statement about what is likely is the other number, 5.83 percentage points away.
The cost is a figure presented to somebody as a likelihood when the model never claimed it was one, and the arithmetic will never flag the error. Every step of the arithmetic is valid. The mistake happened at the point the word chance was typed next to the number, before any calculation began and after every calculation has finished checking itself.
References
| Source | Document | Where |
|---|---|---|
| arXiv, Quantitative Finance | Preprints on probability spaces and measure change in derivative pricing | arxiv.org |
| Social Science Research Network | Working papers on measure-theoretic foundations for pricing models | ssrn.com |
The standard process and its locked path are invented.
Educational material. Not advice on any investment, tax, budget or market position.
