Sigma-Algebra: Which Events Are Measurable, and Why It Matters
A sigma-algebra is the collection of subsets that a probability may be attached to. Three properties define it: it holds the whole outcome set, it holds the opposite of everything it holds, and it holds the joining of any listable sequence of its members. Beneath that formal job sits a second one, and it is the one this subject leans on: the collection states what is known.
The three properties are not a matter of taste, and they are not a list somebody chose to be careful. The list is the smallest set of closures that let the word probability behave. If it is possible to ask whether an event happened, it must be possible to ask whether it did not. If it is possible to ask about each of a listable sequence of events, it must be possible to ask whether at least one of them happened. With either closure removed, ordinary reasoning about chance breaks down halfway through a sentence.
The probability space named three objects and opened only the first. The outcome set was settled there. The measure comes later. Between the outcome set and the measure sits the collection, the object that decides which questions the measure will ever be handed. Fixing the collection fixes, in advance and independently of any number, the entire set of questions the model is capable of answering.
What three properties define a sigma-algebra?
Begin with the outcome set. Written with a capital omega, the outcome set lists everything that could happen. Any subset of it is a candidate event. The natural guess is that probability is a rule handing a number to every subset. It is not. A sigma-algebraThe collection of subsets that may be assigned a probability. is a named collection of subsets, and probability is a rule handing a number only to the members of that collection.
Here is the everyday version. A weighing scale on a grain counter reads to the nearest kilogram. Ask it whether the sack is heavier than three kilograms and it answers. Ask it whether the sack is exactly 3.4 kilograms and it has nothing to say. The silence is not shyness about the answer and it is not a small probability. The instrument simply has no such reading, and no amount of staring at the dial produces one. A collection of subsets is an instrument, and its members are the readings it can take.
| \(\Omega\) | the outcome set, listing every way the world could turn out |
| \(\mathcal{F}\) | the collection under discussion, the object being defined here |
| \(A\) | one subset of the outcome set, that is, one candidate event |
| \(A^{c}\) | the complement of \(A\), meaning every outcome not in \(A\) |
| \(\bigcup\) | the union, the outcomes lying in at least one of the listed sets |
Two useful things follow immediately and cost nothing extra. Because the whole outcome set is a member and complements are members, the empty set is a member too. Because a finite list can be padded out to an endless one by repeating the empty set, closure under a countable unionThe set of outcomes lying in at least one member of an endless but listable collection. already gives closure under a union of two or of ten. Nothing has to be added for the finite case.
Intersection comes free as well, and intersection is the closure most often expected as a fourth property. Intersection is not one. Where complements and unions can be taken, intersections can be taken too. The outcomes lying in every one of the sets are exactly the outcomes not lying outside at least one of them. De Morgan's identity does the work. The collection is closed under and, or and not with only three requirements stated, and that is why the list is three long rather than four or five.
| \(\bigcap\) | the intersection, the outcomes lying in every one of the listed sets |
| \(A_{i}^{\,c}\) | the complement of the \(i\)th set in the sequence |
A collection contains an event but does not contain everything outside that event. Is it a sigma-algebra?
Why can a probability not simply be given to every subset?
On a finite outcome set it can. Take four paths and the power setThe collection of every subset, which is the finest collection there is. has sixteen members, every one of them can carry a number, and the three properties hold without effort. So on any lattice that can be drawn, the question does not bite. The reason the collection is named separately is that the outcome sets that matter are not finite.
A continuous quantity lives on the line. There the demands on an assignment of probability can be written down. First, every subset should get a number. Second, length does not care where anybody stands, so sliding a subset along the line should not change its number. Third, the pieces should add: if a set is cut into a listable sequence of parts that overlap nowhere, the numbers of the parts should sum to the number of the whole, and the unit interval should read one.
| \(P\) | the physical measure, the rule handing a number to a set |
| \(A_{i}\) | pieces that overlap nowhere, listed one after another |
| \(A + x\) | the set \(A\) slid bodily along the line by an amount \(x\) |
| \([0,1]\) | the unit interval, taken as the reference set of size one |
The closing clause of that summary is a theorem, not a caution. There exists a subset of the line to which no consistent number can be given under those demands, a result due to Vitali in 1905. The collection of allowed subsets exists to exclude that subset, not to encourage general carefulness. Once such a set is known to exist, naming a collection stops looking fussy and starts looking forced, exactly like the three closures above.
The construction of that set picks one representative from each of uncountably many groups at once, and watching it done teaches the construction and nothing new about the collection. Only the existence of such a set does any work. Because one bad subset exists, the collection of allowed subsets has to be part of the model rather than an afterthought about it.
Why is the collection restricted at all, rather than simply taking every subset of the outcome set?
What has any of this to do with knowing rather than with size?
Everything after this point in the subject uses a second reading of the same object, and the second reading is not a metaphor laid over the first. The second reading is what the definition already says, read from a different angle. A set is a member of the collection exactly when an observer standing at that point could answer yes or no to the question did this happen. The collection is therefore a list of answerable questions, and a list of answerable questions is a statement about what is known.
Consider a lift in an eight storey building. The indicator reads the floor and nothing else. Asked whether the lift is on the third floor, it answers. Asked whether the lift is on the third floor or the fourth, it answers. One floor or another floor is a union of two readings the indicator already has. Asked whether the lift is in the corner room, it cannot answer, ever, no matter how long the wait. The sets the indicator can resolve are the floors and their unions, and that collection is the exact content of what the indicator knows.
A set that belongs to the collection is called measurableBelonging to the collection, so capable of carrying a probability at all. with respect to it. In this subject the collection carried by a moment in time is written as a script F with that time as a subscript, and it is read aloud as the information available at that time rather than as an inventory of sets. Both readings describe the same object. The formal one names what a probability may attach to and the informational one names what somebody standing there could answer.
How does a coarser collection say that less is known?
A coarserContaining fewer sets, so drawing fewer distinctions and representing less knowledge. collection is one with fewer members. Every collection in this subject sits somewhere between the two extremes. At one end sits the collection holding only the empty set and everything, a collection that distinguishes nothing at all. At the other end, on a finite outcome set, sits the power set. The power set distinguishes everything. Coarsening means moving toward the first, and moving toward the first means giving up distinctions.
Here is the part that trips people. Coarsening does not change any answer. Coarsening removes questions. The measure never depended on which other sets happened to be in the collection alongside a given event, so an event that survives the coarsening keeps precisely the number it had. The scale that reads to the nearest kilogram and the scale that reads to the nearest gram agree exactly on whether the sack is over three kilograms. The two scales differ only in what else can be asked.
The number 0.553908 in that figure is the weight the two step lattice puts on an up move under the risk-neutral measure Q, and 0.446092 is the weight it puts on a down move. Both follow from the four locked parameters rather than from any reading of a market. The same answer appears on both sides of the coarsening, and neither side of the figure has more or less probability in it than the other. One side simply has fewer questions.
A coarser collection is adopted. Does an event that was already a member of both collections get a different probability?
What do the three properties look like on four paths?
Everything above becomes countable on a small lattice. Take the standard process, an invented two step model that starts at Rs 100/-, and let it move twice over one year. The two step lattice locked for this subject multiplies by 1.151910 on an up move and by 0.868123 on a down move, and those two factors are exact reciprocals, so an up followed by a down returns to precisely where it started. Four paths result, and the outcome set has four members.
| Path | After the first move | At the horizon | Written as |
|---|---|---|---|
| up, then up | Rs 115.19/- | Rs 132.69/- | the first path |
| up, then down | Rs 115.19/- | Rs 100.00/- | the second path |
| down, then up | Rs 86.81/- | Rs 100.00/- | the third path |
| down, then down | Rs 86.81/- | Rs 75.36/- | the fourth path |
Counting the collection at three moments makes this concrete. Before anything happens, the outcome set has not been divided at all, so the only members are the empty set and all four paths together: two sets. After the first move the two paths that began up can be told from the two that began down. The split leaves the outcome set in two pieces, and the collection has four members. After both moves every path is distinguishable, and the collection is the whole power set of four paths: sixteen members.
| \(|\mathcal{F}|\) | the number of sets in the collection |
| \(k\) | the number of pieces in the split of the outcome set the collection is built on |
| \(\mathcal{F}_{t}\) | the collection carried at time \(t\), read as the information available then |
The event that the process finishes at exactly Rs 100/- holds on the second path and on the third path, one path from each of the two halfway pieces. The event is therefore not a union of those pieces, so it is one of the sixteen and not one of the four. Halfway through, that event cannot be given a probability, and the reason is not that the chance is small but that the collection at that moment cannot express the set at all.
Working with the halfway collection, can a probability be assigned to the event that the process finishes at exactly Rs 100/-?
Four paths. Before the control below is moved, how many sets can be assigned a probability at the halfway stage, when only the first move is known?
Watch the paths group, and count what can be asked
One control, three positions. Nothing about likelihood moves. The control changes which questions exist rather than which answers they get. Before anything happens the four paths form one piece and the collection holds 2 sets. After the first move they form two pairs, up first and down first, and the collection holds 4 sets. After both moves they form four pieces and the collection holds 16 sets. At the halfway stage the event of finishing at Rs 100/- takes one path from each pair, so it is not among the 4.
After the first move the four paths fall into two pieces, so the collection holds 4 sets, and the event of finishing at exactly Rs 100/- is not one of them.
A quantity takes only two different values across the whole outcome set. How many sets are in the collection it generates?
What is the collection generated by a quantity?
Most collections in practice are not written down set by set. A collection arrives attached to a quantity: somebody says the only thing observed is the reading of this instrument, and the collection follows from that. The generated collectionThe smallest collection in which a stated quantity can be read. is built by a procedure with a stopping rule, and both the procedure and the stopping rule matter.
| \(X\) | a quantity read off an outcome, here a general one rather than the standard process |
| \(\omega\) | one outcome, one member of the outcome set |
| \(B\) | a range of values on the line, such as an interval |
| \(\mathcal{B}(\mathbb{R})\) | the Borel collection, the smallest one on the line holding every interval, named for Borel |
| \(X^{-1}(B)\) | the outcomes whose reading lands inside \(B\) |
The stopping rule is where the meaning sits. Taking the smallest collection is what makes the phrase this is all that is observed precise. Any larger collection would credit the observer with distinctions the instrument never made, and a model built on it would answer questions the observer had no business answering. The word smallest in that definition is not tidiness, it is the honesty condition of the whole apparatus.
Two different quantities can generate the very same collection, and the four path lattice shows it plainly. Take the event that the process finishes above Rs 100/-, the region where the at-the-money contract at Rs 100/- would be in the money, and the event that it finishes above Rs 110/-, the region where the out-of-the-money contract at Rs 110/- would be. On this lattice the terminal values are Rs 132.69/-, Rs 100.00/-, Rs 100.00/- and Rs 75.36/-, so both events pick out the first path and only the first path. Two different contracts, identical generated collections of four sets each, and therefore identical information. The lattice is simply too coarse to tell the two apart.
Why does any of this change what can be said about a price?
One consequence pays for all of this. The collection is fixed before any probability is named, and it does not move when the probability changes. The physical measure P and the risk-neutral measure Q sit on the same four paths and on the same sixteen sets. Swapping one for the other changes every number and changes not one membership. A change of measure alters the answers; only a change of collection alters the questions.
So when somebody reports a number about a price, there are two separate things to check and they fail in different ways. The first is whether the arithmetic is right under the measure being used. The second check comes first in logic and is skipped far more often. Ask whether the event being priced is even a member of the collection the information supports. A number can survive the first check perfectly and fail the second completely.
The two numbers in that figure are the weights the locked two step lattice puts on the remaining move under the risk-neutral measure Q, one weight for each halfway block. Their shape matters rather than their size. One value per block is the most detailed thing the halfway collection can hold, and offering that pair is a complete and honest reply where a single number would have been a fabrication.
Somebody asks for the probability that the process finishes at exactly Rs 100/-, using only what is known halfway. What is the honest reply?
The physical measure P is swapped for the risk-neutral measure Q on the same four paths. How many sets can now be assigned a probability at the end?
Who reaches for this, and what do they actually do with it?
The arithmetic is the part a machine already got right, so somebody reviewing a model does not usually check it first. A reviewer checks the pairing between a reported quantity and the information it claims to rest on. The question they ask, in whatever words their team uses, is which collection this number was computed against, and whether that collection is the one the situation actually supplies.
The check has a clean shape and applies to anything. The outcome set is named. The moment at which the number is supposed to be known is named. The split of the outcome set into pieces that could be told apart is then written down. The question is then whether the event at issue is a union of those pieces. If it is, the number is a statement about that moment. If it is not, the number belongs to some other, finer collection and has been quietly relabelled.
The reason this check is worth its minute is that nothing else will ever flag the failure. A quantity attached to the wrong collection is not out of range, does not break a constraint, does not produce an error and does not look odd. A misattached quantity looks like a number and is read like a number while behaving like an assumption. The same discipline lets somebody say honestly that a question is not answerable yet. Such a sentence is far more useful than a confident figure with nothing underneath it.
The error that gets made, and what it costs
Assuming that anything describable in a clear sentence is an event that can be priced. Finishing at exactly Rs 100/- is plain English, it names a definite set of paths, and at the halfway stage it carries no probability at all. The collection is fixed by what is known, and a perfectly clear description can sit outside it.
The figure 0.494188 above is not a mistake in itself. The figure is correct on the collection that separates all four paths, and every step producing it is valid. The error is the label. Reported as something known halfway, it attaches a quantity to information that could not support it, and from that moment it reads as a number while behaving as an assumption.
The cost is that nothing downstream ever catches it. No step of the arithmetic was wrong at any point, so there is no failed check, no range violation and no warning. The mistake was made before the arithmetic began, in deciding which collection the question belonged to, and that decision leaves no trace in the output.
References
| Source | Document | Where |
|---|---|---|
| arXiv, Quantitative Finance | Preprints on measure-theoretic probability in derivative pricing | arxiv.org |
| Social Science Research Network | Working papers on probability foundations for pricing theory | ssrn.com |
| Giuseppe Vitali, 1905 | Sul problema della misura dei gruppi di punti di una retta, the construction of a subset of the line that carries no consistent measure | Gamberini e Parmeggiani, Bologna |
| Steven E. Shreve | Stochastic Calculus for Finance II: Continuous-Time Models, the treatment of a collection as the information available at a moment | Springer |
The standard process and the two step lattice are invented.
Educational material. Not advice on any investment, tax, budget or market position.
