Filtration: How Information Accumulates Through Time
A filtration is a list of sigma-algebras, one entry for each time, carrying one condition: every entry contains all the earlier ones. The nesting is the whole content, and it says that information accumulates and is never lost. An information set is a single entry in that list. The filtration is the entire sequence, and nearly every statement in this subject is made relative to one.
One item of vocabulary is already in place. A sigma-algebraA collection of subsets of the outcome set, closed under the operations that let it stand for a settled body of knowledge. is a statement of what is known: the collection of questions that can be answered at a given moment, written as a collection of sets. The sigma-algebra is settled vocabulary here. A filtration needs one of them at every moment on the clock, together with a statement about how the moments relate to each other.
One trap governs everything that follows. Somebody supplies twelve collections, one for each month, and calls the result a filtrationA list of sigma-algebras, one for each time, with each entry containing all the earlier ones.. The result is not a filtration yet. Twelve collections on the same outcome set is just twelve collections, and there is nothing in that object that says anything about time at all. The list could be shuffled and nothing would break. The word filtration adds exactly one requirement to a list of collections, and that requirement is the entire subject of this guide.
What is a filtration, and what makes it more than a list of collections?
The requirement is that each entry sits inside the next. Formally the entry at an earlier time is a subset of the entry at a later time, for every pair of times that can be named. Nothing else is added. No probability is involved, no process is involved yet, and no statement about how fast the collections grow appears anywhere. Just containment, and containment in one direction only.
The everyday version is a notebook that can be written in but from which no sheet may be torn out. Opened in April, it shows everything written in March, plus whatever April added. Opened in December, March is still there. The notebook has a state at every moment, those states are ordered, and the ordering has a direction because sheets only ever go in. A ring binder, where somebody could pull March out in December, is a perfectly ordinary object and it is not a notebook of this kind. The direction of the containment is the whole difference, and a list of collections with no direction in it is not a filtration however carefully it was indexed.
| \(\mathbb{F}\) | the filtration itself, meaning the whole list rather than any one of its members |
| \(\mathcal{F}_t\) | the entry at time \(t\), a sigma-algebra on the outcome set, standing for what is known by then |
| \(\mathcal{F}_s\) | the entry at any earlier time \(s\), which the condition places inside the later one |
| \(\subseteq\) | set containment, read here as every member of the left collection is also a member of the right one |
| \(\mathcal{F}\) | the master collection on the outcome set, which every entry sits inside and which nothing exceeds |
| \([0,T]\) | the index set, here every instant from today to the horizon of one year |
Notice how little that says. The condition does not say the collections get bigger. The condition says they do not get smaller. A list in which every entry equals the one before it satisfies the condition perfectly, so the two claims are different and the difference matters. Somebody who never opens the notebook still has a notebook. The degenerate case is a real filtration, and this subject uses it deliberately whenever a model needs a party who learns nothing.
A model gives the same collection of sets at every one of twelve months, never adding anything. Is that list a filtration?
What does the nesting condition actually forbid?
The nesting condition forbids one thing: a set that is a member of an earlier entry and is not a member of a later one. Most readers arrive expecting the condition to be about growth, and the condition is about loss, so the only prohibition in the definition is worth turning over slowly.
Being a member of the entry at a given time means the corresponding statement can be settled by that time. So a set leaving the list would mean a question that used to have an answer no longer having one. NestingThe requirement that every earlier entry sits inside every later one, so no set ever leaves the list. rules that out and rules out nothing else. Two entries can be identical. An entry can leap from two sets to sixteen in one instant. An entry can grow by exactly one set over a whole year. Each of those is allowed, and the word guards against none of them.
| \(A\) | one set of outcomes, standing for one statement about how things turn out |
| \(A \in \mathcal{F}_s\) | that statement can be settled by the earlier time \(s\) |
| \(A \notin \mathcal{F}_t\) | that statement cannot be settled by the later time \(t\), which is the forbidden case |
| \(\exists\) | there exists, read here as the condition fails the moment a single such pair can be found |
| \(s < t\) | any two times in order, with \(s\) the earlier and \(t\) the later |
The condition quietly supplies a second thing. Two sets are in every entry of every filtration whatever else is going on: the empty set and the whole outcome setThe set of every way the situation could turn out, with one member for each complete outcome.. A sigma-algebra is required to contain both, so no entry can be smaller than two members. The two compulsory members are why the counts on the worked instance below start at two rather than at zero. A reader who expects an empty starting entry will find the arithmetic confusing for no good reason.
On the four path lattice worked below, the halfway entry holds 4 sets and the starting entry holds 2. What must be true of those 2?
Is information accumulating a fact, or a choice somebody made?
Accumulation is a choice. The sentence information accumulates sounds like an observation about the world and is nothing of the sort, so the part of the subject that gets waved through fastest deserves better. Accumulation is a property somebody decided to impose on a model when they decided to call the list a filtration.
Models where knowledge is lost are ordinary objects. A lift knows which floor it is on and has forgotten every floor it visited on the way there. A queue that keeps only its current length has forgotten who joined it and in what order. A weighing scale that shows the current reading and stores nothing knows less at four o'clock about the morning than it did at nine. Each of those is a perfectly coherent description of a system evolving through time. Not one of them is a filtration, and calling any of them one would be an error of description rather than a fault in the system.
AccumulationInformation only ever being added to and never removed, which is a modelling assumption rather than an observation. is therefore an assumption the modeller makes. Naming it as one is the difference between a reader who can question a model and a reader who cannot. The reason for assuming it nearly always is practical rather than deep. The machinery built on top, the conditional averages and the fair game conditions and the stopping rules, all take the containment as given, and every one of them would have to be restated for a list that lost sets. Convenience is a legitimate reason to assume something. Pretending the assumption was never made is not.
Is a model in which something known at one time is later forgotten incoherent?
What is one entry called, and why is the whole list a different object?
The Information Set, and Why It Is Only One Entry
One entry in the list has a name of its own. An information setOne entry in the list, meaning the collection of statements settled at one particular time. is the collection standing at one particular time, and it answers exactly one question: what is known now. An information set is a photograph. The photograph has no direction in it, no before and no after, and it carries no claim about anything except the moment it belongs to.
Most confusion in this area comes from sources that use the two words as though they were interchangeable, and they are not interchangeable in either direction. Handing somebody the entry at the halfway stage tells them nothing about whether the model forgets things afterwards. Handing them the whole list tells them that and everything else. The accumulation lives in the list and nowhere else, so an information set on its own cannot state it, cannot deny it, and cannot even raise the question.
Filtration vs Information Set, Side by Side
Set against each other on four points, the difference stops being a matter of wording. The four points are the object itself, what the object answers, what the object can forbid, and what breaks when one is used in place of the other.
| Point of difference | An information set | A filtration |
|---|---|---|
| The object itself | One sigma-algebra, standing alone | A whole indexed list of them, one for each time |
| What it answers | What is known at this one moment | What is known at every moment, and how those moments relate |
| What it can forbid | Nothing about time, because there is no second moment to compare with | Forgetting, which is the only prohibition the definition carries |
| What is lost by swapping it in | Used in place of a filtration, the accumulation assumption disappears silently | Used in place of an information set, nothing is lost, only more is said |
| The one line version | A photograph of one moment | The ordered album, with the rule that no photograph is ever removed |
An information set is what, relative to a filtration?
What does the whole list look like written out on four paths?
Small enough to check by hand is the only size worth learning this at, so take the two step lattice locked for this subject on the invented standard processThe single invented traded quantity this subject runs on throughout, written S with a time subscript., written S with a time subscript and starting at Rs 100/-. Two moves, an up factor of 1.151910 and a down factor of 0.868123 which is its exact reciprocal, and four paths result. Because an up followed by a down returns precisely to the start, the terminal values are Rs 132.69/-, Rs 100.00/-, Rs 100.00/- and Rs 75.36/-.
| Path | Two moves | Value halfway | Value at the horizon |
|---|---|---|---|
| 1 | up, then up | Rs 115.19/- | Rs 132.69/- |
| 2 | up, then down | Rs 115.19/- | Rs 100.00/- |
| 3 | down, then up | Rs 86.81/- | Rs 100.00/- |
| 4 | down, then down | Rs 86.81/- | Rs 75.36/- |
Now write the three entries out completely. At the start nothing has happened, so the four paths cannot be told apart at all, and the entry contains only the two sets every entry must contain: the empty set and all four paths together. Two sets. At the halfway stage the first move has landed, so paths 1 and 2 can be separated from paths 3 and 4, and the entry contains the empty set, the pair {1, 2}, the pair {3, 4} and all four together. Four sets. At the horizon every path is distinguishable and the entry contains all sixteen subsets.
| \(\mathcal{P}_t\) | the split of the four paths into groups that cannot be told apart at time \(t\) |
| \(\preccurlyeq\) | is refined by, meaning every group of the earlier split is a union of groups of the later one |
| \(\sigma(\mathcal{P}_t)\) | the entry generated by that split, being every set that can be built as a union of its groups |
| \(k_t\) | the number of groups in the split at time \(t\), here one, then two, then four |
The containment is the point, and at this size it can be checked set by set rather than taken on trust. Consider the two sets in the starting entry. The empty set is in the halfway entry. All four paths together is in the halfway entry. Both accounted for, so the first containment holds. Now the four sets in the halfway entry. The empty set, {1, 2}, {3, 4} and the whole outcome set are all subsets of the four paths, and the horizon entry contains every subset there is, so all four are in it. Every set in an earlier entry is in every later one, checked here individually rather than argued, and that check is what the word filtration is claiming on every model that uses it.
The time marker in the control below is about to advance. Can a set that was in the earlier collection be missing from the later one?
Advance the clock and watch the earlier entry stay put
Move the marker across the three stages of the four path lattice. The paths regroup, the two move outcomes are revealed one at a time, and the sets of the current entry are drawn below with the ones carried in from the entry before shaded. A filtration records what is known, so nothing about likelihood appears anywhere.
What is the filtration generated by watching a process?
Everything so far has taken the list as given. In practice the list nearly always comes from somewhere, and in this subject it almost always comes from the process itself. The generated filtrationThe one built from watching a process and nothing else, holding at each time exactly what the readings up to then reveal. is the one whose entry at each time holds exactly what watching the process up to that time would reveal, and not one set more.
| \(\mathcal{F}^{S}_{t}\) | the entry at time \(t\) of the filtration generated by the standard process |
| \(S_u\) | the standard process at time \(u\), one reading for each instant up to and including \(t\) |
| \(\sigma(\cdot)\) | the smallest sigma-algebra making everything inside the bracket settled |
| \(0 \le u \le t\) | every instant from the start up to now, and nothing beyond now |
The nesting comes free here rather than being imposed. The readings up to month six are a subset of the readings up to month seven, so anything settled by the first collection of readings is settled by the second. A generated filtration cannot fail the nesting condition, and the guarantee is a large part of why it is the default object in this subject. It is also the smallest list the process fits into, and that smallness is doing real work: it is the honest choice, because any larger list is asserting that somebody knows something beyond the readings, and asserting that without saying so is where a great many wrong arguments begin.
Take the twelve step path locked for this subject: Rs 97.64/-, Rs 107.63/-, Rs 100.35/-, Rs 100.27/-, Rs 101.35/-, Rs 111.08/-, Rs 103.56/-, Rs 101.12/-, Rs 93.74/-, Rs 96.41/-, Rs 102.06/- and Rs 106.18/- across the twelve months. At month six the generated entry pins the first six of those readings exactly and says nothing whatever about the last six. The silence about the last six is not a limitation of the drawing. The silence is the definition doing its job.
One technical note that a reader will meet in any serious treatment and should not be surprised by. The generated list is usually replaced by the augmented filtrationThe generated one with the sets of probability zero added in, which is the usual technical convention., meaning the same list with all the sets of probability zero added into every entry. The replacement is a convenience and not a change of substance: it removes a class of irritating exceptions from later proofs and it changes nothing noticeable on a four path lattice, where no set has probability zero to begin with.
For the filtration generated by the standard process, what does the entry hold at month six?
How is a filtration pinned down in a teaching model?
How to Specify a Filtration in a Teaching Model, in Four Steps
Reading definitions is not the same as being able to write one down, and the gap shows up the first time somebody asks what the filtration in a given model actually is. Four steps close it, and everybody already does the first three.
- Name the outcome setSay what a single complete outcome is, and how many there are if the number is finite. On the two step lattice an outcome is one of the four paths, so the outcome set has four members and can be written out in full. Get this wrong and every later step inherits the error.Check: can one member of the outcome set be listed out loud?
- Name what is observed, and whenSay which quantities are watched and at which times. Here the standard process is watched at the halfway stage and at the horizon, and nothing else is watched at all. This step is where the difference between a coarse and a fine list is actually decided.Check: for each time, what would a watcher have seen by then?
- Say that each entry is generated by everything observed up to thenState plainly that the entry at each time is the smallest collection making every observation up to that time settled. This is the step that produces the nesting for free rather than by decree.Check: does the entry at each time use the readings up to then, and only those?
- State whether anything besides the observed quantities is knownSay explicitly whether the party in the model knows anything the observations do not carry: a parameter, a hidden state, the outcome of a later coin. Usually the answer is no, and saying so out loud is what makes it an assumption instead of an accident.Check: is there a sentence in the model that answers this, or was it assumed silently?
The fourth step is the one that gets left out, and leaving it out is how an argument quietly comes to assume information nobody said was available. The tell is usually a result that looks slightly too clean: a rule that never has a bad period, a value that can be computed at a moment when the ingredient for it has not arrived. An unstated extra fact changes what every entry contains and changes nothing about how the argument reads, so when a derivation looks surprising the fourth step is the first place to look.
Which step of specifying a filtration is usually left out?
Why must every statement in this subject name the filtration it is made against?
Because the statements are relative and they do not look relative. The point is worth stating without hedging: a large share of the vocabulary in this subject is not a property of a quantity at all, but a property of a quantity together with a list. Adapted is such a term, and so is the fair game condition that gives a martingale its name. Both are settled vocabulary here, covered separately. Neither term is a complete claim on its own.
| \(X\) | one quantity, here a single reading of the standard process |
| \(\mathcal{F}_t\) | the entry at time \(t\) of the finer list, the one generated by the process |
| \(\mathcal{G}_t\) | the entry at the same time of a coarser list, which records less |
| \(\nRightarrow\) | does not imply, meaning the property does not survive being carried to the smaller collection |
The error that gets made, and what it costs
The error is short enough to fit in one sentence: saying that a quantity is adapted, or that a process is a martingale, without naming the filtration. The omission is not sloppiness at the edges. Both statements change their truth value when the list changes, so the unqualified version has nothing to be true or false about.
Work it on the published path. The entry at month six of the list generated by the process holds every reading up to month six, so the month six reading of Rs 111.08/- is settled with respect to that entry. Now take a coarser filtrationOne that records less at each time, so fewer questions can be answered at every moment. that records only whether the process stood above Rs 100/- at each date and nothing else. Its entry at month six holds six marks: below, above, above, above, above, above. The six marks are consistent with an enormous number of different paths, and the value Rs 111.08/- is only one of them. The same reading is adapted to the first list and is not adapted to the second, so the bare sentence the reading is adapted is incomplete rather than merely imprecise.
The cost is a claim that is true under one reading and false under another, asserted without either reading being named. The claim survives review because it looks like a finished sentence, it survives recomputation because there is no arithmetic in it to be wrong, and it is caught only by somebody asking which list. Asking which list is the entire defence, and it costs one line.
Somebody says a process is adapted. Which part of the claim is missing?
Who reaches for this, and what do they actually do with it?
Three kinds of reader use this object daily and none of them thinks about it as an abstraction. The first is somebody writing a model down. For them the filtration is the four step routine above, done once, and the payoff is that every later claim in the write up has something concrete to be relative to. The work is entirely in step four, and the habit worth building is to write the sentence nothing beyond the observed readings is known even when it feels too obvious to write.
The second is somebody checking work that already exists. Their tool is a single question applied to every claim in the document: relative to what. A relative statementA claim whose truth depends on which filtration it is made against, so that dropping the filtration removes the claim. with its list missing is the most common defect in this kind of writing, and because it needs no arithmetic it is cheap to find. Read the claims, ask which list, and mark every one that has no answer. A document where the list is named once at the top and then genuinely used throughout is a different quality of document from one where it is never named at all.
The third is somebody reading a result in a paper and trying to decide whether it applies to their situation. The question there is not whether the result is correct but whether the list in the paper matches the list they are in. A fair game condition proved against the list generated by a process does not automatically hold against a list where somebody also observes a second quantity, and it does not automatically hold against a coarser list either. Matching the filtration is what turns a correct result into an applicable one, and it is a check, not a formality.
The everyday version of that third check is close to home. A rule written for somebody who can see the full reading does not transfer to somebody who can only see whether the reading crossed a line, and the two are genuinely different situations rather than the same situation described more or less carefully. Anyone who has tried to follow instructions written for a better instrument than the one they have knows the feeling exactly.
Which subsets may be assigned a probability at all is a different question from how those collections accumulate through time, and is treated separately in this subject area. Adapted and predictable are settled vocabulary here, treated in their own place. Integration against a process is a separate subject, as is what any contract pays. No jurisdiction is named because none applies: the definition of a filtration is fully universal and no authority anywhere sets it.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for information structures and filtrations in continuous time modelling | arxiv.org |
| Social Science Research Network | Working paper repository for the same material, including work on information assumptions in model specification | ssrn.com |
| Hull, Shreve and Wilmott | Standard texts on stochastic calculus, setting out filtrations and the vocabulary built on them |
The standard process, its four parameters, its published twelve step path and the two step lattice built on it are invented.
Educational material. Not advice on any investment, tax, budget or market position.
