Submartingale and Supermartingale: Drifting Up and Drifting Down
A submartingale has a conditional average future value at least as large as its present value, so it drifts up or stands still. A supermartingale has one at most as large, so it drifts down or stands still. The names look reversed because each describes where the process sits relative to its own averages, not which way it travels. A process satisfying both at once is a martingale.
Both properties are the martingale definition with one change: the equality is relaxed to an inequality, and the two directions of relaxation give the two names. Every consequence below follows from that single alteration. One of those consequences is that the weaker statements can be established in situations where equality cannot, and that is why a subject built on fair games spends so much of its time on processes that are not fair.
What does each of the two properties actually say?
Take the martingale definition and read the third condition, the one that sets the conditional averageThe average taken over the outcomes still possible given everything known at the earlier moment. of a later value equal to the present value. Now weaken it in one direction. If the conditional average is allowed to be at least the present value, the process is a submartingale. If it is allowed to be at most the present value, it is a supermartingale. Nothing else changes. The adaptedness requirement stays, the requirement that the average exists stays, and only that one equals sign becomes an inequality sign.
Each of the two properties is one inequality, in one direction, at every pair of times, and the whole of the rest of the definition is unaltered. That is worth stating flatly because most of the trouble readers have with these two words comes from imagining that something more elaborate is going on. Nothing more elaborate is going on. There are three possible relationships between a conditional average and a present value, and the three names simply cover all three.
| \(X_t\) | a general process at time \(t\), adapted to the information by construction |
| \(X_s\) | the same process at the earlier time \(s\), where \(s \le t\) |
| \(\mathcal{F}_s\) | everything known at the earlier time \(s\) |
| \(\mathbb{E}[\cdot\mid\mathcal{F}_s]\) | the average over what is still possible given what is known at \(s\) |
| \(\ge,\ \le\) | at least, and at most, which is the only change from the martingale definition |
Both statements must hold for every pair of times, not merely at the horizon. The inequality has to hold at a pair of times to count, so a process that averages above its present value over a year but dips below it over the first month is not a submartingale. The requirement is easy to skate over and it matters. Holding at every pair of times is what makes the property something that can be carried forward step by step, rather than a single claim about a single date.
Which inequality belongs to the submartingale?
Why does the name that sounds like rising belong to the one that falls?
Almost every reader meets these two words and gets them the wrong way round, and the reason is that both names were borrowed from somewhere else. Neither name was chosen to describe the direction of travel. Both were chosen to match a pair of names that already existed for surfaces, and once that pair is known, the two words stop fighting.
Supermartingale, and Where the Name Comes From
Consider a smooth hill rather than a process. At a point on it, a small circle traced around that point gives a height at every step, and those heights have an average. On a dome, where the ground curves away in every direction, the average of the ring is lower than the height at the point itself. A surface with that property everywhere is called superharmonicA surface whose value at every point is at least the average of the values on a small circle around that point.: the surface sits above its own local averages. On a bowl, the opposite happens, and the surface is called subharmonicA surface whose value at every point is at most the average of the values on a small circle around that point. because it sits below them.
The prefix describes where the object sits relative to its own averages. A surface that sits above its own averages is a surface a wandering point loses height on, so the supermartingale is the one that drifts down. Put a marker on the dome and let it wander without preference in any direction. The dome property says it outright. At each step the average height the marker moves to is lower than the height it left. Repeat that and the marker loses altitude on average. The height of a wandering marker on a superharmonic surface is a supermartingale, and the name was carried over unchanged.
So the word is not a promise about direction at all. The word is a statement of position relative to a set of averages, and the direction of drift is a consequence of that position rather than the position itself. Once the dome is in mind, the direction in which a marker slides can be pictured, and the two words become almost impossible to swap.
Which of the two drifts downward?
What does a one sided property constrain, and what does it leave open?
An inequality is a much weaker instrument than it looks, and the weakness is worth seeing precisely. A submartingale statement fences off one half of the number line for one quantity: the conditional average of the later value may sit anywhere at or above the present value. The conditional average may sit a whisker above. The conditional average may sit fifty rupees above. The conditional average may sit exactly on the present value, since at least includes equal. The property does not name a growth rate, does not name a minimum growth rate, and does not say the average grows by the same amount each period.
The inequality pins one quantity, the conditional average, to one side of one line, and it leaves the size of the gap, the width of the distribution and the behaviour of every individual path completely free. Two of those three are the same freedoms a martingale leaves. The third, the size of the gap, is new, and it is the reason a one sided statement can be true of processes that behave in wildly different ways.
The everyday version runs as follows. A turnstile counter at a gate can only be added to. Whatever happens tomorrow, the average of tomorrow's reading is at least today's reading, and that can be asserted with no model of the crowd at all. The turnstile claim is a submartingale statement and it is completely secure. The turnstile claim is also nearly empty. The claim does not say whether ten people or ten thousand will come through, and it never could. The security and the emptiness are the same fact, seen from two sides.
What is a process that satisfies both at once?
Ask a process to have a conditional average at least its present value and at most its present value at the same time, and only one arrangement survives. The average must be exactly the present value. Exact equality is the martingale condition, arrived at from above and below rather than stated directly. Reaching the condition that way is the cleanest way to hold all three words together at once.
A martingale is not a fourth object standing beside the other two; it is precisely the overlap of the two, the case where both inequalities are tight. This has a practical consequence that shows up constantly. Any statement proved for all submartingales is automatically true of every martingale, and so is any statement proved for all supermartingales. Results about the weaker classes are therefore strictly more general, not weaker in reach, and a great deal of what looks like extra machinery in this subject is really the same result stated for the larger class so that it covers more processes at once.
A process is both a submartingale and a supermartingale. What is it?
Which label does the standard process wear?
All three, depending on what is applied to it and which rule the average is taken under. The standard process, an invented construction, starts at Rs 100/-, carries a drift of 8 per cent a year under the physical measure P, carries a volatility of 20 per cent a year, and sits in a world with a risk-free rate of 5 per cent a year. The horizon is one year. Three quantities are now built from it, and for each one the question is which of the three inequalities holds.
The undiscounted standard process under the model's own rule, at a drift of 8 per cent a year. Which label?
Start with the discounted processThe process multiplied by the price today of a rupee delivered at that later time. under the pricing measure Q. Its conditional average at the horizon is Rs 100.000000/-, exactly where it started. Both inequalities are tight, so it is a martingale, and therefore it is also a submartingale and a supermartingale at the same time. Now keep the same discounted quantity and average it under the model's own rule instead. The average comes to Rs 103.045453/-, or Rs 100/- grown at 3 per cent for a year. Three per cent is the drift of 8 per cent less the rate of 5 per cent exactly. Rs 103.045453/- sits above the present value, so the discounted process is a submartingale and is not a supermartingale. Finally drop the discounting and average under the model's own rule: Rs 108.328707/-, or Rs 108.33/- to two decimals, a submartingale by the wider margin of 8 per cent a year rather than 3.
| The quantity | The rule | Average at one year | Label |
|---|---|---|---|
| Discounted | Pricing measure Q | Rs 100.000000/- | Martingale, so both |
| Discounted | Model's own measure P | Rs 103.045453/- | Submartingale at 3 per cent |
| Undiscounted | Model's own measure P | Rs 108.328707/- | Submartingale at 8 per cent |
One process, one set of paths, three labels, and the process itself did not change between any two rows of that table. What changed was the discounting in one case and the rule for taking the average in the other. Calling a process a submartingale without saying under what is therefore meaningless, and the word on its own is an incomplete sentence.
| \(S_t\) | the standard process at time \(t\), starting at Rs 100/- |
| \(\mu\) | the drift under the model's own measure, 8 per cent a year |
| \(r\) | the risk-free rate, 5 per cent a year, continuously compounded |
| \(e^{-rt}\) | the discount factor to time \(t\), which is 0.951229 at one year |
| \(\mu-r\) | the excess of the drift over the rate, 3 per cent a year exactly |
Notice what does not appear anywhere in that expression. The volatility of 20 per cent a year is absent. Volatility decides how wide the spread of outcomes is, and volatility has no vote at all on which of the three labels applies. A process with a volatility of 2 per cent and a process with a volatility of 200 per cent carry the same label if they carry the same drift and the same discounting. The three labels are statements about the centre and never about the risk.
The drift is about to fall to 2 per cent, below the rate of 5 per cent. Before the control moves: which label does the discounted process take?
Turning the drift changes the label the process carries
The discounted standard process at one year. Move the drift used by the rule and nothing about the paths changes, but the conditional average crosses the present value line and the label flips with it. The crossing is at a drift of exactly 5 per cent, the risk-free rate.
How does a drifting process split into a fair game and a trend?
Here is the result that makes the weaker properties worth having. A submartingale is not a mysterious object that resists analysis. The process comes apart cleanly into two pieces: a fair game that averages nowhere, and a predictable trendThe part of the movement that is settled before the step begins, so it carries no uncertainty at all. whose whole path is settled before any of it happens. The split is named the Doob decomposition in discrete time and the Doob and Meyer decomposition in continuous time, and the result says the split always exists and is unique.
On the standard process the split is not merely available, it is visible in one line of algebra. Write the discounted process out and the exponential separates into a piece with the random driver in it and a piece without.
| \(W_t\) | standard Brownian motion under the model's own measure, the only random thing here |
| \(\sigma\) | the volatility of the standard process, 20 per cent a year |
| \(S_0\) | the starting value, Rs 100/- |
| \(e^{\sigma W_t - \frac{1}{2}\sigma^2 t}\) | the fair game piece, which averages exactly one at every time |
| \(e^{(\mu-r)t}\) | the trend piece, which contains no random quantity at all and is known today for every future date |
Put the numbers in. The fair game piece is an exponential martingaleA process built as e raised to a random driver less half its own variance, constructed so that its average is exactly one at every time. and it averages Rs 100/- at every horizon, contributing nothing whatever to the average and everything to the spread. The trend piece at one year is e raised to 0.03, or 1.030454. The trend piece contributes Rs 3.045453/- and carries no uncertainty at all. Multiplying the two returns Rs 103.045453/-, the number in the table above, reached from a completely different direction.
The submartingale is a martingale with a known trend bolted to it, and once the trend can be named the process holds no remaining mystery about its direction. That is why the decomposition matters more than it first appears. The trend is settled before each step begins, so the two never interfere. Everything hard about a process lives in the fair game piece and everything directional lives in the trend piece.
The discounted process under the model's own rule splits into what two parts?
Where does each of the two turn up in this subject?
More often than the martingale does. The frequency surprises people. The martingale is the special case and special cases are rare; the one sided classes are where most processes actually live.
The submartingale turns up first wherever a quantity grows faster than the rate it is discounted at. The case just worked is exactly that. The submartingale turns up second, and much more interestingly, in the application of a convex functionA function whose graph bends upward, so that a chord joining two points on it lies at or above the curve between them. to a martingale. Take any martingale and square it, or take its absolute value, or apply any other function that bends upward, and the result is a submartingale. The conditional form of Jensen's inequality does the work, and the same bowl from the naming section arrives in a new place. A function that bends upward sits below its own averages, so averaging pushes the value up.
| \(M_t\) | any martingale, here the discounted standard process under the pricing measure |
| \(\varphi\) | any convex function, such as squaring or taking the absolute value |
| \(\ge\) | the conditional form of Jensen's inequality, which supplies this step |
| \(\varphi(M_s)\) | the function applied to the present value, which is what the chain lands on |
Put a number on it. The discounted standard process under the pricing rule is a martingale sitting at Rs 100/-. Square it and the average at one year is 10,408.107742 rather than 10,000, and the margin is exactly e raised to the variance rate of 0.04. The units are squared rupees, a signal that the quantity measures spread rather than price. The average of the square climbs while the average of the thing being squared stands perfectly still, and that is the whole reason spread grows with the horizon while the centre does not.
The supermartingale turns up wherever a quantity has something taken out of it, or grows more slowly than the rate it is discounted at. Set the drift to 2 per cent against a rate of 5 per cent and the discounted process averages Rs 97.044553/- at the horizon: a supermartingale by 3 per cent a year in the other direction. The everyday version is a household water tank that is drawn on daily and refilled only when rain arrives. Whatever the rain does, the average level tomorrow sits below the level today by whatever was drawn, and the property holds without needing any theory of rainfall at all.
Take a martingale and square it. What is the square?
Why is the weaker property often the only one that can be established?
Because an inequality asks far less of the evidence than an equality does. To establish that the discounted standard process is a martingale under some rule, the drift must be pinned to exactly the rate: not near it, exactly it. To establish that it is a submartingale requires only knowing that the drift is at or above the rate. The sign of a difference is enough, and a sign survives a great deal of ignorance that a size does not.
The one sided statement needs only the sign of the gap, and the sign is knowable in a great many situations where the size is not. That is the whole practical argument for these two words existing as separate objects rather than as footnotes to the martingale definition. A statement that can actually be established about a real process beats a stronger statement that cannot.
The second half of the argument is that the inequality is often all a result needs. Many of the proofs in this subject only ever push the inequality in one direction, so a surprising number of the standard results are stated for one of the two one sided classes rather than for martingales. A nonnegative supermartingale settles down to a limit, and the proof of that needs only the downward inequality. A convex function of a martingale is a submartingale, as above. The square of a martingale is essentially never a martingale, so that result would be useless if it demanded equality.
Inequalities also survive operations that equalities do not. The larger of two submartingales is a submartingale; the larger of two martingales is generally not a martingale. The sum of two submartingales is a submartingale; the sum of two martingales is indeed a martingale, but the first statement covers vastly more cases. Working with the one sided classes gives a set of building blocks that stay closed under the operations most often performed on processes.
Why is a one sided property often more useful than it sounds?
The error that gets made, and what it costs
Reading submartingale as a claim that the process will rise. The submartingale property is a statement about the conditional average and it constrains no path whatsoever. The statement points in a direction, so it feels far more like a forecast than the two sided martingale label does, and that is exactly what makes the misreading hard to notice.
Take the locked path, the published twelve step path of the standard process, and discount it at 5 per cent a year. At month six it reads Rs 108.3347/-. At month nine it reads Rs 90.2852/-, a fall of Rs 18.05/- in three months, well over 16 per cent of where it stood. And the submartingale property holds at every single one of those steps. At every step the conditional average of the next reading sits above the current reading by the monthly trend factor of 1.002503, and what actually happened was decided by the random driver.
| Step | Reading now | Conditional average next month | What the path actually did |
|---|---|---|---|
| Month 6 to 7 | Rs 108.3347/- | Rs 108.6059/- | Rs 100.5850/- |
| Month 7 to 8 | Rs 100.5850/- | Rs 100.8368/- | Rs 97.8044/- |
| Month 8 to 9 | Rs 97.8044/- | Rs 98.0492/- | Rs 90.2852/- |
Every row of that table satisfies the submartingale inequality, and every row of it fell. The cost lands when somebody uses the label as a reason to project a level: a submartingale statement is compatible with a conditional average one rupee above the present value and with one fifty rupees above it, so it can carry no projection at all. The second cost is subtler. Somebody who accepts the label as directional will read a fall as evidence against the model, and will start adjusting a model that was never contradicted by anything they saw.
The discounted locked path falls from Rs 108.33/- to Rs 90.29/-. Does that break the submartingale property?
How does somebody reviewing a model use a one sided statement?
The useful move, when reading work built by somebody else, is to notice that the two one sided labels are cheap to check and expensive to misuse, and that the misuse has a visible signature.
Checking is cheap because only the sign of a difference is needed. If a document says a quantity is a supermartingale under a stated rule, that claim can be tested by asking whether anything is being taken out of the quantity, or whether it grows more slowly than the rate it is discounted against. Neither question needs the model's internals and neither needs a figure to be estimated. A household reading a savings balance that has a fixed amount withdrawn every month can make the same judgement with no arithmetic at all: money leaves, nothing arrives, so the average next month sits below this month.
The signature of misuse is an inequality quietly upgraded into an equality, and it shows up as a projected level sitting where only a bound was established. A document or a model that establishes a submartingale property and then projects a number has taken a statement that permits any gap at all and used it as though it named one specific gap. The tell is that the projected figure has to come from somewhere, and if the only thing established was a one sided bound, the figure came from an assumption that was never written down.
The second thing worth doing is asking which of the three labels is being claimed. People who mean submartingale often write martingale, and the two carry very different obligations. A martingale claim commits somebody to the drift being exactly the rate and is falsified by any gap at all. A submartingale claim commits them to nothing but a sign. Knowing which one has actually been asserted shows what would count as evidence against it, and a claim nobody can say how to falsify is not doing any work in the argument.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for one sided martingale properties and their decompositions | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Doob and Meyer | The decomposition of a submartingale into a martingale and a predictable trend | Probability theory literature |
| Jensen | The inequality on convex functions, in its conditional form | Probability theory literature |
| Hull, Shreve and Wilmott | Standard texts on stochastic calculus for finance | Textbook literature |
The standard process and its locked path are invented.
Educational material. Not advice on any investment, tax, budget or market position.
