Markov Process vs Martingale: Two Different Promises
A Markov process promises that the future depends on the past only through the present level. A martingale promises that the average future value equals the present value. One promise is about dependence and the other about centring, so neither contains the other, and every one of the four combinations of the two is occupied by a process that can be written down.
The reason the two get confused is that both are stated as an equation with a conditional expectation on the left, so they look like variations of one idea. The resemblance is only in the notation. One equation restricts what the right hand side is allowed to be a function of. The other restricts what number the right hand side comes out at. The two restrictions are separate demands, and a process can meet either one while failing the other completely. Reading either as implying the other is the commonest error in the subject.
What is a Markov process, defined from scratch?
Freezing any process at a moment leaves two things: the value it is showing right now, and the entire record of how it got there. A Markov processA process whose future distribution depends on the past only through the present state. is one where the second of those does no work. Discarding the record and keeping only the current value leaves every statement about what happens next unchanged.
Think of a lift. The lift knows which floor it is on. Where the lift can travel next is settled entirely by where it is standing, so whether it reached the eighth floor by climbing steadily from the ground or by descending from the fifteenth makes no difference to anything ahead of it. The claim is not that the history has been forgotten or lost; the claim is that conditioning on the full history and conditioning on the present value alone produce identical answers. The record can sit open on the desk. The record simply has nothing left to say.
| \(S_t\) | the standard process at time \(t\), the single invented traded quantity used throughout |
| \(\mathcal{F}_t\) | the information available at time \(t\), which is the whole record up to that moment |
| \(f\) | any question that can be asked about the finish, written as a function of it |
| \(T\) | the horizon, one year for the standard process |
| \(\mathbb{E}[\cdot\mid\cdot]\) | the average of the left quantity given that the right one is known |
The phrase for every bounded f is the load-bearing part and it is what a hurried reading drops. The equality is not asserted for a single question. The equality is asserted simultaneously for the chance of finishing above Rs 90/-, the chance of finishing above Rs 120/-, the average finish, the spread of the finish, and every other question anybody could pose. Holding for every question at once is what makes the property a statement about the conditional distributionThe whole shape of the future given what is known, not just one summary number taken from it. rather than about any one summary of it.
Here it is on the case. The standard process, an invented traded quantity, starts at Rs 100/-, drifts at 8 per cent a year and carries a volatility of 20 per cent a year, over a one year horizon. Standing at Rs 100/- today, the chance it finishes above Rs 90/- is 0.795826, above Rs 100/- is 0.617911 and above Rs 120/- is 0.270399. Every one of those three numbers, and every other number of that kind, is reproduced exactly by somebody who is shown today's level and is shown nothing else at all.
Two completely different histories arrive at the same level at the same moment. What does memorylessness say about their futures?
What is a martingale, defined from scratch?
A martingaleA process whose conditional average future value equals its present value. is a process whose average future value, taken over everything still possible given what is known now, is exactly its present value. There are three conditions rather than one: the process is known by the time it happens, its average exists at all, and that conditional average returns the present level at every moment.
Notice the shape of that demand and how narrow it is. The demand fixes one number, the conditional average, at one value, the present level. The fair game condition says nothing whatsoever about how wide the distribution of future values is, how quickly that width grows, or how far any single path might travel. A weighing scale that reads a little high on one weighing and a little low on the next is fair in exactly this sense: the readings scatter, and the scatter is centred on the true weight. Fairness is a claim about where the readings sit on average, not a claim that any one reading is close.
| \(M_t\) | the process under test at time \(t\), known by that time rather than later |
| \(\mathcal{F}_t\) | the information available at time \(t\) |
| \(\mathbb{E}[|M_T|]<\infty\) | the condition that the average exists at all, which is real rather than bookkeeping |
| \(T\) | any later time, up to the one year horizon |
Because the demand touches only the average, three processes can obey it while looking nothing like each other. Suppose one is sitting at Rs 100/- and the next step has a typical size of Rs 1/-, a second is sitting at Rs 100/- with a typical step of Rs 2.50/-, and a third is sitting at Rs 100/- with a typical step of Rs 4/-. The condition fixes the centre and takes no view at all on the width, so all three satisfy it identically. The counterexample below is built out of exactly that freedom, so the freedom is worth holding on to.
What does each of the two promise, in one sentence each?
Now the two definitions can be set beside each other, and the useful comparison is not which one is stronger. Neither is stronger. The two properties are answers to two different questions, and the fastest way to keep them apart permanently is to hold the pair of words that names those questions.
The first question is about dependenceWhat the future is allowed to be a function of: the memorylessness question.: what is the future allowed to be a function of? Memorylessness answers today's level and nothing else. The second question is about centringWhere the future distribution sits: the fair game question.: where does the future distribution sit? The fair game property answers exactly on the present value. One statement restricts the arguments of the function and the other restricts its output, and no amount of staring at either sentence will produce the other.
| \(X_t\) | a general process, used here because the statement is not about the standard process alone |
| \(\mathcal{F}_t\) | the information available at time \(t\) |
| \(f\) | any bounded question about the finish, on the left only |
| left side | a statement about which quantity may appear after the conditioning bar |
| right side | a statement about the number the average comes out at |
One of the two properties is about dependence. What is the other one about?
If the two really are independent demands, that is not something to assert. Independence is something to settle, and there is only one way to settle it: put a process in every combination. Two properties, each present or absent, gives four boxes. If any box turned out to be empty, one property would be saying something about the other, and the whole point of keeping them apart would collapse.
Four boxes: memoryless only, fair game only, both, neither. How many of them are occupied?
Is a Markov process a martingale, and is a martingale Markov?
No, and no. Here are four processes, one for each box, all built from the same invented case so that nothing is smuggled in. The standard process starts at Rs 100/-, drifts at 8 per cent a year, carries a volatility of 20 per cent a year, against a risk-free rate of 5 per cent a year over one year. The locked path is its published twelve step path. Its twelve driving values sum to nil exactly by construction, and they are minus 0.5, 1.6, minus 1.3, minus 0.1, 0.1, 1.5, minus 1.3, minus 0.5, minus 1.4, 0.4, 0.9 and 0.6.
Memoryless and not a fair game
The standard process itself, under the model's own rule. The standard process is memoryless, for the reason set out above: today's level settles the whole distribution of the finish and the route to that level adds nothing. And the standard process drifts, so it is not a fair game. From Rs 100/- today its average at the horizon is Rs 108.33/-, being Rs 100/- grown continuously at 8 per cent for a year, and its median finish is Rs 106.18/-, being Rs 100/- grown at 6 per cent, the drift less half the variance rate.
The centre of the distribution has moved Rs 8.33/- away from where the process started, so the fair game condition fails at the first moment it is tested, and memorylessness never promised otherwise. This is not a corner case dragged in to make a point. The standard process is the central object of this whole subject area, and it sits in the box readers most often assume cannot exist.
A fair game and not memoryless
A fair game that is not memoryless is the box readers doubt, so the occupant is worth building explicitly rather than gesturing at. Take a running total that starts at Rs 100/-. At each of twelve steps it adds a fresh draw whose average is nil and whose typical size is one unit. A running total of that kind would be both a fair game and memoryless. Now add one rule: multiply each step by one plus a half for every earlier step that came out negative. The multiplier makes each step a history dependent stepA step whose size is set by what has already happened: memorylessness breaks and the average of the step is left untouched., and it changes exactly one of the two answers.
| \(X_k\) | the running total after \(k\) steps, in rupees, starting at Rs 100/- |
| \(Z_k\) | the step's driving value, average nil and typical size one, taken here from the locked path |
| \(n_{k-1}\) | how many of the first \(k-1\) driving values came out negative |
| \(c_k\) | the step multiplier, known one step early because it counts only what has happened |
| \(\mathcal{F}_{k-1}\) | the information available just before step \(k\) |
Work through why the fair game property survives. The multiplier is settled before the step happens: it counts falls that have already occurred, so it is a known number rather than a random one at the moment the step is taken. Multiplying a quantity whose average is nil by a number already known leaves the average at nil. So the average of the next level, given everything known, is the present level, at every step, exactly. No approximation and no limiting argument.
Now work through why memorylessness dies. Run the construction along the locked driving values. The multiplier starts at 1.0 and climbs as falls accumulate: 1.0, then 1.5, 1.5, 2.0, 2.5, 2.5, 2.5, 3.0, 3.5, 4.0, 4.0, 4.0. The running total reads Rs 99.50/-, Rs 101.90/-, Rs 99.95/-, Rs 99.75/-, Rs 100.00/-, Rs 103.75/-, Rs 100.50/-, Rs 99.00/-, Rs 94.10/-, Rs 95.70/-, Rs 99.30/- and Rs 101.70/-.
| Step | Driving value | Falls so far | Multiplier | Step in rupees | Running total |
|---|---|---|---|---|---|
| 1 | minus 0.5 | 0 | 1.0 | minus 0.50 | Rs 99.50/- |
| 2 | 1.6 | 1 | 1.5 | plus 2.40 | Rs 101.90/- |
| 3 | minus 1.3 | 1 | 1.5 | minus 1.95 | Rs 99.95/- |
| 4 | minus 0.1 | 2 | 2.0 | minus 0.20 | Rs 99.75/- |
| 5 | 0.1 | 3 | 2.5 | plus 0.25 | Rs 100.00/- |
| 6 | 1.5 | 3 | 2.5 | plus 3.75 | Rs 103.75/- |
| 7 | minus 1.3 | 3 | 2.5 | minus 3.25 | Rs 100.50/- |
| 8 | minus 0.5 | 4 | 3.0 | minus 1.50 | Rs 99.00/- |
| 9 | minus 1.4 | 5 | 3.5 | minus 4.90 | Rs 94.10/- |
| 10 | 0.4 | 6 | 4.0 | plus 1.60 | Rs 95.70/- |
| 11 | 0.9 | 6 | 4.0 | plus 3.60 | Rs 99.30/- |
| 12 | 0.6 | 6 | 4.0 | plus 2.40 | Rs 101.70/- |
Look at step five. The running total there is Rs 100.00/- exactly, the level it started from, and the multiplier waiting for the next step is 2.5. Now take a different five step history that also lands on Rs 100.00/- at exactly the same moment: driving values of 0.4, 0.6, 0.5, 0.3 and minus 1.8. The five driving values sum to nil as well. No fall occurs until the very last of them, so every multiplier along that route is 1.0, and the running total reads Rs 100.40/-, Rs 101.00/-, Rs 101.50/-, Rs 101.80/- and then Rs 100.00/-. Same moment, same level. The multiplier waiting for the next step on this route is 1.5, not 2.5.
Two routes, one date, one identical level of Rs 100.00/-, and a next step that is two and a half units wide on one and one and a half units wide on the other. Today's level does not carry that difference in the distribution. That is a counterexampleA single case that settles whether one property implies another, by exhibiting one without the other. rather than an argument, and one is all that is needed. The process is a fair game at every step and is not memoryless anywhere.
A running total whose step size is set by the count of earlier falls. Fair game, memoryless, both or neither?
Both at once
Take the standard process again and change nothing about it. Same paths, same starting value, same volatility. Now apply the pricing rule and look at the discounted processThe process multiplied by the price today of a rupee at that time. instead: the standard process multiplied by the price today of a rupee at the horizon. Under that rule its average at the horizon is Rs 100.000000/- exactly. Discounting multiplies by a number that depends on the clock and on nothing in the history, so memorylessness survives untouched.
The process moved from one box to another and nothing about the process changed. No sharper evidence exists that the two properties are answering different questions. The paths are the same paths. The outcome set has not gained or lost a member. Only the rule for taking the average moved, and only the centring answer responded.
| \(\mathbb{P}\) | the physical measure, the model's own rule for weighting the paths |
| \(\mathbb{Q}\) | the risk-neutral measure, the rule used for pricing |
| \(\mu\) | the drift of the standard process, 8 per cent a year |
| \(r\) | the risk-free rate, 5 per cent a year continuously compounded |
| \(e^{-rT}\) | the discount factor over the one year horizon, 0.951229 |
Holding both properties at once buys two different things, and they are needed at two different points in almost every argument in this subject area. Because the quantity is a fair game, its future average can be read straight off today with no growth rate estimated and no view taken. Because it is memoryless, the computation that produces that average can carry one number forward from step to step instead of carrying an entire history.
The first property is what makes an argument provable and the second is what makes the computation possible. Models are built to hold both rather than either for exactly that reason. Drop the fair game property and the average needs a correction nobody can pin down. Drop memorylessness and the calculation has to branch on every route rather than on every level. The difference is a manageable tree against one that doubles at every step.
What does a process holding both properties give an argument that neither one alone gives?
Neither of the two
The fourth box needs no new idea at all. Take the running total with its history set step sizes and add Rs 0.25/- to every step. The average of the next level given everything known is now Rs 0.25/- above the present level rather than equal to it, so the fair game property fails at every step. The step widths still depend on the count of earlier falls, so memorylessness still fails too. Over twelve steps the added drift comes to Rs 3/- exactly, so the process averages Rs 103/- at the horizon from a start of Rs 100/-, and along the locked driving values it finishes at Rs 104.70/-.
Worth noting what the failure in the fourth box is not. A process that drifts upward at every step in this way is a submartingale, and one that drifts downward is a supermartingale, and neither is a martingale. Not a fair game does not mean drifting up: it means the centre moved, in whichever direction it moved. The simulation below turns exactly on that point.
The drift moves on the control below. Which of the two properties responds?
Move the drift and watch one answer refuse to respond
The discounted standard process. The control changes the drift from 0 to 10 per cent a year, the discounted average slides along the scale, and the marker moves between the two boxes. The memorylessness column is drawn at every setting and never once changes.
Run the control down to 2 per cent and watch what happens. The discounted average falls to Rs 97.044553/-, below the starting value rather than above it, and the marker stays in the lower box. A discounted average below the starting value is the supermartingale case, and it fails the fair game test just as completely as the 8 per cent case does. Only one setting out of the twenty one lands in the upper box. There the drift and the rate cancel to nil, so it lands exactly.
The error that gets made, and what it costs
Assuming that a memoryless process must be a fair game, or that a fair game must be memoryless. The standard process is the counterexample to the first and the running total with history set steps is the counterexample to the second, and both counterexamples are one line long.
The first direction is the dangerous one. The process that breaks it is not an exotic construction, but the object the whole subject area is built on: memoryless in its level, and averaging Rs 108.33/- at the horizon from a start of Rs 100/-. Somebody who establishes memorylessness and then quietly helps themselves to an average has moved Rs 8.33/- without writing anything down.
Here is what makes it expensive. The gap is a step in an argument rather than a step in a calculation, so no numerical check anywhere can find it. Every figure downstream is arithmetically correct given the line above it, every reconciliation ties, and every total agrees with its parts. The fault sits at the join between two sentences, in a place where nobody did any arithmetic at all, and it survives every test that operates on numbers because it never touched a number.
The standard process is memoryless. Is it a fair game?
Which of the two does a pricing argument actually need?
Both, at different steps, and the useful skill is telling which step needs which. A short test settles it every time, and it turns on what the step is doing rather than on what it is about.
- Does the step take an average across time?
Any step that says the value of something now equals the average of something later is leaning on the fair game property. Without it the average carries a correction that has to be estimated, and the estimate is exactly the thing the argument was trying to avoid needing.
Tell: the words average, expectation, or a value today set equal to a value later.
- Does the step carry a state forward?
Any step that computes one moment from the moment before, on a grid, a tree or a differential equation, is leaning on memorylessness. Without it the computation has to branch on the whole route rather than on the current level.
Tell: a recursion, a backward induction, a lattice node, or a partial differential equation.
- Does the step do both?
Most of them do, which is why models are built to hold both properties rather than either. Discounted values that satisfy the fair game property are averaged backwards through a tree whose nodes exist only because the process is memoryless.
Tell: a backward average taken node by node.
- Does the step assume one having established the other?
This is the fault. Establishing memorylessness and then averaging without correction, or establishing the fair game property and then collapsing histories into levels, are the two directions of the same error.
Tell: the word therefore, sitting between two sentences that are about different properties.
Notice that path dependenceThe case where the route taken changes the future, so the current level is not a sufficient summary. is where the second property matters most and where its absence is most expensive. A quantity that depends on the route can still be a perfectly good fair game, so nothing about the pricing argument breaks. The computation is what breaks: the state that has to be carried forward is no longer a single number. The two properties fail in different places and cost different things. The practical reason for refusing to treat them as one idea is exactly that.
An argument establishes memorylessness and then takes an average with no correction. What has gone wrong?
How does somebody checking a model rather than building one use this?
The four step test above is worth running on work written by someone else, and its best feature is how little it needs from the model. Following a derivation is not necessary in order to ask which of the two properties a given line is leaning on; the verb is enough. A line that averages is asking for the fair game property; a line that steps forward is asking for memorylessness; and a line that does one having argued the other is where a reviewer earns their morning.
Two tells are worth carrying. The first is a growth rate appearing in a projection of something described as a fair game. If the property really holds under the rule being used, the future average is the present value, so no growth rate is needed at all, and a projection carrying one has either changed rules partway or was never dealing with a fair game. The second is a computation that stores a level per node while the quantity being computed clearly depends on the route: an average taken over a path, a running maximum, a total accumulated so far. A lattice will happily produce a number whether or not the state it carries is sufficient, and the number will look entirely reasonable, so the second tell is the more common of the two in practice.
There is a household version of both tells that makes them easy to remember. Somebody who tracks their spending by keeping only the current balance in a passbook can answer questions about the balance perfectly well. How much was spent on travel this year depends on the route, and the passbook keeps only the level, so that question is beyond them. And a fair coin toss game where the stake doubles after every loss is scrupulously fair at every single toss. What happens next plainly depends on how many losses have already happened. Neither situation is unusual, and neither is a technicality.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for martingale and Markov methods in pricing | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Doob | Stochastic Processes, the source of the martingale results that carry his name | Wiley |
| Hull, Shreve and Wilmott | Standard texts on derivatives pricing and stochastic calculus | Pearson, Springer and Wiley |
The standard process, its four parameters, the twelve driving values and the running total built from them are invented.
Educational material. Not advice on any investment, tax, budget or market position.
