The Markov Property: Memorylessness Precisely Stated
The Markov property says that the distribution of the future, given the whole past, depends on the past only through the present state. The property is a statement about distributions rather than about averages, and it does not say the process forgets where it has been. The level itself carries forward in full. Only the route to it is discarded.
Everything in this guide follows from one claim: that a single quantity, the value the process shows right now, is a complete stand-in for the entire history when the question is what happens next. The whole of the mathematics is an unpacking of what that sufficiency includes and, much more importantly, what it does not. A lift knows which floor it is on and nothing about how it got there, and that is enough to say where it can go next. Knowing the floor is not enough to say the lift has forgotten anything. The lift stands on the eighth floor because somebody pressed eight.
What does the Markov property say, precisely?
Consider a process frozen at some moment. Two things are then in hand: the value it currently shows, and the whole record of how it arrived there. The Markov propertyThe future distribution depends on the past only through the present state. is the claim that if the second is thrown away and the first kept, no statement about the future changes.
The claim of sufficiency is a strong one when it is written out carefully. Sufficiency is not a claim that the record is uninteresting, or that the record is unavailable, or that the record has been lost. The record can sit in full view. The claim is that conditioning on the record and conditioning on the present value produce exactly the same answer, for every question about the future that might be asked. Nothing is being hidden. The extra information is simply not doing any work.
| \(S_t\) | the standard process at time \(t\), the single invented traded quantity used throughout |
| \(\mathcal{F}_t\) | the information available at time \(t\), which is the whole record up to that moment |
| \(A\) | any set of values the process might finish in, for instance everything above Rs 100/- |
| \(\mathbb{P}(\cdot\mid\cdot)\) | the chance of the left statement given that the right one is known |
| \(T\) | the horizon, one year for the standard process |
The phrase for all sets is the load-bearing part of that block and it is the part a hurried reading drops. The equality is not asserted for one convenient question. The equality is asserted simultaneously for the chance of finishing above Rs 100/-, for the chance of finishing between Rs 95/- and Rs 105/-, for the chance of finishing in the top tenth of outcomes, and for every other question anybody could pose. Asserting every question at once is what makes the property a statement about the conditional distributionThe whole shape of the future given what is known, not just its average. rather than about any one summary of it.
Here is the concrete version, and it is the picture worth carrying. The standard process starts at Rs 100/-, drifts at 8 per cent a year and carries a volatility of 20 per cent a year. Its published twelve step path, the locked path, stands at Rs 111.0772/- at month six, having dipped to Rs 97.64/- in month one and jumped to Rs 107.63/- in month two. Now imagine two other histories that also finish month six at Rs 111.0772/-: one that climbs there in six even steps without a single down month, and one that collapses to Rs 84.94/- by month two and then claws back. All three give the identical distribution for the finish, and the identity is exact rather than approximate.
The surprise in that figure is in the drawing rather than in the algebra, and the drawing is worth sitting with. The three lines on the left could hardly be less alike. One of them spends four months below where it started. One never has a down month at all. One loses more than fifteen per cent of its value before recovering. The right half of the picture is drawn a single time, and not because of an economy in the illustration. There is genuinely only one of it.
Is the Markov property a statement about the conditional average, or about the conditional distribution?
Is it a claim about the average, or about the whole shape?
The average version and the distribution version are two different statements, and only one of them is the property. The weaker version constrains a single number. The property constrains the entire shape. Writing both down shows how far apart they are.
| \(\mathbb{E}[\cdot\mid\mathcal{F}_t]\) | the average over what is still possible given the whole record at time \(t\) |
| \(g(\cdot)\) | some function of the current value alone, whatever it happens to be |
| \(S_t\) | the current value of the standard process, the only thing \(g\) is allowed to see |
A process can satisfy that and fail the property outright. Two records showing the same current value can agree on the average of the finish and disagree on almost everything else about it, and the moment they disagree on anything the property has failed. The standard process gives the cleanest possible demonstration, and it uses machinery already locked in for this subject.
Widen the model so that the volatility is itself a moving quantity rather than a constant, using the locked volatility model whose variance starts at 0.04, the square of the locked 20 per cent. Now hold two situations side by side. In both the process stands at Rs 111.0772/- at month six. In the first, the variance also stands at 0.04, so volatility is 20 per cent. In the second, the variance stands at 0.09, so volatility is 30 per cent. Because the average depends on the drift and the level and on nothing else, both answer Rs 115.610377/- for the average of the finish, to the last digit shown. Ask for the chance of finishing above Rs 100/- and the first says 0.830208 while the second says 0.718278.
The contrast does two jobs at once, and both are worth naming. A version stated in terms of averages would wave both of those situations through, and the property therefore has to be stated in terms of distributions. The second situation is also a quiet preview of the hinge below, being exactly a process that fails the property in its level alone. The failure in the level alone is returned to below.
Same level of Rs 111.0772/- at month six, two different volatility readings. Do the two agree on the average of the finish?
Which processes actually have the property?
Whether a process has the property is checked rather than assumed, and the check is short enough to run mentally once it has been seen done. The check feels hard only because the wrong question gets asked first. People ask whether the process looks like it is wandering aimlessly. Wandering is not the test, and the answer it suggests is often the opposite of the truth.
How to Test Whether a Process Is Markov
- Write down the state being proposed
Before anything else, say what quantity is being claimed as sufficient. Almost always somebody means the level and has not said so. Until the proposed stateThe summary of the present that the property claims is sufficient. is written down there is nothing to test, because the property belongs to the process and the state together rather than to the process alone.
Skipping this is the commonest fault, and everything below is meaningless without it.
- Write down what the next move depends on
Go to the rule that generates the next move and list every quantity that appears in it. Not the quantities that happen to be interesting: the quantities the rule actually reads.
A rule that reads only the current value and a fresh disturbance is already halfway to passing.
- Ask whether each of those is recoverable from the state
For every quantity on that list, ask whether somebody holding only the proposed state could reconstruct it. If yes, the state is sufficient. If even one cannot be reconstructed, it is not.
A running average, a running maximum and a time since the last jump all fail this step.
- If it fails, widen the state and run the test again
Add the quantity that could not be reconstructed and retest. This is not a repair to a broken process; it is a correction to a badly chosen state.
Widening always works in principle. What it costs is the subject of the next section.
Run that on a handful of processes and a pattern appears fast. The property is not about wandering, it is about bookkeeping: it holds exactly when the rule for the next move reads nothing the current state cannot supply. A perfectly smooth, entirely predictable process is Markov. A wildly erratic one that reads its own three month average is not.
| Process | Markov in its level alone? | Why |
|---|---|---|
| Standard Brownian motion | Yes | Increments are independent of everything before them, which is stronger than needed |
| The standard process, with constant drift and volatility | Yes | The next move scales off the current level and a fresh disturbance |
| The locked mean reverting rate model | Yes | The pull is toward a fixed long-run mean, computed from the current level |
| A counting process with constant intensity | Yes | The chance of an arrival never reads how long since the last one |
| The running maximum of the standard process | Yes, as a pair with the level | Alone the level cannot say what the peak was, so the pair is the state |
| The standard process with a moving variance | No | The spread of the finish reads the variance, which the level cannot supply |
| A process whose drift reads its three month average | No | The average of months four, five and six is not recoverable from month six |
| A process whose next move reads a coin tossed tomorrow | Not even adapted | It fails an earlier condition than this one and never reaches the test |
Two of those rows deserve a sentence each. Brownian motion has independent incrementsEach move is unrelated to every move before it, which is a stronger condition than memorylessness.. Each move is unrelated to every move that came before. Independence is a strictly stronger condition than the Markov property, and the implication runs one way only. Independent increments give the property immediately. The property does not give independent increments back. A process is perfectly free to have its next move depend on where it currently stands, and a dependence on where it stands is a dependence on the past routed entirely through the present value.
The three month average row is the one that catches people. Look at the three histories from the figure above. At month six the locked path shows a three month average of Rs 104.2333/-, the even climb shows Rs 109.1604/-, and the deep dip shows Rs 101.4436/-. All three stand at Rs 111.0772/-. A process whose drift reads that average would move differently in each case, so the level alone has not summarised the past. Any averaging rule introduces path dependenceThe case where the route taken changes the future, which the property rules out., and nothing about it is exotic.
A process whose next move depends on its average over the last three months. Is it Markov in its level?
What does the property look like on the standard process at month six?
Everything so far has been the statement. Here is the property doing arithmetic, on the locked path, at the month six node. The month six node is the worked instance the rest of this guide is checked against.
Before reading on. The average of the finish, given the process stands at Rs 111.0772/- at month six. Higher or lower than Rs 111.0772/-?
The process stands at Rs 111.0772/- with half a year left to run. Under the physical measure P the logarithm of the finish, conditional on that level, is normally distributed. Its average is the logarithm of the level, 4.710226, plus the drift less half the variance rate over the remaining half year. The addition is 0.03 exactly. Its standard deviation is the volatility times the square root of the remaining time.
| \(S_t\) | the level at month six, Rs 111.0772/- on all three histories |
| \(\mu\) | the drift of the standard process, 0.08 a year |
| \(\sigma\) | the volatility, 0.20 a year, so the variance rate is 0.04 and half of it is 0.02 |
| \(\tau\) | the time still to run, half a year here |
| \(\mathcal{N}(m,v)\) | the normal distribution with average \(m\) and variance \(v\) |
Read the right hand side of that block once more and check what is in it. The level, the drift, the volatility, the time remaining. The drift and the volatility are parameters of the process rather than facts about the history. The time remaining is a clock reading. The level is the only thing on the right hand side that the first six months could have influenced. Replacing the six months with the level therefore changes nothing.
The next step is the average of the finish itself rather than of its logarithm. The half variance correction of 0.02 that appears in the block above cancels exactly when moving from the logarithm back to the level. The average of the finish therefore grows at the full drift.
| \(S_t e^{\mu\tau}\) | the level at month six grown at the drift for the time still to run |
| \(\mu\tau\) | 0.08 times 0.5, which is 0.04 exactly |
| \(e^{0.04}\) | 1.040811, the growth factor over the remaining half year |
Set against the unconditional figure, the gap shows what the level was worth. From today, knowing only that the process starts at Rs 100/-, the average of the finish is Rs 108.3287/- and the chance of finishing above Rs 100/- is 0.617911. From month six, knowing the level is Rs 111.0772/-, the average is Rs 115.610377/- and the chance is 0.830208. Conditioning on the level moved the chance of finishing above Rs 100/- from 0.617911 to 0.830208, so the level is carrying a great deal of information, and the property never claimed otherwise.
| What is given | Average of the finish | Chance above Rs 100/- |
|---|---|---|
| Only the starting value of Rs 100/- | Rs 108.3287/- | 0.617911 |
| The level at month six, by the locked path | Rs 115.610377/- | 0.830208 |
| The level at month six, by the even climb | Rs 115.610377/- | 0.830208 |
| The level at month six, by the deep dip | Rs 115.610377/- | 0.830208 |
| The level at month six plus all six earlier readings | Rs 115.610377/- | 0.830208 |
The last row is the property written as a table. Handing over the six readings on top of the level bought nothing at all, and the first row shows the level itself bought a great deal. The two facts sit together comfortably, and a reader who finds them in tension has read memorylessness as a claim about the level rather than about the route.
Two very different histories both arrive at Rs 111.0772/- at month six. Before switching between them below: does the fan of futures change?
Switch the history and watch what refuses to move
Three histories, all arriving at exactly Rs 111.0772/- at month six. Switch between them and the left half redraws completely. There is only one fan and one density strip, so the right hand side is drawn once. Then move the line and watch the one number that does respond.
The two green panels carry the lesson of that control. However often the history is switched, neither of them moves by so much as a digit in the sixth decimal place. The two white panels beside them swing between Rs 84.94/- and Rs 100/- for the low and between Rs 11.08/- and Rs 41.19/- for the distance travelled. The line, by contrast, does move the chance. Moving the line changes the question rather than the information.
| \(K\) | the level being asked about, Rs 100/- by default on the control above |
| \(\mathcal{N}(\cdot)\) | the standard normal distribution function |
| \(S_t\) | the level at month six, the only input carrying anything from the past |
| \(\sigma\sqrt{\tau}\) | 0.141421, the spread of the logarithm over the remaining half year |
Why does knowing the level buy so much when knowing the route buys nothing?
Because those are two different pieces of information and only one of them was ever relevant. The figure below puts the two conditionings side by side, and the areas are the two probabilities computed above.
The everyday version is a weighing scale. Somebody stands on it and it reads a number. The reading is a genuine summary of a great deal of history, and somebody who knows it can say useful things about that person that they could not say without it. The scale cannot tell them, and does not pretend to tell them, whether the reading was arrived at by gaining and then losing, or by holding steady for a year. The reading is not forgetful, it is compressed, and the whole content of this property is the claim that for this particular question the compression is lossless.
The error that gets made, and what it costs
Reading memoryless as saying the process has forgotten its history, and concluding from that reading that where the process stands now carries no trace of where it has been.
The locked path stands at Rs 111.0772/- in month six precisely because of what happened in the first six months. The first six months are why it is not standing at Rs 93.74/-, the level it reaches three months later. The conditional average of Rs 115.610377/- is computed straight off that level and is therefore computed straight off the history, once removed. The property discards the routeThe particular sequence of values by which the present was reached.. The property keeps entire and unmodified the destinationThe present level itself, which the property keeps in full..
The cost is a mechanism attributed to the wrong property. A reader holding the stronger version will expect the level to revert toward some earlier value, or to ignore its own history in some visible way, and will treat a model that does neither as broken. Worse, the reverse mistake follows. The reader decides that carrying a level is what memorylessness authorises, and a genuinely path dependent quantity gets pushed through machinery that only carries a level. The stronger version survives because it is easier to say. Memorylessness is a two word phrase and the property it names takes a paragraph.
Does the Markov property mean the process has forgotten that it rose to Rs 111.0772/-?
How does the choice of state decide the answer?
The hinge of the whole subject is one sentence, and it is the sentence most worth taking away. The property is not a property of a process. The property belongs to a process together with a proposed state, and asking whether a process is Markov without saying in what is an incomplete question in exactly the way that asking whether a process is a martingale without naming the measure is incomplete.
Go back to the moving variance case from earlier. The standard process with a variance that moves is not Markov in its level alone, and the reason is now visible: a second situation with the same level and a different variance gives a different distribution for the finish. Write the state out as a pair rather than a single number and the failure disappears.
| \(v_t\) | the variance process at time \(t\), 0.04 or 0.09 in the two situations compared |
| \((S_t, v_t)\) | the widened state, a pair rather than a single number |
| \(\neq\) | the failure: the level alone does not reproduce the answer the record gives |
| \(=\) | the repair: the pair does reproduce it, for every set and every horizon |
Widening the stateAdding a quantity so that a process that failed the property becomes Markov again. is the standard remedy and it always works in principle: keep adding whatever the future turns out to depend on until nothing is left out. In practice it is not free, and the cost is the reason state design is a real decision rather than a formality. Each quantity added is one more dimension a numerical method has to cover, and the work of covering a grid grows very roughly by a multiplicative factor per dimension rather than by an addition. A model in one state variable and a model in four are not different by a factor of four.
A process is not Markov in its level. What is the standard remedy?
What does the property let a calculation stop carrying?
Everything above has been about what the property means. The property buys the difference between a calculation that can be run and one that cannot.
The small version comes first. Describing the finish from month six needs the number Rs 111.0772/-. The description does not need Rs 97.64/-, Rs 107.63/-, Rs 100.35/-, Rs 100.27/- or Rs 101.35/-. Six readings collapse to one, and the five discarded were not approximated or compressed with loss. The five were shown to be irrelevant to the question.
Now scale it up, using the locked twelve step lattice, the construction that carries the names of Cox, Ross and Rubinstein. A lattice that tracks only the level recombines: an up move followed by a down move lands on the same node as a down followed by an up, so twelve steps produce thirteen terminal nodes and ninety one nodes in total. A lattice that has to track the route cannot recombine. The two arrivals are now different situations. Twelve steps then produce 4,096 distinct histories and 8,191 nodes in total.
The recombination that makes a lattice usable is not a trick of the construction, it is the Markov property spent. Nothing else licenses treating those two arrivals as one node. Take the property away and the same twelve steps grow by a factor of more than three hundred, and every step beyond twelve doubles it again. The doubling is the whole of the computational argument, and it is why this property, rather than any other, decides what can be built.
Describing the finish from month six. How many numbers are needed from before that moment?
Why does this decide what a pricing model can be built on?
Because the two workhorse numerical methods in this subject both assume it, and neither of them announces the assumption. A lattice assumes it in the recombination just described. A grid method assumes it in a deeper way: the entire idea that the value of a claim can be written as a function of the current state and the current time, and that this function satisfies an equation in those variables, is the Markov property in disguise. If the value depended on the route, there would be no function of the state to solve for, and no equation to write.
The bridge between such an equation and an average taken over the process carries the names of Feynman and Kac, and the reason the bridge exists is that both sides are functions of the state alone. The link is a theorem in its own right and is covered separately. The point for now is narrower and firmer: a pricing method that solves for a function of the state has already assumed that a function of the state is enough, and the Markov property is the name of that assumption.
State design also explains a modelling pattern that otherwise looks arbitrary. When a model needs a moving volatility, or a moving rate, or a term whose value depends on a running average, the standard move is not to abandon the numerical machinery. The standard move is to widen the state until the machinery applies again, and then to pay the dimensional cost. The locked volatility model with its variance process, the locked rate model with its short rate, and every model that carries a running quantity alongside the level are all doing the same thing: restoring the property so that the equipment built on it can be used.
How does somebody checking a model rather than building one use this?
The useful thing about the test is that it needs almost nothing from the model. Asking which quantities a model carries as its state needs no knowledge of how a figure was produced, and that single question resolves a surprising share of disagreements about whether a number is right.
The pattern to watch for is a quantity that reads the route being priced on machinery that carries only a level. The tell is usually in the description rather than in the code. Words like running, average, highest so far, cumulative, to date and since inception are all descriptions of the route, and each one implies a quantity the level cannot supply. When one of those words appears in a description and the state carries one number, either the state has been widened somewhere and not documented, or the calculation is answering a different question from the one it was asked.
The second pattern is subtler and more common. A model is built correctly with a widened state, and then somebody downstream summarises it, keeps the level and drops the second state variable because it looked like an intermediate quantity. Every number afterwards is computed from an insufficient state. Nothing errors, nothing is flagged, and the figures come out looking entirely reasonable. The failure is the same as the one in the block above, arriving through a different door: somebody has decided the level is enough because the process was described as Markov, without noticing in what.
A calculation carrying only the current level is asked for something that depends on the highest value reached so far. What has gone wrong?
No jurisdiction sets the definition of a process, and no rule anywhere makes a process Markov or stops it being one. The property is a mathematical statement and it holds or fails on the arithmetic alone.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for Markov methods in continuous time pricing | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Cox, Ross and Rubinstein, 1979 | The binomial option pricing model and its recombining lattice | Journal of Financial Economics, 1979 |
| Feynman and Kac | The link between a parabolic equation and an expectation over the process | Reviews of Modern Physics, 1948, and Transactions of the American Mathematical Society, 1949 |
| Hull, Shreve and Wilmott | Textbook treatments of Markov processes in continuous time | Pearson, Springer and Wiley |
The standard process, its four parameters, its published path and its moving variance are invented.
Educational material. Not advice on any investment, tax, budget or market position.
