Stochastic Differential Equations: Modelling Change Under Uncertainty
A stochastic differential equation states how a quantity changes across each instant: a predictable part multiplied by the time step, plus a random part multiplied by a Brownian increment. The path it describes has no slope at any point, so the statement is shorthand for an integral equation rather than a derivative equation. The equation fixes a distribution of paths, never one path.
Every equation of change written down before this one came with a quiet promise attached. Given a starting value, it hands back exactly one answer for every future moment. Population grows, a tank drains, a balance compounds, and in each case the equation and the starting point together pin down a single curve that could in principle be drawn with a pencil and checked against a ruler. A single determined curve is what an equation of change has always meant.
The equation in this guide breaks that promise on purpose, and it is worth being precise about what replaces it. Fed its published driving values, the standard process below returns a path that finishes the year at Rs 106.18/-. Fed a different set, the whole path redraws, month by month, and finishes somewhere else entirely. Neither answer is more correct than the other. The equation never picked one path. Its job is to fix the weights on all of them at once, and everything computed from it is a property of that whole set rather than of any single line.
What does a stochastic differential equation actually say?
The shape can be read before any of the meaning arrives, so start with the shape rather than the meaning. A stochastic differential equationA statement of how a quantity changes over each instant, carrying a predictable part and a random part. puts the change in a quantity on the left, and on the right it puts exactly two things added together. Two added terms are the whole architecture, never one term and never three, and the two can be told apart without knowing what either one does.
How? By looking at what each is multiplied by. One term is multiplied by the time step, the little stretch of time the statement is about. The other is multiplied by the increment of a Brownian path over that same stretch. The first is called the drift termThe predictable part of the change, multiplied by the length of the time step. and the second is called the diffusion termThe random part of the change, multiplied by the increment of a Brownian path over the same step., and telling one from the other is genuinely a matter of reading which symbol each is attached to.
Here is the everyday version, and it is about measurement rather than about anything traded. A lift reports only which floor it is on, once a second. Between two reports it has moved by some amount that splits in two: the part that follows the button somebody pressed, and the part that is the cable stretching, the car settling, the sensor rounding. The first part scales with the length of the wait, so waiting twice as long doubles it. The second part does not scale that way at all. The settling and the rounding are a wobble, and wobbles do not add up the way intentions do. The equation below is that split, written down.
| \(X_t\) | the quantity being modelled, read at time \(t\) |
| \(a(X_t,t)\) | the drift coefficient, which may depend on the current level and on the time |
| \(dt\) | the time step, the stretch of time the statement is about |
| \(b(X_t,t)\) | the diffusion coefficient, which may also depend on the level and the time |
| \(dW_t\) | the increment of standard Brownian motion \(W\) over that same step, under the physical measure P |
Two things about that line get lost most often, and both deserve saying immediately. First, the letters on the right can depend on where the quantity currently sits. A model in which the random term is a fixed number is a very different animal from one in which the random term is a percentage of the level, and the notation carries both without complaint. Second, the whole statement is made under a stated rule for weighting outcomes, written P here. The stated rule is not decoration, and it is the item most often left off.
How many terms does a stochastic differential equation carry?
What is an Ordinary Differential Equation, and what does adding a random term change?
A comparison is only sharp when both sides have been defined rather than assumed, so put the ordinary object down properly first. An ordinary differential equationA statement of how a quantity changes with no random term at all, whose solution with a starting value is a single curve. says how a quantity changes with time and says nothing else. There is one term on the right, it is multiplied by the time step, and that is the end of the statement. Given a starting value, it hands back exactly one curve.
Take the same numbers this guide uses throughout. A quantity starts at Rs 100/- and grows at 8 per cent a year in a way that has nothing uncertain in it. The ordinary equation for that says the change over any short stretch is 0.08 multiplied by the current level multiplied by the length of the stretch. Its answer is Rs 100/- multiplied by the exponential of 0.08 times the elapsed years. The exponential lands the year at Rs 108.33/- and lands every intermediate month on a number that can be quoted to the paisa in advance.
| \(Y_t\) | the quantity at time \(t\), with no random term anywhere in its description |
| \(\mu\) | the growth rate, 0.08 a year throughout this guide and invented |
| \(Y_0\) | the starting value, Rs 100/- exactly |
| \(t\) | elapsed time in years, running from nought to the horizon |
The ordinary equation permitted something the stochastic one did not: a rate of change, sitting on the left of the equals sign as a genuine ratio of two things. The quantity has a slope at every moment, the slope is a number, and dividing the change by the length of the step is a legal move that converges to it. The permission to divide is precisely what the random term withdraws, so it is worth holding on to.
Now add the second term. The starting value is unchanged, the growth rate is unchanged, and one line of the statement is longer. The equation does not come back with one curve that has become harder to compute. The answer is not a curve at all: the same starting value now sits at the mouth of a whole spread of paths, and the equation says how much weight each of them carries rather than which of them happens. That is a difference in kind, not a difference in difficulty, and no amount of extra computing power collapses the spread back to a line.
An analyst is handed a stochastic differential equation and a starting value. Is there now one path?
Why is it really an integral equation written in shorthand?
The reading in this section survives every later stage of the subject, so it is worth slowing down for. The general form has a quantity written with a small d in front of it on the left, and it looks exactly like the equations that are meant to be divided through by the time step. The line must not be divided through. Nothing in that line is being differentiated, and nothing in it can be.
The reason was settled under Brownian motion and is spent here in a sentence. A Brownian path is continuous everywhere and has a slope nowhere. Not at awkward corner points, not on a set that careful handling could avoid: nowhere at all, at any moment, on any path. So a quantity built out of a Brownian path cannot have a rate of change either, and a statement that appeared to supply one would be describing an object that was never there.
The notation is abbreviating instead. The honest form of the statement is an integral equationA statement that a quantity equals its starting value plus accumulated amounts, written with integrals rather than with derivatives.: the quantity at any time equals its starting value plus two accumulated amounts. The first is an ordinary integral against time. The second is an Ito integral against the Brownian path, and that object is fully defined, was constructed under the Ito integral, and needs no derivative anywhere in its construction.
| \(X_0\) | the starting value, fixed and known before anything happens |
| \(\int_0^t a\,ds\) | an ordinary integral against time, accumulating the predictable part |
| \(\int_0^t b\,dW_s\) | an Ito integral against the Brownian path, accumulating the random part |
| \(s\) | the running time inside each accumulation, from nought up to \(t\) |
| \(W_s\) | standard Brownian motion under the physical measure P |
So the differential line and the integral line are the same statement. The differential one is shorter, it is what everybody writes, and it is safe as long as it is remembered to be an abbreviation. The moment the small d is treated as an instruction to divide, the abbreviation has been left behind and something with no existence is being described, and the arithmetic gives no warning. Ito, whose integral makes the second accumulation well defined, is the reason the honest form can be written at all.
A useful habit follows. Whenever a step in a derivation looks strange, rewrite the line in its integral form and ask whether the step still makes sense there. Most of the operations that feel unsafe in the differential notation are simply illegal in the integral one, where the illegality is obvious. The habit is the same trick as converting a percentage back into rupees when a percentage argument stops feeling right.
Is anything actually being differentiated in a stochastic differential equation?
What does the equation determine, if not a single path?
One question decides whether the rest of the subject makes sense. If the equation does not hand back one path, what does it hand back? The word solutionWhat satisfies the equation. Here it is a process, meaning a rule for producing a level at every moment, rather than a formula giving one number per moment. is doing quiet work here, and it does not mean what it meant in the ordinary case.
The equation determines a distribution of pathsThe whole set of paths the equation permits, each carrying a weight, rather than any one of them.. Every path the equation permits carries a weight, and the equation fixes those weights completely. Fixed weights are why the answer to almost every question comes back as an average, a probability, a spread or a quantile rather than as a level. There is no answer to what the quantity will be in a year. There is exactly one answer to what its average will be in a year.
For the standard process used throughout this guide, that distribution is written down in closed form. The logarithm of the level at the horizon is normally distributed, with a mean of 4.665170 and a standard deviation of 0.20, and every summary that might be wanted falls out of those two numbers. The average level after one year is Rs 108.33/-. The median is Rs 106.18/-. The most likely neighbourhood is the peak of the density, and it sits lower still at Rs 102.02/-.
| \(S_T\) | the standard process at the horizon, invented and carrying no market meaning |
| \(S_0\) | the starting value, Rs 100/- exactly |
| \(\mu\) | the drift, 0.08 a year under the physical measure P |
| \(\sigma\) | the volatility, 0.20 a year, so the variance rate is 0.04 and half of it is 0.02 |
| \(W_T\) | the Brownian path at the horizon, normally distributed with mean nought and variance \(T\) |
| \(T\) | the horizon, one year throughout this guide |
The half variance correction sitting inside that exponential is where the mean and the median part company, so look at it closely. Half the variance rate is 0.02 exactly here, and subtracting it drops the centre of the logarithm below the drift. The result is that the average of the level and the middle of the level are two different numbers, Rs 108.33/- and Rs 106.18/-, and the Rs 2.15/- between them is that correction made visible in rupees.
The gap between an average and a middle is not a technicality, and here is the everyday version. A road has an average width of four metres, so a four metre lorry might be expected to fit. The lorry does not fit. Half the road is six metres wide and the other half is two, and averages do not drive. An average is a summary of a spread rather than a forecast of anything, so every statement this equation supports has to say which summary it is quoting.
The average at the horizon is Rs 108.33/- and the median is Rs 106.18/-. Why are they not the same number?
A different set of driving values is about to be fed to the same equation. Does the average at the horizon move?
Feed the same equation five different sets of driving values
The drift stays at 8 per cent, the volatility stays at 20 per cent and the horizon stays at one year. The only thing that changes is which published set of twelve driving values goes in. Watch the path redraw completely while the density on the right and the two horizontal markers do not move at all.
| Set | The twelve driving values, published and fixed | Total | Finishes at |
|---|---|---|---|
| Published | minus 0.5, 1.6, minus 1.3, minus 0.1, 0.1, 1.5, minus 1.3, minus 0.5, minus 1.4, 0.4, 0.9, 0.6 | 0.00 | Rs 106.18/- |
| Reversed | the same twelve, read from the last to the first | 0.00 | Rs 106.18/- |
| Tilted up | the published twelve with 0.25 added to every one | 3.00 | Rs 126.26/- |
| Tilted down | the published twelve with 0.25 taken from every one | minus 3.00 | Rs 89.30/- |
| Halved | the published twelve, each one halved | 0.00 | Rs 106.18/- |
Read the last column against the third and one fact jumps out. Every set whose twelve values total nought finishes at Rs 106.18/-, whatever route it took to get there, and the reversed set takes a route that could hardly be more different from the published one. The halved set barely leaves the neighbourhood of Rs 100/- and still finishes in the same place. Only the two sets whose totals are not nought finish anywhere else.
What does the equation for the standard process return when it is fed the published twelve driving values?
How Stochastic Differential Equations Model Financial Variables, and what does that assume?
Now the worked instance, on the standard process this guide uses throughout. The standard process is an invented quantity, written S with a time subscript and observed over one year. Its equation says the change over each instant is the level multiplied by 0.08 multiplied by the time step, plus the level multiplied by 0.20 multiplied by the Brownian increment.
| \(S_t\) | the standard process, an invented traded quantity, in rupees |
| \(\mu S_t\) | the drift coefficient, proportional to the level rather than fixed |
| \(\sigma S_t\) | the diffusion coefficient, also proportional to the level |
| \(dW_t\) | the Brownian increment over the step, under the physical measure P |
| \(S_0\) | the starting value, Rs 100/- exactly |
Both coefficients are proportional to the level, and that is a modelling choice with consequences. The choice says a move is proportional rather than absolute: a one per cent move is exactly as likely when the quantity sits at Rs 50/- as when it sits at Rs 500/-. Because both terms shrink to nothing as the level does, the quantity can never reach nought. Whether either of those is a reasonable thing to assume is covered under geometric Brownian motion.
Read the same equation in its integral form and it says the level at any time equals Rs 100/- plus the integral of 0.08 multiplied by the level against time, plus the integral of 0.20 multiplied by the level against the Brownian path. The second of those is an Ito integral. Fed the published twelve driving values, each multiplied by 0.288675 to give that month's Brownian increment, the exact solution returns twelve readings.
| Month | Driving value | Level of the standard process | Month | Driving value | Level of the standard process |
|---|---|---|---|---|---|
| 1 | minus 0.5 | Rs 97.64/- | 7 | minus 1.3 | Rs 103.56/- |
| 2 | 1.6 | Rs 107.63/- | 8 | minus 0.5 | Rs 101.12/- |
| 3 | minus 1.3 | Rs 100.35/- | 9 | minus 1.4 | Rs 93.74/- |
| 4 | minus 0.1 | Rs 100.27/- | 10 | 0.4 | Rs 96.41/- |
| 5 | 0.1 | Rs 101.35/- | 11 | 0.9 | Rs 102.06/- |
| 6 | 1.5 | Rs 111.08/- | 12 | 0.6 | Rs 106.18/- |
The high is Rs 111.08/- at month six and the low is Rs 93.74/- at month nine, so this is a year with real movement in it. The twelve driving values were constructed to total nought, so the path finishes at Rs 106.18/-, exactly the median of the distribution. The equation supplied the model and the driving values supplied this particular path, and neither of them produced the path on its own.
So how does the equation get used at all, if it never names a level? By turning every question into a question about the distribution. Ask for the chance that the quantity is below Rs 90/- in a year. Ask what spread a stress test should cover. Ask what the average of some function of the quantity comes to. Each of those has one answer, and each of them is computed from the equation without ever choosing a path. Questions of that shape are the whole modelling use, and the four steps below are how the statement gets written down in the first place.
- Name the quantity and its starting value
Say what is being modelled and where it begins. Here it is the standard process, an invented quantity, beginning at Rs 100/- exactly.
Without this, the two accumulations in the integral form have nothing to accumulate on to.
- Write the predictable change
Decide what the drift coefficient is and whether it depends on the level. Here it is 8 per cent a year of the current level, so it is proportional.
A drift that is a fixed number of rupees a year is a different model, not a rounder version of this one.
- Write the random change
Decide what the diffusion coefficient is and whether it too depends on the level. Here it is 20 per cent a year of the current level.
Proportional here means a percentage move is as likely at Rs 50/- as at Rs 500/-, which is an assumption.
- State the measure the whole statement is made under
Say which rule is being used to weight outcomes, P or Q. This is the step that gets skipped, and skipping it makes the previous three unusable.
The drift of 8 per cent is a statement under P. Under Q it would be a different number entirely.
The random term of the standard process is proportional to the level. What does that assume?
What has to be specified before the equation means anything?
An equation with symbols in it is not yet a model. The specificationThe complete set of things that must be fixed before an equation determines anything at all. is a five item list, and every one of the five is load bearing. With any one of them missing, the equation determines nothing that can be computed, compared or checked.
The five are the starting value, the drift, the volatility, the horizon and the measure. The first three carry numbers that appear inside the equation, so they are the ones everybody remembers. The last two are the ones that get lost, and losing them is quiet: nothing looks missing, the equation still reads sensibly, and the answers it produces are simply not answers to anything.
The horizon comes first. The spread of the quantity depends entirely on the length of the wait, so without a stated stretch of time there is no distribution to speak of. Over one year the standard deviation of the logarithm is 0.20. The standard deviation grows with the square root of the elapsed time, so over four years it is 0.40. An equation without a horizon is a shape without a size.
Then the measure. The measure is the item to treat most seriously, and here is why in one line. The drift of 8 per cent is a statement about how outcomes are weighted under the physical measure P. Reweight the same outcomes under the risk-neutral measure Q and the drift becomes a different number. The volatility and the paths themselves are untouched. Quoting a drift without naming the measure is like quoting a distance without naming the unit: the number is real, and it means nothing until somebody says what it is a number of.
An equation gives a starting value, a drift and a volatility, and nothing else. What is missing?
What goes wrong when the equation is divided through by the time step?
The error that gets made, and what it costs
Reading the notation literally and concluding that the quantity has a rate of change. The quantity has no rate of change at any point, and the operation that appears to produce one is the most natural move in the world: divide both sides by the time step and read off a derivative. Dividing through is legal in every ordinary equation of change. There is no warning at the boundary where it stops being legal.
Dividing the drift term by the time step does nothing bad. The drift term was multiplied by the step, so dividing gives back a plain number: for the standard process at Rs 100/-, exactly 8.000000 rupees a year, and that figure is the same however finely the year is cut. Dividing the random term by the step is where the wheels come off. The Brownian increment over a step is of the order of the square root of the step, so dividing by the step leaves one over the square root of the step. At twelve steps in the year that ratio is 3.464102 for a driving value of one. At 252 steps it is 15.874508. At 25,200 steps it is 158.745079. The ratio is not converging on anything at all. Instead it grows without bound, and it grows without bound for every path, not for unlucky ones.
The cost is a quantity that has no limit, produced by an operation that is valid everywhere else, in a calculation that raises no error. Nothing divides by nought, nothing overflows, no square root goes negative. The tell is precise and worth memorising: when refining the time grid makes a quantity larger rather than steadier, a random increment has been divided by a time step somewhere, and no amount of extra steps will fix it because extra steps are what is causing it.
| \(\Delta W\) | the Brownian increment across one step, of the order of the square root of the step |
| \(\Delta t\) | the length of one step, driven toward nought |
| \(\mu S\) | the drift amount at a level of Rs 100/-, being 8.000000 rupees a year |
| \(\infty\) | no limit at all, rather than a large number |
Somebody divides the equation through by the time step to get a rate of change. What happens?
How does somebody building a model actually use this?
Almost nobody rederives one of these equations from the definition. The work is reading a model statement somebody hands over, or writing one, and knowing what makes it checkable. A quantitative researcher reading somebody else's specification does that, so does a risk reviewer looking at a stress test, and so does anybody validating code against the closed form it claims to implement.
The move that pays is asking for the five item specification before asking anything about the output. A model statement missing the measure is not a model that needs a small correction; it is a model whose headline parameter has no meaning yet, and every number downstream of it inherits that. Ask for the horizon too. A spread quoted with no stretch of time attached is a shape with no size, and comparing two of them is comparing nothing.
The second habit is about outputs. When a model built on one of these equations reports a single number as its answer, ask which summary of the distribution that number is. The average, the median and the peak of the density are three different numbers here: Rs 108.33/-, Rs 106.18/- and Rs 102.02/-, all from the same equation and the same year. A model that reports one number without saying which summary it is has thrown away the only thing the equation actually determined.
The third is the refinement test, and it costs one extra run. Take whatever quantity is in doubt and recompute it on a grid twice as fine. A quantity that settles is a quantity the model determines. A quantity that grows steadily as the grid refines has a random increment divided by a time step somewhere in its lineage, and the growth is the signature rather than a bug in the code. Distinguishing those two behaviours by one extra run is cheaper than reading the code.
No authority anywhere sets the form of an equation, no regulator publishes a drift and no market convention alters what an integral converges to. The result is a statement about paths rather than about anything traded, so it holds identically everywhere and nowhere in particular.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for statements of the stochastic differential equation and its integral form | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Ito | The Ito integral, the construction that makes the second accumulation well defined | named in the text only |
| Hull, Shreve and Wilmott | Standard texts on stochastic differential equations and their use in finance | print editions |
The standard process and its four parameters are invented.
Educational material. Not advice on any investment, tax, budget or market position.
