Drift and Diffusion: The Two Terms of an SDE
The drift term contributes a predictable amount that grows in step with the time elapsed. The diffusion term contributes a random amount that grows only with the square root of that time. Because the two grow at different rates, the random term is much the larger over short horizons, the predictable one catches up only after six and a quarter years, and averaging removes the random term entirely.
A stochastic differential equation (SDE) has already been met whole. Taking it apart, holding each half up on its own, and then setting the two back beside each other answers the question a reader always arrives with and rarely asks out loud: which of these two is the real one? The honest reply is that the question is badly posed, but the arithmetic that shows why is short, exact, and worth carrying. At any horizon a reader cares to name, each term has a size that can be stated exactly, and one of the two dominates.
One warning before the mechanics. The vocabulary that has grown up around these two terms does most of the damage. Calling one of them the trend and the other the noise sounds like a description and is in fact a ranking, and it is a ranking nobody checked. Every misreading of these two terms starts by ranking them by their names instead of by their sizes. The sizes are a two line computation, and they run the other way round from the names for a very long time.
Why does one equation need two separate terms?
The equation for the standard process, an invented traded quantity, starts at Rs 100/- with a drift of 8 per cent a year and a volatility of 20 per cent a year. The relative form is the easiest to read: each side is a proportional change rather than a change in rupees, and both coefficients are then the plain percentages just stated.
| \(S_t\) | the standard process at time \(t\) |
| \(\mu\) | the drift coefficient, 0.08 a year under the physical measure P |
| \(\sigma\) | the volatility coefficient, 0.20 a year |
| \(W_t\) | standard Brownian motion under the physical measure P |
| \(dt\) | the time symbol, standing for a term of the order of the step |
| \(dW_t\) | the increment symbol, standing for a term of the order of the square root of the step |
Two terms, and they are not two halves of one idea. The two answer different questions. The drift termThe predictable part of the change, proportional to the time elapsed rather than to its square root. answers where the process is being carried, and it gives the same reply on every occasion the process is run. The diffusion termThe random part of the change, proportional to the square root of the time elapsed. Its average is nil at every horizon. answers how far the process is thrown about while it is being carried there, and it gives a different reply every time. One is a destination, the other is a spread around whatever the destination happens to be.
Here is the whole of the difference, and it is smaller than it looks. Both terms accumulate over a horizon. The time symbol adds up to elapsed time, so accumulating the first term gives the horizon itself. A Brownian increment scales with the square root of the step, a scaling established separately, so accumulating the second term gives something whose typical size is the square root of the horizon. So the first term carries the horizon to the power one and the second carries it to the power one half. One exponent, one half against one, is the entire structural difference between the two terms of this equation, and everything that follows is a consequence of it.
What does the drift term contribute?
Taking the first term on its own, with the second set to nothing for a moment, what is left is an ordinary equation of the kind covered under ordinary differential equations, and it does exactly what an ordinary equation does: it carries the process along a curve with no surprises in it. Accumulating the drift term over a horizon gives the drift coefficient multiplied by the horizon. The product of the two is the whole contribution.
| \(T\) | the horizon, being one year unless a passage states otherwise and says why |
| \(\int_0^{T}\) | accumulation from the start of the horizon to its end |
| \(\mu T\) | the accumulated proportional contribution of the drift term over that horizon |
Two properties follow immediately and both matter later. The first is proportionality. Doubling the horizon doubles the contribution: 0.080000 over a year becomes 0.160000 over two, 0.400000 over five, 0.800000 over ten. The second is that the contribution is a single number rather than a range. There is no spread around it, no better case and worse case, nothing to average. The contribution is the same 0.080000 whether the process spent the year climbing or spent it falling.
Anchor it in rupees so the number is not floating. Starting at Rs 100/-, a drift of 0.08 accumulated over one year is exactly the exponent that produces the expected value at the horizon of Rs 108.33/-, being Rs 100/- multiplied by e to the 0.08. The drift term is the part of the equation that decides the average, and the figure 0.080000 and the figure Rs 108.33/- are the same statement written twice.
An everyday version helps here, and it comes from measurement rather than from any business. A person steps on a bathroom scale each morning. Whatever slow change is genuinely happening to the body is a drift: it accumulates steadily, a little each day, and if it were known exactly it could be multiplied by the number of days to give the honest total. The steady accumulation is the drift term, and it holds no wobble at all. The wobble is the other term, taken up next.
What does the diffusion term contribute?
Now with the drift set to nothing and the second term kept, accumulating it over a horizon does something different: what comes out is not a number that can be named in advance. The accumulation is the volatility coefficient multiplied by where the Brownian path has got to, and where the Brownian path has got to is not settled until the horizon has actually passed. So this term does not contribute a value. The diffusion term contributes a spread, and the only figure quotable in advance is how wide that spread typically is.
| \(\sigma W_T\) | the accumulated proportional contribution of the diffusion term over the horizon |
| \(\mathbb{E}[\;\cdot\;]\) | the average taken across every occasion the process could run |
| \(\operatorname{sd}(\;\cdot\;)\) | the standard deviation, being the typical distance of the quantity from its average |
| \(\sqrt{T}\) | the square root of the horizon, being the scaling of a Brownian increment |
Readers routinely collapse the three statements in that line. Read them as three separate facts. The first says the diffusion term accumulates to something. The second says that something has an average of nil. The third says it is nevertheless typically nowhere near nil on any particular occasion. A quantity can be centred on nothing and still be the biggest thing in the room, and this one usually is.
Put the numbers on it. With volatilityThe coefficient multiplying the random term, 20 per cent a year for the standard process. The coefficient sets the width of the spread, not its direction. at 0.20 a year, the typical accumulated contribution over one year is 0.20 multiplied by the square root of 1, giving 0.200000. Over four years it is 0.20 multiplied by 2, giving 0.400000. Over nine years it is 0.20 multiplied by 3, giving 0.600000. Notice what that pattern does: quadrupling the horizon only doubles the contribution. The drift term behaved in exactly the opposite way. The scalingHow a term grows as the horizon lengthens. The two terms of this equation grow at different rates, and that is what separates them. is lazier, and that laziness is what eventually loses it the race.
The second half of the bathroom scale example is the diffusion term. Every reading carries a wobble that has nothing to do with the body: where the person stood on the plate, whether the floor is level, what the machine rounded to. Across many mornings that wobble averages out to nothing, exactly the property the second line of the formula block states. But on any single morning the wobble is comfortably larger than one day of genuine change. One morning's reading therefore says almost nothing and a season of readings says a great deal. The contrast between the two halves is the whole of the matter, and the rest is putting numbers on them.
What power of the horizon does the diffusion term carry?
How do the two scale as the horizon changes?
With the two accumulations set side by side and the horizon left to vary, one is a straight line through the origin with slope 0.08. The other is a curve that starts vertically and then flattens, reaching 0.20 at one year. At every horizon there is a ratio between them. The ratio identifies which term a reader looking at that horizon is mostly looking at, and no other single number tells as much.
| \(\sigma\sqrt{T}\) | the typical accumulated contribution of the diffusion term |
| \(\mu T\) | the accumulated contribution of the drift term |
| \(\sigma/\mu\) | the ratio of the two coefficients, 0.20 over 0.08, being 2.5 for the standard process |
| \(T^{-1/2}\) | one over the square root of the horizon, the factor by which the ratio decays |
The one expression answers every question of the form which term is bigger. Feed it a horizon and read off the answer. Over one month, the square root of one twelfth is 0.288675, so the ratio is 2.5 divided by 0.288675, or 8.660254. Over one year the square root of 1 is 1, so the ratio is 2.5 exactly. Over ten years the square root of 10 is 3.162278, so the ratio is 0.790569 and the ordering has flipped. The ratio decays with the square root of the horizon. No useful decay is slower, and that slowness is why the random term stays on top for so long.
| Horizon | Drift term | Diffusion term | Ratio, random to predictable |
|---|---|---|---|
| One month | 0.006667 | 0.057735 | 8.660254 |
| Three months | 0.020000 | 0.100000 | 5.000000 |
| One year | 0.080000 | 0.200000 | 2.500000 |
| Two years | 0.160000 | 0.282843 | 1.767767 |
| Five years | 0.400000 | 0.447214 | 1.118034 |
| Six and a quarter years | 0.500000 | 0.500000 | 1.000000 |
| Ten years | 0.800000 | 0.632456 | 0.790569 |
The last column contains the finding, so read down it slowly. At every horizon a reader would ordinarily look at, a month, a quarter, a year, two years, the random term is the larger of the two, and at the shortest of those it is larger by a factor approaching nine. The ratio is still above one at five years. The ratio reaches one only at six and a quarter.
Over ten years, which term is larger?
Drift vs Diffusion Term: which one dominates?
The derivation of the crossing point is one line, and it shows that the answer depends on both coefficients rather than on the volatility alone. Derive it rather than quote it. Set the two accumulated contributions equal to each other and solve for the horizon. The square roots collapse and what falls out is a clean expression: the variance rate divided by the square of the drift.
| \(T^{*}\) | the crossing horizon, being the elapsed time at which the two contributions are equal |
| \(\sigma^{2}\) | the variance rate, 0.04 exactly for the standard process |
| \(\mu^{2}\) | the square of the drift coefficient, 0.0064 exactly |
| \(6.25\) | the crossing horizon in years, at which both terms contribute 0.500000 |
A crossing point that only works one way round is a sign of an algebra slip, so check it in both directions before trusting it. Forwards: the drift contributes 0.08 multiplied by 6.25, giving 0.500000. Backwards: the diffusion contributes 0.20 multiplied by the square root of 6.25, and the square root of 6.25 is 2.5, so that comes to 0.500000 as well. The two agree exactly, and they agree at a round number only because the four parameters of the standard process were chosen so that they would.
The shape of that expression carries a lesson of its own. The crossover horizonThe elapsed time at which the two terms of the equation are equal in size. Beyond it the predictable term is the larger. is not a fixed feature of the mathematics. The crossover moves with the two coefficients, and it moves fast: the drift enters squared. Halving the drift coefficient pushes the crossing out by a factor of four. Halving the volatility pulls it in by a factor of four. So there is no universal statement of the form the random term matters for the first N years. There is only a computation, done with the two coefficients actually in use. The universal part is the shape of the answer, the variance rate over the drift squared, and not any particular number of years.
At what horizon do the two terms become equal?
The worked instance, over one month
The shortest horizon in the table is the one that makes the point hardest, so it is worth working through in full. One month is one twelfth of a year, or 0.083333. The drift term contributes 0.08 multiplied by 0.083333, giving 0.006667. The diffusion term contributes 0.20 multiplied by the square root of 0.083333, and that square root is 0.288675, so the contribution is 0.057735. Dividing the second by the first gives 8.660254.
The figure deserves a moment. Everything else rests on it. Over a single month the random term of this equation is more than eight and a half times the size of the predictable one. Not slightly larger. Not comparable. Larger by nearly an order of magnitude. And a month is not an unusually short horizon; it is the horizon at which the locked path for the standard process takes each of its twelve steps.
Over one month, which term is larger: the predictable one or the random one?
Move the horizon and watch which term is on top
Both bars are computed from the two accumulation formulas above, never sampled, so the reading is identical on every reload. Start at one year, the setting that reproduces the worked instance exactly, then walk the horizon out and find the point where the bars swap places.
| Named horizon | Drift | Diffusion | Ratio |
|---|---|---|---|
| One month | 0.006667 | 0.057735 | 8.660254 |
| One year | 0.080000 | 0.200000 | 2.500000 |
| 6.25 years | 0.500000 | 0.500000 | 1.000000 |
| Ten years | 0.800000 | 0.632456 | 0.790569 |
Which term survives being averaged?
Everything so far has compared the two terms on a single occasion. Now average across occasions. One of the two terms simply is not there any more, and the comparison changes character completely.
| \(\mathbb{E}[\;\cdot\;]\) | the average across every occasion the process could run, taken under the physical measure P |
| \(\mu T\) | the drift contribution, unchanged by averaging because it is the same on every occasion |
| \(0\) | the average of the diffusion contribution, exactly nil at every horizon without exception |
That the accumulated diffusion term averages to zeroHas an average of exactly nil across occasions. The accumulated random term does this at every horizon, short or long. is established separately, so it arrives as a tool rather than as a claim to be proved. The consequence is what matters, and the consequence is structural. Every argument in this subject that is phrased in terms of averages sees one term of a two term equation, and it sees the predictable one. The narrowing is not a simplification anybody chose for convenience. Averaging simply does that.
The narrowing can be watched happening without any randomness at all, using the deterministic construction built for the standard process. Slicing the probability scale into twelve equal pieces and taking the value at the middle of each, then multiplying each by the volatility coefficient, gives twelve possible one year diffusion contributions, running from minus 0.346333 up to plus 0.346333. The twelve are symmetric about nothing by construction, so their average is exactly 0.000000. The twelve drift contributions, by contrast, are twelve copies of 0.080000, so their average is 0.080000. Same twelve occasions, two entirely different fates.
There is a trap sitting immediately beside this result, and it is worth naming before anybody walks into it. A term whose average is nothing is not a term that contributes nothing. The twelve numbers above have an average of exactly nil and a typical size of about 0.200000, and both of those statements are true at once. Averaging is a lossy operation: it throws away the spread and keeps the centre. If the question is about the centre, that loss costs nothing. If the question is about how far the process might land from the centre, averaging has just deleted the only term that could have answered it.
Which term disappears under averaging?
Over one month, what is the ratio between the random contribution and the predictable one?
What does the signal and noise framing cost?
Now the failure, and it is a failure of vocabulary rather than of arithmetic. A fault in the words survives far better than a fault in the sums. Nobody who calls the drift the signal and the diffusion the noise has done a calculation and got it wrong. Nobody in that position has done a calculation at all. The words arrived from a different subject, where a signal is the quantity of interest and noise is the small contamination sitting on top of it, and they were carried across without anybody checking whether the sizes matched the names.
The mistake: ranking the two terms by their names
Somebody describes the model in a note: the process drifts upward at 8 per cent a year, with volatility of 20 per cent a year as the noise around that trend. Nothing in that sentence is false. Every reader of it, including the writer, now carries a picture in which the first quantity is the substance and the second is a haze around it. Over a one year horizon the substance is 0.080000 and the haze is 0.200000, so the haze is two and a half times the substance. Over a month the substance is 0.006667 and the haze is 0.057735, and the haze is 8.660254 times the substance. The signal and noiseA framing borrowed from measurement, in which one quantity is the substance and the other a small contamination. Applied to these two terms it ranks them backwards at ordinary horizons. framing has the two the wrong way round at every horizon shorter than six and a quarter years.
The cost is a specific one and it is not abstract. Somebody looks at a completed period, sees the process ended up somewhere unexpected, and asks how much of that was the model working. With the two terms ranked the wrong way round, the answer that feels right is that the drift explains most of it and the rest was a small disturbance. The arithmetic says the opposite for any horizon under six years, and no arithmetic anywhere in the note corrects the impression. None was done. A framing that is never written as a number is never checked as one.
Somebody calls the drift the signal and the diffusion the noise. Over a one year horizon, is that ranking right?
Can either term be read off an observed path?
Here is the question every reader eventually asks. If the process has been running and has been watched throughout, surely the two terms can be seen in what it did? The answer is no, and the reason is worth having in full: the plausible half answers are the dangerous part.
The drift comes first, taken on the locked path for the standard process, the constructed twelve step path that starts at Rs 100/- and ends at Rs 106.18/-. Its total change in logarithmic terms over the year is 0.060000 exactly. The question is then what drift coefficient produced it. The change in the logarithm is not the drift term of the equation by itself. Converting between the two brings in the half variance correction, supplied by the chain rule for this subject.
| \(\ln(S_T/S_0)\) | the observed change in the logarithm of the process across the horizon |
| \(\tfrac{1}{2}\sigma^{2}\) | the half variance correction, 0.02 exactly for the standard process |
| \(W_T\) | where the Brownian path finished, never observed separately from the process |
| \(0.060000\) | the observed total on the locked path, being the logarithm of Rs 106.18/- over Rs 100/- |
Work three cases and watch the same endpoint come out of each. Hold the volatility at 0.20 and try a drift of 0.04: the predictable part is 0.020000, so the random part must be 0.040000, needing the Brownian path to have finished at plus 0.2. Try a drift of 0.08: the predictable part is 0.060000 and the random part is 0.000000, so the path finished at nil. Try a drift of 0.12: the predictable part is 0.100000 and the random part is minus 0.040000, so the path finished at minus 0.2. Three different drift coefficients, three entirely unremarkable Brownian endings, one identical observed path ending at Rs 106.18/-.
And the endings are genuinely unremarkable. Their ordinariness is what closes the argument. Over one year the Brownian path has a typical size of 1, so finishing at plus or minus 0.2 is a fifth of a typical move. Nothing about the observed path makes any of the three drifts preferable to the others. Watching for longer helps far less than intuition suggests: doubling the horizon doubles the drift contribution but the spread that hides it grows too, and the separation improves only with the square root of the time spent watching.
The second term is not much better off, though the reason is different. A reading of the volatility can be recovered from how much the path jiggled, but the reading obtained depends on how often the path was sampled. Totalling the squared movements of the logarithm of the locked path at various partitions, and converting each total back to an implied volatility, gives the following. At the twelve step partition the total is 0.040300 and the implied figure is 0.200749. At six steps it is 0.041333 and 0.203306. At four steps it is 0.054700 and 0.233880. At two steps it is 0.013067 and 0.114310. At a single step it is 0.003600 and 0.060000, the path happening to finish near where a single step would draw it.
| Partition of the same year | Total of squared log movements | Implied volatility |
|---|---|---|
| One step | 0.003600 | 0.060000 |
| Two steps | 0.013067 | 0.114310 |
| Four steps | 0.054700 | 0.233880 |
| Six steps | 0.041333 | 0.203306 |
| Twelve steps | 0.040300 | 0.200749 |
One observed pathA single realised run of the process. A single run is one draw of the random term, and so cannot separate the two terms on its own., five partitions, five different answers, and the sequence is not even tidy: it rises, falls, rises again. Anybody quoting a single one of those figures as the volatility of the path has quoted the partition they happened to choose as much as anything about the process. Neither term can be read off a single observed path, however long that path runs, and saying so plainly is more useful than a method that pretends otherwise. Doing this properly is a statistical question with its own methods and its own failure modes, and it is covered separately.
Can the drift be read off a single observed path?
How does anybody use this in practice?
Not as theory, in the end. The ratio set out above is a sanity check somebody runs before making a judgement about a completed period, and it takes about ten seconds.
Suppose a reviewer looks back at one year of the standard process and finds it rose by 15 per cent, against a model drift of 8 per cent. The immediate temptation is to say the model was too low, or that something changed. Run the check instead. The drift contribution over that year was 0.080000. The gap between what happened and what the drift explains is 0.070000. A typical diffusion contribution over the same year is 0.200000. So the entire unexplained portion is 0.070000 divided by 0.200000, or 0.35 of one typical random contribution. Rather than being a surprise, it is a third of an ordinary wobble.
The same check runs the other way, and this is where the discipline earns its keep. If a year came in 7 percentage points below the drift instead of above it, the arithmetic is identical and the conclusion is identical: a third of an ordinary wobble, in the other direction. The check does not say what happened, it says how much of what happened is even worth explaining, and at ordinary horizons the answer is much less than it looks.
Three habits fall out of it, and none needs software. The first is to name the horizon before comparing the two terms. The comparison is meaningless without one. The second is to quote a move as a multiple of the typical diffusion contribution at that horizon rather than in raw percentage points, a scale that already accounts for how long the process was watched. The third is to treat any period shorter than a few years as a stretch in which the random term is running the show, and to be suspicious of any story about it that is built entirely on the drift. Every one of these is a habit of arithmetic rather than a rule of thumb, and each of them is checkable on the numbers in the table above.
The two terms of the equation have been opened and their sizes compared. Estimating either term from observations is a statistical question with its own methods, covered separately. Solving the equation turns out not to mean the single thing a reader expects, and that too is covered separately. No regulator anywhere sets the form of an equation, and the mathematics is the same everywhere it is used.
Sources and how they were used
| Source | What it was used for | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository named for the scaling of the two terms of a stochastic differential equation and its consequences | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Ito | The integral and the chain rule that carry his name, both used rather than rebuilt | named in the text only |
| Hull, Shreve and Wilmott | Standard texts on stochastic calculus and derivative pricing, the source of the notation and the ordering used | named in the text only |
The standard process and its locked path are invented.
Educational material. Not advice on any investment, tax, budget or market position.
