The Ornstein-Uhlenbeck Process: Mean Reversion Formalised
The Ornstein-Uhlenbeck process is mean reversion written as an equation: the change in a quantity is a speed multiplied by the gap between where it stands and a fixed long-run level, plus randomness of constant size. The process differs from the standard price process in one term and in nothing else. Its spread does not grow without limit, it settles, here at exactly 0.010000.
Before anything else, one plain statement: this process and the Vasicek Model are one object under two names. The equation set out in this guide is the equation a reader has already met if they have worked through the treatments of volatility, where it appeared with the label Vasicek, 1977, attached to it. Same equation, same three parameters, same solution, same settling spread. The two names come from two different histories, one in physics in 1930 and one in rate modelling in 1977, and neither is a variant of the other. Meeting the same object twice under two names is the ordinary experience of this subject, and a reader who is left to discover it alone has been failed by the sequence rather than tested by it.
For a reader who has been through the volatility material, none of this is new. The same equation arrived there as a named model, and here it stands as a process in its own right, with its solution written down, its distribution stated, and the one number that the whole of rate modelling turns on, the spread it settles at, computed to six decimal places.
What is the Ornstein-Uhlenbeck process?
Mean reversion was established as an idea earlier in this sequence, and here it turns into arithmetic. A quantity is pulled back. The quantity does not wander off and stay away. When it sits above a level something pushes it down, when it sits below that level something lifts it up, and the further away it sits the harder the push. Mean reversion in words is exactly that push, and the ground is settled already.
The Ornstein-Uhlenbeck processMean reversion written as an equation with randomness of constant size. is that description turned into a single line. The process has three parameters and one source of randomness. The three parameters are a level to be pulled toward, a speed at which the pulling happens, and a size for the randomness. Nothing else is in it.
| \(r_t\) | the short rate at time \(t\), the quantity being modelled |
| \(\theta\) | the long-run mean, the fixed level the quantity is pulled toward |
| \(\kappa\) | the speed of mean reversion, in units of one over a year |
| \(\sigma_r\) | the rate volatility, the size of the randomness, in absolute terms |
| \(W_t\) | standard Brownian motion under the physical measure P |
Everything in this guide comes out of the first term, so the first term repays slow reading. The term is a speed multiplied by a gap. When the quantity sits below the long-run level the gap is positive and the term pushes upward. When it sits above, the gap is negative and the term pushes downward. When it sits exactly on the level the gap is nought and the term contributes nothing at all. The level is not a magnet the quantity is stuck to. The level is the place where the pull switches off.
Here is the everyday version, and it is a measurement example rather than anything to do with money. A room with a thermostat set to twenty degrees is not a room that stays at twenty degrees. The door opens, the sun moves, someone switches on an oven, and the reading wanders. The thermostat pushes back in proportion to how far off the reading has drifted: a long way off means a hard push, close to the setting means barely any push at all, and exactly on the setting means the heater does nothing. The temperature in that room is never still and never runs away either. An Ornstein-Uhlenbeck process is a thermostat and a disturbance, written down.
Contrast that with a paper boat let go on a river. Nothing is pulling the boat back to where it was released. Every push it receives is kept, and the boat gets further from its starting point as the day goes on. A price process has that shape, and the two shapes are set against each other below.
The locked parameters of this process
Every number below is a computed consequence of four invented parameters, published in this subject area so that its treatments reconcile with one another rather than each making up its own. The four parameters were chosen rather than measured, and a measured set would move every figure below.
| Parameter | Symbol | Value | What it does |
|---|---|---|---|
| Level at time nought | \(r_0\) | 5 per cent | where the quantity starts, the same 5 per cent as the constant risk-free rate of this subject area, deliberately |
| Long-run mean | \(\theta\) | 6 per cent | the level the pull works toward |
| Speed of mean reversion | \(\kappa\) | 0.5 a year | how hard the pull is per unit of gap |
| Rate volatility | \(\sigma_r\) | 1 percentage point | the size of the randomness, in absolute terms |
| Half-life of the pull | \(t_{1/2}\) | 1.386294 years | the time in which the gap to the level halves, computed rather than assumed |
The starting level sits one percentage point below the long-run mean, and the offset is not decoration. Every quantity in this guide then has two things happening at once: a gap closing, and a spread opening. Keeping those two separate is most of the work, and the figures allow that separation to be checked.
How does its equation differ from the standard price process?
The standard process of this subject area is a single invented traded quantity written S with a time subscript, starting at Rs 100/-, drifting at 8 per cent a year and carrying a volatility of 20 per cent a year over a one year horizon against a risk-free rate of 5 per cent.
| \(S_t\) | the standard process, the single invented traded quantity, at time \(t\) |
| \(\mu\) | the drift of the proportional change, 0.08 on the locked parameters |
| \(\sigma\) | the volatility of the proportional change, 0.20 on the locked parameters |
| \(W_t\) | standard Brownian motion under the physical measure P |
Now hold the two equations next to each other and look only at the first term of each. In the price process the drift is a number multiplied by the level of the quantity itself. In the Ornstein-Uhlenbeck process the drift is a number multiplied by the gap between the level and a fixed target. The single substitution of a gap for a level is the whole difference between the two processes.
Put numbers in and the change becomes obvious. Take the price process at Rs 100/-. Its drift is 0.08 multiplied by 100, or Rs 8/- a year. Double the level to Rs 200/- and the drift doubles to Rs 16/- a year. Halve it to Rs 50/- and the drift halves to Rs 4/- a year. A positive number multiplied by a positive level is always positive, so the drift never changes sign. Where the quantity goes, the drift follows in the same direction and in the same proportion.
Now the same exercise on the rate. At 5 per cent the drift on the gapA drift proportional to the distance from a fixed level, which is the one difference. is 0.5 multiplied by the gap of one percentage point, giving plus 0.500000 percentage points a year. At 3 per cent the gap is three percentage points and the drift is plus 1.500000 percentage points a year, three times as strong. At 6 per cent the gap is nought and the drift is exactly 0.000000. At 10 per cent the gap is minus four percentage points and the drift is minus 2.000000 percentage points a year. The drift changed sign, and the price process cannot do that at any level whatsoever.
What is the drift proportional to here, and what is it proportional to for the standard price process?
What does its solution look like?
The equation describes a change over an instant. A reader usually wants the level at a stated horizon instead, and getting it means solving the equation. The solution is available in closed form, and it comes out as something a reader can read without any calculus at all.
| \(r_0\) | the level at time nought, 5 per cent on the locked parameters |
| \(e^{-\kappa t}\) | the weight carried by the starting level, falling from one toward nought |
| \(1-e^{-\kappa t}\) | the weight carried by the long-run level, rising from nought toward one |
| \(u\) | a time earlier than \(t\), the moment at which a shock of randomness arrived |
| \(dW_u\) | the Brownian increment that arrived at that earlier moment |
The first two terms are a weighted blendThe solution's form, part starting value and part long-run level.. The second weight is defined as one minus the first, so the two weights sum to exactly one at every horizon. The sum is worth pausing on. The expected level is therefore never outside the range set by the starting value and the long-run level, and it cannot overshoot. The randomness carries no such constraint, so the quantity itself can overshoot even though its expected value cannot.
At one year the decaying weight is the exponential of minus 0.5, or 0.606531. The rising weight is therefore 0.393469. Blend 5 per cent and 6 per cent with those weights and the expected level is 5.393469 per cent. The figure 5.393469 is one of the locked figures of this subject area, and it is not a rounded quantity. The second weight is one minus the first and the two levels differ by exactly one, so 0.606531 multiplied by 5 plus 0.393469 multiplied by 6 gives 5.393469 exactly.
| Horizon | Weight on the start | Weight on the level | Sum | Expected level |
|---|---|---|---|---|
| Six months | 0.778801 | 0.221199 | 1.000000 | 5.221199 per cent |
| One year | 0.606531 | 0.393469 | 1.000000 | 5.393469 per cent |
| Two years | 0.367879 | 0.632121 | 1.000000 | 5.632121 per cent |
| Three years | 0.223130 | 0.776870 | 1.000000 | 5.776870 per cent |
| Five years | 0.082085 | 0.917915 | 1.000000 | 5.917915 per cent |
| Ten years | 0.006738 | 0.993262 | 1.000000 | 5.993262 per cent |
The fourth column reads one at every row, and it is one by construction rather than by luck. The last column sets out the whole effect of the pull on expectations: a handover, from the starting level to the long-run level, complete for practical purposes somewhere past five years and effectively finished by ten. The expected level is a handover between two fixed numbers, and the speed decides only how fast the handover runs.
The speed parameter is easiest to read through its half-lifeThe time in which the remaining gap to the long-run level halves, computed as the natural logarithm of two over the speed., the natural logarithm of two divided by the speed. At a speed of 0.5 that is 1.386294 years. In one half-life the gap of one percentage point becomes half a percentage point, putting the expected level at exactly 5.500000 per cent. In two half-lives, 2.772589 years, it is 5.750000 per cent. In three, 4.158883 years, it is 5.875000 per cent. Each halving is exact, and that exactness makes the half-life the honest way to describe a speed to somebody who has to think in years rather than in units of one over a year.
| \(t_{1/2}\) | the half-life, the time in which the remaining gap halves |
| \(\ln 2\) | the natural logarithm of two, 0.693147 |
| \(\kappa\) | the speed of mean reversion, 0.5 a year on the locked parameters |
What are the two weights in the solution at one year, and what do they sum to?
Before reading on: what does the stationary spread work out to at these parameters?
Where does its spread settle, and why?
This is the headline result. The third term of the solution, the accumulated randomnessThe part of the solution the pull does not control., is the only part the pull does not directly control, and it is where the spread lives. Its standard deviation has a closed form too.
| \(\operatorname{sd}\) | the standard deviation of the quantity at that horizon |
| \(\sigma_r\) | the rate volatility, 0.01 on the locked parameters |
| \(\kappa\) | the speed of mean reversion, 0.5 a year |
| \(t\) | the horizon in years |
Put the locked parameters in and read the numbers off. At one year the spread is 0.007951. At two years it is 0.009299. At five years 0.009966. At ten years it reads 0.010000 to six decimal places, and past that it does not move. The spread settles, and the number it settles at is exactly 0.010000.
Why exactly? Because the stationary spreadThe volatility over the square root of twice the speed, exactly 0.010000 here. is the rate volatility divided by the square root of twice the speed, and at these parameters twice the speed is exactly one, whose square root is exactly one. Dividing 0.01 by one leaves 0.01. There is nothing lucky about the arithmetic and nothing hidden in it: the parameters were chosen so that a reader can check the settling point in their head.
| \(\sigma_r\) | the rate volatility, 1 percentage point a year in absolute terms |
| \(2\kappa\) | twice the speed, exactly 1 on the locked parameters |
Now the mechanism, and it matters more than the formula. Look at how the variance itself moves rather than at the spread. Over a small interval the randomness adds a fixed amount of variance, the volatility squared, here 0.000100 a year. Over the same interval the pull removes a share of whatever variance has already built up, and that share is twice the speed multiplied by the current variance.
| \(\operatorname{Var}\) | the variance of the quantity at that horizon, the spread squared |
| \(\sigma_r^{2}\) | the variance added per year by the randomness, 0.000100 here |
| \(2\kappa\) | the fraction of accumulated variance the pull removes per year, 1 here |
The tank picture is the whole explanation and it is worth doing with the numbers. At time nought the variance is nought, so nothing drains and the variance grows at the full 0.000100 a year. By one year the variance is 0.0000632121, so the drain is carrying 0.0000632121 a year away and the net growth has fallen to 0.0000367879. By two years the variance is 0.0000864665 and the net growth is down to 0.0000135335. By five years it is 0.0000006738 and by ten years 0.0000000045. Nothing switches the growth off, it is simply outrun by the removal.
Why does the spread stop growing?
The one year spread is 0.007951. Before the horizon is pushed out on the control below: what does the square-root-of-time rule give at ten years, and what is the truth?
Push the horizon out and watch the two rules separate
Held fixed: the long-run level at 6 per cent, the speed at 0.5 a year, the rate volatility at 1 percentage point. The only thing that moves is the horizon. The dark curve is the true spread of this process and the red curve is the one year figure of 0.007951 scaled by the square root of the horizon, the rule that is exact for the standard process. The wedge between them is the overstatement, and it opens from the one year mark rightward as the control moves.
At a horizon of 1.0 years the true spread is 0.007951 and the square-root-of-time scaling also gives 0.007951, so the rule is 1.000000 times the truth. This is the one horizon at which the two agree, and they agree here because this is the figure the rule was scaled from.
What distribution does it have, and why is that unusual?
The solution is the starting level and the long-run level, both fixed numbers, plus an integral of a deterministic function against a Brownian increment. An integral of that kind is normally distributed, and adding fixed numbers to a normal quantity leaves it normal. So the whole thing is normal, at every horizon, with no approximation anywhere.
| \(\mathcal{N}\) | the normal distribution, given here as mean then variance |
| \(\theta + \left(r_0-\theta\right)e^{-\kappa t}\) | the mean, which is the weighted blend written the other way round |
| \(\sigma_r^{2}\frac{1-e^{-2\kappa t}}{2\kappa}\) | the variance, whose square root is the spread of the previous block |
Being normal at every horizonA property this process has and the standard price process does not. is a genuinely unusual property and it is easy to walk past. The standard price process is not normal at any horizon. Its logarithm is normal, making the process itself lognormal, and the difference between those two statements is the entire reason the standard process cannot go below nought while this one can.
A lognormal quantity has a wall at nought. Whatever the parameters, however long the horizon, the weight it puts below nought is exactly zero. The exponential of a real number is positive, and no real number has a negative exponential. A normal quantity has no wall anywhere. A normal quantity puts some weight on every value on the line, so a process that is normal at every horizon can reach any value at all, including negative ones.
How much weight, on these parameters? At one year the mean is 5.393469 per cent and the spread is 0.795060 percentage points, so nought sits 6.783725 spreads below the mean and the weight there is about 5.855816 in a hundred billion, which is roughly one chance in 170 billion. In the settled distribution the mean is 6 per cent and the spread is exactly 1 percentage point, so nought sits exactly six spreads below the mean and the weight is about 9.865876 in ten billion, roughly one chance in a billion. The size of that number matters far less than the fact that it is not nought. The standard price process puts exactly nought there, and no choice of parameters changes that.
Here is the everyday version, and it is again about measurement rather than money. A kitchen scale that has been knocked can read minus twenty grams with nothing on it. Nobody thinks the object has negative mass; the reading is a measurement that carries error on both sides of the truth, and a symmetric error can put a reading anywhere. A count of objects, by contrast, cannot go below nought however bad the counting is. Counting has a floor built into it. The Ornstein-Uhlenbeck process is a reading, not a count, and that is exactly why the treatment of volatility had to compute how often the reading goes below nought rather than simply declaring that it could.
The process is normal at every horizon. Which of these does normality permit?
What is it, and what is it not, in relation to the rate model?
The plain statement made at the outset is worth making again with the equations in view. The Vasicek Model of the volatility reading order is this process applied to a short rate. Not similar to it, not a special case of it, not built on top of it. The same equation, with the same three parameters and the same letters, arrived at from a different direction by different people at different times.
Uhlenbeck and Ornstein wrote it down in 1930 for the velocity of a particle being knocked about by molecules while friction dragged it back toward rest. Vasicek wrote it down in 1977 for a short rate being knocked about while something pulled it back toward a long-run level. Friction and the pull are the same term. The molecular knocking and the rate randomness are the same term. The physics and the rate modelling are two readings of one equation and neither is a version of the other.
And what is it not? The process is not a bond price. Turning a short-rate process into the price of a zero-coupon bond takes an expectation of a discount factor along the path, and that expectation is set out under the zero-coupon bond, with arithmetic of its own. The process is not a fitted model either. Every number here is a consequence of four invented parameters that were written down in advance; choosing parameters so that a model agrees with something observed is a separate subject entirely, and the difference between a model and a fitted model is what keeps the two apart.
How does this process relate to the rate model of the earlier reading order?
The error that gets made, and what it costs
Assuming the spread grows with the horizon, because that is what it does for the standard process. The square-root-of-time ruleA scaling rule correct for the standard process and wrong here. says a one period spread scales up by the square root of the number of periods, and for the logarithm of the standard process that is exact rather than approximate. Carried across to a mean reverting quantity it is simply wrong, and it is wrong in a direction that always overstates.
The cost is worth working. The one year spread here is 0.007951. Multiplied by the square root of ten it gives 0.025142. The truth at ten years is 0.010000. The rule overstates the ten year spread by a factor of 2.514258, about two and a half times. At twenty years it gives 0.035556 against a truth that is still 0.010000, and it is now overstating by 3.555617 times. One quantity is still growing while the other stopped, so the error does not settle down but opens wider with every year added.
The silent assumption is that nothing removes accumulated spread. The assumption is true of a paper boat on a river and false of a room with a thermostat, and the rule carries no warning label saying which of the two is in hand. Nothing in the arithmetic fails. The multiplication is correct, the square root is correct, the answer is wrong, and the mistake was made before the arithmetic began.
| \(\ln S_T\) | the logarithm of the standard process at the horizon |
| \(\sigma\sqrt{T}\) | the square-root-of-time rule, exact for the left-hand quantity |
| \(r_T\) | this process at the horizon |
| \(\frac{\sigma_r}{\sqrt{2\kappa}}\) | the ceiling this process approaches and does not pass |
Someone scales a rate's one year spread by the square root of ten. Which assumption have they made?
What does a settling spread change about the width of a range?
Everything above is a statement about a distribution. Here is where somebody actually meets it. Anyone who has to say something about where a rate might sit at a stated horizon is, whether they say so or not, quoting a width. The two rules give two widths, and past a year or so they are not close.
Take a two spread band on either side of the expected level, a common way of stating a range without claiming any precision about it. At one year the expected level is 5.393469 per cent and the spread is 0.795060 percentage points, so the band runs from 3.803349 to 6.983590 per cent, a width of 3.180240 percentage points. One year is where the rule was scaled from, so the two rules agree exactly here.
Push it to ten years. The expected level is 5.993262 per cent and the true spread is 0.999977 percentage points, so the true band runs from 3.993307 to 7.993217 per cent, a width of 3.999909 percentage points. Under the square-root-of-time rule the spread would be 2.514201 percentage points, and the band would run from 0.964860 to 11.021664 per cent, a width of 10.056803 percentage points. The same model, the same horizon, and a range two and a half times as wide from one substituted scaling rule.
| Ten year band | Lower | Upper | Width |
|---|---|---|---|
| What this process actually says | 3.993307 | 7.993217 | 3.999909 |
| What the square-root-of-time rule says | 0.964860 | 11.021664 | 10.056803 |
| The rule as a multiple of the truth | 2.514258 |
Consider what each band would do to somebody thinking about a long commitment at a rate, a household weighing a twenty year borrowing against a ten year one, or anybody sizing how much a rate could move against them over a decade. The narrower band says the quantity stays inside a window of about four percentage points. The wider one says it might be near nought or above eleven. The two bands are not two shadings of one answer, they are two different pictures of what the world can do, and only one of them is what the model set out here states. A scaling rule quietly borrowed from a different process reaches all the way into what somebody decides.
Two cautions follow from the object being a model and not a rate. Neither band is a statement about any rate anywhere; both are computed from four invented parameters. And a band drawn two spreads either side of a mean is a convention for showing width, not a claim about how often anything falls inside it. The comparison between the two widths is exact, and that comparison is the point. The widths themselves depend entirely on parameters somebody chose.
Where this holds, and where the rules would come in
The mathematics here is universal. A process, its solution and the spread it settles at are not matters of jurisdiction. Day count conventions, quotation conventions and the way any rate is actually stated are jurisdictional and belong to a later reading order. Conduct duties apply to what anyone does with a model in any particular place and are settled elsewhere.
References
| Source | Document | Where |
|---|---|---|
| arXiv, Quantitative Finance | Preprints on mean reverting processes and short-rate modelling | arxiv.org |
| Social Science Research Network | Working papers on the Ornstein-Uhlenbeck process in rate modelling | ssrn.com |
| Uhlenbeck and Ornstein, 1930 | On the theory of the Brownian motion, the paper the process is named for | Physical Review |
| Vasicek, 1977 | An equilibrium characterisation of the term structure, the same equation applied to a short rate | Journal of Financial Economics |
The standard price process and the locked rate parameters are invented.
Educational material. Not advice on any investment, tax, budget or market position.
