Vasicek Model vs CIR Model: Two Ways to Hold a Rate
Both models pull a rate toward a long-run level at a stated speed, and they agree on that pull term exactly. The two differ in one term only. The Vasicek randomness is the same size at every level. The Cox, Ingersoll and Ross (CIR) randomness shrinks as the rate approaches nought. The usual criticism is that the first can go negative, and at the settings used here that probability is about six in a trillion.
Two models, one difference. Everything said about how the Vasicek model and the CIR model behave, what shape their outcomes take, whether either can produce a negative number and which of the two is harder to work with, comes out of a single term written two different ways. With that term straight, the rest is bookkeeping.
Both models are defined completely before either is set against the other, and the contrast that follows comes with arithmetic rather than with received opinion.
One warning belongs before the definitions. The most repeated sentence about the Vasicek model is that its rate can go negative, and almost nobody who repeats it has computed how likely that is at the settings in front of them. Being told that a floor can be breached, with no number attached, is like being told a bridge can collapse with no mention of the load that does it. The statement is true. On its own it decides nothing. A number is what turns the statement into a decision.
The settings every figure below is built from
Every figure below is built from one set of invented parameters, held fixed so that each number can be checked against the last. The short rateThe rate for borrowing over the next instant of time. Both of these models describe that single quantity and nothing else. starts at 5 per cent, it is pulled toward 6 per cent, the pull acts at a speed of 0.5 a year, and the size of the randomness at the starting level is one percentage point a year. The horizon is one year unless a longer one is named.
| Setting | Symbol | Value | Read it as |
|---|---|---|---|
| Starting level of the rate | r at time zero | 5.000000 per cent | where the rate sits today |
| Long-run level | theta | 6.000000 per cent | where the pull is directed |
| Speed of the pull | kappa | 0.500000 a year | the fraction of the remaining gap closed per year |
| Size of the randomness at the starting level | sigma with subscript r | 1.000000 points | the one-year spread contributed per year, in percentage points |
| Horizon | T | 1.000000 year | the period every distribution below is measured over |
The five settings were chosen so that the results land on figures a reader can check: the gap to close is exactly one percentage point, the speed of 0.5 gives a half-life of 1.386294 years, and the settled spread of the rate works out to exactly 1.000000 percentage points. Each model below is fed these same five numbers, so every difference between the two is caused by the models and not by the inputs.
What is the Vasicek model, defined from scratch?
The Vasicek model, named for Vasicek and published in 1977, says that the short rate changes by two things added together. The first is a pull toward a long-run level, and the size of that pull at any moment is the speed multiplied by however far the rate currently sits from that level. The second is a random push whose size does not depend on the level at all. The two terms are the whole model. There is nothing else in it.
| \(r_t\) | the short rate at time \(t\), as a decimal, which is the quantity being modelled |
| \(\kappa\) | the speed of the pull, 0.5 a year here |
| \(\theta\) | the long-run level the pull is directed at, 0.06 here |
| \(\sigma_r\) | the size of the randomness, 0.01 a year here, in absolute terms rather than as a percentage of the rate |
| \(W_t\) | standard Brownian motion under the physical measure \(\mathbb{P}\), the source of the randomness |
| \(dt\) | an instant of time, in years |
The two terms do different jobs and they are the only two moving parts here, so each one repays reading on its own. The pull term is what mean reversionBeing drawn back toward a stated level at a stated speed. The further a quantity strays, the harder it is pulled back. means: the further the rate strays, the harder it is pulled back, in exact proportion to the distance. At the starting level of 5 per cent the gap to close is one percentage point, so the pull contributes 0.5 multiplied by 0.01. The product is 0.005 a year, or 0.500000 percentage points a year of upward drift. Move the rate to 7 per cent and the same arithmetic gives 0.500000 percentage points a year of downward drift instead. The pull has no preferred direction of its own; it only ever points at the long-run level.
There is a rubber band in this, and it is worth carrying. Stretch a band twice as far and it pulls back twice as hard. The proportionality is exactly what the speed multiplied by the distance is saying, and it is why the rate never simply drifts off: the further it goes, the stronger the force bringing it home. The speed of reversionHow fast the pull acts, expressed as the fraction of the remaining gap closed per year. Here it is 0.5 a year. is the stiffness of the band and the long-run levelThe level the pull is directed at. The rate is drawn toward that level but does not have to sit on it. Here it is 6 per cent. is the point it is anchored to.
The second term is where the argument eventually lands, and it repays close attention. The randomness has a fixed size, one percentage point a year, and that size applies whether the rate is at 12 per cent, at 5 per cent, at 0.1 per cent or at exactly nought. The Vasicek randomness never consults the level of the rate before deciding how large a push to deliver. The fixed size is a modelling choice, and it is the choice everything below turns on.
Because both terms are so simple, the model can be solved rather than only simulated, and the rate at any future date has a distribution that can be written down. The distribution is normal, with a centre that slides from the starting level toward the long-run level and a spread that grows and then settles.
| \(r_T\) | the short rate at the horizon \(T\), the quantity whose distribution is being stated |
| \(r_0\) | the starting level of the rate, 0.05 here |
| \(\theta,\ \kappa,\ \sigma_r\) | the long-run level, the speed and the size of the randomness, as defined above |
| \(\mathcal{N}(m,v)\) | the normal distribution with centre \(m\) and variance \(v\) |
| \(e^{-\kappa T}\) | the fraction of the original gap still uncovered at the horizon, 0.606531 at one year |
Put the five settings into that and the one-year answer is a centre of 5.393469 per cent and a spread of 0.795060 percentage points. The whole of the negative rate argument below is a comparison between those two figures, so both are worth holding. The centre has covered 39.3469 per cent of the one point gap in a year, and the spread has reached 79.5060 per cent of its settled value of 1.000000 percentage points. A normal distribution has no floor of any kind. The missing floor is exactly why the Vasicek rate can in principle be negative, and exactly why a number rather than an opinion is needed next.
In the Vasicek model, what decides how large the next random push is?
How fast does the pull actually close the gap?
The speed of 0.5 a year is a number, and a number on its own is hard to feel. Convert it into a duration and it becomes concrete. The half-lifeThe time the pull takes to close half of whatever gap remains, computed as the natural logarithm of two divided by the speed. of the pull is the natural logarithm of two divided by the speed. At a speed of 0.5 a year the half-life is 1.386294 years. The half-life is the time it takes to close half of whatever gap is left, and it does not depend on how big the gap was.
Start at 5 per cent with a gap of 1.000000 percentage points to the long-run level. After 1.386294 years the expected level is 5.500000 per cent, and the gap remaining is 0.500000 points. After another 1.386294 years the expected level is 5.750000 per cent and the gap is 0.250000 points. The gap halves, then halves again, and it never quite reaches nought. The pull is a geometric decay of the gap rather than a march to the long-run level, so the rate is always still on its way there.
The curve belongs to both models equally, and that is the first thing worth noticing about the comparison. The expected level at any horizon depends only on the starting level, the long-run level and the speed. The randomness does not enter it at all. The Vasicek model and the CIR model produce exactly the same expected path, to every decimal place, at every horizon. Whatever separates them, it is not where they think the rate is going.
At a speed of 0.5 a year, how long does the pull take to close half the remaining gap?
What is the CIR model, defined from scratch?
The CIR model, named for Cox, Ingersoll and Ross and published in 1985, keeps the pull term exactly as it stands and rewrites the randomness. Instead of a fixed size, the size of the random push is a parameter multiplied by the square root of the rate itself. When the rate is high the pushes are large, and when the rate is low they are small. If the rate ever reached nought the random push would be nought too, leaving only the pull. At nought the pull points firmly upward.
| \(r_t\) | the short rate at time \(t\), the same quantity as before |
| \(\kappa,\ \theta\) | the speed and the long-run level, identical to the Vasicek values, 0.5 a year and 0.06 |
| \(\sigma_r^{\,c}\) | the CIR randomness parameter, 0.044721 here. The superscript marks it as a different number from \(\sigma_r\), because it multiplies a square root rather than standing alone |
| \(\sqrt{r_t}\) | the square root of the current rate, which is the whole of the difference between the two models |
| \(W_t\) | standard Brownian motion under the physical measure \(\mathbb{P}\), as before |
There is a household version of this that makes the mechanism obvious. Think of the randomness as a hand shaking while it carries a jug. One model says the hand shakes by the same amount whatever is in the jug, so an almost empty jug still gets a full shake and can end up owing liquid it does not have. The other says the shake scales with what is in the jug, so a nearly empty jug barely trembles and an empty one does not move at all. The CIR model does not forbid the rate from reaching nought by decree; it removes the push that would take it there.
One practical problem is left. The CIR parameter is not the same number as the Vasicek one, so quoting the two models with the same figure for randomness would be comparing nothing to nothing. The fix is to match them at the level the rate actually starts from. On day one the two models then deliver randomness of identical size, and every later difference is caused by the shape of the term rather than by its calibration.
| \(\sigma_r\) | the Vasicek randomness, 0.01, which is one percentage point a year |
| \(\sigma_r^{\,c}\) | the CIR parameter that makes the two agree at the starting level, 0.044721 |
| \(r_0\) | the starting level of the rate, 0.05 |
| \(\sigma_r^{\,c}\sqrt{r_0}\) | the CIR randomness evaluated at the starting level, 0.010000 exactly by construction |
The matching also lands on a clean number. The CIR parameter is 0.044721 and its square is 0.002000 exactly. Squaring 0.01 gives 0.0001, and dividing by 0.05 gives 0.002 with no rounding anywhere. The condition keeping the CIR rate away from nought is stated in terms of that squared value, so the value is about to do real work.
The positivity conditionThe requirement, due to Feller, that the pull at nought is strong enough to overwhelm the randomness there, keeping the rate strictly above nought at all times. asks whether the pull at nought is strong enough to overwhelm the randomness near nought. The condition compares twice the speed multiplied by the long-run level against the squared CIR parameter. If the first is at least as large as the second, the rate is strictly above nought at every future time, with probability one, and not merely unlikely to be negative.
| \(2\kappa\theta\) | twice the speed multiplied by the long-run level, 0.060000 here, which measures how hard the rate is pushed up when it is near nought |
| \(\bigl(\sigma_r^{\,c}\bigr)^{2}\) | the squared CIR parameter, 0.002000 here, which measures how hard the randomness fights back near nought |
| \(\ge\) | the direction of the comparison. When it holds, the rate stays strictly above nought at all times |
At these settings the comparison is 0.060000 against 0.002000, so the condition holds by a factor of thirty and the CIR rate cannot reach nought at all. That is a different kind of statement from a small probability. Reaching nought is not merely unlikely. The set of paths reaching nought has probability nought, so the floor is a property of the model rather than a matter of how the numbers happened to fall.
What keeps the CIR rate away from nought?
What does the one term that differs actually do?
Set the two definitions side by side and count the differences. The pull is the same. The speed is the same. The long-run level is the same. The source of randomness is the same Brownian motion. One thing is written differently: whether the size of the random push is a constant or a constant multiplied by the square root of the rate. Exactly one term separates the Vasicek model from the CIR model, and it is the term controlling how big the wobble is, not the term controlling where the rate is heading.
The single change carries four consequences, and they are worth naming before drawing them. The square root creates a floor at nought under one model and not the other. The square root makes the spread of outcomes depend on the level rather than being the same everywhere. The square root bends the distribution of the rate out of the symmetric normal shape into one that leans to the right. And the square root makes the second model harder to work with. A square root termA randomness term scaled by the square root of the quantity being modelled. The scaling is what makes the CIR distribution harder to work with than a normal one. inside the randomness turns a normal distribution into a scaled non-central chi-square one, a heavier object to carry through every later calculation.
How many terms differ between the Vasicek model and the CIR model?
Before reading on: is the chance of the Vasicek rate being below nought larger at ten years or at one year?
What is the negative rate criticism worth at these parameters?
Here is the criticism, stated fairly. The Vasicek rate at any horizon is normally distributed, a normal distribution puts weight on every value on the line, and so the model assigns a strictly positive probability to the rate being below nought. The positive probability is not a bug anybody overlooked; it follows directly from the randomness being the same size at every level, including at nought where it should arguably be nothing. The criticism is structurally correct and it is not going away.
A probability that is positive can be anything from a certainty to a rounding error, and only the arithmetic says which. The computation settles it. The chance of the rate being below nought at a horizon is the normal distribution function evaluated at minus the centre divided by the spread. Both of those numbers are already established above.
| \(m_T\) | the centre of the distribution at the horizon, 0.05393469 at one year |
| \(s_T\) | the spread at the horizon, 0.00795060 at one year, which is 0.795060 percentage points |
| \(N(\cdot)\) | the standard normal distribution function, the chance of being below the value given to it |
| \(\mathbb{P}\) | the physical measure, the rule under which this chance is taken |
At one year the centre is 0.05393469 and the spread is 0.00795060, so nought sits 6.783725 spreads below the centre. Feed that to the normal distribution function and the answer is 0.000000000005855794. Written shorter, 5.856e-12, or roughly six chances in a trillion. At ten years the centre has risen to 5.993262 per cent and the spread has settled at 0.999977 percentage points, so nought is 5.993398 spreads below the centre and the answer is 1.028e-09, roughly one chance in a billion.
| Horizon | Centre of the rate | Spread | Spreads to nought | Chance below nought |
|---|---|---|---|---|
| One year | 5.393469 per cent | 0.795060 points | 6.783725 | 5.856e-12 |
| Three years | 5.776870 per cent | 0.974789 points | 5.926279 | 1.549e-09 |
| Ten years | 5.993262 per cent | 0.999977 points | 5.993398 | 1.028e-09 |
| Very long run | 6.000000 per cent | 1.000000 points | 6.000000 | 9.866e-10 |
Read the middle rows before the ends. The pattern is not the one most readers expect. The chance rises from one year to three years, peaks at about 1.649e-09 near three and a half years, and then falls back and settles at 9.866e-10. The probability of a negative rate under the Vasicek model does not keep climbing with the horizon; it rises while the spread is still growing, then settles once the pull has taken hold and the spread has stopped growing. The settled value is the chance of being six spreads below a centre, a quantity independent of the horizon once the model has reached its stationary distributionThe distribution a mean reverting process settles into after a long time. Where the process started no longer matters..
Both figures deserve to be said out loud rather than left in scientific notation. Six chances in a trillion is about one chance in one hundred and seventy billion. One chance in a billion is about one chance in nine hundred and seventy three million. The negative rate criticism of the Vasicek model is structurally true and, at these settings, numerically negligible, and both halves of that sentence have to survive. Say only the first half and the reader rejects a serviceable model over nothing. Say only the second half and the reader will be caught out the day somebody hands them a set of parameters where the criticism does bite.
An answer is worth committing to before the control below is used. At these settings, what is the chance of the Vasicek rate being below nought over one year?
Turn the randomness up and watch a negligible criticism become a real one
One control: the size of the randomness at the starting level, in percentage points a year. Everything else is held fixed at the five settings above. The bell is the distribution of the Vasicek rate at one year, drawn to a constant area so the red region below nought is the probability itself rather than a decoration. The answer ranges over forty orders of magnitude and a straight scale would show nothing, so the strip beneath places that probability on a scale of powers of ten. At the default of 1.000000 points the reading is 5.856e-12, the published figure.
Three readings from that control are the whole argument, and they are worth writing down. At 1.000000 points of randomness the chance is 5.856e-12. At 3.000000 points the one-year spread widens to 2.385180 points and the chance is 1.187e-02, about one in eighty four. At 5.000000 points the spread is 3.975300 points and the chance is 8.743e-02, about one in eleven. The negative rate criticism is worth nothing at one point of randomness and a great deal at three, and the only way to know which case applies is to compute it at the settings in hand.
Notice what the CIR reading does across that same range. The positivity condition compares a fixed 0.060000 against the squared CIR parameter, and that squared parameter is 0.002000 at one point, 0.018000 at three points and 0.050000 at five points. The condition holds at every setting the control reaches, and it would first bind at 5.477226 points. So across the entire range where the Vasicek criticism goes from six in a trillion to one in eleven, the CIR floor never fails once. An unbroken floor across that whole range is a real difference between the two models, and it is the one the criticism is actually pointing at.
The failure: repeating the criticism instead of computing it
The sentence saying that the Vasicek rate can go negative is the first thing said about the model almost everywhere it is mentioned, and it is true. A number almost never follows it. A reader who takes the criticism at face value rejects the simpler model over a risk that, at the settings in front of them, is one part in one hundred and seventy billion over a year, and adopts a harder model in exchange for nothing they can measure.
The cost is not dramatic and that is why it survives. The cost is a slightly worse tool chosen on a true statement that does not bite: more difficult arithmetic, a distribution that is no longer normal, a heavier implementation, and a discussion that never happened because everybody agreed the criticism settled it. The fix is one line of arithmetic that almost nobody performs: divide the centre by the spread and look up the tail.
The mirror image of the error is just as costly and is worth naming. Dismissing the criticism because it is small here would be the same mistake in reverse. At three percentage points of randomness the chance is one in eighty four over a single year, and one in eighty four is not a curiosity. The criticism is real. Whether it matters is a computation, and the computation takes about a minute.
At roughly what level of randomness does the negative rate criticism start to bite?
So where do the two genuinely differ, if not at the floor?
The two models are not the same model, so if the floor is not doing the work at these settings, something else must be. The honest answer is the shape of the spread. Under the Vasicek model the size of the randomness is the same at every level of the rate. Under the CIR model it grows with the square root of the level. The Vasicek spread is level independent and the CIR spread widens as the rate rises. The widening changes the whole distribution rather than only its floor.
Put numbers on that and the difference is not subtle. At a rate of 1 per cent the Vasicek randomness is still 1.000000 percentage points a year while the CIR randomness is 0.447214. At 5 per cent the two agree exactly at 1.000000 by construction. At 12 per cent the Vasicek randomness is still 1.000000 while the CIR randomness has grown to 1.549193. The two models disagree about the size of the wobble everywhere except at the one level where they were matched.
The difference in the size of the wobble bends the distribution of the rate. The Vasicek rate is normally distributed at every horizon and therefore perfectly symmetric: the chance of being a given distance above the centre equals the chance of being the same distance below it. The CIR rate is not. Because the wobble grows on the way up and shrinks on the way down, the distribution is squeezed on the left and stretched on the right. Its skewness at one year is 0.262843 against the Vasicek figure of exactly nought, and in the long run it settles at 0.365148.
The clearest way to see this is to line up the same percentiles under both models and subtract. Below the centre the CIR percentiles all sit higher than the Vasicek ones, since the left tail has been pulled in. Above the centre the CIR percentiles sit higher too. The right tail has been stretched out. And the median itself sits slightly lower. The CIR distribution is not the Vasicek one with a floor bolted on; it is a different shape in every part of its range.
Now for the part that most readers find surprising, and it is the strongest evidence that the floor is not where the action is at these settings. A one-year zero-coupon bond priced under each model, treating these same parameters as holding under the risk-neutral measure Q so that one set of numbers serves throughout, comes to Rs 94.921594/- per Rs 100/- of face under Vasicek and Rs 94.921620/- under CIR. The two models agree on a one-year bond to within three hundred-thousandths of a paisa per hundred rupees. A flat 5 per cent rate gives Rs 95.122942/- and is wrong by about twenty paise.
Say that again in yields. Yields are the cleanest form. The one-year zero rate is 5.211896 per cent under the Vasicek model and 5.211869 per cent under the CIR model, a gap of 0.000028 percentage points. Against a flat 5 per cent the gap is 0.211896 percentage points, nearly eight thousand times larger. At ten years the two models separate a little more, to 5.787294 per cent against 5.785316 per cent, but the gap is still only 0.001977 percentage points. The choice between these two models is a much smaller decision than the choice to model the rate at all.
Where do the two models genuinely differ, if the floor is not doing the work at these settings?
Which of the two should be reached for, and on what grounds?
The grounds are a computation and not a preference. Given a set of parameters, the centre and the spread at the horizon in question are worked out, one is divided by the other, and the tail is read off. If the answer is negligible, the simpler model costs nothing, and simpler is worth having: a normal distribution, a bond price that fits on one line, and randomness that behaves the same way everywhere.
If the answer is not negligible, the floor is a genuine constraint and worth paying for. The price is real. The CIR distribution is a scaled non-central chi-square rather than a normal one, its simulation needs care near nought, and every later step carries a heavier object. The price is worth paying when the floor matters and is a waste when it does not.
There is a second ground worth stating. If the rate spends its time near nought, the level-dependent wobble is not just about the floor: it says the model believes a rate near nought moves less than a rate at 12 per cent. The square root term is making a claim about behaviour, and the CIR model is worth holding because that claim is believed, not because somebody said the other model can go negative. The CIR model is chosen for what its square root term asserts about how a rate moves, not as insurance against a tail nobody bothered to measure.
How should the choice between the two models be decided?
How does somebody reviewing a rate model use all this?
Rebuilding anybody's model is not necessary here. Three checks, all cheap, catch most of what goes wrong when these two models are discussed, and each can be run on a printed sheet of output with nothing else to hand.
- Find the centre and the spread, then divide
Any statement of a rate model's one-year distribution carries a centre and a spread. The centre divided by the spread gives the number of spreads to nought, which is the whole of the negative rate question in one figure.
At the settings above that division gives 6.783725, a tail of 5.856e-12 that settles the argument in one line.
- Ask whether the two models were matched anywhere before being compared
A comparison of a fixed randomness against a square root randomness is meaningless unless the two were made to agree at some level first. If no matching level is stated, the comparison is measuring the setting rather than the model.
Here the matching happens at the 5 per cent starting level, and the CIR parameter of 0.044721 is chosen so that both read 1.000000 points there.
- Check whether the claimed difference is about the floor or about the shape
If a note says the models differ because one has a floor, look at the percentiles rather than the floor. At these settings the one-year bond prices differ by three hundred-thousandths of a paisa while the 99th percentile of the rate differs by 0.199349 percentage points.
A difference that shows up in the tails and not in the price is a difference of shape, and it is the one worth arguing about.
All three checks are the same check three times: has anybody put a number next to the claim? The one question catches the error long after the arithmetic has been forgotten, so it is worth carrying away even if every formula above fades. A criticism without a magnitude is not yet a reason, and a model comparison without a matching level is not yet a comparison.
There is a household version of the whole argument too. A shopkeeper who says a scale might be wrong has said something true of every scale ever made, and it decides nothing. A shopkeeper who says it might be wrong by two grams in a kilogram has said enough to settle whether to care, and the answer depends entirely on whether the purchase is rice or gold. The Vasicek criticism is the first sentence. The 5.856e-12 is the second one, and nobody says it.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for work on mean reverting rate models and their transition distributions | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Vasicek, 1977 | An equilibrium characterization of the term structure, the paper introducing the model | Journal of Financial Economics |
| Cox, Ingersoll and Ross, 1985 | A theory of the term structure of interest rates, the paper introducing the model | Econometrica |
| Hull, Shreve and Wilmott | Standard texts on derivatives and stochastic calculus, for notation and the usual ordering of the two models | textbooks, named in the text |
The short rate and the five parameters governing it are invented.
Educational material. Not advice on any investment, tax, budget or market position.
