Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Stochastic Calculus & Derivative Pricing Theory
1Probability Foundations
The Probability SpaceRandom VectorsSigma-AlgebraExpectationSample Space and EventsDensity and Distribution FunctionsRisk-Neutral ProbabilityState Price Density vs…
2Stochastic Processes and Jumps
Properties of a Stochastic ProcessMartingaleBrownian Motion and Its PropertiesBrownian Motion vs Geometric…Stopping TimeThe Markov PropertyState VariablesTransition ProbabilityQuadratic VariationQuadratic Variation vs Ordinary…Submartingale and SupermartingaleMartingale RepresentationMarkov Process vs MartingaleOptional StoppingFiltrationJump ProcessesThe Poisson ProcessLevy ProcessesJump Diffusion
3Ito Calculus
The Ito IntegralThe Ito Integral vs the Riemann IntegralInfinitesimals in Stochastic CalculusQuadratic CovariationIto's LemmaHow to Apply Ito's…The Infinitesimal GeneratorIto Calculus vs Ordinary Calculus
4Stochastic Differential Equations
Stochastic Differential EquationsStochastic Differential Equation vs…Drift and DiffusionStrong and Weak Solutions ComparedDiscretisationGeometric Brownian Motion
5Pricing Theory and No-Arbitrage
No-ArbitrageGirsanov, Radon-Nikodym and Change…Physical and Risk-Neutral Measures…The Fundamental Theorems of…The Law of One PriceThe Pricing KernelDiscount Factors and Zero-Coupon PricesReplication vs HedgingComplete Market vs Incomplete MarketClearing Margin Architecture
6Option Pricing Theory
European and American OptionsMonte Carlo European OptionThe Black-Scholes PDEBlack Scholes and the GreeksThe Payoff FunctionThe Binomial ModelBinomial Option PricingDelta Hedging in TheoryBoundary, Initial and Terminal ConditionsThe Exercise BoundaryHow to Check Put-Call…
7Volatility Models
Constant, Local and Stochastic…Vasicek Model vs CIR ModelThe Heston ModelThe SABR ModelThe Volatility ProcessImplied VolatilityVolatility Smile vs Skew vs Surface
8Interest Rate Models
Interest-Rate DerivativesMean ReversionThe Zero-Coupon BondThe Ornstein-Uhlenbeck ProcessThe Discount CurveZero RatesShort-Rate Model vs Market Model
9Numerical Pricing
Closed Form and Numerical…Monte Carlo PricingEuler and Milstein Schemes ComparedTree MethodsFinite Difference MethodsNumerical Error and StabilityVariance Reduction
10Calibration and Model Risk
Model OverrideMarket Price and Model PriceCalibrationHow to Document a Pricing ModelThe Educational Illustration LabelMarket ConventionsModel Uncertainty and LimitationsBacktesting a Pricing ModelIdentifiabilityCalibrated ParametersThe Calibration Loss Function

Vasicek Model vs CIR Model: Two Ways to Hold a Rate

Both models pull a rate toward a long-run level at a stated speed, and they agree on that pull term exactly. The two differ in one term only. The Vasicek randomness is the same size at every level. The Cox, Ingersoll and Ross (CIR) randomness shrinks as the rate approaches nought. The usual criticism is that the first can go negative, and at the settings used here that probability is about six in a trillion.

Two models, one difference. Everything said about how the Vasicek model and the CIR model behave, what shape their outcomes take, whether either can produce a negative number and which of the two is harder to work with, comes out of a single term written two different ways. With that term straight, the rest is bookkeeping.

Both models are defined completely before either is set against the other, and the contrast that follows comes with arithmetic rather than with received opinion.

One warning belongs before the definitions. The most repeated sentence about the Vasicek model is that its rate can go negative, and almost nobody who repeats it has computed how likely that is at the settings in front of them. Being told that a floor can be breached, with no number attached, is like being told a bridge can collapse with no mention of the load that does it. The statement is true. On its own it decides nothing. A number is what turns the statement into a decision.

The settings every figure below is built from

Every figure below is built from one set of invented parameters, held fixed so that each number can be checked against the last. The short rateThe rate for borrowing over the next instant of time. Both of these models describe that single quantity and nothing else. starts at 5 per cent, it is pulled toward 6 per cent, the pull acts at a speed of 0.5 a year, and the size of the randomness at the starting level is one percentage point a year. The horizon is one year unless a longer one is named.

SettingSymbolValueRead it as
Starting level of the rater at time zero5.000000 per centwhere the rate sits today
Long-run leveltheta6.000000 per centwhere the pull is directed
Speed of the pullkappa0.500000 a yearthe fraction of the remaining gap closed per year
Size of the randomness at the starting levelsigma with subscript r1.000000 pointsthe one-year spread contributed per year, in percentage points
HorizonT1.000000 yearthe period every distribution below is measured over

The five settings were chosen so that the results land on figures a reader can check: the gap to close is exactly one percentage point, the speed of 0.5 gives a half-life of 1.386294 years, and the settled spread of the rate works out to exactly 1.000000 percentage points. Each model below is fed these same five numbers, so every difference between the two is caused by the models and not by the inputs.

What is the Vasicek model, defined from scratch?

The Vasicek model, named for Vasicek and published in 1977, says that the short rate changes by two things added together. The first is a pull toward a long-run level, and the size of that pull at any moment is the speed multiplied by however far the rate currently sits from that level. The second is a random push whose size does not depend on the level at all. The two terms are the whole model. There is nothing else in it.

The Vasicek model, the whole definition
$$ dr_t \;=\; \kappa\bigl(\theta - r_t\bigr)\,dt \;+\; \sigma_r\,dW_t $$
\(r_t\)the short rate at time \(t\), as a decimal, which is the quantity being modelled
\(\kappa\)the speed of the pull, 0.5 a year here
\(\theta\)the long-run level the pull is directed at, 0.06 here
\(\sigma_r\)the size of the randomness, 0.01 a year here, in absolute terms rather than as a percentage of the rate
\(W_t\)standard Brownian motion under the physical measure \(\mathbb{P}\), the source of the randomness
\(dt\)an instant of time, in years
What it says in wordsOver the next instant the short rate moves by the speed multiplied by the distance from the long-run level, which is the pull, plus a fixed amount of randomness whose size is the same whatever the rate happens to be.

The two terms do different jobs and they are the only two moving parts here, so each one repays reading on its own. The pull term is what mean reversionBeing drawn back toward a stated level at a stated speed. The further a quantity strays, the harder it is pulled back. means: the further the rate strays, the harder it is pulled back, in exact proportion to the distance. At the starting level of 5 per cent the gap to close is one percentage point, so the pull contributes 0.5 multiplied by 0.01. The product is 0.005 a year, or 0.500000 percentage points a year of upward drift. Move the rate to 7 per cent and the same arithmetic gives 0.500000 percentage points a year of downward drift instead. The pull has no preferred direction of its own; it only ever points at the long-run level.

There is a rubber band in this, and it is worth carrying. Stretch a band twice as far and it pulls back twice as hard. The proportionality is exactly what the speed multiplied by the distance is saying, and it is why the rate never simply drifts off: the further it goes, the stronger the force bringing it home. The speed of reversionHow fast the pull acts, expressed as the fraction of the remaining gap closed per year. Here it is 0.5 a year. is the stiffness of the band and the long-run levelThe level the pull is directed at. The rate is drawn toward that level but does not have to sit on it. Here it is 6 per cent. is the point it is anchored to.

The whole skeleton, and both models in this guide are built on it. BOX ONE: THE PULL speed times the distance 0.5 times the gap from 6 per cent at 5 per cent: 0.500000 points a year up identical in both models, term for term + BOX TWO: THE RANDOMNESS a size, times the random push Vasicek: the size is fixed at 1.000000 points CIR: the size shrinks as the rate falls this is the only box that changes Box one plus box two is the entire model. Nothing else is in it. The same skeleton carries a variance process later, which is why learning it once pays twice. Changing box one changes where the rate is anchored. Changing box two changes only how it wobbles on the way, which is the whole subject of this guide.
A pull toward a level at a stated speed, plus a random term, is the shared skeleton of both models, and only the random term is written differently between them.

The second term is where the argument eventually lands, and it repays close attention. The randomness has a fixed size, one percentage point a year, and that size applies whether the rate is at 12 per cent, at 5 per cent, at 0.1 per cent or at exactly nought. The Vasicek randomness never consults the level of the rate before deciding how large a push to deliver. The fixed size is a modelling choice, and it is the choice everything below turns on.

Because both terms are so simple, the model can be solved rather than only simulated, and the rate at any future date has a distribution that can be written down. The distribution is normal, with a centre that slides from the starting level toward the long-run level and a spread that grows and then settles.

Where the Vasicek rate is at a horizon
$$ r_T \;\sim\; \mathcal{N}\!\left(\theta + (r_0-\theta)e^{-\kappa T},\;\; \frac{\sigma_r^{2}}{2\kappa}\bigl(1-e^{-2\kappa T}\bigr)\right) $$
\(r_T\)the short rate at the horizon \(T\), the quantity whose distribution is being stated
\(r_0\)the starting level of the rate, 0.05 here
\(\theta,\ \kappa,\ \sigma_r\)the long-run level, the speed and the size of the randomness, as defined above
\(\mathcal{N}(m,v)\)the normal distribution with centre \(m\) and variance \(v\)
\(e^{-\kappa T}\)the fraction of the original gap still uncovered at the horizon, 0.606531 at one year
What it says in wordsAt any horizon the Vasicek rate is normally distributed, its centre sits at the long-run level less whatever fraction of the original gap is still uncovered, and its variance rises from nought toward a settled value of the randomness squared divided by twice the speed.

Put the five settings into that and the one-year answer is a centre of 5.393469 per cent and a spread of 0.795060 percentage points. The whole of the negative rate argument below is a comparison between those two figures, so both are worth holding. The centre has covered 39.3469 per cent of the one point gap in a year, and the spread has reached 79.5060 per cent of its settled value of 1.000000 percentage points. A normal distribution has no floor of any kind. The missing floor is exactly why the Vasicek rate can in principle be negative, and exactly why a number rather than an opinion is needed next.

Try it out

In the Vasicek model, what decides how large the next random push is?

How fast does the pull actually close the gap?

The speed of 0.5 a year is a number, and a number on its own is hard to feel. Convert it into a duration and it becomes concrete. The half-lifeThe time the pull takes to close half of whatever gap remains, computed as the natural logarithm of two divided by the speed. of the pull is the natural logarithm of two divided by the speed. At a speed of 0.5 a year the half-life is 1.386294 years. The half-life is the time it takes to close half of whatever gap is left, and it does not depend on how big the gap was.

Start at 5 per cent with a gap of 1.000000 percentage points to the long-run level. After 1.386294 years the expected level is 5.500000 per cent, and the gap remaining is 0.500000 points. After another 1.386294 years the expected level is 5.750000 per cent and the gap is 0.250000 points. The gap halves, then halves again, and it never quite reaches nought. The pull is a geometric decay of the gap rather than a march to the long-run level, so the rate is always still on its way there.

Half the remaining gap closes every 1.386294 years, and both models share this curve. 6.00 5.00 5.50 half-life 1.386294 years half the gap gone, expected level 5.500000 per cent one year 5.393469 per cent today 1 year 3 years 5 years The curve flattens because the gap it is closing keeps shrinking, not because the pull weakens.
The expected short rate climbs from 5 per cent toward 6 per cent, reaching 5.500000 per cent after one half-life of 1.386294 years and 5.393469 per cent after one year.

The curve belongs to both models equally, and that is the first thing worth noticing about the comparison. The expected level at any horizon depends only on the starting level, the long-run level and the speed. The randomness does not enter it at all. The Vasicek model and the CIR model produce exactly the same expected path, to every decimal place, at every horizon. Whatever separates them, it is not where they think the rate is going.

Try it out

At a speed of 0.5 a year, how long does the pull take to close half the remaining gap?

What is the CIR model, defined from scratch?

The CIR model, named for Cox, Ingersoll and Ross and published in 1985, keeps the pull term exactly as it stands and rewrites the randomness. Instead of a fixed size, the size of the random push is a parameter multiplied by the square root of the rate itself. When the rate is high the pushes are large, and when the rate is low they are small. If the rate ever reached nought the random push would be nought too, leaving only the pull. At nought the pull points firmly upward.

The CIR model, the whole definition
$$ dr_t \;=\; \kappa\bigl(\theta - r_t\bigr)\,dt \;+\; \sigma_r^{\,c}\sqrt{r_t}\;dW_t $$
\(r_t\)the short rate at time \(t\), the same quantity as before
\(\kappa,\ \theta\)the speed and the long-run level, identical to the Vasicek values, 0.5 a year and 0.06
\(\sigma_r^{\,c}\)the CIR randomness parameter, 0.044721 here. The superscript marks it as a different number from \(\sigma_r\), because it multiplies a square root rather than standing alone
\(\sqrt{r_t}\)the square root of the current rate, which is the whole of the difference between the two models
\(W_t\)standard Brownian motion under the physical measure \(\mathbb{P}\), as before
What it says in wordsOver the next instant the CIR rate moves by the same pull as the Vasicek rate, plus a random push whose size is a parameter multiplied by the square root of the rate, so the randomness fades away as the rate approaches nought.

There is a household version of this that makes the mechanism obvious. Think of the randomness as a hand shaking while it carries a jug. One model says the hand shakes by the same amount whatever is in the jug, so an almost empty jug still gets a full shake and can end up owing liquid it does not have. The other says the shake scales with what is in the jug, so a nearly empty jug barely trembles and an empty one does not move at all. The CIR model does not forbid the rate from reaching nought by decree; it removes the push that would take it there.

One practical problem is left. The CIR parameter is not the same number as the Vasicek one, so quoting the two models with the same figure for randomness would be comparing nothing to nothing. The fix is to match them at the level the rate actually starts from. On day one the two models then deliver randomness of identical size, and every later difference is caused by the shape of the term rather than by its calibration.

Matching the two models at the starting level
$$ \sigma_r^{\,c}\sqrt{r_0} \;=\; \sigma_r \qquad\Longrightarrow\qquad \sigma_r^{\,c} \;=\; \frac{\sigma_r}{\sqrt{r_0}} \;=\; \frac{0.01}{\sqrt{0.05}} \;=\; 0.044721 $$
\(\sigma_r\)the Vasicek randomness, 0.01, which is one percentage point a year
\(\sigma_r^{\,c}\)the CIR parameter that makes the two agree at the starting level, 0.044721
\(r_0\)the starting level of the rate, 0.05
\(\sigma_r^{\,c}\sqrt{r_0}\)the CIR randomness evaluated at the starting level, 0.010000 exactly by construction
What it says in wordsChoosing the CIR parameter as the Vasicek randomness divided by the square root of the starting rate makes both models deliver randomness of exactly the same size on day one, so any later difference between them comes from the shape of the term and not from an unfair setting.

The matching also lands on a clean number. The CIR parameter is 0.044721 and its square is 0.002000 exactly. Squaring 0.01 gives 0.0001, and dividing by 0.05 gives 0.002 with no rounding anywhere. The condition keeping the CIR rate away from nought is stated in terms of that squared value, so the value is about to do real work.

The positivity conditionThe requirement, due to Feller, that the pull at nought is strong enough to overwhelm the randomness there, keeping the rate strictly above nought at all times. asks whether the pull at nought is strong enough to overwhelm the randomness near nought. The condition compares twice the speed multiplied by the long-run level against the squared CIR parameter. If the first is at least as large as the second, the rate is strictly above nought at every future time, with probability one, and not merely unlikely to be negative.

The CIR positivity condition
$$ 2\kappa\theta \;\ge\; \bigl(\sigma_r^{\,c}\bigr)^{2} \qquad\Longrightarrow\qquad 0.060000 \;\ge\; 0.002000 $$
\(2\kappa\theta\)twice the speed multiplied by the long-run level, 0.060000 here, which measures how hard the rate is pushed up when it is near nought
\(\bigl(\sigma_r^{\,c}\bigr)^{2}\)the squared CIR parameter, 0.002000 here, which measures how hard the randomness fights back near nought
\(\ge\)the direction of the comparison. When it holds, the rate stays strictly above nought at all times
What it says in wordsWhen twice the speed multiplied by the long-run level is at least the squared randomness parameter, the upward pull near nought beats the randomness near nought, and the CIR rate never reaches nought at any future time.

At these settings the comparison is 0.060000 against 0.002000, so the condition holds by a factor of thirty and the CIR rate cannot reach nought at all. That is a different kind of statement from a small probability. Reaching nought is not merely unlikely. The set of paths reaching nought has probability nought, so the floor is a property of the model rather than a matter of how the numbers happened to fall.

The condition is two computable numbers set beside each other, not a hope about the model. THE PUSH UP NEAR NOUGHT: TWICE THE SPEED TIMES THE LONG-RUN LEVEL 0.060000 THE FIGHT BACK NEAR NOUGHT: THE SQUARED RANDOMNESS PARAMETER 0.002000 thirty times smaller, so the condition holds with room to spare 0.00 0.02 0.04 0.06 0.07 It would only bind if the squared parameter reached 0.060000, thirty times the one in use. Educational illustration built from invented parameters. No rate here describes any market.
Twice the speed multiplied by the long-run level is 0.060000 against a squared randomness parameter of 0.002000, so the CIR positivity condition holds by a factor of thirty at these settings.
Try it out

What keeps the CIR rate away from nought?

Breaking Into Quants Bootcamp — Fin Maverick

What does the one term that differs actually do?

Set the two definitions side by side and count the differences. The pull is the same. The speed is the same. The long-run level is the same. The source of randomness is the same Brownian motion. One thing is written differently: whether the size of the random push is a constant or a constant multiplied by the square root of the rate. Exactly one term separates the Vasicek model from the CIR model, and it is the term controlling how big the wobble is, not the term controlling where the rate is heading.

The single change carries four consequences, and they are worth naming before drawing them. The square root creates a floor at nought under one model and not the other. The square root makes the spread of outcomes depend on the level rather than being the same everywhere. The square root bends the distribution of the rate out of the symmetric normal shape into one that leans to the right. And the square root makes the second model harder to work with. A square root termA randomness term scaled by the square root of the quantity being modelled. The scaling is what makes the CIR distribution harder to work with than a normal one. inside the randomness turns a normal distribution into a scaled non-central chi-square one, a heavier object to carry through every later calculation.

Same geometry, same rows, one row different. Count them yourself. VASICEK, 1977 COX, INGERSOLL AND ROSS, 1985 PULLED TOWARD 6.000000 per cent PULLED TOWARD 6.000000 per cent AT A SPEED OF 0.500000 a year AT A SPEED OF 0.500000 a year EXPECTED LEVEL IN A YEAR 5.393469 per cent EXPECTED LEVEL IN A YEAR 5.393469 per cent RANDOMNESS, THE ONE ROW THAT DIFFERS a fixed 1.000000 points the same at every level, including nought RANDOMNESS, THE ONE ROW THAT DIFFERS 0.044721 times the square root 1.000000 points at 5 per cent, nought at nought Three rows identical, one row different. That is the whole comparison.
Both models pull toward the same level at the same speed and share the same expected path, and only the row describing the size of the randomness is written differently.
Try it out

How many terms differ between the Vasicek model and the CIR model?

Try it out

Before reading on: is the chance of the Vasicek rate being below nought larger at ten years or at one year?

What is the negative rate criticism worth at these parameters?

Here is the criticism, stated fairly. The Vasicek rate at any horizon is normally distributed, a normal distribution puts weight on every value on the line, and so the model assigns a strictly positive probability to the rate being below nought. The positive probability is not a bug anybody overlooked; it follows directly from the randomness being the same size at every level, including at nought where it should arguably be nothing. The criticism is structurally correct and it is not going away.

A probability that is positive can be anything from a certainty to a rounding error, and only the arithmetic says which. The computation settles it. The chance of the rate being below nought at a horizon is the normal distribution function evaluated at minus the centre divided by the spread. Both of those numbers are already established above.

The probability the Vasicek rate is below nought
$$ \mathbb{P}\bigl(r_T < 0\bigr) \;=\; N\!\left(\frac{0 - m_T}{s_T}\right), \qquad m_T = \theta + (r_0-\theta)e^{-\kappa T}, \quad s_T = \sigma_r\sqrt{\tfrac{1-e^{-2\kappa T}}{2\kappa}} $$
\(m_T\)the centre of the distribution at the horizon, 0.05393469 at one year
\(s_T\)the spread at the horizon, 0.00795060 at one year, which is 0.795060 percentage points
\(N(\cdot)\)the standard normal distribution function, the chance of being below the value given to it
\(\mathbb{P}\)the physical measure, the rule under which this chance is taken
What it says in wordsThe chance the Vasicek rate is below nought at a horizon is the normal distribution function applied to how many spreads nought sits below the centre of the distribution at that horizon, and nothing else enters the calculation.

At one year the centre is 0.05393469 and the spread is 0.00795060, so nought sits 6.783725 spreads below the centre. Feed that to the normal distribution function and the answer is 0.000000000005855794. Written shorter, 5.856e-12, or roughly six chances in a trillion. At ten years the centre has risen to 5.993262 per cent and the spread has settled at 0.999977 percentage points, so nought is 5.993398 spreads below the centre and the answer is 1.028e-09, roughly one chance in a billion.

HorizonCentre of the rateSpreadSpreads to noughtChance below nought
One year5.393469 per cent0.795060 points6.7837255.856e-12
Three years5.776870 per cent0.974789 points5.9262791.549e-09
Ten years5.993262 per cent0.999977 points5.9933981.028e-09
Very long run6.000000 per cent1.000000 points6.0000009.866e-10

Read the middle rows before the ends. The pattern is not the one most readers expect. The chance rises from one year to three years, peaks at about 1.649e-09 near three and a half years, and then falls back and settles at 9.866e-10. The probability of a negative rate under the Vasicek model does not keep climbing with the horizon; it rises while the spread is still growing, then settles once the pull has taken hold and the spread has stopped growing. The settled value is the chance of being six spreads below a centre, a quantity independent of the horizon once the model has reached its stationary distributionThe distribution a mean reverting process settles into after a long time. Where the process started no longer matters..

Both figures deserve to be said out loud rather than left in scientific notation. Six chances in a trillion is about one chance in one hundred and seventy billion. One chance in a billion is about one chance in nine hundred and seventy three million. The negative rate criticism of the Vasicek model is structurally true and, at these settings, numerically negligible, and both halves of that sentence have to survive. Say only the first half and the reader rejects a serviceable model over nothing. Say only the second half and the reader will be caught out the day somebody hands them a set of parameters where the criticism does bite.

Try it out

An answer is worth committing to before the control below is used. At these settings, what is the chance of the Vasicek rate being below nought over one year?

Play with it

Turn the randomness up and watch a negligible criticism become a real one

One control: the size of the randomness at the starting level, in percentage points a year. Everything else is held fixed at the five settings above. The bell is the distribution of the Vasicek rate at one year, drawn to a constant area so the red region below nought is the probability itself rather than a decoration. The answer ranges over forty orders of magnitude and a straight scale would show nothing, so the strip beneath places that probability on a scale of powers of ten. At the default of 1.000000 points the reading is 5.856e-12, the published figure.

WHERE THE RATE IS IN ONE YEAR, AND HOW MUCH OF IT SITS BELOW NOUGHT nought centre 5.393469 per cent minus 2 12 per cent THE SAME PROBABILITY, ON A SCALE OF POWERS OF TEN negligible it bites here 1e-45 1e-35 1e-25 1e-15 1e-5 1 5.856e-12
Jump to a setting
Randomness at 5 per cent
1.000000 points
One-year spread
0.795060 points
Spreads to nought
6.783725
Chance below nought
5.856e-12
CIR condition, 0.060000 against
0.002000
At a randomness of 1.000000 points the one-year spread is 0.795060 points, nought sits 6.783725 spreads below the centre of 5.393469 per cent, and the chance of the Vasicek rate being below nought is 5.856e-12, about one in 171 billion. The CIR positivity condition compares 0.060000 against 0.002000, so the CIR rate cannot reach nought at this setting.
Educational illustration. Every reading is computed from the closed-form distribution rather than sampled, so the default reproduces the published figures exactly on every reload. Held fixed throughout: a starting rate of 5 per cent, a long-run level of 6 per cent, a speed of 0.5 a year and a horizon of one year. The bell is drawn to a constant area, so below about 0.7 points its peak runs off the top of the panel and is clipped, which is itself the point: all the weight is piled within a whisker of the centre. The CIR parameter is rescaled with the control so that both models keep matching at the starting level. Every parameter is set for teaching rather than read off a market.

Three readings from that control are the whole argument, and they are worth writing down. At 1.000000 points of randomness the chance is 5.856e-12. At 3.000000 points the one-year spread widens to 2.385180 points and the chance is 1.187e-02, about one in eighty four. At 5.000000 points the spread is 3.975300 points and the chance is 8.743e-02, about one in eleven. The negative rate criticism is worth nothing at one point of randomness and a great deal at three, and the only way to know which case applies is to compute it at the settings in hand.

Notice what the CIR reading does across that same range. The positivity condition compares a fixed 0.060000 against the squared CIR parameter, and that squared parameter is 0.002000 at one point, 0.018000 at three points and 0.050000 at five points. The condition holds at every setting the control reaches, and it would first bind at 5.477226 points. So across the entire range where the Vasicek criticism goes from six in a trillion to one in eleven, the CIR floor never fails once. An unbroken floor across that whole range is a real difference between the two models, and it is the one the criticism is actually pointing at.

The failure: repeating the criticism instead of computing it

The sentence saying that the Vasicek rate can go negative is the first thing said about the model almost everywhere it is mentioned, and it is true. A number almost never follows it. A reader who takes the criticism at face value rejects the simpler model over a risk that, at the settings in front of them, is one part in one hundred and seventy billion over a year, and adopts a harder model in exchange for nothing they can measure.

The cost is not dramatic and that is why it survives. The cost is a slightly worse tool chosen on a true statement that does not bite: more difficult arithmetic, a distribution that is no longer normal, a heavier implementation, and a discussion that never happened because everybody agreed the criticism settled it. The fix is one line of arithmetic that almost nobody performs: divide the centre by the spread and look up the tail.

The mirror image of the error is just as costly and is worth naming. Dismissing the criticism because it is small here would be the same mistake in reverse. At three percentage points of randomness the chance is one in eighty four over a single year, and one in eighty four is not a curiosity. The criticism is real. Whether it matters is a computation, and the computation takes about a minute.

The sentence everybody repeats, and the number almost nobody adds to it. THE ARTEFACT: A LINE IN A REVIEW NOTE "Rejected the Vasicek model: the rate can go negative." True. Complete as a statement. Empty as a decision, because no number follows it. WHAT THE ARITHMETIC SAYS 5.856e-12 over one year, at the settings in this guide about one chance in 171 billion nought sits 6.783725 spreads below the centre The criticism is structurally right. Here it decides nothing; at three points it decides everything. Divide the centre by the spread, look up the tail, and the sentence turns into a decision. Educational illustration. The review note is invented and quotes nobody.
The chance of the Vasicek rate falling below nought at one year is 5.856e-12 at these settings, so the criticism is structurally right and numerically irrelevant here.
Try it out

At roughly what level of randomness does the negative rate criticism start to bite?

Derivatives Foundation Bootcamp — Fin Maverick Bond Pricing and Yield Mechanics — free micro-course from Fin Maverick

So where do the two genuinely differ, if not at the floor?

The two models are not the same model, so if the floor is not doing the work at these settings, something else must be. The honest answer is the shape of the spread. Under the Vasicek model the size of the randomness is the same at every level of the rate. Under the CIR model it grows with the square root of the level. The Vasicek spread is level independent and the CIR spread widens as the rate rises. The widening changes the whole distribution rather than only its floor.

Put numbers on that and the difference is not subtle. At a rate of 1 per cent the Vasicek randomness is still 1.000000 percentage points a year while the CIR randomness is 0.447214. At 5 per cent the two agree exactly at 1.000000 by construction. At 12 per cent the Vasicek randomness is still 1.000000 while the CIR randomness has grown to 1.549193. The two models disagree about the size of the wobble everywhere except at the one level where they were matched.

How big the wobble is, plotted against the level the rate is sitting at. 0.5 1.5 0.0 1.0 matched here, at 5 per cent both read 1.000000 points, by construction VASICEK: flat at 1.000000 CIR: the square root curve 0.447214 at 1 per cent, and nought at nought 1.549193 at 12 per cent 0 4 8 12 per cent the level of the rate The two curves meet at one point and disagree everywhere else. That is the real difference.
The Vasicek randomness stays at 1.000000 percentage points at every level while the CIR randomness follows a square root curve, and the two agree only at the 5 per cent level where they were matched.

The difference in the size of the wobble bends the distribution of the rate. The Vasicek rate is normally distributed at every horizon and therefore perfectly symmetric: the chance of being a given distance above the centre equals the chance of being the same distance below it. The CIR rate is not. Because the wobble grows on the way up and shrinks on the way down, the distribution is squeezed on the left and stretched on the right. Its skewness at one year is 0.262843 against the Vasicek figure of exactly nought, and in the long run it settles at 0.365148.

The clearest way to see this is to line up the same percentiles under both models and subtract. Below the centre the CIR percentiles all sit higher than the Vasicek ones, since the left tail has been pulled in. Above the centre the CIR percentiles sit higher too. The right tail has been stretched out. And the median itself sits slightly lower. The CIR distribution is not the Vasicek one with a floor bolted on; it is a different shape in every part of its range.

The CIR percentile less the Vasicek percentile, at one year, in percentage points. no difference PERCENTILE 0.1st +0.247226 1st +0.115484 5th +0.031890 50th minus 0.035698 95th +0.089889 99th +0.199349 Left tail pulled in, median nudged down, right tail stretched out: a lean, not a floor.
Every CIR percentile below the centre sits above the matching Vasicek one and the top percentiles sit further out still, so the CIR distribution leans right rather than simply stopping at nought.

Now for the part that most readers find surprising, and it is the strongest evidence that the floor is not where the action is at these settings. A one-year zero-coupon bond priced under each model, treating these same parameters as holding under the risk-neutral measure Q so that one set of numbers serves throughout, comes to Rs 94.921594/- per Rs 100/- of face under Vasicek and Rs 94.921620/- under CIR. The two models agree on a one-year bond to within three hundred-thousandths of a paisa per hundred rupees. A flat 5 per cent rate gives Rs 95.122942/- and is wrong by about twenty paise.

Say that again in yields. Yields are the cleanest form. The one-year zero rate is 5.211896 per cent under the Vasicek model and 5.211869 per cent under the CIR model, a gap of 0.000028 percentage points. Against a flat 5 per cent the gap is 0.211896 percentage points, nearly eight thousand times larger. At ten years the two models separate a little more, to 5.787294 per cent against 5.785316 per cent, but the gap is still only 0.001977 percentage points. The choice between these two models is a much smaller decision than the choice to model the rate at all.

Try it out

Where do the two models genuinely differ, if the floor is not doing the work at these settings?

Duration and What It Does Not Tell You teaches you to use duration correctly and to know exactly where it stops being true.

Which of the two should be reached for, and on what grounds?

The grounds are a computation and not a preference. Given a set of parameters, the centre and the spread at the horizon in question are worked out, one is divided by the other, and the tail is read off. If the answer is negligible, the simpler model costs nothing, and simpler is worth having: a normal distribution, a bond price that fits on one line, and randomness that behaves the same way everywhere.

If the answer is not negligible, the floor is a genuine constraint and worth paying for. The price is real. The CIR distribution is a scaled non-central chi-square rather than a normal one, its simulation needs care near nought, and every later step carries a heavier object. The price is worth paying when the floor matters and is a waste when it does not.

One computation decides it, and the computation takes about a minute. Compute the chance of a negative rate at the parameters in hand, over the horizon in question NEGLIGIBLE NOT NEGLIGIBLE Take the simpler model A normal distribution, a bond price in one line, and nothing lost. here: 5.856e-12 over a year Pay for the floor A heavier distribution and more care near nought, bought deliberately. at three points: 1.187e-02 over a year The same two branches, the same arithmetic, and a different answer for different parameters. Educational illustration. This is a way of deciding between two models, not a suggestion about any market.
Deciding between the two models means computing the chance of reaching nought at the settings in hand, which is negligible here and a real constraint at three times the randomness.

There is a second ground worth stating. If the rate spends its time near nought, the level-dependent wobble is not just about the floor: it says the model believes a rate near nought moves less than a rate at 12 per cent. The square root term is making a claim about behaviour, and the CIR model is worth holding because that claim is believed, not because somebody said the other model can go negative. The CIR model is chosen for what its square root term asserts about how a rate moves, not as insurance against a tail nobody bothered to measure.

Try it out

How should the choice between the two models be decided?

How does somebody reviewing a rate model use all this?

Rebuilding anybody's model is not necessary here. Three checks, all cheap, catch most of what goes wrong when these two models are discussed, and each can be run on a printed sheet of output with nothing else to hand.

  1. Find the centre and the spread, then divide Any statement of a rate model's one-year distribution carries a centre and a spread. The centre divided by the spread gives the number of spreads to nought, which is the whole of the negative rate question in one figure.
    At the settings above that division gives 6.783725, a tail of 5.856e-12 that settles the argument in one line.
  2. Ask whether the two models were matched anywhere before being compared A comparison of a fixed randomness against a square root randomness is meaningless unless the two were made to agree at some level first. If no matching level is stated, the comparison is measuring the setting rather than the model.
    Here the matching happens at the 5 per cent starting level, and the CIR parameter of 0.044721 is chosen so that both read 1.000000 points there.
  3. Check whether the claimed difference is about the floor or about the shape If a note says the models differ because one has a floor, look at the percentiles rather than the floor. At these settings the one-year bond prices differ by three hundred-thousandths of a paisa while the 99th percentile of the rate differs by 0.199349 percentage points.
    A difference that shows up in the tails and not in the price is a difference of shape, and it is the one worth arguing about.

All three checks are the same check three times: has anybody put a number next to the claim? The one question catches the error long after the arithmetic has been forgotten, so it is worth carrying away even if every formula above fades. A criticism without a magnitude is not yet a reason, and a model comparison without a matching level is not yet a comparison.

There is a household version of the whole argument too. A shopkeeper who says a scale might be wrong has said something true of every scale ever made, and it decides nothing. A shopkeeper who says it might be wrong by two grams in a kilogram has said enough to settle whether to care, and the answer depends entirely on whether the purchase is rice or gold. The Vasicek criticism is the first sentence. The 5.856e-12 is the second one, and nobody says it.

Every parameter used here was set rather than estimated; choosing parameters from observations is set out under calibration. The term structure these two models generate across many maturities is set out under the discount curve and the zero rate, and changing measure for a rate model is set out under physical and risk-neutral measures, so the bond prices above assume the stated parameters hold under the pricing measure. What any contract pays is set out under interest rate derivatives. The mathematics is identical everywhere, so no jurisdiction sets the definition of either model and no rule, threshold, period or product name is engaged.
Debt Capital Markets Bootcamp — Fin Maverick

References

SourceDocumentWhere
arXiv Quantitative FinancePreprint repository for work on mean reverting rate models and their transition distributionsarxiv.org
Social Science Research NetworkWorking paper repository for the same materialssrn.com
Vasicek, 1977An equilibrium characterization of the term structure, the paper introducing the modelJournal of Financial Economics
Cox, Ingersoll and Ross, 1985A theory of the term structure of interest rates, the paper introducing the modelEconometrica
Hull, Shreve and WilmottStandard texts on derivatives and stochastic calculus, for notation and the usual ordering of the two modelstextbooks, named in the text

The short rate and the five parameters governing it are invented.
Educational material. Not advice on any investment, tax, budget or market position.

Comparison

Other comparisons in Volatility Models

Comparison

Volatility Smile vs Skew vs Surface: Reading the Shape

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.