The SABR Model: Fitting the Smile in Rates and FX
The SABR model is built to reproduce a pattern of prices across strikes directly. Where the Heston model starts from a process and asks what pattern it produces, this one starts from the pattern and works back. Four parameters control it. One of them, the exponent, decides how the randomness scales with the level, and it has no counterpart in the previous model.
The reversal of direction repays a second reading. Everything below follows from it. One model is an argument that runs forwards: here is how variance moves, and the prices that come out are worked out from it. The other is an argument that runs backwards: here is the shape prices make across strikes, and a small set of numbers is then found that reproduces that shape. The two are not rival descriptions of the same thing; they are opposite directions of travel, and a direction of travel decides what a model is good for.
Consider something concrete. Suppose five thermometer readings have been taken along a corridor, at five marked points on the floor, and a rule is wanted for the temperature anywhere in the building. One way is to work out how heat moves: where the vents are, how the air circulates, what the walls do. With that right, something can be said about every room, including rooms nobody walked into. The other way is to draw the smoothest curve through those five marks. The curve will pass through all five marks beautifully. Asked about the basement, it has nothing to say. A curve through five marks was never a theory of heat.
The Heston model is the first way. The stochastic alpha, beta, rho (SABR) modelA model organised around the pattern of prices across strikes rather than around a process. is the second way, done carefully, with a process written down so that the curve is at least consistent with something moving. Being the second way is not a defect; it is a design choice with a specific payoff and a specific bill.
Fitting the model to observations is covered separately
The word fitting sits in the name this model is known by. Choosing parameters so that a model agrees with a set of observations is taken up separately and later, and a model and a fitted model are different objects. Every number below is a computed consequence of parameters written down by hand, never a parameter recovered from an observation.
What is this model built to do?
The standard process running through this whole subject area is a single invented traded quantity, written S with a time subscript, starting at Rs 100/-, carrying a volatility of 20 per cent a year, watched over one year against a risk-free rate of 5 per cent.
Under the constant treatment, that 20 per cent is one number for every strike. The earlier treatment of a single volatility across five strikes showed what happens when one number is asked to serve five contracts at five different strikes: it can be made to agree with any one of them and then misses the other four, cheap below and dear above, in a pattern that swings sign exactly once. Error scatters. A miss that changes sign exactly once does not. So the number is not enough, and something has to replace it.
The SABR model replaces it with a shape. The model sets down a process whose randomness is itself random and scales with the level of the process by an adjustable amount, and an expression then turns those parameters straight into a volatility for every strike. The whole apparatus exists so that a handful of parameters produces a curve across strikes, and the curve is the output the model was designed to deliver.
The ambition is a different one from the previous model's. The previous one wanted variance to behave: to be positive, to pull back to a long-run level, to be correlated with the process. Prices across strikes were a consequence to be worked out afterwards. Here, prices across strikes are the target, and the process is written down in whatever form makes that target reachable.
| \(S_t\) | the standard process, the single invented traded quantity, at time \(t\) |
| \(\alpha_t\) | the level parameter, itself a random process, starting at \(\alpha_0=0.20\) |
| \(\beta\) | the exponent, a fixed number between nought and one, set to 1 here |
| \(\nu\) | the volatility of the level parameter, set to 0.30 here |
| \(\rho\) | the correlation between the two Brownian motions, set to minus 0.7 here |
| \(\tilde W\) | a Brownian motion under the pricing measure Q, one for each equation |
Two things about that pair of equations are worth naming before anything else. First, there is no drift term on the process. The model is conventionally set down on a quantity that does not drift under the measure it is written under. Dropping the drift keeps the two equations as short as they are, and wherever a drift is needed it is carried separately. Second, the level parameter has no long-run level to return to. The level parameter has no mean reversion at all: it wanders, and where it wanders to is where it stays. The missing pull is the second thing this model gives up.
What does it mean to be organised around the pattern rather than the process?
A model organised around a process is judged by what its process does. Does variance stay positive? Does it pull back? Does the shape of its long-run distribution look like something? Each of those questions has an answer inside the model, and the answer is checkable without looking at a single price.
A pattern organisedBuilt to reproduce a shape across strikes rather than derived from a process and its consequences. model is judged by the shape it lays down across strikes. The process underneath is not fictional, and it obeys the same calculus as everything else in this subject area. The point is that the process was chosen for the shape it delivers, and the shape is the deliverable. The model was never organised to answer a question about the path of the process between two dates, even though it has a process and could be made to mumble an answer.
The everyday version runs as follows. A wall chart of measured temperatures at five recorded hours gives the temperature at those five hours perfectly. The chart was made from them. A curve drawn smoothly through those five marks also gives a reading at half past two. The curve connects two readings that were actually taken, so the half past two figure is real work. Asked about three in the morning, before the first mark, the same chart is extending a shape into a region no reading was taken in. The chart is not lying. Extending a shape is all it does, and there is nothing underneath the shape to keep it honest.
Which of the two models starts from a process, and which starts from a pattern?
What are the four parameters, and what does each one control?
Four numbers, and each one has a job that can be stated in a sentence. The four are not interchangeable. Take them one at a time. A reader who blurs two of them will misread every result that follows.
The level parameterThe overall scale of the randomness at the start, set to 0.20 here, matching the locked volatility of the standard process. is the overall size of the randomness at the outset. The level parameter is set to 0.20, exactly the 20 per cent a year the standard process carries throughout this subject area. The match is neither a coincidence nor a fit. The value was written down deliberately, and the model therefore starts life sitting exactly where the constant treatment sits. Unlike the other three, it is not a fixed constant inside the model. The level parameter has its own equation, and it moves.
The exponentThe parameter deciding how the randomness scales with the level of the process, set to 1 here. decides how that randomness changes as the process moves to a different level. The exponent is the one parameter with no counterpart at all in the previous model, and it spans the two treatments set against each other earlier. A block of its own follows below.
The volatility of volatilityHow random the level parameter itself is, set to 0.30 here, the same numeral the previous model carries in the corresponding role. says how random the level parameter itself is. Set it to nought and the level parameter never moves, and the model collapses back to something with no randomness in its randomness. Set it to 0.30, the value used here, and the level parameter wanders. The number 0.30 is deliberate: the previous model carries the same numeral in the role that corresponds to it.
The correlationHow the level parameter moves with the process, set to minus 0.7 here, the same figure the case parameters carry. says how the two move together. At minus 0.7 the level parameter tends to rise when the process falls. The single negative sign tilts the resulting shape across strikes rather than leaving it symmetric. The same minus 0.7 appears elsewhere in this subject area, where the two locked paths have a quadratic covariation of exactly minus 0.028000 and therefore an implied correlation of exactly minus 0.7.
Three of those four numbers were chosen to match the previous model exactly. The exponent is the only thing that is new. Matching them is the whole point of the setting. On two unrelated parameter sets, every difference in output could have come from any of eight numbers, and a comparison like that teaches nothing. With three held fixed, the fourth is isolated.
How many of the four parameters are the previous model numbers, carried over deliberately?
What does the level parameter do once it is allowed to move?
The level parameter is not a constant. Its own equation has no drift and a volatility of 0.30. Those two facts make it the simplest kind of wandering positive quantity: it cannot reach zero, it has no level it is pulled toward, and its average across all the futures the model contemplates stays exactly where it started.
Work that through with the locked numbers and one useful gap appears. The equation carries no drift, so over the one year horizon the average of the level parameter is 0.200000, exactly its starting value. Its middle value is not 0.200000. The middle value is 0.191199. The average of a quantity that wanders multiplicatively sits above its middle, and the gap here is the factor 0.955997. The average and the middle of the level parameter are two different numbers, 0.200000 against 0.191199, and the difference is produced entirely by the volatility of volatility. Its spread after one year comes to 0.061376.
The standard process itself shows the same shape of gap over the same horizon, where the average finish is Rs 108.33/- and the middle finish is Rs 106.18/-. Same mechanism, different quantity. A reader who has met that gap once already does not need it explained again. Notice instead that the gap is now happening to the volatility rather than to the price. Giving volatility its own randomness means exactly that.
| Quantity, one year on | Figure | Where it comes from |
|---|---|---|
| Level parameter at the outset | 0.200000 | written down, matching the locked volatility |
| Its average one year on | 0.200000 | its equation carries no drift |
| Its middle value one year on | 0.191199 | the starting value times 0.955997 |
| Its spread one year on | 0.061376 | computed from a volatility of volatility of 0.30 |
| Gap between average and middle | 0.008801 | produced entirely by the volatility of volatility |
The level parameter has no drift in its equation. What does that make its average one year on?
What does the exponent decide?
Here is the parameter that has no counterpart in the previous model, and the reason it deserves its own block. The randomness in the process at any moment is the level parameter multiplied by the process raised to the exponent. Everything about how that randomness responds to the level of the process is inside that one power.
| \(D(S)\) | the instantaneous randomness of the process when it sits at level \(S\), in rupees a year |
| \(\alpha\) | the level parameter at that moment, 0.20 at the outset |
| \(S\) | the level of the process, in rupees |
| \(\beta\) | the exponent, a fixed number between nought and one |
Putting the two ends in shows what comes out. At an exponent of 1 the randomness is the level parameter times the level, so a process sitting at twice the level carries twice the randomness. An exponent of 1 is exactly the proportional treatment the standard process assumes everywhere in this subject area, and it is why a 20 per cent volatility means the same thing at Rs 50/- as at Rs 200/-. At an exponent of nought the level raised to nought is one, so the randomness is the level parameter and nothing else, the same size wherever the process happens to be. An exponent of nought is the absolute treatment. The Vasicek and Cox, Ingersoll and Ross rate models covered earlier assume it when they add a fixed amount of randomness regardless of where the rate sits.
| \(\beta=1\) | the proportional treatment, randomness scaling with the level |
| \(\beta=0\) | the absolute treatment, randomness of the same size at every level |
Between them the exponent is neither, and it is not a compromise in any vague sense. The exponent is a precise statement about elasticity: multiply the level by any factor and the randomness is multiplied by that factor raised to the exponent.
| \(c\) | any factor the level is multiplied by |
| \(\beta\) | the exponent, read here as an elasticity |
| \(D(S)\) | the randomness at level \(S\), as defined above |
One parameter runs the whole distance between the two treatments set against each other earlier, and that is why it is worth a block of its own. The proportional and the absolute were presented as different worlds with different equations. Here they are two settings of one dial, and every value in between is a real model that is neither.
What does an exponent of 1 correspond to?
What do the four parameters look like when they are written out on the case numbers?
Here is the whole setting in one place, on the standard process, every value written down by hand.
| Parameter | Value | What it does | Where the number came from |
|---|---|---|---|
| Level parameter | 0.20 | overall size of the randomness | matches the locked 20 per cent volatility |
| Exponent | 1 | how randomness scales with the level | no counterpart in the previous model |
| Volatility of volatility | 0.30 | how random the level parameter is | the previous model number, unchanged |
| Correlation | minus 0.7 | how the two move together | the previous model number, unchanged |
Now read the randomness those settings produce, at three levels of the process. At an exponent of 1 and a level parameter of 0.20, a process sitting at Rs 50/- carries a randomness of 10.000000 rupees a year, at Rs 100/- it carries 20.000000, and at Rs 200/- it carries 40.000000. Four times the level, four times the randomness. Set the exponent to nought instead. The level no longer enters, and the three become 0.200000, 0.200000 and 0.200000, all identical. Set it to a half and they become 1.414214, 2.000000 and 2.828427, where four times the level has produced exactly twice the randomness.
The nine figures are the entire content of the exponent, and every one of them is the level parameter multiplied by a level raised to a power. Nothing was sampled and nothing was fitted. Move the exponent in the control below and watch them recompute.
At an exponent of 0.5, how does the randomness at Rs 200/- compare with the randomness at Rs 50/-?
The exponent is about to be set to nought. Before the control moves: what happens to the randomness at the three different levels?
Move the exponent from nought to one
Held fixed: the level parameter at 0.20, the volatility of volatility at 0.30 and the correlation at minus 0.7. Only the exponent moves. The top panel is the randomness in rupees a year at three levels of the process. The bottom panel is the same randomness read as a percentage of the level it sits at, and it is the mirror image. The track at the foot shows where between the absolute treatment and the proportional treatment the current setting sits. All three redraw together.
At an exponent of 1.00 the randomness is 10.000000, 20.000000 and 40.000000 rupees a year at Rs 50/-, Rs 100/- and Rs 200/-. Four times the level gives four times the randomness, which is the proportional treatment the standard process assumes.
Does holding the level parameter at 0.20 hold anything fixed?
Holding the numeral at 0.20 holds nothing fixed, and the trap sits inside the control above. The units of the level parameter change with the exponent, and the same numeral therefore does not mean the same thing at every setting. Reading the equation dimensionally makes the answer fall out.
| \([\alpha]\) | the units the level parameter is measured in |
| \(\beta\) | the exponent, which appears in the units themselves |
Work that through. At an exponent of 1 and a level parameter of 0.20, the randomness at Rs 100/- is 20.000000 rupees a year. Read as a percentage of the level, that is 20.000000 per cent, the locked volatility exactly. At an exponent of nought and the same 0.20, the randomness at Rs 100/- is 0.200000 rupees a year, or 0.200000 per cent of the level: one hundredth as much. The simulation above holds the numeral fixed and lets the meaning change. Holding it that way is what makes the shape visible, and a reader must not mistake it for a like for like comparison.
So what would a like for like comparison look like? Ask instead which level parameter each exponent needs in order to produce the same randomness at Rs 100/- as the locked volatility does, namely 20 rupees a year. Solve it and the answers are exact: 20.000000 at an exponent of nought, 2.000000 at a half, and 0.200000 at one. Feed those back in and the three settings agree exactly at Rs 100/- and fan out differently away from it. The fanning is the honest picture of the work the exponent performs.
Rescaled to agree at Rs 100/-, which exponent setting gives the largest randomness at Rs 50/-, below where the three meet?
What does being organised around the pattern buy, and what does it cost?
The model buys a great deal, and an account that listed only costs would mislead. Take the gains first. A model organised around a pattern across strikes gives a curve with a small number of handles, each of which moves the curve in a way that can be described: one sets the height, one tilts it, one bends it, and one decides how the whole thing shifts when the process moves to a different level. Four handles that behave that predictably are unusual, and they are why the model is reached for so often where a shape across strikes has to be produced quickly and consistently.
The model also buys speed of a specific kind. Because the model was organised to deliver a volatility for a strike, it does so with an expression rather than with an integral or a lattice or a grid of simulated futures. A model that has to be solved before it answers is a different tool from one that answers directly, and the difference is not academic when the same question has to be answered at many strikes at once.
Now the bill. Two items, and both come straight from the direction of travel.
The first is that the model has little to say about the process. Because the previous model was built out of a description of how variance behaves over time, it earned statements about exactly that. The SABR model has a process written down, but the process was chosen for the shape it delivers. The path the quantity takes between two dates is a matter of mechanism, and the mechanism here is scaffolding for a shape rather than a claim about anything.
The second is that the good behaviour is local. InterpolationFilling in between the strikes a model was built on, which is where a pattern organised model does its most defensible work. between the strikes the model was organised around is real work and it is the work the model is best at. ExtrapolationGoing beyond the strikes a model was built on, which is where a shape is being extended with nothing underneath it. beyond them is a different activity wearing the same clothes. Outside the region it was built on the model does not change, and neither does its confidence. The unchanged confidence is precisely the problem.
Someone needs to know the path of the process between two observation dates. Which model does that call for?
Where does the model break down?
The breakdown has a location, and naming the location is more useful than naming a failure mode. The location is wherever the strikes stop.
Inside the range of strikes a pattern organised model has been given, the model is doing something defensible. The model is connecting readings. The four parameters are constrained on both sides, the shape between two marks is pinned at both ends, and if the shape is wrong in the middle it is wrong by a small amount. Inside the range is the region the model is good at, and being good at it is not nothing.
Outside that range there is no reading on one side. The shape is being carried outward on the strength of its own form. Nothing pins it. The same parameters that pinned the middle are now doing all the work at a distance. A small error in the tilt or the bend inside the range becomes a large error far outside it. The model gives no signal that this has happened. It returns a number in the same format with the same apparent authority, and no measure of fit computed inside the range says anything at all about it.
There is a second and quieter breakdown worth naming. The expression that turns the four parameters into a volatility for a strike is an approximation, and approximations have regions where they are good and regions where they are not. Approximations of this kind get worse for strikes far from the current level and for long horizons, and the reason is structural rather than accidental: the further an expansion is asked to reach, the less an expansion holds. The expression itself, and the size of its error, are covered separately.
What does someone actually do with a model organised this way?
Strip out the vocabulary and the practical use is a familiar one. Someone has a handful of readings and needs a value at a point where no reading exists. A pattern organised model turns those readings into a small set of numbers, and those numbers then produce a value anywhere the reader asks. The four parameters are the compressed form of the readings, and compressing is the useful act: five numbers become four handles that move in describable ways, and the four handles can be compared with the four handles from a different set of readings or a different horizon in a way that five raw numbers cannot.
Compression is the honest use, and it has an honest discipline attached. Before taking a value out of a model like this, the reader asks one question: is the point being asked about inside the readings or outside them? Inside, the answer carries most of the authority of the readings themselves. Outside, it carries the authority of a shape. The same discipline applies to the household version. Given a measurement of the electricity used in a two bedroom home and in a four bedroom home, a rule connecting them will do reasonable work for a three bedroom home. Asked about a hospital, it will still return a number, in the same units, with the same confidence, and the number will be worthless.
The whole practice reduces to knowing where the readings stopped, and that is information the model itself does not carry. It has to be kept alongside, deliberately, by whoever reads the output.
The failure: reading a close fit as evidence
The model reproduces the pattern it was organised around, closely. What has the close reproduction established?
Nothing beyond the fact that it did what it was built to do. The failure does not feel like an error at all, and that is why it catches careful readers. A model that matches its target looks like a model that has been validated, and every instinct trained on ordinary testing says a close fit is evidence.
A close fit is not evidence here, for a structural reason. A pattern organised model was constructed so that a small set of parameters could reproduce that shape. Reproducing it is the design specification being met. The match is the same as a curve drawn through five marks passing through those five marks: true, checkable, and empty of information about anything else.
The damage is not the wrong conclusion; it is the extension that follows from it. A reader who takes a close fit as confirmation will start using the model for questions it was never organised for, and the first such questions are always about the region between and beyond the strikes it was built on. Between and beyond those strikes the performance of the model is unmeasured, and no fit statistic computed inside the range will ever raise a flag about it. The failure is silent by construction.
The model reproduces the strikes it was built on closely. What has that established?
Where does this hold, and where would rules come in?
The mathematics in this guide is universal. Whether a model is organised around a process or around a pattern, and what an exponent does to the scaling of randomness, are not matters of jurisdiction. Conduct duties do apply to what anyone does with a model in any particular place, and those are settled elsewhere.
References
| Source | Document | Where |
|---|---|---|
| arXiv, Quantitative Finance | Preprints on stochastic volatility models and their asymptotic expansions | arxiv.org |
| Social Science Research Network | Working papers on smile modelling and parameter interpretation | ssrn.com |
| Heston, 1993 | A closed-form solution for options with stochastic volatility, the process organised model set against this one | Review of Financial Studies |
| Black and Scholes, 1973; Merton, 1973 | The founding papers behind the constant volatility treatment | Journal of Political Economy; Bell Journal of Economics |
| Vasicek, 1977; Cox, Ingersoll and Ross, 1985 | The short rate models whose randomness does not scale with the level | Journal of Financial Economics; Econometrica |
The standard process and its four locked parameters are invented.
Educational material. Not advice on any investment, tax, budget or market position.
