Constant, Local and Stochastic Volatility Compared
Constant volatility assumes one number for all strikes and all times. A local treatment lets volatility depend on the level and the time but stay determined by them. A stochastic treatment gives volatility its own randomness. Each step gives up an assumption, buys the ability to match something the previous one could not, and costs tractability.
The ladder is one assumption being given up in stages. Each rung buys the ability to reproduce a pattern the rung below cannot reproduce at all. Each rung costs closed form: fewer things can be written down, and more of the answer has to be computed rather than read off. Every named model in this subject area stands on one of those three rungs, and knowing which rung a model stands on settles most of what it can and cannot do.
A ruler is a better starting point than a model. Constant volatilityOne number for every strike and every time, which is what the standard model assumes. is one ruler, used for everything, regardless of what is being measured. Local volatilityVolatility determined by the level and the time, with no randomness of its own. is a wall chart of corrections: the correction depends on the position and the moment, but given those two facts it is fixed. Stochastic volatilityVolatility carrying its own randomness, not determined by the level. is a ruler that expands and contracts on its own, for reasons the thing being measured does not explain. Three rulers, three claims about what measurement is, and only the first is a single number.
What does each of the three treatments assume?
The object the whole of this subject area is built on comes first. The standard process is a single traded quantity written S with a time subscript, invented for teaching, starting at Rs 100/-, drifting at 8 per cent a year, carrying a volatility of 20 per cent a year, observed over one year against a risk-free rate of 5 per cent.
Under the first treatment the volatility in that description is a constant. The constant has no arguments. The constant does not know the strike, the constant does not know the date, and the constant does not change when the level changes. A fixed volatility is not a simplification bolted on afterwards; a fixed volatility is the assumption that makes the closed form exist at all.
Under the second treatment the volatility becomes a function of two things the model already carries: the level of the process and the clock. Write down the level and the time, and the volatility is settled. There is nothing else to know. A local treatment enlarges what volatility depends on without giving it any randomness of its own.
Under the third treatment volatility stops being determined by anything the process already carries. Volatility gets its own equation, its own noise, and its own correlation with the process. Once it has that, the level of the process no longer settles it, and knowing exactly where the process is settles strictly less than it did before.
What is the Black-Scholes Model's treatment, and what does a single number buy?
The Black-Scholes ModelThe model built on the constant treatment. Named for Black, Scholes and Merton, 1973., from Black, Scholes and Merton in 1973, takes the constant treatment and takes it seriously. The standard process moves by a drift term and a diffusion term, and the coefficient in front of the diffusion term is one number that never changes.
| \(S_t\) | the standard process, the single invented traded quantity, at time \(t\) |
| \(\mu\) | the drift under the physical measure P, here 0.08 a year |
| \(\sigma\) | the volatility, a constant with no arguments, here 0.20 a year |
| \(W_t\) | standard Brownian motion under the physical measure P |
A single number buys everything that made the model usable. One number means one formula, differentiable in every argument, so the sensitivities fall out as expressions rather than as numerical experiments. On the locked at-the-money contract, strike Rs 100/- and one year, the intermediate quantities are exactly 0.350000 and 0.150000, the two normal probabilities are 0.636831 and 0.559618, and the call comes to Rs 10.450584/-. The matching put is Rs 5.573526/-, and the gap of Rs 4.877058/- between the two agrees with the level less the discounted strike to six decimal places.
The whole of that convenience is bought with the assumption that one number serves every contract, and that assumption is exactly what an observation set can contradict.
Five invented contracts at five strikes follow, priced against one flat volatility. The misses at the low strikes and the misses at the high strikes: same sign, or opposite?
The worked instance: five invented observations at one horizon
Here is the observation set this whole reading order is built against. Five strikes on the standard process, one year, quoted the way such things are quoted, as volatilities rather than as prices. The five numbers were chosen for teaching rather than read off a market, and they exist only so that a model has something to be compared against.
Twenty four per cent at the Rs 80/- strike, twenty two at Rs 90/-, twenty at Rs 100/-, nineteen at Rs 110/- and eighteen and a half at Rs 120/-. Each of those turned into a price with the standard formula produces the middle column below. Pricing the same five contracts with one flat twenty per cent gives the comparison.
| Strike | Invented implied volatility | Price it implies | Price at a flat 20 per cent | Miss |
|---|---|---|---|---|
| Rs 80/- | 24.0 per cent | Rs 25.227000/- | Rs 24.588835/- | minus Rs 0.638165/- |
| Rs 90/- | 22.0 per cent | Rs 17.257579/- | Rs 16.699448/- | minus Rs 0.558131/- |
| Rs 100/- | 20.0 per cent | Rs 10.450584/- | Rs 10.450584/- | nil |
| Rs 110/- | 19.0 per cent | Rs 5.644765/- | Rs 6.040088/- | plus Rs 0.395324/- |
| Rs 120/- | 18.5 per cent | Rs 2.745149/- | Rs 3.247477/- | plus Rs 0.502329/- |
Read the last column downward. Two misses below, one exact, two misses above. The model is cheap where the invented volatilities are high and dear where they are low, and the crossing happens once, in the middle, where the flat number was chosen to sit. The pattern is not error and it is not noise: error scatters, and a run of misses that changes sign exactly once does not.
At a flat twenty per cent, how many of the five invented strikes are matched exactly?
The control below accepts any single volatility from fifteen to twenty five per cent. Before it is moved: can all five strikes be made to match at once?
One number against five observations
Held fixed: the level at Rs 100/-, the horizon at one year, the rate at 5 per cent, and the five invented implied volatilities of 24.0, 22.0, 20.0, 19.0 and 18.5 per cent. The only thing that moves is the single volatility the constant treatment is allowed to use. The top panel is the volatility view, the bottom panel is the price view, and both redraw together.
At a flat 20.0 per cent, one of the five invented strikes is matched exactly. The largest miss is Rs 0.638165/- cheap at the Rs 80/- strike, and the pattern across the five reads cheap, cheap, exact, dear, dear.
What does a local treatment do differently?
The local treatment does the smallest thing that could possibly work. The local treatment leaves the equation alone and replaces the constant with a function of two arguments the model already has in hand: where the process is, and which date it is.
| \(\sigma_{\mathrm{loc}}\) | a function of two arguments only, returning a volatility |
| \(S_t\) | the level of the standard process, the first argument |
| \(t\) | the date, the second argument |
| \(W_t\) | the same single Brownian motion as on rung one, under P |
Notice what has and has not changed. There is still exactly one Brownian motion in the equation. The process is still driven by one source of randomness. The coefficient in front of that single source of randomness is now allowed to vary, and to vary in a way the model can read off from its own state.
The extra freedom is enormous. A function of two arguments has room to hold an entire observation set, so all five of the invented strikes above can be matched exactly rather than one of them. The local treatment buys an exact match at the price of turning volatility from a description into a lookup.
The wall chart is the honest picture of it. A constant treatment stamps the same figure into every cell of the chart. A local treatment fills every cell separately. The chart can be made to agree with anything observed. Agreement is a strength when agreement is wanted and a weakness when an explanation is wanted, and a chart that agrees with everything explains nothing about why.
What does a stochastic treatment add that a local one cannot?
Now give up the last thing. On rung three, volatility is not read off the state at all. Volatility has an equation of its own, driven by a Brownian motion of its own, correlated with the one already in the process. The variance is what appears in the pricing arithmetic, so the quantity modelled is usually the variance rather than the volatility, written v with a time subscript.
| \(v_t\) | the variance process, whose square root is the volatility at time \(t\) |
| \(\kappa_v\) | the speed at which the variance is pulled back toward its long-run level |
| \(\theta_v\) | the long-run level the variance is pulled toward |
| \(\xi\) | the volatility of the variance, the size of its own randomness |
| \(W^{v}_t\) | a second Brownian motion, driving the variance |
| \(\rho\) | the correlation between the two Brownian motions |
Count the Brownian motions. Rung one has one. Rung two has one. Rung three has two, and that single change is the whole difference. A local treatment can make volatility vary; only a stochastic treatment can make it vary for reasons the level does not contain.
The everyday version is the difference between two scales. One reads differently depending on where in the room it is put, and a wall chart can record that. The other drifts on its own overnight, and no chart can record it. Both give a different number tomorrow. Only one of them gives a different number tomorrow with the object in the same place.
The locked parameter set used across this reading order starts the variance at 0.04, whose square root is exactly the 20 per cent volatility of the standard process, and sets the long-run variance to the same 0.04, so the model has nowhere to drift to. The speed is 2.0 a year, the volatility of the variance is 0.30 and the correlation is minus 0.7. With the volatility of the variance turned down to zero the whole apparatus collapses back to the constant treatment and returns Rs 10.450584/- exactly, so any later departure from that figure is attributable to one parameter and not to a coincidence.
What can a stochastic treatment represent that a local one cannot represent at all?
Implied vs Local vs Stochastic Volatility: one word, or three objects?
Three things in this guide have carried the word volatility and they are not the same kind of thing. One is an output, one is a function inside a model, and one is a process. Confusing any two of them is the commonest error in this whole area, and it is easy to make because the word does not change.
Implied volatilityThe number backed out of a price, which is an output rather than an assumption. is not an assumption at all. An implied volatility emerges when a given price is taken, the standard formula is held fixed, and the one number that would have produced that price is solved for backwards. An implied volatility is a restatement of the price in different units.
| \(C^{\mathrm{BS}}\) | the constant volatility formula, taken as a fixed function |
| \(C^{\mathrm{obs}}\) | a given price, here one of five invented observations |
| \(K\) | the strike, one of Rs 80/-, Rs 90/-, Rs 100/-, Rs 110/- and Rs 120/- |
| \(\sigma_{\mathrm{imp}}\) | the number that makes the two sides agree, one per strike |
A set of implied volatilities that differ across strikes is therefore a contradiction rather than a curiosity. Each of the five invented numbers is the volatility the constant treatment would need if that contract were the only one in the world. Five different answers to a question that the model says has one answer is the model stating, in its own units, where it disagrees with the observations.
Implied volatility and local volatility. Are they the same object?
What is a Hedge Ratio under each treatment, and why does it differ?
The hedge ratio is where the difference between the three stops being philosophical. A hedge ratioThe quantity of the process offsetting a contract, which differs under each treatment. is the quantity of the process held against a contract so that small moves cancel. Under the constant treatment it is the derivative of the contract value with respect to the level, and on the locked at-the-money contract that number is 0.636831.
Now suppose the volatility attached to a contract changes when the level changes, exactly as rungs two and three allow. The total change in the contract value for a small move in the level then has two parts, not one: the direct part, and the part that arrives through the volatility.
| \(C\) | the value of the contract, taken as a function whose curvature matters |
| \(\partial C/\partial S\) | the sensitivity to the level with volatility held still, 0.636831 here |
| \(\partial C/\partial \sigma\) | the sensitivity to volatility, Rs 0.375240/- per point here |
| \(d\sigma/dS\) | how much the volatility attached to this contract moves per rupee of level |
Put the locked figures in and it stops being abstract. Across the invented observation set the volatility falls from 22.0 per cent at the Rs 90/- strike to 19.0 per cent at the Rs 110/- strike. The fall is 3.0 points across 20 rupees of strike, or 0.15 points a rupee. If the shape of that set stays put relative to the level rather than relative to the strike, then a one rupee rise in the level lifts the volatility attached to a fixed strike by that same 0.15 points. The sensitivity to volatility on the at-the-money contract is Rs 0.375240/- per point, so the second part contributes 0.375240 multiplied by 0.15. The product is 0.0562860, and the hedge ratio moves from 0.636831 to 0.693117.
Because the two treatments disagree about whether volatility moved, the same contract at the same level over the same horizon needs two different quantities held against it. The figure 0.693117 follows from one stated convention about how the shape moves. Each named model reaches its own figure by its own route, and each is set out separately.
Rungs one and two share something the picture above makes plain. Both are driven by a single Brownian motion, so a position in the process can in principle cancel every random move in the contract. Rung three has two, and one position cannot cancel two independent shocks. The hedge ratio against the level does not become wrong; it becomes insufficient.
Why does the hedge ratio differ under a stochastic treatment?
Heston Model vs SABR Model: what separates them, and what do they borrow?
Four models are named across this reading order and each is covered separately. Placing the four against each other is the useful step, and placement is what a map is for. Two axes do it: what the process being modelled describes, and what each one is remembered for.
The Heston ModelA stochastic treatment with a closed form, opened separately in this reading order., from Heston in 1993, models a variance that is pulled back toward a long-run level and shaken by its own noise, correlated with the process. The Heston Model is remembered for still yielding a price without a full numerical grid: a transform based expression that reduces the work to a single integral rather than a simulation.
The stochastic alpha, beta, rho model (SABR) models a volatility with its own randomness too, but is remembered for something else. The SABR Model is remembered for the shape. Built so that the pattern across strikes comes out in a form that can be written down and reasoned about directly, the SABR Model is the one reached for whenever the object of interest is the shape itself rather than a single contract.
Where the Vasicek Model and the CIR Model sit
The other two describe a rate rather than a volatility, and they are named in this guide for one reason: the machinery is the same. The Vasicek ModelA mean reverting model for a rate, named here and opened separately in this reading order., from Vasicek in 1977, pulls a short rate back toward a long-run mean at a stated speed and is remembered for giving a closed form. The Cox-Ingersoll-Ross model (CIR), published by those three authors in 1985, keeps the same pull and adds a square root in front of the noise. Because the noise fades to nothing exactly where the level does, the process cannot cross below zero, and that enforced shape is what the CIR Model is remembered for.
The square root is the borrowing that puts two rate models into an account of volatility. A variance also cannot be negative, and the same device that keeps a rate above zero keeps a variance above zero. The square root device is precisely the diffusion sitting in the second equation on rung three. The locked parameter set shows the condition holding with room: twice the speed multiplied by the long-run variance is 0.16, against a squared volatility of the variance of 0.09.
Two of the four named models describe something other than a volatility. Which two?
The error that gets made, and what it costs
Treating the three treatments as increasingly accurate versions of one model. The three are different assumptions about what volatility is, and moving up the ladder is replacement rather than refinement. A local treatment can match every observed price exactly while still describing volatility as something the level determines. A stochastic treatment flatly denies exactly that. The two do not differ about how accurate to be. The disagreement is about what volatility is.
A reader who takes the ladder as accuracy will make two further assumptions without noticing. First, that the top rung contains the others as special cases, so nothing is lost by climbing. Second, that a better fit means a truer model, so the comparison between treatments can be settled by measuring residuals.
Neither follows, and the cost is choosing a model on fit alone. Two treatments that agree on every price observed can disagree completely about how the contract value moves when the level moves. A hedge ratio is exactly that quantity. Choosing a model on fit alone is the error that fitting and model risk are built around, and it starts here, with a ladder read as a ranking.
A local treatment matches every observed price exactly. Is it therefore truer than a stochastic one?
How does somebody reading a valuation actually use this ladder?
Almost nobody reading a valuation gets handed the model. A number with the word volatility beside it arrives instead, sitting in a note, and the first useful question is not whether it looks reasonable. The first useful question is which of the three objects that number is.
If it was an input, somebody chose it, and the ladder states what choosing one number commits the chooser to: agreement at one point and a systematic pattern of misses elsewhere. If it was an output backed out of a price, it is not an assumption at all and it carries every assumption of the formula used to extract it, so quoting it as a property of the thing being valued is a category error. If it is one entry from a chart or one path of a process, then a single number in a note has compressed away the object that matters, and the honest thing is to ask for the shape rather than the point.
A valuation and a hedge do not fail in the same way, so the second question is what the number was used for. A valuation reads one number off the middle of an observation set and can survive being a little wrong there. A hedge asks what happens when the level moves, and that is exactly the question on which two treatments that agree on every price can give different answers. The Rs 0.056286 difference in the hedge ratio computed above is small on one contract and is not small on a position that is rebalanced repeatedly.
The third question is the one that costs the most when it is skipped, and it is the plainest: what did this model assume volatility is? Not how well it fitted. A sheet that records the fit and leaves the assumption line blank has compared the wrong thing, and everything downstream inherits that.
Where this holds, and where the rules would come in
The mathematics here is universal. A ladder of assumptions about volatility is not a matter of jurisdiction. Conduct duties do apply to what anyone does with a model in any particular place, and those are settled elsewhere. A traded level, a quoted volatility and an exchange convention are facts about one market on one day; a ladder of assumptions holds whatever those facts turn out to be.
References
| Source | Document | Where |
|---|---|---|
| arXiv, Quantitative Finance | Preprints on volatility modelling and derivative pricing | arxiv.org |
| Social Science Research Network | Working papers on local and stochastic volatility | ssrn.com |
The standard process, the locked contracts and the five observed volatilities are invented.
Educational material. Not advice on any investment, tax, budget or market position.
