The Heston Model: Stochastic Volatility With a Closed Form
The Heston model gives the variance its own mean reverting equation, driven by a second randomness that is correlated with the process itself. Five numbers control it. Unlike almost every other treatment of this kind it yields a price by a single integral rather than by repeated trials, and that tractability is why it became the starting point.
Two earlier subjects did the placing. The volatility model families set the three treatments side by side without opening any of them. The comparison of the Vasicek and CIR models opened the mean reverting skeleton on a rate, and showed how a square root in front of the randomness keeps a quantity above nought. The same skeleton hung on a variance instead of a rate is the whole of what the Heston modelA model giving the variance its own mean reverting process with its own randomness, correlated with the process being valued. is.
A volatility number is a measuring instrument, and the instrument is where the account begins. A kitchen scale that reads a little heavy is a biased instrument: the correction is fixed and can be learned once. A scale whose reading wanders about by itself, more when the shelf is knocked and less when the room is still, is a different object entirely. No single correction will serve. A description of how the wandering behaves is needed instead: where it settles, how fast it returns there, how wildly it moves, and whether it wanders more at the moments the thing being weighed is itself moving. The four questions about the wandering are, almost exactly, four of the five parameters below.
Every number below is computed from one invented process. The standard process is a single traded quantity written S with a time subscript, starting at Rs 100/-, drifting at 8 per cent a year, carrying a volatility of 20 per cent a year, watched for one year against a risk-free rate of 5 per cent.
What does this model add to the constant treatment?
Two things, and only two. The first is that the variance stops being a number and becomes a process, with its own equation and its own source of randomness. The second is that this new randomness is allowed to be correlated with the randomness already driving the process. Every extra pattern the model can produce, and there are several the constant treatment cannot produce at all, comes from one of those two additions.
The first addition is worth being exact about. The constant treatment does not merely assume that volatility is steady; it assumes volatility is not a variable of the system. There is nothing to write an equation for. Once the variance gets an equation, the state of the world at a given moment is no longer one number but two, the level and the variance, and everything downstream inherits that. There is now a second thing that can be unknown, a second thing that has to be hedged, and a second thing that has to be given a starting value before anything can be computed at all.
| \(S_t\) | the standard process, the single invented traded quantity, at time \(t\) |
| \(v_t\) | the variance at time \(t\), now a process rather than a number, starting at 0.04 |
| \(\mu\) | the drift under the physical measure P, here 0.08 a year |
| \(\kappa_v\) | the speed at which the variance is pulled back toward its long-run level, here 2.0 a year |
| \(\theta_v\) | the long-run variance the pull aims at, here 0.04 |
| \(\xi\) | the volatility of the variance process, here 0.30 |
| \(W_t,\;W^{v}_t\) | two standard Brownian motions under the physical measure P |
| \(\rho\) | the correlation between the two Brownian motions, here minus 0.7 |
Look at where the square root sits in the second line. The square root is the same device the mean reverting rate model used to hold a rate above nought: as the variance falls toward nought the randomness pushing it about shrinks toward nought too, so the process runs out of the ability to cross the floor at exactly the point it would need it most. The borrowing is not an accident. A variance cannot be negative and still mean anything at all, and the model takes the part that had already solved that problem on a rate.
The second addition, the correlation, is the one readers underweight. The correlation does not change how much randomness there is. The correlation changes whether the two kinds of randomness tend to arrive together. A correlation of minus 0.7 says that on the occasions the level falls hardest, the variance is most likely to be rising, and that single sentence produces every asymmetry that follows.
What two things does the model add to the constant treatment?
What are the five parameters, and what does each one control?
Five numbers, and the useful thing about them is that each is attached to something that can be named and, in three cases, to something that can be watched moving. The naming matters more than it sounds. A model whose parameters do nothing nameable can only be fitted, never reasoned about, and a reader handed such a model has no way to ask whether a chosen value is sensible.
The starting varianceWhere the variance process begins, 0.04 here, whose square root is the 20 per cent this subject area runs on. is where the variance process opens. Here it is 0.04, whose square root is exactly 0.20, so the model begins standing precisely where the constant treatment stands. The long-run varianceThe level the variance is pulled toward, set to the same 0.04 here. is where the pull aims. Here the long-run variance is also 0.04, the same number. The model has nowhere to drift to. The equality of the two is the single most useful choice in these settings, for reasons the worked instance below makes exact.
The speed of reversionHow fast the variance is pulled back toward its long-run level, 2.0 a year here. is how hard the pull acts. At 2.0 a year, a gap between where the variance is and where it is aiming closes by half in the natural logarithm of two divided by two. The half-life is then 0.346574 years, a little over four months. The volatility of volatilityHow random the variance process itself is, 0.30 here. At nought the variance still reverts, but with nothing pushing it about. is how much the variance is pushed about on the way. At 0.30 the variance has a long-run spread, measured as a standard deviation, of exactly 0.030000 around its level of 0.04. Read as a volatility, the spread is a band running from 10 per cent up to 26.457513 per cent. The correlationWhether the variance randomness arrives alongside the level randomness, minus 0.7 here. The correlation is what carries the asymmetry. settles whether the two randomnesses tend to arrive together, at minus 0.7 here.
| \(v_t\) | the variance at time \(t\), the quantity the second equation moves |
| \(v_0\) | the starting variance, 0.04 |
| \(\theta_v\) | the long-run variance, also 0.04, which is why the second term vanishes |
| \(\kappa_v\) | the speed of reversion, 2.0 a year |
| \(\xi\) | the volatility of volatility, 0.30 |
| \(\ln 2\) | the natural logarithm of two, which turns a speed into a half-life |
The flat expectation makes the rest of the comparison a fair test, and is worth pausing on. The average variance is the same number the constant treatment uses, at every date. Anything the model does differently is therefore not a matter of it assuming more volatility on average. The difference is a matter of the variance being uncertain rather than known, and of that uncertainty arriving alongside the movement of the process itself.
Notice the last column. Setting the volatility of volatility to nought does not merely reduce the model a little; it removes the second randomness entirely, leaving a variance that reverts along a fixed path with nothing pushing it. Setting the correlation to nought leaves all the randomness in place and only cuts the tie between the two sources. Killing the second randomness and cutting the tie between the two are very different kinds of nought, and confusing them is the commonest misreading of the parameter list.
Which parameter controls how random the variance itself is?
Why does it have a closed form when most stochastic volatility models do not?
Give the variance an equation of almost any shape and the pricing problem usually stops being solvable on paper. The answer then has to be computed by stepping many paths forward and averaging. Stepping works, but it is slow, noisy and awkward to differentiate. The Heston model is the famous exception, and the reason is not luck. The reason is that the model was constructed backwards, from the property that makes the mathematics work, rather than forwards from a description of how a variance ought to behave.
Here is the property. Do not ask for the price directly. Ask instead for the transform of the logarithm of the process, the average of a complex exponential of the log level. For this particular pair of equations that average turns out to have the form of an exponential whose exponent is linear in the starting variance and in the starting log level. Because the exponent is linear in the state, the partial differential equation the transform must satisfy collapses into a small pair of ordinary differential equations in time alone, and that pair can be solved in terms of a square root, an exponential and a logarithm. Nothing has to be stepped forward. The transform is a formula.
| \(\phi\) | the frequency the transform is taken at, the variable being integrated over |
| \(C,\;D\) | two functions of frequency and time only, each written with a square root, an exponential and a logarithm |
| \(v_0\) | the starting variance, here 0.04, entering the exponent linearly, which is the whole trick |
| \(\Pi_1,\;\Pi_2\) | two numbers between nought and one, each read off the same transform at a different tilt |
| \(K\) | the strike of the contract being valued, here Rs 100/- or Rs 110/- |
| \(f_j\) | the transform above, taken with the first or the second tilt |
| Q | the risk-neutral measure, the one prices are taken under |
Two honest qualifications belong here. The phrase closed formA formula giving the answer directly. Here it means everything is a formula except one integral over a single variable. gets stretched. First, the price is not literally elementary; the integral above still has to be evaluated numerically. Everything inside the integral is closed. Second, the numerical evaluation has its own traps. A naive rule that samples the integrand at a frequency near nought hits a cancellation in the complex square root and returns a wrong answer while looking perfectly healthy, so the integral here is worked out with a composite Gaussian rule. Every price quoted below was checked against a second scheme before it was written down.
What exactly is closed about the closed form here?
What do the locked settings actually produce?
Now the arithmetic. The settings are the five above: starting variance 0.04, long-run variance 0.04, speed 2.0 a year, volatility of volatility 0.30, correlation minus 0.7. The contract is the at-the-money one, strike Rs 100/-, one year, on the standard process.
First, the check that makes everything after it readable. With the volatility of volatility set to nought and everything else left alone, the variance equation loses its random term, so the variance simply reverts toward its long-run level along a fixed path. But it starts at 0.04 and the level it aims at is also 0.04, so the gap is nought and stays nought. The variance never moves. The total variance accumulated over the year is 0.04 times one, or 0.04. The square root of 0.04 is 0.20. At that setting the model returns exactly Rs 10.450584/-, the constant treatment's own price, not approximately and not to within a rounding.
| \(\xi\) | the volatility of volatility, set to nought for this check only |
| \(v_0,\;\theta_v\) | the starting variance and the long-run variance, both 0.04, deliberately equal |
| \(\int v_t dt\) | the total variance accumulated over the horizon, the only thing the constant treatment needs |
| \(T\) | the horizon, one year |
Hold that thought and turn the volatility of volatility back up. Nothing about the average has changed. Because the starting variance and the long-run variance are equal, the expected variance at every future date is still exactly 0.04, and the expected total variance over the year is still exactly 0.040000. And yet the price moves. At a correlation of nought, where the tilt is switched off and only the randomness of the variance remains, the at-the-money contract comes to Rs 10.274631/- against Rs 10.450584/-. Same expected total variance, different price, and the entire difference is the randomness of the variance rather than its level.
| Volatility of volatility | Price at a correlation of nought | Distance from Rs 10.450584/- | Price at the locked minus 0.7 |
|---|---|---|---|
| 0.30, the locked value | Rs 10.274631/- | 0.175953 | Rs 10.394219/- |
| 0.20 | Rs 10.368876/- | 0.081708 | Rs 10.460960/- |
| 0.10 | Rs 10.429613/- | 0.020971 | Rs 10.480852/- |
| 0.05 | Rs 10.445307/- | 0.005277 | Rs 10.472102/- |
| 0.02 | Rs 10.449738/- | 0.000846 | Rs 10.460734/- |
| Nought | Rs 10.450584/- | nil, exactly | Rs 10.450584/- |
Read the second and third columns downward and the collapse is clean: each halving of the volatility of volatility cuts the distance by rather more than half, and the last row is not a limit being approached but an identity being reached. Read the fourth column downward and it is not clean at all. The fourth column goes 10.394219, then up to 10.460960, up again to 10.480852, then back down through 10.472102 and 10.460734 before arriving. With the correlation switched on the approach is not one-directional, and an account that drew it as a tidy convergence would be drawing something these numbers do not do. Two effects are moving at once there, and they do not shrink at the same rate.
The starting variance and the long-run variance are both 0.04. Why the same number?
What does the positivity condition require, and what happens if it fails?
A variance that reaches nought is a variance whose square root is nought, and a process whose randomness has been switched off. Nothing in the mathematics forbids it outright, but a great deal downstream behaves badly when it happens. So there is a comparison worth making before any of the rest, and it is a comparison of two computable numbers rather than a matter of judgment.
The positivity conditionThe requirement keeping the variance strictly away from nought. The condition carries Feller's name and compares two computable numbers., which carries Feller's name from his work on this class of process, asks whether the pull upward is strong enough to beat the randomness pushing downward. Twice the speed multiplied by the long-run variance measures the pull. The square of the volatility of volatility measures the push. At the settings used here the first is 0.160000 and the second is 0.090000, so the condition holds, and it holds with room rather than by a whisker.
| \(\kappa_v\) | the speed of reversion, 2.0 a year, how hard the variance is pulled back |
| \(\theta_v\) | the long-run variance, 0.04, the level being aimed at |
| \(\xi\) | the volatility of volatility, 0.30, how hard the variance is pushed about |
| \(\xi_{\max}\) | the largest volatility of volatility these other settings would tolerate |
| \(\kappa_{\min}\) | the smallest speed these other settings would tolerate |
A failure of the condition is not a collapse and not an error message. The variance becomes able to touch nought and bounce. The model still runs and still prices contracts. Two things change and both are quiet. The variance spends time in a region where the square root makes it extremely sluggish, and the numerical schemes people use to step the model forward start producing negative variances that have to be patched by hand. The condition failing does not break the model; it moves the difficulty from the mathematics into the code, where it is far easier to overlook.
Twice the speed times the long-run variance is 0.160000 and the squared volatility of volatility is 0.090000. What does that comparison establish?
The locked correlation is minus 0.7. Does that make low strikes carry higher or lower implied volatility than high ones?
What does the correlation do, and why does it carry the asymmetry?
Take the two other parameters out of the argument first. The volatility of volatility decides how much the variance moves. The speed decides how quickly it comes back. Neither of them knows anything about direction, so neither can make a fall behave differently from a rise. Only the correlation says anything about which of the two randomnesses arrives with the other, so only the correlation can.
Set it to minus 0.7 and the model is asserting something specific: when the level drops sharply, the variance is more likely than not to be climbing at the same moment. Play that forward over a year and the distribution of where the process ends up stops being even. The biggest falls happen in exactly the conditions where the process is moving most, so the distribution grows a longer tail downward and a shorter one upward. Contracts struck low are then worth more than a single flat volatility would say, and contracts struck high are worth less. Turn the resulting prices back into implied volatilities and the pattern across strikes leans. The lean is precisely the shape the constant treatment cannot produce at any setting whatsoever.
At the locked settings the five contracts on the standard process at strikes of Rs 80/-, Rs 90/-, Rs 100/-, Rs 110/- and Rs 120/- come to Rs 25.044557/-, Rs 17.075310/-, Rs 10.394219/-, Rs 5.430339/- and Rs 2.332633/-. Read those back as implied volatilities and they are 22.961406, 21.357854, 19.849760, 18.456570 and 17.221087 per cent. The lean is unmistakable, and it runs downward from the low strike to the high one. The invented observation set used throughout this subject leans in the same direction. The agreement in direction is not a fit and must not be read as one. The model was not adjusted to match anything; the numbers came out that way from the five parameters chosen in advance, and fitting is a separate subject, set out under calibration and model risk.
The mirror test, which is what symmetry actually means here
Now the part that is easy to state loosely and worth stating exactly. Set the correlation to nought and the pattern across strikes becomes symmetric. Symmetric about what, though? Not about the middle strike of Rs 100/-. The pattern is symmetric about the forward. On the standard process the forward is Rs 100/- multiplied by the exponential of 5 per cent, or Rs 105.127110/-, and the symmetry is in the logarithm of the strike rather than in the strike itself.
The forward gives a test with no wriggle room in it. Pick a log distance, take the strike that far below the forward and the strike that far above it, and compare their implied volatilities. At a correlation of nought, at a log distance of 0.30, the two strikes are Rs 77.880078/- and Rs 141.906755/-, and both return an implied volatility of 20.550213 per cent. At a log distance of 0.05 the two strikes are Rs 100.000000/- and Rs 110.517092/-, and both return 19.530796 per cent. Every pair matches, to every decimal place worth printing. At the locked minus 0.7 the same two pairs return 23.313489 against 15.322741, and 20.572671 against 17.670460. The mirror closes exactly at a correlation of nought and opens as soon as the correlation moves, in whichever direction the correlation moves.
The correlation is about to be set to nought. What happens to the pattern across strikes?
Move the correlation and fold the pattern back on itself
Held fixed and locked: the starting variance at 0.04, the long-run variance at 0.04, the speed at 2.0 a year, the volatility of volatility at 0.30, the level at Rs 100/-, the horizon at one year and the rate at 5 per cent. Only the correlation moves. The upper panel draws the implied volatility pattern against the log distance from the forward of Rs 105.127110/-, with a dashed copy of the same pattern folded about that forward. When the two lie on top of each other the pattern is symmetric. The lower panel shows what each of the five locked strikes is worth in implied volatility against a flat 20 per cent.
At a correlation of minus 0.70 the pattern leans, with the low wing 7.990748 points above the high wing, and the folded copy sits 7.990748 points away from the pattern at its widest. The at-the-money contract prices at Rs 10.394219/- against the constant treatment's Rs 10.450584/-.
Where does a closed form get mistaken for something it is not?
The error that gets made, and what it costs
Treating the closed form as evidence that the model is right. The closed form is evidence that the model is tractable. Tractability is a different property entirely, and it is the reason this model became the one everybody opens first. A reader who slides from one to the other will prefer it over treatments that describe the same thing better and compute worse, without ever having made that trade in the open.
The tell is easy to spot once looked for. The reasons usually given for choosing this model turn out, on counting, to be mostly about the mathematics. The model solves quickly. The model differentiates cleanly, so the sensitivities come out as expressions. Calibration takes seconds rather than hours. A large body of published work stands behind the model. Every one of those is a statement about how the model behaves under a computer, and not one of them is a statement about whether a variance behaves the way the second equation says it does.
The cost is a model chosen on a criterion nobody wrote down. The cost compounds quietly. The tractability that made the model attractive also makes it the default everything else is measured against, so the alternatives arrive already carrying the burden of proof. One question carries into any model discussion: which of these reasons would still hold if the computation were free?
The model has a closed form. What does that establish about it?
What does the model still not do?
Three things, and they are worth naming plainly because the model is often treated as though it had closed the subject.
The first is a sudden move. Both equations here are continuous: the level and the variance each get from where they were to where they are by passing through everything in between. A move too large to have been reached one small step at a time is outside this model entirely, whatever the parameters are set to. Representing one needs different machinery, a counting process with an arrival rate and a size distribution. The invented settings this subject area uses for that put the rate at 0.5 arrivals a year, making the chance of no arrival at all across the year 0.606531 and the chance of at least one 0.393469. A jump is not a large diffusion, so nothing in the five parameters can imitate one.
The second is anything the five parameters are not allowed to depend on. The speed, the long-run variance, the volatility of volatility and the correlation are all constants here. The four constants do not change with the level, they do not change with the date, and they do not change when conditions do. A model in which the correlation itself moves is a different model, not this one with different numbers put in.
The third belongs to an unfitted account of the model rather than to the model itself. The five settings were chosen in advance so that the arithmetic reconciles, and the pattern they produce leans in the same direction as the invented observation set without having been asked to. Choosing parameters so that a model agrees with a set of prices is a separate task with its own failure modes, and is covered separately. A model and a fitted model are different objects.
Name something this model still cannot represent.
How does somebody choosing between models actually use all this?
Very few readers will ever code this model. Rather more will be handed a valuation, a risk report or a note that says a stochastic volatility treatment was used, and will have to decide how much weight to put on it. There are three questions that do most of the work, and all three come straight out of this guide.
The first is what the five numbers were. Nobody can judge cold whether the numbers look reasonable. The question is whether they were stated at all. Everything the model claims lives in those five values, so a model quoted without its parameters is a model quoted without its content. If the starting variance and the long-run variance differ, the model is asserting that conditions today are not conditions in general, and that assertion should be visible somewhere in the writing rather than buried in a settings file.
The second is whether the positivity comparison was made. The check is two multiplications and one comparison, so there is no excuse for leaving it out, and an absence usually means nobody looked at the parameters as a set. At the settings used here it reads 0.160000 against 0.090000 and it passes, but a set of numbers fitted to a stubborn observation set can quietly walk over that boundary, and when it does the trouble surfaces as odd numerical behaviour rather than as an error.
The correlation carries the shape, so the third question is about the correlation. A model reported at a correlation near nought is reporting a symmetric pattern, and if the prices it is being compared against are visibly asymmetric then either the model has not been fitted or the fit has gone somewhere strange. The question can be asked without any software at all: which way does the reported pattern lean, and does the sign of the correlation agree with it?
A last, plainer point. The whole reason the model can be checked at all is that one setting collapses it onto a figure that was already known. A report using this model that cannot show the case in which it agrees with the simpler treatment offers no anchor at all, and every number in it has to be taken on trust.
Where this holds, and where the rules would come in
The mathematics here is universal. A pair of equations, a correlation and a positivity comparison are not a matter of jurisdiction. Conduct duties do apply to what anyone does with a model in any particular place, and those are settled elsewhere.
References
| Source | Document | Where |
|---|---|---|
| arXiv, Quantitative Finance | Preprints on stochastic volatility models and transform pricing methods | arxiv.org |
| Social Science Research Network | Working papers on the numerical evaluation of transform integrals | ssrn.com |
| Steven L. Heston | A Closed-Form Solution for Options with Stochastic Volatility, 1993, the paper the two equations and the transform are taken from | The Review of Financial Studies |
| William Feller | Two Singular Diffusion Problems, 1951, where the positivity condition on this class of process is set out | Annals of Mathematics |
| John Cox, Jonathan Ingersoll and Stephen Ross | A Theory of the Term Structure of Interest Rates, 1985, the source of the square root mean reverting skeleton borrowed here | Econometrica |
| Fischer Black, Myron Scholes and Robert Merton | The 1973 papers giving the constant volatility price of Rs 10.450584/- that the checks here collapse onto | Journal of Political Economy; Bell Journal of Economics and Management Science |
The standard process, the five volatility settings, the locked contracts and the five observed volatilities are invented.
Educational material. Not advice on any investment, tax, budget or market position.
