The Volatility Process: Modelling Volatility as Random
Treating volatility as a process means giving it its own equation, its own randomness and its own parameters, rather than leaving it as a number attached to the price process. Such a process must stay above nought, because a negative variance means nothing. And unlike a level, it cannot be observed at all: only its consequences can.
Everything difficult about this area comes from one asymmetry, and it is worth stating before any notation appears. A level can be looked at. A volatility cannot. Where the standard process stood at the end of month nine can be written down. The number is simply there. Because nothing anywhere is that number, the volatility at the end of month nine cannot be written down at all. Every quantity that stands in for volatility is an inference from something else, and every inference carries its own error.
The everyday version runs as follows. A kitchen scale reads what an object weighs. The scale gives a number, and the number can be read off. Measuring how shaky the hand is while it holds the object over the scale is another matter entirely. Nothing on the dial is the shakiness. The reading can be watched as it jitters and something computed from the jitter, and that computed something is not the shakiness either: it is a summary of the jitter over whatever stretch happened to be watched, contaminated by how heavy the object was and how long the watching lasted. Volatility is the shakiness. The reading on the dial is the level. Everything below follows from giving the shakiness an equation of its own.
The standard process is the invented single traded quantity this whole subject area is built on, written S with a time subscript. The standard process starts at Rs 100/-, drifts at 8 per cent a year, carries a volatility of 20 per cent a year, and is observed over one year against a risk-free rate of 5 per cent. The locked path is the twelve step path published for this subject area so that any treatment needing a path draws the same one: it starts at Rs 100/-, dips to Rs 93.74/- at month nine, peaks at Rs 111.08/- at month six, and finishes at Rs 106.18/-.
What does it mean to treat volatility as a process?
In the constant treatment, volatility is a slot inside another equation. The slot is a coefficient sitting in front of the noise term of the price equation. The coefficient has no equation of its own, no parameters of its own, and no way of moving except by being replaced by hand. Ask it what it will be next month and the question does not parse. A coefficient is not the kind of object that has a next month.
A volatility processVolatility given its own equation, its own randomness and its own parameters, rather than sitting as a coefficient in another equation. is a different kind of object entirely. Volatility gets promoted out of the coefficient slot and given a line of its own. Once it has a line of its own it needs everything a line needs: a starting value, a drift, a diffusion term, and a source of randomness to drive it. The promotion is not a refinement of the constant treatment; it is a change in what sort of thing volatility is taken to be.
Two conventions are worth settling before the equations arrive. First, the equation usually goes to the variance rather than the volatility, written v with a time subscript. The variance is what the mathematics of the price equation actually consumes, and the square root that keeps it positive is cleaner to write on the variance. Second, the driver of the variance is a second Brownian motion, not the same one that drives the price, and the two are allowed to be correlated. Everything below rests on both conventions.
| \(S_t\) | the level of the standard process at time t, in rupees |
| \(v_t\) | the variance rate at time t, whose square root is the volatility |
| \(\mu\) | the drift of the standard process, 0.08 a year |
| \(\kappa_v\) | the speed at which the variance is pulled back, 2.0 a year |
| \(\theta_v\) | the level the variance is pulled toward, 0.04 |
| \(\xi\) | the size of the randomness in the variance, 0.30 |
| \(W_t\) | the Brownian motion driving the level, under the physical measure P |
| \(W^{v}_t\) | a second Brownian motion, driving the variance |
| \(\rho\) | the correlation between the two drivers, minus 0.7 |
The second line reads the way the first one does. The variance line has a drift, and the drift is not a constant: it is a pull back toward a level, and the further the variance is from that level the harder the pull. The line also has a noise term, and the noise term is scaled by the square root of the variance itself, so a variance near nought gets almost no push. And it has its own driver, related to the price driver but not identical to it.
At the locked settings the pull is toward 0.04, and the variance starts at 0.04. So the model has nowhere to drift to. The coincidence is deliberate rather than lazy. If the size of the randomness is set to nought, the variance never leaves 0.04, the volatility never leaves 20 per cent, and the whole apparatus collapses back to the constant treatment exactly. Every departure from the constant answer is therefore attributable to one parameter, and one attributable parameter is what makes this set of numbers a teaching object rather than a demonstration.
What must a volatility process satisfy that a price process need not?
A price process has to be handled carefully, but it does not have to stay anywhere in particular. A variance process does. A variance is a squared quantity. There is no number whose square is negative, so a negative variance is not an unlikely outcome or an uncomfortable one; it is not an outcome at all. The requirement here is not a preference about behaviour, it is a requirement about the object.
A lift and a thermometer show the difference. A thermometer can read minus four and the reading means something. A lift cannot be on floor minus four in a building whose lowest floor is the ground, and if the display shows it, the display is broken rather than the building being strange. PositivityThe requirement that a process stays above nought. No number has a negative square, so a variance has to satisfy it. for a variance process is the lift, not the thermometer.
The mechanism that enforces it is the square root sitting in front of the noise term. As the variance falls toward nought, the square root falls toward nought with it, so the size of the random push shrinks exactly where a random push would be most dangerous. Meanwhile the pull back toward the long-run level does not shrink; near nought it is at its strongest upward. The two effects race each other, and whether the upward pull wins is a condition that can be written down and checked.
| \(\kappa_v\) | the speed of the pull back toward the long-run level, 2.0 a year |
| \(\theta_v\) | the long-run level the variance is pulled toward, 0.04 |
| \(\xi\) | the size of the randomness in the variance, 0.30 |
| \(2\kappa_v\theta_v\) | twice the strength of the upward pull at the boundary, 0.160000 |
| \(\xi^{2}\) | the strength of the randomness at the boundary, 0.090000 |
The margin matters as much as the verdict. Twice the speed times the level is 0.160000 and the squared size of the randomness is 0.090000, so the left side is 0.070000 clear of the right. The clearance of 0.070000 is not a comfort blanket; it is the answer to a different question, namely how close the variance is allowed to get to nought before the pull takes over. A condition that just barely holds describes a process that hugs the boundary, and a condition that fails describes a process that reaches nought and has to be told what to do next.
Why must a variance process stay positive?
Why can volatility not be observed at all?
Now the hard part, and it is hard in a way that no amount of better data fixes. Point at the observation and ask which number in it is the volatility. There is no such number. The observation carries levels: Rs 97.64/- at month one, Rs 107.63/- at month two, and so on down to Rs 106.18/- at month twelve. Each of those is a reading. None of them is a volatility, and no arithmetic on them recovers one. Volatility is a property of the mechanism generating the readings rather than a property of any reading.
Volatility is therefore latentPresent in the model and absent from the data. A quantity of that kind is hard to pin down.: present in the model and absent from the data. The model has a symbol for it. The observation has no column for it. The mismatch is not a gap in the record. Volatility simply is not the kind of thing a record holds. A process that is unobservableNever directly available in any observation, however complete or frequent that observation is. stays unobservable however finely it is observed.
The two questions sit honestly side by side. Where was the standard process at month nine? Rs 93.74/- exactly, read off the locked path, and there is nothing to argue about. And the volatility of the standard process at month nine? The honest answer is a shrug followed by a method, and the method will produce a number that depends on choices somebody has to make: over what stretch, at what spacing, with what model in the background. Two careful people answer the first question identically and the second question differently.
What, in an observation of the standard process, is the volatility?
What gets measured in its place, and how far off is it?
Since the quantity itself is unavailable, something else is measured and used in its place. Two stand-ins dominate. The first is realised variationA quantity computed from the observed movement of the level over a stretch, used as a stand-in for the volatility over that stretch., computed from the movement of the level: chop the stretch into steps, take the change in the logarithm over each step, square each one, add them up. The second is implied volatilityA number backed out of an observed price by inverting a pricing formula, so it inherits every assumption of that formula., backed out of a price by running a pricing formula in reverse. Neither is the process, and both carry error that belongs to the measurement rather than to the model.
Take the first one. It is the closest available measurement, and its error can be worked out exactly rather than estimated. The exact question is the following. The standard process has a variance rate of 0.040000 a year by construction. The realised sum on the locked path at its twelve step partition is computed next. Do the two agree?
Before the next block. If the realised sum on the locked path does not come to 0.040000 exactly, is the difference a mistake in the arithmetic or a property of how the measurement was taken?
The worked instance: the closest measurement there is, and its gap
The locked path is built from twelve driving values, each multiplied by 0.288675, the square root of one twelfth. The twelve values sum to nought exactly and their squares sum to 12.0 exactly. Both properties are construction rather than luck, and both are what make the arithmetic below exact rather than approximate.
| \(\widehat{QV}_n\) | the realised sum over a partition of n steps, the stand-in for the variance rate times the elapsed time |
| \(n\) | the number of steps the year is divided into, twelve on the locked path |
| \(t_i\) | the end of step i, so that the whole set of them partitions the year |
| \(\ln S_{t_i}\) | the natural logarithm of the level at the end of step i |
Now do it. On the locked path the answer is 0.040300, and the true variance rate over the year is 0.040000. The two do not agree, and they were never going to. The reason is entirely visible once the sum is broken apart.
| \(\sigma^{2}T\) | the variance rate times the elapsed time, 0.040000, which is what the measurement is standing in for |
| \(\mu-\tfrac{1}{2}\sigma^{2}\) | the drift of the logarithm, 0.060000 a year, so 0.005000 over each of twelve steps |
| \(n\) | the number of steps, twelve here |
| \(\Delta W_i\) | the Brownian increment over step i, and the twelve of them sum to nought by construction |
The realised figure at twelve steps is 0.040300, and 0.040000 is the limit rather than the reading. The distinction between a limit and a reading is the whole of it. The excess of 0.000300 is not a rounding artefact and not a slip; it is what the sum of squared steps has to equal when each step carries a drift. Refine the partition and the excess shrinks. The drift per step shrinks faster than the number of steps grows: n copies of a per-step drift of 0.060000 divided by n, each squared, comes to 0.003600 divided by n.
| Steps in the year | Drift carried per step | Excess it contributes | The measurement reads | The quantity wanted |
|---|---|---|---|---|
| 12 | 0.005000 | 0.000300 | 0.040300 | 0.040000 |
| 52 | 0.001154 | 0.000069 | 0.040069 | 0.040000 |
| 252 | 0.000238 | 0.000014 | 0.040014 | 0.040000 |
| 2,520 | 0.000024 | 0.000001 | 0.040001 | 0.040000 |
| In the limit | nought | nought | 0.040000 | 0.040000 |
Read the last row carefully. It is the only row where the two columns agree, and it is not a row anybody ever observes. Every finite partition sits in one of the rows above it, and every one of those rows has a positive excess. The gap goes to nought and never arrives there.
The second stand-in has the same character and a different source of error. A number backed out of a price is an output of whatever formula was inverted to get it, so it carries every assumption in that formula. Invert a formula that assumes constant volatility and the number that comes out is the constant volatility that would have produced the price. The number is well defined, and it is not the value of any process at any instant. An implied number is closer to a quotation convention than to a measurement, and treating it as a reading of the process is the second half of the same mistake.
Name the two quantities most commonly used in place of the volatility itself.
The closest available measurement reads 0.040300 against a true 0.040000. The calculator below offers a control that refines the partition. Before it is used: does the gap vanish?
Refine the ruler and watch the gap shrink without closing
One control: how many steps the year is divided into, taking twelve, fifty two, two hundred and fifty two, and two thousand five hundred and twenty. Everything else is held at the locked settings: drift 8 per cent, volatility 20 per cent, one year, the same locked path throughout. The top strip is the scale anybody would actually draw. The second strip is the same axis magnified one hundred times, where the realised marker walks toward the true one as the partition is refined. The third strip is the ruler itself getting finer. The fourth places the remaining excess on a scale of powers of ten. The excess travels from three ten thousandths to one millionth, and a straight scale would show nothing at the far end.
At twelve steps each step of the logarithm carries a drift of 0.005000, so the measurement reads 0.040300 against a true variance rate of 0.040000 and the excess is 0.000300. That is the figure in the worked instance above.
What does making volatility random do to a hedge?
The promotion stops being a modelling preference here and starts costing something. The argument that produces a price in the constant world works by construction: hold the contract, hold a quantity of the underlying process against it, choose that quantity so the two random moves cancel, and what is left over has no randomness in it at all. One source of randomness, one instrument to offset it, nothing remaining.
Add a second source of randomnessA second Brownian motion in the model. Volatility becomes one of these when it is made a process rather than a coefficient. and the counting changes. There are now two independent pushes and still only one instrument being used against them. Choosing the quantity to cancel the first push does exactly that and leaves the second push untouched. The level of the process has no way of reaching a randomness the level does not drive.
| \(\Pi_t\) | the value of the position holding the contract and the offset together |
| \(f\) | the value of the contract as a function of the level, the variance and time |
| \(\Delta\) | the quantity of the underlying process held against the contract |
| \(\partial f/\partial v\) | how the contract value responds to a move in the variance |
| \(W^{v}_t\) | the second Brownian motion, driving the variance and nothing else |
| \(\xi\) | the size of the randomness in the variance, 0.30 here |
The result is an incomplete hedgeA position that offsets moves in the level but leaves the risk from a second source of randomness still in place.. Notice what has not happened. The offset has not become wrong, and it has not become useless: it still removes exactly what it always removed. The offset has become insufficient, and insufficient is a different complaint needing a different remedy. Removing the second push requires a second instrument whose value responds to the variance, and that is a structural change to what the argument requires rather than a tightening of it.
The household version is straightforward. A household running on one salary that insures the salary has covered one thing that can go wrong. If the cost of living then moves on its own, for reasons the salary does not drive, the insurance is still doing its job and the household is still exposed. One source of trouble was insured and there were two. Nothing about the cover failed; the count of what needed covering was wrong.
Volatility is made a process. What happens to a position that offsets moves in the level?
What does the choice of process decide?
Once volatility has an equation, somebody has to choose which equation. The choice of equation is often presented as a technical detail and it is not one. The equation decides what volatility is allowed to do, and different choices produce entirely different pictures of the same starting point.
The sharpest split is whether the process is pulled back toward a level or not. A pulled-back process treats every excursion as temporary: a period of high variance is a departure, and the pull works against it from the moment it starts. A process with no pull treats every excursion as permanent: wherever it wanders to becomes the new normal. Nothing is arguing for a return. Same starting value, same randomness, two completely different claims about what a quiet stretch or a violent stretch means.
Here it is with the numbers held fixed. Both processes below start at a variance of 0.040000, both have a randomness of size 0.30 with the square root in front, and both are driven by the same twelve values from the locked path, applied in the same order. The only difference is that the first is pulled toward 0.040000 at a speed of 2.0 a year and the second is pulled nowhere at all. Both are constructed from the locked driving values rather than sampled, so both reproduce exactly on every reading.
Both processes dip hard around month nine, to 0.006954 and 0.002247 respectively, and at that point they look like variations on one story. Follow them for three more months and the story splits. The pulled-back one climbs back to 0.039796, a volatility of 19.9490 per cent, essentially where it started. The one with no pull finishes at 0.013611, a volatility of 11.6665 per cent, and stays there because nothing is asking it to leave. The choice of process decided which of those two endings the model was ever capable of producing.
The split is why the choice of equation is a modelling decision rather than a detail. The equation settles what a quiet stretch means for the next stretch, how quickly a violent stretch is expected to subside, and how far ahead the process retains any memory of where it is now. Any use of the model will eventually ask those questions, and they were answered when the equation was chosen rather than when the numbers were.
What does choosing a mean reverting process for volatility decide?
The error that gets made, and what it costs
Treating an estimate of volatility as the volatility. The substitution is the most natural mistake in this area. The estimate arrives as a number with a decimal point, and every other number in the workbook is a reading. Nothing observable is the process, and every quantity standing in for it is an inference carrying error of its own.
Here is the cost in figures. The measurement says 0.040300. The model says 0.040000. A reader who takes the measurement as the quantity concludes that the model is understating variance by 0.000300 and reaches for a parameter to close the difference. The obvious parameter is the volatility itself, and moving it from 20.000000 per cent to 20.074860 per cent makes the model reproduce the measurement exactly. The at-the-money contract at Rs 100/- then prices at Rs 10.478677/- instead of Rs 10.450584/-, a move of Rs 0.028093/-, and the sensitivity to volatility of 0.375240 confirms that move independently.
Every one of those steps is arithmetically correct and the conclusion is wrong. The whole 0.000300 belonged to the partition and none of it belonged to the model. The parameter was moved to absorb a property of how the measurement was taken. The result is a model that agrees with a measurement and disagrees with the process it was built to describe, and the disagreement is now invisible because the residual was tuned away.
The cost compounds because the adjustment is silent. Nothing in the workbook records that the target contained an artefact, so the adjusted volatility gets passed on as though it were an improvement. The remedy is not more precision. The remedy is a line in the working that separates what the measurement can and cannot be held responsible for, and here that line can be written exactly: 0.000300 of the difference is the partition, at twelve steps, on this path.
A model disagrees with a measurement of volatility. Where can the disagreement live?
How does somebody reading a valuation actually use this?
Almost nobody reading a valuation is handed the model. A number with the word volatility beside it arrives in a note, and the useful questions are not about whether it looks sensible. The first question is which of three things that number is: a parameter somebody chose, a measurement somebody computed, or a value of a process at an instant. The third is impossible, so if the note claims it, the note is wrong about what it is holding.
An analyst reading it as a measurement should ask two things immediately: over what stretch, and at what spacing. Both answers change the number, and the second one changes it in the direction computed exactly above. A figure computed at monthly spacing carries more of the drift artefact than the same figure computed at daily spacing, and on this invented process the difference between the two is the difference between 0.040300 and 0.040014. The difference is small on a single valuation and not small when it is the input to a parameter that is then tuned.
A lender or an investment committee reading the same note should ask the question the tuning log never asked: what part of the difference between this model and this measurement is the measurement. The question has an answer, and the answer can be written down before anybody argues about the model. A working file that separates measurement error from model error before adjusting anything is doing something that costs one line and saves the whole downstream chain.
And anybody using the model for a hedge rather than a valuation should read the section above on what the second driver does. The difference between the two uses shows up hardest there. A valuation reads one number and can survive being a little wrong. A hedge asks what is left over after the offset is in place, and once volatility is a process the honest answer is that something is always left over. Knowing the size of what is left is more useful than pretending the count of drivers is still one.
Where this holds, and where the rules would come in
The mathematics here is universal. A variance must stay above nought, a latent quantity is never read off a record, and a sum of squared steps carries the drift of those steps, and none of the three is a matter of jurisdiction. Conduct duties do apply to what anyone does with a model in any particular place, and those are settled separately.
References
| Source | Document | Where |
|---|---|---|
| arXiv, Quantitative Finance | Preprints on latent volatility, realised measures and the effect of the partition | arxiv.org |
| Social Science Research Network | Working papers on stochastic volatility as a process and on incomplete hedging | ssrn.com |
| Black and Scholes, 1973 | The offsetting argument that removes randomness when there is one driver | Journal of Political Economy |
| Cox, Ingersoll and Ross, 1985 | The square root variance process and the condition that keeps it above nought | Econometrica |
| Heston, 1993 | The mean reverting variance process with a correlated second driver | The Review of Financial Studies |
The standard process, the locked path, the locked contracts and the invented settings for the variance process are invented.
Educational material. Not advice on any investment, tax, budget or market position.
