Quadratic Variation: The Property That Breaks Ordinary Calculus
Quadratic variation is what emerges from chopping an interval into pieces, squaring the change over each piece, adding those squares, and then making the pieces smaller without limit. On a smooth path that total collapses to zero. On a Brownian path it settles on the length of the interval itself, for almost every path, not merely on average. The equality with elapsed time is what ordinary calculus cannot absorb.
Everything that follows comes out of one scaling fact established earlier. A Brownian increment over a step of length delta-tThe length of one piece of the grid the interval has been divided into. is of the order of the square root of that step. Squaring it gives something of the order of the step itself. Dividing an interval into n pieces and adding n such squares: each one is of size roughly one over n, there are n of them, and the total refuses to shrink. A smooth increment is of the order of the step, and its square is of the order of the step squared, so changing the path to a smooth one collapses the same arithmetic. The contrast between the two kinds of path is the whole subject, and the rest is careful bookkeeping around it.
How is quadratic variation actually defined?
Quadratic variation is defined as a limit, and the limit is taken over grids rather than over time. Grids rather than time is the part worth slowing down for. Most readers expect a different shape. No time variable is pushed anywhere. The interval is held fixed, and its divisions are made finer and finer.
Start with an interval running from zero to some horizon. Choose a partitionA division of an interval into a finite number of consecutive pieces, marked by the times where one piece ends and the next begins. of it: a finite list of times, starting at zero, ending at the horizon, each one later than the last. Over each consecutive pair of times the path has a change. Square each of those changes and add up the squares. The total of the squares is one number for that one grid.
The same is done again with a finer grid, and again with a finer one still. The grids produce a sequence of numbers, one per grid. Quadratic variation is the number that sequence settles on as the largest piece of the grid is driven to zero, and it exists only as that limit and never as any one reading. The quantity being driven to zero has a name of its own: the meshThe length of the largest piece in a partition. Sending the mesh to zero is what forces every piece to shrink, not just the average piece. of the partition, meaning the length of its longest piece. A grid can have a thousand pieces and still contain one enormous one, so driving the mesh to zero is stronger than driving the number of pieces up.
| \([X]_T\) | the quadratic variation of the path \(X\) over the interval from zero to \(T\) |
| \(\Pi\) | a partition \(0=t_0<t_1<\cdots<t_n=T\), being one particular grid |
| \(\|\Pi\|\) | the mesh, being the length of the longest piece in that grid |
| \(X_{t_i}-X_{t_{i-1}}\) | the change in the path across the step that ends at \(t_i\) |
| \(n\) | how many pieces the grid has, a count that grows as the mesh shrinks |
Notice what step three does. Squaring throws away the sign of every movement, so an upward move and a downward move of the same size contribute identically. Discarding the sign is exactly why quadratic variation is blind to whether the path came back. None of the squares cancels anything, so a path can return precisely to where it began and still have a large quadratic variation. The worked instance below turns on that blindness.
Why does the same sum vanish on a smooth path?
An everyday example comes first. A gently sloping ramp is measured with a ruler, the height marked at each ruler mark and the change in height recorded between marks. The ramp has a slope, and the slope is indifferent to how the run is divided, so halving the ruler roughly halves every recorded change too. Squaring each change changes the accounting. Halving the ruler quarters each square and only doubles how many of them there are. Two times a quarter is a half, so the total halves every time the ruler is halved. Continued far enough, the total goes to nothing.
The ruler argument is the entire argument for a smooth pathA path with a well defined slope at every point, so that the change across a small piece is roughly the slope multiplied by the length of the piece., and three short steps make it exact. If the path has a slope that never exceeds some fixed size, then the change across a piece of length delta-t is at most that size multiplied by delta-t. Its square is at most the size squared multiplied by delta-t squared. Adding n of those over an interval divided evenly into n pieces gives the size squared multiplied by the horizon squared, divided by n. The division by n is the whole result: the total is bounded by something that goes to zero.
| \(f\) | a path with a slope everywhere, in place of a random one |
| \(C\) | the largest size the slope of \(f\) ever reaches on the interval |
| \(T/n\) | the length of each piece when the interval is cut into \(n\) equal ones |
| \(n\) | the number of pieces, sent upward without limit |
The collapse can be watched on a path already published for this subject. The standard processThe single invented traded quantity the whole of this subject runs on, starting at Rs 100/-, with a drift of 8 per cent and a volatility of 20 per cent a year. has a deterministic trend sitting inside it, growing the logarithm at 6 per cent a year in a perfectly straight line. The trend is smooth, so its quadratic variation must be zero. The sum computed on it at a few grids gives 0.003600 at one piece, 0.001800 at two, 0.000900 at four, 0.000300 at twelve and 0.000004 at a thousand and twenty four. Every time the number of pieces doubles, the total on the smooth trend halves, and that halving is the collapse the bound above predicts.
Why does the sum of squared changes vanish for a smooth path?
Why does the limit come out as the elapsed time itself?
Because a Brownian increment does not shrink the way a smooth increment does. Halve the step and a smooth change halves, but a Brownian change only falls by the square root of two, to about seventy one per cent of what it was. The Brownian change shrinks more slowly. Squaring a change that fell to seventy one per cent gives a square that fell to a half, and doubling the number of them returns the total to exactly where it began.
Read that sentence again. The whole subject is that arithmetic. On the smooth path, squaring beat counting and the total collapsed. Here squaring and counting cancel each other exactly, and the total holds still. Nothing else is going on.
| \(\Delta W_i\) | the change in the Brownian path across the step that ends at \(t_i\) |
| \(\Delta t\) | the length of each piece, equal to \(T/n\) on an even grid |
| \(n\Delta t=T\) | the pieces add back to the horizon, whatever n is |
| \(2T^{2}/n\) | the spread of the total, shrinking as the grid refines |
Two separate things are stated there and they do different jobs. The first says the sum is centred on the horizon at every grid, already true at one piece and at two and at three. The second says the sum stops moving around that centre only as the grid refines. The result is not that a coarse reading is near the elapsed time; it is that a coarse reading is centred on the elapsed time while being free to land almost anywhere, and that freedom is what refinement removes. Every honest reading of the locked path below is an instance of that second sentence.
One caution about that picture, and the whole subject is built around it. The flat green line is the scaling, meaning the size the total is centred on. The line is not a plot of any realised path, and no realised path traces it. A realised path at coarse grids does something considerably less tidy, and the worked instance below shows it.
A Brownian change scales with the square root of the step. What does its square scale with?
Is that a claim about the average, or about each path?
About each path, and the difference matters more here than almost anywhere else in this subject. The result is almost sureHolding on every path except a collection of paths carrying probability zero. Stronger than holding on average.: with probability one, the sequence of readings produced by any shrinking grid settles on the elapsed time. Not the average of those readings across many paths. Each path, individually, with the exceptions confined to a collection carrying no probability at all.
The weaker version would be useless, for the following reason. Suppose the result only held on average. The entitled claim would then be that generating a great many paths and reading the sum on each produces readings centred on the horizon. An average across many paths says nothing whatever about the path in hand. Such a claim leaves open that one analyst's path reads 0.4 and another's reads 1.6 and both are fine because the average worked out. No one could point at one realised path and say anything about its variation, and everything built downstream would have to be a statement about a population rather than about a path.
| \(\mathbb{P}\) | the physical measure, being the rule that assigns probability to sets of paths |
| \(\mathbb{P}(\cdot)=1\) | the collection of paths where the statement fails carries probability zero |
| \(T\) | the elapsed time, being the length of the interval |
| \(\|\Pi\|\to 0\) | the mesh being driven to zero, the quantity the limit is taken over |
Almost sure convergence is what licenses the whole of the next stretch of this subject to work with one realised path rather than with a population of them. Every quantity built by integrating along a path inherits that licence. Take it away and there is no such thing as a pathwise statement, and a great deal that reads as ordinary in this subject would have to be rewritten as a statement about expectations.
Does quadratic variation equal the elapsed time on average, or on almost every path?
What does the locked path actually read at each grid?
Now the worked instance, and it is where the arithmetic gets uncomfortable. The locked path is the twelve step Brownian path published for this reading order, built from twelve driving values multiplied by the square root of one twelfth. The twelve values are minus 0.5, 1.6, minus 1.3, minus 0.1, 0.1, 1.5, minus 1.3, minus 0.5, minus 1.4, 0.4, 0.9 and 0.6.
Two facts about that list were built in rather than found. The twelve sum to zero exactly, so the Brownian path finishes the year precisely where it started. And their squares sum to 12.0 exactly, so the twelve step reading of the quadratic variation is exactly one. The squares summing to twelve is what lets the figure reproduce to six decimal places instead of approximately, and a path drawn at random would show neither property.
| \(z_i\) | the twelve locked driving values, one per month |
| \(\sqrt{\Delta t}\) | 0.288675, the square root of one twelfth of a year |
| \(\sum z_i^{2}\) | 12.0 exactly, built into the list on purpose |
| \(\Delta t\) | one twelfth of a year, the length of each of the twelve pieces |
Now the honest part. The same path, unchanged, can be read at coarser grids. Grouping the twelve driving values into six pairs, or four groups of three, or three groups of four, or two halves, or one single block, each grouping gives a different set of changes and therefore a different sum of squares.
| Pieces the year is cut into | Sum of squared changes | Distance from 1.000000 |
|---|---|---|
| One | 0.000000 | 1.000000 below |
| Two | 0.281667 | 0.718333 below |
| Three | 0.031667 | 0.968333 below |
| Four | 1.345000 | 0.345000 above |
| Six | 1.018333 | 0.018333 above |
| Twelve | 1.000000 | exact |
That column does not climb, and nobody should draw it as though it did: it starts at zero, rises, falls to almost nothing, overshoots by a third, then comes back. The reading at three pieces is smaller than the reading at two. The reading at four is larger than the reading at twelve. Any picture that shows a tidy rise toward one has drawn something these numbers do not do.
The reading at one piece deserves its own paragraph. Cut the year into a single block and there is exactly one change to square: the change from the start of the year to the end. The locked path finishes at zero because its driving values sum to zero. So the change is zero, its square is zero, and the sum of squared changes is 0.000000. Meanwhile the standard process built on that path was at Rs 100/- in January, climbed to Rs 111.08/- in June, fell to Rs 93.74/- in September and finished at Rs 106.18/-. The coarsest possible reading of the locked path detects no variation at all in a path that travelled Rs 17.34/- between its high and its low.
The twelve locked driving values have squares that sum to exactly 12.0. Is that a coincidence?
The locked path is about to be read at one single piece covering the whole year. Before the control below is moved: what is the sum of squared changes?
Cut the same year into fewer pieces and watch the total lurch
The locked path never changes. Only the grid does. The grey line is the path, the dark chords are what the grid actually measures, the bars are the squared changes and the meter at the bottom is their total against the elapsed year of 1.000000.
The error that gets made, and what it costs
Reading the result as a promise that the sum of squared changes will sit near the elapsed time at any sensible grid. The result says no such thing. Quadratic variation is a statement about a limit, and a limit constrains the tail of a sequence rather than any particular term in it.
The locked path settles the point on its own. At three pieces it reads 0.031667 against an elapsed year of 1. No standard anyone would accept calls that close. At four pieces it reads 1.345000, overshooting by more than a third. Neither reading is the limit, so both are entirely consistent with the limit being exactly one. Somebody who expects the number to be near one at a coarse grid has quietly converted a statement about refinementDividing an interval into more and smaller pieces, so that the longest piece shrinks toward zero. into a statement about a fixed grid, and those are different statements.
The cost lands as a check that appears to fail. A quantity is sampled a handful of times over a year, a sum of squared changes is computed on it, the answer comes back at a third of what the model implies, and somebody concludes the model is wrong. The model was not wrong; the grid was too coarse to say anything either way. The same person, on a different stretch of the same path, would have got 1.345000 and concluded the model understated. Neither conclusion was available from the data, and nothing in the arithmetic flags that.
At three pieces the locked path gives 0.031667 against an elapsed year of 1. Is the result wrong?
One term in an expansion refuses to vanish as the step shrinks. Before reading on: how much of what follows in this subject does that one term account for?
Why does this one property break ordinary calculus?
Ordinary calculus is built on a habit of discarding. Expand a function around a point, keep the term proportional to the step, and throw away everything proportional to the step squared and beyond. The discarded terms vanish faster than the quantity being divided by. The habit of discarding is not a shortcut. Discarding is the definition of a derivative, and every rule taught on top of it depends on the discarded terms being genuinely negligible.
Here the second order term carries a squared changeThe movement of the path across one piece of the grid, multiplied by itself, so that upward and downward moves count the same., and squared changes on a Brownian path do not vanish. Squared changes add up to the elapsed timeThe length of the interval being examined, which is the number the sum of squared changes settles on for a Brownian path.. The elapsed time is the same order as the first order term, and that is exactly the order ordinary calculus keeps. The term ordinary calculus trains its readers to delete is the same size as the term it trains them to keep.
So the failure is not that ordinary calculus gives a slightly wrong answer here; it is that the step which makes ordinary calculus work has no justification on a path of this kind. Deleting the second order term deletes something of the same magnitude as everything retained. There is no version of the ordinary chain rule that survives this, and the replacement is not a patch on the old rule but a different rule with an extra term in it.
The fix is narrower than it looks, and worth naming. Nothing about probability changed. Nothing about the process changed. One bookkeeping step, familiar to the edge of invisibility, stopped being allowed. Everything built afterwards is the consequence of carrying one extra term through calculations that used to drop it, and that is why a property this small takes up so much room.
How does somebody checking a model rather than building one use this?
The useful move is a question, not a calculation, and it can be asked of somebody else's work without understanding how their figures were produced. Whenever a variation figure is computed from a sum of squared changes, the question is what grid it was computed on, and then what that grid can support.
The reason the question bites is the table above. The very same path yields 0.031667 and 1.345000 depending only on how it was cut, and a reviewer handed either number in isolation would draw a confident and wrong conclusion. A squared change total computed on a coarse grid is not a rough version of the limit, and treating it as one is the mistake this property most often invites.
- Ask what the grid was
Not the horizon, the grid. How many pieces was the interval cut into, and were they even? A figure quoted without its grid is not yet a figure that can be compared with anything.
The locked path returns six different numbers from six different grids over the same year.
- Ask what the longest piece was
The mesh governs the limit, not the count. A grid with many pieces but one long one has not refined in the sense the result requires.
A thousand pieces with one six month gap is still a coarse grid.
- Ask whether the reading is being used as a limit or as a reading
A reading is a fact about one grid. The limit is the quantity the result is about. Conflating them is where the wrong conclusion enters.
If the reading is being compared with a model number, it is being used as a limit.
- Ask what a coarser and a finer grid would have given
Two more computations, and the answer either settles or it does not. If it lurches, the grid is not fine enough to support any conclusion.
Three pieces gives 0.031667 and four pieces gives 1.345000, a lurch by any measure.
None of that requires access to anything. Four questions are asked of a number somebody else produced. The last one resolves most disagreements. A figure that changes by a factor of forty between adjacent grids has announced its own unreliability without anybody having to argue about the model.
A variation figure is computed from a squared change sum over a coarse grid. What is the risk?
Readers of this subject sometimes expect a universal convention, so one last note on universality. No jurisdiction anywhere sets the definition of quadratic variation, no authority publishes a value for it, and no convention alters what a limit over refining grids equals. Quadratic variation is a statement about paths rather than about anything traded, so the result is the same in every market and in none.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for pathwise variation results and their use in pricing | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Ito | The integral and the rule that carry his name, both of which rest on quadratic variation | named in the text only |
| Hull, Shreve and Wilmott | Standard texts on stochastic calculus for finance | named in the text only |
The standard process and the locked path are invented.
Educational material. Not advice on any investment, tax, budget or market position.
