Quadratic Covariation: How Two Random Processes Move Together
Quadratic covariation is the limit of the sum of the products of two processes' increments, step by step, as the grid is refined. The paired quantity is to two processes what quadratic variation is to one. Where the driving randomness is correlated it comes out as the correlation multiplied by both volatilities and by the elapsed time, and where the randomness is independent it is exactly zero.
Only one thing changes from the construction already established. Quadratic variationThe limit of the sum of the squared changes of one process across the pieces of a grid, as the longest piece is shrunk toward zero. takes the change a process makes across a piece of the grid and multiplies it by itself. Quadratic covariation takes the change one process makes across that piece and multiplies it by the change a second process made across the very same piece. Add those products, refine the grid, and take the limit. Everything that follows is a consequence of that single substitution, and the consequences are larger than the substitution looks.
Start with the everyday version, before any notation. A plank rests across a pivot, with one person standing at each end. Every time one end drops a foot, the other end rises a foot. Measuring both ends every minute gives, for each minute, a pair of movements: one number down and one number up. The pair multiplied together is a negative number, every single minute. A year of those minutes adds to a total firmly below zero, even though neither end ever moved a negative distance in any sense anybody could point at. No factor in that product is squared, so no product is forced to be positive, and that is the whole difference between this quantity and quadratic variation.
How is quadratic covariation actually defined?
Take a fixed interval running from zero to a horizon. Choose a partitionA division of an interval into a finite list of consecutive pieces, marked by the times where one piece ends and the next begins. of it, meaning a finite list of times starting at zero and ending at the horizon. Across each piece, the first process has a change and the second process has a change. Multiply those two changes together. Do that for every piece and add the results. The sum gives one number for that one grid, exactly as the single process construction did.
Then refine. Choose a finer grid, get a new number, choose a finer one still. Quadratic covariation is the number that sequence settles on as the longest piece of the grid is driven toward zero, and like its single process ancestor it lives only in that limit and never in any one reading. The quantity being driven to zero is still the length of the longest piece. Because a grid can be crowded almost everywhere and still hide one long stretch, driving that length down is stronger than merely counting more pieces.
| \([X,Y]_T\) | the quadratic covariation of the two processes \(X\) and \(Y\) over the interval from zero to \(T\) |
| \(\Pi\) | a partition \(0=t_0<t_1<\cdots<t_n=T\), one particular grid over the interval |
| \(\|\Pi\|\) | the length of the longest piece in that grid, the quantity driven to zero in the limit |
| \(X_{t_i}-X_{t_{i-1}}\) | the change the first process made across the grid interval ending at \(t_i\) |
| \(Y_{t_i}-Y_{t_{i-1}}\) | the change the second process made across that same piece, over the same clock |
| \(n\) | how many pieces the grid has, growing as the longest piece shrinks |
Set that beside the single process definition and look only at the second bracket. In quadratic variation the second bracket repeats the first. Here it names a different process. The times are the same, the limit is the same, the refinement is the same. One factor changed, and no part of the construction was reorganised around it.
What is the only change from the definition of quadratic variation?
What does it equal when the two sources of randomness are correlated?
Answering that with numbers rather than symbols needs two processes. The first is the standard process, the single invented traded quantity used throughout this reading order, starting at Rs 100/-, carrying a driftThe steady, non random part of how a process grows, quoted here as a rate per year. of 8 per cent a year and a volatility of 20 per cent a year. The second is an invented twin. The twin starts at the same Rs 100/-, carries the same 8 per cent drift and the same 20 per cent volatility, and differs from the standard process in exactly one respect: the randomness driving it is a different Brownian motion, correlated with the first at minus 0.7.
Because every other parameter is identical, nothing that follows can be blamed on one process being larger, faster or wilder than the other. The two are twins in every measurable respect except how their randomness lines up.
| \(S_t\) | the standard process, the invented traded quantity used throughout, starting at Rs 100/- |
| \(U_t\) | the second process, invented for this guide, starting at the same Rs 100/- |
| \(W^{1},W^{2}\) | two standard Brownian motions under the physical measure P, one driving each process |
| \(\mu\) | the drift, 8 per cent a year, the same for both by construction |
| \(\sigma_S,\sigma_U\) | the two volatilities, both 0.20 a year, again the same by construction |
| \(\rho\) | the correlation between the two Brownian motions, locked at minus 0.7 |
Now the paired quantity. Two facts do all the work. First, the drift terms contribute nothing at all, for the same reason they contributed nothing to quadratic variation: a steady term moves by an amount proportional to the length of that piece, so a product involving one of them is proportional to the length squared and collapses under refinement. Second, the product of the two Brownian incrementsThe change a process makes across one piece of the grid, being its value at the end of the piece less its value at the start. over a piece of length delta-t has an average of the correlationA number between minus one and one saying how closely two sources of randomness lean the same way. Zero means neither carries information about the other. multiplied by delta-t, and adding those across the grid gives the correlation multiplied by the elapsed time.
| \([W^{1},W^{2}]_T\) | the paired quantity of the two Brownian motions alone, equal to \(\rho T\), here minus 0.700000 |
| \(\sigma_S\sigma_U\) | 0.04, the two volatilities multiplied together, carried outside the sum because both are constant |
| \(\rho\) | minus 0.7, the locked correlation between the two driving Brownian motions |
| \(T\) | one year, the elapsed time over which the grid is refined |
| \(-0.028000\) | the result, negative because the correlation is |
Read the shape of that expression rather than the number. The correlation enters once and linearly, so it scales the whole quantity in a straight line. Each volatility enters once, so doubling either one doubles the result. Time enters once, so a two year horizon doubles it again. There is nothing curved, nothing squared and nothing to sign. Because the correlation multiplies the expression exactly once, the paired quantity travels a straight line from minus 0.040000 at a correlation of minus one, through exactly 0.000000 at a correlation of zero, to plus 0.040000 at a correlation of plus one.
Notice where the two dashed lines sit. The dashed lines mark plus and minus 0.040000. Each of these two processes has precisely that quadratic variation on its own over the year. So the straight line does not wander off toward some unrelated scale. The line is bounded above and below by the size of the variations it sits beside, touching them only when the correlation reaches one of its two extremes. The bound is not an accident of these numbers. Every pair obeys it, and no paired quantity can ever be large compared with the variations that produce it.
Each process has a quadratic variation of 0.040000. How large is the paired quantity at a correlation of minus 0.7?
What happens when the two are driven independently?
The paired quantity disappears. Not becomes small, not becomes a correction worth carrying. The quantity becomes exactly 0.000000, and every term in every equation that carries it vanishes along with it. Set the correlation to zero in the expression above and the whole product is zero however large the two volatilities are and however long the horizon runs.
| \(\rho=0\) | the two driving Brownian motions carry no information about each other |
| \(\sigma_X,\sigma_Y\) | the two volatilities, of any size at all and still no rescue for the term |
| \(T\) | the elapsed time, likewise no rescue |
| \(0.000000\) | the result, exactly zero rather than merely small |
The vanishing explains why so many readers reach a chain rule for two processes and find a term they have never seen before. Almost every first treatment of the subject works with a single process, and almost every second treatment works with two independentCarrying no information about each other, so that knowing one says nothing about the other. Independent sources of randomness have a correlation of zero. ones. Independent randomness is what a first construction of Brownian motion in several directions produces. In both settings the paired term is either absent or silently zero. A reader who has met only the independent case has not merely under-weighted this quantity. In that setting the term was never printed, so such a reader has never seen it written down at all.
Two processes are driven by independent randomness. What is their quadratic covariation?
Can a quantity built from products really go below zero?
Yes. Readers push back on this claim harder than on any other. The construction they arrived with could not produce a negative total. Quadratic variation adds squares. A square of a real number is never negative, so no amount of downward movement can drag that total below zero, and a path that falls all year produces exactly the same variation as one that rises all year. The intuition is correct, and it is also confined to the single process case.
Here the two factors in each product come from two different processes over the same piece of the grid. When one rises while the other falls, one factor is positive and the other is negative, and their product is negative. There is no squaring anywhere to erase that. Add a year of such products and the total sits below zero. The signWhether a quantity is above or below zero. Here it is inherited directly from the correlation rather than removed by squaring. of the paired quantity is inherited straight from the correlation, and nothing in the construction is available to remove it.
The correlation is about to go negative. Before the control moves: can the paired quantity go below zero?
Move the correlation and watch the paired bar cross the zero line
Both volatilities are held at 20 per cent a year and the horizon is held at one year, so the two grey bars for the processes' own quadratic variations never move. Only the correlation moves. The dark bar in the middle is the paired quantity, the marker on the right rides the straight line, and the meter at the bottom is the quadratic variation of the sum of the two logarithms. The sum is the first place the paired quantity actually shows up.
What do two locked paths actually add up to, month by month?
The expression above is a limit, and a limit is a promise about refinement rather than a number anybody can point at. The sum itself, worked at twelve steps on a pair of constructed paths, arrives at the same minus 0.028000 exactly rather than approximately.
The standard process runs on the locked path published for this reading order, built from twelve driving values multiplied by the square root of one twelfth. The twelve values are minus 0.5, 1.6, minus 1.3, minus 0.1, 0.1, 1.5, minus 1.3, minus 0.5, minus 1.4, 0.4, 0.9 and 0.6. The second process is driven by the same twelve numbers in a different order: 1.6, minus 1.3, 1.5, minus 1.4, minus 0.1, minus 1.3, 0.9, 0.6, 0.1, minus 0.5, 0.4 and minus 0.5. The rearrangement was chosen rather than sampled.
Reusing the same twelve numbers buys something valuable. Their squares sum to 12.0 whichever order they are in, so both processes have a quadratic variation of exactly 0.040000 over the year, with no rounding on either side. The twelve numbers sum to zero in either order, so both processes finish the year at exactly the same Rs 106.18/- they would have reached with no randomness at all. Every property either process has on its own is identical, so anything the pair does that a single one does not must come from the pairing and from nothing else.
| \(z_i\) | the twelve locked driving values for the standard process, one per month |
| \(u_i\) | the same twelve numbers rearranged, driving the second process, invented for this guide |
| \(\sqrt{\Delta t}\) | 0.288675, the square root of one twelfth of a year, appearing once in each bracket |
| \(\sum z_i u_i\) | minus 8.4 exactly, the total the rearrangement was chosen to produce |
| \(\sigma_S\sigma_U\) | 0.04, the two volatilities of 0.20 multiplied together |
Before the arithmetic, look at what the two paths did in rupees. Both start at Rs 100/-. The standard process travels down to Rs 93.74/- in September and up to Rs 111.08/- in June. The second process travels down to Rs 97.26/- in June and up to Rs 112.63/- in March. In ten of the twelve months, one went up while the other went down. And both finish the year at exactly Rs 106.18/-.
The matching endpoints deserve a moment. Given only the two starting values and the two closing values, a reader would report a pair that moved in perfect step. The paired quantity is not built from endpoints. The paired quantity is built from what happened in every piece of the grid, one piece at a time, and it is precisely what endpoint arithmetic throws away.
Now the twelve monthly contributions. Each is the month's driving value for the standard process multiplied by the month's driving value for the second process, then multiplied by both volatilities and by one twelfth of a year. Since both volatilities are 0.20 and the step is one twelfth, that scaling is a division by three hundred applied to each paired product.
| Month | Standard value | Second value | Product | Month's term | Running total |
|---|---|---|---|---|---|
| One | -0.5 | 1.6 | -0.80 | -0.002667 | -0.002667 |
| Two | 1.6 | -1.3 | -2.08 | -0.006933 | -0.009600 |
| Three | -1.3 | 1.5 | -1.95 | -0.006500 | -0.016100 |
| Four | -0.1 | -1.4 | 0.14 | 0.000467 | -0.015633 |
| Five | 0.1 | -0.1 | -0.01 | -0.000033 | -0.015667 |
| Six | 1.5 | -1.3 | -1.95 | -0.006500 | -0.022167 |
| Seven | -1.3 | 0.9 | -1.17 | -0.003900 | -0.026067 |
| Eight | -0.5 | 0.6 | -0.30 | -0.001000 | -0.027067 |
| Nine | -1.4 | 0.1 | -0.14 | -0.000467 | -0.027533 |
| Ten | 0.4 | -0.5 | -0.20 | -0.000667 | -0.028200 |
| Eleven | 0.9 | 0.4 | 0.36 | 0.001200 | -0.027000 |
| Twelve | 0.6 | -0.5 | -0.30 | -0.001000 | -0.028000 |
| Total | 0.0 | 0.0 | -8.40 | -0.028000 | -0.028000 |
Ten of the twelve monthly terms are below zero and two are above. The two positive ones are months four and eleven, the only months in which both driving values happened to carry the same sign, and together they add just 0.001667 back. The largest single contribution is month two at minus 0.006933, almost a quarter of the whole answer from one month. And month five contributes minus 0.000033, rounding to nothing at four decimal places. The contributions are wildly uneven, and no month is representative of the total. The quantity is therefore defined as a sum over a whole grid rather than as a typical monthly reading.
Can it be recovered from variations alone?
Recovery is possible, and it is worth seeing because it settles any lingering doubt that this is a genuinely new quantity rather than a repackaging. Build two new processes out of the pair: their sum and their difference. Each of those is a single process, so each has an ordinary quadratic variation with no pairing in sight. Compute both, subtract, divide by four, and the paired quantity falls out.
The manoeuvre has a name, polarisationRecovering a paired quantity from the variations of the sum and the difference of two objects. The same trick recovers a dot product from lengths in ordinary geometry., and it is the same move that recovers a dot product from lengths in ordinary geometry. Polarisation works here for one reason: the construction is bilinearBehaving like a product in each slot separately, so that scaling or adding inside one slot passes straight through to the answer., meaning it behaves like a product in each slot separately, so the square of a sum expands in the ordinary way and the cross terms are exactly the ones being sought.
| \([X+Y]_T\) | the quadratic variation of the sum of the two logarithms, 0.024000 on the locked pair |
| \([X-Y]_T\) | the quadratic variation of their difference, 0.136000 on the same pair |
| \(\tfrac{1}{4}\) | the constant that falls out of expanding both squares, the cross term appearing twice in each |
| \([X,Y]_T\) | the paired quantity being recovered, minus 0.028000, matching the direct computation |
Look at what those two variations are. Each process alone carries 0.040000, so a reader expecting the sum to carry 0.080000 has assumed the paired quantity away. The sum actually carries 0.024000, less than a third of that, and less even than either process on its own. The difference carries 0.136000, well above the total of the two. Adding two processes that lean against each other produces something quieter than either of them, and that quieting is the paired quantity doing visible work.
Each process carries a quadratic variation of 0.040000. What does their sum carry?
The error that gets made, and what it costs
Assuming the paired quantity cannot be negative because the quantity it grew out of never was. The reasoning feels airtight and it is applied without being noticed: quadratic variation adds squares, squares are never negative, so a variation is never negative, so this variation is never negative either. The last step is the one that fails. Quadratic covariation adds products of two different changes, and a product of two numbers with opposite signs is negative.
The mistake in working is not somebody writing that a negative number is impossible. The mistake is somebody computing 0.7 multiplied by 0.20 multiplied by 0.20, reading 0.028, and writing that down. The size is right. The correlation was even used correctly. Only the sign was quietly discarded on the way from the correlation into the answer, most often because the value was written down as a magnitude first and given its sign afterward, and the second step got missed.
The cost is not the size of the term, it is twice the size of the term. The quantity lands on the wrong side of zero rather than merely missing. On the locked figures the true value is minus 0.028000 and the recorded value is plus 0.028000, so the gap is 0.056000. Set that against the 0.040000 quadratic variation the term is being combined with. The error is larger than the quantity it was supposed to be correcting, and every other line of the arithmetic is perfectly correct, so nothing anywhere flags it.
The paired quantity is minus 0.028000 and somebody records it as 0.028000. How far out is the result?
Why does a term with two processes turn up in a chain rule at all?
Because expanding a function of two variables to second order produces a cross term, and there is nowhere else for that cross term to go. Expand a function of one variable and the second order part carries one squared change. Expand a function of two variables and the second order part carries three pieces: one squared change from each variable, and one term that multiplies one change by the other. In ordinary calculus all three are discarded together. Here the first two survive as the two quadratic variations already established, and the third survives as this.
| \(f\) | any twice differentiable function of two process values, whatever it happens to be |
| \(\partial^{2}f/\partial x\,\partial y\) | the mixed second derivative, being how the slope in one variable changes as the other moves |
| \(dX_t\,dY_t\) | one increment multiplied by the other, exactly the product being accumulated here |
| \(j,k\) | labels running over the two variables, so the double sum has four entries and two of them are the cross term |
Note the half at the front and the two entries. The double sum contains the cross term twice, once with the labels in each order, and the mixed second derivative is the same both times, so the half cancels one of them and the surviving term carries no fraction. The cancellation is a small point of bookkeeping, and it is exactly the point people get wrong when they carry a stray factor of a half through. Ito's lemma is the rule that uses this term, and it is covered separately.
Why does a term involving two processes turn up in a chain rule at all?
How is it computed from what is already in hand?
In one line, from three numbers already in hand. The correlation, the two volatilities, and the elapsed time. Multiply them together. Nothing new is estimated, no fresh observation is required, and no extra parameter enters the working that was not already written down when the two processes were specified.
The absence of any new estimate is worth stating plainly. The quantity has a forbidding name, and readers assume a forbidding computation sits behind it. There is no separate paired parameter to fit. Two processes written down with volatilities, together with a statement of how their randomness is related, already fully determine the paired quantity, whether or not it is ever computed. The paired quantity is not an additional input to a model, it is an output of inputs already chosen, and a model that quietly sets it to zero has not avoided choosing it but has chosen zero.
Both volatilities are 20 per cent and the correlation is minus 0.7, over one year. Compute it.
How does somebody reading a model rather than building one use this?
The useful move is a short sequence of questions asked of somebody else's working, and none of them requires redoing that arithmetic or having access to anything it used. Two of the four catch a lost sign or a missing term, and the other two catch the assumption that hides them.
The exercise is reading a bill rather than adding one up. Nothing is recomputed line by line. The reader is looking for the line that should be there and is not, and the line whose sign nobody checked. The paired quantity is unusual in that its most common failure is total absence rather than an incorrect value, so the first question is always whether it appears at all.
- Ask whether the two are being treated as independent, and whether anybody said so
An assumption of independence sets this quantity to zero and removes every term carrying it. That is a modelling choice, not a simplification, and it is frequently made by default rather than by decision.
On these locked figures, assuming independence discards a term of minus 0.028000 without a line of working appearing anywhere.
- Ask what correlation the paired quantity was computed at, and where that came from
The whole quantity scales in a straight line with that one number, so a correlation that arrived by habit rather than by choice has set the answer by habit too.
Moving the correlation from minus 0.7 to minus 0.5 moves the paired quantity from minus 0.028000 to minus 0.020000.
- Ask which side of the equation the term landed on, and check its sign against the correlation
A negative correlation must produce a negative paired quantity. If the working shows a positive one, the sign was lost between the two, whatever else is right.
A sign flip here costs 0.056000, which is more than the 0.040000 variation the term sits beside.
- Ask what the answer would be at a correlation of zero, and how much of it that removes
This is the fastest way to see how much of the result is riding on the pairing. If setting it to zero barely moves the answer, the pairing was never load bearing. If it moves it a great deal, the correlation deserved an argument.
Setting the correlation to zero here moves the variation of the sum from 0.024000 to 0.080000, more than tripling it.
The fourth question settles most disagreements. Asking it converts an argument about whether a correlation is right into a measurement of how much the answer depends on it. On this pair the dependence is severe. On another pair with a correlation near zero it would be negligible, and the same four questions would have said so in about a minute.
One closing note on universality. Readers sometimes look for a governing rule. No authority anywhere sets the definition of quadratic covariation, no regulator publishes a value for it, and no market convention changes what a limit over refining grids equals. The definition is a statement about paths rather than about anything traded, so it is the same in every jurisdiction and in none. No supervisor governs this arithmetic. Its parameters are chosen rather than observed, and the only thing they are answerable to is internal consistency.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for pathwise covariation results and their use in multi-process settings | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Ito | The integral and the rule that carry his name, in which the cross term set out above is the term the rule carries | named in the text only |
| Hull, Shreve and Wilmott | Standard texts setting out quadratic covariation and the multi-process chain rule | named in the text only |
The standard process, the second process and both twelve step paths are invented.
Educational material. Not advice on any investment, tax, budget or market position.
