Ito's Lemma: The Chain Rule for Random Processes
Ito's lemma is the rule for how a function of a randomly moving quantity changes. Line up its terms against the ordinary chain rule and every one matches except a single extra: one half of the curvature multiplied by the variance rate. The extra term exists because squared movements on a jagged path refuse to disappear, and it changes the answer.
Nothing already established about differentiating a function of a changing quantity is being taken away. Every term the ordinary rule would write down survives, in the same place, meaning the same thing. One line is added underneath them, and that line is invisible on a smooth path and unavoidable on a rough one. Where that line comes from, what it means, when it vanishes, and what it does to a number that can be checked are the four questions that follow.
Why does the ordinary chain rule fail on a random path?
The ordinary chain rule is not a rule about slopes. The ordinary chain rule is a rule about what may be thrown away. For any function whose input moves by a small amount, the change in the output can be written as a series: a term proportional to the movement, then a term proportional to the movement squared, then smaller ones still. The chain rule keeps the first and discards the rest. The discarding is not laziness. Throwing away the tail is the whole content of the rule, and it is legitimate exactly when the discarded terms shrink faster than the term retained.
The fall of a ramp measured with a ruler makes this concrete. Halving the ruler halves each recorded drop. Squaring each drop takes it to a quarter. Only twice as many of them have to be added. Squaring beats counting, so the squared column collapses to nothing and may be ignored. The expansionWriting the change in a function as a running series: a piece proportional to the movement, then a piece proportional to its square, then smaller pieces. keeps the first column and drops the second because the second genuinely goes away.
Now the ramp gives way to a path that shakes. A Brownian movement scales with the square root of the step rather than with the step. Halving the step therefore does not halve the movement. The movement falls only to about seventy one per cent of what it was. Squaring that gives a half, and doubling the count restores the total to exactly where it began. Squaring and counting now cancel each other, so the squared column holds steady instead of collapsing, and the term the ordinary chain rule was built to discard is the same size as the term it was built to keep.
The loss of that cancellation is the whole failure, and it is worth naming what kind of failure it is. The ordinary chain rule does not give a slightly wrong answer here through some approximation getting loose. The ordinary chain rule gives an answer whose justification has been withdrawn. The permission slip that allowed the second column to be crossed out was written for smooth paths and does not apply.
| \(f(t,S)\) | the function being tracked, twice differentiable in \(S\) and once in \(t\) |
| \(S\) | the standard process, the single invented quantity the whole of this reading order runs on |
| \(\Delta S\) | the movement of that process across one small step |
| \(\Delta t\) | the length of the step |
| \(\partial^{2}f/\partial S^{2}\) | the second derivative, being the rate at which the slope itself changes |
What does the lemma actually say, term by term?
Write the standard processThe single invented quantity the whole of this reading order runs on, starting at Rs 100/-, growing at 8 per cent a year and shaking at a volatility of 20 per cent a year. down first. The lemma has to be applied to something. The standard process is the single invented quantity this reading order runs on: it starts at Rs 100/-, its driftThe steady part of a movement, the piece that ticks with the clock rather than with the randomness. Here it is 8 per cent a year. is 8 per cent a year, and it shakes at a volatility of 20 per cent a year. Its movement over any instant has two pieces, one that ticks with the clock and one that ticks with the randomness.
| \(S_t\) | the standard process at time \(t\), starting at Rs 100/-, invented |
| \(\mu\) | the drift, 0.08 a year under the physical measure P |
| \(\sigma\) | the volatility, 0.20 a year, so the variance rate \(\sigma^{2}\) is 0.04 |
| \(W_t\) | standard Brownian motion under the physical measure P |
| \(dt\) | an instant of the clock |
Now take any function of that process and of the clock. The lemma says its movement has three sources rather than two. The first is the passage of time with the process held still. The second is the movement of the process itself, weighted by how sharply the function responds to it. The third is the one that has no counterpart anywhere in ordinary calculus: the bending of the function, meeting the variance of the process.
| \(\partial f/\partial t\) | how the function responds to the clock alone |
| \(\partial f/\partial S\) | how sharply the function responds to the process |
| \(\partial^{2}f/\partial S^{2}\) | how that sharpness itself changes, being the bending |
| \(\sigma^{2}S_t^{2}\) | the variance rate of the process at its current level |
The first two terms are exactly what the ordinary chain rule would have written down without any of this subject. The lemma does not correct the ordinary chain rule, replace it or contradict it; it appends one line to the bottom of it and leaves everything above untouched. The appended line is why the comparison below has two rows in common and one row that stands alone, and why counting the additions is a fair way to measure the whole difference.
How many rows does Ito's lemma add to the ordinary chain rule?
Where does the extra term come from, and why can it not be thrown away?
The extra term comes from one substitution, and the bookkeeping for that substitution has already been settled earlier in this reading order. Look back at the expansion. The third piece is the bending multiplied by the square of the movement of the process. So the only question is what the square of that movement is worth.
Squaring the movement of the standard process gives three products: the clock piece squared, twice the clock piece times the random piece, and the random piece squared. The first two are smaller than a step and go away. The third does not go away. A squared Brownian movement counts as a step. One line, and the extra term is in plain view.
| \((dt)^{2}\) | smaller than a step, so it counts as nothing |
| \(dt\,dW_t\) | also smaller than a step, so it counts as nothing |
| \((dW_t)^{2}\) | counts as \(dt\), which is the fact established earlier |
| \(\sigma^{2}S_t^{2}\) | what is left standing, being the variance rate at the current level |
Substituted into the third piece of the expansion, that finishes the lemma. There is no further argument, no limiting procedure to watch and no cleverness. The work was all done earlier, in establishing that a squared Brownian movement is worth a step of the clock, and the lemma is simply spending what that established.
As for why it cannot be thrown away: it is the same size as the terms being kept. A quantity fit for deletion has to be small compared with what remains, and this one is not small compared with anything. Deleting the bending term is not an approximation with a boundable error. Removing a term the same size as the terms retained is a different act entirely. The result obtained after deleting it is not close to the truth. The result is the answer to a different question.
What does the extra term mean, rather than say?
The extra term is what bending costs, or pays, in the presence of shaking. The everyday version runs as follows. From a mark on a hill, one step uphill, then back to the mark and one equally sized step downhill. On a flat slope what was gained one way was lost the other, so the two heights average back to the height of the mark exactly. On a hill that bends, they do not. If the ground curls upward on both sides, both steps land higher than the straight line predicted, and the average of the two sits above the mark. If it curls downward, the average sits below it.
The gap between the average and the mark is not a rounding error. The gap is a real quantity with a size and a sign, and its size depends on two things only: how hard the ground bends, and how far the two steps went. Since the two steps are supplied by randomness, and the typical distance travelled by randomness is measured by variance, the gap comes out as the bending multiplied by the variance. That product is the extra term, and the one half in front of it is the same one half that sits in front of every squared term in an expansion.
A function that bends upwardConvex: the graph curls upward on both sides, so a straight line joining two of its points sits above the graph between them. therefore gains from shaking, and one that bends downwardConcave: the graph curls downward on both sides, so a straight line joining two of its points sits below the graph between them. loses from it. The direction of the effect is what most readers actually take away, and it survives long after the notation has faded. Randomness is not neutral for a bending function: it pays the one that curls upward and charges the one that curls downward, and the size of the payment is set by the variance rate.
Why does bending interact with randomness at all?
Which functions escape the correction entirely?
The ones that do not bend. If a function is a straight line in the process, its second derivativeThe rate at which the slope itself changes. It is nil for a straight line, below zero where a graph curls downward and above zero where it curls upward. is zero, the extra term is zero multiplied by something, and the lemma collapses back into the ordinary chain rule exactly. Not approximately, not for small volatility, but exactly and at every volatility. The second derivative is the fastest check available in this whole subject: where it is nil, the correction is nil and the ordinary rule is already exact.
Three functions of the standard process, worked all the way through, do more here than another abstraction. Each one gets its bending computed at the starting value of Rs 100/-, then its growth rate under the ordinary chain rule, then its growth rate under the lemma. The variance rateThe square of the volatility. Here the volatility is 0.20 a year, so the variance rate is 0.04 a year exactly. of 0.04 does all the work in the third column.
| Function of the process | Bending | Ordinary rule | Under the lemma | Effect |
|---|---|---|---|---|
| Three times the process | 0 | 0.080000 | 0.080000 | none |
| The logarithm of the process | below zero | 0.080000 | 0.060000 | subtracts 0.020000 |
| The process squared | above zero | 0.160000 | 0.200000 | adds 0.040000 |
Read the last row, because it is the one that stops the correction being remembered as a subtraction. The square of the standard process bends upward, so the lemma pushes its growth rate up rather than down, from 0.160000 to 0.200000. The amount added is 0.040000, the whole variance rate rather than half of it. The addition of a whole variance rate is not a different rule. The addition is the same one half of the bending multiplied by the variance rate, evaluated on a function whose bending happens to be two.
| \(\partial^{2}f/\partial S^{2}=0\) | the function is a straight line in the process |
| \(df\) | what is left, which is the ordinary chain rule unchanged |
| \(\sigma\) | absent from the result, at every value it could take |
A function is a straight line in the process. What does Ito's lemma reduce to?
The square of the standard process grows at 0.160000 a year under the ordinary chain rule. What does the lemma make it?
The standard process drifts at 8 per cent a year. Before reading on: at what rate does its logarithm drift?
What does the lemma look like applied once, to the logarithm?
Now the worked instance, on the function the whole subject reaches for first. Take the logarithm of the standard process. Its first derivative is one over the level, so the slope shrinks as the level rises. Its second derivative is minus one over the level squared, and minus one over a square is below zero everywhere. The logarithm therefore curls downward at every point, so the correction must be a subtraction. The sign is settled before a single number is computed.
Now the size. The bending at a level of Rs 100/- is minus 0.000100. The variance rate at that level is 0.04 multiplied by ten thousand, or four hundred. One half of minus 0.000100 multiplied by four hundred is minus 0.020000. The level squared in the variance rate cancels the level squared under the bending exactly, so every level gives the same answer. The correction on the logarithm is a constant rather than something that drifts around as the process moves.
| \(1/S_t\) | the first derivative of the logarithm, the slope at the current level |
| \(-1/S_t^{2}\) | the second derivative, below zero everywhere, so the function curls downward |
| \(\mu-\tfrac{1}{2}\sigma^{2}\) | 0.08 less 0.02, being 0.06 a year exactly |
| \(\sigma\,dW_t\) | the random piece, which the correction leaves completely alone |
Two true statements now sit side by side and they only look contradictory. The standard process is expected to grow at 8 per cent a year. Its logarithm grows at 6 per cent a year. Both are correct, and the distance between them is precisely the term this guide is about. A reader who insists that only one of them can be right has not yet accepted that the function and its logarithm are two different quantities with two different movements.
Check it on the locked pathThe one twelve step path published for this reading order and drawn wherever a path is needed, so that a worked figure reproduces exactly rather than changing on each reload.. There the correction stops being notation and becomes rupees. The locked path is the twelve step path published for this reading order, built so that its twelve driving values sum to zero exactly. The construction matters here: with the driving values summing to zero, the random piece contributes nothing at all over the year, and the logarithm moves by its drift term alone. Everything the two rules disagree about is therefore laid bare, with nothing random left to hide behind.
Under the lemma the logarithm rises by 0.060000 over the year, so the standard process finishes at Rs 106.18/-. Under the ordinary rule it would rise by 0.080000 and finish at Rs 108.33/-. The gap is Rs 2.15/- on a starting value of Rs 100/-, and it comes from one term in one line, on a path where the randomness contributed exactly nothing.
| Month of the locked path | Under the lemma | Under the ordinary rule | Gap |
|---|---|---|---|
| Three | Rs 100.35/- | Rs 100.85/- | Rs 0.50/- |
| Six, the high point | Rs 111.08/- | Rs 112.19/- | Rs 1.12/- |
| Nine, the low point | Rs 93.74/- | Rs 95.15/- | Rs 1.42/- |
| Twelve, the finish | Rs 106.18/- | Rs 108.33/- | Rs 2.15/- |
What does the correction do as the volatility rises?
The behaviour of the correction as the volatility rises is where a technicality turns into the whole answer, and it turns on one word: squared. The correction is one half of the variance rate, and the variance rate is the volatility multiplied by itself. So the correction does not grow in step with the volatility. The correction grows with the square of the volatility, so it is almost nothing when the shaking is mild and enormous when the shaking is not.
Put numbers on that. At a volatility of 10 per cent the correction is half a percentage point a year, small enough that most readers would never notice it. At 20 per cent it is 2 points. At 30 per cent it is 4.5 points. At 40 per cent it is 8 points, and since the drift of the standard process is 8 per cent a year, the correction has just eaten the entire drift: the corrected growth of the logarithm is exactly zero. Push past that and it goes below zero, so the logarithm drifts downward while the process itself is still expected to grow at 8 per cent a year.
| \(\mu\) | the drift of the standard process, 0.08 a year |
| \(\sqrt{2\mu}\) | the volatility at which the correction equals the drift exactly |
| \(0.40\) | 40 per cent a year, and the equality here is exact rather than rounded |
The volatility is about to reach 40 per cent, with the drift held at 8 per cent. Before the control is moved: what happens to the corrected growth rate of the logarithm?
Turn the volatility up and watch the correction eat the drift
The drift stays at 8 per cent a year and the function stays the logarithm throughout. The only thing that moves is the volatility. The left bars are the two growth rates, the curve on the right is the corrected rate at every volatility, and both use the same vertical scale so a bar height and a curve height mean the same thing.
Why is one extra term worth a whole subject?
Because everything downstream is that term wearing different clothes. Once a bending function and a shaking process are known to interact, and the interaction has a size that can be written down, several other results become available and each of them is covered separately. The point worth carrying away is not the list. The point is that the list exists because of one line.
How Ito's Lemma Supports Derivative Pricing
The lemma is what lets anybody write down how a function of a process moves, and the pricing of a derivative contract begins by treating the contract as exactly such a function. The lemma is the connection between the two. The argument that turns a movement into a value, and the equation carrying the names of Black, Scholes and Merton that comes out of it, are covered separately and much later in this subject.
What does the lemma not do?
The lemma does not say what anything is worth. Both sides of the statement are movements: a small change in the function, expressed through small changes in time and in the process. A movement is not a level. Nothing in the lemma pins down where the function stands, only how it travels, and turning the second into the first takes a separate argument, sitting outside this reading order.
The lemma does not say where the process will finish either, and it never claimed to. Every term in the lemma is a rule about an instant. Getting from instants to a destination means integrating them along a path. Integration is a separate operation with its own machinery, covered earlier in this reading order.
The lemma does not remove the randomness. Look at the applied form on the logarithm: the correction changed the drift term and left the random piece exactly as it found it, still carrying the same volatility. The lemma is a translator rather than a solver: it converts the movement of one quantity into the movement of another, and whatever was uncertain before the translation is just as uncertain after it.
And it does not apply to a function that is not smooth enough. The statement asks for a function that can be differentiated twice in the process and once in the clock. A function with a corner in it fails that test at the corner, and the extension covering such functions is a different result with a different name, covered separately.
Does Ito's lemma say what a function of the process is worth?
The error that gets made, and what it costs
Somebody reaches for the ordinary chain rule on a function of a random process and carries the drift straight through the logarithm. The figure that comes out is too high by one half of the variance rate: 2 percentage points a year at a volatility of 20 per cent, and 8 percentage points at 40 per cent, where it is the entire drift.
Nothing inside the calculation flags it, and that is what makes it durable. Every line of the arithmetic is valid. Every intermediate figure reconciles. The mistake was not made in any step, so every check that compares one step against the next passes. The mistake was made before the arithmetic began, in choosing which of two chain rules to use. There is no cell to inspect and no sum that fails to tie.
The cost compounds. On the locked path the overstatement is Rs 2.15/- after one year, Rs 7.40/- after three and Rs 40.34/- after ten, on a starting value of Rs 100/-. The overstatement also runs in the flattering direction, and the flattering direction gets caught last. A figure that looks better than expected invites less inspection than one that looks worse.
Somebody carries the drift straight through the logarithm. Is the resulting figure too high or too low?
How does somebody checking a growth figure use this?
Nothing has to be built to put this to use. The useful move is a short interrogation of a number somebody else produced, and it works even when their working is invisible. The everyday version is the difference between what a scale reads and what the object weighs: before the weight is argued about, the scale is checked for zeroing. Here the zeroing question is whether the correction was applied, and the answer is usually visible from the figures alone.
- Ask which quantity the rate describes
A growth rate quoted for a process and a growth rate quoted for its logarithm are two different numbers, and they differ by one half of the variance rate. A figure quoted without saying which of the two it is has not yet said anything.
On the standard process the pair is 0.080000 and 0.060000, and either could be correct depending on what is being described.
- Ask what the volatility was
The size of the gap is set entirely by the volatility, so a figure quoted without one cannot be checked at all. At 10 per cent the gap is half a point and nobody would notice. At 40 per cent it is the whole drift.
Doubling the volatility multiplies the gap by four, not by two.
- Do the subtraction yourself and see whether it lands
Subtracting one half of the squared volatility from the quoted rate shows whether the answer is a rounder number than the rate it started from. Corrections are often visible as the difference between a suspiciously round figure and a working one.
A figure of 0.080000 alongside a volatility of 20 per cent is a reason to look for 0.060000 somewhere in the working.
- Ask whether the function bends, and which way
No bending means no correction and the ordinary rule is exact. Bending downward means the rate should have come down. Bending upward means the rate should have gone up, and that is the case people forget.
The square of the process gains 0.040000 while the logarithm of the same process loses 0.020000.
Four questions, none of which needs data, software or access to anybody's working. Half of one squared volatility is arithmetic anybody can do while the conversation is still going, so the third question resolves most disagreements on the spot. The arithmetic is fine, so checking the arithmetic almost never catches this mistake. The mistake is caught by asking which of the two quantities the number was ever meant to describe.
One closing note on universality, for readers of this subject who go looking for a published rule. No authority anywhere sets the form of this result, no jurisdiction publishes a value for the correction, and no market convention changes what one half of a variance rate equals. The mathematics is the same everywhere and in nowhere in particular.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for stochastic calculus and its use in derivative pricing theory | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Ito | The lemma and the integral that carry his name, named wherever either appears, with no text reproduced | named in the text only |
| Black, Scholes and Merton, 1973 | Named only where the equation that carries their names is pointed at, which is covered separately | named in the text only |
| Hull, Shreve and Wilmott | Standard texts, consulted for notation and ordering only, never for copied text | named in the text only |
The standard process, its four parameters and the locked path are invented.
Educational material. Not advice on any investment, tax, budget or market position.
