How to Apply Ito's Lemma Step by Step, in Seven Steps
Applying Ito's lemma is seven steps in a fixed order: write the process in drift and diffusion form, write the function and its arguments, take three derivatives, expand to second order, apply the multiplication table, collect the survivors into a drift and a diffusion, then run two checks. Each step consumes what the one before produced.
The result is already in hand. Ito's lemma itself, and the account of where its extra term comes from, are set out separately. The procedure is the order things are written in, what each step hands to the next, and the two checks that decide whether the answer is fit to use. Each of the seven steps says what to write. None of them says why the writing works.
A procedure and an explanation are different objects with different failure points, and the difference is worth stating plainly. An explanation fails when a reader does not believe it. A procedure fails when a step is skipped, done out of order, or fed something the step before it never produced. The seven steps here are written so that a person who has read the account once can run the calculation cold, at speed, on a function never seen before, and know at the end whether it came out right.
The everyday version is a tailor with a tape measure. Measure, mark, cut, sew: nobody argues about that order, and nobody needs the theory of textiles to follow it. The order is forced by what each step produces. Cloth cannot be marked to a measurement that was never taken, and it cannot be cut to a mark that is not there. The seven steps have exactly that character. There is no judgement anywhere in them once the process and the function are written down.
What makes the order of the seven steps fixed?
Each step takes the output of the one before it as its only input. Step three cannot begin until step two has named the function. Step four cannot begin until step three has produced three derivatives. Step five has nothing to act on until step four has produced products of the increment symbols, and step six has nothing to gather until step five has said which of those products survived. The order is forced by the material rather than chosen by convention, so skipping a step does not make the procedure shorter, it makes the next step empty.
The forced order is why the procedure is worth having as a written sequence rather than as an argument reconstructed each time. The reconstruction gets done under time pressure, on a function more awkward than the one in the account. Reconstructing an argument is slow, and it is where errors enter. A list that can be run straight down is faster and it is auditable: somebody else can look at the working, find the step it had reached, and see whether that step was fed what it needed.
Here is the whole procedure in one place. Each step below gets a section saying what is written down at it. The running example is the natural logarithm of the invented standard processThe single invented traded quantity this whole reading order runs on, written S with a time subscript, starting at Rs 100/- and watched for one year.. The standard process starts at Rs 100/-, drifts at 8 per cent a year and carries a volatility of 20 per cent a year.
- Write the process in drift and diffusion formCopy out the coefficient standing in front of dt and the coefficient standing in front of dW, each on its own line, before anything else is written down.Checking: are both coefficients written as expressions in the level, and does neither of them still contain the other symbol?
- Write the function and name its two slotsWrite the function with both of its slots shown, time and level, then record which of the two it actually uses.Checking: is the time slot present even where it is unused? Step three needs something to differentiate.
- Take the three derivatives, and only threeWrite the derivative by time, the first derivative by the level, and the second derivative by the level. In that order, and nothing beyond them.Checking: are there exactly three expressions written down, and is the third one a second derivative rather than a repeat of the first?
- Expand to second order, then stopWrite the time piece, the first order piece and the squared piece. The squared piece carries the squared increment of the process and must appear in writing.Checking: is there a term containing a squared increment anywhere in the working? If not, the expansion stopped one order too early.
- Apply the multiplication table to every productMultiply out the squared increment, then look up each product in the table and write down what it becomes. Look up, do not derive.Checking: has every product been replaced by either zero or dt, with none left standing in its original form?
- Collect the survivors into two bracketsPut everything carrying dt inside the first bracket and everything carrying dW inside the second. Cancel whatever cancels within each bracket.Checking: are there exactly two brackets, and does the second one contain no dt at all?
- Run the two checks before using the resultSet the volatility to zero and read the drift. Then replace the function with a straight line and read both coefficients.Checking: does the first return the ordinary chain rule, and does the second return the process itself scaled by the slope?
What gets written down before touching the lemma at all?
Two things, and both are written before a single derivative is taken. The first is the process in drift and diffusion formWriting a process as a term in dt plus a term in dW.. The second is the function, with a note saying which of its arguments it actually depends on. Neither of these is a calculation, and that is exactly why they get skipped. A procedure with no judgement anywhere in it still manages to go wrong at the very first step.
| \(S_t\) | the standard process at time \(t\), invented, starting at Rs 100/- |
| \(dS_t\) | the increment of the process over the next instant |
| \(\mu\) | the drift parameter, here 0.08 a year, read off the specification of the process |
| \(\sigma\) | the volatility parameter, here 0.20 a year, read off the same place |
| \(dW_t\) | the increment of Brownian motion under the physical measure P |
| \(dt\) | the increment of clock time |
For the standard process the drift coefficient is 0.08 times the level and the diffusion coefficient is 0.20 times the level. Write those two down as expressions, not as numbers. Steps four and six will multiply them by derivatives, and a number written where an expression belongs is the reason the level fails to cancel later on a function where it should.
| \(Y_t\) | the new process, being the function evaluated along the old one |
| \(f\) | the function, written with both of its slots shown |
| \(t\) | the first slot, clock time, present here and unused |
| \(x\) | the second slot, standing for the level of the process |
The logarithm uses the level slot alone. The time slot is written anyway. Leaving it off saves one symbol now and costs a whole term later on any function that does depend on time, and the whole point of a written procedure is that it runs the same way on every function fed to it.
Before the lemma is touched at all, what two things must already be written down?
Which derivatives are needed, and in what order?
Three, in a fixed order, and never a fourth. The derivative by time, the first derivative by the level, and the second derivative by the level. Naming all three up front, before any of them is used, is what stops the expansion at step four sprawling into terms nobody asked for. Each of the three has a destination: the first two are consumed by step four, and the third is consumed by step four and again by step six.
| \(\partial f/\partial t\) | the derivative by the time slot, zero here because the logarithm does not use it |
| \(\partial f/\partial x\) | the first derivative by the level slot |
| \(\partial^{2} f/\partial x^{2}\) | the second derivative by the level slot |
| \(x\) | the level slot, standing empty until the level of the process is substituted in at step six |
A partial derivativeThe rate of change with respect to one variable, holding the others fixed. by time is written even when it is zero. A written zero is different from a zero that was never considered. Coming back to the working a week later, the written zero shows that the time slot was checked. Nothing shows that if the line is missing.
How many derivatives does the procedure need?
How far does the expansion go?
To second order, and then it stops. Write the time piece, the first order piece and the squared piece, and put the pen down. The second order expansionCarrying the expansion as far as the squared term and no further. is the one place in the seven steps where a reader has to override a habit rather than follow one. Every expansion performed before this one was safely stopped one order earlier.
| \(df\) | the increment of the new process, the object the whole procedure produces |
| \(dS_t\) | the increment of the process, taken unexpanded from step one |
| \(\bigl(dS_t\bigr)^{2}\) | the squared increment, the term that must appear in writing |
| \(\tfrac{1}{2}\) | the coefficient the squared piece carries, written before the second derivative |
The check at this step is a count. There should be a term containing a squared increment in the working. If there is not, the expansion stopped one order too early and the rest of the procedure will run to completion anyway, producing something that looks exactly like an answer. A silent failure of that kind is the whole reason step four carries a check and steps one and two do not.
Going further than second order is not an error in the same sense, but it is wasted writing. A third order piece carries a cubed increment, and step five sends every cubed increment to zero. The truncationStopping an expansion, which here must happen at second order and not first. point is not a matter of taste. One order lower loses a term that survives. One order higher adds terms that do not.
How far does the expansion go?
How is the multiplication table applied to the expansion?
By looking things up. Multiply out the squared increment into its three products, then read each product off the table and write down what it becomes. The multiplication tableThe rule turning each product of the two increment symbols into what survives. is a lookup with four entries, and step five is only ever reading a cell. The step that looks like the hardest part of the procedure is the one part of it requiring no thought at all.
| \((dt)^{2}\) | the time increment against itself, which is written as zero |
| \(dt\,dW_t\) | the time increment against the Brownian increment, written as zero |
| \(dW_t\,dt\) | the same product in the other order, also written as zero |
| \((dW_t)^{2}\) | the Brownian increment against itself, replaced by the time increment |
Applied to the squared increment of the standard process, the multiplication runs like this. The squared increment expands into three products, one carrying the squared time increment, one carrying the cross product, and one carrying the squared Brownian increment. The table sends the first two to zero and turns the third into a time increment, leaving one surviving expression.
| \(\mu^{2} S_t^{2}\) | the coefficient on the squared time increment, sent to zero by the table |
| \(2\mu\sigma S_t^{2}\) | the coefficient on the cross product, sent to zero by the table |
| \(\sigma^{2} S_t^{2}\) | the coefficient on the squared Brownian increment, which survives |
| \(\sigma^{2} S_t^{2}\,dt\) | the one expression step five hands to step six |
There is a conversion chart on the wall of every workshop that deals in two systems of units, and nobody derives a conversion from first principles while standing at the bench. The bench worker reads the cell. Step five has the same character, and treating it as anything more elaborate is how a straightforward substitution becomes an argument that has to be got right.
Which step of the seven requires no thought at all?
How are the surviving terms collected into a drift and a diffusion?
By sorting. Everything carrying dt goes into the first bracket, everything carrying dW goes into the second, and whatever cancels inside each bracket is then cancelled. The word for this step is collectGrouping the surviving terms into the coefficient of dt and the coefficient of dW., and it is a sorting operation rather than a calculation: nothing new is created here, and every symbol in the two brackets arrived from an earlier step.
| first bracket | the drift of the new process, everything that came out carrying the time increment |
| second bracket | the diffusion of the new process, everything carrying the Brownian increment |
| \(\mu S_t\) | the drift coefficient of the process, copied from step one |
| \(\sigma^{2} S_t^{2}\) | the squared diffusion coefficient, arriving from the surviving cell at step five |
Now substitute the three derivatives into the two brackets. The process carries the level in both of its coefficients. On the logarithm the first derivative is one over the level and the second is minus one over the level squared. Every level in the first bracket meets a level in a denominator, and the same happens in the second. The cancellationTerms in the level dividing out, which is what makes the logarithm case clean. is why the logarithm is the case every treatment opens with. What comes out the far side carries no level at all.
| \(0.08\) | what the first order piece leaves once the level cancels |
| \(-0.020000\) | what the second order piece leaves, being half the squared volatility |
| \(0.060000\) | the drift of the new process, the two above added |
| \(0.200000\) | the diffusion of the new process, the volatility with the level cancelled |
On the logarithm, why does the level disappear from the answer?
What check confirms the result before it is used?
Two checks, run in order, and each one catches a different class of error. The first sets the volatility to zero. The second replaces the function with a straight line. Both are limiting caseA setting where the answer is already known, used as a check. tests, meaning the answer is already known before they are run, and a failure names which step to go back to rather than merely reporting that something is wrong.
Set the volatility to zero in the collected drift. The second order piece carries a squared volatility, so it must disappear, and what remains must be the ordinary chain rule. On the logarithm that is 0.080000, and anything else means the fault is in the collection at step six: a coefficient was copied wrongly or a term went into the wrong bracket. Check one catches sorting mistakes.
Now put the function back and make it a straight line in the level instead. A straight line has a second derivative of zero, so the second order piece is zero whatever the volatility is, and the procedure has to return the process itself scaled by the slope. If it does not, the fault is upstream in step three: the derivatives were taken against the wrong slot, or the second derivative was written where the first belonged. Check two catches differentiation mistakes.
The habit this mirrors is the one every laboratory has. A scale that reads a known weight wrongly will read an unknown one wrongly too, and the unknown reading alone never shows it. So a known weight goes on the scale before the weight that matters. Both checks here are known weights, and they weigh differently on purpose.
The volatility is set to zero and the procedure does not return the ordinary chain rule. What does that show?
What does the whole procedure look like run end to end?
Run cold, on the logarithm of the standard process, it fits on a single sheet. Both parameters of the process are stated at the outset rather than measured, and every figure below is a computed consequence of the two.
| Step | What is written | What comes out |
|---|---|---|
| 1 | the process in drift and diffusion form | 0.08 times the level, 0.20 times the level |
| 2 | the function, both slots shown | natural logarithm, level slot only |
| 3 | the three derivatives | 0, one over the level, minus one over the level squared |
| 4 | the expansion to second order | a time piece, a first order piece, a squared piece |
| 5 | the table applied to every product | 0.04 times the level squared, times dt |
| 6 | the survivors sorted into two brackets | 0.080000 less 0.020000, and 0.200000 |
| 7 | the two checks | drift 0.060000, diffusion 0.200000 |
| \(\mu - \tfrac{1}{2}\sigma^{2}\) | the drift of the logarithm, being 0.08 less 0.02 |
| \(\sigma\) | the diffusion of the logarithm, being 0.20 with the level cancelled |
| \(0.060000\) | the drift figure the seven steps produce, an educational illustration |
| \(0.200000\) | the diffusion figure the seven steps produce, an educational illustration |
Now the same seven steps on three other functions. The comparison shows which parts of the working change and which do not. Only step three changes. Steps one, four, five and seven are word for word identical whatever the function is, and step six is the same sorting operation applied to different symbols.
| The function | By time | By the level, once | By the level, twice | Drift | Diffusion |
|---|---|---|---|---|---|
| natural logarithm of the level | 0 | 1 over the level | minus 1 over the level squared | 0.060000 | 0.200000 |
| the level itself | 0 | 1 | 0 | 8.000000 | 20.000000 |
| the square of the level | 0 | twice the level | 2 | 2000.000000 | 4000.000000 |
| three times the level plus twenty | 0 | 3 | 0 | 24.000000 | 60.000000 |
The three lower rows are read in units of the function a year, taken at a level of Rs 100/-. The top row is level free and needs no such reading. Only one column in that table depends on the function at all. The procedure is fixed, and the function only ever enters at step three.
The function is about to be a straight line in the process. Before the switch: what should the procedure return?
Send four different functions through the same seven steps
The process never changes and neither does the frame. Move the control and only the three derivatives change, and everything below them follows from those three. The two bars at the bottom show the drift the procedure produces against the drift a first order stop would have produced.
The two bars respond to the control. A second derivative of zero means step five hands step six nothing at all. On the two straight line settings the bars are the same length. On the logarithm the second bar is shorter, and on the square of the level it is longer. The second order piece is not a correction that always shrinks the answer: it moves the drift in whichever direction the second derivative points, and on the square of the level it adds 400.000000 to a first order reading of 1600.000000.
Where does the procedure usually go wrong?
At step four, and it is done from habit rather than from confusion. The expansion is stopped at first order because every expansion the reader has performed before this one was safely stopped at first order, and the hand does it before the mind objects.
The error that gets made, and what it costs
The first order piece is written, the squared piece is not, and the procedure carries on. Step five is reached with nothing to look up, so it takes no time at all and feels like a step that was always going to be easy. Step six sorts what it was given into two brackets. Step seven, if it is run at all, is run against a result that has already lost a term.
On the logarithm the output is a drift of 0.080000 and a diffusion of 0.200000. The output has a drift and it has a diffusion, and it is written in exactly the shape a finished answer is written in. The error survives review for that reason. It is wrong by 0.020000, which is half the variance rate, and nothing in the shape of the result marks the gap. There is no leftover symbol, no dimension that fails to match, no term standing unpaired at the end of a line.
The cost is a wrong answer that reads as a right one, and a second cost that is easier to miss: a reader who skipped step four finds step five trivial and concludes the procedure is easier than advertised. The false conclusion is then carried to the next function, where the same term goes missing again. The only reliable defence is the count at step four. The count asks one question and takes one second: is there a squared increment written anywhere in the working?
Somebody stops the expansion at first order. What does the procedure give for the logarithm?
How does somebody checking another person's working use this?
Whoever reviews a derivation, whether that is a supervisor reading a research note, a quantitative analyst inheriting somebody else's model documentation, or a student marking their own working from a week ago, has the same problem: the output looks fine and there is no time to redo the whole thing. The seven steps turn that problem into four questions that take under a minute between them.
First, are both panels of step one and step two actually written down? A derivation that opens straight into derivatives has skipped them, and the coefficients it uses later were carried in somebody's head. Second, the derivatives are counted: there should be three, and one of them should be a second derivative. Third, a squared increment should appear somewhere in the working. If there is not one, the expansion stopped at first order and the answer is wrong by half the variance rate times the second derivative, whatever the function was. Fourth, check one can be run directly: setting the volatility to zero in the stated drift should bring the ordinary chain rule back.
Of the four, the third is the one that pays. One glance runs it, and it catches the only error in this procedure that leaves no trace in the output. The other three catch faults that usually announce themselves in some other way: a missing panel shows up as an unexplained symbol, a wrong derivative usually breaks the units, and a sorting mistake usually leaves a dt somewhere it should not be. A missing second order term does none of those things.
The same habit is worth keeping in one's own working, and the cheapest place to keep it is the margin. The count goes in the margin at step four: three pieces, one of them squared. The margin count is the same discipline as writing the tare weight on a jar before filling it. The number is trivial and it is never needed on the occasions when everything went right. Writing it every time, rather than only when something feels off, is what makes the count work.
One last note on how the procedure travels. Nothing above depended on the function being the logarithm or on the process being this particular one. Steps one, two, four, five, six and seven are written the same way for any function of any process in drift and diffusion form, and step three is the only place where the specific function enters at all. A person who has run the seven steps twice can run them on a function they have never seen. A written sequence makes that possible. An argument reconstructed each time does not.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for stochastic calculus and its applications in pricing theory | arxiv.org |
| Social Science Research Network | Working paper repository for the same material, including expositional treatments of the lemma | ssrn.com |
| Hull, Shreve and Wilmott | Standard texts on stochastic calculus setting out the lemma and the order in which it is applied | print editions |
The standard process is invented, as are its parameters and its starting level of Rs 100/-.
Educational material. Not advice on any investment, tax, budget or market position.
