Ito Calculus vs Ordinary Calculus: What Changes and Why
Ordinary calculus is built on increments that shrink faster when they are squared. Ito calculus is built on increments that do not. The behaviour of squared increments sorts every rule into three groups: the ones that survive untouched, the ones that survive with a single extra term, and the ones with no counterpart at all. The third group is the one readers are least prepared for.
The pieces of this subject were built one at a time. The integral came first, then what its symbols mean, then how two random quantities move together, then the chain rule that replaces the familiar one, and then the operator that packages all of it. The comparison below sets the two bodies of mathematics side by side and reads off, rule by rule, what happened to each one on the way across.
There is a measuring idea that carries most of the distance before any notation appears. Pacing out a stretch of coastline with a one kilometre ruler gives a number. The shorter ruler follows inlets the long one stepped over, so pacing the same stretch with a one hundred metre ruler gives a larger number. A one metre ruler makes it grow again. A fence is straight enough that a finer ruler finds nothing new, so a garden fence measured the same three ways gives the same answer all three times. A smooth path is the fence and a random path is the coastline, and every difference in this guide is a consequence of that one distinction.
A comparison entered cold has to define both sides in full before it sets either against the other. The first two sections therefore describe each calculus from its foundation rather than from its rules. A reader who stopped after those two sections would still come away with two complete pictures. Nothing is contrasted until both are on the table.
What is ordinary calculus built on?
Ordinary calculusThe calculus of smooth paths, where squared increments vanish. is built on one assumption about the paths it is willing to work with, and everything else in it is a consequence. The assumption is that a path has a slope at every point. Given that, the change in the path over a short stretch of time is roughly the slope multiplied by the length of that stretch, and the error in that statement is smaller than the stretch itself.
Now square that change. If the change is proportional to the length of the stretch, then the square of the change is proportional to the square of that length. Cut the year into a hundred pieces and each squared change is proportional to one ten thousandth, and there are only a hundred of them to add up, so the total is proportional to one hundredth. Cut it into a thousand pieces and the total is proportional to one thousandth. The sum of squared changes over a smooth path does not merely get small, it goes to zero, and it is the fact that it goes to zero that makes ordinary calculus work the way it does.
Take the straight line that runs from zero to 0.060000 over one year, the growth of the logarithm of the standard process and a genuinely smooth path. Cut the year into twelve equal steps and each step of the line is 0.005000, whose square is 0.000025, and twelve of those add to 0.000300. Cut the year into forty eight steps and the total is 0.000075. At one hundred and ninety two steps it is 0.00001875, and at seven hundred and sixty eight steps it is 0.0000046875. Every quartering of the step divides the total by four exactly, and there is no partition fine enough to stop it.
| \(f\) | a smooth path, here the straight line whose slope is 0.06 a year |
| \(t_k\) | the times cutting the year into \(n\) equal steps, running from 0 to 1 |
| \(n\) | the number of steps in the partition |
| \(0.0036\) | the slope squared, being 0.06 multiplied by itself |
Because that total vanishes, an expansion of any well behaved function of a smooth path can stop after the first derivative. The second derivative multiplies a squared change, and a quantity that is being multiplied by something that vanishes contributes nothing. So the ordinary chain rule has one term. The ordinary product rule has two. Integration by parts has three pieces and no correction. The ordinary link between differentiating and integrating holds in both directions. Every rule taught in ordinary calculus is a rule about first derivatives, and the reason it is safe to stop there is that the squared changes went to zero.
A reader who has only ever worked with smooth paths has never had to notice how much rests on that one fact, so it is worth being blunt about it. The assumption is invisible in ordinary use precisely because it is never violated there.
What is Ito calculus built on?
Ito calculusThe calculus of paths whose squared increments do not vanish. is built the same way, on one statement about the paths it works with, and everything in it is a consequence of that. The statement is the opposite one. For a Brownian path, the sum of squared changes over a partition of the year does not go to zero. It goes to the elapsed time.
The reason sits in the size of a Brownian change. Over a stretch of length one twelfth of a year, a Brownian change has a typical size proportional to the square root of one twelfth, 0.288675, rather than to one twelfth itself. Squaring that returns one twelfth, not one hundred and forty fourth. So each squared change is proportional to the length of its stretch rather than to the square of it, and adding twelve of them gives the whole year. Cutting the year finer makes each squared change smaller in exact proportion. The number of them grows in exact proportion at the same time, and the total does not move.
The locked path is built so that this reads exactly rather than approximately. Its twelve driving values are minus 0.5, 1.6, minus 1.3, minus 0.1, 0.1, 1.5, minus 1.3, minus 0.5, minus 1.4, 0.4, 0.9 and 0.6, each multiplied by 0.288675 to give that month's change. The twelve driving values have a sum of squares of exactly 12.0, a construction choice rather than a coincidence. The twelve squared changes therefore add to one twelfth multiplied by 12.0, and that product is 1.000000. On the locked path the sum of squared changes over the year is exactly 1.000000, the elapsed year, and that single number is the whole foundation of this side of the comparison.
| \(W_t\) | Brownian motion at time \(t\) under the physical measure \(\mathbb{P}\) |
| \(\Delta t\) | the length of one step, one twelfth of a year |
| \(z_k\) | the driving value for month \(k\), from the twelve published numbers |
| \(\sum z_k^{2}\) | 12.0 exactly, a construction choice of the locked path |
| \(T\) | the horizon, one year |
| \(\mathbb{P}\) | the physical measure, the rule under which the path is described here |
The sum reaching the elapsed year is the reason a second order term can no longer be dropped. If a squared change contributes something rather than nothing, then an expansion that stops at the first derivative has thrown away a real quantity, and the size of what it threw away is set by the second derivative multiplied by the elapsed time. The second order term is not a subtlety at the edges. On the standard process it is worth a fifth of the drift, as the worked instance below shows.
The everyday version is a weighing scale that wobbles. If the wobble is small and settles, the reading is the weight and nothing else. If the wobble never settles, no matter how brief the observation, then the average reading is still the weight, but anything that depends on the square of the reading now carries the wobble too. Squared quantities are where a disturbance that averages to nothing stops averaging to nothing.
What single fact accounts for every difference?
Only one, and it has now been stated twice from two directions. On a smooth path the sum of squared changes over a partition goes to zero. On a Brownian path it goes to the elapsed time, and on the one year horizon of the standard process the elapsed time is 1.000000. Every row of the comparison that follows is that one fact applied once.
Put those two behaviours on the same picture and the gap is not a matter of degree. Refine the partition on the straight line and the total drops by a factor of four at every quartering, heading for zero without a floor. Refine it on the Brownian path and there is nothing to refine toward: the value the sum converges to is the elapsed year, and refining does not move it.
One caution belongs here, and it is the honest half of the claim. The flat line above is the value the sum converges to, and a single path read at a coarse partition is noisy rather than obedient. Read the same locked path at one step and the sum is 0.000000. The path returns to exactly where it started. At two steps it is 0.281667, at three steps 0.031667, at four steps 1.345000, at six steps 1.018333 and at twelve steps 1.000000. The sequence overshoots, undershoots and does not climb tidily to one, and any picture that draws a smooth rise is drawing something these numbers do not do. The claim is about the limit, and the twelve step reading lands on it exactly because the path was constructed to make it do so.
| Partition of the same locked path | Sum of squared changes |
|---|---|
| One step, start to finish | 0.000000 |
| Two steps | 0.281667 |
| Three steps | 0.031667 |
| Four steps | 1.345000 |
| Six steps | 1.018333 |
| Twelve steps, the published partition | 1.000000 |
The one step reading is the sharpest line in this table and it is worth sitting with. A partition that looks only at the start and the finish sees no variation at all in a path that travelled between Rs 93.74/- and Rs 111.08/- on the standard process. Coarse measurement does not merely lose precision here. Coarse measurement can report zero for something that is plainly not zero, and a one kilometre ruler does exactly that to a coastline.
What single fact accounts for every row of the comparison that follows?
Three groups follow: unchanged, one extra term, and no counterpart. Which group is the smallest?
Which rules survive completely unchanged?
LinearityConstants factoring out and sums splitting, which survives unchanged., and essentially nothing else. Linearity standing alone is the surprise this comparison delivers, and the bareness of it is worth stating starkly before anything is qualified. A constant multiplying an integrand comes out of the integral untouched. A sum inside an integral splits into a sum of integrals untouched. There is no correction anywhere in either statement, no extra term waiting, and no condition attached beyond the ones the integral already needed.
Additivity over adjoining stretches of time comes through with it, and it is really the same property wearing a different hat: the integral from the start of the year to month six plus the integral from month six to the end equals the integral over the whole year. Additivity is what makes it possible to build an integral step by step at all.
| \(f_t,\,g_t\) | two integrands, each known from the information available at time \(t\) |
| \(a,\,b\) | two constants, not depending on time or on the path |
| \(W_t\) | Brownian motion under the physical measure \(\mathbb{P}\), the integrator |
| \(T\) | the horizon, one year |
The reason linearity survives is that it never touched a squared change in the first place. Ask what the ordinary proof of linearity leans on and the answer is that a finite sum can be rearranged, and rearranging a finite sum is arithmetic with nothing to do with how the path behaves. A rule that does not consult the second order behaviour of the path cannot be disturbed by a change in that behaviour. Asking what a proof leans on is a useful test in its own right, and the last section turns it into three steps.
Notice how thin this group is. A reader arriving with a full working knowledge of ordinary calculus usually expects most of it to carry over with a footnote or two. The property that carries over untouched is the one that would be true of any sensible notion of an integral whatsoever, and almost nothing beyond it survives.
Which rule survives the move completely unchanged, with no correction anywhere?
Which rules survive but gain exactly one term?
Three of them, and the extra term is the same object every time: a quantity built from squared changes. The chain rule gains it, the product ruleThe rule for differentiating a product, which gains the paired term here. gains it, and integration by partsThe rule swapping which factor is differentiated, which gains the same term. gains the identical one the product rule does. Three rules, one correction, and once its origin is clear anybody can work out which rules must carry it before being told.
The chain rule has a number attached that can be checked, so it comes first. The ordinary chain rule says the change in a function of a process is the first derivative multiplied by the change in the process. The Ito version says the same thing plus one half of the second derivative multiplied by the squared change, and the squared change is no longer nothing.
| \(S_t\) | the standard process at time \(t\), in rupees, starting at Rs 100/- |
| \(f\) | a twice differentiable function of the level, here the natural logarithm |
| \(f',\,f''\) | its first and second derivatives with respect to the level |
| \(\sigma\) | the volatility of the standard process, 0.20 a year |
| \(\sigma^{2}S_t^{2}\,dt\) | the squared change of the process over the instant, which is not zero |
Put the natural logarithm through it. Its first derivative is one over the level and its second derivative is minus one over the level squared. The extra term is therefore one half multiplied by minus one over the level squared multiplied by the volatility squared multiplied by the level squared, and the levels cancel exactly, leaving minus half the variance rate. The variance rate of the standard process is 0.04 exactly, so half of it is 0.02 exactly, and the extra term is minus 0.020000 a year regardless of where the process happens to be standing.
The drift of the standard process is 8 per cent a year, and the growth of its logarithm is 6 per cent a year, and the entire gap of 0.020000 is that one extra term. This is not an approximation and it does not shrink at fine partitions. The gap separates the average level at the horizon, Rs 108.33/-, from the middle outcome at the horizon, Rs 106.18/-, where the locked path finishes because its driving values sum to zero.
The product rule gains the paired quantity of the two processes involved. The paired quantity is the running total of the products of their changes rather than the squares of one process alone. Where the two processes move together the paired quantity is positive, where they move against each other it is negative, and where they are unrelated it is zero and the ordinary rule comes back on its own.
| \(X_t,\,Y_t\) | two processes over the same horizon, each driven by its own randomness |
| \([X,Y]_t\) | their paired quantity, the running total of the products of their changes |
| \(d[X,Y]_t\) | its rate, being \(\rho\,\sigma_X\sigma_Y\) a year for each unit of the product |
| \(\rho\) | the correlation between the two, minus 0.7 on the locked pair |
| \(\sigma_X,\,\sigma_Y\) | their volatilities, 0.20 a year each on the locked pair |
Put the locked numbers on it. At a correlation of minus 0.7 and volatilities of 20 per cent a year each, the rate of the paired quantity is minus 0.7 multiplied by 0.20 multiplied by 0.20, giving minus 0.028000 a year. The product rule and integration by parts each gain a term worth minus 0.028000 a year on the locked pair, and it is the same term in both, appearing with a plus sign in one and a minus sign in the other. A reader who computes the ordinary two piece product rule and stops is short by that amount for every year the two processes run together.
Everyday version, and it is the one that makes the sign stick. Two shops on the same street both sell umbrellas and sunglasses. Their daily takings, multiplied together, behave in a particular way. If the days that are good for one are bad for the other, the product is systematically smaller than the product of their averages, and the shortfall is the paired quantity. The ordinary product rule assumes there is no such systematic pairing. The Ito version measures it and puts it back in.
The product rule gains a term. On the locked pair, what is that term worth a year?
Which ordinary operations have no counterpart at all?
The third group is the one readers are least prepared for. The first two groups fit a comfortable story: some rules are fine, some rules need a patch. The third group breaks that story. Some ordinary operations have no Ito version, not because nobody has worked one out, but because the operation itself has no meaning on these paths. There is no term to add, and nothing to add it to. A rule with no counterpartA rule in one calculus matching a rule in the other, which group three lacks. is a different kind of failure from a rule that needs correcting.
The first is the most basic operation in all of ordinary calculus: differentiating the path with respect to time. Ask a Brownian path for its slope at any particular moment and there is no answer. The path is nowhere differentiableHaving no slope at any point, as a Brownian path has.. The slope is not merely hard to compute at a few awkward points. No such number exists at any point of the path at all.
The reason is visible in the arithmetic of the previous sections. A slope is what a difference quotient settles down to as the step shrinks. On a Brownian path the change over a step of length one over n has a typical size of one over the square root of n, so the difference quotient has a typical size of the square root of n, and that grows without limit as the step shrinks. It does not settle. It blows up.
The second member of this group is the ordinary fundamental theoremThe ordinary link between differentiating and integrating, which fails here., and the failure has a number attached to it. Ordinary calculus says that integrating a quantity against its own changes gives one half of its final value squared less one half of its starting value squared, and nothing else. On the locked path the Brownian motion starts at zero and finishes at zero, so the ordinary answer is 0.000000 exactly. The correct answer is minus 0.500000.
| \(W_t\) | the Brownian path, the integrand and the integrator at once |
| \(W_T\) | its value at the horizon, 0.000000 on the locked path by construction |
| \(\tfrac{1}{2}W_T^{2}\) | 0.000000, which is the whole of the ordinary answer |
| \(\tfrac{1}{2}T\) | 0.500000, half the elapsed year, which the ordinary answer has no room for |
The ordinary answer is short by exactly half the elapsed year, and the missing half is the sum of squared changes divided by two, the one fact making its appearance again. The two answers here do not merely differ in size, and that makes this the sharpest single case in the guide. The ordinary answer says nothing happened and the correct answer says something specific happened, on a path that visibly moved.
Why does this belong in the third group rather than the second, when there is plainly a number available to fix it? Because the ordinary theorem is a statement that the two operations undo one another, and here they do not, whatever number is bolted on afterwards. Correcting an answer by hand each time is not the same as having a rule. The Ito version of this relationship exists and is covered separately, but it is a different statement rather than the ordinary one with a term appended, and reading it as the latter is how people end up applying it where it does not hold.
Why is differentiating the path with respect to time in group three rather than group two?
On the locked path, by how much is the ordinary answer for that integral short, and why that amount?
What does the whole comparison look like on one process?
Here is the entire table on the standard process and the locked path, with a number in every row that can be checked. The standard process starts at Rs 100/-, its drift is 8 per cent a year, its volatility is 20 per cent a year, its variance rate is 0.04 exactly and the horizon is one year. The locked path is the twelve step path published for this reading order, whose driving values sum to zero and whose squares sum to 12.0.
| Group | The rule | Ordinary | Ito |
|---|---|---|---|
| One | Linearity of the integral | holds | holds |
| One | Additivity over adjoining stretches | holds | holds |
| Two | Chain rule, growth of the logarithm a year | 0.080000 | 0.060000 |
| Two | Product rule, extra term a year | 0.000000 | minus 0.028000 |
| Two | Integration by parts, extra term a year | 0.000000 | minus 0.028000 |
| Three | Slope of the path with respect to time | exists | no such number |
| Three | Integral of the path against its own changes | 0.000000 | minus 0.500000 |
| The fact | Sum of squared changes over the year | 0.000000 | 1.000000 |
Every row above the bottom row is that row applied once, so read the bottom row first and then read upward. The chain rule row is half the second derivative multiplied by 1.000000 rather than by zero. The product rule row is the paired version of the same thing. The integral row is half of 1.000000 with a minus sign in front. The slope rows are what happens when a quantity that should shrink like the square of the step only shrinks like the step. Nothing else is being remembered anywhere in this table.
The ordinary column deserves a closer look in the last two rows. It is not producing a slightly wrong number. In one row it produces zero for a quantity that is not zero, and in the other it produces an answer to a question that has no answer. A wrong number invites a check and a confident zero does not, so those are the rows that catch people.
The quadratic variation can be set to zero in the control below. What happens to the table?
Move the one number underneath the table
The control sets the sum of squared changes over the year. At 1.000000 the setting is the Brownian path and the worked instance above. As it falls, the rows walk out of group two and group three into group one. At 0.000000 the table is ordinary calculus.
| Setting | Growth of the logarithm | The integral |
|---|---|---|
| 1.000000, the Brownian path | 0.060000 | minus 0.500000 |
| 0.500000, halfway | 0.070000 | minus 0.250000 |
| 0.000000, a smooth path | 0.080000 | 0.000000 |
How is the table read as consequences rather than rules?
By asking, of any rule at all, one question: does its ordinary proof lean on the second order behaviour of the path? If it does not, the rule is in group one. If it does and the operation still makes sense, the rule is in group two and the extra term is whatever the second order piece turns out to be worth. If the operation itself needed a slope, the rule is in group three and there is nothing to correct.
The three questions are three steps, and none of them requires having seen the rule before. The first expands the quantity in question to second order. The second applies the familiar multiplication rules. The multiplication rules say that a squared change of the Brownian path contributes the elapsed time, and that a change in time squared and a change in time multiplied by a Brownian change both contribute nothing. The third looks at what survived. If nothing survived beyond first order the rule sits in group one, if something survived it sits in group two, and if the first order piece never existed it sits in group three.
The group settles what comes next, so knowing which group a rule is in is more useful than knowing the rule. A group one rule is used as it stands. A group two rule is used with one term added and no more, and a second added term is an error. A group three rule is set aside in favour of the statement that replaces it. Those three instructions cover every rule in this subject area, including the ones nobody has written down anywhere.
How does somebody checking a derivation use this?
Using this guide requires no rebuilding of anybody's mathematics. Three checks, all cheap, catch most of what goes wrong when ordinary habits are carried onto a random path, and all three can be run on a printed derivation with nothing else to hand.
- Count the terms in every chain rule step
Find each place where a function of the process has been expanded. An ordinary expansion has one term and an Ito expansion has two. A step with one term where the process is random is a group two rule being used as a group one rule.
On the logarithm of the standard process that omission is worth 0.020000 a year, a quarter of the drift of 0.080000.
- Look for a product of two random quantities with no third term
Wherever two processes are multiplied together, there should be a paired term. If the working shows only the two ordinary pieces, ask what the correlation between them is assumed to be. Writing nothing is the same as asserting zero.
At a correlation of minus 0.7 and volatilities of 20 per cent each, writing nothing costs minus 0.028000 a year.
- Hunt for a slope of the path itself
Any expression that divides a change in the path by a change in time has asked for a number the path never provides. This is a group three case, so there is no term to add and the step has to be rewritten rather than corrected.
The typical difference quotient doubles at every quartering of the step, so no finer partition rescues the step.
All three checks are the same check asked three ways: has somebody used a smooth path rule on a path that is not smooth? The single check is worth carrying away even if every formula in this guide fades. Long after the notation has been forgotten, the check still catches the error. It survives being asked about unfamiliar mathematics too, and the table cannot claim as much.
The error that gets made, and what it costs
Learning the table as a list of rules to memorise rather than as consequences of one fact. The mistake is an attractive one. The table is short, the rows are memorable, and reciting them feels like understanding. A reader who has memorised the seven rows can answer any question about those seven rows perfectly well.
Then they meet the eighth. There are many, and this sequence has already used several: the operator that packages the chain rule, the paired quantity for two processes driven by the same randomness, and the form the chain rule takes for a function of both the level and the time. None of those is a row in this table.
The cost is not a wrong answer, and a wrong answer would at least be visible. It is a wall. The memorising reader produces no answer at all at the first unfamiliar case, and worse, has no way of telling whether the half remembered rule in their head belongs to group one, group two or group three. The reader holding the one fact does it in a line: expand to second order, apply the multiplication rules, and see whether the surviving term is zero.
A rule appears that is not in the table at all. What then?
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for work on stochastic integration and the calculus built on it | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Hull, Shreve and Wilmott | Standard texts on stochastic calculus for finance | in book form |
| Kiyosi Ito | The original papers setting out the stochastic integral and the lemma compared here | journal literature |
The standard process, the locked path and the locked pair of correlated processes are invented.
Educational material. Not advice on any investment, tax, budget or market position.
