The Ito Integral: Integrating Against Randomness
An Ito integral is the limit of a sum of the integrand multiplied by the increment of the random path across each step, with the integrand read at the start of every step. The reading point is a choice rather than a convention, and it changes the answer: on one constructed path the same integral comes out at minus 0.5 read at the start and plus 0.5 read at the end.
Almost everything ordinarily learned about integration is learned on paths that behave. The integral of a function over an interval is a number attached to that function, and every reasonable way of computing it converges on the same number. Chopping finer, using left-hand heights, using right-hand heights, using the height at the middle of each strip: the answer is the same to whatever accuracy is demanded. Robustness that reliable hides its own status. Most people never notice it is a result rather than a definition, and never notice it can fail.
The robustness fails here. In this guide a single fixed path is integrated against its own increments, and three perfectly standard ways of computing the sum return minus 0.500000, 0.000000 and plus 0.500000. Not three approximations to one answer. Three different answers, each exact, from one path that never changed. The difference between them is not sloppiness in the computation but a decision buried inside the definition, and once that decision is made explicit the whole of the subject that follows becomes readable.
What does it mean to integrate against a path at all?
The construction is the same without the randomness, and easier to see, so the randomness comes out first. Suppose something that changes over a year is being recorded, and what is wanted is a running total of some quantity multiplied by how much the thing moved. Not multiplied by how much time passed. Multiplied by how far it moved.
Here is the everyday version. A walker on a hill road carries an altimeter, and every hundred paces writes down two things: what the altimeter reads, and how much it changed over those hundred paces. The two are then multiplied and the products added up over the whole walk. The running product is not a total of altitudes and not a total of climbs. A total of altitude multiplied by climb depends on the shape of the walk in a way neither of the simpler totals does. Every integral below has that shape, and the only thing that changes is what happens when the road stops being a road.
Formally, the recipe is to cut the interval into steps, take the change in the path across each step, multiply that change by the integrandThe quantity that gets multiplied by each change in the path. Here it is a number that can itself move as the path moves., add the products, and then make the steps smaller. The change across a step is called the increment. The list of times marking where one step ends and the next begins is called a partitionA finite list of times, starting at the beginning of the interval and finishing at its end, cutting the interval into consecutive steps., and its longest step is called its meshThe length of the longest step in a partition. The mesh is the quantity driven toward zero when a limit over partitions is taken.. The limit is taken by driving the mesh to zero.
Nothing in that description says where inside a step the integrand should be read. The step runs from one time to the next, the integrand is moving during that step, and somebody has to say which of its values gets multiplied by the increment. In ordinary integration it never matters, so nobody bothers to say. The entire subject of this guide exists because on a random path it matters, and it matters by an amount that can be written down exactly.
How many answers can one integral, over one fixed path, with one fixed integrand, have?
Why can the ordinary construction not be used here?
Because of one scaling fact settled under quadratic variation, and it takes three lines to spend. Read the integrand at the start of a step and read it again at the end, and the two readings differ by however much the integrand moved during the step. The difference between the two readings is then multiplied by the increment. So the whole disagreement between the two ways of computing is a sum of terms, each one being a movement of the integrand multiplied by an increment of the path.
The size of such a term settles the question. On a smooth path, an increment across a step of length delta-t is of the order of delta-t, and the integrand moves by about the same order, so each disagreement term is of the order of delta-t squared. Add a number of them proportional to one over delta-t and the total disagreement is of the order of delta-t, vanishing in the limit. The theorem's conclusion is that the worry can be dropped, so nobody teaches it.
On a Brownian path the increment is of the order of the square root of delta-t. If the integrand is the path itself, then it moves by that same square root order too. Multiply the two and each disagreement term is of the order of delta-t. Add a number of them proportional to one over delta-t and the total does not go to nothing at all. The total stays put. The disagreement between reading at the start and reading at the end is a sum of squared increments, and a sum of squared increments is precisely the quantity that refuses to vanish on a Brownian path.
| \(W_{t_i}\) | the Brownian path read at the time the step begins |
| \(W_{t_{i+1}}\) | the same path read at the time the step ends |
| \(\Delta W_i\) | the increment across that step, being the end reading less the start reading |
| \(\Pi\) | the partition, with mesh \(\|\Pi\|\) being its longest step |
| \(T\) | the elapsed time, one year throughout this guide |
Look at what that algebra did. The two sums differ term by term, and each difference is an increment multiplied by itself. Nothing was approximated and nothing was dropped, so this is not an error estimate. The result is an identity. Whatever the path, the two readings are separated by exactly the sum of squared increments, and the only question is whether that sum survives the limit. On a smooth road it does not, and both readings land on the same number. On a Brownian path it does, and they never meet.
The right-hand panel invites a promise that is not there. The six points come from one path at six coarse grids, and a single path at a coarse grid says nothing about a limit. At three steps the two readings sit 0.031667 apart, close enough to look like agreement. At four steps they sit 1.345000 apart, an overshoot. Neither reading contradicts anything. The claim being made is that the separation tends to the elapsed time as the grid refines without end, and no finite grid on one path is evidence for or against a limit.
Two people compute the same sum, one reading the integrand at the start of each step and one at the end. What exactly separates their two answers?
Where inside each step is the integrand read, and why is that a choice?
A decision with a number attached stops being invisible, so give the decision a name and a number. Call the evaluation pointThe place inside a step at which the integrand is read. Writing it as a fraction turns a hidden decision into a number that can be moved. a fraction of the way through the step, running from zero at the very start to one at the very end. Reading at zero gives the Ito version. Reading at one half gives the Stratonovich version. Reading at one gives the backward version, a perfectly well defined object that nobody in this subject wants.
The path between two grid times is not otherwise pinned down by the grid, so reading at a fraction inside a step means reading along the chord that joins the start value to the end value. The chord convention is not a fudge. At a fraction of one half the chord produces exactly the average of the two end values, and that average is the standard definition of the midpoint version. The whole run from zero to one is a single well behaved sweep between two named objects.
| \(a\) | the evaluation fraction, zero at the start of a step and one at its end |
| \(S(a)\) | the sum produced by reading the integrand at that fraction |
| \(W_{t_i}+a\Delta W_i\) | the integrand read along the chord of the step |
| \(\Delta W_i\) | the increment of the path across the step |
| \(\sum(\Delta W_i)^2\) | the sum of squared increments over the grid |
The single line above carries the whole argument. The answer is not one number but a straight line of numbers, indexed by a decision. The steepness of the line is the sum of squared increments. On a smooth path the line is flat and the decision is invisible. On a Brownian path the line has a slope of one per year of elapsed time. Ordinary calculus is not the general case with randomness added; it is the special case in which this line happens to be flat.
What does the same integral give at three evaluation points?
Now the worked instance. The path is the locked path used throughout: twelve equal steps over one year, each increment being one of twelve driving values multiplied by the square root of one twelfth. The driving values are minus 0.5, 1.6, minus 1.3, minus 0.1, 0.1, 1.5, minus 1.3, minus 0.5, minus 1.4, 0.4, 0.9 and 0.6. The twelve values sum to zero exactly, so the Brownian path finishes the year where it began, and their squares sum to 12.0 exactly, so the sum of squared increments over the year is 1.000000 exactly. Both properties were built in rather than found. Every figure below therefore reproduces to six decimal places instead of approximately.
The integrand is the path itself. The standard process built on that path, an invented illustration with no market meaning, starts at Rs 100/-, climbs to Rs 111.08/- at month six and falls to Rs 93.74/- at month nine before finishing the year at Rs 106.18/-. The movement is large, and worth holding in mind while the integral below comes out at a number as small as a half.
| Step | Increment | Read at start | Product | Read at end | Product |
|---|---|---|---|---|---|
| 1 | minus 0.144338 | 0.000000 | 0.000000 | minus 0.144338 | 0.020833 |
| 2 | plus 0.461880 | minus 0.144338 | minus 0.066667 | plus 0.317543 | 0.146667 |
| 3 | minus 0.375278 | plus 0.317543 | minus 0.119167 | minus 0.057735 | 0.021667 |
| 4 | minus 0.028868 | minus 0.057735 | 0.001667 | minus 0.086603 | 0.002500 |
| 5 | plus 0.028868 | minus 0.086603 | minus 0.002500 | minus 0.057735 | minus 0.001667 |
| 6 | plus 0.433013 | minus 0.057735 | minus 0.025000 | plus 0.375278 | 0.162500 |
| 7 | minus 0.375278 | plus 0.375278 | minus 0.140833 | 0.000000 | 0.000000 |
| 8 | minus 0.144338 | 0.000000 | 0.000000 | minus 0.144338 | 0.020833 |
| 9 | minus 0.404145 | minus 0.144338 | 0.058333 | minus 0.548483 | 0.221667 |
| 10 | plus 0.115470 | minus 0.548483 | minus 0.063333 | minus 0.433013 | minus 0.050000 |
| 11 | plus 0.259808 | minus 0.433013 | minus 0.112500 | minus 0.173205 | minus 0.045000 |
| 12 | plus 0.173205 | minus 0.173205 | minus 0.030000 | 0.000000 | 0.000000 |
| Total | 0.000000 | minus 0.500000 | plus 0.500000 |
The row-by-row contrast is more convincing than the totals, so read the two product columns side by side. Step two contributes minus 0.066667 read at the start and plus 0.146667 read at the end, and the sign has flipped, not merely the size. Step three contributes minus 0.119167 and plus 0.021667, and the sign has flipped again. Nothing was computed differently in the two columns. The increment on each row is identical in both. All that changed is which end of the step supplied the number that multiplied it.
Reading at the middle of each step gives a third column, the average of the two shown, totalling exactly 0.000000. One path, one integrand, one grid, three totals: minus 0.500000, 0.000000 and plus 0.500000, each of them exact rather than rounded.
Two things are worth noticing in that picture. The three totals do not simply drift apart at a constant rate. Separation is fastest in the months where the path moved most, and that is exactly what an accumulation of squared increments does. The middle line also wanders on both sides of nothing before returning to it. Even the reading that agrees with ordinary calculus at the horizon does not agree with it along the way.
| \(W_T\) | the Brownian path at the horizon, which on the locked path is exactly zero |
| \([W]_T\) | the sum of squared increments over the grid, exactly 1.000000 here |
| \(T\) | the horizon, one year |
| \(\tfrac{1}{2}W_T^{2}\) | the answer ordinary calculus would give, with no correction at all |
Substituting the locked path into those three leaves arithmetic short enough to do in the head. The finishing value is zero, so its square is zero. The sum of squared increments is one. Half of nought less one is minus a half. Half of nought is nought. Half of nought plus one is plus a half. A formula that never looked at a single row confirms every figure in the table above, and such a check is worth having on any long summation.
The midpoint reading gives 0.000000, the value ordinary calculus predicts. Is that the Ito integral?
What exactly is the gap between the three answers?
The gap between the start reading and the end reading is 1.000000. The number should feel familiar: 1.000000 is the elapsed year, and also the sum of squared increments of the locked path over that year. The three descriptions are the same number arriving by three routes, and the coincidence is not a coincidence at all: the identity written above says the gap is the sum of squared increments, and the result set out under quadratic variation says that sum settles on the elapsed time.
Watch it accumulate rather than only checking it at the finish. The gap is not a correction applied at the end; it builds step by step alongside the two running totals, and at every intermediate month it equals the squared increments accumulated so far.
| Elapsed | Running start reading | Running end reading | Gap | Squared increments so far |
|---|---|---|---|---|
| Month 3 | minus 0.185833 | plus 0.189167 | 0.375000 | 0.375000 |
| Month 6 | minus 0.211667 | plus 0.352500 | 0.564167 | 0.564167 |
| Month 9 | minus 0.294167 | plus 0.595000 | 0.889167 | 0.889167 |
| Month 12 | minus 0.500000 | plus 0.500000 | 1.000000 | 1.000000 |
The last two columns agree on every row, to every decimal place, and they agree because they are the same sum written twice. The distance between the two ways of reading the integral is not an error, an approximation or a numerical artefact; it is the accumulated squared movement of the path, tracked exactly. Notice also that the running gap is not proportional to elapsed time on this one path. The gap stands at 0.375000 at month three, above one quarter, and moves only to 0.564167 by month six. Proportionality is what the limit delivers over many paths, and this is one path.
The two outer readings differ by 1.000000. What is that number?
What does reading at the start of each step buy?
So far the reading point looks arbitrary: three answers, pick one. The reading point is not arbitrary, and the reason is worth more than the definition itself. Reading at the start of each step means the integrand is fixed before the increment that multiplies it has happened. Reading anywhere else means the integrand is allowed to know something about the increment it is about to be multiplied by.
Here is the everyday version, and it is about counting rather than about anything traded. A weighing scale in a market makes the point. Reading at the start of a step is putting the weight on the scale and then letting the pan settle. Reading at the end is waiting for the pan to settle and then choosing what weight is claimed to have been put on. The second procedure produces a number, and the number is even reproducible, but it is not a measurement of anything that was decided in advance. The second procedure has quietly used the outcome to set the input.
An integrand that uses only information available when the step begins is called non-anticipatingDepending only on what is known up to the moment a step starts, and never on anything that happens after it. Also called adapted in this subject.. Multiply a non-anticipating integrand by an increment whose average is nil and the product has an average of nil too. The two are settled independently: the multiplier was fixed first, and the increment brought no bias. Add up terms each averaging nil and the total averages nil. Averaging to nil is what makes the Ito integral a fair gameA running total whose best forecast of any future level, given everything known now, is its current level. The formal name for it is a martingale., and a fair game is the object every later result in this subject is built on.
| \(H_t\) | the integrand, non-anticipating and square integrable |
| \(dW_t\) | the increment of the Brownian path under the physical measure P |
| \(\mathcal{F}_s\) | the information available at time \(s\) |
| \(\mathbb{E}[\cdot]\) | the average over outcomes, taken under P |
| \(s\) | any time before the horizon \(T\) |
Now see what the end reading does to that. The end reading is the start reading plus the sum of squared increments, and squared increments are never negative, so the end reading carries a built-in upward push of exactly the elapsed time. The end reading cannot average to nil. The choice of the start of each step is not a stylistic preference or an inherited habit; it is the only reading in the whole sweep from zero to one that leaves the resulting integral a fair game. Every other reading buys an unasked-for drift that would then have to be subtracted by hand.
Why is the integrand read at the start of each step rather than at the end?
The integrand is about to be read at the end of each step rather than the start. Does the answer change?
Slide the reading point through the step and watch the answer travel
The locked path never changes and neither does the grid. The only thing that moves is where inside each step the integrand is read. The top strip marks the twelve read points on the path, the middle strip redraws the running total, and the scale at the bottom carries the answer between its three named values.
What is a Riemann Integral, and why does it not care where the integrand is read?
The comparison is only sharp if both sides are defined rather than assumed, so set the ordinary object down properly. A Riemann sumThe sum of the integrand read at a chosen point in each step, multiplied by the width in time of that step. cuts the interval into steps, reads the integrand at some chosen point inside each step, multiplies by the width in time of that step, and adds. The chosen point is called the tag pointThe place inside a step at which the integrand is read when a Riemann sum is formed. The tag point is deliberately left unconstrained by the definition., and the definition deliberately leaves it free.
The Riemann Integral is what those sums converge to as the mesh goes to zero, and the theorem that makes it useful is that for a well behaved integrand every choice of tag point converges to the same limit. Left ends, right ends, midpoints, or a point picked at random inside each step: one number. The tag point is therefore never taught as something to worry about.
| \(f\) | the integrand, a well behaved function of time |
| \(\xi_i\) | the tag point, anywhere inside the step, and the limit does not depend on it |
| \(t_{i+1}-t_i\) | the width in time of the step, which is what the integrand multiplies |
| \(\|\Pi\|\) | the mesh of the partition, driven to zero |
The reason the tag point is free is now easy to see from what came earlier. Two tag choices differ by the movement of the integrand across a step, multiplied by the width of that step. The movement is of the order of the width, so the difference per step is of the order of the width squared. Adding a number of them proportional to one over the width leaves a total of the order of the width itself, and that total vanishes. The insensitivity that ordinary calculus relies on is not a property of integration; it is a property of multiplying by a width in time.
Does an ordinary integral depend on where inside each step the integrand is read?
So what is a Stochastic Integral, stated plainly?
A Stochastic Integral is what results when the thing being multiplied is not a width in time but the increment of a random path. The substitution is the whole of it, and everything else in this guide follows. Because the increment is of the order of the square root of the width rather than of the width, the argument that freed the tag point collapses, and the reading point becomes part of the definition rather than an irrelevance.
The Ito integral is the Stochastic Integral built with the reading point at the start of every step. The Stratonovich version is the one built with the reading point at the midpoint. Both are Stochastic Integrals; they are different Stochastic Integrals, they obey different rules, and neither is an approximation to the other. Every result downstream uses the Ito version, for the fair game reason set out above.
| \(H_{t_i}\) | the integrand read at the moment the step begins, and never later |
| \(W_{t_{i+1}}-W_{t_i}\) | the increment of the Brownian path across that step |
| \(\|\Pi\|\to 0\) | the mesh driven to zero, which is how the limit is taken |
| \(\int_0^T H_t\,dW_t\) | the resulting object, itself random rather than a fixed number |
One thing about the Ito integral is a habit break, and deserves saying out loud. An ordinary integral of a fixed function is a fixed number. A Stochastic Integral is not. Its value depends on which path turned up, so it is a random quantity in its own right, and the minus 0.500000 in this guide is its value on one particular constructed path rather than its value in general. The shape is what holds in general: the answer is always the squared finishing value less the elapsed time, all halved.
How is an Ito Integral checked once it has been handed over?
Three questions, and the first one catches almost everything. None of them requires redoing the calculation. The routine is therefore worth having when the calculation belongs to somebody else.
- Ask where in each step the integrand was read
Not what the integrand was. Where it was read. If the answer is anything other than the start of the step, the number at hand is not an Ito integral, whatever it has been called.
On the locked path the three legitimate readings give minus 0.500000, 0.000000 and plus 0.500000.
- Ask whether the integrand could have known the increment
Anything defined using the maximum, the minimum, the average or the finishing value over the whole interval has looked ahead. The construction does not apply to it, and the fair game property is gone.
An integrand that peeks at the end of the year is not non-anticipating, however carefully the sum is then added up.
- Ask whether the answer matches the start-of-step form
Where the integrand is the path itself, the answer must be the squared finishing value less the elapsed time, all halved. Anything higher by exactly the elapsed time is the end reading wearing the wrong label.
Half of nought less one is minus a half, which is what the twelve row table totals to.
An integrand is defined as the highest level the path reaches over the whole year, and it is integrated against the increments. What is wrong with it?
The error that gets made, and what it costs
The evaluation point did not matter in any integral met before this one, so it is assumed not to matter here. The assumption is not carelessness. A well earned habit is being applied one setting past where it holds, and there is no warning at the boundary.
On the locked path that assumption is wrong by 1.000000 on an integral whose intended value is minus 0.500000. The error is twice the size of the answer and it has the opposite sign. Reading the integrand at the end of each step is a perfectly well defined operation producing a perfectly well defined number, so nothing in the arithmetic flags it. There is no division by zero, no overflow, no negative under a square root and no unstable sum. Both calculations are internally correct, and only one answers the question asked.
The cost is a quantity nobody intended, computed correctly, carried forward and used. The tell is precise and worth memorising: two answers to the same integral that differ by exactly the elapsed time were produced by two different reading points, every single time. When one figure comes out exactly one year larger than another on a one year integral, the summation is not where the fault lies.
Two people compute the same integral over the same year and get answers differing by exactly 1.000000. What happened?
How does somebody reviewing a calculation actually use this?
An Ito integral is rarely built from its definition outside a textbook. A number that came out of one will one day be handed over instead, and the questions that make it checkable have to be known. The handover reaches a risk reviewer looking at a hedging calculation, a quantitative researcher reading somebody else's derivation, and anybody validating numerical code against a closed form.
The move that pays is the one from the check routine above: ask about the construction rather than about the arithmetic. The column was never wrong, so somebody who re-adds a twelve row column is checking the wrong thing. The disagreement lives one level up, in a decision made before any number was written down, and it is invisible in the output. The most expensive errors in this subject are not arithmetic errors; they are correctly executed calculations of the wrong object.
There is a second, quieter use. When a discretised calculation and a closed form disagree by a stable amount that does not shrink as the grid refines, the reading point is the first suspect and the elapsed time is the number to compare the gap against. A gap that halves when the grid doubles is a discretisation effect and will go away. A gap that sits at the elapsed time however fine the grid becomes is the signature described in this guide, and no amount of extra steps will touch it. Distinguishing those two failures by their behaviour under refinement costs one extra run and saves an afternoon.
No authority anywhere sets the definition of an integral, no regulator publishes a reading point, and no market convention alters what a limit over refining partitions equals. The result is a statement about paths rather than about anything traded, so it holds identically everywhere and nowhere in particular.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for constructions of the stochastic integral and their use downstream | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Ito | Originator of the start-of-step construction that carries his name | named in the text only |
| Stratonovich | Originator of the midpoint reading that carries his name | named in the text only |
| Hull, Shreve and Wilmott | Standard textbook treatments of the stochastic integral and its notation | named in the text only |
The standard process and the locked path are invented.
Educational material. Not advice on any investment, tax, budget or market position.
