Infinitesimals in Stochastic Calculus: Reading dt and dW
Neither symbol is a number. Both are shorthand for what happens to a sum as its steps are made smaller without end, and the shorthand earns its place only because the two behave differently: the time symbol is of the order of the step, and the increment symbol is of the order of its square root. Everything in this subject that surprises a reader comes from that single difference in order.
The whole apparatus that follows is bookkeeping. The bookkeeping decides which terms in a sum survive being refined and which are driven to nothing. Once the size of each symbol is fixed, that decision stops being a matter of judgement and becomes arithmetic on exponents. A reader who expects a deep new idea will hunt for one and miss the shallow one that is actually doing the work. The subject holds no deep idea at this point. One number, one half, sits where a reader expects a one, and every strange consequence in this subject comes out of that swap. The word infinitesimalA shorthand for what a term does in a limit, rather than a name for any small quantity. No quantity ever equals it. is unfortunate, because it sounds like the name of a very small number, and it is not the name of any number at all.
Are these two symbols numbers at all?
No, and the trouble begins at the moment they are taken to be. A reader who has just met an expression carrying both symbols, asked what dt is, will honestly answer with some version of a very short interval, a tiny slice of time, perhaps a millionth of a year. Asked what dW is, the answer is usually a very small movement of the process. Both answers feel harmless. Neither is correct, and the second one costs the reader the subject.
Consider what the notation actually does. An expression written in differentials is a compressed way of writing a statement about two sums, each taken over a partitionA division of an interval into a finite list of consecutive pieces, marked by the times where one piece ends and the next begins. of the interval, and each of those sums is then refined without end. The differential form is the abbreviation. The integral form is the statement. A differential is never evaluated. There is nothing there to evaluate. The sums a differential stands for are what gets evaluated, and the differentials are a way of writing down which sums those are without drawing them out every time.
| \(X_t\) | a general process, being whatever quantity the expression is about |
| \(W_t\) | standard Brownian motion under the physical measure P |
| \(a\) | the coefficient multiplying the time symbol, whatever it happens to be |
| \(b\) | the coefficient multiplying the increment symbol |
| \(T\) | the horizon, being one year throughout this reading order |
Read the right hand side and notice that dt and dW have gone. Two integrals are left, and an integral is a limit of a sum over a shrinking grid. The symbols name those two sums without writing them out, and they have no life outside that job. The question of how big dW is has no answer. The answerable question is how a sum of many increments behaves as the grid gets finer.
An everyday version of the same confusion helps here. A household paying a monthly rent knows the rent per month, and can compute a rent per day by dividing, and a rent per hour by dividing again. At no point does anyone imagine a rent per instant sitting in a drawer as a physical amount of money. The rate is a description of how the total accumulates, and the total is what somebody actually pays. Differentials are exactly that: descriptions of how a total accumulates, not amounts sitting anywhere. The difficulty beyond rent is that one of the two descriptions accumulates in a way nobody has any intuition for.
Is dW a small number?
What order is the time symbol, and why start there?
The time symbol comes first because it is the ordinary case, and because what is odd about the other one only shows against something ordinary held beside it. A year divided into n equal pieces gives each piece a length of one over n. Doubling the number of pieces halves each piece. Multiplying the number of pieces by a hundred divides each piece by a hundred. Dividing and halving is the entire behaviour of dt, and no reader has to be persuaded of any of it.
The word for this behaviour is orderHow fast a term shrinks as the grid is refined, expressed as a power of the step. Order decides whether the term survives being added up.. Saying that a quantity is of the order of the step means that when the step is divided by a hundred, the quantity is divided by about a hundred as well. Saying it is of the order of the step squared means that dividing the step by a hundred divides the quantity by about ten thousand. Order is not a statement about how big something is at any particular grid. Order is a statement about how a term responds to refinement. When the whole construction is a limit over refinements, response to refinement is the only thing that matters.
Order alone decides everything, for the following reason. Whatever term is under consideration gets added up over the whole year, and the number of terms in that sum is itself governed by the step: there are one over dt of them. So a term that is of the size of the step, added up one over dt times, gives back the horizon and holds still at every grid. A term that is of the size of the step squared, added up the same number of times, gives back the horizon multiplied by one more factor of the step, and that extra factor drives it to nothing.
| \(\Delta t\) | the length of one piece of the grid, equal to \(T/n\) on an even grid |
| \(n\) | the number of pieces the horizon is divided into |
| \(p\) | the order of one term, being the power of the step it behaves like |
| \(T\) | the horizon, one year throughout this reading order |
| \(p-1\) | what is left over after the count of terms has cancelled one power |
Counting powers is the whole decision procedure, and every verdict below is an exercise in reading off the value of p. The counting is already settled: there are always one over dt terms, and that count cancels exactly one power of the step. So the question is never how small a term looks. The question is whether the term has one power of the step to spare after the counting has taken its share.
Put numbers on it at the grid this reading order uses. Twelve steps in a year makes dt equal to one twelfth, or 0.083333. Adding up twelve copies of 0.083333 returns 1.000000, the elapsed year, exactly as the line above predicts for a term of order one. Twelve copies of 0.083333 squared come to 0.006944 each and add to 0.083333, the step itself. Refine to 252 steps and that second total falls to 0.003968. Refine to 2,520 steps and it falls to 0.000397. The total is going to zero, and going there at exactly the rate of the step.
The year is cut into twelve steps. What is dt numerically on that grid?
What order is the increment symbol, and what fixes it?
The order of the increment symbol was settled in the previous stretch of this subject and is borrowed as a result. A Brownian increment over a step of length dt has a spread equal to the square root of dt. Square root scaling is the defining scaling of the process, and the strange behaviour of everything downstream comes out of it.
Square root scalingThe rule that the spread of a Brownian increment grows with the square root of the span it covers, rather than with the span itself. has a consequence that trips people every time they meet it. Halving the step halves the time symbol. The square root of a half is about seventy one per cent, so the increment falls only that far. Dividing the step by a hundred divides the time symbol by a hundred. The increment divides only by ten. The increment shrinks, but it shrinks lazily, and the time symbol outruns it. So the gap between the two does not close under refinement. It opens.
| \(\Delta W\) | the change in the Brownian path across one piece of the grid |
| \(\sim\) | is of the order of, meaning one standard deviation in size |
| \(\sqrt{\Delta t}\) | the square root of the step, being the scaling settled earlier |
| \(\sqrt{n}\) | the ratio of a typical increment to the time step, which grows without bound |
The coastline is the one everyday case where refining a measurement is known to make things worse rather than better. Pacing out a stretch of coast with a hundred metre rope gives one length. The same stretch paced with a ten metre rope picks up bends the longer rope skipped, so the total goes up. A one metre rule sends it up again. A straight road does not behave like that: refining the ruler settles the answer. The coast behaves like that because at every scale it has more wiggle waiting, and a Brownian path is the same. There is no scale at which it becomes locally straight, so refining never tames it.
The picture most readers carry has dt as the honest workhorse and dW as a small perturbation sprinkled on top. On the numbers the arrangement is the other way round. At twelve steps the random term is already three and a half times the size of the time term, and every refinement of the grid makes that worse. The random term is not a small correction to the time term; at every grid it is the larger of the two, and it becomes relatively larger the finer the grid gets.
The grid is about to refine from 12 steps to 252. Does the random term become relatively larger or smaller than the time step?
Refine the grid and watch the two markers pull apart
Three markers on one logarithmic scale. The top marker is a typical increment, the middle one is the time step, and the bottom one is that same increment squared. Slide the control and watch which pairs move apart and which pair never separates.
Where does each entry of the multiplication table come from?
Now the table everybody remembers from this subject. Presented as a finished object it looks like three rules to memorise, one of which is bizarre. Built entry by entry it is three applications of the line about order already established, and the bizarre one stops being bizarre.
There are four products of two symbols to consider. Take them in order of how easy they are. The first is the time symbol multiplied by itself. Each such term is of the size of the step squared, so its order is two, and two is above one, so the whole sum is driven to nothing. Numerically at twelve steps that sum is 0.083333, at 252 steps it is 0.003968 and at 2,520 steps it is 0.000397. The sum is going to zero at the rate of the step.
The second product is the time symbol multiplied by the increment symbol. Each such term is of the size of the step multiplied by the square root of the step, or the step raised to one and a half. Order one and a half is above one, so this sum is driven to nothing too, though more slowly than the first: it goes at the rate of the square root of the step. On the twelve steps of the locked path the sum of the sizes of these terms is 0.245374, at 252 steps the same construction gives about 0.053545, and at 2,520 steps about 0.016932. Slower, but still going nowhere but zero.
The third product is the increment symbol multiplied by itself, and this is the entry that changes everything. Each such term is the square of something of the size of the square root of the step, and a square root squared is the thing itself. So each term is of the size of the step, its order is one, and one is exactly the order that survives. The increment squared is not a small correction to the time step; it is the time step, at the same order, term for term.
| \(dt\) | the time symbol, of order one in the step |
| \(dW\) | the Brownian increment symbol, of order one half in the step |
| \((dt)^{2}\) | order two, so its sum is the horizon multiplied by one spare power of the step |
| \(dt\,dW\) | order one and a half, so its sum carries a spare half power of the step |
| \((dW)^{2}\) | order one, so its sum holds still at the elapsed time |
| \(=0\) | means the sum of such terms goes to zero, not that any single term is zero |
The equals signs in that table are doing something unusual. Writing that the time symbol squared equals zero does not say that any actual quantity is zero. The entry says that when terms of that kind are added up over the year and the grid is refined, the total goes to zero. Every entry in the table is a statement about a limit of a sum, in exactly the way the whole notation is. Reading the entries as statements about individual quantities returns to treating differentials as numbers, and treating differentials as numbers is the original error.
Which entry of the multiplication table survives the limit?
Is the surviving entry an approximation or an identity?
An identity in the limit, and the strength of that claim is worth more than the claim itself. A reader who has met the increment squared being of the same order as the step will often settle for a comfortable summary: on average, over many paths, the squared increments add up to about the elapsed time. The comfortable summary is true and far weaker than what actually holds, and the difference between the two decides whether the rest of this subject is possible.
The convergence that holds is almost sureHolding on every path except a collection of paths that together carry probability zero. Stronger than holding when readings are averaged. convergence. As the grid is refined without end, the sum of squared increments settles on the elapsed time for every individual path apart from a collection carrying no probability whatever. Not the average of the sums across many paths. Each path, on its own, taken one at a time.
| \(\mathbb{P}\) | the physical measure, being the rule that assigns probability to sets of paths |
| \(\Delta W_i\) | the change in the Brownian path across the grid interval ending at time \(t_i\) |
| \(t\) | the elapsed time up to the point the sum runs to |
| \(=1\) | the collection of paths where the statement fails carries probability zero |
Why does the weaker version not do? Because everything downstream is written along one path. There is one process, one realised history, one set of movements, and the statement required is what a function of that process did over the year. If the surviving entry held only on average, the entitlement would extend to how a great many paths collectively behave, and to nothing at all about the single path at hand. Almost sure convergence is what lets the surviving entry be substituted inside a calculation about one path, and without it there would be no pathwise statement anywhere in this subject.
Squaring does a second thing the increment itself never does, and that second thing makes squares special. Squaring throws away signs. Every squared increment is positive, so nothing in the sum can cancel anything else, and the total simply piles up. The increments themselves carry signs and cancel constantly. A year of them adds to something modest rather than something enormous. Removing the signs converts a sum that cancels into a sum that accumulates, and that conversion is the whole reason this one entry refuses to disappear.
Is the statement that the increment squared equals the time step an approximation or an identity in the limit?
How is an expression carrying both symbols read?
Mechanically, and that is the good news. Once the order of each symbol is fixed at one and one half respectively, no judgement about whether a term looks small is required. The powers are counted, added, and the total compared with one. Term droppingDiscarding what vanishes when the grid is refined, which is the whole purpose of a multiplication table. stops being a matter of taste and becomes an arithmetic check. Two readers will always reach the same verdict.
The rule reads as follows. Every dt in a term contributes one to the total. Every dW contributes one half. Add up the contributions. If the total is exactly one, the term survives and belongs in the answer. If the total is above one, the term is higher orderShrinking faster than the step itself as the grid refines, and therefore driven to nothing when summed over the horizon. and goes. The addition of exponents is the entire procedure, and it disposes of every product this subject presents.
| \(j\) | how many times the time symbol appears in the term |
| \(k\) | how many times the increment symbol appears in the term |
| \(j+k/2\) | the total order of the term, being the power of the step it behaves like |
| \(=1\) | the only total that holds still once the term is summed over the horizon |
| \(>1\) | any larger total, which carries a spare power of the step and vanishes |
A rule stated without its edge is a trap waiting for the reader. The check just given governs products of two or more differentials, and those products are the correction terms the table exists to sort. A single dW standing on its own has a total order of one half. One half is below one, and the check would seem to say its sum should blow up rather than settle. The sum does not blow up, and the reason is the cancellation named a moment ago: the increments carry signs, and a sum of many signed increments does not grow the way a sum of many same-signed terms would. So the increment on its own is handled by that cancellation, and the exponent check handles everything built by multiplying differentials together. Keeping those two cases separate stops the rule from hardening into folklore.
A term carries the time step to the power one and a half. Does it survive?
What does the locked path give when the table is checked on it?
Everything above is a statement about limits, so it deserves a check against something concrete before it is believed. The locked pathThe twelve step Brownian path published for this reading order, built rather than sampled so that its readings reproduce exactly throughout. is the twelve step path this reading order carries, built from twelve driving values multiplied by the square root of one twelfth. The values are minus 0.5, 1.6, minus 1.3, minus 0.1, 0.1, 1.5, minus 1.3, minus 0.5, minus 1.4, 0.4, 0.9 and 0.6, and they were constructed rather than drawn: they sum to zero exactly and their squares sum to 12.0 exactly.
On the locked path, the squared increments summed over the year. What number?
Each driving value multiplied by 0.288675 gives the twelve increments the year is made of: minus 0.144338, 0.461880, minus 0.375278, minus 0.028868, 0.028868, 0.433013, minus 0.375278, minus 0.144338, minus 0.404145, 0.115470, 0.259808 and 0.173205. The standard processThe single invented traded quantity this reading order runs on, starting at Rs 100/-, with a drift of 8 per cent and a volatility of 20 per cent a year. built on those increments started the year at Rs 100/-, reached Rs 111.08/- in June, fell to Rs 93.74/- in September and finished at Rs 106.18/-. Notice the sizes of the increments themselves. Ten of the twelve are larger than 0.083333, the time step on the same grid. The two that are not are the fourth and fifth, at 0.028868 each, and they are the two smallest driving values in the list.
Each row of the table can now be checked on those numbers rather than on a limit. The results are in the next table, and each column shows something different about how a term behaves under refinement.
| What is being summed over the year | At 12 steps | At 252 steps | At 2,520 steps |
|---|---|---|---|
| The time step on its own | 1.000000 | 1.000000 | 1.000000 |
| The time step squared | 0.083333 | 0.003968 | 0.000397 |
| The time step times the increment, in size | 0.245374 | 0.053545 | 0.016932 |
| The increment squared | 1.000000 | 1.000000 | 1.000000 |
Read the rows against each other. The first row is the elapsed year and it never moves. A term of order one does exactly that. The second row falls by a factor of about ten at each refinement shown, the mark of a spare full power of the step. The third row falls more slowly, by a factor of about three, the mark of a spare half power. And the last row sits at 1.000000 at every grid, identical to the first row. The multiplication table is saying in numbers what the notation says in symbols.
| \(z_i\) | the twelve locked driving values, one for each month of the year |
| \(\sqrt{\Delta t}\) | 0.288675, the square root of one twelfth of a year |
| \(\sum z_i^{2}\) | 12.0 exactly, which was built into the list on purpose |
| \(\Delta t\) | 0.083333, the length of each of the twelve pieces |
Two cautions about that reading, both established in the previous stretch of this subject. The exactness is construction, not luck: a path drawn at random would give something near one rather than one itself, and the driving values were chosen so the illustration reproduces to six decimal places. And a reading at one grid is not the limit. The same locked path read at coarser groupings gives readings that lurch rather than climb. Those lurching readings are covered where quadratic variation is covered.
Look at the treads on that staircase and one further point falls out for free. The second month contributes 0.213333 on its own, more than a fifth of the whole year. The fourth and fifth contribute 0.000833 each and are visually flat. The surviving entry of the table does not say that each month contributes its fair share of time; it says the twelve contributions add to the elapsed year. Where they come from is the path business, and the total is the table business.
The error that gets made, and what it costs
Treating dW as a small number and concluding that its square must be negligible. Treating a small thing that way is the most natural move in the world, and every previous context the reader has met has rewarded it. Square a small thing and it gets much smaller: 0.1 becomes 0.01, and 0.01 becomes 0.0001. The habit is correct wherever quantities are ordinary. Here it is false, and false in a particular direction.
Put the locked path numbers to it. A typical increment at twelve steps is 0.288675, and that certainly looks like a small number. Its square is 0.083333. The square is not somewhere far below the time step. The square is the time step, to the last decimal place shown. The habit predicts something around 0.006944, the square of the time step itself, and the habit is out by a factor of exactly twelve. Refine the grid to 252 steps and the habit is out by a factor of 252. Refine to 2,520 and it is out by a factor of 2,520. The size of the mistake is exactly the number of steps the calculation was careful enough to use, so more care makes the error larger.
The cost is the loss of the only term this subject exists to keep. Every result downstream is the consequence of carrying one term that ordinary calculus discards. At the moment of discarding it the arithmetic offers no resistance whatever. The term genuinely looks small, and dropping small terms is what a reader has been trained for years to do. Nothing flags it. The calculation completes and the answer is clean. Ordinary calculus would have given exactly that answer, and it is precisely the answer that does not apply here.
A typical increment on the locked path is about 0.29. What is its square, roughly, and how does that compare with the time step?
How does somebody checking a model use this?
The useful move is a reading habit rather than a calculation, and it can be applied to somebody else's expression without knowing anything about where it came from. Every expression in this subject is a list of terms, some of which belong in the answer and some of which do not, and the exponent check settles the membership question in a few seconds. Most disagreements about a derivation in this subject turn out to be disagreements about one term, and the check names which one.
The reason it bites is that a term dropped in error is silent. Nothing about the resulting expression looks wrong. The shorter expression is cleaner and agrees with what ordinary calculus would produce, so it reads as a tidier version of the same answer rather than as a different answer. A missing second order term never announces itself. The check has to be applied deliberately rather than trusted to instinct.
- Label the order of every symbol before judging any term
Write one over each dt and one half over each dW. Do it on paper before forming any view about which terms look important, because looking important is the judgement the check exists to replace.
At twelve steps the time step is 0.083333 while a typical increment is 0.288675, so looking small is a poor guide.
- Add the labels within each product
A product of two increments totals one. A product of a step and an increment totals one and a half. Three increments total one and a half as well. The addition is the whole check.
Every verdict in the table came from one addition of two exponents and nothing else.
- Keep total order one, discard anything above it
Exactly one is the survival condition. Above one carries a spare power of the step and disappears under refinement. Below one only arises for a lone increment, which is settled by cancellation instead.
The increment squared totals exactly one, which is why it stays and everything else in the table goes.
- Ask what step length the expression will actually be used at
A grid is chosen somewhere downstream, and the choice does not rescue anything. Refining the grid makes the random term relatively larger, so a term dropped because it looked small at a coarse grid is even more wrongly dropped at a fine one.
The ratio of increment to step goes from 3.464102 at twelve steps to 50.199602 at 2,520.
None of those four steps needs data, software or market access. All four are questions asked of an expression somebody else wrote, and the fourth catches the most common failure. A reader who has convinced themselves that a term is negligible almost always has a coarse grid in mind, and has not noticed that refining makes their case weaker rather than stronger.
Readers of this subject sometimes look for a governing rule. No authority anywhere sets the order of a symbol, no jurisdiction publishes a multiplication table, and no convention alters what a sum does under refinement. The result is a statement about limits rather than about anything traded, so it is the same in every market and in none. No regulator governs this arithmetic.
Where does the table go next? The table is the input to the chain rule that carries the name of Ito. There the surviving entry is put to work on a function of the process and produces the extra term that ordinary calculus does not have. Ito's chain rule comes next in this reading order and is covered separately. The table handed forward has three entries at zero and one entry equal to the time symbol, together with the reason each of those four verdicts holds.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for the scaling of Brownian increments and its use in stochastic integration | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Ito | The integral and the chain rule that carry his name, both of which take the multiplication table as their input | named in the text only |
| Hull, Shreve and Wilmott | Standard texts covering the differential notation and the multiplication table | named in the text only |
The standard process and the locked path are invented.
Educational material. Not advice on any investment, tax, budget or market position.
