Jump Diffusion: Combining Continuous Drift With Sudden Moves
A diffusion model moves continuously, carrying a drift and a volatility and nothing else. A jump diffusion model keeps that continuous part and adds sudden discontinuous moves on top of it. The two parts then share one total variance between them, and the jump part supplies something no level of continuous volatility can produce: a move too large to have been reached one small step at a time.
Two consequences follow from bolting a jump part onto a continuous one, and the whole of this guide is an argument for keeping them apart. The first is that variance goes up. Arrivals of a typical size do not average to nothing, so the average goes somewhere too. Adding jumps adds variance and it also adds an average effect, and the drift has to be moved to cancel the second so that only the first is genuinely new. A reader who holds those two consequences separate will get everything else in this guide for free. A reader who lets them merge will build a model whose average growth was chosen by accident.
A picture makes the distinction concrete. A rain gauge in the monsoon collects in two quite different ways. Drizzle raises the level continuously: at any instant a little more water is in the tube than an instant earlier, and a photograph of the tube every second would trace the level rising without the pencil ever lifting. Then a cloudburst arrives and puts an inch in at once. The inch did not pass through any intermediate level in any meaningful sense. Over a season, both mechanisms are at work in the same tube, they contribute different amounts to how much the level varies, and one of them cannot be imitated by simply making the drizzle heavier. A jump diffusion is that tube written down properly.
Everything below uses one invented quantity, called the standard process and written \(S_t\). The standard process starts at Rs 100/-, has a drift of 8 per cent a year and a volatility of 20 per cent a year, and is observed over one year. Where jumps are added they arrive at 0.5 a year and their proportional size has a logarithm with an average of minus 0.05 and a standard deviation of 0.10.
What is a diffusion model, defined from scratch?
A diffusion modelA model whose movement is continuous, with a drift and a volatility and nothing else in it. says that over the next instant the level changes by two things added together, and by nothing else at all. The first is a steady proportional pull, the drift. The second is a random shove whose size is set by the volatility. Both are infinitesimal over an infinitesimal instant, and an infinitesimal step is exactly why the path that results joins up: a pencil can be put on it at the start of the year and traced to the end without lifting.
| \(S_t\) | the level of the standard process at time \(t\), in rupees, starting at Rs 100/- |
| \(\mu\) | the drift, as a decimal a year. Here 0.08 |
| \(\sigma\) | the volatility, as a decimal a year. Here 0.20 |
| \(dt\) | an instant of time, shrunk toward nothing |
| \(dW_t\) | the increment of standard Brownian motion under the physical measure P over that instant |
Because there is no third term, the whole year is the sum of the two. The sum is what gets observed and the parts are what get specified. Readers routinely mistake one for the other, and the split is worth seeing literally rather than taking on trust. The variance rateThe rate at which variance piles up per unit of time. A volatility of 0.20 a year corresponds to a variance rate of 0.04 a year. of this model is 0.04 a year exactly, being the volatility of 0.20 squared, and it is the only place randomness enters.
| \(S_0,\;S_T\) | the level at the start and at the horizon. Here Rs 100/- and whatever the year produces |
| \(T\) | the horizon in years. Here 1.0 throughout |
| \(W_T\) | Brownian motion at the horizon, which has an average of nought and a variance of \(T\) |
| \(\tfrac{1}{2}\sigma^{2}\) | the half variance correction, 0.02 exactly here, which separates the average from the middle outcome |
Notice what the middle panel is not doing. The path never breaks. Between any two instants there is another instant, and the level at that instant lies between the two neighbouring levels. The unbroken path is not a drawing convention, it is the definition. A diffusion cannot get from one level to another without passing through every level in between, and that single property is what the rest of this guide removes.
What is a jump diffusion model, defined from scratch?
A jump diffusion modelA diffusion model with sudden discontinuous moves added on top, so the path has both continuous movement and instant breaks. keeps both terms above and adds a third. At random moments an arrival happens, and when it does the level is multiplied by a random factor. Between arrivals the model is the diffusion, unchanged and unbothered. At an arrival the level takes a step that no amount of elapsed time was needed to produce. Three things now have to be written down rather than two: how the continuous part behaves, how often arrivals come, and how large the multiplier is when one does.
| \(S_{t^-}\) | the level an instant before time \(t\), which is the level an arrival acts on |
| \(\mu\) | the intended average growth rate, 0.08 a year here |
| \(\lambda\) | the arrival intensity, 0.5 a year here |
| \(J\) | the proportional multiplier at an arrival. Its logarithm has an average of \(m=-0.05\) and a standard deviation of \(s=0.10\) |
| \(k\) | the average of \(J-1\), which is the average proportional effect of one arrival |
| \(N^{J}_{t}\) | the counting process, tallying arrivals up to time \(t\). Written with a superscript so it is never read as the normal distribution function |
| \(\lambda k\) | the compensating term subtracted from the drift, computed in full below |
The frame is named after Merton, whose 1976 paper set out this combination, and the notation and ordering here follow the standard texts by Hull, Shreve and Wilmott. The third term is switched off almost all the time. With an intensity of 0.5 a year the chance of no arrival at all across the whole year is 0.606531, so on most years drawn from this model the third term never fires once and the path is indistinguishable from a diffusion.
| \(Y_i\) | the logarithm of the multiplier at the \(i\)th arrival, drawn with an average of \(m\) and a standard deviation of \(s\) |
| \(N^{J}_{T}\) | how many arrivals happened by the horizon. Nought with chance 0.606531, one with chance 0.303265, two with chance 0.075816 |
| \(m,\;s\) | minus 0.05 and 0.10 here, both about the logarithm of the multiplier rather than the multiplier itself |
| \(\sigma W_T\) | the continuous contribution to the year, with a standard deviation of 0.20 |
The shaded strip carries the sharpest version of the point. The continuous part moves a typical distance of \(\sigma\sqrt{\Delta t}\) over a stretch of length \(\Delta t\). To cover 0.05 in logarithmic terms it needs \(\Delta t\) equal to 0.0625 of a year, or 22.8125 days. The arrival covers the same 0.05 in an interval of no length whatever. The two mechanisms are not doing the same work faster or slower; one of them is doing work that has no duration attached to it at all.
Arrivals are about to be added to the standard process at an intensity of 0.5 a year. Does the total volatility rise above 20 per cent by a lot or by a little?
How do the two parts share the total variance between them?
The two parts share it by addition, and the addition happens at the level of variance rather than at the level of volatility. Addition at the level of variance is the arithmetic spine of this guide. The Brownian increments and the arrivals are independent of one another by construction, so the continuous part contributes its variance rate, the jump part contributes its own, and the two simply sum.
| \(\sigma^{2}\) | the continuous variance rate a year. Here 0.040000, being 0.20 squared |
| \(m^{2}+s^{2}\) | the average of the squared jump size, which is 0.0025 plus 0.0100, being 0.012500 |
| \(\lambda\left(m^{2}+s^{2}\right)\) | the jump contribution to the variance rate. Here 0.006250 a year |
| \(\sigma_{\text{total}}\) | the volatility of the combined model a year. Here 0.215058, or 21.5058 per cent |
The locked numbers carry it through, and the whole of this guide hangs on this build reconciling. The jump part contributes 0.5 times 0.012500, or 0.006250 a year. Added to 0.040000, the total variance rate is 0.046250 a year, and the square root of that is 0.2150581, or 21.5058 per cent. The variance splitThe share of the total variance contributed by each part of the model, here 86.4865 per cent continuous and 13.5135 per cent jumps. is 86.4865 per cent continuous and 13.5135 per cent jumps, and that pair of shares is the single number worth carrying away from any jump model.
| Contribution to the variance rate | Amount a year | Share of the total |
|---|---|---|
| Continuous part, being 0.20 squared | 0.040000 | 86.4865 per cent |
| Jump part, being 0.5 times 0.012500 | 0.006250 | 13.5135 per cent |
| Total variance rate | 0.046250 | 100.0000 per cent |
| Total volatility, being the square root | 21.5058 per cent | against 20 per cent |
The red marker is worth dwelling on. The jump part, taken on its own, has a volatility of 7.9057 per cent a year, being the square root of 0.006250. Somebody who reasons that 20 per cent of continuous volatility plus 7.9057 per cent of jump volatility gives 27.9057 per cent has made an error of 6.3999 percentage points on a quantity of about twenty. Volatility is a square root and square roots do not add, so volatilities cannot be added under any circumstances. The rain gauge again: doubling the rainfall does not double the height of the tube's variability in the way a first guess suggests, and the arithmetic that fixes it is the same arithmetic here.
At an intensity of 0.5 arrivals a year, what share of the total variance is contributed by the jumps?
The jump intensity is about to rise from 0.5 to 2.0 a year. Ahead of moving the control below: does the continuous part's contribution to the variance change?
Move the intensity and watch the split move while the continuous amount refuses to
The control changes how often arrivals come. Three things redraw: the stacked variance bar, the marker on the volatility scale and the bar for the drift raise. The dark segment of the top bar is the continuous contribution, nailed at 0.040000 and never moving at any setting. The dashed green marks show where the worked example in this guide sits, so the distance travelled from it stays visible.
At a setting of 2.0 the top bar tells the story better than the numbers. The dark segment has not moved a pixel. Nothing done to the arrivals took anything away from the continuous part, so the dark segment is still 0.040000, exactly what it was. The length of the whole bar changed, and therefore the fraction of it the dark segment represents, down to 61.5385 per cent. A falling share is not a falling amount, and treating the two as the same thing is the most common reading error made against a variance split.
What does the jump part add that continuous volatility cannot produce at any level?
The obvious objection deserves a straight answer. If the jump part only adds 0.006250 to the variance rate, why not simply raise the continuous volatility from 20 per cent to 21.5058 per cent and be done with it? Raising the continuous volatility would land on the identical total variance with one fewer mechanism and three fewer parameters. The answer is that it would land on the identical total variance and a different shape, and the difference lives exactly where it matters.
A large moveA move too big to have been reached one small step at a time, so it can only arrive at once. is the thing at stake. Raising the continuous volatility widens the whole distribution: outcomes near the middle spread out, and so do outcomes far away, all in the same proportion. Adding jumps does something different. A year with one arrival in it is a year whose outcome was shifted bodily by an amount that has nothing to do with how the Brownian part behaved. Adding jumps therefore leaves the middle very nearly alone and puts extra weight specifically far from the centre.
| excess kurtosis | how much more weight sits far from the centre than a normal shape would put there. Nought for any pure diffusion, 0.106647 here |
| skewness | how lopsided the shape is. Nought for any pure diffusion, minus 0.081698 here |
| \(m^{4}+6m^{2}s^{2}+3s^{4}\) | the average of the fourth power of the jump size, being 0.00045625 with these numbers |
| \(m^{3}+3ms^{2}\) | the average of the cube of the jump size, being minus 0.001625 with these numbers |
The place to look is the far part of the distribution, and the honest way to look is to fix a threshold and ask how likely a move past it is. Raising the volatility lifts every threshold together. The jumps lift the distant ones far more than the near ones. Fixing the threshold in absolute terms rather than in standard deviations is what separates the two effects.
Read the two markers together. Near the middle, at a threshold of 0.2, the jump model gives 0.115425280 and the matched diffusion gives 0.116152072, a ratio of 0.9937, so the jump model is very slightly the tamer of the two. Far out, at a threshold of 0.8, the same two give 0.000138309 and 0.000033827, a ratio of 4.0887. Matching on variance makes the two models agree near the centre and disagree by a widening factor far from it, and that far region is exactly the part no setting of continuous volatility can reach without dragging the centre along with it.
The grey dashed curve is worth a glance too. The curve is the original 20 per cent diffusion, and it sits below both at every threshold. Sitting lower is simply the statement that it has less variance. The curve does not change shape. Move it up by raising the volatility and it slides bodily upward; it never bends into the pine curve. Volatility is a scale knob and jumps are a shape knob, and turning a scale knob will never do the work of a shape knob.
Can a high enough continuous volatility reproduce what the jump part does?
Why does the drift have to be adjusted, and by exactly how much?
Because the arrivals do not average to nothing. The drift adjustment is the step that gets skipped, and it gets skipped for a reason worth naming: the specification of the jump size looks symmetric. Its logarithm has an average of minus 0.05 and a standard deviation of 0.10, and a reader who glances at that sees a small tilt and moves on. But the model multiplies by the size, and the average of a multiplier is not the exponential of the average of its logarithm.
| \(J\) | the multiplier at an arrival, whose logarithm has an average of \(m\) and a standard deviation of \(s\) |
| \(m+\tfrac{1}{2}s^{2}\) | minus 0.05 plus 0.005, being minus 0.045 with these numbers |
| \(k\) | the average jump effectThe proportional change one arrival produces on average, which is not the same as the average of its logarithm., being minus 0.044003 here |
| \(\lambda k\) | the average effect a year, being 0.5 times minus 0.044003, which is minus 0.022001 |
Now the everyday version, and it is the sharpest one available. An empty steel vessel goes on a kitchen scale to weigh rice. Before the rice is poured, the tare button is pressed so the scale reads nought with the vessel already on it. Without that press every reading afterwards is heavy by the weight of the vessel, and nothing on the dial says so: the number looks like a perfectly ordinary weight. The compensatorThe adjustment that removes the average effect of the jumps from the drift, leaving the variance changed and the average untouched. is the tare button of a jump model, and \(\lambda k\) is exactly the weight of the vessel.
| \(\mu_{0}\) | the drift as written down, without any adjustment. Here 0.08, the intended 8 per cent |
| \(\lambda k\) | minus 0.022001 a year, the average effect of the arrivals |
| \(\mu_{0}+\lambda k\) | 0.057999 a year, which is what the model actually delivers |
| \(T\) | the horizon in years, 1.0 here |
By how much must the drift be raised at an intensity of 0.5 arrivals a year?
Somebody changes the average logarithm of the jump size from minus 0.05 to nought, leaving the intensity at 0.5 and the spread at 0.10. What happens to the required drift adjustment?
The error that gets made, and what it costs
Adding jumps and leaving the drift exactly as it was.
The jump specification here has a downward average, so switching the arrivals on without the compensator lowers the average growth of the standard process by 0.022001 a year, taking it from 8 per cent to 5.7999 per cent. Nobody decided that. The fall is not a modelling judgement anybody made and defended; it is a consequence of a number chosen for a completely different reason, namely how large a typical arrival is.
The asymmetry in what gets noticed is what makes the failure durable. The variance change is visible, it is what everyone came to discuss, and it appears as a line in the specification. The average change is invisible: no line moved, no parameter was retyped, and the growth rate in the specification still says 8 per cent. So the first change gets reviewed and the second one does not, and it survives every check that consists of reading the specification.
One tell is the only cheap way to catch a missing compensator from the outside, and it is worth memorising. In a compensated model, changing the typical jump size moves the shape and leaves the average growth where it was. In an uncompensated one, changing the typical jump size moves the average growth, behaviour that nothing about the intention of the model would ever suggest. If perturbing the jump size makes the average growth follow it, the compensator is missing.
The cost has a specific shape. The model that remains has an average growth set by an assumption about arrival sizes, and every quantity that depends on average growth has inherited that assumption. The variance discussion was the honest half of the change, and it is the half that got all the attention.
Jumps with a downward average are added and the drift is left alone. What happens?
What does adding jumps cost, and what does it rule out?
Three things, and each is worth being blunt about. A jump model is routinely presented as a strictly better model rather than as a purchase with a price attached.
The first cost is the closed formA formula that gives an answer directly rather than by stepping through a calculation numerically, and jumps usually take it away.. The continuous model answers a good many questions with a formula that evaluates in one line, and once arrivals enter, most of those answers become sums over how many arrivals occurred, or numerical integrations, or lattice work. The questions that a continuous model settles directly about, say, the at-the-money contract at Rs 100/- or the out-of-the-money contract at Rs 110/-, now mostly have to be computed rather than written down. How those contracts work, and what either of them is worth, is covered separately; the point is only that the route to an answer got longer.
The second cost is parameters. A diffusion needs two numbers. A jump diffusion needs five: the drift, the continuous volatility, the intensity, the average jump size and its spread. Going from two to five is not a small change in how much the model can be pinned down by whatever it is fitted to.
The third cost is the one people notice last, and it follows from the second. The variance split is not identified by a single volatility number. Suppose all that is known is the total variance rate, 0.046250. Then an intensity of 0.5 with an average squared jump size of 0.012500 fits perfectly, and so does an intensity of 2.0 with an average squared jump size of 0.003125, for instance an average logarithm of minus 0.025 and a spread of 0.05. Those two settings agree exactly on the total variance and disagree by a factor of four on excess kurtosis, 0.106647 against 0.026662, which is to say they disagree about the whole reason the jumps were added in the first place.
What is the main thing adding jumps costs?
Criterion by criterion, where do the two models actually differ?
Diffusion model vs jump diffusion model, on nine criteria
Both models have now been defined in their own right. Definition on both sides is what makes a comparison worth making rather than a way of introducing the second one. Set them against the same criteria and they part company on every single one, and the variance split is the quantitative anchor that holds the whole comparison together.
Two rows in that grid deserve a second look because they are the ones that carry consequences beyond arithmetic. The row about the path joining up is the reason every rule that watches for a level changes meaning under the second model: a continuous path cannot pass a level without being at it, and an arrival can. The row about the drift adjustment is the one this guide has spent the most time on, and it is the only row where the second model requires something to be done rather than simply behaving differently. Every other row describes a difference in what the model does; that row describes a difference in what somebody must remember to do.
How does somebody reading another person's model use this?
A jump model is met far more often than it is built. Somebody hands over a note, a memorandum or a set of parameters, and the useful question is not whether the model is right but whether it says what its author thinks it says. Four checks answer that, in ascending order of how much access they need.
- What is the total variance rate, and how does it split?
Take the continuous volatility, square it, and put it beside the intensity times the average squared jump size. Those two numbers and their shares are the whole quantitative story of the model.
If the note quotes only one volatility number, it has not said which model it is describing.
- Does the drift carry the compensator?
Look for the intensity times the average jump effect sitting inside the drift term. If the drift is a bare number with nothing subtracted from it, ask what the model's average growth actually is rather than what it says it is.
With these parameters the answer is 5.7999 per cent where the note says 8 per cent.
- What was matched when the parameters were chosen?
If the whole specification was pinned to one total volatility figure, the split between continuous and jump variance was not determined by anything. Two settings that agree on 0.046250 can differ fourfold on excess kurtosis.
Ask what number, other than the total, decided the intensity. If there is no answer, the split is an assumption wearing a fitted parameter's clothes.
- Which answers are now numerical rather than direct?
Every result the note states should be traceable to either a formula or a computation. A jump model that produces its answers as fast as a continuous one has either been approximated somewhere or has not actually switched the arrivals on.
Ask where the sum over arrival counts is. It has to be somewhere.
The second check is the valuable one, and it costs nothing at all to run. The check needs no data, no code and no fitted parameters, only a careful reading of the drift term and a look at whether anything has been taken out of it. Most of the value of understanding the compensator lies in reading other people's models rather than in writing one's own.
A jump diffusion does not depend on location, on holdings or on which rules apply to whom. No jurisdiction sets the definition of a process, so no local rule bears on it. The variance rate adds the same way everywhere, the compensator is the same subtraction everywhere, and 0.022001 is 0.022001 in every country and every decade.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for jump diffusion models and variance decomposition | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Merton, 1976 | The paper that set out the jump diffusion combination | Journal of Financial Economics, 1976 |
| Hull, Shreve and Wilmott | Standard texts on derivatives and stochastic calculus | Pearson, Springer and Wiley |
The standard process is invented.
Educational material. Not advice on any investment, tax, budget or market position.
