Optional Stopping: Why Choosing When to Quit Does Not Help
The result associated with Doob says that a fair game stopped at a rule that could actually have been followed carries the same average as the same game left alone. Choosing the moment does not help, so long as the rule must fire by a fixed date or the game is capped in size. When neither cap is present the result fails, and the shape of that failure is the most instructive part of it.
The oddity of that is worth feeling before it is made precise. Watching is allowed. Waiting as long as the period permits is allowed. Quitting on any reading whatever is allowed, and everything seen so far may be used in deciding. The freedom on offer is enormous, and it buys nothing at all in the average. Not a little; nothing.
The reason sits in one sentence, and everything that follows is that sentence made checkable. A rule that cannot see ahead is standing, at every moment, in front of a set of futures that all carry the same average. Because they are all worth the same, picking one of them cannot change the average. Choosing among things that are worth the same is not a strategy, it is a preference, and preferences do not move averages.
The everyday version is worth keeping beside the mathematics. A sack of rice is being weighed on an old spring scale whose needle will not settle. The needle wanders either side of the true weight, and it does so without favour, drifting up as readily as down. One number has to be written in the register, and it may be written at any moment. The needle is the fair game. The moment chosen is the rule. The result says the average of what is written down is the true weight, whatever waiting policy is adopted, and the rest of the argument is about the exact circumstances under which somebody could claim otherwise.
What does the optional stopping result say?
One version of the result needs fewer conditions than the rest, carries the intuition, and can actually be established in practice, so it comes first. Let the process be a fair game, and let the rule for quitting be followable, meaning that at every moment it can already be said whether the rule has fired. One extra thing is then insisted on: the rule must have fired by some fixed date, come what may. Under that arrangement the average of the value at the quitting moment is exactly the value the process started from.
The extra insistence is what makes this the easy case. A rule that must fire by a fixed date is called a bounded stopping timeA quitting rule that must have fired by some fixed date, whatever the path does. The easiest case of the result to establish., and one can always be manufactured out of any rule at all by adding the clause or the end of the period, whichever comes first. The added clause is not a technicality tucked away in a footnote. The clause is the single most common way the result is made to apply to something real, and it returns three more times below.
| \(X_t\) | a fair game, meaning a process whose average future value given what is known now is its present value |
| \(\tau\) | the moment the quitting rule fires, differing from path to path and settled by what is known at the time |
| \(T\) | a fixed date by which the rule must have fired, one year in the worked instance |
| \(X_{\tau}\) | the value of the process read at that moment, itself a random quantity |
| \(\mathbb{E}[\cdot]\) | the average taken across the whole set of paths, not along any one of them |
Two things in that statement are doing more work than they look. The first is the word average. The claim is about the whole set of paths taken together, and it says nothing whatever about any single path. On one path the rule quits ahead. On another it quits behind. The result is silent on which, and a reader who checks it by running one path has misunderstood what kind of object it is.
The second is the word fair. The result is a statement about a fair game and only about a fair game. Given a process that drifts upwards it says something different, and correctly so: for a process that tends to rise, quitting early on average gets less than waiting. Instinct already says exactly that. The interesting case, and the one that trips people, is the fair one, where instinct says a clever quitting rule ought to be able to extract something and the mathematics says it cannot.
Why does choosing the moment not help?
At any moment before the rule has fired there are exactly two things available: quit now, or carry on and decide later. The two are worth comparing.
Quitting now takes the reading in front. Carrying on takes some reading further along the path, arrived at by continuing to apply the same rule. Because the game is fair, the average of everything that lies ahead of the moment in question is the reading at that moment. So carrying on has the same average as quitting. Not approximately, and not on average over many attempts in some loose sense: the two have the same average from wherever the comparison is made, at every moment, on every path.
The comparison between quitting and carrying on is available at every moment, it always comes out level, and a decision that is level at every moment cannot accumulate into an advantage over the whole period. The formal machinery makes this exact by noting that the stopped process, the one that follows the game until the rule fires and then holds still, is itself a fair game. Once that is in hand, the result is a one line consequence: a fair game read at a fixed date has the average it started with, and the stopped process read at the horizon is precisely the value at the quitting moment.
| \(X^{\tau}_t\) | the stopped process, following the game until the rule fires and then holding the value it had at that moment |
| \(t \wedge \tau\) | whichever comes first, the present moment or the quitting moment |
| \(\mathcal{F}_s\) | the information available at time \(s\), meaning everything that has happened by then |
| \(\mathbb{E}[\,\cdot \mid \mathcal{F}_s]\) | the average taken over what is still unknown, holding fixed everything already seen |
The argument just given is the whole proof in disguise. The rule was never analysed. Nothing was asked about whether it was clever, whether it fired early or late, whether it used one reading or all of them. The argument showed only that applying the rule produces another fair game, and fair games all behave the same way at a fixed date. Any rule at all disappears into that argument. The result therefore covers rules nobody has thought of yet.
Why can a rule that cannot see ahead not improve the average?
What does the result do to the locked path with two levels?
Now the worked instance, and it is small enough to check by hand. The fair game is the locked Brownian path, written W with a time subscript. The path starts at zero and is a fair game under the physical measure P. The locked path is the same path that drives the standard process, so month six of the path reads 0.375278 while the standard process reads Rs 111.08/-, and month nine reads minus 0.548483 while the standard process reads Rs 93.74/-.
The rule is a pair of levels. Put a barrierA level that, when reached, makes the rule fire. Two of them here, one above the start and one below. at plus 0.5 and another at minus 0.5, and quit the first moment either one is reached. The pair of levels is a followable rule: at every moment the current position is known and either a level has been reached or it has not, and no reading from later is consulted. So the result applies, provided the conditions hold, and those conditions are checked rather than assumed.
The answer can be worked out directly, without the result, to give something to compare against. Two questions have to be settled: how often does each level come first, and what reading it gives when it does. Reaching a level means standing on it, so the second question is easy. The first is settled by symmetryThe property that a driftless path is as likely to move up as down, so two levels the same distance away are reached first equally often.: a driftless path is as inclined to go up as down, and the two levels are the same distance away, so each is reached first half the time.
| \(W_t\) | standard Brownian motion under the physical measure, starting at zero, which is the fair game here |
| \(a\) | the distance from the start up to the upper level, 0.5 in the worked instance |
| \(b\) | the distance from the start down to the lower level, also 0.5 in the worked instance |
| \(\tau\) | the first moment either level is reached |
| \(W_{\tau}\) | the reading at that moment, either plus \(a\) or minus \(b\) and nothing in between |
Put the numbers in. Both distances are 0.5, so the chance of the upper level first is 0.5 divided by 1.0, or 0.500000, and the chance of the lower level first is the same. Half of plus 0.5 is plus 0.250000, half of minus 0.5 is minus 0.250000, and the two sum to exactly zero. The average reading at the quitting moment is 0.000000, precisely the reading the path started from, so the elaborate freedom to choose the moment has produced no movement at all in the average.
Seeing a single outcome is what makes the difference between a result and a slogan, so the locked path itself can now be run through the rule. Read at its twelve month ends, the path goes to minus 0.144338, plus 0.317543, minus 0.057735, minus 0.086603, minus 0.057735, plus 0.375278, back to zero, minus 0.144338, and then at month nine it stands at minus 0.548483. Month nine gives the first reading at or beyond either level, so the rule fires there and the reading taken is a loss.
One honest wrinkle, and it is worth a paragraph because it is where the arithmetic and the picture part company. Read monthly, the path is at minus 0.144338 at month eight and minus 0.548483 at month nine, so it is already past the lower level by 0.048483 when it is first seen there. In continuous time the path touched minus 0.500000 exactly, somewhere inside that month, and quit there. The two versions are different rules with different readings, and the exact arithmetic of half of plus 0.5 plus half of minus 0.5 belongs to the continuous one.
A rule capped at month twelve is a rule that must fire by a fixed date, and that is the only condition the easy case asks for, so the result itself survives the monthly version untouched. What does not survive is the tidy pair of readings. Monthly checking gives the level plus whatever overshoot the month happened to carry, the two possible readings are no longer exactly plus 0.5 and minus 0.5, and their chances are no longer exactly half and half. The two readings still average to zero. They just do it with untidier numbers.
So on this one path the rule quits at a loss. Does that count against the result? Not in the slightest, and understanding why is the difference between reading the statement and having it. A single path produces a single number. The result is a claim about the average across every path there is, and one number can no more contradict an average than one reading of a wobbling needle can contradict the weight of the sack. Had the locked path gone the other way and quit at plus 0.5, that would have supported the result exactly as little.
The locked path quits at minus 0.548483, a loss. Does that contradict the result?
What does the waiting actually cost?
There is a second quantity in this problem that nobody asks about and that turns out to carry the whole argument later on: how long the rule takes to fire. The waiting time comes out of a companion fair game. The squared path less the elapsed time is itself a fair game. Put another way, a driftless path spreads out at a rate of exactly one unit of squared distance per unit of time. Read that at the quitting moment and the average waiting time falls straight out.
| \(\tau\) | the first moment either level is reached, measured in years |
| \(W_{\tau}^{2}\) | the squared reading at that moment, either \(a^2\) or \(b^2\) |
| \(a,\,b\) | the distances up to the upper level and down to the lower one |
| \(a\,b\) | the product of the two distances, giving the average waiting time in years |
With both levels at 0.5 the product is 0.25 of a year, or three months. The locked path took nine months, longer than average and entirely unremarkable, in the same way that one household waiting twenty minutes for a bus says nothing about a ten minute average headway. Move the upper level out to 1.5 and the average wait becomes 0.75 of a year, or nine months, so the larger prize is paid for in time rather than in probability. Hold that thought. The rule that appears to beat the result is nothing but this observation pushed to its limit.
The two quantities are locked together more tightly than they look. The chance of the upper level and the size of the upper level move in exact opposition: doubling the distance halves the weight, near enough, and exactly enough that the product never budges. Nothing chose that. The lock falls out of the arithmetic of the two distances, and the simulation below moves the level in both directions while the product refuses to shift.
The upper level is about to move from 0.5 out to 1.5, with the lower one fixed. Before the control is moved: what happens to the average at the quitting moment?
Move the upper level and watch the average refuse to move
The lower level stays at minus 0.5. Only the upper one moves, from 0.2 out to 2.0. Three things redraw: where the levels sit, the split between the two chances, and the two weighted contributions. The two contributions mirror each other exactly at every setting. Three settings are worth stopping at. At 0.2 the chances are 0.714286 and 0.285714. At 0.5 they are 0.500000 and 0.500000, the worked instance, where the standard process behind the same path reads Rs 111.08/- at its high and Rs 93.74/- at its low. At 1.5 they are 0.250000 and 0.750000. The average at the quitting moment is exactly 0.000000 at every one of them.
Sweep the control from one end to the other and watch which readouts move. The chance of the upper level moves a long way, from 0.714286 down to 0.200000. The two contributions grow from 0.142857 each way at the near setting to 0.400000 each way at the far one. The average wait runs from 1.2 months to 12 months, a factor of ten. The one readout that never moves by so much as a digit is the average at the quitting moment, reading 0.000000 at all one hundred and eighty one settings of the control.
The sweep shows the result in a form that can be felt rather than parsed. A more distant level pays more when it arrives and arrives less often, and those two effects are not roughly equal and opposite, they are exactly equal and opposite, by arithmetic rather than by luck.
The result needs conditions. Before they are read: what should they be expected to rule out?
What three conditions does the result need, and why each one?
The statement so far has been carrying an assumption quietly, and it is time to bring it out. Everything above needed the rule to be capped in time. There are two other ways of buying the same conclusion, and in practice whichever of the three is easiest on the problem at hand is the one established. Any one of them on its own is enough. Two are never needed.
The first condition is a cap on time: the rule must have fired by a fixed date. The time cap is the easiest of the three to establish: it can usually be imposed rather than proved, by adding the clause or the end of the period to any rule at all. The cap rules out a rule that quietly relies on being allowed to wait indefinitely for a reading it likes. If the rule has to answer by month twelve, that avenue closes.
The second condition is a cap on size: the game itself must stay inside a bounded range up to the moment the rule fires. Between the levels the path is confined to the band from minus 0.5 to plus 0.5 by construction, and the moment it leaves the band the rule has fired, so the two level rule satisfies this without any help. The size cap rules out a rule that funds a small frequent gain out of a rare loss of unlimited size. If the loss cannot be unlimited, that funding is not available.
The third condition is the technical one, called uniform integrabilityA condition on a set of quantities saying that no vanishingly rare but enormous outcome carries a meaningful share of the average. It stops the average escaping in the limit., and it is what the first two are special cases of. The honest plain words version is this: no vanishingly rare but enormous outcome may carry a meaningful share of the average. The technical condition exists because the first two are sufficient rather than necessary, and there are rules that satisfy neither yet still behave. Uniform integrability is stated last because it is the hardest to check and the one reached for only when the other two have failed.
Name one condition that is enough on its own.
Does quitting only at a gain beat the result?
The objection everybody has deserves to be worked through rather than waved away. The rule below is a real object with a real reason for failing, and the reason is instructive. That rule is a mathematical object under stated assumptions, and its failure is a failure of those assumptions rather than of arithmetic.
The rule is this. Remove the lower level entirely. Quit the first moment the path reaches plus 0.5, and otherwise never quit at all. At every moment it can be said whether plus 0.5 has been reached, so the rule is a perfectly good followable one. And a driftless path has a property called recurrenceThe property of a driftless path that it returns to every level eventually, if it is given unlimited time to do so.: given unlimited time it reaches every level eventually, with certainty. So the rule fires on every path, and when it fires the reading is plus 0.5 exactly. The average at the quitting moment is therefore plus 0.500000, not zero, and the result appears to have been beaten.
The result has not been beaten. Its conditions have failed, and all three of them at once. There is no cap on time: the rule may run for any length. There is no cap on size: before it fires the path may be arbitrarily far below where it started. And the technical condition fails too, the same fact stated in the form that makes it precise.
| \(\tau_a\) | the first moment the path reaches the level \(a\), with no lower level and no end date |
| \(a\) | the level waited for, 0.5 in the worked instance |
| \(\mathbb{P}(\tau_a < \infty)=1\) | the rule does fire eventually on every path, and that is what makes the objection feel solid |
| \(\mathbb{E}[\tau_a] = \infty\) | the average time it takes to fire is infinite, and that is what makes the objection fail |
Putting a floor back under it is the part that settles the argument. With the target kept at plus 0.5 and a lower level added, then pushed further and further away, the chance of reaching the target first climbs towards certainty, exactly as the objection expects. And the loss on the other branch deepens at precisely the rate that keeps the two contributions equal. At every finite depth of floor, however deep, the average at the quitting moment is exactly 0.000000, and the plus 0.5 only appears in the limit where the floor has been removed altogether.
| Floor under the rule | Chance of reaching plus 0.5 first | Gain side contributes | Loss side contributes | Average | Average wait |
|---|---|---|---|---|---|
| minus 0.5 | 0.500000 | 0.250000 | 0.250000 | 0.000000 | 0.25 years |
| minus 2 | 0.800000 | 0.400000 | 0.400000 | 0.000000 | 1 year |
| minus 5 | 0.909091 | 0.454545 | 0.454545 | 0.000000 | 2.5 years |
| minus 20 | 0.975610 | 0.487805 | 0.487805 | 0.000000 | 10 years |
| minus 100 | 0.995025 | 0.497512 | 0.497512 | 0.000000 | 50 years |
| no floor at all | 1.000000 | 0.500000 | none | 0.500000 | infinite |
Read the last two columns down and the trade becomes impossible to miss. The average stays at zero on every row where the floor exists, and the average wait multiplies: three months, one year, two and a half years, ten years, fifty years. The final row is not the end of a trend, it is a different kind of object, arrived at only by removing the floor entirely and accepting an infinite average wait. Every real version of the rule is one of the rows above it, and every one of those has an average of zero.
Quit only at a gain, never at a loss. Does that beat the result?
What does that rule report on the locked path?
Abstract failure is easy to nod along with and hard to remember. So the failing rule goes onto the one path used throughout, to see what it actually produces over the year.
The rule waits for plus 0.5. The locked path never gets there. Its highest reading of the year is 0.375278, at month six, leaving it short by 0.124722 and never closer. So across twelve months the rule fires exactly never. The rule does not report a gain of 0.5. It does not report a loss. Because it has not finished, it reports nothing at all, and what exists at the end of the year is an open positionA position that has not been closed, so what it will eventually come to is not yet a number. with no date attached to it.
The honest accounting stops there, and the whole cost of the failing rule is stated as a fact rather than as a caution: an average of plus 0.5 that is only ever realisedTurned from an open position into a settled figure that can be written down. at a moment nobody can name is not an average of plus 0.5 that anybody has. The plus 0.5 is a number attached to an event whose date is unbounded, and the accounting only closes when the date arrives.
Compare the two rules on the same path, the fairest test available. The two level rule fired at month nine and produced a number, and the number was a loss. The wait for a gain rule produced no number and no loss. The absence sounds better until one notices that it also produced no ending. One rule is finished and one is not, and a rule that is not finished has not avoided anything, it has postponed it.
On the locked path, what does the rule that waits for plus 0.5 report at the end of the year?
The error that gets made, and what it costs
Treating the rule that waits for a gain as a way of turning a fair game into a favourable one. The arithmetic that supports it is not wrong. In the idealised setting the rule really does fire on every path and really does produce plus 0.5 every time, and anybody who checks that step will find it correct.
The error is one level up, in what the setting was allowed to assume. The plus 0.5 is bought with two things that never both exist: unlimited time, and the capacity to be behind by any amount for as long as it takes. Cap either one and the average returns to exactly zero, as the table above shows at every finite depth of floor. The rule does not survive contact with a horizon and it does not survive contact with a floor, and everything real has both.
The cost is a reading that looks settled and is not. A rule with no cap on time reports a certain figure all the way up to the moment a cap arrives, and caps arrive from outside rather than from inside the calculation. At the instant one does, the whole of the accumulated position stops being an open position and becomes a settled figure, and it settles at whatever level the path happened to be standing on rather than at the level the rule was waiting for. On the locked path that instant is month twelve, and the figure is not plus 0.5.
One limit of the claim bears saying once more. The result speaks about the average of a process read at a rule under stated assumptions, and it stops there. An average is not a promise about any single path, and a rule whose wait has no bound is not a rule anybody finishes.
Where does this result get used once anything is priced?
The result is not a curiosity about levels on a driftless path. The result is the licence that lets a value computed at a random moment be compared with a value computed now, and without that licence a large part of the machinery further along would not be permitted.
The chain runs like this. Under the risk-neutral measure Q, the discounted standard process is a fair game, as set out under physical and risk-neutral measures. Any quantity that is a fair game under Q can be read at a followable rule instead of at a fixed date, and the result says the average is unchanged, provided a condition holds. So a payoff whose date is decided by a rule rather than by a calendar can still be valued today by taking an average, and the answer agrees with the one obtained at any other moment.
| \(S_t\) | the standard process, the invented traded quantity these examples run on, worth Rs 100/- at the start |
| \(r\) | the continuously compounded rate used for discounting, 5 per cent a year here, an invented parameter |
| \(\mathbb{Q}\) | the risk-neutral measure, under which the discounted process is a fair game |
| \(\tau\) | the moment a followable rule fires, capped at the horizon \(T\) |
| \(S_0\) | the starting value, Rs 100/- exactly |
Three consequences follow, and they are worth naming because each one is a place where somebody could otherwise go wrong. First, a contract whose settlement date is decided by a level being reached can be valued by averaging at that level rather than at the calendar date, and no separate theory is needed for it. Second, an argument that hedges up to a fixed date can be extended to hedge up to a rule, and that extension is what makes the machinery usable at all. Third, and this is the one people skip, every one of those extensions carries the conditions with it, so the horizon cap has to be imposed explicitly rather than assumed, and a model that reads a value at an uncapped rule has helped itself to something the result does not give.
The standard process itself drifts upward at 8 per cent a year. Does the result apply to it directly?
How does somebody checking a method use this rather than proving it?
An analyst is almost never asked to prove this result. Being handed a written method that reports an average computed at a moment chosen by a rule, and asked whether the number means what it says, is far more common. The check runs to four steps, needs no measure theory, and takes about five minutes.
- Name the fair game
Write down which quantity is claimed to be one, and under which measure. A great many methods never name it, and a method that cannot name its fair game has not earned any of the machinery that follows.
Under the physical measure the driving path is one. Under the risk-neutral measure the discounted process is one. The two are different objects.
- Check the rule is followable
The question at each moment is whether the rule can already be said to have fired. Followability is set out under stopping time and takes one pass over the wording, but it has to be settled before anything else here applies.
A rule quoting the highest reading of the whole period is not followable, so the optional stopping result does not reach it.
- Find the cap
Ask which of the three conditions is being relied on. Is there a fixed date by which the rule must fire? Is the quantity confined to a band until it fires? If neither, the technical condition is being claimed and should be stated rather than assumed.
If nobody can name which condition is doing the work, no condition is doing the work.
- Read the gap as the assumption
If the reported average differs from the starting value, that difference is not an edge the method found. The difference is the size of whatever unbounded assumption the method is carrying, expressed in the units of the answer.
On the failing rule the gap is 0.5, and the assumption it prices is unlimited time with unlimited capacity to be behind.
Step four turns this result into a working tool: it converts a difference in a number into a named assumption that can then be argued about. Somebody reporting an average of plus 0.5 on a fair game is not making an arithmetic mistake that can be pointed at. Such a person is reporting the value of an assumption, and the result names exactly which assumption and exactly how much of it.
The everyday version keeps it honest. Back to the spring scale and the sack of rice. Somebody says their method reads consistently high, and they are right, it does. The question to put is when they write the number down. If the answer is when the needle sits high, then their method contains an unlimited amount of standing about, and the reading they eventually write is bought with time nobody has budgeted for. Asked instead to write the number down by four in the afternoon come what may, they find the high readings vanish. The exchange contains the whole argument, and the cap on time is doing the work in both.
One last thing worth carrying away, smaller than the mathematics and more useful. Whenever a method reports that choosing the moment improved an average, the arithmetic error is not the thing to look for. There usually is not one. Look instead for the quantity that was allowed to be unbounded. Such a quantity is always there and usually stated nowhere, and once it has been named the argument is over.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for optional stopping and first passage results | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Doob | The optional stopping theorem, named for him | named in the text |
| Hull, Shreve and Wilmott | Standard texts on stochastic calculus and derivative pricing | named in the text |
The standard process and the locked path are invented.
Educational material. Not advice on any investment, tax, budget or market position.
