Model Override: Documenting a Departure From the Model
A model override is a decision to use a number other than the one the model produced, together with a record of what was changed, by how much, and on what grounds. Without the record it is not an override at all; it is a different number with no history, and nothing distinguishes it from an error.
The word override sounds like a single action, but the thing it names is a pair. There is the number, and there is the account of the number. Take either half away and what is left has no standing. A number with no account cannot be checked. An account with no number changes nothing. The first half is easy and nobody forgets it, so the second half of the pair is where all the difficulty lives.
The everyday version needs nothing but a kitchen scale. One scale reads two grams with nothing on the pan. Its owner knows this, so every time something is weighed two grams are subtracted. The subtraction is a departure from what the instrument said, and it is entirely correct. Now suppose the correction is never mentioned to anybody. Someone else weighs the same flour on the same scale, gets a reading two grams heavier, and has no way to tell whether the scale drifted, whether an arithmetic slip was made, or whether a fault they know nothing about was being quietly corrected. The correction was right, and it is still unusable. The only thing that would have made it usable was a sentence nobody wrote.
Everything below is worked on the standard process, an invented benchmark carrying four invented parameters: a single traded quantity written S with a time subscript, starting at Rs 100/-, carrying a drift of 8 per cent a year and a volatility of 20 per cent a year, watched over one year, with the risk-free rate at 5 per cent a year continuously compounded. Every figure below is computed from those four numbers and from nothing else.
What is a model override, and what makes it one rather than a mistake?
A model overrideUsing a number other than the model's output, together with a record of why the substitution was made. has two components and they are not interchangeable. The first is the substitution itself: the model produced one figure and a different figure went out under its name. The second is the written account of that substitution. Take the account away and there is nothing to distinguish the substitution from what somebody does when they mistype a cell.
Call the difference the departureThe difference between the number the model produced and the number that was used in its place, measured in the model's own units.. The departure is the plain arithmetic difference between the two numbers, and it has a size, a sign and a unit. All three matter. A departure of plus 0.058294 on a probability is a different animal from a departure of plus 0.058294 on a price in rupees, and a record that gives the size without the unit has recorded almost nothing.
| \(x_{\text{model}}\) | the number the model produced, before anybody touched it |
| \(x_{\text{used}}\) | the number that actually went out in its place |
| \(d\) | the departure, carrying a size, a sign and the units of the output |
Nothing in that arithmetic settles whether a departure was reasonable, and the silence of the arithmetic is the point. The arithmetic of a departure is identical whether the substitution was a careful judgement, a typing slip, or a number somebody preferred, and only the record can tell those three apart. This is why an override is defined by its record rather than by its size, its direction, or the seniority of whoever made it.
When does a departure from a model output need documenting?
The test is not about size and it is not about seniority. The test is a question about the future: will anybody ever need to know where this number came from? Numbers that leave a model get used, and numbers that get used get questioned, so the answer is almost always yes. Any figure that leaves the model and is then used for anything needs a record, and the size of the departure has nothing to do with it.
People reach for a size threshold because it feels tidy. Record anything over one per cent, say, and let the small ones through. A size threshold fails in both directions at once. A small departure repeated across many outputs is not small in aggregate, and the record is exactly what would have revealed the repetition. Meanwhile a large departure that clears the threshold gets a record for the wrong reason, as if the record were a penalty triggered by size rather than the ordinary companion of any substitution.
There is a cleaner way to see it. Sort every number the model touches along two axes: whether it matches the model output or departs from it, and whether it gets used for anything. Only one of the four boxes needs a record. A number that gets produced gets used, so one of the remaining three stays empty in practice.
A number left the model, was adjusted slightly, and was then used in a valuation nobody expects to revisit. Does the adjustment need a record?
What question does Risk-Neutral Valuation answer?
Now to the part that produces more false overrides than any other single cause. The standard process can be written down twice, under two different measures, and the two descriptions are not competitors. The two descriptions answer different questions, and both are correct at the same time.
Under the pricing measureThe measure used for valuing, written Q, under which the growth term of the process is the risk-free rate rather than the drift., written Q, the growth term of the standard process is not the drift at all. The growth term is the risk-free rate. Replacing the drift by the rate is not an approximation and not a convenience. The replacement makes a discounted value behave as a fair game, and that is the entire mechanical purpose of the measure. Risk-Neutral Valuation is the name of the task carried out under Q, and it answers one question only. The question is what something is worth today.
| \(S_t\) | the standard process at time \(t\), starting at Rs 100/- at time zero |
| \(\mu\) | the drift, 0.08 a year, the growth term under the physical measure |
| \(r\) | the risk-free rate, 0.05 a year continuously compounded |
| \(\sigma\) | the volatility, 0.20 a year, and it is the same under both measures |
| \(W_t,\ \tilde{W}_t\) | Brownian motion, standard under \(P\) and under \(Q\) respectively |
| \(P,\ Q\) | the physical measure and the risk-neutral measure |
Read what is unchanged there. The unchanged part is what people forget. The volatility is the same 20 per cent in both lines. The starting value is the same Rs 100/-. The horizon is the same year. Exactly one thing differs between the two descriptions, and it is the growth term. Every difference in every figure computed from them traces back to that single substitution.
| \(S_T,\ S_0\) | the standard process at the horizon and at time zero, Rs 100/- at the start |
| \(T\) | the horizon, one year |
| \(N\) | the standard normal distribution function |
| \(r,\ \sigma\) | 0.05 and 0.20, so the numerator is 0.03 and the ratio is 0.15 exactly |
The figure 0.559618 will already be familiar from the pricing machinery, where the same number appears as the second of the two normal terms in the closed form price of the at-the-money contract on the standard process. The figure is not a forecast and it never was. A component of a price is the whole of what it is.
What question does Real-World Forecasting answer instead?
Under the physical measureThe measure describing how the process actually moves, written P, under which the growth term of the process is the drift., written P, the growth term is the drift of 8 per cent. Real-World Forecasting is the name of the task carried out under P, and it answers a completely different question. The answer describes how the process is likely to move. Nothing about it is a valuation, and no discounting appears in it anywhere.
| \(\mu\) | the drift, 0.08 a year, appearing here and nowhere in the previous block |
| \(\tfrac{1}{2}\sigma^{2}\) | half the variance rate, 0.02 exactly, so the numerator is 0.06 exactly |
| \(\sigma\sqrt{T}\) | 0.20 exactly at the one year horizon, so the ratio is 0.30 exactly |
Set the two blocks beside each other and the shape of the confusion becomes visible. The two formulas are letter for letter identical apart from a single symbol in the numerator. Same distribution function, same volatility, same horizon, same starting value. One has the drift in it and one has the risk-free rate in it, and that is the whole of the difference.
Which of the two figures has the drift of 8 per cent inside it?
How far apart are the two figures, and is the distance big enough to notice?
Subtract them. The physical figure is 0.617911 and the pricing figure is 0.559618, so the difference is 0.058294, or 5.829373 percentage points. Expressed as a share of the pricing figure it is 10.416706 per cent. Every digit of that gap is a computed consequence of the four parameters and of nothing else.
Size is the reason this particular confusion is so durable. A gap of a hundredth of a point would be dismissed as rounding and never investigated. A gap of thirty points would be so obviously wrong that somebody would find the bug in an afternoon. Five and a bit percentage points sits precisely in the range where a careful person concludes they have found something. The false override built on this gap therefore feels like diligence rather than carelessness.
How far apart are the two figures at the locked parameters of the standard process?
One more property of the gap is worth having before the control below is touched, and it settles whether either figure is a corrupted version of the other. Write the gap as a function of the drift and hold everything else fixed. The pricing figure has no drift in it, so it appears in that function only as a constant.
| \(G(\mu)\) | the physical figure less the pricing figure, at the one year horizon |
| \(\mu\) | the drift, the only quantity being varied |
| \(r,\ \sigma\) | held at 0.05 and 0.20 throughout, so the second term is the constant 0.559618 |
A quantity that can be positive, zero or negative depending on a parameter neither figure disagrees about is not a measure of error in anything. At a drift of 5 per cent the two figures coincide exactly, and nobody would call that agreement evidence that the model had suddenly started working.
The drift is about to be doubled from 8 per cent to 16 per cent. Before it moves: what happens to the pricing figure of 0.559618?
Sweep the drift and watch only one figure move
One control: the drift, from nought to 16 per cent. Everything else is held at the locked values, so the volatility stays at 20 per cent, the rate at 5 per cent, the starting value at Rs 100/- and the horizon at one year. Both figures below are computed from the normal distribution function rather than sampled, so a given setting always returns exactly the same reading.
| Drift under P | Real-World Forecasting | Risk-Neutral Valuation | Gap, in points |
|---|---|---|---|
| 0 per cent | 0.460172 | 0.559618 | minus 9.944553 |
| 4 per cent | 0.539828 | 0.559618 | minus 1.978986 |
| 5 per cent | 0.559618 | 0.559618 | 0.000000 |
| 8 per cent, the locked value | 0.617911 | 0.559618 | plus 5.829373 |
| 12 per cent | 0.691462 | 0.559618 | plus 13.184498 |
| 16 per cent | 0.758036 | 0.559618 | plus 19.841866 |
Look down the third column. The column does not change. The third column is the answer Risk-Neutral Valuation gives, and its question never mentioned expectation, so the answer stays the same whatever the process is expected to do. The second column moves by nearly thirty percentage points across the range. A figure that is constant against a parameter and a figure that swings by thirty points against the same parameter cannot be two attempts at the same quantity, and reconciling them is not a task anybody should accept.
What has to be written down when a model is overridden?
Six items. Not a form, not a policy, not a template: six pieces of information without which the record cannot answer the questions that will be put to it. DocumentationThe six items without which a departure has no history and cannot be checked, reversed or learned from later. here means precisely these and nothing more.
- The model outputThe figure the model produced before anybody touched it, at full precision. On the case above this is 0.559618. Recording the replacement without the original makes the departure unrecoverable, because nobody can subtract a number that was never written down.
- The number used insteadWhat actually went out under the model's name, again at full precision. On the case above this is 0.617911.
- The size of the departureBoth the plain difference and its relative size, with units. Here that is plus 0.058294, which is 5.829373 percentage points and 10.416706 per cent of the model output. Writing it out rather than leaving it to be inferred is what lets somebody scan a list of departures and see which ones are large.
- The groundsOne sentence describing what actually happened, in terms of the model. Not the conclusion, and not the feeling. This is the item that gets left out, and it is the one that carries all the information.
- The dateThe day the decision was taken, which fixes what was known at the time. A departure that was reasonable given what was on the desk that morning and unreasonable a month later is a normal outcome, and only the date lets anybody tell the two situations apart.
- Who decidedA named person or a named role. Not for blame. For the far more ordinary reason that somebody reading the record in a year will have a question, and the record should say who can answer it.
The fourth item is where records fail, and they fail in a specific way. GroundsThe stated reason for a departure, which must describe what actually happened rather than restate the conclusion. that read "the model figure looked too low" have recorded a conclusion, not a reason. Grounds naming the comparison that was made, a pricing figure set against a physical one and adjusted toward it, have recorded what happened, and a reader can then decide for themselves whether that was a sound thing to do. The test for a well written ground is that a reader who disagrees with the decision can still tell exactly what the decision was.
How many items make a departure into a documented override?
What does an override built on a measure confusion actually cost?
The two halves of the subject now come together. A measure confusionComparing figures produced under two different measures as though they answered the same question, which is the commonest source of a false override. starts with somebody taking the pricing figure, setting it against what they expect the process to do, judging it too low, and adjusting it. Every step of that reasoning is careful. The two numbers being compared were never comparable, so the conclusion is still wrong.
The cost is computable, and it is not small. Suppose the pricing figure of 0.559618 is replaced by 0.617911 inside the closed form price of the at-the-money contract on the standard process, on the grounds that the higher figure is the one that describes reality. The contract is worth Rs 10.450584/- when the pricing figure is left alone. With the substitution made it comes to Rs 4.905512/-, a fall of Rs 5.545071/- and 53.0599 per cent of the correct price.
The mistake, and what it costs on the locked case
The reasoning runs: the model says the process finishes above Rs 100/- with probability 0.559618, but the drift is 8 per cent a year so it should be nearer 0.617911, therefore the model figure is too low and should be raised. Both figures in that sentence are correct. The word therefore is the error.
Substituting 0.617911 for 0.559618 in the price of the at-the-money contract moves it from Rs 10.450584/- to Rs 4.905512/-, a fall of Rs 5.545071/- or 53.0599 per cent, and it does so while every arithmetic step in the calculation remains valid.
The direction is the last thing anybody expects. Notice it. The adjustment was made to raise a figure that looked too low, and the effect on the price was to cut it by more than half. The probability sits inside a subtracted term, so pushing it up pushes the price down. The record would record an intention to raise something, so nothing in it would ever say so.
There is a second fault hiding inside the first. The other normal term in that price, 0.636831, is also a quantity computed under the pricing measure. Adjusting one and not the other leaves a price built half under one measure and half under the other. Such a price is not a description of anything at all. Grounds that said only "figure looked low" did not describe what happened, so a record carrying them would never surface the second fault.
A pricing figure disagrees with the expected behaviour of the process. Is that disagreement a reason to override the pricing figure?
Why is a documented mistake worse than an undocumented one?
The claim ahead sounds backwards on first reading. Documentation is what makes a departure checkable, and yet a documented error does more damage than a bare one. Both statements are true, and they are true for the same reason: a record carries authority.
An undocumented number invites a question. Somebody looks at 0.617911 where the model says 0.559618, finds nothing explaining it, and asks. The absence of a record is itself a signal, and it is a loud one. A documented number with wrong grounds does the opposite. The grounds answer the question before it is asked, and they answer it incorrectly. The record does not merely fail to prevent the error. Everyone downstream reads the grounds and moves on, so the record actively stops the error being found.
The same authority is why reversibilityBeing able to undo a decision later, which requires a record complete enough to reconstruct what the decision was. matters more than it first appears. Reversing a departure means restoring the model output, and restoring it means knowing what it was. Row one of the record is the only thing that carries it. A wrong departure with a complete record can be undone in a minute by anybody who spots the fault. A wrong departure with grounds that read "figure looked low" cannot be undone at all. Nobody can tell whether the departure was a measure confusion, a data issue or a considered judgement they do not understand.
Taken the other way, those two sentences give the working rule. The value of a record is not that it justifies the decision. The value is that it makes the decision undoable by somebody who was not there. Undoability is a much lower bar than justification and a much more useful one.
Why is a documented mistake worse than an undocumented one?
How does somebody actually use this when a number lands on their desk?
Concretely, and it takes about two minutes without market access, data or software. Somebody handed a figure that departs from a model output runs three checks in order, and none of them requires agreeing or disagreeing with the decision.
First, ask which measure the model output came from. If it came from a valuation it came from the pricing measure, and no comparison against an expectation of what will happen is admissible. If it came from a forecast it came from the physical measure, and no discounting belongs anywhere near it. The measure question disposes of most false overrides before any arithmetic starts.
Second, read the grounds and test them against one sentence: could somebody who disagrees with this decision still say exactly what was decided? Grounds that survive that test describe an action. Grounds that fail it describe a feeling. On the case above, a ground naming the pricing figure, the physical figure and the adjustment between them passes even though the decision it describes is wrong. The pass is exactly the property wanted, and it is what lets the fault be found.
Third, check that row one is present at full precision. Not the departure, not the replacement: the original. Everything else in a record can be reconstructed from the surrounding papers if it goes missing. The model output cannot be reconstructed. The model has since been rerun with different settings and no longer produces it. The single most valuable line in an override record is the number nobody used.
The same three checks work when writing a record rather than reading one, and they are quicker in that direction: the measure is named, the action is described instead of the feeling, and the figure about to be set aside is written down. A household running its budget off a scale that reads two grams heavy needs exactly the same discipline. Record what the scale said, and not only the conclusion drawn from it.
Had any of the nine reading orders before this one fitted a model to an observation?
What does this reading order build that the nine before it did not?
Nine reading orders came before this one in this subject area. Between them the nine built the probability space, the process, the integral and the chain rule, the differential equation, the pricing argument, the option pricing machinery, the volatility models, the rate models and the numerical methods. Every one of them held models. Not one of them fitted a model to anything, and that was deliberate rather than an oversight.
Calibration and model risk is where fitting happens, and everything in it follows from that single act. Once a model is fitted to numbers from outside itself, a new object exists. A fitted model is not the same thing as a model and does not have the same standing. Its parameters are no longer assumptions somebody chose. The parameters are outputs of a procedure, and they carry the date of the observations, the choice of loss function, and the range over which the fit was performed.
The override belongs first in that reading order because it is the earliest place the two things meet. A model produces a number; somebody outside the model produces a different one; a decision gets taken between them. Everything later here is a variation on that meeting. The subjects that follow are a market price and a model price as objects, what fitting is minimising and how three loss functions on identical data return three answers, what a fitted parameter is worth and what dates it, when two parameter sets cannot be told apart, and what has to be documented, what has to be labelled and what has to be said about what a model cannot do.
Covered separately. Calibration itself is covered separately: what fitting minimises, and how three loss functions on identical data return three different answers, comes later in this reading order. Documentation in full comes later too, and the six items above are the minimum that makes a departure checkable rather than a treatment of model documentation as a whole. The payoff of any contract belongs to a different subject area, and the contract arrives already known, appearing here only as the function whose price the substitution moved.
On jurisdiction. No threshold, rate, period or standard is quoted. The mathematics of the two measures holds wherever the measures are defined, and the six items are a description of what makes a record usable rather than a statement of what any authority requires. Where a conduct duty attaches to how a model output or an override is recorded, that duty sits with the authority concerned and must be confirmed at source.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for treatments of measure change, model risk and the standing of fitted parameters | arxiv.org |
| Social Science Research Network | Working paper repository for the same material, including model governance and override practice | ssrn.com |
| Girsanov, and Radon and Nikodym | The change of measure that carries the process from the physical measure to the pricing measure | named in the text only |
| Black, Scholes and Merton, 1973 | The closed form price whose second normal term is the 0.559618 substituted in the failure block | named in the text only |
| Hull, Shreve and Wilmott | Standard texts on stochastic calculus, derivative pricing and model risk | named in the text only |
The standard process and its four parameters are invented.
Educational material. Not advice on any investment, tax, budget or market position.
