Model Uncertainty and Limitations: Stating What the Model Cannot Do
Model uncertainty is everything that could make a model's output wrong other than an error in the arithmetic itself. Model uncertainty divides into two kinds. Parameter risk is a wrong number sitting inside a right equation. Model risk is a wrong equation, whatever numbers are put into it. The two have different sizes, different signatures and different remedies.
Calibration fitted a model. Five invented prices were handed to a search, the search returned a volatility, and the volatility was written down to six decimal places. Everything since has been about what that number is and is not. One question comes after all of it, and a reader eventually has to answer it out loud in front of somebody else: given that the model has been fitted as well as it can be, what can it still be wrong about, and by how much?
The short answer is that it can be wrong in three distinct ways, and the three are routinely collapsed into one vague feeling that models are not to be trusted. A vague feeling cannot be acted on, so it is useless. Three separate numbers can be acted on. The three are kept apart below, each with its own size, in its own units, on the same invented instance.
The units matter more than they look. One of the three is measured in volatility, one in rupees of price, and one in rupees of arithmetic. Set on the same line without their units, the three make a comparison that means nothing. Set on the same line with their units, they make the only honest statement anyone can make about what a fitted model is worth.
Everything below is computed from invented inputs. The quantity being modelled is the standard process, written S with a time subscript, starting at Rs 100/- with a rate of 5 per cent a year over a horizon of one year. The numbers being fitted to are five invented call prices at strikes of Rs 80/-, Rs 90/-, Rs 100/-, Rs 110/- and Rs 120/-, quoted as volatilities of 0.240000, 0.220000, 0.200000, 0.190000 and 0.185000. No market quoted the five. Each was chosen so that all three kinds of uncertainty show up on one instance.
What is model uncertainty?
Model uncertaintyEverything that could make an output wrong other than an error in the arithmetic itself. is the collected answer to a single question: if the arithmetic is done perfectly, what is left that could still make the output wrong? Notice what is excluded. A sign error in a spreadsheet is not model uncertainty; it is a mistake, and mistakes are found by checking. Model uncertainty is what remains after every mistake has been found and removed. Model uncertainty is the part that survives a correct calculation.
The everyday version is exact rather than approximate. A kitchen scale is asked what a stone weighs. Three things can go wrong that have nothing to do with anyone's ability to read a dial. The scale's zero point could be set slightly off, so a right instrument is carrying a wrong setting. The scale could weigh by squashing a spring, which is fine for stones and useless for anything that floats. The instrument itself could be the wrong instrument for the job. The digital display could also have stopped flickering and settled on a figure. A settled display says the electronics have finished thinking and says nothing about the stone. The three faults, in that order, are parameter risk, model risk and numerical tolerance.
Model Limitation
A model limitation is a specific, named, sized statement of one of those three. A limitation is not a mood and it is not a disclaimer. A limitation that does not carry a number is not a limitation; it is a sentence about a limitation. The difference runs through everything below, so it is worth fixing now. Saying that the scale weighs by squashing a spring is a fact about the scale. Saying that it therefore reads a floating object at nothing at all, and that no adjustment of the zero point changes that, is a limitation.
The same distinction in the pricing version. Saying that the model assumes a constant volatility across all strikes is a fact about the model. Every reader of a pricing document has seen that sentence and not one of them can do anything with it. Add that on these five invented targets the model therefore misses the lowest strike by Rs 0.662462/- at the best available setting, and the fact becomes a limitation. The statement now carries a size, and a size can be compared with something.
| \(C(K_i;\sigma)\) | the price the model produces at the \(i\)-th strike when the volatility is set to \(\sigma\) |
| \(C^{\text{obs}}_i\) | the \(i\)-th target, one of the five invented prices the model is being held against |
| \(K_i\) | the strike of the \(i\)-th contract, Rs 80/- to Rs 120/- here |
| \(\sigma\) | the volatility parameter, whatever it has been set to |
| \(\hat{\sigma}\) | the reference setting, 0.198202 here, the answer under squared price error |
Everything below is an unpacking of that one line. The first bracket has a size, the second bracket has a size, and they are not the same size and not even the same kind of thing. An account that says only that there is uncertainty has refused to compute either bracket.
What is the difference between parameter risk and model risk?
Parameter Risk vs Model Risk
Parameter riskA wrong number in a right equation, worth 0.008798 of volatility on the invented instance here. is the risk that the equation is fine and the number fed into it is not. Model riskA wrong equation that no setting of any parameter can repair. is the risk that the equation itself is the wrong shape, so no number would have helped. The two sound close together and behave nothing alike, so the invented instance is the fastest way to separate them.
Start with the parameter. The fit two pieces back returned 0.198202, and it returned that figure because the loss was squared price error. Changing the loss changes the answer, not because the data changed but because the question did. Under squared error in implied volatility the same five targets return 0.207000. Under a price error divided by each target's vega, they return 0.204420. All three losses are ordinary and all three are defensible. Nothing in the five numbers says what close means, so nobody observed which one was correct.
| \(L_{\text{price}}\) | squared error in rupees of price, which returns 0.198202 |
| \(L_{\text{vega}}\) | each price error divided by that strike's vega before squaring, which returns 0.204420 |
| \(L_{\text{vol}}\) | squared error in the quoted volatility, which returns 0.207000, the plain average of the five quotes |
| \(\mathcal{V}_i\) | vega at the \(i\)-th strike, how much its price moves for a small move in volatility |
| \(\sigma^{\text{obs}}_i\) | the volatility the \(i\)-th target was quoted at, from 0.240000 down to 0.185000 |
| \(\Delta\sigma\) | the spread, the width of the band of settings that a defensible procedure could return |
| \(\hat{\sigma}_j\) | the answer returned by the \(j\)-th loss, one of the three above |
The second kind of uncertainty lives on the right-hand side of that picture. Fifteen readings, five strikes at each of the three defensible settings, and not one of them is nought. Move the parameter and the misses rearrange themselves; the low strikes get less wrong and the high strikes get more wrong, or the other way round. The misses never all go away together.
The surviving residue is the signature of model risk: the parameter can shuffle it around and cannot remove it. There is a reason the parameter cannot, and the reason explains the whole thing. The five targets were quoted at five different volatilities, from 0.240000 down to 0.185000. The model has one volatility, used at every strike. One number is being asked to be five different numbers at once. One number cannot be five, and no search, no loss and no amount of computing power changes that.
Back to the scale for a moment. Setting the zero point wrong is parameter risk, and turning the screw fixes it. Weighing a floating object with a spring is model risk. The instrument is answering a different question from the one that was asked, so no position of the screw helps. A better setting does not fix it. A different instrument does.
Which of the two is repaired by a better number?
What do the three come to on the invented instance?
Here is the whole worked instance in one place, and every number in it was computed from the invented inputs rather than asserted. Three kinds of uncertainty, three sizes, three units, all measured on the same five targets and the same one-year horizon.
| Kind | What is wrong | Size | Unit |
|---|---|---|---|
| Parameter risk | The loss function was never observed, so three defensible choices give three answers | 0.008798 | volatility |
| Model risk | One volatility is asked to serve five strikes quoted at five different volatilities | Rs 0.662462/- | price, at the lowest strike |
| Numerical tolerance | The refinement ladder settled without ever refining the space step | Rs 0.055996/- | price, on the grid figure |
Take the parameter risk first. The three fits are 0.198202 under squared price error, 0.204420 under price error divided by vega, and 0.207000 under squared volatility error. The last of the three is just the plain average of the five quoted volatilities, a useful sanity check on the arithmetic: the average of 0.240000, 0.220000, 0.200000, 0.190000 and 0.185000 is 0.207000 exactly. The spread of 0.008798 is the width of the answer that the data left undetermined, and the choice that determined it was made by a person and recorded nowhere unless somebody wrote it down.
Now the model risk, and here is the sharpest way to state it. None of the three settings fits exactly, so ask a fairer question. Ask instead: across every possible setting of the volatility, what is the smallest that the largest of the five misses can be made? The question is a real optimisation with a real answer, and the answer is a floor nothing gets under.
| \(R\) | the floor, in rupees, on the largest of the five misses, over every setting of the one parameter |
| \(\min_{\sigma}\) | taken over every volatility, not only over the three defensible fits |
| \(\max_{i}\) | taken over the five strikes, so this is the worst single miss at a given setting |
| \(\big|\cdot\big|\) | absolute value, because a miss counts the same whichever side of the target it falls |
Set side by side, the two numbers carry the whole distinction. Parameter risk is 0.008798 of volatility, and it exists because somebody had to choose. Model risk is a floor of Rs 0.599082/- on the worst miss, and it exists because one number cannot be five numbers. The first is a decision problem. The second is an arithmetic fact about the shape of the equation. Filing both of them under one heading called model uncertainty and stopping there throws away the only distinction that tells an analyst what to do next.
How large is the parameter risk that comes from the choice of loss function on these five targets?
Which of the two risks is larger, and does that have one answer?
A tempting close to the last section is that model risk is the big one and parameter risk is the small one. On this instance that is not true, and the arithmetic says so plainly. The two are quoted in different units, so the parameter risk has to be converted into rupees before any comparison at all. Ask what the band of 0.008798 is worth as a price move at each strike.
| Strike | Price at 0.198202 | Price at 0.207000 | Parameter swing | Model miss at the best fit |
|---|---|---|---|---|
| Rs 80/- | 24.564539 | 24.687280 | 0.122741 | 0.662462 |
| Rs 90/- | 16.650741 | 16.891505 | 0.240764 | 0.606838 |
| Rs 100/- | 10.383131 | 10.713483 | 0.330352 | 0.067452 |
| Rs 110/- | 5.968944 | 6.317312 | 0.348368 | 0.324179 |
| Rs 120/- | 3.186331 | 3.487726 | 0.301394 | 0.441183 |
| Root mean square | 0.280843 | 0.471464 |
Read the last two columns across. At the lowest strike the model's own miss is Rs 0.662462/- against a parameter swing of Rs 0.122741/-, so the model risk is more than five times the parameter risk and the equation is the problem. At the middle strike the model's miss is Rs 0.067452/- against a parameter swing of Rs 0.330352/-, so the parameter risk is nearly five times the model risk and the choice of loss is the problem. Which uncertainty dominates is not a property of the model; it is a property of the contract being asked about.
The reversal is the reason a document cannot get away with one summary sentence, and it is worth sitting with. A note recording that model risk is the dominant uncertainty here would be correct about the lowest strike and wrong about the middle one, using the same model, the same five targets and the same day. Averaged across the five, the root mean square figures are Rs 0.280843/- of parameter swing against Rs 0.471464/- of model miss, so the average does favour the model risk. Averages hide reversals, and this one hides a factor of five in both directions.
The scale again, and it survives the extension. When stones are being weighed, the spring is fine and the only thing that matters is where the zero point sits. When something that floats is being weighed, the zero point is irrelevant and the spring is everything. The instrument did not change between the two jobs. The question put to it changed.
What is the difference between model error and market risk?
Model Error vs Market Risk
The third distinction is the one most often blurred in writing, usually to somebody's convenience. Model errorThe model being wrong about a fixed set of numbers, measured with those numbers held still. is the model being wrong about a set of numbers that is not moving. Market riskThe numbers themselves moving, outside the model rather than inside it. is those numbers moving. The first is inside the model. The second is outside it, and it was never the model's business.
Freeze the five targets. The five targets are Rs 25.227000/-, Rs 17.257579/-, Rs 10.450584/-, Rs 5.644765/- and Rs 2.745149/-, and for the purposes of measuring model error they do not move at all. Against those frozen five, the fitted model produces Rs 24.564539/-, Rs 16.650741/-, Rs 10.383131/-, Rs 5.968944/- and Rs 3.186331/-. The gaps between the two rows are model error, entirely, and every one of them was computed with nothing in the outside world changing.
Now unfreeze them. Tomorrow the five targets are five different numbers, and the fitted model, still carrying 0.198202, produces a set of gaps that are different again. Nothing about the model became more or less wrong overnight. The question changed. Model error is the answer to a question asked about fixed numbers, and market risk is the numbers not staying fixed, so one is a property of the model and the other is a property of the world.
The reason this matters is not taxonomic. The gap was measured on a frozen set, so a document that reports a large gap and attributes it to market movement has explained nothing. A document that reports a loss and attributes it to model error when the targets moved has explained nothing either, in the opposite direction. Both sentences sound like explanations and neither one is checkable. Being uncheckable is exactly what makes them attractive to write.
Which of the two is inside the model?
Does a numerical tolerance bound the arithmetic or the answer?
What is a numerical tolerance, and what does it bound?
Numerical Tolerance
A numerical toleranceHow far apart two successive refinements of a calculation sit. Bounds the arithmetic, not the distance from the truth. is a rule for deciding when to stop refining a calculation. The answer is computed at one level of refinement, computed again at a finer level, and if the two agree to within some agreed amount the calculation is declared finished. The rule is entirely reasonable and every numerical method needs one. A tolerance is not a statement about accuracy, and the numbers below make the distinction undeniable.
Numerical pricing produced the instance. A grid method was used to price the at-the-money contract, whose closed form is Rs 10.450584/-. The time step was refined while the space step was held at 0.05, and the ladder ran as follows.
| Time steps | The grid figure | Move from the rung before | Distance from the truth |
|---|---|---|---|
| 16 | Rs 10.331273/- | Rs 0.119311/- | |
| 50 | Rs 10.415025/- | Rs 0.083752/- | Rs 0.035558/- |
| 100 | Rs 10.404265/- | Rs 0.010760/- | Rs 0.046319/- |
| 252 | Rs 10.397779/- | Rs 0.006486/- | Rs 0.052805/- |
| 1,000 | Rs 10.394588/- | Rs 0.003191/- | Rs 0.055996/- |
The third column read on its own is all that remains when there is no closed form to check against. The moves shrink: Rs 0.083752/-, then Rs 0.010760/-, then Rs 0.006486/-, then Rs 0.003191/-. Each one is smaller than the one before it. There is no oscillation, no wobble and no sign of anything wrong. Every internal signal says the calculation has settled.
The fourth column is readable only because somebody supplied the closed form. It grows. Rs 0.035558/-, then Rs 0.046319/-, then Rs 0.052805/-, then Rs 0.055996/-. The calculation is settling and moving away from the truth at the same time, and the tolerance is watching only the settling.
| \(V_n\) | the answer the method produced at the \(n\)-th level of refinement |
| \(V_{n+1}\) | the answer at the next level, one rung finer |
| \(V^{\text{true}}\) | the answer actually wanted, Rs 10.450584/- here, and normally unavailable |
| \(\varepsilon\) | the tolerance, the amount of agreement that is treated as enough |
Why did it happen? Because the ladder refined the time step and never touched the space step. The space step was contributing a fixed distance of about Rs 0.057070/- throughout, and refining the other spacing could not see it. A tolerance measures agreement along whatever direction was refined, and is silent about every direction that was not. That is not a defect in the tolerance. Silence about every unrefined direction is what a tolerance is.
The failure: reporting a tolerance as an accuracy
Somebody runs the grid, watches the last few readings agree to a third of a paisa, and writes in the document that the price is Rs 10.39/- to within Rs 0.001/-. Every word of that sentence is produced honestly and the sentence is wrong. The tolerance and the accuracy are answers to different questions, and only one of them was computed.
The reading was Rs 10.394588/-. The truth was Rs 10.450584/-. The distance is Rs 0.055996/-, fifty-six times the figure quoted as the precision. The space step was held at 0.05 the whole time and was carrying the entire error, so nothing in the calculation could have revealed the distance. A tolerance says how much the calculation is still moving, and a reader who hears it as how right the answer is has been handed a precise-sounding claim that is wrong by more than fifty times the precision it claims.
The tell is in the wording, and it is worth learning to hear. A document that says the figure converged to within Rs 0.001/- has made a statement about a sequence. A document that says the figure is accurate to within Rs 0.001/- has made a statement about the truth. The first can be earned by any calculation that stops moving. The second needs something to compare against, and if the document does not say what that something was, it has not been earned at all.
A grid result satisfies a tolerance of Rs 0.001/-. How accurate is it?
How does an analyst find out which uncertainty is in play?
The three kinds have different signatures, and the signatures can be read off a small number of planned runs. The runs are not there to produce a report. Until they have been done, which of the three is at work remains unknown, and the remedies do not overlap at all.
How to Run a Pricing-Model Sensitivity Check
- Fix everything that is not the parameter. The same five targets, the same horizon, the same rate. Any run in which two things moved at once says nothing about either of them, and this is the step people skip because it feels like it does not do anything.
- Refit under at least two more losses and record the answers. Here that gives 0.198202, 0.204420 and 0.207000. The spread of 0.008798 is the parameter risk, and it is a measured band rather than a guess. The three losses go down beside the three answers, because a band without its procedure cannot be reproduced.
- Reprice at every setting in the band and record the swing at each contract. That converts the band from volatility into rupees, which is the only unit in which it can be compared with anything else. Here it runs from Rs 0.122741/- at the lowest strike to Rs 0.348368/- at Rs 110/-.
- Record the miss at every setting, at every target. If some misses survive at every setting, that surviving part is model risk and no further refitting will touch it. Fifteen readings here and not one of them is nought.
- Push past the defensible settings and ask for the floor. The search runs over every volatility, not only the ones anybody would defend, for the smallest the worst miss can be made. Rs 0.599082/- here, at 0.202831. That single number is the cleanest statement of model risk available, because it makes no reference to any fitting procedure.
- Refine each numerical spacing separately, and never only one of them. A ladder that moves one spacing certifies that spacing and nothing else. The grid figure of Rs 10.394588/- settled to a third of a paisa while carrying Rs 0.055996/- of error, entirely because the space step was never refined.
- Compare the three in the same units, per contract, and refuse to average them into one sentence. At the lowest strike the model risk is more than five times the parameter risk; at the middle strike the parameter risk is nearly five times the model risk. One sentence covering both would be wrong about one of them.
The procedure is not a way of making the model better, and running it changes no number anywhere. The procedure finds out which of the three uncertainties is present, so that the effort spent next is spent on the one that is actually large.
The volatility is about to move across all three fitted answers. Before it does: does any of them remove the misses?
Splitting every miss into the part that was chosen and the part that cannot be avoided
Held fixed: the standard process starting at Rs 100/-, a rate of 5 per cent a year, a horizon of one year, and the same five strikes from Rs 80/- to Rs 120/-. The only thing that moves is the volatility. Each bar below is one strike's miss, split into a dark segment and a lime segment. The dark segment is the miss at the reference setting of 0.198202 and never changes; the lime segment is everything the choice of setting added to it. At 0.198202 the misses are minus 0.662462, minus 0.606838, minus 0.067452, plus 0.324179 and plus 0.441183. At 0.204420 they are minus 0.576709, minus 0.437299, plus 0.165950, plus 0.570327 and plus 0.653633. At 0.207000 they are minus 0.539720, minus 0.366074, plus 0.262899, plus 0.672547 and plus 0.742577. The spread of the three answers is 0.008798. The dark segments are the equation rather than the number, so no setting anywhere in the range empties a single one.
At a volatility of 0.198202 the largest of the five misses is Rs 0.662462/- at the strike of Rs 80/- and the smallest is Rs 0.067452/- at the strike of Rs 100/-. The control is sitting exactly on the reference setting, so every parameter part is nought and each bar is its model part alone. Those five dark segments are what the equation cannot do, and they are the same five numbers at every position of the control.
How is a limitation communicated honestly?
How to Communicate Pricing-Model Limitations
An honest limitationA statement naming what is wrong, how large it is, and what would fix it. All three, or it is not one. has three parts and it is not honest with two of them. An honest limitation names what is wrong, gives the size, and names what would fix it. Drop the size and the reader cannot decide whether to care. Drop the remedy and the reader cannot decide what to do. Both omissions are common and both produce a sentence that looks like a disclosure and functions as decoration.
Compare the two sentences directly, on the same model, about the same thing.
| The part | The ritual version | The honest version |
|---|---|---|
| What is wrong | The model assumes a constant volatility across all strikes | The model assumes a constant volatility across all strikes |
| How large | Not stated | It misses the lowest strike by Rs 0.662462/- at the best available setting, and the smallest the worst miss can be made at any setting is Rs 0.599082/- |
| What would fix it | Not stated | An equation whose volatility can vary with the strike. More information about this one parameter would not help |
The first column of that table is the whole of what most documents say, and it is the part that conveys nothing. Every reader already suspected that a model with one volatility assumes one volatility. Only the writer can supply the rest: the assumption costs Rs 0.662462/- at the bottom of the strike range and Rs 0.067452/- in the middle of it. The size is the only part of a limitation that a reader could not have worked out for themselves.
There is a second reason to insist on the size, and it is less obvious. A limitation with a number attached can be checked, and therefore can be wrong, and therefore is worth writing. A limitation without a number cannot be checked by anyone. Such a limitation can never be shown to have been overstated or understated, so writing it costs nothing and proves nothing. A statement that cannot fail is not a disclosure.
What makes a stated limitation useful rather than ritual?
What does stating a limitation not do?
One sentence keeps the procedure above from turning into an empty exercise, and the procedure is exactly the kind of advice that gets followed enthusiastically and then mistaken for a solution. Writing a limitation down does not make it smaller. The miss of Rs 0.662462/- at the lowest strike is Rs 0.662462/- before anybody writes it in a document and Rs 0.662462/- afterwards. Not one figure above moved because it was recorded.
Ritual disclosureStating a limitation without its size, so that the document looks careful and the reader learns nothing. does its real damage here, and the damage is not the empty sentence itself. The damage is the feeling of completion the sentence produces. A document full of limitations reads as diligence. The list reads as somebody having thought carefully. A list of things that are wrong is not a list of things that have been addressed, so a reader who has finished one has not been made safer by a single rupee.
The scale one final time, and it lands the point exactly. A label stuck on the scale records that the instrument reads forty grams light. The error is now visible and useful, and everyone who weighs anything on it can correct for it. The scale still does not read correctly. Peeled off, the label leaves the scale reading exactly as it did before. Left on, it leaves the scale reading forty grams light, only now with a sticker.
So what does documenting achieve? Precisely one real thing. Documenting moves the error from unknown to known, and somebody can then decide whether to accept it, price around it, or replace the model. None of the three decisions can be made about an error nobody has measured. Documentation converts an unknown error into a known one, and every reduction happens afterwards, in the work that follows, or it does not happen at all.
Does writing down a limitation reduce it?
Everywhere: none of it is a rule anyone set
None of the three uncertainties is a rule that a jurisdiction or an authority set. The split of a miss into a parameter part and a model part is arithmetic. The floor of Rs 0.599082/- is an optimisation over one variable. The fact that a tolerance compares two answers of the same calculation to each other rather than to the truth is a property of what a tolerance is. What a document is obliged to say does depend on jurisdiction, and the obligation is set by the authority there rather than by the arithmetic.
References
| Source | Document | Where |
|---|---|---|
| arXiv, Quantitative Finance | Preprints on model risk, parameter uncertainty and the choice of calibration objective in derivative pricing | arxiv.org |
| Social Science Research Network | Working papers on model risk measurement, model validation practice and the reporting of numerical error | ssrn.com |
| Black, Scholes and Merton, 1973 | The pricing map the fitted volatility enters, and the closed form of Rs 10.450584/- that the grid ladder is checked against | Journal of Political Economy; Bell Journal of Economics and Management Science |
The five quoted call prices, the five strikes and the process they are fitted to are invented.
Educational material. Not advice on any investment, tax, budget or market position.
