How to Document a Pricing Model: Assumptions and Review
A pricing model has to be recorded in four parts: what it assumes, what its settings are, what it was fitted to, and what checking it survived. Each part answers a different question and none of them substitutes for another. A document carrying three of the four is not three quarters useful. The missing part makes the document unusable for the one purpose documentation has.
The purpose of documentation is worth naming before anything else. Documentation is adequateWhether someone else can reproduce the number from the document alone, without asking a single question. when someone else can reproduce the number from it alone. Not recognise the number. Not agree that it looks reasonable. Reproduce it, starting from the document and nothing else, with no access to the person who produced it and no access to whatever was on their screen at the time.
Most documents are written to pass the test of existing, and reproduction is a far harder test than that. A document exists the moment somebody saves a file. A document is adequate only when a stranger can rebuild the number out of it. Existing and being adequate are separated by a very large distance, and almost everything in this guide lives inside that distance.
The everyday version is worth holding on to. Somebody weighs a sack of rice on a shop scale and writes 48.6 on a slip of paper. A month later the slip turns up. The number on the slip is plain. Nobody can now tell whether the scale was in kilograms or in some older local unit, whether it had been set to zero before the sack went on, whether the sack was weighed with the jute sacking or without it, or whether anybody ever put a known weight on that scale to see what it said. The number survived. Everything that would let anyone use the number did not. A pricing model is that slip of paper unless four specific things are written beside it.
What has to be recorded about a pricing model?
Four sections, and they are not four subdivisions of one idea. The four sections are four answers to four questions that happen to be about the same model. The four questions are what the model supposed, what values it was run with, what observations it was aimed at, and what testing it came through. Asking any one of those questions of a document that lacks its section produces silence, and silence in a document is indistinguishable from a claim that there was nothing to say.
The first section, assumptionsWhat the model supposes about the world, which the settings do not record anywhere., records what the model supposes about the world. The second, settingsThe parameter values the model was actually run with, at the precision it was run at., records the values it was run with. The third, calibration inputsWhat the model was fitted to, together with the loss, the bounds and the date of the fitting., records what it was fitted to, along with how the fitting was scored, what range it was allowed to search and when it happened. The fourth, the validation reviewWhat checking the model survived, including its behaviour on observations it was never fitted to., records what checking it survived, including the part where somebody deliberately tried it on observations it had never been shown.
The overlap between the four is very small. Nothing in the assumptions section names a single number the model was run with. Nothing in the settings section states what the model supposes. Neither one records what data the fitting aimed at, and none of the first three shows whether anybody ever checked. Four sections, four disjoint contents. Four disjoint contents are why a missing section cannot be inferred from the others, and why the count of sections present, rather than the length of the document, decides whether the document works.
What are the four sections that have to be recorded beside a pricing model?
What goes in the assumptions section?
Everything the model supposes about the world, written as claims rather than as values. An assumption is a statement that could in principle be false; a setting is a number. That distinction sounds pedantic right up to the moment a document arrives with only the numbers in it, at which point it becomes the whole difference between a model that can be judged and a model that can merely be operated.
In the worked instance this guide documents throughout, the model supposes that the level follows the standard process, an assumption about the shape of the randomness. The model also supposes that volatility is constant, an assumption about whether one number is enough. The same list supposes that the rate is constant at 5 per cent, that there are no costs of any kind, and that the horizon is one year. Five statements, and every one of them is something a reader could disagree with, test, or find contradicted by the very observations recorded in the third section.
The settings section says this much about the same model: volatility 0.198202. One line of six digits does not record that constant volatility was supposed. The line is perfectly consistent with a model in which volatility moves and 0.198202 happened to be its starting level, or with a model in which volatility differs by strike and 0.198202 is some average across them. Constant volatility is an assumption and 0.198202 is a setting, and a document recording only the second has hidden the first.
The everyday version again. Somebody reports that a journey took two hours. Two hours is a setting. The report leaves out whether the road was assumed open the whole way, whether the stop for tea was counted, and whether two hours was the outward trip or an average of both directions. Every one of those is an assumption, none is recoverable from the number two, and every one of them changes what the number two is good for.
A document states that volatility does not change over the horizon. Is that an assumption or a setting?
Before the next section. Given the settings section and nothing else, how many of the four questions can be answered from it?
What goes in the settings section?
Every value the model was actually run with, at the precision it was run at. The settings section is the only one of the four that a machine could write unaided, and that is exactly why it is the one people mistake for the whole document. It is easy, it is unambiguous, and it feels like the real content because it contains the numbers.
For the worked instance the settings are five lines: a starting value of Rs 100/-, a volatility of 0.198202, a rate of 0.050000, a horizon of 1.000000 and continuous compounding. The precision is part of the setting, so six decimal places rather than two. A reader who rounds the volatility to 0.20 before rerunning the model is not rerunning the same model, and the difference will show up in the fourth decimal place of every price that comes out.
The five values feed one pricing function, and the document has to name which function. Black, Scholes and Merton, 1973, gave their names to the one used here. Naming the function is part of the settings section because two documents with identical parameter lists and different functions produce different numbers, and nothing in a bare list of parameters reveals which of the two is in hand.
| \(S_0\) | the starting value of the standard process, from the settings section |
| \(K\) | the strike of the contract being valued, from the calibration inputs section |
| \(\sigma\) | the volatility setting, recorded here as 0.198202 |
| \(r\) | the rate setting, recorded here as 0.050000 |
| \(T\) | the horizon setting, recorded here as 1.000000 |
| \(N\) | the standard normal distribution function |
Read that block carefully and the limitation of the settings section becomes visible in the term table itself. Five of the six symbols come from the settings. One does not. What the model was aimed at is not one of the settings section's symbols, so that section is complete enough to run the model and structurally incapable of saying what it was aimed at.
What goes in the calibration inputs section?
Four items go in: what the model was fitted to, how the fitting was scored, what range the search was confined to, and when it happened. The fourth of them surprises people. A fitted number is only meaningful with respect to when it was fitted, so the date belongs with the calibration inputs and not with the settings.
Start with what it was fitted to. The worked instance was fitted to five invented observations on the standard process, all at a one year horizon, quoted as implied volatilities and converted to prices. Every one of the five is an educational illustration and none is an observation of any real market.
| Strike | Quoted as volatility | Observed price used in the fitting |
|---|---|---|
| Rs 80/- | 0.240000 | Rs 25.227000/- |
| Rs 90/- | 0.220000 | Rs 17.257579/- |
| Rs 100/- | 0.200000 | Rs 10.450584/- |
| Rs 110/- | 0.190000 | Rs 5.644765/- |
| Rs 120/- | 0.185000 | Rs 2.745149/- |
Then how it was scored. The fitting minimised squared price error, one choice among several, and that choice has to be written down rather than assumed. Then the range: the search was confined to volatilities between 0.05 and 0.50. Then the date: 23 August 2026. The four items together turn the number 0.198202 from a value the reader is told into a value the reader can derive.
| \(\hat{\sigma}\) | the fitted volatility, which is the number the settings section records as 0.198202 |
| \(C^{\text{obs}}_{i}\) | the observed price at the i-th strike, from the table above, invented |
| \(K_i\) | the i-th strike, from Rs 80/- to Rs 120/- |
| \(n\) | how many observations the fitting used, five here |
| \([\sigma_{lo},\sigma_{hi}]\) | the bounds the search was confined to, 0.05 and 0.50 here |
Now the date. The settings are where the numbers live and the date is a number, so filing the date with the settings is tempting. The instinct is wrong, and the reason is worth stating plainly. The settings are true of the run: they say what went in. The date is true of the fitting: it says what the observations were observations of, and therefore what the fitted number is a fitted number of. Nobody needs a record of when somebody pressed a button, and everybody needs a record of when the fitted value was anchored, so moving the date into the settings section trades the second away for the first.
The household version. A tailor writes a customer's measurements in a book. The measurements are settings. The date beside them is not a note about when the pen touched the paper; it is the only thing telling a later reader whether these are the measurements of a person from this year or from six years ago. Strip the date and the measurements are still numbers. The numbers are just no longer measurements of anybody in particular.
The document records that the fitting was carried out on 23 August 2026. Which section does that date belong in?
What goes in the validation review?
The fourth section records what checking the model survived, and in particular how it behaved on observations it was never fitted to. The validation review is the only section of the four whose figures were not chosen by the person doing the fitting, and so the only section containing a number that can embarrass the model.
Here is what that means for the worked instance. Somebody refits the same procedure to the middle three strikes alone, Rs 90/-, Rs 100/- and Rs 110/-, deliberately holding back Rs 80/- and Rs 120/-. The refit answer is 0.199868. Then the gaps between model price and observed price are measured in two places: across the three strikes the refit was shown, and across the two it was not.
| \(\mathcal{S}\) | a set of strikes the error is measured over, either the three shown or the two held back |
| \(|\mathcal{S}|\) | how many strikes are in that set, three or two here |
| \(\hat{\sigma}\) | the refit volatility, 0.199868, fitted on the three strikes only |
| \(C^{\text{obs}}_{i}\) | the observed price at that strike, invented, unchanged from the table above |
The two readings are 0.394853 inside the fitted set and 0.573319 outside it. The model is 45.20 per cent worse where it had been shown nothing, and that ratio is the single most informative figure in the entire document. It does not say the model is bad. The ratio says how much of the apparent fit was the model describing the observations and how much was the model being aimed at them.
| Strike | Shown to the refit? | Observed | Model at 0.199868 | Gap |
|---|---|---|---|---|
| Rs 80/- | held back | Rs 25.227000/- | Rs 24.587038/- | minus 0.639963 |
| Rs 90/- | shown | Rs 17.257579/- | Rs 16.695864/- | minus 0.561716 |
| Rs 100/- | shown | Rs 10.450584/- | Rs 10.445630/- | minus 0.004953 |
| Rs 110/- | shown | Rs 5.644765/- | Rs 6.034864/- | plus 0.390100 |
| Rs 120/- | held back | Rs 2.745149/- | Rs 3.242980/- | plus 0.497832 |
| Root mean squared | shown, three | 0.394853 | ||
| Root mean squared | held back, two | 0.573319 |
Look at the middle row before anything else. At the Rs 100/- strike the gap is 0.004953, close enough to nothing that a reader shown only that row would conclude the model was very nearly exact. Now look at the two held back rows, where the gaps are 0.639963 and 0.497832. In-sample errorThe miss measured where the model was fitted, which is 0.394853 in this worked instance. and out-of-sample errorThe miss measured where the model was never fitted, which is 0.573319 in this worked instance. are different questions asked of the same model, and reporting only the first is how a document flatters itself without stating anything false.
Which of the four sections contains a number capable of embarrassing the model?
Before touching the control below. The settings section alone records a volatility of 0.198202. Can that number be reproduced from it?
Remove the sections one at a time and watch what stops being answerable
Held fixed: the model, the five invented observations, and the fitted volatility of 0.198202. One thing moves, and it is how many of the four sections the document actually contains. The settings are the part that writes itself, so the sections arrive starting from the settings, in the order a thin document usually grows. Every count below is read off the sections themselves and never sampled, so the reading is identical on every visit.
The static readings, for a reader with no browser. Settings alone: one of four questions answerable, and 0.198202 cannot be reproduced. Settings and assumptions: two of four, and 0.198202 still cannot be reproduced. Settings, assumptions and the calibration inputs: three of four, and 0.198202 becomes reproducible for the first time. All four sections: four of four, and the review figures of 0.394853 and 0.573319 become available as well. The starting value is Rs 100/- and the middle observed price is Rs 10.450584/- throughout.
What makes documentation adequate rather than merely present?
One test, and it is not a checklist. Documentation is adequate when someone else can reproduce the number from it alone, and every other property people ask of a document is either a means to that end or beside the point. Length is beside the point. Tidiness is beside the point. Having a section for each heading is a means, and a poor one. A heading with nothing usable underneath it passes a checklist and fails the test.
Notice what the test demands that a checklist does not. The test demands a specific other person, one who does not have the author to ask. The second demand is a specific output, the number itself rather than an impression of the number. And it demands independence, meaning the reader starts from the document and finishes at the number with nothing borrowed in between. ReproducibilityThe only test of documentation that matters: can a stranger rebuild the number from the document and nothing else. is the whole of it.
| \(\hat{\sigma}\) | the value the reader arrives at by carrying out the fitting themselves |
| \(\stackrel{!}{=}\) | a requirement rather than an observation, meaning the two sides are required to agree |
| 0.198202 | the value the document already claims, invented, fitted 23 August 2026 |
The loop is why the third section is the pivot. Sections one and two tell the reader what to do. Section three tells the reader what to do it to. Neither half closes the loop alone. A document with settings and assumptions and no calibration inputs contains everything except the thing the number was aimed at, and that is precisely the missing half of a reproduction.
And the fourth section answers a different question again, not whether the number can be rebuilt but whether it is worth rebuilding. A model can be perfectly reproducible and still 45.20 per cent worse wherever nobody looked. Reproducibility makes the number checkable; the review is what somebody found when they checked.
What is the test of adequate documentation?
What does the whole document look like on one sheet?
Here is the worked instance in full, all four sections, on the single model this guide has been documenting throughout. Read it as one object rather than four, and notice that the sheet is short: adequacy is not a matter of volume. Every value on it is an educational illustration and none is an observation of any real market.
How does a reviewer actually work through the four sections?
Not in the order they are written. The fourth section is the only one that can quickly tell a reader whether the rest is worth their afternoon, so somebody picking up a model document to judge it reads the fourth section first. That is a habit worth borrowing, and it changes what goes into each section once the order of reading is known.
The working order goes like this. First the validation review, to see whether anybody tested the model on something it had not been fitted to, and by how much it got worse when they did. A review reporting only the in-sample figure of 0.394853 tells the reviewer that the model was scored on its own homework, and that is enough to change the tone of every remaining question. Second the assumptions. A reviewer wants to know what was supposed before looking at any number that was supposed under it. Third the calibration inputs, to see what the number was aimed at and whether five observations is a thin or a reasonable basis for the claim being made. By then the reviewer already knows what they are looking at, so the settings come last and take about a minute.
The reading order explains something about how the sections should be written. The validation review is read cold, by somebody with no context, in a hurry, deciding whether to continue. The validation review has to carry its own comparison inside it. Not 0.573319 on its own, which means nothing to a stranger, but 0.573319 against 0.394853, a claim any reader can act on immediately. A figure without the figure it should be compared against is not a finding, it is a decoration with decimal places.
The everyday version, and it is the one people actually live. A buyer looking at a used scooter does not start with the paperwork. One answer decides how much weight to put on everything the seller wrote down, so the buyer starts with whether anyone independent has looked at the engine. The service record is the validation review. The specification sheet is the settings, and it is the part that took the least effort to produce and gets read last.
The error that gets made, and what it costs
Recording the settings and calling it documentation. The settings section is the easiest to write, it looks the most like real content because it holds the numbers, and a document containing it alone can be circulated without anybody objecting for months. The settings say what the model was run with and say nothing at all about what it assumed, what it was fitted to or what survived checking.
A reader handed only the settings can rerun the model and can do nothing else with it. Such a reader cannot tell whether constant volatility was supposed or merely happened. Nor can they tell whether 0.198202 was fitted to five observations or to five hundred. And they cannot tell whether anybody ever pointed the model at something it had not been shown, the one test that would have surfaced the 45.20 per cent gap between 0.394853 and 0.573319.
The cost is a document that satisfies the requirement to have one and fails the only test that matters. Nothing in it is false. Nothing in it is checkable either. The failure only surfaces on the day somebody asks a question the document was never built to answer, so those two properties can coexist quietly for a very long time.
A document records only the settings section. What can a reader actually do with it?
References
| Source | Document | Where |
|---|---|---|
| arXiv, Quantitative Finance | Preprints on calibration practice and out-of-sample testing of pricing models | arxiv.org |
| Social Science Research Network | Working papers on model risk and the recording of fitted parameters | ssrn.com |
| Black and Scholes, 1973; Merton, 1973 | The papers giving the pricing function the settings are fed into | Journal of Political Economy; Bell Journal of Economics and Management Science |
The standard process and the five observations it is fitted to are invented.
Educational material. Not advice on any investment, tax, budget or market position.
