Calibration: Fitting a Model to Observed Prices
Calibration is choosing a model's parameters so that its outputs come as close as possible to a set of numbers taken from outside it. Calibration has five steps: choose the targets, choose the loss, choose the bounds, search, and record what was done. The fifth is the one everyone skips, and without it the answer cannot be reproduced.
Nine reading orders have come before this one, and not one of them fitted a model to anything. Every parameter met so far was simply stated. A volatility of 20 per cent a year, stated. A rate of 5 per cent a year, stated. A speed of pull of 0.5 a year, stated. A stated parameter lets a model be examined for what it implies, and nine reading orders of implications is exactly what this subject area has been. Calibration turns the arrow around and lets a set of prices choose a parameter, and that single change alters what the model may afterwards be used for.
The quantity being modelled is the standard process, written S with a time subscript, starting at an invented Rs 100/- with a rate of 5 per cent a year over a horizon of one year. The numbers being fitted to are five invented call prices at strikes of Rs 80/-, Rs 90/-, Rs 100/-, Rs 110/- and Rs 120/-. The five prices are not an observation of any market anywhere, and they exist for one reason only: so that fitting has something to work on.
Two sentences carry this guide and they need to be held together. The first is that calibration works: a search really does find the setting that comes closest to a set of numbers, and it finds it to as many decimal places as anyone cares to ask for. The second is that what comes back is a fitted number rather than a discovered one, and the moment the model has been shaped by five prices it stops being able to say anything independent about those five prices. Calibration buys agreement, and it pays for that agreement with independence.
What is calibration?
CalibrationChoosing parameters so a model's outputs come close to numbers taken from outside it. is the act of choosing the free numbers inside a model so that what the model produces lines up with numbers that came from somewhere else. Nothing about the model's equations changes. The relationships stay exactly as they were derived, the assumptions stay exactly as they were stated, and the only thing that moves is the handful of values those equations were always going to need before they could produce anything at all.
The everyday version is exactly right rather than roughly right, and it is worth holding on to. A kitchen scale reads a little wrong. Beside it sit five objects that came with their weights printed on them, so those five weights are already known. The small screw underneath the scale is turned until the five readings are as close to the five printed weights as they can be brought, all five at once. The scale has one screw and there are five objects, and one screw cannot satisfy five demands, so the five readings will not all come out exactly right. Turning that screw is calibration. The important part comes afterwards. The scale has been shaped by those five objects, so it can no longer say anything about them.
The direction in which the information flowed is the thing to hold on to. Before the screw was touched, the scale was an instrument with a claim: put a thing on it and it will report what the thing weighs. After the screw was touched, the scale carries a summary of five weights already known. The scale is still useful, and arguably more useful than before. But the question it can answer has changed, and nothing on the outside of the scale records that change. The whole argument sits in one household object.
In the pricing version, the model is a rule that turns one parameter into five prices. The equations were solved long ago and the solution is a formula, so given a volatility the model returns a price at every strike. Calibration runs that rule backwards. Inverting it is generally impossible, so instead values are tried until the five prices the rule produces sit as close as it allows to the five supplied.
| \(C^{\text{model}}\) | the price the model produces, which depends on the strike and on the parameter |
| \(K_i\) | the strike of the \(i\)-th contract, Rs 80/- to Rs 120/- here |
| \(\sigma\) | the volatility parameter, the only free number in this fit |
| \(S_0\) | the starting value of the standard process, Rs 100/- here |
| \(r\) | the rate, 5 per cent a year here, held fixed and not fitted |
| \(T\) | the horizon, one year here |
| \(N\) | the standard normal distribution function, with \(d_1\) and \(d_2\) its two arguments |
So calibration is not a technique for discovering a volatility. Calibration is a technique for choosing one. The distinction sounds like hair-splitting until the next question is asked: what may the answer then be used for? At that question the distinction is everything.
What are the steps, in order?
How to Calibrate a Pricing Model in a Teaching Setting
Calibration is often described as though it were one act, the act of searching. Calibration is five acts, and the search is the fourth of them. Four of the five are decisions made by a person before any arithmetic happens, and each of those decisions changes the answer. The part worth slowing down for is this: three human choices are made before the data is consulted, so the number that comes out at the end is not determined by the data alone.
Step one is choosing the targetsThe numbers being fitted to. There are five of them here, one at each strike.. Which numbers is the model being asked to match? Five here, one at each strike. The target set could have held three prices, or fifty, at one horizon or across many. Every number included gets a vote in where the answer lands, and every number left out gets no vote at all. A target set is a claim about what matters, made before anything is measured.
Step two is choosing the loss functionThe rule that says what close means. The choice of loss is opened in full under the loss function.. The model will not match the targets exactly, so somebody has to say what close means. Close in rupees of price? Close in points of implied volatility? Close after weighting each target by how sensitive its price is to the parameter? All three are defensible, all three are in ordinary use, and all three give different answers from the same five numbers. Squared price error is used throughout this guide; under squared implied volatility error the same five targets return 0.207000 instead. The effect of each choice, and how to choose between them, is set out under the loss function.
Step three is choosing the boundsA limit on where a parameter may be searched. Volatility between 0.05 and 0.50 here.. Where may the parameter be looked for? Volatility between 0.05 and 0.50 here. Bounds are not decoration. Bounds keep a search away from settings that are arithmetically permitted and physically absurd, and bounds can also silently produce an answer sitting on a wall rather than in a valley. A fit that lands exactly on its bound has not found a minimum; it has found the edge of the box it was allowed to look in.
Step four is the searchThe step that finds the parameter, given the targets, the loss and the bounds. The search returns 0.198202 here.. Given the targets, the loss and the bounds, find the setting that makes the loss as small as it goes. The search is the mechanical step, the one a machine does, and the only one of the five that involves no judgement. The search is also the step readers imagine is the whole of calibration, and that is why the other four are worth naming out loud.
Step five is recording what was done. Write down the targets, the loss, the bounds, the answer and the date. All five, together, in one place. Recording is the step that gets skipped, and skipping it is not a documentation failure, it is an arithmetic failure: the number 0.198202 without the four things that produced it is not a parameter, it is a digit string. Somebody handed that digit string cannot check it, cannot reproduce it, and cannot tell whether it disagrees with their own answer because the model is different or because the loss was.
| \(\hat{\sigma}\) | the fitted parameter, carrying a hat to mark it as chosen rather than stated |
| \(L(\sigma)\) | the loss, a single number saying how far the model sits from the targets |
| \(\arg\min\) | the setting at which the loss is smallest, not the smallest loss itself |
| \([\sigma_{\text{lo}},\sigma_{\text{hi}}]\) | the bounds, the window the search is confined to |
Which of these is the step most often left out, and the one without which the answer cannot be reproduced?
What are the targets, the loss and the bounds here?
Now the same five steps on the invented instance, slowly, with every number shown. Step one, the targets. Five call prices on the standard process, all at a horizon of one year, quoted in the way such things are usually quoted: as an implied volatility rather than as a price. The quoted volatilities rise as the strike falls. The earlier reading order on volatility built and named that pattern. Converting each quote back into a price at its own strike gives the five numbers the model will be asked to match.
| Strike | Quoted as a volatility | The target price |
|---|---|---|
| Rs 80/- | 0.240000 | Rs 25.227000/- |
| Rs 90/- | 0.220000 | Rs 17.257579/- |
| Rs 100/- | 0.200000 | Rs 10.450584/- |
| Rs 110/- | 0.190000 | Rs 5.644765/- |
| Rs 120/- | 0.185000 | Rs 2.745149/- |
Read the middle column before going on. The middle column contains the whole difficulty. The five quotes are five different volatilities. The model has one. The target set is telling the model something the model has no way to say back, and every number that follows in this guide is a consequence of that mismatch. A single constant volatility can pass through any one of those five quotes. The five quotes are not the same number, so one constant volatility cannot pass through all five.
Step two, the loss. Squared price error: the gap between each model price and its target is squared, and the five squares are added up. Squaring does two things worth naming. Squaring makes a miss above and a miss below count the same, and the sameness is intended: a model that is dear at one strike and cheap at another has not cancelled its errors out. Squaring also makes a large miss count for far more than a small one, so the search will accept several small misses rather than one big one. The preference for several small misses is a preference, not a fact, and it is imported into the answer along with the data.
| \(L(\sigma)\) | the loss at a given setting of the parameter, one number |
| \(C_i\) | the \(i\)-th target price, taken from outside the model |
| \(C^{\text{model}}\) | the price the model produces at that strike and that parameter |
| \(\sum_{i=1}^{5}\) | added across all five targets, so every target has a vote |
Step three, the bounds. Volatility between 0.05 and 0.50. The lower bound keeps the search away from settings where the model prices collapse toward the discounted payoff and the arithmetic goes numb; the upper bound keeps it away from settings nothing in the target set suggests. Both are judgements. Neither is in the data. If the answer had come back at 0.050000 or at 0.500000 the correct reading would not be that the volatility is one of those numbers, but that the search ran out of room and the bounds need revisiting.
What does the search find, and what does it leave behind?
Step four. Sweep the parameter across the permitted window, compute the loss at every setting, and take the setting where the loss is smallest. The answer is 0.198202. The value 0.198202 is the number used under market price and model price without saying where it came from, and it comes from three choices and one sweep rather than from any observation of a volatility anywhere.
Now the part that matters more than the answer. At 0.198202 the loss is 1.111392. The loss is not nought. One parameter cannot match five numbers, so the best available fit is still a miss at every strike but one, and the search cannot do anything about it. The remaining miss is not a defect in the search, and running a better search will not improve it. There is simply no setting of a single volatility that reproduces five prices carrying five different implied volatilities, and no amount of optimisation creates one.
| \(\nexists\) | there is no such value, anywhere in the permitted window |
| \(\min_{\sigma} L\) | the smallest the loss goes, which is the value at the fitted setting |
| \(1.111392\) | that smallest value here, in squared rupees, an educational illustration |
Here is what the fitted model actually produces at each strike, against what it was asked to produce. Read the last column as the signed miss, model less target, so a negative entry means the model is cheap there and a positive one means it is dear.
| Strike | Target | Fitted model at 0.198202 | The miss |
|---|---|---|---|
| Rs 80/- | Rs 25.227000/- | Rs 24.564539/- | minus Rs 0.662462/- |
| Rs 90/- | Rs 17.257579/- | Rs 16.650741/- | minus Rs 0.606838/- |
| Rs 100/- | Rs 10.450584/- | Rs 10.383131/- | minus Rs 0.067452/- |
| Rs 110/- | Rs 5.644765/- | Rs 5.968944/- | plus Rs 0.324179/- |
| Rs 120/- | Rs 2.745149/- | Rs 3.186331/- | plus Rs 0.441183/- |
| All five | squared and added | 1.111392 | root mean squared 0.471464 |
The pattern in that last column is worth a moment. The misses are not scattered. The five run cheap, cheap, almost exact, dear, dear, in strike order. A single constant volatility is a straight line laid across a curve, and a straight line laid across a curve leaves its residuals in exactly that shape: under at both ends of one side, over at the other. The residuals are not noise. The residuals are the shape of what the model cannot say.
| \(\text{RMSE}_{\text{cal}}\) | the root mean squared error (RMSE) at the fitted targets, a price miss in rupees |
| \(\hat{\sigma}\) | the fitted parameter, 0.198202 here |
| \(\tfrac{1}{5}\) | averaged across the five targets, so the figure does not grow with the count |
The volatility is about to be swept across the whole window to find the best fit. Before it moves: what is the loss at the best setting?
Sweep the parameter and watch the misses refuse to vanish
Held fixed: the standard process starting at Rs 100/-, a rate of 5 per cent a year, a horizon of one year, the same five strikes from Rs 80/- to Rs 120/-, and squared price error as the loss. The only thing that moves is the volatility. The upper panel is the total miss across the five targets, drawn as a root mean squared price error in rupees so that it reads in the same units as the prices. The lower panel is the five individual misses, model less target, and it redraws with every move. Five settings worth holding on to: at 0.185000 the loss is 1.953633, at 0.190000 it is 1.439128, at 0.198202 it is 1.111392, at 0.207000 it is 1.498165 and at 0.220000 it is 3.527621. The minimum sits at 0.198202 and the loss there is 1.111392 rather than nought.
At a volatility of 0.198202 the squared price error across the five targets is 1.111392 and the root mean squared price miss is Rs 0.471464/-. Three of the five model prices sit below their target and two sit above. The loss is at the lowest value it reaches anywhere in the window, and that lowest value is 1.111392 rather than nought.
At the fitted setting, what is the calibration error across the five targets?
What does the record have to contain?
Step five, and it is short only because it is simple, not because it is small. The record is five lines. The targets, listed, not described. The loss, named. The bounds, stated. The answer, to the precision it was computed at. The date it was run. Anything less and the number is stranded.
The harsh test is this. Somebody handed 0.198202 and nothing else, and asked to check it, cannot. The recipient does not know which five prices were used, so the targets cannot be rebuilt. Nor is it known whether price error or volatility error was minimised, so even with the targets the recipient would land somewhere else. The window is unknown too, so there is no telling whether the answer was a minimum or a wall. Every one of the four missing facts is enough on its own to make the number unreproducible, and reproducibility is the only property that separates a fitted parameter from an assertion.
The date belongs on the record for a reason that only shows up later. A fit is a fit to a set of numbers as they stood on a day. Rerun it against a set gathered later and the answer moves, and without a date on both records there is no way to tell whether the two answers disagree because the method differed or because the days did.
Does fitting a model to a set of numbers make it more useful, or less useful, as evidence about those numbers?
What does calibration change about a model?
Calibration changes what a model may be used for. The statement is bigger than it sounds, and it is worth taking apart.
Before the fit, a model is a set of assumptions plus a set of consequences. Assume the process moves this way, assume no arbitrage, assume the rate is that, and here is what a contract on it must be worth. Every price the model produces is a deduction. Disagreement with the price requires disagreement with one of the assumptions, and the model names exactly which ones are on the table. An unfitted model is an argument, and an argument can be challenged at its premises.
After the fit, the model still contains all those assumptions, but one of its numbers no longer comes from them. The fitted number came from five prices. Feed the model a strike and it now returns something shaped partly by the equations and partly by what it was shown, and there is no line on the output separating the two contributions. The model has become, in part, a summary of the target set.
Both objects are useful. The fitted model is often the more useful of the two: it now agrees with the numbers a reader recognises, and it can extend to strikes and horizons the target set did not cover. Usefulness has not gone. Independence has. Before the fit the model's agreement with a price would have been a finding; after the fit its agreement with that same price is a design feature.
Back to the kitchen scale for a second: the same three sentences apply to it. Before the screw was turned, the scale was an instrument with a claim, and if it read one of the five objects correctly that was mild evidence the scale worked. After the screw was turned, the scale reads all five about right by construction, and reading them about right is evidence of nothing except that the screw was turned carefully. The scale is more useful and less informative, at the same time, for the same reason.
A model has been fitted to a set of five prices. What can it no longer be used as?
What does a fitted model no longer report?
A fitted model no longer reports anything about its own targets. Asked what the price at Rs 100/- should be, a fitted model hands back Rs 10.383131/- against a target of Rs 10.450584/-, and the closeness of those two numbers is a report on the search, not a report on the world. The parameter was chosen to make that gap small. Finding it small afterwards is finding what was put there.
A genuinely careful reader can go wrong here: the check feels like a test, and it has the shape of one. There is a prediction, there is an observation, and the two are compared. The possibility of failure is missing. A test that could not have come out any other way is not a test, and no amount of care in performing it converts it into one.
The error that gets made, and what it costs
Using a calibrated model as evidence about the very prices it was calibrated to. Somebody fits the model to the five targets, then reprices the five targets, then reports that the model reproduces them to within Rs 0.471464/- on average and concludes that the model is sound. Every step of that is arithmetically correct and the conclusion does not follow from any of it.
The reasoning breaks at a place that is easy to miss: the first half is fine. The fit is fine. The repricing is fine. The number 0.471464 is correct. The word sound is the part that is not fine: the model has been shaped by those five numbers, so agreement with them measures how well the search ran, and the search ran well. Had the model been badly chosen, the search would still have found the best available setting, and the check would still have passed, merely with a larger residual. The check has no failing state that means what the reader thinks it means.
The cost is a circular check presented as a test, and it is a particularly expensive kind of mistake because it is self-reinforcing. The circular check never raises an alarm, so nobody investigates, and it always passes, so it accumulates confidence. And it is reported in the same words and the same units as a real test, so a reader downstream cannot tell the two apart from the output. The only defence is knowing which numbers the parameter was chosen against, and step five records precisely that.
A fitted model reprices its own five targets closely. What has that established?
How is calibration different from validation?
Calibration vs Validation
The two words get used as though they were the same activity performed with different levels of formality. Calibration and validation are not the same activity. The two differ in one thing, and that one thing decides everything else about them: the comparison each makes.
ValidationChecking a model against something it was not fitted to. A different act from fitting, with a different question behind it. compares the model to something it has not been shaped by. Calibration compares the model to the set it is being shaped by, over and over, until the comparison is as good as it gets. Same arithmetic, same units, same appearance in a report. Opposite epistemic standing.
The consequence is about failure. A calibration cannot fail informatively. Whatever the model is like, the search returns the best available setting and the residual is whatever it is. There is no threshold at which the calibration says stop. A validation can fail. The model meets a number it was not built to accommodate and either handles it or does not. Only a check that has a way of coming out badly says anything when it comes out well.
The kitchen scale again, and this time it settles the matter in one move. Calibration is turning the screw until the five known objects read right. Validation is putting a sixth object on the scale, one whose weight is known and was never used, and seeing what it says. The first cannot fail: whatever is done with the screw, some position is the best position, and that is where it is left. The second can fail, loudly, and that is why it is worth doing.
What is the difference between calibration and validation?
What is calibration error, and how does it differ from pricing error?
Calibration Error vs Pricing Error
Calibration errorThe miss at the fitted targets. Here a root mean squared price error of 0.471464 across the five. is the miss that remains where the model was fitted. Here it is 0.471464, a root mean squared price error across the five strikes. Calibration error is a statement about how much of the target set the model could absorb, and it is bounded below by the model's own structure: with one parameter and five targets carrying five different implied volatilities, it could never have been nought.
Pricing errorThe miss on something the model was not fitted to. A different quantity, computed on a different set. is the miss on something else. A strike outside the five. A horizon the fit never saw. A contract of a different kind altogether. Pricing error is a statement about whether the model travels, and it is not bounded below by anything the fit did: a model can absorb its targets almost perfectly and still be badly wrong one strike outside them.
| \(\{K_j\}\) | the strikes the miss is measured on now, none of which was fitted |
| \(\{K_i\}\) | the five strikes the parameter was chosen against |
| \(m\) | how many contracts are in the second set |
| \(\varnothing\) | the empty set, so the two collections share nothing |
Now the reason one number cannot report both. Averaged together they give a figure in rupees that is partly a report on the search and partly a report on the model, in a proportion decided by how many contracts happen to sit in each set. Move one contract from the fitted set to the unfitted set and the combined figure changes, without anything about the model having changed at all. A quantity that moves when the bookkeeping is rearranged is not measuring a property of the model.
And the two failure modes point in opposite directions. A large calibration error with a small pricing error says the target set was internally awkward for this model but the model travels anyway. A small calibration error with a large pricing error says the model absorbed its targets and learned nothing that holds outside them. The second combination is the more dangerous of the two, and a single blended number is best at hiding it.
Why can one number not report both calibration error and pricing error?
What does a reader ask when handed a fitted number?
Most readers will never run a calibration. Nearly every reader will at some point be handed the output of one, in a note, a report or a spreadsheet cell, and asked to work with it. The five steps turn straight into five questions. An unsatisfactory answer to an early one makes the later ones moot, so the questions are worth asking in order.
Five questions, in the order the steps were taken
- Which numbers was this fitted to? If the answer is vague, nothing downstream can be relied on. Everything downstream is a statement about a set that cannot be seen, and there is no way to tell what the parameter has absorbed.
- What was minimised? Price error, volatility error and vega-weighted price error give three different answers from one set of numbers. Two people quoting different fitted volatilities may be in complete agreement about the world and disagreeing only about step two.
- What were the bounds, and did the answer land on one? An answer sitting exactly on a bound is a signal to reopen the window, not a result.
- What is the miss at the fitted targets, and is it reported separately from any other miss? If one figure covers both, ask for them split before reading either.
- What is this number now being used for? Pricing something the fit did not see is a legitimate use with a known risk. Offering the fit as evidence about the prices it was fitted to is not a use at all.
The fifth question is the one that catches the most. A fitted parameter is a serviceable input to the next calculation and it is never a finding. The moment somebody offers a calibrated model's agreement with its own targets as support for a conclusion, the conversation has gone in a circle and the number at the centre of it is doing no work.
Where this holds
Calibration is not jurisdictional. The arithmetic of choosing parameters to minimise a stated loss over a stated window is the same wherever it is written down, and the five steps are craft rather than anybody's mandated procedure. The way a contract is quoted, the day count used to turn a horizon into a fraction of a year, and the conventions that decide what a target price even means do vary by place, and those are set out under market convention. A convention of that kind is a rule of the market that writes it, and it is stated there rather than derived from any account of fitting.
References
| Source | Document | Where |
|---|---|---|
| arXiv, Quantitative Finance | Preprints on model calibration, inverse problems and parameter estimation in derivative pricing | arxiv.org |
| Social Science Research Network | Working papers on calibration practice, loss function choice and model risk | ssrn.com |
| Black, Scholes and Merton, 1973 | The option pricing papers that give the map from one volatility to a price at every strike, the map the fitted parameter enters | Journal of Political Economy; Bell Journal of Economics and Management Science |
The standard process, its five strikes and the five quoted volatilities on them are invented.
Educational material. Not advice on any investment, tax, budget or market position.
