The Calibration Loss Function: What You Are Minimising
A calibration loss function is the rule that says what it means for a model to be close to its targets. The person doing the fitting chooses it, nothing in the observed prices selects it, and it decides the answer as surely as the targets do. Three defensible rules on one set of five prices return three different volatilities.
A fit needs three things before it can run. The first is a set of targets. The second is a model that can produce a number comparable with each target. The third is a rule saying which of two near misses is the better one. A model with one free number and five targets is never going to hit all five, so something has to decide where the search stops. The first two are always written down. The third decides the answer and almost never appears in the record.
The quantity being modelled is the standard process, written S with a time subscript, starting at Rs 100/-, over a horizon of one year, with a rate of 5 per cent a year. The targets are five invented observed prices at strikes of Rs 80/-, Rs 90/-, Rs 100/-, Rs 110/- and Rs 120/-, quoted as implied volatilities of 0.240000, 0.220000, 0.200000, 0.190000 and 0.185000.
The five prices never move. Fitting one constant volatility to them under three rules that a careful person could each defend returns 0.198202, 0.204420 and 0.207000. The gap between the outer two is 0.008798, close to a full point of volatility. Nothing observable moved to produce that gap, and nothing observable chooses between the three answers.
What is a calibration loss function?
The case sits away from any market entirely: one thermostat and five rooms. The rooms are at different temperatures and the aim is to have every one of them at twenty degrees. There is a single dial. Whatever it is set to, at most one room lands where it was meant to, and the other four sit above or below. So the question is no longer what the right setting is. The question is what the best setting means, and there is more than one honest answer. One aim is for the total discomfort across all five rooms to be smallest. Another is for the worst room to be as good as possible. A third disregards the storeroom entirely. Each of those is a different rule, each is defensible, and each turns the dial to a different place.
A loss functionThe rule saying what it means for a model to be close to its targets, chosen by the person fitting rather than read off anything. is that rule, written down. The loss function takes a candidate parameter value, works out what the model says under it, compares the result against every target, and returns one number saying how bad the whole fit is. The fit then searches for the parameter that makes that one number as small as it goes. Notice what has happened: five separate comparisons have been collapsed into one number, and a collapse of five things into one is a choice about relative importance whether or not anybody made it deliberately.
| \(\mathcal{L}(\sigma)\) | the loss at a candidate volatility, one number summarising how far the model sits from all five targets at once |
| \(\sigma\) | the single free parameter being fitted, a constant volatility applied to every strike |
| \(K_i\) | the strike of the i-th target, running over Rs 80/-, Rs 90/-, Rs 100/-, Rs 110/- and Rs 120/- |
| \(C(K_i,\sigma)\) | the model price at that strike under the candidate volatility, computed from the pricing formula |
| \(C_i^{\ast}\) | the observed price at that strike, fixed throughout, invented for teaching |
| \(w_i\) | the weight placed on the i-th squared gap, and the part nobody observes |
| \(\hat{\sigma}\) | the fitted volatility, meaning the candidate at which the loss is smallest |
The definition above already contains the whole difficulty, and it is worth pausing on the shape of it rather than the notation. Two of the three inputs to a fit are things that are found. The targets are found, the model is chosen from a small number of named alternatives and then stated. The third input is manufactured on the spot, usually without comment, often by whoever wrote the fitting routine rather than by whoever will read its output. The loss function is the only input to a calibration that is manufactured rather than found.
Is the loss function an input to the fit, or a detail of the method?
What are the usual choices?
Three rules come up constantly, and all three are worked below on one set of numbers. The first is squared price errorAdding up the squared differences between model price and observed price, with every strike counted the same way. It returns 0.198202 on these five prices.: compare prices to prices, count every strike the same way, add the squares. Squared price error is the plainest rule a person can write, and it needs no extra machinery.
The second is vega weightingDividing each price gap by that strike's vega before squaring it, which restates every gap in volatility points rather than rupees. It returns 0.204420 on these five prices.. The same move in volatility produces a different move in price at each strike, so a gap of Rs 0.50/- does not mean the same thing everywhere. So divide each price gap by that strike's vega first, and square only afterwards. The division restates every gap in volatility points. The precise form is the part that most often goes unrecorded, so the name is worth stating exactly. The weight on each squared price gap is one divided by that strike's vega squared.
The third is squared implied volatility errorComparing volatilities to volatilities directly and forgetting prices, which counts every strike the same way. It returns 0.207000 on these five prices.. The targets arrived quoted as volatilities in the first place, so compare volatilities directly and leave prices out of it entirely. Squared implied volatility error is the only one of the three that never evaluates the pricing formula at all.
| \(\mathcal{L}_{P}\) | the squared price error, meaning every weight set to one, minimised at 0.198202 |
| \(\mathcal{L}_{V}\) | the vega weighted squared price error, minimised at 0.204420 |
| \(\mathcal{L}_{I}\) | the squared implied volatility error, minimised at 0.207000 |
| \(\nu_i\) | the vega at the i-th strike evaluated at that strike's own observed volatility, being 18.080359, 28.592755, 37.524035, 39.485723 and 32.862160 per one unit of volatility |
| \(\sigma_i^{\ast}\) | the observed implied volatility at the i-th strike, being 0.240000, 0.220000, 0.200000, 0.190000 and 0.185000 |
| The rule | What it compares | Weight on each squared price gap | The answer |
|---|---|---|---|
| Squared price error | Rupees against rupees | One, at every strike | 0.198202 |
| Vega weighted squared price error | Rupees restated as volatility points | One over that strike's vega squared | 0.204420 |
| Squared implied volatility error | Volatilities against volatilities | No price gap appears at all | 0.207000 |
| The spread between the outer two | Nothing observable changed | The rule changed, and only the rule | 0.008798 |
Which three rules are the same five prices fitted under?
Why does the choice change the answer?
At this point an intuition that sounds right turns out to be wrong, and the mistake is worth walking into fully. The plain squared price error contains no weights at all, so it looks like the neutral rule, the one with no opinion. And because the biggest prices sit at the low strikes, at Rs 25.227000/- against Rs 2.745149/- at the top strike, it looks as though those big numbers must dominate a sum of squared rupee gaps and drag the answer toward whatever fits them. Both halves of that reasoning are wrong, and the arithmetic says so plainly.
Differentiate the loss and look at what appears. Setting the derivative to nought is the condition the fitted value satisfies, and the term that shows up beside every residual is the vega at that strike. Vega was never asked for. Vega arrives because the model price is what is being moved, and how much the model price moves when the volatility moves is the definition of vega.
| \(\nu(K_i,\hat{\sigma})\) | the vega at the i-th strike evaluated at the fitted volatility, meaning how much that strike's model price moves for a move in the volatility |
| \(w_i\) | the stated weight, set to one at every strike for the plain squared price error |
| \(C(K_i,\hat{\sigma})-C_i^{\ast}\) | the residual at the i-th strike, meaning the model price at the fitted volatility less the observed price |
Now look at where vega is large and where it is small on these five strikes. Vega is largest where the log of the ratio of starting value to strike sits near the middle of the distribution and falls away on either side of that. On this invented set it runs 18.080359, 28.592755, 37.524035, 39.485723 and 32.862160 per one unit of volatility, so it peaks near the Rs 110/- strike and is smallest by a wide margin at Rs 80/-. The strike carrying the largest price carries the smallest vega, and those two facts point in opposite directions.
| \(S_0\) | the starting value of the standard process, Rs 100/- here |
| \(\phi\) | the standard normal density, peaking at nought and falling away quickly on both sides |
| \(d_1\) | the standardised distance, being 1.258098, 0.816184, 0.350000, minus 0.143475 and minus 0.622752 at the five strikes under their own observed volatilities |
| \(r\) | the rate, 5 per cent a year here |
| \(T\) | the horizon, one year here |
Work the exact weights out and the intuition inverts completely. Take the share of the total of the five observed prices that each strike carries. The wrong story says that share should drive the fit. Then take the exact share of the fitted answer that each strike actually commands, computed further below. The Rs 80/- strike carries 41.1 per cent of the total price and 4.2 per cent of the fitted answer. The Rs 120/- strike carries 4.5 per cent of the total price and 22.4 per cent of the fitted answer. The squared price error does not weight the largest prices; it very nearly discards them.
So the plain price loss returns the lowest of the three answers, and the reason is the opposite of the tempting one. The Rs 80/- strike carries the highest observed volatility on this set at 0.240000. Moving the volatility barely moves the price at Rs 80/-, so the price loss listens to that strike least of all and has almost nothing to gain by trying. The highest reading is nearly discarded, the answer falls, and the fit ends at 0.198202. The largest miss on the whole fit, at minus Rs 0.662462/-, sits at exactly the strike the wrong story claims the fit is being dragged toward.
Why does squared price error return the lowest of the three answers on this set?
What does each choice implicitly weight?
There is a way of stating all three answers that makes them comparable, and it removes the last of the mystery. Every one of the three fitted volatilities is a weighted average of the same five observed implied volatilities. Not approximately, and not as a rough description. Exactly, to every decimal place. The three rules differ only in the weights they put on those five readings, and once the weights are written down there is nothing left to explain.
| \(W_i\) | the effective weight the i-th reading carries in the answer, and not the stated weight |
| \(\bar{\nu}_i\) | the average slope of price against volatility between that strike's own reading and the fitted answer |
| \(\nu(K_i,\hat{\sigma})\) | the vega at that strike evaluated at the fitted answer |
| \(\sigma_i^{\ast}\) | the observed implied volatility at the i-th strike, unchanged from first line to last |
| Strike | Observed volatility | Weight under price error | Weight under vega weighting | Weight under volatility error |
|---|---|---|---|---|
| Rs 80/- | 0.240000 | 0.041890 | 0.148766 | 0.200000 |
| Rs 90/- | 0.220000 | 0.148362 | 0.199854 | 0.200000 |
| Rs 100/- | 0.200000 | 0.277557 | 0.211964 | 0.200000 |
| Rs 110/- | 0.190000 | 0.308446 | 0.212630 | 0.200000 |
| Rs 120/- | 0.185000 | 0.223744 | 0.226786 | 0.200000 |
| Weighted average of the five | range 0.185 to 0.240 | 0.198202 | 0.204420 | 0.207000 |
Read the third column down and the shape is unmistakable. Under the plain price loss the weight climbs from 0.041890 at the strike carrying the highest volatility to 0.308446 at the Rs 110/- strike, then eases back at Rs 120/-. The weight profile is not the price profile and it is not flat. The weight profile is the vega profile, squared and then bent slightly by the curvature of price in volatility. The heaviest weight sits where vega peaks, and on this set vega peaks near the middle of the strike range, firmly on the low volatility side of it.
Read the fourth column and something different is happening. The vega weighted rule puts 0.148766, 0.199854, 0.211964, 0.212630 and 0.226786 on the five readings. The five weights are nearly equal. Near equality is precisely what dividing by vega squared is for: one vega term in the effective weight is cancelled by the stated weight and the other very nearly is, leaving a rule that treats the five readings almost alike. Vega weighting does not concentrate the fit anywhere; it is a way of undoing the concentration the plain price loss carries. Vega weighting therefore lands at 0.204420, only 0.002580 from the plain mean, while the price loss sits 0.008798 away.
The fifth column is exact by construction and needs no computation at all. Comparing volatilities directly puts 0.200000 on each of the five readings, and a weighted average with equal weights is the plain mean. The five readings of 0.240000, 0.220000, 0.200000, 0.190000 and 0.185000 sum to 1.035000, and a fifth of that is 0.207000. The volatility loss returns the plain average of the five readings because equal weighting and a plain average are the same statement said twice.
Why does the squared implied volatility error return exactly 0.207000?
How much does it change the answer?
The three answers span 0.008798 of volatility. The span is the whole distance between the plain price loss and the volatility loss. The vega weighted rule sits inside that span at 0.204420, or 0.002580 below the plain mean. Set that against the effects this subject treats at length and it is large. The spread is bigger than the gap between a twelve step lattice and the closed form. The spread is bigger than most numerical error a careful implementation carries. The choice of loss is not a small residual sitting under the real answer; on this set it is one of the largest single movements available.
In rupees at the middle strike the same spread is Rs 0.330352/-, the difference between what the Rs 100/- contract prices at under 0.198202 and what it prices at under 0.207000. On a model price of about Rs 10.45/- that is roughly three per cent of the value of the thing, produced entirely by a rule nobody wrote down.
| Strike | Observed price | Miss at 0.198202 | Miss at 0.204420 | Miss at 0.207000 |
|---|---|---|---|---|
| Rs 80/- | Rs 25.227000/- | minus 0.662462 | minus 0.576709 | minus 0.539720 |
| Rs 90/- | Rs 17.257579/- | minus 0.606838 | minus 0.437299 | minus 0.366074 |
| Rs 100/- | Rs 10.450584/- | minus 0.067452 | plus 0.165950 | plus 0.262899 |
| Rs 110/- | Rs 5.644765/- | plus 0.324179 | plus 0.570327 | plus 0.672547 |
| Rs 120/- | Rs 2.745149/- | plus 0.441183 | plus 0.653633 | plus 0.742577 |
| Root mean squared miss | five prices, fixed | Rs 0.471464/- | Rs 0.510661/- | Rs 0.547387/- |
The last row across is the one to take first. The root mean squared miss rises from Rs 0.471464/- to Rs 0.547387/- across the three answers, and that looks like a ranking. The rise is not a ranking. Root mean squared miss in rupees is the squared price error rule wearing a different label, so of course the squared price error rule wins on it. Scored by the vega weighted rule, the same three answers put 0.204420 first, at 0.001875 against 0.002044 and 0.001905. Scored by the volatility rule, 0.207000 wins, at 0.002080 against 0.002467 and 0.002113. Three different questions are on the table and each rule answers its own, so each of the three fits is the best answer available and each of the three is beaten by the other two.
How much of the spread of 0.008798 comes from the five observed prices?
The five prices are about to be held completely fixed while the rule changes. Before it moves: does the fitted volatility change?
Change the rule, watch the answer move, and check that nothing observable moved with it
Held fixed throughout: the standard process starting at Rs 100/-, a rate of 5 per cent a year, a horizon of one year, and five observed prices of Rs 25.227000/-, Rs 17.257579/-, Rs 10.450584/-, Rs 5.644765/- and Rs 2.745149/- at strikes of Rs 80/-, Rs 90/-, Rs 100/-, Rs 110/- and Rs 120/-. Pick a rule and the slider jumps to that rule's own answer, computed on the spot rather than looked up. Squared price error lands at 0.198202 with misses of minus 0.662462, minus 0.606838, minus 0.067452, plus 0.324179 and plus 0.441183. Vega weighting lands at 0.204420 with misses of minus 0.576709, minus 0.437299, plus 0.165950, plus 0.570327 and plus 0.653633. Squared volatility error lands at 0.207000 with misses of minus 0.539720, minus 0.366074, plus 0.262899, plus 0.672547 and plus 0.742577. The spread of 0.008798 across those three answers is produced entirely by the three buttons, and the five prices above are the same five prices in every case.
At a candidate volatility of 0.198202 the five misses are minus 0.662462, minus 0.606838, minus 0.067452, plus 0.324179 and plus 0.441183 rupees, and the root mean squared miss is 0.471464. The selected rule is the squared price error, which stops at 0.198202, so this is exactly where the rule would leave the analyst. That is 0.008798 below the answer the squared volatility error gives on the same five prices.
How should the choice be made and recorded?
Nothing in the five prices has a preference, so the choice follows the use, and the use is the only thing that can sensibly decide it. Ask what the fitted model is for and the rule usually falls out. If the model exists to reproduce prices in rupees across a book, and a rupee of error is a rupee of error wherever it lands, then the plain squared price error is asking the matching question and its concentration on the high vega strikes is a feature rather than a distortion. If the model exists to be compared against volatility readings, or to be handed to somebody who will quote it as a volatility, then the volatility loss asks the matching question and its equal treatment of the five readings is the point.
Vega weighting sits between the two and is the honest choice when the targets are prices but the answer will be read as a volatility. Dividing by vega restates every gap in the units the answer will be quoted in, and that is why vega weighting lands close to the volatility loss. The rule is chosen by what the fitted number will be used for, and only afterwards by anything about the numbers themselves.
Recording it costs one line and there is no defensible reason to leave it out. The line names the rule, the weights, the units the residuals were measured in, and the set of targets they were measured over. If the weights were anything other than one, they are written out rather than described. A phrase like weighted by vega has at least two readings that give different answers, and a reader six months later cannot tell which was meant. The name of a rule is ambiguous in a way the weights never are, so the written form of the rule is the weights themselves.
The fitted volatility will be used to reprice contracts near the middle strikes, where a rupee of error is a rupee of error. Which rule matches that use?
What somebody reviewing a fitted number actually does with all three answers
- Ask what was minimised before asking what came out. A reviewer handed 0.198202 and no rule has been handed one number from a set of three, with no way of telling which. The first question is not whether the number is plausible; it is which of the three questions produced it.
- Refit under the other two rules and keep all three. On this set that costs almost nothing and produces 0.198202, 0.204420 and 0.207000. The width of 0.008798 is the honest statement of how much the fit depends on a choice nobody observed, and it is more informative than any one of the three answers alone.
- Convert the width into the unit the decision is made in. At the Rs 100/- strike the same spread is Rs 0.330352/-, and a reader who thinks in prices rather than volatilities needs that figure in order to judge whether the width matters.
- Check whether the rule matches the use. A number produced by a volatility loss and then used to reprice a book in rupees was fitted to answer a different question from the one it is being asked. The mismatch is invisible in the number and obvious in the record, and that is the argument for keeping the record.
- Refuse to average the three. The midpoint of 0.198202, 0.204420 and 0.207000 is a setting that no rule selected and that none of the three would defend. A range that is honest about its width beats a single number that is not.
A weighted sum of squares depends on no jurisdiction
The mechanism here is arithmetic. A weighted sum of squared differences behaves the same way under every legal system, in every currency and in every market, and there is no threshold, no rate and no reporting rule anywhere in it. Any duty to record what was minimised comes from whatever conduct or model governance obligations apply to the person doing the fitting, and those obligations sit outside the arithmetic entirely. The arithmetic itself carries no jurisdiction.
The failure: treating the fitted number as something the prices decided
The failure is quiet and it looks exactly like success. A number comes out of a fit, to six decimal places, from a routine that ran cleanly and converged. The number is written into a record as 0.198202. Nothing about it announces that it is one of three, and nothing about the six decimal places distinguishes a number the data pinned down from a number a chosen rule selected.
The cost is reproducibilityWhether anyone can recompute a reported number from the record of how it was produced. It needs the rule stated, because the rule is one of the inputs., and it is total rather than partial. A reader handed 0.198202 and the five prices cannot get back to 0.198202. The five prices are consistent with 0.204420 and 0.207000 as well. Neither can its author six months later, once the memory of which routine was used has gone. The number is not wrong, and it is not reproducible, and those two things are usually assumed to be the same thing.
The failure is worse than an ordinary gap in a record because the missing item is an unobserved inputA choice that decides the answer and is never written down anywhere, which the loss function usually is. rather than a missing observation. A missing price can be looked up. A missing rule cannot be looked up anywhere. The rule never existed outside the head of whoever wrote the fitting routine. The record is the only copy there ever was.
A calibrated volatility is reported together with all five target prices, but without the rule that was minimised. Can anyone reproduce it?
What is worth carrying away from all of this?
Three sentences hold the whole subject. The first is that a fit has three inputs and only two of them are ever found: the targets and the model are read off something, and the rule that decides what closeness means is manufactured. The second is that the manufactured input is not small. On five invented prices the manufactured input moves the answer by 0.008798 of volatility, or Rs 0.330352/- at the middle strike, and it does so with every observable quantity held completely still.
The third is the one worth holding on to longest, and it is the one that generalises past this arithmetic. None of the three rules is wrong. Squared price error, vega weighting and squared volatility error each state a coherent view of what it means to be close, and each returns the best answer available under its own view. The failure is a report that states an answer without stating which view produced it. A fit that does not say which question it asked has not reported its own inputs, and the answer it prints is arithmetic without a statement of what was being asked.
References
| Source | Document | Where |
|---|---|---|
| arXiv, Quantitative Finance | Preprints on calibration objective functions, weighting schemes and the stability of fitted parameters in derivative pricing models | arxiv.org |
| Social Science Research Network | Working papers on calibration practice, the reporting of fitted quantities and model governance around fitted inputs | ssrn.com |
| Black, Scholes and Merton, 1973 | The option pricing formula and its volatility sensitivity, the source of the model prices and the vegas used throughout | Journal of Political Economy; Bell Journal of Economics and Management Science |
The standard process, its five strikes and the five observed volatility readings are invented.
Educational material. Not advice on any investment, tax, budget or market position.
