Identifiability: When Two Parameter Sets Fit Equally Well
Identifiability is whether the numbers a fit is aimed at can tell one parameter set apart from another. Two sets are unidentified when they produce the same value for everything the fit looks at. The search still returns one of them, to as many decimal places as anyone asks for, but it chose between them arbitrarily and the data preferred neither.
Backtesting tests a fitted model on prices it never saw. Underneath that test sits a quieter question, one that never announces itself: whether the numbers being fitted to were ever capable of choosing an answer at all. A search always terminates and always returns a value. The value is printed to six decimal places whether or not anything in the targets pointed at it. A search reports where it stopped, and where it stopped is not the same thing as where the data sent it.
Everything below is computed from invented numbers. The quantity being modelled is the standard process, written S with a time subscript, starting at Rs 100/-, with a base volatility of 20 per cent a year, a rate of 5 per cent a year and a horizon of one year. Onto that process a jump component is added, and it is the jump component's parameters that are in question. None of it is an observation of any market anywhere, and every figure below is an educational illustration.
Two facts carry the whole subject. The first is that two genuinely different parameter sets can agree to the last decimal place on the quantity a fit is aimed at. The second is that the agreement is not a property of the parameters, it is a property of the aim. Change the targets a fit is aimed at, and two settings that no search could separate become four times apart.
What is identifiability?
IdentifiabilityWhether the data being fitted to can distinguish one parameter set from another. is a property of a pairing, not of a model. Identifiability pairs a set of parameters with a set of targets and asks one question: if the parameters move, does anything in the targets move? Where the answer is no, the two settings are unidentifiedProducing the same values for everything the fit looks at, so no search aimed at those targets can prefer one over the other. by those targets. There is nothing there to separate, so no amount of care in the search will separate them.
Here is the everyday version, and it is exact rather than approximate. A doorway has a counter on it that reports one number at the end of the day: the total number of passes through the door. Today it reads 100. Now ask what happened. One hundred people could have passed once each. Fifty people could have passed twice. Four people could have passed twenty five times. The counter is not broken, it is not noisy, and it is not imprecise. The counter answers exactly the question it was built to answer, and that question does not include the split. The number of people and the number of passes each are two separate facts about the day, and the counter carries only their product.
The remedies divide sharply into those that work and those that do not. A better counter will not fix it. A counter accurate to the nearest thousandth of a pass will not fix it. Tomorrow has the same structure, so running the count again will not fix it either. A second column fixes it: a timestamp on each pass, or a tally of distinct entrants. The moment the record carries something that the two stories disagree about, the two stories come apart. Applied to a pricing model, that sentence is the whole of identifiability.
| \(\theta\) | a parameter set, meaning one complete choice of every free number the model needs before it can produce anything |
| \(g\) | the target map, meaning the rule that turns a parameter set into exactly the quantities the fit is aimed at, and nothing else |
| \(g(\theta)\) | the values those quantities take under that parameter set, and the whole of what the search can see |
| \(\theta_1, \theta_2\) | two candidate parameter sets, different from each other as settings |
The last point is the one that gets lost. Identification is not a numerical difficulty. A stuck search, a loose tolerance and an unlucky starting guess are real problems with real remedies, and identification is none of them. The failure is a different kind of thing entirely: the information required to choose was never in the room. Computation cannot supply an answer to a question the record does not answer.
Is identification a property of the parameters, or of the targets the fit is aimed at?
How can two different parameter sets fit equally well?
Take the standard process and add a jump component to it, in the manner of Merton, 1976. The process now moves for two separate reasons: a continuous part driven by Brownian motion at a volatility of 20 per cent a year, and a jump part that arrives at random times and moves the process by a random proportion when it does. The jump part needs three numbers before it can produce anything: how often jumps arrive, how large the typical jump is, and how much the jump size varies.
The first of those three is the intensityHow often jumps arrive, measured as an average number per year. It is 0.5 in the first set here and 2.0 in the second., written lambda and measured as an average number of arrivals a year. Set A puts it at 0.5, one arrival every two years on average. Set B puts it at 2.0, one arrival every six months. The two settings are not neighbouring descriptions of the same world. Under set A there is a 0.606531 chance of the year passing with no jump at all. Under set B that chance is 0.135335. An observer watching would know within a year which of the two worlds it was.
| \(S_T, S_0\) | the standard process at the horizon and at the start, Rs 100/- at the start here |
| \(\sigma\) | the base volatility of the continuous part, 0.20 a year here, so \(\sigma^2 T\) is 0.040000 |
| \(T\) | the horizon, one year here |
| \(\lambda\) | the intensity, the average number of jump arrivals a year, counted by \(N^J\) with a time subscript |
| \(\mu_J\) | the average of the logarithm of the jump size, minus 0.050000 in set A and minus 0.025000 in set B |
| \(\sigma_J\) | the spread of the logarithm of the jump size, 0.100000 in set A and 0.050000 in set B |
Work both sets through that formula. In set A the average squared jump is 0.050000 squared plus 0.100000 squared, or 0.012500, and multiplying by an intensity of 0.5 gives a jump variance of 0.006250. In set B the average squared jump is 0.025000 squared plus 0.050000 squared, or 0.003125, and multiplying by an intensity of 2.0 gives a jump variance of 0.006250. Identical. Add the continuous part's 0.040000 to each and the total varianceThe variance of the log return over the whole horizon, continuous part plus jump part. It is 0.046250 under both sets here. is 0.046250 under both, exactly, with nothing rounded away.
| Quantity | Set A | Set B | Do they agree? |
|---|---|---|---|
| Intensity, arrivals a year | 0.500000 | 2.000000 | No, by a factor of four |
| Average log jump size | minus 0.050000 | minus 0.025000 | No, by a factor of two |
| Spread of the log jump size | 0.100000 | 0.050000 | No, by a factor of two |
| Average squared jump | 0.012500 | 0.003125 | No, by a factor of four |
| Jump variance over the year | 0.006250 | 0.006250 | Yes, exactly |
| Continuous variance over the year | 0.040000 | 0.040000 | Yes, by construction |
| Total variance of the log return | 0.046250 | 0.046250 | Yes, exactly |
Read the table downward and the pattern is stark. Every single input differs. Every intermediate quantity differs. And then the last row lands on the same number. The two differences cancel against each other perfectly: the intensity is multiplied by four and the average squared jump is divided by four. A fit aimed at the total variance sees only that last row, and the last row is blind.
| \(c\) | a scaling factor applied to both jump size parameters at once, positive and otherwise free |
| \(\lambda/c^2\) | the arrival rate that compensates for that scaling, rising as the jumps are made smaller |
| \(0.006250\) | the jump variance held fixed along the whole curve, and 0.040000 added to it gives 0.046250 |
The situation is not a near miss. Two settings have not merely landed close together, where a sharper instrument would separate them. An entire one dimensional curve of settings produces the identical target value, so a search aimed at that target has a flat floor to walk on and no reason to prefer any point on it. The minimum is not a point, it is a line, and a search that returns a point has reported its own stopping place.
Which quantity do the two parameter sets agree on exactly?
The two sets do not agree on everything
Now push past the second moment. The jump part of the log return has a fourth cumulant as well, and it is the arrival rate multiplied by the average fourth power of the jump. A normal distribution has no excess in its tail by definition, so the continuous part contributes nothing to it at all. So the whole of the excess kurtosisA measure of how heavy the tail is relative to a normal distribution. It is 0.106647 under set A and 0.026662 under set B, a factor of exactly four. comes from the jumps.
| the top line | the jump part's fourth cumulant, being the arrival rate multiplied by the average fourth power of the log jump size |
| the bottom line | the square of the total variance, 0.046250 squared here and the same under both parameter sets |
| \(\lambda, T\) | the intensity and the horizon, 0.5 or 2.0 arrivals a year over one year here |
| \(\mu_J, \sigma_J\) | the average and the spread of the logarithm of the jump size, as defined above |
| \(\sigma\) | the base volatility of the continuous part, 0.20 a year, which contributes nothing at all to the top line |
Under set A the average fourth power of the jump is 0.000456250, and multiplying by an intensity of 0.5 gives a fourth cumulant of 0.000228125. Under set B the average fourth power is 0.000028516, and multiplying by an intensity of 2.0 gives 0.000057031. Dividing each by the squared total variance of 0.046250 squared gives an excess kurtosis of 0.106647 under set A and 0.026661797, or 0.026662 to six places, under set B. The ratio between them is exactly 4.00, not approximately.
The intensity is about to change from 0.5 to 2.0, with the jump size adjusted alongside it. Before it moves: what happens to the total variance?
Slide the arrival rate along the curve and watch one reading refuse to move
Held fixed: the standard process starting at Rs 100/-, a base volatility of 20 per cent a year, a rate of 5 per cent a year and a horizon of one year. The only thing that moves is the jump arrival rate, and the two jump size parameters are rescaled with it so that the jump variance stays at exactly 0.006250 throughout. Set A sits at the left end of the track, with an intensity of 0.5, an average log jump of minus 0.050000 and a spread of 0.100000, giving a total variance of 0.046250 and an excess kurtosis of 0.106647. Set B sits at the right end, with an intensity of 2.0, an average log jump of minus 0.025000 and a spread of 0.050000, giving the same total variance of 0.046250 and an excess kurtosis of 0.026662, smaller by exactly 4.00. Every setting on this track reproduces a variance target perfectly, and that is why a variance target chooses none of them.
At an intensity of 0.500000 a year, with an average log jump of minus 0.050000 and a spread of 0.100000, the total variance of the log return is 0.046250 and the excess kurtosis is 0.106647. This is set A. A fit aimed only at the total variance would accept every setting on this track equally, so it would have no reason to stop here rather than anywhere else.
What does that mean for the fitted numbers?
A number that came out of a fit carries an implicit claim, and the claim is usually not stated because everybody assumes it. The claim is that the data pushed the search to that value and away from the others. Strip that claim away and what is left is a stopping point. An unidentified fitted number is not a measurement of anything, it is a record of where the search was standing when it ran out of reasons to move.
The everyday version again, and it is the same doorway. Suppose somebody reads the counter, thinks about it for a while, and writes down: forty seven distinct people passed through today. The figure of forty seven is not wrong in the way a miscount is wrong. The wrongness is stranger than that: nothing in the record could have produced the figure, so it came from somewhere else, most likely from whatever the writer happened to assume before looking. Change the assumption and the number changes, and the record sits there unchanged throughout.
A fit in that position performs an arbitrary selectionWhat a fit does when it cannot distinguish between candidates: it returns whichever one it started nearest, because nothing in the targets pushes it away. rather than an estimate. The surface is flat and there is nowhere downhill to go, so a search started at an intensity of 0.6 will report something close to 0.6. Start it at 1.8 and it will report something close to 1.8. Both runs will announce success. Both will report a residual of nought against the target. Neither one has learned anything from the targets that the other did not.
Parameter Risk: what is actually at stake when the fit cannot choose?
Parameter Risk is the exposure that comes from a model's inputs being uncertain rather than from the model's structure being wrong. Parameter Risk usually shows up as a range: the input could reasonably be anywhere in a window, so the output could reasonably be anywhere in a corresponding window, and the width of that second window is the risk. The sharpest form of it appears here, where the window on the input is neither narrow nor the result of noisy data. The window is the whole curve, and the data has nothing to say about where in it the answer lies.
The danger is precisely that an unidentified estimate does not look like a range. A noisy estimate announces itself: rerun it and it moves, and anybody paying attention sees the movement. An unidentified estimate is perfectly stable under everything except a change of starting point, and a change of starting point is not something most reports record. Reproducibility is the property most readers use as a proxy for reliability, so the most dangerous parameter risk is the kind that reproduces perfectly.
The cost lands wherever the model is used for something the fit was not aimed at. If the only use of the fitted model is to reproduce a variance, the choice between the two sets does not matter. Both sets reproduce it. But models are rarely built to reproduce their own targets. Models are built to price something else, or to describe a tail, or to answer a question about how often something large happens. On every one of those uses the two sets part company, and they part company by a factor of four in the tail and by a factor of four and a half in the chance of a quiet year.
| Question put to the fitted model | Set A answers | Set B answers | Does the choice matter? |
|---|---|---|---|
| What is the total variance of the log return? | 0.046250 | 0.046250 | No |
| What is the total volatility over the year? | 0.215058 | 0.215058 | No |
| What is the excess kurtosis? | 0.106647 | 0.026662 | Yes, by 4.00 times |
| What is the skewness? | minus 0.081688 | minus 0.040844 | Yes, by 2.00 times |
| What is the chance of a year with no jump? | 0.606531 | 0.135335 | Yes, and by a wide margin |
| What is the average waiting time between jumps? | 2.000000 years | 0.500000 years | Yes, by 4.00 times |
Read the last four rows and the size of the exposure is plain. A reader who takes the reported intensity of 0.5 and answers a question about how often something large arrives has answered it four times too rarely, or four times too often, depending on which set they were handed. And the fit's own output looked exactly the way a well determined fit looks, so it carried no warning at all.
A fit reports an intensity of 0.500000 to six decimal places. Is that a measurement?
What can separate two sets that a fit cannot?
Everything, as it happens, except the one quantity the fit was aimed at. The surprise is worth drawing out carefully. Two sets like these are often described as though they were somehow the same underneath, and they are not. The two sets agree at exactly one place and disagree everywhere else, and the place they agree is the one place the fit was looking.
The cleanest way to see it is to line up the cumulants of the log return, order by order. A cumulant is a number that describes the shape of a distribution at one particular order: the first is the mean, the second is the variance, the third carries the lean, the fourth carries the weight in the tails. The jump part contributes to every one of them, and its contribution at order n is the arrival rate multiplied by the average of the jump size raised to the power n.
| \(\kappa_n\) | the jump part's contribution to the cumulant of order \(n\) of the log return over the horizon |
| \(Y\) | the logarithm of one jump's size, normally distributed with mean \(\mu_J\) and spread \(\sigma_J\) |
| \(\mathbb{E}[Y^n]\) | the average of that logarithm raised to the power \(n\), halving in scale each time \(n\) rises by one when the jump is halved |
| \(\kappa_n^{\,A}, \kappa_n^{\,B}\) | that contribution under set A and under set B respectively |
Put numbers on it. At the first order the jump contributions are minus 0.025000 and minus 0.050000, a ratio of 2. At the second they are 0.006250 and 0.006250, a ratio of 1. At the third they are minus 0.000812500 and minus 0.000406250, a ratio of one half. At the fourth they are 0.000228125 and 0.000057031, a ratio of one quarter. Plot those four ratios on a doubling scale and they fall along a perfectly straight line that crosses the level of exact agreement at the second order and nowhere else.
One order on that ladder is a trap and needs a note. The first order differs by a factor of two, so the mean of the log return might be expected to separate the two sets. The mean does not separate them, and the reason lies in the pricing theory rather than in the arithmetic. Under the risk-neutral measure Q the drift of the continuous part is not free. The drift is set to make the discounted process a martingale, and it therefore takes whatever value the jump terms leave it. The compensator is minus 0.022001 under set A and minus 0.046940 under set B, and the continuous drift absorbs the difference exactly. Both sets give the same expected value at the horizon, Rs 105.13/-, and they give it for the same reason a fitted parameter is not a discovery: the constraint put it there, not the data.
So the first order is pinned by no-arbitrage, the second is pinned by the construction of the pair, and the third and the fourth are free to differ. The fourth momentThe quantity carrying the weight in the tails of a distribution. The fourth moment distinguishes the two parameter sets here. is therefore the quantity to reach for. The fourth moment is the first target that could be added which carries real separating power and which nothing else has already fixed.
Which quantity distinguishes the two sets here?
How large is the ratio between the two excess kurtosis figures?
How is a lack of identification detected?
The answer looks fine, so looking at the answer detects nothing. The residual is nought, and nought is what a good fit produces, so the residual gives nothing away either. Detection needs the one thing that changes between two runs of the same fit and nothing else: refit from a different starting point, with the same targets and the same bounds, and see whether the same answer comes back.
Run it from an intensity of 0.6 and it reports something near 0.6. Run it from 1.8 and it reports something near 1.8. Both answers carry a residual of nought against the variance target, so neither can be dismissed as a failed run. Two different answers, both reproducing the targets exactly, is the signature, and no single run of any search can produce it. The test therefore needs at least two runs to say anything at all.
There is a second check that costs nothing and catches the same thing from the other side. The answer the search returned is taken, one parameter is moved by a small amount, the others are moved to hold the target fixed, and the loss is watched to see whether it moves. If it does not, that is a direction in the parameter space along which the targets are flat, and any point along it would have served. Here that direction is written down explicitly in the rescaling above, so it does not have to be hunted for, but on a real fit with several parameters at once it usually has to be found by walking.
How is a failure of identification detected?
What does an unidentified fit look like from the outside?
Exactly like an identified one. Every instinct a careful reader has been trained to use fails at this point, so the answer is worth stating baldly. The output sheet carries a parameter, a value to six decimal places, a residual, a date, a note of the bounds and a note of the loss. Every field is filled. Every field is correct. And nothing anywhere on the face of it records that the value in the second field was chosen by the starting guess rather than by the targets.
The error that gets made, and what it costs
A fit is run, an intensity of 0.500000 comes back, and it is written into a document as a finding: the process jumps about once every two years. Nobody lied and nobody miscalculated. The search really did return 0.500000 and it really does reproduce the target exactly. But a different starting point would have returned 2.000000 with precisely the same residual and precisely the same confidence, and the document would then have said the process jumps about twice a year. Two documents, one dataset, opposite descriptions of how the world behaves, and no way to choose between them from anything either document contains.
The cost is a spurious precisionDecimal places carried on a number that nothing in the data selected, which is invisible from outside because the output looks the same either way. that is invisible in every output. The residual is nought, so no wide error bar appears. The same starting point gives the same answer every time, so a rerun shows nothing either. The precision shows itself only where the fitted model is asked a question the targets never covered, and by then the number has usually been copied into three other places and has acquired the authority of something that was measured.
The reader who takes the reported intensity of 0.500000 as a fact about anything has read a number that the fit had no basis for choosing. And the same is exactly as true of a reported 2.000000. The targets prefer neither reading, and nothing else the fit produced prefers one either.
What should be reported when parameters are not identified?
Not one of the answers. The set of them, and the quantity that would have separated them. The report runs shorter than most people expect and does more than a single number with six decimal places ever does. A reader learns exactly what they can and cannot do with what they have been handed.
Four things belong on the sheet. First, the targets, listed rather than described, as the whole finding is about them. Second, the set of parameter settings that reproduce those targets: here the entire curve running from an intensity of 0.5 through to 2.0 and beyond in both directions, usually written down as a relationship rather than enumerated. Third, the quantity that would separate them, named and quantified: here the excess kurtosis at 0.106647 against 0.026662. Fourth, the plain statement that the fit did not choose, and that any single value reported from it is a stopping point.
The test of the sheet is whether a reader who has never seen the fit could tell, from the sheet alone, that a second answer exists. A single number fails that test always. A number with a note saying the search converged fails it too. Convergence is the wrong evidence: an unidentified search converges perfectly well, it just converges wherever it began.
Which of these should be reported when parameters are not identified?
How does somebody actually use this when a fitted number lands on their desk?
The situation is ordinary. A number arrives, attached to a model, with a note saying it was calibrated. The number is going to be used for something. The useful question is not whether to trust it, a question too blunt to help, but which uses it supports and which it does not.
Five questions, in the order that gets to the answer fastest
- What was targeted, exactly? Not the model's subject, but which quantities the loss was computed on. Everything else on this list depends on the answer, and a vague answer here means every question below is unanswerable.
- Was it refitted from more than one starting point? If the answer is no, nobody knows whether the number was chosen by the targets, and the sheet cannot settle it either way. The refit question costs the least and settles the most.
- Is there a rescaling that holds the targets fixed? On a jump component the answer is usually yes and it is usually the one set out here: raise the arrival rate, shrink the jumps, and the variance does not move. Knowing the shape of the flat direction shows which parameters are the risk.
- What is this number about to be used for, and is that quantity in the targets? If the use is inside the target set, the choice does not matter. If it is outside, the choice may matter by a factor of four, and the fit had no opinion about it.
- What would separate the candidates, and can it be added? Often it can, and cheaply. Adding one quantity to the targets is a smaller change than rebuilding the model, and it is the change that turns a stopping point into an answer.
The fourth question is the one that carries the weight in practice. A fitted parameter is a serviceable input to a calculation that stays inside what it was fitted to, and it is a hazard the moment it leaves. Here the boundary is precise and can be stated in one line: any question that depends only on the total variance is answered identically by both sets, and any question that depends on the tail, on the lean, or on how often anything happens is answered differently by a factor of two or four. The use a number is put to comes before whether it is right. For half the uses set out here, which of the two sets is correct never arises at all.
Where this holds
Identifiability is not jurisdictional. Whether a set of targets can distinguish two parameter settings is a property of the arithmetic relating them, and the arithmetic does not change with the place it is written down. The conventions that decide what a target quantity even means, such as how a horizon is turned into a fraction of a year or how a quoted number is defined, do vary by place, and those are set out under market convention. Any such convention should be confirmed at its own source rather than from a treatment of fitting.
References
| Source | Document | Where |
|---|---|---|
| arXiv, Quantitative Finance | Preprints on parameter identification, inverse problems and moment matching in derivative pricing models | arxiv.org |
| Social Science Research Network | Working papers on calibration practice, parameter risk and the reporting of fitted quantities | ssrn.com |
| Robert C. Merton, 1976 | Option pricing when underlying stock returns are discontinuous, the jump component added to the continuous process here | Journal of Financial Economics |
The standard process, the two jump parameter sets and every figure computed from them are invented.
Educational material. Not advice on any investment, tax, budget or market position.
