Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Stochastic Calculus & Derivative Pricing Theory
1Probability Foundations
The Probability SpaceRandom VectorsSigma-AlgebraExpectationSample Space and EventsDensity and Distribution FunctionsRisk-Neutral ProbabilityState Price Density vs…
2Stochastic Processes and Jumps
Properties of a Stochastic ProcessMartingaleBrownian Motion and Its PropertiesBrownian Motion vs Geometric…Stopping TimeThe Markov PropertyState VariablesTransition ProbabilityQuadratic VariationQuadratic Variation vs Ordinary…Submartingale and SupermartingaleMartingale RepresentationMarkov Process vs MartingaleOptional StoppingFiltrationJump ProcessesThe Poisson ProcessLevy ProcessesJump Diffusion
3Ito Calculus
The Ito IntegralThe Ito Integral vs the Riemann IntegralInfinitesimals in Stochastic CalculusQuadratic CovariationIto's LemmaHow to Apply Ito's…The Infinitesimal GeneratorIto Calculus vs Ordinary Calculus
4Stochastic Differential Equations
Stochastic Differential EquationsStochastic Differential Equation vs…Drift and DiffusionStrong and Weak Solutions ComparedDiscretisationGeometric Brownian Motion
5Pricing Theory and No-Arbitrage
No-ArbitrageGirsanov, Radon-Nikodym and Change…Physical and Risk-Neutral Measures…The Fundamental Theorems of…The Law of One PriceThe Pricing KernelDiscount Factors and Zero-Coupon PricesReplication vs HedgingComplete Market vs Incomplete MarketClearing Margin Architecture
6Option Pricing Theory
European and American OptionsMonte Carlo European OptionThe Black-Scholes PDEBlack Scholes and the GreeksThe Payoff FunctionThe Binomial ModelBinomial Option PricingDelta Hedging in TheoryBoundary, Initial and Terminal ConditionsThe Exercise BoundaryHow to Check Put-Call…
7Volatility Models
Constant, Local and Stochastic…Vasicek Model vs CIR ModelThe Heston ModelThe SABR ModelThe Volatility ProcessImplied VolatilityVolatility Smile vs Skew vs Surface
8Interest Rate Models
Interest-Rate DerivativesMean ReversionThe Zero-Coupon BondThe Ornstein-Uhlenbeck ProcessThe Discount CurveZero RatesShort-Rate Model vs Market Model
9Numerical Pricing
Closed Form and Numerical…Monte Carlo PricingEuler and Milstein Schemes ComparedTree MethodsFinite Difference MethodsNumerical Error and StabilityVariance Reduction
10Calibration and Model Risk
Model OverrideMarket Price and Model PriceCalibrationHow to Document a Pricing ModelThe Educational Illustration LabelMarket ConventionsModel Uncertainty and LimitationsBacktesting a Pricing ModelIdentifiabilityCalibrated ParametersThe Calibration Loss Function

Identifiability: When Two Parameter Sets Fit Equally Well

Identifiability is whether the numbers a fit is aimed at can tell one parameter set apart from another. Two sets are unidentified when they produce the same value for everything the fit looks at. The search still returns one of them, to as many decimal places as anyone asks for, but it chose between them arbitrarily and the data preferred neither.

Backtesting tests a fitted model on prices it never saw. Underneath that test sits a quieter question, one that never announces itself: whether the numbers being fitted to were ever capable of choosing an answer at all. A search always terminates and always returns a value. The value is printed to six decimal places whether or not anything in the targets pointed at it. A search reports where it stopped, and where it stopped is not the same thing as where the data sent it.

Everything below is computed from invented numbers. The quantity being modelled is the standard process, written S with a time subscript, starting at Rs 100/-, with a base volatility of 20 per cent a year, a rate of 5 per cent a year and a horizon of one year. Onto that process a jump component is added, and it is the jump component's parameters that are in question. None of it is an observation of any market anywhere, and every figure below is an educational illustration.

Two facts carry the whole subject. The first is that two genuinely different parameter sets can agree to the last decimal place on the quantity a fit is aimed at. The second is that the agreement is not a property of the parameters, it is a property of the aim. Change the targets a fit is aimed at, and two settings that no search could separate become four times apart.

What is identifiability?

IdentifiabilityWhether the data being fitted to can distinguish one parameter set from another. is a property of a pairing, not of a model. Identifiability pairs a set of parameters with a set of targets and asks one question: if the parameters move, does anything in the targets move? Where the answer is no, the two settings are unidentifiedProducing the same values for everything the fit looks at, so no search aimed at those targets can prefer one over the other. by those targets. There is nothing there to separate, so no amount of care in the search will separate them.

Here is the everyday version, and it is exact rather than approximate. A doorway has a counter on it that reports one number at the end of the day: the total number of passes through the door. Today it reads 100. Now ask what happened. One hundred people could have passed once each. Fifty people could have passed twice. Four people could have passed twenty five times. The counter is not broken, it is not noisy, and it is not imprecise. The counter answers exactly the question it was built to answer, and that question does not include the split. The number of people and the number of passes each are two separate facts about the day, and the counter carries only their product.

The remedies divide sharply into those that work and those that do not. A better counter will not fix it. A counter accurate to the nearest thousandth of a pass will not fix it. Tomorrow has the same structure, so running the count again will not fix it either. A second column fixes it: a timestamp on each pass, or a tally of distinct entrants. The moment the record carries something that the two stories disagree about, the two stories come apart. Applied to a pricing model, that sentence is the whole of identifiability.

The definition, stated over a map from parameters to targets
$$ \theta_1 \neq \theta_2 \quad \text{and} \quad g(\theta_1) = g(\theta_2) \;\;\Longrightarrow\;\; \theta_1, \theta_2 \ \text{are unidentified by} \ g $$
\(\theta\)a parameter set, meaning one complete choice of every free number the model needs before it can produce anything
\(g\)the target map, meaning the rule that turns a parameter set into exactly the quantities the fit is aimed at, and nothing else
\(g(\theta)\)the values those quantities take under that parameter set, and the whole of what the search can see
\(\theta_1, \theta_2\)two candidate parameter sets, different from each other as settings
What it says in wordsTwo different parameter settings that produce identical values for every target quantity cannot be told apart by any fit aimed at those quantities. The failure belongs to the map, not to the search. A cleverer optimiser, a tighter tolerance and a longer run are all better ways of descending a surface that has no slope in that direction, so all three leave the failure exactly where it was.

The last point is the one that gets lost. Identification is not a numerical difficulty. A stuck search, a loose tolerance and an unlucky starting guess are real problems with real remedies, and identification is none of them. The failure is a different kind of thing entirely: the information required to choose was never in the room. Computation cannot supply an answer to a question the record does not answer.

Try it out

Is identification a property of the parameters, or of the targets the fit is aimed at?

Two settings in. One value out. The aperture is what the fit is allowed to look at. SET A intensity 0.5 a year jump mean minus 0.050000 jump spread 0.100000 SET B intensity 2.0 a year jump mean minus 0.025000 jump spread 0.050000 THE TARGETS total variance of the log return, and nothing else 0.046250 the same value from both, to every decimal place Widen the aperture and the two arrows separate. Nothing about either set changes when it does. Educational illustration. Invented parameters, invented process.
Two different parameter sets pass through a target set that carries only the total variance, and both arrive at 0.046250. The collapse happens at the aperture rather than inside either set, which is why a better search cannot undo it.
Derivatives Foundation Bootcamp — Fin Maverick

How can two different parameter sets fit equally well?

Take the standard process and add a jump component to it, in the manner of Merton, 1976. The process now moves for two separate reasons: a continuous part driven by Brownian motion at a volatility of 20 per cent a year, and a jump part that arrives at random times and moves the process by a random proportion when it does. The jump part needs three numbers before it can produce anything: how often jumps arrive, how large the typical jump is, and how much the jump size varies.

The first of those three is the intensityHow often jumps arrive, measured as an average number per year. It is 0.5 in the first set here and 2.0 in the second., written lambda and measured as an average number of arrivals a year. Set A puts it at 0.5, one arrival every two years on average. Set B puts it at 2.0, one arrival every six months. The two settings are not neighbouring descriptions of the same world. Under set A there is a 0.606531 chance of the year passing with no jump at all. Under set B that chance is 0.135335. An observer watching would know within a year which of the two worlds it was.

The total variance of the log return over the horizon
$$ \operatorname{Var}\!\left[\ln \frac{S_T}{S_0}\right] \;=\; \sigma^2 T \;+\; \lambda T \left( \mu_J^{\,2} + \sigma_J^{\,2} \right) $$
\(S_T, S_0\)the standard process at the horizon and at the start, Rs 100/- at the start here
\(\sigma\)the base volatility of the continuous part, 0.20 a year here, so \(\sigma^2 T\) is 0.040000
\(T\)the horizon, one year here
\(\lambda\)the intensity, the average number of jump arrivals a year, counted by \(N^J\) with a time subscript
\(\mu_J\)the average of the logarithm of the jump size, minus 0.050000 in set A and minus 0.025000 in set B
\(\sigma_J\)the spread of the logarithm of the jump size, 0.100000 in set A and 0.050000 in set B
What it says in wordsThe total variance is the continuous part's variance plus the jump part's variance, and the jump part's variance is the arrival rate multiplied by the average squared jump. Because the two enter only as that product, an arrival rate and a jump size can be traded against each other without the total moving at all. The counting process is written with a superscript J here to keep it clear of the standard normal distribution function, another user of the letter N.

Work both sets through that formula. In set A the average squared jump is 0.050000 squared plus 0.100000 squared, or 0.012500, and multiplying by an intensity of 0.5 gives a jump variance of 0.006250. In set B the average squared jump is 0.025000 squared plus 0.050000 squared, or 0.003125, and multiplying by an intensity of 2.0 gives a jump variance of 0.006250. Identical. Add the continuous part's 0.040000 to each and the total varianceThe variance of the log return over the whole horizon, continuous part plus jump part. It is 0.046250 under both sets here. is 0.046250 under both, exactly, with nothing rounded away.

QuantitySet ASet BDo they agree?
Intensity, arrivals a year0.5000002.000000No, by a factor of four
Average log jump sizeminus 0.050000minus 0.025000No, by a factor of two
Spread of the log jump size0.1000000.050000No, by a factor of two
Average squared jump0.0125000.003125No, by a factor of four
Jump variance over the year0.0062500.006250Yes, exactly
Continuous variance over the year0.0400000.040000Yes, by construction
Total variance of the log return0.0462500.046250Yes, exactly

Read the table downward and the pattern is stark. Every single input differs. Every intermediate quantity differs. And then the last row lands on the same number. The two differences cancel against each other perfectly: the intensity is multiplied by four and the average squared jump is divided by four. A fit aimed at the total variance sees only that last row, and the last row is blind.

The whole curve of settings that a variance target cannot separate
$$ \lambda \mapsto \frac{\lambda}{c^{2}}, \quad \mu_J \mapsto c\,\mu_J, \quad \sigma_J \mapsto c\,\sigma_J \qquad \Longrightarrow \qquad \lambda\left(\mu_J^{\,2}+\sigma_J^{\,2}\right) = 0.006250 \ \text{throughout} $$
\(c\)a scaling factor applied to both jump size parameters at once, positive and otherwise free
\(\lambda/c^2\)the arrival rate that compensates for that scaling, rising as the jumps are made smaller
\(0.006250\)the jump variance held fixed along the whole curve, and 0.040000 added to it gives 0.046250
What it says in wordsShrink both jump size parameters by any factor and raise the arrival rate by the square of that factor, and the jump variance does not move at all. The two named parameter sets are not a coincidental pair, they are two points on a continuous curve of settings that a variance target treats as identical. Set A sits at a scaling of one and set B sits at a scaling of one half, with an unbroken line of equally acceptable settings between them and beyond them in both directions.

The situation is not a near miss. Two settings have not merely landed close together, where a sharper instrument would separate them. An entire one dimensional curve of settings produces the identical target value, so a search aimed at that target has a flat floor to walk on and no reason to prefer any point on it. The minimum is not a point, it is a line, and a search that returns a point has reported its own stopping place.

Try it out

Which quantity do the two parameter sets agree on exactly?

The two sets do not agree on everything

Now push past the second moment. The jump part of the log return has a fourth cumulant as well, and it is the arrival rate multiplied by the average fourth power of the jump. A normal distribution has no excess in its tail by definition, so the continuous part contributes nothing to it at all. So the whole of the excess kurtosisA measure of how heavy the tail is relative to a normal distribution. It is 0.106647 under set A and 0.026662 under set B, a factor of exactly four. comes from the jumps.

The excess kurtosis of the log return over the horizon
$$ \text{ExKurt}\!\left[\ln \frac{S_T}{S_0}\right] \;=\; \frac{\lambda T \left( \mu_J^{\,4} + 6\,\mu_J^{\,2}\sigma_J^{\,2} + 3\,\sigma_J^{\,4} \right)}{\left( \sigma^2 T + \lambda T \left( \mu_J^{\,2} + \sigma_J^{\,2} \right) \right)^{2}} $$
the top linethe jump part's fourth cumulant, being the arrival rate multiplied by the average fourth power of the log jump size
the bottom linethe square of the total variance, 0.046250 squared here and the same under both parameter sets
\(\lambda, T\)the intensity and the horizon, 0.5 or 2.0 arrivals a year over one year here
\(\mu_J, \sigma_J\)the average and the spread of the logarithm of the jump size, as defined above
\(\sigma\)the base volatility of the continuous part, 0.20 a year, which contributes nothing at all to the top line
What it says in wordsThe excess weight in the tail comes entirely from the jumps, because a continuous Brownian part has none of its own. The arrival rate enters the top line once while the jump size enters it to the fourth power, so trading one against the other no longer cancels the way it did for the variance. The excess kurtosis can therefore separate two settings that the variance treats as identical.

Under set A the average fourth power of the jump is 0.000456250, and multiplying by an intensity of 0.5 gives a fourth cumulant of 0.000228125. Under set B the average fourth power is 0.000028516, and multiplying by an intensity of 2.0 gives 0.000057031. Dividing each by the squared total variance of 0.046250 squared gives an excess kurtosis of 0.106647 under set A and 0.026661797, or 0.026662 to six places, under set B. The ratio between them is exactly 4.00, not approximately.

Same second moment to the last decimal. Fourth moment four times apart. TOTAL VARIANCE OF THE LOG RETURN 0.046250 under both sets 0.046250 0.046250 0.040000 0.040000 SET A SET B intensity 0.5 intensity 2.0 Pale top segment is the jump part, 0.006250 in both. EXCESS KURTOSIS OF THE LOG RETURN 0.106647 0.026662 exactly 4.00 times SET A SET B intensity 0.5 intensity 2.0 The continuous part adds nothing here. All of it is jumps. Educational illustration. Invented parameters. Nothing here is fitted to any observed price.
Both parameter sets give a total variance of 0.046250, drawn as two bars of identical height, while the excess kurtosis is 0.106647 against 0.026662, a ratio of exactly 4.00. One quantity is blind to the difference and the other is not.
Try it out

The intensity is about to change from 0.5 to 2.0, with the jump size adjusted alongside it. Before it moves: what happens to the total variance?

Play with it

Slide the arrival rate along the curve and watch one reading refuse to move

Held fixed: the standard process starting at Rs 100/-, a base volatility of 20 per cent a year, a rate of 5 per cent a year and a horizon of one year. The only thing that moves is the jump arrival rate, and the two jump size parameters are rescaled with it so that the jump variance stays at exactly 0.006250 throughout. Set A sits at the left end of the track, with an intensity of 0.5, an average log jump of minus 0.050000 and a spread of 0.100000, giving a total variance of 0.046250 and an excess kurtosis of 0.106647. Set B sits at the right end, with an intensity of 2.0, an average log jump of minus 0.025000 and a spread of 0.050000, giving the same total variance of 0.046250 and an excess kurtosis of 0.026662, smaller by exactly 4.00. Every setting on this track reproduces a variance target perfectly, and that is why a variance target chooses none of them.

0.500, set Aintensity 0.5002.000, set B
One control. Two readings. Only one of them moves. TOTAL VARIANCE OF THE LOG RETURN, WHAT THE FIT IS AIMED AT 0.046250, and the bar top never leaves this line 0.040 0.000 Dark segment 0.040000, the continuous part. Pale segment 0.006250, the jump part. Neither moves. The shaded strip is the whole set of settings a variance target accepts, and every point in it is a minimiser. EXCESS KURTOSIS, WHAT THE FIT IS NOT AIMED AT 0.1066 0.0267 0.0000 set A, 0.106647 set B, 0.026662 The whole fall from one end of the track to the other is a factor of exactly 4.00. Nothing in a variance target rewards or penalises any of this movement.
Intensity
0.500000
Total variance
0.046250
Excess kurtosis
0.106647
Chance of no jump
0.606531
Average log jump
minus 0.050000
Log jump spread
0.100000
Jump variance
0.006250

At an intensity of 0.500000 a year, with an average log jump of minus 0.050000 and a spread of 0.100000, the total variance of the log return is 0.046250 and the excess kurtosis is 0.106647. This is set A. A fit aimed only at the total variance would accept every setting on this track equally, so it would have no reason to stop here rather than anywhere else.

Educational illustration. Every reading is computed from the moment formulae above on each move of the control, with no sampling anywhere, so the default reproduces the worked instance exactly and returns to it on every reload. The starting value of Rs 100/-, the base volatility of 20 per cent a year, the rate of 5 per cent a year and both jump parameter sets are teaching figures rather than readings from any market. The chance of no jump is shown because it separates the two ends of the track completely, at 0.606531 against 0.135335. The two settings describe visibly different worlds even where the fit's target cannot tell them apart.
Risk Management Program Bootcamp — Fin Maverick

What does that mean for the fitted numbers?

A number that came out of a fit carries an implicit claim, and the claim is usually not stated because everybody assumes it. The claim is that the data pushed the search to that value and away from the others. Strip that claim away and what is left is a stopping point. An unidentified fitted number is not a measurement of anything, it is a record of where the search was standing when it ran out of reasons to move.

The everyday version again, and it is the same doorway. Suppose somebody reads the counter, thinks about it for a while, and writes down: forty seven distinct people passed through today. The figure of forty seven is not wrong in the way a miscount is wrong. The wrongness is stranger than that: nothing in the record could have produced the figure, so it came from somewhere else, most likely from whatever the writer happened to assume before looking. Change the assumption and the number changes, and the record sits there unchanged throughout.

A fit in that position performs an arbitrary selectionWhat a fit does when it cannot distinguish between candidates: it returns whichever one it started nearest, because nothing in the targets pushes it away. rather than an estimate. The surface is flat and there is nowhere downhill to go, so a search started at an intensity of 0.6 will report something close to 0.6. Start it at 1.8 and it will report something close to 1.8. Both runs will announce success. Both will report a residual of nought against the target. Neither one has learned anything from the targets that the other did not.

Two things that print the same way and are not the same thing. A MEASUREMENT One value is correct. The record points at it. Repeats cluster on it. the value repeat runs land near the same place A different starting point changes nothing. The decimal places describe the reading. A SELECTION Every value on the line fits equally. The search lands nearest its start. where the search began repeat runs land wherever they began A different starting point changes the answer. The decimal places describe the search. Both print as a number with six decimal places, and the printing does not say which one it is. Educational illustration. Invented parameters throughout.
A measurement points at one value and repeats cluster on it, while an unidentified fit selects from a whole line of equally acceptable values and lands nearest wherever it started. Nothing in the data preferred an intensity of 0.5 over an intensity of 2.0.

Parameter Risk: what is actually at stake when the fit cannot choose?

Parameter Risk is the exposure that comes from a model's inputs being uncertain rather than from the model's structure being wrong. Parameter Risk usually shows up as a range: the input could reasonably be anywhere in a window, so the output could reasonably be anywhere in a corresponding window, and the width of that second window is the risk. The sharpest form of it appears here, where the window on the input is neither narrow nor the result of noisy data. The window is the whole curve, and the data has nothing to say about where in it the answer lies.

The danger is precisely that an unidentified estimate does not look like a range. A noisy estimate announces itself: rerun it and it moves, and anybody paying attention sees the movement. An unidentified estimate is perfectly stable under everything except a change of starting point, and a change of starting point is not something most reports record. Reproducibility is the property most readers use as a proxy for reliability, so the most dangerous parameter risk is the kind that reproduces perfectly.

The cost lands wherever the model is used for something the fit was not aimed at. If the only use of the fitted model is to reproduce a variance, the choice between the two sets does not matter. Both sets reproduce it. But models are rarely built to reproduce their own targets. Models are built to price something else, or to describe a tail, or to answer a question about how often something large happens. On every one of those uses the two sets part company, and they part company by a factor of four in the tail and by a factor of four and a half in the chance of a quiet year.

Question put to the fitted modelSet A answersSet B answersDoes the choice matter?
What is the total variance of the log return?0.0462500.046250No
What is the total volatility over the year?0.2150580.215058No
What is the excess kurtosis?0.1066470.026662Yes, by 4.00 times
What is the skewness?minus 0.081688minus 0.040844Yes, by 2.00 times
What is the chance of a year with no jump?0.6065310.135335Yes, and by a wide margin
What is the average waiting time between jumps?2.000000 years0.500000 yearsYes, by 4.00 times

Read the last four rows and the size of the exposure is plain. A reader who takes the reported intensity of 0.5 and answers a question about how often something large arrives has answered it four times too rarely, or four times too often, depending on which set they were handed. And the fit's own output looked exactly the way a well determined fit looks, so it carried no warning at all.

Try it out

A fit reports an intensity of 0.500000 to six decimal places. Is that a measurement?

What can separate two sets that a fit cannot?

Everything, as it happens, except the one quantity the fit was aimed at. The surprise is worth drawing out carefully. Two sets like these are often described as though they were somehow the same underneath, and they are not. The two sets agree at exactly one place and disagree everywhere else, and the place they agree is the one place the fit was looking.

The cleanest way to see it is to line up the cumulants of the log return, order by order. A cumulant is a number that describes the shape of a distribution at one particular order: the first is the mean, the second is the variance, the third carries the lean, the fourth carries the weight in the tails. The jump part contributes to every one of them, and its contribution at order n is the arrival rate multiplied by the average of the jump size raised to the power n.

The jump contribution at every order, and the ratio between the two sets
$$ \kappa_n \;=\; \lambda T \,\mathbb{E}\!\left[Y^n\right] \qquad\Longrightarrow\qquad \frac{\kappa_n^{\,B}}{\kappa_n^{\,A}} \;=\; 4 \cdot 2^{-n} \;=\; 2^{\,2-n} $$
\(\kappa_n\)the jump part's contribution to the cumulant of order \(n\) of the log return over the horizon
\(Y\)the logarithm of one jump's size, normally distributed with mean \(\mu_J\) and spread \(\sigma_J\)
\(\mathbb{E}[Y^n]\)the average of that logarithm raised to the power \(n\), halving in scale each time \(n\) rises by one when the jump is halved
\(\kappa_n^{\,A}, \kappa_n^{\,B}\)that contribution under set A and under set B respectively
What it says in wordsGoing from set A to set B multiplies the arrival rate by four and divides every jump size by two, so a cumulant of order n is multiplied by four and divided by two to the power n at the same time. The ratio between the two sets is therefore four divided by two to the power n, and that ratio equals one only when n is two. The second order is the single place in the whole ladder where the two sets agree, and it is agreement by construction rather than by coincidence.

Put numbers on it. At the first order the jump contributions are minus 0.025000 and minus 0.050000, a ratio of 2. At the second they are 0.006250 and 0.006250, a ratio of 1. At the third they are minus 0.000812500 and minus 0.000406250, a ratio of one half. At the fourth they are 0.000228125 and 0.000057031, a ratio of one quarter. Plot those four ratios on a doubling scale and they fall along a perfectly straight line that crosses the level of exact agreement at the second order and nowhere else.

The ratio between the two sets, order by order. It equals one exactly once. SET B DIVIDED BY SET A, ON A DOUBLING SCALE a ratio of one, meaning exact agreement 2.00 1.00 0.50 0.25 2.00 1.00 0.50 0.25 FIRST the mean SECOND the variance THIRD the lean FOURTH the tail The pale marker is the only order the fit was aimed at, and it is the only order that agrees. Educational illustration. Invented parameters, computed from the moment formulae.
The ratio between the two sets falls along a straight line on a doubling scale, from 2.00 at the first order down to 0.25 at the fourth, touching exact agreement only at the second. Identification is a property of which order the targets happen to include.

One order on that ladder is a trap and needs a note. The first order differs by a factor of two, so the mean of the log return might be expected to separate the two sets. The mean does not separate them, and the reason lies in the pricing theory rather than in the arithmetic. Under the risk-neutral measure Q the drift of the continuous part is not free. The drift is set to make the discounted process a martingale, and it therefore takes whatever value the jump terms leave it. The compensator is minus 0.022001 under set A and minus 0.046940 under set B, and the continuous drift absorbs the difference exactly. Both sets give the same expected value at the horizon, Rs 105.13/-, and they give it for the same reason a fitted parameter is not a discovery: the constraint put it there, not the data.

So the first order is pinned by no-arbitrage, the second is pinned by the construction of the pair, and the third and the fourth are free to differ. The fourth momentThe quantity carrying the weight in the tails of a distribution. The fourth moment distinguishes the two parameter sets here. is therefore the quantity to reach for. The fourth moment is the first target that could be added which carries real separating power and which nothing else has already fixed.

The same model. The same search. Two different target sets. TARGETS: TOTAL VARIANCE ONLY. THE MISS AGAINST THE TARGET, AS THE INTENSITY MOVES. 0.000 0.040 0.080 Both panels share this vertical scale, so the emptiness above the line is the finding. set A set B The miss is exactly nought at every setting on the track. There is no downhill direction, so the search stops where it started. 0.5 1.0 1.5 2.0 INTENSITY, ARRIVALS A YEAR TARGETS: TOTAL VARIANCE AND THE FOURTH MOMENT. THE MISS ON THE FOURTH MOMENT. 0.000 0.080 0.040 the two curves cross at 0.800000, where the miss is 0.039993 either way solid curve reaches nought here dashed curve reaches nought here Whichever fourth moment the targets carry, the miss now has one lowest point. Both cases are drawn, because nothing here decides which of the two the targets would carry. Educational illustration. Invented parameters, computed rather than sampled.
With the total variance as the only target the miss is exactly nought at every setting, so the search has no downhill direction anywhere. Adding the fourth moment replaces that flat floor with a single lowest point, whichever of the two fourth moments the targets happen to carry.
Try it out

Which quantity distinguishes the two sets here?

Try it out

How large is the ratio between the two excess kurtosis figures?

How is a lack of identification detected?

The answer looks fine, so looking at the answer detects nothing. The residual is nought, and nought is what a good fit produces, so the residual gives nothing away either. Detection needs the one thing that changes between two runs of the same fit and nothing else: refit from a different starting point, with the same targets and the same bounds, and see whether the same answer comes back.

Run it from an intensity of 0.6 and it reports something near 0.6. Run it from 1.8 and it reports something near 1.8. Both answers carry a residual of nought against the variance target, so neither can be dismissed as a failed run. Two different answers, both reproducing the targets exactly, is the signature, and no single run of any search can produce it. The test therefore needs at least two runs to say anything at all.

There is a second check that costs nothing and catches the same thing from the other side. The answer the search returned is taken, one parameter is moved by a small amount, the others are moved to hold the target fixed, and the loss is watched to see whether it moves. If it does not, that is a direction in the parameter space along which the targets are flat, and any point along it would have served. Here that direction is written down explicitly in the rescaling above, so it does not have to be hunted for, but on a real fit with several parameters at once it usually has to be found by walking.

The test that separates a determined answer from a stopping point. REFIT FROM SEVERAL STARTING POINTS Same targets. Same bounds. Same loss. Only the start changes. One run alone can never answer this question. EVERY START RETURNS THE SAME ANSWER The targets do have a preference, and the search found it from anywhere DIFFERENT STARTS, DIFFERENT ANSWERS and every one of them reproduces the targets just as well as the others IDENTIFIED BY THESE TARGETS Report the answer, and report which targets it is identified by NOT IDENTIFIED BY THESE TARGETS Report the set of answers, and what would separate them Neither outcome is a statement about the parameters. Both are statements about the targets. Educational illustration.
Refitting from several starting points is what separates a determined answer from a stopping point, because two different answers that both reproduce the targets can only mean the targets do not choose. A single run reports its own stopping place either way.
Try it out

How is a failure of identification detected?

Debt Capital Markets Bootcamp — Fin Maverick Regression for Finance — free micro-course from Fin Maverick

What does an unidentified fit look like from the outside?

Exactly like an identified one. Every instinct a careful reader has been trained to use fails at this point, so the answer is worth stating baldly. The output sheet carries a parameter, a value to six decimal places, a residual, a date, a note of the bounds and a note of the loss. Every field is filled. Every field is correct. And nothing anywhere on the face of it records that the value in the second field was chosen by the starting guess rather than by the targets.

The error that gets made, and what it costs

A fit is run, an intensity of 0.500000 comes back, and it is written into a document as a finding: the process jumps about once every two years. Nobody lied and nobody miscalculated. The search really did return 0.500000 and it really does reproduce the target exactly. But a different starting point would have returned 2.000000 with precisely the same residual and precisely the same confidence, and the document would then have said the process jumps about twice a year. Two documents, one dataset, opposite descriptions of how the world behaves, and no way to choose between them from anything either document contains.

The cost is a spurious precisionDecimal places carried on a number that nothing in the data selected, which is invisible from outside because the output looks the same either way. that is invisible in every output. The residual is nought, so no wide error bar appears. The same starting point gives the same answer every time, so a rerun shows nothing either. The precision shows itself only where the fitted model is asked a question the targets never covered, and by then the number has usually been copied into three other places and has acquired the authority of something that was measured.

The reader who takes the reported intensity of 0.500000 as a fact about anything has read a number that the fit had no basis for choosing. And the same is exactly as true of a reported 2.000000. The targets prefer neither reading, and nothing else the fit produced prefers one either.

Two output sheets. One of them was chosen by the data and one by the starting guess. CALIBRATION OUTPUT, RUN ONE Parameter jump intensity Fitted value 0.500000 Residual at the target 0.000000 Loss squared variance error Bounds 0.100000 to 5.000000 Convergence reported as successful Starting point not recorded Refits from other starts not recorded CALIBRATION OUTPUT, RUN TWO Parameter jump intensity Fitted value 2.000000 Residual at the target 0.000000 Loss squared variance error Bounds 0.100000 to 5.000000 Convergence reported as successful Starting point not recorded Refits from other starts not recorded Every field is filled and every field is correct on both sheets. Nothing on either face records that the two blank rows are what decided the fitted value. Educational illustration. Both sheets are invented, and neither one is the right answer.
Two output sheets from the same targets and the same model report an intensity of 0.500000 and 2.000000, each with a residual of nought and a successful convergence. The only rows that would have told the reader anything are the two that were left blank.
The fit looks settled from outside while two parameter sets match. See what moves.

What should be reported when parameters are not identified?

Not one of the answers. The set of them, and the quantity that would have separated them. The report runs shorter than most people expect and does more than a single number with six decimal places ever does. A reader learns exactly what they can and cannot do with what they have been handed.

Four things belong on the sheet. First, the targets, listed rather than described, as the whole finding is about them. Second, the set of parameter settings that reproduce those targets: here the entire curve running from an intensity of 0.5 through to 2.0 and beyond in both directions, usually written down as a relationship rather than enumerated. Third, the quantity that would separate them, named and quantified: here the excess kurtosis at 0.106647 against 0.026662. Fourth, the plain statement that the fit did not choose, and that any single value reported from it is a stopping point.

The test of the sheet is whether a reader who has never seen the fit could tell, from the sheet alone, that a second answer exists. A single number fails that test always. A number with a note saying the search converged fails it too. Convergence is the wrong evidence: an unidentified search converges perfectly well, it just converges wherever it began.

The sheet that says what the fit actually established. WHAT THIS FIT ESTABLISHED, AND WHAT IT DID NOT 1 THE TARGETS, LISTED Total variance of the log return over one year: 0.046250. Nothing else was targeted. 2 THE SET OF SETTINGS THAT REPRODUCE THEM Every setting with an arrival rate times an average squared jump of 0.006250. Two named points: intensity 0.5 with spread 0.100000, and intensity 2.0 with spread 0.050000. Neither is preferred. 3 WHAT WOULD SEPARATE THEM Excess kurtosis: 0.106647 against 0.026662, a ratio of exactly 4.00. Add it and the fit chooses. 4 THE PLAIN STATEMENT These targets did not choose. Any single value quoted from this fit is a stopping point. A reader who has never seen the fit can tell from this sheet alone that a second answer exists. Four rows replace one number, and the four rows are the part that can be checked. Educational illustration. Every value on the sheet is invented.
The honest artefact carries the targets, the whole set of settings that reproduce them, the quantity that would separate them at 0.106647 against 0.026662, and a plain statement that the fit did not choose. Four rows replace one number.
Try it out

Which of these should be reported when parameters are not identified?

How does somebody actually use this when a fitted number lands on their desk?

The situation is ordinary. A number arrives, attached to a model, with a note saying it was calibrated. The number is going to be used for something. The useful question is not whether to trust it, a question too blunt to help, but which uses it supports and which it does not.

Five questions, in the order that gets to the answer fastest

  1. What was targeted, exactly? Not the model's subject, but which quantities the loss was computed on. Everything else on this list depends on the answer, and a vague answer here means every question below is unanswerable.
  2. Was it refitted from more than one starting point? If the answer is no, nobody knows whether the number was chosen by the targets, and the sheet cannot settle it either way. The refit question costs the least and settles the most.
  3. Is there a rescaling that holds the targets fixed? On a jump component the answer is usually yes and it is usually the one set out here: raise the arrival rate, shrink the jumps, and the variance does not move. Knowing the shape of the flat direction shows which parameters are the risk.
  4. What is this number about to be used for, and is that quantity in the targets? If the use is inside the target set, the choice does not matter. If it is outside, the choice may matter by a factor of four, and the fit had no opinion about it.
  5. What would separate the candidates, and can it be added? Often it can, and cheaply. Adding one quantity to the targets is a smaller change than rebuilding the model, and it is the change that turns a stopping point into an answer.

The fourth question is the one that carries the weight in practice. A fitted parameter is a serviceable input to a calculation that stays inside what it was fitted to, and it is a hazard the moment it leaves. Here the boundary is precise and can be stated in one line: any question that depends only on the total variance is answered identically by both sets, and any question that depends on the tail, on the lean, or on how often anything happens is answered differently by a factor of two or four. The use a number is put to comes before whether it is right. For half the uses set out here, which of the two sets is correct never arises at all.

Everywhere

Where this holds

Identifiability is not jurisdictional. Whether a set of targets can distinguish two parameter settings is a property of the arithmetic relating them, and the arithmetic does not change with the place it is written down. The conventions that decide what a target quantity even means, such as how a horizon is turned into a fraction of a year or how a quoted number is defined, do vary by place, and those are set out under market convention. Any such convention should be confirmed at its own source rather than from a treatment of fitting.

Calibration itself, meaning the targets, the loss, the bounds and the search, is set out under calibration, and testing a fitted model on numbers it never saw is set out under backtesting. A full treatment of any jump model belongs with the stochastic processes it is built from, and what any contract pays is settled elsewhere. Which of the two parameter sets is correct is left open, as nothing in the targets used here says so either.
Breaking Into Quants Bootcamp — Fin Maverick

References

SourceDocumentWhere
arXiv, Quantitative FinancePreprints on parameter identification, inverse problems and moment matching in derivative pricing modelsarxiv.org
Social Science Research NetworkWorking papers on calibration practice, parameter risk and the reporting of fitted quantitiesssrn.com
Robert C. Merton, 1976Option pricing when underlying stock returns are discontinuous, the jump component added to the continuous process hereJournal of Financial Economics

The standard process, the two jump parameter sets and every figure computed from them are invented.
Educational material. Not advice on any investment, tax, budget or market position.

Covered in this topic

Subtopics

Parameter Risk
← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.