Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quant Analyst · CoreTrack
1Quantitative Methods, Financial Data & Programming
iProbability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
iiStatistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
iiiCorrelation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
ivTime Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
vSimulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
viOptimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
viiModelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
viiiBacktesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
ixData Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
xProgramming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
xiQuantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity
2Stochastic Calculus & Derivative Pricing Theory
iProbability Foundations
The Probability SpaceRandom VectorsSigma-AlgebraExpectationSample Space and EventsDensity and Distribution FunctionsRisk-Neutral ProbabilityState Price Density vs…
iiStochastic Processes and Jumps
Properties of a Stochastic ProcessMartingaleBrownian Motion and Its PropertiesBrownian Motion vs Geometric…Stopping TimeThe Markov PropertyState VariablesTransition ProbabilityQuadratic VariationQuadratic Variation vs Ordinary…Submartingale and SupermartingaleMartingale RepresentationMarkov Process vs MartingaleOptional StoppingFiltrationJump ProcessesThe Poisson ProcessLevy ProcessesJump Diffusion
iiiIto Calculus
The Ito IntegralThe Ito Integral vs the Riemann IntegralInfinitesimals in Stochastic CalculusQuadratic CovariationIto's LemmaHow to Apply Ito's…The Infinitesimal GeneratorIto Calculus vs Ordinary Calculus
ivStochastic Differential Equations
Stochastic Differential EquationsStochastic Differential Equation vs…Drift and DiffusionStrong and Weak Solutions ComparedDiscretisationGeometric Brownian Motion
vPricing Theory and No-Arbitrage
No-ArbitrageGirsanov, Radon-Nikodym and Change…Physical and Risk-Neutral Measures…The Fundamental Theorems of…The Law of One PriceThe Pricing KernelDiscount Factors and Zero-Coupon PricesReplication vs HedgingComplete Market vs Incomplete MarketClearing Margin Architecture
viOption Pricing Theory
European and American OptionsMonte Carlo European OptionThe Black-Scholes PDEBlack Scholes and the GreeksThe Payoff FunctionThe Binomial ModelBinomial Option PricingDelta Hedging in TheoryBoundary, Initial and Terminal ConditionsThe Exercise BoundaryHow to Check Put-Call…
viiVolatility Models
Constant, Local and Stochastic…Vasicek Model vs CIR ModelThe Heston ModelThe SABR ModelThe Volatility ProcessImplied VolatilityVolatility Smile vs Skew vs Surface
viiiInterest Rate Models
Interest-Rate DerivativesMean ReversionThe Zero-Coupon BondThe Ornstein-Uhlenbeck ProcessThe Discount CurveZero RatesShort-Rate Model vs Market Model
ixNumerical Pricing
Closed Form and Numerical…Monte Carlo PricingEuler and Milstein Schemes ComparedTree MethodsFinite Difference MethodsNumerical Error and StabilityVariance Reduction
xCalibration and Model Risk
Model OverrideMarket Price and Model PriceCalibrationHow to Document a Pricing ModelThe Educational Illustration LabelMarket ConventionsModel Uncertainty and LimitationsBacktesting a Pricing ModelIdentifiabilityCalibrated ParametersThe Calibration Loss Function

The Infinitesimal Generator: How a Process Acts on a Function

The infinitesimal generator is the operator that takes a function of the process and returns the instantaneous rate of change of that function's average. The generator is read straight off the drift and the diffusion: the drift multiplied by the first derivative, plus half the variance rate multiplied by the second. The operator turns a process into something that acts on functions rather than something that merely moves.

The chain rule for random paths comes with a procedure for applying it. Between them they answer a specific question: given a process and a function of it, how does that function move? The answer that comes back has two terms, one predictable and one not, and the unpredictable term is usually the larger of the two. The split is honest and it is also awkward. Most of the questions worth asking are not about one path at all, but about the centre of the whole set of paths.

The infinitesimal generatorThe operator returning the instantaneous rate of change of the average of a function of the process. is what remains when the chain rule's answer has the term with no average deleted. The remainder is a single expression built from two derivatives, and it answers the question about the centre directly. Deleting the random term is allowed rather than sloppy, the expression returns a rate in use, and one reading of that rate is wrong.

Where does this operator come from?

Start with the standard process that Ito calculus uses throughout. The standard process is an invented traded quantity, written S with a time subscript, starting at Rs 100/-, drifting at 8 per cent a year and carrying a volatility of 20 per cent a year. The process is written in drift and diffusion form. One term is proportional to the length of a time step and one term is proportional to the increment of the Brownian motion driving it.

The process, written in the form everything else is read from
$$ dS_t \;=\; \mu\, S_t\, dt \;+\; \sigma\, S_t\, dW_t $$
\(S_t\)the standard process at time \(t\), the invented traded quantity used throughout, starting at Rs 100/-
\(\mu\)the drift, 8 per cent a year, decimal 0.08, an invented parameter
\(\sigma\)the volatility, 20 per cent a year, decimal 0.20, so the variance rate \(\sigma^{2}\) is 0.04 exactly
\(W_t\)standard Brownian motion under the physical measure P, the source of the randomness
\(dt\)the time step symbol, of the order of the step itself
\(dW_t\)the Brownian increment symbol, of the order of the square root of the step
What it says in wordsOver any short stretch the standard process moves by two terms added together. The predictable one is the drift multiplied by the length of the stretch. The other is the volatility multiplied by the Brownian increment over that stretch, and it has no average at all. Both terms are scaled by the level the process currently sits at.

Now put a function on top of it. Ito's chain rule for random paths says what a function of the process does. The rule gives back a two term answer of exactly the same shape as the process itself: a term in the time step and a term in the Brownian increment. The squared increment does not vanish, so the time step term carries three parts rather than one. The novelty is settled under Ito's lemma and is not reopened here.

The chain rule's answer, split into its two parts
$$ df(S_t) \;=\; \underbrace{\left[\mu S_t f'(S_t) + \tfrac{1}{2}\sigma^{2} S_t^{2} f''(S_t)\right]}_{\text{predictable over the step}} dt \;\;+\;\; \underbrace{\sigma S_t f'(S_t)}_{\text{no average at all}}\, dW_t $$
\(f\)any function of the level of the process that can be differentiated twice, such as the level itself, its logarithm or its square
\(f'\)the first derivative, the slope of that function at the level the process currently sits at
\(f''\)the second derivative, the rate at which the slope itself changes, and therefore the curvature
\(\tfrac{1}{2}\sigma^{2}\)half the variance rate, 0.02 exactly for the standard process
What it says in wordsThe movement of a function of the process over a short stretch is a predictable part, made of the drift acting through the slope plus half the variance rate acting through the curvature, added to a random part which is the volatility acting through the slope and which has no average.

Everything now turns on the second bracket. The term in the Brownian increment is an Ito integral in the making, and an Ito integral has an average of nothing. The zero average was settled when the Ito integral was built: reading the integrand at the start of each step is exactly what makes the sum a fair game, and a fair game does not go anywhere on average. So in the average of the whole expression, the second bracket contributes zero and disappears.

The generator is the first bracket, standing on its own, and it is the chain rule with the term that has no average deleted. Nothing has been approximated and nothing has been dropped because it was small. The term was removed because averaging it gives zero exactly, and what remains is a complete answer to a different question from the one the chain rule was answering.

The operator is not a new object. It is the chain rule with one term averaged away. STEP 1: WHAT THE CHAIN RULE GIVES drift on the slope, plus half the variance rate on the curvature volatility on the slope, riding the Brownian increment STEP 2: TAKE THE AVERAGE the first row survives the second row is a fair game, so it averages to zero STEP 3: WHAT IS LEFT THE OPERATOR one expression in two derivatives, returning the rate of change of the average Nothing was approximated at step 2 and nothing was discarded for being small. The red row was removed because its average is zero exactly, which is a property the integral was constructed to have. The two objects therefore agree wherever they overlap. They answer different questions about the same movement. The standard process and its parameters are invented. Educational illustration.
Take the movement the chain rule gives, average away the term that has no average, and what remains is the operator, which is why the two agree wherever they overlap.

How does a process become something that acts on functions?

Turning a process into an operator sounds like a change of subject and is not. Up to now the process has been a thing that moves. An operatorSomething that takes a whole function as its input and returns another whole function as its output, rather than taking a number and returning a number. is a different kind of object entirely: it takes a function in and gives a function back. Not a number in and a number out, but a whole rule in and a whole rule out.

Here is the everyday version. A kitchen scale takes an object and returns a number, so it is an ordinary function. A recipe scaler is not like that. The scaler takes in a whole recipe, every ingredient and every quantity, and hands back a whole recipe. The input was not a number and neither was the output. The recipe scaler is the shape of an operator. Questions can then be asked of the machine itself rather than of any one recipe: which recipes does it leave unchanged, which does it flatten, what does it do to the ratio between two ingredients.

Turning the process into an operator is a deliberate change of viewpoint, from asking where the process goes to asking what it does to every function of itself at once. The payoff is worth naming now so the change of viewpoint does not feel like decoration: once the operator is in hand, a whole class of questions about averages is answered by differentiating twice, with no equation solved anywhere.

The definition, before the convenient form, is a limit. Start the process at a stated level, wait a very short time, take the average of the function over everything the process might have done in that time, subtract the value the function had at the start, and divide by the length of the wait. The quotient is a rate of change of an average. The generator is what it settles down to as the wait shrinks toward nothing.

The definition, before any convenient form
$$ (\mathcal{A}f)(x) \;=\; \lim_{h \downarrow 0} \; \frac{\mathbb{E}\bigl[\, f(S_h) \,\bigm|\, S_0 = x \,\bigr] \;-\; f(x)}{h} $$
\(\mathcal{A}\)the generator, written as a script A throughout this guide, an operator rather than a number
\(f\)the function being fed in, and it must be twice differentiable for the limit to exist
\(x\)the level the process is started from, here Rs 100/- wherever a level is needed
\(h\)the length of the short wait, shrunk toward nothing in the limit
\(\mathbb{E}[\,\cdot\mid S_0 = x\,]\)the average over everything the process might do, given that it started at that level
What it says in wordsThe generator applied to a function, at a stated starting level, is the average value of that function after a very short wait, less its value at the start, divided by the length of the wait, in the limit as the wait shrinks to nothing.

The whole content of the operator is sitting in the numerator. Read it carefully. The numerator is not the change in the function. It is the change in the average of the function. The expectation sign is doing all the work, and quietly dropping it is the single commonest mistake made with the operator.

Notice also how little the definition needs. The definition asks for no solution to the process, no distribution written down and no particular path. The requirements are a starting level and a function that bends smoothly enough to be differentiated twice. The smoothness requirement is the reason the operator cannot be applied to absolutely anything: a function with a sharp corner has no second derivative at the corner, and the limit above will not settle there.

Try it out

What kind of thing does an operator take in, and what does it give back?

How is it read straight off the drift and the diffusion?

The limit above is the honest definition and it is useless for calculation. The chain rule has already done the work, so the limit never has to be computed. The predictable bracket from earlier is exactly that limit, so the generator can be written down by copying two coefficients out of the process and putting them onto two derivatives.

The generator of the standard process, in the form actually used
$$ (\mathcal{A}f)(S) \;=\; \mu\, S\, f'(S) \;+\; \tfrac{1}{2}\,\sigma^{2} S^{2}\, f''(S) \;=\; 0.08\,S\,f'(S) \;+\; 0.02\,S^{2}\,f''(S) $$
\(\mu S\)the coefficient of the time step in the process, copied across unchanged
\(\sigma S\)the coefficient of the Brownian increment in the process, entering squared and halved
\(f'(S)\)the first derivative of the function, evaluated at the level in question
\(f''(S)\)the second derivative, evaluated at the same level
\(0.02\)half the variance rate, being half of 0.04, the number every second derivative in this guide is multiplied by
What it says in wordsThe generator of the standard process is the drift coefficient multiplied by the first derivative of the function, plus half the square of the diffusion coefficient multiplied by the second derivative, so writing it down needs only the two coefficients already written in the process itself.

Two things are worth saying about that. The first is that the drift termThe part of the generator carrying the first derivative, built from the drift coefficient of the process. is the ordinary chain rule and nothing more: rate of change of the input, times the slope. Anybody who has met calculus already has it. The second is that the diffusion termThe part of the generator carrying half the variance rate and the second derivative, which is what randomness contributes. is the entire contribution of the randomness, and it enters through the curvature rather than through the slope. Randomness has no average direction, so it cannot push the average of a straight function anywhere. Randomness can only push the average of a bending one, and which way it pushes is decided by the direction of the bend.

Writing the operator down is a copying exercise rather than a derivation, and every treatment writes it in the same breath as the process. Square the diffusion coefficient, halve it, attach it to the second derivative; take the drift coefficient as it stands and attach it to the first. There is no third step.

Two coefficients, copied down. That is the whole construction. THE PROCESS, AS WRITTEN move in S = 0.08 times S on the time step + 0.20 times S on dW copied across unchanged squared, then halved: 0.04 S squared, then 0.02 THE OPERATOR, WITHOUT A SINGLE LINE OF DERIVATION apply to f = 0.08 S times first slope + 0.02 S squared times the curvature The left half is the ordinary chain rule. The right half is the whole contribution of the randomness, and it reaches the answer only through the curvature, never through the slope. The drift of 8 per cent and the volatility of 20 per cent are invented parameters. Educational illustration.
The drift multiplies the first derivative and half the variance rate multiplies the second, so the operator can be written down the moment the process is written down.

The choice of the curvature rather than the slope is usually accepted without being examined, and it is worth one worked look. Sit at Rs 100/- and suppose the process is equally likely to be at Rs 120/- or at Rs 80/- a moment later. The two outcomes are a fair coin: the average level afterwards is still Rs 100/-, so randomness has moved nothing. Now ask what it did to the average of a function of the level.

For the level itself, nothing at all. Half of 120 and half of 80 is 100, and the average is exactly where it started. For the square, half of 14,400 and half of 6,400 is 10,400 against a starting 10,000, so the average has been lifted by 400. For the logarithm, half of 4.787492 and half of 4.382027 is 4.584759 against a starting 4.605170, so the average has been pushed down by 0.020411. Same symmetric jump, three completely different effects, and the only thing that differed between the three functions was how they bend.

Function of the levelValue at Rs 80/-Value at Rs 120/-Average of the twoShift from Rs 100/-
the level itself, no bend80.000000120.000000100.0000000.000000
the square, bends upward6,400.00000014,400.00000010,400.000000+400.000000
the logarithm, bends downward4.3820274.7874924.584759-0.020411

A symmetric jump leaves the average of a straight function alone, lifts the average of one that bends upward and lowers the average of one that bends downward. Randomness therefore reaches the answer through the second derivative rather than the first. The size of the jump was not chosen at random either. The volatility multiplied by the level is 0.20 times Rs 100/-, or Rs 20/- exactly, so the jump used above is the size the process actually carries over a year. Half the squared jump is 200, and multiplying that by the second derivative reproduces the diffusion term precisely: 200 times 2 is the 400 the square gained, and 200 times minus 0.0001 is the minus 0.02 the logarithm lost. The logarithm's exact shift of minus 0.020411 differs slightly from minus 0.020000 because a jump of Rs 20/- is not small, and the gap between the two is everything the operator discards by working in the limit.

Try it out

What multiplies the second derivative in the generator of the standard process?

What does applying it actually give?

Feed a function in and a number comes back at every level. The number that comes back is a rate, measured per year, and it is the rate at which the average of the function is changing at this instant. Three words in that sentence carry weight and each is worth separating.

It is a rateA quantity per unit of time, so it must be multiplied by a stretch of time before it becomes a movement., so it has a unit of time attached and is not a movement. It is instantaneousHolding at this moment only, rather than over any stretch of time, so it will generally be a different number a moment later., so it holds now and will generally be a different number once the process has moved. And it is a rate of change of an averageThe value taken over everything the process might do from here, rather than the value on the one thing it actually does., taken over everything the process might do from where it currently is, rather than over the one thing it will actually do.

Picture a tray of ball bearings on a table that is vibrating gently while the table itself is being tilted very slowly. No single bearing moves smoothly. Each one rattles, reverses, jumps sideways. But the centre of the whole tray slides steadily downhill, and that steady slide is a real, computable, useful quantity even though no individual bearing is doing it. The generator returns the speed of the centre of the tray. The generator says nothing whatever about any one bearing.

The number the operator returns describes the centre of a distribution of outcomes, and no path is obliged to resemble it. The distinction between the centre and the path is the entire content of the operator, and it is the first casualty of a hurried reading.

Try it out

Applying the operator to the logarithm of the standard process returns 0.060000 a year. Is that the rate the logarithm is changing at on a given path?

Derivatives Foundation Bootcamp — Fin Maverick Value at Risk and What It Hides — free micro-course from Fin Maverick

What do three applications give on the standard process?

Abstraction earns nothing until it produces numbers, so here are three, all exact, all on the same invented process with the same two invented parameters, and all at a level of Rs 100/-. The three functions are the level itself, the logarithm of the level and the square of the level. Nothing else changes between them. The operator is fixed; only what is fed into it moves.

Try it out

The operator applied to the level itself, at Rs 100/-. What number comes back?

Take the level first, meaning the function that hands back whatever it is given. A straight line has no bend at all, so its first derivative is 1 everywhere and its second derivative is 0 everywhere. The diffusion term vanishes completely and only the drift term survives: 0.08 multiplied by Rs 100/-, or 8.000000 rupees a year. Read it in words. At this instant the average level of the process is rising at Rs 8/- a year.

Take the square of the level next. Its first derivative is twice the level, or 200 at Rs 100/-, and its second derivative is 2 everywhere. The drift term is 0.08 times Rs 100/- times 200, or 1,600. The diffusion term is 0.02 times 10,000 times 2, or 400. Adding them gives 2,000.000000 a year, in units of rupees squared. Note how much of that came from the curvature: a quarter of the total, contributed by randomness alone, on a function that bends upward.

The third is the logarithm, and it is the one worth stopping for, so it is held back for one moment.

Try it out

The operator is about to be applied to the logarithm of the level. Ahead of switching to it below: what number should be expected?

Play with it

One operator, five functions, and the answers are nothing like each other

The operator never changes. Only the function fed into it does. The upper panel draws the chosen function as a solid line and its tangent at Rs 100/- as a dashed one, so the space between them is the curvature the second term is paid to notice. The lower panel draws the two terms of the operator against a zero line. The default is the logarithm, and it returns 0.060000 a year, exactly the figure the chain rule reached by a completely different route.

THE FUNCTION (SOLID) AND ITS TANGENT AT Rs 100/- (DASHED). THE GAP IS THE CURVATURE. Rs 50/- Rs 100/- Rs 150/- Rs 200/- bends downward THE TWO TERMS OF THE OPERATOR, SIGNED, DRAWN AGAINST A ZERO LINE DRIFT TERM: 0.08 times the level times the first derivative +0.080000 DIFFUSION TERM: 0.02 times the level squared times the second derivative -0.020000 THE OPERATOR RETURNS 0.060000 A YEAR
First derivative at Rs 100/-
0.010000
Second derivative at Rs 100/-
-0.000100
Drift term
0.080000
Diffusion term
-0.020000
What the operator returns
0.060000
Applied to the logarithm at a level of Rs 100/-, the first derivative is 0.010000 and the second is -0.000100, so the drift term is 0.080000 and the diffusion term is -0.020000, and the operator returns 0.060000 a year. Every level cancels, so this answer is the same at any level, and it is exactly the corrected growth rate the chain rule produced by a completely different route.
Educational illustration. Every reading is computed from the operator formula at a level of Rs 100/-, never sampled from any random draw, so the default reproduces the worked example exactly on every reload. All five answers as static text: the level itself, with derivatives 1 and 0, returns 8.000000 rupees a year. The straight line 2 times the level plus 25, with derivatives 2 and 0, returns 16.000000 rupees a year. The logarithm, with derivatives one over the level and minus one over the level squared, returns 0.060000 a year, matching the chain rule result exactly. The square of the level, with derivatives twice the level and 2, returns 2,000.000000 a year. One over the level cubed, with derivatives minus three over the level to the fourth and twelve over the level to the fifth, returns 0.000000 a year. The standard process, its drift of 8 per cent and its volatility of 20 per cent are all invented and describe no market.

So the logarithm returns 0.060000 a year, and every level cancelled on the way. Its first derivative is one over the level and its second is minus one over the level squared. The drift term is 0.08 times the level times one over the level, or just 0.08, and the diffusion term is 0.02 times the level squared times minus one over the level squared, or just minus 0.02. The remainder is the drift less half the variance rate.

The 0.060000 is exactly the figure the chain rule produced, and it was reached here by a completely different route. There, the correction appeared by expanding a function of the process and collecting the surviving terms of the expansion. Here, no expansion happened at all: two derivatives were written down and multiplied by two coefficients copied out of the process. Two independent routes arriving at the same six decimal places is the strongest evidence available that the machinery is internally consistent, and it is worth more than any single derivation checked twice.

The three applications, side by side
$$ \mathcal{A}\,S \;=\; 8.000000 \qquad\quad \mathcal{A}\,\ln S \;=\; \mu - \tfrac{1}{2}\sigma^{2} \;=\; 0.060000 \qquad\quad \mathcal{A}\,S^{2} \;=\; \bigl(2\mu + \sigma^{2}\bigr)S^{2} \;=\; 2{,}000.000000 $$
\(\mathcal{A}S\)the operator applied to the level, in rupees a year, evaluated at Rs 100/-
\(\mathcal{A}\ln S\)the operator applied to the logarithm, a pure rate per year with no rupee unit, and the same at every level
\(\mathcal{A}S^{2}\)the operator applied to the square of the level, in rupees squared a year, evaluated at Rs 100/-
\(2\mu+\sigma^{2}\)0.20 exactly, being twice 0.08 plus 0.04
What it says in wordsThe same operator applied to the level, to the logarithm and to the square of the level at Rs 100/- returns 8.000000 rupees a year, 0.060000 a year and 2,000.000000 rupees squared a year, and only the logarithm gives an answer that is free of the level altogether.

Each of the three can be checked against a completely separate calculation. Doing that once stops the operator feeling like a trick. The average level at a later time is Rs 100/- multiplied by the exponential of 0.08 times the elapsed time, and differentiating that at the start gives 8. The average logarithm is the logarithm of Rs 100/- plus 0.06 times the elapsed time, whose slope is 0.06. The average of the square is 10,000 multiplied by the exponential of 0.20 times the elapsed time, whose slope at the start is 2,000. The operator reproduced all three without any of those expressions being written down.

One operator, three functions, three answers that share nothing but the machine that made them. THE FUNCTION FIRST DERIVATIVE SECOND DERIVATIVE DRIFT TERM DIFFUSION TERM WHAT COMES BACK THE LEVEL hand back the level 1 0 8 0 8.000000 THE LOGARITHM bends downward 0.01 -0.0001 0.08 -0.02 0.060000 THE SQUARE bends upward 200 2 1,600 400 2,000.000000 All three at a level of Rs 100/-, on the invented standard process. Educational illustration.
Applied to the level the generator returns 8.000000 rupees a year, to the logarithm 0.060000 a year and to the square of the level 2,000.000000 a year.
Value at Risk and What It Hides teaches you to compute value at risk three ways, interpret the figure, and say precisely what it refuses to describe.

What is the gap between the level and its square showing?

Put the first and third answers into proportional terms so they can be compared. The average level grows at Rs 8/- a year on a level of Rs 100/-, or 0.08 a year in proportional terms. The average of the square grows at 2,000 a year on a square of 10,000, or 0.20 a year in proportional terms. Two rates on the same process, and they are obviously different. The question is what the difference means.

Subtracting them naively gives 0.12 and that number means nothing at all. The correct comparison has one extra step in it, and the extra step is where the payoff lives. If the process had no randomness whatever, its level would grow at 0.08 and its square would grow at exactly twice that, 0.16, for the same reason that doubling in length quadruples an area. So 0.16 is what the square would do if nothing were random. The square actually grows at 0.20.

The square of the level grows at 0.20 a year against the 0.16 that twice the level's own rate would predict, and the excess of 0.04 is the variance rate, obtained by applying one operator to two functions and never solving anything. Reading the variance rate off two derivatives is the moment the operator stops being notation and becomes a tool. The spread of the distribution has been read off from two derivative calculations, with no distribution written down, no equation solved and no path examined.

Reading the variance rate straight off two answers, with nothing solved anywhere. THE LEVEL 0.08 a year TWICE THAT RATE 0.16 THE SQUARE 0.20 0.04, the variance rate All three rates are proportional, at a level of Rs 100/- on the invented standard process. Educational illustration.
The average of the level grows at 0.08 a year and the average of its square at 0.20, and the excess of 0.04 over twice the level's own rate is the variance rate obtained without solving anything.

The reason this works is also the reason the second moment is asked for at all, and it is worth one sentence. The average of a squared quantity always exceeds the square of its average by exactly the spread of the quantity. So a rate for the square that runs ahead of twice the rate for the level can only be the spread opening up, and the speed at which it opens is the variance rate. Here it is 0.04, being 0.20 squared, and it agrees with the parameter the process was written with in the first place.

Try it out

The average of the level grows at 0.08 a year in proportional terms and the average of its square at 0.20. What is the 0.04 that appears when the two are compared properly?

Which functions does it send to zero, and why does that matter?

An operator is best understood by asking what it destroys. Feed in a constant and the answer is obviously zero: both derivatives vanish, so both terms vanish. The constant is not interesting on its own, but it points at a question that is. Are there functions that are not constant which the operator still sends to zero?

Setting the whole expression to zero and asking which functions satisfy it gives an equation in the function rather than in the process, and for the standard process it has a clean answer. Besides the constants, the operator sends one over the level cubed to zero, and nothing else of that shape.

The functions the operator destroys, for the standard process
$$ \mathcal{A}f = 0 \;\;\Longleftrightarrow\;\; 0.08\,S f'(S) + 0.02\,S^{2} f''(S) = 0 \;\;\Longleftrightarrow\;\; f(S) = c_1 + c_2\,S^{-3} $$
\(\mathcal{A}f=0\)the requirement that the operator returns nothing at every level, not merely at Rs 100/-
\(c_1\)any constant, sent to zero by the operator because both its derivatives vanish
\(c_2\,S^{-3}\)any multiple of one over the level cubed, the exponent minus 3 being one less twice the drift over the variance rate
\(-3\)from 1 less 2 times 0.08 divided by 0.04, which is 1 less 4, and it is exact for these parameters only
What it says in wordsThe only functions of the level the operator returns nothing for are the constants and the multiples of one over the level cubed, and that exponent of minus three is fixed entirely by the drift and the variance rate of this particular process.

A stated result is worth less than a checked one. Check it directly at Rs 100/-. One over the level cubed has a first derivative of minus three over the level to the fourth, or minus 0.00000003, and a second derivative of twelve over the level to the fifth, or 0.0000000012. The drift term is 0.08 times 100 times minus 0.00000003, giving minus 0.00000024. The diffusion term is 0.02 times 10,000 times 0.0000000012, giving plus 0.00000024. The two terms are equal and opposite, so the total is 0.000000 exactly. Switch the simulation above to the fifth option and watch the two bars come out mirror images of each other.

A function the operator sends to zero is called a harmonic functionA function the generator returns zero for, whose average therefore stays exactly where it starts. for that process, and the reason it matters is immediate from what the operator returns. If the rate of change of the average is zero at every level, then the average never moves, ever. Whatever the process does over the next decade, the average of that particular function of it is the same number today as it will be then.

Spotting a function the operator sends to zero converts a hard question about the future into a statement that can already be written down. Spotting one replaces a calculation over an unknown future with a reading taken now. The average of one over the standard process cubed is 0.000001 today, at a level of Rs 100/-, and it is 0.000001 at every future date as well, not approximately but exactly. The fixed average is worth a great deal more than it looks, and it is the property the later work in this subject is built on rather than something opened here.

Raise the level to a power, and this is the rate its average grows at. Two powers give exactly nothing. power minus 3 power 0 0.20 0.00 0.36 minus 4 minus 2 0 2 The power 1 sits at 0.08 and the power 2 at 0.20, which are the level and the square. The standard process is invented. Educational illustration.
The rate at which the average of a power of the level grows crosses zero at exactly two powers, zero and minus three, and those two are the only harmonic ones.

The numbers on that curve are the ones that actually get quoted, so it is worth reading as a table as well. Raising the level to a power and feeding the result into the operator makes every level cancel out of the answer in proportional terms, so what comes back is one rate per power, and it is a small piece of arithmetic in the power alone: the drift multiplied by the power, plus half the variance rate multiplied by the power and by one less than the power.

The level raised to this powerRate its average grows atWhat that rate says
minus 4+0.080000rising again, having passed through the bottom and back up
minus 30.000000the harmonic one: this average never moves at all
minus 2-0.040000falling, and it is the same rate as the power minus 1
minus 1-0.040000falling, so the average of one over the level shrinks
00.000000a constant, sent to zero for the obvious reason
1, the level itself+0.080000the drift, as it must be, and Rs 8/- a year at Rs 100/-
2, the square of the level+0.200000twice the drift plus the variance rate, and 2,000 a year at Rs 100/-

Two powers give exactly nothing and every other power gives a rate, so the harmonic property is rare rather than typical and is worth checking for rather than assuming. Notice also the pair at minus 1 and minus 2, returning the identical rate of minus 0.040000 by different splits of the two terms. The repeat is not a coincidence either: the arithmetic in the power is a shape that bends, and any shape that bends hits most of its values twice.

Try it out

A function is sent to zero by the generator at every level. What does that say about its average?

Try it out

For the standard process, which of these does the operator send to zero besides the constants?

What does it give that the equation of the process does not?

The equation of the process is a complete description. Everything true about the standard process is a consequence of it, so in a strict sense the operator adds nothing. The difference is not in the information; it is in what can be reached without effort, and that difference is enormous in practice.

The equation is a statement about how the process moves. Getting from it to a statement about the average of some function of the process normally means solving it, or writing down the distribution at each date, or running a great many outcomes and counting. All three are real work and all three have to be redone for each new function of interest. The operator answers the whole class of questions in one differentiation, without solving anything, and it answers a new one every time a new function is handed to it.

The best everyday version is a difference between two kinds of instruction. One kind describes how a machine is wired: complete, correct, and no help at all when somebody asks what it will cost to run for a month. The other is a small table converting any question about the machine straight into an answer. The wiring diagram contains the table, in the sense that the table could be derived from it, but nobody derives it twice.

Three things become immediate that were not. The rate of change of the average of any twice differentiable function of the process follows from writing down two derivatives. The moments of the process, and therefore its spread, follow from feeding in powers, exactly as was done to extract 0.04 above. And a test for functions whose averages never move at all becomes available, reducing a question about the far future to a reading taken now.

One test, applied to any function that can be differentiated twice, and two very different conclusions. APPLY THE OPERATOR TO THE FUNCTION. IS WHAT COMES BACK ZERO AT EVERY LEVEL? YES NO THE AVERAGE NEVER MOVES Today's value is the answer for every future date, exactly. One over the level cubed does this. A RATE, NOT A LEVEL The average is moving at the rate returned, at this instant only, and it will differ a moment later. Neither branch requires the process to be solved, a distribution to be written down, or a single path to be examined. The standard process is invented. Educational illustration.
A function the generator sends to zero has an average that stays where it is, so spotting one turns a hard question about the future into a statement about the present.
Try it out

What does the operator give that the equation of the process does not?

The reading that breaks it: a rate for the average, quoted as a rate for the path

The operator applied to the logarithm returns 0.060000 a year. The mistake is to carry that number away as though the logarithm of the process were changing at 0.060000 a year. The logarithm is not changing at that rate, and it never is, on any path, at any moment.

The locked path is the twelve step path used unchanged throughout Ito calculus wherever one is drawn. Over any one month the logarithm of the standard process moves by two things added together. One is the drift, 0.06 divided by twelve, giving exactly 0.005 every month without exception. The other is the volatility multiplied by the Brownian increment for that month, a different number every month that can point either way.

In the largest month of the locked path the random part moves the logarithm by 0.092376 while the drift contributes 0.005, so the rate the operator returns accounts for 5.13 per cent of what actually happened. In a typical month of that path the drift is about 9 per cent of the movement. Over month four the increment happened to be tiny, so the drift is very nearly half. There is no month in which it is most of the movement, and the number the operator returns describes none of these months.

The everyday version is a kitchen scale with a shaking needle. The needle jumps between readings that are far apart, and a reading taken at one moment says very little. The average of many readings is a real and useful number. But quoting the average as though the needle were reading it right now is a different claim, and it is false.

The cost of the mistake is a rate presented as a description of a path when it describes only the centre of a distribution of paths. Anybody planning against it will find the outcome nothing like the rate, will conclude that the mathematics is wrong, and will be looking in the wrong place. The mathematics was answering a question about a crowd of outcomes. The plan was asking about one of them.

The straight line is what the operator returns. The jagged line is what the path did. month 6, path is 0.075056 above the line month 9, path is 0.109697 below 0.00 0.06 -0.06 start month 12 ONE MONTH BLOWN UP: MONTH TWO, THE LARGEST MOVE ON THE LOCKED PATH the drift, 0.005, which is what the operator accounts for the random part, 0.092376, which the operator says nothing about at all Both bars to the same scale. The drift is 5.13 per cent of the total move. Invented path. Educational illustration.
On the locked path the logarithm moves by 0.005 from the drift over one month and by 0.092376 from the random part, so the rate the generator returns is about five per cent of the movement.

One detail in that picture rewards a second look. The locked path was built with its twelve driving values summing to zero exactly, and that sends the Brownian motion back to where it started. At the horizon the two lines therefore meet, so over the full year the logarithm moves by 0.060000 and the operator's rate is exactly right. The agreement is a property of how the path was built rather than evidence for anything, and reading it as confirmation would be the same mistake in a different costume.

Breaking Into Quants Bootcamp — Fin Maverick

How does somebody checking a model actually use this?

Somebody reviewing a model rarely gets to see how it was built. A set of outputs arrives with a description of the process behind them, and the question is whether the two hold together. The operator gives that reviewer a check that needs no access to anyone else's code, no data, and no agreement about what the right answer is.

The check runs in three lines. Read the drift and the diffusion coefficients off the process as described. Write the operator by copying them onto the two derivatives. Apply it to whatever quantity the model reports an average growth rate for, and see whether the two agree. If a description says the process drifts at 8 per cent and the output reports the average of the square growing at 0.16 a year in proportional terms, the operator says 0.20, and the disagreement of exactly the variance rate says the randomness has been left out somewhere in the reported calculation.

The most useful version of the check is the one that costs nothing: apply the operator to the level itself and confirm the reported average matches the drift. The test is the cheapest possible one, and it catches the commonest error: a rate quoted on the logarithm reported as a rate on the level, or the reverse. On this process those two differ by 0.02 a year, a gap that sounds small and is not: it is the whole distance between the average and the median at the horizon.

There is a second use, less common but worth knowing. When a reviewer is handed a quantity that is claimed to hold steady on average, the operator settles the claim in one calculation rather than in an argument. Applying it decides the matter. If the answer is not zero at every level, the quantity does not hold steady, whatever the description says, and the rate that comes back shows how fast the claim is failing and in which direction.

A third use belongs to whoever is building rather than reviewing, and it is about knowing when to stop. Anybody who has run a great many outcomes to answer a question about an average has produced a number with a wobble on it, and the wobble is often larger than the effect being looked for. Where the question is about the average of a function of the process over a short horizon, the operator answers it exactly and instantly, and running outcomes to get the same answer is effort spent buying noise. The practical rule is to ask whether the question is about a rate of change of an average at a moment. If it is, differentiate twice. If instead it is about the shape of the whole distribution, or about a quantity that depends on the path rather than on where the path ended, the operator will not reach it and there is no shortcut to be had.

All three uses share one thing that is easy to miss: none of them needs the reviewer and the builder to agree on anything except the drift and the volatility of the process. The drift and the volatility are usually the only part of a description that everybody can see. Everything in this guide was computed from exactly those two, plus a starting level, and that is what makes the check portable between people who share no code.

Solving any equation the operator appears in is covered separately and much later in this subject; the operator is written down here and the equations it sits inside are opened there. The formal link between the operator and an expectation over a horizon carries the names of Dynkin and of Feynman and Kac, and it is covered separately. The versions for several processes at once and for processes that jump are covered separately. Pricing models are covered separately, and what any contract pays is covered separately. No jurisdiction sets the definition of an operator: the statement is mathematical and holds everywhere.

References

SourceDocumentWhere
arXiv Quantitative FinancePreprint repository for generators, semigroups and diffusion operatorsarxiv.org
Social Science Research NetworkWorking paper repository for the same materialssrn.com
ItoThe chain rule and the integral that carry his name, named here for structure onlynamed in the text, no text reproduced
KolmogorovThe backward equation the operator appears in, named here and opened separatelynamed in the text, no text reproduced
DynkinThe link between the operator and an expectation, named here and excluded by scopenamed in the text, no text reproduced
Feynman and KacThe same link in its better known form, named here and excluded by scopenamed in the text, no text reproduced
Hull, Shreve and WilmottStandard textbook treatments of the generator of a diffusionnamed in the text, no text reproduced

The standard process and its locked twelve step path are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.