Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Stochastic Calculus & Derivative Pricing Theory
1Probability Foundations
The Probability SpaceRandom VectorsSigma-AlgebraExpectationSample Space and EventsDensity and Distribution FunctionsRisk-Neutral ProbabilityState Price Density vs…
2Stochastic Processes and Jumps
Properties of a Stochastic ProcessMartingaleBrownian Motion and Its PropertiesBrownian Motion vs Geometric…Stopping TimeThe Markov PropertyState VariablesTransition ProbabilityQuadratic VariationQuadratic Variation vs Ordinary…Submartingale and SupermartingaleMartingale RepresentationMarkov Process vs MartingaleOptional StoppingFiltrationJump ProcessesThe Poisson ProcessLevy ProcessesJump Diffusion
3Ito Calculus
The Ito IntegralThe Ito Integral vs the Riemann IntegralInfinitesimals in Stochastic CalculusQuadratic CovariationIto's LemmaHow to Apply Ito's…The Infinitesimal GeneratorIto Calculus vs Ordinary Calculus
4Stochastic Differential Equations
Stochastic Differential EquationsStochastic Differential Equation vs…Drift and DiffusionStrong and Weak Solutions ComparedDiscretisationGeometric Brownian Motion
5Pricing Theory and No-Arbitrage
No-ArbitrageGirsanov, Radon-Nikodym and Change…Physical and Risk-Neutral Measures…The Fundamental Theorems of…The Law of One PriceThe Pricing KernelDiscount Factors and Zero-Coupon PricesReplication vs HedgingComplete Market vs Incomplete MarketClearing Margin Architecture
6Option Pricing Theory
European and American OptionsMonte Carlo European OptionThe Black-Scholes PDEBlack Scholes and the GreeksThe Payoff FunctionThe Binomial ModelBinomial Option PricingDelta Hedging in TheoryBoundary, Initial and Terminal ConditionsThe Exercise BoundaryHow to Check Put-Call…
7Volatility Models
Constant, Local and Stochastic…Vasicek Model vs CIR ModelThe Heston ModelThe SABR ModelThe Volatility ProcessImplied VolatilityVolatility Smile vs Skew vs Surface
8Interest Rate Models
Interest-Rate DerivativesMean ReversionThe Zero-Coupon BondThe Ornstein-Uhlenbeck ProcessThe Discount CurveZero RatesShort-Rate Model vs Market Model
9Numerical Pricing
Closed Form and Numerical…Monte Carlo PricingEuler and Milstein Schemes ComparedTree MethodsFinite Difference MethodsNumerical Error and StabilityVariance Reduction
10Calibration and Model Risk
Model OverrideMarket Price and Model PriceCalibrationHow to Document a Pricing ModelThe Educational Illustration LabelMarket ConventionsModel Uncertainty and LimitationsBacktesting a Pricing ModelIdentifiabilityCalibrated ParametersThe Calibration Loss Function

Transition Probability: Moving From One State to the Next

A transition probability attaches to three things at once: a starting state, an ending state, and the time between them. A single ending value carries no probability, so for a process with a continuous range of values the transition is written as a density rather than a number. Transitions over two steps must equal transitions over one composed with themselves, and that consistency is what makes a set of transitions a process.

Once the present state is a sufficient summary of the past, everything the model knows about the future has to be expressible as a rule that takes one state to another over a stretch of time. There is nowhere else for the knowledge to live. The rule that takes one state to another is the transition probabilityThe chance of moving from one state to another over a stated stretch of time.. Writing that rule out for every pair of states and every gap leaves nothing further to say about the model.

What does a transition probability attach to?

Ask somebody what the probability of Rs 115/- is and the honest answer is that the question is not finished. Probability of Rs 115/- starting from where? And after how long? A probability of Rs 115/- starting from Rs 111.08/- over six months is a completely different number from a probability of Rs 115/- starting from Rs 94/- over one month, and neither of them is the probability of Rs 115/-.

A transition attaches to a pair of states and a gap. A transition is a three input object, not a property of any one state. Drop the starting state and what remains is a marginal probability, which is a different animal. Drop the ending state and there is nothing to attach a number to. Drop the gapThe stretch of time the transition spans, from the moment the starting state is observed to the moment the ending state is read. and what remains is a quantity with no unit, in the same way that a speed quoted without a unit of time is not a speed.

Think of a lift that knows only which floor it is on. The useful question about it is never the chance of the fourth floor on its own. The useful question is the chance of the fourth floor twenty seconds from now, starting on the second floor. Change the starting floor and the answer changes. Change twenty seconds to five minutes and the answer changes again. The lift is a small memoryless machine and its entire behaviour is the table of those answers, one entry for each starting floor, each ending floor and each waiting time.

A transition is a three input object. Nothing less than three makes a number. INPUT 1: STARTING STATE Rs 111.08/- where the process is observed INPUT 2: ENDING STATE Rs 115.00/- the value being asked about INPUT 3: THE GAP half a year the stretch of time it spans THE TRANSITION: one number for a listed set of states, one density for a continuous range of values Remove any one of the three inputs and what is left is not a smaller transition. It is not a transition at all.
A transition needs a starting state, an ending state and the time between them, and dropping any one of the three leaves a quantity that means nothing.
The transition, for a process whose states can be listed
$$ p(x,\,s;\; y,\,t) \;=\; \Pr\bigl(X_t = y \,\mid\, X_s = x\bigr), \qquad s \le t $$
\(X_t\)a general process at time \(t\), used here because the statement is about any process rather than about one in particular
\(x\)the starting state, the value the process is observed to hold at the earlier time
\(y\)the ending state, the value being asked about at the later time
\(s,\,t\)the earlier and later times, so that \(t-s\) is the gap the transition spans
\(p\)the transition rule itself, a function of all four arguments at once
What it says in wordsThe transition is the chance that the process is sitting at the ending state at the later time, given that it was observed at the starting state at the earlier time, and it is a function of the starting state, the ending state and both times together rather than of any one of them alone.

Read as a function of both states at once, the transition rule is often called a kernelAnother name for the transition rule, used when it is being thought of as a function of the starting and ending states together rather than as a list of separate numbers.. The name is worth knowing because it signals the right mental picture: not a list of separate probabilities, but one object with two state arguments, in the same way a matrix is one object with a row argument and a column argument.

One condition comes free with the definition and is worth stating because it is the check people actually apply. Whatever the starting state, the process has to be somewhere at the later time. So the total over every possible ending state is one.

The condition that comes with the definition
$$ \sum_{y} p(x,\,s;\; y,\,t) \;=\; 1 \quad\text{for every } x, \qquad\qquad \int_{-\infty}^{\infty} p(x,\,s;\; y,\,t)\, dy \;=\; 1 $$
\(\sum_{y}\)the total over every ending state, used when the states can be listed
\(\int dy\)the same total over a continuous range of ending values, written as an integral
\(x\)the starting state, held fixed while the total is taken
What it says in wordsHolding the starting state fixed and totalling across every ending state gives one, because the process has to be somewhere at the later time, and the same statement for a continuous range is that the area under the transition curve is one.

The total over every ending state is the row sumThe total over all ending states for one fixed starting state, which must come to one., and it is the first thing anybody checks on a table of transitions. An enormous number of tables describing quite different processes all satisfy it, so the row sum is almost useless as a check on whether a table is right.

Try it out

How many things does a transition probability attach to?

Why does a continuous process need a density instead of a number?

The lift has a listed set of states, so its transitions are numbers and they can be put in a table. The standard process is not like that. The standard process, an invented traded quantity, starts at Rs 100/-, drifts at 8 per cent a year with a volatility of 20 per cent a year, and can finish at Rs 115/-, or at Rs 115.0001/-, or anywhere in between. There is no list.

On a continuous range of values every single ending value carries probability zero, so the transition cannot be a probability and has to be a rate instead. This is not a technicality being tidied away. Writing down the chance of finishing at exactly Rs 115.000000/- and adding decimal places to the requirement shrinks the number being described toward nothing, and in the limit it is nothing. A quantity that is zero for every ending value cannot be the content of a model.

The everyday version runs as follows. A room thermometer reads 24 degrees. The chance that the room is at exactly 24 degrees, to infinite precision, is zero: there is always another decimal place that fails. The chance the room is between 23.5 and 24.5 degrees is a real number that can be acted on. Ranges carry probability; points do not. So the object the model supplies is the transition densityThe continuous version of a transition, written as a rate per unit of value rather than as a probability, so that it must be integrated over a range before it means anything., a rate per rupee of ending value. A probability comes out of the density only by integrating over a stretch.

The transition when the ending values form a continuous range
$$ \Pr\bigl(X_t \in [a,\,b] \,\mid\, X_s = x\bigr) \;=\; \int_{a}^{b} p(x,\,s;\; y,\,t)\, dy $$
\(p(\cdot)\)the transition density, a rate per unit of ending value rather than a probability
\([a,\,b]\)a stretch of ending values, which is the smallest thing that carries a probability
\(y\)the ending value being integrated over, which is why it disappears from the answer
\(x\)the starting state, which stays fixed throughout the integral
What it says in wordsThe chance of finishing anywhere inside a stretch of values, given the starting state, is the area under the transition density across that stretch, and no smaller object than a stretch has a probability attached to it at all.

Two consequences follow immediately and both trip people up. First, a density is measured per unit of value, so its height is not a probability and can be bigger than one. Change the unit and the height changes. Second, comparing two densities at a single point reveals very little on its own; what carries meaning is the area over the stretch in question.

Try it out

Why is a transition for a continuous process written as a density rather than as a probability?

What does the transition from the month six reading actually look like?

Take the locked path, the twelve step path published for this reading order, and stop at its month six reading of Rs 111.08/-. To six decimal places the logarithm of that reading is 4.710226. Ask the transition question: given that reading, and given a gap of half a year to the horizon, what is the rule?

The standard process is normally distributed on the logarithm, so the answer is written there. The logarithm gains the drift less half the variance rate, multiplied by the gap, and it picks up variance equal to the variance rate multiplied by the gap. With a drift of 8 per cent and a variance rate of 0.04, the drift less half the variance rate is 0.06 exactly, and over half a year that is a gain of 0.030000. The variance is 0.04 times half a year, or 0.020000. The spread of the logarithm is therefore 0.141421.

The transition of the standard process, written out
$$ \ln S_T \,\bigl|\, S_t = x \;\;\sim\;\; \mathcal{N}\!\left(\ln x + \Bigl(\mu - \tfrac{1}{2}\sigma^{2}\Bigr)(T-t),\;\; \sigma^{2}\,(T-t)\right) $$
\(S_T\)the standard process at the later time, the invented traded quantity used throughout
\(x\)the starting state, here the month six reading of Rs 111.08/-
\(\mu\)the drift, 8 per cent a year, decimal 0.08, an invented parameter
\(\sigma\)the volatility, 20 per cent a year, decimal 0.20, so \(\sigma^{2}\) is 0.04
\(T-t\)the gap, here half a year, decimal 0.5
\(\mathcal{N}\)the normal distribution, written with its average first and its variance second
What it says in wordsGiven the starting reading, the logarithm of the process at the later time is normally distributed with an average equal to the logarithm of the starting reading plus the drift less half the variance rate times the gap, and with a variance equal to the variance rate times the gap.

Putting the month six reading through that rule gives a conditional average for the logarithm of 4.740226 and a spread of 0.141421, and every other figure in this guide is a consequence of those two numbers. The conditional average of the process itself is Rs 115.61/-, from 115.610377. The middle of the distribution, the value with half the outcomes above and half below, is Rs 114.46/-. The peak of the density, the value where the curve is highest, is Rs 112.19/-, from 112.193574.

Three different numbers, from one rule. The three numbers differ because the shape leans to the right: the logarithm is symmetric, but exponentiating a symmetric shape stretches the upper side and compresses the lower one. The distance from the peak to the average is Rs 3.42/-, and that gap is not a rounding artefact. The distance between peak and average is what leaning looks like, and the distance widens with both the volatility and the length of the gap.

The transition from Rs 111.08/- over half a year. A curve, not a number. Rs 100/- Rs 120/- Rs 140/- Rs 160/- peak Rs 112.19/- conditional average Rs 115.61/- the shaded strip is Rs 110/- to Rs 120/-, an area of 0.241555, and an area is the smallest thing here that is a probability The peak sits Rs 3.42/- below the conditional average because the shape leans right. Educational illustration.
The transition from Rs 111.08/- over half a year is a curve over finishing values peaking at Rs 112.19/-, with the conditional average of Rs 115.61/- sitting to the right of the peak.

To turn the curve into numbers a reader can check, cut the finishing values into stretches and take the area over each. The table below does exactly that from the same rule, and the areas total one because they have to.

Finishing stretchChance from the transitionRead as
Below Rs 100/-0.169792the process ends below where it started the year
Rs 100/- to Rs 110/-0.219547a fall from the month six reading, but a small one
Rs 110/- to Rs 120/-0.241555the stretch containing the peak of the curve
Rs 120/- to Rs 130/-0.185102a clear gain over the half year
Above Rs 130/-0.184005the whole upper tail, unbounded above
Every stretch together1.000000the process has to finish somewhere
The same curve, cut into stretches. Now every bar is a probability. 0.169792 0.219547 0.241555 0.185102 0.184005 below 100 100 to 110 110 to 120 120 to 130 above 130 the five bars total 1.000000, which is the row sum condition made visible
Cutting the transition density into five stretches turns a rate into five probabilities that total one, which is the row sum condition seen as areas.

Notice the last two bars. A ten rupee stretch just above the peak carries 0.185102. The entire unbounded upper tail carries 0.184005, almost the same number. The leaning shape is doing that: the upper side is long but thin, and length does not automatically mean weight.

One more property of this particular rule deserves a name. The property is what makes the arithmetic below so clean. The transition depends on the gap and not on when the gap starts. A transition from Rs 111.08/- over half a year is the same rule whether that half year runs from month six to month twelve or from month two to month eight. A rule with that property is time-homogeneousThe case where the transition depends only on how long the gap is, not on when in the calendar the gap begins., and it collapses the four arguments of the general definition to three.

Try it out

The transition from Rs 111.08/- peaks at Rs 112.19/- and averages Rs 115.61/-. Why are those different numbers?

How do transitions over two steps relate to transitions over one?

Here the object stops being a definition and starts being a constraint. Suppose the transition over three months is known and the transition over six is wanted. Getting from the start to the finish over six months means passing through some value at three months, and the only question is which value. So the six month transition is not open to invention. It is already determined.

So every possible intermediate value is taken, weighted by the chance of reaching it, and the second transition applied from each one. Summing over intermediate values for a listed set of states, or integrating over them for a continuous range, gives the six month transition. The operation just described is the compositionChaining two transitions to obtain the transition over the two gaps combined, by totalling over every possible intermediate state., and composition is the single most important thing a transition rule does.

Composing two transitions into one, the statement Chapman and Kolmogorov carry
$$ p(x,\,t;\; y,\,T) \;=\; \int_{-\infty}^{\infty} p(x,\,t;\; z,\,u)\; p(z,\,u;\; y,\,T)\, dz, \qquad t \le u \le T $$
\(x\)the starting state at the earliest time
\(z\)the intermediate state at the middle time, which is totalled over and therefore disappears
\(y\)the ending state at the latest time
\(u\)the middle time, any time at all between the two ends
\(\int dz\)the total over every intermediate value the process could have held
What it says in wordsThe transition over the whole gap equals the transition to any intermediate value multiplied by the transition on from that value, totalled over every intermediate value the process could have held, and the answer does not depend on which intermediate time is chosen.

The last clause of that sentence carries the whole content. The middle time is arbitrary. The gap can be cut anywhere, and the composed answer has to come out the same. If it does not, the rule as written is not describing one process.

Now do it on the numbers, and watch what adds and what does not. Go from month six to month nine, a quarter of a year, and then from month nine to month twelve, another quarter. On the logarithm each quarter contributes an average gain of 0.06 times a quarter, or 0.015000, and a variance of 0.04 times a quarter, or 0.010000. Two quarters give an average gain of 0.030000 and a variance of 0.020000. The direct half year transition gives an average gain of 0.06 times half a year, or 0.030000, and a variance of 0.04 times half a year, or 0.020000. Identical.

The variances add. The spreads do not. Each quarter has a spread of 0.100000, and the half year spread is 0.141421, not 0.200000. Adding spreads is the single most common arithmetic error in this whole subject, and the reason it is wrong is that spread is the square root of variance and square roots do not add. The everyday version is weighing: weighing a sack in one go leaves a measurement error with some spread, and weighing it in two halves and adding the readings combines the two errors in variance, so the combined spread is smaller than the sum of the two spreads.

Composing two steps. One of these two ledgers is arithmetic and the other is a habit. THE RIGHT LEDGER: VARIANCES ADD month 6 to month 9 0.010000 + month 9 to month 12 0.010000 = half a year direct variance 0.020000 so the spread is 0.141421 THE WRONG LEDGER: SPREADS DO NOT ADD spread of quarter one 0.100000 + spread of quarter two 0.100000 = 0.200000, which is the spread of nothing too wide by 0.058579 Spread is the square root of variance, and square roots do not add. Educational illustration on invented parameters.
Two quarter year transitions each with a logarithm variance of 0.01 compose to 0.02 over half a year, which is exactly what the direct half year transition gives.

The wording of the composition invites one misreading in particular. The month nine reading on the locked path was Rs 93.74/-, and composing through month nine does not mean the process passes through that reading. Composing means totalling over every value month nine could have held, weighted by the chance of holding it. The realised path is one draw; the composition is over all of them.

Try it out

Two quarter year steps, each with a logarithm variance of 0.01. What is the variance over the half year?

What has to hold before a set of transitions is one process?

A table of transitions over one month is a table. Another table of transitions over three months is another table. Neither is a process. A set of tables becomes a process when composing the short ones reproduces the long ones, for every way of cutting every gap. The requirement is the consistency conditionThe requirement that composing short transitions reproduces the longer transition, for every way of cutting the gap., and consistency is the only thing standing between a set of plausible numbers and a model.

Consistency is not a nicety added on top of the transition rule; it is the property that makes the word process apply at all. A rule that fails it gives one answer for the same question depending on the route taken to it, and a quantity with two values is not a quantity.

For the standard process the condition holds by construction, and it is worth seeing why in the simplest possible terms. Cross the half year in one step: the logarithm gains 0.06 times 0.5, or 0.030000, and picks up variance 0.04 times 0.5, or 0.020000. Cross it in three steps: each step gains 0.06 times one sixth, or 0.010000, and picks up 0.04 times one sixth, or 0.006667. Three of each gives 0.030000 and 0.020000. Cross it in twelve steps and each contributes 0.002500 and 0.001667, and twelve of each gives 0.030000 and 0.020000 again.

Why the number of steps cannot change the answer
$$ \sum_{k=1}^{n}\Bigl(\mu - \tfrac{1}{2}\sigma^{2}\Bigr)\frac{\Delta}{n} \;=\; \Bigl(\mu - \tfrac{1}{2}\sigma^{2}\Bigr)\Delta, \qquad\qquad \sum_{k=1}^{n} \sigma^{2}\,\frac{\Delta}{n} \;=\; \sigma^{2}\Delta $$
\(n\)the number of equal steps the gap is cut into, any whole number at all
\(\Delta\)the whole gap, here half a year, decimal 0.5
\(\Delta/n\)the length of one step, which shrinks exactly as fast as the count of steps grows
\(\mu - \tfrac{1}{2}\sigma^{2}\)the drift of the logarithm, 0.06 exactly for the standard process
\(\sigma^{2}\)the variance rate of the logarithm, 0.04 exactly
What it says in wordsCutting the gap into more steps makes each step shorter by exactly the factor by which the number of steps grew, so the average gains total to the same number and the variances total to the same number, whatever the number of steps.

The invariance is now stated, and the calculator below shows it holding. A requirement that has only been read is easy to believe and easy to forget; a curve that refuses to move while a control is dragged is hard to forget.

Try it out

The same half year gap is about to be crossed in twelve steps instead of two. Does the composed curve change?

Play with it

Cut the same half year into more steps and watch what refuses to move

Starting at Rs 111.08/-, with half a year to the horizon and independent increments throughout. The pale band is the direct half year transition, drawn once. The dark line is the transition built by composing the steps one at a time. At every setting the composed average gain is 0.030000 and the composed variance is 0.020000. The number of steps cannot change either. What does move is the bar at the bottom, which adds the step spreads instead of the step variances.

THE GAP, CUT INTO STEPS month 6, Rs 111.08/- month 12, the horizon THE TRANSITION OVER FINISHING VALUES Rs 95/- Rs 105/- Rs 115/- Rs 125/- Rs 135/- Rs 145/- ADDING THE STEP SPREADS INSTEAD, WHICH IS THE THING THAT MOVES the true spread, 0.141421, and it never moves
Steps across the gap
2
Length of one step
0.250000
Composed average gain
0.030000
Composed variance
0.020000
Composed spread
0.141421
Step spreads added up
0.200000
Crossing the half year in 2 steps of 0.250000 years each, the composed transition has an average gain of 0.030000 and a variance of 0.020000, so its spread is 0.141421 and its curve lies exactly on the direct half year curve. Adding the 2 step spreads instead would give 0.200000, which is the spread of nothing.
Educational illustration. Every reading is computed from the transition rule by composing the steps one at a time, never sampled, so the default of two steps reproduces the worked example exactly on every reload. The composed curve lands on the direct curve at all twelve settings, and that is the finding rather than a rounding.
Two routes, one destination. The consistency condition is that they agree. MONTH 6 Rs 111.08/- ROUTE ONE: straight across, half a year average gain 0.030000, variance 0.020000 MONTH 9 every value it could hold, weighted by its chance 0.015000 and 0.010000 0.015000 and 0.010000 ROUTE TWO: by way of the middle BOTH ROUTES ARRIVE HERE gain 0.030000 variance 0.020000
Going from month six to twelve directly and going by way of month nine both give a logarithm variance of 0.02 and an average gain of 0.03, and they have to.
Derivatives Foundation Bootcamp — Fin Maverick

Why is the transition rule the whole content of a memoryless model?

One claim makes the transition rule worth a treatment of its own. Writing down where the process starts, together with the transition rule for every pair of states and every gap, finishes the specification. There is no further ingredient. Everything else anybody might ask of the model is derived from those two things rather than specified alongside them.

Once the transition rule is written down for every pair of states and every gap, nothing further about the model remains to be specified. The average at any horizon comes from integrating the ending value against the transition. The spread comes from integrating the squared deviation. The chance of any stretch of values comes from integrating over that stretch. The average of any function of the future value at all comes from integrating that function against the transition. The rule is not a summary of the model; the rule is the model.

The specification sheet for a memoryless model. It is two lines long. SPECIFIED BY HAND 1. where the process starts 2. the transition rule, for every pair of states and every gap NOTHING FURTHER the sheet ends at line two DERIVED, NEVER SPECIFIED the average at any horizon the spread at any horizon the chance of any stretch of values the transition over any longer gap the average of any function of the future value at all every one of these is an integral against line two A model that needs a third specified line has something the transition rule failed to carry, which means it is not memoryless.
Once the transition rule is written out for every pair of states and every gap, nothing further about the model remains to be specified.

The converse is the diagnostic. If a third thing has to be specified, something the transition rule could not carry, then the present state was not a sufficient summary after all and the model is not the kind of model it was taken to be. Needing to know how the process arrived at its current state, or how long it has been there, is exactly that situation, and it is a signal to widen what counts as the state rather than to bolt an extra rule onto the side.

Try it out

A memoryless model is being written down and the transition rule is complete. What is left to specify?

Where does the transition rule turn up once a model is being solved?

Everywhere, and usually without being named. Three appearances are worth recognising. A reader who has met the transition rule only as a definition will otherwise not notice it wearing working clothes.

  1. As the thing a future quantity is integrated against Any quantity that depends on the future value of the process is turned into a number today by integrating it against the transition density and discounting. That is what an expectation under a stated measure is, mechanically. The contract is treated purely as a function of the finishing value; nothing about what it pays is at issue here.
    On the standard process over one year under the risk-neutral measure, integrating the at-the-money contract against the transition density and discounting at 0.951229 returns Rs 10.45/-, from 10.450584, which is the closed form figure to six decimals.
  2. As the unknown in a partial differential equation Read as a function of the ending state and the later time, the transition density satisfies one equation; read as a function of the starting state and the earlier time, it satisfies another. Kolmogorov carries both, forward and backward. Solving either equation is solving for the transition rule, whether or not the person solving it says so.
    Feynman and Kac carry the statement that links the two views: the solution of the equation is the expectation, and the expectation is the integral against the transition density.
  3. As the weights on a lattice or a grid Any numerical scheme that steps a model forward carries a transition rule inside it, in the form of the weights attached to the moves out of each node. If those weights do not compose to the transition over the longer gap, the scheme is converging to the wrong thing, and no amount of refinement rescues it.
    The transition rule is what those weights are. How a scheme is built, refined and checked is set out under numerical pricing.

The recognisable pattern across all three is that the transition rule is the object being solved for, and the rest is machinery for solving it. A reader who holds that picture will find that a great deal of later material stops looking like unrelated techniques and starts looking like several routes to the same quantity.

Try it out

A quantity depending on the finishing value of the process is being turned into a number today. What is it integrated against?

Bond Pricing and Yield Mechanics — free micro-course from Fin Maverick

What goes wrong when two tables are built without checking that they compose?

The error that gets made, and what it costs

A set of transitions is built by hand, one gap at a time. One table gives the transitions over a month. A second table, built later and separately, gives them over three months. Each table is individually reasonable. Every row of each table totals to one. Each was reviewed and each passed.

Unless the one month table composed three times reproduces the three month table exactly, the two tables describe two different processes. Every check that looks at one table alone is satisfied, so nothing inside either table reveals the failure. The disagreement surfaces later, when somebody computes the same horizon two ways for an unrelated reason and gets two answers.

Numbers make the point. The size of the disagreement is the surprising part. Consider three states, named low, middle and high, and a monthly table where each state mostly stays put. Composing that table with itself three times gives the three month transitions the monthly table implies. Placed beside a separately written three month table with sensible round numbers, every row of both tables totals to one.

Both tables pass every check applied to either one alone. THE MONTHLY TABLE, COMPOSED THREE TIMES TO LOW TO MID TO HIGH LOW 0.431 0.413 0.156 MID 0.248 0.504 0.248 HIGH 0.156 0.413 0.431 vs THE SEPARATE THREE MONTH TABLE TO LOW TO MID TO HIGH LOW 0.500 0.380 0.120 MID 0.240 0.520 0.240 HIGH 0.120 0.380 0.500 Every row of both tables totals to 1.000, and the middle rows agree to within 0.016. The corner cells are 0.069 apart, which is the two tables describing two different processes. Invented illustration.
A one month table and a three month table each with rows summing to one can still disagree, and only composing the first three times exposes it.

The middle row is the part that keeps the error alive. The middle row agrees to within 0.016, close enough that a reviewer scanning the two tables sees a match and moves on. The corners are 0.069 apart, and over a long horizon that difference compounds into a materially different answer. And every row of both tables totals to one, so the check most people actually run has told them nothing at all.

Try it out

A one month table and a three month table are both supplied. What is the one check?

Two tables, each row totalling to one, and neither composes. See which transition breaks.

How does somebody reviewing a model rather than building one use this?

A reviewer does not usually get to see how a model was built. A set of outputs and a description arrive, and the question is whether the description holds together. The transition rule gives that reviewer a check that needs no access to anyone else's code and no agreement about what the right answer is.

The check is to compute one quantity two ways, over the same total gap, cut differently, and see whether the two agree. If a model produces a distribution over a year in one step and also in twelve monthly steps, those two have to match to within the tolerance the scheme claims. Where they do not, either the transition rule is inconsistent or the scheme is not solving the rule it says it is, and both are worth finding before anything is built on top.

The same arithmetic is why a spread computed over one period gets scaled by the square root of the number of periods rather than by the number of periods. Square root scaling is nothing more than the composition rule applied to variances, and it is correct exactly when the increments are independent and their variances are equal. Independent increments of equal variance are precisely the assumption on screen in the simulation above. When somebody scales a spread and the increments are not independent, the composition arithmetic has been applied where its condition fails, and the resulting number is neither the one period spread nor the multi period one.

The third use is the plainest. When a description of a model says the state carries everything relevant, ask what the transition rule is, in full, for every pair and every gap. If the answer needs a further ingredient, the description was wrong about the state, and it is better to widen the state deliberately than to discover the gap through a disagreement nobody can locate.

The property that makes the present state a sufficient summary in the first place is set out under the Markov property, which holds why the history can be dropped. This guide holds how what remains is written down. Solving a model on a grid or a lattice is set out under numerical pricing. Estimating transitions from observed data is covered separately. What any contract pays is a separate subject. No jurisdiction sets the definition of a transition probability: the statement is mathematical and holds everywhere, so no rule, rate or threshold here would need confirming at a regulator.
Breaking Into Quants Bootcamp — Fin Maverick

References

SourceDocumentWhere
arXiv Quantitative FinancePreprint repository for transition densities and Markov methods in pricingarxiv.org
Social Science Research NetworkWorking paper repository for the same materialssrn.com
Chapman and KolmogorovThe composition statement for Markov transitions that carries their namesstandard probability texts
Feynman and KacThe link between the partial differential equation and the expectation that carries their namesstandard probability texts
Hull, Shreve and WilmottStandard texts covering transition densities and the pricing equations built on themtextbook publishers

The standard process, its four parameters, its locked path and the three state table in the failure block are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.