Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Stochastic Calculus & Derivative Pricing Theory
1Probability Foundations
The Probability SpaceRandom VectorsSigma-AlgebraExpectationSample Space and EventsDensity and Distribution FunctionsRisk-Neutral ProbabilityState Price Density vs…
2Stochastic Processes and Jumps
Properties of a Stochastic ProcessMartingaleBrownian Motion and Its PropertiesBrownian Motion vs Geometric…Stopping TimeThe Markov PropertyState VariablesTransition ProbabilityQuadratic VariationQuadratic Variation vs Ordinary…Submartingale and SupermartingaleMartingale RepresentationMarkov Process vs MartingaleOptional StoppingFiltrationJump ProcessesThe Poisson ProcessLevy ProcessesJump Diffusion
3Ito Calculus
The Ito IntegralThe Ito Integral vs the Riemann IntegralInfinitesimals in Stochastic CalculusQuadratic CovariationIto's LemmaHow to Apply Ito's…The Infinitesimal GeneratorIto Calculus vs Ordinary Calculus
4Stochastic Differential Equations
Stochastic Differential EquationsStochastic Differential Equation vs…Drift and DiffusionStrong and Weak Solutions ComparedDiscretisationGeometric Brownian Motion
5Pricing Theory and No-Arbitrage
No-ArbitrageGirsanov, Radon-Nikodym and Change…Physical and Risk-Neutral Measures…The Fundamental Theorems of…The Law of One PriceThe Pricing KernelDiscount Factors and Zero-Coupon PricesReplication vs HedgingComplete Market vs Incomplete MarketClearing Margin Architecture
6Option Pricing Theory
European and American OptionsMonte Carlo European OptionThe Black-Scholes PDEBlack Scholes and the GreeksThe Payoff FunctionThe Binomial ModelBinomial Option PricingDelta Hedging in TheoryBoundary, Initial and Terminal ConditionsThe Exercise BoundaryHow to Check Put-Call…
7Volatility Models
Constant, Local and Stochastic…Vasicek Model vs CIR ModelThe Heston ModelThe SABR ModelThe Volatility ProcessImplied VolatilityVolatility Smile vs Skew vs Surface
8Interest Rate Models
Interest-Rate DerivativesMean ReversionThe Zero-Coupon BondThe Ornstein-Uhlenbeck ProcessThe Discount CurveZero RatesShort-Rate Model vs Market Model
9Numerical Pricing
Closed Form and Numerical…Monte Carlo PricingEuler and Milstein Schemes ComparedTree MethodsFinite Difference MethodsNumerical Error and StabilityVariance Reduction
10Calibration and Model Risk
Model OverrideMarket Price and Model PriceCalibrationHow to Document a Pricing ModelThe Educational Illustration LabelMarket ConventionsModel Uncertainty and LimitationsBacktesting a Pricing ModelIdentifiabilityCalibrated ParametersThe Calibration Loss Function

Expectation: The Probability-Weighted Average

Expectation is the average of a random quantity weighted by the measure, computed as an integral against that measure rather than as a sum over observations. The average is not a value to expect: for the standard process the average finish is reached or beaten in fewer than half of all outcomes. Conditional expectation is the same average taken with information in hand, and it is itself random.

An expectation is built from the measure, not from a sample. Building the average from the measure is why it exists before any observation is made, why it changes the instant the measure changes, and why two people looking at the same quantity under two different measures compute two different averages without either of them making an arithmetical error.

The worked instance throughout is the standard process, written S with a time subscript, an invented traded quantity that starts at Rs 100/-, drifts at 8 per cent a year, carries a volatility of 20 per cent a year and is watched over one year. Every figure below is a computed consequence of those four numbers.

Why is expectation defined as an integral against the measure?

Consider what an average usually means. Given a list of readings, the readings are added up and divided by how many there were. A sum over observations is a perfectly good number, but note what it needs: it needs the observations to exist. Before a single reading is taken it has nothing to work with.

The expectationThe average of a quantity weighted by the measure, taken over the whole outcome set rather than over a list of readings. of this subject is a different object. The expectation does not wait for data. The definition takes every outcome the model admits, reads the value the quantity takes at that outcome, weights it by how much probability the measure puts there, and adds the pieces up. The averaging is done against the measure, so the answer is fixed the moment the measure is fixed, and no observation is needed to produce it.

There is a difference between what a weighing scale reads and what the object actually weighs. The scale gives one reading, then another, and the readings can be averaged. The weight is a property of the object and is there whether or not anybody weighs it. An expectation is the second kind of thing. The expectation is a property of the model, not a reading taken from it.

Expectation, the definition
$$ \mathbb{E}^{P}[X] \;=\; \int_{\Omega} X(\omega)\, dP(\omega) $$
\(X\)a random variable, so a rule assigning a number to each outcome
\(\Omega\)the sample space, the set of every outcome the model admits
\(\omega\)one outcome, one element of that set
\(P\)the physical measure, which assigns the weight to each event
What it says in wordsThe expectation of a random variable is the total of its value at each outcome multiplied by the weight the measure puts on that outcome, added up across the whole outcome set, which is why it is written as an integral rather than as a division by a count.

The integral sign is doing real work here and is not decoration. On a finite outcome set the definition collapses to the schoolroom weighted average, value times probability, summed. The standard process lives on an infinite outcome set. There is no list to sum there and no count to divide by, so the summing has to be done by an integral against a measureThe sum of value times weight, built so that it still makes sense when there are infinitely many outcomes and no list to add up.. Same idea, machinery that survives the infinite case.

When the quantity is continuous and its distribution has a density, the same integral is usually written over values rather than over outcomes. The second form is the one used in actual computation.

The same expectation, written over values
$$ \mathbb{E}^{P}\bigl[g(S_T)\bigr] \;=\; \int_{0}^{\infty} g(x)\, f(x)\, dx $$
\(S_T\)the standard process at the horizon, in rupees
\(g\)any function of that value, including the identity function
\(f\)the density of the finishing value under the measure \(P\)
\(x\)a possible finishing value, the variable being integrated over
What it says in wordsThe average of any function of the finishing value equals that function evaluated at each possible value, multiplied by the density there, integrated over all values, which is the same weighted sum with the outcome set replaced by the values it produces.
Two averages. Only one of them exists before anything is observed. A SUM OVER OBSERVATIONS 98.4 104.1 111.7 not yet add the readings, divide by the count needs data before it can produce anything changes when a new reading arrives carries the accidents of the readings taken AN INTEGRAL AGAINST THE MEASURE OUTCOME VALUE WEIGHT every outcomethe valuefrom P exists the moment the measure is fixed moves only when the measure moves The left number is a summary of what happened. The right number is a property of the model, and it is the one this subject means.
An average of observations needs readings before it can produce anything, while an expectation is a weighted total over the whole outcome set and is fixed as soon as the measure is fixed.
Try it out

Two people hold exactly the same record of past readings of the same quantity, and they disagree about its expectation. Which of these is the only explanation that does not involve an arithmetical mistake?

Try it out

The standard process finishes with an average of Rs 108.33/-. Before the next section: what share of all outcomes reach that figure or beat it?

Why is the average not a value that is likely to be observed?

Because the name is a historical accident and the arithmetic never promised it. The expectation is a weighted total. A weighted total sits wherever the weights and values put it, and that place need not be anywhere near the bulk of the distribution.

Ordinary life already shows this. Take one street and add up every household income on it. If one house on that street earns far more than the rest, the average income of the street is a number that almost nobody on it actually earns. The average has not lied. The average has answered a question about the total. The question asked was about the typical case.

The standard process does exactly this, and the amount by which it does it can be written down exactly. Its finishing value is lognormalThe shape of a quantity whose logarithm is normally distributed, which leans to the right because the quantity cannot fall below zero but has no ceiling above.: the logarithm of the finish is normally distributed. The finish itself therefore cannot go below zero and has no ceiling above. The asymmetry pulls the average to the right of the bulk.

The average and the middle outcome of the standard process
$$ \mathbb{E}^{P}[S_T] = S_0\,e^{\mu T} \qquad\qquad \operatorname{med}\bigl(S_T\bigr) = S_0\,e^{\left(\mu - \frac{1}{2}\sigma^{2}\right)T} $$
\(S_0\)the starting value, Rs 100/- for the standard process
\(\mu\)the drift under the measure \(P\), 0.08 a year here
\(\sigma\)the volatility, 0.20 a year, so the variance rate is 0.04
\(T\)the horizon, one year here
What it says in wordsThe average finishing value grows the starting value at the drift, while the middle finishing value grows it at the drift less half the variance rate, so the two are different numbers whenever the volatility is anything other than zero.

Put the four locked numbers in. The average finish is Rs 100/- times e to the 0.08, or Rs 108.33/-. The medianThe value a quantity falls below half the time and rises above half the time, so the outcome sitting exactly in the middle of the ordering., the middle outcome, is Rs 100/- times e to the 0.06, or Rs 106.18/-. The most likely single value, the peak of the density, is lower again at Rs 102.02/-. Peak, then middle, then average: the ordering is forced by the lean of the shape and not by the particular parameters chosen.

The share of outcomes that reach the average or beat it works out at 0.460172, or 46.0 per cent. Fewer than half. A number described as the average finish is, on this process, a finish that most outcomes fall short of.

The finishing value of the standard process. Three markers, and their order never changes. 60 80 100 120 140 160 180 Rs 46.0 per cent of outcomes sit in the shaded region, at or above the average THE SAME THREE MARKERS, MAGNIFIED NINE TIMES peak Rs 102.02/- middle outcome Rs 106.18/- average Rs 108.33/-
The peak at Rs 102.02/-, the middle outcome at Rs 106.18/- and the average at Rs 108.33/- fall in that order because the shape leans right, and 46.0 per cent of outcomes reach the average or beat it.

The gap is not an accident of these particular numbers. Look at where it comes from. The average grows the starting value at 8 per cent. The middle outcome grows it at 8 per cent less half the variance rateThe square of the volatility, so 0.04 a year when the volatility is 20 per cent a year. Variance accumulates at that rate. of 0.04, which is 6 per cent. The whole distance between the two is half the variance rate. Half the variance rate is a parameter of the model, not a quirk of the arithmetic.

One starting value. Two growth rates. The distance between them is a parameter. THE MIDDLE OUTCOME 6.00% grows Rs 100/- to Rs 106.18/- THE AVERAGE 8.00% grows Rs 100/- to Rs 108.33/- 2.00 points and 2.00 points is exactly half of 0.04 volatility 20 per cent a year, so the variance rate is 0.20 times 0.20, which is 0.04 half of 0.04 is 0.02, and 0.02 is the whole distance between 6 per cent and 8 per cent Educational illustration. Both rates are continuously compounded over the one year horizon of the standard process. Why the correction is half the variance rate is settled elsewhere in this subject. Here it is used, not derived.
The middle outcome grows at 6 per cent and the average at 8 per cent, and the 2.00 percentage points between them are exactly half the variance rate of 0.04.
Try it out

The volatility is about to rise from 20 per cent to 40 per cent, with the drift left at 8 per cent. Before the control below is moved: what happens to the average finishing value?

Play with it

Move the volatility. Watch which marker refuses to move.

The drift stays at 8 per cent, the horizon stays at one year and the starting value stays at Rs 100/-. Only the volatility changes. At the default of 20 per cent the average sits at Rs 108.33/- and the middle outcome at Rs 106.18/-, the two figures worked above. Half of 0.16 is exactly the drift of 0.08. Push the control all the way to 40 per cent and the middle outcome lands on exactly Rs 100.00/-.

0 per cent20 per cent a year40 per cent
The distribution of the finishing value, redrawn from the formula at every setting. 60 80 100 120 140 160 180 200 220 Rs the average, nailed to Rs 108.33/- middle outcome Rs 106.18/- 46.0 per cent at or above the average THE GAP, DRAWN TO ITS OWN SCALE Rs 2.15/- The bar runs from zero to Rs 8.33/- across its full width, so the gap stays readable at every setting. The curve height is rescaled each redraw.
The average
Rs 108.33/-
The middle outcome
Rs 106.18/-
The gap
Rs 2.15/-
Reaching the average
46.0%
At a volatility of 20 per cent a year the average finish is Rs 108.33/- and the middle outcome is Rs 106.18/-, a gap of Rs 2.15/-, and 46.0 per cent of outcomes reach the average or beat it.
Educational illustration. Every reading here is computed from the closed form for the standard process rather than drawn at random, so the same setting always returns the same figures. The drift is held at 8 per cent a year, the horizon at one year and the starting value at Rs 100/-. The average is Rs 100/- times e to the 0.08 at every volatility, which is why its marker never moves. The middle outcome is Rs 100/- times e to the drift less half the volatility squared, which is Rs 107.79/- at 10 per cent, Rs 106.18/- at 20 per cent, Rs 103.56/- at 30 per cent and exactly Rs 100.00/- at 40 per cent.
Derivatives Foundation Bootcamp — Fin Maverick

What is conditional expectation, and why is it not simply a number?

An unconditional expectation averages over every outcome the model admits. A conditional expectationThe average taken over only the outcomes still possible given what is known at a point in time, rather than over the whole outcome set. averages over only the outcomes that are still possible given what is known. The information is supplied by the collection of sets whose questions can already be answered, written as a script F with a time subscript, and the averaging is otherwise identical.

Take a lift in a tall building. The lift only knows which floor it is on. Where it will be in a minute depends entirely on the floor it is standing on now: from the second floor the average destination is low, from the twentieth it is high. There is no single answer to the question until the floor is named. The conditional average is not one number, it is one number for each thing that might be learned.

On the standard process the finishing value equals the value now multiplied by a factor that has nothing to do with how the process arrived there. The conditioning is therefore unusually clean. So the conditional average of the finish is the current value grown at the drift for the time that remains.

Conditional expectation of the standard process
$$ \mathbb{E}^{P}\!\left[\,S_T \;\middle|\; \mathcal{F}_t\,\right] \;=\; S_t\, e^{\mu (T-t)} $$
\(\mathcal{F}_t\)the information available at time \(t\), a collection of sets
\(S_t\)the value of the standard process at time \(t\), itself random
\(T-t\)the time still to run, half a year in the instance below
\(\mu\)the drift under \(P\), 0.08 a year
What it says in wordsGiven everything known at a point in time, the average finishing value of the standard process is its value at that point grown at the drift for the time still to run, so the conditional average inherits its randomness entirely from the current value.

Now work it. The locked path is the twelve step path published for this subject. Along it the standard process stands at Rs 111.08/- at six months. Grow that at 8 per cent for the remaining half year: Rs 111.08/- times e to the 0.04, or Rs 115.61/-. Rs 115.61/- is the average of the finish given that particular information.

Read the right hand side of the formula again. The right hand side contains S with a time subscript, and S with a time subscript is a random variableA quantity that takes a value at each outcome, so a function of the outcome rather than a fixed number.. A random variable multiplied by a constant is still a random variable. The conditional average of the finish is therefore Rs 115.61/- only on the paths standing at Rs 111.08/- at six months, and a different number on every other path. A quantity taking a different number on every path is exactly what a random variable is.

A conditional average is not one number. It is one number at every place the process could stand. WHERE IT STANDS AT SIX MONTHS THE AVERAGE FINISH FROM THERE Rs 85.00/- Rs 88.47/- Rs 95.00/- Rs 98.88/- Rs 105.00/- Rs 109.29/- Rs 111.08/- Rs 115.61/- Rs 120.00/- Rs 124.90/- each one is the left value times 1.040811 The highlighted row is where the locked path actually stands at six months. The other rows are equally real, which is why the conditional average is a function and not a figure.
The conditional average of the finish takes the value Rs 115.61/- where the process stands at Rs 111.08/- and a different value at every other place it could stand, which is what makes it a random variable.
Try it out

The conditional average of the finish, given where the standard process stands at six months, is Rs 115.61/-. Is that figure a number or a random variable?

How is a conditional probability different from a conditional expectation?

Conditional Probability vs Conditional Expectation

The honest answer is that they are not two ideas. One is the other applied to a particular quantity, and once that quantity is identified, the pair collapses into a single operation with two inputs.

The quantity is the indicatorA quantity that reads one when a stated event happens and zero when it does not, so it converts an event into a number. of the event. The indicator reads one on every outcome where the event happens and zero on every outcome where it does not. Average a quantity that is only ever one or zero and the weight on the ones is precisely the probability of the event, so the average is the probability.

A probability is the average of an indicator
$$ P\bigl(A \mid \mathcal{F}_t\bigr) \;=\; \mathbb{E}^{P}\!\left[\, \mathbf{1}_{A} \;\middle|\; \mathcal{F}_t \,\right] $$
\(A\)an event, a set of outcomes the measure can weigh
\(\mathbf{1}_{A}\)the indicator of that event, one on \(A\) and zero elsewhere
\(\mathcal{F}_t\)the information available at time \(t\)
\(P\)the physical measure, supplying the weights
What it says in wordsThe conditional probability of an event is the conditional average of the quantity that reads one when the event happens and zero when it does not, so probability is not a second idea alongside expectation but the same operation applied to an indicator.

Working both on the same node makes the relation visible. The standard process stands at Rs 111.08/- at six months. Feeding the finishing value itself into the averaging gives Rs 115.61/-, a figure in rupees. Rs 100/- is the strike of the at-the-money contract. Feeding in the indicator of finishing above that strike gives 0.830208, a pure number between zero and one. Same measure, same conditioning, same integral, and the only thing that changed was the function being averaged.

One operation, two inputs. That is the whole difference. AVERAGE THE VALUE ITSELF the function fed in finishing value in rupees reads the value the conditional average comes out as Rs 115.61/- AVERAGE THE INDICATOR the function fed in 1 0 Rs 100/- reads one above the level the conditional average comes out as 0.830208 Same measure, same information, same integral, both taken from the node at Rs 111.08/-. A conditional probability is a conditional average that happened to be fed an indicator.
The conditional probability of finishing above Rs 100/- is the conditional average of a quantity reading one above that level and zero below it, so the two ideas are one operation.

One consequence is worth carrying away. Because a probability is an expectation, every rule that holds for expectations holds for probabilities automatically, and no second set of rules is needed. The tower property below is the clearest instance of that.

Try it out

A conditional probability is wanted. Which quantity is the conditional expectation taken of?

Risk Management Program Bootcamp — Fin Maverick

What does the tower property say, and why does every later argument lean on it?

The tower propertyThe rule that averaging a conditional average over the information gives back the plain average, with no correction term. says something that sounds too simple to be useful. Averaging within each group, then averaging the group averages weighted by how likely each group is, gives the plain average. No correction term appears. Nothing is left over.

The tower property is used constantly without being named. The average height of everyone in a town can be found by averaging each street separately and then averaging the street averages, provided each street is weighted by how many people live on it. Nobody expects a correction term for having done it in two stages, and there is none.

The tower property, in its two forms
$$ \mathbb{E}^{P}\Bigl[\ \mathbb{E}^{P}\bigl[X \mid \mathcal{F}_t\bigr]\ \Bigr] = \mathbb{E}^{P}[X] \qquad\qquad \mathbb{E}^{P}\Bigl[\ \mathbb{E}^{P}\bigl[X \mid \mathcal{F}_s\bigr] \;\Big|\; \mathcal{F}_t \Bigr] = \mathbb{E}^{P}\bigl[X \mid \mathcal{F}_t\bigr] $$
\(X\)any random variable whose average exists
\(\mathcal{F}_t\)the information available at the earlier time \(t\)
\(\mathcal{F}_s\)the information available at the later time \(s\), with \(t\) before \(s\)
\(P\)the physical measure, unchanged throughout
What it says in wordsAveraging a conditional average over the information returns the plain average, and averaging a later conditional average down to an earlier information set returns the earlier conditional average, so the coarser of the two conditionings always wins and no correction term ever appears.

Assertions of this kind are worth checking rather than believing, so here is the check on the standard process. Cut the possible positions at six months into five groups of equal probability, one fifth each. For each group compute the average position within it, grow that average at the drift for the remaining half year to get the conditional average of the finish, then weight the five results by one fifth and add.

Group of positions at six monthsWeightAverage position in the groupAverage finish from there
up to Rs 91.48/-0.20Rs 84.72/-Rs 88.18/-
Rs 91.48/- to Rs 99.42/-0.20Rs 95.61/-Rs 99.51/-
Rs 99.42/- to Rs 106.80/-0.20Rs 103.07/-Rs 107.27/-
Rs 106.80/- to Rs 116.07/-0.20Rs 111.13/-Rs 115.66/-
above Rs 116.07/-0.20Rs 125.89/-Rs 131.02/-
weighted total of the five group averages1.00Rs 104.08/-Rs 108.33/-

The bottom right cell is Rs 108.33/-, the unconditional average computed in one step earlier. Two routes, one answer. The agreement holds to six decimal places, not merely to the two shown. Notice also that the locked path node of Rs 111.08/- sits inside the fourth group, whose average position is Rs 111.13/-, which is why that group carries a conditional finish so close to the Rs 115.61/- computed earlier.

Average within each group, then average the group averages. Nothing is lost in the two step route. STEP ONE Rs 88.18/- Rs 99.51/- Rs 107.27/- Rs 115.66/- Rs 131.02/- one conditional average for each group WEIGHT times 0.20 times 0.20 times 0.20 times 0.20 times 0.20 STEP TWO add them up Rs 108.33/- the one step average Rs 108.33/- Rs 100/- times e to the 0.08, computed without any grouping The two routes agree exactly, and there is no correction term anywhere in the collapse.
Averaging the finish within each group of six month positions and then averaging those group averages returns Rs 108.33/-, which is the average taken in a single step.

Why does this small rule carry so much weight later? Because almost every argument in this subject moves a quantity from one time to another, and the tower property is what allows that to be done in stages without accumulating error. The tower property appears again as the engine behind stepping backwards through a lattice, behind the whole treatment of processes whose average never moves, and behind the link between an equation and an average. The rule is used far more often than it is stated.

Try it out

A conditional average is taken over the information at six months, and that result is then averaged over everything that could happen by six months. What does that give?

Building a Revenue Forecast From Drivers — free micro-course from Fin Maverick

What happens to an expectation when the measure changes?

The average moves. The outcome set does not. Separating the two is the whole reason a change of measure is a usable instrument rather than a rewrite of the model, and the separation follows directly from the definition: the outcome set and the function live on one side of the integral, the weights on the other, and only the weights carry the measure.

Two measures run through this subject. P is the physical measure, under which the standard process drifts at 8 per cent. Q is the risk-neutral measure, under which the same process drifts at the risk-free rate of 5 per cent. Why Q exists and how it is built are covered separately; the only feature of Q that matters below is the rate it carries.

The same finishing value, averaged under the second measure
$$ \mathbb{E}^{Q}\bigl[S_T\bigr] \;=\; S_0\, e^{r T} $$
\(Q\)the risk-neutral measure, the second measure carried by this subject
\(r\)the risk-free rate, 0.05 a year, continuously compounded
\(S_0\)the starting value, Rs 100/-, identical under both measures
\(T\)the horizon, one year
What it says in wordsUnder the second measure the average finishing value grows the starting value at the risk-free rate instead of at the drift, so the same quantity on the same outcome set has a different average purely because the weights changed.

Put the numbers in and the pair is Rs 108.33/- under P against Rs 105.13/- under Q, a difference of Rs 3.20/-. Not one path was added, removed or altered between those two figures. Every finishing value that was possible before is still possible, and every finishing value that was impossible before is still impossible. Only the weights on them moved.

Same process, same outcome set, two sets of weights, two averages. Rs 100/- 102 104 106 108 110 112 Rs 105.13/- the average under Q Rs 108.33/- the average under P Rs 3.20/- WHAT CHANGED the weight on each outcome and therefore the average WHAT DID NOT CHANGE the outcome set, the process, the volatility, the horizon
The same finishing value averages Rs 108.33/- under one measure and Rs 105.13/- under the other, with no path added or removed between them.
Try it out

The measure changes from the first rule to the second. Does the outcome set change, does the average change, or do both?

The error that gets made, and what it costs

Reporting the average as the outcome to plan around. On the standard process the average finish of Rs 108.33/- is reached or beaten in 46.0 per cent of outcomes, so a reader who treats it as the typical case has taken a figure that the majority of outcomes fall short of and given it the job of describing the middle.

The gap is not fixed and it widens fast. It is half the variance rate, so doubling the volatility quadruples it. At 20 per cent the average sits Rs 2.15/- above the middle outcome; at 40 per cent it sits Rs 8.33/- above, and the middle outcome has fallen all the way back to the starting value.

The cost is a plan built around a number the majority of outcomes will not reach, produced by an entirely correct calculation of an entirely correct quantity that happened to be the wrong quantity for the question asked. Nothing in the arithmetic is wrong, and nothing in the arithmetic flags it.

Every outcome of the standard process, split by whether it reaches the average. 54.0 per cent finish below the average of Rs 108.33/- 46.0 per cent reach it or beat it the average sits here A figure described as the average finish is described accurately. It is simply not the figure that answers the question most readers are actually asking, which is where the middle of the outcomes sits.
The average finish of Rs 108.33/- is reached or beaten in 46.0 per cent of outcomes, so a plan built on it is a plan built on a number most outcomes do not reach.
Try it out

A projection has been built around the average finish of the standard process. What is the single sentence that would fix it?

The average moves and the outcome set does not. See what the expectation carries.

What does someone reading a projection actually do with all this?

Two habits, and they are cheap.

The first is to demand the pair rather than the single figure. Meet any statement of the form the average outcome is such and such with two more: where does the middle outcome sit, and what share of outcomes reach the average? On the standard process those are Rs 106.18/- and 46.0 per cent, and quoting the three numbers together takes one extra line. A household deciding how much of a projected amount to lean on wants the middle figure. Half the outcomes beat the middle figure. Somebody sizing a worst case wants neither of them and wants a low quantile instead. Different questions want different summaries of the same distribution, and the average answers only one of them.

The second habit is to ask which measure produced the number before comparing two of them. An average of Rs 108.33/- and an average of Rs 105.13/- for the same quantity is not a disagreement to be resolved; it is two correct answers to two different weightings, and putting them in the same column without a label is how a comparison quietly becomes wrong. The label on the measure is part of the number.

The third, if a third is allowed, is noticing when a conditional average is being reported as though it were settled. A conditional average is a random variable, so a statement like the average finish from here is Rs 115.61/- is only true from here. At a different node the whole figure moves with it. The tower property matters in practice for that reason as much as in theory: it shows the pieces still add up once the average is taken across all the places the process might have been standing.

Universal

Which jurisdiction sets these rules

None. The definitions here are mathematics and hold identically everywhere; no regulator sets a probability axiom and no jurisdiction block applies. Where a contract convention rather than a mathematical claim is named, that convention is the one thing on which a venue or clearing body is the authority, and it is confirmed at source where it is used.

Estimating an expectation from observations, and any argument about sampling error, are covered separately. Why the correction to the growth rate of the logarithm is half the variance rate belongs to the chain rule for random processes, covered separately later in this subject; the gap it produces is used here rather than derived. How the second measure is constructed, and why one exists at all, are covered separately. Processes whose average never moves are the subject of the next sequence.
Breaking Into Quants Bootcamp — Fin Maverick

References

SourceDocumentWhere
arXiv, Quantitative FinancePreprints on probability foundations for derivative pricingarxiv.org
Social Science Research NetworkWorking papers on conditional expectation and measure change in financessrn.com
Steven E. ShreveStochastic Calculus for Finance II: Continuous-Time Models, for conditional expectation and the tower propertySpringer

The standard process, the locked path and the at-the-money contract at Rs 100/- are invented.
Educational material. Not advice on any investment, tax, budget or market position.

Covered in this topic

Subtopics

Conditional ExpectationConditional Probability vs Conditional Expectation
← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.