Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Stochastic Calculus & Derivative Pricing Theory
1Probability Foundations
The Probability SpaceRandom VectorsSigma-AlgebraExpectationSample Space and EventsDensity and Distribution FunctionsRisk-Neutral ProbabilityState Price Density vs…
2Stochastic Processes and Jumps
Properties of a Stochastic ProcessMartingaleBrownian Motion and Its PropertiesBrownian Motion vs Geometric…Stopping TimeThe Markov PropertyState VariablesTransition ProbabilityQuadratic VariationQuadratic Variation vs Ordinary…Submartingale and SupermartingaleMartingale RepresentationMarkov Process vs MartingaleOptional StoppingFiltrationJump ProcessesThe Poisson ProcessLevy ProcessesJump Diffusion
3Ito Calculus
The Ito IntegralThe Ito Integral vs the Riemann IntegralInfinitesimals in Stochastic CalculusQuadratic CovariationIto's LemmaHow to Apply Ito's…The Infinitesimal GeneratorIto Calculus vs Ordinary Calculus
4Stochastic Differential Equations
Stochastic Differential EquationsStochastic Differential Equation vs…Drift and DiffusionStrong and Weak Solutions ComparedDiscretisationGeometric Brownian Motion
5Pricing Theory and No-Arbitrage
No-ArbitrageGirsanov, Radon-Nikodym and Change…Physical and Risk-Neutral Measures…The Fundamental Theorems of…The Law of One PriceThe Pricing KernelDiscount Factors and Zero-Coupon PricesReplication vs HedgingComplete Market vs Incomplete MarketClearing Margin Architecture
6Option Pricing Theory
European and American OptionsMonte Carlo European OptionThe Black-Scholes PDEBlack Scholes and the GreeksThe Payoff FunctionThe Binomial ModelBinomial Option PricingDelta Hedging in TheoryBoundary, Initial and Terminal ConditionsThe Exercise BoundaryHow to Check Put-Call…
7Volatility Models
Constant, Local and Stochastic…Vasicek Model vs CIR ModelThe Heston ModelThe SABR ModelThe Volatility ProcessImplied VolatilityVolatility Smile vs Skew vs Surface
8Interest Rate Models
Interest-Rate DerivativesMean ReversionThe Zero-Coupon BondThe Ornstein-Uhlenbeck ProcessThe Discount CurveZero RatesShort-Rate Model vs Market Model
9Numerical Pricing
Closed Form and Numerical…Monte Carlo PricingEuler and Milstein Schemes ComparedTree MethodsFinite Difference MethodsNumerical Error and StabilityVariance Reduction
10Calibration and Model Risk
Model OverrideMarket Price and Model PriceCalibrationHow to Document a Pricing ModelThe Educational Illustration LabelMarket ConventionsModel Uncertainty and LimitationsBacktesting a Pricing ModelIdentifiabilityCalibrated ParametersThe Calibration Loss Function

Variance Reduction: Fewer Paths for the Same Accuracy

Variance reduction is any technique that lowers the spread of a sampling method's answers without changing the quantity being estimated. The technique buys accuracy without buying paths, and with a control variate the gain can be computed in advance rather than discovered afterwards. At the locked settings it leaves 0.145292 of the variance, the same accuracy from about one seventh of the paths.

The sampling method leaves one uncomfortable number behind. The spread of a single discounted payoff on the standard process is 14.719404, and the accuracy of an average falls only with the square root of how many of them are drawn. Halving the error takes four times as many paths. Dividing it by ten takes a hundred times as many. The bargain is a hard one, and it can be refused.

The refusal works because the square root rule is a statement about the spread of the thing being averaged, and that spread is something the analyst is allowed to change. The rule says the standard error is that spread divided by the square root of the path count. Everyone stares at the path count sitting in the denominator. The numerator is the part that can be acted on, and acting on it costs no paths at all.

What is variance reduction?

Variance reductionLowering the spread of a sampling method's answers without changing the quantity it is estimating. is the collective name for every technique that lowers the spread of what is being averaged while leaving what it averages exactly where it was. Both halves of that sentence matter, and the second half is the one people forget. A technique that lowers the spread and quietly moves the target is not variance reduction. Such a technique is a mistake wearing a friendly face.

Here is the everyday version, and it is a counting example rather than anything to do with money. The question is how many people are in a large hall. Counting them all is impossible, so one section is counted carefully and that count is scaled up by the number of sections. The section that happened to be picked might have been unusually full or unusually empty, so the answer is noisy. The only obvious cure is to go and count more sections.

Now add one fact already in hand. While the counting goes on in that section, the chairs are counted too. Someone put the chairs there and wrote the number down, so the number of chairs the whole hall contains is known exactly. If the scaled-up chair count comes out four per cent above the number known to be true, then the section picked was a crowded one, and the scaled-up head count is very likely too high as well. So it is pulled back. Not a single extra person has been counted, and the answer has got better.

The counting trick is the whole idea, and everything below is that idea with arithmetic attached. A quantity whose truth is already known, measured alongside the quantity that is not, shows which way this particular sample leaned, and the lean is corrected for. In the sampling method the sections are paths, the people are the discounted payoff, and the chairs are whatever companion quantity has a true average that can be written down in advance.

Ten thousand paths, both times. The narrow one drew no extra path. 10.25 10.35 10.45 10.55 10.65 the estimate, in rupees the target, Rs 10.450584/- NO CONTROL: one standard error is 0.147194 each way WITH THE CONTROL: 0.056106 each way this much error, on both sides, removed To reach the narrow interval without the control would take 68,827 paths instead of 10,000. Same accuracy, about one seventh of the paths, no extra path drawn.
At ten thousand paths the standard error falls from 0.147194 to 0.056106 with no extra paths drawn, only extra arithmetic on each path, and matching that accuracy without the control would take 68,827 paths.

Read the dark strip at the foot of that figure again. The strip is the honest statement of what the technique buys. The gain is not that the answer becomes right. The answer was already right on average. The gain is that the same accuracy arrives for about a seventh of the work. Where each path is a full walk of a process over many steps, a seventh of the work is often the difference between a calculation that finishes overnight and one that does not.

Try it out

A technique is reported to have lowered the spread of a sampling method's answers. What is the second thing to check before calling it variance reduction?

What is a control variate?

A control variateA companion quantity whose true average is already known, used to remove the part of the error that moved with it. is the chair count. Formally it is a second quantity, computed on the very same paths as the thing actually wanted, whose true average is already known exactly. Because that true average is known, the distance the sample average of the control has strayed on this particular draw is visible. The straying is pure sampling error, visible and measurable, and the only sampling error in the whole exercise that will ever be visible.

A multiple of the visible error is then subtracted from the estimate. The multiple is a chosen number, and choosing it well is the subject of a later section. Call the thing wanted the payoff estimate, call the companion the control, and the corrected estimator is a single line.

The controlled estimator
$$ \widehat{C}(b) \;=\; \overline{Y} \;-\; b\left(\overline{X} - \mathbb{E}^{\mathbb{Q}}[X]\right) $$
\(\overline{Y}\)the average over the drawn paths of the discounted payoff, the quantity actually wanted
\(\overline{X}\)the average over those same paths of the control, computed path by path alongside \(Y\)
\(\mathbb{E}^{\mathbb{Q}}[X]\)the true average of the control under the risk-neutral measure Q, known in advance rather than estimated
\(b\)the coefficient, a number chosen by whoever is doing the calculation
What it says in wordsThe ordinary estimate is taken, the distance the control's sample average has strayed from the true average already known is worked out, that straying is multiplied by a chosen number, and the product is subtracted. The correction uses no new paths and no new information about the payoff.

Look hard at what sits inside the bracket. The bracket holds a sample average minus a known truth, so the bracket holds an error, and the true average of that error is exactly nought. Nothing has been broken, and that nought is the whole reason. Subtracting a multiple of a quantity whose true average is nought leaves the true average of the whole expression exactly where it started, at the price the method is trying to find. The corrected estimator is unbiasedEstimating the right thing on average, whatever any one sample happens to do. for every value of the coefficient without exception, including values that are badly chosen.

The last clause is worth holding on to. The failure set out under the error that gets made grows entirely from it. Every coefficient gives an unbiased answer. The coefficient decides only how noisy the answer is. Nothing anywhere in the arithmetic ever complains about a bad one.

What plays the part of the chairs on the standard process?

The control used here is the level of the standard process at the horizon, written S with a subscript T. The level at the horizon is a good choice for one plain reason. The discounted payoff is a function of it. When the level finishes high the payoff is large, when the level finishes below the strike the payoff is nought, and the two move together by construction rather than by luck.

And its true average is known in closed form. Under the risk-neutral measureThe measure Q under which the standard process is expected to grow at the risk-free rate and at no other rate. the standard process is expected to grow at the risk-free rate and at nothing else, so its average level at the horizon is the starting value compounded at that rate. No simulation is needed and no approximation is involved.

The control's true average, known exactly
$$ \mathbb{E}^{\mathbb{Q}}\!\left[S_T\right] \;=\; S_0\,e^{rT} \;=\; 100 \times e^{0.05} \;=\; 105.127110 $$
\(S_T\)the standard process at the horizon, which is the control used here
\(S_0\)the starting value of the standard process, Rs 100/- exactly
\(r\)the risk-free rate, 0.05 a year, continuously compounded
\(T\)the horizon, one year
What it says in wordsUnder the pricing measure the standard process is expected to grow at the risk-free rate and at no other rate, so the average of its level at the horizon is the starting value multiplied by the compounding factor. That comes to Rs 105.127110/- and it needs no simulation to find.

Set that beside the average of the same level under the physical measure P, Rs 108.33/-. The two figures differ, and because the paths being drawn are drawn under Q, only the first is the control's true average here. Using the other one would push a fixed error into every single estimate and would not be variance reduction at all. The fixed error would be bias, introduced deliberately and hidden completely.

Try it out

What must be true of a quantity for it to be usable as a control?

Derivatives Foundation Bootcamp — Fin Maverick

Why does a control variate work?

Thinking the control adds information about the payoff is easy, so the mechanism is worth stating slowly. The control does no such thing. The control adds nothing whatever about the payoff. The control adds information about this draw, and that turns out to be the only thing missing.

Consider the hall again. The chairs say nothing about how crowded halls are in general. The chairs say that this section was a fuller than average section. The head count came from that same section, so the direction in which that head count is likely to be wrong is now known. The knowledge is entirely about the draw and not at all about the subject.

The control's error is the only error on the whole exercise that can actually be seen. 1. THE CONTROL, THIS DRAW true average, known exactly 105.127110 sample average on this draw the stray 2. SCALE IT multiply the stray by 0.640762 the coefficient, which is not the correlation 3. SUBTRACT IT from the payoff estimate the target does not move at all, because the stray averages nought No extra path is drawn and no new fact about the payoff is used. The control speaks only about this draw. The control works because its true average is known exactly, which is what makes its error visible.
The control works because its true average of Rs 105.127110/- is known exactly, which makes the straying of its sample average visible, so that straying can be scaled and subtracted.

Now the arithmetic. The corrected estimator is one quantity minus a multiple of another, so its variance follows the ordinary rule for the variance of a difference. Written out, the whole result falls out of it.

The variance of the corrected estimator
$$ \operatorname{Var}\!\left[Y - b\,(X - \mathbb{E}[X])\right] \;=\; \operatorname{Var}(Y) \;-\; 2b\operatorname{Cov}(Y,X) \;+\; b^{2}\operatorname{Var}(X) $$
\(\operatorname{Var}(Y)\)the variance of a single discounted payoff, 216.660857 at the locked settings
\(\operatorname{Cov}(Y,X)\)the covariance of the payoff with the control, 289.002263 at the locked settings
\(\operatorname{Var}(X)\)the variance of the control, 451.028808 at the locked settings
\(b\)the coefficient, the one thing here that is chosen
What it says in wordsThe variance of the corrected quantity is the variance started with, less twice the coefficient multiplied by the covariance, plus the squared coefficient multiplied by the control's variance. The middle term is the gain and the last term is the cost, and the coefficient sets how much of each is taken.

The coefficient appears squared with a positive multiplier, so the expression is a parabola in the coefficient, opening upward. A parabola opening upward has exactly one lowest point. The existence of a single best coefficient is not a modelling assumption, it is a property of that shape. And the shape also shows that going too far is punished in exactly the same way as not going far enough, a fact worth meeting twice more.

The middle term is the mechanism in one line. The middle term is a covarianceThe average product of two quantities' departures from their own averages, positive when they tend to stray in the same direction., the average of the product of two departures. When the level at the horizon strays upward the payoff strays upward too, so the covariance is large and positive here. The co-movement is the part of the payoff's error that the control has purchase on, and the correction subtracts it.

One bar, split once. The correction takes the left piece and nothing else. Variance of one discounted payoff: 216.660857 ALL OF IT, BEFORE ANY CORRECTION MOVED WITH THE CONTROL: 185.181759 0.854708 of the total, which is the squared correlation DID NOT: 31.479098 0.145292 of the total, and this is what survives 185.181759 plus 31.479098 is 216.660857 exactly. The split is a computed decomposition, not an approximation.
What survives the correction is the part of the payoff's variance that did not move with the control, 31.479098 out of 216.660857, and the two parts recombine to the total with nothing left over.

The proportion in that split has a name already met. The green part is the squared correlationHow closely two quantities move together, on a scale from minus one to one, ignoring the scales they move on. between the payoff and the control, 0.854708, and the red part is one minus it, 0.145292. The split is why the correlation is the number people reach for. The correlation genuinely does decide how much can be removed. The correlation does not decide how much to subtract, and confusing those two jobs is the mistake set out under the error that gets made.

Try it out

What is the control used here, and what is its true average?

Risk Management Program Bootcamp — Fin Maverick

How much is a control variate worth at the locked settings?

Now put numbers to it. Every figure below is computed from the distribution of the standard process under the pricing measure. Not one of them was measured from a sample, and that is deliberate rather than fastidious. A figure measured from a sample changes when the sample changes, and a worked example that changes on reload is not a worked example. Such an example is a weather report.

How to Apply Variance Reduction in a Teaching Example

The recipe has five steps and none of them involves drawing anything. First, name the control and write down its true average. Here the control is the level at the horizon and its true average is Rs 105.127110/-. Second, compute the variance of the payoff. Third, compute the covariance of the payoff with the control. Fourth, divide that covariance by the control's variance to get the coefficient. Fifth, read off what the corrected spread will be, before a single path exists.

QuantitySymbolValueWhere it comes from
Starting value of the standard process\(S_0\)Rs 100/-a locked parameter, invented
Strike of the contract being priced\(K\)Rs 100/-a locked parameter, invented
Risk-free rate, continuously compounded\(r\)0.05a locked parameter, invented
Volatility\(\sigma\)0.20a locked parameter, invented
Horizon\(T\)1.0a locked parameter, invented
True average of the control\(\mathbb{E}^{\mathbb{Q}}[S_T]\)105.127110the starting value compounded at the risk-free rate
Variance of the control\(\operatorname{Var}(S_T)\)451.028808the lognormal variance at the horizon
Variance of one discounted payoff\(\operatorname{Var}(Y)\)216.660857the second moment of the truncated payoff, less the squared price
Covariance of payoff with control\(\operatorname{Cov}(Y,S_T)\)289.002263the average product, less the product of the averages
Correlation\(\rho\)0.924504the covariance over the product of the two spreads
The best coefficient\(b^{\star}\)0.640762the covariance over the control's variance

The two rows at the bottom are different numbers and the whole of the failure block turns on that. The correlation is also high at 0.924504. One quantity is a function of the other, so a high correlation is what one would expect. High correlation is what makes the technique worth applying here, and it is also what tempts people into the arithmetic mistake.

One move, computed in advance, not discovered afterwards. 14.719404 no control 5.610624 with the control, at the best coefficient the spread of the answers variance 216.660857 becomes 31.479098 a variance ratio of 0.145292 so about one seventh of the paths 0 Standard error at 100 paths: 1.471940 becomes 0.561062. At 1,000: 0.465468 becomes 0.177423. At 10,000: 0.147194 becomes 0.056106.
The spread of the answers falls from 14.719404 to 5.610624, leaving 0.145292 of the variance, which is the same accuracy from about one seventh of the paths.

The strip along the bottom is where the claim becomes concrete. Because the standard errorThe spread of an average, which is the spread of one draw divided by the square root of the number of draws. is the spread of one draw divided by the square root of the path count, both columns fall at the same square root rate. The number they start from is what changes, and it is 5.610624 instead of 14.719404 at every path count without exception.

The path count that buys a given accuracy
$$ \mathrm{SE}(n) \;=\; \frac{s}{\sqrt{n}} \qquad\Longrightarrow\qquad \frac{n_{\text{controlled}}}{n_{\text{plain}}} \;=\; \frac{s_{\text{controlled}}^{2}}{s_{\text{plain}}^{2}} \;=\; 0.145292 $$
\(\mathrm{SE}(n)\)the standard error of the average over \(n\) paths
\(s\)the spread of a single draw, 14.719404 plain and 5.610624 controlled
\(n\)the number of paths drawn
What it says in wordsBecause the standard error falls with the square root of the path count, fixing the accuracy and asking how many paths it takes turns the ratio of the two variances directly into the ratio of the two path counts. A variance ratio of 0.145292 is a path ratio of 0.145292, which is about one seventh.

The equivalence is the reason the technique is quoted in variance rather than in spread, and it is worth pausing on. The spread fell by a factor of about 2.62, a modest sounding number. The variance fell by a factor of about 6.88, and it is the variance that converts one for one into paths saved. The spread is what is seen and the variance is what is paid for.

Try it out

The variance falls to 0.145292 of what it was. How many paths are now needed for the same accuracy?

Try it out

Should the coefficient be set to the correlation between the estimate and the control?

Regression for Finance — free micro-course from Fin Maverick

What is the best coefficient, and why is it not one?

Go back to the parabola. To find the lowest point of a quantity that runs as a constant, minus twice the coefficient multiplied by the covariance, plus the squared coefficient multiplied by the control's variance, differentiate once and set the result to nought. The answer is one ratio.

The best coefficient
$$ b^{\star} \;=\; \frac{\operatorname{Cov}(Y,X)}{\operatorname{Var}(X)} \;=\; \rho\,\frac{s_Y}{s_X} \;=\; \frac{289.002263}{451.028808} \;=\; 0.640762 $$
\(b^{\star}\)the coefficient that makes the variance of the corrected estimator as small as it can be
\(\rho\)the correlation between the payoff and the control, 0.924504 here
\(s_Y\)the spread of the payoff, 14.719404
\(s_X\)the spread of the control, 21.237439
What it says in wordsThe best coefficient is the covariance between the estimate and the control divided by the variance of the control. Written the second way it is the correlation multiplied by the ratio of the two spreads, so it carries information about direction and about scale, while the correlation on its own carries only direction.

The middle form of that line is the whole answer to why the coefficient is not the correlation. The correlation is 0.924504. The ratio of the two spreads is 14.719404 over 21.237439, or 0.693088. Multiplying them gives 0.640762. The correlation says which way to lean and the ratio of spreads says how far, and a coefficient needs both.

The everyday version is worth having. The chair count in the hall came out four per cent high. Does that mean the head count is four per cent high? Only if people and chairs vary in exactly the same proportion from section to section. If the chair layout is nearly uniform and the crowd is wildly uneven, a four per cent chair surplus signals a much larger head count surplus, and the sensible correction is bigger than four per cent. If the chairs are scattered chaotically and the crowd is even, the sensible correction is smaller. The correlation on its own is a pure number that has forgotten the units of both quantities, so the coefficient is what has to encode that translation.

Substituting the best coefficient back into the variance expression collapses it to something clean.

The variance that survives at the best coefficient
$$ \operatorname{Var}\!\left[\widehat{C}(b^{\star})\right] \;=\; \operatorname{Var}(Y)\left(1-\rho^{2}\right) \;=\; 216.660857 \times 0.145292 \;=\; 31.479098 $$
\(\rho^{2}\)the squared correlation, 0.854708, which is the proportion removed
\(1-\rho^{2}\)the proportion that survives, 0.145292
\(\operatorname{Var}(Y)\)the variance before any correction, 216.660857
What it says in wordsAt the best coefficient the surviving variance is the original variance multiplied by one minus the squared correlation. The correlation therefore decides the size of the prize, while the coefficient decides whether it is actually collected, and the two are separate questions with separate answers.

The collapsed expression is the cleanest statement of the division of labour between the correlation and the coefficient. Correlation answers how much is available. The coefficient answers how much is taken. Set right, the coefficient collects the whole prize of 185.181759. Set wrong, it collects part of it, or none of it, or goes past it and starts giving some back.

Try it out

What is the best coefficient made of?

Five candidate coefficients. Only one of them is the answer. 14.719404, the uncontrolled spread 14.719404 0 no correction 5.610624 0.640762 the best coefficient 8.233537 0.924504 the correlation, the mistake 9.470224 1.000000 the other common guess 14.719404 1.281525 exactly twice the best these two are equal, exactly Overcorrecting by a given amount is worth exactly what undercorrecting by that same amount is worth, which is precisely nothing. The curve is symmetric about its lowest point.
The best coefficient is 0.640762 and the correlation is 0.924504, and using the second in place of the first leaves a spread of 8.233537 rather than 5.610624.

The two grey bars at the ends of that figure are the elegant check, and they are equal by algebra rather than by coincidence. At twice the best coefficient the gain term and the cost term are each exactly twice what they were at the best one, so the gain term becomes four times the removed variance divided by the control's variance while the cost term becomes the same quantity with the opposite sign. The two terms cancel completely. Whatever the undercorrection, overcorrecting by the same amount is worth exactly as much, and that is nothing.

Try it out

The best coefficient is 0.640762. Before the control below is moved, what happens at exactly twice that?

Play with it

Move the coefficient and watch the spread bottom out

Held fixed: the variance of the payoff at 216.660857, the covariance with the control at 289.002263, and the variance of the control at 451.028808. The only thing that moves is the coefficient. The curve is the resulting spread, and the red dashed line is the uncontrolled figure of 14.719404. Watch where the curve crosses that line, and notice that it does so twice.

coefficient 00.6407621.5
Jump to
One control, one consequence. The lowest point is not where most people put it. 14.719404, no correction at all 0.640762 0 5 10 15 20 0 1.0 1.5 the coefficient the spread of the answers THE SAME READING AS A BAR, AGAINST THE UNCONTROLLED 14.719404 5.610624
Coefficient
0.640762
Spread of the answers
5.610624
Variance
31.479098
Paths saved, as a multiple
6.882690

At a coefficient of 0.640762 the spread of the answers is 5.610624, so the variance is 31.479098, which is 0.145292 of the uncontrolled 216.660857 and means one controlled path carries the accuracy of 6.882690 plain ones. This is the lowest the curve goes.

Educational illustration. Every reading is computed from the variance expression on each move of the control and nothing is sampled, so the default reproduces the worked instance exactly. The five coefficients read: 0 gives 14.719404, 0.640762 gives 5.610624, 0.924504 gives 8.233537, 1.000000 gives 9.470224, and exactly twice the best gives 14.719404 again. Twice the best coefficient is worth precisely as much as no correction at all. Assumptions on screen: starting value Rs 100/-, strike Rs 100/-, risk-free rate 5 per cent, volatility 20 per cent, one year, and a control whose true average is Rs 105.127110/-. Every one of those is a locked parameter chosen for the worked instance. The coefficient reads 1.281525 at twice the best because 1.281525 is the six decimal rounding of that value; substituting the rounded figure by hand moves the last decimal of the spread, which is why the control carries the unrounded number.
Hypothesis Testing teaches you to run a test, say what it can and cannot support, and recognise a manufactured result.

What does variance reduction not do?

Variance reduction does not make the answer more right. Both the corrected estimate and the plain one sit on the truth in expectation already, so neither is closer to it than the other. The distance a typical answer wanders from that truth is what changes, and nothing else changes at all.

Here is the everyday version. A kitchen scale that reads two grams heavy every single time is biased, and no amount of reweighing will fix it. A kitchen scale that jitters by two grams either way at random is noisy, and reweighing ten times and averaging will fix it almost entirely. Variance reduction is a better way of averaging. Variance reduction has nothing whatever to say to the scale that reads heavy.

Three coefficients, three widths, one centre. The centre never moves. the target, Rs 10.450584/-, unmoved in all three no control coefficient 0.924504 coefficient 0.640762 the estimate at a hundred paths, in rupees The coefficient changes the width and never the centre, which is exactly why a bad one hides.
Every version estimates the same quantity, so a mistake in the coefficient makes the answer noisier rather than wrong, and nothing on the screen looks amiss.

Three more things the technique does not do. The technique does not fix a wrong model. If the process being sampled is not the intended process, a tighter spread around the wrong answer is more convincing than a loose one, and therefore worse. The technique does not fix a wrong measure. The true average of the control must be the true average under the measure the paths are drawn from, and Rs 105.127110/- is that number under Q and not under P. And the technique does not remove the square root rule. The rule still holds, exactly as before. The technique changes the number sitting on top of the square root and leaves the square root itself completely untouched.

There are other techniques with the same aim. Antithetic variatesAnother variance reduction technique, in which paths are drawn in mirrored pairs so their errors partly cancel. pair each draw with its mirror image so their errors partly cancel. Stratifying the probability scale and taking one value from each slice is another, and this subject area uses a deterministic version of it wherever many outcomes are needed. Both are covered separately. The control variate is treated at length here because it is the one whose worth can be written down exactly in advance, and that is a teaching property rather than a practical ranking.

Try it out

A coefficient has been set badly. Is the estimate now wrong?

Who reaches for this, and what does it buy them?

Think about who actually has this problem. Anyone valuing a contract whose payoff has no formula has to reach for a numerical method, and where the payoff depends on the whole path rather than the finish, the sampling method is often the only one of the four that will do at all. Such a calculation is priced in machine time, and machine time is priced in money and in the hour of the day it finishes.

So the question a person in that seat asks is never whether the technique is elegant. The question is whether tonight's calculation will be ready by morning, and at what accuracy. A variance ratio of 0.145292 converts directly into a run about a seventh as long, or into a run of the same length reporting an interval less than half as wide. That is a scheduling decision expressed as arithmetic, and it is settled before any machine is switched on, which is the property that makes the control variate the one people actually reach for.

The same reasoning applies wherever a quantity has to be estimated by sampling and a companion quantity has a known average. Someone auditing a large stack of records and sampling a few hundred of them can use a total they already know, such as a footed sum, as a control on the quantity they are estimating. Someone estimating a total from a survey can use a count they already hold from a complete register. The arithmetic set out here is exactly the arithmetic there. Only the names change.

The error that gets made, and what it costs

Setting the coefficient to the correlation because the correlation is high. Here the correlation is 0.924504 and the best coefficient is 0.640762, and those are two different numbers doing two different jobs. The coefficient is the covariance divided by the control's variance, or equally the correlation multiplied by the ratio of the two spreads, and that ratio is 0.693088 rather than one. Dropping it means subtracting too much.

The cost is the spread coming out at 8.233537 rather than 5.610624, and the variance at 67.791131 rather than 31.479098. A little over eighty per cent of the available gain is collected and the rest is handed back. Setting the coefficient to one instead is worse again, giving 9.470224.

And the estimate remains unbiased throughout, and the unbiasedness is precisely what makes the mistake expensive. The answer still converges to Rs 10.450584/-, no diagnostic fires, no residual looks strange, and the reported interval is honestly computed from the observed spread. Nothing anywhere announces that the same run could have been finished in a seventh of the time. The technique was applied correctly in spirit and wrongly in arithmetic, and the arithmetic is the only place the difference lives.

Try it out

Someone sets the coefficient to the correlation, 0.924504. What happens?

Universal

Where this holds, and where rules would come in

The mathematics here is universal. A variance, a covariance and the coefficient that minimises a quadratic are not matters of jurisdiction. Permission to use a model, the documentation required of a valuation and the question of who may rely on it are conduct matters that belong to a particular place and are settled elsewhere.

Simulating paths to a price is set out under Monte Carlo pricing, along with the square root scaling of its accuracy. Antithetic variates and stratification are set out separately. The corrected price is checked against the locked closed form of Rs 10.450584/-, derived under closed form and numerical approximation compared.
Breaking Into Quants Bootcamp — Fin Maverick

References

SourceDocumentWhere
arXiv, Quantitative FinancePreprints on variance reduction and control variates in derivative pricingarxiv.org
Social Science Research NetworkWorking papers on simulation efficiency and control variate selectionssrn.com

The standard process and the contract struck at Rs 100/- are invented.
Educational material. Not advice on any investment, tax, budget or market position.

Covered in this topic

Subtopics

How to Apply Variance Reduction in a Teaching Example
← Previous
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.