Variance Reduction: Fewer Paths for the Same Accuracy
Variance reduction is any technique that lowers the spread of a sampling method's answers without changing the quantity being estimated. The technique buys accuracy without buying paths, and with a control variate the gain can be computed in advance rather than discovered afterwards. At the locked settings it leaves 0.145292 of the variance, the same accuracy from about one seventh of the paths.
The sampling method leaves one uncomfortable number behind. The spread of a single discounted payoff on the standard process is 14.719404, and the accuracy of an average falls only with the square root of how many of them are drawn. Halving the error takes four times as many paths. Dividing it by ten takes a hundred times as many. The bargain is a hard one, and it can be refused.
The refusal works because the square root rule is a statement about the spread of the thing being averaged, and that spread is something the analyst is allowed to change. The rule says the standard error is that spread divided by the square root of the path count. Everyone stares at the path count sitting in the denominator. The numerator is the part that can be acted on, and acting on it costs no paths at all.
What is variance reduction?
Variance reductionLowering the spread of a sampling method's answers without changing the quantity it is estimating. is the collective name for every technique that lowers the spread of what is being averaged while leaving what it averages exactly where it was. Both halves of that sentence matter, and the second half is the one people forget. A technique that lowers the spread and quietly moves the target is not variance reduction. Such a technique is a mistake wearing a friendly face.
Here is the everyday version, and it is a counting example rather than anything to do with money. The question is how many people are in a large hall. Counting them all is impossible, so one section is counted carefully and that count is scaled up by the number of sections. The section that happened to be picked might have been unusually full or unusually empty, so the answer is noisy. The only obvious cure is to go and count more sections.
Now add one fact already in hand. While the counting goes on in that section, the chairs are counted too. Someone put the chairs there and wrote the number down, so the number of chairs the whole hall contains is known exactly. If the scaled-up chair count comes out four per cent above the number known to be true, then the section picked was a crowded one, and the scaled-up head count is very likely too high as well. So it is pulled back. Not a single extra person has been counted, and the answer has got better.
The counting trick is the whole idea, and everything below is that idea with arithmetic attached. A quantity whose truth is already known, measured alongside the quantity that is not, shows which way this particular sample leaned, and the lean is corrected for. In the sampling method the sections are paths, the people are the discounted payoff, and the chairs are whatever companion quantity has a true average that can be written down in advance.
Read the dark strip at the foot of that figure again. The strip is the honest statement of what the technique buys. The gain is not that the answer becomes right. The answer was already right on average. The gain is that the same accuracy arrives for about a seventh of the work. Where each path is a full walk of a process over many steps, a seventh of the work is often the difference between a calculation that finishes overnight and one that does not.
A technique is reported to have lowered the spread of a sampling method's answers. What is the second thing to check before calling it variance reduction?
What is a control variate?
A control variateA companion quantity whose true average is already known, used to remove the part of the error that moved with it. is the chair count. Formally it is a second quantity, computed on the very same paths as the thing actually wanted, whose true average is already known exactly. Because that true average is known, the distance the sample average of the control has strayed on this particular draw is visible. The straying is pure sampling error, visible and measurable, and the only sampling error in the whole exercise that will ever be visible.
A multiple of the visible error is then subtracted from the estimate. The multiple is a chosen number, and choosing it well is the subject of a later section. Call the thing wanted the payoff estimate, call the companion the control, and the corrected estimator is a single line.
| \(\overline{Y}\) | the average over the drawn paths of the discounted payoff, the quantity actually wanted |
| \(\overline{X}\) | the average over those same paths of the control, computed path by path alongside \(Y\) |
| \(\mathbb{E}^{\mathbb{Q}}[X]\) | the true average of the control under the risk-neutral measure Q, known in advance rather than estimated |
| \(b\) | the coefficient, a number chosen by whoever is doing the calculation |
Look hard at what sits inside the bracket. The bracket holds a sample average minus a known truth, so the bracket holds an error, and the true average of that error is exactly nought. Nothing has been broken, and that nought is the whole reason. Subtracting a multiple of a quantity whose true average is nought leaves the true average of the whole expression exactly where it started, at the price the method is trying to find. The corrected estimator is unbiasedEstimating the right thing on average, whatever any one sample happens to do. for every value of the coefficient without exception, including values that are badly chosen.
The last clause is worth holding on to. The failure set out under the error that gets made grows entirely from it. Every coefficient gives an unbiased answer. The coefficient decides only how noisy the answer is. Nothing anywhere in the arithmetic ever complains about a bad one.
What plays the part of the chairs on the standard process?
The control used here is the level of the standard process at the horizon, written S with a subscript T. The level at the horizon is a good choice for one plain reason. The discounted payoff is a function of it. When the level finishes high the payoff is large, when the level finishes below the strike the payoff is nought, and the two move together by construction rather than by luck.
And its true average is known in closed form. Under the risk-neutral measureThe measure Q under which the standard process is expected to grow at the risk-free rate and at no other rate. the standard process is expected to grow at the risk-free rate and at nothing else, so its average level at the horizon is the starting value compounded at that rate. No simulation is needed and no approximation is involved.
| \(S_T\) | the standard process at the horizon, which is the control used here |
| \(S_0\) | the starting value of the standard process, Rs 100/- exactly |
| \(r\) | the risk-free rate, 0.05 a year, continuously compounded |
| \(T\) | the horizon, one year |
Set that beside the average of the same level under the physical measure P, Rs 108.33/-. The two figures differ, and because the paths being drawn are drawn under Q, only the first is the control's true average here. Using the other one would push a fixed error into every single estimate and would not be variance reduction at all. The fixed error would be bias, introduced deliberately and hidden completely.
What must be true of a quantity for it to be usable as a control?
Why does a control variate work?
Thinking the control adds information about the payoff is easy, so the mechanism is worth stating slowly. The control does no such thing. The control adds nothing whatever about the payoff. The control adds information about this draw, and that turns out to be the only thing missing.
Consider the hall again. The chairs say nothing about how crowded halls are in general. The chairs say that this section was a fuller than average section. The head count came from that same section, so the direction in which that head count is likely to be wrong is now known. The knowledge is entirely about the draw and not at all about the subject.
Now the arithmetic. The corrected estimator is one quantity minus a multiple of another, so its variance follows the ordinary rule for the variance of a difference. Written out, the whole result falls out of it.
| \(\operatorname{Var}(Y)\) | the variance of a single discounted payoff, 216.660857 at the locked settings |
| \(\operatorname{Cov}(Y,X)\) | the covariance of the payoff with the control, 289.002263 at the locked settings |
| \(\operatorname{Var}(X)\) | the variance of the control, 451.028808 at the locked settings |
| \(b\) | the coefficient, the one thing here that is chosen |
The coefficient appears squared with a positive multiplier, so the expression is a parabola in the coefficient, opening upward. A parabola opening upward has exactly one lowest point. The existence of a single best coefficient is not a modelling assumption, it is a property of that shape. And the shape also shows that going too far is punished in exactly the same way as not going far enough, a fact worth meeting twice more.
The middle term is the mechanism in one line. The middle term is a covarianceThe average product of two quantities' departures from their own averages, positive when they tend to stray in the same direction., the average of the product of two departures. When the level at the horizon strays upward the payoff strays upward too, so the covariance is large and positive here. The co-movement is the part of the payoff's error that the control has purchase on, and the correction subtracts it.
The proportion in that split has a name already met. The green part is the squared correlationHow closely two quantities move together, on a scale from minus one to one, ignoring the scales they move on. between the payoff and the control, 0.854708, and the red part is one minus it, 0.145292. The split is why the correlation is the number people reach for. The correlation genuinely does decide how much can be removed. The correlation does not decide how much to subtract, and confusing those two jobs is the mistake set out under the error that gets made.
What is the control used here, and what is its true average?
How much is a control variate worth at the locked settings?
Now put numbers to it. Every figure below is computed from the distribution of the standard process under the pricing measure. Not one of them was measured from a sample, and that is deliberate rather than fastidious. A figure measured from a sample changes when the sample changes, and a worked example that changes on reload is not a worked example. Such an example is a weather report.
How to Apply Variance Reduction in a Teaching Example
The recipe has five steps and none of them involves drawing anything. First, name the control and write down its true average. Here the control is the level at the horizon and its true average is Rs 105.127110/-. Second, compute the variance of the payoff. Third, compute the covariance of the payoff with the control. Fourth, divide that covariance by the control's variance to get the coefficient. Fifth, read off what the corrected spread will be, before a single path exists.
| Quantity | Symbol | Value | Where it comes from |
|---|---|---|---|
| Starting value of the standard process | \(S_0\) | Rs 100/- | a locked parameter, invented |
| Strike of the contract being priced | \(K\) | Rs 100/- | a locked parameter, invented |
| Risk-free rate, continuously compounded | \(r\) | 0.05 | a locked parameter, invented |
| Volatility | \(\sigma\) | 0.20 | a locked parameter, invented |
| Horizon | \(T\) | 1.0 | a locked parameter, invented |
| True average of the control | \(\mathbb{E}^{\mathbb{Q}}[S_T]\) | 105.127110 | the starting value compounded at the risk-free rate |
| Variance of the control | \(\operatorname{Var}(S_T)\) | 451.028808 | the lognormal variance at the horizon |
| Variance of one discounted payoff | \(\operatorname{Var}(Y)\) | 216.660857 | the second moment of the truncated payoff, less the squared price |
| Covariance of payoff with control | \(\operatorname{Cov}(Y,S_T)\) | 289.002263 | the average product, less the product of the averages |
| Correlation | \(\rho\) | 0.924504 | the covariance over the product of the two spreads |
| The best coefficient | \(b^{\star}\) | 0.640762 | the covariance over the control's variance |
The two rows at the bottom are different numbers and the whole of the failure block turns on that. The correlation is also high at 0.924504. One quantity is a function of the other, so a high correlation is what one would expect. High correlation is what makes the technique worth applying here, and it is also what tempts people into the arithmetic mistake.
The strip along the bottom is where the claim becomes concrete. Because the standard errorThe spread of an average, which is the spread of one draw divided by the square root of the number of draws. is the spread of one draw divided by the square root of the path count, both columns fall at the same square root rate. The number they start from is what changes, and it is 5.610624 instead of 14.719404 at every path count without exception.
| \(\mathrm{SE}(n)\) | the standard error of the average over \(n\) paths |
| \(s\) | the spread of a single draw, 14.719404 plain and 5.610624 controlled |
| \(n\) | the number of paths drawn |
The equivalence is the reason the technique is quoted in variance rather than in spread, and it is worth pausing on. The spread fell by a factor of about 2.62, a modest sounding number. The variance fell by a factor of about 6.88, and it is the variance that converts one for one into paths saved. The spread is what is seen and the variance is what is paid for.
The variance falls to 0.145292 of what it was. How many paths are now needed for the same accuracy?
Should the coefficient be set to the correlation between the estimate and the control?
What is the best coefficient, and why is it not one?
Go back to the parabola. To find the lowest point of a quantity that runs as a constant, minus twice the coefficient multiplied by the covariance, plus the squared coefficient multiplied by the control's variance, differentiate once and set the result to nought. The answer is one ratio.
| \(b^{\star}\) | the coefficient that makes the variance of the corrected estimator as small as it can be |
| \(\rho\) | the correlation between the payoff and the control, 0.924504 here |
| \(s_Y\) | the spread of the payoff, 14.719404 |
| \(s_X\) | the spread of the control, 21.237439 |
The middle form of that line is the whole answer to why the coefficient is not the correlation. The correlation is 0.924504. The ratio of the two spreads is 14.719404 over 21.237439, or 0.693088. Multiplying them gives 0.640762. The correlation says which way to lean and the ratio of spreads says how far, and a coefficient needs both.
The everyday version is worth having. The chair count in the hall came out four per cent high. Does that mean the head count is four per cent high? Only if people and chairs vary in exactly the same proportion from section to section. If the chair layout is nearly uniform and the crowd is wildly uneven, a four per cent chair surplus signals a much larger head count surplus, and the sensible correction is bigger than four per cent. If the chairs are scattered chaotically and the crowd is even, the sensible correction is smaller. The correlation on its own is a pure number that has forgotten the units of both quantities, so the coefficient is what has to encode that translation.
Substituting the best coefficient back into the variance expression collapses it to something clean.
| \(\rho^{2}\) | the squared correlation, 0.854708, which is the proportion removed |
| \(1-\rho^{2}\) | the proportion that survives, 0.145292 |
| \(\operatorname{Var}(Y)\) | the variance before any correction, 216.660857 |
The collapsed expression is the cleanest statement of the division of labour between the correlation and the coefficient. Correlation answers how much is available. The coefficient answers how much is taken. Set right, the coefficient collects the whole prize of 185.181759. Set wrong, it collects part of it, or none of it, or goes past it and starts giving some back.
What is the best coefficient made of?
The two grey bars at the ends of that figure are the elegant check, and they are equal by algebra rather than by coincidence. At twice the best coefficient the gain term and the cost term are each exactly twice what they were at the best one, so the gain term becomes four times the removed variance divided by the control's variance while the cost term becomes the same quantity with the opposite sign. The two terms cancel completely. Whatever the undercorrection, overcorrecting by the same amount is worth exactly as much, and that is nothing.
The best coefficient is 0.640762. Before the control below is moved, what happens at exactly twice that?
Move the coefficient and watch the spread bottom out
Held fixed: the variance of the payoff at 216.660857, the covariance with the control at 289.002263, and the variance of the control at 451.028808. The only thing that moves is the coefficient. The curve is the resulting spread, and the red dashed line is the uncontrolled figure of 14.719404. Watch where the curve crosses that line, and notice that it does so twice.
At a coefficient of 0.640762 the spread of the answers is 5.610624, so the variance is 31.479098, which is 0.145292 of the uncontrolled 216.660857 and means one controlled path carries the accuracy of 6.882690 plain ones. This is the lowest the curve goes.
What does variance reduction not do?
Variance reduction does not make the answer more right. Both the corrected estimate and the plain one sit on the truth in expectation already, so neither is closer to it than the other. The distance a typical answer wanders from that truth is what changes, and nothing else changes at all.
Here is the everyday version. A kitchen scale that reads two grams heavy every single time is biased, and no amount of reweighing will fix it. A kitchen scale that jitters by two grams either way at random is noisy, and reweighing ten times and averaging will fix it almost entirely. Variance reduction is a better way of averaging. Variance reduction has nothing whatever to say to the scale that reads heavy.
Three more things the technique does not do. The technique does not fix a wrong model. If the process being sampled is not the intended process, a tighter spread around the wrong answer is more convincing than a loose one, and therefore worse. The technique does not fix a wrong measure. The true average of the control must be the true average under the measure the paths are drawn from, and Rs 105.127110/- is that number under Q and not under P. And the technique does not remove the square root rule. The rule still holds, exactly as before. The technique changes the number sitting on top of the square root and leaves the square root itself completely untouched.
There are other techniques with the same aim. Antithetic variatesAnother variance reduction technique, in which paths are drawn in mirrored pairs so their errors partly cancel. pair each draw with its mirror image so their errors partly cancel. Stratifying the probability scale and taking one value from each slice is another, and this subject area uses a deterministic version of it wherever many outcomes are needed. Both are covered separately. The control variate is treated at length here because it is the one whose worth can be written down exactly in advance, and that is a teaching property rather than a practical ranking.
A coefficient has been set badly. Is the estimate now wrong?
Who reaches for this, and what does it buy them?
Think about who actually has this problem. Anyone valuing a contract whose payoff has no formula has to reach for a numerical method, and where the payoff depends on the whole path rather than the finish, the sampling method is often the only one of the four that will do at all. Such a calculation is priced in machine time, and machine time is priced in money and in the hour of the day it finishes.
So the question a person in that seat asks is never whether the technique is elegant. The question is whether tonight's calculation will be ready by morning, and at what accuracy. A variance ratio of 0.145292 converts directly into a run about a seventh as long, or into a run of the same length reporting an interval less than half as wide. That is a scheduling decision expressed as arithmetic, and it is settled before any machine is switched on, which is the property that makes the control variate the one people actually reach for.
The same reasoning applies wherever a quantity has to be estimated by sampling and a companion quantity has a known average. Someone auditing a large stack of records and sampling a few hundred of them can use a total they already know, such as a footed sum, as a control on the quantity they are estimating. Someone estimating a total from a survey can use a count they already hold from a complete register. The arithmetic set out here is exactly the arithmetic there. Only the names change.
The error that gets made, and what it costs
Setting the coefficient to the correlation because the correlation is high. Here the correlation is 0.924504 and the best coefficient is 0.640762, and those are two different numbers doing two different jobs. The coefficient is the covariance divided by the control's variance, or equally the correlation multiplied by the ratio of the two spreads, and that ratio is 0.693088 rather than one. Dropping it means subtracting too much.
The cost is the spread coming out at 8.233537 rather than 5.610624, and the variance at 67.791131 rather than 31.479098. A little over eighty per cent of the available gain is collected and the rest is handed back. Setting the coefficient to one instead is worse again, giving 9.470224.
And the estimate remains unbiased throughout, and the unbiasedness is precisely what makes the mistake expensive. The answer still converges to Rs 10.450584/-, no diagnostic fires, no residual looks strange, and the reported interval is honestly computed from the observed spread. Nothing anywhere announces that the same run could have been finished in a seventh of the time. The technique was applied correctly in spirit and wrongly in arithmetic, and the arithmetic is the only place the difference lives.
Someone sets the coefficient to the correlation, 0.924504. What happens?
Where this holds, and where rules would come in
The mathematics here is universal. A variance, a covariance and the coefficient that minimises a quadratic are not matters of jurisdiction. Permission to use a model, the documentation required of a valuation and the question of who may rely on it are conduct matters that belong to a particular place and are settled elsewhere.
References
| Source | Document | Where |
|---|---|---|
| arXiv, Quantitative Finance | Preprints on variance reduction and control variates in derivative pricing | arxiv.org |
| Social Science Research Network | Working papers on simulation efficiency and control variate selection | ssrn.com |
The standard process and the contract struck at Rs 100/- are invented.
Educational material. Not advice on any investment, tax, budget or market position.
