Resampled Efficiency: Taming an Unstable Optimiser
Resampled efficiency runs the optimiser many times on inputs drawn around the original estimates, then averages the resulting weights. Resampled efficiency exists because a mean-variance answer moves violently when an input moves slightly. Averaging produces a less jumpy set of weights and a visible spread around them, but it adds no information the original estimates did not already contain.
Why a mean-variance answer moves so far on so little is settled under the mean-variance optimiser and its constraints. Feed the Anantara Multi-Asset Portfolio's own assumption set into a solver, then nudge one assumption by a single point either way: the answer travels from 46.36 per cent equity to 72.86 per cent. More than a quarter of a Rs 500 crore mandate changes hands on a change nobody could defend as knowledge. Something can be done about that, though less than the name of the technique suggests.
The Anantara Multi-Asset Portfolio, an invented discretionary mandate of Rs 500 crore, is run by Faiz Ahmad Ansari for a charitable endowment whose investment committee is chaired by Rukmini Deshpande. Its policy mix is equity 60.0 per cent at Rs 300 crore, fixed income 30.0 per cent at Rs 150 crore and cash 10.0 per cent at Rs 50 crore. Its mandate permits equity between 50 and 70 per cent. Every assumption used below is one the holder wrote down for itself rather than one read off a market or published by anybody.
Why does the answer move so far when the input barely moves?
The blame usually lands in the wrong place. The solver is not unstable and the arithmetic is not fragile: hand a mean-variance problem to two competent people and on the same inputs they return the same weights to six decimal places, every time. Nothing about the machinery wobbles.
The inputs wobble, and the technical name for that wobble is estimation errorThe gap between a quantity as measured or guessed and the quantity itself. Estimation error belongs to the number supplied, not the calculation.: the gap between the number written down and the number that would have been written down with better knowledge. A solver cannot tell a figure measured across forty years of daily observations from one arrived at in a meeting on a Thursday. Both arrive as a decimal, and both are treated as fact.
The everyday version is the same shape at a smaller size. A household is planning a wedding and the caterer charges per plate. The budget, the hall, the tables, the deposit and whether the second cousins get invited all follow from one number: how many people will come. Nobody knows it. Somebody says four hundred, and the whole plan is built on it with beautiful precision. The plan is not fragile, and the arithmetic of four hundred plates is not fragile. The four hundred is fragile, and everything downstream inherited that.
Now translate. The Anantara mandate's own assumption set carries nine numbers: three expected returns, three volatilities and three correlations, with cash taken as uncorrelated. Nothing in the problem statement carries a field for how much anybody believes each one, so the solver treats all nine as facts of equal standing. The instability everybody complains about lives entirely in that missing field. The instability is a problem about estimation, not a problem about solving.
The distinction matters practically. If the trouble were in the arithmetic the fix would be a better solver, and better solvers are cheap. Because the trouble is in the inputs, a better solver changes nothing: it finds the same answer faster and to more decimal places.
Which of the nine inputs is the weak one?
The expected returns, and it is not close. Split the nine by what kind of statement each one is. A volatility says how much something moved about; a correlation says how much two things moved about together. Both describe a stretch of time that has already happened and can be computed from observations that exist, and both move slowly: whatever an equity market's dispersion was over a long stretch, it is unlikely to be a different animal next quarter.
An expected return is a different kind of statement entirely. An expected return is a claim about a period that has not happened yet, and history estimates that claim so poorly that a long run of observations barely narrows it. The reason sits in the statistics layer and is applied here rather than rebuilt: adding observations sharpens a measure of spread far faster than it sharpens a measure of a centre. So the input with the largest effect on the answer is also the one least well known.
The street version is a food stall outside an office gate. The owner knows exactly how much the daily takings bounce around: she has watched them for three years and she has the notebook. Asked for next year's average daily takings she is guessing, and no amount of extra notebook helps. What she is guessing about is whether that office keeps its lease. Input uncertaintyRoom around a number supplied, expressed as a range rather than a single value. Input uncertainty is confidence in the input, not the answer's confidence in itself. is not evenly spread across her figures, and it is not evenly spread across the mandate's nine either.
Of the nine inputs, which is estimated worst, and why does the solver not know that?
What does resampled efficiency actually do?
The procedure has a name. Resampled efficiencyA procedure that solves the same portfolio problem many times on inputs drawn around the original estimates, then reports the average of the resulting weights together with how far they spread. is Richard Michaud's, developed as a response to exactly the instability set out above. The construction being resampled is the mean-variance one Harry Markowitz set out in Portfolio Selection in 1952, worked through under the covariance argument.
The procedure is six steps, and all six fit in the head at once. Take the original estimates, the ones somebody wrote down. Draw a set of inputs around them rather than using them as a single point. Solve the problem on that drawn set and write down the weights that come out. Repeat that many times. Average the weights recorded. Then record how far those weights spread across the whole collection.
Step five is where almost every description of this technique stops, and step six is the one that pays. An average of many runs is a comfortable object: smooth, single-valued, and it fits in a paper. Because the spread shows how much of the answer was never settled, the spread is uncomfortable. The discomfort is why the spread is worth more.
One word in that list is doing quiet work. A drawOne complete set of inputs pulled from around the original estimates, together with the portfolio that comes out when the problem is solved on that set. is one full set of inputs from the neighbourhood of the originals, together with the one portfolio that falls out when the problem is solved on it. Ten thousand draws means ten thousand input sets and ten thousand answers, each internally consistent and none of them the answer.
What happens when it is run on this mandate's own numbers?
A full run varies all nine inputs across thousands of draws. Thousands of draws need a computer and hide the arithmetic. The smallest honest version runs instead, the one that can be checked with a pen. Hold eight of the nine inputs completely still, treat only the equity expected return as uncertain at 12.0 per cent give or take one point, and take three draws, at 11.0, 12.0 and 13.0 per cent.
The problem being solved is the one set out under the mean-variance optimiser: the lowest variance mix that reaches an expected return of 10.05 per cent, with the weights summing to one and no negative weight allowed. At that return target the solver puts nothing at all in cash, a result rather than a convenience, so the answer sits in equity and fixed income only and the arithmetic stays checkable by hand. With cash at nought the return target alone fixes both weights, and the equity weight is 2.55 over the equity assumption less 7.5, where the 2.55 is the target of 10.05 less the fixed income assumption of 7.5.
| Draw | Equity weight | Fixed income | Equity in rupees | Volatility |
|---|---|---|---|---|
| Equity assumed at 11.0 per cent | 72.86 per cent | 27.14 per cent | Rs 3,64,28,57,143/- | 13.45 per cent |
| Equity assumed at 12.0 per cent | 56.67 per cent | 43.33 per cent | Rs 2,83,33,33,333/- | 10.84 per cent |
| Equity assumed at 13.0 per cent | 46.36 per cent | 53.64 per cent | Rs 2,31,81,81,818/- | 9.26 per cent |
| Average of the three equity weights | 58.63 per cent | 41.37 per cent | Rs 2,93,14,57,431/- | see below |
Check one line of it yourself so the rest is trustworthy. At 11.0 the denominator is 11.0 less 7.5, or 3.5, and 2.55 over 3.5 is 0.728571, or 72.86 per cent. On Rs 500 crore that is Rs 3,64,28,57,143/- in equity. Its volatility comes from the two class volatilities of 18.0 and 5.0 and the correlation of 0.20, and works out at 13.45 per cent. Every other row is the same three keystrokes.
The volatility row is where that shows. The three draws carry volatilities of 13.45, 9.26 and 10.84 per cent. Averaging those three gives 11.19 per cent. But the portfolio holding 58.63 per cent equity has a volatility of 11.15 per cent, and 11.15 is not 11.19. An average of portfolios is not the portfolio of the averages, so the averaged weights have to be re-costed as a portfolio in their own right rather than inheriting an average of the parts. The gap here is three hundredths of a point before either figure is rounded, and it is small because only one input moved. The gap does not stay small when nine of them move.
What does averaging the weights buy?
Two things, worth separating carefully: one of them is real and the other is a story people tell.
The real one is smoothness. Nudge an input and a single-point answer lurches; nudge the same input and an average of many runs shuffles. Averaged weightsThe mean of the weights produced across many draws, taken class by class. Averaged weights summarise many answers rather than answering any one problem. also tend to be less concentrated than any single run, for arithmetic reasons rather than wisdom: runs that piled into equity and runs that piled out of it partly cancel when added. On this mandate the extremes at 72.86 and 46.36 land the average at 58.63.
Put the two findings side by side. The average of 58.63 per cent sits 1.96 points from the single point answer of 56.67 per cent, so the averaging changed the recommendation hardly at all. The weight spreadThe distance between the smallest and the largest weight a class took across the draws, reported as a range. The weight spread says how much of the answer the estimates never settled. is 26.5 points, and that is the finding: on one input moving by one point in either direction, more than a quarter of a Rs 500 crore portfolio was never determined at all.
A careful reader will catch the 26.5. Subtracting the unrounded edges, 72.857143 and 46.363636 per cent, gives 26.493506 points, or 26.49 to two places and 26.5 to one. Subtracting the already-rounded figures instead, 72.86 less 46.36, gives 26.50. The one hundredth of a point between those answers is the rounding sitting in the two edges rather than a second quantity, so the spread is named as 26.5 points, with 26.49 shown wherever the subtraction is written out in full.
The averaged answer is 58.63 per cent equity and the single point answer is 56.67 per cent. What did the exercise actually reveal?
Smoothness is real and worth having. A mix that does not lurch every time somebody revises an assumption can survive a committee, a change of analyst and a rewritten paper. Smoothness is a genuine property of the averaged weights, and it is a completely different property from being closer to the truth. The story people tell is the second one, that the averaged mix is closer to the truth. Being closer to the truth is not what happened.
Resampled weights usually come out less concentrated than any single run. Why does that happen?
Suppose ten thousand draws are run instead of three. Does the answer become more accurate?
What can resampling never add, however many draws are taken?
Information. Not a small amount of it, and not some of it. None of it.
Follow the pipe backwards. Every drawn input set was pulled from the neighbourhood of the original estimates, and those estimates were nine numbers somebody wrote down in a meeting. If the equity expected return of 12.0 per cent is two points out, then the neighbourhood is two points out, every single draw inside it is two points out, and the average of ten thousand of them is two points out. Resampling narrows the reported answer without improving the estimate it was drawn from, and anybody presenting it as a route to a better portfolio has confused precisionHow tightly repeated attempts cluster together. Precision says the answers agree with each other, and nothing about whether they agree with the truth. with accuracyHow close an answer sits to the quantity it is trying to measure. An answer can be accurate without being precise, and precise without being accurate..
A full resampling run needs a stated distribution to draw the inputs from and a stated number of draws. The mandate's assumption set fixes neither, so no collection of drawn portfolios follows from it. The three draw version above is therefore not a simplification of a larger run sitting out of sight. Three draws are the whole of what those nine written numbers support.
How should a resampled answer be reported?
Not as a set of weights. The practical claim costs a line in a table. A resampled answer is reported as a set of weights together with the range each weight moved across, and a note saying which inputs were varied and which were held still. Without the range, the reader cannot tell whether the answer was ever determined.
Picture two papers landing on Rukmini Deshpande's desk, both saying 55 per cent equity. In the first, equity moved between 53.5 and 56.5 across the runs; in the second, between 42 and 67. The first says the assumptions genuinely point at a weight in the middle fifties. The second says they do not point at anything, and somebody has averaged a shrug. Reported as weights alone, the two papers are indistinguishable.
Two resampled answers both report 55 per cent equity. One moved across 3 points, the other across 25. Same answer?
Does the mandate admit every draw the assumptions produce?
No, and this is the check the exercise most often skips. Lay the three draws against the constraint the mandate actually carries: equity between 50 and 70 per cent. The middle draw at 56.67 per cent sits comfortably inside. Both of the other two are outside it: 72.86 per cent breaches the ceiling of 70 and 46.36 per cent breaches the floor of 50, so two of the three portfolios being averaged are portfolios this mandate could never have held.
Measure it rather than describing it. The band runs from 46.36 to 72.86. Below the floor sits 3.64 points of it, from 46.36 up to 50.00. Above the ceiling sits 2.86 points, from 70.00 up to 72.86. Together that is 6.49 points of a 26.49 point band, or 24.51 per cent of the whole spread. Just under a quarter of the range those assumptions permit is not a portfolio at all. On Rs 500 crore, the portion below the floor is Rs 18,18,18,182/- and the portion above the ceiling is Rs 14,28,57,143/-.
Turn the same fact round and it becomes more useful still. Instead of asking which weights the mandate admits, ask which equity return assumptions produce an admissible weight at all. A weight of 70 per cent needs an assumption of 11.14 per cent, and a weight of 50 per cent needs 12.60. The mandate therefore admits an assumption anywhere between 11.14 and 12.60 per cent, a window just 1.46 points wide, and the written assumption of 12.0 does not sit in the middle of it. There are 0.86 points of room below 12.0 and only 0.60 points above.
The asymmetry has a practical edge the ceiling usually steals attention from. Because there is less room above 12.0 than below it, widening the uncertainty pushes the band out through the floor before it ever reaches the ceiling: the lower edge lands exactly on 50.00 per cent at plus or minus 0.60 points, and the ceiling holds until about 0.86 points. The first breach happens at the floor, so a committee watching only the ceiling will see the mandate broken later than it actually was.
An inadmissible draw still has to be handled. Averaging it in is the quiet default and the wrong one, because averaging it in mixes portfolios that could never have been held into a figure presented as the answer. The constraint belongs inside every single run, so each draw returns either an admissible portfolio or a statement that none reaches the target on that draw's assumptions. Which of the two happened is itself worth reporting.
One draw wants 72.86 per cent equity and the mandate stops at 70 per cent. Should that draw be averaged in with the others?
Before the control below is touched: the equity return assumption is allowed to move by one point either way. How much of the portfolio does the answer move?
Widen the uncertainty and watch a point become a band
Only the equity return assumption moves. The other eight inputs are held completely still, the return target stays at 10.05 per cent and the equity weight stays 2.55 over the assumption less 7.5. Slide the control from nought to plus or minus two points and watch the answer stop being an answer. The mandate walls at 50 and 70 per cent do not move.
At plus or minus 1.00 points on the equity return assumption, the answer runs from 46.36 to 72.86 per cent equity, a width of 26.49 points, which is Rs 1,32,46,75,325/- of a Rs 500 crore portfolio. Of that band, 24.51 per cent sits outside the mandate's equity range of 50 to 70 per cent.
Both ends of that control teach. At nought the band collapses to a single point at 56.67 per cent, exactly the answer the single mean-variance run produced. A single point is not resampling working well but resampling switched off, and the flat certainty it displays is a claim nobody could defend. At plus or minus two points the upper edge reaches 102.00 per cent. A weight above 100 per cent is not a cautious portfolio. It is not a portfolio at all. Rescaling the chart to fit such a weight would have made an impossible answer look merely like a large one, so the band runs off the chart instead.
When is the whole exercise not worth doing?
Three situations, and naming them is part of knowing the technique rather than a criticism of it.
The first is when the constraints already pin the answer: narrow the Anantara mandate's equity band from 50 to 70 down to 58 to 62, and almost every draw returns a weight at one edge, so what is measured is mostly a spread the mandate was never going to permit. The second is when the reader is going to round. A procedure that refined 56.67 into 58.63 produces nothing that survives a committee writing 60 into the minutes. The third is the one that actually happens. Nobody reads the spread.
A technique whose output is ignored has spent effort and changed no decision, and changing a decision is the only test that matters for any analysis. That is a fair thing to check before starting rather than after finishing.
How does anybody use this in a room, on a Tuesday?
Faiz Ahmad Ansari has to put an equity weight in front of Rukmini Deshpande's investment committee. The useful part is not the procedure. It is the shape of the paper he writes.
He opens with the range, not the number: equity moved between 46.36 and 72.86 per cent across the draws. Then what was varied and what was held. One input of nine moved and the other eight did not, so this range is narrower than an honest one. Then the average, 58.63 per cent, carrying no more information than the single point answer it sits 1.96 points from. Then the admissibility line. Only then the recommendation, by which point the room can see how much of it the numbers actually decided.
An analyst reading somebody else's resampled work runs the same questions in reverse. Which inputs were varied, which were held, what was the range on each weight, and were the constraints applied inside each run or after the averaging. A paper that cannot answer the first two has reported a range that means nothing, and a paper that cannot answer the fourth has probably averaged in portfolios that could never have been held.
A household does a smaller version of this without the vocabulary. Say it has Rs 20,00,000/- of savings and is splitting it between a deposit and a market-linked plan, and the whole split turns on a guess about what the market-linked part returns. The useful move is not to sharpen the guess. The useful move is to work the split out under a guess two points lower and two points higher, and to see how far the answer travels. If it barely moves, the guess did not matter. If it moves by half the savings, the household has learned something real: the decision was never being made by evidence.
When is resampling not worth doing?
The error that gets made, and what it costs
An investment committee is shown a resampled answer of about 58.6 per cent equity. The answer arrived from many runs rather than one, so the room treats it as the more reliable of the two figures in front of it and moves on. Nothing has become more reliable. Every run drew from the same written equity return assumption of 12.0 per cent, so if that figure is two points out then every run is two points out and so is their average.
The second half of the error is the one that costs. The 26.5 point spread that the procedure produced, the only genuinely new thing in the whole exercise, was cut from the paper because it looked like uncertainty rather than analysis. The room is left with a number it is now more confident in for reasons that have nothing to do with evidence, and the room has thrown away the one output that would have told it how much of the answer was ever real.
The fix costs a line. Report the spread first and the average second, and say which inputs were varied and which were held still. An exercise that varies one input of nine reports a range that is far too narrow. Add the admissibility line too: on this invented mandate, two of the three draws being averaged were portfolios the holder could never have held.
What is resampled efficiency not?
Resampled efficiency is not a fix for a bad estimate, and it has no way to reach one. Every draw is taken from the neighbourhood of the number supplied, so the procedure inherits that centre and cannot see past it. Resampling is not a claim that any resulting mix performs better, and no performance figure of any kind follows from it.
Resampling is also not a substitute for looking hard at the assumption itself. If 26.5 points of a portfolio hang on one number, the highest value hour available is the one spent on that number rather than on the procedure wrapped around it. Resampling is a way of showing how much of an answer was never determined, and that is a genuinely useful thing to show, provided nobody mistakes the showing for a cure.
Where the rules around a mandate like this sit
The arithmetic here is universal and depends on no threshold set by anybody, and every constraint quoted is the invented holder's own. Where a real mandate carries a limit, a disclosure duty or a registration requirement, the current wording is published by the Securities and Exchange Board of India at sebi.gov.in, and by the Pension Fund Regulatory and Development Authority at pfrda.org.in where the mandate is a retirement one.
References
| Source | Document | Where |
|---|---|---|
| Richard Michaud | The originating work on resampled efficiency | ideas.repec.org |
| Harry Markowitz | Portfolio Selection, 1952, for the mean-variance construction that is being resampled | ideas.repec.org |
| Securities and Exchange Board of India | Any limit, disclosure duty or registration requirement touching a managed mandate | sebi.gov.in |
| Pension Fund Regulatory and Development Authority | The authority where the mandate is a retirement one | pfrda.org.in |
The Anantara Multi-Asset Portfolio, the charitable endowment that holds it, Rukmini Deshpande and Faiz Ahmad Ansari are invented.
Educational material. Not advice on any investment, tax, budget or market position.
