Performance Appraisal: How to Tell Skill From Luck
Performance appraisal asks whether a result was skill or a draw. One stated year of the Anantara Multi-Asset Portfolio put 1.6 percentage points of gross excess return against 3.7 per cent of tracking error, an information ratio of 0.43. Treated as a signal against noise, that ratio would need roughly twenty two such years before it separates from zero.
Behind every review of a managed portfolio sits one question: is the person running the money any good? One year of numbers carries part of an answer. The part is far smaller than the numbers look, and how much smaller is settled by arithmetic four lines long.
The running example is the Anantara Multi-Asset Portfolio, an invented discretionary mandate of Rs 500 crore run by Faiz Ahmad Ansari for an invented charitable endowment whose investment committee is chaired by Rukmini Deshpande. Over one stated twelve month period it returned 14.2 per cent against a composite benchmark that returned 12.6 per cent, with a risk-free rate of 6.5 per cent standing beside both. Its volatility was 11.8 per cent, the benchmark's was 10.4 per cent, its beta against that benchmark was 1.08 and its tracking error was 3.7 per cent. Every figure belongs to that one twelve month period, and one period is the whole of the record this appraisal has to work with.
What is performance appraisal, and how is it different from attribution?
Appraisal and attribution get used as though they were the same activity done at different depths, and they are not. Attribution takes a finished result and explains where it came from, cutting the excess return into parts that sum back to it. Attribution is assumed here in full. Appraisal takes the same finished result and asks a completely different question: does this number carry any information about what happens next, or is it what a coin would have produced?
A perfect attribution and a completely uninformative record are entirely compatible, and moving from one to the other is a change of question rather than a further calculation. A year can be explained down to the last basis point with nothing learned about whether it repeats. Gary P. Brinson's name sits behind the standard way of cutting a result into parts, and nothing in that cutting was ever meant to answer the repetition question.
The everyday version runs like this. A street vendor outside a wedding hall has his best month of the year in November. Asked why, he explains exactly why: eleven weddings booked in the hall next door, four of them large. The wedding bookings are a complete explanation. Every rupee of the extra takings is accounted for. The second question is whether November says anything about how good he is at running a food stall, and on that the explanation just given is silent. The explanation covers the month. The month says nothing about the man. SkillThe part of a result that came from judgement the person can repeat on purpose. It shows up as a tendency across many attempts rather than as any single outcome. and a good month are different objects, and the arithmetic that explains one cannot reach the other.
Why can careful work and a good year have nothing to do with each other?
Because the outcome contains everything the work never claimed to control. A sound process makes a set of decisions defensible in advance, on the information available in advance. The outcome afterwards was decided by a great many things that were not on anybody's desk: which way one currency moved, whether a monsoon arrived, whether a large holder somewhere else needed cash on a particular Tuesday. Call that luckThe part of a result that came from everything the person did not control and cannot repeat on purpose. It is not a moral word here, only a label for the residue., without any sneering attached to the word, and notice that it is present in every result whether the work was good or bad.
Process quality and outcome are two independent axes, so four cases exist and only one of them is ever investigated. Careless work followed by a poor year is the one everybody expects and nobody argues with. Careless work followed by a strong year never gets examined. Nothing about it looks broken. Sound work followed by a poor year is the one that gets somebody removed. Sound work followed by a strong year is the one that gets somebody promoted, and the promotion is a guess wearing a number.
The number does not say which parts of the work were sound, so judging the work only by the number cannot improve the work. The practical consequence is the reason appraisal exists as a separate activity at all. If a committee only ever sees a result, it can only ever reward or punish. Correction requires knowing which decision was the weak one, and a single outcome figure does not carry that.
Which part of the 1.6 points was never a candidate for skill?
Before appraising anything, throw away the part of the result that arithmetic already explains. The Anantara Multi-Asset Portfolio beat its composite benchmark by 1.6 percentage points over the stated year. The portfolio also carried a beta of 1.08, meaning it moved with about eight per cent more of the benchmark's movement than the benchmark itself did. Over the same year, the benchmark returned 12.6 per cent against a risk-free rate of 6.5 per cent, so the benchmark stood 6.1 points above the risk-free rate.
Now the one line that does the work. Carrying 0.08 of extra beta against a benchmark standing 6.1 points above the risk-free rate delivers 0.08 times 6.1. The product is 0.488 percentage points. Subtract that from 1.6 and 1.112 points remain. On a portfolio of Rs 500 crore those two parts are Rs 2,44,00,000/- and Rs 5,56,00,000/-, and they sum to the Rs 8,00,00,000/- that the full 1.6 points is worth. The 0.488 points is arithmetic rather than judgement, so it is not a candidate for skill at all, and the appraisal question therefore applies to 1.112 points and not to 1.6.
There is more than one honest way to cut 1.6 points and the cuts answer different questions, so which cut is being run has to be said plainly. The exposure cut used here asks how much of the excess came from simply carrying more of the market. The other cut asks where in the portfolio the excess arose. Attribution owns that question, and none of its terms belongs in the exposure cut. Mixing a term from one cut with a term from the other produces a sentence that looks like arithmetic and is not.
Of the 1.6 point gross excess return over the stated year, how much is not a candidate for skill at all?
What does an information ratio actually say?
An information ratio scales a result by the amount of departure taken to produce it. The information ratioThe excess return over a benchmark divided by the tracking error against that same benchmark. It is a ratio of a result to the departure taken to get it, not a return. is the excess return over the tracking errorThe volatility of the difference between a portfolio's return and its benchmark's return, quoted for a stated period. It measures how far the portfolio wanders from the benchmark, in either direction.. For the Anantara portfolio over the stated year, the division is 1.6 by 3.7 and the ratio is 0.43. Nothing more complicated is happening. The numerator is the result obtained. The denominator is how far from the benchmark the portfolio stood while obtaining it.
Why bother? Because two records can show the same 1.6 points and mean quite different things. Suppose a second invented mandate also beat its benchmark by 1.6 points over its own stated year, but ran 8.0 per cent of active riskAnother name for tracking error: the size of the departure a portfolio takes from its benchmark, whichever direction that departure runs in. instead of 3.7. Its ratio is 1.6 divided by 8.0, or 0.20. Same headline. Less than half the ratio. The information ratio makes two records comparable when they departed from their benchmarks by different amounts, and a headline excess return can never do that.
And here is the limit, stated in the same breath so it never gets lost. A ratio is still a ratio computed on one period. William F. Sharpe's name sits behind the older measure that divides excess over the risk-free rate by total volatility, and this one is its cousin measured against a benchmark instead. Neither of them says how much record stands behind the number. The ratio compares two results honestly and still cannot settle whether one year of either result means anything at all. The length of the record behind a ratio needs different arithmetic.
Two records both beat their benchmarks by 1.6 points. One ran 3.7 per cent of tracking error and one ran 8.0 per cent. What does the information ratio do here?
How much record does a ratio of 0.43 need before it means anything?
Now the decisive question, and it is worth guessing at before the answer arrives. Almost nobody guesses high enough. The figure on the table is 0.43 of excess return per unit of active risk over one year. How many such years would have to be lined up before the result could be called something other than a draw?
A mandate beat its benchmark by 1.6 points with 3.7 per cent of tracking error. Roughly how many such years would it take before that result separates from zero on the usual convention?
The statistics layer settles the machinery, so what follows applies it rather than teaching it. Several years of excess return are lined up. The signal being sought is the total excess, and it grows in proportion to the number of years. Independent wobbles partly cancel each other rather than adding up, so the noise around the signal grows in proportion to the square root of the number of years. Dividing one by the other, the signal to noiseA result divided by the typical size of the random wobble around it. Above roughly two, the result is hard to explain as a wobble; below it, easy. ratio after a run of years is the information ratio multiplied by the square root of the number of years.
With 0.43 in that relation, one year gives 0.43 times the square root of one, or 0.43. Four years gives 0.86. Nine years gives 1.29. Sixteen years gives 1.72. Reaching two, the conventional line at which people stop calling a result a wobble, requires the number of years to equal two divided by 0.43, all squared. Two divided by 0.43 is 4.65, and 4.65 squared is 21.6. So this record needs roughly twenty two comparable years before its ratio separates from zero. The unrounded ratio of 0.4311 gives 21.5 years, and the difference changes nothing at all.
The twenty two year figure is worthless without the two assumptions holding it up, and both are assumptions rather than facts about anybody: the information ratio must stay at 0.43 across every one of those years, and each year's excess return must be independent of the others. Neither is observable in advance. If the ratio drifts down, the count rises. If the years are correlated, and a manager with a persistent tilt will produce correlated years, the effective sample sizeHow many genuinely separate pieces of evidence a record contains. Twenty correlated years carry less evidence than twenty independent ones, because they partly repeat each other. is smaller than the count of years and the honest number is worse than twenty two, not better.
The twenty two year figure rests on two assumptions. Which pair?
Drag the years and watch how far it still has to go
The information ratio is held at 0.43, the Anantara portfolio's own figure for its one stated twelve month period. The slider lengthens the record. The upper bar is the signal to noise ratio, 0.43 times the square root of the years, and it grows towards the dashed line at two. The lower band is the range within which a result of plus 1.6 points is still an ordinary draw, and it narrows as the record lengthens. The two fixed marks at plus 1.6 and minus 1.6 never move, so the band can be seen shrinking past them.
With one year of record at an information ratio of 0.43, the signal to noise ratio is 0.43, which is 1.57 short of two. A result of plus 1.6 points is still an ordinary draw from a range of plus or minus 7.40 points, and 7.40 points on a portfolio of Rs 500 crore is Rs 37,00,00,000/- either side.
The gross excess return was plus 1.6 points against 3.7 per cent of active risk over the stated year. How unusual would a result of minus 1.6 points have been?
How wide was the range that the 1.6 points came out of?
A tracking error of 3.7 per cent is usually read as a cost or a warning, and that is the wrong reading for appraisal. Read it as a width. The tracking error says that the difference between this portfolio and its benchmark, over a period like the stated one, typically lands somewhere inside a spread of roughly that size either side of zero. The plus 1.6 points that actually happened is one draw from that spread. So is minus 1.6. So is plus 4, and so is minus 4.
An outcome of plus 1.6 points came out of a spread within which a similar negative figure was an entirely ordinary result, so the sign of one year's excess return is close to uninformative on its own. The sign is exactly what committees and newspapers report, and the sign is the part a single year barely supports. Beat or missed. Ahead or behind. The width says those two words describe a coin landing inside a range wide enough to hold both comfortably.
The household version arrives every year. A cousin puts money into one savings arrangement and a neighbour into another, and after twelve months the cousin is ahead by a bit. How much has the year settled? Almost nothing. The year to year wobble in both arrangements is many times larger than the gap between them. The two would have to be watched for a very long time before the gap said anything about the choice rather than about the year, and everybody in the conversation knows this and nobody acts as though they do.
What does the whole appraisal look like, run end to end?
The Anantara Multi-Asset Portfolio's one stated twelve month period, at a risk-free rate of 6.5 per cent and against its unnamed composite benchmark, runs through the steps in the order that removes the easy answers first. The order matters. Each step throws away a part of the headline that was never going to survive scrutiny, and what is left at the end is smaller and more honest than the headline it began from.
| Step | What is done | Result |
|---|---|---|
| 1 | Remove what arithmetic already explains. The benchmark stood 6.1 points above the risk-free rate, and 0.08 of extra beta times 6.1 is 0.488 points | 1.112 points left |
| 2 | Scale the result by the departure taken to get it. The full 1.6 points of gross excess over 3.7 per cent of tracking error | ratio of 0.43 |
| 3 | Ask what one year of that ratio supports. Two divided by 0.43, all squared | 21.6 years |
| 4 | Read the tracking error as the width of the spread the result was drawn from | plus or minus 3.7 |
| 5 | State the verdict the arithmetic supports, and no more than that | unresolved |
Step two carries a discipline point that is easy to walk past. The 3.7 per cent tracking error is not a fourth free measurement of the stated year. Given the portfolio's 11.8 per cent volatility, the benchmark's 10.4 per cent and the beta of 1.08, it is already determined. The identity fits on one line: 139.24 plus 108.16 less twice 116.8128 is 13.7744, and the square root of 13.7744 is 3.71 per cent, carried here as 3.7. Only three of those four figures are free, so quoting all four as though each were separately observed overstates how much the record contains.
Step five is the one people find unsatisfying, so what it says deserves stating precisely. The record is consistent with a manager who has genuine judgement. The record is equally consistent with a manager who has none and had an ordinary year that landed on the right side of zero. One year of a 0.43 ratio cannot separate those two pictures, and one year out of the roughly twenty two needed is not a small shortfall, it is very nearly the whole distance. Writing down that the evidence does not separate the two possibilities is the finding, not a failure to reach one.
The tracking error quoted here is 3.7 per cent. Is that a fourth independent measurement of the stated year?
What can be appraised this year, when the return record cannot carry it?
The constructive question changes what a review meeting is for. If the return evidence needs two decades, then for the next two decades a committee reviewing only returns is reviewing noise. But there is other evidence, it is about the process rather than the outcome, and almost all of it is available now.
Four things sit on the table this year. First, the constraint record: did the portfolio stay inside the equity band of 50 to 70 per cent, the cap of 5 per cent on any single holding and the rest of what the mandate allowed? A constraint record is a fact, not a draw, and compliance monitoring settles that a clean one is evidence about process rather than about outcome. Second, the consistency between what was said in advance and what was actually held: if the reasoning written down in January describes a portfolio nobody built, that gap is visible immediately. Third, whether the reasons given for positions were checkable at the time they were given, rather than assembled afterwards to fit whatever happened. Fourth, whether the record was reported in a way that removed the choice of what to show, so nobody selected the flattering window after seeing the results.
The four checks above are process evidenceEvidence about how decisions were made rather than about how they turned out. It can be checked as soon as the decisions exist, without waiting for outcomes to accumulate. and every one of the four is available immediately. The return evidence needs decades, so the two are not substitutes waiting on the same clock. A committee that understands this stops asking the return record a question it cannot answer and starts asking the process record questions it can.
The same test is already run on people. In choosing a doctor, nobody waits twenty years to see whether her patients did better than average. Nothing usable would be learned in time. The questions asked instead are whether she examined the patient properly, whether her reasoning was stated before the test results came back, and whether she changed her mind when the evidence changed. Answers to those questions are process evidence, checked on the spot, and the move is exactly the same one.
The return record cannot separate skill from luck for decades. What can a committee examine this year instead?
Where any reporting duty on this actually sits
The arithmetic here is universal and nothing in it comes from regulation. Where a managed portfolio in India carries a duty about how performance is computed, presented or disclosed, that text is published by the Securities and Exchange Board of India at sebi.gov.in, and by the Pension Fund Regulatory and Development Authority at pfrda.org.in where a retirement mandate is the setting. Where an index construction rule is in view, the exchanges publish their own methodology at nseindia.com and bseindia.com, and it belongs to the index provider rather than to any manager.
Who runs this arithmetic, and what changes on the Monday after?
Rukmini Deshpande's investment committee meets to review Faiz Ahmad Ansari's stated year. The useful version of that meeting has three parts and only one of them is about the return. First, somebody states the sample the conclusion would need, out loud, before the result is read: at a ratio of this size, roughly twenty two comparable years under assumptions that flatter the case. Second, the return is read with the exposure part already removed, so the committee is looking at 1.112 points rather than 1.6. Third, the bulk of the hour goes on the constraint record and the written reasoning, the one place where evidence actually exists this year.
After the result is on the table the required sample is unconsciously adjusted to whatever the result supports, so the single change that improves the meeting most is stating the required sample before anybody sees the result. Write the number down in advance and it stops moving.
An analyst covering managed portfolios does the same arithmetic in reverse. Given a track record of some length and a claimed ratio, what does the length support? A five year record at 0.43 reaches 0.96 on this arithmetic, less than half the conventional line, and no amount of confident presentation changes that. A lender or a trustee assessing whether an arrangement is working uses the same discipline for a different purpose: not to grade the manager, but to know what weight the numbers in the pack can carry when a decision is taken on them.
And the household version needs no arithmetic at all, only the habit. Before any yearly result is judged, whether it is a savings arrangement, a small shop's takings or a school report, the question is how much of it one year could possibly settle. Sometimes the answer is a great deal. For a spread as wide as a portfolio's excess return, the answer is almost nothing, and knowing that in advance is what stops a household from switching arrangements every January on the strength of a number that was never evidence.
The error that gets made, and what it costs
A committee reviews one year of the Anantara Multi-Asset Portfolio, notes an information ratio of 0.43 and 1.112 points that the exposure cut did not explain, and records in the minutes that the manager has demonstrated skill. Nobody has made an arithmetical mistake. Every figure in the minutes is correct.
The arithmetic in front of them says the opposite of demonstrated. A ratio of that size needs roughly twenty two comparable years to separate from zero, and that count already assumes a steady ratio and independent years, both of which flatter the case rather than hurting it. The tracking error of 3.7 per cent means a result of minus 1.6 points was an ordinary draw from the same spread. The committee answered the question it wanted answered rather than the one the data could reach. Such a failure is far more common than bad multiplication and much harder to spot in minutes.
The reasoning is symmetric, so the cost lands later and lands hard. The same rule run on a poor year removes somebody for reasons that are equally unsupported, and a process that appoints and removes on one year of noise will do both, repeatedly, at the worst moments. The fix is three lines long. State the sample the conclusion would need before looking at the result. Appraise the constraint record and the stated reasoning where the return record cannot carry the weight. And write down that the evidence does not separate the two possibilities, in those words, whenever it does not.
What must an appraisal never produce?
A grade on a person. An appraisal weighs a body of evidence rather than a person, and a record of one year is evidence about the year rather than about the manager who produced it. The arithmetic reached a boundary and named it, and naming a boundary is a different act from shrugging.
Three things follow. An appraisal that ends in a verdict the evidence cannot support has not been completed, it has been abandoned early and dressed up. An appraisal that refuses to state the support the evidence gives is equally useless. There is always something: the exposure part is settled arithmetic, the ratio is computed, the required sample is computed, and the process record is readable now. A finished appraisal is a statement of exactly what the evidence can and cannot support, written in words that would embarrass nobody if the next year went the other way.
One more absence deserves naming. Everything here is gross of what it costs to run the mandate. Whether the holder kept any of the 1.6 points is a separate question with its own arithmetic, settled elsewhere in this subject area, and no appraisal of judgement answers it. Keeping the two questions apart is part of the discipline.
What is the honest verdict that this one year record actually supports?
References
| Source | Document | Where |
|---|---|---|
| Securities and Exchange Board of India | Any duty on how a managed portfolio's performance is computed, presented or disclosed, named here and not stated | sebi.gov.in |
| Pension Fund Regulatory and Development Authority | The authority where a retirement mandate is the setting, named here and not stated | pfrda.org.in |
| The exchanges | Where index construction methodology is published. The methodology belongs to the index provider rather than to any manager | nseindia.com, bseindia.com |
| William F. Sharpe | The ratio of excess return over the risk-free rate to total volatility, named here and covered in full elsewhere | ideas.repec.org |
| Michael C. Jensen | The residual return against a market model, named here and completed under alpha later in this sequence | ideas.repec.org |
| Gary P. Brinson | The standard cutting of a result into parts, assumed here and covered at the opening of this sequence | ideas.repec.org |
The Anantara Multi-Asset Portfolio, the charitable endowment that holds it, Rukmini Deshpande and Faiz Ahmad Ansari are invented.
Educational material. Not advice on any investment, tax, budget or market position.
