The Halo Effect: One Good Impression Colouring Everything
The halo effect is one judged attribute colouring judgements of every other attribute. An assessor who rates something well on whatever was assessed first will rate it better on everything assessed afterwards, including the things about which there is no information at all. The effect runs in both directions, and the unfavourable version behaves in exactly the same way.
The halo effect is not a strong opinion and it is not enthusiasm. The halo effect is a leak between two judgements that have nothing to do with each other, and the leak happens because the two judgements were made one after another rather than side by side. The halo effect is a fault in the order of judging rather than a fault in the person doing the judging. The order of judging is why the halo survives in careful people, in experienced people, and in people who feel nothing in particular about the thing in front of them.
What is the halo effect, precisely?
Almost nothing is judged in one go. A thing being assessed has parts, and each part is in principle a separate question with a separate answer. Call each of those parts an attributeOne dimension of a thing being assessed, capable in principle of being judged on its own without reference to the others.. The halo effect is what happens when the mark placed against one attribute becomes, quietly and without anybody deciding it should, the mark placed against the rest.
Edward Thorndike set the pattern down in the Journal of Applied Psychology in 1920, in a paper called A Constant Error in Psychological Ratings. He had supervisors rate the people working under them on qualities that were meant to be separate from one another, and found that the ratings agreed far more closely than genuinely separate qualities could possibly agree. Somebody marked highly on the first quality was marked highly on all of them. He called it a constant error because it did not scatter. The error leaned the same way every time, and an error that leans the same way every time is worth naming rather than merely worth apologising for.
Take it out of assessment and into an ordinary afternoon. A household is choosing a caterer for a wedding. The household attends a tasting, one dish is excellent, and by the time they walk back to the car they have formed views on the caterer's punctuality, the staffing, the billing and how the tables will look. Not one of those four was on display at the tasting. The household watched somebody cook, and came away with a rating of somebody's invoicing. The whole mechanism is there, and it took about forty minutes.
The halo effect is one judged attribute colouring the judgement of attributes that were never assessed at all. The word halo carries an unhelpful suggestion of something flattering and mild. The effect is neither. The halo is a measurement problem, and the measurement it damages belongs to the assessor.
What does the halo effect spread from, and what does it spread to?
Why does one impression spread to judgements it has no bearing on?
Because judgements arrive in a queue. Four attributes cannot be assessed simultaneously; one is assessed, then the next, then the next. The assessment orderThe sequence in which the parts of something get judged. It is usually accidental, and it changes the result. is almost never chosen deliberately. The order falls out of whatever was mentioned first, whatever was easiest to look up, or whatever somebody happened to raise.
Once the first judgement exists it does not sit in a sealed box. The first judgement becomes the room the second judgement is made in. Nothing in ordinary assessment requires an independent judgementAn assessment made without knowledge of another judgement that could sway it, whether that other judgement is somebody else's or the assessor's own earlier one. at the second step, and so none is obtained. The result is contaminationAn earlier judgement acting as background context for a later one, so the later one is no longer a separate reading. in the plainest sense: the second reading is taken on an instrument that the first reading has already moved.
A mechanism nobody has measured is only a story, so now put a number on the spread. Take an invented rating scale running from 0 to 100, where 50 is the honest reading for an attribute about which nothing whatever is known. A rating of 50 is honest because it says neither good nor bad, and neither good nor bad is exactly what the evidence supports when there is no evidence. On that illustrative scale, an assessor carrying no impression from anywhere else rates the unknown attribute 50. An assessor carrying a moderately favourable impression from the first attribute rates it 65. An assessor carrying a strong one rates it 80.
Work the spread out rather than accepting it. From 50 to 80 is thirty points of movement, on a scale of a hundred, about an attribute on which not one fact arrived at any point. Thirty points is not a drift and it is not a rounding difference: it is a third of the entire scale, produced by information about something else. The three readings are joined by a straight line, and a straight line is a simplification. The numbers themselves were made up for teaching rather than measured anywhere.
Nothing at all is known about a second attribute. Before the control below is moved: does a strong impression formed on the first attribute change the rating given to the second?
Move the impression, watch a rating with nothing behind it move too
One variable moves: the strength of the impression formed on the first attribute, from none to strong. One consequence follows: the rating given to a second attribute about which nothing at all is known. The three named readings are 50 with no impression, 65 at half strength, and 80 at full strength. The information available about the second attribute is nil at every setting of the control, and that is stipulated rather than assumed.
A moderate impression at a strength of 0.50 lifts the rating of the second attribute from the honest 50.0 to 65.0, a movement of 15.0 points, and not one fact about that second attribute has arrived.
The control, dragged slowly, shows that nothing on the right hand side of the screen is ever fed by anything on the left except an impression. The bar labelled information stays empty at every setting. And yet the rating climbs. Handed over on a sheet of paper, that rating would be read as a finding about the second attribute. A number in that position means a finding, and this rating is not one.
Does the effect run in the unfavourable direction as well?
Yes, identically, and this is the part the name hides. Form a poor impression on the first attribute and every attribute assessed afterwards is marked down by the same distance. On the illustrative scale that is 50 falling to 35 at moderate strength and to 20 at full strength, mirroring the climb exactly. A plumber who arrives forty minutes late is judged worse at pipework, at pricing and at tidying up, none of which anybody has yet seen.
Calling it a halo makes the effect sound flattering, when half of what it does is quietly destroy things that were never examined. There is no separate mechanism for the downward case and no separate name worth learning. The halo is one process with a sign attached, and the sign is set by whichever attribute happened to be assessed first.
Does the effect run in the unfavourable direction too?
How is this different from the affect heuristic, which sounds similar?
Readers merge these two constantly, and the merge costs the reader the ability to spot either one. The affect heuristic, set out by Melissa Finucane and colleagues in the Journal of Behavioral Decision Making in 2000 and covered separately, is a feeling standing in for an analysis. The work has not been done, a feeling is present, and the feeling answers the question the work was going to answer.
The halo effect needs no feeling anywhere. Its input is a completed judgement, arrived at properly, on evidence. The completed judgement then acts as context for the next one. Affect substitutes a feeling for an analysis that never happened. The halo lets an analysis that did happen contaminate a different analysis it has no bearing on. The clearest test separates them cleanly. A person who feels nothing whatever about the thing in front of them, bored by the whole subject, can still show a full halo. A rating on the first attribute is all the mechanism requires.
Somebody has no feeling at all about a holding, finds the whole subject dull, but rated its first attribute highly on the evidence. Halo or affect?
Where does the halo do the most damage in a sequence of judgements?
At the first attribute assessed, and the reason is arithmetic rather than psychology. Whatever goes first is the only judgement made clean, so it is the only one carrying real information, and it then sets the level for every judgement after it. The first attribute is therefore doing several jobs at once while appearing to do one.
Worse, nobody chooses which attribute goes first. The first attribute is chosen by whatever was loudest: a segment on a television channel, a headline, the topic a colleague raised at the start of the meeting. So the attribute that happens to be most visible ends up being the attribute that sets the tone for the ones nobody can see.
The arithmetic works through as follows. Suppose four attributes are to be judged and only one of them has any evidence behind it. The first attribute is rated 80 on the evidence, and 80 is a real finding. The other three have nothing behind them, so the honest reading for each is 50. Averaged honestly, the four come to 57.5. When a strong impression runs through the other three, they read 80 as well, so the average reads 80.0. The overall assessment moves 22.5 points while exactly one quarter of it rests on any information at all.
| Attribute | What arrived about it | Honest rating | Half strength | Strong impression |
|---|---|---|---|---|
| First, the one the evidence covered | a genuine finding, rated on it | 80 | 80 | 80 |
| Second | nothing | 50 | 65 | 80 |
| Third | nothing | 50 | 65 | 80 |
| Fourth | nothing | 50 | 65 | 80 |
| Average across the four | the number that gets carried forward | 57.5 | 68.8 | 80.0 |
The totals stand up to checking rather than having to be taken on trust. Honest: 80 plus 50 plus 50 plus 50 is 230, and 230 divided by 4 is 57.5. At half strength: 80 plus 65 plus 65 plus 65 is 275, and 275 divided by 4 is 68.75. The table shows 68.75 as 68.8. At full strength all four read 80, so the average is 80.0. The movement from 57.5 to 80.0 is 22.5 points, and the movement from 57.5 to 68.75 is 11.25 points.
What did one dated decision actually look like?
Put the mechanism against a record. The Palash decision log is an invented cohortA group of people whose decisions are recorded together over the same stretch of time, so their behaviour can be compared. record of 240 decisions taken by 60 investors over eight quarters, kept by Palash Advisory Services Private Limited, where Devika Rao advises. Meera Sundaram is one of the 60.
On 4 January her holding opened at Rs 12,00,000/-, being four positions of Rs 3,00,000/- each. On the evening of 19 February a television segment named Suvarna Chemicals Limited, and she added Rs 1,00,000/- to that position the same evening. The cost of that one position went from Rs 3,00,000/- to Rs 4,00,000/-, and the cost of the whole holding went from Rs 12,00,000/- to Rs 13,00,000/-. Both sums reconcile in either direction, and the second reconciles against the four positions at 30 September: Rs 3,00,000/- plus Rs 3,00,000/- plus Rs 4,00,000/- plus Rs 3,00,000/- is Rs 13,00,000/-.
Now the point. A television segment lasts a few minutes and can establish something about one aspect of one holding at most. Whatever that segment established, it established nothing about the other aspects. There was no time and no material for it to do so. The impression was formed on one attribute and the money was committed against the whole of it. The gap between what arrived and what was decided about is the halo doing its work inside a dated record rather than inside a rating scale.
The log suggests this is not a peculiarity of one evening. Of the 96 buys recorded across the eight quarters, 41 followed a media mention within three days, or 42.7 per cent. Recompute it: 41 divided by 96 is 0.427, so 42.7 per cent. One case is never evidence that a rule works, and 41 decisions in an invented record are not evidence either. The 41 decisions illustrate what a sequence of judgements looks like when the first one is set by whatever happened to be broadcast.
A segment established something about one aspect of a holding. What did the decision that followed it cover?
Why does adding more assessors not fix it?
Averaging across a group is where most processes think they have already solved the problem. Put six people on it instead of one, average their views, and the individual quirks cancel. Averaging works beautifully when the six are independent. The same averaging does nothing at all when the six are not independent, and in an ordinary process they are not.
Follow what actually happens. Six assessors receive the same material in the same order, so all six form their first impression on the same attribute, from the same input, at the same moment. All six then carry that impression into the attributes about which nothing is known. The six ratings come back tightly clustered, and the cluster is read as confirmation. The cluster is not confirmation. The cluster is apparent agreementAssessors agreeing with one another because they shared an input, rather than because they separately saw the same thing., and it feels exactly the same from outside as the real thing.
Watch the two patterns side by side on the illustrative figures below. Six genuinely separate assessors, none of whom knows anything about the attribute, return ratings scattered from 42 to 78, a spread of 36 points around a mean of 60.0. Six assessors who read in the same order return 63, 64, 65, 65, 66 and 67, a spread of 4 points around a mean of 65.0. The scatter is the honest signal that nothing is known, and the halo replaces it with a tight cluster that looks like knowledge.
So the number that matters is not how many assessors were in the room. The count that matters is how many times the material was looked at without contamination, and each of those is one independent lookOne genuinely separate assessment. A process usually has far fewer of these than it has assessors. the process produced. Six assessors who read in the same order and then discussed it together are much closer to one look repeated six times than to six looks. The two numbers get confused constantly, and the confusion is invisible from inside the room because there really were six people.
The error that gets made, and what it costs
The error is believing that more assessors will average the halo away. More assessors will not. Averaging removes only the variation that differs between the things being averaged, and a shared impression does not differ between them. A shared impression is a common component sitting inside all six ratings, and no amount of adding and dividing removes something every term shares.
Agreement among people who shared an input is not evidence, and it feels exactly like evidence. Feeling like evidence makes it worse than having one assessor. One assessor at least leaves the process aware that it rests on one view. Six agreeing assessors leave it feeling settled. The process has bought confidence and taken delivery of nothing.
The cost is measured in the gap between the two numbers. A process that believes it has six independent looks will accept a conclusion at a strength suited to six, when the material supports the strength of one. The conclusion may well be right. Nobody in the room knows how well supported it is, and everybody thinks they do.
Six assessors read the same material in the same order, discussed it together, and agreed. How many independent looks was that?
What actually breaks the spread?
Independence before discussion, and nothing else reliably does. Every assessor records a rating on every attribute before anybody says a word. Then the discussion happens, with the recorded ratings already on paper and no longer movable. Recording before discussion is the whole correction.
Notice what kind of instruction that is. The instruction is not a request to be more objective, to set the first impression aside, or to be aware of the effect. None of those three work. Awareness of a bias does not remove it, and asking somebody to unknow something is not a procedure. The correction is a scheduling instruction. Changing the order in which two things happen costs nothing, needs no insight from anybody in the room, and can be enforced by whoever is running the meeting without expertise of any kind.
The correction is structural rather than mental, and being structural is exactly why it works. Independence also produces something valuable that the contaminated version destroys: the disagreement. If six people record widely different ratings on an attribute, the process now knows that nothing is known about that attribute. Knowing that is the single most useful thing the process could have learnt. The tight cluster hid precisely that.
What is the correction, stated as a scheduling instruction rather than as an attitude?
When is a first impression legitimately informative?
Often, and this is the part that makes the halo hard to root out rather than merely inconvenient. Attributes frequently do travel together. If somebody has been careless about the one thing that was checked, carelessness elsewhere is genuinely more likely, and treating the first reading as mildly informative about the rest is the correct response rather than a bias. A single correlated attribute is real evidence about its correlated neighbours.
The fault is never that a first impression was used; it is that it was used at the same strength whether the attributes were connected or not. An assessor who moves thirty points on a genuinely correlated attribute may be reading the evidence properly. The same assessor moving thirty points on an attribute with no connection whatever is reading nothing at all, and from the inside the two feel identical. The feeling attached to a spread carries no information about whether the spread was warranted.
Which means the practical question is never whether to let a first impression travel. The question is how far a first impression should be allowed to travel, and the answer depends on a fact about the attributes rather than a fact about the assessor. Asking out loud whether these two attributes actually move together, before the second one is rated, is the only form of the question with an answer.
Is a first impression ever legitimately informative about attributes assessed later?
How does an adviser, or a person deciding alone, actually use this?
What it changes in practice
Devika Rao, advising at Palash Advisory Services Private Limited, does not treat the halo as something to be resisted by willpower. She treats it as a fact about the order papers are read in, and the whole of her response is a rearrangement. Ratings are written down before the meeting rather than in it. The attribute everybody already has an impression about is assessed last rather than first, so it has fewer judgements left to colour. And where two people record very different numbers on the same attribute, that gap is recorded as a finding rather than resolved by whoever speaks with most confidence.
A person deciding alone, with no adviser and no committee, has the harder version of the problem. There is no second person to disagree with, and the sequence of judgements happens inside one head over an evening. The same rearrangement still works. The rating each attribute would receive goes down on paper before anything is looked up, and beside each one goes whatever evidence actually arrived about that specific attribute today. Any attribute where the second column is empty and the first column has moved is the halo, visible on paper, in the assessor's own handwriting.
For a lender, an analyst or anybody assessing on behalf of others, the same discipline shows up as a rule about the order of the file rather than a rule about the reader. The rule delivers the honesty of a process. Any effect on what anybody ends up with is a separate question, and one that eight quarters and sixty invented investors could not settle.
Sources
| Source | Document | Site |
|---|---|---|
| Edward Thorndike | A Constant Error in Psychological Ratings, Journal of Applied Psychology, 1920, the paper in which the effect is first set out | cited to the journal itself |
| Melissa Finucane and colleagues | the paper setting out the affect heuristic, Journal of Behavioral Decision Making, 2000 | ssrn.com |
| Amos Tversky and Daniel Kahneman | Judgment under Uncertainty: Heuristics and Biases, Science, 1974, the paper that opened the heuristics and biases programme | ssrn.com |
| Working paper repositories | a repository of working papers, searchable by author and title | nber.org |
Meera Sundaram, Devika Rao, Palash Advisory Services Private Limited, the Palash decision log, the Vindhya index scheme, the Nilgiri mid-cap scheme, Suvarna Chemicals Limited and Kesari Logistics Limited are invented.
Educational material. Not advice on any investment, tax, budget or market position.
