Investor Sentiment: Measuring Mood at Market Level
Investor sentiment is the aggregate mood of the people trading, and a sentiment indicator is a composite: several measurable quantities, chosen by whoever built it, combined by weights they selected and scaled against a window they picked. All three decisions are invisible in the output. A level therefore carries no natural units, and only its movement means anything at all.
A headline saying that mood is at a two year high, or that fear has taken over, reads like a temperature. Somebody put a thermometer into the market, and out came a number. A mood figure is not a reading off an instrument at all. The figure is something built, by a person, out of parts they chose. Turnover is measured; sentiment is constructed from measurements, and a construction inherits every decision made while building it. The distinction between a measurement and a construction decides how much weight any such figure can carry, and it is why two honest builders can look at the same market on the same afternoon and publish different numbers without either of them being wrong.
Investor Sentiment and the Sentiment Indicator: what is each one made of?
Investor sentimentThe aggregate mood of the people taking part in a market, treated as one quantity rather than as many separate feelings. is a claim about a crowd, not about a person. Meera Sundaram, the invented salaried investor whose decisions run through these notes, was cheerful on 19 February when a television segment named Suvarna Chemicals Limited and she added Rs 1,00,000/- to it the same evening. Meera is one person on one evening. Sentiment is the attempt to say something of that shape about sixty thousand people at once, or six million, when nobody has asked any of them how they feel.
Since they cannot be asked, something they do is measured instead. A measured stand-in of that kind is a proxyA measurable quantity used to stand in for something that cannot be measured directly.: a quantity that can be counted and that plausibly moves when mood moves. Think of a school trying to measure how much its pupils are enjoying a term. The school cannot open their heads. So it counts attendance, library borrowings and how many stay behind after the last bell. Each count is real. None of them is enjoyment. Every one of them is a stand-in that somebody selected because they believed it moved the right way, and a different teacher would have selected different counts.
A sentiment indicatorA single figure assembled from several chosen proxies, weighted and scaled by whoever built it. does exactly this with market quantities. Baker and Wurgler, in Investor Sentiment and the Cross-Section of Stock Returns in the Journal of Finance in 2006, set out the approach that most later work follows: take several proxies believed to move with mood, strip out what plainly belongs to underlying conditions rather than to feeling, and combine what is left into one series. The important word in that description is combine. The output is a compositeOne number assembled from several, using weights that somebody chose rather than weights that were discovered., and a composite is a made thing.
What is a sentiment indicator built from?
What three choices sit inside every indicator, and why does none of them show?
Strip the construction back and three decisions are doing all the work. Which proxies go in. How much each one weighs. And what span the whole thing is scaled against, a step called standardisationExpressing a figure relative to the typical value and the spread of a chosen stretch of history, so that it becomes a comparison rather than a raw count.. Change any one of the three and the same raw counts produce a different published figure.
The third choice is the one readers underrate, so it is worth slowing down on. Standardisation asks: compared to what? A count of 3.1 means nothing until it is stated as 3.1 of what, measured against which stretch of time. Scaled against the last two years, a figure looks extreme. Scaled against the last twenty, the identical count looks ordinary. Neither calculation is dishonest. The two calculations answer two different questions, and the published number does not say which question was asked.
The household version runs like this. A monthly electricity bill comes to Rs 4,200/-. Is that high? Against last month, maybe. Against the same month last year, maybe not. Against the twelve month average of the whole street, something else again. The bill is a fact and the word high is a comparison, and the comparison needs a window that somebody has to pick. Every published sentiment reading has had that pick made already, silently, before the number reached the reader.
What are Market Fundamentals, and what does the word not cover?
Market fundamentalsThe underlying facts a price is supposed to be about: what is earned, what is held, what is owed, and the conditions surrounding all of it. is a category word rather than a calculation. The category covers the underlying facts a price is supposed to be about: what a business earns, what it holds, what it owes, what it pays out, and the wider conditions those sit inside. Computing any of that is valuation. The argument below works perfectly well with fundamentals kept as a category.
The exclusions matter just as much as what the word covers. Fundamentals are not the price. They are not the forecast. And they are emphatically not what people think about them. A shop on a busy corner has fundamentals: its takings, its rent, its stock, the road works starting next month. Whatever the neighbours believe about that shop is a separate fact about the neighbours. The facts and the feelings about the facts are two different objects, and a great deal of confused market commentary consists of quietly swapping one for the other.
Keeping fundamentals as a category also guards against a trap. Because nobody can state the full set of underlying facts, nobody can state exactly what a price should have been. The impossibility of stating the full set is the load-bearing fact that the inefficiency section below turns on.
Investor Sentiment vs Market Fundamentals: which question does each one answer?
Set the two side by side and the confusion clears quickly. Each answers a different question. Fundamentals answer: what are the underlying facts? Sentiment answers: how do the participants feel about them? Both are legitimate questions and neither one is a substitute for the other.
Consider a wedding hall booked eighteen months ahead. The fundamentals are the hall: its capacity, its rate, its kitchen, whether the road to it floods in July. The sentiment is how the two households feel about the booking. Their feeling may swing from delight to panic while the hall itself does not change by a brick. A question about how the wedding is going, answered with the news that everyone is anxious, teaches something real and still says nothing about the hall.
Sentiment describes the people; fundamentals describe the thing. The reason this matters for measurement is that the two get mixed inside a real indicator. Turnover rises when people are excited, and it also rises when genuinely new information arrives and holdings need repricing. Any proxy chosen carries some of both. Serious construction therefore tries to strip the conditions out first. Whether that stripping worked is another decision a reader cannot see.
What question does each of the two answer?
Why does a level mean nothing on its own?
A temperature of 40 degrees means something because degrees are fixed. Everybody using the scale gets the same number for the same room, and two people in two cities can compare readings without discussing method. Sentiment has nothing like that. A reading of 2.0 is 2.0 of whatever the builder's window and weights made it, and a second builder producing 2.0 from different parts has produced a coincidence of digits.
The absence of a fixed scale is why one indicator cannot be called twice another, why a reading cannot be said to have doubled in any meaningful sense, and why today's 2.0 is not comparable to a 2.0 published elsewhere. There is no arithmetic connecting them. The one comparison that survives is variation within one consistently built series: this indicator, built the same way throughout, is higher now than it was, and that comparison is legitimate because the method is held still. Move to a second series and the comparison collapses.
How far does the impression move when only the encoding changes?
Now to the measurements themselves. The Palash decision log, an invented record of 240 decisions taken by 60 investors over eight quarters, sits alongside an invented illustrative index called the Palash 100. Across those eight quarters the index opened at 100.0, peaked at 131.0 in the second quarter and reached its low of 104.0 in the fourth, a fall of 20.6 per cent from peak to trough. And turnover, the one quantity in the log recorded quarter by quarter, ran 3.1 times its eight quarter medianThe middle value of a set once the values are lined up in order, so that half sit above it and half below. in the second quarter and 0.4 times in the fourth. Activity peaked near the top and died near the bottom.
Two facts, then: 3.1 and 0.4. Now watch what happens when nothing changes except the words used to state them.
Stated as a multiple of the median, they are 3.1 and 0.4, exactly as the log itself records them. Stated as a deviation from the median, they are 2.1 above and 0.6 below. The subtraction runs 3.1 less 1.0 is 2.1, and 1.0 less 0.4 is 0.6. Stated as a ratio between the two quarters, they are 3.1 divided by 0.4. The division comes to exactly 7.75, or 7.8 rounded to one decimal place. Stated as a rank among the eight quarters, they are the highest and the lowest.
Read those four sentences back and feel what each one does. Seven point eight times sounds extraordinary. Two point one above the median sounds moderate, almost technical. Highest of eight sounds close to trivial. In any eight quarters some quarter has to be the highest. Four statements, all of them exactly correct, all of them derived from the same two measurements, and four completely different impressions. Nothing in the data changed between those sentences. The only thing that moved was the encoding, and the impression is what a reader carries away.
The pair 3.1 and 0.4 times median can also be stated as a ratio of 7.8 times. Which of the two statements is more accurate?
Before the encoding below is switched: do the underlying turnover figures change as the control moves?
Switch the encoding and watch the picture move while the data sits still
One variable moves: how the two measurements are expressed. The measurements themselves are fixed at 3.1 and 0.4 times the eight quarter median, exactly as the log records them, and every figure below is derived from those two alone. As multiples they read 3.1 and 0.4. As deviations, 3.1 less 1.0 is 2.1 above and 1.0 less 0.4 is 0.6 below. As a ratio, 3.1 divided by 0.4 is exactly 7.75, shown as 7.8 to one decimal place. As ranks among the eight quarters, they are the highest and the lowest. The top panel never redraws; the bottom one redraws every time, and the scale beneath it changes with it. The log holds one proxy at two quarters, so it cannot support a sentiment indicator and none is built from it.
Stated as a multiple of the eight quarter median, the second quarter reads 3.1 and the fourth reads 0.4, which is exactly how the log records them.
What is Market Inefficiency, and does a sentiment reading demonstrate one?
Market inefficiencyA departure of price from what a stated set of information implies it should be. means a price has departed from what a stated information set implies it ought to be. Eugene Fama, in Efficient Capital Markets in the Journal of Finance in 1970, set out the framework in which that sentence has meaning, building on the argument by Paul Samuelson in Proof That Properly Anticipated Prices Fluctuate Randomly in Industrial Management Review in 1965 that anticipated prices should move unpredictably. Read the definition slowly and the problem announces itself. Inefficiency is defined against what the price ought to be. Somebody has to supply that.
The obstacle is the joint hypothesis problemAny test of whether a price is wrong is simultaneously a test of the model used to say what the right price was, so a failure cannot be assigned to one of the two. in a third coat of paint. Claiming a market is mispriced because a mood reading is high requires a statement of what the price should have been. Such a statement is a model of the underlying facts. So the claim rests on the model. If the model is wrong, the mispricing sits in the arithmetic that produced it. The mood figure did no work at all.
And there is a plainer objection sitting underneath. A high reading is entirely consistent with an efficient market in which participants are cheerful for perfectly good reasons. If the underlying conditions genuinely improved, people would feel better, turnover would rise, and every proxy in the composite would climb. A market behaving that way is working, not failing. The indicator was never built to tell the two apart, so it produces the same figure either way.
A published sentiment reading is high. Does that show the market is mispriced?
How to Interpret Investor-Sentiment Data Without Overclaiming: what is the fixed order?
The four questions below are a procedure, and in a procedure the order is the point. Ask these in sequence and stop where the answers run out. Skipping to the last step is exactly the habit that produces overclaiming.
First, name the proxies that went in. Until they can be named, what is being read is not a mood figure but a digit. A composite built from turnover behaves differently from one built from new account openings, and the difference is not a detail. Second, whose weights, and why those. One proxy carrying most of the weight makes the composite a repackaging of that proxy. Repackaging may be fine, so long as somebody says so. Third, what window it was scaled against, and crucially when that window was chosen. The person choosing a window after the data was visible already knew which answer each choice would produce, so a window selected that way is not an independent test of anything. Fourth and only then, the variation rather than the level is what can be read: this series moved, relative to itself, built the same way throughout. Nothing beyond that follows.
What is the first question to ask of any sentiment reading?
What can the Palash log supply, and what exactly is missing?
Showing the gap teaches the construction better than filling it with invention would, so precision is needed now about what the case at hand can and cannot do. The log records turnover against its median at two quarters: 3.1 times in the second, 0.4 times in the fourth. One proxy at two of eight quarters is the whole of it.
The other candidates look promising until they are checked. Of the 96 buys in the log, 41 followed a media mention within three days, or 42.7 per cent, against 11.0 per cent of the eligible list being mentioned at all in a given week. The 42.7 per cent is a striking figure and a whole period total. A written reason was recorded on 84 of the 240 decisions, being 35.0 per cent, and 71 of 240 were taken within 48 hours of a news item, being 29.6 per cent. Both are whole period totals as well. There is no quarterly breakdown of any of them anywhere in the log, so not one can be placed on a timeline, and a series needs points in time.
So the honest statement is short. An indicator needs several proxies at every quarter. The Palash log holds one proxy at two quarters and three totals that cannot be split. The Palash log cannot supply a sentiment index. Inventing the missing quarters would produce a figure that looked exactly like a real one. A figure that looks real and rests on nothing is the failure itself.
Why can the Palash log not supply a sentiment index?
The two errors that get made, and what each one costs
The first error is treating a sentiment reading as evidence of market inefficiency. A sentiment reading is not evidence of inefficiency, for the reason set out above: the claim needs a statement of what the price should have been, that statement is a model of the underlying facts, and so the weight rests entirely on the model rather than on the mood figure. The error costs a false sense of having found something. The market has not been tested. The model has been tested, and its result reported as though it were a fact about other people.
The second error is commoner and quieter: reading a level as though it had units. An indicator scaled against one window and the same indicator scaled against another produce different figures from identical raw counts. A reading of 2.0 is therefore not twice a reading of 1.0, and it is not comparable to a 2.0 from a different builder at all. The four encodings above make the point without needing any indicator: four exactly correct statements of two measurements, four different impressions, and the impression is what survives into somebody's decision.
Both errors share a root. Both treat a constructed figure as though it were an observed one. The cure is not scepticism about sentiment work, careful and useful as that work is. The cure is asking the three construction questions before reading the number at all.
How does a practitioner read a mood figure without acting on it?
Devika Rao, the invented adviser at Palash Advisory Services Private Limited, does look at mood measures. She just does not use them the way the headlines invite. She uses them for conversation timing and expectation setting, a use that survives every objection raised above.
Here is the concrete version. When activity across her clients runs far above its usual level, as it did in the second quarter at 3.1 times the median, she reads that as a signal about her own work rather than about prices: more decisions are being taken quickly, more of them will be taken within 48 hours of a news item, and more of them will arrive without a written reason. The log shows exactly that across the whole period, with a reason recorded on only 84 of 240 decisions. So she puts more time into the checklist conversation in a busy quarter. Twenty of the sixty adopted a written checklist on 4 November, and across quarters five to eight they recorded a reason on 34 of 41 decisions, or 82.9 per cent, against 19 of 63, or 30.2 per cent, for the other forty. Eight quarters and sixty people cannot carry a claim about returns, so none follows from the difference.
For a person deciding alone, with no adviser and no committee, the same reading works inwards. If the news feels loud and the decisions are coming faster than usual, that is a private turnover proxy saying something about the process. The legitimate use of a mood reading is as a prompt to slow a decision down, never as a prompt to change a holding. The two uses feel similar and they are not remotely the same thing.
Why is a mood reading never a position?
Now the sentence that must never follow a mood reading: this reading is high, therefore act on it. A behavioural account of why a price moved is an explanation. An explanation is never an instruction to buy, sell, hold, wait or avoid, and never evidence that acting on it would have paid. Four separate reasons hold that line, and any one of them is sufficient on its own.
First, any measured relationship between mood and later returns is almost always measured before costs. Look at what costs did inside this very log. The five turnover groups of twelve investors each ran annual turnover of 9, 34, 71, 128 and 210 per cent. Their gross returns were 11.2, 11.0, 11.1, 10.9 and 11.0 per cent, all inside 0.3 points of each other. Their net returns were 10.9, 10.4, 9.6, 8.4 and 6.9 per cent, running 4.0 points apart. The difference between the groups is what the trading cost, not what was picked, and 4.0 points a year is larger than most documented effects of this kind.
Second, a published effect has been read by everyone who read the paper, so what it did before publication is not what it does afterwards. Third, the same obstacles that let a mispricing persist are the obstacles that stop a reader capturing it. Andrei Shleifer and Robert Vishny set that out in The Limits of Arbitrage in the Journal of Finance in 1997, and the explanation and the obstacle turn out to be one fact seen twice. Fourth, and specific to sentiment: since a level has no natural units, a current reading cannot even be established as high without choosing a comparison window, and choosing that window after seeing the data settles the answer before the test begins.
Why can a reader not establish, on their own, that a current sentiment reading is high?
Sources
| Source | Document | Site |
|---|---|---|
| Malcolm Baker and Jeffrey Wurgler | Investor Sentiment and the Cross-Section of Stock Returns, Journal of Finance, 2006 | ssrn.com |
| Paul Samuelson | Proof That Properly Anticipated Prices Fluctuate Randomly, Industrial Management Review, 1965 | ssrn.com |
| Eugene Fama | Efficient Capital Markets, Journal of Finance, 1970 | nber.org |
| Andrei Shleifer and Robert Vishny | The Limits of Arbitrage, Journal of Finance, 1997 | ssrn.com |
| Securities and Exchange Board of India | conduct, suitability and disclosure requirements applying to registered intermediaries | sebi.gov.in |
| Association of Mutual Funds in India | investor facing practice material for distributors and advisers | amfiindia.com |
| International Organization of Securities Commissions | principles for the conduct of business with retail customers | iosco.org |
Meera Sundaram, Devika Rao, Palash Advisory Services Private Limited, the Palash decision log, the Palash 100 index and Suvarna Chemicals Limited are invented.
Educational material. Not advice on any investment, tax, budget or market position.
