Spurious Correlation: Patterns With No Mechanism
A spurious correlation is a strong, correctly computed pattern between two series that nothing connects. Across ten invented months, a tally of chairs put out in a community hall tracks the Vasant unit at 0.9502 and accounts for 90.29 per cent of its movement, ahead of the traded input's 75.59 per cent. No sum that can be run on those ten months tells the two apart.
Two things are settled already. A line has already been put through ten pairs of monthly readings and the share of the outcome it accounts for has already been read off. Separately, the notes on what a correlation actually claims set out the readings a large one is consistent with, and one of those readings is the uncomfortable one: the two series have nothing whatever to do with each other, and the pattern belongs to the particular months in hand rather than to anything in the world. The notes on what a correlation claims allowed that reading a single counterexample and named spurious correlation as the place it is worked through. The uncomfortable reading is the one with no arithmetic defence at all, and a reading with no defence deserves more than a footnote.
Three invented objects carry the work, and two of them will be familiar. The Nakshatra unit and the Vasant unit are both a traded unitA vague label for something carrying a market price. Which kind of thing it is never has to be settled. Not one division below would come out differently., and what is written down against each is its monthly change: a price at the close of a month set against the same price at the start of it, expressed as a share of the opening figure. Of the two, the Nakshatra unit sits in the input seat and the Vasant unit is what gets explained. Row by row the two columns cover identical months. Identical months are what make each row a paired observationA row in which two readings belong to one and the same occasion. The shared occasion is the whole licence for comparing them, and a record whose rows have been shuffled has quietly withdrawn it. and what make any comparison between them legitimate at all. Nothing below re-sorts them, and they stay in time orderRows left in the order the months arrived, oldest at the top. Keeping it costs nothing and cannot be undone once a column has been sorted by size instead..
| Month | The Nakshatra unit, per cent | The Vasant unit, per cent |
|---|---|---|
| 1 | 1.00 | 3.00 |
| 2 | 6.00 | 18.50 |
| 3 | minus 4.00 | 2.50 |
| 4 | 11.00 | 17.00 |
| 5 | 1.00 | minus 3.00 |
| 6 | minus 9.00 | minus 13.00 |
| 7 | 6.00 | 6.50 |
| 8 | 1.00 | minus 1.00 |
| 9 | minus 4.00 | minus 7.50 |
| 10 | 1.00 | minus 3.00 |
| Average | 1.00 | 2.00 |
The third object arrives shortly, and it is the one that makes the trouble.
What makes a correlation spurious?
Three things the word does not mean have to be cleared away first. Spurious does not mean weak. Spurious does not mean small. Spurious does not mean miscalculated. The correlation worked here is 0.9502. A figure that high is about as strong as anything an analyst will meet, and it is correct to the last decimal place: anyone who runs the same sums on the same twenty numbers gets the same answer, today and in thirty years.
A correlation is spurious when there is no mechanism: no account, however plain, of how a movement in one series could reach the other.
Spurious is therefore a claim about the world rather than a claim about the number, and that distinction puts it out of reach of every check that can be run on the record. A number can be checked against the record it came from. A mechanism has to be checked against the world, and the record is not the world. The record is twenty figures on a sheet of paper.
The shape is easier to feel at household scale. Try it there. Over six months in one home the number of times the doorbell rang each month climbed steadily, and so did the electricity bill. Line the two columns up and they move together beautifully. Now try to write the sentence that connects them. Not a story about how they might both be caused by something else. A common cause sitting behind both is a different situation with a different name, covered separately. Just the direct sentence: a visitor at the door causes electricity to be consumed, or electricity consumed causes visitors to arrive. Neither sentence survives being said out loud. The pattern is real, the sums are right, and there is nothing there.
Define spurious without leaning on the words weak, small or wrong. Which of these is the working definition?
What sits behind the Kadamba count?
Nothing. The one word is the whole answer, and it has to be accepted completely before the next section starts. The next section is built to shake it.
The Kadamba count is the number of chairs put out each month in a made up community hall that nobody has ever walked into. Somebody unstacks a certain number of chairs before whatever is happening that month and stacks them again afterwards, and a made up register records the figure. Over the ten months it reads 58, 79, 66, 72, 53, 45, 60, 59, 53 and 55, averaging 60 chairs, never dipping below 45 and never passing 79. Each entry is a whole numberA count with no fraction in it, the way 58 chairs is a count. Nothing sits between 58 and 59. Things that are tallied behave that way, and things that are measured do not., because two thirds of a chair cannot be put out.
There is no connection of any kind between that hall and any traded unit, and none is being hinted at. The count is not a disguised version of some well known example half remembered from elsewhere, and it was not picked to resemble one. A reader who starts constructing a story about busy months or seasons is the one supplying that story. The register does not.
Before a single figure in the next section is read: the chair count is about to be fitted against the Vasant unit. What correlation should be expected?
Can a series with nothing behind it beat the real input?
Yes, and on these ten months it does not merely squeak past. The chair count wins on every single measure the arithmetic is capable of producing.
The procedure is the one used on the opening notes, applied without a single change. Take the outcome, take a candidate input, and find the straight line through the ten points that leaves the smallest total when every vertical miss is squared and added up. On the Nakshatra unit that line came out at 0.5000 plus 1.5000 times the input, leaving 218.00 in squared misses. Run the identical procedure with the chair count in the input seat and the line comes out at minus 54.9799 plus 0.9497 times the count, leaving 86.73. The 0.9497 is the coefficientThe number a fitted line multiplies an input by. One step up in the input moves the fitted outcome by that much. The coefficient is a property of the line, not a promise about anything. on the chair count, and it says that the line moves the outcome up by 0.9497 percentage pointsThe unit in which a change to a per cent figure gets counted. A move from 4 per cent up to 7 per cent is three percentage points of movement, and calling that a three per cent rise would be a different and much smaller claim. for every extra chair. At the average of 60 chairs the line reads exactly 2.00 per cent, the average of the outcome. The match is no coincidence: a least squares line always passes through the point where both averages meet.
Here is the whole record, the same ten months as everywhere else and in the same order, with the chair count added as a third column. Twenty numbers became thirty, and nothing else changed.
| Month | The Nakshatra unit, per cent | The Vasant unit, per cent | The Kadamba count, chairs |
|---|---|---|---|
| 1 | 1.00 | 3.00 | 58 |
| 2 | 6.00 | 18.50 | 79 |
| 3 | minus 4.00 | 2.50 | 66 |
| 4 | 11.00 | 17.00 | 72 |
| 5 | 1.00 | minus 3.00 | 53 |
| 6 | minus 9.00 | minus 13.00 | 45 |
| 7 | 6.00 | 6.50 | 60 |
| 8 | 1.00 | minus 1.00 | 59 |
| 9 | minus 4.00 | minus 7.50 | 53 |
| 10 | 1.00 | minus 3.00 | 55 |
| Average | 1.00 | 2.00 | 60 |
The two fits now sit side by side, to be read down slowly. In this worked instance the wrong answer wins every row.
| What was measured | Fitted on the Nakshatra unit | Fitted on the Kadamba count |
|---|---|---|
| The fitted line | 0.5000 plus 1.5000 times the input | minus 54.9799 plus 0.9497 times the count |
| Correlation with the outcome | 0.8694 | 0.9502 |
| R squared | 0.7559 | 0.9029 |
| Explained sum of squares | 675.00 | 806.27 |
| Squared misses left over | 218.00 | 86.73 |
| Root mean squared miss | 4.6690 per cent | 2.9451 per cent |
| Average absolute miss | 3.6000 per cent | 2.5302 per cent |
| Largest single miss | 9.0000 per cent | 5.1980 per cent |
The chairs beat the traded unit on every number the arithmetic can produce. Not on a favourite measure, not on one chosen after the fact: on the correlation, on R squared, on the amount explained, on what is left over, on both averages of the misses, and on the worst month. All of the measures are computed from the same two columns, so a different measure would have gone the same way.
Now the sentence that matters more than any of the figures. Nothing about the chair count has become connected to anything. The hall is exactly as unconnected as it was three paragraphs ago. Not one chair moved. Confidence in the count wobbles somewhere between the bar chart and the table, and that wobble is worth more than the figures that caused it. The arithmetic did not learn anything about the world. A large number simply appeared and started being treated as evidence, and that is what everybody does.
One more line, so nothing is left hanging. The chair count also correlates with the Nakshatra unit, at 0.7145. The 0.7145 is high enough to notice and low enough to rule out the tidy escape route. The chairs are not just a copy of the input wearing a disguise, and the trouble here is not the trouble caused by two inputs repeating each other, a separate subject with a separate name. The chair count is a third column that is genuinely its own thing, genuinely unconnected, and genuinely better at the job.
The chair count reaches an R squared of 0.9029 and the traded input reaches 0.7559. Which is the better model?
Does cutting the record in half catch it?
Splitting the record is the first defence anybody reaches for, and it is a sensible one. If a pattern is an accident, surely it will be an accident of a particular stretch of the record. Cut the ten months into the first five and the last five, measure the correlation separately in each, and a pattern that only lives in one place will show itself by collapsing in the other. Run it and see.
| Half of the record | The chair count against the outcome | The Nakshatra unit against the outcome |
|---|---|---|
| Months 1 to 5 | 0.9308 | 0.7969 |
| Months 6 to 10 | 0.9322 | 0.9863 |
| Distance between the two halves | 0.0014 | 0.1893 |
The check passes cleanly on the series with no mechanism whatever, and it passes it more convincingly than it passes the real input. The chair count comes back at 0.9308 and 0.9322, two readings that agree to within fourteen ten thousandths. The traded input comes back at 0.7969 and 0.9863. The two readings differ by 0.1893, more than a hundred times the gap. On a report that ranked candidates by stability across halves, the chairs would finish first and the input anyone would have guessed at would finish behind them.
So what does the split actually establish? Something narrow and genuinely worth having: that the pattern is not confined to one stretch of the record. The split rules out the case where a single dramatic month, or one unusual quarter, carried the whole relationship and the rest of the record contributed nothing. Such a case exists and the check finds it, so the habit is a good one. The check cannot do anything else. A check that a spurious series passes is not a check for spuriousness, and no amount of enthusiasm for the check changes that.
Consider a pot of dal tasted from the top and again from the bottom. Both spoonfuls taste the same, and that establishes something real: the salt is evenly distributed and nobody stirred badly. The even taste establishes nothing at all about whether this is the dish that was ordered. Two consistent tastings of the wrong dish are still the wrong dish, consistently.
The split passes at 0.9308 and 0.9322 on a series with nothing behind it. What has the check actually proved?
Is there any arithmetic that separates the two cases?
No. Not this arithmetic, not better arithmetic, not arithmetic somebody will invent next year. The reason is structural and it takes one paragraph to see, after which it cannot be unseen.
Every quantity in this guide was computed from two columns of figures. The correlation, the coefficient, R squared, the squared misses, the spreadA one word summary of how widely a column of figures is scattered about its own centre. Two columns can share a centre exactly and still differ completely in this. of each column, the two half record readings: all of them are functions of twenty numbers and nothing else. Now look at what the chair count hands over and what the Nakshatra unit hands over. Both hand over a column of ten figures. Both columns have a centre, a spread and an order. Neither column carries a tag saying what produced it. Which column came from a price and which came from a stack of chairs was never inside either column to begin with, so once the numbers are written down the arithmetic cannot tell them apart.
None of this is a complaint about these particular measures. The correlation is not crude in a way that something more sophisticated would fix. Any quantity that can be defined is defined on the columns, so any quantity that can be defined is blind in exactly the same place. Asking for a cleverer statistic to detect spuriousness is like asking for a more accurate thermometer to settle whether a room is beautiful. The instrument is fine. The thermometer answers a question about temperature, and beauty was never in the reading.
Name the one thing that differs between the two candidates. Where does it live?
Somebody proposes a rule: any correlation above 0.90 should be treated as a real relationship. What is wrong with it?
What separates them, if no number can?
Four things do the separating, and the striking feature of the list is that not one item on it is a quantity that can be computed from the record. The absence is not an accident. If any of them could be computed from the record, the section above would be wrong.
The first is a mechanism, written down as a sentence before the fit is run. Not after, and the ordering is the entire point. Written afterwards, a mechanism is a story assembled to fit a number already seen, and a competent person can assemble one for any pair of columns in about ninety seconds. Written beforehand it is a commitment that the record can then embarrass. A street vendor who says on Monday that fewer people stop at the cart when it rains has made a claim that Tuesday can contradict. A vendor who looks at a slow week and then explains it by the rain has made no claim at all.
The second is a prediction the mechanism makes that the pattern on its own does not. This is the sharp one and it does real work. If chairs reached prices somehow, that account would have to say something further: something about what happens when far more chairs are put out than usual, or about how long the effect takes to arrive, or about what should happen in a hall that closed for two months. Each of those is a separate thing that can be gone out and checked. The bare pattern makes no such commitment, and the missing commitment is precisely what is comfortable and useless about it. A mechanism that predicts nothing beyond the pattern it was invented to explain has added nothing to the pattern.
The third is fresh months: readings the relationship was not chosen on. The chair count was selected because it fit these ten months well. Testing it on these ten months is asking it to repeat the one thing it was picked for. The point is not a subtle one, but it is skipped constantly. The months are already in the file, and fresh ones require waiting.
The fourth is the count of how many candidates were tried, and it is the one that gets left out. A correlation of 0.9502 found on the first series anybody looked at is one kind of finding. The same 0.9502, found by running two hundred series against the outcome and keeping the best, is a different kind of finding entirely, and the two are indistinguishable once written up. The pattern is identical. The difference is how many chances the pattern had, and that number lives in the analyst's process, not in the data. The question is not how strong the pattern is but how many chances it had.
Two people report a correlation of 0.9502 against the same outcome. One tested a single series chosen in advance. The other tested two hundred and reported the best. Should the two reports be read the same way?
Why does adding more months not fix this?
Because more months improve a different property from the one that is broken. Suppose the pattern held at 0.9502 over a hundred and twenty months instead of ten, and then over a thousand. The interval around the estimate tightens: on ten months an interval built around 0.9502 in the usual way runs roughly from 0.80 to 0.99, a width of about 0.19; at a hundred and twenty months it is about 0.04 wide; at a thousand it is about 0.01 wide. The sampleThe part of a record actually in hand, as against everything that could have been written down. More of it steadies an estimate without turning it into an estimate of something else. grows, the estimate steadies, and that is worth having. Notice what kind of gain it is: a gain in precision, and only that.
A spurious relationship measured over a thousand months is a spurious relationship measured precisely, and precision and correctness are different properties. Reading the correlation to a fourth decimal place says how firmly the figure is known. The fourth decimal says nothing at all about whether the figure is measuring a relationship. Nobody has written anything in the empty box beside those shrinking intervals, and nothing done to the columns writes anything in it. The box is the same size at ten months and at a thousand.
Counting the chairs for ten more years yields a very well characterised chair count: its average known to a fine tolerance, its spread known, its month to month behaviour known. Counting something carefully is not the same activity as finding out what it touches, so the count will be no more connected to a traded unit than it was on the first day.
A thousand more months arrive and the correlation still reads 0.9502. What has changed about the problem?
What is written down when no mechanism can be named?
An analyst handing a finding to somebody who will act on it lives in exactly this situation. A mechanism cannot always be named. The situation is common and it is not by itself a disgrace. The write up decides whether the work is honest, and there are four lines to add. Each costs a sentence.
The mechanism goes down as a sentence before anything is run, and if it cannot be written, the file records that it could not. A blank field that somebody left blank on purpose is information. A blank field nobody thought about is a hole with the same shape, and by the time the finding reaches a reader the two look identical. The person who reads the work in six months cannot tell which one they are holding unless the write up says.
Second, state how many series were considered. Two hundred candidates screened and one reported is a perfectly reasonable thing to have done; it becomes misleading only when the two hundred goes unmentioned and the reader takes the survivor for a first guess that happened to land. Third, say plainly whether the relationship was chosen on the same months it is being reported on. If it was, the report is describing a fit rather than a test, and those are different claims.
A pattern is reported as a pattern rather than as a relationship. The distinction costs one word. The word changes what a reader is entitled to do with the finding, and it is the whole of the honesty available once the mechanism field is empty.
Notice what none of this does. None of it makes the finding safe, and saying so is itself the finding. The four lines do not upgrade a pattern into a relationship; they describe accurately what is being handed over, so that the person receiving it can decide how much weight it will bear. A household choosing where its one salary goes, a lender deciding what to lend against, an investor sizing a holding: each of them is entitled to know whether the thing in front of them survived a test or merely fit a file. Writing that down is not modesty. Writing it down is the only part of the job that a stronger correlation could never have done.
The fit is excellent and no mechanism can be named. Which sentence belongs in the report?
The screen that came back clean, and what it cost
An analyst is asked which series moves with the Vasant unit. There is a shared drive with a couple of hundred series on it, so she correlates every one of them against the outcome, sorts the list from high to low, and keeps the top row. The top row is the chair count at 0.9502, with an R squared of 0.9029, comfortably ahead of the traded input anybody in the room would have nominated. She is careful, so she cuts the record in half and checks: 0.9308 and 0.9322. The correlation holds in both halves. She writes it up.
Every step there was competent, and not one figure in the report is wrong. The report is worthless. The reason is in none of the numbers, and that is why nobody catches it in review: no account of how a chair could reach a price was ever written down, and the number two hundred never left her own machine.
The cost is specific and it repeats. The relationship holds on exactly the months it was picked on and nowhere else, so the first month it is used on is the first month it has ever actually been asked to do anything.
The fix is an order of work rather than an extra test, and it is free. Write the mechanism as a sentence before the sums are run. Record how many candidates were considered, in the file, where a reader can see it. And keep calling it a pattern until months nobody chose it on have had a look at it.
What was consulted
| What a reader might expect to find named | What is actually behind the figures |
|---|---|
| A regulator, a supervisor or a trading venue | None, and the emptiness is a finding rather than a gap in the reading. Fitting a line through ten pairs of numbers is arithmetic. It carries no limit, no threshold and no filing period from any market, so a named authority here would be decoration. |
| A dated price record behind the ten months | None. The ten months were constructed so that the sums close exactly, and constructed months carry no as-of date and no venue. |
| A well known published case of a pattern with nothing behind it | None used. The famous ones are somebody's own published work, and a chair count answers the question just as well without borrowing anybody's. |
| The arithmetic behind every number above | Least squares on ten paired rows, and nothing further. Every quantity follows from the three columns printed in the tables, so anybody willing to run the same sums on a sheet of paper reproduces the lot. |
The Nakshatra unit, the Vasant unit, the Kadamba count and the community hall are invented.
Educational material. Not advice on any investment, tax, budget or market position.
