Covariance: Two Series Moving Together, Measured
A covariance takes two series measured over the same periods and returns the average product of their distances from their own means. On ten paired months of the invented Nakshatra unit and the invented Vasant unit it returns 50.00, and its units are per cent squared. Per cent squared is the whole difficulty: the figure carries a readable sign and a size nobody can interpret alone.
Two columns in, and the covariance built one product at a time.
Every cell below is editable and every figure underneath is worked out again the moment one is changed. Nothing is stored, nothing is fetched, and no reading is filled in. The two columns open holding the ten paired months this guide works through, so the panel starts by reproducing the printed answer exactly.
| Period | Series one | Series two |
|---|---|---|
| 1 | ||
| 2 | ||
| 3 | ||
| 4 | ||
| 5 | ||
| 6 | ||
| 7 | ||
| 8 | ||
| 9 | ||
| 10 |
Series one. The monthly change column of the first unit, read off the sheet it was recorded on, oldest reading at the top. A reading below zero is typed with a leading minus sign.
Series two. The monthly change column of the second unit, from the same sheet, in the same order, oldest at the top.
One row is one period in both columns at once. Row seven of the left column and row seven of the right column must be the same month.
Periods used. How many rows from the top of the two columns the panel reads. Rows below that line are ignored and are not deleted.
Divide the total by. One less than the count is what every figure printed elsewhere in this guide uses. The count itself is the other choice, and the divisor in force is printed on the line where the division happens.
Multiplier on series two. Not a figure off any document. The multiplier restates the second column in a unit that many times smaller, and it is the one control this panel exists for.
| Period | Series one | Its distance | Series two | Its distance | Product |
|---|
Educational illustration built on invented data. Both columns open on readings composed for teaching, and neither names an object a person could go out and buy. Every distance, product, total, covariance and correlation above is worked out in this calculator from the cells as they stand, and every figure is shown to two decimal places. The correlation is built from the same two distance columns as the covariance and from nothing else. A multiplier of zero or below would not be a change of unit at all, so the panel refuses it.
As the panel opens it is reading the ten paired months this guide works through. The ten products add to 450.00, that total is divided by nine because ten paired months leave a divisor of one less than the count, and the covariance comes out at 50.00 per cent squared. Beside it the panel prints a correlation of 0.8694, built from those same two distance columns and from nothing else. Choosing the count itself as the divisor instead turns the same 450.00 into 45.00 per cent squared. The divisor in force is therefore named on the line where the division happens. Everything printed elsewhere in this guide uses one less than the count.
Three things arrive from earlier work and not one of them gets rebuilt here. The first is the Nakshatra unit, a nameless traded unitAny object carrying a price that is allowed to move. It is kept nameless on purpose: who issues it and what it represents alter nothing whatsoever in the arithmetic performed on its readings., and the second is its monthly changeHow far a price travelled across one month, stated as a percentage of where it started. Plus 6.00 says the month finished six per cent up on where it opened, and minus 4.00 says four per cent down., the reading that fills every column below. The third is a pair of habits from earlier: taking the arithmetic meanAdd up all the readings, divide by how many there were. Which centre is the right one to quote from a set of numbers, and when this particular one misleads, is covered separately. of a column, and measuring the spreadHow far the readings of a single series typically sit from their own centre, once the distances have been squared and a root taken at the end. How that gets built, and why the squaring happens at all, is covered separately. of one. Both were built already and neither is explained again. The new thing is that two columns are open at once, and the arithmetic multiplies across them.
A fourth thing arrives with them, and it marks the place where the reasoning here runs out. Somebody wrote the generatorA written rule that manufactures readings: the short list of values allowed to occur, each carrying a weight saying how often it turns up. It is composed first, so what it produces on average is settled in advance. behind the Nakshatra unit before anybody counted a single month out of it. Its centre is therefore a known 1.00 per cent, and 5.00 per cent is a known figure for how far its readings scatter, both of them settled by the way the rule was composed rather than discovered by counting. The entire populationEvery reading the object under study could ever produce, as against the handful actually written down. Which of the two any given figure describes, and how that changes what may be claimed from it, is covered separately. standing behind one of the two columns here is known. Behind the Vasant unit there is no such rule at all. The absence turns out to matter a great deal, and a block further down says precisely how.
The panel above, and everything after it, is one tool: two columns go in and one number comes out, and the entire argument of this guide is about what that number is and is not entitled to say.
What does this panel actually compute?
The recipe is short, and everything else in this guide is a consequence of it. For each period, take how far the first series sat from the first series' own mean. Take how far the second series sat from the second series' own mean. The two distances multiplied together give one product for that period. Across every period the products are added, and the total is divided by one less than the number of periods. The figure that comes out is the covariance.
Notice what the recipe never does. The recipe never sets the first series against the second, but compares each series only with itself and then multiplies the two comparisons. Two series in wildly different units can therefore be fed in without the arithmetic complaining: each column is measured against its own centre before the two ever meet.
The sign of each product is the part that carries meaning, and it follows from ordinary multiplication with no statistics at all. If both series sat above their own means in a month, both distances are positive and the product is positive. If both sat below, both distances are negative, and a negative multiplied by a negative gives a positive, so the product is positive again. Only when one sat above and the other below does the product turn negative. Both above or both below gives a positive product, one above with one below gives a negative one, and the average of those products therefore records whether the two series tend to lean the same way.
The everyday version costs nothing to imagine. A tea stall and an umbrella seller work the same street. On rainy days both take more money than they usually do. On hot dry days both take less than they usually do. Measure each against its own ordinary day and the two distances keep landing with the same sign, so their products keep coming out positive. Now put a cold drinks cart next to the umbrella seller: on rainy days the umbrella seller is above his ordinary day and the cart is below its own, so those products come out negative. Nobody needed a price list to see which pairing leans together and which leans apart.
In month 3 the Nakshatra unit sat below its own mean and the Vasant unit sat above its own mean. What sign does that month's product carry?
Where does every number fed into it come from?
There are exactly two inputs. The first is a column of readings for one series, the second is a column of readings for the other. Both are typed in by whoever is using the panel, both must be the same length, and neither is fetched, looked up or inferred from anywhere.
The first column here holds the monthly change of the Nakshatra unit over ten months, oldest reading at the top. The second holds the monthly change of the Vasant unit over the same ten months, in the same order, from the same record. The row number down the side is not an input; it is a label for the reader, and its only job is to make the third requirement visible.
The third requirement is pairing, and pairing is the input that goes wrong most often. Row seven of the first column and row seven of the second column must describe the same month. Not the same length of month, not the same kind of month: the same month. A covariance computed on two columns that are out of step by a single row is arithmetically perfect and completely meaningless, and nothing anywhere in the arithmetic can reveal that it happened. Every product still multiplies cleanly. The total still adds. The division still divides. The figure that comes out answers a question nobody asked, and it looks exactly like the answer to the one that was.
A column of ten numbers pasted next to a column of eleven is something most tools will notice. Ten pasted next to ten with one of them starting a month late is noticed by nothing. A household comparing its electricity bill against its air conditioner hours has the same problem: with the bill for June lined up against the hours for May a number still comes out, and the number is still wrong, and it does not look wrong.
The second column is pasted in one row lower than the first by accident, so both columns still hold ten readings. Will the panel object, and what is the answer worth?
How is the answer built, line by visible line?
A tool that prints one number and hides the working is a tool to be believed rather than checked. So here is the whole build, and every line of it is reproducible with a pen in about four minutes.
Start with the two means. The Nakshatra column, over these ten months, adds to 10.00 and divides to a mean of exactly 1.00 per cent. The Vasant column adds to 20.00 and divides to a mean of exactly 2.00 per cent. The two means are the only thing the panel derives before it starts multiplying, and it derives them separately, one per column.
Then take the ten pairs of distances. In month 4 the Nakshatra unit read 11.00 per cent, or 10.00 above its own mean of 1.00. In the same month the Vasant unit read 17.00 per cent, or 15.00 above its own mean of 2.00. Multiply and that month contributes 150.00. In month 3 the Nakshatra unit read minus 4.00 per cent, or 5.00 below its mean. The Vasant unit that month read 2.50 per cent, or 0.50 above its mean. One distance below and one above gives minus 2.50, the only month of the ten that pulls the total down.
| Month | Nakshatra | Its distance | Vasant | Its distance | Product |
|---|---|---|---|---|---|
| 1 | 1.00 | 0.00 | 3.00 | 1.00 | 0.00 |
| 2 | 6.00 | 5.00 | 18.50 | 16.50 | 82.50 |
| 3 | minus 4.00 | minus 5.00 | 2.50 | 0.50 | minus 2.50 |
| 4 | 11.00 | 10.00 | 17.00 | 15.00 | 150.00 |
| 5 | 1.00 | 0.00 | minus 3.00 | minus 5.00 | 0.00 |
| 6 | minus 9.00 | minus 10.00 | minus 13.00 | minus 15.00 | 150.00 |
| 7 | 6.00 | 5.00 | 6.50 | 4.50 | 22.50 |
| 8 | 1.00 | 0.00 | minus 1.00 | minus 3.00 | 0.00 |
| 9 | minus 4.00 | minus 5.00 | minus 7.50 | minus 9.50 | 47.50 |
| 10 | 1.00 | 0.00 | minus 3.00 | minus 5.00 | 0.00 |
| Ten months | mean 1.00 | adds to 0.00 | mean 2.00 | adds to 0.00 | 450.00 |
The two distance columns each add to exactly zero. The zero is not a coincidence and is worth using as a check every time: distances measured from a mean must cancel, so a distance column that does not add to zero means the mean above it is wrong. The product column adds to 450.00. Ten paired months leave a denominatorThe number that is divided by. Why a count of readings is sometimes reduced by one before dividing, and what that correction is for, is covered separately. of one less than the count, so divide the total by nine and the answer is 50.00.
Every one of those five stages is printed rather than assumed, and printed working is the only thing that makes a tool worth trusting. Printing them also exposes something a single number would have hidden. Four of the ten months, months 1, 5, 8 and 10, contribute exactly 0.00 to the total. In each of those four the Nakshatra unit read 1.00 per cent, its own mean exactly, so its distance is zero and zero multiplied by anything is zero. The Vasant unit moved in all four of those months, sometimes by 5.00 per cent, and none of that movement reached the answer. Six months out of ten are carrying the whole of this covariance.
Four of the ten months contribute exactly 0.00 to the total of 450.00. What put them there?
Which of the inputs does this panel work out for itself rather than take as given?
What are the units of the answer, and what do they cost?
A per cent multiplied by a per cent is a per cent squared. The squaring is not a convention or a bit of notation; it is what multiplication does to units, the same way that metres times metres gives square metres and rupees times rupees would give a quantity nobody has ever needed. Both distance columns in this guide are held in per cent. Every product is therefore in per cent squared, the total of 450.00 is in per cent squared, and dividing by nine does nothing to that, so the answer of 50.00 is 50.00 per cent squared.
Nobody has an intuition for a squared percentage, and that single fact disables almost everything a reader would want to do with the figure 50.00. The figure cannot be called large or small. Per cent squared and per cent are not measured in the same thing at all, so 50.00 cannot be set beside the Vasant spread of 9.96 per cent to settle which is bigger, any more than a length can be bigger than an area. Nor can 50.00 be set beside a covariance computed on a different pair of series held in different units to settle which pair leans together more. And on its own the figure does not carry enough to be read, so it cannot be passed to somebody who does not already have both columns to hand.
A per cent is different. A ruler for a per cent is already to hand. A month reported at 18.50 per cent conveys roughly what happened, and a spread of 9.96 per cent brings to mind the ordinary month wandering about ten per cent either side. A reading of 50.00 per cent squared conveys nothing, and pretending otherwise is where every misuse of this figure begins.
The answer reads 50.00. In what unit is it held, and which sentence does that unit forbid?
What happens when one of the two columns is rescaled?
Here is the failure this guide exists to show. Every figure in the Vasant column is doubled. The first month goes from 3.00 to 6.00, the second from 18.50 to 37.00, the sixth from minus 13.00 to minus 26.00, and so on down all ten rows. Nothing else changes: not the Nakshatra column, not the pairing, and no month is added or removed.
Ask what has changed about the relationship between the two units. Nothing at all. Every month that was above the Vasant mean is still above it, and every month that was below is still below. The mean itself has doubled to 4.00 per cent, so each month's distance from it has doubled too, and every month sits exactly where it sat relative to the others.
Now ask what has changed about the covariance. Every Vasant distance has doubled, so every product has doubled, so the total has doubled from 450.00 to 900.00, so the answer has doubled from 50.00 to 100.00. The covariance doubled and the relationship did not change at all. The figure is measuring the units of the columns as much as it is measuring anything the columns have in common.
The everyday version is a kitchen scale switched from kilograms to grams. Nothing about what is being weighed changed. Every reading became a thousand times larger. A quantity built by multiplying two of those readings became a million times larger. Nobody would look at the new number and say the vegetables had become heavier. The identical mistake with two columns of returns is made constantly: 50.00 and 100.00 both look like plausible answers, and neither carries its unit on its face.
Settle on an answer before going anywhere near the panel underneath. Every figure in the Vasant column is about to be doubled and nothing else at all is changed. What becomes of the covariance, and what becomes of the relationship?
Stretch the second column and watch a number climb past a picture that never moves.
One control moves: a scale factor applied to every figure in the Vasant column, from 0.5 up to 3.0. The Nakshatra column is nailed down at its published readings at every setting, its mean stays at exactly 1.00 per cent, and the pairing between the two columns is never touched. The middle panel is drawn to the Vasant column's own range, so at every setting it is the same picture stem for stem, and that unchanged picture beside a climbing bar is the entire point of the control. The second view button rescales that middle panel onto one fixed ruler instead, and the stretching has been hiding there all along. At the opening setting of 1.0 the covariance reads 50.00 per cent squared, the published figure exactly.
Educational illustration built on invented data. Both units in this panel were composed for teaching, and neither one names an object a person could go out and buy. Only the Vasant column is rescaled here; the Nakshatra column, the pairing and the row order stay exactly where they were published. The covariance shown uses the nine denominator on ten paired months at every setting, and a scale factor of zero or below is refused because it would not be a change of unit at all.
The covariance moved from 50.00 to 100.00. Has the relationship between the two units strengthened?
How wrong is this answer, given that half the truth is known?
A written rule fixed the Nakshatra unit before anybody drew a month out of it, so a centre of 1.00 per cent, and 5.00 per cent for the scatter around it, are settled quantities here rather than estimates of anything. Any figure taken from a record can therefore be laid beside the quantity it was aiming at.
Do that for these ten months and the result is instructive in both directions. Averaged over these ten months the Nakshatra readings come to exactly 1.00 per cent, landing dead on the settled centre. The landing is the luck of these particular ten months and nothing more: a longer record of the same unit, fifty months of it, is used elsewhere in this material and averages 0.50 per cent, a whisker off half. And even with the centre landing perfectly, the spread of these ten months comes out at 5.77 per cent against a true 5.00 per cent, so the sample missed by more than three quarters of a per cent while the mean was sitting dead on. An estimatorA procedure that squeezes a set of readings down to one figure meant to stand in for something nobody can observe directly. How badly such a procedure can miss, and what shrinks the miss, is covered separately. landing on the truth for one quantity says nothing whatsoever about how it did on another.
There is no written rule behind the Vasant unit. The Vasant unit was made up as a second column and nothing else, so no true covariance exists anywhere to place the 50.00 next to. The comparison every other block of this material makes cannot be made here. A covariance computed on two real series is in that same position and not the other one: there is no true answer sitting behind the record waiting to score it.
One further fact about 50.00 is worth stating because it is exact and because it is where this figure goes next. The Nakshatra variance is 33.3333, that column's 300.00 of squared distances over the same nine. Divide the covariance of 50.0000 by it and the result is exactly 1.5000. The ratio of 1.5000 is the slope that fitting a straight line through these same two series produces. The panel computes the top half of that line and stops there deliberately. Fitting anything, drawing a line through a scatter and predicting one column from the other are covered separately and much later, and not one of them is attempted here.
What can be read off a covariance, and what must be left alone?
A covariance gives the sign of a tendency and nothing reliable about the strength of it. The sign is safe because multiplying either column by any positive number multiplies the covariance by that same positive number, and no positive multiplier can turn a positive answer negative. Whatever the change of units, a positive stays positive. The size is not safe: that same multiplier is sitting inside it, and the size is exactly what moved when nothing real did.
So the reading that is warranted is short. The covariance came out positive, so these two units tend to lean the same way. That is it. Nothing warrants saying they lean together strongly, or more strongly than some other pair, or twice as much as they did last year, and every one of those sentences is written somewhere every day.
The repair is a scale free version of the same idea, in which the units are divided back out so the answer cannot move when a column is rescaled, and it is covered separately rather than here. The point at which the output of this tool stops meaning something arrives early.
Which part of a covariance figure survives a change in the units of either column?
With the curved pair loaded in the panel above, where every reading of the second column is the square of the reading beside it, the covariance comes out at 0.00. What does that nil figure entitle anybody to say?
How is a covariance handed over by somebody else checked?
One of these arrives from somebody else far more often than anybody builds one, so here is the order to check it in, five questions deep. A lender comparing a borrower's monthly collections against a district's rainfall, an analyst handed a covariance between two invented units in a note, a household comparing its grocery bill against the number of guests it fed: all three are in the same position, holding a number somebody else built.
First, are the two columns the same length. If one is longer, some rows were paired with nothing and were quietly dropped, and which ones matters. Second, are they aligned on the same periods. Nothing downstream can answer that one, and asking it out loud is the only way to settle it. Third, what unit is each column held in, per cent or a decimal or something already scaled by somebody upstream. Fourth, what would the answer be if one column were expressed in a different unit. If the reply is that the answer would be unchanged, the reply is wrong, and everything built on top of that figure needs looking at.
Fifth, read the sign, and then stop reading. Whether the answer is 50.00 or 100.00 or 6,400.00 says nothing about how tightly anything moves with anything else until somebody divides the units back out, and until they have, the size is not information at all.
The question is how strongly two series move together, not merely whether they do. What kind of measure is needed instead?
Where this goes wrong, and what it costs
Two people compute the covariance between the same pair of units over the same ten months. One reports 50.00 and the other reports 100.00. Neither has made an arithmetic error, and an audit of both sets of working finds every line correct. One of them held the second column in per cent and the other held it in a unit twice as large, and the covariance carried that difference straight through into the answer without leaving a mark.
The disagreement itself is not the expensive part: a disagreement gets noticed and somebody goes and looks. The expensive case is the one where only one of the two figures was ever produced, so nobody disagrees at all. A note goes out describing a covariance of 100.00 between two units, another note describes a covariance of 50.00 between two others, and a reader who sees both concludes that the first pair leans together twice as hard as the second. The two notes may carry the same relationship written twice in different units, and nothing on the face of either figure would show it.
The repair is exactly as blunt as the failure. The sign of a covariance is read and its size refused. Whenever the size is genuinely the question, the scale free measure built for that job is the one to reach for, and it is covered separately. And a published covariance carries the unit of both columns beside it. A figure that cannot be read without them should never travel without them.
The panel computes and it does not argue. The definition of a spread, and the reason distances get squared on the way to building one, was settled earlier and gets used here with no repetition. Choosing between the centres of a column, and noticing the cases where the mean is a poor one to quote, has its own treatment elsewhere. Any scale free measure of how closely two series move together belongs to the material on relationships between two series and is built there; the panel at the top prints one beside the covariance for a single purpose, to show that it holds still when a column is rescaled. Fitting a straight line between two series, obtaining a slope, and predicting one column from the other are covered separately and much later; the ratio of 1.5000 named above is stated as an exact fact about these ten months and is not fitted, tested or extended to anything. Every reading in the two columns, and in each alternative pair the panel can load, was composed so that a division could be watched happening.
Who vouches for the two columns, and what follows from that?
Nobody vouches for them, and that is the honest answer rather than an evasion. A covariance is arithmetic. Arithmetic needs no institution to be correct, only two columns and a division, so there is no rulebook to cite and no register the answer could be checked against. The table below is kept in the shape the rest of these notes use, and its authority column is empty on purpose rather than by oversight.
| What was leaned on | Where it sits | Site |
|---|---|---|
| Covariance itself, worked from its definition | Standard arithmetic belonging to no single text and in ordinary use for well over a century | No outside site is named |
| The twenty readings in the two columns | Printed in full in the build table above, all ten rows of both | No outside site is named |
| The means, distances, products, total and the scale free figure beside it | Recomputed on opening and again on every keystroke, never transcribed | No outside site is named |
| The settled Nakshatra centre and the settled figure for its scatter | Written into the rule that manufactures the readings, decided before the first month was drawn | No outside site is named |
| The exact ratio of 1.5000 | 50.0000 divided by 33.3333, both printed above | No outside site is named |
The Nakshatra unit and the Vasant unit are invented.
Educational material. Not advice on any investment, tax, budget or market position.
