Variance, Standard Deviation and Range Compared
The range is the largest month minus the smallest, 20.00 per cent here, and it is built on two of the fifty months while the other forty eight cannot touch it. Variance averages the squared distances from the mean and comes to 24.74, in per cent squared. The standard deviation is its square root, 4.97 per cent, back in the units of the record.
Two bits of groundwork are already in place, and neither is rebuilt here. The invented Nakshatra unit and its fifty month record are established separately: a traded unitA thing with a market price attached to it, where that price is not fixed. Who issues it and what it is made of do not enter; only the price moving does. whose monthly changeThe percentage difference between one month's price and the previous month's, so a run of prices becomes a run of changes that can be added up and compared. was counted over fifty months. The tallyA count of how many times each distinct value turned up. Fifty months collapse into five lines when only five values ever occurred, and nothing is lost in the collapse. came out as five months at minus 9.00 per cent, nine at minus 4.00, twenty five at 1.00, eight at 6.00 and three at 11.00. The arithmetic mean of that record, worked out separately, came to 0.50 per cent. Every distance here is measured from that 0.50 per cent and from nothing else.
One more thing comes with that groundwork, and it is unusual enough to say out loud rather than tuck into a footnote. The generatorThe written recipe behind the numbers: the list of readings it can hand out, and the weight on each, all fixed on paper before anything came out of it. Both its centre and its spread fall out of the recipe, so neither had to be estimated. behind the fifty month record was fixed on paper before any month came out of it. Its true spread is therefore known to be 5.00 per cent exactly, as a matter of construction rather than measurement. Almost no record ever met in practice arrives that way. A known true spread lets the record's answer and the true answer be set side by side, and the distance between them read off exactly.
Why is a centre on its own only half a sentence?
Two records can report the identical mean and be nothing alike. Take the fifty month record, whose months run all the way from minus 9.00 per cent to 11.00 per cent. Now take a second invented record, also of fifty months, in which twenty five months came in at 0.00 per cent and twenty five at 1.00 per cent. Both records have a mean of exactly 0.50 per cent. Anyone quoting the mean alone has described the two of them in identical words, when one never moved more than a single percentage point and the other had a month at minus 9.00.
The everyday version is easier to feel than to argue. Two households each spend an average of Rs 40,000/- a month. In the first, every month lands between Rs 38,000/- and Rs 42,000/-, and the household can plan. In the second, eleven quiet months at Rs 25,000/- are followed by one month with a wedding in it. The average is the same number and it says almost nothing about what it is like to live in either house. The distance between the months and their own centre is what one number has to capture. A centre quoted with no spread beside it is half a sentence, and every measure that follows is an attempt to put one number on that distance.
What is the range, and how many of the fifty months actually build it?
The range is the largest value in a record minus the smallest. The subtraction is the entire definition and nothing hides inside it. On the fifty month record the largest month came in at 11.00 per cent and the smallest at minus 9.00 per cent, so the range is 11.00 less minus 9.00, or 20.00 per cent. One subtraction, and it is done.
Now count what actually went into that answer. Two months. The range of 20.00 per cent is built on two of the fifty months, and the other forty eight cannot touch it. Those forty eight can be moved anywhere at all between minus 9.00 and 11.00 per cent and the range does not shift by a whisker. Clustered on one value, spread out evenly, reordered: the answer is 20.00 per cent every time. Forty eight months of information sit in that record and the range refuses to look at any of them.
The blindness is a genuine weakness, and the range is nonetheless good for one job. The range is a fast check that nothing absurd is sitting in the record. A record of monthly changes whose range comes back at 400 per cent has a typing mistake in it somewhere, and one subtraction found that out rather than a long stretch of arithmetic. The range is a screening tool, not a description. Every extra month collected is one more chance to find a new extreme, so the range is the only measure here that gets worse as more data arrives, and a record of five hundred months will almost always report a wider range than the same process measured over fifty.
The largest month in the record is 11.00 per cent and the smallest is minus 9.00 per cent. What is the range, and how many of the fifty months went into producing it?
What is variance, and why is every distance squared before it is averaged?
Variance takes the opposite approach to the range: it insists on looking at every month. For each month, take its distance from the mean of 0.50 per cent. Because only five distinct values ever occurred, there are only five distinct distances: minus 9.50, minus 4.50, 0.50, 5.50 and 10.50 percentage points.
The obvious next step is to average those distances, and the obvious next step fails completely. Add them up with their counts, five of the first, nine of the second, twenty five of the third, eight of the fourth and three of the fifth, and the total is exactly zero. Not roughly zero. Exactly. A mean is the balance point where the distances cancel, so the distances below the mean cancel the distances above it by construction. Averaging them measures nothing at all, and it yields the same nothing on any record ever collected, including the two records in the figure above that are so obviously different.
So each distance is squared before anything else happens to it. Squaring turns every distance positive, so nothing can cancel, and the five squares come out as 90.25, 20.25, 0.25, 30.25 and 110.25. Weight each square by how many months carried it, 5, 9, 25, 8 and 3, and the weighted squares are 451.25, 182.25, 6.25, 242.00 and 330.75, adding to 1,212.50. Divide that total by 49 and the variance is 24.74.
Look at those five weighted squares once more. They say something the range could never say. The eight months sitting at the two extreme values, five at minus 9.00 per cent and three at 11.00 per cent, carry 782.00 of the 1,212.50 total between them, or 64.49 per cent of it. The twenty five months at 1.00 per cent, fully half the record, carry 6.25 between them, a bare 0.52 per cent of the total. Variance reads every month, but it does not weigh them equally: squaring means a month twice as far from the centre counts four times as heavily.
Why is each month's distance from the mean squared before the distances are averaged?
Variance vs Standard Deviation: what do the units rule out?
The variance of the fifty month record is 24.74, and here is the part that gets skipped in almost every telling: its units are per cent squared. The units are not a formality and not pedantry. Per cent squared is the reason the variance can never be handed to anyone on its own.
The trouble shows up the moment it is used. Those two numbers are not measured in the same thing, so the record cannot be written up as having a mean of 0.50 per cent and a spread of 24.74. Putting them in one sentence is a category error, like reporting a room as four metres wide and nine square metres long. Nor can the record be said to move typically 24.74 in a month. No month in the record moved 24.74 of anything. And nobody, anywhere, carries an intuition for a squared percentage. Asked what one per cent squared feels like, nobody has an answer waiting. The quantity has no everyday meaning to have an answer about.
So why keep the variance at all, if it cannot be read? Because the very squaring that ruined the units is what made the average work in the first place. There is no version of this where the cancelling is fixed and the units are kept. Variance is the quantity the arithmetic wants and the standard deviation is the quantity a reader wants, and they are the same fact wearing two units. Almost everything built on top of spread, every interval and every test, runs on variances because variances behave well when they are combined, and almost everything reported to a human is a standard deviation because that is the only one of the two that can be laid against a mean.
The variance of the fifty month record is 24.74. What are its units, and what does that rule out doing with the number?
What does taking the square root back actually buy?
The square root of 24.74 is 4.97 per cent. The single square root is the entire relationship between the two measures: the standard deviation is the square root of the variance, and the variance is the standard deviation squared. There is nothing else between them. No extra assumption, no extra data, no judgement call. Either one of the pair gives the other.
The root buys units, and units are what make a sentence possible. 4.97 per cent is in per cent, the same unit the record is in, so it can be set next to a mean of 0.50 per cent and read out loud: a typical month of the fifty sat roughly five percentage points away from the record's own centre. No sentence of that kind can be built on the variance, in per cent squared, at all.
There is a second consequence, and it explains a great deal of confusing reporting. Double the spread of a record and the standard deviation doubles, but the variance quadruples, on exactly the same data. Stretch every month's distance from the centre by a factor of two and the standard deviation moves from 4.97 to 9.95 per cent, while the variance moves from 24.74 to 98.98. The two statements describe the identical change to the identical record. One of them looks like a doubling and the other looks like an explosion. The gap is a fact about squaring, not about the record, and it is worth remembering the next time a variance figure in a note appears to have gone through the roof.
A record's standard deviation doubles. What happens to its variance?
What do all three measures look like off one tally?
Every figure below comes off the same five line tally and nothing else was fed in. Read it downward: the value, how many months carried it, the distance from the mean of 0.50 per cent, that distance squared, and the square multiplied by the count. The last column adds to 1,212.50 and both denominators are applied to that one total.
| Monthly change | Months | Distance from 0.50 | Squared | Squared times months |
|---|---|---|---|---|
| minus 9.00 per cent | 5 | minus 9.50 | 90.25 | 451.25 |
| minus 4.00 per cent | 9 | minus 4.50 | 20.25 | 182.25 |
| 1.00 per cent | 25 | 0.50 | 0.25 | 6.25 |
| 6.00 per cent | 8 | 5.50 | 30.25 | 242.00 |
| 11.00 per cent | 3 | 10.50 | 110.25 | 330.75 |
| Totals across the record | 50 | 0.00 | 1,212.50 |
| Measure | How it is worked out | Answer | Units |
|---|---|---|---|
| Range | 11.00 less minus 9.00, using two months | 20.00 | per cent |
| Variance, 49 denominator | 1,212.50 divided by 49 | 24.7449 | per cent squared |
| Standard deviation, 49 denominator | the square root of 24.7449 | 4.9744 | per cent |
| Variance, 50 denominator | 1,212.50 divided by 50 | 24.2500 | per cent squared |
| Standard deviation, 50 denominator | the square root of 24.2500 | 4.9244 | per cent |
| True spread of the generator | known by construction, not worked out from the record | 5.0000 | per cent |
Every figure in the two tables above comes off the tally of fifty months and nothing else. The true spread of 5.0000 per cent is the only line that did not come out of the record; it comes from the stated rule that produced the record.
Before the panel below runs, commit to a prediction. One month of the fifty falls from minus 9.00 to minus 30.00 per cent and the other forty nine are held exactly where they are. Which moves more, in proportion to where it started?
Drag one month of fifty into the basement and watch the two measures disagree.
The panel opens on the published record: five months at minus 9.00 per cent, nine at minus 4.00, twenty five at 1.00, eight at 6.00 and three at 11.00. One of those months, drawn in lime, is the only thing that moves. The other forty nine are held exactly where they are, so any change that appears is the work of a single month. The range bar and the standard deviation bar underneath are drawn on the same scale as the value axis above them, so the comparison is between two lengths and not two numbers. At the opening setting the lime month sits on top of the four fixed months at minus 9.00 per cent and restores the published record exactly.
Educational illustration on invented data. Exactly one of the fifty months moves; the other forty nine are held at their published values, and the counts are asserted on screen at every setting. The variance and standard deviation shown use the 49 denominator. The moving month has to stay the smallest of the fifty for the range bar to mean what it says, so the panel will not take it above minus 9.00 per cent.
Why are there two denominators, and how far apart are their answers?
The total of 1,212.50, the sum of the weighted squared distances, is not in dispute; it falls straight out of the tally. The divisor is what is in dispute.
Divide by 50, the number of months, and the variance is 24.25 and the standard deviation is 4.92 per cent. Divide by 49, the number of months less one, and the variance is 24.74 and the standard deviation is 4.97 per cent. The two denominators disagree, and the 49 answer is the larger of the two, for the plain reason that dividing by a smaller number gives a bigger result.
The reason the smaller denominator exists at all is worth holding in one sentence. Every distance here was measured from the record's own mean of 0.50 per cent. The record's own mean was worked out from those very fifty months and has been pulled towards them, so it sits closer to the record than the true centre does. So the squared distances come out slightly too small, on purpose and on every record ever measured. Dividing by 49 rather than 50 inflates the answer just enough to correct for it.
The convention that follows is simple to state. If the fifty months are the entire thing in question and there is nothing beyond them, the divisor is 50. If the fifty months are a sample and the real question is about whatever produced them, the divisor is 49. Almost every question anyone asks about a record of monthly changes is the second kind. Most software returns the 49 answer by default for that reason, and it is still worth checking which one was returned.
The same fifty month record produces a variance of 24.25 on one denominator and 24.74 on the other. Which is which, and which of the two is larger?
How wrong was this record's spread against the true one?
An estimate can almost never be checked against the truth. The truth is almost never available. The generator behind the fifty month record was written down first, so its true spread is known to be 5.00 per cent exactly. Put the three numbers on one line and look at them: the truth is 5.00, the 49 denominator gives 4.97 and the 50 denominator gives 4.92.
Both estimates landed below the truth on this record, and the 49 answer landed closer. The 49 answer is low by 0.03 percentage points and the 50 answer is low by 0.08. So on these fifty months the correction did exactly the job it was built to do. Note what has happened in passing: an estimatorA recipe that turns a record of observations into a single figure standing in for something that cannot be observed directly. The word names the recipe, not the answer the recipe produced. was applied to fifty months, it produced a point estimateThe single figure an estimator hands back, with no width around it. One number stands where the honest answer is usually a stretch of numbers, so a width is normally quoted alongside., and for once the answer can be checked against the truth rather than against another estimate.
Now the sentence that keeps this honest, and it matters more than the pleasing result above. One record proves nothing whatsoever about which denominator is better. Another fifty months drawn from the same generator could put both answers above 5.00 per cent, or could make the 50 answer the nearer of the two. The question is settled by how each denominator behaves across many records rather than by how each did on one, and that argument is covered separately. The record can honestly show the size of the disagreement, and the disagreement is small: three hundredths of a percentage point between the truth and the better estimate, and five hundredths between the two estimates themselves.
The true spread is 5.00 per cent and both of this record's answers, 4.97 and 4.92 per cent, landed below it. Does that show the method runs low?
What should be asked before a spread figure from anybody is accepted?
Spread figures travel badly. A figure gets lifted out of one note into another, loses its label on the way, and by the third retelling nobody can say what was divided by what. Five questions catch almost every problem, and the fifty month record answers all five in a line each.
First, is it a variance or a standard deviation, and what are its units? 24.74 and 4.97 per cent are the same fact, and only one of them can be read beside a mean. Second, how many cases is it built on? Fifty months, here, and a spread built on nine months is a different kind of claim from one built on nine hundred. Third, the denominator: 49 here, and the gap to the 50 answer is five hundredths of a percentage point. Fourth, is the figure being driven by two cases or by all of them? The range of 20.00 per cent is two months; the standard deviation of 4.97 per cent is fifty. Fifth, what does it look like beside the centre it belongs to? A spread of 4.97 per cent against a mean of 0.50 per cent says the month to month movement dwarfs the average, and that is a different picture from 4.97 against a mean of 40.00.
A lender sizing a working capital limit for a shop does exactly this, without the vocabulary. Average monthly takings decide how much the shop can service; the spread of those takings decides how large the buffer has to be. A shop averaging Rs 4,00,000/- a month with quiet months at Rs 1,50,000/- needs a facility that survives the quiet months and not the average one. An analyst reading a cost line does the same thing in reverse: a cost that averages steady but swings hard is a cost with something structural inside it worth asking about. A spread quoted without its centre is exactly as incomplete as a centre quoted without its spread, and the two are only ever useful as a pair.
Somebody produces the number 24.74 with no label attached to it. What are the first two questions to ask?
The two ways a spread figure gets misread, and what each one costs
The first is small and common. A note reports that the Nakshatra unit's variance has risen from 24.25 to 24.74, and a reader takes that as the record having become more volatile. Nothing about the record changed at all. The two figures are the identical fifty months divided by 50 and by 49, and somebody switched software between one note and the next. The cost is a paragraph of explanation that should never have been needed.
The second is larger and it is the one worth guarding against. A reader compares a variance of 24.74 taken from one source with a standard deviation of 4.97 per cent taken from another, does not notice that the two are in different units, and concludes that the first record is roughly five times as volatile as the second. The two are the same record. The comparison is wrong by the square root of itself, and nothing in either number carries a label that would have stopped it. Two records really can differ fivefold in spread, so the conclusion is not absurd on its face, and that is exactly what makes it survive a review.
The fix is a habit rather than a check, and it costs one clause. Never quote a spread without saying which of the two it is and what it was divided by, in the same sentence as the number itself. Write it as the standard deviation of 4.97 per cent on the 49 denominator, or as the variance of 24.74 in per cent squared on the 49 denominator, and both failures above become impossible rather than merely unlikely.
Somebody compares a variance of 24.74 taken from one note with a standard deviation of 4.97 per cent taken from another, and concludes that the first record is about five times as volatile. What has gone wrong?
Which centre to quote is settled separately. The shape of a record, meaning whether the months lean to one side or carry a heavier than usual tail, is covered separately, and a record that is not symmetricMatching on both sides of its centre, so the picture folded along the middle would land on itself. A record can be lopsided and still report a perfectly ordinary spread. can still report a perfectly ordinary spread. Computing these measures for inputs supplied by a reader is handled separately by a calculator. How a spread turns into a width around an estimate is covered separately, as is the standard errorHow much an estimate itself would jump about if the record were collected again. The wobble is in the answer, not in the spread of the months, and it is a different quantity with a different formula. that measures how much an estimate would wobble across records. Fitting a line through a set of points is covered separately.
Who published these figures, and why is the honest answer nobody?
Squaring a distance and dividing by a count is arithmetic, and arithmetic has no publisher. The tally of fifty months was set down for teaching before any figure was computed from it. Writing the generator down first is the only reason a true spread of 5.00 per cent can be quoted beside an estimate at all.
| Source | Document | Site |
|---|---|---|
| None named | The fifty month tally of the Nakshatra unit, set down for this guide | Not published anywhere, so there is no site to name |
| The arithmetic itself | The two denominators and the correction between them sit in every statistics text and belong to none of them | Standard in every statistics text, published in particular by none of them |
The Nakshatra unit and its fifty month record are invented.
Educational material. Not advice on any investment, tax, budget or market position.
