The Sampling Distribution: Why Every Estimate Varies
A sampling distribution is the distribution of an estimate across all the records that might have been drawn, not of the months inside one record. Five fifty month records of the Nakshatra unit give means of 0.50, 1.30, 0.90, 0.40 and 1.60 per cent. Its spread is the standard error, 0.7035 per cent here, and it falls only with the square root of the months.
Three things carry over from earlier notes. The invented Nakshatra unit is a traded unitA placeholder for anything with a price that is free to move. None of that changes the arithmetic done to the numbers, so the phrase deliberately leaves the object and its issuer unnamed. whose monthly changeThe gap between one month's price and the month before it, written as a percentage. A reading of plus 6.00 means the month finished six per cent higher than it opened, and a reading of minus 4.00 means four per cent lower. was counted over fifty months. The count of those months, set out as a tallyA line for every reading that can occur, with the number of months that landed on it written beside. Adding the lines back up recovers the length of the record exactly., puts twenty five of the fifty months on 1.00 per cent, with eight more on 6.00 and three on 11.00 above that, and nine on minus 4.00 and five on minus 9.00 underneath. Earlier notes took the average of that record and got 0.50 per cent, and measured how far its months sit from that average and got 4.97 per cent. Four more records of fifty months each were counted the same way. Every figure here is arithmetic that can be redone, and what follows divides one of those numbers by the square root of another.
One further thing carries over, and it changes what can be checked here. The generatorThe list the months were produced from: every reading that can occur, and the weight sitting on each one. Because somebody typed it out before anybody counted anything, its true centre is a fact here rather than a guess. behind all five records was written down first. Its average is a known 1.00 per cent, its spread a known 5.00 per cent, and both of those are facts of construction rather than measurements of anything. Because the truth was fixed on paper before a single month was counted, every estimate below can be put next to the answer it was reaching for and measured against it exactly. Almost no record anybody encounters arrives with that.
What is a sampling distribution a distribution of?
Every distribution built before this one was a distribution of months. Each listed what a month can do and put weight on every possibility, and the shape described a single month. A sampling distribution is not that. A sampling distribution is a distribution of answers.
Everything that follows hangs on the procedure that generates it, so the procedure is worth walking through slowly. A fresh fifty months come out of the generator and are averaged. The average of those fifty months is one number. The number is written down, and the fifty months are thrown away. Then the whole thing happens again: another fifty months, another average, another single number written down. And again. The pile of numbers accumulating is not a pile of months. The pile holds one answer per record, and the distribution of that pile is the sampling distribution of the mean. Each record contributes exactly one point, and the fifty months inside it have collapsed into that point and are gone.
The everyday version is a school with fifty classrooms. Walk into one classroom, measure every child in it, and write down the average height. The average height describes that classroom. Now do it again in another classroom and write down a second average. The heights of individual children live on one scale. The classroom averages live on a different scale. A very tall child and a very short child in the same room partly cancel before the average is taken, so the classroom averages will not spread out anything like as widely. Two scales, two distributions, and confusing them is the single most common way this subject goes wrong.
A sampling distribution is a distribution of what?
What did five records of the same unit actually report?
Abstraction is cheap, so here are five real answers. Five records of fifty months each were counted out of the same generator. Nothing was mis-typed in any of them, no month was dropped, and every tally adds to exactly fifty. Check that against the table below.
| Record | minus 9.00 | minus 4.00 | 1.00 | 6.00 | 11.00 | Months | Its average |
|---|---|---|---|---|---|---|---|
| the fifty month record | 5 | 9 | 25 | 8 | 3 | 50 | 0.50 per cent |
| record two | 3 | 9 | 24 | 10 | 4 | 50 | 1.30 per cent |
| record three | 4 | 10 | 23 | 9 | 4 | 50 | 0.90 per cent |
| record four | 6 | 8 | 25 | 8 | 3 | 50 | 0.40 per cent |
| record five | 2 | 8 | 26 | 10 | 4 | 50 | 1.60 per cent |
| The five together | 20 | 44 | 123 | 45 | 18 | 250 | 0.94 per cent |
Read the last column and then read the true average of 1.00 per cent, and sit with the two for a moment. The lowest answer is 0.40 per cent, less than half the truth. The highest is 1.60 per cent, more than half again above the truth. The closest, record three at 0.90 per cent, still misses by 0.10. Not one of the five records is right, and every one of them was counted correctly.
The bottom row completes the picture. The five answers average 0.94 per cent, within 0.06 of the true 1.00. The pair is a strange one: the procedure is very nearly on target across all five taken together, and it is wrong every single time one of them is taken alone. The centre of the pile is close. The pile itself is wide.
Five records of the same unit report 0.50, 1.30, 0.90, 0.40 and 1.60 per cent against a true 1.00. How many of them are right?
What is Sampling Error, and whose mistake is it?
Sampling error is the gap between what a record reports and what the population actually is. On the fifty month record it is 0.50 less 1.00, so minus 0.50 per cent. On record five it is plus 0.60. On record three it is minus 0.10. Five records, five gaps, all of them different and none of them zero.
The word error is doing real damage here and it deserves to be argued with. In ordinary English an error is something somebody did wrong: a slip, an oversight, a column added up twice. Nobody did anything wrong in any of the five records, and sampling error is not the name of a mistake but the name of a consequence. Sampling error is the consequence of looking at part of something instead of all of it. Avoiding it entirely would mean counting every month the generator will ever produce, and nobody can do that.
The everyday version: ten shops sit in one mall and there is time to walk into three of them. Whatever the visitor concludes about the mall carries the particular three walked into. The three shops were not chosen badly, and nothing inside them was measured badly. But somebody else with a different three comes back with a different conclusion, and the difference between the two answers was created by the walking, not by the measuring. The difference is sampling error, and it is present in every number anybody has ever computed from part of something.
One consequence of that matters more than the rest. Sampling error is not reduced by being more careful. Re-checking the arithmetic on the fifty month record a hundred times still leaves the answer at 0.50 per cent, still 0.50 short of the truth. Care removes the other kind of mistake, the kind where a month really was typed in twice. Care has no purchase on sampling error at all.
Is sampling error a mistake?
Does the wobble lean one way, or does it average out?
Two very different things can make an estimate miss, and they behave in opposite ways under repetition. Sampling error wanders. Across the five records it went minus 0.50, then plus 0.30, then minus 0.10, then minus 0.60, then plus 0.60. Three below the truth, two above it, with no direction it prefers. The five gaps add up to minus 0.30, or minus 0.06 a record. That is nearly nothing, and it would shrink further with more records.
A lean is the other thing, and it does not wander. Imagine a method that reads 0.80 per cent low every single time, whatever record it is handed. Its five answers would be minus 0.30, 0.50, 0.10, minus 0.40 and 0.80 per cent. The list repays a careful look. The same five gaps are sitting under it, so the wobble is identical to the wobble in the five real records. The spread of both sets of five is exactly 0.51 per cent. All five now sit below the truth, and averaging them gives 0.14 per cent. The 0.14 is not drifting towards 1.00 and never will. More records shrink a wobble and leave a lean exactly where it was.
The two must be told apart before anybody argues about either. A single record gives one gap, and one gap has no direction to speak of, so a single record cannot say which one is in play. Repetition separates them: gather more answers and watch whether their centre moves onto the truth or stops somewhere short of it. The one way lean is covered separately.
The five record means average 0.94 per cent against a true 1.00. What would settle whether that 0.06 gap is wobble or a lean?
What is the Standard Error, and where does it come from?
The sampling distribution has a spread like any other distribution, and that spread has a name. The spread is called the standard error, and it is built rather than asserted. Take the record's own spread of 4.97 per cent. Divide it by the square root of the number of months. For fifty months that root is 7.07. The answer is 0.7035 per cent, and that is the standard error of the mean of the fifty month record.
| Step | Where the figure comes from | Value |
|---|---|---|
| Spread of the months in the record | Counted from the fifty months themselves, on the smaller of the two denominators | 4.9744 per cent |
| Square root of the number of months | The square root of 50 | 7.0711 |
| Standard error of the record's average | The first line divided by the second | 0.7035 per cent |
The standard deviation of 4.97 per cent says how far one month typically sits from the average month, and the standard error of 0.70 per cent says how far one record's average typically sits from the true average. Those are two different questions about two different objects, and the names are so nearly identical that people read one and think the other. One is about a month. One is about a whole record. The larger number is roughly seven times the smaller, and the seven is not a coincidence: it is the square root of fifty, precisely what was divided out.
The everyday version again. A household's grocery bill swings between Rs 6,000/- and Rs 14,000/- across the months of a year, a wide monthly swing. A heavy month and a light month partly cancel inside the average, so the household's average monthly bill for the year does not swing anything like that much from one year to the next. The swing of a single month and the swing of a year's average are two separate quantities, and anybody judging the reliability of the yearly figure from the monthly swing alone will be badly misled.
The fifty month record reports a standard deviation of 4.97 per cent and a standard error of 0.70 per cent. Which one describes a month?
How close did the record's own Standard Error get to the truth?
Almost nowhere else can this question be asked at all, so it is worth taking. The 0.7035 per cent above was built from the record's own spread of 4.9744 per cent, and that spread is itself an estimate. The generator's true spread is 5.0000 per cent, and dividing that by the same square root of fifty gives 0.7071 per cent. So the true standard error is 0.7071 and the record's version of it is 0.7035. The two differ by 0.0036 per cent. On this record the estimate landed almost exactly on the answer.
Enjoy that for a second, then be suspicious of it. A single agreement is itself one draw from a distribution and could as easily have been a disagreement, so one record agreeing closely with the truth is not evidence that the method is reliable. Record five, counted from the same generator, has an internal spread of 4.59 per cent, and its standard error would come out at 0.65 rather than 0.70. Nothing is wrong with either. The two are answers to the same question and they differ. Any answer computed from a record behaves exactly that way.
The honest support for a standard error of about 0.70 per cent is not one record agreeing with the truth. The support is the five record averages themselves. The five averages ran 0.50, 1.30, 0.90, 0.40 and 1.60, and how far apart they sit from each other is 0.51 per cent. The 0.51 is measured directly off the sampling distribution rather than derived from anything, and it sits in the same neighbourhood as 0.70. The two are not equal, and nobody should expect them to be: the 0.51 rests on five drawOne trip to the generator. Each trip hands back a single month, and fifty of them stacked together make one record. The word carries no suggestion that anybody chose which month came out. counts and five is a very small number of anything. Three routes to the same quantity, landing at 0.7035, 0.7071 and 0.51, are more convincing together than any one of them alone.
The record's standard error is 0.7035 per cent and the generator's is 0.7071. What does that near agreement establish?
How fast does the Standard Error fall as the record grows?
Everybody already believes that more data helps. Almost nobody knows the exchange rate, and the exchange rate is the interesting part. The standard error is a spread divided by the square root of the count, so lengthening the record does not buy a proportional improvement. The improvement is the square root of the lengthening.
Work it at four lengths, holding the generator's true spread of 5.00 per cent fixed so that only the length is moving.
| Months in the record | The division | Standard error | Against the row above |
|---|---|---|---|
| 50 | 5.00 divided by the square root of 50 | 0.71 per cent | the starting point |
| 200 | 5.00 divided by the square root of 200 | 0.35 per cent | half |
| 800 | 5.00 divided by the square root of 800 | 0.18 per cent | half again |
| 3,200 | 5.00 divided by the square root of 3,200 | 0.09 per cent | half again |
Two readings come out of that table and both matter. The first is the price. Each halving of the standard error costs four times the record, so going from fifty months to two hundred buys one halving and going from fifty to three thousand two hundred buys only three. Three thousand two hundred months is a shade over two hundred and sixty years of monthly readings, and it has bought a standard error of 0.09 per cent rather than 0.71. Better, certainly. Not remotely proportional to the effort.
The second reading is the one that never goes away. The curve below falls steeply, then less steeply, then it flattens, and it keeps flattening for as long as the axis is extended, but it never touches zero at any number of months whatsoever, so no length of record ever converts an estimate into the truth. This is the arithmetic sitting underneath every complaint that a record is too short, and it is also the arithmetic that says a longer record would not have removed the problem, only shrunk it.
To halve a standard error, how much data is needed?
Stretch the record and watch the wobble refuse to fall as fast as the effort.
One control moves: the number of months in a record, from 25 up to 3,200. The generator is held exactly where it is, so its true spread stays at 5.00 per cent at every setting and anything that changes is bought purely with length. The dashed outlines are the picture at fifty months, left on screen at every setting, so the chosen record is always compared against the record the notes above actually used. The five marks on the lower scale are the five published records, each held at the distance from the truth it really landed, then rescaled to the chosen length. At the opening setting of fifty months they sit at 0.50, 1.30, 0.90, 0.40 and 1.60 per cent, the published figures exactly.
Educational illustration on invented data. The generator's true spread of 5.00 per cent is held fixed at every setting while only the number of months moves, so the panel is dividing the true spread and not any record's estimate of it. Dividing the true spread is possible only because somebody typed the generator out first, and an ordinary set of readings would not allow it. The square root of one settles nothing, so the shortest record the calculator draws is two months.
The standard error is 0.71 per cent at fifty months. What is it at two hundred, and by what rule?
What should be asked before treating an estimate as settled?
Here the arithmetic turns into something usable. An analyst reading a summary, a lender sizing up a borrower's reported margin, a household deciding what a year of grocery bills says about next year: all three are handed a single number computed from a finite record, and all three can ask the same four questions of it.
How many cases went into it? Fifty months, on the record above, and that count is what goes under the square root, so it is the first number to hunt for. How far apart are those cases from each other? A spread of 4.97 per cent here, and that is the figure being divided. What is the standard error, and was it reported at all? Here it is 0.70 per cent, and in the great majority of summaries anybody encounters it is simply absent. And how much more data would it take to halve it? Two hundred months against the fifty in hand. The fourth question is the one that converts a vague wish for more data into a number, and the number is almost always four times what anybody has.
The four questions buy a sense of scale, not a verdict. An average of 0.50 per cent with a standard error of 0.70 beside it is a very different object from an average of 0.50 per cent with a standard error of 0.05, and the two look identical until somebody computes the second figure. Turning that scale into a stated range, and turning it into a formal verdict about whether some claim survives the record, are both separate procedures with their own notes. The standard error is the raw material both of them consume.
No length of record drives the standard error to zero. What follows for any estimate anybody will ever be handed?
How this goes wrong, in a meeting where nobody made an error
Two teams study the same invented unit over the same stretch of history. Each team picks a different fifty month window, a perfectly ordinary thing to do. One team reports an average of 0.40 per cent a month. The other reports 1.60. The figures land in the same summary and the gap of 1.20 per cent is four times the first team's answer, so a meeting is called to establish which team made the mistake.
Two weeks go by. Both sets of working are checked twice. The tallies add to fifty in both. No month was entered twice, no sign was flipped, no formula pointed at the wrong column. Nothing was ever wrong, so at the end of it there is nothing to find. Two independent fifty month records of this unit differ from each other with a standard error of exactly 1.00 per cent, so a gap of 1.20 between two honest answers is an entirely ordinary thing to see.
The bill for that comes in two parts. The visible part is a fortnight of senior time spent auditing arithmetic that was already correct. The invisible part is worse and lasts longer: one of the two teams walks out of that meeting quietly marked down as the less careful one, on the strength of a difference that neither team created and neither could have avoided. The fix is not more checking. Two figures cannot be compared at all until somebody knows how far apart two honest answers were expected to fall, so the fix is to print the standard error beside every estimate as a matter of routine.
Several subjects sit right beside the sampling distribution. A population and a sample were built earlier and are used above without re-explanation. Which centre to quote from a record, and when the mean is the wrong one, is covered separately. Turning a standard error into a stated range around an estimate is covered separately, and so is turning it into a formal verdict through a hypothesis testA fixed procedure for asking whether a record is consistent with some stated claim about the population, and reporting how surprising the record would be if that claim held. The procedure itself is covered separately., which is where the notion of an intervalA pair of numbers offered as a range rather than as a single answer, with a stated chance that the truth lies between them. Building one is covered separately; the standard error is the ingredient it is built from. is properly built. A one way lean is covered separately, and no amount of data removes one. The same five records are used in Precision and Accuracy to make a different point about how being off on average differs from being merely noisy, and that argument is not repeated here. Whether a shape is symmetricA shape with the same weight on both sides of its centre, so what sits above the middle mirrors what sits below. Whether a shape leans instead of mirroring is covered separately. or leans is also covered separately. Fitting a line through a scatter, tuning a parameter until some shape agrees with the counts, and saying what the next month will do are three further procedures, each covered separately.
What would it take to prove these figures wrong?
Redoing the counting is the only check available, and also a sufficient one. No institution keeps any of this, so there is nothing to look up and nothing that can go out of date. The tally itself stands in for a citation: five counts adding to fifty, printed in full above, from which anybody with a pen can rebuild every figure above in a few minutes. If a number below does not come back out of that tally, the number below is wrong.
| What is quoted | Where it came from | Site | Last recomputed |
|---|---|---|---|
| The true average of 1.00 per cent and the true spread of 5.00 per cent | Stated properties of an invented five reading generator, fixed before any record existed | None. Nothing external is cited | 19 August 2026 |
| The five record averages of 0.50, 1.30, 0.90, 0.40 and 1.60 per cent | Counted from five invented tallies of fifty months each | None. Nothing external is cited | 19 August 2026 |
| The standard errors of 0.7035 and 0.7071 per cent | Those two spreads each divided by the square root of fifty | None. Nothing external is cited | 19 August 2026 |
| The ladder at 50, 200, 800 and 3,200 months | The same division repeated at four record lengths | None. Nothing external is cited | 19 August 2026 |
| The idea of a distribution of estimates itself | Long settled common property of the subject, carried in no single text | None. No single text carries it | Not applicable |
The Nakshatra unit and its five records are invented.
Educational material. Not advice on any investment, tax, budget or market position.
