Population and Sample: What You Have Versus What You Want to Know
A population is every case a statement is meant to cover. A sample is the handful actually held. The whole subject lives in that gap. The fifty month record of the Nakshatra unit, an invented example, has a mean of 0.50 per cent against a true mean of 1.00 per cent, so the estimate was half the truth, and the interval around it, minus 0.88 to 1.88 per cent, still contained it.
Everything this needs is already in place. A distribution is familiar ground: every outcome that is available, each one carrying its share of the weight. A centre and a spread compress such a thing down to two numbers, and the normal shape is familiar too, well enough to know what reaching 1.96 spreads either side of a centre leaves outside. Nothing new is being built here. The direction of travel changes. Until now a shape was written down and its properties were read off it. From here the work starts at the other end, holding a handful of numbers somebody actually collected, and reasoning backwards to the shape that produced them.
What is a population, and what does this one contain?
A population is every case a statement is meant to cover. Not the cases already seen. Every case, including the ones nobody has looked at yet and the ones that have not happened. In the claim that the vada pav stall outside the station sells three hundred a day, the population behind that sentence is every trading day the stall has and every trading day it will have. Nobody will ever see all of them. The sentence still has to mean something about all of them, or it does not mean anything at all.
Somebody wrote this population down, so it is not a mystery. The traded unitSomething priced, whose quote is refreshed over and again, so the step between successive quotes can be written down as a number. What sort of thing it is has no bearing on any arithmetic here. called the Nakshatra unit has a monthly changeThe change in a quoted price from one month to the next, written as a percentage of where the price started. A rise of one rupee on a hundred is a monthly change of 1.00 per cent. that can take exactly five values, and each value carries a fixed weight:
| Monthly change the rule can produce | Weight it carries | Value times weight |
|---|---|---|
| minus 9.00 per cent | 0.08 | minus 0.72 |
| minus 4.00 per cent | 0.18 | minus 0.72 |
| 1.00 per cent | 0.48 | 0.48 |
| 6.00 per cent | 0.18 | 1.08 |
| 11.00 per cent | 0.08 | 0.88 |
| The whole population | 1.00 | 1.00 |
Adding the right hand column gives the true mean of the population, 1.00 per cent. The same arithmetic on the squared distances from that centre gives a true standard deviationOne number saying how far the values sit from their own centre, on average, measured in the same units as the values themselves. A larger figure means the values are more scattered. of 5.00 per cent. Both of those figures are true by construction rather than by measurement: they were fixed the moment the five values and the five weights were written down, and no amount of collecting months could change them. The order here inverts the situation normally faced. The centre fell out of a weighted averageAdding values up after scaling each one by how much it is meant to count. The scaling factors have to total one between them, or what comes out is not an average of anything. taken over a list that was already whole.
What is a sample, and what exactly is in this record?
A sample is the cases actually held. Not a smaller version of the population, not a fair miniature of it, just the ones that turned up. The stall outside the station knows what it sold on the fifty days somebody bothered to write down. The fifty written down days are the sample. The stall's whole trade, in the sense that matters for a sentence about the stall, is the population, and nobody has it.
The fifty month record is fifty months in which the Nakshatra unit's generatorThe written down rule that produces values, listing what can come out and how much weight each outcome carries. The rule is the thing; the numbers it produces are only its output. was allowed to run. The five values came up 5, 9, 25, 8 and 3 times. The five counts add to 50, and that total is the only sense in which the record is complete. The record is not the unit and it is not the rule: it is fifty draws, and a different fifty months would have produced different counts.
Look at what that does to the shape. The rule puts a weight of 0.08 on minus 9.00 per cent, and the record puts 5 months out of 50 there, a share of 0.10. The rule puts 0.48 on 1.00 per cent and the record puts 0.50 there. On 6.00 per cent the record came in light, 0.16 against a weight of 0.18, and on 11.00 per cent it came in lighter still, 0.06 against 0.08. Not one of those gaps is a mistake. Gaps of that size are what fifty draws look like.
In one sentence, what separates the population from the sample here?
Why is the population never simply available to look at instead?
Because in almost every real situation the population is not something that can be gone to and read. The population is not sitting in a drawer being withheld. Much of it has not happened yet. The stall's next thousand days are part of the population and they do not exist. A hospital's population of patients includes the ones who will walk in next winter. When somebody says they measured the population, they nearly always mean their sample was large. A large sample is a different claim, and a weaker one.
So the ordinary situation looks like this. A record is held. A number is computed from it. Because whatever the number would be scored against is the very quantity the record was collected to reveal, there is no way whatsoever of scoring it. Had the answer been available, the record would have been unnecessary. None of that is a gap in technique. Not knowing is the permanent condition of the work.
The population here was defined rather than measured, so it is known. The five values and their five weights went onto paper first, ahead of any month at all, so the true mean of 1.00 per cent and the true standard deviation of 5.00 per cent are available to look at. Everything below is therefore an exercise in watching a method behave with the answer in plain view. Exactly how wrong an estimate was becomes visible. No real dataset ever shows that.
What is a point estimate, and how wrong was this one?
A point estimate is a single number offered as the best current answer about the population. Adding the fifty months gives a total of 25.00. Dividing by 50 gives the record's mean of 0.50 per cent. The figure of 0.50 per cent is the point estimate, and the arithmetic behind it is perfectly competent. There is nothing better that could have been done with those fifty months.
The true mean is 1.00 per cent. The estimate is half the truth, and nothing went wrong to make it so. No month was mis-recorded. No decimal slipped. Nobody chose a convenient window or dropped an inconvenient month. Fifty draws from a population whose values scatter by 5.00 per cent simply do not pin a centre of 1.00 per cent down to anything like one decimal place, and the record that happened to turn up landed low.
The uncomfortable half of that deserves attention. Handed 0.50 per cent on its own, nobody can tell from the number whether it is a careful estimate that landed half a point low, or a population whose centre genuinely is 0.50 per cent. The two situations produce the identical sentence. A single number carries no information at all about which of the two is in play, and the reason it carries none is that the estimatorThe rule applied to a record to produce an estimate, as distinct from the number it happens to give on this occasion. Add and divide by fifty is an estimator; 0.50 per cent is what it returned here. is a rule that gets applied to whatever record turns up.
What does every number rest on?
One tally, and nothing else. Everything below comes off the fifty month record in order, and nothing is quietly fed in from somewhere else.
| Step | Built from | Comes out at |
|---|---|---|
| The tally itself | minus 9.00 per cent five times, minus 4.00 nine times, 1.00 twenty five times, 6.00 eight times, 11.00 three times | 50 months |
| The total of the fifty months | each value times how often it came up, added | 25.00 |
| The record's mean | 25.00 divided by 50 | 0.50 per cent |
| The record's standard deviation | squared distances from 0.50, added to 1,212.50, divided by 49, square rooted | 4.97 per cent |
| The standard errorHow much the estimate itself would jump about if a fresh record of the same size were collected. Where it comes from and why it shrinks is worked through separately; here it is only a number the interval needs. of the mean | 4.97 divided by the square root of 50 | 0.70 per cent |
| The interval for the mean | 0.50 either side by 1.96 times 0.70 | minus 0.88 to 1.88 per cent |
| The interval for a single month | 0.50 either side by 1.96 times 5.02 | minus 9.35 to 10.35 per cent |
| Beside all of it, from the rule and not the record | the five values and the five weights, written down first | true mean 1.00, true spread 5.00 |
Only the last row came from somewhere the record could not reach, and that row is the scoreboard. Every other line was produced by arithmetic on fifty numbers that were allowed to be whatever they turned out to be.
The fifty month record has a mean of 0.50 per cent. What is the true mean, and how is it available at all?
What does a Confidence Interval claim, and what does it not claim?
A confidence interval is what replaces a single number with a reach. The point estimate is 0.50 per cent. The record says a mean of fifty months wobbles by 0.7035 per cent, and that figure is the standard error. Reaching out 1.96 standard errors on each side gives 1.96 times 0.7035, or 1.38 points. The interval therefore opens at minus 0.88 per cent and closes at 1.88 per cent.
The true mean of 1.00 per cent sits inside that range. Set beside the previous section, that gives the whole of this guide in two lines: the point estimate was half the truth, and the interval was right. The two sentences are not in tension. The pairing is the reason the interval exists. The estimate was doing the only job it can do, giving the single most defensible number, and it landed low. The interval was doing a different job, saying how far the single number might be off, and on this record it reached far enough.
Now the careful part. The sentence people say about intervals is usually the wrong one. The 95 per cent is a property of the method, not of this range. Imagine running the whole exercise again and again: collect fifty fresh months, compute the mean, reach out 1.96 standard errors, write the interval down. Do that a hundred times and about ninety five of the hundred intervals will have closed over the true mean. The ninety five out of a hundred is the claim. The claim is a statement about how the procedure performs in the long runWhat happens when a procedure is repeated a very large number of times, as opposed to what happened on the one occasion it was tried. A long run statement says nothing about which way any single attempt went., and it was true before anything was collected.
The claim is not that the true mean has a 95 per cent chance of lying between minus 0.88 and 1.88 per cent. The true mean is 1.00 per cent. The true mean is not wandering about. A fixed number either sits inside this particular range or it does not. Here it does, and here that can be checked. The uncertainty was never in the population. The uncertainty was in which fifty months would turn up.
Before the panel below: if a narrower interval is obtained by dropping the confidence level, does the claim get stronger or weaker?
Ask for a narrower claim and watch it walk off the truth.
One control. The slider sets the confidence level and nothing else moves: the record's mean of 0.50 per cent and its standard error stay exactly where they are, and the pinned marker at the true mean of 1.00 per cent never shifts. The population does not care what level was asked for. The lower panel turns whatever level is chosen into a count out of a hundred records. The slider opens at 95 per cent, reproducing the interval printed above to the decimal.
Educational illustration. The record's mean of 0.50 per cent and its standard error are recomputed from the tally when the calculator loads and are then held still at every setting; only the level moves. The true mean of 1.00 per cent is drawn here only because the rule was written down first, and on any real record it would not be available to draw at all.
Dragging the slider down tightens the interval around 0.50 per cent. Tightening looks like progress until the pinned marker is taken into account. Somewhere just above 52 per cent the interval lets go of the truth entirely. The panel computes the exact level and prints it, and the reason it lands there is worth naming: the record's mean of 0.50 per cent sits exactly halfway between zero and the true mean of 1.00 per cent, so the interval releases both ends at the same instant. Below that setting a statement is still being made, and it is a tighter looking statement, and it is wrong.
The interval for the mean runs from minus 0.88 to 1.88 per cent. Is the point estimate of 0.50 per cent therefore wrong?
A colleague writes that the true mean stands a 95 per cent chance of lying inside minus 0.88 to 1.88 per cent. What has gone wrong, and what is the repaired sentence?
What does a Prediction Interval claim, and why is it so much wider?
Here is a question the confidence interval cannot answer. How big is next month going to be? Not the average of months. One month.
The two questions are genuinely different, and it is easy to slide between them without noticing. The confidence interval is about the centre of the whole population, and a centre is a quiet thing: it is an average, and averages of fifty months bounce around far less than the months themselves do. A prediction interval is about one single month, and a single month has to carry the entire scatter of the population on its back.
So the reach is built differently. The interval for the mean reaches 1.96 times the standard error of 0.7035, or 1.38 points. The interval for one month reaches 1.96 times 5.0239, or 9.85 points. The 5.0239 is the record's own standard deviation of 4.97 per cent stretched very slightly, by the square root of one plus one over fifty, to admit that the centre is not known either. The interval for a single month therefore runs from minus 9.35 per cent to 10.35 per cent.
The interval for the mean is 2.76 points wide and the interval for a single month is 19.69 points wide, or 7.14 times as much. The ratio is not a rounding curiosity. The ratio says something blunt about what fifty months bought: knowing the average well says almost nothing about any particular month. Fifty more months would tighten the first interval and would barely touch the second, because the second is mostly made of the population's own scatter, and no quantity of collecting shrinks that. Scatter is a property of the thing, not of the effort spent on it.
An everyday version. A wedding caterer who has worked two hundred functions knows the average headcount to within a handful of people. The average does not tell her how many people are coming to Saturday's wedding. The average and the single wedding are different questions with different answers, and the second one stays wide no matter how many functions she works.
One interval is 2.76 per cent wide and the other is 19.69 per cent wide. Which is about the average of months and which is about a single month?
A note claims next month will land somewhere between minus 0.88 and 1.88 per cent. Which interval has been quoted, and what does it actually describe?
How can the estimate be half wrong and the interval still be right?
Because they are answers to different questions, and one tally answers all of them at once without any of the answers contradicting the others. Set the three side by side.
| The question | What this record answers | What it is a statement about |
|---|---|---|
| What is the single best figure for the average month? | 0.50 per cent | One number, chosen to be the most defensible; it says nothing about its own error |
| How far off might that figure be? | minus 0.88 to 1.88 per cent | The centre of the population, with the wobble in the estimate admitted |
| Where might one single month land? | minus 9.35 to 10.35 per cent | One month, carrying the whole scatter of the population as well |
All three are true at the same time, and quoting only the first is what makes an estimate look like a measurement. A measurement is a number taken off an instrument, and it is allowed to stand alone. An estimate is a number inferred about cases mostly unseen, and standing alone is precisely what it cannot do. The three rows above cost the same fifty months. Only the habit of reporting decides how many of them reach the person reading.
A record of fifty months reports a mean of 0.50 per cent. A record of fifty different months reports 1.30 per cent. Has something changed about the unit?
What should be asked of any estimate that arrives without its working?
The habit becomes useful away from invented units. Estimates arrive constantly and almost never with their working. A lender is told a borrower's average monthly collection. An analyst is handed an average ticket size. A household is told a scheme returned some figure. In every one of those the number is a point estimate off a sample, and it is being read as though it were a measurement.
Six questions turn it back into a statement about something. Which population is this meant to describe? How many cases is it built on? How spread out are those cases? Is this an average, or a single case? Which range comes with it? And what would the number have looked like on a different sample from the same source?
The last one is the sharpest and the least often asked. The question is not a hypothetical: a second record of fifty months from this same unchanged rule reported 1.30 per cent. Anybody comparing the two records without their intervals would report a jump of 0.80 points and start looking for what caused it. Nothing caused it. Both records came from a rule whose centre never moved.
A point estimate quoted with no range attached has left out the only part of it that says how much of it can be used. When the range is missing, the honest reading is not that the estimate is wrong. The honest reading is that not enough has been said to know, and asking is not pedantry: asking is the difference between a number a decision can be built on and a number that can only be repeated.
A point estimate arrives with no range attached. What should be asked for first, and why does the answer decide how much of the estimate can be used?
The failure a bare number causes
A reader is handed one sentence: the Nakshatra unit averages 0.50 per cent a month. Because it looks like the answer, they treat it as one. The sentence carries a clean number with two decimals and no hedging, and nothing on its face suggests an inference at all.
The true figure is 1.00 per cent, and those same fifty months would have supported it perfectly comfortably. The interval they were never shown runs from minus 0.88 to 1.88 per cent and contains it. The cost is not abstract, and it compounds in two directions. First, they will describe the unit as half as strong as it is, in every conversation and every note, and that description will sound confident because the number sounded confident. Second, when a later record of fifty months reports 1.30 per cent, they will read it as a change and go looking for a cause. On their picture 0.50 per cent was settled and 1.30 per cent is a departure from it. Both readings are wrong, and the second one wastes real work chasing an event that did not happen.
The fix is small and entirely a matter of habit. Ask for the range in the same breath as the number, and treat a point estimate quoted on its own as an unfinished sentence rather than a firm figure. If the range is not available, that is itself the finding, and it should travel with the number wherever it goes.
Where this guide stops. Choosing a centre for a lopsided record is not settled anywhere above; that argument is taken up next. The origin of the standard error, the reason estimates wander from record to record, and how much a larger count actually buys are all worked through in their own right elsewhere, and the figure of 0.70 per cent is borrowed here as a finished quantity. Interval calculation, and setting the two intervals head to head at the length that contrast deserves, are covered separately. So is fitting: drawing a relationship between two quantities, estimating a parameter to make a shape agree with data, and making a statement about a month that has not happened all come a long way after this and are reached by name at that point. The panel above alters how a finding is reported, and reporting is a different act.
Which outside authority stands behind these figures?
None does, and on this one subject that is the correct answer rather than a gap. Estimation is arithmetic. There is no board, no exchange and no maintained series whose word could make a mean of 0.50 per cent any truer than the tally already makes it. Every figure's basis is written out below, line by line. The whole calculation can be rebuilt with a pen and a sheet of paper.
| What a figure rests on | Where it came from | How to check it |
|---|---|---|
| The Nakshatra unit's five values and their five weights | Written down first, before any month was drawn | Multiply each value by its weight and add the five products. They come to 1.00 per cent. |
| The fifty month record's tally of 5, 9, 25, 8 and 3 | Invented for teaching and then held fixed | The five counts add to 50, and value times count adds to 25.00. |
| The multiplier of 1.96 in both intervals | The normal shape's own arithmetic, built in the reading before this one | It is the reach that leaves 2.5 per cent of the shape beyond each side. |
| Every interval printed in this guide | Recomputed from the tally, never copied from anywhere | Each build is shown in full beside the figure it produces. |
The Nakshatra unit, its tally of fifty months and the Vasant unit are invented.
Educational material. Not advice on any investment, tax, budget or market position.
