Confidence Interval: The Range Your Estimate Allows
A confidence interval takes four numbers, a mean, a spread, a count and a confidence level, and returns a range for the true mean. The build divides the spread by the square root of the count to get a standard error, multiplies the standard error by the level's multiplier, and lays the result either side of the mean. Run on the made up fifty month record, it returns minus 0.88 to 1.88 per cent.
Work out an interval on any four numbers
A mean, a spread and a count go in, a level is chosen, and every rung of the build is printed separately, so every rung can be checked on paper. The defaults hold the made up fifty month record described under the next heading, and on those defaults the interval runs from minus 0.8788 to 1.8788 per cent.
Four fields in, five rungs shown, two endpoints out.
| Rung | What happens | On the numbers in the fields | Result |
|---|---|---|---|
| ONE | Divide the spread by the square root of the count | 4.974424 over 50 square rooted, 7.071068 | 0.703490 |
| TWO | Take the multiplier the chosen level asks for | at a level of 95 per cent | 1.959964 |
| THREE | Multiply the standard error by the multiplier | 0.703490 times 1.959964 | 1.378815 |
| FOUR | Subtract that half width from the mean | 0.50 less 1.378815 | minus 0.878815 |
| FIVE | Add the same half width to the mean | 0.50 plus 1.378815 | 1.878815 |
| OUT | The interval, as it would be reported | two decimal places | minus 0.88 to 1.88 |
What the level says, and what it does not
The interval from minus 0.88 to 1.88 per cent either contains 1.00 per cent or it does not, and it contains it. For this one pair of endpoints there is nothing left to be uncertain about, so the honest answer to the question how likely is it that 1.00 lies in here is one, not 0.95.
The 95 per cent belongs to the recipe, not to this pair. Below, the same recipe is run on 100 fresh records of 50 observations drawn from the generator, and 94 of the 100 intervals cover the true 1.00 per cent.
Push the count field to 800 and watch what the sentence above has to say. The level still reads 95 per cent and the recipe is still the recipe, but this particular interval then sits entirely below 1.00 and the answer flips to zero. Nothing about the arithmetic went wrong. The crossing point, holding this mean and this spread, is a count of 381.
Somebody hands over a stack of monthly figures and asks for the average. Added and divided, they give 0.50 per cent. A single number carries no width at all, so 0.50 per cent is the honest answer to the question asked and also the most confident sentence anybody will utter all week. The figure looks exactly as certain on fifty observations as it does on five. A confidence interval puts the width back. The same record, asked for two more facts about itself, hands back a pair of endpoints instead of a solitary figure. The whole build is four arithmetic operations long and every one of the four can be watched from start to finish, so a calculator is the right shape for it.
One invented setup carries the whole of what follows, and it deserves laying out completely before any reading gets produced. Call the object a traded unitAn invented label for something that has a quoted price and no promise that the price stays put. The four operations below use none of it, so who might issue such a thing, and on what terms, is left blank. and call this particular one the Nakshatra unit, a name that was made up for these notes and belongs nowhere beyond them. Each month it posts one monthly changeHow far a price finished a month from where it opened that month, written as a percentage of the opening figure. One reading per month, and no other kind is needed here., and that reading is only ever one of five numbers: minus 9.00, minus 4.00, 1.00, 6.00 and 11.00 per cent, weighted 0.08, 0.18, 0.48, 0.18 and 0.08 in that order. The five values and their five weights together are the entire generatorThe paper rule behind the readings. The rule fixes what can appear and how often, and the rule was written before a single month existed, so its properties are facts rather than estimates. Assembled in notes that come earlier.. Fifty months were then drawn out of it and stored as a tallyFive counts standing in for fifty readings, one count against each value that can appear. The order the months arrived in plays no part in any sum done here, so nothing is lost in the compression. of 5, 9, 25, 8 and 3. The tally of 5, 9, 25, 8 and 3 is the fifty month record, and the record is the only one computed on here.
The truth about this unit is available, which almost no real record ever is. The generator was typed out first, and only afterwards were fifty months drawn from it. Weighting each of the five values by the number sitting beside it and totalling the five gives the Nakshatra unit a true average month of exactly 1.00 per cent. The same operation on the squared distances gives a true spread of 5.00 per cent exactly. Neither figure was measured and neither is an estimate of anything. Both fall straight out of a definition. When the calculator returns a pair of endpoints, the truth can be set beside the answer and its position read off. A real dataset never permits that move.
Where the standard error comes from, and why adding observations tightens it at the rate they do, is covered separately under the sampling distribution. Testing a stated position against this record is covered separately under hypothesis testing. The full claim an interval makes, as against what these four operations entitle anybody to say out loud, is argued out under populations and samples.
What does this calculator actually compute?
Four numbers go in and one range comes out. The range is the set of values for the true average that the record entered is consistent with, at the level requested. Building that range takes four arithmetic operations, and the operations carry an assumption that the printed answer never shows.
One structural fact explains half of what the answer looks like, so fix it in mind immediately. The tool works out one number, a reach, and then uses that same number twice: once subtracted from the mean and once added to it. The entered mean therefore sits at the exact centre of the answer, always, and that is a fact about the arithmetic rather than a fact about the world.
The centring leaves something out. Nothing in the arithmetic puts the truth at the centre, and nothing in it makes the middle of the range likelier than the edges. Consider a shop assistant asked how long a delivery will take. Four days is a point. Three to five days is a range, and four sits in the middle of it for no better reason than that the assistant reached one day out on either side. The middle of a stated range shows where somebody started from, not where the parcel is.
Where does each of the four numbers come from?
One question is answered about each of the four inputs: where the number is found. The meaning of each input is settled elsewhere.
| Input | Where it is found | On the fifty month record |
|---|---|---|
| The mean | Add every observation in the record and divide by how many there are | 0.50 per cent |
| The spread | The standard deviation of that same record, taken on the corrected denominator | 4.9744 per cent |
| The count | How many observations that record holds. Count the rows, or add the tally up | 50 |
| The level | A convention picked in advance. Nothing in the record chooses it | 95 per cent |
All four are supplied by the reader, so the calculator estimates nothing whatsoever. The calculator has never seen the record. Nor can it check that the entered mean came from the observations it is said to have come from, whether the count is honest, or whether 95 per cent was a sensible level to request. Divide, multiply, subtract, add: those four operations are the complete list of what it does.
One of those four rows is the row where two different answers are both available from the same record, so that row deserves a note. The spread should come from the corrected denominatorDividing the total of squared gaps by one less than the count rather than by the count itself. Which of the two to use, and why they differ, is settled in the notes on dispersion and is only named here.. The fifty month record gives 4.9744 per cent that way and 4.9244 per cent the other way, so the two do not agree, and feeding the wrong one in produces an interval that is quietly a shade too narrow. Which denominator belongs where is covered separately under dispersion. Knowing which of the two is in hand is the whole of what matters here.
Which of the four inputs does this calculator work out for itself?
How is the interval built, one visible step at a time?
A tool that prints only its answer cannot be checked, and a tool that cannot be checked has to be believed instead. So this one shows the standard error, then the multiplier, then their product, then both endpoints, and each of the five lines can be reproduced on paper in under a minute.
On the defaults, the spread of 4.9744 per cent is divided by the square root of the count, and the square root of 50 is 7.0711, so the standard error is 0.7035 per cent. The level of 95 per cent asks for a multiplier of 1.959964. Reaching 1.959964 standard errors out from the middle of the normal shapeThe familiar symmetric bell, assembled in notes that come before this one. The bell appears here only as the place the multiplier is read from, and nothing about it is put together again. leaves 2.5 per cent of the shape outside on each side. Multiply 0.7035 by 1.959964 and the reach is 1.3788 per cent. Subtract the reach from 0.50 and add the reach to 0.50, and the two endpoints are minus 0.8788 and 1.8788 per cent. The printed pair reads minus 0.88 and 1.88. A reader with a pen can reproduce every line above. Nothing else entitles a tool printing an answer to be trusted at all.
A field note on how many decimals to carry. The endpoints are minus 0.8788 and 1.8788 per cent. The familiar pair, minus 0.88 and 1.88, is a presentation of them and is perfectly fine to report. The rounded pair is not fine to do further arithmetic on. The two endpoints have a second job in the notes covering hypothesis testing. At minus 0.8788 the sum there sits precisely on the line. At exactly minus 0.88 the sum tips the other way. Round for reading. Carry the full figures for anything downstream.
The spread on the fifty month record is 4.9744 per cent and the count is 50. What is the standard error, to two decimals?
The standard error is 0.7035 per cent and the multiplier at 95 per cent is 1.959964. What is the reach, and where do the two endpoints land?
How wrong was the estimate, and where did the truth land?
A known generator pays for the comparison that follows. The mean this record produced is 0.50 per cent. The Nakshatra unit's own true average month, taken straight off the rule, is 1.00 per cent. So the point estimateOne number offered as the whole answer, carrying no width of any kind. The claims such a number does and does not make are argued out under populations and samples. was out by 0.50 percentage points, which is to say it was exactly half the truth. Read on its own, 0.50 per cent gives a badly wrong idea of what this unit does in an average month.
Now look at the two endpoints. The interval runs from minus 0.8788 to 1.8788 per cent, and 1.00 sits comfortably between them. The single number missed by half and the range around it still covered the answer, and that pairing is the entire reason this arithmetic exists. The record did not become better when the interval was drawn around it. The change is that the width stopped hiding.
One caution about how to read that, and it matters. Saying the truth landed inside is a description of what happened on this one record. The description says nothing about how often such a thing happens, or about the chance that the truth is inside any particular pair of endpoints. The argument about that chance is genuinely slippery, and it is settled separately, under populations and samples. The fact is checkable and the generator is on the table, so the fact is simply recorded here.
What happens to the interval as the count rises?
One relationship inside the tool is worth watching closely. The count is the only input where a reader's decision changes the answer in a way the formula does not make obvious. Hold the mean at 0.50 per cent, hold the spread at 4.9744 per cent, hold the level at 95 per cent, and change nothing but the count.
| Observations | Standard error | Reach | Interval | Width |
|---|---|---|---|---|
| 50 | 0.7035 | 1.3788 | minus 0.8788 to 1.8788 | 2.7576 |
| 200 | 0.3517 | 0.6894 | minus 0.1894 to 1.1894 | 1.3788 |
| 800 | 0.1759 | 0.3447 | 0.1553 to 0.8447 | 0.6894 |
| 3,200 | 0.0879 | 0.1724 | 0.3276 to 0.6724 | 0.3447 |
Reading the last column downwards: each row holds four times the observations of the row above it, and each width is half the width above it. Four times the record buys half the width, and it buys nothing at all in position. That exchange rate is worth carrying around in words rather than leaving it buried in the square root, because it is the answer to a question people ask constantly: how much more data would I need to halve this. The answer is four times as much, and then four times that again for the next halving.
The second half of that bold sentence is the half people forget. 0.50 per cent is the mean that was typed in, and no amount of counting can move it, so all four rows are centred there and the interval tightens around wherever the starting point was. A householder weighs the same sack of rice on the same shaky kitchen scale twenty times instead of twice, and the average of twenty readings wobbles far less than the average of two. The width is shrinking. But if the scale itself sits two hundred grams heavy, twenty readings and two thousand readings both centre on a figure two hundred grams heavy. A settled look is what makes the tighter answer the more dangerous of the two.
Predict first, then check yourself against the panel further down. The count is quadrupled from 50 to 200 while the mean and the spread are held where they are. What happens to the width, and what happens to the centre?
Drag the number of observations and watch only the width answer.
One control moves: how many observations stand behind the interval. The record's own mean, spread and level stay pinned at 0.50 per cent, 4.9744 per cent and 95 per cent, exactly where the fifty months put them, and all three are printed below the slider, where they can be seen not moving. Two buttons add things the picture cannot show on its own. One puts the true mean of 1.00 per cent on the panel. Only a record built from a rule written in advance has a true mean to show. The other pins a faint copy of the fifty month reading in place so every other setting can be measured against it. The panel opens on a count of 50, the published record.
What is the answer assuming while it prints two decimal places?
Every reading this tool returns rests on one assumption, and the assumption is not in the record. The assumption is in the multiplier. A level of 95 per cent yields 1.959964 because that is how far out the reach must go before only 2.5 per cent of the normal shape is left on each side, and using it is a bet that the normal shape describes how a mean built from the record wobbles about the truth.
On fifty observations that is a reasonable bet, on eight it is a poor one, and the tool will return an answer at eight without a murmur of complaint. This is the part worth staring at. Nothing about the arithmetic changes between the two. Divide the spread by the square root of the count, reach out 1.959964 of those, subtract once and add once. The same four operations, in the same order, producing the same tidy two decimal places. The multiplier's right to be that number at all is what changed silently. The tool never saw the record and cannot count what it was not given, so it has no way to notice.
The consequence is specific: an interval that comes out too narrow, and reads as though it had been measured. The same mean and the same spread entered into the instrument with a count of eight return minus 2.95 to 3.95 per cent, printed to the same two decimals as everything else above. A range invites a firmer conclusion than the record can carry, and invites it in the language of precision, so a narrow interval on eight observations is a worse outcome than no interval at all. The repair is a habit rather than a formula: quote the count in the same sentence as the interval, every time. Written as minus 2.95 to 3.95 per cent the reading looks settled; written as minus 2.95 to 3.95 per cent on eight observations it tells its own story, and the next person to read it can decide what two decimal places deserve.
The second half of the caution is about the question rather than the arithmetic. Every operation above is about an average. The interval says nothing about where a single future month will land. A range for one future month is a different and considerably wider question, and it is covered separately. The record itself makes the point obvious. Within these fifty months, five landed on minus 9.00 per cent and three on 11.00, and neither number is remotely close to the range minus 0.88 to 1.88.
The instrument returns minus 0.88 to 1.88 per cent at a level of 95 per cent, and this generator's true mean is 1.00. What is the chance that this particular pair of endpoints contains 1.00?
The fifty month record returns an interval of minus 0.88 to 1.88 per cent. Does that say where next month will land?
The failure this tool makes easy: reading the level as a chance about the truth
An analyst runs the fifty month record, reads minus 0.88 to 1.88 per cent at a level of 95 per cent, and writes the sentence almost everybody writes the first time: there is a 95 per cent chance the true average month lies between those two figures. The sentence reads like a careful one. No part of these four operations supports it.
Watch it break in the instrument above. The Nakshatra unit's true average month is 1.00 per cent, fixed by the generator before a single month was drawn. Hold the mean at 0.50 and the spread at 4.974424, and push the count field to 800. The level still reads 95 per cent, the five rungs run exactly as before, and the interval comes back as 0.1553 to 0.8447 per cent with 1.00 nowhere inside it. The chance that this pair of endpoints contains the truth is zero. At a count of fifty it was one. The crossing point, holding this mean and this spread, is a count of 381.
The 95 per cent describes the recipe, run again and again on fresh records, and the strip inside the instrument shows exactly that: 100 fresh records of fifty observations, 100 intervals, 94 of them reaching across the true 1.00 per cent and 6 stopping short. Each of the 100 is right or wrong the moment it exists, and the level never says which.
The repair is a change of subject in the sentence rather than any change to the arithmetic: say what proportion of intervals built this way cover the truth, and never what chance this interval has of covering it. A credit officer handed a range for a borrower's monthly takings can act on the first sentence. The second invites a firmer decision than the record carries, in language that sounds more careful than the truthful version does.
How is an interval handed over by somebody else checked?
Here is what a lender, an analyst or anybody signing off on arithmetic they did not perform actually does, and it takes about two minutes. A credit officer looking at a small business borrower whose monthly takings have been summarised into a range, or an analyst handed a supplier's average defect rate with a range attached, runs the same five checks. None of them require the underlying observations. All five can be done on the back of the sheet itself.
First, recompute the standard error: the quoted spread over the square root of the quoted count, and see whether it lands where they landed. Second, confirm the mean sits at the centre. The two endpoints added and halved must give the mean they typed in. Minus 0.8788 plus 1.8788 is 1.0000, halved is 0.5000, the record's mean exactly. Third, quadruple the count mentally and check that the width would halve. Fourth, ask which denominator the spread came from. The same fifty months yield 4.9744 or 4.9244 depending on the answer. Fifth, and most important, ask whether the question was ever about an average at all.
The fifth check saves the most work. If the question was about a single future case, then these four operations are simply the wrong tool, and no amount of correct arithmetic will fix that. Somebody asking what the average monthly change has been is asking a question this calculator answers. Somebody asking what next month will do is asking a different one, and it is covered separately. The two questions produce different ranges of very different sizes, and reaching for the narrow one because it was to hand is the most common way this arithmetic gets misused.
There is one more thing worth asking about, and it sits underneath all five checks. An estimatorThe recipe that turns a record into an answer, held apart from the answer any one record happened to produce. The word appears here just so the recipe and its output do not get muddled. is a recipe, and the interval is the recipe's output on one record. Two people running the same recipe on two different records will get two different intervals, and neither of them is doing anything wrong. The disagreement is the width being honest about itself.
An interval of minus 0.88 to 1.88 per cent arrives with no count beside it. What should be asked for, and what does the answer change?
A colleague presents an interval, and the mean they say they entered is not at its centre. What must have gone wrong?
Four supplied numbers, five visible rungs, and a width that halves on four times the record: that is the whole of an interval for a mean. The full claim an interval makes is settled separately under populations and samples, though the misreading of a level as a chance attaching to one particular pair of endpoints is driven on screen above. The standard error, and why it falls with a square root, is covered separately under the sampling distribution. Testing a stated position against this record is covered separately under hypothesis testing. A range for a single future case is a different and much wider question and is covered separately. No line is put through a scatter here and no parameter is drawn out of a relationship between two things, because an interval for a mean reads one column of numbers and asks only how wide the answer has to be.
Every input here was typed by hand, so what is there to cite?
Nothing, and a calculator is where that answer is easiest to see. Consider what this tool actually consumes. Four numbers, and all four of them were entered by somebody. The tool goes nowhere to fetch them, keeps no copy of anything, and has no idea what the four describe. Divide, multiply, add and subtract is the whole of it. So there is no board that could ratify a reading of minus 0.8788 to 1.8788 per cent and no trading venue that could overturn one. The table sets out instead, for each number printed above, the single operation that puts it back on the screen.
| Number printed above | How it got there | Outside record behind it | Put it back yourself |
|---|---|---|---|
| The five monthly changes, minus 9.00, minus 4.00, 1.00, 6.00 and 11.00 per cent, with weights of 0.08, 0.18, 0.48, 0.18 and 0.08 | Typed out as a definition before any month had been drawn from it | None. Nothing was retrieved | Weight the five changes, total the five results, and 1.00 per cent appears. |
| The tally of 5, 9, 25, 8 and 3 months | Invented once for teaching and not edited since | None. Invented, and labelled so | Add the five counts to 50, then add change times count to 25.00. |
| The spread of 4.9744 per cent | Squared gaps totalling 1,212.50, divided by 49, then square rooted | None. Arithmetic on the row above | Divide 1,212.50 by 49 and take the square root of 24.7449. |
| The standard error of 0.7035 per cent | That spread divided by the square root of the count | None. One division | Divide 4.9744 by 7.0711. |
| The multiplier of 1.959964 | How far out the reach goes before 2.5 per cent of the normal shape is left on each side | A convention carried in many statistics texts and issued by nobody | Confirm the digits in a trusted text before carrying them anywhere else. |
| The endpoints of minus 0.8788 and 1.8788 per cent | 1.959964 multiplied by 0.7035, then laid out either side of 0.50 | None. One multiplication and two additions | Work out 1.3788, then subtract it from 0.50 and add it to 0.50. |
The Nakshatra unit, the Vasant unit and the fifty month record are invented.
Educational material. Not advice on any investment, tax, budget or market position.
