Precision and Accuracy: Two Different Kinds of Being Right
Precision is how tightly a method repeats itself when it is run again. Accuracy is whether it lands on the truth. They are separate properties. Five stretches of fifty months give sample means of 0.50, 1.30, 0.90, 0.40 and 1.60 per cent, averaging 0.94 against a true 1.00: accurate on average, and not precise. A rule that always answers 0.20 never wobbles and is wrong every time.
In ordinary speech the two words are near synonyms, and most readers meet them that way. A watch that is accurate and a watch that is precise sound like the same watch. In careful work they come apart completely, and once they have come apart a method can be flawless on one and a disaster on the other. Being consistent and being correct are answers to two different questions, and only one of the two can be checked at the analyst's own desk.
One invented setup produces every number below, and stating it in full has to come before anything is claimed. There is a traded unitAny single thing that can be bought and sold and whose price moves, taken here purely as something that produces a number each month. Nothing about how it is bought, sold or valued matters here. called the Nakshatra unit, invented for teaching. Its monthly changeThe per cent by which something moved over one month. Plus 6.00 per cent means it ended the month six per cent above where it started. It is one number per month and nothing more. can take exactly five values, minus 9.00, minus 4.00, 1.00, 6.00 and 11.00 per cent, and each value carries a fixed weight: 0.08, 0.18, 0.48, 0.18 and 0.08. That list of values with their weights is the whole generatorThe written down list of what can happen and how heavily each one is favoured. Every record in this guide came out of the same one, and it never changed while the records were being produced.. Multiply and add, and the true centre of that populationThe complete set of outcomes a record is drawn from, as opposed to the handful of months actually collected. Built in earlier notes and carried forward here. comes to 1.00 per cent exactly, and its true spread to 5.00 per cent.
The true centre of that setup is available, and that is unusual. The centre is available because somebody wrote the generator down first and only afterwards drew records out of it. On real numbers there is no such luxury. The record arrives and nothing else, and the true centre stays behind a wall for the whole of the work. The red line marking that true centre below is a line real work would not normally be able to draw. The distinction is learned in a laboratory and applied in the dark.
What does precision mean, on its own?
Precision is a property of repetition. Run the same method again on fresh material and look at how far the answers land from each other. Tight together is precise. Spread out is not. The definition ends there, and what is absent from it is worth noticing. There is no mention of a correct value, no mention of a target, and no mention of whether anybody was pleased with the result. Precision is measured using the answers alone and asks nothing about where they ought to have been.
Consider a shopkeeper with a weighing scale. The same bag of rice goes on it five times and five readings are written down. If they read 2.010, 2.008, 2.012, 2.009 and 2.011 kilograms, the scale is precise: whatever it is doing, it does the same thing every time to within a couple of grams. If instead they read 2.10, 1.94, 2.06, 1.90 and 2.02 kilograms, the scale is not precise, and the true weight of the rice was never needed to say so. The readings were compared with each other. The comparison between readings is available to anybody standing at the counter.
Putting a number on it takes short arithmetic. The five answers have their own average, each sits some distance from it, and those distances summarise as one figure. For the five sample means used throughout, 0.50, 1.30, 0.90, 0.40 and 1.60 per cent, that figure comes to 0.51 per cent. The figure is a standard deviationOne number summarising how far the entries of a list typically sit from their own average. Built in earlier notes. Here it is applied to a list of five answers rather than to a list of months. computed on the five answers, and it used nothing except the five answers. Because precision needs no outside information at all, it is the half that always gets reported, and a reader who is not paying attention will take the report as a report on quality.
Of the two properties, which one can be measured without knowing the true value?
What does accuracy mean, on its own?
Accuracy is a property of position. The answers a method produces sit somewhere on average, and the distance from that place to the true value is the measure. Short distance is accurate. Long distance is not. Again, what is absent matters: accuracy says nothing at all about whether the answers agreed with each other. A method whose answers are flung all over the place can still have an average sitting exactly on the truth, and that method is accurate by this definition, uncomfortable as that sounds on first reading.
Back to the shopkeeper. The scale that read 2.010, 2.008, 2.012, 2.009 and 2.011 kilograms is beautifully consistent. Now weigh a certified one kilogram block on it. If it reads 1.060 every time, the scale is out by sixty grams, permanently, and every consistent reading it produced was a consistently wrong one. The customer paid for that sixty grams on every purchase for as long as the scale was in use. Consistency was never the property doing the protecting, so the consistency of the scale did nothing whatever to protect anybody.
Here is the awkward part. To do that check, the shopkeeper needed a certified block: an object whose true weight was already known from somewhere else. Without it there is no way to discover the sixty grams. The scale will not confess. The scale reads what it reads, calmly and repeatably, and everything about its behaviour is consistent with it being perfect. Accuracy cannot be measured without the truth, and on real numbers the truth is precisely the thing being sought. The asymmetry between the two properties is not a footnote to the subject. It is the subject.
Which is why the Nakshatra unit is built the way it is. Nobody measured its true centre of 1.00 per cent; it was decided, in advance, by writing down five values and five weights and doing the multiplication. The generator is the certified block. Every claim about accuracy below leans on it, and every one of those claims would be unavailable on real numbers.
A shop scale reads fifty grams heavy on every single weighing, without exception. Which is it?
Where exactly do the two part company?
One line holds the whole contrast, and both halves of it have now been built separately, so it can be stated without cheating. Precision is the spread of the answers around their own average, and accuracy is the distance from that average to the truth. Two different distances, measured between two different pairs of things. The first pair sits entirely inside the results themselves. The second pair reaches outside them, to a value that may or may not be visible.
Because they are two questions rather than one, they produce four answers rather than two. A method can be tight and on target, tight and off target, loose and on target on average, or loose and off target. All four exist, all four turn up in real work, and no amount of being good at one of them says anything at all about the other. Strength on one property carries no information whatever about the other. Any single word of praise for a method is therefore almost always praising only one of them.
Two kitchen scales make the pairing concrete. Scale one reads fifty grams heavy on every use: perfectly repeatable, permanently wrong, and it will never say so. Scale two wanders twenty grams either side of the right answer, sometimes over and sometimes under, and averages out correctly across several weighings. The first scale is the tidier instrument and the second is the honest one. On a single weighing, the second scale might mislead by twenty grams and the first will certainly mislead by fifty. Across ten weighings averaged, the second scale converges on the truth and the first stays exactly fifty grams out, for ever, however many weighings are added. The permanent fifty grams is the difference doing real damage.
Two people run one method on different records and get two very different numbers. What has been learned?
How precise and how accurate is a sample mean?
Now score a real method on both properties at once. The method is the ordinary arithmetic mean: add up the fifty monthly changes in a record and divide by fifty. The mean is an estimatorAny fixed procedure that turns a record into a single number offered as the answer. The arithmetic mean is one. The rule further down that ignores the record entirely is also one, which is the point. in the plainest sense: a procedure that takes months in and gives one number out.
Five separate stretches of fifty months were drawn from the identical generator, with nothing changed between them. The tallies say how often each of the five outcomes turned up, and each tally adds to fifty. Checking that by eye takes a moment. The fifty month record ran 5, 9, 25, 8 and 3. Record two ran 3, 9, 24, 10 and 4. Record three ran 4, 10, 23, 9 and 4. Record four ran 6, 8, 25, 8 and 3. Record five ran 2, 8, 26, 10 and 4.
| Record | Tally across the five outcomes | Adds to | Its sample mean |
|---|---|---|---|
| the fifty month record | 5, 9, 25, 8, 3 | 50 | 0.50 per cent |
| record two | 3, 9, 24, 10, 4 | 50 | 1.30 per cent |
| record three | 4, 10, 23, 9, 4 | 50 | 0.90 per cent |
| record four | 6, 8, 25, 8, 3 | 50 | 0.40 per cent |
| record five | 2, 8, 26, 10, 4 | 50 | 1.60 per cent |
| the five together | average of the five means | 5 | 0.94 per cent |
Work the first one through so the rest are believable. Five months at minus 9.00 contribute minus 45. Nine months at minus 4.00 contribute minus 36. Twenty five months at 1.00 contribute 25. Eight months at 6.00 contribute 48. Three months at 11.00 contribute 33. Add them: minus 45 minus 36 plus 25 plus 48 plus 33 comes to 25. Divide by fifty and the sample mean is 0.50 per cent. The other four rows are the same four steps on different counts.
Where these numbers came from. A generator was written down first: five monthly outcomes with five weights attached to them. Its centre of 1.00 per cent was then worked out by multiplying and adding, not measured by watching anything. Five tallies of fifty months each were then set out by hand so that each one adds to exactly fifty, and the five sample means follow from those tallies by arithmetic a reader can redo on paper. The true centre is available at all only because the figures were built rather than collected.
Precision needs nothing extra, so score it first. The five answers are 0.50, 1.30, 0.90, 0.40 and 1.60 per cent. Their own average is 4.70 divided by 5, or 0.94 per cent. Their spread around that average comes to 0.51 per cent, and the widest is 1.20 per cent clear of the narrowest. On a quantity whose true value is 1.00 per cent, answers ranging from 0.40 to 1.60 are not tight in any sense a reader would recognise. The sample mean is not a precise method at fifty months, and it does not pretend to be.
Now score the accuracy. Only the certified block makes that scoring possible. The five answers sit at 0.94 per cent on average, and the truth is 1.00 per cent. The distance is 0.06 per cent, and it is low. Not zero, but nothing in the arithmetic sends it one way rather than the other, so with more records the distance would keep shrinking towards nothing. The sample mean is accurate on average and imprecise, and those two verdicts are about the same method at the same moment.
Then the sharper reading, and it is the one worth carrying away. Not one of the five answers is right. Not one. The closest, record three at 0.90 per cent, is still a tenth of a percentage point out, and the worst two are six tenths out in opposite directions. Calling the sample mean accurate is a statement about the five taken together, not a promise about any one of them. Run once, in real life, the method yields exactly one of those five numbers and no way of telling which. Accuracy on average is a property of the procedure. Averaged accuracy is not a property of the answer sitting on the desk, and confusing those two is how a perfectly sound method gets blamed for a single bad result.
The five means are 0.50, 1.30, 0.90, 0.40 and 1.60 per cent. What do they average, against what truth?
Not one of the five answers is right, yet the method is called accurate. Where is the contradiction?
Can a method be perfectly consistent and always wrong?
To answer that, take the worst method anybody could write down and score it on exactly the same two properties. Here it is, in full: whatever record it is handed, it answers 0.20 per cent. The rule does not look at the months, does not add anything up and does not divide by fifty. One number sits inside the rule, and the rule reads that number out. The fixed rule is a fake, put together purely so that a deliberately poor method could be scored beside a sound one, and nobody uses it or ever should. The rule is about to start looking good, so naming it as a fake the moment it appears matters.
Run it across all five records and write down the five answers: 0.20, 0.20, 0.20, 0.20, 0.20 per cent. Their own average is 0.20 per cent. Their spread around that average is 0.00 per cent, exactly, with no rounding involved. The widest minus the narrowest is 0.00 per cent. On precision, the fixed rule does not merely beat the sample mean; it achieves the best score the measure can produce, and no method will ever beat it.
Now the other property. The truth is 1.00 per cent. The rule answers 0.20 per cent. The distance is 0.80 per cent, and it is low on the first record. The rule is also 0.80 per cent low on the second record, and the third, and the fourth, and the fifth, and it will be 0.80 per cent low on the ten thousandth record. The record has no way of touching the answer. The sample mean misses the truth by 0.42 per cent on average across the five, a good deal less than 0.80. Extra data is not an input the fixed rule uses, so the rule is perfectly precise, thoroughly inaccurate, and no amount of extra data will move it a hair.
Predict before the panel below: as the fixed rule is moved from 0.20 towards 1.00 per cent, what happens to its spread?
Slide the fixed rule across the truth, and watch the property that never changes.
One control moves: the number the fixed rule always answers. The five sample means are held where they are and no reader can adjust them. The spread readout responds as the slider moves. The second button below the slider takes the true centre away and shows the panel as it would look on real numbers, where the red line is not available. The panel opens on the worked reading above, an answer of 0.20 per cent.
Two readings from the panel are worth keeping in prose so they survive without it. At the opening setting of 0.20 per cent the rule has a spread of 0.00 per cent and misses by 0.80 per cent every time. The sample mean has a spread of 0.51 per cent and misses by 0.42 per cent on average. Now drag rightward and something uncomfortable happens on the way. Anywhere between 0.58 and 1.42 per cent, the fixed rule is closer to the truth on average than the sample mean is and still holds a perfect spread of 0.00 per cent. There is a whole stretch of settings where the useless rule beats the sound method on both properties at once, and nothing available from inside the data can reveal whether that stretch is where the work is standing. Set it to exactly 1.00 and the rule is flawless: right every time and never varying. The rule still ignores the record, and the only reason anybody could ever discover it had got lucky is the certified block sitting outside the data.
Why does the useless rule win every ordinary check?
Line up the checks a careful reviewer would actually run on a method, and something quietly alarming becomes visible. Is the answer stable when the work is redone? Do two people handed the same brief get the same number? Does the answer hold steady as months are added? Is the range across reruns tight? Every one of those is a real check, run in real work, and not one of them is silly. Precision is defined as the only property the answers can testify to on their own, so every check that can be run without the truth is a check on precision, without exception.
So the fixed rule sweeps them. The work does nothing, so the rule is perfectly stable when the work is redone. Two people get the same number, always. The months were never consulted, so the answer holds absolutely steady as months are added. Its range across reruns is zero. The sample mean, meanwhile, does honest arithmetic on the material it was given, reports 0.50 on one record and 1.30 on another, and looks flaky by comparison. The honest method is penalised for reporting truthfully how little fifty months can settle, and the fake is rewarded for having nothing to report.
The general form of it, stated plainly, is the sentence this guide exists for. Consistency is the easiest property to fake and the easiest property to measure, and that combination is exactly the wrong way round. The property that is cheap to produce is the one everybody can verify, and the property that actually matters is the one nobody can. Any review process that scores what it can see will therefore drift, over time, towards methods that are quiet rather than methods that are right, and it will do so while every individual step of the review looks sensible.
The fixed rule has a spread of exactly 0.00 per cent. Does that make it the better method?
What can be done when the truth is not available?
A certified block is almost never available. So the useful question is what remains once the red line is taken off the figure, and the honest answer is modest. Four questions help, and all four can be asked without the true value.
The first is whether the record could have moved the answer at all. Changing some months and rerunning shows whether the output shifts. The sample mean shifts immediately; the fixed rule does not budge. The rerun separates a method that is listening from a method that is talking, and it needs no truth whatever. The second is what the method says on material it has never seen. A method that answers the same thing on anything handed to it, including an empty sheet, is not doing work. The third is whether anything inside the method pushes one way rather than wandering both ways. A pinned answer is a permanent push; so is any step that can only ever move a result upward or only ever downward. The fourth is simply to report the spread beside the answer. A reader can then at least see the precision half instead of imagining it.
None of these four recovers accuracy, and it would be dishonest to suggest otherwise: they only stop precision from being mistaken for it. The prize is smaller than a reader wants and it is a real one. A method that fails these questions is not proved wrong. The method is shown to be incapable of being wrong by anything its own record could ever say. Being wrong at least leaves a trace somebody could find, and incapable of being wrong leaves none.
The truth is not visible. Which property can still be reported, and what must be said about the other?
Which corner does a method actually sit in?
An analyst comparing two methods, a lender comparing two scoring procedures, a household comparing two shopkeepers: all three are doing the same task, and the four corners give them a vocabulary for it. Precise and accurate is the goal, and it is rare, and the awkward part is that reaching it cannot be confirmed from inside the results themselves. Imprecise and accurate is where the sample mean sits here, with a spread of 0.51 per cent and a distance of 0.06 per cent from the truth: it is honest work that will review badly, and the correct response to it is more months rather than a different method. Imprecise and inaccurate is loud and gets thrown out quickly, and it is the least dangerous of the four.
The fourth corner is the one the fixed rule sits in. Precise and inaccurate is the dangerous one, precisely because it is the one that looks most professional. The rule produces a clean number, the same number, every time, on schedule, and it survives every review that cannot see the truth. No review can see the truth. A household choosing between two shopkeepers sees the one whose scale always agrees with itself and reads that as care. A reviewer scoring two procedures on stability picks the one that never argues. In both cases the property being rewarded is the one that a rigged instrument can supply for free.
The practical habit that follows is a small one and it is not a technique. When a method is praised, ask which of the two properties the praise is about. Words like stable, repeatable, robust across reruns, consistent between reviewers and low variation are all precision words. Words like unbiased, calibrated, close to the true value and correct on average are accuracy words, and every one of them is a claim that somebody, somewhere, had a certified block. Ask who had it and what it was. If nobody can name the outside value the method was checked against, then no accuracy claim has been made at all, however confident the language sounds.
One question, answerable without the truth, separates the sample mean from the rule that always answers 0.20. Which?
How this goes wrong, in a review where every sentence is true
A reviewer is handed two methods and the same five records, and asked which to keep. The first method reports 0.50, 1.30, 0.90, 0.40 and 1.60 per cent. The second reports 0.20 per cent five times over. The reviewer writes that the second method is dramatically more stable, that its answers agree perfectly across all five records while the first swings across 1.20 per cent of ground, and that the second should be preferred on those grounds. Every sentence in that review is factually correct, and the recommendation is catastrophic.
The cost of that recommendation is easy to state. The rejected method sat 0.06 per cent from the true centre and would have closed that gap further with more months. The data is not an input, so the chosen method sits 0.80 per cent low on every record ever handed to it and will still sit 0.80 per cent low after a thousand years of data. The honest method was rejected for being honest about how little fifty months can settle. The actual comparison needed the true centre, the true centre was not in the room, and the reviewer therefore never saw the comparison at all.
The fix is one question, and it can be asked without the truth. Ask what each method would do if the true value moved. Rewrite the records so the underlying centre is somewhere else entirely, rerun both, and watch. The sample mean follows; the fixed rule does not move at all. A method that cannot follow the truth when the truth moves has not been shown to be wrong, it has been shown to be unable to be right except by accident, and that question never once required knowing where the truth actually is.
What can a reader check these figures against?
Only arithmetic, and the table below is empty on purpose rather than by oversight. Whole areas of study rest on a published document that a reader can go and read. Precision and accuracy rest on none: a method is either correct or it is not, and no authority makes it so. There is no regulator behind the number 0.94 per cent, no exchange behind the tally 5, 9, 25, 8, 3, and no data provider behind a true centre of 1.00 per cent.
Two things follow from an empty table. The first is that a reader who wants to test any claim here needs access to nothing: adding the five means, dividing by five and comparing against 1.00 per cent settles whether the verdict holds. The second is that a reader carrying real numbers gets no such comfort. The true centre behind real numbers is exactly what nobody wrote down first.
| Source | Document | Site | |
|---|---|---|---|
| None named. Every figure here is arithmetic on tallies written down for teaching, so there is no maintained record to cite and no consultation date to record. Redoing the sums is the only check available, and the only one needed. | |||
The Nakshatra unit, the Vasant unit and all five records are invented.
Educational material. Not advice on any investment, tax, budget or market position.
