Reproducibility: Someone Else Gets Your Number From Your Data
Reproducibility means a second person, given the same data and the same steps, reaches the same number. Replication means a second record, run through the same steps, reaches the same finding. The first is a question about the analyst's own work and is achievable by being organised. The second is a question about the world and is in nobody's gift. The two words are routinely swapped for each other.
Three matters are settled elsewhere and are used below without being rebuilt: what an interval is, what a standard errorHow much an estimate would move about if the same amount of data were collected again from the same place. measures, and why several records drawn from one source disagree with each other at all. The narrower pair of questions taken up below is what it costs to make sure somebody else can arrive at the same number, and what it costs to find out whether the world agrees.
Five records are read throughout, and each one holds fifty monthly readings. All five were drawn from the same generatorThe written down mechanism a record was produced by, as opposed to the record it happened to produce. The difference between a generator and any one sample it produces is set out separately., a five value mechanism whose true mean is 1.0000 per cent and whose true spreadThe width of a set of figures, measured outwards from whatever middle they happen to sit around. One middle can carry two very different widths. is 5.0000 per cent. The mechanism was set down on paper before anything was ever drawn out of it, so both of those figures are known by construction. Stating the truth outright is a luxury no real research programme ever has, and it is what lets five honest procedures be watched landing at five different distances from it.
Two arithmetic facts do the work below, and both come straight from that construction. A figure computed from a fixed record by fixed steps has exactly one value, so a second person failing to reach it is always a fault in the written record of the steps. A figure computed from a fresh record by the same steps has a different value almost every time, so a second record failing to reach it is usually a fault in nothing at all.
What exactly are these two words, and why do they get swapped?
Replication vs Reproducibility
Reproduction is same data, same steps, same number. The analyst hands over the record, the written steps and the conventions. A second person works through them and arrives at the same figure. Nothing about the world has been tested and nothing has been learned about it. Only the bookkeeping has been tested.
Replication is new data, same steps, same finding. Somebody takes the procedure, applies it to a record the original analyst did not supply, and asks whether the thing claimed shows up again. The number will not match. Nobody expects the number to match. The question is whether the finding survives.
The household version costs nothing to picture and holds the whole distinction. A shopkeeper adds up Monday's slips, reaches a total, and puts the slips in a box. On Friday somebody pulls out the same box, adds the same slips, and reaches the same total. Adding the same slips a second time is reproduction. The second count confirms the addition was done correctly and says nothing whatsoever about Mondays. Now suppose next Monday comes, a fresh set of slips goes into a fresh box, and the total lands close to the first one. The fresh box is replication, and replication is the only one of the two that has said anything about Mondays.
The cost of swapping the two words is precise: a person who says a result was replicated when they mean somebody reran the code has claimed evidence about the world and produced evidence about a folder. The claim sounds stronger and the work behind it is weaker. And it goes the other way too. A person who says a result failed to reproduce when they mean a second record disagreed has accused somebody of sloppiness when the honest reading is that the world is noisier than fifty months can settle.
Somebody reran the analyst's code on the analyst's file and got the same number back. Which of the two words describes what they just did?
What does a second person actually need?
How to Build an Audit Trail for Quantitative Research
Five things, and every one of them is cheap on the day and impossible six months later. Which versionWhich state of a record was used, so that a record edited afterwards can still be told apart from the one a figure was computed from. Keeping records in dated states is covered separately. of the record was used. Which readings were included and how many of them there were. The conventionA rule fixed about the way a thing gets counted, settled ahead of time so the next person counts it the same way instead of guessing at it. behind every figure, including the denominatorWhatever sits below the line in a division. Under an average it is the tally of things being shared between, and swapping fifty for forty nine turns the answer into a different quantity altogether.. The order the steps ran in. And the figure each step produced.
The figure each step produced is the item people leave out, and recording it is what turns a folder into a trail. Recording the inputs tells a second person where to start. Recording the figure after every step tells them where the two sets of figures part company.
Here is the test that separates a trail from a folder: a second person who gets a different number can point at the exact step where the two sets of figures diverge, and the argument ends in one minute instead of starting again from the beginning. Without the intermediate figures, the only thing either party can compare is the final answer, and a final answer that differs says nothing about which of forty small decisions caused it.
The same discipline already turns up in ordinary life. Consider a disputed electricity bill. If the meter reading and the date it was taken were noted, the conversation at the office takes ninety seconds: one reading, the other reading, one of them is wrong, and everybody can see which line to argue about. If all that remains is a memory that the bill felt too high, there is nothing to point at and the whole month has to be reconstructed from scratch. A trail is not filing for the sake of filing. A trail is what shrinks a disagreement from a whole month down to a single line.
One more thing a trail settles, and it is worth a line of its own because it looks like a coincidence and is one. Record one's sample mean reads 0.5000 per cent. A fitted line drawn in a completely separate ten month exercise elsewhere in these notes has an intercept that also reads 0.5000. The two share no record, no arithmetic and not even the same kind of quantity, and they read alike because two figures sometimes do. Two numbers that match are not thereby related, and the only thing that settles which is which is a trail saying where each of them came from.
A second person arrives at a different number, and neither party wrote down the figure after each step. What does it now cost to settle the disagreement?
Why is a failure to reproduce always somebody's fault?
Record one holds fifty monthly readings. Add them, divide by fifty, and the answer is 0.5000 per cent. Do it again and the answer is 0.5000 per cent. Do it on a different afternoon, on a different machine, in a different order, with a different person holding the pencil, and the answer is 0.5000 per cent. The operation is deterministicLanding on an identical result on every single run, given identical inputs, with no opening anywhere for chance to get in. Arithmetic over a fixed set of figures behaves this way., and a fixed record put through fixed steps has exactly one answer waiting at the end of it.
One answer waiting at the end of the steps is what makes reproduction fully achievable, and therefore fully the original analyst's responsibility. There is no honest reason for a second person to fail. The second person cannot be unlucky. The world cannot have moved underneath them. If they end up somewhere else, one of two things is true: either they did something different from what the original analyst did, or the original analyst did something that was never written down. Both of those are a missing line in the trail, and both belong to whoever kept it.
Set that against what happens when the record changes. Run the same procedure on a second fifty month record from the same generator and the answer is 1.3000 per cent. On a third, 0.9000 per cent. On a fourth, 0.4000 per cent. On a fifth, 1.6000 per cent. Five honest runs of one correct procedure, five different answers, and not one of them is anybody's fault. Nobody chose that spread and nobody could have prevented it.
Why is a failure to reproduce always somebody's fault, when a failure to replicate usually is not anybody's?
What happens when the same steps meet a second record?
The arithmetic is about to say something the instinct fights, and it is worth slowing down for. The generator's true mean is 1.0000 per cent. The true mean is not an estimate and nobody had to work it out from data: it was written down before a single record was drawn. So the finding that the mean sits above zero is true. The finding is true of the mechanism, permanently, and no record can make it untrue.
Five fifty month records come from one generator whose true mean is 1.0000 per cent and is therefore above zero. Before stepping through them: how many of the five should be expected to show an interval that clears zero?
Here is what fifty months actually delivers. The five records read 0.50, 1.30, 0.90, 0.40 and 1.60 per cent, they average 0.9400 per cent, and they spread by 0.5128 per cent around that average. Every one of the five intervals contains the true 1.0000 per cent. Not most of them. All five. The procedure is behaving exactly as it should, and not one of the five records is defective.
And yet the finding that the mean sits above zero shows up on exactly one of the five records, and misses on the other four. Zero sits inside four of the five intervals. Four honest attempts at a true claim came back unable to say anything, and the failure was not in the procedure, not in the records, and not in anybody's care. The failure was in the length of the record.
| Record | Its mean | Its standard error | Its 95 per cent interval | Contains the truth | Clears zero |
|---|---|---|---|---|---|
| record one | 0.50 per cent | 0.70 per cent | minus 0.88 to 1.88 | yes | no |
| record two | 1.30 per cent | 0.69 per cent | minus 0.05 to 2.65 | yes | no |
| record three | 0.90 per cent | 0.72 per cent | minus 0.51 to 2.31 | yes | no |
| record four | 0.40 per cent | 0.72 per cent | minus 1.02 to 1.82 | yes | no |
| record five | 1.60 per cent | 0.65 per cent | 0.33 to 2.87 | yes | yes |
| the five together | 0.94 per cent | spread 0.5128 | against a true 1.00 | five of five | one of five |
Step through the five records one at a time and watch four honest attempts fail to find something that is really there.
One control moves: which of the five records was drawn. Nothing else is allowed to change, so nothing else does. The procedure is identical at every setting, the generator behind all five is identical, and the true mean of 1.0000 per cent stands as one fixed line that never shifts. As the control steps along, the interval bar redraws at its own width and its own position, the marker slides to that record's mean, and the tally along the top fills in. The tally along the top is the reason the panel exists. The tally counts every record looked at, not only the ones that agreed, and by the last record it shows the honest denominator of the whole exercise. The panel opens on record one, the record read earlier.
Look at which record found the effect. The choice is not a random one of the five. Record five reads 1.6000 per cent against a true 1.0000 per cent and overstates the truth by 60.00 per cent. The record that confirms a finding is systematically the one that exaggerates it, and that is arithmetic rather than bad luck.
Work out why. The answer takes three lines. To clear zero, a record's mean has to sit further from zero than its own interval is half wide. Record five's standard error is 0.6490 per cent, so its half width is 1.2721 per cent, and that is the loosest bar any of these five records has to clear. Every other record's bar is higher: 1.3547, 1.3788, 1.4136 and 1.4178 per cent. So any record from this generator that clears zero must read above 1.2721 per cent, and the truth is 1.0000 per cent. The bar sits above the truth. A record cannot both confirm this finding and be accurate about its size. The smallest overstatement a confirming record could possibly show is 27.21 per cent, and record five's 60.00 per cent is what actually turned up.
Four of the five records fail to find the effect, and the effect is genuinely there. What does that show about the records, rather than about the claim?
The one record that does find the effect reads 1.60 per cent. What is the true mean, and by how much does that record overstate it?
Which of the three ways to report this result answers the reader's question?
How to Report Quantitative Results With Uncertainty
Record one produced 0.5000 per cent, and there are three different statements that can be made with it. The three are not three styles of writing the same thing. Each one answers a different question, and two of the three are wrong for any given reader.
Before the figures are put up: which should be wider, the range for an average of fifty readings, or the range for one single next reading?
The first statement is the bare figure: 0.50 per cent. The bare figure answers nothing. A bare figure carries no length of record and no error, so a reader has no way to tell it apart from a figure computed on four readings or on four thousand. A bare number is a claim with its evidence removed.
The second statement is 0.50 per cent with a 95 per cent interval running from minus 0.8788 to 1.8788 per cent. The interval answers a question about the average: where the long run middle of this generator plausibly sits, given fifty readings. Its width comes from the standard error, and the standard error shrinks as the record lengthens.
The third statement is an interval running from minus 9.3467 to 10.3467 per cent. The third statement answers a completely different question: where one single next reading might land. Its width comes from the spread of the readings themselves. A longer record does not make any individual month behave better, so that spread does not shrink at all as the record lengthens.
The second interval is 7.1414 times as wide as the first, and quoting the narrow one to answer the wide question is the commonest reporting error in this whole subject. It is a factor of more than seven, not a rounding difference, and it is the difference between being roughly right and being wrong by an order that would embarrass anybody.
The household version makes it obvious. Somebody works out that the households along one lane spend an average of a certain amount on vegetables each month. The lane's average can be pinned down quite tightly if enough households are counted. One particular household's spending next month is a different matter. A single next month has nothing smoothing it, so the answer is much less certain, and no amount of counting other households will make it more certain. Fifty readings smooth an average. A single reading is not smoothed by anything.
A colleague quotes an interval of minus 0.8788 to 1.8788 per cent and says next month's reading will land somewhere inside it. What is wrong with that?
What would it take before this result should be called confirmed?
Four things, and none of them is expensive on its own. Reproduction by a second person working from the written trail alone, without ringing the author. A replication on a record that person chose rather than one supplied to them. The finding stated with the length of both records printed beside it, so a reader can see what each of them could and could not resolve. And the count of attempts reported rather than the count of successes.
Reporting the count of attempts is the one almost nobody does, and it is the one that changes the meaning of everything above it. One success reported on its own reads as confirmation. One success reported alongside five attempts reads as what it is. On these five records that is exactly the situation: the finding held on record five and missed on the other four, so the honest sentence names five attempts and one success, and a sentence that names only record five is not false but it is not honest either.
A write up says a finding was confirmed on a second record. Which single question comes first?
How does a working analyst use this before signing off?
Six questions, asked in order, before anybody writes the word confirmed. Which version of the record produced this figure. How many readings are behind it. Which denominator is in use, and does every figure quoted use the same one. Was the second run on the same data or on new data. How many attempts were made in total. And is the reported range the one for an average or the one for a single next reading.
Five of those six can be answered in a sentence by anybody who kept a trail, and cannot be answered at all by anybody who did not. The practical argument for the trail is a cost: five of six questions become free, and the person without a trail has to say I do not know five times in a row while somebody senior listens.
The sixth question is different and it is worth naming which one. A trail records the runs somebody wrote down. A run abandoned before anybody wrote anything down leaves no entry to find, so the count of attempts cannot be recovered from the trail of the work that was kept. So the count of attempts needs its own small discipline, and it is a one line habit: an attempt is logged when it starts, not when it works. Nothing else recovers that number afterwards, and it is the number that decides whether one success means anything.
Anybody assessing somebody else's work runs the same six questions, whether they are a lender reading a borrower's own projections, an analyst reading a research note that arrived in an inbox, or somebody on an investment committee being shown a result that is inconvenient to argue with. The questions do not require any technical skill and they do not require access to the data. The six questions require only that the person answering has kept a record of what they did, and the speed of the answer is as informative as the answer itself.
What went wrong: the right word for the wrong run, twice over
A team is handed somebody's analysis and a second fifty month record. The team runs the procedure on the second record, gets a different answer, and reports that the result failed to reproduce. Every word of that sentence is wrong except the first two. The data was not the same, so reproduction was never attempted and nothing whatsoever has been learned about whether the original work was sound. Its trail could be perfect and this exercise would not have touched it.
The team attempted a replication, and a single failed replication on fifty months is close to no evidence. On the five records here the finding holds once and misses four times while the underlying mean really is 1.0000 per cent and really is above zero. So a team reporting one failed attempt has reported an event that happens four times in five when the claim is true. The word they used, reproduce, put the blame on the original analyst; the arithmetic puts it on the length of the record.
The mirror error travels further, so it costs more. A team reruns the same file, gets the same number, and reports that the finding replicated. The team has demonstrated that arithmetic is deterministic. The claim then gets repeated by somebody who was not in the room, arrives at a decision maker as a finding that has been independently confirmed, and by then nobody remembers that the confirmation was a second pass over one file.
The fix is one sentence in the write up and it costs nothing: say whether the second run used the same data or new data. The word that follows depends entirely on that and on nothing else. Not on who ran it, not on how carefully, not on which software. One clause, written once, and neither of these two errors can be made by anybody reading it afterwards.
What is covered elsewhere
Deciding in advance what reading would change the analyst's mind is settled before any of this starts, and it is covered separately. A mechanism set against the record it happened to produce is covered separately too, and so is the reason five records drawn from one source disagree at all. Putting a finished analysis in front of somebody whose job is to attack it is covered under peer review. Writing the claim itself so that a reading could knock it down, and asking whether a result survives when an assumption is changed underneath it, are both covered further on. The ordered cells that make one analysis rerunnable on one machine, and the stages a record passes through before it reaches a table, are covered separately as well.
Three things settle the whole distinction: what the two words mean, what each one costs to achieve, and what each one is a question about.
What stands behind the figures printed above?
The five fifty month records were not extracted from any maintained series. The five records were drawn from a five value generator that was written down first, and every figure above was recomputed from those five values and the counts of each one. There is no institution behind any of it, so there is no site to name and no date on which anything was consulted.
| What produced the figure | Where it is set out | Site |
|---|---|---|
| The five value generator and its weights, fixed before a single record was drawn | printed in full above | none, nothing was read |
| The counts of each value in record one to record five, fifty readings each | printed in full above | none, nothing was read |
| The multiplier standing behind every 95 per cent interval quoted above, 1.959963985 | printed in full above | none, nothing was read |
Record one, record two, record three, record four, record five and the five value generator behind them are invented.
Educational material. Not advice on any investment, tax, budget or market position.
