Versioned Datasets: Knowing Which Data Produced Which Result
A version of a record is its contents at one moment, not its name. The market office put right two cells in the Neelbagh stall record, invented for teaching, and saved the change into the same file. The same arithmetic gives Rs 69,035.45/- over 31 cells on the sheet handed over and Rs 55,723.30/- over 30 cells on the sheet there now. Neither figure is wrong. The pair cannot be traced, and that is worse.
Start with the arithmetic. One quantity is all this whole subject needs, and it never changes shape. The average takings for each stall month is the sum of the takings cells that carry a number, divided by how many of them carry one. A blank cell is not a number. A cell holding the office placeholder codeA number an office types into a cell to mean something that is not a quantity at all, such as no form arrived. The code looks like data and is not. for a return that never came is not a number either. A real zero is a number. The counting convention was settled where this record was first printed. The same record gives a different answer under a different divisor, so every figure below names its denominatorThe number the total is divided by. Change the divisor and the answer moves. The total on top need not budge at all. in the same sentence.
The second thing to fix in place is a single figure somebody already published. Worked on the Neelbagh stall record exactly as the market office handed it over, the sum of Rs 21,40,099/- over the 31 takings cells that carry a number gives Rs 69,035.45/- for each stall month. The average went into a note. Everything that follows happens to that note, and none of it happens because anybody made an arithmetic mistake.
What makes two files two versions of the same record?
Whatever the name on the front says, two files are two versions when their contents differ. The definition is the whole of it, and it repays a slow second reading. The difference sits somewhere the eye cannot reach. A file name is a label on a shelf. The name says where somebody decided to put a thing. The name says nothing at all about what is inside now, as against what was inside on Tuesday.
Think about the shopping list stuck to a fridge door. Somebody writes it on Monday. On Wednesday another person in the household crosses off the rice and adds coriander. On Friday a third person, finding the coriander already bought, rubs it out again. Three different lists have existed on that door and every one of them was called the shopping list. If two people who walked past on different days now argue about whether coriander was on it, neither of them is lying and neither can prove anything. The only copy is the one on the door, and it has been written over twice.
A record kept by an office behaves exactly the same way, with one difference that makes it worse rather than better. On the fridge the rubbing out is at least visible. In a file, nothing on the surface moves. The name stays. The column headings stay. The row headings stay. The number of columns stays. So the file name, the one mark a reader could check, is precisely the mark that never moves. The fault is quiet by construction, not by bad luck.
What makes two files two versions of the same record?
What did the market office change, and was it wrong to change it?
Two cells moved between the day the file went out and today, and both of them moved for a good reason.
The first is NB-08 Peeli Mithai, month 3. The month 3 cell read Rs 4,80,000/- in the file as handed over. The stall took nothing like that; the return slipThe small printed form a stall fills in each month and hands to the office, saying what it took. The office types the record up from a pile of these. NB-08 filed for month 3 reads Rs 48,000/-, and its other three months are Rs 42,000/-, Rs 45,000/- and Rs 47,000/-. The cell was carrying a trailing zeroOne extra zero typed on the end of a number. A trailing zero multiplies the figure by ten, and it does not look like a mistake. Every character in it is a perfectly ordinary digit.. So the office typed Rs 48,000/- in its place. The change from Rs 4,80,000/- to Rs 48,000/- is a correction anybody would make and anybody would defend. The cell was also one of the two largest figures anywhere in that takings column. A figure sitting that far out is the kind people reach for the word outlierA figure sitting a long way from the rest of its column. The word describes where the figure sits and says nothing whatever about whether it is true. to describe. Whether such a cell is an error or a real month is settled separately, never from the takings column alone.
The second change lands on NB-03 Harit Greens, month 2. The stall filed that month twice. One row was filed on day 5 for Rs 36,400/-, the other on day 19 for Rs 39,700/-. One stall and one month should produce one row, so the second filing makes a duplicate rowA second row covering the same thing as a row already in the record. A duplicate row need not be an identical copy of the first. Here it is not, and the two carry different figures. in the sense that matters. The market office works to the rule that a later filing replaces an earlier one, so it dropped the day 5 row. Dropping the earlier row is a correction anybody would defend, and the rule that decides which of two rows survives is covered under duplicate rows.
The distinction between a good correction and an untracked one slips away easily, and it is worth holding on to. Neither change is the fault at issue. The fault is that both of them went into the same file, under the same name, on a machine in the same office, with nobody told and nothing written down. The office improved its record and destroyed the only copy of the record its own published note had been worked out on, in one motion, without noticing that it had done two separate things.
The market office corrected NB-08's month 3 cell from Rs 4,80,000/- to Rs 48,000/- on the evidence of the return slip. Is that the fault at issue?
What does the same arithmetic give on each of the two versions?
The record appears below in both of its versions, side by side. Look for the difference before reading the totals. The layout is identical. The name is identical. The eight columns are identical. Two cells and one row are not.
Now the arithmetic, done the same way on each. On the version handed over on day 12 the takings total is Rs 21,40,099/- and 31 cells carry a number, so the average takings for each stall month is Rs 69,035.45/- over those 31 cells. On the version sitting on the office machine now the takings total is Rs 16,71,699/- and 30 cells carry a number, so the same average is Rs 55,723.30/- over those 30 cells. Same convention, same arithmetic, same file name, and Rs 13,312.15/- between the two answers.
| The record | Rows | Cells with a number | Takings total | That total over those cells |
|---|---|---|---|---|
| As handed over, day 12 | 32 | 31 | Rs 21,40,099/- | Rs 69,035.45/- |
| After change one, day 19 | 32 | 31 | Rs 17,08,099/- | Rs 55,099.97/- |
| As it stands now, day 24 on | 31 | 30 | Rs 16,71,699/- | Rs 55,723.30/- |
The middle line of that table is worth pausing on. The record spent five days in that state and nobody ever saw it. Correcting NB-08's cell takes Rs 4,32,000/- straight out of the takings total and leaves the count of cells exactly where it was, so the figure falls a long way, from Rs 69,035.45/- over 31 cells to Rs 55,099.97/- over the same 31. The drop is Rs 13,935.48/-, almost the whole of the distance between the two published-looking numbers.
Dropping NB-03's earlier month 2 row takes Rs 36,400/- out of the takings total. What happens to the average takings for each stall month?
The second change runs the other way, and this catches almost everybody. Dropping the day 5 filing of Rs 36,400/- takes money out of the total, so the instinct says the average must fall. It rises. Rs 36,400/- was sitting well below the Rs 55,099.97/- the record was averaging at that moment, so pulling it out lifts what is left. The count falls from 31 cells to 30 at the same time, and between them the figure moves up by Rs 623.33/-, from Rs 55,099.97/- over 31 cells to Rs 55,723.30/- over 30. A change that removes money from a record can raise its average, and this one does.
Look at the lower half of that drawing before moving on. The lower panel does something the upper one cannot. On a scale wide enough to hold Rs 69,035.45/-, a move of Rs 623.33/- is about three pixels and the two bars simply look the same length. Redraw the same two numbers on a window one thousand rupees wide and the move is unmissable. Nothing about the record changed between the two panels. A difference that vanishes at one scale and shouts at another is a fact about the drawing, never a fact about the record.
Which of the two figures is the wrong one?
Neither. Rs 69,035.45/- is a correct average of the Neelbagh stall record as it stood on day 12, over the 31 cells that then carried a number. Rs 55,723.30/- is a correct average of the same record as it stands now, over the 30 cells that carry one today. Both were worked out under the same convention. Both would be reproduced by anybody handed the same sheet. Neither contains a slip of any kind.
A figure nobody can trace back to the record that produced it is not wrong, it is untraceable, and untraceable is the worse of the two. A wrong figure has a fix: the error is found, it is corrected, the note is reissued and the matter closes. An untraceable figure has no fix at all. There is nothing to correct, nobody to apologise to, and no procedure that ends the argument. The argument is not about arithmetic. The argument is about which record two people were looking at, and the evidence that would settle it was saved over on day 19.
Two people in a household arguing about what onions cost are in the same position. One checked in March, the other in July. Both remember correctly. Neither wrote down when they looked, so the disagreement has no ending. Add one word to each memory, the month, and the disagreement disappears in a second and turns into a fact about onions. The missing thing was never accuracy. The missing thing was the label saying which day the number came from.
Version one gives Rs 69,035.45/- over 31 cells and version two gives Rs 55,723.30/- over 30. Which of the two is the correct average?
What is the smallest thing that can be written down to fix this?
Two numbers. The row count and the takings total, written beside any published figure, fold the whole problem up.
Version one is 32 and Rs 21,40,099/-. Version two is 31 and Rs 16,71,699/-. Anybody holding a file can produce both of those in about a minute: count the rows, add one column. If the pair matches, they are holding the record the published figure came from. If it does not, they are holding something else, and they know it before a single word is exchanged. Two numbers turn an argument about memory into a check anybody can run without asking permission from anybody.
Two of those four lines carry the weight. The date and the supplier are worth having and neither can be checked against a file in hand; they are somebody's account of events. The row count and the takings total are different in kind. Both are worked out from the record itself. A stamp computed from the data beats a stamp somebody types. The typed one will eventually be last month's, and nothing about the file will contradict it. Anyone who has copied a heading from an old spreadsheet and forgotten to change the month knows exactly how that failure arrives.
Writing those two numbers down is the same habit as writing the meter reading and the date on the back of an electricity bill. The habit costs one line, needs no software, and protects somebody other than the person writing it. The person protected is whoever picks the bill up in fourteen months and wants to know whether the reading came before or after the new fridge.
What two numbers would have made both of the Neelbagh figures traceable?
What does a version stamp not do?
A version stamp does not make either figure right. The stamp does not stop the market office correcting cells, and the office should go on correcting every cell it can put evidence behind. The stamp says nothing about why a cell changed, whether the change was any good, or who decided it. Held up against each of those questions, the two numbers say nothing whatever.
So a stamp answers exactly one question: which record produced this number. One answered question looks like a thin return for the trouble until the alternative is looked at. Every other question about a figure is unaskable until that one has an answer. Nobody can check arithmetic, argue about a correction, or defend a denominator when the two parties may be describing different records. Answering the small question is what puts the large ones back on the table. The day bookThe running written book an office keeps beside its record, where an odd day gets described in sentences rather than in numbers. The day book explains; the record only reports. is where a reason for a change gets written, on a different sheet of paper doing a different job.
Does a version stamp make either of the two figures right?
Why does saving over a file feel harmless at the time?
Because on the afternoon it happens, every single thing about it is fine. The office found a real error. The office fixed the error correctly. The fix went into the record the office keeps and is responsible for. Nobody else was reading that record at four o'clock on a Tuesday. Ask the person at the keyboard to name the harm and they cannot, and they are not being careless: at that moment there is no harm to name.
The cost lands three weeks later, on somebody who was not in the room and who has no way of knowing that a room was ever involved. The person who saves over a file is never the person who pays for it, and that gap in time and in seating is the entire mechanism. The gap is why this cannot be left to judgement in the moment. Judgement in the moment is exactly the faculty that has nothing to work with.
The kitchen drawer works the same way. Whoever tidies the drawer is never the one who goes looking for the scissors a fortnight later. Tidying a shared drawer without telling anybody is a small kindness and a small act of vandalism at once. The fix in both cases is a habit rather than a decision. A habit is the only thing that survives an afternoon when the harm is invisible.
If the same file is taken from the market office on different days between day 12 and day 30, does the figure drift or jump?
Take the same file on a different day and watch the published figure go out of reach.
One control moves: the day the file is asked for at the market office, anywhere from day 12 to day 30. The record, the convention and the arithmetic all stay nailed down. Two further actions are available, and neither is a slider: pinning the note to whichever day is selected, and turning on the two number stamp to see whether the sheet in hand can be told apart from the others. The panel opens on day 12 with the note pinned to day 12. The file gives Rs 69,035.45/- over 31 cells there, exactly the figure the note published.
Educational illustration. The two changes land on day 19 and day 24 and nothing else about the record moves on any day. The simplification is deliberate. A real office would be making small changes all the time, and constant small changes would make the shape below harder to read, not easier. Every figure here is an average takings for each stall month under one convention, the takings cells that carry a number divided by how many carry one, and no other quantity is computed anywhere on the panel.
Three flat stretches and two steps, and nothing at all in between. The shape matters more than any single reading on it. A drifting figure would announce itself, and anybody watching would see it moving and start asking why. A figure that sits perfectly still for six days and then jumps while nobody is looking gives no warning at any point. The stretch from day 12 to day 18 looks exactly as stable as the stretch from day 24 to day 30, and one of them is the record the note was worked out on while the other is not.
Turn the two number stamp on and the three stretches stop being interchangeable. Without the stamp, all three carry the same label. The same label is all the record ever offers. With it, they read 32 and Rs 21,40,099/-, then 32 and Rs 17,08,099/-, then 31 and Rs 16,71,699/-, and no file can match more than one of those pairs. Then move the note onto a later day and watch the whole comparison re-point itself: the published figure is whatever was true on the day somebody quoted it, and everything else is measured from there.
What is actually written down beside a published figure?
Four lines, and this is where a lender, an analyst or anybody who hands a number to somebody else starts using the idea tomorrow morning. The row count. The total of the one column the figure leans on. The day the file was taken. One sentence naming who supplied it.
Every one of those four is known before the figure exists. Knowing them in advance is the practical point, and it is why the habit survives a deadline. Nothing has to be gone back over and reconstructed at the end, when the pressure is on and the note is late. Four things already in front of the analyst are written down at the moment they are there. A result carrying those four lines can be checked by a stranger; a result without them can only be repeated by the person who made it.
An analyst reading somebody else's figure can use the same four lines in reverse, as the four questions to ask. How many rows was that. What did the column add to. When was it taken. Who from. A number that survives all four is a number that can be built on. A number that fails the first two is somebody's recollection wearing a decimal point, and the polite way to say so is to ask which record it came from and wait.
The failure: a note that cannot be defended and was never wrong
The market office publishes a note saying that a Neelbagh stall took Rs 69,035.45/- for each stall month, worked out over the 31 cells that carried a number on the day it took the file. Three weeks later a trader reads the note, thinks the figure looks high, and asks where it came from. The office opens the record, runs the same arithmetic, and gets Rs 55,723.30/- over 30 cells. There is no second copy. The only file was saved over on day 19 and again on day 24.
Now count what has actually been lost, and it is not the figure. The office looks careless about a number it computed perfectly. The trader has no way to check either figure and no reason to believe the second one more than the first. And the two changes that opened the gap were both improvements to the record. The office cannot even undo them without making its own file worse. Every party is behaving reasonably and nobody can get to the bottom of anything.
The habit that prevents all of it takes one line. On the day of publication, the row count and the takings total go beside the figure, and the file used is kept under a name nobody is going to save over. Had the note carried 32 and Rs 21,40,099/-, the trader would have counted 31 rows and Rs 16,71,699/- in the file, seen instantly that they were holding a different record, and asked a completely different and answerable question.
Three weeks after publishing, the market office cannot show which record produced its figure. What has actually been lost?
A version of a record is its contents at one moment, and two versions are told apart by those contents alone. The steps of cleaning a record, and what each step moves, are covered under data cleaning. The Rs 4,80,000/- cell that started all this is covered under outlier treatment. Whether an unusually large figure is an error or a real month is not settled by looking at the takings column at all.
Two larger subjects sit outside versioning. Making a whole body of work traceable from the question asked to the answer given is a discipline in its own right, covered under quantitative research, and it goes far beyond writing two numbers under one figure. How files are stored, named and kept, and how the run that produces a result is built so it can be run again, are covered under programming for finance. Underneath both of them sits the plain idea that a result should be traceable back to the record that produced it.
Telling two versions apart needs no statistical method at all. One arithmetic average, with its denominator named beside it every time, carries the whole of the argument.
Where every number above came from
| Number | Where it is set | How it can be re-derived |
|---|---|---|
| The Neelbagh stall record, in both of its versions | Made up for teaching and printed in full on the opening notes of this set | Count the rows on the two sheets drawn above and re-add the takings column by hand |
| Rs 21,40,099/- over 31 and Rs 16,71,699/- over 30 | Added up from the takings cells shown in this guide | Add the printed cells; the two totals are what the sheets themselves carry |
| Rs 69,035.45/- and Rs 55,723.30/- | Divided in this guide under the convention named beside each one | Rs 21,40,099/- divided by 31, and Rs 16,71,699/- divided by 30 |
| Rs 13,312.15/- and the middle reading Rs 55,099.97/- | The two figures subtracted one from the other, and the average after the first correction only | Subtract one figure from the other; and divide Rs 17,08,099/- by 31 |
The Neelbagh market, the Neelbagh stall record, the market office, the day book and the stalls NB-01 to NB-10 are invented.
Educational material. Not advice on any investment, tax, budget or market position.
