Lead: Looking Forwards, and the Danger of Doing It Accidentally
A lead lines every observation up against a later one, and it turns into a data error the moment a figure built that way is filed under a month it could not have been worked out in. A three month average of the invented Nakshatra unit's price misses the price by Rs 4.9061/- on average using only months that had happened, and by Rs 2.5431/- once next month quietly gets in.
Two things are carried in from earlier and neither is rebuilt below. The first is the invented six year record: 72 dated monthly observations of the Nakshatra unit, running from the month ending 31 January 2019 to the month ending 31 December 2024, measured against an opening markThe value a record is measured back to, fixed just before its first observation. Here it is Rs 100.00/- on 31 December 2018, one day before the first month begins. of Rs 100.00/- on 31 December 2018. The second is the three month moving average, built earlier: three consecutive prices added and divided by three.
The other thing carried in is the mirror of the operation below. The arithmetic of a slide is identical in both directions. Sliding a column of numbers backwards against itself is set out under the lag. When the slide runs the other way the arithmetic does not change. The date the answer may be filed under does, and that single difference is the whole of this guide. Everything below is one price path, two columns of three month averages, and a question about which of them a person sitting at the end of a given month could actually have written down.
What is actually happening when a column is led?
Start with a column that has an order. Duplicate it, then push the duplicate up a single row. The push is the entire operation. Each line of the table then carries a pair: the value for that month, sitting beside the value for the month that followed it.
A household keeping an electricity meter book does the backwards version without thinking about it. The previous leaf of the book is sitting right there, so each month the householder writes this month's units in one column and last month's units beside them. The forward version is a different matter. Filling in the June row needs July's reading, and in June the meter has not reached July yet. The book cannot be kept that way at all. The forward book can only be filled in afterwards, once the year is over and every reading exists, and at that point the householder is no longer keeping a book but describing one.
A lead is a rearrangement and nothing else, and every difficult thing in this guide is about dating rather than about arithmetic. Nothing has been averaged, fitted or estimated by sliding a column. The same numbers have been moved so that a different pair of them sits on the same line.
The rearrangement costs a row, and it costs it at the opposite end from the backwards version. Slide a column down and the first row loses its partner. The record has nothing before its own first month. Slide it up and the last row loses its partner. The record has nothing after its own last month. There is no January 2025 anywhere in the 72 rows, so sliding up by one month leaves December 2024 with nothing beside it. The empty cell is worth holding on to, and it returns at the end of this guide as the cheapest check on the whole subject.
In one sentence, what separates a lead from a lag on the same column of monthly numbers?
When is looking forwards perfectly ordinary work?
Often. A forward reach looks like a crime once the dating fault is in view, and it is not one.
Describing a completed record is the first case. Once all 72 months exist, every one of them is available to describe every other one, and asking what the average of March, April and May came to is a plain question about a finished thing. Labelling what came next is the second. Studying what a large falling month tends to be preceded by means attaching next month's outcome to this month's row on purpose, and the attachment is deliberate rather than an accident inside the work. Smoothing a historical stretch so its shape can be seen is the third, and it is the one that matters here: a great deal of ordinary seasonal adjustmentTaking a repeating calendar pattern out of a record so that what is left can be read without it. Covered separately. work runs a centred average across a completed record precisely because a centred window sits square on the month it describes instead of trailing behind it.
All three are ordinary work, and the centred moving average used below as the villain is a perfectly good tool for every one of them. The tool is not the fault. The fault is in what a person asks the tool to say, and only sometimes.
Name a use of a centred three month average that is completely legitimate.
Where does looking forwards stop being ordinary and become an error in the data?
At one exact point. A figure computed using a later observation is filed under an earlier date, and from that moment the column is making a claim that is not true: that this number existed then.
There is a snack stall outside one office building. The seller wants to know which hour of the week is busiest, so a line goes into the ledger under Tuesday saying the four o'clock hour was the strongest. The line is correct. The seller had time to add up the whole week's slips only on Thursday, so the line was worked out then. Nothing dishonest has happened and no figure is wrong. But the Tuesday sheet of the ledger now carries a sentence that nobody standing at that stall on Tuesday could have written, and if anyone later asks what the seller knew on Tuesday evening, the ledger will answer with more than the truth.
The fault is in the data rather than in a conclusion, and that is exactly why it survives every review of the reasoning. A reviewer who checks whether the arithmetic is right will find that it is. A reviewer who checks whether the argument follows will find that it does. The one thing that went wrong is the timestampThe date, or date and time, stamped on a single observation. The stamp is what tells a program where that observation belongs in the order. on the row, and neither reviewer is looking at it. The name for this fault is lookahead, and it is worth saying plainly that it is almost never done on purpose. Lookahead arrives through a step that looked like tidying.
What do the two columns look like on the six year record?
The demonstration is the reason this guide exists: the price path of the Nakshatra unit across all 72 months, with a three month average built on it twice, two different ways, and nothing else changed at all.
The honest average at any month is that month and the two before it. At November 2022 that means September, October and November 2022. Every one of those three months has finished by the time November closes, so a person sitting at the end of November 2022 could write the answer down that evening.
The centred average at any month is the month before, that month, and the month after. At November 2022 that means October, November and December 2022. December has not happened, so a person sitting at the end of November 2022 cannot write this one down. The centred average can only be written in January, and by then it is being filed under a date a month and a bit in the past.
The only difference between the two columns is one observation: the honest one reaches back two months and the centred one reaches back one month and forward one. Same three numbers added, same division by three, same price path underneath. Here is November 2022 in full.
| November 2022 | Months averaged | The three prices | The average | Miss against the price |
|---|---|---|---|---|
| The price itself | Nov 2022 | Rs 178.8026/- | Rs 178.8026/- | Rs 0.0000/- |
| The honest average | Sep, Oct, Nov 2022 | 158.9074 plus 184.3326 plus 178.8026 | Rs 174.0142/- | Rs 4.7884/- |
| The centred average | Oct, Nov, Dec 2022 | 184.3326 plus 178.8026 plus 180.5906 | Rs 181.2419/- | Rs 2.4393/- |
The two sums come to Rs 522.0425/- and Rs 543.7258/-, and dividing each by three gives the two averages in the table. The centred one lands Rs 2.3491/- closer to the price on this particular month, and it did so by knowing that December 2022 would print Rs 180.5906/-. On 30 November 2022 nobody could have known that figure.
The 72 monthly changes are whole numbers of per cent, and the price path is compounded from them against the opening mark. No series is maintained behind the record, so no observation in it carries an as-of date.
Which three months does each column average at March 2022?
Is the centred column even a different calculation?
No. The answer is the shortest true description of what went wrong, and it is worth stopping on.
Look at what the centred average at November 2024 averages: October 2024, November 2024 and December 2024. Now look at what the honest average at December 2024 averages: October 2024, November 2024 and December 2024. The two windows hold the same three months, so the two averages are the same number, and both come out at Rs 169.9987/-. The centred column at November 2024 is holding the honest column's December 2024 value.
The agreement is not a coincidence of one row. The agreement is arithmetic rather than luck, so it holds on every one of the 70 rows where both columns exist. An average of the month before, this month and the month after is by definition an average of the three months ending next month. The centred column is the honest column slid up by one row. The fault is not a different method at all; it is the same method under the wrong dates.
Seen that way, two other things fall out for free. The centred column runs out at November 2024 because the honest column runs out at December 2024 and there is nothing left to slide into the last slot. And every figure the centred column reports at month t first became available at the end of month t plus one. Written as a date rather than as a warning, lookahead says exactly that.
How much better does the centred column look?
Now measure it. Compare each column against the price it is meant to be tracking, month by month, and take the average size of the difference. Size, not direction, so a miss of Rs 3/- above and a miss of Rs 3/- below both count as Rs 3/-.
Both columns exist together on 69 months, from March 2019 to November 2024. The honest column starts at March 2019 because it needs two earlier months, and the comparison stops at November 2024 because the centred column has nothing after that. Over those same 69 months, on the same price path, with the same arithmetic:
| Across 69 months | Average miss | What it reaches |
|---|---|---|
| The honest average | Rs 4.9061/- | This month and the two before it |
| The centred average | Rs 2.5431/- | The month before, this month, and the month after |
| The fall in the miss | Rs 2.3630/- | 48.1650 per cent of the honest figure |
The last row of the table rewards a slow reading. The tracking miss is very nearly halved. Shown two methods where one of them cut the error by 48.1650 per cent, an analyst would want to know what the better one was doing differently, and would expect the answer to be a cleverer weighting, a longer window, or something worth learning.
Not one rupee of that improvement came from a better method, and all of it came from one month of information that had not happened. The window is the same length. The weights are the same. The record is the same. The only thing that changed is that one of the three months in the window had not occurred yet when the row was dated, and that alone is worth Rs 2.3630/- a month.
The leak also shows up in the shape. Drawn over a stretch of the price path, the centred line sits closer to the price almost everywhere, and it sits closer in a particular way: it turns at the same time the price turns, instead of a month after. Turning on the same month is the visual signature of a column that already knows where the path is going.
An answer is worth settling on before the panel is touched. The setting moves from the honest column to the centred one. Does the average miss rise or fall, and by roughly how much?
Flip between the two columns and watch the miss shrink while the method stays put.
One thing changes the arithmetic in this panel: which column is drawn. The window is three months at both settings, the weights are equal at both settings, and the price path never moves. The bars are that column's miss against the price, month by month, in rupees, and the dashed line across them is the average of the 69 months on which both columns exist. The last slot on the right is the whole argument in one cell. At the honest setting it holds a value. At the centred setting it is empty. December 2024 would need a January 2025 that the record does not have. The slider does not change any number; it only walks a marker along the months so that one row can be read in full, and it opens on November 2022, which is the row written out above.
Educational illustration on invented data. The centred setting is drawn to expose a fault; it is not a column anybody should hand over with a date attached. A miss is the rupee distance between one computed column and the price it claims to describe. Both settings are averaged over the identical 69 rows, so the December 2024 bar at the honest setting is drawn but excluded from the average, and the centred setting has no December 2024 bar to exclude.
The miss improves by Rs 2.3630/- on a base of Rs 4.9061/-. Where did that improvement come from?
Why does reading the rows not catch it?
Because the leak does not flatter every row. The leak flatters the average, and it is perfectly capable of making an individual month look worse.
Of the 69 months where both columns can be compared, the centred column is the closer of the two on 52 and the worse of the two on 17. Seventeen months is roughly one month in four running against the story. November 2024 is one of them, and it runs against it hard: the honest average misses by Rs 2.5862/- and the centred one misses by Rs 5.5654/-, more than twice as far. October 2023 is worse still for the centred column, missing by Rs 10.8681/- against the honest column's Rs 1.7464/-.
Then comes the review. A reviewer is handed a spreadsheet with both columns in it and asked to spot check a few rows, and picks three. On one of them the centred column is plainly the worse of the two. Nobody would expect that from a column that had been quietly given the answer. A leak flatters the average and does not flatter every row, so a reviewer sampling a handful of rows can easily land on one that looks wrong and conclude the column is fine. The fault is a property of the whole column, and the whole column is the only place it is visible.
The gap runs the other way too, and violently. July 2024 is the record's largest falling month, and the honest average misses the price there by Rs 21.5513/- while the centred one misses by Rs 4.0197/-, a gap of Rs 17.5316/- in a single month. The leak is enormous exactly where the path turns sharply, and a sharp turn is exactly where anyone reading the column would care most.
The centred column is closer on 52 months and worse on 17. Why does that make the fault harder to catch rather than easier?
In November 2024 the honest average misses by Rs 2.5862/- and the centred one by Rs 5.5654/-. Does that clear the centred column?
What is the one check that catches it every time?
The check is the last row of the computed column. The last row is the whole of it, the look takes about two seconds, and it requires no understanding of the calculation at all.
The record ends at December 2024. The honest column has a value at December 2024, Rs 169.9987/-. December rose 14.00 per cent and a trailing average cannot see that coming, so the value misses the price that month by Rs 17.4552/-. The centred column has nothing at December 2024 and cannot have anything. A value there would need January 2025, and there is no January 2025 in the record.
Where the record's last row is full and the computed column's last row is empty, the column was built from something that had not happened, and the mismatch is visible without reading a single formula. The arithmetic will be right, so checking the arithmetic finds nothing. The reasoning will follow, so checking the reasoning finds nothing either. The empty cell is the only place where a fault about dates leaves a mark visible to the eye.
The mirror version is worth knowing too, now that both are in hand. A column built by reaching backwards is short at the top: a three month trailing average has nothing at the first two months of a record. So the shape of the gap gives the direction of the reach. Short at the start means it looked back. Short at the end means it looked forward. A column that is short at both ends was centred over a window wider than one month either side.
A computed column in a sheet is blank on its last row while the record itself has a value there. What follows from that, and how long did the check take?
How does a lead get into a dated column by accident?
Four routes cover most of it, and not one of them involves anybody deciding to cheat.
The first is the one used throughout this guide. Running a centred average or a centred smoothing across a completed record is fine. The result is then written into a table under the middle month, and there the fault enters. The second is a record that gets revised. A figure is published for a month, corrected two months later, and the corrected file is loaded back under the original dates, so every row now holds a number that reached its final form after the date on the row. The third is a single figure computed over the whole record and then attached to every row of it: an average, a spread, a calendar adjustment factor. Every row then carries a value that depended on all 72 months, including the ones after it. The fourth is an ordering accident: a column is computed and the sheet is then sorted, or a sheet is sorted after a slide has already been applied, and the values end up beside the wrong dates.
None of the four is dishonest, and all four produce a column that is better than anything that could have been produced at the time. That is the uncomfortable part. In every case something got better, so the fault does not announce itself as an error.
What to ask of every computed column
For an analyst building a table, a lender reading one somebody else built, or a household bookkeeper with a spreadsheet of monthly bills, the same three questions do the work, and they are questions about dates rather than about arithmetic.
First: on what date could this value first have been calculated? Not when was it calculated, a question about somebody's Thursday afternoon, but when did the last item of information it needs become available. Second: is that the date it is filed under? If the answer to the first question is a later month than the row it sits on, the column is claiming knowledge it did not have. Third, and this is the one to reach for when thirty seconds are available rather than thirty minutes: does the column end where the record ends?
The third question is the cheapest and it would have caught every one of the four routes above. A centred smoothing filed under its middle month ends a month early. A whole record figure pasted into every row does end at the right place, but it also starts at the right place with an identical value. The identical value is the same tell in a different shape. A sorted sheet loses its alignment at both ends. The check is not of the calculation. The check is of whether the column's shape matches the record's shape, and where it does not, of why.
There is a habit worth borrowing from the lag calculator covered separately. The lag calculator was written so that it cannot produce a figure dated later than December 2024, the last month the record contains. Not because a person using it would try to, but because a tool that refuses to date a value past its own data removes the whole question. Tools are worth building that way, and a tool built by somebody else is worth asking about, with an answer of no a reason to look at the last row.
Name two of the four ways a forward reach gets into a dated column by accident.
Every number in the sentence was correct. The sentence could not be reproduced.
Somebody is asked a plain question: how closely does a three month average track the Nakshatra unit? The answerer builds a centred three month average across the six year record, compares it with the price month by month, and answers that the average tracks the unit to within about Rs 2.5431/-. Check the arithmetic and it holds. Recompute the whole column and Rs 2.5431/- comes back.
The honest answer to the question asked, on the identical record, is Rs 4.9061/-. The reported figure is 48.1650 per cent better than anything anybody could have worked out at the time, and the gap is not a matter of opinion or method preference. The gap is one month of information that had not happened, priced at Rs 2.3630/- a month.
The figure is not wrong, so the cost is not wrongness. The cost is that nobody working forwards through the record can ever reproduce it, and nothing in the output says so. The output is a column of correct arithmetic. Six weeks later the Rs 2.5431/- has been repeated into a summary with no column attached to it, and there is no longer anything in the sentence to trace back.
The review did happen, and it passed. The reviewer picked three rows and checked them. One of the three was September 2022. There the centred column misses by Rs 11.2629/- against the honest column's Rs 7.3001/-, worse rather than better. The September 2022 row is real, it is one of the 17, and it is exactly the kind of evidence that ends a review early.
The fix is a habit rather than a warning. Before publishing any computed column, look at its last row. If it is empty while the record is not, the column used something that had not happened, and the figure at the bottom of it describes a world nobody was standing in.
Where this guide stops. Pairing a series with its own earlier values is covered separately. Pairing backwards is the reverse of everything above, and the slide is borrowed from it rather than argued again. A figure worked out afresh inside a rolling windowA fixed length stretch of the most recent observations that slides forward one step at a time, so the figure is worked out afresh at every date. Covered separately. as that window travels along a record belongs elsewhere too, including the reason such a figure arrives late. So does what a repeating calendar pattern is, so does the frequencyHow often a record writes something down: daily, weekly, monthly, yearly. The record here is monthly, so each observation covers a whole month. a record is read at, and so is whether a level series ever comes back to a level, which travels under the name unit rootA property of a level series that never pulls back towards a particular level, so a jump stays in the record instead of fading out of it. Covered separately.. The one number that says how strongly a record resembles its own earlier values goes by the name autocorrelationHow closely a series resembles its own earlier values, written as a single number between minus one and one. Covered separately., and no claim above leans on it. The fault here is a dating error in a computed column, and the damage is a tracking miss measured in rupees. A tracking miss says how closely one number followed another, and it is not anything anybody made or lost. Two figures agreeing to the paisaOne hundredth of a rupee. Two figures that agree to the paisa agree to two decimal places of a rupee. is a statement about arithmetic here and about nothing else.
What sits behind these numbers, and how can each be got back?
| What this guide prints | How to get it back by hand | Anybody's authority | Last worked through |
|---|---|---|---|
| The 72 monthly changes and the price path they compound into | For each month, take a fixed drift, add that calendar month's own amount, add that month's irregular amount, and multiply the running price by the result, beginning at Rs 100.00/- | None, and none is possible: the record was typed, not gathered | 21 August 2026 |
| Rs 174.0142/- and Rs 181.2419/- at November 2022 | Add three prices and divide by three, twice, changing only which three | None. Two additions and two divisions | 21 August 2026 |
| Rs 4.9061/-, Rs 2.5431/- and the 48.1650 per cent between them | Take the size of each column's gap from the price on all 69 shared months, average each set, and divide the fall by the larger figure | None. The arithmetic is the whole of it | 21 August 2026 |
| 52 months closer and 17 months worse | Compare the two gaps row by row on the same 69 months and count which way each row went | None. Sixty nine comparisons | 21 August 2026 |
| The blank December 2024 cell | Averaging October, November and December 2024 plus January 2025 stops at the search for January 2025 | None. The record simply ends | 21 August 2026 |
The Nakshatra unit and its six year record are invented.
Educational material. Not advice on any investment, tax, budget or market position.
