Lag: Looking Backwards in a Series, and What It Shows
Lagging a series lines every observation up against an observation a fixed number of periods earlier, and autocorrelation is the correlation between the two columns that pairing produces. At one month over the six year record of the invented Nakshatra unit, the price series answers 0.9650, the change series 0.2011, the seasonally adjusted change series minus 0.0162. One record, three answers.
Three things are already settled before a single lag is taken. The correlation coefficient is already established, including what it means for a figure to sit near one, near zero or below zero. So is the invented six year record: 72 dated monthly observations of the Nakshatra unit, running from the month ending 31 January 2019 to the month ending 31 December 2024, measured against an opening markThe value a record is started from, sitting one step before the first observation. Here it is Rs 100.00/- on 31 December 2018, a chosen starting point rather than anything measured. of Rs 100.00/- on 31 December 2018. And so is the split between the two ways of reading that record. The price series says what one unit is worth at each month end. The change series says what moved between one month end and the next.
One more thing carries over, and it explains why the word may already feel half familiar. When a fitted line was checked earlier, its leftovers were tested for memory: the question was whether one residualThe gap between what a fitted line predicted for a row and what that row actually was. A pile of them is the part of the data the line failed to account for. could be guessed from the one before it, and the name for that check was autocorrelation. The residual check was a special case of the general operation, and the general operation needs only one thing: a column of numbers slid down against itself. No new arithmetic appears anywhere below. The one new habit is asking which column somebody slid.
What does lagging a column actually do?
A column of numbers that has an order is copied, and the copy is slid down by one row. That is it. The slide is the whole operation, and everything else follows from it.
Before the slide, each row of the table held one thing: what happened in that month. After the slide, each row holds two things: what happened in that month, and what happened in the month before it. Nothing has been calculated. Nothing has been averaged, fitted, smoothed or tested. The same numbers have been rearranged so that a question about the past and a question about the present sit on the same line, where they can be looked at together.
A vegetable stall outside one office building does this every evening without calling it anything. The seller writes today's takings in one column and yesterday's takings beside it, on the same line, in the same notebook. Nothing has been worked out yet. But the notebook is now shaped so that a particular question can be asked of it: when yesterday was good, was today usually good too? Before the second column existed, that question had nowhere to live. A lag is a rearrangement rather than a calculation, and every question about memory in a record starts as a question about that rearrangement.
The number of periods slid by is the lag. A slide of one row is a lag of one month on a monthly record. A slide of three puts this month beside the month three months back on every row. A slide of twelve puts this month beside the same month a year earlier. The unit is whatever the rows are: on this record the rows are months, so a lag of twelve means twelve months and nothing else.
What does a lag cost at the top of the column?
A lag costs rows, and the rows are worth counting rather than waving at. The first line of that table shows it. January 2019 is the first month in the record. When the copy slides down, the cell beside January 2019 is empty. The month before January 2019 is December 2018, and December 2018 is not in the record. There is nothing there to fetch.
So that row cannot be used. The row carries a value on one side and a blank on the other, and a pair with one half missing is not a pair. Every lag shortens the usable record by exactly as many rows as the lag, so a record of 72 observations lagged by one month yields 71 usable pairs and not 72. That single row is the honest price of looking backwards, and the price scales: lag by three and three rows go, lag by twelve and twelve rows go, lag by twenty four and twenty four rows go.
Counted out on this record, the cost stops being abstract. At a lag of one month 71 of the 72 rows survive, or 98.61 per cent of what the record started with. At a lag of twelve months 60 survive, or 83.33 per cent. At a lag of twenty four months 48 survive, or 66.67 per cent: a full third of a six year record thrown away in order to ask a question about what happened two years earlier. Nothing gives a warning while this is happening. The table still looks like a table, the arithmetic still runs, and a figure still appears at the end of it.
Here is the consequence people trip on, and it is not obvious until somebody says it out loud. Two figures computed at different lags on the same record are not computed on the same number of rows. A reading at a lag of one month rests on 71 pairs. A reading at a lag of twenty four months rests on 48. The two figures are printed in the same font, to the same number of decimals, in the same column of the same summary, and one of them is standing on a third less data than the other. Two lagged figures from one record are not directly comparable unless the number of surviving pairs is stated beside each of them. This is why the count of pairs is not a footnote. The count is one of the four things worth demanding beside any lagged figure.
The six year record's 72 monthly observations are lagged by one month. How many usable pairs are left, and why is it not 72?
What does one lagged pair look like on this record?
Abstraction is cheap, so the last row of the record is worth looking at directly. December 2024 is the final month of the six years. Lagged by one month, December 2024 sits beside November 2024. Those two months are one pair. But which two numbers land in the pair depends entirely on which reading of the record was slid.
Slide the price series and the pair reads Rs 187.4539/- beside Rs 164.4332/-. Two large numbers, close to each other, both in rupees, both describing what one unit was worth. Slide the change series and the same pair reads 14.00 per cent beside 4.00 per cent. Two small numbers, one more than three times the other, both describing movement rather than worth. The same lag, on the same row, of the same record, produces two pairs that look nothing like each other, and that single fact is why one record carries three answers rather than one.
The last six months, laid out with their lagged columns
Here is the tail of the record with both readings and both of their lagged copies, so the slide can be seen rather than taken on trust. Any row read left to right gives this month's worth, last month's worth, this month's movement and last month's movement, all on one line.
| Month end | Price | Price, lagged one month | Change | Change, lagged one month |
|---|---|---|---|---|
| 31 January 2019 | Rs 102.0000/- | empty | 2.00 per cent | empty |
| 31 July 2024 | Rs 173.6987/- | Rs 199.6537/- | minus 13.00 per cent | minus 6.00 per cent |
| 31 August 2024 | Rs 159.8028/- | Rs 173.6987/- | minus 8.00 per cent | minus 13.00 per cent |
| 30 September 2024 | Rs 162.9988/- | Rs 159.8028/- | 2.00 per cent | minus 8.00 per cent |
| 31 October 2024 | Rs 158.1089/- | Rs 162.9988/- | minus 3.00 per cent | 2.00 per cent |
| 30 November 2024 | Rs 164.4332/- | Rs 158.1089/- | 4.00 per cent | minus 3.00 per cent |
| 31 December 2024 | Rs 187.4539/- | Rs 164.4332/- | 14.00 per cent | 4.00 per cent |
No market was consulted for any number in that table. Every observation is built from three stated parts: a steady drift of 1.00 per cent a month, a twelve month pattern that repeats, and an irregular part. Every correlation quoted below follows from those three parts and nothing else.
The first row is shown deliberately out of order, at the top, in red. January 2019 is the opening month, and both of its lagged cells are empty. January 2019 is the row the lag deletes. Every other row in the table is complete, and 71 of the 72 rows survive rather than all 72.
A summary quotes two figures from this record, one at a lag of one month and one at a lag of twelve months. Name one reason they are not directly comparable as they stand.
What is autocorrelation, and what is it the correlation between?
There are now two columns sitting side by side, and the tool for two columns sitting side by side is already familiar. Correlating them is autocorrelation, in full: the correlation between a column and a shifted copy of itself. The only new thing about it is that both columns came out of the same record.
Everything already established about a correlation carries across without amendment. An autocorrelation cannot go above one or below minus one. A reading near one says the two columns move together. A reading near zero says knowing one gives close to nothing about the other. A reading below zero says they tend to lean opposite ways. None of it changes when the second column happens to be the first column in a different position.
The question an autocorrelation answers is narrow, and worth stating in one sentence so it cannot quietly widen later. The question is this: across the whole record, on average, how much of what a row says is reflected in what the next row says. Not whether the record is going up. Not whether the record is well behaved. Not whether next month can be worked out. One question, averaged over every surviving pair, and the answer is a single number between minus one and one.
One detail worth pinning down, because it changes the fourth decimal
There is a choice hiding inside the phrase correlate the two columns, and an account that does not name it hands over a figure that cannot be reproduced. When the copy slides, the two columns are no longer the same length as the record. The lagged column has lost its bottom row and the original has lost its top row. So which average are the deviations measured against, and which spread is divided by: the whole record's, or the two truncated columns' own?
The convention quoted here, and the one met almost everywhere, centres both columns on the whole record's own average and divides by the whole record's total squared deviation. The convention does that for two reasons that are both about comparability. A figure computed this way does not shift merely because a row fell off an end, and figures at different lags all sit on one common denominator, so they can be laid side by side and read as a shape. The alternative, running an ordinary correlation on just the surviving pairs, is perfectly defensible and gives a slightly different answer.
| Series, at a lag of one month | The convention quoted here | An ordinary correlation of the 71 pairs |
|---|---|---|
| The price series | 0.9650 | 0.9790 |
| The change series | 0.2011 | 0.2114 |
| The seasonally adjusted change series | minus 0.0162 | minus 0.0166 |
The two conventions never disagree about the story and they always disagree about the digits. A gap of exactly that kind makes a figure impossible to check when nobody wrote down which convention was used. Every other reading quoted below, and every reading the panel produces, uses the first column.
An autocorrelation is a correlation between which two columns?
Why does one record give three different answers?
Run that single operation, at a lag of one month, on three readings of the six year record. Nothing changes between the three runs except which column got slid. The lag is one month every time, the surviving pairs number 71 every time, and the convention is the one just named every time.
The price series gives 0.9650. The change series gives 0.2011. The seasonally adjusted change series gives minus 0.0162. The three figures are not three datasets disagreeing. One dataset is answering three different questions, and the spread between them is enormous: from a figure that all but touches the ceiling to one that is indistinguishable from nothing at all.
Start with the largest of the three: the one that looks impressive and means the least. A price series remembers hard for a reason that has nothing to do with anything interesting happening. A price this month is last month's price with a small movement applied to it. In December 2024 the price went from Rs 164.4332/- to Rs 187.4539/-, a large month by this record's standards, and the two figures are still obviously the same sort of number. A price is built out of its own previous price, so a price sits near that previous price by construction. A high autocorrelation on a price series is therefore a description of how prices are made rather than a finding about this particular one. Any price series of anything at all, invented or otherwise, will do the same thing.
Now the smallest of the three. The actual finding sits there. Take the change series' 0.2011, subtract from every month a fixed amount belonging to that calendar month, and the reading falls to minus 0.0162. The fall is not a small adjustment to a number. The fall is the number disappearing. Removing the pattern takes 0.2011 down to minus 0.0162, a reading indistinguishable from a record with no memory at all. All of the month to month memory in the change series was the calendar pattern.
The mechanism behind that fall is ordinary once it is seen, and worth being precise about. The record's calendar values run high in the winter months and low in the middle of the year, and neighbouring calendar values are therefore usually similar to each other. November carries a fixed 3.00 and December carries a fixed 3.00. April carries minus 2.00 and May carries minus 3.00. So a month that was high for calendar reasons is very often followed by another month that is high for calendar reasons, and a plain lag of one month reads that as memory. The reading is not memory in any useful sense. A repeating shape is being detected by a tool that cannot tell a repeating shape from a record that remembers.
The answer is worth committing to before the panel below moves. The record carries a large, real, repeating calendar pattern. With the lag moved out to twelve months and the question put to the change series, will the reading be high?
Move the lag and watch one record answer three ways at once.
One control moves: the lag, from 1 month out to 24 months. Nothing else changes at any setting. The record is the same 72 observations throughout, the three readings are the same three readings, and the convention is the one named above. The strip at the top of the panel is the record itself losing rows as the lag is pulled out, so the cost of looking further back is on screen at the same moment as the reward. The lower chart holds the whole shape at once, all twenty four lags for all three readings, with a marker sitting on whichever lag the control has moved to. At the opening setting, a lag of 1 month, the three readings come out at 0.9650, 0.2011 and minus 0.0162 on 71 pairs, exactly the figures quoted above.
Educational illustration on invented data. All three lines are three readings of one record rather than three separate datasets. The pair count falls as the lag rises, so two readings taken at different settings do not rest on the same amount of data and the panel prints the count at every setting for that reason. A column correlated with an unshifted copy of itself comes to one by definition and teaches nothing, so the panel starts at a lag of one month.
The price series gives 0.9650 and the change series gives 0.2011 at the same lag on the same record. What explains a gap that large?
Removing the calendar pattern takes the change series from 0.2011 to minus 0.0162. What does that indicate about where the memory was?
Does a high autocorrelation mean a series can be predicted?
No, and the price series' 0.9650 is the cleanest possible demonstration of why not. Read carefully what that figure claims. The figure claims that this month's price sits close to last month's price, and nothing more. The claim is true.
The question is what would be needed in order to use it. Saying something about December 2024's price would need November 2024's price of Rs 164.4332/-. But November's price is only known once November has finished. At any moment inside December, the thing the high reading says December is close to is a number that is already in the past, and the only part anybody would actually want, the 14.00 per cent that carried Rs 164.4332/- up to Rs 187.4539/-, is precisely the part the 0.9650 says nothing about. A high reading on a price series says only that the two columns are nearly the same numbers, a statement about how the column was built rather than about what comes next.
The everyday version is a household's electricity meter. A meter only counts upwards from where it already was, so this month's meter reading is very close to last month's and always will be. The correlation between the two readings, computed, comes out enormous. The bill is the difference between the two readings, and the correlation was computed on the readings themselves, so the enormous figure says nothing whatever about next month's bill. A meter that reads 4,180 units after reading 4,050 units has said a great deal about meters and almost nothing about the household.
There is a second thing that a very high reading does say, and it is a warning rather than a discovery. A column whose reading sits up near one is a column that does not come back. Ordinary readings that wander around some settled level pull themselves back toward it, and that pulling back is exactly what stops a lagged reading approaching one. When the reading refuses to fall off as the lag grows, what is on view is a series with no level to return to. The record's price series still reads 0.4982 at a lag of twelve months, a full year later, and that is a long memory by any standard. The property has a name, a proper test and a set of consequences, and every one of them is covered separately: a unit rootThe property of a record that has no settled level to be pulled back toward, so a movement is kept rather than undone. The name for it, the test for it and the consequences of it are covered separately. is a thing to test for rather than a thing to read off a lagged figure. A reading near one licenses a further question, never a conclusion.
An autocorrelation of 0.9650 on a price series. Does it mean the series can be anticipated?
Why is an autocorrelation figure not a test for a calendar pattern?
One shortcut is genuinely tempting here, and this record refuses it flatly.
The reasoning behind the shortcut is sound on its face. If a monthly record repeats itself every twelve months, then each month should resemble the month twelve rows above it, so a lag of twelve should show it. Reach for the figure, read it, and decide.
Do that here and the figure comes back at 0.0022. Two ten thousandths. As close to nothing as a number gets without being nothing. Anybody reading that alone would put down the record and conclude, quite reasonably, that there is no calendar pattern in it.
Now take the same 72 observations and group them by calendar month instead. Every January in the record, averaged: 5.00 per cent. Every June, averaged: minus 3.00 per cent. Every November: 4.00 per cent. Every April: minus 1.00 per cent. The spread is 8.00 percentage points between the strongest calendar month and the weakest, on a record whose months average 1.00 per cent overall. The pattern is not faint, not marginal and not a matter of interpretation. The pattern is enormous, and it accounts for 28.00 per cent of the whole record's variance.
Both facts are correct, and a reader who reached for the correlation figure would have concluded there was no calendar pattern in a record that carries an unmistakable one. The right response is not to distrust arithmetic. The two procedures are not answering the same question, and the one to use is the one that answers the question in hand.
Two things need saying plainly so that neither gets over-learned. The first is that the near zero reading here belongs to this particular record and is not a rule about lagged figures in general. Nothing in the arithmetic forces a twelve month reading to be small when a calendar pattern is present, and on other records it will not be. On this record the reading is small anyway. A test that can quietly return nothing when the thing is present is not a test, and a lagged figure is disqualified from the job on exactly that ground.
The second is what to reach for instead, and it fits in one line: group the observations by calendar month and average each group. Grouping is exact on this record, it used every one of the 72 observations rather than 60 of them, and it produced the twelve bars above. How the resulting pattern is then removed, and what removing it costs, are covered separately.
The twelve month reading is 0.0022 and January averages 5.00 per cent against June's minus 3.00 per cent. Which of the two is evidence about a calendar pattern?
What has to sit beside any autocorrelation figure?
The four fields below transfer directly to other work, and they apply wherever a lagged figure turns up: an analyst reading somebody else's summary, a lender looking at a borrower's month by month record of receipts, a household trying to work out whether a heavy electricity month says anything about the next one. In every case a single number arrives, detached from the operation that produced it, and four things have to arrive with it before it means anything.
Field one is which series, and whether that series is a level seriesA column of what something is worth at each date, as against a column of what moved between one date and the next. Which of the two is in hand is covered separately and at length. or a series of changes. Field two is the lag, in the units the rows are in. Field three is how many pairs survived, a count that follows from the lag and the length of the record and the only one of the four that can be worked out unaided. Field four is whether anything was taken out of the series before the lag was applied, and a seasonal adjustmentSubtracting from each observation a fixed amount belonging to its calendar month, worked out from the record's own history. How it is done, and what it takes away with it, are covered separately. is the commonest thing to have been taken out without anybody mentioning it.
Fields one and four are the two that change the answer most, and they are the two most often left out. What they do on this one record is plain. Moving field one from the price series to the change series turns 0.9650 into 0.2011. Moving field four from nothing to the calendar pattern turns 0.2011 into minus 0.0162. Between those two moves the whole distance from near one to near zero has been covered, and not one observation was added, removed or corrected. Field two changes the answer too, and field three changes only how much the answer can be trusted.
The working habit is smaller than the principle. Put the series name into the sentence that carries the figure, every single time, and if the series is a level, say the word level. The habit takes four extra words, and it removes the entire class of failure described below.
Somebody hands over an autocorrelation of 0.9650 and nothing else. What should be asked for?
The figure was right. The sentence built on it was not.
An analyst computes the one month autocorrelation of the Nakshatra unit's price series, gets 0.9650, and writes a line into a summary: the unit's price is highly predictable from one month to the next. Nobody has made an arithmetic mistake. The calculation rerun any number of times gives 0.9650 every time.
The figure actually says that this month's price sits close to last month's price, and that follows from a price being the previous price with a movement applied. The movement is the only part anybody would need, and the 0.9650 is silent about it. Run the identical calculation on the change series and the answer is 0.2011. Take the calendar pattern out and it is minus 0.0162. Three readings of one record, all correct, and the summary quoted the one that supports the strongest sounding sentence.
The figure is right, so the cost is not a wrong figure. The cost is that a claim about a level travels downstream dressed as a claim about a movement, and nobody who reads it later can see which series it came from. The line gets repeated, and then shortened. Six weeks on it is a sentence with no series attached to it at all, and nobody downstream can trace it back to the column it began in.
The fix is a working habit rather than a warning. The series name goes next to the figure, always, and if the series is a level, that goes in the same sentence. The analyst's line would then have read: the price series, a level, has a one month autocorrelation of 0.9650. The wording already says the figure is about worth rather than movement, and nobody would have built the second sentence on it.
Where this guide stops. Looking forwards rather than backwards is the mirror of the operation above, it carries a dating error of its own, and it is covered separately. Recomputing a figure inside a rolling windowA stretch of fixed length that slides along a record, with a figure recalculated inside it at every step. Covered separately, including why the recalculated figure arrives late. that slides along the record is also left alone. So is what a repeating calendar pattern actually is and how it is taken out. The reason a lagged figure is not the test for one is stated at length above. So are the name and the consequences of a record with no level to come back to, and so is the job of separating a trendThe steady direction a record drifts in across its whole length, as distinct from the month to month wobble around it. Separating the two is covered separately. from the wobble around it.
What are these figures built on?
| What appears here | How it can be rebuilt | External authority | Rechecked |
|---|---|---|---|
| The 72 monthly observations of the Nakshatra unit | Add a drift of 1.00 per cent to a repeating twelve month pattern and an irregular part, month by month, then compound the result forward from an opening mark of Rs 100.00/- to get the prices | None. Invented for teaching | 20 August 2026 |
| 0.9650, 0.2011 and minus 0.0162 at a lag of one month | Slide each of the three columns down one row and correlate it with itself, centring on the whole record's own average as stated above | None. Redo the 71 multiplications instead | 20 August 2026 |
| 0.0022 at a lag of twelve months, and 0.4982 for the price series at the same lag | The same operation at a slide of twelve rows instead of one, on the 60 pairs that survive it | None. The arithmetic is the whole of the authority | 20 August 2026 |
| The twelve calendar month averages, from 5.00 per cent down to minus 3.00 per cent | Group the 72 observations by calendar month and average each group of six. No slide involved anywhere | None. Six additions and a division, twelve times | 20 August 2026 |
The Nakshatra unit and its six year record are invented.
Educational material. Not advice on any investment, tax, budget or market position.
