Stationarity and the Unit Root: The Assumption Most Models Need, and the Test for It
A series is stationary when what sits underneath it holds still: the average does not depend on when it is measured, the spread does not either, and what two observations have to do with each other depends only on how far apart they sit. The six year record of the Nakshatra unit averages Rs 115.0525/- then Rs 190.4591/- as a price across its two halves. As a change it averages 1.00 per cent across both.
Two columns follow. The two columns hold the same 72 numbers, written down two different ways, and one column can be summarised by an average while the other cannot. Nothing about the columns says which is which. The question has to be asked, and there are three ways of asking it, each buying a different amount of certainty.
The Nakshatra unit is an invented traded object written for teaching, and its six year record is 72 monthly observations running from the month ending 31 January 2019 to the month ending 31 December 2024, measured out from an opening markA starting value written down before the first observation exists, so that later rows have something to be counted from. On this record it is dated 31 December 2018 and it was chosen rather than measured. of Rs 100.00/- on 31 December 2018. Compounding those 72 changes forward gives a price for every month, ending at Rs 187.4539/- to the paisaThe hundredth part of one rupee, and the smallest step an ordinary amount of money is written down to.. The prices are one reading of the record and the monthly changes are the other. Neither is more real than the other and both are printed below.
Some of the numbers below are quoted to six decimal places, so where they come from is worth stating first. The record follows a construction rule printed further down, and a check script kept beside these notes rebuilds it from that rule and asserts every quantity. Every number below was rebuilt from the 72 rows during the drafting. The eight figures that also appear on the notes covering differencing were reconciled against them one by one. Two answers to one question makes both answers useless.
What does a series being stationary actually claim?
Three separate things, and all three have to hold at once. Most arguments about the word stationary turn out to be arguments about which of the three conditions somebody had in mind. Taking them in order:
The first condition is that the average does not depend on when it is measured. Any stretch of the record has an average, a different stretch has its own, and the two answers should agree except for the wobble expected from having only so many observations.
The second condition is that the spread does not depend on when it is measured either. How widely the observations scatter about their own average has to be the same quantity early in the record and late in it.
The third condition is the one people skip, and it is about pairs rather than about single observations. What two observations have to do with each other may depend on how far apart they sit and on nothing else. Both pairs sit one month apart, so January against February must be the same kind of relationship as November against December. Where in the record the pair sits is not allowed to matter.
Here is the everyday version, and it is worth holding on to. A vegetable seller works a fixed spot outside one office gate. She takes about the same amount in any given week, the good days beat the bad days by about the same margin in any given month, and a Tuesday says about as much about the Wednesday after it in March as it does in September. Nothing underneath her stall is moving. A second seller has moved to a busier gate halfway through the year. Every one of those three sentences breaks, and no average taken across the whole year describes either half of it.
| The condition | What it is a statement about | How the failure shows |
|---|---|---|
| The average does not depend on when it is measured | The centre of the record | Split the record and the two halves centre in different places |
| The spread does not depend on when it is measured | The scatter about that centre | Split the record and one half is visibly wider than the other |
| What two observations have to do with each other depends only on their distance apart | Pairs of observations rather than single ones | The same one month gap behaves differently early in the record and late in it |
A record satisfies the first two conditions cleanly: both halves centre on the same figure and both scatter by the same amount. Is that enough to call it stationary?
What does Non-Stationarity look like on this record?
Read the six year record as a price and the first condition fails outright, with nothing subtle about it. The 72 prices average Rs 115.0525/- across the first three years and Rs 190.4591/- across the last three, a distance of Rs 75.4066/- between two figures that are supposed to be estimating the same thing. Add the first 36 prices and divide by 36, then do the same to the last 36. The two averages are the whole calculation, and they need no assumption about anything.
Non-Stationarity is worth naming carefully. The failure above is not the same failure as a noisy record. A noisy record scatters widely about a centre that stays put. Here the centre itself is somewhere different depending on which three years are asked about. The second stretch does not merely contain some higher prices; its average is 1.6554 times the first stretch's average.
The consequence is what makes Non-Stationarity more than a curiosity: every figure computed across all 72 months blends two situations together, and the result describes neither one. The average price across all 72 months is Rs 152.7558/-. The addition behind that figure is beyond dispute, and no stretch of this record ever settled there. Read the record as a change instead and the same split gives 1.00 per cent and 1.00 per cent, to the last decimal place. A centre that stays put looks exactly like that.
The six year record averages Rs 115.0525/- as a price over its first three years and Rs 190.4591/- over its last three. Which of the three conditions has that observation settled?
What is a Unit Root, and why does it matter so much?
A series has a Unit Root when this period's value is last period's value plus whatever arrived this period, with no pull back at all towards any level. Written as a sentence rather than as notation: the whole of last month is carried into this month, and then something is added to it. There is no level the series is measured against, so nothing at all is subtracted for being too high or added for being too low.
The reason the term exists at all follows straight from that: in a series with a Unit Root, every movement is permanent. A month that came in high does not get undone the following month, or the month after, or ever. The high month is carried forward in full and it is still sitting inside the level three years later. Such a series has nowhere to come back to, so it wanders, and given enough months it will wander a long way from wherever it started.
The household version is the balance in a savings account against the electricity bill. The balance has a Unit Root: a good month leaves it permanently higher, and there is no force in the account pulling it back to some correct balance. The electricity bill does not: the bill is tethered to what the house actually uses, so a hot month runs high and the next mild month runs lower again. One quantity accumulates its history and the other one forgets it.
The six year record read as a price behaves like the account balance, and its one month autocorrelationOne number for how much a column resembles itself after being slid back a fixed number of rows. How it is worked out, and what it can and cannot detect, is covered separately. of 0.9650 is what that looks like from a distance. A reading sitting that near to one says each price is barely distinguishable from the one before it. Carrying the whole of last month forward produces precisely that reading. Read as a change the same record gives 0.2011. Most of what a month does is not inherited from the month before it.
A series has a unit root. One month comes in unusually high. What has happened to that month three years later, and why does it matter?
Stationarity vs Non-Stationarity: what changes when the same 72 observations are read two ways?
Everything that matters, and none of the data. The two columns below hold the identical 72 observations of the six year record, and every difference between them comes from which reading was taken rather than from any figure being different. Read the card straight down. The left column is a series with no level to come back to. The right column is a series with one.
The bottom row is the one to sit with. One record produced both columns, so nothing in the left column is a property of the Nakshatra unit and nothing in the right column is either: both are properties of the reading somebody chose. A colleague who hands over a column of numbers has already made that choice, usually without mentioning it, and every figure computed afterwards inherits it.
Recall check. What are the one month memory readings on the six year record, taken as a price and as a change?
How to Check a Time Series for Stationarity, and what do the three checks say here?
Three checks, run in ascending order of how much machinery they need. The cheapest check needs no assumptions and the most expensive puts a number on how sure the answer is, so all three are worth running.
Check one: split the record and compare the two averages
Divide the rows into two halves, take the average of each, and set the two figures beside each other. On the six year record read as a price they are Rs 115.0525/- and Rs 190.4591/-. Read as a change they are 1.00 per cent and 1.00 per cent. The split needs nothing except addition, it makes no assumption about how the observations were produced, and it is the check to reach for first every single time.
Check two: read the memory at one month
Slide the column back by a single row and ask how well it lines up with itself. The six year record answers 0.9650 read as a price and 0.2011 read as a change. Carrying the whole of last month forward is exactly what produces a reading close to one, so a reading close to one is the signature of a series with no level to come back to. Two cautions come with this check. The first is that it is a shortcut rather than a proof: a reading near one is strong evidence and not a verdict. The second is stricter, and it is a caution about a different question altogether. An autocorrelation reading is not a test for a repeating calendar pattern. Shift this very record back by twelve rows instead of one and it reads 0.0022, effectively nil. The calendar pattern sitting inside the record is large and real. Seasonal adjustmentThe operation that lifts a record's twelve recurring month by month values out of it, leaving behind whatever the calendar did not cause. Covered separately, along with how those values are found. is tested by grouping the observations by calendar month, and that is covered separately.
Check three: run the pull back regression
The pull back regression is the check with a number attached, and it is one regression run twice. Take a column. Work out its month on month change. Then fit that change against the column's own previous value. A record of 72 rows gives 71 such pairs. A series that returns to a level pulls back harder the further it has strayed, so the coefficient comes out clearly negative. A series with a unit root has nothing to pull it, and the coefficient sits at about zero.
Read as a price, the six year record puts the coefficient at minus 0.026109 and its t at minus 1.0683. The standard errorHow much a fitted figure would move about from one set of observations to another, purely because a different set landed in the calculation. Settled earlier and used here as it stands. behind that t is what turns the coefficient into a verdict. A t of minus 1.0683 says the coefficient sits about one standard error from zero, and one standard error is nowhere near far enough to claim the price series pulls back at all. Read as a change the same record gives minus 0.777838 with a t of minus 6.2898.
The two coefficients say something concrete once turned round. On the change series, one plus minus 0.777838 leaves 0.222162, so a movement has about 22.22 per cent of itself surviving one month later and about 1.10 per cent surviving three months later. The change column is a series with a short memory and a level. On the price series, one plus minus 0.026109 leaves 0.973891, so a movement still has 72.80 per cent of itself twelve months later. Even 72.80 per cent overstates the pull back. The t figure says the coefficient cannot be told apart from zero in the first place.
When the three checks disagree, the split is usually the one to trust. Among the three it alone assumes nothing whatever. The memory reading is a summary that can be pulled about by a single strange stretch, and the regression rests on the fitted line behaving itself. Splitting a record and averaging both halves rests on nothing at all. The other two checks are not therefore worth skipping: the split is the tie breaker rather than the whole answer.
The pull back regression gives minus 0.026109 on the price column and minus 0.777838 on the change column. What is being fitted, and what does a coefficient sitting at about zero say?
Worth settling before the panel below answers it. With the split moved to month 12 instead of month 36, will the change column's two averages pull apart as far as the price column's do?
One handle for where the record gets split, with both readings redrawing underneath
The split drags from month 12 to month 60, and a click anywhere on the chart puts it there. The two panels redraw, each with a level line at its own left average and another at its own right average, and the four averages update as the split moves. The panel opens at a split after month 36, December 2021: the price sides read Rs 115.0525/- and Rs 190.4591/- and the change sides read 1.00 per cent and 1.00 per cent. The selector underneath changes the unit the two distances are reported in. The unit matters more than it sounds like it does.
One reading on that panel is the honest awkward case, and it deserves saying out loud. With the split pushed all the way to month 60, the change column's two sides come out 1.50 per cent and minus 1.50 per cent. The two sides sit 3.00 percentage points apart, plainly not zero. The awkward reading is not the argument falling over. A right hand side of only 12 months, taken from a column whose scatter across those years runs 6.4031 per cent, wobbles about by roughly that much on its own. Switch the panel to standard errors and the two readings never meet: across all 49 splits the panel allows, the price sides sit at least 5.0355 standard errors apart and the change sides never exceed 1.6466. Two distances in different units cannot be compared at all until something puts them on one scale, and that is what the second setting of the selector is for.
What did differencing remove from this record, and what did it leave behind?
Moving from the price column to the change column is one step, and this is where most treatments of the subject quietly overclaim. Taking the change instead of the price took out the unit root and took out the drifting average. Two further problems stayed exactly where they had been. Both halves of that sentence are visible in the comparison card above and in the panel above it.
Two things went. The price column's memory of 0.9650 became 0.2011, its pull back coefficient moved from a figure indistinguishable from zero to minus 0.777838, and its two half averages of Rs 115.0525/- and Rs 190.4591/- became 1.00 per cent and 1.00 per cent. The first condition and the unit root were both genuinely dealt with.
Two things stayed. The change column still carries a repeating calendar pattern in its average, so its January months and its June months do not average the same figure. The third condition is under strain. And its scatter is visibly wider in the last three years than in the first three. The second condition is failing on its own account. Differencing addressed the drifting average and the permanent movement, and the other two conditions have to be checked on their own afterwards. A claim that differencing makes a series stationary, full stop, would have overclaimed on this very record, where the arithmetic is sitting in plain view.
On the six year record, taking the change instead of the price removed the unit root. Name what it did not remove.
Why does the assumption exist at all?
Because almost every figure anybody computes on a record assumes there is one fixed thing being estimated. An average assumes there is a centre to find. A spread assumes there is a scatter to measure. An interval assumes both, and then adds a claim about how precisely the first one has been pinned down. If the centre itself is moving, the average is estimating a quantity that was never there to be estimated, and the interval around it is describing a precision it does not have.
The general statement is easy to nod along to. Here is the specific one, on the record itself. The average price across the whole six year record is Rs 152.7558/-, it is arithmetically correct, and it describes neither the first three years nor the last three. The whole record average sits exactly halfway between Rs 115.0525/- and Rs 190.4591/-, and it sits there for a boring reason: the two halves hold the same number of months, so the whole record's average has to be the plain average of the two. The price path crosses that figure once in 72 months and spends the rest of the record somewhere else.
The everyday version is a household with two earners where one loses work halfway through the year. The average monthly income across the year is a real number, correctly calculated, and it is not what the household lived on in either half. Anyone budgeting off it plans for a household that never existed. The household is the whole of the assumption in miniature, and the same risk is why so many methods start by insisting on stationarity.
The failure: a correct average, an interval around it, and a record that was never sitting still
Somebody working on the six year record computes the average price across all 72 months and reports Rs 152.7558/-. Then, being careful, they put an interval around it. The 72 prices scatter by Rs 43.4565/-, and over 72 observations that scatter gives a standard error of Rs 5.1214/- and a 95 per cent interval of Rs 142.7181/- to Rs 162.7936/-. Every step of that is correct arithmetic.
The interval is worse than the average it surrounds, and here is the measurement. The interval is Rs 20.0755/- wide, and Rs 20.0755/- is 26.62 per cent of the Rs 75.4066/- gap between the two stretches it claims to be describing. Only 7 of the 72 months sit inside it. The narrowness is not precision, it is the arithmetic of a large number of observations being applied to a quantity that was never sitting still long enough to be estimated.
The cost is not the wrong number. The cost is what happens next: the following person treats Rs 152.7558/- as the level the Nakshatra unit belongs at, and every comparison they draw against it carries an assumption nobody ever stated out loud.
A working habit prevents the failure, not a technique. Before anything gets computed across a whole record, cut it in two and compute both halves. If the two halves disagree, either say so in the same sentence as the whole record figure, or do not publish the whole record figure at all.
The average price across the whole six year record is Rs 152.7558/-, and the addition behind it is correct. What is wrong with quoting it?
What does a working analyst do with a record that will not sit still?
Four habits, and the last one is the only one that reliably prevents the failure above.
First, work with the change rather than the level wherever the question allows it. A lender assessing how much a small business's monthly takings bounce about wants the month to month change, not the running balance. The balance carries every past month inside it and the change does not. Converting a level column into a change column, and picking between the ways of doing it, is covered separately.
Second, say which reading every figure came from, every single time. The sentence the average is Rs 152.7558/- and the sentence the average monthly change is 1.00 per cent describe the same 72 observations and mean entirely different things. A figure travelling without its reading attached is a figure somebody will misread within a week.
Third, split the record and compute both halves before quoting anything computed across the whole. The extra split costs one addition and one division. An analyst who does this routinely will catch a moving centre before it reaches a report, and an analyst who does not will publish the whole record figure and find out later.
Fourth, when the three checks disagree, report all three rather than the one that happens to agree with the conclusion already reached. Reporting all three is a working habit rather than a technique, and it is the habit that actually stops the failure. The failure is never that somebody could not do the arithmetic. Note also that splitting a record this way is a description of months that have already happened, so it carries no lookaheadLetting a figure dated to one month be built out of information from a month that had not happened yet. It is a dating fault inside a calculated column rather than a failure of judgement, and it is covered separately. problem: no calculation looks forward from a date the record had not reached.
Three checks disagree about a column somebody has handed over. The split says one thing, the memory reading says another and the pull back regression sits in between. Which one should be leaned on, and why?
Boundaries. Converting a level column into a change column, and choosing between a difference in rupees, a percentage change and a log differenceA change worked out by subtracting the logarithm of one level from the logarithm of the next. It is one of the choices on offer when a level column is converted, and choosing between them is covered separately., is covered separately, on notes that print the same comparison figures as this guide. The name for a spread that changes partway through, and the cost of quoting one figure for it, is covered separately. The nature of a repeating calendar pattern, and the arithmetic that lifts one out, sits elsewhere as well, and so does telling a steady direction apart from the wobble around it. A figure recalculated inside a rolling windowA short block of consecutive rows that moves forward one row at a time, with the same calculation redone inside it at every stop. Covered separately, including why the redone figure arrives late. that slides along a record is covered separately too. Reading the same record at another frequencyHow often a record records: monthly here, and daily, weekly or yearly elsewhere. What changes between those readings is covered separately. than monthly is a separate question again. Methods that require a stationary series are covered separately.
Why does an account this precise name nobody?
Look down the third column of the ledger and the same answer appears on every row. The repetition is the finding rather than a gap somebody meant to come back and fill. Cutting a column of numbers in two and comparing the two averages is something done to data already in hand: nobody issues it, nobody amends it, and there is no office anywhere whose agreement would make Rs 115.0525/- a better answer for the first 36 prices than the addition already makes it. The pull back check is the same story one step further along. The check is the fitted line covered under trend estimation, run on two columns built out of the record itself.
One convention does have to be named. The memory reading is not unique, and two recipes in common use disagree. Every autocorrelation in this guide centres on one average for the whole column and divides the paired products by that column's total sum of squared deviations. Treat the 71 overlapping pairs as two separate columns and correlate them the ordinary way instead, and the same data reads 0.9790 as a price and 0.2114 as a change. The conclusion does not move, the decimals do, and a reading quoted without its recipe is a reading somebody will fail to reproduce.
None of the tidiness above travels. The record was written so that both halves of its change column land on exactly 1.00 per cent. The split check therefore comes out at precisely zero rather than merely small. Data gathered from the world does not oblige anybody like this. Run the same split on a gathered record and the two halves come back close but never equal. A reading like that is harder to make and more honest. The method survives the move to real data. The exactness belongs to this record alone.
| What this guide prints | What was run to get it | What breaks if it is wrong | Recomputed on |
|---|---|---|---|
| The 72 monthly changes, and the price path from Rs 100.00/- to Rs 187.4539/- | A steady 1.00 per cent a month, plus twelve repeating calendar values that add to nothing, plus an irregular part, compounded forward from the opening mark | Everything. Both columns in this guide are readings of this one record | 21 August 2026 |
| Rs 115.0525/- and Rs 190.4591/-, and the distance of Rs 75.4066/- | The first 36 prices added and divided by 36, then the last 36 the same way, then one subtraction | The split check, the comparison card and every reading on the panel | 21 August 2026 |
| 0.9650 and 0.2011 at one month | Each column matched against a copy of itself shifted back one row, on the recipe named above | The memory check, and the row shared with the notes on differencing | 21 August 2026 |
| Minus 0.026109 with a t of minus 1.0683, and minus 0.777838 with a t of minus 6.2898 | Two fitted lines, each of one column's month on month change on that column's own previous value, 71 pairs each | The pull back check and the branch diagram that reads off it | 21 August 2026 |
| Rs 152.7558/-, and the interval of Rs 142.7181/- to Rs 162.7936/- | All 72 prices added and divided by 72, then their scatter of Rs 43.4565/- divided by the square root of 72 | The failure. If the interval width is wrong, so is the claim that it covers a quarter of the gap | 21 August 2026 |
The Nakshatra unit, the vegetable seller, the savings account and the household with two earners are invented.
Educational material. Not advice on any investment, tax, budget or market position.
