Walk-Forward Analysis: Rolling the Test Through Time
Walk forward analysis refits the rule as the record rolls on, and reads each fit only on the months just after it. Eight fits on the six year record, each made on 24 months and read on the next six, pool to 27 of 45 calls. The setting picked moves three times, and the fold readings swing from 2 of 6 up to 5 of 5.
The temptation attached to this subject is enormous, and it is worth naming before anything else. Rolling the fit through time is often described as the grown up version of holding a stretch back once, the thing to graduate to once a single split has been shown to be fragile. The graduation story is half right, and the wrong half is doing the damage. Refitting on a schedule produces more readings than a single split does. More readings are not better readings. Below, one unchanged rule calls five of the five months it can score in one half year and two of six in another, and an entire apparatus of eight separate fits shifts a count by exactly one month.
The record was built from three stated parts, and the rule tested on it has thirteen stated settings. Every count below is a consequence of those two definitions plus one counting convention, and a reader with a spreadsheet can rebuild all of it in an afternoon.
What does walk forward analysis do that a single split does not?
Start outside a bus depot, at a tea stall. Every Monday morning the person running it goes back through the last eight weeks of takings and works out which hour of the day was busiest, then arranges the week around that hour. One week has dropped off the back and a new one has joined the front, so next Monday they do it again on a fresh eight weeks. Keep that up for a year and something worth noticing happens. The stall keeper does not end up holding a better estimate of the busiest hour. The year leaves fifty two estimates behind, and the fifty two are the actual finding. Watching one answer wander from six in the evening to eight in the morning and back tells the stall keeper exactly how much any one of those answers was ever worth.
The version set out here changes almost nothing. A single split asks one question one time: does a setting chosen on this part of the record survive on that part? Walk forward analysis asks the same question eight times, at eight points along the same record, and each of the eight answers comes from a fit that could see only months lying before it. No fit is ever allowed to read a month it is about to be read on. Eight fits do not buy accuracy. Eight fits buy eight answers where there was one, and eight answers are what reveal what the single answer was worth.
Everything below is three steps repeated, so the moving parts are worth naming once, in order. Take a stretchA run of months sitting next to each other, lifted out of a longer record and handled as though it were a small record in its own right. of 24 months and score every one of the Ashwin rule's thirteen settings over it, keeping whichever one calls the most months right. Carry that setting, untouched, onto the six months immediately after the stretch, and count. Then slide the whole arrangement six months down the record and do it again. Eight repetitions of that run cover January 2019 to the end of December 2024, the whole six year record of the Nakshatra unit. The unit and the Ashwin rule are both inventions, built so that a method has numbers to work on.
Two things stay fixed while all that moves, and confusing either of them with the fit is the commonest way to misread a walk forward result. The rule stays fixed: read the finished month's change and call the next month up when that signalThe number a rule consults before it calls. Here that is the change over one month which has already ended, and the word ended is doing real work there. sits above a thresholdThe number a rule holds its reading up against before deciding which way to call. Shift it by a single step and the same months turn into a different set of calls., down otherwise. The grid stays fixed too: thirteen settings, every one of them a whole number, running in single steps from a floor of minus 6.00 per cent to a ceiling of 6.00 per cent. The only thing refitting changes is which of those thirteen the fit reaches for this time, and how often it changes its mind is a finding rather than a nuisance.
What does walk forward analysis give that a single split does not?
What do the eight folds on this record look like?
A scheme described in the abstract sounds far more impressive than the same scheme with its actual numbers attached, so here is the whole exercise with its numbers attached. Each fit and the six months it is read on are called a fold, and the eight of them keep the numbers one to eight throughout. Fold one fits on January 2019 to December 2020 and is read on January to June 2021. Fold eight fits on July 2022 to June 2024 and is read on July to December 2024. In between, the arrangement simply slides six months at a time, and the eight test stretches cover January 2021 to December 2024 end to end with no gap and no overlap.
The picture is already hinting at a counting care, and the care shows up in the headline figure. Six of the 72 months in the six year record stood exactly still, and a month that stood still is dropped from the counting rather than scored. A rule obliged to say up or down has nothing to be right or wrong about when nothing happened. Three of those six still months land in 2021, and 2021 is where the first two test stretches sit. So fold one is read on four scoreable months rather than six, and fold two on five. The eight denominatorThe count sitting under a share. Two shares built on counts of unequal size carry unequal amounts of evidence, so averaging them treats a small one as though it were large.s are not all the same, and that single fact decides which arithmetic the headline figure is allowed to use.
| Fold | Fitted on | Setting picked | Read on | Calls right |
|---|---|---|---|---|
| 1 | Jan 2019 to Dec 2020 | minus 1.00 per cent | Jan to Jun 2021 | 3 of 4 |
| 2 | Jul 2019 to Jun 2021 | minus 1.00 per cent | Jul to Dec 2021 | 5 of 5 |
| 3 | Jan 2020 to Dec 2021 | minus 1.00 per cent | Jan to Jun 2022 | 4 of 6 |
| 4 | Jul 2020 to Jun 2022 | minus 1.00 per cent | Jul to Dec 2022 | 2 of 6 |
| 5 | Jan 2021 to Dec 2022 | 1.00 per cent | Jan to Jun 2023 | 4 of 6 |
| 6 | Jul 2021 to Jun 2023 | minus 1.00 per cent | Jul to Dec 2023 | 2 of 6 |
| 7 | Jan 2022 to Dec 2023 | minus 6.00 per cent | Jan to Jun 2024 | 3 of 6 |
| 8 | Jul 2022 to Jun 2024 | minus 6.00 per cent | Jul to Dec 2024 | 4 of 6 |
| Pooled, hits summed over scoreable months summed | 27 of 45 | |||
Twenty seven calls out of forty five scoreable test months, or 60.00 per cent. Two different figures hide behind the same word, so what the 60.00 per cent means needs saying before anyone quotes it. The pooled figure is a sum of hits divided by a sum of months, so a fold that contributed six months counts for six and the fold that contributed four counts for four. The pooled figure is not the average of the eight fold readings. The two arithmetics disagree on this record, and the practitioner block below shows by exactly how much and why the gap always leans the same way.
Two of the eight folds are read on fewer than six months. Why, and what does that do to a headline computed as the average of the eight fold readings?
What happens to the threshold as the fits roll forward?
Down the setting column of that table, the honest finding of this whole exercise sits in plain sight. Across eight fits the search reaches for minus 1.00 per cent five times, for 1.00 per cent once and for minus 6.00 per cent twice. Three of the thirteen available settings sounds reassuringly stable. The placement of the changes says otherwise. The pick moves at three of the seven steps between one fold and the next, and the three moves run consecutively, at the fourth, fifth and sixth steps. The pick holds still through the first three steps, changes its mind three times in a row, and then sits at the very bottom of the grid for the last two folds.
There is a second movement in the same table that is easy to miss because it is not in a column: what the fit manages on its own fitting window as that window slides. On the first three folds the chosen setting calls 18 of the 20 scoreable fitting months, or 90.00 per cent. By fold five it calls 15 of 21, or 71.43 per cent. By fold eight it calls 13 of 24, or 54.17 per cent. The fit is not getting worse at its job. The fit is being handed a harder stretch of the same record, and it says so honestly on the only months it is allowed to see.
A number that has to be looked up again every six months is a description of the last two years rather than a property of the record. If minus 1.00 per cent were something true about the Nakshatra unit, the fit would keep landing on it whichever 24 months it was handed, and by folds seven and eight it does not: it has walked all the way to the bottom of the grid. A single split simply cannot produce that finding. One fit gives one setting, and with nothing to compare it against, that setting looks exactly as stable as anyone would like it to look.
The fit reaches for three different settings across eight folds. What does a setting that has to be looked up again every six months say about the record it came from?
The control below steps through all eight folds in order. Does the setting the fit picks settle down as the record grows longer behind it?
Slide the fit along the record, one fold at a time, and watch three things move at once.
One control, and it moves one thing: which of the eight folds is on screen. Everything else is recomputed from the record at each setting. The strip along the top slides the 24 month fitting window and its six month test stretch down the record, with the unmoved months marked so that the denominator can be seen to shrink on the first two folds. The ladder in the middle scores all thirteen settings on that fold's fitting window alone and marks the one the fit keeps, with a dashed outline left on the previous fold's pick so that the moment it changes its mind is visible. The row of eight bars at the foot fills in fold by fold, and the line running across them is the running pooled figure through the fold on screen. The 60.00 per cent finally comes from that line. The panel opens on fold three, whose fitting window ends on December 2021, the same date the single split cuts at, and whose reading is 4 of 6.
Educational illustration. Every fold fits on 24 calendar months and is read on the six months immediately after, and no fit ever reads a month outside its own fitting window. A month that stood still is dropped from every count. Two of the eight folds are therefore read on fewer than six months.
One unchanged rule is about to be read eight times, six months at a time. How far apart should the highest and the lowest of those eight readings be expected to sit?
What is a six month reading actually worth?
The sequence is more persuasive spoken than tabulated, so line the eight readings up in order and read them out loud. Three of four. Five of five. Four of six. Two of six. Four of six. Two of six. Three of six. Four of six. As shares those run 75.00, then five calls out of five months, then 66.67, 33.33, 66.67, 33.33, 50.00 and 66.67 per cent. The reading changes at every single one of the seven steps. Not most of them. All seven.
The honest reading of the two extremes is not the flattering one, so sit with them for a moment. Fold two calls all five of the months it can score. Fold four, two folds later, calls two of its six. Nothing about the rule differed between those two folds; the setting was minus 1.00 per cent in both. The difference was which six months of the record happened to fall on the far side of the fitting window. An unchanged rule read in six month slices produced five calls out of five months and two calls out of six within a year of each other, so a six month figure quoted on its own carries almost nothing at all.
The lesson applies well outside walk forward analysis, and it is the part worth carrying away. Somebody produces a figure and says it covers half a year. Before asking how good the figure is, the question to ask is how many months went into it, and then what the same thing read in the half year on either side. A short stretch does not produce a small error around the truth. A short stretch produces a reading dominated by which months it happened to contain, and the eight bars above are that statement drawn to scale on a record where nothing at all was hidden.
Does refitting twice as often change the answer?
An obvious objection to everything above is that six months is simply too fidgety, and that a slower schedule would settle things down. Test it. Roll every twelve months instead of every six: fit on 24 months, read on the next twelve, then slide a whole year. The slower schedule gives four folds rather than eight, and the four fits reach for minus 1.00 per cent, minus 1.00 per cent, 1.00 per cent and minus 6.00 per cent, a different sequence of picks from the eight fold version. The four folds are read on 9, 12, 12 and 12 scoreable months, and the first is nine rather than twelve for the same reason as before: three months of 2021 did not move.
Now the striking part. The two schemes are read on exactly the same 45 scoreable months, January 2021 to December 2024, and they both pool to 27 of 45. Not close to each other. Identical. The equality is arithmetic rather than luck, and the difference between saying so and merely noticing it is the difference between understanding this record and being charmed by it.
The reason sits in the months themselves. Walk both schedules month by month and there is exactly one stretch where they hold different settings at all: July to December 2023, where the six month schedule holds minus 1.00 per cent and the twelve month schedule holds 1.00 per cent. Any other reading sits above both settings or below both, so two whole number settings one step apart can only disagree about a month whose reading is exactly 0.00 per cent or exactly 1.00 per cent. The six readings driving those six calls are minus 6.00, 6.00, 8.00, 6.00, minus 4.00 and 11.00 per cent, and not one of them is 0.00 or 1.00, so the two schedules make an identical call on all 45 months and the equality was never in doubt.
Notice what that does to the objection. Refitting twice as often was supposed to be the dial that mattered, and on this record it moved nothing whatever, not because a slower schedule is wiser but because the settings the two schedules picked were close enough that no month in the record could tell them apart. Reach for the general lesson carefully: this is a fact about a grid of whole numbers meeting a set of readings, and a finer grid or a different record could easily behave otherwise. The habit is what generalises. When two schemes agree, go and find the months where they held different settings, and check whether anything there could ever have separated them.
Refitting twice as often did not change the pooled count. Is that luck, and what would settle the question?
How does walk forward analysis compare with the single split over identical months?
The comparison between the two schemes has to be set up carefully, or it shows the opposite of what it should. The eight fold exercise is read on January 2021 to December 2024 and pools to 27 of 45, or 60.00 per cent. The single split, settled separately, is read on January 2022 to December 2024 and reads 18 of 36, or 50.00 per cent. Ten points apart, in favour of refitting. The two figures are computed on different months, so reaching for that ten point gap is the mistake.
So align them. Take the last three years, exactly the 36 scoreable months the single split is read on, and refit three times across them: fit January 2020 to December 2021 and read 2022, fit January 2021 to December 2022 and read 2023, fit January 2022 to December 2023 and read 2024. The three fits reach for minus 1.00 per cent, 1.00 per cent and minus 6.00 per cent, and they call 6, 6 and 7 of their twelve months. Nineteen of thirty six, or 52.7778 per cent, against the single split's eighteen of thirty six, or 50.00 per cent. Three fits, three searches of thirteen settings each, an entire apparatus of rolling and refitting, and the count moved by one month.
With the two lengths beside each other, the argument finishes itself. The single split showed a fall of 39.6552 points when a setting chosen with every outcome visible was carried onto months nobody had seen. Refitting on a schedule, done properly and three times over, recovers 2.7778 of those points. Walk forward analysis is not a repair for a fitted result. What it produces is more readings, not better ones, and the extra readings are the entire reason to do it. Eight readings that run from two calls out of six months up to five out of five say something a single reading of 50.00 per cent cannot: that any one reading from this record, taken alone, was capable of being almost anything.
The comparison that gets made anyway, and what it costs
Somebody writes up the exercise. Walk forward analysis reads 60.00 per cent, they note, against the single split's 50.00 per cent, so rolling the fit through time repaired the result by ten points. The arithmetic in that sentence is fine. Both figures are correctly computed. And the conclusion is false anyway, for a reason nobody will ever uncover by rechecking either count, however carefully or however often they recheck it.
The two figures are not computed on the same months. The walk forward exercise is read on January 2021 to December 2024, or 45 scoreable months. The single split is read on January 2022 to December 2024, or 36. The walk forward figure therefore contains a full year, the whole of 2021, that the single split never sees, and 2021 sits inside the quieter first half of the record where the months travelled less far from their centre. Line the two schemes up on identical months and the ten point gap becomes 19 of 36 against 18 of 36, a difference of one month.
The cost is not the wrong number; it is the belief that follows it. The person now thinks refitting fixes overfitted results, and carries that belief into the next study, and the one after, in every one of which the two windows will differ slightly and the comparison will be just as broken and just as invisible. The fix is one line of bookkeeping: before comparing any two testing schemes at all, write down the months each of them is read on and make them the same set. The bookkeeping costs a minute, it is skipped constantly, and it is the difference between a comparison and a coincidence.
Over identical months, refitting three times read 19 of 36 against a single split's 18 of 36. What does that one month settle about walk forward analysis?
What should be asked of a walk forward result?
A walk forward figure arrives from somebody else. Seven questions, in the order that saves the most time, and every one of them is answerable from the write up if the write up was done properly.
One, how long is each fitting window and how long is each test stretch? Here, 24 months and six. Eight folds on a short record and eight folds on a long one are not the same exercise, so without both lengths the count of folds says nothing.
Two, how many folds, and how many scoreable months in each, under which counting conventionA rule of counting that is settled by decision rather than by evidence. Two people can hold different ones perfectly honestly, so a figure means little until the one that produced it is known.? Here, eight folds and 45 scoreable months, under a convention that excludes a month which did not move. Change that one convention and every count above moves.
Three, is the headline a sum over a sum, or an average of the fold readings? This question changes headline figures more often than anyone expects. Sum the hits and sum the months and this record gives 27 of 45, or 60.0000 per cent. Average the eight fold readings instead and the same record gives 61.4583 per cent. The averaging hands the four month fold exactly the same weight as a six month one. A gap of 1.4583 points, from nothing but a choice of denominator, and the choice is almost never stated.
Four, did every fold read the same signal, and what does the baselineWhat something holding no information at all would score on the same months. A figure with nothing sitting beside it has nothing to be better or worse than. read over the same test months? Here the signal never changed, and the baseline of calling every month up without reading anything scores 25 of the 45 test months, which is 55.5556 per cent. The pooled 60.00 per cent is therefore two months of work ahead of doing nothing at all. Ask this fold by fold as well. Fold three is read on six months of which five rose, so calling every one of them up would have scored 5 of 6 there against the rule's 4 of 6.
Five, how many distinct settings did the fit pick, and did the pick move? Three settings across eight fits, moving at three of the seven steps and landing at the bottom of the grid for the last two. A pick that never moves is worth asking about too. A frozen pick can mean the grid is too narrow to move within.
Six, what does the same record give under a single split over exactly the same months? Here, 18 of 36 against the refitted 19 of 36. If nobody has run that comparison, the walk forward figure has nothing to be judged against.
Seven, what is the spread of the fold readings, not just their pooled figure? Two calls out of six months at the bottom and five calls out of five months at the top. The spread is the finding. The pooled figure is a summary of it, and summarising a spread that wide into one number and quoting only the number is how a walk forward result ends up being read as reassurance.
Somebody quotes a walk forward figure against a single split figure. What is the first thing to line up before comparing them, and what usually goes wrong when nobody does?
Where does walk forward analysis stop?
A backtest itself, the difference between a backtested figure and one produced afterwards, holding a search down before it starts and holding a stretch back once are all covered separately. A long run of tests and what it does to a threshold is covered separately too, as are the two distinct habits of picking the analysis, or picking the question, once the record is already in front of the analyst. Placing an order, paying a cost or holding a position sits outside this subject altogether and is covered separately.
Two limits inside the exercise itself deserve naming. Everything above is one rule on one invented record, so a fact such as the two refitting schedules agreeing exactly is a fact about this grid meeting these readings, not a general property of refitting schedules. And nothing above makes the Ashwin rule worth testing, worth running, or worth anything. The rule exists so that a testing scheme has something to be applied to, and on this record it mostly fails.
Where do these figures come from?
Three definitions stand in for a citation, and they are enough to rebuild every count above independently.
| What is needed | What it is | Where it was settled |
|---|---|---|
| The record | The Nakshatra unit and its six year record: 72 dated monthly observations, January 2019 to December 2024, mean 1.00 per cent, spread 5.00 per cent, opening mark Rs 100.00/- and closing price Rs 187.4539/- | Invented, and fixed in advance |
| The rule | The Ashwin rule and its thirteen whole number settings, minus 6.00 per cent up to 6.00 per cent in steps of one | Fixed in advance and left alone throughout |
| The counting convention | January 2019 has no month before it to read and six months did not move, so 65 of the 72 months can be scored, and each fold carries its own count | Stated above, and every figure depends on it |
The Nakshatra unit, its six year record and the Ashwin rule are invented.
Educational material. Not advice on any investment, tax, budget or market position.
