How to Document a Backtest So Somebody Can Rerun It
Three things are settled separately and are not rebuilt here: what it means to test a rule against a stretch of history, why a figure arrived at after choosing among options is a different kind of number from one arrived at without choosing, and what fitting on one part of a record and reading on another is for. Every count printed here comes from rerunning one calculation on one invented record of 72 monthly readings, and any word borrowed from finance is explained the first time it turns up.
The argument below rests on three things that can be checked rather than taken on trust. The first is the Ashwin rule, an invented rule settled earlier: it looks at how the Nakshatra unit moved last month, and when that movement clears a set thresholdA cut off. A reading either clears it or it does not, and which side of the cut off it lands on is what makes the rule say up rather than down. the rule calls the coming month up, and otherwise down. The second is the six year record itself, 72 monthly readings of the Nakshatra unit running from January 2019 to December 2024, every one of them made up so that a testing method could be watched from the outside.
The third is counting, and only counting. Nowhere below is a missing field called serious, or careless, or a matter of professional standards. Each one is measured instead, by counting how many different readings a write up lacking that field is still consistent with. Counting is the only currency in use, and counting is why the argument can be checked by anybody with the same 72 readings and an afternoon.
What is a backtest write up actually for?
The purpose of the document decides everything else. A backtest write up is adequate when a second person, holding only that document and the record, sits down and arrives at the same number. The standard stops there. Adequacy is not a matter of effort, or of how carefully the author thought, or of how long they have been doing this. Adequacy is a property of the document, and a document can be tested by handing it to somebody.
The shape is much easier to see away from finance. A recipe card that says bake until done can be followed by the person who wrote it and by nobody else. The person who wrote it knows that until done means the top has stopped wobbling, that this particular oven runs hot, and that the tin goes on the middle shelf rather than the top one. None of that is on the card. Handed to a neighbour, the card produces a different cake. Nobody was careless and nobody lied. The card was never a set of instructions. The card was a reminder, written by somebody who did not need reminding of the parts they left out.
A backtest write up fails in exactly the same way and for exactly the same reason. The author knows which months were counted, having been sitting there while counting them. The author knows that the months in which the Nakshatra unit finished where it started were dropped. At the time that seemed too obvious to write down: there is nothing to be right or wrong about in a month that stood still. The author knows that thirteen settings were tried before one was kept. None of that reaches the document unless somebody puts it there, and every one of those choices moves the number.
Notice what this standard refuses to do. The standard does not grade a write up out of ten, and it does not let a thorough document that is missing one field score nine. The second person either lands on the author's number or does not, and there is no partial credit in that test. The harshness is the useful part. Diligence is a question nobody outside somebody's own head can settle. The standard converts it into a question about a document, and anybody at all can settle that in an afternoon.
When is a backtest write up adequate?
Which six fields does a write up have to carry?
Six fields, and five of them are one line each. Two of the six describe the search rather than the run itself: the gridA written list of the settings a search will work through, decided up front, so that the list is itself a document rather than a recollection. of settings, and the count of attempts. One of them, the counting conventionA rule for counting, settled and put in writing beforehand, so that two people handed the same rows end up with the same total., describes a decision so small that people leave it out for being obvious. And the first of them, the record and its spanHow far a set of readings reaches: the date it opens on and the date it closes on., is the field most likely to be right. The record itself is usually sitting in the same folder. Here are all six, filled in for one honest run.
| Field | What it has to say | This run |
|---|---|---|
| 1. The record and its span | which readings, and between which dates | 72 monthly readings of the Nakshatra unit, January 2019 to December 2024 |
| 2. The rule, in one sentence | what it reads, and what it then calls | the Ashwin rule reads last month's reading and calls the coming month up when that reading clears the threshold |
| 3. The grid searched | every setting on the list, written before the search starts | thirteen whole number settings, from minus 6.00 per cent up to 6.00 per cent |
| 4. The split date | where the record was cut in two | the end of December 2021 |
| 5. The counting convention | how an awkward month is treated | a month in which the unit did not move is left out |
| 6. The count of attempts | how many settings were actually run | thirteen |
Field five is the one people are most surprised to find on a list of six, so look at what it does to the arithmetic rather than at how small it sounds. The record holds 72 months. January 2019 has nothing in front of it, so the rule has nothing to read and that month is never called, leaving 71. Six of those 71 are months in which the Nakshatra unit finished exactly where it started, and a call of up or down cannot be marked right or wrong against a month that did not move, so those six come out as well. The denominatorThe bottom half of a fraction. Twenty six right out of twenty nine and twenty six out of forty five are the same twenty six, and the bottom half is the whole reason they read so differently. is 65, and the only reason a second person knows it is 65 rather than 71 is that field five is in the document.
Field three and field six look like the same thing written twice, and they are not. The grid is the list of settings intended for trial, written down before any of them was tried. The count of attempts is how many were actually run. When a search goes as planned the two match, and on this run they both read thirteen. When a search does not go as planned they part company, and the gap between them is then the most informative line in the whole document. The gap records the settings that were added after somebody had already seen how the first ones came out.
Five of the six fields are one line each, and the sixth is a single word. The completed write up fits on a filing card with room to spare. The reason these fields go missing is never that they were hard to write.
There is time to write only one of the six fields before the analyst is pulled away. Which one, and what does it protect?
What does each missing field cost, measured in readings?
Now the measurement, the part worth the most. Take one honest run. The Ashwin rule at a threshold of minus 1.00 per cent, read over the first three years of the record with the unchanged months left out, calls 26 of 29 months correctly, a reading of 89.6552 per cent. Nothing about that run is wrong. Nothing about it is hidden. The figure is exactly what the arithmetic gives, and it would give the same figure to anybody who ran it that way.
So the question is not whether the run is honest. The question is how much of it a reader can pin down from the document. Take the same run and ask a different thing: given only what the write up actually says, how many other analyses of the same record would also have been fair to run, and what readings do those give?
The menu is short and every item on it is defensible. Five stretches of the record a person might reasonably read the rule over. Two ways of handling a month in which the unit did not move, either dropping it or counting it as a rise. And two signalsWhatever the rule looks at immediately before it decides. Here that is a single number, assembled only from months the record has already finished with. the rule might read, either last month on its own or the average of the last three. Five times two times two is twenty. Nobody has to bend anything to arrive at any one of the twenty; each is a choice somebody would make without blinking, and each is a choice this write up either records or does not.
The three recoverable fields are about to be switched on one at a time. Before the control moves: does the band of readings shrink by roughly a third at each step?
Watch the band close as each field goes in
One run, one record, one threshold, held completely still. The only thing that moves is how much of the run the document actually says. Every tick on the scale is one of the twenty analyses; the pine ticks are the ones this document still allows and the pale ones have been ruled out.
Held still throughout: the record, the rule, the threshold of minus 1.00 per cent, and the run itself. Only the document changes.
Educational illustration. The menu of twenty analyses is the one set out above and no larger, and the threshold is fixed at minus 1.00 per cent throughout so that the band measures documentation and nothing else.
Recording the fields one at a time gives the ladder below, and those four rows carry the whole argument.
| Fields recorded | Analyses still allowed | Band of readings | Width |
|---|---|---|---|
| nothing recorded | 20 | 50.0000 to 89.6552 per cent | 39.6552 points |
| the stretch | 4 | 66.6667 to 89.6552 per cent | 22.9885 points |
| and the counting convention | 2 | 70.3704 to 89.6552 per cent | 19.2848 points |
| and the signal | 1 | 89.6552 per cent only | 0.0000 points |
The band narrows at every step and closes completely at the last, and that is what turns the six fields from an opinion into a list. The top row is the least believable line above, and it deserves a second reading. A write up that names the record, states the rule and prints 89.6552 per cent, and says nothing about the stretch, the convention or the signal, is exactly as consistent with a run that reported 50.0000 per cent. Both documents are true. Both describe real arithmetic on the same 72 readings. Neither reader can tell which one they are holding.
The three widths are not evenly spaced at all, so do not read them as though each field were worth about the same. Naming the stretch closes 16.6667 points of the band. Adding the counting convention closes another 3.7037. Adding the signal closes the remaining 19.2848. The largest saving and the smallest sit on either side of the middle one, in that order, and there is no pattern in it to learn. A field is worth whatever the analyses it rules out happen to disagree about, and that has nothing at all to do with how important the field sounds when it is said out loud.
There is one limit worth naming before the panel above gets over-read. Only three of the six fields move the band, and those three are the ones that change which months got counted and how. The other two, the grid and the count of attempts, do not move the reading by a single month: run the same setting after trying twelve others first and it still calls 26 of 29. Neither the grid nor the count of attempts can show up in the band at all. Neither changes the reading itself; both change how much that reading is worth. Being invisible in the band is precisely why those two go missing.
A write up records the stretch and nothing else. How many analyses does it still allow, and how wide is the band of readings consistent with it?
Which field goes missing most often, and why does nobody notice?
The field is the count of attempts, and the reason it goes missing has nothing to do with dishonesty. The count goes missing because trying thirteen thresholds does not feel like thirteen tests. It feels like one afternoon.
Picture the afternoon. The analyst sits down and tries minus 6.00 per cent; the reading is dull. Then minus 5.00 per cent; still dull. Working up the grid, the reading climbs, it peaks at minus 1.00 per cent, and that is the one written down. From the inside this is a single continuous act with a single result, in the way that looking for a set of keys in six rooms is one search and not six. Nothing in the experience marks a boundary between one attempt and the next, so nothing in the write up marks one either.
Now the part that makes this field different from the other five: the document cannot recover it from anything else in the document. The record is in the document, so a reader can count the months. The rule is in the document, so a reader can reread it. The reading is in the document, and a reader can recompute it from the first two. The number thirteen is in the document only if somebody typed it there. The number leaves no trace in the record, no trace in the rule and no trace in the reading, and once the afternoon is over it exists in one person's memory and nowhere else at all.
The cost falls due later, and it falls on the figure rather than on the person. A reading of 89.6552 per cent from a rule somebody proposed and then tested once is a different claim from the same 89.6552 per cent picked as the best of thirteen. Same number, same record, same arithmetic, different claim. When the write up cannot say which of the two it is, a careful reader has to assume the second. The second is the weaker reading, and it is the one the document does not rule out. A missing attempt count does not make the figure wrong; it leaves the figure unable to defend itself.
Which field goes missing most often, and why does nobody notice it going?
How is it recorded when each number could first have been known?
There is a seventh thing on a good write up, and it is not a field of its own. The seventh thing is a date written beside every number the rule reads. Every number has a date on which it could first have been known, and that date is not always the date the number is filed under. Recording it costs a column and it is the difference between a document a reader can trust and a document a reader has to take on faith.
The earlier reading on this same record already worked the case through, so it does not need rebuilding here. Take a three month average of the price of the Nakshatra unit. Computed honestly, from this month and the two before it, that average sits Rs 4.9061/- away from the price it is meant to describe, averaged over the 69 months where it can be worked out at all. Now compute the same three month average as a centred averageAn average that borrows from both sides of the month it is written against, the ones before it and the ones after. Such an average fits the past neatly and belongs to no single date inside it., using the month before, the month itself and the month after. The centred average now sits Rs 2.5431/- away, an improvement of 48.1650 per cent.
Nothing about the second average is a better idea, and the whole of that improvement came from one month of information that had not happened yet. Where the number stays in a column, this is a data error: a column that describes history beautifully and cannot honestly be dated to any month in it. Inside a backtest the same thing is a backtest error, and the difference is not a quibble. The number did not stay in the column. The number went into a rule, and it came out as a reading somebody might quote.
Here is the backtest version of exactly that mistake, and it is deliberately unglamorous. Leave the Ashwin rule alone. Leave the record alone. Leave all thirteen settings alone. Swap only the signal: instead of reading last month, let the rule read the average of this month and the next. A two month average is perfectly ordinary arithmetic, and this one is dated one month too early. Scored on the same 65 months, the best reading goes from an honest 67.6923 per cent to a leaked 84.6154 per cent, a lift of 16.9231 points.
The leaked ladder reads higher at all thirteen settings and lower at none, and the lift runs from a single month at the lowest setting to thirteen months at a threshold of 1.00 per cent. The shape of the lift is what makes this failure so hard to catch by eye. The leak does not produce one absurd figure sitting next to twelve sensible ones. Anybody would query a figure like that. The leak lifts the whole shape, keeps it looking exactly like a ladder should look, and moves the best setting from minus 1.00 per cent to 1.00 per cent, so even the conclusion about which setting looked best is a leaked conclusion. 84.6154 per cent is the most flattering number in the whole comparison and the least real, so the word leaked stays attached to it every time it is printed.
Swapping in a signal that reaches one month forward lifted the best reading by 16.9231 points. Does that make the rule better, and where did the lift come from?
There is a check that catches a leaked column in about five seconds and needs no arithmetic at all. What is it?
What does a column with an empty last row indicate?
Scroll to the bottom of the column. Scrolling is the whole check, and it costs nothing.
The honest signal on this record is the month's own reading, and the record has one of those for every month. The honest column therefore has 72 filled rows and stops when the record stops. The leaked signal is the average of the month and the one after it, so it can be worked out for 71 months and not for the seventy second. December 2024 needs a January 2025, and the record has not got one. A column whose last row cannot be filled was built from the future, and finding out takes as long as scrolling to the bottom of it.
The check asks a question rather than settling one, and that is worth stating exactly. A column can be honest and still short at the bottom for a dull reason, such as somebody trimming a row while tidying. A signal that reaches two months forward leaves two rows empty rather than one, so the shape of the gap shows how far the reach went. And a column that is short at the top rather than the bottom is usually the ordinary consequence of an honest backward average. A backward average needs earlier months the record does not have. The empty last row gives a specific question, asked of a specific column, at no cost at all. In a folder of forty columns it is the cheapest check of the day.
What this looks like when it goes wrong
An analyst writes the run up carefully. The record is described, the rule is stated in a sentence, the split date is given, the reading of 89.6552 per cent is printed, and the whole write up reads as the work of somebody being thorough. The work was thorough. Two fields are not on it: the grid, and the count of attempts. The grid and the count of attempts were left off because they felt like working notes rather than results. Feeling that about them is completely ordinary, and it is the reason this failure is common rather than rare.
Two years later a colleague picks the document up. The question they need answered is simple. Was one threshold proposed and then tested, or were thirteen tried until one landed? The document cannot say. Neither can the analyst. The number thirteen was never written down anywhere, and two years is longer than that kind of memory lasts. Rereading the document more carefully will not help, and no amount of care on the reader's side can put back a field the writer never wrote. The figure now has to be read as the best of an unknown number of attempts. Reading it that way is the weakest reading available, and the only route back to a stronger one is to rerun the entire search from the start.
The fix costs one line and it has to happen at the right moment: the grid goes down before the search rather than after it. A grid written in advance is a list of thirteen numbers on a card. The same grid written from memory afterwards is a recollection, and it will quietly leave out the two settings that were tried first, disliked, and moved on from. Writing it first is what converts a memory into a document, and there is no later moment at which that conversion can still be made.
How does anybody actually use a write up they have been handed?
A write up arrives from somebody else. The write up will probably not be rerun today and may never be rerun at all, so the practical question is what to read first and what to conclude from whatever is missing.
The six fields are read before the result. A lender looking at a credit team's screening rule, an analyst inheriting a spreadsheet from somebody who has left, a committee member reading a document prepared last year: all three are in the same position. All three are being asked to attach weight to a number they did not produce. The fields tell them how much weight the number can carry. A document with all six is a number they can use the way the author used it. A document missing the stretch is a number that could have been nineteen other numbers. A document missing the attempt count is a number that has to be read as the best of a search whose size nobody knows.
The household version has the same shape, and it makes the question easy to ask out loud. Somebody reports that a particular shop is cheaper. Cheaper on what, over what period, and was it the first shop checked or the eleventh? Nobody is lying in any version of that conversation, and the answer still changes completely depending on which of the three is known.
One sentence is worth carrying away: a second person, holding only the document and the record, arrives at the same number. Six fields, a date beside every number the rule reads, and the file that regenerates the figure. With those handed over, the write up is finished. Without them, what has been written is a reminder, and a reminder only works for the person who did not need it.
An inherited write up carries the record, the rule and the reading, and nothing else. What has to happen before the figure is quoted?
What is left for other reading?
Two things are assumed rather than repeated: what a backtest is at all, and why a reading taken off history is not the same animal as one taken later while the rule is running. Covered separately: how to hold a search down before it starts, how to hold a stretch of the record back, and how to roll a test forward through time, each a separate question from what a document has to carry. Running many tests against one record needs arithmetic of its own and is covered separately too.
Nothing above says that any rule is worth testing, and nothing above says the Ashwin rule is worth anything at all. The Ashwin rule appears only as a thing being documented, and every figure attached to it is a hit rate or a count of months rather than anything a person could receive. Buying anything, selling anything, what a trade costs, and whether any rule deserves running at all: each is covered separately.
What was opened to produce the figures above?
Nothing was opened. Every count above came out of rerunning one calculation on 72 invented monthly readings, so the rows below record the working file that produced the numbers rather than an institution that published a table.
| Source | Document | Site |
|---|---|---|
| No outside source was used | A checking file kept in the same folder as these notes, rebuilding all 72 readings and every count printed above from their stated parts and halting if one has moved | None: no maintained record, published table or filed document is opened |
| The words write up, span, grid, convention, signal, threshold and denominator | Ordinary working vocabulary that turns up wherever people count things and argue about the counting | None required |
The Nakshatra unit, the six year record and the Ashwin rule are invented.
Educational material. Not advice on any investment, tax, budget or market position.
