Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quantitative Methods, Financial Data & Programming
1Probability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
2Statistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
3Correlation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
4Time Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
5Simulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
6Optimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
7Modelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
8Backtesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
9Data Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
10Programming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
11Quantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity

Backtest vs Live Performance: One Rule, Two Records

A backtest and a live stretch read the same rule and differ in what could be chosen. A backtest picks its threshold with the answers visible and can be recounted a second way; a live stretch cannot. Fitted on its first three years the Ashwin rule calls 26 of 29 scoreable months, or 89.66 per cent, and over the three years that followed it calls 18 of 36, or 50.00 per cent.

Two figures, one rule, and a gap of 39.6552 points between them. Most of what goes wrong with a comparison like this goes wrong in the next sentence somebody writes. The gap is real. The subtraction that produced it is correct. The meaning of the gap is a separate question. Two readings cannot answer it and three can at least be argued about, so the argument below builds a third.

Where these counts come from. The Nakshatra unit, an invented series, its six year record and the Ashwin rule that reads it were all built for teaching. Every count printed below was recomputed from the record's own seventy two monthly changes rather than lifted from a table, so there is no vintage to quote and no as of date to check. The recipe can be checked instead. Stating the record, the rule and the grid of thirteen settings brings every count printed below back. A count that refuses to reappear that way is simply wrong.

What makes a live stretch a different kind of record?

Start with the shopkeeper down the road. She wants to know her busiest hour, so she goes through last year's till roll, hour by hour, and finds that the hour after six in the evening beat every other hour in the day. The finding is real and it is about a real year. Now next Tuesday arrives. Whatever happens between six and seven next Tuesday is a number she did not pick, could not have picked, and had no hand in shaping. The number arrives on its own.

The distinction between the two stretches is nothing more than timing, and it is smaller than most people expect. A live stretch is simply the part of a record that arrived after the rule was written down, and what makes it different is not that the months are better but that they are later. Nobody chose them. There was nothing there to choose from at the moment the choosing happened. The months in the fitted stretch were already on the record while the thresholdA cut off figure. The rule compares its reading against this, and calls up on one side of it and down on the other. Where the cut off gets put is somebody's decision, never something the data hands over. was being decided, and the months in the live stretch were not.

The negative is what people get wrong, so say it back to yourself that way. The later months are not cleaner data. The later months are not harder, or fairer, or more honest. The monthly changes in the live stretch were produced by the same record as the earlier ones and carry no special property at all. The only thing separating the two heaps is a date, and the date matters because of what was happening on the day it fell: somebody was making a decision, and half the record was invisible to that decision.

ONE RECORD, CUT ON THE DAY A DECISION WAS MADE The Nakshatra unit, invented. Seventy two monthly observations, January 2019 to December 2024. THE THRESHOLD IS FIXED HERE, END OF DECEMBER 2021 THE FITTED STRETCH 36 months, visible when the choice was made THE LIVE STRETCH 36 months, none of them there yet Jan 2019 Dec 2024 chosen here, then read here 89.66 per cent 26 of 29 scoreable months not chosen anywhere, only read 50.00 per cent 18 of 36 scoreable months
The threshold was fixed at the end of December 2021 against the months to its left, and every month to its right arrived afterwards without being chosen by anybody.
Try it out

What makes the live stretch a different kind of record from the fitted one?

Breaking Into Quants Bootcamp — Fin Maverick

What can a backtest do that a live stretch cannot?

Lists survive being repeated at a meeting in a way ideas do not, so here is the part worth memorising as a list. There are four things a backtest can do that a live stretch cannot do at all. The four are not subtle statistical wrinkles. The four are ordinary practical liberties, and every one of them is the liberty to choose.

One, a backtest can choose its threshold with the answers already in front of it. The Ashwin rule carries thirteen whole number settings, the lowest at minus 6.00 per cent and the highest at 6.00 per cent, and the fit walks all thirteen, scores each, and keeps the best. Every one of those thirteen scores was computed on months whose outcomes were already known. On the live stretch there is nothing to walk. The threshold was fixed in December 2021 and it stays fixed, whatever the months do afterwards.

Two, a backtest can be recounted under a different convention. Six months of this record did not move at all, and an unmoved month is excluded rather than scored. Excluding it is a conventionA counting rule that was agreed rather than discovered. Set it down beforehand and a second person working alone arrives at identical figures from identical rows., not a fact, and somebody could reasonably decide otherwise. Redo the backtest under the other decision and every figure moves. The live stretch was counted once, on the day, by whoever was counting, and there is no second pass to make.

Three, a backtest can be rerun on a different span of the same record. Start it in 2020 instead of 2019. End it in June instead of December. Each span is a different question with a different answer, and all of them are available at once because the whole record is on the table. A live stretch has exactly one span: the one that happened, from the day the rule was fixed to today, and it grows at the rate of one month per month.

Four, a backtest can be repeated after a fault is found. Somebody spots a mistake in the counting, fixes it, and runs the whole thing again in an afternoon. Repeatability is a genuine strength of a backtest and it should be said plainly. A live stretch cannot be repeated at all. If the months were counted wrongly while they were arriving, the fix is to count them again, but the months themselves are what they were.

Put the four together and the pattern is obvious. Every one of them is a form of choice, and a figure arrived at with choice available belongs to a different species from one arrived at with none. That is not a criticism of backtesting. Choice is what a backtest is for. A thing that can be rerun, recut and recounted is enormously more useful than a thing that cannot, and the price of that usefulness is that its headline figure has more of somebody's decisions baked into it.

FOUR QUESTIONS, AND ONE COLUMN ANSWERS NO TO ALL OF THEM CAN IT DO THIS? A BACKTEST A LIVE STRETCH Choose the threshold walk all thirteen settings with the outcomes visible Recount under another convention score the six unmoved months instead of dropping them Change the span start in 2020 instead, or stop in June instead Repeat it after a fault is found fix the counting and run the whole thing again YES YES YES YES NO NO NO NO
A backtest can choose its threshold, recount under a different convention, change its span and be repeated, and a live stretch can do none of the four.
ALL FOUR ARROWS PASS THROUGH ONE GATE, AND IT IS CALLED CHOICE CHOICE choose the threshold, thirteen settings recount the six unmoved months recut the span, any start, any end repeat it once a fault is found THE BACKTEST SIDE every liberty above is available at once THE LIVE SIDE no gate, because there is nothing to pass through it the months simply arrive and get counted once
Everything a backtest can do that a live stretch cannot comes down to choice, and a figure produced with choice is not the same sort of quantity as one produced without it.
Try it out

Somebody supplies a single figure from a backtest and nothing else. Of the four liberties above, which one most needs disclosing first?

What did the Ashwin rule read across the two stretches?

Now the counting. The fit takes the first three years, January 2019 to December 2021, walks the thirteen settings, and keeps the one that calls the most months correctly. The best setting is a threshold of minus 1.00 per cent. Across the 29 scoreable months in those three years it calls 26 of them the right way, or 89.66 per cent.

Then the rule is carried forward without a single change. Same signalWhatever a rule looks at before it commits itself. A signal is always built out of months that have already closed, never out of the month being called, and here it is just the change over the month that has finished., which is the Nakshatra unit's change over the month just ended. Same threshold of minus 1.00 per cent. Same call, up when the reading sits above the threshold and down otherwise. Across the 36 scoreable months of January 2022 to December 2024 it calls 18 of them the right way, or 50.00 per cent.

Notice what each side has to state about itself. Each reading carries its own denominatorThe bottom half of a fraction. The denominator settles what a percentage is a percentage of. Two results carrying the same top half can still be saying quite different things., 29 months on the left and 36 on the right, and the two differ because six months of the record did not move and every one of those six sits in the first three years. January 2019 drops out of the first stretch as well. The signal reads the month just ended, and January 2019 has no month before it, so 36 months come down to 29. A reader who sees only the two percentages has been handed the arithmetic with the denominators removed, and the denominators are half of what those two figures are saying.

One more thing about the second figure, and it needs saying before anybody builds a story on it. The live reading is 18 of 36, exactly one half. Reading something into that half is very tempting, and there is nothing in it to read. Eighteen of thirty six landing on exactly one half is a coincidence of this particular record, not a law that a failed backtest falls to a coin. The count could as easily have come out 17, which reads 47.22 per cent, or 19, which reads 52.78 per cent, and nothing anywhere in the arithmetic was pushing it toward the middle. The number is what it is because thirty six particular months did what they did.

Try it out

The rule calls 18 of 36 on the live stretch, exactly one half. Does landing on exactly one half say something?

Moving the split is worth settling first. December 2021 is not a sacred date. The date is a choice, like everything else on the backtest side, and the honest way to see how much rests on it is to make the same cut at seven different points and read both figures each time.

MOVE THE SPLIT SEVEN TIMES AND READ BOTH SIDES EACH TIME At each split the thirteen settings are searched before it, then that threshold is read after it. 90 80 70 60 50 40 per cent of scoreable months called correctly 90.00 89.66 73.58 57.78 50.00 38.89 Dec 2020 Jun 2021 Dec 2021 Jun 2022 Dec 2022 Jun 2023 Dec 2023 Upper line: the months before the split, where the threshold was chosen. Lower line: the months after it, where nothing was chosen at all.
Across the seven split points the before reading is the higher one at every setting, and the after reading moves up and down rather than falling steadily.
Try it out

As the split moves later and the fitting stretch grows, does the gap between the two readings widen steadily?

Play with it

Move the split through seven dates and watch the gap refuse to behave.

One control, and it moves one thing: the month the record is cut at. At every setting the thirteen settings of the Ashwin rule are searched afresh on the months before the cut, the best one is kept, and that same threshold is then read on the months after it. Three things redraw. The strip along the top reshades to show how much record sits on each side. The two bars in the middle rescale, and the bracket between them relabels itself with the gap in points. The ladder along the foot shows all thirteen settings scored on the fitting stretch, with the one the fit picked marked, so the pick itself can be watched changing. The default setting is the cut after December 2021, and it reproduces 26 of 29 and 18 of 36 exactly.

Dec 2020cut after December 2021Dec 2023
ONE CUT, TWO READINGS, AND A LADDER THAT KEEPS CHANGING ITS MIND The Nakshatra unit and the Ashwin rule, both invented. Every count below is recomputed at each setting. Jan 2019 Dec 2024 shaded: the fitting stretch 89.66 26 of 29 before the cut CHOSEN HERE 50.00 18 of 36 after the cut NOTHING CHOSEN 39.6552 points ALL THIRTEEN SETTINGS, SCORED ON THE FITTING STRETCH ONLY threshold of the Ashwin rule, per cent
Cut after
Dec 2021
Threshold picked
minus 1.00
Before the cut
89.66%
After the cut
50.00%
Gap, points
39.6552
Educational illustration. The Nakshatra unit, the six year record and the Ashwin rule were built for teaching and describe nothing that trades or ever did. A month that did not move is excluded at every setting of the control, so both denominators change as the cut moves and neither side ever adds up to a round thirty six. The same thirteen settings are searched at every cut, and the search is the only thing that ever chooses anything. Two of the seven cuts happen to give an after reading of exactly 50.00 per cent, from 18 of 36 and from 12 of 24; those are different counts on different months and nothing connects them. Every figure on this panel is a count of months called correctly, never a return.
Try it out

Commit before reading on. Suppose the same thirteen settings are fitted on the last three years and then read on those very same months. What sort of reading would that produce?

AI For Finance Bootcamp — Fin Maverick

How much of the drop is the search, and how much is the record?

Here is the move that turns two figures into an argument. Put a third reading between them.

Take the same thirteen settings and fit them on the last three years, the very months the carried rule struggled on. Let the search have every liberty it had the first time round: all thirteen settings, all the outcomes visible, keep the best. The search picks a threshold of minus 6.00 per cent and calls 22 of the 36 scoreable months, or 61.11 per cent.

The 61.11 per cent is the hinge of the argument. A search with full choice over the later stretch reaches only 61.11 per cent there, so the later stretch is a harder stretch to call even for a rule fitted on it. Now the drop of 39.6552 points has somewhere to pass through. From 89.66 per cent down to 61.11 per cent is 28.5441 points, and that step changes the stretch while leaving the search its choice. The carried threshold of minus 1.00 per cent is not what the later stretch would have picked, so the step from 61.11 per cent down to 50.00 per cent, 11.1111 points, holds the stretch still while taking the choice away. The two steps add to 39.6552 points, the whole drop.

THE DROP HAS THREE RUNGS, NOT TWO Same rule, same thirteen settings, same record. Only the fitting stretch and the reading stretch change. 89.66 per cent fitted on the first three years, read there. 26 of 29 61.11 per cent fitted on the last three years, read there. 22 of 36 50.00 per cent fitted on the first three years, read on the last. 18 of 36 28.5441 points the stretch changes, the search keeps its choice 11.1111 points the stretch holds still, the choice is taken away The two steps add to 39.6552 points, the whole drop from the top rung to the bottom. THEY ADD UP BECAUSE OF HOW THEY WERE BUILT. THAT IS NOT THE SAME AS EXPLAINING ANYTHING.
Putting a third reading of 61.11 per cent between the two splits the drop of 39.6552 points into 28.5441 points and 11.1111 points.
Try it out

The drop of 39.6552 points splits into 28.5441 and 11.1111. What would have to be true about the middle reading before those two numbers could be called causes?

What kind of split is that, and what would make it an attribution?

The split is the most dangerous step in the argument, and it carries its own warning. The two parts add to the whole. Two parts will always add to the whole. Adding to the whole is not evidence of anything.

The arithmetic matters more than the story. The starting expression was 89.66 minus 50.00. Then 61.11 was subtracted and added straight back: 89.66 minus 61.11, plus 61.11 minus 50.00. The middle term cancels itself out. The addition is a consequence of how the expression was written and not a finding about the record, so any number at all put in that position would produce two parts adding to 39.6552. That is a telescoping split, and a telescoping split is not an attribution.

A plainly silly middle reading makes the point. With 70.00 per cent in the middle the two parts are 19.6552 points and 20.0000 points, adding to 39.6552. With 55.00 per cent there they are 34.6552 and 5.0000, adding to 39.6552. Nothing about the record changed between those three tellings. Only the number somebody chose to stand in the middle changed, and the two parts obediently rearranged themselves around it.

So what would make it an attribution? A middle reading earns the name on one condition. Somebody has to argue that it holds one thing genuinely fixed while the other genuinely moves, and has to say why that particular hold is the right one. The 61.11 per cent used above has a real argument behind it, which is why it is the one used here: it is what a full search reaches on the later months, so it separates "the later months are harder" from "the threshold was stale". But it is still one choice among possible choices, made by a person, and an account that prints 28.5441 and 11.1111 without saying so has quietly promoted its own arithmetic to a discovery.

The household version is easier to feel. A household electricity bill went up by nine hundred rupees this year. Somebody says four hundred was the tariff and five hundred was the new cooler. Checking that needs a bill for the same house with the new tariff and the old appliances, and if nobody has one, the four hundred and the five hundred are two numbers that add to nine hundred and nothing more. Two such numbers will always add to nine hundred. Whoever picked the middle figure picked the split.

THE SUBTRACTION IS RIGHT. THE SENTENCE IS WRONG. Findings, line 4: overfitting cost 39.6552 points. reported as an attribution when it is a subtraction with the middle missing MISSING: 61.11 PER CENT a full search on the later months reaches only this AND ANY MIDDLE READING SPLITS IT, WHICH IS THE PROBLEM middle at 55.00 per cent 34.6552 and 5.0000 adds to 39.6552 middle at 61.11 per cent 28.5441 and 11.1111 adds to 39.6552 middle at 70.00 per cent 19.6552 and 20.0000 adds to 39.6552 Only the middle row has an argument behind it. All three rows add up perfectly.
The drop of 39.6552 points is a correct subtraction and a wrong explanation, because a rule fitted on the later stretch itself reaches only 61.11 per cent there.
Risk Management Program Bootcamp — Fin Maverick Reading an Option Payoff — free micro-course from Fin Maverick

Did the record itself get harder to call?

The record did get harder to call, and nobody chose that either. A record that changes on its own is what stops the drop from being a story about somebody's carelessness.

The first three years of the six year record carry a spreadA measure of how scattered a set of figures is around its own middle. Scatter them widely and individual entries land a long way out on both sides; scatter them narrowly and they crowd in tight. of 3.00 per cent. The last three years carry a spread of 6.4031 per cent, more than double. And through both stretches the mean sits at exactly 1.00 per cent. The centre of the record did not move at all; only the distance the months travel from it more than doubled, and a rule that calls direction has to work in whatever weather it is given.

Feel it with a street vendor. Two years running, the same juice stall. In the first year daily takings sit between three hundred and five hundred rupees, day after day, and the vendor can guess tomorrow's takings pretty well by looking at today's. A building site opened opposite and its shifts move around, so in the second year the average is exactly the same but the days swing from nothing to twelve hundred. Nothing about the vendor's guessing got worse. The days got wilder. Anyone judging the vendor by the hit rate alone would conclude the vendor lost their touch.

SAME CENTRE, TWICE THE SPREAD, AND NOBODY CHOSE IT Each dot is one month of the Nakshatra unit, invented. Both rows are drawn to the same horizontal scale. MEAN 1.00 PER CENT, BOTH ROWS FIRST THREE YEARS SPREAD 3.00 PER CENT 36 months, none wider than 7 per cent either way LAST THREE YEARS SPREAD 6.4031 PER CENT minus 12 minus 8 minus 4 0 4 8 12 16 monthly change of the Nakshatra unit, per cent
The spread goes from 3.00 per cent to 6.4031 per cent while the mean stays at exactly 1.00 per cent, so the record got harder to call without moving its centre.

And now the uncomfortable part. The wide stretch had not happened yet, so the backtest could not possibly have separated a rule breaking down from a record becoming harder to call. In December 2021, when the threshold was fixed, every month available to look at came from the narrow stretch. There was no signal anywhere in that data saying the next thirty six months would swing twice as far. Missing it was not an oversight anybody could have avoided by being more careful. The limit is on what a backtest can know, and it is permanent.

Try it out

The spread went from 3.00 per cent to 6.4031 per cent while the mean stayed at exactly 1.00 per cent. What does that do, and who chose it?

What to ask when a backtested figure and a live figure disagree

Here is the part that gets used in a room. Somebody has two numbers and a conclusion, and there are about four minutes to test it. Five questions, in this order.

When was the threshold fixed, and against which months? If the answer is vague, or if the date turns out to sit after some of the live months, there is no live stretch here at all and everything downstream is a backtest wearing a different name.

How many months are in the live stretch, and how many of those could be scored? Thirty six months sounds solid; thirty six months of which twelve were dropped under a counting convention is a different object. On this record the two sides carry 29 and 36 scoreable months, and the difference between them is entirely six months that did not move.

What does the baselineWhat a deliberately mindless approach scores on the very same data. A baseline exists to be set alongside a result, giving the result something to be better or worse than. read on the live stretch on its own? Calling every one of those 36 months up and reading nothing at all gets 20 of them right, which is 55.5556 per cent. The baseline sits above the rule's 50.00 per cent. A live figure without its baseline beside it is not yet a finding.

Would a rule fitted on the live stretch itself have done better, and by how much? This is the 61.11 per cent question, and it is the one that decides whether the phrase stale threshold is doing any work. The answer costs one more pass over the same thirteen settings.

And has anything about the record changed between the two stretches that nobody chose? This is the question people skip, and it is the one that decides how much of the disagreement is anybody's fault. Here the spread more than doubled. No amount of care in December 2021 would have caught that, and a report that does not mention it is inviting its reader to blame a person for the weather.

TWO SHEETS, THREE YEARS APART. NEVER ONE SHEET. SHEET ONE, DATED 31 DECEMBER 2021 Rule the Ashwin rule Fitted on Jan 2019 to Dec 2021 Threshold minus 1.00 per cent Live reading left blank, nothing has happened Signed before a single live month arrived. SHEET TWO, DATED 31 DECEMBER 2024 Rule the Ashwin rule Read on Jan 2022 to Dec 2024 Threshold minus 1.00 per cent Live reading 50.00 per cent from 18 of 36 scoreable months Copies the threshold across unchanged. 3 years ONE SHEET CARRYING BOTH FIELDS WAS WRITTEN AFTER THE FACT. The date on sheet one is the only evidence that the threshold was fixed before the live months arrived.
A write up fixing the threshold before the live months begin and a write up reporting the reading afterwards are two separate documents, and a single document carrying both was written after the fact.
Try it out

Somebody reports that a rule tested well and did much worse afterwards. Which pair of figures should be asked for before agreeing with the word afterwards?

The failure: a right subtraction carrying a wrong sentence

Somebody lines up 89.66 per cent against 50.00 per cent, works out that the drop is 39.6552 points, and writes that overfitting cost 39.6552 points. Every digit in that sentence is correct. The sentence is still wrong. The arithmetic is fine, so checking the arithmetic will never catch it.

On this record a rule fitted on the later stretch and read on that same later stretch reaches only 61.11 per cent. So at the very most 11.1111 of those points can be laid at the door of the carried threshold, and even that number is only as good as the argument for putting 61.11 per cent in the middle. The rest, 28.5441 points, belongs to a stretch whose spread went from 3.00 per cent to 6.4031 per cent while nobody was choosing anything at all.

The cost of the wrong sentence is specific and it is expensive. Believing that the whole drop was overfitting leads straight to tightening the search, and to six months spent doing exactly that. Nothing in those six months touches the spread of the record, and the spread is what actually moved. Meanwhile the report has told a room full of people that somebody was careless, when most of what happened was a record getting wilder.

The fix is small and nobody does it. Before deciding what a drop is made of, the rule should be fitted on the later stretch as well and read there. One more pass over the same thirteen settings, on months already in hand. The result is either a high reading, in which case the carried threshold really was most of the story, or a low one, in which case a large part of the drop was never anybody's to prevent.

Holding a stretch back before the work starts is covered separately, and a reading on it means something for exactly that reason. Rolling that hold forward through time, refitting and rereading as the record advances, follows it. The arithmetic of what happens when many separate tests are run against one record, and the two different ways of choosing after the record has already been seen, are covered separately as well. Placing an order, paying a cost, holding anything or moving money is covered separately too, and does not come into a word of the argument above.

The Ashwin rule was written to be put through a test, and mostly it does not survive one. Every figure above is a count of months called correctly out of months that could be counted, never a return, a gain or a loss.

The record got harder to call by itself. See what live months show.

What outside source could settle a count like this?

No rule, no filing requirement, no threshold set by anybody and no product from any market enters the argument, so there is nobody to cite. The whole argument is subtraction performed on an invented record, and subtraction has no publisher. The recipe travels with each figure instead.

SourceDocumentSite
No rule maker namedNone, because no rule is quotedNone
No maintained series namedNone, because no outside record is readNone
No past result namedNone, because every failure here is demonstrated rather than citedNone

The Nakshatra unit, its six year record and the Ashwin rule are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.