Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quantitative Methods, Financial Data & Programming
1Probability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
2Statistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
3Correlation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
4Time Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
5Simulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
6Optimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
7Modelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
8Backtesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
9Data Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
10Programming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
11Quantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity

Backtesting: Testing a Rule on History, and How a Strategy Fits the Past

Somebody hands over a figure. A rule was tried on six years of a record, and it called the direction right about two times in three. The figure is real arithmetic. Nobody made it up, nobody rounded it kindly, and a second count of it would land on exactly the same number. And it still might mean nothing at all. Backtesting lives entirely in the gap between those two sentences. One invented record and one invented rule are enough to walk the gap end to end, and every number in it can be rebuilt.

The three things the counting stands on

The six year record of the Nakshatra unit. 72 dated monthly observationsOne dated entry in a record. The six year record holds 72 of them, one for each month., invented and settled elsewhere: an average month of 1.00 per cent, a spreadHow far the months of a record typically sit from their own average. A wide spread means months land further out in both directions. of 5.00 per cent, an opening mark of Rs 100.00/- and a closing price of Rs 187.4539/-. The record already carries a drift, a repeating calendar pattern and a scatter that more than doubles partway through. Not one figure in it is changed here.

Fitting a rule on one stretch of a record and reading it on another. Splitting a record that way is settled elsewhere in these notes, and it is used below without being explained again.

Counting, and nothing harder. Every quantity below is a count of months, a count of attempts, or one count divided by another. There is no other arithmetic here.

What is a backtest, and what does it actually count?

A backtest takes a rule that can be stated in one sentence, walks it through a record of months that have already finished, and counts how often the rule would have been right. The counting is the whole of the operation. There is no forecasting in it, no judgement in it and nothing clever in it. A backtest is a counting machine, and the only thing it can count is what the record already contains.

Think about a samosa stall outside one office building. The owner writes down a rule at the end of the year: fry early whenever yesterday evening was busy. Then she goes back through forty evenings of her own notebook and marks, evening by evening, whether frying early would have been the right call. At the end she has a count. Forty evenings, thirty one of them marked. The count is a backtest, and a backtest is exactly as good as the notebook, the rule and the marking behind it.

Notice what she cannot get out of the notebook however hard she counts. She cannot get what those evenings would have done under different weather. Different weather is not in the notebook. She cannot get whether the rule will still hold next winter. Next winter is not in the notebook either. And she cannot get whether frying early was worth doing. A count of right calls is not a count of anything else. Everything a backtest cannot answer, it cannot answer for the same reason: a record holds what happened and nothing else.

A BACKTEST IS A COUNTING MACHINE, AND THE BIN BESIDE IT MATTERSThree things go in. One count comes out. Everything else stays outside, because a record cannot hold it.THE RECORD72 dated monthly observationsof the Nakshatra unitTHE RULEOne sentence, fixed before anycounting startsTHE CONVENTIONWhat to do with a month thatdid not moveWALK THE RULETHROUGH THE RECORDONE MONTH AT A TIMEDid the call match thedirection the month actuallytook? Mark it. Move on.no judgement, only marksONE COUNT44 matched calls out of 65scoreable months, at the best of 13 settingsWHAT STAYS IN THE BIN•What the months would have done hadanything been different•Whether the rule keeps calling nextyear•Anything the record does not alreadycontain
A backtest turns a record, a rule and a stated counting convention into one count of matched calls, and everything a record does not already contain stays outside the machine.
Try it out

A backtest produces a count. What exactly is being counted, and what can no amount of counting get out of the record?

Cleaning Financial Data — free micro-course from Fin Maverick

Which of the 72 months can be scored at all?

Before a single call is made, there is a duller question to settle: which months are even eligible? The six year record of the Nakshatra unit holds 72 months. The record does not hold 72 scoreable months, and the difference is not a rounding detail. The scoreable count is the denominatorThe count sitting underneath a share, which decides what the share is a share of. Change it and the share changes without any new information arriving. under every reading below.

Two things take months out. January 2019 has nothing before it to read, so the rule cannot produce a call for it at all. And six months did not move: April 2019, September 2019, August 2020, March 2021, April 2021 and September 2021 each closed exactly where they opened. A rule that calls up or down has no correct answer available on a month that did neither. One month with no signal and six months with no direction leaves 65 scoreable months out of 72.

The treatment of those six months is a conventionA choice about how something gets counted, written down in advance so that somebody else can repeat the count and land in the same place., and the honest version of a convention is one written down before anything is counted rather than settled afterwards when the counts are already on screen. The convention here is that an unmoved month is excluded. An unmoved month could just as defensibly have been counted as a rise. Counting the six as rises would give a different denominator and a different reading. The point is not which choice is better but that the choice was made in advance and is printed where a reader can see it.

All six unmoved months sit in the first three years, and none at all sits in the last three. The clustering is not a detail either. Those six months mean the record does not split into two equal halves of 36 scoreable months. The record splits into 29 and 36, and every later comparison carries those two uneven counts underneath it.

SEVENTY TWO MONTHS, AND ONLY SIXTY FIVE OF THEM CAN BE SCOREDEach cell is one month of the six year record of the Nakshatra unit.JanFebMarAprMayJunJulAugSepOctNovDec2019nosignal21flatminus 2minus 4minus 1minus 1flat2419 scoreable2020642minus 3minus 3minus 2minus 1flat315611 scoreable202173flatflatminus 1minus 3minus 4minus 2flat6359 scoreable202286minus 42531minus 3minus 516minus 3112 scoreable2023448minus 6minus 1minus 6686minus 411minus 312 scoreable20243minus 1minus 11minus 10minus 6minus 13minus 82minus 341412 scoreableOne month with nothing before it to read, and sixmonths that did not move, and all six of those sit inthe first three years.65 in all
The six months that did not move all sit in the first three years, which is why the two stretches carry 29 and 36 scoreable months rather than 36 each.
Try it out

The six year record holds 72 months but every reading below sits over 65. Where did the other seven go?

Try it out

Every one of the six unmoved months landed in the first three years and not one landed in the last three. Accident, or something about that stretch?

Cleaning Financial Data teaches you to find the errors that survive every check and break every model.

What does the Ashwin rule read, and how many settings does it have?

Only one object is still missing, and it is a small one. The Ashwin rule runs like this. At the end of each month, read the change the Nakshatra unit made that month. If that reading sits above a thresholdThe number a reading has to clear before a rule calls one way rather than the other., call the next month up. Otherwise call it down. Nothing else belongs to the rule. The rule fits in one sentence because it has to: a rule that cannot be stated in one sentence is a rule that cannot be checked.

The number the rule reads before it calls is the signalThe number a rule reads before it makes a call, computed only from months that have already finished.. Here the signal is just the month that has finished, taken as it stands. Nothing is averaged, nothing is smoothed and nothing from later is folded in. Every number the rule reads on the last day of a month was already sitting there on the last day of that month. Being able to say that sentence about a rule matters more than it sounds.

ONE MONTH OF THE ASHWIN RULE, END TO ENDThe same four steps repeat for every scoreable month of the record.1. THE MONTH THATJUST FINISHEDNovember 2019 closes.Its reading is4.00 per cent.2. THE SIGNAL READOFF ITThe signal is that samereading, 4.00 per cent,and nothing else.3. COMPARED WITH THETHRESHOLDIs 4.00 per cent aboveminus 1.00 per cent? Itis.4. THE CALL FOR NEXTMONTHCall December 2019 UP.The call is written downbefore the month starts.WHAT DECEMBER 2019 ACTUALLY DIDIt rose, by 1.00 per cent. The call matched, so this month adds one to the count.1 of 65
The Ashwin rule reads the month just finished, compares that signal against one threshold, and writes down a call for the month that has not started yet.

The threshold has to be some number, and there is no principle that supplies one. So the rule comes with a gridThe stated list of settings a search walks through, fixed before the search starts, so that the attempts can be counted afterwards.: thirteen whole number settings running from minus 6.00 per cent up to 6.00 per cent. The grid is stated in advance. A stated grid is a countable grid, and a count of attempts is the single most useful thing anybody can say about a tested figure.

Running all thirteen over the 65 scoreable months of the six year record gives a ladder. The best setting is minus 1.00 per cent, calling 44 of 65 months, or 67.69 per cent, and the worst is 6.00 per cent, calling 28 of 65, or 43.08 per cent. One more figure has to sit beside that ladder. Calling every single month up and reading nothing at all lands 37 of the same 65 months, or 56.92 per cent. The always up count is the baseline, and every reading on this record has to be measured against it.

Threshold settingMatched callsHit ratePoints against the lineWhere it sits
minus 6.00 per cent39 of 6560.00 per cent3.08above the line
minus 5.00 per cent38 of 6558.46 per cent1.54above the line
minus 4.00 per cent38 of 6558.46 per cent1.54above the line
minus 3.00 per cent39 of 6560.00 per cent3.08above the line
minus 2.00 per cent41 of 6563.08 per cent6.15above the line
minus 1.00 per cent44 of 6567.69 per cent10.77above the line
0.00 per cent43 of 6566.15 per cent9.23above the line
1.00 per cent42 of 6564.62 per cent7.69above the line
2.00 per cent40 of 6561.54 per cent4.62above the line
3.00 per cent38 of 6558.46 per cent1.54above the line
4.00 per cent33 of 6550.77 per centminus 6.15below the line
5.00 per cent30 of 6546.15 per centminus 10.77below the line
6.00 per cent28 of 6543.08 per centminus 13.85below the line

Read the last two columns rather than the middle one. Ten of the thirteen settings read above the always up line and three read below it, and not one of them lands exactly on it. The best setting beats reading nothing by 10.77 points, or four extra months out of 65. Four extra months is the honest size of what the rule appears to add on this record, and it is a very different sentence from the one that starts with 67.69 per cent.

WHAT EACH SETTING ADDS OVER READING NOTHING AT ALLPoints of hit rate above or below the always up line of 56.92 per cent. One month is worth 1.5385 points.THRESHOLDminus 6minus 5minus 4minus 3minus 2minus 101234563.081.541.543.086.1510.779.237.694.621.54minus 6.15minus 10.77minus 13.85THE ALWAYS UP LINE, 37 OF 65 MONTHSReading nothing at all and calling every month up sits here.TEN SETTINGS SIT ABOVE THE LINETHREE SIT BELOW ITNone of the thirteen lands exactly on the line, and the widest reading is 24.62 points clear of the narrowest.
Ten of the thirteen settings read above the always up line and three read below it, so no single reading from the ladder means anything until it is placed against that line.
Try it out

The best setting reads 67.69 per cent and the always up line reads 56.92 per cent. How much is the rule actually adding, and what has to be known before treating that as real?

Try it out

The calculator below walks the threshold from minus 6.00 per cent up to 6.00 per cent. Does the count climb steadily to a peak and fall away, or does it wander on the way?

Play with it

Walk the grid, and watch how many attempts it takes to find the tallest bar

One control, one number inside the rule, thirteen places it can sit. The bar redraws to the count of matched calls at that setting, the dashed line across the middle is the always up baseline and never moves, and every setting landed on keeps a small tick above it. The green line marks the best found so far and never comes down. Walking upward from the first setting on the grid shows how many attempts pass before that green line stops moving.

THE THRESHOLD ACROSS THE GRID, AND THE SEARCH IT MAKES VISIBLEThirteen settings, one record, one rule holding one number. The bar is the count of matched calls out of 65.012243648matched calls39minus 638minus 538minus 439minus 341minus 244minus 1430421402383334305286always up: 37 of 65best so farThe line does not move. The rule does not get any bigger. Only the one number inside it changes.
minus 1.00 per cent

Educational illustration. An unmoved month is excluded, so the denominator is 65 rather than 71. The grid was fixed at thirteen settings before any of them was tried. Every bar counts matched calls out of 65 months, never money.

What does the whole read look like, in order?

Four readings, taken in the order anybody would actually take them, on one record with one rule. Read down the table and watch the same rule produce two answers that are nowhere near each other.

StepWhat is being readWhat it lands on
FirstHow many months can be scored at all65 of 72
SecondWhat calling every month up and reading nothing would land37 of 65, 56.92 per cent
ThirdThe best of the thirteen settings on the whole record44 of 65, 67.69 per cent
FourthThat threshold read on the first three years, then on the last three89.66 per cent, then 50.00 per cent

One record, one rule, one threshold, and two readings 39.6552 points apart. Nothing was rerun, nothing was recounted a second way and no arithmetic changed between the third row and the fourth. The only thing that changed is which months the rule was standing on.

Try it out

A threshold picked on the first three years calls 26 of the 29 scoreable months there. On the 36 months that arrived afterwards, does it hold up, slip a few points, or fall a long way?

Live Performance: what did the rule do on the months nobody chose?

Suppose the threshold had actually been settled at the close of December 2021, using the first three years and nothing else. The setting that calls the most months there is minus 1.00 per cent. On those 29 scoreable months it calls 26, or 89.66 per cent. If the work stopped there, that is the figure that would have been written down.

The 36 months that arrived afterwards are the live stretch. Every one of them turned up after the threshold was already fixed. A live stretch cannot be rerun, cannot be recounted a second way, and does not care in the slightest what was chosen before it began. The live stretch is worth something for exactly that reason. Every other reading in this guide can be taken again with a different convention until it comes out nicer. The choosing was over before those months arrived, so the live reading cannot be retaken.

ONE RECORD, TWO STRETCHES, AND ONLY ONE OF THEM WAS CHOSENThe threshold was settled at the close of December 2021 using everything to the left of the mark and nothing to the right.Jan 2019Jul 2019Jan 2020Jul 2020Jan 2021Jul 2021Jan 2022Jul 2022Jan 2023Jul 2023Jan 2024Jul 2024the threshold is chosen hereTHE FIRST THREE YEARS36 months, 29 of them scoreableTHE MONTHS THAT ARRIVED AFTERWARDS36 months, all 36 scoreableSPREAD OF THE MONTHS3.00 per centMonths sitting closer to their own averageSPREAD OF THE MONTHS6.4031 per centMonths sitting more than twice as far outThe average month is exactly 1.00 per cent on both stretches, so the two halves differ in how far the months scatter and not in where they sit.Nobody chose either spread, and nobody chose which stretch would arrive second.
The stretch the threshold was chosen on carries a spread of 3.00 per cent and the stretch that arrived afterwards carries 6.4031 per cent, and nobody chose either of them.

On those 36 months the same threshold, unchanged, calls 18. Eighteen of 36 is 50.00 per cent. The reading falls 39.6552 points, from 89.66 per cent to 50.00 per cent, without one thing about the rule changing.

Two cautions about that second figure, and both matter more than the figure does. The first is that 50.00 per cent is not a law. The figure is 18 matched calls out of 36 on these particular months, and nothing made it land on a round number. A rule whose fitted reading collapses does not have to land there, and reading this one figure as though it were the natural resting place of a failed test would be inventing a pattern out of one observation. The second is that nobody arranged the live stretch to be hostile. Its months simply scatter more than twice as widely as the months the threshold was chosen on, and nobody chose that either.

ONE THRESHOLD, TWO STRETCHES, TWO VERY DIFFERENT READINGSThe threshold is minus 1.00 per cent at both ends of this picture, and it never moves.020406080100hit rate, per cent89.66 per centON THE FIRST THREE YEARS26 of 29 scoreable months50.00 per centON THE MONTHS THAT CAMEAFTER18 of 36 scoreable monthsa drop of39.6552pointsThe rule is identical on both bars. The record is the same record. Only the months underneath the rule are different.
The threshold chosen on the first three years calls 26 of 29 months there and 18 of 36 on the months that arrived afterwards, a drop of 39.6552 points.
Breaking Into Quants Bootcamp — Fin Maverick

Backtest Overfitting: what did searching thirteen settings buy, and what did it give back?

Go back to the moment before the threshold was chosen and watch the choosing happen. Walk the grid in its stated order, from minus 6.00 per cent upward, keeping the best setting found so far. On the first three years the first setting tried calls 17 of 29, or 58.62 per cent. By the sixth attempt the best so far calls 26 of 29, or 89.66 per cent. Six attempts bought 31.0345 points on the stretch the search could see, and the seventh attempt onward bought nothing more.

Now read the same six settings on the 36 months the search could not see, one at a time, in the same order. The first attempt reads 61.11 per cent there. The sixth attempt reads 50.00 per cent. The same six attempts that bought 31.0345 points on the visible stretch gave back 11.1111 points on the invisible one. The two lines walk in opposite directions, and they do it while the rule stays exactly the same size.

SIX ATTEMPTS, AND THE TWO READINGS WALK IN OPPOSITE DIRECTIONSKeeping the best setting so far while walking the grid in its stated order, from minus 6.00 per cent upward.406080100hit rate, per centattempt 1minus 6attempt 2minus 5attempt 3minus 4attempt 4minus 3attempt 5minus 2attempt 6minus 189.66 per cent58.62 per cent50.00 per cent61.11 per centTHE STRETCH THE SEARCH COULD SEEclimbs 31.0345 points and then stops climbingTHE STRETCH IT COULD NOT SEEgives back 11.1111 points over the same six attemptsTHRESHOLD KEPT
Searching six settings bought 31.0345 points on the stretch the search could see and gave back 11.1111 points on the stretch it could not, with the rule holding the same single number throughout.

Backtest overfitting is the name for what just happened, and precision about it matters because the words invite a vaguer story. Nothing was tuned. Nothing was smoothed. No extra term was bolted onto the rule. Somebody looked at thirteen candidates and kept the one that looked best on the months in front of them. Keeping the best of thirteen is a completely reasonable thing to do and is what anybody would do. The searching itself is what moved the reading, and the searching leaves no mark on the rule for anybody to find afterwards.

Try it out

Searching lifted the fitted reading by 31.0345 points. A colleague says that proves the search worked. What single figure answers them?

Backtest Overfitting vs Model Overfitting: what exactly gets added in each?

Overfitting a model and overfitting a backtest get treated as the same worry in a slightly different hat, and they are not. The difference is a count, and once the count is in view it stops being a matter of judgement altogether.

Overfitting a model adds numbers to one rule. Give a rule a second number, then a third, then a tenth, and it gains the flexibility to bend around whatever record it is shown. The rule visibly grows. Anybody looking at it can see it growing. Overfitting a model therefore has a recognisable smell: the rule that fits beautifully is the complicated one.

Backtest overfitting adds attempts at one record. The rule stays exactly as simple as it was. The Ashwin rule held exactly one number at every one of its thirteen attempts. The thirteen candidates differ in that one number and in nothing else. Not one of them is more flexible than any other, not one of them could bend around anything, and every single one of them is a rule anybody would be happy to explain in a sentence. The count of attempts rose from one to thirteen, and the rule that came out at the end is exactly as simple as the rule that went in.

Backtest overfitting is harder to catch for exactly that reason. There is nothing on the finished rule to inspect. The complicated rule announces itself; the well searched one does not. The only trace the search leaves is a count of how many candidates were looked at, and that count lives in somebody's head or in a script nobody reads, not in the rule and not in the figure.

TWO KINDS OF OVERFITTING, SEPARATED BY A COUNT RATHER THAN A MOODThe distinction is not a matter of degree. Look at which number moved.OVERFITTING A MODELOVERFITTING A BACKTESTWHAT GETSADDEDMore numbers inside one rule, until therule can bend around the recordMore attempts at one record, while therule stays exactly as simple as it wasWHATSTAYS THESAMEThe count of attempts can stay at onethroughoutThe count of numbers stays at onethroughout, all thirteen timesWHY IT ISHARD TOSPOTA bending rule usually lookscomplicated, and complicated is visibleEvery rule tried looks defensible,because every rule tried is simpleIN THIS GUIDE, COUNTEDNumbers held by the Ashwin rule: 1, at every one of the thirteen attempts.Attempts made at the record: 1, then 2, then 3, all the way to 13.
Overfitting a model adds numbers to one rule while backtest overfitting adds attempts at one record, and the Ashwin rule held exactly one number at every one of its thirteen attempts.
Try it out

Overfitting a model and overfitting a backtest are not the same thing. What gets added in each?

What should be asked of any backtested figure handed over?

A single habit survives everything else above. A tested figure arrives in a meeting, in a note, in a message at eleven at night, and there is about a minute to weigh it. Here are the seven questions, in the order they cost the least, and what each one is actually protecting against.

The questionWhat it protects against
What record, and how many observations of it?A hit rate off twelve months and a hit rate off seventy two months are not the same kind of claim, and nothing in the figure itself says which one it is.
How many of those could be scored, and where did the rest go?Dropped months change the count sitting underneath the share. Here seven of the 72 went, and all six of the unmoved ones came out of one stretch rather than both.
What is the baseline, and does the figure beat it?Without the line, every reading looks like an achievement. With it, three of these thirteen settings are worse than reading nothing at all.
How many settings were tried before this one was reported?A reading that won a search of thirteen and a reading that was the only one taken are different objects, and they print identically.
Was any stretch held back, and was it looked at first?A stretch that was peeked at before the threshold was chosen is not held back any more, whatever it is called in the write up.
Could every number the rule reads have been known on the day it read it?A signal that quietly includes information from later is the one error that makes a rule look better rather than worse.
What would the same figure be under a different counting convention?Counting the unmoved months as rises instead of dropping them moves the denominator and the count together, and the reading moves with them.

Somebody who works at a lender reads a scoring rule this way as a matter of routine, and the habit is not suspicion. The reason is that a rule chosen from many candidates and a rule written down first are two completely different objects wearing the same clothes, and the only thing separating them is a count nobody volunteers. A backtested figure quoted without its count of attempts is not a small gap in the record; it is a figure nobody can weigh, including the person who quoted it.

The reverse habit is the useful one to build. A tested figure produced in house travels with those five fields every time: the record, the scoreable count, the baseline, the count of settings tried and whether anything was held back. The five fields cost one line. One line buys a figure that somebody who trusts nobody can check, and that is the only kind of checkable worth having.

THE SAME FIGURE, TWICE, AND ONLY ONE OF THEM CAN BE WEIGHEDNothing about the arithmetic changes between these two panels. Five short fields do.AS IT USUALLY ARRIVES67.69 per centhit raterecord: not statedscoreable count: not statedbaseline: not statedsettings tried: not statedheld back stretch: not statedAS IT SHOULD TRAVEL67.69 per centhit raterecord: the six year record, 72 monthsscoreable count: 65 of thembaseline: 56.92 per cent, always upsettings tried: 13 on a stated gridheld back stretch: none, not yetThe panel on the left cannot be argued with, because there is nothing init to argue with. The panel on the right can be checked, and one of itslines is bad news.
The same reading of 67.69 per cent is either a checkable claim or an unweighable one depending entirely on whether five short fields travel beside it.
Try it out

A colleague sends a backtested hit rate with no other detail at all. Which three questions come first?

The failure: writing down 67.69 per cent and walking away with it

Somebody reads the ladder, sees 67.69 per cent at the top of it, and concludes that the Ashwin rule calls direction right about two times in three. The conclusion goes into a note. Nothing they did was dishonest and nothing they computed was wrong. The figure is exactly what it says it is: 44 matched calls out of 65 scoreable months of the six year record.

Two things were skipped, and both of them were sitting alongside the figure itself. The first is the baseline. Calling every month up and reading nothing at all lands 56.92 per cent on those identical 65 months, so the whole of what the rule appears to add is 10.77 points rather than 67.69 per cent. The second is the count of attempts. 67.69 per cent is the best of thirteen readings that run down to 43.08 per cent, and the best of thirteen is not the same kind of number as a single reading, even though it is written in exactly the same way.

The cost is not hypothetical on this record, and that is the uncomfortable part. Carry the same threshold across to the 36 months nobody chose and it calls 18 of them. The note was written, the figure travelled, and the months that would have contradicted it were already in the record when the note was written.

The habit that prevents it costs less than the quoting did. Before any backtested figure goes into a sentence, ask for the baseline and ask how many settings were tried. If neither answer is available, the figure is not ready to be written down yet, and saying so is a great deal cheaper than un-saying it later.

Subjects covered elsewhere. Holding a stretch of the record back before anything is chosen, and reading the rule on it afterwards, is set out separately, and so is rolling that hold out forward through time rather than doing it once. Running many tests against one threshold, the way choices made while analysing one result change what gets reported, and the effect of picking the question itself after looking at the record each have a treatment of their own. Placing an order, paying a cost, moving money, or deciding whether any tested rule is a strategy worth anybody's time, is covered separately.

Risk Management Program Bootcamp — Fin Maverick

Who stands behind the figures printed above?

No outside source stands behind the figures. The recipe is printed instead, and the table below says where each part of it lives.

The input usedWhat it is
The arithmetic check kept with these notesA script that rebuilds the six year record from its three parts and recomputes every count printed above, then refuses to finish if any of them moves
The six year record of the Nakshatra unit72 invented monthly observations, settled where the record was built and reused here without a single figure being changed
The Ashwin rule and its thirteen settingsOne sentence and a stated list of thresholds, both fixed before any count was taken
Everything elseNothing else. No outside body, no maintained series, no published finding and no record of anybody's results stands behind any figure here

The Nakshatra unit, its six year record and the Ashwin rule are invented.
Educational material. Not advice on any investment, tax, budget or market position.

Next →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.