Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quantitative Methods, Financial Data & Programming
1Probability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
2Statistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
3Correlation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
4Time Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
5Simulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
6Optimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
7Modelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
8Backtesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
9Data Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
10Programming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
11Quantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity

Spurious Correlation: Patterns With No Mechanism

A spurious correlation is a strong, correctly computed pattern between two series that nothing connects. Across ten invented months, a tally of chairs put out in a community hall tracks the Vasant unit at 0.9502 and accounts for 90.29 per cent of its movement, ahead of the traded input's 75.59 per cent. No sum that can be run on those ten months tells the two apart.

Two things are settled already. A line has already been put through ten pairs of monthly readings and the share of the outcome it accounts for has already been read off. Separately, the notes on what a correlation actually claims set out the readings a large one is consistent with, and one of those readings is the uncomfortable one: the two series have nothing whatever to do with each other, and the pattern belongs to the particular months in hand rather than to anything in the world. The notes on what a correlation claims allowed that reading a single counterexample and named spurious correlation as the place it is worked through. The uncomfortable reading is the one with no arithmetic defence at all, and a reading with no defence deserves more than a footnote.

Three invented objects carry the work, and two of them will be familiar. The Nakshatra unit and the Vasant unit are both a traded unitA vague label for something carrying a market price. Which kind of thing it is never has to be settled. Not one division below would come out differently., and what is written down against each is its monthly change: a price at the close of a month set against the same price at the start of it, expressed as a share of the opening figure. Of the two, the Nakshatra unit sits in the input seat and the Vasant unit is what gets explained. Row by row the two columns cover identical months. Identical months are what make each row a paired observationA row in which two readings belong to one and the same occasion. The shared occasion is the whole licence for comparing them, and a record whose rows have been shuffled has quietly withdrawn it. and what make any comparison between them legitimate at all. Nothing below re-sorts them, and they stay in time orderRows left in the order the months arrived, oldest at the top. Keeping it costs nothing and cannot be undone once a column has been sorted by size instead..

The record everything in this guide is computed from, before the third column arrives.
MonthThe Nakshatra unit, per centThe Vasant unit, per cent
11.003.00
26.0018.50
3minus 4.002.50
411.0017.00
51.00minus 3.00
6minus 9.00minus 13.00
76.006.50
81.00minus 1.00
9minus 4.00minus 7.50
101.00minus 3.00
Average1.002.00

The third object arrives shortly, and it is the one that makes the trouble.

What makes a correlation spurious?

Three things the word does not mean have to be cleared away first. Spurious does not mean weak. Spurious does not mean small. Spurious does not mean miscalculated. The correlation worked here is 0.9502. A figure that high is about as strong as anything an analyst will meet, and it is correct to the last decimal place: anyone who runs the same sums on the same twenty numbers gets the same answer, today and in thirty years.

A correlation is spurious when there is no mechanism: no account, however plain, of how a movement in one series could reach the other.

Spurious is therefore a claim about the world rather than a claim about the number, and that distinction puts it out of reach of every check that can be run on the record. A number can be checked against the record it came from. A mechanism has to be checked against the world, and the record is not the world. The record is twenty figures on a sheet of paper.

The shape is easier to feel at household scale. Try it there. Over six months in one home the number of times the doorbell rang each month climbed steadily, and so did the electricity bill. Line the two columns up and they move together beautifully. Now try to write the sentence that connects them. Not a story about how they might both be caused by something else. A common cause sitting behind both is a different situation with a different name, covered separately. Just the direct sentence: a visitor at the door causes electricity to be consumed, or electricity consumed causes visitors to arrive. Neither sentence survives being said out loud. The pattern is real, the sums are right, and there is nothing there.

Try it out

Define spurious without leaning on the words weak, small or wrong. Which of these is the working definition?

Breaking Into Quants Bootcamp — Fin Maverick

What sits behind the Kadamba count?

Nothing. The one word is the whole answer, and it has to be accepted completely before the next section starts. The next section is built to shake it.

The Kadamba count is the number of chairs put out each month in a made up community hall that nobody has ever walked into. Somebody unstacks a certain number of chairs before whatever is happening that month and stacks them again afterwards, and a made up register records the figure. Over the ten months it reads 58, 79, 66, 72, 53, 45, 60, 59, 53 and 55, averaging 60 chairs, never dipping below 45 and never passing 79. Each entry is a whole numberA count with no fraction in it, the way 58 chairs is a count. Nothing sits between 58 and 59. Things that are tallied behave that way, and things that are measured do not., because two thirds of a chair cannot be put out.

There is no connection of any kind between that hall and any traded unit, and none is being hinted at. The count is not a disguised version of some well known example half remembered from elsewhere, and it was not picked to resemble one. A reader who starts constructing a story about busy months or seasons is the one supplying that story. The register does not.

THE KADAMBA COUNT, TEN MONTHS IN TIME ORDER chairs put out in an invented community hall, and nothing else 40 50 60 70 80 58 m1 79 m2 66 m3 72 m4 53 m5 45 m6 60 m7 59 m8 53 m9 55 m10 AVERAGE 60 CHAIRS No price, no rate, no market and no cause sits behind any of these ten bars. Hold on to that.
The chair count runs from 45 to 79 with an average of 60, and there is no price, rate, market or cause anywhere behind those ten bars.
Try it out

Before a single figure in the next section is read: the chair count is about to be fitted against the Vasant unit. What correlation should be expected?

Can a series with nothing behind it beat the real input?

Yes, and on these ten months it does not merely squeak past. The chair count wins on every single measure the arithmetic is capable of producing.

The procedure is the one used on the opening notes, applied without a single change. Take the outcome, take a candidate input, and find the straight line through the ten points that leaves the smallest total when every vertical miss is squared and added up. On the Nakshatra unit that line came out at 0.5000 plus 1.5000 times the input, leaving 218.00 in squared misses. Run the identical procedure with the chair count in the input seat and the line comes out at minus 54.9799 plus 0.9497 times the count, leaving 86.73. The 0.9497 is the coefficientThe number a fitted line multiplies an input by. One step up in the input moves the fitted outcome by that much. The coefficient is a property of the line, not a promise about anything. on the chair count, and it says that the line moves the outcome up by 0.9497 percentage pointsThe unit in which a change to a per cent figure gets counted. A move from 4 per cent up to 7 per cent is three percentage points of movement, and calling that a three per cent rise would be a different and much smaller claim. for every extra chair. At the average of 60 chairs the line reads exactly 2.00 per cent, the average of the outcome. The match is no coincidence: a least squares line always passes through the point where both averages meet.

THE TRADED INPUT the Vasant unit against the Nakshatra unit THE CHAIR COUNT the same outcome against the Kadamba count minus 10 0 10 20 the Nakshatra monthly change, minus 11 to 13 per cent the chair count, 43 to 81 chairs RED LENGTHS SQUARED AND ADDED: 218.00 RED LENGTHS SQUARED AND ADDED: 86.73
Every red stick is one month's miss from its own fitted line, and the sticks on the chair count side are plainly shorter: 86.73 in squared misses against 218.00 for the traded input.

Here is the whole record, the same ten months as everywhere else and in the same order, with the chair count added as a third column. Twenty numbers became thirty, and nothing else changed.

The ten months with the chair count added. The right hand column is the only new thing in the table.
MonthThe Nakshatra unit, per centThe Vasant unit, per centThe Kadamba count, chairs
11.003.0058
26.0018.5079
3minus 4.002.5066
411.0017.0072
51.00minus 3.0053
6minus 9.00minus 13.0045
76.006.5060
81.00minus 1.0059
9minus 4.00minus 7.5053
101.00minus 3.0055
Average1.002.0060

The two fits now sit side by side, to be read down slowly. In this worked instance the wrong answer wins every row.

Two fits against the same outcome over the same ten months, in time order. The right hand column is the chair count.
What was measuredFitted on the Nakshatra unitFitted on the Kadamba count
The fitted line0.5000 plus 1.5000 times the inputminus 54.9799 plus 0.9497 times the count
Correlation with the outcome0.86940.9502
R squared0.75590.9029
Explained sum of squares675.00806.27
Squared misses left over218.0086.73
Root mean squared miss4.6690 per cent2.9451 per cent
Average absolute miss3.6000 per cent2.5302 per cent
Largest single miss9.0000 per cent5.1980 per cent
SEVEN WAYS OF ASKING WHICH LINE FITS BETTER WHAT WAS MEASURED NAKSHATRA UNIT KADAMBA COUNT WINS Correlation with the outcome 0.8694 0.9502 chairs R squared 0.7559 0.9029 chairs Explained sum of squares 675.00 806.27 chairs Squared misses left over 218.00 86.73 chairs Root mean squared miss, per cent 4.6690 2.9451 chairs Average absolute miss, per cent 3.6000 2.5302 chairs Largest single miss, per cent 9.0000 5.1980 chairs SEVEN MEASURES, SEVEN WINS FOR THE COLUMN WITH NOTHING BEHIND IT Choosing a different measure of fit would not have changed the verdict, because all seven read the same two columns.
Seven different measures of how well a line fits, and the chair count takes all seven against the traded input on the same ten months.

The chairs beat the traded unit on every number the arithmetic can produce. Not on a favourite measure, not on one chosen after the fact: on the correlation, on R squared, on the amount explained, on what is left over, on both averages of the misses, and on the worst month. All of the measures are computed from the same two columns, so a different measure would have gone the same way.

Now the sentence that matters more than any of the figures. Nothing about the chair count has become connected to anything. The hall is exactly as unconnected as it was three paragraphs ago. Not one chair moved. Confidence in the count wobbles somewhere between the bar chart and the table, and that wobble is worth more than the figures that caused it. The arithmetic did not learn anything about the world. A large number simply appeared and started being treated as evidence, and that is what everybody does.

One more line, so nothing is left hanging. The chair count also correlates with the Nakshatra unit, at 0.7145. The 0.7145 is high enough to notice and low enough to rule out the tidy escape route. The chairs are not just a copy of the input wearing a disguise, and the trouble here is not the trouble caused by two inputs repeating each other, a separate subject with a separate name. The chair count is a third column that is genuinely its own thing, genuinely unconnected, and genuinely better at the job.

Try it out

The chair count reaches an R squared of 0.9029 and the traded input reaches 0.7559. Which is the better model?

Does cutting the record in half catch it?

Splitting the record is the first defence anybody reaches for, and it is a sensible one. If a pattern is an accident, surely it will be an accident of a particular stretch of the record. Cut the ten months into the first five and the last five, measure the correlation separately in each, and a pattern that only lives in one place will show itself by collapsing in the other. Run it and see.

The same split applied to both candidates. The two right hand rows are the ones worth staring at.
Half of the recordThe chair count against the outcomeThe Nakshatra unit against the outcome
Months 1 to 50.93080.7969
Months 6 to 100.93220.9863
Distance between the two halves0.00140.1893

The check passes cleanly on the series with no mechanism whatever, and it passes it more convincingly than it passes the real input. The chair count comes back at 0.9308 and 0.9322, two readings that agree to within fourteen ten thousandths. The traded input comes back at 0.7969 and 0.9863. The two readings differ by 0.1893, more than a hundred times the gap. On a report that ranked candidates by stability across halves, the chairs would finish first and the input anyone would have guessed at would finish behind them.

CUT THE TEN MONTHS IN HALF AND MEASURE EACH HALF SEPARATELY THE KADAMBA COUNT, THE SERIES WITH NOTHING BEHIND IT MONTHS 1 TO 5 0.9308 MONTHS 6 TO 10 0.9322 The two halves differ by 0.0014, so the check passes about as cleanly as a check can pass. THE NAKSHATRA UNIT, THE TRADED INPUT MONTHS 1 TO 5 0.7969 MONTHS 6 TO 10 0.9863 The two halves differ by 0.1893, more than a hundred times the gap above. A CHECK THAT A SPURIOUS SERIES PASSES IS NOT A CHECK FOR SPURIOUSNESS
Both halves of the record confirm the chair count at 0.9308 and 0.9322, while the traded input's halves disagree by 0.1893, so stability across halves ranks the wrong series first.

So what does the split actually establish? Something narrow and genuinely worth having: that the pattern is not confined to one stretch of the record. The split rules out the case where a single dramatic month, or one unusual quarter, carried the whole relationship and the rest of the record contributed nothing. Such a case exists and the check finds it, so the habit is a good one. The check cannot do anything else. A check that a spurious series passes is not a check for spuriousness, and no amount of enthusiasm for the check changes that.

Consider a pot of dal tasted from the top and again from the bottom. Both spoonfuls taste the same, and that establishes something real: the salt is evenly distributed and nobody stirred badly. The even taste establishes nothing at all about whether this is the dish that was ordered. Two consistent tastings of the wrong dish are still the wrong dish, consistently.

Try it out

The split passes at 0.9308 and 0.9322 on a series with nothing behind it. What has the check actually proved?

Bond Pricing and Yield Mechanics — free micro-course from Fin Maverick

Is there any arithmetic that separates the two cases?

No. Not this arithmetic, not better arithmetic, not arithmetic somebody will invent next year. The reason is structural and it takes one paragraph to see, after which it cannot be unseen.

Every quantity in this guide was computed from two columns of figures. The correlation, the coefficient, R squared, the squared misses, the spreadA one word summary of how widely a column of figures is scattered about its own centre. Two columns can share a centre exactly and still differ completely in this. of each column, the two half record readings: all of them are functions of twenty numbers and nothing else. Now look at what the chair count hands over and what the Nakshatra unit hands over. Both hand over a column of ten figures. Both columns have a centre, a spread and an order. Neither column carries a tag saying what produced it. Which column came from a price and which came from a stack of chairs was never inside either column to begin with, so once the numbers are written down the arithmetic cannot tell them apart.

WHAT THE ARITHMETIC IS HANDED, AND WHAT IT IS NEVER HANDED THE OUTCOME 3.0 18.5 2.5 17.0 minus 3.0 minus 13.0 6.5 minus 1.0 minus 7.5 minus 3.0 INPUT ONE 1.0 6.0 minus 4.0 11.0 1.0 minus 9.0 6.0 1.0 minus 4.0 1.0 INPUT TWO 58 79 66 72 53 45 60 59 53 55 EVERY QUANTITY IN THIS GUIDE the centre and the spread of each column the correlation between any two of them the slope and the intercept of a fitted line R squared and the explained sum of squares the squared misses and the absolute misses the same figures again on either half All six read those columns and nothing else. OUTSIDE EVERY COLUMN, AND UNREACHABLE FROM ANY OF THEM Whether anything in the world joins input two to the outcome. It was never written into a column, so no quantity computed from the columns can report it.
Both candidate columns feed the same arithmetic and get the same kinds of answer back, while the question of whether a path exists in the world sits outside both columns entirely.

None of this is a complaint about these particular measures. The correlation is not crude in a way that something more sophisticated would fix. Any quantity that can be defined is defined on the columns, so any quantity that can be defined is blind in exactly the same place. Asking for a cleverer statistic to detect spuriousness is like asking for a more accurate thermometer to settle whether a room is beautiful. The instrument is fine. The thermometer answers a question about temperature, and beauty was never in the reading.

Try it out

Name the one thing that differs between the two candidates. Where does it live?

Try it out

Somebody proposes a rule: any correlation above 0.90 should be treated as a real relationship. What is wrong with it?

Hypothesis Testing teaches you to run a test, say what it can and cannot support, and recognise a manufactured result.

What separates them, if no number can?

Four things do the separating, and the striking feature of the list is that not one item on it is a quantity that can be computed from the record. The absence is not an accident. If any of them could be computed from the record, the section above would be wrong.

THE FOUR THINGS THAT SEPARATE THE TWO CASES SEPARATOR 1 A STATED MECHANISM Written as a sentence before the fit is run. If none can be written, the file records that it could not be. SEPARATOR 2 A PREDICTION IT MAKES Something the account says that the pattern on its own does not. Then go and look at whether it happened. SEPARATOR 3 FRESH MONTHS Readings the relation was not chosen on. The months it was picked from cannot test it. SEPARATOR 4 HOW MANY WERE TRIED One candidate, or two hundred. The pattern looks identical; what it is worth does not. Write the count down. NOT ONE OF THE FOUR IS A NUMBER COMPUTABLE FROM THE RECORD If any of them were, the section above would be wrong.
A stated mechanism, a prediction it makes on its own, months it was not chosen on, and the number of candidates tried, none of them computable from the record.

The first is a mechanism, written down as a sentence before the fit is run. Not after, and the ordering is the entire point. Written afterwards, a mechanism is a story assembled to fit a number already seen, and a competent person can assemble one for any pair of columns in about ninety seconds. Written beforehand it is a commitment that the record can then embarrass. A street vendor who says on Monday that fewer people stop at the cart when it rains has made a claim that Tuesday can contradict. A vendor who looks at a slow week and then explains it by the rain has made no claim at all.

The second is a prediction the mechanism makes that the pattern on its own does not. This is the sharp one and it does real work. If chairs reached prices somehow, that account would have to say something further: something about what happens when far more chairs are put out than usual, or about how long the effect takes to arrive, or about what should happen in a hall that closed for two months. Each of those is a separate thing that can be gone out and checked. The bare pattern makes no such commitment, and the missing commitment is precisely what is comfortable and useless about it. A mechanism that predicts nothing beyond the pattern it was invented to explain has added nothing to the pattern.

The third is fresh months: readings the relationship was not chosen on. The chair count was selected because it fit these ten months well. Testing it on these ten months is asking it to repeat the one thing it was picked for. The point is not a subtle one, but it is skipped constantly. The months are already in the file, and fresh ones require waiting.

The fourth is the count of how many candidates were tried, and it is the one that gets left out. A correlation of 0.9502 found on the first series anybody looked at is one kind of finding. The same 0.9502, found by running two hundred series against the outcome and keeping the best, is a different kind of finding entirely, and the two are indistinguishable once written up. The pattern is identical. The difference is how many chances the pattern had, and that number lives in the analyst's process, not in the data. The question is not how strong the pattern is but how many chances it had.

Try it out

Two people report a correlation of 0.9502 against the same outcome. One tested a single series chosen in advance. The other tested two hundred and reported the best. Should the two reports be read the same way?

Why does adding more months not fix this?

Because more months improve a different property from the one that is broken. Suppose the pattern held at 0.9502 over a hundred and twenty months instead of ten, and then over a thousand. The interval around the estimate tightens: on ten months an interval built around 0.9502 in the usual way runs roughly from 0.80 to 0.99, a width of about 0.19; at a hundred and twenty months it is about 0.04 wide; at a thousand it is about 0.01 wide. The sampleThe part of a record actually in hand, as against everything that could have been written down. More of it steadies an estimate without turning it into an estimate of something else. grows, the estimate steadies, and that is worth having. Notice what kind of gain it is: a gain in precision, and only that.

THE SAME CORRELATION OF 0.9502, MEASURED ON MORE AND MORE MONTHS 0.70 0.80 0.90 1.00 ten months, the record actually in hand width 0.1905 ten years of months, if the pattern held width 0.0359 a thousand months, if the pattern held width 0.0121 THE MECHANISM, AT EVERY ONE OF THOSE ROW COUNTS: STILL EMPTY
The interval around the same 0.9502 shrinks from about 0.19 wide at ten months to about 0.01 wide at a thousand, while the space for a mechanism stays empty at every row.

A spurious relationship measured over a thousand months is a spurious relationship measured precisely, and precision and correctness are different properties. Reading the correlation to a fourth decimal place says how firmly the figure is known. The fourth decimal says nothing at all about whether the figure is measuring a relationship. Nobody has written anything in the empty box beside those shrinking intervals, and nothing done to the columns writes anything in it. The box is the same size at ten months and at a thousand.

Counting the chairs for ten more years yields a very well characterised chair count: its average known to a fine tolerance, its spread known, its month to month behaviour known. Counting something carefully is not the same activity as finding out what it touches, so the count will be no more connected to a traded unit than it was on the first day.

Try it out

A thousand more months arrive and the correlation still reads 0.9502. What has changed about the problem?

What is written down when no mechanism can be named?

An analyst handing a finding to somebody who will act on it lives in exactly this situation. A mechanism cannot always be named. The situation is common and it is not by itself a disgrace. The write up decides whether the work is honest, and there are four lines to add. Each costs a sentence.

THE SAME FINDING, WITH THE THREE FIELDS THAT DECIDE IT FINDING, AS FILED OUTCOME EXAMINED the Vasant unit, ten months in time order CANDIDATE REPORTED the Kadamba count CORRELATION 0.9502 R SQUARED 0.9029 HALF RECORD CHECK passed, 0.9308 and 0.9322 MECHANISM WRITTEN BEFORE THE FIT left empty, and nobody said so CANDIDATE SERIES TRIED 200 MONTHS TESTED SINCE SELECTION none THIS IS A PATTERN. IT IS NOT YET A RELATIONSHIP.
A finding with the mechanism field blank, two hundred candidates tried and no fresh months tested is a pattern rather than a relationship.

The mechanism goes down as a sentence before anything is run, and if it cannot be written, the file records that it could not. A blank field that somebody left blank on purpose is information. A blank field nobody thought about is a hole with the same shape, and by the time the finding reaches a reader the two look identical. The person who reads the work in six months cannot tell which one they are holding unless the write up says.

Second, state how many series were considered. Two hundred candidates screened and one reported is a perfectly reasonable thing to have done; it becomes misleading only when the two hundred goes unmentioned and the reader takes the survivor for a first guess that happened to land. Third, say plainly whether the relationship was chosen on the same months it is being reported on. If it was, the report is describing a fit rather than a test, and those are different claims.

The fourth line, and the cheapest

A pattern is reported as a pattern rather than as a relationship. The distinction costs one word. The word changes what a reader is entitled to do with the finding, and it is the whole of the honesty available once the mechanism field is empty.

Notice what none of this does. None of it makes the finding safe, and saying so is itself the finding. The four lines do not upgrade a pattern into a relationship; they describe accurately what is being handed over, so that the person receiving it can decide how much weight it will bear. A household choosing where its one salary goes, a lender deciding what to lend against, an investor sizing a holding: each of them is entitled to know whether the thing in front of them survived a test or merely fit a file. Writing that down is not modesty. Writing it down is the only part of the job that a stronger correlation could never have done.

Try it out

The fit is excellent and no mechanism can be named. Which sentence belongs in the report?

The screen that came back clean, and what it cost

An analyst is asked which series moves with the Vasant unit. There is a shared drive with a couple of hundred series on it, so she correlates every one of them against the outcome, sorts the list from high to low, and keeps the top row. The top row is the chair count at 0.9502, with an R squared of 0.9029, comfortably ahead of the traded input anybody in the room would have nominated. She is careful, so she cuts the record in half and checks: 0.9308 and 0.9322. The correlation holds in both halves. She writes it up.

Every step there was competent, and not one figure in the report is wrong. The report is worthless. The reason is in none of the numbers, and that is why nobody catches it in review: no account of how a chair could reach a price was ever written down, and the number two hundred never left her own machine.

The cost is specific and it repeats. The relationship holds on exactly the months it was picked on and nowhere else, so the first month it is used on is the first month it has ever actually been asked to do anything.

The fix is an order of work rather than an extra test, and it is free. Write the mechanism as a sentence before the sums are run. Record how many candidates were considered, in the file, where a reader can see it. And keep calling it a pattern until months nobody chose it on have had a look at it.

Where this subject ends. What a correlation claims, and the four readings a strong one is consistent with, are covered separately; the fourth of those readings is the one worked through here. The different trouble caused by two inputs that repeat each other is also covered separately, and it is a distinct failure with a distinct symptom. Screening candidate series against an outcome is dealt with elsewhere, and the practice is itself the failure described just now.
Breaking Into VC Bootcamp — Fin Maverick

What was consulted

What a reader might reasonably look for behind the figures, and what is actually there.
What a reader might expect to find namedWhat is actually behind the figures
A regulator, a supervisor or a trading venueNone, and the emptiness is a finding rather than a gap in the reading. Fitting a line through ten pairs of numbers is arithmetic. It carries no limit, no threshold and no filing period from any market, so a named authority here would be decoration.
A dated price record behind the ten monthsNone. The ten months were constructed so that the sums close exactly, and constructed months carry no as-of date and no venue.
A well known published case of a pattern with nothing behind itNone used. The famous ones are somebody's own published work, and a chair count answers the question just as well without borrowing anybody's.
The arithmetic behind every number aboveLeast squares on ten paired rows, and nothing further. Every quantity follows from the three columns printed in the tables, so anybody willing to run the same sums on a sheet of paper reproduces the lot.

The Nakshatra unit, the Vasant unit, the Kadamba count and the community hall are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.