Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quantitative Methods, Financial Data & Programming
1Probability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
2Statistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
3Correlation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
4Time Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
5Simulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
6Optimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
7Modelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
8Backtesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
9Data Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
10Programming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
11Quantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity

Regression, Correlation and Fit: One Calculator, Ten Months

Two columns of paired readings typed into the table at the top of this calculator return seventeen rows: the fitted line, the correlation, the covariance and four error measures, each step shown as it is built. The ten invented months it opens on give a line of 0.50 plus 1.50 times the input, a correlation of 0.8694, an R squared of 0.7559 and a lag one reading of 0.4862 on the leftover misses.

The two ten month records were built so that the arithmetic lands exactly rather than nearly. Nothing was lifted out of a market, a filing or a database, so no as-of date exists to be quoted.

Paired readings typed in, and the six steps as they build

The table below is the instrument. The table opens on the ten month record this calculator is written around, so the figures it prints the moment it loads are the same ones printed as ordinary text further down. Any cell can be overtyped with fresh readings, rows added and rows dropped, and the six steps, the reconciliation and the chart all recompute as the typing happens. The reconciliation is the part worth watching. The reconciliation puts the input average back into the line the six steps just built and shows the answer landing on the outcome average, with a gap of nothing between them, on every set of pairs that can be typed. The fitted line runs through the crossing of the two averages by construction, never by luck, and that holds on the unrounded figures rather than on the four decimal places shown.

Type it in

Paired readings in, and every step of the fit shown as it is built

Load a record to start from:
OccasionInput readingOutcome reading
1
2
3
4
5
6
7
8
9
10

Input reading. The column being fitted from. One figure per occasion, read off the row of the record where that occasion sits.

Outcome reading. The column being fitted to, taken off the same row of the same record as the input beside it.

Occasion. The row number in the record, counted in the order the readings were taken and never re-sorted.

Both columns. Whatever unit the record itself carries. The tool never converts, so a column kept in decimals and a column kept in per cent will not agree.

The six steps, on whatever is in the table above

Count of pairs
Slope
Intercept
Correlation
R squared
Left over

Every occasion's contribution to the three sums

OccasionInputOutcomeInput distanceOutcome distanceSquared inputMultipliedSquared outcomeFittedMiss
Educational illustration. Every figure printed here is arithmetic on the numbers in the table above and nothing else. A fitted line describes the occasions it was fitted to and reaches no further, so no returned row forecasts an eleventh occasion. Nothing typed in is stored, sent anywhere or kept once the tab closes.

Nothing else is consulted, and no figure below arrives from any source other than the pairs in that table. The published ten describe ten months of an invented traded unitSomething whose price buyers and sellers keep agreeing on, so a record of what it changed by each month can exist at all. Both units named here were invented to be typed into a table. called the Nakshatra unit and a second invented one called the Vasant unit. The meaning of each returned figure is covered separately. The calculator adds the arithmetic behind each one, printed in full with the tool switched off, and a second panel further down that moves a single month so that every row can be watched responding.

What does this calculator take, and what does it hand back?

The calculator takes up to ten paired observationsTwo readings lifted off one and the same occasion and kept side by side. Slip one column against the other by even a single row and every figure underneath describes something that never happened., in the order they actually happened, and it hands back seventeen rows. Three of those seventeen are not really results at all: the count of pairs, and the average of each column. The count and the two averages are handed straight back so that the tool can be confirmed to have read what was intended. The other fourteen are worked out, and every one of them is worked out from the same two columns.

A fit report has no second source of facts standing behind it, and that single origin is the whole reason for gathering fourteen figures a reader usually runs into one at a time into a single table. Met separately, the slope looks like one kind of thing and the correlation looks like another, and it is easy to carry away the impression that a fit report is assembled from several different investigations. It is not. A fit report is one small table of numbers, run through sums, and the sums are then divided by different things.

The everyday version needs no market at all. A household notebook carries two columns for twelve months: what the electricity bill came to, and how many days the air cooler ran. Everything anyone can honestly say about how those two travel together comes off that single sheet of the notebook. There is no thirteenth column arriving from elsewhere to confirm it. If the notebook is wrong, every conclusion is wrong, and no amount of arithmetic afterwards will say so.

Two columns in, seventeen rows out, and one source for all of them the ten paired months of the Nakshatra unit and the Vasant unit, invented for teaching WHAT GOES IN MONTH INPUT OUTCOME 11.003.00 26.0018.50 3minus 4.002.50 411.0017.00 51.00minus 3.00 6minus 9.00minus 13.00 76.006.50 81.00minus 1.00 9minus 4.00minus 7.50 101.00minus 3.00 in time order, never sorted HANDED STRAIGHT BACK 3 rows the count of pairs, and the average of each column THE THREE SUMS 3 rows input movement 300, movement together 450, outcome movement 893 THE LINE 2 rows slope 1.5000, intercept 0.5000 THE STRENGTH 4 rows covariance 50.0000, correlation 0.8694, then 0.7559 and 0.7254 THE MISSES 4 rows 21.8000, then 4.6690, 3.6000 and 5.2202 per cent THE ONE ROW THAT USES THE ORDER 1 row the lag one reading of the misses, 0.4862 3 handed back plus 14 worked out is 17 rows, all from the same two columns
Fourteen of the seventeen returned rows are worked out, and all fourteen come out of the same ten pairs, with nothing arriving from any other source.
Breaking Into Quants Bootcamp — Fin Maverick

How does it get the line out of the pairs?

Two sums do the work. The first asks how much the input column moves about on its own: take each input away from its average of 1.00, square the result so the direction stops mattering, and add the ten squares. The ten squares come to 300. The second asks how much the two columns move together: take each input away from 1.00, take each outcome away from 2.00, multiply the two distances for that month, and add the ten products. The ten products come to 450.

The slope is the second divided by the first. Four hundred and fifty over three hundred is 1.5000. The fitted line always passes through the pair of averages, so the intercept is whatever number puts it there: 2.00 less 1.50 times 1.00 is 0.5000. So the line reads 0.50 plus 1.50 times the input, and the calculator has finished its main job in two divisions and one subtraction.

Both of those figures are exact on the ten default months, not rounded, so a reader who works them out on paper will match the calculator to the last decimal. The ten months were built that way on purpose, leaving no rounding artefact to be explained away. Readings typed in from elsewhere will almost certainly arrive as 1.4997 and 0.5013 instead. A record that was not designed looks exactly like that.

The ten pairs, the fitted line, and the ten misses drawn as gaps the line is 0.50 plus 1.50 times the input, and it passes through the pair of averages by construction the pair of averages, 1.00 and 2.00 per cent months 5 and 10 share this one point the fitted line hit exactly by the line the miss for that month minus 10 minus 5 0 5 10 minus 10 0 10 20 the Nakshatra unit, monthly change in per cent the Vasant unit, per cent
Two months land on the fitted line with nothing left over, two more share a single point, and the line crosses the pair of averages at 1.00 and 2.00 per cent.
Try it out

The slope of 1.5000 is to be reproduced on paper, from the two columns alone. Which set of sums is the smallest one that gets there?

Where do the correlation and the covariance come from?

From the same 450. Reusing that one sum is the part worth slowing down for. The covariance is the movement together divided by nine, one less than the count of pairs, giving 50.0000. Two percentage figures were multiplied together to build the covariance and nothing has taken the square root since, so its units are per cent squared. The correlation is the covariance divided by the two spreadsHow far a column of numbers usually sits from its own centre, written as one number. Covered separately, and used here only as something to divide by., 5.7735 per cent for the input and 9.9610 per cent for the outcome. The answer is 0.8694, and it carries no units at all.

The slope, the covariance and the correlation are one quantity read three times under three different divisors. Swapping any of them for another breaks the reading. 450 divided by the input's own movement of 300 gives the slope. The same 450 divided by nine gives the covariance. Divided by nine and then by both spreads, it gives the correlation. One numerator, three denominators, three answers that mean three different things and are quoted in three different sets of units.

Every one of those denominators is a whole numberA figure that stops at the decimal point, such as ten months or the nine gaps between them. Every divisor named here is one, even where the thing being divided is not. or a length built from the data itself, and not any of them was chosen by preference. The nine is the count of pairs less one. The 300 is a sum that can be added up by hand in a minute.

One numerator, three denominators, three different answers the movement together of the two columns, computed once and then divided three ways 450 the movement together divide by 300 the input column moving about on its own 1.5000 THE SLOPE outcome points per input point divide by 9 one less than the ten pairs entered 50.0000 THE COVARIANCE per cent squared by 9, then both spreads 5.7735 and 9.9610 per cent, multiplied together 0.8694 THE CORRELATION no units at all one 450 on top of all three, so swapping any two of them swaps the denominator
Dividing the movement together of 450 by the input movement gives the slope, by nine gives the covariance, and by nine and both spreads gives the correlation.
Try it out

The slope is 1.5000 and the covariance is 50.0000. Both are built on top of the same 450. Which divisor produced each one?

Try it out

Suppose every figure in the Vasant column were doubled and nothing else changed. Which of these describes what happens to the returned rows?

How is R squared built out of the misses?

Not from a formula that has to be trusted. From a chain of three numbers that can be added up by hand. The worst honest attempt available is the place to start: the same figure guessed for every month, and that figure the outcome column's own average of 2.00 per cent. Each month taken away from 2.00, squared, and the ten squares added, comes to 893. A miss of 893 is how badly the flattest possible answer does.

The line is then fitted and the identical exercise run against it. Each month taken away from what the line predicted for that month, squared, and the ten squares added, comes to 218. The line missed by 218 where the flat guess missed by 893, so the line accounted for the remaining 675. R squared is that difference as a share of the starting point: 675 over 893 is 0.7559.

Adjusted R squared then charges a fee. Fitting a line used up two of the ten pairs, in the sense that two numbers were pulled out of the data to build the line itself, so the adjusted figure divides the leftover 218 by eight rather than by ten and divides the starting 893 by nine rather than by ten. The adjusted figure works out at 0.7254. The gap between the two is 0.0305 on these ten months, and it widens with every extra input added. Widening is the entire reason the second figure exists. On one input the fee is small. The fee stops being small quickly.

R squared is a ratio to add up, not a formula to accept the same ten months, scored twice: once against a flat guess, once against the fitted line GUESS 2.00 PER CENT EVERY MONTH misses by 893 in squared terms FIT THE LINE INSTEAD 218 left 675 accounted for 675 out of 893 is 0.7559 charge for the one input used: divide 218 by 8 and 893 by 9 instead 0.7254 the gap of 0.0305 is on one input alone, and it widens with every input added
The flat guess misses by 893 and the line misses by 218, so the difference of 675 over 893 is the R squared of 0.7559.
Try it out

R squared is 0.7559 and adjusted R squared is 0.7254 on the same ten months. Which of these is that gap of 0.0305 charging for?

AI For Finance Bootcamp — Fin Maverick

Which error measure is read against the outcome?

Four of them come back and they are not interchangeable. The nuisance eases once only two things turn out to separate them: what gets added up, and what it gets divided by. The mean squared error adds the ten squared misses to 218 and divides by ten, giving 21.8000. Its units are per cent squared, so putting it next to a monthly change of 6.50 per cent is comparing an area with a length. The mean squared error is a fine thing to minimise and a poor thing to quote.

The root mean squared error takes the square root of that same 21.8000, giving 4.6690 per cent. A figure in per cent can sit beside a monthly change without embarrassment. The mean absolute error skips the squaring entirely: add the ten miss sizes to 36, divide by ten, and get 3.6000 per cent. The mean absolute error comes in below the root mean squared error because it does not let the two big months of 9.00 and 8.00 dominate. Letting a bad month dominate or not is a choice, never a neutral default.

The fourth is the one that causes the trouble. Two degrees of freedomHow many readings are left to vary after some have already gone into producing a figure. Building the line here uses up two of the ten before any counting starts. were spent fitting the line, so the residual standard error divides the same 218 by eight rather than ten. Its square root is 5.2202 per cent. The residual standard error is the figure most often reported as the typical miss, and on these ten months the typical miss is 4.6690, not 5.2202. The two are 0.5511 percentage pointsThe unit a change in a per cent figure gets counted in. Going from four per cent to six per cent gains two percentage points, and that is not the same claim as gaining two per cent. apart, on the identical set of misses, purely because of the denominator.

Four error measures, one set of ten misses, two things that differ three of the four rows start from the same 218, so any disagreement between them is a denominator THE MEASURE WHAT IS ADDED UP DIVIDED BY ROOT READING UNITS mean squared error the ten squared misses, 218 10 no 21.8000 per cent squared root mean squared error the ten squared misses, 218 10 yes 4.6690 per cent mean absolute error the ten miss sizes, 36 10 no 3.6000 per cent residual standard error the ten squared misses, 218 8 yes 5.2202 per cent 4.6690 and 5.2202 differ only in the 10 against the 8, on identical misses
The root mean squared error of 4.6690 per cent divides by ten and the residual standard error of 5.2202 per cent divides by eight, on the same ten misses.
Try it out

Root mean squared error comes back as 4.6690 per cent and residual standard error as 5.2202 per cent. Both are computed from the very same ten misses. Why do they differ?

What does the last row show that most fit reports leave out?

The lag one reading of the misses comes back as 0.4862 on the default months. The reading is computed by pairing each miss with the one before it, multiplying the two, adding the nine products to 106, and dividing by the 218 already in hand. Nothing exotic. But it is the only row of the seventeen that would give a different answer if the ten pairs were shuffled, and that is precisely why it earns its place.

Sixteen of the seventeen rows throw the running order away, and this one keeps it. Reordered any way at all, the ten months leave the slope at 1.5000, the correlation at 0.8694 and R squared at 0.7559, and every error measure holds to the last decimal. All of them are built from sums, and a sum does not care what order things are added in. Sorted by the outcome column, though, the same ten pairs take the last row down to 0.1422. Sorted by the input column instead, it reads minus 0.2982. Sorted by the misses themselves, smallest to largest, it climbs to 0.6193. Three different orderings, three different readings, one unchanged set of numbers.

On the record as it actually happened, the 0.4862 says that the misses left over from a fit reporting an R squared of 0.7559 are not independent of one another: a month landing above the line is unusually likely to have a month above the line sitting next to it. Whether that finding is fatal, mild or a symptom of something else is a proper argument with its own working, and it is covered separately. Most fit reports never print the number at all.

One set of misses, two orderings, and exactly one row that notices the same ten numbers appear in both panels, rearranged and not altered PANEL ONE: THE ORDER THE MONTHS HAPPENED IN 1.00 9.00 8.00 0.00 minus 5.00 0.00 minus 3.00 minus 3.00 minus 2.00 minus 5.00 m1 m2 m3 m4 m5 m6 m7 m8 m9 m10 the last row reads 0.4862 PANEL TWO: THE SAME TEN PAIRS SORTED BY THE OUTCOME COLUMN 0.00 minus 2.00 minus 5.00 minus 5.00 minus 3.00 8.00 1.00 minus 3.00 0.00 9.00 m6 m9 m5 m10 m8 m3 m1 m7 m4 m2 the last row reads 0.1422 slope, correlation, R squared and all four error measures are identical under both panels sorting changed one row out of seventeen, and nothing on the screen said so
Sorting the ten pairs leaves every returned row unchanged except the last, which falls from 0.4862 to 0.1422 because it is the only one that uses the order.
Try it out

The last row reads 0.4862. The row is built by pairing each miss with the miss before it. Which other returned row could have said the same thing?

What does the whole table read on the ten default months?

Here it is in full, printed as ordinary text so that it exists whether or not the tool is able to run. The ten pairs are the ones in the first figure, in the order they were recorded, and they are never sorted. Every row below was computed from the two columns.

RowReadingHow it was built
Handed straight back
Count of pairs10as typed
Average of the input1.00 per centthe ten inputs added, divided by ten
Average of the outcome2.00 per centthe ten outcomes added, divided by ten
The three sums
Input movement300squared distances of the input from 1.00, added
Movement together450the two distances multiplied each month, added
Outcome movement893squared distances of the outcome from 2.00, added
The line
Slope1.5000450 divided by 300
Intercept0.50002.00 less 1.50 times 1.00
The strength
Covariance50.0000450 divided by 9, in per cent squared
Correlation0.869450.0000 divided by 5.7735 and by 9.9610
R squared0.7559675 divided by 893
Adjusted R squared0.7254218 over 8, against 893 over 9
The misses
Mean squared error21.8000218 divided by 10, in per cent squared
Root mean squared error4.6690 per centthe square root of 21.8000
Mean absolute error3.6000 per cent36 divided by 10
Residual standard error5.2202 per centthe square root of 218 over 8
The one row that uses the order
Lag one reading of the misses0.4862106 divided by 218

And the ten misses themselves, month by month, in the order they happened: 1.00, 9.00, 8.00, 0.00, minus 5.00, 0.00, minus 3.00, minus 3.00, minus 2.00 and minus 5.00 per cent. Add them and nothing is left. The zero reads like a tick in a box and is not one. A fitted line always leaves misses that cancel, so a total of nothing confirms that the sums were done and says not one thing about whether the line is worth having.

Try it out

Before the tool below is touched: month 6 is the worst month in the outcome column at minus 13.00 per cent, and the line happens to hit it exactly. Raising that single figure by five points does what to R squared?

Play with it

Move one month, and watch which of the seventeen rows care

A month is chosen and its Vasant figure dragged, and everything redraws: the scatter, the fitted line, the ten misses in the order they happened, and every returned row. Only one month may be off its published value at a time, so whatever moves is attributable to the month that was moved. The input column never changes and the count stays at ten. The sort button at the end rearranges the ten pairs by outcome and leaves everything alone except the last row. The first worked setting is the one to start from: raising month 6 by five points takes R squared from 0.7559 down to 0.6967 and the slope from 1.5000 down to 1.3333. One month out of ten is worth that much when the line had that month exactly right.

Jump to a worked setting:
Slope
Intercept
Covariance
Correlation
R squared
Adjusted R squared
Mean squared error
Root mean squared
Mean absolute error
Residual standard error
Movement together
Last row, lag one
Educational illustration. The Nakshatra unit and the Vasant unit trade nowhere. The last row assumes the pairs are in the order they happened, and the sort button is there so that the cost of sorting can be watched. Only one month leaves its published value at a time. Nothing is stored, nothing is scored, and no returned row forecasts an eleventh month or proposes anything to anybody about any market.

What is the one thing this calculator cannot settle?

Whether the two columns have anything to do with each other. The calculator will fit a line to whatever is typed in, and it will fit that line just as willingly to two columns with no connection of any kind. The willingness is not a flaw to be patched but the arithmetic itself, and a tool that pretended otherwise would be lying. The misreading this calculator exists to prevent is treating a good fit as a claim that the input causes the outcome, or that the eleventh occasion can be read off the line. Both halves of that are drivable in the instrument at the top, and it is worth pressing the two buttons rather than taking the next four paragraphs on trust.

Press the chair count button and the input column becomes the Kadamba count: how many chairs were set out in a made up community hall that month, running 58, 79, 66, 72, 53, 45, 60, 59, 53 and 55. Nothing links chairs in a hall to a traded price. Fitted against the same Vasant column, the calculator returns a correlation of 0.9502, an R squared of 0.9029 and only 86.7349 left over, against the traded unit's 0.8694, 0.7559 and 218.0000.

The chairs win on every figure this calculator can produce, and there is no arithmetic anywhere in the seventeen rows that can say they should not have. Splitting the record in half is no rescue either. The pattern holds in both halves. The separation being sought is a question about mechanism, and a mechanism is not a number, so it will never appear in a returned table.

The swap button makes the same point from the other side, and it is the harder of the two to argue with. Press it and the Vasant column becomes the input while the Nakshatra column becomes the outcome. A slope is measured in units and the units have just been exchanged, so the slope moves: 450 over 300 becomes 450 over 893, and 1.5000 becomes 0.5039. The correlation stays at exactly 0.8694 and R squared stays at exactly 0.7559. Fit strength reads identically in both directions. Reading the same both ways is the plainest proof available that fit strength carries no direction at all. If the arithmetic cannot tell which column is acting on which, nothing it returns can be evidence that either one is.

The claim about next month fails on related ground. The line was built to sit as close as it can to occasions that have already been recorded, and none of the six steps looked forward. An eleventh month is in none of the sums, so the line has no hold on it, and the R squared of 0.7559 reports how the line did on the very record it was fitted to. The argument about prediction is worked through under patterns with no mechanism, covered separately.

The better fit is the one with nothing behind it the same Vasant column, fitted twice against two different inputs, both invented INPUT: THE NAKSHATRA UNIT Correlation 0.8694 R squared 0.7559 Squared misses left over 218.0000 a traded unit, invented for teaching INPUT: THE KADAMBA CHAIR COUNT Correlation 0.9502 R squared 0.9029 Squared misses left over 86.7349 chairs put out in an invented hall A MECHANISM CONNECTING THE TWO COLUMNS nothing the seventeen rows can see nothing the seventeen rows can see 90.29 per cent explained against 75.59, and the arithmetic prefers the chairs
The invented chair count returns a higher R squared than the traded unit does, and no returned row in the table can tell the two apart.

What is checked before trusting anything it prints?

Four things, and each one has a specific way of going wrong. First, confirm the pairs are in the order they actually happened. Get that wrong and sixteen of the seventeen rows are still perfectly correct while the last one is quietly meaningless, and nothing anywhere on the screen will flag it.

Second, both columns are confirmed to carry the units they are believed to carry. The covariance and the slope both move when units move: a column recorded as 0.065 rather than 6.50 changes the slope by a factor of a hundred and the covariance by the same. The correlation and R squared do not budge. Two people can therefore agree completely about the strength of a relationship and disagree wildly about its size, and neither is making an error of arithmetic.

Third, confirm nothing was sorted, dropped or patched after the first look at the data. A month deleted because it seemed odd is a decision about the answer disguised as a decision about the data, and the returned table cannot see the difference.

Fourth, look at the count. Ten pairs is a thin sampleThe readings that happen to be in hand, as against all the readings that could have been taken. Ten months is a sample, not a complete record of anything., and this calculator prints every figure to four decimal places regardless. The tool has no way of knowing how many pairs deserve that many. Four decimals on ten pairs is a statement about the arithmetic, never about how firmly the answer is known. The printing is exact; the knowing is not, and the tool cannot mark the difference.

Four gates between what goes in and what comes out none of the four is checked by the arithmetic, and none of them raises a warning on screen IS IT IN THE ORDER IT HAPPENED IN? ARE THE UNITS AS BELIEVED? WAS ANYTHING DROPPED LATER? HOW MANY PAIRS ARE THERE REALLY? what goes in 17 ROWS WHAT FAILS IT a sheet sorted for a tidier chart, and the last row now 0.1422 WHAT FAILS IT one column as 6.50, the other as 0.065, slope out by a hundred WHAT FAILS IT a month deleted for looking odd, after the first fit was seen WHAT FAILS IT ten pairs, and every figure still printed to four decimals
Check the order, the units, whether anything was dropped and the count of pairs, because the arithmetic checks none of the four.
Try it out

Sorting made the chart look better, so pairs were pasted in sorted order. Which single returned row is now wrong, and does the tool give a warning?

Try it out

The count of pairs is ten and the correlation is printed as 0.8694. What is misleading about those four decimal places?

Reading an Option Payoff — free micro-course from Fin Maverick

How is the whole table rebuilt by hand?

In six lines, and it is worth doing once by hand even though the instrument above will do it unaided. An analyst who has never done it reads a fit report as a black box that emits figures; an analyst who has done it once can see immediately three things: the figures that move when the units move, the figures that move when a month is added, and the figures that cannot move at all.

Line one: add each column and divide by ten, giving averages of 1.00 and 2.00 per cent. Line two: take each input away from 1.00, square each deviationThe distance from one reading to the centre of its own column, carrying a direction as well as a size., add the ten squares, giving 300. Line three: multiply each month's two distances together and add the ten products, giving 450. Line four: take each outcome away from 2.00, square, add, giving 893. Line five: 450 over 300 is the slope of 1.5000, and 2.00 less 1.50 times 1.00 is the intercept of 0.5000. Line six: take each outcome away from what the line said, giving the ten misses.

Everything else in this calculator is those six lines rearranged. At line three, what happens to the 450 if the outcome column is recorded in decimals rather than percentages? The 450 shrinks by a factor of a hundred, so the covariance shrinks and the slope shrinks with it. Then the correlation divides that shrunken 450 by a spread that shrank by exactly the same factor. Nothing happens to it at all. Written out by hand once, that stops being a fact to memorise.

Six lines of working, and the other fourteen rows are these rearranged nothing below needs anything beyond adding, subtracting, multiplying and dividing 1 Add each column, divide by ten. 1.00 and 2.00 per cent the two averages every distance below is measured from 2 Each input less 1.00, squared, added. 300 how much the input column moves about on its own 3 The two distances multiplied each month, added. 450 the one number three different readings are built on 4 Each outcome less 2.00, squared, added. 893 what the flattest possible guess would have missed by 5 450 over 300, then 2.00 less 1.50 times 1.00. 1.5000 and 0.5000 the slope, and the intercept that puts the line through the averages 6 Each outcome less what the line said. ten misses, adding to zero everything about the fit that is left comes out of these ten two averages, 300, 450 and 893: that is the whole supply of raw material here
Two averages plus the three sums of 300, 450 and 893 generate every other figure printed here.

The two ways this table gets misread

The loud one first. A reader runs the calculator, reports the residual standard error of 5.2202 per cent as the typical miss, and is then asked why a colleague working the same ten months got 4.6690. Both figures are right. Two numbers were consumed building the line, so one divides the 218 by ten and the other divides it by eight, and on ten pairs that choice moves the answer by 0.5511 percentage points. On a bigger record the two converge and nobody notices; on ten they do not, and half a percentage point is easily enough to change what somebody does next.

The quiet one does more damage. Sorting made the chart look tidier, so the same reader pastes their own pairs in sorted order and glances at the last row. On these ten months, sorting by the outcome column drops it from 0.4862 to 0.1422. A drop like that looks like reassurance. The lower reading is not a finding at all: it is a number produced by the sorting, and the sorting threw away the only thing the last row was reading. Sixteen rows are unaffected, so nothing looks broken, and no warning appears anywhere on the screen.

The fix is unglamorous and it works. Print the denominator beside every error measure, so 4.6690 and 5.2202 can never be mistaken for a disagreement about the data. And put the order assumption on the tool itself rather than in a note underneath it. The last row should announce what it depends on before anybody reads it. A calculator that states its own assumptions on screen is doing the one thing a calculator can do about a fault it cannot detect.

The meaning of any of the seventeen returned figures is covered separately, each on the subject that teaches it. The assumptions the fitting method rests on, what to do when a diagnostic reads badly, and how to read the last row properly are also covered separately. The calculator fits one input against one outcome and stops there; fits that carry two or more inputs are treated elsewhere, including the case where two of them say nearly the same thing. Choosing between fitted models, and adjusting one, belong to a different subject area and are covered separately.

One table, two typical misses, both right. See what the divisor decides elsewhere.

What is behind this calculator, and why is no authority named?

Fitting a straight line through paired readings is arithmetic that belongs to nobody: no regulator sets it, no exchange publishes it, and no standard writes it down for a market to follow. The list below is short for that reason, and it names what was actually used.

What was usedWhat it isWhere it sitsWhat it settles
The ten month recordTwo columns of made up monthly changes, kept in the order they were recordedHeld with these notes, not published anywhereEvery figure printed here
The check script beside itA run of assertions that recomputes each printed rowHeld with these notes, beside the recordThat the printed figures and the calculator agree with each other
Ordinary arithmetic of sums and ratiosAdding, squaring, multiplying and dividing, all older than any text that prints themNo single text can safely be credited with itThe method itself, which is why nothing is attributed

The Nakshatra unit, the Vasant unit and the Kadamba count are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.