Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quantitative Methods, Financial Data & Programming
1Probability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
2Statistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
3Correlation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
4Time Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
5Simulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
6Optimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
7Modelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
8Backtesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
9Data Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
10Programming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
11Quantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity

The Data Generating Process: The Mechanism Behind the Numbers

A data generating process is the mechanism a record came out of. The record is one thing that mechanism produced on one occasion, and it is not the mechanism. Every figure computed from a record is an estimate of something the mechanism has. Writing the two down as though they were one quantity is where most confident wrong answers begin.

Somebody has to write a mechanism down before its true mean can be printed at all, so neither object below was read off a record anybody maintains. Two objects do all the work here. One is a five value generator with weights adding to exactly 1.00, and the fifty month record drawn out of it. The other is a six year record of seventy two monthly changes, put together from three parts that were each written down separately. Both were built for teaching in earlier work, and both are rebuilt here from their written parts rather than carried across as totals. Every quantity is carried at full precision, rounded once on its way to the screen, half away from zero, with the direction written as a word. Rounding twice moves an answer, so a printed figure is never the input to the next calculation. A document that rounds twice looks exactly as confident as one that does not.

What is a data generating process, and what is it not?

Here is something easy to picture. A tea stall outside a bus depot has a real busy hour. Not an opinion about a busy hour, a fact about the stall, made of how the depot rosters its drivers and when the buses turn round and how far the next stall is. The busy hour exists on a Sunday when the stall is shut and nobody is counting anything. Now somebody stands there one Tuesday with a tally sheet. The tally that person carries home is a count, and the count is not the busy hour. The count came out of the busy hour, it is the best thing available, and on that particular Tuesday it may be some way off.

The stall and the tally sheet are the whole distinction, and everything below is a longer version of it. A mechanism produces records; a record is one of the things it produced. A mechanism has quantities of its own: a true mean, a true spread and a true shape. Each is a property of the mechanism in the same way the busy hour is a property of the stall. The three quantities sit there whether or not anybody has drawn a record, whether or not anybody ever will, and whether or not anybody has the faintest idea what they are.

A record has its own versions of the same quantities, and those are a different sort of object entirely. A sample mean, a sample spread and a sample shape are computed. The arithmetic that produced each one can be pointed at and, given the record, anybody else gets the same figure. The second set is computed and the first set is not, and no amount of care with the arithmetic on the right hand side turns it into the left hand side.

So how does anybody ever know the left hand side? Almost always, nobody does. A true mean can be printed in exactly one circumstance, and every figure above sits inside it: somebody wrote the mechanism down first. The generator behind the fifty month record was written down as five values with five weights, and its weights add to exactly 1.00, so its true mean works out at 1.00 per cent and its true spread at 5.00 per cent. Because nobody measured those two figures off anything, they carry no error and they are not estimates. A true mean and a true spread are quantities the generator has, not quantities anybody computed. The fifty month record that came out of that same generator reads a sample mean of 0.50 per cent, and it is not wrong.

TWO OBJECTS, AND EVERY FIGURE BELONGS TO ONE OF THEM The five value generator and its fifty month record, both invented. THE MECHANISM, WRITTEN DOWN Five values, five weights adding to exactly 1.00. Nothing here was measured off anything. true mean 1.00 per cent true spread 5.00 per cent true skewness 0.0000 true kurtosis 2.9200 No error attaches to any of these four. THE RECORD, COMPUTED Fifty months drawn out of the mechanism on the left. Every figure here is arithmetic. sample mean 0.50 per cent sample spread 4.9744 per cent sample skewness minus 0.0502 sample kurtosis 2.9998 Each one estimates the figure opposite it. produces, one record at a time never this way Every figure in the right hand box is an estimate of the figure sitting opposite it in the left hand box. Nothing in the left hand box could be printed at all if somebody had not written the mechanism down before any record came out of it.
A mechanism produces a record and a record does not produce a mechanism, so every figure on the record side is an estimate of the figure sitting opposite it on the mechanism side, and the mechanism side can only be printed because somebody wrote it down first.
Try it out

A true spread and a sample spread are two different quantities. Which of the two can actually be computed, and which one is normally never seen?

Breaking Into Quants Bootcamp — Fin Maverick

What have the hundred and fourteen readings before this one been doing?

Here is something worth noticing about the work sitting underneath this guide. Every stretch of it did the same two things in the same order. Each stretch wrote a mechanism down, then drew a record out of that mechanism and estimated from the record. A distribution, then a sample from it. A mechanism run forward, then the records it threw off. A relationship put in on purpose, then a fit that had to go and find it. Ten sets of notes, a hundred and fourteen readings, and the same two step motion under nearly all of it.

The repetition looks like a teaching convenience, and it is not. A miss cannot be measured until the truth is known, so writing the mechanism down first is the only arrangement in which an estimate can be watched missing. Take the truth away and the record still reads 0.50 per cent, still reads it correctly, and there is nothing anywhere to compare it against. The gap between 0.50 and 1.00 per cent is only visible from inside a room where somebody has written 1.00 per cent on the wall in advance.

Ground of that kind is not usually available to an explanation of research method. Nobody has to accept on trust that samples miss their mechanisms; it has been watched happening repeatedly, in different arithmetic, with the answer printed beside it every time. So the distinction can be stated as a definition here rather than argued for as a caution. Notes that connect can do that, and a good answer to a single question cannot.

THE SAME TWO STEPS, TEN TIMES OVER, A HUNDRED AND FOURTEEN READINGS Each panel is one set of notes. The strip along the bottom of each is the motion it repeated. a distribution 1 DEFINE A MECHANISM 2 ESTIMATE FROM A RECORD an estimate and its error 1 DEFINE A MECHANISM 2 ESTIMATE FROM A RECORD a fitted line and its misses 1 DEFINE A MECHANISM 2 ESTIMATE FROM A RECORD order in time 1 DEFINE A MECHANISM 2 ESTIMATE FROM A RECORD a mechanism run forward 1 DEFINE A MECHANISM 2 ESTIMATE FROM A RECORD a best answer found under limits 1 DEFINE A MECHANISM 2 ESTIMATE FROM A RECORD a fit that generalises, or fails to 1 DEFINE A MECHANISM 2 ESTIMATE FROM A RECORD a rule tested on history 1 DEFINE A MECHANISM 2 ESTIMATE FROM A RECORD a record and its faults 1 DEFINE A MECHANISM 2 ESTIMATE FROM A RECORD ordered cells that rerun 1 DEFINE A MECHANISM 2 ESTIMATE FROM A RECORD TAKE STEP ONE AWAY AND STEP TWO STILL RUNS, BUT NOBODY CAN EVER SAY WHETHER IT MISSED
Every set of notes so far defined a mechanism and then estimated from a record it produced, which is the only arrangement in which an estimate can be watched missing, because the truth has to be written down before a miss can be measured at all.
Try it out

What did every earlier set of notes do before it estimated anything at all?

What happens when a mechanism is assembled one part at a time?

Definitions only go so far. The way to feel what a mechanism is, as opposed to what a record is, is to build one and switch its parts on in turn. The six year record does exactly that. The six year record is seventy two monthly changes, and it was not written down as seventy two numbers. The record was written down as three parts that get added together.

The first part is a driftThe steady amount a series moves by on average each period, before anything else is added on top of it. It is the only part of this record that carries a mean. of 1.00 per cent. Every month, without exception, gets 1.00 per cent. The second is a seasonal partA pattern that repeats on a fixed calendar cycle, the same twelve figures over and over, arranged so that a full year of them adds to nothing.: twelve figures that repeat every year and add to exactly zero across a year. The third is an irregular partWhat is left of a series once the steady piece and the repeating piece have been taken out of it. Here it is written down separately and adds to nothing over the whole record., one figure per month, arranged so that all seventy two of them add to exactly zero.

Switching them on one at a time gives the mean and the spread at each setting. Four settings, four readings, and the answer is startling the first time it is laid out.

Which parts are switched onMeanSpreadVariance
the drift alone1.00 per cent0.00 per cent0.0000
the drift and the seasonal part1.00 per cent2.65 per cent7.0000
the drift and the irregular part1.00 per cent4.24 per cent18.0000
all three, which is the six year record1.00 per cent5.00 per cent25.0000

Read the mean column downwards. The mean does not move. Not approximately, not nearly: the mean is exactly 1.00 per cent at every one of the four settings, and the reason is written into the construction. Only the drift carries a mean, so nothing else that gets switched on has any mean to contribute. A part that adds to zero over a year and a part that adds to zero over the record are both, on average, nothing, and adding nothing to 1.00 per cent leaves 1.00 per cent.

Now read the spread column downwards. The spread climbs at every single step, from 0.00 to 2.65 to 4.24 to 5.00 per cent. Everything except the drift carries spread and nothing except the drift carries a mean, and once that is seen, the four rows stop looking like four different mechanisms and start looking like one mechanism with pieces bolted on. Ask a household version of it. A salary that arrives on the first of every month has an average and no wobble. A festival month that pulls spending forward and a quiet month that pushes it back have a great deal of wobble and, across a full year, no effect on the average at all. Both of those things are simultaneously true, and people find that surprising in exactly the way this table is surprising.

Try it out

A seasonal part that adds to exactly zero over a year is about to be switched on. Before the reading appears: does the mean move?

THE SAME SEVENTY TWO MONTHS, WITH THE PARTS SWITCHED ON IN TURN The six year record, invented. The dashed line sits at 1.00 per cent in every panel and never moves. 1. THE DRIFT ALONE mean 1.00 per cent, spread 0.00 per cent 1.00 2. THE DRIFT AND THE SEASONAL PART mean 1.00 per cent, spread 2.65 per cent 1.00 3. THE DRIFT AND THE IRREGULAR PART mean 1.00 per cent, spread 4.24 per cent 1.00 4. ALL THREE PARTS, WHICH IS THE RECORD mean 1.00 per cent, spread 5.00 per cent 1.00 months 1 to 72
Switching on the drift, then the seasonal part, then the irregular part leaves the mean at exactly 1.00 per cent every time, which is why the dashed line sits at the identical height in all four panels while the path around it grows steadily wilder.
Play with it

Switch the parts on yourself, and watch one line refuse to move.

One control, and it moves one thing: which parts of the mechanism are switched on, across the four settings from the drift alone to all three. Three things redraw. The seventy two month path changes shape, the spread bar underneath it grows on its own scale, and the variance bar underneath that fills up and finally splits into the two pieces that made it. The dashed mean line does not move a pixel at any setting, and watching it hold still is the whole exercise. With the box below the control ticked, the setting just left stays on screen as a pale line behind the current one. The gap between the two lines is precisely what the newly switched on part added. The control starts at all three parts, the six year record exactly as the table above prints it.

drift aloneall three partsall three
THE SIX YEAR RECORD AT THE SETTING CHOSEN Invented throughout. 72 months, drift 1.00 per cent, a season adding to zero over a year, an irregular adding to zero over the record. ALL THREE PARTS SWITCHED ON 17 1 minus 15 the mean, fixed at 1.00 per cent month 1 month 72 SPREAD 5.00 per cent 0 5.00 per cent, which is the record VARIANCE 7.0000 18.0000 25.0000 0 25.0000, which is the record The mean has not moved from 1.00 per cent.
Parts switched on
all three
Mean
1.00
Spread
5.00
Variance
25.0000
Lowest month
minus 13.00

The mean, the spread and the lowest month are monthly figures in per cent. The variance is the spread multiplied by itself, so it carries no unit of its own.

Educational illustration only. The true mean and the true spread can be printed here only because the mechanism was written down before any record came out of it. The seasonal part adds to exactly zero over a year and the irregular part adds to exactly zero over the record, both by construction rather than by luck. Every monthly value is a whole number, so the panel holds them as whole numbers and divides once at the end. The variances therefore read exactly 7.0000, 18.0000 and 25.0000 rather than nearly.

AI For Finance Bootcamp — Fin Maverick

Do the spreads of two parts add up to the spread of the whole?

The next step is where a reader loses a figure without noticing, so it is worth slowing down. Look again at the two middle rows of that table. The seasonal part on its own gives a varianceThe average squared distance from the middle of a set of numbers. It is simply the spread multiplied by itself, so a spread of 5.00 gives a variance of 25. of 7.0000. The irregular part on its own gives 18.0000. Put both on and the whole record gives 25.0000.

Seven plus eighteen is twenty five. Not approximately twenty five, not twenty five point something: exactly twenty five, and the exactness is doing a job. Two quantities only add like that when the two parts carry nothing whatsoever about each other, and an inexact sum would have been a measurement of how much they overlap. If the irregular part had leaned even slightly with the seasonal part, the two would have reinforced each other in some months and cancelled in others, the total would have come out above or below twenty five, and the shortfall would have measured how far they were connected. Here it comes out clean, so they are not connected at all, and that was written into the mechanism rather than discovered in the record.

SEVEN AND EIGHTEEN FIT INSIDE TWENTY FIVE WITH NOTHING LEFT OVER Area is variance. The two inner pieces are drawn to scale against the outer one. 7.0000 the seasonal part 18.0000 the irregular part the whole record, 25.0000 A gap inside the outer rectangle, or an overlap between the two inner ones, would have measured how far the two parts move together. There is neither, so on this record they carry nothing at all about each other.
The seasonal part's variance of 7.0000 and the irregular part's of 18.0000 fill the record's 25.0000 exactly, and the exactness is the evidence that the two parts carry nothing about each other rather than a coincidence of the arithmetic.

The same sum on the spreads is the move everybody makes, so try it now. The seasonal part's spread is 2.65 per cent, the irregular part's is 4.24 per cent, and 2.65 plus 4.24 is 6.89 per cent. The record's actual spread is 5.00 per cent. The wrong answer is not a rounding wobble or a near miss; it lands well past the true figure, and it is a correct answer to a question nobody asked. Adding the two spreads answers a different question altogether: what the spread would be if the two parts always moved in perfect step. Perfect step is precisely what the construction ruled out.

THE SPREAD CLIMBS TO 5.00, AND ADDING TWO OF THEM LANDS PAST THE END One axis of spread, in per cent. The solid part stops where the record stops. laying the second spread end to end after the first 0.00 2.65 4.24 5.00 6.89 the record the wrong sum WHAT EACH MARK ON THE AXIS IS 0.00, the drift alone. Every month reads the same figure, so there is nothing to spread. 2.65, the drift and the seasonal part, whose variance is 7.0000. 4.24, the drift and the irregular part, whose variance is 18.0000. 5.00, all three parts together, which is the six year record itself. 6.89, which is 2.65 plus 4.24. Nothing on this record has a spread of 6.89. Add the variances instead: 7.0000 plus 18.0000 is 25.0000, whose root is 5.00.
The spread reads 0.00, 2.65, 4.24 and 5.00 per cent across the four settings, and laying the two middle readings end to end gives 6.89 per cent, which sits well beyond the end of the axis rather than anywhere near the record's actual figure.

The fix is one line long and it is worth memorising in this form. Variances add, spreads do not, and the way back to a spread is to add the variances first and take the root once at the end. Seven plus eighteen is twenty five, and the root of twenty five is 5.00 per cent, the record's own spread. Notice also that the error is not fixed in size. Two parts overstate the spread by a little under two points here. Because every extra part gets laid end to end instead of being folded in, four parts of similar size would overstate it by far more. A mistake of this sort gets worse quietly as a model grows, and quiet growth is precisely what survives review.

Try it out

Two unrelated parts have spreads of 2.65 per cent and 4.24 per cent. A colleague adds them and reports 6.89 per cent for the whole thing. What is wrong, and what is the right answer?

Try it out

Two parts of a mechanism have variances that add to the variance of the whole exactly, rather than approximately. What does that exactness establish?

Risk Management Program Bootcamp — Fin Maverick

What can a record establish about the mechanism behind it, and what can it not?

Turn now to the other object. Written into the generator behind the fifty month record is a mean of 1.00 per cent, and that is its true one. Drawn out of it, the fifty months average 0.50 per cent, and that is a computed one. The record reads half the truth.

The first thing to say is the thing most people will not say out loud. The record is not wrong. The record is exactly what it is: fifty months came out of that mechanism and their average is 0.50 per cent, and anybody who recomputes it gets 0.50 per cent again. The mechanism is also exactly what it is. Nothing has malfunctioned, nobody has made an error, and no better arithmetic on those fifty months would move the answer any closer. The gap between the two is not a fault. The gap is the thing being taught, and an account that calls the record wrong has quietly turned the lesson back into an accusation.

So the record does not give the mechanism's mean. The record gives something else instead, genuinely useful and easy to walk past. The record carries an honest statement of how far off it might be, and that statement is computed from the record alone, without any knowledge of the truth. Here the standard errorA figure computed from a record that says how much a summary like an average would jump about if a fresh record of the same length were drawn from the same place. works out at 0.7035 per cent. Nobody had to know that the truth was 1.00 per cent to get that figure; it falls out of the fifty months and their own spread of 4.9744 per cent.

Set beside each other, the two say this much. The record says 0.50 per cent, and it says in the same breath that a figure of this kind, from fifty months of this material, wobbles by something like 0.7035 per cent. A truth of 1.00 per cent is well inside the reach of that wobble. The record cannot give the truth, but it can say how far from the truth it might be, and that second statement is the honest half of every estimate anybody has ever published. The tally sheet from the bus depot cannot give the stall's real busy hour. The tally sheet can say that one Tuesday is a thin thing to rest a conclusion on, and it can say roughly how thin.

WHAT THE RECORD READS, AND WHAT IT SAYS ABOUT ITS OWN RELIABILITY The fifty month record and the generator behind it, both invented. Monthly figures, in per cent. one standard error of 0.7035 per cent either side of the reading minus 1.00 0.00 2.00 0.50 per cent what the fifty months read computed 1.00 per cent the mechanism's true mean written down, never computed The bracket was worked out from the fifty months and their own spread of 4.9744 per cent. Nothing about the truth went into it.
The record reads 0.50 per cent against a true 1.00 per cent, and its own standard error of 0.7035 per cent was computed from the fifty months alone, so the record cannot give the truth but it can say how far from the truth it might be.
Try it out

A record reads 0.50 per cent and the true mean of the mechanism behind it is 1.00 per cent. Is the record wrong?

Does the shape a record reports belong to the mechanism?

Means and spreads are the easy half. Shape is where the confusion does real damage. A shape figure looks so much like a description of the mechanism behind the record, when it is only a description of the record.

Take skewnessA measure of how lopsided a set of numbers is, reading zero when the two sides of the middle match each other exactly. first. The generator behind the fifty month record is symmetric by construction: five values placed evenly either side of 1.00 per cent, with the two weights on the left matching the two on the right. Its own skewness is exactly zero, and exactly is the right word: the zero comes out of the arrangement rather than out of a calculation that nearly cancelled.

Try it out

The generator behind record one is exactly symmetric. Before the figure appears: what skewness will a fifty month record drawn out of it report?

The fifty month record reports a skewness of minus 0.0502. Small, certainly, but not zero, and it will never be zero. Fifty draws from a symmetric mechanism land where they land, a few more happen to fall on one side than the other, and the arithmetic duly reports a lopsided record. A slightly lopsided sample from a perfectly symmetric mechanism is the ordinary case rather than a finding, and writing it up as a lopsided series describes the fifty months correctly and the mechanism wrongly.

KurtosisA measure of how much of a set of numbers sits far out at the edges rather than bunched near the middle. Its most quoted reference point is three. is worse. The trap there has a number attached to it that people recognise. The fifty month record reports 2.9998. A kurtosis of 2.9998 sits within a thousandth of three, and three is the reading everyone associates with the familiar bell shape. Very few readers can look at 2.9998 and not conclude that the shape behind the record has just been identified.

The reading settles nothing. The generator that produced that record is a list of five values with five weights, and its own kurtosis is 2.9200. The generator is not a bell shape, and it is not a smooth shape of any description. Five outcomes are all it has, and it cannot produce a sixth. A record drawn from that list happened to report a figure a bell shape would also have reported. A reading near three is what a great many quite different mechanisms produce, so it rules almost nothing out, and treating it as identification is how a shape gets assumed rather than established.

TWO SHAPE FIGURES, SIDE BY SIDE, EACH WITH ITS RECORD LENGTH BESIDE IT WHAT IS BEING MEASURED THE MECHANISM'S OWN READ OFF FIFTY MONTHS WHAT THAT ACTUALLY MEANS SKEWNESS how lopsided it is 0.0000 exact, by construction five values, evenly placed minus 0.0502 a reading from 50 months A slightly lopsided record from a symmetric mechanism is the ordinary case. An exact zero would have been the surprise. KURTOSIS how much sits far out 2.9200 a five value list, and not a bell shape at all 2.9998 a reading from 50 months within a thousandth of three The reading everyone associates with a bell shape, produced here by something that is not one. It rules almost nothing out. BOTH READINGS ARE CORRECT ARITHMETIC ON FIFTY MONTHS, AND NEITHER DESCRIBES THE MECHANISM
Record one reports a skewness of minus 0.0502 from a generator whose own skewness is exactly zero, and a kurtosis of 2.9998 from a generator whose own kurtosis is 2.9200, so both readings are right about the fifty months and wrong about the mechanism.
THE MECHANISM THAT PRODUCED A KURTOSIS OF 2.9998 Five outcomes, five weights adding to exactly 1.00. Invented. Monthly figures, in per cent. 0.08 0.18 0.48 0.18 0.08 minus 9 minus 4 1 6 11 0.3989 the smooth bell shape with the same mean and the same spread, read at these same five points. Its five readings add to 0.9909, not to one, because it does not live here. This mechanism has five possible outcomes and cannot produce a sixth. Its own kurtosis is 2.9200. A fifty month record drawn out of it reported 2.9998, which is the bell shape's own reading to three decimals.
A kurtosis of 2.9998 sits within a thousandth of the bell shape's own value and came out of a list of five outcomes that is nothing like a bell shape, so a reading near three identifies almost nothing about the mechanism that produced it.
Try it out

A record reports a kurtosis of 2.9998. A colleague concludes that the mechanism behind it is the familiar bell shape. What is the reply?

What goes wrong when a record's shape gets reported as the mechanism's shape?

Picture the write up. Somebody has the fifty month record in front of them and a section to fill headed distribution of the series. The person computes the two shape figures, both correctly, and types two sentences. The series is slightly negatively lopsided, at minus 0.0502. The series is well behaved, with a kurtosis of 2.9998, close to the familiar bell shape.

Both readings are correct and both descriptions are wrong, and there is no arithmetic error anywhere to find. The lopsidedness is fifty months of draws and nothing else, because the mechanism was built exactly symmetric and its own skewness is zero. The bell shape is not there either: the mechanism is a five value list whose own kurtosis is 2.9200, with five possible outcomes and no sixth.

The expensive part comes next, and it is not the paragraph. A shape assumed from a sample decides which method gets used afterwards. Somebody reads well behaved and picks the tool that suits a bell shape. Somebody reads slightly lopsided and adds a correction for lopsidedness that the mechanism does not have. One misread record quietly chooses the method for everything downstream of it. By then nobody is looking back at the two sentences that started it, and the two sentences contained no error to find.

The fix costs one clause per figure. State every shape figure as a reading from this record, with the length of the record beside it, and never as a property of the mechanism unless the mechanism was written down. A kurtosis of 2.9998 read off fifty months is an honest sentence. The series is well behaved is not, and the difference between them is eight words.

Debt Capital Markets Bootcamp — Fin Maverick Cleaning Financial Data — free micro-course from Fin Maverick

What should be asked about any record that arrives from somebody else?

All of this turns into five questions, and they are worth having by heart because they take about ninety seconds and they are the difference between reading a record and being read to by one. A lender looking at a borrower's twelve months of takings, an analyst handed a spreadsheet of monthly figures, an investor reading a track record, and a household deciding whether last winter's electricity bill says anything about this one are all doing the same job: working out which of the numbers in front of them are properties of the thing and which are properties of the sample.

Five things to establish about a record before concluding anything from it

One. What mechanism is this supposed to have come out of, in one sentence. If nobody can say it in a sentence, every figure below is an estimate of something nobody has described.

Two. How many observations are in it, and were any left out. Fifty months and five months support very different sentences, and a record that quietly dropped its worst stretch is a different mechanism from the one it claims.

Three. Which figures here are properties of the mechanism and which are estimates of them. In this guide that line is easy to draw because the mechanism was written down. On a real desk everything is usually an estimate, and that is itself the answer.

Four. What is the error on each estimate. A figure with no error beside it is half a statement. The fifty month record gives 0.50 per cent, and 0.7035 per cent is the other half of what it said.

Five. If this record had come out differently, which of the conclusions in front of me would change. Ask this one last and ask it slowly. A conclusion that survives every possible record was never about the record at all. Sometimes that is fine and the conclusion rests on the mechanism. Sometimes it means somebody decided the answer before the record arrived and the record is decoration.

THE COVER SHEET THAT GOES ON TOP OF A RECORD Five questions, about ninety seconds, answered before any conclusion is written. BEFORE I CONCLUDE ANYTHING FROM THIS RECORD 1 What mechanism is this supposed to have come out of, in one sentence? 2 How many observations are in it, and were any left out? 3 Which figures here belong to the mechanism, and which estimate them? 4 What is the error on each estimate? 5 If this record had come out differently, which of the conclusions in front of me would change? Ask this one slowly. A CONCLUSION THAT SURVIVES EVERY POSSIBLE RECORD WAS NEVER ABOUT THE RECORD
The last of the five questions asks which conclusions would change if the record had come out differently, and a conclusion that survives every possible record was never about the record at all.
Try it out

A record arrives together with a set of conclusions drawn from it. Which single question separates a conclusion about the mechanism from a conclusion about this particular record?

One distinction is drawn above and nothing more is attempted. Covered earlier: how a question and the reading that would refute it are fixed before the record is opened. Covered later: whether a second person can reach the same number from the same data, whether a second record reaches the same finding, structured challenge, writing a claim that can fail, and whether a result survives a changed assumption. Running a mechanism forward to produce many records, and diagnosing a fault inside a record, are both covered separately.
Cleaning Financial Data teaches you to find the errors that survive every check and break every model.

What was every figure above built from?

A mechanism that somebody wrote down has no keeper, so it has no source to cite and no date of reading. Two written objects produced every figure above, and both were assembled from their own parts rather than carried across as totals. The table names them.

What was usedWhat it holdsSite
The five value generator and its fifty month recordFive values with five weights adding to exactly 1.00, giving a true mean of 1.00 per cent and a true spread of 5.00 per cent, and one record of fifty months drawn out of itNone. Written down for teaching and published nowhere
The six year record and its three named partsSeventy two monthly changes, written as a drift of 1.00 per cent, twelve seasonal figures repeating each year and adding to zero, and seventy two irregular figures adding to zeroNone. Same standing as the object above
The checking script beside these notesRebuilds both objects from those parts and asserts every mean, spread, variance, skewness and kurtosis printed above, along with the geometry of each drawingNone. Kept beside these notes and published nowhere
Nothing else whatsoeverNo regulator, exchange, market, maintained series, company, published work, journal or person was consulted, and none was neededNone

The five value generator, the fifty month record drawn out of it, and the six year record of seventy two monthly changes with its drift, seasonal part and irregular part are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.