Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quant Analyst · CoreTrack
1Quantitative Methods, Financial Data & Programming
iProbability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
iiStatistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
iiiCorrelation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
ivTime Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
vSimulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
viOptimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
viiModelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
viiiBacktesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
ixData Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
xProgramming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
xiQuantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity
2Stochastic Calculus & Derivative Pricing Theory
iProbability Foundations
The Probability SpaceRandom VectorsSigma-AlgebraExpectationSample Space and EventsDensity and Distribution FunctionsRisk-Neutral ProbabilityState Price Density vs…
iiStochastic Processes and Jumps
Properties of a Stochastic ProcessMartingaleBrownian Motion and Its PropertiesBrownian Motion vs Geometric…Stopping TimeThe Markov PropertyState VariablesTransition ProbabilityQuadratic VariationQuadratic Variation vs Ordinary…Submartingale and SupermartingaleMartingale RepresentationMarkov Process vs MartingaleOptional StoppingFiltrationJump ProcessesThe Poisson ProcessLevy ProcessesJump Diffusion
iiiIto Calculus
The Ito IntegralThe Ito Integral vs the Riemann IntegralInfinitesimals in Stochastic CalculusQuadratic CovariationIto's LemmaHow to Apply Ito's…The Infinitesimal GeneratorIto Calculus vs Ordinary Calculus
ivStochastic Differential Equations
Stochastic Differential EquationsStochastic Differential Equation vs…Drift and DiffusionStrong and Weak Solutions ComparedDiscretisationGeometric Brownian Motion
vPricing Theory and No-Arbitrage
No-ArbitrageGirsanov, Radon-Nikodym and Change…Physical and Risk-Neutral Measures…The Fundamental Theorems of…The Law of One PriceThe Pricing KernelDiscount Factors and Zero-Coupon PricesReplication vs HedgingComplete Market vs Incomplete MarketClearing Margin Architecture
viOption Pricing Theory
European and American OptionsMonte Carlo European OptionThe Black-Scholes PDEBlack Scholes and the GreeksThe Payoff FunctionThe Binomial ModelBinomial Option PricingDelta Hedging in TheoryBoundary, Initial and Terminal ConditionsThe Exercise BoundaryHow to Check Put-Call…
viiVolatility Models
Constant, Local and Stochastic…Vasicek Model vs CIR ModelThe Heston ModelThe SABR ModelThe Volatility ProcessImplied VolatilityVolatility Smile vs Skew vs Surface
viiiInterest Rate Models
Interest-Rate DerivativesMean ReversionThe Zero-Coupon BondThe Ornstein-Uhlenbeck ProcessThe Discount CurveZero RatesShort-Rate Model vs Market Model
ixNumerical Pricing
Closed Form and Numerical…Monte Carlo PricingEuler and Milstein Schemes ComparedTree MethodsFinite Difference MethodsNumerical Error and StabilityVariance Reduction
xCalibration and Model Risk
Model OverrideMarket Price and Model PriceCalibrationHow to Document a Pricing ModelThe Educational Illustration LabelMarket ConventionsModel Uncertainty and LimitationsBacktesting a Pricing ModelIdentifiabilityCalibrated ParametersThe Calibration Loss Function

How to Summarise a Financial Dataset With Descriptive Statistics

A complete summary has two halves. The descriptive half says what the record holds: the count, the centres, the spread, the extremes and the shape. The uncertainty half says how much of it is settled: the standard error and two intervals. For the fifty month record of the invented Nakshatra unit that comes to twelve lines, and the three in the uncertainty half are the ones usually missing.

Start with the word in the title. A dataset is a record of cases. Somebody decided what one case would be, went and collected a pile of them, and wrote them down in one place. The cases here are months. The invented Nakshatra unit is a traded unitA nameless something that carries a price, and whose price is allowed to wander. None of it has to be decided before the counting can begin, so what it is, who issues it and what it does are all left unsettled on purpose., and its monthly changeHow far the price travelled across a single month, measured against the level it opened at and written as a percentage. A reading of plus 11.00 marks a month that closed a good way up, and minus 9.00 marks one that closed well down. was written down for fifty months running. Fifty cases, one figure each: that is the whole dataset. Twenty five of the fifty months landed on 1.00 per cent, eight on 6.00 and three on 11.00 above that, nine on minus 4.00 and five on minus 9.00 below. Five counts, adding to fifty.

Every figure in the summary below was built somewhere in the notes that came before. The two means, the median and the mode were built in the notes on centres. The standard deviation and the range were built in the notes on spread. The skewness and the kurtosis were built in the notes on shape. The standard error was built in the notes on the distribution of an estimate, and both intervals were built in the notes on intervals. The list itself has never been assembled until now: which of those figures belong in a summary, what order they go in, and which of them a reader is entitled to see whether or not anybody asked for it.

One thing about this record deserves a sentence rather than a footnote, and it allows a verdict at the end that almost no explanation of the subject can offer. The generatorThe recipe the months came out of, naming each reading it is able to produce and how heavily it favours that one. Because the recipe was on paper ahead of the record, its own centre is something these notes can look up instead of estimate. behind the fifty months was written down first. Its true average is 1.00 per cent and its true spread is 5.00 per cent, both facts of construction rather than measurements. So at the end of this guide the finished summary can be laid beside the answer it was reaching for, line by line, and marked.

What is a summary for, and who is it for?

A summary is not a shorter version of the record. A summary is a substitute for the record, written for somebody who will never see it. The reader of a summary reads the twelve lines, forms a view, and acts on that view without ever counting a month. Everything they will ever know about those fifty months is what the author chose to write down.

The substitution gives a test, and it is a hard one. If a reader would reach a different conclusion from the summary than from the record itself, the summary is incomplete, no matter how many figures it carries. Notice what that test does not say. Correctness is assumed rather than tested. The test says correctness is not enough. A summary can be right on every line it prints and still send its reader somewhere the record would not have sent them, purely through what it left off.

Here is the everyday version. A shopkeeper is asked how business went last year and answers that takings averaged a certain amount a day. Every word true. Suppose a road outside was dug up for six months. Half the year then ran at double that figure and the other half at nothing, and the person who heard only the average now believes something flatly wrong about that shop. Nobody gave them a false number. A true one arriving alone did the misleading.

So the reader is the specification. The reader's own next move settles the list of lines. How much record was there? A reader who does not know that cannot know how seriously to take anything below it. What was a typical month, and how far from typical did the months actually run? How bad was the worst one? And how much of all this would a second record of the same thing have reproduced? Six questions. Twelve lines answer them.

Try it out

By what test is a summary judged complete?

What order are the twelve lines built in?

Count first, then the centres, then the spread, then the extremes, then the shape, then the uncertainty. Six stages, twelve lines. The order is not a filing convention and it is not alphabetical. Each stage exists to make the stage before it readable, so building them out of order produces a summary whose lines cannot be interpreted in the order a reader meets them.

Take the transitions one at a time. The count comes first because it decides what everything below it is worth. A mean drawn from four months and a mean drawn from four hundred are printed identically and are not remotely the same object, and a reader who meets the mean before the count has already formed an impression by the time the count arrives to correct it.

A centre on its own is half an answer, so the centres come next and the spread must follow immediately. Two records can share an average exactly and be nothing alike, one of them barely moving and the other swinging violently around the same middle. The spread is a single summary figure and the extremes are the check on it, so the extremes come after. They say whether that one figure is describing the record fairly or whether one freak month is doing most of the work.

The shape comes next because it answers the question the extremes raise. The lowest and the highest month give how far the record reached in each direction; the shape gives whether the misses pile up on one side and how much weight sits far from the middle. And the uncertainty half comes last because it is a comment on everything above it. The uncertainty half does not describe the record at all. Its three lines say how much of the nine above would have come back the same if somebody had collected a second fifty months instead.

SIX STAGES, TWELVE LINES. THE REASON LIVES ON THE ARROW. Each stage is built so that the stage above it can be read. 1. THE COUNT 1 line: how many cases there are How long the record is decides what every line below it is worth. 2. THE CENTRES 4 lines: two means, median, mode A centre on its own is half an answer, so the spread must follow. 3. THE SPREAD 1 line: the standard deviation The extremes say whether those two describe the record fairly. 4. THE EXTREMES 1 line: the range, lowest to highest The shape says which side the misses fall on and how far out. 5. THE SHAPE 2 lines: skewness and kurtosis The last stage says how much of all of it a second record would repeat. 6. THE UNCERTAINTY 3 lines: standard error, two intervals THE STAGE THAT GETS DROPPED, BECAUSE IT COMES LAST.
Build a summary in order: the count decides what anything after it is worth, and the centres are meaningless until the spread sits beside them.

Notice the small cruelty at the bottom of that ladder. The stage a reader most needs in order to weigh everything else is the stage that arrives last, when the work already looks finished and the person building it is already tired of building it. Order explains the omission. Nothing excuses it.

Try it out

Why does the count come first rather than the average?

What goes in the descriptive half?

Nine lines, and they are simply the output of the first five stages. A summary has rows, so the nine are printed as a list rather than described in prose.

StageLineThe fifty month record
CountNumber of cases50 months
CentresArithmetic mean0.50 per cent
CentresGeometric mean0.38 per cent
CentresMedian1.00 per cent
CentresMode1.00 per cent
SpreadStandard deviation, corrected denominator4.97 per cent
ExtremesRange, minus 9.00 to 11.0020.00 per cent
ShapeSkewnessminus 0.05
ShapeKurtosis3.00

Two of those rows need a word about why they are rows at all rather than a choice between two rows. The first is the spread, printed with its denominator named. There are two ways of dividing when a spread is worked out, they give 4.97 and 4.92 on this record, and which one was used is not something a reader can reverse engineer from the answer. Naming the corrected denominatorOne of the two divisors available when a spread is worked out from a sample rather than from a whole population. Which one to use, and why they differ, is covered separately; a summary only has to say which was used. costs three words and removes an ambiguity that cannot be removed any other way.

The second is the pair of means, and this is the row people cut. A reader who is going to add the figure needs a different line from a reader who is going to compound it, so both means belong in a summary rather than whichever one the author prefers. The arithmetic mean of 0.50 per cent is the one that answers questions about adding and averaging. The geometric mean of 0.38 per cent is the one that answers questions about compoundingGrowth applied on top of growth, so each period's change is worked out on the total the period before left behind rather than on the original amount. Why that needs a different average is covered separately., where each month's change lands on whatever the month before left behind. Adding and compounding are different questions with different answers, and the person writing the summary does not know which one the reader will bring. With both printed, the reader picks. With one printed, the author has picked for them, silently.

Try it out

Why do both means belong in the summary instead of whichever one the author judges better?

What goes in the uncertainty half, and why is it the half that gets dropped?

Three lines. The standard error is 0.70 per cent. The interval for the mean runs from minus 0.88 to 1.88 per cent. The interval for a single month runs from minus 9.35 to 10.35 per cent. All three were built earlier and none of them is rebuilt here.

Together the three lines do what the nine above cannot do for themselves. Without these three lines a reader cannot tell whether the mean of 0.50 per cent is a finding or a shrug, and the entire descriptive half reads as though it had been measured rather than estimated. That is the whole difference. A measured figure is what it is. An estimated figure is one answer out of many that the same procedure could have produced, and the three lines below the rule are the only place in the summary where that is admitted.

Read them in order. Line ten, at 0.70 per cent, gauges the typical distance between this record's average and the average a second batch of fifty months would have thrown up. Line eleven, running from minus 0.88 to 1.88, turns that into a stated range for the true average. The interval for a single month, from minus 9.35 to 10.35, answers a different question again: not where the average sits, but how wild one month can be. Note how far apart those last two ranges are in width. One is under three percentage points wide and the other is nearly twenty. Both are correct and they answer different questions, and which of the two a reader needs is covered separately. Both are computed on the assumption of the normal shapeThe familiar single humped, symmetric curve used here only to supply the multiplier that turns a spread into an interval. The shape itself is covered separately. for the multiplier, which is itself a convention the summary has to declare.

THE COMPLETE SUMMARY: TWELVE LINES, NINE ABOVE THE RULE AND THREE BELOW Every figure below was built earlier and is used here rather than rebuilt. SUMMARY OF THE FIFTY MONTH RECORD Monthly change of the Nakshatra unit, invented for teaching 1 Count 50 months 2 Arithmetic mean 0.50 per cent 3 Geometric mean 0.38 per cent 4 Median 1.00 per cent 5 Mode 1.00 per cent 6 Standard deviation, corrected denominator 4.97 per cent 7 Range, from minus 9.00 to 11.00 per cent 20.00 per cent 8 Skewness minus 0.05 9 Kurtosis 3.00 THE UNCERTAINTY HALF. THESE ARE THE THREE LINES USUALLY MISSING. 10 Standard error of the mean 0.70 per cent 11 Interval for the mean, 95 per cent minus 0.88 to 1.88 per cent 12 Interval for a single month, 95 per cent minus 9.35 to 10.35 per cent
A complete summary of the fifty month record runs to twelve lines, and the three below the rule are the ones usually missing.

Why do they go missing? Not because anybody decided to hide them. Three reasons: nobody asked for them, they arrive last, and above all a summary with its uncertainty half removed looks exactly as finished as one that still has it. There is no ragged edge, no obviously empty box, nothing a reviewer's eye catches. Nine tidy lines of correct arithmetic present themselves as complete work, and they are not.

BOTH OF THESE LOOK FINISHED. ONE OF THEM IS. There is no ragged edge to catch a reviewer's eye. COMPLETE, TWELVE LINES Count50 months Arithmetic mean0.50 per cent Geometric mean0.38 per cent Median1.00 per cent Mode1.00 per cent Standard deviation4.97 per cent Range20.00 per cent Skewnessminus 0.05 Kurtosis3.00 Standard error0.70 per cent Interval for the meanminus 0.88 to 1.88 Interval for a single monthminus 9.35 to 10.35 AS CIRCULATED, NINE LINES Count50 months Arithmetic mean0.50 per cent Geometric mean0.38 per cent Median1.00 per cent Mode1.00 per cent Standard deviation4.97 per cent Range20.00 per cent Skewnessminus 0.05 Kurtosis3.00 Nothing here looks unfinished. No blank row, no gap, no ragged edge. A third of the summary is simply absent. Every figure on both cards is correct. The difference between them is not accuracy. It is completeness.
A summary with its uncertainty half removed looks exactly as complete as one without, which is why the omission is so rarely noticed.
Try it out

Which three of the twelve lines are the ones usually left out, and what does their absence hide?

Try it out

The summary reports a mean of 0.50 per cent, with a spread of 4.97 sitting beside it. Which line tells a reader how much to trust that 0.50?

Breaking Into Quants Bootcamp — Fin Maverick

Which lines move when the record gets longer, and which do not?

One reading makes the two halves worth separating at all, and it is best met as a question before an answer. Suppose the same unit had been recorded for four hundred months instead of fifty, and suppose the longer record came out with the same mean of 0.50 per cent and the same spread of 4.97. Which of the twelve lines would read differently?

The nine descriptive lines would not move at all. The arithmetic mean is still 0.50 per cent. The geometric mean is still 0.38. The median and the mode are still 1.00. The standard deviation is still 4.97 and the range still 20.00. The skewness is still minus 0.05 and the kurtosis still 3.00. The descriptive lines report what is in the record, and by the terms of the question nothing in it has changed.

Every line in the uncertainty half moves. The standard error falls from 0.70 per cent to about 0.25. The interval for the mean tightens from a span of nearly three percentage pointsThe unit used when subtracting one percentage from another. Going from 4 per cent to 6 per cent is a rise of two percentage points, not a rise of two per cent, and keeping the two apart avoids an ambiguity that has cost people money. to a span of about one. The interval for a single month barely moves at all, from a width of 19.69 to 19.52, and that near stillness is itself informative: how wild one month can be is a fact about the unit, not about how long anybody watched it.

The two halves answer different questions, and only one of them is about how much record there is. That single reading is what a summary's split is for. The descriptive half is a report on a pile of months and it says nothing whatever about how confident anybody should be. The uncertainty half is a report on the reporting, and it is the only part of the summary that responds to effort.

THE SAME TWELVE LINES AT TWO RECORD LENGTHS The longer record is imagined to carry the same mean and the same spread. LINE AT 50 MONTHS AT 400 MONTHS Count50400the dial Arithmetic mean0.50 per cent0.50 per centfrozen Geometric mean0.38 per cent0.38 per centfrozen Median1.00 per cent1.00 per centfrozen Mode1.00 per cent1.00 per centfrozen Standard deviation4.97 per cent4.97 per centfrozen Range20.00 per cent20.00 per centfrozen Skewnessminus 0.05minus 0.05frozen Kurtosis3.003.00frozen Standard error0.70 per cent0.25 per centmoved Interval for the meanminus 0.88 to 1.880.01 to 0.99moved Interval for a single monthminus 9.35 to 10.35minus 9.26 to 10.26moved Illustration on invented figures. The longer record is a supposition, not a prediction of what more months would show.
Lengthening the record leaves every descriptive line exactly where it was and changes every line in the uncertainty half.
Try it out

Worth predicting before the panel below is touched: as the record is imagined to grow from fifty months to four hundred, how many of the twelve lines read differently?

Play with it

Turn the length of the record and watch nine of the twelve lines refuse to move.

One control moves: how many months the record is imagined to hold, from 1 to 400. The record's own mean of 0.50 per cent and spread of 4.97 per cent are held fixed by the construction of this panel, and are shown on screen at every setting, so the nine descriptive lines stay frozen throughout. The second control hides the uncertainty half. The card becomes the version that actually gets circulated, and the hidden lines are what a reader can no longer see. At the opening setting of 50 months the three lower lines read 0.70 per cent, minus 0.88 to 1.88 per cent and minus 9.35 to 10.35 per cent, which are the published figures for the fifty month record exactly.

Jump to a setting:
Months
50
Standard error
0.70 per cent
Interval for the mean
minus 0.88 to 1.88
Lines moved from the published card
0 of 12

Educational illustration on invented figures. The nine descriptive lines are held fixed so that only length is doing any work. A second record of four hundred real months would shift them as well, and holding them still is what separates the two effects. Both intervals use the multiplier from the normal shape. The true values used for scoring further down are known only because the generator was typed out before any month was counted.

Two settings are worth visiting deliberately. Push the slider to 1 and the panel refuses: a single month has no spread of its own to divide, so there is no uncertainty half to print, and refusing is the honest output. Push it to 400 and watch the interval for the mean shrink to a span of about one percentage point while the nine lines above it sit exactly where they were at the opening setting. Nothing about the record changed. Only the confidence in the reporting of it did.

What must a summary say about itself?

Three more lines, and none of them is a figure. The three describe the summary rather than the record, and together they make the twelve lines above them auditable by somebody who was not there.

The first says what the cases are and over what stretch. Here: monthly changes of an invented traded unit, fifty months. Without it, the reader cannot tell whether a mean of 0.50 per cent is a month, a quarter or a year, and that ambiguity alone can be worth an order of magnitude.

The second says what was excluded and how much of it there was. Here: nothing was excluded, all fifty months are in. A summary with no exclusion count cannot even be checked by its own author six months later, so the exclusion line is what turns a set of numbers into something anybody can check. Nothing was excluded is a real answer and a strong one. Silence is the fatal answer. On paper it looks identical to nothing was excluded, and the two mean completely different things.

The third says which conventions were used. Here: the spread uses the corrected denominator, and both intervals assume the normal shape for the multiplier. A reader cannot recover either choice from the answers, and stating both costs a line.

WHAT THE SUMMARY MUST SAY ABOUT ITSELF: THREE LINES, NO FIGURES These sit beneath the twelve, and they are what makes the twelve checkable. A. WHAT THE CASES ARE, AND OVER WHAT STRETCH Monthly changes of the Nakshatra unit, invented, over fifty consecutive months. B. WHAT WAS EXCLUDED, AND HOW MUCH OF IT Nothing was excluded. All fifty months are in. This is the auditable line. C. WHICH CONVENTIONS WERE USED Corrected denominator for the spread. Normal shape for both interval multipliers. Silence and the sentence nothing was excluded look identical on paper and mean entirely different things.
A summary must say what the cases are, what was excluded and how much of it, and which conventions were used.

How did this summary score against the truth?

Here is where these notes can do something they have been able to do all the way through and almost nothing else can. The generator behind the fifty month record was typed out before a single month was counted, so its true average of 1.00 per cent and true spread of 5.00 per cent are facts rather than estimates. The completed summary can therefore be marked.

Line by line. The summary reports an arithmetic mean of 0.50 per cent against a true 1.00, so the centre is out by half. The spread of 4.97 against a true 5.00 is very nearly right. The skewness of minus 0.05 against a true 0.00 detects a lean in a population that does not lean at all. And the kurtosis is reported as 3.00, the value the normal shape carries, for a population that takes exactly five values with a hard wall at minus 9.00 and another at 11.00. A population with hard walls has no tails to be heavy or light about.

MARKING A CORRECT SUMMARY AGAINST A KNOWN TRUTH Lime is what the record reported. Pine is what the generator actually holds. THE CENTRE scale runs 0.00 to 1.50 per cent true 1.00 reported 0.50 0.00 1.50 half the truth, 0.50 short THE SPREAD close view: the scale runs only 4.90 to 5.10 per cent true 5.00 reported 4.97 4.90 5.10 0.03 short, very nearly right THE LEAN scale runs minus 0.30 to 0.30 true 0.00 reported minus 0.05 minus 0.30 0.30 a lean reported where there is none THE TAILS IT REPORTED ON Kurtosis reported as 3.00, the value the normal shape carries. The population is five values with a hard wall at each end. wall wall minus 9 minus 4 1 6 11
A correctly computed summary of fifty months reported a mean of 0.50 per cent against a true 1.00 and a kurtosis of 3.00 for a population with no tails at all.

Take that in properly. Nobody made a mistake. Every one of the twelve lines was computed correctly from a record that was collected honestly, and this was a good summary by any standard anybody could apply to it from the outside. A complete, honest, correctly computed summary of fifty months got the centre half wrong and said nothing true about the shape, and its uncertainty half said so.

The last clause is the whole point. Look back at line eleven. The interval for the mean runs from minus 0.88 to 1.88 per cent, and the true average of 1.00 sits comfortably inside it. The point estimatorThe recipe a figure was produced by, as opposed to the figure itself. Taking the average of whatever months are to hand is an estimator; the 0.50 per cent it returned on this record is one output of it. missed by half and the interval was still right. A reader given only the nine descriptive lines would have walked away believing the unit runs at 0.50 per cent a month. A reader given all twelve would have walked away knowing that anything from roughly minus 0.9 to roughly 1.9 was on the table. Less satisfying to be told, and very much more accurate. Neither reader was given a wrong number. Only one of them was given enough.

Try it out

The summary reports a kurtosis of 3.00 and the population it came from takes five values with a hard wall at each end. What does that say about reading a single summary line as a description of a shape?

AI For Finance Bootcamp — Fin Maverick Reading an Option Payoff — free micro-course from Fin Maverick

How is the template used on a record of another kind?

An analyst who is handed somebody else's summary does one thing before reading a single figure: they count the lines. Nine means an incomplete document, whatever else is in it, and the correct next move is to go back and ask for the standard error rather than to start reasoning from the mean. A mean of monthly takings with no spread and no count beside it is not evidence of anything, and reading it carefully cannot turn it into evidence. A lender reading a borrower's trading history does the same count in reverse. And a household deciding whether one year of a side business is enough to plan around is asking exactly the question the uncertainty half answers.

So here is the template as something to copy, twelve lines and three about themselves.

HalfLine to fillFilled for this record
Descriptive1. Count of cases50 months
Descriptive2. Arithmetic mean0.50 per cent
Descriptive3. Geometric mean0.38 per cent
Descriptive4. Median1.00 per cent
Descriptive5. Mode1.00 per cent
Descriptive6. Standard deviation, denominator named4.97 per cent
Descriptive7. Range, lowest to highest20.00 per cent
Descriptive8. Skewnessminus 0.05
Descriptive9. Kurtosis3.00
Uncertainty10. Standard error of the mean0.70 per cent
Uncertainty11. Interval for the meanminus 0.88 to 1.88 per cent
Uncertainty12. Interval for a single caseminus 9.35 to 10.35 per cent
About itselfA. What the cases are, over what stretchMonthly changes of an invented unit, fifty months
About itselfB. What was excluded, and how muchNothing excluded, all fifty months in
About itselfC. Conventions usedCorrected denominator, normal shape for both multipliers

And one instruction is worth carrying away above all the others. A blank line and a line that was never there look identical to everybody except the person who wrote them, so fill every line or write not available beside it. Not available is real information. Those two words tell a reader that somebody considered the line, could not fill it, and said so. A reader can respond to that: they can ask why, they can go and get it, they can discount the summary accordingly. A reader who does not know a line is missing has none of those moves.

A BLANK ROW AND THE WORDS NOT AVAILABLE ARE NOT THE SAME THING Both authors failed to fill the line. Only one of them told the reader. LEFT BLANK Count50 months Arithmetic mean0.50 per cent Standard deviation4.97 per cent Standard error Skewnessminus 0.05 MARKED NOT AVAILABLE Count50 months Arithmetic mean0.50 per cent Standard deviation4.97 per cent Standard errornot available Skewnessminus 0.05 The reader cannot tell whether this was forgotten, withheld or simply not computed. Nothing here can be acted on. The reader now knows the line was considered and could not be filled, so they can go and ask. This is a usable admission.
A blank line and a missing line look identical to everybody except the person who wrote them, so write not available instead.
Try it out

A summary arrives with one row left completely blank. What should have been written there, and why does it matter?

How this goes wrong, with every figure in the document correct

A summary of the fifty month record is circulated. The document carries four things: a count of 50 months, an arithmetic mean of 0.50 per cent, a spread of 4.97 per cent beside it, and the lowest and highest months of minus 9.00 and 11.00. Every one of those figures is correct. The arithmetic is right, so anybody who checks it will find it right.

A reader takes the 0.50 per cent as the unit's monthly figure and builds on it. The true figure is 1.00 per cent, twice what they were told. The three lines that would have warned them, a standard error of 0.70 and the two intervals, are not in the document. Nobody left them out to conceal anything. The omission happened because nobody asked for them and because the four lines already present looked finished.

The summary was not wrong and it was still misleading. With no incorrect figure for anybody to point at, review is the hardest place of all to catch such a failure. Send it back and the author will defend it line by line and win every exchange. The document is unimpeachable and the reader has still been sent somewhere the record does not support.

Two fixes, and they are both about habit rather than skill. Put the uncertainty half inside the summary rather than in an appendix to it. The three lines then go in by default and have to be argued out rather than argued in. And not available goes against any line that cannot be filled. A reader can respond to not available and cannot respond to a line that was never there.

Try it out

Every figure in a summary is correct and the summary is still misleading. Name the mechanism, and say what would have prevented it.

How any single line is computed is covered separately. The two means, the median and the mode, the standard deviation and the range, the skewness and the kurtosis, the standard error and both intervals were each built separately and are used here rather than rebuilt, and which of the two intervals a particular reader needs is covered separately too. Testing a stated position against a record, and reporting a verdict on it, is covered separately. Threading a curve between scattered dots, adjusting a quantity until a shape settles onto the counts, and putting a figure on a month that has not happened yet are three different activities from summarising, and each is covered separately and much later. So is anything that treats a record as a series with an order running through it.

Four correct lines and the record still misled. See what else looks finished.

Twelve lines and no citation. Where does the authority come from?

From the arithmetic itself. Not one of the twelve lines is a rule somebody set, a threshold somebody published or a quantity somebody went out and measured. All twelve can be rebuilt from the five counts printed near the top, giving the same answers on any machine in any year, with nobody's permission required. Arithmetic carries its own warrant, and no maintained record sits underneath a summary template, so there is no date on which one stops being true.

What was usedWhere it came fromHow it can be checkedCarries a date
The five readings and their counts of 5, 9, 25, 8 and 3Typed out for teaching before anything was computedAdd the counts and confirm they reach fiftyNo
The nine descriptive linesRecomputed from those counts aloneRedo the arithmetic on the fifty readingsNo
The three uncertainty linesThe spread and the count, divided and multipliedDivide 4.97 by the square root of fiftyNo
The true average of 1.00 and true spread of 5.00 per centProperties of the generator, fixed before any month existedWeight the five readings by 0.08, 0.18, 0.48, 0.18, 0.08No
The order the six stages are built inOrdinary working practice, nobody's named methodIt is a working habit, not a rule to look upNot applicable

The Nakshatra unit and its fifty month record are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.