Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quantitative Methods, Financial Data & Programming
1Probability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
2Statistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
3Correlation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
4Time Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
5Simulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
6Optimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
7Modelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
8Backtesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
9Data Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
10Programming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
11Quantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity

Variable Types: Continuous, Discrete, Categorical and Ordinal

A continuous variable can land anywhere in a range, so the monthly change of a traded unit is continuous. A discrete variable takes separated values with nothing in between, so a count of months is discrete. A categorical variable is a label with no order. An ordinal variable is a label with an order but no distance. The type decides which summary is legal, and the mean is the one most often taken anyway.

Fifty months of a made-up traded unit sit below, picked so that every count here can be added up on the back of an envelope. Each count, each label and each rank that follows is built out of those fifty months alone.

Underneath the answer sits one habit worth building before any arithmetic starts. Somebody hands the analyst a column. Before a single thing is computed from it, the question to ask is what kind of quantity is sitting in it. Naming the kind of quantity takes about ten seconds, the naming is almost never done out loud, and every summary taken afterwards depends on getting it right.

Very little is needed to follow the argument. Three things are already built and are used below without being explained again: a record of fifty months of a made-up traded unit, the tally that record is stored as, and the three centres that can be taken of a set of numbers. Everything new is a distinction rather than a formula. Fitting a relationship to a record is covered separately, under fitted relationships.

What is a continuous variable?

A continuous variable can take any value in a range, and between any two values it can take there is always another one. The definition is the whole of it, and the second half of the definition is where the work happens.

The monthly change of the Nakshatra unit is continuous. The Nakshatra unit is an invented traded unitAnything with a price that gets quoted over and over, so a change from one quote to the next can be measured. The identity of the thing bought and sold does not matter here., and its monthly changeHow far a price moved over one month, counted against each hundred rupees it started the month at. Going from Rs 100/- to Rs 106/- across a month is a monthly change of 6.00 per cent. is the one finance word this guide needs. A price can move by 1.30 per cent over a month. The same price can move by 1.37 per cent. The same price can move by 1.3712 per cent, and nothing about a price forbids any of those. Between any two changes there sits a third.

The fifty month record shows only five distinct values, and that is a fact about how the record was written down rather than a fact about the quantity underneath it. Somebody decided to store each month as one of five settings. The price itself was never consulted about that decision and was never restricted by it. Five is a property of the file, not of the world.

Think about a child being measured against a doorframe. The pencil marks land on a few dozen centimetre lines over the years, so the record of that child holds a few dozen values. Nobody believes the child grew in centimetre jumps and stood still in between. The child grew continuously and the doorframe recorded coarsely, and those are two different things that happen to live in one place.

What is a discrete variable, and how is it different?

A discrete variable takes separated values with nothing at all in between them. Not nothing that happened to be recorded. Nothing that can exist.

The counts in the fifty month record are discrete. Stored as a tallyA record kept as how many times each value turned up, rather than as a list of every observation in order. Fifty months become five counts, and the order of the months is dropped., the record reads 5, 9, 25, 8 and 3 months against the five values, and those five counts add to 50. There is no such thing as 8.4 months in a tally. A month cannot be half-counted. Between 8 and 9 there is a genuine hole, and the hole belongs to counting rather than to the recording.

Discrete does not mean small and it does not mean whole numbers only; it means the values are separated, and the gap between them is real rather than a rounding choice. Rupee amounts to the nearest paisa are discrete, and the gap is one paisa. Shoe sizes go in halves, and the gap is half a size. A quantity with ten thousand allowed settings is still discrete if nothing is permitted between two neighbouring settings.

Here is the test that separates the two types, and it fits in one line. Take any two values the variable can show, and ask whether something is allowed to sit between them. If yes, always, the variable is continuous. If there is even one pair with a genuine hole between them, it is discrete.

Zooming a continuous scale never runs out of room. Zooming a discrete one reaches a hole. CONTINUOUS: A MONTHLY CHANGE, ZOOMED TWICE 1.00 per cent 2.00 per cent 1.30 and 1.40 1.30 per cent 1.40 per cent 1.37 sits here 1.370 per cent 1.380 per cent 1.3712 sits here, and the zooming could carry on all day DISCRETE: THE COUNTS A TALLY IS ALLOWED TO SHOW 4 5 6 7 8 9 10 11 12 NOTHING LIVES BETWEEN 8 AND 9 Every count in the tally, 5, 9, 25, 8 and 3, lands on a mark like these and never in a shaded strip.
Between any two values of a continuous quantity there is always another value, which is why the zoom never runs out, and between two discrete counts there is nothing at all.
Try it out

The tally holds 5, 9, 25, 8 and 3 months. Is a count of months continuous or discrete, and how would it be shown?

What is a categorical variable?

A categorical variable is a label with no order. Two labels, ten labels, a hundred labels, and no defensible way to say which one comes first.

Build one out of the same fifty months. Set a thresholdA stated cut-off used to sort things into groups. The cut-off is chosen by whoever is doing the sorting, and moving the cut-off moves which cases land on which side. at 0.00 per cent and label every month either a fall or not a fall. The months at minus 9.00 per cent are falls and there are 5 of them. The months at minus 4.00 per cent are falls too and there are 9 of them. Add those: 5 plus 9 gives 14 falls. Everything else is not a fall. Adding 25, 8 and 3 gives 36. The two counts come to 50, and every month is accounted for once.

The labels fall and not a fall cannot be added, averaged or ranked. The labels can be counted, and a count is a perfectly respectable summary in its own right. Fourteen and thirty six is a real finding about the record. The finding just is not a measurement.

Trying the average anyway is where the damage is easiest to see. Nothing stops anybody writing fall as 1 and not a fall as 2, then averaging the codes. Fourteen months at 1 and thirty six at 2 gives 14 plus 72, or 86, and 86 over 50 is 1.72. A perfectly clean piece of arithmetic. Nobody could call it wrong to flip the codes, so let a fall be 2 and not a fall be 1. The same fifty months then give 28 plus 36, or 64, and 64 over 50 is 1.28. With 7 and 3 instead, the same record answers 4.12.

The code that happens to be usedThe arithmeticThe average it produces
A fall is 1, not a fall is 214 times 1 plus 36 times 2, over 501.72
A fall is 2, not a fall is 114 times 2 plus 36 times 1, over 501.28
A fall is 7, not a fall is 314 times 7 plus 36 times 3, over 504.12
The count of falls, under every coding above5 plus 9, straight off the tally14

Three answers, one record, no arithmetic error anywhere. Because an average over a categorical column reports the coding somebody chose rather than anything about the months, the operation is illegal on that type. The count in the last row, by contrast, does not budge whichever codes are used, and that steadiness is what makes it a legitimate summary.

Try it out

A column holds nothing but the labels fall and not a fall. What is the strongest summary that can legally be taken from it?

Risk Management Program Bootcamp — Fin Maverick

What is an ordinal variable, and why is it not just a number?

An ordinal variable is a label that carries an order but no measurable distance between its steps. The ordinal type causes almost all of the trouble, and it causes the trouble quietly.

Rank the same fifty months worst to best in five bands. Band 1 holds the 5 worst months, band 2 holds 9, band 3 holds 25, band 4 holds 8 and band 5 holds the 3 best. The order is defensible and nobody would argue about it. Band 1 sits below band 2, band 2 sits below band 3, and so on up. The column never says how far apart the bands are. Nothing in a rank claims that band 2 minus band 1 equals band 5 minus band 4.

An ordinal column will happily yield a medianThe middle value once everything is sorted, with half the observations on either side. How to compute it and when to prefer it is worked through separately.. Sort the fifty months and look at where the middle lands. The counts run 5, then 14 by the end of band 2, then 39 by the end of band 3. The twenty fifth and twenty sixth months both fall inside band 3, so the median band is 3. The median band was found using the order and nothing else, and using nothing else is precisely why the median is allowed.

Ordinal labels are almost always stored as numbers, so nothing in the file stops them being averaged, and the average is meaningless. The file does not know that 3 is a name. The file sees a number, the software sees a number, the spreadsheet computes happily, and out comes a figure with two decimal places that looks exactly like a measurement.

An everyday version, so the idea lands before the arithmetic does. A cinema hands out first, second and third prize. The order is known. Nobody knows whether second was a whisker behind first and third was miles back, or whether the three finished evenly spaced. Averaging the prize positions of a group of entrants gives 2.4 or 1.8, and no statement about the actual performances comes out the other side.

A mean band of 2.90 is a position on the numbering. It is not a position on the thing being ranked. THE BAND NUMBERS, WHICH ARE EVENLY SPACED BECAUSE SOMEBODY NUMBERED THEM 1 2 3 4 5 THE MEAN BAND, 2.90, WHICH READS AS ALMOST BAND 3 THE SAME FIVE BANDS PLACED AT WHAT THEY COULD STAND FOR, WHICH A RANK NEVER REPORTS 1 2 3 4 5 ? HOW WIDE IS THIS GAP? A RANK NEVER SAID, SO 2.90 HAS NO PLACE HERE THE TOP RULE IS THE NUMBERING. ONLY THE BOTTOM RULE IS THE MONTHS.
Band numbers are evenly spaced because somebody numbered them, so an average of those numbers points at a place on the numbering and at no band on the thing being ranked.
Try it out

A column holds bands 1 to 5, worst to best. Can the median of that column be taken, and can the mean?

Breaking Into Quants Bootcamp — Fin Maverick

What do the same fifty months look like written all four ways?

One record, four notations, and none of them adds a single month or removes one. The table below holds the whole distinction in one place, and every figure in it is worked from the tally rather than carried in.

A made-up record allows one statement that a real record never allows. The generatorThe made-up rule that hands out the values, carrying a fixed weight on each one. Since it existed before any month came out of it, what it does on average is a stated fact and never a guess. behind these fifty months was written down before any month was drawn from it, so its true centre of 1.00 per cent and its true spread of 5.00 per cent are known by construction. The record itself reports a mean of 0.50 per cent, half the true centre. The gap between the two is not a mistake in the record. A gap of that size is what fifty months happen to look like, and seeing the gap at all is a luxury a real dataset never offers.

Written asWhat the column holdsWhat is legal on it
A continuous measurementFifty monthly changes, each of which could in principle have been any value at allThe mean of 0.50 per cent, the spread of 4.97 per cent, the median, the mode and a count
A discrete tally5, 9, 25, 8 and 3 months against the five values, adding to 50Everything above, since the counts themselves are separated numbers with real gaps
A categorical columnFall or not a fall against a threshold of 0.00 per cent, giving 5 plus 9, which is 14, against 25 plus 8 plus 3, which is 36A count of each label and the most common label. Nothing else
An ordinal columnBands 1 to 5, worst to best, carrying 5, 9, 25, 8 and 3 monthsThe median band, which is band 3, plus the counts and the most common band
All fourFifty months, every timeFewer summaries with each step down the table

Check the third row yourself rather than taking it: 5 plus 9 is 14, and 25 plus 8 plus 3 is 36, and 14 plus 36 is 50. Check the fourth row too: the running counts reach 5, then 14, then 39, so the twenty fifth and twenty sixth months sit inside band 3 and the median band is 3.

The mean of 0.50 per cent is a point estimateOne number handed over as the whole answer, standing alone, saying nothing about how far off it might be. How wide a range around such a number ought to run is worked out separately. that survives the first two rows and dies in the third, and that is the entire cost of a change of notation. Nothing about the months changed between row two and row three. Only the writing did.

Try it out

The monthly change shows only five distinct values across this record. Does that make the monthly change a discrete variable?

The same fifty months, four notations, and the summaries permitted shrink from left to right. NOTHING IS ADDED AND NOTHING IS REMOVED. ONLY THE WRITING CHANGES. CONTINUOUS DISCRETE CATEGORICAL ORDINAL minus 9.00 minus 4.00 1.00 6.00 11.00 5 months 9 months 25 months 8 months 3 months FALL 14 NOT A FALL 36 SWAP THEM AND NOTHING CHANGES BAND 5 BAND 4 BAND 3 BAND 2 BAND 1 3 8 25 9 5 No gaps anywhere Real gaps between No order at all Order, no distance Mean allowed Mean allowed Counts only Median, no mean
The same fifty months can be written as a continuous measurement, a discrete tally, a categorical label or an ordinal band, and only the first two of those four permit a mean.

Which summaries is each type allowed?

Five summaries, four types, and one grid that settles every question of permission in this guide. The rule underneath it is short: each summary needs something from the values, and a type either supplies that thing or does not.

A count needs nothing at all. A count needs only that the values be distinguishable from each other, and every type manages that, so a count works everywhere. The mode, meaning the most common value, needs the same and no more: two entries have to be tellable apart or not, and that is it.

The median needs an order. Finding the middle requires sorting, and sorting a column of fall and not a fall is not possible in any defensible way, so a categorical column has no median. Ordinal, discrete and continuous columns all sort, so all three have one.

The arithmetic mean and the standard deviationA single number describing how far the values sit from their centre on average. The standard deviation is worked out separately. Only its demand on a column matters here. need more than an order. Both operations subtract and add, so both need the distance between values to mean something. A mean adds values together. A standard deviation subtracts each value from the centre. Neither operation makes sense unless a difference of one, anywhere on the scale, means the same amount of the thing being measured.

SummaryCategoricalOrdinalDiscreteContinuous
A count of each valueAllowedAllowedAllowedAllowed
The mode, the most common valueAllowedAllowedAllowedAllowed
The median, the middle valueNot allowedAllowedAllowedAllowed
The arithmetic meanNot allowedNot allowedAllowedAllowed
The standard deviationNot allowedNot allowedAllowedAllowed
What the summary needs from the valuesOnly that they can be told apartAn orderA distance that means the same everywhere

The single line worth carrying away from this grid is that a mean requires the distance between values to mean something, and on an ordinal column it does not. Everything else in the table follows from that one sentence and from the two weaker requirements above it.

Read down a column and the permissions run out. Read across to see what a summary needs. SUMMARY CATEGORICAL ORDINAL DISCRETE CONTINUOUS A count of each value The mode, most common The median, the middle The arithmetic mean The standard deviation ALLOWED ALLOWED ALLOWED ALLOWED ALLOWED ALLOWED ALLOWED ALLOWED NOT ALLOWED ALLOWED ALLOWED ALLOWED NOT ALLOWED NOT ALLOWED ALLOWED ALLOWED NOT ALLOWED NOT ALLOWED ALLOWED ALLOWED A COUNT NEEDS NOTHING. A MEDIAN NEEDS AN ORDER. A MEAN NEEDS A DISTANCE.
A count and a mode work on all four types, a median needs an order, and a mean and a standard deviation both need a meaningful distance between the values.
Try it out

Which of the four types allow a standard deviation, and what does a standard deviation need from the values?

What is destroyed when a measurement becomes a label?

Turning the fifty month record into fall and not a fall leaves 14 and 36. A great deal walked out of the room on the way.

The mean of 0.50 per cent is gone. The spread of 4.97 per cent is gone. The five values are gone. And the month at minus 9.00 per cent now sits in the same box as a month at minus 0.01 per cent, entirely indistinguishable from it. Both months are below the threshold, and the label records nothing beyond that fact.

The last loss is worth pausing on, and it is the one people underestimate. Inside the group labelled a fall, the months run from minus 9.00 per cent to minus 4.00 per cent, a span of 5.00 percentage points. Inside the group labelled not a fall, they run from 1.00 per cent to 11.00 per cent, a span of 10.00 percentage points. The label column does not know either span. The label column knows 14 and 36.

The destruction runs one way: labels can always be built out of measurements, and measurements can never be built back out of labels. This is not a limitation of any particular software or any particular effort. The information is not hidden or compressed. The information was never written down.

Fifty measured months go in at the top. Two counts come out at the bottom. Nothing else survives. BOTH SETS OF BARS ARE DRAWN AT 3.2 PIXELS TO ONE MONTH 5 9 25 8 3 minus 9.00 minus 4.00 1.00 6.00 11.00 LABEL AGAINST A THRESHOLD OF 0.00 PER CENT 14 36 FALL NOT A FALL GONE FOR GOOD mean 0.50 per cent spread 4.97 per cent the worst month the five values The two bottom bars still add to fifty. That is the only thing they kept.
Turning the fifty month record into falls and not falls leaves 14 and 36, while the record's mean of 0.50 per cent is gone for good along with its spread of 4.97 per cent.

A household version makes the one way street obvious. A diary that records every month as an amount can always be revisited later and each month marked good or tight against whatever line is chosen, including a line nobody had thought of at the start. A diary kept as good or tight from the beginning hides which of the tight months was the worst and by how much, and no amount of later thought recovers either.

One arrow exists. The other one does not, and no effort or software creates it. FIFTY MEASURED MONTHS every value, in full 14 FALLS AND 36 NOT two counts, nothing else LABEL THEM NOTHING COMES BACK LABELS ARE MADE FROM MEASUREMENTS. MEASUREMENTS ARE NEVER MADE FROM LABELS.
Labels can always be built out of measurements and measurements can never be recovered from labels, which is why the conversion is a decision rather than a formatting step.
Try it out

The fifty measured months become 14 falls and 36 that were not. Which two figures can no longer be computed from what is left?

Try it out

Before the panel below is touched, an answer is worth committing to. As the threshold moves from 0.00 per cent up to 6.00 per cent, what happens to the mean of the record?

Play with it

Slide the threshold and watch fifty measurements collapse into two counts.

The panel opens on the split used above: a threshold of 0.00 per cent, giving 14 months below it and 36 at or above. Check that off the tally as 5 plus 9 against 25 plus 8 plus 3. Moving the slider changes where the line is drawn and nothing else. The five measured bars stay exactly where they are, and only their colour reports which side they landed on. The two category blocks in front of them and the strip underneath redraw. The two counts always add to fifty, so the strip in particular always fills. The strip never reports anything about how far apart the months inside each segment were.

MOVE THE LINE. THE FIVE MEASURED BARS NEVER MOVE. EVERY BAR AND BLOCK IS DRAWN AT 3.4 PIXELS TO ONE MONTH 5 9 25 8 3 minus 10.00 minus 5.00 0.00 5.00 10.00 THE THRESHOLD THE CATEGORICAL COLUMN, AND IT IS EVERYTHING THAT SURVIVES THE LABELLING 14 BELOW 36 AT OR ABOVE The strip always fills exactly once, because the two counts always add to fifty months.
The threshold
0.00 per cent
Months below it
14
Months at or above
36
The two counts add to
50
Span buried in the below group
5.00 per cent
Mean, not in the labels
0.50 per cent
Spread, not in the labels
4.97 per cent
Educational illustration on an invented object. Both the Nakshatra unit and the fifty months behind this panel were made up for these notes. The five values stay pinned at minus 9.00, minus 4.00, 1.00, 6.00 and 11.00 per cent while only the line moves; the record always carries exactly fifty months, so the two category counts always come to 50. The last two readouts are greyed and struck through because both are computed off the measurements and neither is available from the label column at any setting of the line.
AI For Finance Bootcamp — Fin Maverick Bond Pricing and Yield Mechanics — free micro-course from Fin Maverick

Why does anyone throw that away on purpose?

Because labels are cheaper in every way that matters at the moment of collection, and pretending otherwise would leave half the world's records unexplained.

A label is easier to collect. Asking somebody whether the month was a fall gets an answer in a second; asking for the exact change gets a shrug or a wrong number. A label is easier to agree on. Two people will settle on good or tight far faster than they will settle on a figure to two decimal places. And a label is often all that survives from an older file, where somebody long gone recorded a grade because a grade was what the form had space for.

Converting a measurement down to a label is a real decision with a real cost, and the cost belongs in the same sentence as the convenience rather than three paragraphs away from it. The sentence to aim for is something like: these months were recorded as bands because the branch staff could apply bands consistently and could not apply amounts, and the price of that choice is that no mean and no spread exist for this period.

The one salary household again. A household that writes down good month or tight month every month has a record it can actually keep up for years, and keeping it up is worth a great deal. The same household will never afterwards be able to say by how much the tight months were tight. Both halves of that are true at once. The mistake is not the choice. The mistake is forgetting the choice was made, and then quoting an average of the good and tight codes at somebody five years later.

Where this goes wrong: the mean band of 2.90

The older file kept bands, so the record stores every month as a band from 1 to 5, worst to best. A reader opens the record, averages the band numbers and gets 145 over 50, or 2.90. The reader writes it up as a mean band of 2.90.

Nothing in the file objected. The bands were stored as numbers, the software computed happily, and 2.90 with its two decimal places reads exactly like a measurement. The figure is not one. Band 3 covers months at 1.00 per cent and band 2 covers months at minus 4.00 per cent, and 2.90 does not translate back into any monthly change at all. The figure names a position on the numbering, and the numbering was chosen by whoever set up the file.

The cost is a figure with no units that then gets compared across records and across periods as though it had some. A second reader sees 2.90 here and 3.10 in a different record and reports an improvement, without either number having ever measured anything.

A sharper version of the trap sits in these five values, and it has to be said out loud. The five values here are minus 9.00, minus 4.00, 1.00, 6.00 and 11.00 per cent, exactly five percentage points apart at every single step. Under that even spacing the mean band does convert. Working it through, 2.90 maps to minus 9.00 plus 1.90 times 5.00, or 0.50 per cent, and 0.50 per cent is precisely the record's mean. The arithmetic worked. The arithmetic worked because of a property of these particular values that almost no real banding has, and the reader who takes the coincidence as permission has learned the opposite of the lesson.

Test it by changing one thing. Suppose band 5 stood for 31.00 per cent instead of 11.00 per cent. The band column does not change by a single entry, the counts are still 5, 9, 25, 8 and 3, and the mean band is still 2.90. The mean of the months, though, moves to 85 over 50, or 1.70 per cent. One reported figure of 2.90, two entirely different records underneath it.

The fix is to report the median band, band 3, together with the five counts, and to say in plain words that no mean band was available. That sentence looks weaker in a report and is the only one of the three that is true.

Try it out

A colleague reports a mean band of 2.90 across a record stored in bands 1 to 5. What is wrong with that sentence?

How is the type of a column worked out?

Somebody hands over a file with a column headed grade, or score, or level, and no description. Three questions settle it, in order, and none of them needs any software.

Question one: can two values be subtracted to give something that means anything? If yes, the column is discrete or continuous, and a mean is already allowed. If no, the next question follows.

Question two, asked only when the first answered yes: can a value sit between two neighbouring values? If yes it is continuous, and if there is a genuine hole it is discrete. Question two is the only one that separates those two types. Getting the answer wrong costs very little, and that is why the question comes second.

Question three, asked only when the first answered no: can the values be put in a defensible order? If yes the column is ordinal, and a median and a mode are permitted and nothing above them. If no it is categorical, and counts and a mode are permitted.

A column of numbers answers none of these three questions by itself, and that silence is exactly why the mistake keeps getting made. The numbers 1, 2, 3, 4 and 5 in a file are equally consistent with a measurement, with a rank and with a code for five branches, and the file rarely says which. A credit analyst handed a legacy file of internal grades runs into this on the first afternoon, and the honest move is to go and find whoever set the file up, or, failing that, to report counts and a median and say why.

The same three questions are worth running on a household budget sheet. A column of amounts spent passes question one immediately. A column headed how the month felt, scored 1 to 5, fails question one and passes question three, so it gets a median and never an average, however tempting the spreadsheet makes it.

Three questions, asked in this order, land every column on exactly one of the four types. CAN TWO VALUES BE SUBTRACTED AND GET SOMETHING MEANINGFUL? NO YES CAN THE VALUES BE PUT IN A DEFENSIBLE ORDER? CAN A VALUE SIT BETWEEN TWO NEIGHBOURING VALUES? NO YES NO YES CATEGORICAL ORDINAL DISCRETE CONTINUOUS counts and a mode add a median add a mean and a spread add a mean and a spread A COLUMN OF NUMBERS ANSWERS NONE OF THESE THREE QUESTIONS BY ITSELF.
Ask whether subtraction is meaningful, then whether values in between exist, then whether the values can be defensibly ordered, and each path ends on one of the four types.
Try it out

A column of numbers arrives with no description of any kind. What is the first question that separates a measurement from a code?

Covered elsewhere. No particular summary figure is worked through in depth here. A mean, a median, a mode, a spread and an interval, and when each is the right one to quote, are all covered separately. The shape of a record, meaning how lopsided it is and how heavy its extremes are, is covered separately too. Measurement error, a different problem from a variable of the wrong type, is covered separately, as is how a variable of any type would be used inside a fitted relationship.

The mean band held while the months moved. See what each variable type permits.

Where can any of this be checked?

All of it can be checked here, with a pencil. Every figure underneath is either one of the fifty made-up months or an addition performed on those months in plain view, so nothing needs looking up and nothing can go out of date.

The figure statedHow that figure was producedSiteDate read
The five monthly changes and how often each turns upWritten down for teaching before anything was worked out from them. The counts 5, 9, 25, 8 and 3 add to 50None. Nothing was retrievedNot applicable
The 14 falls and the 36 that were notAdded here from those counts, as 5 plus 9 against 25 plus 8 plus 3None. Addition onlyNot applicable
The mean of 0.50 per cent and the spread of 4.97 per centComputed from the same fifty months and recomputed here rather than carried acrossNone. Arithmetic onlyNot applicable
The true centre of 1.00 per cent and true spread of 5.00 per centStated properties of the made-up generator, written down before any month was drawn from itNone. Stated, not measuredNot applicable
The four type names themselvesCommon property in the study of measurement, stated here without being attached to any one author, because no single attribution is safe without checking the textNone. No document quotedNot applicable

The Nakshatra unit and the fifty months recorded against it are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.