Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quant Analyst · CoreTrack
1Quantitative Methods, Financial Data & Programming
iProbability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
iiStatistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
iiiCorrelation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
ivTime Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
vSimulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
viOptimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
viiModelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
viiiBacktesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
ixData Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
xProgramming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
xiQuantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity
2Stochastic Calculus & Derivative Pricing Theory
iProbability Foundations
The Probability SpaceRandom VectorsSigma-AlgebraExpectationSample Space and EventsDensity and Distribution FunctionsRisk-Neutral ProbabilityState Price Density vs…
iiStochastic Processes and Jumps
Properties of a Stochastic ProcessMartingaleBrownian Motion and Its PropertiesBrownian Motion vs Geometric…Stopping TimeThe Markov PropertyState VariablesTransition ProbabilityQuadratic VariationQuadratic Variation vs Ordinary…Submartingale and SupermartingaleMartingale RepresentationMarkov Process vs MartingaleOptional StoppingFiltrationJump ProcessesThe Poisson ProcessLevy ProcessesJump Diffusion
iiiIto Calculus
The Ito IntegralThe Ito Integral vs the Riemann IntegralInfinitesimals in Stochastic CalculusQuadratic CovariationIto's LemmaHow to Apply Ito's…The Infinitesimal GeneratorIto Calculus vs Ordinary Calculus
ivStochastic Differential Equations
Stochastic Differential EquationsStochastic Differential Equation vs…Drift and DiffusionStrong and Weak Solutions ComparedDiscretisationGeometric Brownian Motion
vPricing Theory and No-Arbitrage
No-ArbitrageGirsanov, Radon-Nikodym and Change…Physical and Risk-Neutral Measures…The Fundamental Theorems of…The Law of One PriceThe Pricing KernelDiscount Factors and Zero-Coupon PricesReplication vs HedgingComplete Market vs Incomplete MarketClearing Margin Architecture
viOption Pricing Theory
European and American OptionsMonte Carlo European OptionThe Black-Scholes PDEBlack Scholes and the GreeksThe Payoff FunctionThe Binomial ModelBinomial Option PricingDelta Hedging in TheoryBoundary, Initial and Terminal ConditionsThe Exercise BoundaryHow to Check Put-Call…
viiVolatility Models
Constant, Local and Stochastic…Vasicek Model vs CIR ModelThe Heston ModelThe SABR ModelThe Volatility ProcessImplied VolatilityVolatility Smile vs Skew vs Surface
viiiInterest Rate Models
Interest-Rate DerivativesMean ReversionThe Zero-Coupon BondThe Ornstein-Uhlenbeck ProcessThe Discount CurveZero RatesShort-Rate Model vs Market Model
ixNumerical Pricing
Closed Form and Numerical…Monte Carlo PricingEuler and Milstein Schemes ComparedTree MethodsFinite Difference MethodsNumerical Error and StabilityVariance Reduction
xCalibration and Model Risk
Model OverrideMarket Price and Model PriceCalibrationHow to Document a Pricing ModelThe Educational Illustration LabelMarket ConventionsModel Uncertainty and LimitationsBacktesting a Pricing ModelIdentifiabilityCalibrated ParametersThe Calibration Loss Function

Python for Finance: What It Is Used For, and the Table You Work In

Python here is one thing: a way to hold a record as a table in memory and ask it a question. The Neelbagh stall returns sit as 32 rows and 8 columns. Grouping those rows by the stall gives 8 groups. Grouping them by the stall name gives 9, and the ninth is a spelling rather than a stall.

Underneath that answer sit three things and nothing more. The first is the Neelbagh stall record itself, an office file covering four months of takings at an invented covered market of 14 pitches, along with the faults it already carries and the counting rule its takings figures answer to. The second is what one row of that file is a statement about. One row is one stall in one month. The third is counting. The counting never runs heavier than that: a count of rows, a count of columns, a count of groups, and one average that always travels with the count sitting underneath it.

What is Python used for here?

Four things. Reading a record in. Checking that all of it is there. Shaping it so it answers the question asked. Answering that question and printing the count the answer rests on. Reading, checking, shaping and answering are not a summary of a longer list. The four are the whole list.

Syntax for its own sake, loops explained because loops exist, keeping a value in a name, installing anything, and how software is packaged and shipped are all real subjects, taught extremely well by people whose job that is. None of them is what an analyst was hired for.

The everyday version is a stall keeper buying a calculator. There is a manual in the box with sixty functions in it, and the keeper learns four keys: add, subtract, the total, and clear. Not because the other fifty six are worthless, but because the day's takings need four and learning the rest before opening the stall is a way of never opening the stall. An analyst who has to answer what the average monthly takings were is in exactly that position on the first morning.

Try it out

Python is used here for four things, and a longer list is refused. Which set is the four?

What is a DataFrame, and what does the table hold that the file did not?

A DataFrame is rows and columns held in memory, with a name for every column, a type for every column, and a position for every row. The definition ends there, and it is worth reading twice. The three things a DataFrame promises are exactly the three things a file does not promise. The table the analyst works in is the file plus a declaration about what is in it.

The Neelbagh returns sit in a DataFrame as 32 rows and 8 columns, making 256 cells. The eight column names are not invented by anybody working on the record. The eight names are read straight off the header rowThe first line of a data file, carrying the column names instead of any reading of its own. The header row describes the lines below it and is never one of them. of the file the market office handed over. The names stall_id, stall_name, licence_no, category, month, takings_rupees, pitch_sqft and filed_on_day are the office's words rather than the analyst's.

THE NEELBAGH STALL RECORD, HELD AS A TABLEEight columns, each with a name and a type. Thirty two rows, each with a position.stall_idtextstall_nametextlicence_noan id, held as textcategorytextmonthwhole numbertakings_rupeesrupeespitch_sqftsquare feetfiled_on_dayday number1NB-01Kadamba Idli2104Cooked food142,00012062NB-02Chandan Tea2216Beverages131,0006063NB-03Harit Greens2318Vegetables138,00010064NB-04Peetal Utensils2405Household126,0008065NB-05Bansi Flour2477Groceries155,0001466NB-06Ilaka Fruit2530Vegetables134,00010067NB-07Sundari Chaat2618Cooked food140,0009068NB-08Peeli Mithai2743Cooked food142,00011069NB-01Kadamba Idli2104Cooked food244,000120710NB-02Chandan Tea2216Beverages233,00060711NB-03Harit Greens2318Vegetables236,400100512NB-03Harit Greens2318Vegetables239,7001001913NB-04Peetal Utensils2405Household2080714NB-05Bansi Flour2477Groceries2(blank)14715NB-06Ilaka Fruit2530Vegetables236,000100716NB-07Sundari Chaat2618Cooked food241,00090717NB-08Peeli Mithai2743Cooked food245,000110718NB-01Kadamba Idli2104Cooked food343,000120819NB-02Chandan Tea2216Beverages332,00060820NB-03Harit Green2318Vegetables341,000100821NB-05Bansi Flour2477Groceries354,00014822NB-06Ilaka Fruit2530Vegetables335,000100823NB-07Sundari Chaat2618Cooked food34,80,00090824NB-08Peeli Mithai2743Cooked food34,80,000110825NB-01Kadamba Idli2104Cooked food445,000120626NB-02Chandan Tea2216Beverages499,99960627NB-03Harit Green2318Vegetables440,000100628NB-04Peetal Utensils2405Household428,00080629NB-05Bansi Flour2477Groceries453,00014630NB-06Ilaka Fruit2530Vegetables437,000100631NB-07Sundari Chaat2618Cooked food442,00090632NB-08Peeli Mithai2743Cooked food447,000110632 rows, 8 columns, 256 cells. 31 of the 32 takings cells carry a number.the nine cells the record was already known to be wrong about, tinted so they can be found by eye
All 32 rows of the Neelbagh stall record held as a table, with a name and a type over every one of the 8 columns, giving 256 cells of which 31 takings cells carry a number.

Two of the three promises change the work, and it pays to be precise about which two. The first is that the types are real, so a column can be told what kind of thing it holds before anybody computes with it. The column licence_no is the clean case. The values 2104 and 2318 and 2743 are digits and nothing else, and every value in the column looks like a whole number. But each of those eight values is an identifierA code that points at one particular thing and does no other work. Because it labels rather than measures, a column of them can be sorted and matched but never sensibly totalled. rather than an amount, so their average is real arithmetic that describes nothing whatsoever. Hold that column as text and the mistake becomes impossible rather than merely unlikely.

The second is that the row order is a fact about the table rather than a fact about the market. Row 11 and row 12 both describe NB-03 in month 2, one filed on day 5 and one filed on day 19. The two rows sit next to each other because the office typed them that way, not because the market did anything twice in a row. Sort the table by takings and they move apart. Nothing about Neelbagh changed; only the positions did. Written down, the point sounds obvious, and it is still the single most common source of a wrong answer. A reader who has looked at the same file in the same order for a week starts treating position 12 as though it means something.

The third promise is the least glamorous and the most load bearing. Every column has a name, so the question is asked in the office's own vocabulary. The question does not ask for the sixth thing along. The question asks for takings_rupees, and if somebody inserts a column tomorrow the question still means what it meant yesterday.

ONE RECORD, TWO PLACES TO KEEP IT AS A FILE ON DISK AS A TABLE IN MEMORY DO THE COLUMNS CARRY A TYPE? No. Every value is a run of characters, and 2104 looks exactly like a quantity. Yes. licence_no can be held as text, so no average of it is even offered. DOES THE ROW ORDER MEAN ANYTHING? It is the typing order and nothing else, so no reading depends on it. Each row has a position, and the position belongs to the table, not to the market. WHAT IS AN EMPTY CELL? Two commas with nothing between them. A stated absence that can be counted and tested for.
Held as a table the column types are real and each row carries a position, and an empty takings cell becomes a stated absence that can be counted rather than two commas with nothing between them.
Try it out

In the file, licence_no was just a run of digits like any other. In the table it is held as text. What does the declared type actually prevent?

Breaking Into Quants Bootcamp — Fin Maverick

What is actually done to a table, and how few things are there?

Four moves, and they cover almost every question an analyst is asked of a single record. Look at its shape. Take one column. Keep some rows and drop the rest. Put rows into groups and read something off each group. Everything else on a first pass is one of those four wearing different clothes.

Looking at the shape means asking the table how many rows and how many columns it has, before asking it anything interesting. The table answers 32 and 8. The shape is not a formality. A file that lost four rows on the way in would read 28 and 8, and 28 and 8 look perfectly healthy in every step that follows. The shape check is the only cheap place to catch it.

THE FOUR MOVES, WITH THE ROW COUNT ON BOTH SIDES OF EACH MOVE ONE MOVE TWO MOVE THREE MOVE FOUR Look at the shape Take one column Keep some rows Put rows into groups 32 rows in 32 rows in 32 rows in 32 rows in 32 by 8 out 32 values out 31 rows out 8 groups out nothing moved one per row one blank dropped holding all 32 rows Only move three changes how many rows are held, and it is the only one that can silently lose a figure.
The four moves each carry a row count going in and a row count coming out, and only keeping rows changes the total, dropping the file from 32 rows to 31 usable takings values.

Here is the first of those moves written out. The code runs on the 32 rows printed above and nothing else.

# `rows` holds the 32 lines printed in the figure above, one entry per line.
# `columns` holds the eight names read off the header row of that same file.
shape = (len(rows), len(columns))
print(shape)                     # (32, 8)
print(shape[0] * shape[1])       # 256 cells in all

Asks the table for its shape and gets 32 rows by 8 columns, making 256 cells, before anything at all is computed from it.

Taking one column and keeping some rows are written out below, and notice what the comment is doing. The comment says why the test is there rather than restating the line. The line is already visible.

# takings_rupees is the sixth name in `columns`, so it is position 5 in every row.
takings = [r[5] for r in rows]
print(len(takings))              # 32, exactly one value per row

# a blank cell holds no number, so it cannot join anything that gets averaged
usable = [t for t in takings if t != ""]
print(len(usable))               # 31, and the one that went is NB-05 in month 2

Takes the takings column, gets one value per row, then keeps only the cells carrying a number and reports 31 out of 32.

A result computed from data the reader can see is a result the reader can check. A result computed from data fetched somewhere out of sight has to be taken on trust instead. The 32 rows above are printed in full for exactly that reason, and an argument resting on reading the count before reading the figure cannot then rest on a number nobody was shown.

Try it out

Every code block here runs on the record printed beside it. What does a reader lose if one of them pulls its data from somewhere else instead?

How are rows put into groups, and does the record stay whole?

Grouping is the fourth move and the one worth slowing down for. Grouping is where a language feature and a finance task stop being the same thing. The mechanical description is dull: pick a column, and every row goes into the box named by its value in that column. The finance description is the whole job. The column grouped by is the question being asked, and switching the column switches the question without touching a single figure.

Group the Neelbagh returns by stall_id and the question is how each stall did. Group them by month and the question is how the market did over time. Group them by category and the question is which kinds of trade the market runs on. Same 32 rows, same takings, three different questions, and no arithmetic in between.

Here is the ladder, and it is worth reading the counts before reading anything else.

Grouped byGroupsRows in each groupRows held
stall_id84, 4, 5, 3, 4, 4, 4, 432
stall_name94, 4, 2, 3, 4, 4, 4, 3, 432
month48, 9, 7, 832
category54, 12, 4, 3, 932
filed_on_day51, 16, 7, 7, 132
Every setting4 distinct countsthe sizes always add to 3232

Read the right hand column first: all five groupings hold all 32 rows, so nothing was lost and only the question changed. That column staying still is the point. A grouping that returned 30 rows would be a grouping that quietly threw two away, and that is worth knowing before a single takings figure is read.

FIVE COLUMNS, FIVE GROUP COUNTS, ONE RECORD The same 32 rows every time. Only the column being grouped by changes. 8 = the stalls that traded stall_id stall_name month category filed_on_day 8 9 4 5 5 same count, different rows 0 1 2 3 4 5 6 7 8 9
Five columns give group counts of 8, 9, 4, 5 and 5 from the same 32 rows, and only the count of 8 from stall_id equals the number of stalls that traded.

The code that produces those counts is four lines long and the last three of them are the check rather than the answer.

# every row goes into the box named by its value in one chosen column
groups = {}
for r in rows:
    groups.setdefault(r[0], []).append(r)      # r[0] is stall_id

print(len(groups))                              # 8 groups
print([len(g) for g in groups.values()])       # [4, 4, 5, 3, 4, 4, 4, 4]
print(sum(len(g) for g in groups.values()))  # 32, so nothing was lost

Groups the 32 rows by stall_id into 8 boxes holding 4, 4, 5, 3, 4, 4, 4 and 4 rows, and confirms the sizes still add back to 32.

WHAT THE GROUPING ACTUALLY PRINTSGROUPED BY stall_idNB-01Kadamba Idli4 rowsNB-02Chandan Tea4 rowsNB-03Harit Greens5 rowsNB-04Peetal Utensils3 rowsNB-05Bansi Flour4 rowsNB-06Ilaka Fruit4 rowsNB-07Sundari Chaat4 rowsNB-08Peeli Mithai4 rows8 groups32 rowsEight lines out, and eight stalls traded.The count is the check, and it takes one second.NB-03 shows 5 rows and not 4.That is the month its return was filed twice.The eight counts add back to 32.Grouping moved every row and lost none of them.
Grouping by stall_id prints eight lines holding 4, 4, 5, 3, 4, 4, 4 and 4 rows, and those counts add back to the 32 the record started with.

Grouping by month instead does something different. Four groups, as four months would suggest. But the sizes are 8, 9, 7 and 8, and a market of 8 trading stalls filing one return each should give four groups of 8. Two of those four numbers are the record's own faults arriving as counts rather than as figures. The 9 is month 2, where NB-03 filed a return on day 5 and then filed a revised one on day 19, so the office typed two rows for one stall month. The 7 is month 3, where NB-04 was shut and no return exists at all.

THE SAME ROWS, GROUPED BY MONTH Eight trading stalls filing one return each would give four groups of 8. 8 is what a whole month looks like 0 2 4 6 8 8 9 7 8 month 1 month 2 month 3 month 4 NB-03 filed twice NB-04 never filed
The four months hold 8, 9, 7 and 8 rows, and the 9 is the month one stall filed a revised return while the 7 is the month another stall was shut.
Try it out

Grouping the record by month gives 8, 9, 7 and 8 rows. Two of those four numbers are not 8. What is behind each one?

AI For Finance Bootcamp — Fin Maverick

What do the Neelbagh returns look like grouped five ways?

The shape comes first. Nothing interesting is asked of a table before its shape is known. The record is 32 rows and 8 columns, so 256 cells, and 31 of the 32 takings cells carry a number. Then the five groupings, one after another, on exactly those rows.

By stall_id: 8 groups holding 4, 4, 5, 3, 4, 4, 4 and 4 rows. By stall_name: 9 groups, and the two extras are one stall written two ways, Harit Greens on 3 rows and Harit Green on 2. By month: 4 groups holding 8, 9, 7 and 8. By category: 5 groups holding 4, 12, 4, 3 and 9, the 12 being the three cooked food stalls over four months. By filed_on_day: 5 groups again, holding 1, 16, 7, 7 and 1. Months 1 and 4 were both filed on day 6, and the two odd single rows are NB-03 filing on day 5 and then on day 19. One table, five questions, and all 32 rows present every single time.

Try it out

The same 32 rows are about to be grouped by five different columns in turn, using the control below. Does the number of rows held change as the column switches?

Play with it

Switch the column and watch the same 32 rows fall into different boxes

One control, and all it changes is which column the rows are grouped by. The record is the same 32 rows at every setting, none of them is dropped and no takings figure is touched. Two things redraw together. The scale along the top shows how many boxes the grouping ended up with, against a fixed mark at 8 for the number of stalls that traded. Below it, every group gets a line of 32 cells, and a filled cell means that row of the file landed in that group. The pattern of filled cells shows whether a group takes its rows in a block or scattered right across the file. Clicking any group line holds it and reads out what it holds; clicking it again lets it go. The control starts on stall_id, the grouping worked through above.

Rows in the file
32
Stalls trading
8
Grouped by
stall_id
Groups
8
Largest group
5 rows
Against 8 stalls
matches

Educational illustration. The Neelbagh market, its stall record and its day book were invented for teaching. The 32 rows here are the file as the office handed it over rather than the cleaned table, so the faults are still in it. The counting rule for takings admits only cells carrying a number, and no takings figure is computed in this panel at all.

Try it out

Grouping by stall_id gives 8 groups. Before reading the next block, say what grouping by stall_name will give and why the two would differ at all.

Spotting Quality of Earnings Red Flags — free micro-course from Fin Maverick

What does the count of groups tell the analyst before any figure does?

The count of groups tells the analyst whether the column just grouped by means what it was assumed to mean. The sentence is small and carries a lot of weight, so here it is against the record. Only one of the five group counts is the number of stalls that traded, and it is the 8 from stall_id. The 4 from month is the number of months. The 5 from category is the number of trades. The 5 from filed_on_day is the number of distinct filing days. And the 9 from stall_name is not the number of anything in the market at all.

Grouping by the name gives 9 because NB-03 is written Harit Greens on its months 1 and 2 rows and Harit Green on its months 3 and 4 rows, and month 2 carries two rows. So the long spelling holds 3 rows, the short spelling holds 2, and one stall appears as two groups. The split spelling is a fault found by counting rather than by looking, and that is exactly why it survives. Nobody scanning a printed name column notices a missing letter s halfway down. A count of 9 against an expected 8 takes about a second and does not depend on anybody being sharp that morning.

HOW 8 BECOMES 9 WITHOUT A SINGLE FIGURE MOVING GROUPED BY stall_id GROUPED BY stall_name NB-03 5 rows, one stall plus 7 more groups 8 groups change the column to stall_name same 32 rows, nothing dropped Harit Greens 3 rows Harit Green 2 rows one stall, written two ways 9 groups 3 plus 2 is still 5 rows, so no takings figure changed. Only the count of stalls did.
Grouping by the name splits NB-03 into 3 rows under one spelling and 2 under the other, so a market of 8 stalls reports as 9 while all 32 rows stay present.

Then there is the tie, and it needs labelling in the same breath as it is stated. Category gives 5 groups and filed_on_day gives 5 groups. The equality of the two counts is arithmetic on this particular record and not a finding about anything. The two columns do not group the same rows in any sense: the biggest overlap between any one category group and any one filing day group is 6 rows out of 32. Cooked food and day 6 are not versions of each other, they merely happen to produce the same number of boxes. The check that settles it takes one look: the two groupings set side by side, to see whether the same rows land together. Reading a shared count as a shared meaning invents a relationship out of a coincidence.

Try it out

Grouping by category gives 5 groups, and grouping by filed_on_day also gives 5. Does that say something about the record, and what one check settles it?

Spotting Quality of Earnings Red Flags teaches you to test whether a reported profit is a sound base to forecast from.

What is the answer, and what count sits underneath it?

The question the Neelbagh market office actually asked was what the average takings per stall month were. On the cleaned table that answer is Rs 54,196.55/- over 29 usable stall months, from a total of Rs 15,71,700/-. Both halves of that sentence travel together, always.

The conventionA rule about how something is counted, agreed in advance and put in writing. Its point is that somebody else following the note lands in exactly the same place. the market office settled says which takings cells are allowed into the average: a cell joins only if it carries a number. A blank does not. The office code standing for no return received does not join. The code is a message rather than an amount. A real zero does join. A stall that shut for a month and reported nothing genuinely took nothing. Apply that to the cleaned table and 29 cells qualify.

# `clean` is the 29 row table left after the record's known faults were corrected.
# the counting rule: a cell joins only if it carries a number. a real zero does.
figures = [r[5] for r in clean]
total = sum(figures)
print(total, len(figures))              # 1571700 29
print(round(total / len(figures), 2))   # 54196.55, and the 29 goes with it

Sums the 29 usable takings figures to 15,71,700 and divides by 29 to reach 54,196.55, printing the count beside the total so neither can travel alone. Python prints those totals without the commas an Indian reader writes by hand, so the same two numbers are Rs 15,71,700/- and Rs 54,196.55/-.

Printing the count beside the figure is not a courtesy, it is the figure being defined. The same total of Rs 15,71,700/- divided by 30 rather than 29 gives a different average, and both are correct arithmetic. The two averages differ only in the denominatorThe count sitting under an average, which decides what the average is actually an average of. Two people can divide the same total by two different counts and both be right, so the count has to be stated. used, and a reader can tell which one it was only if the count is stated. An average printed alone is a number somebody has to trust. An average printed with its count is a number somebody can check.

Try it out

The answer is printed as Rs 54,196.55/- over 29 usable stall months rather than as Rs 54,196.55/- on its own. Why does the count travel in the same sentence?

The error that gets made, and what it costs

An analyst groups the returns by stall_name. The names are readable and the codes are not. The grouping runs and returns 9 groups. Nobody looks at the 9. Every per stall figure computed after that point splits NB-03 in two, with 3 rows under the long spelling and 2 under the short one, and the market of 8 stalls is reported as a market of 9.

The cost is not a wrong total. A wrong total would honestly be easier to catch. All 32 rows are still there and none of them changed, and the total across all groups is exactly right. The shape of the answer is what went wrong. Both halves of NB-03 sit in the listing with plausible takings beside them and neither entry is the stall. A ranking of stalls by takings puts one of them mid table and the other near the bottom, and the stall that should have appeared once appears nowhere.

The fix costs one line and it goes before the figures rather than after them: the group count is read and compared against the number expected. 8 against 9 is a one second check. No amount of staring at the takings column would have produced it.

A LISTING WHOSE TOTAL IS RIGHT AND WHOSE SHAPE IS WRONGTAKINGS PER STALL, GROUPED BY stall_nameBansi Flour3 monthsRs 1,62,000/-Chandan Tea3 monthsRs 96,000/-Harit Green2 monthsRs 81,000/-Harit Greens2 monthsRs 77,700/-Ilaka Fruit4 monthsRs 1,42,000/-Kadamba Idli4 monthsRs 1,74,000/-Peeli Mithai4 monthsRs 1,82,000/-Peetal Utensils3 monthsRs 54,000/-Sundari Chaat4 monthsRs 6,03,000/-9 linesRs 15,71,700/-expected 8 lines,counted 9WHAT IT COSTSThe total is right to the rupee, sonothing looks broken. A market of 8stalls is reported as 9, and bothhalves of NB-03 read like real stallswith plausible takings beside them.
Grouping by the name produces a listing of 9 lines totalling Rs 15,71,700/-, which is the correct total in the wrong shape, and only the count of lines reveals it.
Financial Analyst Program Bootcamp — Fin Maverick

Which two counts does an analyst read after every single step?

The row count and the group count. Two counts, and no more. Both are cheap enough to run every time, and both are therefore the first two things dropped by anybody working at speed.

Picture a credit analyst at a lender with a borrower's monthly sales file in front of them, sent across to support a working capital limit. The file arrives with a year of rows. The analyst filters it to one product line and computes an average monthly sale. The row count says whether they still have the record they started with, and the group count says whether the column they filtered on means what they assumed it meant. If the file had 480 rows and the filter left 61 when the borrower described eleven months of trading, something is wrong and it is wrong now, not at the credit committee.

The same discipline is what a household does without naming it. The notes are counted at the bank counter before anyone leaves it, not out of distrust of the cashier, but because that is the last moment the count is cheap. Ten seconds at the counter or an argument tomorrow.

Two counts recorded at every step turn a set of results into a set of results with provenanceWhere a number started and what happened to it at each step, kept in writing. A figure somebody doubts can then be traced back to the cell it came out of.. A step whose row count was never read is a step nobody actually saw, so the trail behind that figure has a hole in it exactly there. The value is not in any single check. The value is that when a figure looks wrong three weeks later, the counts can be walked back through to find the step where the record stopped being the record, instead of starting the whole thing again.

One last thing is worth stating, and it explains why counting carries so much of the work. Every count above is read against one sentence about what a row is: one row of the Neelbagh record describes one stall in one month. The sentence naming what one row stands for is the grainOne line of this table stands for one thing, and this is the sentence naming that thing. Until it is on paper, nobody can say what a count of those lines has counted. of the table, and once it is written down, 9 groups from a column that should describe stalls is visibly wrong rather than merely surprising.

Try it out

A table has just been filtered and a figure is about to be computed from what is left. Which two counts are read first, and what does each protect against?

What is covered elsewhere

The three ways a record arrives, meaning a written down file, a request and its answer, and a stored table, are each covered separately. The effect of a matching key on the row count, when two tables are put side by side, is covered under joining two tables. So is putting an exact question to a stored table, and so is ordering the steps so an analysis can be run again from the top. The language itself, meaning syntax for its own sake, control flow, data structures, setup and packaging, is a separate subject covered elsewhere.

What sits behind each count in the Neelbagh record

What was countedWhat it isWhere it can be checked
The Neelbagh stall recordA made up office file of 32 rows and 8 columns, covering four months of takingsPrinted whole above, row by row
The Neelbagh day bookA made up running log of twelve notes the market office wrote beside those returnsNamed above
The counting rule for takingsThe test a takings cell has to pass before it is allowed into an averageWritten out beside every figure it produces
The five group countsHow many distinct values each of five columns holds across those 32 rowsListed in the ladder, and redrawn by the panel

The Neelbagh market, the Neelbagh stall record, the Neelbagh day book and the market office are invented.
Educational material. Not advice on any investment, tax, budget or market position.

Covered in this topic

Subtopics

DataFrame
← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.