Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quant Analyst · CoreTrack
1Quantitative Methods, Financial Data & Programming
iProbability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
iiStatistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
iiiCorrelation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
ivTime Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
vSimulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
viOptimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
viiModelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
viiiBacktesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
ixData Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
xProgramming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
xiQuantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity
2Stochastic Calculus & Derivative Pricing Theory
iProbability Foundations
The Probability SpaceRandom VectorsSigma-AlgebraExpectationSample Space and EventsDensity and Distribution FunctionsRisk-Neutral ProbabilityState Price Density vs…
iiStochastic Processes and Jumps
Properties of a Stochastic ProcessMartingaleBrownian Motion and Its PropertiesBrownian Motion vs Geometric…Stopping TimeThe Markov PropertyState VariablesTransition ProbabilityQuadratic VariationQuadratic Variation vs Ordinary…Submartingale and SupermartingaleMartingale RepresentationMarkov Process vs MartingaleOptional StoppingFiltrationJump ProcessesThe Poisson ProcessLevy ProcessesJump Diffusion
iiiIto Calculus
The Ito IntegralThe Ito Integral vs the Riemann IntegralInfinitesimals in Stochastic CalculusQuadratic CovariationIto's LemmaHow to Apply Ito's…The Infinitesimal GeneratorIto Calculus vs Ordinary Calculus
ivStochastic Differential Equations
Stochastic Differential EquationsStochastic Differential Equation vs…Drift and DiffusionStrong and Weak Solutions ComparedDiscretisationGeometric Brownian Motion
vPricing Theory and No-Arbitrage
No-ArbitrageGirsanov, Radon-Nikodym and Change…Physical and Risk-Neutral Measures…The Fundamental Theorems of…The Law of One PriceThe Pricing KernelDiscount Factors and Zero-Coupon PricesReplication vs HedgingComplete Market vs Incomplete MarketClearing Margin Architecture
viOption Pricing Theory
European and American OptionsMonte Carlo European OptionThe Black-Scholes PDEBlack Scholes and the GreeksThe Payoff FunctionThe Binomial ModelBinomial Option PricingDelta Hedging in TheoryBoundary, Initial and Terminal ConditionsThe Exercise BoundaryHow to Check Put-Call…
viiVolatility Models
Constant, Local and Stochastic…Vasicek Model vs CIR ModelThe Heston ModelThe SABR ModelThe Volatility ProcessImplied VolatilityVolatility Smile vs Skew vs Surface
viiiInterest Rate Models
Interest-Rate DerivativesMean ReversionThe Zero-Coupon BondThe Ornstein-Uhlenbeck ProcessThe Discount CurveZero RatesShort-Rate Model vs Market Model
ixNumerical Pricing
Closed Form and Numerical…Monte Carlo PricingEuler and Milstein Schemes ComparedTree MethodsFinite Difference MethodsNumerical Error and StabilityVariance Reduction
xCalibration and Model Risk
Model OverrideMarket Price and Model PriceCalibrationHow to Document a Pricing ModelThe Educational Illustration LabelMarket ConventionsModel Uncertainty and LimitationsBacktesting a Pricing ModelIdentifiabilityCalibrated ParametersThe Calibration Loss Function

Databases in Finance: Where Structured Data Sits

A database is a store where a table's shape is declared before any row enters: a name and a type for every column, and a key that says what one row is. Declare the Neelbagh returns on the stall and the month and the store accepts 31 rows of 32 and names the one it turns away.

Somebody hands an analyst a table and says the data is in the database now. The sentence sounds like a statement about a location, as though the rows had been moved from one shelf to another and are otherwise unchanged. A database is not a location at all. Something happened to those rows on the way in. Every row was held up against a shape that somebody wrote down first, and any row that did not match the shape did not get in. What that shape is, who writes it, and what it can and cannot see are the three things worth knowing about any store.

Every figure below is a count of rows, and every count was worked out by offering the same 32 rows to a declared shape and seeing what came back. A count of rows carries no as of date, and nothing in it changes with the day it is read.

What is already settled

Three things arrive here finished, and none of them gets argued again.

  • The Neelbagh stall record itself. An invented covered market, a market office that writes one line per stall per month across four months, 32 lines in eight columns, and the faults that record already carries. The faults were diagnosed in an earlier reading and are used here as material rather than discovered again.
  • The habit of counting rows at every step. A store that refuses a row is worth nothing at all unless a human being reads the number of refusals. Reading that number is a discipline rather than a feature.
  • Counting, and almost nothing else. Every figure below is a count of rows let in, a count of rows kept out, or one single average that exists only to show what a completely empty average looks like.

What is a database, and what does it declare?

The market office started with paper, so start there too. The clerk at Neelbagh used to keep a ruled ledger with the column headings printed at the head of every sheet: stall, licence, month, takings, pitch. The boxes were ruled and the headings were printed before a single figure was ever entered, and that ordering is the entire idea. There is only one box, so the clerk cannot write two takings figures for one stall in one month without visibly writing over something. The shape came first, and the shape is what refuses.

A database is that ledger with the ruling enforced by a machine. Before any row enters, three things are settled and written down. First, the columns and the order they sit in. Second, a declared type for each column. The type says what kind of value may sit in that column. Third, a key. The key states what one row of the table describes. The whole difference between a file and a store sits there: a file carries data and makes no promises about it, and a store refuses data that does not match the promises it was given.

THE THREE THINGS DECLARED BEFORE ANY ROW EXISTS The Neelbagh returns table, written out the way the market office would have to declare it. DECLARATION THREE, THE KEY: stall_id and month One row is one stall's return for one month. Anything else is not a row of this table. DECLARATION ONE: THE COLUMN, IN ORDER DECLARATION TWO: ITS DECLARED TYPE IN THE KEY? stall_id text, five characters YES stall_name text, as the office typed it no licence_no text, because it is a label no category text, one of five words no month whole number, 1 to 4 YES takings_rupees whole number of rupees no pitch_sqft whole number of square feet no filed_on_day whole number, a day of the month no The same eight columns delivered as a plain file declare none of the three. A first line of words is a courtesy from whoever typed it, and nothing more than that.
A store settles the columns and their order, a declared type for each of them, and a key naming what one row describes, and it settles all three before a single row is allowed to enter.

The key on that form is easy to skim past. Notice its job. A key is not a setting for making lookups quicker. A key is a sentence about meaning: one row is one stall's return for one month. Once that sentence is written down, a second row claiming to be the same stall's return for the same month is not extra information, it is a contradiction, and the store is entitled to say so. The two columns named in the key are the two columns that carry that sentence, and every other column in the table is just something the row happens to also say.

THREE QUESTIONS A FILE AND A STORED TABLE ANSWER DIFFERENTLY Same eight columns, same 32 rows of the Neelbagh returns. Only the promises differ. THE QUESTION A PLAIN FILE A STORED TABLE What it settles before a single row arrives Nothing at all. A first line of words is a courtesy, not a promise about what follows. The columns and their order, a type for each, and a key naming one row. What it does with a row that does not fit Takes it. There is nothing inside a file that could refuse a line of text. Turns it away, and writes a line saying which of the declarations it broke. What it says when something is wrong Nothing. It holds no view about its own contents and never will. Which row it refused and what refused it. Whether a figure is true, it cannot say.
A plain file carries the data and makes no promises about it, while a stored table refuses whatever does not match the promises it was handed before the rows arrived.
Try it out

A store settles three things before any row enters. Which set names all three?

What does a declared type do that a file does not?

A declared type says what kind of value a column is allowed to hold, and everything interesting about types happens in the one column where the answer is genuinely arguable. Seven of the Neelbagh columns are easy. A month is a whole number between 1 and 4. Takings are a whole number of rupees. A pitch is a whole number of square feet. A filing day is a whole number naming a day. Names and categories are text. Then there is licence_no, and licence_no is the one that decides whether the store is any help at all.

Try it out

Before reading on, commit to an answer. A licence number can be declared as text or as a whole number. Which declaration makes a meaningless question impossible, rather than merely unwise?

A licence number is an identifierA label with no job except naming a single thing, so that two of them are never confused. Digits are merely the alphabet somebody chose for it, which is why arithmetic on one answers nothing.. Its entire job is to point at one stall. The number happens to be written with digits, an accident of how the office issues them, and the digits carry no quantity. Licence 2743 is not larger than licence 2104 in any sense a person could use, only different. Declare that column as text and the store will sort it, group by it and match on it, and will offer no way at all to add two of them together. Declare it as a whole number and the store will cheerfully do arithmetic on it.

Here is what that costs, on this record, exactly. The eight stalls still trading carry licences 2104, 2216, 2318, 2405, 2477, 2530, 2618 and 2743. Added, they come to 19,411. Divided by the eight stalls, that is 2426.375. The average of eight licence numbers is not money, not a size, not a rate and not anything at all, and nothing in the store objected to producing it. No warning appeared and no line went into any report. The answer came back to three decimal places, exactly the sort of precision that makes a person trust a number. A declared type is not paperwork. The type is how a meaningless question is made impossible to ask rather than merely unwise to ask.

ONE COLUMN, TWO DECLARATIONS, AND ONE EMPTY ANSWER licence_no, for the eight Neelbagh stalls that were still trading in month four. DECLARED AS TEXT: A LABEL DECLARED AS A WHOLE NUMBER 2104 2216 2318 2405 2477 2530 2618 2743 Kadamba Idli Chandan Tea Harit Greens Peetal Utensils Bansi Flour Ilaka Fruit Sundari Chaat Peeli Mithai AVERAGE OF THE COLUMN Eight labels have no average, so the question is refused at the point of being typed rather than answered. 2104 2216 2318 2405 2477 2530 2618 2743 19,411 added over the eight stalls 2426.375 The unit carried by that figure is: none.
Declared as a whole number the licence column averages to 2426.375 without a murmur from the store, and that figure is not money, not a size and not anything at all.
Try it out

The store returns 2426.375 for the average of the licence numbers and raises nothing. What went wrong, and where was it preventable?

Breaking Into Quants Bootcamp — Fin Maverick

What is a key, and what does one turn away?

A key is the declaration of the grainThe sentence saying what one line of a table stands for. Until somebody writes that sentence down, every count taken off the table is a guess about what has just been counted.. The declaration is one sentence, and the sentence takes the form: one row of this table is one of these. One row is one stall's return for one month. One row is one note written on one day. Once that sentence exists, the store has a rule it can apply mechanically, and the rule is that no two rows may claim to be the same one.

The market office's 32 rows are therefore offered to a store five times, changing nothing except that sentence, and what gets in is counted each time. One conventionA decision about how something is to be counted, fixed in advance and recorded, so that the resulting figure does not depend on who happened to produce it. has to be stated before the counts mean anything: the loader here reads the file from the top and keeps the first row carrying a given combination, turning away any later row that repeats it. First one wins. Say it out loud. The convention matters enormously in a moment.

The declared key, in wordsWhat it claims one row isAcceptedTurned away
stall_idOne row is one stall824
stall_nameOne row is one stall name923
stall_id and monthOne row is one stall's return for one month311
stall_name and monthOne row is one stall name's return for one month311
stall_id, month and filed_on_dayOne row is one row320

Read the ladder from the top and it looks like a story of progress: 8, then 9, then 31, then 31, then everything. It is not. The first declaration is a real claim about the world and it is simply wrong for this table, so the store throws out 24 rows of the 32 offered and is right to. The second declaration is the same wrong claim asked in a worse language, and the extra row it lets in is not a ninth stall, it is a second spelling. The third and fourth are the true grain of this table. The fifth is not a key at all: a declaration wide enough to accept every row offered is a row number wearing a key's clothes, and adopting one is how a store is quietly made to agree with a file instead of checking it.

FIVE DECLARED KEYS, ONE FILE OF 32 ROWS, OFFERED FIVE TIMES Each full track is the 32 rows handed over. The pale segment is what the store refused. OUT OF 32 OFFERED stall_id 8 accepted 24 turned away stall_name 9 accepted 23 turned away stall_id and month 31 accepted 1 turned away stall_name and month 31 accepted 1 turned away stall_id, month and filed_on_day 32 accepted 0 turned away At the widest declaration the refused segment has vanished, and so has the check.
The five declared keys accept 8, 9, 31, 31 and 32 rows of the 32 offered, and the widest of the five refuses nothing whatsoever.
Try it out

Somebody widens the key to stall_id, month and filed_on_day, watches all 32 rows load, and calls it a clean run. What is the reply?

Try it out

The panel below widens the key one column at a time. Before it runs: what does the count of turned away rows do across the four steps?

Play with it

Move the declared key and watch the store change its mind about what a row is.

One control, and it moves one thing: which of the five declarations the store was given before the load began. Everything else is recomputed from the same 32 rows. The grid is the record drawn as it actually is, eight stalls across and four months down, so a solid tile is a row the store let in and a pale tile is a row it turned away. The one stall month the office filed twice sits as a split tile, and the one stall month with no row at all sits as a dashed outline. The bar underneath is the same count as a length, and the lines below it name what was refused. The panel opens on the stall and the month, a declaration that accepts 31 rows of the 32 offered.

Narroweststall_id and monthWidest
Rows accepted
31
Rows turned away
1
Rows offered
32

Educational illustration. The Neelbagh market, its stall record and its day book were made up for teaching. The 32 rows offered are the file exactly as the market office handed it over, the loader keeps the first row carrying a combination and turns away any later repeat, and a declared key checks the grain and nothing else. Every quantity in the panel is a count of rows, so no quantity is rounded and none can fall below zero.

Try it out

Declared on the stall and the month, the store refuses exactly one row. Before reading on: which fault does it catch, and which will it certainly miss?

What does a key catch, and what does it walk straight past?

Declared on the stall and the month, the store refuses exactly one row of the 32 and it names which one. NB-03, month 2. The market office filed a return for Harit Greens in month 2 on day 5 reading Rs 36,400/-, and filed a second one for the same stall and the same month on day 19 reading Rs 39,700/-. In the plain file those are two lines that look perfectly ordinary and nothing anywhere says one contradicts the other. In the store they are one contradiction with a name attached to it. A silent fault has been turned loud, and turning a silent fault loud is the single most valuable thing a declared shape does.

Now look closely at which of the two the store kept. The answer is uncomfortable, and it is the reason a refusal is a beginning rather than an ending. The loader keeps the first row carrying the combination, and the first row here is the one filed on day 5. The day book is the market office's running log of what actually happened, and against day 19 it records that the return was revised and the first figure withdrawn. So the store has kept the withdrawn figure and turned away the corrected one. The store did nothing wrong: it was asked whether two rows claim to be the same stall month, and it answered. Nobody asked it which of the two figures is true, and nothing about it could ever answer that.

THE ONE ROW THE STORE REFUSES, PRINTED IN FULL Key declared as stall_id and month. First row carrying a combination wins. 32 rows offered 31 accepted into the table 1 turned away, and named STALL_ID STALL_NAME LICENCE_NO CATEGORY MONTH TAKINGS_RUPEES PITCH_SQFT FILED_ON_DAY KEPT, BECAUSE IT WAS OFFERED FIRST NB-03 Harit Greens 2318 Vegetables 2 36400 100 5 TURNED AWAY, NAMED A REPEAT OF THE STALL MONTH ABOVE NB-03 Harit Greens 2318 Vegetables 2 39700 100 19 The row the store kept is the one the day book records as withdrawn on day 19. A declared key can see that a stall month arrived twice. It has no way of seeing which of the two figures the market office intended to stand.
Declared on the stall and the month, the store accepts 31 rows and names the one it refuses, which is NB-03 in month 2 filed a second time on day 19.

Now the other side of the boundary, and it is the more important side. The very same store, running the very same declaration, accepts both spellings of NB-03 without a murmur. The market office wrote Harit Greens on the returns for months 1 and 2 and Harit Green on the returns for months 3 and 4, so the stored returns table holds 9 distinct stall names for 8 stalls. Not one line of any load report mentions it. Why would it? The key names stall_id, the stall identifier is NB-03 on all four rows, the grain is intact, and stall_name is simply a column the row happens to also carry.

Think of the ruled ledger again. The printed boxes stop the clerk writing two figures where one belongs. The ruling was never a claim about spelling, so the printed boxes do nothing whatsoever about the clerk spelling a stall's name differently in April from how it was spelled in January. A key checks the grain and nothing else, and a fault that leaves the grain intact walks straight past it. Sort the record's nine known faults by whether they break the declared grain and whether they break a declared type, and seven of the nine sit in the quadrant a store cannot see at all.

THE NINE KNOWN FAULTS, AND HOW MANY A DECLARED SHAPE CAN SEE Key declared as stall_id and month. Types declared as in the form further up this reading. BREAKS THE DECLARED GRAIN GRAIN LEFT INTACT EVERY VALUE FITS ITS DECLARED TYPE A VALUE BREAKS A TYPE NB-03 month 2 was filed twice The only fault of the nine that the key refuses, and it is refused by name in the load report. Nothing sits here. No fault on this record breaks both at once. NB-05 month 2 has an empty takings box Refused only if the column was declared as a number that must be present. NB-04 month 3 has no row at all NB-02 month 4 reads 99999, an office code NB-03 is spelled two different ways NB-08 month 3 carries an extra digit NB-05 pitch reads 14 in a square feet box licence_no is a label held as a whole number Two closed stalls were deleted without a trace SEVEN OF THE NINE SIT HERE
A key checks the grain and a type checks the values, so seven of the nine known faults leave both intact and walk past the store entirely.
Try it out

The stored returns table holds 9 distinct stall names for 8 stalls and no load report mentioned it. Why not, and what would have caught it?

AI For Finance Bootcamp — Fin Maverick

Why does the office keep two tables rather than one?

The market office keeps a second thing besides the returns, and it is the running log the clerk writes as events happen: the day book. Twelve notes, one line per note, each carrying a stall, the stall's name as the office wrote it, a month, the day the note was made and the note itself. The twelve notes touch 8 stalls and 11 stall months. Exactly one stall month carries two of them: NB-03 in month 2, precisely the stall month the returns table argued about.

The obvious question is why the office does not just bolt the notes onto the returns and keep one wide table. The answer is one word long and it is the same word the whole of this reading has been circling. The two tables have different grains, and a table can declare only one. One row of the returns is one stall's return for one month. One row of the day book is one note written on one day. Merging them means the very first stall month with two notes forces a choice: either that stall month becomes two rows, and the returns table quietly stops being one row per stall month, or one of the two notes gets thrown away.

Think of a household with a bank passbook and a diary. The passbook has one line per transaction. The diary has one line per day, and some days carry three entries and some carry none. Nobody in the household tries to keep one book, and not because they lack the stationery. The household keeps two because the two books answer different questions, and each book only stays trustworthy while it is answering its own. Putting the two side by side, so a note lands against the return it explains, is a genuinely separate operation with a genuinely separate way of going wrong, and it is covered further on rather than here.

ONE STORE, TWO TABLES, TWO DIFFERENT GRAINS Both invented for teaching, and both kept by the same market office. THE RETURNS TABLE THE DAY BOOK TABLE stall_id, stall_name, licence_no, category, month, takings_rupees, pitch_sqft, filed_on_day stall_id, stall_name, month, note_day, note and 28 more lines like them and 8 more lines like them 32 rows 12 rows Grain: one row is one stall's return for one month. Grain: one row is one note written on one day. They stay apart because a table can declare only one grain, not because of shortage of room.
The returns table holds one row per stall month and the day book table holds one row per note, and no table can declare more than one grain at a time.
Try it out

In one word, why does the market office keep the returns and the day book as two tables, and what is lost by merging them?

Reading an Option Payoff — free micro-course from Fin Maverick

What a store still does not tell

Three things, and each of them is a reason not to mistake a declared shape for a guarantee.

A store does not say whether a figure is right. A store says whether a figure fits. NB-08's month 3 return reads Rs 4,80,000/- where the day book records a slip reading Rs 48,000/-, and no declaration in the world catches that. A whole number of rupees is exactly what the column was promised and exactly what it got. A shape is a filter, not a judgement.

A store does not report what it refused unless somebody goes and looks. The refusal count is written into a load report by a machine that has no idea whether a human being will ever open it, and the entire value of the store collapses to nothing at the moment that report stops being read.

And it carries no provenanceA figure's history: the place it started and each thing done to it since. Nothing inside the figure carries that history, so it survives only where somebody kept notes while the work was happening. of its own. The store knows what it currently holds. The store does not know where those values came from, who typed them, or which of them was corrected on the way. A store will also hand rows back in whatever order suits it, so row orderWhere a line happens to sit in the stack. Shuffle the same lines and the order changes while the record does not, which is why nothing should ever be read off position alone. out of a store is a fact about the store rather than about the market. A declared shape is a filter and not a judgement, and treating a clean load as a verdict on the data is where the expensive mistakes begin.

How this goes wrong: the refusal nobody reads, and the key that was widened until it stopped complaining

A team loads the Neelbagh returns into a store declared on the stall and the month. The loader refuses one row and writes the refusal into the load report, correctly, by name. Nobody opens the report. The load is recorded as done and the work moves on.

Six weeks later somebody notices the loader keeps flagging something and finds the refusal line. The fix chosen is to widen the key to include the filing day. Both NB-03 rows are then legitimately different and the load will run clean. And the load does run clean. All 32 rows go in, the report says nothing, and everybody is satisfied.

The fault has not been fixed. The fault has been declared away. The stall month for NB-03 in month 2 is now in the table twice, contributing Rs 36,400/- and Rs 39,700/- to every total anyone computes. The two together make Rs 76,100/- for a month the market office says took Rs 39,700/-. The withdrawn return now sits permanently alongside the revised one, no part of the system will ever mention it again, and the store has stopped checking the one thing it was checking before. The old load report at least argued with the file. The new one agrees with it.

The fix is not technical, and that is why it keeps not getting done. Read the refusal count on every single load, and treat a widened key as what it is: a change in what a row of the table means, not a change of setting.

THE SAME RECORD, TWO LOAD REPORTS, SIX WEEKS APART Nothing in the market office's file changed between them. Only the declaration did. LOAD REPORT, THE FIRST RUN key: stall_id and month LOAD REPORT, SIX WEEKS LATER key: stall_id, month and filed_on_day rows offered 32 rows accepted 31 rows turned away 1 NB-03 month 2, filed on day 19: this stall month is already present NOBODY OPENED THIS REPORT rows offered 32 rows accepted 32 rows turned away 0 no exceptions EVERYBODY READ THIS ONE WHAT CHANGED IN THE RECORD BETWEEN THE TWO RUNS: NOTHING NB-03 in month 2 is now stored twice and contributes Rs 36,400/- and Rs 39,700/- to every total, so Rs 76,100/- stands for a month the market office recorded as Rs 39,700/-.
Widening a key until the load runs clean does not repair a stall month filed twice, it stops anything in the system from ever mentioning it again.

Four questions to ask of any stored table somebody hands over

Stored tables arrive far more often than they are built, usually with the words it is all in the database and no further detail. The four questions below take about a minute and they are the whole of the job.

  1. What is the grain, in one sentence? If the person handing it over cannot finish the sentence one row of this table is one, then nobody has decided, and every count taken off it will be counting something nobody has named.
  2. What is the key, and how many rows did the last load refuse? Both halves matter. A key with no refusal count attached is a claim with no evidence, and a refusal count of zero on a record anybody has ever typed by hand deserves a second look rather than relief.
  3. Which columns are identifiers stored as numbers? Ask for them by name. Every one of them is an invitation to produce a figure like 2426.375, and the figure will look every bit as respectable as a real one.
  4. What is the convention behind each figure computed from it, with its denominatorWhatever an average was divided by. Change it and the same total becomes a different average without any arithmetic being wrong anywhere. named? An average of takings over the stall months that carry a usable figure and an average over every row in the table are different questions, and the store will answer both without ever mentioning that they differ.

A lender assessing a small market operator, an analyst rebuilding a set of monthly figures and a household reconciling a passbook against a diary are all doing the same four checks in different clothes. A stored table handed over without its refusal count is a table whose faults have been declared away rather than found, and the person handing it over usually does not know that.

Try it out

Somebody hands over a stored table and says it loaded cleanly. Which two questions come first, and what does each protect against?

Where this reading stops. The written down path a record travels from raw source to an analysis ready table, and the shape of a request and the answer that comes back, were both settled earlier and are not repeated here. Working with the table once it has been pulled into memory, and grouping it to get one figure per stall, comes further on. So does setting one table beside another so that a note lands against the return it explains, and that is where the most expensive damage in this subject sits. So does putting a precise question to a stored table, and so do the ordered cells that let a whole analysis be run again from the top. Administering a store, making it answer faster and backing it up are not a finance analyst's work at all.

A store says a figure fits, never that it is right. See what remains.

What was consulted to write this, and why is nothing the honest answer?

No outside document was read, and that is a finding about the subject rather than a gap in the work. The count is a consequence of two things a reader can hold in one hand: a set of rows, and a sentence saying what one row is. Counting how many rows a declared shape lets in therefore requires nobody's authority. A second person given those two will land on 8, 9, 31, 31 and 32 without needing to trust anyone. The table below says where each ingredient of that arithmetic was settled and how it can be tested here.

What a figure above rests onWhere it was settledHow to test it against this reading
The 32 rows, the eight columns and the nine faultsThe Neelbagh stall record, fixed in an earlier reading on record quality and reused here without a single changeCount the tiles in the panel. Eight stalls across, four months down, one stall month split in two and one missing
The five accepted counts of 8, 9, 31, 31 and 32Counting distinct combinations of the named columns across those same 32 rowsTake any declaration, list the combinations it names, strike out the repeats and count what is left
The empty average of 2426.375The eight licence numbers from 2104 to 2743, added to 19,411 and divided by the eight stallsAdd the eight values printed in the figure. The division is exact and stops after three places
The twelve day book notes over 8 stalls and 11 stall monthsThe market office's running log, the one object added to the record beyond the returnsCount the notes named in the two table figure and check that only NB-03 in month 2 carries two
Any rule, rate, threshold, filing period or published standardNone is stated anywhere above. A count of rows in a made up file needs noneThere is no outside value to confirm at source. Such a value met elsewhere is taken from whoever issues it

The Neelbagh covered market, the office that runs it, the stall record it keeps, the day book beside it and every stall named in them are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.