Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quant Analyst · CoreTrack
1Quantitative Methods, Financial Data & Programming
iProbability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
iiStatistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
iiiCorrelation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
ivTime Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
vSimulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
viOptimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
viiModelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
viiiBacktesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
ixData Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
xProgramming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
xiQuantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity
2Stochastic Calculus & Derivative Pricing Theory
iProbability Foundations
The Probability SpaceRandom VectorsSigma-AlgebraExpectationSample Space and EventsDensity and Distribution FunctionsRisk-Neutral ProbabilityState Price Density vs…
iiStochastic Processes and Jumps
Properties of a Stochastic ProcessMartingaleBrownian Motion and Its PropertiesBrownian Motion vs Geometric…Stopping TimeThe Markov PropertyState VariablesTransition ProbabilityQuadratic VariationQuadratic Variation vs Ordinary…Submartingale and SupermartingaleMartingale RepresentationMarkov Process vs MartingaleOptional StoppingFiltrationJump ProcessesThe Poisson ProcessLevy ProcessesJump Diffusion
iiiIto Calculus
The Ito IntegralThe Ito Integral vs the Riemann IntegralInfinitesimals in Stochastic CalculusQuadratic CovariationIto's LemmaHow to Apply Ito's…The Infinitesimal GeneratorIto Calculus vs Ordinary Calculus
ivStochastic Differential Equations
Stochastic Differential EquationsStochastic Differential Equation vs…Drift and DiffusionStrong and Weak Solutions ComparedDiscretisationGeometric Brownian Motion
vPricing Theory and No-Arbitrage
No-ArbitrageGirsanov, Radon-Nikodym and Change…Physical and Risk-Neutral Measures…The Fundamental Theorems of…The Law of One PriceThe Pricing KernelDiscount Factors and Zero-Coupon PricesReplication vs HedgingComplete Market vs Incomplete MarketClearing Margin Architecture
viOption Pricing Theory
European and American OptionsMonte Carlo European OptionThe Black-Scholes PDEBlack Scholes and the GreeksThe Payoff FunctionThe Binomial ModelBinomial Option PricingDelta Hedging in TheoryBoundary, Initial and Terminal ConditionsThe Exercise BoundaryHow to Check Put-Call…
viiVolatility Models
Constant, Local and Stochastic…Vasicek Model vs CIR ModelThe Heston ModelThe SABR ModelThe Volatility ProcessImplied VolatilityVolatility Smile vs Skew vs Surface
viiiInterest Rate Models
Interest-Rate DerivativesMean ReversionThe Zero-Coupon BondThe Ornstein-Uhlenbeck ProcessThe Discount CurveZero RatesShort-Rate Model vs Market Model
ixNumerical Pricing
Closed Form and Numerical…Monte Carlo PricingEuler and Milstein Schemes ComparedTree MethodsFinite Difference MethodsNumerical Error and StabilityVariance Reduction
xCalibration and Model Risk
Model OverrideMarket Price and Model PriceCalibrationHow to Document a Pricing ModelThe Educational Illustration LabelMarket ConventionsModel Uncertainty and LimitationsBacktesting a Pricing ModelIdentifiabilityCalibrated ParametersThe Calibration Loss Function

The Analysis Workflow: From Question to Reproducible Answer, and the Notebook That Stays Auditable

An analysis workflow is the ordered set of steps that turns one question into one answer somebody else can rebuild. Run every cell on the Neelbagh returns and the figure is Rs 54,196.55/-. Skip one cleaning cell of four and it reads Rs 54,196.55/-, Rs 55,723.30/-, Rs 53,603.33/- or Rs 69,093.10/-, and the notebook prints no warning at all.

Every rupee figure below was produced by running the cleaning cells over the thirty two rows the market office typed. The rows themselves are printed below, so the arithmetic can be rerun against them.

The last of eight guides adds no seventh tool. Everything it uses is already in place: the written down path from raw source to a table fit to compute on; the two doorways a record can come through; where a stored table sits, what shape it declares and what a key is for; the table an analyst works in and what a grouping check tells; that putting two tables side by side changes the row count before it changes anything else; and that every answer has a count sitting underneath it. A correct method run in the wrong order, or with one step quietly missing, moves the answer by up to Rs 14,896.55/- on a record of thirty two rows.

Three things hold this guide up. The first is the Neelbagh stall record, the market office file carrying one row for each stall in each of four months, thirty two rows in all, with its faults already in it and already named in the guides before this one. The second is the day book, the office running log of twelve notes. The day book is the only thing that settles which of two identical looking figures in month three is real. The third is plain tallying. Nothing below is more than a tally of surviving rows, a tally of cells that actually ran, or a single average with its denominatorThe number sitting under a division. It decides what the answer is an average of, which is why it gets named in the same breath as the figure it produced. named beside it.

What is an analysis workflow, and what does it produce?

A workflow is not a mood and it is not a habit of tidiness. A workflow is an ordered set of steps with a stated start and a stated end. The start is one question in one sentence, and the end is one answer that somebody who was not in the room can rebuild from the same materials. On this record the question is fixed and short: what were the average takings per stall month at the Neelbagh market across months one to four?

Five cells carry that question to its answer, and the order is the whole of the method. Cell one writes the question down, together with the grainThe unit that one row of a table stands for. Here one row is one stall in one month, which is why thirty two rows is not thirty two stalls and never was. of the table it is being asked of. Cell two reads the source exactly as the office handed it over and prints how many rows arrived. Cell three cleans, in four rules, printing the row count it is left with after each one. Cell four computes the figure and prints the count underneath it. Cell five writes the answer out beside the question and the counting conventionA counting choice written down in advance so a second person can repeat it exactly. Without one, two careful people can both be right and still disagree about the same file. it was reached under.

The product of a workflow is not the figure. The product is the pair: a figure and the path that produced it. A figure without its path cannot be defended and, worse, cannot be corrected. If somebody finds a fault in the record next week, an analyst holding only Rs 54,196.55/- has to start again from nothing. An analyst holding the path changes one cell and reruns.

The everyday version is a household one. Two homes run on one salary each. The first files every receipt in the order it was spent, in one envelope for each month. The second writes the monthly total on the front of the envelope and throws the receipts away. Both can say what they spent last March. Only the first home can answer the real question. The real question is never the total, but why March was higher than February, and whether the school fee has already been counted.

FIVE CELLS, TOP TO BOTTOM, EACH PRINTING WHAT IT LEFT BEHIND ROWS AFTER IT RAN CELL 1 THE QUESTION average takings per stall month, and one row is one stall in one month nothing read yet CELL 2 THE SOURCE AS HANDED OVER read, never edited, and the arriving row count printed at once 32 CELL 3 FOUR CLEANING RULES, IN THIS ORDER drop the row whose takings box was left empty drop the row carrying the office code for no return received drop the withdrawn first return, keeping the revised one correct the typed takings the day book says should read Rs 48,000/- 31 30 29 29 CELL 4 THE FIGURE, WITH ITS COUNT Rs 15,71,700/- over 29 usable takings cells Rs 54,196.55/- 29 29 of them usable CELL 5 THE ANSWER, WRITTEN OUT the figure, the question it answers, and the convention it was reached under 29, unchanged the right hand column is the whole audit trail: a second person checks the sequence from these numbers and nothing else
Five cells run in sequence and each prints the row count it ended with, so a second person can check the sequence from the printed output alone.
Try it out

The workflow has finished and the figure is on the screen. What has the workflow actually produced?

What is a research notebook, and what is it not?

A research notebook is a record of what was actually run, in the sequence it ran, with the output of each step kept beside the step that produced it. The definition is that short, and every word in it is load bearing. Actually run, so a cell that was meant to run and did not is not in the record. In sequence, so the order is part of the record rather than a detail. Output kept beside it, so nobody has to take the author's word for what a step returned.

And here is what it is not: a research notebook is not a scratchpad. A scratchpad is where things get tried, broken, jumped back up to, fixed by hand and tried again, and there is nothing wrong with having one. The fault is letting the scratchpad be the record. A cell that was edited after it ran, or run out of sequence, has quietly turned the notebook into fiction, and the cruel part is that the fiction looks complete. Every cell has a number under it. Every number was really printed by something. The trouble is that the cell that printed it no longer exists anywhere in the file.

The discipline that catches this is one line long, and it is the reason the right hand column of the figure above exists. Every cell prints the row count it ended with. Now the sequence is checkable from the printed output alone: thirty two, thirty one, thirty, twenty nine, twenty nine. If the printed counts do not descend in that order, the cells did not run in that order, and that is visible without asking anybody anything.

ONE NOTEBOOK, TWO HISTORIES. TIME RUNS LEFT TO RIGHT. RUN STRAIGHT THROUGH cell 1 cell 2 cell 3 cell 4 cell 5 prints 32 prints 31, 30, 29, 29 prints 29 the printed record and what happened are the same thing, so the sequence can be checked from the file itself RUN OUT OF SEQUENCE, WITH ONE CELL EDITED AFTERWARDS cell 1 cell 2 cell 4 cell 3 cell 2 edited from here the printed record stops matching what happened every cell still shows a number, so the file looks finished and nothing on it is a warning the two files are indistinguishable on screen; only the printed row counts, in the order they appear, tell them apart
A cell edited after it ran, or run out of sequence, turns the printed record into fiction while leaving it looking complete.
# CELL 1  the question, written before anything is read
question = "average takings per stall month, Neelbagh, months 1 to 4"
grain    = "one row is one stall in one month"

# CELL 2  the source exactly as the office handed it over
rows = office_file_as_handed_over()      # the 32 rows shown in this guide
print("read as handed over:", len(rows))     # 32

# CELL 3  four rules, each printing what it left behind
rows = drop_empty_takings(rows);      print("after the blank drop:", len(rows))      # 31
rows = drop_office_code(rows);        print("after the code drop:", len(rows))       # 30
rows = drop_withdrawn_return(rows);   print("after the duplicate drop:", len(rows))  # 29
rows = correct_typed_takings(rows);   print("after the correction:", len(rows))      # 29

Cells one to three of the notebook. The three cells read the thirty two rows the market office typed, apply the four cleaning rules in order, and print the row count each one ended with.

# CELL 4  the figure, and the count that sits underneath it
usable = [r.takings for r in rows if r.takings is not None]
print(len(usable), sum(usable))          # 29   1571700
print(sum(usable) / len(usable))         # 54196.551724...

# CELL 5  the answer, written out beside what it answers
write_answer(question, grain,
             convention="takings cells carrying a usable number only",
             figure=sum(usable) / len(usable))

Cells four and five. The count is printed before the figure rather than after it, so the denominator is never something a reader has to go looking for.

Try it out

A notebook arrives from somebody else. Every cell shows an output and the file looks finished. One cell, it emerges later, was edited after it ran. What has quietly happened, and what one line of discipline would have shown it?

Breaking Into Quants Bootcamp — Fin Maverick

What does skipping one cleaning cell do to the answer?

Now the part of this guide that nothing before it could have taught. The notebook has four cleaning rules. Each one in turn is switched off, everything else runs, and the result is read. Four skips, and the record is the same record every time.

Skip the blank drop and the answer holds at Rs 54,196.55/- exactly. The answer moves by nothing. The empty takings box was never in the count in the first place. The row survives, so thirty rows go into the calculation instead of twenty nine, but the convention counts takings cells carrying a usable number, and an empty cell has never been one.

Skip the office code drop and the answer reads Rs 55,723.30/-, up Rs 1,526.75/-. The row carrying the office code for no return received now counts as though ninety nine thousand nine hundred and ninety nine rupees of takings had been declared. Skip the withdrawn return drop and the answer reads Rs 53,603.33/-, down Rs 593.22/-. The first NB-03 return for month two was withdrawn and revised, and leaving it in counts one stall month twice at the lower of the two figures. Skip the typing error correction and the answer reads Rs 69,093.10/-, up Rs 14,896.55/-. A stall that took Rs 48,000/- in month three is being read as having taken Rs 4,80,000/-.

One notebook, five settings, the same thirty two rows. Rows means rows surviving into cell four; usable means takings cells carrying a number.
What ranRowsUsableThe answerThe move
every cell, nothing skipped2929Rs 54,196.55/-the reference
skip the blank drop3029Rs 54,196.55/-no change at all
skip the office code drop3030Rs 55,723.30/-up Rs 1,526.75/-
skip the withdrawn return drop3030Rs 53,603.33/-down Rs 593.22/-
skip the typing error correction2929Rs 69,093.10/-up Rs 14,896.55/-

The size of the largest move is not what that table shows. The table shows that one skip moves the answer not at all, one moves it down and two move it up, so the direction of the error cannot be read off which cell was skipped. A reader who arrived expecting an unfinished analysis to read high has already been wrong once on this record, at the withdrawn return. A reader who then decides unfinished analyses read low is wrong twice over. No rule of thumb is available, and the sequence in the notebook therefore has to be printed rather than remembered.

FOUR SKIPS AGAINST ONE FIXED LINE the line is the fully run answer, Rs 54,196.55/-; the vertical scale starts at Rs 50,000/- rather than at zero 50,000 60,000 70,000 FULLY RUN 54,196.55 55,723.30 53,603.33 69,093.10 skip the blank no change skip the office code up Rs 1,526.75/- skip the withdrawn return down Rs 593.22/- skip the typing error up Rs 14,896.55/- THE SAME THREE NEAR SETTINGS, MAGNIFIED: Rs 53,400/- TO Rs 56,000/- 53,400 56,000 53,603.33 54,196.55 55,723.30 the typing error setting is off this strip, further to the right
Skipping one cell of four gives Rs 54,196.55/-, Rs 55,723.30/-, Rs 53,603.33/- and Rs 69,093.10/-, moving nothing once, down once and up twice.
Try it out

The control below switches off one cleaning cell at a time and shows the answer. Before it is touched: does an unfinished run always move the answer away from Rs 54,196.55/- in the same direction?

Play with it

Switch off one cleaning cell and watch where the answer lands

One variable moves: which single cleaning cell the notebook skips. Everything else runs. The default is the worked example, every cell run, Rs 54,196.55/- over twenty nine usable figures. Pinning a reading holds that position on both strips as a dashed line while the control moves.

nothing pinned yet
THE FOUR CLEANING RULES, AND WHICH ONE THIS RUN SKIPPED every rule ran, so the notebook is the worked example drop the blank 32 to 31 drop the office code 31 to 30 drop the withdrawn return 30 to 29 correct the typed takings 29 to 29 THE NEAR STRIP, Rs 53,400/- TO Rs 56,000/- 53,400 54,000 55,000 56,000 fully run, Rs 54,196.55/- Rs 54,196.55/- off this strip to the right THE FAR STRIP, Rs 68,900/- TO Rs 69,200/- 68,900 69,000 69,100 69,200 the untouched file, Rs 69,035.45/- Rs 69,093.10/- this run lands down on the near strip both strips share one direction: further right is a larger figure, further left is a smaller one the two strips are magnified differently, so a step on the far strip covers far less money than the same step on the near one
Rows surviving
29
Usable figures
29
The answer
Rs 54,196.55/-
Move from fully run
no change
Gap to the pin
not pinned
Educational illustration. The cleaning rules and their order are the ones settled in the opening walkthrough of this sequence. The convention counts only takings cells carrying a usable number, and the whole panel recomputes from the thirty two rows the market office typed.
Try it out

Skipping the blank drop moves the answer by nothing at all. Does that make the cell pointless?

Try it out

The notebook that skipped the typing error correction reads Rs 69,093.10/-. The file nobody cleaned at all reads Rs 69,035.45/-. Which of the two sits closer to Rs 54,196.55/-?

AI For Finance Bootcamp — Fin Maverick

Can a half cleaned answer be worse than an uncleaned one?

Set the two side by side. The notebook that ran three of its four cleaning rules, missing only the typing error correction, lands on Rs 69,093.10/-. The file that nobody touched at all reads Rs 69,035.45/-. The answer, when everything runs, is Rs 54,196.55/-. The half cleaned notebook therefore sits Rs 14,896.55/- from the answer, and the untouched file Rs 14,838.90/- from it.

The half cleaned notebook has ended further from the answer than the file it started with, by Rs 57.65/-. Three rules of four ran, and the result went backwards. The result is worth sitting with, and it kills an intuition almost everybody carries into this work. Cleaning is supposed to walk towards the truth one step at a time. A partly cleaned figure is supposed to be a partly correct figure. It is not. Cleaning is not a distance being closed. Each rule moves the figure in whatever direction that particular fault happened to push it, and three rules that pushed one way can leave the figure past where it began once the fourth, the one that mattered most, never ran.

One more thing about that figure, and it matters for the sequence rather than for the arithmetic. Rs 69,093.10/- is not a new result. The figure is the same computation the opening walkthrough reaches at its third stop, the state of the record with everything cleaned except that one typed cell. The opening walkthrough runs the pipeline forward and arrives there; this guide switches one cell off and arrives there. One computation, named the same way from both directions, and not two findings that happen to agree.

THREE READINGS ON ONE LINE, AND THE TWO THAT LOOK IDENTICAL further to the right is a larger figure, and on this line larger means further from the answer 53,000 70,000 Rs 54,196.55/- the answer, every cell run two marks, Rs 57.65/- apart and they cannot be told apart here about fourteen thousand eight hundred rupees of distance, whichever of the two is taken THE SAME TWO MARKS, MAGNIFIED: Rs 68,980/- TO Rs 69,150/- towards the answer further from the answer 68,980 69,150 Rs 69,035.45/- the untouched file Rs 69,093.10/- three of four rules ran Rs 57.65/- the wrong way
The half cleaned notebook lands on Rs 69,093.10/- while the untouched file reads Rs 69,035.45/-, so half cleaned sits further from the answer than not cleaned.

How to reproduce a financial analysis from source data: what does a second person actually need?

Six things, and the list closes. Not seven, and none of the six is optional. The closed list is what makes it a usable test rather than a wish.

  1. The source exactly as it was handed over. Not a cleaned copy. The thirty two rows the market office typed, faults and all.
  2. The question in one sentence, with the grain of the table it is asked of. Average takings per stall month, asked of a table where one row is one stall in one month.
  3. The cells, in the sequence they ran. Not the cells as they now sit on the screen. The two can be different things entirely.
  4. The convention, with its denominator named. Takings cells carrying a usable number, so twenty nine of them, not thirty and not thirty two.
  5. The row count at every step. Thirty two, thirty one, thirty, twenty nine, twenty nine.
  6. The answer. Rs 54,196.55/-, so there is something to check against.

The test that decides whether all six are present is blunt: a stranger holding all six reaches Rs 54,196.55/- without asking a single question, and a stranger missing any one of them cannot. Take away the source and there is nothing to run the cells against. Take away the sequence and they can reach four different figures, as the panel above showed. Take away the convention and they can reach four more, one for each denominator they might reasonably choose. Take away the row counts and they can get the right answer without ever knowing whether they got it the right way. Getting it right without knowing is the same as not knowing. Take away the answer and nobody can tell a successful rebuild from a failed one.

SIX THINGS HANDED ACROSS, AND THE SAME SIX MINUS ONE the source ashanded over the question,with the grain the cells, in thesequence they ran the convention,denominator named the row countat every step the answer, tocheck against Rs 54,196.55/- reached without asking a question the source ashanded over the question,with the grain the cells, in thesequence they ran the convention,denominator named the row countat every step the answer, tocheck against four possible figures and no way to say which one was meant remove any one of the six and the handover breaks in the same way: the second person is left reading the figure back instead of rebuilding it, which is agreement rather than a check the missing item drawn here is the sequence, but the drawing would be the same for any of the six
A stranger with the source, the question, the cells in sequence, the convention, the row counts and the answer reaches Rs 54,196.55/- without asking a question.
Try it out

A stranger has the notebook and the answer, but not the office file in the untouched state it arrived in. Can they reach Rs 54,196.55/-, and what exactly stops them?

How to build a reproducible financial analysis: which three rules never bend?

Three rules, and each one exists because of a specific way that analyses stop being rebuildable.

Never change the source, ever. Every correction is a cell, not an edit. The moment the record the office handed over is opened and the typed takings fixed by hand, the fault vanishes and so does every trace of the decision made about it. Next month, when somebody asks why NB-08 reads Rs 48,000/- when the office file says Rs 4,80,000/-, there is nothing to point at. With the fault kept in the source and the correction in a cell, the question answers itself.

Never type a figure in by hand. A typed figure has no provenanceThe written trail behind a figure: which record it was drawn out of, and which steps touched it before it arrived. A figure carrying none of that can be repeated but never rebuilt.. A typed figure cannot be rebuilt, cannot be traced back to a cell, and does not change when the source changes. Typing a number is faster than computing it, so this is the rule people break most. A typed figure also looks exactly like a computed one. Breaking the rule therefore costs the most.

And rerun the whole thing from a clean start before anything is sent. A notebook that works only in the accidental order it happened to be run in is a notebook that works once. Restarting, running every cell from the top, and reading the printed row counts in order settles it. If the sequence still gives thirty two, thirty one, thirty, twenty nine, twenty nine and Rs 54,196.55/-, it is a workflow. If it gives something else, it was a scratchpad wearing a workflow's clothes, and that has been found out cheaply.

THREE GATES ON ONE PATH, AND WHAT EACH ONE KEEPS OUT SOURCE never change the source never type a figure in by hand rerun from a clean start before sending KEEPS OUT a corrected source in which the fault, and the decision, have vanished KEEPS OUT a figure that traces back to no cell and cannot be rebuilt KEEPS OUT a notebook that works only in the accidental order it was run in sent the path runs left to right and every gate is passed before the answer leaves the desk, never afterwards
Never change the source, never type a figure in by hand, and rerun the whole thing from a clean start before sending anything.
Try it out

Name the three rules that make an analysis rebuildable.

Risk Management Program Bootcamp — Fin Maverick

How to create a reproducible Python financial analysis project: what sits in the one folder?

Four things in one folder, and the arrangement is a layout rather than a purchase. Adopting it costs nothing.

The whole of the project, on a record of thirty two rows and on a record of thirty two million.
What sits thereWhat it is forThe rule on it
The source as handed overThe thirty two rows the market office typed, faults includedNever edited, by anybody, for any reason
The notebookThe cells that carry the question to the answer, in orderReruns from the top and prints a row count at every step
The answer, written out as a fileRs 54,196.55/-, on disk rather than on a screenWritten by a cell, never typed
One plain noteThe question in a sentence and the convention it was answered underShort enough that somebody actually reads it

The reason the answer is written out as a file rather than read off a screen is that a screen keeps no history and a file does. Next month the record gains a fifth month, the notebook reruns, and the answer file changes. The change, and its size, are then visible. A figure that only ever existed in a cell output has nothing to compare against. The two analysts below are standing in exactly that spot.

ONE FOLDER, FOUR THINGS IN IT, AND NOTHING ELSE THE SOURCE AS HANDED OVER 32 rows, 8 columns, every fault still in it never edited, by anybody THE NOTEBOOK five cells, in the order they run each printing the row count it ended with THE ANSWER, WRITTEN OUT Rs 54,196.55/- sitting in a file written by a cell rather than typed ONE PLAIN NOTE the question in one sentence, and the convention the answer carries the same four work on 32 rows and on a record too large to open, because none of them is about size
One folder holds the untouched source, the notebook, the answer written out as a file, and a plain note naming the question and the convention.
Bond Pricing and Yield Mechanics — free micro-course from Fin Maverick

Why do two people get two answers from one file?

Two analysts at the market office are asked the same question and given the same file. The first reports Rs 69,035.45/-. The second reports Rs 55,723.30/-. The two figures sit Rs 13,312.15/- apart, and both people are certain they ran the analysis.

Here is the uncomfortable part. Neither of them is careless. The first read the file, computed the average over every takings cell carrying a number, and reported it. Nobody had made the question precise, so that reading is defensible. The second cleaned, but skipped the office code drop, so a code meaning no return received went into the arithmetic as though it were money. Each of them can describe what they did. Neither of them wrote down which cells ran, in what order, with what count after each, and so there is no artefactSomething a piece of work leaves behind that can be picked up and examined afterwards, as against a memory of having done it. anywhere that can settle the disagreement.

Two clerks in one office both say they totalled the day book. One totalled it, the other totalled it too, and neither wrote down which sheets they turned. There is now no way to find out who missed a sheet short of doing the whole thing again, and by then it is a third total.

ONE FILE NAME, HANDED TO TWO PEOPLE ON THE SAME MORNING THE FIRST NOTEBOOK the figure Rs 69,035.45/- the cells that ran, in sequence blank the row count at each step blank THE SECOND NOTEBOOK the figure Rs 55,723.30/- the cells that ran, in sequence blank the row count at each step blank Rs 13,312.15/- apart, and nothing in either file says which cells were run, so nothing can settle it fill in either of the two blank rows on either notebook and the disagreement becomes a one minute conversation
Two figures Rs 13,312.15/- apart under one file name cannot be settled because neither notebook records which cells ran.

The failure: a figure that is merely wrong is cheap, and a figure that cannot be attributed is not

Rs 69,035.45/- and Rs 55,723.30/-, Rs 13,312.15/- apart, from one file on one morning. The meeting that follows does not settle it. There is nothing in either notebook to settle it with. Both analysts describe what they remember doing. Both descriptions sound reasonable. The meeting ends with somebody promising to look again.

The cost is not the wrong figure. The cost is the meeting, and then the second meeting. A figure that is merely wrong gets corrected in a morning and stays corrected. A figure that cannot be attributed consumes that same morning every time it comes up. Nothing about the situation has changed, so it comes up again. The fix is one habit and not one purchase: restart, run every cell in sequence, and let each cell print the row count it ended with. On this record that habit is worth up to Rs 14,896.55/-, and it costs the time it takes to make tea.

Try it out

Two figures Rs 13,312.15/- apart under one file name. What single missing item makes the disagreement unsettleable, and what would have to be added to settle it in a minute?

Two analysts, one file, two answers, neither careless. See what the workflow settles.

What gets checked in the minute before a figure is sent?

How a lender, an analyst or a market office actually uses this

A lender reading a small trader's monthly takings, an analyst preparing a figure for a committee, and the Neelbagh market office setting a pitch fee are all in the same position: they are about to act on one number and they will be asked where it came from. The habit that protects all three is the same and takes about a minute: restart the notebook, run every cell from the top, read the printed row counts in order, and check the final count against the count expected before the run began.

All of it comes down to one sentence: a figure is only as good as the count beside it and the path behind it. The count states which figures the average was taken over. The path states what was done to get there. A figure with both can be defended, corrected and rebuilt by somebody else. A figure with neither is a rumour with a decimal point.

On this record the two checks are worth stating with prices. Rerunning from a clean start is worth up to Rs 14,896.55/-, the amount the largest single skipped cell moves the answer by. Checking the final count is worth the difference between an answer over twenty nine and an answer over thirty. On the office code skip alone that difference is Rs 1,526.75/-. Neither check requires knowing anything that was not already known at the start.

One caution on row orderThe sequence the lines were keyed in. The keying order belongs to whoever typed the file, and it says nothing at all about the market the file describes., the easiest thing on this record to read a false meaning into. The direction each skipped cell moved the answer is a fact about these thirty two rows and these four faults. Moving the same faults around changes the directions. Four faults on one record license no rule about which way an unfinished analysis reads, and that is exactly why the sequence gets printed instead of remembered.

Try it out

A figure is about to go to somebody who will act on it. Name the two things checked first, and say what each one is worth on this record.

What is covered elsewhere

The written down path that carries a raw handover into a table fit to compute on, the two delivery doorways, where a stored table sits, the table an analyst works in, putting two tables side by side, and asking a stored table a precise question all came earlier in this sequence, and all six are assembled here rather than repeated. Judging whether some entry in the office file is a fault or merely an unusual month is covered separately.

Tools for keeping a history of file changes, running a notebook on a timer, and putting an analysis somewhere other people can reach it are covered elsewhere, and none of them is a finance analyst's daily work. The language itself is covered elsewhere too: its grammar, loops and branching, preparing a machine, and keeping installed libraries current.

What produced every figure above?

Every figure above was recomputed from the three sources listed below and from nothing else. Each row names where a source sits and what it settles.

What was usedWhere it sitsWhat it settles
The Neelbagh stall record, thirty two rows across eight columnsSettled in the guides before this one and reprinted in the panel aboveEvery takings figure the notebook reads
The Neelbagh day book, twelve notes kept by the market officeThe note that decides month three is quoted where it decides itWhich of two identical looking month three figures is a slipped key
The named convention for the average takings per stall monthRestated beside every figure it producesWhat number sits underneath the division

The Neelbagh market, the Neelbagh stall record, the market office and the day book are invented.
Educational material. Not advice on any investment, tax, budget or market position.

Covered in this topic

Subtopics

Research NotebookHow to Reproduce a Financial Analysis From Source DataHow to Build a Reproducible Financial AnalysisHow to Create a Reproducible Python Financial Analysis Project
← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.