Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Quant Analyst · CoreTrack
1Quantitative Methods, Financial Data & Programming
iProbability
Probability in FinanceRandom VariableProbability DistributionsThe Normal DistributionNormal Distribution ProbabilityThe Lognormal DistributionRandomness vs Uncertainty
iiStatistics and Inference
Population and SampleMean, Median and ModePrecision and AccuracyVariable TypesVariance, Standard Deviation and…Dispersion MeasuresStatistical BiasEffect SizeHypothesis TestingThe Sampling DistributionSkewnessKurtosisCovarianceConfidence IntervalArithmetic Mean vs Geometric MeanStatistical Significance vs Economic…Confidence Interval vs Prediction IntervalHow to Summarise a…
iiiCorrelation and Regression
RegressionCorrelation and CausationOrdinary Least SquaresInteraction TermsRegression CoefficientsRegression vs ClassificationHow to Build a…Spurious CorrelationRegression, Correlation and FitResidualsMulticollinearityAutocorrelation and Partial Autocorrelation
ivTime Series
Time Series in FinanceSimple, Weighted and Exponential…Moving Average CalculatorPrice, Return and Level SeriesHow to Prepare Time-Series…LagFrequencySeasonalityTimestampsTrendStationarity and the Unit RootHeteroskedasticityLeadRolling WindowsDifferencing
vSimulation and Numerical Methods
SimulationMonte Carlo SimulationHow to Run a…Numerical MethodsIterationResampling and the BootstrapPseudorandom Numbers and the SeedConvergence and ToleranceNumerical Stability
viOptimisation
OptimisationLocal and Global OptimaConstraintsConvex OptimisationThe SolverLinear ProgrammingThe Objective FunctionConstraint ViolationThe Feasible SetLagrange MultipliersQuadratic Programming
viiModelling Practice
Linear, Logistic, Ridge and…Training, Validation and Test…The ModelModel ErrorDependent and Independent VariablesThe ROC Curve and AUCWhat a Model HoldsMSE, RMSE, MAE and MAPEPrecision and RecallCross Validation and RegularisationOverfitting and UnderfittingReturn Series MeasuresSimple, Compound and Log Return
viiiBacktesting and Research Integrity
BacktestingBacktest vs Live PerformanceHow to Document a…How to Prevent Backtest…Out-of-Sample TestingWalk-Forward AnalysisMultiple TestingP-HackingData Snooping
ixData Quality and Structure
Data QualityThe DatasetSelection and Survivorship BiasVersioned DatasetsData Structures in FinanceData CleaningMissing Data and Null ValuesStructured Data vs Unstructured DataMissing Data vs ZeroData Validation vs Data CleaningOutliersDuplicate Records
xProgramming for Finance
Data PipelinesAPIs for Financial DataAPI vs CSV FileDatabases in FinancePython for FinanceJoinsSQL for FinanceThe Analysis Workflow
xiQuantitative Research
Research DesignThe Data Generating ProcessReproducibilityPeer Review in Analytical WorkThe Research HypothesisRobustness and Sensitivity
2Stochastic Calculus & Derivative Pricing Theory
iProbability Foundations
The Probability SpaceRandom VectorsSigma-AlgebraExpectationSample Space and EventsDensity and Distribution FunctionsRisk-Neutral ProbabilityState Price Density vs…
iiStochastic Processes and Jumps
Properties of a Stochastic ProcessMartingaleBrownian Motion and Its PropertiesBrownian Motion vs Geometric…Stopping TimeThe Markov PropertyState VariablesTransition ProbabilityQuadratic VariationQuadratic Variation vs Ordinary…Submartingale and SupermartingaleMartingale RepresentationMarkov Process vs MartingaleOptional StoppingFiltrationJump ProcessesThe Poisson ProcessLevy ProcessesJump Diffusion
iiiIto Calculus
The Ito IntegralThe Ito Integral vs the Riemann IntegralInfinitesimals in Stochastic CalculusQuadratic CovariationIto's LemmaHow to Apply Ito's…The Infinitesimal GeneratorIto Calculus vs Ordinary Calculus
ivStochastic Differential Equations
Stochastic Differential EquationsStochastic Differential Equation vs…Drift and DiffusionStrong and Weak Solutions ComparedDiscretisationGeometric Brownian Motion
vPricing Theory and No-Arbitrage
No-ArbitrageGirsanov, Radon-Nikodym and Change…Physical and Risk-Neutral Measures…The Fundamental Theorems of…The Law of One PriceThe Pricing KernelDiscount Factors and Zero-Coupon PricesReplication vs HedgingComplete Market vs Incomplete MarketClearing Margin Architecture
viOption Pricing Theory
European and American OptionsMonte Carlo European OptionThe Black-Scholes PDEBlack Scholes and the GreeksThe Payoff FunctionThe Binomial ModelBinomial Option PricingDelta Hedging in TheoryBoundary, Initial and Terminal ConditionsThe Exercise BoundaryHow to Check Put-Call…
viiVolatility Models
Constant, Local and Stochastic…Vasicek Model vs CIR ModelThe Heston ModelThe SABR ModelThe Volatility ProcessImplied VolatilityVolatility Smile vs Skew vs Surface
viiiInterest Rate Models
Interest-Rate DerivativesMean ReversionThe Zero-Coupon BondThe Ornstein-Uhlenbeck ProcessThe Discount CurveZero RatesShort-Rate Model vs Market Model
ixNumerical Pricing
Closed Form and Numerical…Monte Carlo PricingEuler and Milstein Schemes ComparedTree MethodsFinite Difference MethodsNumerical Error and StabilityVariance Reduction
xCalibration and Model Risk
Model OverrideMarket Price and Model PriceCalibrationHow to Document a Pricing ModelThe Educational Illustration LabelMarket ConventionsModel Uncertainty and LimitationsBacktesting a Pricing ModelIdentifiabilityCalibrated ParametersThe Calibration Loss Function

What a Model Holds: Features In, Parameters Learned, Hyperparameters Chosen

A model holds three kinds of thing. Features are the columns it reads. Parameters are the numbers it learns from the record, two on the straight line fitted to the Nakshatra unit and the Vasant unit. Hyperparameters are settings a person chooses before fitting starts and the fitting never touches, and a reported result usually shows the parameters and hides the settings.

Three things are already settled. A fitted line and the numbers sitting inside it were worked out in earlier material, so what it means to draw a rule through a scatter of readings is already established. The idea of charging a fitting routine extra for large numbers, and the idea of a cut off at which a call gets made one way rather than the other, have both been met already. And the ten paired months of the Nakshatra unit and the Vasant unit, both of them invented for teaching, are carried in whole from where they were built. Three names go to the parts of a model, and each part is counted on that one short record. The next result somebody hands over can then be read for exactly what is missing from it.

What are the three kinds of thing a model holds?

The line between the three is not about importance and not about size: it is entirely about who decided the number. Features are what goes in. Parameters are what the fitting step works out. Hyperparameters are what a person fixes before the fitting step is allowed to run at all. Every one of those three sentences answers the question who, and none of them answers the question how big. A setting somebody typed in four seconds can move a headline further than any number the arithmetic laboured over, and that happens twice below.

The everyday version holds all the way through. Nothing later breaks it. A cook puts a tray in the oven. The cook fixes the oven at a temperature, by hand, before anything goes in. The tray then takes as long as it takes, and the cook watches and writes the timing on the recipe card. The recipe card is the thing to notice. Somebody worked the timing out, and the card carries it. Nobody worked the temperature out; somebody just set it. The card very often does not carry the temperature at all. And a person handed that card in a different kitchen, with a different oven, will get a different tray and will not know why.

The third kind of thing is the whole of the problem: it is easy to leave out of a report precisely because nothing computed it, so nothing printed it. A fitting routine is a machine for producing numbers and then reporting them. The routine has nothing to say about the numbers it was handed on the way in. From its point of view the incoming numbers were never in question, but simply the terms on which it was asked to work.

THREE WAYS INTO ONE BOX, AND ONLY ONE WAY OUT The three inlets are told apart by who supplies them, not by how much they matter. FEATURES IN columns the model reads the Nakshatra unit the Chandana unit the Sharad marker the marker times the unit HYPERPARAMETERS CHOSEN a person fixes these before anything runs THE FITTING STEP the arithmetic that actually runs and then reports what it did PARAMETERS LEARNED worked out from the record itself 0.5000 and 1.5000 WHY THE REPORT IS ALWAYS LOPSIDED The box computed the numbers on the right, so it can report them. It did not compute the dials on top and it did not compute the columns on the left, so it says nothing about either, and a reader who sees only the printout is looking at one inlet out of three.
A model holds features it reads, parameters it learns from the record, and hyperparameters a person chooses before fitting starts, and only the middle group ever reaches the printout.
Try it out

What actually separates a parameter from a hyperparameter?

What is a feature, and which four does this record offer?

A feature is a column the fitting step is handed to read, and that is the entire definition. Not a column that is interesting, not a column that turned out to be useful, not a column somebody believes in. A column that was handed over. If it went in, it is a feature; if it was left in the file and never passed across, it is not, however much it might have helped.

The ten-month record offers four features, and all four were built in earlier material. The first is the Nakshatra unit's monthly change, the plain reading. The second is the Chandana unit, the same reading nudged by five hundredths of a percentage point up in four months, down in four more, and left exactly alone in two. The third is the Sharad marker, simply on in five of the ten months and off in the other five. The fourth is the marker multiplied by the Nakshatra unit's change, an interaction columnA column made by multiplying two other columns together, so that what one of them says is allowed to depend on what the other says.: it reads zero in every month the marker is off, and repeats the Nakshatra reading in every month the marker is on.

The fourth column is built entirely out of two columns that were already there, and it is still its own feature. A feature is whatever column is handed to the fitting step, not whatever column carries new information. Think of a form at a lending desk. The form asks for monthly income in one box and rent paid in another. Somebody designing the form can add a third box reading income less rent, and that third box holds nothing the first two did not already hold between them. The third box is still a box: still filled in, still read, still counted, and the clerk reading the form treats it as a question in its own right. The fitting step is exactly that literal.

MonthThe Nakshatra unitThe Chandana unitThe Sharad markerMarker times NakshatraThe Vasant unit
11.001.00off0.003.00
26.005.95off0.0018.50
3minus 4.00minus 3.95onminus 4.002.50
411.0011.05off0.0017.00
51.000.95on1.00minus 3.00
6minus 9.00minus 9.00off0.00minus 13.00
76.005.95on6.006.50
81.001.05on1.00minus 1.00
9minus 4.00minus 4.05off0.00minus 7.50
101.001.05on1.00minus 3.00

The last column is not a feature, and it is worth saying why. The Vasant unit is the column being explained rather than a column doing the explaining, so it is never handed across as an input. Which of two columns should sit on which side of that arrangement, and why swapping them is not a symmetric move, is covered separately. The count is what matters: four columns available to go in, one column waiting to be explained.

FOUR COLUMNS THIS RECORD CAN HAND OVER Two are read straight off the record. Two are built. All four count the same way. 1. THE NAKSHATRA UNIT The plain monthly change, read straight off the record with nothing done to it. Ten readings, from minus 9.00 up to 11.00. 2. THE CHANDANA UNIT BUILT: the Nakshatra reading nudged by five hundredths of a point, up in four months, down in four more, and left exactly alone in two. 3. THE SHARAD MARKER On in five months and off in five, and nothing else. Read straight off the record. five on, five off 4. THE MARKER TIMES THE NAKSHATRA UNIT BUILT FROM TWO COLUMNS THAT ARE ALREADY FEATURES, and still its own feature. row 3 x row 1 zero when the marker is off, the Nakshatra reading when it is on.
This record offers four features: the Nakshatra unit's monthly change, the Chandana unit, the Sharad marker, and the marker multiplied by the Nakshatra unit's change.
Try it out

The Sharad marker multiplied by the Nakshatra unit's change is built from two columns that are already features. Is it a feature in its own right?

What is a parameter, and how many does each model here learn?

A parameter is a number the fitting step chooses for itself, by hunting for the values that make the misses small, and nobody types it in. On the straight line through these ten months there are exactly two of them. One is the interceptThe part of a fitted rule that does not depend on any column at all: what the rule says when every reading handed to it is zero., coming out at 0.5000, and one is the coefficientThe number a fitted rule multiplies a column by. Two boards for every crate is a coefficient of two. on the Nakshatra unit, coming out at 1.5000. Neither was chosen. Both were found. Reading a coefficient off a fitted line is covered separately. A coefficient is one kind of parameter and always has been.

Now hand the fitting step one more column and count again. Add the Sharad marker as a plain input and it learns three numbers instead of two: an intercept of 2.1000, a coefficient of 1.5000 on the Nakshatra unit, and a coefficient of minus 3.2000 on the marker. Look at the middle one. The coefficient on the Nakshatra unit has not shifted a hair, and there is a reason: the marker and the Nakshatra unit line up with each other at exactly zero on this record, so adding one of them leaves the other's number untouched. Such a clean split is a property of this made-up record rather than a rule about models, and it will not usually happen.

Add the interaction column as well and the count goes to four. The four numbers are 1.8800 on the intercept, 1.7200 on the Nakshatra unit, minus 1.8800 on the marker and minus 1.3200 on the interaction. Every one of those four moved when the fourth column arrived, including the two that were 0.5000 and 1.5000 a moment ago. Push further and let the rule bend: a curve of degreeHow curved a rule is allowed to be. Degree one is a straight line, degree two bends once, degree three twice, and so on. four learns five numbers, and on these ten months it runs the leftover misses all the way down to the floor those readings set. The floor itself, and why no rule reading this one column can get under it, is settled separately.

What was handed to the fitting stepNumbers learnedMonths on the recordMonths per learned number
The Nakshatra unit alone, as a straight line2105.00
The Nakshatra unit and the Sharad marker3103.33
Both of those and the interaction column4102.50
The Nakshatra unit bent into a curve of degree four5102.00

The count runs 2, 3, 4 and 5 on one record of ten months, and the record never grew by a single month: every step up that ladder is a choice somebody made. That is the sentence to carry away. The record is a fixed thing, sitting there with ten rows in it, and a person walked up to it four times and asked for a different number of learned numbers each time. Nothing in the ten months requested any of that. Whether pushing the count up is a good idea, and what happens to a short record when the count climbs, is covered separately and deliberately left alone here.

THE COUNT CLIMBS. THE RECORD DOES NOT. Each strip below is the same ten months, drawn at the same width every time. THE MODEL CHOSEN THE RECORD, TEN MONTHS NUMBERS LEARNED MONTHS PER NUMBER a straight line 5.00 plus the Sharad marker 3.33 plus the interaction 2.50 a curve of degree four 2.00 Nothing in the ten months asked for any of this. A person asked, four separate times.
On one record of ten months the parameter count runs 2, 3, 4 and 5 depending only on which model was chosen, and the record never grows.
Try it out

A model carrying the Sharad marker and its interaction learns four numbers from ten months. What question should that raise?

Breaking Into Quants Bootcamp — Fin Maverick

What is a hyperparameter, and who chooses it?

A hyperparameter is one of the terms the fitting was asked to work under. A person fixes it before the fitting step runs, and the fitting step never reads it off the record. Four of them turn up in this arithmetic, and listing them plainly makes the shape easy to recognise. There is the degree of the curve: how much bending the rule is allowed. There is the size of the penaltyAn extra charge added to the fitting arithmetic for large learned numbers, so the fitting settles for smaller ones. Met earlier in these notes., which is how hard large learned numbers are discouraged. There is the number of foldsOne of the equal parts a record is cut into so that each part can take a turn being held back while the rest is used for fitting. a record gets cut into. And there is the threshold at which a call is made one way rather than the other.

Watch what those four have in common. Not one of them can be discovered by staring harder at the ten months. The record has nothing to say about whether the rule should bend once or four times. The record has nothing to say about how much a large number should cost. The ten months have no view on being cut into five parts rather than ten, and none at all on where a cut off should sit. Every one of the four has to be supplied from outside, and every one of the four changes the answer that comes back. That combination, no opinion from the record plus a real effect on the result, is exactly what makes a setting worth reporting and exactly what makes it easy to forget.

The everyday shape is a shopkeeper packing rice. Before any rice is weighed, somebody decides the packet is a kilogram. The packet size is not read off the sack but from a person, and the packet size determines how many packets the sack yields and what each one costs. No amount of weighing the sack will say what size the packet should be. Weighing will only ever say, once the size is decided, how many packets the sack gives.

The settingWhat it controlsRead off the record?What moves when it moves
The degree of the curveHow much the rule is allowed to bendnoThe count of learned numbers, and every one of their values
The size of the penaltyHow hard large learned numbers are discouragednoThe learned numbers themselves, sometimes enormously
The number of foldsHow many parts the record is cut into for testingnoThe figure the testing reports back
The threshold a call is made atWhere a score stops meaning down and starts meaning upnoEvery call, and the score the calling is judged by

The third column answers no four times over, and the repetition is the point of the table. How any one of these ought to be chosen, honestly rather than conveniently, is a whole subject of its own and is covered separately. Here they are named and counted, and none of them is a parameter.

FOUR SETTINGS, AND ONE COLUMN THAT ANSWERS NO EVERY TIME Each of the four changes the answer. None of the four can be found in the ten months. READ OFF THE RECORD? THE DEGREE OF THE CURVE how much the rule may bend no moves the count of learned numbers and all their values THE SIZE OF THE PENALTY how hard large numbers are discouraged no moves the learned numbers, sometimes enormously THE NUMBER OF FOLDS how many parts the record is cut into no moves the figure the testing reports back THE CALLING THRESHOLD where down stops and up begins no moves every call, and the score the calls are judged by THE FITTING STEP READS NONE OF THIS COLUMN. A PERSON SUPPLIES ALL FOUR.
The degree of the curve, the size of the penalty, the number of folds and the calling threshold are all chosen by a person and none is read off the record.

How is a parameter told apart from a setting in practice?

Names are cheap, so here is a test that can be run on a real number, and this record allows it twice over with two very different settings.

Test one, and it is the loudest result here: with the calling threshold moved from 0.50 to 0.75, the accuracyThe share of months a rule called correctly, counting an up called up and a down called down alike. goes from 60.00 per cent to 80.00 per cent with the model completely untouched. The rule scoring these months is a straight line on the up or down label, reading 0.4500 plus 0.0500 times the Nakshatra change, and it was fitted once. Nothing about it is recomputed. The threshold is applied afterwards, to scores that already exist, and that is precisely why moving it refits nothing. Two numbers that nobody fitted sit between the record and the headline, and one of them just moved the headline by twenty percentage points.

The arithmetic here is full of figures that look alike and mean nothing to each other, so one caution comes before the sweep. A threshold of 0.50 and an accuracy of 50.00 per cent are counted in completely different things: one is a cut off on a score that runs from zero to one, the other is a share of ten months called correctly. The two can be written with the same digits, but that is a coincidence of where the decimal point sits, and the resemblance carries no meaning whatever.

Calling thresholdMonths called upCalled up and went upCalled down and went downAccuracy
0.00105050.00 per cent
0.2595160.00 per cent
0.5074260.00 per cent
0.7533580.00 per cent
1.0011560.00 per cent

Every row of that table came out of the same two learned numbers. Not one refit happened anywhere in it. Two other ways of scoring the same calls pull apart differently as the threshold climbs, and both are covered separately.

Test two works from the opposite direction. Raise the penalty from 0.00 to 1.00, and the number sitting on the Nakshatra unit travels from minus 1.0000 all the way to 0.7314. Meanwhile R squaredA single figure saying how much of the movement in the column being explained a fitted rule accounts for. Higher means less left over. moves from 0.756019 to 0.755950, a shift of 0.000069. This one is a fit on two columns at once, the Nakshatra unit and the Chandana unit, which as noted above are near enough the same column twice. With no penalty at all the fitting step splits the work between them as minus 1.0000 and 2.5000, a gap of 3.5000. With a penalty of one it splits it as 0.7314 and 0.7661, a gap of 0.0347. The learned numbers have travelled a very long way and the quality of the fit has barely twitched.

So: the penalty was typed by a person before the fitting ran, and it changed what the fitting returned. The threshold was typed by a person after the fitting ran, and it changed what the result was called. Both are settings. Here is the test in three lines. If the fitting step produced it, it is a parameter. If changing it means fitting again, it is a setting. If changing it changes what the refit gives back, it is a setting, and the fact that it moved a learned number is what proves it sits above the fitting step rather than inside it.

ONE MODEL, TWO THRESHOLDS, TWENTY POINTS OF ACCURACY The two learned numbers are identical in both panels. Only the red line moves. THRESHOLD AT 0.50 0.00 0.25 0.50 0.75 1.00 7 months called up, and 4 of those 7 went up. 2 of the 5 down months were called down. ACCURACY 60.00 per cent learned, and unchanged: 0.4500 and 0.0500 THRESHOLD AT 0.75 0.00 0.25 0.50 0.75 1.00 3 months called up, and all 3 of those went up. All 5 down months were called down. ACCURACY 80.00 per cent learned, and unchanged: 0.4500 and 0.0500 FILLED WENT UP, HOLLOW WENT DOWN. NOTHING BETWEEN THESE PANELS WAS REFITTED.
Moving the calling threshold from 0.50 to 0.75 takes the accuracy from 60.00 per cent to 80.00 per cent with the model completely untouched.
Try it out

The calling threshold is about to move from 0.50 to 0.75. What happens to the two learned numbers?

Play with it

Slide the setting nobody fitted, and watch the numbers somebody did fit refuse to move.

The scoring rule is fixed for the whole of this panel at 0.4500 plus 0.0500 times the Nakshatra change, and it is never fitted again. Only the calling threshold moves. Each of the ten months sits on the score line at its own score, filled if the month actually went up and hollow if it went down, and a month is called up when its marker sits at or to the right of the red line. The panel opens at a threshold of 0.50 and reads an accuracy of 60.00 per cent. Push it to 0.75 and the same untouched model reads 80.00 per cent.

call every month upthreshold 0.50call almost nothing up
Calling threshold
0.50
Accuracy
60.00 per cent
The two learned numbers
0.4500, 0.0500
Times the model was refitted
0
Educational illustration. The two learned numbers, 0.4500 and 0.0500, are held fixed for the whole of this panel while only the setting moves, and the refit counter beside them stays at zero however far the slider is pushed.
Try it out

Moving the penalty from 0.00 to 1.00 changes the Nakshatra coefficient from minus 1.0000 to 0.7314. Is the penalty a parameter or a setting?

ONE QUESTION SEPARATES THEM, AND IT IS NOT ABOUT SIZE WHERE DID THIS NUMBER COME FROM? ask it of every number on the printout The fitting step worked it out by hunting for small misses PARAMETER It cannot be changed by hand. Changing anything else refits it. A person typed it in before the fitting step was allowed to run HYPERPARAMETER a setting, and it belongs in the report Changing it changes what the refit gives back.
If the fitting step produced it, it is a parameter; if a person typed it before fitting ran, it is a setting.
AI For Finance Bootcamp — Fin Maverick

What does a printed result leave out?

A fitting routine prints the parameters because it computed them, and it prints nothing about the settings because somebody typed those before it was called. There is no conspiracy in this and usually no carelessness either. Ask a machine to report on its work and it reports on its work. The degree it was told to use was not its work. The penalty it was handed was not its work. The folds it was given and the threshold applied to its output afterwards were not its work. So a printout carrying two coefficients and a goodness figure is a complete and honest account of the fitting step, and an incomplete account of the model.

Put it beside the cook and the recipe card again and the shape is identical. The card is a complete account of what happened once the tray went in. The card leaves out the one decision made before the tray went in, and that decision is the one most likely to be different in somebody else's kitchen. A result quoting only parameters describes the half of the model the arithmetic chose and stays silent about the half a person chose, and the silent half is the half that will not travel.

THE PRINTOUT, AND THE FOUR THINGS ABOVE IT THAT NEVER REACH IT CHOSEN BEFORE THE RUN. NOT PRINTED, BECAUSE NOTHING COMPUTED THEM. degree of the curve chosen: not stated size of the penalty chosen: not stated number of folds chosen: not stated calling threshold chosen: not stated these went in, and never came back out WHAT THE ROUTINE PRINTED TERM ESTIMATE intercept 0.5000 the Nakshatra unit 1.5000 how much of the movement is accounted for 0.7559 Three numbers, all three computed. Nothing above this card appears on it. Every one of the three printed numbers checks out under audit. That is still one inlet out of three.
A fitting routine prints the parameters because it computed them, and never prints the settings, because somebody typed those.
Try it out

Why does a printed model summary almost never show the settings it was run under?

Regression for Finance — free micro-course from Fin Maverick

What should be asked of a fitted model handed over by somebody else?

Four questions, and they take under a minute between them. Which columns went in? How many numbers were learned, and from how many records? Which settings were fixed before fitting, and what were they? And was any setting chosen after somebody had looked at the result it was going to be judged on?

The four questions are not an academic exercise. A person at a lending desk reading a scoring model, an analyst handed a fitted rule by a colleague, an investor sent a sheet of results, a household comparing two quotes worked out by two different calculators: every one of them is in the same position, holding an output and not the terms it was produced under. The first three questions describe the model. The fourth describes how the model came to be reported, and it is the one that does the real work.

A setting chosen after seeing the figure it was going to be judged on is not a setting any more; it is a parameter that somebody fitted by hand, without saying so and without counting it. On this record the room that leaves is visible exactly. Somebody who tries the threshold at five places and reports the best one has quietly fitted a number and then reported the accuracy as though nothing had been fitted. The two learned numbers are still 0.5000 and 1.5000 on the size model, or 0.4500 and 0.0500 on the label model, and every one of them is still honest. The headline is not.

The second question belongs beside the third. Together they show how much authorship is packed into how little evidence. Four learned numbers fitted on ten months is two and a half rows for every number the arithmetic had to choose. Ten months is a short record carrying a lot of decisions. Knowing the ratio is not the same as knowing what to do about it, and what to do about it is covered separately.

FOUR QUESTIONS, AND WHAT THIS RECORD ANSWERS The first three describe the model. The fourth describes how it came to be reported. 1. WHICH COLUMNS WENT IN? Four were available here: the Nakshatra unit, the Chandana unit, the Sharad marker, the interaction. 2. HOW MANY NUMBERS WERE LEARNED, AND FROM HOW MANY RECORDS? Two, three, four or five, out of ten months. At four numbers that is 2.50 months for each one. 3. WHICH SETTINGS WERE FIXED BEFORE FITTING, AND WHAT WERE THEY? The degree, the penalty, the number of folds and the calling threshold. Ask for the values, not just the names. 4. WAS ANY SETTING CHOSEN AFTER SOMEBODY SAW THE RESULT IT WOULD BE JUDGED ON? If yes, that setting is a number fitted by hand, uncounted and unreported, and the headline is no longer what it says it is. This is the question the other three cannot reach. Under a minute, all four, and they yield more than an hour spent checking the arithmetic would.
Ask which columns went in, how many numbers were learned and from how many records, which settings were fixed, and whether any was chosen after seeing the result.
Try it out

Which of the four questions turns a reported figure into an unreported choice?

An audit that checked every number nobody had chosen

A fitted model arrives with a short note attached. Two coefficients, a goodness figure, and a headline accuracy. The person receiving it does the responsible thing and checks the arithmetic. The coefficients are recomputed from the ten months and they land exactly where the note says. The goodness figure is recomputed and it lands exactly where the note says. Everything reconciles. The note is signed off.

Nothing was verified. The degree was chosen by somebody. The penalty was chosen by somebody. The threshold was chosen by somebody, and on this record moving that last one alone takes the accuracy from 60.00 per cent to 80.00 per cent without a single learned number changing by so much as a decimal place. The coefficients have nothing to do with the setting at all, so every coefficient in the note was correct at every one of those settings. The audit examined the half of the work the arithmetic did and never went near the half a person did, and the half a person did was where the whole of the twenty point swing lived.

The habit that fixes it costs one line. Before any number is checked, the settings are requested, in writing, with their values. Then comes the fourth question: was any of them chosen after somebody had already seen the figure being reported. A model whose settings arrive without argument is a model that can then be checked. A model whose settings arrive after some hesitation about which run this was, or which threshold got used in the end, has said something the arithmetic never could.

Try it out

The coefficients of a reported model are audited and every one of them checks out. What has actually been verified?

How any setting ought to be chosen, honestly rather than conveniently, is covered separately, as is what happens to a short record as the count of learned numbers climbs against it. What a coefficient is, what a penalty does and where a calling threshold came from are each settled in earlier material. Whether any of these four columns would be worth reading anywhere outside this arithmetic is a separate question again. The two ways of scoring calls that pull apart as the threshold climbs are treated separately.
Backtesting a Strategy teaches you to build a backtest, name how it flatters itself, and state what the result establishes.

Where did these figures come from?

Every number above was counted or recomputed from ten made-up months printed in full above, so the whole of it can be redone on paper. Naming what a fitted model holds needs no outside authority at all, and the table below has exactly one row because of that.

DocumentSiteDate consulted
None. The four columns, the learned numbers and the four settings were all made up for teaching and recounted here from the ten months printed above None Not applicable, because no maintained record was opened

The Nakshatra unit, the Vasant unit, the Chandana unit and the Sharad marker are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.