Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Risk, Treasury & Financial Control
1Risk Foundations
Risk Appetite, Tolerance, Capacity…The Risk Taxonomy and UniverseRisk Register vs Risk MatrixStress TestingScenario Analysis vs Stress TestingImpact and LikelihoodLikelihoodThe Risk EventRisk Assessment
2Enterprise Risk Management
Enterprise Risk ManagementThe Four Risk TreatmentsRisk CultureRisk MaturityRisk Monitoring
3Risk Governance
Risk GovernanceHow to set a…The Risk PolicyThe Risk OwnerThe Risk Committee and Its CharterThe Risk Limit FrameworkRisk EscalationHow to set a…
4Credit and Counterparty Risk
Collateral AgreementsCollateral vs NettingProbability of DefaultExposureCounterparty ExposureConcentration Risk vs Wrong Way RiskCounterparty Risk vs Credit RiskHow to assess Counterparty ExposureHow to assess Concentration Risk
5Market Risk
Market RiskSensitivity MeasuresThe Hedging PolicyInterest Rate Risk in the Banking BookIRRBB vs Market RiskExpected ShortfallEconomic Value of EquityVaR BacktestingOpen PositionValue at RiskValue at Risk and Expected ShortfallEconomic Value SensitivityFX ExposureValue at Risk vs Expected ShortfallEarnings at Risk vs…FX Transaction Risk vs…How to measure Interest…How to measure Foreign…
6Liquidity Risk
Liquidity Stress TestingLiquidity Gap vs Liquidity BufferMaturity MismatchThe Debt Maturity ProfileFunding ConcentrationSurvival HorizonThe Contingency Funding PlanNet Stable Funding RatioLiquidity Risk vs Funding RiskLiquidity Coverage RatioLiquidity Gap and BufferHow to run a Liquidity Gap Analysis
7Operational Risk
Operational LossThe Loss EventRisk and Control Self AssessmentException ManagementInformation Security as a…Segregation of DutiesIssue ManagementThe Near MissRoot Cause Analysis in RiskThe Fraud TriangleCyber Risk vs Third Party RiskHow to run a…How to assess Third…
8Risk Reporting, Data and Model Risk
Model RiskModel Validation vs BacktestingHow to run Model ValidationData Governance in RiskModel Risk vs Data RiskKey Risk IndicatorsManagement InformationRisk ReportingRisk ScoreEarnings at RiskRisk Adjusted ReturnEarly Warning IndicatorsHow to build a KRI Dashboard
9Treasury
Corporate TreasuryAsset Liability ManagementIntragroup FundingThe Treasury PolicyThe Treasury Management SystemThe Cash ForecastCash Pooling and ConcentrationHow to build a Cash Forecast
10Financial Controls and Assurance
Control AssuranceThe Control LifecycleThe Assurance MapThe Audit FindingIssue RemediationInternal Financial ControlsControl Design vs Control EffectivenessHow to map Internal Financial ControlsHow to test Control…Control DeficiencyMaterial Weakness
11Operational Resilience
Operational ResilienceBusiness Continuity and Disaster RecoveryBusiness Continuity vs Operational…Crisis ManagementDisaster RecoveryIncident Management

How to run Model Validation: Seven Steps and a Written Report

Validation runs in seven steps. Scope and tier the model, test the data, challenge the assumptions, replicate the implementation, analyse outcomes where any exist, write the limitations as instructions, then produce a report carrying rated findings, a use restriction and a signature. The fifth step is often empty, and a model with no observable outcome still gets the other six.

The seven step sequence is the shape the work takes, and it is not a standard handed down from anywhere. An institution's actual obligation to check its models is set by its own supervisor, and the requirement lives in that supervisor's text rather than in any method. The seven steps are one workable structure that a single person can hold in their head, and what actually binds a given institution is settled against the supervisors named below.

Everything below rests on one idea about method, and it is worth stating before the first step arrives. A validation is not a test, it is a sequence of tests on different objects, and each object has a different owner. The data belongs to a source system. The assumptions belong to a committee. The implementation belongs to whoever built it. The use belongs to whoever runs the model and acts on what comes out. A method that treats a model as one thing produces one verdict, and there is no route to a fix inside a verdict. A method that separates the objects produces findings somebody can pick up. Each one lands on a desk that can do something about it.

The household version comes first, so the shape is familiar before the vocabulary arrives. A household works out what it can afford as a monthly loan instalment: take-home pay in, usual spending out, keep half of what is left. Now suppose the answer looks wrong. Four separate things could be wrong with it and they are fixed by four different people. The pay figure could be stale, and only whoever reads the payslip can settle it. Keeping half could be the wrong rule, and only whoever chose the rule can settle that. The subtraction could have been done wrongly on the paper, and neither of those two people would know. Or the answer could be perfectly correct for a two year loan and be getting used for a seven year one, and use is a question all of its own. Asking whether the number is right produces an argument. Asking the four questions separately produces four answers, three of which are probably fine.

What is a validation for, and what may the person running it never have done?

A validation answers one question: is this model fit for the use it is actually being put to? Not is it clever, not is it the best available, not does it agree with last year. Until somebody writes the use down, the question has no answer at all. The missing statement of use is why the first step produces a document rather than a test result, and why a validation that starts by opening the spreadsheet has started in the wrong place.

The second thing that defines a validation is who runs it. The person doing the checking must have built none of what they check, and that is a condition on the person rather than on the method. At the invented Vindhya Commercial Bank Limited, that person is Kanaka Murthy, the independent validator sitting in the risk function, and the bank's rule is the plain one: she validates nothing she built and nothing she specified. The rule is not a comment on anybody's honesty. Somebody who chose an assumption has already settled once whether it was worth choosing, and will not ask again with an open mind.

The method below has seven steps, numbered MV1 to MV7, worked end to end on one model that has never been checked at all. Each step takes something in, does one thing to it, and hands something on. Nothing in the sequence is clever. The value is in the fact that it is a sequence: running the assumption challenge before the data test costs an afternoon of argument about a choice that was being fed rubbish anyway.

SEVEN STEPS, AND EACH ONE HANDS THE NEXT ITS INPUT Run them out of order and the work still ends in a verdict, but with no route to a fix inside it. STEP WHAT IT IS TAKES IN HANDS ON MV1 SCOPE AND TIER the model record, its use and its tier a written scope, agreed before work starts MV2 DATA the model inputs and their source systems four data tests, each with a verdict MV3 ASSUMPTIONS the specification and the committee record every assumption, each with a named owner MV4 IMPLEMENTATION the same inputs, and the code that runs an independent rebuild and a comparison MV5 OUTCOMES what the model said, and what happened a comparison, or a written absence MV6 LIMITATIONS whatever MV2 to MV5 could not settle limitations written as instructions MV7 THE REPORT the outputs of MV1 to MV6 findings, a use restriction, a signature MV1 to MV7 is this guide's own numbering for the sequence, and every practice shown belongs to one invented bank.
Seven rows for seven steps, and the arrow running down the left edge carries the whole argument: what one step hands on is exactly what the next one needs, so a step taken early is a step taken on an input that does not exist yet.
Risk Management Program Bootcamp — Fin Maverick

How is a validation scoped and tiered, and what does the tier change?

Step one of seven

MV1 Scope and tier

Who does it
The independent validator, working from the model record rather than from the model.
Takes in
The model's row on the register: its name, its stated use, its owner and its tier.
Does
Writes down which model is being checked, for which use, and to what depth.
Hands on
A written scope, agreed with the model owner and the risk function before any testing starts.
Decides
Whether the depth carried on the record is accepted for this validation or challenged in writing.

The validation scopeThe written statement of which model is being checked, for which use, and how deeply, agreed before any work starts. is a short document and it is agreed before anything is opened. The scope names the model, it names every use the model is put to, and it states the depth. Agreeing the scope first stops the argument at the end. The only thing anybody can then dispute is whether the work was done rather than whether it was the right work. A scope written after the testing is a scope written to fit the testing.

The tier is the second field, and the tier decides depth. The tier is a materiality rating the institution sets for itself: how much money moves if this rule is wrong, how many decisions rest on it, whether it feeds a limit, a capital number or a committee paper. Every model gets the same seven steps. The tier decides how far each step goes, how often the whole sequence is repeated, and how senior the person signing at the end has to be. Vindhya Commercial Bank Limited runs three tiers on its own test, and its register at month 12 splits 28 models into 6 in tier 1, 13 in tier 2 and 9 in tier 3, and six plus thirteen plus nine is twenty eight.

Now the part of MV1 that is easy to skip, and it matters on the model this guide works through. The tier on the record is an input to the scope rather than a fact about the world, and MV1 is allowed to challenge it in writing. A depth taken straight off a register is a depth chosen by whoever filled the register in. A validation exists to test exactly that kind of choice. Hold that thought: it comes back the moment the worked instance reaches model V1.

Try it out

What does the model's tier change about the validation?

What is actually tested about the data a model runs on?

Step MV2 is the one people expect to be dull, and it is where a surprising share of the findings come from. MV2 is not a data quality review of the whole institution, and it is not the governance of the data. Data governance is a separate subject with its own owner and its own list of attributes. MV2 asks four questions about the specific inputs this specific model consumes, and each one gets a verdict written down beside it.

Step two of seven

MV2 Data

Who does it
The validator, with the source system contact named against each input.
Takes in
The specification's input list, and the source system each input is drawn from.
Does
Runs the four tests below against every input the model consumes.
Hands on
Four verdicts per input: passed, failed, or unprovable on the evidence available.
Decides
Whether the rest of the validation proceeds on this data or proceeds qualified.
FOUR TESTS, AND EVERY INPUT TAKES ALL FOUR Three of them are about the data. The fourth is about whether the data can carry the question being asked of it. TEST 1: IT EXISTS Every input the written specification names is actually there, for the whole period the model runs on. A field first collected in month 6 cannot feed a rule that looks back twelve. TEST 2: IT HAS A SOURCE It arrives from a named source system and not from a file somebody prepares each month. An input with no source has nobody who can be asked what it means or why it moved. TEST 3: SAME MEANING The field means what the rule assumes it means. A balance that falls to zero while the account stays open is a live account to one system and a closure to the next one along. TEST 4: IT IS ENOUGH There is enough history to support the answer being asked for. A rule about how long a deposit stays needs a run of years and not a run of months, whatever the field quality is. Tests 1 to 3 can be passed by a clean feed. Test 4 can fail on a feed with nothing at all wrong with it.
Three of the four data tests can be passed by any clean feed, and the fourth can fail on a feed with nothing wrong with it, which is why sufficiency gets its own written verdict instead of being folded into quality.

The fourth test is the one that catches people, so it is worth holding apart from the other three. Tests one to three ask whether the data is any good. Test four asks whether there is enough of it to answer the question the model is being asked, and a feed can be flawless and still fail. Two years of balance history is a perfectly good history. Two years is simply not long enough to see how balances behave through a full turn of the rate cycle, and a rule about behaviour through the cycle that has never seen one is a rule fitted to a fragment.

Two more things MV2 hands on rather than settles. MV2 records the source system against every input, and it records how the model behaves when a value does not arrive. Recording the source is what makes the next person's job possible at all. Vindhya Commercial Bank Limited has an instructive gap there. Across the 147 data elements feeding its monthly risk report, the attribute most often missing is the one saying what happens when a value is absent, and it is missing on 105 of them. A data dictionary that describes the value and never its absence describes only the days when everything works. Whose job it is to close that gap is a data governance question, covered separately.

How is an assumption listed, and who does each one belong to?

Step three of seven

MV3 Assumptions

Who does it
The validator, with the written specification and the committee record open together.
Takes in
The specification, and the minutes of whichever body chose each input that was not measured.
Does
Lists every assumption the model makes and names the person or committee that chose it.
Hands on
An assumption list, each entry carrying an owner, a date and the evidence behind it.
Decides
For each entry, whether it can be tested against evidence or only against alternatives.

An assumptionA choice made because evidence does not settle the question, which is why an assumption always has an owner and a date. is a choice made because the evidence does not settle the question. The definition is short and it carries a consequence: somebody chose it, so somebody can be named, and it was chosen on a day, so it can be dated and revisited. An assumption with no owner is not an assumption, it is a habit. Nobody is answerable for a habit, so nobody reviews it. Listing them is most of the work in MV3, and the list is longer than the model's authors expect, because a choice made once and never argued about stops looking like a choice.

At the invented Vindhya Commercial Bank Limited the assumption followed here is a single number. Model V1, the behavioural deposit life model, gives the Rs 36,000 crore of current and savings balances an average behavioural life of 0.5 years. The current and savings balances are contractually repayable on demand, so no contract fixes the answer and something had to be chosen. Committee G4, the asset liability management committee, chose it. Vindhya Commercial Bank Limited makes that committee responsible for behavioural assumptions. So the entry on the list reads: assumption, 0.5 years; owner, committee G4; evidence, the bank's own review. MV3 writes that down and moves to the challenge.

Derivatives Foundation Bootcamp — Fin Maverick

How is an assumption challenged when there is nothing to compare it with?

An assumption with nothing to compare it against is the hard case, and it is the normal case rather than the exception. How long a deposit stays is only visible over years and the record holds no such measurement, so the 0.5 year figure cannot be checked against what actually happened. There is nothing to compare it with. The temptation at this point is to write that the assumption looks aggressive. An opinion is all that is, and an opinion loses an argument with the committee that chose it.

The way out is to stop asking whether the assumption is right and start asking how much it matters. The second question has a name, the range testRecomputing an answer across a range of an assumption to see how far the answer moves and whether it changes sign.: recompute the answer across a plausible range of the assumption and report where the answer goes. The range test converts an argument about a number nobody can settle into a measurement of consequence. Anybody can settle a measurement. And on model V1 it produces something much sharper than a range.

The arithmetic is the bank's own and it is short. Extending the assumed life by one year moves the economic value result by Rs 36,000 crore times the bank's own 2.0 per cent scenario. Rs 36,000 crore at 2.0 per cent is Rs 720 crore. Start from the locked point: at 0.5 years the economic value change is minus Rs 840 crore. Walk it forward. At 1.0 year it is minus Rs 480 crore, at 1.5 years minus Rs 120 crore, at about 1.67 years it is zero, at 2.0 years it is plus Rs 240 crore, at 2.5 years plus Rs 600 crore and at 3.0 years plus Rs 960 crore. The assumption is not adjusting the size of the answer, it is deciding which way the answer points, and that is a finding rather than an opinion.

ONE ASSUMPTION, AND THE ANSWER CROSSES ZERO Economic value change at seven assumed deposit lives, Rs crore, on the invented bank's own 200 basis point scenario. THE BANK'S OWN REPRICING BUCKET RB5: ONE TO THREE YEARS CAP, PLUS Rs 990 CR CAP, MINUS Rs 990 CR ZERO 0.5 1.0 1.5 1.67 2.0 2.5 3.0 VALUE CHANGE, Rs CR SHARE OF THE L8 CAP minus 840 minus 480 minus 120 zero plus 240 plus 600 plus 960 84.8 per cent 48.5 per cent 12.1 per cent 0.0 per cent 24.2 per cent 60.6 per cent 97.0 per cent Assumed average behavioural life of the deposits, in years, along the bottom. Every figure belongs to one invented bank.
Seven recomputations of one number, and the bars flip from below the line to above it, so the choice of deposit life is not tuning the magnitude of this bank's headline value reading, it is choosing the reading's direction.

A reader who skips the picture still needs two things from it. First, the answer stays comfortably inside the relevant limit at every point drawn. Limit L8 caps the economic value sensitivity at Rs 990 crore. The widest reading drawn is plus Rs 960 crore, or 97.0 per cent of the cap. Nothing on any report would look unusual at either end of the range, and a range test is worth running precisely for that reason: the failure mode is not a breach, it is a sign.

Second, and this is the part that makes it a validation finding rather than a curiosity, the range is not invented for the occasion. Vindhya Commercial Bank Limited's own repricing ladder slots the same Rs 36,000 crore at one to three years in bucket RB5. So the bank holds two views of one balance at the same time. Across bucket RB5 alone, from 1.0 year to 3.0 years, the answer runs from minus Rs 480 crore to plus Rs 960 crore and flips sign at about 1.67 years. The minus Rs 840 crore figure is the reading at the 0.5 year life the model itself assumes, and a 0.5 year life sits outside bucket RB5 entirely. The two must not be blurred into one range. The strongest form of a range test is a range the institution has already committed to somewhere else in its own reporting.

Try it out

The 0.5 year deposit life assumption is under challenge and there is no historical outcome to compare it with. What does the validator do?

Investment Banking Analyst Bootcamp — Fin Maverick

What does replicating an implementation mean, and why is it separate from checking the rule?

Step four of seven

MV4 Implementation

Who does it
The validator, building independently and without looking at the original workings.
Takes in
The written specification, the same inputs, and the code or spreadsheet that runs today.
Does
Rebuilds the computation from the specification and compares the two outputs number by number.
Hands on
An independent rebuild, a comparison, and a list of every difference found.
Decides
Whether each difference is a rule question for step MV3 or a build question for the developer.

ReplicationRebuilding a computation independently from the same inputs, to separate a wrong rule from a rule wrongly coded. means rebuilding the computation from the written specification, using the same inputs, without looking at how the original was built, and then comparing. Replication is slow and it feels redundant right up until the moment it is not. The reason it earns a step of its own is that steps MV3 and MV4 catch two faults that look identical from the outside and are repaired by two completely different people.

Think of a household again. Two people work out the monthly instalment they can afford and get different answers. One possibility is that they disagree about the rule, one keeping half of what is left and the other keeping a third. The other possibility is that they agree completely about the rule and one of them added a column wrongly. From the outside both look the same: two numbers that do not match. Inside, the first needs a conversation and the second needs two minutes with a calculator. A wrong rule and a right rule wrongly built produce the same wrong number. One step therefore asks about the rule and a separate step asks about the build.

TWO FAULTS, ONE WRONG NUMBER Step MV3 asks whether the rule is the right rule. Step MV4 asks whether the build does what the rule says. PANEL A: THE RULE IS WRONG THE SPECIFICATION FAULT HERE Says the wrong thing, so the rule is not the right rule. WHAT THE BUILD DOES Exactly what the specification says, line for line. THE NUMBER PRODUCED Wrong, and nothing in the build is at fault. Found by step MV3. Fixed by whoever chose the rule. PANEL B: THE BUILD IS WRONG THE SPECIFICATION Says the right thing, so the rule is the right rule. WHAT THE BUILD DOES FAULT HERE Something else. A sign or a divisor is not the spec. THE NUMBER PRODUCED Wrong, and nothing in the rule is at fault. Found by step MV4. Fixed by whoever built it. THE SAME WRONG NUMBER COMES OUT OF BOTH PANELS Skipping step MV4 produces a long argument about a percentage while a formula that was wrong on the day it was written stays wrong.
Both panels end at an identical wrong figure and the fault sits in a different box in each, so a validation that runs only the assumption challenge has looked at half the places the number could have gone wrong.

MV4 is where the independence condition does real work, and that has one practical consequence. A rebuild that starts by reading the original code will reproduce the original code's mistakes with great fidelity, so the rebuild is made from the specification and not from the original workings. If the specification is too thin to rebuild from, that is itself the finding, and it is a common one. Model V3 at this bank is a spreadsheet scorecard maintained by one person with no written specification, and MV4 cannot be run on it at all. There is nothing to rebuild from and nothing to compare against.

Try it out

Why is replication a separate step from challenging the assumptions?

What is done when the model cannot be tested against any outcome?

Step five of seven

MV5 Outcomes

Who does it
The validator, working from whatever record of realised outcomes the institution actually keeps.
Takes in
What the model said on each past date, and what was afterwards observed on the same date.
Does
Compares the two, counts the misses, and describes their pattern and their size.
Hands on
An outcome comparison where outcomes exist, and a written statement of the absence where they do not.
Decides
Whether the section carries a result, or carries a reason there can be no result.

Outcome analysisComparing what a model said with what happened, which is only possible where the outcome can actually be observed. is the step everybody pictures on hearing the word validation, and it is the step most likely to be empty. Outcome analysis needs two things: the model must have said something on a past date, and the thing it said must have been observable afterwards. Plenty of models fail the second condition through no fault of anybody's.

Model V1 fails it completely. Model V1 predicts how long a balance stays, and that is only visible over a run of years. The bank's record holds no realised deposit life at all, so there is nothing to place beside the model's output. Note carefully what that does and does not mean. Nothing in an empty section says the model has been checked and found sound on outcomes. The absence means the check cannot be performed, and a check that cannot be performed is a different statement altogether. The difference is exactly what the report has to carry.

AN EMPTY SECTION AND A SECTION ABOUT EMPTINESS The same validation, the same absent outcome, and two documents that do completely different things to a reader. VERSION ONE: SECTION LEFT BLANK 5. OUTCOME ANALYSIS nothing written here A READER CONCLUDES Somebody did not get round to this part. VERSION TWO: THE ABSENCE WRITTEN 5. OUTCOME ANALYSIS No outcome analysis is possible for this model. What it predicts, how long a balance stays, is only observable over a run of years. This institution holds no realised deposit life at all, so there is nothing to place beside the output. This is a statement about what can be known and not about work not done. A READER CONCLUDES This one cannot be settled by testing yet. A BLANK SECTION IS AN ABSENCE OF WORK. A WRITTEN ABSENCE IS A FINDING ABOUT WHAT CAN BE KNOWN.
Nothing about the underlying model differs between these two versions, and only the version on the right lets a reader tell an unanswerable question apart from an unanswered one.

There is a second reason to write the absence down rather than skip the section, and it decides how the whole report reads. Outcome analysis is the safety net, and a model whose output cannot be tested against any outcome has none. A model with no safety net most needs the other six steps. Where an outcome exists, a weak assumption eventually shows up as a run of misses and somebody notices. Where no outcome exists, nothing ever shows up, and the only thing standing between a bad assumption and a committee paper is the data test, the assumption challenge and the rebuild. Writing that in section 5 is what tells a reader to weigh sections 2, 3 and 4 more heavily than they otherwise would.

Try it out

The outcome of model V1 is only observable over years, so step MV5 produces nothing. What goes in that section of the report?

How is a limitation written so that somebody downstream can act on it?

Step six of seven

MV6 Limitations

Who does it
The validator, writing for the person who runs the model rather than for the person who built it.
Takes in
Everything steps MV2 to MV5 could not settle, including every written absence.
Does
Turns each unsettled item into a sentence telling a user what to do about it.
Hands on
A numbered list of limitations, each phrased as an instruction with a subject and a verb.
Decides
Which limitations are severe enough to become a restriction on use in step MV7.

A limitationA statement of what a model cannot do, written so that a user knows what to do about it rather than merely that it exists. is a statement of what the model cannot do. Almost everybody writes them as caveats, and a caveat is a sentence everybody nods at and nobody acts on. The test is simple and slightly brutal: read the sentence and ask what a person would do differently on Monday morning because of it. If the answer is nothing, it is a caveat.

Take the one from model V1. The caveat version reads that the economic value figure is sensitive to the assumed deposit life. Perfectly true, entirely inert. Nobody's behaviour changes. The instruction version reads that the economic value figure must be reported with the assumed deposit life printed beside it, every time it appears. An instruction can be followed or not followed, and so it can also be checked. A caveat can be neither. Same fact underneath, two completely different documents.

The pattern generalises. A limitation written as an instruction names who does what and when. The output must always be reported alongside a stated figure. The model must not be used beyond a stated boundary. The answer must be recomputed when a stated input changes. Each of those is checkable a month later by somebody who was not in the room. Checkability is the whole trick, and it is the difference between a report that describes a model and a report that changes how the model is used.

Try it out

What is the difference between a limitation written as a caveat and one written as an instruction?

Debt Capital Markets Bootcamp — Fin Maverick

What goes in the validation report, and what changes because of it?

Step seven of seven

MV7 The report

Who does it
The validator writes and signs it. The model owner and the risk function receive it.
Takes in
The outputs of steps MV1 to MV6, in the order the steps produced them.
Does
Assembles seven parts: scope, rated findings, limitations, a use restriction, actions, a signature and a date.
Hands on
One signed document, and one agreed action per finding sitting on a named person's list.
Decides
What the model may and may not be used for from the date the report is signed.

The report is the output of the whole sequence and it has seven parts. Five of them describe. Two of them act. A report that stops after the findings has described a model, and a report carrying a use restriction has changed what somebody is allowed to do on Monday. That is the only real test of whether MV7 was done or merely written.

The use restrictionThe written statement of what the model may and may not be used for, which is what a validation report exists to produce. is the part most often left out. Writing one feels like overreach. It is not. MV7 closes the sequence MV1 opened: MV1 asked which use the model serves, and MV7 answers whether the model is fit for exactly that use and for nothing wider. On model V1, a restriction that follows from the work would say that the economic value output may be used for internal reporting only when it is accompanied by the assumed deposit life, and may not be used to support a statement about the direction of the institution's rate exposure until the assumption has been tested.

SEVEN PARTS, AND TWO OF THEM ACT The two green bands are the parts that change what somebody may do. The other five describe what was found. VALIDATION REPORT: MODEL V1, THE BEHAVIOURAL DEPOSIT LIFE MODEL 1 SCOPE What was validated, for which use, and to what depth. 2 FINDINGS, RATED Every gap found, rated on the institution's own scale. 3 LIMITATIONS What the model cannot do, written as an instruction to a user. 4 USE RESTRICTION What the model may and may not be used for from today. 5 AGREED ACTIONS One action per finding, with a named owner and a date. 6 SIGNATURE Of the validator, who built none of what is checked here. 7 DATE So the next due date can be counted from something fixed. The rating scale, the depth and the cycle are this invented bank's own policy rather than a requirement of any kind. Bands 1 to 3 tell a reader what was found. Bands 4 and 5 tell a named person what to do and by when.
Five of the seven parts describe a model and two of them change what a person is allowed to do with it, which is why a report that stops after the findings has produced a very careful nothing.

What makes a finding usable rather than merely true?

A findingA gap between what was found and what should have been, carrying a cause, an effect and an agreed action with an owner and a date. is a gap between what was found and what should have been. Written properly it has four parts, and this invented bank uses the same four in its audit work, numbered AF1 to AF4, with the agreed action carried as AF5 alongside. The condition is what was found. The criteria is what should have been. The cause is why the gap exists. The effect is what it could lead to.

The cause decides whether a finding can be fixed or only patched, and the cause is the part most often missing. A finding with a condition and an action but no cause can always be closed, because somebody does the thing once and ticks the box, and then it comes back, because nothing changed about why it happened in the first place. In this invented bank, 9 of the 42 control testing findings carry no cause at all.

FOUR PARTS, AND THE THIRD DECIDES WHETHER IT CAN BE FIXED The same four parts this invented bank numbers AF1 to AF4 in its audit work, applied to model V1. AF1 THE CONDITION what was found Model V1, the behavioural deposit life model, has never been validated, and it sets the assumption that decides the sign of this bank's headline economic value figure. AF2 THE CRITERIA what should have been Policy PL6, this bank's own model risk policy, puts every registered model inside a validation cycle of the bank's own choosing. AF3 THE CAUSE why the gap exists Policy PL6 was approved in month 4 and models V1, V2 and V3 all predate it, so a policy that does not reach backwards left three models outside the cycle it created. AF4 THE EFFECT what it could lead to A figure running from minus Rs 840 crore at the assumed 0.5 year life to plus Rs 960 crore at 3.0 years reaches a committee with no note of which assumption put it there. TAKE AF3 AWAY AND THE FINDING CAN ONLY BE PATCHED Validate model V1 without fixing how a new policy reaches the models that predate it, and the next policy approval creates three more. In this invented bank, 9 of the 42 control testing findings carry no cause at all, which is why they can be closed with nothing changing.
Three of the four parts can be written by anybody holding the evidence, and only the cause requires somebody to work out why the gap happened, which is why it is the part that goes missing.
Try it out

A finding says the model has never been validated, and the agreed action is to validate it. What is missing?

What happens when a finding cannot be closed on time?

Some findings cannot be closed inside the date agreed, and a method that has no answer for that leaves its own output rotting on a list. The answer is an escalation, and an escalation is three things written down in advance: who is told, on what timetable, and what they are being asked to decide. Not who is blamed. The decision is usually whether to accept the position for a further period with a dated plan, or to restrict the use of the model until the work is done.

Where the escalation goes is a question about this institution and not about the method, and it is worth naming because the answer is often surprising. In this invented bank, the risk taxonomy places model risk inside operational risk, so the model inventory and its validation status are reported to committee G6, the operational risk management committee. Meanwhile committee G2, the board risk management committee, sets every limit L1 to L12 and reads the numbers those models produce. The body that receives the validation status is not the body that relies on the models, and that placement is a choice the institution made rather than a rule it followed. How committees are chartered and how a decision is minuted belong to the governance material and are covered separately.

One more thing to keep an eye on. Good methods quietly fail here. An open finding has two different measures on it and they are not the same: how old it is, and whether it is past the date agreed for fixing it. An issue can be old and entirely on track, or raised last week and already late. The bank carries 92 open issues aged in five buckets, and 31 of them are past their agreed remediation date, being 33.7 per cent. A function reporting the first measure and managing on the second is doing it in the right order.

Jurisdiction

Who sets the expectation that a model gets checked at all

The mechanism above holds in any jurisdiction: seven steps, a checker who built none of it, a report and a use restriction. Where the expectation comes from is a separate question with two answers. The Basel Committee at the Bank for International Settlements, at bis.org, is the origin of the supervisory expectation that a model used to measure risk is independently reviewed, and of the vocabulary used for validation and for testing a measure against realised outcomes. The binding requirements on an Indian bank, covering model governance, independent review, outsourced models sitting inside a bought package and the information security around them, come from the Reserve Bank of India, at rbi.org.in.

Try it out

The seven steps are about to be run on model V1, and model V1 has never been validated. Which step will produce nothing at all?

What does the whole method look like run end to end on one model?

The sequence now runs on the model the earlier steps have been circling. Vindhya Commercial Bank Limited runs MV1 to MV7 on model V1, the behavioural deposit life model, and model V1 has never been validated at all.

StepWhat it produced on model V1
MV1Scope: model V1, which sets an average behavioural life for the Rs 36,000 crore of current and savings balances and feeds the economic value of equity computation that runs against limit L8. Depth: challenged rather than accepted, for the reason set out under the table.
MV2The model needs a history of account balances and closures by product and by vintage. Which source system does it come from? What does it mean when a balance goes to zero and the account stays open? Is the history long enough to see a full turn of the rate cycle? Four verdicts, one per test, per input.
MV3One assumption: an average behavioural life of 0.5 years. Chosen by committee G4, the asset liability management committee. Evidence behind it: the bank's own review. No realised outcome exists to test it against.
MV4Rebuild from the same inputs. Rate sensitive assets Rs 84,000 crore at a modified duration of 3.00 years, rate sensitive liabilities Rs 84,000 crore at 2.50 years, so the duration gap is 0.50 years, and 0.50 times 2.0 per cent times Rs 84,000 crore is Rs 840 crore, negative under a rise. Compare that with what the system prints.
MV5Nothing, and the section says so in terms. The outcome the model predicts is only observable over years, this bank holds no realised deposit life, and no test against outcomes is possible.
MV6One limitation, written as an instruction: the economic value figure must be reported with the assumed deposit life printed beside it, every time it appears.
MV7One report. A finding rated on the bank's own scale, a use restriction, one agreed action with a named owner and a date, and a signature from Kanaka Murthy, who built nothing she checks.

Two things in that run deserve pulling out, and the first is the one MV1 produced. The register does not place model V1 in tier 1, and the arithmetic makes that certain rather than likely. This bank's register carries 6 models in tier 1, and of those 6, 5 are validated and current and 1 is overdue. Five plus one is six, so the whole of tier 1 is accounted for, and none of it is a model that has never been validated. Model V1 has never been validated. Model V1 is therefore in tier 2 or tier 3, and the record does not say which. The model that decides the sign of this bank's headline interest rate risk number is not carried at its top materiality tier. MV1 exists to notice exactly that and write it down before any depth is accepted from the record.

The second is what MV3 found, and it is worth restating in one line because it is the payoff of the whole run. The range test moved the answer from minus Rs 840 crore at the model's own 0.5 year life, through zero at about 1.67 years, to plus Rs 960 crore at 3.0 years, and across the bank's own repricing bucket RB5 of one to three years alone it runs from minus Rs 480 crore to plus Rs 960 crore. Every one of those readings sits inside limit L8's Rs 990 crore cap. Nothing breaches. The finding is not about size, it is that one untested choice decides whether this institution reports its value exposure as a loss or a gain under the same scenario.

Building a Revenue Forecast From Drivers — free micro-course from Fin Maverick

What does holding the cycle actually cost, and what happens to the queue?

The failure is not that the method is hard. It is that nobody wrote down what it costs.

Every step above is teachable in an afternoon. The eighth thing is in none of them, and the eighth thing is arithmetic rather than craft. Arithmetic is unforgiving. Vindhya Commercial Bank Limited sets itself a twelve month validation cycle. Holding a twelve month cycle on 28 registered models costs 28 validations a year just to keep the current ones current. Not to improve anything. To stand still.

Now the queue. The validation backlogThe models that are overdue or have never been validated, which only shrinks when throughput exceeds the cycle cost. on the register at month 12 counts 9 models. The 6 that are overdue plus the 3 that have never been validated make 9. So a function completing exactly 28 validations a year clears nothing at all, ever. At exactly the cycle rate the backlog of 9 models is permanent: it is 9 at month 12, 9 at month 24 and 9 at month 36. The first year anything changes is the first year throughput goes above 28. At 31 a year the 9 clears in 3.0 years, at 33 a year in 1.8 years, and at 37 a year in 1.0 year.

Then the sweep lands. The bank went and looked, and found 33 models in use against the 28 on the register. On the population that actually exists, the twelve month cycle costs 33 a year rather than 28, and the queue is 14 rather than 9, being the same 6 overdue plus the same 3 never validated plus the 5 that were never registered and so were never in any cycle. At 37 validations a year the same function clears the registered queue of 9 in 1.0 year and the real queue of 14 in 3.5 years. Nobody worked harder or less hard. The denominator moved, and a plan built from the register was out by two and a half years before anybody opened a model file.

The honest last act of the method is therefore a count of what the method costs. A validation plan built from a register nobody has swept is a plan for a different institution.

AT THE CYCLE RATE, THE QUEUE NEVER MOVES Models still waiting, on the 28 on the register, where the twelve month cycle costs 28 validations a year. AT 28 A YEAR, EXACTLY THE CYCLE COST AT 31 A YEAR, THREE ABOVE THE COST MONTH 12, TODAY 9 9 AFTER ONE YEAR 9 6 AFTER TWO YEARS 9 3 AFTER THREE YEARS 9 0, THE QUEUE IS CLEAR AFTER FOUR YEARS 9 0 The queue of 9 is the 6 overdue plus the 3 never validated at month 12. Clearance takes 9 over throughput less 28 years.
Three validations a year of spare capacity is the entire difference between a queue that empties in three years and a queue that is still exactly as long in the fourth year as it was on the day it was counted.
Try it out

The cycle is twelve months on 28 models and the queue is 9. Before the control moves: how long does a function completing exactly 28 validations a year take to clear it?

Play with it

Move the throughput and find the line below which nothing improves

One control: c, the number of validations the function completes in a year, from 20 to 50. Vindhya Commercial Bank Limited locks no capacity figure for its validation function, so c is left open. Two consequences, side by side. On the 28 on the register, the twelve month cycle costs 28 a year and the queue is 9, so clearance takes 9 over c less 28 years: at c of 28 it never clears, at 29 it is 9.0 years, at 31 it is 3.0 years, at 33 it is 1.8 years, at 37 it is 1.0 year and at 40 it is 0.75 years. On the 33 the sweep found, the cycle costs 33 a year and the queue is 14, so clearance takes 14 over c less 33 years: at c of 34 it is 14.0 years, at 37 it is 3.5 years, at 40 it is 2.0 years and at 47 it is 1.0 year. At or below 28 on the register, and at or below 33 on what is actually running, the queue never clears at all and the control shows no number rather than a very large one. The crossing that matters is at c of 37, where the same year of effort clears the registered queue in 1.0 year and the real one in 3.5 years.

20 A YEAR28 A YEAR, EXACTLY THE CYCLE COST ON THE REGISTER50 A YEAR
TWO QUEUES, ONE FUNCTION, AND A LINE IN THE MIDDLE Years to clear, against validations completed in a year. Nothing is plotted below a curve's own cycle cost. CYCLE COST 28 CYCLE COST 33 0 3 6 9 12 15 20 25 30 35 40 45 50 On the 28 registered: cycle 28, queue 9 On the 33 the sweep found: cycle 33, queue 14
Years to clear, on the register
does not clear
Years to clear, on what runs
does not clear

At 28 validations a year, the queue on the 28 registered models does not clear at all, and the queue on the 33 the sweep found does not clear either, because 28 a year is exactly what holding the twelve month cycle on the register costs.

Educational illustration. Invented figures throughout. The twelve month cycle, the 28 registered, the 6 overdue, the 3 never validated and the 33 found by the sweep are Vindhya Commercial Bank Limited's own invented figures and none of them is a requirement of any kind. The throughput is left open: this case fixes no capacity for the validation function. The control also assumes every model costs the same to validate, which is false in any real function, since a model at the top materiality tier plainly costs more to check than one at the bottom. Below a curve's own cycle cost nothing is plotted, because there is no number there rather than a very large one.
ONE YEAR OF WORK, TWO ANSWERS, 2.5 YEARS APART Both bars are the same function completing 37 validations a year. Only the population being counted differs. ON THE 28 REGISTERED cycle 28 a year, queue 9 1.0 YEAR TO CLEAR ON THE 33 IN USE cycle 33 a year, queue 14 3.5 YEARS 2.5 YEARS OF DIFFERENCE, AND ONLY THE DENOMINATOR MOVED 0 1 2 3 4 Years along the bottom. The 5 unregistered models are unvalidated by definition, which is what takes the queue from 9 to 14.
The same year of effort empties one queue and barely dents the other, and the entire gap of two and a half years was created by counting the models that were actually running rather than the ones on the list.
Try it out

At 37 validations a year, why does the queue take 1.0 year to clear on one measure and 3.5 years on another?

A twelve month validation cycle meets a queue measured in decades. See the arithmetic.

Who reads a validation report, and what do they actually do with it?

A method is easier to hold on to once the audience for the document is clear, and it is read by four different people for four different reasons. Each one reads a different part, so they are taken in turn.

The person running the model reads the limitations and the use restriction, and nothing else. The model runner wants to know what changes on Monday: which figure now has to be published with something printed beside it, and which question may no longer be answered with this output. Parts 3 and 4 of the report are written as instructions rather than as observations for exactly that reason.

The head of internal audit, who at this invented bank is Rustom Batliwala reporting to the audit committee, reads the scope and the findings. Audit is not repeating the validation. Audit is testing whether the validation was done. Testing that means reading the scope to see what was promised and the findings to see what came back, and then checking whether the agreed actions are being closed on time rather than merely being closed.

The committee that relies on the number reads the use restriction and one line of the range. A committee does not want the method. The committee wants to know whether the figure in front of it may be used for the decision in front of it. A report whose fourth part is missing sends a committee a very careful document that answers a question nobody at the table asked.

And the fourth reader is the next validator, two years later. The next validator reads the whole thing. The fastest way to run a validation is to start from the last one and test what has changed. The quiet argument for writing an absence down rather than leaving a section blank sits here: the next person needs to know whether the outcome test was skipped or was impossible, and only one of those two is worth attempting again.

The household version of the same document is a note stuck to the fridge. Not a note that says the affordability sum is uncertain. Such a note changes nothing. A note that says the instalment figure was worked out on a year with no wedding in it and must be redone before anybody commits to a seven year loan. Same facts, and one of the two versions actually stops a decision.

Where this guide stops. What model risk is, why an inventory is a claim rather than a list, how tiering is set and who governs a model are covered separately and are used here rather than re-derived. The difference between validating a model and testing one model's output against realised outcomes is a subject of its own, and step MV5 assumes it: this runbook takes the outcome test as a thing that either exists or does not, and hands the comparison itself on. Which person is accountable for a data element, who stewards it and which attributes it should carry belong to the data governance material, so step MV2 tests the data this model consumes and hands the governance of it straight over. The economic value of equity computation, the repricing ladder and the interest rate risk framework belong to the market risk material, and this guide uses their outputs only to show what an assumption does to an answer. How a committee is chartered, who must be told and on what timetable belong to the governance material. Statistical test procedures are covered separately, as are the requirements set by any particular supervisor.

Sources

SourceDocumentSite
Bank for International SettlementsThe Basel Committee vocabulary for models, independent review and validation, and the approach behind testing a measure against realised outcomes, cited as the origin of the expectation rather than as a requirement in Indiabis.org
Reserve Bank of IndiaWhat actually binds a bank in India on model governance, on independent validation, on outsourced models sitting inside a bought package, and on the information security around themrbi.org.in

Vindhya Commercial Bank Limited, Kanaka Murthy and Rustom Batliwala are invented.
Educational material. Not advice on any investment, tax, budget or market position.

Framework

Other frameworks in Risk Reporting, Data and Model Risk

Framework

How to build a KRI Dashboard: Eight Steps to a Usable Page

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.