Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Risk, Treasury & Financial Control
1Risk Foundations
Risk Appetite, Tolerance, Capacity…The Risk Taxonomy and UniverseRisk Register vs Risk MatrixStress TestingScenario Analysis vs Stress TestingImpact and LikelihoodLikelihoodThe Risk EventRisk Assessment
2Enterprise Risk Management
Enterprise Risk ManagementThe Four Risk TreatmentsRisk CultureRisk MaturityRisk Monitoring
3Risk Governance
Risk GovernanceHow to set a…The Risk PolicyThe Risk OwnerThe Risk Committee and Its CharterThe Risk Limit FrameworkRisk EscalationHow to set a…
4Credit and Counterparty Risk
Collateral AgreementsCollateral vs NettingProbability of DefaultExposureCounterparty ExposureConcentration Risk vs Wrong Way RiskCounterparty Risk vs Credit RiskHow to assess Counterparty ExposureHow to assess Concentration Risk
5Market Risk
Market RiskSensitivity MeasuresThe Hedging PolicyInterest Rate Risk in the Banking BookIRRBB vs Market RiskExpected ShortfallEconomic Value of EquityVaR BacktestingOpen PositionValue at RiskValue at Risk and Expected ShortfallEconomic Value SensitivityFX ExposureValue at Risk vs Expected ShortfallEarnings at Risk vs…FX Transaction Risk vs…How to measure Interest…How to measure Foreign…
6Liquidity Risk
Liquidity Stress TestingLiquidity Gap vs Liquidity BufferMaturity MismatchThe Debt Maturity ProfileFunding ConcentrationSurvival HorizonThe Contingency Funding PlanNet Stable Funding RatioLiquidity Risk vs Funding RiskLiquidity Coverage RatioLiquidity Gap and BufferHow to run a Liquidity Gap Analysis
7Operational Risk
Operational LossThe Loss EventRisk and Control Self AssessmentException ManagementInformation Security as a…Segregation of DutiesIssue ManagementThe Near MissRoot Cause Analysis in RiskThe Fraud TriangleCyber Risk vs Third Party RiskHow to run a…How to assess Third…
8Risk Reporting, Data and Model Risk
Model RiskModel Validation vs BacktestingHow to run Model ValidationData Governance in RiskModel Risk vs Data RiskKey Risk IndicatorsManagement InformationRisk ReportingRisk ScoreEarnings at RiskRisk Adjusted ReturnEarly Warning IndicatorsHow to build a KRI Dashboard
9Treasury
Corporate TreasuryAsset Liability ManagementIntragroup FundingThe Treasury PolicyThe Treasury Management SystemThe Cash ForecastCash Pooling and ConcentrationHow to build a Cash Forecast
10Financial Controls and Assurance
Control AssuranceThe Control LifecycleThe Assurance MapThe Audit FindingIssue RemediationInternal Financial ControlsControl Design vs Control EffectivenessHow to map Internal Financial ControlsHow to test Control…Control DeficiencyMaterial Weakness
11Operational Resilience
Operational ResilienceBusiness Continuity and Disaster RecoveryBusiness Continuity vs Operational…Crisis ManagementDisaster RecoveryIncident Management

Model Risk: Governance, Performance Monitoring and What Goes Wrong

Model risk is what attaches to the measuring rather than to the thing measured: a rule that gives a wrong number, or a right number used for something it was never built for. Model risk is held by an inventory, a tier against each model, a cycle, and a checker who built none of it. At the invented Vindhya Commercial Bank Limited, 28 models are on the register and a sweep found 33 running.

Everything here rests on one distinction, and holding it from the first paragraph makes the rest follow. A market risk model estimates market risk. Model risk is not market risk. Model risk is the risk attached to the estimate itself: that the rule is wrong, that the data feeding it is wrong, that it was built for one purpose and is being used for another, or that nobody has ever checked. A wrong number and a right number arrive in the same report, in the same font, on the same day of the month. So a bank can measure its exposures perfectly well with a broken rule and never find out.

What is model risk, and how is it different from the risk the model measures?

Start with the word. A modelA rule that turns data into an answer somebody relies on, whether it lives in a system, a bought package or a spreadsheet on one desk. here is not a piece of mathematics and not a piece of software. A model is a rule that turns data into an answer somebody relies on. The definition is deliberately wide, and the width is the point. A wide definition catches the rule running inside a purchased package, the rule written in code by a team of four, and the spreadsheet on one person's desk that three departments quietly depend on. If an answer somebody acts on comes out of a rule applied to data, it is a model, whatever it is stored in and whoever built it. The moment the definition is narrowed to things that look technical, half the rules a bank runs on stop being governed, and they are usually the half nobody wrote down.

The everyday version runs like this. A household decides how much it can afford as a monthly loan instalment by taking take-home pay, subtracting what it usually spends, and keeping half of what is left. The household's arithmetic is a rule turning data into an answer somebody relies on, so it is a model. Nobody in the household calls it one. The household rule carries exactly the two failures a bank's models carry. The rule may be wrong: what the household usually spends was measured in a year with no wedding in it. Or the answer may be right and used wrongly: a figure worked out for a two year loan gets reused for a seven year one, and halfway through the seven years the medical costs of an ageing parent arrive. Neither failure announces itself. Both produce a confident number.

Model riskThe risk that the rule gives a wrong answer, or gives a right answer that is then used for something the rule does not support. is that pair. The two halves are found in different ways and fixed by different people, so they are worth separating properly. The first half is a wrong rule: stale data, an assumption nobody tested, a specification that says the wrong thing, an implementation that faithfully does what the specification says. The second half is a right rule put to a use it does not support. Nothing in the machinery is wrong at all, and that is exactly what makes the second half harder to find. The number is correct. The number simply answers a different question from the one being asked of it.

A MODEL IS A RULE THAT TURNS DATA INTO AN ANSWER SOMEBODY RELIES ON Model risk attaches to this chain and not to whatever the chain is measuring. DATA THE RULE AN ANSWER SOMEBODY USES HALF ONE: THE RULE ITSELF IS WRONG Stale data, an assumption nobody tested, a specification that says the wrong thing, or code that does exactly what a wrong specification told it to. The number still arrives on time and still looks like a right one. IN THIS INVENTED BANK Model V1 sets an average life of 0.5 years on the Rs 36,000 crore of current and savings balances, and nobody has ever checked that rule. HALF TWO: THE RULE IS RIGHT, THE USE IS NOT The rule answers the question it was built for, and then somebody puts that answer to a second question it was never built for. Nothing in the output objects, because an answer carries no record of its own purpose. IN THIS INVENTED BANK One 200 basis point rise adds Rs 156 crore to income and takes Rs 840 crore off economic value. Both are right. Either one alone reads as the whole position.
Data through a rule to an answer is the whole chain model risk attaches to, and it breaks in exactly two places: the rule can be wrong, as with the unchecked deposit life assumption in this invented bank, or the rule can be right while its answer is carried across to a question it does not settle.

The half people miss is in the right hand panel. Look at it again. The invented bank's own scenario of a 200 basis point rise adds Rs 156 crore to net interest income over twelve months and takes Rs 840 crore off economic value. Neither figure is a mistake. The two figures settle different questions, one about earnings over a year and one about value today, and how either is computed is set out with the market risk material. Only one fact about the pair matters at this point: a committee handed one of the two, with no statement of which question it settles, has been handed a right number doing a wrong job. Model risk does not require anybody to be wrong about anything; it only requires an answer to travel further than the rule that made it.

Try it out

A market risk model estimates market risk. What does model risk attach to?

Risk Management Program Bootcamp — Fin Maverick

How to create a Model Inventory: where does a row actually come from?

A model inventoryThe list of every rule in use, with a tier, an owner, a validation date and a use restriction against each. is the object that turns model risk from a worry into something that can be counted. The inventory is a table with six columns, and the six are not arbitrary. A specific question gets asked about a model, somebody has to answer it without calling a meeting, and each column holds one of those answers. The name, so two people arguing about a model are arguing about the same rule. The tier, so the amount of checking is decided in advance rather than by whoever shouts. The owner, so there is one person accountable for the model being fit. The last validation date, so an interval can be measured. The next due date, so being late is a fact rather than an opinion. And the use restriction, so a right answer has a written boundary around it.

Now the part that decides whether the whole exercise is worth anything. A system extract can only ever return what the systems already know about, so the list does not start with one. Beginning by asking the technology function for a list of models returns the models that live in things the technology function administers, and every rule running in a spreadsheet on somebody's desk stays invisible, along with every rule embedded in a package the bank bought and never opened. The invisible rules are exactly the ones with no specification, no owner and no cycle. Starting from the extract guarantees that the well governed ones are found and the rest are missed.

So the list starts with a question put to people rather than to systems, and the question is an awkward one: what number does this desk produce, or receive, that somebody acts on, and what rule turns the inputs into it? Asked of every desk, it returns answers a system extract never could. In this invented bank one of the answers that came back was a spreadsheet scorecard maintained by one person. The scorecard is model V3, it has no written specification, and no extract from any system would ever have shown it.

SIX COLUMNS, AND WHAT EACH ONE IS ACTUALLY FOR Every column exists because a question gets asked about a model and somebody has to answer it without calling a meeting. MODEL NAME What the rule is called, so that two people naming it mean one rule and not two. TIER How material the bank says it is. This decides how much checking it gets, in advance. OWNER One named person accountable for the rule being fit for what it is used for. LAST VALIDATED The date somebody independent last checked it. A blank cell is an answer and not a gap. NEXT DUE Last validated plus the cycle. Once that date passes, being late is a fact rather than a view. USE RESTRICTION What the rule may and may not be used for, in writing. The only column that stops a right answer being used wrongly. THE 28 ROWS, AND THE FIVE THINGS THAT SIT ON NO ROW THE REGISTER AT MONTH 12: 28 ROWS validated and current 19 overdue against the cycle 6 never validated at all 3 + ON NO ROW AT ALL: 5 A month 12 sweep looked for rules actually running rather than for rules already on the list, and found 33. Five had no row, so no tier, no owner, no cycle and no restriction. 28 registered against 33 found in use makes inventory completeness 28 over 33, being 84.8 per cent. That 84.8 per cent is model inventory completeness and not limit L8 utilisation, which is a separate 84.8 per cent here.
Six columns exist because six questions get asked about every rule, and the last of them is the only one that writes a boundary around a right answer, while the dashed box holds five rules that were running with none of the six filled in.

How does tiering by materiality decide how much checking a model gets?

Twenty eight models cannot all be checked to the same depth, and pretending otherwise produces a policy nobody follows. So each row carries a model tierA materiality rating that decides how much validation and monitoring a model gets, set by the institution itself.. A tier is a materiality rating the institution sets for itself. Vindhya Commercial Bank runs three tiers on its own test. At month 12 the split is 6 models in tier 1, the tier it calls high materiality, 13 in tier 2 and 9 in tier 3. Six plus thirteen plus nine is twenty eight, and it ties.

Tiering is a governance decision and not a technical one. The numbers 6, 13 and 9 therefore describe the bank rather than the models. Nothing in the mathematics of a rule says whether it is material. Materiality is a statement about consequence: how much money moves if this rule is wrong, how many decisions rest on it, how visible its output is, whether it feeds a limit or a capital number or a committee paper. Two banks with identical models are each answering a question about themselves, so they will tier them differently and both can be right.

Try it out

Of the 6 models this bank puts in tier 1 and calls most material, how many are validated and current?

Against that tiering sits the second split, and it is the one people actually read: status against the bank's own twelve month validation cycleThe stated interval at which each model must be checked again, after which it counts as overdue.. At month 12, 19 of the 28 are validated and current, 6 are overdue and 3 have never been validated at all. Nineteen plus six plus three is twenty eight, and that ties as well. And of the 6 models in tier 1, 5 are current and 1 is overdue. Tier 1 therefore runs at 83.3 per cent current against 67.9 per cent across the whole register.

The comparison is a compliment to the bank, and it is easy to misread as the opposite. Tiering worked. The bank said it would put its effort where materiality was highest, and on its own numbers it did: the tier it calls most material is in better shape than the register as a whole. The problem that remains is not inside the tier 1 population at all. The problem sits somewhere the tiering never reached. The models in question were never on the register at all, so there was nothing there for the tiering to reach.

TWENTY EIGHT ROWS, COUNTED TWO DIFFERENT WAYS Both counts are the invented bank's own and both add to twenty eight. TIERED ON THE BANK'S OWN MATERIALITY TEST TIER 1, HIGH MATERIALITY: 6 TIER 2: 13 TIER 3: 9 6 + 13 + 9 = 28 AGAINST THE BANK'S OWN TWELVE MONTH VALIDATION CYCLE VALIDATED AND CURRENT: 19 OVERDUE: 6 NEVER VALIDATED: 3 19 + 6 + 3 = 28 THE ONE PLACE THE TWO COUNTS ARE JOINED Of the 6 models in tier 1, 5 are validated and current and 1 is overdue, being 83.3 per cent. Across all 28 rows the current share is 67.9 per cent, so the most material tier is in better shape than the register as a whole. The two rows are drawn separately on purpose. This bank's own record fixes the tier split, the status split and the tier 1 detail. It does not fix how the remaining overdue and never validated rows fall across tiers 2 and 3, so no tile here is coloured by both.
The same twenty eight rows counted by tier give six, thirteen and nine, and counted by status give nineteen, six and three, and the only cell where the invented bank joins the two counts is tier one, where five are current and one is overdue.
Derivatives Foundation Bootcamp — Fin Maverick

What is Model Governance, and who may build, check and run a model?

Model Governance sets out the roles around a rule, and it comes down to one uncomfortable requirement: the person who checks a model must not be the person who built it. Everything else is detail hanging off that. There are four roles, and they are four different people. The owner is accountable for the model being fit for the use it is put to, a business accountability and not a technical one. The developer built the rule and wrote down its specification, its data and its limitations. The independent validatorSomebody who checks a model, did not build it, and does not report to the person who did. checks it and reports somewhere other than to the builder. The user runs it and reads its output, inside a written boundary.

The everyday version is a shopkeeper who keeps his own books and also does his own audit. Nothing about him is dishonest. The same reading of the same facts produced the entry and then approved it, and a second look by the same eyes is not a second look. He simply cannot see the mistake. Independence is not a statement about anybody's character; it is a statement about whether two different readings of the same thing have actually happened. That is why the validator's reporting line matters at least as much as the validator's skill: a checker who reports to the person whose work is being checked has only one reading in the room, however good she is.

FOUR ROLES AROUND ONE RULE, AND WHAT EACH ONE MAY NOT DO The lower strip on each box is the part that does the work: a role is defined as much by its bar as by its job. THE OWNER Accountable for the rule being fit for the use it is put to. A business job and not a technical one. MAY NOT pass that accountability to whoever built it. THE DEVELOPER Built the rule, and wrote down its specification, its data and the places where it stops working. MAY NOT sign off their own build as independently checked. THE VALIDATOR Checks the rule, built none of it, and reports somewhere other than to the person who did. MAY NOT have built any model on the list she checks. THE USER Runs the rule and reads its output, inside the written boundary the restriction column sets. MAY NOT carry an answer outside that written boundary. IN THIS INVENTED BANK The independent validator is Kanaka Murthy. She sits in the risk function under chief risk officer Sunanda Ravikumar, and she built no model that she checks. Four roles, and they have to be four different people, because a person checking their own rule is not checking it.
Four roles sit around every rule and each is defined as much by what it may not do as by what it does, and in this invented bank the checking role belongs to somebody who built none of the models she looks at.

Why does policy PL6 leave three models outside the cycle it requires?

Vindhya Commercial Bank holds nine policies numbered PL1 to PL9, and PL6 is the model risk policy. PL6 was approved in month 4. Models V1, V2 and V3 were all running before that. And there, in one dull sentence about dates, is the whole explanation for three models that have never been validated. A policy binds practice forward from the day it is approved, and it does not reach backwards on its own.

The order of those two dates matters more than it looks, and the reason is how a backlog is normally read. Three models with no validation date at all reads as negligence: somebody had a job and did not do it. The negligence reading is wrong here, and getting it wrong sends somebody looking for the wrong fix. Nobody skipped these three. When they were built there was no cycle to be inside, no tier to be given, no report to write and no validator to send it to. The three models became overdue retrospectively, when a policy arrived and drew a line they were already standing on the wrong side of.

The fix is therefore not disciplinary and it is not a reminder email. The fix is a body of work: somebody has to go back over what existed before month 4, decide which of it is a model, register it, tier it, and put it into the cycle. The work costs validator time already committed to the twenty five models the cycle keeps busy. Backlogs of this shape therefore sit still for a long time.

A POLICY APPROVED IN MONTH 4, AND THREE MODELS THAT WERE ALREADY RUNNING The months are numbered rather than dated, and every figure on this line belongs to the invented bank. ALREADY RUNNING V1 deposit life V2 collateral haircut V3 spreadsheet scorecard PL6, THE MODEL RISK POLICY approved in month 4 THE MONTH 12 SWEEP found 33 rules in use 123456789101112 before month 1 PL6 BINDS PRACTICE FROM MONTH 4 FORWARD AND DOES NOT REACH BACKWARDS V1, V2 and V3 therefore sit outside a validation cycle that a policy now requires. That is a fact about dates rather than three people failing to do a job.
The three models with no validation date were already running before the policy that requires validation was approved in month four, so the backlog is produced by the order of two dates rather than by anybody neglecting a task.
Breaking Into Quants Bootcamp — Fin Maverick

How is Model Performance watched between one validation and the next?

Validation is periodic. Model Performance monitoring is continuous, and the difference is not a detail of scheduling. A validation is a deep look at one model at one moment, and on a twelve month cycle it happens once a year. For the other three hundred and sixty four days nobody is looking at the model at that depth. Performance monitoring is the shallow look that happens constantly in between. The two are not competing versions of the same activity: monitoring exists to show when the annual cycle is no longer fast enough for a particular model.

Monitoring watches anything that would move before a validation would catch it. The inputs first: a feed going stale or a field arriving empty is then noticed on the day rather than at the next review. The outputs next: a number moving in a way the model's own history does not support raises a question on the day it moves. Where the model can be tested against what actually happened, that comparison, run continuously. And the population last: a model built on one kind of customer and now being run on another is caught by the change in what goes into it rather than by an argument about the answer coming out.

Think of a delivery van. The annual inspection is the validation: thorough, scheduled, and done by somebody who did not drive the van. The temperature gauge on the dashboard is the monitoring: shallow, constant, and useless as a substitute for the inspection. Neither replaces the other, and a fleet that runs only on annual inspections finds out about an overheating engine eleven months late. The reason a bank sets a trigger level on the gauge is precisely so that the gauge can call for an inspection early rather than waiting for the calendar.

What is a Validation Threshold, and where do its trigger levels come from?

A Validation Threshold is the level on that gauge. The threshold is a performance level the institution writes down in advance, and crossing it starts something. Vindhya Commercial Bank sets two of them on the model behind its trading book value at risk, counted over 250 observation days: at five exceptions the matter is escalated to committee G7, the market risk committee, and at seven a model review is required. In the twelve months of this case, seven were observed, numbered X1 to X7, so both of the bank's own levels were reached.

Every word of that is the bank's own. Neither five nor seven is a requirement, a supervisory band or a zone boundary. The Basel Committee at the Bank for International Settlements, at bis.org, publishes the approach that produces the idea of counting exceptions against a measure at all, and how a supervisor in India treats such a count comes from the Reserve Bank of India at rbi.org.in. A number written in a policy and a number written in a regulation look identical once they are copied into a slide, so reading a case figure as a rule is a serious error.

Now the distinction that actually costs people marks. A validation threshold is not a limit. A limit is set on a position and says how much of something the bank will carry. A threshold is set on a model's performance and says how well the measuring has to work before somebody has to look again. In this bank limit L5 caps the trading book value at risk at Rs 18.0 crore, and on each of the seven exception days the measured reading sat between Rs 14.8 crore and Rs 16.2 crore, so the limit was not crossed on any of those days. The threshold was crossed twice over. One object was breached and the other was not, on the same days, on the same book. The threshold and the limit measure two different things.

A VALIDATION THRESHOLD, ON A MODEL Counted over 250 observation days. Both levels below are the bank's own. REQUIRE A MODEL REVIEW at 7 exceptions ESCALATE TO G7 at 5 exceptions 024681012 OBSERVED: 7, BEING X1 TO X7 Both trigger levels are this invented bank's own policy. No supervisory band or zone boundary is stated anywhere here. A LIMIT, ON A POSITION Trading book value at risk in Rs crore, on the same days. LIMIT L5: Rs 18.0 CRORE the cap on the trading book measure 12.014.016.018.020.0 the seven readings measured on those days, Rs 14.8 to 16.2 crore This scale starts at Rs 12.0 crore and not at zero. None of the seven readings crossed limit L5 on those days. A threshold is set on a model's performance and a limit on a position, and crossing one starts a different thing from crossing the other.
Counting how often a measure was wrong and measuring how large a position is are two separate readings taken on the same book on the same days, and here the first went past both of the bank's own levels while the second stayed inside its cap.
Try it out

Where do the trigger levels of five exceptions and seven come from?

What does a Documentation Standard require before a model can be checked?

A Documentation Standard is the list of things a model file must hold before anybody outside the build team can understand or check the model. A Documentation Standard is not paperwork for its own sake, and one brutally simple test says whether a standard is any good: could a competent person who has never met the developer pick the file up and work out what the rule was built to do, what it was built on, and where it stops working? If not, the file fails, whatever it contains.

In practice the file holds a statement of purpose: what question the model answers, and for whom. The specification, meaning the rule itself written down in a form somebody can check rather than only run. The data: what goes in, from where, in what condition, and what happens when a field is empty. The assumptions, every one of them, stated rather than embedded. The implementation, in enough detail to compare the code against the specification. The testing already done. The known limitations. And the use restriction, the boundary the whole file exists to make enforceable.

The reason this matters more than being late is that documentation is what makes validation possible at all. Model V3 in this bank is a spreadsheet early warning scorecard maintained by one person, and it carries no documented specification. Having nothing written down is a worse position than being overdue, and the difference is worth sitting with. An overdue model has a specification, a history and a previous report, so a validator can pick it up tomorrow. A model with nothing written down has no statement of what it was supposed to do, so there is nothing to test it against. A validator would have to reconstruct the intent by asking the person who maintains it, and a check built out of the builder's account of their own intentions is not independent of anything.

Try it out

Model V3 is a spreadsheet scorecard with no documented specification. Why is that worse than being overdue for validation?

What does a Validation Report carry, and what does its use restriction do?

A Validation Report is the written output of a validation, and it is the artefact everything else in this guide eventually points at. A Validation Report carries eight things: the scope, meaning what was checked and, just as importantly, what was not; the data, and what condition it was found in; the assumptions, each one tested rather than repeated; the implementation, meaning whether the code does what the specification says; the outcomes, being how the model performed against whatever evidence exists; the limitations, being where it stops being reliable; the use restrictionA written statement of what a model may and may not be used for, which is how a right answer stops being used wrongly.; and a signature, from a named person, on a date.

Seven of those eight are a description of the past. The use restriction is the only one that governs the future, and it is the sentence that connects back to the second half of model risk. If a model is fit for one purpose and not another, the report is where somebody writes that down, and the inventory column carries it forward so that the next person to pick up the output reads the boundary before reading the number. A validation with no use restriction has produced an opinion about a model and no instruction to anybody who uses it.

The scope line deserves the same attention and rarely gets it. A validation report that says what was checked has stated only half of what matters. The other half is what was left out and why. A reader who does not know the scope will treat a partial check as a whole one. Reading a partial check as a whole one is the same failure as the second half of model risk, one level up: a right answer about a narrow question, read as an answer to a wide one.

THE VALIDATION REPORT: EIGHT PARTS 1SCOPEwhat was checked, and what was not 2DATAwhere it came from, in what condition 3ASSUMPTIONSevery one of them, stated and tested 4IMPLEMENTATIONwhether the code does what the file says 5OUTCOMEShow it performed against the evidence 6LIMITATIONSwhere it stops being reliable 7USE RESTRICTIONwhat it may and may not be used for 8SIGNATUREa named person, on a date MODEL V3'S FILE V3 is a spreadsheet early warning scorecard maintained by one person, with no documented specification. An overdue model has a specification, a history and a previous report, and somebody can pick it up tomorrow. This one has none of that. NOTHING WRITTEN DOWN TO VALIDATE IT AGAINST The use restriction is the line on this report that governs everybody who later reads the model's output.
Seven parts of a validation report describe what was found and only the eighth governs what happens next, which is why a model with nothing written down cannot be checked at all rather than merely being checked late.

What does Governance Escalation do when a threshold is crossed?

Governance Escalation is what happens after a level is crossed: who is told, in what form, and what they are being asked to decide. The mechanics are ordinary. The performance reading crosses the level the policy wrote down. The model owner is told, and the validator is told. A paper goes to the committee named in the policy, carrying what crossed, by how much, since when, what the owner proposes and by when. The committee either accepts the position with a dated plan or refuses it and requires something. And the model's use restriction may be tightened in the meantime. Tightening the restriction is the fastest control available, and it does not require anybody to fix the model first.

The awkward part in this bank is not the mechanics. The destination is the awkward part. Model risk sits under TX4 operational risk at level two in this bank's own taxonomy, a choice this bank made rather than a law. The placement decides who ever sees the model inventory: committee G6, the operational risk management committee. Meanwhile committee G2, the board risk management committee, sets every limit L1 to L12 and accepts or refuses every breach, and it reads the numbers those models produce. The body that receives the model inventory is not the body that relies on the models, and no rule was broken to arrive at that.

WHERE MODEL RISK SITS DECIDES WHO EVER SEES IT The seven level one categories are this invented bank's own list, and so is the placement below them. TX1 credit TX2 market TX3 liquidity, funding TX4 operational TX5 compliance, conduct TX6 strategic, business TX7 reputational LEVEL TWO MODEL RISK SITS HERE That placement is a choice this bank made and not a law, and everything below it follows from the choice. WHERE THE MODEL INVENTORY IS REPORTED COMMITTEE G6, operational risk management It receives incidents, near misses and the self assessment, and the model inventory arrives alongside them. WHO RELIES ON WHAT THE MODELS PRODUCE COMMITTEE G2, board risk management It sets every limit L1 to L12 and accepts or refuses every breach, and it never sees the model inventory. NO LINE So the body that receives the model inventory is not the body that relies on what the models produce. Nobody broke a rule to arrive here. A placement decision made in a taxonomy decided who would ever be told.
Putting model risk underneath operational risk in a taxonomy quietly routes the inventory to one committee while the committee that sets the limits those models answer never receives it, and that routing was decided by a filing choice.
Try it out

Model risk sits under TX4 in this bank's taxonomy, so the model inventory goes to committee G6. Why is that awkward?

How complete is the inventory, and what does a sweep actually test?

Everything so far has been about the rows on the register. Now the harder question: is the register the right list? At month 12 this bank ran a sweepAn exercise that goes looking for rules actually running, rather than reading the list of rules already recorded.. A sweep goes looking for rules actually running rather than reading the list of rules already recorded. The sweep found 33 in use against the 28 registered. So inventory completenessModels registered divided by models found in use, which is a test of the list rather than of any model on it. is 28 over 33, being 84.8 per cent, and five rules were running with no row. The 84.8 per cent is model inventory completeness, and it is not limit L8's utilisation. Limit L8's utilisation is a separate 84.8 per cent in this same bank.

The purpose of a sweep is the idea this guide is built around. An inventory is not a list of what was built; it is a claim about what is running, and a list can only be tested by going and looking. Every row on the register was put there by somebody who knew about a model. The rules that are missing are missing for one reason: nobody who knew about them was ever asked. A list has no way of testifying about what is not on it, so no amount of careful reading of the register can surface them. A sweep is the only instrument that tests the claim, and it works by asking people what numbers they produce rather than asking systems what they contain.

The household version is a home inventory for insurance. The household lists what it knows it has, and the list is accurate row by row. Then somebody walks through the house opening cupboards, and finds five things nobody thought to list. Not one row on the list was wrong. The list was simply an account of what was remembered, and the walk through was an account of what is there.

Try it out

Somebody says a model inventory is just a list. What is the sentence that corrects them?

Ratio Analysis That Says Something — free micro-course from Fin Maverick

Why do two true validation percentages disagree with each other?

Here is where a sweep stops being a tidying exercise and starts changing what the bank knows about itself. The register reports validation status on 28. The sweep says the population is 33. And an unregistered model was never in any cycle to be validated by, so by definition it has never been validated. So every status percentage has to be recomputed on the population that actually exists.

MeasureReported, on the 28 registeredActual, on the 33 in useGap
Validated and current19 of 28, being 67.9 per cent19 of 33, being 57.6 per cent10.3 points
Overdue6 of 28, being 21.4 per cent6 of 33, being 18.2 per cent3.2 points
Never validated3 of 28, being 10.7 per cent8 of 33, being 24.2 per cent13.5 points
Population28 rows on the register33 rules found running5 unregistered

Nothing about any of the 19 changed. Nobody found a fault in a single model. The numerator sat perfectly still and the denominator grew by five, and two reported figures moved by ten and thirteen percentage points. The register was not lying; it was answering a question about itself. Both columns are true, and each is the answer to a different question: the first asks how well the register is being maintained, and the second asks how much of this bank's actual modelling has ever been checked. Only the second is about the bank.

The shape is recognisable, and this bank already has the same fault elsewhere. Its self assessment rated 196 controls effective out of 214, or 91.6 per cent. Independent testing found 172 effective out of the 198 it tested, or 86.9 per cent. Putting 91.6 beside 86.9 is wrong: the two sit on different denominators. Like for like on the full 214 it is 91.6 per cent against 80.4 per cent. Same shape, different record. A percentage that looks stable while its denominator is unsettled is one of the most reliable ways for an institution to mislead itself without anybody writing down anything false.

ONE NUMERATOR, TWO DENOMINATORS, AND NOT ONE FALSE FIGURE Each block below is one model, drawn to the same width in both bars, so the second bar is longer because the population is. AS REPORTED, MEASURED ON THE 28 ON THE REGISTER 19 6 3 19 of 28 validated and current = 67.9 per cent 3 of 28 never validated = 10.7 per cent THE 5 THE SWEEP FOUND AS IT IS, MEASURED ON THE 33 ACTUALLY RUNNING 19 6 3 5 19 of 33 validated and current = 57.6 per cent 3 on the register plus 5 unregistered gives 8 of 33 never validated, being 24.2 per cent THE TWO GAPS, AND NOT ONE FIGURE IN EITHER BAR IS FALSE Validated and current falls by 10.3 percentage points; never validated rises by 13.5 percentage points. The register was not lying. It was answering a question about itself.
The same nineteen validated models are nineteen in both bars, and two reported shares still move by ten and thirteen percentage points, because the only thing that changed was how many models the count was measured against.
Try it out

A colleague reports that 3 models, being 10.7 per cent, have never been validated. What is wrong with that sentence after the sweep?

Try it out

The register says 19 of 28 models are validated and current, or 67.9 per cent. A sweep then finds 33 in use. Before reaching for the control below: what happens to that 67.9 per cent?

Play with it

Move the sweep and watch the register stop being the population

One control: s, the number of models a sweep finds in use, from 28 to 40, against the 28 on the register. Four consequences: inventory completeness, 28 over s; the count running unregistered, s less 28; the share of what is running that has never been validated, being 3 plus s less 28 over s; and the share currently validated, 19 over s. The solved points are these. At s of 28 the inventory is 100.0 per cent complete, 0 unregistered, 10.7 per cent never validated and 67.9 per cent current. At 30 it is 93.3 per cent, 2, 16.7 per cent and 63.3 per cent. At 33, the count this bank's sweep actually found, it is 84.8 per cent, 5, 24.2 per cent and 57.6 per cent. The 84.8 per cent there is model inventory completeness and not limit L8's utilisation, a different 84.8 per cent in this bank. At 36 it is 77.8 per cent, 8, 30.6 per cent and 52.8 per cent. At 40 it is 70.0 per cent, 12, 37.5 per cent and 47.5 per cent. The share currently validated falls to exactly half of what is running at s of 38.

28 FOUND, THE REGISTER IS THE POPULATION33 FOUND40 FOUND
WHAT IS ACTUALLY RUNNING, AS THE SWEEP FINDS MORE REGISTERED: 28 19 6 3 5 IN USE, FOUND BY THE SWEEP: 33 19 validated and current 6 overdue 3 never validated, on the register unregistered, so never validated THE TWO SHARES AS THE SWEEP FINDS MORE 100 75 50 25 0 PER CENT 28303234363840 MODELS FOUND IN USE BY THE SWEEP at 38 the current share is exactly half inventory completeness, 28 over s never validated, share of what is running validated and current, share running
Completeness
84.8 per cent
Unregistered
5
Never validated
24.2 per cent
Currently validated
57.6 per cent

With 33 models in use against 28 registered, the inventory is 84.8 per cent complete, 5 models are running unregistered, and 24.2 per cent of what is actually running has never been validated.

Educational illustration. Invented figures throughout. The 28 registered, the 19, 6 and 3 status split, the 6, 13 and 9 tiering and the 33 found are Vindhya Commercial Bank Limited's own invented figures and none of them is a requirement. Anything above 33 on the control is a setting chosen at the control and is not a figure from the case. The control assumes every unregistered model is unvalidated, which is true by definition rather than by assumption. Completeness falls slightly faster than the never validated share rises, in the fixed ratio 28 to 25, because one is driven by the 28 registered and the other by the 25 that have been validated at least once. The never validated share would reach half only at 50 found, which is 22 unregistered models against 28 registered.
Ratio Analysis That Says Something teaches you to choose ratios that answer a question rather than fill a template.

Which three models have never been validated, and why is the first one the payoff?

The failure is not that three models are late. It is which three.

The three carry the bank's own numbering V1 to V3, and reading them in order matters. Model V1 is the behavioural deposit life model, and it sets the assumption that decides the sign of this bank's headline interest rate risk number. It gives the Rs 36,000 crore of current and savings balances an average life of 0.5 years. At that life the economic value change under the bank's own 200 basis point scenario is minus Rs 840 crore. At 2.0 years, the midpoint of the bank's own repricing bucket RB5 where the same Rs 36,000 crore is slotted, it is plus Rs 240 crore. The model that decides whether the answer is negative or positive has never been checked by anybody.

Model V2 is the collateral haircut model. Model V2 sets the 35.0 per cent haircut that turns a charge valued at Rs 1,440 crore into Rs 936 crore of eligible collateral on counterparty C1, and that in turn moves C1's expected loss from Rs 3.89 crore uncollateralised to Rs 2.20 crore collateralised. Real money, and a smaller number than V1 moves. Model V3 is the spreadsheet early warning scorecard maintained by one person with no documented specification. Nothing is written down to validate it against, so it cannot be validated even if somebody wanted to.

And the reason all three sit outside the cycle is the dull one from earlier: policy PL6 was approved in month 4 and all three predate it. Nobody skipped anything. Ranking a validation backlog by tier is normal, and ranking it by what the model actually decides is better.

ONE UNVALIDATED ASSUMPTION, AND THE HEADLINE NUMBER CHANGES SIGN Economic value change in Rs crore under this bank's own 200 basis point scenario. Every figure is the invented bank's own. MODEL V1 SETS 0.5 YEARS change: minus Rs 840 crore 84.8 per cent of limit L8, and that 84.8 is L8 utilisation, not completeness BUCKET RB5 MIDPOINT: 2.0 YEARS change: plus Rs 240 crore 24.2 per cent of limit L8, and that 24.2 is L8 utilisation, not a validation share zero OUTSIDE L8 OUTSIDE L8 INSIDE LIMIT L8, WHICH CAPS THE MOVE AT Rs 990 CRORE EITHER WAY A MOVEMENT OF Rs 1,080 CRORE, AND A CHANGE OF SIGN One assumption, never checked by anybody, and the answer moves from minus Rs 840 crore to plus Rs 240 crore. No number anywhere in the bank's reporting shows that the assumption is what moved it.
A single unchecked assumption about how long a deposit stays moves this bank's headline value reading across zero by one thousand and eighty crore, and both readings sit comfortably inside the same limit, so nothing on any report looks unusual at either end.

Two numbers in this guide are literally the same fraction wearing two meanings, and this bank produces both of them. The pair is worth pausing on together. Rs 840 crore against limit L8's cap of Rs 990 crore is 84.8 per cent, and that is limit L8 utilisation. Twenty eight registered against 33 in use is also 84.8 per cent, and that is model inventory completeness. The pair reduces to the same 28 over 33. And the second pair does it again. Rs 240 crore against the same Rs 990 crore cap is 24.2 per cent of limit L8. And 8 of 33 never validated is also 24.2 per cent of the running population. Two exact fraction collisions inside one case is a reason to name the object every single time rather than to treat a recurring figure as a pattern.

One more distinction decides how the backlog is ranked. A backtest takes one model's output and compares it with what actually happened, over a defined number of days. Validation asks a wider question: whether the model is fit for the use it is put to, taking in its data, its assumptions, its implementation and its limitations. Validation can find a model unfit that backtests perfectly. The two are not the same activity, and how they differ is covered separately. One fact about V1 matters here. The outcome V1 predicts, how long a deposit stays, is only observable over years, so V1 cannot be backtested at all. A rule that cannot be tested against outcomes is exactly the rule that most needs validating, and it is the one this bank has never checked.

Try it out

Of the three models that have never been validated, which should be looked at first?

Debt Capital Markets Bootcamp — Fin Maverick

Who actually picks up a model inventory, and what do they do with it?

Four people read a model inventory, and each reads it for a different thing. The four readings together are the fastest way to see what an inventory is for.

The chief risk officer, Sunanda Ravikumar in this invented bank, reads it for the two counts and the difference between them. She is not going to re-derive anybody's arithmetic, and she does not need to. She can ask why the register says 28 and the sweep says 33, ask which five were missing and who had been quietly relying on them, and ask what the validator's capacity is against a backlog that now runs to fourteen models rather than nine. The single most useful question anybody can ask of an inventory is when it was last tested by going and looking.

A credit analyst at another institution, looking at this bank from the outside as a counterparty rather than from inside it, reads it for shape rather than count. Twenty eight models with 19 current is one thing. Twenty eight registered against 33 running, with the unchecked one sitting under the headline interest rate risk reading, is a different thing entirely, and it changes which questions are worth asking on a call. Not because the bank is badly run, but because it tells the analyst which of the bank's reported numbers rest on a rule nobody outside the builder has looked at.

The model owner reads it for the use restriction, and this is the one practitioners underrate. The use restriction is the only column that protects them. When somebody in another part of the bank picks up their model's output and puts it into a paper answering a question the model was never built for, the restriction is the document that says so, in advance, in writing, without anybody having to remember a conversation.

And now the household version. The mechanism does not change with scale. The rule a household actually uses to decide what it can afford each month gets written down: what it is, where its inputs come from, when it was last checked against what really happened, and what it is not for. Writing that down is four of the six columns, it takes twenty minutes, and it will surface at least one rule carried over from a year that no longer resembles the household's life. The bank version needs 28 rows and a sweep. The household version needs an evening. Neither of them discovers a new fact. Both of them turn something relied on into something that can be checked.

India

What is named here, and where the binding version lives

Every tier, cycle length, trigger level, count and percentage in this case belongs to Vindhya Commercial Bank Limited and to no rulebook. The twelve month validation cycle, the three tier scheme, the escalation at five exceptions and the review at seven are that bank's own policy, and a second bank could set all four differently and still satisfy its supervisor.

The vocabulary of models and validation, and the approach that produces the idea of testing a measure by counting exceptions against it, come from the Basel Committee on Banking Supervision at the Bank for International Settlements, bis.org. The same committee publishes the principles on risk data aggregation and risk reporting that stand behind the idea of an inventory being complete. A standard is not what binds an Indian bank, so naming only the global standard is the confident and common error. What actually binds a bank in India on model governance, on what it must compute, validate and report, and on outsourcing and information security where a model sits inside a bought package, comes from the Reserve Bank of India at rbi.org.in, where the binding text sits.

A validation cycle length, an exception band, a zone boundary, a tiering rule and an effective date each have exactly one binding version, and that version sits in the supervisor's text rather than in any bank's policy.

What model governance leaves to other subjects

How to run a validation step by step, from scoping to sign off, is set out under validation method. What separates validation from backtesting is set out under backtesting, beyond the single fact used above, that model V1 cannot be backtested at all. Data quality, data ownership and the attributes a data element should carry are covered separately too. Value at risk, expected shortfall, the repricing ladder and the economic value of equity computation are each derived with the market risk material, and their outputs appear above only as things models produce. The committee charter, the meeting calendar, the risk owner role in general and how an escalation route is built belong with the risk governance material. A swap, a forward, a bond and a deposit product are explained with the product material.

Sources

SourceDocumentSite
Reserve Bank of IndiaWhat actually binds a bank in India on model governance, on what must be computed, validated and reported, and on outsourcing and information security where a model sits inside a bought packagerbi.org.in
Bank for International SettlementsThe Basel Committee vocabulary for models and validation, the approach behind testing a measure by counting exceptions, and the principles on risk data aggregation and risk reportingbis.org

Vindhya Commercial Bank Limited, its counterparties and Sunanda Ravikumar are invented.
Educational material. Not advice on any investment, tax, budget or market position.

Covered in this topic

Subtopics

Model GovernanceModel PerformanceDocumentation StandardValidation ReportValidation ThresholdGovernance EscalationHow to create a Model Inventory
Next →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.