How to run Model Validation: Seven Steps and a Written Report
Validation runs in seven steps. Scope and tier the model, test the data, challenge the assumptions, replicate the implementation, analyse outcomes where any exist, write the limitations as instructions, then produce a report carrying rated findings, a use restriction and a signature. The fifth step is often empty, and a model with no observable outcome still gets the other six.
The seven step sequence is the shape the work takes, and it is not a standard handed down from anywhere. An institution's actual obligation to check its models is set by its own supervisor, and the requirement lives in that supervisor's text rather than in any method. The seven steps are one workable structure that a single person can hold in their head, and what actually binds a given institution is settled against the supervisors named below.
Everything below rests on one idea about method, and it is worth stating before the first step arrives. A validation is not a test, it is a sequence of tests on different objects, and each object has a different owner. The data belongs to a source system. The assumptions belong to a committee. The implementation belongs to whoever built it. The use belongs to whoever runs the model and acts on what comes out. A method that treats a model as one thing produces one verdict, and there is no route to a fix inside a verdict. A method that separates the objects produces findings somebody can pick up. Each one lands on a desk that can do something about it.
The household version comes first, so the shape is familiar before the vocabulary arrives. A household works out what it can afford as a monthly loan instalment: take-home pay in, usual spending out, keep half of what is left. Now suppose the answer looks wrong. Four separate things could be wrong with it and they are fixed by four different people. The pay figure could be stale, and only whoever reads the payslip can settle it. Keeping half could be the wrong rule, and only whoever chose the rule can settle that. The subtraction could have been done wrongly on the paper, and neither of those two people would know. Or the answer could be perfectly correct for a two year loan and be getting used for a seven year one, and use is a question all of its own. Asking whether the number is right produces an argument. Asking the four questions separately produces four answers, three of which are probably fine.
What is a validation for, and what may the person running it never have done?
A validation answers one question: is this model fit for the use it is actually being put to? Not is it clever, not is it the best available, not does it agree with last year. Until somebody writes the use down, the question has no answer at all. The missing statement of use is why the first step produces a document rather than a test result, and why a validation that starts by opening the spreadsheet has started in the wrong place.
The second thing that defines a validation is who runs it. The person doing the checking must have built none of what they check, and that is a condition on the person rather than on the method. At the invented Vindhya Commercial Bank Limited, that person is Kanaka Murthy, the independent validator sitting in the risk function, and the bank's rule is the plain one: she validates nothing she built and nothing she specified. The rule is not a comment on anybody's honesty. Somebody who chose an assumption has already settled once whether it was worth choosing, and will not ask again with an open mind.
The method below has seven steps, numbered MV1 to MV7, worked end to end on one model that has never been checked at all. Each step takes something in, does one thing to it, and hands something on. Nothing in the sequence is clever. The value is in the fact that it is a sequence: running the assumption challenge before the data test costs an afternoon of argument about a choice that was being fed rubbish anyway.
How is a validation scoped and tiered, and what does the tier change?
MV1 Scope and tier
- Who does it
- The independent validator, working from the model record rather than from the model.
- Takes in
- The model's row on the register: its name, its stated use, its owner and its tier.
- Does
- Writes down which model is being checked, for which use, and to what depth.
- Hands on
- A written scope, agreed with the model owner and the risk function before any testing starts.
- Decides
- Whether the depth carried on the record is accepted for this validation or challenged in writing.
The validation scopeThe written statement of which model is being checked, for which use, and how deeply, agreed before any work starts. is a short document and it is agreed before anything is opened. The scope names the model, it names every use the model is put to, and it states the depth. Agreeing the scope first stops the argument at the end. The only thing anybody can then dispute is whether the work was done rather than whether it was the right work. A scope written after the testing is a scope written to fit the testing.
The tier is the second field, and the tier decides depth. The tier is a materiality rating the institution sets for itself: how much money moves if this rule is wrong, how many decisions rest on it, whether it feeds a limit, a capital number or a committee paper. Every model gets the same seven steps. The tier decides how far each step goes, how often the whole sequence is repeated, and how senior the person signing at the end has to be. Vindhya Commercial Bank Limited runs three tiers on its own test, and its register at month 12 splits 28 models into 6 in tier 1, 13 in tier 2 and 9 in tier 3, and six plus thirteen plus nine is twenty eight.
Now the part of MV1 that is easy to skip, and it matters on the model this guide works through. The tier on the record is an input to the scope rather than a fact about the world, and MV1 is allowed to challenge it in writing. A depth taken straight off a register is a depth chosen by whoever filled the register in. A validation exists to test exactly that kind of choice. Hold that thought: it comes back the moment the worked instance reaches model V1.
What does the model's tier change about the validation?
What is actually tested about the data a model runs on?
Step MV2 is the one people expect to be dull, and it is where a surprising share of the findings come from. MV2 is not a data quality review of the whole institution, and it is not the governance of the data. Data governance is a separate subject with its own owner and its own list of attributes. MV2 asks four questions about the specific inputs this specific model consumes, and each one gets a verdict written down beside it.
MV2 Data
- Who does it
- The validator, with the source system contact named against each input.
- Takes in
- The specification's input list, and the source system each input is drawn from.
- Does
- Runs the four tests below against every input the model consumes.
- Hands on
- Four verdicts per input: passed, failed, or unprovable on the evidence available.
- Decides
- Whether the rest of the validation proceeds on this data or proceeds qualified.
The fourth test is the one that catches people, so it is worth holding apart from the other three. Tests one to three ask whether the data is any good. Test four asks whether there is enough of it to answer the question the model is being asked, and a feed can be flawless and still fail. Two years of balance history is a perfectly good history. Two years is simply not long enough to see how balances behave through a full turn of the rate cycle, and a rule about behaviour through the cycle that has never seen one is a rule fitted to a fragment.
Two more things MV2 hands on rather than settles. MV2 records the source system against every input, and it records how the model behaves when a value does not arrive. Recording the source is what makes the next person's job possible at all. Vindhya Commercial Bank Limited has an instructive gap there. Across the 147 data elements feeding its monthly risk report, the attribute most often missing is the one saying what happens when a value is absent, and it is missing on 105 of them. A data dictionary that describes the value and never its absence describes only the days when everything works. Whose job it is to close that gap is a data governance question, covered separately.
How is an assumption listed, and who does each one belong to?
MV3 Assumptions
- Who does it
- The validator, with the written specification and the committee record open together.
- Takes in
- The specification, and the minutes of whichever body chose each input that was not measured.
- Does
- Lists every assumption the model makes and names the person or committee that chose it.
- Hands on
- An assumption list, each entry carrying an owner, a date and the evidence behind it.
- Decides
- For each entry, whether it can be tested against evidence or only against alternatives.
An assumptionA choice made because evidence does not settle the question, which is why an assumption always has an owner and a date. is a choice made because the evidence does not settle the question. The definition is short and it carries a consequence: somebody chose it, so somebody can be named, and it was chosen on a day, so it can be dated and revisited. An assumption with no owner is not an assumption, it is a habit. Nobody is answerable for a habit, so nobody reviews it. Listing them is most of the work in MV3, and the list is longer than the model's authors expect, because a choice made once and never argued about stops looking like a choice.
At the invented Vindhya Commercial Bank Limited the assumption followed here is a single number. Model V1, the behavioural deposit life model, gives the Rs 36,000 crore of current and savings balances an average behavioural life of 0.5 years. The current and savings balances are contractually repayable on demand, so no contract fixes the answer and something had to be chosen. Committee G4, the asset liability management committee, chose it. Vindhya Commercial Bank Limited makes that committee responsible for behavioural assumptions. So the entry on the list reads: assumption, 0.5 years; owner, committee G4; evidence, the bank's own review. MV3 writes that down and moves to the challenge.
How is an assumption challenged when there is nothing to compare it with?
An assumption with nothing to compare it against is the hard case, and it is the normal case rather than the exception. How long a deposit stays is only visible over years and the record holds no such measurement, so the 0.5 year figure cannot be checked against what actually happened. There is nothing to compare it with. The temptation at this point is to write that the assumption looks aggressive. An opinion is all that is, and an opinion loses an argument with the committee that chose it.
The way out is to stop asking whether the assumption is right and start asking how much it matters. The second question has a name, the range testRecomputing an answer across a range of an assumption to see how far the answer moves and whether it changes sign.: recompute the answer across a plausible range of the assumption and report where the answer goes. The range test converts an argument about a number nobody can settle into a measurement of consequence. Anybody can settle a measurement. And on model V1 it produces something much sharper than a range.
The arithmetic is the bank's own and it is short. Extending the assumed life by one year moves the economic value result by Rs 36,000 crore times the bank's own 2.0 per cent scenario. Rs 36,000 crore at 2.0 per cent is Rs 720 crore. Start from the locked point: at 0.5 years the economic value change is minus Rs 840 crore. Walk it forward. At 1.0 year it is minus Rs 480 crore, at 1.5 years minus Rs 120 crore, at about 1.67 years it is zero, at 2.0 years it is plus Rs 240 crore, at 2.5 years plus Rs 600 crore and at 3.0 years plus Rs 960 crore. The assumption is not adjusting the size of the answer, it is deciding which way the answer points, and that is a finding rather than an opinion.
A reader who skips the picture still needs two things from it. First, the answer stays comfortably inside the relevant limit at every point drawn. Limit L8 caps the economic value sensitivity at Rs 990 crore. The widest reading drawn is plus Rs 960 crore, or 97.0 per cent of the cap. Nothing on any report would look unusual at either end of the range, and a range test is worth running precisely for that reason: the failure mode is not a breach, it is a sign.
Second, and this is the part that makes it a validation finding rather than a curiosity, the range is not invented for the occasion. Vindhya Commercial Bank Limited's own repricing ladder slots the same Rs 36,000 crore at one to three years in bucket RB5. So the bank holds two views of one balance at the same time. Across bucket RB5 alone, from 1.0 year to 3.0 years, the answer runs from minus Rs 480 crore to plus Rs 960 crore and flips sign at about 1.67 years. The minus Rs 840 crore figure is the reading at the 0.5 year life the model itself assumes, and a 0.5 year life sits outside bucket RB5 entirely. The two must not be blurred into one range. The strongest form of a range test is a range the institution has already committed to somewhere else in its own reporting.
The 0.5 year deposit life assumption is under challenge and there is no historical outcome to compare it with. What does the validator do?
What does replicating an implementation mean, and why is it separate from checking the rule?
MV4 Implementation
- Who does it
- The validator, building independently and without looking at the original workings.
- Takes in
- The written specification, the same inputs, and the code or spreadsheet that runs today.
- Does
- Rebuilds the computation from the specification and compares the two outputs number by number.
- Hands on
- An independent rebuild, a comparison, and a list of every difference found.
- Decides
- Whether each difference is a rule question for step MV3 or a build question for the developer.
ReplicationRebuilding a computation independently from the same inputs, to separate a wrong rule from a rule wrongly coded. means rebuilding the computation from the written specification, using the same inputs, without looking at how the original was built, and then comparing. Replication is slow and it feels redundant right up until the moment it is not. The reason it earns a step of its own is that steps MV3 and MV4 catch two faults that look identical from the outside and are repaired by two completely different people.
Think of a household again. Two people work out the monthly instalment they can afford and get different answers. One possibility is that they disagree about the rule, one keeping half of what is left and the other keeping a third. The other possibility is that they agree completely about the rule and one of them added a column wrongly. From the outside both look the same: two numbers that do not match. Inside, the first needs a conversation and the second needs two minutes with a calculator. A wrong rule and a right rule wrongly built produce the same wrong number. One step therefore asks about the rule and a separate step asks about the build.
MV4 is where the independence condition does real work, and that has one practical consequence. A rebuild that starts by reading the original code will reproduce the original code's mistakes with great fidelity, so the rebuild is made from the specification and not from the original workings. If the specification is too thin to rebuild from, that is itself the finding, and it is a common one. Model V3 at this bank is a spreadsheet scorecard maintained by one person with no written specification, and MV4 cannot be run on it at all. There is nothing to rebuild from and nothing to compare against.
Why is replication a separate step from challenging the assumptions?
What is done when the model cannot be tested against any outcome?
MV5 Outcomes
- Who does it
- The validator, working from whatever record of realised outcomes the institution actually keeps.
- Takes in
- What the model said on each past date, and what was afterwards observed on the same date.
- Does
- Compares the two, counts the misses, and describes their pattern and their size.
- Hands on
- An outcome comparison where outcomes exist, and a written statement of the absence where they do not.
- Decides
- Whether the section carries a result, or carries a reason there can be no result.
Outcome analysisComparing what a model said with what happened, which is only possible where the outcome can actually be observed. is the step everybody pictures on hearing the word validation, and it is the step most likely to be empty. Outcome analysis needs two things: the model must have said something on a past date, and the thing it said must have been observable afterwards. Plenty of models fail the second condition through no fault of anybody's.
Model V1 fails it completely. Model V1 predicts how long a balance stays, and that is only visible over a run of years. The bank's record holds no realised deposit life at all, so there is nothing to place beside the model's output. Note carefully what that does and does not mean. Nothing in an empty section says the model has been checked and found sound on outcomes. The absence means the check cannot be performed, and a check that cannot be performed is a different statement altogether. The difference is exactly what the report has to carry.
There is a second reason to write the absence down rather than skip the section, and it decides how the whole report reads. Outcome analysis is the safety net, and a model whose output cannot be tested against any outcome has none. A model with no safety net most needs the other six steps. Where an outcome exists, a weak assumption eventually shows up as a run of misses and somebody notices. Where no outcome exists, nothing ever shows up, and the only thing standing between a bad assumption and a committee paper is the data test, the assumption challenge and the rebuild. Writing that in section 5 is what tells a reader to weigh sections 2, 3 and 4 more heavily than they otherwise would.
The outcome of model V1 is only observable over years, so step MV5 produces nothing. What goes in that section of the report?
How is a limitation written so that somebody downstream can act on it?
MV6 Limitations
- Who does it
- The validator, writing for the person who runs the model rather than for the person who built it.
- Takes in
- Everything steps MV2 to MV5 could not settle, including every written absence.
- Does
- Turns each unsettled item into a sentence telling a user what to do about it.
- Hands on
- A numbered list of limitations, each phrased as an instruction with a subject and a verb.
- Decides
- Which limitations are severe enough to become a restriction on use in step MV7.
A limitationA statement of what a model cannot do, written so that a user knows what to do about it rather than merely that it exists. is a statement of what the model cannot do. Almost everybody writes them as caveats, and a caveat is a sentence everybody nods at and nobody acts on. The test is simple and slightly brutal: read the sentence and ask what a person would do differently on Monday morning because of it. If the answer is nothing, it is a caveat.
Take the one from model V1. The caveat version reads that the economic value figure is sensitive to the assumed deposit life. Perfectly true, entirely inert. Nobody's behaviour changes. The instruction version reads that the economic value figure must be reported with the assumed deposit life printed beside it, every time it appears. An instruction can be followed or not followed, and so it can also be checked. A caveat can be neither. Same fact underneath, two completely different documents.
The pattern generalises. A limitation written as an instruction names who does what and when. The output must always be reported alongside a stated figure. The model must not be used beyond a stated boundary. The answer must be recomputed when a stated input changes. Each of those is checkable a month later by somebody who was not in the room. Checkability is the whole trick, and it is the difference between a report that describes a model and a report that changes how the model is used.
What is the difference between a limitation written as a caveat and one written as an instruction?
What goes in the validation report, and what changes because of it?
MV7 The report
- Who does it
- The validator writes and signs it. The model owner and the risk function receive it.
- Takes in
- The outputs of steps MV1 to MV6, in the order the steps produced them.
- Does
- Assembles seven parts: scope, rated findings, limitations, a use restriction, actions, a signature and a date.
- Hands on
- One signed document, and one agreed action per finding sitting on a named person's list.
- Decides
- What the model may and may not be used for from the date the report is signed.
The report is the output of the whole sequence and it has seven parts. Five of them describe. Two of them act. A report that stops after the findings has described a model, and a report carrying a use restriction has changed what somebody is allowed to do on Monday. That is the only real test of whether MV7 was done or merely written.
The use restrictionThe written statement of what the model may and may not be used for, which is what a validation report exists to produce. is the part most often left out. Writing one feels like overreach. It is not. MV7 closes the sequence MV1 opened: MV1 asked which use the model serves, and MV7 answers whether the model is fit for exactly that use and for nothing wider. On model V1, a restriction that follows from the work would say that the economic value output may be used for internal reporting only when it is accompanied by the assumed deposit life, and may not be used to support a statement about the direction of the institution's rate exposure until the assumption has been tested.
What makes a finding usable rather than merely true?
A findingA gap between what was found and what should have been, carrying a cause, an effect and an agreed action with an owner and a date. is a gap between what was found and what should have been. Written properly it has four parts, and this invented bank uses the same four in its audit work, numbered AF1 to AF4, with the agreed action carried as AF5 alongside. The condition is what was found. The criteria is what should have been. The cause is why the gap exists. The effect is what it could lead to.
The cause decides whether a finding can be fixed or only patched, and the cause is the part most often missing. A finding with a condition and an action but no cause can always be closed, because somebody does the thing once and ticks the box, and then it comes back, because nothing changed about why it happened in the first place. In this invented bank, 9 of the 42 control testing findings carry no cause at all.
A finding says the model has never been validated, and the agreed action is to validate it. What is missing?
What happens when a finding cannot be closed on time?
Some findings cannot be closed inside the date agreed, and a method that has no answer for that leaves its own output rotting on a list. The answer is an escalation, and an escalation is three things written down in advance: who is told, on what timetable, and what they are being asked to decide. Not who is blamed. The decision is usually whether to accept the position for a further period with a dated plan, or to restrict the use of the model until the work is done.
Where the escalation goes is a question about this institution and not about the method, and it is worth naming because the answer is often surprising. In this invented bank, the risk taxonomy places model risk inside operational risk, so the model inventory and its validation status are reported to committee G6, the operational risk management committee. Meanwhile committee G2, the board risk management committee, sets every limit L1 to L12 and reads the numbers those models produce. The body that receives the validation status is not the body that relies on the models, and that placement is a choice the institution made rather than a rule it followed. How committees are chartered and how a decision is minuted belong to the governance material and are covered separately.
One more thing to keep an eye on. Good methods quietly fail here. An open finding has two different measures on it and they are not the same: how old it is, and whether it is past the date agreed for fixing it. An issue can be old and entirely on track, or raised last week and already late. The bank carries 92 open issues aged in five buckets, and 31 of them are past their agreed remediation date, being 33.7 per cent. A function reporting the first measure and managing on the second is doing it in the right order.
Who sets the expectation that a model gets checked at all
The mechanism above holds in any jurisdiction: seven steps, a checker who built none of it, a report and a use restriction. Where the expectation comes from is a separate question with two answers. The Basel Committee at the Bank for International Settlements, at bis.org, is the origin of the supervisory expectation that a model used to measure risk is independently reviewed, and of the vocabulary used for validation and for testing a measure against realised outcomes. The binding requirements on an Indian bank, covering model governance, independent review, outsourced models sitting inside a bought package and the information security around them, come from the Reserve Bank of India, at rbi.org.in.
The seven steps are about to be run on model V1, and model V1 has never been validated. Which step will produce nothing at all?
What does the whole method look like run end to end on one model?
The sequence now runs on the model the earlier steps have been circling. Vindhya Commercial Bank Limited runs MV1 to MV7 on model V1, the behavioural deposit life model, and model V1 has never been validated at all.
| Step | What it produced on model V1 |
|---|---|
| MV1 | Scope: model V1, which sets an average behavioural life for the Rs 36,000 crore of current and savings balances and feeds the economic value of equity computation that runs against limit L8. Depth: challenged rather than accepted, for the reason set out under the table. |
| MV2 | The model needs a history of account balances and closures by product and by vintage. Which source system does it come from? What does it mean when a balance goes to zero and the account stays open? Is the history long enough to see a full turn of the rate cycle? Four verdicts, one per test, per input. |
| MV3 | One assumption: an average behavioural life of 0.5 years. Chosen by committee G4, the asset liability management committee. Evidence behind it: the bank's own review. No realised outcome exists to test it against. |
| MV4 | Rebuild from the same inputs. Rate sensitive assets Rs 84,000 crore at a modified duration of 3.00 years, rate sensitive liabilities Rs 84,000 crore at 2.50 years, so the duration gap is 0.50 years, and 0.50 times 2.0 per cent times Rs 84,000 crore is Rs 840 crore, negative under a rise. Compare that with what the system prints. |
| MV5 | Nothing, and the section says so in terms. The outcome the model predicts is only observable over years, this bank holds no realised deposit life, and no test against outcomes is possible. |
| MV6 | One limitation, written as an instruction: the economic value figure must be reported with the assumed deposit life printed beside it, every time it appears. |
| MV7 | One report. A finding rated on the bank's own scale, a use restriction, one agreed action with a named owner and a date, and a signature from Kanaka Murthy, who built nothing she checks. |
Two things in that run deserve pulling out, and the first is the one MV1 produced. The register does not place model V1 in tier 1, and the arithmetic makes that certain rather than likely. This bank's register carries 6 models in tier 1, and of those 6, 5 are validated and current and 1 is overdue. Five plus one is six, so the whole of tier 1 is accounted for, and none of it is a model that has never been validated. Model V1 has never been validated. Model V1 is therefore in tier 2 or tier 3, and the record does not say which. The model that decides the sign of this bank's headline interest rate risk number is not carried at its top materiality tier. MV1 exists to notice exactly that and write it down before any depth is accepted from the record.
The second is what MV3 found, and it is worth restating in one line because it is the payoff of the whole run. The range test moved the answer from minus Rs 840 crore at the model's own 0.5 year life, through zero at about 1.67 years, to plus Rs 960 crore at 3.0 years, and across the bank's own repricing bucket RB5 of one to three years alone it runs from minus Rs 480 crore to plus Rs 960 crore. Every one of those readings sits inside limit L8's Rs 990 crore cap. Nothing breaches. The finding is not about size, it is that one untested choice decides whether this institution reports its value exposure as a loss or a gain under the same scenario.
What does holding the cycle actually cost, and what happens to the queue?
The failure is not that the method is hard. It is that nobody wrote down what it costs.
Every step above is teachable in an afternoon. The eighth thing is in none of them, and the eighth thing is arithmetic rather than craft. Arithmetic is unforgiving. Vindhya Commercial Bank Limited sets itself a twelve month validation cycle. Holding a twelve month cycle on 28 registered models costs 28 validations a year just to keep the current ones current. Not to improve anything. To stand still.
Now the queue. The validation backlogThe models that are overdue or have never been validated, which only shrinks when throughput exceeds the cycle cost. on the register at month 12 counts 9 models. The 6 that are overdue plus the 3 that have never been validated make 9. So a function completing exactly 28 validations a year clears nothing at all, ever. At exactly the cycle rate the backlog of 9 models is permanent: it is 9 at month 12, 9 at month 24 and 9 at month 36. The first year anything changes is the first year throughput goes above 28. At 31 a year the 9 clears in 3.0 years, at 33 a year in 1.8 years, and at 37 a year in 1.0 year.
Then the sweep lands. The bank went and looked, and found 33 models in use against the 28 on the register. On the population that actually exists, the twelve month cycle costs 33 a year rather than 28, and the queue is 14 rather than 9, being the same 6 overdue plus the same 3 never validated plus the 5 that were never registered and so were never in any cycle. At 37 validations a year the same function clears the registered queue of 9 in 1.0 year and the real queue of 14 in 3.5 years. Nobody worked harder or less hard. The denominator moved, and a plan built from the register was out by two and a half years before anybody opened a model file.
The honest last act of the method is therefore a count of what the method costs. A validation plan built from a register nobody has swept is a plan for a different institution.
The cycle is twelve months on 28 models and the queue is 9. Before the control moves: how long does a function completing exactly 28 validations a year take to clear it?
Move the throughput and find the line below which nothing improves
One control: c, the number of validations the function completes in a year, from 20 to 50. Vindhya Commercial Bank Limited locks no capacity figure for its validation function, so c is left open. Two consequences, side by side. On the 28 on the register, the twelve month cycle costs 28 a year and the queue is 9, so clearance takes 9 over c less 28 years: at c of 28 it never clears, at 29 it is 9.0 years, at 31 it is 3.0 years, at 33 it is 1.8 years, at 37 it is 1.0 year and at 40 it is 0.75 years. On the 33 the sweep found, the cycle costs 33 a year and the queue is 14, so clearance takes 14 over c less 33 years: at c of 34 it is 14.0 years, at 37 it is 3.5 years, at 40 it is 2.0 years and at 47 it is 1.0 year. At or below 28 on the register, and at or below 33 on what is actually running, the queue never clears at all and the control shows no number rather than a very large one. The crossing that matters is at c of 37, where the same year of effort clears the registered queue in 1.0 year and the real one in 3.5 years.
At 28 validations a year, the queue on the 28 registered models does not clear at all, and the queue on the 33 the sweep found does not clear either, because 28 a year is exactly what holding the twelve month cycle on the register costs.
At 37 validations a year, why does the queue take 1.0 year to clear on one measure and 3.5 years on another?
Who reads a validation report, and what do they actually do with it?
A method is easier to hold on to once the audience for the document is clear, and it is read by four different people for four different reasons. Each one reads a different part, so they are taken in turn.
The person running the model reads the limitations and the use restriction, and nothing else. The model runner wants to know what changes on Monday: which figure now has to be published with something printed beside it, and which question may no longer be answered with this output. Parts 3 and 4 of the report are written as instructions rather than as observations for exactly that reason.
The head of internal audit, who at this invented bank is Rustom Batliwala reporting to the audit committee, reads the scope and the findings. Audit is not repeating the validation. Audit is testing whether the validation was done. Testing that means reading the scope to see what was promised and the findings to see what came back, and then checking whether the agreed actions are being closed on time rather than merely being closed.
The committee that relies on the number reads the use restriction and one line of the range. A committee does not want the method. The committee wants to know whether the figure in front of it may be used for the decision in front of it. A report whose fourth part is missing sends a committee a very careful document that answers a question nobody at the table asked.
And the fourth reader is the next validator, two years later. The next validator reads the whole thing. The fastest way to run a validation is to start from the last one and test what has changed. The quiet argument for writing an absence down rather than leaving a section blank sits here: the next person needs to know whether the outcome test was skipped or was impossible, and only one of those two is worth attempting again.
The household version of the same document is a note stuck to the fridge. Not a note that says the affordability sum is uncertain. Such a note changes nothing. A note that says the instalment figure was worked out on a year with no wedding in it and must be redone before anybody commits to a seven year loan. Same facts, and one of the two versions actually stops a decision.
Where this guide stops. What model risk is, why an inventory is a claim rather than a list, how tiering is set and who governs a model are covered separately and are used here rather than re-derived. The difference between validating a model and testing one model's output against realised outcomes is a subject of its own, and step MV5 assumes it: this runbook takes the outcome test as a thing that either exists or does not, and hands the comparison itself on. Which person is accountable for a data element, who stewards it and which attributes it should carry belong to the data governance material, so step MV2 tests the data this model consumes and hands the governance of it straight over. The economic value of equity computation, the repricing ladder and the interest rate risk framework belong to the market risk material, and this guide uses their outputs only to show what an assumption does to an answer. How a committee is chartered, who must be told and on what timetable belong to the governance material. Statistical test procedures are covered separately, as are the requirements set by any particular supervisor.
Sources
| Source | Document | Site |
|---|---|---|
| Bank for International Settlements | The Basel Committee vocabulary for models, independent review and validation, and the approach behind testing a measure against realised outcomes, cited as the origin of the expectation rather than as a requirement in India | bis.org |
| Reserve Bank of India | What actually binds a bank in India on model governance, on independent validation, on outsourced models sitting inside a bought package, and on the information security around them | rbi.org.in |
Vindhya Commercial Bank Limited, Kanaka Murthy and Rustom Batliwala are invented.
Educational material. Not advice on any investment, tax, budget or market position.
