Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Risk Management Program · CoreTrack
1Risk, Treasury & Financial Control
iRisk Foundations
Risk Appetite, Tolerance, Capacity…The Risk Taxonomy and UniverseRisk Register vs Risk MatrixStress TestingScenario Analysis vs Stress TestingImpact and LikelihoodLikelihoodThe Risk EventRisk Assessment
iiEnterprise Risk Management
Enterprise Risk ManagementThe Four Risk TreatmentsRisk CultureRisk MaturityRisk Monitoring
iiiRisk Governance
Risk GovernanceHow to set a…The Risk PolicyThe Risk OwnerThe Risk Committee and Its CharterThe Risk Limit FrameworkRisk EscalationHow to set a…
ivCredit and Counterparty Risk
Collateral AgreementsCollateral vs NettingProbability of DefaultExposureCounterparty ExposureConcentration Risk vs Wrong Way RiskCounterparty Risk vs Credit RiskHow to assess Counterparty ExposureHow to assess Concentration Risk
vMarket Risk
Market RiskSensitivity MeasuresThe Hedging PolicyInterest Rate Risk in the Banking BookIRRBB vs Market RiskExpected ShortfallEconomic Value of EquityVaR BacktestingOpen PositionValue at RiskValue at Risk and Expected ShortfallEconomic Value SensitivityFX ExposureValue at Risk vs Expected ShortfallEarnings at Risk vs…FX Transaction Risk vs…How to measure Interest…How to measure Foreign…
viLiquidity Risk
Liquidity Stress TestingLiquidity Gap vs Liquidity BufferMaturity MismatchThe Debt Maturity ProfileFunding ConcentrationSurvival HorizonThe Contingency Funding PlanNet Stable Funding RatioLiquidity Risk vs Funding RiskLiquidity Coverage RatioLiquidity Gap and BufferHow to run a Liquidity Gap Analysis
viiOperational Risk
Operational LossThe Loss EventRisk and Control Self AssessmentException ManagementInformation Security as a…Segregation of DutiesIssue ManagementThe Near MissRoot Cause Analysis in RiskThe Fraud TriangleCyber Risk vs Third Party RiskHow to run a…How to assess Third…
viiiRisk Reporting, Data and Model Risk
Model RiskModel Validation vs BacktestingHow to run Model ValidationData Governance in RiskModel Risk vs Data RiskKey Risk IndicatorsManagement InformationRisk ReportingRisk ScoreEarnings at RiskRisk Adjusted ReturnEarly Warning IndicatorsHow to build a KRI Dashboard
ixTreasury
Corporate TreasuryAsset Liability ManagementIntragroup FundingThe Treasury PolicyThe Treasury Management SystemThe Cash ForecastCash Pooling and ConcentrationHow to build a Cash Forecast
xFinancial Controls and Assurance
Control AssuranceThe Control LifecycleThe Assurance MapThe Audit FindingIssue RemediationInternal Financial ControlsControl Design vs Control EffectivenessHow to map Internal Financial ControlsHow to test Control…Control DeficiencyMaterial Weakness
xiOperational Resilience
Operational ResilienceBusiness Continuity and Disaster RecoveryBusiness Continuity vs Operational…Crisis ManagementDisaster RecoveryIncident Management

Root Cause Analysis in Risk: Finding Why, Not Who

A root cause is a condition which, had it been different, would have changed the outcome, and which is not itself explained by something further back inside the institution's control. A name is not a cause. The question is why and never who: naming somebody stops the enquiry, produces a fix that reaches one person, and leaves the same conditions in place for the next person.

Somewhere in every institution there is a file that reads like this. An event happened. Somebody wrote down what happened, in the right order, with the right dates, and signed it. A committee read it, agreed it was unfortunate, and moved on. Eleven months later the same thing happened again, and the second file reads almost exactly like the first. Nothing went wrong in the writing. The first file explained nothing, and nobody noticed. An accurate account of an event looks so much like an explanation of it that the difference has to be hunted for on purpose.

Hunting for that difference is the work. One event at Vindhya Commercial Bank Limited, an invented lender, runs the whole chain down to the bottom, and each step carries a count of how much of the bank a fix at that depth would actually protect. The counting is the part most treatments leave out, and it is the part that turns a preference for deep answers into an argument anybody can check.

What is a root cause, and how is it different from what obviously went wrong?

The shape is easier to see away from banking, in a kitchen. A pressure cooker whistle stops working and the dal burns. Asked what went wrong, the answer is the whistle. Replacing the whistle solves the burnt dal of that particular Tuesday. Asked instead what would have had to be different for the dal not to burn, a second answer arrives that the first one hid: the cooker was left on a flame with nobody in the kitchen and nothing else in the house to signal that fifteen minutes had passed. The whistle is what failed; the absence of any second signal is what made the failure matter, and only one of those two answers protects the next meal cooked.

The two answers have names. The proximate causeWhat obviously went wrong, which is usually the last thing in the chain and almost never the thing worth fixing. is what obviously went wrong. The proximate cause is nearly always the last event in the chain, the easiest of all to establish, and almost never the thing worth fixing. Repairing it repairs one instance. The root causeThe condition that, had it been different, would have changed the outcome, and which is not itself explained by something further back inside the institution's control. is the condition that, had it been different, would have changed the outcome, and which is not itself explained by something further back that the institution could reach. The distance between the two is not a matter of taste. The distance is measurable, and it gets measured below.

There is a third thing that is neither, and it produces more confused enquiries than the other two combined. A contributing factorSomething that made the outcome more likely or worse without being necessary to it. made the outcome more likely or made it worse without being necessary to it. The kitchen was noisy. The cook was tired. Both are true and both belong in the account. Neither is a cause: a quiet kitchen and a rested cook would not have changed what happened. The test is unforgiving and it is the same test both times: take the thing away, and ask whether the outcome still occurs.

Now the bank. Vindhya Commercial Bank Limited recorded incident I10 in month 10 of its twelve numbered months. The collateral valuation feed was stale for 11 working days, 340 loans were wrongly marked on the back of it, no customer lost money, the gross loss was Rs 1.4 crore, nothing was recovered and the net loss was Rs 1.4 crore. Incident I10 also sits behind the one weakness this invented bank rated as material for the year, in process PR3, collateral management and valuation, and it touches the valuation of Rs 8,640 crore of secured advances. The Rs 8,640 crore is why the event is worth an enquiry at all: Rs 1.4 crore is a small loss, and the thing it happened to is not small.

The two columns in the figure below are worth reading before anything said about them. The left column is a descriptionAn accurate account of what happened in what order, which is not an analysis and is frequently signed off as one. of incident I10, and there is nothing wrong with it. Every line is accurate, every line is in the right order, and it would survive any check on the facts. The right column is an explanation of the same event. Nothing in the left column names anything to change and everything in the right column does. The whole difference between the two documents sits there, and it is why both get written and only one gets used.

TWO DOCUMENTS ABOUT INCIDENT I10, BOTH ACCURATE Vindhya Commercial Bank Limited, invented. One says what happened. The other says what would have had to be different. A DESCRIPTION OF INCIDENT I10 Day 1. The collateral valuation feed did not arrive. Nothing in the arrangement noticed. Days 1 to 11. 340 loans were marked on the last value the feed had delivered. Day 11. The stale value was spotted by eye and the feed was refreshed the same day. Month 10. Gross Rs 1.4 crore, nothing recovered, net Rs 1.4 crore. No customer lost money. Every line is true. None of it is a cause. AN EXPLANATION OF INCIDENT I10 The element had no rule for an absent value, so an absence was not an event anything acted on. That rule is missing across most of the set, so the same silence is available on 105 elements. The standard the dictionary follows describes a value and says nothing about its absence. Each line is a condition. Change any one of them and the eleven days do not happen. Every line is true. Each of these is a cause. Both columns are accurate. The left one says what happened; only the right one says what would have had to be different. A fix can only act on the right hand column, and that is the whole difference.
Set the two accounts of incident I10 beside each other and the left one turns out to carry no cause at all: it is a run of true statements about eleven working days, while the right one names three conditions that each had to hold for those days to pass unnoticed, and a fix can only reach the second sort of sentence.
Risk Management Program Bootcamp — Fin Maverick

Why is the question why and never who?

Why and never who is the rule that gives the method its shape. Stated carelessly, the rule sounds like a plea for kindness, and a plea for kindness is not what it is. The rule is an argument about what a cause is, and it would hold even in an institution with no interest whatsoever in being pleasant to anybody. A name does not satisfy the test a cause has to satisfy. Take the person away and the outcome does not change: the seat they sat in still exists, and somebody else will be in it by Monday.

Watch what happens to the enquiry the moment a name is written in the cause field. The enquiry stops. The stop is not a figure of speech. There is now a subject, an action, and a status of closed, and every question that would have been asked next has become impolite rather than merely unasked. Nobody goes on to ask why the arrangement allowed one person's attention to be the only thing standing between a stale number and 340 loans. The file already contains an answer, and answers are what files are for.

Then count what the closure cost. There are two costs, and only one of them ever gets counted. The first is the fix. Retraining or replacing one person leaves every condition in the right hand column of the previous figure exactly where it was: the missing rule for an absent value is still missing, on this element and on the 104 others, and nothing new would detect the same absence tomorrow. The reachHow many other things a proposed fix would also protect, which is the honest way to compare fixes at different depths. of the fix is one person, and the reach of the failure was never one person to begin with.

The second cost is the one nobody counts, and it is larger. Everybody who watched now knows what an enquiry produces. The next person who notices something odd on a Tuesday afternoon does a very quick and entirely rational calculation about what raising it is likely to lead to. After that the institution stops hearing about the small things and goes on believing it has an open culture. The register still exists, and nobody has said anything to the contrary. An institution that wants to hear about near misses cannot also produce enquiries that end in names. The two are not in tension; they are simply incompatible, and the second wins quietly.

Three separate people at this invented bank did their jobs exactly right. The record makes the argument better than any general statement of it could. In near miss N3 a data quality check caught a stale collateral feed. In near miss N4 a checker refused a trade finance document set carrying a forgery pattern. In near miss N5 a quarterly access review found a leaver's privileged account still live. Three correct actions. And each of the three sits inside a design failure that surrounded it: the catch in N3 was never raised as an issue, the refusal in N4 was never linked to anything, and the access review found the account exactly as late as a quarterly cycle was always going to find it. Not one of those three failures is about a person, and not one of them would be reached by any enquiry that went looking for one.

WHAT A NAME IN THE CAUSE FIELD BUYS, AND WHAT IT LEAVES BEHIND The file on the left is constructed to show the shape. It is not this invented bank's record, and no person is named in it. THE FILE AN ENQUIRY LIKE THAT PRODUCES EVENT Incident I10, collateral feed stale 11 working days CAUSE RECORDED The person on the desk did not check the feed ACTION AGREED Further training for the person concerned STATUS CLOSED WHAT IS STILL TRUE THE NEXT MORNING The rule for an absent value is still undefined for 105 of the 147 risk data elements. Nothing that could detect an absent value has been built, on this element or on any other one. The next person in that seat inherits the identical arrangement, unchanged in every respect. COST ONE: THE FIX THAT WAS NOT MADE A closed file and no change to any condition. The reach of retraining one person is exactly one person, and the failure never was. COST TWO: THE NEXT REPORT THAT NEVER ARRIVES Everybody who watched now knows what an enquiry produces, so the small things stop being mentioned. This is the dearer of the two. Neither cost appears in any loss figure, and the second one is invisible until the register quietly stops filling up.
Naming a person buys a closed file within the week and leaves every condition that produced the eleven days exactly where it was, so the enquiry ends with the rule for an absent value still undefined on 105 of the 147 elements and with everybody who watched now holding a clear view of what raising something leads to.
Try it out

Why is the question why and never who?

Derivatives Foundation Bootcamp — Fin Maverick

What actually qualifies as a cause?

Once names stop being written, something has to go in their place, and this is where most enquiries drift. The temptation is to fill the cause field with whatever is true and sounds serious: the volumes were high, the team was short-staffed, the handover was rushed, the system is old. Every one of those may be entirely accurate. None of them is necessarily a cause, and the difference is settled by two tests applied in order rather than by how weighty the sentence sounds.

The first test is a counterfactual and it is the one that does the work: take the candidate away, put nothing in its place, and ask whether the outcome still happens. If it does, the candidate is background. If it does not, the candidate is a cause. Applied to a busy month at this invented bank, the test fails immediately: with a quiet month the collateral valuation feed still goes silent, nothing is still defined to happen when it does, and 340 loans are still marked on a value that stopped moving. A busy month is a true fact about month 10 that made no difference at all to what month 10 produced.

The second test is what separates a cause from the cause, and it is the one that most enquiries never reach. Ask whether the candidate is itself explained by something further back that this institution could reach. A stale feed passes the first test cleanly. A fresh feed changes the outcome. But a stale feed is explained by something further back: nothing was watching for it. So a stale feed is a real link in the chain and it is not the end of it. The chain ends at the first candidate that passes the first test and has nothing behind it that the institution could change, and that condition, and only that one, is the root.

Four sentences about incident I10 are laid out below with both tests applied. The four are worth reading as a set. All four are true, all four sound like explanations, and they land in three different places.

FOUR TRUE SENTENCES ABOUT INCIDENT I10, PUT THROUGH BOTH TESTS TEST 1: take it away and ask whether the outcome still happens. TEST 2: ask whether it is itself explained by something further back inside this invented bank's control. TEST 1 TEST 2 The month was a busy one for the team. Background. A quiet month produces the same eleven days. NO CHANGE NOT REACHED The collateral valuation feed was stale for 11 working days. A link in the chain. Something further back explains why it sat there. OUTCOME CHANGES YES, FURTHER BACK No rule said what should happen when the value was absent. A link in the chain, and the one whose fix reaches furthest. OUTCOME CHANGES YES, FURTHER BACK The standard the dictionary follows describes a value and not its absence. The root. Ask why once more and the answer leaves this bank entirely. OUTCOME CHANGES NOTHING FURTHER Three of the four are true and only one is the root. Being true was never the qualification.
Put four accurate sentences about incident I10 through the same two tests and they separate into three kinds: one is background because removing it changes nothing, two are real links whose own explanation sits further back, and only the last has nothing behind it that this invented bank could change.
Try it out

A written analysis of incident I10 records this in the cause field: the month was a busy one and the team was short-handed. Is that a cause?

What does the chain look like when it is worked all the way down?

The instrument for getting from the first answer to the last one is embarrassingly simple. Its simplicity is why it is often dismissed and why it keeps working. The answer just given is taken, and why is asked of it. Then that happens again. The repeated whyAsking why of each answer in turn, which is a discipline for not stopping at the first plausible sentence rather than a rule that there are five of them. is not really a technique at all; it is a discipline for not stopping at the first plausible sentence, and it is famous mostly because it removes the excuse that a deeper answer was not available.

The method is borrowed, and where it comes from is worth saying. The habit of asking why five times over belongs to Japanese manufacturing quality practice and is usually credited to Sakichi Toyoda, whose interest was in stopping a machine fault from returning rather than in operational risk. The number five is a rule of thumb from that tradition and nothing more. Five is not a standard, no authority prescribes it, and there is no virtue in reaching exactly five or in stopping short of it. On some events the chain is two links long and on others it is seven. The chain worked here happens to run five, and that is a fact about this event and not about the method.

There is a second borrowed instrument worth naming and handing on rather than teaching here. The cause and effect diagram sorts candidate causes into branches before any of them is tested, and it belongs to Kaoru Ishikawa and the same quality tradition. The diagram is a way of generating candidates. The two tests in the previous block are the way of killing the ones that do not survive, and a workshop that generates candidates without ever killing any produces a diagram and no finding.

Incident I10 now runs all the way down. The ladder below reads one rung at a time, and the right hand column carries the argument that the depth of an answer decides what a fix can reach.

INCIDENT I10 WORKED DOWN FIVE RUNGS, WITH WHAT EACH FIX REACHES Vindhya Commercial Bank Limited, invented. The reach column counts the bank's own 147 risk data elements. WHY 1 Why was a loan marked on the wrong collateral value? Because the collateral valuation feed was stale. Fix: refresh the feed. 1 of 147, being 0.7 per cent WHY 2 Why did a stale value sit there for 11 working days? Because nothing detected that it had stopped moving. Fix: add a monitoring control on that feed. 1 of 147, being 0.7 per cent WHY 3 Why did nothing detect it? Because no rule said what to do when the value was absent. Fix: define that rule for this one element. 1 of 147, being 0.7 per cent WHY 4 Why was no rule defined for this element? Because it is missing for 105 of the 147 risk data elements. Fix: define it across the whole set. 105 of 147, being 71.4 per cent WHY 5 Why is it missing for 105 of them? Because the standard behind the dictionary never mentions absence. Fix: change the standard. 147 of 147, and every later one Five rungs, five different fixes, and three of them protect exactly one of the bank's 147 risk data elements.
Worked all the way down, incident I10 gives five separate answers and five separate fixes, and the count beside each one shows that the first three protect a single data element apiece while the fourth suddenly protects 105 of them.

Look at what the rungs are actually made of. Every one of them is a genuine answer to the question above it, and every one of them would be perfectly acceptable in a written file. Rung one is where most enquiries end. Rung two is where a thorough one ends, and it feels like real depth. Adding a monitoring control sounds like a control improvement rather than a repair. Rung two is still a repair, and the number in the right hand column is the only thing that says so out loud.

Rung three is where the enquiry stops being about this incident and starts being about the bank. Ask why nothing detected an absent value and the answer is not that somebody was careless. The answer is that nothing had ever been defined to happen when the value was absent. The record of that data element carried its meaning, the system it came from, the role accountable for it, the role that maintained it, the values it was allowed to take, how often a new one should arrive, and every hop it made on the way to the report. Seven descriptions, all correct, all present. Every one of the seven describes the value, and not one of them describes its absence, so on the days when nothing arrived the record had nothing to say and neither did the bank.

THE RECORD OF ONE ELEMENT: THE COLLATERAL VALUATION FEED Vindhya Commercial Bank Limited, invented. Eight attribute slots, seven of them filled. SLOT ATTRIBUTE WHAT IT SAYS STATUS T1 Definition What the element means, written in words CARRIED T2 Source system The system each new value arrives from CARRIED T3 Accountable owner The role answerable for the element CARRIED T4 Steward The role that maintains it day to day CARRIED T5 Permitted values The range a valid value may fall inside CARRIED T6 Refresh frequency How often a fresh value should arrive CARRIED T7 Rule for an absent value What happens when no value arrives at all ABSENT T8 Lineage Every hop from source system to report CARRIED SEVEN SLOTS DESCRIBE THE VALUE. THE EIGHTH DESCRIBES ITS ABSENCE. A record built from the first seven alone is a record of the days when everything works.
The single attribute this element did not carry is the only one of the eight describing what to do when nothing arrives, which is precisely the state the feed spent eleven working days in, so a record that was seven eighths complete was of no use at all on the days that actually mattered.

Now read rung four again. Rung four is not a deeper description of the same incident but a different sentence about a different object. Rungs one to three are all sentences about one data element. Rung four is a sentence about the set: the rule for an absent value is missing for 105 of this invented bank's 147 risk data elements, being 71.4 per cent of them. The moment the answer changes what it is a sentence about, the fix changes what it is able to protect, and that is the only mechanical event on the whole ladder. The 42 elements that do carry all eight attributes are exactly the 42 carrying the seventh, and 147 less 105 is 42, being 28.6 per cent, so the two counts in this bank's own record agree with each other and the chain can be checked rather than believed.

One more thing about the ladder before the counting starts, and it is a matter of ordinary honesty. Asking why of each answer in turn was invented somewhere, and by people with names. The habit comes out of the Toyota production system, where Sakichi Toyoda put it to work and Taiichi Ohno wrote it down, and it travelled into quality management and from there into risk. The cause-and-effect diagram that gets drawn beside it, the fish bone, belongs to Kaoru Ishikawa and is a different frame, covered separately. Neither of them was designed for a bank, and the count in the right hand column is the adaptation: in a factory a fix is visible on the shop floor, and in an institution the only way to see how far a fix reaches is to count what else it touches.

Debt Capital Markets Bootcamp — Fin Maverick

What does a fix at each depth actually reach?

One question turns a preference into an argument. Everybody agrees that deep answers are better than shallow ones, and nobody can say by how much, so the conversation ends in taste and the shallow fix gets built because it is cheaper. Counting ends that. Take each of the five fixes in turn and ask a single flat question: if this fix is built tomorrow, how many of the bank's 147 risk data elements are protected from the same failure? Not improved, not reviewed. Protected from this exact failure, being a value that stops arriving with nothing defined to happen.

Try it out

A stale data feed produced this incident. Before the count is read: how many of this invented bank's 147 risk data elements does refreshing that one feed protect?

HOW MANY OF THE 147 RISK DATA ELEMENTS THE FIX AT EACH RUNG PROTECTS Vindhya Commercial Bank Limited, invented. Reach only. This chart prices nothing. 147 105 0 1 1 1 105 147 THE STEP: 1 TO 105 in one rung, being 71.4 per cent WHY 1 WHY 2 WHY 3 WHY 4 WHY 5 refresh the feed monitor the feed define the rule here define it across the set change the standard THREE RUNGS AT ONE ELEMENT, THEN A HUNDRED AND FIVE. The first three plateaus sit on the axis because each protects a single element out of 147.
Plotting reach against depth turns a preference for deep answers into a countable claim, because the first three rungs are indistinguishable from each other at one element apiece and the fourth moves the fix by two orders of magnitude in a single step.

Read the flat part first. The flat part surprises people. Rungs one, two and three look completely different from one another in a written file. Refreshing a feed is housekeeping, adding a monitoring control is a control improvement, and defining a missing rule sounds like proper design work. All three protect exactly one element out of 147, being 0.7 per cent of the set, so all three are the same fix wearing three different levels of respectability. That is the honest reason a patchA fix that addresses one instance and leaves the conditions that produced it in place. is hard to spot: it does not announce itself, and at rung three it can be defended in a committee for twenty minutes by somebody entirely sincere.

Try it out

At which rung does the fix stop being a patch, and what happens to the reach at that point?

Play with it

Walk down the rungs and watch what the fix reaches

Move the control from one why level to five. The grid holds this invented bank's 147 risk data elements, and an element fills in when the fix at the chosen depth would protect it from the same failure. Everything below is the case's own locked count.

RungThe answer at that rungThe fixReachShare
1The collateral valuation feed was staleRefresh the feed1 of 1470.7
2Nothing detected that it had stopped movingAdd a monitoring control on it1 of 1470.7
3Nothing was defined to happen when the value was absentDefine that rule for this element1 of 1470.7
4That rule is missing for 105 of the 147 elementsDefine it across the set105 of 14771.4
5The standard the dictionary follows describes the value and not its absenceChange the standard147 of 147100.0

Reproduced as static text so the chain and the step survive without touching the control. The share column is in per cent of the 147. Vindhya Commercial Bank Limited, invented.

Elements reached
105
Share of the 147
71.4%
Still unprotected
42
THE 147 RISK DATA ELEMENTS, FILLED WHERE THE FIX REACHES Each square is one element of this invented bank's risk data set. 71.4 per cent of the set 1 refresh the feed 2 add a monitoring control 3 define the rule here 4 define it across the set 5 change the standard

Taken to 4 levels, the fix on this incident reaches 105 of 147 risk data elements, being 71.4 per cent, and the fix is to define the rule across the set.

The control shows reach and never cost. None of the five fixes is priced. A fix reaching 105 elements is not automatically the right one, and rung five is a legitimate choice rather than a better one. The chain is one worked example and not a claim that every incident has five rungs. Every figure belongs to Vindhya Commercial Bank Limited, invented. Educational illustration.

How far down does the questioning go, and what says to stop?

Nothing stops the questions on their own. Why asked of rung five produces an answer, and why asked of that answer produces another, until the chain arrives somewhere entirely true and entirely useless: the industry does it this way, people make mistakes, systems are complicated. So the method needs a brake, and the brake it is normally given is a stop ruleThe test for when to stop asking: the last answer that is both inside the institution's control and something somebody could be asked to do.. Keep going while each answer is inside the institution's control and each fix is something somebody could be asked to do, and stop at the last rung that satisfies both.

Apply that rule honestly to this ladder and something awkward happens. Rung one passes both tests. So does rung two. So does rung three, and so does rung four. And so does rung five. Changing the standard the dictionary is built to is inside this bank's control, and it is a bigger piece of work rather than an impossible one, and a bigger piece of work is exactly the kind of thing somebody gets asked to do. Every rung on the ladder passes both tests, so the rule as usually stated lands on rung five and not on rung four, and any version of it that says otherwise is contradicting itself in the same breath. This matters because rung four is where the interesting thing happened, and it is very tempting to write a stop rule that produces the answer already preferred.

The way out is to notice that two different jobs have been quietly handed to one sentence. A rule for when to STOP is a ceiling: it marks where the questions run out of anything an institution can act on. A rule for when the analysis has gone deep ENOUGH is a floor: it marks where the answers stop being about one instance. The ceiling and the floor are different tests, and they land on different rungs. The stop rule sets the ceiling at rung five and the reach test sets the floor at rung four, so the honest output of the method is a band of two rungs and not a single correct depth. Inside that band, the choice between defining the missing rule across 105 elements and rewriting the standard behind all 147 is a question of what each costs, and this invented case prices neither, so nobody reading it can settle the choice and the analysis should say so rather than pretend.

THE TWO TESTS ARE NOT ONE TEST: WHERE EACH RUNG LANDS Incident I10 at Vindhya Commercial Bank Limited, invented. Three questions, five rungs. RUNG AND ITS FIX INSIDE CONTROL? SOMEBODY COULD BE ASKED? REACH ABOVE ONE? VERDICT 1 refresh the feed YES YES NO, reaches 1 PATCH 2 monitor the feed YES YES NO, reaches 1 PATCH 3 define the rule here YES YES NO, reaches 1 PATCH 4 define it across the set YES YES YES, reaches 105 IN BAND 5 change the standard YES YES YES, reaches 147 IN BAND THE FLOOR IS SET BY REACH: RUNG 4 Below it every fix protects one element of 147, so the analysis has not gone deep enough yet. THE CEILING IS SET BY THE STOP RULE: RUNG 5 Above it the next answer leaves the bank, and nobody here can be asked to change it. THE METHOD RETURNS A BAND OF TWO RUNGS, NOT ONE CORRECT DEPTH. Choosing inside the band is a cost question, and this invented case prices neither fix.
Running the three questions across all five rungs shows why a stop rule stated as a single depth cannot hold: the first two questions are answered yes everywhere and separate nothing, and it is the third question, which is not a stopping test at all, that does the work everybody credits to the stop rule.

The two ends are worth keeping apart. The failures run in opposite directions and look nothing like each other. Stopping below the floor produces a closed file with a repair in it and the same event still available on 104 other elements. Running past the ceiling produces a sentence about the state of the industry. Nobody at this bank can be asked to do anything about it, and it will sit in the cause field being unarguable for as long as the record survives. The first failure is common, cheap and quiet, and the second is rare, expensive and loud, and only the first one ever gets repeated.

Try it out

Rung five would change the standard and reach all 147 elements. Is stopping at rung four wrong?

Equity Research Bootcamp — Fin Maverick

How can an analysis be known to have gone deep enough?

A deep-sounding analysis and a deep one read the same. Depth cannot be checked by looking at the analysis. Depth is checked against the record. If the cause written down is real, it should explain more than the event it was written about, and this bank's own register offers a ready-made test that costs nothing to run. Four months before the eleven days, in month 6, the same collateral valuation feed went stale for 2 working days. A data quality check caught it, the check did exactly what it was designed to do, and the event was recorded as a near miss and never raised as an issue.

THE SAME FEED, FOUR MONTHS APART, AT TWO DIFFERENT COSTS Vindhya Commercial Bank Limited, invented. Both events sit in the bank's own register. 4 MONTHS NEAR MISS N3, MONTH 6 The collateral valuation feed Stale for 2 working days A data quality check caught it Recorded, never raised as an issue No loans wrongly marked COST: NOTHING INCIDENT I10, MONTH 10 The collateral valuation feed Stale for 11 working days Nothing detected it at all Raised as the one material weakness 340 loans wrongly marked COST: Rs 1.4 CRORE NET ONE CAUSE EXPLAINS BOTH: NOTHING WAS EVER DEFINED TO HAPPEN WHEN THE VALUE WAS ABSENT. An analysis reaching rung four covers month 6 and month 10 in the same sentence. A cause reading "the feed was stale" explains neither of them, because a feed going stale twice is not explained by going stale.
Setting the near miss beside the incident gives a free test of any analysis, because a cause that genuinely explains the eleven days in month 10 has to explain the two days in month 6 as well, and the proximate one explains neither of them.

The earlier event is the test, and it runs on somebody else's work in about a minute. The cause they wrote is held against the earlier event. A cause that explains one event and not the other is not a cause, it is a description of the event it was written about. Run it on the rung four answer and both events fall out of it immediately: the rule for an absent value was never defined, so on the day the feed stopped in month 6 nothing was defined to happen and on the day it stopped in month 10 nothing was defined to happen, and the only difference between two working days and eleven is Rs 1.4 crore and a check that happened to be looking the first time. The first was luck. Nobody had designed the good outcome. Four months later the same silence produced a bad one.

Try it out

Somebody submits a root cause analysis of incident I10. What tests whether it went deep enough?

What happens when somebody did it deliberately?

Everything so far has assumed nobody meant it. The other case is the year's largest net loss at this invented bank, incident I13: nine letters of credit issued against forged shipping documents over fourteen months ending in month 8, gross Rs 22.4 crore, Rs 7.0 crore recovered, net Rs 15.4 crore, which is 35.2 per cent of the year's net operational loss of Rs 43.8 crore from a single one of thirteen incidents. Somebody meant that. The obvious reaction is that root cause analysis does not apply. The cause is a person, and the person is known.

Root cause analysis does apply, and the question does not change at all. Asking who committed a deliberate act produces a name. A name is a matter for a disciplinary process, a regulator and possibly a court, and not one of those produces a fix. Asking why the act was possible produces something an institution can build against, and on this incident the answer is sitting in the register a month before the discovery. In month 7, a document set carrying the same forgery pattern was refused by a checker. The checker was right, the refusal was correct, and it was written up as a routine refusal and never linked to anything. Nine issuances ran over fourteen months, one refusal spotted the pattern, and nothing in the design of the process turned that refusal into a question about the other eight.

So the causal question here is not about a person's honesty. The question is this: what allows the same pattern to be caught once and issued nine times? And the honest answer, in this case, is that a refusal was a transaction outcome rather than an observation, and nothing existed to compare one refusal against the rest of the book. The three conditions usually said to sit behind a deliberate act come from Donald Cressey's 1953 study Other People's Money. The three ask a different question inside a different frame, and they are covered separately.

Try it out

Incident I13 was a deliberate fraud. Does root cause analysis still apply, and what does it ask?

Ratio Analysis That Says Something — free micro-course from Fin Maverick

What does a finding with no cause recorded leave anybody able to do?

There is a quieter failure than stopping too early, and it is possible to measure it. An independent review at this invented bank raised 42 findings in the year. Every one of them has a description of what was wrong. Nine of those 42, being 21.4 per cent of the findings, carry no recorded cause at all, and for each of those nine nobody in the bank knows what would have to be different for the thing not to recur. That is not a small administrative gap. The gap decides what the institution can do next, and it decides that before anybody has argued about priorities or budgets.

The two kinds of finding differ in what each one permits. Where a cause of recordThe cause actually written on a finding, without which the finding can be patched and cannot be remediated. exists, somebody can propose a change to the condition it names, somebody else can argue that the change is too expensive or reaches too little, and a committee can decide between them. All of that is available because there is a stated condition to point at. Where no cause is recorded, none of that machinery can start. The only available action is to repair the instance and hope, and hoping is a strategy that works exactly as well the second time as it did the first.

42 FINDINGS RAISED, AND WHAT EACH ONE LEAVES ANYBODY ABLE TO DO Vindhya Commercial Bank Limited, invented. One square is one finding. 33 CARRY A CAUSE OF RECORD 9 CARRY NONE, BEING 21.4 PER CENT OF THE FINDINGS 33 findings A condition is named, so it can be argued about, costed, decided and changed. 9 findings Nothing is named, so the only available action is to repair the instance. A FINDING WITH NO CAUSE CAN BE PATCHED AND CANNOT BE REMEDIATED.
Counting the findings that carry no cause turns a vague worry about record quality into a number the bank can be asked about, because those nine are the exact findings on which nobody can propose a change to anything except the instance itself.
Try it out

How many of this bank's 42 findings carry no recorded cause, and what does that leave anybody able to do?

Ratio Analysis That Says Something teaches you to choose ratios that answer a question rather than fill a template.

How is an analysis told apart from a description?

Telling the two apart is the practical skill. Far more written accounts are received than are ever written, and the two kinds look identical from a distance. Both run to two printed sides. Both are accurate. Both have a heading that says root cause analysis. Three tests separate them and none of the three needs any knowledge of the incident.

The first test takes each sentence sitting in the cause field and asks whether changing it would have changed the outcome. A busy month fails: quiet months produce stale feeds too. A missing rule for an absent value passes: define it and the eleven days do not happen. The second test looks for a person in the cause field. A name there is not a cause, and its presence marks an enquiry that stopped at the first thing that could be held responsible. The third test asks how many other things the proposed fix reaches. If the honest answer to the third test is one, the document is a description with a repair attached to it, whatever the heading says.

Try it out

An account of an incident running to two printed sides arrives for review. What three tests settle whether it is an analysis?

Who actually uses this, and for what

The same shape appears at home, where the stakes are smaller. The geyser trips the fuse on a winter morning. Resetting the fuse is rung one, and it trips again on Thursday. Calling somebody to replace the geyser element is rung two, and the fuse holds until the iron and the pump run together in March. The rung four answer is that the flat was wired for a load that nobody has recalculated since two more appliances arrived, and the fix reaches every socket rather than one appliance. Every household has already met this ladder, and everybody has stopped at rung one at least once. Rung one works often enough to feel like it worked.

Now the finance version, and there are three readers who use it for three different purposes. A credit analyst looking at any lender does not have the loss register, so they read what is published and look for repetition: the same kind of event turning up in successive periods is the visible shadow of analyses that stopped at rung one. A cause that had been reached would have removed the class of event rather than the instance. A supervisor or an assurance reader asks a blunter question: how many findings carry a stated cause at all? The answer sets a ceiling on how much of the remediation programme can be more than repair. And the internal reader with the register in front of them uses the reach count as a way of sorting a queue: given twenty open issues and money for six fixes, the six that reach a set beat the six that reach an instance, and the count is what lets that argument be made without anybody appealing to seniority.

There is a household version of the second failure too, and it is worth naming because institutions do it constantly. An answer to the tripping fuse that Indian domestic wiring standards are what they are is perfectly true and changes nothing in the flat. An analysis has to end somewhere a reachable person could be asked to do something, and everything past that point is commentary dressed as depth.

How this goes wrong, in two directions

The first failure is stopping at the person, and it is the common one because it is so satisfying. Write down that somebody did not check the feed and the enquiry is finished: there is a name, there is an action, there is a closed file, and the meeting ends early. The closed file costs two things. The first cost is the fix. Retraining or replacing one person leaves the rule for an absent value undefined on 105 of 147 elements, and the same eleven days and the same Rs 1.4 crore available on any of them. The second cost is the next report. Everybody who watched has now learned what an enquiry produces, and an institution that wants to hear about near misses has just made the price of raising one visible to its whole staff. A report that is not made leaves no record of not being made, so the larger of the two costs never appears in any file.

Look at what this bank's own register says about that. The register carries the whole argument in three lines. The data quality check that caught the feed in month 6 fired exactly as designed. The checker who refused the forged document set in month 7 did the job exactly as designed. The quarterly access review that found a live privileged account belonging to a leaver, 46 days after the person left, ran exactly on time. Three correct actions. Three failures in the design around them: nothing turned the first into an issue, nothing compared the second against the rest of the book, and nothing shortened the 46 days between a leaver leaving and a review looking. The layered defences picture that James Reason set out in Human Error in 1990 is the frame most people reach for here, and this record is the honest version of it: the layers worked and what sat between them did not.

The second failure is stopping too late, and it is rarer and still a failure. An analysis that lands on a culture, an industry practice or the general difficulty of complex systems has produced a sentence rather than a fix. Nobody at this bank can be asked to change it, no committee can decide anything about it, and it will sit in the cause field indefinitely, unarguable and inert. Depth is not a virtue on its own, and an analysis that runs past the last rung somebody could act on has bought unfalsifiability and called it insight.

Where the rules come from

Named, not stated

No authority anywhere prescribes a number of why levels, and the repeated why is a discipline rather than a rule. The framework around event investigation does come from somewhere. The Basel Committee on Banking Supervision at the Bank for International Settlements, bis.org, publishes the operational risk framework and the seven event categories inside which an incident like this one is classified. The Reserve Bank of India, at rbi.org.in, sets what an Indian bank must actually do about investigating an operational risk event, what it must record and what it must report onward. The five rung chain, the reach counts, the 42 findings and the 9 without a cause all belong to one invented bank.

Where this guide stops. The loss event, the near miss and the link between two events are covered separately earlier in this sequence, and the analysis here starts where that link ends. The issue that comes out of an analysis, who holds it, when it is due and how it ages is covered separately in this sequence. The three conditions behind a deliberate act are covered separately later in this sequence and are named here only. How an audit finding is written, the five parts it carries, the scale a deficiency is rated on, what makes a weakness material and how remediation is designed all belong to the controls and assurance sequence, which comes later and checks what this one designs; the count of findings with no cause is used here only as a measurement. The eight data attributes, the data dictionary, the accountable owner and the steward belong to the risk reporting and data sequence; T7 and the counts around it are used here as facts of this invented case. Instruments are covered elsewhere.

Sources

SourceDocumentSite
Bank for International SettlementsThe Basel Committee on Banking Supervision publications setting out the operational risk framework and the seven event categories an incident is classified intobis.org
Reserve Bank of IndiaWhat an Indian bank must actually do about investigating, recording and reporting an operational risk eventrbi.org.in
Taiichi OhnoToyota Production System, the account of the shop floor practice the repeated why is borrowed fromProductivity Press
James ReasonHuman Error, 1990, the source of the layered defences picture the three near misses are read againstCambridge University Press

Vindhya Commercial Bank Limited is invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.