Root Cause Analysis in Risk: Finding Why, Not Who
A root cause is a condition which, had it been different, would have changed the outcome, and which is not itself explained by something further back inside the institution's control. A name is not a cause. The question is why and never who: naming somebody stops the enquiry, produces a fix that reaches one person, and leaves the same conditions in place for the next person.
Somewhere in every institution there is a file that reads like this. An event happened. Somebody wrote down what happened, in the right order, with the right dates, and signed it. A committee read it, agreed it was unfortunate, and moved on. Eleven months later the same thing happened again, and the second file reads almost exactly like the first. Nothing went wrong in the writing. The first file explained nothing, and nobody noticed. An accurate account of an event looks so much like an explanation of it that the difference has to be hunted for on purpose.
Hunting for that difference is the work. One event at Vindhya Commercial Bank Limited, an invented lender, runs the whole chain down to the bottom, and each step carries a count of how much of the bank a fix at that depth would actually protect. The counting is the part most treatments leave out, and it is the part that turns a preference for deep answers into an argument anybody can check.
What is a root cause, and how is it different from what obviously went wrong?
The shape is easier to see away from banking, in a kitchen. A pressure cooker whistle stops working and the dal burns. Asked what went wrong, the answer is the whistle. Replacing the whistle solves the burnt dal of that particular Tuesday. Asked instead what would have had to be different for the dal not to burn, a second answer arrives that the first one hid: the cooker was left on a flame with nobody in the kitchen and nothing else in the house to signal that fifteen minutes had passed. The whistle is what failed; the absence of any second signal is what made the failure matter, and only one of those two answers protects the next meal cooked.
The two answers have names. The proximate causeWhat obviously went wrong, which is usually the last thing in the chain and almost never the thing worth fixing. is what obviously went wrong. The proximate cause is nearly always the last event in the chain, the easiest of all to establish, and almost never the thing worth fixing. Repairing it repairs one instance. The root causeThe condition that, had it been different, would have changed the outcome, and which is not itself explained by something further back inside the institution's control. is the condition that, had it been different, would have changed the outcome, and which is not itself explained by something further back that the institution could reach. The distance between the two is not a matter of taste. The distance is measurable, and it gets measured below.
There is a third thing that is neither, and it produces more confused enquiries than the other two combined. A contributing factorSomething that made the outcome more likely or worse without being necessary to it. made the outcome more likely or made it worse without being necessary to it. The kitchen was noisy. The cook was tired. Both are true and both belong in the account. Neither is a cause: a quiet kitchen and a rested cook would not have changed what happened. The test is unforgiving and it is the same test both times: take the thing away, and ask whether the outcome still occurs.
Now the bank. Vindhya Commercial Bank Limited recorded incident I10 in month 10 of its twelve numbered months. The collateral valuation feed was stale for 11 working days, 340 loans were wrongly marked on the back of it, no customer lost money, the gross loss was Rs 1.4 crore, nothing was recovered and the net loss was Rs 1.4 crore. Incident I10 also sits behind the one weakness this invented bank rated as material for the year, in process PR3, collateral management and valuation, and it touches the valuation of Rs 8,640 crore of secured advances. The Rs 8,640 crore is why the event is worth an enquiry at all: Rs 1.4 crore is a small loss, and the thing it happened to is not small.
The two columns in the figure below are worth reading before anything said about them. The left column is a descriptionAn accurate account of what happened in what order, which is not an analysis and is frequently signed off as one. of incident I10, and there is nothing wrong with it. Every line is accurate, every line is in the right order, and it would survive any check on the facts. The right column is an explanation of the same event. Nothing in the left column names anything to change and everything in the right column does. The whole difference between the two documents sits there, and it is why both get written and only one gets used.
Why is the question why and never who?
Why and never who is the rule that gives the method its shape. Stated carelessly, the rule sounds like a plea for kindness, and a plea for kindness is not what it is. The rule is an argument about what a cause is, and it would hold even in an institution with no interest whatsoever in being pleasant to anybody. A name does not satisfy the test a cause has to satisfy. Take the person away and the outcome does not change: the seat they sat in still exists, and somebody else will be in it by Monday.
Watch what happens to the enquiry the moment a name is written in the cause field. The enquiry stops. The stop is not a figure of speech. There is now a subject, an action, and a status of closed, and every question that would have been asked next has become impolite rather than merely unasked. Nobody goes on to ask why the arrangement allowed one person's attention to be the only thing standing between a stale number and 340 loans. The file already contains an answer, and answers are what files are for.
Then count what the closure cost. There are two costs, and only one of them ever gets counted. The first is the fix. Retraining or replacing one person leaves every condition in the right hand column of the previous figure exactly where it was: the missing rule for an absent value is still missing, on this element and on the 104 others, and nothing new would detect the same absence tomorrow. The reachHow many other things a proposed fix would also protect, which is the honest way to compare fixes at different depths. of the fix is one person, and the reach of the failure was never one person to begin with.
The second cost is the one nobody counts, and it is larger. Everybody who watched now knows what an enquiry produces. The next person who notices something odd on a Tuesday afternoon does a very quick and entirely rational calculation about what raising it is likely to lead to. After that the institution stops hearing about the small things and goes on believing it has an open culture. The register still exists, and nobody has said anything to the contrary. An institution that wants to hear about near misses cannot also produce enquiries that end in names. The two are not in tension; they are simply incompatible, and the second wins quietly.
Three separate people at this invented bank did their jobs exactly right. The record makes the argument better than any general statement of it could. In near miss N3 a data quality check caught a stale collateral feed. In near miss N4 a checker refused a trade finance document set carrying a forgery pattern. In near miss N5 a quarterly access review found a leaver's privileged account still live. Three correct actions. And each of the three sits inside a design failure that surrounded it: the catch in N3 was never raised as an issue, the refusal in N4 was never linked to anything, and the access review found the account exactly as late as a quarterly cycle was always going to find it. Not one of those three failures is about a person, and not one of them would be reached by any enquiry that went looking for one.
Why is the question why and never who?
What actually qualifies as a cause?
Once names stop being written, something has to go in their place, and this is where most enquiries drift. The temptation is to fill the cause field with whatever is true and sounds serious: the volumes were high, the team was short-staffed, the handover was rushed, the system is old. Every one of those may be entirely accurate. None of them is necessarily a cause, and the difference is settled by two tests applied in order rather than by how weighty the sentence sounds.
The first test is a counterfactual and it is the one that does the work: take the candidate away, put nothing in its place, and ask whether the outcome still happens. If it does, the candidate is background. If it does not, the candidate is a cause. Applied to a busy month at this invented bank, the test fails immediately: with a quiet month the collateral valuation feed still goes silent, nothing is still defined to happen when it does, and 340 loans are still marked on a value that stopped moving. A busy month is a true fact about month 10 that made no difference at all to what month 10 produced.
The second test is what separates a cause from the cause, and it is the one that most enquiries never reach. Ask whether the candidate is itself explained by something further back that this institution could reach. A stale feed passes the first test cleanly. A fresh feed changes the outcome. But a stale feed is explained by something further back: nothing was watching for it. So a stale feed is a real link in the chain and it is not the end of it. The chain ends at the first candidate that passes the first test and has nothing behind it that the institution could change, and that condition, and only that one, is the root.
Four sentences about incident I10 are laid out below with both tests applied. The four are worth reading as a set. All four are true, all four sound like explanations, and they land in three different places.
A written analysis of incident I10 records this in the cause field: the month was a busy one and the team was short-handed. Is that a cause?
What does the chain look like when it is worked all the way down?
The instrument for getting from the first answer to the last one is embarrassingly simple. Its simplicity is why it is often dismissed and why it keeps working. The answer just given is taken, and why is asked of it. Then that happens again. The repeated whyAsking why of each answer in turn, which is a discipline for not stopping at the first plausible sentence rather than a rule that there are five of them. is not really a technique at all; it is a discipline for not stopping at the first plausible sentence, and it is famous mostly because it removes the excuse that a deeper answer was not available.
The method is borrowed, and where it comes from is worth saying. The habit of asking why five times over belongs to Japanese manufacturing quality practice and is usually credited to Sakichi Toyoda, whose interest was in stopping a machine fault from returning rather than in operational risk. The number five is a rule of thumb from that tradition and nothing more. Five is not a standard, no authority prescribes it, and there is no virtue in reaching exactly five or in stopping short of it. On some events the chain is two links long and on others it is seven. The chain worked here happens to run five, and that is a fact about this event and not about the method.
There is a second borrowed instrument worth naming and handing on rather than teaching here. The cause and effect diagram sorts candidate causes into branches before any of them is tested, and it belongs to Kaoru Ishikawa and the same quality tradition. The diagram is a way of generating candidates. The two tests in the previous block are the way of killing the ones that do not survive, and a workshop that generates candidates without ever killing any produces a diagram and no finding.
Incident I10 now runs all the way down. The ladder below reads one rung at a time, and the right hand column carries the argument that the depth of an answer decides what a fix can reach.
Look at what the rungs are actually made of. Every one of them is a genuine answer to the question above it, and every one of them would be perfectly acceptable in a written file. Rung one is where most enquiries end. Rung two is where a thorough one ends, and it feels like real depth. Adding a monitoring control sounds like a control improvement rather than a repair. Rung two is still a repair, and the number in the right hand column is the only thing that says so out loud.
Rung three is where the enquiry stops being about this incident and starts being about the bank. Ask why nothing detected an absent value and the answer is not that somebody was careless. The answer is that nothing had ever been defined to happen when the value was absent. The record of that data element carried its meaning, the system it came from, the role accountable for it, the role that maintained it, the values it was allowed to take, how often a new one should arrive, and every hop it made on the way to the report. Seven descriptions, all correct, all present. Every one of the seven describes the value, and not one of them describes its absence, so on the days when nothing arrived the record had nothing to say and neither did the bank.
Now read rung four again. Rung four is not a deeper description of the same incident but a different sentence about a different object. Rungs one to three are all sentences about one data element. Rung four is a sentence about the set: the rule for an absent value is missing for 105 of this invented bank's 147 risk data elements, being 71.4 per cent of them. The moment the answer changes what it is a sentence about, the fix changes what it is able to protect, and that is the only mechanical event on the whole ladder. The 42 elements that do carry all eight attributes are exactly the 42 carrying the seventh, and 147 less 105 is 42, being 28.6 per cent, so the two counts in this bank's own record agree with each other and the chain can be checked rather than believed.
One more thing about the ladder before the counting starts, and it is a matter of ordinary honesty. Asking why of each answer in turn was invented somewhere, and by people with names. The habit comes out of the Toyota production system, where Sakichi Toyoda put it to work and Taiichi Ohno wrote it down, and it travelled into quality management and from there into risk. The cause-and-effect diagram that gets drawn beside it, the fish bone, belongs to Kaoru Ishikawa and is a different frame, covered separately. Neither of them was designed for a bank, and the count in the right hand column is the adaptation: in a factory a fix is visible on the shop floor, and in an institution the only way to see how far a fix reaches is to count what else it touches.
What does a fix at each depth actually reach?
One question turns a preference into an argument. Everybody agrees that deep answers are better than shallow ones, and nobody can say by how much, so the conversation ends in taste and the shallow fix gets built because it is cheaper. Counting ends that. Take each of the five fixes in turn and ask a single flat question: if this fix is built tomorrow, how many of the bank's 147 risk data elements are protected from the same failure? Not improved, not reviewed. Protected from this exact failure, being a value that stops arriving with nothing defined to happen.
A stale data feed produced this incident. Before the count is read: how many of this invented bank's 147 risk data elements does refreshing that one feed protect?
Read the flat part first. The flat part surprises people. Rungs one, two and three look completely different from one another in a written file. Refreshing a feed is housekeeping, adding a monitoring control is a control improvement, and defining a missing rule sounds like proper design work. All three protect exactly one element out of 147, being 0.7 per cent of the set, so all three are the same fix wearing three different levels of respectability. That is the honest reason a patchA fix that addresses one instance and leaves the conditions that produced it in place. is hard to spot: it does not announce itself, and at rung three it can be defended in a committee for twenty minutes by somebody entirely sincere.
At which rung does the fix stop being a patch, and what happens to the reach at that point?
Walk down the rungs and watch what the fix reaches
Move the control from one why level to five. The grid holds this invented bank's 147 risk data elements, and an element fills in when the fix at the chosen depth would protect it from the same failure. Everything below is the case's own locked count.
| Rung | The answer at that rung | The fix | Reach | Share |
|---|---|---|---|---|
| 1 | The collateral valuation feed was stale | Refresh the feed | 1 of 147 | 0.7 |
| 2 | Nothing detected that it had stopped moving | Add a monitoring control on it | 1 of 147 | 0.7 |
| 3 | Nothing was defined to happen when the value was absent | Define that rule for this element | 1 of 147 | 0.7 |
| 4 | That rule is missing for 105 of the 147 elements | Define it across the set | 105 of 147 | 71.4 |
| 5 | The standard the dictionary follows describes the value and not its absence | Change the standard | 147 of 147 | 100.0 |
Reproduced as static text so the chain and the step survive without touching the control. The share column is in per cent of the 147. Vindhya Commercial Bank Limited, invented.
Taken to 4 levels, the fix on this incident reaches 105 of 147 risk data elements, being 71.4 per cent, and the fix is to define the rule across the set.
The control shows reach and never cost. None of the five fixes is priced. A fix reaching 105 elements is not automatically the right one, and rung five is a legitimate choice rather than a better one. The chain is one worked example and not a claim that every incident has five rungs. Every figure belongs to Vindhya Commercial Bank Limited, invented. Educational illustration.
How far down does the questioning go, and what says to stop?
Nothing stops the questions on their own. Why asked of rung five produces an answer, and why asked of that answer produces another, until the chain arrives somewhere entirely true and entirely useless: the industry does it this way, people make mistakes, systems are complicated. So the method needs a brake, and the brake it is normally given is a stop ruleThe test for when to stop asking: the last answer that is both inside the institution's control and something somebody could be asked to do.. Keep going while each answer is inside the institution's control and each fix is something somebody could be asked to do, and stop at the last rung that satisfies both.
Apply that rule honestly to this ladder and something awkward happens. Rung one passes both tests. So does rung two. So does rung three, and so does rung four. And so does rung five. Changing the standard the dictionary is built to is inside this bank's control, and it is a bigger piece of work rather than an impossible one, and a bigger piece of work is exactly the kind of thing somebody gets asked to do. Every rung on the ladder passes both tests, so the rule as usually stated lands on rung five and not on rung four, and any version of it that says otherwise is contradicting itself in the same breath. This matters because rung four is where the interesting thing happened, and it is very tempting to write a stop rule that produces the answer already preferred.
The way out is to notice that two different jobs have been quietly handed to one sentence. A rule for when to STOP is a ceiling: it marks where the questions run out of anything an institution can act on. A rule for when the analysis has gone deep ENOUGH is a floor: it marks where the answers stop being about one instance. The ceiling and the floor are different tests, and they land on different rungs. The stop rule sets the ceiling at rung five and the reach test sets the floor at rung four, so the honest output of the method is a band of two rungs and not a single correct depth. Inside that band, the choice between defining the missing rule across 105 elements and rewriting the standard behind all 147 is a question of what each costs, and this invented case prices neither, so nobody reading it can settle the choice and the analysis should say so rather than pretend.
The two ends are worth keeping apart. The failures run in opposite directions and look nothing like each other. Stopping below the floor produces a closed file with a repair in it and the same event still available on 104 other elements. Running past the ceiling produces a sentence about the state of the industry. Nobody at this bank can be asked to do anything about it, and it will sit in the cause field being unarguable for as long as the record survives. The first failure is common, cheap and quiet, and the second is rare, expensive and loud, and only the first one ever gets repeated.
Rung five would change the standard and reach all 147 elements. Is stopping at rung four wrong?
How can an analysis be known to have gone deep enough?
A deep-sounding analysis and a deep one read the same. Depth cannot be checked by looking at the analysis. Depth is checked against the record. If the cause written down is real, it should explain more than the event it was written about, and this bank's own register offers a ready-made test that costs nothing to run. Four months before the eleven days, in month 6, the same collateral valuation feed went stale for 2 working days. A data quality check caught it, the check did exactly what it was designed to do, and the event was recorded as a near miss and never raised as an issue.
The earlier event is the test, and it runs on somebody else's work in about a minute. The cause they wrote is held against the earlier event. A cause that explains one event and not the other is not a cause, it is a description of the event it was written about. Run it on the rung four answer and both events fall out of it immediately: the rule for an absent value was never defined, so on the day the feed stopped in month 6 nothing was defined to happen and on the day it stopped in month 10 nothing was defined to happen, and the only difference between two working days and eleven is Rs 1.4 crore and a check that happened to be looking the first time. The first was luck. Nobody had designed the good outcome. Four months later the same silence produced a bad one.
Somebody submits a root cause analysis of incident I10. What tests whether it went deep enough?
What happens when somebody did it deliberately?
Everything so far has assumed nobody meant it. The other case is the year's largest net loss at this invented bank, incident I13: nine letters of credit issued against forged shipping documents over fourteen months ending in month 8, gross Rs 22.4 crore, Rs 7.0 crore recovered, net Rs 15.4 crore, which is 35.2 per cent of the year's net operational loss of Rs 43.8 crore from a single one of thirteen incidents. Somebody meant that. The obvious reaction is that root cause analysis does not apply. The cause is a person, and the person is known.
Root cause analysis does apply, and the question does not change at all. Asking who committed a deliberate act produces a name. A name is a matter for a disciplinary process, a regulator and possibly a court, and not one of those produces a fix. Asking why the act was possible produces something an institution can build against, and on this incident the answer is sitting in the register a month before the discovery. In month 7, a document set carrying the same forgery pattern was refused by a checker. The checker was right, the refusal was correct, and it was written up as a routine refusal and never linked to anything. Nine issuances ran over fourteen months, one refusal spotted the pattern, and nothing in the design of the process turned that refusal into a question about the other eight.
So the causal question here is not about a person's honesty. The question is this: what allows the same pattern to be caught once and issued nine times? And the honest answer, in this case, is that a refusal was a transaction outcome rather than an observation, and nothing existed to compare one refusal against the rest of the book. The three conditions usually said to sit behind a deliberate act come from Donald Cressey's 1953 study Other People's Money. The three ask a different question inside a different frame, and they are covered separately.
Incident I13 was a deliberate fraud. Does root cause analysis still apply, and what does it ask?
What does a finding with no cause recorded leave anybody able to do?
There is a quieter failure than stopping too early, and it is possible to measure it. An independent review at this invented bank raised 42 findings in the year. Every one of them has a description of what was wrong. Nine of those 42, being 21.4 per cent of the findings, carry no recorded cause at all, and for each of those nine nobody in the bank knows what would have to be different for the thing not to recur. That is not a small administrative gap. The gap decides what the institution can do next, and it decides that before anybody has argued about priorities or budgets.
The two kinds of finding differ in what each one permits. Where a cause of recordThe cause actually written on a finding, without which the finding can be patched and cannot be remediated. exists, somebody can propose a change to the condition it names, somebody else can argue that the change is too expensive or reaches too little, and a committee can decide between them. All of that is available because there is a stated condition to point at. Where no cause is recorded, none of that machinery can start. The only available action is to repair the instance and hope, and hoping is a strategy that works exactly as well the second time as it did the first.
How many of this bank's 42 findings carry no recorded cause, and what does that leave anybody able to do?
How is an analysis told apart from a description?
Telling the two apart is the practical skill. Far more written accounts are received than are ever written, and the two kinds look identical from a distance. Both run to two printed sides. Both are accurate. Both have a heading that says root cause analysis. Three tests separate them and none of the three needs any knowledge of the incident.
The first test takes each sentence sitting in the cause field and asks whether changing it would have changed the outcome. A busy month fails: quiet months produce stale feeds too. A missing rule for an absent value passes: define it and the eleven days do not happen. The second test looks for a person in the cause field. A name there is not a cause, and its presence marks an enquiry that stopped at the first thing that could be held responsible. The third test asks how many other things the proposed fix reaches. If the honest answer to the third test is one, the document is a description with a repair attached to it, whatever the heading says.
An account of an incident running to two printed sides arrives for review. What three tests settle whether it is an analysis?
Who actually uses this, and for what
The same shape appears at home, where the stakes are smaller. The geyser trips the fuse on a winter morning. Resetting the fuse is rung one, and it trips again on Thursday. Calling somebody to replace the geyser element is rung two, and the fuse holds until the iron and the pump run together in March. The rung four answer is that the flat was wired for a load that nobody has recalculated since two more appliances arrived, and the fix reaches every socket rather than one appliance. Every household has already met this ladder, and everybody has stopped at rung one at least once. Rung one works often enough to feel like it worked.
Now the finance version, and there are three readers who use it for three different purposes. A credit analyst looking at any lender does not have the loss register, so they read what is published and look for repetition: the same kind of event turning up in successive periods is the visible shadow of analyses that stopped at rung one. A cause that had been reached would have removed the class of event rather than the instance. A supervisor or an assurance reader asks a blunter question: how many findings carry a stated cause at all? The answer sets a ceiling on how much of the remediation programme can be more than repair. And the internal reader with the register in front of them uses the reach count as a way of sorting a queue: given twenty open issues and money for six fixes, the six that reach a set beat the six that reach an instance, and the count is what lets that argument be made without anybody appealing to seniority.
There is a household version of the second failure too, and it is worth naming because institutions do it constantly. An answer to the tripping fuse that Indian domestic wiring standards are what they are is perfectly true and changes nothing in the flat. An analysis has to end somewhere a reachable person could be asked to do something, and everything past that point is commentary dressed as depth.
How this goes wrong, in two directions
The first failure is stopping at the person, and it is the common one because it is so satisfying. Write down that somebody did not check the feed and the enquiry is finished: there is a name, there is an action, there is a closed file, and the meeting ends early. The closed file costs two things. The first cost is the fix. Retraining or replacing one person leaves the rule for an absent value undefined on 105 of 147 elements, and the same eleven days and the same Rs 1.4 crore available on any of them. The second cost is the next report. Everybody who watched has now learned what an enquiry produces, and an institution that wants to hear about near misses has just made the price of raising one visible to its whole staff. A report that is not made leaves no record of not being made, so the larger of the two costs never appears in any file.
Look at what this bank's own register says about that. The register carries the whole argument in three lines. The data quality check that caught the feed in month 6 fired exactly as designed. The checker who refused the forged document set in month 7 did the job exactly as designed. The quarterly access review that found a live privileged account belonging to a leaver, 46 days after the person left, ran exactly on time. Three correct actions. Three failures in the design around them: nothing turned the first into an issue, nothing compared the second against the rest of the book, and nothing shortened the 46 days between a leaver leaving and a review looking. The layered defences picture that James Reason set out in Human Error in 1990 is the frame most people reach for here, and this record is the honest version of it: the layers worked and what sat between them did not.
The second failure is stopping too late, and it is rarer and still a failure. An analysis that lands on a culture, an industry practice or the general difficulty of complex systems has produced a sentence rather than a fix. Nobody at this bank can be asked to change it, no committee can decide anything about it, and it will sit in the cause field indefinitely, unarguable and inert. Depth is not a virtue on its own, and an analysis that runs past the last rung somebody could act on has bought unfalsifiability and called it insight.
Named, not stated
No authority anywhere prescribes a number of why levels, and the repeated why is a discipline rather than a rule. The framework around event investigation does come from somewhere. The Basel Committee on Banking Supervision at the Bank for International Settlements, bis.org, publishes the operational risk framework and the seven event categories inside which an incident like this one is classified. The Reserve Bank of India, at rbi.org.in, sets what an Indian bank must actually do about investigating an operational risk event, what it must record and what it must report onward. The five rung chain, the reach counts, the 42 findings and the 9 without a cause all belong to one invented bank.
Sources
| Source | Document | Site |
|---|---|---|
| Bank for International Settlements | The Basel Committee on Banking Supervision publications setting out the operational risk framework and the seven event categories an incident is classified into | bis.org |
| Reserve Bank of India | What an Indian bank must actually do about investigating, recording and reporting an operational risk event | rbi.org.in |
| Taiichi Ohno | Toyota Production System, the account of the shop floor practice the repeated why is borrowed from | Productivity Press |
| James Reason | Human Error, 1990, the source of the layered defences picture the three near misses are read against | Cambridge University Press |
Vindhya Commercial Bank Limited is invented.
Educational material. Not advice on any investment, tax, budget or market position.
