Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Risk, Treasury & Financial Control
1Risk Foundations
Risk Appetite, Tolerance, Capacity…The Risk Taxonomy and UniverseRisk Register vs Risk MatrixStress TestingScenario Analysis vs Stress TestingImpact and LikelihoodLikelihoodThe Risk EventRisk Assessment
2Enterprise Risk Management
Enterprise Risk ManagementThe Four Risk TreatmentsRisk CultureRisk MaturityRisk Monitoring
3Risk Governance
Risk GovernanceHow to set a…The Risk PolicyThe Risk OwnerThe Risk Committee and Its CharterThe Risk Limit FrameworkRisk EscalationHow to set a…
4Credit and Counterparty Risk
Collateral AgreementsCollateral vs NettingProbability of DefaultExposureCounterparty ExposureConcentration Risk vs Wrong Way RiskCounterparty Risk vs Credit RiskHow to assess Counterparty ExposureHow to assess Concentration Risk
5Market Risk
Market RiskSensitivity MeasuresThe Hedging PolicyInterest Rate Risk in the Banking BookIRRBB vs Market RiskExpected ShortfallEconomic Value of EquityVaR BacktestingOpen PositionValue at RiskValue at Risk and Expected ShortfallEconomic Value SensitivityFX ExposureValue at Risk vs Expected ShortfallEarnings at Risk vs…FX Transaction Risk vs…How to measure Interest…How to measure Foreign…
6Liquidity Risk
Liquidity Stress TestingLiquidity Gap vs Liquidity BufferMaturity MismatchThe Debt Maturity ProfileFunding ConcentrationSurvival HorizonThe Contingency Funding PlanNet Stable Funding RatioLiquidity Risk vs Funding RiskLiquidity Coverage RatioLiquidity Gap and BufferHow to run a Liquidity Gap Analysis
7Operational Risk
Operational LossThe Loss EventRisk and Control Self AssessmentException ManagementInformation Security as a…Segregation of DutiesIssue ManagementThe Near MissRoot Cause Analysis in RiskThe Fraud TriangleCyber Risk vs Third Party RiskHow to run a…How to assess Third…
8Risk Reporting, Data and Model Risk
Model RiskModel Validation vs BacktestingHow to run Model ValidationData Governance in RiskModel Risk vs Data RiskKey Risk IndicatorsManagement InformationRisk ReportingRisk ScoreEarnings at RiskRisk Adjusted ReturnEarly Warning IndicatorsHow to build a KRI Dashboard
9Treasury
Corporate TreasuryAsset Liability ManagementIntragroup FundingThe Treasury PolicyThe Treasury Management SystemThe Cash ForecastCash Pooling and ConcentrationHow to build a Cash Forecast
10Financial Controls and Assurance
Control AssuranceThe Control LifecycleThe Assurance MapThe Audit FindingIssue RemediationInternal Financial ControlsControl Design vs Control EffectivenessHow to map Internal Financial ControlsHow to test Control…Control DeficiencyMaterial Weakness
11Operational Resilience
Operational ResilienceBusiness Continuity and Disaster RecoveryBusiness Continuity vs Operational…Crisis ManagementDisaster RecoveryIncident Management

Incident Management: Detect, Contain, Recover, Learn

Incident management is the standing process that runs on every disruption, large or small, in four stages. Detect: something is wrong and somebody records it with a severity. Contain: stop it getting bigger, a different thing from fixing it. Recover: get the service back to the people who use it. Learn: find out what the event shows and change something.

Three of those four stages have somebody waiting for them. The fourth has nobody. The asymmetry between the three stages somebody waits for and the one nobody waits for explains why an institution can run a process correctly thirteen times in a year and still be surprised by an event it had already been warned about. The learn stage is skipped not because anybody decides to skip it but because nothing goes wrong on the day it is skipped. Everything below is worked on Vindhya Commercial Bank Limited, an invented bank, and on the figures its own record carries.

Incident management demands no acquaintance with a data centre, a server room or any banking system. Every technical failure below is described the way a person outside experiences it: a screen that will not load, a counter that cannot pay out cash, a payment that leaves twice. Reading a failure from where the affected person stands is not a simplification of the method. Standing there is the method.

What is incident management, and what does it actually run on?

An incidentAny event that disrupts a service or could have, recorded and worked through in four stages. is any event that disrupts a service or could have. The definition is deliberately wide. The definition catches the four hour outage and it also catches the twenty minute wobble that three people noticed, the file that went out twice, the report that arrived with yesterday's numbers in it.

Incident management is not a plan for a bad day. Incident management is the ordinary, unglamorous machinery that runs on all of them. The test of whether an institution has an incident process is not whether it coped with its worst event; it is whether the same four stages ran on its smallest one. At the invented Vindhya Commercial Bank Limited the process ran on all thirteen operational loss events of the year, numbered I1 to I13, from the card fraud in month 1 to the complaints that were closed without a reply in month 12. Two of those thirteen, incidents I3 and I13, were also declared crises. Eleven were not, and the process still ran on them.

Think about a household for a moment. The shape is identical and the stakes are smaller. The tap in the kitchen starts leaking. Detect is noticing it and saying out loud how bad it is. Contain is putting a bucket under it and shutting the valve. The floor is saved and nothing is repaired. Recover is getting the water back on so that dinner can be cooked. Learn is the conversation nobody has, the one where somebody asks whether the other three taps in the house are the same age and were fitted by the same person on the same afternoon. The bucket gets used every time. The conversation almost never happens.

FOUR STAGES, AND THE QUESTION EACH ONE ANSWERS 1 DETECT Something is wrong. Somebody records it, with a severity. 2 CONTAIN Stop it getting bigger. Available before the cause is known. 3 RECOVER Get the service back to the people who use it. 4 LEARN What does this tell the institution, and what changes because of it? SOMEBODY IS WAITING FOR THIS ONE TO FINISH a customer, a desk, a committee, a supervisor NOBODY IS WAITING so nothing goes wrong today The first three stages are urgent because a person outside the institution is affected while they are unfinished. The fourth is never urgent on any single day, which is the entire mechanism by which it disappears.
The four stages in order, with the pressure that carries the first three and the absence of pressure that quietly drops the fourth.
Try it out

Name the four stages of incident management, in order.

How does something become an incident in the first place?

An event becomes an incident when somebody writes it down. Writing it down sounds trivial, and it is the single most fragile joint in the whole process. Everything downstream, the containment decision, the recovery, the review, the annual count, the picture the board eventually sees, rests on one person deciding that what they are looking at is worth recording.

Notice that the record has to start before anybody knows what happened. Detection is the discovery that something is wrong, not the discovery of why, and waiting for the why before opening a record is how an hour of the response disappears before it begins. At the invented bank, incident I2 in month 2 is the plainest example: a settlement instruction was sent twice and Rs 42 crore left the bank twice. At the moment somebody saw the second debit, nobody knew whether it was a system fault, a keying error, a duplicate approval or something worse. None of that was needed to open the record and none of it was needed to start chasing the money. Rs 41.4 crore came back and Rs 0.6 crore did not.

Detection arrives by four ordinary routes. An institution that has only one of the four is blind in three directions. A monitor notices, meaning a machine watching a number and complaining when it moves. A person inside notices, meaning a clerk seeing a total that cannot be right. A customer complains, the slowest route of the four and the one that has already cost the institution something by the time it arrives. Or a third party reports it, as happened with incident I13: the forgery ran for fourteen months and surfaced only when a beneficiary bank made a claim. A loss found by an outsider was invisible to every layer the institution had built, so the route by which an event is detected is itself information.

Derivatives Foundation Bootcamp — Fin Maverick

How is severity graded, and what does the grade read off?

SeverityHow serious an incident is, graded from what the service impact is rather than from what failed. is the label attached at detection that decides how much of the institution wakes up. The grade decides who is told, how fast, whether anybody is called at home, and whether the event gets a committee's attention or a line in a monthly list.

There is exactly one rule and it is broken constantly. Severity is read off the service impact, meaning which of the named services somebody outside could not use and for how long, and never off the technology impact, meaning how large the broken system was. Vindhya Commercial Bank Limited has named seven important business services, numbered S1 to S7, each with the board's own impact tolerance, and the grade is read against those.

Graded the other way, the answer comes out exactly backwards. A large piece of equipment failing at two in the morning while every service it sits under keeps running on the others is a big technology event and a small incident. Now turn it round. A supplier's equipment fails on a Saturday and tens of thousands of payments do not go through. The institution neither owns nor built that equipment, and the event is a small technology event and a very large incident. Incident I9 at this bank is the second of those: a vendor-hosted payment gateway failed for 9 hours and 48,000 transactions failed. The bank's own machinery was untouched.

THE SAME TWO EVENTS, GRADED TWICE GRADED FROM WHAT BROKE GRADED FROM THE SERVICE EVENT A, at two in the morning One machine among several fails outright. Every service sitting on top of it keeps running on the others. No customer is turned away and no counter stops paying. Minutes any named service was unavailable: zero. HIGH A whole machine is gone. Wake everybody up. LOW Nobody outside the bank went without anything. EVENT B, incident I9, a Saturday A payment gateway run by a supplier fails for 9 hours. 48,000 transactions fail. Nothing the bank built or runs itself was broken at any point in those 9 hours. Net loss booked: Rs 3.2 crore, incident I9. LOW None of the bank's equipment failed at all. HIGH 48,000 payments did not go through, for 9 hours. The left column ranks event A above event B. The right column ranks event B above event A. Only one of the two columns is describing what the outside world went through, and the grade that decides who gets woken up has to be that one.
Grading from the technology and grading from the service put the same two events in opposite orders, and only one order matches what customers experienced.

The two columns of that drawing are not a matter of taste. One of them is answering the question the institution exists to answer and the other one is answering a question about equipment. Note also the gap in the bank's own record: it locks a tolerance measurement for incident I3 and for no other event, so which of the seven services incident I9 touched, and by how much, is nowhere on the record. The measurement was simply not made, and saying so is more honest than producing a number that nobody at the bank ever computed.

Try it out

One server fails and a service that thousands of people use becomes unavailable. What is the severity graded from?

What does containing an incident mean, and why is it not fixing it?

ContainmentStopping an incident getting bigger, which can be done before anybody knows the cause. is stopping the thing getting bigger. Stopping the growth is the entire definition, and everything useful about containment follows from one property: containment is available before the cause is known, and repair is not.

Consider the plainest possible case. A payment file has gone out with the wrong account details on it, and nobody yet has any idea why. What is available this minute? The same file can be stopped from going out again. Whatever sits downstream can be stopped from acting on the entries already sent. The process that produced the file can be frozen so that it does not produce a second one while the fault is being looked for. Not one of those three actions requires anybody to understand the fault. All three of them reduce how large this event finally becomes.

Repair needs a diagnosis, and a diagnosis cannot be had this minute. Somebody has to find the fault, change something, and satisfy themselves that the change is right. Finding the fault and proving the change is real work on a real clock, and the whole of it happens after the cause is understood. An institution that waits for the diagnosis before doing anything at all has decided, without ever saying so, to let the incident keep growing for the length of the investigation.

The street version, again, has the same shape. A food stall's gas cylinder starts leaking. Containment is turning the valve off and moving the flame away. Anyone can do that in four seconds without knowing whether the fault is the washer, the pipe or the regulator. Repair is the mechanic, and the mechanic is not here yet. The people who get hurt in that situation are almost never the ones who could not diagnose the leak. The people who get hurt are the ones who waited to find out what was wrong before turning anything off.

TWO DECISIONS, TWO DIFFERENT STARTING MOMENTS SOMEBODY NOTICES THE CAUSE IS UNDERSTOOD CONTAIN Stop the file going out again. Stop anything downstream acting on it. Freeze the process that made it. None of this needs the cause. REPAIR Find the fault. Change it. Prove the change is right. TIME AN UNCONTAINED INCIDENT SPENDS GROWING Available to containment. Closed to repair. Time runs this way. No minutes are marked, because this bank never recorded how long its detection took. Containment can be decided in a minute by somebody who has no idea yet what went wrong, and it repairs nothing. Repair cannot begin at all until the diagnosis exists, and it is the only one of the two that ends the fault.
Containment and repair are separate decisions with separate starting moments, and the gap between the two is time only containment can use.
Try it out

A payment file has gone out with the wrong account details and nobody knows why yet. What is available right now?

What does the recover stage do, and what does it hand off to?

Recovery is getting the service back to the people who use it. Note the wording. The whole boundary of the recover stage is inside it. The unit is the service, not the machinery, and the finish line is a customer being able to do the thing again rather than a piece of equipment showing a healthy light.

The recover stage of incident management does not itself perform the recovery; it decides which arrangement performs it and it holds the clock while that arrangement works. Two arrangements sit underneath it and they are different things, each covered on its own. Keeping the service running by another route while the usual one is broken is a business arrangement: the counter takes a manual voucher, the payment goes through a second channel, the branch does the thing by hand. Restoring the technology so the usual route works again is a technology arrangement. One keeps the promise while the fault stands; the other removes the fault.

Both of those have designed target times at this invented bank, service by service, and both were tested in month 12. The month 12 test ran on a planned date, with the recovery team on standby, and the failure chosen in advance, so it tested the arrangement and not the surprise. Reading the result without that sentence beside it makes the arrangement look better than the record supports. Incident management adds the sequencing on top: which route is chosen, when, by whom, and at what point the institution says the service is back and starts the clock on everything that follows.

How to build an Incident Management Playbook

A playbookThe written procedure that says who does what at each stage, with the names filled in. is the written procedure that says who does what at each stage. Most institutions have a document with that title. Far fewer have one that would survive contact with a Saturday morning, and the difference between the two is not length or polish. The separation is that every step of a real playbook ends in a name or a number, and every step of a bad one ends in a principle.

"Incidents will be escalated promptly to the appropriate authority" is a sentence that has passed a hundred committee reviews and it tells nobody at three in the morning what to do. "Purnima Ganeshan, head of operational risk, or the duty officer named on the current roster, is reached on the number in section four and may declare" is a sentence that works. Build the second kind, in eight steps.

PB1
Name the services, with a tolerance each. Not the departments and not the equipment. The things somebody outside actually does: send money, take out cash, use the app, get a loan paid out, open an account, have a trade document issued, settle a treasury deal. At this invented bank there are seven, numbered S1 to S7, and the board has set an impact tolerance against each one. Ends in seven names and seven figures.
PB2
Set the severity scale off the service impact. Write the grades as sentences about who is affected and for how long, against those seven services. The person who finds the problem can then apply the grade without understanding the machinery. Ends in a grade that anybody can read off without a technical opinion.
PB3
Name the roles, and name who covers each one out of hours. Covering every hour is what separates a document from a procedure. Every role in the playbook needs a person against it at every hour of every day, including the hours nobody wants. A role with a name only from ten to six is not covered; it is covered for a third of the week.
PB4
Write the rule that says who declares, and at what point. One named role, one trigger, no committee required to convene before anything can start. Almost every plan leaves this step out, and the missing declaration rule is worked through below. Ends in a role and a threshold, both written down.
PB5
List the containment actions that need no diagnosis. Stop the file. Suspend the interface. Block the account. Hold the batch. Take the channel offline. Write the list in advance precisely because the people who will use it are going to be reading it under pressure with incomplete information. Ends in a list somebody can act from without thinking.
PB6
Set the recovery route for each service, and say who chooses it. For each of the seven services, the alternative way of serving the customer while the usual way is broken, and the route back to the usual way. Both of those are built and tested elsewhere; the playbook's job is to say which one is invoked and by whom. Ends in a route per service.
PB7
Define the record kept while the clock is still running. Five fields, filled in as the incident happens rather than written up afterwards, for reasons that get their own section below. Ends in a form somebody is actually holding at the time.
PB8
Define the learn stage, with an owner and a date on every action. When the review happens, who runs it, what it must search, and what it hands over. Without a name and a date attached to each thing it decides, the review produces a document instead of a change. Ends in an owner and a date.
EIGHT STEPS, AND WHAT EACH ONE HAS TO END IN STEP WHAT THE STEP SAYS WHAT IT MUST END IN PB1 Name the services somebody outside relies on, with a tolerance each 7 service names and 7 tolerances PB2 Set the severity scale off service impact, never off what broke A grade anybody can read off PB3 Name the roles, and name who covers each one out of hours A person named at every hour PB4 Write the rule that says who declares, and at what point One role, and one trigger PB5 List the containment actions that need no diagnosis at all A list somebody can act from PB6 Set the recovery route for each service, and who chooses it A route for each of S1 to S7 PB7 Define the record kept while the clock is still running Five fields, filled at the time PB8 Define the learn stage, with an owner and a date attached An owner and a date, per action Not one of the eight ends in a principle. Every one ends in something a reader could point at in the document and check, which is the only difference between a procedure and a description of good practice.
The eight steps of the playbook, each paired with the concrete output it has to produce, because a step ending in a principle cannot be followed under pressure.
Try it out

Of the eight steps, which one is most often missing from a real plan, and where does its absence first show up?

What does a playbook usually leave out, and where does it get tested first?

The missing step is almost always PB4. Plans are generous with structure and stingy with authority. Plans name the team, the meeting and the reporting line, and then leave the one sentence that starts everything unwritten. Writing it means choosing a person, and choosing a person is uncomfortable.

Vindhya Commercial Bank Limited had exactly this hole. Its plan named a crisis team of 7 roles and did not name who could convene it. Nobody noticed. In office hours the question never arises: a senior person is present, somebody takes charge, the machinery starts. A missing declaration rule is invisible for as long as somebody senior happens to be in the building.

Then incident I9 arrived on a Saturday. A vendor-hosted payment gateway failed, 48,000 transactions failed with it, and the disruption ran 9 hours. No crisis was declared, and the reason was not capability, not equipment, not budget and not skill. The failure began on a Saturday and no single named person held the authority to declare a crisis outside working hours. The plan named a crisis team of 7 roles and did not name who could convene it. One unwritten sentence left the longest disruption of the year running without the escalation the bank had already designed for it.

Declaring a crisis, running one, and deciding under pressure with incomplete information are covered separately. The joint between the two is narrower: incident management runs on every one of the thirteen events, and PB4 is the step where it says, in advance and in writing, at what point one of those thirteen stops being handled by the standing process alone. An institution can hold a complete crisis capability and still never use it, if nothing in the standing process is allowed to switch it on.

In practice

Who reads a playbook, and what they are checking for

An internal auditor handed an incident management playbook does not read it front to back. The auditor goes to PB3 and PB4 first and asks two questions. Is there a person's name against every role at three in the morning on a Sunday, and is there one named role that can start the escalation without a meeting? If either answer is no, the rest of the document is a description of intentions. Rustom Batliwala's audit function at this invented bank would have found the gap by reading those two steps.

A supervisor reads it differently again. The supervisor is not looking for elegance but for evidence that the institution can say what happened, when, and to whom it was reported. PB7, the record kept while the clock runs, therefore matters far more to an outsider than it looks: it is the only thing that can be produced afterwards to show how the institution behaved during the event rather than how it describes itself now.

And a household version, genuinely useful in domestic form. If one person in a house handles every emergency, the house has a crisis team of one role and no cover. The test is not whether that person is competent. The test is what happens on the weekend that person is on a train with no signal, and whether anybody else knows which valve to turn, which number to call, and whether they are allowed to spend money without asking first. Spending money without asking is PB4, at a kitchen table.

Breaking Into Quants Bootcamp — Fin Maverick

What has to be in the incident record while the clock is still running?

Step PB7 asks for a record kept during the event rather than after it, and the reason is not administrative tidiness. Four of the five things that record has to carry cannot be recovered afterwards at all, and the fifth would have survived anyway.

Here is what is in it. What was seen, and when. What was believed at each decision point. What was done, and by whom. Which services were affected, and for how long. And what was told to whom. The five fields deserve a second reading with one question in mind: which of them could honestly be written down a week later?

The second one is the killer. A week after an incident everybody knows the cause, and a record written then describes every decision as though the cause was known at the time. The distortion is enough on its own to destroy the learn stage before it starts. A decision that looks obviously wrong in hindsight was often entirely reasonable given what the person actually had in front of them, and a decision that looks fine may have been luck. Neither judgement can be made without knowing what was believed at the moment, and nobody can remember what they believed once they know what was true.

THE FIVE FIELDS, AND WHICH OF THEM SURVIVE BEING WRITTEN LATER INCIDENT RECORD, FILLED IN WHILE THE CLOCK IS RUNNING 1. What was seen, and when the second debit for the same settlement instruction is noticed CANNOT BE RECONSTRUCTED memory quietly reorders what was seen 2. What was believed at each decision this looks like a duplicate release rather than a fault in the system CANNOT BE RECONSTRUCTED a week later everybody knows the cause 3. What was done, and by whom the interface was suspended, then the receiving bank was telephoned CANNOT BE RECONSTRUCTED the order of the actions blurs together 4. Which services were affected, and for how long measured against the named services S1 to S7, minute by minute CANNOT BE RECONSTRUCTED the minutes are gone once the day ends 5. What was told to whom, and when the committee note, the customer message, the report to the supervisor SURVIVES A LATER WRITE UP letters and messages are filed anyway Four of the five fields are worth nothing if the form is filled in next week. The fifth would have survived anyway, which is exactly why a record written afterwards feels complete while carrying almost none of what the learn stage needs.
Four of the five fields in an incident record cannot be reconstructed honestly after the event, which is why the record is a live document rather than a report.
Try it out

Why does the record of what was believed at each decision have to be written while the incident is still running?

What happens in the learn stage, and why is it the one that gets skipped?

The learn stage is a meeting, and the meeting has a name: the post incident reviewThe meeting after the recovery where the learn stage actually happens, with an owner and a date attached to what it decides.. The review happens after the service is back, and nobody is waiting for it. Nobody waiting is the whole problem, and it is worth sitting with rather than rushing past.

Every other stage has a person on the other side of it. Detect has whoever is affected. Contain has the growing loss. Recover has the customer, the desk, the counter, the phone line. Skip any of the three and the consequence lands the same day, loudly, on somebody who will say so. Skip the learn stage and absolutely nothing happens. No part of a process can hold a more dangerous property than that one. The cost arrives months later, attached to an event that looks entirely unrelated, and by then nobody connects the two.

So what does the review actually do, beyond producing a document? Two things. The review decides what changes, with a name and a date against each change. Step PB8 is written the way it is for exactly that reason. And it searches. The search is the part almost nobody does, and it is the part that carries all the value. The search asks a question that has nothing to do with the incident in front of it: has this same failure appeared before, somewhere in the record, in a form that did not cost anything?

What is a near miss worth, and what makes it worth anything?

A near missAn event that would have been an incident if one more thing had not held. is an event that would have been an incident if one more thing had not held. Something went wrong, something else caught it, and no loss was booked. Most institutions record them, and the register is usually short and tidy.

Here is the uncomfortable claim. Recording a near miss is worth nothing at all on its own; the entire value sits in the linkingSearching the record for the same failure appearing before, which is the only thing that makes a near miss worth recording., meaning the search for the same failure appearing before or elsewhere. A register nobody searches is a filing cabinet. A near miss carries almost the same information as a loss and costs nothing to collect, so a register searched at every post incident review is the cheapest source of information an institution has.

Vindhya Commercial Bank Limited recorded five near misses in the year, numbered N1 to N5. In month 2, near miss N1: a second duplicate settlement instruction, this one for Rs 68 crore, was stopped by the four eyes check before release. In month 4, near miss N2: a payment file of 1,240 salary credits was queued against the wrong account and caught at reconciliation before value date. In month 6, near miss N3: the collateral valuation feed was stale for 2 working days and a data quality check caught it. In month 7, near miss N4: a trade finance document set carrying a forgery pattern was refused by a checker. In month 10, near miss N5: a privileged access account belonging to a leaver stayed live for 46 days and was found by the quarterly access review before it was used.

Read that list once more and notice that in every single case something worked. Something working is what a near miss is. And then notice the two that matter: two of the five, near misses N3 and N4, turned out to be the same failure as an incident that followed, and in both cases the bank had the warning in writing before the loss arrived.

Try it out

Two of this bank's five near misses were the same failure as a later incident. Before the answer arrives: in how many of those two had a control failed?

The failure: the warning arrived twice, in writing, and both times a control had worked

In month 6, a data quality check found that the collateral valuation feed was stale for 2 working days. The check did precisely what it was built to do. The event was recorded, it was closed, and nobody raised it as an issue. Four months later, in month 10, the same feed was stale for 11 working days and 340 loans were wrongly marked. The month 10 event is incident I10: gross Rs 1.4 crore, no recovery, net Rs 1.4 crore, and it is the one finding at this bank rated a material weakness. No customer lost money.

In month 7, a checker refused a trade finance document set because the documents did not hold up. The checker did precisely what a checker is for. The refusal was recorded as a routine refusal, exactly what it looked like, and it was never linked to anything. One month later, in month 8, incident I13 surfaced when a beneficiary bank made a claim: 9 letters of credit issued against forged shipping documents over fourteen months, gross Rs 22.4 crore, recovery Rs 7.0 crore, net Rs 15.4 crore. Incident I13 is the largest net loss of the year.

In neither case did a control fail, and that is the point. A check worked and a person worked. Both events were recorded correctly and both were closed correctly, and the closure is where the institution lost the information. The people involved are not the failure here, and it matters to say so plainly: they caught what they were there to catch and they wrote it down. The missing part was a route by which a refusal at a trade finance desk, or a data quality alert on a valuation feed, could reach anybody whose job was to ask whether the same failure was getting through somewhere else. There was no such route and no such question, so a correct action produced a correctly closed record and stopped.

Now the arithmetic, stated carefully. Incident I10 booked a net loss of Rs 1.4 crore and incident I13 booked a net loss of Rs 15.4 crore. Together that is Rs 16.8 crore against the year's total net operational loss of Rs 43.8 crore across the thirteen events, or 38.4 per cent of the year. Rs 16.8 crore is the loss the learn stage would have been looking at, and it is not loss that would automatically have been avoided. Nobody can honestly say the fraud would have been stopped or the feed fixed in time. Linking a near miss to a pattern does not by itself prevent the incident that follows; it puts a question in front of somebody who was otherwise never going to be asked it. The difference between those two statements is the difference between an honest claim and a comfortable one.

One more measurement. The timing is what the simulation below turns into a dial. Near miss N3 is month 6 and incident I10 is month 10. The link is 4 months wide. Near miss N4 is month 7 and incident I13 was discovered in month 8. The link is 1 month wide. A review that searched only one month back would have reached one of the two links, and it happens to be the one sitting in front of Rs 15.4 crore, or 35.2 per cent of the year on its own.

TWO WARNINGS THIS BANK ALREADY HAD IN WRITING NEAR MISS N3, month 6 The collateral valuation feed was stale for 2 working days. A data quality check caught it. THE CHECK WORKED. CLOSED, NO ISSUE RAISED. INCIDENT I10, month 10 The same feed, stale for 11 working days. 340 loans were wrongly marked. Net Rs 1.4 crore 4 MONTHS no route existed between these two records NEAR MISS N4, month 7 A trade finance document set carrying a forgery pattern was refused by a checker. THE CHECKER WORKED. FILED AS A ROUTINE REFUSAL. INCIDENT I13, month 8 9 letters of credit against forged shipping documents, over fourteen months. Net Rs 15.4 crore 1 MONTH no route existed between these two records Both warnings were correct actions correctly recorded. What was missing was not a control and not a person; it was a standing question at the review asking whether this same failure had already appeared somewhere that cost nothing.
Both missed links were cases of a control working, so the failure was never in detection at all but in what the institution did with a success.
Debt Capital Markets Bootcamp — Fin Maverick

Why does a near miss carry the same information as a loss?

There is a model that makes this obvious, and it is worth having because it changes what a near miss looks like. James Reason, in Human Error, 1990, described the protection an institution builds as a stack of layers, each one imperfect and each one with holes in it. A design rule, an automated check, a reconciliation, a supervisor's review, a person who refuses to sign. On any given day the holes sit in different places, and an event reaches the outside world only when the holes in every layer happen to line up.

Layered defencesThe idea that protection is a stack of imperfect layers and an event reaching the outside only when the gaps line up. reframe the near miss completely. A near miss is not a smaller event than a loss; it is the same event, and it found a hole in every layer except one. The near miss has already mapped the holes. Finding out where the gaps are is the expensive part of the work, and the near miss has done it without anybody paying for the finding out. The only thing it did not do was cost money, and money is the one part of a loss event that teaches nothing about the failure.

Which is why the closure of near miss N4 is such a costly piece of tidiness. A document set carrying a forgery pattern reached a checker's desk, having already passed everything in front of the checker. The checker was the last layer, and the last layer held. Every layer before it had a hole in exactly the right place. A refusal form is designed to record a refusal, so the form recorded none of that.

LAYERED DEFENCES, AFTER JAMES REASON, HUMAN ERROR, 1990 NEAR MISS N4 STOPPED BY THE LAST LAYER every layer in front of it had a hole INCIDENT I13 REACHED THE OUTSIDE net Rs 15.4 crore, incident I13 LAYER 1 how the process is built LAYER 2 an automated check LAYER 3 a reconciliation LAYER 4 a supervisor's review LAYER 5 a person who refuses The two paths are the same failure. One of them found a hole in every layer and one of them found a hole in all but the last, so the cheap one has already mapped the holes and the expensive one adds nothing about where they are.
A near miss is the same failure as a loss stopped by the last layer instead of the first, so it maps the holes in every layer at no cost.
Try it out

A control catches a problem before it costs anything, and the record is closed correctly. What should happen next?

Document Extraction in Finance — free micro-course from Fin Maverick

How far back should the learn stage look?

If the search is the valuable part, then the obvious next question is how far back it goes, and this is where the argument stops being a matter of principle and becomes a matter of arithmetic. A review that searches nothing finds nothing, and that is what this bank did. A review that searches a month finds one of this year's two links. A review that searches four months finds both.

Two numbers frame it. Incident I13, at a net loss of Rs 15.4 crore, is 35.2 per cent of the year's Rs 43.8 crore on its own. Incident I10 adds Rs 1.4 crore, taking the pair to Rs 16.8 crore and 38.4 per cent. A one month lookback reaches 35.2 of those 38.4 points, so the short search is not a proportionally small answer; it is nearly the whole of it. The control below sets how far back the learn stage searches.

Try it out

The learn stage searches one month back for the same failure. Before the control is moved: how many of the two real links in this year does it find?

Play with it

How far back does the learn stage search?

One control: how many months back a post incident review searches the near miss record for the same failure. One consequence: how many of the two real links in this year's record it reaches, and how much of the year's net operational loss those links sit in front of. The control starts at zero months, the setting this bank actually ran on.

Lookback: 0 months, from 0 to 12

THE YEAR, MONTH 1 TO MONTH 12 1 2 3 4 5 6 7 8 9 10 11 12 SEARCH AT THE REVIEW OF INCIDENT I13 SEARCH AT THE REVIEW OF I10 INCIDENTS I1 I2 I3 I4 I5 I6 I7 I9 I10 I11 I12 I8 I13 NEAR MISSES N1 N2 N3 N4 N5 Near miss N3, month 6, is the same failure as incident I10, month 10, a gap of 4 months. Near miss N4, month 7, is the same failure as incident I13, discovered month 8, a gap of 1 month. The other three near misses match no incident in this year. SEARCHING 0 MONTHS BACK: 0 OF THE 2 REAL LINKS REACHED
Lookback
0 months
Links reached
0 of 2
Net loss downstream
Rs 0.0 crore
Share of the year
0.0 per cent

Searching 0 months back, the learn stage reaches 0 of the 2 real links in this year's record, sitting in front of Rs 0.0 crore of the year's Rs 43.8 crore of net operational loss. This is what this bank actually did: near miss N3 was closed as routine in month 6 and near miss N4 was closed as a routine refusal in month 7.

Educational illustration. Reaching a link is not the same as preventing the incident that followed: the figure shown is loss the learn stage would have been looking at, and none of it is loss that would automatically have been avoided. The two links, the five near misses N1 to N5, the thirteen incidents I1 to I13 and the Rs 43.8 crore annual net operational loss are the invented bank's own locked figures. The lookback is set by the control above and is not a figure from the case.

A STEP WITH TWO RISERS, AND THE FIRST ONE CARRIES ALMOST ALL OF IT Vertical: how many of the two real links in this year the search reaches. Horizontal: how many months back it searches. 0 1 2 0 1 2 3 4 5 6 7 8 9 10 11 12 0 of 2 links. Rs 0.0 crore. This is what the bank actually did. 1 of 2 links. Rs 15.4 crore, 35.2 per cent of the year, reached from one month back. 2 of 2 links. Rs 16.8 crore, 38.4 per cent of the year, from four months back. MONTHS THE LEARN STAGE SEARCHES BACK The first riser is one month wide and carries 35.2 of the 38.4 points. The second riser takes three more months of searching and adds 3.2 points. This is loss the learn stage would have been looking at, and not loss that would have been avoided.
How far back the search reaches decides how many links it finds, and the first month of lookback carries almost the whole of the value.

Read the shape for a moment rather than the numbers. There are two risers and they are nothing like the same size. The first sits between a lookback of nothing at all and a lookback of one month, and it carries 35.2 percentage points of the year with it. The near miss N4 and incident I13 link is only one month wide, and the net loss behind incident I13 is Rs 15.4 crore. The second riser waits until four months and then adds 3.2 points. The near miss N3 and incident I10 link is four months wide, and the net loss behind incident I10 is Rs 1.4 crore. Almost the whole of what there was to reach in this record is reachable by a search that goes back a single month. Most people guess the opposite when they are asked.

The shape matters because the objection to a linking rule is always cost. Searching a register is somebody's afternoon, and the longer the window the more afternoons it takes, so a rule that demanded five years of searching would lose the argument in the room and deserve to. Here the shape says something else. The cheapest version of the rule, one month, already reaches most of what was there. The limit of that is plain: this is one invented bank's twelve months and not a general law about how far apart a warning and a loss sit. The method travels rather than the answer: measure the gaps in an institution's own record before arguing about the size of the window.

A step chart is persuasive in a way that a paragraph is not, so the caveat has to be repeated beside it. Every rupee on that chart is loss the learn stage would have been looking at, and none of it is loss that would automatically have been avoided. A review in month 8 that had pulled up the month 7 refusal would have had a question in front of it, not a conclusion. The review might have found the fourteen month pattern quickly, or it might have looked at one refused document set, agreed that the checker had done the right thing, and moved on. Linking changes who gets asked, and how early. Linking does not change what they then decide, and letting the two blur sells a comfort rather than a method.

There is also a practical reason the search fails even when somebody actually does it, and near miss N4 shows it exactly. The near miss was filed as a routine refusal. Incident I13 was booked as letters of credit issued against forged shipping documents. The two records describe the same failure and share almost no words, so a search run on the words would have returned nothing and the reviewer would have concluded, correctly by their own method, that there was no prior warning. A register becomes searchable only when the failure is described rather than the outcome. The refusal form has to have somewhere to say what was wrong with the documents and not only that they were refused. Adding that field is a design decision about a form, taken in a quiet month, and it separates a register that can be linked from a register that merely exists.

Backtesting a Strategy teaches you to build a backtest, name how it flatters itself, and state what the result establishes.

How do incident, crisis, continuity and recovery fit together?

Four words get used as though they were interchangeable, and they are not. The four words answer four different questions, they run on different numbers of events, and different people are waiting for each of them. Setting them side by side one final time is the fastest way to place any real event, including one already under way.

Incident management is the standing process and it runs on everything. Its question is: something is wrong, what is done about it, in four stages. At this invented bank it ran on all thirteen recorded events, incident I1 through to incident I13. Crisis management asks a narrower and harder question: the decisions have outrun the facts, so who decides, on what, and with what authority. Crisis management ran twice in the year, on incident I3 and on incident I13, and it sits on top of the standing process rather than replacing it. Business continuity asks how the customer keeps getting served while the usual route is broken: a business arrangement about counters, channels and manual routes. Disaster recovery asks how the usual route itself comes back: a technology arrangement about systems. Each of the last three is covered separately. The joint is this: the recover stage of incident management chooses between the last two and holds the clock while they work, and step PB4 is where the standing process is allowed to switch crisis management on.

FOUR WORDS, FOUR DIFFERENT QUESTIONS, AND ONE MORNING CAN NEED ALL FOUR Read across a row: the question each process answers, what it runs on, and who is waiting for it. THE PROCESS THE QUESTION IT ANSWERS WHAT IT RUNS ON WHO IS WAITING FOR IT INCIDENT MANAGEMENT Something is wrong. What is done about it, in four stages? All thirteen of the year's recorded events, incident I1 to incident I13. Whoever is affected, from the first minute onward. CRISIS MANAGEMENT The decisions have outrun the facts. Who decides, and on what? Two of the thirteen: incident I3 and incident I13 were declared crises. A named person who can say it has started. BUSINESS CONTINUITY The usual route is broken. How does the customer still get served? One service at a time, by a route built in advance. The customer at the counter, the app or the phone. DISASTER RECOVERY The technology is down. How does the usual route come back? The systems sitting underneath the service. Everything above it, until the work is finished. ONE SATURDAY MORNING, INCIDENT I9: A VENDOR RUN PAYMENT GATEWAY, 9 HOURS, 48,000 FAILED TRANSACTIONS INCIDENT MANAGEMENT Running from the first minute, on this one as on all thirteen. CRISIS MANAGEMENT Never declared. No sentence said who could, outside working hours. BUSINESS CONTINUITY How does a customer still pay while the gateway is down? DISASTER RECOVERY How does the gateway itself come back, and by when?
The four processes answer four different questions and run on different numbers of events, and one Saturday morning can put all four questions in play at once.

Now put one morning through the grid, month 9. Incident I9: a payment gateway run by an outside supplier failed for 9 hours and 48,000 transactions failed with it. Incident management runs on every disruption, so it was running from the first minute. Continuity had a live question: how did somebody who wanted to pay for something still pay for it while the gateway was down? Recovery had a live question: how did the gateway itself come back, and by when? And crisis management had a live question too: who was going to hold the decisions that could not wait? Three of the four questions had somebody attached to them and the fourth did not. No crisis was declared, and no sentence in the plan said who could declare one on a Saturday. One morning, four questions, and the one that went unanswered was the one that only needed a line of writing.

Try it out

A payment gateway run by an outside supplier fails for nine hours on a Saturday. Which of the four processes belong in that morning?

In practice

Who asks for this from outside, and what they are really asking

Girish Talwalkar, group treasurer at the invented Nirjhar Industries Limited, is on the other side of the counter from all of this. He does not care which of the four words his bank uses. He has one question: will a payment he has to make on a fixed date be made? And one follow up: who will tell him if it will not, and how soon? The follow up question is answered entirely by two steps of the playbook: PB3, whether a role is covered at the hour his payment runs, and PB4, whether anybody is allowed to start the escalation that ends in somebody calling him. A customer never sees a playbook. A customer sees whether somebody rang.

Purnima Ganeshan, head of operational risk at the invented bank, uses it differently. Her monthly pack carries the thirteen loss events I1 to I13 and the five near misses N1 to N5. For almost no cost she can add one column against every near miss saying what was searched and what was found. The column is the only evidence anywhere that the learn stage happened at all. A review with no record of its search is indistinguishable from a review that did not search, and after twelve months the difference between the two is Rs 16.8 crore of the year sitting in a place nobody looked.

The household version is smaller and the same. A smoke alarm that goes off during cooking has just said something, and the something is not that the alarm works. The alarm has said that the kitchen fills with smoke at a rate that reaches the ceiling sensor, worth knowing before the night nobody is standing there. Nobody in any house has ever written that down, and the reason is exactly the reason a bank does not: nothing went wrong, so there was nothing to record, so there was nothing to search.

India

What the Indian rules require

The four stages are universal and so is the playbook. Local rules decide what an institution must record and report to somebody outside when an operational or technology incident is serious. For an Indian bank the Reserve Bank of India at rbi.org.in is the source of that: what has to be reported, to whom, in what form and inside what time, and what is expected when an outside supplier rather than the institution itself ran what failed. The Bank for International Settlements at bis.org is where the Basel Committee publishes the international standard those expectations implement, including the seven event categories the invented loss record uses as names. Where the entity is a market intermediary rather than a bank, the Securities and Exchange Board of India at sebi.gov.in applies instead.

Reporting windows, thresholds, forms and effective dates must be confirmed directly with the issuing body before any of them is relied on.

Incident management here means the standing process that runs on every disruption, and the playbook that makes it followable. Declaring a crisis, convening a team and deciding under pressure with incomplete information are covered separately. Crisis management sits on top of the standing process and ran on two of the thirteen events. The standing process ran on all thirteen. Keeping a service running by another route, and restoring the technology behind it, are each covered separately, and the recover stage chooses between them and holds the clock rather than performing either. How an operational loss is defined, measured, categorised and booked is covered separately, and the loss record is taken here as given. How a service and a board impact tolerance are set, and what the month 12 continuity test found, are settled separately. Root cause analysis as a method, issue management and remediation tracking, and how a control is designed and tested are all covered separately and are named here only where the process hands over to them. Financial instruments, including the letters of credit named in incident I13, are covered separately.
Risk Management Program Bootcamp — Fin Maverick

Sources

SourceDocumentSite
Reserve Bank of IndiaExpectations on a regulated bank covering the recording and reporting of a serious operational or technology incident, and arrangements run by an outside supplierrbi.org.in
Bank for International SettlementsThe Basel Committee's international standard on operational risk management and operational resilience, and the seven loss event categories the invented record uses as namesbis.org
Securities and Exchange Board of IndiaExpectations where the entity is a market intermediary rather than a banksebi.gov.in
Indian Banks AssociationMaterial on Indian banking operational conventioniba.org.in
James ReasonHuman Error, 1990, the layered defences model of protection as a stack of imperfect layers whose gaps occasionally line upcambridge.org

Vindhya Commercial Bank Limited, Nirjhar Industries Limited and every person named here are invented.
Educational material. Not advice on any investment, tax, budget or market position.

Covered in this topic

Subtopics

How to build an Incident Management Playbook
← Previous
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.