Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Risk Management Program · CoreTrack
1Risk, Treasury & Financial Control
iRisk Foundations
Risk Appetite, Tolerance, Capacity…The Risk Taxonomy and UniverseRisk Register vs Risk MatrixStress TestingScenario Analysis vs Stress TestingImpact and LikelihoodLikelihoodThe Risk EventRisk Assessment
iiEnterprise Risk Management
Enterprise Risk ManagementThe Four Risk TreatmentsRisk CultureRisk MaturityRisk Monitoring
iiiRisk Governance
Risk GovernanceHow to set a…The Risk PolicyThe Risk OwnerThe Risk Committee and Its CharterThe Risk Limit FrameworkRisk EscalationHow to set a…
ivCredit and Counterparty Risk
Collateral AgreementsCollateral vs NettingProbability of DefaultExposureCounterparty ExposureConcentration Risk vs Wrong Way RiskCounterparty Risk vs Credit RiskHow to assess Counterparty ExposureHow to assess Concentration Risk
vMarket Risk
Market RiskSensitivity MeasuresThe Hedging PolicyInterest Rate Risk in the Banking BookIRRBB vs Market RiskExpected ShortfallEconomic Value of EquityVaR BacktestingOpen PositionValue at RiskValue at Risk and Expected ShortfallEconomic Value SensitivityFX ExposureValue at Risk vs Expected ShortfallEarnings at Risk vs…FX Transaction Risk vs…How to measure Interest…How to measure Foreign…
viLiquidity Risk
Liquidity Stress TestingLiquidity Gap vs Liquidity BufferMaturity MismatchThe Debt Maturity ProfileFunding ConcentrationSurvival HorizonThe Contingency Funding PlanNet Stable Funding RatioLiquidity Risk vs Funding RiskLiquidity Coverage RatioLiquidity Gap and BufferHow to run a Liquidity Gap Analysis
viiOperational Risk
Operational LossThe Loss EventRisk and Control Self AssessmentException ManagementInformation Security as a…Segregation of DutiesIssue ManagementThe Near MissRoot Cause Analysis in RiskThe Fraud TriangleCyber Risk vs Third Party RiskHow to run a…How to assess Third…
viiiRisk Reporting, Data and Model Risk
Model RiskModel Validation vs BacktestingHow to run Model ValidationData Governance in RiskModel Risk vs Data RiskKey Risk IndicatorsManagement InformationRisk ReportingRisk ScoreEarnings at RiskRisk Adjusted ReturnEarly Warning IndicatorsHow to build a KRI Dashboard
ixTreasury
Corporate TreasuryAsset Liability ManagementIntragroup FundingThe Treasury PolicyThe Treasury Management SystemThe Cash ForecastCash Pooling and ConcentrationHow to build a Cash Forecast
xFinancial Controls and Assurance
Control AssuranceThe Control LifecycleThe Assurance MapThe Audit FindingIssue RemediationInternal Financial ControlsControl Design vs Control EffectivenessHow to map Internal Financial ControlsHow to test Control…Control DeficiencyMaterial Weakness
xiOperational Resilience
Operational ResilienceBusiness Continuity and Disaster RecoveryBusiness Continuity vs Operational…Crisis ManagementDisaster RecoveryIncident Management

Operational Resilience: Staying Available Through Disruption

Operational resilience is the ability to keep a service available to the people outside the institution while something has gone wrong inside it. Resilience assumes the failure happens. The institution names the services that matter to somebody outside, sets a board level tolerance for how long the outside world can be without each one, maps what the service leans on, and then tests whether the service stays inside that tolerance.

One change of unit carries the whole idea, and it is the thing to carry away above everything else here. Operational risk asks what could go wrong inside the institution and what it would cost. Operational resilience asks what somebody outside experiences while it is going wrong, and for how long. The moment the unit becomes a service rather than a system, a department or a loss figure, every other rule in this guide follows without further argument. One invented bank, Vindhya Commercial Bank Limited, carries every worked figure below.

How is this different from just trying to stop things going wrong?

Most of what an institution does about failure is an attempt to prevent it. The instruction gets checked twice. A second person approves it. The change is tested before it goes live. All of that is worth doing, and none of it is resilience.

Operational resilienceThe ability to keep a service available to people outside the institution while something inside it has failed. starts by conceding the argument. Resilience assumes that prevention will not hold every time, that something will break, and that the interesting question begins at that moment rather than ending there. The premise is not that the failure can be avoided but that the outside world can be kept whole while it happens.

Think about a wedding caterer. Preventing failure means buying good gas cylinders and checking the burners twice. Resilience means something else entirely: the cylinder runs out anyway, and the question is whether four hundred people sit down to dinner at the hour they were promised, or whether they sit down at eleven. Two different questions, two different sets of preparations, and one of them is about the guests rather than about the kitchen. A bank asking a resilience question is asking about the guests.

Measuring from outside the institution is also why a disruptionAny event that stops a service reaching the person who uses it, whatever caused it. is defined without any reference to its cause. A power failure, a supplier who stops answering, a change that was applied badly, a flood at a branch: to the person who could not withdraw cash, these are all the same event. Resilience is measured from where that person stands, so the cause matters enormously for fixing it and not at all for measuring it.

Derivatives Foundation Bootcamp — Fin Maverick

What counts as an important business service, and who decides?

An important business serviceA service somebody outside the institution relies on, named in that person's words rather than in the institution's. is a thing somebody outside the institution actually uses, described the way that person would describe it. Not a department. Not a building. Not a piece of machinery. The test is whether a customer would recognise the name, and if they would not, it is not a service.

Vindhya Commercial Bank Limited has named seven of them, numbered S1 to S7. Read the names and notice what they are: sending money, taking cash out at a counter, using the phone app, getting a loan paid out, opening an account, having a trade document issued, and settling a treasury deal. Every one is a verb somebody outside the bank performs. None of them is the name of a system.

The distinction between a service and a system is not pedantry, and the drawing below is why. A service and a system are not the same shape. One system can sit underneath several services at once, and one service can lean on several systems at once. Making the system the unit protects the machinery and loses sight of the promise. Worse, nobody outside has ever heard of the machinery, so there is no way to say what the outside world lost.

SEVEN SERVICES SOMEBODY OUTSIDE THE BANK USES, NUMBERED S1 TO S7 S1 Payments and remittances S2 Branch counter and cash S3 Internet and mobile banking S4 Loan disbursal S5 Deposit account opening S6 Trade finance issuance S7 Treasury settlement CARD AND COUNTER SWITCH CORE BANKING SYSTEM PAYMENT GATEWAY run by an outside supplier, and still inside the promise DOCUMENT AND WORKFLOW ESTATE THE MACHINERY UNDERNEATH, WHICH NOBODY OUTSIDE THE BANK HAS EVER HEARD OF Solid red lines: the four services the core banking system feeds. Those are the four that went dark together in incident I3. Dashed grey lines: everything else the seven services lean on. One service can sit on more than one of them.
Seven named services sit above the machinery rather than inside it: the core banking system failing in incident I3 reached payments, the branch counter, internet and mobile banking and treasury settlement all at once, because a service is a thing a customer uses and a system is a thing several services lean on.

Notice one more thing in that drawing. The payment gateway is run by an outside supplier and it still sits underneath a service the bank has promised. Handing the work to somebody else moves the machinery out of the building and leaves the promise exactly where it was. Vindhya Commercial Bank Limited found this out the hard way in month 9, when a supplier hosted payment gateway failed for nine hours and 48,000 transactions failed with it. The gateway failure is incident I9 in the bank's loss record.

Try it out

A bank lists its important business services as core banking, the card switch and the payments hub. What has gone wrong?

Building a Client Risk Profile — free micro-course from Fin Maverick

Who sets an impact tolerance, and what makes it a promise?

Once the services are named, somebody has to say how long the outside world can be without each one. The agreed length is the impact toleranceThe longest the board is willing for the outside world to be without a named service., and at Vindhya Commercial Bank Limited it is set by the board, committee G1, and by nobody else. An impact tolerance is a business decision about what the outside world can bear, not a description of what the equipment can currently manage.

The difference between those two sentences is the whole reason the board sets it. If the people who run the machinery set the figure, the figure becomes a statement of what the machinery already does. The figure was written to be met, so it will always be met. A tolerance set by the board can be missed. A missable tolerance is not a defect. A missable tolerance is the only thing that makes the figure a promise rather than a description.

Here are the seven. Every figure is the board's own decision at an invented bank rather than a requirement or an industry norm.

ServiceWhat somebody outside actually doesImpact tolerance
S1Sends money out, or receives it2 hours
S2Walks into a branch and takes cash out4 hours
S3Opens the phone app or the website4 hours
S4Gets an approved loan actually paid out1 working day
S5Opens a new deposit account2 working days
S6Gets a trade document issued1 working day
S7Has a treasury deal settled on time2 hours

The third column is doing more work than it looks like it is doing, so it repays slow reading. Four of the seven are written in hours. Three are written in working days. Nobody sat down and ranked these seven services from most to least important, and yet the ranking is right there in the unit. The unit each tolerance is written in is the sharpest thing in the whole subject.

ONE LINE FROM THE BANK'S OWN TOLERANCE STATEMENT, AND IT HAS EXACTLY FOUR PARTS Payments and remittances shall not be unavailable for longer than 2 hours as agreed by the board, committee G1 THE SERVICE named the way a person outside would name it, never the machinery underneath THE LONGEST TIME the outside world can be without it, chosen so that it can actually be missed THE UNIT the measure the time is in, and the part that leaves S4, S5 and S6 uncomparable WHO AGREED IT a named body with a date against it, so somebody can be asked why it is that A line missing any one of the four cannot be tested against anything. Drop the unit and the sentence still reads perfectly and measures nothing.
An impact tolerance statement is one line with four fixed parts, being the service named as a customer would name it, the longest acceptable time without it, the unit that time is measured in, and the body that agreed it, and the third part is exactly what stops S4, S5 and S6 being comparable to anything.
Try it out

Who sets an impact tolerance?

Building a Client Risk Profile teaches you to turn a client conversation into a documented risk profile, and to separate capacity from tolerance.

How is a tolerance different from a recovery time objective?

Almost everybody trips here, so it is worth going slowly. A recovery time objectiveHow quickly the recovery arrangement is built to bring a service back. is how quickly the recovery arrangement is built to bring a service back. The objective is a design figure. The tolerance is what the institution has promised the outside world, and the objective is what the institution has built to keep that promise, and they are set by different people for different reasons.

Back to the caterer. Dinner is promised at nine. Nine o'clock is the tolerance, and it is what the guests were told. The kitchen plans to have everything on the counter by half past eight. Half past eight is the objective, and no guest has ever been told it. The half hour between the two is not slack anybody is wasting; it is the whole reason a slow burner does not turn into four hundred people waiting.

Which means the two figures fail in completely different ways. Miss the objective and the design has underperformed. The underperformance is worth investigating and it is not yet a broken promise. Miss the tolerance and a promise made by the board has been broken, whatever the design was doing. Quoting the objective outward is an expensive habit for exactly that reason. The quoted figure converts an internal target into a commitment, and the first time the target slips by ten minutes something has been broken that was never actually promised.

TWO OBJECTS THAT LOOK LIKE THE SAME NUMBER AND ARE NOT IMPACT TOLERANCE the promise RECOVERY TIME OBJECTIVE the design WHO SETS IT The board, committee G1, and nobody else The people who build and run the recovery arrangement WHAT IT DESCRIBES How long the outside world can be without the service How quickly the arrangement is built to bring the service back WHICH WAY IT POINTS Outward. It is said to somebody Inward. It is a target for a team S3 INTERNET AND MOBILE BANKING 4 hours 2 hours WHAT A MISS MEANS A promise made by the board has been broken A design underperformed, and the promise may well still be intact The objective is deliberately set inside the tolerance. The distance between the two is the entire safety margin, and it is measured next.
An impact tolerance is the board's promise about how long the outside world can be without the service, a recovery time objective is what the arrangement is built to deliver, and the second is placed inside the first on purpose so that missing the design is not the same event as breaking the promise.
Try it out

S3 internet and mobile banking carries a tolerance of 4 hours and a recovery time objective of 2 hours. Which of those two numbers should be quoted to a customer?

And where does the recovery point objective fit, if it is not about time?

There is a third figure on every row of the bank's service record, and it is the one readers skip. The recovery point objectiveHow much recorded work the institution is willing to lose, measured as data rather than as time. is not about how long the service is away. The recovery point objective is about how much recorded work is gone for good when the service comes back, and it is measured in what was lost rather than in how long the loss took.

Picture the caterer's order book again. The kitchen catches fire, the fire is out in twenty minutes, and dinner is only half an hour late. Excellent recovery time. But the order book burned, and the last two hours of table changes went with it. The service came back fast and the work did not come back at all. The lost half hour and the lost order book are two separate failures, and only one of them is a clock.

Here are all three figures together for the seven services. Every one is again the board's own decision at an invented bank.

ServiceImpact toleranceRecovery time objectiveRecovery point objective
S1 payments and remittances2 hours1 hourZero
S2 branch counter and cash4 hours3 hours15 minutes
S3 internet and mobile banking4 hours2 hours5 minutes
S4 loan disbursal1 working day8 hours1 hour
S5 deposit account opening2 working days12 hours4 hours
S6 trade finance issuance1 working day8 hours1 hour
S7 treasury settlement2 hours1 hourZero

Look at the last column and at S1 and S7 in particular. A recovery point objective of zero is not simply a very small number at the bottom of a range that runs up to S5 at four hours. Zero is a different kind of arrangement altogether. The record has to be written in a second place at the moment it is made rather than copied across afterwards. One minute of gap breaks it just as completely as four hours does. The figure below draws the difference.

THE SAME FAILURE, THE SAME MOMENT, AND TWO COMPLETELY DIFFERENT LOSSES S5 DEPOSIT ACCOUNT OPENING, RECOVERY POINT OBJECTIVE 4 HOURS copy copy next copy every account opened since the last copy is gone THE FAILURE S1 PAYMENTS AND S7 TREASURY SETTLEMENT, RECOVERY POINT OBJECTIVE ZERO written in two places at the moment it is made, so there is no gap to lose Nothing lost at all Both services could be back on their feet in the same hour and still differ completely in what came back with them.
A recovery point objective of zero is a different kind of arrangement rather than a small number, because the whole interval between one usable copy and the next is what a failure takes away, so S5 at four hours can lose an afternoon of account openings while S1 and S7 lose nothing at the same instant.

How much room is there between the promise and the design?

The tolerance less the objective is the safety margin. The margin is the amount by which the recovery arrangement can underperform before anybody outside has been let down, and it is worth computing service by service. Every service has a different margin, and one of the four is running on half the room the others have.

S1 has 60 minutes between its 2 hour tolerance and its 1 hour objective, or 50.0 per cent of the tolerance. S7 is identical. S3 has 120 minutes against a 4 hour tolerance, again 50.0 per cent. S2 has 60 minutes against a 4 hour tolerance, only 25.0 per cent, and that is the thinnest margin of the four. Each figure was set on its own row, so nothing in the board's papers ever compared them.

THE DISTANCE BETWEEN THE PROMISE AND THE DESIGN, IN MINUTES 0 60 120 180 240 minutes S1 PAYMENTS 60 min, 50.0% S2 BRANCH CASH 60 min, 25.0% S3 APP AND WEB 120 min, 50.0% S7 TREASURY 60 min, 50.0% S4, S5 AND S6 no bar can be drawn at all not computable Green: the recovery time objective. Pale grey outline: the impact tolerance. Pink: the margin between them. For S4, S5 and S6 the tolerance is written in working days and the objective in hours, and nothing in the bank's record converts one into the other.
The margin between what the board promised and what the arrangement was built for is the whole safety cushion, it is different for every service, and for three of the seven it cannot be computed at all because the tolerance is written in working days while the objective is written in hours.

The last row is not a drawing problem, it is a governance problem. S4 carries 1 working day against 8 hours. Until somebody writes down how many hours a working day is, those are two numbers that cannot be subtracted. The missing definition looks like a drafting detail in a document nobody reads closely, and it is the reason three of seven services cannot be measured against their own promise. The bank's record never defines a working day in hours, so there is no conversion to take from it.

Try it out

Why can the margin between tolerance and objective not be worked out for S4, S5 and S6?

What does a service actually lean on, and why does that list go stale?

No service can be promised back within two hours unless somebody knows what has to be working for it to be back at all. The written list is the service mapThe written record of everything one service leans on, including the parts run by somebody else., and it is the expensive, boring, unglamorous part of this whole subject. Mapping is where resilience is actually done, and it is also the part that quietly rots.

A useful map goes several layers deep. S1 payments and remittances leans on the core banking system, on a payment gateway run by an outside supplier, on people in a contact centre, and on the network joining them. Each of those leans on something else again. Somewhere down the chain there is a supplier's own supplier that the bank has never had a conversation with, and beyond that there is usually a box with a question mark in it.

Here is the household version. One salary reaches one account. The account pays the rent, the school fee and the electricity. The account has never failed, so nobody has ever written down that all three lean on it. The morning it does, the depth of the chain becomes clear, and it becomes clear in the worst possible order.

WHAT ONE SERVICE LEANS ON, AND HOW FAR THE CHAIN RUNS S1 Payments and remittances The core banking system run inside the bank The payment gateway run by an outside supplier The people who fix it and the network joining them A shared record estate and the building holding it The supplier's own supplier changed after the map was written A settlement arrangement with another institution ? never mapped The red link is the one that moved. Nothing failed, nobody was told, and the map now describes an arrangement that is no longer there.
A service map runs several layers deep and includes the parts run by somebody else, so a supplier who quietly changes their own supplier can leave the written map describing an arrangement that no longer exists while every control still reports as working.

Where does the tolerance clock start, and who is spending it?

Here is the detail that catches out almost every first attempt at this. The clock the board set does not start when the recovery team is called; it starts when the outside world loses the service. It started when the first customer could not pay.

Detection is therefore spending the promise. So is deciding what has happened. If nobody notices for twenty minutes and it takes another half hour to work out which thing broke, then fifty minutes of a 120 minute promise are gone before any recovery arrangement has been switched on at all. The recovery arrangement could then perform perfectly against its own objective and the promise would still break.

Seen that way, the shortest route to a better resilience position is often not a faster recovery route at all, but noticing sooner. Noticing is an unglamorous investment that buys minutes at the front of the clock, and those minutes are worth exactly as much as minutes at the back.

THE PROMISE IS BEING SPENT LONG BEFORE ANYBODY STARTS RECOVERING ANYTHING nobody has noticed yet working out what has actually happened the recovery route is running THE CLOCK STARTS HERE the first customer cannot pay S1 TOLERANCE, 120 MINUTES promise broken here service back The three lengths above are illustrative and are not figures from the bank's record. The record locks how long incident I3 ran in total and says nothing about how that total was divided, so no split is asserted here as a fact about that event. Minutes saved at the front of the clock are worth exactly as much as minutes saved at the back.
The tolerance clock starts when the outside world loses the service and keeps running through noticing and deciding, so a large share of the promise can be spent before any recovery arrangement is switched on and a perfectly performing recovery route can still arrive after the promise has broken.

So how many promises does an outage of a given length actually break?

Everything needed to answer that is now in place, and the answer behaves in a way most people do not expect. Because each service carries its own tolerance, the number of promises broken does not climb steadily as an outage runs longer. The count sits perfectly still and then jumps, and the distances between the jumps are wildly uneven. Predict first, then move the control.

Try it out

An outage runs 4 hours and 20 minutes. How many of the seven services break their tolerance?

Play with it

The outage clock, run against all seven promises at once

One control: how long a single outage runs, from nothing up to 18 hours. One consequence: how many of the seven important business services are outside the impact tolerance the board set for them. The default is 4 hours and 20 minutes, the length of incident I3. At that setting 4 of the 7 services are outside their tolerance, being S1, S2, S3 and S7, and S4, S5 and S6 are inside. The step at that setting runs from just over 4 hours to 8 hours, so anything from 4 hours and 1 minute to a full 8 hours produces exactly the same count of four.

EACH SERVICE AGAINST ITS OWN PROMISE. THE BLACK MARK IS THAT SERVICE'S TOLERANCE. 0 3 h 6 h 9 h 12 h 15 h 18 h S1 PAYMENTS tolerance 2 hours OVER BY 2 h 20 m S2 BRANCH CASH tolerance 4 hours OVER BY 20 m S3 APP AND WEB tolerance 4 hours OVER BY 20 m S4 LOAN DISBURSAL 1 working day, read here as 8 h INSIDE BY 3 h 40 m S5 ACCOUNT OPENING 2 working days, read here as 16 h INSIDE BY 11 h 40 m S6 TRADE FINANCE 1 working day, read here as 8 h INSIDE BY 3 h 40 m S7 TREASURY tolerance 2 hours OVER BY 2 h 20 m Dashed marks: the three tolerances the board wrote in working days. A working day is read here as 8 hours; the board never wrote a conversion down. HOW THE COUNT MOVES AS THE OUTAGE RUNS LONGER 0 2 4 6 7 4 h 20 m, 4 broken 2 hours wide 2 hours 4 hours 8 hours wide and above The width of each flat stretch is the distance between two tolerance figures nobody ever compared with each other.
No outage4 h 20 m18 hours
Outage length
4 h 20 m
Promises broken
4 of 7
Still inside
3 of 7

At an outage of 4 hours and 20 minutes, 4 of the seven services are outside the promise the board made, being S1, S2, S3 and S7.

Educational illustration. Every tolerance here is the board's own invented figure at an invented bank and none of them is a requirement or an industry norm. The bars show tolerances only and not recovery objectives. The tolerances of S4, S5 and S6 are written in working days. The board never wrote down how many hours a working day is, so the two units cannot be subtracted. A working day is therefore read here as 8 hours, a stated assumption of this guide and never the bank's, and those three are marked with dashed lines. The count is the number of promises broken and it is not a measure of what an outage costs.

Two things in that control are worth staring at. First, the count is flat and then jumps: an outage of exactly 2 hours breaks nothing at all, one of 2 hours and 1 minute breaks two promises, and every length from just over 4 hours to a full 8 hours breaks exactly the same four. An extra hour of outage can cost nothing and an extra minute can cost two promises, depending only on where on the clock it falls.

Second, look at the widths along the bottom. The first two flat stretches are two hours wide, the third is four, the fourth is eight. Nobody designed that shape. The shape fell out of seven separate tolerance decisions, each taken on its own row, none of them ever laid beside the others.

Try it out

Four of the seven tolerances broke in incident I3. What do those four have in common?

What did incident I3 actually break?

In month 3 the core banking system at Vindhya Commercial Bank Limited was unavailable for 4 hours and 20 minutes on a working day. In the bank's own record of operational loss events that is incident I3, and 4 hours and 20 minutes is 260 minutes. Now measure that one number against each tolerance in turn. Nobody at the bank performed that exercise at the time.

260 is more than 120, so S1 payments and remittances is broken, and so is S7 treasury settlement. 260 is more than 240, so S2 branch counter and cash is broken, and so is S3 internet and mobile banking. 260 minutes is inside a working day on any reading at all, so S4, S5 and S6 are untouched. Four of the seven promises were broken by one outage, being 57.1 per cent of the named services, and the bank's loss record does not contain that sentence anywhere.

INCIDENT I3 MEASURED AGAINST EVERY ONE OF THE SEVEN PROMISES The core banking system was unavailable for 260 minutes. The bar is the tolerance; the pink is how far past it the outage ran. 0 60 120 180 240 minutes S1 PAYMENTS tolerance 120 min 140 min over S2 BRANCH CASH tolerance 240 min 20 min over S3 APP AND WEB tolerance 240 min 20 min over S7 TREASURY tolerance 120 min 140 min over INCIDENT I3, 260 MINUTES S4, S5 AND S6 working days 260 minutes is inside one working day on any reading anybody could take, so all three stayed inside and no bar needs drawing at all Four broken out of seven, being 57.1 per cent, and the same single outage missed two of them by 20 minutes and two by 140. Saying only that four tolerances were breached flattens a range from 20 minutes to 140 minutes into one word.
One outage of 260 minutes broke four of the seven promises and stayed inside three, and the four were not broken equally, with S1 and S7 running at 2.17 times their tolerance while S2 and S3 ran at 1.083 times theirs.
ServiceToleranceTimes the toleranceMinutes over
S1 payments and remittances120 minutes2.17140
S7 treasury settlement120 minutes2.17140
S2 branch counter and cash240 minutes1.08320
S3 internet and mobile banking240 minutes1.08320
S4, S5 and S6working daysinsidenone

Why are those four the ones that broke, and not some other four?

Go back and look at which services broke. S1, S2, S3 and S7. Now go back to the tolerance table and look at how each of the seven is written. S1, S2, S3 and S7 are exactly the four whose tolerance the board wrote in hours. S4, S5 and S6 are the three it wrote in working days, and all three survived.

The match is not luck and it is not a coincidence: a service whose tolerance is written in hours has already been judged to matter inside a single day, and an outage of 4 hours and 20 minutes is a single-day event. Nobody ever ranked these seven services. The ranking was encoded in the unit at the moment each tolerance was drafted, months before any outage happened, and it turned out to predict exactly which promises a one-day failure would break.

SORT THE SEVEN BY THE UNIT THE BOARD WROTE, AND INCIDENT I3 SORTS ITSELF TOLERANCE WRITTEN IN HOURS TOLERANCE WRITTEN IN WORKING DAYS S1 payments and remittances, 2 hours BROKEN S7 treasury settlement, 2 hours BROKEN S2 branch counter and cash, 4 hours BROKEN S3 internet and mobile banking, 4 hours BROKEN S4 loan disbursal, 1 working day INSIDE S6 trade finance issuance, 1 working day INSIDE S5 deposit account opening, 2 working days INSIDE Nobody ranked these seven services against each other. The unit did it for them. Four in hours, four broken. Three in working days, three intact. The split is perfect and nobody arranged it.
The four services incident I3 broke are exactly the four whose tolerance the board wrote in hours, and the three that survived are exactly the three written in working days, so the unit each tolerance was drafted in already encoded how much the service mattered before any measurement was taken.

How is a service tested for whether it can stay inside its tolerance?

The service is broken on purpose and the clock is watched. In month 12 Vindhya Commercial Bank Limited ran a continuity test across all seven important business services. The test recorded how long each one took to come back. Four met their recovery time objective and three missed it. The result is useful and it is also a smaller claim than it looks, for a reason worth being honest about.

THE MONTH 12 CONTINUITY TEST, EACH SERVICE AGAINST ITS OWN RECOVERY OBJECTIVE Bar length is how far the time achieved sat from the objective, as a share of that objective. THE OBJECTIVE S1 PAYMENTS 13.3% MET by 8 min S2 BRANCH CASH 22.2% MISSED by 40 min S3 APP AND WEB 58.3% MISSED by 70 min S4 LOAN DISBURSAL 25.0% MET by 120 min S5 ACCOUNT OPENING 25.0% MET by 180 min S6 TRADE FINANCE 75.0% MISSED by 360 min S7 TREASURY 20.0% MET by 12 min FASTER THAN THE OBJECTIVE SLOWER THAN THE OBJECTIVE
Four of the seven services met their recovery time objective in the month 12 test and three missed it, with S6 trade finance issuance the worst miss at 75.0 per cent over and S1 payments the thinnest pass at 8 minutes inside.

Now the honest limitation, and it belongs beside every result of this kind. The month 12 test ran on a planned date, with the recovery team on standby, and the failure chosen in advance, so it tested the arrangement and not the surprise. The minutes of not noticing and the minutes of working out what had happened, the same minutes that were spending the promise in the earlier drawing, were removed before the clock started. A real failure spends those minutes. State the limitation beside the result, or the result quietly overstates itself.

Notice also what the test measured and what it did not. The test recorded recovery times. The test recorded nothing at all about how much work was lost, so the recovery point objective column was never tested. Seven services carry a figure in that column and the month 12 test produced no evidence about any of them.

Try it out

Read the month 12 result with its own limitation printed beside it. What did that test actually test?

The error that gets made, and what it costs

Reading the net loss booked on incident I3 as the size of the event. Committee G6, the operational risk management committee, receives the incident record monthly, and incident I3 arrives there as one line: month 3, category 6, business disruption and system failures, gross Rs 3.2 crore, recovery zero, net Rs 3.2 crore, being 7.3 per cent of the year's Rs 43.8 crore of net operational loss. Every figure on that line is correct and correctly computed.

Not one of those figures says that four of seven board promises were broken. A loss figure measures what the institution paid. A tolerance breach measures what the outside world went without. The two are different quantities about different subjects, and a report carrying only the first has hidden the second without anybody deciding to hide anything.

The cost is precise. Ranked by gross loss, incident I3 is the seventh largest of the thirteen events in the year and easy to scroll past; ranked by net loss it is joint fourth, tied with incident I9 which also booked Rs 3.2 crore net, so the reader must always name the incident rather than the figure. On a tolerance measurement it is the one event in the year for which that measurement was made at all, and it found four broken promises. The second report was never produced, so nobody at the bank ever read that sentence.

TWO CORRECT REPORTS ON ONE EVENT, AND ONLY ONE OF THEM WAS WRITTEN THE LOSS REPORT, PRODUCED MONTHLY Incident I3, month 3 Category 6, business disruption and system failures Gross loss Rs 3.2 crore Recovery nil Net loss Rs 3.2 crore 7.3 per cent of the year's Rs 43.8 crore Seventh largest of thirteen ranked by gross loss Every figure here is correct. THE TOLERANCE REPORT, NEVER PRODUCED Incident I3, month 3, 260 minutes without the service S1 payments, promise broken by 140 minutes S7 treasury, promise broken by 140 minutes S2 branch cash, promise broken by 20 minutes S3 app and web, promise broken by 20 minutes S4, S5 and S6 stayed inside Four of seven board promises broken Every figure here would also be correct. Nobody hid anything and no figure was wrong. The second measurement was simply never made, for incident I3 or for any other event in the year. A report that answers only the first question has answered the wrong one without ever being inaccurate.
The same event produces two entirely different and equally correct reports, one measuring what the bank paid in rupees and the other measuring what the outside world went without in minutes, and the bank produced only the first.
Try it out

The month 3 loss report shows incident I3 at Rs 3.2 crore net. What is missing from it?

Risk Management Program Bootcamp — Fin Maverick

Where do the expectations on an Indian bank actually come from?

A service, a tolerance, a map and a test are the same four objects wherever the institution sits, so everything above holds in any jurisdiction. Who tells an institution it has to have them differs, and that question has two answers rather than one.

The Basel Committee, hosted at the Bank for International Settlements at bis.org, publishes the international standard on operational resilience, and it is the origin of the vocabulary used here. But an Indian bank is not held to a standard published in Basel. The Reserve Bank of India at rbi.org.in sets what an Indian bank must actually do about resilience, business continuity, outsourcing and reporting an incident, and naming only the global standard tells an Indian reader nothing about what applies to them.

NAMING THE STANDARD AND NAMING WHAT BINDS ARE TWO DIFFERENT ACTS THE BASEL COMMITTEE hosted at the Bank for International Settlements, at bis.org publishes the standard where the vocabulary used here begins is implemented by THE RESERVE BANK OF INDIA at rbi.org.in sets what actually binds an Indian bank on resilience, continuity, outsourcing and reporting an incident Name the left box alone and an Indian reader still does not know what applies to them. Name the right box alone and the idea has no origin. Nothing here states a requirement, a tolerance, a recovery objective or an effective date from either body. Confirm the current position directly with the issuing body.
The Basel Committee at the Bank for International Settlements publishes the standard on operational resilience while the Reserve Bank of India sets what an Indian bank must actually do about it, so naming only the first tells an Indian reader nothing about what applies to them.
India

What the Indian rules require

The Reserve Bank of India at rbi.org.in is the source of what binds an Indian bank on operational resilience, business continuity, outsourcing arrangements and the reporting of an incident. The Bank for International Settlements at bis.org is where the international standard the Indian requirement implements is published. Where the entity is a market intermediary rather than a bank, the Securities and Exchange Board of India at sebi.gov.in applies instead.

Every figure in the worked case belongs to an invented bank and is its board's own decision, so none of them states a requirement, a reporting window or an effective date. The current position on all of those sits with the issuing body.

Try it out

Which body sets what an Indian bank must actually do about operational resilience?

In practice

Who reads this, and what they do with it

Girish Talwalkar, group treasurer at the invented Nirjhar Industries Limited, has to move a payroll run through a bank on a fixed date. The bank's loss record tells him nothing he can use. The number he wants is the impact tolerance on payments and remittances. The tolerance is the bank's own statement of the longest it thinks the outside world can be without the service, and he is the outside world. If it is two hours, his contingency needs to cover two hours.

A supervisor or an internal auditor reads the same set differently. The tolerance figures are the board's to set; the auditor wants to know whether the board can show that it set them, tested them, and knows how often they were missed. A count of tolerance breaches in a year is a governance answer rather than a technical one. Nobody at this invented bank produced the count, so the count that reached anybody was zero.

And the household version, the same skill on a smaller balance. Where one salary account pays the rent, the school fee and the electricity, an important business service has been named without anybody meaning to. Deciding how many days the household could manage if that account froze is setting an impact tolerance. Keeping a second account with one month of expenses in it is a recovery arrangement. Nobody uses those words at a kitchen table and the structure is identical.

Breaking Into Quants Bootcamp — Fin Maverick

What can operational resilience not promise?

A great deal, and being clear about it is part of doing it honestly. A tolerance is a statement of intent about an outcome nobody fully controls. Setting a two hour tolerance does not make a two hour recovery happen; it states what would count as a broken promise if it does not, and it directs money at making the promise keepable. Those are worth a lot and they are not the same as an outcome.

Tolerances are set against scenarios chosen to be severe but plausibleA supervisory phrase for a scenario chosen to be hard without being fanciful., and that phrase carries its own limit. A scenario has to be imagined before it can be planned for, and the failure that actually arrives is regularly one that was never on the list. Frank Knight drew this line in 1921 between a risk that can be measured and an uncertainty that cannot, and the second kind does not stop existing because the first kind has been carefully tabulated.

Two more limits sit inside this specific case. The map goes stale, so an arrangement that was sound when it was written may not be sound when it is needed, and nothing failed in between to tell anybody. And the test is planned, so the recorded result describes a rehearsal rather than an emergency. Neither of those makes the work pointless. Both mean the claim the evidence will actually carry is narrower than the one usually made from it.

How an operational loss event is defined, measured, categorised or booked is covered separately; the bank's own loss record is used here as a record, with no figure in it re-derived. How a failure is traced back to its cause, how a near miss is logged and how a supplier is assessed as a risk in its own right are covered separately. How a control is designed, tested, found deficient and put right is covered separately too. Keeping a service running by another route, restoring the technology behind it, declaring and running a crisis, and handling one incident end to end are each covered separately. The explanation of any financial instrument is covered elsewhere.

Sources

SourceDocumentSite
Reserve Bank of IndiaExpectations on a regulated bank covering operational resilience, business continuity, outsourcing and the reporting of an incidentrbi.org.in
Bank for International SettlementsThe Basel Committee's international standard on operational resilience, and the operational risk event categories named herebis.org
Securities and Exchange Board of IndiaExpectations where the entity is a market intermediary rather than a banksebi.gov.in
Indian Banks AssociationMaterial on Indian banking operational conventioniba.org.in
Frank KnightRisk, Uncertainty and Profit, 1921, the separation of a measurable risk from an unmeasurable uncertaintyHoughton Mifflin, 1921

Vindhya Commercial Bank Limited, Nirjhar Industries Limited and Girish Talwalkar are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.