Operational Resilience: Staying Available Through Disruption
Operational resilience is the ability to keep a service available to the people outside the institution while something has gone wrong inside it. Resilience assumes the failure happens. The institution names the services that matter to somebody outside, sets a board level tolerance for how long the outside world can be without each one, maps what the service leans on, and then tests whether the service stays inside that tolerance.
One change of unit carries the whole idea, and it is the thing to carry away above everything else here. Operational risk asks what could go wrong inside the institution and what it would cost. Operational resilience asks what somebody outside experiences while it is going wrong, and for how long. The moment the unit becomes a service rather than a system, a department or a loss figure, every other rule in this guide follows without further argument. One invented bank, Vindhya Commercial Bank Limited, carries every worked figure below.
How is this different from just trying to stop things going wrong?
Most of what an institution does about failure is an attempt to prevent it. The instruction gets checked twice. A second person approves it. The change is tested before it goes live. All of that is worth doing, and none of it is resilience.
Operational resilienceThe ability to keep a service available to people outside the institution while something inside it has failed. starts by conceding the argument. Resilience assumes that prevention will not hold every time, that something will break, and that the interesting question begins at that moment rather than ending there. The premise is not that the failure can be avoided but that the outside world can be kept whole while it happens.
Think about a wedding caterer. Preventing failure means buying good gas cylinders and checking the burners twice. Resilience means something else entirely: the cylinder runs out anyway, and the question is whether four hundred people sit down to dinner at the hour they were promised, or whether they sit down at eleven. Two different questions, two different sets of preparations, and one of them is about the guests rather than about the kitchen. A bank asking a resilience question is asking about the guests.
Measuring from outside the institution is also why a disruptionAny event that stops a service reaching the person who uses it, whatever caused it. is defined without any reference to its cause. A power failure, a supplier who stops answering, a change that was applied badly, a flood at a branch: to the person who could not withdraw cash, these are all the same event. Resilience is measured from where that person stands, so the cause matters enormously for fixing it and not at all for measuring it.
What counts as an important business service, and who decides?
An important business serviceA service somebody outside the institution relies on, named in that person's words rather than in the institution's. is a thing somebody outside the institution actually uses, described the way that person would describe it. Not a department. Not a building. Not a piece of machinery. The test is whether a customer would recognise the name, and if they would not, it is not a service.
Vindhya Commercial Bank Limited has named seven of them, numbered S1 to S7. Read the names and notice what they are: sending money, taking cash out at a counter, using the phone app, getting a loan paid out, opening an account, having a trade document issued, and settling a treasury deal. Every one is a verb somebody outside the bank performs. None of them is the name of a system.
The distinction between a service and a system is not pedantry, and the drawing below is why. A service and a system are not the same shape. One system can sit underneath several services at once, and one service can lean on several systems at once. Making the system the unit protects the machinery and loses sight of the promise. Worse, nobody outside has ever heard of the machinery, so there is no way to say what the outside world lost.
Notice one more thing in that drawing. The payment gateway is run by an outside supplier and it still sits underneath a service the bank has promised. Handing the work to somebody else moves the machinery out of the building and leaves the promise exactly where it was. Vindhya Commercial Bank Limited found this out the hard way in month 9, when a supplier hosted payment gateway failed for nine hours and 48,000 transactions failed with it. The gateway failure is incident I9 in the bank's loss record.
A bank lists its important business services as core banking, the card switch and the payments hub. What has gone wrong?
Who sets an impact tolerance, and what makes it a promise?
Once the services are named, somebody has to say how long the outside world can be without each one. The agreed length is the impact toleranceThe longest the board is willing for the outside world to be without a named service., and at Vindhya Commercial Bank Limited it is set by the board, committee G1, and by nobody else. An impact tolerance is a business decision about what the outside world can bear, not a description of what the equipment can currently manage.
The difference between those two sentences is the whole reason the board sets it. If the people who run the machinery set the figure, the figure becomes a statement of what the machinery already does. The figure was written to be met, so it will always be met. A tolerance set by the board can be missed. A missable tolerance is not a defect. A missable tolerance is the only thing that makes the figure a promise rather than a description.
Here are the seven. Every figure is the board's own decision at an invented bank rather than a requirement or an industry norm.
| Service | What somebody outside actually does | Impact tolerance |
|---|---|---|
| S1 | Sends money out, or receives it | 2 hours |
| S2 | Walks into a branch and takes cash out | 4 hours |
| S3 | Opens the phone app or the website | 4 hours |
| S4 | Gets an approved loan actually paid out | 1 working day |
| S5 | Opens a new deposit account | 2 working days |
| S6 | Gets a trade document issued | 1 working day |
| S7 | Has a treasury deal settled on time | 2 hours |
The third column is doing more work than it looks like it is doing, so it repays slow reading. Four of the seven are written in hours. Three are written in working days. Nobody sat down and ranked these seven services from most to least important, and yet the ranking is right there in the unit. The unit each tolerance is written in is the sharpest thing in the whole subject.
Who sets an impact tolerance?
How is a tolerance different from a recovery time objective?
Almost everybody trips here, so it is worth going slowly. A recovery time objectiveHow quickly the recovery arrangement is built to bring a service back. is how quickly the recovery arrangement is built to bring a service back. The objective is a design figure. The tolerance is what the institution has promised the outside world, and the objective is what the institution has built to keep that promise, and they are set by different people for different reasons.
Back to the caterer. Dinner is promised at nine. Nine o'clock is the tolerance, and it is what the guests were told. The kitchen plans to have everything on the counter by half past eight. Half past eight is the objective, and no guest has ever been told it. The half hour between the two is not slack anybody is wasting; it is the whole reason a slow burner does not turn into four hundred people waiting.
Which means the two figures fail in completely different ways. Miss the objective and the design has underperformed. The underperformance is worth investigating and it is not yet a broken promise. Miss the tolerance and a promise made by the board has been broken, whatever the design was doing. Quoting the objective outward is an expensive habit for exactly that reason. The quoted figure converts an internal target into a commitment, and the first time the target slips by ten minutes something has been broken that was never actually promised.
S3 internet and mobile banking carries a tolerance of 4 hours and a recovery time objective of 2 hours. Which of those two numbers should be quoted to a customer?
And where does the recovery point objective fit, if it is not about time?
There is a third figure on every row of the bank's service record, and it is the one readers skip. The recovery point objectiveHow much recorded work the institution is willing to lose, measured as data rather than as time. is not about how long the service is away. The recovery point objective is about how much recorded work is gone for good when the service comes back, and it is measured in what was lost rather than in how long the loss took.
Picture the caterer's order book again. The kitchen catches fire, the fire is out in twenty minutes, and dinner is only half an hour late. Excellent recovery time. But the order book burned, and the last two hours of table changes went with it. The service came back fast and the work did not come back at all. The lost half hour and the lost order book are two separate failures, and only one of them is a clock.
Here are all three figures together for the seven services. Every one is again the board's own decision at an invented bank.
| Service | Impact tolerance | Recovery time objective | Recovery point objective |
|---|---|---|---|
| S1 payments and remittances | 2 hours | 1 hour | Zero |
| S2 branch counter and cash | 4 hours | 3 hours | 15 minutes |
| S3 internet and mobile banking | 4 hours | 2 hours | 5 minutes |
| S4 loan disbursal | 1 working day | 8 hours | 1 hour |
| S5 deposit account opening | 2 working days | 12 hours | 4 hours |
| S6 trade finance issuance | 1 working day | 8 hours | 1 hour |
| S7 treasury settlement | 2 hours | 1 hour | Zero |
Look at the last column and at S1 and S7 in particular. A recovery point objective of zero is not simply a very small number at the bottom of a range that runs up to S5 at four hours. Zero is a different kind of arrangement altogether. The record has to be written in a second place at the moment it is made rather than copied across afterwards. One minute of gap breaks it just as completely as four hours does. The figure below draws the difference.
How much room is there between the promise and the design?
The tolerance less the objective is the safety margin. The margin is the amount by which the recovery arrangement can underperform before anybody outside has been let down, and it is worth computing service by service. Every service has a different margin, and one of the four is running on half the room the others have.
S1 has 60 minutes between its 2 hour tolerance and its 1 hour objective, or 50.0 per cent of the tolerance. S7 is identical. S3 has 120 minutes against a 4 hour tolerance, again 50.0 per cent. S2 has 60 minutes against a 4 hour tolerance, only 25.0 per cent, and that is the thinnest margin of the four. Each figure was set on its own row, so nothing in the board's papers ever compared them.
The last row is not a drawing problem, it is a governance problem. S4 carries 1 working day against 8 hours. Until somebody writes down how many hours a working day is, those are two numbers that cannot be subtracted. The missing definition looks like a drafting detail in a document nobody reads closely, and it is the reason three of seven services cannot be measured against their own promise. The bank's record never defines a working day in hours, so there is no conversion to take from it.
Why can the margin between tolerance and objective not be worked out for S4, S5 and S6?
What does a service actually lean on, and why does that list go stale?
No service can be promised back within two hours unless somebody knows what has to be working for it to be back at all. The written list is the service mapThe written record of everything one service leans on, including the parts run by somebody else., and it is the expensive, boring, unglamorous part of this whole subject. Mapping is where resilience is actually done, and it is also the part that quietly rots.
A useful map goes several layers deep. S1 payments and remittances leans on the core banking system, on a payment gateway run by an outside supplier, on people in a contact centre, and on the network joining them. Each of those leans on something else again. Somewhere down the chain there is a supplier's own supplier that the bank has never had a conversation with, and beyond that there is usually a box with a question mark in it.
Here is the household version. One salary reaches one account. The account pays the rent, the school fee and the electricity. The account has never failed, so nobody has ever written down that all three lean on it. The morning it does, the depth of the chain becomes clear, and it becomes clear in the worst possible order.
Where does the tolerance clock start, and who is spending it?
Here is the detail that catches out almost every first attempt at this. The clock the board set does not start when the recovery team is called; it starts when the outside world loses the service. It started when the first customer could not pay.
Detection is therefore spending the promise. So is deciding what has happened. If nobody notices for twenty minutes and it takes another half hour to work out which thing broke, then fifty minutes of a 120 minute promise are gone before any recovery arrangement has been switched on at all. The recovery arrangement could then perform perfectly against its own objective and the promise would still break.
Seen that way, the shortest route to a better resilience position is often not a faster recovery route at all, but noticing sooner. Noticing is an unglamorous investment that buys minutes at the front of the clock, and those minutes are worth exactly as much as minutes at the back.
So how many promises does an outage of a given length actually break?
Everything needed to answer that is now in place, and the answer behaves in a way most people do not expect. Because each service carries its own tolerance, the number of promises broken does not climb steadily as an outage runs longer. The count sits perfectly still and then jumps, and the distances between the jumps are wildly uneven. Predict first, then move the control.
An outage runs 4 hours and 20 minutes. How many of the seven services break their tolerance?
The outage clock, run against all seven promises at once
One control: how long a single outage runs, from nothing up to 18 hours. One consequence: how many of the seven important business services are outside the impact tolerance the board set for them. The default is 4 hours and 20 minutes, the length of incident I3. At that setting 4 of the 7 services are outside their tolerance, being S1, S2, S3 and S7, and S4, S5 and S6 are inside. The step at that setting runs from just over 4 hours to 8 hours, so anything from 4 hours and 1 minute to a full 8 hours produces exactly the same count of four.
At an outage of 4 hours and 20 minutes, 4 of the seven services are outside the promise the board made, being S1, S2, S3 and S7.
Two things in that control are worth staring at. First, the count is flat and then jumps: an outage of exactly 2 hours breaks nothing at all, one of 2 hours and 1 minute breaks two promises, and every length from just over 4 hours to a full 8 hours breaks exactly the same four. An extra hour of outage can cost nothing and an extra minute can cost two promises, depending only on where on the clock it falls.
Second, look at the widths along the bottom. The first two flat stretches are two hours wide, the third is four, the fourth is eight. Nobody designed that shape. The shape fell out of seven separate tolerance decisions, each taken on its own row, none of them ever laid beside the others.
Four of the seven tolerances broke in incident I3. What do those four have in common?
What did incident I3 actually break?
In month 3 the core banking system at Vindhya Commercial Bank Limited was unavailable for 4 hours and 20 minutes on a working day. In the bank's own record of operational loss events that is incident I3, and 4 hours and 20 minutes is 260 minutes. Now measure that one number against each tolerance in turn. Nobody at the bank performed that exercise at the time.
260 is more than 120, so S1 payments and remittances is broken, and so is S7 treasury settlement. 260 is more than 240, so S2 branch counter and cash is broken, and so is S3 internet and mobile banking. 260 minutes is inside a working day on any reading at all, so S4, S5 and S6 are untouched. Four of the seven promises were broken by one outage, being 57.1 per cent of the named services, and the bank's loss record does not contain that sentence anywhere.
| Service | Tolerance | Times the tolerance | Minutes over |
|---|---|---|---|
| S1 payments and remittances | 120 minutes | 2.17 | 140 |
| S7 treasury settlement | 120 minutes | 2.17 | 140 |
| S2 branch counter and cash | 240 minutes | 1.083 | 20 |
| S3 internet and mobile banking | 240 minutes | 1.083 | 20 |
| S4, S5 and S6 | working days | inside | none |
Why are those four the ones that broke, and not some other four?
Go back and look at which services broke. S1, S2, S3 and S7. Now go back to the tolerance table and look at how each of the seven is written. S1, S2, S3 and S7 are exactly the four whose tolerance the board wrote in hours. S4, S5 and S6 are the three it wrote in working days, and all three survived.
The match is not luck and it is not a coincidence: a service whose tolerance is written in hours has already been judged to matter inside a single day, and an outage of 4 hours and 20 minutes is a single-day event. Nobody ever ranked these seven services. The ranking was encoded in the unit at the moment each tolerance was drafted, months before any outage happened, and it turned out to predict exactly which promises a one-day failure would break.
How is a service tested for whether it can stay inside its tolerance?
The service is broken on purpose and the clock is watched. In month 12 Vindhya Commercial Bank Limited ran a continuity test across all seven important business services. The test recorded how long each one took to come back. Four met their recovery time objective and three missed it. The result is useful and it is also a smaller claim than it looks, for a reason worth being honest about.
Now the honest limitation, and it belongs beside every result of this kind. The month 12 test ran on a planned date, with the recovery team on standby, and the failure chosen in advance, so it tested the arrangement and not the surprise. The minutes of not noticing and the minutes of working out what had happened, the same minutes that were spending the promise in the earlier drawing, were removed before the clock started. A real failure spends those minutes. State the limitation beside the result, or the result quietly overstates itself.
Notice also what the test measured and what it did not. The test recorded recovery times. The test recorded nothing at all about how much work was lost, so the recovery point objective column was never tested. Seven services carry a figure in that column and the month 12 test produced no evidence about any of them.
Read the month 12 result with its own limitation printed beside it. What did that test actually test?
The error that gets made, and what it costs
Reading the net loss booked on incident I3 as the size of the event. Committee G6, the operational risk management committee, receives the incident record monthly, and incident I3 arrives there as one line: month 3, category 6, business disruption and system failures, gross Rs 3.2 crore, recovery zero, net Rs 3.2 crore, being 7.3 per cent of the year's Rs 43.8 crore of net operational loss. Every figure on that line is correct and correctly computed.
Not one of those figures says that four of seven board promises were broken. A loss figure measures what the institution paid. A tolerance breach measures what the outside world went without. The two are different quantities about different subjects, and a report carrying only the first has hidden the second without anybody deciding to hide anything.
The cost is precise. Ranked by gross loss, incident I3 is the seventh largest of the thirteen events in the year and easy to scroll past; ranked by net loss it is joint fourth, tied with incident I9 which also booked Rs 3.2 crore net, so the reader must always name the incident rather than the figure. On a tolerance measurement it is the one event in the year for which that measurement was made at all, and it found four broken promises. The second report was never produced, so nobody at the bank ever read that sentence.
The month 3 loss report shows incident I3 at Rs 3.2 crore net. What is missing from it?
Where do the expectations on an Indian bank actually come from?
A service, a tolerance, a map and a test are the same four objects wherever the institution sits, so everything above holds in any jurisdiction. Who tells an institution it has to have them differs, and that question has two answers rather than one.
The Basel Committee, hosted at the Bank for International Settlements at bis.org, publishes the international standard on operational resilience, and it is the origin of the vocabulary used here. But an Indian bank is not held to a standard published in Basel. The Reserve Bank of India at rbi.org.in sets what an Indian bank must actually do about resilience, business continuity, outsourcing and reporting an incident, and naming only the global standard tells an Indian reader nothing about what applies to them.
What the Indian rules require
The Reserve Bank of India at rbi.org.in is the source of what binds an Indian bank on operational resilience, business continuity, outsourcing arrangements and the reporting of an incident. The Bank for International Settlements at bis.org is where the international standard the Indian requirement implements is published. Where the entity is a market intermediary rather than a bank, the Securities and Exchange Board of India at sebi.gov.in applies instead.
Every figure in the worked case belongs to an invented bank and is its board's own decision, so none of them states a requirement, a reporting window or an effective date. The current position on all of those sits with the issuing body.
Which body sets what an Indian bank must actually do about operational resilience?
Who reads this, and what they do with it
Girish Talwalkar, group treasurer at the invented Nirjhar Industries Limited, has to move a payroll run through a bank on a fixed date. The bank's loss record tells him nothing he can use. The number he wants is the impact tolerance on payments and remittances. The tolerance is the bank's own statement of the longest it thinks the outside world can be without the service, and he is the outside world. If it is two hours, his contingency needs to cover two hours.
A supervisor or an internal auditor reads the same set differently. The tolerance figures are the board's to set; the auditor wants to know whether the board can show that it set them, tested them, and knows how often they were missed. A count of tolerance breaches in a year is a governance answer rather than a technical one. Nobody at this invented bank produced the count, so the count that reached anybody was zero.
And the household version, the same skill on a smaller balance. Where one salary account pays the rent, the school fee and the electricity, an important business service has been named without anybody meaning to. Deciding how many days the household could manage if that account froze is setting an impact tolerance. Keeping a second account with one month of expenses in it is a recovery arrangement. Nobody uses those words at a kitchen table and the structure is identical.
What can operational resilience not promise?
A great deal, and being clear about it is part of doing it honestly. A tolerance is a statement of intent about an outcome nobody fully controls. Setting a two hour tolerance does not make a two hour recovery happen; it states what would count as a broken promise if it does not, and it directs money at making the promise keepable. Those are worth a lot and they are not the same as an outcome.
Tolerances are set against scenarios chosen to be severe but plausibleA supervisory phrase for a scenario chosen to be hard without being fanciful., and that phrase carries its own limit. A scenario has to be imagined before it can be planned for, and the failure that actually arrives is regularly one that was never on the list. Frank Knight drew this line in 1921 between a risk that can be measured and an uncertainty that cannot, and the second kind does not stop existing because the first kind has been carefully tabulated.
Two more limits sit inside this specific case. The map goes stale, so an arrangement that was sound when it was written may not be sound when it is needed, and nothing failed in between to tell anybody. And the test is planned, so the recorded result describes a rehearsal rather than an emergency. Neither of those makes the work pointless. Both mean the claim the evidence will actually carry is narrower than the one usually made from it.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated bank covering operational resilience, business continuity, outsourcing and the reporting of an incident | rbi.org.in |
| Bank for International Settlements | The Basel Committee's international standard on operational resilience, and the operational risk event categories named here | bis.org |
| Securities and Exchange Board of India | Expectations where the entity is a market intermediary rather than a bank | sebi.gov.in |
| Indian Banks Association | Material on Indian banking operational convention | iba.org.in |
| Frank Knight | Risk, Uncertainty and Profit, 1921, the separation of a measurable risk from an unmeasurable uncertainty | Houghton Mifflin, 1921 |
Vindhya Commercial Bank Limited, Nirjhar Industries Limited and Girish Talwalkar are invented.
Educational material. Not advice on any investment, tax, budget or market position.
