Business Continuity vs Operational Resilience
Business continuity and operational resilience are not two names for one thing. Business continuity is an arrangement the institution makes so a service can be delivered another way, and it succeeds when the arrangement works. Operational resilience is an outcome measured against a promise the board made about the outside world, and it succeeds when the promise holds. Continuity is something the institution invokes; resilience is something the institution is measured on whether anything was invoked or not.
Most people meet these two words in the same paragraph of the same document and reasonably conclude that one is a newer, grander way of saying the other. The conclusion is wrong, and the cost of it is that a bank ends up answering a question nobody asked while the question the board actually asked goes unanswered for a whole year. The two words point at different objects: one is a thing the institution builds, and the other is a thing that happens to somebody outside while what was built either holds or does not.
Every figure worked below belongs to Vindhya Commercial Bank Limited, an invented bank with a balance sheet of Rs 96,000 crore. Its services, its objectives and its tolerances are that bank's own decisions rather than facts about Indian banking. None of the difference between the two ideas is technology. The decisions in question were taken by a board and by management and not by anybody near a machine, so knowing what a data centre is is not required.
Before the two ideas can be laid against each other, each one has to stand on its own. So each is defined first on its own terms, with nothing said about the other, and only then does the comparison start.
What exactly is business continuity, taken on its own?
Business continuityAn arrangement for delivering a service another way while the usual way is unavailable. is an arrangement. The word arrangement is doing real work, so hold on to it. An arrangement is a thing that exists before anything goes wrong, that somebody paid for, that somebody wrote down, and that somebody can switch on. Business continuity is the other route plus the written record of who switches to it and what happens next. How that route is built, staffed and paid for is settled under business continuity and is taken as given here.
A tiffin service that delivers lunch boxes to an office block is a good place to feel this. The usual route is one van and one driver. The arrangement is the cousin with a motorcycle who has the route list, the building pass and a phone number that is answered. He is not delivering anything today. He exists so that on the day the van will not start, somebody rings him and lunch still goes out. Continuity in one sentence is a second route, ready, and a person with the authority to call it.
Three things sit inside that description, and each of them matters for the comparison later.
The first is that a continuity arrangement has a scope. The cousin covers this route and not the other three. A written continuity arrangement at a bank says which activities it covers. By silence, it also says which activities it does not. At Vindhya Commercial Bank Limited the arrangement sits under policy PL9, the business continuity policy. PL9 carries the same four fields as the other eight policies PL1 to PL8 and is answerable in the same way. The four fields themselves, and why the placement matters, are settled under business continuity.
The second is that it has to be invoked. Nothing happens by itself. Somebody looks at the situation, decides the usual route is not coming back soon enough, and says the words. The moment somebody says the words is the invocationThe named decision to leave the usual route and use the other one., and an invocation is a decision rather than an event. Nothing else in the comparison matters as much as the invocation being a decision rather than an event.
The third is that the arrangement carries its own performance figure. For each service, management sets a recovery time objectiveThe time the arrangement is designed to take, set inside the promise.. The objective is the time the arrangement is designed to take once it has been switched on. At the invented bank these are set for all seven important business services S1 to S7: one hour on S1 payments and remittances, three hours on S2 branch counter service and cash withdrawal, two hours on S3 internet and mobile banking, eight hours on S4 loan disbursal, twelve hours on S5 deposit account opening, eight hours on S6 trade finance issuance, and one hour on S7 treasury settlement. Every one of those is the invented bank's own design figure and none of them comes from a supervisor.
So continuity, complete and on its own, is this: a defined second route, a named person who can call it, a scope, and a design target for how fast it stands up once called. To judge a continuity arrangement is to judge engineering the institution built for itself.
What exactly is operational resilience, taken on its own?
Operational resilienceAn outcome: whether a service stayed available to the outside world through a disruption. is not an arrangement at all. Resilience is an outcome, and an outcome cannot be built, bought or switched on. An outcome can only be measured after the fact. Operational resilience is the answer to one question asked about one named service: how long was the outside world without it, and was that longer than the board said it could be.
Go and stand in an apartment building. There is a water pump, and next to it there is a standby pump that nobody has ever seen run. The standby pump is the arrangement. Resilience is a different question entirely, and it is asked at the taps rather than in the pump room: did water come out of the taps on the fourth floor, and if it stopped, for how long. The residents' committee has decided that on no day may the taps be dry for more than two hours. The two hours is not a fact about pumps. The two hours is a decision about what the people in the flats can live with.
Two objects make the measurement possible, and without either of them there is nothing to measure.
The first is a named service. Not a system, not a department, not a building: something a person outside the institution would recognise as the thing they came for. At Vindhya Commercial Bank Limited seven of these are named, S1 to S7, and they read the way a customer would say them: payments and remittances, branch counter service and cash withdrawal, internet and mobile banking, loan disbursal, deposit account opening, trade finance issuance, and treasury settlement.
The second is an impact toleranceThe board's promise about the longest the outside world can be without a named service., set by the board and by nobody else. At the invented bank, committee G1, the board, sets one against each of the seven: 2 hours on S1, 4 hours on S2, 4 hours on S3, 1 working day on S4, 2 working days on S5, 1 working day on S6, and 2 hours on S7. Every one of those seven figures is the invented board's own. A tolerance states what the outside world can bear, and the board is free to set one the institution cannot currently meet. Being missable is exactly what makes the figure worth anything.
Notice what is missing from that description. There is no mention of what failed, no mention of whose fault it was, and no mention of whether anything was switched on. Resilience is measured on the outside world's experience whatever caused it, and the measurement is a length of time held against a line the board drew in advance. The case carries exactly one such measurement: incident I3, the core banking outage that ran 4 hours and 20 minutes, breached the tolerance of S1, S2, S3 and S7 and stayed inside the tolerance of S4, S5 and S6.
Read the picture again and something uncomfortable comes out of it. The word breach flattens a range. Two of those promises were missed by twenty minutes and two by more than two hours, and a paper that says four tolerances were breached has told the board a true thing while hiding the shape of it. The unit of a resilience measurement is minutes outside the promise, not a count of promises broken.
Both ideas have now been defined on their own. Here is the hinge. A single continuity test is run across all seven services and produces seven measured times. Can those same seven measurements produce two different and equally correct verdicts on how the bank did?
Are they the same thing under two names?
No, and the cleanest way to see it is to notice that the two definitions above are different parts of speech. One is a noun that can be pointed at: a route, a document, a phone number, a switch. The other is a verdict reachable only afterwards, by measuring something and comparing it with a line. An institution can build a continuity arrangement and never find out whether it is resilient, and it can be resilient on a given day without having invoked a single arrangement.
None of this is a word game, and here is the test that proves it. Ask what each one would look like if it were perfect. A perfect continuity arrangement stands up instantly every time it is called. A perfect resilience position means nobody outside was ever without a service for longer than the board said. The two sentences describe different subjects. The first is about the institution. The second is about the people it serves. The second sentence includes all the time that passed before anybody called anything, so an institution can score full marks on the first and still fail the second.
Six criteria separate them, and each one is a change of subject rather than a change of degree. All six are worth having in view before each is taken in turn.
What is the actual unit each one works on?
The unit decides everything downstream, so start with CR1. A continuity arrangement takes as its unit a plan with a scope. Open one and the first sheet states what it covers: this building, this activity, this process. Organising work that way is sensible, and it is organised around the institution's own shape.
Resilience takes as its unit a named service that somebody outside relies on. The unit is chosen from outside the institution looking in. Nobody standing at a counter has ever asked for a process; they asked to take money out. So S2 at the invented bank is written as branch counter service and cash withdrawal and not as the machinery that makes it possible.
The consequence is arithmetic. A bank can hold thirty continuity plans and seven important business services, and the two sets do not map one to one. One service leans on several plans, and one plan supports several services. Counting plans reveals nothing about how many services are covered. A plan count is never a resilience answer.
Who sets the target in each, and what makes one of them a promise?
CR2 is the criterion that decides whether either figure can ever be missed. At Vindhya Commercial Bank Limited the recovery time objective on each service is management's figure. The objective is a design target: what the arrangement is being built to deliver. The impact tolerance is the board's figure, set by committee G1, and it is a statement about what the outside world can bear.
Why does it matter who holds the pen? Because a target set by the people who will be measured against it tends to describe what they can already do. If the same group set both, the tolerance would quietly become a restatement of the objective, and a promise that cannot be missed is not a promise at all, it is a description. The separation of the two figures across two different sets of people is what creates the gap between them, and that gap is the whole subject of what follows.
The gap has a name. The distance between the promise and the design is the headroomThe distance between the promise and the design, which is what allows a missed objective to still keep a promise.. At the invented bank the headroom is computable for four of the seven services. S1 is a 2 hour promise against a 1 hour design, so 60 minutes of headroom, half the promise. S2 is 4 hours against 3 hours, 60 minutes, a quarter of the promise. S3 is 4 hours against 2 hours, 120 minutes, half. S7 is 2 hours against 1 hour, 60 minutes, half.
For S4, S5 and S6 the headroom cannot be worked out at all, and the reason is worth reading twice. Their promises are written in working days and their design figures in hours. Those are two different units. The board never wrote down how many hours a working day is, so the two units cannot be subtracted. Doing it anyway is not arithmetic, it is an assumption. For three of the seven services, therefore, the headroom, and every reading that depends on it, is stated as unavailable rather than estimated.
Who sets a recovery time objective, and who sets an impact tolerance?
What is the target actually measuring?
CR3. The recovery time objective measures the arrangement: how long the other route took to stand up and start delivering. The impact tolerance measures the outside world: how long a person who wanted the service could not have it, from the first moment they could not have it to the moment they could again.
Both are lengths of time in minutes, and that is exactly why they get confused. The two lengths are about different subjects, on different stopwatches, started by different events. Placed in the same column of the same table without labelling which is which, they make a document that looks complete and answers nothing.
Which of the two counts the time the outside world was without the service?
When does the clock start in each, and why is that the sharpest difference?
Of the six, CR4 is the one worth remembering above all the rest. The two measures start their stopwatches at different moments on the same morning, and the distance between those two moments is where almost every argument about a resilience report actually lives.
The continuity clock starts at the invocation. Timing from the call is the design of the thing: the arrangement is being asked how fast it stands up once it is called. The resilience clock started earlier, and it started without anybody's permission. The resilience clock started the moment the first person outside the institution tried to use the service and could not.
Between those two moments sits a stretch of time that belongs to nobody in the continuity report. Somebody has to notice. Somebody has to work out whether this is a blip or a failure. Somebody with authority has to be found, briefed and persuaded, and at nine on a Saturday morning that somebody may be at a wedding. Every minute of noticing, diagnosing and authorising counts against the board's promise and against nothing else in the entire measurement system.
The two starting points are also why a continuity test result flatters the arrangement without anybody intending it to. In a test, the failure is known in advance, so noticing takes no time; the authoriser is sitting in the room, so deciding takes no time. The forty minutes between the failure and the invocation shrinks to nothing, and the two clocks briefly coincide. In a real failure they do not.
A service fails at 9.00, is noticed at 9.25, the switch is authorised at 9.40 and the other route is live at 10.10. What does each measure record?
What counts as success in each?
CR5. Continuity succeeds when the arrangement performed as it was built to perform: the achieved time landed inside the recovery time objective. Resilience succeeds when the promise held: the time the outside world was without the service landed inside the impact tolerance. Two different sentences, two different subjects, and here is the useful part. Which of the two a document is about can be established without knowing a single technical fact about the institution, just by asking what would count as success.
What does each one assume about the failure?
CR6, and this is the criterion people find hardest to accept. A continuity arrangement is built against failures somebody imagined: loss of the building, loss of the system, loss of the people. Naming those failures is not a criticism. A second route cannot be built without deciding what the second route is a route around, so a plan is necessarily a response to a list.
Resilience assumes nothing about the failure. Resilience asks about the outcome for a person outside, whatever happened, including the causes nobody put on the list and the ones nobody could have. Assuming nothing about the failure is precisely why the measurement is a length of time held against a promise rather than a count of scenarios covered. A list of scenarios can be complete against itself and still be wrong about the world. A time outside a promise is true regardless of what caused it.
Frank Knight drew this distinction in 1921 in Risk, Uncertainty and Profit, separating what can be measured and enumerated from what cannot. A scenario list is an attempt to enumerate. A tolerance is a way of being measured that does not require the list to be complete.
A continuity plan covers loss of the building, loss of the system and loss of staff. Does that make the institution resilient?
Can one test really produce two different verdicts?
One test can produce two, and at this bank one test did. In month 12 Vindhya Commercial Bank Limited ran a continuity test across all seven important business services. Seven achieved times were written down: S1 in 52 minutes, S2 in 3 hours 40 minutes, S3 in 3 hours 10 minutes, S4 in 6 hours, S5 in 9 hours, S6 in 14 hours, S7 in 48 minutes. Seven measurements, taken once, on one day. Nothing below re-measures anything; the only thing that changes is the line the measurements are held against.
Held against the recovery time objectives, the result is four met and three missed. S1 came in 8 minutes inside its 1 hour objective. S2 was 40 minutes over a 3 hour objective, 22.2 per cent over. S3 was 70 minutes over a 2 hour objective, 58.3 per cent over. S4 was 120 minutes inside 8 hours and S5 was 180 minutes inside 12 hours, both 25.0 per cent inside. S6 was 360 minutes over an 8 hour objective, 75.0 per cent over and the worst miss of the seven. S7 came in 12 minutes inside its 1 hour objective. Four met, three missed, being 57.1 per cent and 42.9 per cent of the seven.
Held against the impact tolerances, and only for the four services whose tolerance the board wrote in hours, the result is four of four inside. S1 at 52 minutes was 68 minutes inside a 2 hour promise. S2 at 220 minutes was 20 minutes inside a 4 hour promise. S3 at 190 minutes was 50 minutes inside a 4 hour promise. S7 at 48 minutes was 72 minutes inside a 2 hour promise. Not one of the four measurable promises was broken.
S2 and S3 are the whole point, so look at them for a moment. Both missed the objective and both kept the promise. The headroom between the two figures was put there deliberately by two different sets of people, so missing the design and keeping the promise is not a contradiction. S3 in particular missed its objective by 58.3 per cent and still landed 50 minutes inside the board's line, because that service carries 120 minutes of headroom.
The shape of the result matters more than any single row. Here it is as a table.
| Service | Achieved | Objective | Verdict on the design | Tolerance | Verdict on the promise |
|---|---|---|---|---|---|
| S1 payments and remittances | 52 m | 60 m | Met, by 8 m | 120 m | Inside, by 68 m |
| S2 branch counter and cash | 220 m | 180 m | Missed, by 40 m | 240 m | Inside, by 20 m |
| S3 internet and mobile banking | 190 m | 120 m | Missed, by 70 m | 240 m | Inside, by 50 m |
| S4 loan disbursal | 360 m | 480 m | Met, by 120 m | 1 working day | Not comparable |
| S5 deposit account opening | 540 m | 720 m | Met, by 180 m | 2 working days | Not comparable |
| S6 trade finance issuance | 840 m | 480 m | Missed, by 360 m | 1 working day | Not comparable |
| S7 treasury settlement | 48 m | 60 m | Met, by 12 m | 120 m | Inside, by 72 m |
| The whole test | 7 | 7 | 4 met, 3 missed | 4 | 4 of 4 inside |
The continuity report on that test says three arrangements underperformed. The resilience report says every promise that can be measured was kept. Both statements are correct, both are incomplete on their own, and a bank that writes only one of them has answered only half of what it was asked.
And now the limitation that every honest resilience report has to carry. Without it the four of four reads as far better news than it is. The month 12 test ran on a planned date, with the recovery team on standby, and the failure chosen in advance, so it tested the arrangement and not the surprise. Every minute of noticing and deciding, the forty minutes on the timeline above, was removed before the clock started. In a real failure those minutes are spent, and they are spent against the promise. The four of four is a measurement of a rehearsal.
S2 missed its recovery time objective and stayed inside its impact tolerance. Was that a failure?
What happens to the two counts when the test stops being a test?
The rehearsal problem has a shape that can be watched. Adding time to every achieved figure, standing for the noticing and the deciding that a planned test removes, moves the two counts. The two counts do not move together and they do not move at the same speed. Nothing demonstrates more clearly that these are two measures rather than one measure at two levels of strictness.
The crossings are worked out from the invented bank's own figures and they are these. On the objective, three services are already missing at no added delay, being S2, S3 and S6; a fourth, S1, joins once the delay passes 8 minutes; a fifth, S7, joins past 12 minutes; a sixth, S4, past 120 minutes; and the seventh, S5, past 180 minutes. On the tolerance, and only for the four services where the comparison can be made at all, none is broken at no added delay; S2 breaks past 20 minutes, S3 past 50, S1 past 68 and S7 past 72. Add nine minutes of delay and a fourth service misses its objective while all four promises still hold; add twenty one and the first promise breaks while the objective count has not moved at all.
The month 12 test is re-run with 30 minutes of extra delay on every service. Does the number of broken promises go up?
The surprise premium, read on both lines at once
One control: extra delay in minutes, added to every achieved time from the month 12 test, standing for the difference between a planned rehearsal and a real morning. Two consequences, drawn together: how many of the seven services miss the recovery time objective, and how many of the four measurable promises break. The default of 0 minutes is the month 12 test exactly as it was run, and at that setting 3 services miss their objective and 0 of the 4 measurable promises break. The delay is a dial on the control rather than a figure from the case. The promise count covers only S1, S2, S3 and S7. The promises on S4, S5 and S6 are written in working days, which cannot be compared with a result in hours.
Can one succeed while the other fails?
Both directions happen, and being able to describe each one is the fastest way to prove to yourself that these are genuinely two measures.
Continuity succeeds and resilience fails when a perfect switch happens too late. The other route stands up exactly as designed, well inside the recovery time objective, and the continuity report is immaculate. But the service failed at 6.40 in the morning, nobody was looking until the branches opened, the person who could authorise the switch was reached at 9.15, and by the time the other route was live the outside world had been without the service for longer than the board said it could be. Every minute of that overrun was spent before the arrangement was even asked to do anything. Notice how the case's own record makes this concrete. Incident I9, the vendor payment gateway failure, ran 9 hours and was never declared a crisis: it began on a Saturday, and no single named person held the authority to declare one outside working hours. The plan named a crisis team of 7 roles and did not name who could convene it.
Resilience succeeds and continuity is never tested when the failure is short. Something breaks at 11.10 and is working again at 11.35. Nobody invoked anything, there is no continuity result to report, and every promise held comfortably. The resilience position for that day is good and the continuity arrangement contributed nothing to it. A bank that reports only continuity has nothing to say about that morning at all. Its reporting is blind to the days when the outside world was fine.
Give the cleanest case where continuity succeeds and resilience fails.
Which one does an institution already have, and which does it have to build?
Almost every established bank already has continuity arrangements. The arrangements are old, they are conventional, and they sit under a policy, at the invented bank under PL9. Far less common is the second thing: resilience contains continuity rather than replacing it, so a bank with a tested continuity arrangement has built part of the answer and not the whole of it.
Continuity is one of the arrangements resilience leans on, alongside restoring the technology, handling the incident and declaring a crisis. The four arrangements together are the machinery. Resilience adds the question the machinery is being asked, and the question needs objects that the machinery does not produce by itself.
The third item on that list deserves a name of its own. A service mapThe written record of what one service leans on, without which a tolerance cannot be tested. is the record of everything one service depends on, including the parts somebody else runs. Without it there is no working out which arrangements to invoke for a given service, let alone whether the service can stay inside its promise. An institution that has never drawn a service map does not know what would have to hold for a tolerance to be met, so it has no way of testing one.
A bank has a tested continuity arrangement for every system it runs. What does it still not have?
The failure: a continuity result presented as a resilience position
Here is how it goes wrong, and it goes wrong quietly, in a well written paper produced by competent people. The month 12 result reaches committee G6, the operational risk management committee, as four met and three missed. The paper is titled as a resilience update. Every figure in it is correct, and the paper still cannot answer the question the board actually asked.
The costs run in both directions inside the same document. S2 and S3 missed their objectives and kept their promises, so the paper reads worse than the position actually is. Nothing in it can be measured for S4, S5 and S6, so the paper is silent exactly where the answer is genuinely unknown, and silence reads as nothing to report rather than as nothing measurable. And a board reading that paper is being shown three underperforming arrangements. The same board is not being shown that four of seven promises broke in month 3 during incident I3, a different measurement on a different day for which no paper was ever written.
The same trap catches the loss figure. Incident I3 booked a net loss of Rs 3.2 crore against the year's Rs 43.8 crore of net operational loss. The loss figure is true and it is silent about the four broken promises. The amount an institution paid and the time its customers were left without are two different measurements of two different subjects, and the separation between them is settled under operational resilience. Incident I9 booked a net loss of Rs 3.2 crore too, so the incident has to be named whenever that figure is quoted. The two events are not interchangeable.
Who reads this, and what they do with it
Two banking relationships are on the table in front of Girish Talwalkar, group treasurer at the invented Nirjhar Industries Limited, and both banks have sent him a statement about how they handle disruption. One statement says that its continuity arrangements were tested in the last twelve months and that the test was passed for most services. The other says that for each named service it has published the longest its board is willing for a customer to be without it, and reports how long customers actually were without it. Only the second statement tells him anything about what happens to Nirjhar Industries Limited on a bad day.
His working question is short and can be asked of anyone. Asking whether the bank is resilient invites a yes. The question to ask instead is this: for the service in use, what has been promised, who set that figure, and what did the last measurement against it say. If the answer comes back as a test pass rate, he has been handed a continuity answer to a resilience question, and he now knows exactly which follow-up to send.
The same three questions work for a household choosing where to keep the salary account, an analyst reading an operational risk disclosure, and an internal auditor deciding what to test next year. None of them needs to understand any technology to ask them.
Where do the expectations actually come from?
The Reserve Bank of India at rbi.org.in is the source of what an Indian bank must actually maintain and report on business continuity and on operational resilience. The Bank for International Settlements at bis.org is where the Basel Committee publishes the international standard that an Indian requirement implements. Naming only that standard is the common and confident error: the standard is the origin, and the Indian requirement is the thing a bank here is held to. Whether a requirement uses these two words at all, and how it frames them, differs between requirements and is stated only in the requirement's own text.
Every tolerance and every recovery objective above is a decision of the invented bank's own board and management, and none of them was handed down by a supervisor. What a supervisor requires, from what date, and in which reporting window, is set by the issuing body and changes, so the current position stands only in that body's own current text.
How a continuity arrangement is designed, invoked and written down, and what a plan contains, are covered separately. How an important business service is chosen, how an impact tolerance is set and how a service map is built are covered separately. Deciding under pressure with incomplete information, restoring technology after a failure, and running one incident from detection through to learning are each covered separately too. How an operational loss is measured, and how a control is tested, are covered separately.
What is the one line to carry away from all of this?
Continuity asks whether the arrangement performed as it was built to perform. Resilience asks whether the outside world was harmed. The month 12 test at the invented bank answered the first question with four met and three missed, and the second, where it could be answered at all, with four of four inside, and two of the services that missed the design kept the promise. Both counts came from one test, run once, on one day, and a paper carrying only one of them has answered half a question and looked complete doing it.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | What binds a regulated Indian bank on business continuity, operational resilience, outsourcing arrangements and the reporting of an incident | rbi.org.in |
| Bank for International Settlements | The Basel Committee's international standard on operational resilience, named as the origin of the expectation rather than as the requirement | bis.org |
| Securities and Exchange Board of India | The equivalent expectations where the entity is a market intermediary rather than a bank | sebi.gov.in |
| Indian Banks Association | Material on Indian banking operational convention | iba.org.in |
| Frank Knight | Risk, Uncertainty and Profit, 1921, the separation of what can be measured and enumerated from what cannot | University of Chicago Press |
Vindhya Commercial Bank Limited and Nirjhar Industries Limited are invented.
Educational material. Not advice on any investment, tax, budget or market position.
