Business Continuity and Disaster Recovery: Two Different Jobs
Business continuity is keeping a service running by another route while the usual route is unavailable, and it is a business arrangement. Disaster recovery is restoring the technology the service normally runs on, and it is a technology one. Continuity asks how the customer still gets served today. Recovery asks how the machinery comes back. Both are built to two objectives: how fast, and how much recorded work may be lost.
Continuity and recovery get run together constantly, in board papers and in ordinary conversation, and the merge costs real money. The two are answers to two different questions about the same bad morning, and an institution can do one of them perfectly while failing the other. Everything below is worked on Vindhya Commercial Bank Limited, an invented bank, on its seven services, on the seven objectives set against each of them and on the whole month 12 test result.
The two objectives about to be met look technical and are not. Both of them are decisions somebody at board level took about how much the outside world should be made to put up with, and no technical training is needed to follow either one.
What is business continuity, and what is it actually for?
Business continuityKeeping a service running by another route while the usual route is unavailable. is the arrangement by which a service keeps reaching the person who uses it when the ordinary way of delivering it has stopped. Notice what that definition does not mention. The definition does not mention what broke, it does not mention fixing anything, and it does not mention how long the repair will take. Continuity is entirely about the customer being served, and not at all about mending whatever failed.
Take a school bus that will not start on a Tuesday morning. Continuity is not the mechanic. Continuity is the parent group chat that arranges four cars, a list of who is picking up whom, and one person deciding at seven fifteen that the cars are happening. The children reach school. The bus is still broken and will be broken at lunchtime, and the service was delivered anyway. Somebody else, later, deals with the bus.
A bank does the same thing at a larger scale and with more paperwork. If the branch counter cannot process a withdrawal because the branch has no power, continuity is a nearby branch, or a mobile channel, or a manual register with a per customer cash limit and two signatures. Each of those is a standby arrangementA second way of delivering a service, kept ready but not normally used. or a workaroundA manual or partial route that delivers less than the usual service and delivers something., and both are business decisions with a cost attached, made long before the morning they are needed.
The reason continuity is a business arrangement and not a technical one is that only the business can decide what a reduced service is allowed to look like. Can the manual register go up to Rs 50,000/- per customer or Rs 5,000/-? Are new account applications simply refused for the day, or taken on paper and keyed in later? None of those is a question about machinery, so nobody who runs machinery can answer them. Each one is a question about which customers get told no.
What is disaster recovery, and why is it a different job?
Disaster recoveryRestoring the technology a service normally runs on after it has failed. is the work of bringing the machinery back: the records, the processing, the connections, restored to a state the institution can carry on from. Recovery is measured by whether the machinery that stopped is running again, and it is planned, staffed and rehearsed by the people who run that machinery.
Back to the bus. Recovery is the mechanic, the spare part, and whether the bus runs tomorrow. The repair is real work, it is necessary, and while it is happening not one child is any closer to school. The distinction is that plain, and it survives translation into a bank without a single change.
The two jobs can succeed and fail in any combination. Nothing shows more clearly that they are not one job. A branch with no power has lost nothing technical at all, so there is nothing to recover and the entire answer is continuity. Nobody outside notices a processing failure that has no customer facing effect until the evening batch, so that one is almost entirely recovery. And the case that matters most is the one where recovery finishes cleanly and the service is still not being delivered.
Why are these two jobs and not two words for one job?
Because they are funded, staffed, planned, tested and reported by different people, and because the failure that triggers both of them does not care that they were merged in a document. Keeping the two apart is not a vocabulary exercise. The separation decides whether anybody is answerable for the customer while the repair is going on.
There is a practical test for any plan. Read it and ask: if this works perfectly, is a customer served? If the honest answer is no, and it is only that the machinery is running again, the document is a recovery plan that somebody has labelled continuity. A plan that restores things is not a continuity plan just because the word continuity is on the front of it.
A branch loses power for the whole day. Every system in the bank is running normally. Is that a continuity problem or a recovery problem?
What is a recovery time objective, and who decides it?
A recovery time objectiveHow quickly the arrangement is built to bring one named service back. is a single number attached to one named service: how quickly the arrangement is built to have that service reaching customers again. The objective is a design target. The objective is what the money was spent to buy, and it is written down before anything goes wrong so that the arrangement can be built against it and then tested against it.
Two things about it surprise people. The first is that it belongs to a service and not to a system, so the same piece of machinery can sit under a service asking for one hour and a service asking for twelve, and the arrangement has to satisfy the tighter of the two. The second is that it is a business decision. The number states how long the outside world should be made to wait. Deciding how long customers wait is a judgement about customers, not a description of what the equipment can already do.
The second of the two is where it goes wrong most often. If the people who run the machinery are asked to propose the number, what usually comes back is roughly what the current arrangement already achieves, dressed up as a target. Nothing has been decided at all; the present has been written down and given a name. The board setting it, in the outside world's terms, is the only version that can ever demand a change.
What is a recovery point objective, and why is it not measured in time?
A recovery point objectiveHow much recorded work may be lost, measured as an amount of data rather than an amount of time. answers a completely different question: when the service comes back, how much of the work that was recorded before it stopped is allowed to be missing? The objective is quoted in minutes, and the minutes confuse everybody. The minutes measure an amount of work rather than a delay. Five minutes here means five minutes of instructions, entries and applications that vanished, not five minutes of waiting.
Think about a shopkeeper who writes each sale in a paper book and copies the day's total into a ledger every evening. If the book is lost at four in the afternoon, the ledger is missing everything since last night, and reopening the shop tomorrow does not bring those entries back. How often the copy is taken decides how much can be lost, and that is a decision about the copying and not about the reopening. The two numbers are set by two different considerations and there is no arithmetic that turns one into the other.
At this invented bank the two objectives on the seven services do not sit in any fixed proportion to one another. Divide the second by the first, service by service, and the answer ranges from zero to one third. Knowing how fast a service must come back says nothing about how much of its work may be lost. Both numbers have to be set, and neither can be inferred from the other.
S1 payments and remittances may lose no recorded work at all. S5 deposit account opening may lose four hours of it. What is the actual difference between those two services?
Why is the pair set service by service and never once for the whole institution?
Here is the full set at this invented bank, and read down the two number columns before reading anything else. Every figure is the board's own, and none of it is normal practice anywhere.
| Service | What it is | Back in | May lose |
|---|---|---|---|
| S1 | Payments and remittances | 1 hour | nothing |
| S2 | Branch counter service and cash withdrawal | 3 hours | 15 min |
| S3 | Internet and mobile banking | 2 hours | 5 min |
| S4 | Loan disbursal | 8 hours | 1 hour |
| S5 | Deposit account opening | 12 hours | 4 hours |
| S6 | Trade finance issuance | 8 hours | 1 hour |
| S7 | Treasury settlement | 1 hour | nothing |
The spread across the first column runs from one hour to twelve, a factor of twelve, and the spread across the second runs from nothing at all to four hours. Any single number set once for the whole institution is simultaneously too generous for S1 and far too demanding for S5, and it is wrong in both directions at the same time. Setting it once looks like simplification and is actually two mistakes bought together.
The everyday version is a household that decides everything must be sorted within an hour. The gas cylinder running out has to be sorted within an hour, and so does the leaking tap. One of those is dinner and the other one can wait until Sunday, and a household that treats them alike either spends its weekend on the tap or eats late. Deciding service by service is not bureaucracy; it is how the effort ends up where it changes something.
One boundary constrains every comparison that follows. Behind these seven design numbers sits a separate board promise, the impact tolerance, set out under operational resilience. For S4, S5 and S6 that promise is written in working days rather than in hours. The board never wrote down how many hours a working day is, so the two units cannot be subtracted. The promise on those three cannot be turned into hours, and cannot be compared with anything measured in hours. Every minute figure below is a recovery time objective and never a promise, and the two are not the same object.
A bank sets one recovery time objective of four hours for everything it does, on the grounds that one number is simpler to manage. What has it just done?
What else has to fit inside a recovery time objective?
A recovery time objective catches out almost everybody who meets it for the first time. The objective is not a budget for doing the recovery. The objective is a budget for the whole gap between the service stopping and the service working again, and the doing is only the last third of that gap.
Three things happen in that gap, in this order. Somebody has to notice the service is down. Noticing is not instant, and it is sometimes very slow indeed when the first evidence is a customer telephoning. Somebody with the authority has to decide to switch to the other route. The switch carries a cost, so the decision is not taken in a second. Only then does anybody start switching. The invocationThe moment somebody decides to switch to the other route, which is a decision and not an event. is a decision made by a named person. Until that person makes it, the arrangement standing ready is doing nothing at all.
Read the drawing again and notice where the loss happened. Nothing in it was slow. Fifteen minutes to work out that something is wrong is fast when the first symptom is a customer complaint. Ten minutes to find the person who can authorise a switch is fast if that person is at their desk. Forty minutes to switch is comfortably inside the hour on its own. Three reasonable numbers, one broken objective, and no villain anywhere.
The two most valuable things that can be bought against a tight objective are usually not recovery capability at all. The first is a way of finding out sooner, and the second is a standing rule about who may say go, written down in advance so that nobody spends the third stage of the budget looking for a person. Both are cheap and neither is technical.
Three things happen inside a recovery time objective, one after another. Which set is it?
Seven services were tested at this invented bank in month 12. Which one missed its objective by the largest margin?
What happened when this invented bank actually tested it?
In month 12 Vindhya Commercial Bank Limited ran a continuity test across all seven important business services. The test measured one of the two objectives, the recovery time. Four services met theirs and three did not, so 57.1 per cent met and 42.9 per cent missed. The whole result repays reading row by row rather than as a headline.
| Service | Objective | Achieved | Result | Margin |
|---|---|---|---|---|
| S1 payments and remittances | 60 min | 52 min | Met | 8 min inside, 13.3 pc |
| S2 branch counter and cash | 180 min | 220 min | Missed | 40 min over, 22.2 pc |
| S3 internet and mobile banking | 120 min | 190 min | Missed | 70 min over, 58.3 pc |
| S4 loan disbursal | 480 min | 360 min | Met | 120 min inside, 25.0 pc |
| S5 deposit account opening | 720 min | 540 min | Met | 180 min inside, 25.0 pc |
| S6 trade finance issuance | 480 min | 840 min | Missed | 360 min over, 75.0 pc |
| S7 treasury settlement | 60 min | 48 min | Met | 12 min inside, 20.0 pc |
| Seven services | 4 met, 3 missed | 57.1 pc and 42.9 pc |
Two things in that table are worth more than the headline count. The first is that the two tightest objectives, S1 and S7 at one hour each, both passed, and they passed by 8 minutes and 12 minutes. The services with the least room to spare are the ones that got the attention. The worst failure of the day was S6 trade finance issuance at 75.0 per cent over, and nobody would have called trade finance issuance urgent. That is the ordinary shape of this result, and it is worth expecting rather than being surprised by.
The second is that a count of four out of seven hides a range. The three misses were 40 minutes, 70 minutes and 360 minutes over. One of the three misses is nine times the size of another, and flattening all three into one word loses the only fact a board could act on.
A recovery takes 3 hours in total, measured from the moment the service stopped. How many of the seven services meet their objective?
The recovery time budget, against how many services make it
One control: the whole time from the service stopping to the service working again, in minutes, noticing and deciding included. One consequence: how many of the seven meet the objective they were built to. The result does not fall smoothly. The count sits still for hours and then drops.
Educational illustration. At the default setting of 180 minutes, 4 of the 7 services meet their objective, being S2, S4, S5 and S6, and S1, S3 and S7 do not. The month 12 test also came out at 4 met and 3 missed, and from a different three. The misses there were S2, S3 and S6. The total on this control includes noticing and deciding and not only the switch. How much recorded work may be lost is a separate objective and is not drawn here. The board promise sitting behind these services is a different object again. Its figures for S4, S5 and S6 are written in working days, and none of them is converted into minutes here.
What could that test not test?
The single most important sentence about the month 12 result is not in the table. The month 12 test ran on a planned date, with the recovery team on standby, and the failure chosen in advance, so it tested the arrangement and not the surprise. Those are two different claims, and the one people quote is not the one the test supports.
Look back at the three stages that fit inside a recovery time objective and count how many of them a planned test actually exercises. Noticing takes no time when everybody already knows the failure is coming at eleven o clock. Deciding takes no time when the decision was taken at the planning meeting three weeks earlier. The switch is all that is left, it is the stage a competent arrangement is most likely to handle well, and it is the only stage the clock ran on.
The four passes can now be read honestly. S1 passed by 8 minutes and S7 by 12. The 8 and 12 minute margins are not comfort, they are the entire allowance that noticing and deciding would have to fit inside on a morning nobody planned, and a fifteen minute delay in working out that something is wrong would turn both of those passes into misses without the switch getting one second slower.
None of this makes the test worthless. Testing the arrangement is worth doing, it is the only way to find out that S6 is 75.0 per cent over, and it found exactly that. The honest claim is simply smaller than the claim people usually make from the same evidence, and stating the limitation beside the result is what keeps the two the same size.
The month 12 test met four of seven objectives. Is that a good result?
The failure: everything was restored, and the service was still down
The failure the whole distinction exists to prevent needs nobody to be careless or wrong. In the month 12 test S6 trade finance issuance came back in 14 hours against an 8 hour objective. On that same day a technology report could say, entirely correctly, that every system had been restored. A service report would have to say that trade finance issuance was unavailable for 14 hours, 75.0 per cent over and the worst result of the day.
The failure is not that either report is wrong. The failure is that only the first report gets written. The people who restore machinery know precisely when the machinery is back, because that is their work and they can see it. Nobody has been made answerable for knowing when the service is back, so nobody measures it, so it does not appear. The largest single miss of the day is invisible on the only document that was produced.
The same shape appears at a wedding. The generator is fixed by eight o clock and the caterer confirms it. Whether four hundred people ate is a question nobody at the generator was asked, and if the only person reporting is the one who fixed the generator, the evening is recorded as a success. Two honest answers to two different questions, and only one of them was ever asked.
The repair is unglamorous and cheap: name a person answerable for service availability who is not the person answerable for restoration, and make the service report a separate document with its own row per service. The repair costs one sheet of paper, and in exchange a 75.0 per cent overrun cannot leave the building disguised as a completed restoration.
The technology report says everything was restored by lunchtime. What question has it not answered?
What does a continuity plan actually contain?
A continuity planThe written arrangement naming the other route, who invokes it and what the customer is told. is a short document and most of what makes it useful is a set of names and sentences rather than any capability. Four things have to be in it. Miss any one and what remains is a description rather than a plan, and the difference shows up on the one morning it is opened.
Notice which of the four the money usually goes to. Part one looks like preparation, so it is the expensive one and the one every plan already has. Parts two, three and four cost almost nothing and are missing far more often. The most common reason a continuity arrangement underperforms on the day is not that the other route was absent but that nobody had settled who could say go, and no capability closes that gap.
Part three deserves a sentence of its own because it is the part the customer actually experiences. A person who cannot withdraw cash and is told, at the counter, that the counter is down for about two hours and the branch two streets away is working, has had a bad morning. The same person told nothing at all has had a different kind of morning, and it is the second one that turns up as a complaint and, eventually, as a conduct question. Telling the customer is part of continuity, not a communications afterthought.
Who is answerable for continuity, and where is it written down?
Continuity at this invented bank is not a project that happened once. Continuity is policy PL9, one of nine policies numbered PL1 to PL9, and the placement is what makes it a standing arrangement. Each of the nine names an owner, an approver, a review cycle and the limits it cascades into. A policy carries a named person and a review date, and a completed project carries neither. The difference is between an arrangement somebody keeps current and one that was correct on the day it was built.
There is a second reason the placement matters, and it is about who gets to argue. If continuity lives inside the technology function as an operating practice, the recovery time objective is negotiated by the people who would have to meet it. The failure set out above arrives by a different door. If it lives as a policy with a business owner and a board level approver, the objective is set outward and the arrangement has to move to meet it. Where the document sits decides which way the argument runs.
The invented case does not record who the owner of PL9 is. The case does record the shape: nine policies, each with those four fields, and business continuity as one of them rather than as a folder somebody keeps.
Where is business continuity written down at this invented bank, and why does the answer matter?
What a customer, an analyst and a household each do with it
A large corporate customer does this before it concentrates its payments anywhere. Girish Talwalkar, group treasurer at the invented Nirjhar Industries Limited, will not ask his bank about its technology, and he does not need to. The two questions that get him what he needs are: which of the bank services carries an objective of an hour or less, and what does the last test say was actually achieved on each. Two questions, no technical vocabulary, and an answer that tells him whether his payroll run has a second route on the day something breaks.
An analyst reading a bank from outside uses the same two numbers differently. A published objective is a statement of intent and costs nothing to make. A published test result against that objective is evidence. The gap between an institution that discloses objectives and one that discloses results is the whole distance between an intention and a measurement, and it is visible without any access to the institution at all.
And the household version, on a smaller balance. The case is the single account a household runs every standing payment through. The other route is a second account with a month of expenses in it, and the four parts of a plan translate exactly: the second account, who in the house decides to move to it, what the landlord hears while it is happening, and how the standing instructions get moved back afterwards. Most households have part one and none of the other three, the same shape as most institutions.
What can these two objectives not promise?
Quite a lot, and saying so plainly is part of using them honestly. An objective is a design target and not an outcome. Writing 60 minutes against S1 does not cause a 60 minute recovery to happen; it states what the arrangement was built for and what would count as falling short. That is worth real money and it is not the same as the service coming back in an hour.
The tested figure is smaller than it looks, for the reason drawn earlier: a planned test measures the switch and not the noticing or the deciding. The untested figure is smaller still. The invented case measured recovery time in the month 12 test and did not measure how much recorded work was lost, so all seven of the second objectives stand as decisions and none of them as results. Saying so is itself the finding, and how the technology behind a service is restored, including what an objective of no data loss at all actually asks for, is taken further under disaster recovery.
A control question sits inside the failure block above, and it belongs to a different subject. The missing service availability report is not a control that failed a test. The missing report is a control nobody designed, a different kind of finding, and it is caught by a different process. How a control is designed, how it is tested for operating effectiveness, how a deficiency is recorded and how it is put right is covered separately.
One more limit, and it is about where the other route lives. A standby arrangement can sit with somebody outside the institution, and then its continuity is somebody else business to run and the institution business to be satisfied about. The invented bank has one incident in its year of exactly that shape, incident I9, a vendor-hosted payment gateway failure that ran 9 hours. Assessing a third party as a risk in its own right is covered separately, and so is the record of what each incident cost.
On that record: incident I3, the core banking outage that ran 4 hours and 20 minutes, booked a net loss of Rs 3.2 crore. Incident I9 booked the same Rs 3.2 crore, so both are always named rather than quoted as a single figure. Neither number measures anything about the two objectives, a point settled under operational resilience.
What binds an actual Indian bank here
The mechanism described here is universal. Another route to the same service, and restoring the machinery behind it, are the same two jobs in any country and at any size. Local law decides what an institution is required to maintain, how often it must be tested, what must be reported and to whom.
For a bank in India, the Reserve Bank of India at rbi.org.in is the body that sets what actually binds on business continuity, on testing, on reporting an incident, and on an arrangement that sits with a third party. The Basel Committee, publishing through the Bank for International Settlements at bis.org, is where the international standard behind the Indian requirement comes from, and naming only the international one is the confident and common mistake here. Where the entity is a market intermediary rather than a bank, the Securities and Exchange Board of India at sebi.gov.in applies instead.
None of the seven objectives above comes from any of those bodies. Each of them is a decision taken by the board of one invented bank. Confirm anything binding at the issuing source.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | What a regulated bank in India must maintain, test and report on business continuity, and on an arrangement that sits with a third party | rbi.org.in |
| Bank for International Settlements | The Basel Committee international standard behind the Indian requirement, and the operational risk event categories named in this case | bis.org |
| Securities and Exchange Board of India | What applies instead where the entity is a market intermediary rather than a bank | sebi.gov.in |
| Indian Banks Association | Material on Indian banking operational convention | iba.org.in |
Vindhya Commercial Bank Limited, Nirjhar Industries Limited and Girish Talwalkar are invented.
Educational material. Not advice on any investment, tax, budget or market position.
