Disaster Recovery: Restoring Technology After Failure
Disaster recovery is the job of restoring the technology underneath a service once that technology has stopped working. Two figures, both chosen by the business, define what a good recovery looks like. One fixes how soon the technology must be usable once more. The other fixes the largest quantity of already recorded work the institution will accept losing. The second figure is a volume of data and never a stretch of downtime.
One separation carries this whole guide, and almost every reader arrives having missed it. Time lost asks how long the service was down. Data lost asks how far back the institution has to go to find a usable copy of the record, and therefore how much of what people already did has disappeared. A service can be back in fifteen minutes having lost four hours of records, and a service can take a full day to come back having lost nothing at all. Every worked figure below belongs to Vindhya Commercial Bank Limited, an invented bank.
These decisions need no knowledge of technology. The hard part is not any product, supplier, design or technique, but a business question about the outside world.
What is disaster recovery actually restoring?
Start with a shop. The shape is identical there, and nobody has to know anything to see it. A cloth shop on a busy street sells on credit, and every sale goes into a ruled notebook: name, date, amount. Two things have to work for that shop to trade. The shutter has to open, and the notebook has to be there.
Now break each one. If the shutter jams, the shop is shut. Customers stand outside, nothing is sold, and the moment somebody prises the shutter open the shop is exactly as it was. If instead the shutter opens fine but the notebook was left out in the rain overnight, the shop is open and trading and something else is gone: three weeks of who owes what. Nobody is standing outside. The loss is invisible from the street and it is far worse.
Disaster recoveryPutting technology back into working order after it has broken, so the service riding on it can be delivered once more. is the work of putting both of those back after a failure, and the reason it needs two separate objectives is that they are two separate losses. Restoring the machinery decides how long nobody could use the service, and restoring the record decides how much of what people already did still exists. An institution that plans only for the first has planned for the shutter and not for the notebook.
Somewhere behind every service there is a place the technology can be run from when the usual place is unavailable. A standby siteA second place the technology can be run from, which the people outside never see or hear about while it is working. is that arrangement. The only thing worth holding on to is that the outside world is never supposed to know it exists. None of the decisions here are made at such a site, so how one is built is a separate subject.
What does the recovery time objective say, and why is it not enough on its own?
The recovery time objectiveHow quickly the technology has to be usable again after a failure, set as a design target inside the promise the board made. is the shorter half of the story and it is settled separately. The objective is the figure the recovery arrangement is built to deliver: how quickly the technology has to be usable again. At Vindhya Commercial Bank Limited it runs from 1 hour on S1 payments and remittances and S7 treasury settlement, through 2 hours on S3 internet and mobile banking and 3 hours on S2 branch counter service and cash withdrawal, out to 12 hours on S5 deposit account opening. Every one of those figures was set by the bank's own board.
The recovery time objective cannot say the thing the shopkeeper cared about. A recovery time objective is entirely silent on how much of what people had already done comes back with the service. Bringing the shutter up in one hour meets a one hour objective, whether the notebook is intact, three weeks short, or gone. So the same failure gets a second number, and the second number is the recovery point objective.
What does a recovery point objective measure, and why is it not a time?
The recovery point objectiveA ceiling on vanished work, set as the volume of record lying between the last usable copy and the moment everything stopped. answers a different question: when the record is put back, how far back does it go? Somewhere there is a copy of the record that is usable, and it was made at some moment. Everything that happened between that moment and the failure is not in the copy. The recovery point objective is the institution stating, in advance, the largest amount of that lost work it is willing to accept.
Here is the part that trips people, and it is worth sitting with. The objective is written down as a duration, five minutes or one hour or four hours, and that duration is not a period of downtime. The duration measures the record. Five minutes means the record may be missing up to five minutes worth of what people did. The objective says nothing whatever about how long the service was unavailable. The outage might have been ten seconds or ten hours.
To restoreTo put a copy of the record back into use so the service can run on it again, at the state that copy was in when it was made. a record is therefore to accept a specific, known, quantified hole in it. A known hole is not a defect in the method. The hole is the method. The only institution with no hole is one that arranged, in advance and at real cost, for the record never to exist in a single place, and that arrangement gets its own block further down.
Why are data lost and time lost two different measurements?
Take one service and hold it still. S3 internet and mobile banking at Vindhya Commercial Bank Limited carries a recovery time objective of 2 hours and a recovery point objective of 5 minutes, both of them set by the bank's own board. Both figures are written in minutes. Neither is on the same scale as the other, and the reason is simple: the two are not counting the same thing.
The 2 hours counts unavailability. The unavailability is the stretch during which a customer opens the app, gets nothing, and tries again later. The 5 minutes counts record. The 5 minutes is the largest gap the bank is willing to have between the last usable copy and the failure. At worst a customer who moved money four minutes before everything stopped may find, when the service comes back, that the transfer is not there. One number is about waiting and the other is about work that has vanished, and no amount of hurrying fixes the second.
Read the two figures against each other and the shape of the service appears. 120 minutes divided by 5 minutes is 24, so S3 may be unusable for twenty-four times as long as it may be short of records. The ratio of twenty-four is not a rule, and it is not the same for any other service here. The ratio is a statement about internet and mobile banking specifically: people will wait a couple of hours, and they will not accept that a payment they watched succeed has quietly stopped existing.
A service came back twelve minutes after it failed, and four hours of records were missing when it did. Which objective did it fail?
What do the seven services at this bank actually carry?
Vindhya Commercial Bank Limited has named seven important business services, numbered S1 to S7, and set a recovery point objective on each one. A guess ventured on one familiar service before the table is read is what makes the table stick.
The time objective on S3 internet and mobile banking is 2 hours. Before the table gives it away: is the data objective on that same service tight or loose?
Here are all seven. Every entry below was decided by the bank's own board.
| Service | Recovery time objective | Recovery point objective | What that second figure is saying |
|---|---|---|---|
| S1 payments and remittances | 1 hour | 0 | No lost work at all is acceptable |
| S7 treasury settlement | 1 hour | 0 | No lost work at all is acceptable |
| S3 internet and mobile banking | 2 hours | 5 minutes | Minutes, because the customer saw it happen |
| S2 branch counter service and cash withdrawal | 3 hours | 15 minutes | A short queue of counter work can be redone |
| S4 loan disbursal | 8 hours | 1 hour | An hour of case work can be picked up again |
| S6 trade finance issuance | 8 hours | 1 hour | An hour of case work can be picked up again |
| S5 deposit account opening | 12 hours | 4 hours | The paperwork is still sitting on the counter |
Read down the two number columns and notice that they do not move together. S3 has the tightest time objective after S1 and S7, and the third tightest data objective behind the same two. S5 has the loosest of both, but not in the same proportion. Its time objective is 12 hours against S3 at 2, a ratio of six. Its data objective is 4 hours against S3 at 5 minutes, a ratio of forty-eight. The two columns are set by two different conversations and a reader who assumes one predicts the other will get every service wrong.
What does a recovery point objective of zero actually ask for?
Two of the seven carry a zero: S1 payments and remittances, and S7 treasury settlement. A reader running down the column reads 4 hours, 1 hour, 15 minutes, 5 minutes, zero, and takes zero as the last rung of a ladder. Zero is not a rung. Zero is not the small end of this scale, it is a different arrangement standing on the other side of it, and the price behaves accordingly.
Go back to the shop. Suppose the shopkeeper wants to lose less when the notebook is ruined. She can copy the day's credit entries into a second notebook every evening, and at worst she loses one day. She can copy them at lunch as well, and lose half a day. She can copy after every tenth customer, and lose nine sales. Every one of those is the same arrangement, done more often, and each step costs her more of her own time.
Now suppose she is asked to lose nothing. There is always a gap between one copy and the next, and the fire may land inside it. Copying more often does not get there, no matter how often. The only way to lose nothing is to stop copying and start duplicating: carbon paper under every leaf of the notebook, and the entry exists twice at the instant the pen moves. Carbon paper is a different way of running the shop. She buys different stationery, writes differently, and the cost arrives once and stays.
Duplication is the whole content of a zero. A zero asks for the record to exist in more than one place at the moment it is created, rather than for the record to be copied faster. How an institution arranges that is a separate subject, but the shape of the decision is exactly the shopkeeper's, and so is the shape of the bill.
Two services at this bank carry a recovery point objective of zero. What does that figure ask the institution to build?
Why those two services and not the others? Because of what the lost work is. An instruction to move money that the bank accepted, acknowledged and then lost cannot be reconstructed by anybody, at any price. The customer does not have it. The counterparty does not have it. Nobody in the branch remembers it. The problem is not that losing it would be expensive. There is no procedure at all for getting it back. Zero is what an institution writes when the honest answer to where else does this exist is nowhere.
How does the interval between copies set the worst case?
An objective is what the business asked for. Whether it is actually met is decided by something duller and more physical: how much time passes between one usable copy of the record and the next. Call that the copy intervalThe gap between one usable copy of the record and the next, which sets the largest amount of work a failure can take away.. The copy interval is the only quantity here that an institution controls directly.
Now think about when the failure lands. If copies are taken every fifteen minutes and the failure lands one minute after a copy, one minute of work is gone. If it lands fourteen minutes after a copy, fourteen minutes are gone. Planning is done on the worst case. The worst case is the failure landing immediately before the next copy, so the work lost is taken as the whole interval.
Why the worst case and not the average? Because the objective is a promise about what the institution will accept, and a promise tested against an average is not a promise. The shopkeeper who copies her notebook every evening does not tell a customer she loses half a day on average. She knows that a fire at nine at night takes the whole day, and she plans for the nine o'clock fire.
Put the seven objectives in a row against a single copy interval and something ugly appears. The thresholds are 0 for S1, 0 for S7, 5 minutes for S3, 15 minutes for S2, 1 hour for S4, 1 hour for S6 and 4 hours for S5. A single interval is either above or below each of those, so the count of services that would lose more than the board said they could does not slide upward as the interval grows. The count jumps, and the first jump is in a place nobody looks.
The interval between usable copies is one minute. Before the control below is moved: how many of the seven services breach their recovery point objective?
The copy interval, run against all seven objectives at once
One control sets the gap between one usable copy of the record and the next, and the gap runs from nothing at all out to eight hours. One consequence follows: the number of named services that would end up losing more recorded work than the board allowed. The worst case is used throughout, meaning a failure landing immediately before the next copy. The default is set to 15 minutes. At that setting 3 of the 7 are breached: S1 payments and remittances, S7 treasury settlement and S3 internet and mobile banking. The other 4 are inside, with S2 branch counter service and cash withdrawal sitting exactly on its line at 15 minutes, and S4, S6 and S5 with room to spare. The very first stop on the control shows what happens between zero and one minute.
Copy interval: 15 minutes, from 0 up to 8 hours
With a copy interval of 15 minutes, 3 of the seven services would lose more recorded work than the board said they could, being S1 payments and remittances, S7 treasury settlement and S3 internet and mobile banking. S2 branch counter service and cash withdrawal sits exactly on its line at 15 minutes and is the next one to go.
Educational illustration. All seven recovery point objectives below were chosen by this bank's own board. The worst case is used throughout, meaning the failure is placed immediately before the next copy, so the work lost is taken as the whole interval. The horizontal scale is stretched at the left so the early steps are visible.
Who sets a recovery point objective, and what are they answering?
The figure has to come from the business, and the reason is not a governance nicety. The question behind the figure can only be answered by somebody who knows what the work is. The single question is whether the work that disappeared exists anywhere else, and can therefore be put back.
Everybody has done this arithmetic without naming it. Consider a household for a moment. An electricity bill receipt goes missing. Annoying, and entirely recoverable. The electricity office has its own record and a phone call fixes it. Cash handed to a cousin last Diwali, with nothing written down anywhere by anybody, is a different case: if both sides forget the amount, it is gone, and no amount of searching produces it. Same household, same week, two completely different exposures, and the difference is not about how valuable the item was. The difference is about where else the item exists.
Run that question across the seven services and the objectives fall out of it. S5 deposit account opening carries four hours because the application forms are physically sitting on the counter, so four hours of lost entries costs some re-keying by the branch staff and nothing else. S3 internet and mobile banking carries five minutes because a customer who moved money and watched it confirm cannot sensibly be asked to remember what they did and do it again. S1 payments and remittances and S7 treasury settlement carry zero because an instruction the bank accepted and then lost exists nowhere at all.
A new service arrives with an instruction to set its recovery point objective. What is the one question actually being answered?
The failure: an objective written by whoever runs the technology
An objective written that way arrives quietly, and it never looks like a mistake at the time. Somebody has to put a figure in the plan, and nobody from the business is in the room. The person who is there answers the only question available to them: what the arrangement can currently manage. The number that comes out describes the equipment. The tell is that every service ends up carrying the same one.
Watch what that costs, and watch that it costs in two directions at once. Take a single objective of 15 minutes applied across all seven services at Vindhya Commercial Bank Limited. Against S1 payments and remittances and S7 treasury settlement, whose own objective is zero, fifteen minutes breaks the promise outright: work that exists nowhere else is allowed to disappear. Against S3 internet and mobile banking at 5 minutes, it is three times too loose. Against S4 loan disbursal and S6 trade finance issuance at 1 hour it is four times tighter than the business ever asked for, and against S5 deposit account opening at 4 hours it is sixteen times tighter, all of it bought and paid for to protect paperwork that is still lying on the counter.
So one figure leaves three services with a promise the bank cannot keep, three services protected far beyond anything anybody asked for, and exactly one correct. Nobody in that story was careless. The outcome is simply what happens when a business question is answered by looking at machinery, and the same fact is why the seven honest answers here run all the way from zero to four hours.
A recovery plan gives every service in it the same recovery point objective. What does that say about how the figure was set?
Why is the technology being back not the same as the service being back?
Here is the stage that goes missing from almost every plan, and it is the stage that decides what the customer experiences. The record has been put back. The record is complete up to the moment of the last usable copy and contains nothing after that. The service is technically available. And every bit of work that happened inside the gap is still gone.
Getting that work back is reconciliationWorking out, after a restore, what is missing from the record and having it re-entered or re-checked by people., and it is done by people, not by technology. Somebody has to establish what fell into the gap. Establishing it usually means comparing whatever partial evidence exists: paper at the counter, a queue of instructions that never completed, a customer on the phone insisting a transfer was made. Then somebody has to key it back in, and somebody else has to check it. Reconciliation is the stage nobody is assigned by default, and it starts after the technology report says the job is finished.
Notice how the shopkeeper's evening goes after the notebook is ruined. Getting the shutter open again takes an hour. Walking the street asking eleven customers what they bought last week takes three days, and two of them remember it differently. The first job is the restore; the second job is the reconciliation; and only the second one decides whether the shop actually knows what it is owed.
There is a second event on the other side of all this, and it deserves a name because plans forget it. FailbackReturning from the standby arrangement to the normal one once the failure is fixed, which is its own event rather than the absence of one. is the move back from the standby arrangement to the normal one. Failback is a change like any other change, with its own chance of going wrong. A plan that treats the return journey as automatic has planned half a round trip.
The technology has been restored and the service is technically available. What still has to happen before the service is actually delivering?
What does a restore test on a planned date actually establish?
In month 12 Vindhya Commercial Bank Limited ran a continuity test across all seven important business services. The test produced seven results. S1 came back in 52 minutes against a 60 minute objective and met it. S7 came back in 48 minutes against 60 and met it. S4 came back in 6 hours against 8 and met it, S5 in 9 hours against 12 and met it. Three missed: S2 in 3 hours 40 minutes against 3 hours, S3 in 3 hours 10 minutes against 2 hours, and S6 in 14 hours against 8 hours. Four met and three missed.
Now look at what that list is made of. Every one of the seven results is a time. The month 12 test measured the recovery time objective on all seven services and measured nothing whatever about data, so the recovery point objectives at this bank have never been tested at all. No plausible figure fills that gap. The absence is the finding: half the design was examined, the report was headed with the result of a continuity test, and nobody reading it would know which half.
There are two more things the test could not establish, and they matter as much. No work was actually lost, so none had to be put back, and the test could not show how long the reconciliation would have taken. And it could not show what happens on a day nobody chose. The month 12 test ran on a planned date, with the recovery team on standby, and the failure chosen in advance, so it tested the arrangement and not the surprise. That limitation belongs beside the result every time the result is quoted, or the result quietly claims more than it earned.
The month 12 test measured how long every one of the seven services took to come back. What did it not measure?
What changes when somebody else runs the technology?
Plenty of what a bank runs on is run by somebody else, and the moment that is true a reader tends to assume the objectives move across with it. They do not. Handing the technology to a third party relocates the arrangement. The promise stays put on the bank's own side of the table.
Incident I9 at Vindhya Commercial Bank Limited is the clean instance. A payment gateway run by an outside supplier failed for 9 hours and 48,000 transactions failed with it. Three questions have three answers, and only one of them changes. Who does the restoring? The supplier. Who is measured against the recovery point objective? The bank. Who does the customer deal with, before, during and after? The bank, throughout. The customer never had any relationship with the supplier and in most cases did not know one existed.
Money did move, and it is worth being precise about what that means. Incident I9 booked a gross loss of Rs 4.4 crore and a recovery of Rs 1.2 crore came from the supplier, leaving Rs 3.2 crore net. So the contract did some work. No contract takes the promise off the bank. Whether the supplier can actually deliver what the bank told its customers is a question about the contract and the monitoring behind it, and that whole subject is covered separately.
In incident I9 the payment gateway was run by an outside supplier and it failed for 9 hours. Whose recovery point objective applied to it?
What binds an Indian bank here, and where to check it
Restoring a record, finding the gap and putting the missing work back is the same job wherever the institution sits, so the mechanism above is universal. A supervisor's requirements on a regulated bank are not universal: whether recovery arrangements have to exist, how they have to be tested and reported, what is expected when the technology sits with a supplier, and where records may be kept. For an Indian bank the requirement is issued by the Reserve Bank of India, whose site is rbi.org.in, and the international standard behind it belongs to the Basel Committee, published by the Bank for International Settlements at bis.org. An Indian bank is actually held to the Indian rule, so name the standard first and the Indian rule second. Every requirement, objective, testing frequency, location rule and effective date must be confirmed at the source before it is relied on.
What a corporate treasurer asks a bank, once they know this
Girish Talwalkar is group treasurer at Nirjhar Industries Limited, invented, and he is deciding where his operating balances and his payment runs will sit. The question he used to ask was how quickly the bank could recover from an outage. He can live with a slow morning, and that is the easier half of the question. The question he asks after reading this is different: if the payment system fails at eleven in the morning, how many of the instructions already given to the bank might not exist when it comes back, and has anybody ever measured that rather than estimated it?
The same pair of questions does real work for anybody reading an institution from the outside. A supervisor, an internal auditor, a board member on a risk committee: all three can ask to see the two objectives side by side for each named service, and then ask which of the two the last test actually measured. At Vindhya Commercial Bank Limited the honest answer would be that the last test measured recovery times on all seven services and measured lost work on none of them. The absence is a finding on its own and does not need a number to be worth reporting.
One more use, and it is the one that stops a report misleading its reader. Take incident I3, where core banking sat unavailable for 4 hours and 20 minutes. Its net cost was Rs 3.2 crore, placing it joint fourth largest on a net ranking of the year, level with incident I9, and seventh on a gross one. Whichever way it is ranked, a loss figure describes what the bank paid. A loss figure says nothing about how long the outside world went without the service and nothing at all about how much recorded work came back afterwards. Three different measurements, three different reports, and most institutions only ever write the first one.
Covered separately. What operational resilience is, what an important business service is and how an impact tolerance gets set are covered separately, and so is the argument for why the promise and the design are different lines. Keeping a service running by another route while the technology is still down is a business arrangement rather than a technology one, and it is covered separately; the two are planned together and they are not the same job. Declaring and running a crisis, and the detect, contain, recover and learn cycle that runs on every incident, are each covered separately. How risk is managed when an arrangement sits with a supplier, including what belongs in the contract and how it is monitored, is covered separately, and the promise does not move when the technology does.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated bank covering business continuity and recovery arrangements, the outsourcing of technology, and where records are held | rbi.org.in |
| Bank for International Settlements | The Basel Committee's international standard on operational resilience, which the Indian requirement implements | bis.org |
| Securities and Exchange Board of India | Expectations where the entity is a market intermediary rather than a bank | sebi.gov.in |
| Indian Banks Association | Material on Indian banking operational convention | iba.org.in |
Vindhya Commercial Bank Limited, Nirjhar Industries Limited and Girish Talwalkar are invented.
Educational material. Not advice on any investment, tax, budget or market position.
