Crisis Management: Deciding Under Pressure With Incomplete Information
Crisis management is deciding and acting while the facts are still missing. The difference between a crisis and an incident is not the size of the loss but whether the decisions that cannot wait have outrun the information. A named person with the authority declares one, a standing team takes the decisions the ordinary process cannot take fast enough, and every decision is written down with the assumptions it rested on.
The failure that matters most here is not a failure of capability. An institution can hold every arrangement it needs, staffed and tested and paid for, and still run its longest failure of the year with nobody in charge. The document that describes all of it never said who is allowed to start it. A crisis plan is tested at its weakest sentence, and its weakest sentence is usually the one that begins it. Vindhya Commercial Bank Limited, an invented bank, carries the worked example throughout.
What makes something a crisis rather than a bad incident?
Most failures are handled by people doing their ordinary jobs faster. A payment goes out twice, somebody notices, the recall goes to the other bank, the reconciliation team stays late. Nobody declares anything, and nobody needs to. The decisions that came up were decisions those people were already authorised to take, and the facts they needed arrived quickly enough to take them.
A crisisA situation where the decisions that cannot wait have outrun the facts available to take them. is what remains when that stops being true. The decisions arrive before the information does. Somebody has to choose whether to switch a service to its other route while nobody yet knows what broke, or has to decide what to tell forty thousand people who are already asking, or has to decide whether to stop a whole line of business on a suspicion. The line between an incident and a crisis is drawn on the relationship between decision speed and information speed, not on the size of the loss.
A wedding shows the same line. The caterer's van gets a flat tyre on the way. A flat tyre is an incident. Somebody rings the driver, sends a car, the food arrives forty minutes late and nobody at the mandap ever knows. Now suppose the hall floods an hour before the guests arrive. Four hundred people are already in cars. Nobody yet knows whether the water will drain, whether the electrics are safe, or whether the hall will refund anything. And the decision whether to move the whole event has to be taken in the next ten minutes. Nobody has more information than they did in the first case; what changed is that the decision will not wait for it.
The size of the loss is a poor test for exactly this reason, and the invented bank's own record shows it. Incident I13, the internal fraud discovered in month 8 in which 9 letters of credit were issued against forged shipping documents over fourteen months, booked a net loss of Rs 15.4 crore and was declared a crisis. Incident I9, the vendor-hosted payment gateway that failed for 9 hours in month 9 and failed 48,000 transactions, booked a net loss of Rs 3.2 crore and was not. Incident I2, the settlement instruction that sent Rs 42.0 crore out of the bank twice in month 2, carried the largest gross loss of the whole year and never came close to being called a crisis. The recall worked and Rs 41.4 crore came back.
Incident I13 cost Rs 15.4 crore net and was declared a crisis. Incident I9 cost Rs 3.2 crore net and was not. Does that ordering make sense?
How can anyone tell, at three in the morning, which one it is?
Whatever test is written has to work in the worst possible conditions. Those are the only conditions in which anyone will ever use it. The person applying it will be half asleep, holding a phone, with one fact and three rumours. A test written as a paragraph of description will not survive that. A usable threshold test is three questions with yes or no answers and nothing in between.
The three questions come straight out of the definition. Can the decisions that are waiting be taken by the people who are normally authorised to take them? Can they be taken inside the time actually available? Can they be taken on the information available now? Three yes answers and the situation is an incident, however unpleasant. One no answer anywhere and somebody has to declare.
Notice what is deliberately absent. The loss figure plays no part. Nobody knows it yet and will not know it for weeks. The cause plays no part either. The cause is precisely what is missing. A judgement about how serious the failure feels plays no part. Two tired people will judge that differently, and the whole point of a test is that they do not have to agree first.
Who is allowed to say the word?
Here is where most of the interesting failure lives, and it is not where people look. Everybody who writes a crisis plan thinks hard about what the team should do once it is together. Very few write down who is allowed to tell them to start.
A declarationThe act of somebody with named authority saying that a crisis has started, which is what causes the different arrangement to begin. is not an opinion and it is not a conclusion a group reaches together. A declaration is an act performed by a named person who holds the authority to perform it. The distinction between an act and an opinion sounds legalistic until the alternative is seen in action. Six competent people on a call at half past six on a Saturday morning, all of them worried, none of them able to point at the line in the document that says they may start the arrangement, will spend forty minutes agreeing that this is serious, and not one of those minutes will produce a decision.
So the authority attaches to a role rather than to a person, it has a deputy, the deputy has a deputy, and every hour of the week is covered by one of them. The last clause is the one that gets skipped. The week is very much bigger than the working week, and nobody writing a plan on a Tuesday afternoon quite feels that.
Here is the arithmetic. The case does not record what hours this bank works, so the numbers are illustrative rather than the invented bank's own. A five day week of eight hour days is 40 hours out of the 168 hours in a week, or 23.8 per cent of it. The other 128 hours, being 76.2 per cent of the week, are evenings, nights, weekends and holidays. An authority that exists only in working hours is absent for roughly three quarters of the time in which a failure can start. Incident I9 started on a Saturday. On those numbers that is not bad luck at all.
A plan names a crisis team of seven roles and says the team convenes when a crisis is declared. What is missing?
What is the crisis team actually for?
Not for doing the work. The people restoring a service during a crisis are the same people who would have restored it anyway, and a team of senior people leaning over them makes that slower rather than faster. A crisis teamThe small standing group that takes the decisions during a crisis which the ordinary process cannot take fast enough. exists to take four decisions that the ordinary process cannot take quickly. In ordinary time each of the four sits somewhere else or nowhere at all.
The first is what to protect first. When several services are down at once, two of them cannot both be recovered first, and the choice between them is a business judgement about who outside is hurt worst, not a technical one. In ordinary time that judgement is spread across whoever runs each service, and every one of them believes theirs matters most. Believing so is not a criticism of any of them.
The second is which arrangement to invoke. Keeping a service running by another route and restoring the technology behind it are two different things, run by two different sets of people, and somebody has to choose which is started, in what order, and whether the cost of invoking one is worth paying before the cause is known. The third is what is said to people outside. The fourth is when it ends, and that decision is the one that in ordinary time belongs to nobody.
How many roles should sit in that team? The invented bank's plan named seven, and seven is a reasonable size for a group that has to reach a decision in ten minutes on a bad line. Which seven they were is not in the record. The record holds a more interesting fact instead: the plan named seven roles and named nobody who could convene them.
How is anything decided when the facts are not in yet?
Not by getting better facts. There are none to get. Decisions get taken by sorting them differently instead. Four questions do most of the work, numbered CD1 to CD4 so they can be named later. The four are asked in order, and the order matters more than any individual answer.
CD1 is reversibilityWhether a decision can be undone later if it turns out to have been taken on the wrong facts, and at what cost., and it is worth more than the other three together. Can this be undone if it turns out to be wrong? A decision that can be reversed costs the price of reversing it, and that price is usually small and always knowable. A decision that cannot be reversed costs whatever the wrong answer costs, and that cost is neither small nor knowable. So the reversible ones are taken now, fast, on thin information, and the scarce minutes go on arguing only about the ones that cannot be taken back.
CD2 is the cost of waiting. Does waiting cost more than being wrong? Sometimes it plainly does: every minute the payment service is down is another few hundred people who cannot pay a bill, and the cost of switching to the other route and switching back is an hour of somebody's evening. Sometimes it plainly does not, and the honest answer is to wait twenty minutes for the one fact that settles it. Asking the question out loud is what stops a room defaulting to whichever answer the loudest person prefers.
CD3 turns the argument around. Instead of asking what is true, ask what would have to be true for this to be the right decision, then ask whether anybody in the room believes that thing. Turning the argument around ends a circular discussion startlingly fast. People who cannot agree about the facts can often agree very quickly about which fact would settle it. CD4 is the one everybody forgets in the moment: who has to be told, either way? A decision taken correctly and communicated to nobody is worth about as much as one taken badly.
A service has to be switched to its other route or not, now, and the cause is not yet known. Which test applies first?
Why write the assumption down while the decision is still open?
Every decision taken in a crisis rests on something believed rather than known. Switching the route assumes the other route is healthy. Telling customers it will be back within the hour assumes the restore behaves the way it did in the test. Holding off on the switch assumes the fault will clear itself. Those beliefs are the load-bearing part of the decision, and they are invisible an hour later.
An assumption logThe running record of what was believed to be true at the moment each decision was taken, written at the time and not afterwards. is a running record of what was believed to be true at the moment each decision was taken. The log has four columns and is written by somebody whose only job during the crisis is to write it. The log has to be written at the time. A record written afterwards is written by somebody who now knows the answer.
The difference between the two logs is not small, and it is not about blame. A log written after the event explains the decision using facts nobody had when it was taken. Such a log makes a good decision look obvious and a bad one look careless, and teaches the next crisis nothing at all. A log written at the time lets somebody ask the only fair question: given what was on the table at that minute, was that a reasonable thing to do? A log written at the time also does something more immediately useful. When an assumption is written down, somebody can be sent to check it, and the moment it turns out to be false the decisions resting on it are already listed.
| Time | Decision taken | What is being assumed | Who is checking it |
|---|---|---|---|
| Hour 1 | Do not switch the route yet | The fault will clear without intervention | Named, with a time to report back |
| Hour 2 | Tell customers, no cause stated | More will be known within two hours | Named, with a time to report back |
| Hour 3 | Switch the route | The other route is healthy and has capacity | Named, with a time to report back |
The three rows above illustrate the shape of a log rather than reproducing a record. The invented bank kept no log at all for the failure at issue, and the absence is exactly the finding.
Why is the assumption written down at the moment the decision is taken rather than afterwards?
What can be said while what happened is still unknown?
Something has to be said, and quickly, and this is the part people get wrong out of a decent instinct. The instinct is to stay silent until the facts are in. The trouble is that everybody outside is already speaking. Forty eight thousand payments failed in the invented bank's incident I9. Tens of thousands of people were forming an explanation of their own within minutes, and none of those explanations came from the bank.
A holding statementWhat is said to people outside while the facts are still missing, kept short and issued on a fixed rhythm. solves this with three parts and no fourth. What is not working. What is being done. When the next update will come. The third part is the one everybody wants to leave out and the only one that stops the next round of questions. A person who knows when they will hear again stops asking now.
The parts that people want to add are the cause, a restoration time, and length. None of the three belongs there yet. The cause is not known, and in a technology failure the first guess is wrong often enough that publishing it costs more than silence would have. A restoration time is a decision the crisis team has not taken. And length reads as evasion when the reader wanted three short lines.
Two hours into a failure the cause is still unknown. What is said?
What did nine hours with nobody in charge actually look like?
Now the worked instance, and it is worth reading slowly because everything above turns up in it. Vindhya Commercial Bank Limited declared two crises in the twelve months. Incident I3, the core banking outage that ran 4 hours and 20 minutes in month 3, was one. Incident I13, the internal fraud discovered in month 8, was the other. Two declarations against thirteen recorded loss events, or 15.4 per cent of them.
Incident I9 was not one of the two. In month 9 a payment gateway run by an outside supplier failed for 9 hours and 48,000 transactions failed with it. The bank booked a gross loss of Rs 4.4 crore, recovered Rs 1.2 crore from the supplier and carried Rs 3.2 crore net. At 9 hours, incident I9 was the longest single service disruption recorded in the whole year, running 2.08 times as long as the declared incident I3 at 4 hours and 20 minutes.
Two things about that sentence need care, rather than letting the headline run unqualified. The first is scope. Incident I10, the collateral valuation feed that was stale for 11 working days in month 10, ran far longer in elapsed time; it was a wrong marking on 340 loans rather than a service the outside world could not reach, and no customer lost money. So the claim is about a single continuous service disruption and not about elapsed time. The second is the collision the case record itself flags: incident I3 and incident I9 both booked a net loss of Rs 3.2 crore, so the incident is named every time and never the bare figure.
Why was incident I9 not declared? Incident I9 began on a Saturday, and no single named person held the authority to declare a crisis outside working hours. The plan named a crisis team of 7 roles and did not name who could convene it. Read that with the sizes attached. The bank had a team, a set of roles, a call list and a set of actions, and it did not have a sentence saying who says the word. Everything expensive was in place and the one free thing was missing.
A nine hour failure begins on a Saturday and nobody can be reached who is authorised to declare. How much of that failure ran with no crisis declared?
The declaration dial: how long it takes to reach somebody who can say the word
One control: how many hours pass before a person with the authority to declare a crisis can actually be reached. One consequence: how much of a 9 hour disruption runs with no crisis declared and nobody holding the four decisions. The 9 hours belongs to the invented bank's incident I9, in which a payment gateway run by an outside supplier failed, 48,000 transactions failed with it, and the bank carried Rs 3.2 crore net after recovering Rs 1.2 crore from the supplier. The default is set to 9 hours, or 100.0 per cent, and that is incident I9 exactly as it happened. No authorised person was reached at all. The dial stops at 9 hours because the disruption itself ran 9.
Educational illustration. The 9 hour length of incident I9 is the invented bank's own record. The hours needed to reach an authorised person is a dial on the control and is not a figure from the case. The bar shows time with nobody in charge, and claims nothing about the loss. The loss was Rs 3.2 crore net either way.
The plan that was finished everywhere except at its first step
Read the invented bank's crisis plan in month 8 and it looks complete. Seven roles named. Responsibilities set out against each. An approach to communication. Actions written for several kinds of failure. Nobody reading that document would predict what happened four weeks later.
Then the gateway failed on a Saturday morning in month 9 and the plan never started. Starting it was the one thing nobody had been made responsible for. The cost is not the loss figure. The Rs 3.2 crore net on incident I9 would very likely have been similar either way. The cost is that for 9 hours the longest single service disruption of the year ran on ordinary weekend cover: no team convened, no single person holding the four decisions, no assumption log, and nothing said to anybody outside that had been decided by somebody senior.
And it is only visible at all because somebody wrote the reason down afterwards. The gap is worth pausing on. Nobody here was careless. The design did not name who could declare, so there was nothing for anybody to fail to do. Adding one named convening role, reachable at any hour, is one more line in a document that already named 7 roles.
What would have had to change for incident I9 to be declared a crisis?
How does a crisis end?
Badly, in most institutions. A crisis there does not end. It fades. The service comes back, people drift off the call, the working group carries on for a week doing something nobody has defined, and three days later nobody can say whether the arrangement is still running or who is holding what. A crisis that fades out instead of ending leaves nobody sure who is responsible for anything, and that is how the second failure gets handled worse than the first.
A stand-downThe decision that the crisis has ended and the ordinary process takes the work back, taken by somebody with the authority to take it. is a decision, taken by somebody with the authority to take it, exactly like the declaration that started it. Three things have to happen at it. Somebody declares it over. The ordinary process takes the work back explicitly, by name, rather than by drift. And the record of decisions and the assumptions behind them is closed and kept. Nothing else is inherited by the next crisis.
Which gives a crisis a shape in time with four marked moments, and it is worth seeing that three of the four are decisions somebody has to be authorised to take. The failure starts. Somebody declares. The team decides and communicates while the facts arrive around them. Somebody stands it down. Only the first of the four happens on its own.
What has to happen at a stand-down?
Who hears about a serious technology failure, and how fast?
The route a failure travels to the top looks like a governance question rather than a crisis one, and it is the reason the declaration exists at all. Vindhya Commercial Bank Limited runs eight committees, numbered G1 to G8. Committee G8, information security, has 6 members and meets quarterly. G8 reports into committee G6, operational risk management, a committee of 8 members meeting monthly. Committee G6 in turn feeds committee G1, the board, meeting 6 times a year. Committee G8 does not report to the board directly.
Count the meetings in that path. Four a year, then twelve a year, then six a year. A serious technology failure travelling the standing route passes through two committees before it reaches the top and neither of them is weekly, so the speed at which the board hears is set by the structure rather than by anybody's diligence. Nobody has to do anything wrong for it to be slow. None of that is a criticism of the structure. A committee that meets quarterly is meeting at a sensible rhythm for the work it usually does.
The structure does mean something specific, though. If the only route to the top runs through the standing calendar, then a failure that matters this week will be discussed next quarter. The declaration is the route that does not wait for a meeting. Bypassing the calendar is precisely what a declaration is for, and it is the reason a missing convener is a governance failure and not an operational one.
Committee G8 meets quarterly and reports into committee G6, which meets monthly. How fast does a serious technology failure reach the board by the standing route?
What the Indian rules require
The Reserve Bank of India at rbi.org.in is the source of what an Indian bank must actually report when a serious operational or technology incident occurs, including to whom, in what form and inside what time, and of what is expected where an arrangement is run by an outside supplier rather than in house. The Bank for International Settlements at bis.org is where the Basel Committee publishes the international standard those expectations implement. Where the entity is a market intermediary rather than a bank, the Securities and Exchange Board of India at sebi.gov.in applies instead.
Reporting windows, thresholds, forms and effective dates change, and whether a particular incident is reportable turns on the wording in force on the day rather than on how serious the failure felt. The issuing body's own site carries that wording.
Who reads this, and what they do with it
Rustom Batliwala, head of internal audit at the invented bank and reporting to committee G3, reads a crisis plan differently from the person who wrote it. He does not start at the actions. He turns to the opening sheet and looks for one sentence: who may declare, and who covers that role at every hour of the week. If it is not there, everything after it is a description of something that will never start, and he can write that finding without reading a line further.
Purnima Ganeshan, head of operational risk, has a different use for it. Committee G6 receives the loss record I1 to I13 and the near misses N1 to N5 every month. The pack does not naturally contain a count of how many of the year's events crossed the threshold test and how many were actually declared. Two declarations against thirteen recorded events is a count a governance body can act on, and producing it is the cheapest assurance available on this whole subject.
Girish Talwalkar, group treasurer at the invented Nirjhar Industries Limited, sits on the other side of the counter and cares about exactly one thing: when a service he depends on fails, does anybody senior at the bank actually take hold of it, and how quickly will he be told? He cannot read the bank's plan. He can watch what arrives during a failure. A first message with an update time in it is evidence that somebody is holding the decisions, and silence is evidence that nobody is.
And the household version, at a much smaller scale and in the same shape. An elderly parent falls at eleven at night. Three adults are in the house and everybody agrees this looks serious. How the next twenty minutes go is decided not by who knows first aid, but by whether anybody has ever said out loud who makes the call to go to hospital. Households that have settled that in advance move in four minutes. Households that have not spend twenty minutes agreeing that it looks serious, and those are the same twenty minutes the bank spent, in the same shape, for the same reason.
What can crisis management not do?
Quite a lot, and saying so is part of doing it honestly. A declaration does not fix anything. A declaration does not restore a service, find a cause or return a rupee. All it does is start a different way of taking decisions, and if the underlying arrangements are poor then a well run crisis simply produces well documented decisions about a service that is still down.
Naming a convenerThe named role that can call the crisis team together, at any hour, without waiting for anybody else to agree. would not have made incident I9 shorter or cheaper. The Rs 3.2 crore net was mostly the failed transactions and the remediation, and a supplier's system does not come back faster because a bank has convened a meeting. The honest claim is narrower and still worth a great deal: for 9 hours nobody was holding the four decisions, nobody was recording what was being assumed, and nothing said to anybody outside had been decided by somebody senior.
A plan also cannot anticipate the failure that actually arrives. Frank Knight's separation of the measurable from the unmeasurable, set out under operational resilience, is the reason: a scenario list is a list of the first kind, and the failure that turns up is regularly of the second. The threshold test is written about decision speed for exactly that reason, rather than about a list of scenarios. A list can be incomplete. Asking whether the decisions can wait works on a failure nobody imagined.
One last limit, and it is about the evidence rather than the practice. The reason anybody can say why no crisis was declared for incident I9 is that somebody at this invented bank wrote the reason down afterwards. Most institutions do not have that sentence. The same gap can then sit in a plan for years while every review of it passes. A missing line is not a finding until something starts on a Saturday.
A continuity test is no help on this question either. The month 12 test ran on a planned date, with the recovery team on standby, and the failure chosen in advance, so it tested the arrangement and not the surprise. A planned test says nothing about who could have been reached on a Saturday.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated bank covering the reporting of a serious operational or technology incident, business continuity and arrangements run by an outside supplier | rbi.org.in |
| Bank for International Settlements | The Basel Committee's international standard on operational resilience and operational risk management, and the loss event categories the invented record uses | bis.org |
| Securities and Exchange Board of India | Expectations where the entity is a market intermediary rather than a bank | sebi.gov.in |
| Indian Banks Association | Material on Indian banking operational convention | iba.org.in |
| Frank Knight | Risk, Uncertainty and Profit, 1921, the separation of a measurable risk from an unmeasurable uncertainty | econlib.org |
Vindhya Commercial Bank Limited, Nirjhar Industries Limited, Rustom Batliwala, Purnima Ganeshan and Girish Talwalkar are invented.
Educational material. Not advice on any investment, tax, budget or market position.
