How to assess Third Party Risk, From Onboarding to Exit
Assess it in eight steps and in this order: build the inventory, classify by criticality, assess before contracting, contract for the four terms that carry the risk, onboard and connect, monitor in life against what the arrangement actually does, test the exit, then report. Step one is the one most institutions have not done, and nothing after it works without it.
Start with the ordinary version. The bank version is the same shape at a larger scale. A shop that takes card payments does not run the card machine. Somebody else runs it, somebody the shopkeeper has never met maintains the line it sits on, and on the afternoon it stops working the shop cannot sell anything. Nobody in that shop was careless. The shopkeeper simply bought a service instead of building one. Everybody does that, and in buying it the shopkeeper took on a failure they cannot fix and did not price.
Vindhya Commercial Bank Limited, invented, is that shop at a balance sheet of Rs 96,000 crore. In month 9 a payment gateway hosted by an outside provider stopped working for 9 hours and 48,000 transactions failed. The provider is not named, and the record does not say what caused the failure. The record does state the cost: Rs 4.4 crore gross, Rs 1.2 crore recovered from the provider under the contract, and Rs 3.2 crore net, booked as incident I9. Everything an assessment method is for either happened before that one afternoon or did not happen at all.
In what order should a third party arrangement be assessed?
A third partySomebody the institution contracted with to perform an activity it would otherwise do itself. is anybody the institution contracted with to do something it would otherwise have done itself. The wide definition is deliberate. Every narrow version fails in the same way: an institution that defines third parties as the ones with a technology contract will not have the printer that produces its statements on the list, and an institution that defines them by department will have a different list in every department.
Everything in this method rests on one fact about timing, and holding it makes the order obvious. An arrangement is assessed at two moments, and the second one is far harder than the first. Before contracting there is leverage and no information: nothing has run yet, nobody can say how the service behaves under load, and the institution can still walk away without cost. Afterwards there is information and no leverage: the service is live, customers are on it, the alternative is expensive and slow, and every question has stopped being a condition of the deal and become a request. The eight steps are arranged the way they are so that as much of the assessment as possible happens on the side of the line where the institution can still say no.
What has to exist before any single arrangement can be assessed at all?
Step TP1 is the inventoryThe list of every such arrangement, which has to exist before any of them can be classified or assessed., and it is skipped more often than any other step in this method. The inventory is a list of every arrangement, with the activity it performs, the part of the institution that bought it, and one named person accountable for it. Nothing more elaborate than that, and it is still the step most institutions have not finished.
Here is why it has to come first, and the reason is mechanical rather than tidy-minded. Every step after it operates on a list. Classification operates on a list. Assessment before contracting operates on the classification. Monitoring operates on what classification said mattered. So an institution without an inventory does not skip one step. Such an institution performs all seven of the others on whatever subset it happens to remember. The finance system makes large invoices visible, so the remembered subset is the arrangements with large invoices.
There is a second thing only an inventory can do, and no single contract negotiation can ever reach it. Two separate parts of the institution can buy from the same provider without either knowing, and three separate providers can turn out to sit on the same underlying service. Both of those are concentration in a supplierThe exposure created when many arrangements or many institutions depend on the same provider, which no single contract can reduce and which only the list can reveal.. Each arrangement on its own looks fine, so no individual assessment can see the concentration. The household version is a street of ten shops in one market, all of whom think they have their own electrician, and all of whom have the same one. Concentration is a property of the list and not of any arrangement on it, so an institution with no list cannot see it at all, however carefully it assesses each contract.
Why does nothing in this method work until the inventory exists?
How is criticality decided, and against what?
Step TP2 is classification, and criticalityHow much depends on the arrangement, judged against the services the institution has decided matter, and never against the size of the contract. is how much depends on the arrangement. Criticality is judged against the services the institution has already decided matter, a list settled under important business services: how much stops when this arrangement stops, for whom it stops, and how quickly the stopping starts to hurt. None of the three questions has anything to do with money changing hands under the contract.
The commonest error in the whole method lives here, and it is not stupidity, it is availability. The institution grades what it can see, and what it can see is spending. So the arrangement with the biggest invoice gets the deepest assessment, the annual review, the senior sponsor and the escalation route, and the arrangement with the small invoice gets a form. Run that against this bank's own record and the order comes out backwards. 48,000 failed transactions over 9 hours is 5,333 an hour, so the gateway that failed in month 9 was carrying 5,333 transactions an hour. Nothing about that rate is visible in a purchase order.
Two arrangements: one costs ten times as much as the other. Which is more critical?
What can be assessed before contracting, and what only afterwards?
Step TP3 is the pre-contract assessment, and the useful way to think about it is as a race against the loss of leverage drawn in the first figure. Before signing, the answer to a refusal is that the deal does not happen, so the institution can ask for anything. The institution can ask who else the provider serves, whether it depends on anybody the institution also depends on, how it is financed, what its own failure record looks like, where the work is physically done, and who would actually be doing it.
No pre-contract assessment can see the service run. The institution's volumes have never been on the service, so nobody can say how the arrangement behaves on its worst day. So a pre-contract assessment is always an assessment of the provider and never of the service, and the honest way to hold that is to write down what the assessment could not see. The whole point of the contract step that follows is to buy the right to find out later the things that could not be found out now.
Which contract terms are the ones that actually carry the risk?
A contract for a bought service has a great many terms and most of them are about money, timing and blame. Four of them are about risk, and they are the four that decide whether anything in steps five to eight is possible at all.
Take them one at a time. The right to examineA contract term letting the institution look at how the other side actually operates, without which monitoring is a set of assertions. is the term that converts monitoring from a conversation into evidence, and it is worth noticing that it is the only one of the four whose value shows up in a different step entirely. An institution that cannot get its data back has no exit whatever the exit clause says, so the data term settles where the data sits, who may see it, and above all what happens to it at the end. The sub-contractingThe term governing whether and to whom the other side may pass the work on, without which the institution does not know who is actually doing it. term matters because an arrangement the institution assessed and an arrangement somebody else is now performing are not the same arrangement, and without this term nobody is told. And the exit arrangementThe agreed route for taking the activity back or moving it, whose only real test is having exercised it. is how this ends, on whose notice, in what time, and with what help from the side that is losing the revenue.
Which four contract terms carry the risk?
The fifth term, and it is the one everybody argues about
There is a fifth term that takes up more negotiating time than the other four together, and it is the recovery basisWhether the other side bears a share of a loss or a fixed maximum, which are different contracts that look identical on a small loss.: what the provider pays when its failure costs the institution money. Two shapes are common. One is a share of the loss. The other is a fixed maximum, whatever the loss turns out to be. The two are different contracts, and on a small loss they are impossible to tell apart.
Watch it happen on this bank's own numbers. Incident I9 cost Rs 4.4 crore gross and Rs 1.2 crore came back from the provider. Rs 1.2 crore is 27.3 per cent of Rs 4.4 crore. The recovery is exactly what a 27.3 per cent share would have paid, and exactly what a fixed maximum of Rs 1.2 crore would have paid. The record does not say which contract this was, and on this loss nothing distinguishes them. Now make the failure bigger. At the length that takes this bank's year to its own loss limit, worked out below, the gross loss is about Rs 26.7 crore. A share leaves Rs 19.4 crore of net loss with the bank. A fixed maximum leaves Rs 25.5 crore. A term that makes no difference at all on the loss that happened makes a difference of Rs 6.1 crore on the loss that could.
The contract recovered Rs 1.2 crore of a Rs 4.4 crore loss. Is that a share of the loss or a fixed maximum, and does it matter here?
What changes at the moment the arrangement goes live?
Step TP5 is onboarding, and it is short because most of what it does is plumbing. Access is granted. Connections are opened. People on both sides are named. The reason it earns a step of its own is that this is the moment the assessment stops being about a company and starts being about a running service, and the institution's own record is about to start producing evidence it will either read or ignore.
Two things belong here and are forgotten with great regularity. The first is that whoever is named as accountable inside the institution has to be a person and not a department, for the same reason a risk owner is a person everywhere else in this subject area. The second is that access granted at onboarding has to be recorded somewhere that a later review will actually look at. An arrangement that goes live without a named accountable person on the institution's side has no reader for any of the evidence steps six to eight will produce.
What makes monitoring real rather than an annual questionnaire?
Monitoring in lifeWatching what the arrangement actually does using data it produces anyway, as distinct from asking the other side questions once a year. is step TP6, and it is where most institutions substitute an activity that looks like monitoring for monitoring. The substitute is the annual questionnaire: a set of questions sent to the provider, answered by the provider, filed by the institution, and read by nobody until something goes wrong. The questionnaire is not worthless. A completed questionnaire is evidence of what the other side says, and that is worth having. No questionnaire is evidence of what the arrangement does.
The alternative is not more questions. Monitoring runs instead on the operational data the arrangement produces anyway: how many transactions it handled, how many failed, how long it took, how often it stopped, how quickly it came back, and how many complaints followed. The operational record exists every day whether anybody reads it or not. The household version is the difference between asking the landlord whether the roof leaks and looking at the ceiling after it rains. Only the second kind of monitoring can contradict the first, and a monitoring arrangement that cannot contradict the provider is not monitoring anything.
The institution sends an annual questionnaire and gets it back completed. What does it now know?
Why does the exit have to be tested, and what does an untested exit cost?
Step TP7 is the one institutions defer, and they defer it for a completely rational reason: testing an exit is disruptive, costs real money, annoys the provider, and produces nothing anybody can put in a report except a list of what went wrong. So the exit stays on paper, and everybody involved believes it works.
An untested exit holds an assumption with three parts, and none of them has been checked. How long would the move take. Who exactly would do it, on both sides, with the ordinary work still running. What else depends on this arrangement that nobody has thought of. Every one of the three is a belief until somebody exercises it. The first exercise of an untested exit happens on the day the arrangement has already failed, and that is the worst possible day to discover that the answer to the second question is a person who left last year. Testing an exit does not make the exit work; it converts a belief about the exit into a measurement of it, and that conversion is the only thing that turns a document into a capability.
The exit route is documented and has never been used. What is actually held?
What goes up, to whom, and what does this record not say?
Step TP8 is reporting, and its content follows from the seven steps above: the number of arrangements, the number graded critical, any of the four contract terms missing on any of them, what the operational record showed, the exits exercised and when, and any incident that happened at a provider. The list is short, and it fits on a single sheet.
The record runs out at exactly this step, and the silence is itself part of the answer. Vindhya Commercial Bank Limited has a policy numbered PL8 for outsourcing and third parties. The record names no committee as its recipient. The record also gives no count of third parties, no criticality grading of any of them, and no cause for the failure in month 9. Four silences. The correct response to a record that does not say who receives the reporting is to report that as a gap, not to assume a committee that the record does not name.
Which committee at this bank receives third party and outsourcing reporting?
What does a failure at a third party actually cost, hour by hour?
Everything above is method. The arithmetic below is what makes the method worth doing, and it is built entirely from figures the record already carries. Incident I9, month 9, of the seven event categories the Basel Committee publishes it sits in the sixth, business disruption and system failures. A payment gateway hosted by an unnamed outside provider failed for 9 hours, 48,000 transactions failed, the gross loss was Rs 4.4 crore, Rs 1.2 crore was recovered from the provider and Rs 3.2 crore was the net loss.
One event is an anecdote. Four divisions turn it into rates, and rates are the only thing anybody can reason with about a failure that has not happened yet. 48,000 over 9 gives 5,333 transactions an hour. Rs 4.4 crore over 9 gives Rs 0.4889 crore of gross loss an hour. Rs 4.4 crore over 48,000 gives Rs 917 per failed transaction. And at the recovery actually achieved, Rs 3.2 crore over 9 gives Rs 0.3556 crore of net loss an hour. Every one of those four is a division on two figures already sitting in the record, and none of them is a new fact about anything.
Now run the length out and watch where it goes. At 1 hour the failure costs 5,333 transactions and Rs 0.49 crore gross. At 2 hours, 10,667 and Rs 0.98 crore. At 9 hours, 48,000 and Rs 4.4 crore, the outage that actually happened. At 12 hours, 64,000 and Rs 5.87 crore. At 24 hours, 1,28,000 and Rs 11.73 crore. Two lengths further out are worth marking. At about 45.8 hours the gross loss reaches Rs 22.4 crore. Rs 22.4 crore is the gross loss of incident I13, the letter of credit fraud and the largest single loss of this bank's year: a provider outage has to run nearly two days to match a fourteen month fraud on a gross basis.
The second crossing is the one that changes how an arrangement is graded. The bank's year of operational loss was Rs 43.8 crore net across thirteen incidents, and its own limit L11 caps that at Rs 60.0 crore over a rolling twelve months. Take incident I9 out and the rest of the year is Rs 40.6 crore. Reaching the limit therefore needs Rs 19.4 crore more of net loss, and at Rs 0.3556 crore an hour that is about 54.6 hours. One arrangement, run by somebody else, failing for about two and a quarter days, takes this bank's entire annual operational loss to its own limit. The 54.6 hour figure is what step TP2 exists to produce and step TP7 exists to make survivable, and neither step is about the contract at all.
Did this nine hour outage breach any of the institution's service tolerances?
This outage ran 9 hours and cost Rs 3.2 crore net. Before the control below is moved: how long would it have to run to take the bank's whole year to its own Rs 60.0 crore limit?
Run the outage longer and watch the year walk up to the limit
One control: the length of the failure in hours, from 0 to 60. Three things move together: how many transactions fail, what the outage costs gross, and where the bank's whole year of operational loss ends up with the other twelve incidents held still at Rs 40.6 crore. The default is 9 hours, the outage that actually happened: 48,000 failed transactions, Rs 4.4 crore gross, Rs 1.2 crore recovered from the unnamed provider and Rs 3.2 crore net, inside a year of Rs 43.8 crore at 73.0 per cent of limit L11.
An outage of 9.0 hours fails 48,000 transactions and costs Rs 4.4 crore gross, taking this bank's year to Rs 43.8 crore, being 73.0 per cent of its own limit L11.
Four failures, and the first one causes the other three
The first failure is that step TP1 never finished, and everything downstream inherits it. An institution that cannot list its arrangements cannot classify them, so it classifies the ones it happens to know about, and those are the ones with large invoices.
The second follows immediately and is the commonest of the four: criticality graded by contract value rather than by what depends on the arrangement. A gateway can be a modest contract carrying 5,333 transactions an hour, and a large contract can carry nothing that matters for an afternoon, so grading by invoice size sorts the list in almost exactly the wrong order.
The third is monitoring as a questionnaire. An annual set of questions answered by the other side tells the institution what the other side says. Two things would tell it what is actually happening. The right to examine is settled at step TP4 and cannot be obtained later. The operational data the arrangement itself produces arrives every day and is usually unread.
The fourth is the untested exit, and it is the expensive one. An exit that has never been exercised is an assumption about how long something would take and who would do it, and it is discovered to be wrong at exactly the moment it is needed. None of the four is carelessness. Each one is a reasonable local decision, and each looks wrong only from the list, so the list has to exist first.
Who actually uses this method, and what do they do with it?
Three different readers pick this up for three different things. The head of operational risk inside the institution, Purnima Ganeshan in this invented case, uses it as a queue: the inventory tells her how many arrangements there are, criticality tells her which ten to look at this quarter, and the exit test schedule tells her which of those ten she still cannot say anything honest about. She is not trying to assess everything. She is trying to make sure the order of assessment is not set by invoice size.
A credit analyst at another institution, looking at Vindhya Commercial Bank Limited as a counterparty rather than as an employer, uses it differently. The analyst cannot see the inventory and would not be given it. The analyst can ask one question with a very informative answer: has any exit ever been exercised, and what did it take. An institution that has exercised one can state a duration. An institution that has not will talk about its policy. The gap between those two answers is the whole of what this method delivers, and it is audible in one question.
Take the household version. The mechanism does not change with scale. A person whose salary, savings and borrowing all sit at one place has one arrangement carrying everything, and the size of the fees tells them nothing about that. A caterer whose only refrigeration is a rented unit has one supplier who can end a wedding. In both cases the useful question is not whether the provider is reputable. The useful question is what stops, how fast it starts to hurt, and whether anybody has ever tried the alternative even once.
What is named here, and where the binding version lives
No authority prescribes eight steps, or this order, or these four contract terms as a set, and the labels TP1 to TP8 are a numbering convention rather than anything a regulator publishes. Every figure, incident, policy and limit belongs to Vindhya Commercial Bank Limited and to nothing else.
One international standard sits behind the classification used above: the operational risk framework published by the Basel Committee on Banking Supervision at the Bank for International Settlements, bis.org. The seven event categories that place the month 9 failure in business disruption and system failures come from that framework. The Basel framework is the origin of the classification and not what binds anybody.
The rules that actually bind an Indian bank on outsourcing, on arrangements with third parties and on technology service providers come from the Reserve Bank of India at rbi.org.in. Where the institution is a market intermediary rather than a bank, the equivalent expectations come from the Securities and Exchange Board of India at sebi.gov.in.
The wording of these requirements moves, and only the text the regulator itself has published binds anybody.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | What actually binds a bank in India on outsourcing, arrangements with third parties and technology service providers | rbi.org.in |
| Bank for International Settlements | The Basel Committee operational risk framework and the seven event categories used to classify a loss | bis.org |
| Securities and Exchange Board of India | The equivalent expectations where the regulated institution is a market intermediary rather than a bank | sebi.gov.in |
Vindhya Commercial Bank Limited, Purnima Ganeshan and the unnamed provider behind the failed payment gateway are invented, along with every incident, policy, limit and figure attributed to them.
Educational material. Not advice on any investment, tax, budget or market position.
