Risk and Control Self Assessment: The Process and Its Honesty Problem
A risk and control self assessment asks the people who run a process to list its risks, list the controls against them, and say whether each one works. The answers come back optimistic for reasons that are structural rather than moral. One invented bank called 196 of 214 controls effective, at 91.6 per cent, where a second measurement of the same 214 found 172, at 80.4 per cent.
One feature of this process sits underneath everything else in this guide, and it is worth sitting with before any number appears. The people carrying out the assessment are the people whose work is being assessed. The overlap is not a flaw somebody forgot to design out. The overlap is the reason the process exists at all. Nobody outside a payments team knows what that team actually does at four o'clock on a Friday when a file is late and one checker has already left. Nobody outside a lending unit knows which step gets compressed when volumes double in a quarter. Knowledge of that kind is written down nowhere, and it cannot be bought, sampled or worked out from outside.
So the method goes and asks the only people who hold the knowledge. In doing that, it puts the knowledge and the interest in the same place. The person who can describe most precisely how a control behaves is also the person whose year gets harder if the answer turns out to be that it does not behave at all. Every step of a well built assessment is an attempt to get the first thing without being quietly steered by the second.
The process has a definition, an order its steps must run in, three separate judgements at the end of them, and one honesty problem underneath all of it. At one invented bank, a second and independent measurement of the same controls was later placed beside the business's own answer. The two did not agree. And the first thing almost everybody does with those two numbers is compare them wrongly.
What is a risk and control self assessment, and who actually does it?
The plainest version sits away from banking altogether. A household runs a kitchen. Somebody in that household could sit down one evening and do the whole thing in fifteen minutes: what could go wrong here, what is already done about it, and does what is done actually work? Gas left on. Answer: the last person out checks the knobs. Does that work? Honestly, most nights. Not on the night everybody left in a rush for a wedding. Knife within a toddler's reach. Answer: the block sits at the back of the counter. Does that work? It did until the child got taller.
The kitchen exercise is the entire method. A risk and control self assessmentA structured process in which the people who run a process list its risks, list the controls against them, and rate whether each control works. is a structured version of that fifteen minutes, run by the people inside a process rather than by anybody watching them. The assessment has three raw materials and it produces one output. The raw materials are the process, the risks that live in it, and the controls the process relies on. The output is a rated list.
The phrase that carries the weight is self assessment. In an institution the work is organised so that the people running a process manage the risk in it day to day, and a risk function sets the method, challenges the answers and keeps the record. The assessment belongs to the first of those. A risk function that fills in the ratings itself has produced a document about a process it does not run. The value of the exercise sits entirely in the fact that the people answering it are the people who were there.
Two things are where the process usually goes soft, and both are worth naming immediately. The first is that not everything that reduces risk is a control worth assessing. An institution picks the ones it says it relies on and calls those key controlsA control the institution has decided it relies on, as distinct from every activity that happens to reduce risk a little.. The knife block at the back of the counter is a key control. The general habit of being careful is not. Nobody can test it and nobody can fail it. The second is that the answers have to be written down against something. A conversation about whether the payments process is broadly fine is not an assessment; a list of named controls with a rating against each one is.
What are the steps, and in what order do they have to happen?
The steps are unglamorous and the order matters more than the steps do. Eight of them, and they group into three phases, each of which needs the phase before it to be finished first.
Phase one is description. Nothing is being judged yet. The output is a list of processes, a list of risks inside each, and a list of the controls the process depends on. Description is slow and boring. No other part of the assessment can be produced from outside, so this is the phase that carries the highest value. Phase two is rating, and it asks two different questions about each control, in that order. Phase three is consequence: a rating that is not good enough becomes an issue with a named owner and a date, and the whole assessment is signed by somebody who can be asked about it later.
Signing matters more than it sounds. An unsigned assessment is a document that exists; a signed one is a statement by a person. The difference shows up eighteen months later when somebody asks why a control everybody now knows was broken was rated as working. If a name sits at the bottom, that is a conversation. If no name sits at the bottom, it is an archaeology project. The step by step method for running the whole exercise, meaning who is in the room, what evidence they bring and how the workshop is chaired, is covered separately.
Is the assessment rating the risk, the control, or what is left over?
All three, and the commonest failure in the whole method is collapsing them into one. Three separate judgements are made, in order, and each answers a question the others do not.
The first is inherent riskThe risk in a process before any control is considered.: how bad would this be with nothing in place at all. Not how bad it is now. How bad it would be if the control simply did not exist. In the kitchen, the inherent risk of the knife block is a child reaching a blade. The second judgement is about the control itself: what does the process actually do about that, and does it work. The third is residual riskWhat is left after the controls are taken into account, which is a separate judgement from either of the other two.: given the control and how well it works, what remains.
Here is why the order cannot be shuffled. The value of a control is the distance between the first judgement and the third, so a process that never records the first has no way to show what any of its controls are worth. A workshop that opens with how bad is this really has answered the third question without ever writing down the first two. Everybody nods, the number goes in the file, and the following year nobody can say whether the risk fell because a control got better or because the person filling in the form was in a calmer mood.
There is a second reason to keep them apart, and it shows up in what a reader can do with the result. If inherent risk is recorded and residual risk is recorded, then a control that gets removed next year has a measurable consequence: the residual moves back toward the inherent. If only the residual was ever recorded, removing the control produces no arithmetic at all, and the argument for keeping it comes down to who speaks more confidently in the meeting.
Of the three judgements, one gets skipped far more often than the other two. Which one, and what is lost when it goes?
What does the control population at this bank actually look like?
Vindhya Commercial Bank Limited, an invented mid-sized Indian commercial bank with a balance sheet of Rs 96,000 crore, ran a risk and control self assessment across nine processes. The nine processes, numbered PR1 to PR9, are the scope of the exercise and they are the whole of it.
| Id | Process in scope |
|---|---|
| PR1 | Account opening and customer onboarding |
| PR2 | Lending and disbursal |
| PR3 | Collateral management and valuation |
| PR4 | Payments and settlement |
| PR5 | Trade finance |
| PR6 | Treasury dealing and settlement |
| PR7 | Deposit servicing |
| PR8 | Financial reporting and close |
| PR9 | Access management and information security |
Nine processes at one invented bank. The bank's record does not say how many controls sit in any one of them.
Across those nine processes sit 214 key controls. The population of 214 is the single most important number in the whole exercise, and it is the one most often left out of the summary. A result reported as ninety two per cent effective is unreadable without it. Ninety two per cent of two hundred and fourteen is a different quantity of work from ninety two per cent of thirty. The population comes first; the percentage is a description of the population and cannot be read on its own.
The gaps in the bank's record matter as much as its entries. The 214 are not split across PR1 to PR9. No record shows how many controls sit in trade finance, process PR5, or in access management, process PR9. The bank locked a total and never split it, so the honest answer to a question about any one process is that the measurement was not made.
How many of the 214 key controls sit in process PR5, trade finance?
What did the business actually produce at the end of it?
One rated list. The self assessment covered all 214 key controls and rated 196 of them effectiveA rating meaning the control did its job, and a word wide enough that two people can apply it to the same control and honestly disagree., leaving 18 that the business itself said were not working. As a rate, 196 over 214 is 91.6 per cent. As a piece of work, it is eighteen things somebody has to fix.
Notice how much better the second sentence is than the first. Ninety one point six per cent is a mood. Eighteen controls is a list with names on it, and a list can be sorted, assigned and closed. Every time a control population gets converted into a percentage, something that could be acted on is turned into something that can only be reported. Converting a list into a rate is the whole argument in miniature, and it returns with force once the second measurement is put beside the business's own answer.
What did a second and independent measurement of the same 214 reach?
Later in the same twelve months, the same 214 key controls were measured again, this time by people who do not run the processes. How that second measurement is planned, sampled, evidenced and reported is a craft in its own right, and it belongs to the controls and assurance material. Only its result matters here, stated as a fact about this invented bank.
The second measurement came in two passes, and the two passes asked different questions. The first pass looked at designWhether a control would achieve its objective if it happened every single time, judged from the arrangement itself rather than from what took place. across all 214 controls: would this control achieve its objective if it happened every single time. On that question 198 passed and 16 carried a design gap. The second pass looked at operating effectivenessWhether a control actually did happen every time over the period, which is a different question from design and needs a different kind of evidence., and it looked only at the 198 that had passed design. There is no point asking whether a badly designed control happened reliably. On that question 172 were effective and 26 were not.
The two sums close: 198 plus 16 is 214, and 172 plus 26 is 198. The two kinds of problem added together give 42 control findings, being 26 operating failures plus 16 design gaps. The 42 is a count of findings and nothing else; elsewhere in this bank's year the figure Rs 42.0 crore is the gross loss on one settlement incident, and the two have nothing to do with each other.
One limit has to be stated before anybody reasons further, and it is the kind of limit that gets skipped because it is inconvenient. The record does not say that the 172 controls the second measurement found effective sit inside the 196 the business rated effective. The overlap would be a natural assumption and it is not in evidence here. Some of the 172 may be controls the business itself had already called not effective. Neither set is a known subset of the other.
Why does putting 91.6 per cent beside 86.9 per cent give the wrong answer?
Here is what happens next in almost every institution. Somebody builds a slide. On the left, what the business said: 91.6 per cent. On the right, what the second measurement found: 172 of the 198 it tested, being 86.9 per cent. The gap looks like 4.7 points, close enough to be reassuring. Everyone agrees the business is a little optimistic, the paper is noted, and the meeting moves on.
The comparison is false, and it is false for a reason that has nothing to do with either number being wrong. Each figure is correct about the thing it measures. The problem is that they measure different things. The denominatorThe population a percentage is measured against, which is the single most common place a comparison between two rates goes wrong. under 91.6 per cent is 214. The denominator under 86.9 per cent is 198. Setting them side by side subtracts one rate from another rate that was never taken on the same population.
Look at what falls out of the second denominator. The 16 controls carrying a design gap never reached the operating test at all. There was nothing to test: a control that would not achieve its objective even if it ran perfectly does not need a sample. Dropping them from the denominator is arithmetically tidy and it has a consequence nobody intends. The 16 with a design gap failed earlier and harder than any of the 26, and taking them out of the base quietly credits the bank for the controls that failed first.
Put both measurements on the population the business actually rated, and the picture changes. On the full 214, the business said 196, being 91.6 per cent, and the second measurement reached 172, being 80.4 per cent. The like for like gap is 11.2 percentage points, not 4.7. The comfortable version understates the distance by 6.5 points, more than the comfortable version itself reports.
A reader who wants to rescue the flattering comparison usually tries one more move: perhaps some smaller, fairer population exists on which the two do agree. The arithmetic is worth doing rather than waving at. For 172 controls to read as 91.6 per cent, the denominator would have to be about 188, and more precisely 187.8. A denominator of 188 is roughly ten fewer controls than the 198 the second measurement ever looked at, and eight fewer than the 196 the business itself signed off. No population anywhere in this bank's record closes the gap, and that is the strongest possible statement that the gap is real. Working like for likeTwo measurements expressed on the same population, without which the difference between them is not a difference in performance. is not a stylistic preference here. Like for like is the only reading that survives being checked.
Why is 91.6 per cent against 86.9 per cent the wrong comparison?
How large is the gap once both measurements sit on the same 214?
Eleven point two percentage points is the honest answer, and it is also the least useful form of the honest answer. Take the same gap in counts instead. The business said 196 were working. The second measurement reached 172. The difference is 24 controls. And 24 over 214 is itself 11.2 per cent, so the percentage and the count are not two findings. The rate and the count are one object described twice.
The bank's year carries the figure 24 more than once, so keep the label attached to it. The 24 here is a count of controls. Elsewhere in the same case Rs 24 crore is a movement on a translation reserve, and Rs 24 crore again is a potential future exposure on one counterparty. Three different objects, one figure, and only the words around it say which is meant.
Try that switch on the household example and the point lands immediately. Telling somebody the kitchen is eleven per cent less safe than they thought produces nothing. Telling them the knife block is now within reach and the gas knob check gets skipped when everybody leaves at once produces two jobs before dinner. A percentage is a summary; a count is a list, and only a list can be finished.
The like for like gap is 11.2 percentage points. What is it in controls, and why does that second form matter?
Does a gap of twenty four controls mean somebody was not honest?
The failure is reading the gap as dishonesty
Twenty four controls is a large number, and the reflex on seeing it is to conclude that the business shaded its answers. The dishonesty reading is almost always wrong, and it is expensive to be wrong about. An honesty problem and a dishonesty problem call for completely different fixes, and only one of them would work here. If people were shading, the answer is discipline: challenge, sanction, a harder tone from the top. If the process is structurally optimistic, discipline changes nothing at all and the answer is evidence, separation of the two rating questions, and counting rather than describing. The wrong diagnosis costs a year spent being stern with people who were saying what they honestly believed.
A first line rating its own controls is not lying. The first line is answering a different question, with different information, from a different position in the building. Four ordinary mechanisms produce a gap of this size without anybody saying a word they do not mean, and all four are visible in this invented bank's own record.
Take the first mechanism slowly. The first mechanism does most of the damage. Ask anybody to describe a control they run and they will describe the arrangement: the check exists, it is done daily, here is who does it. The description is accurate about the design and says nothing about the eleven working days when the collateral valuation feed in process PR3 went stale and 340 loans carried the wrong mark. The stale feed is incident I10 in this bank's year. The control was well designed the whole time it was not working. Somebody answering in a workshop is not concealing that. A workshop question asks about the arrangement, and the arrangement is what comes back.
The second is subtler and it is worth naming carefully. Four months before that incident, near miss N3 caught the same collateral feed stale for 2 working days, and it was recorded as a routine catch and raised as nothing. From the inside, that reads as evidence the process works: something went wrong and the check found it. From the outside, the same entry reads as a feed that goes stale and a check whose catching depends on somebody looking. Nobody in the process saw the second reading. The second reading is only available once somebody counts how often the catch was needed, and nobody was counting.
The third and fourth are quieter still. A control nobody has ever tested has never been contradicted, so a rating of effective is not a conclusion, it is a belief that has not yet met any evidence, and it sits in the assessment looking exactly like a conclusion. And the rating word itself is loose. A check that works nine times in ten is not something an ordinary speaker of English would call ineffective, and it is not what a tester means by effective either. Two people can hold the same control in mind, mean what they say, and write down different answers.
A colleague reads the 24 control gap as proof that the business was not being straight. What is wrong with that reading, and what does it cost?
The business rated 196 of 214 controls effective. A second measurement reached 172 of the same 214. How far apart are they, like for like?
Move what the business rates and watch the gap close
One control: b, the number of the 214 key controls the business rates effective, from 172 to 214. The second measurement is held still at 172 of 214, being 80.4 per cent. Holding it still is a teaching device rather than a claim: a business that rated fewer controls effective would very likely be looking harder, and what a second measurement then found might well move too. Two consequences are shown together: the reported rate as b over 214, and the gap in both of its forms, percentage points and controls. The solved points are these. At b of 214 the bank reports 100.0 per cent, a gap of 19.6 points and 42 controls, and 42 is also the number of control findings the second measurement raised. At b of 196, 91.6 per cent, a gap of 11.2 points and 24 controls, exactly what this invented bank did. At b of 188, 87.9 per cent, 7.5 points and 16 controls. At b of 180, 84.1 per cent, 3.7 points and 8 controls. At b of 172, 80.4 per cent, no gap at all, and that is the crossing. Beside all of that sits the false comparison: 91.6 per cent against 86.9 per cent, apparently 4.7 points, false because 91.6 sits on 214 and 86.9 sits on 198. The control starts at b of 196, reproducing the case exactly.
With 196 of the 214 rated effective, this bank reports 91.6 per cent against an independent 80.4 per cent, a gap of 11.2 percentage points and 24 controls.
Educational illustration. Every count here belongs to one invented bank and none of it is a requirement, a norm or a figure from any authority. The 214, the 196 and the 172 are the case's own locked counts; b is the reader's own dial and is not a figure from the case. Holding the second measurement still while the business rating moves is a teaching device and not a claim about how the two would behave together in practice.
At what rating would this business's self assessment agree exactly with the second measurement, and what does that number show?
What is a self assessment still worth once the gap is admitted?
A reader who has followed this far can reasonably ask why anybody bothers. If the first line is structurally optimistic by twenty four controls, why not skip the exercise and go straight to the measurement that turned out to be right? The answer is that the two halves of the work are not the same work, and only one of them can be bought from outside.
Go back to step one. Before anybody rates anything, somebody has to name the risks that live in this process, the controls it actually depends on, and what those controls do on an ordinary day. The description exists nowhere except in the heads of the people who do the job, and every measurement that follows is a measurement of it. A second measurement can establish that a control did not operate on nine days out of a hundred. A second measurement cannot establish that the control exists, that it is the one the process leans on, or that everybody quietly works around it when the queue is long.
So the process is worth running, and there is exactly one thing that makes it worthless. An assessment whose result is never placed beside an independent measurement is a survey of confidence. A survey of confidence reports how the business feels about its controls. The feeling is genuinely useful information about the business and no information at all about the controls. The 11.2 point gap at this invented bank is not the problem with its self assessment. The gap is the only reason anybody knows the process has one, and a bank that never ran the second measurement would be sitting on the same twenty four controls with a much better slide.
Given a structural gap of this size, why run a self assessment at all?
How does anybody outside the process actually use a number like this?
Start with the head of operational risk, who receives both results. Purnima Ganeshan already expects the first line to run optimistic and would be suspicious of a process where it did not, so she is not reading 91.6 per cent against 80.4 per cent as a scoreboard. She reads the size and the direction of the distance, and then whether it moves. A gap of twenty four controls this year and twenty four next year says the rating method never changed. A gap that narrows after evidence is required for a rating says the fix worked. The number she manages is the gap, not either of the two rates that produce it.
An internal auditor uses it differently again. Rustom Batliwala is not trying to close the gap; he is trying to find out where it sits. Twenty four controls is a small enough list to look at one by one, and the interesting question is whether they cluster. Twenty four spread evenly across nine processes is a rating habit. Twenty four concentrated in two processes is a problem with those two processes, and it points somewhere specific. The count survives that question and the percentage does not, and that is the practical reason to keep both forms of the gap in view.
Control results of this kind are internal, so a lender or a counterparty looking at this bank from outside never sees either number, and that is worth saying plainly. An outsider can ask something structural on a call: does the institution measure its controls twice, by different people, and does it publish the distance between the two answers to its own committee. An institution that measures once has no way of knowing it is optimistic. The presence of a second measurement is a better signal than the level of either rate.
And the household version, where the mechanism works in exactly the same way at a much smaller scale. Anybody can rate their own habits: yes, I always lock the back door. The second measurement is somebody else in the house checking it for a fortnight and counting the nights it was open. Almost nobody is lying in that first answer. The count simply knows something the memory does not, and the useful number is never the rating or the count on its own. The useful number is the distance between them, and what that distance does once somebody starts paying attention.
Where the obligation to do any of this actually comes from
The method set out here is jurisdiction free. Listing risks, listing controls, rating design and operation separately and keeping the two rating questions apart works the same way in any country and in any institution. The requirement itself differs: who imposes it, on whom, and in what form.
The operational risk framework within which a self assessment sits originates with the Basel Committee on Banking Supervision at the Bank for International Settlements, bis.org. The same committee publishes the seven operational risk event categories named elsewhere in this material. The Basel Committee is a standard setting body and not an Indian supervisor. Naming only the global standard is the confident and common error in this subject, and it says nothing about what actually binds a bank in India. The Reserve Bank of India at rbi.org.in sets what an Indian bank must actually do about operational risk management and internal control.
Where the institution is a company rather than a bank, the duty to have and to report on internal financial controls sits in the Companies Act, whose text, applicability and exemptions come from the Ministry of Corporate Affairs at mca.gov.in, with the assurance standards and guidance from the Institute of Chartered Accountants of India at icai.org. Section numbers, thresholds, exemptions and effective dates all change, and the issuer's own text is the only place any of them is current.
Sources
| Source | Document | Site |
|---|---|---|
| Bank for International Settlements | The Basel Committee on Banking Supervision publications setting out the operational risk framework within which control self assessment sits | bis.org |
| Reserve Bank of India | What an Indian bank must actually do about operational risk management and internal control | rbi.org.in |
| Ministry of Corporate Affairs | The Companies Act duty on internal financial controls, its applicability and the form of the report | mca.gov.in |
| Institute of Chartered Accountants of India | The assurance standards and guidance behind reporting on internal financial controls | icai.org |
Vindhya Commercial Bank Limited, Purnima Ganeshan and Rustom Batliwala are invented.
Educational material. Not advice on any investment, tax, budget or market position.
