Control Deficiency: Rating a Weakness on a Four Point Scale
A control deficiency is a control that does not do what it was meant to do. Rating it places it on a scale: at Vindhya Commercial Bank Limited, invented, that scale is the bank's own four points, D1 observation, D2 deficiency, D3 significant deficiency and D4 material weakness. A finding is pushed up a step not by the money already lost, but by how much could still go wrong with nothing noticing.
Somebody tested a control and it did not work. Taken by itself, the sentence settles almost nothing. The sentence does not say whether anybody should care, whether the fix goes to the front of a queue or the back of it, or whether the board's audit committee needs to hear about it at all. Every one of those questions is answered by a rating, and a rating is a judgement rather than a measurement.
Vindhya Commercial Bank Limited, invented, tested 214 key controls over twelve numbered months and wrote 42 findings out of the results. The bank then placed all 42 on a scale of its own with four points on it. Eighteen landed at the bottom point, exactly one landed at the top, and the reason the one at the top is the year's most serious weakness has almost nothing to do with what it cost. Why cost and severity come apart is the subject of everything below.
What is a control deficiency, and when does a test result become one?
Start with the object itself. Everything that follows is an opinion about it. A control deficiencyA control that does not do what it was meant to do, whether by design or in operation. is a control that does not do what it was meant to do. The definition is that narrow deliberately. Not a bad outcome. Not a loss. Not an unhappy customer. A shortfall against what the control was there to achieve, and nothing more than that.
Here is the everyday version, and it is worth holding on to. The institutional version behaves exactly the same way at a thousand times the size. A household pays an electricity bill every month, and the arrangement is that one person reads the meter, compares the reading against the bill, and only then pays. The objective is that the household never pays for units it did not use. Now suppose that for four months running the bill got paid and nobody read the meter. The household now has a deficiency. Notice what it is not: it is not the money. The household may have overpaid by nothing at all. The arrangement stopped doing what it was there to do, and it stopped whether or not the electricity company happened to bill correctly.
A test result on its own is not yet a deficiency either. A tester who looks at forty payments and finds that three went out without a second signature has a fact. The fact becomes a deficiency at the moment somebody states what the control was supposed to achieve and asserts that this shortfall means it did not achieve it. A deficiency is a claim about an objective. A control with no stated objective cannot have a deficiency, and cannot be rated either. The objective is settled elsewhere and taken as given here, so the account starts where the shortfall has already been established.
There are two ways to fall short and they are not the same thing. A control can be badly built. Performed perfectly every single time, it would still not achieve the objective. Or it can be well built and simply not performed. Vindhya Commercial Bank Limited found 16 of the first kind and 26 of the second across its 214 key controls, and 16 plus 26 is exactly the 42 findings it wrote for the year. Which of the two a finding is matters enormously for the fix, and by itself it decides nothing about the rating. A design gap can be trivial and an operating failure can be the most serious thing in the record.
What are the four points on this bank's own scale?
One sentence before the four points, and it is not a formality. The scale below is this invented bank's own. Nobody issued it, no body publishes a definition of any point on it, and institutions write their own and differ. A reader who arrives somewhere real expecting these four words in this order will be wrong roughly as often as right. The four words are not what transfers. The shape is.
The shape is this: each point is defined against the point below it rather than by a threshold. There is no rupee figure, no percentage and no count that tips a finding from one step to the next. Between each pair of steps sits a question, and the answer to that question is what moves a finding up.
Read the scale from the bottom. In that direction the questions get harder. An observationThe lowest point on this bank's scale, where the control objective is not actually at risk. is a finding where the control objective is not actually at risk, and that is the only thing the bottom point means. The evidence was filed in the wrong place. The approver signed a day late but did approve. The procedure describes a step everybody performs and nobody has written down since the last version. None of that is nothing, and none of it puts the objective in doubt.
Step up and the question becomes: could something the control was there to stop now get through? If the answer is yes, the finding is a deficiency at D2 whatever else is true about it. The second signature missing on three of forty payments belongs here. A payment leaving without a second pair of eyes is precisely what that arrangement exists to prevent.
Step up again and the question changes shape. A significant deficiencyA weakness serious enough that the institution would have to act on the consequence. is not a bigger version of a deficiency, and reading it as one is where most raters go wrong. A significant deficiency is one whose consequence is something the institution would have to do something about: tell somebody outside, put a set of accounts right, unwind a decision, stop a product. The test is not how uncomfortable the finding is. The test is whether the consequence would force an action that goes beyond fixing the control.
At the top sits material weakness, named here rather than argued. The words that decide the last step carry machinery of their own and they are settled separately. The transferable part is that the top of the scale exists, that this bank put exactly one finding on it out of 42, and that the step into it is a different kind of judgement from the three below.
| Point | What moves a finding on to this point | This bank's count |
|---|---|---|
| D1 observation | Is there a shortfall at all, with the objective still not actually at risk? | 18 |
| D2 deficiency | Could something the control was there to stop or catch now get through? | 16 |
| D3 significant deficiency | Would the consequence force the institution to act beyond fixing the control? | 7 |
| D4 material weakness | Settled separately, and named here rather than argued | 1 |
| All four points | 18 + 16 + 7 + 1 = 42, the year's whole finding count | 42 |
Name the four points on this bank's scale in order, from the lowest to the highest.
What is a rating scale for, and who is it written for?
A scale is not there to describe a finding. The finding already describes itself: the condition, what should have been, why the gap exists and what it could lead to are all written down before anybody rates anything. RatingPlacing a deficiency on a scale, so that finite attention can be sorted. adds exactly one thing to the finding, an ordering.
A rating scale is written for the person who has to decide what gets fixed first, and it is useless to anybody who is not making that decision. The person who ran the control already knows what happened and does not need it. The tester has the evidence and does not need it either. The reader who needs it is the one holding a stack of 42 findings, a finite number of people who can work on them, and a year in which to do it.
Think of a hospital waiting room. The waiting room is the everyday version of exactly this problem. Each person can describe how ill they are perfectly well, so nobody is triaging in order to describe it. The triage happens because there are three doctors and forty patients, and somebody has to decide who goes first. A triage scale that put thirty of the forty in the same top category would have done no work at all. Such a scale would have described the room accurately and helped nobody.
A rating scale should be judged against that standard, and the standard is harsher than it looks. A scale is working when the group of findings at the top is small enough to be acted on and large enough to contain everything that genuinely belongs there. Both halves are real constraints, and a rater who only ever worries about one of them will drift in a predictable direction.
What moves a finding up a step, and what does not?
Three things move a finding up. The first is reachHow much of the institution's activity one control decides, which is the first thing that moves a rating., meaning how much of the institution's activity that one control decides. A control over one desk's stationery orders and a control over how the whole loan book is valued are not the same object, and a failure in each of them has a different ceiling on what it could do.
The second is detectionWhether anything else would catch the failure, which is the second thing that moves a rating., meaning whether anything else in the institution would catch the failure if the control let something through. Controls sit in layers, and James Reason set out the layered defence picture in Human Error, published in 1990: a hole in one layer is survivable when the next layer has no hole in the same place. A failure with a second layer behind it and a failure with nothing behind it are two different severities even when the failure itself is identical.
The third is durationHow long a failure could run before anybody noticed, which is the third thing that moves a rating., meaning how long the failure could have run before anybody noticed. A control that fails and is caught in a day has a bounded consequence. A control that could fail and run for a quarter has a consequence bounded only by how much passes through it in a quarter.
The right hand column of that drawing is the harder half. Every item on it is a real pressure that a rater feels in the room. The money already lost is the one that gets defended out loud, and the next two sections are about why it is wrong. The other three get defended quietly or not at all: nobody says a finding should be rated down because the fix is expensive, and findings still get rated down because the fix is expensive.
Two findings are identical in every respect except that a second control would catch one of the two failures within a day. How should they be rated?
Why is severity a judgement about what could happen rather than what did?
Here is the sentence that carries this guide, and it takes some holding: a rating is a forward looking judgement written about a period that has already closed. Both halves are true at once and neither cancels the other.
The evidence is entirely historical. A control did not operate. On these dates. On this population. The historical evidence is checkable, it can be shown to somebody who disagrees, and it is what makes a finding a finding rather than a worry. The question the rating answers is entirely prospective: how bad could the consequence of this weakness be. A rater who lets the historical half answer the prospective half is rating by the size of the loss that happened to occur. Rating by the size of the loss is the single most common error on this subject, and this bank's own record refutes it in one line.
Take the household again. The four months of unread meters is the historical fact. The rating question is not how much was overpaid. A household that does not read its meter is exposed to every future bill for as long as the arrangement stays broken. If the electricity company happened to bill correctly all four months, the household got lucky. Luck is not a control and it does not lower a rating.
At this bank the gap between the two halves is enormous and it can be drawn. The control at the top of the scale is the one over collateral valuation. The failure cost Rs 1.4 crore net, booked as incident I10. The control decides the valuation of Rs 8,640 crore of secured advances, being 15.0 per cent of net advances of Rs 57,600 crore. The base has to be named there. The same Rs 8,640 crore is also 9.0 per cent of total assets of Rs 96,000 crore, and both of those are exact.
Say the ratio out loud. The ratio is the argument. Rs 8,640 crore divided by Rs 1.4 crore is 6,171 times. The cost of the instance and the size of the thing the control decides are more than three orders of magnitude apart, and a rater who works from the first of those numbers will never arrive anywhere near the second.
A control failed and, by luck, no customer lost money. Does that lower the rating?
What happens to a year of findings when they are ranked by the loss they caused?
The failure worth spending a whole section on is not a careless mistake but a careful one. Ranking findings by the money they produced feels like the responsible thing to do, it is easy to defend in a room, it uses numbers that are already in the record, and it is wrong.
Ranking by loss puts the year's most serious weakness near the bottom of the list
Vindhya Commercial Bank Limited booked 13 operational loss incidents over the twelve months, I1 to I13, with a net loss of Rs 43.8 crore between them. The year's one D4 material weakness is the control behind incident I10, the collateral valuation feed that went stale for 11 working days and left 340 loans wrongly marked. Its net loss was Rs 1.4 crore.
Rs 1.4 crore is the fourth smallest of the thirteen net losses in the year, and it is 3.2 per cent of the Rs 43.8 crore total. Meanwhile the largest net loss of the year was Rs 15.4 crore, being 35.2 per cent of the total, and the largest gross loss booked in the year was Rs 42.0 crore. Neither of those carries the year's only material weakness. A ranking by loss puts the top of the rating scale in tenth place out of thirteen.
So why is the cheap one the serious one? Because reach and detection decide it and cost does not. The valuation control decides how Rs 8,640 crore of secured advances is valued, being 15.0 per cent of net advances, and it failed for 11 working days with nothing in the bank noticing. A control with that reach and no layer behind it could produce a consequence of very nearly any size, and Rs 1.4 crore is simply what this particular instance happened to cost. Every step of the loss based ranking is defensible, arithmetically correct and wrong.
One boundary before moving on. Two records are being read side by side here and they are not the same record. The loss log I1 to I13 counts what money left the bank. The control testing record counts what was found when controls were examined. Exactly one link between the two is recorded in this case: incident I10 and the one D4. No rating is asserted for any other incident. The largest net loss of the year is not rated low; it is not rated on this scale at all.
The bank's one material weakness cost Rs 1.4 crore net and its largest incident cost Rs 15.4 crore net. Why is the cheap one the serious one?
Can several small deficiencies add up to a bigger one?
AggregationTreating several related weaknesses as one bigger one, which is a judgement and not an arithmetic rule. is the question every rater eventually meets and almost nobody is taught. Six findings sit in one process. Each of them on its own is comfortably a D2. Read together they describe a process where supervision is not happening at all. Is that six deficiencies or one significant deficiency with six symptoms?
Aggregation is a judgement that has to be argued rather than an arithmetic rule that can be applied, and the difference between those two things is the whole answer. There is no count at which deficiencies become significant. Six unrelated deficiencies in one process are six deficiencies and nothing more, however uncomfortable the row of them looks in a report. Six findings pointing at one cause may genuinely be one weakness that has surfaced six times. Rating that weakness as six separate findings loses the single cause producing all of them.
The everyday version is a shop with six different stock discrepancies in one month. If they are six different products, six different staff and six different weeks, that is six problems. If all six ran through the same back door on the same shift, that is one problem wearing six faces, and a shopkeeper who writes six notes about six products has described the month accurately and learned nothing from it.
The record of this invented bank settles nothing at all on this point about its own 42 findings. The record holds the counts, 18 at D1, 16 at D2, 7 at D3 and 1 at D4, and it says nothing about whether any of them were aggregated before they were counted. The silence is worth stating rather than filling in. Inventing an answer would teach a rule where the practice has only a judgement.
Six separate deficiencies all sit in one process and all point at the same weak supervision. Can they be treated as one more serious finding?
What does a year of ratings look like when a committee reads it from the top?
Now the artefact itself. Vindhya Commercial Bank Limited rated all 42 of its findings on its own four points: 18 at D1 observation, 16 at D2 deficiency, 7 at D3 significant deficiency and 1 at D4 material weakness. 18 plus 16 plus 7 plus 1 is 42, and the counts tie exactly.
As shares of the 42 those counts are 42.9, 38.1, 16.7 and 2.4 per cent. The four rounded shares sum to 100.1 per cent while the counts sum exactly to 42, so the rounding sits in the shares and never in the counts. Anybody adding the percentage column and finding 100.1 has not found an error in the record; they have found the arithmetic of rounding four numbers to one decimal place.
Now read the same 42 the way a committee actually reads them, cumulatively from the top. D4 alone is 1 finding, being 2.4 per cent. D3 and above is 8 findings, being 19.0 per cent, made of the 7 significant deficiencies plus the one material weakness. D2 and above is 24 findings, being 57.1 per cent. D1 and above is all 42, being 100 per cent. The four readings are 1, 8, 24 and 42.
Stop on that 24 and name it. The bank holds three different 24s, and a bare one has merged two of them. The 24 here is the count of findings at D2 and above, being 16 plus 7 plus 1. A second 24 is the like for like gap between what the business rated effective and what independent testing found effective, being 196 less 172. A third is the 24 open issues sitting in the two oldest ageing buckets, AG4 at 15 plus AG5 at 9. Three different objects, one number, and every one of them has to carry its label.
24 findings sit at D2 and above. What else does the number 24 mean in this bank?
The steps between those four readings are what makes the scale worth having. From 1 to 8 is a rise of 7. From 8 to 24 is a rise of 16. From 24 to 42 is a rise of 18. The list roughly triples at each of the first two steps down, and that unevenness is the whole reason a committee can use the scale at all. A flat distribution, with roughly ten findings at each point, would give a reader no reason to stop anywhere.
And here is what the distribution actually buys. Eight findings sit at D3 and above, so a committee reading top down reaches a workable list on the second step. A scale that had put 24 findings at D3 and above would have produced a list nobody could act on. A scale that had put only the one D4 up there would have produced a list that hid seven things worth knowing about. Neither of those is a better scale. Both of them are a scale that has stopped doing the one job it has.
A year produced 42 findings, one of them at the top of the scale. How many are at significant deficiency and above?
Slide the reporting cut down the scale and watch a workable list turn into an unusable one
One control: which step of the scale a committee decides to read down to, from D4 alone to all four points. One consequence: how many of the year's 42 findings are then in front of it, and what share of the 42 that is. The bar is the year's findings laid out most severe on the left, and the marker is the cut.
Reading down to D3 significant deficiency and above puts 8 of the 42 findings in front of the committee, being 19.0 per cent, which is a list a committee can actually work through in a year.
What is a rating actually spending?
A rating looks free. Nothing moves when a rater writes D3 instead of D2, no money changes hands, and the finding is the same finding either way. The impression that rating is free is the reason scales drift upward, and the impression is wrong for a reason that is easy to state and easy to forget.
Vindhya Commercial Bank Limited carried 92 open issues at the reporting date, aged across five buckets AG1 to AG5, and it closes them at a finite rate because a finite number of people work on them. Every step up the scale is remediation effort taken away from something else, so a rating is a claim on a capacity that is already fully spoken for. Push a finding from D2 to D3 and something that was at the front of the queue is now second.
A scale where everything is serious rations nothing. Such a scale has not made the institution safer, it has not told anybody what to do first, and the work still gets done in whatever order somebody chooses once they stop reading the ratings. The problem has simply moved from the rating into the remediation queue, where it is harder to see and nobody is accountable for it.
An institution rates 24 of its 42 findings as significant deficiencies. What has it achieved?
What can a rating never do?
Two things, and they are worth stating flatly because both get attempted regularly.
The first is that a rating cannot make a finding true. If the condition is wrong, or the criteria the finding is measured against turn out not to apply, no rating rescues it. A weak finding rated D1 is still a weak finding, and rating it low is the polite way of not withdrawing something that should be withdrawn. The right move on a finding that will not stand up is to take it out of the report, not to park it at the bottom of the scale.
The second is that a rating cannot be settled by the person accountable for the control. The control owner has every reason to want a finding rated lower and no standing at all to decide it, and that is not a comment on anybody's integrity, it is the structure. The Institute of Internal Auditors restated the three lines model in 2020. The model puts a control's owner in the first line and the people giving assurance elsewhere, precisely to keep a rating out of the hands of the person being rated. Where the first line rates its own controls, what comes out is a self assessment, and a self assessment carries different weight from an independent test whatever number is on it.
A control owner argues that a finding should be rated one step lower. Who decides?
Who actually reads these ratings, and what do they do with them?
Three people read the same rated list and take three different things from it, and knowing which three is worth having before a rating is written.
The head of the process being rated reads it as a work plan. Purnima Ganeshan, head of operational risk at this invented bank, has the findings themselves and does not need the rating to tell her which controls are weak. She needs it to tell her which two her team starts on in the first quarter, given that 92 issues are already open and the queue is not empty. For her the rating is the only part of the report that changes what happens on Monday.
The independent director on the audit committee reads it as a shape. Not the 42 individual findings, none of which anybody outside the process can hold in their head, but the profile: one at the top, eight at the second step down, and whether that profile has moved since last year. A year in which the D3 and above count went from 8 to 20 is a different institution from a year in which it went from 8 to 6, and the individual findings will not tell them that.
The credit analyst at another institution, looking at this bank as a counterparty rather than as an employer, reads it as a signal about the reporting itself. The analyst cannot verify the findings. The analyst can ask instead whether the distribution looks like a real triage or like a formality: a year with 42 findings and zero at the top says either that the controls are excellent or that nobody is willing to write the word, and no outside reader can tell those apart from the count alone.
And the household version once more. The mechanism is the same at every scale. A person who marks every bill in the drawer as urgent has not made the rent more likely to be paid on time. The marking has made the drawer unreadable, and next month they will pay whatever happens to be on top.
What is named here, and where the binding version lives
The four point scale D1 to D4, every count, every share and every incident belong to Vindhya Commercial Bank Limited. The scale is the bank's own rather than a standard, and no body at all publishes a rating definition, a threshold or a classification for it.
Where a reporting obligation on a control weakness arises in India, the duty on the board and on the auditor in respect of internal financial controls sits under the Companies Act, and the text, the applicability, the exemptions and the form of the report all come from the Ministry of Corporate Affairs at mca.gov.in. The assurance standard and the guidance note that sit behind the work come from the Institute of Chartered Accountants of India at icai.org.
A bank is bound in addition by the risk management and internal control arrangements it must maintain and by what it must report, and both come from the Reserve Bank of India at rbi.org.in. Where an international standard is the origin of something named here, that origin is the Basel Committee at the Bank for International Settlements at bis.org, and what India does with it is set by the Reserve Bank of India rather than by the Committee. The binding text in each case sits with the body named, at the source.
Sources
| Source | Document | Site |
|---|---|---|
| Ministry of Corporate Affairs | The Companies Act duty on the board and the auditor in respect of internal financial controls, its applicability and the form of the report | mca.gov.in |
| Institute of Chartered Accountants of India | The assurance standard and the guidance note behind reporting on internal financial controls | icai.org |
| Reserve Bank of India | What actually binds a bank in India on risk management arrangements, internal control and what must be reported | rbi.org.in |
| Bank for International Settlements | The Basel Committee standards that are the origin of the supervisory expectations an Indian requirement implements | bis.org |
| Institute of Internal Auditors | The three lines model, restated in 2020, which places a control's owner in the first line and assurance outside it | theiia.org |
| James Reason | Human Error, 1990, the layered defence picture behind rating a failure for whether anything else would detect it | Cambridge University Press |
Vindhya Commercial Bank Limited and Purnima Ganeshan are invented.
Educational material. Not advice on any investment, tax, budget or market position.
