How to test Control Design and Effectiveness
Test design first and operation second, and stop the moment design fails. Restate the objective, walk one transaction to judge design, define the population and the period, choose a sample against the failure rate the test is intended to detect, gather evidence, classify every exception by the stage it belongs to, and conclude on the period with the base named. At the invented Vindhya Commercial Bank Limited it took 214 controls to 172.
Testing a control is a method rather than a definition. Control design and control effectiveness, how the two differ, and the bearing of the control under test on a reported number are all settled before any testing starts. The method adds the order of operations, the arithmetic behind the one step everybody gets wrong, and a conclusion that can be defended a year later to a stranger who was not in the room.
In what order should a control be tested?
The everyday case is a shop that weighs rice on a spring balance. There are exactly two questions that can be asked about that scale, and they are not the same question. The first is whether the scale, used exactly as the shopkeeper says he uses it, would give an honest weight at all. The second is whether he actually used it that way on every sale in the last year. The first question is about the instrument. The second is about the habit. How faithfully somebody performed a procedure that could never have worked is not information about anything. So the second question cannot be answered usefully until the first has been.
Control testing runs on exactly that split, and the eight steps below hold it. The eight steps are numbered TS1 to TS8 for teaching, and no bank record carries those labels. Everything from step TS4 onward exists solely to answer the second question, and it never begins until step TS3 has let it.
What has to be settled before any evidence is gathered?
Step TS1 sounds like paperwork and is not. Before anything is looked at, the control objective has to be restated as a sentence that somebody could conceivably fail. The example is the collateral valuation control inside process PR3 at Vindhya Commercial Bank Limited. The objective, written so it can be failed, is this: no loan is marked on a stale valuation. The objective has a subject, a verb and a condition, and any item can be held against it and answered yes or no.
In real life the objective usually gets written as: to ensure that collateral values are appropriately maintained. Nothing can fail that sentence. Appropriately maintained is not a state of the world, it is a mood, so no item can be held up as breaking it. An objective that cannot be failed produces a test that cannot be failed either, and a test that cannot fail is not evidence, it is decoration. Everything from TS2 onward is measured against the sentence written at TS1, so a soft sentence quietly softens all eight steps at once.
The household version is a fair one. A teenager told to be sensible with the scooter leaves no way of establishing later whether they were. A teenager told that the scooter is never taken out after nine at night has been given a sentence with a fact attached to it, and both sides know what checking would even mean.
How is design tested, and what is enough to settle it?
Step TS2 has two parts and they run in this order. Read the written control, then perform a walkthroughFollowing one transaction through a control from start to finish to see what the control would do.: follow one transaction end to end, from the moment the valuation feed arrives to the moment the mark lands on the loan. Nothing is being counted. The question is a conditional one, and the conditional is the whole of it.
The question is: if this activity happened exactly as written, every single time, would no loan be marked on a stale valuation? Notice what that question does not ask. The question does not ask whether the activity did happen. It does not ask how often it happened. The question asks after capability on the best day the activity ever has. If the answer is no even on the best day, the control has a design gap and no amount of diligence in performing it can close that gap.
Design is a question about capability rather than about frequency, and that is exactly why one walkthrough settles it and a sample cannot improve on it. Walking forty transactions through the same written control shows the same capability forty times over. The second walkthrough adds nothing the first did not already establish, unless the control is genuinely different for different item types, in which case there is more than one control and TS1 should have said so.
Why is the design test settled by one walkthrough rather than by a sample?
What happens to the test when design fails?
Step TS3 is the stop ruleEnding the test when design fails. How reliably an unworkable control was performed answers nothing., and it is the step most testing plans do not have. If the design test at TS2 says the control could not meet its objective even performed perfectly, the test ends there. No sample is drawn. No measurement is taken of how diligently it was performed. A high performance rate on an unworkable control establishes only that people are conscientious. Conscientiousness was never the question, so the measurement would be real work producing a real number that carries no information at all.
In this invented bank the design test ran on all 214 key controls. Sixteen of them failed it, being 16 over 214, or 7.5 per cent. The 16 design failures went no further. The operating test then ran on the 198 that remained. Twenty six of those failed, being 26 over 198, or 13.1 per cent. The reason the second denominator is 198 rather than 214 is the stop rule, and that makes it a design decision in the method rather than a gap in the fieldwork. End to end, 172 of the 214 controls that entered came out effective, being 80.4 per cent.
The 80.4 per cent is not the average of the two pass rates. Design passed 92.5 per cent and operation passed 86.9 per cent. Multiplied, the two rates give 80.4 per cent, exactly 172 over 214. Averaged, the two rates give 89.7 per cent, a full 9.3 percentage points too kind. A control has to survive both tests rather than one of them on average, and that is the whole of the difference. The 26 operating failures are 12.1 per cent of the whole 214, and 7.5 plus 12.1 is 19.6, the complement of 80.4.
Why did the operating test in this invented bank run on 198 controls rather than 214?
How are the population and the period defined for the operating test?
Once a control has survived TS3, the operating test begins, and it begins with two definitions rather than with any evidence. The populationEverything the control was supposed to act on in the period, and the set a sample is drawn from. is everything the control was supposed to act on. For the collateral valuation control, that is every loan marked in the twelve months, and not every loan on the books, and not every valuation received. Get the population wrong and every figure downstream is a share of the wrong thing.
The period is the second definition and it is the one people quietly bend. The assertion being tested is that the control operated effectively throughout the period, so the period here is the whole twelve months. Defining the period as the last month because that is where the records are tidiest does not make the test easier, it makes the conclusion smaller, and smaller in exactly the direction that hides things. The period is part of the conclusion rather than part of the convenience, and choosing it for the quality of the evidence settles the answer before any evidence is read.
Vindhya Commercial Bank Limited makes the point. The collateral valuation control did not operate for 11 working days, and those days fell in month 10. A test scoped to month 12 alone would have passed that control cleanly and truthfully, and would have missed the year's single most serious control finding while doing so.
A test defines its period as the last month of the year because that is when the evidence is cleanest. What is wrong?
How big does a sample have to be, and what does the size actually depend on?
Sample size is the step everybody underestimates. A sampleThe items actually examined, whose size is a statement about what failure rate the test could have noticed. is not a quantity of comfort. A sample is a statement about what the test was capable of noticing. Where that statement cannot be said out loud, what has been chosen is not a sample but a number.
The whole of the arithmetic sits in the open. Suppose a control truly fails a share p of the time, and s items are examined, chosen independently of one another. The chance that all s come back clean is 1 less p, raised to the power s. That is it. Nothing else is needed, and independence between the items is an assumption rather than a fact of the world.
Turned around, it becomes a design tool. The failure rate the test is meant to catch is decided first, and then the smallest s that makes a clean result unlikely if the control really is failing at that rate. The rate chosen is the detectable rateThe failure rate a given sample size would be likely to catch, written down before anything is selected., and it goes on paper before a single item is selected. The tolerance used throughout for how often a clean sample from a genuinely broken control is accepted is 5 per cent. The 5 per cent tolerance is a working choice rather than a requirement, and every sample size that follows from it is arithmetic rather than rule.
| Failure rate to be able to detect | Smallest sample on the 5 per cent tolerance used here | Roughly three over the rate |
|---|---|---|
| 25 per cent | 11 items | 12 |
| 13.1 per cent, this invented bank's own operating failure rate | 22 items | 23 |
| 10 per cent | 29 items | 30 |
| 5 per cent | 59 items | 60 |
| 2 per cent | 149 items | 150 |
| 1 per cent | 299 items | 300 |
| 0.5 per cent | 598 items | 600 |
Read from the bottom, the ladder makes the shape jump out. Halving the failure rate to be caught roughly doubles the sample needed: 29 items at 10 per cent, 59 at 5 per cent, 149 at 2 per cent, 299 at 1 per cent, 598 at half of one per cent. The rule of thumb is that the sample is about three divided by the rate, and it holds to within one item across the whole ladder. The exchange rate is brutal, and it is the honest one.
A tester says the sample is 40 items because that is the standard for a daily control. What has not been done?
A tester examines 25 items and finds nothing wrong. Before the control below is moved: what failure rate was that test actually capable of catching?
Drag the failure rate down and watch the sample climb
One variable: the operating failure rate the test is intended to detect, from 25 per cent down to 0.5 per cent. One consequence: the smallest sample that leaves at most a 5 per cent chance of a clean result from a control failing at that rate. The default sits at 13.1 per cent, this bank's own operating failure rate of 26 of the 198 controls tested. Catching a rate that size needs 22 items.
To be confident of catching a failure rate of 13.1 per cent, the smallest sample is 22 items, and at that rate a habitual sample of 25 would come back clean anyway only 3.0 per cent of the time.
At the far right of the control the number stops looking like a sampling question and starts looking like a resourcing question. Resourcing is what it always was. A control tested down to a failure rate of half of one per cent costs 598 items of somebody's attention. The cost may be worth it for the control that feeds the collateral mark on Rs 8,640 crore of secured advances, and it is plainly not worth it for a control over stationery. The judgement belongs to the people accountable for it, and the arithmetic at least forces them to make it out loud.
What counts as evidence, and what does a tester do with an exception?
Step TS6 selects the items and gathers evidence, and evidence here has a specific meaning: a trace that a stranger can check a year later without asking anybody anything. For each selected loan at this bank, the trace has three parts. The valuation date carried on the feed. The date the mark was applied to the loan. And the record that the check itself actually ran on that date. The third one is the one that is usually missing, and it is the only one of the three that speaks to the control rather than to the transaction.
Two ways of gathering that trace answer different questions, and the difference is worth holding. InspectionLooking at the record the control left. It answers whether the control happened. means looking at the record the control left behind, and it answers whether the control happened. Re-performanceDoing the control again yourself. It answers whether the control works. means doing the control again yourself, and it answers whether the control works. A signature is evidence of a signature and not evidence of a judgement. A control judged only by the record it left can be a control nobody ever actually thought about.
The evidence for a sampled item is a screenshot showing the check exists in the system. Is that evidence the control operated?
Step TS7, and the exception that turned out to be two findings
An exceptionAn item in the sample where the control did not do what was claimed. Each one is classified by stage. is an item where the control did not do what was claimed for it, and every exception has to be classified back to the stage it belongs to. Most of them are simple. The control exists, it works, somebody skipped it on this item, and that is an operating failure at stage CL4 of the control lifecycle.
The bank produced an exception that is not simple, and it is the reason the step exists. The collateral valuation control did not operate for 11 working days and 340 loans were wrongly marked. No customer lost money and the net loss booked was Rs 1.4 crore. So far that is an operating failure. But read the second half of the sentence: no monitoring control detected it, for eleven consecutive working days. The eleven days are an operating failure of the control being tested. The silence around them is a question about whether anything was ever designed to notice, and that is a stage CL3 design question about a different control altogether.
Classifying the whole thing as one operating failure loses the second half, and the second half is what made this the year's single material weakness, touching the valuation of Rs 8,640 crore of secured advances. The everyday version is the smoke alarm. If the kitchen caught fire and nobody was in the house, the fire is one problem. The alarm never going off is a completely different problem, and only one of the two gets fixed by being more careful with the stove.
A control did not operate for eleven working days and nothing in the bank noticed. How many stages does that exception touch?
How is the conclusion written, and what does it have to name?
Step TS8 is a sentence, and the sentence has a fixed shape. The conclusion states the result, the period, the population, the sample and the base, all in the same place. Nobody reading it later can then quietly attach it to a different denominator. Something of the form: over the twelve months, on a population of every loan marked in the period, a sample of 22 items chosen to be able to detect a failure rate of 13.1 per cent produced no exception, so the control is concluded effective for the period on that basis.
Every clause in that sentence is load bearing. The one people drop is the basis, and the basis is exactly the clause that says what the test could have seen. Across the whole of this invented bank the same eight steps produced 172 controls effective of the 214 that entered, being 80.4 per cent, and that figure is reached as 214 less 16 design gaps less 26 operating failures. The 80.4 per cent is never reached by multiplying one rate by another and then rounding, and the rate is never quoted without the count beside it.
| Step | What it settles | What it hands to the next step |
|---|---|---|
| TS1 objective | What the control must not let happen | A sentence that can be failed |
| TS2 design | Whether the activity could ever meet it | A pass, or a design gap |
| TS3 stop rule | Whether the test continues at all | 198 of the 214 continued |
| TS4 population and period | What a sample is a sample of, and over what span | Every loan marked in twelve months |
| TS5 sample | What failure rate the test could notice | 22 items at 13.1 per cent |
| TS6 evidence | Whether it happened and whether it works | Three traces for each item |
| TS7 exceptions | Which stage each failure belongs to | CL3 gaps and CL4 failures, kept apart |
| TS8 conclusion | What may honestly be said about the period | 172 effective of 214, being 80.4 per cent |
The sample of twenty five that nobody ever questioned
Here is how this goes wrong, and it goes wrong quietly. A tester picks 25 items because 25 is the size everybody uses. Nothing is found. The working paper records that the control operated effectively. Every review meeting accepts it. Twenty five looks like a serious number, and nobody in the room has ever asked the question the number is an answer to.
The arithmetic on that test shows what it could actually do. On the 5 per cent tolerance used here, a clean sample of 25 items is enough to be reasonably confident about a failure rate of roughly 11 per cent and no better. Against a control failing one time in a hundred, a sample of 25 comes back clean about 78 per cent of the time, or close to eight times in ten. So the working paper that says nothing was found is telling the truth and telling almost nothing.
The awkward part is that at this particular bank, 25 would have been sound. Its own operating failure rate across the tested population is 13.1 per cent, and detecting a rate that size needs 22 items, so 25 clears it comfortably. The fit is luck, not method. The number was chosen before anybody asked what question it was answering. A sample chosen by habit is a test for a badly broken control and nothing else, and the tester almost never says so out loud. Every sample size in the ladder above is arithmetic on a stated tolerance rather than a figure lifted from a standard. The assurance standard and the guidance note behind this kind of work are held by the Institute of Chartered Accountants of India.
Who actually uses this method, and what do they do with the result?
Four people read the output of these eight steps, and they read it for four different things. Knowing that matters before step TS8 is written.
Rustom Batliwala, head of internal audit at this invented bank, uses the method to plan a year of work. His constraint is not curiosity, it is hours. The sample ladder turns a wish into a plan. A testing cycle that wants to detect a 2 per cent failure rate on forty controls has committed to roughly 149 items each, or nearly six thousand items of somebody's attention, and that number arrives before the year starts rather than in month 9. The ladder is a budgeting instrument as much as a statistical one, and the testing plan that never does this arithmetic simply discovers it late.
The independent director on the audit committee reads the conclusion, and reads it for one thing above all: what base was it drawn on. A committee shown 86.9 per cent effective is being shown the pass rate on the 198 controls that reached the operating test. A committee shown 80.4 per cent is being shown the whole 214. Both are true, they differ by 6.5 percentage points, and only one of them answers the question the committee thinks it is asking.
Vivek Anantharaman, the chief financial officer, reads it for the link to a reported number. A stale mark feeds a provision and a provision feeds a reported figure, and that is why the collateral valuation control matters at all. And a credit analyst at another institution, looking at a bank from outside as a counterparty, reads whatever is published for the shape of the failures rather than the count. Sixteen design gaps means sixteen controls that were never going to work; twenty six operating failures means twenty six controls that would have worked if somebody had performed them. The two shapes of failure are two different institutions wearing the same headline percentage.
The mechanism does not change with scale, so here is the household version. A person running a small tuition class wants to know whether the attendance register was actually filled every day. Checking five days out of two hundred is not a light version of checking properly, it is a test that can only catch a teacher who never filled it at all. Catching the one week in the year that got skipped takes more than five days, and knowing that before the checking starts is the whole of step TS5.
What can this method not do?
Three things, and each of them is a limit rather than a weakness.
The method cannot prove a control never failed. A sample of s items on a control failing a share p of the time comes back clean with probability 1 less p raised to s, and that number is never zero for any sample short of the whole population. A clean sample is evidence of absence only to the extent the sample was capable of finding something, and the conclusion at TS8 carries its basis for exactly that reason. A clean 299 items says a failure rate of about 1 per cent or worse is unlikely to have escaped notice, on the tolerance used here. The same clean 299 says nothing whatever about a control that failed twice in a year in a corner the sample never touched.
The method cannot rescue a control that left no record. If the third trace was never kept anywhere, and nobody recorded that the check ran, then the control cannot be tested for operation at all, and that is itself a finding rather than a scheduling problem. An untestable control and an ineffective control are not the same thing, but they land in the same place in a report and they should.
And it says nothing about tomorrow. Every one of these eight steps looks backwards over a defined period. A control that operated perfectly for twelve months and whose one experienced operator left in month 12 has a clean test result and a problem, and no amount of sampling the past will surface it. The test is a statement about a period that has ended.
A control passes its operating test with a clean sample of 299 items. Has it been shown never to have failed?
Where does the standard behind this work come from?
Everything above is arithmetic and craft rather than rule, and where the rules actually live is worth being precise about.
What is named here, and where the binding version lives
Every control count, failure count, rate, loss and rupee figure on this subject belongs to Vindhya Commercial Bank Limited, and none of the arithmetic above is a requirement.
The assurance standard and the guidance note under which work of this kind would be performed in India are held by the Institute of Chartered Accountants of India at icai.org. The reporting duty that such work serves sits in the Companies Act, and its text, its applicability, its exemptions and the form of the report all come from the Ministry of Corporate Affairs at mca.gov.in. Additional binding requirements on a bank, including its risk management and control arrangements, come from the Reserve Bank of India at rbi.org.in.
Every sample figure is arithmetic on the stated 5 per cent tolerance, and that arithmetic assumes each item is examined independently of the others. Section numbers, rule numbers, thresholds, exemptions and effective dates live in the issuing body's own text and nowhere else.
Which body holds the assurance standard and the guidance note this method would be performed under in India?
Sources
| Source | Document | Site |
|---|---|---|
| Institute of Chartered Accountants of India | The assurance standard and the guidance note behind reporting on internal financial controls, including anything on sampling and evidence | icai.org |
| Ministry of Corporate Affairs | The Companies Act duty on internal financial controls, its applicability, its exemptions and the form of the report | mca.gov.in |
| Reserve Bank of India | What actually binds a bank in India on risk management, internal control and assurance arrangements | rbi.org.in |
Vindhya Commercial Bank Limited, Rustom Batliwala and Vivek Anantharaman are invented.
Educational material. Not advice on any investment, tax, budget or market position.
