Control Design vs Control Effectiveness: Two Different Tests
Control design asks whether an activity would meet its objective if it happened exactly as written every time. Control effectiveness asks whether it did happen. The first is judged on the written control, the second on a sample over a period. The two tests run on different populations, they fail for different reasons, and their pass rates multiply rather than average.
Two sentences arrive about the same control and sound like one sentence said twice. The first says the control is well designed. The second says it is operating effectively. A reader who takes those as two ways of saying the thing works has merged two separate tests. A control can pass either test and fail the other, the two are settled by different kinds of evidence, and the fix for one buys nothing against the other.
One distinction carries the whole subject. Control designWhether the activity would meet its objective if it happened exactly as written, every time. asks what the activity would achieve if it were performed perfectly, so it can be answered about a control that has never once been performed. Control effectivenessWhether a properly designed control actually operated as stated throughout the period. asks what happened on days that have already gone, so it cannot be answered about such a control at all. One is about the shape of the thing. The other is about the record of it.
Are these two tests two versions of one question?
No, and the cleanest way to see it is to notice what kind of question each one is. Design asks a conditional questionOne answered under a hypothesis, so it can be settled by reading the control and walking one case.: if this activity happened exactly as written, every single time, would the thing it exists to prevent still be able to happen? A hypothesis does not need history to settle it, so it is answered by reading the control against its objective and following one case through. Operating effectiveness asks a factual questionOne answered by evidence about what happened, so it needs a sample spread across the period.: over a stated stretch of time, did the activity happen as stated? The factual question has no hypothesis in it at all, and only evidence about days already passed can settle it. Every other difference between the two tests, including the populations they run on and the way their results combine, follows from that one split.
The shape is easier to feel on a familiar street. A housing society decides that its night guard must write every visitor's name into a register at the gate, and that written rule is the control. Suppose the guard writes every name faultlessly for a year. If the register records a name and nothing else, it still does not say which flat let the visitor in, when they left, or whether anybody read it the next morning. Filled in perfectly for a year, that register could not tell anybody a stranger stayed inside the building after midnight, and that failure is reached without looking at a single night of it. Whether it was filled in at all is the other question, and no argument settles that one. Settling it takes looking: at which nights, at how many entries, at whether the handwriting stops while the regular guard was on leave. The first question was answered by reading; the second can only be answered by counting.
Get them the wrong way round and both failures are quiet. Answering the conditional question with history produces a sentence that sounds strong and proves nothing: it has worked every month for two years. Answering the factual question with reasoning produces a worse one. The sentence sounds like evidence: the procedure requires it, so it must be happening.
What does the design test actually ask, and what settles it?
The design test starts from the objective, the statement of what must not go wrong. The test then reads the control activity, meaning what somebody actually does, and asks a single question of the pair: is this activity, performed perfectly, capable of meeting that objective? The design test is a test of capability rather than of behaviour, and capability is a property of the written control rather than of the person performing it.
Three things settle it, and none of them is a sample: the objective stated properly, the written control itself, and a walkthroughFollowing one transaction through the control from start to finish to see what the control would do. of one case from one end of the control to the other. A control cannot be judged well designed in the abstract, only well designed for something, and many design failures are a control perfectly capable of meeting an objective nobody actually had. The walkthrough is run not to see whether the control usually happens but to see what it would do with a case it was built to catch, and one case is enough because the exercise inspects a mechanism rather than sampling anything. One walkthrough answers a design question for the same reason that opening one lock reveals what kind of lock it is, and that is why the answer does not improve when a hundred cases are walked instead of one.
Inside Vindhya Commercial Bank Limited, an invented bank of Rs 96,000 crore whose figures are illustrative throughout, the design test was run on all 214 key controlsThe controls an institution has decided are the ones it would notice the absence of, and therefore the only ones on the testing list. across its nine processes PR1 to PR9. The design test produced two numbers and nothing else: 198 controls were judged capable of meeting their objectives, and 16 were not. Not one of the 16 was found by watching anybody work. All 16 were found by reading.
A negative verdict stops the work rather than qualifying it. If the register at the gate cannot say who is still inside the building, there is no point counting how many nights it was filled in. No answer to the operating question could change the conclusion, so a negative design verdict makes that question unanswerable in the only sense that matters. The arithmetic of the two populations is a consequence of that sentence.
A tester says a control is well designed because it has worked every month for two years. What has gone wrong?
What does the operating effectiveness test actually ask, and what settles it?
The operating effectiveness test starts where the design test left off, with a control already judged capable, and asks a question about the world rather than about the control: over months 1 to 12, did this activity happen, in the way the control says, every time it was supposed to? The operating test is a test of behaviour rather than of capability, and behaviour is a property of a period rather than of a document.
Evidence with a particular shape settles it. Not the procedure. The design test already read the procedure. Not a conversation with the person who performs the control. A conversation is a claim and not a record. The evidence is the artefacts the control left behind on the days it ran: the signed reconciliation, the timestamp on the release, the exception log with somebody's initials against it. A control performed diligently in months 1 to 4 and abandoned in month 10 looks flawless to a sample drawn in month 3, so those artefacts are counted on a sample of occasions drawn across the whole period rather than clustered at one end.
The everyday version is a parent checking whether a child did the homework set for the year. The school diary records what was set, the design question, and nothing about what happened. Learning that means opening books from different months, and opening only the first fortnight gives a confident and wrong answer. A sample that does not reach across the period measures the beginning of the period and reports it as the whole of it. An operating conclusion therefore names the period it covers.
At the invented bank the operating test ran across months 1 to 12, and month 12 is the reporting date. The operating test produced two numbers of its own: 172 controls were found to have operated as stated, and 26 were not. The 26 differ from the 16 that fell at the design test not in severity but in what was wrong. Every one of the 26 could have met its objective and did not get performed the way it said.
One more property of the operating test is the source of most of the confusion on this subject. The operating test never runs first, and it cannot. Until somebody has decided that the activity is capable of meeting its objective, a diligent record of the activity happening is a record of something whose value is unknown. The order is not a convention that could have gone the other way, it is forced by what the two questions are.
Which six things separate the two tests?
Both tests are now defined on their own terms, so they can be laid against each other. Six criteria do the work, numbered DE1 to DE6 so that each can be named without being restated. DE1 is the question itself, and the other five are consequences of it rather than independent facts to memorise.
DE1, the question: design asks whether it would work, operation asks whether it did. DE2, what settles it: the written control read against its objective plus one walkthrough, against artefacts from a sample of occasions drawn across the whole period. DE3, the population: all 214 key controls, against only the 198 that passed design. DE4, what a failure is: a control that could not have met its objective, against one that could have and did not happen as stated. DE5, the fix: change the control, against make the existing control actually happen. DE6, what a pass buys: capability with nothing said about whether anybody performed it, against an operating pass that buys nothing at all on a badly designed control.
The columns come before the rows. Each column is a whole instrument: a question, a method, a population, a kind of failure, a kind of fix and a kind of assurance bought. Going across comes second. Only then do the differences stop being a list and start being one difference expressed six times.
Why does the second test run on 198 controls and not on all 214?
DE3, the population, is where most of the arithmetic trouble on this subject begins. The design test ran on all 214 key controls because at that point nothing had been ruled out. The operating test ran on 198, being the 214 less the 16 that carried a design gap. A reader who has met a lot of sampling reads that 198 as a shortcut: somebody had less time, so they tested fewer. The 198 is the opposite of a shortcut, and the reason it is not one settles every rate the testing produces.
The question is what an operating result on one of those 16 would mean. Suppose one is tested anyway, on a sample spread properly across months 1 to 12, and comes back perfectly clean. The finding is that a control which could not meet its objective was performed reliably. The objective is no closer to being met, and the money spent finding out is gone. An operating result on a control that failed its design test is not a weak answer, it is an answer to a question whose value was destroyed by the earlier result. Worse, those controls would then carry a clean operating result and an unmet objective in the same list as controls that genuinely work, and somebody would eventually average the list.
The 16 are easy to overreach on, so be exact about what the record holds. Not one of them was tested for operation. Not tested and failed, not tested and passed: never asked. Any picture that splits those 16 by operating outcome is asserting a result the testing never produced, so they are drawn here as one block with the question left blank.
A stray percentage does more damage on this subject than on almost any other, so the table carries the same six counts as rows, with the denominator named on every rate. The two pass rates sit on different bases, and nothing should put them beside each other without saying so.
| What was measured | Count | Out of | Rate |
|---|---|---|---|
| Designed effectively | 198 | 214 | 92.5 per cent |
| Carried a design gap | 16 | 214 | 7.5 per cent |
| Operating effectively, of those tested for operation | 172 | 198 | 86.9 per cent |
| Operating failures, of those tested for operation | 26 | 198 | 13.1 per cent |
| Operating failures as a share of the whole population | 26 | 214 | 12.1 per cent |
| Came through both tests | 172 | 214 | 80.4 per cent |
Two rows of that table are the same 26 controls at two different rates, 13.1 per cent and 12.1 per cent, and neither is wrong. The two rates answer different questions: what share of the tested population failed to operate, and what share of the whole population did. The rate changed without a single control moving. A rate is a count and a decision about a denominator, and nothing shows it more plainly.
Why was operating effectiveness tested on 198 controls rather than on all 214?
What does a failure of each one look like, and what fixes it?
DE4 and DE5 belong together. In practice one is never met without the other. A design gapA control that could not meet its objective even if performed perfectly, here a failure at stage CL3. is a control that could not have met its objective however well it was performed. The trade finance example in this invented bank is the shape everybody recognises once they have seen it: one person could both check the shipping documents and release the instrument. Performed impeccably by somebody who takes real care, the control still cannot do the thing a second pair of eyes exists to do. There is no second pair of eyes. The failure can be spotted from the written control alone, without ever learning whether anybody performed it.
An operating failureA properly designed control that was not performed as stated, here a failure at stage CL4. is the other creature entirely. The control would have worked, and the design test passed it on its merits. Then in the actual year it did not happen the way it says: the reconciliation was not run in four months of twelve, or the collateral valuation feed went unchecked for 11 working days in month 10 while everybody meant to check it was busy with something that had a deadline. Nothing is wrong with the control. Something was wrong with the year.
Now DE5, the fixes. Closing a design gap means changing the control: rewriting it so a different person releases what a first person checked, resourcing that second person, or building a system rule that refuses a release by the checker. Closing an operating failure means changing nothing about the control and everything about whether it happens: cover for the months the person was on leave, capacity so that deadline work does not eat routine work, and somebody accountable for the months it did not run. The two fixes cost different things, buy different things, and neither touches a single control in the other group.
Getting this backwards is expensive in both directions. Redesigning a control that would have worked solves a problem the institution does not have, and it usually makes the control heavier. A heavier control quietly makes the operating failure more likely next year. Enforcing a control that could never have met its objective makes a defective thing happen more reliably, and produces the most dangerous artefact in this subject: a clean record on a control that does not work.
A control is well designed and was not performed in four months of twelve. What kind of fix does it need?
What does a pass on either test actually buy?
DE6 is the criterion readers underrate, and the one that turns the other five into something usable. A design pass buys capability and nothing else: performed as written, this activity would meet its objective. A design pass says not one word about whether anybody has ever performed the control, and it cannot. No evidence about behaviour went into producing it. A control can hold a clean design verdict and have been performed on none of the days in the period it is about.
An operating pass is stranger. What it buys depends on a result from the other test. On a well designed control it buys the thing everybody wants: the objective was being met, on the evidence, over the period. On a badly designed control it buys nothing. The design test has already established that the activity does not meet the objective even when it happens perfectly. So the pass says the activity happened, and can say nothing at all about the objective.
The asymmetry is the argument for the order of the two tests. The value of an operating result is conditional on a design result; the value of a design result is conditional on nothing. One of these two tests can stand alone and the other cannot, and it is not the one most institutions report first. A pack that leads with an operating percentage has led with the number whose meaning depends on a number it has not shown.
A badly designed control passes its operating test with a clean sample of forty items. What has been learned?
Can a control pass one test and fail the other?
In both directions, and the two combinations look nothing alike on the ground. Start with well designed and badly performed. The monthly reconciliation is a good control: it compares two independent records, it is done by somebody who did not produce either of them, and a difference forces an investigation before the month closes. In the year it happened in eight months and did not happen in four. Everybody involved could describe the control accurately, and the evidence is simply not there for a third of the period. The reconciliation is a design pass and an operating failure, and the fix is a rota rather than a redesign.
The other direction is the one that hurts. Somebody performs a check every single working day, on time, with the evidence filed, and the design test then finds that the person doing the checking is the person who releases the payment. The checker who also releases is a design failure sitting under an exemplary record. The most diligent operating record in the institution can sit on a control that could never have caught the thing it exists to catch. The finding is demoralising to receive, which is exactly why it has to name which stage failed.
The two combinations meet in one rule about language. Calling either of those controls ineffective, with no further word, is accurate and useless. The first needs a rota and the second needs a second person, and a report that says only ineffective sends both fixes to the wrong place half the time. The word ineffective, on its own, destroys the only piece of information the testing actually produced.
What happens when the two rates are put together?
Combining the two results is where careful people produce a wrong number in good faith. The counts alone answer it with addition, and starting from the counts is the whole trick.
16 controls failed design and 26 failed operation. Before any computation: what share of the original 214 came through both tests?
Here is the whole worked instance in one place. Vindhya Commercial Bank Limited holds 214 key controls across nine processes PR1 to PR9, and both tests were run in the twelve months to the month 12 reporting date. The design test, the conditional one, ran on all 214: 198 passed and 16 carried a design gap, so the design pass rate is 198 over 214, being 92.5 per cent. The operating test, the factual one, ran on only the 198 that passed design: 172 passed and 26 did not, so the operating pass rate is 172 over 198, being 86.9 per cent.
Now put the two together. The two rates multiply. 92.5 per cent of 86.9 per cent is 80.4 per cent, and 80.4 per cent of 214 is 172, which is exactly the count that came through both tests. The identity is not a coincidence and not an approximation: the second rate is measured on the survivors of the first, so applying it to the survivors is the only thing that can be done with it. The result is the end to end rateThe share of the original population that came through both tests, which is the product of the two rates., the only one of the three rates that describes the population the bank actually runs.
Read as counts it needs no arithmetic beyond addition. 16 controls fell at the first test. 26 fell at the second. 16 plus 26 is 42. 42 of 214 is 19.6 per cent. 100 less 19.6 is 80.4. A rate carries a denominator somebody chose and a count carries nothing at all, so the counts never lie and the rates can. Every rate that follows carries its count underneath it, so the rate can be checked against the count rather than the other way round.
Why is averaging the two rates never right?
The mistake has one shape: two pass rates arrive in the same paper, somebody wants one number, and the two get averaged. The average of 92.5 and 86.9 is 89.7, and the figure sounds like a reasonable summary of two rates in the high eighties and low nineties. The average is not a summary of anything. 89.7 per cent is not a rounder version of 80.4 per cent, it is a different quantity wearing the same units.
Averaging two sequential pass rates
The usual defence is that averaging is a small approximation, close enough for a summary line, and that nobody makes a decision on the second decimal place. The defence would be worth hearing if the error were small and stable. The error is neither. Averaging is not an approximation at all, and it gets more flattering exactly as the controls get worse.
Hold the operating pass rate at this bank's 86.9 per cent, meaning 172 of the 198 tested for operation, and move only the design pass rate. At a design pass rate of 100 per cent, the true end to end rate is 86.9 per cent and the average says 93.4 per cent, a gap of 6.6 points. At this bank's actual 92.5 per cent, the truth is 80.4 per cent and the average says 89.7 per cent, a gap of 9.3 points. At 50 per cent, the truth is 43.4 per cent and the average says 68.4 per cent, a gap of 25.0 points. The error triples as the institution's controls deteriorate, and that is the last direction in which a measure ought to become kinder.
And there is no rate at which it comes right. Set the product equal to the average and solve for the design pass rate: it would have to be 117.8 per cent. More than every control in the population passing is what averaging would need in order to be correct, so the two readings never meet anywhere a real institution can stand. Averaging is an arithmetic impossibility rather than a matter of degree, and that is why the word approximation cannot be used about it at all.
The cost is not abstract. A committee told 89.7 per cent believes about 192 of its 214 key controls are working. The count that came through both tests is 172. Twenty controls, sitting in nine processes PR1 to PR9, exist in the committee's picture and not in the institution.
A pack reports design effectiveness of 92.5 per cent, operating effectiveness of 86.9 per cent and overall control effectiveness of 89.7 per cent. What is wrong with the third figure?
Drawing the two readings against every possible design pass rate makes the claim visible in a way no single pair of numbers can. Hold the operating pass rate at 86.9 per cent, being the 172 of 198 this bank measured, and let the design pass rate run from 40 per cent to 100 per cent. The truth is a straight line through the origin, being the design rate multiplied by a constant. Half of the average is a number that does not move, so the average is also a straight line, but a flatter one that starts far higher. Two straight lines with different slopes cross exactly once, and the only question is where.
Design passes at 92.5 per cent and operation passes at 86.9 per cent of those. Before the control below is moved: is there any design pass rate at which averaging the two rates gives the right answer?
Move the design pass rate and watch the two readings refuse to meet
One variable: the design pass rate applied to the 214 key controls. Holding the operating pass rate still is the point of the control, so it stays locked at 86.9 per cent, being 172 of the 198 tested for operation.
Which of the two tests produced the number in front of the reader?
Most readers of a controls report never see two rates. Readers see one, in a sentence that says controls effective and gives a percentage, and the sentence does not say which test produced it or what it is a percentage of. Two questions settle it, and both are short enough to ask out loud in a meeting. Ask what it is a percentage of, and ask which of the two tests produced it. No committee paper asks either question on the reader's behalf.
Run them against this bank's own headline. Controls effective, 86.9 per cent. What is it a percentage of? Of 198, not of 214. The operating test only ran on the controls that passed design. Which test produced it? The operating one, so the 16 design gaps are not inside that figure at all, in either direction. The 86.9 per cent is a true statement about a population that is 16 controls smaller than the one the bank actually runs, and the figure for the population it does run is 80.4 per cent.
A second trap sits on that exact number, and it looks like a coincidence and is not. 86.9 per cent is the operating pass rate on the 198 tested. The same fraction is also the ceiling. If every one of the 16 design gaps were closed and the operating failure rate stayed where it is, the end to end result would be 186 of 214 controls, or 86.9 per cent of the whole population. The same fraction arrives twice because closing every design gap makes the design population and the operating population the same 214, so the end to end rate collapses onto the operating rate. A report carrying both must name which 86.9 per cent it means in the same sentence, every time.
A committee paper says controls effective, 86.9 per cent. What two questions settle what that number means?
When does the difference between the two tests change a decision?
A remediation budget arrives that can close the 16 design gaps or the 26 operating failures, and not both. If the two kinds of failure were versions of one problem, the cheaper repair would be the obvious purchase. The two kinds of failure are not versions of one problem, so each purchase has to be priced against what it can deliver, and the arithmetic gives a bound on each.
Buying the design fixes raises a ceiling. Close all 16 design gaps, hold the operating failure rate where it actually sits, and the effective count moves from 172 to 186 out of 214, a gain of 14 controls and 6.5 percentage points, and it stops there. The 186 is the ceiling this bank's current operating behaviour permits, and the share it represents, 86.9 per cent of the whole 214, is the ceiling sense of that figure rather than the operating pass rate on the 198. A design fix cannot make anybody perform anything, so no number of design fixes reaches 90 per cent while the operating failure rate is unchanged.
Buying the operating fixes collects what the current design can already deliver. Close all 26 operating failures with the 16 design gaps left open and the effective count is 198 of 214, or 92.5 per cent. The design pass rate reappears as a ceiling, for the mirror reason: when nothing fails at the second test, the end to end result collapses onto the first test's rate. 92.5 per cent means the design pass rate on 214 in one place and a ceiling produced by perfect operation in another, so the object has to be named every time the figure is written.
198 is larger than 186, and stopping there would be careless. The two figures rest on assumptions of very different strength. The design purchase assumes only that the current operating failure rate holds. The operating purchase assumes that every one of 198 controls operates perfectly for a whole period, a thing no institution has achieved. The arithmetic gives what each purchase can deliver at most, and it does not settle which one to buy. That judgement belongs to the people accountable for the controls. The household version is a flimsy latch that everybody fastens every night and a heavy deadbolt nobody has turned since the monsoon: a lock and the turning of a lock are two different things, and the same sentence carries into an institution with 214 of them.
The same split decides how somebody outside reads a control report. A lender sizing a borrower, or an analyst sizing a bank, is handed one conclusion and has to work out what it can carry. A design conclusion says the institution has built something capable. The statement is about architecture, and it survives a change of staff. An operating conclusion says it happened last year. The statement is about capacity, workload and attention, and it does not survive a bad quarter. The two statements decay at different speeds, and treating them as one number throws away which of the two is being relied on.
The last place it changes a decision is the one closest to money here. The collateral valuation control in process PR3 sits over Rs 8,640 crore of secured advances, being 15.0 per cent of net advances of Rs 57,600 crore, and the denominator has to be named because that same balance is 9.0 per cent of total assets of Rs 96,000 crore and both figures are exact. If that control carries a design gap, no amount of running it more carefully protects the valuation of that book. If it carries an operating failure, the control was capable and the year got in the way. The size of the book does not settle which, and only the two tests do.
A remediation budget can close either the 16 design gaps or the 26 operating failures, not both. How does the difference between the two tests bear on the choice?
Where do the duties behind these two tests actually sit?
Nothing so far is national at all. The split between a conditional question and a factual one is a property of the two questions, and an institution anywhere would recognise both tests and both kinds of failure. The duty is not jurisdiction free: who has to report on internal control, to whom, in what form, and with what independent opinion attached. The duty and the body that holds its text are named below, and the text itself is read at the source.
Keeping to the body and the duty is practical. A section number, an applicability test, an exemption or a date is exactly the sort of thing that changes, and a document carrying one becomes quietly wrong without anybody noticing. Naming the body and the duty stays true, and it points to where to look. A pointer survives the next amendment.
Where an Indian institution's reporting duties on internal control sit
Where these two tests feed an Indian reporting duty on internal financial controls, that duty sits in the Companies Act, and its text, who it applies to, who is exempt and the form the report takes are held by the Ministry of Corporate Affairs at mca.gov.in. The assurance standard and the guidance note that govern how work of this kind is planned, performed and reported sit with the Institute of Chartered Accountants of India at icai.org. The obligations that bind a bank in addition, including its risk management arrangements and the standing of its internal audit function, sit with the Reserve Bank of India at rbi.org.in.
Section numbers, rule numbers, thresholds, applicability tests, exemptions, materiality levels, sampling minima and effective dates change. Name the duty, name the body, then read the current text at the source.
Sources
| Source | Document | Site |
|---|---|---|
| Ministry of Corporate Affairs | The Companies Act duty on internal financial controls: the text, who it applies to, who is exempt, and the form the report takes | mca.gov.in |
| Institute of Chartered Accountants of India | The assurance standard and the guidance note behind independent work on the design and the operating effectiveness of controls | icai.org |
| Reserve Bank of India | What binds a bank in India on internal control, risk management arrangements and the standing of the internal audit function | rbi.org.in |
Vindhya Commercial Bank Limited is invented.
Educational material. Not advice on any investment, tax, budget or market position.
