Bias, Fairness and Explainability Testing
These three tests compare outcomes across groups, ask whether the same file scored twice returns the same answer, and check whether the reason a person was given is the reason that produced their outcome. The firm's own records fix what can be compared. At one invented bank the single group difference found was defined by equipment, in a step the scoring model never touched.
One fact carries everything after it, so start with what one of these tests actually is. A fairness test is not a property a system has or lacks. A fairness test is not a calculation somebody performs quietly on a model and reports as a score. A fairness test is a question somebody asks of a system that is already deciding about people, with a population in mind, a comparison in hand and a record that either supports the comparison or does not. Somebody chooses the question. Somebody finds the two groups. Somebody reads the answer and decides what happens next. Every cost in fairness testing attaches to one of those three people.
Bias itself, where it comes from, and why the competing definitions of fairness cannot all be satisfied at once, are set out under bias and fairness in financial AI. The asking is what matters now: what each test compares, what each one costs, and who bears the answer when it comes back.
What does each of these three tests actually ask?
Consider the weighing scale at a neighbourhood shop. Three complaints could be made about it, and they are not the same complaint. The first is that it seems to weigh differently depending on who is standing at the counter. The second is that it gives two different readings for the same bag of rice, weighed twice within a minute. The third is that when the shopkeeper says the price went up because of transport, transport is not actually why. The three complaints need three different investigations, and settling any one of them leaves the other two exactly where they were.
The three tests are those three investigations, run on a system that decides about people rather than on a scale. The first compares outcomes across groups. The second compares one file against itself. The third compares a stated reason against the outcome it is supposed to explain. A firm that runs one of them and reports that it has tested for fairness has answered one of three questions and left the reader to assume it answered all of them.
What does the explanation test actually do to a file?
What is fairness testing actually comparing, and across which groups?
The first of the three is the outcome distribution testA comparison of outcomes across groups the firm can assemble from what it already records., and it is the one people mean when they say the words. The procedure is short enough to state in a sentence. Take the outcomes a system has already produced, split the people who received them into two or more groups, and compare the shares. The mechanics end there. Everything difficult about the comparison happens before it and after it, never during.
Before it, somebody has to build the groups, and here is the part that decides the result. A group only exists for this purpose if the firm can assemble it from its own records. Not a group that matters, not a group anybody would name in an argument about fairness, but a constructible groupA group a firm can actually assemble from its own records, which is not the same thing as a group that matters.: one whose members the firm can list, today, from what it already holds. A firm's own records therefore decide what it is able to test, and they decide it long before anybody argues about which definition of fairness to adopt.
Picture a shopkeeper who genuinely wants to know whether he serves everybody the same way. He has one record, the till roll. The till roll shows what was bought and whether it was paid by card or in cash. He can compare card payers against cash payers and he can compare mornings against evenings. He cannot compare anything else, not because he refuses to, but because the till roll does not know anything else. If he reports that he checked and found no difference, every word of that is true and it covers two comparisons out of all the ones somebody might have wanted.
What does the record itself limit the comparison to?
Sumeru Bank Limited, invented, runs a retail loan intake chain of nine numbered components, five of them fitted to data and four of them written by people. When the bank ran the outcome distribution test on that chain, it compared outcomes across every group it could construct from what it records, and what it records does not include the attributes people usually mean by the word bias. The bank does not collect those attributes at intake. So the comparison people would ask for first was never available, whatever method the bank might have chosen, and no argument about which fairness measure to use would have made it available.
A comparison that was never available produces the most dangerous line in the whole subject, the line that says the testing found no differences. The sentence covers two completely different situations. In the first, a comparison was made and it came back level. In the second, the comparison was never possible, so nothing was found because nothing was looked at. A summary reports both of those the same way, and only the list of comparisons actually attempted tells them apart. A doctor who says the tests came back clear after running exactly one test is saying the same thing.
A firm reports that its fairness testing found no differences. What is the first question to ask?
This bank compared outcomes across every group it could construct. What defined the one difference it found?
What was the only group difference this bank found?
The bank's intake chain begins with a check on the selfie image an applicant supplies, which is component 2 and is fitted to data rather than written by anybody. In the single month these counts belong to, 10,000 applications were started and that check rejected 620 of them. The bank reviewed 200 of those rejections by hand and found that 31 of the 200 were genuine applicants, being 15.5 per cent. Applying that share to the whole 620 gives about 96 a month, and that number is an extrapolationA figure carried from a sample to the whole population it was drawn from, which is an estimate rather than a count. rather than a count. About 96 of the 10,000 who started is about 1.0 per cent of everybody who reached for a loan.
One detail about those 31 makes the finding useful. All 31 of those genuine applicants shared one condition, and it was the same condition every time: a low-light image taken on a low-specification handsetA device whose camera produces images that a learned check reads less reliably than images from a better one.. Not most of them. All 31. The difference the bank found was not defined by anything about the people; it was defined by the camera in their hand and the light in the room they were standing in.
The distinction between an attribute of people and an attribute of equipment is worth holding on to. Think of a counter where a form has to be filled in blue ink and the only pens on the table are black. Nobody wrote a rule about who may apply. Nobody at the counter intended anything. The room did it, and the room will keep doing it every day until somebody notices that the pens are the wrong colour. The condition that sorted these 31 applicants is a property of the handset's camera and of the light available, and it is not a thing any applicant was in a position to manage. O'Neil describes this pattern in Weapons of Math Destruction, 2016: a deployed component whose errors do not spread evenly over the people it decides about, but land in one part of the population and stay there.
Why could no test of the scoring model have found it?
Here is where the finding turns into a lesson about where testing points. The check that rejected those 620 sits at the very front of the chain. The image check acts before the engineA step that acts on an application before the deciding component ever sees it, so its effects never appear in that component's record.: a file it rejects never reaches the scoring model, is never scored, and appears nowhere in the record of decisions the scoring model produced. Of the 10,000 who started in the month, 86.0 per cent, being 8,600, completed onboarding and reached the decision engine. The other 1,400 did not, and they split into the 620 the image check rejected, 480 who stopped at document upload and 300 who could not complete the consent step.
Think of a hall with a gate. Somebody stands at the gate and turns people away. Inside, a register is kept of everybody who came in and what happened to them, and that register is immaculate. Every audit of the hall reads the register, finds it perfectly balanced, and reports that the hall is running properly. The people turned away at the gate are not in the register, so the more carefully the register is audited the more confidence it produces about a system that never touched them.
The difference sat in a step acting before the decision engine. What does that say about where fairness testing should point?
What does a repeatability test return before anything has changed?
The second test is the quietest of the three and the most misunderstood. A repeatability testScoring the same files twice, some distance apart in time, to see whether the answer comes back the same. takes files that have already been through the system, puts them through again some time later, and compares the two answers. Nothing else. A tailor measures a customer twice on the same afternoon and checks whether the tape says the same thing.
At this bank the reading was taken on 500 files, re-scored 30 days apart with the fitted numbers untouched. Every one of the 500 came back with the identical value: 500 out of 500. The result is clean, and being precise about what it establishes is worth the trouble. The reading establishes that the component repeats itself while its fitted numbers stay where they are. Repeatability says nothing whatsoever about whether any of those 500 answers was right. Repeatability is a reading taken rather than a property that can be read off a design, and the reading holds only until somebody changes the fitted numbers.
500 files re-scored 30 days apart returned the identical value 500 times. What has that shown?
What did the same test show after a correction?
In month 8 of this deployment an upstream income field began arriving in a different format on one channel. Monitoring flagged it in month 9, the reading step was corrected, and the same 500 files were re-scored against the corrected step. On the second reading 41 of the 500 came back with a different value, being 8.2 per cent. Nothing about the fairness of the component changed and nobody had touched the scoring model. A step feeding it had been repaired, and 8.2 per cent of the answers moved.
The pair is the teaching, so hold that 8.2 per cent against a second reading. Over the same episode 1.4 per cent of the month's 8,600 files, being 117, moved out of the accept band and into the referral band. So 8.2 per cent of scores moved and 1.4 per cent of outcomes moved: about six scores moved for every outcome that moved. Most of a score's movement happens inside a band and never crosses the line that decides anything, so far more of the component's answers shift than the number of people whose result actually changes.
The ratio of six to one needs care. The 8.2 per cent sits on a re-scored sample of 500 files and the 1.4 per cent sits on the month's 8,600, so these are two rates on two different populations rather than one count divided by another. Taken together they say something neither says alone. Reporting only the outcome figure leaves the component looking almost untouched, with no indication of how much moved inside it. Reporting only the score figure makes it sound as though one applicant in twelve was affected, and that is not what happened.
8.2 per cent of scores moved and 1.4 per cent of outcomes moved. Which figure should be reported?
What is an explainability test, and how is one actually run?
The third test has two names and they belong to the same thing. Explainability testing asks the general question of whether a reason attached to an outcome is doing any work. The explanation testRe-scoring files that have already been decided, with each named reason removed in turn, to see whether the outcome still holds. is the specific procedure this bank ran to answer it, and the procedure is short enough that anybody can run it.
Take a file that has already been decided and that carries a named reasonThe reason given to a person for the outcome they received., meaning the reason the person was actually told. Remove that reason and nothing else. Re-score the file through the same arrangement. Then ask one question: is the decline still there? If the decline disappears, that reason was carrying the outcome and the explanation was true. If the decline is still standing without it, something else was carrying the outcome and the reason was a description of the file rather than a cause of what happened to the person.
Notice how little the procedure needs. The explanation test does not need anybody to open the component. The test does not need the fitted numbers, the design, or any cooperation from whoever built the thing. All three of these tests work by re-running files and comparing outputs, and that is exactly why every one of them can be run on a component a firm bought rather than built. Needing no access is the practical reason these three are the tests that actually get run in a real building.
The bank ran that procedure over 200 declined files, removing each named reason one at a time. On 168 of them, being 84.0 per cent, the decline rested on the named reason: take the reason away and the outcome went with it. On the other 32, being 16.0 per cent, it did not. 168 plus 32 is 200.
What does it mean when a stated reason does not survive its own test?
The 32 files are the finding that matters most, and they are worse than the number makes them sound. Every one of the 200 people had been given a reason for their decline. On 32 of them, taking the reason away left the decline exactly where it stood, so the reason they were given was not what produced the outcome. Nobody told a lie. The reason was produced by the same arrangement that produced all the other reasons, sitting beside the decision and describing something true about the file, and nobody had ever tested whether describing and causing were the same thing here.
Consider a school admission. The applicant is told the application was refused because the form arrived late. The next year the form goes in three weeks early and the refusal comes again. Lateness was never what decided it. The work was done, and it was the right work on the information given. The early form changed nothing. Nobody in the school will ever know the attempt was made, and a refusal that repeats does not look any different from a refusal that stands for the first time.
Carry the reading across the month and the size of it appears. The intake chain auto-declines 688 files a month with no person touching them, and every one of those people is given a reason. At 16.0 per cent that is about 110 people a month acting on a reason that would not survive its own test. The cost lands entirely on the person who does what the letter suggested: somebody told that a declared income could not be corroborated will go and obtain a better statement, and on those files the effort changes nothing. Across months 6 to 11 alone, 688 a month is 4,128 letters. The deployment was running before month 6 as well, so 4,128 letters is a floor rather than a total.
On 32 of 200 declines the outcome stood with the reason removed. Who pays for that?
When should each of the three be run, and what does each cost?
Each of the three has a different thing that ought to make it fire, and getting those triggers wrong is how a firm ends up running the cheap one often and the useful one never. The outcome distribution test compares across groups, so it should fire whenever the population changes. A new channel, or a shift in the mix of handsets arriving, is precisely a new set of groups. Touching the fitted numbers, or repairing a step that feeds them, is the moment the previous reading stops being true. The repeatability test should fire at exactly that moment. The explanation test has no natural trigger at all, and that is exactly its problem.
On cost, what this deployment knows is narrower than what it does not. The bank's own independent validation at month 12 ran to eleven working days, and exactly one of those days went on reading the outcome distribution across the groups the bank records. One day is the only cost of the three that this deployment recorded. No cost was recorded for the repeatability reading and none for the explanation test. The shape of the cost is clear even where the count is not: all three are re-runs of files the firm already holds, so their cost is mostly the arrangement to run them, not the reading.
| Test | What it needs | What its failure looks like inside the firm |
|---|---|---|
| Outcome distribution | Two groups the record can construct, and outcomes already given | A comparison comes back different, and somebody has to decide what that means |
| Repeatability | The same files, twice, some distance apart | A count moves, and the previous reading stops being true |
| Explanation | Decided files, each named reason removable one at a time | Nothing at all. No complaint, no alert, no measure moves |
Which of the three tests can be run without any access to the inside of the component?
How does a lender, an analyst or a household read a stated reason?
Take the three readers in turn. Each one does something different with the same material. Somebody inside a lending business, reading a pack that says the fairness testing found nothing, has one useful move: ask for the list of comparisons actually attempted. Not the method, not the measure, the list. If the list is short, the finding is short, and both of those are perfectly respectable as long as they are stated together. At this bank the whole finding sat in a step that ends applications before the deciding component ever runs, so a second move is to ask which steps the testing covered.
Somebody analysing a lender from outside cannot see any of the counts, so the useful question is structural rather than numerical: does the firm state what its testing could not compare? A firm that publishes the boundary of its own comparison is stating something real about how it works. There is a related habit worth borrowing from Agrawal, Gans and Goldfarb in Prediction Machines, 2018: a fitted component supplies a prediction, and somebody still has to decide what to do with it. The reason attached to that decision is a separate artefact from the prediction, and nothing makes the two agree unless somebody tests that they do.
And a household on the other side of the letter has the least power and the clearest question. For a person refused something and given a reason, the useful thing to ask is whether fixing the stated problem would actually change the answer. The applicant is entitled to ask what would happen if the reason given were not true, and a firm that has run this test can answer that while a firm that has not cannot. The same question goes to a mechanic who says the noise is the belt: if the belt is replaced, does the noise stop?
The error that gets made, and what it costs
The tempting reading of the 31 is that the bank did something to a group of people. Follow it through and it does not survive the record. The bank found this itself, in its own review. The attribute that sorted those applicants is the class of handset and the light in the room, and neither is an attribute of anybody. Treating this as misconduct produces the wrong response: an apology, a policy line, and a check on the scoring model. The scoring model never saw one of these files, so that check would have found nothing.
The opposite error is quiet, and it is the one that produced this outcome. The error is to read a finding of nothing as evidence of nothing. The outcome distribution test came back with one difference and the natural conclusion was that the chain was broadly fine, when what the result actually said was that one comparison out of many possible ones had been run on the groups the record could build. An absence of findings is reported by almost every arrangement in exactly the same words as a finding of no difference, and the two mean opposite things.
A third error costs more than either, and it is to skip the explanation test because nothing inside the firm ever asks for it. The 32 files generate no complaint anybody can act on, no alert, and no movement in any measure the bank keeps. A control whose failure is invisible to the people funding it will be dropped in the first busy quarter, and the record will show nothing at all happening as a result. About 110 people a month, at this bank's own counts, are what that nothing consists of.
Two of these tests pass and a group still experiences worse outcomes. What has been shown?
What can none of these three tests settle?
All three of these produce comparisons, and a comparison is not a judgement. Suppose the outcome distribution test comes back with a real difference between two groups the record could build. The comparison then says that the outcomes differ. The comparison does not say whether that difference is acceptable, whether it reflects something the firm is entitled to act on, or which of the competing definitions of fairness it should be read against. The remaining questions are arguments about definitions and duties, and people settle them rather than a re-run of files.
The honest limit of all three tests sits there. The three tests show what is happening, and they are the only reliable way to find that out, but not one of them settles what should happen next. A firm that runs all three and never holds the argument has a good set of readings and no position. A firm that holds the argument without running the tests has a position built on what people assumed the system was doing.
What the reader has to confirm at source
Where a regulated lender refuses an applicant and gives a reason, and where a lender is expected to be able to account for how an outcome was reached by an automated arrangement, the applicable expectations sit with the Reserve Bank of India and are published at rbi.org.in. Where the deployer is a market intermediary rather than a bank, the expectations sit with the Securities and Exchange Board of India at sebi.gov.in. The standing international discipline from which model risk practice originates is published by the Bank for International Settlements at bis.org, and what applies to a lender in India is what the Reserve Bank of India states rather than what the international material says. Read the current position at the source.
The three tests, the groups the record could construct and every count attached to them are Sumeru Bank Limited's own choices. A choice one lender makes is not a threshold, a requirement or an effective date set by any authority.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Published expectations on a regulated lender covering digital lending, outsourcing, data and consent, and what a firm should be able to account for where an automated arrangement decides a customer outcome. The expectations stated here are the ones that apply to a lender in India | rbi.org.in |
| Securities and Exchange Board of India | Published expectations where the deployer of such an arrangement is a market intermediary rather than a bank | sebi.gov.in |
| Bank for International Settlements | International supervisory material from which the standing discipline of model risk work originates, named as an origin rather than as the position in India | bis.org |
| O'Neil | Weapons of Math Destruction, 2016, named in the text where a deployed component's errors are shown falling unevenly across the people it decides about. Not quoted | named in the text |
| Agrawal, Gans and Goldfarb | Prediction Machines, 2018, named in the text where a fitted component supplies a prediction that a person still has to act on. Not quoted | named in the text |
Sumeru Bank Limited is invented.
Educational material. Not advice on any investment, tax, budget or market position.
