Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
AI, Automation & Digital Finance
1AI Foundations
Artificial Intelligence in FinanceAlgorithmNeural Networks and Deep LearningMachine LearningArtificial Intelligence vs Machine…Computer Vision in FinanceTraining Data and LabelsNatural Language Processing in Finance
2Generative AI
Generative AIGenerative AI vs Predictive AILarge Language ModelsEmbeddingsHallucinationFine TuningPrompting vs Fine TuningThe PromptThe Context WindowTool CallingGroundingVector DatabasesRetrieval Augmented GenerationRAG vs Fine Tuning
3Automation and Workflow
Workflow AutomationAutomation vs AugmentationHow to Map a…Straight-Through Processing and Exception…Robotic Process AutomationRule EnginesMachine Learning vs Rule-Based…
4Document and Operations AI
Intelligent Document ProcessingBatch vs Real-Time vs…Document Classification vs Entity…Service Level AgreementsCase ManagementHow to Document Data…Reconciliation AutomationOptical Character Recognition and Data ExtractionConfidence Scores
5Customer Systems, Identity and Digital Assets
Digital IdentityConsent ManagementBlockchain and Distributed LedgerChatbots and Conversational AIFrom Use Case to ProductionDigital Assets and TokenisationDigital SignaturesData Sharing in FinanceElectronic KYC and Digital Onboarding
6Credit and Fraud Systems
The Fraud AlertCredit Decisioning SystemsHuman in the Loop…Adverse ActionAnomaly DetectionThe Decision ThresholdCredit Score vs Credit DecisionAlert Triage and EscalationFraud Detection and Transaction MonitoringFraud Model vs Credit ModelHow to Build Human…
7Governance, Data and Vendors
AI Governance and the AI PolicyHow to Create an…Explainability and Interpretability ComparedThe AI VendorBias and Fairness in Financial AIShadow AIAccess Control and Data MinimisationCloud Computing in FinanceData Lineage and Master DataData ResidencyThe AI Use Case Register and Model InventoryThe Model Owner
8Model Performance, Monitoring and Resilience
Model DriftFalse Positives and False NegativesClassification MetricsAdversarial AttacksModel TestingBias, Fairness and Explainability…Stopping an Automated SystemModel ValidationAI Governance vs Model Risk ManagementPrompt InjectionHow to Create an…

Bias, Fairness and Explainability Testing

These three tests compare outcomes across groups, ask whether the same file scored twice returns the same answer, and check whether the reason a person was given is the reason that produced their outcome. The firm's own records fix what can be compared. At one invented bank the single group difference found was defined by equipment, in a step the scoring model never touched.

One fact carries everything after it, so start with what one of these tests actually is. A fairness test is not a property a system has or lacks. A fairness test is not a calculation somebody performs quietly on a model and reports as a score. A fairness test is a question somebody asks of a system that is already deciding about people, with a population in mind, a comparison in hand and a record that either supports the comparison or does not. Somebody chooses the question. Somebody finds the two groups. Somebody reads the answer and decides what happens next. Every cost in fairness testing attaches to one of those three people.

Bias itself, where it comes from, and why the competing definitions of fairness cannot all be satisfied at once, are set out under bias and fairness in financial AI. The asking is what matters now: what each test compares, what each one costs, and who bears the answer when it comes back.

Every relationship here is a count on a fixed population: 200 rejections reviewed, 500 files re-scored, 200 declines tested one reason at a time. Nothing in it is a variable that anybody at this bank ever moved. The three tests are drawn side by side, and the two readings that carry movement are drawn as a movement and as a shape.

What does each of these three tests actually ask?

Consider the weighing scale at a neighbourhood shop. Three complaints could be made about it, and they are not the same complaint. The first is that it seems to weigh differently depending on who is standing at the counter. The second is that it gives two different readings for the same bag of rice, weighed twice within a minute. The third is that when the shopkeeper says the price went up because of transport, transport is not actually why. The three complaints need three different investigations, and settling any one of them leaves the other two exactly where they were.

The three tests are those three investigations, run on a system that decides about people rather than on a scale. The first compares outcomes across groups. The second compares one file against itself. The third compares a stated reason against the outcome it is supposed to explain. A firm that runs one of them and reports that it has tested for fairness has answered one of three questions and left the reader to assume it answered all of them.

THREE TESTS, THREE DIFFERENT QUESTIONS 1 THE OUTCOME DISTRIBUTION TEST IT ASKS Do outcomes differ across the groups this firm can build? IT RETURNS A comparison, group against group, on outcomes given. IT CANNOT SAY Whether a difference that shows up is acceptable. 2 THE REPEATABILITY TEST IT ASKS Does the same file, scored twice, get the same answer? IT RETURNS A count of files whose value moved between two runs. IT CANNOT SAY Whether either of the two answers was the right one. 3 THE EXPLANATION TEST IT ASKS Did the reason a person was given produce the outcome? IT RETURNS A count of files where the outcome survived removal. IT CANNOT SAY Whether the reason was a good one to give at all. Answering one of the three leaves the other two exactly where they were. Educational illustration. Sumeru Bank Limited and every count here are invented.
Each of the three tests answers a question the other two never touch, so a firm reporting that it tested for fairness has settled one of three arguments and left the reader to assume it settled all of them.
Try it out

What does the explanation test actually do to a file?

What is fairness testing actually comparing, and across which groups?

The first of the three is the outcome distribution testA comparison of outcomes across groups the firm can assemble from what it already records., and it is the one people mean when they say the words. The procedure is short enough to state in a sentence. Take the outcomes a system has already produced, split the people who received them into two or more groups, and compare the shares. The mechanics end there. Everything difficult about the comparison happens before it and after it, never during.

Before it, somebody has to build the groups, and here is the part that decides the result. A group only exists for this purpose if the firm can assemble it from its own records. Not a group that matters, not a group anybody would name in an argument about fairness, but a constructible groupA group a firm can actually assemble from its own records, which is not the same thing as a group that matters.: one whose members the firm can list, today, from what it already holds. A firm's own records therefore decide what it is able to test, and they decide it long before anybody argues about which definition of fairness to adopt.

Picture a shopkeeper who genuinely wants to know whether he serves everybody the same way. He has one record, the till roll. The till roll shows what was bought and whether it was paid by card or in cash. He can compare card payers against cash payers and he can compare mornings against evenings. He cannot compare anything else, not because he refuses to, but because the till roll does not know anything else. If he reports that he checked and found no difference, every word of that is true and it covers two comparisons out of all the ones somebody might have wanted.

What does the record itself limit the comparison to?

Sumeru Bank Limited, invented, runs a retail loan intake chain of nine numbered components, five of them fitted to data and four of them written by people. When the bank ran the outcome distribution test on that chain, it compared outcomes across every group it could construct from what it records, and what it records does not include the attributes people usually mean by the word bias. The bank does not collect those attributes at intake. So the comparison people would ask for first was never available, whatever method the bank might have chosen, and no argument about which fairness measure to use would have made it available.

A comparison that was never available produces the most dangerous line in the whole subject, the line that says the testing found no differences. The sentence covers two completely different situations. In the first, a comparison was made and it came back level. In the second, the comparison was never possible, so nothing was found because nothing was looked at. A summary reports both of those the same way, and only the list of comparisons actually attempted tells them apart. A doctor who says the tests came back clear after running exactly one test is saying the same thing.

WHAT THE RECORD ALLOWS, AND WHAT A SUMMARY HIDES CONSTRUCTIBLE FROM THIS BANK'S OWN RECORD Everything else somebody might want to compare. The attributes people usually mean are not collected at intake at all. THE CUTS THIS RECORD CAN ACTUALLY MAKE the channel the class of handset the kind of document the hour of the day the branch geography TWO SUMMARIES THAT READ EXACTLY THE SAME Fairness testing: no differences found. The comparison was made, on groups that existed, and it came back level. Fairness testing: no differences found. The comparison was never possible, so nothing was found or looked at. Educational illustration. Sumeru Bank Limited and every count here are invented.
What a firm writes down sets the outer boundary of what it can compare, so a finding of nothing is very often a fact about the record rather than a fact about the outcomes.
Try it out

A firm reports that its fairness testing found no differences. What is the first question to ask?

Try it out

This bank compared outcomes across every group it could construct. What defined the one difference it found?

What was the only group difference this bank found?

The bank's intake chain begins with a check on the selfie image an applicant supplies, which is component 2 and is fitted to data rather than written by anybody. In the single month these counts belong to, 10,000 applications were started and that check rejected 620 of them. The bank reviewed 200 of those rejections by hand and found that 31 of the 200 were genuine applicants, being 15.5 per cent. Applying that share to the whole 620 gives about 96 a month, and that number is an extrapolationA figure carried from a sample to the whole population it was drawn from, which is an estimate rather than a count. rather than a count. About 96 of the 10,000 who started is about 1.0 per cent of everybody who reached for a loan.

THREE STEPS, THREE SCALES, ONE REVIEW SCALE ONE: THE 10,000 APPLICATIONS STARTED IN THE MONTH 620 rejected by the check on the selfie image, before any scoring happened SCALE TWO: THE 620 THAT CHECK REJECTED 200 reviewed 420 not reviewed about 96 genuine applicants, carried across from the review, an extrapolation SCALE THREE: THE 200 REJECTIONS REVIEWED BY HAND 31 169 were correctly rejected 31 of 200 is 15.5 per cent, and 15.5 per cent of 620 is about 96. About 96 of the 10,000 who started is about 1.0 per cent, told no before the engine. Educational illustration. Sumeru Bank Limited and every count here are invented.
Reading a sample of rejections by hand and carrying the share across to the whole gives about 96 wrongly rejected applicants a month, which is roughly one in every hundred people who began an application.

One detail about those 31 makes the finding useful. All 31 of those genuine applicants shared one condition, and it was the same condition every time: a low-light image taken on a low-specification handsetA device whose camera produces images that a learned check reads less reliably than images from a better one.. Not most of them. All 31. The difference the bank found was not defined by anything about the people; it was defined by the camera in their hand and the light in the room they were standing in.

The distinction between an attribute of people and an attribute of equipment is worth holding on to. Think of a counter where a form has to be filled in blue ink and the only pens on the table are black. Nobody wrote a rule about who may apply. Nobody at the counter intended anything. The room did it, and the room will keep doing it every day until somebody notices that the pens are the wrong colour. The condition that sorted these 31 applicants is a property of the handset's camera and of the light available, and it is not a thing any applicant was in a position to manage. O'Neil describes this pattern in Weapons of Math Destruction, 2016: a deployed component whose errors do not spread evenly over the people it decides about, but land in one part of the population and stay there.

WHERE THE 31 GENUINE APPLICANTS SAT Counts are the genuine applicants wrongly rejected, from the review of 200. A HIGHER-SPECIFICATION HANDSET A LOWER-SPECIFICATION HANDSET ORDINARY LIGHT IN THE ROOM LOW LIGHT IN THE ROOM 0 0 0 31 every one of them, in this cell alone The shared condition belongs to the camera and to the room, not to the people. Educational illustration. Sumeru Bank Limited and every count here are invented.
All 31 wrongly rejected applicants fell into a single cell of the grid, and the attribute that defines that cell is equipment and lighting rather than anything about the applicants themselves.
AI For Finance Bootcamp — Fin Maverick

Why could no test of the scoring model have found it?

Here is where the finding turns into a lesson about where testing points. The check that rejected those 620 sits at the very front of the chain. The image check acts before the engineA step that acts on an application before the deciding component ever sees it, so its effects never appear in that component's record.: a file it rejects never reaches the scoring model, is never scored, and appears nowhere in the record of decisions the scoring model produced. Of the 10,000 who started in the month, 86.0 per cent, being 8,600, completed onboarding and reached the decision engine. The other 1,400 did not, and they split into the 620 the image check rejected, 480 who stopped at document upload and 300 who could not complete the consent step.

Think of a hall with a gate. Somebody stands at the gate and turns people away. Inside, a register is kept of everybody who came in and what happened to them, and that register is immaculate. Every audit of the hall reads the register, finds it perfectly balanced, and reports that the hall is running properly. The people turned away at the gate are not in the register, so the more carefully the register is audited the more confidence it produces about a system that never touched them.

WHERE A TEST OF THE SCORING MODEL CAN LOOK EVERYTHING A TEST OF THE SCORING MODEL CAN SEE 8,600 reached the decision engine 1,400 A MONTH NEVER GOT THIS FAR 620 rejected by the check on the selfie image, before anything was scored 480 stopped at document upload 300 could not complete the consent step 8,600 reached the decision engine, which is 86.0 per cent of the 10,000 About 96 of those 620 were genuine applicants. A test pointed at the scoring model cannot see a single one of them, however well it is run. Educational illustration. Sumeru Bank Limited and every count here are invented.
Testing the component everybody worries about would have found nothing here, because the difference sat two steps upstream of it and its victims never entered that component's record at all.
Try it out

The difference sat in a step acting before the decision engine. What does that say about where fairness testing should point?

What does a repeatability test return before anything has changed?

The second test is the quietest of the three and the most misunderstood. A repeatability testScoring the same files twice, some distance apart in time, to see whether the answer comes back the same. takes files that have already been through the system, puts them through again some time later, and compares the two answers. Nothing else. A tailor measures a customer twice on the same afternoon and checks whether the tape says the same thing.

At this bank the reading was taken on 500 files, re-scored 30 days apart with the fitted numbers untouched. Every one of the 500 came back with the identical value: 500 out of 500. The result is clean, and being precise about what it establishes is worth the trouble. The reading establishes that the component repeats itself while its fitted numbers stay where they are. Repeatability says nothing whatsoever about whether any of those 500 answers was right. Repeatability is a reading taken rather than a property that can be read off a design, and the reading holds only until somebody changes the fitted numbers.

Try it out

500 files re-scored 30 days apart returned the identical value 500 times. What has that shown?

What did the same test show after a correction?

In month 8 of this deployment an upstream income field began arriving in a different format on one channel. Monitoring flagged it in month 9, the reading step was corrected, and the same 500 files were re-scored against the corrected step. On the second reading 41 of the 500 came back with a different value, being 8.2 per cent. Nothing about the fairness of the component changed and nobody had touched the scoring model. A step feeding it had been repaired, and 8.2 per cent of the answers moved.

THE SAME 500 FILES, READ TWICE READING ONE 500 files re-scored 30 days apart, with the fitted numbers unchanged READING TWO the same 500 files re-scored after the month 9 correction 0 of 500 different. 500 of 500 identical. 41 of 500 different, being 8.2 per cent. Educational illustration. Sumeru Bank Limited and every count here are invented.
Repeating itself perfectly for a month says nothing about what a component does after the data feeding it is repaired, and here the repair moved one answer in every twelve.

The pair is the teaching, so hold that 8.2 per cent against a second reading. Over the same episode 1.4 per cent of the month's 8,600 files, being 117, moved out of the accept band and into the referral band. So 8.2 per cent of scores moved and 1.4 per cent of outcomes moved: about six scores moved for every outcome that moved. Most of a score's movement happens inside a band and never crosses the line that decides anything, so far more of the component's answers shift than the number of people whose result actually changes.

The ratio of six to one needs care. The 8.2 per cent sits on a re-scored sample of 500 files and the 1.4 per cent sits on the month's 8,600, so these are two rates on two different populations rather than one count divided by another. Taken together they say something neither says alone. Reporting only the outcome figure leaves the component looking almost untouched, with no indication of how much moved inside it. Reporting only the score figure makes it sound as though one applicant in twelve was affected, and that is not what happened.

HOW MUCH MOVED, AND HOW MANY PEOPLE MOVED SCORES THAT MOVED 41 of 500 re-scored 8.2 per cent OUTCOMES THAT MOVED 117 of the month's 8,600 1.4 per cent 0 2 4 6 8 10 PER CENT OF ITS OWN POPULATION About six scores moved for every outcome that moved. Two rates on two different populations: 500 files re-scored, and the month's 8,600 decided. The comparison is between two rates and is never one count divided by the other. Educational illustration. Sumeru Bank Limited and every count here are invented.
Scores move about six times as often as outcomes do, so quoting only the outcome figure hides how much shifted and quoting only the score figure overstates how many people were affected.
Try it out

8.2 per cent of scores moved and 1.4 per cent of outcomes moved. Which figure should be reported?

Breaking Into Quants Bootcamp — Fin Maverick

What is an explainability test, and how is one actually run?

The third test has two names and they belong to the same thing. Explainability testing asks the general question of whether a reason attached to an outcome is doing any work. The explanation testRe-scoring files that have already been decided, with each named reason removed in turn, to see whether the outcome still holds. is the specific procedure this bank ran to answer it, and the procedure is short enough that anybody can run it.

Take a file that has already been decided and that carries a named reasonThe reason given to a person for the outcome they received., meaning the reason the person was actually told. Remove that reason and nothing else. Re-score the file through the same arrangement. Then ask one question: is the decline still there? If the decline disappears, that reason was carrying the outcome and the explanation was true. If the decline is still standing without it, something else was carrying the outcome and the reason was a description of the file rather than a cause of what happened to the person.

Notice how little the procedure needs. The explanation test does not need anybody to open the component. The test does not need the fitted numbers, the design, or any cooperation from whoever built the thing. All three of these tests work by re-running files and comparing outputs, and that is exactly why every one of them can be run on a component a firm bought rather than built. Needing no access is the practical reason these three are the tests that actually get run in a real building.

THE EXPLANATION TEST, DECIDED ONE FILE AT A TIME 1 TAKE A DECIDED FILE with the reason already written on it 2 REMOVE THE REASON that one named reason, and nothing else 3 RE-SCORE THE FILE through the very same arrangement as before Ask one question: is the decline still there? IT DISAPPEARED IT IS STILL THERE The decline rested on that reason. Removing it changed the outcome, so the reason was doing the work it claimed. 168 files, 84.0 per cent The decline did not rest on it. Removing it changed nothing, so the reason described the file rather than caused it. 32 files, 16.0 per cent THE 200 DECLINED FILES TESTED, TO SCALE 168 32 Educational illustration. 168 plus 32 is 200. Sumeru Bank Limited and every count here are invented.
The whole procedure is remove the reason, re-run the file, and see whether the outcome still stands, which needs no access to the inside of the component and is settled one file at a time.

The bank ran that procedure over 200 declined files, removing each named reason one at a time. On 168 of them, being 84.0 per cent, the decline rested on the named reason: take the reason away and the outcome went with it. On the other 32, being 16.0 per cent, it did not. 168 plus 32 is 200.

What does it mean when a stated reason does not survive its own test?

The 32 files are the finding that matters most, and they are worse than the number makes them sound. Every one of the 200 people had been given a reason for their decline. On 32 of them, taking the reason away left the decline exactly where it stood, so the reason they were given was not what produced the outcome. Nobody told a lie. The reason was produced by the same arrangement that produced all the other reasons, sitting beside the decision and describing something true about the file, and nobody had ever tested whether describing and causing were the same thing here.

Consider a school admission. The applicant is told the application was refused because the form arrived late. The next year the form goes in three weeks early and the refusal comes again. Lateness was never what decided it. The work was done, and it was the right work on the information given. The early form changed nothing. Nobody in the school will ever know the attempt was made, and a refusal that repeats does not look any different from a refusal that stands for the first time.

WHO BEARS A REASON THAT DID NOT SURVIVE ITS OWN TEST WHO WHAT IT COSTS THEM WHERE IT IS COUNTED The bank Nothing. The decline itself was correct. Nothing to count The exception desk Nothing. These files never reach it. Nothing to count The applicant The effort of fixing a stated problem that was never the one deciding it Nowhere at all CARRYING THE READING ACROSS THE MONTH, AND THEN ACROSS THE DEPLOYMENT 32 of the 200 declined files tested 16.0 per cent of the reasons tested about 110 people a month, of the 688 auto-declines at least 4,128 letters across months 6 to 11, a floor Educational illustration. Sumeru Bank Limited and every count here are invented.
A reason that fails its own test costs only the person who acts on it, and that is the one party whose costs appear in no measure the bank keeps.

Carry the reading across the month and the size of it appears. The intake chain auto-declines 688 files a month with no person touching them, and every one of those people is given a reason. At 16.0 per cent that is about 110 people a month acting on a reason that would not survive its own test. The cost lands entirely on the person who does what the letter suggested: somebody told that a declared income could not be corroborated will go and obtain a better statement, and on those files the effort changes nothing. Across months 6 to 11 alone, 688 a month is 4,128 letters. The deployment was running before month 6 as well, so 4,128 letters is a floor rather than a total.

Try it out

On 32 of 200 declines the outcome stood with the reason removed. Who pays for that?

Financial Analyst Program Bootcamp — Fin Maverick

When should each of the three be run, and what does each cost?

Each of the three has a different thing that ought to make it fire, and getting those triggers wrong is how a firm ends up running the cheap one often and the useful one never. The outcome distribution test compares across groups, so it should fire whenever the population changes. A new channel, or a shift in the mix of handsets arriving, is precisely a new set of groups. Touching the fitted numbers, or repairing a step that feeds them, is the moment the previous reading stops being true. The repeatability test should fire at exactly that moment. The explanation test has no natural trigger at all, and that is exactly its problem.

WHAT MAKES EACH TEST FIRE, AND WHAT IT COST HERE THE TEST WHEN IT OUGHT TO FIRE, AND WHEN IT DID WHAT IT COST The outcome distribution test Whenever the population changes. Here it waited until after go-live. month 4 month 12 1 working day inside the 11 day validation The repeatability test Whenever the fitted numbers move, or a step feeding them is repaired. Here it did. either side of the month 9 correction No cost is recorded for it at this bank The explanation test Nothing makes it fire on its own, so it has to be put in the calendar by somebody. no trigger, and no date for it in the record No cost is recorded for it at this bank run not run, though this is where it belonged Two of the three costs here are not recorded. Educational illustration. Sumeru Bank Limited and every count here are invented.
The test whose failure nobody inside the firm ever feels is also the only one of the three with nothing to make it fire, so it has to be put in somebody's calendar deliberately.

On cost, what this deployment knows is narrower than what it does not. The bank's own independent validation at month 12 ran to eleven working days, and exactly one of those days went on reading the outcome distribution across the groups the bank records. One day is the only cost of the three that this deployment recorded. No cost was recorded for the repeatability reading and none for the explanation test. The shape of the cost is clear even where the count is not: all three are re-runs of files the firm already holds, so their cost is mostly the arrangement to run them, not the reading.

TestWhat it needsWhat its failure looks like inside the firm
Outcome distributionTwo groups the record can construct, and outcomes already givenA comparison comes back different, and somebody has to decide what that means
RepeatabilityThe same files, twice, some distance apartA count moves, and the previous reading stops being true
ExplanationDecided files, each named reason removable one at a timeNothing at all. No complaint, no alert, no measure moves
Try it out

Which of the three tests can be run without any access to the inside of the component?

Bond Pricing and Yield Mechanics — free micro-course from Fin Maverick

How does a lender, an analyst or a household read a stated reason?

Take the three readers in turn. Each one does something different with the same material. Somebody inside a lending business, reading a pack that says the fairness testing found nothing, has one useful move: ask for the list of comparisons actually attempted. Not the method, not the measure, the list. If the list is short, the finding is short, and both of those are perfectly respectable as long as they are stated together. At this bank the whole finding sat in a step that ends applications before the deciding component ever runs, so a second move is to ask which steps the testing covered.

Somebody analysing a lender from outside cannot see any of the counts, so the useful question is structural rather than numerical: does the firm state what its testing could not compare? A firm that publishes the boundary of its own comparison is stating something real about how it works. There is a related habit worth borrowing from Agrawal, Gans and Goldfarb in Prediction Machines, 2018: a fitted component supplies a prediction, and somebody still has to decide what to do with it. The reason attached to that decision is a separate artefact from the prediction, and nothing makes the two agree unless somebody tests that they do.

And a household on the other side of the letter has the least power and the clearest question. For a person refused something and given a reason, the useful thing to ask is whether fixing the stated problem would actually change the answer. The applicant is entitled to ask what would happen if the reason given were not true, and a firm that has run this test can answer that while a firm that has not cannot. The same question goes to a mechanic who says the noise is the belt: if the belt is replaced, does the noise stop?

The error that gets made, and what it costs

The tempting reading of the 31 is that the bank did something to a group of people. Follow it through and it does not survive the record. The bank found this itself, in its own review. The attribute that sorted those applicants is the class of handset and the light in the room, and neither is an attribute of anybody. Treating this as misconduct produces the wrong response: an apology, a policy line, and a check on the scoring model. The scoring model never saw one of these files, so that check would have found nothing.

The opposite error is quiet, and it is the one that produced this outcome. The error is to read a finding of nothing as evidence of nothing. The outcome distribution test came back with one difference and the natural conclusion was that the chain was broadly fine, when what the result actually said was that one comparison out of many possible ones had been run on the groups the record could build. An absence of findings is reported by almost every arrangement in exactly the same words as a finding of no difference, and the two mean opposite things.

A third error costs more than either, and it is to skip the explanation test because nothing inside the firm ever asks for it. The 32 files generate no complaint anybody can act on, no alert, and no movement in any measure the bank keeps. A control whose failure is invisible to the people funding it will be dropped in the first busy quarter, and the record will show nothing at all happening as a result. About 110 people a month, at this bank's own counts, are what that nothing consists of.

Try it out

Two of these tests pass and a group still experiences worse outcomes. What has been shown?

The list of comparisons attempted is the finding. See what the testing covered.

What can none of these three tests settle?

All three of these produce comparisons, and a comparison is not a judgement. Suppose the outcome distribution test comes back with a real difference between two groups the record could build. The comparison then says that the outcomes differ. The comparison does not say whether that difference is acceptable, whether it reflects something the firm is entitled to act on, or which of the competing definitions of fairness it should be read against. The remaining questions are arguments about definitions and duties, and people settle them rather than a re-run of files.

The honest limit of all three tests sits there. The three tests show what is happening, and they are the only reliable way to find that out, but not one of them settles what should happen next. A firm that runs all three and never holds the argument has a good set of readings and no position. A firm that holds the argument without running the tests has a position built on what people assumed the system was doing.

India

What the reader has to confirm at source

Where a regulated lender refuses an applicant and gives a reason, and where a lender is expected to be able to account for how an outcome was reached by an automated arrangement, the applicable expectations sit with the Reserve Bank of India and are published at rbi.org.in. Where the deployer is a market intermediary rather than a bank, the expectations sit with the Securities and Exchange Board of India at sebi.gov.in. The standing international discipline from which model risk practice originates is published by the Bank for International Settlements at bis.org, and what applies to a lender in India is what the Reserve Bank of India states rather than what the international material says. Read the current position at the source.

The three tests, the groups the record could construct and every count attached to them are Sumeru Bank Limited's own choices. A choice one lender makes is not a threshold, a requirement or an effective date set by any authority.

Bias, its sources, and the conflict between the competing definitions of fairness are set out under bias and fairness in financial AI. How a component is fitted is set out under machine learning, and how one is evaluated under classification metrics. How an explanation is produced at all for a component whose working cannot be read is set out under explainability and interpretability compared. Testing before go-live in general, and the seven pre-deployment tests it runs through, is set out under model testing. Watching a running system is set out under model drift and monitoring, independent challenge by somebody who did not build the thing is set out separately, and what a firm does once something has been flagged is set out under stopping an automated system. Every count belongs to one invented bank and one deployment.

Sources

SourceDocumentSite
Reserve Bank of IndiaPublished expectations on a regulated lender covering digital lending, outsourcing, data and consent, and what a firm should be able to account for where an automated arrangement decides a customer outcome. The expectations stated here are the ones that apply to a lender in Indiarbi.org.in
Securities and Exchange Board of IndiaPublished expectations where the deployer of such an arrangement is a market intermediary rather than a banksebi.gov.in
Bank for International SettlementsInternational supervisory material from which the standing discipline of model risk work originates, named as an origin rather than as the position in Indiabis.org
O'NeilWeapons of Math Destruction, 2016, named in the text where a deployed component's errors are shown falling unevenly across the people it decides about. Not quotednamed in the text
Agrawal, Gans and GoldfarbPrediction Machines, 2018, named in the text where a fitted component supplies a prediction that a person still has to act on. Not quotednamed in the text

Sumeru Bank Limited is invented.
Educational material. Not advice on any investment, tax, budget or market position.

Covered in this topic

Subtopics

Fairness TestingExplainability Test
← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.