Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
AI, Automation & Digital Finance
1AI Foundations
Artificial Intelligence in FinanceAlgorithmNeural Networks and Deep LearningMachine LearningArtificial Intelligence vs Machine…Computer Vision in FinanceTraining Data and LabelsNatural Language Processing in Finance
2Generative AI
Generative AIGenerative AI vs Predictive AILarge Language ModelsEmbeddingsHallucinationFine TuningPrompting vs Fine TuningThe PromptThe Context WindowTool CallingGroundingVector DatabasesRetrieval Augmented GenerationRAG vs Fine Tuning
3Automation and Workflow
Workflow AutomationAutomation vs AugmentationHow to Map a…Straight-Through Processing and Exception…Robotic Process AutomationRule EnginesMachine Learning vs Rule-Based…
4Document and Operations AI
Intelligent Document ProcessingBatch vs Real-Time vs…Document Classification vs Entity…Service Level AgreementsCase ManagementHow to Document Data…Reconciliation AutomationOptical Character Recognition and Data ExtractionConfidence Scores
5Customer Systems, Identity and Digital Assets
Digital IdentityConsent ManagementBlockchain and Distributed LedgerChatbots and Conversational AIFrom Use Case to ProductionDigital Assets and TokenisationDigital SignaturesData Sharing in FinanceElectronic KYC and Digital Onboarding
6Credit and Fraud Systems
The Fraud AlertCredit Decisioning SystemsHuman in the Loop…Adverse ActionAnomaly DetectionThe Decision ThresholdCredit Score vs Credit DecisionAlert Triage and EscalationFraud Detection and Transaction MonitoringFraud Model vs Credit ModelHow to Build Human…
7Governance, Data and Vendors
AI Governance and the AI PolicyHow to Create an…Explainability and Interpretability ComparedThe AI VendorBias and Fairness in Financial AIShadow AIAccess Control and Data MinimisationCloud Computing in FinanceData Lineage and Master DataData ResidencyThe AI Use Case Register and Model InventoryThe Model Owner
8Model Performance, Monitoring and Resilience
Model DriftFalse Positives and False NegativesClassification MetricsAdversarial AttacksModel TestingBias, Fairness and Explainability…Stopping an Automated SystemModel ValidationAI Governance vs Model Risk ManagementPrompt InjectionHow to Create an…

Bias and Fairness in Financial AI: Where They Conflict

Bias in a financial component is a pattern in its errors, not an attitude. Bias arrives three ways: from the population the component was fitted on, from the data it receives in use, and from the choice of what to measure. A component can be biased about something it never sees. The thing it does see stands in for it, and that is the case a firm is least equipped to notice.

One mechanism sits underneath every case in this guide, and it is worth holding on to before the detail starts. A component measures something. The measurement is connected to a characteristic of the people being measured. The error rate on that measurement differs by characteristic. The characteristic was deliberately kept out of the inputs and never reached the record either, so nobody looks at the errors split that way. Each step is ordinary. The four together produce a result nobody chose and nobody can see.

What does bias actually mean in a component, and what does it not mean?

Start with a machine anybody can walk up to. A shop has an automatic door with a sensor set at the chest height of an average adult. The door opens for most people. The door stays shut for a small child, for someone in a wheelchair, or for anyone carrying a tall box in front of them. Nobody set out to keep those people outside. The sensor was placed at one height, and the people it fails are the people whose chests are not at that height. Bias is the pattern in which people the door fails, not an attitude inside the door.

BiasA pattern in a component's errors, where they fall more heavily on one group than another. in a deployed component means the same thing and nothing more. Bias is a statement about errors, it needs a groupAny set of people a firm might later be asked whether it treated differently. The set does not have to be a characteristic anybody recorded. named before it can be checked, and it is found by counting rather than by asking anyone what they meant. Cathy O'Neil's Weapons of Math Destruction, 2016, put this shape in front of a general reader: errors that do not fall randomly but land repeatedly on the same people, and land hardest exactly where nobody is counting. Landing hardest where nobody is counting is the part a bank has to take personally.

Two things bias does not mean. Bias does not mean intent, so a component built with total care by people of good will can carry one, and a review that asks whether anybody meant it has asked a question with no useful answer. And it does not mean the component is bad overall. A component can be right 95 times in 100 and still put almost all of the other five on the same doorstep. The headline reading and the pattern in the misses are separate facts, and a firm that publishes only the first has published the half that will never embarrass it.

One more distinction, and it is the one people get wrong most often. A difference in refusal rates between two groups is not by itself bias. Two groups can genuinely differ, and a component that refuses more of one may be reading something real. A difference in wrong refusal rates is bias. A wrong refusal is a case a person reviewing it would have accepted, so there is nothing real left for the difference to be reading. Sorting all the refusals says something about the population. Sorting the wrong ones says something about the component.

Try it out

What is bias in a financial component?

AI For Finance Bootcamp — Fin Maverick

Where does bias come from, and why do the three sources need different fixes?

In a deployed chain there are exactly three places it can enter, and a review that cannot say which one it is looking at cannot recommend anything useful. Source 1 is the past the component was built from: the fitting populationThe past cases a component's behaviour was derived from. How a component is fitted, and how that population is chosen, is covered separately., assembled by decisions somebody else made years earlier. Source 2 is the data arriving now, at the moment of capture, from real people on real handsets in real rooms. Source 3 is the choice of what to measure and what counts as a right answer. The choice is made in a meeting and written into a definition.

The three sources need different fixes, and that is why naming the source matters more than naming the pattern. A pattern from source 1 is addressed by changing what the component was built from. A pattern from source 3 is addressed by changing a definition. Changing a definition costs nothing but a decision and a great deal of argument. A pattern from source 2 is already present in the data before the component sees anything, so it cannot be fixed inside the component at all. Two of the three live inside the model. One lives out at the front door, and it is the one that produced the case examined here.

THREE SOURCES, THREE DIFFERENT REMEDIES SOURCE 1 THE FITTING POPULATION Enters before the component exists, inside the past cases it was built from. FIXED BY Changing what it was built from, then building again. SOURCE 2 THE DATA IN USE Enters at the moment of capture, in the handset, the room and the connection. FIXED BY Nothing inside the component. A second route, or capture. SOURCE 3 WHAT IS MEASURED Enters in a definition: what counts as a bad outcome, and over how long it is watched. FIXED BY Changing the definition, and saying so where it is used. Sources 1 and 3 can be addressed by people who work on the component. Source 2 is already inside the data before the component sees anything, so it is answered at the front door or it is not answered.
Three routes in, and only two of them can be closed by anybody working on the component itself, which is why a review names the route before it recommends a remedy.

What does the population a component was built from already decide?

Source 1 comes first, and it is the most familiar of the three. Sumeru Bank Limited, an invented lender, built its scoring model on a past window holding 300,000 applications. Of those, 240,000 were accepted and their outcomes are observable, being 80.0 per cent. The other 60,000 were declined and have no observable outcome at all, being 20.0 per cent. A loan that was never made cannot repay or fail. The component therefore learned the shape of a book that an earlier set of decisions had already filtered.

Think of a tailor who has stitched for one build of customer for twenty years and is excellent at it. Hand him a shape he has never cut for and the confidence is unchanged while the fit is not. The bank's own measured version of this is the channel mix: 62 per cent of the fitting examples came from a single channel, and the readings split 94.1 per cent on that channel against 87.9 per cent on the others, a gap of 6.2 points. Same component, same month, two populations, six points apart.

A fitting population decides not what the component believes, but which people it has any evidence about at all. How a component is fitted, how a fitting population is chosen and how either is validated is covered separately. The part that belongs here is narrower and colder: whatever the older process refused is missing from the record, so the new component inherits the older one's blind spots without inheriting any note that says so.

What arrives in the data once the component is already running?

Source 2 is the source a firm is least equipped to notice. The intake chain begins with a liveness check on a selfie image, component 2 in this bank's numbering, and what a liveness check asserts is covered separately. In the month under review it saw all 10,000 applications and rejected 620 of them.

Split those 10,000 by the handset the application arrived from. 2,400 came from a low-specification handset and 7,600 did not, and 2,400 plus 7,600 is 10,000. The 620 rejections split 384 and 236. Against their own group sizes that is 16.0 per cent of the low-specification applications and 3.1 per cent of the others, about 5.2 times the rejection rate, and 384 plus 236 comes back to 620. So far this is only a difference in refusal rates, and by the distinction drawn earlier it is not yet bias. A gap that size is a reason to look.

The bank looked. Two hundred of the rejections were pulled and reviewed by hand, and 31 of the 200 turned out to be genuine applicants, being 15.5 per cent. Extrapolated across the 620 that is about 96 people in the month. And every one of the 31 shared a single condition: a low-light image taken on a low-specification handset. Read that as a fact about a camera and a room. A small sensor behind a small lens, in a room lit by one bulb after dark, produces a darker and grainier image than a larger sensor near a window at four in the afternoon, and it produces it whoever is holding the handset and however carefully they hold it.

Now put the review's finding back against the two groups. Concentrated on the low-specification group, about 96 wrong rejectionsCases the component refused that a person reviewing the same file would have accepted. against 2,400 applications is 4.0 per cent of those applicants refused in error, against 0.0 per cent of the other 7,600. Not one of the 31 came from outside that group. The refusal gap was a reason to look and the wrong-refusal gap is the finding, and only the second one is a statement about the component rather than about the applicants. Ninety-six people a month is not a rounding error. Ninety-six households asked a bank for money, were told no by a step no person ever saw, and were right to have asked.

PANEL 1 REJECTED BY THE CHECK, PER 100 APPLICATIONS FROM EACH GROUP LOW-SPECIFICATION HANDSET, 2,400 APPLICATIONS 16.0 per cent, being 384 EVERY OTHER HANDSET, 7,600 APPLICATIONS 3.1 per cent, being 236 0 5 10 15 20 per cent of that group PANEL 2 WRONG REJECTIONS, PER 100 APPLICATIONS FROM EACH GROUP LOW-SPECIFICATION HANDSET 4.0 per cent, about 96 people a month EVERY OTHER HANDSET 0.0 per cent, none of the 31 came from here 0 5 10 15 20 per cent of that group
Sixteen per cent against three point one is a reason to look, and four per cent against nothing at all is the finding, because a refusal a reviewer would have overturned has no remaining innocent explanation.

Can a component be biased about something it never sees?

Component 2 has never been given a handset model. The handset is not an input and it is not in the file. Nobody at the bank could have removed a field that was never there. The component receives one image. And an image carries the handset inside it whether anybody intended that or not, in the grain, the exposure and the sharpness of the edges. A proxySomething a component does see that stands in for something it does not. A proxy need not resemble the hidden thing; it only has to move with it. is exactly that: something visible that moves with something invisible.

Everyone has met this outside finance. A school that never asks a candidate's postcode, but does ask how far they can travel each morning, has asked for the postcode in a way that sounds neutral. An employer that never asks about caring responsibilities, but does ask whether a candidate can start at seven, has done the same. Nothing dishonest happened in either case, and the outcome is identical to the outcome of asking directly. Being blind to a characteristicNever receiving it as an input. Blindness is not the same as being unaffected by the characteristic, and not the same as being unable to act on it. removes it from the record, not from the result, and only the second of those is what anybody will ask about afterwards.

The pattern O'Neil describes is exactly this: a measurement standing in for something nobody chose to measure, errors landing on the group behind that stand-in, and no counting anywhere that would reveal it. The additional cruelty of the blind design is that it also destroys the evidence. Because the handset was never recorded against the application, the bank could not have run this analysis from its own files at all. Somebody had to open 200 rejected applications, look at 200 images, and notice that 31 of them looked alike.

WHAT COMES WITH THE IMAGE UNINVITED THE HANDSET sensor size, lens THE ROOM one bulb, after dark IMAGE QUALITY grain, exposure, edges COMPONENT 2 receives one image PASS OR FAIL NEVER AN INPUT, AND THEREFORE NEVER IN THE RECORD EITHER The errors follow image quality. Image quality follows the handset and the room. So the errors follow the handset, which the component is never given and which nobody wrote down against the application. THE SAME DESIGN THAT MAKES IT BLIND MAKES THE PATTERN INVISIBLE IN THE FILES
Component 2 is given an image and never a handset, and every one of the 31 wrong rejections found in review was a low-light image taken on a low-specification handset.
Try it out

A component never receives the handset an application came from. Can it be biased about it?

Breaking Into Quants Bootcamp — Fin Maverick

What does the choice of what to measure quietly decide?

Source 3 is the least visible of the three, and it lives in no data at all but in a definition somebody wrote once. The bank's scoring model needed to know what counted as a bad outcome, and four numbered choices were made: what counts as bad, set at 90 days past due; the observation window, set at 12 months; the population, being accepted applications only; and accounts closed early, excluded. Every one of the four is the bank's own choice and none of them is a standard.

Watch what the second choice alone does to the measured rate on the same 240,000 cases.

Observation windowCases carrying a bad labelMeasured bad rate
6 months5,0402.1 per cent
12 months, the bank's choice8,1603.4 per cent
18 months10,0804.2 per cent
24 months11,5204.8 per cent

Nothing about any borrower changed between those four rows. The same people did the same things with the same loans, and the measured rate more than doubled between the first row and the last because somebody moved a definition. The choice of what to count decides what can be found, and it decides it before a single figure has been collected. A household knows this instinctively: spending measured over a week that includes a wedding supports one conclusion, spending measured over the following week supports the opposite, and neither week lied.

The second measurement choice in this guide is the one that nearly hid the whole case. The liveness review was commissioned about rejection volumes. The question put to it was how many rejections were wrong, and it answered that question well: 31 of 200, being 15.5 per cent, extrapolating to about 96 of the 620. The question it was not asked was who the wrong ones happened to. The pattern appeared only because somebody laid the 31 files beside each other and noticed they looked alike. A measurement designed to count refusals will count refusals faithfully and will never, however carefully it is run, produce a statement about a group.

What are the four definitions of fairness, and what does each one demand?

Here is where the subject stops being intuitive. FairnessA stated property somebody wants a component to have. There are several, they are not the same property, and they make different demands. is not one property that a component either has or lacks. Fairness is at least four properties, all of them defensible, and all of them definitions in use rather than anybody's standard. The formal treatment, and the statistics behind each one, is covered separately. Each definition below has a demand, a cost, and somebody who pays that cost.

Definition 1, equal acceptance. The same share of each group is accepted. Definition 1 demands nothing about who those people are, only how many. Its cost is direct: where the groups genuinely differ in outcome, meeting it means accepting people the evidence expects to fail, and that cost lands on the lender's book and, through pricing over time, on every borrower including the ones the definition was meant to help.

Definition 2, equal accuracy. The component is about as accurate for one group as for another. Its cost is mostly a cost of measurement, and it is a real one. The definition cannot be checked without recording the group, and recording the group sits directly against the principle of collecting the least data needed. The tension is genuine and remains unresolved, and how little a chain may collect and who may see it is set out under access control and data minimisation.

Definition 3, equal acceptance among those who would repay. Among the people who would have been fine, the same share of each group gets in. Definition 3 is the one most people mean when they say a lender should be fair. Its cost is that it is completely silent about the headline acceptance rate, so a firm can satisfy it exactly while accepting 96.6 per cent of one group and 93.2 per cent of another, and the group carrying the higher underlying rate bears that difference.

Definition 4, the same inputs produce the same output. Two identical applications get identical answers whatever group they came from. Its cost is close to nothing, and almost every firm can claim it for precisely that reason. A component that never receives the characteristic satisfies it by construction, on the day it is switched on, without anybody testing anything.

THE DEFINITION WHAT IT DEMANDS WHAT IT COSTS, AND WHO PAYS COMPONENT 2 1 EQUAL ACCEPTANCE The same share of each group gets in. Counts only, nothing about who they are. Accepting people the evidence expects to fail. Paid by the book. FAILS 2 EQUAL ACCURACY About as accurate for one group as another. Errors counted group by group, not overall. Recording the group in order to check it, against collecting the least possible. FAILS 3 EQUAL ACCEPTANCE AMONG THOSE WHO PASS Nobody good refused. Nothing at all about the headline rates. Two acceptance rates that stay apart. Paid by the group with the higher rate. FAILS on the sample 4 SAME INPUTS, SAME OUTPUT Identical files, identical answers. That the group is not an input. That is all. Almost nothing, which is why every firm can claim it. HOLDS
Four definitions demand four different things, and the only one this component satisfies is the one that costs nothing to satisfy and proves nothing once satisfied.
Try it out

Name the four definitions of fairness in the order this guide sets them out.

Which definition does a component satisfy by construction, and why is that not comfort?

Component 2 satisfies definition 4 perfectly, and it is worth stating the reason precisely because the reason is the whole trap. Component 2 never receives the handset, so two identical images cannot receive different answers on the basis of a handset. Send the same image twice and the same answer comes back twice. There is no test to run, no sample to draw and no reviewer to appoint. The property is a consequence of the design, true on day one and true forever, and it says nothing whatever about the month just gone.

Now hold that beside the measurement. In the same month, on the same component, one group was refused at 16.0 per cent and another at 3.1, and the refusals that a reviewer overturned ran at 4.0 per cent of one group and 0.0 per cent of the other. Both of those statements are true at once, and a firm that reports the first and never measures the second has not been dishonest, it has arranged not to find out.

There is a house-buying version of this that people feel immediately. A landlord who says he treats every application identically, and only ever advertises by word of mouth among his existing tenants, is telling the truth about his process and saying nothing at all about who ends up living there. The process test passes. The outcome was decided before the process began.

Try it out

A component gives the same output for the same inputs whoever sent them. Is it fair?

The comfortable reading, and what it costs

The comfortable reading is that a component which never sees a characteristic cannot be biased about it, and it is half right in a way that makes it more dangerous than being simply wrong. Component 2 does satisfy definition 4 exactly. Component 2 also sees image quality, image quality follows the handset and the room, and the errors follow the image quality, so a component blind to the characteristic produced a wrong refusal rate of 4.0 per cent against 0.0 per cent across it.

The cost is not only the roughly 96 people a month. The larger cost is that nobody was looking. The review that found the pattern was commissioned about rejection volumes rather than about groups, and the pattern surfaced only when somebody put the 31 files side by side. The handset was never written down against an application, so the bank could not have found the pattern in its own records at any point.

The other half deserves saying just as plainly. Sumeru Bank commissioned the review itself, its own reviewer found the pattern, and the recommendation that came back was to build the second route that had never existed. The commission, the finding and the recommendation are the process working. Everything before them failed.

Financial Analyst Program Bootcamp — Fin Maverick

Where do two of the definitions conflict, and what does the arithmetic look like?

Definitions 1 and 3 cannot both hold once the underlying rateHow often the outcome actually occurs in a group, before any component sees anybody. The underlying rate is a fact about the world, not about the component. differs between the groups. The conflict is not a defect anybody built in, and no amount of care in construction removes it. The conflict is arithmetic, and almost every reader arrives believing it is avoidable.

Take two groups of 1,000 people each. In group A, 34 of the 1,000 would go bad, being 3.4 per cent, the bank's own locked figure from its fitting work. In group B, twice that many would, being 68 and 6.8 per cent. State the boundary before the arithmetic: this is an illustration, not a measurement. Sumeru Bank Limited never measured a bad rate by group at all, and the second rate is arithmetic chosen to make the shape visible.

Now run a component that does exactly what definition 3 asks, accepting every person who would not go bad and nobody who would. In group A it accepts 966, being 96.6 per cent. In group B it accepts 932, being 93.2 per cent. Definition 3 holds perfectly by construction. Definition 1 fails, and it fails by exactly the difference in the underlying rates, 3.4 percentage points. Now force definition 1. Getting group B up to 96.6 per cent means accepting 966 of them instead of 932, or 34 more people. The component already took everybody who would repay, so all 34 of those extra acceptances go bad. The two definitions are not competing preferences, they are two demands on one set of numbers, and satisfying either exactly makes the other fail by a known amount.

ACCEPTANCE UNDER DEFINITION 3, PER CENT, AS THE UNDERLYING RATES SEPARATE 98 96 94 92 90 GROUP A 96.6 GROUP B 93.2 the illustration used here THE GAP IS THE DIFFERENCE, POINT FOR POINT 0 1 2 3 4 5 6 DIFFERENCE IN UNDERLYING BAD RATE, PERCENTAGE POINTS Illustration on the invented bank's own locked base rate of 3.4 per cent. The second rate is arithmetic, not an observation.
The two lines meet only where the underlying rates are identical, and everywhere else the gap between them equals the difference in those rates exactly, point for point.
Try it out

Before the control moves: two groups, one going bad twice as often as the other. Can a component accept the same share of each and still accept only the people who would repay?

Play with it

Separate the two groups, then try to force the acceptance rates equal

One control: how far group B's underlying bad rate sits above group A's, from 0 to 6 percentage points. Two things redraw: the composition of each group of 1,000, and the two acceptance markers on the scale underneath. The two buttons switch which definition is being enforced, so the cost moves on the screen rather than being described in words. At the default of 3.4 points, acceptance is 96.6 per cent against 93.2 per cent, and forcing the two equal means accepting 34 more people per thousand in group B, every one of whom goes bad. At a difference of 0 both read 96.6 per cent and both definitions hold at once, and that is the only setting on the whole scale where they do.

0.0 points apart3.4 points apart6.0 points apart
TWO GROUPS OF 1,000, UNDER THE DEFINITION BEING ENFORCED GROUP A, 3.4 PER CENT WOULD GO BAD 966 in GROUP B, 6.8 PER CENT WOULD GO BAD 932 in accepted and repays accepted and goes bad refused WHERE THE TWO ACCEPTANCE RATES SIT, WHICH IS THE DEFINITION 1 TEST A 96.6 B 93.2 88 100 per cent Group A is held at the invented bank's own locked bad rate of 3.4 per cent throughout. The second rate is illustrative arithmetic.
Group A accepted
96.6%
Group B accepted
93.2%
Extra per 1,000 to force equal
34
Of those, go bad
34
At a difference of 3.4 percentage points, a component accepting exactly the people who would repay takes 96.6 per cent of group A and 93.2 per cent of group B. Definition 3 holds exactly and definition 1 does not, and closing that gap would mean accepting 34 more people per thousand in group B, every one of whom goes bad.
Educational illustration. Two groups of 1,000 people, a component that accepts exactly those who would not go bad, and group A held at the invented bank's own locked bad rate of 3.4 per cent. THE SECOND GROUP'S RATE IS ILLUSTRATIVE ARITHMETIC AND NOT AN OBSERVATION: this bank never measured a bad rate by group. Equal acceptance and equal acceptance among those who would repay meet at one point only, where the two underlying rates are identical.
Investment Banking Analyst Bootcamp — Fin Maverick

How would this be found in a component already in use?

The method is smaller than people expect and that is its whole virtue. Take the refused cases from one period. Have a person who knows the work go through a sample of them and mark which ones were wrong, meaning cases they would have accepted with the same file in front of them. Then sort only the wrong ones by any characteristic the component never received. That is it. Sorting the wrong ones by an unrecorded characteristic is the whole of the method, and everything before it is an ordinary file review.

Two things make it work. The first is that the method needs no access to the component's internals. Component 2 runs inside a vendor-hosted service, so the bank can see an image go out and a single word come back and nothing in between, a limit set out under the AI vendor. Finding a pattern in what the component did requires no opening of the component. The second is the ordering: the person marks the file wrong before anybody sorts by group, so the judgement cannot be influenced by the very pattern under test.

Be honest about what a sample can carry. Two hundred of 620 is a decent sample and 31 is a small number of wrong cases, so the right claim from this review is not that the wrong-rejection rate is exactly 15.5 per cent. The right claim is that all 31 shared one condition, a statement about concentration rather than about a rate, and concentration is what survives a small sample. The extrapolation to about 96 people a month is a reasonable estimate and is stated as one.

THE REVIEW SHEET, WITH ONE COLUMN ADDED FILE WHAT THE COMPONENT RETURNED A PERSON, SAME FILE ADDED: THE CAPTURE CONDITIONS R-0041 Fail WRONG, genuine low light, low-specification R-0058 Fail Agreed, not genuine daylight, other handset R-0092 Fail WRONG, genuine low light, low-specification R-0117 Fail Agreed, not genuine daylight, other handset R-0143 Fail WRONG, genuine low light, low-specification The reviewer marks each file before anybody sorts by the added column, so the judgement cannot be shaped by the pattern being tested for. Sorting every refusal describes the population that applied; SORTING ONLY THE WRONG ONES IS THE PART THAT DESCRIBES THE COMPONENT.
Take the refused cases, have a person say which were wrong, and sort those wrong ones by a characteristic the component never received.
Try it out

What does it take to find out whether a component already in use does this?

What can be done once it has been found, and what cannot?

Everything depends on which of the three sources produced it. A pattern from source 1 is addressed by changing the population the component was built from and building again, and how that is done is covered separately. A pattern from source 3 is addressed by changing a definition or by adding the missing measurement, and the second of those is usually the honest answer.

A pattern from source 2 cannot be addressed inside the component at all, and this is the part practitioners resist hardest. Rebuilding component 2 on more low-light images might help the reading a little. Rebuilding does not change the fact that a darker, grainier image carries less information than a bright, sharp one, and no component recovers information the image does not hold. The remedy sits at the capture: a second route for anyone the check refuses, guidance on screen before the image is taken, a retry with a different framing, or a person. A pattern that arrives with the data as it is captured is fixed at the capture or it is not fixed.

Three things cannot be done, and saying so is not defeatism. A characteristic cannot be removed from the inputs in a way that also removes its effect. A proxy does the same work with better manners. Definitions 1 and 3 cannot be satisfied together where the underlying rates differ, at any level of care. And definition 2 cannot be checked without recording the group, and recording the group cuts directly against collecting as little as possible, a tension set out under access control and data minimisation. The three limits are the boundary of the craft, and a review that promises to clear all of them has promised something arithmetic does not allow.

WHICH SOURCE PRODUCED IT? ask this before recommending anything SOURCE 1, THE POPULATION Change what it was built from, then build again. INSIDE THE COMPONENT Slow, and it works. SOURCE 3, WHAT IS MEASURED Change the definition, or add the missing measure. INSIDE THE COMPONENT Cheap, and much argued over. SOURCE 2, THE DATA IN USE A second route, guidance before capture, or a person. NOT INSIDE THE COMPONENT No refit recovers the missing light. Two branches end at people who work on the component. The third ends at the front door, and the case examined here is the third.
A fitting population can be rebuilt and a measurement can be redefined, while a pattern arriving in the data as it is captured is fixed at the capture or not at all.
Try it out

The pattern comes from how the image was captured rather than from the fitting. What can be changed inside the component?

Building a Revenue Forecast From Drivers — free micro-course from Fin Maverick

What does the firm owe the group the errors fall on?

Three concrete acts, and none of them is an apology. The first is a second route: anybody the check refuses gets another way to reach a decision. At this bank no such route existed in the month under review, and providing one is precisely what the reviewer recommended. A learned component produces a prediction, not a decision. Agrawal, Gans and Goldfarb's Prediction Machines, 2018, draws that line, and somebody still has to act on the prediction. A second route is what acting on it looks like when the prediction is wrong.

The second is notice. About 96 people a month were refused by a step that was mistaken about them, and a second route that exists only for future applicants leaves everybody already refused exactly where they were. Telling people who were refused that another route now exists is unglamorous, cheap, and the only part of this that reaches the people it was about.

The third is the change that prevents the next occurrence: measure the refusal rate by group from then on, rather than by volume. Volume measurement is what hid this in the first place, so a firm that keeps measuring volume while regretting the pattern has changed nothing at all. Only the third act stops this happening again, and it is the one that costs a report rather than a project. Under the bank's own numbered duties for a named accountable person, seeing the monitoring is duty 4, and at the month 12 validation the named person for the scoring model could not evidence that duty because the monitoring existed and went to the team that built the chain rather than to her.

WHAT IS OWED, IN THREE ACTS 1 2 3 A SECOND ROUTE Anyone the check refuses gets another way to a decision. NOTICE The people already refused are told the route now exists. MEASURE BY GROUP Refusals split by group from now on, rather than counted. Acts 1 and 2 repair what happened. Act 3 is the only one that prevents the next occurrence, and none of the three is an apology.
Give the group a second route, tell the people already refused that one exists, and measure the refusal rate by group from then on rather than by volume.
Try it out

About 96 genuine applicants a month were refused in error, concentrated on one group. What does the firm owe them?

The Risk Management Program bootcamp teaches you to set a limit framework and run it through a breach. Bond Pricing and Yield Mechanics — free micro-course from Fin Maverick

How does somebody actually use this on a working Monday?

For a lender, this becomes three questions put to whoever is accountable for a component, and they take about ten minutes to ask. Which of the three sources has the owner looked for? What do the wrong refusals from last quarter look like, sorted by something the component never received? And which of the four definitions can the owner evidence, as opposed to satisfy by construction? A component owner who can answer the third question with a number rather than a design argument has done the work.

For an analyst reading somebody else's disclosures, the useful test is a subtraction. Count the fairness statements that are about the process and the fairness statements that are about outcomes, and a firm reporting only the first has described what its design does rather than what its month did. For a household on the other side of the screen, the practical translation is short. Where an automated step refuses an application, ask whether a second route exists. Why the refusal happened is often unanswerable, and whether a second route exists always has an answer.

Somebody always asks about cost. The review that found this pattern was a person opening 200 files, days rather than months of effort. Set that against a chain that cost the invented bank Rs 2,40,00,000/- to build once and Rs 65,00,000/- a year to run, and the measurement is not the expensive part of anything described here. The expensive part is a month of refusals nobody can now describe and a route back that did not exist when it was needed.

India

Who would characterise a pattern like this, and where is that written?

Whether a pattern of this kind is permitted, and what a regulated lender must do about one, is for the Reserve Bank of India. The Reserve Bank publishes its expectations on digital lending, customer data, consent, outsourcing and record keeping at rbi.org.in. Where the deployer is a market intermediary rather than a lender, the equivalent expectations sit with the Securities and Exchange Board of India at sebi.gov.in. Where the question is what a board is accountable for, the Ministry of Corporate Affairs publishes at mca.gov.in. The four definitions are definitions in use in the craft and are not anybody's standard. Requirements, thresholds and effective dates move, and the issuing body's own site carries the current position.

How a component is built from past cases, how a fitting population is chosen, and how either is validated and evaluated are covered separately. A liveness check, and the front door of the application journey it sits in, is covered separately. The discipline of reviewing and challenging a model inside a firm is covered separately. How little a chain may collect and who may see it is set out under access control and data minimisation, together with the tension between measuring by group and collecting as little as possible.
Document Extraction in Finance teaches you to design an extraction pipeline for a financial document and set the confidence threshold honestly.

Sources

SourceDocumentSite
Reserve Bank of IndiaPublished expectations on a regulated lender covering digital lending, customer data, consent, outsourcing and record keepingrbi.org.in
Securities and Exchange Board of IndiaEquivalent expectations where the deployer of an automated decisioning step is a market intermediarysebi.gov.in
Ministry of Corporate AffairsMaterial on what a board is accountable for where an automated step affects customersmca.gov.in
Cathy O'NeilWeapons of Math Destruction, 2016, on errors that do not fall randomly but land repeatedly on the same people, and hardest where nobody is countingCrown Publishing Group
Ajay Agrawal, Joshua Gans and Avi GoldfarbPrediction Machines, 2018, on the line between a prediction and the decision somebody still has to makeHarvard Business Review Press

Sumeru Bank Limited is invented.
Educational material. Not advice on any investment, tax, budget or market position.

Covered in this topic

Subtopics

Fairness
← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.