Bias and Fairness in Financial AI: Where They Conflict
Bias in a financial component is a pattern in its errors, not an attitude. Bias arrives three ways: from the population the component was fitted on, from the data it receives in use, and from the choice of what to measure. A component can be biased about something it never sees. The thing it does see stands in for it, and that is the case a firm is least equipped to notice.
One mechanism sits underneath every case in this guide, and it is worth holding on to before the detail starts. A component measures something. The measurement is connected to a characteristic of the people being measured. The error rate on that measurement differs by characteristic. The characteristic was deliberately kept out of the inputs and never reached the record either, so nobody looks at the errors split that way. Each step is ordinary. The four together produce a result nobody chose and nobody can see.
What does bias actually mean in a component, and what does it not mean?
Start with a machine anybody can walk up to. A shop has an automatic door with a sensor set at the chest height of an average adult. The door opens for most people. The door stays shut for a small child, for someone in a wheelchair, or for anyone carrying a tall box in front of them. Nobody set out to keep those people outside. The sensor was placed at one height, and the people it fails are the people whose chests are not at that height. Bias is the pattern in which people the door fails, not an attitude inside the door.
BiasA pattern in a component's errors, where they fall more heavily on one group than another. in a deployed component means the same thing and nothing more. Bias is a statement about errors, it needs a groupAny set of people a firm might later be asked whether it treated differently. The set does not have to be a characteristic anybody recorded. named before it can be checked, and it is found by counting rather than by asking anyone what they meant. Cathy O'Neil's Weapons of Math Destruction, 2016, put this shape in front of a general reader: errors that do not fall randomly but land repeatedly on the same people, and land hardest exactly where nobody is counting. Landing hardest where nobody is counting is the part a bank has to take personally.
Two things bias does not mean. Bias does not mean intent, so a component built with total care by people of good will can carry one, and a review that asks whether anybody meant it has asked a question with no useful answer. And it does not mean the component is bad overall. A component can be right 95 times in 100 and still put almost all of the other five on the same doorstep. The headline reading and the pattern in the misses are separate facts, and a firm that publishes only the first has published the half that will never embarrass it.
One more distinction, and it is the one people get wrong most often. A difference in refusal rates between two groups is not by itself bias. Two groups can genuinely differ, and a component that refuses more of one may be reading something real. A difference in wrong refusal rates is bias. A wrong refusal is a case a person reviewing it would have accepted, so there is nothing real left for the difference to be reading. Sorting all the refusals says something about the population. Sorting the wrong ones says something about the component.
What is bias in a financial component?
Where does bias come from, and why do the three sources need different fixes?
In a deployed chain there are exactly three places it can enter, and a review that cannot say which one it is looking at cannot recommend anything useful. Source 1 is the past the component was built from: the fitting populationThe past cases a component's behaviour was derived from. How a component is fitted, and how that population is chosen, is covered separately., assembled by decisions somebody else made years earlier. Source 2 is the data arriving now, at the moment of capture, from real people on real handsets in real rooms. Source 3 is the choice of what to measure and what counts as a right answer. The choice is made in a meeting and written into a definition.
The three sources need different fixes, and that is why naming the source matters more than naming the pattern. A pattern from source 1 is addressed by changing what the component was built from. A pattern from source 3 is addressed by changing a definition. Changing a definition costs nothing but a decision and a great deal of argument. A pattern from source 2 is already present in the data before the component sees anything, so it cannot be fixed inside the component at all. Two of the three live inside the model. One lives out at the front door, and it is the one that produced the case examined here.
What does the population a component was built from already decide?
Source 1 comes first, and it is the most familiar of the three. Sumeru Bank Limited, an invented lender, built its scoring model on a past window holding 300,000 applications. Of those, 240,000 were accepted and their outcomes are observable, being 80.0 per cent. The other 60,000 were declined and have no observable outcome at all, being 20.0 per cent. A loan that was never made cannot repay or fail. The component therefore learned the shape of a book that an earlier set of decisions had already filtered.
Think of a tailor who has stitched for one build of customer for twenty years and is excellent at it. Hand him a shape he has never cut for and the confidence is unchanged while the fit is not. The bank's own measured version of this is the channel mix: 62 per cent of the fitting examples came from a single channel, and the readings split 94.1 per cent on that channel against 87.9 per cent on the others, a gap of 6.2 points. Same component, same month, two populations, six points apart.
A fitting population decides not what the component believes, but which people it has any evidence about at all. How a component is fitted, how a fitting population is chosen and how either is validated is covered separately. The part that belongs here is narrower and colder: whatever the older process refused is missing from the record, so the new component inherits the older one's blind spots without inheriting any note that says so.
What arrives in the data once the component is already running?
Source 2 is the source a firm is least equipped to notice. The intake chain begins with a liveness check on a selfie image, component 2 in this bank's numbering, and what a liveness check asserts is covered separately. In the month under review it saw all 10,000 applications and rejected 620 of them.
Split those 10,000 by the handset the application arrived from. 2,400 came from a low-specification handset and 7,600 did not, and 2,400 plus 7,600 is 10,000. The 620 rejections split 384 and 236. Against their own group sizes that is 16.0 per cent of the low-specification applications and 3.1 per cent of the others, about 5.2 times the rejection rate, and 384 plus 236 comes back to 620. So far this is only a difference in refusal rates, and by the distinction drawn earlier it is not yet bias. A gap that size is a reason to look.
The bank looked. Two hundred of the rejections were pulled and reviewed by hand, and 31 of the 200 turned out to be genuine applicants, being 15.5 per cent. Extrapolated across the 620 that is about 96 people in the month. And every one of the 31 shared a single condition: a low-light image taken on a low-specification handset. Read that as a fact about a camera and a room. A small sensor behind a small lens, in a room lit by one bulb after dark, produces a darker and grainier image than a larger sensor near a window at four in the afternoon, and it produces it whoever is holding the handset and however carefully they hold it.
Now put the review's finding back against the two groups. Concentrated on the low-specification group, about 96 wrong rejectionsCases the component refused that a person reviewing the same file would have accepted. against 2,400 applications is 4.0 per cent of those applicants refused in error, against 0.0 per cent of the other 7,600. Not one of the 31 came from outside that group. The refusal gap was a reason to look and the wrong-refusal gap is the finding, and only the second one is a statement about the component rather than about the applicants. Ninety-six people a month is not a rounding error. Ninety-six households asked a bank for money, were told no by a step no person ever saw, and were right to have asked.
Can a component be biased about something it never sees?
Component 2 has never been given a handset model. The handset is not an input and it is not in the file. Nobody at the bank could have removed a field that was never there. The component receives one image. And an image carries the handset inside it whether anybody intended that or not, in the grain, the exposure and the sharpness of the edges. A proxySomething a component does see that stands in for something it does not. A proxy need not resemble the hidden thing; it only has to move with it. is exactly that: something visible that moves with something invisible.
Everyone has met this outside finance. A school that never asks a candidate's postcode, but does ask how far they can travel each morning, has asked for the postcode in a way that sounds neutral. An employer that never asks about caring responsibilities, but does ask whether a candidate can start at seven, has done the same. Nothing dishonest happened in either case, and the outcome is identical to the outcome of asking directly. Being blind to a characteristicNever receiving it as an input. Blindness is not the same as being unaffected by the characteristic, and not the same as being unable to act on it. removes it from the record, not from the result, and only the second of those is what anybody will ask about afterwards.
The pattern O'Neil describes is exactly this: a measurement standing in for something nobody chose to measure, errors landing on the group behind that stand-in, and no counting anywhere that would reveal it. The additional cruelty of the blind design is that it also destroys the evidence. Because the handset was never recorded against the application, the bank could not have run this analysis from its own files at all. Somebody had to open 200 rejected applications, look at 200 images, and notice that 31 of them looked alike.
A component never receives the handset an application came from. Can it be biased about it?
What does the choice of what to measure quietly decide?
Source 3 is the least visible of the three, and it lives in no data at all but in a definition somebody wrote once. The bank's scoring model needed to know what counted as a bad outcome, and four numbered choices were made: what counts as bad, set at 90 days past due; the observation window, set at 12 months; the population, being accepted applications only; and accounts closed early, excluded. Every one of the four is the bank's own choice and none of them is a standard.
Watch what the second choice alone does to the measured rate on the same 240,000 cases.
| Observation window | Cases carrying a bad label | Measured bad rate |
|---|---|---|
| 6 months | 5,040 | 2.1 per cent |
| 12 months, the bank's choice | 8,160 | 3.4 per cent |
| 18 months | 10,080 | 4.2 per cent |
| 24 months | 11,520 | 4.8 per cent |
Nothing about any borrower changed between those four rows. The same people did the same things with the same loans, and the measured rate more than doubled between the first row and the last because somebody moved a definition. The choice of what to count decides what can be found, and it decides it before a single figure has been collected. A household knows this instinctively: spending measured over a week that includes a wedding supports one conclusion, spending measured over the following week supports the opposite, and neither week lied.
The second measurement choice in this guide is the one that nearly hid the whole case. The liveness review was commissioned about rejection volumes. The question put to it was how many rejections were wrong, and it answered that question well: 31 of 200, being 15.5 per cent, extrapolating to about 96 of the 620. The question it was not asked was who the wrong ones happened to. The pattern appeared only because somebody laid the 31 files beside each other and noticed they looked alike. A measurement designed to count refusals will count refusals faithfully and will never, however carefully it is run, produce a statement about a group.
What are the four definitions of fairness, and what does each one demand?
Here is where the subject stops being intuitive. FairnessA stated property somebody wants a component to have. There are several, they are not the same property, and they make different demands. is not one property that a component either has or lacks. Fairness is at least four properties, all of them defensible, and all of them definitions in use rather than anybody's standard. The formal treatment, and the statistics behind each one, is covered separately. Each definition below has a demand, a cost, and somebody who pays that cost.
Definition 1, equal acceptance. The same share of each group is accepted. Definition 1 demands nothing about who those people are, only how many. Its cost is direct: where the groups genuinely differ in outcome, meeting it means accepting people the evidence expects to fail, and that cost lands on the lender's book and, through pricing over time, on every borrower including the ones the definition was meant to help.
Definition 2, equal accuracy. The component is about as accurate for one group as for another. Its cost is mostly a cost of measurement, and it is a real one. The definition cannot be checked without recording the group, and recording the group sits directly against the principle of collecting the least data needed. The tension is genuine and remains unresolved, and how little a chain may collect and who may see it is set out under access control and data minimisation.
Definition 3, equal acceptance among those who would repay. Among the people who would have been fine, the same share of each group gets in. Definition 3 is the one most people mean when they say a lender should be fair. Its cost is that it is completely silent about the headline acceptance rate, so a firm can satisfy it exactly while accepting 96.6 per cent of one group and 93.2 per cent of another, and the group carrying the higher underlying rate bears that difference.
Definition 4, the same inputs produce the same output. Two identical applications get identical answers whatever group they came from. Its cost is close to nothing, and almost every firm can claim it for precisely that reason. A component that never receives the characteristic satisfies it by construction, on the day it is switched on, without anybody testing anything.
Name the four definitions of fairness in the order this guide sets them out.
Which definition does a component satisfy by construction, and why is that not comfort?
Component 2 satisfies definition 4 perfectly, and it is worth stating the reason precisely because the reason is the whole trap. Component 2 never receives the handset, so two identical images cannot receive different answers on the basis of a handset. Send the same image twice and the same answer comes back twice. There is no test to run, no sample to draw and no reviewer to appoint. The property is a consequence of the design, true on day one and true forever, and it says nothing whatever about the month just gone.
Now hold that beside the measurement. In the same month, on the same component, one group was refused at 16.0 per cent and another at 3.1, and the refusals that a reviewer overturned ran at 4.0 per cent of one group and 0.0 per cent of the other. Both of those statements are true at once, and a firm that reports the first and never measures the second has not been dishonest, it has arranged not to find out.
There is a house-buying version of this that people feel immediately. A landlord who says he treats every application identically, and only ever advertises by word of mouth among his existing tenants, is telling the truth about his process and saying nothing at all about who ends up living there. The process test passes. The outcome was decided before the process began.
A component gives the same output for the same inputs whoever sent them. Is it fair?
The comfortable reading, and what it costs
The comfortable reading is that a component which never sees a characteristic cannot be biased about it, and it is half right in a way that makes it more dangerous than being simply wrong. Component 2 does satisfy definition 4 exactly. Component 2 also sees image quality, image quality follows the handset and the room, and the errors follow the image quality, so a component blind to the characteristic produced a wrong refusal rate of 4.0 per cent against 0.0 per cent across it.
The cost is not only the roughly 96 people a month. The larger cost is that nobody was looking. The review that found the pattern was commissioned about rejection volumes rather than about groups, and the pattern surfaced only when somebody put the 31 files side by side. The handset was never written down against an application, so the bank could not have found the pattern in its own records at any point.
The other half deserves saying just as plainly. Sumeru Bank commissioned the review itself, its own reviewer found the pattern, and the recommendation that came back was to build the second route that had never existed. The commission, the finding and the recommendation are the process working. Everything before them failed.
Where do two of the definitions conflict, and what does the arithmetic look like?
Definitions 1 and 3 cannot both hold once the underlying rateHow often the outcome actually occurs in a group, before any component sees anybody. The underlying rate is a fact about the world, not about the component. differs between the groups. The conflict is not a defect anybody built in, and no amount of care in construction removes it. The conflict is arithmetic, and almost every reader arrives believing it is avoidable.
Take two groups of 1,000 people each. In group A, 34 of the 1,000 would go bad, being 3.4 per cent, the bank's own locked figure from its fitting work. In group B, twice that many would, being 68 and 6.8 per cent. State the boundary before the arithmetic: this is an illustration, not a measurement. Sumeru Bank Limited never measured a bad rate by group at all, and the second rate is arithmetic chosen to make the shape visible.
Now run a component that does exactly what definition 3 asks, accepting every person who would not go bad and nobody who would. In group A it accepts 966, being 96.6 per cent. In group B it accepts 932, being 93.2 per cent. Definition 3 holds perfectly by construction. Definition 1 fails, and it fails by exactly the difference in the underlying rates, 3.4 percentage points. Now force definition 1. Getting group B up to 96.6 per cent means accepting 966 of them instead of 932, or 34 more people. The component already took everybody who would repay, so all 34 of those extra acceptances go bad. The two definitions are not competing preferences, they are two demands on one set of numbers, and satisfying either exactly makes the other fail by a known amount.
Before the control moves: two groups, one going bad twice as often as the other. Can a component accept the same share of each and still accept only the people who would repay?
Separate the two groups, then try to force the acceptance rates equal
One control: how far group B's underlying bad rate sits above group A's, from 0 to 6 percentage points. Two things redraw: the composition of each group of 1,000, and the two acceptance markers on the scale underneath. The two buttons switch which definition is being enforced, so the cost moves on the screen rather than being described in words. At the default of 3.4 points, acceptance is 96.6 per cent against 93.2 per cent, and forcing the two equal means accepting 34 more people per thousand in group B, every one of whom goes bad. At a difference of 0 both read 96.6 per cent and both definitions hold at once, and that is the only setting on the whole scale where they do.
How would this be found in a component already in use?
The method is smaller than people expect and that is its whole virtue. Take the refused cases from one period. Have a person who knows the work go through a sample of them and mark which ones were wrong, meaning cases they would have accepted with the same file in front of them. Then sort only the wrong ones by any characteristic the component never received. That is it. Sorting the wrong ones by an unrecorded characteristic is the whole of the method, and everything before it is an ordinary file review.
Two things make it work. The first is that the method needs no access to the component's internals. Component 2 runs inside a vendor-hosted service, so the bank can see an image go out and a single word come back and nothing in between, a limit set out under the AI vendor. Finding a pattern in what the component did requires no opening of the component. The second is the ordering: the person marks the file wrong before anybody sorts by group, so the judgement cannot be influenced by the very pattern under test.
Be honest about what a sample can carry. Two hundred of 620 is a decent sample and 31 is a small number of wrong cases, so the right claim from this review is not that the wrong-rejection rate is exactly 15.5 per cent. The right claim is that all 31 shared one condition, a statement about concentration rather than about a rate, and concentration is what survives a small sample. The extrapolation to about 96 people a month is a reasonable estimate and is stated as one.
What does it take to find out whether a component already in use does this?
What can be done once it has been found, and what cannot?
Everything depends on which of the three sources produced it. A pattern from source 1 is addressed by changing the population the component was built from and building again, and how that is done is covered separately. A pattern from source 3 is addressed by changing a definition or by adding the missing measurement, and the second of those is usually the honest answer.
A pattern from source 2 cannot be addressed inside the component at all, and this is the part practitioners resist hardest. Rebuilding component 2 on more low-light images might help the reading a little. Rebuilding does not change the fact that a darker, grainier image carries less information than a bright, sharp one, and no component recovers information the image does not hold. The remedy sits at the capture: a second route for anyone the check refuses, guidance on screen before the image is taken, a retry with a different framing, or a person. A pattern that arrives with the data as it is captured is fixed at the capture or it is not fixed.
Three things cannot be done, and saying so is not defeatism. A characteristic cannot be removed from the inputs in a way that also removes its effect. A proxy does the same work with better manners. Definitions 1 and 3 cannot be satisfied together where the underlying rates differ, at any level of care. And definition 2 cannot be checked without recording the group, and recording the group cuts directly against collecting as little as possible, a tension set out under access control and data minimisation. The three limits are the boundary of the craft, and a review that promises to clear all of them has promised something arithmetic does not allow.
The pattern comes from how the image was captured rather than from the fitting. What can be changed inside the component?
What does the firm owe the group the errors fall on?
Three concrete acts, and none of them is an apology. The first is a second route: anybody the check refuses gets another way to reach a decision. At this bank no such route existed in the month under review, and providing one is precisely what the reviewer recommended. A learned component produces a prediction, not a decision. Agrawal, Gans and Goldfarb's Prediction Machines, 2018, draws that line, and somebody still has to act on the prediction. A second route is what acting on it looks like when the prediction is wrong.
The second is notice. About 96 people a month were refused by a step that was mistaken about them, and a second route that exists only for future applicants leaves everybody already refused exactly where they were. Telling people who were refused that another route now exists is unglamorous, cheap, and the only part of this that reaches the people it was about.
The third is the change that prevents the next occurrence: measure the refusal rate by group from then on, rather than by volume. Volume measurement is what hid this in the first place, so a firm that keeps measuring volume while regretting the pattern has changed nothing at all. Only the third act stops this happening again, and it is the one that costs a report rather than a project. Under the bank's own numbered duties for a named accountable person, seeing the monitoring is duty 4, and at the month 12 validation the named person for the scoring model could not evidence that duty because the monitoring existed and went to the team that built the chain rather than to her.
About 96 genuine applicants a month were refused in error, concentrated on one group. What does the firm owe them?
How does somebody actually use this on a working Monday?
For a lender, this becomes three questions put to whoever is accountable for a component, and they take about ten minutes to ask. Which of the three sources has the owner looked for? What do the wrong refusals from last quarter look like, sorted by something the component never received? And which of the four definitions can the owner evidence, as opposed to satisfy by construction? A component owner who can answer the third question with a number rather than a design argument has done the work.
For an analyst reading somebody else's disclosures, the useful test is a subtraction. Count the fairness statements that are about the process and the fairness statements that are about outcomes, and a firm reporting only the first has described what its design does rather than what its month did. For a household on the other side of the screen, the practical translation is short. Where an automated step refuses an application, ask whether a second route exists. Why the refusal happened is often unanswerable, and whether a second route exists always has an answer.
Somebody always asks about cost. The review that found this pattern was a person opening 200 files, days rather than months of effort. Set that against a chain that cost the invented bank Rs 2,40,00,000/- to build once and Rs 65,00,000/- a year to run, and the measurement is not the expensive part of anything described here. The expensive part is a month of refusals nobody can now describe and a route back that did not exist when it was needed.
Who would characterise a pattern like this, and where is that written?
Whether a pattern of this kind is permitted, and what a regulated lender must do about one, is for the Reserve Bank of India. The Reserve Bank publishes its expectations on digital lending, customer data, consent, outsourcing and record keeping at rbi.org.in. Where the deployer is a market intermediary rather than a lender, the equivalent expectations sit with the Securities and Exchange Board of India at sebi.gov.in. Where the question is what a board is accountable for, the Ministry of Corporate Affairs publishes at mca.gov.in. The four definitions are definitions in use in the craft and are not anybody's standard. Requirements, thresholds and effective dates move, and the issuing body's own site carries the current position.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Published expectations on a regulated lender covering digital lending, customer data, consent, outsourcing and record keeping | rbi.org.in |
| Securities and Exchange Board of India | Equivalent expectations where the deployer of an automated decisioning step is a market intermediary | sebi.gov.in |
| Ministry of Corporate Affairs | Material on what a board is accountable for where an automated step affects customers | mca.gov.in |
| Cathy O'Neil | Weapons of Math Destruction, 2016, on errors that do not fall randomly but land repeatedly on the same people, and hardest where nobody is counting | Crown Publishing Group |
| Ajay Agrawal, Joshua Gans and Avi Goldfarb | Prediction Machines, 2018, on the line between a prediction and the decision somebody still has to make | Harvard Business Review Press |
Sumeru Bank Limited is invented.
Educational material. Not advice on any investment, tax, budget or market position.
