Training Data and Labels: What the Model Learns From, and Inherits
Training data is the set of past cases a component was fitted on, and a label is the outcome attached to each of those cases. The label is written by somebody, and whatever it says is exactly what the component learns to predict, no more and no less. A component also inherits what the examples contain and what they omit, and the largest omission in lending is invisible because it was never observable.
Here is the swap that makes this whole subject slippery. A label looks like a fact about the past and is actually a definition somebody wrote. A label arrives as a column of ones and zeros, it is built out of real events that really happened, and it therefore feels like a reading taken off the world. Change the definition and the same past produces a different component, fitted on the same cases by the same method, with nothing anywhere in the output to show that anything moved. Every claim a deployer makes about what its component predicts has to be traceable back to that definition, and where it is not, the claim is unsupported however well the component works.
What are the examples a component learns from?
Start with a vegetable seller who lets his regular customers take goods through the week and settle on Sunday. He keeps a notebook. Each line holds a name, roughly how long he has known that person, what they usually take, and then, written in later, whether they settled. The seller's notebook is training dataThe set of past cases a component was fitted on, each holding the inputs as they stood at the time and the outcome attached afterwards. in its entirety. Training data is nothing more than a set of past cases, each carrying what was known at the time and what happened after.
A lender's version of the same notebook is longer and colder and it is the same object. Each example is one past application: the details the applicant gave, whatever the lender pulled about them, and the state of the file on the day it was decided. The inputs in an example have to be the inputs as they stood on the day of the decision, not as they stand today. Where a past applicant's record has been updated since, using today's version puts information into the example that nobody had at the time, and the component then learns from a version of the past that never existed. The vegetable seller has the same problem in miniature. Where he writes down what he thinks of a customer now, after two years of settling, he has stopped recording what he knew when he first handed over the goods.
The fitting windowThe period the past cases were drawn from, and so the conditions the examples were recorded under. matters for the same reason. Examples come from a stretch of time, and that stretch had its own conditions: the channels that were live, the products being pushed, the people applying and the people staying away. Examples do not merely contain applicants. Examples contain a period.
What is a label, and who wrote it?
A labelThe outcome attached to each past case, written by somebody as a definition rather than read off the world. is the outcome attached to each past case: the answer written in the margin of the notebook, the thing the component is fitted to reproduce. In the vegetable seller's notebook it is the word settled or the word did not. Nothing is learnable without it. A pile of past applications with no outcome attached gives a component no answer to reach towards, so it teaches nothing at all.
Now take that word settled and try to write it down properly. Settled by Sunday night? By Tuesday? Settled at all, whenever it eventually came? What about the customer who stopped coming without settling and without saying anything? What about the one who settled every week for a year and then went quiet for a month? The moment the label has to be written down as a rule that can be applied to every line in the notebook without further thought, it turns out never to have been a fact at all, but a decision.
So it goes at Sumeru Bank Limited, invented, whose retail loan intake chain runs from an application started on a handset through to a decision. The scoring model in that chain, component 6, needed one written definition of the outcome it was fitted to reproduce. The bank wrote down one sentence: an accepted application that reached 90 days past due within 12 months, excluding accounts closed early. Both numbers in that sentence are Sumeru's own picks, made by people at Sumeru, and neither is a standard, a threshold or a rule handed down by anybody.
Is a label a fact about the past, or a definition somebody wrote?
Which four choices are hidden inside one label?
Read slowly, that sentence yields four separate decisions, numbered here in the order Sumeru worked through them rather than the order they ended up written. Choice 1 is what counts as bad. Sumeru set it at 90 days past due, its own cut and not a standard. Choice 2 is how long the bank waits after the decision before looking, the observation windowHow long after the decision the outcome is looked for before the case is labelled., which Sumeru set at 12 months, again its own pick and again not a standard. Choice 3 is which applications are in the set at all, and Sumeru used accepted applications only. Choice 4 is what to do with accounts that closed early. Sumeru excluded them.
Notice what kind of decisions these are. Not one of them is a modelling decision. Every one is a question about the business, answered by people in a room: Revathi Balan, who heads retail credit and is the named accountable person for component 6, and the people she asked. A different set of people, in a different room, on a different Tuesday, could have answered all four differently and been just as defensible. The examples would not change. The method would not change. The component would.
Which set below names the four choices sitting inside this label?
What does the component actually predict?
Once the label is written, the answer is fixed and it is narrow. A learned component predicts its label, and it predicts nothing else, ever. Component 6 does not predict whether somebody will repay. Component 6 produces a reading of how much an application resembles the past cases that carried a bad label, where bad means what that one sentence says it means at that one bank. Agrawal, Gans and Goldfarb make the point in Prediction Machines that what a learned component produces is a prediction, and that somebody still has to decide what to do with it. Add the label to that and the sentence gets sharper still: what it produces is a prediction of one written definition, and somebody still has to decide what to do with it.
The word bad is doing damage here, so hold it at arm's length. Bad is a name for a row in a table. It marks a bad caseA case the label marks as the outcome the component is being fitted to predict. The word names a definition, not a person., meaning a case matching the definition somebody wrote, and it says nothing about the person the case describes, their circumstances or their intentions. A household can miss three months of payments because a salary stopped, because a hospital bill arrived, or because a transfer failed twice and nobody rang. The label records none of that. The label records only that the definition was met.
The gap between what a component was fitted to and what people say it does opens easily. Nobody in a meeting says the component predicts whether an accepted application reached 90 days past due within 12 months on the bank's own definitions, excluding accounts closed early. People in that meeting say it predicts risk. The short sentence is easier to say and it is a wider claim than anything that was ever measured, and once it is in the room it travels.
In the window the bank used, 80 per cent of applications were accepted. Is that the interesting number here?
What did the examples at one lender actually hold?
Component 6 at Sumeru Bank Limited was fitted on a past window holding 3,00,000 applications. The 3,00,000 is the whole window: everybody who applied in it. Of those, 2,40,000 were accepted, 80.0 per cent, and they are the ones with an outcome anybody can look up. The other 60,000 were declined, 20.0 per cent, and they have no outcome at all.
Inside the 2,40,000 the label divides the set three ways, and the three parts sum back to the whole.
| What the label says about the case | Cases | Share of the 2,40,000 |
|---|---|---|
| Carries a bad label: reached 90 days past due within 12 months, on this bank's own definitions | 8,160 | 3.4 per cent |
| Excluded by choice 4, being accounts that closed early, so carries no label either way | 4,320 | 1.8 per cent |
| Everything else, carrying a good label | 2,27,520 | 94.8 per cent |
| Accepted applications in the window | 2,40,000 | 100.0 per cent |
Two things are worth sitting with in that table. The first is how small 3.4 per cent is: the thing the component is being fitted to find is rare, and almost everything in the examples is an ordinary account that did what accounts usually do. The second is that 4,320 cases were removed by a choice made in a meeting, and if that choice had gone the other way, every one of them would have arrived carrying a good label instead of no label at all.
Which applications are in the examples, and which are not?
Choice 3 looked like the dullest of the four and it is the one that reshapes everything. The examples hold the accepted populationThe applications this lender agreed to, and the only ones whose outcome could later be looked up. and nothing else. Agreeing to an application is what creates the account whose behaviour is later read, so an outcome can only be attached to an application the lender agreed to. The 60,000 declined applications in that window never became accounts, never had a payment due, and therefore never generated the thing the label is asking about.
Look at the drop in the middle of the picture below. One fifth of the window leaves the set at that step, and it does not leave because somebody chose to drop it. Nothing is there to keep.
Why can no amount of data fix the missing outcomes?
The missing outcomes catch people who are used to data problems being fixable. A missing column can be backfilled. A broken feed can be repaired. A short history can be extended by waiting. The declined populationThe applications this lender refused, whose outcome was never allowed to occur and so cannot be recorded anywhere. is none of those. The outcomes of the 60,000 declined applications were not lost, mislaid or badly recorded, they were prevented, and prevention leaves nothing behind to collect.
Say it in the vegetable seller's terms and it is obvious. He turned away four people last year because he did not know them. Would they have settled on Sunday? He never handed anything over, so he has no idea, and no amount of careful notebook keeping will ever tell him. His notebook is a perfect record of the people he said yes to and a complete blank on the people he said no to, and it is the second list that his instinct about strangers is really made of.
Now put that next to how the component is used. Component 6 was fitted on people the bank had already agreed to lend to. Component 6 is deployed on everybody who applies, including the applications a previous rule or a previous person would have refused. The group it was never shown sits closest to the line. At that line its answers matter most, and at that line nothing in the examples can support them. Nobody should read that as a fault somebody committed. Every lender that has ever fitted a component this way faces it, and the honest response is to say so out loud rather than to describe the component as though the gap were not there.
Why can more data collection never fix the missing outcomes for declined applications?
What does the observation window change?
Choice 2 was the waiting time, and it is the choice with the cleanest arithmetic behind it. An account that will eventually fall 90 days past due, on this bank's own cut rather than a standard, does not do it on a schedule. Some reach that point in the fourth month, some in the eleventh, some in the twentieth. Where the clock stops decides how many of them have been seen, so the observation window sets the count of bad cases without touching a single case.
Sumeru measured it on the same fixed set of 2,40,000 accepted applications. At a six month window, 5,040 cases carry a bad label, being 2.1 per cent. At twelve months, 8,160, being 3.4 per cent. At eighteen, 10,080, being 4.2 per cent. At twenty four, 11,520, being 4.8 per cent. Same people, same accounts, same payment histories, four different answers to how many were bad. The bank picked twelve, and twelve months is Sumeru's own choice rather than a rule anybody set for it.
Before the control below moves: the window moves from 12 months to 24. What happens to the number of bad cases?
Move the waiting time, and watch the same cases change label
One control: the observation window, at 6, 12, 18 or 24 months. One consequence: how many of one fixed set of 2,40,000 accepted applications carry a bad label. The default is the twelve month window Sumeru used, giving 8,160 bad cases, being 3.4 per cent of the 2,40,000. Moved left to six months, the same set gives 5,040, being 2.1 per cent. The three other label choices are held still throughout. Every movement therefore belongs to the window alone.
12 months
At a 12 month window, 8,160 of the 2,40,000 accepted applications carry a bad label, which is 3.4 per cent. That is the window Sumeru Bank Limited picked for itself, and it is this bank's own choice rather than a standard.
How many of the labels turned out to be wrong?
A label is written by a process, and processes make mistakes. An account gets flagged against the wrong reference. A settlement lands after the file was already stamped. A restructuring is recorded in one system and not in another. Sumeru ran a data quality pass across the 2,40,000 labelled cases and found 3,120 that had to be changed, a label correctionA case whose recorded outcome was later found to be wrong and changed. rate of 1.3 per cent on the bank's own measurement of its own records.
Read 3,120 as a share and it is a rounding error, read it as a count and it is three thousand past cases the component was fitted on with the wrong answer attached, and both readings are true at once. Which reading applies depends on what is being decided. For whether the overall shape of the component is likely to be badly wrong, 1.3 per cent is reassuring. For whether the label sitting against any particular case is safe to rely on in a dispute, three thousand corrections is a warning that the record is not the last word.
How many of the 2,40,000 labels were later corrected, and what is that as a share?
What else does a component inherit from its examples?
The pattern is not the only thing that travels from the examples into the component. Everything the examples happen to contain travels with it, including things nobody chose and nobody wrote down. A set of examples is not a neutral sample of the world, it is a record of one lender's own past behaviour, and the component inherits that behaviour along with the pattern.
The clearest instance at Sumeru is the channel mix. One channel was where the bank was pushing hardest in the window the examples came from, so 62 per cent of the examples came from that one channel. Nobody decided that the component should learn mostly from one channel. The share simply followed from where the applications had come in. The bank's own measured readings on the two groups differ: 94.1 per cent on cases from that channel against 87.9 per cent on cases from all the others, a gap of 6.2 points. Readings of that kind, and how one is produced, are covered separately. The difference between the two groups came out of the composition of the examples and nowhere else.
The same inheritance runs through everything else in the window. Whichever products were being sold, whichever branches were open, whichever kinds of applicant the previous decision rule let through: all of it is inside the examples, and none of it is written on the component. A serious question opens here about whether the errors fall evenly across different groups of people. Fairness across groups is a subject in its own right and is covered separately. The step before it is simply knowing what the examples contain.
Nobody at the bank decided that 62 per cent of the examples should come from one channel. Where did that share come from?
How does a label definition reach the applicant?
Follow one file to the end. The label is defined in a meeting in month 0. The component is fitted to it during the pilot in months 1 to 3. From month 4 the component scores live applications, and at steady state 688 of them a month are declined with no person touching the file. Each of those applicants gets a letter. The letter is where a definition written in one meeting finally reaches the person it was written about, and at Sumeru nothing between the meeting and the letter ever compared the two.
Sumeru's standard letter said the application had been assessed as likely not to be repaid. Put it beside the definition and the two are not the same statement. One is about applications this lender accepted, one outcome, one waiting time, and one bank's own definitions. The other is about a person and repayment in general, with no population, no outcome definition and no period attached to it at all.
Nobody did anything obviously wrong at any single step, so look closely at how the two ended up so far apart. Each step is reasonable on its own. The definition is written by people who know credit. The component is fitted by people who know fitting. The letter is drafted by people who know letters. The path from the first to the last crosses three sets of hands, and there is no step on it whose job is to hold the sentence up against the definition and check that one supports the other.
The letter says the applicant was assessed as likely not to repay. What is wrong with that sentence?
How does a lender, a reviewer or a household use this?
What each of them does with the definition
A credit head being asked to approve a component asks for the label on one line, in writing, before anything else. Then four questions, one per choice: what counts as bad, who picked it, how long the wait was, and what happened to the cases left out. Revathi Balan can answer all four for component 6 because she was in the room. There is nothing to approve except the definition the component was fitted to. Where nobody in the room can produce the sentence, the component has not been approved, it has been waved through.
A reviewer works the other way round, from the output back. Neelima Rao read the customer letter, then read the label, then held the two side by side. The whole method is that short, and it takes an afternoon. Notice what she did not need: no access to how the component was built, no statistics, no sight of the fitting. Two sentences, and whether the first supports the second.
An analyst reading what any lender says about a component of this kind asks for the population, the outcome and the window before believing any number attached to it. Two bad rates quoted without those three are counting different things under the same word, so neither can be compared with the other. Where the disclosure does not carry the definition, the honest note to write is that the figure is not comparable rather than that it is high or low.
A household does the small version of this every time it lends to a relative. The household decides what counts as a problem, whether it is a week late or never at all. The household decides how long to wait before counting it. And the household only ever finds out about the people it said yes to. The confident view it holds about the ones it turned down rests on nothing it has actually seen.
Who sets expectations on a lender using examples like these?
A bank in India fitting a component on its own past applications sits under the Reserve Bank of India. The Reserve Bank publishes its expectations on digital lending, outsourcing, customer data and consent at rbi.org.in. Where the deployer is a market intermediary rather than a lender, the equivalent expectations come from the Securities and Exchange Board of India at sebi.gov.in. The current position is stated at the issuing body's own site and is to be confirmed there before any reliance on it.
The 90 days and the 12 months in this case are Sumeru's own picks, made by people at an invented bank, and they are not a standard or a definition anybody set. Both numbers resemble figures a reader may have seen quoted as rules elsewhere, and that resemblance is exactly why they must not be carried away from here as one. Where a supervisor defines such a term for a regulated lender, the definition is that supervisor's and is stated at its own site.
The error that gets made, and what it costs
Sumeru's standard decline letter said the application had been assessed as likely not to be repaid. The sentence had been drafted once, approved once, and printed on every letter since. The component behind it predicts one thing: whether an application this lender accepted reached 90 days past due within 12 months, excluding accounts closed early, on this bank's own definitions. The component says nothing about repayment in general, nothing about applications this lender would have declined, and nothing about the applicant beyond that single definition.
Neelima Rao found the wording at the month 12 validation, and she found it by reading two documents next to each other rather than by testing anything. The sentence had stood since go-live in month 4. Take only the six steady months, months 6 through 11, at 688 auto-declines a month: 6 times 688 is 4,128 letters. Months 4 and 5 ran below steady volume, on one channel and then on all of them, and they sit on top of that count rather than inside it. So roughly 4,100 letters is the floor rather than the total.
The component did what it was fitted to do, so the cost was not the component, and it was not the decline either. The cost was that a bank had told several thousand people something it had no basis for, in a sentence nobody had ever traced back to the definition it came from. The letter was unsupported by anything the bank had measured.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering digital lending, outsourcing, customer data and consent | rbi.org.in |
| Securities and Exchange Board of India | Expectations where the deployer of such a component is a market intermediary | sebi.gov.in |
| Bank for International Settlements | International supervisory material on the use of such components by banks | bis.org |
| Agrawal, Gans and Goldfarb | Prediction Machines, on a learned component producing a prediction that somebody still has to act on | Harvard Business Review Press |
Sumeru Bank Limited, its retail loan intake chain, Revathi Balan and Neelima Rao are invented.
Educational material. Not advice on any investment, tax, budget or market position.
