Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
AI, Automation & Digital Finance
1AI Foundations
Artificial Intelligence in FinanceAlgorithmNeural Networks and Deep LearningMachine LearningArtificial Intelligence vs Machine…Computer Vision in FinanceTraining Data and LabelsNatural Language Processing in Finance
2Generative AI
Generative AIGenerative AI vs Predictive AILarge Language ModelsEmbeddingsHallucinationFine TuningPrompting vs Fine TuningThe PromptThe Context WindowTool CallingGroundingVector DatabasesRetrieval Augmented GenerationRAG vs Fine Tuning
3Automation and Workflow
Workflow AutomationAutomation vs AugmentationHow to Map a…Straight-Through Processing and Exception…Robotic Process AutomationRule EnginesMachine Learning vs Rule-Based…
4Document and Operations AI
Intelligent Document ProcessingBatch vs Real-Time vs…Document Classification vs Entity…Service Level AgreementsCase ManagementHow to Document Data…Reconciliation AutomationOptical Character Recognition and Data ExtractionConfidence Scores
5Customer Systems, Identity and Digital Assets
Digital IdentityConsent ManagementBlockchain and Distributed LedgerChatbots and Conversational AIFrom Use Case to ProductionDigital Assets and TokenisationDigital SignaturesData Sharing in FinanceElectronic KYC and Digital Onboarding
6Credit and Fraud Systems
The Fraud AlertCredit Decisioning SystemsHuman in the Loop…Adverse ActionAnomaly DetectionThe Decision ThresholdCredit Score vs Credit DecisionAlert Triage and EscalationFraud Detection and Transaction MonitoringFraud Model vs Credit ModelHow to Build Human…
7Governance, Data and Vendors
AI Governance and the AI PolicyHow to Create an…Explainability and Interpretability ComparedThe AI VendorBias and Fairness in Financial AIShadow AIAccess Control and Data MinimisationCloud Computing in FinanceData Lineage and Master DataData ResidencyThe AI Use Case Register and Model InventoryThe Model Owner
8Model Performance, Monitoring and Resilience
Model DriftFalse Positives and False NegativesClassification MetricsAdversarial AttacksModel TestingBias, Fairness and Explainability…Stopping an Automated SystemModel ValidationAI Governance vs Model Risk ManagementPrompt InjectionHow to Create an…

Fraud Model vs Credit Model: The Label and the Cost

A fraud model and a credit model are both called models and are not the same kind of object. At Sumeru Bank Limited, invented, one is 61 lines somebody wrote and the other is fitted to 2,40,000 past applications. The clock is what separates them: a fraud outcome is known within days, and a credit outcome under this bank's own label is not known for 15 months.

Take that last sentence seriously before anything else. Every other difference between the two hangs off it. A smoke alarm in a kitchen makes it clear within about ninety seconds whether it was right, and a household can be cross with it that evening and move it the next day. A school report saying a nine year old will struggle with mathematics cannot be checked for years, and by the time anybody could check it, the report has already changed what happened. Both are described as assessments. Nobody in a household would treat them as one kind of thing. Two arrangements inside one bank get treated as one kind of thing constantly, and the reason is that somebody wrote the same word in front of both of them.

Sumeru Bank Limited, invented, runs a retail loan intake chain of nine components, and two of them carry the word model in everyday speech inside the bank. Component 7 watches the servicing book for fraud. Component 6 returns a number about a loan application. The two components sit in the same chain, they appear on the same monthly pack, and on almost every question a deployer actually has to answer they point in opposite directions.

What is each of the two actually made of at this bank?

The fraud componentWhat this bank calls its fraud model. It is 61 written lines, each one a condition somebody typed out and can read back. is 61 written lines. Not 61 layers, not 61 features: 61 lines of the form if this and this, then raise an alert. Somebody typed each one. Anybody with the file open can read all of them, and at the rate one reviewer in this bank was measured reading a written rule set, the whole 61 takes about 45 minutes end to end. The 45 minutes is arithmetic on a locked reading rate rather than a measured sitting, not a timing anybody took with a watch.

The credit componentThe fitted scoring model that returns one number about an application, on a scale this bank invented for itself. is fitted. The fitting used a past window of 3,00,000 applications at this invented bank, of which 2,40,000 had been accepted and therefore had an outcome anybody could observe, being 80.0 per cent, and 60,000 had been declined and had no observable outcome at all, being 20.0 per cent. Nobody wrote a line of it. There is no sentence inside it to read back. One of the two is a list somebody wrote and the other is a shape somebody fitted, and reading that sentence first is what stops the rest of the comparison being nonsense.

One more thing has to be said here or a later paragraph will mislead. The two do not act on the same population and they do not act in the same unit. In one steady month the fraud lines raised 18,000 alerts on 20,00,000 servicing transactions, and separately routed 452 application files at intake. The credit component determined the outcome of 5,981 application files, being 69.5 per cent of the month's 8,600. Alerts on transactions and outcomes on applications are different objects counted in different units, and no monthly pack may add them together or set the two rates side by side as though they answered one question.

There is a sharper version of that. Of the month's 8,600 files, the fraud lines determined 452, or 5.3 per cent, and every one of those 452 was a decision about where the file went next. Not one of them was a decision about a loan. The credit component, on the same 8,600, produced 4,902 accepts and 688 declines outright. So on the lending side of this bank the written component decides routes and the fitted component decides outcomes, and that is a difference in what each is allowed to do rather than a difference in how well either does it.

SAME WORD ON BOTH. DIFFERENT KIND OF OBJECT INSIDE EACH. THE FRAUD COMPONENT, BEING COMPONENT 7 THE CREDIT COMPONENT, BEING COMPONENT 6 61 written lines, and every one of them is readable a fitting window of 3,00,000 past applications 44 lines written after a loss the bank had already taken 72.1 per cent of the 61 17 lines written from an expectation of a loss, never taken 2,40,000 accepted, so the outcome could be observed 80.0 per cent of the window 60,000 declined, with no observable outcome at all Nothing in this panel was fitted to anything. Nothing in this panel was written as a line. ONE IS A LIST SOMEBODY WROTE. THE OTHER IS A SHAPE SOMEBODY FITTED.
At this invented bank the fraud component is 61 lines somebody typed and the credit component is a shape fitted to 2,40,000 accepted applications, and only the second of the two was fitted to anything.
Try it out

Of the two components at this bank, which one is actually fitted?

Breaking Into VC Bootcamp — Fin Maverick

What is one written against, and what is the other fitted to?

A shopkeeper who takes a bad note on a Tuesday writes a rule on Tuesday evening. No notes of that series after eight in the evening. The shopkeeper's rule is exact, readable by the boy at the counter, and a description of one event that has already happened to that shop. The rule will catch the next attempt of the same shape and has nothing whatever to say about a kind of trouble the shop has not met yet. Every written line anywhere works this way, and the fraud lines at this invented bank are the same object at a larger scale.

Of the 61 lines, 44 were written after a loss the bank had already taken and 17 from an expectation of one. The 44 and 17 split is the provenanceWhere a written line came from: either a loss the institution has already taken, or an expectation of one that has not happened yet. of the set, and the provenance decides what the set can see. The three alert sources that carry 22 of the month's 27 confirmed cases, being 81.5 per cent of them, all sit among the 44. A rule set is a record of what has already happened to the institution, and the record is at once the source of its accuracy and the exact shape of what it cannot see.

The credit component was made the other way round. The credit component was fitted against a labelThe outcome a fitted component was trained to point at, defined by choices somebody made rather than by anything in the data., and this bank's label carries four choices, all of them its own: what counts as bad, set at 90 days past due; the observation window, set at 12 months; the population, being accepted applications only; and accounts closed early, excluded, a choice that removed 4,320 cases. On those definitions 8,160 of the 2,40,000 carry a bad label, being 3.4 per cent, 4,320 carry no label either way, being 1.8 per cent, and 2,27,520 are good, being 94.8 per cent. Those three sum to 2,40,000 and the shares sum to 100.0.

Now put the two creation stories beside each other and watch the direction reverse. The event a fraud line describes is behind it: the loss happened, then the line was written. The event the credit component will be judged on is ahead of it: it was fitted, then deployed, and the applications it decides today did not exist when it was made. One points backwards at something already survived. The other points forwards at something that has not happened.

THE TWO ARE MADE IN OPPOSITE ORDERS TODAY THE FRAUD LINES a loss the bank took 44 of the 61 begin here a line gets written readable, and dated it waits it fires on the next one of the same shape the event it describes has already happened THE CREDIT COMPONENT 2,40,000 past outcomes accepted files only the component is fitted no line is written it scores a file that did not exist then the outcome arrives 15 months later the event it will be judged on has not happened yet ONE POINTS BACKWARDS AT SOMETHING SURVIVED. THE OTHER POINTS FORWARDS AT SOMETHING NOT YET DONE.
The fraud lines are written after the loss they describe and the credit component is fitted before the applications it now decides existed, so the evidence sits on opposite sides of today.
Try it out

Where did most of this bank's 61 fraud lines come from?

How long does each one wait to find out whether it was right?

The clock decides everything else, so the arithmetic is worth doing in the open. On the fraud side a case is confirmed or it is not, and the answer arrives in days. The desk works a kept case for about 40 minutes, an investigation takes about three hours, and at the end of it somebody writes down whether it was fraud. In one steady month, 27 cases were confirmed out of 18,000 alerts. Every one of those 27 verdicts existed inside the same month that produced the alert.

On the credit side there is no such moment. Sumeru Bank's label needs 12 months of observation on an account, and inside that window an account has to have reached 90 days past due to count as bad. An account that first goes past due right at the end of the observation window takes a further 90 days to arrive at the definition, so the earliest a month's decisions can be scored against real outcomes is 12 months plus 90 days, being 15 months. Decisions taken in month 6 cannot be scored until month 21. The 15 month gap is the label delayHow long after a decision the outcome it was pointing at becomes known, which is set by the label's own definitions and not by anybody's effort., and the delay is a consequence of the four label choices rather than a measurement anybody took.

Nobody can shorten it by working harder. Shortening it means changing a definition, and every way of doing so costs something: a 6 month window at this bank reads 5,040 bad accounts and 2.1 per cent, and a 24 month window reads 11,520 and 4.8 per cent. The two are readings of different things, not better and worse readings of one thing. So the delay is not an inefficiency in the arrangement. The delay is a property of what the arrangement was built to point at.

TWO CLOCKS, DRAWN ON ONE SCALE. MONTHS FROM APPROVAL ALONG THE BOTTOM. FRAUD CASE this is the whole of it, drawn true to scale CREDIT OUTCOME 12 MONTHS OF OBSERVATION 90 DAYS 0 3 6 9 12 15 18 21 the decisions are taken here the earliest they can be scored THE FRAUD CLOCK, STRETCHED THIRTY TIMES DAYS a day here is a month above 12 months of observation plus the 90 days an account takes to reach the bad definition is 15 months, so the month 6 book cannot be assessed until month 21.
A fraud case is confirmed or not within days, while a month's credit decisions cannot be scored for 15 months, so the month 6 book waits until month 21.
Try it out

How long before a month's credit decisions at this bank can be scored against real outcomes?

What does a fifteen month wait do to what a credit component can be?

Watch what falls away. Anything ordinarily done to check an arrangement needs an outcome to check it against, and for 15 months there is no outcome. So for that whole stretch a credit component cannot be argued with about whether the decisions it produced were right, and that is the only question anybody actually cares about. A credit component can be argued with on everything else, and this is the part worth holding on to: its inputs can be checked today, its stability across months can be watched today, and its behaviour on a set of files can be examined today. None of those is the same question.

Sumeru Bank has already seen what that gap allows. An upstream income field changed format in month 8, week 2, and monitoring flagged it in month 9, week 3, six weeks later. About 12,900 files were decided in that window, and 176 of them moved from accept into the referral band. The approval rate fell from 57.0 per cent to 55.6 and referrals rose from 391 a month to 508, a rise of 117 exactly matching the fall in accepts. Nothing about the component changed. The data changed. No outcome existed to catch the change with, so the input watching caught it instead.

For more than a year there is nothing to govern a credit component on except its inputs, its stability and its behaviour, so those three are what the governing has to be built from. An arrangement designed as though outcomes will arrive in time to correct it has been designed against a clock that runs fifteen months slow.

How rare is the thing each one is looking for?

A wedding hall holds three hundred people and somebody has to spot the two who came for the buffet and not for the couple. That is hard and it is doable. A stadium holds seventy four thousand and somebody has to spot the one pickpocket. That is a different job, and no amount of care applied to the first version prepares anybody for the second. The word searching covers both. Nothing else does.

On the fraud side, the bank confirmed 27 cases in a month that carried 20,00,000 transactions on the servicing book. The confirmed rate is about one in 74,000. A fraud nobody detected is not in the count, so the rate is a floor rather than a true incidence. On the credit side, 8,160 of the 2,40,000 accepted applications in the fitting window carry a bad label, being 3.4 per cent, or about one in 29. Divide one base rateHow common the thing being looked for is in the population being watched, before anything is done to look for it. by the other and the credit event is about 2,500 times more common than the confirmed fraud event, being three orders of magnitude.

One reconciliation belongs here before anybody sets those two figures against a third. The 3.4 per cent belongs to a past window of 2,40,000 accepted applications with no cut-off applied to it. The 2.02 per cent this bank expects on the accepts its deployed cut-off produces is a different quantity answering a different question, and the two are never treated as one.

Now the trap the difference creates. The fraud lines speak on 0.9 per cent of transactions and the credit component declines 8.0 per cent of files. The two percentages look like versions of each other and they are answers to unrelated questions, computed on unrelated populations, about events roughly 2,500 times apart in rarity. Putting an alert rate and a decline rate in one table is the single easiest way to make two incomparable things look like two settings of one dial.

HOW RARE IS THE THING BEING LOOKED FOR? RARER TO THE RIGHT. ABOUT 2,500 TIMES RARER, BEING THREE ORDERS OF MAGNITUDE 1 in 10 1 in 100 1 in 1,000 1 in 10,000 1 in 1,00,000 1 2 3 4 5 1 The month's decline rate, 688 of 8,600 files, being 8.0 per cent or one in 12.5 2 The bad label rate, 8,160 of 2,40,000 accepted, being 3.4 per cent or one in 29 3 The alert rate, 18,000 of 20,00,000 transactions, being 0.9 per cent or one in 111 4 Cases triage kept, 540 of 20,00,000 transactions, being about one in 3,704 5 Confirmed cases, 27 of 20,00,000 transactions, being about one in 74,000 MARKERS 1 AND 2 ARE THE CREDIT SIDE. MARKERS 3, 4 AND 5 ARE THE FRAUD SIDE.
The credit component looks for something at about one in 29 and the fraud lines look for something at about one in 74,000, so an alert rate and a decline rate are not comparable numbers even in principle.
Try it out

A monthly pack sets an alert rate of 0.9 per cent beside a decline rate of 8.0 per cent. Are the two comparable?

Which way does the cost asymmetry run on each side?

Both arrangements can be wrong in two directions, and on both sides the two errors cost different amounts. That much is shared. Which error is the expensive one is not shared, and that single reversal is why a tuning principle that serves one of them will be wrong for the other every single time.

Start with the fraud side, where the cost asymmetryThe two errors costing different amounts, and the direction the difference runs in. It is a property of the situation rather than of any component. is severe and openly accepted. A missed case is money that has left the bank and cannot be pulled back by an internal entry, and this bank's own average amount at risk on a confirmed case is Rs 84,000/-. An alert that turns out to be nothing costs 90 seconds of a first read, and sometimes a held payment. So the arrangement is deliberately built to speak 18,000 times in a month in order to be right 27 times. 17,973 of those alerts were not fraud, being 99.85 per cent, and that number is the price the bank decided to pay rather than a fault it failed to fix.

Now the credit side, and here the direction of the asymmetry depends on who is counting. To the bank, the expensive error is an account that goes bad, and that error is countable: moving the accept cut-off from 720 to 760 removes 39 expected bad accounts. To the person, the expensive error is a decline that would have repaid, and that error is not in any cost line anywhere. The same cut-off move puts 760 applicants into a queue for a person instead of a four minute answer, and nothing in the bank's arithmetic prices what that did to any of them. The fraud side pays a visible price to avoid an invisible loss, and the credit side collects a visible saving while the cost lands somewhere nobody counts.

THE SAME QUESTION, ASKED OF BOTH Which error never comes to light? the fraud lines the credit component THE CASE NOBODY ALERTED ON THE DECLINE THAT WOULD HAVE REPAID So the lines are set to speak 18,000 times in a month in order to be right 27 times. 17,973 of those alerts were not fraud, being 99.85 per cent, and a missed case is Rs 84,000/- on average that cannot be pulled back. So tightening the cut-off reads as free. The 39 expected bad accounts a move from 720 to 760 removes are counted, and the 760 applicants it moves into a queue are an entry in nobody's cost line at all. THE COSTLY ERROR AND THE INVISIBLE ERROR ARE THE SAME ONE ON THE FRAUD SIDE AND DIFFERENT ONES ON THE CREDIT SIDE
On the fraud side the expensive error is also the one nobody sees, and on the credit side the invisible error is the one somebody outside the bank pays for.
Try it out

Both components could be tuned to speak less often. Should the two be tuned the same way?

AI For Finance Bootcamp — Fin Maverick

Which of the two errors does each side never get to see?

One more separation is the easiest to miss and the hardest to unsee afterwards. Each of the two arrangements is blind to exactly one of its own errors, and it is not the same one.

The fraud lines produce a verdict on everything they speak about. Every one of the 6,120 alerts a person read got an answer within days, and 6,093 of them were not fraud. The fraud lines never learn about the transaction they said nothing about. A case nobody alerted on generates no record, no verdict and no entry, and the only way it ever surfaces is if the account holder reports it. So the fraud side sees its noise in full and is blind to its misses.

The credit component runs the other way, and its blindness is written into its own label. Choice 3 of the four is that the population is accepted applications only. A declined application has no outcome anybody can observe, and that is why 60,000 of the 3,00,000 in the fitting window carried no label at all. The same structure repeats every month: of the 5,981 files the component scored, the 688 it declined outright will never carry a label either, being 11.5 per cent of the files it scored. The 11.5 per cent is arithmetic on locked counts rather than a measurement. The credit component sees its bad accepts in full, at month 21, and is permanently blind to the declines it should not have made.

O'Neil, in Weapons of Math Destruction, 2016, is the standing reference for errors that fall unevenly across a population and leave no trace where they land, and this is the plain mechanical version of it: an arrangement whose measurement window excludes the people it refused cannot measure what refusing them did. The blindness is not a fault in anybody's arithmetic but a consequence of a definition, and it is why a declined applicant and a stopped payment need somebody watching for them who is not the component's own scorecard.

Try it out

Which error is each side structurally unable to observe?

Is the person on the other side told why, and can they be?

Both arrangements act on people, and the two people get told strikingly different amounts. Neither amount was designed.

On the credit side, a reason code was written to the file on every one of the 688 auto-declines in the month. Six codes cover the whole set: the level of borrowing already outstanding against declared income, 227 files; the length of the credit record available, 151; a recent missed payment on another account, 138; the number of recent applications for credit elsewhere, 97; the declared income against the amount applied for, 51; and no single attribute carrying enough of the decision to name one, 24. The six counts sum to 688. So a reason exists, on the file, for every declined applicant, and somebody who is asked later has something to look at. The letter carried one sentence, the same on all 688.

On the fraud side there is no equivalent object reaching the customer at all. Of the 540 cases triage kept in the month, the bank held the transaction pending on 218 and released 191 of them after review, so 191 people had a lawful payment stopped in one month. Each of those cases carries an eight field record inside the bank naming which line fired and what it compared. The person on the other side is told that a check was run. Which line ran it is never said, and there is a real reason for the silence rather than an oversight: naming the line tells anybody who asks how to sit just underneath it next time. The tension between explaining and not teaching evasion is genuine and has no clean resolution. The silence has to be a decision somebody took rather than a gap nobody noticed.

WHAT IS KNOWN INSIDE, AND WHAT REACHES THE PERSON A DECLINED APPLICANT A CUSTOMER WHOSE PAYMENT WAS HELD One value on a scale of 0 to 1000 61 written lines, and one of them fires One of six reason codes, written to the file An eight field case record, opened inside A letter carrying exactly one sentence A payment held, on 218 cases in the month The same sentence went to all 688 files 191 of the 218 released after review A SENTENCE OUT. A CODE ON THE FILE. A CHECK WAS RUN. NO LINE NAMED.
A declined applicant has a reason code on the file and one sentence in the letter, while a customer whose payment was held is told only that a check was run.
Try it out

Whose file carries a reason for what happened to them, and whose does not?

India

Who states what applies here

Both sides of this comparison sit under named supervisors. The Reserve Bank of India at rbi.org.in states what applies to a regulated lender in India on digital lending, fair practice, the treatment of a borrower and the use of data and consent, and what applies to transaction monitoring at a bank in India is stated there too. Where the deployer is a market intermediary, the Securities and Exchange Board of India at sebi.gov.in is the place to read. The Financial Action Task Force at fatf-gafi.org is the origin of the international standard on transaction monitoring, and the position for a bank in India is again stated by the Reserve Bank of India. Every cut-off, every fraud value and every review interval in this comparison belongs to one invented bank and is that bank's own choice.

Try it out

Both components are put on the same annual review cycle. Which one is the cycle wrong for?

Breaking Into Quants Bootcamp — Fin Maverick

How often can each one be changed, and why is that not a preference?

A review cycleHow often an arrangement's behaviour is examined and can be changed. What is achievable depends on how fast evidence about it arrives. is usually treated as a matter of appetite: some institutions review often, some review annually, and a committee picks. On these two components appetite does not come into it. How fast the evidence arrives is what sets what is achievable, and the evidence arrives at wildly different speeds.

The fraud lines can be argued with every month, and the ingredients for the argument are all present. The month produces 27 confirmed cases with a source attached to each. The whole set of 61 lines can be read end to end in about 45 minutes, at the rate one reviewer here was measured reading a written rule set of comparable length. Anybody can point at a line, name what it caught and what it did not, and propose a change that another person can read before it goes live. An annual cycle on a set like that leaves eleven months of evidence sitting unused.

The credit component cannot be argued with on outcomes at all for 15 months, and the effort involved runs the other way too. The independent review of its behaviour took eleven working days, or 4,620 minutes at this bank's own working day of 420, about a hundred times the reading of the 61 lines. The hundredfold ratio is arithmetic on locked figures rather than a measured comparison. The fraud set is cheap to examine and rich in evidence, the credit component is expensive to examine and starved of evidence for over a year, and no single review interval can be right for both.

What do the seven criteria look like set beside each other?

Here is the whole comparison in one place. Read down the middle column first. The questions are the useful part: they are the seven things a deployer has to settle before either component runs, and every one of them gets a different answer on each side.

SEVEN QUESTIONS A DEPLOYER HAS TO SETTLE, AND SEVEN DIFFERENT ANSWERS ON EACH SIDE THE QUESTION THE FRAUD COMPONENT THE CREDIT COMPONENT What is it made of? 61 lines somebody wrote a shape fitted to 2,40,000 outcomes Measured against what? a case confirmed, or not a label with four choices in it How long until it is known? days 15 months How rare is the thing? about 1 in 74,000 transactions 3.4 per cent of accepted files Which error is invisible? the case nobody alerted on the decline that would have repaid What is the person told? that a check was run one sentence, and a code on the file How often can it change? argued with every month nothing to argue with for 15 months SIX ROWS ARE OPPOSITES. THE FIRST IS NOT, BECAUSE THE TWO ARE DIFFERENT KINDS OF OBJECT.
Seven questions separate the two components at this invented bank, and on six of them the answers are opposites while the first is a difference in kind.

Where does treating the two as one kind of thing go wrong?

A report at this invented bank put both under one heading and called them the bank's two decision models. Nobody who wrote that sentence misunderstood either component. The sentence was accurate in the ordinary sense that both are components, both produce outputs, and both sit in the same chain. The heading handed the reader a category, and a category invites a common treatment.

The common treatment duly arrived as a proposal to review both on the same annual cycle, and it produced two errors at once, pointing opposite ways. For the fraud lines, annual is far too slow: confirmed cases arrive every month, 44 of the 61 lines exist because somebody saw a loss and wrote one, and a set that could be argued with monthly was being reviewed yearly for tidiness. For the credit component, annual sounds right and is not achievable on outcomes: no outcome under its own label exists for 15 months, so an annual review of it is a review of its inputs and its stability rather than of whether the decisions were right.

THE SENTENCE, AND THE TWO OPPOSITE ERRORS IT PRODUCED MONTHLY RISK PACK, SECTION 3.2 The bank's two decision models will move to a common annual review cycle. FAR TOO SLOW FOR THE 61 LINES 44 of them exist because somebody saw a loss and wrote one. Confirmed cases arrive every month. A set that could be argued with monthly was to be reviewed yearly, for tidiness. FASTER THAN THE EVIDENCE FOR THE OTHER No outcome under its own label exists for 15 months. So an annual review of it examines its inputs and its stability, and never whether the decisions it produced were right. NOBODY MISUNDERSTOOD EITHER COMPONENT. THE HEADING DID THE DAMAGE.
One shared heading produced an annual cycle that was far too slow for 61 written lines and faster than the evidence for a component whose outcomes take 15 months.

The error that gets made, and what it costs

The mistake is putting two components under one word and then applying one treatment to the category. The mistake is made by somebody careful, writing a pack that has to be short, and it survives because every individual statement in it is true. Both are components. Both produce outputs. Both need governing. The sentence is unarguable and the treatment it invites is wrong twice.

The cost lands in two places. Inside the bank, eleven months of monthly evidence about a readable rule set goes unused, and the credit component gets a review labelled as a check on its decisions that is nothing of the kind. The second error is the more dangerous of the two. A review of that kind produces a document saying the arrangement was examined. Outside the bank it lands on people. A stopped payment and a declined application are both acts against somebody, and an arrangement reviewed on a schedule that does not match its evidence is one where the person affected has no route to a correction that arrives in time to matter.

Nothing about either component was misunderstood by anybody, and the damage was done entirely by the heading. Naming this failure matters for exactly that reason: it cannot be prevented by knowing more about either component, only by refusing to let one word carry two objects.

Financial Analyst Program Bootcamp — Fin Maverick Bond Pricing and Yield Mechanics — free micro-course from Fin Maverick

What is genuinely the same about the two?

A comparison that only separates does half the work. Three things really are shared, and it is precisely because everything else differs that these have to be stated out loud rather than assumed.

First, both produce an output that somebody acts on, so neither is a study. An alert becomes a held payment. A score becomes an accept or a decline. Second, both need a named person who answers for them and can stop them. Sumeru Bank names Revathi Balan, head of retail credit, for the scoring component, and when Ashok Pillai ran the register sweep in month 10 it found 14 uses in the bank against 9 registered, with only 6 of the 14 carrying a named accountable person, being 42.9 per cent. Third, both need a written record of every value somebody chose and when. On the credit side the two cut-offs are the bank's own choices. On the fraud side, two of the six alert sources compare a transaction against a value the bank set for itself. Neither set of values is a standard, and neither explains itself.

The governance is the same on both sides even though nothing else is, and that is the sentence to carry away from a comparison whose whole business has been separating them. The person affected is the same kind of person too. A declined applicant and a customer whose payment was stopped are both people waiting on an institution to explain itself, and neither of them can be answered by pointing at a component.

Try it out

Name something the two components really do have in common.

Hypothesis Testing teaches you to run a test, say what it can and cannot support, and recognise a manufactured result.

Who needs this distinction, and what does each of them do with it?

Somebody in the position Neelima Rao was in, reviewing an arrangement she did not build, uses it as a sorting question before anything else. Ask of each component, what is it measured against and when does that measurement arrive. If the answer comes back in days, ask what changed in the set since last month and what the month's confirmed cases say about which lines earn their volume. If the answer comes back in months, stop asking about outcomes entirely and ask instead what is being watched in the meantime: the inputs, the stability of the spread across months, and the behaviour on a set of files.

Somebody running a desk, in the position Ismail Sheikh holds, uses it to set expectations about feedback. A fraud desk gets a verdict on almost everything it works, within days, so a person there can improve and can see themselves improving. An exception desk working referral files gets no verdict at all in any timeframe a person feels. The absence is not a failing of that desk, and it does change what supervision there has to look like. Telling one desk it should learn from outcomes the way the other one does is asking for something the clock does not permit.

Somebody reading a monthly pack uses it as a rule about arithmetic: never add or divide across the two. 452 application files routed and 18,000 transaction alerts raised are not 18,452 of anything. An alert rate and a decline rate never go in one column. And somebody on the other side of either arrangement, waiting on a handset, gets the most useful version of all. The useful question is not how good the model is, but what each arrangement is measured against and when that measurement arrives. On one side an answer exists this week; on the other no answer arrives until month 21. Both of those are the institution's to hold, and neither is a question about the person.

How either component works inside is set out under credit decisioning systems and under fraud detection and transaction monitoring, what a declined applicant is told and how a reason is produced under adverse action, how alerts are triaged, escalated and worked under alert triage and escalation, and the cut-off trade under the decision threshold. How a component of either kind is fitted, validated or evaluated is a separate subject.

Sources

SourceDocumentSite
Reserve Bank of IndiaPublished expectations on a regulated lender covering digital lending, fair practice, the treatment of a borrower and the use of data and consent, and what applies to transaction monitoring at a bank in Indiarbi.org.in
Securities and Exchange Board of IndiaEquivalent expectations where the deployer of an automatic decision arrangement is a market intermediarysebi.gov.in
Financial Action Task ForceOrigin of the international standard on transaction monitoring, with the position for a bank in India stated by the Reserve Bank of Indiafatf-gafi.org
Ministry of Corporate AffairsThe accountability of a board for what a company does, and the reporting line a named accountable person for a component ultimately sits onmca.gov.in
Cathy O'NeilWeapons of Math Destruction, 2016, on decisions whose errors fall unevenly across a population and leave no record where they landCrown

Sumeru Bank Limited, its retail loan intake chain, Revathi Balan, Ismail Sheikh, Neelima Rao and Ashok Pillai are invented.
Educational material. Not advice on any investment, tax, budget or market position.

Comparison

Other comparisons in Credit and Fraud Systems

Comparison

Credit Score vs Credit Decision: The Number and the Cut-Off

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.