Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
AI, Automation & Digital Finance
1AI Foundations
Artificial Intelligence in FinanceAlgorithmNeural Networks and Deep LearningMachine LearningArtificial Intelligence vs Machine…Computer Vision in FinanceTraining Data and LabelsNatural Language Processing in Finance
2Generative AI
Generative AIGenerative AI vs Predictive AILarge Language ModelsEmbeddingsHallucinationFine TuningPrompting vs Fine TuningThe PromptThe Context WindowTool CallingGroundingVector DatabasesRetrieval Augmented GenerationRAG vs Fine Tuning
3Automation and Workflow
Workflow AutomationAutomation vs AugmentationHow to Map a…Straight-Through Processing and Exception…Robotic Process AutomationRule EnginesMachine Learning vs Rule-Based…
4Document and Operations AI
Intelligent Document ProcessingBatch vs Real-Time vs…Document Classification vs Entity…Service Level AgreementsCase ManagementHow to Document Data…Reconciliation AutomationOptical Character Recognition and Data ExtractionConfidence Scores
5Customer Systems, Identity and Digital Assets
Digital IdentityConsent ManagementBlockchain and Distributed LedgerChatbots and Conversational AIFrom Use Case to ProductionDigital Assets and TokenisationDigital SignaturesData Sharing in FinanceElectronic KYC and Digital Onboarding
6Credit and Fraud Systems
The Fraud AlertCredit Decisioning SystemsHuman in the Loop…Adverse ActionAnomaly DetectionThe Decision ThresholdCredit Score vs Credit DecisionAlert Triage and EscalationFraud Detection and Transaction MonitoringFraud Model vs Credit ModelHow to Build Human…
7Governance, Data and Vendors
AI Governance and the AI PolicyHow to Create an…Explainability and Interpretability ComparedThe AI VendorBias and Fairness in Financial AIShadow AIAccess Control and Data MinimisationCloud Computing in FinanceData Lineage and Master DataData ResidencyThe AI Use Case Register and Model InventoryThe Model Owner
8Model Performance, Monitoring and Resilience
Model DriftFalse Positives and False NegativesClassification MetricsAdversarial AttacksModel TestingBias, Fairness and Explainability…Stopping an Automated SystemModel ValidationAI Governance vs Model Risk ManagementPrompt InjectionHow to Create an…

Generative AI vs Predictive AI: Text Against a Value

A predictive component returns a value: a score, a class, a probability. A system acts on it, it is checked by comparing many outputs against what actually happened, and it fails by drifting. A generative component returns material. A person acts on it, it is checked one output at a time against the material it was given, and it fails by asserting something nothing supports.

Generative and predictive components get grouped together because both were built by related means, and almost nothing anybody deploying them has to decide follows from that shared origin. Every decision a deployer does have to make follows from something much plainer: what comes out, and who reads it. Every practical difference between the two traces back to the shape of the output rather than to the method that produced it. A difference resting on the output will still hold when the methods have moved on.

What does each of the two actually produce?

A school makes the plainest starting point. Both objects are familiar there. At the end of a term a child comes home with two things on one sheet. There is a mark, say 62 out of 100. And there is a comment: three sentences from a teacher about how the term went. Nobody confuses the two, and nobody would dream of handling them the same way.

The mark is a value. The mark gets used without being read in any meaningful sense: it goes into an average, it crosses a cut-off for the next set, it decides a prize list. A clerk copies it into a column. If the marking got quietly harsher this year, no single mark looks wrong, and the only way anybody would find out is by lining up many years of marks against what those students actually went on to do.

The comment is material. Somebody reads every word of it, and a parent will hold on to one sentence of it for a decade. If the comment says the child won the reading prize and she did not, that is instantly and obviously wrong to exactly one reader, and it never shows up in any average anywhere.

A predictive componentOne that returns a value such as a score, a class or a probability, rather than something written out in words. produces the mark, and a generative componentOne that produces new material, usually text a person can read straight off. produces the comment. That is the whole of the distinction, and the four rows that follow from it are the ones anybody deploying either has to settle. How either was built appears in none of those rows. Grouping them by method therefore puts two very different objects under one heading.

Sumeru Bank Limited, an invented mid-sized Indian bank, runs both inside one retail personal loan intake chain. Component 6, its scoring model, returns a value. Component 8, its drafting assistant, returns paragraphs. Set the two side by side and the differences read off.

Five rows a deployer has to settle. The method appears in none of them. WHAT HAS TO BE DECIDED PREDICTIVE: A VALUE COMES OUT GENERATIVE: MATERIAL COMES OUT 1 What comes out A score, a class, a probability. One number on one file. Sentences somebody can read straight off the screen. 2 Who acts on it Another system, with nobody reading the output at all. A person, who reads it, checks it and puts a signature on it. 3 How it is checked Many outputs set against what actually happened afterwards. One output set against the material it was handed. 4 How it fails It drifts. Same component, moving output. It asserts. It states what nothing supports. 5 What the failure hits A rate, across thousands of files, before anybody looks. One sentence, in one letter, read by one person. How it was built Behaviour derived from data Behaviour derived from data The one row where the two agree is the row every policy heading is written around. Sumeru Bank Limited is invented.
Set on five rows a deployer has to settle, the two agree on nothing except the last row, which describes how each was built and decides none of the five above it.
Try it out

What does each of the two produce?

AI For Finance Bootcamp — Fin Maverick

Who acts on the output, and how long does that take?

Follow one file through Sumeru Bank Limited's chain and time the two outputs. Component 6 returns its value and the workflow router acts on it immediately. Nobody opens it, nobody reads it, nobody forms a view about it. On the straight-through path the whole chain returns an answer to the applicant in about four minutes. No part of that chain has a person considering the score, so consideration takes up none of the four minutes. In month six, the first steady month, the arrangement determined the outcome of 5,981 of the month's 8,600 decisions, or 69.5 per cent.

Component 8 returns its draft and then the file stops moving. A desk officer opens the note, reads it against the material it was drafted from, works through five numbered checks and signs it. The reading takes nine minutes, measured. The desk does it 3,010 times a month, once for every file routed to a person. Multiplied out, the desk spends 27,090 minutes, or 451.5 hours, on reading in a month. The predictive component consumes zero minutes of human attention and the generative one consumes 451.5 hours a month, and that gap is not a detail of staffing, it is the whole governance difference between them.

Ismail Sheikh, who runs the exception desk, describes the difference the way any operations head would. The score is something his desk lives with. His people's names go on the bottom of the draft, so his desk is answerable for the draft sentence by sentence.

The same chain, two outputs, and only one of them waits for a person. PREDICTIVE PATH, COMPONENT 6, THE SCORING MODEL A value is returned one number on one file The router acts on it nobody opens the output The decision stands 5,981 of 8,600 in the month TIME SPENT WITH A PERSON: NONE, ON ANY OF THE 5,981 GENERATIVE PATH, COMPONENT 8, THE DRAFTING ASSISTANT A draft is returned paragraphs on one file A person reads it five checks, nine minutes Signed, then sent 3,010 notes in the month TIME SPENT WITH A PERSON: 9 MINUTES x 3,010 NOTES = 27,090 MINUTES, BEING 451.5 HOURS Month six figures, the invented bank's own. No product, model or supplier is named.
Drawn as two paths from a component to its effect, the score reaches the decision with no person in the way while the draft spends nine minutes with a reader, 3,010 times a month.
Try it out

One of the two outputs costs 451.5 hours of reading a month. What is that number actually saying?

Financial Analyst Program Bootcamp — Fin Maverick

What does each one need supplied before it can work at all?

Think about a wedding. Two jobs land on the same table. The first is working out how many people will actually turn up, so the caterer can be told a number. The second is writing the card that goes to one particular aunt who has been unwell. Both are hard, and they are hard in completely different ways.

The headcount needs the past: how many were invited to the last four weddings in the same circle, and how many walked in. Nothing about this particular aunt helps at all. The card needs the present: her name spelled right, what she is recovering from, whether she came last time. Four previous weddings supply nothing that can go in that card.

A predictive component needs past cases with outcomes attached, and a generative component needs the material for this one output, in front of it, at the moment it runs. Those two requirements do not substitute for each other in either direction, and each one is a real project before anything can be deployed.

Sumeru Bank Limited fitted component 6 on a past window holding 3,00,000 applications. Only 2,40,000 of them, or 80.0 per cent, were accepted and therefore had an outcome anybody could observe. The other 60,000 were declined, and a declined application never shows whether it would have gone bad. The bank then had to choose what bad even meant, and it chose 90 days past due observed over 12 months, both of which are its own choices and are not a standard anybody set. On those definitions 8,160 of the 2,40,000 carried a bad label, or 3.4 per cent. A past that has already finished happening is the supply side of a value.

Component 8 needs nothing of the sort and could not use it. The one thing component 8 needs is the file in front of it now. The material supplied with one exception note averages about 4,800 words, and the longest five per cent run past 12,000. Alongside it sits a store of 11,400 policies, procedures and product terms, cut into 47,000 passages, from which 6 passages are pulled for each question the drafting step has to answer. The store, the cutting and the pulling exist to solve a problem the predictive side does not have: assembling, freshly and for every single output, the material that output is allowed to rest on.

Two supply problems that do not substitute for each other in either direction. THE PREDICTIVE SIDE NEEDS A PAST 3,00,000 past applications in the window 2,40,000 accepted, so an outcome exists 80.0 per cent of the window 60,000 declined, so no outcome to see 12 months before a label exists, on the bank's own definition 8,160 carried a bad label, 3.4 per cent THE GENERATIVE SIDE NEEDS A PRESENT 1 file, this one, right now 4,800 words of material on average 12,000 words passed by the longest five per cent of files 47,000 passages cut from 11,400 policies and product terms 6 passages returned per question ONE WAITS A YEAR FOR ITS SUPPLY. THE OTHER ASSEMBLES ITS SUPPLY AGAIN FOR EVERY OUTPUT. Every count belongs to one invented bank and one deployment. The 90 days and the 12 months are that bank's own choices.
One side waits roughly a year for outcomes to exist before a single label can be written, while the other assembles fresh material for every individual output it produces.
Try it out

What does a generative component need that a predictive one does not?

How is each one checked, and by whom?

Checking sounds like one activity with two settings. Checking is in fact two activities, and they share no equipment. A population checkChecking a component by setting many of its outputs against what actually happened afterwards, rather than looking at any one of them. gathers a lot of outputs and sets them against what happened afterwards. Nobody reads any single output while doing it. An output checkChecking one output on its own, against the material that output was given, without waiting for any later event. takes one output and reads it against the material that output was handed. Nobody waits for anything to happen while doing it.

Component 6 gets the first. Revathi Balan, as the named accountable person for it, does not review a score; there is nothing in a score to review. She reviews a distribution: what share of files were accepted this month against last, how the accepted files behaved once they were on the book, whether the shape of the output has moved. The population check has one awkward property built into it. Nothing can start until outcomes exist, and on this bank's own definition an outcome takes twelve months to arrive.

Component 8 gets the second, and it happens 3,010 times a month at the desk rather than once a quarter in a review. The nine minutes is five numbered checks: every figure traced back to the supplied material, every assertion traced to a passage, nothing asserted that the material does not contain, the required format and statements present, and the decline-to-answer case handled properly where it arises. Asking the reader to notice something that is not there makes the third of those five the slow one, and nothing in any document is harder to notice.

Notice what settles which check a component needs. Not how it was built, not what it is called on the register, not which decade it came from. One question about the output, asked before anybody opens a policy: does a person read this output before it takes effect?

One question, asked of the output, settles which check the component needs. Does a person read this output before it takes effect? Nothing here asks how it was built. YES NO CHECK ONE OUTPUT AT A TIME against the material it was given, at the moment it is produced. 3,010 times a month. CHECK MANY OUTPUTS AT ONCE against what actually happened afterwards, once the outcomes have had time to arrive. The question is about the output, so it can be answered before anybody agrees what the component should be called.
Asking whether a person reads the output before it takes effect places the component with one check or the other, without anybody needing to agree on terminology first.
Try it out

A thousand outputs are compared against what actually happened. Which failure can that find?

Investment Banking Analyst Bootcamp — Fin Maverick

How does each of the two fail?

Two words name the two failures, and holding them apart is most of the work. The predictive side fails by driftA learned component's behaviour moving over time although nothing inside the component itself has changed.. The generative side fails by assertion failureAn output stating something the material it was given does not support, said in the same confident tone as everything else in the output..

Drift means the component did not change and its behaviour did. Nobody edited it, nobody refitted it, nobody approved anything. The change sat upstream of the component, in the data arriving. At Sumeru Bank Limited an income field changed format on one channel in month 8, and the reading step downstream of that change began handing the scoring model a slightly different picture of the same applicants. The component behaved exactly as it always had on what it was given. The input was something else.

Assertion failure means the output states something and nothing in the file supports it. Not a typing error, not a confusion of tone: a sentence that reads exactly like the true sentences around it and is not anchored to anything. Measured across 200 exception notes in month six, 23 drafts contained such a statement, or 11.5 per cent, about one draft in nine. Twenty-one of the 23 were caught by the person doing the nine minutes, and 2 were not, and those 2 reached a customer.

Try it out

Name each one's characteristic failure in one word.

Breaking Into Quants Bootcamp — Fin Maverick

What does each failure look like from outside?

Set side by side on a desk, the two do not look like the same category of problem. Not the same size, not the same shape, not the same kind of evidence.

The drift at Sumeru Bank Limited looked like a line bending. Six weeks passed between the field changing format and monitoring flagging it. About 12,900 files were decided in that window, 430 a working day across 30 working days, and 176 of them moved out of acceptance and into the referral bandThe range of scores the bank routes to a person rather than deciding automatically either way.. The month's approval rate fell from 57.0 per cent to 55.6, and monthly referrals rose from 391 to 508, a rise of 117 that matches the fall in accepts exactly. Every one of those files got a correct-looking decision. Nothing in any individual file was wrong on its face.

The assertion failure looked like one sentence. A paragraph in a letter to one applicant, saying something the file never contained, in the same steady tone as the sentences on either side of it. No aggregate anywhere in the bank moved by so much as a decimal when that sentence went out, and that is precisely what makes it invisible to the machinery built for the other failure.

Two failures in the same guide. Only one of them has a shape that can be plotted. DRIFT: A RATE MOVING WITH NOTHING CHANGED 57.0 56.5 56.0 55.5 57.0 per cent 55.6 per cent week 0 week 6 12,900 files decided, 176 moved into the referral band, nothing in the component changed. ASSERTION: ONE SENTENCE IN ONE LETTER one sentence the file never supported SIGNED, DATED AND SENT 23 of 200 drafts carried one. 21 were caught in the nine minutes and 2 were not, and no aggregate in the bank moved when they went.
One failure appears as an approval rate bending from 57.0 per cent to 55.6 across six weeks, and the other appears as a single sentence in a single letter that moves no total at all.
Reading an Option Payoff — free micro-course from Fin Maverick

Would either one's check have found the other one's failure?

Try it out

Two components, two failures, two checks already running. Would either check have caught the other one's failure?

Neither, in either direction, and it is worth being precise about why rather than treating it as a slogan.

The population check cannot see an assertion failure because the assertion failure does not disturb any population. One wrong sentence in one letter changes no rate, no distribution, no total. The population check could run on every letter the bank sent for a year, and because it never opens a letter, the sentence would sit there untouched.

The output check cannot see drift because drift is not visible in one output. The desk officer reading a note about a file in the referral band sees a note that correctly describes a file in the referral band. She has no way of knowing that the file would have been accepted six weeks earlier. A firm running both kinds of component therefore needs two controls, and a stricter version of one of them is not a substitute for the other.

Each check finds one failure and is blind to the other. There is no diagonal. DRIFT a rate moving, nothing changed ASSERTION FAILURE one sentence nothing supports POPULATION CHECK many outputs against what happened after FINDS IT the rate is what moved BLIND no total moves at all OUTPUT CHECK one output against the material it was given BLIND one file looks correct FINDS IT the sentence is right there Two controls, not one strict one. Drawn from one invented bank's two measured failures.
Comparing many outputs against what happened finds drift and cannot see an unsupported sentence, while reading one draft finds the sentence and cannot see drift.
Cleaning Financial Data teaches you to find the errors that survive every check and break every model.

Where does one deployed chain put each of them?

Position in the chain matters more than either name. Component 6 sits at the decision itself, and in month six it determined the outcome of 5,981 files: 4,902 accepted automatically, 688 declined automatically, and 391 more that landed in the referral band and went to a person. Component 8 sits after the decision, writing about something already settled, and it determined the outcome of none of the month's 8,600 files.

The component that decides nothing is the one whose actual words a customer reads, and the component that decides 5,981 outcomes produces nothing a customer will ever see. Anybody ranking the two by importance is going to rank them the wrong way round at least once, depending on which unit they happen to be holding at the time. So do not rank them. State each one's reach in its own unit and let the two numbers sit beside each other without a comparison operator between them.

Reach, stated in its own unitComponent 6, the scoring modelComponent 8, the drafting assistant
Decisions determined in the month5,981 of 8,600None of the 8,600
Outputs produced in the month8,600 scores3,010 notes
Minutes a person spends on the output09 each
Words of it a customer may readNoneAll of them
Each one's reach in its own unit, with no comparison operator between them. COMPONENT 6, DECISIONS DETERMINED IN THE MONTH 5,981 of 8,600, being 69.5 per cent COMPONENT 8, NOTES PRODUCED IN THE MONTH 3,010 notes, every one read and signed by a person COMPONENT 8, DECISIONS DETERMINED IN THE MONTH None and yet its wording is the part an applicant actually reads Bars share one scale. Month six, one invented bank, its own counts.
The scoring model determined 5,981 of the month's decisions and the drafting assistant determined none of them, while producing 3,010 notes a person had to read.
Try it out

How many of the month's 8,600 decisions did the drafting assistant determine?

Rebalancing: When, Why and What It Costs — free micro-course from Fin Maverick

What happens when one policy heading covers both?

Sumeru Bank Limited's first policy on this subject did the natural thing. Because both components had derived their behaviour from data, the policy put them under one heading and required quarterly performance monitoring of every component in scopeCaught by the wording of a policy, so its approvals, records and reviews apply to that component.. One heading, one control, applied evenly. Read on paper it looks careful.

On component 6 that control is exactly the right one, and it is the control that eventually surfaced the drift. Quarterly is the correct cadence for a component whose behaviour moves slowly across a population, and comparing a quarter's outputs against outcomes is precisely the check drift is visible to.

On component 8 the same control produces a quarterly report about a component whose failures happen one note at a time, in a nine minute window, at a desk. Three months of notes is 9,030 outputs. The report sampled 30 of them, one output in 301. At the 11.5 per cent already measured on 200 notes, a quarter of 9,030 drafts carries in the order of 1,038 with an unsupported statement in them, and a sample of 30 would be expected to contain three or four. The 1,038 is arithmetic on a measured rate rather than a count anybody made.

The error that gets made, and what it costs

The quarterly report was produced correctly and on time, by people doing precisely what the policy asked of them. Nobody skipped a step, nobody missed a deadline, and the sampling was carried out as written. Negligence explains none of it, and reading the episode as negligence makes the same failure far more likely somewhere else.

The fault sits one level up, in the scoping choice. A single heading covered two components that shared nothing except how they had been built, so a single control was written for both. Applied to the scoring model it was the right control and it found the drift. Applied to the drafting assistant, the same control produced a document that could not, even in principle, have noticed the 2 unsupported statements a measurement of 200 notes had already found reaching customers. No quarterly aggregate contains those 2.

The cost is a quarter of false comfort. A signed report existed, saying the component had been monitored. Nobody reading that report could tell that the only control on the failure that component actually has was the nine minutes at the desk, unmeasured by the report and uncounted by anybody. The second version of the policy separated the two headings, and separating them early is a cheaper repair than the one that follows an incident.

One heading, one control, two components. Correct for one of them. POLICY HEADING: QUARTERLY PERFORMANCE MONITORING OF EVERY COMPONENT IN SCOPE ON COMPONENT 6, THE SCORING MODEL A quarter of outputs against outcomes is the check drift is visible to. IT FOUND THE DRIFT ON COMPONENT 8, THE DRAFTING ASSISTANT The failure happens once a note, inside nine minutes, at a desk. IT SAMPLED 30 OF 9,030 SEVEN ROWS OF 43 SLOTS IS THE QUARTER. EACH SLOT IS 30 DRAFTS. ONE SLOT WAS READ: THE 30 IN THE QUARTERLY SAMPLE, BEING ONE OUTPUT IN 301 Seven rows of 43 slots is 301, and 301 slots of 30 drafts is the quarter's 9,030. One invented bank's own counts.
Quarterly monitoring found the drift in the scoring model and read one draft in 301 of the drafting assistant's quarter, which is not a control on a failure happening one note at a time.
Try it out

A quarterly report samples 30 of 9,030 drafts. What has that control established?

India

Where a supervisor sits in all this

Sumeru Bank Limited is a regulated lender, so the expectations covering outsourcing, digital lending, customer data, consent and record keeping sit with the Reserve Bank of India at rbi.org.in, and where such a system is deployed by a market intermediary the Securities and Exchange Board of India at sebi.gov.in is the relevant body. The current material at those sites governs what applies to a given firm.

One policy heading covered both, and quarterly monitoring fitted neither. See what wording misses.

What do the two genuinely have in common?

The two do share one property, and it is worth being fair about. The shared property is a useful one, and it is not the one people usually name.

Neither carries a written procedure a reviewer can read. Sumeru Bank Limited has four components in the same chain that do: the income corroboration rule is 34 lines, and Neelima Rao read all 34 of them end to end in 25 minutes. There is nothing equivalent to open for either component 6 or component 8. Whatever governs their behaviour was derived rather than typed, so a reviewer arriving with an afternoon and a highlighter has nothing to highlight.

The absence of a written procedure has a precise consequence, and the precise wording matters because a loose wording circulates. Repeatability does survive: put the same input in and the same output comes back, for as long as the fitted numbers stay where they are. Repeatability survives in the form that can be tested rather than the form that can be read, and it ends at the next refit; readability does not survive at all. So both components are reviewed by looking at data and behaviour instead of at instructions, and that shared property is what makes it so tempting to write one policy heading over the two. The similarity is genuine, and it has no bearing on anything the checking has to do.

Try it out

What do the two actually have in common?

Why is there nothing to slide between the two?

Much of this subject is a relationship, where one quantity can be taken hold of and a second watched to respond. The split between a value and written material is not a relationship. An output is read by a person before it takes effect or it is not, and there is no position between those two states for anything to slide along. The two failure pictures do the work a dial would otherwise do, and both come from measurements Sumeru Bank Limited made on two of its own components in the same year.

What does this change about the way checking is done?

Three people, three different Mondays

Whoever writes a monitoring plan chooses a unit before choosing a frequency, and the unit is the part that gets skipped. For a component whose output no person reads, the unit is a population and the frequency is whatever lets outcomes arrive; at Sumeru Bank Limited that meant waiting twelve months on the bank's own definition before a label existed to compare anything against. For a component whose output a person reads, the unit is one output and the frequency is every single time. A plan that says quarterly for both has picked a frequency without ever picking a unit.

Whoever staffs a desk should read the 451.5 hours as a running cost rather than a transition cost. Because the reader cannot know in advance which note is the faulty one, reading time on a generative output does not taper as the component improves. The entire point of the check is that the fault is invisible until read. When Sumeru Bank Limited's error rate fell from 11.5 per cent to 3.5 in month ten, the verification time per note fell from nine minutes to six, a real gain. Six minutes is not zero, and the time was never going to reach zero.

And whoever answers a complaint needs to know which of the two produced the thing being complained about, before picking up the file. A complaint about an outcome is answered from the decision and whatever can be attributed about it. A complaint about a sentence in a letter is answered from that note and the material it was drafted from, and no monitoring report will contain either. Getting that wrong costs a fortnight and a second letter, and only what came out of the component settles it, never the component's name.

How either kind of component works is covered elsewhere in this subject area, along with the statistics beneath each. How a model is fitted, validated or evaluated is covered separately. Model risk work as an organised discipline, and where its boundary with this comparison falls, is covered separately. What causes drift, and how a monitoring programme catches it, is covered under monitoring.

Sources

SourceDocumentSite
Reserve Bank of IndiaExpectations on a regulated lender covering outsourcing, digital lending, customer data, consent and record keepingrbi.org.in
Securities and Exchange Board of IndiaExpectations where the deployer of such a system is a market intermediarysebi.gov.in
Agrawal, Gans and GoldfarbPrediction Machines, 2018, on a fitted component producing a prediction somebody still has to act onHarvard Business Review Press

Sumeru Bank Limited, Revathi Balan, Ismail Sheikh and Neelima Rao are invented.
Educational material. Not advice on any investment, tax, budget or market position.

Comparison

Other comparisons in Generative AI

Comparison

Prompting vs Fine Tuning: Changing the Instruction or the Model

Comparison

RAG vs Fine Tuning: Supplying Material or Changing the Model

← PreviousNext →
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.