Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
AI, Automation & Digital Finance
1AI Foundations
Artificial Intelligence in FinanceAlgorithmNeural Networks and Deep LearningMachine LearningArtificial Intelligence vs Machine…Computer Vision in FinanceTraining Data and LabelsNatural Language Processing in Finance
2Generative AI
Generative AIGenerative AI vs Predictive AILarge Language ModelsEmbeddingsHallucinationFine TuningPrompting vs Fine TuningThe PromptThe Context WindowTool CallingGroundingVector DatabasesRetrieval Augmented GenerationRAG vs Fine Tuning
3Automation and Workflow
Workflow AutomationAutomation vs AugmentationHow to Map a…Straight-Through Processing and Exception…Robotic Process AutomationRule EnginesMachine Learning vs Rule-Based…
4Document and Operations AI
Intelligent Document ProcessingBatch vs Real-Time vs…Document Classification vs Entity…Service Level AgreementsCase ManagementHow to Document Data…Reconciliation AutomationOptical Character Recognition and Data ExtractionConfidence Scores
5Customer Systems, Identity and Digital Assets
Digital IdentityConsent ManagementBlockchain and Distributed LedgerChatbots and Conversational AIFrom Use Case to ProductionDigital Assets and TokenisationDigital SignaturesData Sharing in FinanceElectronic KYC and Digital Onboarding
6Credit and Fraud Systems
The Fraud AlertCredit Decisioning SystemsHuman in the Loop…Adverse ActionAnomaly DetectionThe Decision ThresholdCredit Score vs Credit DecisionAlert Triage and EscalationFraud Detection and Transaction MonitoringFraud Model vs Credit ModelHow to Build Human…
7Governance, Data and Vendors
AI Governance and the AI PolicyHow to Create an…Explainability and Interpretability ComparedThe AI VendorBias and Fairness in Financial AIShadow AIAccess Control and Data MinimisationCloud Computing in FinanceData Lineage and Master DataData ResidencyThe AI Use Case Register and Model InventoryThe Model Owner
8Model Performance, Monitoring and Resilience
Model DriftFalse Positives and False NegativesClassification MetricsAdversarial AttacksModel TestingBias, Fairness and Explainability…Stopping an Automated SystemModel ValidationAI Governance vs Model Risk ManagementPrompt InjectionHow to Create an…

Confidence Scores: What the Number Means and How to Threshold It

A confidence value is the component's own reading of how sure it is about one output, on its own scale. A confidence value is not the probability the answer is right, and it is not comparable across components or across field types. A threshold turns that value into a decision to accept or to route, and moving it trades wrong answers accepted against work sent to a person.

The only thing that connects a confidence value to correctness is somebody checking a sample of accepted answers by hand. Without that check the number is an ordering and nothing more: genuinely useful for deciding which answers to look at first, and worthless for saying how many of the accepted ones are wrong. Everything difficult about thresholds follows from that one gap between a number a component produces about itself and a number a person produced by looking.

What is a confidence value actually reporting?

An ordinary case makes the point. A clerk hands a photocopied form to a colleague across the desk and asks her to read out the account number. She reads it out, and then she adds something she was not asked for: "I am fairly sure, though that third digit could be a five or a six." That second sentence is a confidence value. Her sureness is a reading of her own reading. It is not evidence about the account number. Nobody has looked at the original. If she is the sort of person who is fairly sure about everything, the sentence carries almost no information at all, and that could only be found out by pulling the original and checking.

A component behaves the same way. At Sumeru Bank Limited, an invented mid-sized Indian bank running a retail loan intake chain, the field reading step returns a value between 0.00 and 1.00 beside every field it produces. The month holds 8,600 completed files at 14 fields each, so 120,400 field values arrive with 120,400 confidence valuesA component's own reading of how sure it is about one output, on its own scale. attached. The figures are the bank's own and describe one deployment in one month.

What is the number honestly good for? Ordering. Sorted by confidence value, the month's fields at the bottom of the list really are the ones most worth a person's attention first. The ordering survives almost every complaint that can be made about the numbers themselves, and it is not a small thing: an ordering is what makes a queue of 120,400 items workable at all. A confidence value is an ordering before it is anything else, and almost everything a chain does well with one is an ordering rather than a measurement.

The trouble starts the moment somebody draws a line across that ordering and calls everything above it correct. Drawing the line is unavoidable and completely reasonable. Calling what sits above it correct is the error, and it is an error that survives review meetings because nothing in the paperwork looks wrong.

Try it out

What is a confidence value a reading of?

Debt Capital Markets Bootcamp — Fin Maverick

What are the four things a confidence value is not?

Four denials do most of the teaching here, and they are worth learning as a numbered list because each one is a different mistake made by a different person in a different meeting.

  1. A confidence value is not the probability the answer is right. A value of 0.95 is a point on a scale the component made up while producing the answer. Nothing in that arithmetic forces it to correspond to being right 95 times in 100.
  2. The value is not comparable across components. The document classifier and the reading step both emit numbers between 0.00 and 1.00. The two numbers sit on separate scales that happen to share a range, in the way that two different examiners both mark out of ten.
  3. The value is not comparable across field types. A six character amount and a twenty character employer name are not the same reading problem, and a value of 0.93 on each does not describe the same situation.
  4. The value is not stable across a refit. Refit the component on newer documents and the values move. The acceptance barThe confidence level above which an output is taken without a person looking at it. written into the configuration does not move, because a bar is a number in a file and nothing refits it.

Every one of the four denials says the same thing in a different costume: the value is internal to one component at one moment, and every use anybody wants to make of it is external. Somebody wants to compare it to a target, to another team's figure, to last quarter, to a supplier's claim. None of those comparisons is supported by the number itself.

One thing it is. Four things it is not. WHAT IT IS The component's own reading of how sure it is about one output, on a scale that belongs to that component and to nothing else. NOT 1: A PROBABILITY 0.95 does not mean right 95 times in 100. Only a hand check on accepted output can say that. NOT 2: PORTABLE 0.90 from a classifier and 0.90 from a reading step are two scales. A bar cannot be carried between them. NOT 3: ONE FIGURE A short amount and a long employer name are different problems. 0.93 on each does not describe one thing. NOT 4: STABLE Refit the component and the values move. The bar in the configuration file does not. All four denials are the same denial: the value is internal to one component at one moment, and every use anybody wants to make of it reaches outside that component.
A confidence value reports one component's sureness on its own scale, and it is not a probability of being right, not portable between components, not comparable across field types and not stable when the component is refitted.

The second denial looks like a saving, and that is exactly why it is the one that costs money. A team that has already done the work of choosing a bar for one component has a number in hand, and carrying that number to the next component looks like reuse. Carrying the number across is not reuse. It is a fresh decision taken without any of the evidence, wearing the decimal places of an old decision that had some.

Two components. The same printed number. Two different scales. THE DOCUMENT CLASSIFIER 0.90 on one document Sureness about a statement of the form: this is a pay slip. One answer for the whole document, on scale A. THE FIELD READING STEP 0.90 on one field Sureness about a statement of the form: the net pay is 38,000. Fourteen answers for each file, on scale B. NO BAR CARRIES This bank chose a bar for the reading step, on evidence about the reading step, and never carried it anywhere else.
A value of 0.90 from the document classifier and a value of 0.90 from the field reading step sit on two separate scales, so a bar chosen on evidence about one component says nothing about the other.
Try it out

Can a bar chosen for one component be reused on another?

Why does a value of 0.95 not mean right 95 times in 100?

Two people mark the same pile of examination scripts. Both are asked to write a percentage sureness beside each mark. One is temperamentally cautious and writes 70 when he is nearly certain. The other writes 95 for everything he has ever got right. Their judgement might be identical. Their numbers are not. The numbers were never anchored to anything outside the marker's head, so no arithmetic on the numbers alone will show which of the two to trust.

A component is in exactly that position, and there is nothing negligent about it. The value came out of the same machinery that produced the answer. Whether a value of 0.95 corresponds to being right about 95 times in 100 is a separate property with its own name, calibrationWhether a value of 0.9 corresponds to being right nine times in ten, which is a separate subject in its own right., and how a component is calibrated, fitted or evaluated is covered separately. The plain consequence is what matters: nothing in the arithmetic of a confidence value forces 0.95 to mean right 95 times in 100, and at this bank nobody ever claimed it did.

A great deal still works without that property. The ordering holds. The queue of fields most worth checking is still the right queue. One kind of sentence does not work at all: fields are accepted above 0.92, so at most 8 per cent of accepted fields can be wrong. The sentence is arithmetic performed on a quantity that does not support arithmetic, and it appears in more threshold papers than anybody would like.

What connects a confidence value to whether the answer is right?

One thing only, and it is unglamorous. A person takes a sample of the outputs the chain accepted, opens the underlying documents, and checks each one. The exercise is called a hand checkA person checking a sample of accepted outputs against the source, which is what connects the value to correctness., and it is the only bridge between a component's opinion of itself and a fact about the world.

The word doing the work in that paragraph is accepted. Sampling the routed fields instead is tempting: they are already in front of people and the sampling costs almost nothing. The routed sample answers a question nobody needs answered. Routed fields already go to a person. Nobody knows what is quietly going wrong inside the 95 per cent that never stopped, and nothing except a hand check on accepted output will say. A hand check on routed output measures the work already being done, and a hand check on accepted output measures the risk already being carried.

A sweet shop keeps the same discipline when it opens one box from a batch of two hundred rather than tasting the ones the packer already set aside as doubtful. Checking what was rejected describes the rejecting. Checking what passed describes what the customers are about to eat.

AI For Finance Bootcamp — Fin Maverick

What did one bank find when it checked 400 accepted fields?

Sumeru Bank Limited set its own acceptance bar at 0.92. The bar of 0.92 is this bank's own choice, taken for this bank's own chain, and a bar chosen that way is not a standard, a benchmark or anything anybody else is expected to match. At that bar, of the month's 120,400 fields, 114,380 are accepted without a person, being 95.0 per cent, and 6,020 are routed to a person, being 5.0 per cent.

Then the bank did the only thing that makes those numbers mean anything. The bank pulled 400 accepted fields, opened the documents behind them, and checked. Seven were wrong. Seven in 400 is 1.75 per cent. Projected across 114,380 accepted fields, that rate implies about 2,002 wrong field values sitting inside the month's accepted work, untouched, unflagged and already used.

The month at the bank's own bar of 0.92FieldsShare
Field values produced by the reading step1,20,400100.0 per cent
Accepted without a person, above 0.921,14,38095.0 per cent
Routed to a person, below 0.926,0205.0 per cent
Of the accepted, wrong on the hand check reading of 1.75 per centabout 2,0021.75 per cent of accepted

Every one of the 400 checked fields had cleared a bar of 0.92 and 7 of them were wrong, which is a fact no reading of the confidence values on their own would ever have produced. The two sentences set next to each other hold the whole subject. The chain reports 95.0 per cent accepted. The chain does not report 1.75 per cent wrong, and cannot. The 1.75 per cent came from a person with a document open beside a screen.

Two honest caveats belong beside that finding, and the bank stated both. Seven is a small count, and how a sample size is chosen is covered separately. And 1.75 per cent is a reading of this component, on these documents, at this bar, in this month. The rate is not a property of reading steps in general.

400 accepted fields. Every one of them cleared 0.92. One square, one field. Seven came back wrong when a person opened the document behind it. WHAT THE CHECK PRODUCED 7 of 400 1.75 per cent of accepted fields Projected across the 1,14,380 fields accepted in the month, that is about 2,002 wrong field values accepted, used and never flagged. The confidence values alone said nothing about any of the seven. All seven were above the bar and looked exactly like the other 393.
All 400 sampled fields had cleared the bank's own bar of 0.92 and 7 of them were wrong, being 1.75 per cent, which projects to about 2,002 wrong field values accepted across the month.
Try it out

Every one of 400 fields cleared a bar of 0.92 and 7 were wrong. What does that establish about what the bar means?

How a bar is chosen, and who chooses it

Five steps, in this order, and the order is the teaching. Most threshold papers start at step two and skip step one entirely. Skipping step one is why most threshold papers carry a wrong-field figure that nobody measured.

  1. Hand check a sample of accepted outputs. Not routed outputs. The hand check is the step that produces a wrong-field rate at all, and without it every later figure is an assumption with a decimal point.
  2. Sweep the bar. Run the month's fields against several candidate bars and record, at each one, how many are accepted, how many of those are wrong on the checked rate, and how many are routed.
  3. Price the extra routing in desk minutes. A routed fieldA field sent to a person because its confidence sat below the bar. is not free. A routed field becomes a stopped file, a case, a person's minutes and eventually a post somebody has to fund.
  4. Price the wrong outputs in what they cause downstream. A wrong field is not free either, and what it costs depends entirely on which field it is and what reads it next.
  5. Choose, and record who chose it. A bar with no name against it is a setting rather than a decision, and settings drift.

Who chooses? Not the team that built the component, whose evidence is essential and whose incentive is to accept more. The choice belongs with the people who carry the two consequences. At this bank those are Revathi Balan, head of retail credit, who carries what a wrong accepted field does to a lending decision, and Ismail Sheikh, head of the exception desk, who carries every routed field as minutes on a queue. A threshold is the point where two different people's problems are traded against each other, so it cannot sensibly be set by either of them alone or by anybody who carries neither.

Choosing a bar is five steps, and the first one is a hand check rather than a sweep. 1 Hand check a sample of ACCEPTED outputs 2 Sweep the bar and record what is accepted, wrong and routed at each 3 Price the extra routing in desk minutes and then in posts 4 Price the wrong outputs in what they cause further down the chain 5 Choose, and record who chose it and on what Step 1 is the one that gets skipped, and skipping it does not stop the paper being written. It stops the paper being about anything: without a hand check, every wrong-field figure in steps 2 and 4 is an assumption somebody typed, and the sweep becomes a distribution of a number nobody has checked.
The five steps for choosing an acceptance bar run from a hand check of accepted outputs through a sweep and two pricings to a recorded decision, and skipping the first step leaves every later figure unmeasured.
Try it out

A team proposes a bar based on a sweep alone. What is missing?

Financial Analyst Program Bootcamp — Fin Maverick

What does the bank's own sweep across four bars show?

A threshold sweepMeasuring what is accepted, what is wrong and what is routed at several candidate bars. is the second step, and it is mechanical once the first is done. Run the same month of 120,400 fields against four candidate bars, apply the hand-checked wrong rate measured at each, and write down three columns. Sumeru Bank Limited did exactly that, and every row of the result sums back to 120,400. Summing back to the month's total is the first thing to check on anybody's sweep.

Acceptance barAcceptedOf those, wrongRouted to a person
0.851,17,7513,415 at 2.9 per cent2,649
0.92, the bank's own choice1,14,3802,002 at 1.75 per cent6,020
0.961,05,952954 at 0.9 per cent14,448
0.9985,484256 at 0.3 per cent34,916

Read the two right-hand columns as a pair rather than one after the other, because that is where the decision lives. As the bar rises the wrong accepted fields fall from 3,415 to 256, a fall of 3,159. Over the same range the routed fields rise from 2,649 to 34,916, a rise of 32,267. The wrong column falls by about three thousand fields across the whole range and the routed column rises by more than thirty two thousand, so the two sides of a threshold are not two views of one quantity, they are two quantities moving at completely different speeds.

Fields a month, against the acceptance bar. Both lines drawn on one scale. 0 10,000 20,000 30,000 0.85 0.92 0.96 0.99 the bank sits here FIELDS ROUTED TO A PERSON 2,649 at 0.85, rising to 34,916 at 0.99 WRONG FIELDS ACCEPTED 3,415 at 0.85, falling to 256 at 0.99 The red line falls 3,159 fields across the whole range. The green line rises 32,267, more than ten times as far.
Across the four bars the wrong accepted fields fall from 3,415 to 256 while the routed fields rise from 2,649 to 34,916, so on one scale the red line is nearly flat and the green line climbs off the top.
Try it out

Before the control is moved: the bar goes from 0.92 to 0.96. Wrong accepted fields fall by 1,048. How much more work reaches a person?

Play with it

Move the acceptance bar, and watch the two sides refuse to move together

One input moves: the acceptance bar, from 0.85 to 0.99. Two consequences redraw against one shared scale, the wrong fields accepted in a month and the fields routed to a person, plus what the routed fields become as posts on the exception desk. The dashed outlines behind every bar stay where the bank sits today at 0.92, so the movement is visible without remembering anything.

0.850.920.99
The acceptance bar 0.92 today 0.96 proposed 0.85 0.88 0.92 0.96 0.99 A month at this bar, in fields Wrong fields accepted 2,002 Fields routed to a person 6,020 0 10,000 20,000 30,000 and at the exception desk, in posts Posts needed 6.81 6.81 posts, the desk today Every dashed outline is where the bank sits at 0.92. Everything solid is the bar currently chosen.
Wrong accepted
2,002
Routed to a person
6,020
Posts at the desk
6.81
Against 0.92
the bar today

At the bank's own bar of 0.92 the month accepts 1,14,380 fields, about 2,002 of them wrong on the hand-checked rate, and routes 6,020 to a person, which is the 1,264 stopped files and the 6.81 posts the exception desk actually runs on.

The readings that anchor this control. At 0.92, the bank's own bar, 2,002 wrong fields accepted and 6,020 routed. At 0.96, 954 wrong and 14,448 routed. Between the two, 1,048 wrong fields prevented for 8,428 more fields sent to a person, being 8.04 extra fields for a person for every wrong field prevented, and at an assumed fully loaded Rs 9,00,000/- a post that difference is Rs 36,00,000/- a year.
Educational illustration. Every figure here is the bank's own, describing one deployment in one month: 1,20,400 fields, four measured bars with wrong-field shares taken from the bank's own hand checks, and straight-line interpolation between those four points, which is a simplification rather than a measurement. Routed fields become files at the month's observed clustering of 4.76 routed fields to a stopped file, the other five exception causes are held at their measured 1,746 files, handling time is held at 19 minutes and a working month at 8,400 minutes a person. The bar of 0.92 is the bank's own choice and is never a standard.
Cleaning Financial Data — free micro-course from Fin Maverick

What is the exchange rate between a wrong field prevented and work sent to a person?

Both sides of a threshold are now measured, so they can be put in one number. Between the bank's own bar of 0.92 and the bar of 0.96 above it, wrong accepted fields fall from 2,002 to 954, so 1,048 wrong fields are prevented. Routed fields rise from 6,020 to 14,448, so 8,428 more fields go to a person. Dividing the second by the first gives 8.04.

The figure 8.04 is an exchange rateExtra work sent to a person for each wrong output prevented by raising the bar., and it is the single most useful number on a threshold paper. An exchange rate turns an argument into a question anybody can answer. Moving from 0.92 to 0.96 costs 8.04 extra fields for a person for every wrong field it prevents, and a person who cannot say whether that is worth paying has not been given a decision to take. The virtue of the number is that it is arguable. Somebody can say yes, a wrong income figure in a lending file is worth eight fields of somebody's attention. Somebody else can say no. Both are having the right conversation. Neither could have had it from a slide reporting only that errors fall by half.

Raising the bar from 0.92 to 0.96, both sides drawn on one scale. WRONG FIELDS ACCEPTED at 0.92 2,002 at 0.96 954 1,048 wrong fields prevented FIELDS ROUTED TO A PERSON at 0.92 6,020 at 0.96 14,448 8,428 more fields for a person 8.04 extra fields for a person for every wrong field prevented. 8,428 divided by 1,048. That is the whole trade in one number.
Between 0.92 and 0.96 the bank prevents 1,048 wrong accepted fields and sends 8,428 more fields to a person, which is 8.04 extra fields of work for every wrong field prevented.
Cleaning Financial Data teaches you to find the errors that survive every check and break every model.

Why does the exchange rate get worse as the bar rises?

Because the rate is not constant, and treating it as one is the second big error on this subject. Work the same division on each pair of adjacent bars in the sweep and three very different numbers come out. From 0.85 to 0.92 the rate is 2.39. From 0.92 to 0.96 it is 8.04. From 0.96 to 0.99 it is 29.3. The price of preventing one wrong field rises more than twelve times across the usable range, so where a bar already sits changes what moving it means far more than how far it moves.

The reason is worth building an intuition for rather than memorising. Picture the gate of a housing society where the guard is told to telephone a flat whenever he is not fully sure the person at the gate lives there. At a loose rule he phones only about the genuinely unfamiliar faces, and a fair share of those calls catch somebody who should not be walking in. Tighten the rule and he starts phoning about residents he has seen a hundred times but cannot place today. He is still catching the occasional outsider, but now he makes forty calls to do it. Nothing about the outsiders changed. The composition of everybody he is now stopping is what changed.

The same composition effect runs the sweep. Fields sitting just under a low bar are a mixed population with plenty of genuinely bad readings in them. Almost everything the component was uncertain about has already been routed at the lower bars, so fields sitting just under a high bar are overwhelmingly correct. Each step up the scale therefore sweeps in many more right answers for each additional wrong one it catches, and the ratio deteriorates accordingly.

Extra fields for a person, for every wrong field prevented. 0 10 20 30 2.39 0.85 to 0.92 8.04 0.92 to 0.96 29.3 0.96 to 0.99 What sits just below a high bar is overwhelmingly correct already, so each step up sweeps in far more right answers for each wrong one it catches.
The exchange rate is 2.39 extra routed fields for each wrong field prevented between 0.85 and 0.92, 8.04 between 0.92 and 0.96, and 29.3 between 0.96 and 0.99.
Try it out

The rate is 2.39 from 0.85 to 0.92 and 29.3 from 0.96 to 0.99. Why?

Fund Waterfalls and Carry — free micro-course from Fin Maverick

What does the choice of bar cost the desk, in people?

Nobody funds a field, so fields are the wrong unit for the second half of the trade. Convert. The month's routed fields do not spread evenly across files; they sit together, at an observed clusteringFailures sitting together in a few files rather than spreading evenly across many. of 4.76 routed fields to each stopped file. At 0.92 the clustering puts the 6,020 routed fields inside 1,264 stopped files, an exception cause the desk already knows. At 0.96 the 14,448 routed fields sit in about 3,035 files instead.

Hold the other five exception causes at their measured 1,746 files and the desk's month rises from 3,010 exceptions to 4,781. At the desk's handling time of 19 minutes a file, that is 90,839 minutes against 57,190. At an assumed working month of 8,400 minutes a person, that is 10.81 posts against 6.81. Moving the bar from 0.92 to 0.96 is a request for four more people, and stating it that way is the only version of the sentence anybody can actually approve or refuse.

The exception deskAt 0.92, the bank's own barAt 0.96
Routed fields in the month6,02014,448
Stopped files, at 4.76 routed fields each1,264about 3,035
Total exceptions, with the other causes held at 1,7463,0104,781
Desk minutes, at 19 minutes a file57,19090,839
Posts, at 8,400 minutes a person a month6.8110.81
The same threshold move, stated three ways. The three rows are one movement rescaled. Exceptions a month 0.92 3,010 0.96 4,781 Desk minutes a month 0.92 57,190 0.96 90,839 Posts on the desk 0.92 6.81 0.96 10.81 +4.00 posts, four more people At an assumed fully loaded Rs 9,00,000/- a post, four posts is Rs 36,00,000/- a year. Across the year that prevents 12,576 wrong fields, being about Rs 286/- a field.
Raising the bar from 0.92 to 0.96 takes the desk from 3,010 exceptions and 6.81 posts to 4,781 exceptions and 10.81 posts, which is four more people at about Rs 36,00,000/- a year.
Try it out

State the move from 0.92 to 0.96 in people rather than in fields.

The paper that carried one side of the trade

The bar was raised to 0.96 in a review that presented a single number. Wrong accepted fields would fall from 2,002 a month to 954, a reduction of 1,048. The reduction of 1,048 was correct, it came from a real hand check, and nobody in the room had any reason to doubt it.

The paper did not carry the other column. Routed fields rise from 6,020 to 14,448, an increase of 8,428, being 8.04 extra fields for a person for every wrong field prevented. Held at the month's clustering of 4.76, those 14,448 fields sit in about 3,035 stopped files against 1,264, so the desk's month rises from 3,010 exceptions to 4,781, and at 19 minutes each that is 90,839 minutes against 57,190. At the assumed working month of 8,400 minutes a person that is 10.81 posts against 6.81. Four more people. At the bank's assumed fully loaded Rs 9,00,000/- a year, that is Rs 36,00,000/- a year to prevent 12,576 wrong fields across the year, about Rs 286/- a field.

The failure is not that raising the bar was wrong. Keeping a wrong income figure out of a lending decision may well be worth Rs 286/-. The failure is narrower and harder: a threshold paper carrying one side of the trade is not a decision, it is half of one, and the half that was missing is the half somebody would have had to fund.

Nobody funds a field, so the unit is wrong. See what the desk pays.

What has to be recorded about a threshold, and when is it revisited?

The record is short enough to fit on one screen, and being that short is exactly why nobody writes it. Five lines: the bar itself, who chose it, the hand check it rested on with its date and its sample size, the sweep it was chosen from, and the date it is next revisited. The five lines are the whole artefact.

The date and sample size are not bureaucratic decoration, and they are the two lines most often dropped. The bar means nothing without the check behind it, and the check ages. Refit the component and the values shift underneath a bar that did not move. Start receiving a new kind of document and the population the check was drawn from is no longer the population running through the chain. A threshold with no recorded check behind it is a setting, and a setting nobody revisits will eventually be defended in a review as though somebody had chosen it.

Recording the sample size does one more thing, and it is the honest one. Seven wrong in 400 is a small count. A reader of the record needs to be able to see that for themselves rather than meeting 1.75 per cent as though it arrived with three decimal places of authority. Write the 7 and the 400 down, not just the percentage.

The threshold record. Five lines, and it is the thing nobody writes. ACCEPTANCE BAR, FIELD READING STEP 1 The bar 0.92 2 Chosen by Revathi Balan, with Ismail Sheikh 3 Hand check it rests on 7 wrong in 400 accepted, month 6 4 Chosen from the sweep at 0.85, 0.92, 0.96 and 0.99 5 Revisited On any refit, and at the annual review WHY EACH LINE IS THERE 1  The bar is the bank's own choice, not a standard, and the record says so. 2  A name, so it is a decision rather than a setting somebody inherited. 3  The count and the date, because the bar means nothing without the check and the check ages. 4  The alternatives, so a later reader sees what was rejected and at what price. 5  A refit moves the values underneath a bar that did not move, so a refit is itself a trigger to look again.
A threshold record carries the bar, who chose it, the hand check with its date and sample size, the sweep it came from and what triggers a revisit, which together are what stop a chosen bar decaying into an inherited setting.
Try it out

Why record the date and sample size of the hand check beside the bar?

India

Who sets the expectations where a read field feeds a lending decision

Where an automated reading feeds a credit decision at a regulated lender, the expectations on controls, outsourcing, customer data and record keeping sit with the Reserve Bank of India, published at rbi.org.in, and with the Securities and Exchange Board of India at sebi.gov.in where the deployer is a market intermediary. The acceptance bar of 0.92 is the invented bank's own choice for its own chain and is not a standard, a benchmark or an expectation of any authority. Both authorities amend those expectations from time to time, so what binds a lender on any given day is what stands at rbi.org.in and sebi.gov.in that day.

Breaking Into Quants Bootcamp — Fin Maverick

What does a person handed one of these numbers actually do with it?

Three people meet a confidence figure in the ordinary course of a week, and each of them wants a different thing from it.

The head of an operations desk uses it as a staffing forecast. Ismail Sheikh does not care what 0.92 means philosophically; he cares that the bar and the clustering together produce 1,264 stopped files, 3,010 exceptions and 6.81 posts, and that a change to the bar arrives on his desk as a change to his roster with a lag. He uses one arithmetic chain: routed fields, over clustering, into files, times handling time, over the working month, into posts. The last number in that chain is the number being asked for, so anybody proposing a threshold change to an operations head should arrive with it already worked out.

A reviewer in a risk function uses it as a question. Neelima Rao, reviewing this chain, would ask one thing first: show me the hand check on accepted output, its date and its sample size. If the answer is that there isn't one, then every wrong-field figure in the paper is an assumption, and the review is about assumptions rather than about a threshold. The question about the hand check separates a measured threshold from a decorated one faster than any other.

Somebody buying a reading service from outside uses it as a translation test. A supplier's claim of high accuracy is an accuracy on the supplier's documents at the supplier's bar. The buyer's question is not whether the claim is true but what it becomes on their own documents at their own bar, and the only way to find out is to run their own hand check on accepted output after go-live. Ajay Agrawal, Joshua Gans and Avi Goldfarb's framing in Prediction Machines, 2018, is useful here: a prediction is only worth what somebody does differently because of it, and a confidence value that changes nobody's routing and nobody's checking has cost money and bought nothing.

All three uses share one habit worth taking away whole. Each person converts the number out of the component's units into their own before arguing about it: into posts, into a question about evidence, into what happens on their own documents. The confidence value itself never left the component, and it was never going to.

How characters are recovered from an image and assembled into fields, the step just before this in the same chain, is covered separately. How the exception desk handles what is routed to it sits with the rest of the routing and workflow material. How a component is fitted, calibrated or evaluated, and how a sample size is chosen, are each subjects in their own right.

Sources

SourceDocumentSite
Reserve Bank of IndiaExpectations on a regulated lender covering digital lending, outsourcing, customer data, consent and record keeping, which apply to the lender whatever component reads the documents behind a decision and whatever bar it accepts them atrbi.org.in
Securities and Exchange Board of IndiaExpectations on a market intermediary where the same kind of automated reading sits inside a regulated intermediary rather than a lendersebi.gov.in
Ajay Agrawal, Joshua Gans and Avi GoldfarbPrediction Machines, 2018, on a prediction being worth only what somebody does differently because of it, the test a confidence value has to meetHarvard Business Review Press

Sumeru Bank Limited, Revathi Balan, Ismail Sheikh and Neelima Rao are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← Previous
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.