Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryFinancial LiteracyInvestment Banking Analyst
Private Equity AnalystHedge Funds AnalystBreaking Into VCBreaking Into QuantsAI For Finance
Financial Analyst ProgramRisk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Internships
Equity Research InternMutual Fund Intern
Portfolio Management InternFinancial Literacy Intern
Explore Micro Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
Courses
Explore Career Roadmaps
Investment Banking AnalystEquity Research AnalystVC AnalystPrivate Equity AnalystHedge Funds Analyst
Quant AnalystAI For FinanceFinancial Analyst ProgramPrivate Wealth ManagementDebt Capital Markets
Risk Management ProgramDerivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
AI, Automation & Digital Finance
1AI Foundations
Artificial Intelligence in FinanceAlgorithmNeural Networks and Deep LearningMachine LearningArtificial Intelligence vs Machine…Computer Vision in FinanceTraining Data and LabelsNatural Language Processing in Finance
2Generative AI
Generative AIGenerative AI vs Predictive AILarge Language ModelsEmbeddingsHallucinationFine TuningPrompting vs Fine TuningThe PromptThe Context WindowTool CallingGroundingVector DatabasesRetrieval Augmented GenerationRAG vs Fine Tuning
3Automation and Workflow
Workflow AutomationAutomation vs AugmentationHow to Map a…Straight-Through Processing and Exception…Robotic Process AutomationRule EnginesMachine Learning vs Rule-Based…
4Document and Operations AI
Intelligent Document ProcessingBatch vs Real-Time vs…Document Classification vs Entity…Service Level AgreementsCase ManagementHow to Document Data…Reconciliation AutomationOptical Character Recognition and Data ExtractionConfidence Scores
5Customer Systems, Identity and Digital Assets
Digital IdentityConsent ManagementBlockchain and Distributed LedgerChatbots and Conversational AIFrom Use Case to ProductionDigital Assets and TokenisationDigital SignaturesData Sharing in FinanceElectronic KYC and Digital Onboarding
6Credit and Fraud Systems
The Fraud AlertCredit Decisioning SystemsHuman in the Loop…Adverse ActionAnomaly DetectionThe Decision ThresholdCredit Score vs Credit DecisionAlert Triage and EscalationFraud Detection and Transaction MonitoringFraud Model vs Credit ModelHow to Build Human…
7Governance, Data and Vendors
AI Governance and the AI PolicyHow to Create an…Explainability and Interpretability ComparedThe AI VendorBias and Fairness in Financial AIShadow AIAccess Control and Data MinimisationCloud Computing in FinanceData Lineage and Master DataData ResidencyThe AI Use Case Register and Model InventoryThe Model Owner
8Model Performance, Monitoring and Resilience
Model DriftFalse Positives and False NegativesClassification MetricsAdversarial AttacksModel TestingBias, Fairness and Explainability…Stopping an Automated SystemModel ValidationAI Governance vs Model Risk ManagementPrompt InjectionHow to Create an…

How to Create an AI Incident-Response Playbook

A playbook is eight stages written down before anything happens: detect, confirm, contain, tell, fix the cause, work out who was affected, put it right for them, and change what let it happen. At one bank the fix took four working days. The record did not carry what the corrected step would have produced, so working out who was affected took eight more.

A deployed chain accumulates a great deal of discipline long before anything goes wrong. The nine numbered components of one deployed chain, the nine monitoring signals and the three the bank actually watched, the three levels of stopping and the six numbered actions a stop requires, the seven pre-go-live tests and the three that were skipped, the six numbered entries in the audit trail, the eight data requirement items, the eleven working days of independent challenge. Each of those was settled on its own, and not one of them says what happens on the morning something has already gone wrong. A playbook is the document that puts all of it in an order and attaches a clock to it, and the clock is the only part that ever gets measured afterwards.

What counts as an incident when nothing actually broke?

Picture a chemist on a busy street. On a Tuesday morning the wholesaler sends a batch in which one strip has been mislabelled. The shop does not shut. The till works, the shutters go up on time, the staff are polite, the queue moves. For six weeks the shop trades exactly as it always has and sells that strip to whoever asks for it. Nothing in the shop failed. The whole of the problem is in what left the shop, in the hands of people who have gone home.

That is the shape of an incidentAn episode where a deployed arrangement produced outcomes it should not have, whether or not anything stopped working. in an automated decision chain, and it is why an ordinary technology incident procedure covers only the first half of it. An ordinary procedure is written for the case where a service stops answering: the alert fires because something is down, the containment is a restart or a switch to another site, and when the service answers again the episode is over. Here the service never stopped answering. The chain answered every time, in about four minutes, politely and at volume, and every answer went to a person who then arranged their finances around it.

So what makes this different is not severity and not technical difficulty. The difference is that the episode arrives with people attached to it. Because nothing stopped, the episode has an affected populationThe people whose outcomes were produced inside the suspect window, which is what makes this different from a system failure. rather than an outage window, and every stage after containment is about those people rather than about the system. An outage has a start time, an end time and a status notice. An episode of this kind has a list of names, and somebody has to produce that list.

At Sumeru Bank Limited, invented, that is exactly what happened. In month 8, week 2, an upstream income field began arriving in a changed format on one channel. In month 9, week 3, the monitoring flagged it. The bank's month holds twenty working days, so six weeks is thirty of them, and at the chain's four hundred and thirty files a working day that window held about 12,900 decided files. Of those, 176 came out in a different place from where they would otherwise have come out. Nobody logged a fault, because there was no fault to log. Every fitted number inside the scoring model sat precisely where somebody had put it. The chain was in perfect health throughout.

Notice what that does to the vocabulary. Nothing was ever unavailable and nothing needed restoring, so there is no downtime to report, no availability figure to restate and no restoration time. The measures that remain are a count of files, a count of people, and a date by which each of those people was made whole. A firm that measures its response in minutes of downtime will report a flawless month here, and it will be telling the truth about a question nobody asked.

Try it out

An automated decision chain produced worse outcomes for six weeks and never stopped running. What makes that different from an ordinary system failure?

Breaking Into Quants Bootcamp — Fin Maverick

What are the eight stages, and in what order do they happen?

Eight, in this order, and the order is not negotiable because each stage needs the output of the one before it. Nothing can be contained before it has been confirmed. Nothing can be put right for anybody until it is known who they are. A design cannot sensibly be changed until the cost of the episode is known. A stage with no named output is a stage nobody can ever declare finished, so the eight are written down with an output beside each one.

  1. DetectThe moment a signal speaks. Output: a dated flag naming which signal moved, what it read, and what it was compared against.
    Finished when: somebody who can act on it has the flag in front of them.
  2. ConfirmEstablish that the reading is real and not the monitoring misfiring. Output: one written statement naming the suspected fault and the date it is believed to have started.
    Finished when: a named person has written down what they think is wrong and since when.
  3. ContainStop the affected population growing. Output: the arrangement no longer producing the suspect outcomes, by fallback, by narrowing, or by a full stop.
    Finished when: no new file can join the affected population.
  4. TellSay what is happening, inside the firm and outside it. Output: a dated list of who was told, when, and what they were told.
    Finished when: the people whose files are held know their files are held.
  5. Fix the causeCorrect the component, the rule or the feed. Output: the corrected version in service and the containment lifted.
    Finished when: new files are being decided correctly and somebody has checked a sample of them by hand.
  6. Work out who was affectedProduce the list. Output: a numbered set of files decided inside the suspect window whose outcome the fault actually changed.
    Finished when: a list with a count on it exists and the count can be defended.
  7. Put it right for themRe-decide each file on that list and act on the result. Output: every file on the list carrying a fresh decision, and the customer told.
    Finished when: the list is empty and each entry has a dated outcome against it.
  8. Change what let it happenClose the specific gaps this episode exposed. Output: named changes with dates against them, not principles.
    Finished when: somebody could point at each change and say which stage it would have shortened.

Notice where the effort actually sits: stages 1 to 4 are about the system and are the ones every firm rehearses, and stages 5 to 8 are about the people and are the ones nobody has ever performed. Most incident procedures in most firms end somewhere around stage 5. The older procedures were written for outages, and an outage produces no wrong outcomes to unwind, so fixing the cause genuinely is the end of it. Carrying that habit into an arrangement that decides things about customers produces half a procedure with a whole procedure's name on it.

The three levels of stopping and the six numbered actions a stop actually requires are set out under stopping an automated system. The playbook's concern is where those levels and those actions attach. ContainmentStopping the arrangement producing more of the same outcomes, which is not the same as fixing the cause. is stage 3 and it is where the levels and the first five of those six actions live. The sixth action, deciding what happens to the files already decided in the suspect window, does not live at stage 3 at all. The sixth action belongs to stages 6 and 7, and it is the reason those two stages exist as separate entries.

Eight stages. The four that get rehearsed, and the four that do not. STAGES 1 TO 4, ABOUT THE SYSTEM 1 Detect out: a dated flag naming the signal that moved 2 Confirm out: the fault named and the date it started 3 Contain out: the affected count stops growing 4 Tell out: a dated list of who was told and when STAGES 5 TO 8, ABOUT THE PEOPLE 5 Fix the cause out: corrected version in service, hold lifted 6 Who was hit out: a numbered list with a defensible count 7 Put it right out: every file re-decided and the customer told 8 Change it out: named changes with dates, not principles Most written procedures end somewhere in the top row. Stage 6 is where the real ones stall.
The eight stages run in a fixed order because each needs the output of the one before it, and the four that concern the affected people all sit after the four that concern the system, which is the opposite of where rehearsal usually goes.
Try it out

A firm's incident procedure ends at fixing the cause. What has it left out?

AI For Finance Bootcamp — Fin Maverick

What does each stage need written down before anything happens?

A playbook is not the eight stage names. Anybody can produce the eight stage names in an afternoon, and a document containing only the eight stage names is exactly as useful during an episode as a menu is during a famine. A document becomes a playbook only when each stage carries four things, decided in advance, at a time when nothing is at stake and there is room to argue about them.

Name the person who does it, not the team. A team is nobody at half past ten at night. Name what they need in front of them, and be specific enough that somebody could go and check today whether that thing exists. Name what decides the stage is finished, in a form that admits a yes or a no. Name who is told when it is. A stage missing any one of those four is the stage that will stall, and which one that is can be found out on any ordinary Tuesday, without waiting for an incident.

The fourth of those, who is told, is what stops a playbook becoming a relay of people waiting for each other. At this bank Ismail Sheikh runs the exception desk, Revathi Balan is the named accountable person for the scoring model, Neelima Rao is in the risk function and did not build any part of the chain, and Ashok Pillai is in technology risk. The playbook's value is not that those four are competent. The value is that on the day it happens, nobody has to work out which of them stage 6 belongs to.

One playbook entry. Four things named against one stage. STAGE 6, WORK OUT WHO WAS AFFECTED WHO DOES IT a named person, never a team WHAT THEY NEED IN FRONT OF THEM a trail entry saying what the previous version would have produced WHAT DECIDES IT IS FINISHED a list with a defensible count WHO IS TOLD: THE DESK THAT WILL RE-DECIDE THE SAME ENTRY AT GO-LIVE WHO DOES IT a named person, never a team WHAT THEY NEED IN FRONT OF THEM blank WHAT DECIDES IT IS FINISHED a list with a defensible count WHO IS TOLD: THE DESK THAT WILL RE-DECIDE Three of the four are identical. The one blank line is what turned a query into eight working days. Sumeru Bank Limited, invented. The entry and its four items are that bank's own design.
Every playbook entry names who does the stage, what they need in front of them, what decides it is finished and who is told, and the entry that stalls is the one whose second line was never filled in.

What did the clock look like on one six week episode?

Here is the whole of it at this bank, counted in working days from the flag. Durations can be compared and stage names cannot, so the clock is what turns a list of eight nouns into something arguable.

StageWhat happenedWorking day from the flag
1 detectThe monitoring flags a movement. It is thirty working days lateday 0
2 confirmThe reading is real and the changed input format is named as the causeday 1
3 containThe affected channel falls back to manual decisioningday 1
4 tellInside the bank on day 1, the customers whose files were held on day 3day 1 and day 3
5 fix the causeThe reading step is corrected and the fallback ends after its 4 working daysday 4
6 work out who was affectedAbout 12,900 files re-scored against the corrected step to find 176day 12
7 put it right for themThe 176 re-decided by people, carried alongside normal volumeday 19
8 change what let it happenFour named changes made in the same monththat month

Read the third column as widths rather than as dates and the shape of the thing appears at once. Stages 2, 3 and 4 all landed inside the first working day. A competent operations function moves that fast when a flag arrives. Stage 5, correcting the reading step and ending the fallback, took three more working days. And then stage 6, a stage involving no code and no vendor at all, took eight. The stage that consumed the most of the response was the one asking a question of a record, and the record could not answer it.

Nothing in that clock is a criticism of the people on it. Every stage from 2 onwards was done by somebody working at pace on the day the work reached them. The reflex when a clock looks bad is to ask who was slow, and on this clock the honest answer is nobody. The slow part was a decision taken months earlier about what the record would carry.

The month 9 clock. Working days from the flag. 1 detect day 0, and thirty working days late 2 confirm 1 working day 3 contain day 1, the channel falls back 4 tell inside on day 1, customers on day 3 5 fix the cause done on day 4 6 who was affected 8 working days of re-scoring 7 put it right 7 working days 8 change it four named changes, same month 0 2 4 6 8 10 12 14 16 18 Sumeru Bank Limited, invented. Working days from the flag.
Drawn against a real clock the stages are wildly uneven: confirming and containing took one working day, fixing the cause took four, working out who was affected took twelve and putting it right took nineteen.

Why was fixing the cause the quick part of the whole thing?

Go back to the chemist. Once he knows which strip was mislabelled, taking it off the shelf is a two minute job, and it is a two minute job whether the batch has been on the shelf for one day or for six weeks. The size of the mistake does not touch the size of the repair. The flat cost of the repair is the first thing to understand about stage 5, and it is why every firm gets it right and every firm then assumes the rest will be similar.

Correcting the reading step at this bank was work of a fixed size. Somebody had to see what the changed format looked like, change how the field was read, test it, and put it into service. The correction took three working days from the confirmation on day 1, ending on day 4, and it would have taken three working days if the format had changed yesterday. The files are not an input to the repair, so the number that went through in the meantime does not touch how long it takes.

Stage 6 is the opposite. Every file decided inside the suspect windowThe stretch between a fault starting and being confirmed, and the files decided inside it. is an input to it, because in principle any of them might have come out differently, and the only way to know is to look. So the fix has a flat cost and the identification has a cost proportional to the window. A detection lag therefore charges twice: once in the outcomes produced while nobody knew, and once again in the working days it takes to find out whose outcomes they were.

Put the numbers on it and the two lines cross somewhere useful. At this bank the window ran to thirty working days and about 12,900 files, and the identification took eight working days, which is a re-scoring rate of about 1,612 files a working day, being 12,900 over 8, or 1,612.5 exactly. Hold that rate and the identification takes about 0.27 of a working day for every working day the window ran. Set that against a flat three working day repair and the two are equal at a window of about 11.25 working days. Below that, the fix is the longer half of the response. Above it, finding people is, and it gets worse without limit.

The fix does not grow with the window. Finding people does. stage 5, fix the cause: a flat 3 working days stage 6, find who was affected they are equal at about 11 working days of window this episode: 30 working days of window, 8 of identification 3 8 12 working days of work 10 20 30 40 length of the detection window, in working days Sumeru Bank Limited, invented. Drawn at that bank's own 430 files a working day and its own re-scoring rate of about 1,612 a working day.
Correcting the reading step took four working days whatever the window had been, and finding the affected files took eight because the window held about 12,900 of them, so the two costs behave completely differently as the lag grows.
Financial Analyst Program Bootcamp — Fin Maverick

Why did working out who was affected take twice as long as the fix?

Because of the question stage 6 asks. It is not "which files went through while this was broken", which any system can answer in seconds, and it is not "which files look wrong", because none of them looked wrong. It is a strange, backwards question: which of these files would have come out somewhere else if the fault had not been there? That is a question about something that did not happen, and a record built to say what did happen has no way of answering it.

So the bank did the only thing left. The bank took the corrected reading step and put about 12,900 already-decided files back through it, one by one, to see which of them landed somewhere different from where they had landed the first time. That is what re-scoringRunning decided files through a corrected step to find out which ones would have come out differently. means, and it is not a clever technique. Re-scoring is brute force applied because the cheap route was not available.

The answer, when it came, was 176 files. The 176 are 1.4 per cent of the window, and every one of the other 12,724 files was untouched by the fault. Eight working days of machine time and people's attention were spent establishing, for 98.6 per cent of the files, that nothing had happened to them. Those eight working days are not a measure of how bad the fault was. The eight days measure the size of the haystack, and the size of the haystack was set by the detection lag.

An account of an incident can easily leave a reader believing something worse than what happened, so the 176 are worth stating precisely. The 176 files moved out of the accept region and into the referral region, so a person looked at each of them. Of the 176, 152 were accepted by that person and 24 were not, being 13.6 per cent of the 176 and 0.19 per cent of the 12,900 files in the window. Not one file was declined by the chain itself that would not otherwise have been declined. What the fault produced was not a wave of wrong refusals; it was a shift of work onto people, and for 24 applicants an outcome that a corrected chain would not have produced.

There is a second cost in that shift, and it is worth naming because it lands on the same 176 files twice. Carrying them as referrals inside the window cost the exception desk 19 minutes each, being 3,344 minutes of unplanned handling. Re-deciding the same 176 at stage 7 cost 19 minutes each again, another 3,344 minutes. The identical figure appears on both sides of the correction: 6,688 minutes in total, for files that should never have reached a person at all.

Which stage does the record decide, and what has to be in it?

Only one of the eight. Detection is decided by the monitoring, containment by whether a fallback exists and has capacity, and putting it right by whether anybody funded the desk time. Stage 6 is the one stage whose whole duration is set by what somebody chose to write down months earlier, and that is why it is the stage worth arguing about while the arrangement is still being designed.

The audit trail on this chain carries six numbered entries against every decision: 1 which version of which component acted, 2 what it read, 3 what it produced, 4 what the previous version would have produced where that is known, 5 who could have intervened and did not, and 6 the time. Four of the six were recorded from go-live. Entries 4 and 5 were not. The two that were missing are precisely the two that answer what would have happened otherwise, and that is the only question stage 6 ever asks.

Entry 4 is the counterfactual fieldThe trail entry saying what the previous version of a component would have produced. and it is the expensive one. With it present, finding the affected files is a query: show me every file in this window where the current answer and the previous answer disagree. That is minutes. Without it, the same request becomes a re-scoring exercise across every file in the window, and at this bank that ran to eight working days. The gap between minutes and eight working days is one column in a table that nobody thought was worth the storage.

Entry 5, who could have intervened and did not, is the quieter one, and it decides something different. Entry 5 establishes, for each affected file, whether a person was in a position to catch the fault and let it pass. The question is not a hunt for somebody to blame. Entry 5 separates a fault that no arrangement could have caught from one where a review step existed and produced nothing, and those two findings lead to completely different changes at stage 8.

Six entries in the trail. Two of them blank on the day it went live. WHAT THE TRAIL CARRIES AGAINST EVERY DECISION AT GO-LIVE 1 which version of which component acted recorded 2 what it read recorded 3 what it produced recorded 4 what the previous version would have produced blank 5 who could have intervened and did not blank 6 the time recorded With entry 4 present, the 176 are a query. Without it, they are 12,900 files re-scored over 8 working days.
The trail did not carry what the previous version would have produced, so the 176 affected files had to be found by re-scoring about 12,900 of them at about 1,612 a working day, and two blank columns became eight working days of work.
Try it out

Why did finding the 176 affected files take eight working days?

The error that gets made, and what it costs

A playbook is written, approved and circulated. The playbook names eight stages, names an accountable person against each one, and is a genuinely good document. Nobody notices that stage 6 says "identify the affected population" and stops there. The sentence reads like an instruction and is in fact a wish. Nothing anywhere in the design says what the record must hold in order for anybody to carry it out.

The cost is measurable and it landed here as eight working days between a corrected system and a corrected customer, on a fault that took three working days to repair. In money the cost is small. Eight working days of re-scoring is machine time and one person's attention. The real cost is the eight working days of waiting borne by people who had no idea they were waiting.

A playbook that names eight stages and does not say what the record must hold has written down, in confident language, the one stage it cannot perform.

Investment Banking Analyst Bootcamp — Fin Maverick

What does putting it right actually cost, in minutes and in money?

Asked what it would take to put 176 wrongly handled loan applications right, a room will answer somewhere between a fortnight and a programme of work. The phrase sounds unbounded. The unbounded sound of the phrase is exactly why the work gets deferred, and why a playbook has to force somebody to multiply before anything happens rather than after.

Multiply. The exception desk at this bank measures its own handling at 19 minutes a case, so the figure is known rather than assumed. So putting it rightRe-deciding the affected files and acting on the result, which is work rather than a notification. for 176 files is 176 times 19, being 3,344 minutes. At the bank's assumed working day of 420 minutes that is 7.96 working days of one person's time. Carried alongside normal volume by a desk of seven people, the work ran across 7 working days, and stage 7 landed on day 19 having started on day 12.

Now price it. At the bank's assumed fully loaded cost of Rs 9,00,000/- a year a post, and 240 working days in its year, a working day of one person is Rs 3,750/-. So 7.96 working days is about Rs 29,850/-. Set that beside the Rs 65,00,000/- a year it costs this bank to run the whole chain and putting 176 people right came to about 0.46 per cent of one year's running cost. The stage that sounds unbounded is the cheapest thing on the entire clock, and the stage that sounds like paperwork, the two blank columns in the trail, is what actually cost eight working days.

Agrawal, Gans and Goldfarb, in Prediction Machines, 2018, argue that a learned component supplies a prediction and a person still has to act on it. Stage 7 is where that lands with full force. The corrected reading step can establish that the 176 files would have come out somewhere else. The corrected step cannot re-decide them, cannot write to the applicant, and cannot arrange what happens next. Re-deciding is a person doing 19 minutes of work, one hundred and seventy six times, and no version of stage 7 avoids it.

Putting it right, multiplied out before anybody has to argue about it. 176 files on the list from stage 6 x 19 min the desk's measured handling time a case 3,344 minutes of desk work in total /420 7.96 working days of one person at 420 minutes Carried alongside normal volume by a desk of seven, it ran across 7 working days. About Rs 29,850/- of desk time, being 0.46 per cent of one year of running the chain. Sumeru Bank Limited, invented. Priced at that bank's assumed fully loaded Rs 9,00,000/- a year a post over 240 working days.
Putting it right is ordinary work with ordinary arithmetic, 176 files at 19 minutes being 3,344 minutes or 7.96 working days of one person, and stating it that way is what stops the stage being deferred indefinitely.
Try it out

176 files have to be re-decided by people at 19 minutes each. Is that a project?

What did changing what let it happen actually produce?

Stage 8 is where most incident write-ups go to die, because the natural output of a review is a principle, and a principle costs nothing and changes nothing. Monitoring will be strengthened. Documentation will be improved. Nobody can point at a principle a year later and say whether it was done.

Sumeru Bank produced four changes, each of them a specific thing that either exists or does not. Trail entries 4 and 5 were added, so the counterfactual question becomes a query. The wrong-arrival test, the one asking how the chain behaves when a field is absent or arrives in an unexpected format, was added to the set of tests run before go-live, having previously been run only afterwards. The eight data requirement items were completed for the fourteen extracted fields, including the sixth item, what happens when the field is absent, which had been missing for eleven of the fourteen. And the input format signal, the share of each input field arriving in the expected format, was added to the monitoring.

Not one of the four changes is about the model, because nothing about the model was ever wrong. Every fitted number in the scoring component sat exactly where it had been put, before the episode, during it and after. Two of the four changes are about a record, one is about a test, and one is about a signal. A playbook that leaves a reader thinking the model failed has taught the wrong reflex, and the wrong reflex here is expensive: it sends people to inspect the one thing that was working.

Look at the fourth change and read it against the clock again. The input format signal would have spoken on the day the format changed, in month 8, week 2. Thirty working days of detection would have been close to none. The signal was not added because somebody was clever afterwards; it was added because an episode had made the absence of it visible, and most monitoring gets built in that honest and rather uncomfortable way.

Four changes. Two records, one test, one signal. None about the model. WHAT WAS CHANGED WHICH GAP IT CLOSES Trail entries 4 and 5 added the counterfactual, and who could act Stage 6 becomes a query rather than eight working days of re-scoring The wrong-arrival test moved into the set run before go-live A field arriving in an unexpected format is now tested for, not met live The absent-field item written down for all fourteen extracted fields It had been missing for eleven of the fourteen, being 78.6 per cent The input format signal added to what the monitoring watches This is the one that attacks the thirty working days of detection Sumeru Bank Limited, invented. All four are that bank's own changes.
The last stage produced four specific changes with dates against them rather than four principles, and each one closes a named gap the episode exposed instead of restating an intention.
Try it out

Four changes came out of this episode. What did none of them touch?

Hypothesis Testing — free micro-course from Fin Maverick

How much of the whole episode was detection, and what follows from that?

Put the two halves together. Thirty working days passed between the fault starting and the monitoring speaking. Nineteen working days passed between the monitoring speaking and the last affected file being put right. The episode ran forty nine working days end to end, of which detection was thirty, being 61.2 per cent. The longest stretch of the whole episode is the stretch during which nobody was doing anything at all, because nobody knew there was anything to do.

Sit with that for a moment, because it inverts where attention usually goes. Every hour of those nineteen response days had somebody working on it. People were confirming, containing, telling, correcting, re-scoring and re-deciding, at pace, and the clock still came to nineteen. Every one of the thirty detection days passed with the chain running normally, the desk working normally, the pack being produced on its usual cycle and not one person aware that anything needed attention. One of those two stretches can be compressed by working harder. The other cannot be compressed by working at all.

A firm reviewing this episode and concluding that it must respond faster has read its own clock backwards. Suppose the response were somehow halved, from nineteen working days to ten. The episode falls from forty nine to forty, an improvement of about 18 per cent, at the cost of everything a halved response would take. Now suppose instead the detection lag falls from thirty working days to two, which is what watching the input format signal actually buys. The episode falls from forty nine to twenty one. One of those two is a change of arrangement and the other is a change of column in a monitoring pack.

There is a second, stranger reading in the same numbers, and the simulation below is built on it. Shortening the response does not reduce detection's share of the episode; it increases it. With a complete trail, stage 6 collapses to a query, the response falls from nineteen working days to eleven, the whole episode falls from forty nine to forty one, and detection rises from 61.2 per cent to 73.2 per cent of the elapsed time. The better the response gets, the more nakedly the detection lag stands out as the thing left unfixed.

One episode, two halves. Both bars drawn on the same day scale. AS IT HAPPENED: 49 WORKING DAYS 30 working days detecting nobody was working on it 19 responding people working throughout 61.2 per cent 38.8 per cent WITH A COMPLETE TRAIL: 41 WORKING DAYS 30 working days detecting, unchanged 11 73.2 per cent 26.8 per cent 8 working days saved Sumeru Bank Limited, invented. Shortening the response makes detection a larger share of the episode, not a smaller one.
Detection dominated the whole episode at 61.2 per cent of the elapsed time and it is the only stretch nobody was working on, which reverses where a firm would otherwise put its next rupee.
Try it out

Detection was 61.2 per cent of the elapsed time. Where should the next rupee go?

Hypothesis Testing teaches you to run a test, say what it can and cannot support, and recognise a manufactured result.

What happens to the customer's date when the record is complete?

The clock has one movable part and it is not the one people reach for. The detection lag belongs to the monitoring arrangement and is covered under model drift and monitoring. The fix is a fixed size. The re-deciding is 176 files at 19 minutes and will not compress. The movable part is the middle: the number of working days between the system being right and the list of affected people existing, and that number is set entirely by what the trail carries.

Try it out

Before the control is moved. The cause was fixed on working day 4. On which working day was the last affected file put right?

Play with it

Move the identification stage and watch the customer's date move

One variable: how many working days stage 6 takes to produce the list of affected files. Everything else is held. The fix still completes on working day 4 from the flag, because correcting a reading step does not depend on how many files went through it, and re-deciding the 176 still takes 7 working days. Because the thirty working days of detection belong to the monitoring arrangement rather than to the playbook, they do not move.

0 days, a complete trail8 working days12 days
Working days from the fault. The flag falls on day 30. 30 working days before anybody knew the flag last file put right, day 19 from the flag 0 10 20 30 40 50 detecting stages 2 to 5, the fix stage 6, finding who was affected stage 7, re-deciding Sumeru Bank Limited, invented. 176 affected files at 19 minutes each, a 420 minute working day, a fix completing on working day 4.
Stage 6 takes
8 days
Last file put right
day 19
Whole episode
49 days
Detection share
61.2%
At 8 working days for stage 6, which is what happened, the last of the 176 affected files is put right on working day 19 from the flag, the whole episode runs to 49 working days and detection is 61.2 per cent of it.
Educational illustration. Sumeru Bank Limited, an invented lender, and every figure attached to it describe one deployment. Held constant: the fix completes on working day 4 from the flag, because correcting a reading step does not depend on how many files went through it, and re-deciding the 176 files at 19 minutes each is 3,344 minutes, being 7.96 working days of one person at 420 minutes, carried alongside normal volume across 7 working days. The thirty working days of detection are held constant here because they belong to the monitoring arrangement rather than to the playbook. THE MEASURED READING is stage 6 at 8 working days, giving working day 19 from the flag and a 49 working day episode of which detection is 61.2 per cent. THE READING AT ZERO, where the trail carries what the previous version would have produced and the affected files are a query rather than a re-scoring, is working day 11 and a 41 working day episode of which detection is 73.2 per cent.

Move it to zero and watch what happens to the two readouts on the right. The customer's date moves from working day 19 to working day 11, which is eight working days earlier for one hundred and seventy six people. The episode falls from forty nine working days to forty one. And detection's share rises from 61.2 per cent to 73.2 per cent, because the only thing that got shorter was the part somebody was working on. The single column nobody thought worth the storage is worth eight working days of somebody else's waiting, every time this happens.

How is a playbook tested without waiting for an incident?

The test is a drillRunning the playbook against an invented incident to find which stage nobody can actually perform., and where it is started decides whether the exercise is worth the morning. Most drills start at stage 1, because that is where an incident starts and it feels wrong to begin in the middle. Everybody then performs beautifully, because detecting, confirming, containing and telling are the four stages every operations function does every week for ordinary faults. The drill passes. Nothing is learned.

A drill that starts at stage 6 works differently. The drill takes a component from the nine, invents a plausible fault, picks a date it started, and puts one question to the room: who can say which files came out differently, and how long will that take? The question can be asked on any Tuesday, costs one meeting, and has only two possible answers. Either somebody names themselves, describes the query and gives a number in hours, or the room goes quiet and somebody eventually says they would have to run everything again.

The second answer is the finding, and it has been obtained for the price of a meeting rather than for the price of an episode. A quiet room shows that the record cannot answer the counterfactual question and therefore what to add, and it shows this before the eight working days are being counted by people who are waiting. The same exercise applies to stage 7 and asks who has the capacity to re-decide a list of that size while normal volume keeps arriving. Then to stage 3, and what the fallback's capacity actually is.

One more thing worth building into the drill: make somebody state the count before anybody looks. If the room believes an episode of this kind would touch a hundred files and the honest answer is a hundred and seventy six, that gap is itself a finding about how well the arrangement is understood. Independent validation at this bank ran for eleven working days at month 12 and produced seven findings, and one of those seven was precisely the absence of any procedure for files already decided in a suspect window. A drill is not a substitute for that challenge, and validation is not a substitute for detection. Each of the three finds a different kind of thing.

Where the drill starts decides whether anything is learned. RUN A DRILL Start at stage 1, detect the way an incident actually starts Start at stage 6, who was affected the stage nobody has ever performed Everybody performs. The drill passes. These four are done every week anyway. One question to the room: which files came out differently? Tested what the function was already good at a name and a query: answered in hours silence, or re-score all: that is the finding The right hand branch costs one meeting. The same finding, obtained in an episode, cost eight working days of other people waiting.
A drill that starts at stage one tests the stages people already perform every week, and the useful drill starts at stage six by asking who could say which files came out differently and how long it would take them.
Try it out

Which stage should a drill start at?

How does a lender, an analyst or a household read a promise to respond?

Start with the household, because the shape is the same and the stakes are legible. A school runs a bus. One morning the office discovers that for six weeks a form has been recording the wrong pick-up point for some children. Nothing crashed. The bus ran, on time, every day. Correcting the form takes an afternoon. Working out which children were affected takes as long as it takes to go back through six weeks of paper, and the parents are waiting for a phone call the whole time. Any parent asked which of those two numbers matters gives the same answer every time.

A lender buying a decisioning service from somebody else should ask for exactly that second number and nothing else. Not the policy document, not the incident procedure, not the availability figure. The last episode's clock: when it started, when the supplier knew, when it was fixed, when the affected were identified, and when the last of them was put right. Five dates. A supplier who cannot produce five dates has never performed stages 6 and 7, whatever the procedure says, and a supplier who produces five dates with a long gap between the third and the fourth is disclosing that their record does not carry the counterfactual column.

An analyst reading a firm from the outside has a harder job, because none of this is published. The shape of what does get disclosed is observable from outside. A firm that reports operational incidents in minutes of unavailability is measuring the first four stages. A firm that reports the number of customers affected and the date they were made whole is measuring all eight. The measure a firm chooses to publish reveals which half of the procedure it actually operates, and it costs nothing to notice.

O'Neil, in Weapons of Math Destruction, 2016, shows that a model's errors do not fall evenly across a population. Uneven error is exactly the shape here, and it matters for who does the asking. The changed format arrived on one channel, so the entire affected population came in through that one channel, and every one of the 176 files sat there. A firm looking at an approval rate for the whole month would have seen a movement of 1.4 percentage points and shrugged. The people carrying the whole of the movement were a subset of applicants none of whom knew they had anything in common.

India

Where does this sit with an Indian supervisor?

Telling customers that their files are held, putting affected outcomes right, and reporting an episode of this kind inside a regulated lender all sit under published expectations rather than under a firm's own preference. For a bank or a lender in India those expectations are stated by the Reserve Bank of India at rbi.org.in, across its material on digital lending, outsourcing, data and consent, and the oversight of arrangements that decide customer outcomes. The standing discipline of model risk work, from which the idea of independent challenge originates, comes from supervisory material published by the Bank for International Settlements at bis.org, named here as an origin rather than as the position in India.

The eight stages, the four items against each stage, the six trail entries and every figure attached to them are that bank's own arrangement, not a requirement, a threshold, a reporting window or an effective date. The current position should be read at the source before any of it is relied on.

Rebalancing: When, Why and What It Costs — free micro-course from Fin Maverick

What does somebody writing their first playbook do this week?

Not draft the document. The document is the easy part and it will be better written after the four things below than before them, because each one changes what has to be in it.

The first step is to list the arrangements that decide something about a customer without a person deciding. At this bank the answer was seven of the nine components once the test was set by consequence rather than by technology. The second is to ask the stage 6 question out loud about the largest of them and write down the answer it gets. The third is to take whichever record backs that arrangement and check whether it carries what the previous version would have produced; if it does not, that is the first change to make, and it is a column, not a programme. The fourth is to multiply: a plausible affected count, times the desk's own handling time, divided by a working day, with the number written into the playbook so that nobody ever has to guess at it under pressure.

Then write the eight stages, with four items against each, and put a date in the calendar for a drill that starts at stage 6. A playbook is worth exactly as much as the stage in it that nobody has ever performed, and the whole craft is finding out which stage that is before an episode finds out first.

The last finding of this episode is the one most likely to be lost, and it belongs in front of whoever writes the playbook. Nothing about the model was ever wrong. Every fitted number sat where somebody put it, throughout. The data arriving changed, the record failed because it could not answer a question about something that did not happen, and two blank columns in a table cost one hundred and seventy six people eight working days of waiting. Two blank columns are where the craft actually lives, and that is a long way from where the attention usually goes.

How a fault gets detected in the first place, and which signals can speak how fast, is covered under model drift and monitoring. The three levels of a stop and its six numbered actions are covered under stopping an automated system. Independent challenge before approval is covered under model validation. Ordinary technology incident management, the kind written for a service that stops answering, is a separate subject.
Retrieval and Grounding for Finance teaches you to design a retrieval setup over a document set and to say what grounding does and does not prevent.

Sources

SourceDocumentSite
Reserve Bank of IndiaPublished expectations on a regulated lender covering digital lending, outsourcing, data and consent, telling customers about outcomes affecting them, and the oversight of arrangements that decide those outcomes. The expectations applying to a lender in India are stated in this materialrbi.org.in
Bank for International SettlementsInternational supervisory material from which the standing discipline of model risk work and independent challenge originates, named as an origin rather than as the position in Indiabis.org
Agrawal, Gans and GoldfarbPrediction Machines, 2018, on a learned component producing a prediction that a person still has to act onnamed in the text
O'NeilWeapons of Math Destruction, 2016, on a model's errors falling unevenly across a populationnamed in the text

Sumeru Bank Limited, Revathi Balan, Ismail Sheikh, Neelima Rao and Ashok Pillai are invented.
Educational material. Not advice on any investment, tax, budget or market position.

← Previous
Fin Maverick Micro CoursesExplore Micro Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsCareersShowdown
RESOURCES
All CoursesMicro CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.