From Use Case to Production: Defining It, Piloting It, Running It
A use case is one decision somebody currently makes, written down before anything is built. The use case travels six stages from that sentence to a live deployment, and the fourth is a shadow run in which the component decides and nothing acts on the decision. Between the fourth stage and the fifth, nothing about the component changes and everything about the consequence does.
Everything else here rests on one sentence that sounds obvious and is routinely ignored. A pilot measures the deployment it was allowed to be. Anything its shape excludes, whether that is a consequence nobody acts on or a population it never saw, produces no number at all rather than a wrong number. And a missing number is invisible in a business case in a way a wrong number never is. A wrong number gets argued about, and an empty cell gets skipped over.
What is a use case, written so that somebody can actually build it?
A use caseOne decision somebody currently makes, named and written down before anything is built. is one decision somebody currently makes. The sentence is the whole definition, and every word in it is load bearing. One, so it can be scoped. Decision, so there is a right answer to compare against. Somebody currently makes, so the thing already happens today and it can be watched happening before anything is spent.
The shape is easier to see when there is no technology in the room, so start with a vegetable stall. The person running the stall makes a decision every morning: how much coriander to buy at the wholesale market. The coriander decision is a use case. The stallholder makes one decision, it happens today, and after three days each morning's answer can be judged right or wrong. The unsold bunches are sitting there at closing time saying so. Now compare it with the sentence a consultant would write about the same stall: modernise procurement. Modernising procurement is not a decision. Nobody makes it on a Tuesday, nobody can be wrong about it on a Wednesday, and there is nothing at closing time that says whether it happened.
The test of a use case is whether somebody could have been wrong about it yesterday, and whether that could be checked. Sumeru Bank Limited is an invented bank running a retail loan intake chain that begins on an applicant's handset and ends in a decision. Written properly, one of its use cases fits on a single line. Written the way most of them arrive, it fits on a single line too, and that is exactly the problem.
Which of these is a use case: automate underwriting, or decide whether a declared income is corroborated?
What is a use case not, and why does that difference bite so hard?
Three things get written on proposal forms in place of a decision, and each one fails in its own way. A department is not a use case. A floor of a building does not make decisions, the people on it do, and they make dozens of different ones. A method is not a use case. A method answers how before anybody has said what, so the choice of component has already been made by somebody who never wrote down the problem. An ambition is not a use case. No moment ever arrives at which an ambition is right or wrong.
The reason this matters is not tidiness, it is that every later stage of delivery needs a right answer to compare against, and a sentence that names no decision never produces one. There is no checking whether the data exists for improve the customer journey. There is no shadow test for automate underwriting. There is nothing to shadow. And when the thing goes live, nobody can say whether it worked, so what gets reported instead is activity: how many files it touched, how fast it ran, how many people used it. None of those is an outcome, and anybody presenting one of those numbers can be asked what the decision was and how often it was right.
What turns a list of candidates into a strategy?
Digital Strategy, and the two things a list is missing
Most of what gets called a strategy in this area is a list. Twelve things a firm would like to do, each with a paragraph, usually arranged in the order somebody thought of them. A list is not useless, and it is the raw material for everything that follows. But it is missing the two properties that would make it a strategy, and both of them are unglamorous.
The first is a shared scoring. If twelve candidates have been described by twelve different people in twelve different ways, they cannot be compared at all, and the one that gets built is the one whose sponsor argues best. The second is an order, with a reason attached to it, and a statement of what would make the firm stop. A strategy is a scored, ordered set of candidates with a stated reason for the order, and anything without those two properties is a list with a cover sheet.
The household version is a home that needs work. Fixing the leaking tap, painting the second bedroom and replacing the wiring are three candidates. The list says all three. The strategy says wiring first because a fault there burns the place down, tap second because it is cheap and it is wasting water every day, paint last because nothing bad happens if it waits a year. Notice what the strategy added: not more items, but an order and the reasoning behind it. The reasoning is the part somebody can disagree with, and that is exactly what makes it worth writing.
What turns a list of candidates into a strategy?
What are the stages between a stated problem and a live deployment?
AI Delivery Stages, the six of them, in order
Six of them, and this bank numbered its own so that anybody in a meeting could say which one a piece of work had reached without a paragraph of explanation. The stages are the bank's own design rather than anybody's published standard, and the value of numbering them is entirely practical: a question like which stage is this at has a one word answer, and a question like how is it going does not.
Stage 1 is the problem stated as a decision somebody currently makes. Stage 2 is the data check: does the input exist in a readable form, and was the outcome ever observed. Stage 3 is the build. Stage 4 is the shadow run, where the component decides and nothing acts on the decision. Stage 5 is the live run on one channel. Stage 6 is the full run with monitoring. There are five boundaries between those six stages, four of them are engineering boundaries, and exactly one changes who bears the cost of a mistake.
Each of those six is a delivery stageOne of six ordered steps between a stated decision and a monitored live run., and the honest thing to say about the calendar underneath is that the four months before go-live are where all the work sits and none of the risk does. Months 1 to 3 carried the shadow run at about 400 files a month. Month 4 put the chain live on one channel. Month 5 extended it to every channel and month 6 is the first steady month. Every volume quoted below comes from that steady month.
Why does the data check ask whether the outcome was ever observed?
Because it decides whether the thing is buildable at all, and no amount of care later can rescue a no. There are two halves to stage 2 and people usually remember only the first. The first half asks whether the input exists in a readable form: are the bank statements there, can the fields be pulled off them, is the history long enough. The first half gets attention because it is a plumbing question and plumbing questions have owners.
The second half is the one that kills projects, and it asks whether anybody ever observed the outcome. Not whether the outcome happened. Whether it was recorded in a form somebody can go back and read. Think about a shopkeeper who wants to know which customers would have paid on time if he had extended credit. He has years of records about the customers he did extend credit to. Refusing somebody generates no history, so he has nothing at all about the ones he refused. The outcome for that second group was never observed, and no cleverness recovers it.
Without an observed outcome there is nothing to fit against and nothing to measure against, so an input that exists in a readable form is not enough on its own. An answer of no is why the question sits at stage 2 rather than at stage 5. At stage 2 the answer costs a fortnight of somebody's time. At stage 5 it costs the build, and the answer has not changed.
Stage 2 asks whether the outcome was ever observed. Why does that question sit before the build rather than after it?
What is a shadow run, and what does it actually prove?
Pilot, and the safest arrangement anybody has found
A shadow runA stage where the component decides and nothing acts on the decision. is the fourth stage, and it is the cleverest thing in the sequence. The component runs on real files, in real time, and produces a real decision for each one. The component's decision is written to a record. Then nothing happens. The file continues down the existing manual process exactly as it would have done if the component did not exist, a person decides it, and the two answers sit side by side afterwards.
At Sumeru the shadow run covered months 1 to 3 on one channel, at about 400 files a month, for 1,200 shadow files in total. Of those 1,200, the chain would have sent 852 straight through with no person touching them, or 71.0 per cent. The 71.0 per cent is the single most quoted number in this entire case, and it should never have been quoted on its own.
A shadow run proves something genuinely valuable and genuinely narrow. The component decides well on the files it saw. Agrawal, Gans and Goldfarb make the point in Prediction Machines that a learned component produces a prediction while a person still has to decide what to do about it, and a shadow run is that separation built deliberately into an operating arrangement: the prediction runs at full volume and the action stays with the existing process. Nothing at any point acts on a single one of a shadow run's decisions, so a shadow run tests the prediction completely and tests the consequence not at all. The complete absence of consequence is what makes a shadow run safe, and it is also the precise reason two of this bank's four pilot measures do not exist.
What changes when a pilot becomes a live deployment?
Production Deployment, and the one thing that moves
A production deploymentThe live run where the output reaches a customer or a record without anything else in the way. is the run where the output reaches a customer or a record with nothing else in between. Ask what changed at that moment and the honest answer is almost nothing. The smallness of the change is why the moment is so easy to underestimate.
Pilot vs Production Deployment, held side by side
The component is the same component. The same fitted numbers, the same written lines, the same inputs arriving in the same format. The chain produces the same decision on the same file that it would have produced the week before. The change is in what happens next: on Friday the decision became a row in a log that a person read later, and on Monday it became a message on an applicant's handset saying the loan is refused.
The distinction between a pilot and a live deployment is entirely a distinction about what happens to the output, and never about the component itself. The distinction inverts the usual instinct, and it is worth sitting with. People prepare for go-live by testing the component harder. The component is not what changed. The change is that a category of person who did not exist last week now exists this week: the person who is worse off because the component got one wrong.
A shadow run produces excellent decisions for three months. What has it proved about the live deployment?
How are candidates chosen when only a few can be built?
How to Build a Digital-Finance Use-Case Portfolio
A portfolioThe set of candidates a firm is choosing between, scored on the same axes. here means nothing more than the set of candidates a firm is choosing between, scored on the same axes so that they can be compared. Sumeru held 12 of them and built 4 in the period. All 12 are candidates that were considered, not systems in use, and the distinction matters because a count of things a firm thought about and a count of things a firm is running are two entirely different numbers that get confused constantly.
The bank scored each candidate on three numbered axes, one to three each, so three to nine in total. The axes are the bank's own invention rather than anybody's published method, and the interesting one is the third, the one people skip.
| Axis | The question it asks | Score for the intake chain |
|---|---|---|
| 1 Value if it works | How much is at stake if this works as intended, in money, in time or in applicants served | 3 |
| 2 Confidence it will work | How sure is the firm that a component can do this at all, given the data and the decision | 3 |
| 3 Containment | How well is the cost of being wrong held, if it is wrong | 2 |
| Total | Quoted afterwards as a single number by almost everybody | 8 |
ContainmentHow well the cost of being wrong is limited, which is the axis usually scored last and read least. is scored last, read least, and it is the only one of the three that describes what happens on the bad days. Notice also what the total hides. Eight out of nine sounds like a strong candidate and it says nothing whatever about which axis is low. Two of the three axes describe the upside and one describes who gets hurt, so a total treats them as interchangeable when they are not.
A candidate scores 3 on value, 3 on confidence and 2 on containment. What does the total of 8 hide?
What did the containment score of two say in advance?
It said, in the bank's own handwriting, before a line of the build existed, that the cost of being wrong on this candidate was not well held. Somebody sat down at month 0, looked at a chain that would decide retail loan applications with nobody in the room, and wrote 2. Then the total was calculated, the number 8 went into the paper, and the row was not read again.
Here is what turned up later, and every figure is the invented bank's own month 6 reading. Handling time on an exception ran at 19 minutes against a business case of 11. The chain saved 5 posts on the exception desk where the business case had promised 10. And 14.0 per cent of the applicants who started an application never reached the decision engine at all, being 1,400 people in the month. Every one of those three is a cost of being wrong, and the cost of being wrong is precisely what the row scored 2 was about.
None of that is negligence. The people who scored the sheet did the harder thing and wrote a low number next to a project everybody wanted. The failure is smaller and much more common: a scoring sheet is treated as an entry gate rather than as a document to go back to. Nobody diarised a date to re-read row 3 after go-live, so the one piece of paper that had predicted the problem sat in a shared folder while the problem happened.
What did one bank's own portfolio rule actually stop?
Sumeru wrote itself a single sentence of a rule, and it is worth reading twice because it only bites when two things are true at once. No more than one use case at a time may sit where the correct outcome cannot be stated in advance and the output reaches a customer with no person in between.
Two conditions, joined by and. Plenty of components fail the first: anything whose right answer can only be found out later, by waiting. Plenty fail the second: anything whose output goes straight to a person outside the bank. Almost nothing fails both. Of the chain's components, exactly one comes out.
The scoring model is component 6 of nine in this chain. Its correct outcome is whether the applicant repays, and nobody can state that in advance because it takes months to find out. Its output reaches the applicant directly as an acceptance or a refusal. The classifier fails the second condition because its output never leaves the system. The drafting assistant fails it too. A person signs every output. A rule joined by an and is only as good as how rarely the pair occurs, and the whole craft of writing one is choosing conditions that catch the thing being guarded against and almost nothing else.
The rule allows one use case at a time where the correct outcome cannot be stated in advance and the output reaches a customer with nobody in between. Which component was that one?
How is a pilot measured without flattering it?
How to Measure an AI Pilot Responsibly
Start by writing down the measures before the pilot runs, then ask one question of each: could a pilot of this shape produce this number at all. Not will it be a good number. Could it exist. A pilot measureA number the pilot produced, as distinct from a number the pilot's shape made impossible. that the shape makes impossible is not a gap to fill in later. An impossible measure is a finding about the pilot, and it belongs in the report on the day the pilot is designed rather than on the day the business case is signed.
Two questions decide it, and both are about the shape rather than about the component. Does anything act on the output? If nothing does, every measure that depends on a consequence will be missing. Does the pilot see the whole population the live run will see? If it does not, every measure that depends on the population will be missing. A no to either question names a measure that will be absent rather than wrong, and absent is the harder one to notice.
The bank set four measures for this pilot. Two of them came back with numbers, one of which moved and one of which did not, and two came back empty. Look at the empty column rather than the filled one.
The straight through rate fell 6.0 percentage points between the pilot and month 6, and the bank never split that fall into causes. The pilot ran on one channel and month 5 extended the chain to all of them. The extension is a change in the population and not a change in the component.
A pilot report has four measures, two of which are blank. What does that call for?
What could this pilot never have measured at all?
Measures 3 and 4 were not measured badly. Both were unmeasurable by shapeA measure the pilot could never have produced, whatever care was taken, because of how the pilot was set up.. Being unmeasurable is a different thing from being measured badly, and a much more useful thing to be able to say out loud.
Handling time on an exception could not be measured because nothing acted on a shadow decision. A shadow run creates no work for anybody. No file was ever routed to the exception desk because of it, so no exception existed, so there was nothing to put a clock on. The measure was not skipped. The measure had no possible value.
The share of applicants who never reached the decision engine could not be measured because of who the pilot's files were. Every file it saw had already been onboarded through the existing manual process before the component ever looked at it. By construction, every single one had survived the front door. Asking that population how many people fall at the front door is like standing inside a hall and counting how many guests were turned away at the gate.
The two absent measures are exactly the two that produced the business case miss, and that is not a coincidence: a pilot's shape hides precisely the costs that shape protects it from. A shadow run protects against consequence, so it hides the cost of consequence. A pilot on pre-onboarded files protects against the front door, so it hides the cost of the front door. Read backwards, that sentence becomes a design tool: naming what the pilot protects against names its blind spots before it starts.
Why could the pilot not measure handling time on an exception?
Before the control below is moved: the pilot reported 71.0 per cent of files running straight through, and it only ever saw applicants who had already been onboarded by hand. What share of all applicants got an answer with no person touching the file once the chain went live?
The measure the pilot could not produce
One control moves: the share of applicants lost before they ever reach the decision engine. The straight through rate is held at one of the two values this bank actually measured, so only one thing changes at a time. The bar on top is the pilot corner, the reading a pilot that never met the front door reports.
Lost before the decision engine: 14.0 per cent of the 10,000 who started
Educational illustration. Figures are the invented bank's own and describe one deployment. Held constant: 10,000 applications started in the month, and the straight through rate held at one of the two values the bank measured rather than modelled. The default reproduces month 6 exactly: a 14.0 per cent front door loss and a 65.0 per cent straight through rate give 55.9 per cent of applicants, being 5,590 of the 10,000 started, against a pilot corner of no front door loss and 71.0 per cent, being 7,100. The pilot never saw an applicant who failed onboarding, so its own shape fixed the front door loss at zero.
How a lender, an analyst or an investor actually uses this
A lender's operations head reads a pilot report by turning to the measures table first and reading the empty cells before the full ones. For each blank, one question: was this not taken, or could it not have been taken. If it could not have been taken, that is a finding about the pilot's design, it goes in the report in those words, and the business case is written knowing that two of its inputs are guesses rather than readings.
An analyst reads a deployment's business case the same way, and this one is worth reading in full because it did not pay back where it said it would. Sumeru spent an invented Rs 2,40,00,000/- to build the chain once and Rs 65,00,000/- a year to run it. Five posts were saved on the exception desk, and at an assumed fully loaded Rs 9,00,000/- a post a year that is Rs 45,00,000/- a year, so the running cost exceeds the headcount saving by Rs 20,00,000/- a year. The chain does not pay back on headcount at all, and it pays back on the 65.0 per cent of applicants who now get an answer in about 4 minutes instead of 2 working days. An analyst who reads only the headcount line concludes the wrong thing about a deployment that is working.
And an investor or a lending committee reading any proposal of this kind has one cheap question available before any of the technical detail: what population will the pilot see, and how does it differ from the population the live run will see. The answer takes a minute and it predicts which numbers the report will be missing six months later.
India, in one paragraph
Where a regulated lender puts a component live so that its output reaches a customer, the expectations on outsourcing, digital lending, records and customer data sit with the Reserve Bank of India at rbi.org.in, and where the deployer is a market intermediary the equivalent expectations sit with the Securities and Exchange Board of India at sebi.gov.in. Instructions arising from a live retail deployment travel over rails whose operating material is published by the National Payments Corporation of India at npci.org.in. Requirements, thresholds and effective dates change and must be read at the source. The six stages, the three scoring axes and the portfolio rule described here are one invented bank's own design and are nobody's standard.
The failure: an empty cell reads as a zero and nobody notices
Measures 3 and 4 were not measured badly. Neither could be measured at all by a pilot of that shape, and nobody wrote that sentence down anywhere. Handling time on an exception was unmeasurable because nothing acted on a shadow decision, so no exception ever reached the desk. The share of applicants who never reached the decision engine was unmeasurable because the pilot took files that had already been onboarded through the existing manual process, so every file it saw had by definition survived the front door.
Here is the cost. The two absent measures are exactly the two that produced the business case miss: handling time rose to 19 minutes against the 11 the case assumed, and 14.0 per cent of applicants never arrived at all. The case had promised 10 posts and 5 were saved. Decompose the shortfall on the bank's own figures and it splits in two: holding handling time at 11 minutes and taking the actual 3,010 exceptions costs 18,920 desk minutes more than the case allowed, and taking those same 3,010 at the actual 19 minutes costs a further 24,080. The two together are 43,000 minutes, being 5.12 posts at an assumed working month of 8,400 minutes a person, and 56.0 per cent of it comes from the handling time that the pilot's shape made unmeasurable.
A wrong number in a business case gets challenged in the meeting, and an empty cell gets read as though the risk it describes were zero. The containment score of 2 had said in advance that the cost of being wrong was not well held here, and after go-live nobody went back and read it. Two people could have caught this and neither had a reason to look: Revathi Balan, who is the named accountable person for the scoring model, was accountable for a component rather than for a pilot's shape, and Ismail Sheikh, who runs the exception desk, was never in the pilot at all because the pilot sent him nothing.
Covered elsewhere. How a component is fitted, validated or evaluated is covered separately. The work that reaches the exception desk once a deployment is live is covered separately. Who approves a use case, who is accountable for it and what a register of them holds is covered separately.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering outsourcing, digital lending, records and customer data where a deployed component reaches a customer | rbi.org.in |
| Securities and Exchange Board of India | Equivalent expectations where the institution deploying such a component is a market intermediary | sebi.gov.in |
| National Payments Corporation of India | Operating material on the rails that instructions from a live retail deployment travel over | npci.org.in |
| Agrawal, Gans and Goldfarb | Prediction Machines, on a learned component producing a prediction that a person still has to act on | Harvard Business Review Press |
Sumeru Bank Limited, its retail loan intake chain, Revathi Balan and Ismail Sheikh are invented.
Educational material. Not advice on any investment, tax, budget or market position.
