Prompting vs Fine Tuning: Changing the Instruction or the Model
Prompting changes the instruction the component is given each time it runs. Fine tuning changes the component's fitted numbers, once, using worked examples. The first is configuration and the second is a new version of the component. Both change the shape of what comes out, and neither changes whether an assertion is supported by the material supplied at the time. The comparison that decides anything is cost, reversal and approval.
Almost every argument between these two is conducted as though they were competing ways of making an output better. Prompting and fine tuning are not competing on quality. Both are ways of changing the form of an output, and on any one task the two land close enough together that nobody would choose between them on quality alone. Three other things separate them: what each one costs before it has improved anything, how quickly it can be put back, and whose signature it needs. Cost, reversal and approval differ by roughly an order of magnitude while the quality of the words barely differs at all. A comparison built on quality ages within a year. A comparison built on cost and reversal does not.
Everything below is drawn from one invented bank, Sumeru Bank Limited, and one deployed thing: component 8 of its retail personal loan intake chain, the drafting assistant that writes the first draft of an exception note and of the explanation paragraph inside a decline letter. A person signs both. The assistant runs on 3,010 exception notes a month.
What does each of the two actually change?
The same shape appears in a kitchen, and it is a familiar one. A tiffin service runs out of a two room flat and delivers about sixty boxes a day. There are two ways to change what comes out of that kitchen. A fresh note can be pinned above the stove every morning: less oil today, no chilli in the boxes going to the school, packing closes at eleven. Or the cook can be put through a month of training so that the way she cooks changes at the source, note or no note.
Both are real changes and both change what arrives in the boxes. The note and the training are still not the same kind of thing. The note is re-written tonight for tomorrow, and a fortnight later there is no trace it ever existed. The training cannot be un-done in an evening, it applies to everything the cook makes whether that was intended or not, and getting the old cooking back is a matter of tasting rather than of reading.
PromptingChanging the instruction a component is given each time it runs, rather than changing the component. is the note above the stove. What changes is the instructionThe text supplied with every request, telling the component what it is being asked to do and how the answer should look. that travels with each request, and the component itself is untouched. At Sumeru Bank Limited that instruction is 340 words in six numbered parts, and part 6 was added in month 9. Part 6 tells the component what to do when the material it has been given does not support an answer. The afternoon it was added, every note drafted after that used it. Nothing about the component had changed. Only the instruction had changed.
Fine tuningAdjusting a component's own fitted numbers using worked examples, so its behaviour changes at the source rather than in what it is told. is the month of training. Worked examples of the wanted output are used to adjust the component's own fitted numbers. The component that comes back afterwards is not the old one doing as it is told. Fine tuning produces a different component. How that adjustment is done is covered separately; what matters for the comparison is only that a new thing exists at the end of it and did not exist at the start.
The distinction between changing an instruction and changing a component sounds like a technicality and decides everything that follows. One of them is configurationSomething changed without producing a new version of the thing being changed, so there is nothing new to describe, record or approve afterwards.: something changed with nothing new produced. The other is a new versionA changed component. Because it is a different thing from the one before it, it usually has to be recorded and approved as such. of a component. Every practical difference between the two, the cost, the reversal, the approvals and the record, falls out of that one line and not out of anything either of them does to the text.
Which of the two produces a new version of the component?
What does each one cost before anything improves?
The phrasing is doing work. Not what does it cost, but what does it cost before anything improves. The cost before anything improves is the amount handed over whether or not the change turns out to help, and that cost decides whether an option is ever tried rather than merely discussed. A change costing an afternoon gets tested by somebody on a Tuesday. A change costing a fortnight gets a meeting, then a paper, then another meeting.
At Sumeru Bank Limited the two costs were not close. Adding part 6 to the instruction was an afternoon's work: draft the wording, read it against a handful of notes, put it in. Call it half a working day. The fine tuning proposal asked for 4,000 worked examples, and assembling them was estimated at 11 working days. Sumeru Bank Limited's working day is 7 hours, so 11 working days is 77 hours of somebody's time spent before a single note had been drafted any differently. Half a working day against eleven is twenty-two times, and it is paid up front, in advance of knowing whether the thing works at all.
The bank did not go straight to 4,000. The bank ran a trial on 300 worked examples first. At the same rate of assembly that is 0.8 of a working day, and the trial is where the measured figure below comes from. Running the trial was the sensible thing to do, and what it costs is worth noticing: finding out whether the eleven day option is worth eleven days takes most of a day. Finding out whether the instruction change is worth an afternoon takes the afternoon, and at the end of it nothing is being estimated.
Adding part 6 took about half a working day, and the full set of 4,000 worked examples was estimated at 11 working days. Which comparison does that gap settle first?
How is each one undone, and how fast?
ReversibilityHow quickly and how completely a change can be put back, including whether any trace of it survives afterwards. is the criterion teams discover last and should look at first. Undoing part 6 of the instruction is one action. The paragraph is deleted, and the next request that goes out carries the shorter instruction. The deletion takes minutes, nothing is left behind, and no version of anything sits around afterwards waiting to be described.
Undoing a fine tuning is not one action and it is not one afternoon. The previous version of the component has to be restored. Restoring it is easy enough if somebody kept that version and a serious problem if nobody did. The restoration is itself recorded. A version changed twice is two changes rather than none. Then the component has to be measured again. Everything known about how the component behaves was measured on the version just taken out. At this bank a measurement round is 200 notes checked by a person against the file that produced them. At the 9 minutes verification took at the time, that is 1,800 minutes, being 30 hours or about 4.3 working days. The 4.3 days are arithmetic on the bank's own numbers rather than a separately measured cost. The 4.3 days are also the honest price of putting a fine tuning back. Restoring the old version is quick. The price sits in re-establishing what was already thought to be known.
Undoing part 6 takes minutes. What does undoing a fine tuning take at this bank, and why?
Why does reversal speed change how a team works, rather than just how it recovers?
A wall can be repainted on a Sunday and repainted back the following Sunday, so nobody holds a meeting about the colour. A wall cannot be un-knocked, so knocking one down produces three conversations, a drawing, and somebody's brother-in-law who knows about these things. The two decisions are not different in importance. The two decisions differ in how expensive it is to be wrong, and that single difference decides whether the household argues or simply tries it.
Teams optimise for the cost of being wrong without ever noticing they are doing it. A change that takes days to undo generates weeks of meetings, and a change that takes minutes to undo generates evidence. Where reversal is cheap, trying the thing is the cheapest way to find out, so people try it and come back with a reading. Where reversal is expensive, arguing is cheaper than testing, so people argue, and the argument runs on opinion because nobody can afford to produce a fact. Arguing rather than testing is not a comment on anybody's character. Arguing rather than testing is arithmetic people do in their heads without writing it down.
Why does reversal speed change how a team works, rather than just how it recovers?
Who is able to make each change, and what approval does it need?
Approval is where the difference stops being about engineering and starts being about a bank. At Sumeru Bank Limited the 340 word instruction is held as configuration. One person can change it. Because nothing new is produced, there is nothing new to describe, and unless somebody separately writes down what they did, the change leaves no record of its own. The fine tuned version is the other case entirely. Being a new version of a recorded component, it has to be recorded as a version and approved before it runs, and the name against the approval belongs to a person rather than to the team.
The intake chain already carried one instance of exactly this asymmetry. The asymmetry is worth taking seriously rather than treating as paperwork. The waiting time in the workflow router, before a file escalates, was changed in month 7 with no approval and no record. Nobody hid it. The waiting time was configuration, so nothing had to be written down, and the change surfaced only in the month 10 sweep that Ashok Pillai ran across the bank. The cheaper and more reversible of any two changes is very often also the one that leaves no trace, and that is a property of the change, not a failing of the person who made it.
There is a second turn here that is easy to miss. The bank's own scoping test asks whether an output reaches a customer or a reported figure without a person deciding, and on that test component 8 sits outside the scope, on the ground that a person signs every note and every decline letter. So the approval standing over a new version of the drafting assistant is the bank's internal change arrangement rather than that scope. Both routes into component 8 therefore run outside the scoped list: the fine tuning through an internal approval, and the instruction through nobody at all.
Where the record and the approval come from
A change to a component that a regulated lender is running sits inside that lender's own obligations on outsourcing, digital lending, customer data and record-keeping, and those expectations are set by the Reserve Bank of India and published at rbi.org.in. Expectations of that kind carry thresholds and effective dates that change from time to time, so the position that applies to any particular lender is whatever the Reserve Bank of India currently publishes. The split described above, one person for the instruction and a recorded approval for a new version, is Sumeru Bank Limited's own arrangement.
What do both of them leave completely untouched?
Back to the tiffin kitchen for the sentence that matters most. Suppose nobody in the kitchen has been told how many boxes are due today, so the boxes keep going out with the wrong count. A better note above the stove does not fix that. A month of training does not fix that either. Both of them change how the cook works. Neither of them tells her the number, and the number is what is missing.
A missing count is the shape of a fault of supportAn output that asserts something the material it was given does not contain, so the problem is what was supplied rather than how it was written., and it is why this comparison is so often held over the wrong thing. Neither prompting nor fine tuning changes what material the component is given. Neither changes whether that material contains the answer to the question being asked. And neither changes who carries the residue afterwards. A person still reads the draft and a person still signs it. The material, its sufficiency and the residue are properties of the arrangement around the component rather than of the component, and both options on the table are changes to the component's behaviour.
There is one apparent counter-example and it repays a careful look. Getting it wrong wastes a lot of everybody's time. Part 6 of the instruction says what to do when the material does not support an answer, so it plainly does touch the third item. Part 6 changes the component's conduct when the material falls short. Part 6 does not change whether the material falls short, and the gap between those two sentences is where the whole error lives. Part 6 helped, and it helped by making a shortfall visible rather than by removing it: the rate at which the component declined to answer went from 0 to 12 of 200, being 6.0 per cent. The 12 declined notes are not free. Each one goes back to a person to write from nothing.
Name one thing neither prompting nor fine tuning changes.
One bank compared these two on this task. Which won, and by how much?
What did the measured comparison at one bank actually favour?
Here are the three readings, and they are all this invented bank's own. The unsupported statement rate on component 8 was measured on 200 exception notes in month 6 and stood at 23 of 200, being 11.5 per cent. The fine tuning trial on 300 worked examples moved it to 9.5 per cent, an improvement of 2.0 percentage points. For the full 4,000 examples, never run, the bank projected 3.5 percentage points against those 11 working days of assembly. And over the same period, the three grounding changes made between month 8 and month 9 moved the rate from 11.5 per cent to 3.5, being 8.0 percentage points.
| What was changed | Cost before anything improves | Improvement | Measured or projected |
|---|---|---|---|
| The trial on 300 worked examples | 0.8 of a working day | 2.0 points | Measured |
| The full set of 4,000 worked examples | 11 working days | 3.5 points | The bank's own projection |
| The three grounding changes together | Spread over two months | 8.0 points | Measured |
Thirteen times the examples were projected to buy less than twice the improvement, and even that projected figure is less than half of what the option nobody was arguing about actually delivered. Read the middle row twice. The middle row decided the matter and the middle row was never measured. Deciding against an unmeasured projection is the ordinary situation. Measuring an option costs exactly what the option costs.
Why did the instruction win on this task?
Not because it wrote better sentences. The instruction won on three things and none of the three is about the text. The instruction cost half a working day against eleven, so it could be tried before it was believed. Taking it out again took minutes, so trying it risked almost nothing. And it produced no new version, so nobody had to wait for an approval to find out whether it helped. Put those together and the instruction change had already been measured by the time the other option had finished being described.
Now the honest part, and this bank was careful about it. Three changes were made between the two measurements and they cannot be pulled fully apart. In month 8 a retrieval arrangement was added, so different material reached the component. In month 9 part 6 was added to the instruction, and a connection was added letting the component read the file's structured fields instead of restating figures out of prose. In month 10 the second measurement was taken, and the rate stood at 7 of 200, being 3.5 per cent. Anybody crediting the whole 8.0 point improvement to one of those three changes is guessing, and this bank never split it.
There is exactly one clean attribution inside it, and it is worth having because it shows what a clean attribution looks like. Of the 23 faulty drafts in the first measurement, 6 were restated figures and 17 were other statements. Thirty-one of the 200 notes restated a figure at all, so 6 of 31 were wrong, being 19.4 per cent. In the second measurement 34 notes restated a figure and none of them was wrong. No other change touched figures, so the connection to the structured fields is the cause of that part, and it is neither prompting nor fine tuning. The cause is a connection to a system that already held the number.
Restated figures went from 6 wrong out of 31 notes to 0 out of 34. What is that clean attribution evidence for?
When would the answer be the other way round?
A verdict about this pair is worth about eighteen months. The test that produced the verdict travels a great deal further. Fine tuning is the better answer where three conditions hold together, and the interesting thing about this bank is not that it chose the instruction but that all three conditions were readable in advance and none of them held.
The first condition is that the fault is the form of the output rather than its support. Examples are extremely good at teaching a shape: this is how a note opens, this is the order the paragraphs go in, this is the register a note is written in. Examples cannot demonstrate the absence of a claim about a file that did not exist when the example was written. At this bank the fault was 23 notes in 200 asserting something the file did not contain. A note asserting something the file does not hold is a fault of support, not of form. The second condition is that the worked examples already exist as a by-product of the work. Assembling them is then collecting rather than writing. Here they did not: 4,000 had to be made, at 11 working days.
The third condition is the one people forget. An instruction is supplied with every single request, and so is the material. Where the instruction has to grow long enough to crowd the material out, moving that behaviour into the component buys back room. At Sumeru Bank Limited the instruction is 340 words against material averaging 4,800 words a note. The instruction is 7.1 per cent of the material's length, nowhere near crowding anything out. The bank was already losing material at the far end of the range, where the longest five per cent of files exceed 12,000 words and the oldest part was dropped silently, and shortening a 340 word instruction would not have saved a single one of the 9 notes in 200 that were drafted from a truncated file.
When would fine tuning be the better choice on this kind of task?
Why was the three week argument the wrong argument?
Sumeru Bank Limited spent three weeks choosing between these two, and the argument was conducted properly. Both sides prepared. The trial was run rather than asserted. The decision that came out of it was the right one. And the argument was still the wrong argument. The fault lay in what it was being held about.
The error that gets made, and what it costs
Both options were put on the table as the way to stop the component asserting things the file did not contain. Asserting what the file did not contain is a fault of support. Neither option touches support. The instruction can tell the component what to do when the material does not carry an answer, and part 6 helped for exactly that reason. The instruction cannot make the material carry one. The worked examples can show what a good note looks like, and they cannot demonstrate the absence of a claim about a file that did not exist when the example was written.
Both boxes on the table were aimed at the form of the output. The fault sat in what was being supplied. Three weeks went into choosing between them. The 8.0 points came from better material plus an instruction covering the case where the material falls short, and neither option as it was framed was that.
The cost is not the three weeks. Three weeks are recoverable. The cost is that a decision reached for the wrong reason leaves nothing reusable: the bank was right to decline the fine tuning, it declined it on assembly cost, and it would have declined a cheap fine tuning too, for a reason nobody in the room had written down. Nobody in that argument behaved badly and nobody was careless. Everybody in the room was answering a well-posed question that was not the question in front of them.
Both options were proposed to stop the component asserting things the file did not contain. What was wrong with that argument?
Why is there no control to move here?
How this gets used on a change somebody proposes tomorrow
Whoever runs a deployment asks the four questions in order and refuses to discuss quality until they are answered. What does the change actually change. What does it cost before it has improved anything. How is it undone, and how long does that take. Who has to sign it. Four answers take twenty minutes to collect and settle most of these arguments before anybody prepares a slide. In the ordinary case one option is an afternoon and reversible, and the other is a fortnight and is not.
Whoever reviews change records reads this backwards, and it is the more useful direction. The cheap reversible change is the one most likely to be missing from the record, precisely because nothing was produced that had to be written down. So a review that reads only the version history of a component has read the expensive changes and none of the cheap ones, and at this invented bank the change that went unrecorded for three months was configuration on a written rule set, not a new version of anything.
And anybody sitting in a meeting that has run past its second week can ask the one question that ends it. What fault is being fixed, and does either option on the table touch it. If the answer to the second half is no, the meeting is not close to a decision, it is close to the wrong decision, and the cost of asking is one sentence.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering outsourcing, digital lending, customer data, consent and record-keeping | rbi.org.in |
| Securities and Exchange Board of India | Expectations where the deployer of such a component is a market intermediary | sebi.gov.in |
| Bank for International Settlements | International supervisory material on the deployment of such systems by banks | bis.org |
Sumeru Bank Limited, its intake chain, component 8 and Ashok Pillai are invented.
Educational material. Not advice on any investment, tax, budget or market position.
