Fine Tuning: Adapting a Model, and When It Is Worth It
Fine tuning adjusts a component's fitted numbers using worked examples of the task to be done. Fine tuning reliably changes the shape of the output: the format, the section order, the register, the required statements. Support is a property of the material in front of the component rather than of anything an example can demonstrate, and fine tuning does not reliably change whether a sentence it writes is carried by that material.
The distinction between shape and support decides almost every argument ever held about fine tuning, and the argument is almost never held in those terms. The argument is usually held as a question of degree. Somebody says the component makes things up, somebody else says fine tuning will fix that, and a third person asks how many examples it would take. All three are treating the problem as a dial that can be turned far enough. The problem is not a dial but a pair of different problems, and only one of them is the sort of thing an example can teach.
Sumeru Bank Limited, an invented lender, runs a retail personal loan intake chain of nine numbered components. Component 8 is a drafting assistant: it writes the first draft of the exception note a desk officer writes on a stopped file, and the explanation paragraph inside a decline letter. A person signs both. In month 6 it ran on 3,010 exception notes. Somebody proposed fine tuning it, and the proposal was declined. The interesting part is not the decision but the reason that was recorded, and the better reason that was available and not written down.
What does fine tuning actually change about a component?
Think about how a new officer on a desk learns to write a note. Nobody hands them a theory of note writing. Somebody hands them a folder of thirty notes that previous officers wrote and that the desk head was happy to sign, and says read these. After thirty notes the new officer writes notes that look like those notes. The paragraphs come in the same order. The tone is the same flat, unexcited register. Because all thirty carried it, the sentence the desk always puts at the end is there too.
Fine tuningAdjusting a component's fitted numbers using worked examples of a task, so that its later outputs resemble those examples. is that folder, applied to a component instead of a person. Examples of the task are assembled. The component's fitted numbersThe values inside a component that were derived from data rather than written by a person, and which decide what it produces. are adjusted so that what it produces afterwards resembles what those examples contain. How that adjustment is computed is covered separately; for the decision at hand, the arithmetic of it does not matter at all. Timing is what matters: the adjustment happens once, and it happens before the component ever sees the file it will be asked to write about.
Hold that thought against the new officer. Reading thirty notes teaches them the shape of a note. The file that lands on the desk next Tuesday did not exist while they were reading, and the folder teaches them nothing whatsoever about it. If next Tuesday's file has no salary credit in it, no amount of folder reading tells the officer that. The folder was closed before the file arrived.
What does fine tuning reliably change about a component?
What has to be assembled before any of it can happen?
A worked exampleA single case showing the input material and the output somebody with authority considered correct for it. is a pair. On one side, the material a person was given. On the other, the output that somebody with the authority to say so considered correct for that material. Not approximately correct, not a draft somebody later fixed. Correct, as signed.
The second half of that pair is where the work is, and it is worth being blunt about it. A system's own output is the thing being changed, so examples cannot be assembled out of it. An example needs output somebody stands behind. At this bank that means going to Ismail Sheikh, who runs the exception desk, and pulling notes his officers wrote and he was willing to sign, then pairing each one with the file it was written from. The examples are not found, they are made, and the making is done by the people whose time is the scarcest thing in the building.
A worked example holds a file and the note somebody signed as correct for it. Which of these can it not demonstrate?
What does assembling the examples actually cost?
The proposal that went to the bank's committee asked for 4,000 worked examples, and the estimate attached to assembling them was 11 working days of somebody's time. Eleven days is the whole of the assembly costThe working days spent producing worked examples, almost always the real cost of this kind of work. that mattered, and nothing technical appeared anywhere in the costing. Notice how ordinary that is. The cost is a person, a screen and a fortnight.
Converted into units that can be felt, using this bank's own assumptions and nothing else: a working day here is 7 hours, or 420 minutes. Eleven of them is 4,620 working minutes, or 77 hours. Spread across 4,000 examples that is about 1.2 minutes an example. A rate of 1.2 minutes says something the estimate does not say out loud: the 11 days assumes the examples are being selected and paired from work already done, not written from scratch, and if the notes have to be authored the estimate is not the right estimate. That is the first question to put to any proposal of this shape.
Now price it. The bank assumes a fully loaded cost of Rs 9,00,000/- a year for a desk post. Its month holds 20 working days, so its year holds 240. A working day is then Rs 3,750/-. Eleven working days is therefore Rs 41,250/-. Set that beside the Rs 2,40,00,000/- it cost to build the intake chain once and the Rs 65,00,000/- a year it costs to run, and the proposed fine tuning comes to 0.17 per cent of the build. Hold on to that figure. The 0.17 per cent becomes uncomfortable later.
What was the real cost of the fine tuning this bank was asked to approve?
Which faults do worked examples fix extremely well?
A fault of formA defect in the shape of an output: its format, section order, length, register, or a statement the process requires it to carry. is a defect in the shape of what comes out. The note is right but arrives with its sections in the wrong order. The register drifts into something apologetic when the desk writes flat. The closing line the desk requires on every note is missing. The paragraph runs to four hundred words where the desk allows one hundred and fifty. Nothing in any of that is about whether the note is true.
Examples are close to ideal for this, and the reason is almost embarrassingly simple. There is no way to write an example without demonstrating the correct form, so every single example demonstrates it. Three hundred examples is three hundred demonstrations of the section order, three hundred demonstrations of the register and three hundred demonstrations of the closing line, without anybody having had to describe any of them in words. Showing beats telling in exactly that situation, and the same reason explains why the folder of thirty notes works on a new officer.
There is a second, quieter reason form is the right target. Form is checkable by somebody who does not know the case. A person who has never seen the file can be handed a note and say whether the sections are in order and whether the closing line is present. The same person cannot say whether the third sentence is supported. Faults that can be checked without the case are the faults an assembled pile of examples can carry.
Which faults can no number of examples fix?
A fault of supportAn assertion in an output that the material supplied at the time does not contain. is different in kind, not in size. The note says the applicant has a salary credit of Rs 38,000/- in each of the last three months when the file shows only two. The note gives a reason for the stop that is a perfectly reasonable reason and is not the reason in the file. The sentence reads well, sits in the right section and is not in the material.
Building an example that teaches a component not to do that is the difficulty. Such an example would need a pair whose correct output demonstrates the absence of a particular claim. But the claim that would have been unsupported concerns a file that did not exist when the example was written, so there is nothing in the pair for that absence to be about. A thousand notes can go in the folder and every one of them will be a note that says true things about its own file. Absence is not a feature an example can carry, so none of them is a demonstration of not saying a false thing about a different file.
The household version is easy to feel. Every laying shows the fork on the left, so a child can be taught to lay a table by having it laid in front of them thirty times. Thirty conversations in which nobody repeated anything hold nothing for a child to copy, so no number of them teaches a child not to repeat something overheard. Support, like discretion, is a property of the situation in front of the writer and not of the examples behind.
Why can no worked example teach a component not to assert something the material does not support?
What did the trial at this bank actually deliver?
Before committing to 4,000 examples, the bank ran a trial on 300, and measured the result on the same 200 note test it had used before. The rate at which a draft contained a statement the file did not support moved from 11.5 per cent to 9.5 per cent, an improvement of 2.0 percentage points. The 2.0 point movement is the entire measured result, and every other figure attached to this proposal was projected rather than measured.
Turn the percentages back into notes. Percentages hide small numbers and this one is small. Out of 200 drafts, 11.5 per cent is 23 drafts carrying an unsupported statement, and 9.5 per cent is 19. Three hundred assembled examples moved four notes out of two hundred. Four notes is not a disappointment and not a scandal. Four notes is a result, and the useful question is not whether four is a lot but which four they were.
The trial moved the rate 2.0 points and the bank read most of that movement as form. Why does that matter?
How much of those two points was really about form?
The reading of those 2.0 points is worth slowing down on. A number stops being a number and starts being evidence at exactly this step. The measurement counted a draft as faulty if it contained a statement the file did not support. Some of the drafts it counted were faulty because the closing statement the desk requires on every note was absent, and the file itself required that statement. Others had their sections out of order in a way that put an assertion under the wrong heading. Both are faults of form wearing the coat of a fault of support, and both are precisely the sort of thing worked examples fix.
So the honest reading of the trial is not that fine tuning delivered two points against unsupported statements. The honest reading is that fine tuning delivered its two points mostly by fixing form. The mechanism says it should do exactly that, so the result generalises rather than being a fluke. A result that matches the mechanism can be extended, and extension makes it far more useful than a bigger result that does not. A result that contradicts the mechanism shows only that something in the measurement needs looking at.
The bank did not split the four notes further, so the split between form and support is not known. A split that nobody counted is not a finding. The direction is what was recorded, and the direction is the teaching.
A right decision, carrying a reason that will not hold next year
The proposal reached the committee describing fine tuning as the way to stop the component making things up. Stopping a component making things up is the one thing this tool cannot be relied on to do. The trial result was not a disappointment but the mechanism behaving exactly as it behaves. Whoever drafted the proposal did the right thing by running a 300 example trial before asking for 4,000; the trial is the reason anybody knows anything at all.
The problem is the reason that got written down. The decline was recorded on cost. The cost was 11 working days, being Rs 41,250/- of somebody's time at the bank's own assumed loading, against a build of Rs 2,40,00,000/- and a run cost of Rs 65,00,000/- a year. A reason that collapses the moment anybody divides one number by another is not a reason a decision can rest on.
Here is what that costs, and it is not money. The fault it named is a real fault and nobody has been told why this tool does not address it, so the same proposal will come back. Next time it will arrive with the cost objection pre-answered, and the room will have nothing left to say. A decision that is right for a reason nobody wrote down is a decision with a shelf life.
What did the comparison in front of the committee look like?
Over roughly the same period, three separate changes were made to the arrangement around this component, and together they moved the unsupported statement rate from 11.5 per cent to 3.5 per cent, an improvement of 8.0 percentage points. The three changes cannot be pulled apart. The bank never split the credit between them, and a split nobody measured is a guess wearing a number. The three changes are covered separately. Here they are only the figure on the other side of the comparison.
Set the three quantities side by side and the shape of the decision is immediate. The trial delivered 2.0 points, measured. The full 4,000 example build was projected by the bank at 3.5 points, never tested. The three changes together delivered 8.0 points, measured. The first two are the same tool at two scales and the third is a different tool. The honest comparison is therefore not fine tuning against nothing, but fine tuning against the three changes that were already working.
One trap in those numbers, and it is a trap because the same digits appear twice. The 4,000 example projection would leave the unsupported statement rate at 8.0 per cent. The three changes delivered an improvement of 8.0 percentage points. The two 8.0s are not the same quantity. One is a rate that remains, the other is a movement that happened, and confusing them makes the projection look like it matched. In notes out of 200 it separates cleanly: the projection leaves 16 faulty drafts standing, and the three changes left 7.
Before the control below is moved: 300 examples gave 2.0 points. What did the bank project for 4,000?
Move the number of worked examples and watch the improvement flatten
One input: how many worked examples get assembled, from none to 4,000. Everything else is held: the same component, the same 200 note test, the same starting rate of 11.5 per cent. The default is the trial as it was actually run. At 300 examples the improvement is 2.0 percentage points for about 0.8 working days of assembly. At the full 4,000 the bank projected 3.5 points for 11 working days. The reference line at 8.0 points is what the three changes together delivered, and the curve never reaches it.
Assembling 300 worked examples takes about 0.8 working days.
How is the adjustment undone, and how fast?
Reversibility sounds like housekeeping and is not. ReversalReturning to the component as it was before an adjustment, including re-running whatever was measured on the adjusted version. is the property that decides how a team is allowed to work.
Reversing an instruction is a small act. The instruction that governs this component runs to 340 words in six numbered parts. Somebody opens it, changes a part, saves it, and the next draft comes out under the new wording. If it was wrong, they change it back. The whole loop is minutes, and because it is minutes, anybody can try something on a Tuesday afternoon without asking permission.
Reversing a fine tuning is a larger act. The component is put back as it was before the adjustment. A measurement taken on the adjusted version does not describe the restored one, so whatever was measured has to be measured again. At this bank that means going back to Neelima Rao in the risk function and asking for the test to be run again, and her calendar is not a Tuesday afternoon. The cost of undoing a change decides whether the change gets tried at all, and a change that takes days to undo has to be argued for before it can be attempted rather than after.
Why does reversibility belong in a comparison of approaches at all?
So when is fine tuning genuinely the right tool?
Two questions settle it, and neither of them asks how much improvement is wanted. The first: is the fault to be fixed a fault of form? The second: do enough worked examples of the actual task already exist, or can they be produced by people whose signature means something? Two yeses make this the right tool and a single no makes it the wrong one, whatever the size of the problem.
The best case for it is unglamorous and common. The outputs are correct and inconsistently shaped. Three teams write the same note four different ways, the closing statement appears sometimes, the length wanders. And two thousand past outputs sit in a system, already signed. The fault is form, the examples exist, so the assembly cost is mostly paid and the answer is yes. The worst case is the one this bank met: the fault is support, the examples have to be made, and even the bank's own projected improvementAn estimate of what a larger effort would deliver, extrapolated from a smaller measured trial rather than measured itself. at full scale falls short of what a different tool had already delivered.
The outputs are correct and inconsistently formatted, and two thousand past ones sit signed in a system. Fine tune?
What does this change for the person sitting in the room?
For a committee member, for somebody running a lending operation, and for somebody reviewing these arrangements on behalf of a risk function, three questions cost nothing to ask and change most conversations.
The first is which fault the proposal is aimed at, and somebody should have to say the sentence out loud: is the output the wrong shape, or does it assert things the material does not hold? If the answer is the second and the proposal is fine tuning, the proposal has not found its own problem, and saying so is not obstruction. The second is what assembling the examples costs, in working days, and whose days those are. A proposal of this shape that does not carry that number has not been costed, and the number will be uncomfortable precisely because it is somebody senior. The third is how the change gets undone and how long that takes. The answer settles whether what is on the table is an experiment or a commitment.
There is a fourth habit worth taking away, and it is about how decisions get recorded rather than how they get made. When a decision is right, spend the extra five minutes writing down the reason that will still be true next year. The reason is what defends the decision when the same proposal comes back. Cost changes. Prices fall, an estimate gets sharpened, somebody finds the examples already sitting in a system. The shape of the fault does not change. Ajay Agrawal, Joshua Gans and Avi Goldfarb, in Prediction Machines, make the related point that these components produce something a person still has to act on, and the judgement about which fault is in front of the decision maker stays firmly on the human side of that line.
Is there anything a rule requires here?
No rule book settles the choice between fine tuning and any other approach. The decision is a commercial and technical one. The notes and letters concerned reach customers of a regulated lender, so the output itself does sit under supervision. The Reserve Bank of India states what it expects of such a lender on digital lending, outsourcing, customer data, consent and record keeping at rbi.org.in.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering digital lending, outsourcing, customer data, consent and record keeping | rbi.org.in |
| Ajay Agrawal, Joshua Gans and Avi Goldfarb | Prediction Machines, 2018, on a component producing something a person still has to act on | Harvard Business Review Press |
Sumeru Bank Limited, Ismail Sheikh and Neelima Rao are invented.
Educational material. Not advice on any investment, tax, budget or market position.
