RAG vs Fine Tuning: Supplying Material or Changing the Model
Retrieval supplies material at the moment of answering. Fine tuning adjusts the component once, before it runs at all, using worked examples. Because the material now arrives with the question, retrieval addresses assertions the material does not support. Examples demonstrate shape, so fine tuning addresses the shape of the output. The two repair different faults, and treating them as rivals is what leaves the argument unwinnable.
The unwinnable argument, or one shaped exactly like it, is a familiar one. Two capable people, both prepared, both quoting evidence, and after three meetings neither has moved an inch. When that happens it is rarely because one of them is being difficult. An argument that cannot be settled is almost always an argument in which the two sides are answering different questions, and this particular pair is the clearest example of it in the whole of this subject area. One person is worried that the output says things nothing supports. The other is worried that the output comes out in the wrong shape. Both worries are real. The two worries are not the same, and the two approaches on the table repair one each.
Sumeru Bank Limited, an invented lender, deployed component 8 of its retail personal loan intake chain: the assistant that produces the first draft of an exception note and of the explanation paragraph inside a decline letter. A person signs both. There are about 3,010 exception notes in a steady month. How either approach works is covered separately; what is compared below is the fault each one is built to repair.
Which fault does each of the two actually repair?
Start at a counter, where this happens often enough to be familiar. A clerk writes letters to customers all day. Two complaints come back about his work, and they arrive from different people for different reasons.
The first complaint is that a letter said something that is not true of that customer: it referred to a missed instalment in a month where the account was clean. The question is what fixes that. The clerk did not write badly, so sending him on a letter writing course fixes nothing at all. He wrote confidently about a file he did not have in front of him. The clerk needs the file on the desk, open, at the moment he writes. The second complaint is different in kind. The letters are accurate, and they begin in the wrong place, carry no reference number, run to three sheets where one would do, and read like a note to a colleague rather than a letter to a customer. Putting more papers on his desk fixes none of that. The clerk needs twenty good letters put in front of him, until the shape of a good letter is the shape his hand makes.
RetrievalSupplying relevant material to a component at the moment of answering, so the answer is produced with that material in front of it. is the file on the desk. Fine tuningAdjusting a component's own fitted numbers once, using worked examples, so its behaviour changes at the source rather than in what it is given. is the twenty good letters. Retrieval and fine tuning are not two speeds of the same repair. The two are repairs to two different faults, and the reason the argument between them never ends is that each side has correctly diagnosed a fault the other side is not talking about.
The two faults have names, and they are worth holding on to. A fault of supportAn assertion in the output that the supplied material does not contain. The sentence may even be true; nothing in front of the component establishes it. is an assertion the supplied material does not contain. At Sumeru Bank Limited that was measured: of 200 drafted notes checked against the files that produced them, 23 carried a statement that was not in the file. A fault of formA defect in the shape of the output rather than its substance: the format, the order, the register, or a required statement that is missing. is a defect in the shape of the output, meaning its format, its order, its register, or a required sentence that is missing. Both are worth repairing. Neither repair reaches the other one.
Which of the two repairs an assertion the supplied material does not contain?
Why can no quantity of worked examples repair a fault of support?
The claim carries a great deal of weight, and the obvious objection to it is fair. If thousands of examples of careful, well supported notes were fed in, would the component not learn to be careful?
The component would learn the manner of care. The manner is not nothing, and it is genuinely worth having. The component cannot learn the content of a file it has never been shown. Think about what a worked example actually is: a request that was made in the past, paired with the output somebody was happy with. Every one of those examples was written about a file that existed at the time. The file that will arrive on Thursday has not been created yet, and the customer it belongs to has not applied. None of the examples can say anything about it. A worked example can demonstrate what a good note looks like, and it cannot demonstrate the absence of a claim about material that did not exist when the example was written.
The everyday version is a schoolteacher's problem. A pupil can be trained to write beautifully structured essays, and the structure will hold in every essay she writes afterwards. She cannot be trained into knowing what happened at a meeting she did not attend. If she needs to write about that meeting, somebody has to hand her the minutes. Training and minutes are not competing solutions. Training and minutes answer two different questions, and a school arguing about which to spend money on has not yet worked out which question it has.
The same thing holds from the other side. No quantity of supplied material teaches a component the register a particular bank writes in, or which sentence has to appear in every decline explanation, or where the reference number goes. Handing somebody the correct file does not tell them how the letter should read. The material settles what may be asserted. The material says nothing about how the words are arranged.
When does each one act, and why does that decide everything else?
Placed on a line, the two separate, and everything else falls out of that picture. Fine tuning acts once, before the component has drafted a single note, in a room with nobody from the exception desk in it. Retrieval acts during every request. At this bank that means about 3,010 times a month, every time somebody stops a file and needs a note written.
Acting once and acting on every request is not a small scheduling detail. The moment at which each one acts is the origin of every other difference between the two. Because fine tuning happens before deployment, whatever it knows is whatever was true when the examples were assembled, and the only way to move that is to do it again. Because retrieval happens at the moment of answering, whatever is in the store at that moment is what the answer rests on, and the way to move it is to change the store. One of them is a photograph taken before the run started, and the other is a window open while the run is happening. The lopsided answer to the updating question below follows from exactly that.
The timing explains one more thing. The two do not interfere with each other. A fine tuned component supplied with retrieved material behaves exactly as expected: it writes in the shape the examples taught it, about the material in front of it. Nothing about the timing puts them in competition. The two act at different moments on different things.
A file arrives at the exception desk this morning. At what moment does a fine tuning that was done last month act on the note drafted for it?
What does each one need in place before it can run at all?
Both approaches ask for something to exist before they deliver anything, and the two somethings are not alike. Retrieval needs a store of material worth retrieving. At Sumeru Bank Limited that store holds 11,400 documents, being policies, procedures and product terms, cut into 47,000 passages, a little over four a document. Fine tuning needs worked examples. The proposal at this bank asked for 4,000 of them, and assembling those was estimated at 11 working days.
The evidence here is asymmetric, and the asymmetry is worth stating plainly. A tidy comparison that hid it would mislead. The bank priced one side and not the other. The 11 working days is a stated estimate, and at this bank's 7 hour working day that is 77 hours, or about 1.2 minutes an example. At 1.2 minutes an example, the examples were being selected and paired from work already signed rather than written from scratch. Priced at the assumed fully loaded Rs 9,00,000/- a year over 240 working days, a working day is Rs 3,750/- and the eleven days come to Rs 41,250/-. Nobody produced the equivalent figure for building the store. So the up-front comparison, the one everybody reaches for first, is the one where this bank's evidence is weakest, and a decision resting on it would have been resting on a number that exists for one option and not the other.
| Before it runs | Retrieval | Fine tuning |
|---|---|---|
| What has to exist first | A store of material worth supplying: 11,400 documents in 47,000 passages | Worked examples of the output wanted: 4,000 asked for |
| What that cost here | Never priced separately at this bank | 11 working days of assembly, being 77 hours, about Rs 41,250/- |
| What is spent on each answer | Material is supplied again with every question | Nothing further, the adjustment is already inside the component |
| What was actually tried | The store was built and the arrangement was deployed | A trial on 300 examples, being 0.825 of a working day |
The last row is the honest shape of the decision. Nobody at this bank spent 11 working days. The bank spent about four fifths of one, on a trial of 300 examples, precisely so that it would not have to guess. Running that trial was the right instinct. The reading it produced is the one everybody remembers, and that is where the trouble starts.
Which of the two up-front costs did this invented bank actually put a number against?
What happens to each one when the policy underneath changes?
Here is where the two stop looking like near neighbours. Policies change. Procedures are replaced. Product terms are re-issued, and at a bank they are re-issued rather often. A bank is a written set of promises that keeps being amended.
The school notice board is the example to hold. A school changes its uniform rule. If the rule lives on the board by the gate, somebody takes down the old sheet, pins up the new one, and from that afternoon every parent who reads the board gets the current answer. If instead every teacher memorised the rule in June, the school has a different job on its hands: call them back in, tell them all, and hope nobody answers from memory in the meantime.
UpdatingWhat has to happen when the underlying material changes, so that answers given afterwards reflect the change rather than the version it replaced. under retrieval is loadingPutting a new or replaced document into the store, and taking the version it replaces out.. The replacement document goes in, the withdrawn one comes out, and the next question is answered from the new one. Loading is somebody's task on a Tuesday. Updating under fine tuning is not a task at all, it is the original project again: assemble examples reflecting the change, adjust the component, and then measure it. Everything known about how it behaved was measured on the version just replaced. One of them is an operational job with a name on a rota, and the other is a project that recurs every time the material moves, and no amount of good engineering turns the second into the first.
Loading is not automatic, and at this bank it was not fast. Publishing a document to the intranet and loading it into the store were two unconnected activities that nobody had joined up, so a replaced procedure reached the store eleven days after it was published. For eleven days, answers about that procedure came from the withdrawn document rather than the replacement. The eleven day gap is a real failure and it is worth naming plainly. The failure is also a loading problem with a loading fix: join the two activities, and the gap closes. Under fine tuning there is no fix of that shape at all. The same change needs another adjustment.
Procedures are re-issued every few weeks. Which option does that favour, and on what grounds?
Why does the updating question settle most real cases before anybody measures anything?
Because of what it compares. Every other criterion here needs evidence that has to be gone out and collected. The updating question needs one fact already known inside the institution: how often the material underneath moves. Answering that takes no trial, only thirty seconds and honesty, and in a bank the answer is almost never that it sits still.
As a single sentence it becomes a test anybody can apply in a meeting. How often does the material underneath change? If the answer is continuously, or every few weeks, then one option on the table is a loading task and the other is a project that recurs on every change, and the answer is settled before anybody prepares a slide. If the answer is genuinely rarely, once a year or less, then the updating criterion stops disqualifying fine tuning and the decision moves on to the other three. The updating question is the criterion to ask first: it is free to evaluate, it is decisive in most institutions, and it is the one that a measured trial on one task says nothing about.
What does each one leave for a checker to look at?
Somebody has to check the output before it is signed. At this bank that check is five numbered steps, and the slowest of them asks the reader to notice something that is not there: an assertion the file does not carry. Noticing an absence is hard work, and how hard it is depends entirely on what the checker has been handed.
Two shopkeepers illustrate the point. One hands over an itemised bill, and the line being queried can be pointed to directly. The other states the total from memory, confidently and probably correctly, and querying it means adding the whole thing up again. Neither shopkeeper is dishonest. One of them has made checking a ten second job and the other has made it a ten minute one.
Retrieval leaves a pointerThe reference from a sentence in the output back to the passage that supports it, which a checker can open and read.: the passage a sentence rests on can be opened and read. A fine tuned component leaves nothing to point at. The sentence rests on an adjustment made inside the component weeks earlier, and no passage sits behind it to open. Checking then means establishing the support the way it was established before any of this existed, by going back to the file. The gap between the two is not a difference of degree in how convenient checking is, it is a difference in whether there is an object to check against at all.
At Sumeru Bank Limited verification fell from 9 minutes a note to 6 across the same period the three grounding changes were made. Two things have to be said about that in the same breath. The 3 minutes is real and worth having: across about 3,010 notes a month it is 9,030 minutes, being about 150 hours of somebody's reading time. And the bank never separated the three changes, so those 3 minutes belong to all three together rather than to the pointer alone. The pointer does not establish the size of the saving. The pointer establishes that a fast check is available at all, and that is what the other side cannot offer at any price.
What does a fine tuned component give a checker to look at?
The bank compared the two properly. Which one won, and was accuracy the reason?
What did the measured comparison at this bank actually show?
The comparison showed a large gap. The percentages in this case are a trap that has already caught people, so the gap is worth stating precisely, in notes rather than in percentages.
The test is 200 drafted notes checked by a person against the files that produced them. Before any change, 23 of those 200 carried a statement the file did not contain, being 11.5 per cent. The fine tuning trial on 300 worked examples brought that to 19 of 200, being 9.5 per cent, an improvement of 2.0 percentage points, or exactly four notes. The bank's own projection was that the full set of 4,000 examples would improve it by 3.5 points, leaving 16 of the 200 standing. The three grounding changes together improved it by 8.0 percentage points and left 7 of the 200. Sixteen notes still carrying an unsupported statement against seven is the clean way to state this comparison, and it is a real result that pointed the same way as everything else in this comparison.
One warning attaches to those numbers. Two different 8.0s live in this case and they are not the same quantity. The projection leaves the rate at 8.0 per cent. The three grounding changes produced a movement of 8.0 percentage points. One is a rate that remains and one is a distance travelled, and reading them as the same number turns a clear comparison into nonsense.
The measured gap was 8.0 points against 2.0. Why is that the weakest of the four reasons for the choice?
Why was a well conducted comparison still won on the wrong ground?
Everything about this decision was done properly. Both options were prepared. A trial was run rather than asserted. The reading was honest, the arithmetic was right, and the decision that came out of it was the correct one. And the ground it was won on was the ground that will not survive.
The error that gets made, and what it costs
The comparison was conducted on accuracy, and accuracy is the one criterion here that could come out differently somewhere else. 8.0 points against 2.0 is a real reading from a real test, and it is a reading about one task, one set of 300 examples, and one month. A different task would move it. A better assembled set of examples would move it. The projection to 4,000 examples was itself an admission that the trial figure was not the ceiling.
The other three reasons follow from what each approach is rather than from what a trial happened to return. Retrieval reaches a fault of support and fine tuning cannot, and that is a statement about the mechanism, so it holds on tasks nobody has ever measured. Updating is loading a document against assembling examples and adjusting again, and in an institution amending its procedures continuously that is the difference between a rota task and a recurring project. And retrieval leaves a passage a checker can open while the other leaves nothing to point at.
Here is the cost, and it is not the three weeks. Had the trial come back the other way, the decision built on it would have reversed, and all three durable reasons would have gone on being true the whole time. A decision reached on the least stable of four reasons carries nothing to the next question, and the next question is coming, because this material dates faster than anything else in this subject area.
When is fine tuning the better answer instead?
None of this is an argument against fine tuning. There are conditions under which fine tuning is the right choice, they are statable, and where they hold everything above points the other way.
There are three of them and they have to hold together. First, the fault is the shape of the output rather than its support: the notes are accurate and they come out wrong. Second, the worked examples already exist, as a by-product of work somebody has already signed, so nothing has to be assembled from scratch. Third, the material underneath changes rarely enough that re-adjusting is not a recurring project. Where all three hold, fine tuning does the job better than anything that can be bolted around the outside of the component, and it does it without adding a step to every single request.
At Sumeru Bank Limited none of the three held, and it is worth walking them. The fault was support, measured at 23 drafts of 200 asserting what the file did not carry. The examples did not exist: 4,000 were asked for and 11 working days of assembly was the estimate. And the material was anything but stable, in a store of 11,400 documents where a single replaced procedure took eleven days just to be loaded. Zero of three. Zero of three is a finding about this task at this invented bank, not a verdict on the approach, and a reader who takes away the verdict rather than the test has taken away the wrong thing.
Name the three conditions under which fine tuning is the better answer here.
Are the two of them alternatives at all?
Only in the sense that they were once written on the same slide. Set them against each other on four criteria rather than one and the comparison stops being close. On three of the four the two are not on the same scale. The fault each repairs is not a scale, it is two different repairs. How each is updated is not a scale, it is a rota task and a project. The object a checker gets is not a scale, it is a passage against nothing. Only on measured accuracy do the two sit on one axis where a number for each can be put side by side. The argument gravitates there for exactly that reason. Accuracy is the only place a scoreboard fits.
Used together each has a job, and naming the jobs is what stops the pair being treated as a contest. Retrieval supplies what the answer has to rest on. Fine tuning makes the answer come out in the shape the institution requires. A deployment can want both, for reasons that never meet, and the sentence to hold on to is that they are not rivals but colleagues who do not share a task.
The practical version of that is easy to say. If the notes are accurate and badly shaped, the shape is the work. If the notes are well shaped and occasionally assert what nothing supports, the material is the work. If both are true, both are the work, and the order to do them in is decided by which fault is reaching a customer.
Can one deployment sensibly want both of them?
Why is there no slider between the two?
How a deployment head, a reviewer or an internal auditor actually uses this
Whoever runs a deployment asks the fault question before anything else and refuses to let the meeting move on until it is answered. The method is to take twenty recent outputs somebody has complained about and sort them into two heaps: things asserted that nothing supports, and things shaped wrongly. The heaps show which approach is even relevant. A meeting choosing between two repairs without having sorted the complaints can run for weeks and never conclude. The question on the table has not been asked yet.
Whoever reviews the arrangement afterwards reads it in the opposite direction, and the useful question there is what the checker was handed. If an output cannot be traced to a passage, the check is not a light one, whatever the record says about how long it takes, and a verification time that falls without a pointer appearing is a time that fell for some other reason and is worth asking about. At this invented bank the check is five numbered steps, and the slow one asks a reader to notice an absence. Noticing an absence is the hardest thing to ask of anybody at four in the afternoon.
And anybody who has to write the decision down should write the reasons in the order of how long they will stay true. Put the fault repaired, the updating cost and what the checker gets above the accuracy figure, and note beside the figure that it belongs to one task. Do that, and when somebody re-runs the trial next year on a different task and gets a different answer, the decision does not have to be re-argued from the beginning. Separating a durable reason from a measured one has exactly that practical value, and it costs one extra sentence in the record.
Whose expectations stand over this choice at a regulated lender?
The choice between the two is a craft choice, not a regulated one, and no rule anywhere tells a bank which to use. The surrounding conduct does sit under supervision: what a lender may say to a customer, what has to be traceable to a source, what records are kept and what may be outsourced. For a regulated lender in India those expectations are set by the Reserve Bank of India, at rbi.org.in. The reason the point is worth making at all is that one of these approaches leaves a traceable pointer behind each assertion and the other leaves none, so a choice made purely on craft grounds can quietly change how easily an answer given to a customer can be traced afterwards.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering outsourcing, digital lending, customer data, consent and record-keeping | rbi.org.in |
| Securities and Exchange Board of India | Expectations where the deployer of such a component is a market intermediary | sebi.gov.in |
| Bank for International Settlements | International supervisory material on the deployment of such systems by banks | bis.org |
Sumeru Bank Limited, its intake chain and every figure attached to them are invented.
Educational material. Not advice on any investment, tax, budget or market position.
