Retrieval Augmented Generation: Grounding in an Institution's Own Documents
Retrieval augmented generation means finding passages relevant to a question and supplying them to a generative component along with the question, so the answer is produced from material rather than from the component's own patterns alone. The component answers whatever it is given, and answers anyway when it is given nothing useful. The accuracy of the whole arrangement is therefore mostly the accuracy of the retrieval.
Two components sit in series here, and in most deployments only one of them has ever been examined. The generative one is the interesting one, the one people demonstrate and argue about and write policies for. The one in front of it decides the result. Sumeru Bank Limited, an invented bank, measured the two halves separately instead of reporting their combination, and its numbers settle that sentence rather than assert it. How passages are turned into positions, how a store is prepared, and how a component produces text are each covered separately. Retrieval augmented generation is what happens when the three are put in series.
What are the five steps, in the order they happen?
Consider a familiar exchange. A customer rings a bank helpline and asks whether a fee applies to an account. The person on the line does not answer from memory. The person types something, pulls up two or three screens, reads what is in front of them, and then answers. The quality of what the customer hears is decided in the seconds before the person speaks, when they either found the right screen or did not. If they pulled up the wrong product's fee schedule, they will read it out with exactly the same confidence.
Retrieval augmented generationFinding relevant passages and supplying them to a generative component along with the question, so the answer is produced from that material. is that arrangement, built. The arrangement has five steps, and all five are worth naming. People usually name two, and then wonder why the answers behave as they do.
Sumeru Bank Limited runs this arrangement beside the drafting assistant on its retail loan intake chain. A desk officer stops on an exception, asks what the written procedure says, and the arrangement produces a paragraph. The store holds 11,400 documents cut into 47,000 passages, and six passages come back to every question. Six is a setting somebody chose. The contents of those six are not.
The arrows in the figure all point one way. There is no arrow going back from step 4 to step 2. The missing arrow is what makes the second step decisive. A generative component that has been handed six passages will work extremely hard on those six, and it has no way at all of asking for a seventh.
Where do the passages come from, relative to the question being answered?
Why are the passages supplied with the question rather than looked up afterwards?
Because the order decides where the sentences of the answer come from. Hand the question over alone and the component produces something from a very large quantity of general writing that contains not one line of this bank's procedures. Hand it the passages first and the answer is assembled from material that exists, sitting in a store somebody at the bank is accountable for.
Supplying the passages is the whole of the word grounding, and the word is worth being blunt about. Grounding does not mean the answer is correct. Grounding means the answer had something in front of it. Supplying material changes where the sentences come from; it does not install any obligation to use them.
Nothing in the writing separates the fourth sentence from the first three. The fourth sentence is not vaguer, not hedged, not shorter. A sentence produced from nothing supplied looks precisely like a sentence produced from passage 4.
Supplying six passages with the question changes one thing for certain. Which?
Where does the accuracy of the whole arrangement actually come from?
Sumeru Bank measured this properly, and that is rarer than it sounds. Somebody took 200 questions, and for each one recorded two things instead of one: whether the passage actually needed came back among the six, and whether the answer was right. RetrievalThe step that searches the store and returns passages for a question. and generation were scored apart before they were scored together.
The passage actually needed was among the six for 178 of the 200 questions. The hit rateHow often the passage actually needed is among the passages returned. The searching step owns this reading, not the component that writes the answer. is 89.0 per cent. On those 178, the answer was right 171 times, being 96.1 per cent. On the other 22 the answer was right 3 times, being 13.6 per cent. Add them: 171 plus 3 is 174 right out of 200, being 87.0 per cent overall.
| The 200 questions | Questions | Right | Wrong | Right, per cent |
|---|---|---|---|---|
| The needed passage was among the six | 178 | 171 | 7 | 96.1 |
| It was not, and the answer came anyway | 22 | 3 | 19 | 13.6 |
| Both together, which is what gets reported | 200 | 174 | 26 | 87.0 |
Splitting the 200 questions into the ones where the needed passage came back and the ones where it did not is the diagnostic method of this whole subject. The 87.0 per cent is not a property of the generative component. The 87.0 per cent is a weighted average of two wildly different numbers, 96.1 and 13.6, and the weights are set by the retrieval. Move the weights and the headline moves, with nothing whatsoever having changed about the component that writes the sentences.
Now watch what happens when the retrieval improves and nothing else does. Take the hit rate from 89.0 to 95.0 per cent, hold the two conditional readings where they are, and the overall figure goes to 92.0 per cent. Take it down to 80.0 and the overall figure goes to 79.6. A six point swing in the headline, with the generative component untouched throughout.
Overall accuracy on the 200 questions is 87.0 per cent. What does that establish about the generative component?
The retrieval is improved and the hit rate goes from 89.0 to 95.0 per cent. Nothing about the generative component changes. What happens to overall accuracy?
What happens on the questions where the needed passage was returned?
On those 178, the arrangement did what it was built to do. The answer was right 171 times. The remaining 7 are the interesting ones, and they are only 3.9 per cent of the population where everything upstream worked.
The 7 wrong answers are the honest residue of the generative step. The material was there, in front of it, and the answer still came out wrong: a condition read past, a figure restated from the wrong line, an obligation stated more broadly than the passage stated it. Seven questions in 178 is the size of the problem that is genuinely about the component writing the answer, and at this bank it was a small problem sitting underneath a much larger one.
The size of that residue decides where effort goes. A team that hears 13.0 per cent wrong will start rewriting instructions, trialling adjustments to the component, and arguing about how the answer is produced. All of that work aims at 3.5 points of the 13.0. The other 9.5 points sit in a step nobody in that meeting has opened.
On the 178 questions where the passage was found, 7 answers were still wrong. What is the right way to read those 7?
What happens on the questions where it was not?
Twenty two questions, and on every one of them the store held nothing useful among the six that came back. Not nothing at all: six passages did arrive, on roughly the right subject, looking exactly like the six that arrive when the answer is present. The needed one simply was not among them.
The component answered all 22. The answer was right on 3 of them, being 13.6 per cent, and those three are luck rather than skill: general knowledge of how such procedures usually read, landing on the version this bank happened to have written. The answer was wrong on 19.
Nineteen wrong answers out of a total of 26 wrong answers means 73.1 per cent of everything this arrangement got wrong came from 11.0 per cent of the questions. Errors cluster the same way exception causes cluster on the intake chain: a small, identifiable population producing most of the damage. A population that can be named can be acted on; an average on its own leaves nothing to act on.
What does answering anyway buy, and what does it pay?
Answering anywayProducing an answer when nothing useful was retrieved, rather than saying that nothing useful was found. is the default behaviour. There is no state inside the arrangement corresponding to having found nothing, and no step in the five that tests for it. Six passages arrived, they were handed over, an answer was produced from them. As far as the arrangement is concerned that is a completed job.
So price it. Answering anyway got 3 questions right that would otherwise have produced no answer, being 1.5 points of the 200. Answering anyway got 19 questions wrong that would otherwise have produced no answer, being 9.5 points. For every one extra right answer it buys, it pays six and a third wrong ones. Counted rather than shared out, the number of wrong answers goes from 7 to 26, or 3.71 times as many.
The failure: a trade nobody chose, running by default
Set the trade out plainly and it is absurd. Buy 1.5 percentage points of correct answers, pay 9.5 percentage points of confidently wrongAn answer that is wrong and carries no signal in its wording, length or tone that it might be. ones. Nobody would sign such a trade knowingly. The trade is the default behaviour of every arrangement of this kind unless somebody writes the instruction that stops it, and at Sumeru Bank that instruction was part 6 and it did not exist for the first three months the arrangement ran.
The reason the trade survives is not that anyone defends it. The reason is that the trade improves the only number usually reported. 87.0 per cent correct reads better than 85.5 per cent correct, and on that measure it genuinely is better. The 11.0 per cent of questions that would have produced no answer have quietly become 9.5 points of wrong answers and 1.5 points of right ones, and a single accuracy figure cannot see the difference. The trade has the same shape as a shopkeeper who never says he is out of stock and hands the customer something similar instead: his sales figure improves, and every complaint that follows arrives at a different desk.
Who makes this mistake: a project team reporting one accuracy number to a steering group, and a steering group that asks for one. The cost: the reported figure improves while the count of wrong answers reaching a person nearly quadruples, and the objectively worse arrangement is the one that gets approved.
Why does answering anyway survive in most deployments?
What would declining have cost instead?
DecliningProducing no answer, and saying so, because nothing useful was retrieved. means the component says it did not find material that supports an answer, and stops. Run the same 200 questions through an arrangement that declines and the readings are 85.5 per cent right, 3.5 per cent wrong and 11.0 per cent declined.
Notice what has happened to each column. The right answers fall by 1.5 points, a fall that is real and should be stated. The wrong answers fall from 13.0 per cent to 3.5, a fall of nearly three quarters. And an eleven point block appears that did not exist before: questions with no answer. A question with no answer is a state somebody can see, count, route to a person and put in a report.
Put the two reported figures side by side and the problem is obvious: 87.0 and 85.5. A committee choosing between them on that line alone chooses the worse arrangement, and does so reasonably. Nothing on that line says the first arrangement has 26 wrong answers in it and the second has 7. The count of declined answers is not a nice extra to report; without it the accuracy figure is not interpretable at all.
Before the control below is moved: the component declines instead of answering when nothing useful is found. What happens to the share of right answers?
Move the retrieval hit rate and watch 200 questions change colour
Each grid is the same 200 questions. The left grid answers anyway; the right grid declines. Only the hit rate moves.
Retrieval hit rate: 89.0 per cent
At the measured hit rate of 89.0 per cent, answering anyway gives 87.0 per cent right and 13.0 per cent wrong with nothing declined, while declining gives 85.5 per cent right, 3.5 per cent wrong and 11.0 per cent declined.
Educational illustration. The worked default is the measured reading: a hit rate of 89.0 per cent giving 87.0 per cent right and 13.0 per cent wrong when the component answers anyway, against 85.5 right, 3.5 wrong and 11.0 declined when it does not. Assumptions held constant: 200 questions at one invented bank, six passages returned throughout, and the two conditional readings of 96.1 and 13.6 per cent held fixed as the hit rate moves, which a real change to the retrieval would not do exactly. Dot counts are the percentages rounded to whole questions. Every figure is Sumeru Bank Limited's own and describes one deployment.
Why must the two numbers be measured separately?
Because a single figure moves for opposite reasons and offers no way to tell which. Suppose the reported accuracy falls from 87.0 to 84.0 per cent one quarter. Working the arithmetic backwards on the two conditional readings gives two clean stories, and both fit.
Story one: the retrieval got worse. Hold the two conditional readings fixed and a hit rate falling from 89.0 to about 85.3 per cent produces exactly 84.0 overall. Nothing about the generative component changed. Something upstream did: new documents added without being cut properly, a whole class of question arriving that the store was never stocked for, an old version competing for the six places.
Story two: the retrieval was untouched and the answers got worse. Holding the hit rate at 89.0 and letting the accuracy given a hit fall from 96.1 to about 92.7 per cent lands on 84.0 as well. Two repairs with nothing in common, and the reported figure is identical under both. Both of these are arithmetic on the bank's locked readings rather than measurements the bank took.
Overall accuracy fell by three points this quarter. What could have caused it?
What does this arrangement not fix?
Deployments usually assume away the second half of this list. The arrangement supplies material and attaches a pointer, and supplying and pointing is the entire scope of the arrangement.
The arrangement does not decide whether the passage that came back is the version in force today. The arrangement does not check that the passage supports the sentence written from it. The arrangement has no behaviour at all for the case where nothing useful came back, unless somebody wrote one. Each of those three is an arrangement built around this one, and not a single one of them arrives with it.
Which of these does the arrangement fix on its own?
Which two numbers show whether it is working?
Two, reported side by side, always. How often the needed passage is among those returned, and how often the answer is right given that it was. At Sumeru Bank those were 89.0 per cent and 96.1 per cent. Every diagnosis of this kind depends on having the two numbers apart, and none is possible from their combination.
A third line belongs beside them: how many questions produced no answer. Without that line the first two can be reconstructed only by someone who already knows they exist. An arrangement reported as one accuracy figure is an arrangement nobody can diagnose, improve or safely compare against anything.
Reporting the parts is a general habit rather than a technical trick, and it is the same habit that stops a household misreading its own month. Total spending went up by six thousand rupees. Was that more trips to the market, or the same trips costing more? The total cannot say, and the two have different answers. DecompositionSplitting one reported figure into the separate parts that produced it, so that a movement in the figure can be traced to a cause. is what turns a number that can be watched into a number that can be acted on.
Who at this bank actually needs the decomposition, and for what?
Four people, four different uses of the same two numbers
Revathi Balan, accountable for the scoring model on the same intake chain, reads the decomposition as a scoping question rather than a performance one. If the wrong answers concentrate in the questions where nothing was found, then the control that matters is the behaviour when nothing is found, and that is a line of written instruction rather than a change to any component. She does not need to understand how the answer is produced to sign off that control, and that is exactly why the decomposition is worth insisting on before the meeting rather than during it.
Neelima Rao, doing an independent review, goes at it from the measurement rather than from the output. She asks a single question: were the two numbers recorded separately at the time, or reconstructed afterwards? If nobody wrote down whether the needed passage came back, the accuracy figure cannot be decomposed later and the whole quarter of measurement produces one uninterpretable number. Neelima Rao can write that finding without reading a single answer.
Ismail Sheikh, running the exception desk, uses it to decide what he is asking his officers to do. If the arrangement declines when it finds nothing, an officer is checking a draft against a passage. If it answers anyway, that officer is also silently being asked to notice the answers that came from nowhere. Spotting something that is not there is the slowest kind of checking there is. At the month 10 reading, verification stood at 6 minutes a note across about 3,010 exception notes a month, being 18,060 minutes or about 301 hours of somebody's reading time, and at the bank's assumed fully loaded cost of Rs 9,00,000/- a post that is about Rs 19,35,000/- a year. The yearly figure is arithmetic on the bank's locked figures rather than a separately measured amount, and it is the budget line that pays for the third of the five checks.
Ashok Pillai, in technology risk, uses it as a register question rather than a measurement question. An arrangement of this kind is two components, and if only one of them is written down, the register is wrong in a way that no amount of monitoring on the recorded one will reveal.
Where an answer produced from internal documents reaches a customer
Where a regulated lender in India states a position to a customer that was produced from its own internal documents, the expectations on that lender covering records, disclosure, outsourcing and the use of customer data sit with the Reserve Bank of India, which publishes its position at rbi.org.in. Where the deployer is a market intermediary rather than a lender, the equivalent expectations are set by the Securities and Exchange Board of India at sebi.gov.in. Requirements, thresholds and effective dates change, and the current position appears at the issuing body's own site.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering records, disclosure, outsourcing and the use of customer data | rbi.org.in |
| Securities and Exchange Board of India | Equivalent expectations where the deployer is a market intermediary rather than a lender | sebi.gov.in |
| Bank for International Settlements | International supervisory material on the deployment of such arrangements by banks | bis.org |
| Ajay Agrawal, Joshua Gans and Avi Goldfarb | Prediction Machines, for the framing of a learned component as producing something a person still has to act on | Harvard Business Review Press |
Sumeru Bank Limited, its retail loan intake chain, Revathi Balan, Neelima Rao, Ismail Sheikh and Ashok Pillai are invented.
Educational material. Not advice on any investment, tax, budget or market position.
