Large Language Models: How They Work and Where They Break
A large language model produces text one piece at a time, choosing each piece from everything that came before it, using patterns fitted from a very large quantity of written material. No step in that loop consults a source, checks a fact, or tests whether the material it was handed supports what it is writing. Producing the next piece is the only thing it does, so it produces something every time.
The property that causes every difficulty with a large language model is not error. Errors are ordinary and every component makes them. The difficulty is that this one has no state that corresponds to not knowing. A person who cannot remember a figure hesitates, hedges, or goes and looks. There is no step in this process at which hesitation could be produced, so what a reader experiences as confidence is not a tone the model chose. The confident tone is the absence of any mechanism that could have produced anything else.
How does a large language model actually produce its text?
Everything else falls out of the loop, so the loop comes first. A large language modelA component that produces text a piece at a time using patterns fitted from a large quantity of written material. is handed some text. The text handed over is whatever was supplied plus whatever the model has written so far. It chooses one more piece of text to put at the end. Then it looks at the whole text again, including what it just added, and chooses one more. Then again. A paragraph is a few hundred repetitions of that same step, and the step never changes.
The unit it chooses is smaller than a word and it does not matter much here. The point is that the next pieceThe unit of text the component chooses at each step, given everything before it. It is usually smaller than a whole word. is selected on the strength of the pattern, and on nothing else. There is no register of facts alongside the loop. There is no moment where the process pauses and asks whether the sentence forming is true. Follow the loop round once, then round again, and try to point at the step where a source could be consulted. There is not one. The loop has three operations and none of them is a lookup, so anything resembling a lookup has to be built around the loop rather than found inside it.
At which step in producing the text is a source consulted?
What was it fitted on, and what does that mean it holds?
It was fitted on a very large quantity of written material. How that fitting was done, how it was checked and how anybody would know whether it worked are covered separately. The narrower and more useful question is what the fitting leaves behind. The fitting leaves patterns over that material. It does not leave a searchable copy of the material, and the difference between a pattern and a copy is where most of the trouble starts.
A household comparison holds here. Somebody who has read a few hundred tenancy agreements over twenty years can write a tenancy clause that sounds exactly right, in the correct register, with the correct rhythm, without opening a single one of those agreements. The fluency is real and it is worth something. It is also not the same thing as reading the tenant's own agreement. The clause produced is shaped by the pattern of hundreds of documents and by none of the particulars of that one.
Two consequences follow and both bite at a bank. The first is that nothing an institution wrote privately was ever in the material, so Sumeru Bank Limited, an invented lender whose retail loan intake chain supplies every figure below, has none of its own credit policy, none of its exception procedures and none of its product terms sitting inside the component it deployed. The second is that the component can still write a fluent paragraph about a bank income tolerance. The pattern of such paragraphs is thoroughly present in general written material. Fluency about a subject and access to a document are separate things, and the output gives no way to tell which one is in play.
Everything that matters at Sumeru Bank Limited therefore has to arrive as supplied materialWhat the component is handed for one particular output, as distinct from the material it was fitted on., handed over at the moment the output is produced. The file, the statement summary, the relevant procedure. If a fact is not in the supplied material, and not reachable by some connection built around the loop, then it is not available for this output, however confidently the sentence about it reads.
The component was fitted on a very large quantity of written material. Does that mean the bank's own exception procedure is inside it?
Why does it produce something whether or not it has anything to go on?
Because producing the next piece is the entire operation. There is no second operation to fall through to. Compare it with the written rule that sits a few steps away in the same intake chain. The income corroboration rule at Sumeru Bank Limited is 34 lines long, and Neelima Rao in the risk function read all 34 of them end to end in 25 minutes. The rule compares a declared income against a corroborated one and either routes the file or does not. Give it a statement it cannot parse and it stops. Stopping is one of the things it is built to do.
The drafting component has no equivalent of stopping. Hand it a file with the income section missing and it will still produce an exception note about the income, in the same length, with the same headings. The pattern of exception notes says that an exception note has an income paragraph. Nothing about the missing section changes the operation being performed. A written rule can fail to produce an answer and a text producing loop cannot, and that single asymmetry is why one of them can be trusted to go quiet and the other cannot.
The component can be made to go quiet, and Sumeru Bank Limited did so. In month nine a sixth part was added to the written instruction the component works from, telling it what to do when the supplied material does not support an answer. The rate at which it declined to answer moved from nothing at all to 12 of 200 questions, being 6.0 per cent, on that bank measurement. The result repays careful reading. It does not show that the component acquired a sense of its own limits. It shows that somebody wrote an instruction, and that the instruction changed the pattern the loop was following. Declining to answer at Sumeru Bank Limited is a piece of construction, and construction can be removed, mis-specified or forgotten in the next version of the instruction.
Why is there no state inside it corresponding to not knowing?
Set the two paths beside each other and the answer becomes obvious rather than philosophical. A colleague asked for a figure they do not have runs a process that branches. There is a moment, however brief, where they check whether the thing is held, and the branch has a second arm: say so, ask somebody, or go and find it. The branch is what produces every hedge a competent person offers.
The loop has no branch. The question arrives, the pattern selects the next piece, the sentence forms. There is no place to attach a hesitation because there is no test whose answer could trigger one. The essential point is this: the confident tone is not a stylistic choice the component made and not a defect in one version of one product, it is the shape of a process that has exactly one thing it can do. Every remedy that treats it as a tone problem, including asking the component in writing to be more cautious, is working on the wrong layer.
Why does the output always arrive in the same confident tone?
Break one: what happens when it asserts what nothing supports?
The first break is the one everybody has met. An assertionA statement in the output presented as fact, whether or not anything supplied supports it. appears in the output that is not in the file, not in any procedure, and not anywhere else the component could have got it. The assertion is not flagged, not softened and not set apart. It sits in the middle of a paragraph in exactly the register of the sentences around it.
Look at one draft. On a file stopped by the income rule, the desk at Sumeru Bank Limited receives a first draft of the exception note. The applicant declared Rs 45,000/- a month. The three month statement shows a median salary credit of Rs 38,000/-. The gap is Rs 7,000/-, or 15.6 per cent of the declared figure, and that is above the tolerance the bank has set for itself. All of that is in the file. The draft then states the gap as Rs 7,500/- and adds a sentence saying the employer confirmed a salary revision in the current quarter. Neither of those is in the file. Nothing about the way those two sentences are written distinguishes them from the three true ones above, and the check that catches them is slow for exactly that reason.
Break two: what happens when more is supplied than it can hold at once?
The second break is quieter and it is the one people miss. There is a limit on how much material can be handed over with a single request. Where the material handed to the drafting component at Sumeru Bank Limited went past that limit, the oldest part of it was dropped, and dropped silently, with no marker in the output saying that anything had gone.
The numbers there belong to that bank and to that month. The material supplied with an exception note averaged 4,800 words, comfortably inside the limit. The longest five per cent of files went past 12,000 words. On the first measurement of 200 notes, 9 of them, being 4.5 per cent, had been drafted from a file whose earliest section had already been cut away before the loop ever saw it. The output on those 9 was not wrong about the material it had; it was complete and confident about a file it had only partly been given, and no reader of the note could have known which 9 they were. How much can be held at once is covered separately, and it moves. Going past that limit removes material rather than raising an objection, and that is the part worth carrying away.
Break three: what happens when the answer needs something it was never given?
The third break has the cleanest fix, so it is the cleanest of the three. Some things simply are not in any document that could be handed over. The current outstanding balance on an account sits in a system. So does the count of active obligations, the date of the last credit, and the exact figure in a structured field that a person typed into a form. If none of that is supplied, the loop still produces a sentence containing a number. The pattern of exception notes contains numbers in that position.
Sumeru Bank Limited was able to measure this one cleanly, and it is the only clean attribution in the whole exercise. Of the 23 unsupported statements found in the first 200 drafts, 6 were figures the component had restated from prose. Across the whole 200, 31 notes restated a figure at all. So 6 of 31 restatements were wrong, being 19.4 per cent on that bank measurement. In month nine a connection was added letting the component read the structured fields on the file instead of copying figures out of prose. At the second measurement 34 notes restated a figure and none of them was wrong. Three separate changes were made between the two measurements and the bank could never fully separate them, but no change other than that connection touched figures. A figure that is fetched rather than written cannot be restated wrongly, and that is the one part of this problem that can be closed rather than reduced.
How can the three breaks be told apart?
An output arrives with something wrong in it and somebody asks what happened. Three questions sort it, and the order matters. The second and third questions can hide the first.
First, was any part of the supplied material dropped before the loop read it? If yes, the fault is break two, and the fix is in what was supplied rather than anywhere near the wording. Second, did the answer need a value that lives in a system rather than in a document? If yes, that is break three. The value was never in the room, so no amount of better instruction will help. Only if both answers are no is the fault break one, the plain assertion of something nothing supports. The three can also happen together, and when a figure is cut away by truncation and then confidently supplied by the pattern, the fault belongs to break two even though the symptom looks exactly like break one.
A draft states a figure that appears nowhere in the file and nowhere in any procedure supplied with it. Nothing was truncated. Which break is that?
Six of the 23 unsupported statements were figures the component had restated from prose. What does that suggest about the fix?
What did the 200 drafted notes actually show?
Component 8 in the intake chain at Sumeru Bank Limited is the drafting assistant. The assistant produces the first draft of two things, the exception note a desk officer writes on a stopped file and the explanation paragraph inside a decline letter, and a person signs both. Component 8 runs on 3,010 exception notes a month. Ismail Sheikh, who runs the exception desk, had it measured on 200 notes in month six, before anything at all had been built around it.
Drafting time fell from 18 minutes to 4. Verification added 9 minutes, so the net saving was 5 minutes a note. Across 3,010 notes that is 15,050 minutes, being about 250.8 hours a month. Against that, 23 of the 200 drafts, being 11.5 per cent, contained a statement that was not in the file. Of those 23, six were restated figures and seventeen were other statements. Twenty one were caught in verification and two reached a customer, and 21 plus 2 is 23.
| The first measurement, 200 notes, month six | Figure |
|---|---|
| Drafting time before | 18 minutes |
| Drafting time after | 4 minutes |
| Verification added | 9 minutes |
| Net saving a note | 5 minutes |
| Drafts holding a statement that was not in the file | 23 of 200 |
| Of those 23, restated figures | 6 |
| Of those 23, other statements | 17 |
| Caught in verification by a person | 21 of 23 |
| Reached a customer | 2 of 23 |
The 9 minutes of verification is not one act of reading. Verification is five numbered checks: every figure traced back to the supplied material, every assertion traced to a passage, nothing asserted that the material does not contain, the required format and required statements present, and the decline to answer case handled properly where it arises. The third of those is the one that catches the 21. Asking the reader to notice something that is not there rather than something that is makes it the slowest of the five. Every minute taken out of verification comes out of one of those five checks, and only one of them is doing the catching.
Can it check its own output, and how far does that get?
The obvious next move is the one the desk made. If a person catches 21 of 23, why not ask the component to read its own draft first and mark anything it cannot support? Machine time is cheap and human time is not, so a self-reviewAsking the same component to check its own output before a person sees it, using the same loop that produced it. pass looks like something for nothing. Sumeru Bank Limited ran it on the same 200 notes.
One pass flagged 4 of the 23. A second pass took it to 6. A third took it to 7. A fourth added nothing at all, leaving it at 7 of 23, being 30.4 per cent. The climb has reached a plateauThe point past which more of the same effort adds nothing measurable., and it sits at under a third of the faults, against 21 of 23 caught by a person working the nine minute check. The self-review counts are one bank measurement on 200 of its own notes and they are not a statement about what such components can do in general, but the shape of the curve is worth taking from them.
Before the control below is moved: the component reviews its own draft once. How many of the 23 unsupported statements does that one pass flag?
Add self-review passes, and watch where the climb stops
One control: the number of times the component is asked to review its own draft, from none to four. One consequence: how many of the 23 unsupported statements get flagged, drawn as 23 cells with a fixed marker where a person lands. Two views are available. Flagged shows what self-review finds. Goes through unmarked shows what is left, and that is the number that decides anything. The default below is the first thing the desk tried: one pass, flagging 4 of the 23, being 17.4 per cent, against 21 of 23 caught by a person in nine minutes. The climb stops at 7 of 23 on the third pass and the fourth adds nothing.
Self-review passes: 1
One self-review pass flags 4 of the 23 statements that were not in the file, being 17.4 per cent, and leaves 19 of them unmarked. A person working the full nine minute check catches 21 of the 23.
Why does a third and a fourth pass buy nothing?
Because the second look is the same loop as the first. The second look has the same fitted patterns, the same absence of a lookup step, and the same material in front of it. Whatever made a statement look plausible enough to write in the first place makes it look plausible enough to leave alone on review, and asking again does not introduce any new information into the room. The first pass picks up the faults that are internally awkward, the second picks up a few more, and after that the passes are producing text about text.
Now the part that matters for anybody costing a control. Self-review flags 7 of 23 and a person catches 21 of 23. Nobody knows which 7 they are until everything has been checked, so the 7 cannot be subtracted from the person work. Finding out whether a flagged statement is genuinely unsupported requires the check. Finding out whether an unflagged statement is fine requires the same check. A control that finds some of the faults does not reduce the work of finding the rest, so the flags change the order the queue is worked in and change nothing else. Reordering is worth something. The notes most likely to be wrong get read while the reader is fresh. Reordering is just not what anybody was hoping to buy.
Self-review flags 7 of the 23. Can the nine minute verification be shortened by those 7?
The error that gets made, and what it costs
The self-review result was first read as encouraging. A second component catching about 30 per cent of the faults, for machine time and no human time at all, looks like a control obtained for free. The proposal that followed was to take the verification down from 9 minutes to 6, on the grounds that some of the work had already been done.
Read the two counts again before agreeing. Self-review flags 7 of 23 and a person catches 21 of 23. The 7 are not a subset anybody can point at, so the person still reads every assertion in every note, and nothing has been taken off the desk. Push the proposal to its limit and self-review on its own would leave 16 of the 23 unmarked, against the 2 that got past the full check. Sumeru Bank Limited never measured a six minute check without anything else built around it, so the honest answer sits somewhere between those two counts and the bank could not say where. Proposing a number in that gap, rather than the range, is the error.
The cost of getting this wrong is not measured in desk minutes. Two of the 23 reached a customer under the full check. The two that reach a customer under a shortened one are letters and notes carrying a statement about that customer that nobody could support, and a person signed each one.
Who sets the expectations where a drafted output reaches a customer?
Where a component drafts material that goes to a customer of a regulated lender in India, the expectations on that lender are set by the Reserve Bank of India. The Reserve Bank publishes its position on outsourcing, digital lending, customer data and consent, and record keeping at rbi.org.in. Where the deployer is a market intermediary rather than a lender, the Securities and Exchange Board of India sets the equivalent expectations at sebi.gov.in. Requirements, thresholds and effective dates move, so the position to rely on is the current one published by the issuing body.
Which of these properties are permanent, and which are only current practice?
Notes on this subject go out of date faster than almost anything else in finance. An understanding stops dating when every claim is sorted into one of two columns before anything is done with it.
In the left column go the structural propertySomething true of how the component works, as distinct from what current practice happens to be. statements. The component produces text one piece at a time. No step in that loop consults anything. The component produces an output in every case and has no state corresponding to not knowing. The component holds patterns over what it was fitted on rather than a retrievable copy of it. Nothing about how a sentence reads says whether anything supports it. The five statements describe the component itself rather than how well it is currently doing, so all five will still be true when everything around them has moved.
In the right column go the arrangements. How much material can be supplied at once. Whether a connection to another system exists and what it can reach. Whether anything is retrieved and placed in front of the loop before it writes. Whether the arrangement declines to answer when nothing supports an answer. And every measured error rate, all of them from one invented bank, one month and 200 notes. A claim in the left column is about the component itself; a claim in the right column has a date on it, whether or not the date is printed beside it.
Which of these will still be true in five years: that the component produces text with no lookup step, or that it can hold a particular amount of material at once?
What a reviewer at a lender actually does with this
The practical use of everything above fits into one habit. When somebody proposes taking time out of the verification on a drafted output, ask which of the five numbered checks is being shortened. At Sumeru Bank Limited the five are: trace every figure to the supplied material, trace every assertion to a passage, confirm that nothing is asserted which the material does not contain, confirm the required format and required statements, and handle the decline to answer case. The third one is the slowest and it is the only one catching the unsupported statements.
If the saving is coming out of checks four and five, it is a genuine efficiency and worth having. If it is coming out of check three, the saving is being taken directly out of the one control that works, and the proposal should say so in those words rather than presenting a minute count.
The same habit works on any claim about a component of this kind, whether the person holding it is a reviewer, an analyst reading a deployment note, or somebody in a household deciding how much to trust a drafted summary of an agreement. The questions are what would have to be true for the statement to be checkable, and then whether anybody checked it. Agrawal, Gans and Goldfarb put the general shape of this well in Prediction Machines, published in 2018: a fitted component supplies something, and a person still has to decide what to do about it. On a drafted note the decision is not whether the writing is good; it is whether every sentence in it is supported, and nothing about the writing helps answer that.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering outsourcing, digital lending, customer data and consent, and record keeping, where a drafted output reaches a customer | rbi.org.in |
| Securities and Exchange Board of India | Equivalent expectations where the deployer of such an arrangement is a market intermediary rather than a lender | sebi.gov.in |
| Bank for International Settlements | International supervisory material on the deployment of automated and assisted processes by banks | bis.org |
| Agrawal, Gans and Goldfarb | Prediction Machines, 2018, on a fitted component as something that supplies a prediction a person still has to act on | Harvard Business Review Press |
Sumeru Bank Limited, Neelima Rao and Ismail Sheikh are invented.
Educational material. Not advice on any investment, tax, budget or market position.
