Hallucination: Confident, Fluent and Wrong
A hallucination is an output asserting something the material supplied to the component does not contain. Unsupported is not the same as untrue: an unsupported statement can happen to be correct, and a supported one can be wrong because the material was wrong. Defining it against the supplied material rather than against the truth is what makes it something a person can check in a fixed number of minutes.
The word makes it sound like a malfunction, and that is the first thing to put down. The component has no step at which it stops and consults something, and no state corresponding to not knowing. Producing a sentence the file does not support is therefore the ordinary operation of the thing, observed in the cases where the material it was handed happened not to contain what the sentence needed. Everything practical here follows from treating that as a rate to be measured rather than a fault to be removed.
The measurements here come from Sumeru Bank Limited, an invented mid-sized Indian bank whose retail loan intake chain supplies every figure in this guide. A drafting component writes the first draft of two things there: the note a desk officer puts on a stopped file, and the explanation paragraph inside a decline letter. A person signs both. The rate was measured twice, on 200 notes each time. One reading without the other is either a sales pitch or a scare, so both are set out below.
What exactly is a hallucination, and what is it not?
A hallucinationAn output asserting something the material supplied for that output does not contain. is an assertion in the output that the supplied material does not contain. The comparison against the supplied material is the whole test. The test asks nothing about how the sentence was produced, nothing about how confident it reads, and nothing about whether the sentence is true. One question, and a person holding the file can answer it: does this material say this?
The everyday version is a reference letter. A neighbour asks for one to be written for a young man applying for a room. The writer has met him four times on the stairs. The letter says he is quiet, tidy and reliable with money. The first is something observed. The second is a reasonable guess. The third has no basis behind it at all, and it may well be correct. Nothing the writer was given supports the third, so it is still the sentence that should not be in the letter. The landlord reading it cannot tell which of the three sentences is which. The problem in a drafted note has exactly that shape.
An unsupportedNot present in the material the component was given for this particular output. statement and an untrueContrary to fact, which is a different test from support and not the one used here. one are two different tests, and they cut across each other rather than lining up. A statement can be unsupported and correct. A statement can be supported and wrong. Where the material itself was wrong, the failure is real and a completely different job to fix. The two tests produce four cases, and only one of the two tests can actually be run by somebody holding one file and a clock.
Is an unsupported statement the same thing as an untrue one?
Why define it against the supplied material rather than against the truth?
Because one of those definitions can be turned into a task with a name, a duration and a person attached, and the other cannot. A desk officer asked to confirm that every sentence in a draft is true has been handed something with no end to it. She would need to know the applicant, the market, the history and the future. Asked instead whether every sentence in the draft is present in the material in front of her, she has been handed something she can finish, and finish in a measurable number of minutes.
A bank already makes the same move everywhere else. A cashier does not verify that a cheque is honest; she verifies that it matches the mandate on record. A stores clerk does not verify that a delivery was fairly priced; he verifies that it matches the order. A workable control is almost always a matching test against something on the table rather than a judgement about the world, and defining this fault against the supplied material is exactly that move applied to text.
There is a cost to the choice and it should be stated. Defining the fault this way means the bank counts some sentences that were in fact perfectly accurate, and misses nothing that was unsupported but happens to be right. The definition is deliberately stricter than truth in one direction and silent in the other. A check has to make that trade to be finishable, and pretending otherwise is where most verification designs quietly fall apart.
Why is producing one not a malfunction of the component?
Because there is no moment in its operation where it could have done otherwise. The component continues text. The component has no register that fills up with what it does not know, and nothing in it corresponds to reaching for a file and finding the drawer empty. When the material contains what the sentence needs, the sentence comes out supported. When the material does not, the sentence still comes out, in the same tone, at the same length, with the same fluency. Fluency is the one property the arrangement produces reliably, and it is entirely detached from support.
The detachment of fluency from support is the whole difficulty. In almost every other control a bank runs, the signal of trouble is roughness: a form filled in badly, a figure that does not foot, a statement that arrives late. Here the faulty output looks exactly like the sound one. The faulty sentence is the same length, uses the same vocabulary and sits in the same place in the paragraph. Nothing about an unsupported sentence advertises itself, so the fault has to be found by a procedure rather than noticed by a reader.
Why measure the rate at all, rather than simply checking every output and moving on?
How often did it actually happen, in one measured deployment?
In month 6, before any change was made to how the component was fed, Sumeru Bank Limited took 200 exception notes and had every draft compared line by line against the file it was drafted from. Twenty three of the 200 drafts, being 11.5 per cent, contained at least one statement that was not in the file. Twenty one of those 23 were caught by the person doing the check. Two were not, and they went out. 21 plus 2 is 23, and the bank published all three numbers together.
Both halves of that hold at once. One hundred and ninety eight of the 200 signed outputs, 99.0 per cent, carried nothing the file did not support, and the check found 21 of the 23 faults that were there to find. The arrangement was doing most of its job. And two people received a letter about themselves containing a sentence their own file never said. Two is not an acceptable residual, and the 198 do not make it acceptable. Both sentences are true at once, and an account that drops either one is doing something other than teaching.
| The first measurement, month 6 | Count | Of 200 |
|---|---|---|
| Drafts checked against their file | 200 | 100.0 per cent |
| Drafts carrying a statement not in the file | 23 | 11.5 per cent |
| Of those, caught by the person verifying | 21 | 10.5 per cent |
| Of those, not caught, and sent to a customer | 2 | 1.0 per cent |
| Drafts signed with nothing unsupported in them | 198 | 99.0 per cent |
Every figure in that table belongs to one invented bank and one deployment of one component on one kind of document. The rate is not a property of anything in general, and the most common misuse of a number like 11.5 per cent is to carry it to a different task as though it travelled. A measured rate is a fact about a deployment, never a fact about a technology, and the only honest way to get one for a particular arrangement is to measure that arrangement.
What kinds of unsupported statement are there?
The bank split its 23 into two groups and the split turned out to matter more than the total did. Six of the 23 were restated figuresA number written out into the text rather than read from the system that holds it., meaning a number written out into the sentence when a number was sitting in a field somewhere. Seventeen were everything else: a characterisation, a reason, a piece of context, a sentence about the applicant's circumstances. 6 plus 17 is 23.
The reason the split matters is that the two groups have different ceilings. A figure has a definite right answer sitting in a definite place, and the component can be connected to that place and made to read it rather than write it. Once that connection exists nothing is left to guess at, and no residue is left over. A characterisation has no such place. There is no field in any system holding the sentence about whether this applicant shows signs of financial strain, so the best available fix reduces the fault and does not abolish it.
Why do the two kinds need different fixes?
Because a fix can only be as complete as the thing it appeals to. The fix for figures appeals to a record: there is a monthly salary credit sitting in a field, and if the component reads that field the sentence containing it is correct by construction. The fix for the other 17 appeals to a body of text: there are policies, procedures and product terms, and if the relevant passage is put in front of the component the sentence is more likely to be supported. The two fixes are not the same kind of promise. One ends in a value, the other ends in a likelihood.
The difference in what the two fixes can promise is the practical reason to do the split before designing anything. A team that treats the 23 as one problem will fund one fix. The difficult group dominates the conversation, so the fix funded will be the harder, slower, weaker one. A team that splits them first sees that a quarter of its faults have a definite ending available for the price of a connection to a system it already runs. The cheapest large improvement available here came from removing a class of fault rather than from making the component better at avoiding it.
Six of the 23 were restated figures. Which fix removes that group entirely?
Why are the correct outputs the thing that makes the wrong ones dangerous?
Because a check is not performed by a procedure. A check is performed by a person, and what a person does with the four hundredth item in a queue is shaped by what happened with the first three hundred and ninety nine. VigilanceA checker's ability to keep noticing faults when almost everything being checked is correct. is the name for the ability to keep noticing faults when almost nothing is faulty, and it is not a virtue. Vigilance is a resource that a task design either protects or spends.
The household version is familiar enough. A smoke alarm with a weak battery chirps once an hour for a week and then somebody takes the battery out, and the person who took it out is not careless. Removing the battery was the correct response to a very long run of signals that meant nothing. A security guard at a residential gate who has waved through nine thousand residents and stopped nobody is in exactly the same position at the nine thousand and first. The base rateHow often the thing being looked for actually occurs in the population being checked. of the thing being looked for is what determines how hard the looking is, and nobody involved chose that rate.
Here the base rate was 11.5 per cent. A rate that high sounds easy to catch until it is seen from the checker's chair, where it means about eight correct drafts, then one with a problem, then about eight more. Nothing in the eight announces that it is one of the eight. The 177 clean drafts are not the background against which the 23 stand out. The 177 clean drafts are what makes the 23 hard to see.
Two unsupported statements reached a customer. Were those two notes checked less carefully than the others?
What did the two that got through actually look like?
Two letters, and why neither of them looked wrong
Both of the two were decline letters. Both read well. Both carried the correct outcome, the one the scoring component had actually produced. Neither cited a document that did not exist, neither contradicted anything else in the letter, and neither contained a number that was out by a rupee. Each contained one sentence characterising the applicant's circumstances in a way the file did not support.
The person signing spent the same nine minutes on each of those two notes as on the other 198. By the time they reached that sentence they had already read and accepted every sentence before it, and the unsupported sentence read as a summary of what they had just accepted rather than as a new claim. The mechanism is worth holding on to. The obvious remedy does not touch it.
The cost falls on two people who were told something about themselves that their own record never said, in a letter refusing them credit, over the bank's signature. The cost is not an acceptable residual, and no arithmetic on the other 198 makes it one. Nor is it evidence of an inattentive employee, and every fix that follows is structural for exactly that reason: not one of them asks anybody to try harder.
What catches them, and what do the nine minutes buy?
VerificationThe checking a person does on an output before it takes effect or reaches anybody. at this bank is not a general instruction to read carefully. Verification is five numbered checks, and the whole nine minutes is the time those five take. Agrawal, Gans and Goldfarb, in Prediction Machines, 2018, make the general point that a fitted component produces something a person still has to act on; the five checks are what acting on it actually consists of when the something is a paragraph rather than a score.
| The check | What the person does |
|---|---|
| 1. Figures traced | Every number in the draft traced back to the supplied material |
| 2. Assertions traced | Every assertion traced to a passage that carries it |
| 3. Nothing asserted that the material does not contain | Read for what is present but should not be. Check 3 caught the 21 |
| 4. Format and required statements | The required wording and the required statements are all present |
| 5. The decline-to-answer case | Where the component said it could not answer, that case is handled properly |
Notice which of the five is doing the work and why it is the slow one. Checks 1, 2, 4 and 5 all give the reader something to look for: a number, a citation, a required phrase, a flagged case. Check 3 asks the reader to notice something that is not there. Noticing an absence is a much harder cognitive task, and it has no natural stopping point except reaching the end of the draft. Twenty one of the 23 faults were caught by the single check that asks a person to spot an absence, and that is the check every shortened verification procedure drops first.
Did grounding remove them, or only reduce them?
Three changes were made between the two measurements and they cannot be separated after the fact. In month 8 a retrieval arrangement was added, so the component was handed passages from the bank's own document store rather than working from the file alone. In month 9 a sixth part was added to the instruction, telling it what to do when the material does not support an answer, and a connection was added letting it read the file's structured fields instead of restating figures out of prose. In month 10 the second measurement was run.
The result was 7 of 200, being 3.5 per cent, against 11.5 in month 6. The drop is an improvement of 8.0 percentage points, and verification time fell from 9 minutes to 6 because there was less to trace. Eight percentage points is a large, real improvement. Seven of 200 is also not zero, and the bank did not expect zero. GroundingTying an assertion in an output to a passage somebody can open and read. makes a supported sentence more likely rather than making an unsupported one impossible.
Did the grounding changes remove unsupported statements from the drafts?
One part of the improvement can be attributed cleanly and the rest cannot, and saying so is what keeps the whole claim honest. In the first measurement, 31 of the 200 notes restated a figure at all, and 6 of those 31 were wrong, being 19.4 per cent. In the second, 34 notes restated a figure and none was wrong. No change other than the connection touched figures, and that improvement therefore belongs to the connection alone.
The remaining improvement is shared between the retrieval arrangement and the sixth part of the instruction, and the bank never ran the trials that would have separated them. The bank could have run them, and chose speed instead. Choosing speed is defensible and expensive at once. The next time somebody asks which change to make first at another site, there is no answer on file. An improvement that cannot be attributed is still an improvement, but it is not yet knowledge, and the difference shows up the moment somebody tries to repeat it somewhere else.
Why does an improving component make the check harder to sustain?
The least intuitive fact in the whole subject follows from arithmetic rather than from psychology. At an 11.5 per cent rate, one draft in about every nine carries a fault, so the checker meets a problem roughly once a morning. At 3.5 per cent, one draft in about every twenty nine carries a fault, so the checker now reads through nearly three times as long a run of entirely correct drafts before anything is wrong.
Every one of those correct drafts is a small piece of evidence that reading closely is not paying. The evidence is honest. The world really is showing the checker one clean draft after another. And the improvement to the component that produced it is genuine and worth having. So the situation is not that the improvement was bad; it is that the improvement moved the difficulty from the component to the person, and nobody redesigned the person's task at the same time.
Before the control below is moved: the rate falls from 11.5 per cent to 3.5 per cent. Roughly how often does an unsupported statement now turn up?
The vigilance relationship: a better component, a harder check
Two hundred drafts, one measurement round, drawn as a grid. Move the unsupported statement rate and watch how many of the 200 are marked and how long a run of correct drafts sits between two marks. Held constant: the round is always 200 drafts, and the marks are drawn evenly spaced.
At 11.5 per cent, 23 of the 200 drafts carry an unsupported statement, which is one in every 8.7, so the person checking reads a run of about 8 correct drafts between two of them.
What can be done about that, since vigilance cannot be exhorted?
Start by ruling out the response everybody reaches for. Telling the desk to concentrate harder does not work, has never worked, and in this case has an additional problem: the two notes that got through were checked exactly as long as the other 198, so there is no slack in the task to appeal to. Any fix that begins with a person trying harder is addressing a cause that was measured and found absent.
The remaining responses are structural, and they come in three shapes. Reduce what has to be checked, so the check has less surface. Put the evidence beside the claim, so the check gets easier rather than longer. Change who checks and how often, so no single person is asked to hold attention across a long correct run. All three change the task. None of them changes the person.
Vigilance cannot be exhorted. Which of these is a structural response rather than an appeal to effort?
What does the checking cost, and who decides whether it is worth paying?
How an operations head reads these two numbers
Ismail Sheikh runs the exception desk, and the component drafts about 3,010 exception notes a month for him. The measured net saving in month 6 was 5 minutes a note, so 3,010 notes gave back 15,050 minutes, about 250.8 hours. The measured verification was 9 minutes a note, so the same 3,010 notes cost 27,090 minutes, about 451.5 hours. On the bank's own assumed working month of 8,400 minutes a person, that check is 3.23 posts, and at an assumed fully loaded Rs 9,00,000 a year it is about Rs 29,02,500 a year.
Read those two together and the month 6 arrangement is checking itself for roughly 1.8 times what it saves. The 1.8 is not an argument against the arrangement. The saving was never the reason for it, and the check was never optional. The 1.8 is an argument for knowing the figure before signing anything. A business case built on the 5 minutes and silent on the 9 has understated the running cost of the arrangement by more than the whole of its benefit.
By month 10 the same arithmetic turns over. Net saving 8 minutes a note is 24,080 minutes, and verification at 6 minutes is 18,060 minutes, or 2.15 posts and about Rs 19,35,000 a year. The check now costs about three quarters of what the arrangement saves rather than nearly twice, and that reversal, not the fall in the error rate, is the number an operations head actually needs.
All four of those money figures are arithmetic on the bank's own measurements rather than separately measured amounts. Neelima Rao, in the risk function, made the same point from the other side during her review: nobody can say what a further improvement would be worth, so a rate with no cost attached to it cannot be argued about. The bank refused a fine tuning proposal on exactly that basis, comparing a trial that moved the rate by 2.0 percentage points against three grounding changes that moved it by 8.0, and the comparison is the teaching rather than the conclusion.
In month 6 the arrangement saved 15,050 minutes a month and its verification cost 27,090. What does that tell an operations head?
Where the expectations sit
Where a statement an applicant's own record does not support reaches that applicant in a letter from a regulated lender, the expectations that apply sit with the Reserve Bank of India and are published at rbi.org.in. Where the deployer is a market intermediary rather than a lender, the Securities and Exchange Board of India at sebi.gov.in is the relevant body. The current material at those sources governs, and it has to be read there before any control is designed around it.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering outsourcing, digital lending, customer communication and record keeping | rbi.org.in |
| Securities and Exchange Board of India | Expectations where the deployer of such an arrangement is a market intermediary | sebi.gov.in |
| Bank for International Settlements | International supervisory material on banks deploying automated and assisted decision processes | bis.org |
| Agrawal, Gans and Goldfarb | Prediction Machines, 2018, on a fitted component producing something a person still has to act on | Harvard Business Review Press |
Sumeru Bank Limited, Ismail Sheikh and Neelima Rao are invented.
Educational material. Not advice on any investment, tax, budget or market position.
