Optical Character Recognition and Data Extraction: From Image to Field
Optical character recognition recovers characters from an image. Data extraction assembles those characters into named fields and validates them. A field is right only if every character in it is right, so a character reading of 99.2 per cent gives a twelve character field a 90.8 per cent chance of being right. Field length, not the reading step, sets the result.
Two numbers about a reading chain sound like the same measurement and are not. One is a statement about characters. The other is a statement about whole fields, and it is the first one compounded once for every character the field holds. Nothing about the components explains the distance between them. The length of what is being read explains all of it, and once that is clear, a great many confident reports about reading systems stop being reassuring.
Because it is the mistake most often made here, one warning belongs before any number arrives. Sumeru Bank Limited, an invented bank, reports that 95.0 per cent of its fields cleared its own confidence setting in a month. The 95.0 per cent is the share of fields accepted above a bar, and it is not a field accuracy at all. Accepted and correct are two different statements about a value, they are produced by two different mechanisms, and the difference is worked in full below. The two stay apart from here on.
What happens between an image arriving and a field being populated?
The situation is a familiar one. A relative sends a photograph of an electricity bill on a messaging app, and asks for the amount due. The photograph was taken at an angle in poor light, so the picture gets opened, squinted at, the phone turned, perhaps pinched to zoom. The printing is read. Then the part of the bill that carries the amount has to be found. Finding it is not the same act as reading the printing. The bill is covered in numbers, and only one of them is the answer. Last comes the judgement of whether what has been read makes sense as an amount, and if the reading gives a figure with a stray mark in the middle, the looking starts again.
Reading a bill by hand is five acts, and a document reading chain performs the same five in the same order. Only the third of them is optical character recognition, and three of the four ways this step fails happen outside it. The five are: acquire the image; assess and improve it; recover the characters; locate the field on the document; assemble and validate the value. Optical character recognitionRecovering the characters in an image so they can be read as text. is step three and nothing else. The reading step does not choose the image, it does not decide where a value begins, and it does not judge whether the result is sensible.
Step two deserves a moment because it is the step people forget exists. Before a single character is recovered, the image is scored and cleaned: straightened where it was photographed at an angle, cropped, contrast lifted, and given an image quality indexThe bank's own score for an image, which its fitted relationship turns into a count of fields readable.. At Sumeru Bank Limited the index is the bank's own score for how much the reading step is being asked to work with. A poor score at step two is not corrected by anything at step three. It is inherited.
Which of the five steps is character recognition?
What does the character reading actually measure?
Character accuracyThe share of individual characters recovered correctly. is the share of individual characters the step recovers correctly. At Sumeru Bank Limited the reading step recovers 99.2 per cent of individual characters correctly. The 99.2 per cent is the bank's own measurement on its own document supply in one month. Turned around, it says something concrete: out of every thousand characters it reads, about eight come back as something other than what was printed. A five becomes a six. An O becomes a zero. A one becomes a seven where the printing was thin.
Eight in a thousand is a genuinely good reading, and the step is doing what it was measured to do. The trouble is not that 99.2 per cent is wrong, it is that 99.2 per cent is an answer to a question nobody in the room was actually asking. Nobody wants to know about characters. A character is not stored anywhere, nobody makes a decision on one, and no part of the chain downstream has a place to put one. The chain stores fields, and a field is what a person eventually reads.
What does data extraction assert, and why is its measure a different one?
Data extractionAssembling recovered characters into named fields and validating each one. is the work of steps four and five taken together. Handed a document full of recovered characters, data extraction decides where one named value begins and ends, pulls the characters between those two points, assembles them into a value, and then checks that value against whatever the chain expects of it. Out of a wall of text it produces a small number of named things: the employer name is this, the net pay is this, the account number is this. At this bank it produces 14 such fields on every file, and across the month's 8,600 completed files that is 120,400 fields.
The measure that belongs to that work is field accuracyThe share of whole fields that are entirely correct, which compounds with length.: the share of whole fields that come back entirely correct. One rule decides everything that follows. A field is not partly right. An account number with one wrong digit is not a slightly imperfect account number, it is a wrong account number, and it will fail to match, or worse, it will match something else. A field is right only when every single character in it is right, so field accuracy is character accuracy multiplied by itself once for every character the field holds.
field accuracy = 0.992 raised to the power of the number of characters
| 0.992 | the bank's own measured character reading, being 99.2 per cent, held fixed throughout |
| the power | how many characters the field holds, counted including spaces and separators |
| the result | the share of fields of that length that come back entirely correct |
Why does a reading of 99.2 per cent not survive into a longer field?
A household chore has the same shape. Ten children have to get onto one school bus, and for each child the chance that everything they need has been remembered is 99 in 100. Very good. But the morning only goes right if all ten are fine. The chance of that is 99 in 100 multiplied by itself ten times, or about 90 in 100. One morning in ten goes wrong, and no individual child was the problem. Compounding is what turns a very good per-item rate into an ordinary per-outcome rate, and the only thing that decides how ordinary is how many items the outcome depends on.
Applied to the fields at this bank, and holding the character reading at 99.2 per cent throughout, the readings are these. A 6 character field comes back entirely right 95.3 per cent of the time. A 12 character field, 90.8 per cent. A 20 character field, 85.2 per cent. A 30 character field, 78.6 per cent. Read that last one against the headline: a step that recovers more than 99 characters in every 100 correctly returns a completely correct thirty character value only about four times in five.
| Characters in the field | Field accuracy at 99.2 per cent a character | Fields wrong if all 120,400 were this length |
|---|---|---|
| 1 | 99.2 per cent | 963 |
| 6 | 95.3 per cent | 5,665 |
| 8 | 93.8 per cent | 7,493 |
| 12 | 90.8 per cent | 11,063 |
| 20 | 85.2 per cent | 17,868 |
| 30 | 78.6 per cent | 25,781 |
The right hand column is a what-if and not a count of anything that happened. The column answers one question only: if every one of the month's 120,400 fields happened to be that length, how many would come back wrong? The bank's fields are of many different lengths, so no single row of that column is the month. Length alone moves the count of wrong fields from 963 to 25,781.
Before the control below moves: the reading is right on 99.2 per cent of characters. What share of 20 character fields comes back entirely right?
Stretch one field, and watch a hundred fields change colour
One input moves: how many characters the field holds, from 1 to 30. Three things redraw together. The field at the top grows a box for every character. The marker rides down the curve. And the hundred squares at the bottom recolour, so the fields of that length that come back entirely right can be counted one by one. The character reading is held at 99.2 per cent the whole way, so nothing that moves is a change in the step.
Static readings to check the control against. The character reading is 99.2 per cent throughout and the month holds 120,400 fields. At the default of 12 characters a field comes back entirely right 90.8 per cent of the time, and if every field in the month were 12 characters long, 11,063 would be wrong. At 6 characters the reading is 95.3 per cent. At 20 characters it is 85.2 per cent.
At 12 characters a field comes back entirely right 90.8 per cent of the time, so 91 of every 100 such fields are correct and 9 are not, and if every one of the month's 120,400 fields were 12 characters long, 11,063 would be wrong.
Which fields suffer most, and what decides that?
A room asked which of the fourteen fields worries it will name the hard ones: the handwritten bits, the smudged bits, the employer letter that arrived as a photograph of a screen. Reasonable, and not the answer. Ranking the fourteen fields by character count ranks them by field accuracy, whatever anybody thinks about how hard each one looks. Difficulty affects the character reading. Length affects how many times that reading has to hold.
No report the bank produced ever separated its fields by length, and that omission is what the sign-off below turned on. So the lengths below are ordinary lengths for values of that sort rather than the bank's own measurement, and they settle the ordering rather than any count. The ordering is the part that does not depend on the lengths being exactly right: the short codes sit at the top of the list at more than 95 per cent, and the long free text values, an account number or an employer name, sit at the bottom, and no plausible set of lengths reverses that.
Two fields are read by the same step. One holds 6 characters and one holds 20. Which one needs a person more often?
What does the image itself do to the number of fields that come back?
Everything so far has held the image constant. Now let the image move. In practice the image is where most of the trouble at this bank actually arrived. The bank scores every image it receives and relates that score to how much of the file comes back readable using its own fitted relationship: the share readable is 0.50 plus 0.0062 times the index, capped at 1.00, and multiplied by the 14 fields on the file, then rounded down to a whole field. The relationship is the bank's own, fitted on the bank's own images, and it is not a general property of anything.
The bank's median image quality index is 74. Put 74 through the relationship: 0.50 plus 0.0062 times 74 is 0.9588, times 14 is 13.4232, rounded down to 13. So on a file of median image quality, 13 of the 14 fields come back above the bank's acceptance barThe confidence level above which a field is taken without a person, chosen by this bank at 0.92. and one does not. Drop the index to 52 and the same relationship returns 11.5136. Rounded down that is 11, so three fields stop instead of one. Lift it to 81 and the relationship reaches its cap: every one of the 14 comes back, and improving the image beyond that point buys nothing at all on this file.
812 of the 1,264 stopped files carried a low quality image. What does that say about where to spend?
How does one file's reading fit with the month's figures?
Two readings are now in play and they look as though they disagree, so settle that before either is used for anything. The per-file relationship says a median file returns 13 of 14, or about one stopped field per file. The month says 6,020 fields were routed to a person, and those 6,020 sat in only 1,264 files, an average of 4.76 fields each. One stopped field a file against 4.76 in far fewer files. One reading describes a single file at a given index and the other is an aggregate over a month of files that mostly had none, so the two are not in conflict and must never be merged.
Work it through in both directions and it settles cleanly. Taken continuously rather than rounded to whole fields, the relationship at the median index gives 95.88 per cent of fields readable. On 120,400 fields that would be about 4,960 routed. Rounded up to whole files, if every one of the 8,600 files sat exactly at the median with one stopped field each, that would be 8,600 routed fields spread across 8,600 files. The month sits between the two on count, at 6,020, and nowhere near either on spread: those 6,020 fields touched 1,264 files, or 14.7 per cent of the month rather than all of it. The concentration is clusteringFailures sitting together in a few files rather than spreading evenly across many., and it is the reason an exception desk is very much smaller than a field level error rate makes it sound.
The fitted relationship gives about one routed field on a median file, and the month has 4.76 routed fields per stopped file. Is that a contradiction?
Why is the month's 95.0 per cent not a field accuracy at all?
Now the figure quoted at the outset, in full. Of the month's 120,400 fields, 114,380 cleared the bank's own chosen acceptance bar of 0.92 and were taken without a person looking. The 114,380 accepted are 95.0 per cent of the month. The remaining 6,020 were routed. The 95.0 per cent is reported everywhere in the bank, and it is often the only figure anyone quotes about the reading step. Being the only figure quoted is exactly why it deserves this much care. The figure answers the question how many fields were accepted, and it does not answer the question how many fields were right.
The distinction is a familiar one. A shopkeeper accepting a hundred rupee note without holding it to the light has accepted it. Whether it was genuine is a separate matter settled by a separate mechanism, and the acceptance rate describes the shopkeeper rather than the notes. Here the separate mechanism is a hand check: 400 accepted fields were pulled and read by a person, and 7 of them were wrong. Seven in 400 is 1.75 per cent of accepted fields. Across the 114,380 accepted in the month that projects to about 2,002 wrong values sitting quietly inside the accepted pile.
The month shows 95.0 per cent of fields accepted. Is the reading step right 95.0 per cent of the time?
What is the month's implied field accuracy, and why is it only a floor?
With the two numbers held apart, the month's field accuracy can finally be built, and it is one line of subtraction. Start with all 120,400 fields. Take out the roughly 2,002 that were accepted and are wrong, the number the hand check implies. Take out all 6,020 that were routed to a person. Left over are 112,378 fields that were both accepted and right, or 93.3 per cent of the month.
The subtraction treats every one of the 6,020 routed fields as wrong, and most of them were not wrong at all: they were merely unreadable. So the 93.3 per cent is a floor rather than a reading. A field the step declined to guess at is not an error in the value; it is an absence of a value. The true field accuracy is therefore somewhere above 93.3 per cent, and the honest treatment of a number that can only be bounded is to say which side it has been bounded from. Called a floor, it is defensible. Called the answer, it quietly reports the step as worse than it was.
One more reading falls out of the floor, and it closes the circle with the figures at the outset. If a character reading of 99.2 per cent produces a field accuracy of 93.3 per cent, what average field length would do that? About 8.6 characters answers it. Multiply 0.992 by itself 8.6 times and the result is 93.3 per cent. The 8.6 is arithmetic on the bank's own figures rather than a measurement of its fields, so it is a consistency check rather than a count. The month behaves as though its typical field is between eight and nine characters long. For a set of fields holding dates, codes, amounts, an account number and an employer name, eight or nine characters is entirely plausible.
Why is the month's implied field accuracy of about 93.3 per cent described as a floor?
The error that gets made, and what it cost here
The reading step at Sumeru Bank Limited was signed off on a character accuracy of 99.2 per cent. The figure was correctly measured, correctly reported and correctly minuted. The next four seconds in a room full of competent people did the damage. Nobody said the word character out loud, so everybody heard a system that is right 99.2 per cent of the time.
A chain right 99.2 per cent of the time is not what they had. At an average field length of about 8.6 characters the same 99.2 per cent gives a field accuracy of about 93.3 per cent, a floor built from one line of subtraction on the month's own numbers, and the difference between those two readings across 120,400 fields is thousands of fields a month. Expectations about how much would stop, and therefore how many people the exception desk needed, were carried out of that room on the character figure. The business case sat at 2 posts of exception work. The desk that steady state actually required was 7.
Nobody misreported anything, and the reading step did not underperform. A measure was carried across a boundary it does not cross, and the boundary is that characters compound into fields. Where the damage landed is entirely predictable in hindsight: on the long fields, an account number or an employer name, the very fields no report ever separated out. Nothing was mismeasured, so no measurement discipline catches this. One question asked out loud before the sign-off catches it: this figure is a share of what?
What are the four ways this step fails, and how is each one caught?
A single accuracy figure conceals four quite different events, and they are caught by four different things, cost four different amounts and are fixed by four different teams. Setting them out separately is the most useful half hour anybody running such a chain can spend. Three of the four happen outside the character recognition step entirely, and so the character figure can be excellent while the chain is not.
The third of them deserves particular attention because nothing at character level can ever see it. Suppose the characters are all recovered perfectly and the field boundaryWhere a value starts and stops on the document, which can be wrong even when the characters are right. is drawn in the wrong place, so the value assembled runs one character short or swallows a stray mark from the line above. The result is a well formed, confident, wrong value. Every character in it was read correctly. By its own definition that event was a perfect success, so a character accuracy measure scores it as one and is right to. The fourth kind runs the other way: the value is entirely correct and a validation ruleA written check that a value must pass, such as a format, a range or a match against another record. refuses it because the format was unusual, and a person is asked to confirm something that was never wrong.
The characters were recovered correctly and the field is still wrong. Which failure kind is that?
What has to be recorded so the four failures stay separable?
A chain that records one accuracy figure per month makes all four of those events look identical afterwards, and the cost of that is not theoretical: it is somebody spending a quarter improving the reading step when 64.2 per cent of the stopped files were carrying an inherited image problem. The record has to keep the steps apart in the same way the diagram does.
Six things, and none of them is expensive to keep. The image as received, unaltered, alongside the improved version actually read. The image quality index at the moment of reading. A file that failed at index 41 and a file that failed at index 79 are two different problems. The characters recovered, with the confidence value attached to each field. The boundary used for each field, so a boundary fault can be told apart from a character fault. The validation outcome, separately from the reading outcome. And the person intervention, saying what was changed and what was merely confirmed. Keep those six and every one of the four failure kinds can be counted separately after the fact; keep one accuracy figure instead and none of them can.
Where the accountability sits
Where document images supporting a lending decision are read and retained by a regulated lender, the expectations covering record keeping, outsourcing, customer data and consent sit with the Reserve Bank of India, and its material is published at rbi.org.in. The acceptance bar of 0.92 and the image quality bar are the bank's own choices rather than a standard of any kind.
How does a lender, an analyst or an operations head actually use this?
The use is a single question, asked before any money moves, and it takes about a minute. Somebody presents an accuracy figure for a reading step. The question is whether that is a share of characters or a share of fields. If the answer is characters, the next request is for the average number of characters in the fields the chain actually stores, and the character figure is raised to that power. At this bank that is 0.992 to the power of about 8.6, giving about 93.3 per cent. The number people were budgeting against was 99.2, so the conversation changes shape immediately.
Then two follow-ups. First, how many fields a month, so the percentage becomes a count of exceptions and the count becomes desk minutes. At this bank the exception desk runs 7 posts, and at the fully loaded Rs 9,00,000/- a year for one post that the bank assumes, that desk is about Rs 63,00,000/- a year, arithmetic on the assumed figure rather than a payroll reading. Set that beside the Rs 65,00,000/- a year this chain costs to run and the sizing question stops being an operations detail. Second, do the failures cluster? Here 6,020 stopped fields sat in 1,264 files, so the desk works 1,264 items and not 6,020, and a chain whose failures spread evenly would have needed a very different desk on identical field level numbers.
Ajay Agrawal, Joshua Gans and Avi Goldfarb, in Prediction Machines, 2018, put the general form of it: a component of this sort produces a prediction, and its value is set entirely by what somebody does differently as a result. Here the prediction is a field value, and what somebody does differently is either accept it or go and look at the document. A character figure cannot say how often somebody has to go and look, and that is the only quantity the desk, the budget and the customer waiting for a decision actually feel. Cathy O'Neil, in Weapons of Math Destruction, 2016, makes the companion point that a model's errors rarely fall evenly, and this case is a small clean instance of it: the errors here fall on whoever has the longest employer name.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering digital lending, outsourcing, customer data, consent and the retention of records supporting a credit decision, which apply to the lender whatever reads the documents behind that decision | rbi.org.in |
| Ajay Agrawal, Joshua Gans and Avi Goldfarb | Prediction Machines, 2018, for the value of a prediction being set entirely by what somebody does differently as a result | Harvard Business Review Press |
| Cathy O'Neil | Weapons of Math Destruction, 2016, for a model's errors falling unevenly rather than spreading across everybody alike | Crown |
Sumeru Bank Limited is invented.
Educational material. Not advice on any investment, tax, budget or market position.
