Digital Identity: The Layers, Liveness and the Error Nobody Counts
A digital identity is four separate checks that everyday speech collapses into one word. Identification asks who a person claims to be. Verification asks whether that claim matches a real record and belongs to the person presenting it. Authentication asks whether the person coming back is the one who was verified. Authorisation asks what that person may now do. The first two happen once. The third happens every single time.
All four rest on something that is easy to say and hard to hold on to. A biometric check does not answer yes or no. The check compares two images and returns a number saying how alike they are, and somebody at the deploying institution then chooses the bar that number is read against. Every property of the check that an applicant actually experiences, including how often a real person is turned away, follows from where that chosen bar sits rather than from anything intrinsic to the component. The bar is a decision. Decisions have owners, and this one usually does not have a name attached to it.
What is a digital identity actually made of?
All four checks happen at a wedding, and nobody there calls them by any name. Start at the gate. A man arrives at the gate on the first evening and says he is the bride's uncle from Nagpur. The claim is the whole of the first check: he has said who he is, and nothing has been tested. The person at the gate then looks at the printed invitation in his hand, and glances at the bride's mother, who nods. The invitation and the nod together are the second check, and notice the check has two halves rather than one: the invitation is real, and it belongs to this man rather than to whoever else might be carrying one.
By the third evening the same man walks in and the same person at the gate waves him through without looking at anything. Nobody has re-read the invitation. The third evening is a different check entirely: the gate recognised him as the man it admitted on the first evening. Then he wanders towards the room where the jewellery is being kept, and somebody stops him. Being a recognised guest and being allowed into that particular room are two separate permissions, and nobody had confused the two until he tried.
The four moments at that gate are the four layers of a digital identity, and the only thing a bank adds is that each one is now carried out by a system rather than by a person. Sumeru Bank Limited, an invented lender, runs all four inside a retail loan chain that begins on an applicant's handset. Numbering the four layers lets any one of them be pointed at later without ambiguity:
Layer 1 is identificationThe claim a person makes about who they are, before anything at all has been checked.. A claim costs nothing to make, so layer 1 is free. Layer 2 is verificationChecking that a claimed identity matches a real record and that it belongs to the person presenting it., and it is the expensive one. Verification has to reach outside the applicant to find evidence, and reaching outside costs money. Layer 3 is authenticationChecking that a returning person is the same one who was verified at the start., and it has to be cheap. Layer 3 runs thousands of times for one customer, and a price paid thousands of times is a different kind of price. Layer 4 is authorisationDeciding what an authenticated person is permitted to do, instruction by instruction., and it is the one most people never notice is a separate question at all.
Which of the four layers runs on every single return, and which ones run only once?
Why does the timing column decide how each layer is built?
Because a check that runs once and a check that runs ten thousand times cannot be the same check, however similar the question sounds said out loud. The household version is familiar. Getting a new gas connection involves a stack of documents, a visit and a wait, and it happens once. Opening a front door with a key takes two seconds and happens four times a day. Nobody would accept the gas connection process as a way of getting into a house in the evening, and nobody would accept a house key as proof of identity at the gas office. The two are not a strong version and a weak version of one thing. The two checks answer different questions, and they are priced accordingly.
A bank meets the same arithmetic with larger numbers. One applicant is verified once. The same applicant, if they take the loan, will sign in to check a balance, change a mandate or pull a statement for the next four years. Verification is a cost the institution pays once per person and authentication is a cost it pays per visit, so a design that confuses the two is not slightly wrong, it is wrong by a multiple of however many times the customer comes back. A multiple that size turns a small design fault into a permanent one.
What goes wrong when verification and authentication are merged?
Identity Verification vs Authentication, and why one cannot stand in for the other
Merging the two is the commonest fault in a front door, and it is worth seeing why it is so easy to commit. Both checks look like the sentence "make sure this is really them". The difference is what each one compares against. Verification compares a claim against the world outside the applicant: a record held by somebody else, an attribute at its source, a face against a document. Authentication compares a person against what this institution already stored about them at verification. One reaches outward, once. The other reaches inward, repeatedly, and for years.
Collapsing them produces one of two failures. The uncomfortable part is that they are opposite failures, and a firm usually has one of them without ever noticing that the other is the same fault wearing different clothes.
The two failures are not a trade between security and convenience, they are one design fault producing opposite symptoms, and the reason a firm rarely spots it is that only one of the two symptoms generates a complaint. The first failure makes customers cross. Cross customers produce calls, and calls produce a project to fix it. The second failure makes nobody cross at all. The person it admits is delighted, and the person it robs does not know yet. The asymmetry between a fault that complains and a fault that stays quiet returns in a much sharper form below, where one of two errors is counted every month and the other is never counted at all.
A system asks a returning customer to upload their identity document every single time they sign in. Which layer is it running, and which one has it skipped?
What are the ways an identity can be verified, and are they equally strong?
Identity Verification Methods, the five of them, and what each one actually proves
There are five ways in ordinary use to satisfy layer 2. No official list contains exactly these five, and the useful way to read them is not by how modern each one sounds. Read them by one question only: what would somebody pretending to be another person have to get hold of in order to pass this?
Methods 1 to 4 all require an impostor to obtain something genuinely difficult to obtain: a convincing artefact, a way into a record at its source, a body, or the nerve to sit in front of a person and be asked follow-up questions. Method 5 requires only information about the person being impersonated, and information about a person circulates entirely without that person taking part.
Why the knowledge check is the weakest of the five
A past address is not a secret. A past address sat on an electricity bill, a delivery slip, a school form and a rental agreement, and every one of those passed through the hands of somebody else. The amount of a person's last loan sat in a statement, a message and a call centre note. A knowledge check treats information as if it were possession, and information is the one thing on this list that can be in two places at once. The argument ends there. A harder question is simply a rarer fact rather than a different kind of thing, so harder questions do not make a knowledge check any stronger.
None of that makes method 5 useless. Method 5 is a top-up rather than a foundation. A firm that uses it alongside a credentialAnything a person holds or knows that a system will accept as standing for them, such as a password, a one time code or a card. the person also holds is doing something reasonable. A firm that uses it as the whole of layer 2 has built a front door that opens for whoever did the reading.
A firm verifies identity by asking a caller for a past address and the amount of their last loan. How strong is that, and why?
The three things layer 3 can ask for
Authentication runs on a shorter list, and it is worth holding in mind because every sign in anybody has ever done is one of these three or a combination of them. Something known, such as a password or an answer. Something held, such as a handset receiving a one time code or a card in a machine. Something inherent, such as a face or a fingerprint. The reason a serious front door asks for two of the three rather than two of the same is that two things from the same column fail together: two known things both leak in the same breach, and asking for a second password buys almost nothing.
| What layer 3 asks for | Everyday instance | How it is usually lost |
|---|---|---|
| Something known | A password, a passcode, an answer to a question | It leaks, and it leaks in bulk rather than one at a time |
| Something held | A handset receiving a code, a card, a key | It is lost, borrowed, or handed over willingly to somebody persuasive |
| Something inherent | A face, a fingerprint, a voice | It cannot be lost, and it also cannot be changed once copied |
The third row holds the trade nobody mentions, and it is the one to read slowly. A password that leaks can be replaced by five o'clock. A face that is copied is copied for good. The permanence of a copied face is why a face check deserves a closer look than a password does, and it is why a firm should be slower to add a biometric than to add a code.
What does a biometric check actually return?
A biometricA measurement of the body used as an identifier, such as a face, a fingerprint or a voice. comparison does not return matched or not matched. The comparison returns a similarity scoreThe number a biometric comparison returns, saying how alike two images are, which somebody must then read against a chosen bar., and a similarity score is a number saying how alike the two images are. Nothing in the component decides what that number means. Somebody at the bank draws a line across the range of possible scores and says everything on this side is accepted and everything on that side is refused.
Picture two spreads of scores rather than one. Feed the check a pile of comparisons from genuine applicants and the scores bunch high. Not all of them do, and a genuine comparison in bad conditions can score low. Feed the same check comparisons from everybody else and the scores bunch low. Again not all of them do, and two faces can resemble each other closely. The two spreads overlap. The two spreads always overlap. And the line somebody draws sits inside that overlap, wherever they put it.
The one property of this picture that never goes away is the overlap, so the only question a deployer is actually settling when they place the bar is which of two errors they would rather make. Push the line right and fewer impostors get through and more genuine applicants are refused. Push it left and the reverse. There is no position that removes both, and any supplier or team that describes their setting as removing both has described the picture wrongly rather than built a better one.
A team says their face check returns a match or no match, and nothing else. What have they left out of that description?
Where does the liveness check sit in a real chain, and what is it for?
An image of a face is a face as far as pixels are concerned, so a photograph can fool a biometric comparison. So a second test is bolted on beside it. A liveness checkA test on whether the image presented is of a person present at that moment rather than a photograph, a screen or a recording. asks whether what was submitted came from a person present at that moment rather than from a photograph, a screen or a recording. A liveness check does not ask who the person is. A liveness check asks what kind of thing the camera was pointed at.
At Sumeru the check is component 2 of the intake chain, and the chain's onboardingThe sequence of steps a new customer completes before a relationship can begin, from first contact to submission. path runs over six numbered steps on an applicant's handset. Component 2 sits at step 3, and where it sits is the whole reason it matters as much as it does.
Step 3 is early. A check that keeps an impostor out is worth more before the bank has spent effort on the file than after, so the early placement is deliberate and sensible. The cost of putting a check early is that everything downstream of it inherits its errors, and an applicant refused at step 3 never appears in any number describing steps 4, 5 or 6. A refused applicant is not a rejected applicant. A refused applicant is simply not there any more, and the difference matters enormously to what the bank's reports can see.
What are the two ways a liveness check can be wrong?
Exactly two, and they are numbered here so that either can be pointed at later. Error 1 is a genuine applicant refused: a real person, applying for a real loan, turned away. Error 2 is an impostor accepted: somebody who should not have got through, getting through. Every biometric check that has ever run makes both, at some rate, and where the bar sits decides the mix.
Now ask a duller question than which error is worse. Ask which error a bank can see. Error 1 announces itself the same day. An application visibly stops, and a counter somewhere increments. Error 2 announces itself only if something later goes wrong, months on, and only if somebody connects that later problem back to a check that ran at step 3 in month 6. Very often it announces itself never.
Sumeru had a count for error 1 and no count whatsoever for error 2, and that is not a gap in the reporting, it is a property of the two errors themselves. The bank never had a number for error 2, and a plausible invented figure would be worse than an absence. The absence is the finding.
Consider what that does to a management report. The 620 refusals appeared every month under a heading about fraud stopped at the front door, and the number was read as evidence that the check was earning its keep. Nobody was lying. A refusal count is a real count of a real thing. A refusal count still cannot say whether the check is set correctly. Two numbers are needed for that, and the bank holds only one of them.
Which of the two errors is easy for a firm to count, and what follows from that?
620 applicants were refused by the liveness check in month 6, and 200 of those refusals were later read by hand. How many of the 200 turned out to be genuine applicants?
What did reading two hundred refusals by hand find?
Neelima Rao, in the risk function, who did not build any part of the chain, pulled 200 of the month's 620 refusals and had each one looked at. Thirty-one of the 200 were genuine applicants. Thirty-one out of 200 is 15.5 per cent of what was read, and applying that share to the whole 620 puts about 96 genuine applicants inside the month's refusal count.
The important thing about that review is not the number it produced. The review was work somebody chose to do. No report generated it, no dashboard flagged it, and no threshold was breached to trigger it. The counted error came free with the system and the finding that mattered had to be paid for in somebody's time. A difference in price is the whole reason one of them shaped the bank's view of the check for six months and the other did not exist until month 6.
What did all thirty one of them have in common?
One condition, shared by every single one of the 31. Not most of them and not a tendency: all of them. Each had submitted an image taken in low light on a low-specification handset.
Read that carefully. There is a wrong way to read it, and the wrong reading costs people a loan application. The shared condition is not a statement about how carefully those applicants took a photograph. The shared condition is a statement about two objects: a room with the light it had, and a handset with the camera it came with. An applicant photographing themselves in the light available to them, on the equipment they could afford, has done every single thing the process asked of them, and what failed was a step built to work in one set of conditions and then offered as the only way in. The condition belongs to the equipment and to the room. The condition never belonged to the person.
One condition shared by all 31 also says something structural about where this error lands. The error is not sprinkled evenly over applicants like a small tax. The error concentrates, and it concentrates on the same people every month. The same people keep the same handset and stand in the same room. Cathy O'Neil describes the same pattern in Weapons of Math Destruction, and the group it picks out at Sumeru is defined by what an applicant's handset cost.
All 31 wrongly refused applicants had a low-light image on a low-specification handset. What does that say about where the error falls?
Put the 96 next to the number of people who started rather than next to the number who were refused, because that is the comparison an applicant would make. About 96 people a month out of the 10,000 who started an application is about 1.0 per cent of everybody who came to the front door, told no before the decision engine ever saw their file. Not declined on their income, not declined on their record. Removed at step 3, over the light in a room.
Move the share, and watch the people appear one square at a time
One control: the share of the month's 620 refusals that are genuine applicants, from 0 to 30 per cent. One consequence: how many people that is, drawn as one square each, with the desk time a second route would need shown underneath. The default is the measured reading of 15.5 per cent, or 31 genuine applicants found in 200 refusals read by hand. At that setting the grid shades about 96 squares, or about 1.0 per cent of the 10,000 people who started an application, and the desk time is 1,152 minutes a month, or 0.14 of one post.
Share of the 620 refusals that are genuine applicants: 15.5 per cent
Educational illustration. 620 refusals a month is a measured figure at that bank; the 15.5 per cent comes from 200 refusals read by hand rather than from all 620; the 12 minutes a case at an assisted live session and the 8,400 minute working month are assumptions.
What does a second route cost, and what does it restore?
Something is wrong at step 3. Fixing it is a different problem from finding it, and the two obvious fixes both fail.
The instinct is to go to the bar and move it, or to go shopping for a component that is better in low light. Both are worse answers than they look. Moving the bar trades error 1 against error 2, and Sumeru has no count at all for error 2, so moving it is trading a number that exists against a number that does not. Replacing the component leaves a different overlap in the same shape, a new bar somebody still has to place, and a fresh set of conditions in which it happens to be weak, discovered in about six months.
A route out changes the outcome for the people the check turns away without touching the check at all, so the response that works is not a better component at step 3 but a second way through step 3. The applicant in month 6 did not meet a strict test. The applicant met a closed door with one button on it.
The review recommended four steps, and the striking thing about all four is that Sumeru already had every one of them working somewhere else in the bank. Not one of the four had to be built from nothing. The second route is a set of existing parts, connected so that the closed screen becomes a door.
Price it. About 96 people a month at an assumed 12 minutes each is 1,152 minutes of desk time, or 0.14 of an assumed 8,400 minute working month. At the bank's assumed fully loaded cost of Rs 9,00,000/- a post a year, that is roughly Rs 1,23,000/- a year. The figure is arithmetic on two assumed numbers rather than a measured cost. Set it beside the Rs 65,00,000/- a year the whole intake chain costs to run and the second route is about 1.9 per cent of it. The build was six working days, or 42 hours of work.
For roughly one fiftieth of what the chain costs to run, about 96 people a month stop being a refusal and become an application again, and the check at step 3 is left exactly as it was. Leaving the check exactly as it was is the point. Nobody has to agree about where the bar belongs, nobody has to find a number for error 2, and nobody has to argue that the component is bad. The route is cheap precisely because it settles none of those arguments.
What does a second route change about the 620, and what does it leave completely untouched?
The error that gets made, and what it costs
Sumeru had one number and treated it as two. The 620 refusals were real, and they were reported every month as fraud stopped at the front door, and that reading was accepted for six months because nothing in any report contradicted it. Nothing could have. A report is built from what a system records, and the system recorded refusals because refusals happen in front of it. The system recorded nothing about impostors accepted because there was nothing to record.
The cost of the error was about 96 people a month, being about 1.0 per cent of everybody who started, told no at step 3 with nowhere to go. The condition they shared was the light in the room they were standing in and the camera on the handset they could afford, and not one of them did anything the process had not asked for. Indifference did not make it last six months. The only number anybody had said the check was working, and the number that said otherwise did not exist until somebody spent a week making it exist.
How a lender, an analyst or an applicant actually uses this
A lender's operations head reads a front door by asking four questions in order, and they take about ten minutes. Which of the four layers is this step, in the numbering above? Does the check return a score or a verdict, and who chose the bar? How large is the count for each of the two errors, and if one of them has no count, say so out loud rather than treating the missing one as zero. And what happens to somebody this step refuses: is there a second route, or is there a button that says close?
An analyst comparing two lenders reads a refusal rate the same way. One firm refusing 6.2 per cent at a liveness step and another refusing 2.0 per cent says nothing on its own about which is tighter. The two numbers could equally reflect two different bars, two different applicant populations, or two different handsets in two different price brackets. A completion rate is never read alone; it is read beside what happened to the people it does not count.
And an applicant refused at a step like this is entitled to a plainer thing than either of those: a way to finish that does not depend on the light in their room. A second route is exactly that, and it costs less than anything else at the front door.
Where these expectations are published
The requirement to identify a customer before opening a relationship is set by the Reserve Bank of India at rbi.org.in, and is a separate subject from the digital mechanism described here. Where the institution deploying such a check is a market intermediary rather than a lender, the equivalent expectations sit with the Securities and Exchange Board of India at sebi.gov.in. Instructions authenticated at layer 3 and layer 4 commonly travel on payment rails operated by the National Payments Corporation of India at npci.org.in. Thresholds, requirements and effective dates in these instruments change, and an institution is bound by whatever is in force on the day it acts.
Covered separately. The Reserve Bank of India sets what a regulated lender must check before opening a relationship, and that requirement is covered separately. How a learned component is fitted or evaluated is established elsewhere. Digital onboarding as a whole path, and where applicants are lost along the rest of it, is set out under digital onboarding.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Expectations on a regulated lender covering customer identification, digital lending and record keeping | rbi.org.in |
| Securities and Exchange Board of India | Equivalent expectations where the institution deploying such a check is a market intermediary | sebi.gov.in |
| National Payments Corporation of India | Material on the payment rails that authenticated customer instructions travel over | npci.org.in |
| Cathy O'Neil | Weapons of Math Destruction, on how a system's errors concentrate on an identifiable group rather than scattering across a population | Crown, 2016 |
Sumeru Bank Limited, Revathi Balan, Ismail Sheikh, Neelima Rao and Ashok Pillai are invented.
Educational material. Not advice on any investment, tax, budget or market position.
