Adversarial Attacks: Manipulating a Model Deliberately
Deliberate manipulation is not drift. Somebody is choosing the input in order to produce an outcome. One invented bank saw four kinds in the period: shaping a declared figure to clear a written line, showing a false image to a learned check, altering a document after it was issued, and putting an instruction inside text a component reads. The cheapest kind is the one aimed at the rule the bank published.
Underneath those four kinds sits one asymmetry, and it decides how each kind can be defended. A rule somebody wrote down can be read, and anything that can be read can be worked backwards. A component fitted to data cannot be read, so it cannot be worked backwards, but it also cannot be argued with or explained. So the readable part of a deployed chain attracts cheap manipulation in volume, the unreadable part attracts rare and expensive manipulation, and no single defence covers both. Naming which part is under pressure is therefore the first decision, not a detail of the write-up.
What is the difference between drift and somebody choosing the input on purpose?
Picture the weighing scale at a vegetable stall. Over two years the spring inside it tires, and the scale starts reading forty grams heavy on everything. Nobody decided that. Get the scale recalibrated and the problem is gone for good, and it stays gone. Now picture a different stall where the scale is perfect and the vendor rests a thumb on the pan. Take the thumb away and the reading is correct. The person wanted a particular reading and still wants it, so come back next week and something else is holding the pan down.
The thumb on the pan is the whole distinction, and the distinction survives translation into a deployed system without losing anything. Drift is a change in what arrives, with no intention behind it, and it is repaired by correcting the input or refitting the component. An adversarial inputAn input chosen to produce a particular outcome rather than to describe a fact. is chosen by a person who wants a specific outcome. The person adapts to whatever defence is put in the way, so neither correcting the input nor refitting the component repairs it. The practical consequence is that a fix which ends drift permanently only relocates manipulation.
Drift and deliberate shaping look identical in a monitoring pack. At Sumeru Bank Limited, invented, the retail loan intake chain produced a movement in its approval rate that nobody had asked for, and the pack showed a rate that had moved and a referral count that had risen. A movement of that shape can come from a changed input field, from a shifting mix of applicants, or from a group of people who have worked out where a line sits. The numbers on the front of the pack do not say which, and the response to each is different enough that guessing is expensive.
What separates deliberate manipulation from drift?
What kinds of deliberate manipulation did one deployed chain actually see?
Sumeru Bank Limited recorded four kinds across the period, and the useful thing about the record is that it carries counts rather than adjectives. Kind 1 is shaping a declared figure so that it lands on the permitted side of a written line. Kind 2 is presenting a false image to a learned check. Kind 3 is altering a document after it was issued. Kind 4 is putting an instruction inside a passage of text that a component reads. Each is aimed at a different part of the chain and each was stopped, or not stopped, by a different part of it.
The counts run in the opposite direction to the alarm. Kind 1 shows up as an excess of about 280 files a month. Kinds 2, 3 and 4 together produced 16 attempts in the whole period: 11, 4 and 1. The gap between the ordinary case and the dramatic ones is two orders of magnitude, and both halves of the comparison are worth saying. Deliberate manipulation of a learned component is real, it happened, and somebody had to deal with it. Manipulation of a learned component is also rare. The monthly cost to this bank comes from a person rounding a number up to sit under a published line.
Kind 1 shows an excess of 280 files a month, and kinds 2, 3 and 4 together produced 16 attempts. What is the safe thing to say about those two numbers?
How does somebody shape a declared figure to clear a written line?
Start with the shop near a bus stop that offers free delivery on any order above Rs 500/-. Nobody has to be told what that does to the order book. A customer with a Rs 470/- basket adds a packet of biscuits rather than pay for delivery. So orders do not spread evenly across Rs 400/- to Rs 600/-; they pile up at Rs 505/- and Rs 510/-. Not one of those customers has done anything wrong. The shop published a line and people arranged themselves around it. People arrange themselves around every published line.
Sumeru Bank Limited runs a written income rule inside its intake chain. The rule compares the income an applicant declares against the median salary credit corroborated from three months of statement, and it sends the file to a person when the declared figure sits more than the bank's own chosen toleranceThe gap a firm allows between a declared figure and a corroborated one before it routes the file to a person. of 10 per cent above the corroborated one. A file declaring Rs 45,000/- against a corroborated Rs 38,000/- carries a gap of Rs 7,000/-. A gap that size is comfortably more than a tenth of the corroborated figure, so the file is routed. The 10 per cent tolerance is this invented bank's own choice, not a standard, a norm or a requirement of anybody.
Now put the two together. The bank has to be able to tell an applicant why the file was routed, so a rule of that kind is not a secret and cannot practically be one. So the line is knowable, and gamingShaping a declared figure so that it lands on the permitted side of a written line. it requires no skill, no equipment and no understanding of anything technical. Gaming a written line requires arithmetic. The cheapest manipulation available against a deployed chain is not an attack on the clever part; it is a declared figure adjusted to sit just inside a line the firm published itself.
946 files a month at this bank declare an income somewhere between five and ten per cent over the corroborated figure. How many of those 946 sit in the single percentage point immediately under the line?
What does shaping look like in the numbers, before anybody has confessed?
Nobody at Sumeru Bank Limited ever received an admission. The bank has instead the shape of its own population, and the shape is enough to say something specific. The bank's tolerance sweep records 1,548 files routed a month if the allowed gap were set at 5 per cent and 602 routed at 10 per cent. The 946 files sitting between those two counts declare an income above the corroborated figure by something more than five per cent and less than ten, and at the actual setting every one of them clears the rule.
The 946 files are not spread evenly across that stretch. 412 of them sit in the single percentage point immediately below the line, against about 132 that the bank's own smooth reference would put there. The four percentage-point bands further away hold 534 between them against an expected 814. The pile of files against the line is what bunchingMore cases sitting just inside a line than a smooth spread would put there. looks like, and it is visible in figures the bank already holds without asking anybody a single question.
Two cautions belong in the same breath as that finding. The smooth reference is the bank's own construction and not a measurement: it is a statement about the shape the band would have had if nobody had known where the line sat, and the whole excess depends on it. And the split of the 534 across those four bands was never recorded, only the total, so the four bands appear below as a single block.
How much of that band was always going to be there?
One question decides whether the finding is analysis or accusation, and the question has an arithmetic answer. The band holds 946 files. The smooth reference puts about 132 in the point below the line and 814 across the four bands further away, and the recorded counts are 412 and 534. So 412 less 132 is 280, and 814 less 534 is also 280. The same 280 files that are missing from the body of the band are the ones sitting hard against the line. Shaping does exactly that: it does not create files, it moves them.
The reconciliation is worth carrying because it puts a boundary on the claim. The excessThe count in a band above what a smooth spread would predict, which is the part attributable to shaping. is 280 files a month. The excess is 29.6 per cent of the 946 in the band, and 3.3 per cent of the 8,600 files the chain decided that month. The other 666 files in the band are people whose declared income happens to sit near a line they may never have heard of, and there is nothing in the arithmetic that distinguishes any individual file. The finding is about a population and it stays about a population.
412 files sit in the point just under the line. How many of them is the bank entitled to treat as shaped?
How is a learned check attacked, and what stopped the attempts?
Kind 2 is aimed at a component nobody can read. The liveness checkA learned component deciding whether an image is of a person present at that moment rather than a picture of one. at the front door of the intake chain judges whether the image it is given belongs to a person who is actually there. Eleven attempts in the period presented it with a photograph of a screen instead. Attacking a learned check is a different proposition from adjusting a declared figure. Nobody can state the line, so there is no line to arrange yourself around, and whatever the component has learned about images is not written down in a form a person could work backwards from.
The record of what happened is short. Nine of the eleven were stopped by that component itself, and two got past it and were stopped by a person. The split of nine and two is a useful shape rather than a comforting one. The learned check did most of the work, and the arrangement only held because a person was standing behind it. Agrawal, Gans and Goldfarb, in Prediction Machines, 2018, are the names attached to the framing that a fitted component produces a prediction while somebody else still has to act on it, and two of eleven is precisely where that framing earns its keep.
The everyday version is the security guard who is shown a picture of an identity card on a phone screen rather than the card, and who has to decide whether a picture of a thing counts as the thing. The guard is doing something no written rule handles well. The judgement was given to a learned component for that reason, and for the same reason the component cannot explain what it did.
Who found the altered documents, and was it the component they were shown to?
Kind 3 is altering a document after it was issued and putting the altered version in front of the chain. Four files in the period carried one. The four are worth a block of their own for who found them rather than for the count. All four were caught by a downstream checkA later step that catches something an earlier component let through. further along the chain, and none by the document classifier they were actually presented to. The classifier sorted them, passed them along, and something later in the sequence disagreed with what they contained.
Put that beside kind 2 and a pattern appears that ought to change how money is spent. Across the 16 attempts on the three kinds aimed at components that cannot be read, the component under attack was the one that caught it in 9 cases, and something else entirely caught the other 7. The defence in this chain is a property of the sequence rather than of any single component, so hardening the component that was aimed at is not obviously where the next rupee should go. A design in which every component is assumed to defend itself, with no later check disagreeing with anything, has nothing at all to fall back on.
All four altered documents were found by a later check rather than by the classifier they were put in front of. What follows from that?
What is the fourth kind, and why does a single instance prove so little?
Kind 4 is putting an instruction inside a passage of text that a component reads. The component then treats the instruction as something it should carry out. Sumeru Bank Limited saw one instance in the period. One instance is a count and never a rate, and it supports neither a division nor a projection forward.
Kind 4 belongs in the list of four, and a single instance cannot carry an argument about defence economics or a shape about bunching. The single instance does illustrate the same finding as kind 3: the component that received the text was not what stopped it. A check sitting after the draft was written did that, and the check had been built and funded for an entirely different purpose.
How many instances of an instruction placed inside text did this bank see in the period?
Why does the readable part of a chain attract the cheap manipulation?
Put the four kinds on two axes and the shape of the whole subject appears in one look. The vertical axis is how often this bank saw each kind. The horizontal axis is what it takes to attempt: effort, equipment, and the risk of being caught doing it. Arithmetic is free and the line was published, so kind 1 sits at the top left. Each of kinds 2, 3 and 4 requires somebody to prepare something and to present it in person or in a file that will be looked at, so all three sit far lower.
The reason is not that fitted components are stronger. A rule somebody wrote can be read, and reading is the cheap step. The property that lets a bank tell an applicant exactly why a file was routed is the same property that lets somebody else work out where the line is, and no version of that rule has one without the other. Explainability has an uncomfortable half, and the uncomfortable half is why a firm cannot buy its way out of kind 1 by making the rule better.
Notice that kind 4 breaks the neat story, and it is worth letting it. Putting an instruction inside text is cheap, and it still happened once. Cost is not the only thing setting the count; the opening has to exist and be reachable as well, and at this bank that opening was narrow. So the honest reading is that cheapness is necessary for volume and not sufficient for it.
Which part of a deployed chain attracts the cheap, high volume manipulation?
Why is hiding the written rule the wrong defence?
The instinct after seeing a bunching chart is immediate: stop publishing the tolerance. Concealment fails twice over, and both failures are worth stating separately because they are different kinds of failure.
The first is that it does not work. A line can be found by trying. Somebody routed at one declared figure and accepted at another has learned roughly where the line sits without anybody telling them, and a line inferred from experience is exactly as usable as a line published in a leaflet. Concealment removes the statement of the line, and the line itself carries on doing what it did.
The second is that it takes something real away from people who were doing nothing. In a steady month at Sumeru Bank Limited, 1,290 files reach an outcome the applicant would call a refusal. 602 of them were routed by the written income rule, and behind each of those sits a printed procedure that says what happened and why. The other 688 were declined outright by the scoring model, and behind those sits an attributed reasonA reason stated for an outcome that was produced by a component whose working cannot be read. with no procedure at all. The 688 are 53.3 per cent of the 1,290. Concealing the tolerance moves the 602 into the position of the 688, so the bank would end up unable to give a procedural reason to anybody.
The arithmetic already set out says who pays for that. 946 files a month sit in the band and only about 280 are the excess, so roughly seven of every ten people who would lose the explanation were not shaping anything. O'Neil, in Weapons of Math Destruction, 2016, is the name attached to the observation that a model's errors do not land evenly on the people it is applied to, and a defence whose whole cost falls on the majority who did nothing is a clean instance of it.
What a reader has to confirm at source
Where a lender in India must state a reason for an outcome an applicant would call a refusal, the expectations sit with the Reserve Bank of India and are published at rbi.org.in. The standing discipline of model risk work that sits behind independent challenge originates in material from the Bank for International Settlements at bis.org. Material from that source is an origin rather than the position in India. The 10 per cent tolerance is the invented bank's own choice, and so is the smooth reference the excess is measured against. Requirements, thresholds and effective dates move, and the current position at the source governs.
A firm proposes to stop publishing its tolerance. What has it bought, and what has it spent?
How is manipulation told apart from drift when only the outputs are available?
Most of the time the first thing anybody sees is a rate that moved, and nothing else. Three checks separate the two causes using figures a firm already holds. None of the three requires knowing what anybody intended, and that independence from intention is what makes them usable.
The first check asks whether the movement is concentrated at a line or spread across the population. Shaping aims at a specific line, so cases pile up against that line. A change in what arrives touches every part of the population, so drift moves the whole shape. The second check asks whether anything upstream changed its format, its source or its mix in the same window. If something did, drift is the ordinary explanation and is the one to test first. The third check asks whether the pile follows the line when the line is moved. A population that has arranged itself around a line will rearrange itself around the new one; a distribution that has genuinely shifted will not care where the line goes.
No single check settles it, and a reviewer acts when two of the three point the same way. Two of the three is a lower standard of proof than a court would want and a higher one than a monitoring pack alone can give. The two-of-three standard gets used because the response to each cause is different enough that being wrong is expensive in both directions.
An approval rate moves and the input fields all look normal. Manipulation or drift?
What does a person running or reviewing a deployed chain do with all this?
Ismail Sheikh runs the exception desk at Sumeru Bank Limited, and what reaches him is a queue of files rather than a chart. For him the practical output is a triage question he can actually ask: when the queue grows, is it growing evenly across the six causes it records, or is one cause growing on its own? A queue growing evenly is a volume story or a drift story. A queue where files routed by one written rule are growing while nothing else moves is the shape that ought to prompt somebody to plot the band.
Neelima Rao sits in the risk function and did not build any part of the chain, and her use of it is different. When she reviews a written rule set she is not only reading it for lines that contradict each other; she is reading it as somebody on the other side would read it, asking which lines are cheap to arrange yourself around and whether the bank has ever plotted the population near each of them. A rule set that has never been plotted against its own lines has an unmeasured exposure rather than no exposure.
For anybody applying for anything, the reading is simpler and worth stating plainly. A firm that publishes the line it uses is giving the applicant a reason that can be checked and argued with, and that is worth more to an ordinary applicant than the small amount of shaping the published line invites. When a firm stops saying why, the thing it has protected is a number, and the thing it has taken is the applicant's ability to know what happened to their own file.
The error that gets made, and what it costs
The mistake is not seeing the bunching. The mistake happens in the next meeting, when the chart of 412 against about 132 is presented and somebody says the bank has found 412 people gaming the income rule. Saying so converts an arithmetic result about a population into an allegation about individuals, and every file in the band is a person who applied for a loan.
The finding is 280 files a month, and it is a count above a constructed reference rather than a list of anybody. 946 sit in the band, so about seven in ten of the people who would be caught by any action aimed at the band did nothing at all. And the reference itself is the bank's own construction, so the 280 moves if somebody argues successfully about what shape the band should have had.
The cost of getting this wrong is paid in two places. The first is by the people whose files are treated as suspect on the strength of where they happened to sit. The second is paid by the bank. A finding stated as an accusation cannot survive being challenged, and a finding that cannot survive challenge is one nobody will fund a fix for.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Published expectations on a regulated lender covering digital lending, outsourcing, data and consent, and the reason a lender states for an outcome an applicant would call a refusal | rbi.org.in |
| Bank for International Settlements | International supervisory material on model risk, from which the standing discipline of independent challenge originates | bis.org |
| Agrawal, Gans and Goldfarb | Prediction Machines, 2018. A fitted component produces a prediction that somebody else still has to act on | Harvard Business Review Press |
| O'Neil | Weapons of Math Destruction, 2016. The errors of a deployed system fall unevenly across the people it is applied to | Crown |
Sumeru Bank Limited, its intake chain, Ismail Sheikh and Neelima Rao are invented, and so is every figure attached to them.
Educational material. Not advice on any investment, tax, budget or market position.
