Robustness and Sensitivity: Whether the Result Survives, and Which Assumption Carries It
What has to be settled before a check like this means anything?
A check like this comes at the end of a long stretch of work. The question has been framed, the mechanism has been separated from the record it produced, somebody else has rerun the analyst's steps and reached the same figure, a reviewer has put six structured questions to it, and the claim has been written so that a reading could refute it. All of that is settled elsewhere. In every real analysis somebody then leans over and says, fine, but what if it had been done slightly differently?
Three things carry the whole of what follows, and it is worth naming them before the arithmetic starts. The first is that a result is almost never one number. A fitted relationship yields a direction, a size, a measure of how much of the movement it accounts for, and a set of gaps between what happened and what it said should have happened. The four travel together and they do not have to move together. An assumption can shove one of them across the room and leave another where it stood. An analyst who has quietly been treating a result as a single figure has been reading one of the four and calling it the answer.
The second is that changing exactly one thing is what makes a move attributableTraceable to one named change rather than to any of several at once.. Changing one thing is not a nicety borrowed from a laboratory. The same discipline governs a panel with one control on it: turn two knobs and the reading moves, and what arrives is a moved reading with no indication of which knob did it.
The third is the data. Everything below runs on ten paired months of an invented input and an invented outcome that earlier work in this subject area already fitted a line to, on a second input built by nudging the first slightly, called here the near duplicateA second input that carries almost the same information as one already in the relationship., and on a count of chairs, invented along with the hall they were set out in, one figure for each of those same ten months. The three sets are used rather than diagnosed again, and the chairs are present because a check which passes cleanly on a count of chairs is the sharpest possible evidence that checks have limits.
What is a robustness check, and what is a sensitivity test?
Start with a shopkeeper. The distinction lives on a counter long before it ever lives in a spreadsheet. Two shopkeepers each close their books for the month and each wonders whether the month was profitable, and each is unsure how to count the rent. The landlord took it late, so it could reasonably sit in either month.
The first shopkeeper works it both ways, finds the month profitable either way, and stops. She has answered one question: does the answer survive? Yes. The second shopkeeper works it both ways, also finds the month profitable either way, and then keeps going. She notices that counting the rent one way leaves her with a comfortable margin and counting it the other way leaves her with almost nothing, and that no other decision she made all month comes close to mattering that much. She has answered a second question: which decision carries this? Both shopkeepers have a profitable month, and only the second one knows what to watch next month.
Survival and ranking are two questions, not two phrasings of one. Whether the result survives has a yes or a no for an answer. Which assumption carries it has a ranking for an answer. A report that says the result is robust to five checks has answered the first and skipped the second entirely, and the second is usually where the information a reader could act on is sitting. The word robust reads like a grade, as though five checks were better than three. Robustness is not a grade. Robustness is a yes or a no, and the count of checks a result survived is not a score.
A report ends with the line: this result is robust to five separate checks. Which of the two questions has that report answered?
Why does only one assumption change at a time?
Five settings run on the same ten paired months. Each one changes exactly one thing and everything else stays where it was. Change nothing at all. The unchanged run is the starting line and belongs on the list. Drop the month with the largest gap between what happened and what the line said. Fit the line that makes the average absolute missThe size of a gap between what happened and what was predicted, counted without regard to direction. as small as possible instead of the one that makes the squared misses as small as possible. Add the near duplicate as a second input. Use only the first five of the ten months.
Every one of those is a decision somebody could defend in a meeting. None of them is a trick. A robustness check built out of settings nobody would have chosen is theatre: it varies things that would never have been done, finds the result standing, and teaches nothing about the choices actually made.
One at a time is not fussiness, and here is the whole of why: two changes in one run produce a move nobody can attribute, and a move nobody can attribute is worth exactly as much as no move at all. Suppose a month were dropped and the near duplicate added in the same run, and the direction of the relationship reversed. Which change did it? The run cannot say, only that something in there reverses it. The two would then have to be separated anyway, a run spent to discover that two runs were needed. Ten shops in one shopping centre share one entrance; when takings fall in all ten, nobody yet knows whether the entrance or the trade did it, and finding out means looking at a shop with its own door.
A colleague, saving time, changes two assumptions in a single run. The answer moves. What have they learned?
Does the result survive, and how many settings does it take to say no?
Here is the whole run, both readings at every setting. Neither of the two readings means much without the other. The slope is the direction and size of the relationship. The fit is the share of the outcome's movement the relationship accounts for.
| What was changed | The slope | The fit |
|---|---|---|
| Nothing changed | 1.5000 | 0.7559 |
| Drop the month with the largest miss | 1.3163 | 0.7988 |
| Fit the smallest average miss instead | 1.5000 | 0.7559 |
| Add the near duplicate input | minus 1.0000 | 0.7560 |
| Use the first five months only | 1.3385 | 0.6351 |
Read the slope column downward. The column reads 1.5000, then 1.3163, then 1.5000, then minus 1.0000, then 1.3385. The direction of the relationship holds at four of the five settings and turns over at one of them. One setting is enough: a result that reverses under a single stated, defensible assumption has not survived, and no amount of arithmetic about the other four changes that. This is where the language does real damage. Four out of five sounds like a pass mark. Four out of five is not a mark at all. The question was whether the direction survives every setting the analyst was willing to state in public, and it did not.
Two of the five settings, the two shaded green above, moved neither reading by anything at all. Changing nothing moves nothing, and the unchanged run belongs on the list for exactly that reason. Fitting the smallest average miss instead lands on the very same line on this data, slope and intercept both. Landing on the same line is a real finding, and it cost a run to get. A setting that finds nothing is what earns the right to say a result is insensitive to it.
Before touching the panel below. Five settings, one of which adds a second input carrying almost the same information as the first. Does the direction of the relationship hold at every setting?
Change one assumption, and watch two readings disagree about how much happened
One control, five stops, and it moves one thing: which assumption is changed. Both bars redraw together. The slope bar is drawn from a fixed line at zero, so crossing that line is the direction turning over. The fit bar is drawn from zero with a dashed marker at the starting 0.7559. The strip underneath ranks the three settings that move something, once by the slope and once by the fit, and the current setting lights up in both rankings.
Educational illustration on invented data. Each stop changes exactly one assumption from the same starting point, and a stop that moves nothing is a result rather than a blank.
Two of the five settings moved neither reading. When the check is written up, should those two appear in it?
What does a Sensitivity Test rank, and which assumption comes first?
Now the second question, and it is a different exercise on the same table. Three of the five settings moved something. Take the size of the move each one produced, ignore its direction, and put the three in order, largest first. The resulting orderingThe ranking of several things by how much they move a stated quantity, largest first. is the answer to which assumption carries the result, and it is the whole content of a sensitivity test. Notice that it does not have a yes or a no anywhere in it.
| The setting | Move in the slope | Rank | Move in the fit | Rank |
|---|---|---|---|---|
| Add the near duplicate input | 2.5000 | 1 | 0.0001 | 3 |
| Drop the month with the largest miss | 0.1837 | 2 | 0.0429 | 2 |
| Use the first five months only | 0.1615 | 3 | 0.1208 | 1 |
The two orderings are exact reverses of each other. Whatever the slope says matters most, the fit says matters least, and they are looking at the same three settings on the same ten months at the same moment. The middle setting is the only thing the two readings agree about, and they agree about it because there is nowhere else for a middle to go.
An exact reversal of the two orderings is what makes quoting one column alone indefensible. An analyst who ran this check and quoted the fit would have written that the assumption this result leans on hardest is the length of the record, and that adding a second input barely touches it. An analyst who quoted the slope would have written the opposite, word for word. Neither of those is a mistake in arithmetic. Both are correct readings of their own column. The mistake is the idea that either column, on its own, is the result.
A prediction before looking back at the table. Three settings move something. Do the slope and the fit rank those three the same way?
How can one change move a coefficient and leave the fit alone?
Take the setting that produced the reversal and look at it properly. Adding the near duplicate moves the slope from 1.5000 to minus 1.0000, a move of 2.5000 which is larger than the slope it started from, and the very same change moves the fit by 0.0001. The move of 0.0001 is not a rounded stand in for something bigger: the fit goes from 0.7559 to 0.7560 and that is the whole of it. No single fitted valueWhat a fitted relationship says an observation should have been, as against what it actually was. moves by more than 0.1250 either.
The reason is that the two inputs carry almost the same information, correlating at 0.999967, so the relationship as a whole is left exactly where it was while the share of it attributed to each input becomes arbitrary. Picture two people carrying one table through a doorway. How much of the weight each of them is taking can swing wildly from one end of the corridor to the other, and can even end up with one of them pushing down while the other lifts harder, without the table rising a single inch higher or lower. Anyone watching the table sees nothing. Anyone watching the two people sees everything.
The fit watches the table. The fit is a statement about the outcome as a whole, and about the outcome as a whole nothing happened. The coefficients watch the two people, and by that measure the answer inverted. So both readings belong at every setting, and neither is ever quoted alone. A single column is not a shortened version of the check. A single column is a different check, answering a question nobody asked.
Under one and the same change, the fit moved by 0.0001 while the slope moved by 2.5000. What does the tiny move in the fit indicate?
What can a robustness check never see?
Now the limit, and it is not a caution. The limit is arithmetic, and it is why the count of chairs has been present from the beginning.
A split half checkRefitting on the first half of a record and on the second half separately, to see whether the two halves agree with each other. is about as respectable as a robustness check gets. Cut the record in two, measure the relationship on each half on its own, and see whether the halves tell the same story. A relationship that holds in the first half and vanishes in the second was probably never there. Run on both of the relationships carried through this guide, it yields four readings. Calling one of the two the real relationship separates the one with a mechanism behind it from the one with nothing behind it, and nothing more than that.
| The relationship | First five months | Last five months | The gap |
|---|---|---|---|
| The real relationship, input against outcome | 0.7969 | 0.9863 | 0.1893 |
| The count of chairs against the same outcome | 0.9308 | 0.9322 | 0.0014 |
The invented count of chairs passes this check far more cleanly than the real relationship does, and it is not close: the chairs' two halves agree to 0.0014 where the real relationship's agree only to 0.1893. Nothing connects a hall's chairs to that outcome. There is no mechanism at all. And the check does not merely fail to catch it: running the check harder makes the chairs look better, not worse. The trap works like this. The analyst reaches for a tougher standard, the honest relationship struggles against it and the invented one sails through, and the afternoon ends with more confidence in the wrong thing than it began with.
The reason sits in the definition. A robustness check varies an assumption and asks whether the result holds. A missing mechanism is not an assumption. A missing mechanism is not a dial on the panel, not a row that can be dropped and not a criterion that can be swapped, so there is no setting at which it changes and nothing for a check to vary. The question gets answered somewhere else entirely, by somebody stating in plain words why one of these things should move the other, with no number in the sentence.
Precision has to survive the carrying as well as the arithmetic. Every quantity here was held at full precision and rounded exactly once, on its way to the screen. Take the two halves of the real relationship as printed, 0.9863 and 0.7969, and subtract one from the other: the answer is 0.1894, where the gap worked at full precision is 0.1893. Rounding both sides and subtracting afterwards moves the answer in the fourth place. A fit that moves by 0.0001 lives in that same fourth place.
On the split half check, the chair count's two halves agree to 0.0014 and the real relationship's agree only to 0.1893. Which of the two looks the sounder on this check?
A result passes every check anyone can think of, at every setting that can be defended. What has still not been established?
How is a robustness check reported so that it means something?
Almost nobody reads the check. A lender, an analyst or a household reads the paragraph about the check, so the reporting rules matter more than the technique, and there are five of them.
- Name every setting tried, including the ones that moved nothing. The two settings above that moved neither reading are not filler. They are the only evidence that the result is insensitive to those two things, and they are also the only proof that nobody went looking for settings that behaved.
- Report both readings at every setting. Not the one that behaved. The reading that misbehaved is the one a reader needs, and it is the one almost every write up drops.
- Give the ranking as well as the verdict. The verdict is a word. The ranking tells a reader which assumption to argue with, which is the only thing a reader can actually do with the work.
- Say which single setting would have to be wrong for the conclusion to change. Here that sentence writes itself: the conclusion turns entirely on whether the second input belongs in the relationship at all.
- Say what was not varied. This is the line everybody leaves out, and it is where the next argument starts. Nothing above varied who decided which months counted as usable, and nothing above varied the choice of these ten months over some other ten.
A report that names only the checks which passed is a selection rather than a check, and the difference is invisible to every reader who was not in the room. The same reader who would never accept an average without knowing how many things it was divided among will accept the words robust to five checks without asking which five.
In writing up a robustness check, which line does almost every write up leave out altogether?
Where does this go wrong, and what does it cost?
An analyst runs all five settings, writes them up honestly, and reports that the fit stayed between 0.6351 and 0.7988 throughout, a spread of about a tenth from end to end, with the check that disturbed it least moving it by a single ten thousandth. The conclusion filed is that the result is robust.
Every figure in that report is correct. The fit really did stay in that range. The setting that barely moved it really did move it by 0.0001, and the 2.5000 that came with that same setting appears nowhere in the fit column. Nothing has been rounded kindly, nothing has been dropped, and a reviewer recomputing the whole thing would reconcile to the last place and sign it off.
The cost is that the setting which moved the fit least is the one that reversed the answer. Adding the near duplicate takes the slope from 1.5000 to minus 1.0000, a move of 2.5000, and the fit registers 0.0001 of it. So the report is arithmetically true and its conclusion is backwards, and every reader who was handed the fit column will agree with it, confidently, for the rest of the quarter.
There is a second cost underneath the first. The count of chairs passes the split half check at 0.9308 and 0.9322 where the real relationship manages 0.7969 and 0.9863, so a careful analyst who responds to doubt by checking harder ends up with the invented count looking better than the real thing. Checking harder was the wrong lever.
The fix is a reporting rule and not a technique: print both readings at every setting, publish the ranking beside the verdict, and end with the line naming what was never varied. None of that requires a single extra computation. All of it requires giving up the sentence that says the result is robust.
Covered elsewhere. Five things are settled before a check like this begins: deciding in advance which reading would change a mind, telling a generator apart from the record it threw off, handing a set of steps to somebody else and watching them reach the same figure, putting a fixed set of challenges to finished work, and writing a claim in a shape an observation could knock over. All five are covered earlier. Refitting a rule on rolling stretches of history, and choosing between two competing fitted relationships, are covered separately.
What would have to be wrong for any of these figures to move?
Every figure above was computed from the same twenty invented readings, so a figure moves only if one of those readings, or one of the rules applied to them, is different. The right hand column below names, for each figure, exactly which of the two it would have to be.
| The figure | How it was arrived at | What would have to be wrong for it to move |
|---|---|---|
| The line on the ten paired months, slope 1.5000 with intercept 0.5000 | Fitted to twenty invented readings, every one of them recomputed here rather than carried across | One of those twenty readings would have to be a different number |
| Both readings at each of the five settings | Each refitted from those same twenty readings by the rule named in its own row | A setting would have to change more than the one thing its row names |
| The move of 2.5000 in the slope beside the move of 0.0001 in the fit | Taken at full precision from the two settings either side of it, then rounded once | The second input would have to carry information the first one does not |
| The two halves, 0.7969 and 0.9863 against 0.9308 and 0.9322 | Four correlations worked separately, on the first five months and on the last five | The split would have to fall somewhere other than halfway |
| Anything drawn from outside those twenty readings | There is none | Nothing, because there is nothing outside them to be wrong about |
The ten paired months, the outcome fitted to them, the near duplicate input and the monthly count of chairs in a hall nobody has ever walked into are invented.
Educational material. Not advice on any investment, tax, budget or market position.
