The Research Hypothesis: Making a Claim That Can Fail
Here is a sentence that appears somewhere, in some form, in a great deal of written work. This study investigates whether the outcome is related to the input. The sentence sounds like the start of real work, careful and even modest. And it is not a claim at all. No result could be put in front of the person who wrote it that would make them say, out loud, that they were wrong. The whole difference between a hypothesis and a topic is that somebody can lose a hypothesis, and the losing is arranged in advance rather than argued about afterwards.
The same thing happens outside any research setting. Two people at a wedding argue about whether the caterer undercharged. One of them says the food was worth more than what was billed. The other says it was not. Neither of them has said what would settle it, so the argument can go around all evening. Now one of them says: the caterer served four hundred plates and billed under Rs 250/- a plate, and if the plate count turns out to be under three hundred then I am wrong. The plate count sentence can be lost. Somebody can go and count.
Several earlier subjects are assumed here, and none of them is taught again. How long a record has to be before it can settle anything, what a mechanism is against the record it produced, whether a second person can reach the same number from the same data, and how finished work gets challenged, are all covered separately and are taken as given here. Only one thing is at issue below: the wording of the claim itself.
The fifty month record and the rule behind it. Somebody wrote down a generatorA rule written out in full that says which values come out and how often each one does. Because it is written down, its own average is known before any record is drawn from it.: five possible monthly values, each with a weight, the weights adding to exactly one. The rule has an average of 1.00 per cent and a spread of 5.00 per cent, and both are known only because somebody wrote the rule out. Fifty months were then drawn from it. The fifty months average 0.50 per cent, with a standard errorA measure of how far a figure taken from a record is likely to sit from the figure the record was drawn from. It shrinks as the record gets longer. of 0.7035 per cent, and a 95 per cent intervalA range reported around an estimate, saying which values the record is consistent with rather than which single value it produced. running from minus 0.8788 per cent to 1.8788 per cent. Every one of those readings was worked out separately and is used here rather than rederived.
The ten paired months, and the two things laid against them. Separately, and on a different and much shorter record, ten months of an input and an outcome sit side by side. A straight line through them has a fitA single number between zero and one saying how much of the outcome's variation the line accounts for. Nearer one means the line tracks the outcome more closely. of 0.7559. An invented count of chairs put out each month in a hall, averaging 60 chairs and connected to nothing, reads 0.9029 against the same ten months. And a near duplicate of the input, nudged by five hundredths in eight of the ten, brings the fit to 0.7560. Ten months is not fifty, and nothing from one record is ever added to the other.
The rule for calling a claim refuted, written before anything below. A claim here says the true mean is at least some stated figure. The fifty month record refutes such a claim when that stated figure sits outside its 95 per cent interval, and leaves the claim standing when the figure sits inside. The convention is a decision about how to count, not a discovery, so it travels beside every verdict below.
What makes a claim a hypothesis rather than a topic?
A claim that can be tested has three parts, and most written claims have the first two. The first part is a quantity a record can actually produce. Not an impression, not a mood, but something with a value: the average monthly change over fifty months, say. The second part is a stated value or direction for that quantity. Not larger, not meaningful, but a figure: at least 1.00 per cent. The two parts feel like a claim, and they are what most people write down and stop at.
The third part is the one that turns the first two into something testable, and it is the part almost everybody leaves out: the reading that would end the claim. Written in advance, before the record is opened, and written specifically enough that a second person reading the sentence would write down the same reading. Here that third part reads: the claim is over when 1.00 per cent sits outside the record's 95 per cent interval. Very little judgement is left in that. Somebody works out the interval, looks at where 1.00 per cent falls, and the answer is not up for discussion.
The everyday version is a street food stall two streets away. The owner is thinking about moving the stall to a new pitch outside a college gate. Saying the new pitch is better is a position: on a slow day at the new pitch he will say the college was on holiday, and on a good day at the old one he will say a wedding party came past. Nothing settles it. Now he says something else. He says the new pitch takes more than Rs 4,000/- a day on average, and that if the first full month there averages under Rs 4,000/- a day he goes back to the old pitch. The stall owner can lose that claim, and the losing was arranged before the first day of trading.
A colleague offers this claim for testing: the average monthly change is positive. What is missing before anybody can test it at all?
What does a claim that cannot fail look like?
A claim that cannot fail is easier to see than to describe. Take a real one, put things in front of it, and count how many of them it turns away. So here is the claim, in the exact wording an unhurried person might write: the outcome is related to something. Three candidatesThe things being tried against a claim, one at a time, before any of them has been chosen or ruled out. were tried against it, all three read on the same ten paired months, and none of them was picked in advance to make a point.
The claim is that the outcome is related to something. Three candidates get tried against it: the ordered input, a count of chairs put out each month in a hall, and a near duplicate of the input. How many of the three refute the claim?
The ordered input reads 0.7559. The count of chairs reads 0.9029. The near duplicate reads 0.7560. All three satisfy the claim, and the number of candidates that refute it is zero. The count of refuting candidates is the figure worth sitting with. Not one, not a small number: zero. There was no observation available anywhere in this exercise that could have gone against the sentence. Running the work was never going to change what anybody believed.
Notice what the count of chairs does to the argument. The count of chairs is not a weak candidate that squeaked through. A count of furniture put out in a hall, connected to the outcome by nothing whatsoever, reads higher than the ordered input on the very measure the claim was checked with. If a claim is satisfied by that, the trouble is not that the claim is weak. The set of things that would have refuted the sentence is empty, so the trouble is that the sentence is not a claim. A weak claim is one a record could go against and probably will not. A record could not go against this one at all.
A colleague writes down that the input matters. Which rewriting turns it into a claim a reading could actually end?
How is a claim stated so that a reading can refute it?
Write the claim as a stated floor. Not the outcome is related to something, but the true mean is at least some figure, with the figure written out. Then fix the conventionA choice about how something will be counted, written down in advance so that a second person can repeat it and reach the same verdict. for refuting it, in the same breath, before anything is opened. The convention used throughout is one sentence long: this record refutes a claim that the true mean is at least some stated figure when that figure sits outside the record's 95 per cent interval. The interval runs from minus 0.8788 per cent to 1.8788 per cent, and it was fixed by the length of the record and the spread inside it, not by anybody's preference.
Six claims run down that convention give the following verdicts. Not one of the six verdicts is a judgement call. Somebody with the interval in front of them and no opinion about the subject would produce the same six answers.
| The claim, as written | Does the stated floor sit inside the interval? | This record says | And the claim is |
|---|---|---|---|
| the true mean is at least 0.00 per cent | Yes, inside | survives | true |
| the true mean is at least 0.50 per cent | Yes, inside | survives | true |
| the true mean is at least 1.00 per cent | Yes, inside | survives | true |
| the true mean is at least 1.50 per cent | Yes, inside | survives | false |
| the true mean is at least 2.00 per cent | No, it sits above the upper end | refuted | false |
| the true mean is at least 2.50 per cent | No, it sits above the upper end | refuted | false |
The verdict turns exactly once, at 1.8788 per cent, and it never turns back. Four claims survive and two are refuted, and the place where the answer changes was settled by the record's length before a single claim had been written down. The turning point is worth holding on to. Nobody chose 1.8788 per cent. The figure fell out of fifty months and the spread inside them, and it would have sat where it sits whichever six claims anybody had decided to line up against it.
One more thing about the ladder, and it is hygiene rather than a finding. The nearest claim on the ladder, the one stating at least 2.00 per cent, sits 0.1212 per cent away from the turning point. Nothing on this ladder is decided by a hair. If a claim had been placed at, say, 1.88 per cent, its verdict would turn on the fourth decimal place of a figure nobody can measure that finely, and the honest thing then would be to say the record cannot separate that claim from its neighbour rather than to print a verdict.
The record's 95 per cent interval runs from minus 0.8788 per cent to 1.8788 per cent. Before anything in the panel below is moved, is a claim that the true mean is at least 2.50 per cent refuted, or left standing?
Slide the claim's floor and watch the verdict flip at a point nobody chose
One control, one consequence. The band is the record's 95 per cent interval and it never moves. The record is already written, and nothing done here can lengthen it. The claim's stated floor is what moves. As the floor slides through the six settings, two things happen at once: the verdict flips, exactly once, at a place fixed before any claim existed; and the grid underneath fills in. One cell of that grid stays empty on this record, and one fills up in a way that should give pause.
Educational illustration. The fifty month record and the generator behind it are inventions built for teaching, and every figure in this panel is illustrative. The band is the record's own 95 per cent interval, worked from a mean of 0.50 per cent and a standard error of 0.7035 per cent, and it is the same band at every setting because the record does not change when the claim does. The true mean of 1.00 per cent is knowable here only because the generator was written down first. No real record offers that luxury. The convention for calling a claim refuted was fixed before any of the six claims was written, and nothing in this panel decides it.
Does surviving a test count as evidence for the claim?
Look again at two of the six claims, slowly. The claim that the true mean is at least 1.00 per cent survives, and it is true: the generator's mean is exactly 1.00 per cent, and somebody wrote that generator down. The claim that the true mean is at least 1.50 per cent also survives, and it is false, for the same reason and by the same figure. A true claim and a false claim came through the same record with the same verdict, so the verdict separated nothing.
Most readers of published research get the next sentence wrong, and it is worth putting bluntly. Surviving is not a small amount of evidence. Surviving is not weak support, not a hint, not a promising sign that firms up when somebody runs it again. On this record surviving is a statement about where a claim's floor sits relative to 1.8788 per cent, and 1.8788 per cent is a fact about how many months were collected. A claim at 1.50 per cent survived because fifty months cannot see the difference between 1.00 per cent and 1.50 per cent, not because there is anything to be said for 1.50 per cent.
Set the six claims out against two questions instead of one. Is the claim true? Only the generator can answer that. Does this record refute the claim? Only the convention and the interval can answer that. Two questions give four cells to fill. Only three of the four fill up on this record, and the empty one is the cell where a true claim gets refuted. Everything true about this generator sits comfortably inside the band, so nothing lands in that cell at any setting on the ladder. The record never makes the loud, obvious mistake. The record makes the quiet one instead.
The quiet one is the cell holding the claim of at least 1.50 per cent: false, and standing. A survived claim cannot be reported as a supported claim, and that one cell is why. Had 1.50 per cent been written down before the record was opened, the verdict now in hand would read exactly like the verdict a correct claim gets. Nothing in the arithmetic distinguishes them, and nothing in a longer write up would either.
A claim of at least 1.00 per cent survives on this record, and so does a claim of at least 1.50 per cent. Exactly one of the two is true. What does surviving actually say about either of them?
What does failing to reject a claim of nothing actually mean?
A second kind of claim turns up constantly, and it is the flattest one available: the true mean is zero, and nothing is going on. Run against the fifty month record under the usual arrangement, it produces a reading of 0.7107, with a two sidedCounting a departure in either direction as evidence against a claim, rather than only a departure one way. chance of about 0.4772. So the claim of nothing is not rejected. A great deal of careless writing begins with that sentence.
The record's reading against a mean of zero comes out at 0.7107. Before anything else is examined, is the two sided chance attached to that reading likely to come out large or small?
Failing to reject a claim of nothing does not say the mean is zero, and on this record it is flatly not zero: the generator's mean is 1.00 per cent, positive, and written down before a single month was drawn. The failure says something narrower and much less interesting: fifty months cannot tell 1.00 per cent apart from zero. The statement is about the length of the record, and the design knew it before the record was opened. How far apart two values have to be before a record of a given length can separate them is arithmetic available in advance.
Here is the household version. A shopkeeper weighs a sack of rice on a bathroom scale that reads to the nearest kilogram, and it shows 50 kilograms both before and after a customer takes a handful. The scale has failed to detect a change. Nobody sensible concludes that no rice left the sack. The reasonable conclusion is that a handful is smaller than what this instrument can see, and that if the question really matters somebody should fetch a better scale rather than write down that the sack is unchanged.
The record fails to reject a mean of zero. Is the mean zero?
Why is the claim written before the record is opened?
Wording is only half of the discipline. The other half is order, and order costs nothing at all: the same words, written at a different time. Once the interval has been seen the turning point is visible, so a claim written afterwards can always be placed on the surviving side of it. Nothing dishonest has to happen for this to go wrong. Somebody works out the interval, looks at it, thinks about what they always suspected, and writes down a claim that fits. Nobody is lying. The person genuinely believes they suspected it.
The smallest version of the problem is right here on this ladder. A claim of at least 1.50 per cent survives. A claim of at least 2.00 per cent does not. Anybody who has seen that the interval stops at 1.8788 per cent can choose which of those two they had believed all along, and both choices will look reasonable in writing. Neither wording is any more informative than the other, and both cost nothing to produce. A verdict that costs nothing to obtain is worth nothing when obtained.
There is a very ordinary version of this in any household that has ever bet on a cricket match after the fact. Claiming afterwards to have known the chase was gone once the fourth wicket fell costs nothing and cannot be checked, and everybody in the room has an equally good memory of having known it. Saying it before the fourth wicket falls costs something, and it is the only version anybody can be held to. Research works the same way and for the same reason, and writing the claim down first is the whole of the mechanism.
Somebody reads the interval first, and then writes down the claim they say they had held all along. Which claims are available to them?
What does a written claim have to pass before any data is read?
A working analyst, a credit reviewer or anybody signing off an internal note can run the check that follows on a Tuesday morning without any arithmetic whatsoever. Four questions, asked of the wording alone, before the record is opened. Each question takes about a minute, and every one can be answered from the sentence itself. Nothing in research is cheaper quality control.
One. Is there a reading that would end this claim? Not in principle, not eventually: it has to be written out. If that sentence is vague, the claim is vague, and no amount of careful work downstream will fix that. Two. Is that reading one this record could actually produce? A claim ended only by something a fifty month record cannot deliver is untestable here even if it is testable somewhere. Three. Would somebody else, handed the same wording, write down the same refuting reading? If two competent readers write down different endings, the claim has not been stated, it has been gestured at, and whichever ending suits the result will be the one that gets used.
Four, and this is the one that catches a claim written to survive: is the claim narrow enough that a count of chairs could not satisfy it? The count of chairs is the sharpest test available, so the question is worth asking in exactly those words. If something unrelated can be imagined coming through the claim, the claim is admitting everything, and the work about to be signed off will be reported as holding no matter what the record says. Tests one to three catch sloppiness. Test four catches a sentence that was written so that it could not lose, perhaps without anybody meaning to.
A lender's credit committee uses the same four questions without ever calling them that. When somebody brings a paper saying a borrower's collections have improved, the useful question in the room is never whether the paper is well argued. The question is: what would the committee have had to see this quarter to be told the collections had not improved, and was that written down last quarter or is it being invented now? A household saves the same way. Deciding to try a cheaper vegetable market for a month is only a decision if somebody says in advance what the monthly bill would have to come to before the household goes back to the old one.
A written claim passes the first three tests cleanly and fails the fourth. What has gone wrong with it?
The failure: a claim that held, and could not have done anything else
An analyst frames the work as a question about whether the outcome is related to anything at all, does the analysis carefully, and reports that the claim held. And it did hold. The holding is what makes this failure so hard to catch from the outside: nothing in the write up is false, no arithmetic is wrong, and the person is not being careless.
The claim held because it could not have done anything else. Three candidates were tried and all three satisfy it: the ordered input at 0.7559, the invented count of chairs at 0.9029, and the near duplicate at 0.7560. The number of candidates that would have refuted the claim is zero. There was no observation available anywhere in the exercise that could have gone the other way, and a verdict with no losing case behind it is not a result.
The cost lands later and lands hard. The work gets cited as a confirmed finding. Somebody builds the next study on top of it, and somebody after that builds on them. When the whole line eventually falls over, nobody can point to the cell, the month or the reading that should have stopped it. There never was one. A missing reading is a much worse position than a wrong number. Somebody can at least find a wrong number.
And the fix is a wording change that costs nothing. State the claim as a stated floor with a figure in it, and state the convention for refuting it in the same breath. Before the record is opened, both the writer and the reader can then name the readings that would have ended the claim. Same work, same record, same afternoon. Only the sentence changes, and the sentence is what made the difference between a finding and a formality.
Where this guide stops. Fixing the whole design, choosing the record and working out what a record of a given length can settle are covered separately. The difference between a mechanism and the record it produced, and whether a second person can reach the same number from the same data, is also covered separately. Structured challenge of finished work is covered separately. Whether a result survives a changed assumption, and which assumption it is most sensitive to, is covered separately. The arithmetic of a test statistic, how a threshold gets chosen and what happens when many claims are tested at once are all covered separately.
And what none of this claims. A claim that survives is not therefore probably wrong, and a refuted claim was not therefore better written. The verdict and the truth are two different questions, three of their four combinations actually occur on this record, and the wording of the claim decides whether the verdict was ever capable of telling a reader anything.
Is there a source to check, and what happens when there is not?
Every reading in this guide was produced by writing a rule down, drawing records from it and working the arithmetic; those derivations are covered separately. A figure with no outside source still has to be checkable, and one of these is checked by rerunning the script that made it rather than by looking it up. Putting the rule on paper first is the whole point: the true mean of 1.00 per cent is knowable only because somebody wrote it, and no record anywhere outside a lesson arrives with its own answer key attached.
| Reading used here | What produced it | Where it is rechecked | Outside source |
|---|---|---|---|
| The five value rule, its mean of 1.00 per cent and its spread of 5.00 per cent | Written out as five values and five weights adding to exactly one | The checking script kept beside these notes | None used |
| Record one's mean of 0.50 per cent, its standard error of 0.7035 per cent, and its interval from minus 0.8788 to 1.8788 per cent | Fifty months drawn from that rule | The same script, recomputed rather than copied across | None used |
| The reading against zero of 0.7107 and its two sided chance of about 0.4772 | The same fifty months, worked against a mean of zero | The same script | None used |
| The three candidate readings of 0.7559, 0.9029 and 0.7560 | Ten paired months, with a count of chairs and a near duplicate input laid against the same ten | The same script, which also checks that all three sit on ten months and not fifty | None used |
The fifty month record, the five value rule behind it, the ten paired months of an input and an outcome, the count of chairs put out in a hall and the near duplicate input are invented.
Educational material. Not advice on any investment, tax, budget or market position.
