Research Design: Deciding What Would Change Your Mind
A research design fixes four things before any data is read: the outcome, the input, the record that will be used, and the reading that would overturn the claim. The last one is the whole of it. A design that names no reading capable of ending the claim has arranged for the claim to survive. Arranging for a claim to survive is not the same as testing it.
Five worked cases sit underneath what follows, and none of them is new. Each one was computed, checked and drawn earlier in this subject area, and every one of them is research that went wrong in a way nobody could see from the answer. There is a manufactured pattern. There is a count of chairs put out in a hall that beats a real relationship on every arithmetic measure anybody thought to apply. There is a fifty month record that reads half the truth. There is a near duplicate input that flips a coefficient while the fit barely moves. And there is a figure computed with a month of information that had not happened yet. None of them is diagnosed again here. The five failures are put to work instead. An argument that can point at a failure already watched happening is one that no explanation of research method can borrow.
One sentence carries the whole of research design. Before the record is read, the reading that would force the claim to be abandoned is written down. Everything below is either what that sentence costs, what it protects against, or the arithmetic that settles whether the record could ever produce such a reading at all.
What is a research design, and when is it written?
A research design is a written statement of four things: the question being asked, the record it will be asked of, the method to be used, and the reading that would settle it. A design differs from a description entirely in its timing. The design is written while none of the answers are known. Once the record has been opened, whatever is written down is no longer a design. Nobody can tell it apart from a report of what the record happened to say.
Think about a stall keeper trying a new pitch. One of them decides in the morning, before the first customer, that a day taking less than a stated figure means the new pitch was a mistake and the old one comes back tomorrow. By evening that person knows the answer, whichever way it goes, and knows it without arguing with themselves. The other stall keeper decides in the evening. Whatever the day took, there is a reason it was fine: it rained, it was a working day, the regular customers had not found the new spot yet. The second stall keeper is not being dishonest. The second stall keeper is doing something structurally impossible. Judging a result against a line drawn after the result arrived is not judging at all.
Timing is the whole of it, and the rest is detail. A design is not paperwork and it is not a formality completed so that somebody senior can sign it. A design is the only mechanism anyone has found for making a claim capable of losing.
A research design and a description of what a record turned out to say are not the same thing. What does the first one have that the second one cannot have?
How to Define a Quantitative Research Question: what are the four parts?
A question that data could answer has four parts, and any question can be checked against them in about a minute. The first is the outcome, named and measurable. Two people reading the same record must produce the same figure for it. Not "how things went" but a stated quantity with a stated way of arriving at it.
The second is the input, named and available at the time it is attached to. Availability at the time it is attached to matters far more than it looks. An input that is perfectly well defined but only knowable a month later is not an input at all. One of the five cases is exactly that failure: a figure computed with a month of information that had not happened yet, and it improved a fit by 48.1650 per cent purely by knowing the future.
The third is the record, fixed by length and by which observations are in it. Not "whatever data happens to be available" but a stated set: this many months, these observations, and any observation dropped is dropped for a reason written down at the same time.
The fourth is the reading that would settle the question in either direction. The fourth part is the one almost nobody writes, and its absence is invisible. A question missing only the fourth part still reads like a question. Such a question has a subject, a verb and a quantity, and it has no way to lose. A question with three of the four parts is a topic.
A colleague writes down an outcome, an input and a record of sixty months, and asks for it to be run. Which of the four parts is missing, and what goes wrong without it?
What would overturn the claim, and why is that decided first?
Here is the sentence this whole sequence exists for. Before the record is read, the reading that would force the claim to be abandoned is written down. Then the record is read. A line drawn after the answer is known is not a line at all, so deciding it first is the only version that works.
The fifty month record, one of the five cases behind this guide, works as an example. The claim is that the true mean sits above zero. The refuting reading is written down first, in one sentence: a 95 per cent intervalA pair of ends reported alongside an estimate, marking off the values the record leaves open instead of the single value it happened to give. containing zero ends the claim. The claim and the refuting reading are the whole design, and nothing has been read when they are written.
Now the record is opened. The record reads 0.50 per cent, its interval contains zero, and the claim is abandoned exactly as written. Was the claim true?
The record reads 0.50 per cent. Its standard errorA figure saying how far an estimate would typically wander if the same quantity of data were gathered all over again. is 0.7035 per cent, so the range around it runs from minus 0.8788 per cent to 1.8788 per cent. The range contains zero. The claim is abandoned, exactly as the design said it would be.
And the generator behind that record has a true mean of 1.00 per cent, a figure above zero. A correctly designed study reached the correct decision and got the world wrong, and that is not a fault in the design. Fifty months can resolve only so much, and how much they can resolve is a different question entirely. If this feels uncomfortable, sit with the discomfort for a moment. The alternative is worse: a design that could not have abandoned the claim would have been right here by accident, and would have been right by accident every other time too.
What can a record already in hand actually settle?
How far a record can reach is arithmetic rather than judgement, and the arithmetic is available before anything is read. The length of the record decides what the record could ever settle, and it decides it in advance. The generator behind that fifty month record has a true spread of 5.00 per cent, known here by construction and never known in practice. Work the half widthHalf the width of a reported range, measured from the middle out to one end, so doubling it gives the whole range. of the 95 per cent range at five different lengths and a ladder appears.
| Months in the record | Half width of the range | Could it separate 1.00 per cent from zero, even at the truth? |
|---|---|---|
| 25 | 1.96 per cent | No |
| 50 | 1.39 per cent | No |
| 100 | 0.98 per cent | Yes |
| 200 | 0.69 per cent | Yes |
| 400 | 0.49 per cent | Yes |
The third column repays slow reading. At 25 and at 50 months the range is wider than the effect being hunted, so even a reading that landed exactly on the truth of 1.00 per cent would still reach down past zero. At 100, 200 and 400 months it does not. The ladder falls at every step, it never lands on 1.00 per cent, and it passes below the true mean exactly once. None of this needed the record. All of it was knowable from the length and the spread alone.
The true mean is 1.00 per cent and the spread is 5.00 per cent. Before the control below moves, what does a record of 400 months do to the range: narrower than 1.00 per cent, or wider?
Move the length of the record and watch the range close around a fixed truth
One control, and all it changes is the number of months the design commits to. Nothing is read. The estimate is placed on the true mean of 1.00 per cent, the kindest thing that could happen to it, so what appears is the best case rather than a result. Two views redraw together. The top one shows the whole range against a fixed line at zero and a fixed mark at the truth. At 100 months the range clears zero by 0.02 per cent, and no full scale drawing can show a gap that small, so the bottom one magnifies a narrow strip either side of zero. The hold control leaves the current length on screen as a faint outline for comparison with the next one.
Educational illustration. The unit, its generator and every record drawn from it were invented for teaching. The true spread of 5.00 per cent is known here by construction and is never known in practice. The range drawn assumes the estimate lands exactly on the truth. No estimate ever does, so every setting flatters itself.
A design has fifty months and is meant to separate an effect of 1.00 per cent from zero. What should the person who asked for it be told?
So how long would it have to be? The shortest record that could push the range clear of zero here is 97 months, and a condition travels with that figure everywhere it goes: 97 months works only if the sample mean lands exactly on the truth, and 96 months does not work even then. A record of 97 months buys a single best case, not a working design. To find an effect of that size reliably rather than once takes 197 months, more than double.
The shortest record that could push the range clear of zero here is 97 months, and a condition travels with it. What is the condition?
There are two separate questions to put to any design, and mixing them up is common. The first is whether the reading named would actually settle the question. The second is whether the record in hand could ever produce that reading. A design can pass the first question perfectly and fail the second one completely. The fifty month design is exactly that case: its refuting reading is well chosen, unambiguous and decided in advance, and the record it sits on could never have produced anything else.
How to Test a Financial Hypothesis Responsibly: what do the four conditions ask?
Responsibly is not a mood and it is not a matter of care. Responsibility means four specific things, and each one can be checked by anybody who was not there when the work was done. The claim was written before the record was read. The refuting reading was named in the same breath. The record was fixed rather than extended until the answer arrived. And the result is reported with the range around it rather than as a single point.
The third is the hardest, and it is the one that fails quietly. A record that grows until the answer appears has no stated length. A line only means something when the thing being measured against it was fixed first, so without a stated length no thresholdA line put down ahead of time, so that whichever side a reading lands on, the side was chosen before the reading existed. means anything. Nobody experiences this as cheating. Extending the record feels like diligence: one more month of data, one more observation, everything available may as well be used. The stall keeper who keeps the new pitch open one more day, and then one more, until a good day arrives, has not run a longer test. The stall keeper has run no test.
The second condition is the cheapest of the four to satisfy and it catches the most, especially read as a question about dates rather than about statistics. The question, put to every single input, is whether it could have been known at the moment it is attached to. One of the five cases is a figure computed with one month of information that had not happened yet, and it improved a fit by 48.1650 per cent on that basis alone. A design that wrote down when each input became available would have caught it before the record was ever opened, at a cost of one extra column on a sheet of paper.
What does one whole design look like, walked through in order?
Here is the fifty month design end to end, in the order it actually happened, showing which steps needed the record and which did not. Only the third step reads anything.
| Step | What happens | What it needed |
|---|---|---|
| 1 | The design is written with nothing read. The outcome is the average monthly change of an invented unit, the record is fifty months, the claim is that the true mean sits above zero, and the refuting reading is a 95 per cent range containing zero. | Nothing |
| 2 | The arithmetic the design could have done in advance. With a spread of 5.00 per cent, fifty months give a half width of 1.39 per cent, and an effect of 1.00 per cent is smaller than that. | The length and the spread |
| 3 | The record is opened. The record reads 0.50 per cent with a standard error of 0.7035 per cent, so the range runs from minus 0.8788 per cent to 1.8788 per cent. | The record |
| 4 | The verdict. The range contains zero, so the claim is abandoned exactly as written. | Steps 1 and 3 |
| 5 | The truth, never available to the design. The generator's mean is 1.00 per cent, and it is positive. | Nothing available in practice |
| 6 | What the design should have said at step 1. Ninety seven months at best and only if the sample mean lands on the truth, 197 months to find it reliably, and fifty months answers a smaller question or none. | Step 2, and it was free |
Look at the last column. Step 2 was available before anything was read and it costs nothing, and had anybody done it at step 1 the whole exercise would have been redesigned rather than run and abandoned. The most valuable arithmetic in a research project is usually the arithmetic that can be done before any data exists.
What happens when the question is chosen after the record is read?
Now the other direction of failure, and it is the one that looks most like good work. Instead of the question being written first, the record is opened, whatever moves with the outcome is hunted out, and the question is written around it. Asked which of two inputs explains an outcome, and choosing on the fit alone, an invented count of chairs put out each month in a hall wins at 0.9029 against the real relationship's 0.7559.
The chair count beats the real relationship on the fit. Of six arithmetic measures anybody might apply, how many does the chair count win?
The chair count does not win narrowly, and it does not win on a matter of judgement. The chairs also win on the misses left over, on the correlation with the outcome, on the slope's own reading against zero, on the misses at lag one where the chairs' reading of minus 0.1908 is smaller in size than the real relationship's 0.4862, and on how closely the two halves of the record agree, where the chairs agree to 0.0014 and the real relationship agrees only to 0.1893. Six measures, and the chairs win all six.
Two candidate inputs are on offer. One has a fit of 0.9029 and the other 0.7559. What should be asked before choosing between them?
The failure: the question written around whatever the record offered
An analyst opens the record first, hunts for what moves with the outcome, finds the chair count at a fit of 0.9029 against the real relationship's 0.7559, and writes the question around it. The fit is not wrong, and the cost lies somewhere else. The fit reads 0.9029 and it reconciles, and anyone who recomputes it will get 0.9029 again. The cost is that no arithmetic is left that could catch it. The chairs also beat the real relationship on the misses left over at 86.7349 against 218.0000, on the correlation at 0.9502 against 0.8694, on the slope's own reading against zero at 8.6236 against 4.9770, on the misses at lag one at minus 0.1908 against 0.4862 in size, and on how closely the two halves agree at 0.0014 against 0.1893.
Six measures and the chairs win all six, so the reviewer who checks the arithmetic signs it off, correctly. The fix costs nothing and has to happen first: write the question and the refuting reading before the record is opened. Do that, and a count of chairs can never become the answer to a question nobody asked.
What does a design write down before the first figure is computed?
Here is the artefact rather than the advice. A working design sheet has seven lines on it, and every one of them can be filled in before a single figure exists. Line one, the outcome, in one sentence. Line two, the input, and whether it could have been known at the time. Line three, which observations are in the record and how many of them. Line four, the reading that would end the claim. Line five, the reading that would leave it standing, and whether that reading is actually different from the fourth. Line six, the conventionAn agreed way of counting something, set down beforehand so that the next person counts it the same way. sitting behind every figure, including the denominatorWhatever the total was divided by. The denominator sits below the line and settles what the figure above the line is a figure of. underneath every average. And line seven, what would have been said if the record had come out the other way.
The seven lines are not a form for researchers. A lender deciding whether a borrower's takings have genuinely improved has exactly this sheet to fill in, whether they call it that or not: which months count, what improvement would be enough, and what reading would send the file back. A household deciding whether the new shop location is working has the same seven lines. So does an analyst asked whether a change in policy did anything. The people who write the seven lines down are not more rigorous than the ones who do not; they are the ones who will still be able to say next month what they thought before they knew.
A design arrives with all seven lines filled in except the last. What does that missing line say about when the design was written?
What sits outside a research design
The difference between a generator and the record it produced is covered separately, as is whether somebody else arrives at the same figure from the same data, and whether a fresh record arrives at the same finding. How a claim is written so that a reading can refute it, and how structured challenge is run against finished work, are covered separately too. Whether a result survives a changed assumption, and which assumption carries it, is covered separately. Testing a rule on history without self deception is covered separately as well.
What a reader might have expected to find cited here
| What might be expected | What stands there instead | Whether it can be checked |
|---|---|---|
| A kept record of monthly changes | Fifty months drawn from a five value generator built for teaching, with a true mean of 1.00 per cent and a true spread of 5.00 per cent | No, and it was never meant to be |
| A published finding about how long a record needs to be | The distance recomputed at five lengths from the generator's own spread, reading 1.96, 1.39, 0.98, 0.69 and 0.49 per cent | Yes, by doing the arithmetic again |
| A named account of a research failure | Five failures already computed earlier in this subject area, including the chair count and the leaked month | Yes, by rereading the stretch that computed them |
| A rule, a rate, a period or a line drawn by somebody | A research design carries no rule, rate, period or line, so none appears above | Nothing to look up |
The five value generator, the fifty month record of the unit drawn from it, the count of chairs put out each month in a hall and the ten paired months of an input and an outcome are invented.
Educational material. Not advice on any investment, tax, budget or market position.
