Hypothesis Testing: The Null, the P-Value and the Threshold
A hypothesis test asks whether a record could plausibly have come from a stated position. Test the fifty month record of the Nakshatra unit against a monthly mean of zero: the statistic is 0.7107 and the p-value is about 0.4772, so the record cannot reject zero. The true mean is 1.00 per cent and positive. Failing to reject is a statement about the record, not about the world.
Somebody puts fifty months of figures in front of an analyst and asks one question. Does this thing drift upward, or does it only wander? A hypothesis test is the machine built to answer that question, and it answers a far narrower version of it than most people reading the output believe. The machine is assembled below out of two numbers already in hand and run once. Then the verdict is checked against the answer, a step almost no explanation of testing manages.
Two objects carry everything below, both of them invented for teaching. The Nakshatra unit is a generatorA rule set down on paper saying which values can come out and how heavily each one is weighted. Nothing about it was measured, so everything about it can be worked out by hand. for a traded unitSomething with a price attached that moves about from month to month. Here the phrase is a label on an invented object, so no instrument anybody could buy is meant.: its monthly changeThe move in a price across a single month, expressed against where that price started. The record in this guide is fifty of them and nothing else. can come out at any of five values: minus 9.00, minus 4.00, 1.00, 6.00 or 11.00 per cent, carrying weights of 0.08, 0.18, 0.48, 0.18 and 0.08. The fifty month record is one draw of fifty months from that rule, landing 5, 9, 25, 8 and 3 times on those five values. Its own mean comes to 0.50 per cent, and its standard errorHow much a summary figure such as a mean bounces about from one sample of the same size to the next. Where the figure of 0.70 per cent comes from is worked through separately. is 0.70 per cent, carried at 0.7035 through every calculation here. Both figures are stated once and used, and each is built separately.
A rule written down in advance carries an advantage no real dataset offers: the population is known, so the test can be watched missing something that is genuinely there. Somebody set the rule on paper, and only afterwards were fifty months pulled out of it, so a true monthly mean of 1.00 per cent and a true spread of 5.00 per cent follow from a definition rather than estimating anything. Ordinarily nobody has access to the answer a test was reaching for, so the test is the last word. Here the answer is sitting on the table the whole time. Everything turns on it: the truth is 1.00 per cent and positive, and by the end the test will have failed to see it.
What is a null hypothesis, and why start from one?
The null hypothesis is the position provisionally held while the record is examined. Here it is the statement that the true monthly mean of the Nakshatra unit is zero. Not the position anybody believes, not the position anybody hopes for, and not the position the record suggests. A null hypothesis is simply a stated position. Writing it down before the arithmetic starts gives the arithmetic something to be surprising about.
Starting from a position nobody believes strikes almost everybody as backwards on first meeting. The reason is worth stating bluntly. Impressive against what is not a question arithmetic can answer, so a test cannot ask whether a record is impressive on its own. A test can only ask whether a record is surprising under a stated position, and somebody has to state that position first. The convention is to state the dull one. Nothing is happening. The mean is zero. The two groups are the same. The check that follows is whether the record actually held sits comfortably with the dull position or awkwardly against it.
An everyday version helps here. A shopkeeper claims the new till queue is faster than the old one. Faster cannot be tested in the abstract, so the dull claim is written down first: the new queue takes exactly as long as the old one. Forty customers are then timed, and the question is how odd those forty timings would look if the dull claim were true. The dull claim is doing real work. Without it there is nothing to be surprised by, only forty numbers.
Choosing the null hypothesis is a decision a person makes, and different nulls give different verdicts on exactly the same record. The null arrives on most test reports looking like a fact about the data. A null is nothing of the kind. Somebody typed it. Below, the same fifty months are tested against a mean of zero and against a mean of 1.00 per cent, and both tests come back with the identical p-value and the identical verdict. If the null were something the record decided, that could not happen.
Who or what settles the null hypothesis in a test on the fifty month record?
How is the test statistic built from the record?
Three quantities go in and one number comes out, and all three are already in hand. The record's mean is 0.50 per cent. The null hypothesis states zero. Subtracting that leaves a gap of 0.50 per cent. Dividing that gap by 0.7035 per cent, the wobble in a fifty month mean from one draw to the next, gives 0.7107.
The statistic is the record's mean, less the position, divided by the wobble
| What goes in | Where it is already in hand | Value |
|---|---|---|
| The record's mean | The mean of the fifty month record, computed earlier and carried in unchanged | 0.50 per cent |
| The position being tested | Written down by the person running the test, before the arithmetic | 0.00 per cent |
| The standard error, the wobble | How far a mean of fifty months moves from one draw to the next | 0.7035 per cent |
| The statistic | 0.50 less 0.00, all divided by 0.7035 | 0.7107 |
Said in words. The statistic is the distance between what the record said and what the null hypothesis said, measured in units of how much a record like this wobbles anyway. A statistic of 0.7107 says the gap is smaller than the ordinary wobble.
A gap only means something once it is compared with how much the measurement moves about on its own, and the division is the whole idea. A gap of 0.50 per cent sounds like something on its own. A fresh fifty months from the same rule would routinely produce a mean 0.70 per cent away from the last one. Held against that, 0.50 is less than one ordinary step. The statistic of 0.7107 is saying exactly that and nothing more: the record landed about seven tenths of a normal step away from the position being tested.
Consider a bathroom scale that reads two kilograms differently depending on where a person stands on it. A weighing comes out one kilogram above last week's. The one kilogram is real in the sense that it was read off the dial, but it is smaller than the noise in the instrument, so it will not support a claim on its own. Dividing by the wobble is how a test refuses to be impressed by a difference smaller than the thing measuring it.
The record's mean is 0.50 per cent and its standard error is 0.7035 per cent. What is the statistic against a null hypothesis of zero?
What is a p-value, and what does it actually measure?
The p-value is the chance of seeing a gap at least this large if the null hypothesis were true. On the fifty month record, tested against a mean of zero, that chance is about 0.4772. Read the definition again and notice where the weight falls: everything to the right of the word if is doing the work. The p-value is computed inside a world where the null hypothesis is assumed correct, and it never leaves that world. A number that never leaves that world can say nothing at all about whether the null hypothesis is correct.
Here is the picture behind the number. Assume the true mean really is zero. Fifty month records drawn under that assumption would still not average exactly zero; they would scatter, and the normal shapeThe familiar single humped curve, tallest in the middle and thinning away symmetrically on both sides. The curve is built separately and carried in unchanged. built earlier describes how they scatter. Mark the point 0.7107 wobbles to the right of centre, mark the same distance to the left, and shade everything further out than those two marks. The shaded area comes to 0.4772, not far off half of everything. A gap this size is thoroughly ordinary in a world where the mean is zero.
Three misreadings of a p-value reach print constantly. A p-value is not the chance the null hypothesis is true, not the chance the result came about by luck, and not a measure of how big anything is. The first two both make the same mistake in different clothes: they turn the definition around. The p-value runs from an assumed position to the chance of the record, and it cannot be read backwards from the record to the chance of the position. The third is a different mistake and a more expensive one. A p-value can be pushed arbitrarily low by adding cases, whatever the size of the difference.
The p-value on this record is 0.4772. Is that the chance the Nakshatra unit has a true monthly mean of zero?
What does statistical significance mean, and where does the usual threshold come from?
Statistical significance is a label applied when a p-value falls below a threshold that somebody chose in advance. The threshold in widest use is 0.05. Below it, a result is conventionally called statistically significant; above it, it is not. The threshold and the comparison are the entire content of the phrase.
The threshold is a convention people adopted rather than a property anybody derived, so nothing whatever changes in the arithmetic at 0.05. A p-value of 0.049 and a p-value of 0.051 describe practically identical records. The curve does not kink there, the record does not improve as it crosses, and no calculation behaves differently on the two sides. A word in a report changes. Nothing in the arithmetic does. Treat any account that describes the line as a rule handed down by some authority with suspicion: it is a habit of practice, widely shared and useful as a common shorthand, and it is not issued by anybody and applies nowhere by force.
The words mislead in ordinary English, so the label needs stating exactly. Statistically significant means the record would have been surprising under the stated null hypothesis, and it does not mean the difference is large, important, useful, or reliable enough to act on. Significant in daily speech means it matters. Significant here means it was unlikely under one specific assumption. The two meanings came from different places and only share a spelling. A threshold also promises something about the long runWhat happens across a great many repetitions rather than in the one case at hand. A rule can behave well in the long run and still get the case in front of it wrong. rather than about the case at hand: holding to a 0.05 line for years means wrongly crying discovery on a small fraction of the occasions where nothing was happening. The line says nothing about whether this record is one of them.
Two records give p-values of 0.049 and 0.051. What changes in the arithmetic as a p-value crosses 0.05?
What does this record say about a mean of zero?
The record says nothing that would rule zero out. The p-value of 0.4772 is not near 0.05, not approaching it, and not suggestive of it. The p-value sits almost ten times higher. The fifty month record of the Nakshatra unit cannot reject a monthly mean of zero, and the whole test runs end to end in seven lines.
| The step | Where the figure comes from | The figure |
|---|---|---|
| The record's mean | The fifty monthly changes added and divided by fifty months | 0.50 per cent |
| The position being tested | Chosen by the person running the test, not by the record | 0.00 per cent |
| The gap | 0.50 less 0.00 | 0.50 per cent |
| The wobble, the standard error | Carried in from the record's own spread over fifty months | 0.7035 per cent |
| The statistic | 0.50 divided by 0.7035 | 0.7107 |
| The p-value, both sides counted | The chance of a gap this wide if the mean were zero | 0.4772 |
| The verdict at the usual line of 0.05 | 0.4772 is nowhere near 0.05 | cannot reject |
The true mean is 1.00 per cent and positive, by construction, and the test cannot see it. Not because the arithmetic slipped. Not because the record was drawn badly. Not because the method is broken. Every step above is correct, and the verdict it produced is the right verdict for a record of this length. Fifty months, on a rule whose values scatter with a spread of 5.00 per cent, simply cannot resolve a monthly effect of 1.00 per cent. The instrument is not fine enough for the thing being measured, and a test issues no warning about that. A test just returns cannot reject, in the same untroubled tone it would use if the mean really were zero.
Sit with what that means for a moment. The point estimateA lone figure put forward as the answer, carrying no width of its own. The missing width is what makes 0.50 per cent look more settled than fifty months can support. of 0.50 per cent was pointing in the right direction the whole time, and it was half the truth. The verdict came back cannot reject. Both of those are honest outputs of correct arithmetic on a genuine record, and together they would leave a reader with no idea that the answer is 1.00 per cent and positive. None of that is a failure of statistics. The size of what fifty months can settle is simply much smaller than people expect.
The test cannot reject a mean of zero. The true mean is 1.00 per cent. What has the test just reported on?
As the position being tested is slid up from zero towards 1.00 per cent, does the p-value rise, fall, or do both?
Slide the position being tested and watch a whole band of them survive.
One control. The control moves the null hypothesis, the position being tested, from minus 2.00 to 3.00 per cent in steps of 0.01. The record never changes: its mean stays 0.50 per cent and its standard error stays 0.7035 per cent at every setting, and only the position moves. The curve reshades as the statistic moves, the chance bar underneath rescales against the 0.05 line, the verdict strip at the foot shows every position at once, and the sentence below restates the current reading in words. The default of zero reproduces the worked test above exactly.
Educational illustration. The Nakshatra unit and the fifty month record exist only as teaching objects. Both figures the test needs stay put wherever the control is set, a mean of 0.50 per cent and a wobble of 0.7035 per cent, and the position being tested is the only thing that moves. The chance is worked out from the normal shape and is approximate to the four decimals shown. The true mean of 1.00 per cent can be marked on the strip only because somebody set the rule on paper ahead of drawing anything out of it. No real record extends that courtesy.
Why is failing to reject not evidence that the mean is zero?
Because the same fifty months also fail to reject a mean of 1.00 per cent, with exactly the same p-value. Run it. The record's mean of 0.50 per cent less a null value of 1.00 per cent leaves a gap of minus 0.50 per cent. Divide by the standard error of 0.7035 and the statistic is minus 0.7107, the same distance from the centre as before and on the other side. The chance of a gap at least that wide, counted on both sides, is 0.4772 again. The verdict is cannot reject again.
The reason the two match exactly is worth seeing rather than accepting. A record averaging 0.50 per cent sits precisely halfway between zero and 1.00 per cent, so the gap is the same size measured towards either of them, and the test only ever looks at the size of the gap. The fifty month record cannot tell a mean of zero from the truth, and a note reporting only the first of those two tests would leave a reader believing it had.
So the two verdicts are not symmetric in what they buy. Rejecting a null hypothesis establishes something specific. The record would have been surprising under that position, leaving it awkward to hold. Failing to reject is a verdict shared by a great many positions at once, so it singles out none of them and establishes much less. Cannot reject zero is true. Cannot reject 1.00 per cent is equally true. Cannot reject 1.50 per cent is true as well. A verdict that hundreds of positions all pass is not evidence for any one of them.
The same record fails to reject zero and fails to reject 1.00 per cent, with the same p-value. What does that show about what fifty months can settle?
Which positions can this record not reject?
All of them between minus 0.8788 and 1.8788 per cent, and that is a far more useful thing to report than any single verdict. Work out where the boundary falls. A position gets rejected at the usual line once the statistic passes about 1.96. The statistic passes 1.96 once the gap passes 1.96 multiplied by the standard error of 0.7035, or 1.3788 per cent. Measure that distance out on both sides of 0.50 per cent and the two edges land at minus 0.8788 and 1.8788 per cent. Rounded to two places, minus 0.88 and 1.88.
Look at what is sitting inside that band. Zero sits inside the band, and that verdict was already reached above. The truth of 1.00 per cent is inside it too. So is every position in between and a good deal on either side. A test and an interval are the same arithmetic asked in two directions, so the band from minus 0.8788 to 1.8788 per cent has the same pair of endpoints as the interval built earlier from these two figures. A test asks whether one named position survives. An interval asks which positions survive. Same mean, same standard error, same 1.96, one question turned inside out.
Which positions does the fifty month record fail to reject, and what else here shares those two endpoints?
What should be asked of any test result handed over by somebody else?
Five questions, and they take about a minute. What was the null hypothesis, and who chose it? How large is the difference, stated in units somebody cares about? How many cases were there? What size of difference could a record this small have detected at all? And if the report says nothing was found, what else would have produced that same verdict? A p-value printed with no size of difference and no count of cases beside it cannot be read by anybody, however carefully it was computed.
The fourth question is the one nobody asks and the one that matters most. A record too short to detect the effect it was pointed at will report cannot reject, cleanly and confidently, exactly as the fifty month record did here. Asking what could this have detected turns a null result from a finding into a measurement of the record's own reach. How big a difference is, and how many months it would take to settle one, are both worked out separately, under effect sizeHow big a difference is in units somebody cares about, as against whether it turned up in a test at all. Effect size is measured separately..
A lender, an analyst or a household should read any claim of the same shape the same way, and the shape turns up constantly outside statistics. A clinic tests forty patients and reports no difference between two treatments. A school tries a new timetable in one classroom for a term and reports no improvement. A household tracks its electricity bill for three months after changing a habit and concludes the habit made no difference. In every one of those, the honest reading is the same: the study was not large enough to see the thing it was looking for, so it did not look for it in any meaningful sense. Nothing was found is a statement about the size of the search, not about the world.
A colleague reports that no effect was found. Which two figures should be asked for before reading that as evidence of anything?
The note that says nothing was found, read as though it said something
A short note circulates. A test on fifty months of the Nakshatra unit failed to reject a monthly mean of zero, p-value 0.4772, no significant effect found. A reader takes that at face value and concludes the unit goes nowhere. Everything in the note is true and correctly computed. The true mean is 1.00 per cent and positive, so the conclusion drawn from the note points in exactly the wrong direction.
The same fifty months equally fail to reject 1.00 per cent, and 1.50 per cent, and every position from minus 0.8788 to 1.8788 per cent. The note mentioned none of them. The note reported one verdict from a band of positions that all share it, and a verdict shared by a band cannot pick out a member of that band. The cost is a reader who now believes something false and believes it on the strength of a test. An error held on that basis is far harder to dislodge than a hunch.
The fix is two lines long. Report the whole surviving band rather than a single verdict, so the reader sees at once that zero and 1.00 per cent are both in it. Then state, beside every result of this kind, the size of difference the record could have detected. A record that could not have seen the effect has not looked for it, and saying so is the difference between an honest null result and a misleading one.
How large the difference is, how to state it in units a person cares about, and how many months it would take to settle one are covered separately. Where the standard error of 0.7035 per cent comes from, and why it shrinks the way it does as months are added, is covered separately. Setting a statistical verdict against an economic one, head to head, is covered separately. Fitting a line through points, estimating a parameter from the relationship between two things, and forecasting a month that has not happened are covered separately.
Where do these numbers come from?
A test is arithmetic, and arithmetic answers to nobody. The Nakshatra unit is a rule somebody wrote down first: five monthly outcomes with five stated weights, and nothing else. The rule existed before the record did, so the true monthly mean of 1.00 per cent is not an estimate of anything. The true mean is arithmetic performed on a definition. The fifty month record was then produced from that rule, and every statistic quoted above is recomputed from those fifty months.
| What is quoted | Where it comes from | What kind of source that is | Checked on |
|---|---|---|---|
| True mean 1.00 per cent and true spread 5.00 per cent | Arithmetic on the five value definition of the Nakshatra unit | Invented teaching object, no outside authority | 19 August 2026 |
| Record mean 0.50 per cent and standard error 0.7035 per cent | The fifty months of the record, recomputed rather than copied | Invented teaching object, no outside authority | 19 August 2026 |
| Statistic 0.7107 and p-value 0.4772 | Those two figures put through the definition of the statistic | Arithmetic anybody can redo by hand | 19 August 2026 |
| The band from minus 0.8788 to 1.8788 per cent | 1.96 multiplied by 0.7035, taken either side of 0.50 | Arithmetic anybody can redo by hand | 19 August 2026 |
| The 0.05 line | A convention in wide use, derived from nothing and issued by nobody | Named as a convention, no authority claimed | 19 August 2026 |
The Nakshatra unit, the fifty month record and the Vasant unit are invented.
Educational material. Not advice on any investment, tax, budget or market position.
