Effect Size: Why Significance Without Magnitude Is Empty
An effect size is how big a difference is, stated in units a reader can act on and kept apart from whether any record could detect it. The Nakshatra unit's true monthly mean of 1.00 per cent is 0.2000 of a standard deviation. As a share of the spread that is small. Compounded over twelve months the same mean comes to 12.68 per cent, and 12.68 per cent is not small. Detecting an effect that size reliably would take 197 months.
Everything below runs on one invented setup, and it is worth laying out before a single claim is made. There is a traded unitSomething that can be bought and sold, whose price is written down at fixed intervals. Here it matters only as a thing that produces one number a month, and nothing about how it is bought or valued is used. named the Nakshatra unit, made up for these notes alone. Each month it posts one monthly changeThe percentage by which a price ended a month above or below where it started that month. One figure per month, and no more than that., and that figure can land on exactly five values, minus 9.00, minus 4.00, 1.00, 6.00 and 11.00 per cent, weighted in that same order 0.08, 0.18, 0.48, 0.18 and 0.08. Values with weights are the whole generatorA fixed roster of outcomes with a weight beside each one, settled in advance. Every record used in these notes came out of that roster after it was settled.. Multiply each value by its weight and add, and the true centre comes to 1.00 per cent. Do the matching arithmetic on the distances from that centre and the true standard deviationA single figure for how widely the entries of a list scatter around their middle. Built in notes covered separately, and divided by here rather than rebuilt. comes to 5.00 per cent.
The true centre and the true spread are known here rather than estimated. Knowing both makes it possible to state the true size of an effect and then ask, separately, how much record it would take to see it. On real numbers the true centre is precisely what the whole exercise was trying to uncover, so nobody can check the second question against the first. Here the generator was fixed first and the months were drawn out of it afterwards, so both sides are on the table at once. A fifty month record was then drawn from that generator, with the five values turning up 5, 9, 25, 8 and 3 times. Its own average comes to 0.50 per cent, half the truth. A known truth of 1.00 per cent beside a record reporting 0.50 per cent is the pairing to hold on to.
What is an effect size, and what is it actually for?
An effect size is the magnitude of a difference, stated on its own, with no reference whatever to how many cases somebody managed to collect. The definition stops there, and the useful part of it is what has been left out. The definition mentions no length of record. Nothing in it concerns how confident anybody is. Nothing in it says whether the difference could be established. How big a thing is, and whether a particular record can see it, are two different questions, and only the first is a fact about the world; the second is a fact about the record.
Take a vendor who moves his cart twenty metres closer to the station entrance and afterwards sells four more plates a day. Four plates is the effect. Four plates is a fact about the new position, and it stays four plates whether he counts for a week or for three years. Now ask a different question: could he tell? His daily sales already swing by thirty plates depending on weather, day of the week and whether it rained at four in the afternoon. Four plates sits inside that swing. He might count for two months and see nothing at all, and the four plates would still be there the whole time, quietly, every day. The two questions have separate answers and separate arithmetic.
The distinction matters because the two questions get reported as though they were one. A note that says a difference could not be established has answered the second question and said nothing about the first. A note that says a difference is large has answered the first and said nothing about the second. A result carrying only one of the two answers is unreadable. The reader cannot tell which of the two questions was actually settled.
How is an effect stated as a share of the spread?
Divide the effect by the spread it sits inside. Here the true monthly mean is 1.00 per cent and the true standard deviation is 5.00 per cent. The effect as a share of the spread is 1.00 divided by 5.00, or 0.2000 exactly. Nothing has been added to the two figures and nothing has been assumed about them. One number is divided by another. Per cent cancels against per cent, and the answer carries no unit at all.
The division compares the signal against the ordinary month to month wobble the signal is buried in. A share of 0.2000 says the effect is one fifth the size of the noise surrounding it. The share states how hard the effect is to notice, not how much it does. Go back to the vendor: four extra plates against a daily swing of thirty plates is a share of about 0.13, and the reason he cannot tell is right there in the ratio. The bigger the wobble, the smaller the share, and the same effect gets harder to see without getting any smaller.
By the conventions taught for reading these shares, 0.2000 sits at the small end. The conventions are rough bands rather than laws, and they vary between texts. The sensible practice is to say where a share sits and let the reader see the number itself, so 0.2000 stands above in place of a label. Only the direction is worth taking from the convention: a fifth of the wobble is on the faint side of anything anyone would call obvious.
The true monthly mean is 1.00 per cent and the true standard deviation is 5.00 per cent. What is the effect stated as a share of the spread?
What does economic significance look like in units somebody actually counts?
Economic significance is the same effect written in units a person keeps track of, rather than in units of noise. Nobody counts standard deviations. People count months, plates, rupees and years. So take the identical 1.00 per cent a month and ask what it comes to over a year of such months. One idea has to be defined first, and no finance is assumed anywhere in the working.
Compounding means each period's change works on whatever the period before it left standing, rather than on the sum first put in. Month one takes 1.00 per cent of 100 and leaves 101. Month two takes 1.00 per cent of 101, not of 100, so it adds 1.01 rather than 1.00 and leaves 102.01. Month three works on 102.01. Repeat that twelve times and the total is not twelve lots of 1.00. The total is 12.6825 per cent, rounded here to 12.68. The extra 0.68 per cent is nothing but the later months working on ground the earlier months added, and no assumption of any kind was needed to produce it.
| Month | Applied to | Added this month | Total so far, per cent |
|---|---|---|---|
| 1 | 100.0000 | 1.0000 | 1.0000 |
| 2 | 101.0000 | 1.0100 | 2.0100 |
| 3 | 102.0100 | 1.0201 | 3.0301 |
| 6 | 105.1010 | 1.0510 | 6.1520 |
| 9 | 108.2857 | 1.0829 | 9.3685 |
| 11 | 110.4622 | 1.1046 | 11.5668 |
| 12 | 111.5668 | 1.1157 | 12.6825 |
Put a quantity against it and the second reading stops being abstract. Carry Rs 1,00,000/- through twelve months of this invented generator's true centre and the arithmetic ends at Rs 1,12,683/- in whole rupees. Adding the twelve months instead of compounding them gives Rs 1,12,000/-, short by Rs 683/-. The same 1.00 per cent that looked like a rounding error against the noise is 12.68 per cent and Rs 12,683/- on a lakh once it repeats, and nothing was added to the 1.00 per cent along the way.
Twelve months at 1.00 per cent each come to 12.68 per cent rather than 12.00. Where does the extra 0.68 per cent come from?
How can one effect be small and large at the same time?
Set the two readings beside each other and look at them without deciding which to keep. The Nakshatra unit's true monthly mean is 0.2000 of a standard deviation. The same monthly mean is 12.68 per cent compounded over twelve months. There is one underlying number, 1.00 per cent a month, and both readings are exact arithmetic on it. The two readings answer different questions, so neither conflicts with the other and neither is the correct one: the first says how hard the effect is to see, the second says how much it changes.
The confusion comes from the words attached to them. A share of 0.2000 gets called small. A twelve month figure of 12.68 per cent gets called large. Both labels are fair descriptions of what they measure, and they are measuring different things, so the labels cannot be compared with each other any more than a temperature can be compared with a weight. A reader handed only the share will conclude the effect hardly matters. A reader handed only the twelve month figure will conclude it is obvious and easily established. Both readers have been misled, and neither was told anything false.
One effect reads as 0.2000 of a spread and as 12.68 per cent over twelve months. Which of the two is the real one?
A colleague reads the share of 0.2000 and calls the effect negligible. What have they answered, and what have they left untouched?
How many months would it take to detect an effect this size?
The question is about the record rather than about the unit, and it has a plain arithmetic answer. The number of months asked for is built from three things: the spread squared, 25.00; the effect squared, 1.00 here; and a pair of conventional cut-off numbers from the testing machinery covered separately. The two cut-offs are 1.959964 and 0.841621. Add them to get 2.801585, square that to get 7.848879, multiply by the squared spread of 25.00, and divide by the squared effect of 1.00.
The result is 196.22 months. Months are whole things, and a fraction of a month buys nothing, so the requirement rounds up to 197. Dividing by twelve gives 16.4 years. Detecting a true monthly effect of 1.00 per cent against a spread of 5.00 per cent, at the conventional cut-offs, asks for 197 months of record. That sentence should be read exactly as written. The 197 is arithmetic about how much record a method needs before it can resolve a difference of this size, and it is not a rule about what anybody should do, wait for, expect or hold. The month count is the figure most likely to be misquoted as though it were such a rule.
Notice which of the three inputs did the damage. The spread enters squared, so noisy material is punishingly expensive. The effect enters squared in the denominator, so faint effects are expensive in the same violent way. The thresholdA cut-off fixed in advance against which a result is compared. The same rule then applies before anybody has seen the answer. The two numbers used here come from that convention and are built in notes covered separately. numbers are conventions rather than laws, and treating them as physics misreads them: change the convention and the 197 changes with it. The shape of the relationship does not change.
Detecting this effect reliably asks for 197 months. Is that a statement about the Nakshatra unit or about the record?
Commit to an answer before the panel underneath. If the effect halved from 1.00 to 0.50 per cent while the spread stayed at 5.00, the months asked for would
Why does halving the effect quadruple the record needed?
Because the effect enters the arithmetic squared, and halving a number quarters its square. Work it through on the exact figures rather than the rounded ones. At an effect of 1.00 per cent the requirement is 196.22 months. At 0.50 per cent it is 784.89 months. The only thing that changed was a division by one quarter, so dividing the second by the first gives 4.0000 exactly. Rounded up to whole months those two requirements read 197 and 785, whose ratio is 3.98 rather than 4.00, and that small discrepancy is the rounding and nothing else.
There is a plainer way to hold it. The resolving power of a record improves with the square root of its length, so halving the size of the effect to be resolved costs four times as much record. Four times the months buys twice the resolution, never four times. The square root rule is why faint effects are so expensive to establish and why so many of them sit unestablished for decades while being perfectly real the whole time.
Move the size of the effect, and watch two answers move in opposite directions.
One control moves: the true monthly effect. The spread stays at 5.00 per cent throughout, so the size of the effect is the only thing moving. Three things redraw. The upper bar is the months of record the arithmetic asks for, with a fixed marker at the fifty months these notes actually hold. The lower bar is the same effect compounded over twelve months. The strip beneath both shows the effect against one spread. The faint dashed outline marks the setting held before the last jump, so one press of the halving control shows the leap rather than describing it. The opening setting is the worked reading above, an effect of 1.00 per cent.
Two settings on the panel deserve the arithmetic written out. Press the halving control once from the opening setting and the months asked for go from 197 to 785. Press it again and they go to 3,140 months, or 261.7 years. Two halvings of the effect multiplied the exact requirement by sixteen, from 196.22 months to 3,139.55. The twelve month figure only fell from 12.68 per cent to 3.04 per cent. That asymmetry is the whole of the square law, and it is why an effect can go on being economically substantial long after it has stopped being establishable by anybody with a normal length of record.
What does that arithmetic mean for a record of fifty months?
Set the two lengths beside each other. The arithmetic asks for 197 months. The record holds fifty. Fifty divided by 197 is 0.2538, so the record is roughly a quarter of the length that would be asked for. The record is not long enough to settle the question, and that is a statement about the record rather than about the Nakshatra unit. The unit's true monthly mean is 1.00 per cent and it is positive, and it was positive the whole time the fifty months were being collected.
The pairing from the opening becomes useful here. The fifty month record cannot rule out a true monthly mean of zero. The true monthly mean is not zero. Both of those sentences are true at once, and the second is knowable at all only because the five outcomes were settled ahead of the record. The machinery that produces the first sentence, the null, the p-value and the significance threshold, is covered separately. The other half is knowing how big the effect actually is and how much record it would take to see it. Both facts together explain why the record stayed quiet.
The record holds fifty months and the arithmetic asks for 197. What follows about something this record fails to show?
What should be reported beside every result of a test?
Four lines, and an analyst, a lender or anybody reading a research note can ask for all four in one sentence. First, the size of the difference in its own units: 1.00 per cent a month here, and 12.68 per cent over twelve months. Second, the size relative to the spread: 0.2000. Third, the number of cases the result rests on: fifty months. Fourth, the smallest difference this record could have resolved at all. Turning the months arithmetic around, with fifty months and a spread of 5.00 per cent, the smallest effect resolvable is about 1.98 per cent a month, nearly twice the effect that is actually there.
The fourth line is the one almost nobody reports, and it is the line that decides whether a null resultA result where the difference being looked for did not come through as established. The machinery that produces one is covered separately; the word is used here only as the thing the fourth reporting line qualifies. carries any information at all. Without it, a note saying nothing was established is unreadable: the reader cannot tell whether the record was capable of establishing anything. With it, the same note becomes precise and useful. Saying that nothing above 1.98 per cent a month showed up in fifty months tells a reader exactly what has been ruled out and exactly what has not, and it does so without a single extra observation being collected.
An investor reading a research note, a lender reading a credit study or a household reading a claim about a scheme can all use the same habit. Ask what size of difference the material could have caught. If the answer is that it could only have caught something twice the size of the thing being discussed, the study has not failed and it has not succeeded either. The study has simply been asked a question it did not have the material to answer, and knowing that is worth more than any single point estimateA lone figure quoted as the estimate, carrying no width with it. The 0.50 per cent that the fifty months produced is one, and what belongs either side of it is covered separately. in the note.
Of the four reporting lines, which one is almost always missing, and what does its absence hide?
Two notes, both correct, and a reader who concludes one author is wrong
The first note reports that the unit's monthly effect is 0.2000 of a standard deviation and describes it as negligible. The second note reports that the same unit compounds to 12.68 per cent over twelve months and describes it as substantial. A reader holding both concludes that somebody has made a mistake and starts working out which author to trust. Neither note contains an error, both are describing the identical 1.00 per cent a month, and the disagreement the reader can see is entirely an artefact of each note reporting one half of the effect.
The cost is the argument itself. The discussion moves to which author is more careful, an unanswerable question that generates heat for weeks. Meanwhile nobody asks the one question that settles it: how much record did either of them have? Both notes were written on the same record, so in both cases the answer is fifty months against the 197 the arithmetic asks for. Once the record length is on the table the two notes stop disagreeing. One was describing detectability and the other magnitude, and the record was short for the first and irrelevant to the second.
A habit rather than a technique prevents the whole muddle. Report the effect in both units inside the same sentence, and add the length of record that would have been asked for to resolve it. A sentence reading that the effect is 0.2000 of the spread and 12.68 per cent over twelve months, on a fifty month record where 197 months would be asked for, cannot be misread in either direction, and it is no longer than the sentence that can.
Two notes describe the same 1.00 per cent a month, one calling it negligible and the other substantial. Which question settles the disagreement?
What sits behind these figures, given that no institution is named?
Every number here is arithmetic performed on an invention, and the invention itself is set out where a reader can inspect it. Inspectable arithmetic is a different kind of evidence from a citation, and in some ways a firmer one.
Where each figure comes from
Statistical method is taught here from mechanism on invented arithmetic, so no regulator, exchange, index or published series stands behind any number. Naming one would borrow authority the arithmetic neither needs nor can claim. Each row below can be rebuilt by hand from the definition in the first row, a harder check than a citation.
| What is used here | How it was produced | Outside document | What confirms it |
|---|---|---|---|
| The Nakshatra unit's five monthly outcomes and their weights | Written down as a definition inside these notes, before any record was drawn from it | None | The weights add to one, and everything else follows from the definition by arithmetic |
| The true monthly mean of 1.00 per cent and the true spread of 5.00 per cent | Computed from that definition rather than measured | None | Recomputed here, and again by a checking script kept beside these notes |
| The fifty month record | Invented counts of 5, 9, 25, 8 and 3 against the five outcomes | None | The five counts add to fifty |
| The share of the spread, the twelve month figure and the months asked for | Arithmetic on the two rows above, printed step by step in the blocks that use them | None | Every step appears above, so a reader can redo it without trusting it |
| The two conventional cut-offs inside the months arithmetic | Named as the widely taught pair and printed to six decimal places beside the working | Any standard statistics text | Both values are the pair printed in standard statistics texts, to six decimal places |
The bands that decide when a share of the spread gets called small or large are common property, taught in many texts and belonging to no single author. The bands differ between texts, so a borrowed threshold quoted as though it were settled does more harm than the convenience is worth.
The Nakshatra unit and its fifty month record are invented.
Educational material. Not advice on any investment, tax, budget or market position.
