Statistical Significance vs Economic Significance
Statistical significance asks whether a record this short could have thrown up a result this large by chance alone. Economic significance asks whether an effect of that size would matter to anybody. The two take different inputs and disagree routinely. The fifty month record of the Nakshatra unit cannot reject a mean of zero, at a p of about 0.4772. The same effect compounds to 12.68 per cent over twelve months.
Both of those sentences are true at the same time, about the same fifty months, and neither is a rounding or a softening of the other. Most people have been trained to treat one of the two as the verdict and the other as colour, so the pair is uncomfortable on first encounter. Each of the two is defined on its own below, in full, with its inputs listed. The contrast only makes sense once it is clear that the two lists are not the same list.
Two invented objects carry the arithmetic. The Nakshatra unit is a traded unitA stand in for any object whose price is quoted and moves about, which is all that is needed for one month to yield one figure. The arithmetic depends on the run of prices and not on the sort of object behind it. whose monthly changeOne month of price movement, expressed against the level the month started from. Something that opens at Rs 200/- and closes at Rs 212/- has changed by 6.00 per cent. is restricted to a list of exactly five outcomes. Minus 9.00 per cent comes up eight times in every hundred months, minus 4.00 per cent eighteen times, 1.00 per cent forty eight times, 6.00 per cent eighteen times and 11.00 per cent eight times. Draw fifty months from that generatorThe rule that produces the numbers, written down in advance. Here the rule is a short list of outcomes with a weight against each. Every month is one draw from that list. and the fifty month record appears. Every figure below is worked on that sample. In it, the five outcomes came up 5, 9, 25, 8 and 3 times respectively.
The unusual thing about these two objects is that the rule was written down before any month came out of it, so the true monthly effect of 1.00 per cent and the true spread of 5.00 per cent are known here rather than estimated. That almost never happens. With a real record there are the fifty months and nothing else, and what sits behind them can be argued over forever. Here the record and the truth can be set side by side and the distance between them measured exactly. Only a known truth lets anybody say with a straight face that a real effect went undetected.
What is statistical significance, taken on its own?
Statistical significance is the answer to one narrow question: if there were genuinely nothing there, how often would a record like this one throw up a result at least this far from nothing, purely by luck? If the answer is often, the record has shown nothing that shuffling could not have produced. If the answer is rarely, the record is hard to explain away as luck.
Start with a shopkeeper rather than a record of prices. The shopkeeper suspects Tuesdays are busier than Wednesdays, counts three of each, and Tuesdays come out ahead. Nobody would take that seriously, and everybody knows why without being taught: three days of anything wobbles enough to produce that gap by accident. Counting three hundred of each and finding the same size of gap is a different sentence entirely. Nothing about Tuesdays changed between the two counts. The amount of luck available to produce the gap changed.
Three things decide whether a result comes out statistically significant, and they are the size of the effect, the spreadHow far apart the individual cases sit from each other, boiled down to one number. How that number is built, and the several ways of building it, are covered separately. of the individual cases, and the number of cases. A larger effect is harder to fake. A tighter spread means less wobble to fake it with. Averages of many things settle down, so more cases mean the wobble in the average shrinks. The three inputs go into one arithmetic and one number comes out.
Now notice something about that list that almost nobody says out loud: two of those three inputs are properties of the record rather than of the world. The spread of the cases is partly the world and partly what was chosen for measurement. The number of cases is entirely the record keeper's own. Nothing about the Nakshatra unit changes when somebody keeps the ledger open for another two hundred months. The unit does what it always did. Only the record grows. So a statistical verdict is a joint statement about the world and about how much of it somebody bothered to write down, and it never separates the two.
The shopkeeper counts three Tuesdays and three Wednesdays and finds Tuesdays busier. Then she counts three hundred of each and finds the same size of gap. What changed between the two counts?
What is economic significance, taken on its own?
Economic significance is the answer to a different narrow question: if a difference of this size is real, would it be worth anything in units somebody actually counts? Rupees, months, seconds at a counter, litres of milk. Not units of luck. Units of consequence.
Economic significance is the reading a household does instinctively and a report often forgets. One per cent a month sounds like nothing. Say it out loud and it sounds like a rounding error. But compoundingApplying a change to whatever the last change left behind, rather than to the starting amount every time. Each round works on a slightly larger base, so twelve rounds of one per cent do not add up to twelve per cent. is what turns one per cent a month into something visible. Rs 10,00,000/- with 1.00 per cent applied twelve times over, each time to whatever the last round left behind, finishes at Rs 11,26,825/-. The gap is Rs 1,26,825/- over one year, from a monthly figure most people would have called noise. In percentage terms it is 12.68 per cent, and that figure runs through this whole guide.
Two things decide whether a difference is economically significant: the size of the effect, and the setting it is applied in. The same 1.00 per cent a month is one figure when it runs on Rs 10,000/- and another when it runs on Rs 10,00,000/-, and a different one again over one month than over twelve. Size and setting. The list ends there.
The number of cases does not appear in that list at all, and the whole contrast rests on that absence. It is not there. How many months of record somebody happened to keep has no bearing whatsoever on what 1.00 per cent a month does to Rs 10,00,000/- over a year. The arithmetic of compounding does not consult the ledger. A hundred years of record and a single month of record leave 12.68 per cent exactly where it was.
Which of the two kinds of significance depends on how many cases somebody happened to collect?
Where exactly do the two part company?
One line does it. The two kinds of significance differ because one of them depends on how much record there is and the other does not. Everything else in this guide is a consequence of that single asymmetry.
Laid next to each other, the two lists overlap in exactly one place. Both consult the size of the effect. After that they go their separate ways. Statistical significance reaches for the spread and the count. Economic significance reaches for the setting. The two are not two opinions about the same calculation but two calculations that happen to share one ingredient.
The two can disagree without either of them being wrong, so asking which is the real verdict is asking the wrong question. There is no school of thought here to pick a side in. Different numbers fed into different arithmetic give different answers, and that is not a controversy, it is subtraction. A reader who demands to know which one counts is really asking somebody else to make the judgement for them, and no arithmetic in this guide or anywhere else will do that.
The practical version is simpler still. A statistical verdict can be changed by collecting more. An economic reading cannot. When a result flips from nothing to something and the only thing that changed was the length of the record, nothing about the world moved. When a result flips because the effect turned out to be twice the size, something did.
What does the fifty month record say on each of the two?
The record itself now goes through both, slowly.
On the statistical side, the fifty months average 0.50 per cent. The spread across those months is 4.9744 per cent. Divided by the square root of fifty, that spread gives a standard errorHow much the average of a record wobbles from record to record, purely because a record is a handful of cases and not the whole story. The standard error shrinks as the count grows, and how it is built is covered separately. of 0.7035 per cent. Divide 0.50 by 0.7035 and the statistic is 0.7107. Under the normal shapeThe single humped, symmetric shape that turns up whenever many small independent wobbles are averaged together. The normal shape judges how surprising a statistic is, and how it is built up is covered separately. used to judge it, a statistic that far from zero or further turns up about 47.72 times in every hundred records when nothing is there at all. So the record cannot reject a mean of zero. Not narrowly. By a distance.
On the economic side, the true monthly effect behind those months is 1.00 per cent, and that is known here only because the rule was written down first. As a share of the true spread of 5.00 per cent it is 0.20. Compounded through twelve months it is 12.68 per cent. On Rs 10,00,000/- that is Rs 1,26,825/- across a year. Nobody looking at that figure in rupees would call it nothing.
Both statements are true, and neither is a rounding of the other: the test says the record cannot see the effect, and the arithmetic says the effect is large. The honest one line summary is not that the effect is absent and not that the test failed. The summary is that fifty months is too short a record to see an effect of a size that would matter. Sitting between the two figures is a third one, computed separately, that reconciles them: resolving an effect this size reliably would take about 197 months, and the record holds 50.
| What the reader is asking | Could this be luck? | Would this size matter? |
|---|---|---|
| The effect it consults | 1.00 per cent a month | 1.00 per cent a month |
| The spread it consults | 4.9744 per cent | does not consult it |
| The number of cases it consults | 50 months | does not consult it |
| The setting it consults | does not consult it | twelve months of compounding |
| What it computes | statistic 0.7107 | 12.68 per cent |
| What comes out | p value 0.4772 | Rs 1,26,825/- on Rs 10,00,000/- |
| The verdict | cannot reject a mean of zero | large in units somebody counts |
Read the two columns downwards and count how many rows they share. One. The size of the effect, at the top. Every other row has one column answering and the other saying it never asked. The two columns are not two readings of one calculation but two calculations.
The record gives a p value of about 0.4772 and the same effect compounds to 12.68 per cent over twelve months. Which of the two statements is the real one?
A prediction is worth settling on before the panel below is touched. As the record gets longer while its mean stays 0.50 per cent and its spread stays 4.9744 per cent, what happens to the p value and what happens to the twelve month figure?
Stretch the record and watch one verdict march while the other refuses to move.
One control, the number of months in the record, from 12 up to 500. The record's mean of 0.50 per cent and its spread of 4.9744 per cent are held exactly where they are, so the only thing changing is how much of it there is. The default of 50 months is the published record and reads a statistic of 0.7107 and a p value of 0.4772. Every setting visited leaves a lime dot behind on the curve, so the shape of the relationship is painted in as the control is dragged.
How can a real effect be statistically invisible?
A real effect going statistically invisible is the first of the two failure directions, and the fifty month record demonstrates it. There is a genuine 1.00 per cent a month sitting behind those months. The test cannot see it.
The third input is the reason. Statistical significance rises with the number of cases, and a short record fails to reject almost everything put to it. With fifty months the wobble in the average is 0.7035 per cent, which is larger than the whole effect being looked for. The signal is smaller than the noise around the measurement of it. Nothing is broken. The instrument simply is not fine enough for the thing it is pointed at.
A household version helps. The question is whether switching vegetable sellers saves money. One week of bills is compared against one week of bills, and the new seller comes out Rs 40/- cheaper. Weekly grocery bills swing by hundreds of rupees depending on what was in season and who came to dinner. One week against one week cannot separate Rs 40/- of saving from ordinary variation, and that is true even if the Rs 40/- is completely real and repeats every single week for a lifetime. Over a year that Rs 40/- is Rs 2,080/-. Real, worth having, and invisible to a one week comparison.
So the consequence for reading a report is sharp: a result of nothing found, taken from a short record, is not evidence that there is nothing there. It is a statement about what that record was capable of resolving. The fifty month record would have needed roughly 197 months to resolve an effect of 1.00 per cent a month reliably, and 381 months before its own reading would drop under the 0.05 line. Where those two figures come from is covered separately. The shape of the situation is what matters. The record was asked a question it did not have the length to answer. A test always answers, so the record answered anyway.
A colleague reports that no effect was found in fifty months of a record. What two things should be asked for before that is read as evidence that nothing is there?
How can an effect too small to matter come out significant?
Now run the same machinery the other way. The fix for one direction quietly creates the other. A reader who learns to distrust null results will start treating every significant finding as important, and that is the second error wearing the first one's clothes.
Take a completely different record, invented for this block. A bank chain tries two counter layouts and measures how long each visit takes. Call it the Tarang counter record. Some people cash a cheque and some people open an account, so the average visit runs 360 seconds and the spread across visits is 210 seconds. Across 2,00,000 visits, one layout comes out 1.2 seconds faster than the other.
Put that through the arithmetic. The standard error is 210 divided by the square root of 2,00,000, or 0.4696 seconds. The statistic is 1.2 divided by 0.4696, or 2.5555, and the p value is 0.0106. Comfortably under 0.05. Statistically significant, and nobody could argue with the arithmetic.
The effect itself is 1.2 seconds off a 360 second visit. The difference is one third of one per cent of the wait. Nobody standing in that queue would notice it even after being told in advance to watch for it. Hold an effect tiny, raise the number of cases far enough, and any difference that is not exactly zero will eventually cross the threshold. The wobble in the estimate keeps shrinking while the effect does not. That same 1.2 seconds first crosses the 0.05 line at about 1,17,645 visits. Below that count it is not significant. Above it, it is. The seconds never changed.
On a very large record, statistical significance stops carrying information about whether anything matters. The verdict becomes a statement about how much data somebody collected. That is not a criticism of the arithmetic, which did exactly what it was asked. The warning is about what the word significant is doing in a sentence. Significant has never meant important. Significant means hard to explain as luck. On a record of two lakh cases, almost nothing is hard to explain as luck, so almost everything passes.
A study of two lakh cases reports a statistically significant difference between two options. What is the first thing to ask about it?
Who pays for each of the two errors?
The two errors are not symmetric in their consequences, so the two directions deserve separating rather than being folded into one warning about being careful.
The first error discards something real, and the person who pays is whoever needed to know it existed. A note says nothing was found. The finding goes in a drawer. Nobody looks again. Looking again costs money and somebody already looked. The 1.00 per cent a month carries on being 1.00 per cent a month, unrecorded by anybody, and the cost is silent: it is the thing nobody ever learned. Silent costs are the hardest to argue for a budget against, and the difficulty is exactly why this error survives so well.
The second error dresses up something trivial, and the person who pays is whoever acts on it. A significant result is louder than a null one. A significant result gets a slide. A significant result gets a rebuild of the counter layout, at real expense, to recover a third of one per cent of a wait. Here the cost is not silent at all, it is an invoice, and it still gets paid because the word significant was doing work in the room that it was never entitled to do.
A note that gives only one of the two figures leaves the reader unable to tell which of the two errors is being made. That is the practical reason to care. A p value on its own does not say whether the reader is in the corner where something real is being thrown away or the corner where something trivial is being polished. A p value names which half of the picture the reader is standing in and says nothing at all about the other half.
Name the two failure directions and say who pays in each.
What goes into a note so a reader can judge both?
The distinction becomes usable inside a note. An analyst reading somebody else's note, a lender reading a portfolio review, a household reading a comparison of two plans, all face the same problem: the note tells them one thing and they need two. Four lines fix it, and three of them are usually already there.
Line one is the size of the effect in its own units. Here, 0.50 per cent a month measured, against a true 1.00 per cent a month behind it. Line two is the same size translated into units somebody counts. For a monthly figure that usually means a year of compounding: 12.68 per cent, or Rs 1,26,825/- on Rs 10,00,000/-. Line three is the number of cases, here 50 months. Line four is the one almost every note leaves blank, and it is the smallest effect that record could have detected reliably. For fifty months at this spread, that is about 1.97 per cent a month, or 26.39 per cent over twelve months.
Line four does the most work. Line four says that this particular record was only ever capable of reliably flagging an effect roughly twice the size of the one that was actually there. Once that line is in the note, a null result stops sounding like an answer and starts sounding like a description of the instrument. A p value on its own answers neither of the two questions, and a note that carries only a p value has told the reader almost nothing.
What this looks like at a desk
An analyst who reads notes for a living builds a reflex out of those four lines. Handed a null result, the first question is how much record it rests on and what the smallest detectable effect was. A null from a short record is a shrug rather than an answer. Handed a significant result on an enormous record, the first question is the size in rupees or seconds or litres. On a record that large the word significant has almost stopped filtering anything. A household comparing two loan offers does the identical thing without the vocabulary: how much is the difference in rupees a month, and did somebody check enough months to be sure. The pair of questions is the same, asked in a kitchen.
Which single input do the two kinds of significance share, and what does that say about the rest of the inputs?
Detecting this effect reliably would need about 197 months and the record holds 50. Which of the two questions does that comparison answer?
The reading that goes wrong, in both directions
A note reports that no monthly effect was found in fifty months of the Nakshatra unit, and a reader concludes there is nothing there. Every word in the note is defensible. The statistic really is 0.7107, the p value really is 0.4772, and the record really cannot reject a mean of zero. The reader has converted a statement about the record into a statement about the world, and the two are not the same statement. The true effect is 1.00 per cent a month, or 12.68 per cent over twelve months, and this record would have needed roughly 197 months to resolve it. The note is correct and the conclusion drawn from it is not supported by it.
Now run it the other way. The cure for the first error manufactures the second. A reader who has learned to distrust null results starts treating every significant finding as important. Hand that reader the counter study, 1.2 seconds across 2,00,000 visits at a p value of 0.0106, and they will read significant as meaningful and sign off a rebuild for one third of one per cent of a wait. The cost in the first case is a real effect discarded, unnoticed, with nobody to complain. The cost in the second is money spent on a difference nobody in the queue can feel.
The fix is the same in both directions and it is not a rule of thumb about thresholds. Read the size and the count together, and never either one alone. A p value with no size beside it cannot show whether something real is being thrown away. A size with no count beside it cannot show whether the record was long enough to have seen it. Two figures, always, in both directions.
Adjacent subjects. The machinery of the test itself, the position being tested, how a p value is built and where the 0.05 line came from, is covered separately. The arithmetic of effect size, including how a share of the spread is computed and where the figure of 197 months comes from, is covered separately too. How the standard error is built out of a spread and a count is covered separately again. Drawing a line through a set of points, turning a point estimateOne figure put forward as the whole answer, an average of 0.50 per cent for instance, rather than a stretch of values the record cannot separate. Which of the two a question needs is covered separately. into a parameter of a model, and forecasting a month are each covered separately. The worth of an effect of any size to a particular reader is a judgement rather than arithmetic.
Which of these numbers rests on a source somebody else keeps current?
None of them does, and the two verdicts rest on different kinds of nothing. The statistical reading rests on a convention that people adopted and nobody issued, and the economic reading rests on a rule that was typed out for teaching. Neither has an institution standing behind it, neither can go stale, and neither could be settled by producing a document. Every figure below is set out with how it was made and what could possibly overturn it.
| Figure in this guide | How it was made | What could overturn it |
|---|---|---|
| The 0.05 line | A habit that spread through practice rather than a rule anybody issued | Nobody issued it, so nobody can withdraw it. Any other line would be equally unofficial |
| The true monthly effect of 1.00 per cent and the true spread of 5.00 per cent | Weighted arithmetic on five values typed out before any month came out of them | Only an arithmetic slip. Weight the five values and total them to check |
| The statistic of 0.7107 and the p value of 0.4772 | 0.50 divided by 4.9744 over the square root of 50, then read off the normal shape | Only an arithmetic slip. The division gives 0.7035 for the wobble |
| The 12.68 per cent and the Rs 1,26,825/- on Rs 10,00,000/- | 1.00 per cent applied twelve times, each round on what the last one left | Only an arithmetic slip. Twelve rounds, not twelve additions |
| The 381 months where this record crosses, and the 197 it would need | Recomputed here from the record's own mean and spread | Only an arithmetic slip, and the panel above redoes it at every setting |
| The Tarang counter record and its p value of 0.0106 | 1.2 seconds against a spread of 210 across 2,00,000 invented visits | Only an arithmetic slip. The wobble is 210 over the square root of 2,00,000 |
The Nakshatra unit, the fifty month record and the Tarang counter record are invented.
Educational material. Not advice on any investment, tax, budget or market position.
