Resampling and the Bootstrap: Using Your Own Data to Estimate Uncertainty
The bootstrap estimates how much a number moves about by drawing new sets of fifty months out of the fifty already held, with replacement, and watching the answer shift. Two thousand such draws say a fifty month mean typically sits 0.7015 per cent away from what it estimates. The formula says 0.7035, the truth is 0.7071, and the bootstrap was told neither.
Every reading below is manufactured by one arithmetic rule, x becomes 16,807 times x with the remainder taken on 2,14,74,83,647, started at a stated number and turned over in the browser as the reading goes on. Redone on paper, it lands in the same place.
Two things arrive already settled and neither is rebuilt here. The first is the Nakshatra unit, an invented thing that carries a price and nothing else, whose monthly change is the percentage by which its price ended a month away from where it started. The monthly change takes exactly five values, minus 9.00, minus 4.00, 1.00, 6.00 and 11.00 per cent, with weightsNumbers attached to a list of outcomes saying how much of the whole each one accounts for, adding to one across the list. 0.08, 0.18, 0.48, 0.18 and 0.08 fixed on it by whoever invented it. The second is the fifty month record, fifty months of that unit holding those five values 5, 9, 25, 8 and 3 times, whose mean is 0.5000 per cent, whose standard deviation is 4.9744 per cent, and whose standard errorHow far a calculated figure typically misses the quantity it was calculated to stand in for. of 0.7035 per cent was worked out from a formula earlier. A swap pass started at the number 4321 puts the record in a stated order, rearranging the fifty months and altering none of them.
The third thing settled in advance is the truth itself. Because the weights on the Nakshatra unit were written down rather than measured, its true mean of 1.00 per cent and its true standard deviation of 5.00 per cent are known exactly. So every estimate below gets printed next to the answer it was trying to reach. No study on a real record can ever do that, and it is the reason a method that sounds like a conjuring trick can be caught in the act here.
What is resampling, and what is it for?
Begin with something concrete. Suppose a walk to the station has been timed on fifty mornings and the fifty times average out at eighteen minutes. Eighteen minutes is the answer, but it is a point estimateA lone number handed over as the answer, carrying no width and no range beside it. and it carries no width. The honest follow up question is how different that eighteen would have been had fate handed over a different fifty mornings. Getting another fifty mornings means waiting two months, and by then the traffic has changed. Walking cannot settle it.
Resampling is the answer to that question, and it is startling on first hearing. The fresh sets of mornings are built out of the fifty mornings already in hand. Not out of a description of what mornings are like, not out of a shape somebody drew, but out of the actual timings sitting in the notebook. The source of the new data is the whole difference between resampling and drawing from a stated shape: one method builds its new data from the record in hand, the other builds it from an assumption about where the record came from.
The idea gets oversold here, so the assumptions are worth following carefully. Resampling has not removed an assumption. Drawing from a stated shape assumes the shape is right. Resampling assumes instead that the fifty mornings are a fair stand in for mornings in general. Neither assumption is free and neither is checkable from inside the method. One has been swapped for the other, and the reason to swap is that on many days the record is easier to defend than an invented shape.
Resampling builds new sets of data out of data already held, rather than out of a stated shape. What has that traded away, and what has it gained?
What is the bootstrap, and how is one actually run?
The bootstrap answers that question in four instructions. Nothing sits underneath the four. One reading of them is the whole method.
- Take the record in hand. Here that is the fifty month record, fifty monthly changes of the Nakshatra unit.
- Draw fifty months out of it at random, with replacement, meaning a month already drawn goes straight back in the bag and can be drawn again. That set of fifty is one resample.
- Compute whatever number is actually wanted from that resample. Here it is the mean of the fifty drawn months.
- Do it two thousand times and look at how far apart the two thousand answers landed.
Look closely at the first resample from the number 2468. It is nothing like a copy of the record. Eighteen of the fifty months never got drawn at all. Twenty were drawn once, eight were drawn twice, two were drawn three times and two were drawn four times. Its mean came out at 0.2000 per cent against the record's own 0.5000, and eighteen of its fifty drawn months were falling ones, a share of 36.0000 per cent against the record's 28.0000. One resample is a set of fifty months that the same source could plausibly have produced, and it is emphatically not the fifty months the run started with.
Now run the fourth instruction. Two thousand resamples give two thousand means. Each resample got a different mix of the fifty months, so the two thousand means are not all the same. The spreadHow widely a set of figures sits apart from one another, boiled down to a single number. of those two thousand numbers is the answer, and no formula was consulted at any point to produce it. A draw, a computation, and a look at how much the computed figure jumped about are the entire method.
Why does the bootstrap draw with replacement?
Drawing with replacement is the step readers refuse. Refusing it is the right instinct, so the step is worth working through rather than accepting. Consider what happens without replacement. There are fifty months. Fifty of them are drawn without any being put back. The result can only be one thing: all fifty months, every time, in a different order. A mean does not know what order its inputs arrived in, so the mean of the fifty is 0.5000 per cent before the shuffle and 0.5000 per cent afterwards. Two thousand repetitions give two thousand identical answers and a spread of exactly nothing.
So the replacement is the entire mechanism, not a technicality tucked into the instructions. Take it out and the bootstrap reduces to shuffling a deck and reporting that the deck still has fifty two cards. Putting each month back is what allows a resample to over represent some months and miss others, which is exactly what a fresh fifty months from the same source would have done. The variation across the two thousand answers is a stand in for the variation that two thousand genuinely fresh records would have shown, and it was obtained without collecting a single extra month.
Fifty months are drawn out of fifty without replacement and the mean is taken. What comes back, and what does that establish?
What did two thousand resamples of the fifty month record give?
The run is fixed and repeatable: the fifty month record in its stated order, the arithmetic rule started at the number 2468, two thousand resamples of fifty months each. Every figure in the table below came out of that one run, and starting the rule at 2468 again reproduces all of them exactly.
| What was measured | Reading | What it is |
|---|---|---|
| The record's own mean | 0.5000 per cent | The single number the record hands over before any resampling starts |
| The average of the two thousand resample means | 0.4863 per cent | Where the resampling settles, which is on the record and not on anything else |
| The spread of the two thousand resample means | 0.7015 per cent | The bootstrap standard error, obtained without a formula |
| The lowest and highest resample mean seen | minus 2.0000 to 3.1000 per cent | How far a fifty month mean wandered in two thousand tries |
The third row is doing more work than it looks. The record on its own says one thing: the mean of these fifty months is 0.5000 per cent. The record has no way of saying how seriously to take its own mean. Put together, the two say that a fifty month mean typically lands about 0.70 per cent away from whatever it is estimating. The record could not make that statement on its own, and nobody had to derive it. The record supplied the number; the resampling supplied the width around it.
The record's mean is 0.5000 per cent, and the bootstrap says a fifty month mean typically sits about 0.70 per cent from what it estimates. What does that pair establish that the record alone did not?
How good is that answer, when the truth can be worked out exactly?
Here is where the invented setting pays for itself. The Nakshatra unit's true standard deviation of 5.00 per cent was written down rather than measured, so the true standard error of a fifty month mean is 5.00 divided by the square root of fifty. The division comes to 0.7071 per cent exactly, and 0.7071 is the quantity every method here is chasing. Four different answers to it can now be laid side by side.
| The answer | Reading | Where it came from |
|---|---|---|
| The bootstrap, two thousand resamples | 0.7015 per cent | Counting how far apart two thousand resample means landed |
| The bootstrap's own limit | 0.6964 per cent | Where the bootstrap reading settles as the resamples are pushed up without end |
| The formula from the record | 0.7035 per cent | The record's standard deviation divided by the square root of fifty |
| The truth | 0.7071 per cent | 5.00 per cent divided by the square root of fifty, computable because 5.00 was stated |
All four sit inside a band 0.0107 per cent wide. Handed fifty numbers and an instruction to draw from them, the bootstrap landed within six thousandths of a per cent of a truth it was never told and could not have derived. The whole case for the method rests on that single fact. The bootstrap did not need to know that a standard error involves a square root. Nor did it need to know that the divisor is fifty rather than forty nine. It counted.
Are 0.6964 and 0.7035 the same number?
The two are not the same number, and both read 0.70 when rounded to two decimals. A rounding artefact gets taught as a fact in exactly that way. The bootstrap's limit divides the record's squared departures by fifty. The formula divides the same squared departures by forty nine. One denominatorThe count written underneath a division, which settles what the result is an average across. is fifty and the other is forty nine, and nothing else differs between them. The two sums compute different quantities rather than wobble around one.
Rounding both sides of a comparison and then declaring them equal proves nothing except that the display carries two decimals. This is why every reading in the table above is printed to four places. At two places the four answers collapse into 0.70, 0.70, 0.70 and 0.71 and the interesting structure vanishes. A reader would conclude that the bootstrap merely reproduced the formula, and the bootstrap did no such thing.
The bootstrap's limit is 0.6964 per cent and the formula gives 0.7035. Both read 0.70 at two decimals. Are they the same number, and what differs?
What does the interval add, and why does it cost almost nothing?
Once the two thousand resample means are sitting in a list, one more thing falls out for free. Put the list in order and read off the value at the two and a half percentileA list put in order and walked along: the value reached once a stated share of the list has gone past is the percentile for that share. and the value at the ninety seven and a half. On this run those are minus 0.9000 per cent and 1.8000 per cent, and the stretch between them holds the middle ninety five per cent of everything the resampling produced. No extra computation was required. The numbers were already there; they only had to be sorted.
Now look at what that stretch contains. The true mean of the Nakshatra unit is 1.00 per cent, and 1.00 sits comfortably inside minus 0.9000 to 1.8000. Meanwhile the point estimate this record handed over was 0.5000 per cent. The truth halved is a poor figure to hand anybody on its own. The single number was badly off and the interval was still right, and that is the entire reason to report an interval instead of a number. An interval is not a decorated point estimate. An interval is a different claim, and one this record can actually support.
The point estimate was 0.5000 per cent, the truth is 1.00, and the interval ran minus 0.9000 to 1.8000. Was the estimate wrong, was the interval wrong, or neither?
The fifty month record holds fourteen falling months out of fifty. If the falling share is bootstrapped two thousand times, what will those resamples centre on?
Bootstrap vs Monte Carlo Simulation: what is each method blind to?
Both methods produce thousands of made up sets of data and count what happens in them, so from a distance they look like the same thing done twice. The two methods are not the same, and the difference is one line. Monte Carlo simulation draws from a stated shape. The bootstrap draws from the record already in hand. Everything else follows from that, including what neither method can be argued out of.
Point both at one question and watch. The question is how often a month falls. The fifty month record holds fourteen falling months out of fifty, a share of 28.0000 per cent. The truth, from the stated weights on the Nakshatra unit, is 26.0000 per cent. Nothing is wrong with the record; fifty months is simply fifty months, and this particular fifty happened to catch one more falling month than the truth would hand over on average.
Bootstrap the falling share two thousand times and the resamples average 28.1800 per cent. Draw one lakh single months from the stated shape instead, starting the rule at 12345, and they fall 25.9750 per cent of the time. The bootstrap settles on the record and the drawn shape settles on the truth, and the reason is not that one method is cleverer: it is that every bootstrap resample is assembled out of those same fourteen falling months and there is nowhere else for it to get information from.
Turn it round and the drawn shape is the one exposed. Suppose whoever wrote those five weights had made them slightly wrong. Monte Carlo simulation would then converge, beautifully and with a very tight spread, on an answer inherited from the mistake. The drawn months are made out of the mistake, so no number of them would ever reveal it. Each method is blind to exactly one thing and it is not the same thing: the bootstrap inherits every way in which the record is not representativeStanding in for the wider group fairly, rather than by whatever happened to get caught., the drawn shape inherits every error in the shape, and choosing between them is choosing which blindness is easier to live with.
A bootstrap of the falling share centres on 28.1800 per cent and the truth is 26.0000. Would two lakh resamples instead of two thousand close that gap?
What does one record look like with both methods and the truth beside it?
Everything above, collected. One record, two methods pointed at it, and a true answer printed against each estimate because the weights that generate the truth were written down rather than measured. Read the last column of every row: it is the column that no study on a real record ever gets to print.
| Question | Method | Estimate | The truth |
|---|---|---|---|
| How far does a fifty month mean sit from what it estimates? | The bootstrap, two thousand resamples | 0.7015 per cent | 0.7071 per cent |
| The same question | The formula on the record | 0.7035 per cent | 0.7071 per cent |
| The same question | The bootstrap's own limit | 0.6964 per cent | 0.7071 per cent |
| What is the mean monthly change? | The record, as a single number | 0.5000 per cent | 1.0000 per cent |
| The same question, as a range | The bootstrap interval | minus 0.9000 to 1.8000 | 1.0000 per cent |
| How often does a month fall? | The bootstrap, two thousand resamples | 28.1800 per cent | 26.0000 per cent |
| The same question | Drawing one lakh months from the stated shape | 25.9750 per cent | 26.0000 per cent |
Two rows in that table are worth pausing on together. The bootstrap wins the first question and loses the last one, and it is the same method with the same seed on the same record both times. Between those two rows the method's quality did not change. The change is in whether the quantity being estimated depends on how representative the fifty months happened to be. The spread of the record is close to the spread of the source, so the bootstrap's answer about spread is close. The falling share of the record is two points above the source, so the bootstrap's answer about falling months is two points above.
The resample count is about to be stepped down to 50 and back up to 2,000. Does the bootstrap standard error walk steadily towards the truth, or does it wander?
Move the resample count and watch the shape firm up.
One control, the number of resamples, stepping through 50, 100, 250, 500, 1,000 and 2,000. The record never changes and the arithmetic rule is started at 2468 every time, so the settings are nested: the first fifty resamples are the first fifty of the hundred, and moving the control up refines one single run rather than starting a new one. The histogram redraws, the interval bar under it redraws, and the reading slides along the fine scale below while two marks stay welded in place. At the opening setting of 2,000 resamples the panel reads a bootstrap standard error of 0.7015 per cent against a formula of 0.7035 and a true 0.7071. The tables above print the same three figures.
Educational illustration, worked start to finish on figures invented for it. The standard deviation of 5.00 per cent was written down rather than measured, so the true standard error of 0.7071 per cent can be printed beside every estimate. Outside a teaching setting no such column can be filled in. The arithmetic rule is started at 2468 at every setting, so a setting returned to hands over the reading it handed over before. Fifty months is a short record, and the shortness of the record is what fixes how far any reading here reaches.
How many resamples are enough, and what will more resamples never fix?
An analyst with fifty months of anything faces this question the moment the method is suggested, and it usually arrives as a budget question: how many resamples are worth paying for? The panel above answers the first half. The bootstrap standard error on the fifty month record reads 0.8095 per cent at fifty resamples, 0.7969 at a hundred, 0.7164 at two hundred and fifty, 0.7249 at five hundred, 0.7190 at a thousand and 0.7015 at two thousand.
Look at the third and fourth of those and the pattern breaks: two hundred and fifty resamples reads 0.7164 and five hundred reads 0.7249, so the higher setting sits further from the truth than the lower one. That is not an error in the run and it will not be smoothed away by picking nicer settings. The reading wanders towards the answer rather than marching towards it. A reading taken at fifty resamples is not a rough draft of the right answer. A fifty resample reading carries a great deal of its own noise, and the noise happens to be pointing somewhere.
So the practical answer on count is unglamorous: enough that the reading has stopped moving noticeably, and then some more. Resamples are cheap, and a run of two thousand on fifty months finishes before the instruction has been read through. The far more important half of the question is what all that counting is buying. The counting buys a smaller amount of one particular kind of noise, the noise the resampling itself introduced by drawing rather than enumerating. More resamples do nothing whatever about whether the fifty months in hand were representative in the first place, and that second thing is usually the larger of the two by a wide margin.
A lender looking at fifty months of a small shop's takings meets both halves at once. The bootstrap will tell them, honestly and cheaply, how much the average monthly takings figure would jump about if the same shop had handed over a different fifty months. The bootstrap will say nothing at all about whether those fifty months happened to include a festival season the next fifty will not, and pushing the resample count from two thousand to two lakh will not add a syllable on that point. The figure the resampling is heading towards, however long it runs, is a fact about the fifty months and not a fact about the shop.
State it in one sentence. What does the bootstrap measure, and what does it not?
The mistake this method invites, and what it costs
An analyst bootstraps the fifty month record for the falling share, gets 28.1800 per cent with a tight spread around it, and reports that the chance of a falling month is 28.18 per cent, adding that two thousand resamples make the figure reliable. The true answer is 26.0000 per cent. Running two lakh resamples instead would have made the report tighter and not one step closer: the average falling share across two lakh resamples comes out at 28.0008 per cent, and the standard error of that average falls from 0.14285 to 0.01421, a tenfold improvement in a quantity that was never the source of the error.
The error came from the fifty months. The fifty months contain fourteen falling ones where the truth would deliver thirteen on average, and every one of the two lakh resamples was assembled out of those same fourteen. The resampling improved its own sampling noise. The two point gap between the record and the world stayed exactly where it was.
The cost is a number carrying a confidence the method cannot supply, and a reader downstream who reads a tight spread as accuracy. The fix is a sentence rather than a calculation, and it belongs in the write up next to the figure: the bootstrap measures how much the answer moves, never how far the record sits from the world.
Four things are deliberately left alone here. What a standard error is and how its formula is built is covered separately and taken as read here. Drawing from a stated shape is covered separately too, and is used above only for the contrast. Where the drawn numbers themselves come from, and how many resamples a stated tolerance would cost, are both covered separately as well.
Where can these figures be checked?
Each block of figures has a recipe behind it: the rule, the number it was started at, and how many times it was turned. Set those three the same way and the same readings appear on any machine. A document has to be trusted and a recipe can be run, so a recipe is the weaker authority and the stronger check.
| Source | Document | Site |
|---|---|---|
| The arithmetic in this guide | The stated rule started at 2468, two thousand resamples of fifty months each | None, because the arithmetic is worked out in full above |
| The five monthly values and the weights on them | The invented Nakshatra unit, fixed earlier in these notes and carried in unaltered | None, invented |
| The fifty month record and its reading order | Fifty months holding the five values 5, 9, 25, 8 and 3 times, ordered by a swap pass started at 4321 | None, invented |
| The truth each estimate is printed against | The stated true mean of 1.00 per cent and true standard deviation of 5.00 per cent | None, invented |
The Nakshatra unit, the fifty month record and the Nakshatra ladder are invented.
Educational material. Not advice on any investment, tax, budget or market position.
