MSE, RMSE, MAE and MAPE: Four Error Measures Compared
All four read the same ten misses and disagree only about how to add them up. Across the ten months where the Nakshatra unit is paired with the Vasant unit the mean squared error is 21.80, its square root is 4.6690, the mean absolute error is 3.60, and the mean absolute percentage error is 110.81 per cent. One unchanged straight rule, one unchanged column of misses, four different answers.
Two invented columns run down a ten month record. The Nakshatra unit is the input and the Vasant unit is the outcome, and earlier reading already put a straight rule through the ten pairs and worked out how far that rule sits from the truth in each month. Each of those distances is a miss. The ten misses are what all four measures read. Where the misses came from is covered under the fitting of the rule. The ten misses now have to be squashed into one number for the top of a report.
There is no neutral way to do that. Four measures are in common use. All four take exactly the same ten misses as their input. None of them touches the rule that produced those misses. And yet one of them announces 21.80, one announces 4.6690, one announces 3.60, and one announces 110.81 per cent. Each of those is a correct answer to a slightly different question, and the differences between the questions are decisions somebody made about what counts as bad.
What do all four measures actually read?
Here are the ten misses, in time orderThe order the months actually arrived in, first to tenth. Records like this one are cut along that order when a rule is tested on months it has not seen, so shuffling the rows would destroy something. and in percentage pointsThe unit that results from subtracting one per cent figure from another. A move from 4 per cent to 7 per cent is three percentage points, and calling it three per cent would mean something quite different.: a miss of 1 in month 1, 9 in month 2, 8 in month 3, nothing at all in month 4, minus 5 in month 5, nothing again in month 6, minus 3 in months 7 and 8, minus 2 in month 9 and minus 5 in month 10. A positive miss means the rule came in under the outcome and a negative one means it came in over. The ten add to exactly nothing, a property of how the rule was fitted rather than a coincidence.
All four measures start from that same row of ten. The rule does not change between the four measures. Every disagreement below is manufactured entirely by the adding up. When two reports about the same work quote different error figures, the honest first question is not which team built the better rule. The question is which of the four measures each team reached for.
A bus runs a ten stop route. At each stop it is a few minutes early or a few minutes late, and by the end of the day there are ten numbers on the conductor's sheet. Now somebody upstairs wants one number for how the timetable held up. One sheet adds up how far off the bus was, ignoring early and late. A second adds up the squares so that a fifteen minute mess counts for far more than five three minute wobbles. A third takes a square root at the end so the answer comes back in minutes. A fourth divides each delay by the gap it was meant to leave. Being three minutes off then matters more on a five minute headway than on an hourly one. Four sheets, four numbers, one bus, one day.
Two reports describe the same fitted rule on the same ten months. One quotes an error of 3.60 and the other quotes 21.80. What has happened?
What does squaring do that taking the size does not?
Take each of the ten misses, multiply it by itself, and average the results. Multiplying a negative number by itself gives a positive one, so direction disappears on its own without anybody having to strip it out. The ten squares are 1, 81, 64, 0, 25, 0, 9, 9, 4 and 25. The ten squares add to 218, and dividing by ten gives the mean squared error of 21.80. The 218 and the 21.80 are the same quantity with a division by ten between them, and the digits agreeing is arithmetic rather than a misplaced decimal point.
Now look at what the squaring did to month 2. Its miss was 9, the largest on the record. As a plain size it is a quarter of the total: 9 out of a summed size of 36. As a square it is 81 out of 218, a share of the totalOne item's contribution divided by the sum of every item's contribution, written as a per cent. The share answers how much of the whole one row is responsible for, and the shares always add to one hundred across all rows. of 37.16 per cent. The same month, the same miss, and its weight in the headline figure jumped by half again simply because somebody chose to square before averaging.
Push it further with a cleaner case. Compare one miss of 9 against three separate misses of 3. By plain size those are identical: nine points of error either way. By squaring they are not remotely identical. The single miss of 9 squares to 81. The three misses of 3 square to 9 each, 27 in total. Squaring rules that one big miss is three times as bad as three small ones adding to the same amount, and nothing in the formula announces that it is making such a ruling.
One miss of 9, or three separate misses of 3. Which arrangement does squaring prefer?
Mean Squared Error vs Mean Absolute Error: where exactly do the two part company?
Take the size of each miss, throw the direction away, and average the ten. The sizes are 1, 9, 8, 0, 5, 0, 3, 3, 2 and 5. The ten sizes add to 36, and the mean absolute error is 3.60. Under this measure month 2 carries 9 of the 36, or 25.00 per cent, against the 37.16 per cent it carried under squaring.
Now the honest part, and it matters more than the headline. On this record the squared measure and the size measure agree about the whole order, from the worst month to the best, and that has to be said plainly rather than dressing the pair up as rivals. Month 2 is first under both. Month 3 is second under both. Months 5 and 10 tie for third under both. A bigger size always squares to a bigger square, so squaring a set of sizes cannot reshuffle them. The two measures do not part company over which month they point at. They part company over how much of the blame that month is made to carry: 37.16 per cent against 25.00 per cent, half as much again.
Month 2 carries 37.16 per cent of the squared error and 25.00 per cent of the absolute error. Do the two measures disagree about which month went worst?
So if the two never reorder the months, is the choice between them empty? The choice is not empty, and there is a sharper way to see the difference than any share of a total. Put to each measure the simplest question a measure can be asked: forget the fitted rule entirely, and pick one single number to use as the answer in all ten months. Which single number would each measure choose?
Feed that question to the squared measure and it lands on the plain average of the ten outcomes, 2.00 per cent, and it lands there uniquely. Guess 2.00 and the mean squared deviation is 89.30. Guess anything else at all and it goes up: 90.30 at a guess of 1.00 or 3.00, 93.30 at nothing, 98.30 at minus 1.00. There is exactly one winner and every step away from it is punished.
Feed the same question to the size measure and something quite different happens. Guess minus 1.00 per cent and the mean absolute deviation is 7.50. Guess nothing at all and it is 7.50. Guess 1.00, or 2.00, or 2.50, and it is still 7.50. Every guess from minus 1.00 per cent up to 2.50 per cent ties for first place under the size measure, and only outside that stretch does the figure start to climb. Those two ends are the fifth and sixth outcomes when the ten are put in order, so what the size measure has actually chosen is the medianThe middle reading once a column is placed in order, smallest to largest. With an even count of rows there is no single middle, so anything between the two central readings sits equally in the middle., and with ten readings there is no single middle to land on.
The two measures punish different things. The squared measure punishes distance, so one outcome far away drags the answer toward itself, and the answer it settles on is the balance point of the whole column. The size measure punishes presence, so an outcome far away counts once, the same as one nearby, and the answer it settles on is whatever has half the column on either side. Month 2's outcome of 18.50 per cent pulls the squared answer up hard; it barely troubles the size answer at all. The two measures never disagree about the order of the misses, and they disagree completely about what number to aim at in the first place.
Under the size measure, every guess from minus 1.00 per cent to 2.50 per cent scores exactly 7.50. What does that stretch correspond to?
Why take a square root at the end?
The mean squared error of 21.80 has a problem that has nothing to do with weighting. Its unit is squared percentage points. Nobody has any feel for a squared percentage point. A squared percentage point is not a distance, it is not comparable with any other figure here, and setting it beside a miss of 9 in the same sentence compares quantities that are not the same kind of thing at all.
The square root of 21.80 is 4.6690. The root is back in percentage points, the same unit as the misses themselves, and now it can be read: the rule is off by about four and two thirds percentage points in a typical month. Along the row of ten sizes, 4.6690 sits sensibly among them, above the 3s and the 2 and below the 8 and the 9. Taking the root reorders nothing, improves nothing and changes no ruling about which miss counts more; it changes only whether a reader can interpret the number in front of them.
The root is easy to oversell here. Taking a square root of a positive number never flips an order. If one rule has a lower mean squared error than another, it has a lower root as well, always. The root is not a better judge. The root is a better label. The judging was all done by the squaring, several steps earlier, and it stays exactly as done.
Taking the square root of 21.80 gives 4.6690. What exactly does that step change?
What does a percentage measure add, and what does it require?
The three measures so far all answer in points. Answering in points is fine when every month is the same size, and it stops being fine the moment they are not. Being off by 3 points in a month where the outcome was 17.00 per cent is a different kind of error from being off by 3 points where the outcome was 1.00 per cent, and none of the first three measures can tell those two apart.
The fourth measure can. Dividing each miss by the size of the outcome it missed and averaging those ten ratios gives the mean absolute percentage error. Month 1 was off by 1 on an outcome of 3.00 per cent, so its ratio is a third. Month 4 was off by nothing at all, so its ratio is nothing. Month 2 was off by 9 on an outcome of 18.50 per cent, so its ratio is a little under a half, and the biggest miss on the whole record turns into one of the smaller ratios. The measure is doing exactly what it was built to do: judging a miss against the size of what it was trying to hit.
Picture a vegetable seller. On a busy Sunday the stall takes Rs 4,000/-, and being wrong about that by Rs 300/- is a small mistake. On a wet Wednesday the stall takes Rs 100/-, and being wrong about that by the same Rs 300/- is not a small mistake at all. A measure in rupees calls those two errors identical. A measure in per cent calls the second one twelve times worse, closer to how the seller feels about it.
But that step relocated the outcome. The outcome now sits in the denominatorThe number underneath in a division, the one being divided by. As it shrinks toward nothing the answer to the division grows without limit. A denominator near nothing is always worth checking before the result can be trusted., and nothing whatever in the formula insists that it stay away from nothing. That is the requirement this measure carries and never states. Every outcome on the record must be comfortably clear of nothing, or the division will do something violent, and the formula will do it silently. Two of the ten outcomes on this record are not comfortably clear of nothing at all.
Settle this one before the control below moves at all. Month 8 was missed by 3 percentage points on an outcome of minus 1.00 per cent. If that outcome shrinks toward nothing while the miss of 3 is held exactly where it is, what happens to month 8's percentage error?
Shrink one outcome toward nothing and watch a single month swallow the whole measure.
One control, and it touches one number. The control moves month 8's outcome from minus 10.00 per cent up to minus 0.10 per cent, and every other figure on the record is pinned: month 8's miss stays at 3 percentage points throughout, the other nine months keep their own outcomes and their own misses, and the rule itself is never refitted. As the control slides, the bar for the whole record's percentage measure grows while the block belonging to the other nine months sits stubbornly still, the pair of bars on the right shows the miss holding steady against a vanishing outcome, and the strip along the foot reorders itself live as month 8 climbs the ranking. The default setting of minus 1.00 per cent is the record's own reading and reproduces 110.81 per cent exactly.
Why does a rule that explains three quarters of the movement score above one hundred per cent?
The mean absolute percentage error on this record is 110.81 per cent. The R squaredA score running from nothing up to one that reports how much of an outcome column's variation a fitted rule managed to track. Its construction is covered separately. of the same rule on the same ten months is 0.7559. Both figures are correct. Neither is a typing error. And a reader meeting them in the same paragraph with no explanation is entitled to think one of them must be wrong.
Here is what is actually happening. Month 3's outcome was 2.50 per cent and month 8's was minus 1.00 per cent, the two smallest readings on the whole record. Month 3 was missed by 8, and 8 divided by 2.50 gives a percentage error of 320.00 per cent for that month alone. Month 8 was missed by 3, and 3 divided by 1.00 gives 300.00 per cent. Month 3 and month 8 carry 55.95 per cent of the entire percentage measure between them. The headline figure is more than half decided by the two months where the least was happening.
Notice what did not go wrong. The rule did not miss badly in those months. A miss of 3 is smaller than the record's average miss of 3.60, and it would not stand out at all in a list of sizes. The miss became enormous only after being divided by a very small number. The measure is not reporting a fault in the rule. The measure is reporting a fault in its own arithmetic, in exactly the same tone of voice it uses for everything else.
A rule reaching an R squared of 0.7559 on these ten months scores a percentage error of 110.81 per cent. Which of the two figures is wrong?
Do the four ever disagree about which month went worst?
One finding makes the choice of measure more than a matter of taste. Rank the ten months three times over. By squared error the worst is month 2, carrying 37.16 per cent of the total. By plain size the worst is month 2 again, carrying 25.00 per cent. By percentage error the worst is month 3, carrying 28.88 per cent, and month 2 has dropped to fifth place with 4.39 per cent of the total.
Two of the three measures agree on the entire order and the third reshuffles it, so the disagreement is not a difference of degree but a different answer about where to go and look. That distinction matters more than any change in share. If a share moves, the headline changes and the investigation goes to the same place. If the rankingThe list of items placed in order from worst to best by some measure. Two measures can hand back lists in different orders even when both are computed correctly from the same rows. moves, two people reading two correctly computed reports about the same rule will walk to two different months and start asking questions there.
Under the percentage measure month 2 falls from first to fifth. Why does a reordering matter more than the change in its share of the total?
What should a report actually carry?
Somebody has to write one line at the top of a report and let a reader who will not check anything draw a conclusion from it. Four rules make that line honest, and none of them costs more than a few extra words.
Quote a measure in the units of the thing being predicted. That means the root rather than the raw squared figure: 4.6690 percentage points, not 21.80 squared percentage points. A reader can hold the first one against a miss of 9 and understand it. The second is a number in a unit that exists only inside the arithmetic.
Quote a second measure that weights the misses differently. Put the 3.60 beside the 4.6690. When the two sit close together, as they do here, the summary is not resting on how heavily one bad month was counted. When they sit far apart, one month is doing most of the work and the reader deserves to know that before drawing anything from either figure.
Never quote a percentage measure without also stating the smallest outcome in the record. On this record the smallest is minus 1.00 per cent, and once that is stated the 110.81 per cent stops looking like a verdict on the rule and starts looking like what it is. A percentage measure without its smallest denominator is a figure a reader cannot audit.
Name the worst single case under each measure quoted. Month 2 under the first two, month 3 under the third. Naming the worst case stops two readers walking away with two different investigations from one correctly computed report.
What single fact has to travel alongside any percentage error quoted?
The failure: the figure that got quoted alone
Somebody puts 110.81 per cent at the top of a report and concludes that the rule is worthless. Or, worse, puts it in a summary two lines under an R squared of 0.7559 and leaves the reader to reconcile them. Neither figure is wrong and neither was computed carelessly. The rule really does account for about three quarters of the movement in the Vasant unit, and it really does miss month 8 by 3 percentage points on an outcome of minus 1.00 per cent, a percentage error of three hundred on that month by itself.
The measure was built for records whose outcomes stay well away from nothing, and this record is not one of them. Two of its ten outcomes sit at 2.50 per cent and minus 1.00 per cent, and those two months between them decide 55.95 per cent of the whole percentage figure. The problem is not in the rule. No amount of improving the rule fixes it. Improve every other month to a perfect hit and month 3 and month 8 would carry a hundred per cent of what remained rather than 55.95 per cent of it.
Here is the habit that prevents it, and it takes one look. Before a percentage error is quoted at all, the smallest outcome in the record has to be found and looked at. If it is anywhere near nothing, the measure is about to be dominated by the cases that matter least, and two choices remain: state the smallest outcome next to the figure so the reader can judge it, or quote a different measure. Printing the figure on its own and letting it stand as a verdict is not honest.
Where the misses come from in the first place, and which parts of them more months would shrink, is covered under the fitting of the rule. So is how a fitted rule is judged on months it has never seen, a different question from how the misses on months it has seen are added up. The measures used when the answer is a yes or a no rather than a size, where counting hits and misses replaces averaging distances, are covered separately, and none of the four here applies to them.
Whether an error of 4.6690 percentage points is large or small in any setting outside itself, and whether a rule fitting these ten invented months would be worth pointing at anything, are separate questions. Questions of that shape carry conditions of their own and belong where those conditions can be set out properly.
Which of these figures could be knocked down, and how?
Every number above is either an invented reading or something a calculator settles in a minute, so the honest citation is an instruction to recompute rather than a link to follow. Each headline figure below carries the shortest route to catching it out.
| The figure | How it was got | The shortest way to catch it out |
|---|---|---|
| The ten misses, month by month | Composed for teaching, then subtracted: each outcome less what the straight rule returned for that month | Add the ten together. They come to nothing at all. A column of misses that does not is not this column |
| 21.80 and 4.6690 | The ten misses squared and averaged, then the square root of that average taken once | Multiply 4.6690 by itself and watch whether 21.80 comes back |
| 3.60 | The size of each miss, direction thrown away, averaged over the ten months | Add the sizes: one, nine, eight, nothing, five, nothing, three, three, two, five. Divide by ten |
| 110.81 per cent | Each miss divided by the size of the outcome it missed, those ten ratios averaged | Do month 8 alone: three divided by one. If that is not three hundred per cent, nothing else here will close either |
| 7.50, and the stretch that ties | The ten outcomes scored against one flat guess, at every guess from minus 4.00 to 6.00 per cent | Try a guess of nothing and a guess of 2.00 by hand. Both give a summed size of 75 across the ten months |
| Any rate, threshold, period or standard | None is used anywhere above, because averaging a column of misses needs nothing issued by anybody | Nothing to confirm, because no claim of that kind is made |
The Nakshatra unit and the Vasant unit are invented.
Educational material. Not advice on any investment, tax, budget or market position.
