Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
076You must predict a quantity with a single constant, and your data are 1, 2, 3, 4 and 40. Which constant minimises the mean squared error, which minimises the mean absolute error, and what does the difference tell you?Tower Research CapitalPrinceton · 2018
Try it first
Before you calculate: which pair of constants is right?
Show the worked solution
The mean, 10, minimises squared error; the median, 3, minimises absolute error. Squared error grows with the square of a miss, so the single 40 pulls the best constant towards it. Absolute error charges every unit of miss equally, so the best constant sits in the middle of the pack. Choosing the loss is choosing which of the two statistics you estimate.
Why does each loss land on a different constant?
Picture five friends choosing a meeting point on a straight road: four live at the 1, 2, 3 and 4 km marks and one at the 40 km mark. If the goal is the smallest total travel, you meet at 3 km: moving towards the far friend saves them one km per km moved but costs the other four one km each. If the goal is the smallest total of squared travel, the far friend's 37 km squared, 1,369, dominates everything and the meeting point slides to 10 km. Absolute error balances the count of points on each side, which is the median; squared error balances the total distance on each side, which is the mean.
Squared error is lowest at the mean, 226 at c = 10, and the median scores 275 on it; absolute error is lowest at the median, 8.2 at c = 3, and the mean scores 12 on it, so each constant is poor under the other loss. The relationshipx_i the five data points c the constant you predict #{x_i < c} how many points sit below c What it says in wordsSquared error is flat where the total distance above equals the total below, which is the mean; absolute error is flat where the count above equals the count below, which is the median.So which constant should you actually use?
It depends on what the 40 is, and that is the answer the interviewer wants to hear. If the 40 is a typing error or a one-off glitch in a price feed, the median is the honest summary: remove the 40 and the mean falls to 2.5, while the median barely moves. If the 40 is real, say the one big winning day in a strategy's P&L, the mean is the number that matters, because your total profit is the sum of the days and the sum is five times the mean. A robust estimator that ignores the big day would tell you the strategy earns 3 a day when it actually earns 10.
Where does this show up in a quant job?
Every regression makes this choice silently. Ordinary least squares minimises squared error and so fits conditional means; one extreme observation can swing the line. Least absolute deviation, or quantile regression at the 50% level, fits conditional medians and shrugs off the same extreme point. In practice, desks winsorise or clip returns before fitting squared-error models, or switch to a Huber loss that is squared near zero and linear in the tails. The limitation is that none of these fixes is free: every one of them throws away some of the information in genuine large moves.
Where candidates lose it
The common slip is to answer 10 for both, because the mean feels like the default best guess. The two losses answer different questions, and the interviewer is checking that you know squared error chases the outlier and absolute error does not.
The second loss is stopping at the arithmetic. Say what the 40 might be, a data error or a real big day, and which loss fits each case. That judgement is the point of the question.
What the interviewer asks next
- Which constant minimises the maximum absolute error, and what is that loss called?
- Add a sixth point at 5. What happens to the median, and to the set of absolute-error minimisers?
- Why does Lasso use an absolute-value penalty, and how is that related to this puzzle?
Asked at Tower Research Capital, Trading, Princeton, 2018 (Wall Street Oasis):
What if instead of minimizing mean squared error we look at mean absolute error?
