Case 008Model risk and validationCore
A quant team presents a signal with a backtested Sharpe ratio of 2.5, the best of 200 variants it tried. How much of that is the search, and what figure would you approve it on?
1The situation
Tessaro Capital's quant team wants capital for a new equity signal. Its backtest over five years of daily data shows an annualised Sharpe ratio of 2.5. In the appendix you find that the team tested 200 variants of the signal, changing look-back windows, thresholds and rebalancing rules, and is presenting the best one.
For the estimate, assume that a strategy with no skill has a backtested annualised Sharpe ratio that is roughly normal around zero with a standard error of one over the square root of the number of years. Treat the variants as independent to begin with, then consider what correlation between them changes.
2Your task
What is the best Sharpe ratio you would expect from 200 variants with no skill, and what should the presented 2.5 be discounted to before approval?
Quick check
If none of the 200 variants had any skill, roughly what Sharpe ratio would the best of them show over five years?
Worked solution
Try it on paper, then open one step at a time.
30-second answerThe answer to give first
The best of 200 no-skill variants would show a Sharpe ratio of about 1.23, so roughly half of the presented 2.5 can be the search. Treating the variants as independent, the skill left is about 1.27; because variants of one signal are correlated, the true figure is likely between 1.3 and 1.7. Approve for small live trading on about 1.3, not 2.5.
Step 1Why does trying many variants inflate the best result?
Ask 200 people to toss a coin ten times and the best of them will get eight or nine heads. Nobody concludes that person is skilled at tossing coins. When you choose the best of many noisy results, you choose the one that got luckiest, so its number is biased upward by the search itself. Backtests are noisy estimates: over five years, a Sharpe ratio has a standard error of about 1 over root 5, which is 0.447.
Step 2What would the best of 200 look like with no skill?
The expected maximum of 200 independent standard normal draws is about 2.75. Multiply by the standard error of 0.447 and the best no-skill backtest shows a Sharpe of about 1.23. That is a very presentable number for a signal that does nothing, which is why the size of the search belongs on the first slide, not in the appendix.
| Z | a standard normal draw, one per variant |
| T | years of backtest, here 5 |
| 1/\sqrt{T} | the standard error of an annualised Sharpe ratio for a strategy with no skill |
Step 3What figure would you approve it on?
If every variant shared the same true Sharpe, the best would beat it by about 1.23 on luck, so the presented 2.5 implies a true figure near 1.27. The 200 variants are tweaks of one signal and move together, so the effective number of independent tries is smaller; at 20 effective tries the luck premium is 0.84 and the implied true Sharpe about 1.66. Either way, the honest range is roughly 1.3 to 1.7, not 2.5.
| Effective independent tries | Expected best with no skill | Implied true Sharpe |
|---|---|---|
| 200 | 1.23 | 1.27 |
| 20 | 0.84 | 1.66 |
| 1 (no search) | 0.00 | 2.50 |
Then the decision. The signal is probably real, since 2.5 is far beyond what luck alone produces, but its size is overstated by about half. Approve a small live or paper allocation sized on a Sharpe of about 1.3, with a pre-agreed review after a year, and ask the team for the full list of variants tried. The limitation: this correction handles selection only; it does not fix look-ahead bias, costs left out of the backtest or a regime that has ended.
Where candidates lose it
Candidates either accept 2.5 at face value or throw the signal out as data mining. The interviewer wants a number in between, reached by asking how many things were tried.
The second miss is using the average no-skill Sharpe, zero, as the benchmark. The right benchmark is the best of the search, because that is what was presented.
What the interviewer asks next
- The team says only 10 of the 200 variants were genuinely different. What changes?
- How would a hold-out period have helped, and how would you stop the team peeking at it?
- Why do Sharpe ratios in live trading usually come in below the backtest even without selection bias?
Company names and figures are illustrative.
