Portfolio Management puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 31
- Topics
- 13
- Hard
- 30
014You backtest 20 independent trading strategies, none of which has any real edge, and test each one at the 5% significance level. What is the chance at least one looks significant, and what per-test threshold would hold that overall false alarm rate at 5%?Quantitative researchSystematic investing
Try it first
Chance that at least one of the 20 worthless strategies passes?
Show the worked solution
About 64%, and a per-test threshold of about 0.25%. Each worthless strategy passes by luck 5% of the time, so all 20 fail with chance 0.95 to the 20th, 35.8%, and at least one passes 64.2% of the time. To hold the overall rate at 5%, test each at 5% divided by 20, which is 0.25%; the exact version, 1 minus 0.95 to the power of one twentieth, is 0.256%.
Why does testing more ideas throw up a false winner?
Ask a room of 20 people to each flip a coin five times, and there is a fair chance someone gets five heads. Nobody in the room has a lucky hand; there were simply enough tries. A 5% test lets one worthless idea in twenty through by chance, so a researcher who tests twenty ideas should expect about one false winner, not be impressed by it. The expected number of false positives here is 20 x 0.05, exactly 1, and the chance of at least one is 64.2%.
Testing each worthless strategy at 5%, the chance that at least one looks significant reaches 64.2% at 20 strategies, while testing each at 0.25% holds it near 4.9%. The relationship0.95 the chance a worthless strategy fails a 5% test 20 the number of independent strategies tested \alpha_{each} the per-test threshold that caps the overall false alarm rate near 5% What it says in wordsThe chance of at least one false winner is one minus the chance that every test correctly fails; dividing the level by the number of tests caps it.What does a quant desk actually do about it?
Dividing the threshold by the number of tests is called the Bonferroni correctionA rule that divides the significance level by the number of tests run, so the chance of any false positive across all of them stays near the original level.. The real discipline is counting every test you ran, including the ones you dropped quietly, because the correction is only as honest as that count. A researcher who tried 200 variants and reports the best 20 has a far bigger multiple testing problem than the 20 suggest. Desks also hold out data the research never touched and demand a reason for the edge before the backtest. Say the limitation: the correction assumes independent tests, and for correlated strategies it is too strict, which costs real ideas.
Where candidates lose it
The fast wrong answers are 5%, which ignores that there are 20 tests, and 100%, which adds the chances. Say the complement and the answer arrives in one line.
The second trap is naming the fix without the cost. A tighter threshold throws away some genuine strategies too, and an interviewer on a systematic desk expects you to say that trade-off out loud.
What the interviewer asks next
- If the 20 strategies are highly correlated, is the true chance of a false winner higher or lower than 64%?
- Of 1,000 strategies, how many worthless ones pass at 5%?
- Why is out-of-sample testing a better defence than a stricter threshold?
032A loan scoring model catches 90% of applicants who will go on to default, but it also flags 15% of good borrowers. If 4% of applicants default, what share of flagged applicants actually default?BlackRockWilmington · 2025
Try it first
Instinct first: what share of flagged applicants will default?
Show the worked solution
Only 20%. Take 1,000 applicants. 40 will default and the model flags 36 of them. 960 are good and the model wrongly flags 15% of them, 144 people. The flagged pile holds 180, and 36 of those default: one in five. The model is good at catching defaulters, but defaulters are so rare that false alarms outnumber them four to one.
Why is 90% the wrong answer when the model catches 90% of defaulters?
Picture a smoke alarm that always sounds when there is a fire and also sounds now and then for burnt toast. Because toast burns far more often than houses do, most alarms in a year are toast. A flag's meaning depends on how common the thing it looks for is: when defaulters are 4% of applicants, even a modest false alarm rate on the other 96% produces more wrong flags than right ones. The 90% is the chance a defaulter gets flagged; the question asks the reverse, the chance a flag is a defaulter.
Of 1,000 applicants, 36 defaulters and 144 good borrowers are flagged, so only 36 of the 180 flagged applicants, 20%, actually default. The cleared pile is much cleaner: 4 defaulters in 820. Why count people instead of using the formula?
Bayes' rule gives the same answer, but natural frequencies are faster to say and harder to get wrong under pressure. Turn every percentage into a count of people out of a round number, and the answer is simply the flagged defaulters over everyone flagged. The formula version is 0.04 x 0.9 over (0.04 x 0.9 plus 0.96 x 0.15), which is 0.036 over 0.18, or 20%. Offer it as the check after the counts.
The relationshipD the applicant will default G the applicant is a good borrower F the model flags the applicant P(F|D) the catch rate, 90% P(F|G) the false alarm rate, 15% What it says in wordsThe chance a flag is real is the true flags divided by all flags, true and false.Then say what a lender does with it. A flag at 20% is a reason for a closer look, not a rejection. The cleared pile, by contrast, holds only 4 defaulters in 820, about 0.5%, against 4% before the model, so the model is most useful for waving through the safe majority. Cutting the false alarm rate from 15% to 5% would lift the flagged default share to about 43%.
Where candidates lose it
Most candidates answer 90% or something close, swapping the chance of a flag given default for the chance of default given a flag. It is the same slip as reading a medical test's accuracy as the chance you are ill.
Say the base rate first, then count 1,000 people through the tree out loud. The interviewer mostly wants to hear that you know the base rate drives the answer.
What the interviewer asks next
- What false alarm rate would make half of all flags real defaulters?
- If the lender rejects every flagged applicant, how many good borrowers does it turn away per defaulter avoided?
- How does the answer change for a riskier segment where 15% of applicants default?
Asked at BlackRock, Generalist, Wilmington, 2025 (Wall Street Oasis):
Questions are pretty straightforward and test about statistics models about loan application and loan origination.
082Let X be the market's daily return, symmetric around zero, and let Y equal X squared, the shape of a long straddle's profit. What is the correlation between X and Y? Are they independent?Quantitative asset managementHedge funds
Try it first
What is the correlation?
Show the worked solution
The correlation is zero, but X and Y are completely dependent. Because X is symmetric, a 2% fall and a 2% rise both give Y of 4, so the up-slope and the down-slope cancel and the covariance is zero. Yet Y is fixed exactly by X. Zero correlation rules out a straight-line link only, not a link.
How can a perfect relationship show zero correlation?
An umbrella seller does well when it pours and an ice cream seller when it is scorching; a stall selling both does well on any extreme day and badly on mild ones. Its takings depend completely on the weather, but not in a straight line. Correlation asks only whether a straight line fits, so a U-shaped link, rising on both sides, averages out to zero.
Plotting Y equal to X squared for returns of minus 2% to plus 2% gives a U of 4, 1, 0, 1, 4, and the best-fit straight line through those points is flat at 2, so the correlation is zero although Y is fixed exactly by X. What does the algebra say?
Covariance is the average of X times Y, less the product of the averages. With X symmetric around zero, the average of X is zero and the average of X cubed is zero, so the covariance, and with it the correlation, is exactly zero. Using five equally likely days of minus 2, minus 1, 0, 1 and 2 per cent gives the same answer by hand: the covariance works out to 0.
The relationshipE[X] the average return, zero by symmetry E[X^3] the average cubed return, also zero by symmetry What it says in wordsFor a symmetric variable, the covariance with its own square is zero.Now the portfolio point. A long straddleBuying a call and a put at the same strike, so the position profits from a large move in either direction. profits roughly with the square of the move, so its returns can show near-zero correlation with the market while depending heavily on it. A risk model that reads zero correlation as no exposure will treat the straddle as a diversifier and miss that it is a pure bet on the size of market moves.
Where candidates lose it
The trap is equating zero correlation with independence. Independence implies zero correlation; the reverse fails, and this question is the standard counter-example.
The second trap is saying the correlation is high because Y is a function of X. Correlation measures a straight-line fit, and a symmetric U has no slope.
What the interviewer asks next
- What if X is skewed, with bigger falls than rises? Is the correlation still zero?
- Name a measure that would detect this dependence.
- Why can hedge fund returns look uncorrelated with equities and still lose money in a crash?
