Risk Management puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 17
- Topics
- 13
- Hard
- 30
011Two loans each have a 5% one-year probability of default and a 2% chance of defaulting together. What is the probability that at least one defaults, and how does it compare with the 9.75% you would get if they were independent?Quant riskBank credit risk
Try it first
What is the probability that at least one of the two loans defaults?
Show the worked solution
8%, lower than the 9.75% for independent loans. The chance of at least one default is 5% plus 5% minus the 2% where both default, which would otherwise be counted twice. Independent loans default together only 0.25% of the time, so they give 9.75%. Correlation makes any single default slightly less likely but a double default eight times more likely.
Why subtract the joint probability?
Think of a class where 5 students play cricket and 5 play football, and 2 play both. If you ask how many play at least one game, adding 5 and 5 counts the two all-rounders twice, so the answer is 8. The chance of at least one default is the sum of the single chances minus the chance of both, because the joint case sits inside each single case. Only when the joint case is tiny does adding the two chances get close.
Out of 400 equally likely outcomes, correlated loans put 8 cells in both-default and 32 cells in any-default, 8%, while independent loans put only 1 cell in both-default and spread defaults over 39 cells, 9.75%. The relationshipP(A), P(B) each loan's probability of default, 5% P(A and B) the chance both default together, 2% What it says in wordsAdd the chances of each loan defaulting, then take away the overlap you counted twice.Why does correlation cut the chance of any default but raise the chance of both?
Because the defaults are bunched into the same outcomes. When defaults tend to happen together, the bad outcomes overlap, so fewer outcomes contain a default at all, but the ones that do are worse. Independent loans default together only 5% of 5%, 0.25% of the time. Here it is 2%, eight times more often, and the implied default correlationThe correlation between two yes-or-no default outcomes, computed from the joint and single default probabilities. is about 0.37.
That is the lesson a credit portfolio manager takes from the puzzle. The expected number of defaults is 0.1 in both cases, because expectations add regardless of correlation. What correlation changes is the shape of losses: fewer mild years, more years where everything goes wrong at once. Capital is held for those years, which is why correlated books need more of it even when the expected loss is identical.
Where candidates lose it
The first trap is answering 10%, adding the two 5% figures and forgetting the overlap. The second is using the independent formula, 9.75%, when the question has handed you a joint probability that is not 0.25%.
The quieter miss is stopping at 8% and not saying what it means. The interviewer wants to hear that correlation shifts risk from single defaults to joint defaults.
What the interviewer asks next
- What is the probability that exactly one loan defaults?
- What joint default probability would make the two loans perfectly correlated?
- With 100 such loans, how does correlation change the distribution of the number of defaults?
012A Monte Carlo VaR at 99% uses 10,000 simulated paths. How many scenarios sit in the tail that defines the VaR, and how many paths do you need to halve the standard error of the estimate?MSCISan Francisco · 2018
Try it first
How many paths halve the standard error of the VaR estimate?
Show the worked solution
100 scenarios define the tail, and you need 40,000 paths to halve the error. At 99%, 1% of 10,000 paths is 100 scenarios, and the VaR is read at the edge of that group. Monte Carlo error falls as one over the square root of the number of paths, so halving it means four times the paths. For a book with Rs 10 crore of daily volatility, the error is about Rs 37 lakh at 10,000 paths.
Why does a 99% VaR rest on so few scenarios?
Picture an exit poll that interviews 10,000 voters but reports only on the one voter in a hundred who picked a small party. Your estimate for that party rests on about 100 people, not 10,000. A 99% VaR is read at the boundary of the worst 1% of outcomes, so only about 100 of 10,000 paths carry information about where that boundary is. For a normal book with Rs 10 crore of daily volatility the true VaR is Rs 23.26 crore, and the estimate wobbles around it by about Rs 37 lakh from one run of 10,000 paths to the next.
The relationshipp the tail probability, 1% n the number of paths f(q) the height of the P&L density at the VaR point What it says in wordsThe error of a quantile estimate shrinks with the square root of the paths and grows where the tail is thin.For a book with Rs 10 crore of daily volatility, the standard error of a 99% Monte Carlo VaR is about Rs 37.3 lakh at 10,000 paths and Rs 18.7 lakh at 40,000, because the error falls with the square root of the number of paths. Why does halving the error cost four times the paths?
Because the paths sit under a square root. To halve the error, the square root of the path count has to double, which means the path count itself must quadruple. That is why precision in Monte Carlo is expensive: going from about 1.6% error to 0.8% of the VaR costs four times the computing time. It also explains why desks use variance reductionSimulation techniques, such as antithetic paths or importance sampling, that give a more precise estimate from the same number of paths., especially importance sampling, which pushes more paths into the tail where the VaR is decided.
Say the limit plainly. More paths shrink sampling error only; they do nothing about model error. If the simulated distribution has the wrong tail, a million paths give a very precise estimate of the wrong number. And expected shortfall at 97.5% averages the worst 250 paths, so it is usually more stable than the 99% VaR read from a single boundary.
Where candidates lose it
The trap is answering 20,000 paths, assuming error falls in proportion to the path count. It falls with the square root, and the difference is a factor of two in cost that a quant interviewer will not let pass.
The other miss is thinking all 10,000 paths inform the VaR. Say the number 100 out loud; it shows you know why tail estimates are noisy.
What the interviewer asks next
- How many paths would you need for the same precision on a 99.9% VaR?
- How does importance sampling reduce the error without adding paths?
- Why is a Monte Carlo VaR often less stable from day to day than a historical VaR on 500 days of data?
Asked at MSCI, Risk Management, San Francisco, 2018 (Wall Street Oasis):
the next round was paper test, materials were similar to CFA level 1, and it mostly focused on monte carlo, dividend, and risk
013For a normally distributed P&L, compare the 99% VaR with the 97.5% expected shortfall. Why would switching from one to the other barely change capital for a normal book but raise it for a fat-tailed one?UBSZurich · 2021
Try it first
For a normal P&L, how do the 99% VaR and the 97.5% expected shortfall compare?
Show the worked solution
For a normal P&L they are almost identical: 2.33 against 2.34 standard deviations. On Rs 10 crore of daily volatility that is Rs 23.3 crore against Rs 23.4 crore. The 97.5% level was chosen so the switch would be neutral for a normal book. For a fat-tailed book with the same volatility, VaR is Rs 26.2 crore and expected shortfall Rs 29.1 crore, about 11% higher.
What does each number actually measure?
Think of a river's flood level. VaR is the line on the wall that the water passes one year in a hundred. Expected shortfallThe average loss on the days when losses exceed the VaR at a chosen confidence level. is how deep the water gets, on average, in the years it passes a lower line. VaR reads a single point in the tail; expected shortfall averages everything beyond its cut-off, so it responds to how far the tail stretches. For a normal distribution the 97.5% cut-off is 1.96 standard deviations, and the average of the losses beyond it is 2.338 standard deviations.
The relationshipsigma the standard deviation of daily P&L phi(1.96) the height of the standard normal curve at the 97.5% cut-off 0.025 the probability of being beyond that cut-off What it says in wordsFor a normal book, the average loss beyond 1.96 standard deviations lands almost exactly on the 99% VaR.With the same Rs 10 crore volatility, a normal P&L puts the 99% VaR at Rs 23.3 crore and the 97.5% expected shortfall at Rs 23.4 crore, almost the same point, while a fat-tailed P&L puts them at Rs 26.2 crore and Rs 29.1 crore, 11% apart. Why does the fat-tailed book pay more under expected shortfall?
Because its extreme losses are larger even though its everyday volatility is the same. Expected shortfall averages the tail, so a book that sells protection against crashes, whose losses are rare but very large, shows a much higher number under ES than under VaR. Using a Student t with three degrees of freedom and the same Rs 10 crore volatility, the 99% VaR rises only to Rs 26.2 crore, but the 97.5% expected shortfall reaches Rs 29.1 crore. That gap is exactly the risk VaR was criticised for ignoring.
Give both sides, because the question asks for advantages and disadvantages. Expected shortfall sees the tail's depth and adds up sensibly across desks, since it is subadditive. But it is harder to backtest, because you are checking an average of rare events rather than a count of breaches, and it needs more data to estimate. VaR is easy to backtest and explain, and blind beyond its own line.
Where candidates lose it
The trap is assuming that a 97.5% measure must be smaller than a 99% measure. It compares the confidence levels and forgets that expected shortfall averages beyond its line while VaR stops at it.
The second miss is stating that ES is always much larger than VaR. For a normal book it is not; say the 2.33 and 2.34 and explain that the gap only opens when the tail is fat.
What the interviewer asks next
- Why is VaR not subadditive, and can you build an example with two bonds?
- How would you backtest an expected shortfall model?
- A desk sells deep out-of-the-money puts. Which measure shows its risk better, and why?
Asked at UBS, Risk Management, Zurich, 2021 (Wall Street Oasis):
what are the advantages and disadvantages of ES compared to VaR?
014A stock goes from 100 to 150 and back to 100. Compare the average of the simple returns with the average of the log returns.Asset manager riskQuant risk
Try it first
What is the average simple return over the two periods?
Show the worked solution
The simple returns average +8.3%, the log returns average 0%. The simple returns are +50% and -33.3%, and their average suggests a gain although the price is back at 100. The log returns are +40.5% and -40.5%, which add to zero and match what happened. Log returns add over time; simple returns do not.
Why does the simple average show a gain that never happened?
Imagine a shop that raises a price from Rs 100 to Rs 150 and then cuts it back to Rs 100. The rise is 50 on a base of 100, and the cut is 50 on a base of 150. Simple returns are each measured against a different starting price, so averaging them mixes percentages of different bases and overstates growth whenever prices bounce around. The arithmetic mean of +50% and -33.3% is +8.3%, but the investor has exactly the money they started with.
For a price that goes 100, 150, 100, simple returns of +50% and -33.3% average +8.3% although the price made nothing, while log returns of +40.5% and -40.5% add to zero and average zero. The relationshipr_t the log return in period t P_t the price at the end of period t What it says in wordsLog returns add up across periods to the log of the total change, so a round trip sums to zero.Which one should a risk manager use?
It depends on what you are adding up. Log returns add across time, so they are the natural choice for compounding a return over many days and for most statistical models of a single asset. Simple returns add across assets, so a portfolio's return for one period is the weighted average of its holdings' simple returns, which log returns do not give you. The gap between the two averages is the volatility dragThe shortfall of compound growth below the arithmetic average return, which grows with the variance of returns.: roughly half the variance, and here the swings are wild enough to make it 8 points.
Close with the practical warning. A fund that reports the arithmetic average of its yearly returns will look better than the growth its investors actually got, and the more volatile the fund, the larger the flattering gap. The honest single number for growth over time is the geometric average, which here is zero.
Where candidates lose it
The trap is quoting +8.3% as the average return and calling it performance. It is a correct average of the wrong thing: percentages taken on different bases.
The second miss is saying log returns are simply better. They add over time but not across assets, so a risk manager needs both and should say when each applies.
What the interviewer asks next
- What is the geometric average return here, and how does it relate to the log returns?
- A fund returns +20% and -20% in alternate years. What is its compound growth rate?
- Why do most VaR models use log returns for single assets but simple returns to aggregate a portfolio?
015A 60/40 portfolio holds equities with 15% volatility and bonds with 6% volatility, and the correlation between them is minus 0.2. What is the portfolio volatility?Asset manager risk
Try it first
Pick the portfolio volatility.
Show the worked solution
About 8.84%. The equity term is 0.6 times 15, squared, which is 81. The bond term is 0.4 times 6, squared, 5.76. The cross term is 2 times 0.6 times 0.4 times minus 0.2 times 15 times 6, which is minus 8.64. The variance is 78.12 and its square root is 8.84%, well below the 11.4% weighted average.
Why is the answer not the weighted average of 15% and 6%?
Think of two friends walking a dog on separate leads. If they always pull the same way, the dog is dragged as far as their combined pull. If one sometimes pulls left while the other pulls right, the pulls partly cancel. Volatilities only add in a straight line when the correlation is exactly one; any lower correlation lets the swings offset, and a negative one subtracts risk outright. The weighted average of 15% and 6% is 11.4%, the answer for a correlation of one.
The relationshipw_1, w_2 the weights, 0.6 in equities and 0.4 in bonds sigma_1, sigma_2 the volatilities, 15% and 6% rho the correlation, minus 0.2 What it says in wordsAdd each asset's own variance contribution, then add the cross term, which is negative when the correlation is negative.Equities contribute 81 and bonds 5.76 squared percentage points of variance, and the minus 0.2 correlation subtracts 8.64, leaving 78.12, a volatility of 8.84% against 9.31% at zero correlation and 11.4% at a correlation of one. What do you add after the number?
Two things. First, equities carry almost all the risk: 81 of the 86.76 squared points before the cross term, so a 60/40 portfolio is mostly an equity risk portfolio with a bond cushion. A risk contributionThe share of a portfolio total variance that comes from one holding, including its share of the cross terms. breakdown makes that visible and is usually the next question. Second, the minus 0.2 is an estimate from history, and correlations between equities and bonds have changed sign across decades. If it turned positive at plus 0.3, the volatility would rise to about 10%.
Say the limit too. Volatility treats upside and downside swings alike and assumes the correlation holds in a crisis. In a sharp sell-off correlations can move together, so the diversification shown here is the benefit in normal conditions, not a promise for the worst month.
Where candidates lose it
The trap is answering 11.4%, the weighted average of the volatilities, which ignores diversification entirely. The second is getting the sign of the cross term wrong and adding 8.64 instead of subtracting it.
Say the formula before the numbers, and square the weighted volatilities first: 9 squared and 2.4 squared are easier out loud than 0.36 times 225.
What the interviewer asks next
- What bond weight minimises the portfolio volatility?
- What is the portfolio volatility if the correlation is plus 0.3?
- What share of the portfolio's risk comes from equities once the cross term is split between the two?
016A swap has an expected positive exposure of about Rs 5 crore across its five-year life. The counterparty's credit spread is 200 basis points. Roughly what is the credit valuation adjustment?Counterparty riskQuant risk
Try it first
Which rough estimate is closest?
Show the worked solution
Roughly Rs 45 lakh. The credit spread is what the market charges each year for bearing the counterparty's default risk. Rs 5 crore of expected exposure at 2% is Rs 10 lakh a year. Over five years that is Rs 50 lakh, and discounting at an assumed 4% brings it to about Rs 44.5 lakh. That is the value you give up for trading with a counterparty that can default.
Why does a spread times an exposure give you CVA?
Think of a friend who owes you money and a moneylender who would charge that friend 2% more a year than a safe borrower. That 2% is the market's price for the chance your friend does not pay. The credit valuation adjustmentThe reduction in the value of a derivative to reflect the chance that the counterparty defaults while owing you money. is roughly the counterparty's credit spread charged on the amount it is expected to owe you, year by year, discounted to today. The spread already folds together the chance of default and the loss if it happens.
Rs 5 crore of expected exposure charged at a 2% credit spread costs Rs 10 lakh a year, and discounting each year at 4% gives Rs 9.62, 9.25, 8.89, 8.55 and 8.22 lakh, a rough CVA of about Rs 44.5 lakh. The relationships the counterparty's credit spread, 2% a year EPE_t expected positive exposure in year t, Rs 5 crore D_t the discount factor for year t, at an assumed 4% What it says in wordsCharge the spread on each year's expected exposure, discount it to today, and add the years.What does the rough version leave out?
Three things, each worth a sentence. First, the rough formula ignores that a counterparty that defaults in year two cannot default again in year three, so it slightly overstates the charge. Modelled explicitly, with a 60% loss given default and a constant default rate implied by the spread, the CVA comes to about Rs 41 lakh. Second, exposure on a swap is rarely flat; it usually rises and then falls as payments are made. Third, if the counterparty's credit worsens exactly when your exposure grows, the charge is larger, which is wrong-way risk.
The discount rate here is an assumption for the illustration, and the market spread would come from the counterparty's bonds or credit default swaps where they trade. The structure is what the interviewer wants: exposure profile, times default cost, discounted and summed.
Where candidates lose it
The trap is answering Rs 10 lakh, one year of the spread, or treating the spread as a one-off probability of losing the whole exposure. The spread is an annual rate, and the exposure lasts five years.
The second loss is not naming the assumptions. Say that you used a flat exposure, a 4% discount rate and the spread as a stand-in for default probability times loss given default.
What the interviewer asks next
- How does a collateral agreement with a zero threshold change this CVA?
- The exposure rises from Rs 2 crore in year one to Rs 8 crore in year five, with the same average. Is CVA higher or lower?
- What is debit valuation adjustment, and why do some people dislike booking it as profit?
017A Rs 1,000 crore loan pool is tranched into equity from 0 to 5%, mezzanine from 5 to 15% and senior from 15 to 100%. The pool loses 12%. How much does each tranche lose as a share of its size, and what pool loss wipes out the mezzanine?Moody'sNew York · 2024
Try it first
What share of the mezzanine tranche is lost when the pool loses 12%?
Show the worked solution
Equity loses 100%, mezzanine 70% and senior nothing; the mezzanine is wiped out at a 15% pool loss. The Rs 120 crore loss fills the tranches from the bottom. Equity absorbs its full Rs 50 crore. The remaining Rs 70 crore falls on the Rs 100 crore mezzanine. The senior tranche starts losing only once pool losses pass 15%, the point where the mezzanine is gone.
How do losses move through a tranche stack?
Picture a building flooding from the ground up. The ground floor is soaked before a drop reaches the first floor, and the top floors stay dry until the water climbs to them. Losses fill the tranches from the bottom: each tranche loses nothing until the pool loss passes its attachment pointThe level of pool loss at which a tranche starts to lose money., and everything once the loss passes its detachment point. The equity attaches at 0% and detaches at 5%; the mezzanine attaches at 5% and detaches at 15%.
A 12% loss on the Rs 1,000 crore pool wipes out the Rs 50 crore equity tranche, takes Rs 70 crore, or 70%, of the Rs 100 crore mezzanine, and leaves the senior tranche untouched until pool losses pass 15%. The relationshipL the pool loss, 12% A the attachment point, 5% for the mezzanine D the detachment point, 15% for the mezzanine What it says in wordsThe part of the pool loss that falls between a tranche's lower and upper edges, divided by the tranche's thickness.Why does thickness decide how risky a tranche is?
Because a thin tranche goes from untouched to wiped out over a small range of pool losses. The mezzanine is only 10 points thick, so a pool loss moving from 5% to 15% takes it from zero to total loss, while the same move barely registers on the pool as a whole. That is the leverage inside structured finance: the mezzanine's loss share moved 7 times as far as the pool's 12% average suggests from 5% onwards. A rating analyst evaluating the deal asks how likely the pool loss is to cross each attachment point, which depends heavily on how correlated the loans are.
Name the risks the structure does not remove. Correlation among the loans decides whether pool losses cluster at a few percent or occasionally jump past 15%. The collateral data may be weak. And the waterfall rules in the documents, such as when cash is diverted to protect senior holders, can shift losses between tranches in ways this simple loss-only picture does not show.
Where candidates lose it
The trap is answering 12% for every tranche, as if losses were shared in proportion. The whole point of tranching is that they are not.
The second miss is saying the mezzanine loses 7%, the points above its attachment, and forgetting to divide by its 10 point thickness. Loss share is always relative to the tranche's own size.
What the interviewer asks next
- What pool loss would cost the senior tranche 10% of its value?
- How does rising correlation among the loans change the risk of the equity versus the senior tranche?
- Why might a mezzanine tranche be rated well below the pool's average credit quality?
Asked at Moody's, Credit Risk, New York, 2024 (Wall Street Oasis):
What is structured finance, how would you evaluate it, and what are the credit risks?
018A bond callable at 102 trades at 101. If yields fall 100 basis points, does its price rise as much as an otherwise identical non-callable bond, and what happens to its duration?Treasury and ALMBank market risk
Try it first
When yields fall 100 basis points, roughly how does the callable bond move?
Show the worked solution
No. The callable bond rises only about 1.0 point, to about 102.0, against about 7.2 points for the straight bond, and its duration collapses. The issuer will call the bond once refinancing is cheaper, so investors will not pay much above 102. In this stylised example its effective duration is about 3.4 against 6.8 for the straight bond, and it shrinks further as yields fall.
Why can the callable bond not keep rising?
Think of a home loan with no prepayment penalty. When rates fall, the borrower refinances, and the lender who was enjoying a high rate gets the money back. A call option lets the issuer do the same: when yields fall enough, it buys the bond back at 102, so no investor will pay much more than 102 for it. The price is effectively capped at the call price, while a straight bond with the same coupon and maturity keeps gaining as yields fall.
In this stylised 10 year 8% bond, a 100 basis point fall in yield lifts the straight bond from 101.7 to 108.9 but the callable bond only from 101.0 to 102.0, because its price flattens under the 102 call price. What happens to duration and convexity?
Duration is how much the price moves for a yield change, and near the cap the callable bond hardly moves. As yields fall towards the level where the call is exercised, the callable bond's effective duration shrinks, and its price curve bends the wrong way: negative convexityWhen a bond gains less from a fall in yields than it loses from an equal rise, because its price curve bends downward.. Here a 100 basis point rise costs both bonds about the same, roughly 6 points, while a 100 basis point fall gives the callable bond only about 1. That lopsided payoff is the price of the call the investor has sold to the issuer, and it is paid for through a higher yield.
Be clear about the model. The curve here is a stylised cap chosen for illustration, not a full option pricing model; a desk would use an interest rate model to value the call and compute effective duration by bumping the whole yield curve. The shape is what matters for the interview: gains capped, losses intact, duration that shortens exactly when you would want it long.
Where candidates lose it
The trap is applying the straight bond's duration and predicting a 7 point gain. Standard duration assumes the cash flows are fixed, and a callable bond's cash flows change when the call is exercised.
The second miss is saying duration rises because the price is near par. Effective duration falls as the call becomes more likely, because the bond starts to behave like a short bond ending at the call date.
What the interviewer asks next
- Why do mortgage-backed securities show the same pattern?
- How would you hedge a portfolio of callable bonds against falling yields?
- What would a putable bond's price curve look like against the same straight bond?
019Three traders, Asha, Bilal and Chen, are asked who breached a limit. Exactly one of them did, and exactly one of them tells the truth. Asha says Bilal did it, Bilal says he did not, and Chen says he did not. Who breached the limit?Rating agencyRisk GCC
Try it first
Who breached the limit?
Show the worked solution
Chen breached the limit, and Bilal is the one telling the truth. Assume each trader in turn is the culprit and count the true statements. Asha as culprit makes two statements true, Bilal's and Chen's. Bilal as culprit also makes two true, Asha's and Chen's. Only Chen as culprit leaves exactly one true statement, Bilal's denial, which matches the rules.
Where do you start when every statement might be a lie?
Think of a lost house key and three rooms. Rather than arguing about who saw what, you check each room in turn and stop when one fits. With three possible culprits, the fastest route is to assume each one did it, work out which statements are then true, and keep the case that matches the rule of exactly one truth. That turns an argument about who is lying into three short checks.
Assuming Asha is the culprit gives two true statements and assuming Bilal gives two, both breaking the rule, while assuming Chen gives exactly one true statement, Bilal's, so Chen breached the limit. Is there a shortcut that finds it faster?
Yes: look for two statements that cannot both be false. Asha says Bilal did it and Bilal says he did not, so exactly one of those two is true whoever the culprit is. That uses up the single true statement, which means Chen's statement must be false. Chen saying he did not do it is false, so Chen did it. Saying this pairing out loud shows the interviewer you look for structure before grinding through cases.
Then connect it to the job, in one line. A limit breach investigation often starts with conflicting accounts, and the same discipline applies: list the possibilities, check each against the hard evidence, and discard the ones that contradict it. The trade records play the part of the rule here; the accounts are only as good as what they agree with.
Where candidates lose it
The trap is trusting the first statement and following it: Asha accuses Bilal, so candidates test Bilal first and then get tangled. Starting from an accusation lets the story steer you.
The other loss is solving it silently. Say each assumption and its count of true statements; the reasoning is what the interviewer scores.
What the interviewer asks next
- Change the rule to exactly one liar. Who breached now?
- Add a fourth trader, Dev, who says Asha is lying. Does the answer change?
- How would you set up a truth table for this in a spreadsheet?
020An auditor samples 60 of 3,000 trades and finds no booking errors. What error rate can you rule out at 95% confidence?Operational riskBig Four risk advisory
Try it first
What can you conclude from zero errors in 60?
Show the worked solution
You can rule out error rates above about 4.9%, and nothing below. If the true error rate were p, the chance of 60 clean trades would be one minus p, to the power 60. That falls to 5% at p of about 4.87%. The rule of three gives the same answer quickly: 3 divided by 60 is 5%. With a real error rate of 1%, a clean sample happens 55% of the time.
Why does a clean sample not prove a clean book?
Think of tasting one spoonful from a large pot to check the salt. A good spoonful tells you the pot is not badly over-salted, but a few salty patches could easily be missed. Zero errors in a sample puts an upper bound on the error rate; it never proves the rate is zero. At a true error rate of 1%, one trade in a hundred is wrong, yet 60 random trades would all be clean 55% of the time. The sample simply is not large enough to see errors that rare.
The chance that a sample of 60 trades shows no errors is 55% when the true error rate is 1% and falls to 5% only at an error rate of 4.87%, so a clean sample rules out rates above about 4.9% but nothing below. The relationshipp the true error rate in the population of trades 60 the sample size 0.05 the chance you accept of being wrong, for 95% confidence What it says in wordsFind the error rate at which a clean sample would be a one-in-twenty event; anything higher is ruled out.Where does the rule of three come from, and does the 3,000 matter?
The chance of zero errors is roughly e to the power of minus n times p, and e to the minus 3 is about 5%. So n times p equal to 3 marks the 95% bound, which gives the rule of three: divide 3 by the sample size. The population of 3,000 barely matters here because the sample is only 2% of it. Sampling without replacement tightens the bound slightly: counting exactly, the book could hold at most 144 errors, about 4.80%, instead of the 4.87% the simple formula gives.
Then give the practical point. If the firm's tolerance for booking errors is 1%, a sample of 60 cannot confirm it; you need about 300 clean trades, 3 divided by 1%, to rule out 1%. The limit is that the sample must be random; a sample of the easiest trades to check says little about the ones that go wrong.
Where candidates lose it
The trap is reporting that the error rate is zero, or that it is below one in 60. Neither follows. A clean sample of 60 is quite likely even when one trade in a hundred is wrong.
The second miss is freezing on the exact formula. Give the rule of three first, then refine it to 4.87% if asked.
What the interviewer asks next
- How many trades must you sample, all clean, to rule out a 0.5% error rate?
- The sample of 60 finds one error. What upper bound can you now give?
- Why might a random sample still understate errors in complex trades?
