Hedge Funds puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 38
- Topics
- 14
- Hard
- 30
013A researcher tests 20 unrelated trading signals, each at a 5% significance level, and none of them truly works. What is the chance that at least one of them looks significant?Quant and systematic funds
Try it first
Your instinct: the chance of at least one false discovery?
Show the worked solution
About 64.2%. A useless signal clears a 5% bar by luck one time in twenty. The chance that all 20 stay insignificant is 0.95 to the twentieth, about 35.8%, so the chance that at least one looks like a discovery is 64.2%. On average the search turns up one false signal, as 20 x 5% suggests, but at least one appears in roughly two searches out of three.
Why does testing more ideas manufacture a winner?
Ask twenty friends to flip a coin five times each. Any one of them flips five heads only once in 32 tries, yet the chance that at least one of the twenty does is 47.0%, and that friend will look gifted. Each test is a lottery ticket for a false discovery, and buying twenty tickets makes a win likely even when nothing works. A signal chosen because it looked best among twenty has not passed a 5% test; it has passed a 64.2% one.
The chance of at least one false positive rises from 5% for one useless signal to 40.1% for ten and 64.2% for twenty, passing even odds at 14 tests, although every individual test is run at 5%. The relationship\alpha the significance level of each test, 5% m the number of independent tests, 20 What it says in wordsThe chance that every test stays quiet shrinks with each test added, so the chance of a false winner grows.How do you correct for it?
Tighten the bar to match the number of tries. The Bonferroni correction tests each signal at 5% divided by 20, which is 0.25%, and that brings the chance of any false discovery back to 4.9%. The cost is power: a real but modest signal now struggles to get through. The other defence is data the search never touched: choose the best signal on one period, then test it once on another.
What does a quant fund take from this?
Research teams run thousands of tests, and the ones that get presented are the survivors. Count every test, including the ones you ran and forgot, because the significance of the survivor depends on how many were tried. That is why systematic funds keep research logs and hold data back, and why a backtest with a t-statistic of 2 means much less after a large search than after one planned test. Say the limitation: the 64% assumes independent tests; correlated signals give a lower figure, but rarely a comfortable one.
Where candidates lose it
The fast wrong answer adds the probabilities: 20 x 5% = 100%, a certainty. Adding only works for events that cannot happen together; here several false positives can appear at once, so go through the complement.
The quieter error is answering 5%, treating the batch as one test. The interviewer wants you to see that the error rate of the search is not the error rate of each test inside it.
What the interviewer asks next
- How many tests at 5% before a false positive is more likely than not?
- What significance level per test keeps the family-wide chance at 5% across 100 tests?
- Why does out-of-sample testing help, and what can still go wrong with it?
022Two independent, unbiased forecasts of next quarter's GDP growth have error standard deviations of 2% and 3%. How should you combine them, and how accurate is the blend?Quant and systematic funds
Try it first
What weight should the 2% forecast get?
Show the worked solution
Weight them 9/13 and 4/13, about 69% and 31%, and the blend's error falls to about 1.66%. For independent unbiased forecasts the best weights are proportional to one over each error variance: 1/4 for the 2% forecast and 1/9 for the 3% one. The blended error variance is 1 over (1/4 + 1/9), which is 36/13, so its standard deviation is 1.66%, better than either forecast alone.
Why does blending two forecasts beat the better one?
Ask two people to guess the weight of a pumpkin at a village fair. One is usually closer, but their mistakes are unrelated, so averaging tends to cancel part of each. Independent errors partly cancel when you average, so even a weaker forecast adds information, provided it gets a smaller weight. Throwing the 3% forecast away leaves you at 2%; blending it in well gets you to 1.66%.
Blending with inverse-variance weights of 9/13 and 4/13 gives an error of 1.66%, lower than the better forecast's 2.00%, while equal weights give 1.80% and weights of 3/5 and 2/5 give 1.70%. The relationship\sigma_A, \sigma_B the two forecasts' error standard deviations, 2% and 3% w_A the weight on the sharper forecast What it says in wordsEach forecast is weighted by its precision, one over its variance, and the blend's precision is the sum of the two.Why inverse variance and not inverse error?
The blend's error variance is w squared times 4 plus (1 minus w) squared times 9. Setting its slope to zero gives w = 9/(4 + 9), so the weights follow one over the variance, which penalises the noisier forecast harder than one over the standard deviation would. Check the alternatives: equal weights give an error of 1.80%, weights of 3/5 and 2/5 give 1.70%, and the inverse-variance weights give the minimum, 1.66%.
What would you check before trusting the blend?
Two assumptions carry the answer. The forecasts must be unbiased and their errors independent; if both forecasters lean on the same survey, their errors are correlated and the gain from blending shrinks. With an error correlation of 0.5, the best blend gives the sharper forecast 6/7 of the weight and improves the error only from 2.00% to 1.96%. Ask where each forecast comes from before you average them.
Where candidates lose it
Candidates either average equally, which overweights the noisier forecast, or keep only the better one, which throws information away. Both miss that the right weights come from the variances.
The subtler slip is weighting by one over the standard deviation, 3/5 and 2/5. It is close but not optimal; state the inverse-variance rule and show that the blended error beats 2%.
What the interviewer asks next
- What if the two forecast errors have a correlation of 0.5?
- How would you estimate each forecaster's error variance in practice?
- One forecast is biased upwards by 0.5%. What do you do?
025Four observations, 3.1, 7.4, 5.2 and 9.0, come from a uniform distribution on 0 to theta. What is the maximum likelihood estimate of theta, is it biased, and how would you correct it?AQR Capital ManagementTown of Greenwich · 2022
Try it first
Which statement is right?
Show the worked solution
The MLE is 9.0, the largest observation; it is biased low, and the unbiased correction is 5/4 x 9.0 = 11.25. The likelihood is 1 over theta to the fourth for any theta of at least 9.0 and zero below it, so it peaks at the sample maximum. But the maximum of n draws averages n/(n + 1) of theta, here 4/5, so scaling by (n + 1)/n removes the bias.
Why is the MLE the largest observation?
A friend draws four raffle tickets numbered from 1 up to some unknown top number, and the highest you see is 90. The top number is at least 90; guessing higher only spreads your belief over tickets nobody drew. The likelihood, 1 over theta to the n, is zero for any theta below the largest observation and falls as theta rises above it, so it is maximised exactly at the sample maximum, 9.0. This is a case where you do not differentiate: the maximum sits on a boundary, not where a slope is zero.
Four uniform draws on 0 to theta cut it into five gaps of equal expected size, so the largest draw averages four fifths of theta; the MLE of 9.0 therefore sits below theta, and scaling by 5/4 gives the unbiased 11.25, against 12.35 from doubling the sample mean. The relationshipL(\theta) the likelihood of the four observations n the number of observations, 4 \max the largest observation, 9.0 What it says in wordsThe likelihood peaks at the largest observation, which on average falls short of theta by a factor n/(n + 1).Why is it biased, and by how much?
The largest draw can never be above theta and is almost always below it. Four points dropped at random on 0 to theta cut it into five gaps of equal expected length, so the largest point sits on average four fifths of the way up, and the MLE underestimates theta by a fifth on average. Multiplying by 5/4 fixes it: 9.0 becomes 11.25. The bias shrinks as n grows, since n/(n + 1) tends to 1, but with four points it is large.
How does it compare with the obvious alternative?
The method of moments doubles the sample mean, since a uniform on 0 to theta averages theta/2: the mean here is 6.175, giving 12.35. Both 11.25 and 12.35 are unbiased, but the corrected maximum has a much smaller variance, theta squared over n(n + 2) against theta squared over 3n, because the largest draw carries the most information about the top of the range. With four points that is theta squared over 24 against theta squared over 12: half the variance.
Where candidates lose it
Candidates differentiate the log-likelihood, get minus n over theta, set it to zero and find no solution. The likelihood only falls on the allowed range, so the maximum sits at the boundary, the largest observation; say that before reaching for calculus.
The second miss is calling the MLE unbiased because maximum likelihood estimates are often well behaved. Here it is biased low by construction, and the interviewer expects the (n + 1)/n correction.
What the interviewer asks next
- What is the variance of the corrected estimator with four observations?
- What is the MLE if the distribution is uniform on theta to 2 theta?
- Derive the expected value of the maximum of n uniform draws.
Asked at AQR Capital Management, Trading, Town of Greenwich, 2022 (Wall Street Oasis):
Derive the mle for some given distribution. Explain linear regression intuitively and derive the ols estimate.
038X and Y are independent random variables with the same variance. What is the correlation between X and X + Y?Squarepoint CapitalMontreal · 2026
Try it first
Pick one:
Show the worked solution
1 over root 2, about 0.71. The covariance of X with X + Y is Var(X) plus Cov(X, Y), which is sigma squared plus zero. The standard deviation of X + Y is root 2 times sigma because the variances add. So the correlation is sigma squared over (sigma x root 2 sigma), which is 1/root 2. X explains half the variance of the sum, and the correlation is the square root of that half.
What is the fastest way to set it up?
A two-member team's score is the sum of both players' scores. If the players are equally good and play independently, knowing one player's score tells you something about the team total, but only half the story. Split the covariance: Cov(X, X + Y) = Cov(X, X) + Cov(X, Y) = sigma squared + 0. The variance of the sum is sigma squared + sigma squared = 2 sigma squared, because independent variances add. Correlation is covariance over the product of standard deviations: sigma squared over (sigma x root 2 sigma) = 1/root 2.
Drawn as arrows, independent X and Y sit at right angles and their sum lies at 45 degrees to X, so the correlation is cos 45, about 0.707; equivalently, X supplies half of the variance of X + Y, and the correlation is the square root of one half. Why is the answer not 0.5?
Because 0.5 is the R squaredThe share of one variable variance explained by another; for a simple regression it is the correlation squared., not the correlation. X explains exactly half of the variance of X + Y, and correlation is the square root of the share of variance explained, so it is root 0.5, about 0.707. The geometric picture makes it stick: treat independent variables as arrows at right angles, and correlation as the cosine of the angle between arrows. X + Y sits at 45 degrees to X, and cos 45 is 0.707.
Give the general version to show you own it. If Y has variance k times X's, the correlation is 1/root(1 + k): the more noise you add, the lower it falls. That is the logic behind a noisy signal: a forecast that is half signal and half independent noise, by variance, correlates about 0.71 with the signal, not 0.5.
Where candidates lose it
The common loss is answering 0.5 because X is half of the sum. That is the share of variance, and correlation is its square root.
The other loss is saying zero because X and Y are independent. The sum contains X, so it cannot be independent of X. Split the covariance in one line and the answer falls out.
What the interviewer asks next
- What is the correlation between X + Y and X - Y?
- Y has four times the variance of X. What is corr(X, X + Y) now?
- What is the correlation between the sum of the first 10 and the sum of the first 20 of a series of independent returns?
Asked at Squarepoint Capital, Desk Quant Analyst Interview, Montreal, 2026 (Wall Street Oasis):
There were also 3-4 basic math/stats questions about mean, covariance, correlation, etc.
048A strategy's true annual Sharpe ratio is 1.0. How many years of monthly returns do you need before its average return shows a t-statistic of 2? What if the true Sharpe is 0.5?Viking Global InvestorsNew York · 2014
Try it first
Years needed for a Sharpe of 0.5:
Show the worked solution
About 4 years for a Sharpe of 1.0 and about 16 years for a Sharpe of 0.5. The t-statistic of a mean return is the mean over its standard error, which works out to the annual Sharpe ratio times the square root of the number of years, whatever the data frequency. Setting Sharpe x root(years) = 2 gives years = (2 / Sharpe) squared: 4 for 1.0, 16 for 0.5 and just 1 for 2.0.
Why does the t-statistic grow with the square root of time?
A coin that lands heads 55% of the time looks fair after 20 tosses; you need hundreds before the bias shows through the noise. The average return grows in proportion to time, but the noise around it grows only with the square root of time, so the signal-to-noise ratio, the t-statistic, grows with root time. With monthly data, the t-statistic is the monthly Sharpe times root(12 x years), and the monthly Sharpe is the annual Sharpe divided by root 12, so the twelves cancel: t = annual Sharpe x root(years).
Because the t-statistic equals the Sharpe ratio times the square root of years, a Sharpe of 2.0 clears t = 2 after 1 year, a Sharpe of 1.0 after 4 years and a Sharpe of 0.5 only after 16 years. Why does monthly data not shorten the wait?
More frequent data gives more observations but each is noisier relative to its mean. Sampling the same years more often does not add information about the mean return; only more years do. This is why a {term('t-statistic', 'An estimate divided by its standard error; a value around 2 is the usual threshold for saying an effect is unlikely to be pure noise.')} on the average return depends on the span of the data, not the number of rows. Frequency helps you estimate volatility, not the mean.
The relationshipSR the true annual Sharpe ratio Y years of data 2 the target t-statistic What it says in wordsThe years needed to prove a strategy grow with the inverse square of its Sharpe ratio.Say the practical point. Most real strategies have Sharpe ratios well below 1, so their track records are too short to separate skill from luck with any confidence. An allocator looking at a three-year record with a Sharpe of 0.8 sees a t-statistic of about 1.4. The limitation of the rule: it assumes returns are independent and stable over the whole sample, and fat tails or regime changes make the real uncertainty larger.
Where candidates lose it
The common loss is thinking monthly data gives twelve times the evidence, which leads to answers like four months. The twelve cancels, because the monthly Sharpe is smaller by root 12.
The second loss is saying a Sharpe of 0.5 needs twice as long as 1.0. The dependence is on the square: half the Sharpe, four times the data.
What the interviewer asks next
- How many years for a Sharpe of 0.3?
- You test 20 strategies and pick the best one with t = 2.2. How much do you trust it?
- Would daily data change the answer for estimating the Sharpe ratio itself rather than the mean?
Asked at Viking Global Investors, Quantitative Research, New York, 2014 (Wall Street Oasis):
how to reject a hypothesis test, what's your structure of your code, what's the sample size
050A researcher regresses 12-month forward returns on a signal using monthly observations, so consecutive observations overlap by 11 months, and reports a t-statistic of 4.0 from ordinary least squares. Roughly what is the honest t-statistic?Quant and systematic funds
Try it first
The honest t-statistic is closest to
Show the worked solution
Roughly 1.2, not 4.0. Consecutive 12-month returns share 11 months, so 240 monthly rows over 20 years hold only about 20 independent observations. OLS standard errors assume independence and come out too small by roughly the square root of the overlap, root 12, about 3.5. Dividing 4.0 by 3.46 gives about 1.15: the result is no longer significant. A Newey-West or Hansen-Hodrick standard error does this properly.
What does the overlap do to the regression?
Asking twelve friends for restaurant advice sounds like twelve opinions, but if eleven of them only repeat what the first one said, you have heard about one. Each 12-month return shares 11 months with its neighbour, so the rows are mostly the same data counted again, and the regression thinks it has twelve times more independent evidence than it does. The slope estimate is not biased by the overlap. What breaks is the standard error, because the residuals are strongly correlated from one row to the next, and that breaks one of the {term('OLS assumptions', 'The conditions under which ordinary least squares standard errors are correct, including residuals that are uncorrelated across observations.')}.
Monthly observations of 12-month returns share 11 of every 12 months, so 240 rows over 20 years hold only about 20 independent observations, and the reported t-statistic of 4.0 shrinks to about 1.2 once divided by root 12. Why divide by root 12 and not by 12?
The standard error scales with one over the square root of the number of independent observations. If the effective sample is twelve times smaller, the standard error is root 12, about 3.46, times larger, and the t-statistic is 3.46 times smaller: 4.0 becomes about 1.15. This is a rough correction. The exact factor depends on how persistent the signal is: for a slow-moving signal, such as a valuation ratio, it is close to root 12; for a fast-moving one it can be smaller.
The relationshiph the overlap horizon, 12 months t_OLS the t-statistic from plain OLS standard errors, 4.0 What it says in wordsWith overlapping returns of horizon h, the plain t-statistic is too large by about the square root of h.Say how you would fix it properly: use Newey-West standard errors with at least 11 lags, or Hansen-Hodrick errors built for exactly this overlap, or run the regression on non-overlapping annual data and accept the smaller sample. Any of those should give a t-statistic well below 4.0, and a researcher who reports only the OLS number has not yet shown the signal works.
Where candidates lose it
The common loss is accepting the 4.0 because the slope looks economically sensible. The overlap does not move the slope; it fakes the precision, and the interviewer wants to see you spot that.
The second loss is overcorrecting, dividing by 12 instead of root 12. Standard errors shrink with the square root of the sample, so the correction is the square root of the overlap.
What the interviewer asks next
- How many Newey-West lags would you use here, and why?
- Would non-overlapping annual regressions give the same slope but a bigger standard error?
- Why do long-horizon return predictability studies often report very high R squared values?
063The sample variance computed with n minus 1 in the denominator is an unbiased estimator of the population variance. Is its square root an unbiased estimator of the standard deviation?Squarepoint CapitalLondon · 2026
Try it first
Is the square root of the unbiased sample variance unbiased for the standard deviation?
Show the worked solution
No. The square root of the unbiased variance underestimates the standard deviation on average. The square root is concave, so by Jensen's inequality the average of the square roots is below the square root of the average. For normal data with two observations the estimate averages about 0.80 sigma; the bias shrinks as the sample grows, to about 6% at five observations and under 1% at thirty.
Why does taking a square root break unbiasedness?
Two square rooms have floor areas of 4 and 16 square metres, so their sides are 2 and 4 metres. Average the areas, 10, and take the root: 3.16 metres. Average the sides instead: 3 metres. Averaging and then taking a square root gives a bigger answer than taking square roots and then averaging, because the square root bends downwards. The sample variance is right on average, so the average of its square roots must fall short of the true standard deviation.
Two equally likely variance estimates of 0.04 and 1.96 average to the true variance of 1.0, but their square roots, 0.2 and 1.4, average only 0.8, below the true standard deviation of 1.0, because the square-root curve bends downwards. How big is the bias?
It depends on the sample size and on the distribution. For normal data the expected sample standard deviation is c4 times sigma, with c4 about 0.80 at n = 2, 0.94 at n = 5, 0.97 at n = 10 and 0.99 at n = 30. At n = 2 you can check it directly: the sample standard deviation is the gap between the two draws divided by the square root of 2, and the average gap between two normal draws is 2 sigma over the square root of pi, which leaves the square root of 2/pi, about 0.798.
The relationships the square root of the unbiased sample variance sigma the true standard deviation c4(n) the correction factor for normal data, below 1 for every n What it says in wordsThe average sample standard deviation is a fixed fraction of the true one, and that fraction is below one.Does it matter in practice?
Sometimes. With a year of daily returns the bias is a rounding error; with a handful of monthly returns for a new fund it is not. A manager with five monthly returns has a volatility estimate that averages about 6% too low under normality, which flatters a Sharpe ratioAverage excess return divided by the standard deviation of returns, a measure of return per unit of risk. before anyone has looked at fat tails. Dividing by c4 removes the bias for normal data, but the fix depends on the distribution, so name the assumption. And unbiased is not the same as most accurate.
Where candidates lose it
The trap is assuming unbiasedness carries through any function of an estimate. It carries through straight-line transformations only; the square root is curved, so the property is lost.
The second loss is saying it is biased without the direction or the size. Say biased low, give the Jensen reason in one sentence, and quote about 0.80 at two observations, shrinking towards 1 as the sample grows.
What the interviewer asks next
- Is the square of an unbiased estimator of the standard deviation unbiased for the variance?
- Why does the sample variance divide by n minus 1 rather than n?
- Which estimator of sigma has the lowest mean squared error for normal data?
Asked at Squarepoint Capital, Quantitative Research, London, 2026 (Wall Street Oasis):
Is the square root of the unbiased estimator for sample variance unbiased for standard deviation?
073A thousand fund managers have no skill at all: each has a 50% chance of beating the market in any year, independently. How many will beat it five years running, and what does that say about track records?Quant and systematic funds
Try it first
How many of the 1,000 unskilled managers beat the market five years in a row?
Show the worked solution
About 31 managers, 1,000 halved five times. Each year roughly half the unbeaten managers beat the market by luck, so 500 survive year one, 250 year two, then 125, 62.5 and 31.25. A perfect five-year record is something luck hands to about 3 managers in every 100, so in a large crowd it cannot on its own separate skill from chance.
Why does a crowd produce streaks even without skill?
Fill a stadium with a thousand people and ask each to toss a coin five times. Someone will throw five heads, and about 31 will. A result that is rare for one person is almost certain somewhere in a large group, so the question is never whether a flawless record exists but how many you would expect by chance. Each manager's chance is 1 in 32; across 1,000 managers that is 31.25 expected perfect records.
Starting from 1,000 unskilled managers, half fall away each year, leaving 500, 250, 125, 62.5 and finally about 31 with a flawless five-year record produced by chance alone. The relationship1,000 the number of managers 1/2 each manager's chance of beating the market in a year 5 the number of years What it says in wordsMultiply the crowd by the chance that one member gets the streak.If some managers really are skilled, how much does a perfect record tell you?
Suppose 5% of the thousand are skilled and beat the market 60% of the time. They produce about 3.9 perfect records, while the 950 unskilled produce about 29.7, so a manager with five perfect years is skilled only about 12% of the time. A 60% manager has only a 8% chance of five perfect years, so most skilled managers do not have flawless records either. The record is weak evidence in both directions.
What should you look at instead?
Longer records, more decisions per year and a reason. Skill shows up more reliably in many independent decisions than in a handful of annual outcomes, and in a process that explains where the edge comes from. Allocators also check how many managers were in the starting pool, because the funds still reporting are the ones that survived; the ones that were closed after bad years have dropped out of the data. That is survivorship biasThe distortion that comes from studying only the survivors of a process, whose results look better than those of the whole starting group., and it makes every surviving record look stronger than it is.
Where candidates lose it
The first loss is saying none, or very few, because five in a row sounds impressive. The interviewer wants the crowd arithmetic: rare for one, expected for many.
The second is stopping at 31 without the conclusion. The number is only half the answer; say what it means for reading a track record, and name survivorship bias.
What the interviewer asks next
- How many of the 1,000 beat the market in at least four of the five years?
- How many years of beating the market would one unskilled manager in 1,000 be expected to reach?
- How would you design a test that separates a 60% manager from a 50% one?
075A stock-selection signal has an information coefficient of 0.05, and you can make 400 independent bets a year with it. What information ratio should you expect, and how many independent bets would you need for an information ratio of 1.5?Quant and systematic funds
Try it first
How many independent bets a year does an IC of 0.05 need for an information ratio of 1.5?
Show the worked solution
An information ratio of about 1.0, and about 900 independent bets a year for 1.5. The fundamental law of active management says the information ratio is roughly the information coefficient times the square root of breadth: 0.05 x the square root of 400 = 0.05 x 20 = 1.0. To reach 1.5 the square root must be 30, so breadth must be 900, more than double, because breadth enters under a square root.
Why do many weak calls add up to a strong result?
Picture a cricket pundit who calls the winner right 52.5% of the time. On one match that is nearly useless; over hundreds of independent matches, the small edge becomes a steady record. With independent bets, the expected gain grows in proportion to the number of bets while the noise grows only with its square root, so the ratio of the two grows with the square root of the number of bets. An information coefficientThe correlation between a signal's forecasts and the returns that follow; for a simple up or down call it equals twice the hit rate minus one. of 0.05 is roughly that pundit's edge: a hit rate of 52.5%.
The relationshipIR the information ratio: active return per unit of active risk IC the information coefficient, the skill of each forecast BR breadth, the number of independent bets a year What it says in wordsExpected information ratio is the skill per bet times the square root of the number of independent bets.With an information coefficient of 0.05 the information ratio rises with the square root of breadth, reaching 1.0 at 400 independent bets and 1.5 only at 900, while doubling the coefficient to 0.10 reaches 1.5 with just 225 bets. What does the square root mean for building a strategy?
Skill and breadth are not equal levers. Doubling the information coefficient doubles the information ratio; doubling breadth raises it only by about 41%, so matching a doubling of skill needs four times the bets. Going from 1.0 to 1.5 on breadth alone means 2.25 times as many independent bets, 900 against 400. That is why quant funds chase breadth across many stocks and short horizons, and why a small gain in forecast quality is worth so much.
What does the law leave out?
Two things that usually cut the answer. Independence is the hard part: 400 bets on stocks in one sector, or rebalanced so often that they repeat the same view, are far fewer than 400 independent bets. And constraints on position size, shorting and turnover stop a portfolio from fully expressing the signal; a transfer coefficientA number between 0 and 1 measuring how fully a constrained portfolio reflects the signal; it multiplies the fundamental law. of 0.6 would take the expected information ratio from 1.0 to 0.6. State the law, then say which of these you would check first.
Where candidates lose it
The common slip is scaling linearly: 1.5 is one and a half times 1.0, so 600 bets. Breadth sits under a square root, so the bets needed rise with the square of the target: 2.25 times, or 900.
The second loss is treating 400 bets as 400 independent bets without comment. The interviewer wants to hear that correlated positions and portfolio constraints shrink the effective breadth, and that the law is an upper guide rather than a forecast.
What the interviewer asks next
- Your 400 bets are 100 stocks rebalanced quarterly with a signal that barely changes. What is the real breadth?
- What information coefficient would give an information ratio of 1.5 with the original 400 bets?
- The signal's IC decays by half after one month. How should that change the rebalancing frequency?
088A stock's true model is: stock return = 0.5 x market return + 1.0 x sector return + noise. Regressing the sector's return on the market gives a slope of 0.4. If you regress the stock on the market alone, what slope do you get?Quant and systematic funds
Try it first
What does the market-only regression report?
Show the worked solution
About 0.9. The market reaches the stock by two paths: directly, with a coefficient of 0.5, and through the sector, which moves 0.4 for each unit of market and passes all of it on with a coefficient of 1.0. A regression on the market alone cannot separate the two and reports the total, 0.5 + 1.0 x 0.4 = 0.9. The extra 0.4 is omitted variable bias.
Why does leaving the sector out change the market slope?
Suppose you measure how much ice cream sales rise on hot days, but hot days also tend to be holidays, and holidays sell ice cream too. Leave holidays out and the heat gets the credit for both. A regression gives a left-out variable's effect to whichever included variable moves with it, in proportion to how strongly the two move together. Here the sector moves with the market, so the market's slope absorbs part of the sector's effect.
The market reaches the stock directly with a coefficient of 0.5 and through the sector with 0.4 x 1.0 = 0.4, so a regression of the stock on the market alone reports 0.9, of which 0.4 is the sector's effect credited to the market. How do you compute the bias?
Write the sector as 0.4 x market plus a part unrelated to the market, then substitute. Stock = 0.5 x market + 1.0 x (0.4 x market + other) + noise = 0.9 x market + (1.0 x other + noise). The bracket is unrelated to the market, so a regression on the market alone recovers 0.9. The bias is the omitted coefficient times the slope of the omitted variable on the included one, 1.0 x 0.4. A simulation of 20,000 days with these coefficients gives a slope of 0.897, matching the algebra.
The relationshipbeta_M the stock's true direct loading on the market, 0.5 beta_S the stock's loading on the sector that was left out, 1.0 delta the slope of the sector's return on the market's, 0.4 What it says in wordsThe short regression's slope is the true slope plus the left-out variable's effect times how much that variable moves with the one you kept.Is 0.9 wrong, or answering a different question?
It depends on what you use it for. If you want to hedge the stock with the market alone, 0.9 is the right hedge ratio, because it captures everything the market drags along with it. If you want the stock's exposure holding the sector fixed, say to build a sector-neutral book, 0.9 overstates it and 0.5 is the number you need. The bias can also run the other way: if the sector moved against the market, or the stock loaded negatively on the sector, the short slope would sit below 0.5. Naming both uses is what the interviewer is listening for.
Where candidates lose it
The fast wrong answer is 0.5: candidates assume a regression recovers the true coefficient whatever else is left out. It does so only when the omitted variable is unrelated to the included one.
The second loss is getting 0.9 and calling it simply wrong. It is the correct total effect of the market and the right number for a market-only hedge; it is wrong only as an estimate of the direct effect.
What the interviewer asks next
- What slope do you get if the sector's slope on the market is minus 0.4?
- You add the sector to the regression. What happens to the standard error of the market coefficient if the two are highly correlated?
- How does this bias show up when you estimate a stock's factor exposures with too few factors?
