Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
Explore NISM prep
Series-VIII · Equity DerivativesSeries-XII · Securities Markets FoundationSeries-V-A · Mutual Fund DistributorsSeries-XV · Research AnalystSeries-XIX-E · Category III AIF ManagersSeries-XIX-D · Category I & II AIF ManagersSeries-XIX-C · Alternative Investment Fund ManagersSeries-XVI · Commodity DerivativesSeries-VI · Depository OperationsSeries-II-A · Registrars & Transfer AgentsSeries-I · Currency DerivativesSeries-VII · Securities Operations & Risk Management
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies

Hedge Funds puzzles, solved step by step

Puzzles
100
Traced to a firm
38
Topics
14
Hard
30
Topic
All topicsBetting and sizing5Conditional probability and Bayes7Continuous probability and distributions7Counting and combinatorics7Estimation and mental maths4Expected value and dice games8Logic and brainteasers10Market making and trading games6Options and payoffs5Portfolio and risk maths8Random walks and Markov chains7Returns, compounding and fees7Statistics and estimation11Valuation, accounting and macro riddles8
Level
AnyWarm upCoreHard
Source
AnyReported at a firmStandard
Showing 1–4 of 4 · filtered from 100Clear filters
  1. 025Four observations, 3.1, 7.4, 5.2 and 9.0, come from a uniform distribution on 0 to theta. What is the maximum likelihood estimate of theta, is it biased, and how would you correct it?Statistics and estimationHardACAQR Capital ManagementTown of Greenwich · 2022

    Try it first

    Which statement is right?

    Show the worked solution

    The MLE is 9.0, the largest observation; it is biased low, and the unbiased correction is 5/4 x 9.0 = 11.25. The likelihood is 1 over theta to the fourth for any theta of at least 9.0 and zero below it, so it peaks at the sample maximum. But the maximum of n draws averages n/(n + 1) of theta, here 4/5, so scaling by (n + 1)/n removes the bias.

    Why is the MLE the largest observation?

    A friend draws four raffle tickets numbered from 1 up to some unknown top number, and the highest you see is 90. The top number is at least 90; guessing higher only spreads your belief over tickets nobody drew. The likelihood, 1 over theta to the n, is zero for any theta below the largest observation and falls as theta rises above it, so it is maximised exactly at the sample maximum, 9.0. This is a case where you do not differentiate: the maximum sits on a boundary, not where a slope is zero.

    The largest draw always sits below theta; scaling by 5/4 corrects it0246810123.17.45.29.0The data, and three estimates of thetaMLE 9.0corrected 11.252 x mean 12.35On average, four draws cut 0 to theta into five equal gapsgapgapgapgapgap04/5 theta = 9.0theta = 11.25So E[max] = 4/5 theta, and theta = 5/4 x 9.0 = 11.25
    Four uniform draws on 0 to theta cut it into five gaps of equal expected size, so the largest draw averages four fifths of theta; the MLE of 9.0 therefore sits below theta, and scaling by 5/4 gives the unbiased 11.25, against 12.35 from doubling the sample mean.
    The relationship
    L(θ)=∏i=141θ=θ−4    (θ≥9.0)E[max⁡]=nn+1θ  ⇒  θ^=n+1nmax⁡=54×9.0=11.25L(\theta) = \prod_{i=1}^{4}\frac{1}{\theta} = \theta^{-4} \;\;(\theta \ge 9.0) \qquad E[\max] = \frac{n}{n+1}\theta \;\Rightarrow\; \hat\theta = \frac{n+1}{n}\max = \frac{5}{4}\times 9.0 = 11.25
    L(\theta)the likelihood of the four observations
    nthe number of observations, 4
    \maxthe largest observation, 9.0
    What it says in wordsThe likelihood peaks at the largest observation, which on average falls short of theta by a factor n/(n + 1).

    Why is it biased, and by how much?

    The largest draw can never be above theta and is almost always below it. Four points dropped at random on 0 to theta cut it into five gaps of equal expected length, so the largest point sits on average four fifths of the way up, and the MLE underestimates theta by a fifth on average. Multiplying by 5/4 fixes it: 9.0 becomes 11.25. The bias shrinks as n grows, since n/(n + 1) tends to 1, but with four points it is large.

    How does it compare with the obvious alternative?

    The method of moments doubles the sample mean, since a uniform on 0 to theta averages theta/2: the mean here is 6.175, giving 12.35. Both 11.25 and 12.35 are unbiased, but the corrected maximum has a much smaller variance, theta squared over n(n + 2) against theta squared over 3n, because the largest draw carries the most information about the top of the range. With four points that is theta squared over 24 against theta squared over 12: half the variance.

    Where candidates lose it

    Candidates differentiate the log-likelihood, get minus n over theta, set it to zero and find no solution. The likelihood only falls on the allowed range, so the maximum sits at the boundary, the largest observation; say that before reaching for calculus.

    The second miss is calling the MLE unbiased because maximum likelihood estimates are often well behaved. Here it is biased low by construction, and the interviewer expects the (n + 1)/n correction.

    What the interviewer asks next

    • What is the variance of the corrected estimator with four observations?
    • What is the MLE if the distribution is uniform on theta to 2 theta?
    • Derive the expected value of the maximum of n uniform draws.

    Asked at AQR Capital Management, Trading, Town of Greenwich, 2022 (Wall Street Oasis): Derive the mle for some given distribution. Explain linear regression intuitively and derive the ols estimate.

  2. 050A researcher regresses 12-month forward returns on a signal using monthly observations, so consecutive observations overlap by 11 months, and reports a t-statistic of 4.0 from ordinary least squares. Roughly what is the honest t-statistic?Statistics and estimationHardQuant and systematic funds

    Try it first

    The honest t-statistic is closest to

    Show the worked solution

    Roughly 1.2, not 4.0. Consecutive 12-month returns share 11 months, so 240 monthly rows over 20 years hold only about 20 independent observations. OLS standard errors assume independence and come out too small by roughly the square root of the overlap, root 12, about 3.5. Dividing 4.0 by 3.46 gives about 1.15: the result is no longer significant. A Newey-West or Hansen-Hodrick standard error does this properly.

    What does the overlap do to the regression?

    Asking twelve friends for restaurant advice sounds like twelve opinions, but if eleven of them only repeat what the first one said, you have heard about one. Each 12-month return shares 11 months with its neighbour, so the rows are mostly the same data counted again, and the regression thinks it has twelve times more independent evidence than it does. The slope estimate is not biased by the overlap. What breaks is the standard error, because the residuals are strongly correlated from one row to the next, and that breaks one of the {term('OLS assumptions', 'The conditions under which ordinary least squares standard errors are correct, including residuals that are uncorrelated across observations.')}.

    Monthly 12-month windows share 11 of every 12 months123456789101112131415161718MonthObs 1Obs 2Obs 3Obs 4Obs 5Obs 6Each window adds one new month (lime) and repeats 11 months already counted (green)20 years of data240 monthly observationsabout 20 independent ones4.0 / root 12 = 4.0 / 3.46t about 1.2below 2: not significanton this rough correctionPlain OLS standard errors treat all 240 rows as independent, so the t-statistic is too big by about root 12
    Monthly observations of 12-month returns share 11 of every 12 months, so 240 rows over 20 years hold only about 20 independent observations, and the reported t-statistic of 4.0 shrinks to about 1.2 once divided by root 12.

    Why divide by root 12 and not by 12?

    The standard error scales with one over the square root of the number of independent observations. If the effective sample is twelve times smaller, the standard error is root 12, about 3.46, times larger, and the t-statistic is 3.46 times smaller: 4.0 becomes about 1.15. This is a rough correction. The exact factor depends on how persistent the signal is: for a slow-moving signal, such as a valuation ratio, it is close to root 12; for a fast-moving one it can be smaller.

    The relationship
    thonest≈tOLSh=4.012≈1.15t_{\text{honest}} \approx \frac{t_{\text{OLS}}}{\sqrt{h}} = \frac{4.0}{\sqrt{12}} \approx 1.15
    hthe overlap horizon, 12 months
    t_OLSthe t-statistic from plain OLS standard errors, 4.0
    What it says in wordsWith overlapping returns of horizon h, the plain t-statistic is too large by about the square root of h.

    Say how you would fix it properly: use Newey-West standard errors with at least 11 lags, or Hansen-Hodrick errors built for exactly this overlap, or run the regression on non-overlapping annual data and accept the smaller sample. Any of those should give a t-statistic well below 4.0, and a researcher who reports only the OLS number has not yet shown the signal works.

    Where candidates lose it

    The common loss is accepting the 4.0 because the slope looks economically sensible. The overlap does not move the slope; it fakes the precision, and the interviewer wants to see you spot that.

    The second loss is overcorrecting, dividing by 12 instead of root 12. Standard errors shrink with the square root of the sample, so the correction is the square root of the overlap.

    What the interviewer asks next

    • How many Newey-West lags would you use here, and why?
    • Would non-overlapping annual regressions give the same slope but a bigger standard error?
    • Why do long-horizon return predictability studies often report very high R squared values?
  3. 075A stock-selection signal has an information coefficient of 0.05, and you can make 400 independent bets a year with it. What information ratio should you expect, and how many independent bets would you need for an information ratio of 1.5?Statistics and estimationHardQuant and systematic funds

    Try it first

    How many independent bets a year does an IC of 0.05 need for an information ratio of 1.5?

    Show the worked solution

    An information ratio of about 1.0, and about 900 independent bets a year for 1.5. The fundamental law of active management says the information ratio is roughly the information coefficient times the square root of breadth: 0.05 x the square root of 400 = 0.05 x 20 = 1.0. To reach 1.5 the square root must be 30, so breadth must be 900, more than double, because breadth enters under a square root.

    Why do many weak calls add up to a strong result?

    Picture a cricket pundit who calls the winner right 52.5% of the time. On one match that is nearly useless; over hundreds of independent matches, the small edge becomes a steady record. With independent bets, the expected gain grows in proportion to the number of bets while the noise grows only with its square root, so the ratio of the two grows with the square root of the number of bets. An information coefficientThe correlation between a signal's forecasts and the returns that follow; for a simple up or down call it equals twice the hit rate minus one. of 0.05 is roughly that pundit's edge: a hit rate of 52.5%.

    The relationship
    IR≈IC×BR=0.05×400=1.0,BR=(1.50.05)2=900\text{IR} \approx \text{IC} \times \sqrt{\text{BR}} = 0.05 \times \sqrt{400} = 1.0, \qquad \text{BR} = \left(\frac{1.5}{0.05}\right)^2 = 900
    IRthe information ratio: active return per unit of active risk
    ICthe information coefficient, the skill of each forecast
    BRbreadth, the number of independent bets a year
    What it says in wordsExpected information ratio is the skill per bet times the square root of the number of independent bets.
    Skill counts once, breadth counts under a square root0.51.01.5002004006008001,000Independent bets a year (breadth)IC 0.10IC 0.05400 bets: IR 1.0900 bets: IR 1.5225IR = IC xroot of breadthDouble IC =4x the bets
    With an information coefficient of 0.05 the information ratio rises with the square root of breadth, reaching 1.0 at 400 independent bets and 1.5 only at 900, while doubling the coefficient to 0.10 reaches 1.5 with just 225 bets.

    What does the square root mean for building a strategy?

    Skill and breadth are not equal levers. Doubling the information coefficient doubles the information ratio; doubling breadth raises it only by about 41%, so matching a doubling of skill needs four times the bets. Going from 1.0 to 1.5 on breadth alone means 2.25 times as many independent bets, 900 against 400. That is why quant funds chase breadth across many stocks and short horizons, and why a small gain in forecast quality is worth so much.

    What does the law leave out?

    Two things that usually cut the answer. Independence is the hard part: 400 bets on stocks in one sector, or rebalanced so often that they repeat the same view, are far fewer than 400 independent bets. And constraints on position size, shorting and turnover stop a portfolio from fully expressing the signal; a transfer coefficientA number between 0 and 1 measuring how fully a constrained portfolio reflects the signal; it multiplies the fundamental law. of 0.6 would take the expected information ratio from 1.0 to 0.6. State the law, then say which of these you would check first.

    Where candidates lose it

    The common slip is scaling linearly: 1.5 is one and a half times 1.0, so 600 bets. Breadth sits under a square root, so the bets needed rise with the square of the target: 2.25 times, or 900.

    The second loss is treating 400 bets as 400 independent bets without comment. The interviewer wants to hear that correlated positions and portfolio constraints shrink the effective breadth, and that the law is an upper guide rather than a forecast.

    What the interviewer asks next

    • Your 400 bets are 100 stocks rebalanced quarterly with a signal that barely changes. What is the real breadth?
    • What information coefficient would give an information ratio of 1.5 with the original 400 bets?
    • The signal's IC decays by half after one month. How should that change the rebalancing frequency?
  4. 098With orthonormal regressors, ordinary least squares gives coefficients of 0.6 and 0.15. What do ridge and lasso with a penalty of 0.2 give for each, and why does only lasso set a coefficient to zero?Statistics and estimationHardCitadelLondon · 2026

    Try it first

    Using half the residual sum of squares plus the penalty, what do the two methods give?

    Show the worked solution

    Ridge gives 0.5 and 0.125; lasso gives 0.4 and exactly 0. With orthonormal regressors each coefficient is shrunk on its own. Ridge divides each by 1 + 0.2 = 1.2, so it scales both down and never reaches zero. Lasso subtracts 0.2 from each size and stops at zero, so the small coefficient, 0.15, is removed. This uses the scaling of half the residual sum of squares plus 0.2 times the penalty.

    Why do orthonormal regressors make this a one-line problem?

    When the regressors are uncorrelated and scaled to unit length, the fit for each coefficient does not depend on the others, so the penalised problem splits into separate one-variable problems. Think of adjusting the volume on two speakers that are not wired together: turning one down does not change the other. For each coefficient you minimise half of (beta minus b) squared plus the penalty, where b is its least squares value, and the answer depends only on that one number. Ridge's penalty is half of beta squared times 0.2; lasso's is the size of beta times 0.2.

    Ridge scales every coefficient down; lasso subtracts and stops at zeroRidge: divide by 1 + 0.20.6 to 0.50.15 to 0.12500.20.40.60.8least squares coefficientpenalisedLasso: subtract 0.2, floor at 00.6 to 0.40.15 is inside the reddead zone below 0.2,so it goes to exactly 000.20.40.60.8least squares coefficientpenalised
    Ridge scales every least squares coefficient by 1/1.2, so 0.6 becomes 0.5 and 0.15 becomes 0.125, while lasso subtracts 0.2 and stops at zero, so 0.6 becomes 0.4 and 0.15, inside the dead zone below 0.2, becomes exactly 0.

    What does each penalty do to a coefficient?

    Ridge's squared penalty pulls hard on big coefficients and gently on small ones. Setting the slope to zero gives beta x (1 + 0.2) = b, so ridge divides by 1.2: 0.6 becomes 0.5 and 0.15 becomes 0.125, shrunk but never zero. Lasso's absolute-value penalty pulls with the same force, 0.2, whatever the size. Its answer is the sign of b times the larger of (the size of b minus 0.2) and zero: 0.6 becomes 0.4, and 0.15, weaker than the pull of 0.2, lands exactly on 0.

    The relationship
    β^ridge=b1+λβ^lasso=sign⁡(b) max⁡(∣b∣−λ, 0)\hat\beta^{\text{ridge}} = \frac{b}{1 + \lambda} \qquad \hat\beta^{\text{lasso}} = \operatorname{sign}(b)\,\max\big(|b| - \lambda,\ 0\big)
    bthe least squares coefficient, 0.6 or 0.15
    lambdathe penalty weight, 0.2
    What it says in wordsRidge divides every coefficient by the same factor; lasso takes the same amount off every coefficient and never goes past zero.

    Why does only lasso select variables?

    At zero, the squared penalty is flat: its slope is zero, so any small coefficient still earns its place by improving the fit a little. The absolute-value penalty has a corner at zero with a slope of 0.2 on each side, so a coefficient stays at zero unless the fit improves by more than 0.2 per unit, and 0.15 does not. That is why lasso gives sparseHaving many coefficients exactly equal to zero, so the model uses only a few of the available variables. models and ridge does not. The limitation to state: with correlated regressors lasso tends to keep one of a group arbitrarily and drop the rest, which is why desks often blend the two penalties in an elastic net.

    Where candidates lose it

    The common slip is to swap the two, saying ridge sets small coefficients to zero because it penalises harder. Ridge's penalty is heavy on big coefficients and almost nothing on small ones, which is exactly why it never zeroes them.

    The second loss is getting the lasso numbers off by a factor of two. Written as the full residual sum of squares plus 0.2 times the absolute values, without the half, the threshold is 0.1, giving 0.5 and 0.05. State your scaling before you give numbers.

    What the interviewer asks next

    • At what penalty does lasso set the 0.6 coefficient to zero as well?
    • The two regressors now have a correlation of 0.9. How do ridge and lasso behave differently?
    • How would you choose the penalty in practice without fitting it to noise?

    Asked at Citadel, Quantitative Research, London, 2026 (Wall Street Oasis): very detailed and difficult questions about regularisation ridge and lasso

Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.