Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies
099

Case 099Signal research and data tasksHard

At Tessorin, a random forest on 50 features scores a training R-squared of 35% and a test R-squared of -0.8% on daily returns, while a three-feature ridge model scores 0.6% and 0.4%. Explain the gap, choose a model, and say how you would set the number of trees and the depth.

Tower Research CapitalPrinceton · 2018

1The situation

A researcher at Tessorin Capital, an invented systematic fund, is predicting next-day returns for 500 stocks from 750 days of training data, 375,000 stock-days, with a further year held out as the test set. Daily returns have a standard deviation of about 2%.

Two models are on the table. A random forest of 500 trees grown to full depth on 50 features reports a training R-squared of 35% and a test R-squared of -0.8%. A ridge regression on three hand-picked features reports 0.6% in training and 0.4% in test. The researcher's note says the forest is the stronger model because its fit is fifty times better.

2Your task

Explain where the 35% came from and why the test is negative, pick the model for production, and describe how you would set the forest's number of trees and its depth if you kept it.

Quick check

What does a test R-squared of -0.8% mean?

Worked solution

Try it on paper, then open one step at a time.

30-second answerThe answer to give first

The forest's 35% is memory, not skill: 500 full-depth trees on 375,000 rows can carve about 75,000 leaves each and reproduce the training noise, and the 35.8 point gap to a test R-squared of -0.8% is the measure of that overfit. The ridge's 0.4% out of sample, with a gap of 0.2 points, is a real if small signal: a correlation of about 0.06 with next-day returns. Ship the ridge. If the forest stays, set depth by walk-forward validation, which peaks at depth 4 here, and add trees until the validation score stops moving, around 200.

Step 1Where did the 35% come from?

A student who memorises last year's exam paper scores full marks on it and fails this year's. A tree grown until its leaves are pure keeps splitting until it has isolated the noise of each training day, and with 375,000 stock-days and leaves of 5 observations it can carve about 75,000 leaves per tree; 500 such trees on 50 features have enough freedom to describe the training set almost exactly, so 35% in training says nothing about tomorrow. Daily returns are mostly noise: a 2% daily standard deviation with a true predictable component worth 0.4% of variance means the forecastable part has a standard deviation of only 0.13% a day. A model claiming 35% claims to predict a 1.18% a day component that does not exist. The out-of-sample R-squaredOne minus the ratio of the model squared errors to the squared errors of a constant forecast, measured on data the model never saw. It can be negative. of -0.8% then says the forest's forecasts add error on new data: predicting zero every day would have beaten it.

The relationship
Roos2=1−∑(rt−r^t)2∑(rt−rˉ)2σpred=σrR2=2%×0.004≈0.13% a dayR^2_{\text{oos}} = 1 - \frac{\sum (r_t - \hat r_t)^2}{\sum (r_t - \bar r)^2} \qquad \sigma_{\text{pred}} = \sigma_r \sqrt{R^2} = 2\% \times \sqrt{0.004} \approx 0.13\% \text{ a day}
r tthe realised next-day return
r hat tthe model's forecast made before the day
sigma predthe standard deviation of the forecastable part of returns implied by an R-squared
What it says in wordsOut-of-sample R-squared measures the model against a constant forecast on unseen data; a usable daily signal explains a fraction of a percent of variance, so 35% in training is a claim about noise.
At full scale only the forest's memory shows; zoom in and the ridge is the one that generalisesFull scale, R-squared 0 to 35%10%20%30%0foresttrain35.0%foresttest-0.8%ridgetrain0.6%ridgetest0.4%Zoomed, R-squared -1% to +1%-1.0%-0.5%+0.5%+1.0%035%off scaleforesttrain-0.8%foresttest+0.6%ridgetrain+0.4%ridgetestGap between train and test: forest 35.8 points, ridge 0.2 points. The gap is the overfit.
At full scale only Tessorin's forest shows, at 35% in training against -0.8% in test; zoomed to one percent either side of zero, the ridge's 0.6% and 0.4% sit almost level, and that small gap is the mark of a model that generalises.
Step 2Which model goes to production, and why?

Judge a model by its test score and by the gap. The ridge's test R-squared of 0.4% is a correlation of about 0.06 with next-day returns, which across 500 stocks every day is a tradeable signal, and its gap of 0.2 points says the training estimate was honest; the forest's test score is negative and its gap is 35.8 points, so it goes nowhere. Three features with four parameters cannot memorise 375,000 rows; that is the point of the ridge penalty and of the small feature set. Two checks before shipping: that the three features were chosen on the training years only, since picking them with the test year in view is leakage of the same kind the forest commits, and that the 0.4% holds across sub-periods and sectors rather than coming from one quarter. A test R-squared is one number from one year; report it with its range across walk-forward folds.

ModelTrain R-squaredTest or validation R-squaredGapPredictable daily sd, train / test
Random forest, 50 features, full depth35.0%-0.8%35.8 points1.18% / none
Ridge, 3 features0.6%0.4%0.2 points0.15% / 0.13%
Random forest, depth 4, leaves of 2,0003.0%0.5%2.5 points0.35% / 0.14%
Tessorin's three candidates. The full-depth forest claims a 1.18% daily predictable component in training and delivers none; the ridge and a shallow forest each deliver a real component of about 0.13% a day with small gaps.
Step 3How would you set the trees and the depth?

The two settings do different jobs and are tuned differently. Depth controls how much each tree can memorise, so it is tuned by walk-forward validation inside the training years: Tessorin's sweep peaks at depth 4, where leaves hold thousands of stock-days and a split has to improve the fit on a large sample to be made; beyond that, training R-squared keeps climbing and validation falls to -0.8% at full depth. The number of trees does not overfit: each tree is a noisy forecast and averaging more of them lowers the variance of the average, so validation improves with count and then flattens, here past about 200 trees. Add trees until the out-of-bag score stops moving, then stop for cost. Set a minimum leaf size as well as a depth, and subsample features per split, so no single feature dominates. Then compare the tuned forest with the ridge on the same folds: a depth-4 forest at 0.5% validation is close to the ridge's 0.4%, and the ridge is simpler to explain and to monitor, which decides it. The limitation: the validation scores are themselves noisy at this level, so prefer the simplest model inside one standard error of the best.

Depth buys training fit all the way up and validation fit only to depth 4-1.0%-0.5%+0.5%02train 1.2%+0.3%3train 2.0%+0.4%4train 3.0%+0.5%6train 6.0%+0.4%8train 10.0%+0.2%12train 18.0%-0.2%nonetrain 35.0%-0.8%Training R-squared:best on validation: depth 4Maximum tree depth (none = grown until leaves are pure)worse thanpredicting zeroTrees: validation R-squared stops improving past about 200 trees; depth, not count, is what overfits.
Tessorin's walk-forward sweep: validation R-squared peaks at 0.5% at depth 4 and falls to -0.8% with no depth limit, while training R-squared climbs from 1.2% to 35%; depth is what overfits, and tree count only reduces variance.

Where candidates lose it

The common loss is comparing the 35% with the 0.6% and calling the forest fifty times better. Training fit on a model with hundreds of thousands of leaves is a measure of memory. The test score is the only comparison, and on it the forest is below zero.

The second is tuning depth and trees on the test year, which turns the test into another training set. Depth comes from validation folds inside the training data; the test year is touched once, at the end.

What the interviewer asks next

  • How would you decide whether the ridge's 0.4% is statistically distinguishable from zero across walk-forward folds?
  • What would you do with the forest's 50 features if the ridge only uses 3?
  • How does gradient boosting change the overfitting picture compared with a random forest?
  • If the forest's test R-squared were +0.5% and the ridge's 0.4%, would you switch?

Asked at Tower Research Capital, Trading, Princeton, 2018 (Wall Street Oasis): Primarily technical, which heavily focuses on machine learning/statistics; some algorithm and brain teasers.

← Case 098Corvinta runs two markets on the same pair of dice: A settles at their sum and B at the first die minus the second. Another player quotes A at 6.5 bid, 7.5 offer and B at -0.5 bid, 0.5 offer. The first die is revealed as 5. Update both fair values and find the trades.Case 100 →Vindhavan's two assets have equilibrium expected returns of 6% and 7%. A manager believes A will beat B by 3% and holds that view with 50% confidence. In a simple two-asset Black-Litterman setting, how do the blended expected returns move, and what happens to the weights?

Company names and figures are illustrative.

Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.