Case 099Signal research and data tasksHard
At Tessorin, a random forest on 50 features scores a training R-squared of 35% and a test R-squared of -0.8% on daily returns, while a three-feature ridge model scores 0.6% and 0.4%. Explain the gap, choose a model, and say how you would set the number of trees and the depth.
1The situation
A researcher at Tessorin Capital, an invented systematic fund, is predicting next-day returns for 500 stocks from 750 days of training data, 375,000 stock-days, with a further year held out as the test set. Daily returns have a standard deviation of about 2%.
Two models are on the table. A random forest of 500 trees grown to full depth on 50 features reports a training R-squared of 35% and a test R-squared of -0.8%. A ridge regression on three hand-picked features reports 0.6% in training and 0.4% in test. The researcher's note says the forest is the stronger model because its fit is fifty times better.
2Your task
Explain where the 35% came from and why the test is negative, pick the model for production, and describe how you would set the forest's number of trees and its depth if you kept it.
Quick check
What does a test R-squared of -0.8% mean?
Worked solution
Try it on paper, then open one step at a time.
30-second answerThe answer to give first
The forest's 35% is memory, not skill: 500 full-depth trees on 375,000 rows can carve about 75,000 leaves each and reproduce the training noise, and the 35.8 point gap to a test R-squared of -0.8% is the measure of that overfit. The ridge's 0.4% out of sample, with a gap of 0.2 points, is a real if small signal: a correlation of about 0.06 with next-day returns. Ship the ridge. If the forest stays, set depth by walk-forward validation, which peaks at depth 4 here, and add trees until the validation score stops moving, around 200.
Step 1Where did the 35% come from?
A student who memorises last year's exam paper scores full marks on it and fails this year's. A tree grown until its leaves are pure keeps splitting until it has isolated the noise of each training day, and with 375,000 stock-days and leaves of 5 observations it can carve about 75,000 leaves per tree; 500 such trees on 50 features have enough freedom to describe the training set almost exactly, so 35% in training says nothing about tomorrow. Daily returns are mostly noise: a 2% daily standard deviation with a true predictable component worth 0.4% of variance means the forecastable part has a standard deviation of only 0.13% a day. A model claiming 35% claims to predict a 1.18% a day component that does not exist. The out-of-sample R-squaredOne minus the ratio of the model squared errors to the squared errors of a constant forecast, measured on data the model never saw. It can be negative. of -0.8% then says the forest's forecasts add error on new data: predicting zero every day would have beaten it.
| r t | the realised next-day return |
| r hat t | the model's forecast made before the day |
| sigma pred | the standard deviation of the forecastable part of returns implied by an R-squared |
Step 2Which model goes to production, and why?
Judge a model by its test score and by the gap. The ridge's test R-squared of 0.4% is a correlation of about 0.06 with next-day returns, which across 500 stocks every day is a tradeable signal, and its gap of 0.2 points says the training estimate was honest; the forest's test score is negative and its gap is 35.8 points, so it goes nowhere. Three features with four parameters cannot memorise 375,000 rows; that is the point of the ridge penalty and of the small feature set. Two checks before shipping: that the three features were chosen on the training years only, since picking them with the test year in view is leakage of the same kind the forest commits, and that the 0.4% holds across sub-periods and sectors rather than coming from one quarter. A test R-squared is one number from one year; report it with its range across walk-forward folds.
| Model | Train R-squared | Test or validation R-squared | Gap | Predictable daily sd, train / test |
|---|---|---|---|---|
| Random forest, 50 features, full depth | 35.0% | -0.8% | 35.8 points | 1.18% / none |
| Ridge, 3 features | 0.6% | 0.4% | 0.2 points | 0.15% / 0.13% |
| Random forest, depth 4, leaves of 2,000 | 3.0% | 0.5% | 2.5 points | 0.35% / 0.14% |
Step 3How would you set the trees and the depth?
The two settings do different jobs and are tuned differently. Depth controls how much each tree can memorise, so it is tuned by walk-forward validation inside the training years: Tessorin's sweep peaks at depth 4, where leaves hold thousands of stock-days and a split has to improve the fit on a large sample to be made; beyond that, training R-squared keeps climbing and validation falls to -0.8% at full depth. The number of trees does not overfit: each tree is a noisy forecast and averaging more of them lowers the variance of the average, so validation improves with count and then flattens, here past about 200 trees. Add trees until the out-of-bag score stops moving, then stop for cost. Set a minimum leaf size as well as a depth, and subsample features per split, so no single feature dominates. Then compare the tuned forest with the ridge on the same folds: a depth-4 forest at 0.5% validation is close to the ridge's 0.4%, and the ridge is simpler to explain and to monitor, which decides it. The limitation: the validation scores are themselves noisy at this level, so prefer the simplest model inside one standard error of the best.
Where candidates lose it
The common loss is comparing the 35% with the 0.6% and calling the forest fifty times better. Training fit on a model with hundreds of thousands of leaves is a measure of memory. The test score is the only comparison, and on it the forest is below zero.
The second is tuning depth and trees on the test year, which turns the test into another training set. Depth comes from validation folds inside the training data; the test year is touched once, at the end.
What the interviewer asks next
- How would you decide whether the ridge's 0.4% is statistically distinguishable from zero across walk-forward folds?
- What would you do with the forest's 50 features if the ridge only uses 3?
- How does gradient boosting change the overfitting picture compared with a random forest?
- If the forest's test R-squared were +0.5% and the ridge's 0.4%, would you switch?
Asked at Tower Research Capital, Trading, Princeton, 2018 (Wall Street Oasis): Primarily technical, which heavily focuses on machine learning/statistics; some algorithm and brain teasers.
Company names and figures are illustrative.
