Case 022Systematic research and dataHard
In a timed data exercise you get 5,000 stock-days with 40 features and must build a predictor of next-day returns. Your model has an in-sample R-squared of 4% and an out-of-sample R-squared of 0.5%. What happened, how do you present it to the researcher, and what would you try next?
1The situation
On the superday you get the Kestrov data exercise: 5,000 rows, each a stock on a day, 50 stocks over 100 trading days, with 40 anonymised features and the stock's next-day return. You have a few hours to analyse the data, build a predictive model and then present to a researcher for an hour.
You fit a linear regression on all 40 features using a random 80% of the rows and test on the other 20%. In sample the model explains 4% of the variance of next-day returns. On the held-out rows it explains 0.5%.
2Your task
What does the gap between 4% and 0.5% tell you, how do you present the result honestly, and what would you try next?
Quick check
What is the most likely reason the fit collapses out of sample?
Worked solution
Try it on paper, then open one step at a time.
30-second answerThe answer to give first
The model is overfitting: 40 features fitted to noisy daily returns learn patterns that do not repeat, so fit falls from 4% to 0.5% on new rows. Present the out-of-sample number first, explain the gap, and flag that a random split leaks because stocks on the same day move together. Next, split by time, shrink the model with regularisation or fewer features, and check whether the 0.5%, a correlation of about 0.07, survives.
Step 1What does a 4% in-sample and 0.5% out-of-sample fit tell you?
A student who memorises last year's exam paper scores brilliantly on it and poorly on this year's. A model with 40 free coefficients fitted to daily returns, which are mostly noise, does the same: it memorises quirks of the training rows, so in-sample fit overstates what it knows. Even pure noise would give an in-sample R-squared of about 40 over 5,000, 0.8%, and adjusting for the number of features cuts the 4% to about 3.2%. The out-of-sample 0.5% is the honest number.
Step 2Was the test itself honest?
Probably not, and saying so is what earns credit. The 5,000 rows are 50 stocks on 100 days, and stocks on the same day share a market move. A random split puts some of a day's stocks in training and others in testing, so the model is tested on market moves it has already seen, and even 0.5% may be flattering. The honest design splits by time: train on the early days, leave a gap of a few days, tune on the next block, and test once on the final days. It also means the effective sample is closer to 100 independent days than to 5,000 rows.
Step 3How do you present it to the researcher?
Lead with the out-of-sample result, then the gap, then what you learned. An R-squared of 0.5% on next-day returns is a correlation of about 0.07 between forecast and outcome, which would be useful across many stocks and days if it held up on a clean test. Say which features carried the signal and whether their signs make economic sense. Say what you did not have time to do. Researchers who run these exercises are listening for judgement about the result, not for the biggest number, and a candidate who oversells a 4% in-sample fit fails the room.
Step 4What would you try next?
Four things, in order. Redo the split by time with a gap. Shrink the model with regularisationA penalty on the size of model coefficients, such as ridge or lasso, that pulls weak ones towards zero so the model fits less noise. or keep only a few features chosen on the training block. Clean the inputs: rank or winsorise features and returns so a few extreme days do not drive the fit. Then test stability: does the signal hold in each half of the test period, and does it survive a realistic trading cost? A signal that works only in one month or vanishes after costs is not a signal.
Where candidates lose it
The trap is presenting the 4%. In-sample fit on noisy returns is almost meaningless, and leading with it tells the researcher you do not know the difference between fitting and predicting.
The second is calling the model worthless because 0.5% sounds tiny. For daily returns, a stable correlation near 0.07 can matter; the right move is to test whether it is real, not to dismiss it or inflate it.
What the interviewer asks next
- How would you choose the ridge penalty without touching the test set?
- Two features have a correlation of 0.95. What does that do to the regression, and what do you do?
- How would you turn the forecast into positions, and what cost level kills it?
Asked at Squarepoint Capital, Quantitative Research, Paris, 2025 (Wall Street Oasis): You are basically handed a dataset, and you are ask to both analyse it and construct a predictive model from it.
Asked at Squarepoint Capital, Quantitative Research, Paris, 2025 (Wall Street Oasis): The data task is definitely hard, especially is the time allowed.
Company names and figures are illustrative.
