Case 075Systematic research and dataHard
You receive the Project Tamarind dataset of 30 anonymised columns and a target of next-day returns, to be tested on data you cannot see. Column 17 correlates 0.35 with the target in-sample. What do you suspect, how do you check it, and how do you build a model that survives the hidden test?
1The situation
As part of the process, you receive Project Tamarind: a table of about 2,500 trading days, 30 columns with their names removed, and a target column, the next day's return on an equity index future. You must hand in a model; the firm will score it on later data you never see.
A first pass shows that 29 columns have in-sample correlations with the target between -0.04 and 0.04. Column 17 has 0.35.
2Your task
What do you suspect about column 17, how would you check it, and how would you build and validate a model that holds up on the hidden test?
Quick check
What is the most likely explanation for column 17's 0.35 correlation with next-day returns?
Worked solution
Try it on paper, then open one step at a time.
30-second answerThe answer to give first
Column 17 is almost certainly leakage, a feature that contains part of the next day's return, so check it before trusting it. A real daily signal correlates a few hundredths with next-day returns; 0.35 is about 18 standard errors from zero and implies a Sharpe ratio near 6. Lag it one day: if the correlation collapses, as it does here to 0.01, drop it. Then build a simple, regularised model tested in time order with purge gaps, and expect an out-of-sample correlation of a few hundredths.
Step 1Why is 0.35 a warning rather than a find?
Because of what it would imply. A student who scores 100% on every practice test has probably seen the answer key. With 2,500 days, noise alone gives correlations within about plus or minus 0.04; the other 29 columns sit there, and column 17 is about 18 standard errors out. A genuine predictor of daily index returns at 0.35 would explain 12% of next-day moves and support a Sharpe ratio near 5.6, roughly 0.35 times the square root of 252. No such signal survives in a traded market. The likely cause is leakageWhen a feature contains information that would not have been available at the time of the prediction, often because it was computed using the target period itself.: a column computed with data from the target's own day.
Step 2How do you check it?
Test whether it could have been known in time. Shift column 17 back one day and recompute: a real predictor loses a little strength, a leaked one collapses; here it falls from 0.35 to about 0.01. Then look at how it relates to the target itself. If column 17 correlates about 0.97 with the part of the target's return that happens after the market opens, it is probably tomorrow's opening move or a price stamped at the wrong time. Check its timestamps against the target's, and plot its correlation by year; leakage usually looks unnaturally stable.
| Check on column 17 | Result | Reads as |
|---|---|---|
| In-sample correlation with next-day return | 0.35 | too strong for a daily signal |
| Standard errors from zero | 17.5 | not noise |
| Correlation after a one-day lag | 0.01 | collapses: not usable in time |
| Implied Sharpe ratio if genuine | 5.6 | not credible |
| Decision | drop | treat as leakage |
Step 3How do you build a model that survives the hidden test?
Assume the hidden test is later in time, and make every choice as if you were already there. Split the data in time order, train on the past, leave a purge gap of a few days so overlapping information cannot bleed across, and test on the next block; repeat the walk forward five times. Lag every feature so it is known before the target period starts. Prefer a simple regularised model, such as ridge regression or a small ensemble, over a complex one tuned until the folds look good. An honest out-of-sample correlation of about 0.03 is a success: at that level, traded daily, it supports a Sharpe ratio of about 0.5 before costs.
Close with what you hand in. Report column 17 and why you dropped it, the fold-by-fold results, not only their average, and the out-of-sample correlation you expect on the hidden data, a few hundredths. Interviewers who set this task are often testing whether you notice the planted leak; the candidate who submits a model built on it scores brilliantly in-sample and fails on the data that matters.
Where candidates lose it
The frequent loss is building the model around column 17 because it is the strongest feature. It will look superb in-sample and collapse on the hidden test, which is exactly what the test is designed to catch.
The second is using random shuffled cross-validation on time series. Shuffling lets the model train on days after the ones it is tested on, which flatters every feature and hides leakage.
What the interviewer asks next
- Column 17 lagged one day still correlates 0.06. Would you keep it?
- How would you choose the purge gap length?
- The hidden test period is a crisis. What would you change about the model or its expected performance?
Asked at Point72, Hedge Fund, new york, 2026 (Wall Street Oasis): was given a data project. the columns/features were removed, apparently it's something they are currently working on.
Asked at Point72, Hedge Fund, new york, 2026 (Wall Street Oasis): had total of 7 rounds. each person is super smart and capable. plus a data project.
Company names and figures are illustrative.
