Case 089Signal research and data tasksCore
Parvanta asks you to design the train, validation and test split for eight years of daily data with 20-day forward-return labels. How many days must be purged and embargoed around each boundary, and how many walk-forward folds with one-year test windows remain?
1The situation
Parvanta Quant has eight years of daily data, about 2,016 trading days, for a universe of stocks. Each row holds features known at the close and a label: the stock's return over the next 20 trading days. The researcher wants to fit a model, tune its settings and report an honest out-of-sample result.
The interviewer asks you to partition the data: where the training, validation and test blocks go, what must be removed around each boundary, and how many folds a walk-forward test with one-year test windows allows if the model needs three years of history before its first test.
2Your task
Lay out the split, say how many days to purge and embargo at each boundary and why, and count the folds.
Quick check
Why can't you simply put the boundary between training and test on one day and use every row?
Worked solution
Try it on paper, then open one step at a time.
30-second answerThe answer to give first
Purge the 20 training days before each boundary, because their 20-day labels reach into the next block, and embargo 20 days after a test block before any later data is used to train. With three years of warm-up, eight years leave 5 walk-forward folds, testing years 4 to 8. Each fold trains on everything before it, purges 20 days, validates on half a year, purges 20 more and tests a year, losing about 2% of its days to purges.
Step 1Why do overlapping labels leak?
Suppose a teacher sets a test on chapters 5 to 8 but the last homework sheet, set the week before, covered chapters 4 to 6. A student graded on that homework has already practised part of the test. A 20-day forward label on day t is built from prices on days t+1 to t+20, so a training row in the last 20 days before the test block carries test-period returns inside its label. Train on those rows and the model learns from the very moves it will be scored on. The fix is to purge them: drop every training row whose label window reaches past the boundary. Strictly 19 rows overlap and the 20th ends on the boundary price; purging 20 is the clean rule and costs one day.
Step 2What is the embargo for?
The purge handles labels that look forward into the test block. The embargo handles the other direction: rows just after a test block, if they are ever used for training, have features built from test-period prices and labels correlated with the test labels, so a gap of about one label length, 20 days, is left after each test block. In a pure walk-forward design, the model never trains on data after the test block, so the embargo only matters when later data is reused, as in cross-validation that rotates the test block through the middle of the sample, or when the final model is refitted on everything. If features use long look-back windows, for example a 60-day momentum, the embargo should grow with them.
Step 3How many folds does eight years allow?
With 2,016 days and a three-year warm-up, 1,260 days remain, exactly 5 one-year test windows. Fold 1 trains on years 1 to 3 and tests year 4; fold 5 trains on years 1 to 7 and tests year 8. Inside each fold, the half-year before the test block is the validation set used to tune settings, with a 20-day purge on each side: train, purge 20, validate 106 days, purge 20, test 252. Each fold gives up 40 days, about 2% of the sample, and the test years never touch the data used to tune the model they score.
| Fold | Train, days | Purge | Validation | Purge | Test |
|---|---|---|---|---|---|
| 1 | 0 to 609 | 20 | 630 to 735 | 20 | 756 to 1007 |
| 2 | 0 to 861 | 20 | 882 to 987 | 20 | 1008 to 1259 |
| 3 | 0 to 1113 | 20 | 1134 to 1239 | 20 | 1260 to 1511 |
| 4 | 0 to 1365 | 20 | 1386 to 1491 | 20 | 1512 to 1763 |
| 5 | 0 to 1617 | 20 | 1638 to 1743 | 20 | 1764 to 2015 |
Step 4What should you warn the researcher about?
Overlap shrinks the information as well as leaking it. A year of 20-day labels holds only about 12.6 non-overlapping periods, so each test year is a small sample even though it has 252 rows per stock. Standard errors that treat every day as independent will be far too small; compute them on non-overlapping labels or adjust for the overlap. Also keep the universe as it stood on each date, including companies later delisted, and make sure every feature uses data available at that day's close. The split is only as clean as the timestamps behind it.
Where candidates lose it
The common loss is shuffling the rows into random folds, as a textbook would for images. With 20-day labels, neighbouring rows share most of their returns, so random folds put near-copies of test rows in training and the score is inflated.
The second is purging the training rows but tuning the model on the test years anyway, by trying settings until the walk-forward result looks good. The validation block exists so that the test years are scored once.
What the interviewer asks next
- How would the purge change if the labels were 5-day returns and the features used a 60-day window?
- How would you build a cross-validation scheme that uses more of the eight years for testing?
- What standard error would you put on a test-year correlation of 0.05 with these labels?
- When would you choose a rolling training window instead of an expanding one?
Asked at Optiver, Quant Research Interview, San Francisco, 2026 (Wall Street Oasis): Design a ML data pipeline. Partition the data.
Company names and figures are illustrative.
