Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies
089

Case 089Signal research and data tasksCore

Parvanta asks you to design the train, validation and test split for eight years of daily data with 20-day forward-return labels. How many days must be purged and embargoed around each boundary, and how many walk-forward folds with one-year test windows remain?

OptiverSan Francisco · 2026

1The situation

Parvanta Quant has eight years of daily data, about 2,016 trading days, for a universe of stocks. Each row holds features known at the close and a label: the stock's return over the next 20 trading days. The researcher wants to fit a model, tune its settings and report an honest out-of-sample result.

The interviewer asks you to partition the data: where the training, validation and test blocks go, what must be removed around each boundary, and how many folds a walk-forward test with one-year test windows allows if the model needs three years of history before its first test.

2Your task

Lay out the split, say how many days to purge and embargo at each boundary and why, and count the folds.

Quick check

Why can't you simply put the boundary between training and test on one day and use every row?

Worked solution

Try it on paper, then open one step at a time.

30-second answerThe answer to give first

Purge the 20 training days before each boundary, because their 20-day labels reach into the next block, and embargo 20 days after a test block before any later data is used to train. With three years of warm-up, eight years leave 5 walk-forward folds, testing years 4 to 8. Each fold trains on everything before it, purges 20 days, validates on half a year, purges 20 more and tests a year, losing about 2% of its days to purges.

Step 1Why do overlapping labels leak?

Suppose a teacher sets a test on chapters 5 to 8 but the last homework sheet, set the week before, covered chapters 4 to 6. A student graded on that homework has already practised part of the test. A 20-day forward label on day t is built from prices on days t+1 to t+20, so a training row in the last 20 days before the test block carries test-period returns inside its label. Train on those rows and the model learns from the very moves it will be scored on. The fix is to purge them: drop every training row whose label window reaches past the boundary. Strictly 19 rows overlap and the 20th ends on the boundary price; purging 20 is the clean rule and costs one day.

A 20-day label reaches across the boundary; purge the days whose label doestest blockpurge zoneday -28day -24day -20day -16day -12day -8day -4day -1safe: label ends before day 0leaks: label overlaps test returns-30-20-100+10+20+30Trading days relative to the start of the test block; bar = the 20 days a label covers
At a boundary, every training day within 20 days of the test block has a label that reaches into the test period, so those days are purged, while days further back have labels that end before the boundary and are safe.
Step 2What is the embargo for?

The purge handles labels that look forward into the test block. The embargo handles the other direction: rows just after a test block, if they are ever used for training, have features built from test-period prices and labels correlated with the test labels, so a gap of about one label length, 20 days, is left after each test block. In a pure walk-forward design, the model never trains on data after the test block, so the embargo only matters when later data is reused, as in cross-validation that rotates the test block through the middle of the sample, or when the final model is refitted on everything. If features use long look-back windows, for example a 60-day momentum, the embargo should grow with them.

Step 3How many folds does eight years allow?

With 2,016 days and a three-year warm-up, 1,260 days remain, exactly 5 one-year test windows. Fold 1 trains on years 1 to 3 and tests year 4; fold 5 trains on years 1 to 7 and tests year 8. Inside each fold, the half-year before the test block is the validation set used to tune settings, with a 20-day purge on each side: train, purge 20, validate 106 days, purge 20, test 252. Each fold gives up 40 days, about 2% of the sample, and the test years never touch the data used to tune the model they score.

Five walk-forward folds, with a 20-day purge before every boundaryFold 1Fold 2Fold 3Fold 4Fold 5yr 0yr 1yr 2yr 3yr 4yr 5yr 6yr 7yr 8trainpurge 20 daysvalidationtest, one yearEach fold loses 40 of its days to purges, about 2% of the sample; later test years become training for later folds.
Eight years of Parvanta's data give 5 walk-forward folds after a three-year warm-up, each with an expanding training block, a 20-day purge, a validation block, a second 20-day purge and a one-year test.
FoldTrain, daysPurgeValidationPurgeTest
10 to 60920630 to 73520756 to 1007
20 to 86120882 to 987201008 to 1259
30 to 1113201134 to 1239201260 to 1511
40 to 1365201386 to 1491201512 to 1763
50 to 1617201638 to 1743201764 to 2015
Day ranges for each of Parvanta's 5 folds, numbering days from 0. Every boundary into validation or test is preceded by a 20-day purge.
Step 4What should you warn the researcher about?

Overlap shrinks the information as well as leaking it. A year of 20-day labels holds only about 12.6 non-overlapping periods, so each test year is a small sample even though it has 252 rows per stock. Standard errors that treat every day as independent will be far too small; compute them on non-overlapping labels or adjust for the overlap. Also keep the universe as it stood on each date, including companies later delisted, and make sure every feature uses data available at that day's close. The split is only as clean as the timestamps behind it.

Where candidates lose it

The common loss is shuffling the rows into random folds, as a textbook would for images. With 20-day labels, neighbouring rows share most of their returns, so random folds put near-copies of test rows in training and the score is inflated.

The second is purging the training rows but tuning the model on the test years anyway, by trying settings until the walk-forward result looks good. The validation block exists so that the test years are scored once.

What the interviewer asks next

  • How would the purge change if the labels were 5-day returns and the features used a 60-day window?
  • How would you build a cross-validation scheme that uses more of the eight years for testing?
  • What standard error would you put on a test-year correlation of 0.05 with these labels?
  • When would you choose a rolling training window instead of an expanding one?

Asked at Optiver, Quant Research Interview, San Francisco, 2026 (Wall Street Oasis): Design a ML data pipeline. Partition the data.

← Case 088The interviewer asks for a market on the maximum of four fair dice, then lifts your offer twice in a row at 5.4. What is the contract worth, how wide should the market be given its skewed distribution, and how do you requote?Case 090 →Sthiram's two assets have expected returns of 8% and 8.5%, volatilities of 15% and 16%, and correlation 0.9. Show how a half-point change in one expected return swings the mean-variance weights, and propose a fix.

Company names and figures are illustrative.

Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.