Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies
022

Case 022Systematic research and dataHard

In a timed data exercise you get 5,000 stock-days with 40 features and must build a predictor of next-day returns. Your model has an in-sample R-squared of 4% and an out-of-sample R-squared of 0.5%. What happened, how do you present it to the researcher, and what would you try next?

SCSquarepoint CapitalParis · 2025SCSquarepoint CapitalParis · 2025

1The situation

On the superday you get the Kestrov data exercise: 5,000 rows, each a stock on a day, 50 stocks over 100 trading days, with 40 anonymised features and the stock's next-day return. You have a few hours to analyse the data, build a predictive model and then present to a researcher for an hour.

You fit a linear regression on all 40 features using a random 80% of the rows and test on the other 20%. In sample the model explains 4% of the variance of next-day returns. On the held-out rows it explains 0.5%.

2Your task

What does the gap between 4% and 0.5% tell you, how do you present the result honestly, and what would you try next?

Quick check

What is the most likely reason the fit collapses out of sample?

Worked solution

Try it on paper, then open one step at a time.

30-second answerThe answer to give first

The model is overfitting: 40 features fitted to noisy daily returns learn patterns that do not repeat, so fit falls from 4% to 0.5% on new rows. Present the out-of-sample number first, explain the gap, and flag that a random split leaks because stocks on the same day move together. Next, split by time, shrink the model with regularisation or fewer features, and check whether the 0.5%, a correlation of about 0.07, survives.

Step 1What does a 4% in-sample and 0.5% out-of-sample fit tell you?

A student who memorises last year's exam paper scores brilliantly on it and poorly on this year's. A model with 40 free coefficients fitted to daily returns, which are mostly noise, does the same: it memorises quirks of the training rows, so in-sample fit overstates what it knows. Even pure noise would give an in-sample R-squared of about 40 over 5,000, 0.8%, and adjusting for the number of features cuts the 4% to about 3.2%. The out-of-sample 0.5% is the honest number.

More features, better fit in sample, worse out of it1%2%3%4%1.20.95 features2.60.820 features4.00.540 features1.51.040 with ridgein sampleout of sampleR-squared in per cent; figures other than the 40 feature model are illustrative reruns.
As the model grows from 5 to 40 features, in-sample R-squared rises from 1.2% to 4.0% while out-of-sample R-squared falls from 0.9% to 0.5%; the 40 feature model with ridge regularisation fits least in sample and best out of sample, 1.0%.
Step 2Was the test itself honest?

Probably not, and saying so is what earns credit. The 5,000 rows are 50 stocks on 100 days, and stocks on the same day share a market move. A random split puts some of a day's stocks in training and others in testing, so the model is tested on market moves it has already seen, and even 0.5% may be flattering. The honest design splits by time: train on the early days, leave a gap of a few days, tune on the next block, and test once on the final days. It also means the effective sample is closer to 100 independent days than to 5,000 rows.

Split by time, with a gap, or the test set leaksRandom splitleaksSplit by timehonestTrain, days 1 to 60ValidateTest once5 day gaptraining daystest days mixed in among themStock-days on the same date share one market move, so a random split tests on days the model has seen.
A random split scatters test days among training days, so stocks on the same date and the same market move appear on both sides; a split by time trains on days 1 to 60, leaves a 5 day gap, validates on days 66 to 80 and tests once on days 81 to 100.
Step 3How do you present it to the researcher?

Lead with the out-of-sample result, then the gap, then what you learned. An R-squared of 0.5% on next-day returns is a correlation of about 0.07 between forecast and outcome, which would be useful across many stocks and days if it held up on a clean test. Say which features carried the signal and whether their signs make economic sense. Say what you did not have time to do. Researchers who run these exercises are listening for judgement about the result, not for the biggest number, and a candidate who oversells a 4% in-sample fit fails the room.

Step 4What would you try next?

Four things, in order. Redo the split by time with a gap. Shrink the model with regularisationA penalty on the size of model coefficients, such as ridge or lasso, that pulls weak ones towards zero so the model fits less noise. or keep only a few features chosen on the training block. Clean the inputs: rank or winsorise features and returns so a few extreme days do not drive the fit. Then test stability: does the signal hold in each half of the test period, and does it survive a realistic trading cost? A signal that works only in one month or vanishes after costs is not a signal.

Where candidates lose it

The trap is presenting the 4%. In-sample fit on noisy returns is almost meaningless, and leading with it tells the researcher you do not know the difference between fitting and predicting.

The second is calling the model worthless because 0.5% sounds tiny. For daily returns, a stable correlation near 0.07 can matter; the right move is to test whether it is real, not to dismiss it or inflate it.

What the interviewer asks next

  • How would you choose the ridge penalty without touching the test set?
  • Two features have a correlation of 0.95. What does that do to the regression, and what do you do?
  • How would you turn the forecast into positions, and what cost level kills it?

Asked at Squarepoint Capital, Quantitative Research, Paris, 2025 (Wall Street Oasis): You are basically handed a dataset, and you are ask to both analyse it and construct a predictive model from it.
Asked at Squarepoint Capital, Quantitative Research, Paris, 2025 (Wall Street Oasis): The data task is definitely hard, especially is the time allowed.

← Case 021Kovil Realty's promoters own 55% and have pledged 70% of that stake. Lenders call for more collateral if the stock falls 30% and start selling at a 40% fall. The free float is 45% and daily volume is 0.3% of shares. Build the short thesis and the risk to it.Case 023 →Voltrail Charging wants to justify a valuation of Rs 3,000 crore. Size its market: 20 lakh electric cars in its region, 30% of charging done at public stations, 2,000 kWh charged per car a year and a margin of Rs 4 per kWh. What EBITDA pool exists, and what share must Voltrail win?

Company names and figures are illustrative.

Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.