Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
055When would you use gradient boosting on market data, and when would you stick with a linear model?Quant researchQuant trading
Say this
Boosting earns its keep when you have a lot of data, genuine non-linearity and interactions, and a target with enough signal to support the extra capacity. For low-frequency return prediction with a few hundred monthly observations I would use a regularised linear model almost every time.
Then walk it
- The case for trees: they capture interactions and thresholds automatically, handle mixed feature types, are insensitive to monotone transforms, and do not care about outliers in the features. On microstructure problems with millions of observations they genuinely win.
- The case against on returns: the signal-to-noise is so low that a flexible learner mostly memorises noise, and the model cannot extrapolate beyond the range it saw, which is exactly where the interesting market states live. A boosted tree trained through 2019 has no representation of March 2020.
- Data volume is the deciding variable. Daily cross-sectional data with 3,000 names times 20 years is 15 million rows and trees are viable. Monthly aggregate time series with 300 observations is not, no matter how you tune it.
- If I use boosting, I use it with heavy constraints: shallow trees of depth three to five, low learning rate, strong subsampling, early stopping on a purged time-series split, and monotonic constraints where I have a prior on the sign.
- And I would always run the regularised linear baseline first and report both. In practice the boosted model often adds a modest amount of out-of-sample R squared over a good linear model on financial data, which is a real gain but far from the step change people expect. Knowing that the gain is modest rather than transformative is the useful piece of experience here.
Where candidates lose it
Defaulting to whatever is fashionable with no reference to data volume or signal-to-noise. The interviewer wants a judgement, not a preference. Also name the extrapolation limitation of trees, because that is the specific reason they fail in a regime the training set never saw, which is when you most need the model.
Expect next
- How would you stop a boosted model overfitting on financial data?
- Why can trees not extrapolate, and when does that hurt you?
- How much out-of-sample improvement would make you switch from the linear model?
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

