Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
039What are the differences between Lasso and Ridge regression?Tower Research CapitalTrading · Princeton · 2018
Say this
Both add a penalty on coefficient size to trade variance for bias. Ridge penalises the sum of squares and shrinks everything smoothly towards zero without eliminating anything. Lasso penalises the sum of absolute values and sets coefficients exactly to zero, so it selects features.
Then walk it
- The geometry explains it. The L1 constraint region is a diamond with corners on the axes, so the solution tends to land on a corner, which means a zero coefficient. The L2 region is a ball with no corners, so solutions are interior and nothing is exactly zero.
- Ridge has a closed form, beta equals (X'X plus lambda I) inverse X'y, which is why it also fixes a singular X'X. Lasso has no closed form and needs coordinate descent or LARS.
- Correlated predictors behave very differently. Ridge splits the weight across a group of correlated features, which is stable. Lasso arbitrarily picks one and zeroes the rest, which is unstable across samples. Elastic net, which mixes both penalties, exists precisely to get sparsity without that instability.
- In a Bayesian reading, ridge is a Gaussian prior on the coefficients and lasso is a Laplace prior. The Laplace prior's spike at zero is what produces exact zeros.
- What I would say about which to use on financial data: predictors are usually highly correlated and the signal-to-noise ratio is awful, so ridge or elastic net typically beats pure lasso out of sample. Lasso is attractive when you need an interpretable short list of factors, but do not confuse the features it selected with the features that matter, because a slightly different sample gives you a different list.
Where candidates lose it
Stopping at L1 gives sparsity, L2 does not. Everyone says that. The differentiators are the diamond-versus-ball geometry, the behaviour under correlated predictors, and the Bayesian priors. Also always say that both require standardised features, because the penalty is scale-dependent and forgetting to standardise silently ruins the fit.
Expect next
- What is elastic net for?
- How do you choose lambda?
- Why do you have to standardise your features first?
Reported by candidates at Tower Research Capital (Trading, Princeton, 2018). Source: Wall Street Oasis.
055When would you use gradient boosting on market data, and when would you stick with a linear model?Quant researchQuant trading
Say this
Boosting earns its keep when you have a lot of data, genuine non-linearity and interactions, and a target with enough signal to support the extra capacity. For low-frequency return prediction with a few hundred monthly observations I would use a regularised linear model almost every time.
Then walk it
- The case for trees: they capture interactions and thresholds automatically, handle mixed feature types, are insensitive to monotone transforms, and do not care about outliers in the features. On microstructure problems with millions of observations they genuinely win.
- The case against on returns: the signal-to-noise is so low that a flexible learner mostly memorises noise, and the model cannot extrapolate beyond the range it saw, which is exactly where the interesting market states live. A boosted tree trained through 2019 has no representation of March 2020.
- Data volume is the deciding variable. Daily cross-sectional data with 3,000 names times 20 years is 15 million rows and trees are viable. Monthly aggregate time series with 300 observations is not, no matter how you tune it.
- If I use boosting, I use it with heavy constraints: shallow trees of depth three to five, low learning rate, strong subsampling, early stopping on a purged time-series split, and monotonic constraints where I have a prior on the sign.
- And I would always run the regularised linear baseline first and report both. In practice the boosted model often adds a modest amount of out-of-sample R squared over a good linear model on financial data, which is a real gain but far from the step change people expect. Knowing that the gain is modest rather than transformative is the useful piece of experience here.
Where candidates lose it
Defaulting to whatever is fashionable with no reference to data volume or signal-to-noise. The interviewer wants a judgement, not a preference. Also name the extrapolation limitation of trees, because that is the specific reason they fail in a regime the training set never saw, which is when you most need the model.
Expect next
- How would you stop a boosted model overfitting on financial data?
- Why can trees not extrapolate, and when does that hurt you?
- How much out-of-sample improvement would make you switch from the linear model?
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

