Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
039What are the differences between Lasso and Ridge regression?Tower Research CapitalTrading · Princeton · 2018
Say this
Both add a penalty on coefficient size to trade variance for bias. Ridge penalises the sum of squares and shrinks everything smoothly towards zero without eliminating anything. Lasso penalises the sum of absolute values and sets coefficients exactly to zero, so it selects features.
Then walk it
- The geometry explains it. The L1 constraint region is a diamond with corners on the axes, so the solution tends to land on a corner, which means a zero coefficient. The L2 region is a ball with no corners, so solutions are interior and nothing is exactly zero.
- Ridge has a closed form, beta equals (X'X plus lambda I) inverse X'y, which is why it also fixes a singular X'X. Lasso has no closed form and needs coordinate descent or LARS.
- Correlated predictors behave very differently. Ridge splits the weight across a group of correlated features, which is stable. Lasso arbitrarily picks one and zeroes the rest, which is unstable across samples. Elastic net, which mixes both penalties, exists precisely to get sparsity without that instability.
- In a Bayesian reading, ridge is a Gaussian prior on the coefficients and lasso is a Laplace prior. The Laplace prior's spike at zero is what produces exact zeros.
- What I would say about which to use on financial data: predictors are usually highly correlated and the signal-to-noise ratio is awful, so ridge or elastic net typically beats pure lasso out of sample. Lasso is attractive when you need an interpretable short list of factors, but do not confuse the features it selected with the features that matter, because a slightly different sample gives you a different list.
Where candidates lose it
Stopping at L1 gives sparsity, L2 does not. Everyone says that. The differentiators are the diamond-versus-ball geometry, the behaviour under correlated predictors, and the Bayesian priors. Also always say that both require standardised features, because the penalty is scale-dependent and forgetting to standardise silently ruins the fit.
Expect next
- What is elastic net for?
- How do you choose lambda?
- Why do you have to standardise your features first?
Reported by candidates at Tower Research Capital (Trading, Princeton, 2018). Source: Wall Street Oasis.
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

