Quant interview preparation
Prop market making and quantitative research, weighted the way the interviews actually are: probability and expected value, statistics and machine learning, market making logic, programming and options. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it, and every probability answer shows the reasoning path rather than just the number.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 53
- Firms
- 15
- Updated
- September 2026
022If X and Y are dependent, does that tell you anything about the relationship between X and Z?Tower Research CapitalProp Trading · New York · 2019
Say this
Nothing at all. Dependence is not transitive and it says nothing about a third variable you have not mentioned. X can be dependent on Y and completely independent of Z.
Then walk it
- Trivial counterexample: let X and Y be the same fair coin and let Z be a separate independent coin. X and Y are maximally dependent, X and Z are independent.
- The deeper point is that even if X depends on Y and Y depends on Z, X need not depend on Z. Let Y be X plus Z with X and Z independent. Y is dependent on both, and X and Z remain independent of each other.
- Correlation is a bit more constrained than dependence because the correlation matrix must be positive semi-definite. If corr(X,Y) is 0.9 and corr(Y,Z) is 0.9, then corr(X,Z) is bounded below by about 0.62. So high correlations do restrict the third pair, but only through that PSD constraint, and dependence in general carries no such bound.
- The formula for the bound: rho_xz is at least rho_xy times rho_yz minus the square root of (1 minus rho_xy squared)(1 minus rho_yz squared). Plug in 0.9 and 0.9 and you get 0.81 minus 0.19, which is 0.62.
- Why this matters on a desk: people assume that if two assets both correlate with a factor they must correlate with each other. If the loadings are moderate, say 0.5 and 0.5, the bound is minus 0.5, so they can be strongly negatively correlated. That mistake shows up in risk models constantly.
Where candidates lose it
Answering yes because it feels like dependence should chain. Give the counterexample in one breath, then earn the extra credit with the correlation bound, because the interviewer's follow-up is almost always the correlation version. And be precise that zero correlation does not mean independence, only the converse holds.
Expect next
- Now with correlations. If corr(X,Y) is 0.9 and corr(Y,Z) is 0.9, what do you know about corr(X,Z)?
- Give me an example of zero correlation with strong dependence.
- What is conditional independence and why does it matter for factor models?
Reported by candidates at Tower Research Capital (Prop Trading, New York, 2019). Source: Wall Street Oasis.
039What are the differences between Lasso and Ridge regression?Tower Research CapitalTrading · Princeton · 2018
Say this
Both add a penalty on coefficient size to trade variance for bias. Ridge penalises the sum of squares and shrinks everything smoothly towards zero without eliminating anything. Lasso penalises the sum of absolute values and sets coefficients exactly to zero, so it selects features.
Then walk it
- The geometry explains it. The L1 constraint region is a diamond with corners on the axes, so the solution tends to land on a corner, which means a zero coefficient. The L2 region is a ball with no corners, so solutions are interior and nothing is exactly zero.
- Ridge has a closed form, beta equals (X'X plus lambda I) inverse X'y, which is why it also fixes a singular X'X. Lasso has no closed form and needs coordinate descent or LARS.
- Correlated predictors behave very differently. Ridge splits the weight across a group of correlated features, which is stable. Lasso arbitrarily picks one and zeroes the rest, which is unstable across samples. Elastic net, which mixes both penalties, exists precisely to get sparsity without that instability.
- In a Bayesian reading, ridge is a Gaussian prior on the coefficients and lasso is a Laplace prior. The Laplace prior's spike at zero is what produces exact zeros.
- What I would say about which to use on financial data: predictors are usually highly correlated and the signal-to-noise ratio is awful, so ridge or elastic net typically beats pure lasso out of sample. Lasso is attractive when you need an interpretable short list of factors, but do not confuse the features it selected with the features that matter, because a slightly different sample gives you a different list.
Where candidates lose it
Stopping at L1 gives sparsity, L2 does not. Everyone says that. The differentiators are the diamond-versus-ball geometry, the behaviour under correlated predictors, and the Bayesian priors. Also always say that both require standardised features, because the penalty is scale-dependent and forgetting to standardise silently ruins the fit.
Expect next
- What is elastic net for?
- How do you choose lambda?
- Why do you have to standardise your features first?
Reported by candidates at Tower Research Capital (Trading, Princeton, 2018). Source: Wall Street Oasis.
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

