Risk Management interview preparation
Market, credit and operational risk, plus model validation, regulatory capital, liquidity and ALM, the statistical foundations and the Indian regulatory syllabus. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it — and answers lead with the point, then the mechanism, then the limitation.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 37
- Firms
- 12
- Updated
- September 2026
066What does it mean if an estimator is BLUE?UBSRisk Management · Zurich · 2021
Say this
Best Linear Unbiased Estimator. Among all estimators that are linear in the data and unbiased, it has the smallest variance. That's the Gauss-Markov result: ordinary least squares is BLUE provided a specific set of assumptions holds.
Then walk it
- Unpack each word, because that's what the question is testing. Linear in the observations. Unbiased, so its expected value equals the true parameter. Best, meaning minimum variance within that class.
- The Gauss-Markov conditions: correct linear specification, errors with zero conditional mean, homoskedasticity, no autocorrelation, and no perfect multicollinearity. Notice normality is not required for BLUE. You need normality for the t and F tests in small samples, not for OLS to be efficient.
- The restriction that matters is 'linear'. A biased or non-linear estimator can easily beat OLS on mean squared error. Ridge and lasso are deliberately biased and often predict better, and James-Stein shrinkage famously dominates the sample mean. So BLUE is optimality within a box, not optimality.
- In finance the assumptions fail routinely. Returns are heteroskedastic and volatility clusters, so OLS stays unbiased but the standard errors are wrong, which means your t-statistics lie. That's the practical consequence and it's the one to lead with when asked what breaks.
- The fixes: White or Newey-West robust standard errors for heteroskedasticity and autocorrelation, generalised least squares if you know the error structure, and instrumental variables if the regressor is endogenous. Endogeneity is the serious one, because it destroys unbiasedness rather than just efficiency.
- The distinction to keep straight: heteroskedasticity and autocorrelation cost you efficiency and valid inference. Omitted variables, measurement error in a regressor and simultaneity cost you unbiasedness. Those are different problems needing different fixes.
- So the answer I'd close with: BLUE is a useful benchmark and a weak guarantee. In a risk model I care more about whether the specification is right and whether the relationship is stable than about being efficient within the linear unbiased class.
Where candidates lose it
Expanding the acronym and stopping, or claiming normality is a Gauss-Markov requirement. It isn't. The two answers that separate candidates are that 'best' is only within linear unbiased estimators, so biased shrinkage estimators can beat it, and that in finance the binding violation is heteroskedasticity making your standard errors wrong.
Expect next
- Is normality required for OLS to be BLUE?
- Which Gauss-Markov assumption fails most often in financial data?
- Can a biased estimator ever be preferable?
Reported by candidates at UBS (Risk Management, Zurich, 2021). Source: Wall Street Oasis.
067Would you model equity returns as normal or Student's t, and what difference does it make?Bank market riskBuy-side risk
Say this
Student's t, with something like four to six degrees of freedom for daily equity returns. Real returns are leptokurtic, so a normal assumption systematically understates the tail, and the understatement gets worse the further out you go.
Then walk it
- The evidence is simple and worth quoting: daily equity index returns have excess kurtosis well above zero, often 3 to 8, against zero for a normal. The 1987 crash was more than 20 standard deviations under a normal fitted to prior data, which that distribution says should never happen in the age of the universe.
- A t with low degrees of freedom has polynomial rather than exponential tail decay, which fits observed extremes far better. Around four to six degrees of freedom is the usual empirical range for daily equity data, and it converges to normal as the degrees of freedom rise.
- The practical consequence for VaR: at 95 percent the two distributions give similar answers, and at 99.9 percent the normal can understate by a large multiple. So the choice barely matters for routine reporting and matters enormously for capital and stress.
- The second issue is that the unconditional fat tail is partly a mixture effect. Returns are closer to normal conditional on the current volatility regime, and it's volatility clustering that produces fat unconditional tails. So a GARCH model with normal innovations gets you a long way, and GARCH with t innovations gets you further.
- For a multivariate problem it gets harder. A multivariate t gives you tail dependence, which a multivariate normal does not, and tail dependence is the property that matters in a crisis. That's the bridge to copulas.
- Where the t is not enough: extreme value theory, fitting a generalised Pareto distribution to the exceedances above a threshold, is the right tool if you genuinely only care about the far tail. It uses less data more efficiently for that specific job.
- The caveat I'd give: a t distribution has more parameters and the degrees of freedom estimate is unstable in small samples. And skew matters too, since equity downside tails are fatter than upside ones, so a skewed t or an empirical distribution is often the pragmatic choice.
Where candidates lose it
Saying 'returns aren't normal' without knowing what to use instead or how big the error is. Naming a plausible degrees-of-freedom range, and saying that the choice matters at 99.9 percent but hardly at 95, shows you've fitted these to real data rather than repeating a slogan.
Expect next
- How would you estimate the degrees of freedom?
- Does GARCH with normal innovations solve the problem?
- When would you use extreme value theory instead?
068Correlation versus dependence. Why does correlation fail exactly when you need it?Bank market riskModel validation
Say this
Pearson correlation measures only the linear component of dependence, and it's a single average number. Dependence is the whole joint distribution. So two variables can be uncorrelated and completely dependent, and correlation tells you nothing about whether they move together in the tail.
Then walk it
- The clean counterexample: Y equals X squared with X symmetric around zero. Perfectly dependent, zero linear correlation. Any non-monotonic relationship defeats Pearson.
- It's also not invariant to non-linear transforms, which matters because option payoffs are non-linear transforms of the underlying. Rank measures like Kendall's tau and Spearman's rho are invariant and are what you should use for dependence.
- The failure that costs money is that correlation is an average over the whole distribution, dominated by the many ordinary days. Tail dependence, the probability that both variables are extreme together, is a separate property that Pearson does not capture at all.
- Two distributions can share the same correlation matrix and have utterly different joint tails. A Gaussian copula has zero asymptotic tail dependence: extreme joint events are asymptotically independent. A t copula has positive tail dependence. Same correlation, completely different stress behaviour.
- Then correlations are non-stationary and they rise in a crisis. Equity pair correlations that sit around 0.3 in calm markets go to 0.8 in a sell-off. Diversification built on the calm number evaporates precisely when you were counting on it.
- Part of that rise is a statistical artefact worth knowing about: conditioning on large moves mechanically increases measured correlation even with a stable underlying joint distribution. So not all of the observed increase is a regime change, and that subtlety is a genuinely strong thing to say.
- What I'd do instead: use rank correlations for dependence, look at exceedance correlations conditional on large moves, use a copula with tail dependence when I need a joint distribution, and stress correlations to one in the scenario rather than trusting the estimate.
- Real cost: the 2007 quant equity deleveraging and the 2008 structured credit losses were both failures of dependence assumptions rather than of the individual marginal distributions.
Where candidates lose it
Saying 'correlation goes to one in a crisis' as the whole answer. That's true and it's the easy half. The deeper point is that correlation and dependence are different objects, that two joint distributions can share a correlation matrix and differ entirely in the tail, and that part of the observed crisis rise is a conditioning artefact.
Expect next
- What is tail dependence?
- Why do Gaussian and t copulas differ if they have the same correlation?
- How would you stress correlations in a scenario?
069What is a copula, and what went wrong with the Gaussian copula in 2008?Model validationBank credit risk
Say this
A copula separates the marginal distributions from the dependence structure. Sklar's theorem says any joint distribution can be written as a copula applied to its marginals, so you can model each variable's tail properly and then choose how they move together. The 2008 failure was choosing a Gaussian copula, which has no tail dependence, for a problem that is entirely about tail dependence.
Then walk it
- Mechanically: transform each variable to a uniform through its own cumulative distribution, then model the joint behaviour of those uniforms. That's the copula. It's a clean separation of 'how fat is each tail' from 'do they go bad together'.
- Gaussian copula: parameterised by a correlation matrix, and its asymptotic tail dependence is zero for any correlation below one. So as you go further into the tail, extreme joint events become asymptotically independent. That's the mathematical fact behind the failure.
- t copula: has positive tail dependence controlled by the degrees of freedom, so extremes cluster. Gumbel gives you upper tail dependence only, Clayton lower tail only, which is useful for credit where you care about joint defaults and not joint survivals.
- What actually happened with CDOs. David Li's Gaussian copula model let you price a tranche off a single correlation number, calibrated to historical data from a benign period, often 0.3 for mortgage pools. The senior tranche's value depends almost entirely on the probability that many defaults happen together, which is exactly the quantity the Gaussian copula sets too low.
- So the model said AAA was safe because widespread simultaneous default across regions was nearly impossible. In a national housing downturn, correlation went towards one and losses blew through tranches the model said were remote.
- The market knew before the maths admitted it. The base correlation skew, needing a different correlation for each tranche to fit observed prices, was the model telling you it was wrong, and it was read as a market quirk instead.
- The honest lesson, and the one to say: the failure was not really the copula. It was calibrating a tail parameter to data with no tail in it, using one number to describe dependence across a whole system, and treating a pricing convention as a risk model. A t copula with a bad correlation input would have failed too.
- Practical use today: copulas are standard in economic capital, portfolio credit models and multi-asset VaR. You'd use a t copula, calibrate dependence with attention to stress periods, and stress the dependence parameter rather than point-estimating it.
Where candidates lose it
Blaming 'the formula that killed Wall Street' without being able to say what property was missing. The specific answer is zero asymptotic tail dependence, plus a correlation parameter calibrated on benign data. And the mature closing point is that a better copula with the same bad calibration would also have failed.
Expect next
- Why does the Gaussian copula have zero tail dependence?
- What was base correlation telling the market?
- Would a t copula have prevented it?
070What's the tracking error formula?MSCIFinancial Tools · Monterrey · 2013
Say this
Tracking error is the standard deviation of the difference between portfolio and benchmark returns. Compute active return each period, take its standard deviation, then annualise by multiplying by the square root of the number of periods in a year.
Then walk it
- Formula in words: active return equals portfolio return minus benchmark return each period. Tracking error is the standard deviation of that series. On monthly data you annualise by root twelve, on daily by root 252.
- The ex-post version uses realised returns. The ex-ante version uses a risk model: active weights transposed times the covariance matrix times active weights, then square root. Those two numbers routinely disagree, and explaining the gap is a real part of a buy-side risk job.
- Be careful about mean adjustment. Some definitions use the standard deviation of active returns and others use the root mean square of active returns, which includes the average outperformance. They differ, and you should say which one you mean.
- Feel for the numbers: an index fund runs 5 to 50 basis points, an enhanced index strategy 0.5 to 2 percent, an active core equity fund 3 to 6 percent, and a concentrated high-conviction fund 8 percent or more. Quoting a range is what shows you've looked at real funds.
- What it's used for: mandate limits, since most institutional mandates cap tracking error. And the information ratio, active return divided by tracking error, which is the risk-adjusted measure of skill and the number that actually matters.
- Decomposition is the useful part. Break tracking error into factor and specific contributions: how much comes from a systematic style or sector tilt versus stock selection. A manager paid for stock picking who is running most of their tracking error on an unintended country bet has a problem, and that's exactly the conversation a risk analyst has with a PM.
- Limitations to volunteer: it's symmetric, so it penalises outperformance identically to underperformance. It assumes a stable covariance structure, so realised tracking error jumps in a crisis. And it's backward-looking, so a manager who has just changed their positioning has a stale number.
Where candidates lose it
Giving the formula and nothing else. This question is a screen: the formula takes five seconds, and the rest of your answer is what's being assessed. Knowing typical ranges by strategy, the ex-ante versus ex-post gap, and factor decomposition is what converts a definition into an answer.
Expect next
- What tracking error would you expect from an index fund?
- Why do ex-ante and ex-post tracking error differ?
- What is the information ratio?
Reported by candidates at MSCI (Financial Tools, Monterrey, 2013). Source: Wall Street Oasis.
071Here is a book of credit exposures and a default history. Analyse it using dynamic probability metrics.Jane StreetCredit Risk · London · 2025
Say this
Dynamic means the probabilities have to be conditional and updating, not a static average. So I'd build a hazard-rate view: probability of default in the next period given survival so far, conditioned on observable state, and then update it as new information arrives.
Then walk it
- Start with the data audit before any modelling. Definition of default, observation window, censoring, survivorship, and whether exposures enter and leave the sample. Most of the wrong answers in credit analytics come from the panel being constructed badly, not from the maths.
- Then the right framing: survival analysis rather than a single-period classification. Estimate a hazard function, the instantaneous default rate conditional on having survived, using a Cox proportional hazards or discrete-time hazard model. That handles censoring properly and gives you a term structure of default rather than one number.
- Make it conditional on state. Time-varying covariates: rating migration, spread level, utilisation, macro variables. A default probability that changes when the world changes is what 'dynamic' means here.
- Then update sequentially. Bayesian updating, or a Kalman-filtered latent credit factor, so each new month of data revises the estimate rather than triggering a full refit. For sparse default data, a Bayesian approach with an informative prior from external data is far more stable than maximum likelihood on twelve observations.
- Dependence, which is what makes a credit book different from a set of single names. Estimate a common factor and its loading, because portfolio loss is driven by correlation, not by average PD. Report the loss distribution, the 99th percentile and expected shortfall, not just expected loss.
- Validation appropriate to sparse data: time-series calibration tests, a binomial or Vasicek test given the low default counts, discriminatory power via a time-dependent AUC, and a comparison against market-implied hazard rates from CDS where they exist.
- Then say what you'd report. Not a single PD. A term structure of conditional default probabilities, the portfolio loss distribution with tail measures, the top contributors to tail loss, and an explicit statement of how much of the answer is driven by the correlation assumption.
- And the honest limitation: with a short history and few defaults, the tail of the loss distribution is an assumption rather than an estimate. I'd say that up front rather than present a confident 99.9th percentile.
Where candidates lose it
Reaching for a classifier and reporting AUC. The word 'dynamic' is doing the work in the question: it's asking for conditional, time-varying, updating probabilities, which means hazard models and Bayesian updating, not a static logistic fit. And in a credit book the correlation assumption drives the tail more than the PD does.
Expect next
- How would you handle the sparsity of defaults?
- How would you estimate the common factor loading?
- How would you present the uncertainty in the tail?
Reported by candidates at Jane Street (Credit Risk, London, 2025). Source: Wall Street Oasis.
072A fraud test is 99 percent accurate and fraud occurs in one transaction in ten thousand. The test flags a transaction. What is the chance it's really fraud?Operational riskModel validation
Say this
About one percent. Out of a million transactions, 100 are fraud and the test catches 99 of them, but it also falsely flags one percent of the 999,900 clean ones, which is about 9,999. So 99 true positives against roughly 10,000 flags means a 1 percent hit rate.
Then walk it
- Set it up with counts rather than Bayes' formula. A million transactions, 100 frauds, 999,900 clean. It's faster and you won't fumble the algebra out loud.
- True positives: 99 percent of 100, so 99. False positives: 1 percent of 999,900, so 9,999. Total flags about 10,098, of which 99 are real. That's 0.98 percent.
- The lesson is the base rate. When the event is rare, even a very accurate test produces overwhelmingly false alarms, because the clean population is so much larger. This is the base rate fallacy and it's the entire point of the question.
- This is not a puzzle, it's the daily reality of transaction monitoring and sanctions screening. Real AML alert systems run false positive rates above 95 percent, which is why banks employ thousands of people to clear alerts, and it's a genuine operational risk and cost problem.
- So the design conclusion: for rare events, headline accuracy is the wrong metric. You care about precision and recall, and about the cost asymmetry between a missed fraud and an investigated false alarm. Then you tune the threshold to that cost, not to accuracy.
- How you improve it in practice: raise the prior before you apply the test by segmenting on risk, so you're testing a population with a much higher base rate. Or stack models so an expensive accurate check only runs on things a cheap screen flagged. Both raise precision far more than improving the test itself would.
- And the number worth remembering as a reasonableness check: with a 1-in-10,000 base rate you need a false positive rate around 1 in 10,000 to get to a 50 percent hit rate. That is a far harder test than 99 percent accurate.
Where candidates lose it
Answering 99 percent. That's the reflex answer and it's what the question is designed to catch. Use counts on a million, and then draw the operational conclusion about alert volumes, because in a risk interview the business implication is worth as much as the arithmetic.
Expect next
- What accuracy would you need for a 50 percent hit rate?
- How would you reduce false positives in practice?
- How do you set the threshold if a missed fraud costs 500 times an investigation?
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

