Risk Management interview preparation
Market, credit and operational risk, plus model validation, regulatory capital, liquidity and ALM, the statistical foundations and the Indian regulatory syllabus. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it — and answers lead with the point, then the mechanism, then the limitation.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 37
- Firms
- 12
- Updated
- September 2026
047What is model risk?UBSRisk Management · Zurich · 2021
Say this
Model risk is the risk of loss from using a model that's wrong, or from using a right model in the wrong place. Two sources, and the second is the bigger one in practice: fundamental errors in the model itself, and correct models applied outside the conditions they were built for.
Then walk it
- The US Federal Reserve's SR 11-7 definition is the one to quote, because it splits it exactly that way: errors in design, and incorrect or inappropriate use.
- The error side includes bad theory, bad data, coding bugs and bad calibration. It's the side people think of and it's the side validation catches most easily.
- The misuse side is the one that hurts. A model calibrated on investment grade credit applied to high yield. A pricing model used for risk. A VaR model built for a linear book applied once options were added. Nothing is wrong with the model; the use is wrong.
- It compounds through the chain. Models feed models: a PD model feeds ECL, which feeds capital planning, which feeds the dividend decision. An error at the bottom is unrecognisable four steps up, which is why model inventories and dependency maps exist.
- Real examples worth naming: the Gaussian copula in structured credit, where the model was fine and the correlation assumption was not. The 2012 JPMorgan CIO losses, where a spreadsheet error and a newly approved VaR model both featured. Long-Term Capital Management, where the model was right about relationships and wrong about liquidity and leverage.
- How you manage it: an inventory of every model with a tier, independent validation proportionate to that tier, ongoing performance monitoring, documented limitations, and an owner. And the control that matters most is the simplest, writing down what the model may not be used for.
- The limitation to volunteer: you can't eliminate model risk, only bound it. The mitigant with the best return is not more validation, it's a stated range of applicability and a human who understands the model sitting between it and a decision.
Where candidates lose it
Defining it as 'the model being wrong'. That's half of it, and the smaller half. The answer that lands names misuse of a correct model as the larger source, and gives a concrete case. If you can cite SR 11-7, do, because it signals you've worked near a validation function.
Expect next
- Give me an example of a correct model used wrongly.
- How would you tier a model inventory?
- Can you eliminate model risk?
Reported by candidates at UBS (Risk Management, Zurich, 2021). Source: Wall Street Oasis.
048How would you validate a model?Model validationGlobal capability centres
Say this
Three pillars: conceptual soundness, outcomes analysis and ongoing monitoring. So does the theory make sense for this use, does it perform against reality, and will you know when it stops working. And it has to be done by someone independent of whoever built it.
Then walk it
- Conceptual soundness first, and it's the part that gets skipped. Read the documentation, check the theory is appropriate for the intended use, check the assumptions are stated and reasonable, and review the data: source, quality, representativeness, and whether the development sample looks like today's population.
- Then replicate. Independently rebuild at least the core of it from the documentation. If you can't reproduce the results from the document, the documentation fails, and that's a finding in itself.
- Outcomes analysis: backtesting against realised outcomes, benchmarking against an alternative model or a simpler challenger, and sensitivity analysis to see which inputs the output actually depends on. Stress the inputs to the edge of plausibility and see if it breaks gracefully or catastrophically.
- Then the boundary work: what is this model not valid for. A validated model with no stated limitations is a hazard, because the next user will apply it to something new and assume it's approved.
- Ongoing monitoring: performance thresholds, population stability, and a revalidation cycle tiered by materiality. Tier 1 models annually, lower tiers less often, and any material change triggers a revalidation regardless of the cycle.
- Governance: findings rated by severity, owners and deadlines, and a model approval that can be conditional or refused. A validation function that has never refused an approval isn't independent, and that's the question I'd ask about any validation team I joined.
- Sizing it honestly: full validation of a Tier 1 pricing model is weeks of work for two people. Proportionality is the whole design problem, because validating everything to the same depth means validating nothing well.
Where candidates lose it
Going straight to backtesting. Backtesting is one third of it and it's the third that needs data you often don't have. Conceptual soundness and the explicit statement of limitations are what prevent the misuse that causes most model losses. And say the word independent, because organisational independence is the first thing a supervisor checks.
Expect next
- What would you do if you couldn't backtest because there was no data?
- How do you validate a vendor model you can't see inside?
- How would you tier models for validation intensity?
049What's the difference between backtesting and benchmarking a model?Model validationBank market risk
Say this
Backtesting compares the model to reality. Benchmarking compares it to another model. One tells you whether you're right, the other tells you whether you're different, and you need benchmarking precisely when reality doesn't give you enough observations to backtest.
Then walk it
- Backtesting is the gold standard because the comparator is truth: forecast versus realised outcome. VaR exceptions against actual P&L, predicted default rates against observed defaults, predicted prepayment against actual.
- Its constraint is data. At 99% confidence you get about 2.5 exceptions a year, so a year of data can't distinguish a good model from a mediocre one. For a low-default portfolio, sovereigns or large corporates, you might have zero defaults in a decade, so backtesting is simply unavailable.
- Benchmarking fills that gap. Run a challenger model, a vendor model, or a simple closed-form approximation on the same portfolio and compare. Large unexplained divergence is a finding even if you can't say which one is right.
- It's also how you test parts of a model you can't observe. You can't observe a 20-year lifetime PD, but you can compare your curve to an agency cumulative default table or to CDS-implied hazard rates.
- The important limitation: benchmarking two models that share an assumption tells you nothing. If both assume a normal distribution, they'll agree and both be wrong. The benchmark has to be structurally different to be informative, and that's the part people get wrong.
- So they answer different questions. Backtesting: is the model calibrated? Benchmarking: is the model an outlier, and can I explain why? A full validation uses both plus sensitivity analysis, and leans on benchmarking exactly where data is thin.
- Practical example: for a low-default corporate portfolio, you'd benchmark the PD model against external ratings and market-implied PDs, use a binomial or Vasicek test for the little default data you have, and lean on the qualitative review. That combination is what a supervisor expects to see.
Where candidates lose it
Using the words interchangeably. And missing the key insight that benchmarking is only informative if the benchmark makes different assumptions. Two models with the same flawed assumption will agree beautifully, and a candidate who says that has thought about validation rather than memorised a checklist.
Expect next
- How would you validate a PD model for a portfolio with no defaults?
- What makes a good benchmark model?
- Your model and the benchmark differ by 40 percent. Now what?
050Explain what a Kalman filter is.UBSRisk · London · 2022
Say this
It's a recursive estimator for a hidden state you can only observe with noise. Each period you predict the state forward with your model, then correct that prediction with the new observation, weighting the two by how much you trust each. Under linear-Gaussian assumptions it's the optimal estimator.
Then walk it
- Two equations. A state equation for how the unobserved thing evolves, and a measurement equation linking the state to what you actually see, each with its own noise.
- Two steps per period. Predict: roll the state and its uncertainty forward. Update: compute the surprise, the difference between the observation and what you expected, and move your estimate toward it by the Kalman gain.
- The gain is the whole intuition. If measurement noise is large relative to state uncertainty, the gain is small and you mostly trust your model. If your state uncertainty is large, the gain is large and you mostly trust the new data. It's Bayesian updating with the arithmetic done for you.
- Where it's used in finance: extracting a time-varying beta or hedge ratio, estimating a stochastic volatility or unobserved factor, filtering a fair-value or pairs-trading spread, term structure models where the factors are latent, and nowcasting a macro variable from noisy high-frequency data.
- Why a risk function cares: it gives you an estimate that adapts without the jumpiness of a rolling window. A 60-day rolling beta lurches when an old observation drops out; a Kalman-filtered beta moves smoothly and quantifies its own uncertainty.
- The assumptions and their cost: linear dynamics and Gaussian noise. For non-linear problems you need the extended or unscented variants or a particle filter. And you have to specify the two noise covariances, which are rarely known, so in practice you estimate them by maximum likelihood and the result is sensitive to them.
- The limitation to volunteer: it's optimal given the model, and it has no way to tell you the state equation is wrong. Feed it a misspecified process and it will produce confident, smooth, wrong estimates, which is a particularly dangerous failure mode.
Where candidates lose it
Reciting matrix equations. Nobody wants the algebra; they want the predict-then-correct intuition, the gain as a trust weighting, and one concrete financial use. If you can't name a use case, the answer reads as memorised from a signal-processing course.
Expect next
- How would you use it to estimate a time-varying hedge ratio?
- What happens if the noise covariances are misspecified?
- How does it compare to a simple exponentially weighted estimate?
Reported by candidates at UBS (Risk, London, 2022). Source: Wall Street Oasis.
051You are validating a gradient boosting credit model that beats the existing logistic scorecard by eight Gini points. Do you approve it?Model validationBank credit risk
Say this
Not on the Gini alone. Eight points of discrimination is worth having, but I'd need calibration, stability, explainability and fair-lending testing before approving it, and I'd want to know whether the gain survives out of time rather than just out of sample.
Then walk it
- First question: is the eight points real? Check for leakage, which is the most common cause of a suspiciously strong challenger. Any feature that encodes the outcome, a post-application field, a collections flag, a date artefact, and the gain evaporates.
- Second: out of time, not just out of sample. Boosted models overfit to the period as well as to the sample. If the gain is eight points on a random split and two points on a later year, the story changes completely.
- Third: calibration. Gradient boosting ranks well and is often badly calibrated in the extremes, which is where pricing and provisioning live. Check the predicted-versus-observed curve by decile and consider isotonic or Platt scaling.
- Fourth: monotonicity and explainability. A credit model has to survive being explained to a customer who was declined and to a regulator. Unconstrained boosting can learn that higher income increases risk in some segment, which is a spurious interaction you can't defend. Monotonic constraints usually cost very little Gini and buy a lot of defensibility.
- Fifth: fairness. Test outcomes across protected characteristics and proxies for them. Complex models find proxies more efficiently than simple ones, so this risk genuinely rises with model power.
- Sixth: operational reality. Feature pipeline stability, retraining cadence, latency, reproducibility, version control, and whether anyone can support it in three years when the builder has left. Model risk includes the risk that nobody understands the production model.
- So my recommendation would be conditional approval with constraints: monotonic constraints on the key variables, calibration layer on top, capped score-level overrides, tightened monitoring thresholds, and the logistic model retained as a live benchmark. That is a real validation outcome rather than a yes or a no.
- And the commercial framing to say out loud: eight Gini points on a large retail book is worth real money, so the answer isn't to refuse complexity. It's to price the governance cost and decide deliberately.
Where candidates lose it
Picking a side. Reflexively rejecting machine learning makes you look like an obstacle; approving it on Gini alone makes you look like you've never validated anything. The answer is conditional approval with named conditions, and leakage plus out-of-time degradation are the two checks that must come first.
Expect next
- How would you test for leakage?
- What would you tell a declined customer?
- How much Gini would you give up for monotonicity?
052How do you detect overfitting in a risk model?Model validationBuy-side risk
Say this
The signature is a large gap between in-sample and out-of-sample performance, and for financial models specifically between in-sample and out-of-time. If accuracy collapses on data from a later period, you've fitted the era, not the relationship.
Then walk it
- Standard test: hold out data, or cross-validate. But for time series you need forward-chaining rather than random k-fold, because random folds leak the future into the training set and make everything look good.
- The financial-specific test is out of time. Fit on 2015 to 2019, test on 2021 to 2023. Random splits from the same period share the same regime, so they flatter the model badly.
- Structural warning signs: parameter count relative to observations, coefficients with implausible signs or magnitudes, a variable that only works in one sub-period, and results that change materially when you drop a single year.
- For strategy or signal work, the killer is multiple-testing bias. If you tried 500 specifications, the best one will look great by construction. The deflated Sharpe ratio and the idea of testing on a truly held-out period exist precisely for that.
- Stability diagnostics: rolling-window coefficient estimates. A genuine relationship has coefficients that wobble; an overfitted one has coefficients that flip sign. And bootstrap the performance metric to see how wide the confidence interval actually is.
- The cultural control, which matters more than any statistic: limit how many times you look at the holdout. Every peek at the test set converts it into training data, and in practice that's how overfitting enters a model that passed every formal test.
- And the counterweight to state: underfitting is also a failure, and a heavily constrained model that misses real non-linearity costs money too. The judgement is whether added complexity buys performance that survives out of time.
Where candidates lose it
Naming cross-validation and stopping. For financial models, random cross-validation is itself a source of false confidence because of regime and look-ahead leakage. Out-of-time testing and multiple-testing bias are the two points that show you've done this on real data.
Expect next
- Why is random k-fold cross-validation dangerous for time series?
- How would you account for having tried 200 specifications?
- How would you tell overfitting apart from a regime change?
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

