Risk Management interview preparation
Market, credit and operational risk, plus model validation, regulatory capital, liquidity and ALM, the statistical foundations and the Indian regulatory syllabus. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it — and answers lead with the point, then the mechanism, then the limitation.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 37
- Firms
- 12
- Updated
- September 2026
044Design three key risk indicators for a payments operation, and tell me what makes a KRI good.Operational riskGlobal capability centres
Say this
A good KRI is leading, measurable without manual effort, and has a threshold that triggers a specific action. For payments I'd use the unreconciled item count and ageing, the manual intervention rate on straight-through processing, and the failed or returned payment rate by corridor.
Then walk it
- Unreconciled items over two days old, by value and count. It's leading, because loss events start as breaks nobody chased, and it's cheap to produce from the reconciliation system.
- Manual touch rate on payments that should be straight-through. Every manual touch is a keystroke error waiting to happen, so this is a direct proxy for the frequency of process losses. A rise from 2 to 6 percent is a red flag before any loss appears.
- Failed and returned payment rate, split by corridor and by cause. It picks up upstream data quality problems, sanctions-screening false positives and correspondent bank issues, each of which needs a different fix.
- What makes a KRI good: leading not lagging, objectively measurable from a system rather than from a survey, sensitive enough that it actually moves, owned by someone with the authority to act, and attached to an amber and red threshold with a pre-agreed response.
- What makes a bad one: loss count, which is lagging and tells you the failure already happened. Headcount, which is context rather than risk. Anything requiring a manual monthly collection, because it degrades into a copy-paste exercise.
- Thresholds should be calibrated off the historical distribution, not picked round. Amber at roughly the 90th percentile of the last two years, red at the 99th, then reviewed annually. And each threshold needs a named action, or breaching it changes nothing.
- The failure mode to name: KRI inflation. A dashboard with 120 indicators gets ignored. Eight to twelve real ones per business, reviewed monthly by someone who can act, beats a hundred reported to nobody.
Where candidates lose it
Proposing lagging indicators. Loss count, number of incidents and audit findings are all after the fact, and they are what most candidates offer. A KRI is supposed to give you time to act, so lead with something that moves before the loss, and attach a threshold and an action to each.
Expect next
- How would you calibrate the thresholds?
- What do you do when a KRI turns red?
- Give me a KRI for cyber risk.
046A trader has breached his VaR limit three times this month, and each time got approval after the fact. What do you do?Operational riskBank market risk
Say this
Three retrospective approvals isn't a limit breach problem, it's a control failure. The limit has effectively been replaced by a negotiation. I'd document the pattern, escalate it as a governance issue rather than three separate incidents, and force a decision: either the limit is wrong or the behaviour is.
Then walk it
- First establish the facts precisely: what the limit is, the size and duration of each breach, who approved each one, whether the approver had authority, and whether the approvals were documented at the time or reconstructed afterwards. That last detail changes the nature of the issue completely.
- Then separate the two possible root causes. Either the limit is miscalibrated for a legitimate business, in which case the fix is a properly approved limit increase through the right committee. Or the trader is running more risk than the firm sanctioned, in which case it's a discipline matter.
- The crucial reframing: three ad hoc approvals in a month means the limit is no longer a control. A pre-approved excess is a limit; a post-approved excess is an apology. Say that sentence in the interview, because it's the point of the question.
- Escalation route: my head of risk and the market risk committee, not a quiet conversation with the desk head who has been signing the approvals. The approver is part of what needs reviewing, so escalating to them alone is the mistake.
- Then the pattern question. Look for other symptoms: end-of-day position reductions that reverse the next morning, P&L volatility inconsistent with reported risk, stale or hard-to-verify marks on illiquid positions. Limit breaches with cooperative approvals are a classic precursor, and every large rogue trading loss has this shape in hindsight.
- Consequences and record. In a bank this is reportable to the risk committee, it should appear in the operational risk event log, and it belongs in the trader's performance file. If nothing happens, the next breach is certain.
- And the systemic fix: hard-coded pre-trade blocks rather than post-trade reporting, a rule that excesses need pre-approval at a level above the desk, and an automatic escalation after a second breach in a rolling period.
Where candidates lose it
Treating it as three separate breaches to be logged. It's one control failure, and the interviewer is testing whether you'll escalate past the person who authorised it. The other failure is going straight to a disciplinary framing without checking whether the limit is simply miscalibrated for a legitimate business.
Expect next
- What if the approver is the head of the desk and outranks your boss?
- What other red flags would you look for?
- How would you redesign the limit framework so this can't happen?
048How would you validate a model?Model validationGlobal capability centres
Say this
Three pillars: conceptual soundness, outcomes analysis and ongoing monitoring. So does the theory make sense for this use, does it perform against reality, and will you know when it stops working. And it has to be done by someone independent of whoever built it.
Then walk it
- Conceptual soundness first, and it's the part that gets skipped. Read the documentation, check the theory is appropriate for the intended use, check the assumptions are stated and reasonable, and review the data: source, quality, representativeness, and whether the development sample looks like today's population.
- Then replicate. Independently rebuild at least the core of it from the documentation. If you can't reproduce the results from the document, the documentation fails, and that's a finding in itself.
- Outcomes analysis: backtesting against realised outcomes, benchmarking against an alternative model or a simpler challenger, and sensitivity analysis to see which inputs the output actually depends on. Stress the inputs to the edge of plausibility and see if it breaks gracefully or catastrophically.
- Then the boundary work: what is this model not valid for. A validated model with no stated limitations is a hazard, because the next user will apply it to something new and assume it's approved.
- Ongoing monitoring: performance thresholds, population stability, and a revalidation cycle tiered by materiality. Tier 1 models annually, lower tiers less often, and any material change triggers a revalidation regardless of the cycle.
- Governance: findings rated by severity, owners and deadlines, and a model approval that can be conditional or refused. A validation function that has never refused an approval isn't independent, and that's the question I'd ask about any validation team I joined.
- Sizing it honestly: full validation of a Tier 1 pricing model is weeks of work for two people. Proportionality is the whole design problem, because validating everything to the same depth means validating nothing well.
Where candidates lose it
Going straight to backtesting. Backtesting is one third of it and it's the third that needs data you often don't have. Conceptual soundness and the explicit statement of limitations are what prevent the misuse that causes most model losses. And say the word independent, because organisational independence is the first thing a supervisor checks.
Expect next
- What would you do if you couldn't backtest because there was no data?
- How do you validate a vendor model you can't see inside?
- How would you tier models for validation intensity?
051You are validating a gradient boosting credit model that beats the existing logistic scorecard by eight Gini points. Do you approve it?Model validationBank credit risk
Say this
Not on the Gini alone. Eight points of discrimination is worth having, but I'd need calibration, stability, explainability and fair-lending testing before approving it, and I'd want to know whether the gain survives out of time rather than just out of sample.
Then walk it
- First question: is the eight points real? Check for leakage, which is the most common cause of a suspiciously strong challenger. Any feature that encodes the outcome, a post-application field, a collections flag, a date artefact, and the gain evaporates.
- Second: out of time, not just out of sample. Boosted models overfit to the period as well as to the sample. If the gain is eight points on a random split and two points on a later year, the story changes completely.
- Third: calibration. Gradient boosting ranks well and is often badly calibrated in the extremes, which is where pricing and provisioning live. Check the predicted-versus-observed curve by decile and consider isotonic or Platt scaling.
- Fourth: monotonicity and explainability. A credit model has to survive being explained to a customer who was declined and to a regulator. Unconstrained boosting can learn that higher income increases risk in some segment, which is a spurious interaction you can't defend. Monotonic constraints usually cost very little Gini and buy a lot of defensibility.
- Fifth: fairness. Test outcomes across protected characteristics and proxies for them. Complex models find proxies more efficiently than simple ones, so this risk genuinely rises with model power.
- Sixth: operational reality. Feature pipeline stability, retraining cadence, latency, reproducibility, version control, and whether anyone can support it in three years when the builder has left. Model risk includes the risk that nobody understands the production model.
- So my recommendation would be conditional approval with constraints: monotonic constraints on the key variables, calibration layer on top, capped score-level overrides, tightened monitoring thresholds, and the logistic model retained as a live benchmark. That is a real validation outcome rather than a yes or a no.
- And the commercial framing to say out loud: eight Gini points on a large retail book is worth real money, so the answer isn't to refuse complexity. It's to price the governance cost and decide deliberately.
Where candidates lose it
Picking a side. Reflexively rejecting machine learning makes you look like an obstacle; approving it on Gini alone makes you look like you've never validated anything. The answer is conditional approval with named conditions, and leakage plus out-of-time degradation are the two checks that must come first.
Expect next
- How would you test for leakage?
- What would you tell a declined customer?
- How much Gini would you give up for monotonicity?
059A bank reports CET1 of 11 percent against a 9 percent requirement. Is it safe?Regulatory reportingBank credit risk
Say this
Not from that number alone. A capital ratio tells you about solvency under the RWA model, and banks fail from liquidity and from concentration, not from a ratio. I'd want to know the composition of the denominator, the funding profile and the trajectory before answering.
Then walk it
- First, what's the 9 percent made of? Pillar 1 minimum, plus the conservation buffer, plus Pillar 2, plus any systemic surcharge. If the 9 percent is mostly buffer, breaching it restricts dividends rather than triggering resolution, which is a different kind of 2 percent of headroom.
- Second, interrogate the denominator. RWA density against total assets, how much is IRB-modelled, and single-name and sector concentration. An 11 percent ratio on a book with 25 percent in one sector is far weaker than the same ratio on a granular one, and the IRB formula won't show it.
- Third, the trajectory, which is what actually matters. Was it 13 percent two years ago? Is it falling through loan growth, buybacks or rising provisions? The direction and the stress path matter more than the level.
- Fourth, and this is the real answer, liquidity. SVB had a capital ratio comfortably above requirement the week it failed. Look at LCR, NSFR, deposit concentration, the uninsured deposit share, and unrealised losses in held-to-maturity securities that don't touch CET1 until they're sold.
- Fifth, asset quality and provision adequacy. Coverage ratio, NPL ratio, Stage 2 share, and whether provisioning looks light relative to peers. A thin provision stock means the capital ratio is borrowing from the future.
- Sixth, the stress result. What does CET1 do in the adverse ICAAP or supervisory scenario? If it drops to 8 percent, the 2 percent buffer is already spoken for and the bank is effectively at its constraint.
- So my answer would be: 11 against 9 is adequate headroom on a granular, well-funded, well-provisioned book with a stable trajectory, and thin on a concentrated book with a volatile funding base. And I'd say what I'd need to see rather than guess, because the interviewer is testing whether I'll commit to a number without the information.
Where candidates lose it
Answering yes or no. There isn't enough information, and the interviewer is testing whether you know that solvency ratios don't capture liquidity or concentration. SVB is the example that proves it, and naming unrealised held-to-maturity losses is the detail that lands.
Expect next
- What would you want to see to be comfortable?
- SVB had a fine capital ratio. Why did it fail?
- Does breaching the buffer requirement mean the bank fails?
063How would you build a survival horizon for a bank treasury?Treasury and ALMIndian bank risk and treasury
Say this
Build a daily cash flow ladder under stress and find the first day the counterbalancing capacity runs out. That day count is the survival horizon, and the useful output isn't the number itself but which assumption drives it.
Then walk it
- Start with contractual cash flows by day: maturing loans in, maturing deposits and wholesale funding out, coupon and interest flows, and known commitments. That's the easy part and the least informative.
- Then layer behavioural assumptions, which is where the answer is actually made. Deposit run-off rates by segment, drawdown on committed undrawn facilities, no rollover of wholesale funding, and rating-trigger-driven collateral calls.
- Then the counterbalancing capacity: unencumbered HQLA at stressed haircuts, central bank facilities and what collateral qualifies, committed lines you can genuinely draw, and asset sales with a realistic time-to-cash. Repo is the fast one, whole-loan sales are not.
- Then run at least three severities: an idiosyncratic name-specific stress where markets function but nobody will lend to you, a market-wide stress where everyone is short cash, and a combined scenario. The combined one is what regulators require and it's the one that binds.
- Read off the first day of negative cumulative net cash. A typical internal appetite is 30 days for a combined stress and 90 days for an idiosyncratic one, plus the regulatory 30-day LCR as a floor.
- Then the genuinely valuable step: sensitivity. If the horizon goes from 45 days to 12 when retail run-off moves from 10 to 20 percent, the number is an assumption, not a fact. I'd present the horizon as a range with the binding driver named.
- Two things people forget. Intraday and currency granularity: being liquid in rupees and short dollars on day three is a failure even if the aggregate is fine. And encumbrance, because assets already pledged in repo are not available however liquid they look.
- And the escalation link: each horizon threshold should map to a contingency funding plan trigger with named actions, or the measurement is an academic exercise.
Where candidates lose it
Building a contractual maturity ladder and calling it done. Contractual flows tell you almost nothing, because the risk lives in behaviour. And failing to split by currency is the classic error: an aggregate survival horizon can hide a dollar funding gap that kills you first.
Expect next
- What run-off rate would you assume on uninsured corporate deposits?
- How do you treat central bank facilities in the counterbalancing capacity?
- How would this differ for a non-bank finance company?
064How would you model a bank's savings deposits, which have no contractual maturity?Treasury and ALMIndian bank risk and treasury
Say this
Split the balance into a stable core and a volatile portion, assign a behavioural maturity to the core, and estimate a deposit beta for how much of a policy rate move you pass through. Those two parameters, core share and beta, drive the entire banking book rate risk answer.
Then walk it
- Volume modelling first. Look at the historical balance series per segment, strip out trend and seasonality, and take a low percentile of the remaining distribution as the stable core. Many banks use something like the balance exceeded 95 percent of the time over five years.
- Assign a repricing or behavioural maturity to the core. If it has been sticky for a decade, you might treat it as five to seven years of effective duration, capped by supervisory limits. Basel caps the average repricing maturity of core retail non-maturity deposits at five years in its standardised framework, and that cap exists because banks were assuming longer and flattering their rate risk.
- Deposit beta: regress the rate you actually paid against the policy rate. Indian savings rates have historically been very sticky, so betas on savings accounts are low, maybe 0.2 to 0.4, while term deposits and bulk deposits run much higher, 0.6 to 0.9. Segment or the average is meaningless.
- Betas are asymmetric and non-linear, and that's the point most candidates miss. They're low on the way up until competition bites and then they jump, and they're low on the way down because you can't pay less than zero. Modelling a single symmetric beta understates the squeeze in a rising cycle.
- Segmentation matters more than technique: retail versus corporate, insured versus uninsured, digitally active versus branch-only, relationship versus rate-shopping. A digitally active uninsured corporate depositor behaves nothing like a pensioner with a branch passbook.
- Then validate the assumptions against a real episode rather than trusting the regression. What actually happened to your balances and your pass-through in the 2022 to 2023 hiking cycle? That's the out-of-time test.
- The limitation to volunteer, and it's the important one: these models are all calibrated on a world where moving money was slow. Mobile banking and instant transfers have shortened behavioural maturities in a way the history doesn't contain. SVB lost a quarter of its deposits in a day, which no core-stability model would have produced. So I'd stress the assumption hard rather than trust it.
- RBI's 2024 draft LCR revisions add a run-off add-on for internet and mobile banking enabled deposits for exactly this reason, which is a good example of supervisors updating faster than the models.
Where candidates lose it
Treating it as purely a statistical problem. The regression is the easy part; the judgement is in segmentation, in the asymmetry of beta, and in recognising that the historical data predates instant digital withdrawal. A candidate who says 'the history no longer applies' and then says what they'd do about it stands out.
Expect next
- What deposit beta would you assume for Indian savings accounts?
- How would you stress the core assumption?
- Why does Basel cap the assumed maturity at five years?
065A bank funds long-dated fixed-rate securities with uninsured corporate deposits. Walk me through everything that can go wrong.Treasury and ALMBank market risk
Say this
That's the Silicon Valley Bank structure, and it fails in a specific sequence: rates rise, the asset side loses economic value without showing it in the accounts, depositors leave because they have better options, and selling the assets to pay them crystallises the loss and destroys the capital.
Then walk it
- Step one, the duration mismatch. Long fixed-rate assets and overnight liabilities means economic value of equity falls hard when rates rise, even while net interest income looks fine for a while because deposit rates lag.
- Step two, the accounting shield that becomes a trap. Securities classified as held-to-maturity aren't marked through capital, so the loss is invisible in the reported ratios. SVB had roughly $15bn of unrealised HTM losses against about $16bn of equity at the end of 2022. The capital ratio said nothing was wrong.
- Step three, the depositor incentive. Uninsured corporate treasurers are rate-sensitive and professional. When T-bills yield 5 percent and your account pays 0.5, they leave for economic reasons before there's any fear. That's a slow outflow that forces asset sales.
- Step four, the crystallisation. Selling HTM securities to fund outflows moves the loss from a footnote into the income statement and the capital ratio. Worse, selling any of the portfolio can force reclassification of the whole HTM book under the accounting rules, which is why banks resist it until they can't.
- Step five, the run. Once the loss is public, uninsured depositors with a 100 percent loss-given-failure have every incentive to leave first, and they can now do it in an afternoon from a phone with a group chat coordinating them. SVB lost about $42bn in a single day.
- Step six, concentration as the accelerant. A depositor base drawn from one industry with shared advisers and shared venture investors is not a diversified funding book. It's one depositor with many accounts.
- What the risk function should have done: report EVE alongside NII and escalate the gap, treat unrealised HTM losses as economic capital regardless of accounting, model uninsured deposits with far faster run-off, set a concentration limit on depositor type, and hedge the duration with swaps. The last one is the cheapest and SVB had almost none on.
- And the governance point: SVB had no chief risk officer for part of 2022 and its interest rate stress scenarios had reportedly been changed to be less severe. The measurement failure was downstream of a governance failure, which is almost always the case.
Where candidates lose it
Describing it as a liquidity problem only, or a rate problem only. It's the interaction, plus an accounting classification that hid the loss, plus a concentrated and professional depositor base. Candidates who name the HTM accounting treatment and the depositor concentration show they've read the post-mortem rather than the headline.
Expect next
- Why didn't the capital ratio show the problem?
- What single hedge would have changed the outcome?
- How would you set a depositor concentration limit?
071Here is a book of credit exposures and a default history. Analyse it using dynamic probability metrics.Jane StreetCredit Risk · London · 2025
Say this
Dynamic means the probabilities have to be conditional and updating, not a static average. So I'd build a hazard-rate view: probability of default in the next period given survival so far, conditioned on observable state, and then update it as new information arrives.
Then walk it
- Start with the data audit before any modelling. Definition of default, observation window, censoring, survivorship, and whether exposures enter and leave the sample. Most of the wrong answers in credit analytics come from the panel being constructed badly, not from the maths.
- Then the right framing: survival analysis rather than a single-period classification. Estimate a hazard function, the instantaneous default rate conditional on having survived, using a Cox proportional hazards or discrete-time hazard model. That handles censoring properly and gives you a term structure of default rather than one number.
- Make it conditional on state. Time-varying covariates: rating migration, spread level, utilisation, macro variables. A default probability that changes when the world changes is what 'dynamic' means here.
- Then update sequentially. Bayesian updating, or a Kalman-filtered latent credit factor, so each new month of data revises the estimate rather than triggering a full refit. For sparse default data, a Bayesian approach with an informative prior from external data is far more stable than maximum likelihood on twelve observations.
- Dependence, which is what makes a credit book different from a set of single names. Estimate a common factor and its loading, because portfolio loss is driven by correlation, not by average PD. Report the loss distribution, the 99th percentile and expected shortfall, not just expected loss.
- Validation appropriate to sparse data: time-series calibration tests, a binomial or Vasicek test given the low default counts, discriminatory power via a time-dependent AUC, and a comparison against market-implied hazard rates from CDS where they exist.
- Then say what you'd report. Not a single PD. A term structure of conditional default probabilities, the portfolio loss distribution with tail measures, the top contributors to tail loss, and an explicit statement of how much of the answer is driven by the correlation assumption.
- And the honest limitation: with a short history and few defaults, the tail of the loss distribution is an assumption rather than an estimate. I'd say that up front rather than present a confident 99.9th percentile.
Where candidates lose it
Reaching for a classifier and reporting AUC. The word 'dynamic' is doing the work in the question: it's asking for conditional, time-varying, updating probabilities, which means hazard models and Bayesian updating, not a static logistic fit. And in a credit book the correlation assumption drives the tail more than the PD does.
Expect next
- How would you handle the sparsity of defaults?
- How would you estimate the common factor loading?
- How would you present the uncertainty in the tail?
Reported by candidates at Jane Street (Credit Risk, London, 2025). Source: Wall Street Oasis.
082A new trading desk is being set up. How would you build its limit framework?Bank market risk
Say this
Start from the mandate, not from the metrics. Understand what the desk is supposed to do and how it makes money, then set limits that allow that strategy and block everything else. Then layer loss limits on top, because risk limits alone don't stop a desk bleeding.
Then walk it
- First the mandate: which products, which currencies, which maturities, which counterparty types, and where the edge comes from. A market-making desk and a relative-value desk with identical VaR need completely different limits.
- Then the aggregate risk limit, usually VaR or expected shortfall, sized off the capital allocated and the expected return. A desk allocated 100 crore of economic capital with a target return might get 5 crore of VaR, and you sanity-check that the implied risk-return is credible.
- Then granular sensitivity limits, because aggregate VaR is nettable and hides structure. DV01 by curve bucket, gamma, vega by expiry and by strike bucket, credit spread sensitivity by rating, single-name concentration, and FX delta by pair. These are what a trader actually manages to.
- Then loss limits, which are separate and essential. A daily stop, a month-to-date and a year-to-date drawdown trigger, each with a defined action: reduce, review, or stop trading. Risk limits control exposure; loss limits control the bleed when the strategy is simply wrong.
- Then stress limits, because VaR won't catch the scenario that matters. A cap on loss under the prescribed stress scenarios, which is often the binding constraint for a desk selling tail options.
- Then the boundaries that aren't about size: a product whitelist, tenor caps, a concentration cap as a share of market open interest or average daily volume, and a liquidity limit on days-to-exit. New products need explicit approval, which is how you stop mandate creep.
- Calibration approach for a brand new desk with no history: use a comparable desk's profile, size conservatively, and review after three and six months against actual usage. A limit used at 20 percent is wasted capital; a limit at 95 percent every day is being managed to rather than respected.
- And the governance wrapper: who can approve an excess and at what level, pre-trade blocks where possible rather than post-trade reports, and an automatic escalation after a second breach in a rolling period. Say that, because it's where limit frameworks actually fail.
Where candidates lose it
Setting a VaR limit and calling it done. VaR nets, so a desk can sit inside it with enormous concentrated positions. The complete answer has aggregate, sensitivity, loss, stress and liquidity limits, plus a product whitelist. And the loss limit is the one candidates most often forget.
Expect next
- Why do you need loss limits if you already have VaR limits?
- How would you calibrate the limits with no trading history?
- The desk is at 95 percent of its limit every day. Good or bad?
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

