Risk Management interview preparation
Market, credit and operational risk, plus model validation, regulatory capital, liquidity and ALM, the statistical foundations and the Indian regulatory syllabus. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it — and answers lead with the point, then the mechanism, then the limitation.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 37
- Firms
- 12
- Updated
- September 2026
005Calculate the one-day 99% VaR of a $100 million portfolio with 2 percent daily volatility.UBSRisk Management · Remote · 2020UBSRisk Management · Zurich · 2021
Say this
Parametric, assuming zero mean and normality: 2.33 times 2 percent times $100m, so about $4.66 million. Round it to $4.7m and say the assumptions out loud as you go.
Then walk it
- The z-score at 99% one-tailed is 2.326. I'd use 2.33. At 95% it's 1.645, worth knowing cold because interviewers switch between them mid-question.
- Daily volatility in money terms is 2 percent of $100m, so $2m. Multiply by 2.33 and you get $4.66m.
- I'm assuming a zero expected return over one day, which is standard for a short horizon, and that returns are normal, which is the assumption doing all the work.
- To scale to ten days, multiply by the square root of ten, about 3.16, so roughly $14.7m. That scaling needs independent and identically distributed returns, so it understates the number if volatility clusters or the market trends.
- And the caveat I'd volunteer: because the real distribution has fatter tails than normal, this figure is probably too low at 99%. Empirically the 99th percentile of equity returns sits closer to 2.6 or 2.7 sigma than 2.33.
Where candidates lose it
Using 1.645 for 99% or 2.33 for 95%, or reaching for a calculator. Know both z-scores by heart, state your two assumptions before you multiply, and flag that fat tails make the answer conservative in the wrong direction.
Expect next
- Now give me the ten-day number.
- What if the portfolio had a 10 percent annual expected return, does that change it?
- How would the answer change if the returns were t-distributed with five degrees of freedom?
Reported by candidates at UBS (Risk Management, Remote, 2020); UBS (Risk Management, Zurich, 2021). Source: Wall Street Oasis.
012Your 99 percent one-day VaR model produced nine exceptions in the last 250 days. Walk me through what you do.Bank market riskModel validation
Say this
Nine is amber, one short of red, so two things happen in parallel: the capital multiplier steps up and I open a model investigation. But before either, I check that the exceptions are real and not a data or P&L-attribution problem.
Then walk it
- Step one, validate the exceptions. Bad marks, a stale curve, a missing trade feed, or backtesting against actual instead of hypothetical P&L can all manufacture breaches. I've seen a whole amber month turn out to be one mispriced illiquid bond.
- Step two, look at clustering. Nine breaches spread evenly across the year says the model is calibrated too low. Nine in a three-week window in March says the model is fine in normal times and slow to react to a volatility regime shift. Completely different fixes.
- Step three, attribute. Which desk, which risk factor, which side. If eight of the nine come from one credit desk, it's not a firmwide VaR problem, it's a missing risk factor or a proxy that stopped working.
- Step four, size them. Breaches at 1.1 times VaR are a calibration issue. Breaches at three times VaR mean the tail shape is wrong, which points at normality or at unmodelled optionality.
- Step five, the regulatory and capital consequence. Under the Basel backtesting framework nine exceptions sits in the amber zone with a multiplier around 3.65 rather than 3.0, and it's a disclosable model performance issue. I'd tell the CRO and the supervisor rather than wait to be asked.
- Step six, the fix, and it should be the smallest defensible one: reweighting the window or moving to volatility-scaled historical simulation for clustering, adding a missing factor for a desk problem, moving to full revaluation for an optionality problem. Then re-run the backtest on the corrected model over the same period.
- And the interim control while the fix is validated: a VaR add-on or a tightened desk limit. You don't get to run unlimited with a broken model while the remediation is in flight.
Where candidates lose it
Jumping straight to 'recalibrate the model'. The first move is always to check whether the exceptions are real, and the second is to look at their pattern. A candidate who recalibrates without diagnosing has just fitted the model to a data error, and that is the exact failure the interviewer is probing for.
Expect next
- What if all nine were in the same fortnight?
- What's the capital consequence of amber versus red?
- Would you tell the regulator before or after you had a fix?
016Design a stress scenario for a book that is long Indian corporate bonds and short interest rate futures.Indian bank risk and treasuryBank market risk
Say this
The scenario has to break the hedge, not just move the market. The position is long credit and short duration, so the pain case is spreads widening while the risk-free curve rallies, which is precisely what a flight to quality does.
Then walk it
- Start by naming the real exposures. Net duration is small by design, so a parallel shift is not the risk. The live risks are credit spread, the government-bond-to-swap basis, the futures-to-cash basis, and liquidity in the corporate leg.
- So the core shock: AAA and AA corporate spreads widen 150 to 250 basis points, while the ten-year G-sec yield falls 75 basis points. You lose on both legs at once. That is the textbook flight-to-quality asymmetry and it happened in March 2020.
- Layer in the basis. The bond futures may not track the cash bond you hold, and the cheapest-to-deliver can switch. Add 25 to 50 basis points of adverse basis independent of the spread move.
- Layer in liquidity. Indian corporate bond secondary volumes are thin outside the top names, so add a bid-offer widening of two to four times normal and assume you can only exit 20 percent of the position in a week. Then mark the rest at the stressed exit price, not the matrix price.
- Layer in funding. Repo haircuts on corporate paper rise, margin on the futures short goes up as volatility spikes, and both hit the same day. That is the mechanism that turns a mark-to-market loss into a forced sale.
- Add a name-specific tail: one issuer in the book is downgraded below investment grade, which triggers forced selling by mandate-constrained funds and moves the whole rating bucket. The IL&FS episode in 2018 is the live Indian precedent, and the credit-fund redemption spiral that followed is the second-round effect.
- Then report it properly: P&L by leg, the funding call in rupees, days to unwind, and which limits break. A scenario that produces one aggregate number is not decision-useful.
Where candidates lose it
Designing a parallel rate shock. The book is deliberately hedged against that, so the scenario shows nothing and you have proved you didn't look at the position. A good stress scenario attacks the assumption the hedge relies on, which here is spread-to-rate correlation, and it must include liquidity and funding, not just price.
Expect next
- How would you calibrate the size of the spread move?
- What second-round effects would you add?
- How would you present this to a treasurer who says the book is hedged?
018How would you run an ICAAP, and how does it relate to the supervisory stress tests?Regulatory reportingBank credit risk
Say this
ICAAP is the bank's own answer to 'how much capital do you actually need', as opposed to the minimum the rules prescribe. You identify all material risks, quantify them including the ones Pillar 1 ignores, stress the plan, and conclude with a capital number and a plan to hold it.
Then walk it
- Start with a risk identification and materiality assessment across everything, not just credit, market and operational. Concentration, interest rate risk in the banking book, pension, reputational, strategic and model risk are the Pillar 2 gaps, and IRRBB and concentration are usually the two biggest.
- Quantify each, then stress the three-year business plan under a baseline and at least one severe but plausible adverse scenario. The adverse case has to be internally consistent: if GDP falls, credit costs rise, fee income falls and RWAs inflate through downgrades, all at once.
- Project the capital path, not just the end point. The trough quarter is what matters, and it usually sits in year two because provisions lag the macro.
- Set the internal capital requirement above the regulatory minimum, with a management buffer sized so that you don't breach the buffer requirement in the adverse case and get dividend restrictions.
- Then the management actions, with triggers. Which of those are credible under stress is the question a supervisor will push on hardest, because cutting dividends works and issuing equity in a crisis usually doesn't.
- The relationship with supervisory tests: the regulator's exercise, CCAR in the US, the EBA's in Europe, and the RBI's stress-testing guidance in India, uses common prescribed scenarios so banks can be compared. ICAAP is idiosyncratic and covers risks the common scenario ignores. Under the SREP the supervisor uses your ICAAP to set a Pillar 2 requirement on top of Pillar 1.
- Governance is half the assessment. An ICAAP the board has clearly never read fails regardless of the modelling quality. The board's sign-off on the risk appetite and the capital plan is the artefact supervisors look for first.
Where candidates lose it
Describing ICAAP as a document rather than a process, and forgetting Pillar 2 risks. If you can't name interest rate risk in the banking book and concentration as the two big risks outside Pillar 1, you have not understood why ICAAP exists at all.
Expect next
- Which Pillar 2 risk is usually the largest?
- How would you size a management buffer?
- What management actions would a supervisor refuse to credit?
021A trader tells you his book is delta neutral. It lost $4 million yesterday on a 3 percent market move. What happened?Bank market riskDerivatives risk
Say this
Almost certainly short gamma. Delta neutral only holds for an infinitesimal move; if he's short options, delta turns against him as the market runs, so he's rehedging at worse and worse prices all the way. The move being 3 percent is the clue.
Then walk it
- The mechanism: short gamma means delta moves against you. Market rallies, your delta goes short, you buy to rehedge, market falls back, your delta goes long, you sell. You buy high and sell low mechanically all day.
- Rough size check: for a book with gamma of minus $2m per percent, a 3 percent move costs about half times gamma times move squared, so around $9m of gamma P&L. A $4m loss is entirely consistent with a modest short gamma position.
- Second candidate, vega. A 3 percent move usually comes with implied vol up several points. If he's short vol, that's a separate loss on top, and on a big book vega loss can dwarf gamma loss.
- Third, delta neutral in what. Neutral to the index but long a basket of single names is a beta hedge, not a delta hedge. Dispersion or a basis between the hedge instrument and the underlying gives you exactly this.
- Fourth, the hedge was neutral at the close and not during the day. Intraday delta drift with no rehedging looks flat on both snapshots and loses money in between.
- So the questions I'd ask him, in order: what's your gamma and vega, what did implied vol do, what instrument are you hedged in, and when was the last rehedge. And the control conclusion: a delta limit alone was never going to catch this, which is why you need gamma and vega limits.
Where candidates lose it
Saying 'he must have been wrong about being delta neutral'. He probably wasn't. The whole point is that delta neutrality is a local property and says nothing about second-order risk. Name gamma first, vega second, and then draw the control conclusion about limits.
Expect next
- How would you size a gamma limit?
- How do you explain to a trader that delta neutral isn't neutral?
- What if implied vol had fallen instead?
027How would you approach building a delinquency model?Neuberger BermanRisk · Chicago · 2024
Say this
Define the target first, then build backwards. Delinquency is not default, so I'd fix the bad definition, say 90 days past due within twelve months, set an observation and performance window, and only then worry about features and model form.
Then walk it
- Target definition is the decision that determines everything else. 30, 60 or 90 days past due, and over what horizon. Roll-rate analysis tells you where delinquency becomes effectively irreversible, and that's where you draw the line.
- Sampling: pick an observation point, take the borrower's state as at that date, then observe outcomes over the following twelve months. Strict separation, or you leak future information into features and get a model that looks brilliant in development and fails in production.
- Features in three families. Behavioural: utilisation trend, minimum-payment behaviour, recent missed payments, bounced mandates. Bureau: enquiry velocity, existing delinquency elsewhere, thin-file flags. Loan and demographic: loan-to-value, instalment-to-income, vintage, product, channel of origination.
- Model form: start with logistic regression on coarse-classified, weight-of-evidence binned variables. It's monotonic, explainable and passes validation. Then run a gradient boosting challenger to see how much signal the simple model leaves on the table. If the gap is small, ship the simple one.
- Validation: out-of-time as well as out-of-sample, because credit models degrade through the cycle not through the sample. Report Gini or AUC for ranking, and a calibration curve for whether the predicted rates match observed. A model can rank perfectly and be badly calibrated.
- Two traps specific to credit. Survivorship and selection bias: you only observe outcomes for people you approved, so the model is blind to the rejected population, and you need reject inference. And macro sensitivity: a model built on 2021 data has never seen a rate cycle, so the absolute PD level will be wrong even if the ranking holds.
- Then monitoring. Population stability index on the score distribution, drift on each feature, and a monthly actual-versus-expected. Most delinquency models fail from population shift rather than bad maths.
Where candidates lose it
Going straight to algorithms. In credit, the target definition, the observation window and the reject-inference problem are worth more than model choice, and interviewers who build these for a living are listening for exactly those. Also say the word calibration; ranking power alone doesn't let you price or provision.
Expect next
- How would you handle reject inference?
- How would you know the model had degraded?
- Would you use gradient boosting in production for this?
Reported by candidates at Neuberger Berman (Risk, Chicago, 2024). Source: Wall Street Oasis.
028How do you build a credit scorecard, and how do you prove it works?Bank credit riskGlobal capability centres
Say this
Bin every variable, convert to weight of evidence, fit a logistic regression, then scale the log odds into points. You prove it works on three axes: discrimination, calibration and stability, tested out of time, not just out of sample.
Then walk it
- Coarse classification first. Bin each variable so the bad rate is monotonic across bins and each bin has enough volume, usually at least 5 percent of the population. Then replace the bin with its weight of evidence, the log of the good-to-bad odds ratio.
- Information value tells you which variables to keep. Below about 0.02 is useless, 0.1 to 0.3 is useful, above 0.5 and I'd check for leakage rather than celebrate.
- Fit logistic regression on the WOE variables, then scale: points equal offset plus factor times log odds, calibrated so a chosen score doubles the odds every 20 points. That scaling is cosmetic but it's how credit officers read the output.
- Discrimination: Gini, or equivalently AUC, where Gini equals two times AUC minus one. A retail behavioural scorecard should hit 0.55 to 0.70 Gini; an application scorecard on a thin-file population might only get 0.35, and that can still be commercially valuable. Kolmogorov-Smirnov is the other standard, the maximum gap between the cumulative good and bad distributions.
- Calibration: plot predicted against observed bad rate by score band, and run a Hosmer-Lemeshow style test. Discrimination decides who you approve; calibration decides what you charge and what you provision. You need both.
- Stability: population stability index between development and current, per variable and on the score. Above 0.25 and the population has shifted enough that the model needs rebuilding, not just recalibration.
- And the governance point: build on a development sample, validate on a holdout, then validate again on a later time period the model never saw. An out-of-sample test on a random split proves almost nothing for a credit model, because the whole failure mode is time.
Where candidates lose it
Quoting a Gini target as if it were universal. A good Gini depends entirely on the population and the product, and someone who has built these knows that. The other failure is testing only discrimination. A model with 0.7 Gini and broken calibration will approve the right people and price them all wrong.
Expect next
- What Gini would you expect on a prime mortgage book?
- The Gini is stable but the bad rate has doubled. What happened?
- When do you recalibrate versus rebuild?
032What is structured finance, how would you evaluate it, and what are the credit risks?Moody'sCredit Risk · New York · 2024
Say this
Structured finance is taking a pool of cash-flow-generating assets, putting it in a bankruptcy-remote vehicle, and slicing the cash flows into tranches of different seniority. You evaluate it in three layers: the collateral, the structure, and the parties.
Then walk it
- Layer one, the collateral. Pool composition, weighted average life, seasoning, geographic and obligor concentration, historical default and prepayment behaviour, and how the underwriting was done. Everything downstream depends on this, and it's where the 2007 failure actually was.
- Layer two, the structure. Where does the cash go, and in what order. Credit enhancement comes from subordination, excess spread, overcollateralisation and reserve accounts. Then the triggers: performance triggers that turn a pro-rata waterfall sequential, and cash-trapping mechanics.
- Layer three, the parties. Originator, servicer, trustee, swap counterparty. Servicer quality drives recoveries, and servicer failure has broken deals whose collateral was fine. Then the legal question: is the true sale robust, and is the SPV actually bankruptcy remote?
- How I'd analyse it: model the pool, run default and prepayment scenarios, and see at what cumulative loss each tranche takes its first rupee of loss. That break-even loss compared with the expected loss is the real measure of a tranche's safety.
- The credit risks specific to tranching. Correlation risk: a senior tranche is a bet on correlation, not just on average defaults, because it only fails if losses cluster. Cliff risk: a mezzanine tranche goes from untouched to wiped out over a narrow loss range, so it's far more convex than its rating suggests.
- Then prepayment and extension risk on the timing, basis risk if the assets and liabilities reprice off different benchmarks, and originator alignment. Skin in the game is why post-crisis rules require the sponsor to retain a slice.
- The Indian version worth naming: pass-through certificates and direct assignments on NBFC loan pools, where the live risks are servicer concentration, priority-sector motivation on the buyer side, and the 2018 to 2019 NBFC liquidity episode showing how quickly refinancing assumptions fail.
- And the honest limitation: the rating of a structured tranche is far more model-dependent than a corporate rating. Small changes in a correlation assumption move a AAA to a BBB, and that is exactly what happened to CDOs.
Where candidates lose it
Explaining tranching and stopping. The two things a credit risk interviewer at a rating agency wants are the sensitivity of senior tranches to correlation rather than to average default rates, and the cliff-risk convexity of mezzanine. Naming the servicer and the true-sale question shows you've read a deal document, not a textbook.
Expect next
- Why is a senior tranche a bet on correlation?
- What actually went wrong with CDO ratings in 2007?
- How would you analyse an Indian NBFC pass-through certificate?
Reported by candidates at Moody's (Credit Risk, New York, 2024). Source: Wall Street Oasis.
033Your loan book has 22 percent of exposure in commercial real estate. How do you think about that?Bank credit riskIndian bank risk and treasury
Say this
Concentration is the risk that diversification math ignores, and 22 percent in one cyclical sector is a real number. I'd measure it, work out how much of a tail loss it drives, then decide between limits, hedging and pricing rather than just flagging it.
Then walk it
- Measure it properly first. Sector share is the headline, but the useful metrics are a Herfindahl index on single names, the top-20 exposure as a share of CET1, and the correlated cluster, because CRE developers, construction firms and building materials all fail together.
- Then quantify the capital impact. The IRB formula assumes an infinitely granular, single-factor portfolio, so it systematically understates concentrated books. A granularity adjustment or a multi-factor economic capital model is how you show the board the real number, and this is a classic Pillar 2 add-on.
- Stress it specifically. Property values down 30 percent, vacancy up, refinancing unavailable at maturity. CRE defaults are refinancing events far more than they are cash-flow events, so the maturity profile matters more than current interest coverage.
- Look at what's inside the 22 percent. Office in one city is a different animal from warehousing and retail across ten. Loan-to-value distribution, debt-service coverage, single-tenant concentration, and how much matures in the next 18 months.
- Then the actions, in order of cost. Tighten new-origination limits by sub-sector, price the concentration into new deals, syndicate or sell down the largest names, and buy protection or securitise if a market exists. In India that last option is thin, so limits do most of the work.
- And the governance line: 22 percent may be entirely within appetite if the board decided that deliberately and is paid for it. Concentration isn't automatically a fault. What is a fault is concentration that accumulated without anyone setting a limit.
Where candidates lose it
Saying 'that's too high' without a benchmark or a measurement. And forgetting that the IRB capital formula assumes a granular portfolio, so regulatory capital alone will not show the concentration. That granularity point is what a credit risk interviewer is waiting for.
Expect next
- How would you measure concentration in a single number?
- Why doesn't the IRB formula capture it?
- What limit would you set, and on what basis?
043What is an RCSA, and how would you actually run one?Operational riskGlobal capability centres
Say this
A risk and control self-assessment is the business identifying its own risks, rating them before and after controls, and owning the gap. Run badly it's a spreadsheet nobody reads. Run well it's the only forward-looking view of operational risk you have.
Then walk it
- Start from processes, not from a risk list. Map the end-to-end process, find the failure points, and derive risks from those. Starting from a generic taxonomy produces generic risks that nobody recognises as theirs.
- For each risk, rate inherent likelihood and impact, then identify the controls, test whether they actually work, and rate residual risk. The distinction between design effectiveness and operating effectiveness matters: a beautifully designed control that is performed late every month is not effective.
- Then compare residual risk to appetite. Anything above appetite needs an action with an owner and a date, or a formal risk acceptance signed at the right level. That is the actual output; the ratings are just how you get there.
- Who does it: the first line owns it, the second line facilitates and challenges. If risk management fills in the RCSA, the business hasn't assessed anything and you've built a document rather than a control.
- Where it fails, and I'd say this without prompting. Everything gets rated amber, because nobody wants to own a red. Ratings never change year to year. The workshop runs after a fine rather than before. And it never reconciles against actual loss events, so a process with twelve losses last year is rated low risk.
- So the tests I'd apply to an RCSA: does it reconcile to the loss database, does it reconcile to audit findings, has anything moved since last year, and can a process owner explain their own top risk without reading the sheet.
- Practical numbers: for a mid-sized operation, expect 15 to 40 risks per process area. Hundreds means it's a control inventory dressed up as a risk assessment, and nobody will use it.
Where candidates lose it
Describing the template instead of the process. Two things separate a real answer: saying the first line must own it with the second line challenging, and naming the amber-everywhere failure mode. And reconciling the RCSA against actual loss data is the check almost no candidate mentions.
Expect next
- Who should own the RCSA?
- How do you stop everything being rated amber?
- How does the RCSA connect to your loss data?
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.

