Case 021Model risk and validationHard
You are asked to validate a bank's retail PD scorecard. Its Gini has fallen, its population stability index is high and observed defaults run well above predicted. Interpret each result and decide whether to recalibrate, redevelop or keep it with an overlay.
1The situation
Varenth Bank's retail probability of default scorecard for personal loans was built four years ago. You are running its annual validation on the latest 50,000 accounts. The Gini coefficient, which measures how well the score separates defaulters from non-defaulters, was 62% at development and is 48% now.
The population stability index comparing today's applicants with the development sample is 0.28. The model predicted an average default rate of 2.5% for the year; 4.1% was observed. The bank's model policy uses the common rules of thumb that a PSI above 0.25 signals a major population shift; confirm the bank's own thresholds before applying them.
2Your task
What does each of the three results tell you, and what do you recommend the bank does with the model?
Quick check
Gini has fallen from 62% to 48%. Does the model still rank borrowers?
Worked solution
Try it on paper, then open one step at a time.
30-second answerThe answer to give first
Varenth's scorecard still ranks borrowers, but the population has moved and the level of risk is badly understated. Gini of 48% is weaker but usable; a PSI of 0.28 means today's applicants differ materially from the development sample; and observed defaults are 1.64 times predicted. Recalibrate now, as an overlay, so provisions and capital reflect 4.1%, and schedule a redevelopment, because the shift in population will keep eroding the ranking.
Step 1What are the three questions a validation asks?
Think of a bathroom scale. It can put three people in the right order of weight and still read five kilos light for all of them; or it can read correctly for adults and badly for children, because it was never built for them. A PD model is validated on three separate questions: does it rank borrowers, is it being used on the population it was built for, and is its level of risk right. These are discrimination, stability and calibration, and each can fail while the others pass.
Step 2What does each result mean?
Discrimination: a Gini of 48% still ranks borrowers usefully, but a 23% relative fall in four years says the model is ageing. Stability: the population stability indexA measure of how far the score distribution of current applicants has moved from the development sample; larger values mean a bigger shift. of 0.28 comes mostly from applicants crowding into the lowest score band, 37% of today's book against 20% at development. Calibration is the sharpest failure: on 50,000 accounts the model expected 1,250 defaults and 2,050 arrived, about 23 standard deviations away, far beyond chance.
| Score band | Share at development | Share today | Predicted PD | Observed PD |
|---|---|---|---|---|
| Band 1 | 20% | 37% | 4.44% | 7.49% |
| Band 2 | 20% | 25% | 2.22% | 3.53% |
| Band 3 | 20% | 19% | 1.11% | 1.66% |
| Band 4 | 20% | 11% | 0.60% | 0.86% |
| Band 5 | 20% | 8% | 0.30% | 0.42% |
| Portfolio | 100% | 100% | 2.50% | 4.10% |
Step 3Recalibrate, redevelop or overlay?
Match the fix to the failure. Recalibration fixes the level without touching the ranking, which is exactly what the calibration failure needs, and it can be done in weeks. Scaling predicted PDs so the portfolio averages 4.1%, about 1.64 times, brings provisions and capital into line now. But recalibration cannot fix a population shift. Applicants the model was not built on are a growing share of the book, and the falling Gini is what that looks like as it happens.
So the recommendation has three parts: an immediate recalibration, documented as a model overlay with an owner and an expiry date; a redevelopment on recent data that includes the new applicant segments; and, until then, a review of the approval cut-off, since a cut-off set on understated PDs is approving loans the bank would otherwise decline. The limitation: one year of observed defaults may partly reflect a weak economy rather than the model, so check whether defaults rose across other models too before blaming this one alone.
Where candidates lose it
Candidates see three bad numbers and say redevelop, stopping there. That leaves the bank running an understated model for the six to twelve months a redevelopment takes; the interviewer wants the immediate fix as well.
The opposite miss is recalibrating and declaring the model fixed. Calibration and stability are different failures, and scaling the PDs does nothing about a population the model was never built for.
What the interviewer asks next
- How would you tell whether the calibration gap is the model or the economy?
- What would you check in the data before trusting a PSI of 0.28?
- Who should own the overlay, and what makes an overlay acceptable to a regulator?
Asked at UBS, Risk Management, Zurich, 2021 (Wall Street Oasis): How would you validate a model?
Company names and figures are illustrative.
