Case 045Model risk and validationHard
A bank's fraud model flags 2% of transactions and catches 70% of fraud; a challenger flags half as many and catches 60%. Given the cost of a missed fraud and of reviewing an alert, which model should the bank run?
1The situation
Ivrosa Bank screens 10 lakh card and payment transactions a month, of which about 500 are fraudulent. Its current vendor fraud model flags 2% of transactions for review and catches 70% of the fraud. An in-house challenger model flags only 1% and catches 60%.
Each missed fraud costs the bank Rs 50,000 on average in reimbursements and write-offs. Each alert costs Rs 150 to review, in analyst time and customer contact. The operations head favours the challenger because it halves the alert queue.
2Your task
What does each model cost a month in total, which should the bank run, and at what loss per fraud would the answer flip?
Quick check
Which model is cheaper in total?
Worked solution
Try it on paper, then open one step at a time.
30-second answerThe answer to give first
Keep the current model: it costs about Rs 1.05 crore a month against Rs 1.15 crore for the challenger. The challenger saves Rs 15 lakh of review work but misses 50 more frauds, costing Rs 25 lakh. Its alerts are more precise, 3.0% against 1.75%, yet it is the more expensive model. The answer flips only if a missed fraud costs under Rs 30,000, which is the number to check.
Step 1What does each model cost in total?
A smoke alarm that goes off for burnt toast is annoying, and one that stays silent in a real fire is ruinous; the right setting depends on what each mistake costs. The current model raises 20,000 alerts, Rs 30 lakh to review, and misses 150 frauds, Rs 75 lakh: Rs 1.05 crore a month. The challenger raises 10,000 alerts, Rs 15 lakh, and misses 200 frauds, Rs 100 lakh: Rs 1.15 crore.
| Per month | Current model | Challenger |
|---|---|---|
| Alerts raised | 20,000 | 10,000 |
| Frauds caught of 500 | 350 | 300 |
| Precision: frauds per alert | 1.75% | 3.00% |
| Review cost at Rs 150, Rs lakh | 30 | 15 |
| Missed fraud at Rs 50,000, Rs lakh | 75 | 100 |
| Total, Rs lakh | 105 | 115 |
Step 2Why is the more precise model the worse choice?
Precision and catch rate each describe one kind of mistake. The decision needs both, weighted by what each mistake costs. Look at the margin: the current model's extra 10,000 alerts catch 50 extra frauds, one fraud per 200 alerts. Reviewing 200 alerts costs Rs 30,000, and each fraud caught saves Rs 50,000. Spending Rs 30,000 to save Rs 50,000 is a good trade, so the extra alerts pay for themselves.
Step 3What would you check before closing the question?
The inputs. The Rs 50,000 average hides a skew: if a few large frauds make up most of the loss, compare which model catches those, not just how many each catches. Check whether 20,000 alerts a month is within the review team's capacity, because alerts that are not reviewed catch nothing. And this is a champion-challengerRunning a proposed new model alongside the current one on the same live data, so the two can be compared on real outcomes before one replaces the other. comparison on one month: run both in parallel for longer, and try the challenger at a lower threshold, since a model flagging 2% might beat both. Finally, a vendor model the bank cannot explain carries its own model risk; the validation file should record that too.
Where candidates lose it
The common error is to pick on one metric: fewer alerts, which operations likes, or higher precision, which looks cleaner. Neither tells you the cost; the missed frauds do.
The second is to compare only catch rates and ignore review cost. At a higher alert cost, say Rs 300, the crossover moves to Rs 60,000 and the answer flips, so the costs are part of the answer, not a detail.
What the interviewer asks next
- Review cost rises to Rs 300 an alert. Which model wins now?
- How would you pick the alert threshold for the current model to minimise total cost?
- The challenger catches more of the largest frauds. How would you redo the comparison?
Company names and figures are illustrative.
