Chatbots and Conversational AI: What They Handle and Escalate
A conversational assistant answers questions in words. The assistant sits beside a process rather than inside it: it can tell a customer what happens next, and it does not decide what happens next. Two numbers describe it and they are not the same number. One is the share of conversations that end without a person. The other is the share where the customer's question was actually answered.
Measuring an assistant runs into one uncomfortable property at the root. The number a firm reports for an assistant is the number the assistant can produce by itself, and the only thing it can see about a conversation is whether the conversation stopped. Whether the person on the other end got what they came for is not visible to the thing being measured. A record of whether they did exists only if somebody reads what was said, line by line, after the fact. A firm that never reads its own conversations has one number and believes it has the other.
What is Conversational AI actually doing with a question?
Stand behind the enquiry counter at a large railway station on a Sunday evening. A queue of people, one question each. The clerk does three things with every one of them, in the same order, and never announces any of them. First he decides which of the questions he already knows the answer to this one resembles. Second he says something. Third he decides whether that person is finished with him or whether they need to be sent to the reservation office at the far end of the concourse.
A conversational assistantA component that answers questions in words, sitting beside a process rather than acting inside it. does exactly those three things and nothing else. The assistant takes a question, produces a match to something it holds, and decides whether to keep the conversation or hand it over. How the words themselves get assembled is a separate subject, set out elsewhere. For a firm deploying one, the middle step is a prediction rather than an understanding. The component predicts which of the things it holds is the right thing to say, and the customer then acts on that prediction. Ajay Agrawal, Joshua Gans and Avi Goldfarb describe exactly that arrangement in Prediction Machines: the prediction is cheap, and the judgement about what to do with it is not.
The first is observable and the second is a guess about a machine, so describe it by what it does with a question and never by what it understands. The moment a deployment note says the assistant understands the customer, every design question that follows gets asked in the wrong shape. Understanding does not have a poor version, so nobody asks what happens when the match is poor. Either it understands or it does not. No component behaves that way.
Where do conversational interfaces sit, beside the process or inside it?
Sumeru Bank Limited, an invented bank, runs a retail loan intake chain built from nine numbered components, and beside that chain it runs an assistant that talks to customers. The distinction in that sentence is the whole of this block. The chain decides things. The assistant talks about the things the chain decided. Chain and assistant touch, and they are not the same object.
The arrangement is familiar from a large hospital. The doctors decide, the tests measure, the pharmacy dispenses. At the entrance there is a help desk that can tell a visitor which floor cardiology is on, when visiting hours are, and roughly how long the queue for the blood test has been running. The help desk is genuinely useful and it is not part of anybody's treatment. Nobody at the hospital would describe the help desk as a clinical step, and nobody would let it change a prescription because it happened to be nearest to the patient asking.
The assistant at Sumeru answers questions about the process and takes no part in deciding anything inside it, so it is that help desk rather than one of the nine numbered components. The distinction is not a technicality about a numbering scheme. Where the assistant sits settles what governance it needs, what record has to exist afterwards, and above all what it is allowed to say. A component that decides an outcome needs an accountable person, an approval, a stated reason for what it did and a way to be challenged. A help desk needs an accurate answer and a clean way to hand the customer on.
Is the conversational assistant one of the nine numbered components of the chain?
What are the three things it can do, and which one must it not be given?
An assistant can be given three kinds of capability and they escalate sharply in what they demand of the firm. The first is stating a fact that is the same for everybody: how long a decision usually takes, what a document upload step needs, what happens after a file goes to a person. The second is stating a fact about one customer specifically: where that customer's own application currently sits. The third is the ability to change something: move a repayment date, cancel a standing instruction, restart an application.
The first two are answering. The third is acting, and it is the one that must not be handed over casually. The moment an assistant can change something rather than describe something, it has stopped being beside the process and started being inside it, and everything a component needs must now be built for it. A record of what changed and on whose instruction. An approval for the change being made this way at all. A named person accountable for what it does. A route for a customer to say the change was not what they asked for.
None of that arrives with the conversation, and nothing in the conversation announces that it is missing. The same interface produces the same kind of sentence in both cases, so the capability to act feels like a small extension of the capability to answer. The extension is not small. Acting is a different object with a different obligation attached, and the sentence is the only part that looks familiar.
| What it is asked to do | What it is | What has to exist behind it |
|---|---|---|
| State a fact true for everybody | Answering | The answer being current, and a date on it |
| State where the customer's own file sits | Answering, about one person | The above, plus knowing it is really the customer |
| Change something on the customer's account | Acting | Record, approval, named accountability, a route to challenge |
A firm wants its assistant to be able to change a customer's repayment date. What has just changed?
Why does it hand a customer to a person, and for which five reasons?
Every escalationHanding the conversation to a person, with the reason recorded. at Sumeru carries a numbered reason. Each reason is a different kind of message about the design, and the five together are the most informative output of the whole deployment. In month 6, of 24,000 conversations, 7,200 went to a person, being 30.0 per cent. Every figure here is the invented bank's own and describes one deployment in one month.
| Reason | What triggered it | Count | Share |
|---|---|---|---|
| 1 | The customer asked for a person | 2,448 | 34.0% |
| 2 | The assistant could not match the question to anything it holds | 2,160 | 30.0% |
| 3 | The question was about the decision on one named application | 1,440 | 20.0% |
| 4 | The customer asked the same question three times | 720 | 10.0% |
| 5 | The conversation carried a word on the complaint list | 432 | 6.0% |
| Total handed to a person | 7,200 | 100.0% |
The five say completely different things, so read them one at a time. Reason 1 is a preference and it is legitimate: a customer who wants a person is allowed to want a person, and a design that makes that hard is a design that has confused containment with service. Reason 2 is a coverage gap: there is a question customers keep asking that nothing in the assistant addresses, and the fix is content rather than anything to do with the component. Reason 3 is a deliberate wall, and it is the most important one on the list. Reason 5 is a complaint listA set of words whose presence sends a conversation to a person regardless of what else it says., a set of words whose presence sends the conversation to a person regardless of anything else, and it exists because a complaint mishandled by an assistant is a conduct matter rather than a service one.
Reason 4 is the only escalation reason that is also direct evidence about the quality of the answers, and it is the one worth watching on its own. A repeat questionThe same question asked again, which is the assistant signalling that it did not answer it. is the assistant signalling that it did not answer, and a customer asking the same thing a third time is not being difficult. The assistant answered twice, and the customer did not accept either answer. Neither answer landed. Somebody who returns with the same question was failed the first time. Any deployment note that describes reason 4 as an impatience threshold has quietly changed the subject from the assistant to the customer.
Reason 4 is the customer asking the same question three times. What is that reason actually reporting?
Why is reason 3 a wall rather than a gap?
Reason 3 covers 1,440 conversations a month: somebody asking about the decision on one named application. Their own. Nothing is more natural for a customer to ask, and the assistant is built never to answer it. The refusal is a design choice made in advance rather than a limitation discovered later.
Here is why. A decision on a named application is the output of the scoring model and the routing rules, and it carries an accountable person and a stated reason for what it did. If an assistant answers a question about that decision, it is producing a second account of a decision it did not make and cannot see the workings of. If the two accounts differ, the customer now holds a sentence from the bank that the bank cannot support. An assistant that answers a question about a live decision has stopped describing the process and started speaking for it. The crossing is the same one as being allowed to act.
So the wall is not caution about accuracy. The wall is a boundary about authority. The assistant may say the file is with a person, it may say what usually happens next, and it may say when the customer can expect to hear. The assistant may not say whether the answer will be yes.
An assistant is asked whether a particular application will be approved. What should it do?
What is a containment rate, and what exactly does it count?
The number every firm reports for an assistant is the containment rateThe share of conversations that ended without a person, which is what the assistant can count by itself.. At Sumeru in month 6 it was 70.0 per cent: 16,800 of 24,000 conversations closed without a person. Nobody made that number up and nobody manipulated it. The containment rate is exactly what it says.
The rate says, precisely, that 16,800 conversations stopped and no member of staff was involved when they did. A household example sits usefully beside that sentence. A customer calls a shop about a delivery. In one call they are told the parcel arrives Thursday and they hang up satisfied. In another call they go round three times, get nowhere, and hang up because dinner needs cooking. Both calls ended. Neither involved anybody being transferred. To a counter that measures how many calls ended without a transfer, those two calls are the same event.
A containment rate counts endings, and an ending caused by a customer giving up increments it exactly as much as an ending caused by an answer. None of that is a flaw in the arithmetic. The arithmetic is correct. The blindness is a property of what the assistant can observe about itself. A counter has no way to see the difference between those two calls, so no amount of care in building the assistant changes the number.
A firm reports 70.0 per cent containment. What is the one question to ask about it?
What is a resolution rate, and what is the only way to get one?
The number that matters is the resolution rateThe share of conversations where the question was actually answered, which only a person reading transcripts can produce.: the share of conversations where the customer's question was actually answered. There is exactly one way to obtain it and it is unglamorous, and it has a name: transcript readingSampling closed conversations by hand, which is the only route to a resolution rate.. Somebody takes a sample of closed conversations, reads each one, marks whether the question was answered, and scales the share to the month.
That is it. Four steps of ordinary work, no technique in any of them, and the reason most firms do not have a resolution rate is not that it is hard. The containment rate was already on the report and looked like the same number, so nobody funded those four steps.
What did reading three hundred closed conversations find?
Sumeru did the four steps. Three hundred of the 16,800 closed conversations were pulled at random and read by hand, and 27 of them had been closed with the customer's question unanswered. The 27 are 9.0 per cent of the sample. Applied to the whole closed population it is about 1,512 conversations of the 16,800.
Take those out and 15,288 conversations were actually resolved. Against the month's 24,000, that is 63.7 per cent rather than 70.0. The 27 were a sample, and the 9.0 per cent they produced applies to the whole of the closed population, so the headline moved by 6.3 percentage points on a finding of 27 conversations.
There is a small arithmetic point worth noticing, and it is not a coincidence. The gap between the two rates is 6.3 percentage points, and about 1,512 conversations is also 6.3 per cent of the month's 24,000. The match is arithmetic on the locked counts rather than a second finding. The containment rate and the resolution rate share the same base of 24,000, so every conversation moved out of the resolved count moves the two rates apart by exactly its own share of that base.
Before the control below is moved: 300 closed conversations were read by hand to see whether the question had been answered. How many had not?
Move the unanswered share, and watch one rate slide while the other refuses to
One control: the share of the 16,800 closed conversations in which the customer's question went unanswered, from 0 to 20 per cent. One consequence: the resolution rate, drawn as a marker on the same track as the containment rate. Nothing the control moves is visible to the containment rate, so it stays pinned at 70.0 per cent. The default is Sumeru's measured reading of 9.0 per cent, taken from 300 transcripts read by hand: about 1,512 unanswered, 15,288 resolved, and a resolution rate of 63.7 per cent against a containment rate of 70.0. At 0 per cent the two markers sit on top of each other. A firm reporting containment alone is silently assuming exactly that. At 20 per cent the resolution rate is 56.0 per cent.
Share of closed conversations left unanswered: 9.0 per cent
Educational illustration. On screen: 24,000 conversations in the month, 16,800 closed without a person, 7,200 handed to a person, and an unanswered share estimated from a sample of 300 read by hand rather than from all 16,800. Nothing the control moves is visible to the containment rate, so it holds flat at 70.0 per cent throughout. The 63.7 per cent is never set beside the chain's straight through rate. The two measure different populations. Figures are the invented bank's own and describe one deployment in one month.
Why can a resolution rate never be set beside a straight through rate?
Here is a report that gets written every quarter somewhere. The assistant resolved 63.7 per cent of conversations and the intake chain decided 65.0 per cent of files without a person. Two numbers of similar size, printed one under the other, inviting a reader to conclude something about the relationship between them.
There is no relationship. The 63.7 per cent counts conversations, of which there were 24,000 in the month. The 65.0 per cent counts loan files that reached a decision engine, of which there were 8,600. Neither population is a subset of the other: a customer can hold four conversations and no application, or an application and no conversation at all, and both are ordinary. Setting the two side by side compares nothing, and the closeness of the two figures is arithmetic coincidence rather than evidence.
Both numbers arrived from the same deployment in the same month, and both are shares out of a hundred. Setting them side by side is the most common reporting error around assistants, and the easiest one to make. The test is one question long: what is the denominator, and is it the same denominator? If two rates do not share a base, they do not belong on the same line of a report without a sentence explaining that they measure different things.
A report sets a 63.7 per cent resolution rate beside a 65.0 per cent straight through rate. What is wrong with that?
How to Evaluate a Financial-Services Chatbot: the six questions to ask
Everything above collapses into a single sheet to take into a room with whoever is proposing one. Six questions, in order, and they are ordered so that a poor answer to an early one makes the later ones pointless.
Question 2 is the load bearing one. Every other question on the sheet can be answered from the design of the thing: what it counts, how it splits its escalations, what it does at the edge of what it holds, whether it can act, what it writes down. Question 2 cannot. A resolution rate can only be produced by somebody reading what actually happened in the firm's own conversations. So it is never on a slide when a system is first proposed, and it is never anybody's job until somebody is given it.
Question 6 connects to consent management, and deserves a word of its own. A conversation is a record. If the assistant told a customer something about a retention period, a sharing arrangement or how to withdraw a permission, then what it said, when it said it, and which version of the wording it was working from all have to survive being asked about eighteen months later. A conversation that vanishes at the end of the session cannot support any of that.
Which of the six evaluation questions cannot be answered by the supplier alone?
What should it do when it does not know the answer?
There are exactly three things an assistant can do when the match at move 1 found nothing close, and choosing between them is a design decision somebody makes deliberately, or a design decision somebody makes by accident.
The assistant can say it does not know and offer a person. A second option is to answer anyway, producing a confident sentence from a poor match. A third is to ask the customer to rephrase and go round again. Most of the harm comes from the second option. The sentence it produces reads exactly like the sentence a good match produces, and nothing on the customer's screen distinguishes them. DeflectionEnding a conversation without answering it, which a containment rate counts as a success. is the third one repeated until the customer stops, and it is the mechanism that generated most of Sumeru's 1,512.
Only the first option is safe where the answer affects a customer's money, and it is also the option that lowers the containment rate. Choose it on purpose, rather than leaving it to whatever optimises the reported number. Charles Goodhart described the shape in the nineteen seventies and everyone has been rediscovering it since: once containment becomes the target, the cheapest way to raise it is to make giving up easier than being handed on, and nothing in the reported number objects.
The error that gets made, and what it costs
Sumeru reported 70.0 per cent containment and treated it as a service result. The 70.0 per cent was not a service result but an operational one, and the two were never separated on any report anybody read. The assistant was performing well by the only measure the deployment produced, and the measure was blind to the thing the deployment existed to do.
Reading 300 closed conversations found 27 that had been closed with the question unanswered, being 9.0 per cent. About 1,512 conversations a month therefore sit in a state the reported number counted as a success. Set that beside what the bank could already see and did not connect: 720 escalations a month were customers asking the same question a third time. The 720 and the 1,512 are separate populations, one escalated and one closed, and together they are 2,232 conversations a month, being 9.3 per cent of the 24,000, in which an answer did not land. The 720 were sitting in the escalation report the whole time.
The cost is not a number of rupees, and inventing one would be worse than leaving it out. The cost is about 1,512 people a month who came with a question, left without an answer, and were counted as having been served. Some of them will have called instead, at a cost nobody attributed to the assistant. Some will simply have stopped asking. The fault is in the measure rather than in the component: no amount of care in building the assistant produces a number that can see the difference between an answer and a departure.
How a lender, an analyst or a customer actually uses this
An operations head at a lender reads a monthly assistant report in about ten minutes, in this order. First, is the denominator the same for every rate in the report? Second, is a resolution rate printed anywhere near the containment rate? Third, how do the escalations split by numbered reason, and which way is reason 4 moving month on month? Reason 4 is a quality reading dressed as a volume reading. And is anybody reading transcripts this month, or did that stop after the first review? At Sumeru's volumes the daily arithmetic is easy to hold: at the bank's own 20 working days, 24,000 conversations a month is 1,200 a working day, of which 840 close without a person and 360 go to somebody.
An analyst comparing two firms treats a containment rate the way they would treat any self reported operational figure. A firm reporting 85 per cent containment and no resolution rate has said less than a firm reporting 70.0 per cent with a transcript reading behind it, and a rising containment rate with no resolution rate beside it is as likely to be a design that made giving up easier as a design that got better at answering. Ask what the number counts before asking whether it is high.
And the customer's version of all this is short. A customer who has asked the same question twice and got nothing usable is not a difficult customer and has not misunderstood anything. The next step is to ask for a person. Asking for a person is escalation reason 1 at Sumeru, 2,448 conversations a month, and it is the single most legitimate thing on the list.
Which supervisor sets the expectations, and on whom
Where an assistant speaks to a customer of a regulated lender, the conduct and record keeping expectations on that lender are set by the Reserve Bank of India at rbi.org.in, and complaint handling in particular sits under a named supervisor rather than under any service target a firm sets for itself. Where the institution deploying such an assistant is a market intermediary rather than a lender, the equivalent expectations sit with the Securities and Exchange Board of India at sebi.gov.in. Where a conversation is handled by a service bought in rather than built, the outsourcing expectations are set by the same supervisor.
Subjects covered elsewhere. How a language model produces words, and what it does when it has nothing to draw on, is established separately. The drafting component inside the intake chain writes a first draft of an exception note rather than talking to a customer, and is covered separately. Digital onboarding as a whole path is also covered separately. The invented bank's own numbers hold no cost for the assistant itself. For the chain it sits beside they hold Rs 2,40,00,000/- to build once and Rs 65,00,000/- a year to run, and no part of either belongs to the assistant.
Sources
| Source | Document | Site |
|---|---|---|
| Reserve Bank of India | Conduct, complaint handling, outsourcing and record keeping expectations on a regulated lender | rbi.org.in |
| Securities and Exchange Board of India | Equivalent expectations where the institution deploying such an assistant is a market intermediary | sebi.gov.in |
| Bank for International Settlements | Material on the supervision of technology deployed in customer facing processes at banks | bis.org |
| Ajay Agrawal, Joshua Gans and Avi Goldfarb | Prediction Machines, on a component that produces a cheap prediction somebody else still has to judge | Harvard Business Review Press |
Sumeru Bank Limited and its intake chain are invented.
Educational material. Not advice on any investment, tax, budget or market position.
