AI Academy · Book
Executives & Directors · Module 05 · Chapter 006

AI in Customer Service

A service conversation can end without the customer's problem ending. Customer-service AI earns its place when it resolves the issue with less effort, takes only the actions it is allowed to take, and hands over to a person, with context, when it should. The company owns every word it says, in chat and now on the phone.

≈ 16 min read

After this chapter you can

  • Distinguish answering a customer's question from resolving the customer's issue, using the service workflow.
  • Compare self-service, agent assist, workflow automation and bounded service agents by access and control.
  • Use the field evidence on agent assist to choose a pattern contact reason by contact reason.
  • Explain why grounded knowledge, action limits and designed handover are part of the service design.
  • Replace deflection with repeat contact, effort and quality measures, and state the EU duty to disclose AI interaction.

In January 2024 a London musician went looking for a missing parcel. He opened the chat on the website of DPD, the parcel carrier, and asked for help. By his own account, as IBTimes UK reported it, the assistant could not give him any information about the parcel, could not pass him to a human, and could not give him the number of the call center1.

Stuck, he started to play. He asked the assistant to ignore its rules and swear, and it did. He asked for a poem about how bad the company was, and it wrote one. His post of the exchange reportedly reached about a million views within 24 hours1. DPD said an error had occurred after a system update, disabled the AI element of the chat and noted that it had run AI in that chat successfully for years2.

The swearing made the headlines. The more instructive failure came first. The assistant replied to every message the customer sent, so on most service dashboards the conversation would have counted as handled. The customer still had no parcel, no answer and no way out.

Answering is not resolving

That gap is the central idea of this chapter. An answer tells the customer something. A resolution changes the customer’s situation and confirms it. Many service systems are built, measured and celebrated on the first while the customer only cares about the second.

Consider one illustrative contact at an online furniture retailer: a customer asks to change the delivery address on her order. One design replies with five clear, accurate steps to make the change in the app. The other checks her identity, updates the address and sends a confirmation. Both conversations end politely, and both count as handled by AI. In the first, the app asks for a security code she never receives, and she calls the contact center the next morning, annoyed. In the second there is nothing left to do.

Answering ends the conversation; resolving ends the customer's issue.ANSWEREDThe customer is told how tofix it.The conversation endsRESOLVEDThe problem is fixedand confirmed.The issue endsvs
Figure 5.6.1 An answer ends the conversation. A resolution ends the customer’s problem.

Customers rarely care whether the outcome came from a person, a chatbot or an automated workflow. They care whether the problem was fixed, how much effort it took, and whether they can trust what they were told. Effort is the part of service that most reliably drives customers away, and repeat contact is its most common and most measurable form3. So the design target for service AI is resolution with less effort and trust intact. It is not a channel, and it is not a smaller headcount.

Service is a workflow, not a list of answers

Customer service covers far more than questions: troubleshooting, orders, returns, complaints, account changes and service recovery. Behind each contact sits the same short workflow, and AI can take part at almost every step.

One service contact runs from understand to retrieve, act and resolve, and every contact feeds learning upstream.UnderstandIntent andcontextRetrieveApprovedknowledgeActIn approvedsystemsResolveConfirm withthe customerLearn: every contact shows what is broken upstream
Figure 5.6.2 One contact, five steps. Most first-generation chatbots stopped after the second.

Walk the delivery-address contact through it. Understand: the customer wants her address changed, and the system already knows who she is from her login. Retrieve: the current policy for address changes, including which identity check it needs. Act: update the order in the delivery system. Resolve: confirm the change and tell her the delivery date still holds. Learn: if many customers ask the same thing, the app’s own flow is broken, and the product team should hear about it.

The first-generation design stops after Retrieve. It finds the right article and reads it out. That is why the question “should we have a chatbot?” is the wrong place to start. It chooses the answer before anyone has looked at the work. The better first question is: for our most common contact reasons, what does resolved mean, and which steps of the workflow stand between the customer and that outcome?

Four patterns, each with more access

Service AI comes in four common patterns. Each one gives the system more access to customer data and more power to act, and each one needs more control.

Four service AI patterns rise from self-service answers to bounded agents; each step adds access and action and needs more control.Self-service answersApproved informationAgent assistAI suggests; the agent decidesHUMAN STAYS RESPONSIBLEWorkflow automationFixed, rule-based stepsBounded service agentPlans, acts and checks within limits
Figure 5.6.3 Value rises up the ladder, and so do the controls each rung needs.

Self-service answers give the customer approved information directly. For simple informational requests, such as opening hours or what a tariff includes, that is often enough. Agent assist keeps a person in the conversation and puts AI beside them: it summarizes the history, suggests a reply, finds the policy and proposes the next step, while the agent reviews and stays responsible. Workflow automation runs fixed steps, such as sending a password-reset link; not every step needs AI, and the simplest reliable mechanism is usually the right one. A bounded service agent plans, uses tools, acts and checks the result inside limits you set. “Reschedule my technician visit” becomes: verify identity, check eligibility, find a slot, book it, confirm.

Agent assist has the strongest field evidence. Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied the staggered rollout of a generative AI assistant to 5,172 customer-support agents. Access to the assistant raised productivity, measured as issues resolved per hour, by 15 percent on average4.

Across 5,172 support agents, an AI assistant raised issues resolved per hour by 15 percent on average, and customers asked for a manager less often.5,172Support agents studiedStaggered rollout of anAI assistant15%More issues resolvedper hourAverage gainFewerRequests for a managerCustomers were also more politeSource: Brynjolfsson, Li and Raymond, QJE · 2025
Figure 5.6.4 The study counted resolved issues, not conversations, and the gains were uneven: largest for less experienced agents.

The average hides the most useful finding. Less experienced and lower-skilled agents improved in both speed and quality. The most experienced agents gained a little speed and lost a little quality. Customers were more polite and less likely to ask to speak to a manager4. Agent assist is therefore a strong option where issues are complex, judgment matters or the knowledge base is large, and especially where turnover keeps the team junior. It is not the only right place to start. Simple, well-documented requests can go straight to self-service, and the right rung is a decision you make contact reason by contact reason.

The evidence thins as you climb. In a Gartner survey of 5,728 customers, run in December 2023, 73 percent had used self-service at some point in their journey, yet only 14 percent of issues were fully resolved there. Even for issues customers called very simple, the figure was 36 percent. The most common reasons were that the company did not understand what the customer was trying to do and that the customer could not find relevant content5. For bounded service agents the public record is thinner still. Most published figures are company announcements of conversations handled, not independent measures of problems solved; Where AI Reduces Cost tells how Klarna, after one such announcement in 2024, began hiring people again so customers could always reach a human. Treat the top rung as a capability to pilot on a few contact reasons until your own repeat-contact data show that it resolves.

A fluent answer hides what makes it right

Service knowledge is usually scattered across product documents, policies, old tickets, wikis and release notes. AI can give agents and customers one way to search all of it and answer with sources. That is real value, and it carries a specific danger. A generative model can produce a plausible answer when it does not know the right one, so in service the failure looks like the wrong policy, the wrong price or the wrong refund rule, delivered politely. Accuracy, Hallucination and Reliability, in Module 06, explains why.

A fluent answer is only the visible tip; approved sources, owners, review dates, policy version and a not-found rule make it right.WHAT THE CUSTOMER SEESA fluent answerWHAT MAKES IT RIGHTApproved sourceNamed ownerReview dateCurrent policy versionA 'not found' rule
Figure 5.6.5 The customer sees only the answer. Everything that makes it trustworthy sits below the waterline.

What makes the answer right is not the model. It is an approved source for each topic, a named owner for that source, a review date, the current version of the policy, and a rule for what the system says when it finds nothing. The last item matters most. “I could not find that; let me connect you to a colleague” is a good answer. A confident guess is not. The service leader’s job is narrow and hard: make sure every topic the assistant may speak about has an owner who will notice when the policy changes.

Before it acts, ask what it may change

Informational service tells the customer something. Transactional service changes something: an address, an appointment, a refund, a contract. The design has to tell the two apart, because a wrong answer misleads one customer while a wrong action changes a record, moves money or cancels a service.

Information requests get a sourced answer; actions proceed only when identity and limits pass, otherwise the case goes to a person with full context.Does it changeanything?NoInformation onlyAnswer fromapproved sourcesShow the sourceYesAn actionChecks pass: act and confirmOutside limits: hand over
Figure 5.6.6 Information gets a sourced answer. An action needs two checks: is this the customer, and is it within our limits?

When a request would change something, two checks come first. Is the customer who they say they are? Is the action inside the rules and limits the company has set, such as a refund ceiling or an eligibility rule? If both pass, the system acts and confirms. If either fails, or the system is unsure, it hands the case to a person with the full context, so the customer never has to start again.

Two principles follow. First, permissions follow the action, not the label. “It is only a service assistant” is no reason to give it broad access to customer records or payment systems, because every extra permission is another way for one wrong answer to become one wrong action. Second, escalation is a feature, not a failure. The US Consumer Financial Protection Bureau found that all ten of the largest US commercial banks use chatbots in customer service, and it has received complaints about “doom loops”: repetitive, unhelpful replies with no off-ramp to a human6. The DPD customer was in one. A designed handover, with the conversation attached, is what separates a service system from a wall.

Voice agents: capability, not yet outcome

Until recently, the honest view of the phone channel was that AI could transcribe and summarize calls while a person handled them. That view is out of date as a statement of capability. Research speech models can now listen and speak at the same time, cope with interruptions and overlapping speech, and respond in about 200 milliseconds7. Connected to the same knowledge and systems as a chat assistant, a voice agent can in principle handle a complete call: identify the caller, look things up, take an approved action and confirm it.

What that evidence does not show is results. A latency figure from a research model says nothing about how often a voice agent resolves the caller’s problem, how many callers ring back, or how many hang up or ask for a person. As of October 2026, independent published measures of those outcomes in live service are scarce, and most figures in circulation come from vendors. Given that only one issue in seven is fully resolved in self-service today, a voice agent should be judged on the same repeat-contact test as chat before it takes a whole contact reason.

Voice agents can now handle full calls, but must disclose that they are AI, hand over with context and keep the same limits as chat.Now possible, not yet provenWhole callsInterruptions handledActions in approved systemsStill requiredSay it is AIHandover with contextSame limits as chatEU AI Act, Article 50: transparency duties apply from 2 Aug 2026
Figure 5.6.7 Voice changes what the channel can do, not yet what it is proven to achieve. It does not change the rules.

The rules do not change with the channel. The same approved sources, the same action limits and the same handover apply. One duty becomes sharper. A natural-sounding voice makes it harder for a caller to tell they are talking to a machine, so the opening line of every AI call is a design decision, and in the European Union it is also a legal one.

Deflection is not success

Many service dashboards lead with deflection: the share of contacts that never reached a person. It is easy to report and easy to game. A conversation that ends because the customer gave up looks exactly like one that ends because the problem was fixed. The DPD conversation, and the first delivery-address design, would both have counted as deflected.

Better measures follow the customer after the conversation ends. Did they contact the company again about the same issue within a week? How much effort did it take to get it fixed? Add a regular quality sample: was the answer accurate, did it follow current policy, and was the handover timely? Deflection is valuable only when the problem is actually resolved, so report it next to repeat contact, never alone.

A small calculation shows why. Suppose, as an illustration, a contact center receives 100,000 contacts a month and the assistant contains 60 percent of them. At 5 per human contact, the dashboard reports 300,000 a month saved. Now suppose 30 percent of the contained customers come back within a week about the same issue, and that a returning contact costs 7, because it is longer and the customer is already annoyed. That is 18,000 returns costing 126,000. The real saving is 174,000, about 42 percent less than reported, before the assistant’s own running cost.

An illustrative reported saving of 300 thousand a month falls to 174 thousand once 18,000 repeat contacts at 7 each are counted.300 kSaving the dashboard reports−126 kReturns within a week (18,000 at 7)174 kSaving after repeat contactILLUSTRATIVE NUMBERS
Figure 5.6.8 Illustrative numbers. Contacts times repeat-contact rate times cost per return: the term deflection leaves out.

Two executive decisions hide inside these numbers. The first is what real recovered capacity buys: shorter waits, more time on hard cases, proactive outreach or lower cost. Someone has to choose. The second is what the contacts teach. Every contact reason is evidence about something broken upstream, and fixing the cause produces the cheapest contact of all: the one that never happens.

Story: the bereavement fare

In November 2022 a British Columbia man’s grandmother died. Before booking a flight to Toronto, he asked the chatbot on Air Canada’s website about bereavement fares, the reduced fares some airlines offer for travel after a death in the family. The chatbot told him he could book at the normal fare and apply for the bereavement rate afterwards, within 90 days of the ticket being issued. It also linked to the airline’s bereavement policy page, which said the opposite: bereavement rates could not be claimed for travel already completed. He booked, flew and applied. The airline refused10.

Timeline of Moffatt v. Air Canada - the chatbot's wrong advice in November 2022, the refused claim, and the February 2024 ruling against the airline.Nov 2022The chatbot answersBook now, claim the rate within90 daysAfter travelThe claim is refusedThe policy page said noretroactive claimsFeb 2024The tribunal rulesThe airline is liable
Figure 5.6.9 A fluent answer, a contradicting policy page and no handover: the airline paid for what its chatbot said.

He took the airline to British Columbia’s Civil Resolution Tribunal. The airline argued, in effect, that the chatbot was a separate legal entity responsible for its own actions. The tribunal called this “a remarkable submission” and rejected it. The chatbot was part of the airline’s website, and the airline was responsible for all the information on that website, whether it came from a static page or a chatbot. Nor should a customer have to check one part of the site against another. The tribunal found that the airline had not taken reasonable care to make its chatbot accurate, held it liable for negligent misrepresentation, and ordered it to pay 650.88 Canadian dollars in damages, plus interest and tribunal fees, about 812 Canadian dollars in all10.

The sum was small. The reasoning is what boards remember, though it comes from a Canadian small-claims tribunal and is not a binding precedent elsewhere. The chatbot answered. Nothing grounded the answer in the approved policy, which existed one click away. No owner checked what the assistant said about a sensitive, money-related topic. And nothing told it to hand a bereavement question to a person. Each missing control is cheap to add before launch and expensive to explain afterwards. The company owns every word its AI says to a customer.

What this means for leaders

Customer service is where AI meets customers at scale, every day, and where its mistakes are made in the company’s name. The leadership work is less about choosing a tool than about four decisions. Define resolved for each of the main contact reasons, and design back from it. Choose the pattern contact reason by contact reason, with agent assist as a strong option where the work is complex and the team is junior. Set what the AI may change, the limits on each action and the exact point at which a person takes over. And measure the customer’s outcome, not the conversation’s end.

Check yourself

  1. The best customer-service AI is the one that deflects the most contacts.
  2. A handover to a human means the AI failed.
  3. If its chatbot gives wrong information, a company can argue the chatbot was responsible.
  4. In a large field study, AI assistance helped less experienced support agents most.
  5. Fast, natural-sounding voice agents are proven to resolve more calls.
  6. In the EU, people must be told they are interacting with an AI system unless it is obvious.

Reflection: follow one customer

What comes next

Customer service shows AI at the edge of the enterprise, where every answer is a promise to a customer. The next chapter moves to its operational core, where a prediction is not a message on a screen but a decision about stock, vehicles and machines: AI in Operations and Supply Chain.

Laws referenced

EU AI Act · EU

Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744

Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.

  • 2024-08-01 — Entered into force
  • 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
  • 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
  • 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
  • 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
  • 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
  • 2028-08-02 — High-risk obligations for AI in products regulated under Annex I

Last verified 2026-10-06 · official text

References

  1. International Business Times UK. DPD Disables AI Chatbot After It Swears And Calls Company 'Worst Delivery Firm'. International Business Times UK. 2024.
  2. The Register. DPD chatbot goes rogue: swears and writes poems criticizing the company. The Register. 2024.
  3. Matthew Dixon, Karen Freeman and Nicholas Toman. Stop Trying to Delight Your Customers. Harvard Business Review 88(7/8). 2010.
  4. Erik Brynjolfsson, Danielle Li and Lindsey Raymond. Generative AI at Work. The Quarterly Journal of Economics 140(2). 2025.
  5. Gartner. Gartner Survey Finds Only 14% of Customer Service Issues Are Fully Resolved in Self-Service. Gartner press release. 2024.
  6. Consumer Financial Protection Bureau. Chatbots in consumer finance (Issue Spotlight). Consumer Financial Protection Bureau. 2023.
  7. Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave and Neil Zeghidour. Moshi: a speech-text foundation model for real-time dialogue. arXiv:2410.00037. 2024.
  8. European Parliament and Council of the European Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union. 2024.
  9. European Union. Regulation (EU) 2026/1744 (Digital Omnibus on AI) amending Regulation (EU) 2024/1689. Official Journal of the European Union. 2026.
  10. Civil Resolution Tribunal (British Columbia). Moffatt v. Air Canada, 2024 BCCRT 149. CanLII. 2024.

Further reading

Sources last verified 2026-10-08.