AI Agent - Intelligent task automation and workflow optimization

AI Agent Use Cases Customer aiagent.app

AI Customer Service: Liability, Deflection, and What Really Resolves

AI Customer Service: Liability, Deflection, and What Really Resolves

Most vendor pages sell AI customer service as a staffing replacement. The public record is less tidy. Air Canada was ordered to pay a customer after its chatbot gave wrong advice. Klarna’s own numbers show both real volume gains and a later course-correction when cost became the only scoreboard. The FTC has already punished a company for overclaiming what its AI could do. And MIT’s Project NANDA found that a large majority of enterprise generative-AI pilots never show up on the P&L.

If you run a two-person company, you do not need another “trends” roundup. You need the failure modes, the metrics that get gamed, and a cost model that does not pretend AI is a flat seat upgrade. That is what this piece is for: what is on the record, what is vendor-sourced and should be labeled as such, and where the evidence is still thin.

What AI customer service covers — and what it does not

AI customer service usually means conversational systems that answer tickets, chat, or voice contacts before a human agent does. In practice that spans scripted chatbots, LLM-backed assistants, and newer agentic setups that take multi-step actions inside your tools. The shared promise is the same: resolve more contacts without adding headcount, cut average handle time, and keep CSAT from falling.

The shared risk is also the same: the system can sound confident while being wrong, and your company still owns the answer. That is not a philosophical point. It is how a Canadian tribunal treated an airline chatbot, and how U.S. consumer protection has started treating overstated AI claims.

For a small team, the useful split is operational, not branding. Tier 1 work — password resets, order status, policy lookups, refund eligibility — is where automation can help if the knowledge base is accurate and escalation is fast. Ambiguous edge cases, goodwill decisions, and anything that looks like legal or financial advice are where “autonomous” becomes a liability label. Human-in-the-loop is not a soft preference; it is how you keep negligent misrepresentation and deceptive-capability claims off your desk.

Klarna’s numbers — and why the “reversal” story is oversold

Klarna is the case every AI customer service pitch cites, so start with Klarna’s own press release rather than a third-party paraphrase. One month after launch, Klarna said its OpenAI-powered assistant handled conversations — two-thirds of its customer-service chats — did work equivalent to 700 full-time agents, matched human-agent CSAT, cut repeat inquiries by 25%, and reduced average resolution time from 11 minutes to under 2 minutes across 23 markets and 35+ languages, with an estimated $40M USD profit improvement for 2024.

That is a real deployment at scale. It is also one company’s self-reported month-one snapshot. Treat it as evidence that high-volume, relatively structured support can absorb a large share of chats — not as a guarantee that your backlog will behave the same way.

In May 2025 the story got noisier. Forbes, drawing on Bloomberg reporting from May 8, 2025 and Klarna’s on-record response, reported that CEO Sebastian Siemiatkowski said overweighting cost produced “lower quality,” and that Klarna began piloting a hybrid model with a small number of new human remote agents (). The same coverage captures Klarna’s pushback on the “reversal” framing: the AI assistant’s workload had grown to the equivalent of over 800 full-time roles, only two new agents were hired in a flexible remote pilot, and Klarna said it never eliminated human support, having kept several thousand outsourced agents. Forrester analysts Kate Leggett and Christina McAllister characterized the episode as overpivoting to cost containment and underestimating support complexity. Separately, CNBC headlined that Klarna’s CEO said AI helped the company shrink its workforce by .

For a founder, the useful takeaway is narrower than either camp’s slogan. Cost-only evaluation can degrade quality. Hybrid staffing can still sit next to a growing AI workload. And “we rehired everyone because AI failed” is not what Klarna put on the record.

Liability: Air Canada and the FTC already drew bright lines

In Moffatt v. Air Canada, the British Columbia Civil Resolution Tribunal held Air Canada liable after its chatbot gave a customer inaccurate information. The tribunal rejected the idea that the chatbot was a “separate legal entity,” found negligent misrepresentation, and ordered damages of CA$ plus CA$36.14 pre-judgment interest and CA$125 in tribunal fees, as summarized by Dentons. BBC Travel covered the same episode for travelers: the airline remained responsible for information on its site regardless of format ().

The dollar amounts are small. The holding is not. If your bot invents a refund policy, bereavement fare, or cancellation rule, “the model said it” is not a defense.

The FTC’s final order in the DoNotPay matter (February 11, 2025) is the complementary U.S. signal on overclaiming. The order prohibits DoNotPay from claiming its product is an “AI lawyer” without evidence, requires in monetary relief, and requires notice to subscribers who signed up from 2021 to 2023. The underlying complaint alleged DoNotPay never tested whether its AI performed to the standard of a human professional.

You do not need to be selling legal services for the pattern to matter. Marketing copy that implies professional-grade judgment, guaranteed outcomes, or “replaces your support team” without evidence is the same failure mode: capability claims ahead of proof.

Deflection rate is easy to inflate

Deflection rate is usually defined as contacts resolved via self-service or automation without reaching a human, divided by total contact attempts. On paper it is the north-star metric for AI customer service ROI. In practice it is easy to game.

Vendor documentation from Decagon and Gladly — both AI customer-service vendors with a stake in how buyers read this metric — describes the same pattern: deflection can look perfect when chatbots let sessions time out, article views go unrated, or portal sessions end without confirming the issue was solved (; ). The warning sign both point to is high deflection paired with declining CSAT or rising repeat contacts.

There is a real evidence gap here. Those glossaries explain how the metric can be gamed; there is not a named independent study in the public record (at least not one verified for this article) that quantifies how often or by how much deflection is inflated across the industry. If a vendor leads with deflection and will not show CSAT, first-contact resolution, and repeat-contact rate on the same cohort, treat the number as incomplete.

What “good” resolution still looks like

If you want a non-AI yardstick for resolution quality, SQM Group’s 2024 First Call/Contact Resolution benchmark is one of the few named, methodologically disclosed baselines. SQM sells CX benchmarking services, so it is commercially interested — but it discloses the method: annual Voice-of-Customer post-call surveys across North American call centers, with a minimum 400-survey sample per participating center. Aggregate FCR across industries averaged 69% (range 43%–88%). “Good” is 70–79%; “world-class” is 80%+, achieved by only 5% of call centers. Every 1-point FCR improvement is associated with about $286,000/year in savings for a typical midsize call center and a 1.4-point NPS gain. SQM also notes that no telecom company has hit world-class FCR in more than 25 years of tracking.

Those figures are not AI-specific. That is the point. When a demo promises that automation will “beat your team,” ask beat what. If your human FCR is already in the high sixties, the bar is not “bot answered.” It is resolved without a second contact.

One related gap deserves a plain statement: there is no credible, source-named study in the pack used for this article that quantifies the CSAT drop from a failed bot interaction followed by slow escalation. Vendor blogs assert the pattern constantly. The measured curve is under-served. Operate as if that failure mode is real; do not invent a percentage for it.

What customers say they want from AI in support

Zendesk’s 2026 CX Trends report — Zendesk sells CX and AI software, so label the commercial interest — is the eighth annual edition, based on a survey of consumers and business leaders across 22 countries, announced November 18, 2025. Selected findings: 85% of CX leaders say a single unresolved first-contact issue is enough to lose a customer; 95% of consumers expect an explanation when AI makes a decision about them, but only 37% of CX leaders currently provide one; 74% of consumers expect 24/7 availability; 81% want a rep to pick up where they left off across channels; and 76% would choose a company offering text, voice, and video in one thread.

The transparency gap is the sharpest operational warning. If your AI declines a refund, changes a plan, or routes someone to a dead end, customers expect an explanation. Most organizations, by Zendesk’s own survey, are not giving one.

Salesforce’s public preview of its “State of the AI Connected Customer” research — again vendor-commissioned, and gated beyond the landing-page stats — surveys consumers and business buyers worldwide and reports that 72% want to know when they are talking to an AI agent, while only 17% are comfortable with an AI agent making a financial decision for them, versus 46% comfortable with an AI agent handling a request for faster service. 61% say AI advances make company trustworthiness more important; 64% believe companies are “reckless” with customer data; 71% feel increasingly protective of personal data.

Speed is welcome. Silent automation of money decisions is not.

Build your own AI agent, tailored to your needs

Create customized AI agents to automate tasks and enhance productivity.

Most GenAI pilots still miss the P&L

MIT Media Lab’s Project NANDA report, “The GenAI Divide: State of AI in Business 2025” (July 2025), is the clearest cross-industry check on optimism. Despite –$40B in enterprise GenAI spend, about 95% of pilots show no measurable P&L impact; only about 5% achieve rapid revenue acceleration. Buying from a specialized vendor and partnering succeeds about 67% of the time, versus roughly a third as often for in-house builds. The report names customer support and administrative roles as areas where workforce effects are already concentrated — often through not backfilling vacated roles rather than mass layoffs (report PDF; ).

That finding travels. If your AI customer service project cannot name a P&L line — fewer paid support hours, lower cost per resolved contact, retained revenue from faster refunds — you are in the common failure set, not the rare success set.

Salesforce’s own State of Sales Report 2026 (survey of sales professionals) is unusually blunt for vendor research: 87% of sales organizations use some form of AI, but only 33% of AI initiatives hit their own ROI targets; 62% worry about unpredictable AI costs; only 21% strongly agree they have adequate agentic-AI governance (Salesforce stats; ). Support tooling and CRM AI are not identical markets, but the ROI and governance pattern is the same warning: adoption is cheap to announce and expensive to prove.

What AI support actually costs on major platforms

For small teams evaluating whether AI “justifies the price,” the honest answer on major CRM platforms is that AI is often usage-metered on top of seats — not a single percentage uplift.

Salesforce’s list pricing for CRM suites includes Free Suite at /user/month (2 licenses, assistive AI not included), Starter Suite at $25/user/month (first tier with Assistive AI with Employee Agent), and Pro Suite at $100/user/month. Agentforce — Salesforce’s AI-agent layer — is priced separately: Flex Credits at per 100,000 credits; pay-per-conversation at $2 per conversation for customer-facing agents; and per-user licensing starting at $125+/month via enterprise agreement.

HubSpot follows a similar credit pattern. HubSpot Credits cost per 1,000 when paid annually. The AI Customer Agent costs 50 credits per resolved conversation (about $0.45 per resolution) and requires Professional tier or above. HubSpot’s own, not independently audited, claims on the same pricing surface include “70%+ of conversations resolved automatically” and “39% faster ticket resolution vs. teams not using Customer Agent.” Marketing Hub seat tiers run Free / Starter $7/mo / Professional $800/mo / Enterprise $3,600/mo.

There is no independent (non-vendor) study in the verified pack that directly compares CRM-with-AI versus CRM-without-AI outcomes for support. What you can say from primary pricing pages is simpler: cost scales with conversation volume, so a low-volume founder team and a Klarna-scale queue are not buying the same product economically.

Bad data makes confident systems worse

MIT Sloan Management Review, drawing on research by Experian and consultants James Price and Martin Spratt, reported that bad data costs most companies –25% of revenue through time spent correcting errors, re-verifying sources, and absorbing mistakes. That research predates the agentic-AI wave and is about data quality generally, not AI customer service specifically. The inference — that an LLM will propagate a wrong SKU, entitlement, or address faster and with more polish — is logical, not something that source measured directly.

If your contact and account records are duplicated, your knowledge base contradicts your refund policy, or your CRM fields are half-empty, automation does not fix the mess. It performs it at scale.

Vendor lock-in when the AI layer sits inside your CRM is another under-served question. Consultancy blogs often float migration-cost figures without a named study behind them; those dollar claims are excluded here for that reason. The practical risk you can still plan for without a fake statistic: behavioral preferences and prompt-tuned workflows are often harder to export than raw tickets. Keep canonical policies and customer history in systems you can leave.

What to measure before you buy

Instrument the baseline for two weeks before any bot goes live. You need first-contact resolution (or a close proxy: resolved without a follow-up in X days), CSAT on resolved contacts, repeat-contact rate, time-to-human-escalation, and cost per resolved contact. Deflection without those peers is a vanity metric — as Decagon and Gladly both warn in their own glossaries (; ).

Compare post-launch results to something like SQM’s FCR bands so “good” has an external meaning, not just an internal dashboard win (). Track explanation rate for AI decisions against the transparency gap Zendesk reports (). Log every policy statement the bot makes that a human would need a manager to approve.

For cost, model usage, not seats alone. Use the Salesforce and HubSpot meter examples as templates even if you buy elsewhere: conversations × unit price, plus the human hours that remain after automation (; ). If unpredictable AI cost already worries a majority of sales orgs in Salesforce’s survey, assume your finance person will ask the same question ().

Kill the pilot on a written rule. If FCR does not move, CSAT falls, or repeat contacts rise while deflection climbs, you are looking at the gaming pattern, not a win. MIT NANDA’s finding that about of GenAI pilots lack measurable P&L impact is the prior you should bring into that review.

Limits a small team should accept

Do not claim the bot is a professional substitute. The FTC’s DoNotPay order is the cautionary tale for capability theater (). Do not let the bot invent policy; Air Canada’s tribunal outcome shows the company owns the words (; ). Do not optimize only for cost; Klarna’s 2025 comments are a live example of quality loss when cost is the sole evaluation factor ().

Disclose when the customer is talking to AI. Salesforce’s connected-customer preview finds most people want that disclosure (). Keep humans for financial judgment calls; comfort with AI on money decisions is low in that same dataset. Prefer specialized tools with a partner path over a heroic in-house build if you care about the success-rate gap NANDA reports ().

For founders using a platform like to delegate operations work across connected apps, the same constraints apply: agents can draft replies, triage tickets, and update records, but policy truth and escalation ownership stay with you. Automation without a clean knowledge base and a fast handoff is how you buy deflection and lose customers on the first unresolved contact — the outcome of CX leaders in Zendesk’s survey already fear.

AI customer service works when the queue is repetitive, the answers are grounded, the handoff is fast, and the metrics include resolution quality — not just avoidance of humans. The public cases and vendor-critical surveys above are enough to design that system. They are also enough to reject the version sold as a full replacement for people who still have to clean up the mess.

How the work divides

Focus areaWhat the agent doesWhat stays with a personWhat breaks without review
What AI customer service covers, and what it does notAI Agent can draft replies, triage tickets, and update records across connected apps. This fits repetitive requests such as password resets, order status, policy lookups, and refund eligibility.A person handles ambiguous edge cases, goodwill decisions, legal or financial advice, policy truth, and escalation ownership.A confident but wrong answer can create negligent misrepresentation or deceptive capability claims, especially when a refund or cancellation rule is misstated.
Klarna's numbers, and why the "reversal" story is oversoldStructured customer service work can absorb high chat volume and reduce repeat inquiries. AI Agent's documented role is narrower, drafting replies, triaging tickets, and updating records across connected apps.People compare quality with cost and retain support capacity. Klarna reported 2.3 million conversations in one month, work equivalent to 700 full-time agents, a 25% reduction in repeat inquiries, and later a hybrid model with several thousand outsourced agents and two new remote agents.A cost-only scoreboard can lower quality. Calling the episode a complete reversal ignores Klarna's reported AI workload equivalent to over 800 full-time roles and its continued human support.
Liability: Air Canada and the FTC already drew bright linesAI Agent can draft a customer reply and update the related ticket or record, including a response about a refund or policy.A person checks policy language, approves consequential replies, and owns the company's capability claims. Air Canada remained liable for chatbot information, and the FTC acted against unsupported "AI lawyer" claims.An invented refund policy can support negligent misrepresentation. Marketing that implies professional-grade judgment or guaranteed outcomes can create deceptive-capability exposure.
Deflection rate is easy to inflateAI Agent can triage tickets and draft replies, while the team measures whether the contact was actually resolved rather than treating an unanswered session as success.A person reviews CSAT, first-contact resolution, repeat-contact rate, and time to human escalation on the same cohort.Timeouts, unrated article views, and sessions ending without confirmation can make deflection look high while unresolved contacts and repeat inquiries rise.
What "good" resolution still looks likeAI Agent should help produce a resolution that needs no second contact by grounding drafted replies in the knowledge base and moving unresolved tickets to a fast human handoff.A person validates the answer, manages escalation, and compares results with an external resolution baseline. SQM reports aggregate first-contact resolution at 69%, with "good" at 70% to 79%.Counting "bot answered" as resolution hides repeat contacts and weak first-contact resolution. A failed bot interaction followed by slow escalation can damage customer experience even without a measured percentage.

Explore all AI agent use cases or start building on AI Agent.

ai customer serviceai customer supportcustomer service automationdeflection ratechatbot liabilityfirst contact resolution

Frequently Asked Questions