Skip to main content
E-commerce / Customer Service AI

A composite engagement: real patterns, real numbers from our audit work, details merged and anonymised so no client is identifiable.

The Chatbot That Invented a Refund Policy

AI customer-service chatbot grounded against real systems

The situation

A 60-person online homewares retailer launched a support chatbot to cut ticket volume. It demoed beautifully. Six months later, support headcount was unchanged, and the team had developed a new daily task: cleaning up after the bot.

What the Reality Check found

The bot had no connection to the systems it was answering questions about. It couldn’t see stock, orders or the returns policy — so it improvised. It quoted discontinued prices, promised next-day delivery the courier didn’t offer, and on one memorable occasion granted a customer a "90-day no-questions refund" from a policy document that did not exist. Containment (queries resolved without a human) was 22%. Nobody had agreed what number the bot was supposed to move, so nobody had noticed it wasn’t moving one. Support agents had quietly built a "bot said what?" macro.

The verdict

Fix. The use case was right — 60% of tickets were order-status, returns and delivery questions with answers that live in systems. The build had skipped the wiring.

The fix (Recovery Sprint, 6 weeks, two integrations)

Grounded the bot on the live product catalogue, order API and the actual returns policy — answers cite source, and anything it can't source goes to a human, stated plainly. Added an evaluation set of 400 real past queries; every change scored against it before shipping. Hard guardrails: no invented commitments, price and delivery answers only from the API. One success metric agreed in writing before work started: containment on order-status and returns queries, measured weekly against the pre-fix baseline.

The numbers

  • Containment: 22% → 61% overall; 78% on order-status queries
  • "Bot said what?" corrections: ~40/week → under 3/week
  • Support hours returned: ~340 hours/quarter, redeployed to the trade-customer queue
  • First month post-fix: zero invented policies. The eval harness catches drift before customers do.
Is your AI paying for itself?
Book an AI Reality Check →