AI Chatbot Hallucinations: Real Examples and How to Catch Them
By Muhammad Usman · · 7 min read
What a hallucination looks like in customer service
In customer service, hallucinations are rarely dramatic. They look like normal, helpful answers:
- A shipping price or delivery time that is out of date
- A refund or cancellation rule that doesn't exist
- A discount, warranty or feature the business never offered
- A made-up link, phone number or opening hours
- Legal, medical or financial statements the business would never approve
Because the answer sounds right, customers believe it, and the business often only finds out when someone complains.
Well-known cases
Several chatbot mistakes became public and show the business risk:
- Air Canada (2024): a customer was told by the airline's chatbot that he could claim a bereavement discount after travelling. That wasn't the airline's policy. A Canadian tribunal ruled that the airline was responsible for what its chatbot said and had to compensate the customer.
- A Chevrolet dealership (2023): users got a dealership's website chatbot to "agree" to sell a car for one dollar, and screenshots spread widely online.
- DPD (2024): a parcel company's chatbot was prompted into swearing and criticising the company, and the company disabled part of the bot.
- New York City's MyCity chatbot (2024): journalists reported that the city's business-advice chatbot gave answers that contradicted local laws.
The lesson from all of them: the business owns what its bot says, and problems are found by customers or the press unless you check first.
Why chatbots make things up
- Missing information: when the help articles don't answer the question, the model fills the gap with something plausible.
- Outdated information: old prices or policies still sit in the knowledge base.
- Conflicting information: two articles say different things and the bot picks one.
- Over-helpful prompts: instructions like "always give a helpful answer" push the bot to answer instead of saying "I don't know".
- Manipulation: users deliberately push the bot to say things it shouldn't.
How to catch hallucinations
- Ground every check in your own docs. A general fact-check doesn't help; the question is whether the answer matches *your* policies.
- Separate "made up" from "not in docs". A made-up answer contradicts your docs; a "not in docs" answer might be true but isn't supported. Both need attention, the first urgently.
- Check every answer, not a sample. Hallucinations are rare per answer but common across thousands of conversations.
- Alert on high-risk topics. Prices, refunds, legal and medical topics deserve an instant alert.
- Add rules. "Never say lifetime warranty", "always hand over to a human when a customer mentions a chargeback", "when pricing comes up, say prices include VAT".
- Fix the cause. Group problems by the article behind them and update that article.
How to reduce hallucinations
- Keep help articles current and remove contradictions.
- Tell the bot to say "I don't know, let me connect you with our team" when the docs don't cover a question.
- Hand over to a human for complaints, legal threats and refund disputes.
- Re-test after every change to the prompt, model or knowledge base.
How ProofMyAI helps
ProofMyAI checks each chatbot answer against your help articles, labels it correct, not in docs, made up, should escalate or off policy, and alerts you about the dangerous ones. For regulated businesses there is a compliance setup with masking and results-only storage.
✅ Key takeaways
- A hallucination is a confident answer that isn't backed by your business's own information.
- Courts and the public hold businesses responsible for what their chatbot says.
- Most hallucinations come from missing, outdated or conflicting help articles.
- Check every answer against your docs, alert on high-risk topics, and fix the source article.
❓ Frequently asked questions
Can you stop an AI chatbot from hallucinating completely?
No model is hallucination-free today. You can reduce them with current, consistent help articles and a clear 'I don't know' instruction, and catch the rest quickly with automatic checks and alerts.
Is a business responsible for what its chatbot says?
In the Air Canada case in 2024, a Canadian tribunal held the airline responsible for incorrect information its chatbot gave a customer. Treat your chatbot's answers as statements from your business.
What's the difference between 'made up' and 'not in docs'?
A made-up answer contradicts your documentation or invents specifics. A 'not in docs' answer may be true but nothing in your documentation supports it. Both should be reviewed; made-up answers are more urgent.
Find your chatbot's made-up answers for free
Free plan, no card needed. Set up in minutes.
Start free →