How to Test an AI Chatbot for Accuracy (Step-by-Step Guide)
By Muhammad Usman · · 8 min read
Why chatbot accuracy needs its own test
AI chatbots don't fail like normal software. They rarely crash; instead they answer confidently and wrongly. A bot can quote an old price, invent a refund rule, or promise a delivery time your business never offered, and nothing in your logs will look like an error. The only way to know is to check what the bot actually says against what is actually true for your business.
That is why accuracy testing has two parts: a fixed test set you run on purpose, and monitoring of real conversations after launch.
Step 1: Collect real questions
Start with questions customers really ask, not questions you imagine. The best sources are:
- Exports from your chat platform (Intercom, Zendesk, Tidio, Crisp, Chatbase or your own logs)
- Your support inbox and the questions your team answers every day
- Search terms on your help centre
Pick 20 to 50 questions that cover your most important topics: pricing, delivery, refunds, cancellations, opening hours, product limits, and anything with legal or financial risk.
Step 2: Write the expected facts, not the expected sentence
For each question, write the facts a correct answer must contain, not the exact wording. The bot can phrase things in many ways; what matters is whether the facts are right.
- Question: "How much is shipping to Germany?"
- Must contain: "€9.90", "3 to 5 business days"
- Must NOT contain: "free shipping"
Adding "must not say" facts is important: it catches the most dangerous mistakes, such as promising something you don't offer.
Step 3: Ask the bot and grade every answer
Send each question to the bot and compare the answer with the expected facts and with your help articles. Use a small set of clear verdicts:
- Correct: matches your docs.
- Not in docs: the claim might be true, but nothing in your documentation supports it.
- Made up: contradicts your docs or invents specifics like prices, dates or policies.
- Should escalate: the customer needed a human (a complaint, a legal threat, a refund dispute) and the bot kept talking.
- Off policy: breaks one of your rules, for example mentioning a competitor.
Grading by hand works for a first test. For anything larger, use an automated grader (a tool that uses an AI model to compare each answer against your documentation), and review a sample of its verdicts yourself.
Step 4: Fix the source, not the symptom
When an answer is wrong, find why. In most cases the bot is repeating something outdated or missing from your help articles. Group the wrong answers by the article that caused them, and fix the articles that cause the most errors first. One outdated "Shipping rates" page can be behind dozens of wrong answers.
Step 5: Re-test after every change
Run the same test set again after you change the bot's prompt, model, settings or knowledge base. This is regression testing: it shows whether a fix in one place broke something elsewhere. Running it automatically every night catches problems caused by changes you didn't make, such as a platform update.
Step 6: Monitor real conversations
A test set only covers the questions you thought of. Real customers ask everything else. After launch, check a sample of live conversations, or better, check every answer automatically and get an alert when the bot makes something up, fails to hand over to a human, or stops replying.
How often should you test?
- Before launch: the full test set.
- After any change: the full test set.
- Every night: the most important questions (pricing, refunds, legal topics).
- Always: live monitoring with alerts.
How ProofMyAI does this
ProofMyAI grades every chatbot answer against your own help articles, runs your test questions every night, and alerts you by email, Slack or WhatsApp when an answer is wrong. Start with a free audit of a chat export.
✅ Key takeaways
- Test with real customer questions and the facts a right answer must contain.
- Grade answers against your own documentation: correct, not in docs, made up, should escalate, off policy.
- Fix the help article behind the wrong answers, then re-test after every change.
- Monitor live conversations, because customers ask questions your test set doesn't cover.
❓ Frequently asked questions
How many test questions do I need?
Start with 20 to 50 real questions covering pricing, delivery, refunds, cancellations and anything with legal or financial risk. Add a new question every time a customer finds a wrong answer.
Can I test a chatbot automatically?
Yes. Send each test question to the bot's API and grade the answers automatically against the expected facts and your help articles. Tools like ProofMyAI run this every night and alert you when an answer gets worse.
What is the most common cause of wrong chatbot answers?
Outdated or missing information in the help articles the bot uses. Fixing the article usually fixes many wrong answers at once.