Skip to main content

Why your support bot hallucinates, and how to fix it

· 8 min read
Documentation studio

When a support bot tells a customer something false, the instinct is to blame the model and start adjusting the prompt. Sometimes that is the problem. More often, in our experience, the bot is faithfully summarizing content that was missing, stale, contradictory or impossible to retrieve, and the model is the messenger. This guide is a diagnosis method: how to trace each bad answer to its cause, and which fix belongs to which cause.

It applies to most retrieval-based bots, whether you built one in-house or use the AI agent built into your help desk. For the strategic side, meaning what to decide before you ever deploy a bot, see is your documentation strategy ready for AI. This piece assumes the bot is live and misbehaving.

How a support bot produces an answer​

Most support bots follow the same basic pattern, usually called retrieval-augmented generation. The customer's question is used to search an index of your content. The top matching passages, often called chunks, are handed to a language model along with instructions. The model writes an answer from those passages.

That pipeline has three places where things go wrong: what is in the index, what the search step retrieves, and what the model does with it. A wrong answer is the visible symptom of a failure in one of them, and the fixes are completely different for each.

The five content failures​

When we review a bot's bad answers, nearly every one falls into one of five categories. The first question for any wrong answer is which one.

FailureWhat happenedTypical signFix belongs to
MissingThe answer is not in the content at allThe bot improvises something plausibleContent: write it, or route to a human
MissedThe answer exists but was not retrievedRetrieved passages are about something nearbyStructure and vocabulary
ConflictingRetrieved passages disagree, or one is staleThe answer mixes old and new instructionsContent governance
Out of scopeThe retrieved passage applies to a different plan, version, platform or audienceCorrect instructions for the wrong customerMetadata and explicit applicability
MisreadThe right passage was retrieved but is unusable out of contextSteps missing, wrong values from a tablePage structure

A sixth cause sits outside the content: instructions. A bot told to always be helpful and never say it does not know will fill every gap with invention. That one is a configuration fix, and it only works once the content failures are under control, because a bot that says "I don't know" to half the questions is not a success either.

Step 1: Collect the bad answers​

Gather a set of real failures, not hypothetical ones. Good sources are conversations customers rated poorly, conversations that escalated to an agent right after the bot answered, and a sample your support team reviews each week. Fifty well-documented failures teach you more than five hundred unread ones.

For each failure, record:

  • The customer's question, exactly as asked.
  • The passages the bot retrieved, with their source pages.
  • The answer it gave.
  • The correct answer, and the page where it should have come from, if one exists.

If your bot platform does not show you what it retrieved for each answer, fix that first. Without retrieval logs you cannot tell a missing-content problem from a retrieval problem, and you will end up rewriting prompts for issues that live in the content.

Step 2: Classify each failure​

Walk through each failure with three questions, in order:

  1. Does the correct answer exist anywhere in the indexed content? If not, it is missing.
  2. Was the correct passage among the retrieved passages? If not, it is missed.
  3. If it was retrieved, was anything else retrieved that disagreed with it, or applied to a different product, plan or version? If yes, it is conflicting or out of scope. If no, and the answer is still wrong, it is misread, or an instruction problem.

Count the categories. The distribution tells you where to spend the next month. A bot whose failures are mostly missing content needs writers; one whose failures are mostly missed retrievals needs restructuring; one with lots of conflicts needs an audit and a retirement pass.

Fixing missing content​

Missing content is the easiest to diagnose and the most tempting to fix badly. Do not respond by indexing everything you can find, such as internal wikis, old PDFs and ticket transcripts. That adds conflicts faster than it adds answers.

Instead, write the missing articles deliberately, in the customer's vocabulary, starting from the questions that failed. Support tickets are the best source for both the gaps and the wording, which is the method we describe in turning support tickets into a knowledge base. For questions the bot should never answer, such as refunds, legal terms or account-specific investigations, give it an explicit escalation path instead of an article.

Fixing missed retrievals​

The answer exists, but search did not find it. Common causes and fixes:

  • Vocabulary mismatch. Customers say "cancel my subscription", the docs say "manage plan lifecycle". Retitle and rewrite headings in customer words, and keep a list of synonyms. The handbook page on naming and vocabulary covers how to keep terms consistent.
  • The answer is buried. One long page covering twelve topics produces chunks that each look relevant to nothing. Split pages so each answers one question, with a heading that states it.
  • Headings that say nothing. "Overview", "Details" and "Other" carry no retrieval signal. A heading should work as a question or a task.
  • Content the indexer cannot read. Text inside images, PDFs with broken text layers, content rendered only by client-side scripts, and pages behind a login the crawler does not have. Check what the index actually contains, not what the site shows.

Fixing conflicts and staleness​

Two passages disagree when the same thing is documented twice and only one copy was updated. Find duplicates, pick one owner and one page, redirect or delete the other, and remove drafts, archived pages and internal-only notes from the index. Then connect documentation updates to releases, so that a change in the product produces a change in the docs in the same cycle. A content audit is the systematic version of this; our content audit post walks through it.

Fixing out-of-scope answers​

The bot retrieved accurate instructions for the wrong customer: the enterprise plan's SSO setup for a starter-plan user, the desktop app's menu for a mobile user. Two fixes work together. Add metadata to pages, such as product, plan, platform, version and audience, and configure retrieval to filter on it where your platform allows. And state applicability in the text itself, near the top: "Available on the Business and Enterprise plans." A model can only respect a limitation it can see.

Fixing misread passages​

Here the right passage was retrieved, but out of context it misleads. Typical culprits are procedures that depend on a previous page ("as described above"), tables whose header row lands in a different chunk, pronouns with no referent, and critical values shown only in screenshots. Make sections self-contained, keep tables simple with clear headers, and put key values in text. This is the same discipline that makes docs readable for humans, which we argue in why documentation structure beats writing style.

Build an evaluation set and keep it​

Once you have classified your failures, turn them into a regression test. A plain file of real questions, each with the page that should answer it and the facts the answer must contain, is enough:

- question: "How do I download last month's invoice?"
expected_source: /billing/download-an-invoice
must_include: ["Billing", "Invoices", "Download PDF"]
- question: "Can I use SSO on the Starter plan?"
expected_source: /security/single-sign-on
must_include: ["Business", "Enterprise"]
must_not_include: ["Starter plan includes SSO"]

Run it after every significant content change and every bot configuration change. The handbook page on building a content inventory explains how to tie each expected source back to an owned page, so failures have someone to go to.

Get a diagnosis​

If your bot is live and you are not sure where its failures come from, an outside review is often faster than another round of prompt changes. Our AI-ready documentation service starts with a $2,500 audit that classifies real failures and maps each to a content fix, with implementation from $6,000. Send us a short description of your bot setup and where the content comes from through the contact page. We reply within 1 business day.