Skip to main content

Turning support tickets into a knowledge base

· 8 min read
Documentation studio

Your support queue is the most honest documentation backlog you will ever get. Every ticket is a customer telling you, in their own words, what the product and the docs failed to explain. The trouble is that a queue of several thousand tickets is not a backlog yet. This is the method we use to turn it into one: a ranked list of articles to write, fix or retire, each backed by real tickets and written in the words customers actually use.

The method works the same on Zendesk and Zoho Desk exports, and on most other help desks that let you get tickets and their conversation threads out.

Step 1: Decide the question before you export​

Ticket mining can answer several different questions, and mixing them produces a muddy result. We frame it as one question: which tickets could the customer have resolved alone with the right article, and what would that article need to say?

That framing excludes, on purpose, tickets that need a human: refunds, account-specific investigations, bugs that need engineering, and anything involving judgment about a contract. Those still matter, but they belong in a different report for a different owner.

Pick a time window long enough to cover a full release cycle and any seasonal patterns, such as annual renewals or tax season. Six to twelve months is typical.

Step 2: Export tickets and threads​

The subject line alone is not enough. You need the customer's first message and the agent reply that resolved the ticket, because the resolution is the first draft of the article.

On Zendesk, the cursor-based incremental export returns tickets changed since a start time, including the description field, which holds the first comment:

curl -s -u "$ZENDESK_EMAIL/token:$ZENDESK_TOKEN" \
"https://example.zendesk.com/api/v2/incremental/tickets/cursor.json?start_time=1735689600" \
> tickets-001.json
# Follow after_url in each response until end_of_stream is true.

Full conversation threads come from each ticket's comments endpoint. A quick flattening pass gives you a table to work from:

jq -r '.tickets[]
| [.id, .created_at, .status, .via.channel, (.tags | join(";")),
((.subject // "") | gsub("[\r\n]+"; " "))]
| @csv' tickets-*.json > tickets.csv

Zoho Desk exposes tickets and their conversation threads as separate API resources in the same way, so the shape of the job is identical. Both tools also offer exports from the admin interface; what those include varies by plan and role, so check before you promise anyone a date.

Step 3: Scrub personal data first​

Tickets are full of names, email addresses, phone numbers, account IDs, billing details and occasionally things customers should never have pasted into a support form. Scrub them before anyone reads the data for analysis and before any text goes near an AI tool.

Our scrubbing pass replaces emails, phone numbers, URLs with account identifiers and anything that looks like a key or token with typed placeholders, such as [EMAIL] or [ACCOUNT_ID]. Names are harder to catch automatically, so we also replace the requester's and agent's names using the ticket's own metadata. Keep the raw export somewhere access-controlled, and delete it when the project ends.

Step 4: Filter out the noise​

Remove tickets that are not support conversations: spam, auto-replies, out-of-office bounces, sales inquiries routed to the wrong queue, internal test tickets, and tickets merged into others. Keep a count of what you removed and why. When someone later asks why the analysis covers fewer tickets than the help desk dashboard shows, you want a one-line answer.

Step 5: Cluster by intent, not by tags​

Agent-applied tags and ticket form fields are the obvious grouping, and they are usually wrong in useful ways. Tags reflect how the support team routes work, not what the customer was trying to do, and they drift as agents come and go.

So we cluster by intent: what was the customer trying to accomplish, and where did they get stuck? The process we use:

  1. Read a random sample of a few hundred tickets and write a short intent label for each, in the customer's terms: "can't find last month's invoice", not "billing".
  2. Merge labels that mean the same thing into a draft taxonomy.
  3. Assign every remaining ticket to one intent, adding new ones only when a ticket genuinely fits nothing.
  4. Review the largest clusters by reading a handful of tickets from each, to catch clusters that are really two problems.

AI tools can help with the first-pass assignment on large volumes, but we treat their output as a draft and review each cluster by reading actual tickets. A cluster name that sounds right and groups the wrong tickets is worse than no cluster at all.

Step 6: Classify each cluster by what fixes it​

This is where ticket mining turns into a plan. Every cluster gets exactly one of these labels:

LabelWhat it meansOwnerFix
Content gapNo article covers thisDocsWrite a new article
Findability gapAn article exists, customers do not find itDocsRetitle in customer words, add synonyms, link from where they get stuck
Content defectAn article exists and is wrong or outdatedDocs and productCorrect it; retire the stale version
Product issueThe docs cannot fix a confusing UI or a bugProductReport with ticket evidence
Needs a humanAccount-specific or policy decisionSupportDocument the escalation path, not an answer

The findability gap is the one teams underestimate. It is common to find a cluster of tickets about something the help center already explains, under a title nobody would search for. Those are the cheapest wins in the whole project. Our content audit post covers how to classify the existing articles you will be comparing against.

Step 7: Rank and write​

Rank clusters by volume, by how completely an article could resolve them, and by effort. Then write, following a few rules:

  • Start from the resolution. The agent reply that closed the ticket is usually the right answer, tested on a real customer. Turn it into steps.
  • Title in the customer's words. Use the phrasing from the tickets, not the feature name from the product spec.
  • One problem per article. A support agent should be able to paste a single link that answers the question completely.
  • Use a template. Troubleshooting articles need symptom, cause, fix and "still stuck?" sections. The handbook page on content types and templates has the templates we use.

A worked example​

What follows is a worked example, not a client result. The product, the clusters and the counts are invented to show the shape of the output.

Imagine one batch of 500 tickets, after filtering, for a fictional billing product. The five largest clusters might look like this:

Intent clusterTicketsLabelAction
Can't find last month's invoice41Findability gapRetitle "Billing documents" to "Download an invoice"; link from the billing page
Card declined on renewal33Content gapNew troubleshooting article with decline reasons and next steps
Changing the billing email27Content defectArticle shows the old settings screen; update steps
Refund for unused seats19Needs a humanDocument how to request it and what to include
Tax ID not shown on invoice12Product issueReport to product with ticket links

Five rows, five owners, and a clear first sprint of docs work. That is the deliverable.

Step 8: Close the loop​

Articles only reduce tickets if they reach customers at the moment of need. Update the support macros for each cluster to link the new article, add the articles to the help widget or in-product help where the problem occurs, and tag new tickets in each cluster so you can compare the next quarter against this one. Then repeat the mining on a fresh window. The second pass is faster and usually surfaces the next layer of gaps.

The same clusters are also the best test set for a support bot, which we cover in why your support bot hallucinates. The handbook page on building a content inventory explains how to line the clusters up against what already exists.

Get it done for you​

If you would rather hand this off, our support-ticket knowledge base service runs this method on Zendesk and Zoho Desk exports, from $1,800 per 500 tickets, and delivers the cluster report plus the written articles. Send us your help desk, rough ticket volume and time window through the contact page. We reply within 1 business day.