Turning support tickets into a knowledge base
Your support queue is the most honest documentation backlog you will ever get. Every ticket is a customer telling you, in their own words, what the product and the docs failed to explain. The trouble is that a queue of several thousand tickets is not a backlog yet. This is the method we use to turn it into one: a ranked list of articles to write, fix or retire, each backed by real tickets and written in the words customers actually use.
The method works the same on Zendesk and Zoho Desk exports, and on most other help desks that let you get tickets and their conversation threads out.
Step 1: Decide the question before you export
Ticket mining can answer several different questions, and mixing them produces a muddy result. We frame it as one question: which tickets could the customer have resolved alone with the right article, and what would that article need to say?
That framing excludes, on purpose, tickets that need a human: refunds, account-specific investigations, bugs that need engineering, and anything involving judgment about a contract. Those still matter, but they belong in a different report for a different owner.
Pick a time window long enough to cover a full release cycle and any seasonal patterns, such as annual renewals or tax season. Six to twelve months is typical.
Step 2: Export tickets and threads
The subject line alone is not enough. You need the customer's first message and the agent reply that resolved the ticket, because the resolution is the first draft of the article.
On Zendesk, the cursor-based incremental export returns tickets changed since a start time, including the description field, which holds the first comment:
curl -s -u "$ZENDESK_EMAIL/token:$ZENDESK_TOKEN" \
"https://example.zendesk.com/api/v2/incremental/tickets/cursor.json?start_time=1735689600" \
> tickets-001.json
# Follow after_url in each response until end_of_stream is true.
Full conversation threads come from each ticket's comments endpoint. A quick flattening pass gives you a table to work from:
jq -r '.tickets[]
| [.id, .created_at, .status, .via.channel, (.tags | join(";")),
((.subject // "") | gsub("[\r\n]+"; " "))]
| @csv' tickets-*.json > tickets.csv
Zoho Desk exposes tickets and their conversation threads as separate API resources in the same way, so the shape of the job is identical. Both tools also offer exports from the admin interface; what those include varies by plan and role, so check before you promise anyone a date.
Step 3: Scrub personal data first
Tickets are full of names, email addresses, phone numbers, account IDs, billing details and occasionally things customers should never have pasted into a support form. Scrub them before anyone reads the data for analysis and before any text goes near an AI tool.
Our scrubbing pass replaces emails, phone numbers, URLs with account identifiers and anything that looks like a key or token with typed placeholders, such as [EMAIL] or [ACCOUNT_ID]. Names are harder to catch automatically, so we also replace the requester's and agent's names using the ticket's own metadata. Keep the raw export somewhere access-controlled, and delete it when the project ends.
Step 4: Filter out the noise
Remove tickets that are not support conversations: spam, auto-replies, out-of-office bounces, sales inquiries routed to the wrong queue, internal test tickets, and tickets merged into others. Keep a count of what you removed and why. When someone later asks why the analysis covers fewer tickets than the help desk dashboard shows, you want a one-line answer.
Step 5: Cluster by intent, not by tags
Agent-applied tags and ticket form fields are the obvious grouping, and they are usually wrong in useful ways. Tags reflect how the support team routes work, not what the customer was trying to do, and they drift as agents come and go.
So we cluster by intent: what was the customer trying to accomplish, and where did they get stuck? The process we use:
- Read a random sample of a few hundred tickets and write a short intent label for each, in the customer's terms: "can't find last month's invoice", not "billing".
- Merge labels that mean the same thing into a draft taxonomy.
- Assign every remaining ticket to one intent, adding new ones only when a ticket genuinely fits nothing.
- Review the largest clusters by reading a handful of tickets from each, to catch clusters that are really two problems.
AI tools can help with the first-pass assignment on large volumes, but we treat their output as a draft and review each cluster by reading actual tickets. A cluster name that sounds right and groups the wrong tickets is worse than no cluster at all.
Step 6: Classify each cluster by what fixes it
This is where ticket mining turns into a plan. Every cluster gets exactly one of these labels:
| Label | What it means | Owner | Fix |
|---|---|---|---|
| Content gap | No article covers this | Docs | Write a new article |
| Findability gap | An article exists, customers do not find it | Docs | Retitle in customer words, add synonyms, link from where they get stuck |
| Content defect | An article exists and is wrong or outdated | Docs and product | Correct it; retire the stale version |
| Product issue | The docs cannot fix a confusing UI or a bug | Product | Report with ticket evidence |
| Needs a human | Account-specific or policy decision | Support | Document the escalation path, not an answer |
The findability gap is the one teams underestimate. It is common to find a cluster of tickets about something the help center already explains, under a title nobody would search for. Those are the cheapest wins in the whole project. Our content audit post covers how to classify the existing articles you will be comparing against.
Step 7: Rank and write
Rank clusters by volume, by how completely an article could resolve them, and by effort. Then write, following a few rules:
- Start from the resolution. The agent reply that closed the ticket is usually the right answer, tested on a real customer. Turn it into steps.
- Title in the customer's words. Use the phrasing from the tickets, not the feature name from the product spec.
- One problem per article. A support agent should be able to paste a single link that answers the question completely.
- Use a template. Troubleshooting articles need symptom, cause, fix and "still stuck?" sections. The handbook page on content types and templates has the templates we use.
A worked example
What follows is a worked example, not a client result. The product, the clusters and the counts are invented to show the shape of the output.
Imagine one batch of 500 tickets, after filtering, for a fictional billing product. The five largest clusters might look like this:
| Intent cluster | Tickets | Label | Action |
|---|---|---|---|
| Can't find last month's invoice | 41 | Findability gap | Retitle "Billing documents" to "Download an invoice"; link from the billing page |
| Card declined on renewal | 33 | Content gap | New troubleshooting article with decline reasons and next steps |
| Changing the billing email | 27 | Content defect | Article shows the old settings screen; update steps |
| Refund for unused seats | 19 | Needs a human | Document how to request it and what to include |
| Tax ID not shown on invoice | 12 | Product issue | Report to product with ticket links |
Five rows, five owners, and a clear first sprint of docs work. That is the deliverable.
Step 8: Close the loop
Articles only reduce tickets if they reach customers at the moment of need. Update the support macros for each cluster to link the new article, add the articles to the help widget or in-product help where the problem occurs, and tag new tickets in each cluster so you can compare the next quarter against this one. Then repeat the mining on a fresh window. The second pass is faster and usually surfaces the next layer of gaps.
The same clusters are also the best test set for a support bot, which we cover in why your support bot hallucinates. The handbook page on building a content inventory explains how to line the clusters up against what already exists.
Get it done for you
If you would rather hand this off, our support-ticket knowledge base service runs this method on Zendesk and Zoho Desk exports, from $1,800 per 500 tickets, and delivers the cluster report plus the written articles. Send us your help desk, rough ticket volume and time window through the contact page. We reply within 1 business day.