Confluence to Docusaurus: a migration guide
Moving a Confluence space to Docusaurus is mostly a conversion problem with three traps in it: macros that have no Markdown equivalent, a page tree that does not map cleanly to folders, and URLs that carry page IDs your redirects have to honor. This guide covers each step in the order we do it, with the commands and config we actually use.
It assumes Confluence Cloud and a Docusaurus 3 target. Data Center works the same way in principle, with different URL patterns, which we note where it matters.
Step 1: Decide what is leaving
Confluence spaces accumulate meeting notes, retrospectives, drafts and half-finished specs next to the documentation that customers or support agents actually use. Only the second group belongs in Docusaurus.
Before exporting, build an inventory of the space: page ID, title, parent page ID, last updated date, author and whether the page has view restrictions. The restrictions column matters more than it looks. A page that only three people can see in Confluence becomes public the moment it lands in a public Git repository or a public site, so restricted pages need an explicit decision, not a default. The handbook page on building a content inventory describes the columns we use, and the content audit post covers how to decide what each page earns.
Step 2: Pick an export path
Confluence Cloud gives you three realistic ways out. They are not equivalent.
| Path | What you get | Good for | Watch out for |
|---|---|---|---|
| Space export to HTML | A zip of HTML pages, with attachments in per-page folders | Small spaces, one-off moves | Space admin permission needed; blog posts and comments are not included |
| Single-page export | Word or PDF, one page at a time | A handful of pages | Not a migration path at any real size |
| REST API, storage format | Each page body as XHTML with Confluence macro elements, plus metadata | Anything you may need to rerun | Macros arrive as structured elements you must convert yourself |
Atlassian also offers PDF, CSV and XML space exports. PDF is for reading, not converting. XML and CSV are designed for importing into another Confluence instance, and we do not use them as conversion sources.
For anything over a few dozen pages we use the API, because a migration takes weeks and the source keeps changing while you work. An HTML export is a snapshot; a script against the API can be rerun the night before cutover. With the v2 API, page bodies are only returned when you ask for them:
# Look up the numeric space ID (v2 uses IDs, not space keys).
curl -s -u "$ATLASSIAN_EMAIL:$ATLASSIAN_TOKEN" \
"https://your-site.atlassian.net/wiki/api/v2/spaces?keys=DOCS" | jq '.results[0].id'
# Fetch pages in that space with their storage-format bodies.
# Follow _links.next in each response until it is absent.
curl -s -u "$ATLASSIAN_EMAIL:$ATLASSIAN_TOKEN" \
"https://your-site.atlassian.net/wiki/api/v2/spaces/$SPACE_ID/pages?body-format=storage&limit=100" \
> pages-001.json
Save the raw JSON to disk before converting anything. Conversion scripts change; the raw export should not.
Step 3: Convert macros before you convert HTML
The storage format is XHTML with Confluence-specific elements in the ac: and ri: namespaces. A generic HTML-to-Markdown converter will either drop those elements or flatten them into plain text, which is how info panels turn into orphaned sentences and code macros lose their language.
So we run conversion in two passes. The first pass rewrites known macros into plain HTML or placeholder markers that survive the second pass. The second pass converts the cleaned HTML to Markdown, for example with pandoc:
pandoc --from=html --to=gfm --wrap=none page-1183.clean.html -o page-1183.md
The macro mapping is where the judgment is. This is the table we start from:
| Confluence macro | Docusaurus equivalent | Notes |
|---|---|---|
| Info, Note, Tip, Warning panels | Admonitions such as :::info, :::note, :::tip, :::warning | Map by meaning, not by color |
| Code block | Fenced code with a language | Keep the language parameter; default to text |
| Expand | A details element with a summary | Check the content inside still renders |
| Table of contents | Remove | Docusaurus renders a table of contents from headings |
| Children display, page tree | Remove, use a generated category index | See step 4 |
| Include page, excerpt include | An MDX partial imported where it is used | Decide one owner for the shared text |
| Jira issue links | A plain link, or remove | Customers usually cannot open internal Jira anyway |
| Status lozenge | Plain text, or a small custom component | Rarely worth a component |
| User mention | The team or role name | Personal names go stale fastest |
| Anchor | A heading ID or an explicit anchor | Keep IDs that other pages link to |
Two Docusaurus details save pain here. First, Docusaurus 3 parses .md files as MDX by default, and MDX treats { and < as syntax, which Confluence content is full of. Setting markdown.format: 'detect' makes .md files use CommonMark and keeps MDX for .mdx files, so converted pages build without escaping every brace. Second, admonitions are native, so panels need no custom components:
:::warning[Token scope]
Tokens created before the permissions change keep their old scope until they are rotated.
:::
Step 4: Rebuild the page tree as folders
Confluence hierarchy is a parent-child tree with no depth limit and no distinction between a page and a section. Docusaurus uses folders for sections and files for pages, and sidebars are generated from folders or defined explicitly.
The mapping we use: a Confluence page that has children becomes a folder with a category index, and its own content either becomes that folder's index page or a normal page inside it if it is substantial. Flatten anything deeper than three levels; deep trees in Confluence are usually the result of people nesting pages to control permissions or to hide drafts, not a real information structure. The handbook page on navigation models covers how to choose the new top level.
Slugs come from titles, lowercased and hyphenated, but check them by hand. Titles such as "FAQ (new)" or "Copy of Setup" produce slugs you will regret.
Step 5: Attachments and images
Attachments are referenced in the storage format by filename relative to the page, not by a stable URL. Download them through the API or take them from the HTML export's per-page folders, rename them to lowercase hyphenated names, store them next to the page that uses them, and rewrite the references. Add alt text while you are there; a migration is the cheapest moment to fix it. The handbook page on images and attachments has the naming rules we use.
Large binary attachments, such as installers or long PDFs, usually belong in object storage or a release page rather than in the docs repository.
Step 6: Internal links and redirects
Confluence links between pages by page title or ID. Rewrite every internal link to the new relative Markdown path using your inventory as the lookup table, then let the Docusaurus build check them: with onBrokenLinks: 'throw', a missed link fails the build instead of reaching readers.
For external traffic, Confluence Cloud page URLs look like /wiki/spaces/DOCS/pages/1183/Authentication. The numeric ID is the stable part; the title segment changes when someone renames the page. Key your redirect map on the ID. Data Center adds older patterns such as /display/DOCS/Authentication and /pages/viewpage.action?pageId=1183, which you will find in bookmarks and support macros for years. The full method, including tests, is in our post on redirect mapping.
If the Confluence space stays alive for internal use, you cannot serve redirects from Atlassian's domain. In that case replace the migrated pages' content with a short notice and a link to the new location, and make the space read-only.
Step 7: Cut over and close the loop
Freeze edits in the space, rerun the export and conversion scripts once more to pick up late changes, build, run link checks, deploy, then switch the redirect layer. Afterward, set the space to read-only so nobody keeps editing the old copy. Two sources of truth is the most common way a migration quietly unwinds.
When to hand this off
If the space is small and mostly text, this is a manageable in-house project. It gets expensive when there are hundreds of pages, heavy macro use, several spaces with overlapping content, or a hard cutover date tied to a Confluence renewal. That is the work our Confluence to Docusaurus migration service covers, including the inventory, conversion scripts, redirect map and contributor handover. Tell us roughly how many pages and spaces you have through the contact page. We reply within 1 business day.