Skip to main content

Confluence to Docusaurus: a migration guide

· 8 min read
Documentation studio

Moving a Confluence space to Docusaurus is mostly a conversion problem with three traps in it: macros that have no Markdown equivalent, a page tree that does not map cleanly to folders, and URLs that carry page IDs your redirects have to honor. This guide covers each step in the order we do it, with the commands and config we actually use.

It assumes Confluence Cloud and a Docusaurus 3 target. Data Center works the same way in principle, with different URL patterns, which we note where it matters.

Step 1: Decide what is leaving​

Confluence spaces accumulate meeting notes, retrospectives, drafts and half-finished specs next to the documentation that customers or support agents actually use. Only the second group belongs in Docusaurus.

Before exporting, build an inventory of the space: page ID, title, parent page ID, last updated date, author and whether the page has view restrictions. The restrictions column matters more than it looks. A page that only three people can see in Confluence becomes public the moment it lands in a public Git repository or a public site, so restricted pages need an explicit decision, not a default. The handbook page on building a content inventory describes the columns we use, and the content audit post covers how to decide what each page earns.

Step 2: Pick an export path​

Confluence Cloud gives you three realistic ways out. They are not equivalent.

PathWhat you getGood forWatch out for
Space export to HTMLA zip of HTML pages, with attachments in per-page foldersSmall spaces, one-off movesSpace admin permission needed; blog posts and comments are not included
Single-page exportWord or PDF, one page at a timeA handful of pagesNot a migration path at any real size
REST API, storage formatEach page body as XHTML with Confluence macro elements, plus metadataAnything you may need to rerunMacros arrive as structured elements you must convert yourself

Atlassian also offers PDF, CSV and XML space exports. PDF is for reading, not converting. XML and CSV are designed for importing into another Confluence instance, and we do not use them as conversion sources.

For anything over a few dozen pages we use the API, because a migration takes weeks and the source keeps changing while you work. An HTML export is a snapshot; a script against the API can be rerun the night before cutover. With the v2 API, page bodies are only returned when you ask for them:

# Look up the numeric space ID (v2 uses IDs, not space keys).
curl -s -u "$ATLASSIAN_EMAIL:$ATLASSIAN_TOKEN" \
"https://your-site.atlassian.net/wiki/api/v2/spaces?keys=DOCS" | jq '.results[0].id'

# Fetch pages in that space with their storage-format bodies.
# Follow _links.next in each response until it is absent.
curl -s -u "$ATLASSIAN_EMAIL:$ATLASSIAN_TOKEN" \
"https://your-site.atlassian.net/wiki/api/v2/spaces/$SPACE_ID/pages?body-format=storage&limit=100" \
> pages-001.json

Save the raw JSON to disk before converting anything. Conversion scripts change; the raw export should not.

Step 3: Convert macros before you convert HTML​

The storage format is XHTML with Confluence-specific elements in the ac: and ri: namespaces. A generic HTML-to-Markdown converter will either drop those elements or flatten them into plain text, which is how info panels turn into orphaned sentences and code macros lose their language.

So we run conversion in two passes. The first pass rewrites known macros into plain HTML or placeholder markers that survive the second pass. The second pass converts the cleaned HTML to Markdown, for example with pandoc:

pandoc --from=html --to=gfm --wrap=none page-1183.clean.html -o page-1183.md

The macro mapping is where the judgment is. This is the table we start from:

Confluence macroDocusaurus equivalentNotes
Info, Note, Tip, Warning panelsAdmonitions such as :::info, :::note, :::tip, :::warningMap by meaning, not by color
Code blockFenced code with a languageKeep the language parameter; default to text
ExpandA details element with a summaryCheck the content inside still renders
Table of contentsRemoveDocusaurus renders a table of contents from headings
Children display, page treeRemove, use a generated category indexSee step 4
Include page, excerpt includeAn MDX partial imported where it is usedDecide one owner for the shared text
Jira issue linksA plain link, or removeCustomers usually cannot open internal Jira anyway
Status lozengePlain text, or a small custom componentRarely worth a component
User mentionThe team or role namePersonal names go stale fastest
AnchorA heading ID or an explicit anchorKeep IDs that other pages link to

Two Docusaurus details save pain here. First, Docusaurus 3 parses .md files as MDX by default, and MDX treats { and < as syntax, which Confluence content is full of. Setting markdown.format: 'detect' makes .md files use CommonMark and keeps MDX for .mdx files, so converted pages build without escaping every brace. Second, admonitions are native, so panels need no custom components:

:::warning[Token scope]
Tokens created before the permissions change keep their old scope until they are rotated.
:::

Step 4: Rebuild the page tree as folders​

Confluence hierarchy is a parent-child tree with no depth limit and no distinction between a page and a section. Docusaurus uses folders for sections and files for pages, and sidebars are generated from folders or defined explicitly.

The mapping we use: a Confluence page that has children becomes a folder with a category index, and its own content either becomes that folder's index page or a normal page inside it if it is substantial. Flatten anything deeper than three levels; deep trees in Confluence are usually the result of people nesting pages to control permissions or to hide drafts, not a real information structure. The handbook page on navigation models covers how to choose the new top level.

Slugs come from titles, lowercased and hyphenated, but check them by hand. Titles such as "FAQ (new)" or "Copy of Setup" produce slugs you will regret.

Step 5: Attachments and images​

Attachments are referenced in the storage format by filename relative to the page, not by a stable URL. Download them through the API or take them from the HTML export's per-page folders, rename them to lowercase hyphenated names, store them next to the page that uses them, and rewrite the references. Add alt text while you are there; a migration is the cheapest moment to fix it. The handbook page on images and attachments has the naming rules we use.

Large binary attachments, such as installers or long PDFs, usually belong in object storage or a release page rather than in the docs repository.

Confluence links between pages by page title or ID. Rewrite every internal link to the new relative Markdown path using your inventory as the lookup table, then let the Docusaurus build check them: with onBrokenLinks: 'throw', a missed link fails the build instead of reaching readers.

For external traffic, Confluence Cloud page URLs look like /wiki/spaces/DOCS/pages/1183/Authentication. The numeric ID is the stable part; the title segment changes when someone renames the page. Key your redirect map on the ID. Data Center adds older patterns such as /display/DOCS/Authentication and /pages/viewpage.action?pageId=1183, which you will find in bookmarks and support macros for years. The full method, including tests, is in our post on redirect mapping.

If the Confluence space stays alive for internal use, you cannot serve redirects from Atlassian's domain. In that case replace the migrated pages' content with a short notice and a link to the new location, and make the space read-only.

Step 7: Cut over and close the loop​

Freeze edits in the space, rerun the export and conversion scripts once more to pick up late changes, build, run link checks, deploy, then switch the redirect layer. Afterward, set the space to read-only so nobody keeps editing the old copy. Two sources of truth is the most common way a migration quietly unwinds.

When to hand this off​

If the space is small and mostly text, this is a manageable in-house project. It gets expensive when there are hundreds of pages, heavy macro use, several spaces with overlapping content, or a hard cutover date tied to a Confluence renewal. That is the work our Confluence to Docusaurus migration service covers, including the inventory, conversion scripts, redirect map and contributor handover. Tell us roughly how many pages and spaces you have through the contact page. We reply within 1 business day.