Skip to content

feat(import): add a VitePress importer - #42

Merged
btopro merged 1 commit into
haxtheweb:mainfrom
SanikaA3:feat/2923-vitepress-importer
Sep 25, 2026
Merged

btopro merged 1 commit into
haxtheweb:mainfrom
SanikaA3:feat/2923-vitepress-importer

Conversation

@SanikaA3

Copy link
Copy Markdown
Contributor

Refs haxtheweb/issues#2923, implementing the plan in this comment. This is step 1 of 3: the converter, the dispatcher case and the OpenAPI entry. The @system/vitepressToSite micro-frontend registration (webcomponents) and the CLI wiring (create) follow as their own PRs, the same split as #40.

curl -X POST http://localhost:8080/system/api/v1/site/import/vitepress \
  -H 'content-type: application/json' \
  -d '{"repoUrl": "https://github.com/dmd-program/dmd-100-book"}'

What it does

src/systemRoutes/v1/routes/imports/convertVitepressToSite.js takes a VitePress repository URL and returns { items, filename, files, site, truncated, unmappedComponents }, the shape the other converters in the directory return. (The plan lists siteFiles in the return shape; no converter here emits one, so this follows the code.)

The outline comes out of the config, not a SUMMARY.md

The repository and its default branch are resolved through the GitHub API and the file tree is read once. .vitepress/config.{mjs,mts,js,ts} is located in that tree, preferring a config under docs/ over one at the repository root, then themeConfig.sidebar, base, title, description, license, defaultAuthor and workTitle are read out of it.

The config is parsed, never executed: a small literal reader walks the object and returns nothing for anything that is not plain data — a spread, a function call, an import, a template literal. Sidebar groups become indented children, link: '/' maps to index.md. When the config cannot be read, or parses but names no pages, the outline falls back to the markdown file tree the way VitePress's own auto-sidebar does.

Pages

Each page is rendered locally with markdown-it plus markdown-it-container and markdown-it-footnote — the two plugins the source repository itself uses, added as dependencies (10 lines of lockfile, no other churn). Frontmatter is split off first and supplies title, description, license and author.

VitePress HAX
![alt](x.png) <media-image source="files/x.png" alt="…">
<VideoEmbed src type title caption> <video-player source media-title> with a slot="caption"
::: learning-objective, assessment, practice, learning-component, instructional-pattern <oer-schema typeof oer-property> using the matching vocabulary, skill/forCourse/additionalType/gradingFormat/assessing carried as properties
frontmatter license, else themeConfig.license one <license-element license title creator source> per page
any other Vue component wrapper dropped, inner content kept, name reported in unmappedComponents

Every item carries metadata.vitepress (repo, branch, path, license, author, workTitle, accessed) and a metadata.source pointing at the file on GitHub, the way the OpenStax import carries metadata.openstax.

Assets and links

Media referenced by pages is collected into build.files as files/<name> → raw URL and createSite stages it. That is only possible because of #41 (haxtheweb/issues#3060): until that landed createSite rejected URL-valued build.files, which is why the OpenStax importer downloads and stages every image itself. This importer hands over URLs and lets createSite fetch them through its SSRF-guarded staging path, so an import that is never turned into a site writes nothing to disk.

Both VitePress reference styles are resolved — /assets/… and ../assets/… against the docs root and the page that contains them, and docs/public/x.png referenced from the site root as /x.png. Names are de-duplicated (x.png, x-1.png) and limited to the extensions the bulk import accepts. Links between pages are resolved against the containing page and rewritten to the imported slug, and the configured base (/dmd-100-book/) is stripped, since HAX manages its own.

The plan's open questions, as implemented

Q2 config parsing Static parse plus the filesystem fallback, the recommended option. The config is never executed and the repository's dependencies are never installed.
Q3 video One <video-player> per embed. Not media-playlist: praw pairs media-playlist with audio-player as the audio pattern. Local .webm/.mp4 are collected into build.files and played from files/; YouTube and Vimeo sources stay remote, which video-player already handles.
Q4 footnotes markdown-it-footnote's own output — numbered references linking to an end-of-page list with back-links. There is no footnote element in webcomponents to map to, and this output is plain editable HTML.
Q5 license Both: site.license for the site and a <license-element> closing each page with its title, creator and source. Frontmatter wins over themeConfig, and cc-by-sa style codes are normalized to the by-sa codes license-element takes.
Q6 theme Left alone. No converter picks a theme; that belongs to the site-creation flow.
Q7, Q9 scope Import only — no round-trip, no embed mode, no Common Cartridge.
Q8 other components Stripped with their content kept, and each name comes back in unmappedComponents, so "log unmapped constructs" is visible to the caller instead of buried in a log.

Q1 (how the CLI reaches the endpoint) is answered in step 3, the same way as OpenStax: option (a), @system/vitepressToSite resolved through MicroFrontendRegistryConfig.base.

The GitHub API needs a User-Agent

The first live run got a 403 on every call. GitHub rejects API requests without a User-Agent; convertGitbookToSite documents this and sends HAXcms-Import/1.0, so this importer sends the same header and a test pins it.

Worth flagging separately: convertNotionToSite sends no User-Agent on its GitHub calls, so it cannot be reading anything from that API today. Not touched here.

Safety valves

500 pages, 2,000 files, a 900 second budget and 100 ms between requests, in one exported LIMITS object, matching the OpenStax importer. When a limit bites the response carries truncated: true and the pages that were not fetched link to their source instead of arriving empty.

Testing

test/unit/convertVitepressToSite.test.cjs adds 43 tests with the network stubbed at safeFetch against an in-memory fixture repository, so the suite never leaves the machine: request validation and every error path, config discovery including the docs/ preference, sidebar nesting, both fallbacks (unparseable config, and a config that parses but names no pages), frontmatter, the license normalization and per-page element, each OER container, footnotes, VideoEmbed local and remote, images and alt text, link rewriting and base stripping, unmapped components, duplicate file names, the User-Agent, and each limit.

  • 43/43 pass. Full unit suite 1315/1315; API conformance 160 passed, 4 skipped, 0 failed.
  • Coverage of the new file: 94.4% lines, 82.3% branches, 100% functions.
  • Eight deliberate mutations, all caught. One was not at first — a config that parsed but named no pages could 422 instead of falling back — which is where that test came from.

Live against dmd-program/dmd-100-book, the repository in the issue:

  • Import: 127 items, 56 files, 86 media-image, 13 video-player, 104 license-element, 46 footnote references; no :::, VideoEmbed or {{ left in the output; 0 orphaned items, 0 duplicate slugs, 0 file keys createSite would reject.
  • Through a real createSite over HTTP: 200, with 56 requested = 56 saved = 56 file entities, 51 of 51 pages with references linked to their entities, 0 references without a file on disk, and the staging directory empty afterwards.

🤖 Generated with Claude Code

Adds convertVitepressToSite and registers it with the
/system/api/v1/site/import/:platform dispatcher as the "vitepress"
platform, so a VitePress documentation repository can be imported into a
HAXcms site (issue #2923).

How the import works

- Resolves the repository and its default branch through the GitHub API,
  then reads the file tree once and works from that listing.
- Finds .vitepress/config.{mjs,mts,js,ts}, preferring a config under docs/
  over one at the repository root, and reads themeConfig.sidebar, base and
  title from it. The config is parsed, never executed: a small literal
  reader walks the object and gives up on anything that is not plain data.
- Builds the outline from the sidebar, nesting sidebar groups as indented
  children. When the config cannot be read, or names no pages, the outline
  falls back to the markdown file tree.
- Renders each page with markdown-it plus the container and footnote
  plugins VitePress itself uses, so ::: containers and footnotes survive.
  Frontmatter supplies the title, description and license.

HAX element mapping

- images                  -> media-image (source, alt)
- <VideoEmbed>            -> video-player (source, media-title, caption)
- ::: learning-objective, assessment, practice, learning-component and
  instructional-pattern containers -> oer-schema (typeof, oer-property)
- a per page license-element from frontmatter or the site license

Assets and links

- Assets referenced by pages are collected into build.files as raw URLs for
  createSite to stage. Names are de-duplicated, and only the extensions
  createSite accepts are collected.
- Relative links are resolved against the page that contains them and
  rewritten to the imported slug, with the configured base prefix stripped.

GitHub API requests send a User-Agent, without which GitHub answers 403 to
every call; this follows the convention already documented in
convertGitbookToSite. Limits mirror the OpenStax importer: 500 pages, 2000
files, a 900 second fetch budget and a 100ms delay between requests.

Tests

test/unit/convertVitepressToSite.test.cjs adds 43 unit tests covering
outline building and both fallbacks, the page pipeline, asset collection
and link rewriting, the User-Agent, the limits and every error path. The
network is stubbed at safeFetch with an in-memory fixture repository.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@SanikaA3
SanikaA3 requested a review from btopro as a code owner September 24, 2026 20:15
@codesandbox

codesandbox Bot commented Sep 24, 2026

Copy link
Copy Markdown

Review or Edit in CodeSandbox

Open the branch in Web Editor • VS Code • Insiders

Open Preview

@btopro
btopro merged commit e969655 into haxtheweb:main Sep 25, 2026
3 checks passed
@github-actions github-actions Bot locked and limited conversation to collaborators Sep 25, 2026
@btopro

btopro commented Sep 25, 2026

Copy link
Copy Markdown
Member

@SanikaA3 excellent! next up as follow ups to this would be:

  • create accepting this as a param
  • haxcms-php port
  • webcomponents front-end importer code for UI so end users can do it on the dashboard

Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants