This project crawls a website, parses each page with Defuddle CLI using English markdown output, and writes markdown files enriched with Obsidian-friendly frontmatter.
npm installnpm run crawl -- https://example.com/docs --max-depth 2 --output-dir ./output/docsOr directly:
node ./src/cli.js crawl https://example.com/docs --max-depth 2 --output-dir ./output/docsNo global install required (from this repo directory):
npm exec -- wiki-clipper https://example.com --max-depth 5 --output-dir ./output/exampleOr (also from this repo directory):
./src/cli.js https://example.com --max-depth 5 --output-dir ./output/exampleIf you want wiki-clipper available globally without sudo, use a user-level npm prefix:
mkdir -p ~/.npm-global
npm config set prefix ~/.npm-global
export PATH="$HOME/.npm-global/bin:$PATH"
npm linkThen:
wiki-clipper https://example.com --max-depth 5 --output-dir ./output/example(You can also still use: defuddle-kb crawl <url> ...)
- Uses
npx defuddle parse <url> --markdown --json --lang enfor every page. - Fetches HTML separately to extract canonical URLs and internal links.
- Deduplicates pages by canonical URL.
- Generates filenames from URL paths such as
docs-getting-started.md. - Writes YAML frontmatter with
tags,wikilinks,backlinks, andsitemap_path. - Persists crawl progress in
crawl-state.jsonand failures infailures.json.
By default the crawler stays within the start URL path prefix (path scope). This prevents parsing an entire site when the page contains global navigation links.
By default the crawler writes markdown files incrementally as each page is parsed, so you can see output appear while the crawl is still running. At the end it performs a final pass to update wikilinks/backlinks using the full discovered graph.
- Disable incremental writes (only write at the end):
wiki-clipper https://example.com --write-mode final- Default (path scoped):
wiki-clipper https://www.tensorflow.org/guide/data_performance --output-dir ./output/tf- Crawl the whole domain intentionally:
wiki-clipper https://www.tensorflow.org/guide/data_performance --scope domain --output-dir ./output/tf-all- Ignore boundaries entirely (not recommended unless you really want it):
wiki-clipper https://www.tensorflow.org/guide/data_performance --allow-externalBy default the tool prints INFO-level progress logs to stderr.
- Silence logs:
wiki-clipper https://example.com --quiet- More detail:
wiki-clipper https://example.com --log-level debug- Root URLs are written to
index.md. - By default, existing markdown files are overwritten. Use
--no-overwriteto preserve existing files.