Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

Defuddle Recursive Knowledge Base Crawler

This project crawls a website, parses each page with Defuddle CLI using English markdown output, and writes markdown files enriched with Obsidian-friendly frontmatter.

Install

npm install

Usage

npm run crawl -- https://example.com/docs --max-depth 2 --output-dir ./output/docs

Or directly:

node ./src/cli.js crawl https://example.com/docs --max-depth 2 --output-dir ./output/docs

Run as a terminal command (wiki-clipper)

No global install required (from this repo directory):

npm exec -- wiki-clipper https://example.com --max-depth 5 --output-dir ./output/example

Or (also from this repo directory):

./src/cli.js https://example.com --max-depth 5 --output-dir ./output/example

If you want wiki-clipper available globally without sudo, use a user-level npm prefix:

mkdir -p ~/.npm-global
npm config set prefix ~/.npm-global
export PATH="$HOME/.npm-global/bin:$PATH"
npm link

Then:

wiki-clipper https://example.com --max-depth 5 --output-dir ./output/example

(You can also still use: defuddle-kb crawl <url> ...)

Behavior

  • Uses npx defuddle parse <url> --markdown --json --lang en for every page.
  • Fetches HTML separately to extract canonical URLs and internal links.
  • Deduplicates pages by canonical URL.
  • Generates filenames from URL paths such as docs-getting-started.md.
  • Writes YAML frontmatter with tags, wikilinks, backlinks, and sitemap_path.
  • Persists crawl progress in crawl-state.json and failures in failures.json.

Crawl boundaries (scope)

By default the crawler stays within the start URL path prefix (path scope). This prevents parsing an entire site when the page contains global navigation links.

Write behavior

By default the crawler writes markdown files incrementally as each page is parsed, so you can see output appear while the crawl is still running. At the end it performs a final pass to update wikilinks/backlinks using the full discovered graph.

  • Disable incremental writes (only write at the end):
wiki-clipper https://example.com --write-mode final
  • Default (path scoped):
wiki-clipper https://www.tensorflow.org/guide/data_performance --output-dir ./output/tf
  • Crawl the whole domain intentionally:
wiki-clipper https://www.tensorflow.org/guide/data_performance --scope domain --output-dir ./output/tf-all
  • Ignore boundaries entirely (not recommended unless you really want it):
wiki-clipper https://www.tensorflow.org/guide/data_performance --allow-external

Logging

By default the tool prints INFO-level progress logs to stderr.

  • Silence logs:
wiki-clipper https://example.com --quiet
  • More detail:
wiki-clipper https://example.com --log-level debug

Notes

  • Root URLs are written to index.md.
  • By default, existing markdown files are overwritten. Use --no-overwrite to preserve existing files.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages