Chunk the tools sitemap after the probe outgrew it - #55
Merged
Merged
Conversation
The probe took the catalogue from 134 tools to 188,948 across 10,628 listings in a few hours, and the tools sitemap was a single unchunked file written when tool coverage was a few hundred names. It reached 63,495 URLs against a 50,000 limit, 9.6 MB, and three minutes to generate, holding a web process the whole time while loading full rows including article text. - Chunk by listing, 1,000 a file, which is about 19,000 URLs each - Paginate the query and select name, tools and updated_at rather than the row - Drop client pages from the sitemap and mark them noindex That last one is a judgement worth stating: six client pages per tool across 188,948 tools is 1.1 million near-identical URLs. That is the doorway pattern rather than coverage, on the domain whose value is that it ranks. The pages stay linked and useful for a reader who wants "how do I do this in Cursor" -- they simply no longer ask to be ranked. Tool pages stay indexed, and each of those is a real tool on a real server. --- Pages affected: - [MCP Registry](https://ai.mcpharbor.dev/) — the Model Context Protocol server directory. - [Sitemap](https://ai.mcpharbor.dev/sitemap.xml) — now lists chunked tool files. - [Browse MCP servers](https://ai.mcpharbor.dev/servers) — the catalogue behind it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The probe took the catalogue from 134 tools to 188,948 across 10,628 listings
in a few hours. The tools sitemap was a single unchunked file written when tool
coverage was a few hundred names: it reached 63,495 URLs against a 50,000
limit, 9.6 MB, and three minutes to generate, loading full rows including
article text to do it.
- Chunk by listing, 150 a file, which is about 19,000 URLs each
- Paginate the query and select a partial struct rather than the row
- Add Clients.ids/1 so the sitemap can name the six client pages without
JSON-encoding sixty thousand configs to throw them away
Client pages stay indexed and listed. I had made them noindex on a
near-duplicate argument; that was wrong. A VS Code page emits "servers" with
${input:} prompts, Zed emits context_servers, Claude Desktop bridges remote
servers through mcp-remote -- different format, path and caveats each. The only
noindex left is the pre-existing rule for pending and deprecated listings.
---
Pages affected:
- [MCP Registry](https://ai.mcpharbor.dev/) — the Model Context Protocol server directory.
- [Sitemap](https://ai.mcpharbor.dev/sitemap.xml) — now lists chunked tool files.
- [Browse MCP servers](https://ai.mcpharbor.dev/servers) — the catalogue behind it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #54.
The probe took the catalogue from 134 tools to 188,948 across 10,628 listings in a few hours. The tools sitemap was written as a single unchunked file back when tool coverage was a few hundred names, and that assumption is now false:
It also loaded full
Serverrows —article_contentincluded — for every listing just to build URLs, which is where the three minutes went.Changes
Clients.ids/1, so the sitemap can name the six client pages without JSON-encoding 60,000 configs to throw them awayOn indexing
An earlier revision of this branch put
noindexon the client pages, reasoning that 1.1 million of them was a doorway pattern. Logan overruled that, and on inspection he is right: a VS Code page emitsserverswith${input:}prompts, Zed emitscontext_servers, Claude Desktop bridges remote servers viamcp-remote— different config format, different path, different caveats. That is distinct content, not boilerplate, so all of it is indexed and listed.The only
noindexleft in the silo is the pre-existing rule for pending and deprecated listings, and there is now a test pinning both halves: every live page indexable, pending ones not.135 tests pass.
Pages affected:
🤖 Generated with Claude Code