Some websites answer AI crawlers with HTTP 200 and an empty page. The agent believes it
read the article, and reports on something it never read.
Botvue fetches one URL as a browser, as Googlebot, and as each major AI crawler, then compares what came back.
Live: botvue.onrender.com — check any URL yourself. Full archive: botvue.onrender.com/archive — all 126 properties this scan measured, not just the examples below.
flowchart LR
A1["Agent reads /openapi.json<br/>to find the check operation"] --> A2["POST /check, no payment"]
A2 --> G2["402 — price, payTo,<br/>Blocky402's fee-payer account"]
G2 --> A3["Agent signs a Hedera transfer<br/>as the paying account only"]
A3 --> G3["POST /check again,<br/>X-PAYMENT: signed transaction"]
G3 --> G4["Blocky402 /verify"]
G4 --> G5["Blocky402 /settle<br/>broadcasts, pays the network fee"]
G5 --> G6["Cross-check: independent lookup<br/>on Hedera's public mirror node"]
G6 --> G7["Fetch the URL as a browser,<br/>Googlebot, and four AI crawlers"]
G7 --> G8["Compare: word ratio,<br/>status codes, sha256 of each body"]
G8 --> G9["Grade: soft-blocked / substituted /<br/>refused / clean / ..."]
G9 --> R1["Write finding + hashes to a<br/>Hedera Consensus Service topic<br/>(no submit key)"]
G9 --> Resp["200 — verdict returned to the agent"]
R1 --> R2["Public mirror node:<br/>anyone re-verifies, no credentials"]
R2 -.->|"npm run verify -- <topic> --live"| V["/archive and / re-check the live<br/>page against what was recorded"]
Two things settle independently and are compared, not trusted on their own word: Blocky402's
/settle response against the public mirror node for a payment, and the recorded evidence
hash against a fresh fetch for a finding.
Sixty-four news properties gate AI crawlers through the same vendor, which returns the same 161-byte message every time — "You are not authorized to access this content without a valid TollBit Token."
| Status it is sent with | Responses |
|---|---|
402 Payment Required |
126 |
200 OK |
56 |
Identical bytes. Identical vendor. The status code is a setting.
And in one case the two behaviours sit inside the same company:
| Publisher group | Properties | AI crawlers receive | Status |
|---|---|---|---|
| Hearst newspapers | 22 | an empty challenge page | 200 |
| Lee Enterprises | 9–14 | a "not authorized" message | 200 |
| Gannett | 24 | a refusal | 402 |
| Advance Local | 11 | a refusal | 403 |
| Hearst magazines | 19 | the article | 200 |
| McClatchy | 15 | the article | 200 |
Googlebot receives the real article on every soft-blocking property. Existing cloaking checkers compare Googlebot against a browser, so all of them come back clean.
Hearst is exact and repeatable. Lee moves between runs for a reason described under Limitations.
No accounts, no keys, nothing to sign up for:
python -m venv .venv && .venv/Scripts/python -m pip install -r requirements.txt
.venv/Scripts/python -m scanner.probe https://www.houstonchronicle.com/agent st bytes blocks new gone bps sha
claudebot 402 — — — — —
gptbot 200 3,036 1 1 6 10000 32ed63159c77
oai-searchbot 200 1,099,486 6 0 0 0 218a146ccfcf (identical)
perplexitybot 200 3,036 1 1 6 10000 32ed63159c77
googlebot 200 1,099,486 6 0 0 0 218a146ccfcf (identical)
>> Googlebot gets the human version while gptbot, perplexitybot do not.
>> Every cloaking checker that compares Googlebot to a browser returns clean here.
Your run will show a different hash and byte count on the Googlebot and browser rows — the
homepage changes constantly. 32ed63159c77 on the GPTBot row will be the same as above.
Or run the service and open http://localhost:8000:
.venv/Scripts/python -m uvicorn scanner.gateway:app --port 8000Every number above is a hash of a response body anyone can fetch again.
curl -s -H 'User-Agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.1; +https://openai.com/gptbot)' \
https://www.houstonchronicle.com/ | sha256sumThat still returns, byte for byte, what was recorded on 6 September:
32ed63159c77e21ee19ca1b9aa3213ccf0218eb59539560b132a8e68ef0e18ea 3,036 bytes
Now run the same command as Googlebot:
curl -s -H 'User-Agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Googlebot/2.1; +http://www.google.com/bot.html)' \
https://www.houstonchronicle.com/ | wc -cYou will get roughly 1.1 MB, and a different hash from the recorded one every time — because the homepage is alive and its headlines rotate.
That contrast is the point. The page a person gets keeps changing. The page GPTBot gets has not moved in hours. Transient failures do not hold still; configured responses do.
Two bodies were served byte-identically across many properties, which is what makes this one configuration rather than many coincidences:
| sha256 | Bytes | Served identically on |
|---|---|---|
32ed63159c77… |
3,036 | 23 properties |
ad3a87823990… |
161 | 14 properties |
Both are in evidence/stubs/, verbatim.
evidence/manifest.json has all 126 properties with the status,
byte length, word count and hash of every response — browse it filtered by finding at
botvue.onrender.com/archive, or pull the same data
as JSON from /evidence/all.
If a publisher has since changed configuration, the hashes will stop matching. That is why they are written down.
Findings are anchored to a Hedera consensus topic, because these responses carry
cache-control: no-store — nothing is archived, and a configuration can change in an
afternoon.
Topic 0.0.10395053
— open that link. It is Hedera's public mirror node, not us.
cd packages/chain && npm run verify -- 0.0.10395053 --liveverify uses no credentials. It reads the public mirror node and re-fetches pages from the
publishers directly, then compares each recorded hash against the site as it is now.
That comparison produced the most useful result in the project:
statesman.com soft-blocked
live gptbot unchanged — still serving the recorded body
live perplexitybot unchanged — still serving the recorded body
live googlebot CHANGED — now 7b3308e38330
The stub is byte-stable; the real page is not. Hours later the crawler still receives the identical recorded body while the human page has moved on. Transient failures do not behave that way. A configured response does.
The topic has no submit key — anyone can append to it, including a publisher who disputes a finding. A record only this project could write to would prove only that this project wrote something down.
verify above answers "is the recorded hash still what this site serves." A different
question is whether the decision changed at all — did a soft-blocked property start
answering honestly, or start soft-blocking crawlers it used to let through cleanly.
scanner/rescan.py answers that one: a scheduled job
(.github/workflows/rescan.yml, daily) re-fetches every
property in the corpus live, rebuilds the manifest, and diffs the new decision against the
old one for every domain. A same-decision change — a different verdict label that still
resolves to pass either side — is filtered out; it is not something a caller needs to
act on. What is left is written to evidence/changelog.json and served at
/evidence/changelog, and the homepage
shows it directly under the consensus-topic facts, with how long ago the archive was last
re-observed next to it. The manifest a reader is looking at was never a photograph from one
afternoon; it says how old it is, and what has moved since.
An agent asks before it trusts a page.
curl -X POST localhost:8000/check -H 'content-type: application/json' \
-d '{"url":"https://www.houstonchronicle.com/"}'{
"decision": "block",
"verdict": "soft-blocked",
"explanation": "This page returned HTTP 200 to AI crawlers with almost none of its content...",
"word_ratio": {"gptbot": 0.031, "googlebot": 1.0},
"evidence": {"gptbot": "32ed63159c77…", "googlebot": "<changes between fetches>"}
}refused returns pass, not block. A 4xx already tells the agent it got nothing, so
there is nothing to protect it from. Only responses that look like success get a decision.
That distinction is the whole project: this measures deception, not crawler blocking.
The service is x402-gated and settles through the Blocky402 facilitator on Hedera
testnet. The caller signs a transfer as the paying account and hands over the frozen,
unsubmitted transaction bytes — Blocky402 names itself as fee payer, adds its own signature,
and broadcasts. The server never holds a private key, and a caller never has to trust the
server's own word that payment settled: Blocky402's /settle response is checked a second
time against Hedera's public mirror node, and a disagreement is recorded, not hidden — see
CrossCheck in scanner/payments.py.
And it answers 402 when it means "pay me" — which is the thing 56 of those news responses
do not do.
/check is the agent API and takes payment. The page at
botvue.onrender.com runs the same scan through a free path
capped at 15 checks an hour per visitor, because a reader who has not signed a Hedera
transfer should still be able to see what the tool does. That path is deliberately absent
from /openapi.json — that document is what a Bazantic gateway wraps and prices, and a
free twin of the paid operation listed beside it would make the gate decorative. Anonymous
checks are not written to the consensus topic either; that record carries findings this
project stands behind, not every URL a visitor happened to type.
| Consensus topic | 0.0.10395053 |
| Service payee | 0.0.10395128 |
| Facilitator | Blocky402 (api.testnet.blocky402.com) |
| Network | Hedera testnet |
Full flow — discovery, payment, facilitator settlement, independent cross-check:
cd packages/agent && npx tsx src/run.ts https://www.houstonchronicle.com/ \
--service https://botvue.onrender.comThe agent reads the OpenAPI document to find the endpoint and price, and the 402 body for the facilitator's fee-payer account. Nothing is hardcoded — a stale copy of any of those would be a silent-wrong-network failure, which is exactly the failure class this project spends the rest of its time measuring in other people's services.
The same OpenAPI document is also what a Bazantic gateway wraps — Botvue's own x402 logic stays inside this repo; Bazantic builds an MCP server and a second payment rail on top of it rather than being handed a pre-built gate.
gateway.py's own docstring makes a claim: "a second adapter — an MCP server, say — reuses
the same code rather than reimplementing it." scanner/mcp_server.py
is that claim, made true rather than left as a comment. It wraps the exact same
scanner.service.check the HTTP gateway calls — same fetch, same diff, same classifier, same
verdict — behind three MCP tools instead of a REST route:
.venv/Scripts/python -m pip install -r requirements-mcp.txt
.venv/Scripts/python -m scanner.mcp_servercheck_url(url, fresh=false) — the check itself: decision, verdict, explanation, per-crawler
word ratios, a SHA-256 of every response body, and — when the
verdict rests on specific text — the sentences themselves,
quoted, not described.
fetch_url(url) — a fetch-tool replacement: returns page content only when it
matches what a browser sees, withholding it with a reason
otherwise. See "A third adapter" below.
list_agents() — the exact user-agent strings sent, so a result can be
reproduced with a plain curl rather than taken on trust.
Point any MCP client at it over stdio — Claude Desktop, mcp dev, an agent framework's own
tool loader. Free and unattested: an MCP call carries no x402 payment leg the way /check
does, and writing every anonymous query to the consensus topic would turn a demonstration
tool into unbounded, uncontrolled evidence — the same reasoning behind the web page's own
free /check/preview path, which this mirrors. This is a separate process from the HTTP
gateway (mcp is not a dependency of gateway.py, and is not installed in the deployed
service), so it does not add a stdio server to a web dyno that has no business running one.
Every fetch-tool MCP server — agentfetch, web-retrieval-mcp, the built-in WebFetch —
competes on reliability: did the page load, was it parsed cleanly. None of them ask whether
what loaded is what a human would see.
fetch_url in scanner/mcp_server.py answers that question first,
then returns content only when it's clean:
fetch_url("https://www.houstonchronicle.com/")
# → {"status": "blocked", "reason": "...", "content": None, "warning": "..."}
fetch_url("https://www.elle.com/")
# → {"status": "clean", "content": "<the actual article text>", "warning": None}Point any MCP client's fetch tool at this instead of a generic one, and it silently gets the
same authenticity check every other verdict in this project goes through: content is withheld,
not silently handed back, on the same block/flag/pass decision the web page's own checker
uses.
A finding here is only as good as the vocabulary it's reported in. scanner/classify.py
sorts crawler-only content into five severities — an ordinal, not a flag — and that
ordering is what decides whether a difference gets reported at all: grade.py only marks a
verdict when the worst block reaches PROMOTIONAL or above, precisely because "the text
differs" and "the text differs and carries a promotional or instructional marker" are
different claims, and only the second one holds up.
That taxonomy is published as a JSON Schema — spec/classification.schema.json,
served live at /spec/classification.schema.json
— so a different project measuring the same failure class (an AI crawler receiving content,
or an instruction, a browser does not) can adopt the same five levels instead of inventing
its own. One of them, POLICY_VIOLATION, is documented as reserved and currently unused —
no rule in the classifier assigns it yet. It stays in the spec anyway: dropping an
unimplemented level to make the taxonomy look more finished than it is would be the exact
gap between declared and actual this project reports in other people's services. Full
account of what's live, what isn't, and how to version against it: spec/README.md.
Detecting added content is unreliable, and it does not drive decisions. Measured
against hand-checked samples, that test is right about 58% of the time — it mistakes
navigation rails and page metadata for new content. So a substituted verdict requires an
explicit promotional marker, an instruction aimed at the reader, or a response header the
browser never receives. Text alone gets crawler-only-text, which is reported and never
acted on.
That costs real findings. docs.frends.com serves crawlers genuine machine-only
instructions, and it is not counted, because it carries no marker distinguishing it from
Forbes' navigation rails. Losing recall to keep precision is the deliberate trade: wrongly
blocking a page would end this tool's credibility.
Soft-block detection needs no classifier, which is why it leads. It is a ratio of extracted words. Across 1,483 crawler responses the distribution is bimodal — 6 below 5%, 1,477 above 60%, and nothing in between — so there is no threshold to argue about.
Seven properties flip verdict between runs, and all of them are Lee Enterprises. The verdict needs Googlebot to have received the human page, and Lee's homepages rotate enough that the control drifts across the 0.93 threshold — their control similarities sit at 0.847 to 1.0, right on the boundary. So Lee's soft-block count moves between roughly 9 and 14 depending on the minute you run it.
This does not touch Hearst, whose control sits at 1.0 on every property and does not move.
And it does not affect whether Lee is soft-blocking, only whether this test can say so from
the comparison: Lee's response is a 200 whose body reads "You are not authorized to
access this content", which is self-evident without any control at all.
Intent is never asserted. A response can be a policy, a vendor default, or a mistake.
Nothing here distinguishes them, and the harm does not depend on which it is: an agent
receiving 200 cannot tell any of them from success.
A Graph cross-check for the Bazantic/Base payment leg was attempted and dropped, not
shipped half-working. The same "two independent sources must agree" pattern used for
Hedera settlement — CrossCheck in scanner/payments.py — was tried
against Base: confirm a Bazantic USDC settlement independently, via an indexed subgraph,
the way the mirror node confirms Blocky402. Three specific dead ends, in order. The only
purpose-built base-usdc subgraph on The Graph Network carries zero curation signal, so no
indexer serves it ("subgraph not found: no allocations"). The Graph's own x402 query
gateway is real and was called live — gateway.thegraph.com/api/x402/subgraphs/id/...
returns a genuine 402 with Base's real USDC contract as the asset — but its challenge
arrives in a payment-required header, not the JSON body every x402 client in this project
expects, and settling it would in any case need signing with a wallet key this project's
custodial payer never holds. The Token API, the product actually shaped for "any wallet's
transfers," authenticates through a separate Pinax key unrelated to a Subgraph Studio key.
None of these is a bug in this project; they are the shape of generic ERC-20 lookups on The
Graph right now. Asserting a cross-check that whoever reads this could not reproduce would
be exactly the failure class this project spends its time finding in other people's
services.
Re-verified 6 September, ~1 hour after the first scan. The numbers that carry the
argument were byte-identical on both runs: 182 TollBit responses split 126 × 402 and
56 × 200, and 46 Hearst stub responses. Re-run
.venv/Scripts/python -m scanner.scan --chains --fresh to check it yourself.
Corrections welcome. If a record is wrong, open an issue and it will be annotated rather than removed.
scanner/ fetch matrix → normalise → diff → classify → grade
probe.py one URL, printed
scan.py fetch a corpus grade.py verdicts, from cache
archive.py freeze evidence eval.py measure our own precision
payments.py x402 gate, settled through Blocky402, cross-checked on the mirror node
gateway.py the FastAPI adapter — transport and payment only, no scan logic
apps/web/ index.html the public page and URL checker
archive.html every property in the frozen scan, filterable and searchable
packages/
chain/ Hedera consensus attestation, topic creation, mirror-node verify
agent/ a demo agent that discovers, pays through Blocky402, and verifies
evidence/ frozen observations with hashes, and the stub bodies verbatim
spec/ the classification taxonomy as a JSON Schema, for another project to adopt
Every response is cached to disk, so re-grading the whole corpus costs no requests — which is what made recovering from a bad run cheap.
Setup, including Hedera credentials: SETUP.md
Objections and the evidence that answers them: DEFENCE.md
MIT