SiteScan
Evidence-based scan of what crawlers and AI agents can fetch, read, and cite from a public URL.
Use it when you need to know how a page reads to crawlers and AI agents, or to prove a change fixed it.
Access
This tool runs behind a sign-in. If you have access, open it; if not, ask and we will say yes or explain why not.
What it does
Fetches a public URL the way a crawler or an AI agent would and reports what it could actually read: robots and sitemap directives, metadata, structured data, readable text, and how much of the page needs a browser to render.
The analysis engine is the product. The CLI is the primary way to run it, the web view is one view over the same report, and callers inside the workspace import the engine rather than going through HTTP. A consumer outside it calls the gated web API and pins the projection, not the whole report. The overview does not infer missing facts; if the page does not state something, the report says so and shows the evidence it used.
Features
- One implementation runs in Node, in Cloudflare Workers, and against on-disk fixtures; fetch and the clock are injected
- Optional browser render (Playwright) and PageSpeed Insights lab and field data, kept as separate sources that are never blended
- Re-fetch as each major crawler to compare what they are served
- JSON and Markdown report output, and a saved-run compare that leads with findings resolved and introduced
- Refuses to diff two runs taken under different conditions, and exits 2 when it does
- --fail-on <severity> for gating in CI
- The web API is gated before anything is fetched: our own pages by same-origin, a rate limit per caller, and a ceiling on scans open at once. A refused call answers 403, 429 or 503 with a JSON error
- A service caller that is not a browser presents a key issued to it instead of satisfying same-origin. Production and preview hold separate keys, and the limits still apply to whoever holds one
- The rate limit counts in a Durable Object rather than in whichever isolate served the request, so spreading load across cold starts does not reset it
Inputs
- Public URL
- Refresh (boolean). Skip the cached report and scan again
- Run options. Browser render, PageSpeed Insights, crawler probes, a check profile
- Two saved runs (JSON). For sitescan compare
Outputs
- Web report. Verdict, coverage cards, scorecard and a plain-language overview
- Scan report (JSON)
- Surface facts projection (JSON). The versioned slice to pin, rather than the whole report
- Report document (Markdown)
- Compare report (Markdown). Findings resolved and introduced between two runs
- Exit code. Non-zero at or above the --fail-on severity, for CI