Skip to content
Emergent LabsOpen source, emerging technology
← All tools
website-intelligence
Active

Snapshot

Crawls a public site, screenshots every page, and exports the run as labelled evidence.

Use it when you need the before picture of a site ahead of a redesign, a migration or a client onboarding.

Access

This tool runs behind a sign-in. If you have access, open it; if not, ask and we will say yes or explain why not.

Open toolRequest access

What it does

Maps a public site, captures each page at desktop, tablet and mobile widths in one run, builds a copyable architecture tree, and exports the whole run as labelled evidence. It is the before picture for a redesign, a migration, or a client onboarding, and the visual record that a later run can be compared against.

One codebase, two capture engines: local Playwright for day-to-day use and Cloudflare Browser Rendering for the hosted deployment. The crawl result records which engine produced it.

Features

  • Crawl mode, an explicit page list, or a knowledge-base mode
  • Desktop, tablet and mobile in one run, each viewport saved before the next starts so the browser holds one at a time
  • Full-page or viewport screenshots
  • capture-run, a terminal driver over the app's own endpoints, so a CLI run and a form run write identical output; it checks every requested page came back
  • Directory tree view and copyable architecture tree text
  • External dependency detection, from real network requests locally and from rendered HTML on Cloudflare
  • JSON, CSV, full-run ZIP and labelled screenshot ZIP exports, generated statelessly from the current crawl result
  • The whole tool sits behind one password, held as a signed session cookie; the API answers 401 and /crawls redirects without it
  • Runtime chosen in the form. Auto prefers CAPTURE_PROVIDER, else Cloudflare when its credentials are present, else local Playwright; an explicit choice is honoured or refused, never swapped. The crawl result records which engine ran

Inputs

  • Root URL. Crawled to a page and depth limit, or paired with a list of specific pages
  • Crawl limits. Up to 100 pages and 10 levels deep running locally; a hosted crawl is capped at 6 pages and 180 seconds, and says so in the result
  • Viewports. Any of desktop 1440, tablet 768 and mobile 390, captured one after another
  • Capture options. Full page or viewport, query parameters, include and exclude patterns
  • Help Center URL. Knowledge-base mode: a Zendesk Help Center's articles instead of screenshots

Outputs

  • Screenshot gallery. One capture per page, with status and depth; files named from URL and viewport
  • Saved run. Locally, captures/<date>-<host>/ with a manifest and an index.html contact sheet. Hosted, each viewport downloads as a ZIP instead
  • Page tree and site architecture. Copyable as text
  • Third-party dependency map. External domains each page loads from
  • Crawl result (JSON)
  • Page list (CSV)
  • Labelled screenshots (ZIP)
  • Full run (ZIP)

Related

SiteScanObjectReader