Skip to content

AI Page Discovery — I/O Definition

Overview

Discovers public pages from a homepage or reads URLs directly from a sitemap. Discovery does not import pages or generate Figma layers. Users select discovered URLs from a path tree before creating a Page Import batch.

Input

Field Type Required Description
discovery_id string Yes UUID job ID
organization_id integer Yes Organization ID
project_id integer Yes Project ID
nonce string Yes Job identity fence
attempt integer Yes 1
url string Yes Public HTTP(S) homepage or .xml sitemap URL
max_pages integer No 1–1,000; default 500
max_depth integer No 0–5; default 3
include_paths string[] No Path prefixes included in results
exclude_paths string[] No Path prefixes excluded from crawling
group_id string Yes SHA-256 of the origin; FIFO message group

Output

S3 checkpoint and preview are stored under {organization_id}/{project_id}/page-discovery/{discovery_id}/. The preview contains status (0 processing, 1 completed, 2 failed), pages (url, title, depth), groups, warnings, truncated, and updated_at. Each group contains an ID, label, representative URL, member URLs, confidence, and grouping reason. It is returned through the authenticated backend, never read directly by the plugin. Shared batch and batch_item rows track submission and dispatch; no additional batch tables or DocumentDB collections are required.

Processing

  1. Validate public origin, options, and job identity.
  2. For a homepage, read robots.txt and discover same-origin links. For a .xml sitemap, parse sitemap and sitemap-index URLs without requesting the listed pages.
  3. Reject private addresses, credentials, unsafe ports, external redirects, action links, and non-HTML downloads. Pin connections to validated public IP addresses.
  4. Remove fragments and tracking parameters; preserve meaningful query parameters.
  5. Process a bounded chunk, persist the frontier and visited URLs, then enqueue continuation if needed. FIFO grouping serializes work per origin; wait at least one second between requests and respect longer robots crawl delays.
  6. Group pages by their exact parent URL path. Page HTML and AI are not used for grouping.
  7. Publish a selectable preview. The plugin displays URLs as a nested path tree and does not preselect pages.

Runtime Configuration

Variable Description
RESULT_BUCKET S3 checkpoint and preview bucket
DISCOVERY_QUEUE_URL FIFO queue used for continuations
MAX_CHUNK_PAGES Maximum tasks per invocation; default 10

Limitations

  • No JavaScript navigation, authentication, form submission, or asset downloads.
  • Same origin only, including redirects and sitemap locations.
  • Maximum 1,000 discovered URLs, 1,000 queued candidates, and 20 sitemap documents per job. Page Import remains limited to 500 selected URLs per batch.
  • Homepage discovery fails closed when robots.txt cannot be read. Direct sitemap input does not request robots.txt or the listed pages.