AI Page Discovery — I/O Definition
Overview
Discovers public pages from a homepage or reads URLs directly from a sitemap. Discovery does not import pages or generate Figma layers. Users select discovered URLs from a path tree before creating a Page Import batch.
Input
| Field | Type | Required | Description |
|---|---|---|---|
| discovery_id | string | Yes | UUID job ID |
| organization_id | integer | Yes | Organization ID |
| project_id | integer | Yes | Project ID |
| nonce | string | Yes | Job identity fence |
| attempt | integer | Yes | 1 |
| url | string | Yes | Public HTTP(S) homepage or .xml sitemap URL |
| max_pages | integer | No | 1–1,000; default 500 |
| max_depth | integer | No | 0–5; default 3 |
| include_paths | string[] | No | Path prefixes included in results |
| exclude_paths | string[] | No | Path prefixes excluded from crawling |
| group_id | string | Yes | SHA-256 of the origin; FIFO message group |
Output
S3 checkpoint and preview are stored under {organization_id}/{project_id}/page-discovery/{discovery_id}/. The preview contains status (0 processing, 1 completed, 2 failed), pages (url, title, depth), groups, warnings, truncated, and updated_at. Each group contains an ID, label, representative URL, member URLs, confidence, and grouping reason. It is returned through the authenticated backend, never read directly by the plugin. Shared batch and batch_item rows track submission and dispatch; no additional batch tables or DocumentDB collections are required.
Processing
- Validate public origin, options, and job identity.
- For a homepage, read robots.txt and discover same-origin links. For a
.xmlsitemap, parse sitemap and sitemap-index URLs without requesting the listed pages. - Reject private addresses, credentials, unsafe ports, external redirects, action links, and non-HTML downloads. Pin connections to validated public IP addresses.
- Remove fragments and tracking parameters; preserve meaningful query parameters.
- Process a bounded chunk, persist the frontier and visited URLs, then enqueue continuation if needed. FIFO grouping serializes work per origin; wait at least one second between requests and respect longer robots crawl delays.
- Group pages by their exact parent URL path. Page HTML and AI are not used for grouping.
- Publish a selectable preview. The plugin displays URLs as a nested path tree and does not preselect pages.
Runtime Configuration
| Variable | Description |
|---|---|
| RESULT_BUCKET | S3 checkpoint and preview bucket |
| DISCOVERY_QUEUE_URL | FIFO queue used for continuations |
| MAX_CHUNK_PAGES | Maximum tasks per invocation; default 10 |
Limitations
- No JavaScript navigation, authentication, form submission, or asset downloads.
- Same origin only, including redirects and sitemap locations.
- Maximum 1,000 discovered URLs, 1,000 queued candidates, and 20 sitemap documents per job. Page Import remains limited to 500 selected URLs per batch.
- Homepage discovery fails closed when robots.txt cannot be read. Direct sitemap input does not request robots.txt or the listed pages.