AI Page Discovery — Overview
Page Discovery finds public pages from a homepage, sitemaps, and HTML links. Users review the discovered URLs before creating a separate Page Import batch.
Scope
- One homepage per discovery; up to 500 pages.
- Same-origin HTTP(S), robots.txt rules, deduplication, depth and path filters.
- Checkpointed SQS processing; closing the plugin does not stop discovery.
- No JavaScript navigation, asset capture, authentication, or Figma generation.
Security boundary
The worker rejects non-public DNS addresses, unsafe ports, credentials, external redirects, and known action links. Connections use validated IP addresses. Responses are bounded to 2 MB and requests are throttled. robots.txt retrieval errors stop discovery.
Ownership
| Layer | Responsibility |
|---|---|
| Backend | Project access, shared batch submission, FIFO dispatch, authenticated preview |
| Page Discovery worker | Bounded discovery, S3 checkpoints/results, continuation messages |
| Figma plugin | Preview, URL selection, separate import action |
The local implementation requires separate FIFO SQS and Lambda resources before AWS deployment.