AI Page Import — Overview
Page Import is an independent, reusable public-URL capture system. POST /page-imports creates a PostgreSQL page_import row and sends one attempt-fenced SQS message. apps/page-import opens the URL in Playwright Chromium, captures the fully rendered page, writes immutable evidence to S3, and reports a terminal manifest through POST /api/v1/webhooks/ai-status.
Users do not start this worker from the Figma plugin. Public URL capture is started separately through REST. The plugin lists completed Page Imports, and only after the user selects one does it call the separate Code2Des API and materialize that new result.
Scope
- One public HTTP(S) URL and one viewport per import.
- Post-JavaScript DOM, computed style, document geometry, paint order, SVG evidence, screenshot, and fetched image assets.
- Optional
readySelector; bounded navigation, readiness, and asset-download timeouts. - No LLM, embedding, PostgreSQL, or DocumentDB access in the worker.
- No crawling, authenticated browser session, form submission, or multi-viewport capture.
Security boundary
The worker uses a fresh browser context. URLs with credentials or nonstandard ports are rejected. DNS results must be public, private/link-local/loopback requests are blocked during navigation, and every asset redirect target is revalidated before download. Asset and result writes are immutable and tenant-scoped.
Ownership
| Layer | Responsibility |
|---|---|
| Backend | Authorization, Page Import PostgreSQL lifecycle, nonce/attempt generation, queue dispatch, and webhook row transitions |
| Page Import worker | Browser capture, normalization, immutable S3 artifacts, terminal webhook |
| Code2Des worker | Deterministic conversion of the pinned capture into a native-Figma spec |
| Figma plugin | Completed-job selection and native node creation only |
The separately documented repository/CLI upload API remains a planned future source adapter. It is not the source path used by the implemented public-URL pipeline.