Skip to content

AI Page Import — I/O Definition

SQS input

{
  "page_import_id": "019ffa75-01c0-75a2-8123-456789abcdf0",
  "attempt": 1,
  "nonce": "transport-fence",
  "organization_id": 1,
  "project_id": 7,
  "source_url": "https://example.com/",
  "viewport": { "width": 1440, "height": 900, "device_scale_factor": 1 },
  "ready_selector": null
}

Unknown fields are rejected. Tenant IDs must be positive; attempt is 1; source URL must use HTTP(S); viewport is width/height 240..7680 and device scale factor 0.5..4; ready_selector is nullable and at most 512 characters. The backend persists the nonce and accepts a callback only when it matches the row.

Uploaded pages use source_kind: "codebase_route", source_manifest_url, and source_manifest_hash. Their source URL identifies a locally rendered page and is not fetched by the worker. Only static capture is supported. Public URL messages use source_kind: "public_url" and omit manifest fields.

The script uploads objects directly to S3 using backend-authorized PUT URLs before submitting Page Import. The queue carries references, not file bytes. The manifest URL must be s3://{result_bucket}/{organization_id}/{project_id}/page-import/{page_import_id}/source/manifest.json. Its SHA-256, tenant/job IDs, source URL, and viewport must match the queue message.

The manifest contains source_bundle_version: "1.0", captured_at, source, viewport, evidence, screenshot, and assets. Object references contain key, hash, content_type, and byte_length; assets also contain source_url. Keys use source/objects/{sha256}.{ext}. The worker verifies length and hash before persisting canonical artifacts.

Evidence uses application/vnd.guinness.page-evidence+json and contains evidence_version: "1.0", document bounds, nodes, font faces, warnings, and truncation flags. Screenshots must be PNG. The manifest limit is 2 MB; evidence and screenshots are limited to 20 MB each. Image assets are limited to 3 MB each or the lower configured limit, with configured count/total-byte limits. Videos are rejected. Public URL capture skips videos and oversized images.

Capture output

The worker writes under:

{organization_id}/{project_id}/page-import/{page_import_id}/
Key Content
screenshot.png Full-page PNG
assets/{sha256}.{ext} Immutable fetched image/SVG-derived assets
capture.json RichPageCapture: source/final URL, viewport, document bounds, ordered nodes, assets, screenshot hash/URL, warnings
result.json Success summary pointing to capture and screenshot
manifest.json Successful backend row effect
failed.json / failed-manifest.json Safe terminal failure artifact/effect

JSON is serialized with sorted keys and compact separators. Immutable writes accept identical existing bytes and reject conflicting bytes. capture_hash is the SHA-256 of the exact capture.json bytes.

Captured node evidence

Nodes contain stable IDs and parent IDs, kind/tag/name/role/text, safe attributes, document-relative bounds, selected computed styles, optional asset hash or inline SVG, and paint order. The collector omits hidden/zero-size nodes, enforces a configured node limit, and emits node_limit_reached when truncated.

Images are copied into S3. SVG that Figma cannot reliably import is rasterized in the browser; supported inline SVG remains vector evidence. CSS image saturation/grayscale is normalized for native Figma image filters.

Terminal webhook

{
  "type": "page_import",
  "job_id": "019ffa75-01c0-75a2-8123-456789abcdf0",
  "attempt": 1,
  "nonce": "transport-fence",
  "status": "succeeded",
  "result_manifest_url": "s3://bucket/1/7/page-import/.../manifest.json",
  "manifest_schema_version": 1
}

The manifest has job_type: "page_import" and row_effects.page_import with status "1" plus capture/result pointers, or status "2" plus the failure pointer. The backend verifies bucket, tenant/job prefix, job ID, attempt, nonce, and row-effect shape. A successful PostgreSQL compare-and-set occurs once; duplicate delivery is idempotent. The callback never creates or dispatches Code2Des.

Database boundary

Workers have no PostgreSQL or DocumentDB credentials. PostgreSQL stores lifecycle and capture pointers; S3 stores the detailed capture. The backend is the sole database writer.