Code2Des Page Import
This folder imports rendered application routes into Guinness Page Import. By default it imports all static Next.js App Router pages under the configured folder, sequentially.
code2des-page-import/
├── .env.example Runtime configuration template
├── README.md This guide
├── capture.mjs Playwright page capture
├── cookie_banners.js Cookie consent cleanup (mirrored from the AWS worker)
├── import.mjs Upload and Page Import registration
├── routes.mjs Folder discovery, page selection, and batch execution
└── __tests__/ Import contract tests
The importer:
- Discovers the selected pages, then opens each route in Chromium and waits for client-side rendering.
- Scrolls the page to load lazy content, then captures DOM, computed styles, SVG, and a screenshot.
- Reads image assets from the browser session, including assets served only from localhost.
- Requests short-lived upload URLs, then sends the evidence, screenshot, and image bytes directly from this script to S3.
- Registers the source bundle as a Page Import and waits for processing to complete.
CSS and JavaScript files are not uploaded. Their rendered result is captured after the page settles. Page navigation and interaction-state discovery are outside this importer.
The importer only creates a Page Import; conversion, organization, and component-library generation are separate actions.
Capture follows the worker's visibility rules: fully clipped or transparent subtrees and offscreen fixed overlays are excluded, while below-fold page content is retained. Visible pseudo-element background images are included in the upload bundle.
Cookie banners are automatically hidden before DOM and screenshot capture, including banners inserted after page load. Cleanup uses known consent-provider containers or cookie copy plus consent buttons in overlays/dialogs; it never clicks consent buttons. Ordinary forms and uncertain matches remain visible. Removals appear as cookie_banner_removed warnings. There is no configuration option. Unknown backdrops, shadow-root content, and cross-origin iframe contents are not traversed. Existing imports must be recaptured. Keep cookie_banners.js identical to guinness-ai-v2/apps/page-import/src/page_import/cookie_banners.js.
The backend receives only upload metadata and the source-manifest reference. Image assets larger than 3 MB and all video resources are skipped. Evidence and screenshots may use up to 20 MB per object, and the complete bundle is limited to 50 MB; oversized DOM evidence and screenshots are clipped before upload.
Configuration
Start the application route first:
Create the local environment file:
Set GUINNESS_USER_TOKEN to a current Cognito user access token. Set the application URL and page folder in .env:
The folder is relative to the frontend repository root. For example, src/app/proto-pages/faq discovers only FAQ pages but retains /proto-pages/faq/ in their URLs. With no folder configured, discovery uses all of src/app or app. Route groups are omitted from URLs; private folders, parallel slots, and intercepting routes are excluded. Dynamic templates such as [id] are reported as skipped: supply concrete routes using --pages or a routes file. A routes file replaces folder discovery and supports other frameworks. The legacy PAGE_IMPORT_URL is accepted as a base-URL fallback, not as a single-page selector; use --url for one page.
| Variable | Purpose | Default |
|---|---|---|
GUINNESS_API_URL |
Guinness backend stage or /api/v1 URL |
AWS dev |
GUINNESS_USER_TOKEN |
Cognito user bearer token | Required |
GUINNESS_ORGANIZATION_ID |
Target organization | 1 |
GUINNESS_PROJECT_ID |
Target project | 1 |
PAGE_IMPORT_BASE_URL |
Running application base URL | http://127.0.0.1:3000 |
PAGE_IMPORT_PAGES_PATH |
Folder beneath src/app or app |
Auto-detect app root |
PAGE_IMPORT_ROUTES_FILE |
JSON array or newline list of concrete routes/URLs | Empty |
PAGE_IMPORT_READY_SELECTOR |
Optional visible readiness selector | Empty |
PAGE_IMPORT_SETTLE_MS |
Wait after rendering and lazy loading | 750 |
PAGE_IMPORT_OUTPUT_DIR |
Local evidence and report | code2des-page-import/output |
PLAYWRIGHT_CHANNEL |
Chrome channel, or bundled for Playwright Chromium |
chrome |
Run
This imports all pages in the configured folder. Preview the selection without capturing or uploading:
Import selected routes instead (repeat --pages if needed):
Override the folder with --pages-path src/app/proto-pages/faq, or use --routes-file routes.json containing a JSON array such as ["/product/123", "/product/456"]. --url is repeatable and also disables discovery. An explicitly empty --pages value is rejected.
Override the route or viewport when needed:
npm run guinness:code2des-page-import -- \
--url http://127.0.0.1:3000/race/ \
--name Race \
--ready-selector main \
--width 1440 \
--height 900
To validate browser capture without uploading:
Each run writes a unique run-* directory under code2des-page-import/output, with numbered page directories containing evidence, screenshots, source bundles, and individual completion reports. batch-report.json is updated after each page. A failed page is recorded and processing continues; any failure makes the command exit nonzero. HTTP error pages are rejected. Re-running creates new Page Import IDs; it does not resume an earlier batch. --dry-run captures the entire selected set without uploading.
Run the import contract and browser capture regression tests with npm run test:guinness-code2des-page-import. The browser tests use Chrome by default; set PLAYWRIGHT_CHANNEL=bundled to use an installed Playwright Chromium.