Skip to content

Code2Des Page Import

This folder imports rendered application routes into Guinness Page Import. By default it imports all static Next.js App Router pages under the configured folder, sequentially.

code2des-page-import/
โ”œโ”€โ”€ .env.example       Runtime configuration template
โ”œโ”€โ”€ README.md          This guide
โ”œโ”€โ”€ capture.mjs        Playwright page capture
โ”œโ”€โ”€ cookie_banners.js  Cookie consent cleanup (mirrored from the AWS worker)
โ”œโ”€โ”€ import.mjs         Upload and Page Import registration
โ”œโ”€โ”€ routes.mjs         Folder discovery, page selection, and batch execution
โ””โ”€โ”€ __tests__/         Import contract tests

The importer:

  1. Discovers the selected pages, then opens each route in Chromium and waits for client-side rendering.
  2. Scrolls the page to load lazy content, then captures DOM, computed styles, SVG, and a screenshot.
  3. Reads image assets from the browser session, including assets served only from localhost.
  4. Requests short-lived upload URLs, then sends the evidence, screenshot, and image bytes directly from this script to S3.
  5. Registers the source bundle as a Page Import and waits for processing to complete.

CSS and JavaScript files are not uploaded. Their rendered result is captured after the page settles. Page navigation and interaction-state discovery are outside this importer. The importer only creates a Page Import; conversion, organization, and component-library generation are separate actions. Capture follows the worker's visibility rules: fully clipped or transparent subtrees and offscreen fixed overlays are excluded, while below-fold page content is retained. Visible pseudo-element background images are included in the upload bundle. Cookie banners are automatically hidden before DOM and screenshot capture, including banners inserted after page load. Cleanup uses known consent-provider containers or cookie copy plus consent buttons in overlays/dialogs; it never clicks consent buttons. Ordinary forms and uncertain matches remain visible. Removals appear as cookie_banner_removed warnings. There is no configuration option. Unknown backdrops, shadow-root content, and cross-origin iframe contents are not traversed. Existing imports must be recaptured. Keep cookie_banners.js identical to guinness-ai-v2/apps/page-import/src/page_import/cookie_banners.js.

The backend receives only upload metadata and the source-manifest reference. Image assets larger than 3 MB and all video resources are skipped. Evidence and screenshots may use up to 20 MB per object, and the complete bundle is limited to 50 MB; oversized DOM evidence and screenshots are clipped before upload.

Configuration

Start the application route first:

npm run dev:next

Create the local environment file:

cp code2des-page-import/.env.example code2des-page-import/.env

Set GUINNESS_USER_TOKEN to a current Cognito user access token. Set the application URL and page folder in .env:

PAGE_IMPORT_BASE_URL=http://127.0.0.1:3000
PAGE_IMPORT_PAGES_PATH=src/app/proto-pages

The folder is relative to the frontend repository root. For example, src/app/proto-pages/faq discovers only FAQ pages but retains /proto-pages/faq/ in their URLs. With no folder configured, discovery uses all of src/app or app. Route groups are omitted from URLs; private folders, parallel slots, and intercepting routes are excluded. Dynamic templates such as [id] are reported as skipped: supply concrete routes using --pages or a routes file. A routes file replaces folder discovery and supports other frameworks. The legacy PAGE_IMPORT_URL is accepted as a base-URL fallback, not as a single-page selector; use --url for one page.

Variable Purpose Default
GUINNESS_API_URL Guinness backend stage or /api/v1 URL AWS dev
GUINNESS_USER_TOKEN Cognito user bearer token Required
GUINNESS_ORGANIZATION_ID Target organization 1
GUINNESS_PROJECT_ID Target project 1
PAGE_IMPORT_BASE_URL Running application base URL http://127.0.0.1:3000
PAGE_IMPORT_PAGES_PATH Folder beneath src/app or app Auto-detect app root
PAGE_IMPORT_ROUTES_FILE JSON array or newline list of concrete routes/URLs Empty
PAGE_IMPORT_READY_SELECTOR Optional visible readiness selector Empty
PAGE_IMPORT_SETTLE_MS Wait after rendering and lazy loading 750
PAGE_IMPORT_OUTPUT_DIR Local evidence and report code2des-page-import/output
PLAYWRIGHT_CHANNEL Chrome channel, or bundled for Playwright Chromium chrome

Run

npm run guinness:code2des-page-import

This imports all pages in the configured folder. Preview the selection without capturing or uploading:

npm run guinness:code2des-page-import -- --list

Import selected routes instead (repeat --pages if needed):

npm run guinness:code2des-page-import -- --pages /proto-pages/faq/cm_ftr_faq01,/des2code/AC_RGST01

Override the folder with --pages-path src/app/proto-pages/faq, or use --routes-file routes.json containing a JSON array such as ["/product/123", "/product/456"]. --url is repeatable and also disables discovery. An explicitly empty --pages value is rejected.

Override the route or viewport when needed:

npm run guinness:code2des-page-import -- \
  --url http://127.0.0.1:3000/race/ \
  --name Race \
  --ready-selector main \
  --width 1440 \
  --height 900

To validate browser capture without uploading:

npm run guinness:code2des-page-import -- --dry-run

Each run writes a unique run-* directory under code2des-page-import/output, with numbered page directories containing evidence, screenshots, source bundles, and individual completion reports. batch-report.json is updated after each page. A failed page is recorded and processing continues; any failure makes the command exit nonzero. HTTP error pages are rejected. Re-running creates new Page Import IDs; it does not resume an earlier batch. --dry-run captures the entire selected set without uploading.

Run the import contract and browser capture regression tests with npm run test:guinness-code2des-page-import. The browser tests use Chrome by default; set PLAYWRIGHT_CHANNEL=bundled to use an installed Playwright Chromium.