コンテンツにスキップ

Code2Des Page Import

This folder imports rendered application routes into Guinness Page Import. By default it imports all static Next.js App Router pages under the configured folder, sequentially.

code2des-page-import/
├── .env.example       Runtime configuration template
├── README.md          This guide
├── capture.mjs        Playwright page capture
├── cookie_banners.js  Cookie consent cleanup (mirrored from the AWS worker)
├── import.mjs         Upload and Page Import registration
├── routes.mjs         Folder discovery, page selection, and batch execution
└── __tests__/         Import contract tests

The importer:

  1. Discovers the selected pages, then opens each route in Chromium and waits for client-side rendering.
  2. Scrolls the page to load lazy content, then captures DOM, computed styles, SVG, and a screenshot.
  3. Reads image assets from the browser session, including assets served only from localhost.
  4. Requests short-lived upload URLs, then sends the evidence, screenshot, and image bytes directly from this script to S3.
  5. Registers the source bundle as a Page Import and waits for processing to complete.

CSS and JavaScript files are not uploaded. Their rendered result is captured after the page settles. Page navigation and interaction-state discovery are outside this importer. The importer only creates a Page Import; conversion, organization, and component-library generation are separate actions. Capture follows the worker's visibility rules: fully clipped or transparent subtrees and offscreen fixed overlays are excluded, while below-fold page content is retained. Visible pseudo-element background images are included in the upload bundle. Cookie banners are automatically hidden before DOM and screenshot capture, including banners inserted after page load. Cleanup uses known consent-provider containers or cookie copy plus consent buttons in overlays/dialogs; it never clicks consent buttons. Ordinary forms and uncertain matches remain visible. Removals appear as cookie_banner_removed warnings. There is no configuration option. Unknown backdrops, shadow-root content, and cross-origin iframe contents are not traversed. Existing imports must be recaptured. Keep cookie_banners.js identical to guinness-ai-v2/apps/page-import/src/page_import/cookie_banners.js.

The backend receives only upload metadata and the source-manifest reference. Image assets larger than 3 MB and all video resources are skipped. Evidence and screenshots may use up to 20 MB per object, and the complete bundle is limited to 50 MB; oversized DOM evidence and screenshots are clipped before upload.

Configuration

Start the application route first:

npm run dev:next

Create the local environment file:

cp code2des-page-import/.env.example code2des-page-import/.env

Set GUINNESS_USER_TOKEN to a current Cognito user access token. Set the application URL and page folder in .env:

PAGE_IMPORT_BASE_URL=http://127.0.0.1:3000
PAGE_IMPORT_PAGES_PATH=src/app/proto-pages

The folder is relative to the frontend repository root. For example, src/app/proto-pages/faq discovers only FAQ pages but retains /proto-pages/faq/ in their URLs. With no folder configured, discovery uses all of src/app or app. Route groups are omitted from URLs; private folders, parallel slots, and intercepting routes are excluded. Dynamic templates such as [id] are reported as skipped: supply concrete routes using --pages or a routes file. A routes file replaces folder discovery and supports other frameworks. The legacy PAGE_IMPORT_URL is accepted as a base-URL fallback, not as a single-page selector; use --url for one page.

Variable Purpose Default
GUINNESS_API_URL Guinness backend stage or /api/v1 URL AWS dev
GUINNESS_USER_TOKEN Cognito user bearer token Required
GUINNESS_ORGANIZATION_ID Target organization 1
GUINNESS_PROJECT_ID Target project 1
PAGE_IMPORT_BASE_URL Running application base URL http://127.0.0.1:3000
PAGE_IMPORT_PAGES_PATH Folder beneath src/app or app Auto-detect app root
PAGE_IMPORT_ROUTES_FILE JSON array or newline list of concrete routes/URLs Empty
PAGE_IMPORT_READY_SELECTOR Optional visible readiness selector Empty
PAGE_IMPORT_SETTLE_MS Wait after rendering and lazy loading 750
PAGE_IMPORT_OUTPUT_DIR Local evidence and report code2des-page-import/output
PLAYWRIGHT_CHANNEL Chrome channel, or bundled for Playwright Chromium chrome

Run

npm run guinness:code2des-page-import

This imports all pages in the configured folder. Preview the selection without capturing or uploading:

npm run guinness:code2des-page-import -- --list

Import selected routes instead (repeat --pages if needed):

npm run guinness:code2des-page-import -- --pages /proto-pages/faq/cm_ftr_faq01,/des2code/AC_RGST01

Override the folder with --pages-path src/app/proto-pages/faq, or use --routes-file routes.json containing a JSON array such as ["/product/123", "/product/456"]. --url is repeatable and also disables discovery. An explicitly empty --pages value is rejected.

Override the route or viewport when needed:

npm run guinness:code2des-page-import -- \
  --url http://127.0.0.1:3000/race/ \
  --name Race \
  --ready-selector main \
  --width 1440 \
  --height 900

To validate browser capture without uploading:

npm run guinness:code2des-page-import -- --dry-run

Each run writes a unique run-* directory under code2des-page-import/output, with numbered page directories containing evidence, screenshots, source bundles, and individual completion reports. batch-report.json is updated after each page. A failed page is recorded and processing continues; any failure makes the command exit nonzero. HTTP error pages are rejected. Re-running creates new Page Import IDs; it does not resume an earlier batch. --dry-run captures the entire selected set without uploading.

Run the import contract and browser capture regression tests with npm run test:guinness-code2des-page-import. The browser tests use Chrome by default; set PLAYWRIGHT_CHANNEL=bundled to use an installed Playwright Chromium.