Skip to content

AI Code Import โ€” Test Case Design

This document is the Single Source of Truth for testing the code-import worker. It defines:

  1. Test Strategy โ€” categories, layers, and areas of responsibility.
  2. Test Environments โ€” local, CI, dev (deployed Lambda), staging.
  3. Test Fixtures โ€” sample payloads, sample code, sample images.
  4. Test Catalog โ€” all test cases (ID, purpose, input, expected output, priority).
  5. Coverage Goals โ€” definition of "done".
  6. CI Integration โ€” how to run on every PR and deployment.

The official I/O contract for what is being tested is in I/O Definition; technical details and processing flow are in Overview.

The paths below assume apps/code-import/ in the repository (Python package code_import under src/code_import/). __tests__/ is the target layout for the current scaffold.


1. Test Strategy

Pyramid

     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
     โ”‚  E2E / Smoke (dev)   โ”‚   ~3 cases  (real AWS, real OpenAI, real DocDB)
     โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
     โ”‚  Integration (CI)    โ”‚   ~10 cases (process_record with mocks)
     โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
     โ”‚  Contract (CI)       โ”‚   ~6 cases  (schema validation, Webhook format)
     โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
     โ”‚  Unit (CI)           โ”‚   ~25 cases (pure functions, agent outputs)
     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Categories

Category Scope Runs In Mocks
Unit Single function or class CI (every push) All external (S3, OpenAI, DocDB)
Contract Schema / payload validation CI (every push) None โ€” pure data
Integration Full process_record(record) CI (every push) S3, OpenAI, DocDB, Webhook (HTTP)
Smoke (dev) Real Lambda invocation Manual / nightly None
Load Throughput & latency Manual / pre-release OpenAI (recorded playback)

Areas of Responsibility

Layer Owner Tools
Unit + Contract + Integration PLABS.ID (worker development) pytest, pytest-asyncio, moto, mongomock, respx
Smoke (dev) PLABS.ID operations aws lambda invoke + pytest assertions
Load PLABS.ID + 4D infrastructure locust against dev Lambda

2. Test Environments

Environment Purpose OpenAI S3 DocDB Webhook
local (developer PC) Unit + Integration Mock moto mongomock respx
ci (GitHub Actions) Same as local Mock moto mongomock respx
dev (deployed AWS) Smoke Real (gpt-5.4-nano) Real bucket Real cluster Real backend
staging Pre-release smoke + load Real Real Real Real

Cost guard: Tests that call real OpenAI must declare the @pytest.mark.real_openai marker and are skipped by default. Enable with pytest -m real_openai.


3. Test Fixtures

All fixtures are located under apps/code-import/__tests__/fixtures/.

3.1 Sample SQS Records

Fixture File Variant
record_minimal.json Required fields only (no css_code, no img_url)
record_with_screenshot.json All fields including img_url
record_with_css.json All fields except img_url
record_full.json All fields including css_code and img_url
record_page.json type=0 (page)
record_based_on_design.json based_on=1
record_based_on_wireframe.json based_on=2
record_invalid_uuid.json Invalid code_id
record_missing_field.json Missing name
record_invalid_type.json type=2
record_empty_source_code.json source_code=""
record_huge_source_code.json 250 KB source_code
record_url_encoded_s3.json img_url with %E3%83%95-style escaping
record_double_wrapped.json Body nesting a Records array (workaround for legacy backend bug)

3.2 Sample Source Code

Fixture File Description
code_button.tsx Simple Tailwind button
code_modal.tsx Modal dialog with form
code_table.tsx Sortable data table
code_page_login.tsx Full page (type=0)
code_with_inline_styles.tsx No classes, inline styles only

3.3 Sample Screenshots

Fixture File Format Notes
button_primary.png PNG 320x80, normal case
button_primary.jpg JPEG Same scene, with compression artifacts
large_image.png PNG 4 MB, near upload limit
corrupt.png โ€” Truncated byte sequence

3.4 Sample LLM Responses (Recorded)

Recorded with pytest-recording (or hand-written) under fixtures/llm_responses/. Used to deterministically replay agent outputs in CI.

Fixture Input Target Return Value
semantic_button.json Code Analysis LLM (arch node C) ["submit","action","primary","cta","button"]
context_button.json Deterministic source parser component level/family/variant/import/export/reference/class/token/text evidence
embedding_512.json Configured embedding provider 2 float vectors of 512 dimensions

4. Test Catalog

4.1 Unit Tests โ€” apps/code-import/__tests__/unit/

Schema Validation (test_schemas.py)

ID Purpose Input Expected
U-SCH-001 Accept required fields only record_minimal.json body CodeImportMessage instance, css_code=None, img_url=None
U-SCH-002 All fields assigned correctly record_full.json body All fields match expected values
U-SCH-003 Reject invalid UUID record_invalid_uuid.json body ValidationError on code_id
U-SCH-004 Reject missing required field record_missing_field.json body ValidationError on name
U-SCH-005 Reject out-of-range type record_invalid_type.json body ValidationError on type
U-SCH-006 Reject out-of-range based_on based_on=3 ValidationError
U-SCH-007 Reject empty source_code record_empty_source_code.json ValidationError
U-SCH-008 Optional css_code accepts None minimal + css_code=None OK
U-SCH-009 Optional img_url accepts None minimal + img_url=None OK

S3 URL Parsing (test_s3_helpers.py)

ID Purpose Input Expected
U-S3-001 Parse s3:// URL s3://bucket/key/file.png ("bucket", "key/file.png")
U-S3-002 Parse virtual-hosted HTTPS https://bucket.s3.us-east-1.amazonaws.com/file.png ("bucket", "file.png")
U-S3-003 Parse path-style HTTPS https://s3.us-east-1.amazonaws.com/bucket/file.png ("bucket", "file.png")
U-S3-004 URL-decode key s3://bucket/%E3%83%95%E3%82%9A%E3%83%AD.png Key contains ใƒ•ยฐใƒญ.png (unquoted)
U-S3-005 Reject unsupported scheme ftp://bucket/key ValueError

Image Format Detection (test_image_helpers.py)

ID Purpose Input Expected
U-IMG-001 Detect PNG URL canonical .png URL "png"
U-IMG-002 Detect JPEG URL canonical .jpg URL "jpeg"
U-IMG-003 Detect GIF URL .gif URL "gif" (helper compatibility; worker canonical validation rejects it)
U-IMG-004 Detect WebP URL canonical .webp URL "webp"
U-IMG-005 Unknown extension fallback URL without known extension "png"

Embedding Helpers (test_embedding.py)

ID Purpose Input Expected
U-EMB-001 2 inputs produce 2 vectors of 512 dimensions Mock embedding client Tuple of 2 list[float] of length 512
U-EMB-002 Single batch call Client spy Exactly 1 API call
U-EMB-003 Semantic input preserves component identity name/kind/context/agent terms deterministic identity string

Agent Output Models (test_agent_outputs.py)

ID Purpose Input Expected
U-AGT-001 CodeSemantics accepts 5 words List of 5 strings OK
U-AGT-002 CodeSemantics accepts 10 words List of 10 strings OK
U-AGT-003 CodeSemantics rejects 4 words List of 4 strings ValidationError
U-AGT-004 CodeSemantics rejects 11 words List of 11 strings ValidationError
U-AGT-005 CodeSemantics strips blanks before length validation mixed blank/nonblank terms canonical bounded list

4.2 Contract Tests โ€” __tests__/contract/

ID Purpose Input Expected
C-WH-001 Success Webhook matches OpenAPI-style schema Success payload from worker Conforms to ai-status webhook schema
C-WH-002 Failure Webhook matches schema Failure payload from worker Conforms
C-WH-003 Webhook type is always "code-import" Any Webhook type == "code-import"
C-WH-004 Webhook recordId exactly echoes code_id Success path recordId == code_id
C-WH-005 Webhook scope echoes SQS tenant/project Success and failure payloads organizationId == organization_id and projectId == project_id
C-DOC-001 DocDB document has strict visual metadata, semantics, context, and processing blocks Sample run Extra legacy fields rejected; no visual vector
C-DOC-002 semantic/context vectors are exactly 512 Sample run Both array lengths == 512

4.3 Integration Tests โ€” __tests__/integration/test_process_record.py

All tests use mongomock, moto, respx, and recorded LLM responses. Each test calls process_record(record) end-to-end.

Happy Path

ID Purpose Fixture Expected DB Write Expected Webhook
I-HP-001 Minimal payload, no screenshot, no CSS record_minimal.json css_code=null, visual.metadata.image_url=null, semantic/context embeddings 512 dimensions success, keywords length 5โ€“10
I-HP-002 All fields with screenshot record_full.json visual.metadata.image_url == img_url success
I-HP-003 With CSS, no screenshot record_with_css.json css_code saved, visual.metadata.image_url=null success
I-HP-004 Page (type=0) record_page.json type=0 saved success
I-HP-005 based_on=1 (design) record_based_on_design.json based_on=1 saved success
I-HP-006 based_on=2 (wireframe) record_based_on_wireframe.json based_on=2 saved success
I-HP-007 Canonical S3 reference configured bucket/exact preview key Reference accepted; no image download success
I-HP-008 Idempotent reprocessing Run record_full.json twice Second run reuses fingerprint/vectors; paid work skipped Success Webhook twice
I-HP-009 Double-wrapped body (legacy) record_double_wrapped.json Inner body extracted, normal processing success

Failure Path

Repository regression tests cover transient write conflicts: retry the complete database transaction with a fresh session, preserve atomic index invalidation, and do not repeat AI calls. Permanent failures are not retried locally; retry exhaustion remains an SQS record failure.

ID Purpose Fixture / Mock Expected Behavior
I-FP-001 Empty body string {"body":""} ValueError, no Webhook (no code_id), re-raise exception
I-FP-002 Invalid JSON body {"body":"{not json"} JSONDecodeError, no Webhook, re-raise
I-FP-003 Pydantic validation failure record_invalid_uuid.json ValidationError, partial Webhook attempt (best-effort), re-raise
I-FP-004 Wrong preview key Reference does not equal canonical key Webhook failed, record failure
I-FP-005 Wrong preview bucket Reference does not target S3_BUCKET_NAME Webhook failed, record failure
I-FP-006 LLM semantic timeout Semantic agent raises AgentRunError Webhook failed, re-raise
I-FP-007 Invalid embedding response Missing or wrong-sized second vector Webhook failed, record failure
I-FP-008 Embedding API rate limit openai.RateLimitError Webhook failed, re-raise
I-FP-009 DocumentDB write failure mongomock raises PyMongoError Webhook failed, re-raise
I-FP-010 Webhook itself fails (5xx) respx returns 500 Log and re-raise so SQS retries
I-FP-011 Webhook 4xx (bad auth) respx returns 401 Log and re-raise so SQS retries

Edge Cases

ID Purpose Fixture Expected
I-EC-001 Source code at 200 KB limit record_huge_source_code.json Accepted, processed
I-EC-002 Source code with non-ASCII (Japanese, Vietnamese) Unicode source_code Correctly round-trips to DocDB
I-EC-003 Preview is JPEG canonical .jpg reference Stored unchanged; no download
I-EC-004 Preview is WebP canonical .webp reference Stored unchanged; no download
I-EC-005 Preview URL changes with same fingerprint same source, new canonical extension Only image URL patched; index remains valid
I-EC-006 2 records in 1 SQSEvent Batch of 2 records Both processed, attempt second even if first fails

4.4 Smoke Tests โ€” __tests__/smoke/

Run against deployed dev Lambda only. Marked with @pytest.mark.smoke.

ID Purpose Steps Expected
S-001 Real end-to-end (with preview) 1. Send record_full.json to SQS. 2. Wait โ‰ค 3 minutes. 3. Query DocDB by scope + code_id. 4. Query PG code row. DocDB has semantic/context vectors and processing hash; PG status completed
S-002 Real end-to-end (no screenshot) Send record_with_css.json. Same wait + verification. Same as above, visual.metadata.image_url=null
S-003 Real end-to-end (failure path) Send payload with invalid preview reference. PG row failed within 3 minutes; webhook log shows failed

4.5 Load Tests โ€” __tests__/load/

Run before each release. Marked with @pytest.mark.load.

ID Purpose Profile Pass Criteria
L-001 Steady throughput 10 msg/min for 30 minutes p50 < 90 s, p95 < 180 s, error rate < 1%
L-002 Burst processing 100 msg in 1 minute, then idle All processed within 10 minutes, no DLQ messages
L-003 Sustained burst 30 msg/min for 1 hour No DLQ, no OpenAI rate limit exceeded

5. Coverage Goals

Layer Line Coverage Branch Coverage Notes
apps/code-import/src/code_import/handler.py โ‰ฅ 90% โ‰ฅ 85% Lambda entry (lambda_handler) โ€” must cover all try/except branches
apps/code-import/src/code_import/schemas.py 100% 100% Pydantic models โ€” test exhaustively
packages/agentic/agents/code_semantics.py โ‰ฅ 80% โ€” Builder only, LLM runtime is mocked
packages/des2code-core/code_context.py โ‰ฅ 90% โ‰ฅ 85% Deterministic parser/identity text
shared embedding builder 100% โ€” Provider selection and one two-input call
packages/guinness-ai-sdk webhook โ‰ฅ 90% โ‰ฅ 85% Including retry path
packages/models/documentdb/code.py โ‰ฅ 85% โ€” Index creation idempotency

Run:

pytest --cov=code_import --cov=packages \
       --cov-report=term-missing --cov-fail-under=85

6. CI Integration

Per-PR Pipeline (.github/workflows/test-code-import.yml)

on:
  pull_request:
    paths:
      - 'apps/code-import/**'
      - 'packages/agentic/agents/code_*.py'
      - 'packages/helpers/**'
      - 'packages/models/documentdb/code.py'

jobs:
  test:
    steps:
      - uv sync
      - pytest -m "not real_openai and not smoke and not load" \
               --cov --cov-fail-under=85
      - ruff check apps/code-import packages
      - mypy apps/code-import packages

Nightly (against dev)

on:
  schedule:
    - cron: '0 18 * * *'   # 03:00 JST

jobs:
  smoke:
    steps:
      - uv sync
      - pytest -m smoke

Pre-Release (against staging)

A manually triggered workflow that runs -m "smoke or load".

Failure Policy

CI Stage On Failure
Lint / type / unit / contract / integration Block merge
Below coverage threshold Block merge
Nightly smoke Notify #guinness-ai Slack, open issue after 3 consecutive failures
Load tests Block release, do not block merge

7. Quick Test Matrix (Cheat Sheet)

Scenario Test IDs
Required fields only U-SCH-001, I-HP-001
All optional fields present U-SCH-002, I-HP-002
Page vs code I-HP-004
All based_on variants I-HP-005, I-HP-006
Screenshot missing I-HP-001, I-HP-003
Invalid input (validation) U-SCH-003 .. U-SCH-007, I-FP-003
External service failures I-FP-004 .. I-FP-009
Webhook retry behavior I-FP-010, I-FP-011
Idempotency I-HP-008
Real-world end-to-end S-001, S-002
Throughput L-001, L-002

8. Definition of Done (Per PR)

A code-import PR is mergeable when:

  • [ ] All applicable test cases are implemented and passing.
  • [ ] Coverage thresholds met (see ยง5).
  • [ ] At least one new integration test added for any behavioral change.
  • [ ] If SQS schema, Webhook schema, or DocDB document shape changed, manually verify S-001โ€“S-003 smoke tests on dev.
  • [ ] PR description links to the relevant test IDs.

For repository-wide release readiness, see docs/overview/non-functional.en.md (SLA / SLO targets).