AI Code Import โ Test Case Design
This document is the Single Source of Truth for testing the code-import worker. It defines:
- Test Strategy โ categories, layers, and areas of responsibility.
- Test Environments โ local, CI, dev (deployed Lambda), staging.
- Test Fixtures โ sample payloads, sample code, sample images.
- Test Catalog โ all test cases (ID, purpose, input, expected output, priority).
- Coverage Goals โ definition of "done".
- CI Integration โ how to run on every PR and deployment.
The official I/O contract for what is being tested is in I/O Definition; technical details and processing flow are in Overview.
The paths below assume apps/code-import/ in the repository (Python package code_import under src/code_import/). __tests__/ is the target layout for the current scaffold.
1. Test Strategy
Pyramid
โโโโโโโโโโโโโโโโโโโโโโโโ
โ E2E / Smoke (dev) โ ~3 cases (real AWS, real OpenAI, real DocDB)
โโโโโโโโโโโโโโโโโโโโโโโโค
โ Integration (CI) โ ~10 cases (process_record with mocks)
โโโโโโโโโโโโโโโโโโโโโโโโค
โ Contract (CI) โ ~6 cases (schema validation, Webhook format)
โโโโโโโโโโโโโโโโโโโโโโโโค
โ Unit (CI) โ ~25 cases (pure functions, agent outputs)
โโโโโโโโโโโโโโโโโโโโโโโโ
Categories
| Category |
Scope |
Runs In |
Mocks |
| Unit |
Single function or class |
CI (every push) |
All external (S3, OpenAI, DocDB) |
| Contract |
Schema / payload validation |
CI (every push) |
None โ pure data |
| Integration |
Full process_record(record) |
CI (every push) |
S3, OpenAI, DocDB, Webhook (HTTP) |
| Smoke (dev) |
Real Lambda invocation |
Manual / nightly |
None |
| Load |
Throughput & latency |
Manual / pre-release |
OpenAI (recorded playback) |
Areas of Responsibility
| Layer |
Owner |
Tools |
| Unit + Contract + Integration |
PLABS.ID (worker development) |
pytest, pytest-asyncio, moto, mongomock, respx |
| Smoke (dev) |
PLABS.ID operations |
aws lambda invoke + pytest assertions |
| Load |
PLABS.ID + 4D infrastructure |
locust against dev Lambda |
2. Test Environments
| Environment |
Purpose |
OpenAI |
S3 |
DocDB |
Webhook |
local (developer PC) |
Unit + Integration |
Mock |
moto |
mongomock |
respx |
ci (GitHub Actions) |
Same as local |
Mock |
moto |
mongomock |
respx |
dev (deployed AWS) |
Smoke |
Real (gpt-5.4-nano) |
Real bucket |
Real cluster |
Real backend |
staging |
Pre-release smoke + load |
Real |
Real |
Real |
Real |
Cost guard: Tests that call real OpenAI must declare the @pytest.mark.real_openai marker and are skipped by default. Enable with pytest -m real_openai.
3. Test Fixtures
All fixtures are located under apps/code-import/__tests__/fixtures/.
3.1 Sample SQS Records
| Fixture File |
Variant |
record_minimal.json |
Required fields only (no css_code, no img_url) |
record_with_screenshot.json |
All fields including img_url |
record_with_css.json |
All fields except img_url |
record_full.json |
All fields including css_code and img_url |
record_page.json |
type=0 (page) |
record_based_on_design.json |
based_on=1 |
record_based_on_wireframe.json |
based_on=2 |
record_invalid_uuid.json |
Invalid code_id |
record_missing_field.json |
Missing name |
record_invalid_type.json |
type=2 |
record_empty_source_code.json |
source_code="" |
record_huge_source_code.json |
250 KB source_code |
record_url_encoded_s3.json |
img_url with %E3%83%95-style escaping |
record_double_wrapped.json |
Body nesting a Records array (workaround for legacy backend bug) |
3.2 Sample Source Code
| Fixture File |
Description |
code_button.tsx |
Simple Tailwind button |
code_modal.tsx |
Modal dialog with form |
code_table.tsx |
Sortable data table |
code_page_login.tsx |
Full page (type=0) |
code_with_inline_styles.tsx |
No classes, inline styles only |
3.3 Sample Screenshots
| Fixture File |
Format |
Notes |
button_primary.png |
PNG |
320x80, normal case |
button_primary.jpg |
JPEG |
Same scene, with compression artifacts |
large_image.png |
PNG |
4 MB, near upload limit |
corrupt.png |
โ |
Truncated byte sequence |
3.4 Sample LLM Responses (Recorded)
Recorded with pytest-recording (or hand-written) under fixtures/llm_responses/. Used to deterministically replay agent outputs in CI.
| Fixture |
Input Target |
Return Value |
semantic_button.json |
Code Analysis LLM (arch node C) |
["submit","action","primary","cta","button"] |
context_button.json |
Deterministic source parser |
component level/family/variant/import/export/reference/class/token/text evidence |
embedding_512.json |
Configured embedding provider |
2 float vectors of 512 dimensions |
4. Test Catalog
4.1 Unit Tests โ apps/code-import/__tests__/unit/
Schema Validation (test_schemas.py)
| ID |
Purpose |
Input |
Expected |
| U-SCH-001 |
Accept required fields only |
record_minimal.json body |
CodeImportMessage instance, css_code=None, img_url=None |
| U-SCH-002 |
All fields assigned correctly |
record_full.json body |
All fields match expected values |
| U-SCH-003 |
Reject invalid UUID |
record_invalid_uuid.json body |
ValidationError on code_id |
| U-SCH-004 |
Reject missing required field |
record_missing_field.json body |
ValidationError on name |
| U-SCH-005 |
Reject out-of-range type |
record_invalid_type.json body |
ValidationError on type |
| U-SCH-006 |
Reject out-of-range based_on |
based_on=3 |
ValidationError |
| U-SCH-007 |
Reject empty source_code |
record_empty_source_code.json |
ValidationError |
| U-SCH-008 |
Optional css_code accepts None |
minimal + css_code=None |
OK |
| U-SCH-009 |
Optional img_url accepts None |
minimal + img_url=None |
OK |
S3 URL Parsing (test_s3_helpers.py)
| ID |
Purpose |
Input |
Expected |
| U-S3-001 |
Parse s3:// URL |
s3://bucket/key/file.png |
("bucket", "key/file.png") |
| U-S3-002 |
Parse virtual-hosted HTTPS |
https://bucket.s3.us-east-1.amazonaws.com/file.png |
("bucket", "file.png") |
| U-S3-003 |
Parse path-style HTTPS |
https://s3.us-east-1.amazonaws.com/bucket/file.png |
("bucket", "file.png") |
| U-S3-004 |
URL-decode key |
s3://bucket/%E3%83%95%E3%82%9A%E3%83%AD.png |
Key contains ใยฐใญ.png (unquoted) |
| U-S3-005 |
Reject unsupported scheme |
ftp://bucket/key |
ValueError |
| ID |
Purpose |
Input |
Expected |
| U-IMG-001 |
Detect PNG URL |
canonical .png URL |
"png" |
| U-IMG-002 |
Detect JPEG URL |
canonical .jpg URL |
"jpeg" |
| U-IMG-003 |
Detect GIF URL |
.gif URL |
"gif" (helper compatibility; worker canonical validation rejects it) |
| U-IMG-004 |
Detect WebP URL |
canonical .webp URL |
"webp" |
| U-IMG-005 |
Unknown extension fallback |
URL without known extension |
"png" |
Embedding Helpers (test_embedding.py)
| ID |
Purpose |
Input |
Expected |
| U-EMB-001 |
2 inputs produce 2 vectors of 512 dimensions |
Mock embedding client |
Tuple of 2 list[float] of length 512 |
| U-EMB-002 |
Single batch call |
Client spy |
Exactly 1 API call |
| U-EMB-003 |
Semantic input preserves component identity |
name/kind/context/agent terms |
deterministic identity string |
Agent Output Models (test_agent_outputs.py)
| ID |
Purpose |
Input |
Expected |
| U-AGT-001 |
CodeSemantics accepts 5 words |
List of 5 strings |
OK |
| U-AGT-002 |
CodeSemantics accepts 10 words |
List of 10 strings |
OK |
| U-AGT-003 |
CodeSemantics rejects 4 words |
List of 4 strings |
ValidationError |
| U-AGT-004 |
CodeSemantics rejects 11 words |
List of 11 strings |
ValidationError |
| U-AGT-005 |
CodeSemantics strips blanks before length validation |
mixed blank/nonblank terms |
canonical bounded list |
4.2 Contract Tests โ __tests__/contract/
| ID |
Purpose |
Input |
Expected |
| C-WH-001 |
Success Webhook matches OpenAPI-style schema |
Success payload from worker |
Conforms to ai-status webhook schema |
| C-WH-002 |
Failure Webhook matches schema |
Failure payload from worker |
Conforms |
| C-WH-003 |
Webhook type is always "code-import" |
Any Webhook |
type == "code-import" |
| C-WH-004 |
Webhook recordId exactly echoes code_id |
Success path |
recordId == code_id |
| C-WH-005 |
Webhook scope echoes SQS tenant/project |
Success and failure payloads |
organizationId == organization_id and projectId == project_id |
| C-DOC-001 |
DocDB document has strict visual metadata, semantics, context, and processing blocks |
Sample run |
Extra legacy fields rejected; no visual vector |
| C-DOC-002 |
semantic/context vectors are exactly 512 |
Sample run |
Both array lengths == 512 |
4.3 Integration Tests โ __tests__/integration/test_process_record.py
All tests use mongomock, moto, respx, and recorded LLM responses. Each test calls process_record(record) end-to-end.
Happy Path
| ID |
Purpose |
Fixture |
Expected DB Write |
Expected Webhook |
| I-HP-001 |
Minimal payload, no screenshot, no CSS |
record_minimal.json |
css_code=null, visual.metadata.image_url=null, semantic/context embeddings 512 dimensions |
success, keywords length 5โ10 |
| I-HP-002 |
All fields with screenshot |
record_full.json |
visual.metadata.image_url == img_url |
success |
| I-HP-003 |
With CSS, no screenshot |
record_with_css.json |
css_code saved, visual.metadata.image_url=null |
success |
| I-HP-004 |
Page (type=0) |
record_page.json |
type=0 saved |
success |
| I-HP-005 |
based_on=1 (design) |
record_based_on_design.json |
based_on=1 saved |
success |
| I-HP-006 |
based_on=2 (wireframe) |
record_based_on_wireframe.json |
based_on=2 saved |
success |
| I-HP-007 |
Canonical S3 reference |
configured bucket/exact preview key |
Reference accepted; no image download |
success |
| I-HP-008 |
Idempotent reprocessing |
Run record_full.json twice |
Second run reuses fingerprint/vectors; paid work skipped |
Success Webhook twice |
| I-HP-009 |
Double-wrapped body (legacy) |
record_double_wrapped.json |
Inner body extracted, normal processing |
success |
Failure Path
Repository regression tests cover transient write conflicts: retry the complete database transaction with a fresh session, preserve atomic index invalidation, and do not repeat AI calls. Permanent failures are not retried locally; retry exhaustion remains an SQS record failure.
| ID |
Purpose |
Fixture / Mock |
Expected Behavior |
| I-FP-001 |
Empty body string |
{"body":""} |
ValueError, no Webhook (no code_id), re-raise exception |
| I-FP-002 |
Invalid JSON body |
{"body":"{not json"} |
JSONDecodeError, no Webhook, re-raise |
| I-FP-003 |
Pydantic validation failure |
record_invalid_uuid.json |
ValidationError, partial Webhook attempt (best-effort), re-raise |
| I-FP-004 |
Wrong preview key |
Reference does not equal canonical key |
Webhook failed, record failure |
| I-FP-005 |
Wrong preview bucket |
Reference does not target S3_BUCKET_NAME |
Webhook failed, record failure |
| I-FP-006 |
LLM semantic timeout |
Semantic agent raises AgentRunError |
Webhook failed, re-raise |
| I-FP-007 |
Invalid embedding response |
Missing or wrong-sized second vector |
Webhook failed, record failure |
| I-FP-008 |
Embedding API rate limit |
openai.RateLimitError |
Webhook failed, re-raise |
| I-FP-009 |
DocumentDB write failure |
mongomock raises PyMongoError |
Webhook failed, re-raise |
| I-FP-010 |
Webhook itself fails (5xx) |
respx returns 500 |
Log and re-raise so SQS retries |
| I-FP-011 |
Webhook 4xx (bad auth) |
respx returns 401 |
Log and re-raise so SQS retries |
Edge Cases
| ID |
Purpose |
Fixture |
Expected |
| I-EC-001 |
Source code at 200 KB limit |
record_huge_source_code.json |
Accepted, processed |
| I-EC-002 |
Source code with non-ASCII (Japanese, Vietnamese) |
Unicode source_code |
Correctly round-trips to DocDB |
| I-EC-003 |
Preview is JPEG |
canonical .jpg reference |
Stored unchanged; no download |
| I-EC-004 |
Preview is WebP |
canonical .webp reference |
Stored unchanged; no download |
| I-EC-005 |
Preview URL changes with same fingerprint |
same source, new canonical extension |
Only image URL patched; index remains valid |
| I-EC-006 |
2 records in 1 SQSEvent |
Batch of 2 records |
Both processed, attempt second even if first fails |
4.4 Smoke Tests โ __tests__/smoke/
Run against deployed dev Lambda only. Marked with @pytest.mark.smoke.
| ID |
Purpose |
Steps |
Expected |
| S-001 |
Real end-to-end (with preview) |
1. Send record_full.json to SQS. 2. Wait โค 3 minutes. 3. Query DocDB by scope + code_id. 4. Query PG code row. |
DocDB has semantic/context vectors and processing hash; PG status completed |
| S-002 |
Real end-to-end (no screenshot) |
Send record_with_css.json. Same wait + verification. |
Same as above, visual.metadata.image_url=null |
| S-003 |
Real end-to-end (failure path) |
Send payload with invalid preview reference. |
PG row failed within 3 minutes; webhook log shows failed |
4.5 Load Tests โ __tests__/load/
Run before each release. Marked with @pytest.mark.load.
| ID |
Purpose |
Profile |
Pass Criteria |
| L-001 |
Steady throughput |
10 msg/min for 30 minutes |
p50 < 90 s, p95 < 180 s, error rate < 1% |
| L-002 |
Burst processing |
100 msg in 1 minute, then idle |
All processed within 10 minutes, no DLQ messages |
| L-003 |
Sustained burst |
30 msg/min for 1 hour |
No DLQ, no OpenAI rate limit exceeded |
5. Coverage Goals
| Layer |
Line Coverage |
Branch Coverage |
Notes |
apps/code-import/src/code_import/handler.py |
โฅ 90% |
โฅ 85% |
Lambda entry (lambda_handler) โ must cover all try/except branches |
apps/code-import/src/code_import/schemas.py |
100% |
100% |
Pydantic models โ test exhaustively |
packages/agentic/agents/code_semantics.py |
โฅ 80% |
โ |
Builder only, LLM runtime is mocked |
packages/des2code-core/code_context.py |
โฅ 90% |
โฅ 85% |
Deterministic parser/identity text |
| shared embedding builder |
100% |
โ |
Provider selection and one two-input call |
packages/guinness-ai-sdk webhook |
โฅ 90% |
โฅ 85% |
Including retry path |
packages/models/documentdb/code.py |
โฅ 85% |
โ |
Index creation idempotency |
Run:
pytest --cov=code_import --cov=packages \
--cov-report=term-missing --cov-fail-under=85
6. CI Integration
Per-PR Pipeline (.github/workflows/test-code-import.yml)
on:
pull_request:
paths:
- 'apps/code-import/**'
- 'packages/agentic/agents/code_*.py'
- 'packages/helpers/**'
- 'packages/models/documentdb/code.py'
jobs:
test:
steps:
- uv sync
- pytest -m "not real_openai and not smoke and not load" \
--cov --cov-fail-under=85
- ruff check apps/code-import packages
- mypy apps/code-import packages
Nightly (against dev)
on:
schedule:
- cron: '0 18 * * *' # 03:00 JST
jobs:
smoke:
steps:
- uv sync
- pytest -m smoke
Pre-Release (against staging)
A manually triggered workflow that runs -m "smoke or load".
Failure Policy
| CI Stage |
On Failure |
| Lint / type / unit / contract / integration |
Block merge |
| Below coverage threshold |
Block merge |
| Nightly smoke |
Notify #guinness-ai Slack, open issue after 3 consecutive failures |
| Load tests |
Block release, do not block merge |
7. Quick Test Matrix (Cheat Sheet)
| Scenario |
Test IDs |
| Required fields only |
U-SCH-001, I-HP-001 |
| All optional fields present |
U-SCH-002, I-HP-002 |
| Page vs code |
I-HP-004 |
All based_on variants |
I-HP-005, I-HP-006 |
| Screenshot missing |
I-HP-001, I-HP-003 |
| Invalid input (validation) |
U-SCH-003 .. U-SCH-007, I-FP-003 |
| External service failures |
I-FP-004 .. I-FP-009 |
| Webhook retry behavior |
I-FP-010, I-FP-011 |
| Idempotency |
I-HP-008 |
| Real-world end-to-end |
S-001, S-002 |
| Throughput |
L-001, L-002 |
8. Definition of Done (Per PR)
A code-import PR is mergeable when:
- [ ] All applicable test cases are implemented and passing.
- [ ] Coverage thresholds met (see ยง5).
- [ ] At least one new integration test added for any behavioral change.
- [ ] If SQS schema, Webhook schema, or DocDB document shape changed, manually verify
S-001โS-003 smoke tests on dev.
- [ ] PR description links to the relevant test IDs.
For repository-wide release readiness, see docs/overview/non-functional.en.md (SLA / SLO targets).