AI Des2WF — Test Case Design
Companion to the I/O Definition; a contract change lands in both before code
merges. Cases are grouped by what they can be written against today versus what runs once its
input exists.
Test verbatim copy, source geometry, guarded graphic simplification, conservative fallback and
scoped review transport. A registered kit is needed for kit-specific cases; visual recognition
quality additionally needs source PNGs and real-provider runs. These cases describe expectations,
not a claim that every manual or downstream check has already passed.
D-P — Parse (specifiable now)
| ID |
Case |
Expectation |
| D-P-01 |
A design FRAME parses |
tree emitted; node count and depth recorded |
| D-P-02 |
Instances carry identity |
every INSTANCE node carries componentId and its componentProperties verbatim |
| D-P-03 |
Named styles survive |
every style reference is captured as a publish key; zero dangling keys |
| D-P-04 |
Determinism |
the same snapshot parsed twice is byte-identical; no LLM in the parse path |
| D-P-05 |
Cache key excludes position |
moving the frame on the canvas does not invalidate the parse — the WF2Des parse hash covered absoluteBoundingBox and a pure canvas move invalidated a byte-identical extraction |
| D-P-06 |
Snapshot maps retained |
Capture stores referenced components/componentSets/styles; verify publish identity survives a scoped fetch |
| D-P-07 |
Masters are not read by the tree extractor |
reading a COMPONENT_SET through the tree extractor is rejected; derive_component_context returns the variant partition |
| D-P-08 |
Hidden flow state |
Invisible content draws no paint; zero-opacity children that occupy auto-layout space retain transparent source-layout slots |
| ID |
Case |
Expectation |
| D-V-01 |
Not a FRAME (COMPONENT, GROUP, page) |
invalid_input.not_a_frame, no partial write |
| D-V-02 |
Frame with no visible authored descendants |
invalid_input.empty_frame |
| D-V-03 |
Snapshot above the byte cap |
invalid_input.snapshot_too_large; the cap is enforced on the des2wf capture path specifically |
| D-V-04 |
Wireframe submitted as design |
No design-ness detector; provenance is not itself a downstream enforcement guarantee (see D-O-03) |
D-C — Control plane (specifiable now)
| ID |
Case |
Expectation |
| D-C-01 |
Row before enqueue |
the row exists with edge=des2wf before the message is sent |
| D-C-02 |
Worker never writes PostgreSQL |
no PG credentials in the worker; asserted structurally |
| D-C-03 |
Explicit dispatch |
DES2WF messages retain required edge=des2wf; generation and kind=render_review use the dedicated DES2WF queue |
| D-C-04 |
Attempt fencing |
a stale-attempt webhook is ignored; DocDB commits fenced on (job_id, attempt) |
| D-C-05 |
Terminal re-emit |
a second succeeded for the same attempt is a no-op |
| D-C-06 |
Open-generation fence |
a second run for the same (project, file, node, edge) is rejected while one is open |
| D-C-07 |
Fence release |
Terminal state or explicit cancel frees the key; do not assume a periodic reaper exists |
| D-C-08 |
Idempotency across edges |
a WF2Des job and a Des2WF job never collide in the plugin's jobId → rootNodeId index |
| D-C-09 |
Scoped capture takes the fast path |
with a client-supplied sectionNodeId containing the frame, the capture fetches only that subtree and never issues the whole-file GET /files/{key} — the call measured at 255 MB / 67.5s with a socket-close failure at 140s |
| D-C-10 |
A wrong scope degrades, never corrupts |
a sectionNodeId that does not contain the frame, or names a deleted node, falls back to the whole-file walk with a log line, and the resulting source_hash equals the unscoped capture's — the hash keys the parse cache, so a stale scope must never serve another frame's parse |
D-O — Output document (specifiable now)
| ID |
Case |
Expectation |
| D-O-01 |
WireframeDoc shape |
based_on = 1, type = 0, memos = [], summaries computed |
| D-O-02 |
_id separation |
Des2WF writes …_{node_id}~des2wf, wf_parse writes the unscoped id; the same frame processed by both edges yields two documents, neither overwritten |
| D-O-03 |
Downstream provenance guard (follow-up) |
Verify each consumer before claiming machine-generated based_on=1 wireframes are refused; output provenance alone is not that guard |
| D-O-04 |
Unmatched is visible |
an inexpressible node becomes a placeholder and drives the score gate to 0 |
| D-O-05 |
Artifacts |
Per-job parse/result/manifest exist before terminal success; version-pinned source PNG keys include source and render hashes |
| D-O-06 |
Kit miss falls back to primitives |
a node with no kit expression emits layout_frame/text primitives — never unmatched, never a failure |
| D-O-07 |
Kit swap degrades cleanly |
after a different kit is registered, entries referencing now-missing components are dropped at consult and re-voted; emission degrades to primitives; no crash and no placeholder flood |
| D-O-08 |
Drawn-row totality |
Every drawn row has emitted lineage or explicit composite coverage; confirmed ornament may intentionally emit no paint, retaining flow slots when needed |
| D-O-09 |
Guarded kit replacement |
A kit replacement accounts for copy and independently drawn descendants; this is separate from explicit ornament suppression |
| D-O-10 |
layout: absolute |
Drawn rows use source coordinates under the root; depth gate is 2. Still the API default when layout is omitted, not the current plugin setting |
| D-O-11 |
layout: auto |
Retain native hierarchy and one-child wrappers, itemSpacing, padding, baseline, wrapping and per-axis sizing; do not infer new gaps after flattening |
| D-O-12 |
Both layouts carry the same copy |
one frame emitted absolute and emitted auto yields the same strings the same number of times; the choice changes structure, never content |
| D-O-13 |
A refused render is recorded, not retried |
the design frame and the built wireframe export as a PNG pair scaled so neither edge exceeds the vision model's 8000px limit; a render the model would refuse writes status: "skipped" with the reason and a null finding count, so nothing can read "no findings" off a pass that never ran |
D-S — Score (partly specifiable now)
| ID |
Case |
Expectation |
| D-S-01 |
Content preservation |
every source string appears in the output verbatim and with the same multiplicity — a label the design says three times is emitted three times. Copy that lands in a kit instance's slots counts as emitted |
| D-S-02 |
The metrics are measured against the design |
engine/measure.py computes integrity, flatness, ink_covered and boundary_fidelity from the design tree and the emitted spec, and reports them at weight 0 while they are calibrated. It may not import the emitter, the census or selection — a denominator supplied by the thing being measured is not a measurement, and the import ban is asserted |
| D-S-03 |
Gate |
the score is 0 when any emitted node is unmatched, and when any of strings_invented, text_displaced, cross_over_text or boxes_coincident is non-zero. strings_lost is deliberately not gated — content preservation already grades it, and gating it too collapses the score into a pass/fail bit |
| D-S-04 |
Band, not scalar |
the score is reported with a margin, and the ratchet refuses an improvement inside the noise band |
| D-S-05 |
The number reaches the server |
score and score_version travel to the shared generation row in the manifest's row_effects.wf2des block, which is keyed by the affected table rather than the edge; they are strictly additive, so the WF2Des edge omits them and the same backend reader still applies the row. The full block — gate, band, terms — lives in the result document that row points at |
| D-S-06 |
Model-assisted band |
Margin is 0.07 when kit selection or graphic classification sets llm_assisted, otherwise 0.0; primitives is not a no-model guarantee |
| D-S-07 |
Depth is gated only where flatness is the contract |
the depth term applies to the absolute layout, where flat is what was asked for; an auto wireframe nests on purpose, so its depth is reported as flatness and gates nothing. Measured across the corpus: 6 to 11 |
Cases requiring the kit or the accepted-result wireframe
Written now, runnable once their input exists — listed so the gap is visible rather than absent. A
dash in Requires means the case needs nothing that is not already here.
| ID |
Case |
Requires |
| D-K-01 |
Every emitted component reference resolves in the target file |
the installed, registered kit |
| D-K-02 |
Vocabulary coverage — the share of design components with a target expression |
the chosen kit's inventory |
| D-K-03 |
Variant selection picks the evidenced variant — the selectable unit is a variant, not a component: a component set cannot be instanced, and every variant of a set shares the registry _id, so the vote and the resolution both key on a unique per-variant key |
the installed kit |
| D-K-04 |
De-styling preserves source semantics and typography: neutral literal text and surfaces, no design style bindings; source font/weight/geometry and protected glyph colors can remain; rules and state markers can be filled |
— |
| D-K-05 |
Source layout, copy and guarded image policy: meaningful glyphs stay, approved imagery becomes ink, decoration becomes ornament; no cross over text |
— |
| D-K-06 |
Similarity against a human-made wireframe of the same screen |
an accepted-result wireframe |
| D-K-07 |
Selection guards — a voted atom is rejected by unreachable, copy_truncated, slot_overflow, imagery_heavy, size_implausible or unknown_atom, and the node falls back to a primitive box rather than emitting a bad instance. Slot capacity is the number of distinct slot layer paths a variant declares, since the materializer addresses a slot by path — two declarations of one path are one destination |
the installed kit |
| D-K-08 |
Selection stability: repeat real-provider runs and report variant/graphic differences; majority voting and a score band do not guarantee identical output |
Installed kit and real provider |
| D-K-09 |
Mode split: both modes can call the graphic classifier; only kit adds sampled component-context selection |
Installed kit; stubbed model for call assertions |
| D-K-10 |
Mode default: omitted mode uses primitives, with the same conservative graphic policy |
— |
D-G — Graphics and source geometry
| ID |
Case |
Expectation |
| D-G-01 |
Arrow, action/hamburger, state, SNS/identity, list marker or unknown |
Preserve actual source shape; names or raster/vector format alone never authorize replacement |
| D-G-02 |
Simplification guards |
Only high-confidence, non-information-bearing content_image / redundant_artwork / decoration can simplify; redundant_artwork requires a real local label_node_id |
| D-G-03 |
Button flourish and footer backdrop |
Confirmed decoration emits no crossed placeholder or paint; preserve required flow space and surrounding button/text layout |
| D-G-04 |
Missing/invalid evidence and replies |
Missing/corrupt/oversized/aspect-mismatched PNG, provider failure, incomplete/unknown/duplicate IDs or timeout preserve affected artwork and record reasons |
| D-G-05 |
Batch bounds and coverage |
6 candidates per batch, 3 concurrent calls, 180-second classification budget; no first-120 cap; recursively interleave branches/controls; prepare crops per active batch |
| D-G-06 |
Context and evidence |
Overview, graphic/control crops, labels, variants and ancestors are supplied per appearance; source/evidence hashes and prompt/model recorded |
| D-G-07 |
Overrides |
Scope by organization/project/file; exact component ID and variants; node precedence; conflicting duplicates preserve; simplify override means ink, not ornament; no legacy verdict consumption |
| D-G-08 |
Actual rendered policy |
Audit distinguishes a proposal from a child covered by a preserved source parent or kit instance; counts do not imply independently visible changes |
| D-G-09 |
Source glyph materialization |
Clone actual instance and overrides, proportional scale and x/y centering; clone failure exports actual source; wrong-file/missing source gives visible fallback and diagnostics |
| D-G-10 |
Layout sizing |
Retain one-child wrappers, FIXED/HUG/FILL per axis, min/max, absolute overlays and opacity-zero slots; center sole SPACE_BETWEEN child only with source evidence |
| D-G-11 |
Text geometry |
Preserve font family/weight, line height, spacing, auto-resize, trim, alignment and paragraph geometry; source cloning retains PALT; verify fallback/reflow diagnostics |
| D-G-12 |
Legacy compatibility |
Specs without source_layout/text_geometry retain existing WF2DES behavior |
| D-G-13 |
Canvas naming and plugin settings |
Plugin requests auto; DES2WF canvas uses des2wf · source_name (version fallback); explicit spec.name takes precedence, including CODE2WF |
D-R — Capture and review safety
| ID |
Case |
Expectation |
| D-R-01 |
Version-aligned source evidence |
Export the version from the response that supplied the frame, absolute bounds; PNG key includes source/render hashes |
| D-R-02 |
Capture fallback |
Missing version skips export; export/upload failure retains valid snapshot without PNG URL; no newer unpinned image substitution |
| D-R-03 |
Review row validation |
Require project write access and DES2WF row scope; mismatched body jobId is 400 and non-COMPLETED status is 409, before upload/enqueue |
| D-R-04 |
Source/project changes during generation |
Capture pair remains bound to original source and backend client/project; materialized job ID must match |
| D-R-05 |
Review worker scope |
Reject missing/wrong-tenant/wrong-project result before download, model call or result write; never upsert an unknown result |
| D-R-06 |
Review status and history |
Overlarge pair is skipped with null finding_count; provider failure is not a clean review; retain previous block; generation remains completed |
Verification locations and limits
- AI:
apps/des2wf/__tests__/unit/test_graphic_classification.py, test_graphics.py,
test_source_geometry.py, test_engine.py, test_service.py; shared packages/figma-ir/__tests__.
- Plugin: materializer/source-geometry tests and
des2wfPanel.review.test.tsx.
- Backend:
apps/app/__tests__/unit/services/des2wf.test.ts, des2wf-snapshot.test.ts,
and apps/app/__tests__/unit/repositories/figma-rest-render.test.ts.
- Mocked model tests prove policy and transport safety, not recognition accuracy. The frozen
JSON corpus has no source PNGs; it tests conservative fallback and layout invariants.
- Before claiming visual improvement, run real-provider classification and inspect regenerated
plugin output: arrows/SNS/actions, button decoration, footer backgrounds, icon centering and text.
Compare both source and result; do not infer correctness from a 1.0 score.
- Compare repeated runs and per-screen metrics; a mean must not hide a regression on another screen.
Keep general mechanisms in tests, not client-specific names or IDs.