Skip to content

AI WF2Des — Test Cases

Library fidelity regression gate

  • Reject overfull compound copy even when VARIANT/TEXT property definitions exist; genuine BOOLEAN controls retain their path. Numeric units cannot split alphanumeric identifiers, unrelated container slots, or exact whole-string matches.
  • Test local identical-control state consistency, explicit-state priority, ambiguous conflicts, disabled/state differences, and unrelated groups. Repeated-action reuse requires exact copy/style/state/structure and one compatible captured size variant; conflicting donors and unknown sizes abstain.
  • Unverified textless whole-control substitution preserves the source; verified artwork bindings retain the native path. Bound single-label tabs grow with padding, while fitting labels, compound controls and source clones remain unchanged.

  • Whole-control obligations cover arbitrary sizes, source roles and repeated families: painted owner plus atomic stateful glyph or sole label requires its own decision when compatible native candidates exist. A leaf decision cannot hide an omitted owner; table cells, generic cards, multi-control rows, unknown opaque leaves and candidate-free units remain untouched. Source variants/graphic overrides survive payload construction. Sparse review honors valid enclosing native ownership and explicit fallback, with no increased request budget or automatic guessed match.

  • Style capture includes local tokens and exact remote bindings used by frozen board roots, referenced masters/variants, mixed text runs and paints. Repeated IDs are resolved once, local tokens win, ambiguous remote aliases are omitted, and inaccessible references do not erase readable captures. It must not import libraries, mutate Figma, traverse unrelated pages, or infer bindings from raw colors.

  • A stalled style/master/variable read must not block component-sweep enqueue: each read is bounded to 5 seconds within a 60-second capture budget; retain readable bindings and emit an incomplete-capture runtime message. Tests cover a never-settling read and total-budget exhaustion.

  • Render review permits pinned native compound restyling/reorganization while requiring all source copy, independent actions and states. Correction tests change/remove an owning component or hide its action and restore previously covered source navigation; unchanged or newly proven coverage remains dropped. Sparse merges retain valid ownership, legacy deterministic drops replay safely, and unrelated annotation drops remain unchanged.

  • Footer promotion before independent one/multiple-action strips requires unique exact copy and native-action coverage; arbitrary following prose, real fields, missing glyph evidence and ambiguous optional configurations abstain. Hidden or variant-excluded footer actions never remove source navigation; source wording and strip action count are not hard-coded.

  • Single-brand top bars work with header/navigation roles and unrelated brand copy/IDs. Promotion requires captured brand-media selection and one hash-verified graphic-only header variant. Independent menus/actions/copy, non-brand graphics, ambiguous/absent profiles, unknown visibility, painted extra units, explicit unmatched and parent-specific styling abstain. Complete single-brand proof may reconcile an unstyled compose decision. Native whole-header ownership emits the brand once without claiming atomic wordmark text equivalence.

  • Single-row table headers synchronize mixed composed/native cell heights; partial, spanning, duplicate and reordered tracks reject without mutation. Existing multi-row bodies remain unchanged.

  • Nested table surfaces follow their explicit row only through source-proven, same-bounds transparent single-child wrappers. Direct/multirow tracks grow painted FRAME surfaces, never native icons; inset, padded, multi-child, absolute, foreign, duplicate or unpainted surface paths reject or remain unmarked. Legacy rows omit empty surface metadata.
  • Atomic control height floors preserve hit-area geometry without clipping taller content; non-control sections remain HUG.
  • Partial border edges survive compose, nested frames and materialization; omitted edges retain legacy behavior.
  • Covered geometry is subtracted only from its intersection with a retained sibling gap, unioned to avoid double subtraction. Real whitespace, annotations and overlays remain intact.

  • Same-named components with distinct keys, snapshots or axis names remain independently selectable; keyless names cannot borrow identity. Proven same-source metadata adoption remains supported.

  • Atomic textless media, including secondary media classification, passes scalable-artwork scope checks; non-media compounds remain rejected.
  • Proxy INSTANCE BOOLEAN values and visibility references are available without master definitions. Unknown, ambiguous and nested-scope bindings cannot prove content disappearance.
  • Optional controls may remove sample content, never required source copy. Duplicate navigation requires visible copy and, for a label-plus-icon action, captured native action structure; unrelated artwork is retained.
  • Native-action coverage must retain extra opaque controls inside the button or elsewhere in its wrapper, including INSTANCE leaves with absent/default-false ink flags. Only explicitly empty FRAME/GROUP layout cells may be ignored; missing ink census cannot prove emptiness.
  • Covered native actions in component sets and root components retain their proven-true visibility bindings through ScreenPlan and serialized component_properties. Existing overrides survive; false, unknown, defaultless or ambiguous actions cannot enable coverage, and bound-action deduplication without a serializable owner abstains.
  • Zero-opacity text, roots, ancestors, nested instances and native-action labels/glyphs cannot supply visible copy or coverage, even with a true BOOLEAN binding. Positive nonzero opacity remains eligible visible evidence.
  • Verify two fresh production plugin renders before the nine-screen regression batch. Audit native footer identity, all login-provider icons, copy/state preservation, spacing and overflow. Include a held-out screen; never count synthetic or replay-only output as visual acceptance.

Source-fidelity regression matrix (September 2026)

Case Required evidence
Component snapshots Before paid selection, missing/unsupported/corrupt/hash-mismatched snapshots produce an actionable resync failure. Transient storage errors remain retryable. Integrity verification and inspection have separate bounded byte budgets; a valid large component does not fail merely because its detail response must be paginated.
Filtered flow Dropping the last overlapping sibling produces correctly ordered vertical/horizontal output and aligned gaps, without changing layer paths. Remaining modal overlaps retain coordinates/z-order; explicit auto-layout retains authored order.
Fallback overrides Same-file fallback clones the source instance, preserving hidden parts, text and nested state. Missing/cross-file source IDs cannot clone unrelated nodes.
Graphic override evidence Exact visible painted graphic targets set has_graphic_overrides for explicit fill/stroke/opacity/style overrides, including inherited fill-style overrides. Gray paint alone, text/geometry-only overrides, hidden ancestors, missing targets and malformed entries do not. Bounded traversal terminates; default false is omitted and survives old-artifact parsing. Role assignment carries true evidence without changing tree shape.
Selection fidelity Passive table content cannot become an action. Repeated controls preserve cardinality and distinguishable state. Explicit checked/selected state overrides defaults; incompatible compound replacements compose.
Repeated painted cells Partial or complete compact-badge replacements cannot erase full-height painted table cells. Equal source master/state/surface groups cannot become differently painted variants through text-only overrides. Compatible full-height native cells, legitimate differing source states, inset badges, missing profile evidence and unrelated lists remain untouched; sparse review and direct stitch use the same guard.
Shared table tracks Proven column-major tables share row minimum/intrinsic heights, including a trailing total row. Missing middle cells, duplicate paths, row spans, overlapping columns and ambiguous/padded grids remain unmarked. Unmarked historical specs omit the additive field.

Synthetic assertions and read-only replay are not a substitute for a fresh full plugin render.

Atomic fallback tests cover scope, unknown-component, source-copy visibility, surplus-copy and fidelity rejection, plus explicit/legacy empty composition, at top-level and nested boundaries. A native source instance must retain identity, state and lineage; opaque FRAME/GROUP artwork must retain its glyph source; image plus solid fill must still reach image preservation. Plain text, decomposable instances and proven empty layout cells must keep their existing behavior. A sparse review's unknown component must not overwrite its retained initial choice.

Parse cache regression: unchanged source hashes with a generation_parse@0.3 (or older) stamp must re-extract under @0.4; a current-version cache still bypasses download and paid role labeling. Component and rule processor versions remain unchanged.

Semantic section-layout regression matrix (September 2026)

Use generic source IDs and multiple screen families; no production-screen/node allowlist. This matrix is the acceptance contract for section-layout@1, not a claim of visual parity.

Case Required evidence
Grounded grouping Contiguous direct siblings in an eligible vertical flow form one semantic section; original text, controls, assets and member order survive exactly once.
Invalid scope Foreign/missing/repeated/reordered/noncontiguous members, overlapping groups, and grid/absolute/instance boundary crossings produce visible findings and source-layout fallback.
Post-selection coverage A component replacement, dropped member or changed emitted structure cannot make a formerly valid grouping silently absorb unrelated content.
Layout ownership Component internals and painted padding are unchanged; section joints account for existing edge padding without double insets. Chrome boundaries and intentional source gaps are preserved.
Structural clipping Fully contained explicit vertical wrappers can carry section spacing and native height authority without removing their clip. Overflow/crops, artwork, rounded or fixed viewports and overlays do not qualify. Rule-proven excess neutral edge padding has one joint owner; unruled whitespace and nested independently painted insets remain unchanged.
Rule authority Exact applicable rules govern effective section joints; ambiguous/conflicting policies fall back. Missing rules preserve measured geometry; ungrounded role guesses never invent spacing.
Viewport authority Positive explicit viewport_width_px may govern output width; zero/negative values reject. Advisory base viewport, content width and padding alone never resize the screen.
Component height Eligible instances use actual component height without former-footprint spacers; unmarked instances and nonvertical/absolute layouts retain legacy behavior.
Replay identity Pinned plan plus identical inputs yields byte-identical spec in generation and no-op render review, with no new candidate or state write.
Sparse correction A changed selection revalidates affected section coverage; the candidate persists its reconciled plan and new immutable state hash. Existing one-correction and verification-only limits still hold.
Legacy compatibility A state without layout_plan and a spec without height_authority retain prior replay/materialization behavior. Existing five node kinds remain sufficient.

Exercise deterministic engine tests plus test_render_review.py's pinned-scope/no-op review fixture and production generation tests. Plugin tests must check actual flow sizing and unchanged legacy footprint behavior. Actual Figma renders on held-out wireframes remain required to assess spacing, overflow and visual completeness beyond structural assertions.

Render-review and generalization regression coverage (September 2026)

test_render_review.py / test_render_review_repo.py cover exact job/tenant/attempt and parent identity, immutable spec/state/evidence hashes, PNG budgets/completeness, rooted manifest validation, sparse state/property merging, the one-correction bound, style-only verification failures, unavailable providers, stale runs, duplicate claims and fenced writes. Backend tests independently cover authentication/scope and evidence upload/queue contracts. Plugin tests cover original-frame retention, adjacent candidate placement, review identity, cancellation, polling failure, verification-only mode and reinforcement gating.

Generic counterexamples (not production node IDs) cover vector-only components, stale placeholder inference, empty vs absent variant slots, source text/assets surviving compose, explicit/inferred reading order, typography/wrapping, captured optional BOOLEAN visibility, and exact variant inspection with pagination and untrusted snapshot text. The replay harness keeps registry/rule/source pins and evidence provenance, does not write jobs or decision ledgers, and can test held-out wireframes. These tests and text/registry counts do not establish visual parity: actual Figma renders remain the acceptance evidence.

Additional counterexamples cover file-local node-ID collisions, unknown source-file identity, published-key resolution across files, and geometry inference preserving every explicit variant axis. A missing text slot or unresolved warning cannot pass even when the model says pass; successful literal-style preservation remains informational. Wordmark bindings must reject hidden artwork, visibility overrides, blank sources, changed components and non-brand copy while preserving the exact represented text and its evidence hashes.

Control-artwork bindings require actual source/render images, successful exact-variant inspection, painted artwork and one exact source control label. Synthetic non-brand lettering is supported; component names/shape alone, hidden ink, editable text, extra labels/controls, unknown ink census, changed hashes/variants and nonvisual requests must retain source copy. Accepted bindings preserve copy metadata, collapse only the generated sole-instance with_text wrapper, and remain stable on pinned replay; unrelated table/painted wrappers and source trees must not be changed. New lineage defaults to schema 1.4, while explicitly stored older version strings remain readable.

Multi-node correction findings use synthetic, unrelated regions: explicit affected_node_ids must permit all reported regions, preserve singular node_id behavior, and reject root/foreign IDs as scope expansion. IDs mentioned only in detail confer no authority, informational findings confer none, and malformed/oversized ID lists fail schema validation. Verification must not call any explicitly reported original region a new regression, while a current finding mixing old and unrelated regions is still a regression. Legacy unscoped diagnostic categories remain comparable without widening node scope.

This page is the test plan for apps/wf2des/. Every generation and internal-event contract, DocumentDB/S3 write, webhook effect, failure row, and fence must map to a case below.

Convention: Test IDs follow the format WF-<layer>-<NN>. Layers are U (unit — pure functions, no I/O; PydanticAI / Figma REST / webhook faked at the SDK boundary), C (component — end-to-end process_record per run family against real localstack S3 + a testcontainer Mongo / DocumentDB-compatible image, with the LLM [PydanticAI], Figma REST, and the ai-status webhook mocked at the SDK boundary), and E (end-to-end — real dev cloud, smoke only). Each case's Expected column is its assertion contract: anything not written there is not verified.


Test Layers & Tooling

These cases are a manual / local plan run before merging or as a release checklist; they are not an automated CI gate.

Layer Scope Tooling
Unit (U) Pure functions: message-schema validation, deterministic WFNode extraction, memo-status derivation, screen_id grammar + variation_label split, role/kind candidate gate, decision-ledger consult + score-ratchet write-back, K-sample self-consistency vote, deterministic selection guards, stitching, rules validator, deterministic rule-fragment merge, computed confidence, content_hash canonicalization, manifest + webhook payload builders pytest, no network. PydanticAI / Figma REST / httpx webhook faked at the SDK boundary
Component (C) End-to-end process_record per run family. Real localstack S3 + testcontainer Mongo (DocumentDB-compatible image); mock PydanticAI agents (role/intent, section-selection, per-board rule-extraction), mock Figma REST, mock ai-status webhook httpx client pytest, testcontainers, localstack, respx / unittest.mock
End-to-end (E) Real dev AWS — real generation SQS + wf2des-events queue + EventBridge, S3, DocumentDB, LLM dev key, real backend ai-status webhook / PostgreSQL pytest @e2e marker, smoke only

Fixtures

build_parse_msg(...) and build_assemble_msg(...) are the single sources of truth for the two generation payload shapes — always synthesize from them rather than hand-writing dicts.

Fixture Purpose
build_parse_msg(**overrides) Well-formed generation parse SQS message (wf2des_id, attempt, nonce, wireframe_img_url, wireframe_json_url, snapshot_scope, wf_content_hash, screen_id, optional pin URLs, auto_confirm)
build_assemble_msg(**overrides) Well-formed generation assemble message (wf2des_id, attempt, fresh nonce, confirm{attempt, parse_artifact_hash})
event_rule_upload(...) wf2des-events rule-upload event (event_type, design_rule_id, figma_file_key, board_node_ids[] — the selected guideline nodes: board frames or a CANVAS/SECTION container expanded by the worker, min 1 — org/project)
event_frame_reg(...) wf2des-events plugin frame-registration event (event_type, figma_file_key, node_id, snapshot URL, org/project)
event_resync(...) wf2des-events plugin resync event routing to component_sweep
eventbridge_invoke() Bodyless EventBridge schedule invocation for component_sweep
wireframe_snapshot_s3 A WF snapshot JSON in S3 (frame + enclosing board section, memo nodes, a 取込済 ("imported") status frame, hidden layers, image-fill rects, vector leaves)
figma_component_json A Figma COMPONENT / COMPONENT_SET REST JSON for the sweep walk (published + local, variant props, text/image slot layers)
figma_rule_boards(...) Mocked Figma REST for rule_process: get_file_nodes JSON carrying the selected guideline board titles + a PNG render per board behind the images API
component_registry(...) design_component context docs + the pinned platform design (type=component) set — instancing context by variant_properties, slots, default_size
design_resolution_doc(...) A design_resolution decision-ledger doc keyed _id = {org}_{project}_{section_signature} — {case, platform_design_id, component_name, variant_policy, provenance (mined \| voted), score}; mined and voted variants at controllable scores
design_rule_doc(...) A compiled design_rule revision doc (merged spacing / typography / colors / component_specs / screen_group_policies + boards[] board-image refs) keyed (design_rule_id, content_hash); plus a no-rules variant
config_json(design_area=...) Per-project config.json in S3 (referenced by project_figma_file.config_url) — a known DESIGN-area rect (design_area) + memo markers (memo_resolved_markers) + optional variation_separator + optional annotation_section_markers
mock_llm(role_intent=..., selection=..., rule_extraction=...) Structured-output mock for the three LLM steps: role/intent, section-selection, per-board rule-extraction — never returns a tree shape or a confidence self-report
mock_ai_status(http_status=200) Captures the outbound POST /v1/webhooks/ai-status payload (envelope + manifest key) and returns the given status

Coverage Matrix

Rows = contract areas from the I/O Definition. Columns = outcomes. Each cell names one or more test IDs; a — means the outcome does not apply to that area.

Contract area Happy path Failure path Idempotency / Fencing
Generation message schema (parse + assemble) WF-U-01, WF-U-02 WF-U-03, WF-U-04 WF-U-05
WFNode extraction & summaries WF-U-06, WF-U-07, WF-U-08, WF-U-09, WF-U-10 — —
Memo status / screen_id / variation_label WF-U-11, WF-U-12, WF-U-13, WF-U-14 — —
Parse pipeline (extraction + LLM roles/intent + parse cache) WF-C-01, WF-C-01a, WF-C-02, WF-C-04 WF-C-03, WF-C-03a, WF-C-04a, WF-C-04d, WF-C-05 WF-C-01b, WF-C-02, WF-C-11a, WF-U-06
Confirm / reject transitions WF-C-04b, WF-C-10, WF-E-01 WF-C-05a, WF-C-11b WF-E-05
Candidate set / stitch / validator / confidence WF-U-15, WF-U-16, WF-U-17, WF-U-18, WF-U-19, WF-U-20 — —
Selection trust boundary + project config WF-U-29, WF-U-30 WF-U-29, WF-U-31, WF-U-32 —
Role-assignment fence + terminal-surface guards WF-U-34 WF-U-35 WF-U-35
Section-parallel selection WF-U-36 WF-U-36 —
Decision ledger / K-sample vote / deterministic guards WF-U-37, WF-U-39, WF-U-40, WF-C-06f WF-U-40 WF-U-37, WF-U-38, WF-C-06g
Five spec node cases WF-U-21, WF-U-22, WF-U-23, WF-U-24, WF-U-25 — —
Assemble (rehydrate pins + candidate set + ledger consult + K-sample selection + guards + stitch + validator + confidence + ledger write-back) WF-C-06, WF-C-06b, WF-C-06c, WF-C-06d, WF-C-06e, WF-C-06f, WF-E-02 WF-C-07, WF-C-08, WF-C-06h, WF-C-06i WF-C-06a, WF-C-06g, WF-C-11
Result artifact + manifest + ai-status webhook WF-C-06, WF-C-09a, WF-U-27, WF-U-28 WF-C-08, WF-C-09 WF-C-11, WF-E-04
wf_parse (wireframe doc only) WF-C-12 WF-C-13 WF-C-14
rule_process (immutable revision per hash) WF-C-15, WF-C-15a, WF-U-26, WF-U-33 WF-C-15a, WF-C-16 WF-C-17
component_sweep (+ registry manifest, single-flight) WF-C-18, WF-C-20a, WF-E-03 WF-C-20 WF-C-19
Figma OAuth credential (Admin token provider, refresh, fallback) WF-C-28, WF-C-34 WF-C-32, WF-C-33, WF-C-35 WF-C-29, WF-C-30, WF-C-31
Figma OAuth connect flow (Admin API authorize + callback) WF-C-36 WF-C-37, WF-C-38, WF-C-39, WF-C-40 —
Snapshot scoping (sectionNodeId fast path + fallback) WF-C-41 WF-C-42 —
Selection guard tiers (kind vs size gating) WF-C-43, WF-C-45 WF-C-44, WF-C-46 —
project_figma_file self-registration WF-C-47 WF-C-47 —
Leftover-text preservation WF-C-48 — WF-C-48
Intake dispatch (event_type → run family) WF-C-27 — —
Lineage / provenance envelope WF-C-26 — —
Tenancy (organization_id + project_id) WF-C-24 — WF-C-24
Workers never write PostgreSQL WF-C-25 WF-C-25 —

Unit (U) Cases

Pure Python, no network. LLM / Figma / webhook faked at the SDK boundary.

Message schema (schemas/)

ID Intent Expected
WF-U-01 Accept a well-formed parse message Valid model; wf2des_id, attempt, nonce, wireframe_img_url, wireframe_json_url, snapshot_scope{frame_node_id, section_node_id}, wf_content_hash, screen_id, auto_confirm all present; optional pin URLs default absent
WF-U-02 Accept a well-formed assemble message Valid model; confirm{attempt, parse_artifact_hash} present; no snapshot/URL fields (all rehydrated later)
WF-U-03 Reject a parse message missing any required field Parametrized over each required field → ValidationError; the field name appears in the error
WF-U-04 Reject a bad snapshot_scope / non-s3:// URL ValidationError on malformed snapshot_scope and on a wireframe_json_url that is not an s3:// URL
WF-U-05 Open-generation dedupe shape The (job_id, attempt, nonce) tuple derivable from the message is well-formed; attempt is a monotonic integer; nonce is a per-send token distinct across two build_parse_msg calls

Deterministic WFNode extraction (figma.py)

ID Intent Expected
WF-U-06 Extract the WFNode tree from a snapshot Every visible FRAME / GROUP / INSTANCE / TEXT descendant + image-fill rectangle → a WFNode; the tree shape is byte-stable across two runs of the same snapshot
WF-U-07 Vector / shape leaves collapse into parent A vector / non-image shape leaf produces no WFNode; it collapses into its parent node
WF-U-08 Hidden layers skipped A layer with visibility off yields no WFNode
WF-U-09 Image-fill rectangles become nodes A rectangle with an image fill becomes a WFNode; a plain fill rectangle collapses
WF-U-10 Derived summaries summaries.element_type_histogram + summaries.node_count computed from the tree, deterministic

Memo status, screen_id, variation_label

ID Intent Expected
WF-U-11 Memo status is deterministic A memo inside a 取込済 frame → status=resolved; a resolved memo is excluded from live intent; status is not an LLM output
WF-U-12 screen_id grammar match [A-Z]{2,3}_[A-Z0-9]+ unanchored, searched frame → board → page name; first match returned
WF-U-13 screen_id grammar miss No match → screen_id=null and an ingest issue is logged to staging/issues.json; not an error (wf_parse)
WF-U-14 variation_label split 会員登録TOP|案1 → 案1; deterministic frame-name split, not LLM

Deterministic candidate set, stitching, validator, confidence

ID Intent Expected
WF-U-15 Deterministic role/kind candidate gate Registry is tenant-scoped, deduped, and hygiene-collapsed; with the production-default gate, every section receives only role-admissible kinds. Size remains a signal, not a gate; no similarity/vector machinery is invoked; counts/truncation are recorded
WF-U-16 Deterministic stitching Given fixed per-section decisions, the spec tree + spec_nodes_flat + style_bindings map are reproduced byte-for-byte; style_bindings derived FROM project_figma_file.style_captures
WF-U-17 Validator auto-fix / flag A safe violation is auto_fixed; an unsafe one is flagged; validator_report.status = ok, violations[] carry {rule, layer_path, action, detail}
WF-U-18 Validator no-rules path No rules registered → no-op; validator_report.status = no_rules; every node flagged rules_unvalidated
WF-U-19 Confidence is COMPUTED, never self-report Confidence derives from validator violations + unmatched/composed counts + slot ambiguity + component-fit; a model-supplied confidence field is ignored; formula_version stamped
WF-U-20 Compose capped below instance For otherwise-equal inputs, a compose node's computed confidence is strictly below the equivalent instance node's

The five spec node cases

ID Intent Expected
WF-U-21 layout_frame node Emits {auto_layout{direction,gap,padding,sizing}, lineage_wf_node_ids, children}
WF-U-22 instance node Emits component_key, platform_design_id (= design_component._id, the stable validator key), variant_props, text_slots, bbox, source{kind: registry \| wireframe} — source.kind = wireframe is the FALLBACK (the wireframe's own componentId + variants), always flagged, never worse than the wireframe
WF-U-23 compose node ALWAYS flagged flagged == true unconditionally; source{kind:composed}; children inlined; source fill/stroke/radius/clipping retained
WF-U-24 text node Emits verbatim content, grounded literal typography and per-field variable bindings. Captured vocabulary alone never assigns a global style_token; explicit text styles remain supported and type-checked at materialization
WF-U-25 unmatched never dropped Emits a placeholder{role,text,bbox}, flagged=true, source{kind:none}; the node is never silently dropped

content_hash canonicalization + builders

ID Intent Expected
WF-U-26 Rule content_hash canonicalization sha256 over canonical JSON with sorted keys, computed over the MERGED RuleSet; two byte-different-but-canonically-equal merged rule sets hash equal — identical extractions converge (rule_process)
WF-U-27 Result manifest builder Builds {manifest_schema_version, job_id, attempt, nonce, job_type:"generation", result_url, flag_count, row_effects.wf2des{status, result_url, flag_count}}; job_type distinct from the envelope type
WF-U-28 ai-status envelope builders parse_done envelope carries no result_manifest_url; succeeded/failed carry it; discriminator type = "wf2des"; manifest_schema_version present; error present on failed only

Selection trust boundary + project config

ID Intent Expected
WF-U-29 Selection echo sanitized to registry truth A hallucinated variant axis/value, or a text_slots/image_slots key not among the registry doc's text_slots/image_slots layer paths, is DROPPED and the instance is FLAGGED (any drop ⇒ flag; a faithful echo passes unflagged); an invented platform_design_id (not among the section's candidates) downgrades the whole section to unmatched
WF-U-30 Project config overrides reach the pipeline A config.json override table (memo_resolved_markers, variation_separator, annotation_section_markers) replaces the platform default and reaches the deterministic derivations (memo status, variation_label split, annotation-section detection); the pipeline itself stays design-system agnostic (no client vocabulary hard-coded). Empty style_captures flags missing configuration without minting an unresolvable text-style stub. Captured tokens never imply one style for every text node
WF-U-31 config.json absent vs broken ABSENT config ⇒ {} (platform defaults: markers, separator, annotation-section markers); present-but-malformed (invalid JSON / non-object / bad override value such as a non-list memo_resolved_markers or empty variation_separator) RAISES — never a silent fallback to defaults
WF-U-32 Validator never mints unbound or wrongly typed tokens A typography auto-fix never swaps to an uncaptured token or variable ID. A variables-only vocabulary produces one unvalidated-typography finding, not one false violation per text. Numeric font scales accept variable IDs only; styles with numeric-looking names are excluded. Legacy variable-valued whole-text references are removed while literals and per-field bindings survive. Plugin binding rejects mismatched style fields and FLOAT/STRING typography types without mutation.
WF-U-33 Deterministic rule-fragment merge merge_rule_fragments is a pure function over the per-board RuleSet fragments, boards processed sorted by node_id: spacing.scale and the allowed colors are UNIONED and sorted, scalar fields (section_gap) take the first non-null value, typography is first-occurrence-wins per style_token; empty input yields the default RuleSet
WF-U-34 Role-assignment fence The LLM fills role/intent ONLY: an id the model failed to echo degrades to role="other" (logged), a minted id adds NO node, and the tree shape/bbox/text are byte-identical to the deterministic extraction
WF-U-35 Terminal-surface guards A duplicate delivery of a terminal attempt re-emits manifest + webhook from the STORED doc without re-running (parse and assemble); a failure terminal re-emits nothing; a superseded attempt skips with no stale surface; an assemble scope mismatch fails closed with the -failed.json + failed-manifest + failed webhook surface
WF-U-36 Section-parallel selection Sections run concurrently on one event loop (observed overlap — wall-clock is the slowest section, never the sum), bounded by SECTION_SELECTION_CONCURRENCY; a failing section lets its siblings SETTLE (no cancellation) and is logged with its section ids before the transient re-raise (SQS redelivery, no terminal surface)

Decision ledger, self-consistency vote, deterministic guards

The design_resolution ledger is the AUTONOMOUS quality ratchet — no human-correction loop exists, so quality converges machine-side. The section_signature is STRUCTURAL and LLM-free (sha256 over kit-component names, normalized name, child-type shape, text count), so the same section KIND hits the same entry on every screen and rerun.

ID Intent Expected
WF-U-37 Decision-ledger consult LOCK vs FLOOR Consult by structural section_signature, applied BEFORE the guards: a mined entry OR an entry with score ≥ the trust threshold LOCKS the section (no LLM call, decision taken verbatim); a voted entry below threshold is a FLOOR — the section re-votes and the strictly-higher score wins (a lower re-vote keeps the floor); a default-variant instance entry is scope-revalidated at consult, and a stale/mis-scoped entry falls through to the vote
WF-U-38 Ledger write-back + score-ratchet fence Every non-unmatched section (and compose-child) decision upserts design_resolution with a DETERMINISTIC score; the RATCHET filter lands a voted write ONLY over a held entry that is NOT mined AND scores STRICTLY lower; a voted write is REJECTED over an equal-or-higher entry and over any mined entry; a mined write is UNFENCED; a DuplicateKeyError on the fenced upsert means KEPT (returned as success, never surfaced as an error); identical evidence scores identically ⇒ reruns never churn updated_at
WF-U-39 K-sample self-consistency vote The section-selection step draws K samples and takes the MAJORITY on (case, platform_design_id); the tallied winner (not any single sample) is the section decision; a tie is arbitrated by the deterministic guards, never by picking the first sample; K is read from config and the vote is a pure function of the K structured outputs
WF-U-40 Deterministic guards over the vote Over the voted decision (echo never trusted): variant_props CLAMPED to the registry axes before anything reads them; symmetric SCOPE guards (a component grossly larger than the section → unmatched; a small component claiming a multi-unit section → compose so its units re-select); duplicate-claim dedupe on wireframe-name evidence; kit-name override (a section that IS one named kit element takes the registry component of that name, carrying the kit's own variant values). Guards run AFTER any ledger floor and keep the final say

Component (C) Cases

Real localstack S3 + testcontainer Mongo. LLM (PydanticAI), Figma REST, and the ai-status webhook mocked at the SDK boundary. Grouped by run family; each case drives process_record.

Generation — Parse

ID Intent Expected
WF-C-01 Cache-MISS happy path Pins inputs FIRST; parse-cache lookup by composite _id = {project_id}_{figma_file_key}_{node_id} misses; downloads the WF snapshot; deterministic extraction; mock_llm.role_intent assigns role + intent only; writes the wireframe doc (LWW) + result-doc parse block + immutable parse.json (CAS on (_id, attempt)); parse_done webhook fires once with no result_manifest_url; timings.created_at + parse_done_at + per-phase <phase>_ms durations written (created_at ≠ lineage.processed_at)
WF-C-02 Cache-HIT copy path Same wf_content_hash present in wireframe; prior parse artifacts copied into this run's own parse.json; no LLM call and no S3 snapshot download; result-doc parse block written; parse_done webhook still fires
WF-C-03 Snapshot GET error → SQS retry wireframe_json_url GET 404/500 → exception re-raised so SQS redelivers; no wireframe write, no result-doc parse block, no webhook
WF-C-04 auto_confirm=false parks awaiting_confirm parse_done webhook is the terminal act of the invocation; assemble is NOT chained; the captured envelope has status="parse_done"
WF-C-05 Webhook send raises → SQS retry mock_ai_status(500) on parse_done → send raises → record re-raised for redelivery; DocDB parse write already durable (re-converges on redelivery)
WF-C-10 auto_confirm=true chains assemble Parse then assemble run back-to-back in one invocation; no parse_done park; only the succeeded webhook fires
WF-C-11 CAS fence rejects a superseded / zombie write A write for a stale (_id, attempt) (lower attempt than the doc) is rejected; the zombie worker's DocDB commit and its late webhook are no-ops
WF-C-01a Parse pins the latest rule revision Parse pins the project's LATEST design_rule revision into inputs (highest version, ties broken by newest lineage.processed_at); no revisions registered ⇒ pin None (assemble takes the validator no_rules path)
WF-C-03a Memo Figma-REST fallback Snapshot missing memo nodes → the worker falls back to the service-account Figma REST for memos (not the media path); extraction proceeds; a REST failure re-raises for redelivery
WF-C-04a Missing-row read is a real error A parse message whose result doc cannot be pinned because the run identity is unknown is a hard error (row-first-then-SQS), not an eventual-consistency retry-forever
WF-C-04d Fail-closed parse terminalizes A ContractFailure during parse (unresolvable pin, tenancy-scope mismatch — never resolvable on redelivery) emits the terminal …-failed.json + …-failed-manifest.json + failed webhook (status="failed" + error), exactly like assemble; a TRANSIENT parse error still re-raises with NO terminal surface (contrast WF-C-03)
WF-C-01b Memo text snapshot survives reparse After parse writes result-doc parse.memo_influences, a reparse of the same wireframe _id with changed memos (bumps the LWW wireframe doc) leaves the FIRST run's result-doc parse.memo_influences text and its parse.json UNCHANGED — the memo text is snapshotted, surviving the wireframe overwrite

Generation — Confirm / Reject

The confirm/cancel CAS is applied by the backend endpoint before enqueue; these cases assert what the worker does with the resulting message (the enum transitions themselves are backend-owned — see Out of Scope).

ID Intent Expected
WF-C-04b Confirm enqueues assemble with a fresh nonce The assemble message carries the same wf2des_id, the confirm attempt, a fresh nonce, and confirm.parse_artifact_hash; the worker rehydrates against that pinned parse.json
WF-C-05a Reject is NOT a webhook A parse rejection is written by the plugin to wf2des-api (parse.rejected = {reason_code, note}) — the worker emits no reject webhook; assemble is never enqueued for a rejected run
WF-C-11b Completed generation is never reprocessed A redelivered message for a run with a terminal manifest applies that manifest rather than re-running; the worker never overwrites a completed generation

Generation — Assemble

ID Intent Expected
WF-C-06 Assemble happy path Rehydrates pins; applies the deterministic role/kind gate; consults the decision ledger; runs section selection as a K-sample self-consistency vote with deterministic guards over the vote, deterministic stitch/validator, and computed confidence; writes back the ledger (score-ratchet fence); writes spec, screen-plan/layout evidence, hashes, assembly artifact, result and manifest under CAS, plus design_resolution upserts; sends one succeeded webhook
WF-C-07 Rehydrate fails closed on hash mismatch A pin in inputs whose stored hash no longer matches → fail closed; no fallback to older rules/components; run ends failed, …-failed.json + failed webhook
WF-C-08 Generation fails after assemble An assemble-stage error → …-failed.json + …-failed-manifest.json; failed webhook carries error + the failure manifest URL; no succeeded webhook
WF-C-09 Validator no-rules path (assemble) No design_rule registered → validator_report.status = no_rules, every node flagged rules_unvalidated; run still succeeds; flag_count reflects the flags
WF-C-06a Redelivery reuses pins, re-selects only if absent On redelivery the existing inputs pins are reused; selection re-runs only if the block is absent — a crash-retry can never silently select different components or rule revisions
WF-C-06b Large spec spills to S3 spec > ~1MB → the spec block spills to …-spec.json and artifact_urls.spec is set; spec_nodes_flat + summaries stay inline; the result doc still validates
WF-C-06c Candidate truncation is flagged When a section's candidate set is truncated, selection.truncated records it and the affected nodes surface as flagged, not silently dropped
WF-C-09a flag_count mirrors to the row via manifest confidence.flag_count is copied into the manifest row_effects.wf2des.flag_count and the succeeded envelope's manifest → the backend row's flag_count
WF-C-06d placement.design_area resolved from config.json Seed config.json (via project_figma_file.config_url) with a known DESIGN-area rect; the written placement.design_area {x,y,w,h} equals that rect — proving the worker RESOLVES it, never fabricates it
WF-C-06e Rule-revision boards attach to selection The guideline boards ride the pinned design_rule revision: parse pins the LATEST revision, assemble rehydrates BOTH the RuleSet and the revision's N board images from that ONE pin and attaches them as N vision BinaryContent items to EVERY per-section selection call (advisory reference; the deterministic stitch/validator/confidence are unchanged); a missing pinned doc or a missing board image ⇒ fail closed
WF-C-06f Decision-ledger consult LOCKs / FLOORs a section Seed design_resolution for a section's section_signature: a mined entry (or one with score ≥ the trust threshold) LOCKS the section — mock_llm.selection is not called for it, the ledger decision is taken verbatim; a weaker voted entry is a FLOOR — the section re-votes and the strictly-higher score wins (a lower re-vote keeps the floored decision); a stale/mis-scoped default-variant entry falls through to the vote. Consult filters organization_id + project_id
WF-C-06g Ledger write-back + ratchet reconverges After a successful assemble, each non-unmatched section upserts design_resolution with a deterministic score; a re-run over a lower-scoring voted entry REPLACES it (monotonic improvement), a re-run whose score is equal-or-lower is REJECTED and the held entry is KEPT (a DuplicateKeyError on the fenced upsert returns success, not an error); a held mined entry is never overwritten by a voted write; identical evidence reruns leave updated_at unchanged (no churn)
WF-C-06h Terminal provider configuration/account error Exhausted credits, invalid credentials/access, or unavailable model writes failed artifact + manifest, sends failed webhook, exposes a safe error, and consumes the record
WF-C-06i Transient provider failure Ordinary 429 rate limit or 5xx emits no terminal surface and returns the item for SQS redelivery

wf_parse

ID Intent Expected
WF-C-12 Writes the wireframe doc only Frame-registration event → deterministic extraction + memo status + screen_id grammar + variation_label + mock_llm.role_intent; writes the wireframe doc only; no result doc, no parse.json, no preview, no webhook, no PG effect
WF-C-13 Snapshot GET error → SQS retry The registration snapshot URL GET fails → exception re-raised for wf2des-events redelivery; no wireframe write
WF-C-14 LWW fence on (source_hash, processed_at) A redelivery with an older/equal processed_at and unchanged source_hash does not overwrite; a new source_hash or newer processed_at does; generation-parse + wf_parse on the same _id is intended non-collision

rule_process

ID Intent Expected
WF-C-15 Immutable revision per (design_rule_id, content_hash) Resolves the selected guideline boards via get_file_nodes (titles); renders each board via the Figma images API (service-account PAT) and stores the PNG to S3 (wf2design/rule-boards/{file_key}/{node_id}.{sha256}.png); one mock_llm.rule_extraction vision call per board (slot_constraints always empty); deterministic merge of the fragments; canonical content_hash over the MERGED RuleSet; assigns version = max+1; writes ONE immutable design_rule revision doc carrying the merged rules AND boards[] {node_id, title, image_url, content_hash, mime} + extraction provenance {extraction_model_id, extraction_prompt_version, extracted_at} + draft_source="llm_extracted"; no webhook, no PG effect
WF-C-15a Container selection expands to board frames A selected CANVAS (page) or SECTION expands to its child FRAMEs (nested sections recurse; non-frame children are not boards); the revision's boards[] carries the frames, never the container; an EMPTY container is a hard error — resolution of the designer's selection, never curation
WF-C-16 Missing / unrenderable board fails the record A selected board node absent from the file (get_file_nodes miss) or a null render from the images API is a hard error → the record fails (re-raised for redelivery); no partial design_rule revision is written
WF-C-17 Replay of identical merged rules = same revision Re-registering boards whose MERGED rules are identical maps to the existing (design_rule_id, content_hash) revision — no new ordinal minted; changed rules mint the next version (accepted version churn); concurrent identical registrations serialized by the single event

component_sweep

ID Intent Expected
WF-C-18 Full sweep happy path Schedule/resync → acquires the sweep_marker; walks Figma REST (/components published + local via full-file fallback); no LLM; per COMPONENT/COMPONENT_SET writes a content-addressed S3 snapshot + upserts a design_component context doc (_id = platform design id) with deterministic variant_properties/slots/default_size; captures project_figma_file.style_captures; emits the discovered-component sweep manifest; stamps components_synced_at LAST
WF-C-19 Single-flight serializes concurrent sweeps A second sweep while a live sweep_marker holds is blocked until release/expiry; a marker past expires_at is reclaimable (crash-safe); design_component upsert is on _id
WF-C-20 Sweep failure stamps the file doc A Figma REST error mid-walk → project_figma_file.sweep_error set field-level, marker released/expired; a missed manifest apply re-converges on the next sweep; NOT the ai-status webhook
WF-C-20a Emits a registry manifest, not ai-status component_sweep is not webhook-free: it emits the discovered-component manifest (id · name · type=component · image_url · json_schema_url · status); it never posts /v1/webhooks/ai-status; the worker does not touch PG (backend applies the manifest)

Figma OAuth credential (the scheduler's authority)

The scheduler is a bodyless EventBridge invoke — no user, no request. Figma supports only the authorization-code flow, so the credential is a grant an admin authorizes ONCE. The Admin API stores and refreshes it; the worker requests only a valid access token. Most failures at this boundary are otherwise easy to misdiagnose, which is what these pin.

ID Intent Expected
WF-C-28 Healthy token costs nothing Admin finds an access token more than REFRESH_SKEW (1 day) from expiry → returns it as-is; no Figma round trip and no write to figma_oauth_grant
WF-C-29 Near-expiry refresh preserves the refresh token Inside the skew → Admin holds the organization advisory lock, calls /v1/oauth/token with grant_type=refresh_token, and stores a NEW encrypted access_token + expires_at; the stored refresh_token remains encrypted and unchanged when the response carries none
WF-C-30 The refresh race has exactly one winner Two provider calls inside the skew serialize on the PostgreSQL transaction advisory lock. Exactly one calls Figma; the second reads the refreshed row and returns it without refreshing again
WF-C-31 A failed refresher releases the lock The refresh call fails or the Admin request exits → the transaction ends and PostgreSQL releases the advisory lock; no persistent marker can strand the grant
WF-C-32 A missing grant is an operator error, not a retry No PostgreSQL figma_oauth_grant row for the org → the Admin provider returns an actionable error naming the connect flow; the worker does not silently fall back to PAT when OAuth is configured
WF-C-33 A rotated encryption key fails loudly Admin cannot decrypt a grant written with another key → returns a specific provider failure rather than garbage that would surface as a confusing Figma 403
WF-C-34 Credential selection and fallback OAuth configured → Authorization: Bearer, no X-Figma-Token; OAuth unconfigured but figma_pat set → X-Figma-Token (an un-migrated deploy keeps working); neither → a configuration error naming both options
WF-C-35 Tokens stay inside the required boundary No log line, webhook, manifest, public API response or DocumentDB document contains a token; PostgreSQL stores ciphertext; the internal provider returns only the access token + expiry and never the refresh token, client secret or encryption key

Figma OAuth connect flow (Admin API)

The callback is unauthenticated by necessity — a browser redirect from Figma carries no Cognito token — so state is the entire security boundary.

ID Intent Expected
WF-C-36 Only an organization admin may connect Admin Cognito session with organization read/write → accepted; a role without organization read/write → 403; the org-wide credential cannot be replaced by an application user
WF-C-37 state cannot be forged or tampered A state signed with another secret, a state whose payload was swapped to a different organization_id keeping the signature, a malformed state, and an expired state all → 401; the organization is taken from the signed payload, never from a query parameter
WF-C-38 The callback is not a forging oracle Signature-failure and malformed-input return the same message, so an attacker cannot distinguish which check failed
WF-C-39 Declines and expired codes are reported honestly Figma returns error=access_denied → 400 naming the decline; an authorization code older than its 30-second life → 400 telling the admin to start again — neither is reported as a server fault
WF-C-40 Admin encrypts before the PostgreSQL write The callback encrypts both tokens with the platform ENCRYPTION_KEY and upserts one row per organization. With ENCRYPTION_KEY unset, storage is refused rather than writing plaintext; AI receives neither the key nor the refresh token

Snapshot scoping, guard tiers, and the bookkeeping doc

ID Intent Expected
WF-C-41 A supplied sectionNodeId scopes the snapshot fetch With the field set and the frame inside that subtree, only GET /files/{key}/nodes?ids= is called and GET /files/{key} is never called. Measured motivation: the whole-file response was 255 MB / 67.5s on a real 31-page file, and a trigger died at 140s with a closed socket
WF-C-42 A wrong or stale scope degrades, never corrupts A sectionNodeId that does not contain the frame, or names a deleted node, logs wf2desSnapshot.scopeRejected / fullDocumentFetch and falls back to the whole-file walk. The resulting snapshot_scope and wf_content_hash are the same as if the field had been omitted — the hash keys the parse cache, so a wrong scope must never serve another screen's parse
WF-C-43 The KIND guard judges variant-named picks An instance pick that echoes variant_props and whose component kind the region's role does not admit is downgraded to compose and logged selection_kind_mismatched. Regression guard: the check previously sat behind the size tier's not variant_props condition, leaving it inert for 57% of real picks (33 of 57 measured)
WF-C-44 The SIZE guards stay gated to no-variant picks An overscoped component whose kind the role admits passes when a variant is named, and is caught (unmatched) when none is. default_size describes only the default variant, so widening this tier would drop correct slim-variant picks
WF-C-45 kind = other is never a category error A component filed other is admitted at every role. ROLE_CANDIDATE_KINDS lists other for no role, so treating it as a mismatch rejects all 19 such registry components everywhere — measured live as logo rejected at a header region
WF-C-46 The ledger revalidates kind, not just scope A stored entry whose component kind the role cannot admit is rejected at consult time with ledger_entry_misselected reason=kind_mismatched and the section re-votes — even when the entry names a variant and its score clears the trust threshold. Without this a 0.88-scored bad pick locked three sections and two consecutive runs produced byte-identical output
WF-C-47 The bookkeeping doc is created, not assumed With no project_figma_file doc, the registry-enrichment pass creates one ($setOnInsert, role = 1) and its slot_roles write LANDS. Re-running is idempotent and never overwrites a role / config_url set by registration. Regression guard: absent doc made all five writers silent no-ops, costing 104 discarded LLM calls per sync
WF-C-48 A leftover text is preserved, not dropped A section whose text count exceeds the chosen component's slot capacity emits a wrapper carrying [preserved_text…, instance] (heading above when it precedes every slot-consumed text). The log reports text_leftover — at INFO for one, WARNING at the flag threshold — and text_preserved names the outcome. The event must not claim a drop for text the caller then preserves

Cross-family component checks

ID Intent Expected
WF-C-24 Every read/write filters tenancy For each family, the DocDB query + write filter both include organization_id + project_id; a doc for another project is never read or written
WF-C-25 Workers never open a PG connection Run every family with the PG driver unreachable / a guard env var that fails loudly at import; no path opens a PostgreSQL connection or holds PG credentials; all effects land in DocDB / S3 / the webhook only
WF-C-26 Lineage envelope populated Every written DocDB doc (wireframe / design_rule / design_component / result) carries a populated lineage{source_url, source_hash, processor_version, index_schema_version, processed_at, job_id}; job_id = the producing run id
WF-C-27 Intake event_type dispatch Each wf2des-events event_type routes to exactly the right run family (rule→rule_process, frame-reg→wf_parse, resync→component_sweep) and an EventBridge invoke → component_sweep; a rule event never triggers component_sweep

End-to-end (E) Cases

Smoke set only, against real dev. Cost is non-zero — keep it small.

ID Intent Steps Expected
WF-E-01 Interactive generation round-trip Backend creates the wf2des row + enqueues parse → confirm endpoint → assemble Within budget: parse_done then succeeded webhooks received; result artifact + manifest in S3; wf2des row → status '1' completed + result refs + flag_count, phase cleared
WF-E-02 auto_confirm one-shot Enqueue a parse message with auto_confirm=true Parse chains straight into assemble in one invocation; only the succeeded webhook is observed; row reaches status '1' with no awaiting_confirm park
WF-E-03 Real component_sweep Trigger a resync (or wait for the EventBridge schedule) on a project with registered files design_component docs + content-addressed snapshots written; the discovered-component manifest is applied to platform design (type=component) rows; components_synced_at advances
WF-E-04 DLQ when the webhook keeps failing Force the backend ai-status webhook to 5xx across all retries The generation message DLQs after retries and the terminal S3 manifest remains the durable record. There is no second net: the stuck-generation sweep that would replay the manifest is not implemented, so the row stays non-terminal until cleared by hand
WF-E-05 Awaiting-confirm safety nets Park a run at awaiting_confirm; let it exceed the confirm timeout The backend awaiting-confirm sweep (e.g. 72h) fails the run closed; no assemble is enqueued

Cross-cutting Checks

Invariants that must hold across all applicable cases. These are not standalone rows — verify them while running the cases above.

  • Workers NEVER write PostgreSQL (WF-C-25): every family runs with the PG driver unreachable / a guard env var; no path opens a PostgreSQL connection or holds PG credentials. Every PG effect in the contract is backend-side.
  • ai-status webhook is generation-only and exactly once per terminal transition: it fires for generation only; wf_parse / rule_process never emit it, and component_sweep emits the discovered-component manifest instead. Within one process_record, exactly one of parse_done (park) or succeeded/failed (terminal) is sent — never a reject webhook.
  • The LLM fence: role/intent (parse + wf_parse), per-section selection as a K-sample self-consistency vote (assemble), and per-board rule extraction (rule_process, ingestion-time). Extraction, memo status, screen_id, variation_label, the role/kind candidate gate, the ledger consult, the deterministic guards over the vote, the ledger score + write-back, stitching, validator, the rule-fragment merge, content_hash, and the sweep are deterministic; component_sweep calls no LLM. Confidence is COMPUTED (formula_version), never a model self-report; the section_signature is LLM-free.
  • Compose ALWAYS flagged, unmatched NEVER dropped (WF-U-23, WF-U-25, WF-C-06): every compose node carries flagged=true; every unmatched node becomes a visible placeholder.
  • Tenancy (WF-C-24): every read and every write filters organization_id + project_id; cross-project access never happens by default.
  • Fencing: result-doc CAS on (_id, attempt) (generation, WF-C-11); wireframe LWW on (source_hash, processed_at) (WF-C-14); single-flight sweep_marker (WF-C-19); rule content-addressed uniqueness (WF-C-17); design_resolution score-ratchet upsert — a voted write lands only over a non-mined, strictly-lower-scoring entry, a DuplicateKeyError on the fenced upsert means KEPT (WF-U-38, WF-C-06g). A superseded attempt's DocDB write and late webhook are no-ops.
  • Idempotent content-addressed writes reconverge on redelivery (WF-C-05, WF-C-14, WF-C-17): a redelivered internal-run event or a retried generation write converges on the same doc; completed generations are never reprocessed — a terminal row is final and re-running a wireframe creates a NEW row.

Intentionally Out of Scope

Listed so future tests land in the right bucket. These are owned elsewhere.

  • The backend wf2des PostgreSQL row state machine + all endpoint auth: status/phase/attempt transitions, the confirm/cancel/placement/feedback endpoints, and their authorization live in backend tests. The worker only reports via the webhook.
  • The ai-status webhook handler's manifest↔payload validation + 5xx-on-manifest-read: the handler validating the manifest's job_id / job_type / S3-prefix against the payload (and returning 5xx on a manifest-read failure) is backend logic — the worker only builds the manifest + envelope (WF-U-27, WF-U-28).
  • The plugin's Figma materialization + placement writes: reading the spec, building native Figma, intrinsic-height footprint preservation, bounded correction/promotion, and the seven materializer events live in plugin tests.
  • The platform design (type=component) registry upsert from the sweep manifest: applying the discovered-component manifest as design row upserts is a backend concern — the worker only emits the manifest.
  • Figma service-account auth / secret management: the PAT in wf2design's own secret store and its rotation are infrastructure, not worker logic.
  • DLQ / SQS / EventBridge infrastructure: queue wiring, redrive policies, and the schedule cadence are infrastructure tests; the worker only verifies re-raise / redelivery behavior.
  • Cross-tenant RBAC: purely upstream authorization; worker tests assume the message is already authorized for its organization_id + project_id.
  • The LLM's actual selection QUALITY: mocked here with structured outputs. Whether the chosen component/variant/label is good is evaluated by a separate human-review protocol, not these tests.

  • I/O Definition — the contract these tests enforce.
  • Overview — the direct-assembly pipeline and internal event paths.