Skip to content

AI WF2Des — Assembly Architecture

Bounded planning and artwork fidelity (September 23, 2026)

Screen planning retains a hard total of 20 provider requests and 48 inspection calls. Exploration reserves three requests for tool-free finalization using the same verified history and usage counter. Repeated identical inspections return a compact reference; unexecuted over-budget calls supply no evidence. Failed calls retain their usage records. If finalization still fails, the baseline survives with cannot_verify findings and a screen_planning validator warning, not an implicit visual pass.

Hash-verified snapshots identify explicitly placeholder-labelled icon slots containing only empty/container/rectangle artwork, separately for each captured variant. Real vectors, images and hidden slots do not count as empty placeholders. Such slots cannot replace source artwork; uncertain variants preserve the source rather than inventing a filled icon. Binary selection-state guards inspect the actual captured axes even when a toggle was tagged as button or other, preventing generic arrow toggles from replacing opaque check artwork.

Adjacent one-action navigation wrappers can be covered by a footer only when captured native label-plus-glyph evidence proves the action remains visible after overrides. A parsed form label alone neither authorizes nor prohibits coverage; real fields and unrelated artwork stay. The renderer counts padding once when a sole instance and wrapper share the same exact source owner and bounds. Remaining footprint is measured after width-dependent label wrapping. Independent source slots, compound siblings, explicit minimum heights and table tracks remain unchanged. These fixes require no project-specific IDs or registry migration; new actual renders are still required to verify visual quality.

Asset identity and optional-content proof

Registry hygiene no longer combines assets by display name. Separate captured sources retain their own axes and defaults, avoiding impossible variant combinations. Snapshot preflight also derives read-only BOOLEAN/visibility evidence for proxy instances whose master definitions were not captured. The planner receives these controls, and grounding checks the resulting visible copy before allowing a compound replacement or duplicate-navigation coverage. Observed proxy values are explicitly applied rather than mistaken for master defaults. A textless atomic mark can be media even when its primary classification is button. These checks are independent of screen IDs, project IDs, component names and provider labels. Fresh plugin renders remain the acceptance gate; successful tests/replay alone do not establish visual parity.

Generation prefers a unique complete compound when its proven visible-text multiset covers all source copy (at least three labels), with only captured optional parts disabled and any extra navigation accounted for by a nearby matching native action. Ambiguous candidates, unresolved variants, independent source images and partial copy matches abstain. Actual-render corrective review is not forced back to this choice. The chosen identity/properties/coverage are recorded in the existing screen plan with a deterministic rationale.

Semantic section layout and authority (section-layout@1)

Component selection answers which design-system units represent the source. The additional section-layout layer answers which contiguous source siblings form a meaningful vertical section and who owns spacing at its boundaries. The model may request bounded sections via ScreenPlan.semantic_sections; a deterministic compiler validates source membership/order, current emitted coverage and layout eligibility before applying them. It cannot rewrite copy, cross grids/absolute regions/instances or silently discard content. Unsupported or ambiguous requests retain measured source geometry and record findings.

The authority order is explicit: a selected component owns its internal geometry; an unambiguously applicable rule owns stated section joints and an explicitly stated viewport_width_px; source measurements supply the fallback. base_viewport_px is advisory, and content width/padding cannot be converted into an invented screen width. Effective joints include existing boundary padding. Painted containers, chrome edges and intentional spacing are protected from blanket normalization. Only eligible flow instances receive height_authority="component" to avoid obsolete wireframe-footprint spacers.

The immutable assembly state records the validated layout_plan (version, requests, sections, joints, findings and viewport). Generation and actual-render correction share its deterministic application, with current selection coverage revalidated before correction. No-op review reproduces the same spec; old states without a plan retain old replay behavior. The existing five-node spec union is unchanged, and the plan is not a claim that the final render passed visual review. See the additive I/O contract and regression matrix.

Screen planning and materialization fidelity (September 2026)

Generation now supports a bounded whole-screen review after component matching, using the source tree, captured screenshot, registry and rules together. It returns sparse corrections to existing source-node/component references, never arbitrary Figma code or rewritten copy. Unchanged matcher decisions survive. A blank-slate planner was tested and produced worse component choices despite improved wrapping; it is an experiment, not the production default. screen_planning_enabled=false disables the review for comparison. Unknown IDs, variant values and style references are rejected; invalid review output leaves the matcher intact and records an explicit unavailable status. Registry choices still pass deterministic scope guards.

The plan also assigns existing rule colors to authored text and layout surfaces. It cannot invent colors, font sizes, spacing or content. The deterministic stitch and validator remain responsible for text-slot completion, content preservation, geometry and confidence. Plan decisions and rejected proposals are recorded on the result as screen_plan for replay and diagnosis. Review and matcher model IDs are separately pinned at parse time. New runs copy screenshot bytes to a content-addressed run artifact and verify the hash on read. Older runs without a screenshot pin remain tree-only unless a replay explicitly supplies an advisory image and records its provenance.

The additive shared TextNode fields font_family, line_height, line_height_pct and max_lines carry explicit typography into the plugin. Null/absent line limits mean no truncation. WF2Des always supplies a positive wrapping width when the source has a measured width; explicit wrap_width=null is reserved for the Des2WF single-line contract. Existing saved specs still load. No destructive database migration is required: older inferred empty-component semantics are quarantined on read and corrected by subsequent capture/enrichment.

InstanceNode's additive component_properties carries optional-part booleans (for example footer navigation). Planning may set only names captured in the registry; the plugin resolves the live hashed property name and verifies the property's actual type before applying it. Text overrides continue through source-grounded text_slots, not arbitrary property values.

The comparison harness replays stored source/rule/component pins without updating existing jobs or the decision ledger. Structural metrics are diagnostics, not visual-accuracy percentages; rendered output must additionally be checked for overflow, missing assets, placeholders and duplicated chrome before claiming parity with a designer/MCP result.

The reviewer now has a compact component index plus bounded, on-demand inspection tools. Each detail response verifies the pinned snapshot hash and names the exact existing variant, ordered children, geometry, typed property defaults, hidden content, and evidence hashes. Missing Cartesian variant combinations are rejected and actual combinations are offered; the tool never guesses the first child as a default. The default reviewer is GPT-5.4 with high reasoning over Responses; the section matcher remains independently configurable. Neither a larger model nor a successful tool call is a visual-quality guarantee.

After the plugin materializes the spec, it sends the actual PNG, node/path/text/bounds manifest and diagnostics through the scoped render-review endpoint. The worker proposes at most one deterministic reassembly, and a second actual render is verification-only. The original is retained; failures, stale attempts and unresolved defects remain unverified. An unresolved warning, error or failed text-slot write blocks a visual pass; informational notes about successfully preserved literal styles do not. No diagnostic count is an accuracy score. The plugin rechecks current PG job and result metadata before accepting terminal review results. Frozen assembly_state.section_ids keep validation at the same scope as generation: using every source node would silently rewrite nested spacing even in a no-change review.

A source-proven empty structural leaf stays a transparent, fixed-size layout slot. This does not hide unresolved controls, text, artwork, component identities or legacy unknown ink evidence. A separate narrow wordmark binding can represent a sole short brand label with inspected graphical artwork during actual-render review. It requires exact source text/identity, positive painted wordmark evidence and successful current-session inspection; the instance retains exact copy, source/component/evidence hashes as metadata. Only its duplicate leftover sibling is removed, never source text globally or arbitrary headings, paragraphs, navigation or controls. The new render still must be checked for equivalence.

Held-out evaluation inputs include payment-method selection and login-complete screens, not only the original three examples. Their runs use the same library/rule pins without reference-output recipes, job updates or decision-ledger writes. The first held-out renders preserved payment/identity artwork but still exposed duplicate branding and footer defects; they are regression evidence, not an assertion that visual parity has been reached.

The actual-render correction removed duplicate branding and the empty footer placeholder. Its verification then exposed an explicit default input selection being changed to active by geometry-based variant inference. Geometry may now infer only compatible unselected axes; it cannot overwrite an explicit selected variant. This regression applies to arbitrary component states and axes, not a named screen or input component.

Component and source-node identities also include their Figma file scope. Local node IDs may be used only when the originating file is proven to match; cross-file components use their published key. A coincidentally equal node ID in another file is not a match. Legacy references without sufficient identity fail explicitly instead of cloning unrelated content.

apps/wf2des/scripts/replay_generation.py supports stored-spec export, matcher-only replay, production review, and an explicitly experimental standalone planner. Matcher state can be saved and reused with identical input pins, avoiding a second round of section voting when iterating on review. guinness-figma-plugin/scripts/build-mcp-materializer.mjs renders the unchanged app spec through the real materializer in MCP for visual QA; its shims adapt unsupported host APIs only.

This page is the implementation-facing deep dive into the direct-assembly engine that powers WF2Des generation. Where the Overview sketches the pipeline in eight steps, this page defines where the boundaries are: which decisions are deterministic transforms, which are the LLM's, how the two phases hand off, how each spec node is produced, and why the whole thing is reproducible.

The defining property is stated once here and holds throughout: there is no retrieval, no embeddings, no vector search, and no learning corpus. The engine reads the project's own component registry and rule set directly and assembles from them. This is the structural contrast with des2code, which retrieves a broad candidate pool from indexed vectors and reranks it. WF2Des never indexes, never scores by similarity, and never ranks against a corpus — its candidate step is a straight, tenant-scoped filter.


The LLM Fence

The engine is deterministic-first. The LLM makes only bounded, evidence-pinned decisions inside a fence; everything on the other side of the fence is a reproducible transform. The fence is the single most important architectural line in this worker, so it is drawn explicitly here.

Inside the fence — bounded model stages:

Phase LLM decides LLM never decides
Parse (fast tier) Each WFNode's role (from the shared element-type vocabulary — the closed set of WFNode role values, pinned in prompts.py as ROLE_VOCABULARY) and intent (from open memos + variant labels) The tree shape — WFNode extraction is deterministic
Assemble (strong tier) Per section: the case (instance / compose / unmatched), the component key, variant props, and text-slot fills Candidate admissibility, the deterministic guards over the vote, the stitch, the validation, or the confidence
Source-screen review (separately pinned reviewer) Sparse corrections after matching, grounded in source, rules and inspected component evidence Arbitrary code, invented copy, components or variant combinations
Actual-render review (same pinned reviewer) Findings and at most one correction, followed by verification of its actual render A visual pass based only on an unrendered spec or an unbounded correction loop

These are stages, not a fixed count of provider requests: selection voting and bounded inspection tools can require multiple requests. Rule ingestion separately makes a vision extraction call per guideline board (rule_process, vision tier) — see Guideline boards on the rule revision.

Outside the fence — everything else is deterministic:

  • WFNode extraction — the wireframe snapshot is walked into a WFNode tree by a fixed rule (visible FRAME / GROUP / INSTANCE / TEXT descendants plus image-fill rectangles; vector/shape leaves collapse into their parent; hidden layers skipped). The tree shape is a function of the snapshot, not of the model.
  • Memo status — a memo inside a 取込済 ("imported") status frame is resolved by containment; resolved memos are excluded from live intent. Pure geometry/containment, no model.
  • screen_id grammar — deterministic name grammar ([A-Z]{2,3}_[A-Z0-9]+, unanchored) in wf_parse; at generation screen_id comes from the request. variation_label is a deterministic frame-name split.
  • Candidate set — the project registry is deduped and hygiene-collapsed, then a deterministic role→admissible-kind gate narrows each section. Size is evidence for the model, not a gate. No scoring or similarity is involved (see Candidate Set).
  • Deterministic guards over the vote — the LLM's selection echo is clamped and scope-checked by a fixed program before it is trusted (see Deterministic Guards).
  • Decision ledger — design_resolution is consulted before selection and written back after the guards by deterministic rules (see The design_resolution Decision Ledger).
  • Spec stitching — the per-section decisions are assembled into the DesignSpec by a fixed transform (see Deterministic Stitching).
  • Rules validator — the compiled rule set runs as a program over the stitched spec (see The Rules Validator).
  • Computed confidence — a formula over structural signals; never the model's self-report (see Computed Confidence).

The fence has a hard consequence: everything the LLM decides is stored with its evidence (source, lineage_wf_node_ids, and the pinned inputs), and everything the LLM doesn't decide can be recomputed. There is no place in the pipeline where a similarity score, an embedding, or a corpus lookup enters — those simply do not exist in this engine.


Two-Phase Pipeline

Generation runs in two phases separated by a confirm checkpoint. Parse produces the parsed wireframe and parks (interactive) or chains straight through (auto-confirm); assemble rehydrates the pinned inputs and produces the terminal DesignSpec.

flowchart TD
  subgraph PARSE["Phase 1 — Parse"]
    P0["PIN inputs into the result doc<br/>(first act — the pinning moment)"]
    P1{"parse cache hit?<br/>composite _id + same wf_content_hash"}
    P2["DETERMINISTIC WFNode extraction<br/>visible FRAME / GROUP / INSTANCE / TEXT<br/>+ image-fill rects; vectors collapse; hidden skipped"]
    P3["DETERMINISTIC memo status<br/>取込済 containment ⇒ resolved (excluded from intent)"]
    P4[["LLM · fast tier · roles + intent ONLY<br/>never the tree shape"]]
    P5["copy prior parse artifacts<br/>(no LLM)"]
    P6["write wireframe cache doc (LWW)<br/>+ parse block + immutable parse.json (CAS)"]
    P0 --> P1
    P1 -->|miss| P2 --> P3 --> P4 --> P6
    P1 -->|hit| P5 --> P6
  end
  P6 --> CONF{"confirm checkpoint<br/>auto_confirm?"}
  CONF -->|"no · interactive"| PARK["park at awaiting_confirm<br/>plugin previews → designer confirms"]
  CONF -->|"yes · auto"| CHAIN["assemble in the same invocation"]
  PARK --> A0
  CHAIN --> A0
  subgraph ASM["Phase 2 — Assemble"]
    A0["REHYDRATE pins from result doc inputs<br/>FAIL CLOSED on any hash mismatch"]
    A1["DETERMINISTIC candidate set<br/>registry hygiene + per-section role/kind gate<br/>(no scoring / no similarity; size is a signal)"]
    AL0["DECISION-LEDGER consult (design_resolution)<br/>mined or score ≥ trust → LOCK (no LLM)<br/>weaker entry → FLOOR (re-vote)"]
    A2[["LLM · strong tier · SECTION-PARALLEL selection<br/>K-SAMPLE self-consistency vote per section<br/>instance / compose / unmatched · component key ·<br/>variant props · text-slot fills"]]
    AG["DETERMINISTIC guards over the vote<br/>variant clamp · symmetric scope guards ·<br/>dup-claim dedupe · kit-name override ·<br/>variant + screen-facts resolution"]
    AR[["LLM · pinned whole-screen reviewer<br/>source image + rules + inspected components<br/>sparse corrections, deterministic grounding"]]
    A3["DETERMINISTIC stitch → DesignSpec<br/>spec tree · spec_nodes_flat · style_bindings"]
    A4["DETERMINISTIC rules validator<br/>auto-fix where safe, else flag"]
    A5["COMPUTED confidence<br/>never model self-report"]
    AL1["DECISION-LEDGER policy<br/>reviewed output is not automatically learned"]
    A6["result doc spec/selection/validator/confidence (CAS)<br/>+ S3 result artifact + manifest"]
    A0 --> A1 --> AL0 --> A2 --> AG --> AR --> A3 --> A4 --> A5 --> AL1 --> A6
  end

Double-bracket nodes ([[…]]) mark model stages, not provider-request counts. After generation, the plugin captures its actual materialized output for the separately fenced render-review cycle described above. A successful assembly is not a visual pass.

Phase 1 — Parse

The first act of the invocation is to pin the inputs into the result doc's inputs block (wireframe identity + source_hash, plus the rule/component/LLM pins that assemble will rehydrate). Pinning first is what makes the run reproducible: everything downstream references the pins, not live state.

Parse then checks the parse cache — a primary-key lookup in the wireframe collection by composite _id = {project_id}_{figma_file_key}_{node_id}, matched against the message's wf_content_hash. A hit copies the prior run's parse artifacts into this run's own parse.json with no LLM call; a miss downloads the snapshot and runs the deterministic extraction, then the single fast-tier LLM step for roles + intent. The cache is deliberately skipped on retry-after-reject — a fresh parse is the point in that case.

Parse writes the wireframe cache doc (LWW-fenced) and the result doc's parse block plus the immutable parse.json. The parse.json artifact is authoritative: the wireframe collection doc may be overwritten by a later reparse, so the preview and memo-review surfaces read the result doc + parse.json, never the overwritable cache doc.

Confirm Checkpoint

Between phases sits the confirm checkpoint. Interactively, parse posts parse_done and the row parks at phase = awaiting_confirm; the plugin previews the parse and the designer confirms (or rejects). Confirm is a backend endpoint that CAS-checks awaiting_confirm + attempt and enqueues assemble. On the auto_confirm path there is no park — assemble runs back-to-back inside the same invocation and only the terminal webhook fires. The checkpoint exists because parse is where memo intent is interpreted; letting the designer confirm the interpretation before committing to component selection keeps the expensive strong-tier step from acting on a misread wireframe.

Phase 2 — Assemble

Assemble's first act is to rehydrate the pins from the result doc's inputs block and fail closed on any hash mismatch — a pinned rule revision or component snapshot that no longer matches its hash aborts the run rather than silently falling back to a newer version. From the rehydrated pins it builds deterministic per-section admissible candidate sets, consults the design_resolution decision ledger, runs the single strong-tier section-parallel selection step (a K-sample self-consistency vote), applies the deterministic guards over the vote and the source-screen review, then runs deterministic stitch, validation, confidence and the ledger policy, then writes the terminal result doc + S3 artifact + manifest.

On redelivery, assemble reuses existing pins and re-runs selection only if the inputs block is absent — a crash-retry can never silently select different components or rule revisions than the attempt it is retrying.


The Five Spec Node Cases

The stitched DesignSpec is a tree of exactly five node cases. The section-selection LLM chooses among instance / compose / unmatched for content nodes; layout_frame and text fall out of the deterministic structure. Each case is produced as follows.

Case Produced by Flagged? Contents
layout_frame Deterministic — an auto-layout or coordinate-positioned container derived from the WFNode structure No auto_layout {direction, gap, padding, sizing} (direction=none for overlapping free-form layers), fill_opacity, clips_content, children, lineage_wf_node_ids
instance LLM selects a registered component from the section's role/kind-admissible set; stitch resolves slots Only on low confidence component_key, platform_design_id (the stable validator key), variant_props, text_slots, bbox, source {kind: registry \| wireframe}
compose LLM decides no single component fits and composes a subtree from children Always flagged inlined children, source {kind: composed}
text Deterministic — a text WFNode with content bound to a style token Only on low confidence content, style_token, source
unmatched the LLM section-selection finds no confident match → a visible placeholder Always flagged placeholder {role, text, bbox}, source {kind: none}

Two of these are bright lines the engine will not cross:

  • compose is always flagged. A composed subtree is the least-grounded outcome — no registered component matched, so the engine built one from parts. It is never presented as a confident result; it is always surfaced for review.
  • unmatched is never silently dropped. A node with no confident match becomes a visible placeholder carrying its role, text, and bbox, flagged for the designer. A gap in the design is made visible, not hidden.

Every node — all five cases — carries lineage_wf_node_ids back to the WFNodes it came from, and each content node also appears once in spec_nodes_flat ({layer_path, case, component_key, wf_role, confidence, flagged, lineage_wf_node_ids}) for flat review and feedback attribution.


Candidate Set

Before the LLM selects, the engine builds the candidate set — this is the step that in des2code would be a vector retrieval, and here is nothing of the sort.

  • Source — the project's component registry, which is the shared platform design table (type=component) pinned in inputs.components. The assembly/instancing context for each candidate is read from the design_component collection (keyed by _id = {project_id}_{figma_file_key}_{node_id}): variant_properties, text_slots, image_slots, default_size, component_key.
  • Role-aware admissibility. After registry-wide dedupe and hygiene collapse, a deterministic role→kind map removes categories that cannot structurally satisfy the section. The production default is CANDIDATE_KIND_GATE=true; it may be disabled only for controlled evaluation. The selection LLM sifts the admissible candidates by each candidate's full structure. Size remains a signal, not a gate, so a valid component is not rejected solely because its default variant has different dimensions.
  • No scoring, no similarity, no ranking. There is no embedding of the WF section, no cosine distance, no learned reranker. The set is a straight, tenant-scoped read over registry metadata, scoped by organization_id + project_id.
  • Truncation is recorded, not silent. Per-section candidate counts land in the selection block's candidates_considered map, and any truncation of an oversized candidate set is flagged there.

Because both hygiene and role→kind admissibility are deterministic, each per-section candidate set is reproducible and recorded. The LLM sees a known, bounded subset of exactly the components the project registered, keeping every decision auditable.


Selection — K-Sample Voting

Selection is the strong-tier assemble call, and it does not trust a single completion. Per section it runs a K-sample self-consistency vote: the same evidence-pinned, low-temperature prompt is sampled K times, and the decision is the majority on (case, platform_design_id) — the same section KIND therefore converges on the same component even when a single draw would wobble. This is one LLM step, not K fence calls: the K samples ride the single per-section selection call, so the LLM fence is unchanged. Ties and every downstream trust concern are settled deterministically by the guards below; the vote proposes, the guards dispose.


Deterministic Guards over the Vote

The vote is a proposal. The engine never trusts the LLM's selection echo as final: a fixed program sits between the vote and the stitch (the trust boundary) and re-derives everything that must be exact. Every guard is deterministic and recorded — none is a second model call.

  • Variant clamp. Variant props are clamped to the registry axes before anything reads them: a value the model invented or an axis it dropped is snapped back to the component's real variant_properties, so no downstream step ever sees an off-axis variant.
  • Symmetric scope guards. Fit is checked both ways. A component grossly larger than the section it claims is downgraded to unmatched (a visible placeholder, never a silent misfit); a small component claiming a multi-unit section is downgraded to compose so its units re-select individually against the registry rather than collapsing into one wrong instance.
  • Duplicate-claim dedupe. When two sections claim the same component, the tie is broken with wireframe-name evidence — the section whose own wireframe name matches keeps the claim; the other re-resolves.
  • Kit-name override. A section that IS one named kit element takes the registry component of that name, carrying the kit's own variant values, rather than whatever the vote drifted to.
  • Focused variant-resolution pass. Under-specified instances get a second, narrow resolution of their variant axes only — the section and its one candidate, nothing else.
  • Screen-facts pass. Variant axes that are shared across a screen (a logged-in/out state, a platform) are resolved once per screen and applied consistently, so sibling sections don't disagree on a screen-wide fact.

Ledger floors (below) apply before these guards, so the guards always keep the final say: whatever the vote or a ledger floor proposed, the clamped, scope-checked, deduped result is what reaches the stitch.


Which guards may judge which picks

Guards fall into two tiers, and the split is a contract rather than an implementation detail:

Tier Guards Runs on
Size-based overscoped, underscoped Only picks that named no variant
Variant-independent kind_mismatched, heading_as_compound Every instance pick

The size tier is restricted because it reads default_size, which describes only the component's DEFAULT variant — a pick that names another variant may honestly be a fraction of that size (a slim footer variant on a set whose default is the full 786px footer), so judging it on the default's area would drop correct picks. Kind is different: a button_group is a button_group in every variant, so nothing about a named variant can excuse a category error.

This distinction was learned the hard way. The kind guard originally sat inside the size tier's not selection.variant_props block, which made it inert for any pick that echoed variant props — measured at 33 of 57 stored instance picks, 57%. The consequence on a live screen: a header region holding the page title 「ログイン」 took button container (kind button_group, which header does not admit) at 0.9 confidence, unflagged, and the generated design rendered the page heading as a second large dark login button above the input fields. Content alignment is what made it so confident — the title text matches a login button's own default label exactly, and the slot-fill guard only catches the opposite failure, text that rejects every slot.

other is not a category error. A component whose filed kind is other is the semantics tagger saying it does not know, exactly as an absent kind does, and ROLE_CANDIDATE_KINDS lists other for no role. Scoring it as a mismatch would therefore reject every such component everywhere — 19 of 279 registry components are filed other, among them logo, icon-slot and list dot. The guard abstains on it. (Measured: the moment the kind guard began running on variant-named picks, it rejected logo at a header region, where a logo plainly belongs.)

The ledger revalidates the variant-independent tier too. A stored decision is re-checked against these guards at consult time, so an entry written before a guard existed — or under a since-fixed bug — cannot lock its junk in. Kind was missing from that re-check, and that is what froze the defect above in place: the bad pick scored 0.88, cleared the trust threshold, locked three sections, and two consecutive runs produced byte-identical output. Fixing the vote path alone changed nothing, because selection never re-voted. Rejections are logged as ledger_entry_misselected with a reason naming the guard.


The design_resolution Decision Ledger

design_resolution is the sixth DocumentDB collection (see design_resolution) and the engine's autonomous quality ratchet. It exists because there is no human-correction loop in this worker — quality has to converge machine-side, so the engine remembers its own best decision per section KIND and lets a better decision replace a worse one over time. Each entry is keyed by a structural, LLM-free section_signature (sig@1:… — a versioned hash over the section's kit-component names, normalized name, child-type shape, and text count; never the LLM-assigned roles), so the same section KIND hits the same entry on every screen and every rerun.

The ledger touches assemble at two points.

  • Consult, before selection. Keyed by section_signature, an entry either locks its section or floors it. An entry with provenance = mined (read from designer-made designs in the file — the only above-vote authority) or a voted entry whose score ≥ the trust threshold LOCKS the section: no LLM runs for it. A weaker voted entry is a FLOOR — the section is re-voted and the higher-scoring decision wins. Floors are applied before the deterministic guards, so the guards keep the final say on anything the floor let through.
  • Write-back, after the guards. Every non-unmatched section decision (and each compose-child) upserts into design_resolution with a deterministic score (structural quality — base + modeled + variant-completeness + scope-fit; mined = 1.0). The write is ratchet-fenced: a voted write lands only over a non-mined entry that scores strictly lower; mined writes are unfenced. Identical evidence scores identically, so reruns never churn, and better decisions monotonically replace worse ones.

The ledger is deterministic on both sides — consult and write-back are program logic, and the score is a pure function of structural signals — so it adds no LLM calls and rides inside the existing assemble step. It is the mechanism by which the engine improves without a human in the loop.


Native Materialization and Render Review

The plugin consumes the self-contained spec and creates native component instances. Width is pinned to the planned region while component height stays intrinsic. In vertical auto-layout, a transparent spacer preserves the source/spec minimum footprint when an instance is shorter; this prevents later siblings from collapsing without stretching the component or divider.

Variant props, style/variable bindings, images, and text slots are applied only after the relevant component and fonts resolve. The materializer emits one of seven explicit diagnostics: name_fallback, ordinal_fallback, build_error, font_fallback, prop_rejected, unmatched, or preserved. Details are stored before the backend placement flag changes.

Actual-render correction is deliberately bounded. The plugin uploads a PNG, node/bounds manifest, and diagnostic report under a hash-derived render_review identity. A first review may pass or return one grounded candidate spec. The original remains untouched while that candidate is built adjacent and rendered again. The second event is verification-only: pass promotes the candidate in place; fail or unverifiable evidence deletes the candidate and retains the original. A parent review cannot spawn another correction, so recursive drift and unbounded model cost are impossible.


Deterministic Stitching

Once the LLM has returned per-section decisions, stitching assembles them into a self-contained DesignSpec by a fixed transform — the same decisions always stitch to the same spec.

The stitch produces three things:

  1. The spec tree (spec.root) — the five-case node tree above, with layer_path on every node so the plugin can address it deterministically.
  2. spec_nodes_flat — one flat entry per content node for review and feedback.
  3. The style_bindings map — token → Figma style/variable id, derived from project_figma_file.style_captures (the file-level local text styles + color variables captured by the sweep). The spec references style tokens; style_bindings resolves each token to a concrete Figma style so the plugin never has to look styles up itself.

The spec is self-contained by design. Every node carries its layer_path and embedded bbox; every style reference resolves through style_bindings; every image placeholder is a structural image_slot the component itself carries. The plugin can therefore materialize the design deterministically from the spec alone — same spec, same canvas result — without re-reading the registry, the rules, or the wireframe. This is what makes materialization a pure function of the spec rather than a second interpretation pass.


The Rules Validator

The compiled rule set is the validator's program. The rule set is not hand-authored as a strict rule file: at ingestion it is extracted from the project's selected Figma guideline boards — one vision extraction per board, the per-board fragments deterministically merged into a single RuleSet and stored with draft_source: "llm_extracted" pending designer review. However the rules were extracted, the validator itself is unchanged: a design_rule doc (one immutable doc per (design_rule_id, content_hash)) carries four rule classes, and the validator runs each as a deterministic pass over the stitched spec:

Rule class Binds to Enforces
spacing layout_frame auto-layout Allowed spacing scale + section gaps
typography text nodes Allowed style tokens + max lines per token
color_tokens style bindings The allowed token set
screen_group_policies layout_frame geometry Per screen-FAMILY layout: content width, paddings, section/element gaps, column bounds — selected by matching the policy's stated scope against the screen name
component_specs instance slot roles The guideline's per-component rules (variants, states, slot rules), joined to the slot-role classifier

The validator's outcome per violation is auto-fix where safe, else flag. A spacing value that snaps cleanly to the allowed scale is auto-fixed; a geometry the engine cannot safely resolve is flagged. The validator_report records every violation as {rule, layer_path, action (auto_fixed | flagged), detail}.

slot_constraints was removed. It bound a wf_role to an allowed set of platform_design_ids, but no extraction pass ever produced one, so the field was always empty and the binding never ran. Component-level rules now arrive as component_specs, which the guideline actually states, and the kind guard rejects a component whose filed category the region cannot be (a selection control chosen for a decorative badge).

The no-rules path is explicit. With no rules registered for the project, the validator is a no-op and sets validator_report.status = no_rules, and every node is flagged rules_unvalidated. Unvalidated is never silently treated as valid — the absence of a rule set is itself surfaced on every node.

The validator always overrides the LLM. The rules digest is fed to the selection step as advisory input, but a model choice that violates a rule is corrected or flagged by the validator regardless of what the model preferred. The deterministic pass is the authority.


Computed Confidence

Confidence is computed from structural signals by a versioned formula. The model's self-report is banned — the engine never asks the LLM "how confident are you?" and never records such an answer.

The inputs to the formula:

  • Validator violations — flagged (unfixed) violations lower confidence; auto-fixed ones count less.
  • Unmatched / composed counts — unmatched and compose outcomes are low-ground signals.
  • Slot ambiguity — a component whose text_slots / image_slots were flagged ambiguous at registration carries ambiguity into any instance that fills them.
  • Component-fit (role agreement) — how well the selected component's structure agrees with the section's WF role.

A hard rule lives in the formula itself: compose is capped below instance. A composed subtree can never score as confident as a clean registered-component instance, no matter how the other signals fall. This is the numeric expression of the "compose is always flagged" bright line.

The output confidence block is {min, avg, flag_count, formula_version}. flag_count counts every node the engine flagged for review — scores below the flag threshold, compose/unmatched outcomes, selection echoes clamped at the trust boundary (a dropped variant axis or slot path), and text nodes whose style token cannot resolve through style_bindings. It mirrors to wf2des.flag_count via the completion webhook, and formula_version (e.g. cf@0.6) pins the formula so a confidence number is always interpretable against the exact formula that produced it. Because every input is a structural signal recorded on the spec, confidence is reproducible from the spec + the pinned formula version — like the validator, it is recomputable rather than remembered.


Fencing & Reproducibility

The engine is built to be reproducible from its pinned decisions. Three mechanisms combine to guarantee it.

1. Result-doc CAS on (_id, attempt). Every generation DocDB write is a compare-and-swap on the result doc's _id (= the wf2des row id) and attempt. A zombie worker from a superseded attempt cannot commit its writes, and its late webhook is a no-op — the fencing pair (wf2des_id, attempt) is echoed back through the webhook and the backend advances the row attempt-guarded. There is no other CAS: internal runs carry no row and are fenced only by their content-addressed / LWW output docs.

2. Pinned inputs, fail-closed. The inputs block is pinned as the first act of parse and rehydrated as the first act of assemble; any hash mismatch on rehydrate aborts rather than falling back. A run therefore always assembles against exactly the rule revision, component snapshots, and LLM/prompt version it pinned — never a drifted "latest". Completed generations are never reprocessed into a full replace, because a re-run would overwrite committed human-event data with a different nondeterministic spec. A terminal row is final; re-running a wireframe creates a NEW row rather than replacing the old one.

3. Deterministic stitch + self-contained spec. Given the same LLM decisions, the stitch, the validator, and the confidence formula all produce the same output — only the two fenced LLM steps are nondeterministic, and their outputs are stored with evidence. The resulting spec is self-contained (style_bindings, layer_path, embedded bbox, image_slots), so the plugin materializes it deterministically without re-interpreting anything.

Put together: a generation run is reproducible from its pinned model decisions, and everything those decisions feed is a deterministic transform. The nondeterminism is confined to the fence; the rest of the pipeline is replayable.


Guideline boards on the rule revision

The per-section selection calls attach design-system guideline board images as vision reference — and the boards are not a separately curated set. They are the source boards of the pinned design_rule revision: the same guideline boards the rule extraction read at ingestion time, stored on the revision as boards[]. One pin, two consumers — at assemble, rehydrating the pinned revision yields both the merged RuleSet (the deterministic validator's program) and the revision's board images (loaded from S3 and attached as a vision prefix to every per-section selection call). Rehydration is fail-closed: a missing pinned revision doc or a missing board image is a ContractFailure — the run aborts rather than assembling without its pinned reference. With no rule revision registered there are no boards and no rules: the pin is None and the validator takes its explicit no_rules path.

The boards are placed first in each selection prompt so they form a provider-cached prefix shared across sections, and they are advisory only: like the rules digest, they inform selection — the LLM can prefer the component/variant whose styling and usage match the design system — but they never override the deterministic stitch, validator, or computed confidence, and the LLM may never invent ids or slots to satisfy them. When the strong tier is given boards it must be a vision-capable model. Reproducibility needs no separate set hash — the boards ride the rule revision's immutability, so the pinned revision fixes both the rules and the reference images. The boards ride the existing section-selection requests rather than a separate model stage.

Section-Parallelism

The strong-tier selection step decides sections independently and in parallel. Two consequences follow directly.

  • Wall-clock is the slowest section, not the sum. Because sections don't depend on one another during selection, they are decided concurrently; the phase's latency is bounded by the slowest single section rather than the total. (Soft p50 targets: confirm→spec under 2 minutes.)
  • Prompt caching shares the static prefix. Every section's prompt shares the same static prefix — the instructions and the rules digest — which is cached and reused across the section requests. Only the per-section evidence (the WF section, its filtered candidate set, memo intent) varies. This makes many small section decisions cheaper than one monolithic prompt, and it keeps each decision scoped to one section's evidence for clean review attribution.

Section-parallelism is why the engine can afford a bounded LLM call per section rather than one giant selection pass: the parallelism hides the per-call latency, and the shared cached prefix keeps the cost of many calls close to the cost of one.