AI WF2Des — Overview
WF2Des turns a Figma wireframe into a hi-fi native Figma design by direct assembly under the project's design rules. For each wireframe section, an LLM selects a registered component (plus its variant, text, and image assets); a deterministic validator enforces the rules; and the Figma plugin materializes the result on the canvas. There is no retrieval, no embeddings, no vector search, and no learning corpus — the engine reads the project's own component registry and rule set directly and assembles from them. This is the defining difference from des2code, which reranks a candidate pool from indexed vectors.
One Lambda worker serves generation (two-phase parse → assemble) and six discriminated
wf2des-events messages: frame_registration, rule_upload, resync, component_upload,
component_capture, and render_review. Generation is product-row tracked; internal events
publish queued/processing/done/failed progress in wf2des_event_status when they carry an
event_run_id.
The backend owns the product contract — a wf2des PostgreSQL row plus the endpoints that create, read, confirm, cancel, place, and give feedback on a run — while the AI side owns the runs. Workers NEVER write PostgreSQL and hold no PG credentials. Generation completion flows back through the generation-only ai-status webhook, which the backend applies to the row inside one transaction; internal events never touch the webhook or flip the wf2des row.
- Queues: backend-owned generation SQS queue (parse / assemble); AI-owned
wf2des-eventsintake queue (rule / plugin events); AI-owned EventBridge schedule (component resync only) - Source:
guinness-ai-v2—apps/wf2des - Primary output: native Figma design materialized by the plugin; generation additionally writes an S3 result artifact + manifest and posts the
ai-statuswebhook - Stores: 1 PostgreSQL table
wf2des(backend-owned) + 6 DocumentDB collections + S3 artifacts - Status reporting:
ai-statuswebhook for generation only; internal events are webhook-free and usewf2des_event_statuswhen tracked
Tech Stack
| Layer | Technology |
|---|---|
| Runtime | Python 3.12 on AWS Lambda (container images; partial-batch SQS with ReportBatchItemFailures) |
| AI framework | PydanticAI (agent orchestration, provider-agnostic) — used only where a step explicitly calls an LLM |
| LLM tiers | Fast tier (parse / wf_parse roles + intent), strong tier (assemble section selection); model ids configured per deployment |
| Database | Amazon DocumentDB (guinness_v2, DocumentDB 5.0) — 6 wf2des collections |
| Object storage | Amazon S3 — snapshots, result artifacts + manifests, config |
| Queues | AWS SQS: backend-owned generation queue + AI-owned wf2des-events intake queue; AWS EventBridge schedule for component resync |
| Figma access | Figma REST via a service-account PAT (component sweep, rule_process board renders, snapshot / memo fallback) |
| Status notification | Webhook POST /v1/webhooks/ai-status with X-API-Key — generation only |
| Data-plane API | wf2des-api route group hosted in the internal-api app (DocDB + S3 only) |
| Logging | Structured logs via the shared observability package |
| RDB | No access — the worker never writes PostgreSQL / MySQL and holds no PG credentials |
| Retrieval | None — no embeddings, no vector index, no similarity search, no learning corpus |
The Generation Engine — Direct Assembly
Generation is the signature capability. It is a deterministic-first pipeline in which the LLM makes only bounded, evidence-pinned decisions and everything else is a reproducible transform. There is no similarity machinery anywhere: no embeddings, no vector search, no candidate reranking against an index, no learning corpus. The worker reads the project's own component registry and rule set and assembles directly from them.
The pipeline, in order:
- Deterministic WFNode extraction — the wireframe snapshot is walked into a WFNode tree (visible FRAME / GROUP / INSTANCE / TEXT descendants plus image-fill rectangles; vector leaves collapse into their parent; hidden layers skipped). The tree shape is deterministic — the LLM never invents or reshapes it.
- LLM assigns roles + intent only (fast tier) — each node gets a
rolefrom the shared element-type vocabulary and anintentderived from open memos and variant labels. This is the only LLM step in parse, and it never touches the tree shape. - Deterministic candidate set — the project registry is deduped and hygiene-collapsed, then the production-default role→kind gate removes structurally inadmissible component kinds for each section. There is no vector scoring or similarity retrieval — a straight, tenant-scoped set; size remains a selection signal, never a gate.
- LLM section-parallel selection (strong tier) — sections are decided independently and in parallel. Per section the model chooses one of a small, closed set of cases (below), a component key, variant props, and text-slot fills from that section's admissible candidates, judged by each candidate's full structure through K-sample self-consistency voting (the majority resolution across K samples wins). Memo intent outranks the request prompt; the rules digest is advisory input — the validator always overrides. Before selection the run consults the
design_resolutiondecision ledger (the autonomous quality ratchet) and writes the chosen resolution back after — because there is no human-correction loop, quality converges machine-side. - Deterministic stitching — the per-section decisions are stitched into a self-contained
DesignSpec(spec tree +spec_nodes_flat+ astyle_bindingsmap). The spec is reproducible from the same decisions. - Deterministic rules validator — violations are auto-fixed where safe, otherwise flagged. With no rules registered, the validator is a no-op and every node is flagged
rules_unvalidated. - Computed confidence — confidence is computed from validator violations, unmatched/composed counts, slot ambiguity, and component-fit (role agreement). The model's self-report is never used.
- The plugin materializes the spec into native Figma on the canvas. A bounded actual-render review may retain it or test one adjacent correction; only a verified correction replaces the original.
The spec has exactly five node cases:
| Spec node | Meaning |
|---|---|
layout_frame |
An auto-layout container (direction, gap, padding, sizing) holding child nodes |
instance |
A registered component instance — component key, variant props, text-slot and asset fills |
compose |
A composed subtree (no single component fits) — always flagged as least-grounded |
text |
A text node with content bound to a style token |
unmatched |
A placeholder for a node with no confident match — flagged, never silently dropped |
The bright lines: compose sections are always flagged, unmatched nodes always become visible placeholders, and confidence is never a model self-report. Together these keep every low-confidence decision surfaced for the designer's review rather than hidden.
Run Families
One worker handles generation plus five internal event paths. Generation is the interactive,
job-tracked flow. Internal events use wf2des_event_status, not the generation ai-status
webhook; wf_parse and rule_process are also PG-free. component_sweep alone also produces
a backend-side registry effect (a design type=component upsert from its sweep manifest) —
never a wf2des-row flip; render_review evaluates plugin-rendered evidence and permits at
most one correction plus one verification. See the I/O Definition for
message and field detail.
| Run | Trigger | What it does | Primary output |
|---|---|---|---|
| Generation (parse + assemble) | Backend-owned generation SQS queue | Parses a wireframe frame, parks for confirm (or auto-confirms), then selects components, stitches + validates the spec, computes confidence | Native Figma design (plugin) + S3 result artifact + manifest + ai-status webhook |
wf_parse |
wf2des-events — plugin frame-registration event |
Same parse pipeline as generation-parse, but writes only the wireframe cache doc — no result doc, preview, or webhook |
wireframe cache document |
rule_process |
wf2des-events — rule-upload event |
Renders the selected guideline boards (Figma REST, PAT), vision-extracts each board, and deterministically merges the fragments into one immutable design_rule revision (boards + rules on one doc) |
design_rule document |
component_sweep |
EventBridge schedule, or wf2des-events plugin resync event |
Deterministic Figma REST full-file walk; refreshes component instancing context + emits a discovered-component manifest for the registry (no LLM) | design_component docs + component snapshots + sweep manifest |
component_capture |
wf2des-events — plugin capture event |
Merges plugin-only publish keys, axes, and variant evidence without downgrading sweep-derived fields | Enriched design_component docs + event status |
render_review |
wf2des-events — plugin render evidence |
Reviews actual PNG + node geometry + diagnostics; may propose one grounded correction, then accepts verification only | Review result/candidate artifacts + event status |
Backend ↔ AI Split
Ownership is split along a clean control-plane / data-plane line.
Backend — the product control plane. The backend owns the wf2des PostgreSQL row (status / phase / attempt) and every product-facing endpoint: create, get, confirm, cancel, placement, and feedback. It also hosts the ai-status webhook handler, which is the sole completion-time PG writer — the only place a generation run's row is advanced to a terminal state. The backend writes the row first (before enqueuing a generation job) and applies the completion webhook attempt-guarded inside one transaction.
AI side — the worker and its runs. The AI side owns the apps/wf2des worker and the run plumbing: the backend-owned generation queue drives generation (parse / assemble), the AI-owned wf2des-events intake queue drives the internal runs (the backend holds send permission only), and an AI-owned EventBridge schedule drives component resync. The AI side also owns the wf2des-api data plane — a route group inside the internal-api app that reads and writes DocumentDB + S3 only, used by the plugin for previews, placement, and file registration.
The hard rule. Workers NEVER write PostgreSQL and hold no PG credentials. Workers fence their DocumentDB commits on (wf2des_id, attempt) and report both back through the webhook; the backend advances the row attempt-guarded. Every PG effect in the system is applied backend-side — by the ai-status handler or by the confirm / placement / feedback endpoints — never by the worker.
Processing Flow
The generation happy path (interactive confirm shown, with the auto-confirm shortcut):
flowchart TD
A["Backend: INSERT wf2des row<br/>status '0' phase parse (row FIRST)<br/>+ capture WF snapshot"] --> B["SendMessage → generation queue<br/>(parse: wf2des_id, attempt, nonce, URLs, hash)"]
B --> C["Worker: parse<br/>deterministic WFNode extraction<br/>+ LLM roles/intent + parse.json"]
C --> D["Webhook: parse_done"]
D --> E["ai-status handler:<br/>wf2des.phase → '1' awaiting_confirm"]
E --> F{"auto_confirm?"}
F -->|no| G["Plugin previews via wf2des-api<br/>→ designer confirms"]
G --> H["Backend confirm endpoint<br/>CAS on awaiting_confirm + attempt<br/>→ phase '2' assemble → enqueue assemble"]
F -->|yes| I["assemble chained in the<br/>same parse invocation<br/>(no awaiting_confirm park)"]
H --> J["Worker: assemble<br/>rehydrate pins → role/kind candidate gate<br/>→ LLM section selection → stitch spec<br/>→ rules validator → computed confidence"]
I --> J
J --> K["S3: result artifact + manifest"]
K --> L["Webhook: succeeded + manifest key"]
L --> M["ai-status handler:<br/>wf2des status → '1' completed<br/>+ result refs + flag_count (one PG txn)"]
M --> N["Plugin materializes native Figma<br/>on the canvas"]
| Phase | Processing | Output |
|---|---|---|
| Row + snapshot | Backend inserts the wf2des row (status '0', phase parse) and captures the trigger-time wireframe snapshot, then enqueues |
Row exists before the job; parse message on the generation queue |
| Parse | Deterministic WFNode extraction + memo statuses; LLM assigns roles + intent; write the wireframe cache doc + immutable parse.json + result-doc parse block |
parse.json artifact + parse_done webhook |
| Awaiting confirm | Handler parks the row at phase awaiting_confirm; the plugin previews via wf2des-api and the designer confirms (skipped on auto-confirm) |
phase '1' awaiting_confirm; assemble message on confirm |
| Assemble | Rehydrate pins; build role/kind-admissible candidate sets; LLM section-parallel selection (K-sample voting, design_resolution ledger consulted then written back); stitch spec; run the rules validator; compute confidence |
Result-doc spec / validator / confidence blocks |
| Complete | Write the S3 result artifact + manifest; post the succeeded webhook; the handler flips the row to completed in one PG transaction |
status '1' completed + result refs + flag_count |
| Materialize | The plugin reads the spec and builds the native Figma design on the canvas; placement fields are written back through wf2des-api |
Native Figma design |
The auto_confirm path skips the awaiting_confirm park entirely — assemble runs back-to-back inside the same parse invocation, and only the succeeded webhook fires.
Stores
WF2Des touches one PostgreSQL table (backend-owned) and six DocumentDB collections (AI-owned), plus S3 for artifacts and snapshots. See the I/O Definition for full document shapes and fencing.
| Store | Purpose |
|---|---|
PostgreSQL wf2des (backend-owned) |
The generation run row — status / phase / attempt + result refs + flag_count. Written only by the backend; internal runs carry no row |
DocumentDB wireframe |
Parsed wireframe cache doc (WFNode tree + memos + summaries); written by generation-parse and wf_parse |
DocumentDB design_rule |
Immutable rule revision doc — one per (design_rule_id, content_hash); each revision holds the merged extracted ruleset and the source guideline board images (the selection LLM's vision reference); the collection is the rule's revision history |
DocumentDB design_component |
Component instancing context only (variant props, text/image slots, default size) — the registry lifecycle lives on the platform design row |
DocumentDB project_figma_file |
Per-file style captures + sweep watermark / marker; scopes the component sweep |
DocumentDB design_generation_result |
The generation run's own document — pinned inputs, parse, spec, selection, validator_report, confidence, placement |
DocumentDB design_resolution |
The decision ledger — structural section-signature → component resolution; mined/voted provenance + deterministic score (the autonomous quality ratchet consulted before selection and written back after). See the I/O Definition for the full shape |
The component registry is the shared platform design table (type=component) — the same catalog des2code consumes — not a wf2des-only table. component_sweep emits a manifest that the backend applies as design row upserts; the design_component collection holds only the assembly/instancing context keyed to those rows.
Module Structure
Follows the platform worker recipe (handler / service / schemas / repo, as in design-import and des2code), with the deterministic engine and the rule-ingestion pipeline split into their own packages.
apps/wf2des/src/wf2des/
handler.py # Lambda entry — SQS / event parsing, batchItemFailures, warm-start clients
service/ # one module per RUN FAMILY (domains decided by call-graph reachability)
dispatch.py # process_record() — one record in, one run family out
generation.py # parse + board comprehension + assemble
rules.py # rule_process
sweep.py # component_sweep
registration.py # wf_parse
render_review.py # render_review
shared.py # helpers reached by more than one family
engine/ # deterministic: no I/O, no network
candidates.py # the candidate set offered to selection
stitch.py # decisions in, DesignSpec out
validate.py # rules validator + computed confidence
ingest/ # the NODE rule pipeline: comprehend → per-role extract → deterministic merge
schemas/ # Pydantic models, one module per domain; the package IS the contract surface
llm.py # the LLM fence — the ONLY Agent constructions in the worker
nodes.py # wireframe-node helpers shared by the engine and the run phases
repo.py # DocumentDB + S3 + Figma REST + ai-status webhook I/O
figma.py # snapshot parse / WFNode extraction / component-sweep full-file walk
prompts.py # every prompt + its PROMPT_VERSION pin
constants.py # domain constants (thresholds, vocabularies, formula pins)
usage.py # per-stage LLM usage accounting
config.py # Config (pydantic-settings)
| Module | Responsibility |
|---|---|
handler.py |
Lambda handler across all triggers (SQS + EventBridge); batchItemFailures |
service/ |
One module per run family; dispatch.process_record() routes generation and each discriminated internal event |
engine/ |
Deterministic candidate selection, stitch and validation — pure functions, no I/O |
ingest/ |
The node rule pipeline (live): comprehend, per-role extract, authority-gated merge |
schemas/ |
Pydantic message + document models = the contract surface (bump index_schema_version on shape change) |
llm.py |
Every Agent construction — model ids, prompt pins and settings in one place |
repo.py |
DocumentDB upserts, S3 get/put, Figma REST reads, webhook POST |
figma.py |
Snapshot parsing, WFNode extraction, and the component-sweep walk |
usage.py |
Per-stage LLM usage, persisted onto the rule and generation documents |
config.py |
Environment variables (pydantic-settings) |
Environment & Configuration
The wf2des-relevant environment and configuration. Concrete model ids, queue URLs, schedule rate, and secret ARNs are configured per deployment.
| Variable / Config | Description | Example / Default |
|---|---|---|
DOCUMENTDB_CONNECTION_STRING |
DocumentDB connection string | mongodb://user:pass@host:27017/?tls=true |
DOCUMENTDB_NAME |
DocumentDB database name | guinness_v2 |
DOCUMENTDB_CA_PATH |
TLS CA bundle path | /var/task/global-bundle.pem |
| Per-collection names | Env-overridable collection names for the 6 collections | wireframe, design_rule, design_component, project_figma_file, design_generation_result, design_resolution |
| Generation queue | Backend-owned SQS queue driving parse / assemble | queue URL per deployment |
wf2des-events queue |
AI-owned intake queue (rule / plugin events; backend send-only) | queue URL per deployment |
| EventBridge schedule | AI-owned schedule invoking component_sweep resync sweeps |
cron / rate per deployment |
WEBHOOK_BASE_URL |
Backend ai-status webhook endpoint (private VPC) |
https://api.internal/v1/webhooks/ai-status |
WEBHOOK_API_KEY |
Shared service key sent as X-API-Key |
Secrets Manager value |
| S3 result bucket | Backend-owned AI bucket for result artifacts + manifests | bucket name per deployment |
| S3 result key prefix | Client-facing artifact prefix | {org}/{proj}/wf2des/ |
| Figma service-account PAT | wf2design's own secret store — used by component_sweep, rule_process board renders, + snapshot / memo fallback |
secret ARN per deployment (never the platform figma_token) |
DEFAULT_ORGANIZATION_ID |
Single-tenant default; organization_id on every doc |
1 |
FAST_MODEL |
Parse roles/intent | openai:gpt-4o-mini |
STRONG_MODEL |
Section matching and rule consolidation; the matching model is pinned into inputs.llm per run |
openai:gpt-5.4 |
VISION_MODEL |
Rule-board comprehension/extraction | openai:gpt-5.4-mini |
SCREEN_REVIEW_MODEL / SCREEN_REVIEW_REASONING_EFFORT |
Whole-screen planning and actual-render review; pin per run | openai:gpt-5.4 / high; use none for non-GPT models |
CANDIDATE_KIND_GATE / SCREEN_PLANNING_ENABLED |
Production selection/review switches | true / true |
PROMPT_VERSION |
Pinned prompt version, per the platform worker recipe | pinned |
All environment variables are documented in AI Infrastructure — Environment Variables.
Official field-level contract: I/O Definition.