guinness-ai-v2 Coding Rules
Core Principles
- Use Python 3.12
- Define all schemas with Pydantic โ avoid bare
dictorTypedDict - Build agents with PydanticAI โ do not use LangChain or LangGraph
- Add type hints to every function (
Anyis banned in principle) - Leverage warm starts โ initialize clients that can live outside the Lambda context at module level
- Respect DB isolation โ AI workers must not access PostgreSQL or MySQL directly; all communication with the backend goes through Webhook only
Directory Structure
Each worker app follows this file layout.
apps/<worker>/src/<worker>/
handler.py # Lambda entry point โ SQS parsing, batchItemFailures
service.py # process_record() โ main processing pipeline
schemas.py # Pydantic models (SQS input, DocumentDB document, Webhook payload)
repo.py # I/O operations (S3 download, DocumentDB upsert, Webhook POST)
config.py # Environment variables via pydantic-settings
prompts.py # LLM prompt constants and PROMPT_VERSION
packages/
agentic/ # PydanticAI agent definitions (shared across workers)
models/ # DocumentDB collection definitions and index creation
utils/ # Shared S3, Webhook, and logging utilities
Separation of Concerns
| File | What belongs here | What does NOT belong here |
|---|---|---|
handler.py |
SQS parsing, batchItemFailures construction, log initialization |
Business logic, DB operations |
service.py |
Orchestration of the full processing pipeline | Direct DB queries, prompt strings |
repo.py |
S3 / DocumentDB / Webhook I/O | Business rules |
schemas.py |
Pydantic model definitions | Logic |
config.py |
Environment variable validation (pydantic-settings) | Logic |
prompts.py |
Prompt strings and version constants | Logic |
Naming Conventions
| Target | Convention | Example |
|---|---|---|
| Files | snake_case | code_import/service.py |
| Classes | PascalCase | CodeImportMessage, CodeDocument |
| Functions / variables | snake_case | process_record, vector_embedding |
| Constants | SCREAMING_SNAKE_CASE | PROMPT_VERSION, VECTOR_DIMENSIONS |
| PydanticAI agents | build_<name>_agent() |
build_vision_agent() |
| Pydantic settings class | Config |
class Config(BaseSettings) in config.py |
Lambda Handler Pattern
Warm Start Optimization
Initialize heavy clients at module level so they are reused across invocations, not recreated on every call.
# handler.py
from .config import Config
from .repo import DesignRepo
from packages.agentic.vision import build_vision_agent
config = Config()
repo = DesignRepo(config) # reused across warm starts
agent = build_vision_agent(config) # same for PydanticAI agents
def lambda_handler(event: dict, context: object) -> dict:
...
SQS batchItemFailures
Report partial failures at the record level โ do not retry the entire batch.
def lambda_handler(event: dict, context: object) -> dict:
batch_item_failures = []
for record in event["Records"]:
try:
message = CodeImportMessage.model_validate_json(record["body"])
service.process_record(message, repo, agent)
except Exception as e:
logger.error("record failed", extra={"message_id": record["messageId"], "error": str(e)})
batch_item_failures.append({"itemIdentifier": record["messageId"]})
return {"batchItemFailures": batch_item_failures}
Pydantic Schemas
SQS Input Schema
Use model_validate_json to validate the SQS body. Enforce required invariants with model_validator.
from pydantic import BaseModel, model_validator
import uuid
class CodeImportMessage(BaseModel):
code_id: str
organization_id: int
project_id: int
name: str
source_code: str
css_code: str | None = None
img_url: str | None = None
@model_validator(mode="after")
def validate_code_id(self) -> "CodeImportMessage":
uuid.UUID(self.code_id) # enforce UUID format
return self
DocumentDB Document Schema
Define every document written to DocumentDB as a Pydantic model.
class EmbeddingBlock(BaseModel):
metadata: dict
vector_embedding: list[float]
class CodeDocument(BaseModel):
id: str # maps to _id
organization_id: int
project_id: int
name: str
source_code: str
visual: EmbeddingBlock
semantics: EmbeddingBlock
index_schema_version: str
PydanticAI Agents
Agent Definition
Place all agents in packages/agentic/ and import them from worker apps.
# packages/agentic/src/agentic/vision.py
from pydantic import BaseModel
from pydantic_ai import Agent
class DesignDescription(BaseModel):
layout: str
component_types: list[str]
color_palette: list[str]
typography_style: str
semantic_words: list[str]
def build_vision_agent(config: Config) -> Agent[None, DesignDescription]:
return Agent(
config.desc_model, # provider-prefixed format: "openai:gpt-5.4-nano"
output_type=DesignDescription,
system_prompt=VISION_INSTRUCTIONS,
)
Model Specification
Always specify models via environment variables โ never hard-code them. Use the provider:model-name format.
# Correct
config.desc_model # "openai:gpt-5.4-nano"
# Forbidden
"gpt-5.4-nano" # no provider prefix, hard-coded
Using run_sync
On Lambda, use synchronous execution (run_sync).
result = agent.run_sync(user_prompt, message_history=None)
description: DesignDescription = result.output
Embedding Generation
Always Use a Single Batch Call
Generate multiple embeddings in one embeddings.create call. Splitting into multiple calls risks dimension drift.
# Correct โ single batch
response = openai_client.embeddings.create(
model=config.embedding_model,
input=[visual_text, semantic_text], # pass together
dimensions=config.embedding_dimensions,
)
visual_vector = response.data[0].embedding
semantic_vector = response.data[1].embedding
# Forbidden โ two separate calls
visual_vector = openai_client.embeddings.create(input=visual_text, ...).data[0].embedding
semantic_vector = openai_client.embeddings.create(input=semantic_text, ...).data[0].embedding
Repository (I/O) Patterns
DocumentDB Upsert
Use the record ID as _id and write idempotently.
def upsert_code(self, doc: CodeDocument) -> None:
self.collection.replace_one(
{"_id": doc.id},
doc.model_dump(by_alias=True),
upsert=True,
)
Webhook Notification
Always send a Webhook notification for both success and failure โ even when an exception occurs.
def notify_success(self, record_id: str, **payload) -> None:
self._post_webhook({"status": "success", "record_id": record_id, **payload})
def notify_failure(self, record_id: str, error: str) -> None:
self._post_webhook({"status": "failed", "record_id": record_id, "error": error})
Environment Variables
Manage all environment variables through BaseSettings from pydantic-settings to enforce validation at startup.
from pydantic_settings import BaseSettings
class Config(BaseSettings):
desc_model: str
embedding_model: str
embedding_dimensions: int = 512
openai_api_key: str
documentdb_connection_string: str
documentdb_name: str
webhook_base_url: str
class Config:
env_file = ".env"
- Required variables have no default โ they fail fast at startup if missing
- Optional variables have defaults
- Never log secrets (API keys, connection strings)
Error Handling
Retryable vs Non-Retryable Failures
| Error type | Action |
|---|---|
| Malformed SQS payload | Discard without adding to batchItemFailures (retrying won't help) |
| Target resource not found | Send Webhook failed and return |
| OpenAI timeout / rate limit | Add to batchItemFailures to trigger SQS retry |
| DocumentDB connection error | Add to batchItemFailures to trigger SQS retry |
| Webhook delivery failure | Add to batchItemFailures to trigger SQS retry |
try/except Granularity
Wrap the full process_record call and capture per-record errors in handler.py.
# service.py โ business errors become Webhook failures
def process_record(message: CodeImportMessage, repo: CodeRepo, agent: Agent) -> None:
try:
...
except DesignNotFoundError:
repo.notify_failure(message.code_id, error="Design not found")
return # do NOT add to batchItemFailures โ retrying won't help
# handler.py โ infrastructure errors are left to SQS retry
try:
service.process_record(message, repo, agent)
except Exception as e:
batch_item_failures.append({"itemIdentifier": record["messageId"]})
Logging
Use packages/observability or the Python standard logging module for structured output. print() is forbidden.
import logging
logger = logging.getLogger(__name__)
# Correct
logger.info("process_record started", extra={"code_id": message.code_id, "project_id": message.project_id})
logger.error("webhook failed", extra={"code_id": message.code_id, "status_code": resp.status_code})
# Forbidden
print(f"processing {message.code_id}")
Testing
Test Types and Layout
apps/<worker>/
__tests__/
unit/ # Unit tests with mocked external dependencies
contract/ # Contract tests for SQS / Webhook schemas
integration/ # Integration tests against a real DocumentDB container
Unit Test Guidelines
- Mock
repoand testservice.process_recordin isolation - Use
TestModelfrom PydanticAI to mock agents - Each test must be runnable independently
from pydantic_ai.models.test import TestModel
def test_process_record_success(mocker):
agent = build_vision_agent.__wrapped__(TestModel())
repo = mocker.Mock(spec=CodeRepo)
message = CodeImportMessage(code_id=str(uuid.uuid4()), ...)
service.process_record(message, repo, agent)
repo.upsert_code.assert_called_once()
repo.notify_success.assert_called_once()
Contract Tests
Write contract tests to catch schema changes in SQS messages and Webhook payloads.
def test_sqs_message_schema():
raw = {"code_id": "660e...", "organization_id": 1, ...}
msg = CodeImportMessage.model_validate(raw)
assert msg.code_id == raw["code_id"]
Code Review Checklist
- [ ] Lambda clients are initialized at module level (warm start)
- [ ] SQS input is validated through a Pydantic model
- [ ] Multiple embeddings are generated in a single batch call
- [ ] DocumentDB writes use upsert (idempotent)
- [ ] Both success and failure paths send a Webhook notification
- [ ] Retryable errors are added to
batchItemFailures - [ ] No
print()calls โ structured logging is used - [ ] Environment variables are accessed through
Config(no hard-coding) - [ ] Model names use the
provider:modelformat - [ ] No direct access to PostgreSQL or MySQL