Skip to content

guinness-ai-v2 Coding Rules


Core Principles

  • Use Python 3.12
  • Define all schemas with Pydantic โ€” avoid bare dict or TypedDict
  • Build agents with PydanticAI โ€” do not use LangChain or LangGraph
  • Add type hints to every function (Any is banned in principle)
  • Leverage warm starts โ€” initialize clients that can live outside the Lambda context at module level
  • Respect DB isolation โ€” AI workers must not access PostgreSQL or MySQL directly; all communication with the backend goes through Webhook only

Directory Structure

Each worker app follows this file layout.

apps/<worker>/src/<worker>/
  handler.py      # Lambda entry point โ€” SQS parsing, batchItemFailures
  service.py      # process_record() โ€” main processing pipeline
  schemas.py      # Pydantic models (SQS input, DocumentDB document, Webhook payload)
  repo.py         # I/O operations (S3 download, DocumentDB upsert, Webhook POST)
  config.py       # Environment variables via pydantic-settings
  prompts.py      # LLM prompt constants and PROMPT_VERSION

packages/
  agentic/        # PydanticAI agent definitions (shared across workers)
  models/         # DocumentDB collection definitions and index creation
  utils/          # Shared S3, Webhook, and logging utilities

Separation of Concerns

File What belongs here What does NOT belong here
handler.py SQS parsing, batchItemFailures construction, log initialization Business logic, DB operations
service.py Orchestration of the full processing pipeline Direct DB queries, prompt strings
repo.py S3 / DocumentDB / Webhook I/O Business rules
schemas.py Pydantic model definitions Logic
config.py Environment variable validation (pydantic-settings) Logic
prompts.py Prompt strings and version constants Logic

Naming Conventions

Target Convention Example
Files snake_case code_import/service.py
Classes PascalCase CodeImportMessage, CodeDocument
Functions / variables snake_case process_record, vector_embedding
Constants SCREAMING_SNAKE_CASE PROMPT_VERSION, VECTOR_DIMENSIONS
PydanticAI agents build_<name>_agent() build_vision_agent()
Pydantic settings class Config class Config(BaseSettings) in config.py

Lambda Handler Pattern

Warm Start Optimization

Initialize heavy clients at module level so they are reused across invocations, not recreated on every call.

# handler.py
from .config import Config
from .repo import DesignRepo
from packages.agentic.vision import build_vision_agent

config = Config()
repo = DesignRepo(config)        # reused across warm starts
agent = build_vision_agent(config)  # same for PydanticAI agents

def lambda_handler(event: dict, context: object) -> dict:
    ...

SQS batchItemFailures

Report partial failures at the record level โ€” do not retry the entire batch.

def lambda_handler(event: dict, context: object) -> dict:
    batch_item_failures = []

    for record in event["Records"]:
        try:
            message = CodeImportMessage.model_validate_json(record["body"])
            service.process_record(message, repo, agent)
        except Exception as e:
            logger.error("record failed", extra={"message_id": record["messageId"], "error": str(e)})
            batch_item_failures.append({"itemIdentifier": record["messageId"]})

    return {"batchItemFailures": batch_item_failures}

Pydantic Schemas

SQS Input Schema

Use model_validate_json to validate the SQS body. Enforce required invariants with model_validator.

from pydantic import BaseModel, model_validator
import uuid

class CodeImportMessage(BaseModel):
    code_id: str
    organization_id: int
    project_id: int
    name: str
    source_code: str
    css_code: str | None = None
    img_url: str | None = None

    @model_validator(mode="after")
    def validate_code_id(self) -> "CodeImportMessage":
        uuid.UUID(self.code_id)  # enforce UUID format
        return self

DocumentDB Document Schema

Define every document written to DocumentDB as a Pydantic model.

class EmbeddingBlock(BaseModel):
    metadata: dict
    vector_embedding: list[float]

class CodeDocument(BaseModel):
    id: str  # maps to _id
    organization_id: int
    project_id: int
    name: str
    source_code: str
    visual: EmbeddingBlock
    semantics: EmbeddingBlock
    index_schema_version: str

PydanticAI Agents

Agent Definition

Place all agents in packages/agentic/ and import them from worker apps.

# packages/agentic/src/agentic/vision.py
from pydantic import BaseModel
from pydantic_ai import Agent

class DesignDescription(BaseModel):
    layout: str
    component_types: list[str]
    color_palette: list[str]
    typography_style: str
    semantic_words: list[str]

def build_vision_agent(config: Config) -> Agent[None, DesignDescription]:
    return Agent(
        config.desc_model,  # provider-prefixed format: "openai:gpt-5.4-nano"
        output_type=DesignDescription,
        system_prompt=VISION_INSTRUCTIONS,
    )

Model Specification

Always specify models via environment variables โ€” never hard-code them. Use the provider:model-name format.

# Correct
config.desc_model  # "openai:gpt-5.4-nano"

# Forbidden
"gpt-5.4-nano"  # no provider prefix, hard-coded

Using run_sync

On Lambda, use synchronous execution (run_sync).

result = agent.run_sync(user_prompt, message_history=None)
description: DesignDescription = result.output

Embedding Generation

Always Use a Single Batch Call

Generate multiple embeddings in one embeddings.create call. Splitting into multiple calls risks dimension drift.

# Correct โ€” single batch
response = openai_client.embeddings.create(
    model=config.embedding_model,
    input=[visual_text, semantic_text],   # pass together
    dimensions=config.embedding_dimensions,
)
visual_vector = response.data[0].embedding
semantic_vector = response.data[1].embedding

# Forbidden โ€” two separate calls
visual_vector = openai_client.embeddings.create(input=visual_text, ...).data[0].embedding
semantic_vector = openai_client.embeddings.create(input=semantic_text, ...).data[0].embedding

Repository (I/O) Patterns

DocumentDB Upsert

Use the record ID as _id and write idempotently.

def upsert_code(self, doc: CodeDocument) -> None:
    self.collection.replace_one(
        {"_id": doc.id},
        doc.model_dump(by_alias=True),
        upsert=True,
    )

Webhook Notification

Always send a Webhook notification for both success and failure โ€” even when an exception occurs.

def notify_success(self, record_id: str, **payload) -> None:
    self._post_webhook({"status": "success", "record_id": record_id, **payload})

def notify_failure(self, record_id: str, error: str) -> None:
    self._post_webhook({"status": "failed", "record_id": record_id, "error": error})

Environment Variables

Manage all environment variables through BaseSettings from pydantic-settings to enforce validation at startup.

from pydantic_settings import BaseSettings

class Config(BaseSettings):
    desc_model: str
    embedding_model: str
    embedding_dimensions: int = 512
    openai_api_key: str
    documentdb_connection_string: str
    documentdb_name: str
    webhook_base_url: str

    class Config:
        env_file = ".env"
  • Required variables have no default โ€” they fail fast at startup if missing
  • Optional variables have defaults
  • Never log secrets (API keys, connection strings)

Error Handling

Retryable vs Non-Retryable Failures

Error type Action
Malformed SQS payload Discard without adding to batchItemFailures (retrying won't help)
Target resource not found Send Webhook failed and return
OpenAI timeout / rate limit Add to batchItemFailures to trigger SQS retry
DocumentDB connection error Add to batchItemFailures to trigger SQS retry
Webhook delivery failure Add to batchItemFailures to trigger SQS retry

try/except Granularity

Wrap the full process_record call and capture per-record errors in handler.py.

# service.py โ€” business errors become Webhook failures
def process_record(message: CodeImportMessage, repo: CodeRepo, agent: Agent) -> None:
    try:
        ...
    except DesignNotFoundError:
        repo.notify_failure(message.code_id, error="Design not found")
        return  # do NOT add to batchItemFailures โ€” retrying won't help

# handler.py โ€” infrastructure errors are left to SQS retry
try:
    service.process_record(message, repo, agent)
except Exception as e:
    batch_item_failures.append({"itemIdentifier": record["messageId"]})

Logging

Use packages/observability or the Python standard logging module for structured output. print() is forbidden.

import logging

logger = logging.getLogger(__name__)

# Correct
logger.info("process_record started", extra={"code_id": message.code_id, "project_id": message.project_id})
logger.error("webhook failed", extra={"code_id": message.code_id, "status_code": resp.status_code})

# Forbidden
print(f"processing {message.code_id}")

Testing

Test Types and Layout

apps/<worker>/
  __tests__/
    unit/          # Unit tests with mocked external dependencies
    contract/      # Contract tests for SQS / Webhook schemas
    integration/   # Integration tests against a real DocumentDB container

Unit Test Guidelines

  • Mock repo and test service.process_record in isolation
  • Use TestModel from PydanticAI to mock agents
  • Each test must be runnable independently
from pydantic_ai.models.test import TestModel

def test_process_record_success(mocker):
    agent = build_vision_agent.__wrapped__(TestModel())
    repo = mocker.Mock(spec=CodeRepo)
    message = CodeImportMessage(code_id=str(uuid.uuid4()), ...)

    service.process_record(message, repo, agent)

    repo.upsert_code.assert_called_once()
    repo.notify_success.assert_called_once()

Contract Tests

Write contract tests to catch schema changes in SQS messages and Webhook payloads.

def test_sqs_message_schema():
    raw = {"code_id": "660e...", "organization_id": 1, ...}
    msg = CodeImportMessage.model_validate(raw)
    assert msg.code_id == raw["code_id"]

Code Review Checklist

  • [ ] Lambda clients are initialized at module level (warm start)
  • [ ] SQS input is validated through a Pydantic model
  • [ ] Multiple embeddings are generated in a single batch call
  • [ ] DocumentDB writes use upsert (idempotent)
  • [ ] Both success and failure paths send a Webhook notification
  • [ ] Retryable errors are added to batchItemFailures
  • [ ] No print() calls โ€” structured logging is used
  • [ ] Environment variables are accessed through Config (no hard-coding)
  • [ ] Model names use the provider:model format
  • [ ] No direct access to PostgreSQL or MySQL