Skip to content

Shared Packages

packages/ holds logic shared by more than one Lambda. Each package is a uv workspace member and is imported as packages.<name>.src.

Package Role Main consumers
models Beanie models for DocumentDB All Python Lambdas
race_result_facts Deterministic interpretation and checking of race results article_generation, blog_generation, odds_poc_app
race_analysis_agent Combined analysis of race video, course image and race result article_generation, blog_generation, blog_rewriter
media_analysis_sdk Low-level SDK for media analysis with Bedrock race_analysis_agent
article_generation_agent Writes the 10-second race result summary article_generation
blog_generation_agent Writes race recap / prediction blogs blog_generation
blog_parts Block-JSON handling plus results table and race card assembly blog_generation, blog_rewriter
thumbnail_generation_agent Selects the best moment for a thumbnail thumbnail_generation
ng_word_filter NG word screening (Aho-Corasick) article_generation, blog_generation, blog_rewriter
tone_of_voice Shared tone-of-voice guideline text article_generation_agent, blog_rewriter

models

Defines the DocumentDB document models (Beanie Documents), one file per collection under packages/models/documentDB/.

See Database for the details of each model.


race_result_facts

Interprets race_result (the official race result) deterministically, taking over the work that a machine can do reliably and that used to be left to the LLM.

Use Function Purpose
Eligibility void_reason / is_void Prevent generation for void / cancelled races (no finishing order)
Fact preparation build_fact_block Resolve lap- and corner-leaders down to "number + name" before handing them over
Post-generation checking find_fact_errors Verify number-to-name mapping and payout figures against the official data
Value substitution placeholders Fill entrant names, payouts and the like from the official data via placeholders
Popularity correction article_checks Correct directly written popularity and "the Nth favourite finished Mth" mix-ups
Provenance measurement source_mix Measure how much the article drew on the race result versus the video
Prediction facts prediction For race-prediction blogs. Because the reference is the race card (race.prediction) rather than the official result, it carries its own fact-block builder and checks (used by blog_generation_agent.generate_prediction)

Design principle

Asking an LLM to transcribe values or do multi-step matching produces mix-ups, so the division of labour is: the machine interprets, matches and transcribes; the LLM only writes the Japanese.

find_fact_errors only checks what can be asserted mechanically. Narrative descriptions (who led on which lap) are checked too, but only when a single entrant can be pinned down in the sentence (check_lead / check_position / check_lap_range_lead / check_never_led_lap / check_unsupported_running). Sentences whose subject cannot be resolved are left unjudged โ€” the design errs towards missing a real error rather than raising a false one, because a single false positive burns the regeneration budget and the article never gets published.

Prevention (handing over pre-matched facts via build_fact_block) and checking (cross-checking the output via find_fact_errors) are not alternatives โ€” both are used, in that order.


race_analysis_agent

An AI agent that analyses the race video, course image and race result data together into a single summary, built on a LangGraph ReAct agent and Amazon Bedrock (Nova Pro).

Input Handled by
Race video (mp4) The analyze_race_video tool (multimodal analysis)
Course image (horse racing only) The analyze_course_image tool
Race result JSON The agent itself

The agent combines all three and emits one summary.

Tools are memoised and capped at two calls each and four in total. Beyond that the loop is cut off and the agent is pushed to produce output; if the agent run fails, it falls back to a single tool-less call.


media_analysis_sdk

A low-level SDK for analysing media (video, images) with Bedrock, called by the race_analysis_agent tools.

Videos are passed by S3 URI, since they exceed the Converse API's 25 MB inline payload limit.


article_generation_agent

Writes the "10-second race result summary" (one headline plus three sentences).

  • The primary source is the fact block assembled by race_result_facts.build_fact_block
  • The combined analysis (output of race_analysis_agent) is used only as a stylistic reference
  • Entrant "number + name", payout amounts and popularity are never written by the LLM โ€” placeholders are filled in from the official data
  • After writing, NG word, fact-check, placeholder-resolution and length guards run; anything detected is fed back and the text is rewritten (up to twice)
  • Since it is tool-less text processing, it uses a single structured-output call rather than a ReAct agent

blog_generation_agent

Writes the body (block JSON) and the 10-second summary for race recap / prediction blogs. Its design mirrors article_generation_agent: the fact block is the primary source, values are substituted via placeholders, and guards run after generation.

The model call is injected by the caller (apps/blog_generation) as a converse function.


blog_parts

Shared building blocks for the payload handed to blog_builder, used by both blog_rewriter (rewrites) and blog_generation (original writing).

Module Role
blocks.py Parsing, validating, repairing and flattening the AI's "block JSON". Block kinds and shapes must stay in sync with blog_builder's articleBlockSchema
race_card.py Builds the full finishing-order table and the related-information race card for race-linked blogs

thumbnail_generation_agent

An AI agent that analyses the race video, race result and course image (horse racing only) to select the best moment for an article thumbnail (seconds from the start of the video), built on LangGraph and Amazon Bedrock (Nova Pro).

Frame extraction and S3/DB access belong to the caller (apps/thumbnail_generation). The output is {"timestamp_seconds": float, "reason": str}.


ng_word_filter

Screens generated text for prohibited words using Aho-Corasick.

Function Role
load_ng_words() Loads prohibited words from the ng_word collection in DocumentDB
build_matcher() Narrows the word set and builds the automaton
NgWordMatcher.find() Returns the prohibited words present in a text

Some words are excluded to avoid false positives (see the comments in filter.py).


tone_of_voice

Provides the oddspark tone-of-voice guideline text injected into generation prompts. Both race result articles (article_generation) and rewritten blogs (blog_rewriter) reference the same copy.

When the guideline changes, only src/prompts/tone_of_voice.yml needs editing.