Shared Packages
packages/ holds logic shared by more than one Lambda. Each package is a uv workspace member and is
imported as packages.<name>.src.
| Package | Role | Main consumers |
|---|---|---|
| models | Beanie models for DocumentDB | All Python Lambdas |
| race_result_facts | Deterministic interpretation and checking of race results | article_generation, blog_generation, odds_poc_app |
| race_analysis_agent | Combined analysis of race video, course image and race result | article_generation, blog_generation, blog_rewriter |
| media_analysis_sdk | Low-level SDK for media analysis with Bedrock | race_analysis_agent |
| article_generation_agent | Writes the 10-second race result summary | article_generation |
| blog_generation_agent | Writes race recap / prediction blogs | blog_generation |
| blog_parts | Block-JSON handling plus results table and race card assembly | blog_generation, blog_rewriter |
| thumbnail_generation_agent | Selects the best moment for a thumbnail | thumbnail_generation |
| ng_word_filter | NG word screening (Aho-Corasick) | article_generation, blog_generation, blog_rewriter |
| tone_of_voice | Shared tone-of-voice guideline text | article_generation_agent, blog_rewriter |
models
Defines the DocumentDB document models (Beanie Documents), one file per collection under
packages/models/documentDB/.
See Database for the details of each model.
race_result_facts
Interprets race_result (the official race result) deterministically, taking over the work that a
machine can do reliably and that used to be left to the LLM.
| Use | Function | Purpose |
|---|---|---|
| Eligibility | void_reason / is_void |
Prevent generation for void / cancelled races (no finishing order) |
| Fact preparation | build_fact_block |
Resolve lap- and corner-leaders down to "number + name" before handing them over |
| Post-generation checking | find_fact_errors |
Verify number-to-name mapping and payout figures against the official data |
| Value substitution | placeholders |
Fill entrant names, payouts and the like from the official data via placeholders |
| Popularity correction | article_checks |
Correct directly written popularity and "the Nth favourite finished Mth" mix-ups |
| Provenance measurement | source_mix |
Measure how much the article drew on the race result versus the video |
| Prediction facts | prediction |
For race-prediction blogs. Because the reference is the race card (race.prediction) rather than the official result, it carries its own fact-block builder and checks (used by blog_generation_agent.generate_prediction) |
Design principle
Asking an LLM to transcribe values or do multi-step matching produces mix-ups, so the division of labour is: the machine interprets, matches and transcribes; the LLM only writes the Japanese.
find_fact_errors only checks what can be asserted mechanically. Narrative descriptions (who led on
which lap) are checked too, but only when a single entrant can be pinned down in the sentence
(check_lead / check_position / check_lap_range_lead / check_never_led_lap /
check_unsupported_running). Sentences whose subject cannot be resolved are left unjudged โ the
design errs towards missing a real error rather than raising a false one, because a single false
positive burns the regeneration budget and the article never gets published.
Prevention (handing over pre-matched facts via build_fact_block) and checking (cross-checking the
output via find_fact_errors) are not alternatives โ both are used, in that order.
race_analysis_agent
An AI agent that analyses the race video, course image and race result data together into a single summary, built on a LangGraph ReAct agent and Amazon Bedrock (Nova Pro).
| Input | Handled by |
|---|---|
| Race video (mp4) | The analyze_race_video tool (multimodal analysis) |
| Course image (horse racing only) | The analyze_course_image tool |
| Race result JSON | The agent itself |
The agent combines all three and emits one summary.
Tools are memoised and capped at two calls each and four in total. Beyond that the loop is cut off and the agent is pushed to produce output; if the agent run fails, it falls back to a single tool-less call.
media_analysis_sdk
A low-level SDK for analysing media (video, images) with Bedrock, called by the
race_analysis_agent tools.
Videos are passed by S3 URI, since they exceed the Converse API's 25 MB inline payload limit.
article_generation_agent
Writes the "10-second race result summary" (one headline plus three sentences).
- The primary source is the fact block assembled by
race_result_facts.build_fact_block - The combined analysis (output of
race_analysis_agent) is used only as a stylistic reference - Entrant "number + name", payout amounts and popularity are never written by the LLM โ placeholders are filled in from the official data
- After writing, NG word, fact-check, placeholder-resolution and length guards run; anything detected is fed back and the text is rewritten (up to twice)
- Since it is tool-less text processing, it uses a single structured-output call rather than a ReAct agent
blog_generation_agent
Writes the body (block JSON) and the 10-second summary for race recap / prediction blogs. Its design
mirrors article_generation_agent: the fact block is the primary source, values are substituted via
placeholders, and guards run after generation.
The model call is injected by the caller (apps/blog_generation) as a converse function.
blog_parts
Shared building blocks for the payload handed to blog_builder, used by both blog_rewriter (rewrites) and blog_generation (original writing).
| Module | Role |
|---|---|
blocks.py |
Parsing, validating, repairing and flattening the AI's "block JSON". Block kinds and shapes must stay in sync with blog_builder's articleBlockSchema |
race_card.py |
Builds the full finishing-order table and the related-information race card for race-linked blogs |
thumbnail_generation_agent
An AI agent that analyses the race video, race result and course image (horse racing only) to select the best moment for an article thumbnail (seconds from the start of the video), built on LangGraph and Amazon Bedrock (Nova Pro).
Frame extraction and S3/DB access belong to the caller (apps/thumbnail_generation). The output is
{"timestamp_seconds": float, "reason": str}.
ng_word_filter
Screens generated text for prohibited words using Aho-Corasick.
| Function | Role |
|---|---|
load_ng_words() |
Loads prohibited words from the ng_word collection in DocumentDB |
build_matcher() |
Narrows the word set and builds the automaton |
NgWordMatcher.find() |
Returns the prohibited words present in a text |
Some words are excluded to avoid false positives (see the comments in filter.py).
tone_of_voice
Provides the oddspark tone-of-voice guideline text injected into generation prompts. Both race result articles (article_generation) and rewritten blogs (blog_rewriter) reference the same copy.
When the guideline changes, only src/prompts/tone_of_voice.yml needs editing.