article_generation Lambda I/O Definition
Schemas for the article_generation Lambda. See the overview for the processing flow and responsibilities.
SQS event payload (input)
Messages received from the App Lambda over the SQS article-generation queue.
{
"article_id": "string",
"race_id": "string",
"racing_type": "horse_racing",
"title": "string",
"race_result": {},
"custom_prompt": "string"
}
| Field | Type | Required | Notes |
|---|---|---|---|
| article_id | string | โฏ | _id of the article row (in generating state) |
| race_id | string | โฏ | _id of one of the racing tables |
| racing_type | enum | โฏ | auto_racing / bicycle_racing / horse_racing |
| title | string | โฏ | Title / topic of the article to generate |
| race_result | object | โฏ | Race result, read from the racing table by the App Lambda and embedded in the payload |
| custom_prompt | string | Extra instructions on regeneration (from POST /v1/articles/{article_id}/regeneration) |
|
| category_id | string | Kept for backwards compatibility; the current App Lambda does not send it |
The race video and course image are not in the payload โ this Lambda fetches them from S3.
| Media | S3 key | How it is passed |
|---|---|---|
| Race video | {racing_type}/video/{race_id}.mp4 |
As an S3 URI (s3://{bucket}/{key}) |
| Course image (horse racing only) | horse_racing/course_image/{race_id}.gif |
As inline bytes |
If either is missing, analysis continues as if it did not exist.
Message sent to the article-builder queue (output)
{
"article_id": "string",
"race_id": "string",
"racing_type": "horse_racing",
"content": {
"headline": "string",
"summary": ["string", "string", "string"]
},
"accuracy_test_data": {}
}
| Field | Type | Required | Notes |
|---|---|---|---|
| article_id | string | โฏ | _id of the article row to update |
| race_id | string | โฏ | Race _id |
| racing_type | enum | โฏ | Selects the summary colour theme (horse racing = green, keirin = blue, auto racing = red) |
| content | object | โฏ | The generated 10-second race result summary |
| accuracy_test_data | object | Accuracy-verification data, stored as article.accuracy_test_data |
content schema
Mirrors ArticleContent in packages/article_generation_agent/src/schemas.py.
| Field | Type | Required | Notes |
|---|---|---|---|
| headline | string | โฏ | Summary headline |
| summary | string[] | โฏ | Exactly three sentences covering the key points; target 60 characters or fewer each |
accuracy_test_data schema
{
"generated_article": {
"title": "string",
"contents": ["string", "string", "string"]
},
"race_result_in_generated_article": {},
"race_analysis": "string"
}
| Field | Contents |
|---|---|
| generated_article | The generated article (headline + each summary line) |
| race_result_in_generated_article | The numeric facts the article used, declared by the generating AI in the same shape as race_result |
| race_analysis | Output of stage 1, kept so the race-result vs. video provenance mix can be measured |
scripts/article_accuracy_test/audit.py uses this data for matching and provenance aggregation.
Stage 1: race_analysis_agent
A LangGraph ReAct agent that analyses the race video, course image and race result together.
Input
| Field | Type | Required | Notes |
|---|---|---|---|
racing_type |
string | โฏ | auto_racing / bicycle_racing / horse_racing |
race_result |
object | โฏ | Race result |
media_files |
MediaFile[] | โฏ | Race video (s3_uri) and course image (data) |
client |
boto3 client | โฏ | Bedrock client initialised by the Lambda |
model_id |
string | โฏ | Nova Pro inference profile ARN |
A MediaFile carries file_name, mime_type (image / video) and either data (inline bytes) or
s3_uri.
Output
| Item | Contents |
|---|---|
summary |
A single combined summary of video, course image and race result |
Tools
| Tool | Input | What it does | Output |
|---|---|---|---|
analyze_race_video |
Race video (S3 URI) | Multimodal analysis of the race footage with Bedrock | Description of the analysis |
analyze_course_image |
Course image (bytes) | Course image analysis with Bedrock (horse racing only) | Description of the analysis |
Tools are memoised, so calling the same tool again returns the same result. Each tool is capped at two calls and the agent at four in total; anything beyond that is cut off and the agent is pushed to produce its output. If the agent run fails, it falls back to a single tool-less call.
Stage 2: article_generation_agent
A single structured-output call that writes the article from the confirmed facts and the stage-1 summary.
Not a ReAct agent
Stage 2 is pure text processing with no tools, so it uses a single with_structured_output call
rather than a ReAct agent.
Input
| Field | Type | Required | Notes |
|---|---|---|---|
title |
string | โฏ | Article title / topic |
race_result |
object | โฏ | Race result (used to build the fact block and fill placeholders) |
race_analysis |
string | โฏ | The stage-1 summary |
racing_type |
string | โฏ | Racing type |
ng_word_matcher |
NgWordMatcher | Matcher used for NG word screening | |
custom_prompt |
string | Extra instructions on regeneration | |
client / model_id / provider / base_model_id |
โฏ | Bedrock client and model selection |
Output
| Item | Contents |
|---|---|
headline / summary |
The generated 10-second summary |
race_result_in_generated_article |
The numeric facts the article mentioned, extracted by a separate structured-output call |
For race_result_in_generated_article, the LLM only outputs a selection of which facts it used;
the actual values are transcribed from the input race_result by code, because letting the LLM
transcribe them causes single-character corruption.