Skip to content

article_generation Lambda I/O Definition

Schemas for the article_generation Lambda. See the overview for the processing flow and responsibilities.


SQS event payload (input)

Messages received from the App Lambda over the SQS article-generation queue.

{
  "article_id": "string",
  "race_id": "string",
  "racing_type": "horse_racing",
  "title": "string",
  "race_result": {},
  "custom_prompt": "string"
}
Field Type Required Notes
article_id string โ—ฏ _id of the article row (in generating state)
race_id string โ—ฏ _id of one of the racing tables
racing_type enum โ—ฏ auto_racing / bicycle_racing / horse_racing
title string โ—ฏ Title / topic of the article to generate
race_result object โ—ฏ Race result, read from the racing table by the App Lambda and embedded in the payload
custom_prompt string Extra instructions on regeneration (from POST /v1/articles/{article_id}/regeneration)
category_id string Kept for backwards compatibility; the current App Lambda does not send it

The race video and course image are not in the payload โ€” this Lambda fetches them from S3.

Media S3 key How it is passed
Race video {racing_type}/video/{race_id}.mp4 As an S3 URI (s3://{bucket}/{key})
Course image (horse racing only) horse_racing/course_image/{race_id}.gif As inline bytes

If either is missing, analysis continues as if it did not exist.


Message sent to the article-builder queue (output)

{
  "article_id": "string",
  "race_id": "string",
  "racing_type": "horse_racing",
  "content": {
    "headline": "string",
    "summary": ["string", "string", "string"]
  },
  "accuracy_test_data": {}
}
Field Type Required Notes
article_id string โ—ฏ _id of the article row to update
race_id string โ—ฏ Race _id
racing_type enum โ—ฏ Selects the summary colour theme (horse racing = green, keirin = blue, auto racing = red)
content object โ—ฏ The generated 10-second race result summary
accuracy_test_data object Accuracy-verification data, stored as article.accuracy_test_data

content schema

Mirrors ArticleContent in packages/article_generation_agent/src/schemas.py.

Field Type Required Notes
headline string โ—ฏ Summary headline
summary string[] โ—ฏ Exactly three sentences covering the key points; target 60 characters or fewer each

accuracy_test_data schema

{
  "generated_article": {
    "title": "string",
    "contents": ["string", "string", "string"]
  },
  "race_result_in_generated_article": {},
  "race_analysis": "string"
}
Field Contents
generated_article The generated article (headline + each summary line)
race_result_in_generated_article The numeric facts the article used, declared by the generating AI in the same shape as race_result
race_analysis Output of stage 1, kept so the race-result vs. video provenance mix can be measured

scripts/article_accuracy_test/audit.py uses this data for matching and provenance aggregation.


Stage 1: race_analysis_agent

A LangGraph ReAct agent that analyses the race video, course image and race result together.

Input

Field Type Required Notes
racing_type string โ—ฏ auto_racing / bicycle_racing / horse_racing
race_result object โ—ฏ Race result
media_files MediaFile[] โ—ฏ Race video (s3_uri) and course image (data)
client boto3 client โ—ฏ Bedrock client initialised by the Lambda
model_id string โ—ฏ Nova Pro inference profile ARN

A MediaFile carries file_name, mime_type (image / video) and either data (inline bytes) or s3_uri.

Output

Item Contents
summary A single combined summary of video, course image and race result

Tools

Tool Input What it does Output
analyze_race_video Race video (S3 URI) Multimodal analysis of the race footage with Bedrock Description of the analysis
analyze_course_image Course image (bytes) Course image analysis with Bedrock (horse racing only) Description of the analysis

Tools are memoised, so calling the same tool again returns the same result. Each tool is capped at two calls and the agent at four in total; anything beyond that is cut off and the agent is pushed to produce its output. If the agent run fails, it falls back to a single tool-less call.


Stage 2: article_generation_agent

A single structured-output call that writes the article from the confirmed facts and the stage-1 summary.

Not a ReAct agent

Stage 2 is pure text processing with no tools, so it uses a single with_structured_output call rather than a ReAct agent.

Input

Field Type Required Notes
title string โ—ฏ Article title / topic
race_result object โ—ฏ Race result (used to build the fact block and fill placeholders)
race_analysis string โ—ฏ The stage-1 summary
racing_type string โ—ฏ Racing type
ng_word_matcher NgWordMatcher Matcher used for NG word screening
custom_prompt string Extra instructions on regeneration
client / model_id / provider / base_model_id โ—ฏ Bedrock client and model selection

Output

Item Contents
headline / summary The generated 10-second summary
race_result_in_generated_article The numeric facts the article mentioned, extracted by a separate structured-output call

For race_result_in_generated_article, the LLM only outputs a selection of which facts it used; the actual values are transcribed from the input race_result by code, because letting the LLM transcribe them causes single-character corruption.