blog_generation Lambda Overview
Triggered by messages on the SQS blog-generation queue, this Lambda writes a blog body from
scratch out of race data. The result is passed to the blog_builder Lambda over the SQS
blog-builder queue, which renders it to HTML and stores it in the blog collection.
Unlike rewriting a past blog (blog_rewriter), there is no source article.
Kinds of generation
| kind | Content | Input | Race video analysis |
|---|---|---|---|
recap (default) |
Race recap (post-race) | Race result, race video, course image (horse racing only) | Yes |
prediction |
Race prediction (pre-race) | Entry list and AI prediction marks (race.prediction) |
No โ the race has not run yet |
For predictions the reference data is the entry list rather than the official race result, so result-derived values (finishing order, payouts, times) are not put into the fact block. Distance, race name and start time are known before the race, so those are filled in from the result data.
Trigger
| Sender | When |
|---|---|
App Lambda (odds_poc_app) |
On a blog generation request (POST /v1/blogs/generate) |
SQS event payload (input)
| Field | Type | Required | Notes |
|---|---|---|---|
| blog_id | string | โฏ | _id of the blog row (in generating state) |
| race_id | string | โฏ | Race _id |
| racing_type | enum | โฏ | auto_racing / bicycle_racing / horse_racing |
| kind | enum | recap / prediction; defaults to recap for compatibility with existing queued messages |
Processing flow (recap)
- Receive the SQS event and parse the payload
- Connect to DocumentDB
- Check the target
blogis stillstatus=generating; otherwise exit without calling Bedrock - Fetch the race result (error if missing)
- Build the full finishing-order table and the related-information race card (
blog_parts.race_card) - Fetch the race video and course image (horse racing only) from S3
- Stage 1: analyse the race with
race_analysis_agent - Stage 2:
blog_generation_agentwrites the body (block JSON) and the 10-second summary, grounded in the fact block - Embed the generated body and store the vector in
blog.embedding - Send to the SQS
blog-builderqueue with the category "ใฌใผใน็ตๆใปๅ้กง"
flowchart TD
AppLambda([App Lambda]) -->|SQS| Start
Start[Receive SQS trigger] --> DB[Connect to DocumentDB]
DB --> Guard[Check blog is generating]
Guard -->|not generating| Skip[Exit without work]
Guard -->|OK| Race[Fetch race result]
Race -->|prediction| PredBranch[Write prediction from entries + AI marks]
Race -->|recap| Table[Build results table and race card]
Table --> Media[Fetch video / course image from S3]
Media --> Analyze[Stage 1: race_analysis_agent]
Analyze --> Write[Stage 2: blog_generation_agent writes body + summary]
PredBranch --> Embed
Write --> Embed[Embed the body and store it on the blog]
Embed --> SendSQS[Send to blog-builder queue]
SendSQS -.->|SQS| Builder([blog_builder Lambda])
Processing flow (prediction)
Differences from a recap:
- No race video analysis (stage 1 is skipped)
- Facts are checked against the entry list (
race.prediction) instead of the race result race_result_tableis sent asnullโ finishing order does not exist before the race- The category sent is "ใฌใผในไบๆณใปๆค่จผ"
If prediction data (race.prediction) is not registered, generation fails โ register it with
PUT /v1/races/prediction first.
Message sent to the blog-builder queue (output)
{
"blog_id": "string",
"race_id": "string",
"racing_type": "horse_racing",
"category": "ใฌใผใน็ตๆใปๅ้กง",
"race_result_table": {},
"race_info": {},
"content": {
"title": "string",
"body_html": "",
"body_blocks": [],
"summary": ["string"]
}
}
| Field | Type | Required | Notes |
|---|---|---|---|
| blog_id | string | โฏ | _id of the blog to update |
| race_id | string | โฏ | Race _id |
| racing_type | enum | โฏ | Racing type |
| category | string | โฏ | Drives template selection (ใฌใผใน็ตๆใปๅ้กง / ใฌใผในไบๆณใปๆค่จผ) |
| race_result_table | object | null | Full finishing-order table; null for predictions |
|
| race_info | object | Contents of the race card shown in the related-information section | |
| content | object | โฏ | Body (block JSON) and the 10-second summary |
For originally written blogs, content.body_html is always empty and content.body_blocks is used.
Generation guards
After writing, blog_generation_agent runs the following guards; anything detected is fed back and
the text is rewritten (up to twice). If it still cannot be resolved, generation fails and nothing is
published.
| Guard | What it checks |
|---|---|
| Block JSON | The output parses as block JSON (blog_parts.blocks) |
| Placeholders | All placeholders resolved, and no approximate payouts ("over 100,000 yen", โฆ) |
| Fact checking | No contradiction with the official race result (race_result_facts.find_fact_errors) |
| NG words | Screened, excluding matches inside entrant and jockey names |
Entrant "number + name", payout amounts, combinations and popularity are never written by the LLM: they are placeholders filled in from the official data.
Models
| Use | Model |
|---|---|
| Race video / course image analysis (stage 1) | Amazon Nova Pro (supports video input) |
| Body writing and summary (stage 2) | Claude Sonnet (Nova Pro when BEDROCK_CLAUDE_SONNET_PROFILE_ARN is unset) |
| Embeddings | Titan Text Embeddings v2 (amazon.titan-embed-text-v2:0) |
The output token limit is 16384 on Claude and 8192 on Nova, because block JSON that gets truncated mid-output cannot be parsed.
Error handling
| Situation | Behaviour |
|---|---|
| Failure on a non-final attempt | Re-raise and let SQS redelivery recover |
| Failure on the final attempt | Hard-delete the status=generating blog and finish without re-raising |
This differs from article_generation, which keeps failures as status=failed. Since the exception
is swallowed, the stack trace is logged for investigation.