Skip to content

blog_generation Lambda Overview

Triggered by messages on the SQS blog-generation queue, this Lambda writes a blog body from scratch out of race data. The result is passed to the blog_builder Lambda over the SQS blog-builder queue, which renders it to HTML and stores it in the blog collection.

Unlike rewriting a past blog (blog_rewriter), there is no source article.


Kinds of generation

kind Content Input Race video analysis
recap (default) Race recap (post-race) Race result, race video, course image (horse racing only) Yes
prediction Race prediction (pre-race) Entry list and AI prediction marks (race.prediction) No โ€” the race has not run yet

For predictions the reference data is the entry list rather than the official race result, so result-derived values (finishing order, payouts, times) are not put into the fact block. Distance, race name and start time are known before the race, so those are filled in from the result data.


Trigger

Sender When
App Lambda (odds_poc_app) On a blog generation request (POST /v1/blogs/generate)

SQS event payload (input)

{
  "blog_id": "string",
  "race_id": "string",
  "racing_type": "horse_racing",
  "kind": "recap"
}
Field Type Required Notes
blog_id string โ—ฏ _id of the blog row (in generating state)
race_id string โ—ฏ Race _id
racing_type enum โ—ฏ auto_racing / bicycle_racing / horse_racing
kind enum recap / prediction; defaults to recap for compatibility with existing queued messages

Processing flow (recap)

  1. Receive the SQS event and parse the payload
  2. Connect to DocumentDB
  3. Check the target blog is still status=generating; otherwise exit without calling Bedrock
  4. Fetch the race result (error if missing)
  5. Build the full finishing-order table and the related-information race card (blog_parts.race_card)
  6. Fetch the race video and course image (horse racing only) from S3
  7. Stage 1: analyse the race with race_analysis_agent
  8. Stage 2: blog_generation_agent writes the body (block JSON) and the 10-second summary, grounded in the fact block
  9. Embed the generated body and store the vector in blog.embedding
  10. Send to the SQS blog-builder queue with the category "ใƒฌใƒผใ‚น็ตๆžœใƒปๅ›ž้กง"
flowchart TD
    AppLambda([App Lambda]) -->|SQS| Start
    Start[Receive SQS trigger] --> DB[Connect to DocumentDB]
    DB --> Guard[Check blog is generating]
    Guard -->|not generating| Skip[Exit without work]
    Guard -->|OK| Race[Fetch race result]
    Race -->|prediction| PredBranch[Write prediction from entries + AI marks]
    Race -->|recap| Table[Build results table and race card]
    Table --> Media[Fetch video / course image from S3]
    Media --> Analyze[Stage 1: race_analysis_agent]
    Analyze --> Write[Stage 2: blog_generation_agent writes body + summary]
    PredBranch --> Embed
    Write --> Embed[Embed the body and store it on the blog]
    Embed --> SendSQS[Send to blog-builder queue]
    SendSQS -.->|SQS| Builder([blog_builder Lambda])

Processing flow (prediction)

Differences from a recap:

  • No race video analysis (stage 1 is skipped)
  • Facts are checked against the entry list (race.prediction) instead of the race result
  • race_result_table is sent as null โ€” finishing order does not exist before the race
  • The category sent is "ใƒฌใƒผใ‚นไบˆๆƒณใƒปๆคœ่จผ"

If prediction data (race.prediction) is not registered, generation fails โ€” register it with PUT /v1/races/prediction first.


Message sent to the blog-builder queue (output)

{
  "blog_id": "string",
  "race_id": "string",
  "racing_type": "horse_racing",
  "category": "ใƒฌใƒผใ‚น็ตๆžœใƒปๅ›ž้กง",
  "race_result_table": {},
  "race_info": {},
  "content": {
    "title": "string",
    "body_html": "",
    "body_blocks": [],
    "summary": ["string"]
  }
}
Field Type Required Notes
blog_id string โ—ฏ _id of the blog to update
race_id string โ—ฏ Race _id
racing_type enum โ—ฏ Racing type
category string โ—ฏ Drives template selection (ใƒฌใƒผใ‚น็ตๆžœใƒปๅ›ž้กง / ใƒฌใƒผใ‚นไบˆๆƒณใƒปๆคœ่จผ)
race_result_table object | null Full finishing-order table; null for predictions
race_info object Contents of the race card shown in the related-information section
content object โ—ฏ Body (block JSON) and the 10-second summary

For originally written blogs, content.body_html is always empty and content.body_blocks is used.


Generation guards

After writing, blog_generation_agent runs the following guards; anything detected is fed back and the text is rewritten (up to twice). If it still cannot be resolved, generation fails and nothing is published.

Guard What it checks
Block JSON The output parses as block JSON (blog_parts.blocks)
Placeholders All placeholders resolved, and no approximate payouts ("over 100,000 yen", โ€ฆ)
Fact checking No contradiction with the official race result (race_result_facts.find_fact_errors)
NG words Screened, excluding matches inside entrant and jockey names

Entrant "number + name", payout amounts, combinations and popularity are never written by the LLM: they are placeholders filled in from the official data.


Models

Use Model
Race video / course image analysis (stage 1) Amazon Nova Pro (supports video input)
Body writing and summary (stage 2) Claude Sonnet (Nova Pro when BEDROCK_CLAUDE_SONNET_PROFILE_ARN is unset)
Embeddings Titan Text Embeddings v2 (amazon.titan-embed-text-v2:0)

The output token limit is 16384 on Claude and 8192 on Nova, because block JSON that gets truncated mid-output cannot be parsed.


Error handling

Situation Behaviour
Failure on a non-final attempt Re-raise and let SQS redelivery recover
Failure on the final attempt Hard-delete the status=generating blog and finish without re-raising

This differs from article_generation, which keeps failures as status=failed. Since the exception is swallowed, the stack trace is logged for investigation.