blog_rewriter Lambda Overview
Triggered by messages on the SQS blog-rewriter queue, this Lambda rewrites the body of a past blog
(old_blog) without changing its content. When the blog is tied to a race, the race result, race
video and course information (horse racing only) are taken into account as well.
The rewritten body is passed to the blog_builder Lambda over the SQS blog-builder queue, which
renders it to HTML and stores it in the blog collection.
Trigger
| Sender | When |
|---|---|
App Lambda (odds_poc_app) |
On a rewrite request (POST /v1/old_blogs/{old_blog_id}/rewrite) |
App Lambda (odds_poc_app) |
On a blog generation request (POST /v1/blogs/generation, via the blog-rewrite queue) |
SQS event payload (input)
{
"blog_id": "string",
"old_blog_id": "string",
"race_id": "string",
"racing_type": "horse_racing"
}
| Field | Type | Required | Notes |
|---|---|---|---|
| blog_id | string | โฏ | _id of the blog row (in generating state) |
| old_blog_id | string | โฏ | old_blog._id being rewritten |
| race_id | string | _id of the linked race; optional, since not every past blog has one |
|
| racing_type | enum | auto_racing / bicycle_racing / horse_racing |
Processing flow
- Receive the SQS event and parse the payload
- Connect to DocumentDB
- Check the target
blogis stillstatus=generating; otherwise exit without calling Bedrock - Fetch the
old_blog - If the blog is tied to a race
- Fetch the race result and build the full finishing-order table and the related-information race card
- Fetch the race video and course image (horse racing only) from S3
- Stage 1: analyse the race with
race_analysis_agent - If it is not tied to a race, this stage is skipped and the rewrite proceeds with "no race information"
- Stage 2: rewrite the body in light of the analysis, structured as block JSON
- Stage 3: generate the "10-second summary" from the rewritten body
- Send to the SQS
blog-builderqueue
flowchart TD
AppLambda([App Lambda]) -->|SQS| Start
Start[Receive SQS trigger] --> DB[Connect to DocumentDB]
DB --> Guard[Check blog is generating]
Guard -->|not generating| Skip[Exit without work]
Guard -->|OK| Old[Fetch old_blog]
Old -->|race linked| Race[Fetch race result, build table and race card]
Old -->|no race| NoRace[Continue with no race information]
Race --> Media[Fetch video / course image from S3]
Media --> Analyze[Stage 1: race_analysis_agent]
Analyze --> Rewrite
NoRace --> Rewrite[Stage 2: rewrite body into block JSON]
Rewrite --> Summary[Stage 3: generate the 10-second summary]
Summary --> SendSQS[Send to blog-builder queue]
SendSQS -.->|SQS| Builder([blog_builder Lambda])
Rewriting policy
- Do not change the content โ preserve the original information and the author's voice
- The shared tone-of-voice text from
packages/tone_of_voiceis injected into the prompt - The body is structured as an array of "blocks" matched to the content (see blog_builder for the block types)
- Relative image
srcvalues are absolutized against the source article URL (old_blog.source_url)- Body images of past blogs are served from the same host as the source article. Only when
source_urlis missing does it fall back tohttps://www.oddspark.com/
- Body images of past blogs are served from the same host as the source article. Only when
- For race-linked blogs, entry-list enumerations are deleted rather than rewritten (entrant and finishing-order data is rendered separately from the official data)
- The prompt says so, and the code drops enumeration blocks as a safety net when the LLM does not comply
- The colour of a prediction number badge is determined by the bracket number in horse racing. Using the horse number would give the wrong colour, so the bracket number is looked up from the race result by horse name; if it cannot be resolved, no colour is applied (no colour is more accurate than a wrong one)
Regeneration guards
| Guard | What it checks | Limit |
|---|---|---|
| Block JSON | The output parses as block JSON. If it is broken or truncated by the token limit, the error is attached as a correction instruction and the text is rewritten | 2 retries |
| NG words | The rewritten text contains no prohibited words | 2 retries |
Words from the source blog are excluded
A rewrite preserves the content, so words already in the source (e.g. a jockey's name) were not introduced by the AI and cannot be removed by regenerating. Words present in the original body are therefore excluded from screening.
Broken JSON is never published as HTML
If the model attempted block JSON and produced something broken, falling back to HTML would publish the raw JSON text as the article body. Instead, nothing is published: the error is attached as a correction instruction and the text is rewritten in-process โ faster than waiting for SQS redelivery, and more likely to succeed because a correction can be supplied.
If the limit is reached without resolution, generation fails and nothing is published.
Related information is out of scope
This Lambda does not generate the related-information section.
| Concern | Owner |
|---|---|
| Attaching embeddings | odds_poc_app (on POST /v1/old_blogs, attached to old_blog) |
| Carrying vectors over | odds_poc_app (when creating the blog for a rewrite) |
| Generating related information | odds_poc_app (POST /v1/blogs/{blog_id}/related_generate) |
Models
| Use | Model |
|---|---|
| Race video / course image analysis (stage 1) | Amazon Nova Pro (supports video input) |
| Rewriting and summary (stages 2 and 3) | Claude Sonnet (Nova Pro when BEDROCK_CLAUDE_SONNET_PROFILE_ARN is unset) |
The output token limit is 16384 on Claude and 8192 on Nova, because block JSON for a long column that gets truncated mid-output cannot be parsed.
Error handling
| Situation | Behaviour |
|---|---|
| Failure on a non-final attempt | Re-raise and let SQS redelivery recover |
| Failure on the final attempt | Hard-delete the status=generating blog and finish without re-raising |