Skip to content

old_blog Table

Overview

Item Details
Table name old_blog
Description Holds past blog articles targeted for migration. Referenced by blog.old_blog_id and used as the input for AI rewriting (blog_rewriter)
Primary use Storing past blog articles, feeding rewrites, and keeping vectors for related-blog search

Column Definitions

Column Type PK FK Unique Nullable Default Notes
_id ObjectId โ—ฏ โ—ฏ Auto-generated by DocumentDB
name string Blog title
race_id string auto_racing:_id / bicycle_racing:_id / horse_racing:_id โ—ฏ Linked race; not every past blog is tied to one. Indexed
race_type enum โ—ฏ auto_racing / bicycle_racing / horse_racing. Set only when race_id is present
blog_type string โ—ฏ Blog series type (e.g. jockey interview, Saito column)
category string โ—ฏ Broad blog category (e.g. interview, race recap, prediction). Used by blog_builder to pick the article template
source_url string โ—ฏ URL of the source blog article. Also the base for absolutizing relative image src values in the body
published_at number โ—ฏ Publish date/time of the source article (UNIX ms)
body_html string HTML. Keeps the source body as-is
embedding number[] โ—ฏ Embedding of category + title + body. Attached on registration and carried over to blog on rewrite
created_at number
updated_at number Lists are sorted by update date/time by default
deleted_at number โ—ฏ

category is a string

old_blog stores category as a string, not a category_id. On rewrite, a category record with a matching name is looked up and linked as blog.category_id.

Embedding

On registration (POST /v1/old_blogs), "category + title + body (tags stripped)" is flattened to text and embedded with Amazon Bedrock Titan Text Embeddings v2. Updating the category (PUT /v1/old_blogs/{id}) regenerates the vector.

Embedding is supplementary: if it fails, registration/update still succeeds and embedding stays null.