Skip to content

Infrastructure

Environments

Environment URL Purpose
dev Frontend: https://interview-arrange-web.heineken.dev.4digit.ai / API: https://interview-arrange-api.heineken.dev.4digit.ai Verification environment, deployed automatically on push to dev. All external side effects are disabled
stg Frontend: https://interview-arrange-web.heineken.stg.4digit.ai / API: https://interview-arrange-api.heineken.stg.4digit.ai Deployed automatically on push to stg. Email and Google Calendar writes actually happen here
Local Frontend: http://localhost:3000 / API: http://localhost:8080 Local development

Note

No production environment exists yet. Only the prod branch has been prepared.

dev and stg are fully independent down to the VPC and the database. The only things they share are the Databricks service principal and the Google service account / OAuth client.

Feature flags per environment

Variable dev stg Effect
PROTOTYPE_MODE true (default) false Setting it to false enables the email and calendar side effects
EMAIL_ENABLED unset true Mail is only sent when PROTOTYPE_MODE=false and this is true
GMAIL_SENDER_EMAIL empty sender address The service account impersonates this address to send
ALLOWED_SURVEY_IDS unset (no limit) the survey to work with Limits the survey list, detail and metadata sync to these IDs

PROTOTYPE_MODE gates email and calendar together. Enabling email on stg necessarily enables calendar writes as well.

Terraform (genai-infrastructure)

AWS resources live in a separate Terraform monorepo, genai-infrastructure. One directory is one tfstate.

Stack tfstate
aws/envs/dev/heineken-interview-arrange-backend 4d-dev-genai-terraform / dev.heineken-interview-arrange-backend.tfstate
aws/envs/dev/heineken-interview-arrange-web same bucket (...-web.tfstate)
aws/envs/stg/heineken-interview-arrange-backend 4d-stg-genai-terraform / staging.heineken-interview-arrange-backend.tfstate
aws/envs/stg/heineken-interview-arrange-web same bucket (staging....-web.tfstate)
  • VPC CIDR is 10.20.0.0/16 for dev and 10.21.0.0/16 for stg
  • Secrets are SOPS-encrypted (KMS) in each stack's secrets/values.yml. dev and stg use different KMS keys, so copying the dev file is not enough โ€” it cannot be decrypted
  • Non-secret settings such as ALLOWED_SURVEY_IDS go in locals.tf as plain text, so changing them needs no decrypt/re-encrypt cycle
  • Lambda images are updated by the app repositories' GitHub Actions, so Terraform carries lifecycle { ignore_changes = [image_uri] }

The first apply takes two passes

Neither the Lambda functions nor the App Runner service can be created while ECR is empty. Run terraform apply -target=module.ecr first, push the images, then apply everything. The App Runner custom domain needs one more pass: apply -target=aws_apprunner_custom_domain_association.web first so the certificate validation records exist, because the count that depends on them cannot be resolved before apply.

Architecture

graph TD
  User[Console user]
  Candidate[Candidate]
  AR[App Runner<br/>&lt;env&gt;-heineken-interview-arrange-web]
  LambdaApp[Lambda<br/>&lt;env&gt;-heineken-interview-arrange-backend]
  LambdaMig[Lambda<br/>...-migration]
  ECS[ECS Fargate task<br/>...-sync]
  RDS[(PostgreSQL)]
  DBX[Databricks<br/>Unity Catalog]
  Google[Google APIs<br/>OAuth / Calendar / Gmail]
  ECR[ECR<br/>3 repositories]

  User --> AR
  Candidate --> AR
  AR -->|REST| LambdaApp
  LambdaApp --> RDS
  LambdaApp --> Google
  LambdaApp --> DBX
  LambdaApp -->|RunTask| ECS
  LambdaMig --> RDS
  ECS --> DBX
  ECS --> RDS
  ECR -.image.-> AR
  ECR -.image.-> LambdaApp
  ECR -.image.-> LambdaMig
  ECR -.image.-> ECS

Services

Service Purpose
AWS Lambda (container image) The API itself. Ships aws-lambda-adapter as an extension so Bun's HTTP server runs unchanged
AWS Lambda (migration) Dedicated function that runs scripts/migrate.ts
Amazon ECS (Fargate) Candidate import and survey metadata refresh. The script path and JSON argument come from the container override on Dockerfile.sync
AWS App Runner Hosting for the Next.js frontend
Amazon ECR Stores four images (backend app / migration / sync / frontend web)
PostgreSQL Persistence
Databricks Answer data for Creative Survey / Ask One. Read-only (Databricks integration)
Google Workspace OAuth, Calendar, Gmail

Container Images

Repository names carry a per-environment prefix (dev- / stg-).

Dockerfile ECR repository Base Command
Dockerfile <env>-heineken-interview-arrange-backend-lambda oven/bun:1-alpine + aws-lambda-adapter bun run src/server.ts
Dockerfile.migration <env>-heineken-interview-arrange-backend-migration oven/bun:1-alpine bun run scripts/migrate.ts
Dockerfile.sync <env>-heineken-interview-arrange-backend-sync oven/bun:1-alpine ENTRYPOINT ["bun", "run"]; the script path comes from the override
frontend Dockerfile <env>-heineken-interview-arrange-web Next.js Started by App Runner

Branching

Changes flow one way: dev โ†’ stg โ†’ prod.

graph LR
  F[feature branch] -->|PR| D[dev]
  D -->|PR| S[stg]
  S -->|PR| P[prod]
  • All three branches are protected: PR required, one approving review, no force pushes, no deletion. Administrators are exempt
  • dev is the default branch; cut feature branches from it
  • delete_branch_on_merge is on, so merging a PR removes the feature branch automatically

Deployment Flow

Branch Target Trigger Workflow
dev Two dev Lambdas + ECS task definition / App Runner Automatic on push deploy-dev.yml
stg The same resources on stg Automatic on push deploy-stg.yml
prod Not built yet - -

The two workflows differ only in the resource names (dev- / stg-) and the concurrency group; the steps are identical.

Backend (.github/workflows/deploy-dev.yml / deploy-stg.yml)

graph TD
  A[Push to dev or stg] --> B[AWS auth / ECR login]
  B --> C[Build & push app image]
  B --> D[Build & push migration image]
  B --> E[Build & push sync image]
  C --> F[update-function-code for app Lambda]
  D --> G[update-function-code for migration Lambda]
  E --> H[Register new ECS task definition revision]
  F --> I[Invoke migration Lambda]
  G --> I
  H --> I
  I --> J{FunctionError?}
  J -->|yes| K[Job fails]
  J -->|no| L[Done]
  1. Push three images (app / migration / sync) to ECR under both github.sha and latest
  2. Update the app and migration Lambdas via update-function-code, then wait function-updated
  3. For sync, describe-task-definition โ†’ swap the image โ†’ register-task-definition to create a new revision
  4. invoke the migration Lambda to run migrations; fail the job if the output contains FunctionError

Frontend (.github/workflows/deploy-dev.yml / deploy-stg.yml)

  1. Build the image with NEXT_PUBLIC_API_URL baked in as a build argument
  2. Push to ECR (github.sha / latest)
  3. Look up the App Runner service by name and call start-deployment

NEXT_PUBLIC_API_URL is fixed at build time

The frontend's API URL is embedded via a build argument and cannot be changed with a runtime environment variable. Rebuild the image to point it elsewhere.

Required GitHub Secrets

Secret Purpose
AWS_ACCESS_KEY_ID Operating ECR / Lambda / ECS / App Runner
AWS_SECRET_ACCESS_KEY Same

Both workflows hard-code the region as ap-northeast-1.

Running Batches

src/batch/*.ts runs independently of the API. It opens its own pg pool from DATABASE_URL alone rather than going through getEnvConfig(), so it works in environments where the app's required variables are absent. Arguments are passed as a single JSON string.

bun run src/batch/import.ts '{"action":"import","importLogId":1,"projectId":1,"surveyId":"123","mapping":{}}'
bun run src/batch/databricks-meta-sync.ts '{"dryRun":true}'
bun run src/batch/pii-cleanup.ts '{"dryRun":true}'

services/s3-sync.ts is the kicker for batch/import.ts: it spawns it locally in development and launches an ECS Fargate task in production. (The filename still says s3-sync for historical reasons; it has nothing to do with S3.)

databricks-meta-sync.ts is never launched by the app. The intent is to run it as an ECS task after the Databricks pipeline, but that hand-off is not wired up yet.

Deploys are serialised

Two concurrent deploys make UpdateFunctionCode fail with ResourceConflictException. The workflows declare concurrency: deploy-dev / deploy-stg with cancel-in-progress: false. Cancelling could leave the app updated with the migration unrun, so later runs queue instead. The groups differ per environment, so dev and stg can deploy in parallel.

Migration success is decided by MIGRATION_COMPLETE

The migration image runs scripts/migrate.ts directly rather than a Lambda handler, so Lambda reports an error even on success once the container exits. CI instead looks for MIGRATION_COMPLETE in the CloudWatch logs.

Never run git push from automation

By operational rule, agents and automation do not run git push. A push to dev or stg deploys, so it must be done deliberately by a person.

Google Workspace Setup

The service account never uses its own calendar or mailbox. It impersonates a user in the domain, which means domain-wide delegation has to be granted in the Workspace admin console.

Operation Impersonated account
Sending email GMAIL_SENDER_EMAIL
Creating the project calendar, reading and writing its events The project creator
Free/busy lookups Each interviewer

In the admin console (Security โ†’ Access and data control โ†’ API controls โ†’ Domain-wide delegation), authorise the service account's client ID for these scopes:

https://www.googleapis.com/auth/gmail.send
https://www.googleapis.com/auth/gmail.compose
https://www.googleapis.com/auth/calendar
https://www.googleapis.com/auth/calendar.events
https://www.googleapis.com/auth/calendar.freebusy

On the GCP side, enable the Gmail API and the Google Calendar API, and register each environment's /api/v1/auth/google/callback as an authorised redirect URI on the OAuth client.

The logs tell you which part of the setup is missing

  • unauthorized_client: Client is unauthorized to retrieve access tokens... โ†’ a scope is missing from the delegation. A single missing scope only breaks that one API, which is why calendar can work while email fails
  • Delegation denied for <address> โ†’ the address does not exist
  • ... API has not been used in project ... before or it is disabled โ†’ the API is not enabled in GCP