Infrastructure
Environments
| Environment | URL | Purpose |
|---|---|---|
| dev | Frontend: https://interview-arrange-web.heineken.dev.4digit.ai / API: https://interview-arrange-api.heineken.dev.4digit.ai |
Verification environment, deployed automatically on push to dev. All external side effects are disabled |
| stg | Frontend: https://interview-arrange-web.heineken.stg.4digit.ai / API: https://interview-arrange-api.heineken.stg.4digit.ai |
Deployed automatically on push to stg. Email and Google Calendar writes actually happen here |
| Local | Frontend: http://localhost:3000 / API: http://localhost:8080 |
Local development |
Note
No production environment exists yet. Only the prod branch has been prepared.
dev and stg are fully independent down to the VPC and the database. The only things they share are the Databricks service principal and the Google service account / OAuth client.
Feature flags per environment
| Variable | dev | stg | Effect |
|---|---|---|---|
PROTOTYPE_MODE |
true (default) |
false |
Setting it to false enables the email and calendar side effects |
EMAIL_ENABLED |
unset | true |
Mail is only sent when PROTOTYPE_MODE=false and this is true |
GMAIL_SENDER_EMAIL |
empty | sender address | The service account impersonates this address to send |
ALLOWED_SURVEY_IDS |
unset (no limit) | the survey to work with | Limits the survey list, detail and metadata sync to these IDs |
PROTOTYPE_MODE gates email and calendar together. Enabling email on stg necessarily enables calendar writes as well.
Terraform (genai-infrastructure)
AWS resources live in a separate Terraform monorepo, genai-infrastructure. One directory is one tfstate.
| Stack | tfstate |
|---|---|
aws/envs/dev/heineken-interview-arrange-backend |
4d-dev-genai-terraform / dev.heineken-interview-arrange-backend.tfstate |
aws/envs/dev/heineken-interview-arrange-web |
same bucket (...-web.tfstate) |
aws/envs/stg/heineken-interview-arrange-backend |
4d-stg-genai-terraform / staging.heineken-interview-arrange-backend.tfstate |
aws/envs/stg/heineken-interview-arrange-web |
same bucket (staging....-web.tfstate) |
- VPC CIDR is
10.20.0.0/16for dev and10.21.0.0/16for stg - Secrets are SOPS-encrypted (KMS) in each stack's
secrets/values.yml. dev and stg use different KMS keys, so copying the dev file is not enough โ it cannot be decrypted - Non-secret settings such as
ALLOWED_SURVEY_IDSgo inlocals.tfas plain text, so changing them needs no decrypt/re-encrypt cycle - Lambda images are updated by the app repositories' GitHub Actions, so Terraform carries
lifecycle { ignore_changes = [image_uri] }
The first apply takes two passes
Neither the Lambda functions nor the App Runner service can be created while ECR is empty. Run terraform apply -target=module.ecr first, push the images, then apply everything. The App Runner custom domain needs one more pass: apply -target=aws_apprunner_custom_domain_association.web first so the certificate validation records exist, because the count that depends on them cannot be resolved before apply.
Architecture
graph TD
User[Console user]
Candidate[Candidate]
AR[App Runner<br/><env>-heineken-interview-arrange-web]
LambdaApp[Lambda<br/><env>-heineken-interview-arrange-backend]
LambdaMig[Lambda<br/>...-migration]
ECS[ECS Fargate task<br/>...-sync]
RDS[(PostgreSQL)]
DBX[Databricks<br/>Unity Catalog]
Google[Google APIs<br/>OAuth / Calendar / Gmail]
ECR[ECR<br/>3 repositories]
User --> AR
Candidate --> AR
AR -->|REST| LambdaApp
LambdaApp --> RDS
LambdaApp --> Google
LambdaApp --> DBX
LambdaApp -->|RunTask| ECS
LambdaMig --> RDS
ECS --> DBX
ECS --> RDS
ECR -.image.-> AR
ECR -.image.-> LambdaApp
ECR -.image.-> LambdaMig
ECR -.image.-> ECS
Services
| Service | Purpose |
|---|---|
| AWS Lambda (container image) | The API itself. Ships aws-lambda-adapter as an extension so Bun's HTTP server runs unchanged |
| AWS Lambda (migration) | Dedicated function that runs scripts/migrate.ts |
| Amazon ECS (Fargate) | Candidate import and survey metadata refresh. The script path and JSON argument come from the container override on Dockerfile.sync |
| AWS App Runner | Hosting for the Next.js frontend |
| Amazon ECR | Stores four images (backend app / migration / sync / frontend web) |
| PostgreSQL | Persistence |
| Databricks | Answer data for Creative Survey / Ask One. Read-only (Databricks integration) |
| Google Workspace | OAuth, Calendar, Gmail |
Container Images
Repository names carry a per-environment prefix (dev- / stg-).
| Dockerfile | ECR repository | Base | Command |
|---|---|---|---|
Dockerfile |
<env>-heineken-interview-arrange-backend-lambda |
oven/bun:1-alpine + aws-lambda-adapter |
bun run src/server.ts |
Dockerfile.migration |
<env>-heineken-interview-arrange-backend-migration |
oven/bun:1-alpine |
bun run scripts/migrate.ts |
Dockerfile.sync |
<env>-heineken-interview-arrange-backend-sync |
oven/bun:1-alpine |
ENTRYPOINT ["bun", "run"]; the script path comes from the override |
frontend Dockerfile |
<env>-heineken-interview-arrange-web |
Next.js | Started by App Runner |
Branching
Changes flow one way: dev โ stg โ prod.
graph LR
F[feature branch] -->|PR| D[dev]
D -->|PR| S[stg]
S -->|PR| P[prod]
- All three branches are protected: PR required, one approving review, no force pushes, no deletion. Administrators are exempt
devis the default branch; cut feature branches from itdelete_branch_on_mergeis on, so merging a PR removes the feature branch automatically
Deployment Flow
| Branch | Target | Trigger | Workflow |
|---|---|---|---|
dev |
Two dev Lambdas + ECS task definition / App Runner | Automatic on push | deploy-dev.yml |
stg |
The same resources on stg | Automatic on push | deploy-stg.yml |
prod |
Not built yet | - | - |
The two workflows differ only in the resource names (dev- / stg-) and the concurrency group; the steps are identical.
Backend (.github/workflows/deploy-dev.yml / deploy-stg.yml)
graph TD
A[Push to dev or stg] --> B[AWS auth / ECR login]
B --> C[Build & push app image]
B --> D[Build & push migration image]
B --> E[Build & push sync image]
C --> F[update-function-code for app Lambda]
D --> G[update-function-code for migration Lambda]
E --> H[Register new ECS task definition revision]
F --> I[Invoke migration Lambda]
G --> I
H --> I
I --> J{FunctionError?}
J -->|yes| K[Job fails]
J -->|no| L[Done]
- Push three images (app / migration / sync) to ECR under both
github.shaandlatest - Update the app and migration Lambdas via
update-function-code, thenwait function-updated - For sync,
describe-task-definitionโ swap the image โregister-task-definitionto create a new revision invokethe migration Lambda to run migrations; fail the job if the output containsFunctionError
Frontend (.github/workflows/deploy-dev.yml / deploy-stg.yml)
- Build the image with
NEXT_PUBLIC_API_URLbaked in as a build argument - Push to ECR (
github.sha/latest) - Look up the App Runner service by name and call
start-deployment
NEXT_PUBLIC_API_URL is fixed at build time
The frontend's API URL is embedded via a build argument and cannot be changed with a runtime environment variable. Rebuild the image to point it elsewhere.
Required GitHub Secrets
| Secret | Purpose |
|---|---|
AWS_ACCESS_KEY_ID |
Operating ECR / Lambda / ECS / App Runner |
AWS_SECRET_ACCESS_KEY |
Same |
Both workflows hard-code the region as ap-northeast-1.
Running Batches
src/batch/*.ts runs independently of the API. It opens its own pg pool from DATABASE_URL alone rather than going through getEnvConfig(), so it works in environments where the app's required variables are absent. Arguments are passed as a single JSON string.
bun run src/batch/import.ts '{"action":"import","importLogId":1,"projectId":1,"surveyId":"123","mapping":{}}'
bun run src/batch/databricks-meta-sync.ts '{"dryRun":true}'
bun run src/batch/pii-cleanup.ts '{"dryRun":true}'
services/s3-sync.ts is the kicker for batch/import.ts: it spawns it locally in development and launches an ECS Fargate task in production. (The filename still says s3-sync for historical reasons; it has nothing to do with S3.)
databricks-meta-sync.ts is never launched by the app. The intent is to run it as an ECS task after the Databricks pipeline, but that hand-off is not wired up yet.
Deploys are serialised
Two concurrent deploys make UpdateFunctionCode fail with ResourceConflictException. The workflows declare concurrency: deploy-dev / deploy-stg with cancel-in-progress: false. Cancelling could leave the app updated with the migration unrun, so later runs queue instead. The groups differ per environment, so dev and stg can deploy in parallel.
Migration success is decided by MIGRATION_COMPLETE
The migration image runs scripts/migrate.ts directly rather than a Lambda handler, so Lambda reports an error even on success once the container exits. CI instead looks for MIGRATION_COMPLETE in the CloudWatch logs.
Never run git push from automation
By operational rule, agents and automation do not run git push. A push to dev or stg deploys, so it must be done deliberately by a person.
Google Workspace Setup
The service account never uses its own calendar or mailbox. It impersonates a user in the domain, which means domain-wide delegation has to be granted in the Workspace admin console.
| Operation | Impersonated account |
|---|---|
| Sending email | GMAIL_SENDER_EMAIL |
| Creating the project calendar, reading and writing its events | The project creator |
| Free/busy lookups | Each interviewer |
In the admin console (Security โ Access and data control โ API controls โ Domain-wide delegation), authorise the service account's client ID for these scopes:
https://www.googleapis.com/auth/gmail.send
https://www.googleapis.com/auth/gmail.compose
https://www.googleapis.com/auth/calendar
https://www.googleapis.com/auth/calendar.events
https://www.googleapis.com/auth/calendar.freebusy
On the GCP side, enable the Gmail API and the Google Calendar API, and register each environment's /api/v1/auth/google/callback as an authorised redirect URI on the OAuth client.
The logs tell you which part of the setup is missing
unauthorized_client: Client is unauthorized to retrieve access tokens...โ a scope is missing from the delegation. A single missing scope only breaks that one API, which is why calendar can work while email failsDelegation denied for <address>โ the address does not exist... API has not been used in project ... before or it is disabledโ the API is not enabled in GCP