runbook.md
December 28, 2025 · View on GitHub
# PATH: docs/ops/runbook.md
<!-- // --- REPLACE START: Ops Runbook cleanup (canonical index + incident actions; consistent language/links) --- -->
# Ops Runbook (Loventia)
> **Purpose:** A single “index + quick actions” page for incidents and production operations.
> **Scope:** CI/CD, rollback, Stripe cutover, security hardening, test plan, admin endpoints.
> **Rule:** During an active incident, prefer **rollback** over “hot fixes”.
---
## 1) Links (canonical docs)
### CI/CD
- **CI/CD guide:** [`../ci-cd.md`](../ci-cd.md)
### Rollback (canonical)
- **Rollback playbook (canonical):** [`./rollback-playbook.md`](./rollback-playbook.md)
> Note: `docs/rollback-playbook.md` should be a **pointer** file that redirects here so we avoid maintaining two conflicting playbooks.
### Stripe
- **Stripe live cutover:** [`../stripe-live-cutover.md`](../stripe-live-cutover.md)
### Security
- **Security hardening:** [`../security-hardening.md`](../security-hardening.md)
### Testing
- **Test plan:** [`../test-plan.md`](../test-plan.md)
### Admin
- **Admin endpoints:** [`../admin-endpoints.md`](../admin-endpoints.md)
---
## 2) Quick commands (when prod is broken)
> Use these as a minimal first response: confirm → scope → rollback if needed.
### 2.1 Identify environment & AWS identity
```bash
aws sts get-caller-identity
aws configure list
2.2 Frontend “is the site up” (CloudFront / static)
# Replace with your real domain when you have one.
# For CloudFront default domain, use: https://<dist>.cloudfront.net/
curl -I https://<frontend-domain>/
curl -I https://<frontend-domain>/discover
If the page loads but looks broken:
- Open DevTools → Console + Network
- Hard refresh (Ctrl+F5)
- Confirm the newest JS/CSS files load (no 403/404)
2.3 Backend “is the API up” (health/ready)
curl -fsS https://<api-domain>/health
curl -fsS https://<api-domain>/ready
If /health is OK but /ready fails:
- Suspect DB connectivity, migrations, or a downstream dependency.
2.4 ECS quick checks (if backend runs on ECS)
AWS_REGION="eu-north-1"
CLUSTER_NAME="<your-ecs-cluster>"
SERVICE_NAME="<your-ecs-service>"
# Service status + current task definition
aws ecs describe-services \
--region "$AWS_REGION" \
--cluster "$CLUSTER_NAME" \
--services "$SERVICE_NAME" \
--query 'services[0].{Status:status,Running:runningCount,Desired:desiredCount,Pending:pendingCount,TaskDef:taskDefinition,Deployments:deployments}' \
--output json
# Recent running tasks
aws ecs list-tasks \
--region "$AWS_REGION" \
--cluster "$CLUSTER_NAME" \
--service-name "$SERVICE_NAME" \
--desired-status RUNNING \
--max-items 10
# If you know a task ARN, describe it
TASK_ARN="<task-arn>"
aws ecs describe-tasks \
--region "$AWS_REGION" \
--cluster "$CLUSTER_NAME" \
--tasks "$TASK_ARN" \
--output json
2.5 Rollback fast (link)
If a deploy likely caused the incident, jump here:
- Canonical rollback playbook:
./rollback-playbook.md
2.6 CloudFront invalidate (frontend stale/broken assets)
DISTRIBUTION_ID="<your-cloudfront-distribution-id>"
aws cloudfront create-invalidation --distribution-id "$DISTRIBUTION_ID" --paths "/*"
3) Where are logs?
Frontend (FE)
-
Browser DevTools
- Console errors (JS runtime)
- Network tab (failed asset loads, 403/404/5xx)
-
CloudFront
- Distribution metrics (Requests, 4xx/5xx)
- (Optional) Access logs to S3 (if enabled)
-
Sentry / client error tracking
- If enabled, use Sentry to correlate spikes with release version.
Backend (BE)
Depends on deployment target:
-
ECS / CloudWatch Logs
-
CloudWatch log groups for the service/task
-
Look for:
- auth failures (401/403 spikes)
- webhook errors
- DB connection errors
- unhandled exceptions
-
-
Load balancer (if used)
- ALB target health
- ALB 5xx/4xx spikes
-
Local dev
- Server console output
- Any
logs/folder (if configured)
AWS services (infra)
- Route 53: DNS record health / propagation
- ACM: cert validation / status (CloudFront requires certs in us-east-1)
- S3: object existence + permissions + CORS
- CloudFront: caching, invalidations, origin errors (403/404/5xx)
4) Common incidents (what to check first)
4.1 Stripe webhook failures
Symptoms
- Users pay but premium does not activate
POST /api/billing/syncfails or premium state “lags”- Webhook endpoint returns 400/500
Checklist
-
Stripe Dashboard → Developers → Logs (find the event)
-
Confirm webhook secret and endpoint URL are correct:
STRIPE_WEBHOOK_SECRETon server
-
Check server logs for:
- signature verification errors
- parsing errors (raw body / middleware order)
-
Confirm webhook returns 2xx quickly
-
If needed: run a manual sync flow (authenticated/admin endpoint)
Helpful commands
curl -fsS https://<api-domain>/health
# If you have a billing sync endpoint:
# curl -fsS -H "Authorization: Bearer <token>" https://<api-domain>/api/billing/sync
4.2 Login failures / auth refresh loop
Symptoms
- Users cannot log in
- Login returns 401/500
- Frontend stuck refreshing tokens
Checklist
-
Confirm backend is healthy (
/health,/ready) -
Check env:
- JWT secrets (access + refresh)
- cookie settings (SameSite/Secure/Path)
- CORS allowed origins + credentials
-
Confirm time is sane on server (clock drift can break JWT)
-
Inspect server logs around:
- login
- refresh
- cookie parsing
4.3 5xx spike / timeouts
Symptoms
- ALB/CloudFront shows a 5xx spike
- API requests time out
- DB connection errors
Checklist
-
Check whether the spike correlates with a deploy time
-
Check ECS:
- running vs desired tasks
- last deployment events
-
Check DB availability / pool saturation
-
If deploy-related: rollback
4.4 S3 images broken (403/404 / slow)
Symptoms
- Profile images do not load
<img>shows broken or 403/404- Some users affected, others not
Checklist
-
Confirm the URL is correct and the object exists:
- bucket, key, region
-
Check S3 permissions / bucket policy:
- public read vs signed URLs
-
Check CORS rules (browser blocks)
-
If behind CloudFront:
- invalidate cache
- confirm origin access settings
-
Verify from a client machine:
curl -I <image-url>should be 200
5) Incident notes template (optional but useful)
When the incident is resolved, capture:
- Start time / end time
- Impact
- Root cause (short)
- Fix (rollback / config change / code patch)
- Prevention action (one concrete item)
::contentReference[oaicite:0]{index=0}