Operations

July 24, 2026 ยท View on GitHub

This section is for the platform team after the first Forge tenant is running. It covers tenant support, image management, artifacts, cleanup, secrets, upgrades, and troubleshooting.

Day-2 Loop

CadenceActionDoc
Every tenant requestCollect tenant values, add config, run plan, run smoke workflow.Tenant Onboarding
Every image releaseBuild base/custom images, share AMIs, update tenant runner specs.Runner Images
WeeklyRun example apply/destroy for helpers, infra, platform, integrations.Workflow Blueprints
WeeklyRun cleanup and policy jobs.Cloud Custodian
MonthlyReview module refs, Renovate output, AMI age, stale ECR tags, and secrets.Upgrades
Planned ARC upgradeRebuild blue/green EKS clusters and move tenants one at a time.Move ARC Tenants
IncidentTriage queued jobs, failed runner registration, IAM, webhooks, or ARC.Troubleshooting
IncidentDebug Terraform, OpenTofu, or Terragrunt plan/apply that appears stuck.Terraform/Terragrunt Stuck Runbook
IncidentUse Splunk dashboards to identify the failing subsystem and severity.Splunk Dashboard Runbook
IncidentTriage Forge metrics, dependencies, resource pressure, and detectors.Splunk Observability Dashboard Runbook

Operating Repos

If you need an end-to-end operating model, copy from Operations Repo Blueprints. The blueprints include real folders for Packer, Ansible, containers, Renovate, Cloud Custodian, Terragrunt, reusable actions, and weekly example deployments.

For Splunk-based operations, start with the Splunk Dashboard Runbook. Use the panel reference when you need to map a dashboard panel back to its operational question.

For Splunk Observability metrics and detectors, use the Splunk Observability Dashboard Runbook and Splunk Observability Dashboard Panel Reference.

For installations that do not deploy Splunk, use Troubleshooting Without Splunk as the baseline support runbook.