GitLab Runner VM
September 24, 2026 · View on GitHub
Provisions GitLab Runner VMs on OpenShift Virtualization or GCE (Google Compute Engine) with a Podman custom executor and OpenShell gateway for fullsend agent jobs.
Architecture
Each runner VM runs:
- gitlab-runner (custom executor) — receives CI jobs from GitLab.
The runner is a systemd system service running as the VM user, so it
does not go through
pam_systemdand does not inherit a login session.setup.shenables lingering for that user and writesXDG_RUNTIME_DIR/DBUS_SESSION_BUS_ADDRESSinto the gitlab-runner drop-in with the runner user's numeric UID (resolved at setup time — systemd%Uon a system-scope unit expands to 0, notUser=).systemctl --user(OpenShell gateway start/stop) can then reach the user bus.executor/gateway.shalso pins those variables itself, so a job still works if the unit environment is missing (#7453, #7696). - Podman (rootless) — creates per-job containers. A user systemd
timer (
fullsend-podman-prune.timer) plus a prepare/cleanup hook reclaim unused images and stopped leftovers so the ~30 GiB root disk cannot fill with superseded job layers (#7663). In-flightrunner-*/openshell-*containers are never stopped; the warm-cache runner and supervisor images listed in~/.config/fullsend-gitlab-runner/keep-imagesare never removed. - OpenShell gateway — started per job in
prepare.sh, torn down incleanup.sh. The VM does not keep a long-livedsystemd --usergateway: that accumulated a stale profile registry, a baked-in OpenShell version, and leaked sandboxes (#7218). Each job recreates the gateway with an empty registry, matching GitHub Actions (fresh, version-matched install, throw the runner away). The ~6 GB image cache stays warm; only the CLI (~39 MB) and a version-skewed supervisor (~30 MB) are fetched when the job's pin differs from the host.
Job containers use --network=host to reach the gateway. An OCI
createRuntime hook injects the host CA trust bundle into every container
(Debian and RHEL-family layouts) so the OpenShell supervisor and sandboxed
processes can verify internal TLS endpoints. That hook is the sandbox-host
half of private-CA support; GitLab job containers separately consume
CI_SERVER_TLS_CA_FILE (see Private CA (self-hosted GitLab)).
The Kubernetes executor does not run this hook — do not treat a job-local
CI_SERVER_TLS_CA_FILE path as available on a remote sandbox host.
Gateway mTLS credentials are mounted read-only from the runner user's
OpenShell config. prepare.sh reaps leftover openshell-* / openshell.managed
containers from an abruptly-killed prior job before starting the new gateway.
This is a deployment variant of the container isolation model described in ADR-0036. It uses a Podman custom executor instead of Docker/Kubernetes executors but maintains equivalent container-level isolation for agent jobs.
See also:
Quick start — OpenShift Virtualization
Requires:
oc,virtctl(client >= 1.5 — ships with OpenShift Virtualization >= 4.19),python3, and a localsshbinary (virtctl wraps it via ProxyCommand).
# 1. Create and provision a VM — group-scoped runner (recommended):
GL_TOKEN=glpat-xxx GROUP_ID=12345 \
GITLAB_URL=https://gitlab.example.com \
NAMESPACE=my-namespace \
RUNNER_IMAGE=ghcr.io/org/runner:v1.2.3 \
./create-openshift-vm.sh
# Or project-scoped runner (for single-project deployments):
GL_TOKEN=glpat-xxx PROJECT_ID=12345 \
GITLAB_URL=https://gitlab.example.com \
NAMESPACE=my-namespace \
RUNNER_IMAGE=ghcr.io/org/runner:v1.2.3 \
./create-openshift-vm.sh
# Or join an existing runner pool (runner-hub — multiple VMs, one registration):
RUNNER_TOKEN=glrt-xxx \
GITLAB_URL=https://gitlab.example.com \
NAMESPACE=my-namespace \
RUNNER_IMAGE=ghcr.io/org/runner:v1.2.3 \
./create-openshift-vm.sh 05
# 2. Delete a VM (individual runner — drains, deregisters, deletes):
GL_TOKEN=glpat-xxx \
GITLAB_URL=https://gitlab.example.com \
NAMESPACE=my-namespace \
./delete-openshift-vm.sh fullsend-gitlab-runner-01
# Or delete a fleet VM (no GL_TOKEN — drains and deletes, does not deregister):
RUNNER_TOKEN=glrt-xxx NAMESPACE=my-namespace \
./delete-openshift-vm.sh fullsend-gitlab-runner-05
# Or skip the drain step and delete immediately:
GL_TOKEN=glpat-xxx \
GITLAB_URL=https://gitlab.example.com \
NAMESPACE=my-namespace \
./delete-openshift-vm.sh --no-drain fullsend-gitlab-runner-01
# 3. List VMs:
NAMESPACE=my-namespace ./delete-openshift-vm.sh --list
Quick start — GCE (Google Compute Engine)
Requires:
gcloudCLI (authenticated),python3,curl. By default, VMs have no public IP — SSH access is tunneled via IAP (Cloud IAP API must be enabled; operator needsroles/iap.tunnelResourceAccessor). The VPC subnet must have Cloud NAT configured — without it, VMs created with--no-addresscannot reach package mirrors or container registries anddnf installwill fail. SetGCP_USE_IAP=falseto create the VM with an external IP and SSH directly.
# 1. Create and provision a VM — group-scoped runner (recommended):
GL_TOKEN=glpat-xxx GROUP_ID=12345 \
GITLAB_URL=https://gitlab.example.com \
GCP_PROJECT=my-gcp-project \
RUNNER_IMAGE=ghcr.io/org/runner:v1.2.3 \
./create-gcp-vm.sh
# Or project-scoped runner (for single-project deployments):
GL_TOKEN=glpat-xxx PROJECT_ID=12345 \
GITLAB_URL=https://gitlab.example.com \
GCP_PROJECT=my-gcp-project \
RUNNER_IMAGE=ghcr.io/org/runner:v1.2.3 \
./create-gcp-vm.sh
# Or join an existing runner pool (runner-hub — multiple VMs, one registration):
RUNNER_TOKEN=glrt-xxx \
GITLAB_URL=https://gitlab.example.com \
GCP_PROJECT=my-gcp-project \
RUNNER_IMAGE=ghcr.io/org/runner:v1.2.3 \
./create-gcp-vm.sh 05
# 2. Delete a VM (individual runner — drains, deregisters, deletes):
GL_TOKEN=glpat-xxx \
GITLAB_URL=https://gitlab.example.com \
GCP_PROJECT=my-gcp-project \
./delete-gcp-vm.sh fullsend-gitlab-runner-01
# Or delete a fleet VM (no GL_TOKEN — drains and deletes, does not deregister):
RUNNER_TOKEN=glrt-xxx GCP_PROJECT=my-gcp-project \
./delete-gcp-vm.sh fullsend-gitlab-runner-05
# Or skip the drain step and delete immediately:
GL_TOKEN=glpat-xxx \
GITLAB_URL=https://gitlab.example.com \
GCP_PROJECT=my-gcp-project \
./delete-gcp-vm.sh --no-drain fullsend-gitlab-runner-01
# 3. List VMs:
GCP_PROJECT=my-gcp-project ./delete-gcp-vm.sh --list
Environment variables
Shared
| Variable | Required | Default | Description |
|---|---|---|---|
RUNNER_TOKEN | yes (create)² | — | GitLab runner authentication token (glrt-*). In create mode, the VM joins an existing runner pool and GL_TOKEN / PROJECT_ID / GROUP_ID are not required. In delete mode, setting it selects fleet mode: the shared fleet runner is not deregistered and GL_TOKEN is not required |
GL_TOKEN | yes (create²/delete³) | — | GitLab PAT (Owner role on the target group or project, scopes: create_runner + manage_runner + api). In create mode, required unless RUNNER_TOKEN is set. In delete mode, required unless RUNNER_TOKEN is set (fleet mode) — used to look up and deregister an individually-registered runner |
PROJECT_ID | yes (create)¹ | — | GitLab project ID — registers a project-scoped runner (locked=true) |
GROUP_ID | yes (create)¹ | — | GitLab group ID — registers a group-scoped runner (locked=false). Recommended for platform-service deployments |
GITLAB_URL | yes (create/delete³) | — | GitLab instance URL. In delete mode, required only when GL_TOKEN is used to look up / deregister a runner (not required in fleet mode) |
RUNNER_IMAGE | yes (create) | — | Image pre-pulled as a warm cache; jobs must still set image: in .gitlab-ci.yml |
RUNNER_TAG | no | fullsend-gitlab-runner | Runner tag for job matching |
DRAIN_TIMEOUT_SEC | no | 600 | Delete mode only: cap in seconds on draining in-flight jobs (SIGQUIT + wait for runner-* containers to exit) before deleting the VM. On cap overrun, the script warns and proceeds. See --no-drain to skip draining entirely |
RUNNER_ACCESS_LEVEL | no | not_protected | ref_protected restricts the runner to protected branches and tags, so merge-request pipelines on unprotected source refs never match and sit pending. In project mode, not_protected means any job on any branch of the project can run; in group mode, any tag-matched job on any branch of any project invited into the group tree runs on this VM (see Security below) |
OPENSHELL_VERSION | no | from .github/scripts/openshell-version.sh | OpenShell version used at VM provision (Renovate-tracked). Per-job prepare.sh re-reads the job image's openshell --version and upgrades the host CLI + supervisor when they differ, so a stale VM pin cannot stick. |
GITLAB_RUNNER_VERSION | no | 19.2.1 | gitlab-runner version |
REGISTRATION_TOKEN | setup only | — | GitLab runner registration token |
¹ Exactly one of PROJECT_ID or GROUP_ID must be set when registering a new runner with GL_TOKEN (mutually exclusive). Not required when RUNNER_TOKEN is set.
² Create mode: RUNNER_TOKEN joins an existing pool (and takes precedence if both are set). GL_TOKEN registers a new runner. One of the two is required.
³ Delete mode: RUNNER_TOKEN selects fleet mode (skip deregistration, GL_TOKEN/GITLAB_URL not required). Otherwise GL_TOKEN + GITLAB_URL are required so the script can look up and deregister an individually-registered runner — omitting them is not authorization to skip deregistration.
OpenShift-specific
| Variable | Required | Default | Description |
|---|---|---|---|
NAMESPACE | yes (create/delete) | — | OpenShift namespace |
VM_USER | no | fedora | Cloud-image login user (cloud-user on RHEL/CentOS Stream images) |
SSH_PUBLIC_KEY | no | contents of ~/.ssh/id_rsa.pub or id_ed25519.pub | SSH public key contents (not a path) |
RUNNER_USER | no | VM_USER | Delete mode only: Unix account gitlab-runner/podman run as on the VM (setup.sh's RUNNER_USER). Used to drain as the correct identity when it differs from VM_USER |
GCE-specific
| Variable | Required | Default | Description |
|---|---|---|---|
GCP_PROJECT | yes | — | GCP project ID |
GCP_ZONE | no | us-east1-b | GCE zone |
GCP_MACHINE_TYPE | no | e2-standard-4 | GCE machine type (4 vCPU, 16 GB — closest to the 4 CPU / 14 GiB KubeVirt spec) |
GCP_NETWORK | no | gitlab-runners | VPC network (must have IAP ingress and egress firewall rules) |
GCP_SUBNET | no | — | VPC subnet (required for custom-mode VPCs; omit for auto-mode) |
GCP_USE_IAP | no | true | Use IAP tunneling for SSH. Set to false to create the VM with an external IP and SSH directly. |
GCP_IMAGE_FAMILY | no | fedora-cloud-43-x86-64 | GCE image family |
GCP_IMAGE_PROJECT | no | fedora-cloud | GCE image project |
RUNNER_USER | no | unset | Delete mode only: Unix account gitlab-runner/podman run as on the VM (setup.sh's RUNNER_USER, i.e. whichever identity ran setup.sh). Used to drain as the correct identity when gcloud compute ssh connects as someone else. GCE has no fixed login user equivalent to OpenShift's VM_USER, so unlike there this has no safe default — without it, the drain runs as the connecting identity and can under-report idle if that identity differs from the one gitlab-runner runs as |
Files
create-openshift-vm.sh— end-to-end VM creation on OpenShift + runner registration + setupdelete-openshift-vm.sh— drain in-flight jobs, then OpenShift VM teardown + runner deregistrationcreate-gcp-vm.sh— end-to-end VM creation on GCE + runner registration + setupdelete-gcp-vm.sh— drain in-flight jobs, then GCE VM teardown + runner deregistrationsetup.sh— standalone VM configuration (called by create-openshift-vm.sh / create-gcp-vm.sh). Idempotent and safe to re-run in place as a debug convenience; recreation is the compliance path (see #7257). Re-running it on an already-provisioned VM also installs/refreshes the Podman prune timer.setup_test.sh— unit tests for setup.sh idempotency hygiene (backup, gateway seed skip)podman-prune.sh— reclaims unused rootless Podman containers and images; installed as a user systemd timer by setup.sh and invoked from prepare/cleanuppodman-prune_test.sh— unit tests for the prune script and timer installgitlab-runner-version.sh— central pin for the gitlab-runner versionvm.yaml— KubeVirt VirtualMachine template (OpenShift only)executor/job_id.sh— shared helper resolving the trusted job IDexecutor/prepare.sh— custom executor prepare stage (reaps leftover OpenShell containers, prunes unused images, starts a per-job gateway matched to the job image's OpenShell version)executor/run.sh— custom executor run stageexecutor/cleanup.sh— custom executor cleanup stage (stops the gateway, wipes~/.local/state/openshell/{gateway,tls}, reaps sandboxes)executor/gateway.sh— shared per-job gateway helpers sourced by prepare/cleanup
Executor script layout
install_executor in setup.sh does not glob executor/*.sh;
it copies an explicit five-file allowlist — job_id.sh, prepare.sh,
run.sh, cleanup.sh, gateway.sh — into a flat EXECUTOR_DIR
(~/gitlab-runner-executor) when the VM is provisioned. create-gcp-vm.sh
and create-openshift-vm.sh independently hard-code the same five-file list
to stage and chmod +x the scripts on the VM. A script under executor/
that is not on all three lists (e.g. gateway_test.sh,
prepare_validation_test.sh) is never installed at all, regardless of what
paths it references — a new script must be added to all three lists first.
Once a script is on those lists, parent directories from the source tree are
not preserved when it lands in EXECUTOR_DIR, except those explicitly
seeded by install_executor. Today that exception is only
.github/scripts/openshell-version.sh. Any such script that needs a file
outside its own directory must either:
- Reference only same-directory siblings (the path still works after flattening), or
- Have
install_executorexplicitly copy the needed file intoEXECUTOR_DIR, mirroring theopenshell-version.shprecedent.
Guessing a BASH_SOURCE-relative path across the flattening boundary
silently fails at per-job runtime: prepare.sh/cleanup.sh source the
flattened copy, not the source-tree file. gateway.sh
is the current example of (2); job_id.sh, prepare.sh, run.sh, and
cleanup.sh only reference same-directory siblings.
podman-prune.sh is not an executor script. setup.sh installs it
to ~/.local/lib/fullsend/podman-prune.sh and prepare.sh/cleanup.sh
invoke that path via prune_unused_podman_storage in gateway.sh. Do
not add it to the five-file executor allowlist.
Disk / image prune
These VMs are long-lived. Without periodic reclaim, unused Podman images
accumulate until podman pull fails with no space left on device.
setup.sh therefore:
- Writes
~/.config/fullsend-gitlab-runner/keep-imageswith the pre-pulledRUNNER_IMAGEand OpenShell supervisor tag. - Installs
~/.local/lib/fullsend/podman-prune.shand a user systemd timer (fullsend-podman-prune.timer, hourly, Nice=19) that skips the run when arunner-*oropenshell-*container is in-flight. - Invokes the same helper from
prepare.sh(before the job image pull) andcleanup.sh(after the job container is gone) so a busy runner still reclaims space between jobs.
To apply this to an already-provisioned VM, copy the updated
hack/gitlab-runner-vm/ files onto the VM and re-run setup.sh with
the same GITLAB_URL / RUNNER_IMAGE used at provision time. The
install is idempotent. For an immediate reclaim without waiting for the
first timer tick:
systemctl --user start fullsend-podman-prune.service
Security notes
-
By default (
GCP_USE_IAP=true), GCE VMs are created with--no-address(no public IP) and SSH access is tunneled through IAP, which requires the operator to haveroles/iap.tunnelResourceAccessorand authenticates via GCP IAM — mirroring the OpenShift model where SSH is tunneled through the K8s API server. WhenGCP_USE_IAP=false, the VM gets an external IP and SSH connects directly with trust-on-first-use host-key verification (StrictHostKeyChecking=accept-new) — the first connection accepts the key and subsequent connections within the same run reject changes. The script prints a command to remove the external IP afterward. -
GCE VMs are created with
--no-service-account --no-scopes. The runner does not need a Compute Engine service account: operatorgcloudruns on the workstation, and inference auth is GitLab OIDC → WIF. Attaching the default Compute SA (roles/editor) would expose a stealable OAuth token via the metadata server (169.254.169.254) to the orchestration container (--network=host, no L7 egress policy). Existing VMs created without these flags should have the SA removed (stop, set-service-account, start) or be recreated withcreate-gcp-vm.sh:gcloud compute instances stop "${vm}" --project="${GCP_PROJECT}" --zone="${GCP_ZONE}" gcloud compute instances set-service-account "${vm}" \ --project="${GCP_PROJECT}" --zone="${GCP_ZONE}" \ --no-service-account --no-scopes gcloud compute instances start "${vm}" --project="${GCP_PROJECT}" --zone="${GCP_ZONE}" -
The CA trust bootstrap uses trust-on-first-use (TOFU). For higher assurance, provide the CA bundle out-of-band before running setup.sh.
-
The OCI CA-injection hook fires for all containers on the host. It only copies a CA bundle file and is scoped to the
createRuntimestage. -
Job containers share the host network namespace (
--network=host) to reach the OpenShell gateway. The gateway binds to0.0.0.0(required for the Podman compute driver — sandbox containers register viahost.containers.internal). mTLS protects the endpoint. The gateway process itself is per-job:prepare.shstarts it against a wiped store,cleanup.shstops it. A leftover from a killed job is reaped at the next prepare. -
Job containers receive read-only access to the runner's gateway mTLS credentials (
~/.config/openshell). This is required for the fullsend agent inside job containers to authenticate to the gateway. Credential exposure depends on the runner scope:- Project mode (
PROJECT_ID): the runner is scoped to one project byrunner_type=project_typeandlocked=true, so only jobs from that project can access these credentials. - Group mode (
GROUP_ID): the runner is scoped to the group byrunner_type=group_type, so any project invited into the group tree can run tag-matched jobs on this VM and access the mounted gateway credentials. This is a wider trust boundary than project mode — credential access extends from a single project to every project in the group tree. For group-scoped runners, consider settingRUNNER_ACCESS_LEVEL=ref_protectedas a compensating control to restrict jobs to protected branches and tags. In both modes, access is narrowed byrun_untagged=false(only jobs tagged with the runner's tag are matched) and optionally byref_protected(restricting to protected branches/tags). If job-scoped credential minting is added to the gateway, this mount should be replaced with short-lived per-job tokens.
- Project mode (