sbx-integration.md
August 12, 2026 · View on GitHub
This document explains Docker Sandboxes (the sbx CLI) and how AWF uses it
to run an agent inside a hypervisor-isolated microVM while keeping AWF's own
egress-filtering infrastructure on the host. It is written for two audiences:
- Engineers who want to understand how the existing sbx integration works.
- Ourselves, if we later want to add another microVM backend that runs on top of KVM (e.g. Firecracker, Cloud Hypervisor, or a bespoke krun-based runner).
:::note
This is distinct from Sandbox design, which explains why
AWF's default network sandbox uses Docker containers + Squid rather than a
microVM. This document covers the optional --container-runtime sbx path,
where the agent additionally runs inside a real microVM.
:::
Part 1 — What is Docker Sandboxes (sbx)?
Docker Sandboxes is a Docker product (CLI: sbx, distributed via
docker/sbx-releases) that runs an AI
coding agent inside a lightweight microVM. Where a container shares the host
kernel, each sandbox boots its own Linux kernel, filesystem, and Docker
Engine, so the agent cannot see or touch the host except through explicitly
shared channels.
Isolation model
Docker Sandboxes stacks five isolation layers:
- Hypervisor isolation — a separate kernel per sandbox. On Linux this rides
on KVM (
/dev/kvm; the installer adds the user to thekvmgroup). On macOS/Windows it uses the platform hypervisor. Agent processes inside the VM are invisible to the host. - Network isolation — all HTTP/HTTPS egress is forced through a host-side proxy that enforces the sandbox's configured network policy (Open, Balanced, or Locked Down). Raw TCP, UDP, and ICMP are blocked at the network layer; DNS resolution goes through the proxy. AWF separately applies its deny-by-default Squid ACL.
- Docker Engine isolation — each sandbox runs its own Docker daemon inside
the VM.
docker build/docker compose upexecute against the in-VM engine, with no path to the host daemon. - Workspace isolation — the host workspace is shared via a virtiofs
filesystem passthrough, mounted at the same absolute path as on the host.
Default is a read-write direct mount (edits are live on the host);
--clonegives a read-only mount plus a private in-VM clone. - Credential isolation — the host-side proxy injects auth headers into outbound API requests. Raw credential values never enter the VM.
The host-side proxy and DOCKER_SANDBOXES_PROXY
Every sandbox's outbound traffic exits through a proxy running on the host. That proxy is where policy enforcement and credential injection happen. Two env vars matter to AWF:
DOCKER_SANDBOXES_PROXY— sets the upstream proxy for sandbox traffic only. UnlikeHTTP_PROXY/HTTPS_PROXY, it does not affect image pulls or the daemon's own requests. It acceptshttp://,https://,socks5://, andsocks5h://URLs. It must be set in the environment where the sandbox daemon starts — if the daemon is already running, it must be restarted for a change to take effect.DOCKER_SANDBOXES_NO_PROXY— excludes destinations from the above.
This upstream-proxy hook is the single most important integration point for AWF: it lets AWF interpose its own Squid proxy underneath Docker's sandbox proxy.
Lifecycle and CLI surface
| Command | Purpose |
|---|---|
sbx login | Authenticate the local daemon (required before use). |
sbx create --name <n> <agent> <workspace> [mounts...] | Create a VM with workspace + extra mounts. |
sbx exec [--workdir] [--env K=V] [--tty] <n> <cmd> | Run a command inside the VM. |
sbx run <agent> | Convenience: create + start an agent in one step. |
sbx stop <n> / sbx rm --force <n> | Stop / delete the VM and its contents. |
sbx ls | List sandboxes (also usable as an auth probe). |
sbx daemon status | Inspect the background daemon. |
VMs persist until explicitly removed; stopping an agent does not delete the VM.
Enclave runtimes are independent
container.containerRuntime: "sbx" selects the primary agent's execution
model. enclaves[].runtime: "sbx" on a keyed enclave entry is a separate enclave backend
behind the AWF-owned MCP server and must never reuse the primary agent VM,
agent-ingress capability, gateway capability, or agent credentials.
Both enclave sbx backends are currently fail-closed previews. Docker Sandboxes
v0.37.1 has CPU/memory limits and read-only same-path mounts, but it still
does not prove the full network, PID, disk, per-file size, guest mount-target,
and lifecycle controls AWF requires for unified enclaves. AWF therefore blocks
both enclave sbx backends before staging or Compose assembly, with no
Docker/gVisor fallback.
Promotion is gated on a digest-pinned template plus real-VM proof of network/lateral denial, PID/memory/CPU/disk/file-size enforcement, explicit guest mount targets, credential and cross-invocation isolation, canonical failure bytes, timing buckets, and interruption cleanup. See AWF configuration spec §14 and Unified Enclave Architecture and Migration.
Part 2 — How AWF uses sbx
AWF's default backend runs the agent as a Docker Compose service alongside Squid and the api-proxy. The sbx backend instead runs the agent inside an sbx microVM, while Squid and the api-proxy stay in Docker Compose on the host. Only the agent crosses the hypervisor boundary; its egress is chained back down through AWF's host-side Squid for domain filtering (and, optionally, through the api-proxy for credential injection / model routing / token logging).
The Docker Compose agent resource settings --pids-limit and
container.pidsLimit do not apply to sbx. Docker's agent cgroup is not passed
into the microVM, so pids.max and pids.current are unavailable; AWF emits
an incompatibility warning and ignores the setting.
The executionModel abstraction (src/container-runtime.ts)
AWF centralizes runtime differences in a small registry. Each runtime declares
an executionModel of either compose or microvm:
const RUNTIME_REGISTRY = {
gvisor: { executionModel: 'compose', dockerRuntime: 'runsc', needsStaticDns: true, usesIptables: false },
sbx: { executionModel: 'microvm', dockerRuntime: undefined, needsStaticDns: false, usesIptables: false },
firecracker: { executionModel: 'microvm', dockerRuntime: undefined, needsStaticDns: false, usesIptables: false },
};
Three capability queries drive the rest of the codebase:
resolveDockerRuntime(name)— maps a user-facing name to a Docker OCI runtime (gvisor→runsc); returnsundefinedfor microVM backends (they don't use Docker'sruntime:field).runtimeNeedsStaticDns(name)— whether AWF must inject static/etc/hostsentries (gVisor needs this; sbx manages its own DNS).runtimeUsesComposeAgent(name)— the key switch:falseformicrovmmodels. When false, the agent is not emitted intodocker-compose.yml, and lifecycle is driven by the microVM CLI instead ofdocker logs/docker wait. Infrastructure services (Squid, api-proxy) are generated regardless.
The lifecycle wrapper (src/sbx-manager.ts)
sbx-manager.ts is a thin, well-documented wrapper over the sbx CLI:
createSandbox(config)— probes auth (sbx ls, falling back tosbx daemon statusfor diagnostics), then runssbx create --name <n> shell <workspace> [mounts...]. It translates AWF's Docker-stylehost:container:modemount strings into sbx's positionalpath[:ro]form, deduplicates paths, and always adds/tmp,/usr/local/bin, and$HOMEso agent runtime files and installed CLIs (e.g. Copilot) are reachable.execInSandbox(name, cmd, opts)— runssbx execwith optional--workdir,--tty, and--envflags, streams stdout/stderr, maps timeouts to exit code124, and returns the command's exit code.removeSandbox(name)—sbx stopthensbx rm --force(best-effort).isSbxAvailable()—sbx versionpresence check.
Secret sanitization
sanitizeEnvForSbx() strips secret-looking keys from the inherited process.env
used to launch the sbx CLI. Explicit entries passed through
execInSandbox(..., { environment }) become --env arguments without this
filtering, so callers must sanitize that assembled guest environment separately.
:::caution
sbx create is deliberately not run with the sanitized env. The management
CLI needs some of those vars to look up credentials against the local daemon,
and those never enter the VM (the interior env is controlled separately by
execInSandbox). During create, AWF also temporarily unsets
DOCKER_SANDBOXES_PROXY (so registry/daemon auth isn't forced through a
not-yet-ready Squid) and XDG_CONFIG_HOME (so the sbx CLI finds credentials in
$HOME/.config, not wherever the Copilot harness pointed it).
:::
Volume mounting & toolchain sharing
The most important structural difference from AWF's Docker/gVisor path is that
the microVM does not use a chroot. In compose mode the agent is bind-mounted
under /host/... and then chroots into it; sbx instead uses positional path
mounts where the host path maps to the identical path inside the VM
(/home/runner/work/... on the host is /home/runner/work/... in the guest).
The generated command is:
sbx create --name <name> shell <workspaceDir> [extraMount...] /tmp /usr/local/bin $HOME
(shell is sbx's generic agent image, which supplies the guest base OS.)
What createSandbox() shares, in order:
- Workspace —
workspaceDir = $GITHUB_WORKSPACE || process.cwd(), the first positional mount, read-write. This is the repo checkout the agent edits. - Extra mounts —
config.volumeMounts(from--volume). AWF stores these Docker-style (host:container:mode), but sbx only accepts a positionalhostPathwith an optional:rosuffix, so the manager parses out the host path and mode and discards the container-path segment (host path = guest path). Default is read-write;romode becomeshostPath:ro. - Three always-added system mounts that carry the toolchain and runtime state:
-
/usr/local/bin— the toolchain seam. The Copilot CLI and other host-installed tools live here and are mounted straight in. Note it is narrow: only/usr/local/bin, not/usr,/lib,/lib64, or/opt. -
/tmp— agent runtime files (rendered prompts, logs). -
$HOMEtool dirs — a curated whitelist of writable agent dirs, not the whole home directory. The manager mounts only the subdirs that exist on the host fromHOME_TOOL_SUBDIRS(.cache,.config,.local,.azure,.anthropic,.claude,.cargo,.rustup,.npm,.nvm) plus the agent state dirs.copilotand.gemini. Credential-store dirs such as.aws,.ssh,.docker,.kube, and.gnupgare never whitelisted, so they never enter the VM. Each whitelisted dir is mounted wholesale (as a directory — sbx positional mounts cannot target an individual file, so its loose files like~/.copilot/mcp-config.jsonare preserved).:::note
.azureis a credential-bearing exception.azureis mounted to provide Azure CLI config and account metadata. However, its live token caches (msal_token_cache.bin,msal_token_cache.json,accessTokens.json,service_principal_entries.json) are treated as credential stores and scrubbed before sandbox creation (sbx) or masked with/dev/nulloverlays (compose). Agents cannot read host Azure auth tokens directly. Azure API authentication must be handled by the api-proxy's sidecar-only OIDC exchange or by an external trusted service. The separateADO_MCP_AUTH_TOKENenvironment variable remains available for ADO MCP. :::
-
Scrubbing nested credential stores. Several whitelisted dirs legitimately
hold tool settings but also stash a secret in a well-known child — e.g.
.config/gh, .config/gcloud, .cargo/credentials, .claude/.credentials.json,
.gemini/oauth_creds.json, and the Azure CLI token caches under .azure
(msal_token_cache.bin, msal_token_cache.json, accessTokens.json,
service_principal_entries.json). Because the parent is mounted
wholesale and sbx cannot overlay or mask a nested path, the manager instead
moves those credential paths aside on the host before sbx create and restores
them after the sandbox is torn down (scrubHomeCredentials /
restoreHomeCredentials in sbx-manager.ts). The move target is a
.awf-sbx-cred-backup-<pid> dir at the home root — never a mounted subdir — so
the secrets are absent from the VM while the benign tool state stays available.
This is the sbx analog of compose mode's /dev/null credential overlays, and the
central credential list in sandbox-mount-policy.json is shared between backends
to prevent drift. The agent accesses OIDC-backed providers through requests
routed to the api-proxy, or receives separately allowed environment credentials
such as ADO_MCP_AUTH_TOKEN, not by reading the host's on-disk auth store.
Provider credentials and Actions OIDC request variables remain in the api-proxy
or another trusted external service, so removing these paths is safe.
A seenPaths set deduplicates so no path is mounted twice, and
execInSandbox(..., { workDir }) passes --workdir so commands run inside the
mounted workspace.
| Aspect | sbx (microVM) | Docker / gVisor (compose agent) |
|---|---|---|
| Path model | Host path == guest path | Bind-mounted under /host, then chroot /host |
| System libraries | From the sbx shell guest image | Host /usr,/bin,/lib,/lib64,/opt mounted read-only |
| Toolchain binaries | Host /usr/local/bin mounted in | /usr from the host or sysroot; optional chroot.binariesSourcePath overlay at /host/tmp/awf-runner-bin (ro) |
| Workspace | workspaceDir positional (rw) | <workspaceDir>:/host<workspaceDir>:rw |
| Home | Curated $HOME tool-dir whitelist (rw), nested credential stores scrubbed before create | Empty home volume with only whitelisted subdirs |
:::caution Toolchain portability
Because host system libraries are not shared into the VM, a binary in
/usr/local/bin that dynamically links against host-specific libraries — or
expects an interpreter/runtime under /usr — can fail inside the microVM unless
the sbx guest image already provides a compatible base. This is the sharpest
contrast with compose mode's broad read-only /usr+/lib mounts, and the single
most important detail to plan for when building a KVM-based microVM backend:
you must either ship a guest image whose base matches the tools you mount in, or
widen the mount set to include the libraries those tools need.
:::
Wiring into the main workflow (src/commands/main-action.ts)
main-action.ts decides useSbx = !runtimeUsesComposeAgent(config.containerRuntime)
and, when true, substitutes two functions into the shared workflow runner:
sbxStartContainerswraps the normalstartContainers(which brings up the infra-only compose: Squid + api-proxy, no agent service), then:- Verifies
isSbxAvailable(). - Builds the agent environment (
buildAgentEnvironment) using microVM-specific network targets (see below), merging credential env (buildAgentCredentialEnv) when the api-proxy is enabled. - When enclaves are enabled, waits for mcpg to prove the enclave MCP backend is registered and reachable before primary-agent startup. No enclave endpoint, capability, repository list, or private state enters the VM.
- Calls
createSandbox({ workspaceDir, squidIp: SQUID_IP, extraMounts }). - When the api-proxy is enabled, runs
assertSbxApiProxyReflect: creates a privateHOSTALIASESresolver file mappingapi-proxyto a loopback HTTP bridge. The bridge forwards tohost.docker.internal:10000with the host header expected by docker-sbx, then AWF probeshttp://api-proxy:10000/reflectwith Node's built-infetchand up to 30 retries. Startup aborts (fail-closed) if the endpoint is unreachable after all retries, because an isolated runtime that cannot reach/reflectcannot do model auto-resolution or credit accounting. - Runs a Squid connectivity diagnostic (
curl --proxy ... https://api.github.com).
- Verifies
sbxRunAgentCommandruns the actual agent command withexecInSandbox, honoring the agent timeout, workdir, TTY, and computed environment, and dumps api-proxy logs on non-zero exit for debugging.
Cleanup: when the runtime is microVM and --keep-containers is not set,
removeSandbox(SBX_DEFAULT_NAME) runs before compose teardown.
Networking: crossing the VM boundary
Inside the microVM, AWF's internal Docker network (172.30.0.0/24) is not
reachable — the VM is on its own network. AWF compensates with two indirections:
- Squid is reached at the sbx gateway IP (
172.17.0.0in code, i.e. the docker0 bridge range) on its published port3128, rather than the internal172.30.0.10. AWF setsHTTP_PROXY/HTTPS_PROXYinside the sandbox tohttp://172.17.0.0:3128, so proxy-aware agent tools route through AWF's Squid domain ACL.DOCKER_SANDBOXES_PROXY(the sbx daemon-level upstream knob) is not currently set by AWF — it must be present when the sbx daemon starts, a point AWF does not control — so the sbx daemon's own egress is not filtered by AWF's Squid. - The api-proxy (credential injection) is reached via
host.docker.internal, which resolves to the docker0 bridge from inside the VM.COPILOT_*/ proxy env vars are pointed there instead of at172.30.0.30. - The enclave control plane is not mounted into the sbx primary sandbox.
When unified enclaves are enabled, the primary agent reaches them only
through the externally launched
gh-aw-mcpggateway after AWF proves the run-labelled handoff and end-to-end readiness. The AWF-ownedenclave-mcp-serverstays on its own private control network, outsideawf-netandawf-ext.
The net effect: agent tools that respect HTTP_PROXY/HTTPS_PROXY route through
AWF's Squid domain ACL; credentials are injected by AWF's api-proxy. Tools that
bypass proxy env vars are subject only to sbx's own host-side policy, not AWF's
ACL.
Configuration surface
- CLI:
--container-runtime sbx. - gh-aw workflow frontmatter:
sandbox.agent.runtime: docker-sbx(see thesmoke-docker-sbx*workflows under.github/workflows/). - Strict-security note (
src/commands/validators/security-mode.ts): sbx enforces isolation at the hypervisor layer viaDOCKER_SANDBOXES_PROXY, so AWF's Docker network-isolation topology is not forced on for it (isMicroVmRuntimeskips that override). The Firecracker microVM backend is the exception: it explicitly attaches its host-side veth to AWF's proven internal bridge, so network-isolation topology is still forced on for--container-runtime firecracker. The api-proxy is still always enabled for both.
End-to-end traffic flow
flowchart TB
subgraph host["Host (CI runner / dev machine)"]
cli["awf CLI (main-action.ts)"]
sbxproxy["sbx host-side proxy<br/>(daemon-level)"]
subgraph compose["Docker Compose (infra only)"]
squid["Squid proxy<br/>domain ACL"]
apiproxy["api-proxy<br/>credential injection"]
end
subgraph vm["sbx microVM (KVM)"]
agent["Agent command"]
end
end
cli -->|sbx create/exec| vm
agent -->|"HTTP_PROXY/HTTPS_PROXY (proxy-aware)"| squid
agent -->|proxy-unaware tools| sbxproxy
agent -->|COPILOT_* via host.docker.internal| apiproxy
apiproxy --> squid
squid -->|allowed domains| internet["Internet"]
sbxproxy -->|not filtered by AWF ACL| internet
Part 3 — Adding another KVM-based microVM backend
Because the microVM path is abstracted behind a small set of seams, adding a new
KVM backend is mostly a matter of implementing a manager and registering it.
Firecracker (src/firecracker-runtime-backend.ts, src/firecracker/) is a real,
fail-closed workload preview built on this same seam (gated behind
--firecracker-preview) — unlike sbx, it attaches its host-side veth directly to
AWF's proven internal awf-net bridge and reaches Squid/api-proxy at their
normal internal IPs (SQUID_IP/API_PROXY_IP from
src/config/network-policy.ts) rather than a docker0 gateway IP, and it keeps
strict network-isolation topology forced on (see
src/commands/validators/security-mode.ts). Cloud Hypervisor, krun/libkrun,
QEMU/KVM, etc. remain hypothetical. Here is the
checklist.
1. Register the runtime
Add an entry to RUNTIME_REGISTRY in src/container-runtime.ts:
myvm: {
executionModel: 'microvm',
dockerRuntime: undefined, // not a Docker OCI runtime
needsStaticDns: false, // set true only if the VM can't reach an expected DNS
usesIptables: false, // microVM manages its own network egress
},
That entry makes runtimeUsesComposeAgent('myvm') return false, which omits
its agent from docker-compose.yml and skips the Docker network-isolation
override in strict mode. It does not select the new manager: register an
ExternalRuntimeBackendFactory in src/external-runtime-backend-resolver.ts.
2. Implement an external runtime backend
Implement ExternalAgentRuntimeBackend from
src/external-runtime-backend.ts, following SbxRuntimeBackend in
src/sbx-runtime-backend.ts (docker0-gateway addressing) or
FirecrackerRuntimeBackend in src/firecracker-runtime-backend.ts
(jailed microVM attached directly to awf-net, using internal
SQUID_IP/API_PROXY_IP addressing instead). The backend owns preflight,
startup, execution, diagnostics, and idempotent stop state. Concretely, a KVM
backend must:
- Boot a microVM on
/dev/kvmwith a kernel + rootfs. Confirm KVM is available (/dev/kvmpresent, user in thekvmgroup). On stock GitHub-hosted runners/dev/kvmis not exposed — this backend is only for self-hosted / nested-virt-capable environments. - Mount the workspace at its host absolute path (virtiofs is the sbx choice;
Firecracker typically uses a virtio-blk device or a virtiofsd sidecar). Also
surface
/tmp,/usr/local/bin, and$HOMElikecreateSandboxdoes. - Inject the agent environment, and sanitize secrets first — reuse the
sanitizeEnvForSbx()pattern (stripTOKEN|SECRET|KEY|...). - Return the agent's exit code faithfully (AWF propagates it), mapping
timeouts to
124.
3. Chain egress through AWF's Squid (the critical seam)
This is the property that keeps AWF's guarantees intact. Your VMM must force all sandbox egress through AWF's host-side Squid:
- If the VMM offers an enforced upstream-proxy knob analogous to
DOCKER_SANDBOXES_PROXY, point it athttp://<squidGatewayIp>:3128. - Otherwise, enforce egress outside the guest (for example at the VMM/TAP or
host firewall) so the guest can reach only Squid. Guest
HTTP_PROXY/HTTPS_PROXYsettings may improve client compatibility, but are not an enforceable security boundary. - Reproduce the boundary-crossing addressing pattern: sbx reaches Squid at the
bridge gateway IP + published port (not the internal
172.30.0.x) viaSBX_GATEWAY_IP/SBX_HOST_DOCKER_INTERNALinsrc/sbx-runtime-backend.ts, because the sbx microVM sits outsideawf-net. Firecracker instead attaches its veth directly toawf-net, so it addresses Squid/api-proxy at their normal internal IPs (SQUID_IP/API_PROXY_IPfromsrc/config/network-policy.ts) — pick whichever addressing model matches how your VMM's network attaches to the host.
4. Register the backend
Add the factory to EXTERNAL_RUNTIME_BACKENDS in
src/external-runtime-backend-resolver.ts. main-action.ts resolves exactly one
backend instance, adapts it to WorkflowDependencies, and uses that same
instance for cleanup and signal handling. Compose-managed Docker and gVisor
runtimes bypass this adapter.
5. Things to get right (lessons from the sbx path)
- Daemon/proxy ordering — don't force registry/daemon auth through Squid
before Squid is healthy (sbx unsets
DOCKER_SANDBOXES_PROXYduringcreate). - Cross-boundary health checks — compose
depends_ondoes not extend into a VM. The current sbx path polls api-proxy and probes Squid, but logs/warns and proceeds on failure; a backend requiring a true health gate must abort startup. - Env leakage during management vs. interior — management commands may need host env for auth; the interior env must be sanitized. Keep those two paths separate.
- DNS — decide whether the VM resolves DNS itself or via the proxy; set
needsStaticDnsaccordingly. - Exit code + timeout fidelity — CI relies on accurate propagation.
Division of responsibility
| Concern | Docker Sandboxes (sbx) | AWF |
|---|---|---|
| Hypervisor / kernel isolation | ✅ owns it | delegates to backend |
| In-VM Docker engine | ✅ owns it | — |
| Workspace mount | ✅ virtiofs passthrough | passes mount list |
| Domain egress ACL | its own default policy | ✅ enforced by AWF Squid (chained upstream) |
| Credential injection | its own proxy | ✅ AWF api-proxy (COPILOT_* retargeted) |
| Secret scrubbing of interior env | credential isolation | ✅ sanitizeEnvForSbx() |
| Lifecycle orchestration | CLI primitives | ✅ sbx-manager + main-action |
The takeaway for a new backend: AWF keeps ownership of egress filtering and credential injection; the microVM backend only needs to provide isolation plus a way to force egress through AWF's Squid.
References
- Docker Sandboxes docs: https://docs.docker.com/ai/sandboxes/ (architecture, security / isolation)
sbxreleases: https://github.com/docker/sbx-releases- AWF source:
src/container-runtime.ts,src/sbx-manager.ts,src/commands/main-action.ts,src/commands/validators/security-mode.ts,src/firecracker-runtime-backend.ts(second microVM backend built on this seam) - Related: Sandbox design, Architecture