Agent runtimes

September 5, 2026 · View on GitHub

A runtime is the agent program fullsend runs inside the sandbox — the thing that talks to the model and executes tool calls. fullsend run delegates to it and owns everything around it: the sandbox, the credentials, and the verdict.

RuntimeUse it forStatus
claudeProduction agent runs (Claude Code)Default
piSecond runtime, opt-in per repo — Claude, Grok and Gemini on Vertex; GPT via OpenAI WIF (wired, not yet exercised live)Supported for all roles
codexThird runtime, opt-in per repo or agent — OpenAI models only, via the same secretless credential path (wired, not yet exercised live)Opt-in
dummyBehaviour tests — scripted ops, no inferenceInternal
dummy-playbackBehaviour tests — replays canned agent results from a playlist, no inferenceInternal
opencodeNot yet functionalStub

Pick one with runtime: in .fullsend/config.yaml, or per run with --runtime.

fullsend run triage --runtime pi --model xai-vertex/xai/grok-4.6

How a run uses the runtime

The runner owns the sandbox, credentials and verdict; the runtime owns what happens between "start" and "event stream".

sequenceDiagram
  autonumber
  participant R as Runner
  participant S as Sandbox
  participant A as Runtime
  participant M as Model
  R->>R: pick runtime (config.yaml)
  R->>S: .env, host files
  R->>S: Bootstrap
  R->>S: OIDC token (4-min refresh)
  R->>S: clean up stray processes (between iterations)
  R->>S: Run (per iteration)
  S->>A: start + hook wiring
  loop tool-use loop
    A->>M: request (WIF)
    M-->>A: response
    A->>A: Pre → tool → Post hooks
  end
  A-->>R: event stream
  R->>S: extract artifacts
  R->>R: verdict, metrics.json

Choosing a runtime

Claude Codepicodex
ModelsAnthropic on VertexClaude, Grok and Gemini on Vertex; GPT via OpenAI WIF (opt-in, not yet exercised live)GPT only, via OpenAI WIF (not yet exercised live)
Sub-agentsNative (Agent tool)Agent/Task via a fullsend extensionNot available
Fallback model chainFULLSEND_FALLBACK_MODELS, tried in orderIgnored with a warningIgnored with a warning
RolesAllAll; review/retro at --thinking medium by defaultSame recommendation as before — no sub-agent roster on codex
Effort--effort low..max--thinking, same levels (high when unset)model_reasoning_effort, same levels
ToolsNative Claude permission syntax--tools (strict) + a first-token Bash allowlistShell + apply_patch only; tools: is recorded, not enforced (the allowlist hook is opt-in)
Security controlsFull matrixFull matrix; stricter on failed-call sanitizingFull matrix; post-tool hooks detect and block but cannot rewrite output
Cost in metrics.jsonReportedReportedNot reported — codex sends none

All three run unattended in the same sandbox, behind the same egress allowlist. Stay on claude when you need a fallback chain. Choose pi when you want a non-Anthropic model, several vendors from one runtime, or its Agent/Task sub-agent roster. Choose codex when you want OpenAI models specifically and codex's shell-centric way of working.

Selecting a runtime and model

First non-empty wins — the usual flag > env var > config > default. fullsend run resolves this once, validates it, prints the source, and records it in metrics.json; runtimes never read the override variables themselves.

flowchart LR
  F["--runtime / --model<br/>(flag)"] --> E["FULLSEND_RUNTIME<br/>FULLSEND_MODEL"]
  E --> R["config.yaml<br/>agents: entry for the agent"]
  R --> C["config.yaml runtime:<br/>harness model:"]
  C --> A["agent frontmatter<br/>model:"]
  A --> D["default<br/>claude · opus"]
  classDef s fill:#e3e9fb,stroke:#2d5be3,color:#1b2230;
  classDef d fill:#eceee8,stroke:#a9afa4,color:#1b2230;
  class F,E,R,C,A s;
  class D d;
SettingFlagEnvConfig (per-agent)Config (repo-wide)
Runtime--runtimeFULLSEND_RUNTIMEruntime: on the agent's agents: entryruntime: in .fullsend/config.yaml (repo default)
Model--modelFULLSEND_MODEL (FULLSEND_PI_MODEL on pi and FULLSEND_CODEX_MODEL on codex are lower-precedence aliases)model: on the agent's agents: entryharness model:, then agent frontmatter model:; models.aliases in .fullsend/config.yaml remaps the alias any of these resolve to
Effort--effortFULLSEND_EFFORTeffort: on the agent's agents: entryharness effort:

In CI these are repository variables of the same name, plain or role-prefixed (TRIAGE_FULLSEND_MODEL), so a repo can switch one role's model without a pull request. For durable per-agent configuration that lives in the repository and is reviewable, use the agent's agents: entry in .fullsend/config.yaml instead. Harness env.runner does not reach the fullsend process.

Per-agent runtime, model and effort

The agents: list is the per-agent place in config.yaml: an entry names an agent and can set its runtime, model, effort and subagents. A built-in agent (triage, code, review, fix, retro, prioritize) is tuned with a name-only entry; a custom agent carries the settings on its source: entry.

runtime: pi                    # repo default for agents that set none
agents:
  - name: triage
    model: xai-vertex/xai/grok-4.6
  - name: code
    runtime: claude
    model: sonnet
    effort: high
  - name: review
    subagents:
      default: haiku           # personas that name no model, and children that name no persona
      correctness: opus
  - source: https://raw.githubusercontent.com/acme/agents/<sha>/harness/lint.yaml#sha256=…
    model: haiku

Or from the CLI, which validates the entry before writing it:

  1. fullsend agent set code --fullsend-dir .fullsend --runtime claude --model sonnet --effort high
  2. fullsend agent list --fullsend-dir .fullsend shows the settings next to each agent — code (built-in) [runtime=claude model=sonnet effort=high], or the source: path for a custom agent.
  3. The next fullsend run code names the entry as the source — Runtime: claude (from <config path> agents.code) — and a --runtime/--model flag on that run still wins.

An invalid value is refused before the write — invalid effort "turbo": must be one of low, medium, high, xhigh, max — and the same check runs on every fullsend run, so a hand-edited entry fails the run before a sandbox starts rather than being skipped.

A source: entry needs no name: — the agent's name is derived from the source file (harness/lint.yamllint, ADR 0058), and that is the name the settings, fullsend run lint and fullsend agent set lint all use; add name: only to override it.

Names are agent names as passed to fullsend run <agent>not harness role: values (code and fix both carry role: coder) — matched case-insensitively. A name-only entry for anything that is not a built-in agent fails validation (coder gets a "did you mean code" hint); a custom agent gets its settings on its own entry.

Precedence: flag > env var > the agent's agents: entry > repo-wide runtime: / harness model: effort: > default. Entries merge per field across the layered config (config.yaml over config.base.yaml), so a preset base can tune agents too. fullsend run validates the whole agents: list in every layer (names, runtime, model syntax, effort) and fails the run with an error naming the file and entry rather than silently skipping a mistyped entry or handing a bad value to the runtime.

A value that came from here shows up as <config path> agents.<name> wherever the selection is surfaced (plan block, stderr, metrics.json — see below); the path is the effective config file.

provider/id is pi's model form. The syntax is accepted for every runtime (model ids are not a closed set), but an entry that pairs runtime: claude with a provider/id model gets a warning in the plan block — Claude Code expects an alias (opus, sonnet, …) or an Anthropic model id.

Migrating from repository variables. A repo that carries <ROLE>_FULLSEND_MODEL / <ROLE>_FULLSEND_RUNTIME variables can move them onto agents: entries one-to-one: the variable prefix is the agent name (CODE_FULLSEND_RUNTIME=claude- name: code / runtime: claude). Delete the variable afterwards — while it exists it still wins, so the config entry would be silently shadowed. Bump the workflow's fullsend pin to a version that carries per-agent settings before adding them: an older pinned CLI rejects an enabled agents: entry without a source, whereas a current CLI validates the settings on every run.

Set the runtime per repo with fullsend github setup <owner/repo> --runtime pi. Repos on pi need a sandbox image that carries PI_VERSION; repos on codex need one that carries CODEX_VERSION.

Models

On Claude Code, pass an alias (opus, sonnet, haiku, fable) or a model id.

On pi, a model is provider/id — aliases and bare ids still work, and the provider comes from FULLSEND_PI_PROVIDER (default anthropic-vertex). pi reaches Claude, Gemini and Grok, each through its own provider; see Pi › Models and providers.

On codex, a model is an OpenAI id — openai/<id> or the bare id. The Claude aliases above do not apply: codex serves the OpenAI Responses API only, so opus and friends are refused rather than remapped to a GPT model. Because the fleet harnesses say model: opus, a repo moving to codex names its model either once on the runner with FULLSEND_CODEX_MODEL=openai/gpt-5.6-luna or per agent with model: openai/<id> on the agents: entry — no harness needs editing either way. See Codex › Models.

Harness model: and agents: entry model: values accept provider-qualified provider/id syntax (e.g. google-vertex/gemini-3.8-flash). On pi, a harness can also select a provider with a bare model: plus FULLSEND_PI_PROVIDER.

Per-repo alias overrides

Point an alias at a different model for one repo with models.aliases in .fullsend/config.yamlsonnet: claude-sonnet-5 changes sonnet and leaves the other aliases alone. Works on both runtimes; the override applies to the parent run and to sub-agent dispatch on pi (children resolve aliases through the same merged table). See Pi › Per-repo alias overrides for the syntax and what the plan block shows, and Claude Code › Models for the one limit there (sub-agent model: frontmatter is not remapped).

Where the selection appears

SurfaceWhat it shows
Run plan blockRuntime: <name> (from <source>) next to Model and Effort; <source> is the flag, the variable, or <config path> (suffixed agents.<name> when the agent's entry decided)
stderrruntime: selected "<name>" from <source>
Status comment / ::notice::Runtime · Model: <requested → reported> · Effort · Cost
OTel spanfullsend.runtime, next to gen_ai.request.model
metrics.jsonruntime, requested_runtime, runtime_source, requested_model, override_source

requested_model is the model after the per-run overrides (an alias stays the alias name) and override_source says where it came from, with , remapped by <config path> models.aliases appended when a per-repo alias override applied — so a silent override is visible after the fact. The reported model is the provider-stripped id (claude-opus-4-6); for a provider whose ids are publisher-qualified it keeps that segment (xai/grok-4.6), since that is the wire id.

Harness config keys per runtime

Harness keys are runtime-neutral in YAML; each runtime owns the translation. Test-only runtimes (dummy, dummy-playback) ignore all harness config keys and are omitted from this table.

Harness keyClaude Codepicodex
model--modelalias table (merged with models.aliases), then provider/id; see Models--model <id>; OpenAI ids only
effort--effort--thinking (superset of the harness levels; high when unset)model_reasoning_effort (same levels)
tools:Native Claude permission syntax--tools (strict) + a first-token Bash allowlistNo native allowlist. Bash(...) lists are recorded but not enforced, entries with no codex tool are dropped with a warning, and the tool-allowlist hook is opt-in (FULLSEND_TOOL_ALLOWLIST)
skillsCLAUDE_CONFIG_DIR/skills/PI_CODING_AGENT_DIR/skills/, discovered nativelyCODEX_HOME/skills/, discovered natively
pluginsLoads the plugin.json directories (marketplace layout)Loads the extension directories: uploaded to PI_CODING_AGENT_DIR/extensions/, tree-hash preflight, -e (Plugins, ADR 0094)Unsupported — warned and skipped
security.sandbox_hookshooks.json via --settingsHook scripts + manifest + adapter extensionhooks.json + adapter script under CODEX_HOME
validation_loop.feedback_modeReplaces the prompt on retrySameSame

Full per-key detail, including the exact --tools mapping and allowlist parsing rules, is in Implementing an agent runtime.