Threat Model

September 11, 2026 · View on GitHub

Last reviewed: 2026-09-04 | Applies to: v1.9.1+

Rampart is a policy engine for AI agents — not a sandbox, not a hypervisor, not a full isolation boundary. This document describes what Rampart protects against, what it doesn't, and why.

What Rampart Is

A firewall for AI agent tool calls. It evaluates agent-reported shell commands, file operations, fetch requests, and related metadata against YAML policies. Rampart sees what the framework sends it, not raw syscalls or network traffic. It is designed to reduce common accidental and prompt-injection-driven tool misuse; no universal detection-rate percentage is claimed.

Primary Threat: Misbehaving AI Agents

Rampart's target threat is an AI agent that:

  • Hallucinated a destructive command (rm -rf /, DROP TABLE)
  • Was manipulated by prompt injection (malicious content in a file or webpage told it to exfiltrate data)
  • Made a well-intentioned mistake (wrong environment, wrong file, wrong server)
  • Escalated beyond its intended scope (sub-agent spawning unrestricted tool calls)

These agents aren't adversarial — they're confused, manipulated, or wrong. Rampart catches them reliably.

Not the Target: Adversarial Human Attackers

Rampart does not claim to stop a skilled human who has already compromised your system. If an attacker has shell access, they can bypass Rampart the same way they'd bypass any userspace tool. Rampart is one layer in defense-in-depth, not a replacement for OS hardening, network segmentation, or access control.

Trust Boundaries

┌─────────────────────────────────────────────┐
│ Trusted                                      │
│  • Policy files (admin-authored YAML)        │
│  • Rampart binary                            │
│  • rampart serve process                     │
│  • Audit log directory (when user-separated) │
│  • Policy registry sources (when verified)   │
│  • HMAC signing key (~/.rampart/signing.key) │
├─────────────────────────────────────────────┤
│ Untrusted                                    │
│  • AI agent tool calls (all input)           │
│  • Agent-generated commands                  │
│  • MCP tool call parameters                  │
│  • Webhook response payloads (validated)     │
│  • Project-local .rampart/policy.yaml files  │
│  • Community policies (verified by SHA-256)  │
└─────────────────────────────────────────────┘

Policy files are the security boundary. If an attacker can modify policy files, Rampart's guarantees do not hold. This is why user separation is recommended for production.

Known Limitations

1. Interpreter Bypass

For an exec call that reaches a Rampart integration, Rampart evaluates the command string exposed at that integration boundary. Native hooks (Claude Code, Codex, Cline, Gemini CLI, GitHub Copilot), wrap mode, preload mode, and the HTTP API can all evaluate such a command, but their interception coverage differs. If an agent runs python3 script.py through a covered boundary, Rampart sees and evaluates python3 script.py — but cannot inspect what script.py does internally.

Mitigations:

  • LD_PRELOAD cascade (v0.1.9+): rampart preload intercepts supported child-process calls made by compatible dynamically linked programs. rampart wrap is a separate cooperative boundary: it covers $SHELL and PATH-resolved shell launches, not absolute shell paths or direct process APIs.
  • Interpreter one-liner blocking (v0.6.9+): Patterns like python3 -c, node -e, ruby -e, perl -e with dangerous system calls are blocked by default in standard/paranoid policies.
  • The optional rampart-verify sidecar uses LLM classification to assess intent of ambiguous commands via action: webhook.

Remaining surface: LD_PRELOAD cascade applies only to preload mode, not wrap or native hooks (Claude Code, Codex, Cline, Gemini CLI, GitHub Copilot). Programs that use native file I/O without shelling out, or setuid binaries that drop LD_PRELOAD, are not covered.

Multi-step sequences: With file tool coverage enabled (native hooks, OpenClaw native plugin, or legacy --patch-tools), write-then-execute sequences are evaluated at both steps independently — the write is checked against file policies and the exec against command policies.

2. Audit Log Rewrite

The hash-chained audit trail detects inconsistent event hashes and broken links between retained records. It cannot prove that every action was logged. A complete rewrite with a new valid chain, or deletion of a valid suffix beyond the retained checkpoint, is not detectable from the remaining local history alone.

Rampart v1.6 also validates older logs containing chronological chain epochs created by legacy service restarts. Recovery verifies every event hash, internal link, continuation header, epoch order, and whole-graph connectivity without rewriting the existing records. Forks, missing references, disconnected events, and overlapping epochs fail verification.

Audit records are capped at 2 MiB. If JSON escaping or decision guidance would exceed that limit, Rampart replaces large fields with their original encoded size and SHA-256 digest, then computes the event hash over the compacted record. This preserves a bounded, correlatable decision record instead of silently dropping the audit event.

Rampart redacts common credential patterns, sensitive-key values, and encoded command copies before audit events reach local JSONL, syslog, or CEF sinks. Redaction is defense in depth rather than a data-classification guarantee: unusual secret formats or sensitive source text can still appear in request content, so operators must continue to protect audit storage and exports.

Mitigations:

  • Run rampart serve as a separate user so the agent can't access audit files
  • Configure SIEM export (--syslog or --cef) and a separately controlled collector to retain received events outside the agent's write authority. Export is best effort; the transport and collector determine delivery and retention guarantees.
  • Webhook notifications can retain selected alerts in a separate system; they are not a complete audit trail.
  • audit-anchor.json is a local checkpoint beside the audit files, updated every 100 events by default. Verification checks it against retained events, but an identity that can replace both logs and checkpoint can rewrite them consistently. Append-mode files are not immutable storage.

For independently retained evidence, configure collection, access control, and retention outside the agent's authority. Rampart's local verifier does not retrieve or validate remote SIEM evidence. See the SIEM integration guide.

3. Token Exposure in Wrap Mode

The rampart wrap shim needs a bearer token to authenticate with rampart serve. The token is stored in a 0600 file (not embedded in the shim script). However, the agent user can still read this file.

Mitigations:

  • Run rampart serve as a separate user and store the token in that user's home directory with restricted permissions
  • Use native hooks (rampart setup claude-code) instead of wrap mode — hooks don't require a token file

Resolution: Use user separation for production deployments, or prefer native hooks which don't require a token file.

4. Encoding and Obfuscation

An agent could encode commands to bypass pattern matching:

  • Base64-encoded payloads
  • Variable expansion and command splitting
  • Subshell nesting, heredocs, process substitution

Mitigations (v0.6.9+):

  • Rampart decodes base64 commands before pattern matching
  • Leading shell comments and ANSI escape sequences are stripped
  • Null bytes and control characters are removed
  • Subcommand extraction: $(cmd), backticks, eval 'cmd' — inner commands are matched independently
  • Common obfuscation patterns (base64 *, eval *, xxd -r | bash) trigger deny rules in standard policy
  • The semantic verification sidecar classifies intent regardless of encoding

Coverage: The two-layer approach (pattern matching + LLM classification) significantly reduces the obfuscation surface. Pattern matching catches known encodings; the LLM layer catches intent regardless of how the command is formatted. v0.6.9 closed 10 specific bypass vectors identified in a security audit.

5. OpenClaw Integration Boundaries

OpenClaw has a native plugin path, which is the preferred integration. On older supported hosts Rampart keeps global tools.exec.ask off; on current hosts it preserves the canonical tools.exec.mode policy, removes only equivalent retired ask/security siblings, and refuses conflicting mixed policies. Rampart evaluates tool calls first and returns OpenClaw's native requireApproval result for calls that match a Rampart ask rule. A stricter operator-selected OpenClaw exec mode can add another approval or block after Rampart, but cannot grant past the Rampart hook.

What this means in practice:

  • Allow rules pass through without a Rampart-originated approval prompt; a stricter OpenClaw exec mode may still gate them
  • Deny rules short-circuit before native approval
  • Ask rules request native approval for the matched supported tool call only when its complete redacted review fits the host's 512-character description limit after escaping; otherwise the plugin blocks before creating an approval

Native plugin approvals offer allow-once and deny. They do not create permanent command/path rules. OpenClaw owns the resume operation; Rampart's resolution callback cannot inspect or veto the resumed parameters. See the OpenClaw integration limits for the trusted-plugin composition boundary.

--patch-tools still exists as a compatibility path for older OpenClaw setups and broader file-tool interception, but it remains fragile because it modifies installed framework files.

Security implications of the legacy patch path:

  • Timing window: Between framework upgrade and re-patch, patched file tools can bypass Rampart
  • Silent degradation: If a new OpenClaw version changes the patched integration points, the patch can fail-open until setup is checked and re-applied

Trade-off: The native plugin path is cleaner and more durable. The legacy patch path still closes real gaps on older setups, but it should be treated as compatibility machinery, not the long-term design.

6. Degraded-Mode Behavior

Rampart does not behave identically across every integration when policy evaluation becomes unavailable. That difference is a real security boundary and has to be understood clearly.

Current behavior:

  • rampart wrap --mode enforce runs its own local policy service and denies shell commands when that service cannot confirm enforcement. Monitor mode permits them with a warning. This applies only to commands that reach the cooperative shell boundary.
  • rampart preload defaults to fail-open for transport and server failures unless explicitly configured fail-closed.
  • The native OpenClaw plugin defaults to fail closed for every tool when its policy service is unavailable. Operators can explicitly opt tools into degraded fail-open behavior with failOpenTools or the deprecated failOpen setting; manual setup can retain those choices. rampart protect openclaw clears the exceptions and configures every tool to fail closed.
  • Native hook integrations (Claude Code, Codex, Cline, Gemini CLI, GitHub Copilot) evaluate policies locally in-process, so they do not depend on rampart serve for the core allow/deny path. Codex and Gemini approval-required actions still need the external Rampart queue and deny if it is unavailable; Copilot uses its native ask prompt.

Mitigations:

  • Monitor the Rampart service and alert on downtime
  • Use systemd/launchd to auto-restart on failure (rampart serve install does this)
  • Prefer native hooks or the native OpenClaw plugin when you want less reliance on a long-running local service
  • For OpenClaw, use rampart protect openclaw for the strict fail-closed posture, or manage failOpenTools explicitly in advanced setups

Trade-off: Fail-open improves availability but creates a temporary security gap during outages. Fail-closed reduces bypass risk but can break agent workflows when the policy service is sick. Rampart makes that trade-off explicit per integration rather than pretending one answer fits everything.

7. Regex Complexity Limits

Rampart imposes limits on regex patterns used for response matching to prevent ReDoS.

Current limits:

  • Maximum pattern length: 500 characters
  • Nested quantifiers: Rejected at load time (patterns like (a+)*)
  • Execution timeout: 100ms per regex match
  • Response cap: 1MB maximum for response-side evaluation; an oversized response fails closed when an applicable response rule exists

These limits protect against both accidental performance degradation and malicious patterns. They prevent policy authors from creating DoS conditions, and prevent attackers from injecting malicious regex patterns via webhook-driven policy updates. Patterns exceeding these limits are rejected at policy load time with clear error messages.

Glob matching is bounded separately: values used by glob conditions are limited to 8 KiB, ordinary patterns to 8 KiB, double-star patterns to 256 bytes, and each pattern to two ** occurrences. Oversized match inputs are denied as whole values; Rampart never checks a truncated prefix. The policy loader and linter reject patterns outside these limits.

call_count conditions retain at most 1,000 calls per tool, 1,024 active tool identities, and a 30-day window. Long-running proxy mode keeps this state in memory. One-shot native hooks share a locked state file at ~/.rampart/hook-call-counts.json, so thresholds apply across separate hook processes. Corrupt, unavailable, or capacity-exhausted state fails closed in enforce mode.

8. TLS on HTTP API

As of v0.7.4, rampart serve supports TLS via --tls-auto (self-signed ECDSA P-256) or --tls-cert/--tls-key (bring your own). On localhost, plaintext is still acceptable; for remote or team deployments, enable TLS.

Notes:

  • Default bind is 127.0.0.1 (localhost only). Use --addr 0.0.0.0 or another explicit interface only when you intend remote access.
  • --tls-auto generates a self-signed cert stored in ~/.rampart/tls/ (1-year validity)
  • The SHA-256 fingerprint is printed on startup for manual verification
  • For production, use proper certs via --tls-cert/--tls-key or a reverse proxy

9. Approval Persistence Limits

Pending approvals are now persisted to a local JSONL journal in normal rampart serve setups, so a routine service restart no longer necessarily wipes the queue. That said, approvals are still a live runtime workflow, not a durable transaction system.

Remaining limits:

  • Older or custom setups that disable persistence can still lose pending approvals on restart
  • A corrupted or deleted persistence file can drop pending approval state
  • An approval request that times out or restarts mid-flow can still surface to the agent as a denial/timeout

Mitigations:

  • Keep the default approval persistence path intact
  • Avoid unnecessary restarts during active approval flows
  • Treat approvals as short-lived human decisions, not long-running queued work

One-time (once: true) allow rules are claimed synchronously before an allow decision is returned. Rampart coordinates policy read-modify-write operations with a per-file cross-process lock and atomic replacement, including native hooks that run as separate processes. A failed claim is a denial; it is never allowed optimistically.

10. Project Policy Trust

Project-local .rampart/policy.yaml files are loaded automatically when present. A malicious repository could include a permissive project policy.

Mitigations (v0.6.9+):

  • Project policies cannot weaken global policies or restrictive global defaults; repository webhook actions are rejected
  • Set RAMPART_NO_PROJECT_POLICY=1 to skip project policy loading in untrusted repos
  • Project policy denials are prefixed with [Project Policy] for visibility

11. Community Policy Supply Chain

rampart policy fetch downloads policies from the registry with SHA-256 verification. However, the registry itself is hosted in the main repo — a compromise of the repository could introduce malicious policies.

Mitigations:

  • SHA-256 verification prevents modification after registry publication
  • --dry-run flag allows inspection before installation
  • Policy linting (rampart policy lint) validates syntax and flags suspicious patterns

Integration-Specific Notes

IntegrationExec CoverageFile CoverageResponse ScanningCascade
Native hooks (Claude Code)✅ (via hooks)✅ PostToolUse
Native hooks (Codex CLI/IDE/desktop)✅ (via hooks)✅ PostToolUse
Native hooks (GitHub Copilot CLI/VS Code)✅ (via hooks)✅ PostToolUse
Antigravity CLI/IDE plugin✅ (via PreToolUse)❌ host omits result
Native hooks (Cline)✅ (via hooks)
rampart wrap✅ cooperative shell calls
rampart preload✅ LD_PRELOAD
rampart protect openclaw
rampart setup openclaw --patch-tools✅ (shim)✅ (patched)
rampart setup codex✅ (native hooks)✅ (native hooks)✅ PostToolUse
HTTP proxy
MCP proxy

Platform Notes: macOS

v0.4.4 added 17 macOS-specific built-in policies to the standard and paranoid profiles. These cover:

  • Keychain access — blocks unauthorized reads from the macOS Keychain (security tool abuse)
  • Gatekeeper bypass — blocks attempts to disable or circumvent Gatekeeper (spctl, xattr -d com.apple.quarantine)
  • Persistence mechanisms — blocks writes to ~/Library/LaunchAgents/, ~/Library/LaunchDaemons/, and login items
  • User management — blocks dscl and sysadminctl commands that create or elevate user accounts
  • AppleScript shell execution — blocks osascript -e "do shell script …" patterns used to run commands via AppleScript

These policies are active automatically when using the standard or paranoid profile on macOS.

Platform Notes: Windows

v0.6.6 added Windows policy parity. Key differences from Linux/macOS:

  • No LD_PRELOAD or wrap moderampart preload and rampart wrap are not available. Use native hooks, the HTTP API, or the MCP proxy instead.
  • No POSIX file permissionschmod 0600 is not enforced by the OS. Rampart protects its persisted token with an owner-only Windows DACL derived from the current process SID; other sensitive files need explicit Windows ACL hardening.
  • Binary upgrade — in-process self-upgrade is intentionally disabled on Windows. Rerun the checksum-verifying PowerShell installer to replace rampart.exe; the Windows binary and installer are not currently code-signed. rampart upgrade --no-binary remains available for policy-only refreshes.
  • Path separators — Rampart normalizes backslashes to forward slashes internally for consistent policy matching.
  • Service management — automatic service installation is currently supported on Linux and macOS only. On Windows, run rampart serve directly or configure Task Scheduler/NSSM.

Deployment Recommendations

SetupAgent reads audit?Agent modifies policy?Best for
Same user (default)✅ Yes✅ YesDevelopment, testing
Separate user❌ No❌ NoProduction, unsupervised agents
Separate user + SIEM❌ No❌ NoEnterprise, compliance

Prerequisite: The agent must run as a non-root user. If the agent runs as root, user separation provides no protection — root can read and modify all files regardless of ownership.

Sudo caveat: Many real-world deployments grant the agent user sudo access for system administration tasks. An agent with unrestricted sudo (e.g., NOPASSWD: ALL) can bypass user separation by running sudo cat /etc/rampart/policy.yaml or sudo rm -rf /var/lib/rampart/audit/. Rampart still catches the common case — a hallucinating or prompt-injected agent won't think to sudo around a deny rule — but it's not a hard boundary.

Best practice: Restrict sudo to the specific commands your agent needs (e.g., apt, systemctl, k3s) rather than granting blanket access. This limits the blast radius regardless of Rampart.

12. API Self-Approval

Rampart now supports per-agent tokens with explicit scopes. Eval-only tokens can submit tool calls but cannot approve requests, reload policy, or mutate rules. That closes one big part of the old self-approval story.

The remaining risk is narrower but still real: in same-user deployments, any integration that exposes a readable admin-capable token to the agent process can still let that agent approve or mutate its own policy state by calling administrative endpoints directly.

Where this still matters most:

  • rampart wrap
  • rampart preload
  • ad hoc HTTP clients using the shared admin token from ~/.rampart/token

Mitigations:

  • Use user separation so the agent cannot read the admin token
  • Use per-agent eval-only tokens for HTTP/MCP clients whenever possible
  • Prefer native hook/plugin integrations where the agent is not handed a general-purpose admin bearer token

Current status: Better than before, not magically solved. Scoped per-agent tokens reduce the blast radius, but same-user deployments with readable admin tokens are still not a hard security boundary.

13. Temporal Allow Expiry

v0.7.4 introduced temporal allows (--for, --once). Expired rules are skipped during evaluation, so an expired entry cannot authorize a call.

Security implications:

  • Expired rules exist in the YAML but are inert — the engine checks expires_at before matching
  • Enforcement paths claim and remove a matching --once rule synchronously, under both an in-process lock and a cross-process policy-file lock, before returning an allow decision
  • If the one-time rule cannot be durably claimed or the policy cannot be refreshed, enforcement denies the call instead of allowing it
  • Expired rules may remain in policy YAML until a cleanup path removes them; this is a maintenance concern, not authorization
  • Clock skew: expiry is evaluated against the system clock. If the system clock is set backwards, an expired rule could become active again. Use NTP.

Self-Modification Protection

Rampart protects its own configuration from agent tampering through two layers:

  1. Exec-level: The standard policy blocks rampart allow, rampart block, rampart init, and shell redirects to .rampart/ directories. This prevents agents from running CLI commands that modify policy.
  2. Write/Edit-level: The standard policy blocks write and edit tool calls targeting **/.rampart/**. This prevents agents from directly overwriting policy files, config, or audit logs via file tools.

Both layers are active by default in the standard and paranoid profiles. The yolo profile disables these protections — it's named that way for a reason.

Remaining surface: An agent with exec access could modify Rampart's binary on disk (if file permissions allow), or kill the rampart serve process (triggering fail-open). User separation mitigates both: run rampart serve as a different user than the agent.

Philosophy

Rampart is a seatbelt, not a roll cage. It catches the vast majority of dangerous situations an AI agent will encounter — accidental or manipulated. It doesn't claim to stop every possible attack vector, and we're honest about what falls outside its scope.

If you need full isolation, use a sandbox (container, VM, or a tool like nono). Rampart and sandboxes are complementary — use both for defense in depth.


Reporting Security Issues

If you've found a vulnerability not covered here, please email rampartsec@pm.me. We'll acknowledge within 48 hours and work with you on coordinated disclosure. Please do not open public issues for security vulnerabilities.