README.md
July 31, 2026 · View on GitHub
Production-grade agents.
Secure. Sandboxed. Self-improving.
A TypeScript toolkit for shipping AI agents that are multi-tenant-safe, without complex infrastructure. Curated memory, human-gated learning, and a real execution sandbox.
Why · How it works · Try it · Install · Set up · Wire into your app · Security · Docs
Example
Why
Most agent harnesses are built for local use: one developer, one machine, one trust boundary. Shipping agents into production is different. One mistake can delete files, leak secrets, or cross tenant boundaries.
You stay in control of auth, tenancy, and the tools the agent can reach. Your users get a real agent that can run code and improve over time, inside an isolated home.
No huge cloud bill. No fancy infrastructure.
| Pillar | What you get |
|---|---|
| Secure by default | Prompt-injection, promptware, and exfiltration scanning on every memory and skill write. Threats never reach the system prompt. |
| Sandboxed execution | Per-tenant AgentFS volumes and bash-tool guardrails. Destructive commands, secret exfil, and non-allowlisted network egress are blocked before they run. |
| Production multi-tenancy | One isolated filesystem, memory, skill library, transcript store, and audit trail per tenant. A bug in tenant A cannot touch tenant B. |
| Self-improving under approval | A background curator distills sessions into durable memory and reusable skills. Writes stage for human review by default. Hosts can set curator.autoApprove when end users are not the right reviewers. |
Learning without a sandbox is a liability. A sandbox without learning is just a cage. agent-kit is both.
It is a library, not a hosted service. You authenticate users, map each one
to a tenantId, and open a per-tenant home. The kit supplies the volume,
sandbox, session runtime, and learning loop.
Hosting shape today: one Node process and one local SQLite volume file per tenant. Multi-machine hosting is not ready yet (roadmap).
Built on
agent-kit composes existing libraries. The Vercel AI SDK
shapes most of the live API (ModelMessage, session.run / session.stream,
toolApproval, and AI SDK UI useChat).
| Layer | Library | What you feel in the API |
|---|---|---|
| Model loop | ai (Vercel AI SDK) | Messages, run / stream, tools, UI approval |
| Model providers | AI SDK providers or @ai-sdk/gateway | Pass a LanguageModel, or a string id via the Gateway |
| Tenant volume | AgentFS | One SQLite filesystem per tenant |
| Sandbox shell | bash-tool + just-bash | bash / readFile / writeFile behind guardrails |
If you already use the AI SDK, agent-kit slots in as the tenant home, memory, skills, and sandbox around that loop.
How it works
Install and wiring make more sense once you see the loop.
Write the agent once as files: identity (SOUL.md), rules (AGENTS.md), and
optional skills or memories. Then the session loop is:
- Open. Start a session for one tenant. Load or compile the agent so those files seed the tenant volume. The runtime builds the system prompt once from identity, rules, skills, and a frozen memory snapshot. Memory does not change mid-chat.
- Guard. Scan content before it can enter memory or skills. Block dangerous shell commands before they run.
- Curate. After each turn,
createTenantHomecan propose durable memory or skills in the background. Setconfig.curator: falseto turn this off. - Approve. Proposals stay under
pending/until a human accepts them, or apply immediately whenconfig.curator.autoApproveis true. - Recall. The next chat includes approved memory. Past chats for that tenant are searchable. Other tenants stay isolated.
Your app owns auth and tenantId. The kit owns isolation, scanning, the
sandbox, and the approval gate.
Try it
Clone an example, or jump to Install to wire the package into your own app.
| Example | What it shows |
|---|---|
examples/example-app | Streaming Next.js chat + /code-runner page with js-exec |
git clone https://github.com/socialrobot-io/agent-kit.git
cd agent-kit && bun install
cd examples/example-app
cp .env.sample .env.local # set DEEPSEEK_API_KEY or AI_GATEWAY_API_KEY
npx nx dev example # http://localhost:3000
# Code runner (js-exec): http://localhost:3000/code-runner
Install
Requirements
- Node.js 20+ (or Bun)
- Durable local disk for per-tenant SQLite volumes (one machine today)
- A model provider for live turns
Happy path package (volume + transcripts + sandbox + live loop):
npm i @socialrobot-io/agent-kit-node ai
# also pulls core, agent-kit-ai, sessions, sandbox, curator
# `ai` is a peer (Vercel AI SDK); install it next to the kit
| Package | Job |
|---|---|
…-node | createTenantHome (volume, sandbox, sessions, curator) |
…-core | Definition, memory, skills, approval |
…-ai | AgentSession.run / .stream |
…-sessions | Transcripts + session_search |
…-sandbox | Volume + guarded bash |
…-curator | Background review (wired by …-node) |
Package manager notes
-
Bun: Bun blocks lifecycle scripts by default. Trust
@mongodb-js/zstdso itsprebuild-installcan fetch a binary:{ "trustedDependencies": ["@mongodb-js/zstd"] }Then reinstall (
bun install), or runbun pm trust @mongodb-js/zstd. Do not trustnode-liblzmaunless your machine can compile it (pkg-config+ systemliblzma). A failed trusted install can delete the optional package under Bun. -
npm:
Cannot find native bindingoutside a bundler means a missing platform package (known npm bug npm/cli#4828). Removenode_modulesandpackage-lock.json, then install again. -
Next.js / Turbopack: native packages must stay outside the bundle. See Next.js (App Router).
Full notes: Getting started.
Model provider
Pass a ready LanguageModel from any AI SDK provider
(@ai-sdk/openai, @ai-sdk/anthropic, @ai-sdk/deepseek, …). That is the
usual path.
import { anthropic } from "@ai-sdk/anthropic";
const home = await createTenantHome({
tenantId: "brand-123",
agent,
model: anthropic("claude-sonnet-4-5"),
});
Or pass a "provider/model" string and set AI_GATEWAY_API_KEY so the
Vercel AI Gateway resolves it.
Set up
Three steps: author files, compile them into your app, run a turn.
1. Author the agent as files
The agent is a directory of markdown, not a large config object.
agent/
SOUL.md who the agent is (always in the system prompt)
AGENTS.md house rules
skills/ reusable how-to procedures (optional)
memories/ USER.md and MEMORY.md (optional; behind approval when learned)
Example SOUL.md:
You are a concise research assistant for a fintech startup.
Example AGENTS.md:
Prefer short, factual answers.
Cite a source for every non-obvious claim.
Never invent numbers.
Skills under agent/skills/ are mutable unless you mark them locked
(locked: true / pinned / bundled in frontmatter, or a .locked marker).
See Skills & learning.
2. Compile the agent, open a tenant home, run a turn
createTenantHome only installs identity and skills when you pass agent.
Compile agent/ in CI / predev into an importable module so Next, Docker, and
workers ship the content without a runtime agent/ directory on disk.
// scripts/compile-agent.mjs — wire into predev / prebuild
import { compileAgent } from "@socialrobot-io/agent-kit-node";
await compileAgent({
dir: "./agent",
outFile: "./src/generated/agent.ts",
});
import { createTenantHome } from "@socialrobot-io/agent-kit-node";
import { agent } from "./generated/agent"; // output of compileAgent
const tenantId = "brand-123"; // from your auth layer — never from the client body alone
const sessionId = "chat-abc";
// Default: ./data/tenants/${tenantId}.db + transcripts + sandbox.
// Pass model: a LanguageModel from any AI SDK provider (recommended).
const home = await createTenantHome({ tenantId, agent });
// Memory freezes when openSession returns. Reuse that AgentSession for the
// life of the chat (cache by sessionId in your process). Calling openSession
// again rebuilds the snapshot from disk.
const session = await home.openSession(sessionId);
const turn = await session.run([
{ role: "user", content: "Summarize /workspace; prefer short answers going forward." },
]);
Plain Node scripts that can read ./agent at runtime may use
loadAgent("./agent") instead of compile + import.
3. Override only what you need
import { anthropic } from "@ai-sdk/anthropic";
const home = await createTenantHome({
tenantId,
agent,
dataDir: "/var/lib/agents", // or volumePath: "/data/acme.db"
model: anthropic("claude-sonnet-4-5"), // or "provider/model" + AI_GATEWAY_API_KEY
interactiveApproval: true, // UI Approve applies writes
workspaceFiles: { "README.md": "# hi\n" },
sandbox: {
// Hostnames only (not full URLs). Or sandbox: false to disable.
allowedHosts: ["api.example.com"],
secrets: [process.env.TENANT_API_KEY!],
},
});
const session = await home.openSession(sessionId, {
addTools: [myTool],
disableTools: ["skill_manage"],
});
You now have a working turn. Next: put auth, session cache, and transcripts around it.
Wire into your app
Your app authenticates the user. The kit only trusts the tenantId you pass.
import type { AgentSession } from "@socialrobot-io/agent-kit-ai";
import { createTenantHome } from "@socialrobot-io/agent-kit-node";
import { agent } from "./generated/agent"; // from compileAgent in predev / CI
// Reuse the same AgentSession for a chat so memory stays frozen.
// Key includes tenantId so two tenants never share a session handle.
const sessions = new Map<string, AgentSession>();
async function handleTurn(opts: {
tenantId: string; // from your auth layer — never from the request body alone
sessionId: string; // one id per chat conversation
userText: string;
userMessageId: string;
}) {
// Opens (or reuses) volume + transcripts + sandbox for this tenant.
const home = await createTenantHome({ tenantId: opts.tenantId, agent });
const key = `${opts.tenantId}:${opts.sessionId}`;
let session = sessions.get(key);
if (!session) {
session = await home.openSession(opts.sessionId);
sessions.set(key, session);
}
// Persist both sides so session_search can browse past chats.
await home.transcripts!.createSession({
id: opts.sessionId,
tenantId: opts.tenantId,
source: "api",
createdAt: Date.now() / 1000,
});
await home.transcripts!.appendMessage({
id: opts.userMessageId,
sessionId: opts.sessionId,
role: "user",
content: opts.userText,
createdAt: Date.now() / 1000,
});
const turn = await session.run([{ role: "user", content: opts.userText }]);
await home.transcripts!.appendMessage({
id: `asst_${Date.now()}`,
sessionId: opts.sessionId,
role: "assistant",
content: turn.text || "(no text)",
createdAt: Date.now() / 1000,
});
return turn;
}
For streaming Next.js chat, copy examples/example-app.
On Next.js App Router, keep the native packages outside the bundle with
serverExternalPackages: Next.js (App Router).
More detail: Hosting.
Security
| Layer | Stops |
|---|---|
| Threat scanning | Injection and exfil patterns in memory/skills before they reach the prompt. Bad on-disk entries show as [BLOCKED]. |
| Write approval | Silent self-edits. Background and skill writes wait for a human. |
| Sandbox | Destructive shell, secret dumps, hosts you did not allow. |
| Tenant isolation | One volume and audit trail per tenant. Search never crosses tenants. |
Before you ship:
- Resolve
tenantIdonly from trusted auth. Use an opaque id safe for paths. - Pass
agentso company identity and skills are installed on the volume. - Lock company-owned skills; unlocked skills stay mutable behind approval.
- Pass sandbox
secretsand hostname-onlyallowedHostsat home creation. - Do not hand tools the raw volume write handle.
Details: Security guide.
Docs
Read in this order when you integrate:
| Guide | Answers |
|---|---|
| Getting started | Install, agent/ files, first turn |
| Hosting | Auth, volume, session, approve in your app |
| Security | Scans, approval, isolation |
| Tools | Host tools vs sandbox vs skills |
| Sandbox | Curl, js-exec, python3, custom bash cmds |
| Models | Pick a model, run or stream a turn |
| Memory | What is remembered across chats |
| Skills & learning | Skills, curator, human approve |
| Publishing | npm release (maintainers) |
Not ready yet: Multi-machine.
Commands (contributors)
bun install
npx nx run-many -t test --all
npx nx run-many -t build --all
Before a commit: npx nx run-many -t typecheck test build --all must be green.
License
MIT. See NOTICE for third-party attribution.
agent-kit: agents you can ship.