The Hive

June 4, 2026 · View on GitHub

How Munder Difflin turns a room full of independent claude processes into a collaborating, self-coordinating team with persistent memory, a shared blackboard, and a "god" orchestrator that runs the floor.

This document is the design source of truth for the agent-collaboration layer. It sits alongside SPEC.md (terminal/event plane) and DESIGN.md (visual system). Code is the source of truth for what's built; this is the source of truth for what we're building toward.


1. What we're building (and what it's called)

Each spawned agent is a real claude CLI process with a filesystem, a system prompt, and a hook lifecycle. We layer four classic patterns on top:

Behaviour the user asked forPattern (the name)
Per-agent memory file made at spawn, that the agent reads and updatesAgent long-term memory (MemGPT/Letta-style self-managed memory)
Writing a requirement into another agent's fileStigmergy — coordinating by modifying a shared environment
A shared plan multiple agents editBlackboard architecture (Hearsay-II)
"Check after finishing every task"Mailbox / actor model — drain an inbox at a lifecycle point
A "god" agent that runs the floor and clarifies for othersOrchestrator / supervisor (LangGraph-supervisor-style)

The umbrella term is a multi-agent system (MAS) with autonomous agent loops. The closest academic analogue to this app is Stanford's Generative Agents (Park et al., 2023): Sims-style avatars in a 2D world with a memory stream, retrieval, reflection, and planning.


2. Locked design decisions

  1. Git as the coordination/audit layer, single committer. Everything the hive knows is files in one local git repo. To avoid .git/index.lock corruption with many concurrent agents, only the Electron main process commits. Agents never call git — they write plain files. (Research: GitHub Desktop's commit-queue pattern; lazygit/git-retry backoff.)
  2. Single-writer-per-file. Each agent writes only inside its own agents/<id>/ directory. Cross-agent delivery happens by the router (main process) moving messages from a sender's outbox/ into a recipient's inbox/. No file is ever written by two processes.
  3. God-mode autonomy, native HITL. A privileged god agent (lives in Michael's room) adjudicates cross-agent traffic. Routine requests (clarifications, data asks, plan tweaks) it resolves itself and the system keeps running fully autonomously. Critical items (destructive ops, spend, scope changes, unresolvable conflicts) route to the god, who surfaces them to the human natively in his own Claude Code session — there is no separate approval queue. Tool-permission prompts are the HITL gate, and they're approvable remotely from a phone via /remote-control.
  4. Memory: markdown first. Per-agent memory.md + shared blackboard, with a SQLite FTS index when keyword recall isn't enough. A heavyweight vector layer (Letta/Mem0/Zep) is not needed at 5–15 agents and is architecturally wrong here (they want to own the agent runtime; our runtime is the claude CLI). Optional future upgrade: MemPalace over MCP (validate its retrieval first — its public benchmarks are overstated per independent audit).
  5. Autonomous loop = Stop hook. An agent that finishes drains its inbox via a Stop hook that returns {"decision":"block","reason":…} to keep it working, guarded by stop_hook_active to prevent infinite loops.

3. On-disk layout — the "hive"

Lives under <harnessHome>/hive/, a git repo committed only by the main process.

hive/
  PROTOCOL.md            # the agent-facing contract (how to remember + message)
  registry.json          # roster: every agent, role, capabilities, status, seat
  board.md               # shared blackboard / co-authored plans
  tasks.json             # task ledger (id, assignee, spec, status, result ref)
  log.jsonl              # append-only event feed (drives the UI activity stream)
  agents/<agentId>/
    identity.md          # who am I, my role, my capabilities  (read at start)
    memory.md            # my long-term memory  (I read at start, append as I learn)
    inbox/               # messages delivered TO me — <ts>-<msgid>.json
    inbox/.done/         # processed messages (kept for audit, not deleted)
    outbox/              # messages I want to SEND — router drains these
    cursor.json          # { lastProcessed: <msgid> }  — avoids reprocessing

Design rules that make this robust:

  • One JSON file per message, written via temp-file + atomic rename — never a co-edited shared mailbox file (those conflict under git).
  • Append-only log.jsonl; consumers track their own cursor.
  • board.md is the one genuinely co-edited file — it goes through the god agent (single scribe) to avoid conflicts.

4. Message schema (FIPA-lite)

Borrow the one useful idea from FIPA-ACL/KQML — the speech act — and drop the LISP syntax. Seven semantic fields:

{
  "id":            "2026-05-30T14-03-11-123Z-a1b2",  // unique, time-sortable
  "conversation":  "conv-7f3",                        // groups a thread
  "in_reply_to":   "<prev msgid> | null",
  "from":          "agent.researcher",
  "to":            "agent.coder | god | broadcast",
  "act":           "request | inform | propose | query | agree | refuse | done",
  "subject":       "short human-readable summary",
  "body":          "free text / markdown / structured payload",
  "hops":          3,            // ++ per reply; capped to kill ping-pong loops
  "requires_reply": true,        // only request/query/propose obligate a reply
  "needs_human":   false,        // router/god may flip this to escalate
  "created_at":    "ISO-8601"
}

Anti-livelock rules: only request/query/propose obligate a reply (pure inform/done are terminal); every reply increments hops; past a hop cap the god agent escalates instead of letting two agents loop forever; re-seeing a processed id is a no-op (idempotent via cursor).


5. Control flow

agent B mid-task needs something from agent C
        │ writes  agents/B/outbox/<msg>.json   (act:request, to:C)

┌─────────────────────── main process (the harness) ───────────────────────┐
│  Router watches every outbox/                                             │
│    → deliver to agents/C/inbox/   (to:"human" → routed to the god proxy;  │
│       the god surfaces critical calls natively in its own session)        │
│    → append to log.jsonl → git commit (single committer, retry+backoff)   │
└──────────────────────────────────────────────────────────────────────────┘
        │ delivered to C's inbox

agent C finishes its current turn → Stop hook fires
        │ hook POSTs to the hive socket; main process checks C's inbox
        │ unread messages?  → reply {"decision":"block","reason": <messages>}

agent C keeps working: reads the messages, acts, replies via its own outbox

The same hook socket drives the avatars: PreToolUse/PostToolUse payloads move an agent to the right station (replacing today's mockEvents.ts / PTY-scraping).


6. The god agent (orchestrator)

A fixed, always-on agent seated at desk-ceo (Michael's room), character: michael, flagged isGod. It is an ordinary claude process — the intelligence — while the main process is the mechanism (git, sockets, routing). It owns:

  • Roster & routing (registry.json): who exists, their capabilities, status.
  • Adjudication: read each outbound request; resolve routine ones itself (answer clarifications, route to the right specialist with a self-contained task spec), escalate only critical ones. This is "god mode."
  • Blackboard scribe: the single writer of board.md, so shared plans never conflict.
  • Task ledger (tasks.json): assign, track, retry, checkpoint.

Its escalation policy (what counts as "critical") lives in its system prompt and is the primary control surface — tune the prompt, not the code.


7. Phased plan

  • Phase 0 — Foundation ✅: hive.ts on-disk layer + spawn injection (identity, protocol, env) + IPC to read hive state. Agents are hive-aware: they read their memory/inbox at task start and send via outbox; the router delivers; everything is committed and visible.
  • Phase 1 — Autonomy ✅: hooks.ts UDS server + cth-hook shim (attached per agent via --settings) + Stop-loop so agents drain their inbox automatically and keep running (guarded by stop_hook_active + cursor); hook events stream to the renderer to drive avatars.
  • Phase 2 — God mode ✅: the god agent auto-spawns into Michael's room (desk-ceo reserved) and, on a fresh spawn, is started with /remote-control (best-effort) plus an orientation prompt so it begins running the floor on its own. The router routes to:"human" traffic to the god (the human's proxy); there is no separate approval queue — human-in-the-loop is native to each agent's Claude Code session (permission prompts, approvable remotely from a phone). Idle agents are woken when they hold unread inbox messages.
  • Phase 3 — Semantic memory ✅ (CLI integration): memory.ts wraps the MemPalace CLI (not MCP, by decision). The harness keeps one shared palace under harnessHome, points every agent's MEMPALACE_PALACE_PATH at it, mines each agent's memory.md into its own wing (mtime-gated), and agents recall via mempalace search / wake-up. Detect-and-degrade: a no-op when mempalace isn't installed (markdown memory still works). Default model minilm (light, for low-RAM Macs); embeddinggemma is the multilingual opt-in. A MemoryPanel lets the human search the same palace.
    • Still open: reflection/summarization to bound memory.md; needs a live mempalace install to validate retrieval end-to-end.

8. Key risks & mitigations

RiskMitigation
index.lock corruptionSingle committer (main process), retry+backoff, stale-lock cleanup
Infinite Stop-hook loopGuard on stop_hook_active; hops cap; CLAUDE_CODE_STOP_HOOK_BLOCK_CAP
Two agents ping-pongingOnly request/query/propose obligate replies; hop cap → god escalates
Reprocessing messagesPer-agent cursor.json; processed messages move to inbox/.done/
memory.md unbounded growthPhase 3 reflection/summarization
Modifying the user's repo with hooksWrite hooks to <cwd>/.claude/settings.local.json (gitignored convention)

9. References

  • Anthropic — Building a multi-agent research system (lead/subagent, plan-to-memory).
  • LangGraph supervisor (structured routing + handoff registry + checkpoints).
  • FIPA-ACL / KQML (speech acts).
  • Stanford Generative Agents (memory stream, reflection, 2D world).
  • Claude Code hooks reference (Stop, PreToolUse, UserPromptSubmit; stop_hook_active).