dsh-reset-handoff
August 30, 2026 · View on GitHub
DSH never restarts itself. A host plugin that hands reset requests to an external ops agent over a versioned JSON protocol — preflight-snapshot → restart → health-check → recover — then delivers the result back to the requesting session after reboot.
Why
A long-lived DeepSeek Harness (DSH) instance needs to restart for many reasons: reload plugins/config, apply settings, recover from a wedged state. But the agent inside DSH should not restart DSH itself:
- restarting kills the very process that issued it, so the agent has no chance to see the outcome;
- the agent cannot see what business is running (other live sessions, the relay channels, pending jobs);
- if the reboot fails, nobody is left to diagnose and recover.
The safe pattern is a handoff: DSH writes a request, a separate, independent ops agent (here: Hermes Agent) reads it, runs the restart with preflight/health/recovery, and writes a result that DSH reads back after it comes up.
How it works
[DSH] agent calls reset_handoff(reason)
│ writes request.json (JSON protocol)
│ (optional) triggers the external executor
▼
[ext] ops agent reads request.json
│ 1. preflight — snapshot live sessions, relay state, pending jobs
│ 2. GATE — pre-restart maturity gate (see below)
│ 3. restart — restart the dsh web service (macOS launchd)
│ 4. health — poll http://127.0.0.1:3080 until 200 (with timeout)
│ 5. recover — verify relay/auth-proxy self-heal, list interrupted sessions
│ 6. result — write result.json (status done/failed + per-stage detail)
▼
[DSH] after reboot, the plugin reads result.json and delivers a readable
summary back into the requesting session (followup), so the agent
that asked can resume its interrupted work.
Pre-restart maturity gate
The reference executor refuses to restart unless it is safe to do so. It checks:
- No pending approvals/questions — relay
pending.jsonhas an emptypendinglist (a restart would otherwise drop the approval stack). - Enough free disk — at least
MIN_FREE_DISK_MB(default 500 MB). - Cooldown — at least
RESTART_COOLDOWN_SEC(default 60 s) since the previous restart, to break crash loops.
If any condition fails, the executor writes result.json with status: "failed", restart: { ok: false, gated: true }, and a gate array listing each failed condition with its detail — and does not restart. The requesting agent (or user) sees the exact reason and decides when it is safe to retry.
Tools
| Tool | Purpose |
|---|---|
reset_handoff(reason, scope?) | Submit a reset request to the external ops agent. Never restarts DSH in-process. |
reset_status() | Query the latest request and its result (read-only). |
Both tools are registered host-wide, so every session's agent can call them when a reset is needed.
Protocol (v1)
The plugin and the executor are decoupled — they only share two JSON files under ~/.dsh/reset-handoff/ (override with DSH_RESET_HANDOFF_DIR):
request.json (written by DSH):
{
"schema": "dsh-reset-handoff/request",
"version": 1,
"id": "<uuid>",
"requestedAt": "2026-08-30T12:00:00+08:00",
"reason": "重新加载插件配置",
"sessionId": "<requesting session id>",
"requester": "dsh-reset-handoff",
"scope": { "restartDshWeb": true, "healthCheck": true, "recoverInterrupted": true }
}
result.json (written by the executor):
{
"schema": "dsh-reset-handoff/result",
"version": 1,
"requestId": "<uuid>",
"status": "done",
"startedAt": "...",
"finishedAt": "...",
"preflight": { "liveSessions": ["..."], "relay": { }, "hermesJobs": ["..."] },
"restart": { "ok": true },
"health": { "ok": true, "checks": [ { "name": "dsh-web http :3080", "ok": true, "detail": "200" } ] },
"recovery": { "resumed": ["..."], "report": "..." },
"recoveryAction": {
"ok": true,
"attempts": [ { "attempt": 1, "restart": { "ok": true }, "time": "..." } ],
"diag": { "keyErrors": [], "logTail": "..." }
},
"gate": [ { "name": "relay 无待审批/待回答诉求", "ok": true } ]
}
Executor recovery contract (the part that makes "recover DSH itself" real): if the health check fails after restart, the executor must attempt recovery, not just report failure:
- Diagnose — read the dsh web error log tail and extract key errors (loader failures, missing deps like undici,
EADDRINUSEmulti-instance, OOM). - Retry — restart up to
MAX_RESTART_ATTEMPTS(default 3) times with a cooldown between attempts. - Observe — after each restart, wait an initialization window (default 120 s) before judging success.
- Report — write the outcome in
recoveryAction(attempts+diag), so the requesting agent and the human see why it failed and how many tries were made.
Any executor that reads/writes these two files can drive the reset — Hermes, a custom script, a cloud function. The protocol is the contract.
Install
dsh plugin --profile <profile> add github:<owner>/dsh-reset-handoff
Optional executor trigger: configure triggerCommand so reset_handoff also wakes the external agent (default: none — the executor may poll request.json instead). See cordis.patch.yml for the config shape.
Executor (Hermes example)
A reference executor is included under hermes/reset_agent.py (pure Python, no deps). It is meant to live inside a Hermes profile (reset-agent) and be triggered by hermes cron run <job>:
python3 reset_agent.py # run the five-step flow
python3 reset_agent.py --dry-run # print the flow, don't restart
Requirements
- DeepSeek Harness with the web profile (host plugin).
- The external executor must be able to restart the dsh web service (macOS
launchctl kickstart -k com.dsh.web, or equivalent for your OS/init).
Ops guardrails (read this before restarting anything)
Learned the hard way from a real 7-hour restart loop (2026-08-30). These rules are mandatory for any agent that manages a DSH host:
- Never create suicide/unconditional restart jobs. No
launchctl submitjobs containingkickstart -k, nokill -9on the DSH port, no unconditional restart logic. Restart only via thereset_handofftool (which goes through the executor's gate) or DSH's own mechanism. - Verify plugin dependencies before restart. The DSH loader resolves from
~/.dsh/profiles/web/node_modules— a missing transitive dep (e.g. undici) makes the whole plugin tree fail to load. Confirm deps exist and the tree loads cleanly before restarting. - Health-check first, observe after. Before any restart:
curlthe port, check for single instance (lsof -i :3080). After restart: wait a 2-minute observation window and confirm the PID is stable before proceeding. - Make plugins degrade gracefully. Missing config / bad fields should fall back to defaults with a friendly error, so calling agents never feel the need to edit plugin source or kill services to work around bugs.
License
MIT