dsh-safe-restart
A macOS-only DSH host plugin that safely restarts the DSH process from inside
an Agent session. It terminates only the exact DSH listener process through a
detached helper, lets the browser reconnect through DSH's native
ConnectionController (no manual refresh), resumes the original persisted
session, and starts exactly one automatic continuation turn to verify the
restart and continue the interrupted task.
The plugin registers a single Agent tool, restart_dsh. There is no Web
button and no settings UI.
What it does
- One Agent tool:
restart_dsh— schedules a safe restart of the current DSH host, bound to the invoking Agent's current session. - macOS only — every other platform is rejected before any side effect.
- Verified on DSH
0.1.0-rc.8— the recovery coordinator relies on host internals that are still moving in the rc cycle (agent/created,ctx.agents.resumesetup,installModelSelection). End-to-end restart + auto-continuation was verified on0.1.0-rc.8(standalone profile, macOS). The declared range is>=0.1.0-rc.8 <0.2.0; anything outside it is untested and may need adapter work. After any DSH upgrade, re-verify:pnpm test, then one realrestart_dshround — confirm PID change, HTTP 200, journalconsumed, and a continuation turn whoseturn/endreason iscompleted(noterror). - Detached helper — the journal write is durable first, then a helper process is spawned detached and unref'd; it survives host termination, replaces the DSH process, and records the outcome.
- Native browser reconnect — after the host comes back, the existing Web client reconnects through DSH's built-in connection/resync loop. No manual page refresh and no "continue" message are needed.
- One automatic continuation turn — the same persisted session resumes with its preset and model selection, and exactly one plugin-sourced continuation message is delivered per restart.
- Interrupted streams are not resumed — the exact token stream that was running when the host dropped is not continued. DSH closes that turn as interrupted; recovery starts a new turn in the same session.
Tool schema
restart_dsh({
"continue_after_restart": true, // boolean, default true
"continuation": "optional instruction" // string, optional, max 2000 chars
})
| Field | Type | Default | Meaning |
|---|---|---|---|
continue_after_restart |
boolean | true |
Whether to deliver one automatic continuation turn after recovery. Set false for a pure restart with no continuation. |
continuation |
string (optional) | — | Task-specific instruction for the single continuation turn (max 2000 chars). When omitted, the default instruction is: "Verify that DSH recovered successfully, then continue the pre-restart task using the current session and durable workspace state." |
The target session is always the invoking Agent's current session; a caller cannot name another session. The tool returns:
{ "restartId": "<uuid>", "scheduled": true, "continueAfterRestart": true }
Installation
Install from the plugin repository root:
dsh plugin --profile standalone add /path/to/dsh-safe-restart
The plugin activates on the next DSH start — restart_dsh is not loaded
until then. Use an already-validated external restart path (e.g. a codex /
safe-restart script that stops and relaunches DSH and verifies HTTP 200) for
that one activation restart, then verify http://127.0.0.1:<port> returns
HTTP 200 before continuing.
Usage
Ask the Agent to call restart_dsh. With the default
continue_after_restart: true, what the user sees is:
- The host connection drops briefly (the tool response may be cut short).
- The browser reconnects automatically through DSH's native
ConnectionController— no manual refresh. - The original persisted session reappears with its preset and model.
- Exactly one new assistant turn appears, carrying the restart ID marker
[restartId: <uuid>], verifying recovery and continuing the task.
A restart with "continue_after_restart": false performs the same safe
restart and browser recovery but adds no continuation turn.
Safety
- Never
killall— the helper terminates only the recorded PID, and only after verifying its exact command line against the launch identity recorded at schedule time (the current Node executable plus the DSH entry argument). An unrelated or unknown process is never touched. - One active restart — a second
restart_dshwhile a record is in flight is rejected (restart already active). - 60-second cooldown — a restart within 60 s of the most recent one is
rejected (
restart in cooldown). - 30-second health timeout — after spawning the replacement, the helper polls the health endpoint for at most 30 s (PID exit and port release use a 12 s budget).
- No retry loop — each
restartIdis executed at most once; a record that is notscheduledis refused, and a failed restart is never retried. - Exactly-once delivery — before delivering the continuation, the
coordinator scans the session's durable log for the restart ID marker; a
previous delivery that crashed before the
consumedtransition is detected and consumed without re-delivering. An ambiguous window becomesdelivery-uncertaininstead of a blind retry. - No secrets persisted — the journal and log carry no credentials and no environment values.
- Interrupted token stream is not resumed — the continuation is a fresh turn in the same session, never a replay of the interrupted generation.
Storage
All durable state lives under ~/.dsh/storages/dsh-safe-restart/:
| Path | Purpose |
|---|---|
active.json |
The single in-flight restart record (atomic writes: temp file → fsync → rename). |
.lock |
Exclusive restart lock, created with open(..., 'wx'); one active restart at a time. |
history/<restartId>.json |
Terminal records, archived when a record reaches a terminal phase. |
restart.log |
Timestamped, non-secret lifecycle lines: restartId, old/new PID, port-release result, HTTP result, terminal phase. |
Phases
| Phase | Meaning |
|---|---|
scheduled |
The tool request was persisted; the detached helper has not claimed it yet. |
restarting |
The helper validated and terminated the old PID, spawned the replacement, and is waiting for health. |
ready |
The new host answered HTTP 200; newPid and readyAt are set. Waiting for the resumed session's Agent. |
delivering |
The continuation message was created and handed to the Agent; awaiting durable confirmation. |
consumed |
Terminal: exactly one continuation delivered (or, for continueAfterRestart=false, none) and archived. |
failed |
Terminal: the restart failed with a bounded, non-secret error message; never retried. |
delivery-uncertain |
Terminal: the continuation window was ambiguous (e.g. the host died between delivering and confirmation). Not blindly redelivered. |
Troubleshooting
failedrecord — checkrestart.logand the record'serrorfield (e.g. PID validation failure, port not released, health timeout). The record is already archived; no action is required for future restarts.delivery-uncertainrecord — the continuation may or may not have been delivered. Check the session's durable log for the marker[restartId: <uuid>]:- marker present → the message was delivered; the record is archived, done.
- marker absent → it was not delivered and will not be redelivered; decide whether to re-run the task manually.
Stale lockless
scheduledrecord — if the helper crashed before it could claim the record (or the claim failed), ascheduledrecord with no helper running and no lock can be left behind. Recovery polls it for a bounded window and then deliberately leaves it for inspection. It blocks new restarts (restart already active).Ready but no Agent (blocks subsequent restarts) — if a record reaches
readybut the matching session's Agent never reappears (for example the session was deleted or the browser tab was never reopened),active.jsonstays inreadyand every laterrestart_dshis rejected until the record is resolved.Remedy (manual, after inspection):
- Inspect
~/.dsh/storages/dsh-safe-restart/active.json— note itsrestartIdandphase— and the latest lines of~/.dsh/storages/dsh-safe-restart/restart.logto see what happened. - If the continuation marker
[restartId: <restartId>]already exists in the session's durable log, the continuation was delivered; archive the record asconsumed: moveactive.jsontohistory/<restartId>.jsonwith"phase": "consumed"and removeactive.json. - If the continuation was never delivered and never will be, archive the
record with a terminal phase instead (
"failed"or"delivery-uncertain"), then removeactive.json. - Check
~/.dsh/storages/dsh-safe-restart/.lock— it records the owningrestartId. Remove the lock only when itsrestartIdmatches the archived record; never delete a lock owned by a different restart. - Leave
history/<restartId>.jsonandrestart.login place — they are the audit trail. The nextrestart_dshis then allowed again.
- Inspect
Verification script
scripts/verify-real-restart.mjs is an external, read-only evidence collector
that runs outside the DSH process (plain Node ESM, zero dependencies). It
never performs a restart and never touches the live listener. It reads the
journal and probes the HTTP endpoint, then reports one JSON object:
{
"oldPid": 111, // replaced DSH PID, from the selected record
"newPid": 222, // healthy replacement PID (null when unknown)
"pidChanged": true, // newPid is set and differs from oldPid
"http200": true, // live GET http://127.0.0.1:<port>/ returned 200
"phase": "consumed", // lifecycle phase of the selected record
"continuationCount": 1 // consumed records that delivered a continuation
}
node scripts/verify-real-restart.mjs --check-latest # latest history record, default port 3081
node scripts/verify-real-restart.mjs --journal /path/to/journal --port 3000
node scripts/verify-real-restart.mjs --journal /tmp/fixture --port 0 --dry-run # detection only, no HTTP
--dry-run prints the detection logic (listener PID via lsof -ti tcp:<port>
and a journal contents summary) without probing HTTP, so it can be exercised
without a live restart. It exits nonzero on malformed journal JSON, an
unusable listener detection, or unexpected errors. continuationCount counts
consumed records whose continueAfterRestart is true — per the design each
such record delivered exactly one continuation turn; pure-restart records
(false) delivered none and are not counted.
Uninstall
Remove the plugin with the remove counterpart of the install command, e.g.:
dsh plugin --profile standalone remove dsh-safe-restart
then restart DSH once so the tool is dropped. If you also want to remove the
durable state, delete ~/.dsh/storages/dsh-safe-restart/ after confirming no
restart is in flight (no active.json).
Recovery limitation
If the browser tab was closed during the restart, simply reopen it: DSH's native reconnect restores the same persisted session (same preset, model, and history). The continuation is delivered exactly once regardless of how many times the tab reconnects.
No comments yet. Be the first to write one.