dsh-verified-progress
A verified-progress ledger for long-running DSH work — where a verdict only counts if the thing that produced it provably stayed read-only.
A long-running task fails in a specific way: the agent reports progress it cannot prove, the report enters the transcript, and the next turn builds on a claim that was never checked. Nothing downstream can tell a verified step from an optimistic one, so the error compounds until the final answer is confidently wrong.
This plugin adds the missing check. When the agent claims a step is done, an independent verifier with a fresh context inspects the real workspace and returns a structured verdict. Only a clean, aligned, complete verdict enters the ledger. A rejected step stays in the ledger as evidence — it is never promoted to progress.
claim ──▶ independent verification ──▶ verdict
│
complete + clean + aligned ──▶ ledger (progress)
│
anything else ──▶ ledger (evidence only)
The part that is not a prompt
Read-only verifiers are usually enforced by asking the verifier to behave and by removing the write tools from its tool set. Both are prevention, and both are defeated by anything the tool filter does not name.
This plugin detects instead. The workspace is fingerprinted — every file's
size and mtime, plus a SHA-256 for files under the hash limit — immediately
before and after each verification episode. If anything changed, the verifier
mutated the workspace during a read-only audit, so its own findings are voided by
construction: the round is recorded as blocked / violation and can never be
promoted to progress, no matter how clean its report reads.
That matters because a verifier that writes is not a neutral witness. It can create the evidence it then reports, and no amount of prompt discipline or tool allow-listing proves it did not. A fingerprint comparison does not have to trust the verifier at all — which is why the ledger's admission rule can be absolute:
integrity === "violation" || contract !== "aligned" ⟹ status !== "complete"
This is enforced in code, not requested in a prompt. A verifier that omits a
control line degrades conservatively — an unstated integrity is read as
suspect, never as clean — because a placeholder must never be the thing that
certifies work as done.
What it does
1. Independent verification instead of self-report. The verifier is a separate subagent that inherits no conversation context. It reads the files and answers with three control lines:
Status: complete | incomplete | blocked
Integrity: clean | suspect | violation
Contract audit: aligned | unknown | needs_revision | invalid
2. A durable verified-progress ledger. Accepted steps are appended to
ledger.jsonl (one JSON object per line, fsync'd per append) with a last-wins
projection in state.json. A crashed process, a compacted context, or a fresh
session can reopen the run and answer: what is verified, what was rejected and
why, and what is left. Rejected rounds are preserved as evidence rather than
discarded, so a later turn can see which claims were already tried and failed.
3. Mutation detection. Described above; it is the admission rule for the ledger, not a side feature.
Tools
| Tool | Purpose |
|---|---|
verified_progress_verify |
Verify one claimed step against the workspace; records the verdict in the ledger. |
verified_progress_ledger |
Read back verified progress, rejected claims, and what remains. |
verified_progress_state |
Current run state: verified rounds, pending claims, open items. |
Declaring what the guard may judge
verified_progress_verify takes a guard_scope: a workspace-relative path or
glob whose non-glob prefix becomes the fingerprint root.
verified_progress_verify claim="the parser handles CRLF input"
guard_scope="src/parser/**"
This is not optional decoration. The guard originally fingerprinted the whole
workspace root, and in the desktop profile that root is the harness home itself —
sessions/**, storages/** and plugin state are written continuously by the
very session the plugin runs inside. The first two verifications came back
workspace mutated (+3 ~5 -3), and every changed path belonged to the harness.
The verifier had written nothing. A guard that accuses the verifier of the
host's writes fails every round while looking like a conservative verdict, so
the scope is now declared rather than assumed:
- the scope must resolve inside the workspace; absolute paths and
..are rejected instead of silently widened; - when the workspace is the harness home, a scope is required, and
harness-owned roots (
sessions,storages,logs,telemetry,cache,profiles, …) are excluded only at the root, so a real project containing a nestedsessions/is still guarded; - the result prints
Fingerprinted: <path>, so a verdict never leaves its scope ambiguous.
Install
dsh plugin --profile web add dsh-verified-progress
Or install a released tarball, which needs no registry:
dsh plugin --profile web add \
https://github.com/lion231226/dsh-verified-progress/releases/download/v0.3.1/dsh-verified-progress.tgz
Then restart the harness: a plugin that failed to activate is not retried, and the module cache holds the previous version until the process restarts. No build step — the package ships plain ESM and has no runtime dependencies.
Storage
Everything lives under the harness state directory, not in your project:
<state>/verified-progress/runs/<runId>/ledger.jsonl
<state>/verified-progress/runs/<runId>/state.json
<runId> is derived from the session id when there is one, otherwise from a hash
of the working directory, so reopening the same session reopens the same run.
Design limits — read these
- The verifier is a separate agent, not a separate process. It cannot see the executor's conversation, which is what makes it independent, and it runs with a read-only tool set. It runs inside the harness and is subject to the same permissions; the fingerprint check is what makes a violation detectable rather than impossible.
- The mutation guard is a correctness tool, not a security boundary. It compares file fingerprints. It does not defend against a local attacker, it is not a sandbox, and it does not cover paths excluded from the snapshot.
- Unhashed large files are reported as an evidence gap. Files above the hash limit are compared by size and mtime, and the snapshot says so rather than claiming a comparison it did not make.
- Verification costs tokens. Each verified step adds one verifier episode. The ledger is what you get for it; if your task is a single short turn, you do not need this plugin.
Why this exists
The loop design is ported from LongHorizon-Harness (AMAP-ML, MIT) — manager/executor/auditor rounds, a durable round ledger, and an auditor that cannot certify a dirty result.
That project is a Python orchestrator that wraps an agent CLI. It does not run on
Windows: its persistent layer requires os.O_NOFOLLOW, os.O_DIRECTORY and
supports_dir_fd (all absent on win32) and it shells out with POSIX
VAR=value cmd templates. It also, by its own design, drives one agent CLI per
role episode — so on DeepSeek Harness it can only read the final answer of each
dsh --profile headless run, not the intermediate tool events.
This plugin takes the parts that carry the value — independent verification, the downgrade invariant, the verified-state ledger, the mutation guard — and implements them natively in the harness, where the events are available and no POSIX-only primitive is needed. Its workspace-mutation detection is its own contribution; the upstream project guards against verifier writes only through the same prompt-and-allow-list prevention described above.
Verifying this plugin
npm test # 37 tests: verdict grammar, guard, ledger, tools, extraction
npm run verify # acceptance probe against the packed artifact
npm run verify is not a unit test run. It rebuilds the package from the
files list in package.json, loads the tool set out of that artifact, and
runs three scenarios against a real file system with a scripted verifier:
| Scenario | What it must show |
|---|---|
clean-verdict-is-progress |
a clean verdict becomes progress, in one verifier episode |
positive-control-mutation-voids-verdict |
a verifier that edits a file is caught; the round is blocked / violation and never becomes progress |
ledger-survives-a-fresh-context |
a brand-new context reads verified progress back off disk |
The middle scenario is a positive control: if the fingerprint guard ever
stops noticing a mutation, positiveControlTripped turns false and the probe
exits non-zero, so a broken detector cannot be reported as a pass. The probe also
proves it ran — it records the artifact it loaded, the tool names it found, and
that all three scenarios reached a deciding assertion.
The result of the last run is committed: verify/acceptance-report.json
and the raw tool results and ledger contents in
verify/acceptance-evidence.json.
Requirements and known environment issues
A subagent provider that does not inherit the parent context. The plugin injects
toolsandsubagents; thespawnprovider shipped with the web and headless profiles satisfies this. If only theforkprovider is available, the plugin refuses to certify anything and says so in the tool result rather than silently falling back to a verifier that can see the executor's reasoning.Windows: a PowerShell that can actually be spawned. DSH resolves
pwshthrough a fixed candidate list (@deepseek-ai/dsh-pwsh-local); a Microsoft Store PowerShell can leave a zero-byte execution alias in%LOCALAPPDATA%\Microsoft\WindowsApps\pwsh.exe. That alias is a regular file tolstat, so it is selected, and Node then fails withspawn ...\WindowsApps\pwsh.exe ENOENT. If your harness will not start for that reason, point the resolver at a real binary:# profile cordis.patch.yml - id: pwsh-local config: pwshPath: 'C:\Program Files\WindowsApps\Microsoft.PowerShell_7.6.6.0_x64__8wekyb3d8bbwe\pwsh.exe'This is a host resolution issue, not a plugin one, but it is recorded here because it is the first thing that stops a Windows acceptance run.
License
MIT
No comments yet. Be the first to write one.