dsh-taskforce
Multi-agent orchestration plugin for DeepSeek Harness (dsh).
Splits long tasks across multiple purpose-built agents so no single agent's context accumulates the whole task: a planner decomposes work into a task graph, each worker runs in a fresh one-shot context holding only what its task needs, and workers return structured reports instead of transcripts.
Core loop: plan → dispatch (parallel) → structured reports → review → pass | redo.
Status: P3 review loop working end to end — plan → dispatch → structured
reports → review (pass/fail) → reroute → re-dispatch → finalize, verified with
local fake-host tests and live headless runs. State is keyed per conversation;
worker tool access is whitelist-constrained; worker profiles are configurable.
See ../DESIGN.zh.md.
Install
From a local checkout (requires dsh 0.1.7-rc.2; run pnpm install inside this
directory after cloning — the bundle has one npm dependency):
git clone https://github.com/LiSheng5/dsh-taskforce.git
pnpm install
dsh plugin --profile web add <path-to>/dsh-taskforce
<path-to> must be an absolute path. dsh plugin add also accepts git specs
directly (github:LiSheng5/dsh-taskforce) — pnpm then resolves the dependency at
install time.
Compatibility
| dsh version | status |
|---|---|
0.1.7-rc.2 |
developed against, verified |
dsh is a developer preview and officially announces breaking changes. After upgrading dsh, run the offline regression suite once — it needs no dsh, no API key, no network:
node test/local-loop.mjs
Upgrading from a version that predates the restart-recovery fix? Conversations it touched need a one-time repair — see Upgrading from a version that predates the fix.
Safety
This plugin spawns real sub-agents with real powers: file reads/writes and shell access inside the session workspace, bounded by each worker profile's tool restrictions. Read dsh's own SAFETY notes before running orchestrations on workspaces you care about, keep the per-task timeout and attempt caps enabled, and treat worker reports as claims — that is what the reviewer is for.
Workers also lose the host's agent-plane control tools by default
(send_message, interrupt_agent, list_agents, subagent, subagent_fork):
send_message would post a worker's text straight into the orchestrator's
session, which is exactly what this plugin exists to prevent, and the others
let a worker disturb its parent or start more agents. This list was calibrated
by measurement (a real researcher worker reported exactly these tools); set
denyControlPlane: false to keep them.
Layout
index.js plugin entry: name / inject / apply
client.js browser half: the panel above the composer (lazy __ModuleLoader__ factory)
cordis.patch.yml bundle patch inserted into the profile composition
src/protocol.js Task / Report / Verdict JSON schemas (host JSON-Schema subset)
src/graph.js DAG validation, readiness, status transitions (CAS)
src/service.js orchestration state (ctx.provide('orchestra'))
src/events.js restart recovery: folds the host's tool/call + tool/result back into a graph
src/router.js context router: the task book a worker receives
src/profiles.js agent_type → persona / reasoning effort
src/json-schema.js argument validator (subset semantics)
src/tools/ orchestrate_plan / _status / _dispatch (+ ping)
scripts/ maintenance: repair session logs written by older versions
test/local-loop.mjs offline regression suite (fake host, no network)
Restart recovery
The graph lives in memory, keyed per conversation, so a restart loses it. What brings it back is the session log the host already keeps:
- dsh logs every tool call as
tool/call(full arguments) andtool/result(rendered output), so orchestration state is durable in the host's own vocabulary — the plugin writes no session events of its own. - The plugin registers an
orchestraGraphsession projection (underctx.inject([...]), so a composition without the projection registry is unaffected) that purely folds those events into the task graph, and loads it the first time a fresh process touches the conversation. - A call's effect lands when its result settles. The host writes
tool/callbefore the tool body runs, so applying eagerly would apply the call being executed right now a second time. A dispatch that was cut short by a restart is therefore never half-applied either: its task stays where the orchestrator last saw it — ready to run again. - Without the projection service the plugin simply runs in-memory, as before.
orchestrate_status answers with the recovered graph after a restart.
Upgrading from a version that predates the fix
Early versions appended their own event types, which made those conversations
unloadable after a dsh restart (contains event type "orchestra/graph" … refusing to interpret the log). The repair script adds the harness's own
compatibility marker (ignorable) so they load again, backing up first:
node scripts/repair-session-logs.mjs --dry-run # report first
node scripts/repair-session-logs.mjs # repair (self-checked, backup kept)
node scripts/repair-session-logs.mjs --restore # undo
Web UI panel
Conversations that have an orchestration show a panel above the composer (a conversation that never planned anything renders nothing at all):
▾ ▣ orchestra #1 · 1/2 · review 1 PASS
─────────────────────────────────────
✓ t1 researcher completed att 1
read package.json — name dsh-taskforce, version 0.1.0 …
· t2 coder pending → ready
- It reads the wire view of the same
orchestraGraphprojection (client.jscallsuseProjection('orchestraGraph')), so it shows exactly whatorchestrate_statusreports, updates itself as events arrive, and after a restart shows the recovered graph. - Only ids, agent types, statuses, attempts, readiness, report summaries, diagnostics and issue counts cross to the browser; worker personas stay host-side.
- The browser half is a build-free plain script (
client.js, declared bydsh.client.platform=webwithimmediately). It inherits the host theme viacurrentColorand pulls in no stylesheet or third-party dependency. - Uninstalling the plugin removes the panel with its loader row; the rest of the conversation UI is untouched.
Tools
orchestrate_plan— submit the task graph (mission_brief + tasks); validates fail-loud and returns the ready set.orchestrate_status— compact one-row-per-task view plus dispatchable ids, and the diagnostic for tasks that failed without producing a report.orchestrate_dispatch— spawn ready tasks as isolated one-shot workers, block until they settle, return one report digest per task (≤4 concurrent).orchestrate_review— run the reviewer over the reports; returns a verdict.orchestrate_reroute— send review issues back to workers with guidance.orchestrate_finalize— record the delivery summary and lock the graph.orchestra_ping— liveness probe.
Configuration (worker personas, models, limits)
The orchestrator writes per-task personas itself: each task in a plan may
carry a persona field — a role definition A tailors to that task ("你是数据库
迁移专家,只改 SQL 文件"). It becomes the worker's primary system prompt, layered
above the profile's discipline (non-negotiable constraints that A cannot
override: researchers don't write, reviewers don't fix, …). model and
reasoning_effort are likewise per-task fields A sets at plan time.
Two more ways to tune what each worker role is like — both take effect after restarting the dsh process:
Edit
src/profiles.jsdirectly. Each role hasrole(identity line),discipline(permanent constraints),effort(off|low|high|max), andtoolDenyPatterns— regexes matched (case-insensitively) against the host's global tool names to hard-restrict that role's tools.Config overrides (no code edit). Add a
config:block to the plugin entry in the profile'scordis.patch.yml(e.g.~/.dsh/profiles/web/cordis.patch.yml):- id: dsh-taskforce name: dsh-taskforce config: taskTimeoutMs: 1200000 maxConcurrentWorkers: 6 profiles: researcher: role: | 你是资深代码考古学家…… discipline: | 只读,不修改任何文件。 effort: off toolDenyPatterns: ["write", "bash"] coder: model: deepseek-chatConfig
role/discipline/model/effortreplace the builtin values;toolDenyPatternsare concatenated with the builtin ones. Engine knobs:provider(defaultspawn),maxConcurrentWorkers(4),maxAttempts(3),maxReviewRounds(3),taskTimeoutMs(600000).
Plan-time worker programming
The orchestrator is not limited to the four builtin worker types. In a plan it
may declare new worker types — orchestrate_plan accepts a workers field
where A programs one agent spec per type:
"workers": [
{
"type": "ui-builder",
"role": "你是界面构建师,专注像素级还原。", // identity (required for new types)
"discipline": "不得改动 src/api/ 下的文件。", // optional; generic floor otherwise
"effort": "high",
"model": "deepseek-flash", // optional defaults
"toolDenyPatterns": ["bash"]
}
]
- Tasks then reference
"agent_type": "ui-builder". Shadowing a builtin type ("type": "coder", no role) replaces only the parts it names. - The layering stays: a task-level
personastill layers above the declared role, and the declareddisciplineis always appended — A cannot write it away. - The review loop is not reprogrammable from plans: the reviewer always runs on the operator-configured profile.
- Engine-level guarantees are unaffected: orchestration tools are always denied to workers, delegation depth prevents grandchild spawning, tasks are one-shot isolated sessions.
Design notes
- Hand-built tool definitions, near-zero imports: tool definitions match
what
defineToolcompiles to; the only dependency is the npm-published@deepseek-ai/schemasteryfor the Config schema, so pnpm-local-path installs stay trivial. - Context isolation: workers are spawned (never forked) one-shot subagents; upstream information travels only as structured reports, never as history.
- Per-conversation state: the orchestration graph is keyed by the calling agent's session id — simultaneous conversations never share state.
- Layered worker restriction: orchestration tools are always denied to
workers; profile deny-patterns are resolved against the host's global tool
view at dispatch time (unknown names are never passed, so a bad pattern
degrades instead of throwing). Grandchild delegation is impossible by
delegation depth. Control-plane tools can only be denied by PATTERN — they
are not ours and may not exist in every composition — so
resolveToolFilteralways enumerates first and matches second. - Only the host's vocabulary reaches the log: orchestration state rides the
host's own tool-call records (
tool/call+tool/result); the plugin adds host-log audit lines (orchestra/graph, …) and no session event type the harness would refuse to read back. - Recovery replays the service itself: the fold calls the same
submitPlan/markDispatched/settleReportmethods with the recorded arguments instead of reimplementing the state machine, so a recovered graph cannot drift from live semantics.
License
MIT — see LICENSE.
No comments yet. Be the first to write one.