
dsh-pilot
Give DeepSeek Harness hands and eyes.
Out of the box, a DeepSeek Harness session can reason about your desktop but cannot touch it: you describe what is on screen, you take the screenshots, you do the clicking, and you report what happened. dsh-pilot closes that loop. The model looks at the real pixels and then operates them:
| tool | what it does |
|---|---|
screen_view |
captures the live desktop and returns the PNG as a real image block the model can see — not OCR, not a text summary |
desktop_control |
moves the pointer, clicks, drags, scrolls, types, sends keys, focuses and lists windows — and can capture the result in the same call |
desktop_sequence |
runs a whole batch of those actions in one call and looks once at the end |
You : 把浏览器里那个报表的第三列改名成 "Q3 收入",然后保存
DSH : desktop_sequence([focus "报表.xlsx", click 812,430, type "Q3 收入", key enter, key ctrl+s])
→ 1 round trip, 1 screenshot, per-step timings, done
Why this exists
A model that cannot see the screen has to be told what is on it. A model that cannot touch the screen has to be told what to do next. That turns every desktop task into a conversation about the desktop instead of work on the desktop — and the user becomes the hands, the eyes, and the error message.
The expensive part of desktop automation was never the mouse. It is how many times the model and the computer have to talk:
| a five-step task: focus a window, click a field, type, press Enter, look at the result | DSH alone | DSH + dsh-pilot |
|---|---|---|
| who reads the screen | you | the model, from real pixels |
| who clicks and types | you | the model |
| model ↔ computer round trips | 5 (one per step, plus yours) | 1 |
| screenshots shipped to the model | 4–5 | 1 |
| where the model spends its time | waiting, re-describing, re-asking | thinking about the task |
| when a step fails | you notice | the same result says which step, and why |
What it feels like
sequenceDiagram
participant U as You
participant M as Model
participant P as dsh-pilot
participant W as Warm worker (PowerShell)
participant D as Your desktop
U->>M: 把第三列改名成 Q3 收入并保存
M->>P: screen_view(window: "报表")
P->>D: capture (per-monitor DPI aware)
D-->>M: real PNG + cursor, foreground window, monitor bounds
M->>P: desktop_sequence([focus, click, type, key, key])
P->>W: 5 requests over one pipe — 6–12 ms each
W->>D: SetCursorPos, SendInput, clipboard paste
D-->>M: one end frame (with coordinate rulers) + per-step outcome
Why the coordinates can be trusted
The classic way a vision-driven click goes wrong is a coordinate that is almost right. dsh-pilot removes the three ways that happens:
1. Both halves are per-monitor DPI aware. A DPI-unaware process is lied to by
Windows: it sees a 2560×1600 panel as 1707×1067, so every click computed from such a
screenshot lands in the wrong place. The sensor and the effector each declare
SetProcessDpiAwarenessContext(PER_MONITOR_AWARE_V2) before doing anything, so one
image pixel is one physical screen pixel.
2. Frames carry rulers. Ticks every 200 px along the top and left edge, each
labelled with the screen coordinate it sits at, plus an origin x,y step n note.
The model reads a position off the image instead of counting pixels — and because the
labels are screen coordinates, they survive any downscaling the attachment store does.
3. The envelope says what it actually delivered. If the store re-encodes a frame smaller, the model is told the ratio and the multiplier, and the coordinate contract becomes conditional — it only promises "use x and y as they are" when the delivered frame is the capture:
<image_size delivered="1974x873" captured="2261x1000">1974x873</image_size>
<delivered_scale>delivered 1974x873 = 0.873x capture; multiply image readings by 1.1454 to get screen pixels</delivered_scale>
Why it is fast
Batching is the big win (round trips, not clicks). The second win is that PowerShell stops being restarted: one warm worker process answers every action and every frame.
| request | one-shot script | warm worker |
|---|---|---|
action (move, click, type, …) |
~0.9–2.0 s | 6–12 ms |
| capture (primary monitor) | ~1.2–2.3 s | ~0.5 s |
Measured on a 2560×1600 panel at 150% scaling with the harness UI animating. The worker is started in the background on the first tool call, so the ~1 s it needs overlaps the first capture instead of delaying the first action.
Identical frames are not resent. Every capture returns a frameHash; when a frame
is byte-identical to the previous one of the same session, the image is not attached
(or stored) again — the result says unchanged: true and the coordinates from the
frame the model already has still apply. forceImage: true overrides it.
Why it is safe
Giving a model hands is the part worth getting right:
- Risk is declared, then enforced. Low risk (default) batches freely. A batch
declared
medium/high, or one carrying a step markedrisk: "high", is refused before anything touches the desktop unless the caller passesconfirm: true— which it does only after you agreed. dryRun: truevalidates a batch and lists the plan without executing.- Stop on the first failure (default): a refused
focuscan never be followed by typing into whatever window happens to be in front. - Targeted input. When a step names a window (
hwnd), the effector re-asserts it as foreground immediately before injecting, so a click cannot land in a neighbour. - An unknown outcome is never a silent retry. If the worker dies after a request was delivered, the action is reported as failed rather than repeated — it may already have taken effect. A capture has no side effects, so it is retried.
- Interruption works. Pressing Esc writes a cancel token, the action loop polls it
between every step (during drags, during the settle wait, before the capture) and a
drag releases the button in a
finallyblock, so an aborted turn cannot leave the mouse latched down or the remaining clicks playing out. - It cannot lie about privilege. Windows refuses input injection into a higher-integrity (elevated) window; the tool reports that failure instead of pretending it worked.
desktop_sequenceis bounded bymaxSequenceSteps(default 24).
Install
# from the repository
dsh plugin --profile <name> add github:momasiku/dsh-pilot
# or this repository's prebuilt tarball, which skips the build-approval step
dsh plugin --profile <name> add https://github.com/momasiku/dsh-pilot/releases/latest/download/dsh-pilot.tgz
# or a checkout you already have
dsh plugin --profile <name> add "file:E:\path\to\dsh-pilot"
Install it by repository, not by bare name. dsh plugin add asks the npm
registry first, and the name dsh-pilot on npm belongs to an unrelated
browser-automation plugin.
The package ships its own loader row, so plugin add places dsh-pilot in
dsh.profile.bundles and the row comes from the package (dsh.bundle.patch →
cordis.patch.yml) rather than being hand-written into your patch layer. Restart
the desktop app once afterwards, then start a new conversation: the three tools
join the tool list.
Requires DSH 0.1.5-rc.2 through 0.2.x — that is the peer range in package.json,
and what dsh plugin add checks before it installs anything. Requirements:
Windows, PowerShell 5.1 or 7, and a model route that declares image input
(deepseek-flash does).
Tools
screen_view
| parameter | type | meaning |
|---|---|---|
screen |
string | primary (default), all (whole virtual desktop), or a 0-based monitor index |
window |
string | case-insensitive substring of a window title or process name; frames that window |
region |
string | crop as x,y,width,height in physical screen pixels |
includeCursor |
boolean | draw a crosshair at the pointer (default true) |
rulers |
boolean | draw the coordinate rulers (default: the rulers setting) |
forceImage |
boolean | attach the frame even when it repeats the previous frame of the session |
The result carries the image plus a <screen_capture> envelope: delivered size,
capture size when they differ, physical origin, scale, every monitor's bounds, cursor
position, and the foreground window (title, process, class, bounds) — then the
coordinate contract for that frame.
desktop_control
| parameter | type | meaning |
|---|---|---|
action |
string, required | click, doubleClick, rightClick, middleClick, move, drag, scroll, type, key, focus, windows |
x, y |
integer | target in physical screen pixels; required by the pointer actions and by scroll-at-a-point |
toX, toY |
integer | drag end point, or a nonzero value to make scroll horizontal |
text |
string | text to insert for type |
key |
string | a character, enter/tab/esc/f5, or a chord like ctrl+s, alt+tab, win |
amount |
integer | wheel notches for scroll (default 3, positive scrolls up) |
button |
string | left (default), right, middle |
title |
string | window title or process substring; required by focus, filters windows |
hwnd |
string | lock input to one window handle (0x1094C or decimal) |
capture |
boolean | capture a fresh frame after acting (default true) |
settleMs |
integer | wait before that capture so the screen can repaint (default 750) |
forceImage |
boolean | attach the frame even when it repeats the previous frame of the session |
type delivers text as one paste (clipboard set, Ctrl+V, clipboard restored), so
applications see a single insertion — multi-line text and CJK included — instead of a
burst of per-character keystrokes. windows returns every visible titled top-level
window with its bounds, minimized state, and which is foreground: useful before
clicking anything.
desktop_sequence
| parameter | type | meaning |
|---|---|---|
steps |
array, required | ordered steps; each is one desktop_control action without its own screenshot, plus optional settleMs, risk, riskNote; action wait takes ms |
risk |
string | low (default) / medium / high — the risk you declare for the whole batch |
confirm |
boolean | required for a medium/high batch; without it the tool refuses before touching the desktop |
dryRun |
boolean | validate the batch and list the plan without executing anything |
capture |
string | end (default) captures one frame after the last step, none skips it |
target |
string | window (default: the window that was foreground during the batch), screen, or all |
window |
string | frame this window instead (title/process substring), overriding target |
region |
string | crop the end frame as x,y,width,height in physical pixels |
forceImage |
boolean | attach the end frame even when it repeats the previous frame |
stopOnError |
boolean | stop at the first failing step (default true) |
settleMs |
integer | delay before the end capture so the screen can repaint |
Every step comes back with its own action, duration in milliseconds, and outcome — so one call still tells the model exactly where a batch went wrong.
Settings
- insert:
- id: pilot
name: dsh-pilot
config:
captureRetention: 30 # frames kept under .dsh-pilot/
timeoutMs: 20000 # per-call kill deadline for a helper
settleMs: 750 # default delay before the automatic capture
stepSettleMs: 120 # per-step delay inside a desktop_sequence
maxSequenceSteps: 24 # hard cap on the steps of one batch
alwaysSaveFile: true # keep the PNG even when no image block is attached
useWorker: true # keep one PowerShell process warm (6-12 ms actions)
rulers: true # draw coordinate rulers on every frame
How it works
flowchart LR
subgraph JS["lib/index.js (the plugin)"]
T1[screen_view] --> C[capture]
T2[desktop_control] --> A[runAction]
T3[desktop_sequence] --> A
C --> R[frame hash + dedup + envelope]
end
C -->|warm request| W[desktop-worker.ps1]
A -->|warm request| W
C -.->|fallback| PB[desktop-probe.ps1]
A -.->|fallback| AB[desktop-action.ps1]
W --> WC["_dsh-win32.ps1 · _dsh-action.ps1 · _dsh-capture.ps1"]
PB --> WC
AB --> WC
Two halves, one coordinate space: the sensor reports the geometry, the effector
consumes the same physical pixels, and GetCursorPos/SetCursorPos report back, so
the loop is self-checking.
The shared code sits in _dsh-win32.ps1 (the Win32 bridge and DPI awareness,
idempotent), _dsh-action.ps1 (Invoke-DshAction) and _dsh-capture.ps1
(Invoke-DshCapture). The one-shot scripts stay as thin bootstraps — they are what
the plugin falls back to when the worker cannot be used. The worker speaks one
base64 JSON request per line, answers with one JSON line per request, and keeps
serving after a failed action, a failed capture, or a cancelled step.
Captured frames land in <session cwd>/.dsh-pilot/ and are pruned to
captureRetention. Pruning never breaks an image the conversation already carries,
because attachments are stored content-addressed.
Troubleshooting
| symptom | cause and fix |
|---|---|
no visible top-level window matches "..." |
the window is hidden to the tray or the title changed; desktop_control { action: "windows" } lists what is really there, and hwnd targets a window that has no title match |
| clicks land next to the target | the frame was delivered smaller than the capture — check <delivered_scale> and apply the multiplier |
| a click does nothing | the target window may be elevated; Windows refuses input injection across integrity levels |
| the tools are missing after an update | restart the desktop app: a package's code is re-imported on load, and the profile's plugin rows are composed at boot |
| a batch refuses to run | it was declared medium/high (or contains a risk: "high" step): re-send with confirm: true after the user agreed, or dryRun: true to see the plan |
Developing
The smoke tests load the plugin against the real @deepseek-ai/dsh-tools from the
desktop checkout, so they exercise every schema and renderer without a harness restart:
# the package resolves @deepseek-ai/* from the desktop app's node_modules
New-Item -ItemType Junction -Path node_modules\@deepseek-ai `
-Target "D:\AGAENT\DSH Desktop\resources\app\node_modules\@deepseek-ai"
node tools/smoke.mjs "D:\AGAENT\DSH Desktop\resources\app" # load, schemas, renderers
node tools/render-smoke.mjs "D:\AGAENT\DSH Desktop\resources\app" # coordinates: 1:1 vs shrunk, contract, rulers
node tools/sequence-smoke.mjs "D:\AGAENT\DSH Desktop\resources\app" # end-to-end batch on a throwaway window
powershell -NoProfile -ExecutionPolicy Bypass -File tools\script-smoke.ps1 # PowerShell half + worker protocol
sequence-smoke.mjs drives a real batch against a throwaway window it creates
itself (never one of your applications) and the window writes back what it received —
so "focus, then type, then Enter" is proven to have happened in that order, not
merely attempted. script-smoke.ps1 drives the PowerShell half directly: the one-shot
paths, the worker handshake, warm latency, a failing action, a cancellation, a failing
capture, and that the worker survives each of them. The last two need a full-access
shell: the plugin reads its helpers through pipes, which a restricted sandbox denies
with EPERM.
The helper scripts can be driven by hand while debugging — they take a UTF-8 JSON params file and print one JSON result:
'{"out":"E:\tmp\shot.png","screen":"primary","rulers":true}' | Set-Content params.json -Encoding utf8
powershell -NoProfile -File lib\scripts\desktop-probe.ps1 -ParamsPath params.json
powershell -NoProfile -File lib\scripts\desktop-action.ps1 -ParamsPath params.json
Keep the .ps1 files saved with a UTF-8 BOM: Windows PowerShell reads BOM-less
files as ANSI, which mangles non-ASCII window titles and paths.
Reinstalling after an edit
dsh plugin add file:... installs a copy, and a repeat add of an unchanged
file: spec reports "Already up to date" without re-copying even when the files on
disk changed. After editing, copy the changed files into the installed package
yourself and restart the desktop app:
$src = "E:\path\to\dsh-pilot"
$dst = "$env:USERPROFILE\.dsh\profiles\<profile>\node_modules\dsh-pilot"
Copy-Item "$src\lib\index.js" "$dst\lib\index.js" -Force
Why the smoke test validates result shapes
A tool result carrying a key its closed output schema does not declare is rejected by
the runtime after execute() has already performed the side effect — so the action
happens yet the caller sees an error. A schema-only check cannot see that failure
mode, so smoke.mjs validates a realistic result object against each tool's compiled
output schema, and additionally asserts that the validator itself catches the
undeclared ok/actedAt keys that shipped once.
License
MIT © 2026 momasiku
中文说明
给 DeepSeek Harness 装上眼睛和手。
DSH 原本只会"说":它能推理、能写代码、能规划,但看不见你的屏幕,也动不了你的鼠标。于是看屏幕的是你、截图的是你、点按钮的是你、出错了汇报的也是你——你成了它的手、它的眼,还兼任它的报错信息。
dsh-pilot 把这个环闭上:模型直接看真实像素,然后自己动手。
| 工具 | 作用 |
|---|---|
screen_view |
抓取当前屏幕,把 PNG 作为真正的图像块交给模型(不是 OCR、不是文字描述),连光标位置、前台窗口、各显示器范围一起给出 |
desktop_control |
移动鼠标、单击/双击/右键、拖拽、滚轮、打字、按键、聚焦与列出窗口;可以在同一次调用里顺手截一张,让效果立刻可见 |
desktop_sequence |
一次调用跑完一整串动作,最后只看一眼结果 |
你 :把浏览器里那张报表的第三列改名成 "Q3 收入",然后保存
DSH:desktop_sequence([focus "报表.xlsx", click 812,430, type "Q3 收入", key enter, key ctrl+s])
→ 1 次往返、1 张截图、每步耗时,完成
它到底解决了什么
桌面自动化的瓶颈从来不是鼠标点得快不快,而是 "模型和电脑要来回说多少次话":
| 一个五步任务(聚焦窗口 → 点输入框 → 打字 → 回车 → 看结果) | 只用 DSH | DSH + dsh-pilot |
|---|---|---|
| 谁在看屏幕 | 你 | 模型,看真实像素 |
| 谁在点击打字 | 你 | 模型 |
| 模型 ↔ 电脑往返次数 | 5 次(每步一次,外加你的) | 1 次 |
| 送给模型的截图 | 4–5 张 | 1 张 |
| 出错时 | 你发现、你描述 | 同一次结果里就写明哪一步、为什么 |
为什么坐标可信(这是"点得准"的关键)
视觉驱动点击最经典的翻车方式是"坐标差一点点",这里堵掉了三个来源:
- 两端都是 per-monitor DPI 感知。非 DPI 感知的进程会被 Windows 骗——它把 2560×1600 的屏幕看成 1707×1067,于是按这种截图算出来的点击必然偏。传感器和执行器在做任何事之前都声明
PER_MONITOR_AWARE_V2,所以一个图像像素就是一个物理屏幕像素。 - 每张图都带坐标刻度尺:上沿与左沿每 200px 一条刻度,标注的是屏幕坐标(不是图像偏移),还有
origin x,y step n说明。模型直接读刻度,不用数像素;刻度是屏幕坐标,所以即使图片被压缩也依然有效。 - 信封如实说明交付尺寸。如果宿主把图压小了,模型会收到比例与乘数;坐标契约也变成条件式——只有在"交付尺寸 == 采集尺寸"时才承诺"照用 x、y 即可",否则明确要求乘多少:
<delivered_scale>delivered 1974x873 = 0.873x capture; multiply image readings by 1.1454 to get screen pixels</delivered_scale>
为什么快
批处理是大头(省的是往返,不是点击);第二个大头是不再每次重启 PowerShell——一个常驻 worker 同时服务动作与截图:
| 请求 | 一次性脚本 | 常驻 worker |
|---|---|---|
| 动作(移动/点击/打字…) | ~0.9–2.0 秒 | 6–12 毫秒 |
| 截图(主屏) | ~1.2–2.3 秒 | ~0.5 秒 |
(2560×1600、150% 缩放、DSH 界面还在动的情况下实测)。worker 在第一次调用时后台预热,它需要的约 1 秒与第一张截图重叠,不会拖慢第一个动作。
没变化的帧不重复发送:每次截图都带 frameHash,与同会话上一帧逐字节相同时就不重复附图(也不重复存储),结果里写 unchanged: true,模型手上那张图的坐标继续有效;想看就传 forceImage: true。
为什么安全
给模型装上手,最该讲究的就是这一块:
- 风险先声明、后强制:默认
low可自由批量;声明medium/high的批次、或含risk: "high"步骤的批次,在碰桌面之前就被拒绝,除非调用方带confirm: true(而它只会在你同意之后带)。 dryRun: true只校验并列出计划,不执行。- 默认遇错即停:
focus失败绝不会有"接着往面前那个窗口里打字"的后续。 - 定向输入:步骤里指定了窗口(
hwnd)时,注入前会再次把该窗口置前,点击不会落进旁边的窗口。 - 结果不明就绝不静默重试:worker 在请求送达后死掉时,动作按失败上报而不是重放——它可能已经生效了;截图没有副作用,才会自动回退重试。
- 中断是真的能中断:按 Esc 会写取消令牌,动作循环在每一步之间轮询(拖拽中、等待中、截图前都查),拖拽在
finally里松开按键——中断的回合不会留下按住的鼠标,也不会把剩下的点击跑完。 - 不会假装成功:Windows 禁止向更高完整性级别(提权)的窗口注入输入,工具会如实报错。
- 批量有上限:
maxSequenceSteps(默认 24)。
安装
# 从仓库装
dsh plugin --profile <你的profile> add github:momasiku/dsh-pilot
# 或用本仓库 Release 里的预构建包,免构建授权
dsh plugin --profile <你的profile> add https://github.com/momasiku/dsh-pilot/releases/latest/download/dsh-pilot.tgz
# 或你本地已有的检出
dsh plugin --profile <你的profile> add "file:E:\path\to\dsh-pilot"
⚠️ 请用仓库地址安装,不要用裸包名。 dsh plugin add 会先去 npm registry 查,而 npm 上的 dsh-pilot 属于另一款无关的浏览器操控插件。
包自带 loader 行:plugin add 会把 dsh-pilot 放进 dsh.profile.bundles,行本身来自包内的 cordis.patch.yml(由 dsh.bundle.patch 声明),所以不会像手写补丁行那样被插件管理器重写时弄丢。装完重启一次 DSH Desktop,然后新开一个会话,三个工具就会出现在工具列表里。
要求 DSH 0.1.5-rc.2 至 0.2.x(即 package.json 里的 peer 范围,也是 dsh plugin add 在安装前的门禁)。环境要求:Windows、PowerShell 5.1 或 7、以及声明了图像输入的模型线路(deepseek-flash 可以)。
使用要点
- 先
desktop_control { action: "windows" }看清有哪些窗口、哪个在前台,再动手。 - 多步任务优先用
desktop_sequence:把"聚焦 → 点击 → 输入 → 回车"写成一次调用。 - 打字支持多行与中文(走剪贴板一次粘贴,之后恢复剪贴板)。
- 想更省 token:
capture: "none"不附图,或依赖"未变化不重发";想看得更清:rulers: true(默认开)。 - 常用设置(
captureRetention、settleMs、stepSettleMs、maxSequenceSteps、useWorker、rulers)写在配置行里,见上面的 Settings 一节。
常见问题
| 现象 | 原因与处理 |
|---|---|
| 提示找不到匹配窗口 | 窗口缩到托盘了,或标题变了:用 action: "windows" 列出来,或改用 hwnd 指定 |
| 点击总是差一点 | 图片被宿主压小了:看 <delivered_scale> 的乘数并乘上去 |
| 点了没反应 | 目标窗口可能是提权窗口,Windows 不允许跨完整性级别注入输入 |
| 升级后工具不见了 | 重启 DSH Desktop:包的代码在加载时重新导入,profile 的插件行在启动时合成 |
| 批次被拒绝执行 | 声明了 medium/high:经用户同意后带 confirm: true 重发,或先用 dryRun: true 看计划 |
License
MIT © 2026 momasiku
No comments yet. Be the first to write one.