dsh-computer-control
A DeepSeek Harness bundle that turns Windows desktop operation into discrete tool calls. Every call performs exactly one action and returns the frames captured from just before that action until a settle period after it, so the model never has to remember to take a screenshot.
Installing
dsh plugin forwards its arguments verbatim to pnpm in the profile directory, so a
checkout, a git specifier, or a published tarball are all installed the same way:
dsh plugin --profile web add <path-to-this-checkout>
The peer dependencies (@deepseek-ai/dsh-tools, @deepseek-ai/dsh-llm,
@deepseek-ai/schemastery) are supplied by the harness, so this package declares them
rather than depending on them, and a checkout normally has no node_modules of its own.
DeepSeek Harness does not accept external pull requests; the documented way for a
community plugin to be found is a public repository carrying the dsh-plugin GitHub
topic, which is why package.json lists it under keywords.
The suites under test/ resolve their imports against this checkout and against the
harness packages, so they run from any working directory once the peers are installed:
npm test # or: node test/run-all.mjs
What you get
Eleven tools, registered on the host tools service:
| Tool | One call does | Returns |
|---|---|---|
computer_probe |
Nothing; reports screen geometry, pointer position and shape, and the window under the pointer | Text only, no image |
computer_screenshot |
Nothing; captures the screen, or a rectangle at an integer zoom | One image |
computer_move |
Moves the pointer | Frames around the move |
computer_click |
Presses and releases one button once | Frames around the click |
computer_drag |
Presses, sweeps, releases | Frames across the gesture |
computer_scroll |
One wheel gesture | Frames around the gesture |
computer_type |
Types a string as Unicode keystrokes | Frames while the text lands |
computer_key |
Presses and releases one named key | Frames around the key |
computer_chord |
Holds modifiers, presses one key, releases everything | Frames around the chord |
computer_wait |
Injects nothing and watches | Frames over the wait |
computer_frames |
Reads the recorded frames between two earlier frames | Up to the tool's own maxFrames (default 16); at most maxAttachedImages of them arrive as images and the rest are printed as paths |
computer_frames is the only tool that takes no new capture: it re-reads the recorder ring between two frames a previous call returned, which is how a two-frame envelope is enough to reconstruct the motion between them without paying for more images on every call.
Reading between two returned frames
A call captures ten to thirty frames and returns two of them. computer_frames reaches back into the
recorder ring for the rest of that span:
<computer-control action="frames">
result: ok
from: <the earlier frame the call printed>
to: <the later frame the call printed>
span: 1995 ms
recorded between them: 21 frame(s); returned: 4
achieved rate: 1.78 fps
attached: 2 of 4 returned frame(s), the bound that keeps one call's cost predictable
note: sampled 4 of 21 recorded frames at 2 fps; a higher fps returns more of them; stopped at the 4-frame bound; narrow the span or lower fps to see all of it
frames (offset relative to the earlier frame):
0. +0 ms | 1024x576 | 51399 bytes | image attached | <path>
1. +587 ms | 1024x576 | 51389 bytes | image attached | <path>
2. +1121 ms | 1024x576 | 51304 bytes | path only, read it with read_image | <path>
3. +1686 ms | 1024x576 | 51303 bytes | path only, read it with read_image | <path>
</computer-control>
Two things in that envelope are deliberate. recorded between them reports what the ring held, so a
thin answer is distinguishable from a span the recorder never covered, and attached names how many
of the listed frames arrived as images: the bound is maxAttachedImages, which exists to keep one
call's cost predictable, so frames past it are saved and printed as paths rather than attached. An
unattached frame can be read later with read_image instead of being asked for again, for as long as
the deployment's screenshot directory keeps it: that directory holds the most recent
SHOT_RETENTION files, after which the oldest saved paths stop resolving and a later request for one
says so rather than reporting a missing file.
Discrete by design
There is no hold, repeat, macro, loop, or "run this until" parameter, and no way to leave a button or a key down across calls:
computer_clicksends its press and release in one injected batch.computer_keyandcomputer_chordrelease every key they press inside the same call.computer_dragis one atomic gesture; its step count is bounded bymaxDragSteps.- Every tool is registered as concurrency-unsafe, so the registry serializes desktop input.
computer_scrollrefuses a zero-notch gesture instead of reporting a no-op as success.
To click twice, call computer_click twice. Each call then comes back with its own frames.
Configuration
Every value below is a Config field; override the row in your profile's
cordis.patch.yml, restating the keys you keep, because a patch replaces a row's
whole config.
| Key | Default | Meaning |
|---|---|---|
capturesPerSecond |
10 |
Frame rate the recorder produces |
maxFrames |
2 |
Frames returned by one call (1 to 20): the pre-action reference and the settled end state. computer_frames is capped lower, at 16 |
maxAttachedImages |
2 |
Frames attached as images (0 to 19); further frames are still saved to disk and their paths printed |
changeThreshold |
0.003 |
Signature movement that counts as a change |
settleMs |
700 |
Extra capture after the action completes |
settleRunFrames |
2 |
Consecutive unchanged tail frame pairs required before the envelope calls the window settled |
maxWindowMs |
15000 |
Ceiling on one capture window (200 to 120000); also sizes the tool timeout |
maxWaitSeconds |
10 |
Longest passive watch |
maxTypeChars |
2000 |
Longest string computer_type accepts (up to 20000) |
maxScrollClicks |
10 |
Largest wheel gesture, in notches |
maxDragSteps |
120 |
Largest drag step count |
maxZoom |
8 |
Largest integer zoom (1 to 16) |
maxZoomedPixels |
16000000 |
Pixel ceiling for one zoomed capture, checked before any allocation |
frameMaxEdge |
1024 |
Longest edge of a captured frame |
frameQuality |
75 |
JPEG quality, 1 to 100, for the in-process fallback capture only; recorded frames are encoded by ffmpeg under recorderQuality |
ffmpegPath |
ffmpeg |
Executable that records the screen |
tempRoot |
`` | Directory for the staged helper, recorder ring, screenshots, and markers; empty means a directory under the system temp directory. A drive root and an extended-length path are refused when the plugin activates, since with either one the recorder cannot be kept and every call would fall back to the in-process capture |
recorderQuality |
5 |
ffmpeg -q:v for recorded frames, 1 to 31; lower is better quality and larger files |
recorderBufferMs |
4000 |
Recorded video a call can reach back through for its own first frame |
recorderPrimeMs |
400 |
Recorder lead time before input is injected |
recorderIdleStopMs |
30000 |
Grace period before an idle recorder is stopped |
imageMode |
auto |
auto attaches frames when the route accepts images and the attachment service is mounted; text never attaches. The route's own capability is still resolved, so the envelope can say whether read_image would work for a frame already on disk |
captureTimeoutMs |
20000 |
Timeout budget for a capture or probe call |
actionTimeoutSlackMs |
10000 |
Added to maxWindowMs for an action tool timeout |
compactionCleanupEnabled |
false |
Delete this conversation's images once a compaction has summarised them. Off by default: the cleanup is implemented and tested, but eight adversarial reviews each found a real defect in it, so it ships unarmed. With it false a compaction deletes nothing and every image stays on disk. Arming it is a decision, not a default — see DELIVERY.md |
desktopTarget |
local |
Which desktop the screen and input work happens on. agent is the relayed instance's desktop — the agent's own, by default the deputy account; user is this process's own Windows session — the person at the keyboard; auto prefers agent and falls back to user only when the relay cannot be reached at all, never retrying a command the relay actually answered, because a relay that was reached may already have performed the action and repeating it on the user's desktop would be a second action nobody asked for (every fallback is logged). local and relay are mechanism-named spellings of user and agent. A deployment fact rather than a preference: ffmpeg's gdigrab and SendInput act only on the calling process's own session, so an instance that must drive a different account's desktop has to delegate to one running there |
relayUrl |
"" |
Origin of the relay instance, used when desktopTarget is agent, relay, or auto, for example http://127.0.0.1:3079. Only loopback is accepted |
relayToken |
"" |
Bearer token for the relay protocol. Setting it on an instance installs that instance's /relay route; empty means the route does not exist at all |
relayTimeoutMs |
60000 |
Ceiling for one forwarded command, so a hung relay cannot hold a tool call open indefinitely |
agentAccount |
"" |
Windows account this deployment runs automated sessions as. Together with userAccount it enables the account paragraph; while either is empty the paragraph is absent, which is the default |
userAccount |
"" |
Windows account the person at the keyboard uses, named in the account paragraph |
userDesktop |
"" |
That person's desktop path, named in the account paragraph. Empty leaves it unmentioned, which is right when it is the profile's own Desktop folder |
userDocuments |
"" |
That person's documents path, named in the account paragraph. Empty leaves it unmentioned |
Implementation
Capture is delegated to ffmpeg (-f gdigrab), which holds a real frame rate at full
screen size: measured 9.9-10.7 fps at 1600x900 and at 1024 wide, with no dropped frames.
ffmpeg scales to frameMaxEdge itself, so no frame is ever encoded by this bundle. The
in-process GDI+ path exists only as a fallback and costs 100-300 ms per frame, because a
PowerShell downscale-and-JPEG-encode dominates it; that cost is scaling, not capture.
lib/desktop-helper.ps1 is a long-lived PowerShell helper that owns input injection and
supervises the recorder. It speaks newline-delimited JSON over stdin/stdout and is
delivered as a BOM-encoded copy under a per-process temp directory, because Windows
PowerShell 5.1 reads a BOM-less script as ANSI. lib/desktop.mjs supervises that
process, lib/frames.mjs selects the frames worth returning, and lib/tools.mjs
declares the tools.
Frame records come from the files the recorder writes. Their instant is the file's
creation time, not the poll that noticed them: NTFS keeps sub-millisecond creation times,
and a poll would quantize every frame in one sweep onto the same instant. Frames are
emitted in ordinal name order, because that is capture order while creation timestamps can
tie or step backwards by a few milliseconds. The helper also runs a bare blocking
ReadLine, because Console.In.Peek() blocks on a redirected pipe in this PowerShell and
an "idle slice" loop around it hangs the helper.
Two lifecycle rules matter more than they look:
- The recorder must be stopped, not just the helper. ffmpeg is the helper's child, so
killing the helper alone leaves a recorder writing frames forever; a measured leak of 16
orphaned recorders reached 2.4 GB of frames. The watchdog covers a host that dies, and
sweepOrphanRecorders()covers a helper that was killed. - The ring is trimmed by frame count, never by age. A caller may still need a frame that no age window would keep, so the count is the safe bound, and it is applied only after a call has handed its frames over.
Returned captures are downscaled, so the result publishes the mapping back to real screen
pixels: the envelope names the screen rectangle and each frame carries a scale factor.
A model that measures a feature in a returned image multiplies by scale to get the
coordinate it should click.
Frame selection anchors on the injection instant rather than on the largest difference: it keeps the newest frame captured before input, the settled end frame, and then the changed frames nearest the injection.
Model Experience
Request context and condition
What the model sees
One system-prompt section, registered through ctx.systemPrompt.section() in the slot the harness reserves for computer-use plugins (TOOL_COMPUTER_USE, order 3000) and resolved by a provider on each assembly; the code is lib/prompt.mjs, the registration is in index.js. It has three parts, of which only the first is unconditional:
- The standing description of the tools and of what a call returns.
- The account paragraph, present only when this process runs as the paired agent account: which Windows account the session is, which account the person at the keyboard works as, and that their files may be read and written freely.
COMPUTER_CONTROL_ACCOUNTS=agent:useroverrides the pairing. - The desktop-target paragraph, present only when
desktopTargetis notlocaloruser: that commands act on the paired account's desktop rather than this session's, and that a shell command and a UAC prompt follow whoever executes them.
Nothing is added to the conversation, so the guidance neither competes with the user's own messages nor grows the history.
Verbatim standing description
Computer-control tools drive a desktop: computer_probe and computer_screenshot read the screen, and computer_move, computer_click, computer_drag, computer_scroll, computer_type, computer_key, computer_chord and computer_wait operate it.
A call returns two images by default: the last frame at or before the input was injected, and the settled frame after it. Read the first as the state the action started from and the second as its result; the envelope marks which returned frame is which. Every captured frame is also written to disk and the result prints an absolute path for each frame it returns, and those paths stay valid across a restart.
When two frames are not enough to see what happened between them, call computer_frames with those two printed paths and a target fps. Each returned frame costs vision tokens, so ask for the rate the question needs rather than the highest one.
One call performs exactly one action: to click twice, call twice. A call never leaves a key or button held down.
Token effect
Conditional, and fixed within each condition. The standing description is 1005 characters in every request. The account paragraph adds 1196 characters and appears only in the paired agent account. The desktop-target paragraph adds 628 characters, or 690 when desktopTarget is auto because that value also states the fallback, and appears only when the target is not local or user. A deployment that runs as the user's own account with the default target therefore pays the standing description alone.
KV Cache effect
Prefix-stable. The section is part of the system prompt and its text is identical across the requests of a deployment, so it extends the reusable prefix rather than replacing earlier tokens. Three package-owned changes invalidate reuse from that point on: moving a session between Windows accounts, which adds or removes the account paragraph; changing desktopTarget between local or user and any other value, which adds or removes the desktop-target paragraph; and switching between agent, relay, and auto, which changes that paragraph's length. All three are deployment changes, and a profile-patch edit that alters them is applied live by the loader without a restart.
Tool schemas for computer_*
What the model sees
Present in every request of a session whose profile mounts this bundle; the eleven tool definitions and their parameter descriptions are sent on every turn, and nothing here is conditional or capped. One entry per tool in the request's tool list, in registration order: the eight action tools (computer_move, computer_click, computer_drag, computer_scroll, computer_type, computer_key, computer_chord, computer_wait), then computer_screenshot, computer_probe, and computer_frames. The descriptions state that each call performs exactly one action, and the parameter descriptions name every bound (maxTypeChars, maxDragSteps, maxScrollClicks, maxZoom, maxWaitSeconds) so a rejected argument is predictable rather than surprising.
Token effect
A fixed cost per request for as long as the bundle is mounted, proportional to the eleven schemas. The parameter maps are deliberately narrow: coordinates are integers, a type string is capped, and no tool exposes a free-form object.
KV Cache effect
Append-only and stable. The tool list is byte-identical across turns of a session, so it forms part of the reusable request prefix and does not invalidate it.
Frames and envelope returned by an action tool
What the model sees
Only on a turn where the model actually called one of the eight action tools or the screenshot tool, and only when the calling route accepts images: imageMode: text suppresses attachment entirely, and a route without image input receives the envelope text and no attachments. Up to maxFrames frames come as images, each carrying its offset in milliseconds relative to the instant input was injected, plus a text envelope naming the screen rectangle, the image-to-screen scale, the cursor, the window under the pointer, the captured and returned frame counts, and the gap between returned frames. One frame is the deliberate pre-action reference, and the envelope labels it so the gap from it to the rest is not read as a sampling gap.
Verbatim envelope
<computer-control action="click">
result: ok
detail: {...}
captured: 12 frame(s), of which 2 were not returned
showing: 2 frame(s) over 874 ms, 84-383 ms apart (sampled, so gaps vary); recorded at 10 fps and returned as 2 image(s), not one image per recorded frame
first change: +81 ms after injection
last change: +180 ms after injection — the capture saw no movement after this
settled: no — only 1 unchanged frame pair(s) end the window (2 required), so the last returned frame is a mid-transition state rather than a settled one; a follow-up computer_wait is the way to see the outcome
screen: 1600x900 at 0,0; every returned image is a downscaled capture of it, so multiply a measured pixel by scale to get real screen pixels
frames (offset relative to the moment input was injected):
note: frame 0 is the pre-action reference (1 frame(s) at or before injection); the frames after it are the action and its result, spaced 84-383 ms
0. -32 ms | 1024x576 scale 1.5625 | cursor not recorded | pointer over not recorded | image attached, saved at <path> <- pre-action reference
1. +180 ms | 1024x576 scale 1.5625 | cursor not recorded | pointer over not recorded | image attached, saved at <path> <- settled end state
</computer-control>
Token effect
The dominant cost of this package. An image is billed by its pixel dimensions, so frameMaxEdge and maxAttachedImages are the two knobs that control it: lowering either cuts tokens per call directly. A 1024x576 frame is far cheaper than the 1600x900 screen it came from, and is why capture downscales at all.
KV Cache effect
Replaces earlier request tokens rather than appending to a stable prefix, so it does not extend a reusable prefix. Frames are never re-sent: they belong to the tool result of one turn, and a later turn that no longer shows them restores the shorter request.
Known Limitations and Deferred Work
- The recorder is not literally free while idle. It is started by the first call of a
burst and stopped by a watchdog once
recorderIdleStopMselapses with no further call. During that grace window the screen is still being captured; setrecorderIdleStopMsto a small value to shrink the window at the cost of re-priming (each cold start costs about 1.5-2 s before the recorder's first frame). - Frame time is the moment the file was written, not the moment the pixels were read.
gdigrabcapture, JPEG encode, and DWM composition put a lag of roughly one frame interval between the screen and the frame that represents it. A frame's offset is therefore accurate to about one interval, not exact. - A live pointer is composited into recorded frames. The cursor may be drawn at a position that does not match the moment the rest of the frame was captured.
- Recording is Windows-only. It depends on
gdigrab; there is no equivalent path for macOS or Linux, and no plan for one in this package. - The in-process fallback is much slower. When the recorder cannot start, the helper falls back to GDI+ capture, which this machine measures at 2-3 fps at 1024 px wide because the downscale and JPEG encode dominate — an order of magnitude below the 9.9-10.7 fps the ffmpeg recorder holds. The envelope reports the measured gap either way, so a fallback call is visibly slower rather than silently wrong.
- Route capability decides what the envelope can suggest. When the calling route rejects
images, the envelope says so and the frames stay on disk; the route's own capability is still
resolved, so the envelope distinguishes "this deployment turned images off" — where
read_imagecan still display a saved frame — from "this route refuses images", where nothing on that route can. Recovering the evidence then requires a route that accepts images. - Returned frames are a sampled subset. A call captures far more than
maxFrames, and the selection keeps two anchors — the last frame at or before injection and the final frame — then fills the remaining budget with a contiguous run starting at injection, and only then works backwards from the pre-action reference. Frames that changed most are considered before that padding, so a one-frame flash survives the cap; a rule that ranked purely by position would spend the whole budget on the first frames that twitched. The envelope reports two repeat counts —repeatedFramesby signature andidenticalFramesby JPEG bytes — because repetition there means the sampler had nothing new to add, not that the screen was watched continuously. Only the byte count is reproducible from the images a caller received, and the two differ by a few frames whenever re-encoding shifts pixels below the signature threshold. - A recorded frame's signature is computed by the helper. The frame triage needs one signature per frame to rank by change, and nothing else can supply it: the recorder writes JPEG files, no pixels pass through the helper, and the Node side has no image decoder. Each new frame therefore costs one 32x32 thumbnail decode — a few milliseconds — on the call path. The grid is sampled point by point rather than averaged, because an averaged cell hides a pointer, a checkbox, or a highlighted menu row inside its own mean.
- The two capture paths reach the same threshold through different grids. A recorded frame
carries a 32x32 signature sampled point by point, so one cell is 0.000977 of the frame and the
default
changeThresholdadmits three; the in-process fallback computes 16x16 averaged cells, where one cell is 0.0039 and the same threshold admits less than one. A single noisy averaged cell can therefore read as a change on the fallback path, and a feature small enough to live inside one recorded cell can be missed on the recorder path. Raising the zoom is the way to see such a feature; it does not change either grid. - A recorded frame carries no pointer metadata.
cursorandundercome from the GDI+ capture path, which samples them per capture; the ffmpeg recorder writes pixels only, so those members are absent from action frames and the envelope printsnot recordedrather than a placeholder coordinate. That data is available fromcomputer_probeandcomputer_screenshot, anddetailcarries the helper's own result for the call, which includes the window under the pointer for any call whose helper sampled it. - The image-to-screen factor points one way only.
scalealways means "multiply a pixel measured in the returned image by this to get a screen pixel", so it is greater than 1 for a downscaled full-screen frame and less than 1 for a zoomed region. The two capture paths measured it in opposite directions, so it is derived from the frame's own source rectangle and the returned width rather than taken from either one. waitcannot observe on its own terms. It injects nothing, so its frames are the screen changing by itself; a genuinely static screen yields frames that differ only in the composited pointer.- Frames are sampled, so the returned set is usually full. Leftover budget goes to a
contiguous run from injection, which means a call that captured more than
maxFramesnormally returns exactlymaxFrames, most of them identical to their predecessor when the screen was quiet. The envelope's repeat counts say how many. - Windows only. The helper compiles
user32.dllinterop through PowerShell and refuses to load on other platforms. - One desktop. All sessions in one machine share one pointer and one keyboard. Two concurrent calls are serialized by the registry, but a call in one session still moves the pointer the user is holding.
- Input goes to the focused window for keyboard actions.
computer_type,computer_key, andcomputer_chordact on whatever holds focus; only the pointer tools target a coordinate. A Chinese IME in the target window can capture injected keys and turn them into composition text. - Capture is best-effort at the requested rate. The recorder holds about 10 fps at this screen size; the in-process fallback captures one full-screen frame in roughly 100-300 ms, so on that path a short window yields two or three frames rather than a strict sample.
- A frame is one screen, not one window. There is no per-window capture, so a window that is covered cannot be observed without bringing it forward.
- Region zoom multiplies pixels.
maxZoomedPixelsbounds the allocation, but a large region at high zoom is still an expensive capture. - Screenshots are kept on disk under the deployment's scratch root, capped at 200 files per scratch root, and are not deleted when the plugin unloads.
engines.dshis declarative only in this Harness vintage: nothing reads it.peerDependenciesis the field the plugin manager actually checks.- The scratch profiles used for automated verification are not part of this bundle; the package ships no test suite of its own yet.
No comments yet. Be the first to write one.