DSH HUB
HomePlugin StorePlugin PacksCommunityRankingsResourcesPublish Guide
Plugin source
Back to catalog

dao-awa /

dao-awa/dsh-computer-use-native

Verified

Windows computer use for DeepSeek Harness, built for a vision model: screenshot the real desktop, read the interface from pixels, and drive it through Win32 FFI with no accessibility tree. Input posts to a window's own queue so the desktop focus is left alone.

★ 0 Stars0 Forks0 IssuesN/A Community rating0 Confirmed installs
View on GitHub
READMESource: master@2005a001

dsh-computer-use-native

Windows desktop control for DeepSeek Harness, built for a vision model.

The agent takes a screenshot, reads the interface from the pixels, and clicks what it sees. Input tools accept coordinates measured on that screenshot and convert them to screen coordinates themselves, so the model never does the scaling arithmetic that makes vision-driven clicking unreliable.

Why this exists

Accessibility-tree automation fails on modern Windows applications. Chromium, WebView2, and Electron windows expose almost no accessible controls — a UI Automation tree over the new Outlook or a Vivaldi window returns a handful of nodes and no content. A tool built on that tree reports an empty window and the agent cannot read the interface at all.

Screenshots do not have that problem. Whatever draws the pixels, the pixels are there. A model that can see them can operate the application.

This provider therefore does the opposite of a tree-based tool: it captures the real rendered surface with PrintWindow and PW_RENDERFULLCONTENT, which is the only route that returns actual pixels for a DirectComposition surface, falls back to a screen-region capture when that yields a blank frame, and hands the image to the model with the coordinate mapping attached.

It also aims at a window without taking the desktop over. An agent that must raise every window it touches makes the machine unusable for whoever is sitting at it, so input posts messages to the target's own queue by default: nothing is raised, the real cursor does not move, and the person at the keyboard keeps working. See Input routes for the measured behaviour and the two limits that come with it.

Requirements

  • Windows 10 or later, x64.
  • Node.js 22 or later.
  • A model route that accepts image input.
  • DSH with the computerUse, tools, systemPrompt, and attachments services.

Install

dsh plugin add dsh-computer-use-native

Or install straight from the repository, pinned to a commit so a later push cannot change what you are running:

dsh plugin --profile web add github:dao-awa/dsh-computer-use-native#1354d1d1481619c2cc0a5ecccfa8180f1a34452a

The bundle's patch layer mounts two rows: @deepseek-ai/dsh-computer-use, which owns the provider registration and is not part of the base bundle, and this provider. Restart DSH and start a new session; bundles are mounted before a session is created or resumed.

Coexistence with an MCP desktop server

computer-use-win and similar packages are MCP servers bridged through @deepseek-ai/dsh-mcp-client. They are not computer-use providers, so they do not conflict with this one at registration. They do offer the model a second, overlapping set of desktop tools, which is confusing in practice. Mount one or the other.

Tools

Tool Purpose
computer_screenshot Capture the desktop or one window; returns the image and a viewport id.
computer_window List windows, read one window's geometry, or raise a window.
computer_click Click, double-click, or right-click at a point on the last screenshot.
computer_move Move the pointer without clicking, for hover menus and tooltips.
computer_type Type Unicode text — any character, regardless of keyboard layout.
computer_key Send a key or chord such as enter, f5, ctrl+shift+t.
computer_scroll Scroll the wheel at a point on the last screenshot.
computer_drag Drag between two points, for sliders, selection, and drag-and-drop.

The five input tools all take the same dispatch parameter, described under Input routes.

How coordinates work

This is the part that decides whether the agent hits what it aims at.

A screenshot is downscaled before it reaches the model, so the image is smaller than the screen region it came from. On a 2560×1600 display at 150% scaling, capturing a 2582×1550 window and downscaling to fit 1568 pixels gives a scale factor of about 1.65. A model that assumes image pixels are screen pixels misses its target by 65% of the distance from the origin — hundreds of pixels on a wide window.

So the plugin absorbs the conversion:

  1. computer_screenshot records a viewport: the image size, the screen rectangle it came from, the scale factors, and the window it belongs to.
  2. The model reads a target's position off the image and passes those numbers to an input tool.
  3. The tool converts image pixels to screen coordinates through that viewport, exactly and reversibly, and reports where the action actually landed.

Two guards make this safe rather than merely convenient:

  • Staleness. The viewport fingerprints the source window's frame. If the window moved or resized after the capture, input is refused instead of landing where the model never looked.
  • Bounds. A point outside the image is refused rather than clamped. Clamping would act on an unrelated part of the desktop.

Looking is cheap, so look often

A screenshot costs mostly its encode. Measured on a 2560x1600 desktop: the GDI capture is about 43 ms, converting and encoding to PNG is 122–208 ms, and comparing pixels is a fraction of a millisecond. Re-encoding a screen that has not moved is the whole cost of looking twice, and looking twice is most of what an agent does.

So each look is compared against the last look at the same target, and when the view has not moved the tool says so, sends no image, and points back at the viewport that already describes it. Measured end to end through the shipped tool: 238 ms for the first look, 41 ms for a repeat — 17% of the cost, with no attachment written and no image in the request.

The comparison is deliberately not byte identity. Five samples of a live desktop taken a fifth of a second apart produce five different frames, because a caret blinks, a spinner turns, and a clock ticks. A person watching that screen sees one unchanged view. Each frame is reduced to a 32x20 grid of average luminances and a look counts as unchanged while fewer than 2% of those cells move, which sits above a spinner and below a window opening. On a real desktop the animated indicator in a running agent loop moves 0.0–0.2% of the grid.

Repeat looks read a reduced capture rather than the frame the model would be shown. Both cost the same to obtain — GDI's screen readback dominates and does not shrink with the destination, so a 160x100 probe costs the same 43 ms a full capture does — but the probe is 63 KiB instead of 15.6 MiB, so looking repeatedly does not churn the heap, and its resolution loss averages away exactly the flickers that are not changes.

waitForChangeMs uses the same comparison to make an action's result observable. Instead of capturing a half-drawn frame or sleeping for a guessed delay, the tool watches the screen and reports whether it moved and after how long.

The floor is the screen readback, not the encoding. Going below it needs the Desktop Duplication API, which reads only the rectangles the compositor reports as dirty; GDI's BitBlt has no such option and always copies the whole surface.

Input routes: background and foreground

Every input tool takes a dispatch parameter, and the two routes trade opposite costs.

background is the default. It posts window messages — WM_LBUTTONDOWN, WM_CHAR, WM_MOUSEWHEEL — straight to the window the screenshot came from. That window is never raised, the real cursor never moves, and whoever is using the machine keeps their focus and their typing. Mouse messages carry client-area coordinates, so the tool converts the screenshot point back through ScreenToClient before posting; WM_MOUSEWHEEL is the one exception, whose lParam holds screen coordinates.

A posted message goes to the control that owns the point, not to the frame the screenshot showed. Windows sends real mouse input to the deepest window under the pointer, and most classic Win32 software — dialogs, property sheets, anything built from control windows — keeps its edit boxes and buttons in child windows. A frame that receives a message meant for one of its controls discards it, silently, while PostMessage still reports success. So the module hit-tests the point with WindowFromPoint, confirms the hit is a descendant of the addressed window, and posts there instead. Characters and keys go to the focused control, which is reachable only after AttachThreadInput shares the target's input queue with ours, since a thread's keyboard focus is private to that thread.

A window that owns its whole surface — a browser, or a XAML application with no child windows — is its own hit-test result, so nothing changes for it.

foreground raises the target and moves the real cursor through SendInput. Every window receives it, at the cost of taking the desktop over.

Background is the default because a silent no-op the model can see and retry is a smaller failure than commandeering the machine. The route is chosen per call, so a model that finds one step did nothing retries that step with dispatch="foreground" and leaves the rest of the run silent.

Measured behaviour — Windows 11, 150% DPI, each target backgrounded behind a real application and driven through the shipped tools:

Case Result
Click a backgrounded Chromium page Delivered; the page's own click counter incremented.
Type into the focused field of a backgrounded Chromium page Delivered; all 15 characters arrived.
Foreground after a click into body text Unchanged.
Foreground after a click into a text field Chromium activates its own window, because the field needs keyboard focus. The result note reports it.
Click and type into a WinForms text box, addressed by its frame's coordinates Delivered to the child control; all 19 characters arrived, and the foreground stayed where it was.
Click and type into Windows 11 Notepad Nothing arrived. Notepad is a XAML app with no child windows, and it does not read posted messages. dispatch="foreground" is the route for it.

Three consequences follow. Posted characters reach whatever already holds focus inside the target window, so a field has to be clicked before text lands in it. A control that needs keyboard focus — a text field, not body text — makes its own window activate, and the tool reports that rather than hiding it. And an application that reads hardware input instead of its message queue, which includes XAML and WinUI software, will ignore the posted route entirely; the result says so, and the retry is dispatch="foreground".

If a window will not come forward on the foreground route, the foreground lock is the usual cause. The tool raises through AttachThreadInput to share the target's input queue, which is the documented way around it. When even that fails it reports the refusal instead of sending input to the wrong place.

Capture routes

captureWindowAuto tries PrintWindow with PW_RENDERFULLCONTENT first, counts the distinct colours in the result, and falls back to a screen-region capture when the frame is near-uniform. No single route covers every window class: an IME host window returns one colour through PrintWindow while a Chromium window returns a blank frame through a naive BitBlt. A capture that stays blank through every route is reported to the model as unreliable rather than presented as a valid image.

Configuration

Set these under the plugin's config in cordis.yml:

Field Default Meaning
maxEdge 1568 Longest edge of a returned screenshot, in pixels. Lower spends fewer tokens per capture; higher makes small UI text legible.
compressionLevel 6 PNG effort, 0–9. Measured difference between 6 and 9 is about 1% of file size, so raising it rarely pays.
typeDelayMs 0 Pause between typed characters. Raise it for applications that drop fast input.
providerName native-win32 Name recorded in the computer-use registration.

Measured on a 2582×1550 browser window, PNG bytes by maxEdge:

maxEdge Image Bytes
1024 1024×615 533 KB
1280 1280×768 762 KB
1568 1568×941 1.08 MB
1920 1920×1153 1.49 MB

A route that normalizes images for its own token budget may reduce these before they reach the model.

How it is built

Direct Win32 through koffi FFI — no native addon, no C++ toolchain, no compiled artifact to match a Node ABI. The published package is JavaScript plus a prebuilt FFI binding.

src/
  index.ts          plugin entry: calibration, registration, prompt
  guidance.ts       the model-facing system-prompt section
  viewport.ts       image <-> screen coordinate mapping and staleness
  encode.ts         BGRA -> RGBA, downscale, PNG
  tools/            the eight tools
  win32/
    structs.ts      koffi structure layouts
    dll.ts          user32/gdi32/kernel32/shcore bindings and constants
    native.ts       layout assertions and UTF-16 decoding
    window.ts       DPI awareness, enumeration, geometry, foreground
    capture.ts      PrintWindow/BitBlt with fallback, colour sampling
    input.ts        SendInput: pointer, wheel, keys, Unicode text
    post.ts         posted messages: the background input route

Five details are load-bearing and would be silent failures if wrong:

  • DPI awareness is set before anything else. SetProcessDpiAwarenessContext with PER_MONITOR_AWARE_V2 must run before the first capture or coordinate query. Without it GetSystemMetrics reports scaled logical pixels while the capture returns physical pixels, so every click is off by the scale factor.
  • The INPUT structure is asserted to be 40 bytes. A wrong stride makes SendInput deliver corrupt events with no error. The plugin refuses to start rather than send them.
  • Text uses KEYEVENTF_UNICODE. Characters are delivered independently of the active keyboard layout, so Chinese, accented, and emoji input work without changing layouts. Surrogate pairs are kept adjacent.
  • Posted mouse messages carry client-area coordinates, not screen ones, and the client origin is not the frame origin — a browser draws its own toolbar inside the client area. Posting a screen point mis-clicks by the height of that toolbar. WM_MOUSEWHEEL is the exception and does take screen coordinates.
  • A posted message belongs to the control under the point, not to the frame. Posting to the frame is accepted and then discarded by any window whose controls are separate child windows, which is most classic Win32 software. The point is hit-tested with WindowFromPoint and the message goes to the deepest descendant of the addressed window.

Limitations

  • Windows only. The provider refuses to load elsewhere.
  • No session isolation. Input goes to the interactive desktop. A locked or disconnected session has nothing to capture.
  • Protected windows. Windows with DRM or elevated-integrity protection may refuse capture or reject synthesized input. The failure is reported, not hidden.
  • Posted input can be ignored. The background route is the default because it is quiet, not because it is universal. A window that reads the hardware input queue rather than its message queue sees nothing — a XAML or WinUI application such as Windows 11 Notepad is the measured case — and the tool reports the refusal so the model can retry with dispatch="foreground".
  • No accessibility data. This provider deliberately does not read the UI Automation tree. Reading a control's value without seeing it is a different capability; use an MCP UI Automation server alongside if you need it.

Development

npm install
npm run typecheck                        # compiles against the harness declarations
npm run build
npm run pipeline                         # capture -> encode -> viewport -> coordinates
npm run pipeline:click                   # the same, and sends a real click
npm run compose                          # mounts the plugin in a real Cordis context
npm run compose:real                     # mounts it against the real harness services
npm run bg-tools                         # drives a backgrounded Chromium page through the tools
npm run child-routing                    # posts into a WinForms control window
npm run unchanged-look                   # looking twice at the real desktop
npm run latency                          # where a screenshot's time goes
npm run bench                            # image size by maxEdge and PNG effort

Three tsconfigs exist because they have different jobs. tsconfig.json typechecks against the harness declaration files, so a build fails when the host's API changes. tsconfig.test.json clears that mapping so the spike scripts resolve the real packages from node_modules at runtime; a loader would otherwise follow the mapping to a .d.ts file. tsconfig.harness.json does the opposite and maps the specifiers onto the harness sources, which is the only way a test can mount a real harness service instead of a stand-in. The spike scripts must be started with the matching one.

compose:real is the composition check that matters. It mounts the plugin against the real @deepseek-ai/dsh-computer-use, @deepseek-ai/dsh-system-prompt, and @deepseek-ai/dsh-tools rather than stand-ins, and each of the three has an API a stub would have hidden:

  • The provider registry reserves its slot through a ctx.effect call made inside register(). Cordis resolves a service's this.ctx to its caller, so that effect binds to this plugin's fiber and disposal releases the slot — a claim about framework behaviour, and one whose failure would leave the slot occupied after an unload.
  • The prompt service owns the section order this plugin asks for, and assembling is what the model actually receives, so the check asserts the guidance text appears in the assembly and leaves it again on disposal.
  • The tool registry validates each definition and projects its schema to lossless JSON. Reading the catalog back is what shows the eight tools are registrable; a stub register() that accepts anything cannot.

The profile side is verifiable without starting the app:

dsh --profile web --dump-config          # composes the tree, patches applied
dsh --profile web --dump-config-schema   # imports every entry to read its Config

The schema dump is the stronger of the two, because printing an entry's config schema requires importing its module. This plugin's maxEdge, compressionLevel, and typeDelayMs appearing there, with their defaults, is what proves the patch row resolves and the built module loads.

Harness schema rules the tools follow

These are easy to get wrong and the error messages do not explain the rule:

  • Tool parameters and tool output.schema are different DSLs. parameters marks a required field with required: true on the field. output.schema does too — but only on fields. A required key at the root of an output schema is rejected, whether it is true or an array of names, so requiredness there is expressed on each property.
  • Every nested object in either DSL must state additionalProperties explicitly as true or false. Omitting it is an error.
  • A function plugin exports name, inject, Config, and apply, and must have no default export; a default export makes the loader discard the namespace.
  • Every registration must be yielded inside ctx.effect(), including systemPrompt.section(). A registration made outside the effect survives disposal.
  • spike/schema-probe.ts and spike/schema-probe2.ts record which forms the installed harness accepts. Re-run them after a harness upgrade.

spike/win32-smoke.ts exercises enumeration, capture, and the foreground workaround without the plugin layer.

License

MIT

—/ 5

No ratings yet

Verified DSH bundle

Commit 2005a001f40b

Community comments

No comments yet. Be the first to write one.

DSH HUB

A community index for DSH plugins. Not an official GitHub or DeepSeek AI product.

CommunityResourcesAPIAbout