TokenSqueezer
A DeepSeek Harness plugin that reduces what an agent costs to run. It works on the model's own output — not on the input it is fed.
English | 中文
Read this first: what actually saves tokens
Two mechanisms with completely different economics live in this plugin, and conflating them is the easiest way to be disappointed.
| Mechanism | Saves | Why | |
|---|---|---|---|
| Prevention | brevityInstruction |
output | The model writes less. The answer stays complete. |
reasoningBrevityInstruction |
output | Reasoning is ~87% of everything a model emits; asking it to think in a straight line is the largest lever here. | |
outputTokenCap |
output | A ceiling set through agent/request before the call. Tokens never generated are never billed. |
|
reasoningEffort |
output | Pins the reasoning tier. Applied only when the model advertises it. | |
| Compression | llm/stream rewrite |
replay only | By the time a chunk arrives, those tokens are generated and paid for. |
session/event backstop |
replay only | Same. | |
answer / reasoning / per-kind folds |
replay only | Same. |
Compression cannot make the current answer cheaper. That money is spent. What it does is shorten the transcript, so every later request in the turn carries fewer bytes — a saving that compounds and is invisible at the moment it is made.
The plugin never touches usage, so reported token counts stay truthful.
Install
# 1. Put this repository somewhere permanent, then link it into a profile.
# The profile manages its own node_modules; a link: install keeps the plugin editable.
cd <your DSH profile directory>
pnpm add link:<path-to-this-repository>
# 2. Enable it. The bundle ships cordis.patch.yml, which inserts the row for you.
# In <DSH_HOME>\profiles\<profile>\cordis.patch.yml:
- id: tokensqueezer
disabled: false
Restart the application. The two halves reload by different mechanisms, and the asymmetry is worth knowing before it wastes an afternoon:
| How it loads | Editing it while DSH runs | |
|---|---|---|
Host half (index.js, lib/) |
a loader entry, read once at application start | does nothing until a full restart |
Client half (client.js) |
a module artifact whose revision is derived from mtimeMs/ctimeMs/size |
hot-reloads immediately |
So a UI that updates while behaviour does not is the normal symptom of a stale Host. status prints
the file the Host was loaded from, which makes the question a comparison rather than a guess:
Host code: <path>\index.js @ 2026-09-30T04:26:12.732Z.
How it works
Host half — five surfaces, each installed independently
| Surface | Hook | Role |
|---|---|---|
| Stream rewrite | llm/stream (waterfall) |
Buffers each text/reasoning block to its block-end, compresses, re-emits. This is the only leg that runs before the text is displayed, and because the loop builds its assistant/message from the chunks it receives, it also decides what is committed. |
| Durable backstop | session/event |
Replaces an over-budget assistant/message that reached the log by another route, using the replay-safe compaction/prune + surfaceOp: replace protocol. A no-op in the normal path. |
| Output budget | agent/request (waterfall) |
The prevention half: sets maxTokens, optionally pins reasoningEffort, before prepareCall. |
| Tool results | tools/post-execute (waterfall) |
Compresses a tool result before the model reads it — with the code guarantee below. |
| Status route | webServer.register |
A read-only GET /api/tokensqueezer/status so the UI can show a number. GET only; anything else gets 405. The payload is this plugin's own counters: no session content, no paths, no credentials. |
Each is its own installation step, so a failure names itself in status (Installed: …) and does not
take the other four down with it.
Client half — two seats
conversation.composer.dock— a readout below the composer. It polls the status route every 3s and shows the session's saving. When no number is available it says so instead of showing a placeholder.settings.section— a settings page built from the shipped primitives (Tag,Switch,SegmentedControl,Input,CodeBlock,Button), grouped by what each setting actually saves, localized through the shipped locale service.
The pipeline
lib/engine/output.js splits a payload into fenced-code and prose segments first, so no code block
ever reaches the semantic layer. Everything else flows through:
- L1 lossless — ANSI and control removal, carriage-return redraws, trailing whitespace, blank-line runs, repeated line/block folding.
- L2 normalisation — path relativisation against the session cwd, optional symbol interning. Both skipped for executable kinds.
- L3 semantic — one compressor per detected kind:
json,yaml,diff,stacktrace,log,terminal,testreport,filetree,searchresults,deptree,markdown,html,code,prose, plusanswerandreasoning, which fold at answer scale.
Reversibility
Everything the lossy layers remove is interned into a content-addressed store (8 hex characters of
sha1) and replaced by a bare ⟦TS:id⟧ marker, with a legend line at the end of the payload. The same
text always yields the same id, so a second pass over compressed output is a no-op. The model
retrieves originals through the tokensqueezer tool's expand action.
The code guarantee
A tool result that is file content or a patch is not compressed semantically. This is enforced in three places, and the reason there are three is a mistake worth recording:
The first implementation called
engine.compress(text, { lossy: false })and assumed that was enough. Measured on a 40-line repeated body, it returned three lines. The repeat-line folding that damages code runs in the lossless stage, so switching the lossy layer off happens upstream of the damage.
So:
- The tool name decides the kind.
toolContentKind()pinsread/read_file/view/cat/write/edit/str_replace→code, anddiff/patch→diff. Heuristics are the wrong instrument here: a source file thatdetectKindhappens to read aslogwould have its repeated lines folded. - Those kinds skip the pipeline entirely and receive
losslessCleanuponly.EXECUTABLE_KINDSis not the guard — it only protects L2, whilecompressCode/compressDiffdo fold code bodies. - The result must be net-shorter. A payload whose final text — marker and legend included — is not shorter is discarded unchanged. This is what makes "compress everything" safe rather than self-defeating, because a 13-character marker plus a ~100-character legend would otherwise grow small payloads.
Measured on the shipped engine:
read_file, 40 repeated code lines 1079 → 1079 40/40 lines survive
patch, 30 added lines 1231 → 1231 30/30 lines survive
code with ANSI escapes and blanks 1122 → 1111 41/41 lines survive, escapes stripped
Configuration
Every key has a default; nothing needs configuring. The floors default to 0, so an ordinary
payload always enters the compressor — and the net-shorter rule above decides whether the result is
kept.
| Key | Default | Meaning |
|---|---|---|
compressOutput |
true |
Master switch for output compression. |
compressStreamOutput |
true |
The llm/stream leg — the only one that runs before display. |
compressDurableOutput |
true |
The session/event backstop. |
compressReasoning |
true |
Fold reasoning blocks. false skips them before any compressor runs. |
compressToolResults |
true |
Compress tool results, behind the code guarantee. |
outputMinChars |
0 |
Floor for a visible text block. |
reasoningMinChars |
0 |
Floor for a reasoning block. |
toolMinChars |
0 |
Floor for a tool result. |
minSavedChars |
0 |
Worth-it floor. Safe at 0 because of the net-shorter rule. |
segmentMinChars |
200 |
Floor for one prose segment when an answer is split around code. |
reasoningElideMinChars |
0 |
Compressor-internal floor for reasoning. |
reasoningFoldMinChars |
0 |
Compressor-internal fold floor for reasoning. |
answerElideMinChars |
0 |
Compressor-internal floor for answers. |
answerFoldFloor |
0 |
Compressor-internal fold floor for answers. |
outputTokenCap |
unset | Hard ceiling on generated tokens. Too low truncates the answer (finish: max-tokens). Never raises an existing stricter limit. |
reasoningEffort |
unset | Reasoning tier to force. Quality tradeoff; applied only when the model advertises the value. |
brevityInstruction |
true |
Ask for lean answers. The one writing-style change the plugin makes. |
reasoningBrevityInstruction |
true |
Ask for lean reasoning. |
level |
1 |
0 lossless only, 1 balanced, 2 aggressive. |
relativizePaths |
true |
Relativise repeated absolute paths. Never on code or diffs. |
internSymbols |
false |
Intern long repeated identifiers. Never on code or diffs. |
protectCodeFences |
true |
Keep fenced code out of the semantic layer. |
toolSkip |
[] |
Tool names whose results are never compressed. |
archMaxChars / archMaxDepth |
1200 / 3 |
Bounds for the architecture digest. |
outputTokenCap and reasoningEffort apply to every agent, subagents included — deliberately, so
a delegation cannot quietly outspend the budget it was given.
Tooling
The plugin ships its own verification, because the engine layer was never the part that broke:
$node = '<node executable>'
& $node scripts\linkcheck.mjs # every relative import resolves to a real export
& $node scripts\selftest.mjs # engine, 9 sample kinds: 122537 → 32812 chars (73.2%)
& $node scripts\hosttest.mjs # Host half, fake ctx and a session as strict as the real one
& $node scripts\clienttest.mjs # Client half, fake loader, a real render, and the schema the settings form needs
& $node scripts\reasontest.mjs # skeleton folding: skeleton kept, references reversible
& $node scripts\zstdtest.mjs # session-log reader: concatenated frames, truncated tails
& $node scripts\archtest.mjs # architecture digest
Three diagnostics exist for the questions the UI cannot answer:
| Script | Question it answers |
|---|---|
scripts/audit.mjs |
Did the plugin actually rewrite committed output? It reads the durable session logs and counts, per session, how many assistant/message rows carry a marker — split into answers and reasoning. The UI is a derived view; this is the log. |
scripts/sessionrepair.mjs |
Which rows would format-v4 admission refuse, and offers to fix two shapes of them. Previews by default; --write backs up first. Its second rule only applies to v4 logs, so v3 sessions in other projects are never touched. |
scripts/asar-find.mjs · asar-extract.mjs |
What the shipping Harness actually does. Both read the application archive's header and extract named entries, which is how the contracts below were read instead of guessed. |
hosttest.mjs uses a fake session that enforces the real rules — an assistant/message replacement
must not carry sourceEventSeqs, must carry turn/step while that step is open, and surfaceOp
must be exactly {op, startSeq, endSeq}.
Platform constraints
These were all found the hard way, and each one cost at least one restart. They are properties of the Harness, not of this plugin.
A profile-installed bundle cannot import
@deepseek-ai/*. The packages live inside the application archive; Node's resolution never reaches it, and the entry fails with a barefailed to import. So the Host half is a functional plugin with no framework imports, and the two helpers it needs are vendored.cordis resolves a service on a context only when that context injects it. An undeclared service is not merely unavailable — reading it as a property throws. That throw is what silently killed one installation step here.
injectis a gate, and both halves declare what they need (['tools', 'webServer']and['slots', 'locale', 'configForms']).A plugin
Configis two different protocols. cordis activation validates through Standard Schema (Config['~standard'].validate, synchronous); DSH's settings projection wants a schemastery brand (Symbol.for('schemastery'),type,meta). Supplying only the brand satisfies the inspect provider — which reports a healthystatus: "schema"— while the entry goes tofiberPhase: "failed". An inspect provider is a read-only projection, not a validator.A client-visible projection needs compile-time declaration merging into
SessionProjectionMapplus awirecodec, and a typert Remote namespace needs generated code. Neither is available to a plain-JS plugin, which is why the UI reads a read-only HTTP route instead.SessionReferencedocuments behaviour, not a content API. It has paging, minima and lifetime semantics, but no public way to read the loaded messages.agent/request,llm/streamandtools/post-executeare waterfalls and can legitimately replace what they return.agent/assistant-streamisemitonly and can never be rewritten.
Honest limitations
Short content does not compress, and should not. With every floor at zero the compressor is still reached; it simply finds nothing to remove. A 46-character answer has no foldable middle, and an 11-character tool result has no redundancy. Pushing further would be deleting information, not saving it.
Measured savings on real sessions are modest — single-digit to low-double-digit percent of the blocks that qualify, with roughly half of all blocks having no redundancy at all. The large numbers are in the prevention knobs, and two of those (
outputTokenCap,reasoningEffort) default to off because they change what the model produces.The settings page's write path is verified against the contract, not against a browser. The controls call
mutate(ops, revision)— read from the shippedSettingsFormModel, not guessed. Two gates decide whether a hand-built schema is served at all, andlib/config.jsnow satisfies both: the entry'sConfigmust answer"toJSON" in schema, and at least one field must carrymeta.volatile. Miss the first anddescribe()drops the row before it reads anything else; miss the second andvolatileForm()drops it one step later. Either way no namespace appears,configForms.get()finds nothing, and every switch and number input renders disabled.clienttest.mjspins both gates, so the page cannot silently go grey again. What is still unexercised is the round trip through a live browser: change a value, watch forSaved, and confirm the Host applies it.toolSkipis not editable from the settings page. It is the only non-scalar setting, and an array node that failed to rehydrate would drop the page back tostatus: loading— greying every control, the exact failure above. It stays editable through thetokensqueezertool'ssetaction and the profile patch row.The dock readout depends on the status route. If the route is not reachable it says so; it never renders a number it does not have.
The reference store is per session and bounded.
expandresolves within the session that produced a marker, and the LRU keeps 32 sessions warm. Originals always remain in the durable log.
Layout
index.js Host half: five surfaces, the tool, the prompt sections, the status route
client.js Client half: the composer dock and the settings page
cordis.patch.yml bundle patch: inserts the tokensqueezer row
package.json bundle and client declarations
lib/config.js the plugin Config (Standard Schema + schemastery brand)
lib/session-message.js the v4 producer-source rule still relied on
lib/arch.js architecture digest
lib/engine/
compress.js orchestrator
output.js fenced-code split, code protection, answer-scale wiring
detect.js content-type detection
textsafe.js L1 lossless
refs.js content-addressed reference store
dedupe.js line and block folding
normalize.js path and symbol normalisation
tokens.js token estimation
kinds/ one compressor per content kind
locale/ i18n: en.json, zh.json
scripts/ the seven suites, the shared session-log reader, and the three diagnostics
License
MIT.
No comments yet. Be the first to write one.