English | 中文
dsh-token-ledger — cross-project token accounting for DSH
A DSH plugin that totals the tokens consumed across every project in the installation and presents them in a graphical panel in Settings.
Where to find it: Sidebar → Settings → "Token ledger" (it sits after Models / Plugins / Agent Presets).
Requirements
This is a DSH profile bundle, not a standalone program. It registers with a DSH profile (through plugin_manager or the dsh plugin CLI) and reads that installation's session corpus — the persisted sessions belonging to the profile it is loaded into. Every figure it shows is derived on demand from that corpus rather than from a store of its own.
The host half registers one HTTP route (GET /token-ledger/summary) and the panel is a web client, so DSH's web server must be running for the UI to appear at all. The package declares "engines": { "node": ">=22.15" } — the host half itself only uses built-ins, and the corpus tests additionally need Node's zstd decoding, available from 22.15.
What it counts
Every session in DSH carries a tokenUsage projection that records four billing buckets:
| Field | Meaning |
|---|---|
uncachedInputTokens |
input tokens that missed the cache |
cacheReadTokens |
input tokens read from the cache |
cacheWriteTokens |
tokens written to the cache |
outputTokens |
tokens generated by the model |
These numbers come from the usage the provider reports on each call. They are the real billing figures, not heuristic estimates (heuristic estimates are used only for the context-occupancy display).
The plugin groups every session into a project by the session's working directory (header.cwd) and then reports the totals.
Table columns
| Column | Meaning |
|---|---|
| Sessions | total sessions in that project |
| Subagents | how many of them were produced by a subagent (header.origin === 'subagent'). Most of these are fresh contexts and do not double-count tokens anywhere else |
| Inherited | how many of them "copied a parent session's history at the start of the log". This is the column that actually represents duplicated volume |
| Input | uncached + cache read + cache write — all three are billed on the input side |
| Output | outputTokens — tokens generated by the model |
| Tokens | deduplicated tokens; in fast mode the column is headed "Recorded tokens" and shows the un-deduplicated value |
| Cache rate | cache read ÷ input — the same definition the chat stats strip uses, and a partial hit is never shown as 100% |
| Share | bar share relative to the largest value in that project |
Why "inherited history" rather than "subagents": the two are not the same thing. In the sample corpus only 40 of 105 subagent sessions carry an inherited prefix; conversely, manually forking a conversation counts as "inherited history" even though it is not a subagent. This column always measures one thing: whether the log starts with somebody else's history.
The criterion is the header's isSeeded — the session's own declaration of that fact, which both read paths get for free — whereas the exact inheritedEventCount is only knowable by reading the log (deduplication itself still uses that measured value; isSeeded is not substituted for it).
Two accounting modes (switchable in the panel)
The panel has two buttons at the top:
Accurate (default)
For each session, only its own events are counted (seq >= inheritedEventCount), skipping the log prefix inherited from the parent session.
Why this is necessary: a DSH fork / subagent session copies the parent session's history, and the tokenUsage projection is folded over the whole log, inherited prefix included. A fork session's own recorded token count therefore already contains the parent's share. Adding up the records of 120+ sessions directly would count the same billed calls more than once.
Sample corpus (about 120 sessions, roughly a third of them inherited):
summed directly 3.00B tokens ← about 70% too high
deduplicated 1.75B tokens ← every billed call counted once
The cost: every session log has to be read (zstd-compressed; decompressed it is roughly 4× the compressed size). The first tally takes about 3 seconds at sample scale. After that each session is reused incrementally against sessionPersistence's revision — only sessions whose log changed are read again — so a repeated tally normally takes a bit over a hundred milliseconds, with no TTL wait. Pressing "Recount" forces the cache to be cleared and every log to be re-read.
Fast
Reads the durable projection cache (ctx.sessionProjectionCache) directly, reading no logs at all, and returns in about 60 milliseconds.
The cost: the inheritedEventCount boundary is unavailable, so nothing can be deduplicated and there is no model-route attribution; the numbers come out high. The panel says so explicitly.
The deduplicated figures are trustworthy
The fold logic in this plugin mirrors the tokenUsage projection in @deepseek-ai/dsh-token-meter. Self-check results:
- 65 / 65 sessions with no inherited prefix: the locally computed figure and the projection value DSH persisted itself are exactly identical (byte for byte).
- 40 / 40 fork sessions: the locally computed value is ≤ the recorded one.
- The panel footer displays
foldMismatches; a non-zero value usually means some session is being written to.
Two entry points
| Entry point | Location | Notes |
|---|---|---|
| Sidebar icon | the sidebar's top global panel row (below Plugins) | clicking it switches the centre column to the ledger page; the page has "Back to conversation" in its top-left corner |
| Settings page | Settings → Token ledger | the same component as the sidebar panel, only in a different host container |
The sidebar entry uses the same mechanism as two official DSH plugins (Plugins, Task manager): register an icon row in sidebar.panellist, then register a key with the same id in the main slot. The row button, the label, the tooltip and the click behaviour all belong to the sidebar itself; the plugin supplies only the icon glyph and the panel content — so do not wrap it in a span of your own, which would break that row's icon baseline.
"Back to conversation" calls ctx.layout.selectPanel(null), and it reads ctx.get('layout') optionally rather than declaring it in inject: a missing declared service leaves the fiber pending, and a client plugin that is pending, or that throws in apply, takes the whole page's boot down with it — not a risk worth taking for a back button.
The four views on the page
| Control | Effect |
|---|---|
| Accurate / Fast | switches the accounting mode (see above) |
| All time / Last 30 days / Last 7 days | filters by session creation time; all four views recompute with it |
| Recount | clears the incremental cache and re-reads everything (for troubleshooting). Repeated clicks within 5 seconds are throttled (see below) |
| Export CSV | exports the current view as a session-level long table (the four billing buckets, own/recorded totals, origin flags), with a BOM so Excel opens it directly without mojibake |
The tables:
- Project table — columns: Project / Sessions / Subagents / Inherited / Input / Output / Tokens / Cache rate / Share. Clicking a project name expands that project's session detail (time, type flags,
in X · out Y · cache Z%, total), listing at most 50 sessions. - By model route — columns: Model route / Input / Output / Tokens / Cache rate / Share. Every billed call is attributed to the provider/model that actually served it (preferring the route that the settlement itself declares in
message.source, and otherwise carrying forward the newestrequest/header/request/contextroute). Available only in accurate mode, because fast mode reads no logs and cannot attribute anything. - Every column's meaning is in Table columns above; the Input / Output columns are sums of the four billing buckets, and hovering expands the exact value of each bucket. The cards at the top list the four buckets separately (uncached / cache read / cache write / output), so you can read both the "input vs output" bottom line and the cache-hit situation. Buckets whose value is 0 are not rendered (when a provider does not report cache writes, that bucket is always 0).
- All three row kinds (project, route, session) have the same shape in the view model: each flattens the four buckets, and
totalTokensis the figure the current view actually displays (accurate mode = the deduplicated value, fast mode = the recorded value), so "input + output = total" holds on every row. - Every view on the page is aggregated locally from the session-level detail the host returns, and cross-checked against the host's own totals; a disagreement shows a warning in the footer instead of being silently ignored.
The two containers differ twofold in width, so the layout is responsive
The same component is dropped into two places whose widths are completely different:
| Entry point | Container | Width |
|---|---|---|
| Sidebar icon → centre-column panel | the app's main column | varies with the window (the author measured roughly 960px and up on a wide screen) |
| Settings → Token ledger | the Settings dialog's content column | ≈564px |
564px is a measured figure, not a guess — the Settings dialog's own CSS is width:800px, its left nav is width:188px, and the content area is padding:0 24px 24px, so 800 − 188 − 48 = 564.
A 9-column table naturally wants about 716px, so inside Settings it gets clipped (which is why "the layout is wrong when you go in through Settings"). The Settings dialog is the app's own shell, fixed at 800px and not draggable, and the plugin cannot change it — it can only make the content adapt. The approach is to turn .dtl-section into a CSS container query container (container-type:inline-size) and step the layout down by container width:
| Container width | Columns | Estimated width |
|---|---|---|
| > 760px | all 9 columns | ~716px |
| ≤ 760px | hide the Share bar (header and cells hidden together, leaving no empty column) | ~620px |
| ≤ 640px | additionally hide the two diagnostic columns "Subagents" and "Inherited" | ~490px |
Checked against the measured 564px container: it lands in the third tier, which needs ~490px, so there is 74px of headroom and no horizontal scrolling is required. The headroom is 74 / 80 / 84px per tier, enough to absorb error in the column-width estimates.
Fallback: every table is wrapped in an overflow-x:auto layer — in a genuinely narrow case you get a horizontal scrollbar rather than clipped content; and a browser without container query support ignores these rules and falls straight through to the fallback, so nothing breaks.
Columns that are hidden still appear in the host panel, and their full values are always in the hover tooltips and in the exported CSV.
As an aside: the card grid's minimum width drops from 180px to 150px, so that at 564px the two card rows hold 3 cards each and the third one does not end up stranded on a line of its own.
How the cache rate is displayed
The definition is deliberately identical to DSH's own: cache read ÷ total billed input, i.e. cacheRead / (uncached + cache read + cache write). That way the panel's figure is necessarily the same as the cache hit displayed by the chat's session stats and this-turn usage, and two differently computed values are never both called "cache rate".
Three details are copied from DSH's implementation (they were not invented here):
- A partial hit is never displayed as 100%. If rounding at the usual precision would yield 100 (99.996%, say), it automatically raises the number of decimal places to stay honest — a test asserts that 999999/1000000 displays as
99.9999rather than100. - With no billed input it displays
—instead of making up a 0%. An entry that only produces output and has no prompt should not be displayed as "cache missed everything". - One decimal place, with a trailing 0 dropped:
93.0%displays as93%,97.6%stays97.6%. This matches the precision of the chat's this-turn usage panel.
Where it appears:
| Where | Form |
|---|---|
| Top cards | Cache hit rate 94.4%, with the numerator/denominator spelled out in the subtitle (cache read 944M / input 1.00B) |
| Both tables | a dedicated Cache rate column, showing the exact numerator/denominator on hover |
| Session detail | in X · out Y · cache Z% on every row |
Why give it a whole column instead of just an overview number: the cache rate is the cost lever, and it varies enormously between projects and routes — the same corpus contains routes at a 99% hit rate and routes at only 44%, a difference of more than 2× in per-unit input cost; looking only at totals never reveals this.
Units are 亿/万 in Chinese or B/M/K in English, switching with the UI language (the plugin registers its own locale dictionary, dsh-token-ledger, one copy each for Chinese and English).
Two layers of request protection
- Same-mode concurrency coalescing: concurrent requests in the same mode share a single tally instead of each running its own.
- Forced-refresh throttling: the
?refresh=1full re-read has a 5-second cooldown. The first one always runs (and starts the cooldown from there); later refreshes inside the cooldown still return a result, just from the cache, and flagrefreshThrottledin the payload.
Why the second layer is needed: one full tally reads the entire corpus (about 70 MiB of zstd at sample scale). This route is not protected by the /api auth gate (no custom DSH route is; even the official open-in-app only gets it by opting in), so any process or page on the local loopback can hit it. The throttle confines both misuse such as "hammering refresh" and a malicious loop to one run per cooldown window.
Tests
npm test # equivalent to node --test test/*.test.mjs
npm run verify:live # end-to-end acceptance against a running DSH (see below)
npm run verify:live
The host half only reloads when DSH restarts (Node caches ESM modules by URL), so every host-side change needs an acceptance run. This command turns that into a single instruction:
- fetch the live payload and check that
schemaVersion,sessions[]androutesare the new version (otherwise it says "host not reloaded, please restart" straight away); - compare the live result session by session against a local fold of the same logs (sessions still being written are excluded by mtime and listed);
- verify route totals = grand total, project totals = grand total, and
reused + re-read = session count; - fetch once more and require
sessionsRead === 0— direct evidence that the incremental cache is working; - finally confirm that fast mode still works.
Failures are listed one by one and the command exits with code 1. The address is taken from --url, else the DSH_WEB_URL environment variable (a DSH session exports it), else the desktop default port.
Eight test groups:
| File | Cases | Content | Needs a local corpus |
|---|---|---|---|
test/fold.test.mjs |
11 | Fold semantics: synchronous replacement, llm/retry-started retiring an attempt, inherited-prefix skipping, route attribution, byRoute totals equal to the grand total |
No (synthetic events) |
test/aggregate.test.mjs |
8 | Aggregation invariants: project totals = grand total, route totals = grand total, overlap = un-deduplicated − deduplicated, count classification, fast-mode shape | No |
test/sweep.test.mjs |
10 | Incremental caching and lease discipline (driven by a fake ctx): an unchanged revision is not re-read, only sessions that changed are re-read, sessions not yet durable are read every time, refresh clears the cache, every observation is disposed, a failed read is not cached but retried next round, the cache is pruned to the corpus, one session's failure does not abort the whole round, fast mode reads no logs |
No (fake ctx) |
test/runner.test.mjs |
6 | Request layer: same-mode concurrency coalesces into one tally, the two modes do not coalesce with each other, a forced refresh inside the cooldown is throttled and still returns a result, and it is let through once the cooldown ends | No (fake ctx + controllable clock) |
test/client-view.test.mjs |
8 | Client-side derivation logic, the old-host degradation path, CSV escaping, row-shape consistency | No |
test/cache-hit.test.mjs |
9 | Cache hit rate definition: the denominator is the input side rather than the total, no billed input returns empty, a full hit is 100, precision escalation on a partial hit, decimal-place rounding | No |
test/client-host-agreement.test.mjs |
3 | The two halves must arrive at the same numbers on the same real corpus (project/route/total/CSV aligned item by item) | Yes |
test/projection-parity.test.mjs |
4 | Session-by-session comparison against DSH's own persisted projection | Yes (skips automatically when the corpus is missing; DSH_HOME can point at it) |
One of these deserves a mention of its own: "every observation is disposed". observeSession returns a lease, and failing to dispose it pins the entire parsed log forever (the query service's cache capacity is 5, and a pinned entry is never evicted). So this invariant has a dedicated assertion.
The key trick in the equivalence tests: fold the log only up to the water mark of that session's tokenUsage row (row.seq) before comparing. A full fold also counts events the projection has not observed yet — which is normal for a session being written to — so once the comparison is aligned to the water mark the test is deterministic and does not flake because of concurrent writes.
Data sources
The host half uses only public services:
ctx.sessionQuery.listSessions()— lists every session (includingheader.cwd).ctx.sessionQuery.observeSession(id, { projectionMode: 'all' })— reads one session's events and projections; the lease is released immediately afterwards (not releasing it pins an entire parsed log forever).ctx.sessionProjectionCache.cachedSnapshot(header, ['tokenUsage'])— the zero-I/O read that fast mode uses.
Composition
dsh-token-ledger/
├── package.json # dsh.bundle.patch + dsh.client declarations
├── cordis.patch.yml # the single row: the host half
└── lib/
├── index.js # host half: aggregation + GET /token-ledger/summary
└── client.js # client half: the settings.section panel
The data channel is a same-origin HTTP JSON route rather than a generated Remote: that way the client half needs no code generation and does not have to pull in DSH-internal packages.
Install / Uninstall
You need a working DSH installation (which provides plugin_manager and the dsh plugin CLI) and a checkout of this repository.
Registering it with one profile can be done either way:
- Through the plugin manager — have the agent call
plugin_manager'sinstall_bundlewith this repository's directory as the target, or do the same from the Web Settings → Plugins page. - Manually — add this package directory as a dependency of that profile (the package manager writes it as a
link:), and list it in the profilepackage.json'sdsh.profile.bundles.
Once registered it appears under Settings → Plugins, and can be disabled or uninstalled at any time.
Registration as a local bundle is two declarations and nothing else:
the profile depends on this package through a
link:to the package directory, so the profile loads the code straight from the checkout instead of from a published copy;package.json'sdsh.bundle.patchpoints atcordis.patch.yml, which inserts exactly one row:- insert: - id: token-ledger name: dsh-token-ledger
That single row is the host half. The client half needs no row of its own: the module scanner finds it through the package's dsh.client declaration and serves it as part of the same plugin.
To change this plugin's own code, see CONTRIBUTING.md — it covers how each half reloads, and the repository's commit conventions.
Known limits
- Only sessions whose logs are still on disk are counted. Sessions whose logs were deleted or archived do not appear —
listSessions()goes by the durable list. - Sessions with no
cwdare grouped into a single "(no project directory)" row and still counted in the totals. - Accurate mode is O(all log bytes): a full cold read takes about 3 seconds and cannot be parallelised (
observeSessionhas no batch interface). As the number of sessions grows, the time rises linearly. - The numbers are not a bill: they are the usage the provider reported, and DSH performs no price conversion.
License
MIT © u-fw · Repository: github.com/u-fw/dsh-token-ledger
No comments yet. Be the first to write one.