dsh-llm-fidelity
Your AI agent breaks mid-task because a model rejected one optional parameter. Your usage dashboard shows numbers that don't add up. dsh-llm-fidelity fixes both — automatically, invisibly, and honestly.
A DeepSeek Harness plugin. Install it, change nothing else.
What changes after you install it
A real scenario, before and after:
You asked your Agent to fix a bug. Three steps in, the model service suddenly returns:
400 — 'stop' is not supported with this model.Without the plugin — the conversation dies on the spot. You dig through provider docs, edit a config, start a new session, and re-explain the whole task. Switch providers next week? Same wall, different parameter.
With the plugin — that 400 is absorbed in milliseconds: the offending parameter is dropped and the identical request re-sent. What you see is the answer simply continuing. And the plugin remembers this provider's temperament — the next request never even hits that error. You never learn the error existed.
The same promise applies to 'max_tokens' is not supported…, 'temperature' is not supported…, reasoning effort is not supported…, and friends — plus the quieter pain of usage statistics being wiped to zero or counted twice by gateways. All of it stops at the plugin.
- 🔧 Conversations that don't break — a rejected parameter is fixed and re-sent within the same request. You never see the error.
- 💯 Numbers you can trust — usage statistics land in your accounting exactly as the provider reported them. No wiping, no double-counting, no invented figures.
- 🚀 Zero config — smart defaults work out of the box; the plugin registers no model routes and never conflicts with official or community adapters.
Sound familiar?
If you have ever pointed an AI coding agent at a third-party model or gateway, you have almost certainly met one of these:
Unsupported parameter: 'max_tokens' is not supported with this model.
Use 'max_completion_tokens' instead.
Unsupported parameter: 'stop' is not supported with this model.
'temperature' is not supported with this model.
reasoning effort is not supported by this model
These are real, current, and everywhere. OpenAI's reasoning models (GPT-5 and GPT-6 series) reject max_tokens, stop, and temperature — and the rejections hit users through LiteLLM, Azure OpenAI, crewAI, and every OpenAI-compatible gateway in between. Passing stop: null does not help (crewAI #1908). The max_tokens → max_completion_tokens rename even carries a trap: the new field includes hidden reasoning tokens, so a naively converted budget returns empty responses that you still pay for.
What actually happens to you: your agent is three tool calls into a task, hits one of these errors, and the whole turn dies. You dig into provider docs, edit a config, restart, and pray the next provider doesn't reject something else.
And the quieter problem — your usage stats. Gateways re-send a zeroed usage snapshot at stream end, wiping the real numbers. The same snapshot arrives twice and gets counted twice. Cache-hit fields go missing. Your dashboard says 3.2M tokens; your invoice says 5.1M; nobody can explain the difference.
What dsh-llm-fidelity does about it
When a model explicitly rejects an optional parameter, the plugin fixes it and re-sends — within the same request, in milliseconds, invisibly.
You send: { temperature: 0.7, stop: ["END"], ... }
Model: 400 — 'temperature' is not supported with this model.
Plugin: (removes temperature, re-sends immediately)
Model: 200 — here is your answer.
You see: the answer. Nothing else. Ever.
The fix list is deliberately short: temperature, reasoningEffort, stop, maxTokens (omitted, or clamped down when the error states an explicit output limit). Each rejected field is read from the provider's own error message — no guessing, no brute-force retries. And the plugin remembers which provider accepts which combination (1 hour, in-process), so the failure happens at most once per route. The second request already speaks that provider's dialect.
For the usage numbers, the plugin guards the statistics stream frame by frame: a later frame that omits fields can't erase earlier ones, an all-zero placeholder can't wipe real totals, a duplicate can't be counted twice. And one rule above all: if a number was never reported, the plugin never invents it. One missing reference number beats a fake one.
Why this one
There are a dozen ways to paper over a 400 error. This plugin chooses the boring, correct ones — on purpose:
| dsh-llm-fidelity | |
|---|---|
| Fixes only what the provider explicitly rejected | ✅ Reads the rejection from the error itself; never guesses. Context-overflow, quota, auth, and network errors are passed through honestly — a real problem should reach you. |
| Never touches your content | ✅ Messages, system prompt, tools, and constraints are untouchable. Only optional scalars (temperature / reasoningEffort / stop / maxTokens) may be dropped or clamped. |
| Bounded, no magic | ✅ At most 4 adaptations per request, each of which must actually change the request. No infinite retry loops, no silent protocol switching. |
| Learns | ✅ Per provider+model, in-process, 1-hour TTL. Fail once, never again on that route. |
| Plays well with others | ✅ Registers zero model routes and replaces zero services. Official adapters, community adapters, and the official dsh-llm-retry keep working untouched — 429 rate limits and server errors are its job, and this plugin steps aside for them. |
Install
Compatibility: works with DSH
0.1.0-rc.6and later, across the whole0.2.0-rcline (tested against DeepSeek desktop 0.2.0-rc.2). If your DSH reports an incompatibility on install, check your version first.
dsh plugin add dsh-llm-fidelity
or any npm-based workflow:
pnpm add dsh-llm-fidelity
Then enable it in the plugin settings. Defaults are production-ready; every switch is optional:
| Option | Default | Meaning |
|---|---|---|
usageFidelity |
true |
usage guarding (dedup, anti-wipe, field merging) |
rectify |
true |
automatic parameter-rejection adaptation |
rectifyStatuses |
[400, 422] |
statuses treated as explicit parameter rejections |
maxAdaptations |
4 |
maximum automatic re-sends per request |
strategyTtlMs |
3600000 |
how long a learned provider preference is kept (ms) |
strategyCacheSize |
256 |
maximum learned preferences kept in memory |
rectifyAuxiliary |
false |
whether background tasks (e.g. auto titling) are adapted too |
Changing the configuration
The plugin installs as a profile bundle layer, and its configuration lives in the profile's patch file (on the desktop app: ~/.dsh/profiles/desktop/cordis.patch.yml). Override the row by its id llm-fidelity:
- id: llm-fidelity
config:
maxAdaptations: 2
rectifyAuxiliary: true
Note: a patch's config replaces wholesale rather than merging — restate every key you want to keep (keys you omit fall back to the built-in defaults).
How it works (for the curious)
The plugin listens on the llm/stream event — the one road every model request travels. It registers no model routes and replaces no service; it adds two checks beside the road:
- Parameter adaptation: watches for explicit parameter rejections, adjusts the offending optional field, and re-enters the queue through the public
ctx.llm.stream()entry (the exact path a normal request takes, so internal state stays consistent). A re-send may only happen while nothing has been delivered downstream — once content reached you, errors pass through verbatim. No replay, no fabricated success. - Usage guarding: walks the statistics frames, merges missing fields, drops duplicates and placeholders, and hands the final numbers to the accounting layer right before the turn closes.
Honest boundaries
- The plugin works at the unified data layer and sees standardized frames only. If a service never reported a number (say, cache hits), it cannot be recovered — this plugin keeps data clean, it cannot conjure data.
- OAuth-based services, oversized contexts, and exhausted balances are out of adaptation scope by design.
- The session log records the original request parameters; adaptation is transparent, and replay provenance follows the winning attempt.
License
MIT
No comments yet. Be the first to write one.