DSH HUB
HomePlugin StorePlugin PacksCommunityRankingsResourcesPublish Guide
Plugin source
Back to catalog

d3vmeh /

d3vmeh/dsh-llm-gate

Verified

Per-provider concurrency gate for DeepSeek Harness model requests

★ 1 Stars0 Forks0 IssuesN/A Community rating0 Confirmed installs
View on GitHub
READMESource: main@2a7acb42

dsh-llm-gate

Per-provider concurrency gate for DeepSeek Harness model requests.

If a provider can only serve a fixed number of requests at once (e.g a local llama-server with --parallel 1), every extra request is deferred by the server with nothing sent back. The client cannot tell "waiting for a slot" from "dead", and Node HTTP layer times out after 300 seconds with terminated. In practice this happens when there is overlap between a subagent and the main agent or compaction and the agent.

This plugin holds surplus requests inside dsh instead. A request waits in a FIFO queue before any HTTP request is made so no timeout is running while it waits. When a slot frees, the next request is dispatched.

Install

dsh plugin --profile web add dsh-llm-gate

Then configure the providers to gate in ~/.dsh/profiles/web/cordis.patch.yml:

- id: llm-gate
  config:
    providers:
      llamacpp:
        maxConcurrent: 1
        maxQueued: 16
        queueTimeoutMs: 3600000

The provider key is the route name from your llm-pi-ai.providers (or other adapter) settings. Providers not listed are not gated. Restart dsh web and open a new session.

Check the composed config with dsh --profile web --dump-config.

Settings

Setting Required Meaning
maxConcurrent yes Requests allowed in flight to this provider. For llama.cpp, match --parallel.
maxQueued no Requests allowed to wait. Beyond this, a request fails at once with QUEUE_FULL. Default: unlimited.
queueTimeoutMs no Longest a request may wait for a slot before failing with QUEUE_TIMEOUT. Default: wait indefinitely.

Queue failures end the turn with the code shown. They are not retried by dsh-llm-retry.

What you will see

The plugin prints a line to the dsh terminal only when a request has to wait:

llm-gate: llamacpp session=a61e6e40 queued (depth 1)
llm-gate: llamacpp session=a61e6e40 dispatched after 5730ms

purpose=compaction or purpose=session-title is added for auxiliary requests. Requests that get a slot immediately print nothing.

Notes

  • This gate serializes requests so it does not make a single-slot server faster. For parallelizing, give llama.cpp more slots (--parallel 2 --kv-unified) and raise maxConcurrent to match.
  • Waiting time is not counted by the adapter's streamIdleTimeoutMs because the adapter is not called until the slot is acquired. You still need streamIdleTimeoutMs large enough for your prompt processing time (see the llm-pi-ai provider settings).
  • A queued request is cancelled through its abort signal. Dropping the stream without aborting leaves the request queued until a slot frees, at which point it dispatches and is closed immediately.
  • Requires the llm service; hooks the llm/stream waterfall, so it covers every model request in the host: agents, subagents, compaction, and title generation.

License

MIT

—/ 5

No ratings yet

Verified DSH bundle

Commit 2a7acb42a7dd

Community comments

No comments yet. Be the first to write one.

DSH HUB

A community index for DSH plugins. Not an official GitHub or DeepSeek AI product.

CommunityResourcesAPIAbout