DSH HUB
HomePlugin StorePlugin PacksCommunityRankingsResourcesPublish Guide
Plugin source
Back to catalog

Soulize /

Soulize/dsh-continue-on-limit-host

Verified

Auto-continue for DeepSeek Harness: when a local model hits its output-token cap, automatically send "continue" so the reply keeps flowing

★ 1 Stars0 Forks0 IssuesN/A Community rating0 Confirmed installs
View on GitHub
READMESource: main@a4e2fea1

dsh-continue-on-limit-host

DeepSeek Harness / DSH 的 Host 侧全局 max-tokens 自动续写插件。

模式

插件提供两个严格互斥的实现:

Follow-up(默认)

provider finish=max-tokens
        ↓
AgentLoop 正常记录 max-tokens
        ↓
agent/turn-stopping
        ↓
agent.followup("继续")
        ↓
下一 Turn

Follow-up 保留 DSH 原生语义:

  • Assistant stream 中仍是 finish=max-tokens
  • 当前 Turn 仍是 turn/end(max-tokens)
  • “继续”进入 next-turn

Steer(透明同 Turn)

0.4.0 起,Steer 不再等 max-tokens 进入 AgentLoop 后补救,而是在官方 llm/stream waterfall 直接截获终止 chunk:

Provider
finish=max-tokens
        ↓
dsh-continue-on-limit-host
        ↓
max-tokens 仅作为插件内部触发信号
        ↓
需要时 agent.steer("继续") -> next-step
        ↓
对 AgentLoop 输出 finish=stop
        ↓
AgentLoop 从未进入 sticky max-tokens 状态

0.4.2 起,Steer 对“max-tokens 时已经输出的 tool-call”提供两种可选策略。

discard:0.4.1 安全策略(默认)

DSH 自己的 BlockAssembler 在 finish=max-tokens 时会丢弃该响应中的全部 tool-call。这个模式保持相同语义:

Provider finish=max-tokens
+ zero or more tool-call blocks
        ↓
从第一个 tool-call 起暂存后续 stream
        ↓
最终确认 max-tokens
        ↓
丢弃该响应中的全部 tool-call chunks
保留 text / reasoning
        ↓
agent.steer("继续")
        ↓
对 AgentLoop 输出 finish=stop

若最终不是 max-tokens,暂存内容会按原顺序全部释放,所以正常 finish=tool-calls 不受影响。

passthrough:0.4.0 兼容策略

这个模式不暂存 tool-call。chunk 会直接进入 Harness:

tool-call chunks
        ↓
直接进入 Harness

最终 finish=max-tokens
        ↓
如果此前至少有一个 tool-call 已 block-end
    → 对外改成 finish=tool-calls
    → 不额外 steer
否则
    → 对外改成 finish=stop
    → agent.steer("继续")

它的用途是保留模型在达到输出上限前已经生成的工具调用,避免续写后的模型认为该调用已经发生而跳过它。

这是 0.4.0 的兼容行为:因为 tool-call chunk 已经提前交给 Harness,最终发现 max-tokens 后无法再撤回,因此它会绕过 DSH 原生的“max-tokens 一律丢弃 tool-call”策略。用户需要自行选择这一取舍。

两种策略都满足同一个目标:

  • AgentLoop 不会看到 FinishReason { kind: "max-tokens" }
  • 持久化 assistant/message.stream 中不会记录 max-tokens
  • 不会触发 DSH 的 sticky turnEnds=max-tokens
  • 不需要修改 Session.append()
  • 不需要改写 turn/end
  • 不需要 tool-loop bridge

对 shipped pi-ai adapter,discard 会把 replay stopReason: "length" 改成 stop 并删除 tool-call replay entries;passthrough 会按最终伪装结果改成 stop 或 toolUse。DeepSeek Messages 的 replay metadata 本身不保存 stop reason。

配置

enabled: true
continuationMode: followup   # followup | steer
steerToolCallPolicy: discard  # discard | passthrough
continueText: 继续
maxConsecutive: 3
minIntervalMs: 0
includeSubagents: true
debugLogging: false

字段说明:

字段 默认值 说明
enabled true 总开关
continuationMode followup Follow-up 或透明 Steer
steerToolCallPolicy discard Steer 遇到 max-tokens + tool-call 时:discard=0.4.1 暂存并丢弃;passthrough=0.4.0 兼容透传
continueText 继续 自动 Steer / Follow-up 的文本
maxConsecutive 3 最多连续注入多少次续写;0 不限
minIntervalMs 0 自动注入之间的最小间隔
includeSubagents true 是否处理 Subagent
debugLogging false 详细日志

Steer 模式中,maxConsecutive 限制的是插件主动注入的 agent.steer() 次数。discard 每次截获 max-tokens 都会尝试 Steer;passthrough 在已经有 closed tool-call 时改走原生工具执行,不额外消耗一次 Steer 计数。

配置为 volatile,可在 Plugins 页面热更新;模式切换会清空上一模式的连续计数和已处理状态。

Steer 为什么要在 llm/stream 做

DSH AgentLoop 对 max-tokens 有 sticky 语义:

const stepEnd = await this.step(decision)

if (turnEnds === null || turnEnds.kind !== 'max-tokens') {
  turnEnds = stepEnd
}

一旦某个 Step 把 turnEnds 设成 max-tokens,后面的正常 tool-calls / completed 都不能覆盖它。

所以 Steer 0.4.0 的策略是:不让这个标识进入 AgentLoop。

插件只在 llm/stream 内部看到真实 max-tokens,然后按 steerToolCallPolicy 把对外终止原因替换为普通 stop 或 tool-calls。这样 AgentLoop 不会进入 sticky max-tokens 状态。

Follow-up 检测

Follow-up 仍在 agent/turn-stopping 读取最新 Assistant settlement:

latest assistant/message | assistant/attempt
        ↓
lastAssistantStreamChunk(stream, "finish")
        ↓
finish.reason.kind === "max-tokens"

然后调用 agent.followup()。

队列处理

自动注入前会检查:

agent.inbox.nextTurn
agent.inbox.nextStep

已有待处理工作时不再重复插入。

Steer 已截获的 max-tokens 即使因为 inbox、次数上限或运行中配置变化而没有再注入,也不会重新暴露给 AgentLoop。

Subagent

includeSubagents=true 时,live Subagent 与主 Agent 使用同一策略。

  • Follow-up:max-tokens Turn 结束前排入 next-turn
  • Steer:max-tokens 在模型流边界被隐藏,同 Turn 按原生 Step/tool loop 继续

Steer 模式最终正常完成时,Subagent 看到的是普通 completed Turn,而不是 max-tokens settlement。

日志

日志:

$DSH_HOME/logs/dsh-continue-on-limit-host.log

Windows 通常为:

%USERPROFILE%\.dsh\logs\dsh-continue-on-limit-host.log

Follow-up 重点标记:

MAX_TOKENS_PRESTOP_CAPTURED
FOLLOWUP_MAX_TOKENS_CAPTURED
FOLLOWUP_QUEUED

Steer 重点标记:

STEER_MAX_TOKENS_INTERCEPTED ... policy=discard
STEER_MAX_TOKENS_INTERCEPTED ... policy=passthrough ... mask=tool-calls
STEER_QUEUED
STEER_SEND_SKIPPED reason=closed-tool-call
TURN_END_CAPTURED ... reason=completed

安装

dsh plugin --profile web add github:Soulize/dsh-continue-on-limit-host

更新时建议固定到最新 commit:

dsh plugin --profile web add github:Soulize/dsh-continue-on-limit-host#<commit-sha>

自检

npm test

Smoke test 覆盖:

  • Follow-up 保留 max-tokens
  • Steer 将 max-tokens 映射为 stop
  • discard:max-tokens 响应中的 tool-call 全部丢弃
  • passthrough:closed tool-call 在 max-tokens 下按 0.4.0 行为保留
  • 正常非截断 tool-calls 原样通过
  • pi-ai replay length 同步隐藏
  • Steer 次数上限下仍不向 AgentLoop 暴露 max-tokens
  • Follow-up / Steer 热切换隔离
  • Plugins 配置页注册

兼容性

  • DSH:>=0.1.7-rc.1
  • 不设置最高 DSH 版本上限
  • Node.js:>=20

关键接口:

  • llm/stream
  • isAgentLoopRequest()
  • agent/turn-stopping
  • Agent.followup()
  • Agent.steer()
  • Agent.inbox
  • session/event
  • volatile Config
  • plugins.row.config

License

MIT

—/ 5

No ratings yet

Verified DSH bundle

Commit a4e2fea17b10

Community comments

No comments yet. Be the first to write one.

DSH HUB

A community index for DSH plugins. Not an official GitHub or DeepSeek AI product.

CommunityResourcesAPIAbout