dsh-sensevoice-npu
Local speech-to-text for DSH that runs SenseVoiceSmall on the Intel NPU through OpenVINO, plugged into the official voice-input stack.
It registers one more provider into the stock
@deepseek-ai/dsh-experimental-speech-to-text service, so the voice-input
button, the settings card, the preparation UI and the speech API keep working
unchanged — only the recognizer behind them changes.
Why a separate plugin
The bundled provider (dsh-experimental-speech-to-text-sensevoice) is CPU-only
by construction:
lib/worker.jshard-codesprovider: "cpu"for the recognizer and the VAD;- the shipped
sherpa-onnx-win-x64package is a CPU-only ONNX Runtime build, so sherpa-onnx'sopenvinoprovider logs "OpenVINOExecutionProvider is not available … Fallback to cpu"; - ONNX Runtime publishes no OpenVINO build for Windows;
- and OpenVINO crashes the NPU compiler on this model's default graph
(unbounded dynamic dimensions from the
x_lengthattention mask).
This plugin sidesteps all four: it bakes x_length into a constant for one
frame bucket (which the NPU compiler accepts), compiles that graph for the NPU,
and serves it through a small Python worker.
Layout
| Path | Role |
|---|---|
lib/index.js |
cordis plugin: registers the provider, spawns the worker, runs preparation |
lib/worker.py |
authenticated loopback HTTP worker (one WAV per request) |
lib/sensevoice_npu.py |
frontend (fbank + LFR + CMVN), Silero VAD, NPU session, CTC decode |
runtime/bootstrap.ps1 |
creates the virtual environment the worker needs |
The frontend mirrors sherpa-onnx exactly: kaldi_native_fbank with the same
options, ApplyLfr (7/6), and (x + neg_mean) * inv_stddev with the CMVN
vectors read from the ONNX metadata — so transcripts match the CPU provider.
Install
Prerequisite: model files
The worker reads three files (this plugin does not download models — it registers with an empty downloadSources):
Path relative to dataRoot |
Size |
|---|---|
models/sensevoice-onnx/model.int8.onnx |
~228 MB |
models/sensevoice-onnx/tokens.txt |
~0.3 MB |
models/silero/silero_vad.onnx |
~1.7 MB |
The official dsh-experimental-speech-to-text-sensevoice provider downloads them into the same dataRoot (<DSH home>/speech-to-text/sensevoice), so let that provider prepare once, or place the files at those paths yourself.
Build the Python environment (needs a 64-bit Python 3.11–3.13 and the Intel NPU driver):
pwsh -File runtime/bootstrap.ps1Install the package into the profile that owns the voice-input bundle (the DSH desktop app uses the
desktopprofile):dsh plugin --profile desktop add "file:C:/path/to/dsh-sensevoice-npu"pnpm copies the package, so while iterating on it replace the copy with a junction to your working tree:
$d = "$env:USERPROFILE\.dsh\profiles\desktop\node_modules\dsh-sensevoice-npu" Remove-Item -Recurse -Force $d New-Item -ItemType Junction -Path $d -Target C:\path\to\dsh-sensevoice-npuThe bundle's own
cordis.patch.ymlalready selects the provider and registers the plugin, and it contains no machine-specific path:pythonPathanddataRootdefault to thebootstrap.ps1venv and DSH's speech root. Append an override to~/.dsh/profiles/desktop/cordis.patch.yml(back it up first) only when you moved either one:- id: speech-to-text-npu name: 'dsh-sensevoice-npu' config: pythonPath: <custom venv>\Scripts\python.exe # only with -VenvDir dataRoot: <custom speech root> # only with a non-default DSH_HOMERestart DSH. The voice-input card then lists SenseVoiceSmall · Intel NPU (NPU) and uses it by default; the first preparation compiles the NPU graph once (~2.5 min), after which starts take ~0.2 s. If the card still shows the CPU provider, pick the NPU one in the voice-input settings (an explicit choice persists over the patch default).
Configuration
| Key | Default | Meaning |
|---|---|---|
pythonPath |
%LOCALAPPDATA%\dsh-sensevoice-npu\venv\Scripts\python.exe |
worker interpreter (runtime/bootstrap.ps1 target) |
dataRoot |
<DSH home>/speech-to-text/sensevoice |
DSH speech root; models are read from <dataRoot>/models |
precision |
int8 |
int8 (model.int8.onnx) or fp32 (model.onnx) |
device |
NPU |
OpenVINO device for the recogniser (NPU, GPU, CPU) |
vadDevice |
CPU |
Silero VAD device — tiny model, dynamic shapes, NPU rejects it |
useVad |
true |
trim silence and split long speech into segments |
buckets |
[64, 96, 128, 192, 256, 384, 512, 768, 1024] (plugin config sets [256]) |
LFR-frame buckets the engine may compile |
prewarmBuckets |
[256] |
buckets compiled before the worker reports ready |
idleTimeoutMs |
600000 |
release the worker after idling; compiled blobs stay cached |
Each bucket costs roughly 800 MB on disk (baked IR + OpenVINO NPU cache) and one ~2.5 minute compile the first time. One bucket is enough for dictation: 256 frames ≈ 15 s of audio, and the NPU handles the padding cheaply.
Measured behaviour (Core Ultra X7 358H, Arc B390, NPU driver 32.0.100.5540)
| Case | Result |
|---|---|
| Warm start (compiled blob cached) | 0.21 s |
| 4.5 s Chinese clip → 256-frame bucket | 今天天气不错,我们去公园散步吧, 0.66 s inference |
| 5.2 s clip | 明天上午10点开会,记得带上项目进度报告, 0.48 s |
| 14.5 s clip | full sentence correct, 0.65 s |
| First compile per bucket | 100–145 s, cached afterwards |
CPU is used only for the frontend (fbank + VAD, a few ms) and JSON plumbing; the encoder/decoder graph runs on the NPU. Note that the NPU is not faster than sherpa-onnx on CPU for clips this short (~0.5 s vs ~0.2 s) — the win is moving the heavy work off the CPU/GPU entirely.
Development
node --test # 6 tests, no dependencies
They cover the registration contract (provider id, device label, diagnostic file),
the portable pythonPath/dataRoot defaults, and the two config guards that
would otherwise fail much later: an empty bucket list and a relative path.
Troubleshooting
Confirm the profile really loaded the plugin and see the last phase/device:
Get-Content "$env:USERPROFILE\.dsh\speech-to-text\sensevoice\npu\provider-status.json"NPU device not visible→ update the Intel NPU driver (32.0.100.5540 tested).NPU compile failed→ delete<dataRoot>/npuand retry; check free disk (~1 GB per bucket).Transcription slow on first use of a new bucket → it is compiling; later runs reuse the cache.
No comments yet. Be the first to write one.