DSH HUB
HomePlugin StorePlugin PacksCommunityRankingsResourcesPublish Guide
Plugin source
Back to catalog

sakuraboy9128-cmd /

sakuraboy9128-cmd/dsh-sensevoice-npu

Verified

DSH speech-to-text provider that runs SenseVoiceSmall on the Intel NPU via OpenVINO

★ 0 Stars0 Forks0 IssuesN/A Community rating0 Confirmed installs
View on GitHub
READMESource: main@301f8afd

dsh-sensevoice-npu

Local speech-to-text for DSH that runs SenseVoiceSmall on the Intel NPU through OpenVINO, plugged into the official voice-input stack.

It registers one more provider into the stock @deepseek-ai/dsh-experimental-speech-to-text service, so the voice-input button, the settings card, the preparation UI and the speech API keep working unchanged — only the recognizer behind them changes.

Why a separate plugin

The bundled provider (dsh-experimental-speech-to-text-sensevoice) is CPU-only by construction:

  • lib/worker.js hard-codes provider: "cpu" for the recognizer and the VAD;
  • the shipped sherpa-onnx-win-x64 package is a CPU-only ONNX Runtime build, so sherpa-onnx's openvino provider logs "OpenVINOExecutionProvider is not available … Fallback to cpu";
  • ONNX Runtime publishes no OpenVINO build for Windows;
  • and OpenVINO crashes the NPU compiler on this model's default graph (unbounded dynamic dimensions from the x_length attention mask).

This plugin sidesteps all four: it bakes x_length into a constant for one frame bucket (which the NPU compiler accepts), compiles that graph for the NPU, and serves it through a small Python worker.

Layout

Path Role
lib/index.js cordis plugin: registers the provider, spawns the worker, runs preparation
lib/worker.py authenticated loopback HTTP worker (one WAV per request)
lib/sensevoice_npu.py frontend (fbank + LFR + CMVN), Silero VAD, NPU session, CTC decode
runtime/bootstrap.ps1 creates the virtual environment the worker needs

The frontend mirrors sherpa-onnx exactly: kaldi_native_fbank with the same options, ApplyLfr (7/6), and (x + neg_mean) * inv_stddev with the CMVN vectors read from the ONNX metadata — so transcripts match the CPU provider.

Install

Prerequisite: model files

The worker reads three files (this plugin does not download models — it registers with an empty downloadSources):

Path relative to dataRoot Size
models/sensevoice-onnx/model.int8.onnx ~228 MB
models/sensevoice-onnx/tokens.txt ~0.3 MB
models/silero/silero_vad.onnx ~1.7 MB

The official dsh-experimental-speech-to-text-sensevoice provider downloads them into the same dataRoot (<DSH home>/speech-to-text/sensevoice), so let that provider prepare once, or place the files at those paths yourself.

  1. Build the Python environment (needs a 64-bit Python 3.11–3.13 and the Intel NPU driver):

    pwsh -File runtime/bootstrap.ps1
    
  2. Install the package into the profile that owns the voice-input bundle (the DSH desktop app uses the desktop profile):

    dsh plugin --profile desktop add "file:C:/path/to/dsh-sensevoice-npu"
    

    pnpm copies the package, so while iterating on it replace the copy with a junction to your working tree:

    $d = "$env:USERPROFILE\.dsh\profiles\desktop\node_modules\dsh-sensevoice-npu"
    Remove-Item -Recurse -Force $d
    New-Item -ItemType Junction -Path $d -Target C:\path\to\dsh-sensevoice-npu
    
  3. The bundle's own cordis.patch.yml already selects the provider and registers the plugin, and it contains no machine-specific path: pythonPath and dataRoot default to the bootstrap.ps1 venv and DSH's speech root. Append an override to ~/.dsh/profiles/desktop/cordis.patch.yml (back it up first) only when you moved either one:

    - id: speech-to-text-npu
      name: 'dsh-sensevoice-npu'
      config:
        pythonPath: <custom venv>\Scripts\python.exe   # only with -VenvDir
        dataRoot: <custom speech root>                 # only with a non-default DSH_HOME
    
  4. Restart DSH. The voice-input card then lists SenseVoiceSmall · Intel NPU (NPU) and uses it by default; the first preparation compiles the NPU graph once (~2.5 min), after which starts take ~0.2 s. If the card still shows the CPU provider, pick the NPU one in the voice-input settings (an explicit choice persists over the patch default).

Configuration

Key Default Meaning
pythonPath %LOCALAPPDATA%\dsh-sensevoice-npu\venv\Scripts\python.exe worker interpreter (runtime/bootstrap.ps1 target)
dataRoot <DSH home>/speech-to-text/sensevoice DSH speech root; models are read from <dataRoot>/models
precision int8 int8 (model.int8.onnx) or fp32 (model.onnx)
device NPU OpenVINO device for the recogniser (NPU, GPU, CPU)
vadDevice CPU Silero VAD device — tiny model, dynamic shapes, NPU rejects it
useVad true trim silence and split long speech into segments
buckets [64, 96, 128, 192, 256, 384, 512, 768, 1024] (plugin config sets [256]) LFR-frame buckets the engine may compile
prewarmBuckets [256] buckets compiled before the worker reports ready
idleTimeoutMs 600000 release the worker after idling; compiled blobs stay cached

Each bucket costs roughly 800 MB on disk (baked IR + OpenVINO NPU cache) and one ~2.5 minute compile the first time. One bucket is enough for dictation: 256 frames ≈ 15 s of audio, and the NPU handles the padding cheaply.

Measured behaviour (Core Ultra X7 358H, Arc B390, NPU driver 32.0.100.5540)

Case Result
Warm start (compiled blob cached) 0.21 s
4.5 s Chinese clip → 256-frame bucket 今天天气不错,我们去公园散步吧, 0.66 s inference
5.2 s clip 明天上午10点开会,记得带上项目进度报告, 0.48 s
14.5 s clip full sentence correct, 0.65 s
First compile per bucket 100–145 s, cached afterwards

CPU is used only for the frontend (fbank + VAD, a few ms) and JSON plumbing; the encoder/decoder graph runs on the NPU. Note that the NPU is not faster than sherpa-onnx on CPU for clips this short (~0.5 s vs ~0.2 s) — the win is moving the heavy work off the CPU/GPU entirely.

Development

node --test        # 6 tests, no dependencies

They cover the registration contract (provider id, device label, diagnostic file), the portable pythonPath/dataRoot defaults, and the two config guards that would otherwise fail much later: an empty bucket list and a relative path.

Troubleshooting

  • Confirm the profile really loaded the plugin and see the last phase/device:

    Get-Content "$env:USERPROFILE\.dsh\speech-to-text\sensevoice\npu\provider-status.json"
    
  • NPU device not visible → update the Intel NPU driver (32.0.100.5540 tested).

  • NPU compile failed → delete <dataRoot>/npu and retry; check free disk (~1 GB per bucket).

  • Transcription slow on first use of a new bucket → it is compiling; later runs reuse the cache.

DSH HUB

A community index for DSH plugins. Not an official GitHub or DeepSeek AI product.

CommunityResourcesAPIAbout
—/ 5

No ratings yet

Verified DSH bundle

Commit 301f8afdb30c

Community comments

No comments yet. Be the first to write one.