DSH HUB
HomePlugin StorePlugin PacksCommunityRankingsResourcesPublish Guide
Plugin source
Back to catalog

lpeixin /

lpeixin/dsh-qualityforge

Verified

QualityForge is an enterprise-grade DeepSeek Harness QA plugin that automates systematic end-to-end testing and delivers P0–P3 severity-based, actionable reports and quality gates to help developers identify issues and prioritize fixes. QualityForge 是一个面向企业级项目的 DeepSeek Harness QA 插件,自动执行系统性端到端测试,并通过 P0–P3 严重度分级、可勾选的测试报告与质量门禁,帮助开发者快速识别问题并确定修复优先级。

★ 3 Stars1 Forks2 IssuesN/A Community rating0 Confirmed installs
View on GitHub
READMESource: main@0f27bf4d
QualityForge — systematic QA for finished projects

English | 简体中文

License DSH Plugin Node Tests Catalog


QualityForge is a DeepSeek Harness plugin that runs a systematic, enterprise-grade test pass over a finished project and produces a per-item checkable report with P0–P3 severity — a quality verdict, and the artifact developers and the Harness use to agree on what to fix and in which order.

Question What QualityForge gives you
Can this ship? A verdict: ⛔ Blocked / ⏳ Audit incomplete / ⚠️ Conditional / ✅ Ready
What exactly is wrong? 45 test domains, 290 methods, 1363 independently verifiable checks — each with a conclusion, evidence and a fix suggestion
What first? Fix waves (Wave 1 clears blockers → Wave 2 criticals → …) plus a handoff sheet you can paste back into the Harness
What after fixing? Tick items in the report → qf_update sync=true reads them back → qf_exec re-tests → only then does an item become verified

The report, catalog, tool descriptions and the documents under docs/ are written in Chinese today; an English report locale is on the roadmap. Identifiers (QF-001), statuses and the JSON artifacts are language-neutral.


Why this exists

"The code is written" and "this is shippable" are two different claims, and the second one needs evidence. What usually gets in the way:

  • Not knowing what to test — enterprise test practice lives in specs, checklists and people's heads; improvising leaves blind spots, and the blind spots are exactly where releases break.
  • No trace of what was checked — commands get run, conclusions evaporate, and two weeks later nobody can say what was actually verified.
  • Verdicts that cannot be negotiated — reports read like sentences; there is no structured way to say "this must be fixed, that risk we accept".
  • Unclear fix scope — with dozens of findings, which come first, which are in this round, and how do we know a fix is really done?

QualityForge turns that into a repeatable pipeline: recon → exhaustive checklist → run what can be run → agent analyses the rest → checkable report → agree the fix scope → re-test to close. Every intermediate artifact stays inside the audited project under .qualityforge/, so the report can be re-rendered, reviewed in a pull request, and diffed.


Contents

  • Why this exists
  • Features
  • Install
  • Quick start
  • Usage
  • Tools
  • Coverage
  • Artifacts
  • Security boundaries
  • FAQ
  • Development
  • Compatibility
  • Roadmap
  • License

Features

Feature Detail
Enterprise test-method catalog 45 domains / 290 methods / 1363 cases: build & dependencies, static quality, unit and advanced testing, integration & contracts, API layer, UI & accessibility, E2E & acceptance, data, performance & capacity, security & compliance, reliability, observability & ops, engineering process, plus DSH-plugin-specific and open-source-readiness domains
Four depths smoke 331 / standard 984 / deep 1295 / exhaustive 1363 items, expanded by tier, filterable by domain
Project-type aware Detects library / cli / service / web / desktop / mobile / plugin / data / ml / monorepo and skips what cannot apply (a CLI has no multi-region failover) — skips are reported with reasons, never counted as passes
Evidence first Reconnaissance conclusions come from disk evidence only; every defect needs evidence, reproduction steps, files and a recommendation
Automated runs with exact write-back 11 preset commands (build, typecheck, lint, format, test, coverage, e2e, dependency audit, secret scan, install, outdated); results are written back to the exact test item by method + title. A missing tool is skipped, never a failure
Checkable report + read-back One checkbox per item in Markdown; qf_update sync=true writes the ticks back into structured data (WONTFIX / DEFER / NA markers supported)
Fix-scope negotiation qf_fixplan builds fix waves and a handoff sheet: fix everything, only P0-P1, specific ids, or everything except some ids
Zero runtime dependencies Node built-ins only, no build step, no install scripts — it cannot break because a host package version or symlink layout changed
Read-only recon, narrow execution surface Recon never writes; qf_exec only runs whitelisted preset commands; qf_probe only reaches loopback addresses

Install

In-session (needs danger-full-access):

plugin_manager install_bundle target=dsh-qualityforge
plugin_manager install_bundle target=github:lpeixin/dsh-qualityforge
plugin_manager install_bundle target=/absolute/path/to/dsh-qualityforge

CLI:

dsh plugin --profile <profile> add dsh-qualityforge
dsh plugin --profile <profile> add github:lpeixin/dsh-qualityforge   # lib/ is committed, no build needed
dsh plugin --profile <profile> add /absolute/path/to/dsh-qualityforge
dsh --profile <profile> --dump-config | grep -A3 qualityforge       # verify the bundle layer

Restart DSH (or start a new session) after installing so the module is loaded.

Disable without uninstalling — in ~/.dsh/profiles/<profile>/cordis.patch.yml:

- id: qualityforge
  disabled: true

Quick start

Ask, in a session, with the target project as the working directory:

Run a full QA audit on this project at standard depth, then give me a report I can tick off item by item.

The Harness then works through:

qf_scan   → recon: stack, project type, runnable commands, confirmed engineering gaps
qf_plan   → generate the checklist (depth=standard, applicability-filtered by default)
qf_exec   → build / typecheck / lint / test / coverage / dependency audit / secret scan
qf_probe  → loopback HTTP probes against a locally started service (optional)
qf_record → record the agent's deep-analysis conclusions (batched, with evidence)
qf_report → render Markdown + JSON, with verdict and fix waves
then: developer ticks → qf_update sync=true → qf_fixplan → fix → qf_exec to re-test

Artifacts live inside the audited project, never in the plugin:

<project>/.qualityforge/
├── audit.json                  structured data (single source of truth)
├── QUALITYFORGE-REPORT.md      the checkable report
└── report.json                 machine-readable snapshot (for CI)

See what a report looks like: sample report excerpt (a real self-audit of this very repository).


Usage

1. Recon — qf_scan

{ "projectPath": "/path/to/project", "force": true }

Produces a project profile: stack and package manager, project type, languages, runnable commands (each with its source, e.g. package.json#scripts.test), test frameworks, CI/container/IaC/security tooling, documentation completeness, and confirmed engineering gaps (a missing lockfile, suspected hardcoded credentials, and so on). Secret findings record location and kind only — never the value.

2. Plan — qf_plan

{ "depth": "standard", "categories": ["sec-input", "api-semantics"], "includeAll": false }
depth items use
smoke 331 pre-release blocker list
standard 984 normal delivery audit (default)
deep 1295 deep audit of an important release
exhaustive 1363 first full audit, compliance evidence

Every item gets a stable id (QF-001), a priority (P0–P3), a severity, a judgement mode (auto / agent analysis / human confirmation) and an explicit pass criterion. Re-running is safe: existing items keep their status and history; only reset=true rebuilds from scratch.

3. Run and probe — qf_exec / qf_probe

{ "presets": ["build", "typecheck", "lint", "test", "coverage", "audit", "secrets"] }
{ "urls": ["http://127.0.0.1:3000/health"], "expectStatus": 200, "itemId": "QF-512" }

qf_exec spawns only the preset commands discovered by recon (no shell, no arbitrary command strings); install is opt-in because it mutates the working tree. qf_probe is restricted to loopback http(s), GET/HEAD, with a truncated body.

4. Record conclusions — qf_record

{
  "items": [
    {
      "id": "QF-142",
      "status": "fail",
      "actual": "`ORDER BY ${sort}` is interpolated into SQL at src/api/orders.ts:88",
      "recommendation": "Map sort fields through an allowlist; return 400 otherwise",
      "files": ["src/api/orders.ts:88"],
      "repro": ["GET /api/orders?sort=id;DROP TABLE users--"],
      "evidence": ["SQL log shows the interpolated statement"],
      "cwe": ["CWE-89"],
      "effort": "S"
    }
  ]
}

Judgement discipline: fail (with reproduction) / pass (with evidence) / blocked (environment or dependency missing — do not record as fail) / na (does not apply, with a reason) / wontfix (risk accepted).

5. Report — qf_report

{ "reportPath": ".qualityforge/QUALITYFORGE-REPORT.md", "waveScope": "defects" }

Sections: how to use the report · verdict summary · priority & severity distribution · fix waves · problem list with one checkbox per defect · full checklist · skipped items · command evidence · project profile · re-test and sign-off table.

6. Negotiate fix scope — qf_update / qf_fixplan

Developers tick - [x] in the report (or `WONTFIX` / `DEFER` / `NA`); the Harness runs qf_update sync=true to read the ticks back. Ticking a defect sets it to fixed, pending re-test — only a re-test makes it verified.

Then pick a scope:

scope meaning
all everything still open
p0 / p0-p1 blockers only / the minimum shippable fix set
defects (default) every failing or blocked item
pending only items not judged yet
ids + excludeIds explicit include/exclude

Output is a fix handoff sheet — per item: expectation, observation, location, suggestion — plus the constraints (touch only what is in scope, every fix must be verifiable, write ticks back when done).


Tools

Tool Purpose Key parameters
qf_scan Recon + confirmed gaps projectPath, force
qf_plan Generate the checklist depth, categories, includeAll, withRecon, reset
qf_exec Run preset checks, write results back presets, timeoutMs, maxOutputBytes, autoRecord, recordPass
qf_probe Loopback HTTP probes urls, method, expectStatus, expectBodyContains, itemId
qf_record Record conclusions (batched) items[], createMissing
qf_list Query and filter items priorities, statuses, category, method, search, onlyOpen, limit
qf_update Change status / read ticks back ids+status, or sync=true
qf_report Render Markdown + JSON reportPath, waveScope, includePassedDetail, writeJson
qf_fixplan Fix waves + handoff sheet scope, ids, excludeIds, maxPerWave, format, writeFile
qf_reset Delete audit data confirm=true

The plugin also registers a methodology skill, qualityforge-audit (workflow, judgement discipline, anti-patterns), whose body lives in lib/SKILL.md and can be edited after installation without touching code.


Coverage

45 domains / 290 methods / 1363 cases — counts are verified by npm run catalog:stats.

Group Domains
Build & release build, deps, release
Static quality static-analysis, lint-format, code-quality
Testing unit-test, test-isolation, advanced-testing, integration, contract, test-double
Interfaces api-semantics, api-errors, api-evolution, ui-interaction, ui-a11y, ui-compat, ui-i18n
End-to-end e2e-journey, acceptance, regression
Data data-schema, data-integrity, data-quality
Performance performance, scalability, efficiency
Security sec-authn, sec-authz, sec-input, sec-data, sec-supply
Reliability & ops fault-tolerance, resilience-testing, recovery, observability, deployment, continuity
Process & ecosystem docs, workflow, maintainability, dsh-plugin, oss-ready, distribution

Methods map to real standards where genuinely applicable: ISO/IEC 25010, ISTQB, ISO/IEC/IEEE 29119, OWASP ASVS / Top 10, CWE Top 25, WCAG 2.2 AA, NIST SSDF, SLSA, Twelve-Factor, SRE golden signals, ITIL, SemVer, Keep a Changelog, OpenSSF Scorecard, GDPR / PCI-DSS.


Artifacts

File Purpose
.qualityforge/audit.json Single source of truth: profile, items, runs, probes, skips, full history
.qualityforge/QUALITYFORGE-REPORT.md Human-readable checkable report (deterministic render — safe to diff)
.qualityforge/report.json Machine-readable snapshot (verdict, defects, waves, counters) for CI

Minimal CI gate:

node -e "
const r = require('./.qualityforge/report.json');
if (r.summary.p0Open > 0) { console.error('blockers:', r.summary.p0Open); process.exit(1); }
if (r.summary.verdict === 'incomplete') { console.error('audit incomplete:', r.summary.pending); process.exit(1); }
console.log('quality gate passed:', r.summary.verdictLabel);
"

Security boundaries

QualityForge runs inside the DSH host process, outside the workspace sandbox, so its surface is deliberately narrow (see SECURITY.md):

Capability Boundary
Filesystem Read-only recon of the audited project; writes only inside its .qualityforge/
Execution Only recon-discovered preset commands, spawned without a shell; install is opt-in
Network Loopback http(s) only, GET/HEAD, truncated responses
Credentials Never read or stored; secret findings record location and kind only
Session No custom session events are written
Dependencies No external imports, no install scripts

Recommended reproductions in the report are harmless probes (random-token echo, timing, errors) — never destructive exploitation.


FAQ

How is this different from asking an agent to review the code?

A code review depends on improvisation; coverage is neither reproducible nor comparable. QualityForge draws items from an explicit catalog, so the same project at the same depth yields the same ids — you can answer "are the 300 items that passed last time still passing?" and the report diffs cleanly.

1363 items sounds like a lot.

Pick your depth: 331 / 984 / 1363. You can also narrow by categories; inapplicable items are filtered by project type; and most automation-judged items are settled in bulk by qf_exec.

Will it modify my project or install dependencies?

No. Recon is read-only, qf_exec runs only discovered presets (install off by default), qf_probe only reaches loopback, and every write goes into .qualityforge/.

Is a missing tool reported as a project defect?

No — it is recorded as skipped and the item stays "not judged", with the reason attached.

Why were some items skipped, and does that hide problems?

Items are skipped only when the project type cannot have that surface (a CLI has no multi-region failover). Every skip is listed with its reason in section 6. Use qf_plan includeAll=true to force them.

Does ticking a box mean "verified"?

No: ticking sets fixed, pending re-test; only a re-test sets verified. The two states are counted separately.

Can we fix only part of the findings?

Yes — that is what scope is for: p0, p0-p1, defects, pending, or ids + excludeIds. "Fix everything except QF-003 and QF-017" is scope=defects with excludeIds=["QF-003","QF-017"]. Whatever is left stays visible via qf_list onlyOpen=true.


Development

git clone https://github.com/lpeixin/dsh-qualityforge.git
cd dsh-qualityforge

node --test                       # 27 tests (catalog + recon + end-to-end)
npm run validate                  # catalog contract + preset-mapping checks
npm run catalog:stats

Zero dependencies, no build step: the plugin uses Node built-ins only, so a clone is immediately runnable.

Three hard rules (see CONTRIBUTING.md):

  1. No external imports — nothing to resolve at install time, nothing to break when host packages move;
  2. Named exports only, no export default — the Loader's unwrapExports would otherwise drop inject;
  3. Never modify the audited project and never write custom session events.

Adding test cases is the most valuable contribution: edit a pure-data module under lib/catalog/, follow docs/CATALOG-AUTHORING.md, then run npm run validate.


Compatibility

Item Requirement
DSH Any version supporting Cordis plugins and the dsh.bundle.patch contract
Node.js ≥ 20.11 (CI runs 20.x and 22.x)
Platform macOS / Linux / Windows (built-ins and spawn only, no native modules)
Install size npm tarball ~215 KB (~726 KB unpacked), excluding audit artifacts of audited projects

If a composition lacks the skills service, skill registration is skipped and the tools still work.


Roadmap

  • qf_diff — compare two audits and report newly introduced / fixed / still present, for pre-release regression;
  • Coverage parsing — read real numbers from lcov / coverage.py output into the matching test item;
  • English report and catalog locale (lang=en);
  • Industry-specific domains: financial reconciliation, healthcare data compliance, IoT firmware updates;
  • Export to SARIF / JUnit XML so results can flow into existing quality platforms.

License

MIT © 2026 Peixin

DSH HUB

A community index for DSH plugins. Not an official GitHub or DeepSeek AI product.

CommunityResourcesAPIAbout
—/ 5

No ratings yet

Verified DSH bundle

Commit 0f27bf4d96d2

Community comments

No comments yet. Be the first to write one.