gba-mep-docx-toc
A DeepSeek Harness (dsh) bundle that registers one skill: insert a real Word table of contents into an existing .docx.
The skill writes a genuine TOC field with _Toc bookmarks straight into word/document.xml — graded indent per heading level, dot leaders, right-aligned page numbers — using only the Python standard library. It does not need Microsoft Word, LibreOffice, COM, or any third-party package.
Install
dsh plugin --profile web add gba-mep-docx-toc
This repository is also installable straight from source through pnpm's GitHub spec, and it is listed in the awesome-dsh-plugin catalog, so dsh-market can install it with one click.
After installing, restart dsh web once so the bundle layer is composed.
What it does
Given a .docx whose headings use the Heading 1–Heading 9 styles (or carry w:outlineLvl):
- finds the headings in document order,
- inserts
_Toc…bookmarks so every entry is clickable, - builds the
TOC \o "1-3" \h \z \ufield with one entry per heading — hyperlink anchor → title text → right-aligned tab with dot leader →PAGEREFfield, - indents level 2 by
100and level 3 by200character units, - rewrites only
word/document.xmland copies every other part of the package byte for byte, - verifies the result with
verify(entry count = bookmark count =PAGEREFcount).
Requirements
- A DSH build that exposes the
skillsservice — the entry imports nothing from the harness, so it works on any host publishing that contract (checked against DSH0.2.0-rc.2). - Python 3.9+ on the machine that runs the script. No packages to install.
- Node.js only because it is a dsh plugin; the entry itself uses
node:fs,node:pathandnode:url.
Usage
python scripts/docx_toc.py build 输入.docx -o 输出.docx --title "目 录" --page-break both
python scripts/docx_toc.py verify 输出.docx
python scripts/docx_toc.py selftest
In a session you can also just describe the task — the skill's description is what the agent routes on.
Boundaries
- Page numbers are placeholders. Python does not paginate, so real numbers appear only after Word updates the field (
Ctrl+A, thenF9, choose Update page numbers only). Until then every entry reads1. - No section-based page numbering. With sections that restart numbering (roman front matter, arabic body), the
PAGEREFresult depends on the Word section setup and is not guaranteed. - It does not edit an existing TOC. A document that already carries a
TOCfield gets a second one; delete the old one first. - Paginating is not guaranteed to match Word. Rendering checks used LibreOffice, whose page breaks differ from Microsoft Word's.
- Headings must use heading styles. Manually bolded and enlarged text is invisible to the script.
Layout
lib/index.js bundle entry — registers the skill via ctx.skills.register()
cordis.patch.yml loader patch: one insert row pointing at this package by name
SKILL.md skill body (frontmatter: name / description / whenToUse)
scripts/docx_toc.py the implementation (stdlib only)
references/OOXML-TOC.md notes on the OOXML table-of-contents field
tests/ offline tests: entry behaviour and installability contract
screenshots.json screenshots shown by plugin storefronts
Tests
node --test tests/
tests/entry.test.mjs runs apply() against a minimal context and asserts the registered fields, the resourceBase, and that the referenced files ship with the package — including a CRLF regression for the frontmatter parser. tests/bundle-contract.test.mjs asserts the installability contract: dsh.bundle declared, patch row naming the package, files whitelist complete, entry importing only node: builtins.
License
MIT — see LICENSE.md.
No comments yet. Be the first to write one.