DSH HUB
HomePlugin StorePlugin PacksCommunityRankingsResourcesPublish Guide
Plugin source
Back to catalog

ReGMeIoN /

ReGMeIoN/cancer-meta-pipeline

Topic repository only

AI-assisted, reproducible pipeline that turns a paper list into publication-ready extraction tables for cancer systematic reviews

★ 1 Stars0 Forks0 IssuesN/A Community rating0 Confirmed installs
View on GitHub
READMESource: main@c5972278

cancer-meta-pipeline

A reproducible, AI-assisted pipeline that turns a list of papers into publication-ready extraction tables for cancer systematic reviews / meta-analyses.

Give it three things — a paper list, a screening protocol written in plain language, and an abstrackr project id — and it runs the whole literature workflow: normalise → deduplicate → compile a machine-readable rubric → title/abstract screening in batches → write decisions back to the platform → full-text chase → data extraction → Table 1 / Table 2.

It ships as a DeepSeek Harness Skill: a SKILL.md plus plain Python scripts. No framework, no database, no service to run.


What it does (8 stages, each with a human checkpoint)

# Stage Output
1 Ingest — CSV / Excel / RIS / BibTeX / EndNote / bare PMID list work/records_local.csv + import report
2 Pull platform — full export with citation_id, status, tags work/records_all.csv, status snapshot
3 Compile rubric — plain-language protocol → E-code table + inclusion clauses work/_compiled/rubric.json, PROMPT_batch.md, PROMPT_fulltext.md
4 Batch screening — LLM subagents judge title/abstract in 250-record batches work/batches/ → work/decisions/
5 Validate & merge — one central validator, never self-written checks work/validation_report.md, screening_decisions.csv
6 Submit — dry-run first, audit tags, rollback path platform labels + logs/submit_audit.log
7 Full text — three-channel open-access chase (Europe PMC / Unpaywall / PMC) article/oa/*.txt, needs-a-human list
8 Extract — Table 1 (characteristics + reported effects), Table 2 (effect sizes) out/table1_*, out/table2_*, .docx, coverage_audit.md

Every stage ends with a checkpoint: the skill stops and asks the reviewer before moving on.

Design rules baked in

  • Batches of 200–300 records (small batches multiply subagent startup cost).
  • Subagents may not write their own scripts or validators — one central validator.
  • Prompts live in files, so a hundred batches share one identical wording.
  • Everything is traceable: every decision carries a reason + a verbatim quote; every platform write leaves an audit tag (ft:corrected-from-*) and a documented rollback.
  • E-code ⇒ excluded is enforced in code: a study that reports only a composite outcome (sensitivity-only) or is a Mendelian randomisation study (triangulation stream) can never enter the primary synthesis, even if a subagent labels it "Include".

Requirements

  • Python 3.10+ (standard library only; pypdf for PDF→text, python-docx for the Word table)
  • Optional: R + metafor for the pooling step (see reference/HANDOFF_to_analysis.md)
  • Optional: abstrackr-api-toolkit for the platform stage, and an ABSTRACKR_HOME directory holding your credentials

Quick start

projects/<name>/
├── config.json      # platform project id + batch sizes
├── protocol.md      # your PECO + inclusion clauses + E-code table (see templates/)
└── input/           # the paper list, in any supported format
$skill = '<this repo>'
$py    = 'python'
$env:PYTHONIOENCODING = 'utf-8'
$proj  = '<absolute path to projects/<name>>'

& $py "$skill\scripts\ingest.py"             --project $proj
& $py "$skill\scripts\pull_platform.py"      --project $proj
& $py "$skill\scripts\compile_rubric.py"     --project $proj      # ⛳ confirm the E-code table
& $py "$skill\scripts\prep_batches.py"       --project $proj --size 250
#   ... one subagent per batch, using work/_compiled/PROMPT_batch.md ...
& $py "$skill\scripts\validate_decisions.py" --project $proj
& $py "$skill\scripts\merge_decisions.py"    --project $proj
& $py "$skill\scripts\submit_decisions.py"   --project $proj      # dry-run, then --execute
& $py "$skill\scripts\verify_write.py"       --project $proj
& $py "$skill\scripts\build_tables.py"       --project $proj --audit

Offline smoke test (fictional data, no network, no platform):

& $py "$skill\scripts\ingest.py"             --project "$skill\examples\smoke"
& $py "$skill\scripts\compile_rubric.py"     --project "$skill\examples\smoke"
& $py "$skill\scripts\prep_batches.py"       --project "$skill\examples\smoke" --size 2
& $py "$skill\scripts\build_tables.py"       --project "$skill\examples\smoke" --audit

Layout

SKILL.md                 entry point: stages, hard rules, checkpoints
stages/DETAILS.md        step-by-step operating notes
scripts/                 13 scripts, one per stage (+ common.py)
templates/               protocol / rubric / prompt / config / handover templates
reference/PITFALLS.md    26 real pitfalls (read this before running anything)
reference/checklist.md   per-stage deliverables and the questions to ask
examples/                fictional smoke projects (CSV / RIS / BibTeX)

Read this first

reference/PITFALLS.md is the most valuable file here — every entry is a real failure from a real review (platform throttling, silently truncated snapshots, BOM-broken JSON, a case-sensitive regex that quietly disabled a central rule, metafor's I² already being a percentage, and more).

License

Not yet chosen — add one before relying on this in production.

—/ 5

No ratings yet

Manifest verification required

Commit c59722784c8c

Community comments

No comments yet. Be the first to write one.

DSH HUB

A community index for DSH plugins. Not an official GitHub or DeepSeek AI product.

CommunityResourcesAPIAbout