cancer-meta-pipeline
A reproducible, AI-assisted pipeline that turns a list of papers into publication-ready extraction tables for cancer systematic reviews / meta-analyses.
Give it three things — a paper list, a screening protocol written in plain language, and an abstrackr project id — and it runs the whole literature workflow: normalise → deduplicate → compile a machine-readable rubric → title/abstract screening in batches → write decisions back to the platform → full-text chase → data extraction → Table 1 / Table 2.
It ships as a DeepSeek Harness Skill: a SKILL.md plus plain
Python scripts. No framework, no database, no service to run.
What it does (8 stages, each with a human checkpoint)
| # | Stage | Output |
|---|---|---|
| 1 | Ingest — CSV / Excel / RIS / BibTeX / EndNote / bare PMID list | work/records_local.csv + import report |
| 2 | Pull platform — full export with citation_id, status, tags |
work/records_all.csv, status snapshot |
| 3 | Compile rubric — plain-language protocol → E-code table + inclusion clauses | work/_compiled/rubric.json, PROMPT_batch.md, PROMPT_fulltext.md |
| 4 | Batch screening — LLM subagents judge title/abstract in 250-record batches | work/batches/ → work/decisions/ |
| 5 | Validate & merge — one central validator, never self-written checks | work/validation_report.md, screening_decisions.csv |
| 6 | Submit — dry-run first, audit tags, rollback path | platform labels + logs/submit_audit.log |
| 7 | Full text — three-channel open-access chase (Europe PMC / Unpaywall / PMC) | article/oa/*.txt, needs-a-human list |
| 8 | Extract — Table 1 (characteristics + reported effects), Table 2 (effect sizes) | out/table1_*, out/table2_*, .docx, coverage_audit.md |
Every stage ends with a checkpoint: the skill stops and asks the reviewer before moving on.
Design rules baked in
- Batches of 200–300 records (small batches multiply subagent startup cost).
- Subagents may not write their own scripts or validators — one central validator.
- Prompts live in files, so a hundred batches share one identical wording.
- Everything is traceable: every decision carries a reason + a verbatim quote; every
platform write leaves an audit tag (
ft:corrected-from-*) and a documented rollback. E-code ⇒ excludedis enforced in code: a study that reports only a composite outcome (sensitivity-only) or is a Mendelian randomisation study (triangulation stream) can never enter the primary synthesis, even if a subagent labels it "Include".
Requirements
- Python 3.10+ (standard library only;
pypdffor PDF→text,python-docxfor the Word table) - Optional: R +
metaforfor the pooling step (seereference/HANDOFF_to_analysis.md) - Optional: abstrackr-api-toolkit for the
platform stage, and an
ABSTRACKR_HOMEdirectory holding your credentials
Quick start
projects/<name>/
├── config.json # platform project id + batch sizes
├── protocol.md # your PECO + inclusion clauses + E-code table (see templates/)
└── input/ # the paper list, in any supported format
$skill = '<this repo>'
$py = 'python'
$env:PYTHONIOENCODING = 'utf-8'
$proj = '<absolute path to projects/<name>>'
& $py "$skill\scripts\ingest.py" --project $proj
& $py "$skill\scripts\pull_platform.py" --project $proj
& $py "$skill\scripts\compile_rubric.py" --project $proj # ⛳ confirm the E-code table
& $py "$skill\scripts\prep_batches.py" --project $proj --size 250
# ... one subagent per batch, using work/_compiled/PROMPT_batch.md ...
& $py "$skill\scripts\validate_decisions.py" --project $proj
& $py "$skill\scripts\merge_decisions.py" --project $proj
& $py "$skill\scripts\submit_decisions.py" --project $proj # dry-run, then --execute
& $py "$skill\scripts\verify_write.py" --project $proj
& $py "$skill\scripts\build_tables.py" --project $proj --audit
Offline smoke test (fictional data, no network, no platform):
& $py "$skill\scripts\ingest.py" --project "$skill\examples\smoke"
& $py "$skill\scripts\compile_rubric.py" --project "$skill\examples\smoke"
& $py "$skill\scripts\prep_batches.py" --project "$skill\examples\smoke" --size 2
& $py "$skill\scripts\build_tables.py" --project "$skill\examples\smoke" --audit
Layout
SKILL.md entry point: stages, hard rules, checkpoints
stages/DETAILS.md step-by-step operating notes
scripts/ 13 scripts, one per stage (+ common.py)
templates/ protocol / rubric / prompt / config / handover templates
reference/PITFALLS.md 26 real pitfalls (read this before running anything)
reference/checklist.md per-stage deliverables and the questions to ask
examples/ fictional smoke projects (CSV / RIS / BibTeX)
Read this first
reference/PITFALLS.md is the most valuable file here — every entry is a real failure from a
real review (platform throttling, silently truncated snapshots, BOM-broken JSON, a
case-sensitive regex that quietly disabled a central rule, metafor's I² already being a
percentage, and more).
License
Not yet chosen — add one before relying on this in production.
No comments yet. Be the first to write one.