DSH HUB
HomePlugin StorePlugin PacksCommunityRankingsResourcesPublish Guide
Plugin source
Back to catalog

wind342 /

wind342/dsh-gfg

Verified

Lightweight generation-fact graphs for tracing actual DeepSeek Harness tool formation

★ 0 Stars0 Forks0 IssuesN/A Community rating0 Confirmed installs
View on GitHub
READMESource: main@b7b798a1

GFG for DeepSeek Harness

A lightweight generation-fact graph for tracing how agent actions and tool results were actually formed.

Final Tool Result
       ↑
Execution Occurrence
       ↑
Effective Tool Action
       ↑
Pre-execute / Policy
       ↑
Original Tool Call
       ↑
Assistant Message Settlement

This is not log correlation. Relations are captured from concrete runtime occurrences and preserved as exact formation facts. This is formation provenance, not a claim of causal explanation or access to a model's internal reasoning.

gfg_trace({"target_id":"read-1"}) returns a structured formation subgraph. Ordinary tool results carry no added GFG metadata.

Formation diagram

Try the demo

Requires Node 24+, pnpm, and the pinned Harness 0.2.0-rc.2 packages.

git clone https://github.com/wind342/dsh-gfg.git
cd dsh-gfg
pnpm install --frozen-lockfile
pnpm build
pnpm check
pnpm test
pnpm demo

The demo uses the real, unmodified Harness AgentLoop, session store and tool runtime, with a scripted LLM adapter for repeatability; no API key is needed. The Agent dispatches an actual file-reading tool, receives hello, then dispatches gfg_trace. A second session demonstrates a policy denial with zero tool-body calls. The script checks capture-OFF/ON equality for ordinary outputs. This is not a remote DeepSeek-model test and not the product CLI launcher.

Committed synthetic examples:

  • Demo results
  • Read-file graph / trace
  • Denied graph / trace

Install in a Harness profile

Build a package locally (not published to npm):

pnpm pack
dsh plugin --profile headless add /absolute/path/to/dsh-gfg-0.1.0.tgz
dsh --profile headless "Read a.txt, then use gfg_trace with the read tool's call ID to inspect how its result formed."

The package declares dsh.bundle and contains cordis.patch.yml, so the profile plugin manager can enable its row. Use the official dsh launcher, not a separate home-made app bootstrap. Set up your model provider/credentials in Harness itself. The official CLI 0.2.0-rc.2 profile install smoke also runs without a provider: it packs this plugin, installs the tarball with dsh plugin, checks the enabled bundle and --dump-config row, and boots a minimal profile through the official launcher until its application-ready callback observes the GFG service. A live provider remains unverified here. Do not silently override the pinned compatibility range on another Harness version.

Run pnpm profile:smoke /absolute/path/to/@deepseek-ai/dsh/lib/bin.js with an independently installed official CLI. The smoke uses a fresh ignored .profile-smoke/ home, not your real Harness profile. It records failure as product_cli_profile_test: false; packing alone is not a profile PASS. Actual CLI smoke result includes source hashes checked by pnpm verify to detect stale evidence.

The tarball was separately installed into an isolated package directory and its registered tools exercised with node examples/package-smoke.mjs <install-directory>. That check loads dependencies from the installed package, not the source checkout.

Optional override in the profile's cordis.patch.yml:

- id: dsh-gfg
  config:
    directory: /absolute/private/path/to/gfg

By default, each process/session has a fresh run directory under ~/.dsh/gfg/. It contains append-only receipts.jsonl and a gfg.json snapshot written at Harness session flush and plugin disposal. Receipt writes are synchronous; session flush also calls fsync. Memory-only mode is used by tests, not the default profile.

Raw receipts can contain private prompts, file content, arguments and results. Keep the directory private, check OS ACLs (especially on Windows), and never publish real-user graphs automatically. Only public fixture artifacts are committed here.

Model-callable tools

{"target_id":"read-1", "direction":"backward", "max_depth":32}
  • gfg_trace: backward/forward local graph traversal; returns nodes, edges, formation_path, evidence_refs, target, direction, complete, and truncated.
  • gfg_get_node({"id":"..."}): the structural projection of one exact node.

Both queries use a positive field allowlist. They never return raw receipt/node payloads, canonical values, arguments, metadata or error text, including data hidden by post-execute policy. Receipt references retain the ID/SHA-256, evidence availability and payload_visibility: "private". Full evidence remains in the private journal and internal RuntimeGFG, unchanged. There is no model-callable raw-payload option. The simple read demo's compact JSON trace is 30,792 bytes, versus 44,055 bytes for the same internal full-receipt trace (about 30% smaller); graph nodes and edges are retained. This is not a compressed graph protocol.

Targets may be a fact/occurrence/outcome ID, a durable tool-result message ID, or the original provider tool-call ID. result:<callId> addresses the final runtime result before a durable result exists. Reused aliases produce AMBIGUOUS_TARGET; they are never resolved by guessing the newest result. Use an exact node ID instead. Queries are scoped to the calling session; there is no model-controlled run_id. Query tools themselves are excluded from capture to avoid recursively storing graphs.

What is captured

Harness boundary Recorded material
assistant/message with a tool call Full durable event, stream and proposed call blocks
assistant/attempt Stream with explicit no_surface_message disposition
tool/call Full durable event, original raw argument string
tools/pre-execute Observed entry and actual policy decision
tools/execute Observed dispatch entry and normalized returned result
tools/post-execute Incoming result and actual downstream decision
tools/result Authoritative frozen result, including structured value
tool/result Full durable model-facing event

Entry observations represent actual middleware invocations, not invented tool effects. Identity binding uses native call IDs and the registry's opaque execution tokens; nested dispatches use the actual parent token. No relation is inferred from timestamps, text similarity, or proximity. Runtime handles (agent/context/functions, private schema and cancellation object) are not serialized; their relevant identity and signal state are recorded. Native JSON arguments and results are retained in full.

Denial, cancellation, execution failure, suppression and unobserved settlement are ExplicitDispositions. Missing upstream history becomes an explicit source record, not a fabricated stage. On capture failure, ordinary results still pass through; the graph is marked incomplete and query tools refuse to present it as complete. An error without observed tools/execute is conservatively classified as policy_or_pre_dispatch_unobserved, unless a known denied/cancelled/not-executed classification is available. It is never guessed to be execution_failed.

Graph structure

Each fact keeps (origin, transform, occurrence, outcome; relation_role), plus fact_id, run_id, sequence and receipt_sha256.

Source / GeneratedOrigin → OccurrenceNode ──realizes_fact──→ FactNode → Outcome
                                                                      │
                       next GeneratedOrigin ←──────────────────────────┘

Actual continuation is Outcome → GeneratedOrigin → next occurrence/fact. The compiler only uses supplied bindings. Multi-source/multi-result events do not become a Cartesian product. Traversal respects fact-specific origin bindings even when facts share an occurrence. formation_path is a BFS visitation order with depths; edges, not adjacent rows in that array, define the formation relationships.

Canonical JCS + SHA-256 identifies receipts, facts and graph snapshots. The schema is dsh-gfg/jcs/1, not Core-v3's Python wire/hash format. The offline validator checks hashes, references, receipt sequence/chain, aliases and recompilation. It does not validate the entire graph on each query. Indexes are incremental and queries visit only the requested neighborhood; no global transitive closure.

Tests and measured overhead

pnpm test covers ordinary-output invariance, real AgentLoop formation, denied/ failed/cancelled calls, explicit missing outcomes, multi-source exact bindings, concurrency, nesting, repeated IDs, session isolation, byte-identical deterministic graphs, journal replay, tampering and capture failures. pnpm check typechecks both the product and fixtures. Regression tests also cover post-policy secret redaction (including content-only redaction), monotonic guard rejection with zero body calls, unobserved pre-dispatch failures, and compact query size with unchanged internal evidence. GitHub Actions runs frozen install, build, typecheck, tests, demo and verification on Node 24. The separately recorded real CLI smoke is source-hash bound.

Run pnpm test:report && pnpm demo && pnpm verify to regenerate the machine-readable checks and test results.

pnpm benchmark runs 3 × 500 synthetic zero-I/O tool calls per mode. On the recorded Windows/i5-12490F run, median time per call was approximately 0.21 ms OFF, 0.70 ms memory capture, and 0.85 ms journal capture. A local trace was roughly 0.21–0.27 ms, over a graph with 3,000 facts. These figures exclude LLM time, setup, final snapshot and offline validation; fsync is reported separately. This is measurable overhead, not free capture or a production throughput claim. All samples and scope.

Scope and limitations

  • This v1 observes exposed runtime boundaries, not provider internals, arbitrary tool-body file dependencies, or every private retry attempt. An outer middleware that bypasses our listener is outside that listener's capture scope. Load before work starts; no retrospective inference of unobserved events.
  • Native tool mode is exercised end to end. Explicit parent-token nesting is tested; PTC-specific durable bridge events and subagent cross-session joins are not implemented. No claims of complete PTC/body-level dataflow.
  • Observer cost can affect wall time and timeout-sensitive behavior. Registering query tools changes the tool catalog; deterministic ordinary-output invariance does not imply unchanged stochastic model decisions.
  • Full receipts consume RAM/disk proportional to captured payload size. No automatic retention/deletion, compression or background upload.
  • An incomplete/torn journal is detectable, but an attacker rewriting all records and hashes cannot be defeated without a separately trusted root. These hashes are integrity identities, not signatures or proof that the runtime was truthful.
  • Abrupt process termination may omit the final snapshot/close disposition; a recovered journal is a captured prefix, not proof of a completed execution.
  • The strict JSON boundary rejects non-JSON objects, cycles, lone surrogates, non-finite numbers and unsafe integers rather than silently summarizing them.

Provenance and license

Independent Apache-2.0 plugin; not an official DeepSeek product. Reuses official Harness services and existing GFG organization, without importing private research code or the heavy Core-v3 validator. Reuse/license audit.

Full research evidence remains in the separate frozen research-evidence repository. The public GFG Core extraction is also independent. Neither repository nor any frozen tag is changed by this plugin.

Prepared Discussion post — not published automatically.

—/ 5

No ratings yet

Verified DSH bundle

Commit b7b798a126e9

Community comments

No comments yet. Be the first to write one.

DSH HUB

A community index for DSH plugins. Not an official GitHub or DeepSeek AI product.

CommunityResourcesAPIAbout