GFG for DeepSeek Harness
A lightweight generation-fact graph for tracing how agent actions and tool results were actually formed.
Final Tool Result
↑
Execution Occurrence
↑
Effective Tool Action
↑
Pre-execute / Policy
↑
Original Tool Call
↑
Assistant Message Settlement
This is not log correlation. Relations are captured from concrete runtime occurrences and preserved as exact formation facts. This is formation provenance, not a claim of causal explanation or access to a model's internal reasoning.
gfg_trace({"target_id":"read-1"}) returns a structured formation subgraph.
Ordinary tool results carry no added GFG metadata.
Try the demo
Requires Node 24+, pnpm, and the pinned Harness 0.2.0-rc.2 packages.
git clone https://github.com/wind342/dsh-gfg.git
cd dsh-gfg
pnpm install --frozen-lockfile
pnpm build
pnpm check
pnpm test
pnpm demo
The demo uses the real, unmodified Harness AgentLoop, session store and tool
runtime, with a scripted LLM adapter for repeatability; no API key is needed.
The Agent dispatches an actual file-reading tool, receives hello, then dispatches
gfg_trace. A second session demonstrates a policy denial with zero tool-body calls.
The script checks capture-OFF/ON equality for ordinary outputs. This is not a
remote DeepSeek-model test and not the product CLI launcher.
Committed synthetic examples:
Install in a Harness profile
Build a package locally (not published to npm):
pnpm pack
dsh plugin --profile headless add /absolute/path/to/dsh-gfg-0.1.0.tgz
dsh --profile headless "Read a.txt, then use gfg_trace with the read tool's call ID to inspect how its result formed."
The package declares dsh.bundle and contains cordis.patch.yml, so the profile
plugin manager can enable its row. Use the official dsh launcher, not a separate
home-made app bootstrap. Set up your model provider/credentials in Harness itself.
The official CLI 0.2.0-rc.2 profile install smoke also runs without a provider:
it packs this plugin, installs the tarball with dsh plugin, checks the enabled
bundle and --dump-config row, and boots a minimal profile through the official
launcher until its application-ready callback observes the GFG service.
A live provider remains unverified here.
Do not silently override the pinned compatibility range on another Harness version.
Run pnpm profile:smoke /absolute/path/to/@deepseek-ai/dsh/lib/bin.js with an
independently installed official CLI. The smoke uses a fresh ignored .profile-smoke/
home, not your real Harness profile. It records failure as
product_cli_profile_test: false; packing alone is not a profile PASS.
Actual CLI smoke result includes source
hashes checked by pnpm verify to detect stale evidence.
The tarball was separately installed into an isolated package directory and its
registered tools exercised with node examples/package-smoke.mjs <install-directory>.
That check loads dependencies from the installed package, not the source checkout.
Optional override in the profile's cordis.patch.yml:
- id: dsh-gfg
config:
directory: /absolute/private/path/to/gfg
By default, each process/session has a fresh run directory under ~/.dsh/gfg/.
It contains append-only receipts.jsonl and a gfg.json snapshot written at Harness
session flush and plugin disposal. Receipt writes are synchronous; session flush
also calls fsync. Memory-only mode is used by tests, not the default profile.
Raw receipts can contain private prompts, file content, arguments and results. Keep the directory private, check OS ACLs (especially on Windows), and never publish real-user graphs automatically. Only public fixture artifacts are committed here.
Model-callable tools
{"target_id":"read-1", "direction":"backward", "max_depth":32}
gfg_trace: backward/forward local graph traversal; returnsnodes,edges,formation_path,evidence_refs,target,direction,complete, andtruncated.gfg_get_node({"id":"..."}): the structural projection of one exact node.
Both queries use a positive field allowlist. They never return raw receipt/node
payloads, canonical values, arguments, metadata or error text, including data hidden
by post-execute policy. Receipt references retain the ID/SHA-256, evidence availability
and payload_visibility: "private". Full evidence remains in the private journal
and internal RuntimeGFG, unchanged. There is no model-callable raw-payload option.
The simple read demo's compact JSON trace is 30,792 bytes, versus 44,055 bytes
for the same internal full-receipt trace (about 30% smaller); graph nodes and
edges are retained. This is not a compressed graph protocol.
Targets may be a fact/occurrence/outcome ID, a durable tool-result message ID,
or the original provider tool-call ID. result:<callId> addresses the final runtime
result before a durable result exists. Reused aliases produce AMBIGUOUS_TARGET;
they are never resolved by guessing the newest result. Use an exact node ID instead.
Queries are scoped to the calling session; there is no model-controlled run_id.
Query tools themselves are excluded from capture to avoid recursively storing graphs.
What is captured
| Harness boundary | Recorded material |
|---|---|
assistant/message with a tool call |
Full durable event, stream and proposed call blocks |
assistant/attempt |
Stream with explicit no_surface_message disposition |
tool/call |
Full durable event, original raw argument string |
tools/pre-execute |
Observed entry and actual policy decision |
tools/execute |
Observed dispatch entry and normalized returned result |
tools/post-execute |
Incoming result and actual downstream decision |
tools/result |
Authoritative frozen result, including structured value |
tool/result |
Full durable model-facing event |
Entry observations represent actual middleware invocations, not invented tool effects. Identity binding uses native call IDs and the registry's opaque execution tokens; nested dispatches use the actual parent token. No relation is inferred from timestamps, text similarity, or proximity. Runtime handles (agent/context/functions, private schema and cancellation object) are not serialized; their relevant identity and signal state are recorded. Native JSON arguments and results are retained in full.
Denial, cancellation, execution failure, suppression and unobserved settlement are
ExplicitDispositions. Missing upstream history becomes an explicit source record,
not a fabricated stage. On capture failure, ordinary results still pass through;
the graph is marked incomplete and query tools refuse to present it as complete.
An error without observed tools/execute is conservatively classified as
policy_or_pre_dispatch_unobserved, unless a known denied/cancelled/not-executed
classification is available. It is never guessed to be execution_failed.
Graph structure
Each fact keeps (origin, transform, occurrence, outcome; relation_role), plus
fact_id, run_id, sequence and receipt_sha256.
Source / GeneratedOrigin → OccurrenceNode ──realizes_fact──→ FactNode → Outcome
│
next GeneratedOrigin ←──────────────────────────┘
Actual continuation is Outcome → GeneratedOrigin → next occurrence/fact.
The compiler only uses supplied bindings. Multi-source/multi-result events do not
become a Cartesian product. Traversal respects fact-specific origin bindings even
when facts share an occurrence. formation_path is a BFS visitation order with
depths; edges, not adjacent rows in that array, define the formation relationships.
Canonical JCS + SHA-256 identifies receipts, facts and graph snapshots. The schema
is dsh-gfg/jcs/1, not Core-v3's Python wire/hash format. The offline validator
checks hashes, references, receipt sequence/chain, aliases and recompilation.
It does not validate the entire graph on each query. Indexes are incremental and
queries visit only the requested neighborhood; no global transitive closure.
Tests and measured overhead
pnpm test covers ordinary-output invariance, real AgentLoop formation, denied/
failed/cancelled calls, explicit missing outcomes, multi-source exact bindings,
concurrency, nesting, repeated IDs, session isolation, byte-identical deterministic
graphs, journal replay, tampering and capture failures. pnpm check typechecks both
the product and fixtures.
Regression tests also cover post-policy secret redaction (including content-only
redaction), monotonic guard rejection with zero body calls, unobserved pre-dispatch
failures, and compact query size with unchanged internal evidence.
GitHub Actions runs frozen install, build, typecheck, tests, demo and verification
on Node 24. The separately recorded real CLI smoke is source-hash bound.
Run pnpm test:report && pnpm demo && pnpm verify to regenerate the machine-readable
checks and test results.
pnpm benchmark runs 3 × 500 synthetic zero-I/O tool calls per mode. On the recorded
Windows/i5-12490F run, median time per call was approximately 0.21 ms OFF,
0.70 ms memory capture, and 0.85 ms journal capture. A local trace was roughly
0.21–0.27 ms, over a graph with 3,000 facts. These figures exclude LLM time,
setup, final snapshot and offline validation; fsync is reported separately.
This is measurable overhead, not free capture or a production throughput claim.
All samples and scope.
Scope and limitations
- This v1 observes exposed runtime boundaries, not provider internals, arbitrary tool-body file dependencies, or every private retry attempt. An outer middleware that bypasses our listener is outside that listener's capture scope. Load before work starts; no retrospective inference of unobserved events.
- Native tool mode is exercised end to end. Explicit parent-token nesting is tested; PTC-specific durable bridge events and subagent cross-session joins are not implemented. No claims of complete PTC/body-level dataflow.
- Observer cost can affect wall time and timeout-sensitive behavior. Registering query tools changes the tool catalog; deterministic ordinary-output invariance does not imply unchanged stochastic model decisions.
- Full receipts consume RAM/disk proportional to captured payload size. No automatic retention/deletion, compression or background upload.
- An incomplete/torn journal is detectable, but an attacker rewriting all records and hashes cannot be defeated without a separately trusted root. These hashes are integrity identities, not signatures or proof that the runtime was truthful.
- Abrupt process termination may omit the final snapshot/close disposition; a recovered journal is a captured prefix, not proof of a completed execution.
- The strict JSON boundary rejects non-JSON objects, cycles, lone surrogates, non-finite numbers and unsafe integers rather than silently summarizing them.
Provenance and license
Independent Apache-2.0 plugin; not an official DeepSeek product. Reuses official Harness services and existing GFG organization, without importing private research code or the heavy Core-v3 validator. Reuse/license audit.
Full research evidence remains in the separate frozen research-evidence repository. The public GFG Core extraction is also independent. Neither repository nor any frozen tag is changed by this plugin.
Prepared Discussion post — not published automatically.
No comments yet. Be the first to write one.