Fusion MoA
Fusion MoA is a model-, GPU-, and harness-independent Mixture-of-Agents runtime for coding agents. Bring a main model and an optional pool of read-only experts, configure them in one recipe, and expose one stable API to Codex, Claude Code, OpenCode, DeepSeek Harness, or another client.
coding agents / DeepSeek Harness
|
OpenAI Chat + Responses + Anthropic Messages
|
protocol-neutral runtime
policy + experts + budgets + fallback
|
vLLM / llama.cpp / cloud APIs / plugins
Status
The v0.1 contract includes:
- strict
fusion/v1recipes with cross-reference and secret validation; - declared model capabilities and per-model concurrency limits;
- OpenAI-compatible, llama.cpp, and Anthropic-compatible providers;
direct,main-critic, and parallelreview-boardpolicies;- OpenAI Chat, OpenAI Responses, and Anthropic Messages endpoints;
- portable function/tool-call round trips and
/v1/modelsdiscovery; - native final-model SSE for all three public protocols;
- Python entry points for third-party providers and policies;
- an evaluation-gated promotion command for controlled RSI;
- a version-pinned DeepSeek Harness profile bundle.
Native final-model streaming means expert orchestration completes first, then text and tool-call deltas from the one authoritative main-model call are forwarded without buffering the full answer. Expert output is never exposed as the public stream. Non-function built-in tools, multimodal parity across every provider, and online training are not claimed in v0.1.
Quick start
git clone https://github.com/xiaohou521/fusion-moa.git
cd fusion-moa
python -m venv .venv
. .venv/bin/activate
pip install .
cp recipes/local-main-critic.yaml my-recipe.yaml
# Edit endpoints/model ids and export referenced keys. Never put keys in YAML.
fusion-runtime --config my-recipe.yaml --port 18888
Point an OpenAI-compatible client at http://127.0.0.1:18888/v1 and select
fusion-coding:
curl http://127.0.0.1:18888/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"model":"fusion-coding","messages":[{"role":"user","content":"Review this patch"}]}'
Use recipes/review-board.yaml to mix providers
and define role-named experts. A pool can use local models, hosted APIs, or both;
the runtime never inspects GPU type or guesses capability from model names.
Configuration model
A recipe has five explicit layers:
providers: transport endpoints and environment-variable credential refs;models: provider model ids plus declared capacity and capabilities;pools: one authoritative main model and role-to-expert assignments;policy: orchestration, expert-call budget, and policy-specific options;serve: one public model name and enabled client protocols.
Experts are advisory-only. They receive no coding tools, their output is
bounded and marked untrusted, and only the main model can produce the public
answer or tool call. A failed expert is surfaced through x-fusion-fallback;
it cannot silently become the writer.
Plugin contract
Provider packages register a factory under fusion_runtime.providers; policy
packages use fusion_runtime.policies:
[project.entry-points."fusion_runtime.providers"]
my-provider = "my_package:MyProvider"
[project.entry-points."fusion_runtime.policies"]
my-policy = "my_package:MyPolicy"
A provider implements async complete(model, request) -> ModelResponse and
stream(model, request) -> AsyncIterator[ModelStreamEvent]. A policy implements
async prepare(runtime, pool_name, request) -> PreparedCall: experts finish in
prepare, while the runtime owns the sole final call in complete or streaming
mode. Protocol translation stays at the gateway boundary and must not own
routing. A provider without stream fails a streaming request visibly instead
of silently falling back to buffered output.
DeepSeek Harness
integrations/deepseek-harness is a real
DeepSeek Harness profile bundle using its official generic OpenAI-compatible
LLM seam. It is pinned to the current developer-preview contract; see that
directory for installation and compatibility notes. The integration is
community-maintained and does not claim upstream endorsement.
Controlled RSI
RSI in this project means evaluation-gated improvement of recipes, prompts, budgets, stopping rules, and completion gates. It does not mean an online model may rewrite production code, configuration, or weights.
After evaluating a candidate and baseline on the same frozen task set, seed, and environment, compare their JSON summaries:
fusion-runtime-gate \
--baseline cards/direct-summary.json \
--candidate cards/review-board-summary.json
The command exits 0 only if every quality, latency, cost, infrastructure, and
reproducibility gate passes; otherwise it exits 2 with explicit reasons.
Provider-neutral runtime code belongs in core. Vendor SDKs, custom routers,
training backends, and additional harness adapters belong in plugins. See
CONTRIBUTING.md and SECURITY.md.
Performance claims
Fusion is not automatically better than direct inference. Publish a recipe only with a reproducible card comparing direct and fusion modes on identical tasks, seed, environment, latency, token cost, and infrastructure-failure accounting. If a candidate does not clear its declared objective, keep the direct route.
License
Apache-2.0.
No comments yet. Be the first to write one.