BigMoeOnEdge/docs/seam.md
Helldez bea5a0b99e
feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95)
* feat(moe): --drop-cold-experts, spend quality only where it buys I/O

Turbo top-k drops the tail of a routing whether or not those experts were
already in RAM. A resident expert costs no flash read, so that trade pays
quality for nothing on the ~80% of decode routings that are cache hits.

This skips a routed expert only when it is a cache MISS and the router
weighted it below frac x (1/top-k). Replayed over the committed route
traces at frac 1.0, decode phase, that avoids 66% of flash reads for 9.5%
of the router's weight mass, where --n-expert-used 5 avoids 23% for a
comparable 10.6% -- about 3x the reads at the same quality cost.

Implementation. The decision needs the FINAL router weights, which arrive
several nodes after the topk where the streamer normally loads, so
load_layer() is deferred to the terminal node of the layer's weight chain.
Which node that is depends on the model's gating, so the hook learns it
from the graph rather than carrying an architecture table; if it fails to
arrive the hook forgets it and re-learns rather than re-betting. A dropped
slot has its weight zeroed and its expert id repointed at the routing's
top-weighted expert: an unread expert can sit in reserved-but-uncommitted
VM and mul_mat_id would touch it anyway, so the kernel is given memory
that is certainly resident and multiplies it by exactly zero.

Requires the LRU cache -- with --cache-mb 0 residency reads all-miss and
the policy would silently degenerate into an unconditional weight cut.
Prefill is excluded by default. The top expert is always pinned, so no
routing can be emptied at any threshold.

Gates: G8a/G8a' prove the deferral and the learned terminal node are
transparent (byte-identical output, zero drops, at a threshold below any
producible weight); G8b that full strength against a constantly-evicting
cache never reaches an unloaded slot; G8c that at top-k 1 dropping is a
no-op, pinning both the top-expert guarantee and the threshold tracking
the effective top-k.

Three existing metrics shift meaning under dropping and the docs now say
so: cache_hit_pct rises without the cache serving more (a dropped routing
is a miss that is never looked up), and token/layer_demand measure what
was staged rather than routed. prefetch.md's "cannot change output" is
scoped, limitations.md gains the non-reproducibility entry, and
benchmark-method.md warns that reversing the run order cannot distinguish
a moved drop rate from a contaminated cell.

Off by default in the CLI and in the app. The output is not reproducible
-- what gets dropped depends on what the cache held -- so it carries no
rows in the README tables, and switching it on by default waits on a
published on-device A/B rather than on the replay argument alone.

* feat(app): default cache-aware dropping to 75%, measured on device

Qwen3.6-35B-A3B (top-k 8 of 256), in-app, cache 3000, one variable
changed: 2.549 tok/s off, 3.938 at F=0.75 (+55%), 4.702 at F=1.0 (+84%),
with flash reads falling 248 -> 163 -> 48 GiB. Per-token bootstrap
intervals separate every pair except off vs 0.50, which overlaps -- at
half the uniform share the policy drops 2.7% of routings and buys
nothing, which doubles as a negative control that the machinery is free
when it does not fire.

Run order was 1.0, off, 0.5, 0.75, so the two fastest cells are the first
and the LAST; thermal drift would have made the last the worst. The
mechanism orders by threshold even though the run order does not.

The replay turned out conservative rather than optimistic. It is
documented as an upper bound because it cannot model the cache changing
in response to dropping: at F=0.75 it was accurate (37% predicted, 34%
measured), at F=1.0 it understated (66% predicted, 81% measured). Avoided
reads free cache capacity, which raises the hit rate, which leaves fewer
misses to drop.

75% rather than 100% is deliberate: it takes the larger part of the win
for half the discarded routings (14% against 28%). Quality is still
unquantified -- no perplexity number and no side-by-side exists -- so the
conservative end of a measured range is the defensible default. The CLI
stays off; the byte-identity gates need a deterministic default.

Also records cache_hit_pct rising 67.8 -> 90.7% as the documented
accounting artefact rather than the cache serving more, and majflt/token
as dominated by each run's starting memory state, not by the threshold.
2026-07-22 17:21:55 +02:00

10 KiB

The seam: how we hook llama.cpp without forking it

Everything that connects BigMoeOnEdge to llama.cpp goes through two public mechanisms. This file documents the exact contract so it can be re-verified when the submodule is updated.

1. The eval-callback

llama_context_params.cb_eval / cb_eval_user_data (public) install a function called by ggml_backend_sched for every graph node:

  • callback(node, ask=true, ud) is called for each node. Returning true isolates that node: the scheduler computes it alone, ggml_backend_synchronizes, then calls callback(node, ask=false, ud).
  • Returning false groups the node with its neighbours for normal computation (no non-ask callback).

We use both phases:

Capture phase (one warm-up decode). ask is called for every node, so we scan each node's src[] for expert weight tensors (blk.<il>.<suffix>.weight, where the suffixes come from the arch's recipe — ffn_{gate,up,down}_exps for the split layout, a fused ffn_gate_up_exps for others) and record the live ggml_tensor*. We return false throughout — capture observes, it does not isolate. ggml_tensor is a public struct, so reading ->name, ->ne, ->nb and writing ->data is public API surface.

Stream phase (real generation). We return true for ffn_moe_topk-<il>. The non-ask callback then hands us that node with the selected expert ids materialized; we gather them (stride-aware) and trigger the slice reads.

Two optional jobs ask for more: the route trace and cache-aware dropping also want each layer's ffn_moe_weights*-<il> chain, which is another barrier per node but no new kind of access — same public struct, same read of ->data.

Dropping does go one step further, and it is the only place the engine writes into a graph tensor's contents rather than repointing ->data at its own buffer: at the terminal node of the weight chain it zeroes a dropped slot's weight and repoints that slot's expert id. Both tensors are scratch the graph produced and has not yet consumed, so this alters the values flowing through the run — deliberately, that is what the lossy policy is — and never llama.cpp's own state, its weights, or its control flow. It stays inside the same callback contract; nothing is patched.

2. gguf offsets

gguf_init_from_file(..., no_alloc=true) + gguf_get_data_offset + gguf_get_tensor_offset (all public) give each tensor's absolute byte offset in the file, without loading any tensor data. We match these to the captured tensors by name.

3. The expert-ready hook (fork extension)

Sections 1 and 2 are enough for the serial streamer: block on the expert reads, then let the layer compute. Overlapping the two — reading a token's experts while the same token's expert matmuls are running — needs a wait point that no public API exposes. That is the one place where BigMoeOnEdge carries a llama.cpp extension.

What it is. A single optional hook, ~25 lines, living on the fork branch bmoe/expert-ready-hook of Helldez/llama.cpp as a single commit on top of the upstream pin. It adds nothing to the model files and changes no data layout; it is a callback the CPU MoE kernel invokes.

Exact API and call site.

// ggml-cpu.h
void ggml_cpu_set_expert_ready_hook(ggml_expert_ready_hook_t hook, void * user_data);

ggml_compute_forward_mul_mat_id calls the hook at the top of its per-expert loop, right after the "expert not routed this token" skip, before it consumes that expert's weight slice. Every compute thread calls it for every routed expert; the hook may block. There is no barrier inside the expert loop, so a thread blocking on one expert cannot deadlock the threadpool — other threads proceed to the experts whose slices are already resident. The streamer's hook blocks until the requested expert slice has been read in, then returns. When no hook is registered (stock upstream, or --overlap off) the call is a single null check — zero cost.

Why it exists. The topk eval-callback (section 1) is the only public hook near routing, and it can only fire before the expert matmuls of a layer run — it cannot pause partway through them. Overlapping expert reads with expert compute requires a per-expert wait point inside the kernel, which the public API does not provide. Hence the extension.

Graceful degradation. CMake probes ggml-cpu.h for the hook symbol and, when present, defines BMOE_HAVE_EXPERT_READY_HOOK. Built against stock upstream (symbol absent) the whole project still compiles and runs — the serial streaming path is unchanged; only --overlap is affected, and it fails with a clear runtime error instead of silently falling back.

Sunset condition. This fork exists solely for this one hook. The moment upstream ships an equivalent per-expert readiness/residency callback, the branch is dropped and the submodule bumps straight back to ggml-org/llama.cpp. It is a tide-me-over until the wait point is public, not a divergence we intend to maintain.

The chat glue: llama.cpp common (not the streaming seam)

Separate from the two streaming hooks above, session.cpp links llama.cpp's common library for one thing: rendering the model's own chat template and parsing reasoning output. common_chat_templates_init / common_chat_templates_apply run the real Jinja template the gguf ships (so Gemma's channel format, Qwen ChatML, etc. all format correctly, driven by the model rather than hardcoded), and common_chat_parse extracts a reasoning model's thinking so it can be reported apart from the answer. The parser-params wiring lives in its own translation unit, chat_parse.cpp — the PEG parser arena has to be loaded explicitly or common_chat_parse throws on the first token, which is how issue #49 stayed invisible; keeping it separate makes that seam unit-testable without a model.

A second translation unit, thinking_control.cpp, crosses the same boundary for "thinking off". enable_thinking is only a request to the template, and many templates never read it, so the engine renders the template to find out (three renders at open, no model names involved) and, where the flag is inert and reasoning is a structural section of the format, asks for a continuation instead: the continue_final_message field of common_chat_templates_inputs, plus a synthetic trailing assistant message, makes llama.cpp's own per-template handler emit that family's "reasoning is over" span into the prompt. This is why no <think> or harmony channel marker appears anywhere in core/ — the markers stay upstream, where a submodule bump keeps them current.

Whether the continuation is binding is read off common_chat_params::thinking_start_tag/ thinking_end_tag: a model that declares a reasoning span owns it, so a pre-closed empty one is a suggestion it can decline (LFM2.5 does), while a model that declares none separates reasoning structurally and cannot. Both facts come from the loaded model, never from its name. tests/think_control_test.cpp pins all of it against the vendored templates, again with no model.

Unlike the public-C-API streaming seam, common is not a stable API — it can change between upstream versions. So a submodule bump may require updating this chat glue in session.cpp / chat_parse.cpp / thinking_control.cpp; the build and gates catch a break at compile time rather than at runtime (tests/chat_parse_test.cpp and tests/think_control_test.cpp cover these seams directly). This trade-off is deliberate and is also noted at the link site in the root CMakeLists.txt. The gates themselves run with the template off (raw prompt), so they stay deterministic and are unaffected by this dependency.

The one ggml behaviour we depend on

That a node marked "needed" is computed and synchronized before the non-ask callback, and that the batch containing the dependent expert matmul runs after the callback returns. This is how ggml_backend_sched implements the eval-callback today (ggml/src/ggml-backend.cpp). It is not a stability-guaranteed contract, so:

  • the byte-identity gates assert lossless output, which fails loudly if the ordering ever changes;
  • CI runs the gates on every submodule bump.

Upgrading llama.cpp

Because the submodule pins the bmoe/expert-ready-hook fork branch (section 3), a bump rebases that 1-commit branch onto the new upstream tag, re-pushes it, and re-pins:

# in a Helldez/llama.cpp checkout: rebase the single hook commit onto the new tag
git fetch upstream && git checkout bmoe/expert-ready-hook
git rebase <newer-upstream-tag> && git push --force-with-lease origin bmoe/expert-ready-hook

# in this repo: move the submodule to the rebased commit, rebuild, run the gates
cd third_party/llama.cpp && git fetch origin && git checkout <rebased-commit>
cd ../.. && git add third_party/llama.cpp && scripts/build-host.sh
cd build && ctest --output-on-failure     # gates must stay green

When the sunset condition lands (upstream ships the readiness callback) this collapses back to a plain git checkout <newer-tag> against ggml-org/llama.cpp with no branch to carry. Either way the gates are the enforcement. The fragility is not the public API but the internal naming conventions the seam attaches to — the tensor suffixes and the ffn_moe_topk node name are how llama.cpp happens to build MoE graphs today, not a guaranteed contract, so upstream can rename or restructure them (Gemma 4's fused ffn_gate_up_exps is one such evolution we absorbed with a recipe row). The gates are the enforcement: a rename breaks byte-identity before merge instead of silently corrupting output. Each supported architecture adds one more gate to keep green across a bump.

If a future release moves the two hooks (a stable expert-residency API, say) upstream, this seam shrinks further or disappears — core/ does not change.

Pinned submodule at the time of writing: Helldez/llama.cpp branch bmoe/expert-ready-hook, commit 5236140 — the single expert-ready-hook commit (section 3) on top of upstream ggml-org/llama.cpp master 22b69b6 (see .gitmodules / git submodule status for the current pin).