BigMoeOnEdge/docs/seam.md
Helldez 49c72e7ce5
feat(engine): self-speculative decoding — the model's own MTP head, or n-gram lookup (#134)
* feat(engine): MTP self-speculative decoding for Qwen3.5/3.6 (proposal)

Qwen3.5/3.6 ship a trained multi-token-prediction block inside the gguf. With
--mtp that head drafts --mtp-draft continuation tokens and the target verifies
all of them in one wider decode, confirming the longest prefix whose argmax
equals what the target itself would have produced. Nothing is approximated and
no weight is skipped, so the quality is the full model's — but it is NOT
byte-identical the way --overlap and --prefetch are, and must not be used in a
byte-identity gate: a verify pass evaluates 1+N positions in one batch, and a
batched matmul is not bit-identical to N single-token ones, so a near-tie can
flip. Off by default.

The prize is that a decode's dominant cost, moving the dense weights and the
routed expert slices, is paid once per group instead of once per token. The
counterweight is that the verify positions route independently, so a layer's
read set widens toward N*k wherever adjacent tokens disagree, and the draft
pass routes through the MTP block's own expert layer on top. Measured on the
desktop host (DRAM-bandwidth-bound, model streamed at ~1.4x RAM): +15.1% at
draft 3 with the host's best recipe (7.12 -> 8.19 tok/s), +29% without the
lossy drop knob, acceptance falling from 71% at draft 2 to 52% at draft 4, and
flash bytes per token rising 19.7 -> 33.7 MiB as the widening predicts. Draft 3
is the optimum here; 4 is worse than 2. On a flash-I/O-bound phone that balance
can invert, so the flag ships off pending the device A/B.

The orchestration is llama.cpp's own (common/speculative.h, public headers
only): no fork, no patch, no submodule bump. Self-speculation is one model with
two contexts over it — the target, created with n_rs_seq so a rejected tail is
rewound from a bounded snapshot rather than replayed, and a draft context
created with ctx_type = LLAMA_CONTEXT_TYPE_MTP. The engine builds the draft
context itself rather than through common_speculative_init_from_params because
the eval callback is per-context: the streamer only sees the MTP block's expert
layer if the draft context carries the same cb_eval.

The MTP block is streamed like any other layer. It sits at layer index n_layer,
contiguous with the trunk and using the same tensor naming, so the hook and the
expert source are sized n_layer + n_layer_nextn; left at n_layer its experts
stay silently mmap-resident. Two consequences that are easy to get wrong: the
capture warm-up has to run on the draft context too (the MTP graph is built
nowhere else), and prefill is fed through the driver so the draft context's KV
reaches the last prompt position.

The loop accepts BEFORE catching the draft context up, so the catch-up runs on
the accepted prefix instead of the whole verify batch. Acceptance depends only
on the target's logits, which are already in hand once the decode returns, and
the rejected tail was being computed only to be deleted a few statements later.
The resulting state is identical — the driver seeds from row
min(n_accepted, n_rows-1), the same row under either batch, and the surviving
KV is exactly the range the rollback used to carve out — while skipping
n_draft - n_accepted positions through the MTP block per group. Since that
block carries its own MoE FFN, on a streamed device those are expert reads that
no longer happen. It also removes the draft context's rollback entirely: it is
never given a tail to drop.

Requires an MTP-converted gguf (most quantisations strip the nextn tensors) and
greedy decoding; both are rejected at load with a message rather than silently
ignored, as is a n_ubatch narrower than the verify batch, which would split the
graph back into single-token passes and spend the draft for nothing.

Telemetry: an "mtp:" summary line, an mtp_batch per-token CSV column (a verify
decode's whole cost is charged to its group's first row, the rest carry zeros),
mtp_drafted / mtp_accepted / mtp_decodes in the CSV trailer and in BMOE_DONE,
and mtp / mtp_draft_max in the CSV preamble. The Android app exposes the flag
and the draft width, off by default.

Host gates pass. Validated on Qwen3.6-35B-A3B-MXFP4 with the streamed recipe:
draft 1 and draft 3 produce identical text, which is the invariant a broken
accept/rollback path would violate. Device A/B still owed.

* perf(mtp): shrink the draft context, make its cost measurable, record the device verdict

The first on-device A/B says MTP loses at every draft width, and the counters
say why. Same gguf with the flag on and off, shipping recipe (overlap, 3000 MiB
cache, pinned dense, drop 0.75), Qwen3.6-35B-A3B-Q4_K_M streamed:

    off             5.82 / 6.14 tok/s   69.3 MiB/tok    69-109 majflt/tok
    --mtp-draft 2   5.59                93.8            230
    --mtp-draft 3   4.38               106.6            633

Speculation is working - 2.35-2.52 tokens per verify decode, 52-69% acceptance
- and still losing, because the prize does not exist in this regime.
stall_s/tok is 0.025-0.027 in every one of those runs, MTP on or off: 11-16% of
the token. This configuration is compute-bound, and what MTP amortises is weight
movement. The costs meanwhile are real and monotonic in the draft width: the read
set widens (+35%, +54% flash bytes per token), CPU per token rises (+28%, +67%),
and the draft context's memory tips the device into a fault storm.

Two things follow, and both are engine bugs rather than facts of nature.

The draft context's graph width drops from 256 to 32. Compute buffers are
reserved for the widest ubatch and the dominant term scales with
ubatch x vocabulary; on device that reservation measured 493 MiB - for a context
that evaluates ONE token per draft step and is handed at most 1 + draft_max
positions by the catch-up, with no logits asked for. Only prefill ever feeds it a
wide batch, and that is one layer, so splitting it costs very little. On this
engine memory is never free: it is the expert cache's, and the cache is what
decides whether the widened verify read set is a hit or a flash read.

And the cost of speculation is now measured instead of inferred. Drafting happens
between decodes, so it never entered wall_ms and tok/s never included it - a
speculated run could report a rate the user was not experiencing. New
mtp_draft_ms per-token column (a slice of loop_overhead_ms, not an addition),
mtp_draft_s/tok in the CSV trailer, mtp_draft_s_tok and loop_overhead_s_tok in
BMOE_DONE, and a second "mtp:" summary line printing the effective rate next to
the reported one.

Adds --mtp-p-min F, which stops drafting once the head's confidence in what it is
proposing falls below F. The draft loop already had this floor and the engine was
passing 0, so it always drafted the full width however unsure the head was - with
roughly half the drafts rejected at draft 3, that is the cheapest waste available
to cut. On a streamed device it pays twice: a draft not made is a pass through the
MTP block (which carries its own MoE FFN, so its own expert reads) that never
happens, AND one fewer independently routed position in the verify batch. Default
0, the setting the host numbers were measured at; the useful value is a property
of a device's balance between drafting cost and acceptance, so it is a knob to
measure rather than a constant to guess.

The Android app now reads the mtp_* keys it was already being sent: acceptance,
tokens per pass, and the effective rate. Before this the UI could not tell whether
speculation had run at all - only the session CSV could - which made the A/B this
commit reports impossible to run from the phone.

Neither mitigation changes the regime. The honest expectation is nearer
break-even, not a win, and the flag stays off by default.

Host gates pass. Note the noise floor: the two off runs did byte-identical work
and still differ by 5.6% in tok/s, and the runs were back-to-back without thermal
gating - the mechanism counters are the trustworthy part, not the exact deltas.

* perf(mtp): split the drafting flash cost from the widened verify batch

A speculated run streams more bytes per token for two unrelated reasons: the
MTP block carries its own MoE FFN, so every draft pass routes experts of its
own, and the verify batch widens the trunk's read set wherever adjacent
positions disagree. They need opposite fixes -- a narrower draft attacks the
first, only better agreement attacks the second -- and the route trace can
separate neither, since its framing brackets the target decode while the head
only ever runs in the draft context.

Measure the head's share directly by bracketing both drafting passes with the
expert source's byte counter, and report it as a third mtp: summary line.

Also record the branch-deletion rule in AGENTS.md: a branch list should only
show work in flight, and a rejected PR loses nothing.

* feat(engine): n-gram prompt-lookup draft source, and the measurement that closes it

The flash split added last commit said where MTP's cost actually is: at draft 3 on
the host, the head's own routing was 2.9% of the extra bytes a speculated run
streams and the widened verify batch was the other 97.1%. So a cheaper draft
producer is worth almost nothing, and the only property that could matter is one
the head does not have -- the ability to decline to draft at zero cost.

--ngram is that source. It takes the last few tokens, finds where that run occurred
before in the prompt or in what has been generated, and proposes whatever followed.
No head, no draft context, no decode, no expert read, and it works on any gguf
including the ones --mtp refuses for want of a nextn block. Below --ngram-min-match
it proposes nothing and the step falls through to a plain single-token decode.

Measured on the host, Qwen3.6-35B-A3B-MXFP4 streamed, 256 greedy tokens, cells
back-to-back with off run twice:

    prose        off 5.80 / 6.59    mtp3 7.32 eff    ngram3 6.51  (cov 7.4%)
    copy-heavy   off 5.45 / 5.65    mtp3 6.43 eff    ngram3 5.24  (cov 15%)

The zero-cost claim holds exactly -- mtp_draft_s/tok reads 0.0000 in every n-gram
cell, against 0.020-0.023 for the head plus the ~500 MiB of expert cache its draft
context takes. But the floor turns out to be per STEP, not per run: the 15% of steps
that did draft widened the read set to 67.2 MiB/token against 48-58 at baseline and,
at 44% acceptance, bought 1.20 tokens per decode. That is not enough to earn the
widening back, and a modest fraction of such steps sinks the run.

A --ngram-min-match sweep settles it rather than leaving it open. Raising the gate
3 -> 5 -> 8 lifts acceptance 44% -> 75% while coverage collapses 15% -> 3.4%, and
narrowing to --draft 1 reaches 82.6% -- the head's own figure on this prompt. Every
cell climbs toward baseline from BELOW and none crosses it; the best configuration
found lands on the floor. A knob whose optimum is its own disablement is not a
tuning problem. Acceptance, not drafting cost, is what pays for a widened batch, and
what a trained head buys is being right often enough to justify a batch that has
already been widened.

--ngram ships off. It is kept because it is the only speculation available on a
model with no head, because the per-step floor is real, and because the counters it
adds make the next speculation claim falsifiable.

Wiring. MtpConfig became SpecConfig with DraftSource {none, mtp, ngram}, and
--mtp-draft became --draft: the width belongs to the verify batch, not to whoever
filled it. --mtp and --ngram are rejected together rather than resolved by flag
order. In the session the gate split in two -- spec_on (wide batch, acceptance,
rollback: both sources) against mtp_on (draft context, common/speculative.h, the
catch-up: the head only) -- which is what lets the n-gram source reuse the whole
verify half while allocating nothing.

A step that drafts nothing now takes the plain path: llama_batch_get_one with a
logits row of -1, byte for byte the unspeculated decode. It used to build the wide
batch anyway. Required for --ngram, and it tightens --mtp-p-min's zero-draft steps
for free.

The matcher is pure policy over token ids with no llama.cpp at all -- not even
llama.h, since llama_token is int32_t -- so it sits on the clean side of the seam,
adds no dependency on the common layer, and is unit-tested with no model
(tests/ngram_test.cpp covers tie-breaks, clipping, self-match exclusion and the gate
boundary). Telemetry: spec= / spec_draft_max= / ngram_min_match= in the CSV
preamble, a new drafted_steps key in the trailer and BMOE_DONE, and an ngram: line
reporting coverage -- without which a delta cannot be divided by the fraction of the
run it applies to. The per-token and trailer counters keep their mtp_ names: they
always described the loop rather than a source, spec_* already means the temporal
prefetch in that trailer, and renaming would break every CSV already holding a
measurement. The Android setting became a three-way picker, migrating the old
boolean preference.

The device A/B agrees and adds a cost the host could not show. Thermally gated cells
(a 120 s settle, then a battery-temperature gate, so all six start between 35.3 and
36.4 C): prose 4.90 inside a 4.59-5.17 band, copy-heavy 3.14 against 4.43 -- a 29%
loss, worse than MTP's 18%. Major faults per token go 126 -> 1427 for a source that
allocates no draft context at all, and that is the rollback snapshots: n_rs_seq =
draft_max is asked for by ANY speculation, since rejecting a draft means rewinding the
KV, and on a hybrid attention/SSM model that snapshot is a real allocation scaling with
the context. The n-gram source escapes MTP's draft context but not the loop's own
memory, and on device that memory is the expert cache's.

The same run re-measured MTP with the thermal confound removed -- 3.64 effective
against 4.43, so the earlier device verdict was not an artefact of benching without a
cooldown gate -- and reproduced the flash split at 3.7% head against 96.3% widened
verify batch, matching the host's 2.9-3.0%.

Byte-identity gates pass; speculation stays out of them for the reason docs/mtp.md
gives.

The app's CSV configuration surface follows: the three new preamble keys get their own
glossary entries rather than falling through to the unexplained-key renderer, and the
draft source joins the short run label. A speculated run is not the same KIND of run --
under speculation a decode confirms a whole group, so its per-token rows are not even
accounted the same way -- and two compare legends differing by it must not read alike.
2026-08-02 00:09:39 +02:00

12 KiB

The seam: how we hook llama.cpp without forking it

Everything that connects BigMoeOnEdge to llama.cpp goes through two public mechanisms. This file documents the exact contract so it can be re-verified when the submodule is updated.

1. The eval-callback

llama_context_params.cb_eval / cb_eval_user_data (public) install a function called by ggml_backend_sched for every graph node:

  • callback(node, ask=true, ud) is called for each node. Returning true isolates that node: the scheduler computes it alone, ggml_backend_synchronizes, then calls callback(node, ask=false, ud).
  • Returning false groups the node with its neighbours for normal computation (no non-ask callback).

We use both phases:

Capture phase (one warm-up decode). ask is called for every node, so we scan each node's src[] for expert weight tensors (blk.<il>.<suffix>.weight, where the suffixes come from the arch's recipe — ffn_{gate,up,down}_exps for the split layout, a fused ffn_gate_up_exps for others) and record the live ggml_tensor*. We return false throughout — capture observes, it does not isolate. ggml_tensor is a public struct, so reading ->name, ->ne, ->nb and writing ->data is public API surface.

Stream phase (real generation). We return true for ffn_moe_topk-<il>. The non-ask callback then hands us that node with the selected expert ids materialized; we gather them (stride-aware) and trigger the slice reads.

Two optional jobs ask for more: the route trace and cache-aware dropping also want each layer's ffn_moe_weights*-<il> chain, which is another barrier per node but no new kind of access — same public struct, same read of ->data.

Dropping does go one step further, and it is the only place the engine writes into a graph tensor's contents rather than repointing ->data at its own buffer: at the terminal node of the weight chain it zeroes a dropped slot's weight and repoints that slot's expert id. Both tensors are scratch the graph produced and has not yet consumed, so this alters the values flowing through the run — deliberately, that is what the lossy policy is — and never llama.cpp's own state, its weights, or its control flow. It stays inside the same callback contract; nothing is patched.

2. gguf offsets

gguf_init_from_file(..., no_alloc=true) + gguf_get_data_offset + gguf_get_tensor_offset (all public) give each tensor's absolute byte offset in the file, without loading any tensor data. We match these to the captured tensors by name.

3. The expert-ready hook (fork extension)

Sections 1 and 2 are enough for the serial streamer: block on the expert reads, then let the layer compute. Overlapping the two — reading a token's experts while the same token's expert matmuls are running — needs a wait point that no public API exposes. That is the one place where BigMoeOnEdge carries a llama.cpp extension.

What it is. A single optional hook, ~25 lines, living on the fork branch bmoe/expert-ready-hook of Helldez/llama.cpp as a single commit on top of the upstream pin. It adds nothing to the model files and changes no data layout; it is a callback the CPU MoE kernel invokes.

Exact API and call site.

// ggml-cpu.h
void ggml_cpu_set_expert_ready_hook(ggml_expert_ready_hook_t hook, void * user_data);

ggml_compute_forward_mul_mat_id calls the hook at the top of its per-expert loop, right after the "expert not routed this token" skip, before it consumes that expert's weight slice. Every compute thread calls it for every routed expert; the hook may block. There is no barrier inside the expert loop, so a thread blocking on one expert cannot deadlock the threadpool — other threads proceed to the experts whose slices are already resident. The streamer's hook blocks until the requested expert slice has been read in, then returns. When no hook is registered (stock upstream, or --overlap off) the call is a single null check — zero cost.

Why it exists. The topk eval-callback (section 1) is the only public hook near routing, and it can only fire before the expert matmuls of a layer run — it cannot pause partway through them. Overlapping expert reads with expert compute requires a per-expert wait point inside the kernel, which the public API does not provide. Hence the extension.

Graceful degradation. CMake probes ggml-cpu.h for the hook symbol and, when present, defines BMOE_HAVE_EXPERT_READY_HOOK. Built against stock upstream (symbol absent) the whole project still compiles and runs — the serial streaming path is unchanged; only --overlap is affected, and it fails with a clear runtime error instead of silently falling back.

Sunset condition. This fork exists solely for this one hook. The moment upstream ships an equivalent per-expert readiness/residency callback, the branch is dropped and the submodule bumps straight back to ggml-org/llama.cpp. It is a tide-me-over until the wait point is public, not a divergence we intend to maintain.

The chat glue: llama.cpp common (not the streaming seam)

Separate from the two streaming hooks above, session.cpp links llama.cpp's common library for two things: rendering the model's own chat template and parsing reasoning output, and (with --mtp) driving speculative decoding. common_chat_templates_init / common_chat_templates_apply run the real Jinja template the gguf ships (so Gemma's channel format, Qwen ChatML, etc. all format correctly, driven by the model rather than hardcoded), and common_chat_parse extracts a reasoning model's thinking so it can be reported apart from the answer. The parser-params wiring lives in its own translation unit, chat_parse.cpp — the PEG parser arena has to be loaded explicitly or common_chat_parse throws on the first token, which is how issue #49 stayed invisible; keeping it separate makes that seam unit-testable without a model.

A second translation unit, thinking_control.cpp, crosses the same boundary for "thinking off". enable_thinking is only a request to the template, and many templates never read it, so the engine renders the template to find out (three renders at open, no model names involved) and, where the flag is inert and reasoning is a structural section of the format, asks for a continuation instead: the continue_final_message field of common_chat_templates_inputs, plus a synthetic trailing assistant message, makes llama.cpp's own per-template handler emit that family's "reasoning is over" span into the prompt. This is why no <think> or harmony channel marker appears anywhere in core/ — the markers stay upstream, where a submodule bump keeps them current.

Whether the continuation is binding is read off common_chat_params::thinking_start_tag/ thinking_end_tag: a model that declares a reasoning span owns it, so a pre-closed empty one is a suggestion it can decline (LFM2.5 does), while a model that declares none separates reasoning structurally and cannot. Both facts come from the loaded model, never from its name. tests/think_control_test.cpp pins all of it against the vendored templates, again with no model.

Speculative decoding

Only the MTP source crosses into common. The verify half of the loop — the wide batch, the argmax acceptance, the KV rollback — is written against public llama.h alone, and so is the n-gram draft source, which is why --ngram needs nothing from common at all. core/include/bmoe/ngram_draft.h does not even include llama.h: it is written over int32_t token ids, which is what llama_token is, so the drafting policy stays on the pure-policy side of the seam and is unit-tested with no model and no native backend (tests/ngram_test.cpp). See ngram.md.

--mtp reaches common/speculative.h, which is a header of the same common library and includes only llama.h and common.h. The draft/verify orchestration — running the trained MTP head, moving hidden states from the target to it, seeding the next draft from the accepted position — is entirely upstream's; the engine supplies the loop around it and the two contexts it works on. Worth stating explicitly because the internals it needs (llama_set_embeddings_nextn and the nextn hidden-state getters) live in src/llama-ext.h, a staging header: they are used inside speculative.cpp, which is already compiled into the llama-common the engine links, so nothing in core/ includes a private header and no in-tree patch is involved.

What the engine does own is the pair of contexts. Self-speculation is one model with two contexts over it — the target, and a draft created with ctx_type = LLAMA_CONTEXT_TYPE_MTP — and the engine builds the draft one itself rather than through common_speculative_init_from_params, for a reason that belongs to this project: the eval callback is per-context. The streamer only sees the MTP block's expert layer if the draft context carries the same cb_eval, and init_from_params derives its context parameters from a full common_params with no way to inject one. See mtp.md.

Unlike the public-C-API streaming seam, common is not a stable API — it can change between upstream versions. So a submodule bump may require updating this chat glue in session.cpp / chat_parse.cpp / thinking_control.cpp; the build and gates catch a break at compile time rather than at runtime (tests/chat_parse_test.cpp and tests/think_control_test.cpp cover these seams directly). This trade-off is deliberate and is also noted at the link site in the root CMakeLists.txt. The gates themselves run with the template off (raw prompt), so they stay deterministic and are unaffected by this dependency.

The one ggml behaviour we depend on

That a node marked "needed" is computed and synchronized before the non-ask callback, and that the batch containing the dependent expert matmul runs after the callback returns. This is how ggml_backend_sched implements the eval-callback today (ggml/src/ggml-backend.cpp). It is not a stability-guaranteed contract, so:

  • the byte-identity gates assert lossless output, which fails loudly if the ordering ever changes;
  • CI runs the gates on every submodule bump.

Upgrading llama.cpp

Because the submodule pins the bmoe/expert-ready-hook fork branch (section 3), a bump rebases that 1-commit branch onto the new upstream tag, re-pushes it, and re-pins:

# in a Helldez/llama.cpp checkout: rebase the single hook commit onto the new tag
git fetch upstream && git checkout bmoe/expert-ready-hook
git rebase <newer-upstream-tag> && git push --force-with-lease origin bmoe/expert-ready-hook

# in this repo: move the submodule to the rebased commit, rebuild, run the gates
cd third_party/llama.cpp && git fetch origin && git checkout <rebased-commit>
cd ../.. && git add third_party/llama.cpp && scripts/build-host.sh
cd build && ctest --output-on-failure     # gates must stay green

When the sunset condition lands (upstream ships the readiness callback) this collapses back to a plain git checkout <newer-tag> against ggml-org/llama.cpp with no branch to carry. Either way the gates are the enforcement. The fragility is not the public API but the internal naming conventions the seam attaches to — the tensor suffixes and the ffn_moe_topk node name are how llama.cpp happens to build MoE graphs today, not a guaranteed contract, so upstream can rename or restructure them (Gemma 4's fused ffn_gate_up_exps is one such evolution we absorbed with a recipe row). The gates are the enforcement: a rename breaks byte-identity before merge instead of silently corrupting output. Each supported architecture adds one more gate to keep green across a bump.

If a future release moves the two hooks (a stable expert-residency API, say) upstream, this seam shrinks further or disappears — core/ does not change.

Pinned submodule at the time of writing: Helldez/llama.cpp branch bmoe/expert-ready-hook, commit 5236140 — the single expert-ready-hook commit (section 3) on top of upstream ggml-org/llama.cpp master 22b69b6 (see .gitmodules / git submodule status for the current pin).