BigMoeOnEdge/docs/moe-streaming.md
Raffaele 47924565c1
feat(io): release the model file's mapping after load (--release-mmap) (#185)
llama.cpp maps every gguf it loads and keeps the mapping for the model's lifetime. On
Windows that is expensive in a way nothing had attributed: while a section of a file is
alive, NTFS serialises concurrent unbuffered reads on that file, and a lane opened while
the section existed keeps serialising against it after the section is gone. Four I/O lanes
therefore delivered exactly one lane's throughput, which is why lanes and threads have
always measured dead on the desktop host and why the engine read at about a third of what
the drive can serve.

--release-mmap hands the mapping back after load: unmap the file, close its section, reopen
the reader lanes. Both halves are needed. Whether it is safe is decided by looking rather
than by reasoning, the engine asks the OS whether any weight the capture pass observed still
points inside a mapping of the model files, and declines if any does. Off by default,
because the check answers for the pointers the capture saw and for no others.

Host A/B on Qwen3.6-35B-A3B Q4_K_M: 3.16 to 4.63 tok/s (+46%), flash stall per token 0.182
to 0.074, with bytes read, hit rate, evictions and re-reads identical to the digit and the
generated text byte-identical. On the phone the read path is flat (f2fs does not serialise)
but CPU per token falls about 9%; that cell is two short runs per variant and is recorded as
a direction, not a number.

Also fixes a bug this uncovered, independent of the flag: when a gguf carries no
output.weight, llama.cpp builds the output head from the token embedding table and the model
holds two identically named tensors over the same bytes. The capture pass keyed its map by
name, so --dense-weights anon and ahwb rebound one and left the twin reading the mmap for
the whole run, which on a model past RAM means the output projection served by page faults
from flash. The capture now records every distinct leaf object by address and the dense
policy rebinds every tensor over one file range onto the same buffer.

Adds bmoe-iobench --mmap / --reopen-lanes / --range-mb / --fresh, the cells that isolate the
mechanism, a mapping_release unit test on both platforms, the app switch "Release the model
mapping", and the bench findings. README, architecture, AGENTS and roadmap updated, the last
correcting a diagnosis this refutes.
2026-09-07 20:43:02 +02:00

7.2 KiB
Raw Permalink Blame History

MoE expert-selective streaming

The lever

A MoE layer stores n_expert experts (128 for Qwen3-30B-A3B) but each token is routed to only its top-k (8). The other experts' weights are never read for that token. If the model does not fit in RAM, streaming just the routed experts from flash turns "the whole expert bank per token" into "top-k/n_expert of it" — about 6% for that model.

This sparsity is real only for autoregressive, one-token-at-a-time decoding. A batch of T tokens routes the union of T × k experts, which approaches all of them for any useful T. So streaming deliberately runs at n=1 and is incompatible with speculative decoding or a canvas — the engine keeps decode single-token by construction.

Mechanism

  1. Bind. After a one-token warm-up capture (see seam.md), every layer's three expert tensors (ffn_{gate,up,down}_exps) are rebound onto streaming buffers and never read from the mmap again.
  2. Route. The eval-callback marks only the routing node ffn_moe_topk-<il> as needed. ggml computes it alone, synchronizes, and calls back with the selected expert ids. They are gathered respecting the view strides — selected_experts is a view of the full argsort with row stride nb[1], so a flat read would grab the wrong experts and corrupt the KV cache.
  3. Load. The expert source reads exactly those experts' slices from the gguf (O_DIRECT, page cache bypassed) into each expert's canonical offset inside the bound tensor, just before that layer's expert matmul runs.

Ordering is guaranteed by ggml's eval-callback loop: the node we mark is computed and ggml_backend_synchronize'd before the non-ask callback fires, and the following compute (the expert matmul) runs only after our load returns. The next layer cannot overwrite the buffers until this layer's matmul has synchronized. Correct on any backend.

The result is lossless: byte-identical to running with every expert resident, asserted by the gates. That is the streaming path itself; two opt-in knobs deliberately trade output for speed on top of it — --n-expert-used (fewer experts per token) and --drop-cold-experts (skip an expert that would cost a read and was barely weighted). Both are off unless asked for, which is what keeps the sentence above true by default.

Residency modes

  • Cache off (shared slots). Three heap buffers (full n_expert size) are shared across layers — one layer computes at a time. Routed slices are re-read fresh every token. Lowest RAM, highest I/O.
  • LRU cache (--cache-mb N). Each (layer, projection) gets a reserved, lazily-committed address range. A routed expert already resident is a hit (no read); a miss is read once and kept; over budget, the coldest (layer, expert) is evicted and its pages physically released (madvise(MADV_DONTNEED) / MEM_DECOMMIT). RAM is bounded for real.

The cache rule: 0 or ≥ ~2 GB

Expert reuse is broad, not skewed: hit rate rises roughly linearly with budget, with no small-cache plateau. A budget below one token's routed working set (~1 GB for Qwen3-30B-A3B) yields zero hits and pays eviction overhead — measurably slower than no cache. So validate() rejects a budget in the 1..1499 MiB band unless you force it. Use 0, or ≥ 2000.

Parallel reads (--io-threads N)

Routed slices are read across N lanes, each with a private fd and bounce buffer; the calling thread participates as lane 0. On UFS 4.x, 4 lanes roughly triples effective read bandwidth over serial. Compute threads (-t) show a U-shape — 4 is the measured optimum; 8 regresses badly because ggml's spin-wait contends with the synchronous reads.

The model file's mapping (--release-mmap)

llama.cpp maps the gguf and keeps it mapped for the model's lifetime. On Windows that mapping serialises the lanes above: while a section of the file is alive, concurrent unbuffered reads on it are taken one at a time, so --io-threads 4 reads at one lane's rate. It is not the drive and it is not the engine — bmoe-iobench --model M.gguf --lanes 4 --slice-kb 576 measures 2400-2660 MiB/s, the same command with --mmap measures 895-930, and adding --reopen-lanes recovers the full rate. A lane opened while the section existed stays serialised after it is gone, which is why the recovery needs both halves.

--release-mmap does exactly that inside the engine: after load, once nothing reads through the mapping any more, the file is unmapped, its section closed and the reader lanes reopened.

Whether releasing is safe is decided by looking, not by reasoning about which tensors ought to have been rebound: the engine asks the OS whether any weight the capture pass observed still points inside a mapping of the model files, and declines if any does. That one question covers every residency policy — a dense set left mmap'd under mmap or warm, a table held back as oversized, or a tensor no name-based accounting could have found. Run with --dense-weights mmap and the engine reports the count and stands down.

On Windows the run ends with warning: UnmapViewOfFile failed, printed by llama.cpp rather than by the engine. That is the designed outcome, not a defect: llama.cpp is not patched and still believes it owns the mapping, so at teardown it unmaps a base the engine has already released. An inert reservation is left in that range precisely so the call finds a placeholder and fails harmlessly, instead of finding whatever was allocated there next.

It is still opt-in, because the check answers for the pointers the capture pass saw and for no others. A graph shape this session never builds could hold another one, and llama.cpp exposes no way to enumerate a loaded model's tensors and settle it. The MTP draft is the concrete case: it builds a second graph, so a gguf tensor no policy owns blocks the release there as well. Worth +46% decode on the desktop host, byte-identical.

Android is a different story with the same conclusion. The iobench cells above are flat there — f2fs does not serialise, so there is no read bandwidth to recover, and the engine's flash stall is unchanged with the flag and without. What changes is CPU: a 20 GB mapping the kernel still has to account for costs about 9% of the decode's CPU time on a device under memory pressure, and dropping it is worth 5-9% of throughput. That measurement is two short cells per variant and is a direction, not a number. The flag stays off by default on both platforms.

Why repack must stay off

The streamer rebinds tensor->data to a buffer it fills from the file's native byte layout. use_extra_bufts=true would repack Q4_K weights into a different in-memory layout (e.g. q4_K_8x8), so the file offsets would no longer describe what the matmul reads. The engine loads with use_extra_bufts=false; this is load-bearing, not a tuning knob.

Assumptions to re-check on a submodule bump

  • The routing node is named ffn_moe_topk-<il> and the expert tensors blk.<il>.ffn_{gate,up,down}_exps.weight. The recipe isolates these names.
  • The eval-callback fires per decode (not skipped by graph reuse) and computes a marked node alone before the non-ask callback. The gates catch a regression here.