* feat(moe): --drop-cold-experts, spend quality only where it buys I/O Turbo top-k drops the tail of a routing whether or not those experts were already in RAM. A resident expert costs no flash read, so that trade pays quality for nothing on the ~80% of decode routings that are cache hits. This skips a routed expert only when it is a cache MISS and the router weighted it below frac x (1/top-k). Replayed over the committed route traces at frac 1.0, decode phase, that avoids 66% of flash reads for 9.5% of the router's weight mass, where --n-expert-used 5 avoids 23% for a comparable 10.6% -- about 3x the reads at the same quality cost. Implementation. The decision needs the FINAL router weights, which arrive several nodes after the topk where the streamer normally loads, so load_layer() is deferred to the terminal node of the layer's weight chain. Which node that is depends on the model's gating, so the hook learns it from the graph rather than carrying an architecture table; if it fails to arrive the hook forgets it and re-learns rather than re-betting. A dropped slot has its weight zeroed and its expert id repointed at the routing's top-weighted expert: an unread expert can sit in reserved-but-uncommitted VM and mul_mat_id would touch it anyway, so the kernel is given memory that is certainly resident and multiplies it by exactly zero. Requires the LRU cache -- with --cache-mb 0 residency reads all-miss and the policy would silently degenerate into an unconditional weight cut. Prefill is excluded by default. The top expert is always pinned, so no routing can be emptied at any threshold. Gates: G8a/G8a' prove the deferral and the learned terminal node are transparent (byte-identical output, zero drops, at a threshold below any producible weight); G8b that full strength against a constantly-evicting cache never reaches an unloaded slot; G8c that at top-k 1 dropping is a no-op, pinning both the top-expert guarantee and the threshold tracking the effective top-k. Three existing metrics shift meaning under dropping and the docs now say so: cache_hit_pct rises without the cache serving more (a dropped routing is a miss that is never looked up), and token/layer_demand measure what was staged rather than routed. prefetch.md's "cannot change output" is scoped, limitations.md gains the non-reproducibility entry, and benchmark-method.md warns that reversing the run order cannot distinguish a moved drop rate from a contaminated cell. Off by default in the CLI and in the app. The output is not reproducible -- what gets dropped depends on what the cache held -- so it carries no rows in the README tables, and switching it on by default waits on a published on-device A/B rather than on the replay argument alone. * feat(app): default cache-aware dropping to 75%, measured on device Qwen3.6-35B-A3B (top-k 8 of 256), in-app, cache 3000, one variable changed: 2.549 tok/s off, 3.938 at F=0.75 (+55%), 4.702 at F=1.0 (+84%), with flash reads falling 248 -> 163 -> 48 GiB. Per-token bootstrap intervals separate every pair except off vs 0.50, which overlaps -- at half the uniform share the policy drops 2.7% of routings and buys nothing, which doubles as a negative control that the machinery is free when it does not fire. Run order was 1.0, off, 0.5, 0.75, so the two fastest cells are the first and the LAST; thermal drift would have made the last the worst. The mechanism orders by threshold even though the run order does not. The replay turned out conservative rather than optimistic. It is documented as an upper bound because it cannot model the cache changing in response to dropping: at F=0.75 it was accurate (37% predicted, 34% measured), at F=1.0 it understated (66% predicted, 81% measured). Avoided reads free cache capacity, which raises the hit rate, which leaves fewer misses to drop. 75% rather than 100% is deliberate: it takes the larger part of the win for half the discarded routings (14% against 28%). Quality is still unquantified -- no perplexity number and no side-by-side exists -- so the conservative end of a measured range is the defensible default. The CLI stays off; the byte-identity gates need a deterministic default. Also records cache_hit_pct rising 67.8 -> 90.7% as the documented accounting artefact rather than the cache serving more, and majflt/token as dominated by each run's starting memory state, not by the threshold.
4.4 KiB
MoE expert-selective streaming
The lever
A MoE layer stores n_expert experts (128 for Qwen3-30B-A3B) but each token is routed to
only its top-k (8). The other experts' weights are never read for that token. If the
model does not fit in RAM, streaming just the routed experts from flash turns "the whole
expert bank per token" into "top-k/n_expert of it" — about 6% for that model.
This sparsity is real only for autoregressive, one-token-at-a-time decoding. A batch
of T tokens routes the union of T × k experts, which approaches all of them for any
useful T. So streaming deliberately runs at n=1 and is incompatible with speculative
decoding or a canvas — the engine keeps decode single-token by construction.
Mechanism
- Bind. After a one-token warm-up capture (see seam.md), every layer's
three expert tensors (
ffn_{gate,up,down}_exps) are rebound onto streaming buffers and never read from the mmap again. - Route. The eval-callback marks only the routing node
ffn_moe_topk-<il>as needed. ggml computes it alone, synchronizes, and calls back with the selected expert ids. They are gathered respecting the view strides —selected_expertsis a view of the full argsort with row stridenb[1], so a flat read would grab the wrong experts and corrupt the KV cache. - Load. The expert source reads exactly those experts' slices from the gguf
(
O_DIRECT, page cache bypassed) into each expert's canonical offset inside the bound tensor, just before that layer's expert matmul runs.
Ordering is guaranteed by ggml's eval-callback loop: the node we mark is computed and
ggml_backend_synchronize'd before the non-ask callback fires, and the following compute
(the expert matmul) runs only after our load returns. The next layer cannot overwrite the
buffers until this layer's matmul has synchronized. Correct on any backend.
The result is lossless: byte-identical to running with every expert resident, asserted
by the gates. That is the streaming path itself; two opt-in knobs deliberately trade output for
speed on top of it — --n-expert-used (fewer experts per token) and
--drop-cold-experts (skip an expert that would cost a read and was barely
weighted). Both are off unless asked for, which is what keeps the sentence above true by default.
Residency modes
- Cache off (shared slots). Three heap buffers (full
n_expertsize) are shared across layers — one layer computes at a time. Routed slices are re-read fresh every token. Lowest RAM, highest I/O. - LRU cache (
--cache-mb N). Each(layer, projection)gets a reserved, lazily-committed address range. A routed expert already resident is a hit (no read); a miss is read once and kept; over budget, the coldest(layer, expert)is evicted and its pages physically released (madvise(MADV_DONTNEED)/MEM_DECOMMIT). RAM is bounded for real.
The cache rule: 0 or ≥ ~2 GB
Expert reuse is broad, not skewed: hit rate rises roughly linearly with budget, with no
small-cache plateau. A budget below one token's routed working set (~1 GB for
Qwen3-30B-A3B) yields zero hits and pays eviction overhead — measurably slower than no
cache. So validate() rejects a budget in the 1..1499 MiB band unless you force it. Use
0, or ≥ 2000.
Parallel reads (--io-threads N)
Routed slices are read across N lanes, each with a private fd and bounce buffer; the
calling thread participates as lane 0. On UFS 4.x, 4 lanes roughly triples effective read
bandwidth over serial. Compute threads (-t) show a U-shape — 4 is the measured optimum;
8 regresses badly because ggml's spin-wait contends with the synchronous reads.
Why repack must stay off
The streamer rebinds tensor->data to a buffer it fills from the file's native byte
layout. use_extra_bufts=true would repack Q4_K weights into a different in-memory layout
(e.g. q4_K_8x8), so the file offsets would no longer describe what the matmul reads.
The engine loads with use_extra_bufts=false; this is load-bearing, not a tuning knob.
Assumptions to re-check on a submodule bump
- The routing node is named
ffn_moe_topk-<il>and the expert tensorsblk.<il>.ffn_{gate,up,down}_exps.weight. The recipe isolates these names. - The eval-callback fires per decode (not skipped by graph reuse) and computes a marked node alone before the non-ask callback. The gates catch a regression here.