docs: correct content overtaken by the code

roadmap.md listed three shipped capabilities as future work: overlapping I/O with
compute (--overlap), prefetching the next layer's experts (--prefetch), and the
routing predictor behind them -- which was in fact built and then removed. It also
opened on 'decode is ~79% flash I/O', true only with the cache off; with a sized
cache and overlap decode is compute-bound, which changes what is worth doing next.
Rewrite it against the measurements, and say plainly why speculative gating is not
coming back.

benchmarks.md described Gemma as 'A4B (4 experts active)' at the top and corrected
that same claim 95 lines further down; keep the correction. CONTRIBUTING.md offered
the merged ffn_gate_up_exps layout as a good first contribution -- it shipped as
gemma4. Flag the archived spec-gate findings as describing a removed flag.
This commit is contained in:
Helldez 2026-07-15 07:50:47 +02:00
parent a5a86e58df
commit 44c2c2db73
4 changed files with 48 additions and 23 deletions

View file

@ -23,8 +23,7 @@ cd build && ctest --output-on-failure
## Good first contributions
- A new MoE architecture recipe + its gate.
- Support for the merged `ffn_gate_up_exps` expert layout.
- A new MoE architecture recipe + its gate (`docs/adding-a-model.md`).
- Read coalescing / expert-contiguous layout experiments (see `docs/roadmap.md`).
## PRs

View file

@ -1,5 +1,10 @@
# On-device A/B for temporal prefetch (PR2) and speculative gating (PR3)
> **Archived, 2026-07-12.** Speculative gating has since been removed from the engine to restore
> the modular seam; `--spec-gate` is no longer a flag. The `spec-gate` rows and the advice about
> tuning it below are kept as a record of why that decision was made — they are not actionable.
> See [../README.md](../README.md).
OnePlus 15R (11.3 GB), 256-token decode, prompt-processing + steady-state. Each config on top of
its model's best measured baseline (Qwen: cache 4000, lane 4, overlap; Gemma: cache 2000, lane 4,
overlap). See `summary.md` for the full tables. Device was cooler than the 2026-07-12 baseline run,

View file

@ -132,7 +132,9 @@ there is too much I/O — 1.6 s/tok of lane-busy reads — to hide behind ~0.4 s
### Gemma-4-26B-A4B-it-Q4_K_M
- **File:** `Gemma-4-26B-A4B-it-Q4_K_M.gguf` — 17.0 GB on disk, Q4_K_M quantization.
- **Shape:** fused gate+up expert layout, A4B (4 experts active). ≈1.51× device RAM (11.3 GB).
- **Shape:** fused gate+up expert layout. The "A4B" label is ~4 **billion active parameters**, not
4 active experts — the measured I/O ratio puts the default routing width at 8 (see the
active-expert override section below). ≈1.51× device RAM (11.3 GB).
| Config | mean | min | max | median | p5 | p95 | cache hit | flash read/token | decode: compute + I/O (s/tok) |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|

View file

@ -1,37 +1,56 @@
# Roadmap
Themes, not deadlines. Ordered roughly by expected impact on the flash-I/O-bound decode.
Themes, not deadlines. Ordered roughly by expected impact.
## Read bandwidth
The starting point has moved: with a well-sized expert cache and `--overlap`, decode on the
reference device is **compute-bound, not I/O-bound**. Flash I/O is ~79 % of decode only with the
cache off; at Qwen's cache 4000 it inverts, and an infinite cache would still cap at 1/compute
≈ 6.2 tok/s — this SoC's in-RAM decode speed ([benchmarks.md](benchmarks.md#reading-the-numbers)).
So the streaming path has already recovered most of what streaming can recover, and the themes
below are ordered by that fact: read bandwidth still matters for the models that do not fit a
useful cache, but throughput on the ones that do now depends on compute.
Decode is ~79% flash I/O, and effective O_DIRECT bandwidth is well below the drive's
sequential ceiling because routed expert slices are scattered. The largest untapped lever:
## Read bandwidth, where it still binds
- **Expert-contiguous gguf layout.** An offline repack that stores each layer's experts
contiguously (and/or groups the three projections per expert) so a routed set becomes
one coalesced read instead of many scattered ones.
Effective O_DIRECT bandwidth is well below the drive's sequential ceiling because routed expert
slices are scattered. This still dominates the models too large for a useful cache (gpt-oss-120b
at 5.2× RAM), and the cold-cache warm-up on every model.
- **Expert-contiguous gguf layout.** An offline repack storing each layer's experts contiguously
(and/or grouping the projections per expert) so a routed set becomes one coalesced read instead
of many scattered ones.
- **Read coalescing** of adjacent routed slices at runtime.
## Overlap I/O with compute
## Warm-up
Today the read phase and the expert matmul are serial. Prefetching the next layer's likely
experts (or overlapping the read of layer N+1 with the compute of layer N) could hide a
large share of the I/O. Requires a routing predictor or speculative prefetch.
Dense weights are warmed into the page cache at load, which removes the >RAM fault storm. The
expert cache still fills from cold, so the first tokens pay for it and no warm-up flag can change
that ([warmup-analysis.md](warmup-analysis.md)). Worth exploring: preloading experts by routing
frequency rather than by arrival, and a cross-run persistent cache so a second session starts warm.
## More architectures
Beyond `qwen3moe`: other `build_moe_ffn` models are one recipe row each. The merged
`ffn_gate_up_exps` layout and shared-expert models are supported too (`gemma4` / Gemma 4
MoE exercises both). Remaining frontier: architectures whose routing node is not the shared
`ffn_moe_topk` (e.g. custom gating), which the capture/stream hook would need to learn. See
[adding-a-model.md](adding-a-model.md).
`qwen3moe`, `gemma4` (merged `ffn_gate_up_exps` plus shared experts) and OpenAI `gpt-oss` (MXFP4,
purely routed) are supported; other `build_moe_ffn` models are one recipe row each. The remaining
frontier is architectures whose routing node is not the shared `ffn_moe_topk` — custom gating,
which the capture/stream hook would need to learn. See [adding-a-model.md](adding-a-model.md).
## Expert quantization on the fly
Storing streamed experts at a lower precision than the resident parts to cut read volume,
if it can stay within an acceptable quality/lossless boundary.
Storing streamed experts at a lower precision than the resident parts to cut read volume, if it
can stay within an acceptable quality boundary. Most valuable exactly where read bandwidth still
binds, above.
## Bigger, smarter cache
The cache is capacity-bound, not policy-bound (reuse is broad, not skewed), so the simplest
win is more budget. A cross-run persistent cache and admission policies are worth exploring.
The cache is capacity-bound, not policy-bound (reuse is broad, not skewed), so the simplest win is
more budget — which `--cache-mb auto` now takes automatically, capped by `--cache-ceil-mb`
([adaptive-cache.md](adaptive-cache.md)). Admission policies and a persistent cross-run cache
remain unexplored.
## Not on this list
Routing prediction and speculative expert gating were built and **removed**: the recall/latency
trade never paid on-device, and the predictor coupled the streamer to model internals, which cost
more in modularity than it returned in throughput. The archived measurements are in
[bench-data/2026-07-12-pr23/](bench-data/2026-07-12-pr23/).