mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
docs: correct content overtaken by the code
roadmap.md listed three shipped capabilities as future work: overlapping I/O with compute (--overlap), prefetching the next layer's experts (--prefetch), and the routing predictor behind them -- which was in fact built and then removed. It also opened on 'decode is ~79% flash I/O', true only with the cache off; with a sized cache and overlap decode is compute-bound, which changes what is worth doing next. Rewrite it against the measurements, and say plainly why speculative gating is not coming back. benchmarks.md described Gemma as 'A4B (4 experts active)' at the top and corrected that same claim 95 lines further down; keep the correction. CONTRIBUTING.md offered the merged ffn_gate_up_exps layout as a good first contribution -- it shipped as gemma4. Flag the archived spec-gate findings as describing a removed flag.
This commit is contained in:
parent
a5a86e58df
commit
44c2c2db73
4 changed files with 48 additions and 23 deletions
|
|
@ -23,8 +23,7 @@ cd build && ctest --output-on-failure
|
|||
|
||||
## Good first contributions
|
||||
|
||||
- A new MoE architecture recipe + its gate.
|
||||
- Support for the merged `ffn_gate_up_exps` expert layout.
|
||||
- A new MoE architecture recipe + its gate (`docs/adding-a-model.md`).
|
||||
- Read coalescing / expert-contiguous layout experiments (see `docs/roadmap.md`).
|
||||
|
||||
## PRs
|
||||
|
|
|
|||
|
|
@ -1,5 +1,10 @@
|
|||
# On-device A/B for temporal prefetch (PR2) and speculative gating (PR3)
|
||||
|
||||
> **Archived, 2026-07-12.** Speculative gating has since been removed from the engine to restore
|
||||
> the modular seam; `--spec-gate` is no longer a flag. The `spec-gate` rows and the advice about
|
||||
> tuning it below are kept as a record of why that decision was made — they are not actionable.
|
||||
> See [../README.md](../README.md).
|
||||
|
||||
OnePlus 15R (11.3 GB), 256-token decode, prompt-processing + steady-state. Each config on top of
|
||||
its model's best measured baseline (Qwen: cache 4000, lane 4, overlap; Gemma: cache 2000, lane 4,
|
||||
overlap). See `summary.md` for the full tables. Device was cooler than the 2026-07-12 baseline run,
|
||||
|
|
|
|||
|
|
@ -132,7 +132,9 @@ there is too much I/O — 1.6 s/tok of lane-busy reads — to hide behind ~0.4 s
|
|||
### Gemma-4-26B-A4B-it-Q4_K_M
|
||||
|
||||
- **File:** `Gemma-4-26B-A4B-it-Q4_K_M.gguf` — 17.0 GB on disk, Q4_K_M quantization.
|
||||
- **Shape:** fused gate+up expert layout, A4B (4 experts active). ≈1.51× device RAM (11.3 GB).
|
||||
- **Shape:** fused gate+up expert layout. The "A4B" label is ~4 **billion active parameters**, not
|
||||
4 active experts — the measured I/O ratio puts the default routing width at 8 (see the
|
||||
active-expert override section below). ≈1.51× device RAM (11.3 GB).
|
||||
|
||||
| Config | mean | min | max | median | p5 | p95 | cache hit | flash read/token | decode: compute + I/O (s/tok) |
|
||||
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
||||
|
|
|
|||
|
|
@ -1,37 +1,56 @@
|
|||
# Roadmap
|
||||
|
||||
Themes, not deadlines. Ordered roughly by expected impact on the flash-I/O-bound decode.
|
||||
Themes, not deadlines. Ordered roughly by expected impact.
|
||||
|
||||
## Read bandwidth
|
||||
The starting point has moved: with a well-sized expert cache and `--overlap`, decode on the
|
||||
reference device is **compute-bound, not I/O-bound**. Flash I/O is ~79 % of decode only with the
|
||||
cache off; at Qwen's cache 4000 it inverts, and an infinite cache would still cap at 1/compute
|
||||
≈ 6.2 tok/s — this SoC's in-RAM decode speed ([benchmarks.md](benchmarks.md#reading-the-numbers)).
|
||||
So the streaming path has already recovered most of what streaming can recover, and the themes
|
||||
below are ordered by that fact: read bandwidth still matters for the models that do not fit a
|
||||
useful cache, but throughput on the ones that do now depends on compute.
|
||||
|
||||
Decode is ~79% flash I/O, and effective O_DIRECT bandwidth is well below the drive's
|
||||
sequential ceiling because routed expert slices are scattered. The largest untapped lever:
|
||||
## Read bandwidth, where it still binds
|
||||
|
||||
- **Expert-contiguous gguf layout.** An offline repack that stores each layer's experts
|
||||
contiguously (and/or groups the three projections per expert) so a routed set becomes
|
||||
one coalesced read instead of many scattered ones.
|
||||
Effective O_DIRECT bandwidth is well below the drive's sequential ceiling because routed expert
|
||||
slices are scattered. This still dominates the models too large for a useful cache (gpt-oss-120b
|
||||
at 5.2× RAM), and the cold-cache warm-up on every model.
|
||||
|
||||
- **Expert-contiguous gguf layout.** An offline repack storing each layer's experts contiguously
|
||||
(and/or grouping the projections per expert) so a routed set becomes one coalesced read instead
|
||||
of many scattered ones.
|
||||
- **Read coalescing** of adjacent routed slices at runtime.
|
||||
|
||||
## Overlap I/O with compute
|
||||
## Warm-up
|
||||
|
||||
Today the read phase and the expert matmul are serial. Prefetching the next layer's likely
|
||||
experts (or overlapping the read of layer N+1 with the compute of layer N) could hide a
|
||||
large share of the I/O. Requires a routing predictor or speculative prefetch.
|
||||
Dense weights are warmed into the page cache at load, which removes the >RAM fault storm. The
|
||||
expert cache still fills from cold, so the first tokens pay for it and no warm-up flag can change
|
||||
that ([warmup-analysis.md](warmup-analysis.md)). Worth exploring: preloading experts by routing
|
||||
frequency rather than by arrival, and a cross-run persistent cache so a second session starts warm.
|
||||
|
||||
## More architectures
|
||||
|
||||
Beyond `qwen3moe`: other `build_moe_ffn` models are one recipe row each. The merged
|
||||
`ffn_gate_up_exps` layout and shared-expert models are supported too (`gemma4` / Gemma 4
|
||||
MoE exercises both). Remaining frontier: architectures whose routing node is not the shared
|
||||
`ffn_moe_topk` (e.g. custom gating), which the capture/stream hook would need to learn. See
|
||||
[adding-a-model.md](adding-a-model.md).
|
||||
`qwen3moe`, `gemma4` (merged `ffn_gate_up_exps` plus shared experts) and OpenAI `gpt-oss` (MXFP4,
|
||||
purely routed) are supported; other `build_moe_ffn` models are one recipe row each. The remaining
|
||||
frontier is architectures whose routing node is not the shared `ffn_moe_topk` — custom gating,
|
||||
which the capture/stream hook would need to learn. See [adding-a-model.md](adding-a-model.md).
|
||||
|
||||
## Expert quantization on the fly
|
||||
|
||||
Storing streamed experts at a lower precision than the resident parts to cut read volume,
|
||||
if it can stay within an acceptable quality/lossless boundary.
|
||||
Storing streamed experts at a lower precision than the resident parts to cut read volume, if it
|
||||
can stay within an acceptable quality boundary. Most valuable exactly where read bandwidth still
|
||||
binds, above.
|
||||
|
||||
## Bigger, smarter cache
|
||||
|
||||
The cache is capacity-bound, not policy-bound (reuse is broad, not skewed), so the simplest
|
||||
win is more budget. A cross-run persistent cache and admission policies are worth exploring.
|
||||
The cache is capacity-bound, not policy-bound (reuse is broad, not skewed), so the simplest win is
|
||||
more budget — which `--cache-mb auto` now takes automatically, capped by `--cache-ceil-mb`
|
||||
([adaptive-cache.md](adaptive-cache.md)). Admission policies and a persistent cross-run cache
|
||||
remain unexplored.
|
||||
|
||||
## Not on this list
|
||||
|
||||
Routing prediction and speculative expert gating were built and **removed**: the recall/latency
|
||||
trade never paid on-device, and the predictor coupled the streamer to model internals, which cost
|
||||
more in modularity than it returned in throughput. The archived measurements are in
|
||||
[bench-data/2026-07-12-pr23/](bench-data/2026-07-12-pr23/).
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue