diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index d5e4f92..d462890 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -23,8 +23,7 @@ cd build && ctest --output-on-failure ## Good first contributions -- A new MoE architecture recipe + its gate. -- Support for the merged `ffn_gate_up_exps` expert layout. +- A new MoE architecture recipe + its gate (`docs/adding-a-model.md`). - Read coalescing / expert-contiguous layout experiments (see `docs/roadmap.md`). ## PRs diff --git a/docs/bench-data/2026-07-12-pr23/findings.md b/docs/bench-data/2026-07-12-pr23/findings.md index 0a34cde..3ab9446 100644 --- a/docs/bench-data/2026-07-12-pr23/findings.md +++ b/docs/bench-data/2026-07-12-pr23/findings.md @@ -1,5 +1,10 @@ # On-device A/B for temporal prefetch (PR2) and speculative gating (PR3) +> **Archived, 2026-07-12.** Speculative gating has since been removed from the engine to restore +> the modular seam; `--spec-gate` is no longer a flag. The `spec-gate` rows and the advice about +> tuning it below are kept as a record of why that decision was made — they are not actionable. +> See [../README.md](../README.md). + OnePlus 15R (11.3 GB), 256-token decode, prompt-processing + steady-state. Each config on top of its model's best measured baseline (Qwen: cache 4000, lane 4, overlap; Gemma: cache 2000, lane 4, overlap). See `summary.md` for the full tables. Device was cooler than the 2026-07-12 baseline run, diff --git a/docs/benchmarks.md b/docs/benchmarks.md index 49be0ad..1ba9939 100644 --- a/docs/benchmarks.md +++ b/docs/benchmarks.md @@ -132,7 +132,9 @@ there is too much I/O — 1.6 s/tok of lane-busy reads — to hide behind ~0.4 s ### Gemma-4-26B-A4B-it-Q4_K_M - **File:** `Gemma-4-26B-A4B-it-Q4_K_M.gguf` — 17.0 GB on disk, Q4_K_M quantization. -- **Shape:** fused gate+up expert layout, A4B (4 experts active). ≈1.51× device RAM (11.3 GB). +- **Shape:** fused gate+up expert layout. The "A4B" label is ~4 **billion active parameters**, not + 4 active experts — the measured I/O ratio puts the default routing width at 8 (see the + active-expert override section below). ≈1.51× device RAM (11.3 GB). | Config | mean | min | max | median | p5 | p95 | cache hit | flash read/token | decode: compute + I/O (s/tok) | |---|---:|---:|---:|---:|---:|---:|---:|---:|---:| diff --git a/docs/roadmap.md b/docs/roadmap.md index 7275a86..b4628be 100644 --- a/docs/roadmap.md +++ b/docs/roadmap.md @@ -1,37 +1,56 @@ # Roadmap -Themes, not deadlines. Ordered roughly by expected impact on the flash-I/O-bound decode. +Themes, not deadlines. Ordered roughly by expected impact. -## Read bandwidth +The starting point has moved: with a well-sized expert cache and `--overlap`, decode on the +reference device is **compute-bound, not I/O-bound**. Flash I/O is ~79 % of decode only with the +cache off; at Qwen's cache 4000 it inverts, and an infinite cache would still cap at 1/compute +≈ 6.2 tok/s — this SoC's in-RAM decode speed ([benchmarks.md](benchmarks.md#reading-the-numbers)). +So the streaming path has already recovered most of what streaming can recover, and the themes +below are ordered by that fact: read bandwidth still matters for the models that do not fit a +useful cache, but throughput on the ones that do now depends on compute. -Decode is ~79% flash I/O, and effective O_DIRECT bandwidth is well below the drive's -sequential ceiling because routed expert slices are scattered. The largest untapped lever: +## Read bandwidth, where it still binds -- **Expert-contiguous gguf layout.** An offline repack that stores each layer's experts - contiguously (and/or groups the three projections per expert) so a routed set becomes - one coalesced read instead of many scattered ones. +Effective O_DIRECT bandwidth is well below the drive's sequential ceiling because routed expert +slices are scattered. This still dominates the models too large for a useful cache (gpt-oss-120b +at 5.2× RAM), and the cold-cache warm-up on every model. + +- **Expert-contiguous gguf layout.** An offline repack storing each layer's experts contiguously + (and/or grouping the projections per expert) so a routed set becomes one coalesced read instead + of many scattered ones. - **Read coalescing** of adjacent routed slices at runtime. -## Overlap I/O with compute +## Warm-up -Today the read phase and the expert matmul are serial. Prefetching the next layer's likely -experts (or overlapping the read of layer N+1 with the compute of layer N) could hide a -large share of the I/O. Requires a routing predictor or speculative prefetch. +Dense weights are warmed into the page cache at load, which removes the >RAM fault storm. The +expert cache still fills from cold, so the first tokens pay for it and no warm-up flag can change +that ([warmup-analysis.md](warmup-analysis.md)). Worth exploring: preloading experts by routing +frequency rather than by arrival, and a cross-run persistent cache so a second session starts warm. ## More architectures -Beyond `qwen3moe`: other `build_moe_ffn` models are one recipe row each. The merged -`ffn_gate_up_exps` layout and shared-expert models are supported too (`gemma4` / Gemma 4 -MoE exercises both). Remaining frontier: architectures whose routing node is not the shared -`ffn_moe_topk` (e.g. custom gating), which the capture/stream hook would need to learn. See -[adding-a-model.md](adding-a-model.md). +`qwen3moe`, `gemma4` (merged `ffn_gate_up_exps` plus shared experts) and OpenAI `gpt-oss` (MXFP4, +purely routed) are supported; other `build_moe_ffn` models are one recipe row each. The remaining +frontier is architectures whose routing node is not the shared `ffn_moe_topk` — custom gating, +which the capture/stream hook would need to learn. See [adding-a-model.md](adding-a-model.md). ## Expert quantization on the fly -Storing streamed experts at a lower precision than the resident parts to cut read volume, -if it can stay within an acceptable quality/lossless boundary. +Storing streamed experts at a lower precision than the resident parts to cut read volume, if it +can stay within an acceptable quality boundary. Most valuable exactly where read bandwidth still +binds, above. ## Bigger, smarter cache -The cache is capacity-bound, not policy-bound (reuse is broad, not skewed), so the simplest -win is more budget. A cross-run persistent cache and admission policies are worth exploring. +The cache is capacity-bound, not policy-bound (reuse is broad, not skewed), so the simplest win is +more budget — which `--cache-mb auto` now takes automatically, capped by `--cache-ceil-mb` +([adaptive-cache.md](adaptive-cache.md)). Admission policies and a persistent cross-run cache +remain unexplored. + +## Not on this list + +Routing prediction and speculative expert gating were built and **removed**: the recall/latency +trade never paid on-device, and the predictor coupled the streamer to model internals, which cost +more in modularity than it returned in throughput. The archived measurements are in +[bench-data/2026-07-12-pr23/](bench-data/2026-07-12-pr23/).