mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-02 19:15:51 +00:00
The NPU prefill's expert arena read every expert of every layer ahead of its routing. It now reads, ahead of a layer's routing, the experts the previous graph routed there, and at the routing node whatever the routing adds. The matmul reads only routed experts, so the output is bit for bit the same. A layer routing more than --prefill-routed-full (0.85) of its experts gets the next one read whole; --no-prefill-routed restores whole layers everywhere. Phone, Hexagon v81 NPU, top-4, same session, every answer identical: Qwen3.6-35B-A3B Q4_0 7.68 -> 4.16 s, Q4_K_M 9.95 -> 5.37 s, Gemma 4 26B-A4B Q4_K_M 6.69 -> 3.70 s, Nemotron 3.5 30B-A3B Q4_0 7.42 -> 6.81 s. Also: --decide-probe (experimental per-decision expert usage and layer-exit answers), BMOE_DECIDE prefill_dev_* counters, gates G17f/G17g, app 0.28.0. |
||
|---|---|---|
| .. | ||
| assets | ||
| bench-data | ||
| adding-a-model.md | ||
| android-memory.md | ||
| architecture.md | ||
| benchmark-method.md | ||
| benchmarks-gpt-oss.md | ||
| benchmarks.md | ||
| cache-aware-substitution.md | ||
| cache-sizing.md | ||
| community-benchmarks.md | ||
| decide.md | ||
| expert-dropping.md | ||
| expert-prediction.md | ||
| limitations.md | ||
| moe-streaming.md | ||
| mtp.md | ||
| ngram.md | ||
| npu-prefill.md | ||
| prefetch.md | ||
| pressure.md | ||
| README.md | ||
| roadmap.md | ||
| route-ahead.md | ||
| row-gathered-tables.md | ||
| seam.md | ||
| session.md | ||
| telemetry.md | ||
| warmup-analysis.md | ||
Documentation
Start with architecture.md for the layer map, or moe-streaming.md for the idea the project is built on.
Understanding the design
| Doc | What it answers |
|---|---|
| architecture.md | How the layers fit together, and why llama.cpp is not forked. |
| moe-streaming.md | Why streaming experts from flash makes a >RAM model run at all. |
| seam.md | The exact contract with llama.cpp's public API, and how to upgrade the submodule. |
| limitations.md | What this does not do, what it cannot do, and the prior art it builds on. |
| roadmap.md | Themes worth exploring next. |
Using and extending it
| Doc | What it answers |
|---|---|
| adding-a-model.md | How to support a new MoE architecture (a recipe row plus a gate). |
| telemetry.md | The BMOE_* line protocol and CSV schema — the integration contract. |
| session.md | Session lifecycle, KV prefix reuse, cancellation. |
| decide.md | --decide: picking one of a list of choices from a single prefill, with no decode, and keeping the state after a shared prefix between calls. |
| cache-sizing.md | --cache-mb auto, the cache ceiling, and dense warm-up. |
| prefetch.md | --prefetch K: the design and why it cannot change output (with the lossy knobs off). |
| expert-dropping.md | --drop-cold-experts F: spending quality only where it buys a flash read, and why a cache-dependent setting's output is not reproducible. |
| row-gathered-tables.md | --row-stream: serving a dense table the graph only gathers rows from out of flash, and how the engine decides which tables those are without naming one. |
| cache-aware-substitution.md | --expert-substitute L: re-ranking a routing toward the experts already in the cache, the paper it comes from, and why a wide scoring batch cannot price it. |
| mtp.md | --mtp: drafting with the model's own MTP head and verifying a whole group per decode — lossless by construction, and why a wider decode can lose in the streamed regime. |
| ngram.md | --ngram: drafting from text that repeats, with no head and no draft decode — and why a source that abstains costs exactly an unspeculated step. |
| expert-prediction.md | --predict-log: how much of a routing can be known a layer early, measured against the predictor --prefetch already bets on — and why a good score still would not mean a faster decode. |
| route-ahead.md | --route-ahead N: committing the routing to the N-layers-early prediction, so a prefetch of it can never miss — and what that costs in quality (experimental, lossy). |
| npu-prefill.md | --prefill-device: prefill on the NPU and decode on the CPU, a model larger than RAM fed to the NPU two layers at a time, and why decode stays on the CPU. |
| android-memory.md | What reclaims the engine's memory on a phone, which levers exist (almost none), and why the cache hit rate is what the kernel judges you by. |
| pressure.md | Cache policy under memory pressure: why an unaffordable budget starts a reclaim war, why the adaptive governor was retired, and what the fixed --cache-mb / --dense-weights levers do. |
Measurements
| Doc | What it answers |
|---|---|
| benchmarks.md | Measured results per model on Android, with device-pressure numbers. |
| benchmarks-gpt-oss.md | gpt-oss-120b: a 58 GB model at 5.2× device RAM, and what it costs. |
| benchmark-method.md | How to measure this engine on any machine: what each knob does, when to move it, and the rules that keep a matrix honest. |
| community-benchmarks.md | Results on hardware we do not own, the one-command protocol (scripts/bench-report.sh), and how to submit a row. |
| warmup-analysis.md | Why first tokens are slow, and the two regimes behind it. |
| bench-data/ | Raw per-run CSVs and session notes. A dated archive — see its README. |
Every benchmark figure in these docs is measured on the hardware named beside it. If you re-measure, update benchmark-method.md and the affected table together.