Find a file
Helldez 4d57da5edb feat(moe): temporal prefetch of the next layers' experts on idle lanes
While a token computes layer l, the previous token's routing at layers l+1..l+K is a strong
predictor of this token's — so speculatively read those experts on the otherwise-idle I/O lanes,
turning the next layer's read into a cache hit. RouterHook records each layer's last-token
routing and calls the new IExpertSource::prefetch() hint; ExpertStreamSource services it with a
low-priority speculative queue that never delays real work: all LRU mutation stays on the eval
thread (prefetch commits + enqueues, quiesce integrates completed reads and discards the rest at
the next real load), while workers only read bytes and yield the moment a real batch arrives. A
speculative slice is the identical read (same file offset, same buffer) a real miss would issue,
so integration is byte-safe by construction. --prefetch K (env BMOE_PREFETCH), cache required;
telemetry reports speculative MiB and useful-hit rate. Gates G5a/b/c pass byte-identically on
qwen3moe and gemma4 — G5c forces synchronous completion to exercise integrate-then-hit
deterministically (observed 3 experts prefetched, 1 useful).
2026-07-12 20:05:28 +02:00
.github/workflows fix(android): stage libllama-common.so into jniLibs 2026-07-11 22:44:43 +02:00
cli feat(moe): temporal prefetch of the next layers' experts on idle lanes 2026-07-12 20:05:28 +02:00
core feat(moe): temporal prefetch of the next layers' experts on idle lanes 2026-07-12 20:05:28 +02:00
docs docs: session mode protocol, architecture, and CHANGELOG 2026-07-12 19:48:49 +02:00
examples/android feat(android): drive one persistent engine session per model 2026-07-12 19:47:22 +02:00
scripts feat(moe): surface cache-management cost as a mgmt telemetry term 2026-07-12 19:28:03 +02:00
tests test: session byte-identity gates (S1/S2) 2026-07-12 19:36:56 +02:00
third_party build: point llama.cpp submodule at Helldez fork with expert-ready hook 2026-07-12 08:54:19 +02:00
.clang-format chore: bootstrap repository 2026-07-10 18:17:31 +02:00
.gitattributes chore: enforce LF line endings via .gitattributes 2026-07-10 18:18:10 +02:00
.gitignore docs: full Qwen+Gemma benchmark matrix with compute/IO decode split 2026-07-11 22:44:43 +02:00
.gitmodules build: point llama.cpp submodule at Helldez fork with expert-ready hook 2026-07-12 08:54:19 +02:00
AGENTS.md docs: architecture, seam, MoE streaming, benchmarks, and agent guide 2026-07-10 18:18:24 +02:00
CHANGELOG.md docs: session mode protocol, architecture, and CHANGELOG 2026-07-12 19:48:49 +02:00
CLAUDE.md docs: architecture, seam, MoE streaming, benchmarks, and agent guide 2026-07-10 18:18:24 +02:00
CMakeLists.txt build: point llama.cpp submodule at Helldez fork with expert-ready hook 2026-07-12 08:54:19 +02:00
CONTRIBUTING.md docs: architecture, seam, MoE streaming, benchmarks, and agent guide 2026-07-10 18:18:24 +02:00
LICENSE chore: bootstrap repository 2026-07-10 18:17:31 +02:00
README.md docs: document overlap, mmap baseline, reasoning control and the fixes 2026-07-12 08:54:19 +02:00

BigMoeOnEdge

Run Mixture-of-Experts language models that are larger than a device's RAM at usable speed, by streaming only the experts each token actually routes to.

A MoE layer holds many experts but a single token uses only its top-k of them. For Qwen3-30B-A3B that is 8 of 128 per layer — about 6% of the expert weights. BigMoeOnEdge keeps the small, dense parts of the model resident and reads just those routed expert slices from flash on demand, so an 18.5 GB model runs on an 11 GB phone, losslessly.

  • Measured: Qwen3-30B-A3B-Q4_K_M on a OnePlus 15R (11.3 GB RAM) → 1.8 tok/s with no cache, up to 3.75 tok/s with a 4 GiB expert cache and 4 read lanes, byte-identical to a full in-memory run. See the table below, or docs/benchmarks.md for the full matrix (Qwen + Gemma, mean/min/max/p95).
  • Public-API streaming seam. Expert streaming is driven entirely through llama.cpp's public eval-callback and public gguf accessors and works against stock upstream. The one exception is the optional intra-layer overlap feature (--overlap), which needs a ~25-line per-expert readiness hook carried as a single-commit fork branch — an explicit tide-me-over that is dropped the moment upstream ships an equivalent callback. Everything else, and the serial streamer in full, is a submodule pointer bump away from upstream. See docs/architecture.md and docs/seam.md § 3.
  • Modular. A ports-and-adapters engine: the streaming strategy, the metrics sink and the target are interfaces. Adding a MoE architecture is one registry row (docs/adding-a-model.md).

Prior art, credited: this is an engineering package of ideas from AirLLM, Apple's "LLM in a flash", FlexGen, PowerInfer and EdgeMoE — not a novel technique. See docs/limitations.md.

Benchmarks

Qwen3-30B-A3B-Q4_K_M (18.5 GB, 128 experts, top-8, 48 layers) on a OnePlus 15R (11.3 GB RAM, UFS 4.x), 4 compute threads, 256-token steady-state runs:

Expert cache I/O lanes tok/s (mean) flash read/token cache hit
mmap only — 1.86 (unstable) — —
off (stream) 4 1.78 1051 MB —
2000 MiB 4 2.47 480 MB 53%
4000 MiB 2 3.38 225 MB 76%
4000 MiB 4 3.75 225 MB 76%

Cache size is the dominant lever (2000 → 4000 MiB nearly doubles throughput as the hit rate climbs); read lanes help mainly when the cache is small. mmap-only looks comparable on average but is unstable (single tokens from 0.36 to 8.34 tok/s) and evicts other apps — streaming with a bounded cache stays responsive. The cache rule is 0 or ≥ ~2 GB: a smaller budget thrashes and is slower than no cache. Gemma-4-26B-A4B-Q4_K_M reaches 3.48 tok/s at the same best setting. Full matrix and method: docs/benchmarks.md, docs/benchmark-method.md.

An optional --overlap mode pipelines each layer's expert reads with its compute (via the fork's per-expert readiness hook) instead of blocking on them, aiming to hide flash I/O behind FFN compute. It is byte-identical to the serial path (gate G4); its on-device throughput is being measured and the table will gain an overlap row once those numbers land.

The same works on desktop for a model larger than the machine's RAM. Qwen3-30B-A3B-Q4_K_M (17.3 GiB) on a Windows PC with 14.8 GiB RAM (1.17× RAM, so it cannot be held resident), cache 4000 MiB, 4 I/O lanes, 4 threads → 2.58 tok/s, 861 MiB/token, 44.8% cache hit, 1.28 GiB/s O_DIRECT, coherent output. Streaming is what makes an over-RAM MoE runnable at all here; if a model fits in RAM, run it resident instead (faster).

Quickstart (host)

git clone --recursive https://github.com/Helldez/BigMoeOnEdge.git
cd BigMoeOnEdge
scripts/build-host.sh

# stream a MoE model, cache 4 GiB, 4 read lanes
build/cli/bmoe-cli -m Qwen3-30B-A3B-Q4_K_M.gguf --moe-stream \
  --cache-mb 4000 --io-threads 4 -t 4 -n 48 --chatml -p "Explain MoE routing."

Add --overlap to pipeline expert reads with compute (needs the fork submodule), and --no-think to render the chat template with reasoning off. Omit --moe-stream entirely for the plain mmap baseline the streaming modes are compared against.

Run the byte-identity gates (needs python3 with the gguf package):

cd build && ctest --output-on-failure

Quickstart (Android)

A minimal chat app with a live telemetry panel is in examples/android. Build the CLI for arm64 with scripts/build-android.ps1, then build the APK and push a model. Its settings expose the streaming knobs (expert cache, I/O lanes, O_DIRECT, I/O–compute overlap), a reasoning toggle, and an mmap baseline switch that turns streaming off entirely so you can compare modes on the same device; the panel reports per-token compute-vs-flash split, cache hit rate and the aggregate tok/s at the end of a run.

How it works, briefly

  1. Load the model file-backed (mmap on, weight repack off).
  2. A one-token warm-up capture reads the expert tensor pointers from the compute graph via the eval-callback, then rebinds them onto streaming buffers.
  3. Each token, the callback sees the routing node, reads the selected expert ids, and the expert source reads exactly those slices from flash (O_DIRECT), with an optional LRU cache and a parallel read pool — just before that layer's expert matmul runs.

Details: docs/moe-streaming.md.

License

Apache-2.0. See LICENSE.