BigMoeOnEdge/docs
Helldez 2a7efae487
feat(moe): hold back dense tensors larger than RAM under every mode (#174)
The dense policy assumed the dense set fits in RAM. Qwen3.8-Flash-Next breaks
that with its 51B n-gram table (per_layer_token_embd, ~28.8 GB at IQ4_NL),
which every mode failed on in its own way. A dense tensor larger than the
kernel's MemAvailable is now held back under every mode: it stays mmap'd
with MADV_RANDOM, leaves the warm sweep, the residency sensor and the auto
cache budget, and one stderr line names it. The bound is size, not access
shape: a row-gathered table that fits keeps its mode. Inert on every other
supported model (Qwen3.6-35B matches its baselines to the decimal under
anon, warm, mmap and auto); gates pass; the guard fires on device and pinned
dense weights survive load.
2026-08-28 10:20:13 +02:00
..
assets feat(moe): Qwen3.8-Flash-Next support (#172) 2026-08-28 10:07:43 +02:00
bench-data docs(bench): record the 2026-07-24 desktop campaign — the bottleneck flips 2026-07-24 10:34:14 +02:00
adding-a-model.md refactor(moe): drop the llada-moe recipe (diffusion, out of scope) 2026-07-13 16:02:32 +02:00
android-memory.md feat(moe): hold back dense tensors larger than RAM under every mode (#174) 2026-08-28 10:20:13 +02:00
architecture.md feat(moe): stream split multi-shard ggufs natively + DeepSeek V4 Flash recipe (#144) 2026-08-01 23:34:35 +02:00
benchmark-method.md feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95) 2026-07-22 17:21:55 +02:00
benchmarks-gpt-oss.md docs: record the expert-sidecar negative result and ship the --scatter microbench 2026-07-20 17:02:48 +02:00
benchmarks.md docs: record the expert-sidecar negative result and ship the --scatter microbench 2026-07-20 17:02:48 +02:00
cache-sizing.md feat(moe): warn when the expert cache budget sits below one token's cycle (#167) 2026-08-25 12:48:22 +02:00
community-benchmarks.md ci: prebuilt bmoe-cli for Linux, macOS and Windows on every release (#178) 2026-08-27 22:27:59 +02:00
expert-dropping.md docs: professional README and canonical AGENTS.md (#132) 2026-07-28 17:40:36 +02:00
expert-prediction.md feat(moe): measure how predictable expert routing is, then act on it — and refute it (#101) 2026-07-28 09:21:41 +02:00
limitations.md feat(moe): Qwen3.8-Flash-Next support (#172) 2026-08-28 10:07:43 +02:00
moe-streaming.md feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95) 2026-07-22 17:21:55 +02:00
mtp.md feat(engine): self-speculative decoding — the model's own MTP head, or n-gram lookup (#134) 2026-08-02 00:09:39 +02:00
ngram.md feat(engine): self-speculative decoding — the model's own MTP head, or n-gram lookup (#134) 2026-08-02 00:09:39 +02:00
prefetch.md feat(moe): measure how predictable expert routing is, then act on it — and refute it (#101) 2026-07-28 09:21:41 +02:00
pressure.md feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95) 2026-07-22 17:21:55 +02:00
README.md feat(bench): community benchmarks: one-command host protocol, issue form, results table (#176) 2026-08-27 21:29:19 +02:00
roadmap.md feat(moe): warn when the expert cache budget sits below one token's cycle (#167) 2026-08-25 12:48:22 +02:00
route-ahead.md feat(engine): --route-ahead N — commit decode routing to the N-layers-early prediction (#142) 2026-08-02 00:26:38 +02:00
seam.md feat(moe): Qwen3.8-Flash-Next support (#172) 2026-08-28 10:07:43 +02:00
session.md perf(cli): BMOE_PROGRESS carries the answer as a delta, not cumulatively (#127) 2026-07-28 15:38:34 +02:00
telemetry.md fix(telemetry): attribute critical-path stall explicitly (#169) 2026-08-27 17:45:51 +02:00
warmup-analysis.md docs: correct the benchmark recipe, the arch list and mismatched sizes 2026-07-19 11:27:18 +02:00

Documentation

Start with architecture.md for the layer map, or moe-streaming.md for the idea the project is built on.

Understanding the design

Doc What it answers
architecture.md How the layers fit together, and why llama.cpp is not forked.
moe-streaming.md Why streaming experts from flash makes a >RAM model run at all.
seam.md The exact contract with llama.cpp's public API, and how to upgrade the submodule.
limitations.md What this does not do, what it cannot do, and the prior art it builds on.
roadmap.md Themes worth exploring next.

Using and extending it

Doc What it answers
adding-a-model.md How to support a new MoE architecture (a recipe row plus a gate).
telemetry.md The BMOE_* line protocol and CSV schema — the integration contract.
session.md Session lifecycle, KV prefix reuse, cancellation.
cache-sizing.md --cache-mb auto, the cache ceiling, and dense warm-up.
prefetch.md --prefetch K: the design and why it cannot change output (with the lossy knobs off).
expert-dropping.md --drop-cold-experts F: spending quality only where it buys a flash read, and why it is the one setting whose output is not reproducible.
mtp.md --mtp: drafting with the model's own MTP head and verifying a whole group per decode — lossless by construction, and why a wider decode can lose in the streamed regime.
ngram.md --ngram: drafting from text that repeats, with no head and no draft decode — and why a source that abstains costs exactly an unspeculated step.
expert-prediction.md --predict-log: how much of a routing can be known a layer early, measured against the predictor --prefetch already bets on — and why a good score still would not mean a faster decode.
route-ahead.md --route-ahead N: committing the routing to the N-layers-early prediction, so a prefetch of it can never miss — and what that costs in quality (experimental, lossy).
android-memory.md What reclaims the engine's memory on a phone, which levers exist (almost none), and why the cache hit rate is what the kernel judges you by.
pressure.md Cache policy under memory pressure: why an unaffordable budget starts a reclaim war, why the adaptive governor was retired, and what the fixed --cache-mb / --dense-weights levers do.

Measurements

Doc What it answers
benchmarks.md Measured results per model on Android, with device-pressure numbers.
benchmarks-gpt-oss.md gpt-oss-120b: a 58 GB model at 5.2× device RAM, and what it costs.
benchmark-method.md How the numbers are produced, so you can reproduce them.
community-benchmarks.md Results on hardware we do not own, the one-command protocol (scripts/bench-report.sh), and how to submit a row.
warmup-analysis.md Why first tokens are slow, and the two regimes behind it.
bench-data/ Raw per-run CSVs and session notes. A dated archive — see its README.

Every benchmark figure in these docs is measured on the hardware named beside it. If you re-measure, update benchmark-method.md and the affected table together.