BigMoeOnEdge/docs
Raffaele 47924565c1
feat(io): release the model file's mapping after load (--release-mmap) (#185)
llama.cpp maps every gguf it loads and keeps the mapping for the model's lifetime. On
Windows that is expensive in a way nothing had attributed: while a section of a file is
alive, NTFS serialises concurrent unbuffered reads on that file, and a lane opened while
the section existed keeps serialising against it after the section is gone. Four I/O lanes
therefore delivered exactly one lane's throughput, which is why lanes and threads have
always measured dead on the desktop host and why the engine read at about a third of what
the drive can serve.

--release-mmap hands the mapping back after load: unmap the file, close its section, reopen
the reader lanes. Both halves are needed. Whether it is safe is decided by looking rather
than by reasoning, the engine asks the OS whether any weight the capture pass observed still
points inside a mapping of the model files, and declines if any does. Off by default,
because the check answers for the pointers the capture saw and for no others.

Host A/B on Qwen3.6-35B-A3B Q4_K_M: 3.16 to 4.63 tok/s (+46%), flash stall per token 0.182
to 0.074, with bytes read, hit rate, evictions and re-reads identical to the digit and the
generated text byte-identical. On the phone the read path is flat (f2fs does not serialise)
but CPU per token falls about 9%; that cell is two short runs per variant and is recorded as
a direction, not a number.

Also fixes a bug this uncovered, independent of the flag: when a gguf carries no
output.weight, llama.cpp builds the output head from the token embedding table and the model
holds two identically named tensors over the same bytes. The capture pass keyed its map by
name, so --dense-weights anon and ahwb rebound one and left the twin reading the mmap for
the whole run, which on a model past RAM means the output projection served by page faults
from flash. The capture now records every distinct leaf object by address and the dense
policy rebinds every tensor over one file range onto the same buffer.

Adds bmoe-iobench --mmap / --reopen-lanes / --range-mb / --fresh, the cells that isolate the
mechanism, a mapping_release unit test on both platforms, the app switch "Release the model
mapping", and the bench findings. README, architecture, AGENTS and roadmap updated, the last
correcting a diagnosis this refutes.
2026-09-07 20:43:02 +02:00
..
assets docs(assets): label the Qwen3.8-Flash-Next hero clip like the DeepSeek one 2026-08-30 17:35:24 +02:00
bench-data feat(io): release the model file's mapping after load (--release-mmap) (#185) 2026-09-07 20:43:02 +02:00
adding-a-model.md refactor(moe): drop the llada-moe recipe (diffusion, out of scope) 2026-07-13 16:02:32 +02:00
android-memory.md feat(moe): hold back dense tensors larger than RAM under every mode (#174) 2026-08-28 10:20:13 +02:00
architecture.md feat(io): release the model file's mapping after load (--release-mmap) (#185) 2026-09-07 20:43:02 +02:00
benchmark-method.md docs: correct the macOS platform status after F_NOCACHE and the CI job 2026-08-29 19:11:31 +02:00
benchmarks-gpt-oss.md docs: record the expert-sidecar negative result and ship the --scatter microbench 2026-07-20 17:02:48 +02:00
benchmarks.md docs: benchmark pages as a contributor guide, and the Apple platform limits 2026-08-28 16:53:22 +02:00
cache-aware-substitution.md feat(moe): cache-aware expert substitution (--expert-substitute, experimental) and --ppl (#171) 2026-08-29 10:13:47 +02:00
cache-sizing.md feat(cli): say which mode a run used, and give --moe-stream a cache by default (#187) 2026-08-29 20:42:17 +02:00
community-benchmarks.md fix(scripts): bench-report model size and storage probe (du on the raw path) (#194) 2026-08-31 12:48:09 +02:00
expert-dropping.md feat(moe): cache-aware expert substitution (--expert-substitute, experimental) and --ppl (#171) 2026-08-29 10:13:47 +02:00
expert-prediction.md feat(moe): measure how predictable expert routing is, then act on it — and refute it (#101) 2026-07-28 09:21:41 +02:00
limitations.md feat(io): uncached reads on macOS via F_NOCACHE, and o_direct from the open's outcome (#182) 2026-08-29 16:05:11 +02:00
moe-streaming.md feat(io): release the model file's mapping after load (--release-mmap) (#185) 2026-09-07 20:43:02 +02:00
mtp.md feat(engine): self-speculative decoding — the model's own MTP head, or n-gram lookup (#134) 2026-08-02 00:09:39 +02:00
ngram.md feat(engine): self-speculative decoding — the model's own MTP head, or n-gram lookup (#134) 2026-08-02 00:09:39 +02:00
prefetch.md feat(moe): measure how predictable expert routing is, then act on it — and refute it (#101) 2026-07-28 09:21:41 +02:00
pressure.md feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95) 2026-07-22 17:21:55 +02:00
README.md feat(moe): cache-aware expert substitution (--expert-substitute, experimental) and --ppl (#171) 2026-08-29 10:13:47 +02:00
roadmap.md feat(io): release the model file's mapping after load (--release-mmap) (#185) 2026-09-07 20:43:02 +02:00
route-ahead.md feat(engine): --route-ahead N — commit decode routing to the N-layers-early prediction (#142) 2026-08-02 00:26:38 +02:00
row-gathered-tables.md feat(moe): serve row-gathered dense tables from flash (--row-stream) (#180) 2026-08-29 10:04:02 +02:00
seam.md feat(moe): cache-aware expert substitution (--expert-substitute, experimental) and --ppl (#171) 2026-08-29 10:13:47 +02:00
session.md perf(cli): BMOE_PROGRESS carries the answer as a delta, not cumulatively (#127) 2026-07-28 15:38:34 +02:00
telemetry.md feat(cli): say which mode a run used, and give --moe-stream a cache by default (#187) 2026-08-29 20:42:17 +02:00
warmup-analysis.md docs: correct the benchmark recipe, the arch list and mismatched sizes 2026-07-19 11:27:18 +02:00

Documentation

Start with architecture.md for the layer map, or moe-streaming.md for the idea the project is built on.

Understanding the design

Doc What it answers
architecture.md How the layers fit together, and why llama.cpp is not forked.
moe-streaming.md Why streaming experts from flash makes a >RAM model run at all.
seam.md The exact contract with llama.cpp's public API, and how to upgrade the submodule.
limitations.md What this does not do, what it cannot do, and the prior art it builds on.
roadmap.md Themes worth exploring next.

Using and extending it

Doc What it answers
adding-a-model.md How to support a new MoE architecture (a recipe row plus a gate).
telemetry.md The BMOE_* line protocol and CSV schema — the integration contract.
session.md Session lifecycle, KV prefix reuse, cancellation.
cache-sizing.md --cache-mb auto, the cache ceiling, and dense warm-up.
prefetch.md --prefetch K: the design and why it cannot change output (with the lossy knobs off).
expert-dropping.md --drop-cold-experts F: spending quality only where it buys a flash read, and why a cache-dependent setting's output is not reproducible.
row-gathered-tables.md --row-stream: serving a dense table the graph only gathers rows from out of flash, and how the engine decides which tables those are without naming one.
cache-aware-substitution.md --expert-substitute L: re-ranking a routing toward the experts already in the cache, the paper it comes from, and why a wide scoring batch cannot price it.
mtp.md --mtp: drafting with the model's own MTP head and verifying a whole group per decode — lossless by construction, and why a wider decode can lose in the streamed regime.
ngram.md --ngram: drafting from text that repeats, with no head and no draft decode — and why a source that abstains costs exactly an unspeculated step.
expert-prediction.md --predict-log: how much of a routing can be known a layer early, measured against the predictor --prefetch already bets on — and why a good score still would not mean a faster decode.
route-ahead.md --route-ahead N: committing the routing to the N-layers-early prediction, so a prefetch of it can never miss — and what that costs in quality (experimental, lossy).
android-memory.md What reclaims the engine's memory on a phone, which levers exist (almost none), and why the cache hit rate is what the kernel judges you by.
pressure.md Cache policy under memory pressure: why an unaffordable budget starts a reclaim war, why the adaptive governor was retired, and what the fixed --cache-mb / --dense-weights levers do.

Measurements

Doc What it answers
benchmarks.md Measured results per model on Android, with device-pressure numbers.
benchmarks-gpt-oss.md gpt-oss-120b: a 58 GB model at 5.2× device RAM, and what it costs.
benchmark-method.md How to measure this engine on any machine: what each knob does, when to move it, and the rules that keep a matrix honest.
community-benchmarks.md Results on hardware we do not own, the one-command protocol (scripts/bench-report.sh), and how to submit a row.
warmup-analysis.md Why first tokens are slow, and the two regimes behind it.
bench-data/ Raw per-run CSVs and session notes. A dated archive — see its README.

Every benchmark figure in these docs is measured on the hardware named beside it. If you re-measure, update benchmark-method.md and the affected table together.