BigMoeOnEdge/docs
Helldez bb318cbd2e docs: split the benchmarks, and say what the engine grew
benchmarks.md was two documents in one file: 460 lines, two H1s, and four
section names appearing twice. roadmap.md's link to #reading-the-numbers
resolved to whichever came first, which happened to be the intended one — a
coincidence, not a design. Split the gpt-oss-120b half into its own doc and
cross-link the two.

That also fixes the anchor at the old benchmarks.md:30: GitHub slugs an em-dash
heading to a double hyphen, so #device-pressure-not-just-tokens never jumped
anywhere. The README already had this right for its own gpt-oss link, so the
convention was there — this one was just wrong.

README: route traces and the app's Markdown answers have been in main for
several commits with no mention, and a feature nobody can find is a feature
nobody has. Trimmed the O_DIRECT and cache bullets in exchange: both re-taught
mechanism the linked docs already own, which is what a landing page delegates.
2026-07-15 20:43:41 +02:00
..
bench-data docs: point the telemetry contract and the data index at the route-trace session 2026-07-15 09:53:18 +02:00
adaptive-cache.md docs: what Android reclaim actually does to a >RAM engine 2026-07-15 20:10:08 +02:00
adding-a-model.md refactor(moe): drop the llada-moe recipe (diffusion, out of scope) 2026-07-13 16:02:32 +02:00
android-memory.md docs: what Android reclaim actually does to a >RAM engine 2026-07-15 20:10:08 +02:00
architecture.md docs: session mode protocol, architecture, and CHANGELOG 2026-07-12 19:48:49 +02:00
benchmark-method.md docs(benchmark): gpt-oss-120b on-device streaming results + drivers 2026-07-14 12:29:16 +02:00
benchmarks-gpt-oss.md docs: split the benchmarks, and say what the engine grew 2026-07-15 20:43:41 +02:00
benchmarks.md docs: split the benchmarks, and say what the engine grew 2026-07-15 20:43:41 +02:00
limitations.md refactor(moe): drop the llada-moe recipe (diffusion, out of scope) 2026-07-13 16:02:32 +02:00
moe-streaming.md docs: architecture, seam, MoE streaming, benchmarks, and agent guide 2026-07-10 18:18:24 +02:00
prefetch.md docs+android: temporal prefetch guide, telemetry, and settings row 2026-07-12 20:08:14 +02:00
README.md docs: split the benchmarks, and say what the engine grew 2026-07-15 20:43:41 +02:00
roadmap.md docs: correct content overtaken by the code 2026-07-15 07:50:47 +02:00
seam.md docs: document the expert-ready fork hook and overlap telemetry 2026-07-12 08:54:19 +02:00
session.md docs: multi-turn chat, per-turn metrics, and the new telemetry fields 2026-07-13 12:15:52 +02:00
telemetry.md feat(metrics): --compute-trace and --io-trace decompose the decode 2026-07-15 20:38:56 +02:00
warmup-analysis.md docs(telemetry): document the compute decomposition and stall floor 2026-07-15 07:03:04 +02:00

Documentation

Start with architecture.md for the layer map, or moe-streaming.md for the idea the project is built on.

Understanding the design

Doc What it answers
architecture.md How the layers fit together, and why llama.cpp is not forked.
moe-streaming.md Why streaming experts from flash makes a >RAM model run at all.
seam.md The exact contract with llama.cpp's public API, and how to upgrade the submodule.
limitations.md What this does not do, what it cannot do, and the prior art it builds on.
roadmap.md Themes worth exploring next.

Using and extending it

Doc What it answers
adding-a-model.md How to support a new MoE architecture (a recipe row plus a gate).
telemetry.md The BMOE_* line protocol and CSV schema — the integration contract.
session.md Session lifecycle, KV prefix reuse, cancellation.
adaptive-cache.md --cache-mb auto, the cache ceiling, and dense warm-up.
prefetch.md --prefetch K: the design and why it cannot change output.
android-memory.md What reclaims the engine's memory on a phone, which levers exist (almost none), and why the cache hit rate is what the kernel judges you by.

Measurements

Doc What it answers
benchmarks.md Measured results per model on Android, with device-pressure numbers.
benchmarks-gpt-oss.md gpt-oss-120b: a 58 GB model at 5.2× device RAM, and what it costs.
benchmark-method.md How the numbers are produced, so you can reproduce them.
warmup-analysis.md Why first tokens are slow, and the two regimes behind it.
bench-data/ Raw per-run CSVs and session notes. A dated archive — see its README.

Every benchmark figure in these docs is measured on the hardware named beside it. If you re-measure, update benchmark-method.md and the affected table together.