BigMoeOnEdge/docs/README.md
Helldez 5f19289288
docs: DeepSeek V4 Flash leads the README, and the gate list stops lying (#150)
The flagship demo is now the 284B model generating on a 12 GB phone at about
1 tok/s, with its own recording, instead of gpt-oss carrying that slot. The
quoted 0.94 tok/s is the app's own reading and the sentence next to it says
what produced it: cold-expert dropping at full strength, which trades quality.
gpt-oss keeps its lossless and knob-on figures one paragraph down, and the
three-model clip stays where it was.

The model is named by its exact release, 0731, everywhere it appears rather
than only in the opening line. Anyone reproducing this needs to know which
DeepSeek V4 Flash it was.

The list of what ctest gates claimed the LRU cache, evictions, overlap,
temporal prefetch, the dense rebind and multi-turn sessions. It omitted
predictive prefetch and the split multi-shard model, both of which are gated,
and it did not mention that a lossy knob only gets machinery gates. Fixed to
match what the suite actually runs today.

Also: docs/ngram.md was missing from the documentation index despite being a
250-line document for a shipped flag, and the methodology caveat carried an
editorialising aside about future storage in the one paragraph that has to
stay dry.
2026-08-02 01:15:33 +02:00

3.4 KiB
Raw Blame History

Documentation

Start with architecture.md for the layer map, or moe-streaming.md for the idea the project is built on.

Understanding the design

Doc What it answers
architecture.md How the layers fit together, and why llama.cpp is not forked.
moe-streaming.md Why streaming experts from flash makes a >RAM model run at all.
seam.md The exact contract with llama.cpp's public API, and how to upgrade the submodule.
limitations.md What this does not do, what it cannot do, and the prior art it builds on.
roadmap.md Themes worth exploring next.

Using and extending it

Doc What it answers
adding-a-model.md How to support a new MoE architecture (a recipe row plus a gate).
telemetry.md The BMOE_* line protocol and CSV schema — the integration contract.
session.md Session lifecycle, KV prefix reuse, cancellation.
cache-sizing.md --cache-mb auto, the cache ceiling, and dense warm-up.
prefetch.md --prefetch K: the design and why it cannot change output (with the lossy knobs off).
expert-dropping.md --drop-cold-experts F: spending quality only where it buys a flash read, and why it is the one setting whose output is not reproducible.
mtp.md --mtp: drafting with the model's own MTP head and verifying a whole group per decode — lossless by construction, and why a wider decode can lose in the streamed regime.
ngram.md --ngram: drafting from text that repeats, with no head and no draft decode — and why a source that abstains costs exactly an unspeculated step.
expert-prediction.md --predict-log: how much of a routing can be known a layer early, measured against the predictor --prefetch already bets on — and why a good score still would not mean a faster decode.
route-ahead.md --route-ahead N: committing the routing to the N-layers-early prediction, so a prefetch of it can never miss — and what that costs in quality (experimental, lossy).
android-memory.md What reclaims the engine's memory on a phone, which levers exist (almost none), and why the cache hit rate is what the kernel judges you by.
pressure.md Cache policy under memory pressure: why an unaffordable budget starts a reclaim war, why the adaptive governor was retired, and what the fixed --cache-mb / --dense-weights levers do.

Measurements

Doc What it answers
benchmarks.md Measured results per model on Android, with device-pressure numbers.
benchmarks-gpt-oss.md gpt-oss-120b: a 58 GB model at 5.2× device RAM, and what it costs.
benchmark-method.md How the numbers are produced, so you can reproduce them.
warmup-analysis.md Why first tokens are slow, and the two regimes behind it.
bench-data/ Raw per-run CSVs and session notes. A dated archive — see its README.

Every benchmark figure in these docs is measured on the hardware named beside it. If you re-measure, update benchmark-method.md and the affected table together.