BigMoeOnEdge/docs
Helldez 0091a90a48
feat(app): settings grouped by purpose, and four defects an audit found (#151)
* feat(app): settings grouped by purpose, and four defects an audit found

Settings now show the recommended configuration first and fold everything else
into a collapsed Experimental group per category. That is a statement about
evidence, not about how finished the code is: inside are the levers measured on
one device, measured once, or still owed a measurement. They stay in the
release build, because testing them on other hardware is what this app is for
and a lever nobody can reach is a lever nobody can refute. The caveat is stated
once in the group header instead of leaking into some descriptions and not
others.

Every description was rewritten to say what the setting does for the person
reading it. Out went the measured figures, which need the device, the model and
the day beside them to mean anything and have none of that room under a switch,
and out went the implementation names: O_DIRECT, top-k, dma-buf, mmap and KV
cache are not what someone deciding whether to turn something on needs to know.
The metrics screen keeps the flag names, deliberately: there the reader is
matching the UI against a CSV column and the technical name IS the vocabulary.

Four defects, all found by auditing rather than by anything failing:

The session signature is now derived from the argv instead of being a
hand-written list beside it. Those two had to be kept in step with nothing
enforcing it, and forgetting a field is a silent bug: the setting appears to
change while the engine keeps running the old configuration. Three of four
rebases this week collided on exactly that list.

A malformed end-of-turn summary no longer strands the UI. The whole handler sat
inside a catch with no failure branch, so a parse error left the state in
GENERATING with no turn committed and nothing said. The streamed answer is now
kept, the reason is shown, and the state returns to READY.

MainActivity drops from about 1050 lines to under 700: the model download and
import UI moves to ModelPickerUi.kt, which shares nothing with the chat screen.
No logic moved, only its address.

Dead code removed: a field whose own comment described a use it did not have,
two functions nobody called, and five string resources describing a UI two
rewrites ago.

* docs: record the settings regrouping and the signature fix

Rule 6: the changelog and the docs a change invalidates ship with it. The app README described the settings screen as it was before the regrouping, and explained one experimental lever in terms of predictor accuracy percentages that the UI no longer shows.

* docs: the improved DeepSeek hero recording

* feat(app): keep the flag vocabulary in Settings, and make Experimental read as a boundary

The first pass at rewriting the descriptions went too far: it renamed the controls into consumer phrasing and lost the vocabulary that lets a setting here be matched against the CLI, the CSV preamble and the docs. Labels are back to the flag's own names; the descriptions are shorter than the originals rather than longer, and still carry no measured figures. The Experimental group gets a divider and a tonal bar: collapsed, it is the only thing between the recommended configuration and the levers that can change the reply, so it has to look like a boundary rather than one more row.
2026-08-02 01:50:15 +02:00
..
assets feat(app): settings grouped by purpose, and four defects an audit found (#151) 2026-08-02 01:50:15 +02:00
bench-data docs(bench): record the 2026-07-24 desktop campaign — the bottleneck flips 2026-07-24 10:34:14 +02:00
adding-a-model.md refactor(moe): drop the llada-moe recipe (diffusion, out of scope) 2026-07-13 16:02:32 +02:00
android-memory.md feat(dense): --dense-weights ahwb — dense weights in memory Android cannot reclaim (#93) 2026-07-21 11:09:51 +02:00
architecture.md feat(moe): stream split multi-shard ggufs natively + DeepSeek V4 Flash recipe (#144) 2026-08-01 23:34:35 +02:00
benchmark-method.md feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95) 2026-07-22 17:21:55 +02:00
benchmarks-gpt-oss.md docs: record the expert-sidecar negative result and ship the --scatter microbench 2026-07-20 17:02:48 +02:00
benchmarks.md docs: record the expert-sidecar negative result and ship the --scatter microbench 2026-07-20 17:02:48 +02:00
cache-sizing.md docs: measure the cache and I/O levers, and correct what the measurements contradict (#88) 2026-07-20 11:44:48 +02:00
expert-dropping.md docs: professional README and canonical AGENTS.md (#132) 2026-07-28 17:40:36 +02:00
expert-prediction.md feat(moe): measure how predictable expert routing is, then act on it — and refute it (#101) 2026-07-28 09:21:41 +02:00
limitations.md feat(moe): stream split multi-shard ggufs natively + DeepSeek V4 Flash recipe (#144) 2026-08-01 23:34:35 +02:00
moe-streaming.md feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95) 2026-07-22 17:21:55 +02:00
mtp.md feat(engine): self-speculative decoding — the model's own MTP head, or n-gram lookup (#134) 2026-08-02 00:09:39 +02:00
ngram.md feat(engine): self-speculative decoding — the model's own MTP head, or n-gram lookup (#134) 2026-08-02 00:09:39 +02:00
prefetch.md feat(moe): measure how predictable expert routing is, then act on it — and refute it (#101) 2026-07-28 09:21:41 +02:00
pressure.md feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95) 2026-07-22 17:21:55 +02:00
README.md docs: DeepSeek V4 Flash leads the README, and the gate list stops lying (#150) 2026-08-02 01:15:33 +02:00
roadmap.md docs(roadmap): record why a slice cannot be read straight into its cache slot (#117) 2026-07-28 11:50:09 +02:00
route-ahead.md feat(engine): --route-ahead N — commit decode routing to the N-layers-early prediction (#142) 2026-08-02 00:26:38 +02:00
seam.md feat(engine): self-speculative decoding — the model's own MTP head, or n-gram lookup (#134) 2026-08-02 00:09:39 +02:00
session.md perf(cli): BMOE_PROGRESS carries the answer as a delta, not cumulatively (#127) 2026-07-28 15:38:34 +02:00
telemetry.md feat(engine): --route-ahead N — commit decode routing to the N-layers-early prediction (#142) 2026-08-02 00:26:38 +02:00
warmup-analysis.md docs: correct the benchmark recipe, the arch list and mismatched sizes 2026-07-19 11:27:18 +02:00

Documentation

Start with architecture.md for the layer map, or moe-streaming.md for the idea the project is built on.

Understanding the design

Doc What it answers
architecture.md How the layers fit together, and why llama.cpp is not forked.
moe-streaming.md Why streaming experts from flash makes a >RAM model run at all.
seam.md The exact contract with llama.cpp's public API, and how to upgrade the submodule.
limitations.md What this does not do, what it cannot do, and the prior art it builds on.
roadmap.md Themes worth exploring next.

Using and extending it

Doc What it answers
adding-a-model.md How to support a new MoE architecture (a recipe row plus a gate).
telemetry.md The BMOE_* line protocol and CSV schema — the integration contract.
session.md Session lifecycle, KV prefix reuse, cancellation.
cache-sizing.md --cache-mb auto, the cache ceiling, and dense warm-up.
prefetch.md --prefetch K: the design and why it cannot change output (with the lossy knobs off).
expert-dropping.md --drop-cold-experts F: spending quality only where it buys a flash read, and why it is the one setting whose output is not reproducible.
mtp.md --mtp: drafting with the model's own MTP head and verifying a whole group per decode — lossless by construction, and why a wider decode can lose in the streamed regime.
ngram.md --ngram: drafting from text that repeats, with no head and no draft decode — and why a source that abstains costs exactly an unspeculated step.
expert-prediction.md --predict-log: how much of a routing can be known a layer early, measured against the predictor --prefetch already bets on — and why a good score still would not mean a faster decode.
route-ahead.md --route-ahead N: committing the routing to the N-layers-early prediction, so a prefetch of it can never miss — and what that costs in quality (experimental, lossy).
android-memory.md What reclaims the engine's memory on a phone, which levers exist (almost none), and why the cache hit rate is what the kernel judges you by.
pressure.md Cache policy under memory pressure: why an unaffordable budget starts a reclaim war, why the adaptive governor was retired, and what the fixed --cache-mb / --dense-weights levers do.

Measurements

Doc What it answers
benchmarks.md Measured results per model on Android, with device-pressure numbers.
benchmarks-gpt-oss.md gpt-oss-120b: a 58 GB model at 5.2× device RAM, and what it costs.
benchmark-method.md How the numbers are produced, so you can reproduce them.
warmup-analysis.md Why first tokens are slow, and the two regimes behind it.
bench-data/ Raw per-run CSVs and session notes. A dated archive — see its README.

Every benchmark figure in these docs is measured on the hardware named beside it. If you re-measure, update benchmark-method.md and the affected table together.