mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
The dense policy assumed the dense set fits in RAM. Qwen3.8-Flash-Next breaks that with its 51B n-gram table (per_layer_token_embd, ~28.8 GB at IQ4_NL), which every mode failed on in its own way. A dense tensor larger than the kernel's MemAvailable is now held back under every mode: it stays mmap'd with MADV_RANDOM, leaves the warm sweep, the residency sensor and the auto cache budget, and one stderr line names it. The bound is size, not access shape: a row-gathered table that fits keeps its mode. Inert on every other supported model (Qwen3.6-35B matches its baselines to the decimal under anon, warm, mmap and auto); gates pass; the guard fires on device and pinned dense weights survive load. |
||
|---|---|---|
| .. | ||
| assets | ||
| bench-data | ||
| adding-a-model.md | ||
| android-memory.md | ||
| architecture.md | ||
| benchmark-method.md | ||
| benchmarks-gpt-oss.md | ||
| benchmarks.md | ||
| cache-sizing.md | ||
| community-benchmarks.md | ||
| expert-dropping.md | ||
| expert-prediction.md | ||
| limitations.md | ||
| moe-streaming.md | ||
| mtp.md | ||
| ngram.md | ||
| prefetch.md | ||
| pressure.md | ||
| README.md | ||
| roadmap.md | ||
| route-ahead.md | ||
| seam.md | ||
| session.md | ||
| telemetry.md | ||
| warmup-analysis.md | ||
Documentation
Start with architecture.md for the layer map, or moe-streaming.md for the idea the project is built on.
Understanding the design
| Doc | What it answers |
|---|---|
| architecture.md | How the layers fit together, and why llama.cpp is not forked. |
| moe-streaming.md | Why streaming experts from flash makes a >RAM model run at all. |
| seam.md | The exact contract with llama.cpp's public API, and how to upgrade the submodule. |
| limitations.md | What this does not do, what it cannot do, and the prior art it builds on. |
| roadmap.md | Themes worth exploring next. |
Using and extending it
| Doc | What it answers |
|---|---|
| adding-a-model.md | How to support a new MoE architecture (a recipe row plus a gate). |
| telemetry.md | The BMOE_* line protocol and CSV schema — the integration contract. |
| session.md | Session lifecycle, KV prefix reuse, cancellation. |
| cache-sizing.md | --cache-mb auto, the cache ceiling, and dense warm-up. |
| prefetch.md | --prefetch K: the design and why it cannot change output (with the lossy knobs off). |
| expert-dropping.md | --drop-cold-experts F: spending quality only where it buys a flash read, and why it is the one setting whose output is not reproducible. |
| mtp.md | --mtp: drafting with the model's own MTP head and verifying a whole group per decode — lossless by construction, and why a wider decode can lose in the streamed regime. |
| ngram.md | --ngram: drafting from text that repeats, with no head and no draft decode — and why a source that abstains costs exactly an unspeculated step. |
| expert-prediction.md | --predict-log: how much of a routing can be known a layer early, measured against the predictor --prefetch already bets on — and why a good score still would not mean a faster decode. |
| route-ahead.md | --route-ahead N: committing the routing to the N-layers-early prediction, so a prefetch of it can never miss — and what that costs in quality (experimental, lossy). |
| android-memory.md | What reclaims the engine's memory on a phone, which levers exist (almost none), and why the cache hit rate is what the kernel judges you by. |
| pressure.md | Cache policy under memory pressure: why an unaffordable budget starts a reclaim war, why the adaptive governor was retired, and what the fixed --cache-mb / --dense-weights levers do. |
Measurements
| Doc | What it answers |
|---|---|
| benchmarks.md | Measured results per model on Android, with device-pressure numbers. |
| benchmarks-gpt-oss.md | gpt-oss-120b: a 58 GB model at 5.2× device RAM, and what it costs. |
| benchmark-method.md | How the numbers are produced, so you can reproduce them. |
| community-benchmarks.md | Results on hardware we do not own, the one-command protocol (scripts/bench-report.sh), and how to submit a row. |
| warmup-analysis.md | Why first tokens are slow, and the two regimes behind it. |
| bench-data/ | Raw per-run CSVs and session notes. A dated archive — see its README. |
Every benchmark figure in these docs is measured on the hardware named beside it. If you re-measure, update benchmark-method.md and the affected table together.