mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
Experimental, off by default. Before a decode routing is committed, every expert already in the LRU cache gets its score raised by L times the token's score range and the top-k is taken again, so a near-tie goes to the expert already in RAM (Skliar et al., arXiv:2412.00099). The same number of experts runs; fewer are read from flash. Scores are read from the tensor the graph itself sorted, exact for any gating function. Desktop, Qwen3.6-35B Q4_K_M at L=0.15: 258 to 119 MiB of flash per token, 2.37 to 3.84 tok/s, perplexity +1 to 4 %, tinyMMLU 88 to 84/100, HumanEval-50 42 = 42. The on-device A/B is still owed, hence experimental. Also: --ppl / --ppl-step / --ppl-list / --ppl-choices (teacher-forced perplexity, one token per decode so cache-dependent policies are priced where they act), scripts/tinymmlu-bench.py, scripts/humaneval-bench.py, gates G8d/G8e, app switch "Prefer cached experts" under Experimental, docs/cache-aware-substitution.md.
4.1 KiB
4.1 KiB
Documentation
Start with architecture.md for the layer map, or moe-streaming.md for the idea the project is built on.
Understanding the design
| Doc | What it answers |
|---|---|
| architecture.md | How the layers fit together, and why llama.cpp is not forked. |
| moe-streaming.md | Why streaming experts from flash makes a >RAM model run at all. |
| seam.md | The exact contract with llama.cpp's public API, and how to upgrade the submodule. |
| limitations.md | What this does not do, what it cannot do, and the prior art it builds on. |
| roadmap.md | Themes worth exploring next. |
Using and extending it
| Doc | What it answers |
|---|---|
| adding-a-model.md | How to support a new MoE architecture (a recipe row plus a gate). |
| telemetry.md | The BMOE_* line protocol and CSV schema — the integration contract. |
| session.md | Session lifecycle, KV prefix reuse, cancellation. |
| cache-sizing.md | --cache-mb auto, the cache ceiling, and dense warm-up. |
| prefetch.md | --prefetch K: the design and why it cannot change output (with the lossy knobs off). |
| expert-dropping.md | --drop-cold-experts F: spending quality only where it buys a flash read, and why a cache-dependent setting's output is not reproducible. |
| row-gathered-tables.md | --row-stream: serving a dense table the graph only gathers rows from out of flash, and how the engine decides which tables those are without naming one. |
| cache-aware-substitution.md | --expert-substitute L: re-ranking a routing toward the experts already in the cache, the paper it comes from, and why a wide scoring batch cannot price it. |
| mtp.md | --mtp: drafting with the model's own MTP head and verifying a whole group per decode — lossless by construction, and why a wider decode can lose in the streamed regime. |
| ngram.md | --ngram: drafting from text that repeats, with no head and no draft decode — and why a source that abstains costs exactly an unspeculated step. |
| expert-prediction.md | --predict-log: how much of a routing can be known a layer early, measured against the predictor --prefetch already bets on — and why a good score still would not mean a faster decode. |
| route-ahead.md | --route-ahead N: committing the routing to the N-layers-early prediction, so a prefetch of it can never miss — and what that costs in quality (experimental, lossy). |
| android-memory.md | What reclaims the engine's memory on a phone, which levers exist (almost none), and why the cache hit rate is what the kernel judges you by. |
| pressure.md | Cache policy under memory pressure: why an unaffordable budget starts a reclaim war, why the adaptive governor was retired, and what the fixed --cache-mb / --dense-weights levers do. |
Measurements
| Doc | What it answers |
|---|---|
| benchmarks.md | Measured results per model on Android, with device-pressure numbers. |
| benchmarks-gpt-oss.md | gpt-oss-120b: a 58 GB model at 5.2× device RAM, and what it costs. |
| benchmark-method.md | How to measure this engine on any machine: what each knob does, when to move it, and the rules that keep a matrix honest. |
| community-benchmarks.md | Results on hardware we do not own, the one-command protocol (scripts/bench-report.sh), and how to submit a row. |
| warmup-analysis.md | Why first tokens are slow, and the two regimes behind it. |
| bench-data/ | Raw per-run CSVs and session notes. A dated archive — see its README. |
Every benchmark figure in these docs is measured on the hardware named beside it. If you re-measure, update benchmark-method.md and the affected table together.