mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
Below one token's worst-case routed bytes, global LRU evicts precisely what it is about to read: the hit rate is 0% while the run still pays the cache's management time and its RAM. The only protection so far was cache_min_mb, a fixed floor that happens to sit above the cycle for the shipped models at their default top-k and stops holding the moment --n-expert-used widens the routing without touching the budget. The cycle is priced at init from the model's shape alone — every bound layer's entry_bytes times min(top_k, n_expert), at the top-k the run actually applies — and recorded as cache_cycle_mb in the metrics preamble next to the budget it should be judged against, so a committed CSV answers on its own whether its cache could ever have hit. The engine prints one stderr line when the resolved budget falls under it. It warns rather than refuses: the budget is legal and the output byte-identical, and the engine states the fact and leaves the choice — the argument that --cache-mb 0 is strictly better below the cliff stays in docs/cache-sizing.md, not in the engine's output. Docs updated in the same commit: cache-sizing.md (the guard section, past tense), roadmap.md (the item moves from still-worth-doing to shipped) and telemetry.md (the new preamble key). |
||
|---|---|---|
| .. | ||
| assets | ||
| bench-data | ||
| adding-a-model.md | ||
| android-memory.md | ||
| architecture.md | ||
| benchmark-method.md | ||
| benchmarks-gpt-oss.md | ||
| benchmarks.md | ||
| cache-sizing.md | ||
| expert-dropping.md | ||
| expert-prediction.md | ||
| limitations.md | ||
| moe-streaming.md | ||
| mtp.md | ||
| ngram.md | ||
| prefetch.md | ||
| pressure.md | ||
| README.md | ||
| roadmap.md | ||
| route-ahead.md | ||
| seam.md | ||
| session.md | ||
| telemetry.md | ||
| warmup-analysis.md | ||
Documentation
Start with architecture.md for the layer map, or moe-streaming.md for the idea the project is built on.
Understanding the design
| Doc | What it answers |
|---|---|
| architecture.md | How the layers fit together, and why llama.cpp is not forked. |
| moe-streaming.md | Why streaming experts from flash makes a >RAM model run at all. |
| seam.md | The exact contract with llama.cpp's public API, and how to upgrade the submodule. |
| limitations.md | What this does not do, what it cannot do, and the prior art it builds on. |
| roadmap.md | Themes worth exploring next. |
Using and extending it
| Doc | What it answers |
|---|---|
| adding-a-model.md | How to support a new MoE architecture (a recipe row plus a gate). |
| telemetry.md | The BMOE_* line protocol and CSV schema — the integration contract. |
| session.md | Session lifecycle, KV prefix reuse, cancellation. |
| cache-sizing.md | --cache-mb auto, the cache ceiling, and dense warm-up. |
| prefetch.md | --prefetch K: the design and why it cannot change output (with the lossy knobs off). |
| expert-dropping.md | --drop-cold-experts F: spending quality only where it buys a flash read, and why it is the one setting whose output is not reproducible. |
| mtp.md | --mtp: drafting with the model's own MTP head and verifying a whole group per decode — lossless by construction, and why a wider decode can lose in the streamed regime. |
| ngram.md | --ngram: drafting from text that repeats, with no head and no draft decode — and why a source that abstains costs exactly an unspeculated step. |
| expert-prediction.md | --predict-log: how much of a routing can be known a layer early, measured against the predictor --prefetch already bets on — and why a good score still would not mean a faster decode. |
| route-ahead.md | --route-ahead N: committing the routing to the N-layers-early prediction, so a prefetch of it can never miss — and what that costs in quality (experimental, lossy). |
| android-memory.md | What reclaims the engine's memory on a phone, which levers exist (almost none), and why the cache hit rate is what the kernel judges you by. |
| pressure.md | Cache policy under memory pressure: why an unaffordable budget starts a reclaim war, why the adaptive governor was retired, and what the fixed --cache-mb / --dense-weights levers do. |
Measurements
| Doc | What it answers |
|---|---|
| benchmarks.md | Measured results per model on Android, with device-pressure numbers. |
| benchmarks-gpt-oss.md | gpt-oss-120b: a 58 GB model at 5.2× device RAM, and what it costs. |
| benchmark-method.md | How the numbers are produced, so you can reproduce them. |
| warmup-analysis.md | Why first tokens are slow, and the two regimes behind it. |
| bench-data/ | Raw per-run CSVs and session notes. A dated archive — see its README. |
Every benchmark figure in these docs is measured on the hardware named beside it. If you re-measure, update benchmark-method.md and the affected table together.