mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#94)
* feat(moe): --drop-cold-experts, spend quality only where it buys I/O Turbo top-k drops the tail of a routing whether or not those experts were already in RAM. A resident expert costs no flash read, so that trade pays quality for nothing on the ~80% of decode routings that are cache hits. This adds the cache-aware version: skip a routed expert only when it is a cache MISS and the router weighted it below frac x (1/top-k). Replayed over the committed route traces at frac 1.0, decode phase, that avoids 66% of flash reads for 9.5% of the router's weight mass, against 59%/37% for --n-expert-used 3 — about 3x the reads avoided at a comparable cost. The threshold is a curve, not a switch: 0.75 trades 4.4% of the mass for 37% of the reads, better than --n-expert-used 5 on both axes. Implementation. The decision needs the FINAL router weights, which arrive several nodes after the topk where the streamer normally loads, so with the policy armed load_layer() is deferred to the terminal node of the layer's weight chain. Which node that is depends on the model's gating, so the hook learns it from the graph rather than carrying an architecture table; until it is known a layer loads at its topk node undropped. A dropped slot has its weight zeroed and its expert id repointed at the routing's top-weighted expert: an expert we decline to read may sit in reserved-but-uncommitted VM and mul_mat_id would still touch it, so the kernel is given memory that is certainly resident and multiplies it by exactly zero. Survivors are rescaled by default, since a systematically shrunk expert output perturbs the residual stream more than the missing contribution does. Prefill is excluded by default (cold cache, ~4x the weight mass discarded, and compute-bound anyway). The largest weight in a routing is always at least the uniform share, so frac <= 1 can never empty a layer; validate() enforces the bound and the top expert is pinned regardless. Gates: G8a proves the deferral and the learned terminal node are transparent (a threshold below any producible weight leaves the output byte-identical), G8b that full strength with the cache off never reaches an unloaded expert. Unlike every other knob this one is state-dependent: what gets dropped depends on what the cache held, so output is not reproducible across runs. Off by default, not in the app's settings, and NOT yet measured on device — the numbers above are a static replay and an upper bound. docs/expert- dropping.md states what is owed before it is recommended anywhere. * feat(app): expose cache-aware expert dropping in Settings Speed / quality -> Drop cold experts, as a percentage of the uniform share (off / 50 / 75 / 100). The engine takes a fraction; the app stores integer rungs, so the setting divides by 100 on the way to the flag. Disabled in mmap mode: the policy asks the expert source what is resident, and there is no expert source without the streamer. Included in the session signature, so changing it reopens the session rather than being ignored by a process already loaded. Off by default. This exists so the A/B can be run where the engine actually ships -- through the app, not a pushed CLI binary. * fix(moe): require the cache for dropping, and correct what it reports Review of the first two commits found the policy could be armed in a configuration where it is not cache-aware at all, and that two of the numbers it reports were wrong. - Require the LRU cache. With --cache-mb 0 query_residency answers all-miss, so the policy silently degenerated into an unconditional weight cut -- exactly what --n-expert-used already does, under a flag claiming to consult residency. validate() now rejects it, as it already did for --prefetch, and the app gates the setting on the same condition. - Fix experts_routed. It was incremented inside apply_drop, so it counted what the policy examined rather than what the router selected: layers before the terminal weight node is learned, and every un-armed phase, were missing from the denominator. The reported drop rate was a fraction of the wrong thing. - Re-learn instead of re-betting. If the node learned as terminal does not arrive, the deferral now also forgets it, so the next graph loads at the topk node while it re-learns. Deferring again on a stale guess would repeat the fault every token against a graph that had moved. - Point the gates at a real cache. G8a/G8b ran with the cache off, where the shared-slot path has no reserved-but-uncommitted memory -- so the id repointing, which is the design's whole safety argument, was never exercised. They now run against a constantly-evicting budget. Adds G8a' (asserts routings were examined and none dropped, so an inert-threshold flake fails legibly) and G8c (at top-k 1 dropping is a proven no-op, pinning both the top-expert guarantee and the threshold being taken against the effective top-k). Docs: three metrics change meaning under dropping and none of them said so. A dropped routing is a miss that is never looked up, so cache_hit_pct rises without the cache serving more, and token/layer_demand measure what was staged rather than routed -- documented in telemetry.md, pressure.md (size the cache with dropping off, then turn it on) and metrics.h. prefetch.md's "cannot change output" is scoped: under dropping a correct guess un-drops an expert. limitations.md gains the non-reproducibility entry, benchmark-method.md the axis plus a warning that reversing the run order cannot distinguish a moved drop rate from a contaminated cell, and architecture.md/runtime.h no longer claim unconditional determinism. Fixes two anchors the README rename broke, and a changelog sentence that quoted the equal-I/O row while drawing the equal-quality conclusion. App: Drop cold experts defaults to 75%. The default is a product decision taken on the maintainer's device; no benchmark for it is published here, and docs/expert-dropping.md says that plainly instead of implying a measured figure. The CLI stays off by default -- the byte-identity gates need a deterministic default.
This commit is contained in:
parent
719478908f
commit
45a90a2df5
31 changed files with 981 additions and 39 deletions
|
|
@ -21,7 +21,8 @@ for the idea the project is built on.
|
|||
| [telemetry.md](telemetry.md) | The `BMOE_*` line protocol and CSV schema — the integration contract. |
|
||||
| [session.md](session.md) | Session lifecycle, KV prefix reuse, cancellation. |
|
||||
| [cache-sizing.md](cache-sizing.md) | `--cache-mb auto`, the cache ceiling, and dense warm-up. |
|
||||
| [prefetch.md](prefetch.md) | `--prefetch K`: the design and why it cannot change output. |
|
||||
| [prefetch.md](prefetch.md) | `--prefetch K`: the design and why it cannot change output (with the lossy knobs off). |
|
||||
| [expert-dropping.md](expert-dropping.md) | `--drop-cold-experts F`: spending quality only where it buys a flash read, and why it is the one setting whose output is not reproducible. |
|
||||
| [android-memory.md](android-memory.md) | What reclaims the engine's memory on a phone, which levers exist (almost none), and why the cache hit rate is what the kernel judges you by. |
|
||||
| [pressure.md](pressure.md) | Cache policy under memory pressure: why an unaffordable budget starts a reclaim war, why the adaptive governor was retired, and what the fixed `--cache-mb` / `--dense-weights` levers do. |
|
||||
|
||||
|
|
|
|||
|
|
@ -45,8 +45,11 @@ Streaming experts serially needs three things from the inference engine. All thr
|
|||
already public in llama.cpp:
|
||||
|
||||
1. **A hook at routing time.** `llama_context_params.cb_eval` is called for every graph
|
||||
node. We ask for only the routing nodes (`ffn_moe_topk-<il>`); ggml computes and
|
||||
node. We ask for the routing nodes (`ffn_moe_topk-<il>`); ggml computes and
|
||||
synchronizes each alone, then calls us back with the selected expert ids materialized.
|
||||
The route trace and [cache-aware dropping](expert-dropping.md) additionally ask for each
|
||||
layer's `ffn_moe_weights*-<il>` chain — and dropping is the one path that *writes into* a
|
||||
graph tensor's contents rather than only rebinding `->data`. See [seam.md](seam.md).
|
||||
2. **The expert tensor pointers.** During a one-token warm-up we scan each graph node's
|
||||
sources for tensors named `blk.<il>.ffn_{gate,up,down}_exps.weight` and record the
|
||||
live `ggml_tensor*`. We then rebind their `->data`.
|
||||
|
|
@ -89,4 +92,6 @@ The composition root is `Session` (core/src/engine/session.cpp):
|
|||
so the gates and the interactive session share the same code path.
|
||||
|
||||
Greedy sampling makes the output a deterministic function of the graph — the property the
|
||||
[byte-identity gates](../tests/moe_gates.cpp) assert.
|
||||
[byte-identity gates](../tests/moe_gates.cpp) assert. That holds with the lossy knobs off. Under
|
||||
[`--drop-cold-experts`](expert-dropping.md) the hook edits routing weights from live cache state,
|
||||
which is not in the graph, so output becomes a function of the graph *and* the run's history.
|
||||
|
|
|
|||
|
|
@ -47,6 +47,7 @@ Vary one axis at a time:
|
|||
| threads (-t) | 2, 4, 8 | U-shape, 4 optimal, 8 regresses |
|
||||
| overlap | off, on | net gain **only over a warm cache** (hides residual flash wait behind compute); a net loss on a cold cache-0 stream, where I/O dwarfs compute |
|
||||
| n-expert-used | default, 6 | fewer active experts cut compute + I/O ~linearly (8→6 ≈ −25%), changes the output |
|
||||
| drop-cold-experts | off, 0.75, 1.0 | the second lossy axis, and the only **non-deterministic** one: what is skipped depends on cache state, so cells are noisier and the drop rate must be reported with the tok/s. Needs the cache on |
|
||||
| dense-weights | warm, anon | decisive well past RAM, near-neutral near it: on gpt-oss (5.2× RAM) `anon` drops majflt/token from the hundreds to **6–10** and compute with it; on Qwen (1.64× RAM) there is little dense-fault pressure to remove. Watch `majflt/token`, not just tok/s |
|
||||
|
||||
When sweeping `--n-expert-used`, run it as a **matched A/B against the model's own default**
|
||||
|
|
@ -125,6 +126,16 @@ So:
|
|||
Re-run such a cell; do not publish it. And sanity-check any matrix by **reversing the run order** —
|
||||
cells that move were measuring device state.
|
||||
|
||||
**The reversal check does not work under `--drop-cold-experts`.** There a cell can move because the
|
||||
*drop rate* moved — the policy reads live cache state, so the same command legitimately discards a
|
||||
different number of experts on a different run. That is the feature working, not the device
|
||||
contaminating the cell, and the two tells above cannot tell them apart. Always record
|
||||
`experts_dropped`/`experts_routed` (or the `moe-drop:` line) next to the tok/s: a dropping cell
|
||||
without its drop rate is uninterpretable, because the flag fixes a threshold and not a rate. Note
|
||||
also that a dropping run pays the same extra per-MoE-layer barriers a route-traced run does, so an
|
||||
A/B against `--n-expert-used` is not overhead-matched — see
|
||||
[expert-dropping.md](expert-dropping.md).
|
||||
|
||||
### Caveats
|
||||
|
||||
- **Thermal.** Sustained decode throttles. Warm up, then measure a steady window; discard
|
||||
|
|
|
|||
172
docs/expert-dropping.md
Normal file
172
docs/expert-dropping.md
Normal file
|
|
@ -0,0 +1,172 @@
|
|||
# Cache-aware expert dropping
|
||||
|
||||
`--drop-cold-experts F` skips a routed expert when it is **not in the cache** *and* the router
|
||||
weighted it below `F × (1 / top-k)` — that is, below `F` of the uniform share each of the `k`
|
||||
selected experts would get if the router split its mass evenly. Off by default.
|
||||
|
||||
It is the second lossy knob in the engine, after
|
||||
[turbo top-k](../README.md#turbo-top-k--the-measured-lossy-option), and it exists because the first one
|
||||
spends quality in a place it does not have to.
|
||||
|
||||
## Why cache state belongs in the decision
|
||||
|
||||
`--n-expert-used k` drops the routing's tail unconditionally: slot 7 and slot 8 go, whether or not
|
||||
they were already sitting in RAM. But an expert that is already resident costs **no flash read** —
|
||||
and on a streamed decode, flash reads are what the token is waiting for
|
||||
([decode is I/O-bound](benchmarks.md)). Dropping a resident expert pays quality for nothing.
|
||||
|
||||
Turn that around and the policy writes itself: **spend quality only where it buys I/O**. Keep every
|
||||
resident expert however small its weight; consider dropping only the ones that would cost a read,
|
||||
and only when the router says they barely matter.
|
||||
|
||||
## What it costs and what it buys
|
||||
|
||||
Replayed over the committed route traces (`docs/bench-data/2026-07-15-route-trace/`), decode phase,
|
||||
threshold at the uniform share (`F = 1.0`):
|
||||
|
||||
| policy | flash reads avoided | router weight discarded |
|
||||
|---|---|---|
|
||||
| `--drop-cold-experts 1.0` | **66%** | **9.5%** |
|
||||
| `--n-expert-used 5` | 23% | 10.6% |
|
||||
| `--n-expert-used 3` | 59% | 36.8% |
|
||||
|
||||
(Qwen3-30B-A3B at k=6; Gemma-4-26B-A4B is within a point and a half on both columns: 67.4% / 8.2%. On gpt-oss-120b at k=2 the
|
||||
policy matches `--n-expert-used 1`'s read saving while discarding 25% of the weight mass instead of
|
||||
42%.)
|
||||
|
||||
At a comparable quality cost the cache-aware policy avoids roughly **three times** the reads. The
|
||||
reason is visible in the third column of the trace: about 80% of decode routings are cache hits, and
|
||||
the policy leaves every one of them alone.
|
||||
|
||||
`F` is a curve, not a switch. At `F = 0.75` the same model trades 4.4% of the weight mass for 37% of
|
||||
the reads — still better than `--n-expert-used 5` on **both** axes.
|
||||
|
||||
These are replay numbers and an **upper bound**: skipping a read changes what the cache holds later,
|
||||
so the real hit pattern drifts from the recorded one. The on-device A/B is what settles it.
|
||||
|
||||
## Two properties worth knowing
|
||||
|
||||
**A routing is never emptied.** The largest weight in a routing is always at least the uniform
|
||||
share, so at `F ≤ 1.0` the top expert can never fall below the threshold. `validate()` rejects
|
||||
`F > 1.0` for that reason, and the implementation additionally pins the top-weighted expert, so the
|
||||
guarantee does not rest on the bound alone.
|
||||
|
||||
**It requires the expert cache.** With `--cache-mb 0` every expert reads as a miss, so the policy
|
||||
would stop being cache-aware and become an unconditional weight cut — which is what
|
||||
`--n-expert-used` already does, without claiming to consult residency. `validate()` rejects the
|
||||
combination, the same way it rejects `--prefetch` without a cache.
|
||||
|
||||
**It changes what `--prefetch` means.** Speculation is normally output-neutral by construction. Here
|
||||
residency is an *input* to the policy, so a correct guess un-drops an expert that would otherwise
|
||||
have been discarded: prefetch depth becomes an output-affecting setting. The decision point also
|
||||
settles pending speculation a few nodes after it was issued, which shortens the overlap window the
|
||||
prefetch exists for — treat the two as interacting, not composable.
|
||||
|
||||
**Prefill is excluded by default.** With a cold cache almost every expert is a miss, and the same
|
||||
threshold discards ~42% of the weight mass instead of ~9%. Prefill is compute-bound anyway, so there
|
||||
is little to win. `--drop-in-prefill` arms it for experiments.
|
||||
|
||||
## The output is no longer reproducible
|
||||
|
||||
This is the real novelty, and the reason the flag is off by default and named the way it is.
|
||||
|
||||
`--n-expert-used` is lossy but **deterministic**: same prompt, same config, same tokens. Dropping is
|
||||
lossy and **state-dependent** — what gets discarded depends on what the cache happened to hold,
|
||||
which depends on everything decoded before it. The same prompt can produce different text across
|
||||
runs, and a benchmark cell is noisier because the drop rate itself varies.
|
||||
|
||||
The greedy byte-identity gates therefore do not cover the policy's output, and cannot: there is
|
||||
nothing stable to compare against. They cover the machinery instead (see below).
|
||||
|
||||
## How it is implemented
|
||||
|
||||
The decision needs the **final** router weights, and those are produced several graph nodes after
|
||||
the topk node where the streamer normally loads. So with the policy armed, `load_layer()` is
|
||||
postponed from the topk node to the terminal node of the layer's weight chain — the last node before
|
||||
the expert matmul consumes either the ids or the weights.
|
||||
|
||||
Which node is terminal depends on the model's gating (`_norm`, `_softmax`, `_scaled`, or none), so
|
||||
the hook **learns** it from the graph instead of carrying a per-architecture table: the first graph
|
||||
of a run records the chain, and dropping starts from the second. A layer whose shape has not been
|
||||
seen yet simply loads at its topk node, undropped. That costs a run its first token's dropping and
|
||||
nothing else, and it keeps [hard rule 4](../CLAUDE.md) — no model-specific constants in the
|
||||
streaming path.
|
||||
|
||||
At the decision point two edits happen, both before anything reads them:
|
||||
|
||||
1. the dropped slot's **weight is zeroed**, and with `drop_renorm` (default on) the survivors are
|
||||
scaled so the routing keeps its original total mass;
|
||||
2. the dropped slot's **id is repointed** at the routing's top-weighted expert.
|
||||
|
||||
The second edit is not cosmetic. An expert the engine declines to read may sit in a
|
||||
reserved-but-uncommitted slot, and `mul_mat_id` would still touch it. Pointing the slot at an expert
|
||||
that is certainly resident makes the kernel read valid memory and multiply it by exactly zero. It
|
||||
costs a duplicate matmul — the right trade on a decode bound by flash rather than arithmetic.
|
||||
|
||||
Renormalisation matters more than it looks: without it the layer's expert output is systematically
|
||||
scaled down by the discarded mass, a perturbation of the residual stream the model never sees in
|
||||
training. `--drop-no-renorm` exists to A/B that claim.
|
||||
|
||||
Cost of the extra barriers: the policy asks for each layer's weight nodes, a handful more
|
||||
synchronisation points per MoE layer on tensors of a few floats. The same asks a route trace makes.
|
||||
|
||||
## Measuring it
|
||||
|
||||
The engine reports what the policy actually did, which the flag alone cannot tell you — the
|
||||
threshold is fixed, the drop rate is not:
|
||||
|
||||
```
|
||||
moe-drop: <dropped>/<routed> routed experts dropped (<pct>%), threshold <F> x uniform
|
||||
```
|
||||
|
||||
The route trace gains a `dropped` column: `weight` and `residency` stay as the **router** produced
|
||||
them, `dropped` records what the policy then did, and `expert_bytes` is 0 for a dropped routing
|
||||
because it costs no read. That is enough to replay a real run against the offline model and check
|
||||
whether the upper bound held. See [telemetry.md](telemetry.md).
|
||||
|
||||
## Gates
|
||||
|
||||
`bmoe_moe_gates` covers the machinery, not the policy's output:
|
||||
|
||||
- **G8a** — with a threshold below any weight the router can produce, nothing is dropped and the
|
||||
output is **byte-identical** to the undropped stream. This proves the deferral and the learned
|
||||
terminal node are transparent, separating "the plumbing is correct" from "the policy is lossy" —
|
||||
a regression in the first would otherwise hide behind the expected difference. **G8a'** asserts
|
||||
the count separately (`experts_routed > 0`, `experts_dropped == 0`), so "a weight happened to fall
|
||||
under the threshold" fails legibly instead of as a mysterious byte mismatch.
|
||||
- **G8b** — at full strength against a cache small enough to be evicting constantly, so dropped
|
||||
experts really do land on slots the cache has released. Generation still completes: the id
|
||||
repointing means no matmul ever reads reserved-but-uncommitted memory. (The gates deliberately do
|
||||
*not* run this with the cache off — there the shared-slot path has no uncommitted memory, so the
|
||||
safety property the repointing exists for would go untested.)
|
||||
- **G8c** — forcing top-k to 1 makes every routed expert the top one, so dropping must be a no-op at
|
||||
any threshold and the output must match the undropped k=1 run byte for byte. This pins both the
|
||||
top-expert guarantee and the fact that the threshold is taken against the **effective** top-k
|
||||
discovered at runtime — a hardcoded width would not survive the override.
|
||||
|
||||
## Defaults, and where the numbers do and do not come from
|
||||
|
||||
The **CLI defaults it off**, and will keep doing so: the byte-identity gates need a deterministic
|
||||
default, and an instrument should not quietly change the thing it measures.
|
||||
|
||||
The **app ships it at 75%** — under **Speed / quality → Drop cold experts**, with rungs 50 / 75 /
|
||||
100 as percentages of the uniform share. It is disabled there in mmap mode and with the cache off,
|
||||
the same two conditions `validate()` enforces.
|
||||
|
||||
That default is a product decision taken on the maintainer's own device measurement. **It is not
|
||||
backed by a published benchmark in this repository**, and the tables in the README deliberately
|
||||
carry no rows for it — they are a deterministic protocol and this knob is not deterministic. Nothing
|
||||
here should be read as "75% is worth X%"; the honest claim is narrower: the replay above says the
|
||||
shape of the trade is favourable, and the default was chosen after checking it on hardware.
|
||||
|
||||
What is still owed before this is recommended beyond that:
|
||||
|
||||
- a published decode A/B against `--n-expert-used` at matched tok/s, with the device state recorded
|
||||
the way [benchmark-method.md](benchmark-method.md) requires;
|
||||
- a quality comparison at that matched speed — the whole thesis is that this knob buys the same
|
||||
throughput for less damage, and only a side-by-side can support it;
|
||||
- a re-run of the replay against a real traced run with the `dropped` column, to see how far the
|
||||
static upper bound overstated the win.
|
||||
|
||||
The [`layer-lfu` entry in the roadmap](roadmap.md) is the standing reminder for why the third one
|
||||
matters: it simulated exactly as predicted and was ~30% slower in reality.
|
||||
|
|
@ -18,6 +18,11 @@ serial path, and only a single ~25-line hook (with an explicit sunset) for the o
|
|||
|
||||
## Limitations
|
||||
|
||||
- **One setting makes output non-reproducible.** Every other knob is deterministic given a
|
||||
configuration: `--n-expert-used` changes the output, but changes it the same way on every run.
|
||||
[`--drop-cold-experts`](expert-dropping.md) decides per routing from live cache state, so the
|
||||
same prompt and the same flags can decode differently run to run, and the byte-identity gates
|
||||
cannot cover its output — only its machinery. Off by default in the CLI.
|
||||
- **n=1 only.** The expert sparsity exists only for single-token decode, so streaming is
|
||||
incompatible with speculative decoding or batching. Prefill streams the union of the
|
||||
prompt's routed experts (still far below the full bank, but larger than one token's).
|
||||
|
|
|
|||
|
|
@ -32,7 +32,10 @@ Ordering is guaranteed by ggml's eval-callback loop: the node we mark is compute
|
|||
buffers until this layer's matmul has synchronized. Correct on any backend.
|
||||
|
||||
The result is **lossless**: byte-identical to running with every expert resident, asserted
|
||||
by the gates.
|
||||
by the gates. That is the streaming path itself; two opt-in knobs deliberately trade output for
|
||||
speed on top of it — `--n-expert-used` (fewer experts per token) and
|
||||
[`--drop-cold-experts`](expert-dropping.md) (skip an expert that would cost a read and was barely
|
||||
weighted). Both are off unless asked for, which is what keeps the sentence above true by default.
|
||||
|
||||
## Residency modes
|
||||
|
||||
|
|
|
|||
|
|
@ -32,7 +32,11 @@ captures most of the benefit. Prefetch requires the LRU cache to be on — eithe
|
|||
|
||||
## How it stays correct and out of the way
|
||||
|
||||
The speculative path never delays real work and never changes output:
|
||||
The speculative path never delays real work and never changes output — with the lossy knobs off.
|
||||
(Under [`--drop-cold-experts`](expert-dropping.md) residency is an *input* to the routing policy,
|
||||
so a correct guess un-drops an expert that would otherwise have been discarded. Prefetch depth
|
||||
becomes output-affecting there; everything below still holds for the bytes themselves.)
|
||||
|
||||
|
||||
- **Same bytes.** A speculative read is the *identical* read a real miss would issue — same file
|
||||
offset, same destination buffer (`lbuf_[p][il] + e*slice`). A prefetched expert is therefore
|
||||
|
|
|
|||
|
|
@ -30,7 +30,10 @@ is:
|
|||
|
||||
## Why a budget cannot be a constant
|
||||
|
||||
The expert cache is the one lever that trades RAM for flash reads, so the temptation is to set it as
|
||||
The expert cache is the one lever that trades RAM for flash reads (
|
||||
[`--drop-cold-experts`](expert-dropping.md) is the other kind of trade — quality for flash reads —
|
||||
and the two interact: a squeezed cache raises the miss rate, which raises the drop rate, so memory
|
||||
pressure degrades output quality there instead of only throughput). The temptation is to set it as
|
||||
large as the device seems to allow. On a phone that is the wrong shape of decision, for three
|
||||
reasons that are measured rather than argued:
|
||||
|
||||
|
|
@ -115,6 +118,11 @@ not a floor), `layer_demand_MiB` (the mechanical floor), `cache_budget_MiB` (the
|
|||
effect), `cache_hit_pct`. Per token, `dense_resident_frac` says whether the dense set is holding in
|
||||
RAM (the live signal now that the cache-residency governor sensor is gone).
|
||||
|
||||
This sizing procedure assumes dropping is off. With
|
||||
[`--drop-cold-experts`](expert-dropping.md) on, dropped routings are misses that never reach the
|
||||
cache, so `cache_hit_pct` reads high and `token_demand_MiB` reads low for the same budget — size
|
||||
the cache first, then turn dropping on.
|
||||
|
||||
Reading `cache_hit_pct` against `token_demand_MiB` is how you tell whether a fixed `--cache-mb N` is
|
||||
earning its RAM: a budget near or below one token's demand holds no history between tokens and its
|
||||
hits are only inter-token correlation; well above it, a high hit rate means real reuse.
|
||||
|
|
|
|||
|
|
@ -113,6 +113,19 @@ routed) are supported; other `build_moe_ffn` models are one recipe row each. The
|
|||
frontier is architectures whose routing node is not the shared `ffn_moe_topk` — custom gating,
|
||||
which the capture/stream hook would need to learn. See [adding-a-model.md](adding-a-model.md).
|
||||
|
||||
## Skipping reads the router barely wants — built, unmeasured
|
||||
|
||||
`--drop-cold-experts` ([expert-dropping.md](expert-dropping.md)) is the first lever that treats
|
||||
quality and I/O as a *joint* budget rather than two separate knobs: an expert already in the cache
|
||||
runs however small its weight, and only a routing that would cost a flash read can be dropped. On
|
||||
the recorded traces that is worth ~3× the reads of turbo top-k for a comparable weight cost, which
|
||||
is the strongest offline case any remaining lever has shown.
|
||||
|
||||
What it does **not** have is a device measurement, and the previous entry on this page is the reason
|
||||
that matters: `layer-lfu` simulated well and was ~30% slower in reality. The open questions are the
|
||||
device A/B against turbo top-k at matched throughput, the quality comparison at that speed, and how
|
||||
far the static replay overstated the win once dropping starts changing what the cache holds.
|
||||
|
||||
## Expert quantization on the fly
|
||||
|
||||
Storing streamed experts at a lower precision than the resident parts to cut read volume, if it
|
||||
|
|
|
|||
14
docs/seam.md
14
docs/seam.md
|
|
@ -24,10 +24,22 @@ come from the arch's recipe — `ffn_{gate,up,down}_exps` for the split layout,
|
|||
throughout — capture observes, it does not isolate. `ggml_tensor` is a public struct, so
|
||||
reading `->name`, `->ne`, `->nb` and writing `->data` is public API surface.
|
||||
|
||||
**Stream phase** (real generation). We return true only for `ffn_moe_topk-<il>`. The
|
||||
**Stream phase** (real generation). We return true for `ffn_moe_topk-<il>`. The
|
||||
non-ask callback then hands us that node with the selected expert ids materialized; we
|
||||
gather them (stride-aware) and trigger the slice reads.
|
||||
|
||||
Two optional jobs ask for more: the route trace and
|
||||
[cache-aware dropping](expert-dropping.md) also want each layer's `ffn_moe_weights*-<il>` chain,
|
||||
which is another barrier per node but no new kind of access — same public struct, same read of
|
||||
`->data`.
|
||||
|
||||
Dropping does go one step further, and it is the only place the engine **writes into** a graph
|
||||
tensor's contents rather than repointing `->data` at its own buffer: at the terminal node of the
|
||||
weight chain it zeroes a dropped slot's weight and repoints that slot's expert id. Both tensors are
|
||||
scratch the graph produced and has not yet consumed, so this alters the values flowing through the
|
||||
run — deliberately, that is what the lossy policy *is* — and never llama.cpp's own state, its
|
||||
weights, or its control flow. It stays inside the same callback contract; nothing is patched.
|
||||
|
||||
## 2. gguf offsets
|
||||
|
||||
`gguf_init_from_file(..., no_alloc=true)` + `gguf_get_data_offset` +
|
||||
|
|
|
|||
|
|
@ -51,6 +51,10 @@ BMOE_PROGRESS {"step":<int>,"steps":<int>,"wall_ms":<float>,"io_ms":<float>,
|
|||
can be a large share of the token; at steady state it is near zero. Surfacing it stops the "all
|
||||
compute" reading on warm-up tokens where the real cost is cache churn, not matmul.
|
||||
- `cache_hit_pct` is the cumulative cache hit rate, or `-1` when no cache is used.
|
||||
**Under [`--drop-cold-experts`](expert-dropping.md) read it with care:** a dropped routing is a
|
||||
miss that is never looked up, so it leaves both sides of the ratio and the reported hit rate
|
||||
rises without the cache having served anything more. Compare runs at the same drop rate, or read
|
||||
`experts_dropped` next to it.
|
||||
- `majflt` / `cpu_ms` **decompose the `compute_ms` residual** — the whole point being that "compute"
|
||||
above is a catch-all that silently absorbs page faults and scheduler stalls, not just matmul.
|
||||
They are measured directly around `llama_decode` (no submodule patch needed): `majflt` is the
|
||||
|
|
@ -101,6 +105,16 @@ moe-prefetch: <mib> MiB speculative, <useful>/<prefetched> experts useful (<pct>
|
|||
`<prefetched>` the experts fully read ahead, and `<useful>` how many of those a later routing
|
||||
actually hit. See [prefetch.md](prefetch.md).
|
||||
|
||||
With `--drop-cold-experts F` a `moe-drop:` line is added:
|
||||
|
||||
```
|
||||
moe-drop: <dropped>/<routed> routed experts dropped (<pct>%), threshold <F> x uniform
|
||||
```
|
||||
|
||||
The flag fixes a *threshold*, not a rate: how much is actually discarded depends on what the cache
|
||||
held, so this line — not the flag — is what a run traded. See
|
||||
[expert-dropping.md](expert-dropping.md).
|
||||
|
||||
Under `--overlap` the `moe-stream:` line additionally reports `stall_s/tok=<s>` — the mean
|
||||
wall time per token that compute threads waited for expert reads to complete. It is `0` in
|
||||
serial mode (where the read wait is already folded into decode time).
|
||||
|
|
@ -130,7 +144,10 @@ sampled dense-weight residency, `-1` when unmeasured. All are additive: older CS
|
|||
so consumers must read by column NAME (from the header row) and treat any as optional. The `# summary`
|
||||
line likewise gains `stall_s/tok=<s>`, `mgmt_s/tok=<s>`, `majflt/tok=<f>`, `cpu_s/tok=<s>`,
|
||||
`token_demand_MiB=<f>` (the expert bytes one token routes, measured — where cache hits start, NOT a
|
||||
floor to defend; see [pressure.md](pressure.md)) and `layer_demand_MiB=<f>` (the widest layer's routed
|
||||
floor to defend; see [pressure.md](pressure.md)), `experts_routed=<n>` / `experts_dropped=<n>` (what
|
||||
[cache-aware dropping](expert-dropping.md) actually discarded during generation — the flag sets a
|
||||
threshold, not a rate, so this is the only record of the trade a run made) and
|
||||
`layer_demand_MiB=<f>` (the widest layer's routed
|
||||
bytes: the mechanical floor the cache must be able to stage); see the `io_ms` note above for how the
|
||||
read-time columns are reinterpreted under overlap.
|
||||
|
||||
|
|
@ -161,6 +178,10 @@ per routed expert. **A traced run is not a benchmark run** — the numbers in th
|
|||
traced run are slower than the real thing, and `mgmt_ms` in particular shifts, because settling
|
||||
speculative prefetch moves outside the window that times it.
|
||||
|
||||
Columns are **append-only** within `v1`, like the metrics CSV: `dropped` was added after
|
||||
`expert_bytes`, so consumers must read by column NAME and treat any column as optional rather than
|
||||
indexing by position.
|
||||
|
||||
The file is long format: a `#` preamble carrying the run's static facts, then one row per routed
|
||||
expert. Conceptually it is a matrix — rows are steps, columns are layers — and a **cell** is the
|
||||
`n_expert_used` rows sharing `(turn, phase, step, layer)`.
|
||||
|
|
@ -169,7 +190,7 @@ expert. Conceptually it is a matrix — rows are steps, columns are layers — a
|
|||
# route_trace v1
|
||||
# model=<path> arch=<string> n_layer=<int> n_expert=<int> n_expert_used=<int>
|
||||
# layer=<int> expert_bytes=<int> dense_bytes=<int> (one per layer)
|
||||
turn,phase,step,layer,slot,expert,weight,residency,expert_bytes
|
||||
turn,phase,step,layer,slot,expert,weight,residency,expert_bytes,dropped
|
||||
```
|
||||
|
||||
| column | meaning |
|
||||
|
|
@ -183,6 +204,7 @@ turn,phase,step,layer,slot,expert,weight,residency,expert_bytes
|
|||
| `weight` | the final applied routing weight, after whatever softmax/normalise/scale the architecture uses. `nan` when the graph exposed no weight node — "unknown", never `0`. |
|
||||
| `residency` | `0` = miss (this routing reads from flash), `1` = hit, `2` = hit on a speculative prefetch's first touch. |
|
||||
| `expert_bytes` | flash bytes this routing reads; `0` unless `residency=0`. |
|
||||
| `dropped` | `1` when [cache-aware dropping](expert-dropping.md) discarded this routing — a miss weighted below the threshold, never read, weight zeroed. Always `0` with `--drop-cold-experts` off. |
|
||||
|
||||
`(turn, phase, step, layer, slot)` is unique. Two asymmetries are deliberate:
|
||||
|
||||
|
|
@ -195,6 +217,11 @@ turn,phase,step,layer,slot,expert,weight,residency,expert_bytes
|
|||
streamed, so there is nothing to measure per step: `dense_bytes` is what a cold layer costs to
|
||||
page in, stated once. Per-layer *I/O time* is absent for the same kind of reason — under
|
||||
`--overlap` reads complete asynchronously, so any per-layer timing would be fiction.
|
||||
- **`weight` and `residency` describe the router; `dropped` describes the policy.** When dropping is
|
||||
on, a discarded routing keeps the weight the router gave it and the residency it faced — the trace
|
||||
records the routing that was *chosen* — while `expert_bytes` falls to `0`, because a dropped
|
||||
expert is never read. Summing `expert_bytes` therefore still measures real flash traffic, and
|
||||
`dropped` is what explains the gap against `residency==0`.
|
||||
|
||||
**The last layer has only one prefill step, and that is real.** Before the final layer's FFN,
|
||||
llama.cpp gathers only the tokens whose logits were asked for (`inp_out_ids`; see `il == n_layer
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue