feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#94)

* feat(moe): --drop-cold-experts, spend quality only where it buys I/O

Turbo top-k drops the tail of a routing whether or not those experts were
already in RAM. A resident expert costs no flash read, so that trade pays
quality for nothing on the ~80% of decode routings that are cache hits.

This adds the cache-aware version: skip a routed expert only when it is a
cache MISS and the router weighted it below frac x (1/top-k). Replayed over
the committed route traces at frac 1.0, decode phase, that avoids 66% of
flash reads for 9.5% of the router's weight mass, against 59%/37% for
--n-expert-used 3 — about 3x the reads avoided at a comparable cost. The
threshold is a curve, not a switch: 0.75 trades 4.4% of the mass for 37% of
the reads, better than --n-expert-used 5 on both axes.

Implementation. The decision needs the FINAL router weights, which arrive
several nodes after the topk where the streamer normally loads, so with the
policy armed load_layer() is deferred to the terminal node of the layer's
weight chain. Which node that is depends on the model's gating, so the hook
learns it from the graph rather than carrying an architecture table; until
it is known a layer loads at its topk node undropped. A dropped slot has its
weight zeroed and its expert id repointed at the routing's top-weighted
expert: an expert we decline to read may sit in reserved-but-uncommitted VM
and mul_mat_id would still touch it, so the kernel is given memory that is
certainly resident and multiplies it by exactly zero. Survivors are rescaled
by default, since a systematically shrunk expert output perturbs the residual
stream more than the missing contribution does.

Prefill is excluded by default (cold cache, ~4x the weight mass discarded,
and compute-bound anyway). The largest weight in a routing is always at least
the uniform share, so frac <= 1 can never empty a layer; validate() enforces
the bound and the top expert is pinned regardless.

Gates: G8a proves the deferral and the learned terminal node are transparent
(a threshold below any producible weight leaves the output byte-identical),
G8b that full strength with the cache off never reaches an unloaded expert.

Unlike every other knob this one is state-dependent: what gets dropped
depends on what the cache held, so output is not reproducible across runs.
Off by default, not in the app's settings, and NOT yet measured on device —
the numbers above are a static replay and an upper bound. docs/expert-
dropping.md states what is owed before it is recommended anywhere.

* feat(app): expose cache-aware expert dropping in Settings

Speed / quality -> Drop cold experts, as a percentage of the uniform
share (off / 50 / 75 / 100). The engine takes a fraction; the app stores
integer rungs, so the setting divides by 100 on the way to the flag.

Disabled in mmap mode: the policy asks the expert source what is resident,
and there is no expert source without the streamer. Included in the session
signature, so changing it reopens the session rather than being ignored by
a process already loaded.

Off by default. This exists so the A/B can be run where the engine actually
ships -- through the app, not a pushed CLI binary.

* fix(moe): require the cache for dropping, and correct what it reports

Review of the first two commits found the policy could be armed in a
configuration where it is not cache-aware at all, and that two of the
numbers it reports were wrong.

- Require the LRU cache. With --cache-mb 0 query_residency answers
  all-miss, so the policy silently degenerated into an unconditional
  weight cut -- exactly what --n-expert-used already does, under a flag
  claiming to consult residency. validate() now rejects it, as it already
  did for --prefetch, and the app gates the setting on the same condition.

- Fix experts_routed. It was incremented inside apply_drop, so it counted
  what the policy examined rather than what the router selected: layers
  before the terminal weight node is learned, and every un-armed phase,
  were missing from the denominator. The reported drop rate was a fraction
  of the wrong thing.

- Re-learn instead of re-betting. If the node learned as terminal does not
  arrive, the deferral now also forgets it, so the next graph loads at the
  topk node while it re-learns. Deferring again on a stale guess would
  repeat the fault every token against a graph that had moved.

- Point the gates at a real cache. G8a/G8b ran with the cache off, where
  the shared-slot path has no reserved-but-uncommitted memory -- so the id
  repointing, which is the design's whole safety argument, was never
  exercised. They now run against a constantly-evicting budget. Adds G8a'
  (asserts routings were examined and none dropped, so an inert-threshold
  flake fails legibly) and G8c (at top-k 1 dropping is a proven no-op,
  pinning both the top-expert guarantee and the threshold being taken
  against the effective top-k).

Docs: three metrics change meaning under dropping and none of them said
so. A dropped routing is a miss that is never looked up, so cache_hit_pct
rises without the cache serving more, and token/layer_demand measure what
was staged rather than routed -- documented in telemetry.md, pressure.md
(size the cache with dropping off, then turn it on) and metrics.h.
prefetch.md's "cannot change output" is scoped: under dropping a correct
guess un-drops an expert. limitations.md gains the non-reproducibility
entry, benchmark-method.md the axis plus a warning that reversing the run
order cannot distinguish a moved drop rate from a contaminated cell, and
architecture.md/runtime.h no longer claim unconditional determinism.
Fixes two anchors the README rename broke, and a changelog sentence that
quoted the equal-I/O row while drawing the equal-quality conclusion.

App: Drop cold experts defaults to 75%. The default is a product decision
taken on the maintainer's device; no benchmark for it is published here,
and docs/expert-dropping.md says that plainly instead of implying a
measured figure. The CLI stays off by default -- the byte-identity gates
need a deterministic default.
This commit is contained in:
Helldez 2026-07-22 15:08:49 +02:00 • committed by GitHub
parent 719478908f
commit 45a90a2df5
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
31 changed files with 981 additions and 39 deletions

View file

@ -21,7 +21,8 @@ for the idea the project is built on.
| [telemetry.md](telemetry.md) | The `BMOE_*` line protocol and CSV schema — the integration contract. |
| [session.md](session.md) | Session lifecycle, KV prefix reuse, cancellation. |
| [cache-sizing.md](cache-sizing.md) | `--cache-mb auto`, the cache ceiling, and dense warm-up. |
| [prefetch.md](prefetch.md) | `--prefetch K`: the design and why it cannot change output. |
| [prefetch.md](prefetch.md) | `--prefetch K`: the design and why it cannot change output (with the lossy knobs off). |
| [expert-dropping.md](expert-dropping.md) | `--drop-cold-experts F`: spending quality only where it buys a flash read, and why it is the one setting whose output is not reproducible. |
| [android-memory.md](android-memory.md) | What reclaims the engine's memory on a phone, which levers exist (almost none), and why the cache hit rate is what the kernel judges you by. |
| [pressure.md](pressure.md) | Cache policy under memory pressure: why an unaffordable budget starts a reclaim war, why the adaptive governor was retired, and what the fixed `--cache-mb` / `--dense-weights` levers do. |

View file

@ -45,8 +45,11 @@ Streaming experts serially needs three things from the inference engine. All thr
already public in llama.cpp:
1. **A hook at routing time.** `llama_context_params.cb_eval` is called for every graph
node. We ask for only the routing nodes (`ffn_moe_topk-<il>`); ggml computes and
node. We ask for the routing nodes (`ffn_moe_topk-<il>`); ggml computes and
synchronizes each alone, then calls us back with the selected expert ids materialized.
The route trace and [cache-aware dropping](expert-dropping.md) additionally ask for each
layer's `ffn_moe_weights*-<il>` chain — and dropping is the one path that *writes into* a
graph tensor's contents rather than only rebinding `->data`. See [seam.md](seam.md).
2. **The expert tensor pointers.** During a one-token warm-up we scan each graph node's
sources for tensors named `blk.<il>.ffn_{gate,up,down}_exps.weight` and record the
live `ggml_tensor*`. We then rebind their `->data`.
@ -89,4 +92,6 @@ The composition root is `Session` (core/src/engine/session.cpp):
so the gates and the interactive session share the same code path.
Greedy sampling makes the output a deterministic function of the graph — the property the
[byte-identity gates](../tests/moe_gates.cpp) assert.
[byte-identity gates](../tests/moe_gates.cpp) assert. That holds with the lossy knobs off. Under
[`--drop-cold-experts`](expert-dropping.md) the hook edits routing weights from live cache state,
which is not in the graph, so output becomes a function of the graph *and* the run's history.

View file

@ -47,6 +47,7 @@ Vary one axis at a time:
| threads (-t) | 2, 4, 8 | U-shape, 4 optimal, 8 regresses |
| overlap | off, on | net gain **only over a warm cache** (hides residual flash wait behind compute); a net loss on a cold cache-0 stream, where I/O dwarfs compute |
| n-expert-used | default, 6 | fewer active experts cut compute + I/O ~linearly (8→6 ≈ −25%), changes the output |
| drop-cold-experts | off, 0.75, 1.0 | the second lossy axis, and the only **non-deterministic** one: what is skipped depends on cache state, so cells are noisier and the drop rate must be reported with the tok/s. Needs the cache on |
| dense-weights | warm, anon | decisive well past RAM, near-neutral near it: on gpt-oss (5.2× RAM) `anon` drops majflt/token from the hundreds to **6–10** and compute with it; on Qwen (1.64× RAM) there is little dense-fault pressure to remove. Watch `majflt/token`, not just tok/s |
When sweeping `--n-expert-used`, run it as a **matched A/B against the model's own default**
@ -125,6 +126,16 @@ So:
Re-run such a cell; do not publish it. And sanity-check any matrix by **reversing the run order** —
cells that move were measuring device state.
**The reversal check does not work under `--drop-cold-experts`.** There a cell can move because the
*drop rate* moved — the policy reads live cache state, so the same command legitimately discards a
different number of experts on a different run. That is the feature working, not the device
contaminating the cell, and the two tells above cannot tell them apart. Always record
`experts_dropped`/`experts_routed` (or the `moe-drop:` line) next to the tok/s: a dropping cell
without its drop rate is uninterpretable, because the flag fixes a threshold and not a rate. Note
also that a dropping run pays the same extra per-MoE-layer barriers a route-traced run does, so an
A/B against `--n-expert-used` is not overhead-matched — see
[expert-dropping.md](expert-dropping.md).
### Caveats
- **Thermal.** Sustained decode throttles. Warm up, then measure a steady window; discard

172
docs/expert-dropping.md Normal file
View file

@ -0,0 +1,172 @@
# Cache-aware expert dropping
`--drop-cold-experts F` skips a routed expert when it is **not in the cache** *and* the router
weighted it below `F × (1 / top-k)` — that is, below `F` of the uniform share each of the `k`
selected experts would get if the router split its mass evenly. Off by default.
It is the second lossy knob in the engine, after
[turbo top-k](../README.md#turbo-top-k--the-measured-lossy-option), and it exists because the first one
spends quality in a place it does not have to.
## Why cache state belongs in the decision
`--n-expert-used k` drops the routing's tail unconditionally: slot 7 and slot 8 go, whether or not
they were already sitting in RAM. But an expert that is already resident costs **no flash read** —
and on a streamed decode, flash reads are what the token is waiting for
([decode is I/O-bound](benchmarks.md)). Dropping a resident expert pays quality for nothing.
Turn that around and the policy writes itself: **spend quality only where it buys I/O**. Keep every
resident expert however small its weight; consider dropping only the ones that would cost a read,
and only when the router says they barely matter.
## What it costs and what it buys
Replayed over the committed route traces (`docs/bench-data/2026-07-15-route-trace/`), decode phase,
threshold at the uniform share (`F = 1.0`):
| policy | flash reads avoided | router weight discarded |
|---|---|---|
| `--drop-cold-experts 1.0` | **66%** | **9.5%** |
| `--n-expert-used 5` | 23% | 10.6% |
| `--n-expert-used 3` | 59% | 36.8% |
(Qwen3-30B-A3B at k=6; Gemma-4-26B-A4B is within a point and a half on both columns: 67.4% / 8.2%. On gpt-oss-120b at k=2 the
policy matches `--n-expert-used 1`'s read saving while discarding 25% of the weight mass instead of
42%.)
At a comparable quality cost the cache-aware policy avoids roughly **three times** the reads. The
reason is visible in the third column of the trace: about 80% of decode routings are cache hits, and
the policy leaves every one of them alone.
`F` is a curve, not a switch. At `F = 0.75` the same model trades 4.4% of the weight mass for 37% of
the reads — still better than `--n-expert-used 5` on **both** axes.
These are replay numbers and an **upper bound**: skipping a read changes what the cache holds later,
so the real hit pattern drifts from the recorded one. The on-device A/B is what settles it.
## Two properties worth knowing
**A routing is never emptied.** The largest weight in a routing is always at least the uniform
share, so at `F ≤ 1.0` the top expert can never fall below the threshold. `validate()` rejects
`F > 1.0` for that reason, and the implementation additionally pins the top-weighted expert, so the
guarantee does not rest on the bound alone.
**It requires the expert cache.** With `--cache-mb 0` every expert reads as a miss, so the policy
would stop being cache-aware and become an unconditional weight cut — which is what
`--n-expert-used` already does, without claiming to consult residency. `validate()` rejects the
combination, the same way it rejects `--prefetch` without a cache.
**It changes what `--prefetch` means.** Speculation is normally output-neutral by construction. Here
residency is an *input* to the policy, so a correct guess un-drops an expert that would otherwise
have been discarded: prefetch depth becomes an output-affecting setting. The decision point also
settles pending speculation a few nodes after it was issued, which shortens the overlap window the
prefetch exists for — treat the two as interacting, not composable.
**Prefill is excluded by default.** With a cold cache almost every expert is a miss, and the same
threshold discards ~42% of the weight mass instead of ~9%. Prefill is compute-bound anyway, so there
is little to win. `--drop-in-prefill` arms it for experiments.
## The output is no longer reproducible
This is the real novelty, and the reason the flag is off by default and named the way it is.
`--n-expert-used` is lossy but **deterministic**: same prompt, same config, same tokens. Dropping is
lossy and **state-dependent** — what gets discarded depends on what the cache happened to hold,
which depends on everything decoded before it. The same prompt can produce different text across
runs, and a benchmark cell is noisier because the drop rate itself varies.
The greedy byte-identity gates therefore do not cover the policy's output, and cannot: there is
nothing stable to compare against. They cover the machinery instead (see below).
## How it is implemented
The decision needs the **final** router weights, and those are produced several graph nodes after
the topk node where the streamer normally loads. So with the policy armed, `load_layer()` is
postponed from the topk node to the terminal node of the layer's weight chain — the last node before
the expert matmul consumes either the ids or the weights.
Which node is terminal depends on the model's gating (`_norm`, `_softmax`, `_scaled`, or none), so
the hook **learns** it from the graph instead of carrying a per-architecture table: the first graph
of a run records the chain, and dropping starts from the second. A layer whose shape has not been
seen yet simply loads at its topk node, undropped. That costs a run its first token's dropping and
nothing else, and it keeps [hard rule 4](../CLAUDE.md) — no model-specific constants in the
streaming path.
At the decision point two edits happen, both before anything reads them:
1. the dropped slot's **weight is zeroed**, and with `drop_renorm` (default on) the survivors are
scaled so the routing keeps its original total mass;
2. the dropped slot's **id is repointed** at the routing's top-weighted expert.
The second edit is not cosmetic. An expert the engine declines to read may sit in a
reserved-but-uncommitted slot, and `mul_mat_id` would still touch it. Pointing the slot at an expert
that is certainly resident makes the kernel read valid memory and multiply it by exactly zero. It
costs a duplicate matmul — the right trade on a decode bound by flash rather than arithmetic.
Renormalisation matters more than it looks: without it the layer's expert output is systematically
scaled down by the discarded mass, a perturbation of the residual stream the model never sees in
training. `--drop-no-renorm` exists to A/B that claim.
Cost of the extra barriers: the policy asks for each layer's weight nodes, a handful more
synchronisation points per MoE layer on tensors of a few floats. The same asks a route trace makes.
## Measuring it
The engine reports what the policy actually did, which the flag alone cannot tell you — the
threshold is fixed, the drop rate is not:
```
moe-drop: <dropped>/<routed> routed experts dropped (<pct>%), threshold <F> x uniform
```
The route trace gains a `dropped` column: `weight` and `residency` stay as the **router** produced
them, `dropped` records what the policy then did, and `expert_bytes` is 0 for a dropped routing
because it costs no read. That is enough to replay a real run against the offline model and check
whether the upper bound held. See [telemetry.md](telemetry.md).
## Gates
`bmoe_moe_gates` covers the machinery, not the policy's output:
- **G8a** — with a threshold below any weight the router can produce, nothing is dropped and the
output is **byte-identical** to the undropped stream. This proves the deferral and the learned
terminal node are transparent, separating "the plumbing is correct" from "the policy is lossy" —
a regression in the first would otherwise hide behind the expected difference. **G8a'** asserts
the count separately (`experts_routed > 0`, `experts_dropped == 0`), so "a weight happened to fall
under the threshold" fails legibly instead of as a mysterious byte mismatch.
- **G8b** — at full strength against a cache small enough to be evicting constantly, so dropped
experts really do land on slots the cache has released. Generation still completes: the id
repointing means no matmul ever reads reserved-but-uncommitted memory. (The gates deliberately do
*not* run this with the cache off — there the shared-slot path has no uncommitted memory, so the
safety property the repointing exists for would go untested.)
- **G8c** — forcing top-k to 1 makes every routed expert the top one, so dropping must be a no-op at
any threshold and the output must match the undropped k=1 run byte for byte. This pins both the
top-expert guarantee and the fact that the threshold is taken against the **effective** top-k
discovered at runtime — a hardcoded width would not survive the override.
## Defaults, and where the numbers do and do not come from
The **CLI defaults it off**, and will keep doing so: the byte-identity gates need a deterministic
default, and an instrument should not quietly change the thing it measures.
The **app ships it at 75%** — under **Speed / quality → Drop cold experts**, with rungs 50 / 75 /
100 as percentages of the uniform share. It is disabled there in mmap mode and with the cache off,
the same two conditions `validate()` enforces.
That default is a product decision taken on the maintainer's own device measurement. **It is not
backed by a published benchmark in this repository**, and the tables in the README deliberately
carry no rows for it — they are a deterministic protocol and this knob is not deterministic. Nothing
here should be read as "75% is worth X%"; the honest claim is narrower: the replay above says the
shape of the trade is favourable, and the default was chosen after checking it on hardware.
What is still owed before this is recommended beyond that:
- a published decode A/B against `--n-expert-used` at matched tok/s, with the device state recorded
the way [benchmark-method.md](benchmark-method.md) requires;
- a quality comparison at that matched speed — the whole thesis is that this knob buys the same
throughput for less damage, and only a side-by-side can support it;
- a re-run of the replay against a real traced run with the `dropped` column, to see how far the
static upper bound overstated the win.
The [`layer-lfu` entry in the roadmap](roadmap.md) is the standing reminder for why the third one
matters: it simulated exactly as predicted and was ~30% slower in reality.

View file

@ -18,6 +18,11 @@ serial path, and only a single ~25-line hook (with an explicit sunset) for the o
## Limitations
- **One setting makes output non-reproducible.** Every other knob is deterministic given a
configuration: `--n-expert-used` changes the output, but changes it the same way on every run.
[`--drop-cold-experts`](expert-dropping.md) decides per routing from live cache state, so the
same prompt and the same flags can decode differently run to run, and the byte-identity gates
cannot cover its output — only its machinery. Off by default in the CLI.
- **n=1 only.** The expert sparsity exists only for single-token decode, so streaming is
incompatible with speculative decoding or batching. Prefill streams the union of the
prompt's routed experts (still far below the full bank, but larger than one token's).

View file

@ -32,7 +32,10 @@ Ordering is guaranteed by ggml's eval-callback loop: the node we mark is compute
buffers until this layer's matmul has synchronized. Correct on any backend.
The result is **lossless**: byte-identical to running with every expert resident, asserted
by the gates.
by the gates. That is the streaming path itself; two opt-in knobs deliberately trade output for
speed on top of it — `--n-expert-used` (fewer experts per token) and
[`--drop-cold-experts`](expert-dropping.md) (skip an expert that would cost a read and was barely
weighted). Both are off unless asked for, which is what keeps the sentence above true by default.
## Residency modes

View file

@ -32,7 +32,11 @@ captures most of the benefit. Prefetch requires the LRU cache to be on — eithe
## How it stays correct and out of the way
The speculative path never delays real work and never changes output:
The speculative path never delays real work and never changes output — with the lossy knobs off.
(Under [`--drop-cold-experts`](expert-dropping.md) residency is an *input* to the routing policy,
so a correct guess un-drops an expert that would otherwise have been discarded. Prefetch depth
becomes output-affecting there; everything below still holds for the bytes themselves.)
- **Same bytes.** A speculative read is the *identical* read a real miss would issue — same file
offset, same destination buffer (`lbuf_[p][il] + e*slice`). A prefetched expert is therefore

View file

@ -30,7 +30,10 @@ is:
## Why a budget cannot be a constant
The expert cache is the one lever that trades RAM for flash reads, so the temptation is to set it as
The expert cache is the one lever that trades RAM for flash reads (
[`--drop-cold-experts`](expert-dropping.md) is the other kind of trade — quality for flash reads —
and the two interact: a squeezed cache raises the miss rate, which raises the drop rate, so memory
pressure degrades output quality there instead of only throughput). The temptation is to set it as
large as the device seems to allow. On a phone that is the wrong shape of decision, for three
reasons that are measured rather than argued:
@ -115,6 +118,11 @@ not a floor), `layer_demand_MiB` (the mechanical floor), `cache_budget_MiB` (the
effect), `cache_hit_pct`. Per token, `dense_resident_frac` says whether the dense set is holding in
RAM (the live signal now that the cache-residency governor sensor is gone).
This sizing procedure assumes dropping is off. With
[`--drop-cold-experts`](expert-dropping.md) on, dropped routings are misses that never reach the
cache, so `cache_hit_pct` reads high and `token_demand_MiB` reads low for the same budget — size
the cache first, then turn dropping on.
Reading `cache_hit_pct` against `token_demand_MiB` is how you tell whether a fixed `--cache-mb N` is
earning its RAM: a budget near or below one token's demand holds no history between tokens and its
hits are only inter-token correlation; well above it, a high hit rate means real reuse.

View file

@ -113,6 +113,19 @@ routed) are supported; other `build_moe_ffn` models are one recipe row each. The
frontier is architectures whose routing node is not the shared `ffn_moe_topk` — custom gating,
which the capture/stream hook would need to learn. See [adding-a-model.md](adding-a-model.md).
## Skipping reads the router barely wants — built, unmeasured
`--drop-cold-experts` ([expert-dropping.md](expert-dropping.md)) is the first lever that treats
quality and I/O as a *joint* budget rather than two separate knobs: an expert already in the cache
runs however small its weight, and only a routing that would cost a flash read can be dropped. On
the recorded traces that is worth ~3× the reads of turbo top-k for a comparable weight cost, which
is the strongest offline case any remaining lever has shown.
What it does **not** have is a device measurement, and the previous entry on this page is the reason
that matters: `layer-lfu` simulated well and was ~30% slower in reality. The open questions are the
device A/B against turbo top-k at matched throughput, the quality comparison at that speed, and how
far the static replay overstated the win once dropping starts changing what the cache holds.
## Expert quantization on the fly
Storing streamed experts at a lower precision than the resident parts to cut read volume, if it

View file

@ -24,10 +24,22 @@ come from the arch's recipe — `ffn_{gate,up,down}_exps` for the split layout,
throughout — capture observes, it does not isolate. `ggml_tensor` is a public struct, so
reading `->name`, `->ne`, `->nb` and writing `->data` is public API surface.
**Stream phase** (real generation). We return true only for `ffn_moe_topk-<il>`. The
**Stream phase** (real generation). We return true for `ffn_moe_topk-<il>`. The
non-ask callback then hands us that node with the selected expert ids materialized; we
gather them (stride-aware) and trigger the slice reads.
Two optional jobs ask for more: the route trace and
[cache-aware dropping](expert-dropping.md) also want each layer's `ffn_moe_weights*-<il>` chain,
which is another barrier per node but no new kind of access — same public struct, same read of
`->data`.
Dropping does go one step further, and it is the only place the engine **writes into** a graph
tensor's contents rather than repointing `->data` at its own buffer: at the terminal node of the
weight chain it zeroes a dropped slot's weight and repoints that slot's expert id. Both tensors are
scratch the graph produced and has not yet consumed, so this alters the values flowing through the
run — deliberately, that is what the lossy policy *is* — and never llama.cpp's own state, its
weights, or its control flow. It stays inside the same callback contract; nothing is patched.
## 2. gguf offsets
`gguf_init_from_file(..., no_alloc=true)` + `gguf_get_data_offset` +

View file

@ -51,6 +51,10 @@ BMOE_PROGRESS {"step":<int>,"steps":<int>,"wall_ms":<float>,"io_ms":<float>,
can be a large share of the token; at steady state it is near zero. Surfacing it stops the "all
compute" reading on warm-up tokens where the real cost is cache churn, not matmul.
- `cache_hit_pct` is the cumulative cache hit rate, or `-1` when no cache is used.
**Under [`--drop-cold-experts`](expert-dropping.md) read it with care:** a dropped routing is a
miss that is never looked up, so it leaves both sides of the ratio and the reported hit rate
rises without the cache having served anything more. Compare runs at the same drop rate, or read
`experts_dropped` next to it.
- `majflt` / `cpu_ms` **decompose the `compute_ms` residual** — the whole point being that "compute"
above is a catch-all that silently absorbs page faults and scheduler stalls, not just matmul.
They are measured directly around `llama_decode` (no submodule patch needed): `majflt` is the
@ -101,6 +105,16 @@ moe-prefetch: <mib> MiB speculative, <useful>/<prefetched> experts useful (<pct>
`<prefetched>` the experts fully read ahead, and `<useful>` how many of those a later routing
actually hit. See [prefetch.md](prefetch.md).
With `--drop-cold-experts F` a `moe-drop:` line is added:
```
moe-drop: <dropped>/<routed> routed experts dropped (<pct>%), threshold <F> x uniform
```
The flag fixes a *threshold*, not a rate: how much is actually discarded depends on what the cache
held, so this line — not the flag — is what a run traded. See
[expert-dropping.md](expert-dropping.md).
Under `--overlap` the `moe-stream:` line additionally reports `stall_s/tok=<s>` — the mean
wall time per token that compute threads waited for expert reads to complete. It is `0` in
serial mode (where the read wait is already folded into decode time).
@ -130,7 +144,10 @@ sampled dense-weight residency, `-1` when unmeasured. All are additive: older CS
so consumers must read by column NAME (from the header row) and treat any as optional. The `# summary`
line likewise gains `stall_s/tok=<s>`, `mgmt_s/tok=<s>`, `majflt/tok=<f>`, `cpu_s/tok=<s>`,
`token_demand_MiB=<f>` (the expert bytes one token routes, measured — where cache hits start, NOT a
floor to defend; see [pressure.md](pressure.md)) and `layer_demand_MiB=<f>` (the widest layer's routed
floor to defend; see [pressure.md](pressure.md)), `experts_routed=<n>` / `experts_dropped=<n>` (what
[cache-aware dropping](expert-dropping.md) actually discarded during generation — the flag sets a
threshold, not a rate, so this is the only record of the trade a run made) and
`layer_demand_MiB=<f>` (the widest layer's routed
bytes: the mechanical floor the cache must be able to stage); see the `io_ms` note above for how the
read-time columns are reinterpreted under overlap.
@ -161,6 +178,10 @@ per routed expert. **A traced run is not a benchmark run** — the numbers in th
traced run are slower than the real thing, and `mgmt_ms` in particular shifts, because settling
speculative prefetch moves outside the window that times it.
Columns are **append-only** within `v1`, like the metrics CSV: `dropped` was added after
`expert_bytes`, so consumers must read by column NAME and treat any column as optional rather than
indexing by position.
The file is long format: a `#` preamble carrying the run's static facts, then one row per routed
expert. Conceptually it is a matrix — rows are steps, columns are layers — and a **cell** is the
`n_expert_used` rows sharing `(turn, phase, step, layer)`.
@ -169,7 +190,7 @@ expert. Conceptually it is a matrix — rows are steps, columns are layers — a
# route_trace v1
# model=<path> arch=<string> n_layer=<int> n_expert=<int> n_expert_used=<int>
# layer=<int> expert_bytes=<int> dense_bytes=<int> (one per layer)
turn,phase,step,layer,slot,expert,weight,residency,expert_bytes
turn,phase,step,layer,slot,expert,weight,residency,expert_bytes,dropped
```
| column | meaning |
@ -183,6 +204,7 @@ turn,phase,step,layer,slot,expert,weight,residency,expert_bytes
| `weight` | the final applied routing weight, after whatever softmax/normalise/scale the architecture uses. `nan` when the graph exposed no weight node — "unknown", never `0`. |
| `residency` | `0` = miss (this routing reads from flash), `1` = hit, `2` = hit on a speculative prefetch's first touch. |
| `expert_bytes` | flash bytes this routing reads; `0` unless `residency=0`. |
| `dropped` | `1` when [cache-aware dropping](expert-dropping.md) discarded this routing — a miss weighted below the threshold, never read, weight zeroed. Always `0` with `--drop-cold-experts` off. |
`(turn, phase, step, layer, slot)` is unique. Two asymmetries are deliberate:
@ -195,6 +217,11 @@ turn,phase,step,layer,slot,expert,weight,residency,expert_bytes
streamed, so there is nothing to measure per step: `dense_bytes` is what a cold layer costs to
page in, stated once. Per-layer *I/O time* is absent for the same kind of reason — under
`--overlap` reads complete asynchronously, so any per-layer timing would be fiction.
- **`weight` and `residency` describe the router; `dropped` describes the policy.** When dropping is
on, a discarded routing keeps the weight the router gave it and the residency it faced — the trace
records the routing that was *chosen* — while `expert_bytes` falls to `0`, because a dropped
expert is never read. Summing `expert_bytes` therefore still measures real flash traffic, and
`dropped` is what explains the gap against `residency==0`.
**The last layer has only one prefill step, and that is real.** Before the final layer's FFN,
llama.cpp gathers only the tokens whose logits were asked for (`inp_out_ids`; see `il == n_layer