mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#94)
* feat(moe): --drop-cold-experts, spend quality only where it buys I/O Turbo top-k drops the tail of a routing whether or not those experts were already in RAM. A resident expert costs no flash read, so that trade pays quality for nothing on the ~80% of decode routings that are cache hits. This adds the cache-aware version: skip a routed expert only when it is a cache MISS and the router weighted it below frac x (1/top-k). Replayed over the committed route traces at frac 1.0, decode phase, that avoids 66% of flash reads for 9.5% of the router's weight mass, against 59%/37% for --n-expert-used 3 — about 3x the reads avoided at a comparable cost. The threshold is a curve, not a switch: 0.75 trades 4.4% of the mass for 37% of the reads, better than --n-expert-used 5 on both axes. Implementation. The decision needs the FINAL router weights, which arrive several nodes after the topk where the streamer normally loads, so with the policy armed load_layer() is deferred to the terminal node of the layer's weight chain. Which node that is depends on the model's gating, so the hook learns it from the graph rather than carrying an architecture table; until it is known a layer loads at its topk node undropped. A dropped slot has its weight zeroed and its expert id repointed at the routing's top-weighted expert: an expert we decline to read may sit in reserved-but-uncommitted VM and mul_mat_id would still touch it, so the kernel is given memory that is certainly resident and multiplies it by exactly zero. Survivors are rescaled by default, since a systematically shrunk expert output perturbs the residual stream more than the missing contribution does. Prefill is excluded by default (cold cache, ~4x the weight mass discarded, and compute-bound anyway). The largest weight in a routing is always at least the uniform share, so frac <= 1 can never empty a layer; validate() enforces the bound and the top expert is pinned regardless. Gates: G8a proves the deferral and the learned terminal node are transparent (a threshold below any producible weight leaves the output byte-identical), G8b that full strength with the cache off never reaches an unloaded expert. Unlike every other knob this one is state-dependent: what gets dropped depends on what the cache held, so output is not reproducible across runs. Off by default, not in the app's settings, and NOT yet measured on device — the numbers above are a static replay and an upper bound. docs/expert- dropping.md states what is owed before it is recommended anywhere. * feat(app): expose cache-aware expert dropping in Settings Speed / quality -> Drop cold experts, as a percentage of the uniform share (off / 50 / 75 / 100). The engine takes a fraction; the app stores integer rungs, so the setting divides by 100 on the way to the flag. Disabled in mmap mode: the policy asks the expert source what is resident, and there is no expert source without the streamer. Included in the session signature, so changing it reopens the session rather than being ignored by a process already loaded. Off by default. This exists so the A/B can be run where the engine actually ships -- through the app, not a pushed CLI binary. * fix(moe): require the cache for dropping, and correct what it reports Review of the first two commits found the policy could be armed in a configuration where it is not cache-aware at all, and that two of the numbers it reports were wrong. - Require the LRU cache. With --cache-mb 0 query_residency answers all-miss, so the policy silently degenerated into an unconditional weight cut -- exactly what --n-expert-used already does, under a flag claiming to consult residency. validate() now rejects it, as it already did for --prefetch, and the app gates the setting on the same condition. - Fix experts_routed. It was incremented inside apply_drop, so it counted what the policy examined rather than what the router selected: layers before the terminal weight node is learned, and every un-armed phase, were missing from the denominator. The reported drop rate was a fraction of the wrong thing. - Re-learn instead of re-betting. If the node learned as terminal does not arrive, the deferral now also forgets it, so the next graph loads at the topk node while it re-learns. Deferring again on a stale guess would repeat the fault every token against a graph that had moved. - Point the gates at a real cache. G8a/G8b ran with the cache off, where the shared-slot path has no reserved-but-uncommitted memory -- so the id repointing, which is the design's whole safety argument, was never exercised. They now run against a constantly-evicting budget. Adds G8a' (asserts routings were examined and none dropped, so an inert-threshold flake fails legibly) and G8c (at top-k 1 dropping is a proven no-op, pinning both the top-expert guarantee and the threshold being taken against the effective top-k). Docs: three metrics change meaning under dropping and none of them said so. A dropped routing is a miss that is never looked up, so cache_hit_pct rises without the cache serving more, and token/layer_demand measure what was staged rather than routed -- documented in telemetry.md, pressure.md (size the cache with dropping off, then turn it on) and metrics.h. prefetch.md's "cannot change output" is scoped: under dropping a correct guess un-drops an expert. limitations.md gains the non-reproducibility entry, benchmark-method.md the axis plus a warning that reversing the run order cannot distinguish a moved drop rate from a contaminated cell, and architecture.md/runtime.h no longer claim unconditional determinism. Fixes two anchors the README rename broke, and a changelog sentence that quoted the equal-I/O row while drawing the equal-quality conclusion. App: Drop cold experts defaults to 75%. The default is a product decision taken on the maintainer's device; no benchmark for it is published here, and docs/expert-dropping.md says that plainly instead of implying a measured figure. The CLI stays off by default -- the byte-identity gates need a deterministic default.
This commit is contained in:
parent
719478908f
commit
45a90a2df5
31 changed files with 981 additions and 39 deletions
53
CHANGELOG.md
53
CHANGELOG.md
|
|
@ -4,6 +4,59 @@ All notable changes to this project are documented here. The format follows
|
||||||
[Keep a Changelog](https://keepachangelog.com/), and the project aims to follow
|
[Keep a Changelog](https://keepachangelog.com/), and the project aims to follow
|
||||||
Semantic Versioning.
|
Semantic Versioning.
|
||||||
|
|
||||||
|
## [Unreleased]
|
||||||
|
|
||||||
|
### Added
|
||||||
|
- **`--drop-cold-experts F` — cache-aware expert dropping.** Skips a
|
||||||
|
routed expert when it is a cache **miss** *and* the router weighted it below `F × (1/top-k)`. An
|
||||||
|
expert already resident costs no flash read, so it always runs however small its weight: quality
|
||||||
|
is spent only where it buys I/O. Replayed over the committed route traces at `F = 1.0`, decode
|
||||||
|
phase, this avoids **66% of flash reads for 9.5% of the router's weight mass**, where
|
||||||
|
`--n-expert-used 5` avoids 23% for a comparable 10.6% — roughly 3x the reads at the same quality
|
||||||
|
cost. (At equal *reads* instead, `--n-expert-used 3` avoids 59% but discards 37% of the mass.)
|
||||||
|
`--drop-no-renorm` and `--drop-in-prefill` are the A/B switches. Requires the LRU cache:
|
||||||
|
`validate()` rejects it with `--cache-mb 0`, where every expert reads as a miss and the policy
|
||||||
|
would silently degenerate into an unconditional weight cut.
|
||||||
|
Unlike turbo top-k the output is **not reproducible** — what gets dropped depends on what the
|
||||||
|
cache held — so it carries no rows in the README benchmark tables, which are a deterministic
|
||||||
|
protocol. See [docs/expert-dropping.md](docs/expert-dropping.md).
|
||||||
|
- `scripts/route-drop-replay.py`: the offline model the numbers above come from, including the
|
||||||
|
static-`k` baseline replayed on the same rows so the two policies are comparable at equal I/O.
|
||||||
|
- Route trace gains a `dropped` column, and the metrics summary `experts_routed` /
|
||||||
|
`experts_dropped` — the flag fixes a threshold, not a drop rate, so only these say what a run
|
||||||
|
actually traded. New CLI summary line `moe-drop:`.
|
||||||
|
- Example app: **Speed / quality → Drop cold experts** (off / 50% / 75% / 100% of the uniform
|
||||||
|
share), **defaulting to 75%**, disabled in mmap mode and with the cache off. The default is a
|
||||||
|
product decision taken on the maintainer's device; it is not backed by a benchmark published here,
|
||||||
|
and `docs/expert-dropping.md` says so plainly rather than implying a measured figure. The CLI
|
||||||
|
keeps defaulting to off, since the byte-identity gates need a deterministic default.
|
||||||
|
- Gates **G8a/G8a'/G8b/G8c**: a threshold below any producible weight leaves the output
|
||||||
|
byte-identical to the undropped stream (the deferred load and the learned terminal weight node are
|
||||||
|
transparent) and is asserted to have examined routings while dropping none; at full strength
|
||||||
|
against a constantly-evicting cache generation still completes, so no matmul reaches a
|
||||||
|
reserved-but-uncommitted slot; and at `--n-expert-used 1` dropping is a proven no-op, which pins
|
||||||
|
both the top-expert guarantee and the threshold being taken against the *effective* top-k.
|
||||||
|
|
||||||
|
### Changed
|
||||||
|
- With the policy armed, `load_layer()` moves from the topk node to the terminal node of the layer's
|
||||||
|
weight chain — the decision needs the final router weights. Which node that is depends on the
|
||||||
|
model's gating, so the hook **learns** it from the graph rather than carrying an architecture
|
||||||
|
table; until it is known a layer loads at its topk node undropped, exactly as before. No behaviour
|
||||||
|
changes when `--drop-cold-experts` is off.
|
||||||
|
- README no longer calls turbo top-k "the one lossy knob" — it is now the *measured* one.
|
||||||
|
- Docs that assumed a deterministic engine are scoped: `prefetch.md` ("cannot change output" holds
|
||||||
|
only with the lossy knobs off — under dropping, a correct guess un-drops an expert),
|
||||||
|
`moe-streaming.md`, `architecture.md`, `limitations.md` (new entry for non-reproducibility) and
|
||||||
|
`runtime.h`'s contract.
|
||||||
|
- `cache_hit_pct`, `token_demand_MiB` and `layer_demand_MiB` shift meaning under dropping — a
|
||||||
|
dropped routing is a miss that is never looked up, so the hit rate rises without the cache serving
|
||||||
|
more, and the demand figures measure what was *staged* rather than routed. Documented in
|
||||||
|
`telemetry.md`, `pressure.md` (which tells you to size the cache first, dropping off) and
|
||||||
|
`metrics.h`; `benchmark-method.md` gains the axis plus a warning that its reverse-the-run-order
|
||||||
|
check cannot distinguish a moved drop rate from a contaminated cell.
|
||||||
|
- `scripts/route-analyze.py` reports when a trace was recorded with dropping on, so its
|
||||||
|
working-set figures are not misread as flash traffic.
|
||||||
|
|
||||||
## [0.14.0] - 2026-07-21
|
## [0.14.0] - 2026-07-21
|
||||||
|
|
||||||
### Added
|
### Added
|
||||||
|
|
|
||||||
29
README.md
29
README.md
|
|
@ -41,7 +41,7 @@ Highlights:
|
||||||
|
|
||||||
- **gpt-oss-120b (Q4_K_M), ~5× device RAM**: **1.3 tok/s** at the model's own routing width against
|
- **gpt-oss-120b (Q4_K_M), ~5× device RAM**: **1.3 tok/s** at the model's own routing width against
|
||||||
0.09 tok/s for the same file loaded the ordinary way (mmap), a **14×** difference at matched settings.
|
0.09 tok/s for the same file loaded the ordinary way (mmap), a **14×** difference at matched settings.
|
||||||
**2.2 tok/s** with the one lossy knob on (fewer experts).
|
**2.2 tok/s** with the measured lossy knob on (fewer experts).
|
||||||
- **Lossless on models past RAM**: Qwen3-30B-A3B (Q4_K_M, 18.5 GB) up to **5.2 tok/s**,
|
- **Lossless on models past RAM**: Qwen3-30B-A3B (Q4_K_M, 18.5 GB) up to **5.2 tok/s**,
|
||||||
Qwen3.6-35B-A3B (Q4_K_M, 22.3 GB) up to **5.0 tok/s** and Gemma-4-26B-A4B (Q4_K_M, 17.0 GB) up to
|
Qwen3.6-35B-A3B (Q4_K_M, 22.3 GB) up to **5.0 tok/s** and Gemma-4-26B-A4B (Q4_K_M, 17.0 GB) up to
|
||||||
**4.1 tok/s** on the same phone, output identical to the resident model.
|
**4.1 tok/s** on the same phone, output identical to the resident model.
|
||||||
|
|
@ -95,8 +95,15 @@ and a manual copy to the device: steps in the
|
||||||
device.
|
device.
|
||||||
- **I/O–compute overlap** (`--overlap`): hides flash latency behind compute. Byte-identical;
|
- **I/O–compute overlap** (`--overlap`): hides flash latency behind compute. Byte-identical;
|
||||||
needs a small optional add-on to llama.cpp (see [docs/seam.md](docs/seam.md)).
|
needs a small optional add-on to llama.cpp (see [docs/seam.md](docs/seam.md)).
|
||||||
- **Turbo top-k** (`--n-expert-used N`): the one lossy knob. Fewer experts per token, ~+22–24%
|
- **Turbo top-k** (`--n-expert-used N`): the measured lossy knob. Fewer experts per token, ~+22–24%
|
||||||
speed, output quality is yours to judge.
|
speed, output quality is yours to judge.
|
||||||
|
- **Cache-aware expert dropping** (`--drop-cold-experts F`): skips a routed expert only when it would
|
||||||
|
cost a flash read *and* the router barely weighted it, so quality is spent only where it buys I/O.
|
||||||
|
Replayed against recorded traces it avoids ~3× the reads of turbo top-k at a comparable weight
|
||||||
|
cost. It is the one setting whose output is **not reproducible** — what gets skipped depends on
|
||||||
|
what the cache held — so it has no rows in the tables below, which are a deterministic protocol.
|
||||||
|
The app ships it at 75%; the CLI defaults it off. See
|
||||||
|
[docs/expert-dropping.md](docs/expert-dropping.md).
|
||||||
- **Multi-turn sessions and live telemetry**: the model stays loaded across chat turns, and every
|
- **Multi-turn sessions and live telemetry**: the model stays loaded across chat turns, and every
|
||||||
run can emit a per-token breakdown of where the time went.
|
run can emit a per-token breakdown of where the time went.
|
||||||
- **Android demo app** ([`examples/android`](examples/android)): a chat app with a live telemetry
|
- **Android demo app** ([`examples/android`](examples/android)): a chat app with a live telemetry
|
||||||
|
|
@ -127,7 +134,7 @@ phone.
|
||||||
ordinary way (no streaming), which is what the streamed rows are compared against. *k* is how many
|
ordinary way (no streaming), which is what the streamed rows are compared against. *k* is how many
|
||||||
experts each token routes to — for us the number of experts, i.e. `n_expert_used` (set with
|
experts each token routes to — for us the number of experts, i.e. `n_expert_used` (set with
|
||||||
`--n-expert-used`). Each table shows the model's default width and, where measured, a reduced *k*,
|
`--n-expert-used`). Each table shows the model's default width and, where measured, a reduced *k*,
|
||||||
the one lossy setting — see [Turbo top-k — the one lossy option](#turbo-top-k--the-one-lossy-option)
|
the measured lossy setting — see [Turbo top-k — the measured lossy option](#turbo-top-k--the-measured-lossy-option)
|
||||||
below.
|
below.
|
||||||
- **tok/s**: generation speed; higher is better.
|
- **tok/s**: generation speed; higher is better.
|
||||||
- **Flash/token**: data read from storage per generated token; lower means the cache is working.
|
- **Flash/token**: data read from storage per generated token; lower means the cache is working.
|
||||||
|
|
@ -167,7 +174,7 @@ dense weights kept out of the page cache (`--dense-weights anon`) runs it stably
|
||||||
| **streamed, k=6, cache 3000 MiB, 4 lanes, overlap** | **5.8** | 91 MiB | 68% |
|
| **streamed, k=6, cache 3000 MiB, 4 lanes, overlap** | **5.8** | 91 MiB | 68% |
|
||||||
|
|
||||||
All streamed rows use `--overlap --dense-weights anon`. A larger cache is the main lossless lever
|
All streamed rows use `--overlap --dense-weights anon`. A larger cache is the main lossless lever
|
||||||
(cache 3000 is worth +16% over 2000); the k=6 rows are the one lossy option (turbo top-k, below),
|
(cache 3000 is worth +16% over 2000); the k=6 rows are the measured lossy option (turbo top-k, below),
|
||||||
worth a further ~16% by routing to six experts instead of eight. The lossless best here is cache
|
worth a further ~16% by routing to six experts instead of eight. The lossless best here is cache
|
||||||
3000 at the model's own width, **5.0 tok/s** — output byte-identical to the resident model.
|
3000 at the model's own width, **5.0 tok/s** — output byte-identical to the resident model.
|
||||||
|
|
||||||
|
|
@ -206,7 +213,7 @@ Gemma keeps more of itself permanently resident, so the 4000 MiB cache fits only
|
||||||
free at launch; cache 2000 + overlap is the dependable everyday setting on this device. Turbo top-k
|
free at launch; cache 2000 + overlap is the dependable everyday setting on this device. Turbo top-k
|
||||||
(k=6) is the fastest here (+22%) but changes the output.
|
(k=6) is the fastest here (+22%) but changes the output.
|
||||||
|
|
||||||
### Turbo top-k — the one lossy option
|
### Turbo top-k — the measured lossy option
|
||||||
|
|
||||||
Every model here ships a routing width — the number of experts each token uses (8 for the Qwen and
|
Every model here ships a routing width — the number of experts each token uses (8 for the Qwen and
|
||||||
Gemma models, 4 for gpt-oss). Forcing it lower with `--n-expert-used` cuts both compute and flash
|
Gemma models, 4 for gpt-oss). Forcing it lower with `--n-expert-used` cuts both compute and flash
|
||||||
|
|
@ -214,9 +221,15 @@ reads; the `k=6` rows folded into the tables above are that knob, measured A/B a
|
||||||
own width. It is worth **+22–24%** on the Qwen and Gemma models, and takes gpt-oss from 1.3 to
|
own width. It is worth **+22–24%** on the Qwen and Gemma models, and takes gpt-oss from 1.3 to
|
||||||
**2.2 tok/s** (k=2).
|
**2.2 tok/s** (k=2).
|
||||||
|
|
||||||
Everything else in this README changes *how* weights are fetched, never the math. This knob changes
|
Every benchmarked setting other than this one changes *how* weights are fetched, never the math.
|
||||||
*what* the model computes: output differs from the full model and quality can degrade. Judge it on
|
This knob changes *what* the model computes: output differs from the full model and quality can
|
||||||
your own task before relying on it.
|
degrade. Judge it on your own task before relying on it.
|
||||||
|
|
||||||
|
It also spends quality indiscriminately: the tail of the routing goes whether or not those experts
|
||||||
|
were already in RAM, and a resident expert costs no flash read at all.
|
||||||
|
[Cache-aware dropping](docs/expert-dropping.md) is the experimental answer to that — same kind of
|
||||||
|
trade, but only where it buys I/O. It has no measured rows here yet, which is why the tables above
|
||||||
|
are still turbo top-k's.
|
||||||
|
|
||||||
### What to expect in the app
|
### What to expect in the app
|
||||||
|
|
||||||
|
|
|
||||||
20
cli/main.cpp
20
cli/main.cpp
|
|
@ -363,6 +363,12 @@ static void print_usage(const char * argv0) {
|
||||||
" --force-cache allow a cache-mb in the pathological band\n"
|
" --force-cache allow a cache-mb in the pathological band\n"
|
||||||
" --overlap overlap async expert reads with FFN compute (needs the fork)\n"
|
" --overlap overlap async expert reads with FFN compute (needs the fork)\n"
|
||||||
" --prefetch K temporally prefetch the next K layers' experts (needs the cache)\n"
|
" --prefetch K temporally prefetch the next K layers' experts (needs the cache)\n"
|
||||||
|
" --drop-cold-experts F skip a routed expert that is a cache MISS and carries less than\n"
|
||||||
|
" F x (1/top-k) of the routing's weight. F in (0, 1]; 1.0 is the\n"
|
||||||
|
" uniform share and the useful maximum. LOSSY and cache-dependent:\n"
|
||||||
|
" it changes the output, and not reproducibly. Off by default.\n"
|
||||||
|
" --drop-no-renorm do not rescale the surviving weights after a drop (A/B)\n"
|
||||||
|
" --drop-in-prefill drop during prefill too (off: the cold cache makes it expensive)\n"
|
||||||
" --list-archs print supported MoE architectures and exit\n"
|
" --list-archs print supported MoE architectures and exit\n"
|
||||||
"\n"
|
"\n"
|
||||||
" Env overrides (flag wins): BMOE_CACHE_MB, BMOE_IO_THREADS, BMOE_PROGRESS, BMOE_OVERLAP, BMOE_PREFETCH, "
|
" Env overrides (flag wins): BMOE_CACHE_MB, BMOE_IO_THREADS, BMOE_PROGRESS, BMOE_OVERLAP, BMOE_PREFETCH, "
|
||||||
|
|
@ -478,6 +484,12 @@ int main(int argc, char ** argv) {
|
||||||
cfg.moe.prefetch_layers = std::atoi(next("--prefetch"));
|
cfg.moe.prefetch_layers = std::atoi(next("--prefetch"));
|
||||||
else if (a == "--prefetch-sync") // debug: complete speculative reads synchronously
|
else if (a == "--prefetch-sync") // debug: complete speculative reads synchronously
|
||||||
cfg.moe.prefetch_sync = true;
|
cfg.moe.prefetch_sync = true;
|
||||||
|
else if (a == "--drop-cold-experts")
|
||||||
|
cfg.moe.drop_cold_frac = (float) std::atof(next("--drop-cold-experts"));
|
||||||
|
else if (a == "--drop-no-renorm")
|
||||||
|
cfg.moe.drop_renorm = false;
|
||||||
|
else if (a == "--drop-in-prefill")
|
||||||
|
cfg.moe.drop_prefill = true;
|
||||||
else if (a == "--list-archs") {
|
else if (a == "--list-archs") {
|
||||||
std::printf("supported MoE architectures:\n");
|
std::printf("supported MoE architectures:\n");
|
||||||
for (int k = 0; k < n_moe_recipes(); ++k)
|
for (int k = 0; k < n_moe_recipes(); ++k)
|
||||||
|
|
@ -622,6 +634,14 @@ int main(int argc, char ** argv) {
|
||||||
std::printf("moe-prefetch: %.1f MiB speculative, %lld/%lld experts useful (%.0f%%)\n", s.moe_spec_read_mib,
|
std::printf("moe-prefetch: %.1f MiB speculative, %lld/%lld experts useful (%.0f%%)\n", s.moe_spec_read_mib,
|
||||||
s.moe_spec_useful, s.moe_spec_experts,
|
s.moe_spec_useful, s.moe_spec_experts,
|
||||||
s.moe_spec_experts > 0 ? 100.0 * s.moe_spec_useful / s.moe_spec_experts : 0.0);
|
s.moe_spec_experts > 0 ? 100.0 * s.moe_spec_useful / s.moe_spec_experts : 0.0);
|
||||||
|
// How hard the policy actually bit. The flag sets a threshold, not a drop rate: what gets
|
||||||
|
// discarded depends on what the cache held, so this is the only honest report of the trade
|
||||||
|
// a given run made.
|
||||||
|
if (cfg.moe.drop_cold_frac > 0.0f)
|
||||||
|
std::printf("moe-drop: %lld/%lld routed experts dropped (%.1f%%), threshold %.2f x uniform\n",
|
||||||
|
s.experts_dropped, s.experts_routed,
|
||||||
|
s.experts_routed > 0 ? 100.0 * s.experts_dropped / s.experts_routed : 0.0,
|
||||||
|
(double) cfg.moe.drop_cold_frac);
|
||||||
}
|
}
|
||||||
return 0;
|
return 0;
|
||||||
}
|
}
|
||||||
|
|
|
||||||
|
|
@ -94,6 +94,35 @@ struct MoeStreamConfig {
|
||||||
// and the policy the Android app ships by default — the CLI matches it here.
|
// and the policy the Android app ships by default — the CLI matches it here.
|
||||||
DenseWeightsMode dense_weights = DenseWeightsMode::Anonymous;
|
DenseWeightsMode dense_weights = DenseWeightsMode::Anonymous;
|
||||||
|
|
||||||
|
// ── cache-aware expert dropping (lossy; opt-in) ──────────────────────────────────
|
||||||
|
// Skip a routed expert when it is a cache MISS *and* the router weighted it below
|
||||||
|
// drop_cold_frac × (1 / n_expert_used) — i.e. below that fraction of the uniform share a
|
||||||
|
// top-k routing would give each expert. 0 (the default) disables it and the engine is
|
||||||
|
// bit-exact as before.
|
||||||
|
//
|
||||||
|
// The asymmetry is the whole idea: an expert already resident costs no flash read, so it
|
||||||
|
// always runs however small its weight. Quality is spent only where it buys I/O. Because
|
||||||
|
// the largest weight in a routing is always >= the uniform share, a frac of 1.0 can never
|
||||||
|
// empty a routing; validate() rejects anything above it, and the implementation additionally pins the
|
||||||
|
// top-weighted expert so no cell is ever left with nothing to compute.
|
||||||
|
//
|
||||||
|
// This changes the output — it is a quality/throughput trade like n_expert_used, not an
|
||||||
|
// optimisation. Unlike n_expert_used it is *state-dependent*: the same prompt can decode
|
||||||
|
// differently depending on what the cache happened to hold, so a run is no longer
|
||||||
|
// reproducible token-for-token. See docs/expert-dropping.md.
|
||||||
|
float drop_cold_frac = 0.0f;
|
||||||
|
|
||||||
|
// Rescale the surviving weights so the routing still sums to what it did before the drop.
|
||||||
|
// Without it the layer's expert output is systematically scaled down by the discarded mass
|
||||||
|
// (~10% at frac 1.0), which perturbs the residual stream more than the missing expert does.
|
||||||
|
bool drop_renorm = true;
|
||||||
|
|
||||||
|
// Apply dropping during prefill too. Off by default and deliberately so: the cache is cold
|
||||||
|
// there, so nearly every expert is a miss and the same threshold discards ~4x the weight
|
||||||
|
// mass it does in decode (measured; see docs/expert-dropping.md). Prefill is also
|
||||||
|
// compute-bound, so there is little to win.
|
||||||
|
bool drop_prefill = false;
|
||||||
|
|
||||||
// Test/debug only: complete each prefetch's speculative reads synchronously, on the eval
|
// Test/debug only: complete each prefetch's speculative reads synchronously, on the eval
|
||||||
// thread, before returning. This defeats the latency-hiding purpose (the reads no longer
|
// thread, before returning. This defeats the latency-hiding purpose (the reads no longer
|
||||||
// overlap compute) but makes speculative integration deterministic, so the byte-identity
|
// overlap compute) but makes speculative integration deterministic, so the byte-identity
|
||||||
|
|
|
||||||
|
|
@ -108,6 +108,17 @@ struct RunSummary {
|
||||||
double token_demand_mib = 0.0;
|
double token_demand_mib = 0.0;
|
||||||
// The widest layer's routed bytes: the mechanical floor a cache must be able to stage.
|
// The widest layer's routed bytes: the mechanical floor a cache must be able to stage.
|
||||||
double layer_demand_mib = 0.0;
|
double layer_demand_mib = 0.0;
|
||||||
|
// Both measure what reached the streamer. Under MoeStreamConfig::drop_cold_frac that is what a
|
||||||
|
// token STAGES, not what it routed — a dropped expert is never handed over — so the "floor a
|
||||||
|
// cache must clear" reading stops being mechanical there: the floor shrinks because the cache
|
||||||
|
// was small. Size the cache with dropping off, then turn it on.
|
||||||
|
|
||||||
|
// Cache-aware expert dropping (zero when --drop-cold-experts is off). Routed counts what the
|
||||||
|
// router selected across the generation, dropped how much of it the policy declined to read;
|
||||||
|
// their ratio is the lever's actual bite, which depends on the cache and so cannot be read off
|
||||||
|
// the flag. Both cover generation only — prefill drops nothing unless armed for it.
|
||||||
|
long long experts_routed = 0;
|
||||||
|
long long experts_dropped = 0;
|
||||||
|
|
||||||
// Temporal prefetch (zero when --prefetch is off): speculative bytes read during generation,
|
// Temporal prefetch (zero when --prefetch is off): speculative bytes read during generation,
|
||||||
// experts successfully prefetched, and how many of those a later routing actually used.
|
// experts successfully prefetched, and how many of those a later routing actually used.
|
||||||
|
|
@ -142,6 +153,7 @@ struct RunInfo {
|
||||||
bool overlap = false;
|
bool overlap = false;
|
||||||
int prefetch_layers = 0;
|
int prefetch_layers = 0;
|
||||||
std::string dense_weights = "anon"; // dense (non-expert) policy: "mmap" | "warm" | "anon"
|
std::string dense_weights = "anon"; // dense (non-expert) policy: "mmap" | "warm" | "anon"
|
||||||
|
float drop_cold_frac = 0.0f; // cache-aware expert dropping threshold (0 = off)
|
||||||
};
|
};
|
||||||
|
|
||||||
// Optional per-token sink (e.g. CSV for benchmarks). The engine calls on_run_info once before the
|
// Optional per-token sink (e.g. CSV for benchmarks). The engine calls on_run_info once before the
|
||||||
|
|
|
||||||
|
|
@ -38,6 +38,12 @@ struct RouteTraceRow {
|
||||||
float weight = 0.0f; // final applied routing weight; NaN if the graph exposed none
|
float weight = 0.0f; // final applied routing weight; NaN if the graph exposed none
|
||||||
uint8_t residency = 0; // RouteResidency
|
uint8_t residency = 0; // RouteResidency
|
||||||
uint64_t expert_bytes = 0; // flash bytes this routing reads; 0 unless residency == route_miss
|
uint64_t expert_bytes = 0; // flash bytes this routing reads; 0 unless residency == route_miss
|
||||||
|
// Cache-aware dropping (MoeStreamConfig::drop_cold_frac) discarded this routing: the expert was
|
||||||
|
// a miss weighted below the threshold, so it was never read and its weight was zeroed. `weight`
|
||||||
|
// and `residency` stay as the ROUTER produced them — the trace records the routing that was
|
||||||
|
// chosen, and this flag records what the policy then did with it — but expert_bytes is 0,
|
||||||
|
// because a dropped expert costs no read. Always 0 when dropping is off.
|
||||||
|
uint8_t dropped = 0;
|
||||||
};
|
};
|
||||||
|
|
||||||
// Run-level facts the rows cannot carry. Emitted once, before any row.
|
// Run-level facts the rows cannot carry. Emitted once, before any row.
|
||||||
|
|
|
||||||
|
|
@ -5,7 +5,8 @@
|
||||||
// them to the expert source, then greedily generates n_predict tokens — reporting each
|
// them to the expert source, then greedily generates n_predict tokens — reporting each
|
||||||
// token to the optional callback/sink and returning a RunSummary. Greedy sampling makes
|
// token to the optional callback/sink and returning a RunSummary. Greedy sampling makes
|
||||||
// the output a deterministic function of the graph, which is what the byte-identity
|
// the output a deterministic function of the graph, which is what the byte-identity
|
||||||
// gates rely on.
|
// gates rely on — true of every configuration except MoeStreamConfig::drop_cold_frac, which
|
||||||
|
// decides from live cache state and so is reproducible only within a single run.
|
||||||
#pragma once
|
#pragma once
|
||||||
|
|
||||||
#include "bmoe/config.h"
|
#include "bmoe/config.h"
|
||||||
|
|
|
||||||
|
|
@ -85,6 +85,17 @@ ValidationResult validate(const RunConfig & cfg) {
|
||||||
"speculative reads land in the per-layer cache buffers, which do not exist "
|
"speculative reads land in the per-layer cache buffers, which do not exist "
|
||||||
"with the cache off.");
|
"with the cache off.");
|
||||||
}
|
}
|
||||||
|
if (m.drop_cold_frac > 0.0f && !cache_on) {
|
||||||
|
return fail("moe.drop_cold_frac requires the LRU cache (cache_mb > 0 or cache_auto): with the "
|
||||||
|
"cache off every expert is a miss, so the policy stops being cache-aware and "
|
||||||
|
"degenerates into an unconditional weight cut — which is what n_expert_used already "
|
||||||
|
"does, without pretending to consult residency.");
|
||||||
|
}
|
||||||
|
if (m.drop_cold_frac < 0.0f || m.drop_cold_frac > 1.0f) {
|
||||||
|
return fail("moe.drop_cold_frac must be in [0, 1] (0 = off). Above 1.0 the threshold can "
|
||||||
|
"exceed the largest weight in a routing, which would discard every expert of a "
|
||||||
|
"layer; 1.0 is the uniform share 1/n_expert_used and the useful maximum.");
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
return r;
|
return r;
|
||||||
|
|
|
||||||
|
|
@ -370,6 +370,7 @@ std::unique_ptr<Session> Session::open(const SessionConfig & cfg,
|
||||||
|
|
||||||
im.hook = std::make_unique<RouterHook>(recipe ? *recipe : MoeRecipe{}, im.n_layer);
|
im.hook = std::make_unique<RouterHook>(recipe ? *recipe : MoeRecipe{}, im.n_layer);
|
||||||
im.hook->set_prefetch_layers(cfg.moe.prefetch_layers);
|
im.hook->set_prefetch_layers(cfg.moe.prefetch_layers);
|
||||||
|
im.hook->set_drop_policy(cfg.moe.drop_cold_frac, cfg.moe.drop_renorm, cfg.moe.drop_prefill);
|
||||||
|
|
||||||
llama_context_params cparams = llama_context_default_params();
|
llama_context_params cparams = llama_context_default_params();
|
||||||
cparams.n_ctx = cfg.n_ctx;
|
cparams.n_ctx = cfg.n_ctx;
|
||||||
|
|
@ -561,6 +562,7 @@ std::unique_ptr<Session> Session::open(const SessionConfig & cfg,
|
||||||
ri.o_direct = cfg.moe.enabled && cfg.moe.o_direct;
|
ri.o_direct = cfg.moe.enabled && cfg.moe.o_direct;
|
||||||
ri.overlap = cfg.moe.enabled && cfg.moe.overlap;
|
ri.overlap = cfg.moe.enabled && cfg.moe.overlap;
|
||||||
ri.prefetch_layers = cfg.moe.enabled ? cfg.moe.prefetch_layers : 0;
|
ri.prefetch_layers = cfg.moe.enabled ? cfg.moe.prefetch_layers : 0;
|
||||||
|
ri.drop_cold_frac = cfg.moe.enabled ? cfg.moe.drop_cold_frac : 0.0f;
|
||||||
// The CSV keeps the two familiar flags, derived from the resolved dense-weights policy.
|
// The CSV keeps the two familiar flags, derived from the resolved dense-weights policy.
|
||||||
ri.dense_weights = cfg.moe.dense_weights == DenseWeightsMode::Mmap ? "mmap"
|
ri.dense_weights = cfg.moe.dense_weights == DenseWeightsMode::Mmap ? "mmap"
|
||||||
: cfg.moe.dense_weights == DenseWeightsMode::Anonymous ? "anon"
|
: cfg.moe.dense_weights == DenseWeightsMode::Anonymous ? "anon"
|
||||||
|
|
@ -757,6 +759,9 @@ RunResult Session::generate(const GenerateRequest & req,
|
||||||
// The frame the I/O rows are stamped with at flush; the other traces carry their own.
|
// The frame the I/O rows are stamped with at flush; the other traces carry their own.
|
||||||
int trace_phase = 0, trace_step = 0;
|
int trace_phase = 0, trace_step = 0;
|
||||||
auto trace_begin = [&](int base_pos, int n_tokens, int phase) {
|
auto trace_begin = [&](int base_pos, int n_tokens, int phase) {
|
||||||
|
// Not a trace concern, but the same per-decode frame: the drop policy is decode-only
|
||||||
|
// unless armed for prefill, so it has to be told which phase this batch is.
|
||||||
|
im.hook->set_batch_phase(phase);
|
||||||
if (im.route_trace) im.hook->begin_trace_batch(base_pos, n_tokens, phase, im.turn);
|
if (im.route_trace) im.hook->begin_trace_batch(base_pos, n_tokens, phase, im.turn);
|
||||||
// A node is computed once for the whole batch, not per token, so a prefill chunk's graph is
|
// A node is computed once for the whole batch, not per token, so a prefill chunk's graph is
|
||||||
// attributed to its last position rather than pretending to split across the chunk.
|
// attributed to its last position rather than pretending to split across the chunk.
|
||||||
|
|
@ -842,6 +847,10 @@ RunResult Session::generate(const GenerateRequest & req,
|
||||||
long long prev_spec_bytes = moe.enabled ? (long long) im.source.stats().spec_read_bytes : 0;
|
long long prev_spec_bytes = moe.enabled ? (long long) im.source.stats().spec_read_bytes : 0;
|
||||||
long long prev_spec_experts = moe.enabled ? im.source.stats().spec_experts : 0;
|
long long prev_spec_experts = moe.enabled ? im.source.stats().spec_experts : 0;
|
||||||
long long prev_spec_useful = moe.enabled ? im.source.stats().spec_useful : 0;
|
long long prev_spec_useful = moe.enabled ? im.source.stats().spec_useful : 0;
|
||||||
|
// Taken after prefill, so the drop counters describe generation — the phase the policy is armed
|
||||||
|
// for and the one the tok/s number is about.
|
||||||
|
const long long prev_routed = im.hook->experts_routed();
|
||||||
|
const long long prev_dropped = im.hook->experts_dropped();
|
||||||
|
|
||||||
for (int t = 0; t < req.n_predict; ++t) {
|
for (int t = 0; t < req.n_predict; ++t) {
|
||||||
// Greedy stays argmax (byte-identical to the resident reference the gates check); with a
|
// Greedy stays argmax (byte-identical to the resident reference the gates check); with a
|
||||||
|
|
@ -931,6 +940,8 @@ RunResult Session::generate(const GenerateRequest & req,
|
||||||
s.moe_spec_experts = st.spec_experts - prev_spec_experts;
|
s.moe_spec_experts = st.spec_experts - prev_spec_experts;
|
||||||
s.moe_spec_useful = st.spec_useful - prev_spec_useful;
|
s.moe_spec_useful = st.spec_useful - prev_spec_useful;
|
||||||
}
|
}
|
||||||
|
s.experts_routed = im.hook->experts_routed() - prev_routed;
|
||||||
|
s.experts_dropped = im.hook->experts_dropped() - prev_dropped;
|
||||||
if (sink) sink->on_summary(s);
|
if (sink) sink->on_summary(s);
|
||||||
|
|
||||||
{
|
{
|
||||||
|
|
|
||||||
|
|
@ -23,9 +23,9 @@ public:
|
||||||
r.model.c_str(), r.arch.c_str(), r.n_layer, r.n_expert, r.n_expert_used, r.n_threads, r.n_ctx);
|
r.model.c_str(), r.arch.c_str(), r.n_layer, r.n_expert, r.n_expert_used, r.n_threads, r.n_ctx);
|
||||||
std::fprintf(f_,
|
std::fprintf(f_,
|
||||||
"# moe_stream=%d cache_mb=%d cache_auto=%d cache_ceil_mb=%d force_cache=%d "
|
"# moe_stream=%d cache_mb=%d cache_auto=%d cache_ceil_mb=%d force_cache=%d "
|
||||||
"io_threads=%d o_direct=%d overlap=%d prefetch=%d dense_weights=%s\n",
|
"io_threads=%d o_direct=%d overlap=%d prefetch=%d dense_weights=%s drop_cold_frac=%.4g\n",
|
||||||
r.moe_stream, r.cache_mb, r.cache_auto, r.cache_ceil_mb, r.force_cache, r.io_threads, r.o_direct,
|
r.moe_stream, r.cache_mb, r.cache_auto, r.cache_ceil_mb, r.force_cache, r.io_threads, r.o_direct,
|
||||||
r.overlap, r.prefetch_layers, r.dense_weights.c_str());
|
r.overlap, r.prefetch_layers, r.dense_weights.c_str(), (double) r.drop_cold_frac);
|
||||||
write_header();
|
write_header();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
@ -49,13 +49,14 @@ public:
|
||||||
"n_prompt=%d load_s=%.3f prefill_s=%.3f prefill_tps=%.2f stall_s/tok=%.3f mgmt_s/tok=%.3f "
|
"n_prompt=%d load_s=%.3f prefill_s=%.3f prefill_tps=%.2f stall_s/tok=%.3f mgmt_s/tok=%.3f "
|
||||||
"cache_resident_MiB=%.1f cache_budget_MiB=%.1f cache_resizes=%lld "
|
"cache_resident_MiB=%.1f cache_budget_MiB=%.1f cache_resizes=%lld "
|
||||||
"spec_read_MiB=%.1f spec_experts=%lld spec_useful=%lld "
|
"spec_read_MiB=%.1f spec_experts=%lld spec_useful=%lld "
|
||||||
"majflt/tok=%.2f cpu_s/tok=%.4f token_demand_MiB=%.1f layer_demand_MiB=%.1f\n",
|
"majflt/tok=%.2f cpu_s/tok=%.4f token_demand_MiB=%.1f layer_demand_MiB=%.1f "
|
||||||
|
"experts_routed=%lld experts_dropped=%lld\n",
|
||||||
s.n_generated, s.s_per_token, s.tokens_per_second, s.moe_read_mib, s.moe_io_seconds,
|
s.n_generated, s.s_per_token, s.tokens_per_second, s.moe_read_mib, s.moe_io_seconds,
|
||||||
s.moe_compute_s_per_token, s.moe_io_s_per_token, s.cache_hit_pct, s.n_prompt, s.load_seconds,
|
s.moe_compute_s_per_token, s.moe_io_s_per_token, s.cache_hit_pct, s.n_prompt, s.load_seconds,
|
||||||
s.prefill_seconds, s.prefill_seconds > 0 ? s.n_prompt / s.prefill_seconds : 0.0,
|
s.prefill_seconds, s.prefill_seconds > 0 ? s.n_prompt / s.prefill_seconds : 0.0,
|
||||||
s.moe_stall_s_per_token, s.moe_mgmt_s_per_token, s.cache_resident_mib, s.cache_budget_mib,
|
s.moe_stall_s_per_token, s.moe_mgmt_s_per_token, s.cache_resident_mib, s.cache_budget_mib,
|
||||||
s.cache_resizes, s.moe_spec_read_mib, s.moe_spec_experts, s.moe_spec_useful, s.majflt_per_token,
|
s.cache_resizes, s.moe_spec_read_mib, s.moe_spec_experts, s.moe_spec_useful, s.majflt_per_token,
|
||||||
s.cpu_s_per_token, s.token_demand_mib, s.layer_demand_mib);
|
s.cpu_s_per_token, s.token_demand_mib, s.layer_demand_mib, s.experts_routed, s.experts_dropped);
|
||||||
std::fflush(f_);
|
std::fflush(f_);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -26,7 +26,7 @@ public:
|
||||||
std::fprintf(f_, "# layer=%zu expert_bytes=%llu dense_bytes=%llu\n", il,
|
std::fprintf(f_, "# layer=%zu expert_bytes=%llu dense_bytes=%llu\n", il,
|
||||||
(unsigned long long) s.expert_bytes_per_layer[il], (unsigned long long) dense);
|
(unsigned long long) s.expert_bytes_per_layer[il], (unsigned long long) dense);
|
||||||
}
|
}
|
||||||
std::fprintf(f_, "turn,phase,step,layer,slot,expert,weight,residency,expert_bytes\n");
|
std::fprintf(f_, "turn,phase,step,layer,slot,expert,weight,residency,expert_bytes,dropped\n");
|
||||||
std::fflush(f_);
|
std::fflush(f_);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
@ -36,12 +36,13 @@ public:
|
||||||
// A weight the graph never exposed prints as nan, not 0: "unknown" must not read as
|
// A weight the graph never exposed prints as nan, not 0: "unknown" must not read as
|
||||||
// "the router gave this expert no mass".
|
// "the router gave this expert no mass".
|
||||||
if (std::isnan(r.weight))
|
if (std::isnan(r.weight))
|
||||||
std::fprintf(f_, "%d,%d,%d,%d,%d,%d,nan,%u,%llu\n", r.turn, r.phase, r.step, r.layer, r.slot,
|
std::fprintf(f_, "%d,%d,%d,%d,%d,%d,nan,%u,%llu,%u\n", r.turn, r.phase, r.step, r.layer, r.slot,
|
||||||
(int) r.expert, (unsigned) r.residency, (unsigned long long) r.expert_bytes);
|
(int) r.expert, (unsigned) r.residency, (unsigned long long) r.expert_bytes,
|
||||||
|
(unsigned) r.dropped);
|
||||||
else
|
else
|
||||||
std::fprintf(f_, "%d,%d,%d,%d,%d,%d,%.6g,%u,%llu\n", r.turn, r.phase, r.step, r.layer, r.slot,
|
std::fprintf(f_, "%d,%d,%d,%d,%d,%d,%.6g,%u,%llu,%u\n", r.turn, r.phase, r.step, r.layer, r.slot,
|
||||||
(int) r.expert, (double) r.weight, (unsigned) r.residency,
|
(int) r.expert, (double) r.weight, (unsigned) r.residency,
|
||||||
(unsigned long long) r.expert_bytes);
|
(unsigned long long) r.expert_bytes, (unsigned) r.dropped);
|
||||||
}
|
}
|
||||||
std::fflush(f_); // once per decode, not per row
|
std::fflush(f_); // once per decode, not per row
|
||||||
}
|
}
|
||||||
|
|
|
||||||
|
|
@ -655,8 +655,10 @@ bool ExpertStreamSource::load_layer(int il, const int32_t * ids, int n_ids) {
|
||||||
// Already staged in this batch: still promote so the LRU order reflects the LAST
|
// Already staged in this batch: still promote so the LRU order reflects the LAST
|
||||||
// token that used this expert (ids arrive token-major), not its first touch — this
|
// token that used this expert (ids arrive token-major), not its first touch — this
|
||||||
// keeps the prompt tail's experts hot across prefill. Reads are scheduled only once
|
// keeps the prompt tail's experts hot across prefill. Reads are scheduled only once
|
||||||
// (the seen_ guard below), so this is bookkeeping only. In decode (n=1) the top-k ids
|
// (the seen_ guard below), so this is bookkeeping only. The router's own top-k ids are
|
||||||
// are distinct, so this branch never runs and behaviour is unchanged.
|
// distinct, so in decode (n=1) this used to be unreachable — but cache-aware dropping
|
||||||
|
// repoints a dropped slot's id at the routing's top expert, which makes duplicates the
|
||||||
|
// normal case there. Promoting the same entry twice is idempotent, so it stays correct.
|
||||||
if (cache_max_) {
|
if (cache_max_) {
|
||||||
const int32_t id = il * n_expert_ + e;
|
const int32_t id = il * n_expert_ + e;
|
||||||
lru_unlink(id);
|
lru_unlink(id);
|
||||||
|
|
@ -822,7 +824,9 @@ bool ExpertStreamSource::load_layer_async(int il, const int32_t * ids, int n_ids
|
||||||
// Promote every touched expert in raw id order (token-major) so the LRU order reflects
|
// Promote every touched expert in raw id order (token-major) so the LRU order reflects
|
||||||
// the LAST token that used it, not the sorted/first-touch order staged_ imposes — this
|
// the LAST token that used it, not the sorted/first-touch order staged_ imposes — this
|
||||||
// keeps the prompt tail's experts hot across prefill. Bookkeeping only; every id staged
|
// keeps the prompt tail's experts hot across prefill. Bookkeeping only; every id staged
|
||||||
// above is now valid and linked. In decode (n=1, distinct top-k) this is a no-op reshuffle.
|
// above is now valid and linked. In decode (n=1) this is a no-op reshuffle when the ids are
|
||||||
|
// distinct, and an idempotent re-promote of the same entry when cache-aware dropping has
|
||||||
|
// repointed a slot at the routing's top expert.
|
||||||
// Skipped in load_all (everything is resident, so LRU order is meaningless).
|
// Skipped in load_all (everything is resident, so LRU order is meaningless).
|
||||||
if (!load_all_) {
|
if (!load_all_) {
|
||||||
for (int i = 0; i < n_ids; ++i) {
|
for (int i = 0; i < n_ids; ++i) {
|
||||||
|
|
|
||||||
|
|
@ -62,15 +62,27 @@ static bool match_weights(const char * name, int & il_out) {
|
||||||
// (token j at nb[2], slot k at nb[1]) — except the norm variant, whose callback fires on the
|
// (token j at nb[2], slot k at nb[1]) — except the norm variant, whose callback fires on the
|
||||||
// pre-reshape 2-D [nu, nt] (token j at nb[1], slot k at nb[0]). ne[0] == 1 tells the two apart.
|
// pre-reshape 2-D [nu, nt] (token j at nb[1], slot k at nb[0]). ne[0] == 1 tells the two apart.
|
||||||
// As with the topk ids, these are views: only the strides say where a token's row really starts.
|
// As with the topk ids, these are views: only the strides say where a token's row really starts.
|
||||||
static void gather_weights(const ggml_tensor * t, int nu, int nt, std::vector<float> & out) {
|
// Where token j's slot k lives inside a weight node. Single source of truth for the layout: the
|
||||||
|
// reader below and the drop policy's writer must agree, or the policy would zero another token's
|
||||||
|
// slot. Returns a mutable pointer; the gather takes a const tensor and only reads through it.
|
||||||
|
static float * weight_at(const ggml_tensor * t, int j, int k) {
|
||||||
const bool three_d = t->ne[0] == 1;
|
const bool three_d = t->ne[0] == 1;
|
||||||
const size_t tok_nb = three_d ? t->nb[2] : t->nb[1];
|
const size_t tok_nb = three_d ? t->nb[2] : t->nb[1];
|
||||||
const size_t slot_nb = three_d ? t->nb[1] : t->nb[0];
|
const size_t slot_nb = three_d ? t->nb[1] : t->nb[0];
|
||||||
|
return (float *) ((char *) t->data + (size_t) j * tok_nb + (size_t) k * slot_nb);
|
||||||
|
}
|
||||||
|
|
||||||
|
static void gather_weights(const ggml_tensor * t, int nu, int nt, std::vector<float> & out) {
|
||||||
out.assign((size_t) nu * nt, 0.0f);
|
out.assign((size_t) nu * nt, 0.0f);
|
||||||
for (int j = 0; j < nt; ++j)
|
for (int j = 0; j < nt; ++j)
|
||||||
for (int k = 0; k < nu; ++k)
|
for (int k = 0; k < nu; ++k)
|
||||||
out[(size_t) j * nu + k] =
|
out[(size_t) j * nu + k] = *weight_at(t, j, k);
|
||||||
*(const float *) ((const char *) t->data + (size_t) j * tok_nb + (size_t) k * slot_nb);
|
}
|
||||||
|
|
||||||
|
// Same, for the selected-expert ids. The node is a VIEW of the full argsort, so the row stride is
|
||||||
|
// nb[1] (n_expert * 4), not n_expert_used * 4 — see the gather at the topk node.
|
||||||
|
static int32_t * id_at(ggml_tensor * t, int j, int k) {
|
||||||
|
return (int32_t *) ((char *) t->data + (size_t) j * t->nb[1] + (size_t) k * t->nb[0]);
|
||||||
}
|
}
|
||||||
|
|
||||||
void RouterHook::begin_capture() {
|
void RouterHook::begin_capture() {
|
||||||
|
|
@ -83,6 +95,114 @@ void RouterHook::end_capture() {
|
||||||
capturing_ = false;
|
capturing_ = false;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
void RouterHook::set_drop_policy(float frac, bool renorm, bool in_prefill) {
|
||||||
|
drop_frac_ = frac > 0.0f ? frac : 0.0f;
|
||||||
|
drop_renorm_ = renorm;
|
||||||
|
drop_prefill_ = in_prefill;
|
||||||
|
term_node_.assign(n_layer_ > 0 ? n_layer_ : 0, std::string{});
|
||||||
|
drop_ = PendingDrop{};
|
||||||
|
chain_last_.clear();
|
||||||
|
experts_routed_ = experts_dropped_ = 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Is dropping live for the batch being decoded? Needs a source to ask about residency, a non-zero
|
||||||
|
// threshold, and — unless armed for prefill — a decode batch: with a cold cache the same threshold
|
||||||
|
// discards several times the weight mass for a phase that is not I/O-bound anyway.
|
||||||
|
bool RouterHook::drop_armed() const {
|
||||||
|
return drop_frac_ > 0.0f && source_ != nullptr && (batch_phase_ == 1 || drop_prefill_);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Apply the policy to the layer held in drop_, then load only what survives.
|
||||||
|
//
|
||||||
|
// `wt` is the terminal node of the weight chain: the weights as the expert matmul will apply them.
|
||||||
|
// Both edits happen here, before any node consumes them:
|
||||||
|
// - the dropped slot's weight is zeroed (and, with renorm, the survivors are scaled back up so
|
||||||
|
// the routing keeps the total mass it had);
|
||||||
|
// - the dropped slot's ID is repointed at the routing's top-weighted expert. That second edit is
|
||||||
|
// not cosmetic. An expert we decline to read may sit in a reserved-but-uncommitted slot, and
|
||||||
|
// mul_mat_id would still touch it; pointing the slot at an expert that is certainly resident
|
||||||
|
// makes the kernel read valid memory and multiply it by exactly zero. It costs a duplicate
|
||||||
|
// matmul, which is the right trade on a decode bound by flash rather than arithmetic.
|
||||||
|
// The top-weighted expert is never dropped, so a routing always keeps at least one live expert
|
||||||
|
// whatever the threshold — the guarantee does not rest on frac <= 1 alone.
|
||||||
|
void RouterHook::apply_drop(ggml_tensor * wt) {
|
||||||
|
PendingDrop & D = drop_;
|
||||||
|
const int nu = D.nu, nt = D.nt;
|
||||||
|
gather_weights(wt, nu, nt, drop_w_);
|
||||||
|
|
||||||
|
// Classify against the cache BEFORE anything is loaded; settle landed prefetches first, or an
|
||||||
|
// expert a prefetch correctly guessed would look like a miss and be dropped for nothing.
|
||||||
|
source_->settle_spec();
|
||||||
|
drop_res_.assign(drop_ids_.size(), (uint8_t) 0);
|
||||||
|
source_->query_residency(D.layer, drop_ids_.data(), (int) drop_ids_.size(), drop_res_.data());
|
||||||
|
|
||||||
|
const float thr = drop_frac_ / (float) nu; // frac of the uniform share each of k experts would get
|
||||||
|
const bool tracing = trace_on_ && pending_.layer == D.layer;
|
||||||
|
drop_mask_.assign((size_t) nu * nt, (uint8_t) 0);
|
||||||
|
|
||||||
|
for (int j = 0; j < nt; ++j) {
|
||||||
|
const size_t row = (size_t) j * nu;
|
||||||
|
int best = 0;
|
||||||
|
float total = 0.0f;
|
||||||
|
for (int k = 0; k < nu; ++k) {
|
||||||
|
total += drop_w_[row + k];
|
||||||
|
if (drop_w_[row + k] > drop_w_[row + best]) best = k;
|
||||||
|
}
|
||||||
|
const int32_t best_id = drop_ids_[row + best];
|
||||||
|
|
||||||
|
float kept = 0.0f;
|
||||||
|
int n_dropped = 0;
|
||||||
|
for (int k = 0; k < nu; ++k) {
|
||||||
|
const size_t idx = row + k;
|
||||||
|
const bool drop = k != best && drop_res_[idx] == route_miss && drop_w_[idx] < thr;
|
||||||
|
if (!drop) {
|
||||||
|
kept += drop_w_[idx];
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
*weight_at(wt, j, k) = 0.0f;
|
||||||
|
*id_at(D.ids, j, k) = best_id;
|
||||||
|
drop_ids_[idx] = best_id;
|
||||||
|
drop_mask_[idx] = 1;
|
||||||
|
++n_dropped;
|
||||||
|
}
|
||||||
|
experts_dropped_ += n_dropped;
|
||||||
|
|
||||||
|
// Restore the routing's total mass. Without this the layer's expert output is scaled down
|
||||||
|
// by whatever was discarded, which perturbs the residual stream in a direction the model
|
||||||
|
// never sees in training — a systematic shrink, unlike the one missing contribution.
|
||||||
|
if (drop_renorm_ && n_dropped > 0 && kept > 0.0f) {
|
||||||
|
const float g = total / kept;
|
||||||
|
for (int k = 0; k < nu; ++k)
|
||||||
|
if (!drop_mask_[row + k]) *weight_at(wt, j, k) *= g;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
if (tracing) pending_.dropped = drop_mask_;
|
||||||
|
source_->load_layer(D.layer, drop_ids_.data(), (int) drop_ids_.size());
|
||||||
|
D.deferred = false;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Finish with the layer whose topk we last saw: record which node ended its weight chain, so the
|
||||||
|
// next graph can decide there, and make sure nothing was left waiting on a node that never came.
|
||||||
|
void RouterHook::close_drop_layer() {
|
||||||
|
PendingDrop & D = drop_;
|
||||||
|
if (D.layer < 0) return;
|
||||||
|
if (D.layer < (int) term_node_.size() && term_node_[D.layer].empty() && !chain_last_.empty())
|
||||||
|
term_node_[D.layer] = chain_last_;
|
||||||
|
if (D.deferred && source_) {
|
||||||
|
// The node we learned as terminal did not appear this time, so the deferral was never
|
||||||
|
// honoured and this layer's matmul has already run against slots nothing loaded. Load the
|
||||||
|
// routing now to keep the cache's accounting straight, and — more importantly — FORGET the
|
||||||
|
// terminal node, so the next graph re-learns it and loads at the topk node meanwhile.
|
||||||
|
// Deferring again on the same stale guess would repeat the fault every single token; one
|
||||||
|
// bad layer in one token is recoverable, a standing bet against a graph that moved is not.
|
||||||
|
source_->load_layer(D.layer, drop_ids_.data(), (int) drop_ids_.size());
|
||||||
|
if (D.layer < (int) term_node_.size()) term_node_[D.layer].clear();
|
||||||
|
D.deferred = false;
|
||||||
|
}
|
||||||
|
D.layer = -1;
|
||||||
|
}
|
||||||
|
|
||||||
void RouterHook::set_trace(bool on) {
|
void RouterHook::set_trace(bool on) {
|
||||||
trace_on_ = on;
|
trace_on_ = on;
|
||||||
pending_ = PendingLayer{};
|
pending_ = PendingLayer{};
|
||||||
|
|
@ -143,7 +263,11 @@ void RouterHook::flush_pending() {
|
||||||
r.expert = P.ids[idx];
|
r.expert = P.ids[idx];
|
||||||
r.weight = have_w ? P.weights[idx] : std::numeric_limits<float>::quiet_NaN();
|
r.weight = have_w ? P.weights[idx] : std::numeric_limits<float>::quiet_NaN();
|
||||||
r.residency = idx < P.residency.size() ? P.residency[idx] : (uint8_t) 0;
|
r.residency = idx < P.residency.size() ? P.residency[idx] : (uint8_t) 0;
|
||||||
if (r.residency == route_miss && charged_.insert(r.expert).second) r.expert_bytes = ebytes;
|
r.dropped = idx < P.dropped.size() ? P.dropped[idx] : (uint8_t) 0;
|
||||||
|
// A dropped routing is never read, so it is neither charged nor allowed to claim the
|
||||||
|
// charge for its expert: if another token of the batch routes the same expert and keeps
|
||||||
|
// it, that routing pays the read.
|
||||||
|
if (!r.dropped && r.residency == route_miss && charged_.insert(r.expert).second) r.expert_bytes = ebytes;
|
||||||
trace_rows_.push_back(r);
|
trace_rows_.push_back(r);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
@ -277,9 +401,12 @@ bool RouterHook::on_eval(ggml_tensor * t, bool ask) {
|
||||||
// ── stream: the routing nodes get the single-node barrier so we see the selected ids ──
|
// ── stream: the routing nodes get the single-node barrier so we see the selected ids ──
|
||||||
int il = -1;
|
int il = -1;
|
||||||
const bool is_topk = std::sscanf(t->name, "ffn_moe_topk-%d", &il) == 1 && il >= 0;
|
const bool is_topk = std::sscanf(t->name, "ffn_moe_topk-%d", &il) == 1 && il >= 0;
|
||||||
// Only a traced run asks for the weight nodes: each extra ask is another barrier.
|
// The weight nodes are asked for by a traced run, and by the drop policy, which decides on the
|
||||||
|
// weights the matmul will actually apply. Each extra ask is another barrier — a handful per MoE
|
||||||
|
// layer, on tensors of a few floats — so neither is on by default.
|
||||||
int wl = -1;
|
int wl = -1;
|
||||||
const bool is_weights = trace_on_ && match_weights(t->name, wl);
|
const bool want_weights = trace_on_ || drop_armed();
|
||||||
|
const bool is_weights = want_weights && match_weights(t->name, wl);
|
||||||
// The compute trace wants every node isolated — or, at layer granularity, only the first
|
// The compute trace wants every node isolated — or, at layer granularity, only the first
|
||||||
// node of each layer: the cursor advances on the ask stream (every node passes through
|
// node of each layer: the cursor advances on the ask stream (every node passes through
|
||||||
// here), so one isolation request per layer transition. Layerless names (embeddings, the
|
// here), so one isolation request per layer transition. Layerless names (embeddings, the
|
||||||
|
|
@ -302,15 +429,29 @@ bool RouterHook::on_eval(ggml_tensor * t, bool ask) {
|
||||||
}
|
}
|
||||||
|
|
||||||
// Weights follow their layer's topk, so the pending record is already open; keep the last
|
// Weights follow their layer's topk, so the pending record is already open; keep the last
|
||||||
// one offered (match_weights explains why) and let the flush read it.
|
// one offered (match_weights explains why) and let the flush read it. This runs BEFORE the drop
|
||||||
|
// policy edits the same tensor, so the trace records the routing the router produced, not the
|
||||||
|
// one the policy left behind — `dropped` is what says which is which.
|
||||||
if (is_weights && t->data && t->type == GGML_TYPE_F32 && pending_.layer == wl && pending_.nu > 0)
|
if (is_weights && t->data && t->type == GGML_TYPE_F32 && pending_.layer == wl && pending_.nu > 0)
|
||||||
gather_weights(t, pending_.nu, pending_.nt, pending_.weights);
|
gather_weights(t, pending_.nu, pending_.nt, pending_.weights);
|
||||||
|
|
||||||
|
// Learn which node ends this layer's weight chain, and — once known — use it as the point where
|
||||||
|
// the routing is decided: everything the drop policy needs is final here, and nothing has
|
||||||
|
// consumed it yet. The learning pass and the deferral are the same walk, so a layer whose chain
|
||||||
|
// shape the hook has not seen yet simply keeps the undropped behaviour.
|
||||||
|
if (is_weights && t->data && t->type == GGML_TYPE_F32 && drop_.layer == wl) {
|
||||||
|
chain_last_ = t->name;
|
||||||
|
if (drop_.deferred && wl >= 0 && wl < (int) term_node_.size() && term_node_[wl] == t->name) apply_drop(t);
|
||||||
|
}
|
||||||
|
|
||||||
if (source_ && is_topk && t->data && t->type == GGML_TYPE_I32) {
|
if (source_ && is_topk && t->data && t->type == GGML_TYPE_I32) {
|
||||||
// selected_experts is [n_expert_used, n_tokens] but a VIEW of the full argsort
|
// selected_experts is [n_expert_used, n_tokens] but a VIEW of the full argsort
|
||||||
// [n_expert, n_tokens]: its row stride is nb[1] (= n_expert*4), not
|
// [n_expert, n_tokens]: its row stride is nb[1] (= n_expert*4), not
|
||||||
// n_expert_used*4. Gather respecting the strides — a flat read would grab token
|
// n_expert_used*4. Gather respecting the strides — a flat read would grab token
|
||||||
// 0's sorted tail as token 1's experts, corrupting the KV cache.
|
// 0's sorted tail as token 1's experts, corrupting the KV cache.
|
||||||
|
// The previous layer's weight chain has been fully offered by now.
|
||||||
|
close_drop_layer();
|
||||||
|
|
||||||
gathered_.clear();
|
gathered_.clear();
|
||||||
const int nu = (int) t->ne[0], nt = (int) t->ne[1];
|
const int nu = (int) t->ne[0], nt = (int) t->ne[1];
|
||||||
for (int j = 0; j < nt; ++j)
|
for (int j = 0; j < nt; ++j)
|
||||||
|
|
@ -325,6 +466,7 @@ bool RouterHook::on_eval(ggml_tensor * t, bool ask) {
|
||||||
pending_.nt = nt;
|
pending_.nt = nt;
|
||||||
pending_.ids = gathered_;
|
pending_.ids = gathered_;
|
||||||
pending_.weights.clear();
|
pending_.weights.clear();
|
||||||
|
pending_.dropped.clear();
|
||||||
// Classify against the cache BEFORE load_layer makes these experts resident —
|
// Classify against the cache BEFORE load_layer makes these experts resident —
|
||||||
// afterwards everything reads as a hit. Settle landed prefetches first, or an expert
|
// afterwards everything reads as a hit. Settle landed prefetches first, or an expert
|
||||||
// a prefetch correctly guessed would be recorded as a miss.
|
// a prefetch correctly guessed would be recorded as a miss.
|
||||||
|
|
@ -333,7 +475,26 @@ bool RouterHook::on_eval(ggml_tensor * t, bool ask) {
|
||||||
source_->query_residency(il, gathered_.data(), (int) gathered_.size(), pending_.residency.data());
|
source_->query_residency(il, gathered_.data(), (int) gathered_.size(), pending_.residency.data());
|
||||||
}
|
}
|
||||||
|
|
||||||
source_->load_layer(il, gathered_.data(), (int) gathered_.size());
|
// Count what the ROUTER selected, here rather than inside apply_drop: a layer that is not
|
||||||
|
// deferred yet (the first graph, or a phase the policy is not armed for) still routed these
|
||||||
|
// experts, and a denominator that skipped them would report the drop rate as a fraction of
|
||||||
|
// the wrong thing.
|
||||||
|
if (drop_frac_ > 0.0f) experts_routed_ += (long long) gathered_.size();
|
||||||
|
|
||||||
|
// Open the layer for the drop policy. Deferring the load is only safe once this layer's
|
||||||
|
// terminal weight node is known — otherwise there is no callback left to decide in, and the
|
||||||
|
// expert matmul would run against slots nothing loaded. First graph of a run: load here.
|
||||||
|
const bool defer = drop_armed() && il >= 0 && il < (int) term_node_.size() && !term_node_[il].empty();
|
||||||
|
drop_.layer = il;
|
||||||
|
drop_.nu = nu;
|
||||||
|
drop_.nt = nt;
|
||||||
|
drop_.ids = t;
|
||||||
|
drop_.deferred = defer;
|
||||||
|
chain_last_.clear();
|
||||||
|
if (defer)
|
||||||
|
drop_ids_ = gathered_;
|
||||||
|
else
|
||||||
|
source_->load_layer(il, gathered_.data(), (int) gathered_.size());
|
||||||
|
|
||||||
// Temporal prefetch: hint the next K layers with what the PREVIOUS token routed there,
|
// Temporal prefetch: hint the next K layers with what the PREVIOUS token routed there,
|
||||||
// to be read on idle lanes while this layer computes; then record this layer's routing
|
// to be read on idle lanes while this layer computes; then record this layer's routing
|
||||||
|
|
|
||||||
|
|
@ -73,6 +73,22 @@ public:
|
||||||
// has returned. Off by default: asking for the extra nodes costs a barrier per layer.
|
// has returned. Off by default: asking for the extra nodes costs a barrier per layer.
|
||||||
void set_trace(bool on);
|
void set_trace(bool on);
|
||||||
|
|
||||||
|
// ── cache-aware expert dropping (lossy; see MoeStreamConfig::drop_cold_frac) ──────
|
||||||
|
// `frac` > 0 arms the policy: a routed expert that is a cache MISS and carries less than
|
||||||
|
// frac × (1/n_expert_used) of the routing's weight is discarded — not read, weight zeroed,
|
||||||
|
// its slot pointed at the routing's top-weighted expert so the matmul still reads memory
|
||||||
|
// that is certainly resident. `renorm` rescales the survivors to preserve the routing's
|
||||||
|
// total mass. Off (frac == 0) the hook behaves exactly as before, bit for bit.
|
||||||
|
void set_drop_policy(float frac, bool renorm, bool in_prefill);
|
||||||
|
|
||||||
|
// Which phase the batch being decoded belongs to (0 = prefill, 1 = decode). The drop policy
|
||||||
|
// is decode-only unless armed for prefill, and unlike the traces it must know this on every
|
||||||
|
// run, so it cannot ride on begin_trace_batch.
|
||||||
|
void set_batch_phase(int phase) { batch_phase_ = phase; }
|
||||||
|
|
||||||
|
long long experts_routed() const { return experts_routed_; }
|
||||||
|
long long experts_dropped() const { return experts_dropped_; }
|
||||||
|
|
||||||
// ── compute trace (diagnostics; see bmoe/decode_trace.h) ────────────────────────
|
// ── compute trace (diagnostics; see bmoe/decode_trace.h) ────────────────────────
|
||||||
// When on, the hook asks for EVERY node, which makes ggml compute and synchronize each one
|
// When on, the hook asks for EVERY node, which makes ggml compute and synchronize each one
|
||||||
// alone — so the wall delta between consecutive callbacks is that node's real compute time,
|
// alone — so the wall delta between consecutive callbacks is that node's real compute time,
|
||||||
|
|
@ -119,8 +135,12 @@ private:
|
||||||
std::vector<int32_t> ids;
|
std::vector<int32_t> ids;
|
||||||
std::vector<float> weights;
|
std::vector<float> weights;
|
||||||
std::vector<uint8_t> residency;
|
std::vector<uint8_t> residency;
|
||||||
|
std::vector<uint8_t> dropped; // set by the drop policy; all zero when it is off
|
||||||
};
|
};
|
||||||
void flush_pending();
|
void flush_pending();
|
||||||
|
bool drop_armed() const;
|
||||||
|
void apply_drop(ggml_tensor * weights);
|
||||||
|
void close_drop_layer();
|
||||||
void ctrace_close_segment(int interval_layer, const char * tail_name);
|
void ctrace_close_segment(int interval_layer, const char * tail_name);
|
||||||
|
|
||||||
// Stored by value, not by reference: the caller often constructs us from a temporary
|
// Stored by value, not by reference: the caller often constructs us from a temporary
|
||||||
|
|
@ -140,6 +160,36 @@ private:
|
||||||
int prefetch_layers_ = 0;
|
int prefetch_layers_ = 0;
|
||||||
std::vector<std::vector<int32_t>> prev_ids_;
|
std::vector<std::vector<int32_t>> prev_ids_;
|
||||||
|
|
||||||
|
// Cache-aware dropping. Inert unless drop_frac_ > 0.
|
||||||
|
//
|
||||||
|
// The decision needs the FINAL router weights, and those are produced several nodes after the
|
||||||
|
// topk that opens the layer — so load_layer() is postponed from the topk node to the terminal
|
||||||
|
// node of the layer's weight chain, and the ids/weights are edited there, before the expert
|
||||||
|
// matmul consumes either. Which node is terminal depends on the model's gating (norm, softmax,
|
||||||
|
// scaled, or none of them), so it is LEARNED from the graph rather than tabulated per
|
||||||
|
// architecture: term_node_[il] fills in on the first graph, and until it does the layer loads
|
||||||
|
// at its topk node undropped, exactly as with the policy off. That costs the first token of a
|
||||||
|
// run its dropping and nothing else.
|
||||||
|
float drop_frac_ = 0.0f;
|
||||||
|
bool drop_renorm_ = true;
|
||||||
|
bool drop_prefill_ = false;
|
||||||
|
int batch_phase_ = 1; // 0 prefill, 1 decode
|
||||||
|
long long experts_routed_ = 0, experts_dropped_ = 0;
|
||||||
|
|
||||||
|
struct PendingDrop {
|
||||||
|
int layer = -1;
|
||||||
|
int nu = 0, nt = 0;
|
||||||
|
ggml_tensor * ids = nullptr; // the topk view, rewritten in place for dropped slots
|
||||||
|
bool deferred = false; // true when load_layer is waiting for the terminal weight node
|
||||||
|
};
|
||||||
|
PendingDrop drop_;
|
||||||
|
std::vector<std::string> term_node_; // per layer, "" until learned
|
||||||
|
std::string chain_last_; // last weight node seen for drop_.layer while its chain runs
|
||||||
|
std::vector<int32_t> drop_ids_; // this layer's routed ids, kept across the deferral
|
||||||
|
std::vector<float> drop_w_; // scratch: the final weights
|
||||||
|
std::vector<uint8_t> drop_res_; // scratch: residency of each routed id
|
||||||
|
std::vector<uint8_t> drop_mask_; // scratch: which slots this layer dropped
|
||||||
|
|
||||||
// Route trace. All of this is inert unless trace_on_.
|
// Route trace. All of this is inert unless trace_on_.
|
||||||
bool trace_on_ = false;
|
bool trace_on_ = false;
|
||||||
int trace_base_pos_ = 0, trace_batch_n_ = 1, trace_phase_ = 0, trace_turn_ = 0;
|
int trace_base_pos_ = 0, trace_batch_n_ = 1, trace_phase_ = 0, trace_turn_ = 0;
|
||||||
|
|
|
||||||
|
|
@ -21,7 +21,8 @@ for the idea the project is built on.
|
||||||
| [telemetry.md](telemetry.md) | The `BMOE_*` line protocol and CSV schema — the integration contract. |
|
| [telemetry.md](telemetry.md) | The `BMOE_*` line protocol and CSV schema — the integration contract. |
|
||||||
| [session.md](session.md) | Session lifecycle, KV prefix reuse, cancellation. |
|
| [session.md](session.md) | Session lifecycle, KV prefix reuse, cancellation. |
|
||||||
| [cache-sizing.md](cache-sizing.md) | `--cache-mb auto`, the cache ceiling, and dense warm-up. |
|
| [cache-sizing.md](cache-sizing.md) | `--cache-mb auto`, the cache ceiling, and dense warm-up. |
|
||||||
| [prefetch.md](prefetch.md) | `--prefetch K`: the design and why it cannot change output. |
|
| [prefetch.md](prefetch.md) | `--prefetch K`: the design and why it cannot change output (with the lossy knobs off). |
|
||||||
|
| [expert-dropping.md](expert-dropping.md) | `--drop-cold-experts F`: spending quality only where it buys a flash read, and why it is the one setting whose output is not reproducible. |
|
||||||
| [android-memory.md](android-memory.md) | What reclaims the engine's memory on a phone, which levers exist (almost none), and why the cache hit rate is what the kernel judges you by. |
|
| [android-memory.md](android-memory.md) | What reclaims the engine's memory on a phone, which levers exist (almost none), and why the cache hit rate is what the kernel judges you by. |
|
||||||
| [pressure.md](pressure.md) | Cache policy under memory pressure: why an unaffordable budget starts a reclaim war, why the adaptive governor was retired, and what the fixed `--cache-mb` / `--dense-weights` levers do. |
|
| [pressure.md](pressure.md) | Cache policy under memory pressure: why an unaffordable budget starts a reclaim war, why the adaptive governor was retired, and what the fixed `--cache-mb` / `--dense-weights` levers do. |
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -45,8 +45,11 @@ Streaming experts serially needs three things from the inference engine. All thr
|
||||||
already public in llama.cpp:
|
already public in llama.cpp:
|
||||||
|
|
||||||
1. **A hook at routing time.** `llama_context_params.cb_eval` is called for every graph
|
1. **A hook at routing time.** `llama_context_params.cb_eval` is called for every graph
|
||||||
node. We ask for only the routing nodes (`ffn_moe_topk-<il>`); ggml computes and
|
node. We ask for the routing nodes (`ffn_moe_topk-<il>`); ggml computes and
|
||||||
synchronizes each alone, then calls us back with the selected expert ids materialized.
|
synchronizes each alone, then calls us back with the selected expert ids materialized.
|
||||||
|
The route trace and [cache-aware dropping](expert-dropping.md) additionally ask for each
|
||||||
|
layer's `ffn_moe_weights*-<il>` chain — and dropping is the one path that *writes into* a
|
||||||
|
graph tensor's contents rather than only rebinding `->data`. See [seam.md](seam.md).
|
||||||
2. **The expert tensor pointers.** During a one-token warm-up we scan each graph node's
|
2. **The expert tensor pointers.** During a one-token warm-up we scan each graph node's
|
||||||
sources for tensors named `blk.<il>.ffn_{gate,up,down}_exps.weight` and record the
|
sources for tensors named `blk.<il>.ffn_{gate,up,down}_exps.weight` and record the
|
||||||
live `ggml_tensor*`. We then rebind their `->data`.
|
live `ggml_tensor*`. We then rebind their `->data`.
|
||||||
|
|
@ -89,4 +92,6 @@ The composition root is `Session` (core/src/engine/session.cpp):
|
||||||
so the gates and the interactive session share the same code path.
|
so the gates and the interactive session share the same code path.
|
||||||
|
|
||||||
Greedy sampling makes the output a deterministic function of the graph — the property the
|
Greedy sampling makes the output a deterministic function of the graph — the property the
|
||||||
[byte-identity gates](../tests/moe_gates.cpp) assert.
|
[byte-identity gates](../tests/moe_gates.cpp) assert. That holds with the lossy knobs off. Under
|
||||||
|
[`--drop-cold-experts`](expert-dropping.md) the hook edits routing weights from live cache state,
|
||||||
|
which is not in the graph, so output becomes a function of the graph *and* the run's history.
|
||||||
|
|
|
||||||
|
|
@ -47,6 +47,7 @@ Vary one axis at a time:
|
||||||
| threads (-t) | 2, 4, 8 | U-shape, 4 optimal, 8 regresses |
|
| threads (-t) | 2, 4, 8 | U-shape, 4 optimal, 8 regresses |
|
||||||
| overlap | off, on | net gain **only over a warm cache** (hides residual flash wait behind compute); a net loss on a cold cache-0 stream, where I/O dwarfs compute |
|
| overlap | off, on | net gain **only over a warm cache** (hides residual flash wait behind compute); a net loss on a cold cache-0 stream, where I/O dwarfs compute |
|
||||||
| n-expert-used | default, 6 | fewer active experts cut compute + I/O ~linearly (8→6 ≈ −25%), changes the output |
|
| n-expert-used | default, 6 | fewer active experts cut compute + I/O ~linearly (8→6 ≈ −25%), changes the output |
|
||||||
|
| drop-cold-experts | off, 0.75, 1.0 | the second lossy axis, and the only **non-deterministic** one: what is skipped depends on cache state, so cells are noisier and the drop rate must be reported with the tok/s. Needs the cache on |
|
||||||
| dense-weights | warm, anon | decisive well past RAM, near-neutral near it: on gpt-oss (5.2× RAM) `anon` drops majflt/token from the hundreds to **6–10** and compute with it; on Qwen (1.64× RAM) there is little dense-fault pressure to remove. Watch `majflt/token`, not just tok/s |
|
| dense-weights | warm, anon | decisive well past RAM, near-neutral near it: on gpt-oss (5.2× RAM) `anon` drops majflt/token from the hundreds to **6–10** and compute with it; on Qwen (1.64× RAM) there is little dense-fault pressure to remove. Watch `majflt/token`, not just tok/s |
|
||||||
|
|
||||||
When sweeping `--n-expert-used`, run it as a **matched A/B against the model's own default**
|
When sweeping `--n-expert-used`, run it as a **matched A/B against the model's own default**
|
||||||
|
|
@ -125,6 +126,16 @@ So:
|
||||||
Re-run such a cell; do not publish it. And sanity-check any matrix by **reversing the run order** —
|
Re-run such a cell; do not publish it. And sanity-check any matrix by **reversing the run order** —
|
||||||
cells that move were measuring device state.
|
cells that move were measuring device state.
|
||||||
|
|
||||||
|
**The reversal check does not work under `--drop-cold-experts`.** There a cell can move because the
|
||||||
|
*drop rate* moved — the policy reads live cache state, so the same command legitimately discards a
|
||||||
|
different number of experts on a different run. That is the feature working, not the device
|
||||||
|
contaminating the cell, and the two tells above cannot tell them apart. Always record
|
||||||
|
`experts_dropped`/`experts_routed` (or the `moe-drop:` line) next to the tok/s: a dropping cell
|
||||||
|
without its drop rate is uninterpretable, because the flag fixes a threshold and not a rate. Note
|
||||||
|
also that a dropping run pays the same extra per-MoE-layer barriers a route-traced run does, so an
|
||||||
|
A/B against `--n-expert-used` is not overhead-matched — see
|
||||||
|
[expert-dropping.md](expert-dropping.md).
|
||||||
|
|
||||||
### Caveats
|
### Caveats
|
||||||
|
|
||||||
- **Thermal.** Sustained decode throttles. Warm up, then measure a steady window; discard
|
- **Thermal.** Sustained decode throttles. Warm up, then measure a steady window; discard
|
||||||
|
|
|
||||||
172
docs/expert-dropping.md
Normal file
172
docs/expert-dropping.md
Normal file
|
|
@ -0,0 +1,172 @@
|
||||||
|
# Cache-aware expert dropping
|
||||||
|
|
||||||
|
`--drop-cold-experts F` skips a routed expert when it is **not in the cache** *and* the router
|
||||||
|
weighted it below `F × (1 / top-k)` — that is, below `F` of the uniform share each of the `k`
|
||||||
|
selected experts would get if the router split its mass evenly. Off by default.
|
||||||
|
|
||||||
|
It is the second lossy knob in the engine, after
|
||||||
|
[turbo top-k](../README.md#turbo-top-k--the-measured-lossy-option), and it exists because the first one
|
||||||
|
spends quality in a place it does not have to.
|
||||||
|
|
||||||
|
## Why cache state belongs in the decision
|
||||||
|
|
||||||
|
`--n-expert-used k` drops the routing's tail unconditionally: slot 7 and slot 8 go, whether or not
|
||||||
|
they were already sitting in RAM. But an expert that is already resident costs **no flash read** —
|
||||||
|
and on a streamed decode, flash reads are what the token is waiting for
|
||||||
|
([decode is I/O-bound](benchmarks.md)). Dropping a resident expert pays quality for nothing.
|
||||||
|
|
||||||
|
Turn that around and the policy writes itself: **spend quality only where it buys I/O**. Keep every
|
||||||
|
resident expert however small its weight; consider dropping only the ones that would cost a read,
|
||||||
|
and only when the router says they barely matter.
|
||||||
|
|
||||||
|
## What it costs and what it buys
|
||||||
|
|
||||||
|
Replayed over the committed route traces (`docs/bench-data/2026-07-15-route-trace/`), decode phase,
|
||||||
|
threshold at the uniform share (`F = 1.0`):
|
||||||
|
|
||||||
|
| policy | flash reads avoided | router weight discarded |
|
||||||
|
|---|---|---|
|
||||||
|
| `--drop-cold-experts 1.0` | **66%** | **9.5%** |
|
||||||
|
| `--n-expert-used 5` | 23% | 10.6% |
|
||||||
|
| `--n-expert-used 3` | 59% | 36.8% |
|
||||||
|
|
||||||
|
(Qwen3-30B-A3B at k=6; Gemma-4-26B-A4B is within a point and a half on both columns: 67.4% / 8.2%. On gpt-oss-120b at k=2 the
|
||||||
|
policy matches `--n-expert-used 1`'s read saving while discarding 25% of the weight mass instead of
|
||||||
|
42%.)
|
||||||
|
|
||||||
|
At a comparable quality cost the cache-aware policy avoids roughly **three times** the reads. The
|
||||||
|
reason is visible in the third column of the trace: about 80% of decode routings are cache hits, and
|
||||||
|
the policy leaves every one of them alone.
|
||||||
|
|
||||||
|
`F` is a curve, not a switch. At `F = 0.75` the same model trades 4.4% of the weight mass for 37% of
|
||||||
|
the reads — still better than `--n-expert-used 5` on **both** axes.
|
||||||
|
|
||||||
|
These are replay numbers and an **upper bound**: skipping a read changes what the cache holds later,
|
||||||
|
so the real hit pattern drifts from the recorded one. The on-device A/B is what settles it.
|
||||||
|
|
||||||
|
## Two properties worth knowing
|
||||||
|
|
||||||
|
**A routing is never emptied.** The largest weight in a routing is always at least the uniform
|
||||||
|
share, so at `F ≤ 1.0` the top expert can never fall below the threshold. `validate()` rejects
|
||||||
|
`F > 1.0` for that reason, and the implementation additionally pins the top-weighted expert, so the
|
||||||
|
guarantee does not rest on the bound alone.
|
||||||
|
|
||||||
|
**It requires the expert cache.** With `--cache-mb 0` every expert reads as a miss, so the policy
|
||||||
|
would stop being cache-aware and become an unconditional weight cut — which is what
|
||||||
|
`--n-expert-used` already does, without claiming to consult residency. `validate()` rejects the
|
||||||
|
combination, the same way it rejects `--prefetch` without a cache.
|
||||||
|
|
||||||
|
**It changes what `--prefetch` means.** Speculation is normally output-neutral by construction. Here
|
||||||
|
residency is an *input* to the policy, so a correct guess un-drops an expert that would otherwise
|
||||||
|
have been discarded: prefetch depth becomes an output-affecting setting. The decision point also
|
||||||
|
settles pending speculation a few nodes after it was issued, which shortens the overlap window the
|
||||||
|
prefetch exists for — treat the two as interacting, not composable.
|
||||||
|
|
||||||
|
**Prefill is excluded by default.** With a cold cache almost every expert is a miss, and the same
|
||||||
|
threshold discards ~42% of the weight mass instead of ~9%. Prefill is compute-bound anyway, so there
|
||||||
|
is little to win. `--drop-in-prefill` arms it for experiments.
|
||||||
|
|
||||||
|
## The output is no longer reproducible
|
||||||
|
|
||||||
|
This is the real novelty, and the reason the flag is off by default and named the way it is.
|
||||||
|
|
||||||
|
`--n-expert-used` is lossy but **deterministic**: same prompt, same config, same tokens. Dropping is
|
||||||
|
lossy and **state-dependent** — what gets discarded depends on what the cache happened to hold,
|
||||||
|
which depends on everything decoded before it. The same prompt can produce different text across
|
||||||
|
runs, and a benchmark cell is noisier because the drop rate itself varies.
|
||||||
|
|
||||||
|
The greedy byte-identity gates therefore do not cover the policy's output, and cannot: there is
|
||||||
|
nothing stable to compare against. They cover the machinery instead (see below).
|
||||||
|
|
||||||
|
## How it is implemented
|
||||||
|
|
||||||
|
The decision needs the **final** router weights, and those are produced several graph nodes after
|
||||||
|
the topk node where the streamer normally loads. So with the policy armed, `load_layer()` is
|
||||||
|
postponed from the topk node to the terminal node of the layer's weight chain — the last node before
|
||||||
|
the expert matmul consumes either the ids or the weights.
|
||||||
|
|
||||||
|
Which node is terminal depends on the model's gating (`_norm`, `_softmax`, `_scaled`, or none), so
|
||||||
|
the hook **learns** it from the graph instead of carrying a per-architecture table: the first graph
|
||||||
|
of a run records the chain, and dropping starts from the second. A layer whose shape has not been
|
||||||
|
seen yet simply loads at its topk node, undropped. That costs a run its first token's dropping and
|
||||||
|
nothing else, and it keeps [hard rule 4](../CLAUDE.md) — no model-specific constants in the
|
||||||
|
streaming path.
|
||||||
|
|
||||||
|
At the decision point two edits happen, both before anything reads them:
|
||||||
|
|
||||||
|
1. the dropped slot's **weight is zeroed**, and with `drop_renorm` (default on) the survivors are
|
||||||
|
scaled so the routing keeps its original total mass;
|
||||||
|
2. the dropped slot's **id is repointed** at the routing's top-weighted expert.
|
||||||
|
|
||||||
|
The second edit is not cosmetic. An expert the engine declines to read may sit in a
|
||||||
|
reserved-but-uncommitted slot, and `mul_mat_id` would still touch it. Pointing the slot at an expert
|
||||||
|
that is certainly resident makes the kernel read valid memory and multiply it by exactly zero. It
|
||||||
|
costs a duplicate matmul — the right trade on a decode bound by flash rather than arithmetic.
|
||||||
|
|
||||||
|
Renormalisation matters more than it looks: without it the layer's expert output is systematically
|
||||||
|
scaled down by the discarded mass, a perturbation of the residual stream the model never sees in
|
||||||
|
training. `--drop-no-renorm` exists to A/B that claim.
|
||||||
|
|
||||||
|
Cost of the extra barriers: the policy asks for each layer's weight nodes, a handful more
|
||||||
|
synchronisation points per MoE layer on tensors of a few floats. The same asks a route trace makes.
|
||||||
|
|
||||||
|
## Measuring it
|
||||||
|
|
||||||
|
The engine reports what the policy actually did, which the flag alone cannot tell you — the
|
||||||
|
threshold is fixed, the drop rate is not:
|
||||||
|
|
||||||
|
```
|
||||||
|
moe-drop: <dropped>/<routed> routed experts dropped (<pct>%), threshold <F> x uniform
|
||||||
|
```
|
||||||
|
|
||||||
|
The route trace gains a `dropped` column: `weight` and `residency` stay as the **router** produced
|
||||||
|
them, `dropped` records what the policy then did, and `expert_bytes` is 0 for a dropped routing
|
||||||
|
because it costs no read. That is enough to replay a real run against the offline model and check
|
||||||
|
whether the upper bound held. See [telemetry.md](telemetry.md).
|
||||||
|
|
||||||
|
## Gates
|
||||||
|
|
||||||
|
`bmoe_moe_gates` covers the machinery, not the policy's output:
|
||||||
|
|
||||||
|
- **G8a** — with a threshold below any weight the router can produce, nothing is dropped and the
|
||||||
|
output is **byte-identical** to the undropped stream. This proves the deferral and the learned
|
||||||
|
terminal node are transparent, separating "the plumbing is correct" from "the policy is lossy" —
|
||||||
|
a regression in the first would otherwise hide behind the expected difference. **G8a'** asserts
|
||||||
|
the count separately (`experts_routed > 0`, `experts_dropped == 0`), so "a weight happened to fall
|
||||||
|
under the threshold" fails legibly instead of as a mysterious byte mismatch.
|
||||||
|
- **G8b** — at full strength against a cache small enough to be evicting constantly, so dropped
|
||||||
|
experts really do land on slots the cache has released. Generation still completes: the id
|
||||||
|
repointing means no matmul ever reads reserved-but-uncommitted memory. (The gates deliberately do
|
||||||
|
*not* run this with the cache off — there the shared-slot path has no uncommitted memory, so the
|
||||||
|
safety property the repointing exists for would go untested.)
|
||||||
|
- **G8c** — forcing top-k to 1 makes every routed expert the top one, so dropping must be a no-op at
|
||||||
|
any threshold and the output must match the undropped k=1 run byte for byte. This pins both the
|
||||||
|
top-expert guarantee and the fact that the threshold is taken against the **effective** top-k
|
||||||
|
discovered at runtime — a hardcoded width would not survive the override.
|
||||||
|
|
||||||
|
## Defaults, and where the numbers do and do not come from
|
||||||
|
|
||||||
|
The **CLI defaults it off**, and will keep doing so: the byte-identity gates need a deterministic
|
||||||
|
default, and an instrument should not quietly change the thing it measures.
|
||||||
|
|
||||||
|
The **app ships it at 75%** — under **Speed / quality → Drop cold experts**, with rungs 50 / 75 /
|
||||||
|
100 as percentages of the uniform share. It is disabled there in mmap mode and with the cache off,
|
||||||
|
the same two conditions `validate()` enforces.
|
||||||
|
|
||||||
|
That default is a product decision taken on the maintainer's own device measurement. **It is not
|
||||||
|
backed by a published benchmark in this repository**, and the tables in the README deliberately
|
||||||
|
carry no rows for it — they are a deterministic protocol and this knob is not deterministic. Nothing
|
||||||
|
here should be read as "75% is worth X%"; the honest claim is narrower: the replay above says the
|
||||||
|
shape of the trade is favourable, and the default was chosen after checking it on hardware.
|
||||||
|
|
||||||
|
What is still owed before this is recommended beyond that:
|
||||||
|
|
||||||
|
- a published decode A/B against `--n-expert-used` at matched tok/s, with the device state recorded
|
||||||
|
the way [benchmark-method.md](benchmark-method.md) requires;
|
||||||
|
- a quality comparison at that matched speed — the whole thesis is that this knob buys the same
|
||||||
|
throughput for less damage, and only a side-by-side can support it;
|
||||||
|
- a re-run of the replay against a real traced run with the `dropped` column, to see how far the
|
||||||
|
static upper bound overstated the win.
|
||||||
|
|
||||||
|
The [`layer-lfu` entry in the roadmap](roadmap.md) is the standing reminder for why the third one
|
||||||
|
matters: it simulated exactly as predicted and was ~30% slower in reality.
|
||||||
|
|
@ -18,6 +18,11 @@ serial path, and only a single ~25-line hook (with an explicit sunset) for the o
|
||||||
|
|
||||||
## Limitations
|
## Limitations
|
||||||
|
|
||||||
|
- **One setting makes output non-reproducible.** Every other knob is deterministic given a
|
||||||
|
configuration: `--n-expert-used` changes the output, but changes it the same way on every run.
|
||||||
|
[`--drop-cold-experts`](expert-dropping.md) decides per routing from live cache state, so the
|
||||||
|
same prompt and the same flags can decode differently run to run, and the byte-identity gates
|
||||||
|
cannot cover its output — only its machinery. Off by default in the CLI.
|
||||||
- **n=1 only.** The expert sparsity exists only for single-token decode, so streaming is
|
- **n=1 only.** The expert sparsity exists only for single-token decode, so streaming is
|
||||||
incompatible with speculative decoding or batching. Prefill streams the union of the
|
incompatible with speculative decoding or batching. Prefill streams the union of the
|
||||||
prompt's routed experts (still far below the full bank, but larger than one token's).
|
prompt's routed experts (still far below the full bank, but larger than one token's).
|
||||||
|
|
|
||||||
|
|
@ -32,7 +32,10 @@ Ordering is guaranteed by ggml's eval-callback loop: the node we mark is compute
|
||||||
buffers until this layer's matmul has synchronized. Correct on any backend.
|
buffers until this layer's matmul has synchronized. Correct on any backend.
|
||||||
|
|
||||||
The result is **lossless**: byte-identical to running with every expert resident, asserted
|
The result is **lossless**: byte-identical to running with every expert resident, asserted
|
||||||
by the gates.
|
by the gates. That is the streaming path itself; two opt-in knobs deliberately trade output for
|
||||||
|
speed on top of it — `--n-expert-used` (fewer experts per token) and
|
||||||
|
[`--drop-cold-experts`](expert-dropping.md) (skip an expert that would cost a read and was barely
|
||||||
|
weighted). Both are off unless asked for, which is what keeps the sentence above true by default.
|
||||||
|
|
||||||
## Residency modes
|
## Residency modes
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -32,7 +32,11 @@ captures most of the benefit. Prefetch requires the LRU cache to be on — eithe
|
||||||
|
|
||||||
## How it stays correct and out of the way
|
## How it stays correct and out of the way
|
||||||
|
|
||||||
The speculative path never delays real work and never changes output:
|
The speculative path never delays real work and never changes output — with the lossy knobs off.
|
||||||
|
(Under [`--drop-cold-experts`](expert-dropping.md) residency is an *input* to the routing policy,
|
||||||
|
so a correct guess un-drops an expert that would otherwise have been discarded. Prefetch depth
|
||||||
|
becomes output-affecting there; everything below still holds for the bytes themselves.)
|
||||||
|
|
||||||
|
|
||||||
- **Same bytes.** A speculative read is the *identical* read a real miss would issue — same file
|
- **Same bytes.** A speculative read is the *identical* read a real miss would issue — same file
|
||||||
offset, same destination buffer (`lbuf_[p][il] + e*slice`). A prefetched expert is therefore
|
offset, same destination buffer (`lbuf_[p][il] + e*slice`). A prefetched expert is therefore
|
||||||
|
|
|
||||||
|
|
@ -30,7 +30,10 @@ is:
|
||||||
|
|
||||||
## Why a budget cannot be a constant
|
## Why a budget cannot be a constant
|
||||||
|
|
||||||
The expert cache is the one lever that trades RAM for flash reads, so the temptation is to set it as
|
The expert cache is the one lever that trades RAM for flash reads (
|
||||||
|
[`--drop-cold-experts`](expert-dropping.md) is the other kind of trade — quality for flash reads —
|
||||||
|
and the two interact: a squeezed cache raises the miss rate, which raises the drop rate, so memory
|
||||||
|
pressure degrades output quality there instead of only throughput). The temptation is to set it as
|
||||||
large as the device seems to allow. On a phone that is the wrong shape of decision, for three
|
large as the device seems to allow. On a phone that is the wrong shape of decision, for three
|
||||||
reasons that are measured rather than argued:
|
reasons that are measured rather than argued:
|
||||||
|
|
||||||
|
|
@ -115,6 +118,11 @@ not a floor), `layer_demand_MiB` (the mechanical floor), `cache_budget_MiB` (the
|
||||||
effect), `cache_hit_pct`. Per token, `dense_resident_frac` says whether the dense set is holding in
|
effect), `cache_hit_pct`. Per token, `dense_resident_frac` says whether the dense set is holding in
|
||||||
RAM (the live signal now that the cache-residency governor sensor is gone).
|
RAM (the live signal now that the cache-residency governor sensor is gone).
|
||||||
|
|
||||||
|
This sizing procedure assumes dropping is off. With
|
||||||
|
[`--drop-cold-experts`](expert-dropping.md) on, dropped routings are misses that never reach the
|
||||||
|
cache, so `cache_hit_pct` reads high and `token_demand_MiB` reads low for the same budget — size
|
||||||
|
the cache first, then turn dropping on.
|
||||||
|
|
||||||
Reading `cache_hit_pct` against `token_demand_MiB` is how you tell whether a fixed `--cache-mb N` is
|
Reading `cache_hit_pct` against `token_demand_MiB` is how you tell whether a fixed `--cache-mb N` is
|
||||||
earning its RAM: a budget near or below one token's demand holds no history between tokens and its
|
earning its RAM: a budget near or below one token's demand holds no history between tokens and its
|
||||||
hits are only inter-token correlation; well above it, a high hit rate means real reuse.
|
hits are only inter-token correlation; well above it, a high hit rate means real reuse.
|
||||||
|
|
|
||||||
|
|
@ -113,6 +113,19 @@ routed) are supported; other `build_moe_ffn` models are one recipe row each. The
|
||||||
frontier is architectures whose routing node is not the shared `ffn_moe_topk` — custom gating,
|
frontier is architectures whose routing node is not the shared `ffn_moe_topk` — custom gating,
|
||||||
which the capture/stream hook would need to learn. See [adding-a-model.md](adding-a-model.md).
|
which the capture/stream hook would need to learn. See [adding-a-model.md](adding-a-model.md).
|
||||||
|
|
||||||
|
## Skipping reads the router barely wants — built, unmeasured
|
||||||
|
|
||||||
|
`--drop-cold-experts` ([expert-dropping.md](expert-dropping.md)) is the first lever that treats
|
||||||
|
quality and I/O as a *joint* budget rather than two separate knobs: an expert already in the cache
|
||||||
|
runs however small its weight, and only a routing that would cost a flash read can be dropped. On
|
||||||
|
the recorded traces that is worth ~3× the reads of turbo top-k for a comparable weight cost, which
|
||||||
|
is the strongest offline case any remaining lever has shown.
|
||||||
|
|
||||||
|
What it does **not** have is a device measurement, and the previous entry on this page is the reason
|
||||||
|
that matters: `layer-lfu` simulated well and was ~30% slower in reality. The open questions are the
|
||||||
|
device A/B against turbo top-k at matched throughput, the quality comparison at that speed, and how
|
||||||
|
far the static replay overstated the win once dropping starts changing what the cache holds.
|
||||||
|
|
||||||
## Expert quantization on the fly
|
## Expert quantization on the fly
|
||||||
|
|
||||||
Storing streamed experts at a lower precision than the resident parts to cut read volume, if it
|
Storing streamed experts at a lower precision than the resident parts to cut read volume, if it
|
||||||
|
|
|
||||||
14
docs/seam.md
14
docs/seam.md
|
|
@ -24,10 +24,22 @@ come from the arch's recipe — `ffn_{gate,up,down}_exps` for the split layout,
|
||||||
throughout — capture observes, it does not isolate. `ggml_tensor` is a public struct, so
|
throughout — capture observes, it does not isolate. `ggml_tensor` is a public struct, so
|
||||||
reading `->name`, `->ne`, `->nb` and writing `->data` is public API surface.
|
reading `->name`, `->ne`, `->nb` and writing `->data` is public API surface.
|
||||||
|
|
||||||
**Stream phase** (real generation). We return true only for `ffn_moe_topk-<il>`. The
|
**Stream phase** (real generation). We return true for `ffn_moe_topk-<il>`. The
|
||||||
non-ask callback then hands us that node with the selected expert ids materialized; we
|
non-ask callback then hands us that node with the selected expert ids materialized; we
|
||||||
gather them (stride-aware) and trigger the slice reads.
|
gather them (stride-aware) and trigger the slice reads.
|
||||||
|
|
||||||
|
Two optional jobs ask for more: the route trace and
|
||||||
|
[cache-aware dropping](expert-dropping.md) also want each layer's `ffn_moe_weights*-<il>` chain,
|
||||||
|
which is another barrier per node but no new kind of access — same public struct, same read of
|
||||||
|
`->data`.
|
||||||
|
|
||||||
|
Dropping does go one step further, and it is the only place the engine **writes into** a graph
|
||||||
|
tensor's contents rather than repointing `->data` at its own buffer: at the terminal node of the
|
||||||
|
weight chain it zeroes a dropped slot's weight and repoints that slot's expert id. Both tensors are
|
||||||
|
scratch the graph produced and has not yet consumed, so this alters the values flowing through the
|
||||||
|
run — deliberately, that is what the lossy policy *is* — and never llama.cpp's own state, its
|
||||||
|
weights, or its control flow. It stays inside the same callback contract; nothing is patched.
|
||||||
|
|
||||||
## 2. gguf offsets
|
## 2. gguf offsets
|
||||||
|
|
||||||
`gguf_init_from_file(..., no_alloc=true)` + `gguf_get_data_offset` +
|
`gguf_init_from_file(..., no_alloc=true)` + `gguf_get_data_offset` +
|
||||||
|
|
|
||||||
|
|
@ -51,6 +51,10 @@ BMOE_PROGRESS {"step":<int>,"steps":<int>,"wall_ms":<float>,"io_ms":<float>,
|
||||||
can be a large share of the token; at steady state it is near zero. Surfacing it stops the "all
|
can be a large share of the token; at steady state it is near zero. Surfacing it stops the "all
|
||||||
compute" reading on warm-up tokens where the real cost is cache churn, not matmul.
|
compute" reading on warm-up tokens where the real cost is cache churn, not matmul.
|
||||||
- `cache_hit_pct` is the cumulative cache hit rate, or `-1` when no cache is used.
|
- `cache_hit_pct` is the cumulative cache hit rate, or `-1` when no cache is used.
|
||||||
|
**Under [`--drop-cold-experts`](expert-dropping.md) read it with care:** a dropped routing is a
|
||||||
|
miss that is never looked up, so it leaves both sides of the ratio and the reported hit rate
|
||||||
|
rises without the cache having served anything more. Compare runs at the same drop rate, or read
|
||||||
|
`experts_dropped` next to it.
|
||||||
- `majflt` / `cpu_ms` **decompose the `compute_ms` residual** — the whole point being that "compute"
|
- `majflt` / `cpu_ms` **decompose the `compute_ms` residual** — the whole point being that "compute"
|
||||||
above is a catch-all that silently absorbs page faults and scheduler stalls, not just matmul.
|
above is a catch-all that silently absorbs page faults and scheduler stalls, not just matmul.
|
||||||
They are measured directly around `llama_decode` (no submodule patch needed): `majflt` is the
|
They are measured directly around `llama_decode` (no submodule patch needed): `majflt` is the
|
||||||
|
|
@ -101,6 +105,16 @@ moe-prefetch: <mib> MiB speculative, <useful>/<prefetched> experts useful (<pct>
|
||||||
`<prefetched>` the experts fully read ahead, and `<useful>` how many of those a later routing
|
`<prefetched>` the experts fully read ahead, and `<useful>` how many of those a later routing
|
||||||
actually hit. See [prefetch.md](prefetch.md).
|
actually hit. See [prefetch.md](prefetch.md).
|
||||||
|
|
||||||
|
With `--drop-cold-experts F` a `moe-drop:` line is added:
|
||||||
|
|
||||||
|
```
|
||||||
|
moe-drop: <dropped>/<routed> routed experts dropped (<pct>%), threshold <F> x uniform
|
||||||
|
```
|
||||||
|
|
||||||
|
The flag fixes a *threshold*, not a rate: how much is actually discarded depends on what the cache
|
||||||
|
held, so this line — not the flag — is what a run traded. See
|
||||||
|
[expert-dropping.md](expert-dropping.md).
|
||||||
|
|
||||||
Under `--overlap` the `moe-stream:` line additionally reports `stall_s/tok=<s>` — the mean
|
Under `--overlap` the `moe-stream:` line additionally reports `stall_s/tok=<s>` — the mean
|
||||||
wall time per token that compute threads waited for expert reads to complete. It is `0` in
|
wall time per token that compute threads waited for expert reads to complete. It is `0` in
|
||||||
serial mode (where the read wait is already folded into decode time).
|
serial mode (where the read wait is already folded into decode time).
|
||||||
|
|
@ -130,7 +144,10 @@ sampled dense-weight residency, `-1` when unmeasured. All are additive: older CS
|
||||||
so consumers must read by column NAME (from the header row) and treat any as optional. The `# summary`
|
so consumers must read by column NAME (from the header row) and treat any as optional. The `# summary`
|
||||||
line likewise gains `stall_s/tok=<s>`, `mgmt_s/tok=<s>`, `majflt/tok=<f>`, `cpu_s/tok=<s>`,
|
line likewise gains `stall_s/tok=<s>`, `mgmt_s/tok=<s>`, `majflt/tok=<f>`, `cpu_s/tok=<s>`,
|
||||||
`token_demand_MiB=<f>` (the expert bytes one token routes, measured — where cache hits start, NOT a
|
`token_demand_MiB=<f>` (the expert bytes one token routes, measured — where cache hits start, NOT a
|
||||||
floor to defend; see [pressure.md](pressure.md)) and `layer_demand_MiB=<f>` (the widest layer's routed
|
floor to defend; see [pressure.md](pressure.md)), `experts_routed=<n>` / `experts_dropped=<n>` (what
|
||||||
|
[cache-aware dropping](expert-dropping.md) actually discarded during generation — the flag sets a
|
||||||
|
threshold, not a rate, so this is the only record of the trade a run made) and
|
||||||
|
`layer_demand_MiB=<f>` (the widest layer's routed
|
||||||
bytes: the mechanical floor the cache must be able to stage); see the `io_ms` note above for how the
|
bytes: the mechanical floor the cache must be able to stage); see the `io_ms` note above for how the
|
||||||
read-time columns are reinterpreted under overlap.
|
read-time columns are reinterpreted under overlap.
|
||||||
|
|
||||||
|
|
@ -161,6 +178,10 @@ per routed expert. **A traced run is not a benchmark run** — the numbers in th
|
||||||
traced run are slower than the real thing, and `mgmt_ms` in particular shifts, because settling
|
traced run are slower than the real thing, and `mgmt_ms` in particular shifts, because settling
|
||||||
speculative prefetch moves outside the window that times it.
|
speculative prefetch moves outside the window that times it.
|
||||||
|
|
||||||
|
Columns are **append-only** within `v1`, like the metrics CSV: `dropped` was added after
|
||||||
|
`expert_bytes`, so consumers must read by column NAME and treat any column as optional rather than
|
||||||
|
indexing by position.
|
||||||
|
|
||||||
The file is long format: a `#` preamble carrying the run's static facts, then one row per routed
|
The file is long format: a `#` preamble carrying the run's static facts, then one row per routed
|
||||||
expert. Conceptually it is a matrix — rows are steps, columns are layers — and a **cell** is the
|
expert. Conceptually it is a matrix — rows are steps, columns are layers — and a **cell** is the
|
||||||
`n_expert_used` rows sharing `(turn, phase, step, layer)`.
|
`n_expert_used` rows sharing `(turn, phase, step, layer)`.
|
||||||
|
|
@ -169,7 +190,7 @@ expert. Conceptually it is a matrix — rows are steps, columns are layers — a
|
||||||
# route_trace v1
|
# route_trace v1
|
||||||
# model=<path> arch=<string> n_layer=<int> n_expert=<int> n_expert_used=<int>
|
# model=<path> arch=<string> n_layer=<int> n_expert=<int> n_expert_used=<int>
|
||||||
# layer=<int> expert_bytes=<int> dense_bytes=<int> (one per layer)
|
# layer=<int> expert_bytes=<int> dense_bytes=<int> (one per layer)
|
||||||
turn,phase,step,layer,slot,expert,weight,residency,expert_bytes
|
turn,phase,step,layer,slot,expert,weight,residency,expert_bytes,dropped
|
||||||
```
|
```
|
||||||
|
|
||||||
| column | meaning |
|
| column | meaning |
|
||||||
|
|
@ -183,6 +204,7 @@ turn,phase,step,layer,slot,expert,weight,residency,expert_bytes
|
||||||
| `weight` | the final applied routing weight, after whatever softmax/normalise/scale the architecture uses. `nan` when the graph exposed no weight node — "unknown", never `0`. |
|
| `weight` | the final applied routing weight, after whatever softmax/normalise/scale the architecture uses. `nan` when the graph exposed no weight node — "unknown", never `0`. |
|
||||||
| `residency` | `0` = miss (this routing reads from flash), `1` = hit, `2` = hit on a speculative prefetch's first touch. |
|
| `residency` | `0` = miss (this routing reads from flash), `1` = hit, `2` = hit on a speculative prefetch's first touch. |
|
||||||
| `expert_bytes` | flash bytes this routing reads; `0` unless `residency=0`. |
|
| `expert_bytes` | flash bytes this routing reads; `0` unless `residency=0`. |
|
||||||
|
| `dropped` | `1` when [cache-aware dropping](expert-dropping.md) discarded this routing — a miss weighted below the threshold, never read, weight zeroed. Always `0` with `--drop-cold-experts` off. |
|
||||||
|
|
||||||
`(turn, phase, step, layer, slot)` is unique. Two asymmetries are deliberate:
|
`(turn, phase, step, layer, slot)` is unique. Two asymmetries are deliberate:
|
||||||
|
|
||||||
|
|
@ -195,6 +217,11 @@ turn,phase,step,layer,slot,expert,weight,residency,expert_bytes
|
||||||
streamed, so there is nothing to measure per step: `dense_bytes` is what a cold layer costs to
|
streamed, so there is nothing to measure per step: `dense_bytes` is what a cold layer costs to
|
||||||
page in, stated once. Per-layer *I/O time* is absent for the same kind of reason — under
|
page in, stated once. Per-layer *I/O time* is absent for the same kind of reason — under
|
||||||
`--overlap` reads complete asynchronously, so any per-layer timing would be fiction.
|
`--overlap` reads complete asynchronously, so any per-layer timing would be fiction.
|
||||||
|
- **`weight` and `residency` describe the router; `dropped` describes the policy.** When dropping is
|
||||||
|
on, a discarded routing keeps the weight the router gave it and the residency it faced — the trace
|
||||||
|
records the routing that was *chosen* — while `expert_bytes` falls to `0`, because a dropped
|
||||||
|
expert is never read. Summing `expert_bytes` therefore still measures real flash traffic, and
|
||||||
|
`dropped` is what explains the gap against `residency==0`.
|
||||||
|
|
||||||
**The last layer has only one prefill step, and that is real.** Before the final layer's FFN,
|
**The last layer has only one prefill step, and that is real.** Before the final layer's FFN,
|
||||||
llama.cpp gathers only the tokens whose logits were asked for (`inp_out_ids`; see `il == n_layer
|
llama.cpp gathers only the tokens whose logits were asked for (`inp_out_ids`; see `il == n_layer
|
||||||
|
|
|
||||||
|
|
@ -34,6 +34,10 @@ data class AppSettings(
|
||||||
val overlap: Boolean = true, // read the next experts while the current layer computes
|
val overlap: Boolean = true, // read the next experts while the current layer computes
|
||||||
val denseWeights: DenseWeights = DenseWeights.ANON, // dense (non-expert) weight residency policy
|
val denseWeights: DenseWeights = DenseWeights.ANON, // dense (non-expert) weight residency policy
|
||||||
val prefetchLayers: Int = 0, // temporal prefetch depth K (0 = off); needs the cache
|
val prefetchLayers: Int = 0, // temporal prefetch depth K (0 = off); needs the cache
|
||||||
|
// Cache-aware expert dropping, as a PERCENTAGE of the uniform share 1/top-k (0 = off, 100 = the
|
||||||
|
// share itself). Stored as an Int because the settings are integer rungs; the flag takes a
|
||||||
|
// fraction. LOSSY and cache-dependent — it changes the output, and not reproducibly.
|
||||||
|
val dropColdPct: Int = 75,
|
||||||
val thinking: Boolean = false, // reasoning; off passes --no-think (enable_thinking=false)
|
val thinking: Boolean = false, // reasoning; off passes --no-think (enable_thinking=false)
|
||||||
val metricsCsv: Boolean = true, // write the engine's per-token CSV for this session (--csv)
|
val metricsCsv: Boolean = true, // write the engine's per-token CSV for this session (--csv)
|
||||||
) {
|
) {
|
||||||
|
|
@ -89,6 +93,11 @@ data class AppSettings(
|
||||||
// Auto sizing is a live LRU cache, so it satisfies the prefetch cache requirement.
|
// Auto sizing is a live LRU cache, so it satisfies the prefetch cache requirement.
|
||||||
val cacheOn = cacheMb == CACHE_AUTO || cacheMb > 0
|
val cacheOn = cacheMb == CACHE_AUTO || cacheMb > 0
|
||||||
if (prefetchLayers > 0 && cacheOn) a += listOf("--prefetch", prefetchLayers.toString())
|
if (prefetchLayers > 0 && cacheOn) a += listOf("--prefetch", prefetchLayers.toString())
|
||||||
|
// Cache-aware dropping needs a live cache to ask about residency — with the cache off
|
||||||
|
// every expert reads as a miss and the engine rejects the combination outright, so the
|
||||||
|
// same cacheOn condition that guards prefetch guards this. The engine takes a fraction
|
||||||
|
// of the uniform share; the setting is stored as a percentage.
|
||||||
|
if (dropColdPct > 0 && cacheOn) a += listOf("--drop-cold-experts", (dropColdPct / 100.0).toString())
|
||||||
}
|
}
|
||||||
return a
|
return a
|
||||||
}
|
}
|
||||||
|
|
@ -101,7 +110,7 @@ data class AppSettings(
|
||||||
*/
|
*/
|
||||||
fun sessionSignature(modelPath: String): String =
|
fun sessionSignature(modelPath: String): String =
|
||||||
listOf(modelPath, mmap, cacheMb, cacheCeilMb, ioThreads, threads, nExpertUsed, oDirect,
|
listOf(modelPath, mmap, cacheMb, cacheCeilMb, ioThreads, threads, nExpertUsed, oDirect,
|
||||||
overlap, denseWeights, prefetchLayers)
|
overlap, denseWeights, prefetchLayers, dropColdPct)
|
||||||
.joinToString("|")
|
.joinToString("|")
|
||||||
|
|
||||||
fun save(ctx: Context) {
|
fun save(ctx: Context) {
|
||||||
|
|
@ -114,6 +123,7 @@ data class AppSettings(
|
||||||
.putBoolean("overlap", overlap)
|
.putBoolean("overlap", overlap)
|
||||||
.putString("denseWeights", denseWeights.name)
|
.putString("denseWeights", denseWeights.name)
|
||||||
.putInt("prefetchLayers", prefetchLayers)
|
.putInt("prefetchLayers", prefetchLayers)
|
||||||
|
.putInt("dropColdPct", dropColdPct)
|
||||||
.putBoolean("thinking", thinking)
|
.putBoolean("thinking", thinking)
|
||||||
.putBoolean("metricsCsv", metricsCsv)
|
.putBoolean("metricsCsv", metricsCsv)
|
||||||
.apply()
|
.apply()
|
||||||
|
|
@ -178,6 +188,10 @@ data class AppSettings(
|
||||||
// 0 = model default (top-k as trained). 6/4/3/2 trade output quality for tok/s (fewer routed experts).
|
// 0 = model default (top-k as trained). 6/4/3/2 trade output quality for tok/s (fewer routed experts).
|
||||||
val N_EXPERT_CHOICES = intArrayOf(0, 6, 4, 3, 2)
|
val N_EXPERT_CHOICES = intArrayOf(0, 6, 4, 3, 2)
|
||||||
val PREFETCH_CHOICES = intArrayOf(0, 1, 2, 4)
|
val PREFETCH_CHOICES = intArrayOf(0, 1, 2, 4)
|
||||||
|
// Percent of the uniform share 1/top-k. 100 is the share itself and the useful maximum:
|
||||||
|
// above it the threshold could exceed every weight in a routing. The rungs below it are the
|
||||||
|
// conservative half of the curve, where the replay already beats a top-k cut on both axes.
|
||||||
|
val DROP_COLD_CHOICES = intArrayOf(0, 50, 75, 100)
|
||||||
val THREAD_CHOICES = intArrayOf(2, 4, 6, 8)
|
val THREAD_CHOICES = intArrayOf(2, 4, 6, 8)
|
||||||
val NPREDICT_CHOICES = intArrayOf(16, 32, 48, 64, 128, 256, 512, 1024, 2048)
|
val NPREDICT_CHOICES = intArrayOf(16, 32, 48, 64, 128, 256, 512, 1024, 2048)
|
||||||
|
|
||||||
|
|
@ -206,6 +220,7 @@ data class AppSettings(
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
prefetchLayers = p.getInt("prefetchLayers", d.prefetchLayers),
|
prefetchLayers = p.getInt("prefetchLayers", d.prefetchLayers),
|
||||||
|
dropColdPct = p.getInt("dropColdPct", d.dropColdPct),
|
||||||
thinking = p.getBoolean("thinking", d.thinking),
|
thinking = p.getBoolean("thinking", d.thinking),
|
||||||
metricsCsv = p.getBoolean("metricsCsv", d.metricsCsv),
|
metricsCsv = p.getBoolean("metricsCsv", d.metricsCsv),
|
||||||
)
|
)
|
||||||
|
|
|
||||||
|
|
@ -144,6 +144,25 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
|
||||||
"but the output changes — a speed/quality trade-off.",
|
"but the output changes — a speed/quality trade-off.",
|
||||||
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
|
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||||
)
|
)
|
||||||
|
IntSetting(
|
||||||
|
"Drop cold experts (% of even share)", AppSettings.DROP_COLD_CHOICES, current.dropColdPct,
|
||||||
|
format = { if (it == 0) "off" else "$it%" },
|
||||||
|
// Unlike top-k, this one asks the expert source what is resident, so it needs
|
||||||
|
// both the streamer and a live cache — the same condition prefetch is under.
|
||||||
|
enabled = !current.mmap &&
|
||||||
|
(current.cacheMb == AppSettings.CACHE_AUTO || current.cacheMb > 0),
|
||||||
|
) { onChange(current.copy(dropColdPct = it)) }
|
||||||
|
Text(
|
||||||
|
"Experimental. What slows a token down is reading an expert that is not already in RAM. " +
|
||||||
|
"This skips such an expert when the router barely wanted it anyway — below the chosen " +
|
||||||
|
"share of an even split. With 8 active experts an even split is 12.5% each, so 75% " +
|
||||||
|
"means \"skip it if it carries less than 9.4% of the routing\".\n\n" +
|
||||||
|
"Experts already in RAM always run, however small their weight: they cost no read. The " +
|
||||||
|
"strongest expert of each routing is never skipped.\n\n" +
|
||||||
|
"Higher is faster and rougher. Like Active experts, the reply changes — but unlike it, " +
|
||||||
|
"not the same way twice: what gets skipped depends on what the cache happened to hold.",
|
||||||
|
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||||
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
Section("Compute") {
|
Section("Compute") {
|
||||||
|
|
|
||||||
|
|
@ -52,6 +52,7 @@ class Trace:
|
||||||
self.expert_bytes = {} # layer -> bytes of one expert
|
self.expert_bytes = {} # layer -> bytes of one expert
|
||||||
self.dense_bytes = {} # layer -> non-streamed bytes
|
self.dense_bytes = {} # layer -> non-streamed bytes
|
||||||
self.rows = [] # (step, layer, slot, expert, weight, residency, ebytes)
|
self.rows = [] # (step, layer, slot, expert, weight, residency, ebytes)
|
||||||
|
self.n_dropped = 0 # routings discarded by --drop-cold-experts (0 unless it was on)
|
||||||
self._read(path, phase, turn)
|
self._read(path, phase, turn)
|
||||||
|
|
||||||
def _read(self, path, phase, turn):
|
def _read(self, path, phase, turn):
|
||||||
|
|
@ -72,6 +73,12 @@ class Trace:
|
||||||
continue
|
continue
|
||||||
w = float("nan") if p[6] == "nan" else float(p[6])
|
w = float("nan") if p[6] == "nan" else float(p[6])
|
||||||
self.rows.append((int(p[2]), int(p[3]), int(p[4]), int(p[5]), w, int(p[7]), int(p[8])))
|
self.rows.append((int(p[2]), int(p[3]), int(p[4]), int(p[5]), w, int(p[7]), int(p[8])))
|
||||||
|
# `dropped` (appended after expert_bytes) marks a routing cache-aware dropping
|
||||||
|
# discarded: the router asked for it, the engine never read it. Counted, not folded
|
||||||
|
# into the rows, so every figure below keeps meaning "what the ROUTER asked for" —
|
||||||
|
# see docs/expert-dropping.md for why the two diverge.
|
||||||
|
if len(p) > 9 and p[9].strip() not in ("", "0"):
|
||||||
|
self.n_dropped += 1
|
||||||
|
|
||||||
def _preamble(self, line):
|
def _preamble(self, line):
|
||||||
kv = kv_tokens(line)
|
kv = kv_tokens(line)
|
||||||
|
|
@ -391,6 +398,13 @@ def main():
|
||||||
print("arch %s n_layer %d n_expert %d n_expert_used %d"
|
print("arch %s n_layer %d n_expert %d n_expert_used %d"
|
||||||
% (tr.arch, tr.n_layer, tr.n_expert, tr.n_expert_used))
|
% (tr.arch, tr.n_layer, tr.n_expert, tr.n_expert_used))
|
||||||
print("phase %s rows %d" % (args.phase, len(tr.rows)))
|
print("phase %s rows %d" % (args.phase, len(tr.rows)))
|
||||||
|
if tr.n_dropped:
|
||||||
|
# Say it once, loudly: on such a trace "routed" and "read" are no longer the same set, and
|
||||||
|
# every reuse/working-set figure below describes the router's demand, not flash traffic.
|
||||||
|
print("NOTE cache-aware dropping was ON: %d of %d routings (%.1f%%) were discarded and "
|
||||||
|
"never read.\n Figures below are what the ROUTER asked for; use "
|
||||||
|
"scripts/route-drop-replay.py for what it cost."
|
||||||
|
% (tr.n_dropped, len(tr.rows), 100.0 * tr.n_dropped / len(tr.rows)))
|
||||||
|
|
||||||
names = DEFAULT_VIEWS if args.view == "default" else (
|
names = DEFAULT_VIEWS if args.view == "default" else (
|
||||||
list(VIEWS) if args.view == "all" else [args.view])
|
list(VIEWS) if args.view == "all" else [args.view])
|
||||||
|
|
|
||||||
159
scripts/route-drop-replay.py
Normal file
159
scripts/route-drop-replay.py
Normal file
|
|
@ -0,0 +1,159 @@
|
||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Replay a route trace against the cache-aware expert-dropping policy (docs/expert-dropping.md).
|
||||||
|
|
||||||
|
The policy skips a routed expert when it is a cache MISS and the router weighted it below
|
||||||
|
`frac x (1 / n_expert_used)`. A resident expert costs no flash read, so it is never dropped:
|
||||||
|
quality is spent only where it buys I/O. This script answers, per threshold, what that trade
|
||||||
|
would have been on an already-recorded run:
|
||||||
|
|
||||||
|
io_saved fraction of MISS BYTES the policy never reads -- the win
|
||||||
|
mass_lost fraction of total router weight discarded -- the proxy for the damage
|
||||||
|
|
||||||
|
The static-k baseline (`--n-expert-used`) is replayed on the same rows, because the only question
|
||||||
|
that matters is comparative: at equal io_saved, which policy discards less weight?
|
||||||
|
|
||||||
|
Two limits, both deliberate:
|
||||||
|
* This is a STATIC replay. Skipping a read changes what the cache holds later, so the real hit
|
||||||
|
pattern drifts from the recorded one. io_saved is an UPPER BOUND, not a prediction.
|
||||||
|
* mass_lost is a proxy. It says how much of the router's mass went away, not what that did to
|
||||||
|
the output. Only a quality A/B answers that.
|
||||||
|
|
||||||
|
A trace recorded with dropping already ON reports its `dropped` column instead of re-deriving it,
|
||||||
|
which is how the upper bound above gets checked against a real run.
|
||||||
|
|
||||||
|
Usage: route-drop-replay.py <route.csv> [<route.csv> ...]
|
||||||
|
Stdlib only, like the other analysis scripts here.
|
||||||
|
"""
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
from collections import defaultdict
|
||||||
|
|
||||||
|
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||||
|
from trace_io import read_preamble_csv # noqa: E402
|
||||||
|
|
||||||
|
MISS = 0
|
||||||
|
DECODE, PREFILL = 1, 0
|
||||||
|
|
||||||
|
|
||||||
|
def cells(rows, phase):
|
||||||
|
"""Group rows into routing cells: (turn, step, layer) -> [(slot, weight, residency, bytes, dropped)]."""
|
||||||
|
out = defaultdict(list)
|
||||||
|
for r in rows:
|
||||||
|
if int(r["phase"]) != phase:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
w = float(r["weight"])
|
||||||
|
except ValueError:
|
||||||
|
continue # 'nan': the graph exposed no weight node, so no threshold can be applied
|
||||||
|
out[(r["turn"], r["step"], r["layer"])].append(
|
||||||
|
(int(r["slot"]), w, int(r["residency"]), int(r["expert_bytes"]), int(r.get("dropped", 0) or 0))
|
||||||
|
)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def replay_threshold(cs, thr):
|
||||||
|
"""Drop a miss weighted below thr, never the cell's top expert (which the engine also pins)."""
|
||||||
|
miss_bytes = dropped_bytes = 0
|
||||||
|
total_mass = lost_mass = 0.0
|
||||||
|
kept_hist = defaultdict(int)
|
||||||
|
for entries in cs.values():
|
||||||
|
best = max(range(len(entries)), key=lambda i: entries[i][1])
|
||||||
|
kept = 0
|
||||||
|
for i, (_slot, w, res, nb, _d) in enumerate(entries):
|
||||||
|
total_mass += w
|
||||||
|
if res == MISS:
|
||||||
|
miss_bytes += nb
|
||||||
|
if i != best and w < thr:
|
||||||
|
dropped_bytes += nb
|
||||||
|
lost_mass += w
|
||||||
|
continue
|
||||||
|
kept += 1
|
||||||
|
kept_hist[kept] += 1
|
||||||
|
return miss_bytes, dropped_bytes, total_mass, lost_mass, kept_hist
|
||||||
|
|
||||||
|
|
||||||
|
def replay_static_k(cs, keep_k):
|
||||||
|
"""Baseline --n-expert-used: keep the top keep_k slots whatever the cache holds."""
|
||||||
|
miss_bytes = dropped_bytes = 0
|
||||||
|
total_mass = lost_mass = 0.0
|
||||||
|
for entries in cs.values():
|
||||||
|
for slot, w, res, nb, _d in entries:
|
||||||
|
total_mass += w
|
||||||
|
if res == MISS:
|
||||||
|
miss_bytes += nb
|
||||||
|
if slot >= keep_k:
|
||||||
|
lost_mass += w
|
||||||
|
if res == MISS:
|
||||||
|
dropped_bytes += nb
|
||||||
|
return miss_bytes, dropped_bytes, total_mass, lost_mass
|
||||||
|
|
||||||
|
|
||||||
|
def observed(cs):
|
||||||
|
"""What a trace recorded with dropping ON actually did. (dropped rows, weight mass, miss bytes)."""
|
||||||
|
n_dropped = 0
|
||||||
|
lost_mass = total_mass = 0.0
|
||||||
|
for entries in cs.values():
|
||||||
|
for _slot, w, _res, _nb, d in entries:
|
||||||
|
total_mass += w
|
||||||
|
if d:
|
||||||
|
n_dropped += 1
|
||||||
|
lost_mass += w
|
||||||
|
return n_dropped, lost_mass, total_mass
|
||||||
|
|
||||||
|
|
||||||
|
def pct(num, den):
|
||||||
|
return 100.0 * num / den if den else 0.0
|
||||||
|
|
||||||
|
|
||||||
|
def report(path):
|
||||||
|
meta, rows = read_preamble_csv(path)
|
||||||
|
k = int(meta.get("n_expert_used", 0) or 0)
|
||||||
|
if k <= 0:
|
||||||
|
print(f"{path}: no n_expert_used in the preamble; cannot express a threshold")
|
||||||
|
return
|
||||||
|
print("=" * 78)
|
||||||
|
print(f"{os.path.basename(path)} arch={meta.get('arch')} n_expert={meta.get('n_expert')} k={k}")
|
||||||
|
print("=" * 78)
|
||||||
|
|
||||||
|
for phase, label in ((DECODE, "DECODE"), (PREFILL, "PREFILL")):
|
||||||
|
cs = cells(rows, phase)
|
||||||
|
if not cs:
|
||||||
|
print(f"\n[{label}] no rows")
|
||||||
|
continue
|
||||||
|
n_tot = sum(len(e) for e in cs.values())
|
||||||
|
n_miss = sum(1 for e in cs.values() for x in e if x[2] == MISS)
|
||||||
|
print(f"\n[{label}] {len(cs)} routing cells, {n_tot} routed experts, {pct(n_miss, n_tot):.1f}% misses")
|
||||||
|
|
||||||
|
n_drop, lost, total = observed(cs)
|
||||||
|
if n_drop:
|
||||||
|
print(f" recorded: dropping was ON for this run -- {n_drop} routings dropped "
|
||||||
|
f"({pct(n_drop, n_tot):.1f}%), {pct(lost, total):.2f}% of the weight mass")
|
||||||
|
|
||||||
|
uniform = 1.0 / k
|
||||||
|
print(f"\n cache-aware threshold, as a fraction of the uniform share 1/k = {100 * uniform:.2f}%")
|
||||||
|
print(f" {'frac':>6} {'thr':>8} {'io_saved':>9} {'mass_lost':>10} surviving experts per cell")
|
||||||
|
for frac in (0.25, 0.5, 0.75, 1.0):
|
||||||
|
thr = uniform * frac
|
||||||
|
mb, db, tm, lm, hist = replay_threshold(cs, thr)
|
||||||
|
h = " ".join(f"{kk}:{pct(v, len(cs)):.0f}%" for kk, v in sorted(hist.items()))
|
||||||
|
print(f" {frac:6.2f} {100 * thr:7.2f}% {pct(db, mb):8.1f}% {pct(lm, tm):9.2f}% {h}")
|
||||||
|
|
||||||
|
print(f"\n static-k baseline (--n-expert-used), same rows, for comparison at equal io_saved")
|
||||||
|
print(f" {'keep_k':>6} {'io_saved':>9} {'mass_lost':>10}")
|
||||||
|
for keep in range(k - 1, 0, -1):
|
||||||
|
mb, db, tm, lm = replay_static_k(cs, keep)
|
||||||
|
print(f" {keep:6d} {pct(db, mb):8.1f}% {pct(lm, tm):9.2f}%")
|
||||||
|
print()
|
||||||
|
|
||||||
|
|
||||||
|
def main(argv):
|
||||||
|
if len(argv) < 2:
|
||||||
|
print(__doc__)
|
||||||
|
return 2
|
||||||
|
for p in argv[1:]:
|
||||||
|
report(p)
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main(sys.argv))
|
||||||
|
|
@ -143,6 +143,28 @@ int main() {
|
||||||
expect_ok("out-of-range sampling knobs are inert under greedy", c);
|
expect_ok("out-of-range sampling knobs are inert under greedy", c);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Cache-aware expert dropping. The upper bound is not cosmetic: above the uniform share the
|
||||||
|
// threshold can exceed every weight in a routing, and a config that can empty a layer must not
|
||||||
|
// be accepted just because the implementation happens to guard against it too.
|
||||||
|
{
|
||||||
|
RunConfig c = ok_base();
|
||||||
|
c.moe.enabled = true;
|
||||||
|
c.moe.drop_cold_frac = 0.5f;
|
||||||
|
expect_fail("dropping without a cache is rejected (nothing to be aware of)", c);
|
||||||
|
|
||||||
|
c.moe.cache_mb = MoeStreamConfig::cache_min_mb;
|
||||||
|
c.moe.drop_cold_frac = 0.0f;
|
||||||
|
expect_ok("dropping off is the default and valid", c);
|
||||||
|
c.moe.drop_cold_frac = 0.5f;
|
||||||
|
expect_ok("a threshold below the uniform share is valid", c);
|
||||||
|
c.moe.drop_cold_frac = 1.0f;
|
||||||
|
expect_ok("the uniform share itself is valid", c);
|
||||||
|
c.moe.drop_cold_frac = 1.01f;
|
||||||
|
expect_fail("a threshold above the uniform share is rejected", c);
|
||||||
|
c.moe.drop_cold_frac = -0.1f;
|
||||||
|
expect_fail("a negative threshold is rejected", c);
|
||||||
|
}
|
||||||
|
|
||||||
if (failures == 0) {
|
if (failures == 0) {
|
||||||
std::printf("all config checks passed\n");
|
std::printf("all config checks passed\n");
|
||||||
return 0;
|
return 0;
|
||||||
|
|
|
||||||
|
|
@ -350,6 +350,85 @@ int main(int argc, char ** argv) {
|
||||||
std::printf("[SKIP] S2 (expert-ready hook not built)\n");
|
std::printf("[SKIP] S2 (expert-ready hook not built)\n");
|
||||||
#endif
|
#endif
|
||||||
|
|
||||||
|
// G8 — cache-aware expert dropping, plumbing vs policy.
|
||||||
|
//
|
||||||
|
// Arming the policy moves load_layer() from the topk node to the terminal node of the layer's
|
||||||
|
// weight chain, and has the hook learn which node that is. That machinery must be transparent:
|
||||||
|
// with a threshold below any weight the router can produce, nothing is dropped and the output
|
||||||
|
// must stay byte-identical to the undropped stream. This separates "the deferral is correct"
|
||||||
|
// from "the policy is lossy" — only the second is allowed to change bytes, and a regression in
|
||||||
|
// the first would otherwise hide behind the expected difference.
|
||||||
|
// The policy needs a real LRU cache: with the cache off every expert reads as a miss, so it
|
||||||
|
// would degenerate into an unconditional weight cut and the repointing below would never face
|
||||||
|
// the reserved-but-uncommitted slot it exists to avoid. The small forced budget is the same one
|
||||||
|
// G2 uses to provoke evictions, so misses and hits both occur.
|
||||||
|
RunConfig drop_inert = base(model);
|
||||||
|
drop_inert.moe.enabled = true;
|
||||||
|
drop_inert.moe.cache_mb = 2;
|
||||||
|
drop_inert.moe.force_cache = true;
|
||||||
|
drop_inert.moe.io_threads = 4;
|
||||||
|
drop_inert.moe.drop_cold_frac = 1e-6f;
|
||||||
|
std::string s_drop_inert;
|
||||||
|
if (!gen(drop_inert, s_drop_inert, err)) {
|
||||||
|
std::fprintf(stderr, "drop(inert threshold) run failed: %s\n", err.c_str());
|
||||||
|
return 2;
|
||||||
|
}
|
||||||
|
fails += check("G8a drop(threshold below any weight) == streaming(cached, undropped)", s_sc, s_drop_inert);
|
||||||
|
// The identity above only means anything if the policy really was armed and really dropped
|
||||||
|
// nothing. Asserting the count separately turns "a weight happened to fall under the threshold"
|
||||||
|
// from a mysterious byte mismatch into a legible failure.
|
||||||
|
{
|
||||||
|
RunResult r = run(drop_inert);
|
||||||
|
if (!r || r.summary.experts_dropped != 0 || r.summary.experts_routed <= 0) {
|
||||||
|
std::printf("[FAIL] G8a' inert threshold must examine routings and drop none (routed=%lld dropped=%lld)\n",
|
||||||
|
r.summary.experts_routed, r.summary.experts_dropped);
|
||||||
|
++fails;
|
||||||
|
} else {
|
||||||
|
std::printf("[PASS] G8a' inert threshold examined %lld routings, dropped none\n", r.summary.experts_routed);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// G8b — the same at full strength, against a cache small enough to be evicting constantly, so
|
||||||
|
// dropped experts really do land on slots the cache has released. There is no reference output
|
||||||
|
// to compare against (it is lossy by design), so the gate is that the engine survives it: a
|
||||||
|
// dropped expert's slot is repointed at one that is certainly resident, so the matmul must never
|
||||||
|
// read reserved-but-uncommitted memory and generation must still complete.
|
||||||
|
RunConfig drop_hard = drop_inert;
|
||||||
|
drop_hard.moe.drop_cold_frac = 1.0f;
|
||||||
|
drop_hard.moe.drop_prefill = true;
|
||||||
|
std::string s_drop_hard;
|
||||||
|
if (!gen(drop_hard, s_drop_hard, err)) {
|
||||||
|
std::fprintf(stderr, "drop(full strength) run failed: %s\n", err.c_str());
|
||||||
|
return 2;
|
||||||
|
}
|
||||||
|
if (s_drop_hard.empty()) {
|
||||||
|
std::printf("[FAIL] G8b drop(full strength) produced no output\n");
|
||||||
|
++fails;
|
||||||
|
} else {
|
||||||
|
std::printf("[PASS] G8b drop(full strength) generates without touching unloaded experts\n");
|
||||||
|
}
|
||||||
|
|
||||||
|
// G8c — the top-weighted expert is pinned, so a routing can never be emptied. Forcing top-k to
|
||||||
|
// 1 makes every routed expert the top one, and dropping must then be a no-op at ANY threshold:
|
||||||
|
// the output has to match the same k=1 run with the policy off, byte for byte. This also pins
|
||||||
|
// down that the threshold is taken against the EFFECTIVE top-k discovered at runtime — a
|
||||||
|
// hardcoded width would not survive the override.
|
||||||
|
RunConfig k1 = base(model);
|
||||||
|
k1.moe.enabled = true;
|
||||||
|
k1.moe.cache_mb = 2;
|
||||||
|
k1.moe.force_cache = true;
|
||||||
|
k1.moe.io_threads = 4;
|
||||||
|
k1.n_expert_used = 1;
|
||||||
|
RunConfig k1_drop = k1;
|
||||||
|
k1_drop.moe.drop_cold_frac = 1.0f;
|
||||||
|
k1_drop.moe.drop_prefill = true;
|
||||||
|
std::string s_k1, s_k1_drop;
|
||||||
|
if (!gen(k1, s_k1, err) || !gen(k1_drop, s_k1_drop, err)) {
|
||||||
|
std::fprintf(stderr, "top-k=1 drop run failed: %s\n", err.c_str());
|
||||||
|
return 2;
|
||||||
|
}
|
||||||
|
fails += check("G8c drop(full strength, top-k=1) == top-k=1 undropped (top expert pinned)", s_k1, s_k1_drop);
|
||||||
|
|
||||||
if (fails == 0) std::printf("\nall MoE byte-identity gates passed\n");
|
if (fails == 0) std::printf("\nall MoE byte-identity gates passed\n");
|
||||||
return fails == 0 ? 0 : 1;
|
return fails == 0 ? 0 : 1;
|
||||||
}
|
}
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue