* feat(moe): --drop-cold-experts, spend quality only where it buys I/O Turbo top-k drops the tail of a routing whether or not those experts were already in RAM. A resident expert costs no flash read, so that trade pays quality for nothing on the ~80% of decode routings that are cache hits. This skips a routed expert only when it is a cache MISS and the router weighted it below frac x (1/top-k). Replayed over the committed route traces at frac 1.0, decode phase, that avoids 66% of flash reads for 9.5% of the router's weight mass, where --n-expert-used 5 avoids 23% for a comparable 10.6% -- about 3x the reads at the same quality cost. Implementation. The decision needs the FINAL router weights, which arrive several nodes after the topk where the streamer normally loads, so load_layer() is deferred to the terminal node of the layer's weight chain. Which node that is depends on the model's gating, so the hook learns it from the graph rather than carrying an architecture table; if it fails to arrive the hook forgets it and re-learns rather than re-betting. A dropped slot has its weight zeroed and its expert id repointed at the routing's top-weighted expert: an unread expert can sit in reserved-but-uncommitted VM and mul_mat_id would touch it anyway, so the kernel is given memory that is certainly resident and multiplies it by exactly zero. Requires the LRU cache -- with --cache-mb 0 residency reads all-miss and the policy would silently degenerate into an unconditional weight cut. Prefill is excluded by default. The top expert is always pinned, so no routing can be emptied at any threshold. Gates: G8a/G8a' prove the deferral and the learned terminal node are transparent (byte-identical output, zero drops, at a threshold below any producible weight); G8b that full strength against a constantly-evicting cache never reaches an unloaded slot; G8c that at top-k 1 dropping is a no-op, pinning both the top-expert guarantee and the threshold tracking the effective top-k. Three existing metrics shift meaning under dropping and the docs now say so: cache_hit_pct rises without the cache serving more (a dropped routing is a miss that is never looked up), and token/layer_demand measure what was staged rather than routed. prefetch.md's "cannot change output" is scoped, limitations.md gains the non-reproducibility entry, and benchmark-method.md warns that reversing the run order cannot distinguish a moved drop rate from a contaminated cell. Off by default in the CLI and in the app. The output is not reproducible -- what gets dropped depends on what the cache held -- so it carries no rows in the README tables, and switching it on by default waits on a published on-device A/B rather than on the replay argument alone. * feat(app): default cache-aware dropping to 75%, measured on device Qwen3.6-35B-A3B (top-k 8 of 256), in-app, cache 3000, one variable changed: 2.549 tok/s off, 3.938 at F=0.75 (+55%), 4.702 at F=1.0 (+84%), with flash reads falling 248 -> 163 -> 48 GiB. Per-token bootstrap intervals separate every pair except off vs 0.50, which overlaps -- at half the uniform share the policy drops 2.7% of routings and buys nothing, which doubles as a negative control that the machinery is free when it does not fire. Run order was 1.0, off, 0.5, 0.75, so the two fastest cells are the first and the LAST; thermal drift would have made the last the worst. The mechanism orders by threshold even though the run order does not. The replay turned out conservative rather than optimistic. It is documented as an upper bound because it cannot model the cache changing in response to dropping: at F=0.75 it was accurate (37% predicted, 34% measured), at F=1.0 it understated (66% predicted, 81% measured). Avoided reads free cache capacity, which raises the hit rate, which leaves fewer misses to drop. 75% rather than 100% is deliberate: it takes the larger part of the win for half the discarded routings (14% against 28%). Quality is still unquantified -- no perplexity number and no side-by-side exists -- so the conservative end of a measured range is the defensible default. The CLI stays off; the byte-identity gates need a deterministic default. Also records cache_hit_pct rising 67.8 -> 90.7% as the documented accounting artefact rather than the cache serving more, and majflt/token as dominated by each run's starting memory state, not by the threshold.
4.3 KiB
Cache-aware expert dropping — measured POSITIVE on device (2026-07-22)
Verdict: --drop-cold-experts is a real throughput win, and the offline replay understated it at
full strength. 0.75 and 1.00 are each separable from the baseline and from each other; 0.50 does
nothing measurable. Quality was not measured and remains the open question.
Setup
In-app, one device, Qwen3.6-35B-A3B-Q4_K_M (qwen35moe, 40 layers, 256 experts, top-k 8), cache
3000 MiB, --dense-weights ahwb, overlap on, prefetch off, 4 compute + 4 read lanes, ~1350-1470
token generations. Every cell differs in drop_cold_frac only; the four # moe_stream=…
preambles in this directory are identical apart from that field.
Whole-run numbers
drop_cold_frac |
tok/s | flash read | I/O s/token | cache hit | routings dropped |
|---|---|---|---|---|---|
| 0 (off) | 2.549 | 248.2 GiB | 0.299 | 67.8% | 0 |
| 0.50 | 2.564 | 231.4 GiB | 0.279 | 69.1% | 2.7% |
| 0.75 | 3.938 | 162.8 GiB | 0.190 | 77.2% | 14.2% |
| 1.00 | 4.702 | 48.1 GiB | 0.070 | 90.7% | 28.4% |
Per-token, first 50 tokens trimmed (cache warm-up, plus the one token the hook spends learning the graph), 2000-sample bootstrap on the mean:
drop_cold_frac |
median ms/token | mean ms/token | 95% CI on the mean |
|---|---|---|---|
| 0 | 249.5 | 391.1 | [375.3, 407.1] |
| 0.50 | 263.2 | 391.6 | [375.3, 408.5] |
| 0.75 | 202.8 | 253.5 | [246.9, 260.4] |
| 1.00 | 153.4 | 214.9 | [206.2, 225.0] |
Every pair is disjoint except 0 vs 0.50, which overlaps. At half the uniform share the policy finds almost nothing to drop (2.7% of routings) and buys nothing — exactly what the replay predicted, and a useful negative control: the machinery costs nothing measurable when it is not firing.
Why this is not the device drifting
Run order was 1.00 (13:42), off (13:51), 0.50 (14:04), 0.75 (14:23) — deliberately not in threshold order. The two fastest cells are the first and the last. Thermal drift or accumulated memory pressure would make the last cell the worst; it is the second best. The baseline also ran early, on a relatively cool device, and still lost.
The mechanism is monotone in the threshold even though the run order is not: flash read
248 → 231 → 163 → 48 GiB and I/O 0.299 → 0.279 → 0.190 → 0.070 s/token order themselves perfectly by
drop_cold_frac. That dose-response is what run order cannot fake.
The replay was conservative, not optimistic
docs/expert-dropping.md calls the offline replay an upper bound, because it cannot model the
cache changing as a result of dropping. At F = 0.75 it was accurate (predicted ~37% of reads
avoided, measured 34%). At F = 1.0 it understated: predicted 66%, measured 81%.
The extra comes from a compounding effect the static replay is blind to: reads avoided free cache capacity, which raises the hit rate, which leaves fewer misses to drop in the first place. The caveat in the doc should be read as "the bound holds where dropping is light, and is pessimistic where it is heavy".
Read cache_hit_pct with the documented caveat
The rise from 67.8% to 90.7% is not the cache serving more. A dropped routing is a miss that is
never looked up, so it leaves both sides of the ratio (see docs/telemetry.md). read_MiB is the
honest number here, and it is unaffected by that accounting: 248 GiB → 48 GiB is real.
majflt/token (224 / 178 / 14 / 611) does not order by threshold and is dominated by what the
device's memory looked like when each run started. Nothing should be concluded from it here.
What this does not show
- Quality is unmeasured. At
F = 0.75the policy discards 14.2% of routings and atF = 1.028.4%. There is no perplexity number and no side-by-side comparison in this experiment. Throughput is settled; whether the output holds up is not. - One model, one device, one run per cell. The confidence intervals are within-run (over ~1300
tokens), not across repeats. The reversed-order pair is owed, as it was for
ahwb. - Top-k 8 only. Qwen3.6 routes 8 of 256. A model routing 2 or 4 experts has a much larger uniform
share, so the same
Fbites differently there. gpt-oss in particular is unmeasured. - Prefetch was off. The interaction between speculation and dropping (a correct guess un-drops an expert) is untested.