BigMoeOnEdge/examples
Helldez 6441494b76
feat(moe): warn when cache-aware dropping meets a narrow routing (#99)
The threshold is a fraction of the uniform share 1/top-k, so what it
removes scales with how wide the routing is. At top-k 8 -- where every
number in docs/expert-dropping.md was collected -- 0.75 means "below 9.4%
of the routing", a tail trim. At top-k 4 it means "below 18.8%", and at
top-k 2 "below 37.5%", which on a miss discards the whole minority
expert: closer to halving the routing than trimming it, and unmeasured.

The engine now says so once at load when dropping is armed and the
effective top-k is 4 or fewer, quoting the actual share for the model in
hand. It warns rather than clamping or refusing: it cannot know whether
that trade is acceptable for a given model and task, and silently
adjusting a number the caller chose would be worse than a loud caveat.
MoeStreamConfig::drop_low_topk_warn is documented as an EVIDENCE
boundary, not a physical one -- nothing in the streaming path reads it.

The app shows the same caveat inline under the setting, in the error
colour, computed from the width the loaded model reports rather than
assumed -- and only once a model is loaded, since guessing would be worse
than staying quiet. gpt-oss is the case this exists for: it routes 4 of
128, and the app default is 75%.

To make that possible, BMOE_READY gains n_expert_used (the effective
width after any override, 0 on a non-MoE model) and Session exposes
n_expert_used() for embedders. Additive: older consumers ignore it.
2026-07-23 10:08:25 +02:00
..
android feat(moe): warn when cache-aware dropping meets a narrow routing (#99) 2026-07-23 10:08:25 +02:00