mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
The threshold is a fraction of the uniform share 1/top-k, so what it removes scales with how wide the routing is. At top-k 8 -- where every number in docs/expert-dropping.md was collected -- 0.75 means "below 9.4% of the routing", a tail trim. At top-k 4 it means "below 18.8%", and at top-k 2 "below 37.5%", which on a miss discards the whole minority expert: closer to halving the routing than trimming it, and unmeasured. The engine now says so once at load when dropping is armed and the effective top-k is 4 or fewer, quoting the actual share for the model in hand. It warns rather than clamping or refusing: it cannot know whether that trade is acceptable for a given model and task, and silently adjusting a number the caller chose would be worse than a loud caveat. MoeStreamConfig::drop_low_topk_warn is documented as an EVIDENCE boundary, not a physical one -- nothing in the streaming path reads it. The app shows the same caveat inline under the setting, in the error colour, computed from the width the loaded model reports rather than assumed -- and only once a model is loaded, since guessing would be worse than staying quiet. gpt-oss is the case this exists for: it routes 4 of 128, and the app default is 75%. To make that possible, BMOE_READY gains n_expert_used (the effective width after any override, 0 on a non-MoE model) and Session exposes n_expert_used() for embedders. Additive: older consumers ignore it. |
||
|---|---|---|
| .. | ||
| android | ||