mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
The NPU prefill's expert arena read every expert of every layer ahead of its routing. It now reads, ahead of a layer's routing, the experts the previous graph routed there, and at the routing node whatever the routing adds. The matmul reads only routed experts, so the output is bit for bit the same. A layer routing more than --prefill-routed-full (0.85) of its experts gets the next one read whole; --no-prefill-routed restores whole layers everywhere. Phone, Hexagon v81 NPU, top-4, same session, every answer identical: Qwen3.6-35B-A3B Q4_0 7.68 -> 4.16 s, Q4_K_M 9.95 -> 5.37 s, Gemma 4 26B-A4B Q4_K_M 6.69 -> 3.70 s, Nemotron 3.5 30B-A3B Q4_0 7.42 -> 6.81 s. Also: --decide-probe (experimental per-decision expert usage and layer-exit answers), BMOE_DECIDE prefill_dev_* counters, gates G17f/G17g, app 0.28.0. |
||
|---|---|---|
| .. | ||
| config.h | ||
| decide.h | ||
| decode_trace.h | ||
| expert_source.h | ||
| metrics.h | ||
| ngram_draft.h | ||
| predict_stats.h | ||
| recipe.h | ||
| route_trace.h | ||
| row_source.h | ||
| runtime.h | ||
| session.h | ||
| version.h | ||