mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
* feat(engine): --route-ahead N — commit decode routing to the N-layers-early prediction Every prefetch lives under the same ceiling: layer L's routing needs layer L-1's output, so any predictor working earlier is approximate and every speculated read can miss. This inverts the bet. The expert selection of decode layer L is REPLACED by the ranking layer L's own gate matrix produced on the hidden state N layers back in the same forward pass, so the selection is known N layers early and cannot miss. With the cache on, those reads are issued the moment the selection is fixed. Lossy by construction: it changes the output, and roughly a fifth of slots route to a different expert than the router chose at N=1. Off by default, mutually exclusive with both prefetchers and with the prediction probe, since each would speculate on a future this policy has already decided. Quality is measured rather than assumed: the committed generations match the baseline on a four-prompt objective battery, on a long essay, and on a second model of a different generation and quantization; output stays deterministic across repeated runs, which expert dropping does not. See docs/route-ahead.md. Squashed from the seventeen commits of exp/route-ahead: the branch predated the multi-shard and speculative-decoding work, and replaying it commit by commit meant resolving the same two collisions seventeen times over. The history is preserved on the pull request; what lands here is what a squash-merge would have produced anyway. * fix(engine): refuse route-ahead alongside self-speculation, and say why Running the two together on a real model committed NOTHING: 0 routings taken, 249 passed through. A verify decode is several positions wide and the policy correctly declines each one. But it still charged for itself — the prediction GEMVs ran (2.6 ms/token of worker CPU) and its early reads degraded into ordinary speculation, falling from 100% useful to 81%. Cost with no commitment is worse than either feature alone, and nothing told the user. validate() now rejects the pair the way it already rejects route-ahead beside the two prefetchers, the app stops emitting the flag and greys the row out, and the config test covers both draft sources. Making the combination work is a different change: commit the whole verify batch to one selection. That is written and measured, and it costs draft acceptance (70% to 53%) while route-ahead alone still won, so the exclusion is the honest state today rather than a limitation to be worked around. Found by the desktop smoke run, not by the gates. |
||
|---|---|---|
| .. | ||
| chat_parse_test.cpp | ||
| CMakeLists.txt | ||
| config_test.cpp | ||
| moe_gates.cpp | ||
| ngram_test.cpp | ||
| think_control_test.cpp | ||