BigMoeOnEdge/examples
Helldez 4645dc6b09
feat(engine): --route-ahead N — commit decode routing to the N-layers-early prediction (#142)
* feat(engine): --route-ahead N — commit decode routing to the N-layers-early prediction

Every prefetch lives under the same ceiling: layer L's routing needs layer L-1's
output, so any predictor working earlier is approximate and every speculated read
can miss. This inverts the bet. The expert selection of decode layer L is REPLACED
by the ranking layer L's own gate matrix produced on the hidden state N layers back
in the same forward pass, so the selection is known N layers early and cannot miss.
With the cache on, those reads are issued the moment the selection is fixed.

Lossy by construction: it changes the output, and roughly a fifth of slots route to
a different expert than the router chose at N=1. Off by default, mutually exclusive
with both prefetchers and with the prediction probe, since each would speculate on a
future this policy has already decided.

Quality is measured rather than assumed: the committed generations match the
baseline on a four-prompt objective battery, on a long essay, and on a second model
of a different generation and quantization; output stays deterministic across
repeated runs, which expert dropping does not. See docs/route-ahead.md.

Squashed from the seventeen commits of exp/route-ahead: the branch predated the
multi-shard and speculative-decoding work, and replaying it commit by commit meant
resolving the same two collisions seventeen times over. The history is preserved on
the pull request; what lands here is what a squash-merge would have produced anyway.

* fix(engine): refuse route-ahead alongside self-speculation, and say why

Running the two together on a real model committed NOTHING: 0 routings taken,
249 passed through. A verify decode is several positions wide and the policy
correctly declines each one. But it still charged for itself — the prediction
GEMVs ran (2.6 ms/token of worker CPU) and its early reads degraded into
ordinary speculation, falling from 100% useful to 81%.

Cost with no commitment is worse than either feature alone, and nothing told
the user. validate() now rejects the pair the way it already rejects
route-ahead beside the two prefetchers, the app stops emitting the flag and
greys the row out, and the config test covers both draft sources.

Making the combination work is a different change: commit the whole verify
batch to one selection. That is written and measured, and it costs draft
acceptance (70% to 53%) while route-ahead alone still won, so the exclusion
is the honest state today rather than a limitation to be worked around.

Found by the desktop smoke run, not by the gates.
2026-08-02 00:26:38 +02:00
..
android feat(engine): --route-ahead N — commit decode routing to the N-layers-early prediction (#142) 2026-08-02 00:26:38 +02:00