* feat(moe): --predict-log, measure how predictable expert routing is
Temporal prefetch was built on a predictor nobody had priced. The
previous-token bet turned out to be right ~38% of the time on Qwen and
~18% on gpt-oss, which cannot pay for the reads it speculates. The
lesson was not that prefetch is impossible but that a predictor should
be measured before it is wired into anything. This adds the instrument,
not a policy.
--predict-log ranks each layer's experts a layer early, by running the
NEXT layer's router matrix on the CURRENT layer's gate input. The
residual stream barely moves between layers, so the stale input ranks
nearly as the real one will -- the mechanism FATE (arXiv 2502.12224)
reports 78.8% for. It is training-free and changes no model: the router
matrix is a dense weight already resident, and the prediction is one
GEMV per layer. The matrix is learned from the graph (the gate matmul's
first source) rather than looked up by tensor name, so no architecture
is named anywhere in the path; being a weight leaf, its pointer stays
valid into layers the current token has not reached.
Three predictors are scored against the routing the router actually
produced, so they are comparable on one run: the stale gate, the
previous-token bet --prefetch already places, and a zero-staleness
control. The control is the load-bearing part. It shares every line of
code with the prediction under test and differs only in using the
layer's own matrix, so it must reproduce the selection llama.cpp
computes from those same two tensors. A transposed matrix, a mis-strided
row or the wrong token of the batch collapses it toward chance while the
stale figure would stay superficially plausible; an architecture that
selects by something other than raw-logit ranking (an additive bias,
group-limited routing) puts it below 100% and by that much the stale
figure understates the method. The CLI says so rather than letting the
gap be blamed on staleness.
Reported per layer as well as in aggregate, because an aggregate
flatters a prefetch: what a prefetch costs is set by the layers it gets
wrong, and a MoE model's first layers route far less predictably than
its last. Denominators are printed per predictor -- the stale gate
structurally cannot rank layer 0 (nothing precedes it) or the first
token of a run, and those routings are counted as unscored rather than
folded in, since a routing that was not ranked is not a wrong guess. A
predictor with no routings at a layer prints "-", never 0.0.
Diagnostics only: nothing it computes reaches load_layer, the cache or
the graph, so a probed run reads exactly the bytes an unprobed one does.
G9a gates that byte identity and G9b gates the control, which reads
100.0% on the tiny model. It is not free -- one isolated node and two
GEMVs per layer on the eval thread -- so a probed run is not a benchmark
run, and it requires --moe-stream since routing does not depend on how
the weights reached memory.
docs/expert-prediction.md also records the caveat the number will need:
a high score would say the routing is knowable earlier, not that knowing
it earlier makes decode faster. On a flash already saturated, starting a
read sooner adds no bandwidth -- which is why prefetch, layer-LFU and
the expert sidecar all lost despite improving the metric each was
designed around.
* feat(moe): --predict-prefetch, speculate on the stale-gate prediction
The probe said the routing is knowable a layer early (~89% of routed
slots on a 128-expert model, vs ~43% for the previous-token bet the
temporal prefetch acts on). This wires that prediction into the existing
speculative read path -- same cache buffers, same accounting, same
settle, same moe-prefetch summary line (tagged [stale-gate]) -- so the
only thing that changes is which guess rides the idle lanes.
Two decisions carry the design:
Speculation is issued AFTER the current layer's load, not at prediction
time. Every load path begins by quiescing speculation, and a layer's own
load sits a few graph nodes after its gate matmul -- reads queued at
prediction time would be cancelled before a lane picked them up. Each of
the three load sites (plain topk, deferred drop, drop fallback) issues
the pending next-layer prediction right after its load_layer, restoring
the same read-ahead window the temporal prefetch gets. On the tiny-model
gates this is the difference between 0% and 27% of speculated experts
proving useful -- the latter matching the probe's measured accuracy on
that model, which is the accounting agreeing with itself.
It is drop-aware. With --drop-cold-experts armed, a predicted expert
whose predicted routing weight (softmax over the predicted top-k)
falls below the drop threshold is not speculated: if it misses, the
policy discards it unread, so reading it ahead would spend the exact
I/O the policy exists to save. The top prediction is always kept,
mirroring the policy's own pin of the top-weighted expert. The known
interplay is inherited from the temporal prefetch and deliberate: a
correct guess un-drops an expert, buying quality at the same threshold
rather than speed.
The routing width the prefetch predicts at is learned from the topk
node, not read from config, so an --n-expert-used override stays honest
with no extra plumbing. Mutually exclusive with --prefetch (two
predictors would double-speculate the same future); requires the LRU
cache; decode only. The control GEMV remains probe-only, so the
production path costs one gate GEMV per MoE layer per token.
Gates: G10a proves byte-identity through the speculative path
(prefetch-sync, forced small cache, hits and evictions both occur);
G10b proves the run actually speculated and that useful-hit accounting
tracks the probe's accuracy. Off by default, pending an on-device A/B.
* feat(moe): cap predictive speculation at the top 2 predicted misses per layer
Speculating the whole predicted routing was measured on device at -38%
against its own baseline despite every intermediate metric improving
(hit 77.6->88.9%, stall 32->9 ms, 84% of speculations useful): the +33%
flash bytes and the vm commits fighting a full cache (major faults x3)
cost several times the stall removed. The stall a prefetch can remove is
head-of-line only -- overlap already hides the tail behind the expert
matmul -- so the cap keeps the part of the bet that can pay and drops
the part that provably cannot. Residency-aware: a predicted expert
already in cache does not burn a slot of the cap.
* perf(moe): rebuild the predictive prefetch around its measured costs
The observer-tax run priced the naive implementation: ~35-45 ms per
GEMV pass (ggml_fp16_to_fp32 is a function call per weight element --
21M calls/token) and ~20 ms/token for the extra isolated node, against
a speculation machinery that costs ~15-25. The GEMV and the barrier
were the feature; this commit removes both.
- gate_scores converts F16 natively on aarch64 (one instruction,
vectorizable) instead of a function call per element.
- The prefetch no longer isolates the gate matmul: the ask pass hands
over its source pointers for free, and the gate-input row is read at
the topk callback with no barrier of its own. A sampled watchdog
(the zero-staleness control, every 512 routings) validates the
barrier-less read and disarms the prefetch out loud if the memory
planner ever reuses that buffer -- without it, a future llama.cpp
bump could silently turn the predictor into a noise generator.
- The GEMV runs on a dedicated worker at a TWO-layer horizon: one
layer ahead has no landing spot (the callbacks between a layer's
gate and its own load are microseconds apart), while at l+2 the
worker has a whole layer for a ~0.5M-MAC job and the result inherits
the same post-load issue window as before. The probe now also scores
stale-2, so the extra layer of staleness is priced per model rather
than assumed.
- The prediction's residents are RETAINED (new IExpertSource::retain:
move-to-MRU, deliberately not a cache hit so the hit-rate metric
stays honest) -- protecting a predicted expert costs zero bytes,
unlike prefetching it. Only predicted misses are speculated, still
capped at 2.
The probe path keeps its barrier and both GEMVs: a probed run is
diagnostics, and its job is to be right, not fast.
* feat(moe): --predict-spec-max N — how much flash the prediction may spend (0 = retention only)
The cap was a constant; the retention-only point (0) is the config the
whole experiment now hinges on -- the prediction protecting predicted
residents from eviction while spending no flash at all -- and a
measured constant that cannot be varied is not a mechanism. Validated
[0, 8]; retention happens at every value because it is free.
* docs(predict): record the 2026-07-23 campaign — accuracy confirmed, throughput verdict open
Accuracy on device: stale-gate 88.6% (Qwen3-30B) / 80.7% (Qwen3.6),
control exactly 100.0% on every layer of both, prev-token 43/35% --
corroborating the offline route-trace estimates.
Throughput: the day's C-vs-B losses are recorded WITH their
invalidation. Re-running the reference on the by-then-hot device gave
3.93 tok/s against the cool morning's 6.54 with byte-identical I/O,
hit and drop counts -- the engine is deterministic, the -40% was
silent thermal capping, and every variant had been compared against
the cool number. The one thermally matched pair that was measured
(speculation on 8 io lanes, -28%, effective flash bandwidth 585->392
MiB/s) kills the more-lanes hypothesis specifically; spec-max 2 and
retention-only still owe a matched cool pair.
What did survive: the observer-tax decomposition (109 ms/token: the
per-element exported-function F16 conversion at 21M calls/token, plus
the barrier), and retention moving hit rate 0.1pp on a 3000 MiB cache
-- the offline replay bound confirmed from inside the engine.
* feat(app): expose the predictive prefetch as an experimental Streaming toggle
Off by default. Gated on streaming + a live cache like the temporal
prefetch, and the two settings disable each other in the UI -- the
engine refuses the pair, and a control the engine will reject is worse
than one that cannot be set. The spec-max rung selector (0/1/2/4)
surfaces the retention-only point, which is the configuration the open
throughput question most needs measured from the app. Session
signature includes both fields so flipping them reopens the process.
No versionCode bump: this is a PR-branch test build, not a release.
* docs(predict): record the 2026-07-24 matched pairs — read-ahead refuted, retention hit-neutral
A four-cell session (B, retention-only, spec-2, B sentinel) run at fixed 30 s
spacing re-proved the thermal-contamination mechanism (sentinel −17%, clusters
silently capped from cell 2) and yielded one genuinely matched pair: spec-2 vs
the B sentinel at the same caps and battery temperature, 3.14 vs 3.96 tok/s
(−21%) with hit rate up 4.3pp and 79% of speculations useful. Speculation
improves every metric it owns and still loses the wall clock — the flash has no
spare bandwidth to spend. Retention-only again moved the hit rate by nothing
(77.2% vs 77.6%), as the offline replay bound predicted.
* feat(app): contrast the two prefetch predictors in the UI, default spec-max to 0
The temporal and predictive toggles now say what actually differs — the bet
("repeats the previous token", ~40%) vs the question ("ask the next router a
layer early", ~85%) — instead of describing mechanisms side by side. Spec-max
defaults to 0 (retention only): the matched-pair A/B showed the read-ahead
losing −21% on a saturated flash, so 0 is the only rung the measurements did
not refute, and the helper text says so.
* chore(predict): file the changelog under Unreleased, drop a dead member, record the verdict
Three loose ends found reviewing the branch for merge:
- The changelog entries had been appended to the already-released 0.15.1 section;
they belong under [Unreleased], where the release commit carves them out.
- nu_hint_ was written on every routing and read nowhere: a leftover of the
synchronous first design, whose successor passes the routing width straight into
the prediction job.
- docs/roadmap.md still closed the routing-prediction question on the 2026-07-12
removal. It now records what reopening it with a training-free predictor found:
the accuracy is real and the throughput is not, for the same reason more lanes
and the sidecar lost.
13 KiB
Expert prediction: measuring it before building it
Measured, 2026-07-23/24 — accuracy confirmed, read-ahead refuted in a matched pair, retention neutral on hit rate, and a thermal lesson that invalidated half a day of comparisons. On device, the stale-gate predictor scores 88.6% of routed slots on Qwen3-30B (48 layers, top-8 of 128) and 80.7% on Qwen3.6-35B (top-8 of 256), with the zero-staleness control at exactly 100.0% on every layer of both — Qwen selects by pure logit ranking, so these numbers need no discount. The previous-token predictor scored 43.3% / 35.4% on the same runs, corroborating the offline route-trace estimates.
Acting on the prediction (
--predict-prefetch) was then measured across the whole cost spectrum against a drop-0.75 + pinned-dense baseline ("B", 6.54 tok/s cool): full-routing speculation 4.02 (−38%: +33% flash bytes, major faults ×3), top-2-miss cap 4.46, the rebuilt implementation (native-F16 GEMV ~40× faster, no barrier + self-validating watchdog, worker at an L+2 horizon, retention of predicted residents) 5.30, retention-only (zero extra bytes) 5.44. Every variant improved its intermediate metrics — hit rate up, stall down to 9-22 ms, 77-86% of speculations useful — while the wall clock stayed behind.But those deltas are contaminated: re-running B itself at the end of the day, on a device that had heated through the whole campaign (mid cores silently capped ~8%), gave 3.93 tok/s with byte-for-byte identical I/O, hit rate and drop counts — the engine is deterministic; the −40% was throttling. Each C variant ran on a progressively hotter device against the cool-morning 6.54, so the pairwise conclusions do not stand. The one thermally matched pair of the day — B-hot 3.93 vs prefetch on EIGHT io lanes 2.83 (−28%, effective flash bandwidth collapsing from 585 to 392 MiB/s) — kills the more-lanes hypothesis specifically, consistent with the measured lane ceiling, but the spec-max 2 and retention-only configurations still owe a matched cool-pair A/B before any verdict. Two structural facts did survive the day: the observer tax was 109 ms/token before the rebuild (~35-45 ms per GEMV pass from a per-element exported-function conversion — 21M calls/token — plus ~20 ms of barrier), and retention moved the hit rate by 0.1pp on a 3000 MiB cache, confirming offline replay: at that size, eviction order within a two-layer horizon is irrelevant.
The 2026-07-24 redo settled most of what stayed open. A four-cell session (B, retention-only, spec-2, then B again as a drift sentinel) run with fixed 30 s spacing first re-proved the contamination mechanism on purpose — the closing B lost 17% against the opening B (4.77 → 3.96 tok/s) with the CPU clusters silently capped from the second cell on — and then yielded the one pair the spacing left genuinely matched: spec-2 vs the B sentinel, adjacent cells at the same caps and the same battery temperature, 3.14 vs 3.96 tok/s (−21%), with hit rate up 4.3pp, 79% of speculations useful, and +37% flash bytes per token. Speculation improves every metric it owns and still loses the wall clock; the top-2 read-ahead is refuted in the shipping configuration, by the same mechanism that killed more-lanes — the flash has no spare bandwidth to spend. Retention-only again moved the hit rate by nothing (77.2% vs 77.6%), as the offline replay bound predicted; its own wall-clock cell started inside the deepest throttle window and is not readable, but with zero hit-rate movement there is no mechanism left for it to win by — only the residual gate-GEMV cost to lose by.
--predict-log answers one question and refuses to answer any other: how much of a layer's
routing could be known before that layer runs? It predicts, scores itself against what the router
actually chose, and prints the result. Nothing it computes reaches the loading path — a probed run
reads exactly the bytes an unprobed one does, only slower.
It exists because temporal prefetch was built on a predictor nobody had priced. That predictor — "the previous token routed these, the next one probably will too" — turned out to be right about 38% of the time on Qwen and 18% on gpt-oss, which is not enough to pay for the reads it speculates. The lesson was not that prefetch is impossible; it was that a predictor should be measured before it is wired into anything.
The three predictors
All three are scored on the same routings, on the same run, so they are directly comparable.
stale-gate is the one under test. Each layer's router is a matrix multiplied by that layer's gate input; the input comes from the residual stream, which every layer only nudges. So while layer l is computing, we take the input it is about to consume and multiply it by layer l+1's router matrix. The ranking that comes out is a guess at what layer l+1 will select, available a full layer early. It needs no training and no change to the model — the matrix is already resident (it is a dense weight, not a streamed expert) and the multiply is one GEMV.
This is the mechanism FATE uses, where it scored 78.8% on DeepSeek-V2-Lite.
prev-token is the incumbent: what this layer routed for the previous token — exactly the bet
--prefetch places. Reported here so the new predictor is judged against the one already shipped,
not against zero.
ctrl is not a predictor. It is the stale-gate with the staleness removed: the same row read, the same GEMV, the same ranking, but using layer l's own matrix on layer l's own input. It therefore has to reproduce the selection llama.cpp is about to compute from those very tensors, and it should read 100%. Anything less is this probe being wrong, not routing being hard:
- a GEMV that disagrees with ggml's (a transposed matrix, a mis-strided row, the wrong token of the batch) collapses it toward chance;
- an architecture that does not select by raw-logit ranking — one that adds a per-expert selection bias, or masks whole expert groups before the argsort — puts it somewhere below 100%, and by roughly that much the stale-gate number understates the method.
Read the other two against ctrl, never against 100%. The CLI says so out loud when it drops.
Reading the report
moe-predict: stale-gate <pct>% of routed slots (<pct>% whole routings) | prev-token <pct>% (<pct>%) | fresh-gate control <pct>% (<pct>%)
moe-predict: scored — stale-gate <rows> routings/<slots> slots, prev-token <rows>/<slots>, control <rows>/<slots>; <n> routings the stale-gate could not rank
layer stale prev ctrl routings
The headline figure is slot overlap: of the k experts a routing selected, how many the predictor had named. That is the number a prefetch cares about — each hit is one read turned into a hit. The parenthesised figure is whole routings predicted exactly, i.e. the fraction of layers that would have needed no on-demand read at all. They are far apart, and papers differ on which one they quote, so both are printed. FATE's 78.8% is the first kind; the ETH pre-attention work's 93–97% is the second.
The per-layer table is the substantive half. An aggregate flatters a prefetch, because what a prefetch costs is set by the layers it gets wrong, and those are not evenly spread: a MoE model's first layers route far less predictably than its last, their routing scores sitting close enough together that a slightly stale input reorders them. Both published results above report the same shape.
A predictor with no routings at a layer prints -, never 0.0.
What it structurally cannot do
- Layer 0 is out of reach. There is no previous layer to borrow an input from, so the stale gate
never predicts it — the table shows
-. This is the method's limit, not a bad score. It is also precisely the gap the ETH work exists to close, by training a small per-layer predictor that reads the layer's own pre-attention state instead. That needs training data and a shipped artifact per model, which is a different project from this one. - The first token of a run is unscored, because a layer only learns the next layer's matrix once the graph has reached it. Both cases are counted and reported, never folded into the denominator: a routing that was not ranked is not a wrong guess.
- Decode only. During prefill every token's routing is known at once and there is nothing to predict; it is also not the phase a prefetch would be built for.
- Routings wider than 32 experts are declined rather than scored against a truncated ranking.
Cost, and why it is not a benchmark run
The probe isolates one extra graph node per MoE layer (a barrier, as --route-trace costs) and runs
two GEMVs per layer on the eval thread — the prediction and its control. On a 48-layer model with
128 experts that is a few tens of millions of scalar multiply-accumulates per token, single
threaded. Correctness is unaffected and the byte-identity gates cover it (G9a), but do not read
tok/s off a probed run.
Requires streaming (--moe-stream): the probe attaches to the routing nodes the streamer already
isolates, and routing does not depend on how the weights reached memory, so a dense run would only
reproduce the same numbers more slowly.
Acting on it: --predict-prefetch
--predict-log measures; --predict-prefetch bets. It hands the stale-gate prediction for layer
l+1 to the same speculative read path --prefetch uses — same cache buffers, same
accounting, same moe-prefetch: summary line (tagged [stale-gate]) — replacing the previous-token
predictor with one measured roughly twice as accurate on the same run. Mutually exclusive with
--prefetch, and it needs the LRU cache for the same reason.
Two design points matter more than the swap itself:
- The speculation is issued after the current layer's load, not when the prediction is made. Every load path begins by quiescing speculation, and the layer's own load sits a few graph nodes after its gate — reads queued at prediction time would be cancelled before a lane picked them up. Issuing after the load gives a queued read the layer's whole expert-matmul window, the same window the temporal prefetch gets.
- It is drop-aware. With
--drop-cold-expertsarmed, a predicted expert whose predicted routing weight (softmax over the predicted top-k) falls below the drop threshold is not speculated: if it misses, the policy would drop it unread — prefetching it would spend the exact I/O the policy exists to save. The top prediction is always kept, mirroring the policy's own pin of the top-weighted expert. - It speculates at most the top 2 predicted misses per layer, not the routing. The stall a prefetch can remove is head-of-line — the wait on the first missing expert's slices — and overlap already hides the tail behind the expert matmul. Speculating a full top-8 was measured to read 33% more flash per token and to fight a full cache hard enough to triple major faults, costing several times the stall it removed. Predicted experts already resident do not burn a slot of the cap.
Byte-identity is the same contract as the temporal prefetch (gate G10): speculation only warms the
cache, so output is unchanged — except under --drop-cold-experts, where residency is an input to
the routing policy and a correct guess un-drops an expert. That interplay is deliberate: prediction
spent there buys back quality at the same threshold, not speed.
Cost: one gate GEMV per MoE layer per decoded token on the eval thread (the control GEMV is probe-only and skipped here), plus one isolated node per layer.
The honest caveat about what a good number would mean
A high stale-gate accuracy would say the routing is knowable earlier. It would not say that knowing it earlier makes decode faster. Those are different claims, and on this engine the second one is the doubtful one: when the flash is already saturated — which the decode attribution shows it is on Qwen at top-k 8 — starting a read sooner adds no bandwidth, and speculating extra reads spends the bandwidth that was the binding constraint. That is why temporal prefetch, layer-LFU caching and the expert sidecar all lost despite improving the metric each was designed around.
The use that does not spend bandwidth is eviction: knowing what the next layer will ask for is a reason not to discard it. That is a cache-policy change, not a prefetch, and offline replay bounds the whole class of online cache policies at roughly 5 percentage points over LRU. So even a predictor that scores well should be expected to buy little — which is exactly why it gets measured here first, at the cost of one flag, rather than built.