mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
feat(moe): measure how predictable expert routing is, then act on it — and refute it (#101)
* feat(moe): --predict-log, measure how predictable expert routing is
Temporal prefetch was built on a predictor nobody had priced. The
previous-token bet turned out to be right ~38% of the time on Qwen and
~18% on gpt-oss, which cannot pay for the reads it speculates. The
lesson was not that prefetch is impossible but that a predictor should
be measured before it is wired into anything. This adds the instrument,
not a policy.
--predict-log ranks each layer's experts a layer early, by running the
NEXT layer's router matrix on the CURRENT layer's gate input. The
residual stream barely moves between layers, so the stale input ranks
nearly as the real one will -- the mechanism FATE (arXiv 2502.12224)
reports 78.8% for. It is training-free and changes no model: the router
matrix is a dense weight already resident, and the prediction is one
GEMV per layer. The matrix is learned from the graph (the gate matmul's
first source) rather than looked up by tensor name, so no architecture
is named anywhere in the path; being a weight leaf, its pointer stays
valid into layers the current token has not reached.
Three predictors are scored against the routing the router actually
produced, so they are comparable on one run: the stale gate, the
previous-token bet --prefetch already places, and a zero-staleness
control. The control is the load-bearing part. It shares every line of
code with the prediction under test and differs only in using the
layer's own matrix, so it must reproduce the selection llama.cpp
computes from those same two tensors. A transposed matrix, a mis-strided
row or the wrong token of the batch collapses it toward chance while the
stale figure would stay superficially plausible; an architecture that
selects by something other than raw-logit ranking (an additive bias,
group-limited routing) puts it below 100% and by that much the stale
figure understates the method. The CLI says so rather than letting the
gap be blamed on staleness.
Reported per layer as well as in aggregate, because an aggregate
flatters a prefetch: what a prefetch costs is set by the layers it gets
wrong, and a MoE model's first layers route far less predictably than
its last. Denominators are printed per predictor -- the stale gate
structurally cannot rank layer 0 (nothing precedes it) or the first
token of a run, and those routings are counted as unscored rather than
folded in, since a routing that was not ranked is not a wrong guess. A
predictor with no routings at a layer prints "-", never 0.0.
Diagnostics only: nothing it computes reaches load_layer, the cache or
the graph, so a probed run reads exactly the bytes an unprobed one does.
G9a gates that byte identity and G9b gates the control, which reads
100.0% on the tiny model. It is not free -- one isolated node and two
GEMVs per layer on the eval thread -- so a probed run is not a benchmark
run, and it requires --moe-stream since routing does not depend on how
the weights reached memory.
docs/expert-prediction.md also records the caveat the number will need:
a high score would say the routing is knowable earlier, not that knowing
it earlier makes decode faster. On a flash already saturated, starting a
read sooner adds no bandwidth -- which is why prefetch, layer-LFU and
the expert sidecar all lost despite improving the metric each was
designed around.
* feat(moe): --predict-prefetch, speculate on the stale-gate prediction
The probe said the routing is knowable a layer early (~89% of routed
slots on a 128-expert model, vs ~43% for the previous-token bet the
temporal prefetch acts on). This wires that prediction into the existing
speculative read path -- same cache buffers, same accounting, same
settle, same moe-prefetch summary line (tagged [stale-gate]) -- so the
only thing that changes is which guess rides the idle lanes.
Two decisions carry the design:
Speculation is issued AFTER the current layer's load, not at prediction
time. Every load path begins by quiescing speculation, and a layer's own
load sits a few graph nodes after its gate matmul -- reads queued at
prediction time would be cancelled before a lane picked them up. Each of
the three load sites (plain topk, deferred drop, drop fallback) issues
the pending next-layer prediction right after its load_layer, restoring
the same read-ahead window the temporal prefetch gets. On the tiny-model
gates this is the difference between 0% and 27% of speculated experts
proving useful -- the latter matching the probe's measured accuracy on
that model, which is the accounting agreeing with itself.
It is drop-aware. With --drop-cold-experts armed, a predicted expert
whose predicted routing weight (softmax over the predicted top-k)
falls below the drop threshold is not speculated: if it misses, the
policy discards it unread, so reading it ahead would spend the exact
I/O the policy exists to save. The top prediction is always kept,
mirroring the policy's own pin of the top-weighted expert. The known
interplay is inherited from the temporal prefetch and deliberate: a
correct guess un-drops an expert, buying quality at the same threshold
rather than speed.
The routing width the prefetch predicts at is learned from the topk
node, not read from config, so an --n-expert-used override stays honest
with no extra plumbing. Mutually exclusive with --prefetch (two
predictors would double-speculate the same future); requires the LRU
cache; decode only. The control GEMV remains probe-only, so the
production path costs one gate GEMV per MoE layer per token.
Gates: G10a proves byte-identity through the speculative path
(prefetch-sync, forced small cache, hits and evictions both occur);
G10b proves the run actually speculated and that useful-hit accounting
tracks the probe's accuracy. Off by default, pending an on-device A/B.
* feat(moe): cap predictive speculation at the top 2 predicted misses per layer
Speculating the whole predicted routing was measured on device at -38%
against its own baseline despite every intermediate metric improving
(hit 77.6->88.9%, stall 32->9 ms, 84% of speculations useful): the +33%
flash bytes and the vm commits fighting a full cache (major faults x3)
cost several times the stall removed. The stall a prefetch can remove is
head-of-line only -- overlap already hides the tail behind the expert
matmul -- so the cap keeps the part of the bet that can pay and drops
the part that provably cannot. Residency-aware: a predicted expert
already in cache does not burn a slot of the cap.
* perf(moe): rebuild the predictive prefetch around its measured costs
The observer-tax run priced the naive implementation: ~35-45 ms per
GEMV pass (ggml_fp16_to_fp32 is a function call per weight element --
21M calls/token) and ~20 ms/token for the extra isolated node, against
a speculation machinery that costs ~15-25. The GEMV and the barrier
were the feature; this commit removes both.
- gate_scores converts F16 natively on aarch64 (one instruction,
vectorizable) instead of a function call per element.
- The prefetch no longer isolates the gate matmul: the ask pass hands
over its source pointers for free, and the gate-input row is read at
the topk callback with no barrier of its own. A sampled watchdog
(the zero-staleness control, every 512 routings) validates the
barrier-less read and disarms the prefetch out loud if the memory
planner ever reuses that buffer -- without it, a future llama.cpp
bump could silently turn the predictor into a noise generator.
- The GEMV runs on a dedicated worker at a TWO-layer horizon: one
layer ahead has no landing spot (the callbacks between a layer's
gate and its own load are microseconds apart), while at l+2 the
worker has a whole layer for a ~0.5M-MAC job and the result inherits
the same post-load issue window as before. The probe now also scores
stale-2, so the extra layer of staleness is priced per model rather
than assumed.
- The prediction's residents are RETAINED (new IExpertSource::retain:
move-to-MRU, deliberately not a cache hit so the hit-rate metric
stays honest) -- protecting a predicted expert costs zero bytes,
unlike prefetching it. Only predicted misses are speculated, still
capped at 2.
The probe path keeps its barrier and both GEMVs: a probed run is
diagnostics, and its job is to be right, not fast.
* feat(moe): --predict-spec-max N — how much flash the prediction may spend (0 = retention only)
The cap was a constant; the retention-only point (0) is the config the
whole experiment now hinges on -- the prediction protecting predicted
residents from eviction while spending no flash at all -- and a
measured constant that cannot be varied is not a mechanism. Validated
[0, 8]; retention happens at every value because it is free.
* docs(predict): record the 2026-07-23 campaign — accuracy confirmed, throughput verdict open
Accuracy on device: stale-gate 88.6% (Qwen3-30B) / 80.7% (Qwen3.6),
control exactly 100.0% on every layer of both, prev-token 43/35% --
corroborating the offline route-trace estimates.
Throughput: the day's C-vs-B losses are recorded WITH their
invalidation. Re-running the reference on the by-then-hot device gave
3.93 tok/s against the cool morning's 6.54 with byte-identical I/O,
hit and drop counts -- the engine is deterministic, the -40% was
silent thermal capping, and every variant had been compared against
the cool number. The one thermally matched pair that was measured
(speculation on 8 io lanes, -28%, effective flash bandwidth 585->392
MiB/s) kills the more-lanes hypothesis specifically; spec-max 2 and
retention-only still owe a matched cool pair.
What did survive: the observer-tax decomposition (109 ms/token: the
per-element exported-function F16 conversion at 21M calls/token, plus
the barrier), and retention moving hit rate 0.1pp on a 3000 MiB cache
-- the offline replay bound confirmed from inside the engine.
* feat(app): expose the predictive prefetch as an experimental Streaming toggle
Off by default. Gated on streaming + a live cache like the temporal
prefetch, and the two settings disable each other in the UI -- the
engine refuses the pair, and a control the engine will reject is worse
than one that cannot be set. The spec-max rung selector (0/1/2/4)
surfaces the retention-only point, which is the configuration the open
throughput question most needs measured from the app. Session
signature includes both fields so flipping them reopens the process.
No versionCode bump: this is a PR-branch test build, not a release.
* docs(predict): record the 2026-07-24 matched pairs — read-ahead refuted, retention hit-neutral
A four-cell session (B, retention-only, spec-2, B sentinel) run at fixed 30 s
spacing re-proved the thermal-contamination mechanism (sentinel −17%, clusters
silently capped from cell 2) and yielded one genuinely matched pair: spec-2 vs
the B sentinel at the same caps and battery temperature, 3.14 vs 3.96 tok/s
(−21%) with hit rate up 4.3pp and 79% of speculations useful. Speculation
improves every metric it owns and still loses the wall clock — the flash has no
spare bandwidth to spend. Retention-only again moved the hit rate by nothing
(77.2% vs 77.6%), as the offline replay bound predicted.
* feat(app): contrast the two prefetch predictors in the UI, default spec-max to 0
The temporal and predictive toggles now say what actually differs — the bet
("repeats the previous token", ~40%) vs the question ("ask the next router a
layer early", ~85%) — instead of describing mechanisms side by side. Spec-max
defaults to 0 (retention only): the matched-pair A/B showed the read-ahead
losing −21% on a saturated flash, so 0 is the only rung the measurements did
not refute, and the helper text says so.
* chore(predict): file the changelog under Unreleased, drop a dead member, record the verdict
Three loose ends found reviewing the branch for merge:
- The changelog entries had been appended to the already-released 0.15.1 section;
they belong under [Unreleased], where the release commit carves them out.
- nu_hint_ was written on every routing and read nowhere: a leftover of the
synchronous first design, whose successor passes the routing width straight into
the prediction job.
- docs/roadmap.md still closed the routing-prediction question on the 2026-07-12
removal. It now records what reopening it with a training-free predictor found:
the accuracy is real and the throughput is not, for the same reason more lanes
and the sidecar lost.
This commit is contained in:
parent
7eff9ee47b
commit
9e04e04d53
22 changed files with 1293 additions and 17 deletions
|
|
@ -94,3 +94,13 @@ expert cache at 4000 MiB, 4 I/O lanes and 4 compute threads, decode settles arou
|
|||
sweep point from the benchmark protocol, not the app default: the app ships a fixed 2000 MiB
|
||||
expert cache. See `../../docs/benchmark-method.md` for the full procedure and the cache/thread
|
||||
sweep.
|
||||
|
||||
The Streaming section also exposes the **predictive prefetch** (experimental, off by default):
|
||||
the engine predicts each layer's experts one layer early and reads ahead / retains what the
|
||||
prediction names (`--predict-prefetch`, with "predicted misses to read ahead" mapping to
|
||||
`--predict-spec-max`; 0 = retention only, the app default). It needs the cache on and replaces
|
||||
the temporal prefetch — the two are mutually exclusive, and they differ only in the predictor:
|
||||
temporal bets each layer repeats the previous token's experts (~40% right), predictive asks the
|
||||
next layer's own router one layer early (~85% right). A better guess did not buy throughput:
|
||||
in thermally matched pairs the read-ahead **lost** (−21% at spec-max 2), because the flash is
|
||||
already saturated — see `../../docs/expert-prediction.md` before drawing conclusions from a run.
|
||||
|
|
|
|||
|
|
@ -34,6 +34,15 @@ data class AppSettings(
|
|||
val overlap: Boolean = true, // read the next experts while the current layer computes
|
||||
val denseWeights: DenseWeights = DenseWeights.ANON, // dense (non-expert) weight residency policy
|
||||
val prefetchLayers: Int = 0, // temporal prefetch depth K (0 = off); needs the cache
|
||||
// Predictive prefetch (experimental): run the NEXT layer's router on the current layer's
|
||||
// input and speculate/retain on that prediction instead of the previous token's routing.
|
||||
// Needs the cache; the engine rejects it combined with the temporal prefetch above, so
|
||||
// sessionArgv only emits it when prefetchLayers == 0.
|
||||
val predictPrefetch: Boolean = false,
|
||||
// How many predicted MISSES per layer it may read ahead (0 = retention only: the prediction
|
||||
// spends no flash and only protects predicted residents from eviction). Defaults to 0 because
|
||||
// the matched-pair A/B showed read-ahead losing on a saturated flash (docs/expert-prediction.md).
|
||||
val predictSpecMax: Int = 0,
|
||||
// Cache-aware expert dropping, as a PERCENTAGE of the uniform share 1/top-k (0 = off, 100 = the
|
||||
// share itself). Stored as an Int because the settings are integer rungs; the flag takes a
|
||||
// fraction. LOSSY and cache-dependent — it changes the output, and not reproducibly.
|
||||
|
|
@ -93,6 +102,13 @@ data class AppSettings(
|
|||
// Auto sizing is a live LRU cache, so it satisfies the prefetch cache requirement.
|
||||
val cacheOn = cacheMb == CACHE_AUTO || cacheMb > 0
|
||||
if (prefetchLayers > 0 && cacheOn) a += listOf("--prefetch", prefetchLayers.toString())
|
||||
// Predictive prefetch shares the cache requirement and excludes the temporal one —
|
||||
// two predictors would double-speculate the same future, and the engine refuses the
|
||||
// pair. Temporal wins when both are somehow set; the UI keeps them exclusive anyway.
|
||||
if (predictPrefetch && cacheOn && prefetchLayers == 0) {
|
||||
a += "--predict-prefetch"
|
||||
a += listOf("--predict-spec-max", predictSpecMax.toString())
|
||||
}
|
||||
// Cache-aware dropping needs a live cache to ask about residency — with the cache off
|
||||
// every expert reads as a miss and the engine rejects the combination outright, so the
|
||||
// same cacheOn condition that guards prefetch guards this. The engine takes a fraction
|
||||
|
|
@ -110,7 +126,7 @@ data class AppSettings(
|
|||
*/
|
||||
fun sessionSignature(modelPath: String): String =
|
||||
listOf(modelPath, mmap, cacheMb, cacheCeilMb, ioThreads, threads, nExpertUsed, oDirect,
|
||||
overlap, denseWeights, prefetchLayers, dropColdPct)
|
||||
overlap, denseWeights, prefetchLayers, predictPrefetch, predictSpecMax, dropColdPct)
|
||||
.joinToString("|")
|
||||
|
||||
fun save(ctx: Context) {
|
||||
|
|
@ -123,6 +139,8 @@ data class AppSettings(
|
|||
.putBoolean("overlap", overlap)
|
||||
.putString("denseWeights", denseWeights.name)
|
||||
.putInt("prefetchLayers", prefetchLayers)
|
||||
.putBoolean("predictPrefetch", predictPrefetch)
|
||||
.putInt("predictSpecMax", predictSpecMax)
|
||||
.putInt("dropColdPct", dropColdPct)
|
||||
.putBoolean("thinking", thinking)
|
||||
.putBoolean("metricsCsv", metricsCsv)
|
||||
|
|
@ -188,6 +206,10 @@ data class AppSettings(
|
|||
// 0 = model default (top-k as trained). 6/4/3/2 trade output quality for tok/s (fewer routed experts).
|
||||
val N_EXPERT_CHOICES = intArrayOf(0, 6, 4, 3, 2)
|
||||
val PREFETCH_CHOICES = intArrayOf(0, 1, 2, 4)
|
||||
// Speculated predicted misses per layer. 0 = retention only (zero flash spent) and the app
|
||||
// default — the matched-pair A/B showed 2 losing −21% on a saturated flash; anything above
|
||||
// it re-buys the measured full-speculation pathology.
|
||||
val PREDICT_SPEC_CHOICES = intArrayOf(0, 1, 2, 4)
|
||||
// Percent of the uniform share 1/top-k. 100 is the share itself and the useful maximum:
|
||||
// above it the threshold could exceed every weight in a routing. The rungs below it are the
|
||||
// conservative half of the curve, where the replay already beats a top-k cut on both axes.
|
||||
|
|
@ -220,6 +242,8 @@ data class AppSettings(
|
|||
}
|
||||
},
|
||||
prefetchLayers = p.getInt("prefetchLayers", d.prefetchLayers),
|
||||
predictPrefetch = p.getBoolean("predictPrefetch", d.predictPrefetch),
|
||||
predictSpecMax = p.getInt("predictSpecMax", d.predictSpecMax),
|
||||
dropColdPct = p.getInt("dropColdPct", d.dropColdPct),
|
||||
thinking = p.getBoolean("thinking", d.thinking),
|
||||
metricsCsv = p.getBoolean("metricsCsv", d.metricsCsv),
|
||||
|
|
|
|||
|
|
@ -119,12 +119,40 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
|
|||
IntSetting(
|
||||
"Temporal prefetch (layers)", AppSettings.PREFETCH_CHOICES, current.prefetchLayers,
|
||||
format = { if (it == 0) "off" else "$it" },
|
||||
enabled = stream && cacheOn,
|
||||
// Mutually exclusive with predictive prefetch: two predictors would speculate
|
||||
// the same future twice, and the engine refuses the pair.
|
||||
enabled = stream && cacheOn && !current.predictPrefetch,
|
||||
) { onChange(current.copy(prefetchLayers = it)) }
|
||||
Text(
|
||||
"Experimental. Prefetch the next K layers' likely experts on idle read lanes. Needs the cache on.",
|
||||
"Experimental. Bets each layer will reuse the experts it picked for the previous " +
|
||||
"token and reads them ahead on idle lanes — right ~40% of the time. Needs the cache on.",
|
||||
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||
)
|
||||
SwitchRow(
|
||||
"Predictive prefetch (experimental)",
|
||||
"Instead of betting on the previous token, asks the next layer's own router one " +
|
||||
"layer early — right ~85% of the time. Reads ahead within the budget below and " +
|
||||
"keeps what the prediction names. Needs the cache; replaces temporal prefetch.",
|
||||
current.predictPrefetch, enabled = stream && cacheOn && current.prefetchLayers == 0,
|
||||
) { onChange(current.copy(predictPrefetch = it)) }
|
||||
if (current.predictPrefetch) {
|
||||
IntSetting(
|
||||
"Predicted misses to read ahead", AppSettings.PREDICT_SPEC_CHOICES, current.predictSpecMax,
|
||||
format = {
|
||||
when (it) {
|
||||
0 -> "0 — retention only (default)"
|
||||
else -> "$it"
|
||||
}
|
||||
},
|
||||
enabled = stream && cacheOn,
|
||||
) { onChange(current.copy(predictSpecMax = it)) }
|
||||
Text(
|
||||
"How much flash the prediction may spend per layer. 0 reads nothing ahead and only " +
|
||||
"protects predicted experts already in RAM from eviction. Measured: reading ahead " +
|
||||
"lost its matched A/B — the flash has no spare bandwidth — so 0 is the honest default.",
|
||||
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
Section("Speed / quality") {
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue