docs+android: temporal prefetch guide, telemetry, and settings row

Add docs/prefetch.md (the temporal-locality bet, the correctness argument, and telemetry),
document the moe-prefetch summary line, record the feature in the changelog, extend
bench-analyze with prefetch config rows for a device A/B, and expose a prefetch-depth row in the
Android settings (session argv, so changing it reopens the session). Kotlin compiles.
This commit is contained in:
Helldez 2026-07-12 20:08:14 +02:00
parent 4e02eb7584
commit e1f8b977fc
6 changed files with 92 additions and 3 deletions

View file

@ -7,6 +7,15 @@ Semantic Versioning.
## [Unreleased]
### Added
- **Temporal prefetch** (`--prefetch K`, env `BMOE_PREFETCH`): while a token computes layer *l*,
the experts the previous token routed at layers *l+1…l+K* are read speculatively on the idle I/O
lanes, so a correct guess turns the next layer's read into a cache hit. Requires the LRU cache.
The speculative path never delays real work (workers drain it only as spare capacity and yield
to real batches; all cache-state mutation stays on the eval thread) and never changes output (a
speculative read is the identical read a real miss would issue). Gates G5a/b/c prove
byte-identity, including the integrate-then-hit path. A `moe-prefetch:` summary line reports the
speculative bytes and useful-hit rate; an Android settings row exposes the depth. See
`docs/prefetch.md`.
- **Session mode**: the engine can now load a model once and serve many prompts against it, with
the expert LRU cache staying warm between prompts, instead of re-paying the model load and the
cold-cache ramp on every generation. `run()` splits into a `Session` (`open` / `generate` /

50
docs/prefetch.md Normal file
View file

@ -0,0 +1,50 @@
# Temporal prefetch
The expert LRU cache is filled reactively: a layer's experts are read only once its router has
selected them. That leaves the read on the critical path — the first tokens of a generation miss
often and stall on flash (see [benchmarks.md](benchmarks.md)). Temporal prefetch reads ahead.
## The bet
MoE routing has strong temporal locality: the experts a token selects at layer *l* overlap
heavily with the experts the **previous** token selected at the same layer. So while a token
computes layer *l*, we speculatively read — on the otherwise-idle I/O lanes — the experts the
previous token routed at layers *l+1 … l+K* (`--prefetch K`). A correct guess turns the next
layer's read into a cache hit; a wrong guess only wastes a read. For the very first generated
token, the predictor is the last prompt token's routing, recorded during prefill.
`K` is a depth, not a certainty: recall falls as you look further ahead, so small `K` (1–2)
captures most of the benefit. Prefetch requires the LRU cache (`--cache-mb > 0`); the speculative
slices land in the per-layer cache buffers.
## How it stays correct and out of the way
The speculative path never delays real work and never changes output:
- **Same bytes.** A speculative read is the *identical* read a real miss would issue — same file
offset, same destination buffer (`lbuf_[p][il] + e*slice`). A prefetched expert is therefore
bit-for-bit what a real read would produce; a later routing that hits it gets identical bytes.
Integration cannot change output by construction. Gates **G5a/b/c** assert this.
- **One writer of cache state.** All LRU mutation stays on the eval-callback thread.
`prefetch()` (eval thread) commits pages and enqueues per-projection reads; **workers only read
bytes**; `quiesce_spec()` (eval thread, at the next real `load_layer`) integrates the entries
whose every projection finished and releases the rest. No lock on the hot path.
- **Real work wins.** Workers drain the real batch first and only spend spare capacity on the
speculative queue, yielding the instant a real batch appears. At each real load, in-flight
speculation is cancelled (its reads disowned) before staging touches the cache, so speculation
can never race the reads a token actually depends on, nor hold a page an eviction wants.
Because speculation only pays off when there is real flash latency to hide, it does nothing on a
fast host with the model in page cache (the reads are cancelled before they start) — which is
exactly correct. It is measured on-device.
## Telemetry
With `--prefetch K` the summary gains a `moe-prefetch:` line: speculative MiB read this
generation, and how many prefetched experts a later routing actually used (`useful / prefetched`).
A low useful rate means the look-ahead is too deep for this model, or the cache is too small to
hold the speculation alongside the working set.
`--prefetch-sync` (debug) completes speculative reads synchronously on the eval thread — it
defeats the latency hiding but makes integration deterministic, which the gates use to exercise
the integrate-then-hit path on a host where the timing race otherwise never fires.

View file

@ -46,6 +46,16 @@ moe-cache: <pct>% hit, resident <mib> MiB
`compute` in the `moe-stream:` line is the same residual described for `compute_ms` above; `cache
mgmt` is the per-token mean of `mgmt_ms`.
With `--prefetch K` a `moe-prefetch:` line is added:
```
moe-prefetch: <mib> MiB speculative, <useful>/<prefetched> experts useful (<pct>%)
```
`<mib>` is the flash read done speculatively this generation (a subset of the total read),
`<prefetched>` the experts fully read ahead, and `<useful>` how many of those a later routing
actually hit. See [prefetch.md](prefetch.md).
Under `--overlap` the `moe-stream:` line additionally reports `stall_s/tok=<s>` — the mean
wall time per token that compute threads waited for expert reads to complete. It is `0` in
serial mode (where the read wait is already folded into decode time).

View file

@ -14,6 +14,7 @@ data class AppSettings(
val nPredict: Int = 48,
val oDirect: Boolean = true, // bypass the page cache
val overlap: Boolean = false,// prefetch experts while the layer computes (experimental)
val prefetchLayers: Int = 0, // temporal prefetch depth K (0 = off); needs the cache
val thinking: Boolean = false,// reasoning; off passes --no-think (enable_thinking=false)
val loadAll: Boolean = false,// debug: read ALL experts each token (A/B baseline)
) {
@ -42,6 +43,7 @@ data class AppSettings(
a += listOf("--io-threads", ioThreads.toString())
if (!oDirect) a += "--no-odirect"
if (overlap) a += "--overlap"
if (prefetchLayers > 0 && cacheMb > 0) a += listOf("--prefetch", prefetchLayers.toString())
if (loadAll) a += "--load-all"
}
return a
@ -54,14 +56,15 @@ data class AppSettings(
* excluded — they vary per request without touching the loaded model.
*/
fun sessionSignature(modelPath: String): String =
listOf(modelPath, mmap, cacheMb, ioThreads, threads, oDirect, overlap, loadAll).joinToString("|")
listOf(modelPath, mmap, cacheMb, ioThreads, threads, oDirect, overlap, prefetchLayers, loadAll)
.joinToString("|")
fun save(ctx: Context) {
ctx.prefs().edit()
.putBoolean("mmap", mmap)
.putInt("cacheMb", cacheMb).putInt("ioThreads", ioThreads).putInt("threads", threads)
.putInt("nPredict", nPredict).putBoolean("oDirect", oDirect)
.putBoolean("overlap", overlap)
.putBoolean("overlap", overlap).putInt("prefetchLayers", prefetchLayers)
.putBoolean("thinking", thinking).putBoolean("loadAll", loadAll)
.apply()
}
@ -76,6 +79,7 @@ data class AppSettings(
// thrashes and is slower than no cache — the engine rejects it. Use 0 or >= 2000.
val CACHE_CHOICES = intArrayOf(0, 2000, 3000, 4000, 5000, 6000)
val IO_CHOICES = intArrayOf(1, 2, 4, 8)
val PREFETCH_CHOICES = intArrayOf(0, 1, 2, 4)
val THREAD_CHOICES = intArrayOf(2, 4, 6, 8)
val NPREDICT_CHOICES = intArrayOf(16, 32, 48, 64, 128, 256, 512, 1024, 2048)
@ -90,6 +94,7 @@ data class AppSettings(
nPredict = p.getInt("nPredict", d.nPredict),
oDirect = p.getBoolean("oDirect", d.oDirect),
overlap = p.getBoolean("overlap", d.overlap),
prefetchLayers = p.getInt("prefetchLayers", d.prefetchLayers),
thinking = p.getBoolean("thinking", d.thinking),
loadAll = p.getBoolean("loadAll", d.loadAll),
)

View file

@ -71,6 +71,16 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
"Prefetch experts while the layer computes (experimental)",
current.overlap, enabled = stream,
) { onChange(current.copy(overlap = it)) }
IntSetting(
"Temporal prefetch (layers)", AppSettings.PREFETCH_CHOICES, current.prefetchLayers,
format = { if (it == 0) "off" else "$it" },
enabled = stream && current.cacheMb > 0,
) { onChange(current.copy(prefetchLayers = it)) }
Text(
"Read the next K layers' likely experts on idle lanes, predicted from the " +
"previous token. Needs the cache on.",
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
Section("Compute") {

View file

@ -10,7 +10,8 @@ import os, sys, statistics
BENCH = sys.argv[1] if len(sys.argv) > 1 else r"C:\Users\raffa\Documents\BigMoeOnEdge\.bench"
ORDER = ["mmap", "stream", "c2000_l2", "c2000_l4", "c4000_l2", "c4000_l4",
"stream_ov", "c4000_l4_ov"]
"stream_ov", "c2000_l4_ov", "c4000_l4_ov",
"c4000_l4_pf1", "c4000_l4_pf2", "c4000_l4_pf4"]
LABEL = {
"mmap": "solo mmap (no streaming)",
"stream": "streaming O_DIRECT, cache 0, lane 4",
@ -19,7 +20,11 @@ LABEL = {
"c4000_l2": "streaming + cache 4000 MiB, lane 2",
"c4000_l4": "streaming + cache 4000 MiB, lane 4",
"stream_ov": "streaming O_DIRECT + overlap, cache 0, lane 4",
"c2000_l4_ov": "streaming + cache 2000 MiB, lane 4, overlap",
"c4000_l4_ov": "streaming + cache 4000 MiB, lane 4, overlap",
"c4000_l4_pf1": "streaming + cache 4000 MiB, lane 4, prefetch 1",
"c4000_l4_pf2": "streaming + cache 4000 MiB, lane 4, prefetch 2",
"c4000_l4_pf4": "streaming + cache 4000 MiB, lane 4, prefetch 4",
}
def pct(v, q):