mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
docs+android: temporal prefetch guide, telemetry, and settings row
Add docs/prefetch.md (the temporal-locality bet, the correctness argument, and telemetry), document the moe-prefetch summary line, record the feature in the changelog, extend bench-analyze with prefetch config rows for a device A/B, and expose a prefetch-depth row in the Android settings (session argv, so changing it reopens the session). Kotlin compiles.
This commit is contained in:
parent
4e02eb7584
commit
e1f8b977fc
6 changed files with 92 additions and 3 deletions
|
|
@ -7,6 +7,15 @@ Semantic Versioning.
|
|||
## [Unreleased]
|
||||
|
||||
### Added
|
||||
- **Temporal prefetch** (`--prefetch K`, env `BMOE_PREFETCH`): while a token computes layer *l*,
|
||||
the experts the previous token routed at layers *l+1…l+K* are read speculatively on the idle I/O
|
||||
lanes, so a correct guess turns the next layer's read into a cache hit. Requires the LRU cache.
|
||||
The speculative path never delays real work (workers drain it only as spare capacity and yield
|
||||
to real batches; all cache-state mutation stays on the eval thread) and never changes output (a
|
||||
speculative read is the identical read a real miss would issue). Gates G5a/b/c prove
|
||||
byte-identity, including the integrate-then-hit path. A `moe-prefetch:` summary line reports the
|
||||
speculative bytes and useful-hit rate; an Android settings row exposes the depth. See
|
||||
`docs/prefetch.md`.
|
||||
- **Session mode**: the engine can now load a model once and serve many prompts against it, with
|
||||
the expert LRU cache staying warm between prompts, instead of re-paying the model load and the
|
||||
cold-cache ramp on every generation. `run()` splits into a `Session` (`open` / `generate` /
|
||||
|
|
|
|||
50
docs/prefetch.md
Normal file
50
docs/prefetch.md
Normal file
|
|
@ -0,0 +1,50 @@
|
|||
# Temporal prefetch
|
||||
|
||||
The expert LRU cache is filled reactively: a layer's experts are read only once its router has
|
||||
selected them. That leaves the read on the critical path — the first tokens of a generation miss
|
||||
often and stall on flash (see [benchmarks.md](benchmarks.md)). Temporal prefetch reads ahead.
|
||||
|
||||
## The bet
|
||||
|
||||
MoE routing has strong temporal locality: the experts a token selects at layer *l* overlap
|
||||
heavily with the experts the **previous** token selected at the same layer. So while a token
|
||||
computes layer *l*, we speculatively read — on the otherwise-idle I/O lanes — the experts the
|
||||
previous token routed at layers *l+1 … l+K* (`--prefetch K`). A correct guess turns the next
|
||||
layer's read into a cache hit; a wrong guess only wastes a read. For the very first generated
|
||||
token, the predictor is the last prompt token's routing, recorded during prefill.
|
||||
|
||||
`K` is a depth, not a certainty: recall falls as you look further ahead, so small `K` (1–2)
|
||||
captures most of the benefit. Prefetch requires the LRU cache (`--cache-mb > 0`); the speculative
|
||||
slices land in the per-layer cache buffers.
|
||||
|
||||
## How it stays correct and out of the way
|
||||
|
||||
The speculative path never delays real work and never changes output:
|
||||
|
||||
- **Same bytes.** A speculative read is the *identical* read a real miss would issue — same file
|
||||
offset, same destination buffer (`lbuf_[p][il] + e*slice`). A prefetched expert is therefore
|
||||
bit-for-bit what a real read would produce; a later routing that hits it gets identical bytes.
|
||||
Integration cannot change output by construction. Gates **G5a/b/c** assert this.
|
||||
- **One writer of cache state.** All LRU mutation stays on the eval-callback thread.
|
||||
`prefetch()` (eval thread) commits pages and enqueues per-projection reads; **workers only read
|
||||
bytes**; `quiesce_spec()` (eval thread, at the next real `load_layer`) integrates the entries
|
||||
whose every projection finished and releases the rest. No lock on the hot path.
|
||||
- **Real work wins.** Workers drain the real batch first and only spend spare capacity on the
|
||||
speculative queue, yielding the instant a real batch appears. At each real load, in-flight
|
||||
speculation is cancelled (its reads disowned) before staging touches the cache, so speculation
|
||||
can never race the reads a token actually depends on, nor hold a page an eviction wants.
|
||||
|
||||
Because speculation only pays off when there is real flash latency to hide, it does nothing on a
|
||||
fast host with the model in page cache (the reads are cancelled before they start) — which is
|
||||
exactly correct. It is measured on-device.
|
||||
|
||||
## Telemetry
|
||||
|
||||
With `--prefetch K` the summary gains a `moe-prefetch:` line: speculative MiB read this
|
||||
generation, and how many prefetched experts a later routing actually used (`useful / prefetched`).
|
||||
A low useful rate means the look-ahead is too deep for this model, or the cache is too small to
|
||||
hold the speculation alongside the working set.
|
||||
|
||||
`--prefetch-sync` (debug) completes speculative reads synchronously on the eval thread — it
|
||||
defeats the latency hiding but makes integration deterministic, which the gates use to exercise
|
||||
the integrate-then-hit path on a host where the timing race otherwise never fires.
|
||||
|
|
@ -46,6 +46,16 @@ moe-cache: <pct>% hit, resident <mib> MiB
|
|||
`compute` in the `moe-stream:` line is the same residual described for `compute_ms` above; `cache
|
||||
mgmt` is the per-token mean of `mgmt_ms`.
|
||||
|
||||
With `--prefetch K` a `moe-prefetch:` line is added:
|
||||
|
||||
```
|
||||
moe-prefetch: <mib> MiB speculative, <useful>/<prefetched> experts useful (<pct>%)
|
||||
```
|
||||
|
||||
`<mib>` is the flash read done speculatively this generation (a subset of the total read),
|
||||
`<prefetched>` the experts fully read ahead, and `<useful>` how many of those a later routing
|
||||
actually hit. See [prefetch.md](prefetch.md).
|
||||
|
||||
Under `--overlap` the `moe-stream:` line additionally reports `stall_s/tok=<s>` — the mean
|
||||
wall time per token that compute threads waited for expert reads to complete. It is `0` in
|
||||
serial mode (where the read wait is already folded into decode time).
|
||||
|
|
|
|||
|
|
@ -14,6 +14,7 @@ data class AppSettings(
|
|||
val nPredict: Int = 48,
|
||||
val oDirect: Boolean = true, // bypass the page cache
|
||||
val overlap: Boolean = false,// prefetch experts while the layer computes (experimental)
|
||||
val prefetchLayers: Int = 0, // temporal prefetch depth K (0 = off); needs the cache
|
||||
val thinking: Boolean = false,// reasoning; off passes --no-think (enable_thinking=false)
|
||||
val loadAll: Boolean = false,// debug: read ALL experts each token (A/B baseline)
|
||||
) {
|
||||
|
|
@ -42,6 +43,7 @@ data class AppSettings(
|
|||
a += listOf("--io-threads", ioThreads.toString())
|
||||
if (!oDirect) a += "--no-odirect"
|
||||
if (overlap) a += "--overlap"
|
||||
if (prefetchLayers > 0 && cacheMb > 0) a += listOf("--prefetch", prefetchLayers.toString())
|
||||
if (loadAll) a += "--load-all"
|
||||
}
|
||||
return a
|
||||
|
|
@ -54,14 +56,15 @@ data class AppSettings(
|
|||
* excluded — they vary per request without touching the loaded model.
|
||||
*/
|
||||
fun sessionSignature(modelPath: String): String =
|
||||
listOf(modelPath, mmap, cacheMb, ioThreads, threads, oDirect, overlap, loadAll).joinToString("|")
|
||||
listOf(modelPath, mmap, cacheMb, ioThreads, threads, oDirect, overlap, prefetchLayers, loadAll)
|
||||
.joinToString("|")
|
||||
|
||||
fun save(ctx: Context) {
|
||||
ctx.prefs().edit()
|
||||
.putBoolean("mmap", mmap)
|
||||
.putInt("cacheMb", cacheMb).putInt("ioThreads", ioThreads).putInt("threads", threads)
|
||||
.putInt("nPredict", nPredict).putBoolean("oDirect", oDirect)
|
||||
.putBoolean("overlap", overlap)
|
||||
.putBoolean("overlap", overlap).putInt("prefetchLayers", prefetchLayers)
|
||||
.putBoolean("thinking", thinking).putBoolean("loadAll", loadAll)
|
||||
.apply()
|
||||
}
|
||||
|
|
@ -76,6 +79,7 @@ data class AppSettings(
|
|||
// thrashes and is slower than no cache — the engine rejects it. Use 0 or >= 2000.
|
||||
val CACHE_CHOICES = intArrayOf(0, 2000, 3000, 4000, 5000, 6000)
|
||||
val IO_CHOICES = intArrayOf(1, 2, 4, 8)
|
||||
val PREFETCH_CHOICES = intArrayOf(0, 1, 2, 4)
|
||||
val THREAD_CHOICES = intArrayOf(2, 4, 6, 8)
|
||||
val NPREDICT_CHOICES = intArrayOf(16, 32, 48, 64, 128, 256, 512, 1024, 2048)
|
||||
|
||||
|
|
@ -90,6 +94,7 @@ data class AppSettings(
|
|||
nPredict = p.getInt("nPredict", d.nPredict),
|
||||
oDirect = p.getBoolean("oDirect", d.oDirect),
|
||||
overlap = p.getBoolean("overlap", d.overlap),
|
||||
prefetchLayers = p.getInt("prefetchLayers", d.prefetchLayers),
|
||||
thinking = p.getBoolean("thinking", d.thinking),
|
||||
loadAll = p.getBoolean("loadAll", d.loadAll),
|
||||
)
|
||||
|
|
|
|||
|
|
@ -71,6 +71,16 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
|
|||
"Prefetch experts while the layer computes (experimental)",
|
||||
current.overlap, enabled = stream,
|
||||
) { onChange(current.copy(overlap = it)) }
|
||||
IntSetting(
|
||||
"Temporal prefetch (layers)", AppSettings.PREFETCH_CHOICES, current.prefetchLayers,
|
||||
format = { if (it == 0) "off" else "$it" },
|
||||
enabled = stream && current.cacheMb > 0,
|
||||
) { onChange(current.copy(prefetchLayers = it)) }
|
||||
Text(
|
||||
"Read the next K layers' likely experts on idle lanes, predicted from the " +
|
||||
"previous token. Needs the cache on.",
|
||||
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||
)
|
||||
}
|
||||
|
||||
Section("Compute") {
|
||||
|
|
|
|||
|
|
@ -10,7 +10,8 @@ import os, sys, statistics
|
|||
|
||||
BENCH = sys.argv[1] if len(sys.argv) > 1 else r"C:\Users\raffa\Documents\BigMoeOnEdge\.bench"
|
||||
ORDER = ["mmap", "stream", "c2000_l2", "c2000_l4", "c4000_l2", "c4000_l4",
|
||||
"stream_ov", "c4000_l4_ov"]
|
||||
"stream_ov", "c2000_l4_ov", "c4000_l4_ov",
|
||||
"c4000_l4_pf1", "c4000_l4_pf2", "c4000_l4_pf4"]
|
||||
LABEL = {
|
||||
"mmap": "solo mmap (no streaming)",
|
||||
"stream": "streaming O_DIRECT, cache 0, lane 4",
|
||||
|
|
@ -19,7 +20,11 @@ LABEL = {
|
|||
"c4000_l2": "streaming + cache 4000 MiB, lane 2",
|
||||
"c4000_l4": "streaming + cache 4000 MiB, lane 4",
|
||||
"stream_ov": "streaming O_DIRECT + overlap, cache 0, lane 4",
|
||||
"c2000_l4_ov": "streaming + cache 2000 MiB, lane 4, overlap",
|
||||
"c4000_l4_ov": "streaming + cache 4000 MiB, lane 4, overlap",
|
||||
"c4000_l4_pf1": "streaming + cache 4000 MiB, lane 4, prefetch 1",
|
||||
"c4000_l4_pf2": "streaming + cache 4000 MiB, lane 4, prefetch 2",
|
||||
"c4000_l4_pf4": "streaming + cache 4000 MiB, lane 4, prefetch 4",
|
||||
}
|
||||
|
||||
def pct(v, q):
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue