mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
feat(io): release the model file's mapping after load (--release-mmap) (#185)
llama.cpp maps every gguf it loads and keeps the mapping for the model's lifetime. On Windows that is expensive in a way nothing had attributed: while a section of a file is alive, NTFS serialises concurrent unbuffered reads on that file, and a lane opened while the section existed keeps serialising against it after the section is gone. Four I/O lanes therefore delivered exactly one lane's throughput, which is why lanes and threads have always measured dead on the desktop host and why the engine read at about a third of what the drive can serve. --release-mmap hands the mapping back after load: unmap the file, close its section, reopen the reader lanes. Both halves are needed. Whether it is safe is decided by looking rather than by reasoning, the engine asks the OS whether any weight the capture pass observed still points inside a mapping of the model files, and declines if any does. Off by default, because the check answers for the pointers the capture saw and for no others. Host A/B on Qwen3.6-35B-A3B Q4_K_M: 3.16 to 4.63 tok/s (+46%), flash stall per token 0.182 to 0.074, with bytes read, hit rate, evictions and re-reads identical to the digit and the generated text byte-identical. On the phone the read path is flat (f2fs does not serialise) but CPU per token falls about 9%; that cell is two short runs per variant and is recorded as a direction, not a number. Also fixes a bug this uncovered, independent of the flag: when a gguf carries no output.weight, llama.cpp builds the output head from the token embedding table and the model holds two identically named tensors over the same bytes. The capture pass keyed its map by name, so --dense-weights anon and ahwb rebound one and left the twin reading the mmap for the whole run, which on a model past RAM means the output projection served by page faults from flash. The capture now records every distinct leaf object by address and the dense policy rebinds every tensor over one file range onto the same buffer. Adds bmoe-iobench --mmap / --reopen-lanes / --range-mb / --fresh, the cells that isolate the mechanism, a mapping_release unit test on both platforms, the app switch "Release the model mapping", and the bench findings. README, architecture, AGENTS and roadmap updated, the last correcting a diagnosis this refutes.
This commit is contained in:
parent
92a1114917
commit
47924565c1
29 changed files with 1303 additions and 18 deletions
|
|
@ -170,6 +170,13 @@ Two worth knowing before you turn them on:
|
|||
- **"Stream row-gathered tables"** (`--row-stream`) serves the token embedding table out of flash
|
||||
instead of RAM. Lossless, and which tables it applies to is read off the model's own graph, so
|
||||
on a model where none qualify it does nothing. See `../../docs/row-gathered-tables.md`.
|
||||
- **"Release the model mapping"** (`--release-mmap`) hands the model file's mapping back to the
|
||||
kernel once every weight has been copied into the app's own memory, so it needs **Dense weights**
|
||||
on Anon or Pinned and is disabled otherwise. Lossless. The mechanism that makes it worth +46% on
|
||||
a Windows desktop does not exist here — on f2fs the read lanes measure the same either way — but
|
||||
keeping a 20 GB mapping registered costs a kernel under memory pressure, and dropping it took
|
||||
~9% off CPU per token. Two 48-token cells on a device that spreads 20%: a direction, not a
|
||||
number, which is why it is off by default.
|
||||
- **"Prefer cached experts"** (`--expert-substitute`) steers each routing toward experts already
|
||||
in RAM, so the same number of experts runs but fewer are read from flash. It changes the reply,
|
||||
and past 20% the reply keeps reading well while the model behind it is much worse: judge it on
|
||||
|
|
|
|||
|
|
@ -65,6 +65,18 @@ data class AppSettings(
|
|||
// so the only question it raises is whether the reads cost more than the RAM is worth - which
|
||||
// is why it is off until the on-device A/B says otherwise.
|
||||
val rowStream: Boolean = false,
|
||||
// Hand the model file's mapping back to the kernel once every weight has been rebound onto the
|
||||
// app's own memory. Needs a dense policy that does that rebinding (Anon or Pinned), which is why
|
||||
// the switch is disabled under Mmap and Warm — under those the engine looks at its own pointers,
|
||||
// sees weights still reading the mapping, and stands down anyway.
|
||||
//
|
||||
// The mechanism that makes this worth +46% on a Windows desktop (a live mapping serialises the
|
||||
// streamer's unbuffered reads) does NOT exist here: on f2fs the read lanes measure the same with
|
||||
// the mapping and without. What it buys on device is CPU — keeping a 20 GB mapping registered
|
||||
// costs a kernel under memory pressure, and dropping it took ~9% off CPU per token. That is two
|
||||
// 48-token cells against a device whose cells spread 20%, so it is a direction and not a number,
|
||||
// and the switch stays off until a 256-token A/B earns it.
|
||||
val releaseMmap: Boolean = false,
|
||||
// Cache-aware substitution, as a PERCENTAGE of the router's score range (0 = off). Before a
|
||||
// routing is committed, every expert already resident gets its score raised by this fraction of
|
||||
// the range and the top-k is taken again, so a resident expert wins a slot only when it was
|
||||
|
|
@ -187,6 +199,12 @@ data class AppSettings(
|
|||
// discovered by the streamer's capture pass; independent of the cache and of the
|
||||
// dense-weight mode, since what it changes is which tensors that mode applies to.
|
||||
if (rowStream) a += "--row-stream"
|
||||
// Only the policies that rebind every weight into the app's own memory can leave the
|
||||
// mapping unreferenced. Sending it under Mmap or Warm is not unsafe — the engine checks
|
||||
// its own pointers and declines — but it would be a switch that silently does nothing.
|
||||
if (releaseMmap && (denseWeights == DenseWeights.ANON || denseWeights == DenseWeights.AHWB)) {
|
||||
a += "--release-mmap"
|
||||
}
|
||||
// Same cacheOn guard and for the same reason: with no cache there is nothing resident
|
||||
// to substitute toward, so the policy would re-rank against an all-miss mask.
|
||||
if (substitutePct > 0 && cacheOn) a += listOf("--expert-substitute", (substitutePct / 100.0).toString())
|
||||
|
|
@ -236,6 +254,7 @@ data class AppSettings(
|
|||
.putInt("routeAhead", routeAhead)
|
||||
.putInt("dropColdPct", dropColdPct)
|
||||
.putBoolean("rowStream", rowStream)
|
||||
.putBoolean("releaseMmap", releaseMmap)
|
||||
.putInt("substitutePct", substitutePct)
|
||||
.putInt("sessionCtx", sessionCtx)
|
||||
.putString("spec", spec).putInt("mtpDraft", mtpDraft).putInt("mtpPMinPct", mtpPMinPct)
|
||||
|
|
@ -379,6 +398,7 @@ data class AppSettings(
|
|||
routeAhead = p.getInt("routeAhead", d.routeAhead),
|
||||
dropColdPct = p.getInt("dropColdPct", d.dropColdPct),
|
||||
rowStream = p.getBoolean("rowStream", d.rowStream),
|
||||
releaseMmap = p.getBoolean("releaseMmap", d.releaseMmap),
|
||||
substitutePct = p.getInt("substitutePct", d.substitutePct),
|
||||
sessionCtx = p.getInt("sessionCtx", d.sessionCtx),
|
||||
spec = run {
|
||||
|
|
|
|||
|
|
@ -127,6 +127,17 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
|
|||
enabled = stream,
|
||||
) { onChange(current.copy(rowStream = it)) }
|
||||
|
||||
SwitchRow(
|
||||
"Release the model mapping",
|
||||
"Once every weight has been copied into the app's own memory, the model file " +
|
||||
"does not need to stay mapped. Handing the mapping back frees the kernel " +
|
||||
"from tracking it, which showed up as less CPU per token. Lossless - the " +
|
||||
"output is identical. Needs Dense weights on Anon or Pinned.",
|
||||
current.releaseMmap,
|
||||
enabled = stream && (current.denseWeights == DenseWeights.ANON ||
|
||||
current.denseWeights == DenseWeights.AHWB),
|
||||
) { onChange(current.copy(releaseMmap = it)) }
|
||||
|
||||
ExperimentalGroup {
|
||||
IntSetting(
|
||||
"Temporal prefetch (layers)", AppSettings.PREFETCH_CHOICES, current.prefetchLayers,
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue