Commit graph

4 commits

Author SHA1 Message Date
Raffaele
47924565c1
feat(io): release the model file's mapping after load (--release-mmap) (#185)
llama.cpp maps every gguf it loads and keeps the mapping for the model's lifetime. On
Windows that is expensive in a way nothing had attributed: while a section of a file is
alive, NTFS serialises concurrent unbuffered reads on that file, and a lane opened while
the section existed keeps serialising against it after the section is gone. Four I/O lanes
therefore delivered exactly one lane's throughput, which is why lanes and threads have
always measured dead on the desktop host and why the engine read at about a third of what
the drive can serve.

--release-mmap hands the mapping back after load: unmap the file, close its section, reopen
the reader lanes. Both halves are needed. Whether it is safe is decided by looking rather
than by reasoning, the engine asks the OS whether any weight the capture pass observed still
points inside a mapping of the model files, and declines if any does. Off by default,
because the check answers for the pointers the capture saw and for no others.

Host A/B on Qwen3.6-35B-A3B Q4_K_M: 3.16 to 4.63 tok/s (+46%), flash stall per token 0.182
to 0.074, with bytes read, hit rate, evictions and re-reads identical to the digit and the
generated text byte-identical. On the phone the read path is flat (f2fs does not serialise)
but CPU per token falls about 9%; that cell is two short runs per variant and is recorded as
a direction, not a number.

Also fixes a bug this uncovered, independent of the flag: when a gguf carries no
output.weight, llama.cpp builds the output head from the token embedding table and the model
holds two identically named tensors over the same bytes. The capture pass keyed its map by
name, so --dense-weights anon and ahwb rebound one and left the twin reading the mmap for
the whole run, which on a model past RAM means the output projection served by page faults
from flash. The capture now records every distinct leaf object by address and the dense
policy rebinds every tensor over one file range onto the same buffer.

Adds bmoe-iobench --mmap / --reopen-lanes / --range-mb / --fresh, the cells that isolate the
mechanism, a mapping_release unit test on both platforms, the app switch "Release the model
mapping", and the bench findings. README, architecture, AGENTS and roadmap updated, the last
correcting a diagnosis this refutes.
2026-09-07 20:43:02 +02:00
Helldez
631d98f01a docs: record the expert-sidecar negative result and ship the --scatter microbench 2026-07-20 17:02:48 +02:00
Helldez
783d7b8f17 fix(tools): make the diagnostics fail instead of reporting a confident wrong number
Both of these tools produce figures quoted as evidence in docs/roadmap.md, and both had a path
that answered rather than errored when its input could not support an answer.

route-replay.py cost() defaulted a layer with no recorded expert_bytes to zero bytes. A free layer
is admitted without charge, never counts against the budget and is never evicted, so a trace whose
per-layer preamble is missing did not fail -- it printed a complete table in which every policy
scored the same near-perfect number. Reproduced on a gate-model trace with the preamble stripped:
the old code prints 96.9% across all six policies, the new code names the layer and exits. The
recorded traces behind the published curves all carry complete preambles, so no result in docs/
changes.

An unrecognised --policies name fell through every branch of victim() to the LRU default and was
tabulated under its own column header, claiming to compare a policy that never ran. argparse
choices= cannot express a comma-joined list, so the names are checked after the split.

bmoe-iobench hardcoded align = 4096 while alignment is the variable it exists to characterise; on
a 16 KiB-page device it measured the wrong one. It now calls pio::vm_page(), the engine's own
source. Its usage text also described --slice-kb as "bytes per read, default 4096" for a value
multiplied by 1024, leaving every printed figure open to being read off by 1024x.

Verified: replay accepted a real qwen3moe gate trace and reproduced the known LRU-cliff shape
(0% below one token cycle); rejected `--policies lur`; failed loudly on the stripped preamble.
bmoe-iobench rebuilt and swept, still negotiating O_DIRECT at the queried page size. Host gates
7/7, clang-format 18 clean.

No engine, CLI or app code is touched, so the Android version is unchanged.
2026-07-20 12:25:58 +02:00
Helldez
c8ca0b01d2
docs: measure the cache and I/O levers, and correct what the measurements contradict (#88)
Adds the two instruments the measurements were taken with, the data, and fixes to
the maintained docs that turned out to assert things that do not hold. No engine
code changes: the one candidate that was implemented is a measured regression and
stays on its branch.

tools/bmoe-iobench (new, BMOE_BUILD_TOOLS=OFF by default) sweeps flash read
bandwidth against lane count and read size. It drives bmoe::FileReader -- the
engine's own read path, so alignment, the bounce buffer and the O_DIRECT
verify/fallback are part of what is measured -- but links nothing else, since the
I/O layer has no llama.cpp dependency. It therefore cross-compiles in seconds and
cannot perturb the streamer. --compute-load adds CPU contention, because the
streamer reads while ggml's threads spin and an idle-CPU number is not the
condition it operates under.

scripts/route-replay.py (new, stdlib only) replays the committed route traces
through hypothetical cache policies at zero device cost. It reproduces the
recorded on-device hit rate to the decimal on all three captures and independently
predicts a historical budget-shrink measurement it was not calibrated against.

What they found, and what it invalidated:

- roadmap.md opened its read-bandwidth theme on the premise that effective
  O_DIRECT bandwidth sits far below the drive's ceiling because routed slices are
  scattered. Measured, reads saturate at 2 lanes and are flat above 256 KiB, so
  scatter is cheap here. Runtime read coalescing is retired (0.6-4 % of a layer's
  routed experts are id-adjacent); the expert-contiguous repack survives but needs
  a different justification.
- prefetch.md opened on "MoE routing has strong temporal locality". It is 17.9 %
  on gpt-oss, beaten there by a static hot list, and --prefetch 1 is a 2x
  slowdown. The mechanism and its correctness argument stand; the bet is annotated
  with what it returns and with the case still open (top-6 models).
- cache-sizing.md said a too-small budget makes the hit rate "collapse". It makes
  it exactly 0.0 %, and the boundary is one token cycle -- not cache_min_mb, which
  is a different quantity roughly n_layer smaller and protects today's models only
  by coincidence. Reproduced on device at a budget the CLI accepts.

bench-data/2026-07-20-cache-replay/ carries the three notes, the curves and every
run CSV, indexed by a README stating the six verdicts.
2026-07-20 11:44:48 +02:00