A first command with nothing but -m, -p and -t ran plain llama.cpp on mmap: streaming
off, cache off, and a dense policy that only applies once streaming is on. Nothing in
the report said so, because the moe-stream: block only prints when streaming is enabled,
so a baseline run read as a measurement of this engine and got reported as one (#186).
A mode: line is now printed on every run. With streaming off on a MoE architecture the
build has a recipe for it says the run is a baseline and names the flag; on any other
model it says the architecture is not one this build streams. RunSummary carries the
model's arch so the CLI can tell those apart.
--cache-mb defaults to auto whenever --moe-stream is on. The previous default of 0 meant
the cache was off, which re-reads every routed expert from flash every token. On a 16 GB
host streaming Qwen3.6-35B-A3B Q4_K_M, 63 tokens, -t 8 --overlap: 1.164 to 2.351 tok/s
and 585 to 238 MiB per token, same output. The default is resolved in the CLI, not in
the library, so an embedder passing 0 still means no cache; an explicit --cache-mb or
BMOE_CACHE_MB still wins, including an explicit 0.
Reported by @eiffel31.
Below one token's worst-case routed bytes, global LRU evicts precisely
what it is about to read: the hit rate is 0% while the run still pays
the cache's management time and its RAM. The only protection so far was
cache_min_mb, a fixed floor that happens to sit above the cycle for the
shipped models at their default top-k and stops holding the moment
--n-expert-used widens the routing without touching the budget.
The cycle is priced at init from the model's shape alone — every bound
layer's entry_bytes times min(top_k, n_expert), at the top-k the run
actually applies — and recorded as cache_cycle_mb in the metrics
preamble next to the budget it should be judged against, so a committed
CSV answers on its own whether its cache could ever have hit. The engine
prints one stderr line when the resolved budget falls under it. It warns
rather than refuses: the budget is legal and the output byte-identical,
and the engine states the fact and leaves the choice — the argument that
--cache-mb 0 is strictly better below the cliff stays in
docs/cache-sizing.md, not in the engine's output.
Docs updated in the same commit: cache-sizing.md (the guard section,
past tense), roadmap.md (the item moves from still-worth-doing to
shipped) and telemetry.md (the new preamble key).
Adds the two instruments the measurements were taken with, the data, and fixes to
the maintained docs that turned out to assert things that do not hold. No engine
code changes: the one candidate that was implemented is a measured regression and
stays on its branch.
tools/bmoe-iobench (new, BMOE_BUILD_TOOLS=OFF by default) sweeps flash read
bandwidth against lane count and read size. It drives bmoe::FileReader -- the
engine's own read path, so alignment, the bounce buffer and the O_DIRECT
verify/fallback are part of what is measured -- but links nothing else, since the
I/O layer has no llama.cpp dependency. It therefore cross-compiles in seconds and
cannot perturb the streamer. --compute-load adds CPU contention, because the
streamer reads while ggml's threads spin and an idle-CPU number is not the
condition it operates under.
scripts/route-replay.py (new, stdlib only) replays the committed route traces
through hypothetical cache policies at zero device cost. It reproduces the
recorded on-device hit rate to the decimal on all three captures and independently
predicts a historical budget-shrink measurement it was not calibrated against.
What they found, and what it invalidated:
- roadmap.md opened its read-bandwidth theme on the premise that effective
O_DIRECT bandwidth sits far below the drive's ceiling because routed slices are
scattered. Measured, reads saturate at 2 lanes and are flat above 256 KiB, so
scatter is cheap here. Runtime read coalescing is retired (0.6-4 % of a layer's
routed experts are id-adjacent); the expert-contiguous repack survives but needs
a different justification.
- prefetch.md opened on "MoE routing has strong temporal locality". It is 17.9 %
on gpt-oss, beaten there by a static hot list, and --prefetch 1 is a 2x
slowdown. The mechanism and its correctness argument stand; the bet is annotated
with what it returns and with the case still open (top-6 models).
- cache-sizing.md said a too-small budget makes the hit rate "collapse". It makes
it exactly 0.0 %, and the boundary is one token cycle -- not cache_min_mb, which
is a different quantity roughly n_layer smaller and protects today's models only
by coincidence. Reproduced on device at a budget the CLI accepts.
bench-data/2026-07-20-cache-replay/ carries the three notes, the curves and every
run CSV, indexed by a README stating the six verdicts.
--dense-weights anon allocates the whole dense set into anon buffers, and it
does so after auto has already sized the cache from MemAvailable. The document
claimed the budget was the engine's only large pinned allocation, which stopped
being true when anon became the default.
The 'adaptive' name outlived the retired runtime governor and read as if a
control loop were still there. The document describes load-time sizing, so
name it for that.