Below one token's worst-case routed bytes, global LRU evicts precisely
what it is about to read: the hit rate is 0% while the run still pays
the cache's management time and its RAM. The only protection so far was
cache_min_mb, a fixed floor that happens to sit above the cycle for the
shipped models at their default top-k and stops holding the moment
--n-expert-used widens the routing without touching the budget.
The cycle is priced at init from the model's shape alone — every bound
layer's entry_bytes times min(top_k, n_expert), at the top-k the run
actually applies — and recorded as cache_cycle_mb in the metrics
preamble next to the budget it should be judged against, so a committed
CSV answers on its own whether its cache could ever have hit. The engine
prints one stderr line when the resolved budget falls under it. It warns
rather than refuses: the budget is legal and the output byte-identical,
and the engine states the fact and leaves the choice — the argument that
--cache-mb 0 is strictly better below the cliff stays in
docs/cache-sizing.md, not in the engine's output.
Docs updated in the same commit: cache-sizing.md (the guard section,
past tense), roadmap.md (the item moves from still-worth-doing to
shipped) and telemetry.md (the new preamble key).
Adds the two instruments the measurements were taken with, the data, and fixes to
the maintained docs that turned out to assert things that do not hold. No engine
code changes: the one candidate that was implemented is a measured regression and
stays on its branch.
tools/bmoe-iobench (new, BMOE_BUILD_TOOLS=OFF by default) sweeps flash read
bandwidth against lane count and read size. It drives bmoe::FileReader -- the
engine's own read path, so alignment, the bounce buffer and the O_DIRECT
verify/fallback are part of what is measured -- but links nothing else, since the
I/O layer has no llama.cpp dependency. It therefore cross-compiles in seconds and
cannot perturb the streamer. --compute-load adds CPU contention, because the
streamer reads while ggml's threads spin and an idle-CPU number is not the
condition it operates under.
scripts/route-replay.py (new, stdlib only) replays the committed route traces
through hypothetical cache policies at zero device cost. It reproduces the
recorded on-device hit rate to the decimal on all three captures and independently
predicts a historical budget-shrink measurement it was not calibrated against.
What they found, and what it invalidated:
- roadmap.md opened its read-bandwidth theme on the premise that effective
O_DIRECT bandwidth sits far below the drive's ceiling because routed slices are
scattered. Measured, reads saturate at 2 lanes and are flat above 256 KiB, so
scatter is cheap here. Runtime read coalescing is retired (0.6-4 % of a layer's
routed experts are id-adjacent); the expert-contiguous repack survives but needs
a different justification.
- prefetch.md opened on "MoE routing has strong temporal locality". It is 17.9 %
on gpt-oss, beaten there by a static hot list, and --prefetch 1 is a 2x
slowdown. The mechanism and its correctness argument stand; the bet is annotated
with what it returns and with the case still open (top-6 models).
- cache-sizing.md said a too-small budget makes the hit rate "collapse". It makes
it exactly 0.0 %, and the boundary is one token cycle -- not cache_min_mb, which
is a different quantity roughly n_layer smaller and protects today's models only
by coincidence. Reproduced on device at a budget the CLI accepts.
bench-data/2026-07-20-cache-replay/ carries the three notes, the curves and every
run CSV, indexed by a README stating the six verdicts.
--dense-weights anon allocates the whole dense set into anon buffers, and it
does so after auto has already sized the cache from MemAvailable. The document
claimed the budget was the engine's only large pinned allocation, which stopped
being true when anon became the default.
The 'adaptive' name outlived the retired runtime governor and read as if a
control loop were still there. The document describes load-time sizing, so
name it for that.