BigMoeOnEdge/scripts
Helldez c8ca0b01d2
docs: measure the cache and I/O levers, and correct what the measurements contradict (#88)
Adds the two instruments the measurements were taken with, the data, and fixes to
the maintained docs that turned out to assert things that do not hold. No engine
code changes: the one candidate that was implemented is a measured regression and
stays on its branch.

tools/bmoe-iobench (new, BMOE_BUILD_TOOLS=OFF by default) sweeps flash read
bandwidth against lane count and read size. It drives bmoe::FileReader -- the
engine's own read path, so alignment, the bounce buffer and the O_DIRECT
verify/fallback are part of what is measured -- but links nothing else, since the
I/O layer has no llama.cpp dependency. It therefore cross-compiles in seconds and
cannot perturb the streamer. --compute-load adds CPU contention, because the
streamer reads while ggml's threads spin and an idle-CPU number is not the
condition it operates under.

scripts/route-replay.py (new, stdlib only) replays the committed route traces
through hypothetical cache policies at zero device cost. It reproduces the
recorded on-device hit rate to the decimal on all three captures and independently
predicts a historical budget-shrink measurement it was not calibrated against.

What they found, and what it invalidated:

- roadmap.md opened its read-bandwidth theme on the premise that effective
  O_DIRECT bandwidth sits far below the drive's ceiling because routed slices are
  scattered. Measured, reads saturate at 2 lanes and are flat above 256 KiB, so
  scatter is cheap here. Runtime read coalescing is retired (0.6-4 % of a layer's
  routed experts are id-adjacent); the expert-contiguous repack survives but needs
  a different justification.
- prefetch.md opened on "MoE routing has strong temporal locality". It is 17.9 %
  on gpt-oss, beaten there by a static hot list, and --prefetch 1 is a 2x
  slowdown. The mechanism and its correctness argument stand; the bet is annotated
  with what it returns and with the case still open (top-6 models).
- cache-sizing.md said a too-small budget makes the hit rate "collapse". It makes
  it exactly 0.0 %, and the boundary is one token cycle -- not cache_min_mb, which
  is a different quantity roughly n_layer smaller and protects today's models only
  by coincidence. Reproduced on device at a budget the CLI accepts.

bench-data/2026-07-20-cache-replay/ carries the three notes, the curves and every
run CSV, indexed by a README stating the six verdicts.
2026-07-20 11:44:48 +02:00
..
bench-analyze.py chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
bench-lib.ps1 chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
bench-matrix-rework.ps1 chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
bench-matrix.ps1 chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
bench-pr23-c2000.ps1 chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
bench-pr23-summary.py chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
bench-prefetch.ps1 chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
bench-run.sh refactor(moe): remove speculative gating to restore the modular seam 2026-07-14 10:41:27 +02:00
bench-warmonly.sh refactor(android): rename the dev shared model dir shardllm -> bmoe 2026-07-17 12:43:23 +02:00
build-android.ps1 build(android): drop i8mm so the CLI runs on pre-armv8.6 SoCs 2026-07-13 13:56:50 +02:00
build-host.sh build: add llama.cpp submodule and CMake skeleton 2026-07-10 18:17:31 +02:00
decode-analyze.py feat(trace): add layer-granularity compute trace (--compute-trace-layers) 2026-07-19 09:21:39 +02:00
gptoss-matrix.sh refactor(android): rename the dev shared model dir shardllm -> bmoe 2026-07-17 12:43:23 +02:00
gptoss-mmap.sh refactor(android): rename the dev shared model dir shardllm -> bmoe 2026-07-17 12:43:23 +02:00
make-tiny-moe.py test(moe): add a gemma4 fused-layout byte-identity gate 2026-07-11 17:01:14 +02:00
route-analyze.py chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
route-replay.py docs: measure the cache and I/O levers, and correct what the measurements contradict (#88) 2026-07-20 11:44:48 +02:00
route-viewer.py feat(scripts): route-viewer.py — read a route trace without a spreadsheet 2026-07-15 09:52:09 +02:00
trace_io.py chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00