mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 11:35:50 +00:00
Adds the two instruments the measurements were taken with, the data, and fixes to the maintained docs that turned out to assert things that do not hold. No engine code changes: the one candidate that was implemented is a measured regression and stays on its branch. tools/bmoe-iobench (new, BMOE_BUILD_TOOLS=OFF by default) sweeps flash read bandwidth against lane count and read size. It drives bmoe::FileReader -- the engine's own read path, so alignment, the bounce buffer and the O_DIRECT verify/fallback are part of what is measured -- but links nothing else, since the I/O layer has no llama.cpp dependency. It therefore cross-compiles in seconds and cannot perturb the streamer. --compute-load adds CPU contention, because the streamer reads while ggml's threads spin and an idle-CPU number is not the condition it operates under. scripts/route-replay.py (new, stdlib only) replays the committed route traces through hypothetical cache policies at zero device cost. It reproduces the recorded on-device hit rate to the decimal on all three captures and independently predicts a historical budget-shrink measurement it was not calibrated against. What they found, and what it invalidated: - roadmap.md opened its read-bandwidth theme on the premise that effective O_DIRECT bandwidth sits far below the drive's ceiling because routed slices are scattered. Measured, reads saturate at 2 lanes and are flat above 256 KiB, so scatter is cheap here. Runtime read coalescing is retired (0.6-4 % of a layer's routed experts are id-adjacent); the expert-contiguous repack survives but needs a different justification. - prefetch.md opened on "MoE routing has strong temporal locality". It is 17.9 % on gpt-oss, beaten there by a static hot list, and --prefetch 1 is a 2x slowdown. The mechanism and its correctness argument stand; the bet is annotated with what it returns and with the case still open (top-6 models). - cache-sizing.md said a too-small budget makes the hit rate "collapse". It makes it exactly 0.0 %, and the boundary is one token cycle -- not cache_min_mb, which is a different quantity roughly n_layer smaller and protects today's models only by coincidence. Reproduced on device at a budget the CLI accepts. bench-data/2026-07-20-cache-replay/ carries the three notes, the curves and every run CSV, indexed by a README stating the six verdicts. |
||
|---|---|---|
| .. | ||
| bench-analyze.py | ||
| bench-lib.ps1 | ||
| bench-matrix-rework.ps1 | ||
| bench-matrix.ps1 | ||
| bench-pr23-c2000.ps1 | ||
| bench-pr23-summary.py | ||
| bench-prefetch.ps1 | ||
| bench-run.sh | ||
| bench-warmonly.sh | ||
| build-android.ps1 | ||
| build-host.sh | ||
| decode-analyze.py | ||
| gptoss-matrix.sh | ||
| gptoss-mmap.sh | ||
| make-tiny-moe.py | ||
| route-analyze.py | ||
| route-replay.py | ||
| route-viewer.py | ||
| trace_io.py | ||