BigMoeOnEdge/scripts
Helldez f6600993f0 feat(scripts): route-analyze.py + document the route trace format
The trace is a long-format CSV; this reads it and answers the questions that shape
streaming speed: the step x layer matrix itself (with a '*' on each expert id the same
layer also routed on the previous step), routing concentration per layer (what a warm-up
should preload), reuse distance per (layer, expert) (LRU vs pinning, and how big a cache
buys what), overlap with earlier steps (whether temporal prefetch can predict), the
cumulative unique-expert curve against the cache budget, hit rate and prefetch
usefulness, routing entropy, and bytes per step and layer.

Stdlib only, like bench-analyze.py — nothing to install.

docs/telemetry.md gains the format (v1), the column semantics, and the two asymmetries
that would otherwise be misread: residency is per routing while expert_bytes is per read
(prefill dedups), and the last layer legitimately has a single prefill step because
llama.cpp gathers only the output token before its FFN.
2026-07-15 08:44:28 +02:00
..
bench-analyze.py refactor(moe): remove speculative gating to restore the modular seam 2026-07-14 10:41:27 +02:00
bench-matrix-rework.ps1 refactor(moe): remove speculative gating to restore the modular seam 2026-07-14 10:41:27 +02:00
bench-matrix.ps1 docs: document the expert-ready fork hook and overlap telemetry 2026-07-12 08:54:19 +02:00
bench-pr23-c2000.ps1 bench(android): device A/B for prefetch and spec-gating (PR2/PR3) 2026-07-12 21:01:10 +02:00
bench-pr23-summary.py bench(android): device A/B for prefetch and spec-gating (PR2/PR3) 2026-07-12 21:01:10 +02:00
bench-prefetch.ps1 refactor(moe): remove speculative gating to restore the modular seam 2026-07-14 10:41:27 +02:00
bench-run.sh refactor(moe): remove speculative gating to restore the modular seam 2026-07-14 10:41:27 +02:00
bench-warmonly.sh docs(warmup): document the dense warm-up fix with on-device results 2026-07-14 18:44:15 +02:00
build-android.ps1 build(android): drop i8mm so the CLI runs on pre-armv8.6 SoCs 2026-07-13 13:56:50 +02:00
build-host.sh build: add llama.cpp submodule and CMake skeleton 2026-07-10 18:17:31 +02:00
gptoss-matrix.sh docs(benchmark): gpt-oss-120b on-device streaming results + drivers 2026-07-14 12:29:16 +02:00
gptoss-mmap.sh docs(benchmark): gpt-oss-120b on-device streaming results + drivers 2026-07-14 12:29:16 +02:00
make-tiny-moe.py test(moe): add a gemma4 fused-layout byte-identity gate 2026-07-11 17:01:14 +02:00
route-analyze.py feat(scripts): route-analyze.py + document the route trace format 2026-07-15 08:44:28 +02:00