BigMoeOnEdge/scripts
Helldez 3db882d278 feat(trace): add layer-granularity compute trace (--compute-trace-layers)
The per-node compute trace pays ~3000 barriers per token, which serializes
the graph against the expert stream: on a model that streams heavily the
trace mostly measures its own serialization (Qwen3-30B: 9.4 s/token traced
vs 0.39 untraced), so its absolutes cannot be compared across models.

Layer granularity isolates only the first node of each layer (~n_layer
barriers per token). Operator coalescing and the async expert prefetch
survive, so the traced numbers stay close to an untraced run. Rows share
the per-node schema with op LAYER: name blk.<il> aggregates one layer''s
segment, pre the embedding lookup, post the last layer''s tail plus
final norm and LM head (closed by the session right after llama_decode,
since the tail has no successor boundary to observe it).

The granularity flows RunConfig -> SessionConfig -> RouterHook; the routing
nodes the streamer isolates anyway also close a segment, a barrier that
exists untraced too. decode-analyze.py detects the granularity and prints
the per-segment table.

Gates: all 6 pass (byte-identity qwen3moe + gemma4). Smoke-tested on the
tiny-moe models with streaming on.
2026-07-19 09:21:39 +02:00
..
bench-analyze.py chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
bench-lib.ps1 chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
bench-matrix-rework.ps1 chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
bench-matrix.ps1 chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
bench-pr23-c2000.ps1 chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
bench-pr23-summary.py chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
bench-prefetch.ps1 chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
bench-run.sh refactor(moe): remove speculative gating to restore the modular seam 2026-07-14 10:41:27 +02:00
bench-warmonly.sh refactor(android): rename the dev shared model dir shardllm -> bmoe 2026-07-17 12:43:23 +02:00
build-android.ps1 build(android): drop i8mm so the CLI runs on pre-armv8.6 SoCs 2026-07-13 13:56:50 +02:00
build-host.sh build: add llama.cpp submodule and CMake skeleton 2026-07-10 18:17:31 +02:00
decode-analyze.py feat(trace): add layer-granularity compute trace (--compute-trace-layers) 2026-07-19 09:21:39 +02:00
gptoss-matrix.sh refactor(android): rename the dev shared model dir shardllm -> bmoe 2026-07-17 12:43:23 +02:00
gptoss-mmap.sh refactor(android): rename the dev shared model dir shardllm -> bmoe 2026-07-17 12:43:23 +02:00
make-tiny-moe.py test(moe): add a gemma4 fused-layout byte-identity gate 2026-07-11 17:01:14 +02:00
route-analyze.py chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00
route-viewer.py feat(scripts): route-viewer.py — read a route trace without a spreadsheet 2026-07-15 09:52:09 +02:00
trace_io.py chore(scripts): consolidate the bench drivers, prune the retired framing 2026-07-17 10:41:13 +02:00