mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 11:35:50 +00:00
The per-node compute trace pays ~3000 barriers per token, which serializes the graph against the expert stream: on a model that streams heavily the trace mostly measures its own serialization (Qwen3-30B: 9.4 s/token traced vs 0.39 untraced), so its absolutes cannot be compared across models. Layer granularity isolates only the first node of each layer (~n_layer barriers per token). Operator coalescing and the async expert prefetch survive, so the traced numbers stay close to an untraced run. Rows share the per-node schema with op LAYER: name blk.<il> aggregates one layer''s segment, pre the embedding lookup, post the last layer''s tail plus final norm and LM head (closed by the session right after llama_decode, since the tail has no successor boundary to observe it). The granularity flows RunConfig -> SessionConfig -> RouterHook; the routing nodes the streamer isolates anyway also close a segment, a barrier that exists untraced too. decode-analyze.py detects the granularity and prints the per-segment table. Gates: all 6 pass (byte-identity qwen3moe + gemma4). Smoke-tested on the tiny-moe models with streaming on. |
||
|---|---|---|
| .. | ||
| bench-analyze.py | ||
| bench-lib.ps1 | ||
| bench-matrix-rework.ps1 | ||
| bench-matrix.ps1 | ||
| bench-pr23-c2000.ps1 | ||
| bench-pr23-summary.py | ||
| bench-prefetch.ps1 | ||
| bench-run.sh | ||
| bench-warmonly.sh | ||
| build-android.ps1 | ||
| build-host.sh | ||
| decode-analyze.py | ||
| gptoss-matrix.sh | ||
| gptoss-mmap.sh | ||
| make-tiny-moe.py | ||
| route-analyze.py | ||
| route-viewer.py | ||
| trace_io.py | ||