mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
The trace is a long-format CSV; this reads it and answers the questions that shape streaming speed: the step x layer matrix itself (with a '*' on each expert id the same layer also routed on the previous step), routing concentration per layer (what a warm-up should preload), reuse distance per (layer, expert) (LRU vs pinning, and how big a cache buys what), overlap with earlier steps (whether temporal prefetch can predict), the cumulative unique-expert curve against the cache budget, hit rate and prefetch usefulness, routing entropy, and bytes per step and layer. Stdlib only, like bench-analyze.py — nothing to install. docs/telemetry.md gains the format (v1), the column semantics, and the two asymmetries that would otherwise be misread: residency is per routing while expert_bytes is per read (prefill dedups), and the last layer legitimately has a single prefill step because llama.cpp gathers only the output token before its FFN. |
||
|---|---|---|
| .. | ||
| bench-analyze.py | ||
| bench-matrix-rework.ps1 | ||
| bench-matrix.ps1 | ||
| bench-pr23-c2000.ps1 | ||
| bench-pr23-summary.py | ||
| bench-prefetch.ps1 | ||
| bench-run.sh | ||
| bench-warmonly.sh | ||
| build-android.ps1 | ||
| build-host.sh | ||
| gptoss-matrix.sh | ||
| gptoss-mmap.sh | ||
| make-tiny-moe.py | ||
| route-analyze.py | ||