mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 11:35:50 +00:00
The per-token CSV reports compute as a residual (wall - io - mgmt), so everything the engine does not itself clock is pooled into it: page faults, scheduler stalls, and the matmuls. A residual cannot say which. That is the whole reason gpt-oss-120b reads as "compute-bound" at 1.7 s/token while faulting 8.4k pages per token — the flash wait is billed to compute because it happens under llama_decode. --compute-trace measures it instead. Asking the eval callback to isolate a node makes ggml compute exactly up to it and synchronize, so the wall delta between consecutive boundaries is that node's real compute time; sampling major faults across the same boundaries attributes the >RAM stall to the node that paid it. Still no llama.cpp patch: this rides the public cb_eval ABI, whose ask/no-ask contract already specifies the isolation. It costs a barrier per node and forbids operator coalescing, so it is a diagnostic — a traced run is not a benchmark run, and only the shares are meaningful. Unlike the other traces it does not need --moe-stream: it times the graph, so a dense mmap baseline can be traced and compared against a streamed run. --io-trace records one row per pread: latency, requested vs aligned bytes, lane, and the (layer, expert, projection) it serves — values already computed at every enqueue site and until now discarded. This is where the flash floor is: the aggregate 760 MiB/s sits far below the drive's sequential ceiling because routed slices are scattered, and the trace says whether that is per-read latency, request size, or lanes idling. It also measures the adjacency the roadmap's read-coalescing and expert-contiguous-layout items assume. Node classification stays out of the engine: which node is attention vs dense FFN vs expert matmul is naming policy that varies by architecture, so the rows carry the raw op and name and scripts/decode-analyze.py classifies. Verified on both gate models that the generated text, cache hit rate and bytes read are identical with the traces on. |
||
|---|---|---|
| .. | ||
| CMakeLists.txt | ||
| main.cpp | ||