BigMoeOnEdge/cli
Helldez f9e408f542 feat(metrics): --compute-trace and --io-trace decompose the decode
The per-token CSV reports compute as a residual (wall - io - mgmt), so everything the
engine does not itself clock is pooled into it: page faults, scheduler stalls, and the
matmuls. A residual cannot say which. That is the whole reason gpt-oss-120b reads as
"compute-bound" at 1.7 s/token while faulting 8.4k pages per token — the flash wait is
billed to compute because it happens under llama_decode.

--compute-trace measures it instead. Asking the eval callback to isolate a node makes
ggml compute exactly up to it and synchronize, so the wall delta between consecutive
boundaries is that node's real compute time; sampling major faults across the same
boundaries attributes the >RAM stall to the node that paid it. Still no llama.cpp patch:
this rides the public cb_eval ABI, whose ask/no-ask contract already specifies the
isolation. It costs a barrier per node and forbids operator coalescing, so it is a
diagnostic — a traced run is not a benchmark run, and only the shares are meaningful.
Unlike the other traces it does not need --moe-stream: it times the graph, so a dense
mmap baseline can be traced and compared against a streamed run.

--io-trace records one row per pread: latency, requested vs aligned bytes, lane, and the
(layer, expert, projection) it serves — values already computed at every enqueue site and
until now discarded. This is where the flash floor is: the aggregate 760 MiB/s sits far
below the drive's sequential ceiling because routed slices are scattered, and the trace
says whether that is per-read latency, request size, or lanes idling. It also measures
the adjacency the roadmap's read-coalescing and expert-contiguous-layout items assume.

Node classification stays out of the engine: which node is attention vs dense FFN vs
expert matmul is naming policy that varies by architecture, so the rows carry the raw op
and name and scripts/decode-analyze.py classifies. Verified on both gate models that the
generated text, cache hit rate and bytes read are identical with the traces on.
2026-07-15 20:38:56 +02:00
..
CMakeLists.txt build: add llama.cpp submodule and CMake skeleton 2026-07-10 18:17:31 +02:00
main.cpp feat(metrics): --compute-trace and --io-trace decompose the decode 2026-07-15 20:38:56 +02:00