BigMoeOnEdge/examples
Helldez c9c5d37367
feat(metrics): record the whole run configuration, and what tok/s omits (#116)
Two gaps, both about a benchmark file being unable to explain itself.

The CSV preamble had eighteen keys and had fallen behind several releases.
Missing: n_ubatch, which sets the compute-buffer reservation and therefore
moves the very memory columns printed underneath it; predict_log,
predict_spec_max and prefetch_sync, so a probed run was indistinguishable from
a benchmark run the docs explicitly say it is not; drop_renorm and
drop_prefill, which change how much mass dropping discards; cache_floor_mb, the
input behind an auto-sized cache_mb; load_all, whose read_bytes mean something
else entirely; every sampling parameter, so a stochastic run read as a greedy
one; and compute_trace_layers. All are recorded now, under a bumped
"# bmoe_metrics v2" banner.

`think` stays out on purpose: it is a property of a request, not of a session,
so one value in a session-wide preamble would be wrong for every turn that
asked for the other.

The file also now names the build that wrote it. There was no version string
anywhere in the engine — the CMake project had none — so a committed CSV could
only be dated by the commit that copied it in. Added as a project VERSION, a
BMOE_VERSION define, bmoe::version(), `bmoe-cli --version`, and engine= in the
preamble. It sits on its own line so the model= line still BEGINS with model=,
which is how the app's CSV reader finds a run's name; the app's lookup is made
order-independent too, so the next key to be appended cannot break it again.

Second gap: wall_ms brackets llama_decode and nothing else. That is what makes
compute_ms a clean residual, and it also means sampling, detokenization,
rendering and the sink writes are outside wall_ms, outside gen_seconds and
outside the reported tok/s. Work moved into or out of that region was
unmeasurable by the number the project optimizes. loop_overhead_ms now reports
it per token, and loop_overhead_s/tok in the summary closes the accounting with
the tail after the last token that no row can carry.

docs/telemetry.md documents the preamble — it specified every other # block but
not this one — the new column, and the stall_ms divisor: stall is a per-thread
mean, so a stall that is not simultaneous across threads is under-stated and
compute_ms absorbs the difference.

Gates 7/7; preamble, column and --version verified against the tiny gate model.
Both readers checked: scripts/bench-analyze.py reads columns by name and skips
unknown # lines; the app's Csv.read keeps unknown keys and picks the new column
up automatically.
2026-07-28 11:49:34 +02:00
..
android feat(metrics): record the whole run configuration, and what tok/s omits (#116) 2026-07-28 11:49:34 +02:00