BigMoeOnEdge/scripts/trace_io.py
Helldez b1f42cbdc7 chore(scripts): consolidate the bench drivers, prune the retired framing
scripts/ had drifted into a pile of near-duplicate one-off drivers. Four PS1
files each carried a byte-identical Run-Cfg (cooldown, adb into bench-run.sh,
filter the perf lines, pull the artifacts); three Python tools each re-derived
the same readers, one of them saying so out loud ("Mirrors scripts/
route-analyze.py's reader").

bench-lib.ps1 now holds the shared driver plumbing, so a driver is its config
matrix and nothing else. trace_io.py holds the reading contract every artifact
shares: `key=value` tokens on `#` lines, rows below, unknown keys kept.

The device model paths were hardcoded in scripts that were otherwise
parameterised; they are now defaults behind -Qwen/-Gemma params (PS1) and
${VAR:-default} (sh), which is what they always were in spirit.

Retired framing, pruned from the tools that still run:
  - bench-analyze.py listed `sg_ov` (speculative gating — removed from the
    engine, PR #15) and labelled --cache-mb auto "adaptive cache" with a
    `resizes` column. The governor that resized is gone, so that count is 0 for
    every run bmoe-cli can produce, and a column that is structurally always 0
    reads as a finding rather than a blank. Dropped; auto is "auto-sized".
  - bench-pr23-c2000.ps1 grepped for `moe-spec-gate:`, a line the engine no
    longer emits. Its --prefetch A/B is still valid, so it moves to the lib.

bench-matrix-rework.ps1 and bench-pr23-summary.py are NOT retrofitted: they only
re-derive tables already published under docs/bench-data, and their spec-gate
cells cannot run against a current build. They are marked ARCHIVED with why —
deleting them would leave published numbers with no visible derivation.

bench-lib.ps1 also documents what the copies silently carried: the cooldown is a
timer, and a timer does not guarantee a thermal baseline (see the contaminated
matrices where tok/s tracked run order). Fixing that needs device calibration;
saying so beats leaving the next reader to rediscover it.

Verified against committed data, not just by inspection:
  - route-analyze --view hot/reuse/overlap/cache and decode-analyze: output
    byte-identical before and after (docs/bench-data/2026-07-15-route-trace).
  - bench-analyze on docs/bench-data/2026-07-13: every figure matches the
    committed summary.md digit for digit; only labels changed and sg_ov is gone.
  - All five PS1 files parse; the dot-sourced Invoke-BenchCfg builds the exact
    same adb command string the copies did.
2026-07-17 10:41:13 +02:00

53 lines
2 KiB
Python

#!/usr/bin/env python3
"""Shared reading for the analysis scripts. Stdlib only, like everything else here.
Every artifact the engine writes — the per-token CSV, the route trace, the decode/IO traces, the
device .metrics file — carries its context as `key=value` tokens on `#` lines, with the rows below.
That contract is one thing, and it was being re-implemented per script (decode-analyze.py's reader
said as much: "Mirrors scripts/route-analyze.py's reader"). It lives here now.
The contract is deliberately forgiving in one direction: unknown keys are kept and missing keys are
absent rather than an error, so a file from an older or newer engine still reads.
"""
import csv
import os
def kv_tokens(line):
"""`# a=1 b=2` -> {"a": "1", "b": "2"}. Non-`key=value` tokens are skipped."""
return dict(tok.split("=", 1) for tok in line.lstrip("#").split() if "=" in tok)
def read_preamble_csv(path):
"""Split a `#`-preamble CSV into (meta, rows): meta merges every `#` line's key=value tokens,
rows are dicts keyed by the header. Use when the rows are wanted by column NAME."""
meta, body = {}, []
with open(path, newline="", encoding="utf-8") as f:
for ln in f:
if ln.startswith("#"):
meta.update(kv_tokens(ln.rstrip()))
else:
body.append(ln)
return meta, list(csv.DictReader(body))
def read_kv_file(path):
"""A flat `key=value` per line file (the device .metrics sidecar). {} when absent."""
d = {}
if os.path.exists(path):
with open(path) as f:
for line in f:
if "=" in line:
k, v = line.strip().split("=", 1)
d[k] = v
return d
def percentile(values, q):
"""Linear-interpolated percentile of an ALREADY SORTED list; 0.0 when empty."""
if not values:
return 0.0
i = q * (len(values) - 1)
lo = int(i)
hi = min(lo + 1, len(values) - 1)
return values[lo] + (values[hi] - values[lo]) * (i - lo)