BigMoeOnEdge/scripts/bench-lib.ps1
Helldez b1f42cbdc7 chore(scripts): consolidate the bench drivers, prune the retired framing
scripts/ had drifted into a pile of near-duplicate one-off drivers. Four PS1
files each carried a byte-identical Run-Cfg (cooldown, adb into bench-run.sh,
filter the perf lines, pull the artifacts); three Python tools each re-derived
the same readers, one of them saying so out loud ("Mirrors scripts/
route-analyze.py's reader").

bench-lib.ps1 now holds the shared driver plumbing, so a driver is its config
matrix and nothing else. trace_io.py holds the reading contract every artifact
shares: `key=value` tokens on `#` lines, rows below, unknown keys kept.

The device model paths were hardcoded in scripts that were otherwise
parameterised; they are now defaults behind -Qwen/-Gemma params (PS1) and
${VAR:-default} (sh), which is what they always were in spirit.

Retired framing, pruned from the tools that still run:
  - bench-analyze.py listed `sg_ov` (speculative gating — removed from the
    engine, PR #15) and labelled --cache-mb auto "adaptive cache" with a
    `resizes` column. The governor that resized is gone, so that count is 0 for
    every run bmoe-cli can produce, and a column that is structurally always 0
    reads as a finding rather than a blank. Dropped; auto is "auto-sized".
  - bench-pr23-c2000.ps1 grepped for `moe-spec-gate:`, a line the engine no
    longer emits. Its --prefetch A/B is still valid, so it moves to the lib.

bench-matrix-rework.ps1 and bench-pr23-summary.py are NOT retrofitted: they only
re-derive tables already published under docs/bench-data, and their spec-gate
cells cannot run against a current build. They are marked ARCHIVED with why —
deleting them would leave published numbers with no visible derivation.

bench-lib.ps1 also documents what the copies silently carried: the cooldown is a
timer, and a timer does not guarantee a thermal baseline (see the contaminated
matrices where tok/s tracked run order). Fixing that needs device calibration;
saying so beats leaving the next reader to rediscover it.

Verified against committed data, not just by inspection:
  - route-analyze --view hot/reuse/overlap/cache and decode-analyze: output
    byte-identical before and after (docs/bench-data/2026-07-15-route-trace).
  - bench-analyze on docs/bench-data/2026-07-13: every figure matches the
    committed summary.md digit for digit; only labels changed and sg_ov is gone.
  - All five PS1 files parse; the dot-sourced Invoke-BenchCfg builds the exact
    same adb command string the copies did.
2026-07-17 10:41:13 +02:00

60 lines
2.9 KiB
PowerShell

# Shared plumbing for the on-device benchmark drivers. Dot-source it:
#
# . "$PSScriptRoot\bench-lib.ps1"
#
# Every driver does the same thing: cool the SoC, run one config through the device-side
# bench-run.sh, echo the lines worth watching, pull the CSV + metrics back. What differs between
# drivers is the config matrix and which perf lines matter — that IS the driver. This is everything
# else, and it used to be copied verbatim into each of them.
# Where bench-run.sh and the per-run artifacts live on the device.
$script:DEV_DIR = "/data/local/tmp"
# Model paths on the test device. Defaults, not constants: pass -Model whatever you are benching.
$script:BENCH_QWEN = "/sdcard/Download/Qwen3-30B-A3B-Q4_K_M.gguf"
$script:BENCH_GEMMA = "/sdcard/Download/google_gemma-4-26B-A4B-it-Q4_K_M.gguf"
# The perf lines worth echoing for any run. A driver exercising a specific feature appends its own
# (e.g. "|moe-prefetch:") via -Match rather than re-stating the whole pattern.
$script:BENCH_MATCH_DEFAULT =
"generation:|prefill:|moe-stream:|moe-cache:|peak_rss|mem_avail_floor|batt_temp_max|cpu_temp_max|charge_"
<#
.SYNOPSIS
Run one benchmark config on the device and pull its artifacts.
.DESCRIPTION
One cell of a matrix: cooldown, run, echo, pull. The tag names the run and its files.
KNOWN LIMITATION — the cooldown is a timer, and a timer is the wrong instrument. A fixed wait does
not reliably reach a thermal baseline, so a long matrix can drift and tok/s starts tracking run
order rather than the config under test. Read a suspicious matrix with that in mind: check the
temperature columns before believing a trend. Making this a condition (wait until the SoC is
actually cool) is the fix, and it needs device time to calibrate.
#>
function Invoke-BenchCfg {
param(
[Parameter(Mandatory)][string]$Tag,
[Parameter(Mandatory)][string]$Model,
[Parameter(Mandatory)][string]$OutDir,
[string]$Flags = "",
[int]$NPred = 256,
[int]$CooldownSec = 45,
[string]$Match = $script:BENCH_MATCH_DEFAULT
)
Write-Host "==================== $Tag ===================="
Start-Sleep -Seconds $CooldownSec # see the note above: this is a timer, not a guarantee
$t0 = Get-Date
$log = & adb shell "sh $script:DEV_DIR/bench-run.sh $NPred $Model $script:DEV_DIR/$Tag.csv $script:DEV_DIR/$Tag.metrics $Flags" 2>&1
$dt = [math]::Round(((Get-Date) - $t0).TotalSeconds, 1)
$log | Where-Object { $_ -match $Match } | ForEach-Object { Write-Host " $_" }
Write-Host " wall(incl load)=${dt}s"
$log | Out-File -FilePath "$OutDir\$Tag.log" -Encoding utf8
& adb pull "$script:DEV_DIR/$Tag.csv" "$OutDir\$Tag.csv" 2>&1 | Out-Null
& adb pull "$script:DEV_DIR/$Tag.metrics" "$OutDir\$Tag.metrics" 2>&1 | Out-Null
if (Test-Path "$OutDir\$Tag.csv") {
Write-Host " -> $Tag.csv + .metrics pulled"
} else {
Write-Host " !! no CSV for $Tag"
}
}