BigMoeOnEdge/scripts/bench-prefetch.ps1
Helldez b1f42cbdc7 chore(scripts): consolidate the bench drivers, prune the retired framing
scripts/ had drifted into a pile of near-duplicate one-off drivers. Four PS1
files each carried a byte-identical Run-Cfg (cooldown, adb into bench-run.sh,
filter the perf lines, pull the artifacts); three Python tools each re-derived
the same readers, one of them saying so out loud ("Mirrors scripts/
route-analyze.py's reader").

bench-lib.ps1 now holds the shared driver plumbing, so a driver is its config
matrix and nothing else. trace_io.py holds the reading contract every artifact
shares: `key=value` tokens on `#` lines, rows below, unknown keys kept.

The device model paths were hardcoded in scripts that were otherwise
parameterised; they are now defaults behind -Qwen/-Gemma params (PS1) and
${VAR:-default} (sh), which is what they always were in spirit.

Retired framing, pruned from the tools that still run:
  - bench-analyze.py listed `sg_ov` (speculative gating — removed from the
    engine, PR #15) and labelled --cache-mb auto "adaptive cache" with a
    `resizes` column. The governor that resized is gone, so that count is 0 for
    every run bmoe-cli can produce, and a column that is structurally always 0
    reads as a finding rather than a blank. Dropped; auto is "auto-sized".
  - bench-pr23-c2000.ps1 grepped for `moe-spec-gate:`, a line the engine no
    longer emits. Its --prefetch A/B is still valid, so it moves to the lib.

bench-matrix-rework.ps1 and bench-pr23-summary.py are NOT retrofitted: they only
re-derive tables already published under docs/bench-data, and their spec-gate
cells cannot run against a current build. They are marked ARCHIVED with why —
deleting them would leave published numbers with no visible derivation.

bench-lib.ps1 also documents what the copies silently carried: the cooldown is a
timer, and a timer does not guarantee a thermal baseline (see the contaminated
matrices where tok/s tracked run order). Fixing that needs device calibration;
saying so beats leaving the next reader to rediscover it.

Verified against committed data, not just by inspection:
  - route-analyze --view hot/reuse/overlap/cache and decode-analyze: output
    byte-identical before and after (docs/bench-data/2026-07-15-route-trace).
  - bench-analyze on docs/bench-data/2026-07-13: every figure matches the
    committed summary.md digit for digit; only labels changed and sg_ov is gone.
  - All five PS1 files parse; the dot-sourced Invoke-BenchCfg builds the exact
    same adb command string the copies did.
2026-07-17 10:41:13 +02:00

36 lines
1.4 KiB
PowerShell

# On-device A/B for temporal prefetch, on top of each model's best measured config from
# docs/bench-data/2026-07-12 (Qwen: cache 4000, lane 4, overlap; Gemma: cache 2000, lane 4,
# overlap). One pair per model: base / +prefetch.
param(
[string]$OutDir = (Join-Path (Split-Path $PSScriptRoot -Parent) ".bench-pr23"),
[int]$NPred = 256,
[int]$CooldownSec = 45,
[string]$Qwen,
[string]$Gemma
)
$ErrorActionPreference = "Continue"
. "$PSScriptRoot\bench-lib.ps1"
New-Item -ItemType Directory -Force -Path $OutDir | Out-Null
if (-not $Qwen) { $Qwen = $BENCH_QWEN }
if (-not $Gemma) { $Gemma = $BENCH_GEMMA }
# This A/B is about overlap and prefetch, so watch their lines too.
$match = $BENCH_MATCH_DEFAULT + "|moe-overlap:|moe-prefetch:"
function Run-Cfg($tag, $model, $flags) {
Invoke-BenchCfg -Tag $tag -Model $model -Flags $flags -OutDir $OutDir -NPred $NPred `
-CooldownSec $CooldownSec -Match $match
}
# Qwen — best baseline: cache 4000 MiB, lane 4, overlap
$QB = "--moe-stream --cache-mb 4000 --io-threads 4 --overlap"
Run-Cfg "qwen_base" $Qwen $QB
Run-Cfg "qwen_pf2" $Qwen "$QB --prefetch 2"
# Gemma — best baseline: cache 2000 MiB, lane 4, overlap
$GB = "--moe-stream --cache-mb 2000 --io-threads 4 --overlap"
Run-Cfg "gemma_base" $Gemma $GB
Run-Cfg "gemma_pf2" $Gemma "$GB --prefetch 2"
Write-Host "ALL DONE -> $OutDir"