mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 11:35:50 +00:00
scripts/ had drifted into a pile of near-duplicate one-off drivers. Four PS1
files each carried a byte-identical Run-Cfg (cooldown, adb into bench-run.sh,
filter the perf lines, pull the artifacts); three Python tools each re-derived
the same readers, one of them saying so out loud ("Mirrors scripts/
route-analyze.py's reader").
bench-lib.ps1 now holds the shared driver plumbing, so a driver is its config
matrix and nothing else. trace_io.py holds the reading contract every artifact
shares: `key=value` tokens on `#` lines, rows below, unknown keys kept.
The device model paths were hardcoded in scripts that were otherwise
parameterised; they are now defaults behind -Qwen/-Gemma params (PS1) and
${VAR:-default} (sh), which is what they always were in spirit.
Retired framing, pruned from the tools that still run:
- bench-analyze.py listed `sg_ov` (speculative gating — removed from the
engine, PR #15) and labelled --cache-mb auto "adaptive cache" with a
`resizes` column. The governor that resized is gone, so that count is 0 for
every run bmoe-cli can produce, and a column that is structurally always 0
reads as a finding rather than a blank. Dropped; auto is "auto-sized".
- bench-pr23-c2000.ps1 grepped for `moe-spec-gate:`, a line the engine no
longer emits. Its --prefetch A/B is still valid, so it moves to the lib.
bench-matrix-rework.ps1 and bench-pr23-summary.py are NOT retrofitted: they only
re-derive tables already published under docs/bench-data, and their spec-gate
cells cannot run against a current build. They are marked ARCHIVED with why —
deleting them would leave published numbers with no visible derivation.
bench-lib.ps1 also documents what the copies silently carried: the cooldown is a
timer, and a timer does not guarantee a thermal baseline (see the contaminated
matrices where tok/s tracked run order). Fixing that needs device calibration;
saying so beats leaving the next reader to rediscover it.
Verified against committed data, not just by inspection:
- route-analyze --view hot/reuse/overlap/cache and decode-analyze: output
byte-identical before and after (docs/bench-data/2026-07-15-route-trace).
- bench-analyze on docs/bench-data/2026-07-13: every figure matches the
committed summary.md digit for digit; only labels changed and sg_ov is gone.
- All five PS1 files parse; the dot-sourced Invoke-BenchCfg builds the exact
same adb command string the copies did.
36 lines
1.4 KiB
PowerShell
36 lines
1.4 KiB
PowerShell
# On-device A/B for temporal prefetch, on top of each model's best measured config from
|
|
# docs/bench-data/2026-07-12 (Qwen: cache 4000, lane 4, overlap; Gemma: cache 2000, lane 4,
|
|
# overlap). One pair per model: base / +prefetch.
|
|
param(
|
|
[string]$OutDir = (Join-Path (Split-Path $PSScriptRoot -Parent) ".bench-pr23"),
|
|
[int]$NPred = 256,
|
|
[int]$CooldownSec = 45,
|
|
[string]$Qwen,
|
|
[string]$Gemma
|
|
)
|
|
$ErrorActionPreference = "Continue"
|
|
. "$PSScriptRoot\bench-lib.ps1"
|
|
New-Item -ItemType Directory -Force -Path $OutDir | Out-Null
|
|
|
|
if (-not $Qwen) { $Qwen = $BENCH_QWEN }
|
|
if (-not $Gemma) { $Gemma = $BENCH_GEMMA }
|
|
|
|
# This A/B is about overlap and prefetch, so watch their lines too.
|
|
$match = $BENCH_MATCH_DEFAULT + "|moe-overlap:|moe-prefetch:"
|
|
|
|
function Run-Cfg($tag, $model, $flags) {
|
|
Invoke-BenchCfg -Tag $tag -Model $model -Flags $flags -OutDir $OutDir -NPred $NPred `
|
|
-CooldownSec $CooldownSec -Match $match
|
|
}
|
|
|
|
# Qwen — best baseline: cache 4000 MiB, lane 4, overlap
|
|
$QB = "--moe-stream --cache-mb 4000 --io-threads 4 --overlap"
|
|
Run-Cfg "qwen_base" $Qwen $QB
|
|
Run-Cfg "qwen_pf2" $Qwen "$QB --prefetch 2"
|
|
|
|
# Gemma — best baseline: cache 2000 MiB, lane 4, overlap
|
|
$GB = "--moe-stream --cache-mb 2000 --io-threads 4 --overlap"
|
|
Run-Cfg "gemma_base" $Gemma $GB
|
|
Run-Cfg "gemma_pf2" $Gemma "$GB --prefetch 2"
|
|
|
|
Write-Host "ALL DONE -> $OutDir"
|