BigMoeOnEdge/docs/bench-data/README.md
Raffaele 4334c89616
Some checks are pending
ci / changes (push) Waiting to run
ci / format (push) Blocked by required conditions
ci / host-linux (push) Blocked by required conditions
ci / host-windows (push) Waiting to run
ci / android-apk (push) Waiting to run
feat(moe): cache-aware expert substitution (--expert-substitute, experimental) and --ppl (#171)
Experimental, off by default. Before a decode routing is committed, every
expert already in the LRU cache gets its score raised by L times the
token's score range and the top-k is taken again, so a near-tie goes to
the expert already in RAM (Skliar et al., arXiv:2412.00099). The same
number of experts runs; fewer are read from flash. Scores are read from
the tensor the graph itself sorted, exact for any gating function.

Desktop, Qwen3.6-35B Q4_K_M at L=0.15: 258 to 119 MiB of flash per
token, 2.37 to 3.84 tok/s, perplexity +1 to 4 %, tinyMMLU 88 to 84/100,
HumanEval-50 42 = 42. The on-device A/B is still owed, hence experimental.

Also: --ppl / --ppl-step / --ppl-list / --ppl-choices (teacher-forced
perplexity, one token per decode so cache-dependent policies are priced
where they act), scripts/tinymmlu-bench.py, scripts/humaneval-bench.py,
gates G8d/G8e, app switch "Prefer cached experts" under Experimental,
docs/cache-aware-substitution.md.
2026-08-29 10:13:47 +02:00

2.3 KiB
Raw Permalink Blame History

Benchmark data archive

Raw per-run CSVs and the session notes written alongside them, one directory per measurement session. Each directory is a historical record of the build it was captured on, kept for provenance so the tables in ../benchmarks.md can be traced back to data.

These files are not maintained. They are not corrected when the code moves on, and a recommendation inside one is only valid for the build it was measured against. Read the maintained docs for current guidance; read these to check where a number came from.

Session Captured against Note
2026-07-12/ baseline cache/lane sweep Feeds the cache-and-lanes tables in benchmarks.md.
2026-07-12-pr23/ temporal prefetch + speculative gating Speculative gating no longer exists — it was removed to restore the modular seam. The --spec-gate rows and the advice about it describe a flag the CLI no longer accepts.
2026-07-13/ adaptive cache budget Source of the capped-auto recipe numbers.
2026-07-14/ per-token warm-up Feeds warmup-analysis.md.
2026-07-14-warmup/ dense warm-up A/B Feeds cache-sizing.md and warmup-analysis.md.
2026-07-15-route-trace/ first --route-trace capture (Qwen / Gemma / gpt-oss) Routing data, not throughput: every run had the trace on, so its tok/s are not comparable with benchmarks.md and must not feed those tables.
2026-07-17/ all-O_DIRECT dense weights (--dense-weights anon) Source of the gpt-oss steady-state numbers. Several cells are contaminated by run order (a fixed cooldown does not reach baseline) — NOTES.md marks exactly which, and which claims survive. Do not read a top-k or lane claim out of it.
2026-07-24-desktop-qwen36/ desktop (x86 laptop) streaming, Qwen3.6-35B at ~1.5× RAM Source of the README Desktop table. Decode is DRAM-bandwidth-bound there, the opposite of the phone; round 1 carries a documented --cache-mb auto confound that round 2 (fixed cache) resolves.
2026-08-26-cache-aware-substitution/ desktop streaming, Qwen3.6-35B, --expert-substitute at 0 / 0.15 / 0.30 / 0.60 Source of the tables in cache-aware-substitution.md. Perplexity is token-by-token (--ppl-step); the one-batch cells are kept as the record of why that mode exists, and must not be read as the policy's cost.