BigMoeOnEdge/docs/bench-data
Raffaele 3170385fad
feat(prefill): the NPU prefill reads only routed experts (+ --decide-probe); 0.28.0 (#208)
The NPU prefill's expert arena read every expert of every layer ahead of its
routing. It now reads, ahead of a layer's routing, the experts the previous
graph routed there, and at the routing node whatever the routing adds. The
matmul reads only routed experts, so the output is bit for bit the same. A layer
routing more than --prefill-routed-full (0.85) of its experts gets the next one
read whole; --no-prefill-routed restores whole layers everywhere.

Phone, Hexagon v81 NPU, top-4, same session, every answer identical:
Qwen3.6-35B-A3B Q4_0 7.68 -> 4.16 s, Q4_K_M 9.95 -> 5.37 s, Gemma 4 26B-A4B
Q4_K_M 6.69 -> 3.70 s, Nemotron 3.5 30B-A3B Q4_0 7.42 -> 6.81 s.

Also: --decide-probe (experimental per-decision expert usage and layer-exit
answers), BMOE_DECIDE prefill_dev_* counters, gates G17f/G17g, app 0.28.0.
2026-09-29 16:08:23 +02:00
..
2026-07-12 docs(bench): commit the 2026-07-12 device matrix and align the benchmark docs 2026-07-13 11:41:21 +02:00
2026-07-12-pr23 docs: correct content overtaken by the code 2026-07-15 07:50:47 +02:00
2026-07-13 bench(android): 2026-07-13 device matrix — adaptive cache and reworked spec-gating 2026-07-13 10:10:00 +02:00
2026-07-14 docs(benchmarks): per-token warm-up analysis for Qwen/Gemma and gpt-oss 2026-07-14 17:25:00 +02:00
2026-07-14-warmup docs: rename adaptive-cache.md to cache-sizing.md 2026-07-19 10:56:16 +02:00
2026-07-15-route-trace docs: record the expert-sidecar negative result and ship the --scatter microbench 2026-07-20 17:02:48 +02:00
2026-07-17 docs: gpt-oss leads the README; all-O_DIRECT dense weights measured 2026-07-17 11:31:35 +02:00
2026-07-20-cache-replay docs: point the layer-lfu record at its tag, not a branch that is being deleted 2026-07-20 11:47:27 +02:00
2026-07-20-sidecar docs: record the expert-sidecar negative result and ship the --scatter microbench 2026-07-20 17:02:48 +02:00
2026-07-21-pinned-dense-ab feat(dense): --dense-weights ahwb — dense weights in memory Android cannot reclaim (#93) 2026-07-21 11:09:51 +02:00
2026-07-21-pinned-memory test(memory): reclaim-exempt memory for the dense weights — bandwidth gate passes, 2047 MiB cap (#92) 2026-07-21 11:08:24 +02:00
2026-07-22-drop-cold-experts feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95) 2026-07-22 17:21:55 +02:00
2026-07-22-drop-quality test(moe): GSM8K quality check for --drop-cold-experts finds no loss (#97) 2026-07-22 20:04:01 +02:00
2026-07-24-desktop-qwen36 docs(bench): record the 2026-07-24 desktop campaign — the bottleneck flips 2026-07-24 10:34:14 +02:00
2026-08-26-cache-aware-substitution feat(moe): cache-aware expert substitution (--expert-substitute, experimental) and --ppl (#171) 2026-08-29 10:13:47 +02:00
2026-08-29-mmap-serialisation feat(io): release the model file's mapping after load (--release-mmap) (#185) 2026-09-07 20:43:02 +02:00
2026-09-29-prefill-routed feat(prefill): the NPU prefill reads only routed experts (+ --decide-probe); 0.28.0 (#208) 2026-09-29 16:08:23 +02:00
README.md feat(moe): cache-aware expert substitution (--expert-substitute, experimental) and --ppl (#171) 2026-08-29 10:13:47 +02:00

Benchmark data archive

Raw per-run CSVs and the session notes written alongside them, one directory per measurement session. Each directory is a historical record of the build it was captured on, kept for provenance so the tables in ../benchmarks.md can be traced back to data.

These files are not maintained. They are not corrected when the code moves on, and a recommendation inside one is only valid for the build it was measured against. Read the maintained docs for current guidance; read these to check where a number came from.

Session Captured against Note
2026-07-12/ baseline cache/lane sweep Feeds the cache-and-lanes tables in benchmarks.md.
2026-07-12-pr23/ temporal prefetch + speculative gating Speculative gating no longer exists — it was removed to restore the modular seam. The --spec-gate rows and the advice about it describe a flag the CLI no longer accepts.
2026-07-13/ adaptive cache budget Source of the capped-auto recipe numbers.
2026-07-14/ per-token warm-up Feeds warmup-analysis.md.
2026-07-14-warmup/ dense warm-up A/B Feeds cache-sizing.md and warmup-analysis.md.
2026-07-15-route-trace/ first --route-trace capture (Qwen / Gemma / gpt-oss) Routing data, not throughput: every run had the trace on, so its tok/s are not comparable with benchmarks.md and must not feed those tables.
2026-07-17/ all-O_DIRECT dense weights (--dense-weights anon) Source of the gpt-oss steady-state numbers. Several cells are contaminated by run order (a fixed cooldown does not reach baseline) — NOTES.md marks exactly which, and which claims survive. Do not read a top-k or lane claim out of it.
2026-07-24-desktop-qwen36/ desktop (x86 laptop) streaming, Qwen3.6-35B at ~1.5× RAM Source of the README Desktop table. Decode is DRAM-bandwidth-bound there, the opposite of the phone; round 1 carries a documented --cache-mb auto confound that round 2 (fixed cache) resolves.
2026-08-26-cache-aware-substitution/ desktop streaming, Qwen3.6-35B, --expert-substitute at 0 / 0.15 / 0.30 / 0.60 Source of the tables in cache-aware-substitution.md. Perplexity is token-by-token (--ppl-step); the one-batch cells are kept as the record of why that mode exists, and must not be read as the policy's cost.