BigMoeOnEdge/docs/bench-data
Helldez fbdf755561 docs(bench): record the 2026-07-24 desktop campaign — the bottleneck flips
Qwen3.6-35B (22.3 GB, ~1.5x RAM) on a Windows x86 laptop, engine
unmodified: baseline 4.78 tok/s, best recipe --overlap --cache-mb auto
--drop-cold-experts 0.75 = 7.33 (+55%), above the phone's 5.0-5.8 on
the same model.

Streamed decode is DRAM-bandwidth-bound there, the opposite of the
phone: compute sits at ~0.11 s/token in every cell (a ~9 tok/s ceiling
at zero I/O), doubling compute threads buys 2 ms, and io8 = io4 to the
millisecond once the cache is fixed — round 1's apparent +17% from
lanes was entirely a --cache-mb auto budget confound (5.3-7.4 GiB
across cells), recorded with its resolution. Aggregate flash read
stays ~900 MiB/s on a ~3 GB/s NVMe regardless of lanes, and with
overlap on it drops to ~330 MiB/s: async reads compete with FFN
compute for the same DRAM bandwidth.

README gets a dedicated Desktop section (table + the flip) replacing
the old one-line quick check; raw CSVs and findings.md land under
docs/bench-data/2026-07-24-desktop-qwen36/.
2026-07-24 10:34:14 +02:00
..
2026-07-12 docs(bench): commit the 2026-07-12 device matrix and align the benchmark docs 2026-07-13 11:41:21 +02:00
2026-07-12-pr23 docs: correct content overtaken by the code 2026-07-15 07:50:47 +02:00
2026-07-13 bench(android): 2026-07-13 device matrix — adaptive cache and reworked spec-gating 2026-07-13 10:10:00 +02:00
2026-07-14 docs(benchmarks): per-token warm-up analysis for Qwen/Gemma and gpt-oss 2026-07-14 17:25:00 +02:00
2026-07-14-warmup docs: rename adaptive-cache.md to cache-sizing.md 2026-07-19 10:56:16 +02:00
2026-07-15-route-trace docs: record the expert-sidecar negative result and ship the --scatter microbench 2026-07-20 17:02:48 +02:00
2026-07-17 docs: gpt-oss leads the README; all-O_DIRECT dense weights measured 2026-07-17 11:31:35 +02:00
2026-07-20-cache-replay docs: point the layer-lfu record at its tag, not a branch that is being deleted 2026-07-20 11:47:27 +02:00
2026-07-20-sidecar docs: record the expert-sidecar negative result and ship the --scatter microbench 2026-07-20 17:02:48 +02:00
2026-07-21-pinned-dense-ab feat(dense): --dense-weights ahwb — dense weights in memory Android cannot reclaim (#93) 2026-07-21 11:09:51 +02:00
2026-07-21-pinned-memory test(memory): reclaim-exempt memory for the dense weights — bandwidth gate passes, 2047 MiB cap (#92) 2026-07-21 11:08:24 +02:00
2026-07-22-drop-cold-experts feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95) 2026-07-22 17:21:55 +02:00
2026-07-22-drop-quality test(moe): GSM8K quality check for --drop-cold-experts finds no loss (#97) 2026-07-22 20:04:01 +02:00
2026-07-24-desktop-qwen36 docs(bench): record the 2026-07-24 desktop campaign — the bottleneck flips 2026-07-24 10:34:14 +02:00
README.md docs(bench): record the 2026-07-24 desktop campaign — the bottleneck flips 2026-07-24 10:34:14 +02:00

Benchmark data archive

Raw per-run CSVs and the session notes written alongside them, one directory per measurement session. Each directory is a historical record of the build it was captured on, kept for provenance so the tables in ../benchmarks.md can be traced back to data.

These files are not maintained. They are not corrected when the code moves on, and a recommendation inside one is only valid for the build it was measured against. Read the maintained docs for current guidance; read these to check where a number came from.

Session Captured against Note
2026-07-12/ baseline cache/lane sweep Feeds the cache-and-lanes tables in benchmarks.md.
2026-07-12-pr23/ temporal prefetch + speculative gating Speculative gating no longer exists — it was removed to restore the modular seam. The --spec-gate rows and the advice about it describe a flag the CLI no longer accepts.
2026-07-13/ adaptive cache budget Source of the capped-auto recipe numbers.
2026-07-14/ per-token warm-up Feeds warmup-analysis.md.
2026-07-14-warmup/ dense warm-up A/B Feeds cache-sizing.md and warmup-analysis.md.
2026-07-15-route-trace/ first --route-trace capture (Qwen / Gemma / gpt-oss) Routing data, not throughput: every run had the trace on, so its tok/s are not comparable with benchmarks.md and must not feed those tables.
2026-07-17/ all-O_DIRECT dense weights (--dense-weights anon) Source of the gpt-oss steady-state numbers. Several cells are contaminated by run order (a fixed cooldown does not reach baseline) — NOTES.md marks exactly which, and which claims survive. Do not read a top-k or lane claim out of it.
2026-07-24-desktop-qwen36/ desktop (x86 laptop) streaming, Qwen3.6-35B at ~1.5× RAM Source of the README Desktop table. Decode is DRAM-bandwidth-bound there, the opposite of the phone; round 1 carries a documented --cache-mb auto confound that round 2 (fixed cache) resolves.