BigMoeOnEdge/docs/bench-data/README.md
Helldez fbdf755561 docs(bench): record the 2026-07-24 desktop campaign — the bottleneck flips
Qwen3.6-35B (22.3 GB, ~1.5x RAM) on a Windows x86 laptop, engine
unmodified: baseline 4.78 tok/s, best recipe --overlap --cache-mb auto
--drop-cold-experts 0.75 = 7.33 (+55%), above the phone's 5.0-5.8 on
the same model.

Streamed decode is DRAM-bandwidth-bound there, the opposite of the
phone: compute sits at ~0.11 s/token in every cell (a ~9 tok/s ceiling
at zero I/O), doubling compute threads buys 2 ms, and io8 = io4 to the
millisecond once the cache is fixed — round 1's apparent +17% from
lanes was entirely a --cache-mb auto budget confound (5.3-7.4 GiB
across cells), recorded with its resolution. Aggregate flash read
stays ~900 MiB/s on a ~3 GB/s NVMe regardless of lanes, and with
overlap on it drops to ~330 MiB/s: async reads compete with FFN
compute for the same DRAM bandwidth.

README gets a dedicated Desktop section (table + the flip) replacing
the old one-line quick check; raw CSVs and findings.md land under
docs/bench-data/2026-07-24-desktop-qwen36/.
2026-07-24 10:34:14 +02:00

2 KiB
Raw Blame History

Benchmark data archive

Raw per-run CSVs and the session notes written alongside them, one directory per measurement session. Each directory is a historical record of the build it was captured on, kept for provenance so the tables in ../benchmarks.md can be traced back to data.

These files are not maintained. They are not corrected when the code moves on, and a recommendation inside one is only valid for the build it was measured against. Read the maintained docs for current guidance; read these to check where a number came from.

Session Captured against Note
2026-07-12/ baseline cache/lane sweep Feeds the cache-and-lanes tables in benchmarks.md.
2026-07-12-pr23/ temporal prefetch + speculative gating Speculative gating no longer exists — it was removed to restore the modular seam. The --spec-gate rows and the advice about it describe a flag the CLI no longer accepts.
2026-07-13/ adaptive cache budget Source of the capped-auto recipe numbers.
2026-07-14/ per-token warm-up Feeds warmup-analysis.md.
2026-07-14-warmup/ dense warm-up A/B Feeds cache-sizing.md and warmup-analysis.md.
2026-07-15-route-trace/ first --route-trace capture (Qwen / Gemma / gpt-oss) Routing data, not throughput: every run had the trace on, so its tok/s are not comparable with benchmarks.md and must not feed those tables.
2026-07-17/ all-O_DIRECT dense weights (--dense-weights anon) Source of the gpt-oss steady-state numbers. Several cells are contaminated by run order (a fixed cooldown does not reach baseline) — NOTES.md marks exactly which, and which claims survive. Do not read a top-k or lane claim out of it.
2026-07-24-desktop-qwen36/ desktop (x86 laptop) streaming, Qwen3.6-35B at ~1.5× RAM Source of the README Desktop table. Decode is DRAM-bandwidth-bound there, the opposite of the phone; round 1 carries a documented --cache-mb auto confound that round 2 (fixed cache) resolves.