mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 11:35:50 +00:00
Qwen3.6-35B (22.3 GB, ~1.5x RAM) on a Windows x86 laptop, engine unmodified: baseline 4.78 tok/s, best recipe --overlap --cache-mb auto --drop-cold-experts 0.75 = 7.33 (+55%), above the phone's 5.0-5.8 on the same model. Streamed decode is DRAM-bandwidth-bound there, the opposite of the phone: compute sits at ~0.11 s/token in every cell (a ~9 tok/s ceiling at zero I/O), doubling compute threads buys 2 ms, and io8 = io4 to the millisecond once the cache is fixed — round 1's apparent +17% from lanes was entirely a --cache-mb auto budget confound (5.3-7.4 GiB across cells), recorded with its resolution. Aggregate flash read stays ~900 MiB/s on a ~3 GB/s NVMe regardless of lanes, and with overlap on it drops to ~330 MiB/s: async reads compete with FFN compute for the same DRAM bandwidth. README gets a dedicated Desktop section (table + the flip) replacing the old one-line quick check; raw CSVs and findings.md land under docs/bench-data/2026-07-24-desktop-qwen36/. |
||
|---|---|---|
| .. | ||
| 2026-07-12 | ||
| 2026-07-12-pr23 | ||
| 2026-07-13 | ||
| 2026-07-14 | ||
| 2026-07-14-warmup | ||
| 2026-07-15-route-trace | ||
| 2026-07-17 | ||
| 2026-07-20-cache-replay | ||
| 2026-07-20-sidecar | ||
| 2026-07-21-pinned-dense-ab | ||
| 2026-07-21-pinned-memory | ||
| 2026-07-22-drop-cold-experts | ||
| 2026-07-22-drop-quality | ||
| 2026-07-24-desktop-qwen36 | ||
| README.md | ||
Benchmark data archive
Raw per-run CSVs and the session notes written alongside them, one directory per measurement session. Each directory is a historical record of the build it was captured on, kept for provenance so the tables in ../benchmarks.md can be traced back to data.
These files are not maintained. They are not corrected when the code moves on, and a recommendation inside one is only valid for the build it was measured against. Read the maintained docs for current guidance; read these to check where a number came from.
| Session | Captured against | Note |
|---|---|---|
2026-07-12/ |
baseline cache/lane sweep | Feeds the cache-and-lanes tables in benchmarks.md. |
2026-07-12-pr23/ |
temporal prefetch + speculative gating | Speculative gating no longer exists — it was removed to restore the modular seam. The --spec-gate rows and the advice about it describe a flag the CLI no longer accepts. |
2026-07-13/ |
adaptive cache budget | Source of the capped-auto recipe numbers. |
2026-07-14/ |
per-token warm-up | Feeds warmup-analysis.md. |
2026-07-14-warmup/ |
dense warm-up A/B | Feeds cache-sizing.md and warmup-analysis.md. |
2026-07-15-route-trace/ |
first --route-trace capture (Qwen / Gemma / gpt-oss) |
Routing data, not throughput: every run had the trace on, so its tok/s are not comparable with benchmarks.md and must not feed those tables. |
2026-07-17/ |
all-O_DIRECT dense weights (--dense-weights anon) |
Source of the gpt-oss steady-state numbers. Several cells are contaminated by run order (a fixed cooldown does not reach baseline) — NOTES.md marks exactly which, and which claims survive. Do not read a top-k or lane claim out of it. |
2026-07-24-desktop-qwen36/ |
desktop (x86 laptop) streaming, Qwen3.6-35B at ~1.5× RAM | Source of the README Desktop table. Decode is DRAM-bandwidth-bound there, the opposite of the phone; round 1 carries a documented --cache-mb auto confound that round 2 (fixed cache) resolves. |