mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 11:35:50 +00:00
The NPU prefill's expert arena read every expert of every layer ahead of its routing. It now reads, ahead of a layer's routing, the experts the previous graph routed there, and at the routing node whatever the routing adds. The matmul reads only routed experts, so the output is bit for bit the same. A layer routing more than --prefill-routed-full (0.85) of its experts gets the next one read whole; --no-prefill-routed restores whole layers everywhere. Phone, Hexagon v81 NPU, top-4, same session, every answer identical: Qwen3.6-35B-A3B Q4_0 7.68 -> 4.16 s, Q4_K_M 9.95 -> 5.37 s, Gemma 4 26B-A4B Q4_K_M 6.69 -> 3.70 s, Nemotron 3.5 30B-A3B Q4_0 7.42 -> 6.81 s. Also: --decide-probe (experimental per-decision expert usage and layer-exit answers), BMOE_DECIDE prefill_dev_* counters, gates G17f/G17g, app 0.28.0. |
||
|---|---|---|
| .. | ||
| 2026-07-12 | ||
| 2026-07-12-pr23 | ||
| 2026-07-13 | ||
| 2026-07-14 | ||
| 2026-07-14-warmup | ||
| 2026-07-15-route-trace | ||
| 2026-07-17 | ||
| 2026-07-20-cache-replay | ||
| 2026-07-20-sidecar | ||
| 2026-07-21-pinned-dense-ab | ||
| 2026-07-21-pinned-memory | ||
| 2026-07-22-drop-cold-experts | ||
| 2026-07-22-drop-quality | ||
| 2026-07-24-desktop-qwen36 | ||
| 2026-08-26-cache-aware-substitution | ||
| 2026-08-29-mmap-serialisation | ||
| 2026-09-29-prefill-routed | ||
| README.md | ||
Benchmark data archive
Raw per-run CSVs and the session notes written alongside them, one directory per measurement session. Each directory is a historical record of the build it was captured on, kept for provenance so the tables in ../benchmarks.md can be traced back to data.
These files are not maintained. They are not corrected when the code moves on, and a recommendation inside one is only valid for the build it was measured against. Read the maintained docs for current guidance; read these to check where a number came from.
| Session | Captured against | Note |
|---|---|---|
2026-07-12/ |
baseline cache/lane sweep | Feeds the cache-and-lanes tables in benchmarks.md. |
2026-07-12-pr23/ |
temporal prefetch + speculative gating | Speculative gating no longer exists — it was removed to restore the modular seam. The --spec-gate rows and the advice about it describe a flag the CLI no longer accepts. |
2026-07-13/ |
adaptive cache budget | Source of the capped-auto recipe numbers. |
2026-07-14/ |
per-token warm-up | Feeds warmup-analysis.md. |
2026-07-14-warmup/ |
dense warm-up A/B | Feeds cache-sizing.md and warmup-analysis.md. |
2026-07-15-route-trace/ |
first --route-trace capture (Qwen / Gemma / gpt-oss) |
Routing data, not throughput: every run had the trace on, so its tok/s are not comparable with benchmarks.md and must not feed those tables. |
2026-07-17/ |
all-O_DIRECT dense weights (--dense-weights anon) |
Source of the gpt-oss steady-state numbers. Several cells are contaminated by run order (a fixed cooldown does not reach baseline) — NOTES.md marks exactly which, and which claims survive. Do not read a top-k or lane claim out of it. |
2026-07-24-desktop-qwen36/ |
desktop (x86 laptop) streaming, Qwen3.6-35B at ~1.5× RAM | Source of the README Desktop table. Decode is DRAM-bandwidth-bound there, the opposite of the phone; round 1 carries a documented --cache-mb auto confound that round 2 (fixed cache) resolves. |
2026-08-26-cache-aware-substitution/ |
desktop streaming, Qwen3.6-35B, --expert-substitute at 0 / 0.15 / 0.30 / 0.60 |
Source of the tables in cache-aware-substitution.md. Perplexity is token-by-token (--ppl-step); the one-batch cells are kept as the record of why that mode exists, and must not be read as the policy's cost. |