mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 19:45:46 +00:00
Closes the open item the throughput measurement left behind. 15 questions taken verbatim from the GSM8K test split, same model and config as the throughput A/B (Qwen3.6-35B-A3B, cache 3000, 4 lanes, overlap, ahwb, no-think), four cells differing only in the threshold: off 12/15 0 routings dropped 0.50 13/15 ~3% 0.75 13/15 ~14% 1.00 13/15 ~28% Twelve of the fifteen questions give an identical final answer in all four cells. All variation sits on two questions, and it flips in BOTH directions as the threshold rises rather than worsening with it -- the signature of a perturbation on problems already at the edge of the model's competence, not of accumulating damage. Reply length is flat and no cell produced a missing #### marker or an empty reply. Decoding is greedy, so there is no sampling noise: every difference between cells is caused by the policy. That makes the twelve identical answers a real statement rather than a coincidence. It does not make 13 against 12 an improvement -- perturbing a question the model already got wrong can land either side of the right answer, and at 6.7 points per question this sample cannot establish the sign of the effect. It rules out a collapse, not a subtle cost, and the write-up says so. The grading rule was fixed before any output was read (number after the last ####, falling back to the last number in the reply, applied identically to every cell) and answers.csv carries every reply in full so it can be audited. The prompt set, driver and grader are committed with the results. README gains a section stating how the three kinds of evidence are kept apart -- gates assert correctness, device runs measure speed, a public benchmark checks quality -- and where each comes from, including the GSM8K source and licence. |
||
|---|---|---|
| .. | ||
| 2026-07-12 | ||
| 2026-07-12-pr23 | ||
| 2026-07-13 | ||
| 2026-07-14 | ||
| 2026-07-14-warmup | ||
| 2026-07-15-route-trace | ||
| 2026-07-17 | ||
| 2026-07-20-cache-replay | ||
| 2026-07-20-sidecar | ||
| 2026-07-21-pinned-dense-ab | ||
| 2026-07-21-pinned-memory | ||
| 2026-07-22-drop-cold-experts | ||
| 2026-07-22-drop-quality | ||
| README.md | ||
Benchmark data archive
Raw per-run CSVs and the session notes written alongside them, one directory per measurement session. Each directory is a historical record of the build it was captured on, kept for provenance so the tables in ../benchmarks.md can be traced back to data.
These files are not maintained. They are not corrected when the code moves on, and a recommendation inside one is only valid for the build it was measured against. Read the maintained docs for current guidance; read these to check where a number came from.
| Session | Captured against | Note |
|---|---|---|
2026-07-12/ |
baseline cache/lane sweep | Feeds the cache-and-lanes tables in benchmarks.md. |
2026-07-12-pr23/ |
temporal prefetch + speculative gating | Speculative gating no longer exists — it was removed to restore the modular seam. The --spec-gate rows and the advice about it describe a flag the CLI no longer accepts. |
2026-07-13/ |
adaptive cache budget | Source of the capped-auto recipe numbers. |
2026-07-14/ |
per-token warm-up | Feeds warmup-analysis.md. |
2026-07-14-warmup/ |
dense warm-up A/B | Feeds cache-sizing.md and warmup-analysis.md. |
2026-07-15-route-trace/ |
first --route-trace capture (Qwen / Gemma / gpt-oss) |
Routing data, not throughput: every run had the trace on, so its tok/s are not comparable with benchmarks.md and must not feed those tables. |
2026-07-17/ |
all-O_DIRECT dense weights (--dense-weights anon) |
Source of the gpt-oss steady-state numbers. Several cells are contaminated by run order (a fixed cooldown does not reach baseline) — NOTES.md marks exactly which, and which claims survive. Do not read a top-k or lane claim out of it. |