BigMoeOnEdge/docs/bench-data
Helldez b40fa5ba73
test(moe): GSM8K quality check for --drop-cold-experts finds no loss (#97)
Closes the open item the throughput measurement left behind. 15 questions
taken verbatim from the GSM8K test split, same model and config as the
throughput A/B (Qwen3.6-35B-A3B, cache 3000, 4 lanes, overlap, ahwb,
no-think), four cells differing only in the threshold:

  off   12/15   0 routings dropped
  0.50  13/15   ~3%
  0.75  13/15   ~14%
  1.00  13/15   ~28%

Twelve of the fifteen questions give an identical final answer in all four
cells. All variation sits on two questions, and it flips in BOTH
directions as the threshold rises rather than worsening with it -- the
signature of a perturbation on problems already at the edge of the model's
competence, not of accumulating damage. Reply length is flat and no cell
produced a missing #### marker or an empty reply.

Decoding is greedy, so there is no sampling noise: every difference
between cells is caused by the policy. That makes the twelve identical
answers a real statement rather than a coincidence. It does not make 13
against 12 an improvement -- perturbing a question the model already got
wrong can land either side of the right answer, and at 6.7 points per
question this sample cannot establish the sign of the effect. It rules out
a collapse, not a subtle cost, and the write-up says so.

The grading rule was fixed before any output was read (number after the
last ####, falling back to the last number in the reply, applied
identically to every cell) and answers.csv carries every reply in full so
it can be audited. The prompt set, driver and grader are committed with
the results.

README gains a section stating how the three kinds of evidence are kept
apart -- gates assert correctness, device runs measure speed, a public
benchmark checks quality -- and where each comes from, including the
GSM8K source and licence.
2026-07-22 20:04:01 +02:00
..
2026-07-12 docs(bench): commit the 2026-07-12 device matrix and align the benchmark docs 2026-07-13 11:41:21 +02:00
2026-07-12-pr23 docs: correct content overtaken by the code 2026-07-15 07:50:47 +02:00
2026-07-13 bench(android): 2026-07-13 device matrix — adaptive cache and reworked spec-gating 2026-07-13 10:10:00 +02:00
2026-07-14 docs(benchmarks): per-token warm-up analysis for Qwen/Gemma and gpt-oss 2026-07-14 17:25:00 +02:00
2026-07-14-warmup docs: rename adaptive-cache.md to cache-sizing.md 2026-07-19 10:56:16 +02:00
2026-07-15-route-trace docs: record the expert-sidecar negative result and ship the --scatter microbench 2026-07-20 17:02:48 +02:00
2026-07-17 docs: gpt-oss leads the README; all-O_DIRECT dense weights measured 2026-07-17 11:31:35 +02:00
2026-07-20-cache-replay docs: point the layer-lfu record at its tag, not a branch that is being deleted 2026-07-20 11:47:27 +02:00
2026-07-20-sidecar docs: record the expert-sidecar negative result and ship the --scatter microbench 2026-07-20 17:02:48 +02:00
2026-07-21-pinned-dense-ab feat(dense): --dense-weights ahwb — dense weights in memory Android cannot reclaim (#93) 2026-07-21 11:09:51 +02:00
2026-07-21-pinned-memory test(memory): reclaim-exempt memory for the dense weights — bandwidth gate passes, 2047 MiB cap (#92) 2026-07-21 11:08:24 +02:00
2026-07-22-drop-cold-experts feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95) 2026-07-22 17:21:55 +02:00
2026-07-22-drop-quality test(moe): GSM8K quality check for --drop-cold-experts finds no loss (#97) 2026-07-22 20:04:01 +02:00
README.md docs: rename adaptive-cache.md to cache-sizing.md 2026-07-19 10:56:16 +02:00

Benchmark data archive

Raw per-run CSVs and the session notes written alongside them, one directory per measurement session. Each directory is a historical record of the build it was captured on, kept for provenance so the tables in ../benchmarks.md can be traced back to data.

These files are not maintained. They are not corrected when the code moves on, and a recommendation inside one is only valid for the build it was measured against. Read the maintained docs for current guidance; read these to check where a number came from.

Session Captured against Note
2026-07-12/ baseline cache/lane sweep Feeds the cache-and-lanes tables in benchmarks.md.
2026-07-12-pr23/ temporal prefetch + speculative gating Speculative gating no longer exists — it was removed to restore the modular seam. The --spec-gate rows and the advice about it describe a flag the CLI no longer accepts.
2026-07-13/ adaptive cache budget Source of the capped-auto recipe numbers.
2026-07-14/ per-token warm-up Feeds warmup-analysis.md.
2026-07-14-warmup/ dense warm-up A/B Feeds cache-sizing.md and warmup-analysis.md.
2026-07-15-route-trace/ first --route-trace capture (Qwen / Gemma / gpt-oss) Routing data, not throughput: every run had the trace on, so its tok/s are not comparable with benchmarks.md and must not feed those tables.
2026-07-17/ all-O_DIRECT dense weights (--dense-weights anon) Source of the gpt-oss steady-state numbers. Several cells are contaminated by run order (a fixed cooldown does not reach baseline) — NOTES.md marks exactly which, and which claims survive. Do not read a top-k or lane claim out of it.