Commit graph

11 commits

Author SHA1 Message Date
Helldez
631d98f01a docs: record the expert-sidecar negative result and ship the --scatter microbench 2026-07-20 17:02:48 +02:00
Helldez
4be2e177c1
feat(cli): default --dense-weights to anon, matching the Android app (#55)
The library/CLI default was `warm` while the Android app has always shipped
`anon` ("the default — wins on >RAM models"). Align the CLI to the app: set
`DenseWeightsMode::Anonymous` as the config default so `bmoe-cli` without a
`--dense-weights` flag now reads the dense weights via O_DIRECT into anon
buffers, the policy that pays on the >RAM models the engine targets (3.2x on
gpt-oss). `warm` and `mmap` stay available for RAM-fitting models.

Byte-identity gates re-run green — anon rebind is already proven identical to
the mmap reference (G6/G7), so the default flip changes no output. Docs and the
CLI --help updated to reflect anon as the default.
2026-07-18 09:39:18 +02:00
Helldez
00bf75aea9 docs: gpt-oss leads the README; all-O_DIRECT dense weights measured
Sending the dense (non-expert) weights through O_DIRECT into anonymous buffers
(--dense-weights anon) is decisive well past RAM: on gpt-oss-120b at 5.2x device
RAM it cuts major faults per token from the hundreds to 6-10 and compute from
0.948 to 0.156 s/tok -- 0.687 -> 2.191 tok/s, a 3.2x, measured at 256-token
steady state instead of the old 24-token probes.

That also overturns the cache advice for this model: with the dense set out of
the page cache the two claimants on RAM no longer fight, and a 2000 MiB expert
cache now beats cache-off (0.998 vs 0.711 at matched lanes, while clocked
lower). "cache-off is the ceiling past RAM" was true of a configuration, not of
the model; pressure.md and the README say so now.

README is restructured to lead with gpt-oss and quotes best-observed figures
across all sessions (Qwen 5.23, Gemma 4.09 lossless; 5.01/4.99 at k=6), with an
explicit note that these are best-observed and that device state moves them. New
"What to expect in the app" section: the demo APK reads 1.91 tok/s on the same
gpt-oss config vs 2.191 over adb (~13%), and the gap is the protocol (short
turns inside the cache warm-up window, live device state, co-resident UI) --
at true parity the app measured slightly faster.

Method: a fixed cooldown does not return the device to baseline, so a matrix
silently measures its own run order -- it inverted the Turbo top-k result
outright. benchmark-method.md now prescribes gating each cell on a measured
condition, reading the CPU sensor rather than the lagging battery one, and
logging the entry state; it also lists the two tells of a contaminated cell.
The Qwen/Gemma top-k pairs and the gpt-oss lane pair from this session are
contaminated and are published as data only, with no claim drawn from them.

Stale docs corrected: adaptive-cache.md still described the retired runtime
governor (and contradicted itself), --no-warm-dense was still presented as the
primary flag, and the docs index still carried the old cache-off claim.

Raw CSVs, metrics, generated text and drivers in docs/bench-data/2026-07-17/,
with per-cell confounds recorded in its NOTES.md.
2026-07-17 11:31:35 +02:00
Helldez
bb318cbd2e docs: split the benchmarks, and say what the engine grew
benchmarks.md was two documents in one file: 460 lines, two H1s, and four
section names appearing twice. roadmap.md's link to #reading-the-numbers
resolved to whichever came first, which happened to be the intended one — a
coincidence, not a design. Split the gpt-oss-120b half into its own doc and
cross-link the two.

That also fixes the anchor at the old benchmarks.md:30: GitHub slugs an em-dash
heading to a double hyphen, so #device-pressure-not-just-tokens never jumped
anywhere. The README already had this right for its own gpt-oss link, so the
convention was there — this one was just wrong.

README: route traces and the app's Markdown answers have been in main for
several commits with no mention, and a feature nobody can find is a feature
nobody has. Trimmed the O_DIRECT and cache bullets in exchange: both re-taught
mechanism the linked docs already own, which is what a landing page delegates.
2026-07-15 20:43:41 +02:00
Helldez
44c2c2db73 docs: correct content overtaken by the code
roadmap.md listed three shipped capabilities as future work: overlapping I/O with
compute (--overlap), prefetching the next layer's experts (--prefetch), and the
routing predictor behind them -- which was in fact built and then removed. It also
opened on 'decode is ~79% flash I/O', true only with the cache off; with a sized
cache and overlap decode is compute-bound, which changes what is worth doing next.
Rewrite it against the measurements, and say plainly why speculative gating is not
coming back.

benchmarks.md described Gemma as 'A4B (4 experts active)' at the top and corrected
that same claim 95 lines further down; keep the correction. CONTRIBUTING.md offered
the merged ffn_gate_up_exps layout as a good first contribution -- it shipped as
gemma4. Flag the archived spec-gate findings as describing a removed flag.
2026-07-15 07:50:47 +02:00
Helldez
b640e8d432 docs(benchmarks): per-token warm-up analysis for Qwen/Gemma and gpt-oss
Add docs/warmup-analysis.md dissecting the streamed warm-up transient from the
per-token --csv, in two regimes by model/RAM ratio:

- Models near RAM (Qwen ~1.6x, Gemma ~1.5x): I/O-bound warm-up. compute_ms is
  flat; the cache-hit climb (4.5% -> 77-83%) carries the warm-up and overlap
  hides it, so tok/s recovers within ~3 tokens.
- gpt-oss-120b (~5.2x RAM): memory-residency-bound warm-up. Streaming still
  bounds expert memory via O_DIRECT, but the mmap-resident non-expert set faults
  in under near-zero free RAM, surfacing inside compute_ms (~18s -> ~0.3s), which
  overlap cannot hide; a 24-token probe is mostly this cold head.

Raw per-token CSVs in docs/bench-data/2026-07-14/warmup/. The gpt-oss runs were
on a thermally-degraded device and are flagged as a floor, not a headline number.
Cross-linked from both Reading-the-numbers sections in benchmarks.md.
2026-07-14 17:25:00 +02:00
Helldez
01390689f1 docs(benchmark): gpt-oss-120b on-device streaming results + drivers
Document the first on-device run of a 58 GB / 120B MoE (gpt-oss-120b at
5.2x device RAM on the OnePlus 15R): a top-k x lanes x prefetch sweep and
an mmap baseline, streaming at up to 7.7x a plain mmap load (top-k 2).

- docs/benchmarks.md: full 12-cell matrix + 2 mmap rows, with honest
  caveats (24-token probe, not steady state; the k=4 rows straddled USB
  interruptions and are marked/excluded) and a quality note -- --no-think
  drops gpt-oss's reasoning, so default top-4 answers 17x23 wrong while
  k=2/3 answer right.
- docs/benchmark-method.md: how to benchmark harmony models (--no-think to
  reach the answer, the /data path for O_DIRECT, the reasoning trade-off).
- README + CHANGELOG: headline (a 120B on a phone) and the gpt-oss section.
- scripts/gptoss-matrix.sh + gptoss-mmap.sh: the exact drivers.
- docs/bench-data/2026-07-14/: raw summary log + the two mmap per-token CSVs.
2026-07-14 12:29:16 +02:00
Helldez
d2ecb83758 docs: lead intro with vs-mmap headline for both models, explain O_DIRECT, reconcile the Gemma cache-4000 note 2026-07-13 16:29:46 +02:00
Helldez
92a2f8b024 docs(bench): measured Turbo top-k (--n-expert-used 6) on OnePlus 15R
Matched A/B (default vs k=6, same session, same 4000 MiB cache / 4 lanes,
thermally comparable) for Qwen3-30B-A3B and Gemma-4-26B-A4B:

- Qwen  4.03 -> 5.01 tok/s (+24.3%), flash read 225 -> 165 MiB/token (-26.8%)
- Gemma 4.09 -> 4.99 tok/s (+22.1%), flash read 144 -> 98 MiB/token (-31.8%)

The -26.8% I/O on Qwen matches 6/8, confirming its default top-8. Output
diverges under greedy decode but stayed coherent and accurate on the essay
prompt. Adds the --n-expert-used axis to the method sweep, with a note to
run it as a same-session matched A/B (cool-vs-warm baseline swamps it) and
to inspect text quality alongside tok/s.
2026-07-13 11:52:38 +02:00
Helldez
fbab83cd06 docs(bench): commit the 2026-07-12 device matrix and align the benchmark docs
The 2026-07-12 base matrix (Qwen and Gemma, cache/lane/overlap sweep) that the
README and benchmarks.md cite was left untracked; commit it so the cited numbers
have their provenance in-tree, and align the three benchmark docs with it.
2026-07-13 11:41:21 +02:00
Helldez
d862300d59 docs: full Qwen+Gemma benchmark matrix with compute/IO decode split
256-token steady-state runs on the OnePlus 15R for two MoE families across the
cache/lane matrix. Adds benchmarks.md (the measured tables), the committed
drivers (bench-run.sh, bench-matrix.ps1, bench-analyze.py), and per-row flag +
model-file provenance so every number reproduces from the logs.

Records the decode compute+I/O split: at a 4 GiB cache decode is compute-bound
(I/O ~0.1 s/tok), so the streaming ceiling is the SoC's in-RAM speed (~7 tok/s
Qwen, ~5.3 Gemma). Ignores .bench/ and local *.log scratch.
2026-07-11 22:44:43 +02:00