The library/CLI default was `warm` while the Android app has always shipped
`anon` ("the default — wins on >RAM models"). Align the CLI to the app: set
`DenseWeightsMode::Anonymous` as the config default so `bmoe-cli` without a
`--dense-weights` flag now reads the dense weights via O_DIRECT into anon
buffers, the policy that pays on the >RAM models the engine targets (3.2x on
gpt-oss). `warm` and `mmap` stay available for RAM-fitting models.
Byte-identity gates re-run green — anon rebind is already proven identical to
the mmap reference (G6/G7), so the default flip changes no output. Docs and the
CLI --help updated to reflect anon as the default.
Sending the dense (non-expert) weights through O_DIRECT into anonymous buffers
(--dense-weights anon) is decisive well past RAM: on gpt-oss-120b at 5.2x device
RAM it cuts major faults per token from the hundreds to 6-10 and compute from
0.948 to 0.156 s/tok -- 0.687 -> 2.191 tok/s, a 3.2x, measured at 256-token
steady state instead of the old 24-token probes.
That also overturns the cache advice for this model: with the dense set out of
the page cache the two claimants on RAM no longer fight, and a 2000 MiB expert
cache now beats cache-off (0.998 vs 0.711 at matched lanes, while clocked
lower). "cache-off is the ceiling past RAM" was true of a configuration, not of
the model; pressure.md and the README say so now.
README is restructured to lead with gpt-oss and quotes best-observed figures
across all sessions (Qwen 5.23, Gemma 4.09 lossless; 5.01/4.99 at k=6), with an
explicit note that these are best-observed and that device state moves them. New
"What to expect in the app" section: the demo APK reads 1.91 tok/s on the same
gpt-oss config vs 2.191 over adb (~13%), and the gap is the protocol (short
turns inside the cache warm-up window, live device state, co-resident UI) --
at true parity the app measured slightly faster.
Method: a fixed cooldown does not return the device to baseline, so a matrix
silently measures its own run order -- it inverted the Turbo top-k result
outright. benchmark-method.md now prescribes gating each cell on a measured
condition, reading the CPU sensor rather than the lagging battery one, and
logging the entry state; it also lists the two tells of a contaminated cell.
The Qwen/Gemma top-k pairs and the gpt-oss lane pair from this session are
contaminated and are published as data only, with no claim drawn from them.
Stale docs corrected: adaptive-cache.md still described the retired runtime
governor (and contradicted itself), --no-warm-dense was still presented as the
primary flag, and the docs index still carried the old cache-off claim.
Raw CSVs, metrics, generated text and drivers in docs/bench-data/2026-07-17/,
with per-cell confounds recorded in its NOTES.md.
benchmarks.md was two documents in one file: 460 lines, two H1s, and four
section names appearing twice. roadmap.md's link to #reading-the-numbers
resolved to whichever came first, which happened to be the intended one — a
coincidence, not a design. Split the gpt-oss-120b half into its own doc and
cross-link the two.
That also fixes the anchor at the old benchmarks.md:30: GitHub slugs an em-dash
heading to a double hyphen, so #device-pressure-not-just-tokens never jumped
anywhere. The README already had this right for its own gpt-oss link, so the
convention was there — this one was just wrong.
README: route traces and the app's Markdown answers have been in main for
several commits with no mention, and a feature nobody can find is a feature
nobody has. Trimmed the O_DIRECT and cache bullets in exchange: both re-taught
mechanism the linked docs already own, which is what a landing page delegates.
roadmap.md listed three shipped capabilities as future work: overlapping I/O with
compute (--overlap), prefetching the next layer's experts (--prefetch), and the
routing predictor behind them -- which was in fact built and then removed. It also
opened on 'decode is ~79% flash I/O', true only with the cache off; with a sized
cache and overlap decode is compute-bound, which changes what is worth doing next.
Rewrite it against the measurements, and say plainly why speculative gating is not
coming back.
benchmarks.md described Gemma as 'A4B (4 experts active)' at the top and corrected
that same claim 95 lines further down; keep the correction. CONTRIBUTING.md offered
the merged ffn_gate_up_exps layout as a good first contribution -- it shipped as
gemma4. Flag the archived spec-gate findings as describing a removed flag.
Add docs/warmup-analysis.md dissecting the streamed warm-up transient from the
per-token --csv, in two regimes by model/RAM ratio:
- Models near RAM (Qwen ~1.6x, Gemma ~1.5x): I/O-bound warm-up. compute_ms is
flat; the cache-hit climb (4.5% -> 77-83%) carries the warm-up and overlap
hides it, so tok/s recovers within ~3 tokens.
- gpt-oss-120b (~5.2x RAM): memory-residency-bound warm-up. Streaming still
bounds expert memory via O_DIRECT, but the mmap-resident non-expert set faults
in under near-zero free RAM, surfacing inside compute_ms (~18s -> ~0.3s), which
overlap cannot hide; a 24-token probe is mostly this cold head.
Raw per-token CSVs in docs/bench-data/2026-07-14/warmup/. The gpt-oss runs were
on a thermally-degraded device and are flagged as a floor, not a headline number.
Cross-linked from both Reading-the-numbers sections in benchmarks.md.
Document the first on-device run of a 58 GB / 120B MoE (gpt-oss-120b at
5.2x device RAM on the OnePlus 15R): a top-k x lanes x prefetch sweep and
an mmap baseline, streaming at up to 7.7x a plain mmap load (top-k 2).
- docs/benchmarks.md: full 12-cell matrix + 2 mmap rows, with honest
caveats (24-token probe, not steady state; the k=4 rows straddled USB
interruptions and are marked/excluded) and a quality note -- --no-think
drops gpt-oss's reasoning, so default top-4 answers 17x23 wrong while
k=2/3 answer right.
- docs/benchmark-method.md: how to benchmark harmony models (--no-think to
reach the answer, the /data path for O_DIRECT, the reasoning trade-off).
- README + CHANGELOG: headline (a 120B on a phone) and the gpt-oss section.
- scripts/gptoss-matrix.sh + gptoss-mmap.sh: the exact drivers.
- docs/bench-data/2026-07-14/: raw summary log + the two mmap per-token CSVs.
Matched A/B (default vs k=6, same session, same 4000 MiB cache / 4 lanes,
thermally comparable) for Qwen3-30B-A3B and Gemma-4-26B-A4B:
- Qwen 4.03 -> 5.01 tok/s (+24.3%), flash read 225 -> 165 MiB/token (-26.8%)
- Gemma 4.09 -> 4.99 tok/s (+22.1%), flash read 144 -> 98 MiB/token (-31.8%)
The -26.8% I/O on Qwen matches 6/8, confirming its default top-8. Output
diverges under greedy decode but stayed coherent and accurate on the essay
prompt. Adds the --n-expert-used axis to the method sweep, with a note to
run it as a same-session matched A/B (cool-vs-warm baseline swamps it) and
to inspect text quality alongside tok/s.
The 2026-07-12 base matrix (Qwen and Gemma, cache/lane/overlap sweep) that the
README and benchmarks.md cite was left untracked; commit it so the cited numbers
have their provenance in-tree, and align the three benchmark docs with it.
256-token steady-state runs on the OnePlus 15R for two MoE families across the
cache/lane matrix. Adds benchmarks.md (the measured tables), the committed
drivers (bench-run.sh, bench-matrix.ps1, bench-analyze.py), and per-row flag +
model-file provenance so every number reproduces from the logs.
Records the decode compute+I/O split: at a 4 GiB cache decode is compute-bound
(I/O ~0.1 s/tok), so the streaming ceiling is the SoC's in-RAM speed (~7 tok/s
Qwen, ~5.3 Gemma). Ignores .bench/ and local *.log scratch.