Commit graph

6 commits

Author SHA1 Message Date
Helldez
d10d10f1b8 docs: correct the benchmark recipe, the arch list and mismatched sizes
benchmark-method contradicted its own protocol - it requires a fixed -n of
at least 256 so the expert cache reaches steady state, then demonstrated
with -n 48 - and staged the model in the app external files dir, which is
FUSE-backed and where O_DIRECT silently returns wrong data. A run measured
there falls back to buffered I/O, so it does not measure the path the rest
of the method assumes.

roadmap listed only qwen3moe, gemma4 and gpt-oss as supported while
qwen2moe and qwen35moe are registered and shipped, and still credited the
page-cache warm-up with removing the >RAM fault storm - anon is the default
and the mechanism that actually holds past RAM.

Sizes: gpt-oss-120b was 62 GB in one place and 58.5 elsewhere; the
android-memory hit-rate table gave Qwen and Gemma in GiB under a GB label
while the gpt-oss row on the same table was decimal GB.
2026-07-19 11:27:18 +02:00
Helldez
4be2e177c1
feat(cli): default --dense-weights to anon, matching the Android app (#55)
The library/CLI default was `warm` while the Android app has always shipped
`anon` ("the default — wins on >RAM models"). Align the CLI to the app: set
`DenseWeightsMode::Anonymous` as the config default so `bmoe-cli` without a
`--dense-weights` flag now reads the dense weights via O_DIRECT into anon
buffers, the policy that pays on the >RAM models the engine targets (3.2x on
gpt-oss). `warm` and `mmap` stay available for RAM-fitting models.

Byte-identity gates re-run green — anon rebind is already proven identical to
the mmap reference (G6/G7), so the default flip changes no output. Docs and the
CLI --help updated to reflect anon as the default.
2026-07-18 09:39:18 +02:00
Helldez
00bf75aea9 docs: gpt-oss leads the README; all-O_DIRECT dense weights measured
Sending the dense (non-expert) weights through O_DIRECT into anonymous buffers
(--dense-weights anon) is decisive well past RAM: on gpt-oss-120b at 5.2x device
RAM it cuts major faults per token from the hundreds to 6-10 and compute from
0.948 to 0.156 s/tok -- 0.687 -> 2.191 tok/s, a 3.2x, measured at 256-token
steady state instead of the old 24-token probes.

That also overturns the cache advice for this model: with the dense set out of
the page cache the two claimants on RAM no longer fight, and a 2000 MiB expert
cache now beats cache-off (0.998 vs 0.711 at matched lanes, while clocked
lower). "cache-off is the ceiling past RAM" was true of a configuration, not of
the model; pressure.md and the README say so now.

README is restructured to lead with gpt-oss and quotes best-observed figures
across all sessions (Qwen 5.23, Gemma 4.09 lossless; 5.01/4.99 at k=6), with an
explicit note that these are best-observed and that device state moves them. New
"What to expect in the app" section: the demo APK reads 1.91 tok/s on the same
gpt-oss config vs 2.191 over adb (~13%), and the gap is the protocol (short
turns inside the cache warm-up window, live device state, co-resident UI) --
at true parity the app measured slightly faster.

Method: a fixed cooldown does not return the device to baseline, so a matrix
silently measures its own run order -- it inverted the Turbo top-k result
outright. benchmark-method.md now prescribes gating each cell on a measured
condition, reading the CPU sensor rather than the lagging battery one, and
logging the entry state; it also lists the two tells of a contaminated cell.
The Qwen/Gemma top-k pairs and the gpt-oss lane pair from this session are
contaminated and are published as data only, with no claim drawn from them.

Stale docs corrected: adaptive-cache.md still described the retired runtime
governor (and contradicted itself), --no-warm-dense was still presented as the
primary flag, and the docs index still carried the old cache-off claim.

Raw CSVs, metrics, generated text and drivers in docs/bench-data/2026-07-17/,
with per-cell confounds recorded in its NOTES.md.
2026-07-17 11:31:35 +02:00
Helldez
584e96ca6a docs(telemetry): document the compute decomposition and stall floor
Add majflt / cpu_ms to the BMOE_PROGRESS, BMOE_DONE and CSV schemas and explain
how they attribute the compute_ms residual (fault stall vs. throttled core vs.
genuine matmul), plus the new compute: summary line. Document why stall_ms has a
structural floor above zero: the router picks a token's experts only just before
the FFN needs them, so a cache miss forces an on-demand read the overlap cannot
hide, and the residual stall tracks the miss rate (never zero below ~100% hit).
Extend the warmup-analysis CSV schema table with the two columns and record the
CHANGELOG entries.
2026-07-15 07:03:04 +02:00
Helldez
4c19838dd7 docs(warmup): document the dense warm-up fix with on-device results
Close the loop opened by the per-token warm-up analysis: add a "The fix" section
to warmup-analysis.md and fold the warm-up into adaptive-cache.md. Include the
before/after per-token CSVs (bench-data/2026-07-14-warmup/) and the reproducer
script (scripts/bench-warmonly.sh).

Also record the rejected alternative — reserving the dense bytes in the auto
budget — which lowered Gemma's hit rate (4000->2909 MiB, 83%->73%) without the
warm-up's benefit, and was dropped in its favour.
2026-07-14 18:44:15 +02:00
Helldez
b640e8d432 docs(benchmarks): per-token warm-up analysis for Qwen/Gemma and gpt-oss
Add docs/warmup-analysis.md dissecting the streamed warm-up transient from the
per-token --csv, in two regimes by model/RAM ratio:

- Models near RAM (Qwen ~1.6x, Gemma ~1.5x): I/O-bound warm-up. compute_ms is
  flat; the cache-hit climb (4.5% -> 77-83%) carries the warm-up and overlap
  hides it, so tok/s recovers within ~3 tokens.
- gpt-oss-120b (~5.2x RAM): memory-residency-bound warm-up. Streaming still
  bounds expert memory via O_DIRECT, but the mmap-resident non-expert set faults
  in under near-zero free RAM, surfacing inside compute_ms (~18s -> ~0.3s), which
  overlap cannot hide; a 24-token probe is mostly this cold head.

Raw per-token CSVs in docs/bench-data/2026-07-14/warmup/. The gpt-oss runs were
on a thermally-degraded device and are flagged as a floor, not a headline number.
Cross-linked from both Reading-the-numbers sections in benchmarks.md.
2026-07-14 17:25:00 +02:00