Commit graph

12 commits

Author SHA1 Message Date
Helldez
d9381f6363 docs: benchmark pages as a contributor guide, and the Apple platform limits
The benchmark call is pinned, and both pages were written for the maintainer
rather than for the people landing on them. community-benchmarks.md opened with
a list of hardware we want, which reads as an entry requirement, and neither
page ever answered the first question a contributor has: what do I set?

community-benchmarks.md:
- a "start here" for the three cases someone is actually in (PC or laptop,
  Android phone, Apple hardware), each with the command and what to paste back
- the settings-override table, which used to be one buried sentence
- an explicit adb protocol for phones
- "hardware we want to see" moved to the end as open questions: any hardware is
  a useful row, the list is what we cannot answer ourselves

benchmark-method.md:
- reference device out of the opening, named once at the end as the provenance
  of the published numbers
- new "choosing the parameters": lossless, lossy and experimental knobs kept in
  separate tables, each with its default and the telemetry field that says
  whether moving it worked
- the hard-won rules kept as method rather than as the story of one session

The community protocol now pins --ubatch 512, which the app has always done and
bench-report.sh never did: prefill width costs resident memory the expert cache
would otherwise get, so a host row was running a different configuration from
the app it is compared against. UBATCH= overrides it.

Two platform limits documented for the first time: macOS has no O_DIRECT and
the engine does not call the F_NOCACHE equivalent, so expert reads there go
through the page cache while the metrics still report o_direct=1; and there is
no iOS target at all. Both were already true.
2026-08-28 16:53:22 +02:00
Helldez
bea5a0b99e
feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95)
* feat(moe): --drop-cold-experts, spend quality only where it buys I/O

Turbo top-k drops the tail of a routing whether or not those experts were
already in RAM. A resident expert costs no flash read, so that trade pays
quality for nothing on the ~80% of decode routings that are cache hits.

This skips a routed expert only when it is a cache MISS and the router
weighted it below frac x (1/top-k). Replayed over the committed route
traces at frac 1.0, decode phase, that avoids 66% of flash reads for 9.5%
of the router's weight mass, where --n-expert-used 5 avoids 23% for a
comparable 10.6% -- about 3x the reads at the same quality cost.

Implementation. The decision needs the FINAL router weights, which arrive
several nodes after the topk where the streamer normally loads, so
load_layer() is deferred to the terminal node of the layer's weight chain.
Which node that is depends on the model's gating, so the hook learns it
from the graph rather than carrying an architecture table; if it fails to
arrive the hook forgets it and re-learns rather than re-betting. A dropped
slot has its weight zeroed and its expert id repointed at the routing's
top-weighted expert: an unread expert can sit in reserved-but-uncommitted
VM and mul_mat_id would touch it anyway, so the kernel is given memory
that is certainly resident and multiplies it by exactly zero.

Requires the LRU cache -- with --cache-mb 0 residency reads all-miss and
the policy would silently degenerate into an unconditional weight cut.
Prefill is excluded by default. The top expert is always pinned, so no
routing can be emptied at any threshold.

Gates: G8a/G8a' prove the deferral and the learned terminal node are
transparent (byte-identical output, zero drops, at a threshold below any
producible weight); G8b that full strength against a constantly-evicting
cache never reaches an unloaded slot; G8c that at top-k 1 dropping is a
no-op, pinning both the top-expert guarantee and the threshold tracking
the effective top-k.

Three existing metrics shift meaning under dropping and the docs now say
so: cache_hit_pct rises without the cache serving more (a dropped routing
is a miss that is never looked up), and token/layer_demand measure what
was staged rather than routed. prefetch.md's "cannot change output" is
scoped, limitations.md gains the non-reproducibility entry, and
benchmark-method.md warns that reversing the run order cannot distinguish
a moved drop rate from a contaminated cell.

Off by default in the CLI and in the app. The output is not reproducible
-- what gets dropped depends on what the cache held -- so it carries no
rows in the README tables, and switching it on by default waits on a
published on-device A/B rather than on the replay argument alone.

* feat(app): default cache-aware dropping to 75%, measured on device

Qwen3.6-35B-A3B (top-k 8 of 256), in-app, cache 3000, one variable
changed: 2.549 tok/s off, 3.938 at F=0.75 (+55%), 4.702 at F=1.0 (+84%),
with flash reads falling 248 -> 163 -> 48 GiB. Per-token bootstrap
intervals separate every pair except off vs 0.50, which overlaps -- at
half the uniform share the policy drops 2.7% of routings and buys
nothing, which doubles as a negative control that the machinery is free
when it does not fire.

Run order was 1.0, off, 0.5, 0.75, so the two fastest cells are the first
and the LAST; thermal drift would have made the last the worst. The
mechanism orders by threshold even though the run order does not.

The replay turned out conservative rather than optimistic. It is
documented as an upper bound because it cannot model the cache changing
in response to dropping: at F=0.75 it was accurate (37% predicted, 34%
measured), at F=1.0 it understated (66% predicted, 81% measured). Avoided
reads free cache capacity, which raises the hit rate, which leaves fewer
misses to drop.

75% rather than 100% is deliberate: it takes the larger part of the win
for half the discarded routings (14% against 28%). Quality is still
unquantified -- no perplexity number and no side-by-side exists -- so the
conservative end of a measured range is the defensible default. The CLI
stays off; the byte-identity gates need a deterministic default.

Also records cache_hit_pct rising 67.8 -> 90.7% as the documented
accounting artefact rather than the cache serving more, and majflt/token
as dominated by each run's starting memory state, not by the threshold.
2026-07-22 17:21:55 +02:00
Helldez
79611654cd Revert "feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#94)"
This reverts 45a90a2.

The feature merged before the evidence for its shipping default did. The
replay numbers argue the shape of the trade is favourable, but no on-device
A/B is published in this repository, and the app default it landed with
(75%) changes model output for every user of the demo app -- and changes it
non-reproducibly, which no other setting in this engine does.

Nothing was found wrong with the code. This is a sequencing decision: the
work returns as a pull request, with the app default back to off, so the
measurement lands before the default does.

Reverted rather than force-pushed: main is public and this commit was
already pushed, so the history stays honest about what happened.
2026-07-22 15:47:14 +02:00
Helldez
45a90a2df5
feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#94)
* feat(moe): --drop-cold-experts, spend quality only where it buys I/O

Turbo top-k drops the tail of a routing whether or not those experts were
already in RAM. A resident expert costs no flash read, so that trade pays
quality for nothing on the ~80% of decode routings that are cache hits.

This adds the cache-aware version: skip a routed expert only when it is a
cache MISS and the router weighted it below frac x (1/top-k). Replayed over
the committed route traces at frac 1.0, decode phase, that avoids 66% of
flash reads for 9.5% of the router's weight mass, against 59%/37% for
--n-expert-used 3 — about 3x the reads avoided at a comparable cost. The
threshold is a curve, not a switch: 0.75 trades 4.4% of the mass for 37% of
the reads, better than --n-expert-used 5 on both axes.

Implementation. The decision needs the FINAL router weights, which arrive
several nodes after the topk where the streamer normally loads, so with the
policy armed load_layer() is deferred to the terminal node of the layer's
weight chain. Which node that is depends on the model's gating, so the hook
learns it from the graph rather than carrying an architecture table; until
it is known a layer loads at its topk node undropped. A dropped slot has its
weight zeroed and its expert id repointed at the routing's top-weighted
expert: an expert we decline to read may sit in reserved-but-uncommitted VM
and mul_mat_id would still touch it, so the kernel is given memory that is
certainly resident and multiplies it by exactly zero. Survivors are rescaled
by default, since a systematically shrunk expert output perturbs the residual
stream more than the missing contribution does.

Prefill is excluded by default (cold cache, ~4x the weight mass discarded,
and compute-bound anyway). The largest weight in a routing is always at least
the uniform share, so frac <= 1 can never empty a layer; validate() enforces
the bound and the top expert is pinned regardless.

Gates: G8a proves the deferral and the learned terminal node are transparent
(a threshold below any producible weight leaves the output byte-identical),
G8b that full strength with the cache off never reaches an unloaded expert.

Unlike every other knob this one is state-dependent: what gets dropped
depends on what the cache held, so output is not reproducible across runs.
Off by default, not in the app's settings, and NOT yet measured on device —
the numbers above are a static replay and an upper bound. docs/expert-
dropping.md states what is owed before it is recommended anywhere.

* feat(app): expose cache-aware expert dropping in Settings

Speed / quality -> Drop cold experts, as a percentage of the uniform
share (off / 50 / 75 / 100). The engine takes a fraction; the app stores
integer rungs, so the setting divides by 100 on the way to the flag.

Disabled in mmap mode: the policy asks the expert source what is resident,
and there is no expert source without the streamer. Included in the session
signature, so changing it reopens the session rather than being ignored by
a process already loaded.

Off by default. This exists so the A/B can be run where the engine actually
ships -- through the app, not a pushed CLI binary.

* fix(moe): require the cache for dropping, and correct what it reports

Review of the first two commits found the policy could be armed in a
configuration where it is not cache-aware at all, and that two of the
numbers it reports were wrong.

- Require the LRU cache. With --cache-mb 0 query_residency answers
  all-miss, so the policy silently degenerated into an unconditional
  weight cut -- exactly what --n-expert-used already does, under a flag
  claiming to consult residency. validate() now rejects it, as it already
  did for --prefetch, and the app gates the setting on the same condition.

- Fix experts_routed. It was incremented inside apply_drop, so it counted
  what the policy examined rather than what the router selected: layers
  before the terminal weight node is learned, and every un-armed phase,
  were missing from the denominator. The reported drop rate was a fraction
  of the wrong thing.

- Re-learn instead of re-betting. If the node learned as terminal does not
  arrive, the deferral now also forgets it, so the next graph loads at the
  topk node while it re-learns. Deferring again on a stale guess would
  repeat the fault every token against a graph that had moved.

- Point the gates at a real cache. G8a/G8b ran with the cache off, where
  the shared-slot path has no reserved-but-uncommitted memory -- so the id
  repointing, which is the design's whole safety argument, was never
  exercised. They now run against a constantly-evicting budget. Adds G8a'
  (asserts routings were examined and none dropped, so an inert-threshold
  flake fails legibly) and G8c (at top-k 1 dropping is a proven no-op,
  pinning both the top-expert guarantee and the threshold being taken
  against the effective top-k).

Docs: three metrics change meaning under dropping and none of them said
so. A dropped routing is a miss that is never looked up, so cache_hit_pct
rises without the cache serving more, and token/layer_demand measure what
was staged rather than routed -- documented in telemetry.md, pressure.md
(size the cache with dropping off, then turn it on) and metrics.h.
prefetch.md's "cannot change output" is scoped: under dropping a correct
guess un-drops an expert. limitations.md gains the non-reproducibility
entry, benchmark-method.md the axis plus a warning that reversing the run
order cannot distinguish a moved drop rate from a contaminated cell, and
architecture.md/runtime.h no longer claim unconditional determinism.
Fixes two anchors the README rename broke, and a changelog sentence that
quoted the equal-I/O row while drawing the equal-quality conclusion.

App: Drop cold experts defaults to 75%. The default is a product decision
taken on the maintainer's device; no benchmark for it is published here,
and docs/expert-dropping.md says that plainly instead of implying a
measured figure. The CLI stays off by default -- the byte-identity gates
need a deterministic default.
2026-07-22 15:08:49 +02:00
Helldez
d10d10f1b8 docs: correct the benchmark recipe, the arch list and mismatched sizes
benchmark-method contradicted its own protocol - it requires a fixed -n of
at least 256 so the expert cache reaches steady state, then demonstrated
with -n 48 - and staged the model in the app external files dir, which is
FUSE-backed and where O_DIRECT silently returns wrong data. A run measured
there falls back to buffered I/O, so it does not measure the path the rest
of the method assumes.

roadmap listed only qwen3moe, gemma4 and gpt-oss as supported while
qwen2moe and qwen35moe are registered and shipped, and still credited the
page-cache warm-up with removing the >RAM fault storm - anon is the default
and the mechanism that actually holds past RAM.

Sizes: gpt-oss-120b was 62 GB in one place and 58.5 elsewhere; the
android-memory hit-rate table gave Qwen and Gemma in GiB under a GB label
while the gpt-oss row on the same table was decimal GB.
2026-07-19 11:27:18 +02:00
Helldez
00bf75aea9 docs: gpt-oss leads the README; all-O_DIRECT dense weights measured
Sending the dense (non-expert) weights through O_DIRECT into anonymous buffers
(--dense-weights anon) is decisive well past RAM: on gpt-oss-120b at 5.2x device
RAM it cuts major faults per token from the hundreds to 6-10 and compute from
0.948 to 0.156 s/tok -- 0.687 -> 2.191 tok/s, a 3.2x, measured at 256-token
steady state instead of the old 24-token probes.

That also overturns the cache advice for this model: with the dense set out of
the page cache the two claimants on RAM no longer fight, and a 2000 MiB expert
cache now beats cache-off (0.998 vs 0.711 at matched lanes, while clocked
lower). "cache-off is the ceiling past RAM" was true of a configuration, not of
the model; pressure.md and the README say so now.

README is restructured to lead with gpt-oss and quotes best-observed figures
across all sessions (Qwen 5.23, Gemma 4.09 lossless; 5.01/4.99 at k=6), with an
explicit note that these are best-observed and that device state moves them. New
"What to expect in the app" section: the demo APK reads 1.91 tok/s on the same
gpt-oss config vs 2.191 over adb (~13%), and the gap is the protocol (short
turns inside the cache warm-up window, live device state, co-resident UI) --
at true parity the app measured slightly faster.

Method: a fixed cooldown does not return the device to baseline, so a matrix
silently measures its own run order -- it inverted the Turbo top-k result
outright. benchmark-method.md now prescribes gating each cell on a measured
condition, reading the CPU sensor rather than the lagging battery one, and
logging the entry state; it also lists the two tells of a contaminated cell.
The Qwen/Gemma top-k pairs and the gpt-oss lane pair from this session are
contaminated and are published as data only, with no claim drawn from them.

Stale docs corrected: adaptive-cache.md still described the retired runtime
governor (and contradicted itself), --no-warm-dense was still presented as the
primary flag, and the docs index still carried the old cache-off claim.

Raw CSVs, metrics, generated text and drivers in docs/bench-data/2026-07-17/,
with per-cell confounds recorded in its NOTES.md.
2026-07-17 11:31:35 +02:00
Helldez
01390689f1 docs(benchmark): gpt-oss-120b on-device streaming results + drivers
Document the first on-device run of a 58 GB / 120B MoE (gpt-oss-120b at
5.2x device RAM on the OnePlus 15R): a top-k x lanes x prefetch sweep and
an mmap baseline, streaming at up to 7.7x a plain mmap load (top-k 2).

- docs/benchmarks.md: full 12-cell matrix + 2 mmap rows, with honest
  caveats (24-token probe, not steady state; the k=4 rows straddled USB
  interruptions and are marked/excluded) and a quality note -- --no-think
  drops gpt-oss's reasoning, so default top-4 answers 17x23 wrong while
  k=2/3 answer right.
- docs/benchmark-method.md: how to benchmark harmony models (--no-think to
  reach the answer, the /data path for O_DIRECT, the reasoning trade-off).
- README + CHANGELOG: headline (a 120B on a phone) and the gpt-oss section.
- scripts/gptoss-matrix.sh + gptoss-mmap.sh: the exact drivers.
- docs/bench-data/2026-07-14/: raw summary log + the two mmap per-token CSVs.
2026-07-14 12:29:16 +02:00
Helldez
92a2f8b024 docs(bench): measured Turbo top-k (--n-expert-used 6) on OnePlus 15R
Matched A/B (default vs k=6, same session, same 4000 MiB cache / 4 lanes,
thermally comparable) for Qwen3-30B-A3B and Gemma-4-26B-A4B:

- Qwen  4.03 -> 5.01 tok/s (+24.3%), flash read 225 -> 165 MiB/token (-26.8%)
- Gemma 4.09 -> 4.99 tok/s (+22.1%), flash read 144 -> 98 MiB/token (-31.8%)

The -26.8% I/O on Qwen matches 6/8, confirming its default top-8. Output
diverges under greedy decode but stayed coherent and accurate on the essay
prompt. Adds the --n-expert-used axis to the method sweep, with a note to
run it as a same-session matched A/B (cool-vs-warm baseline swamps it) and
to inspect text quality alongside tok/s.
2026-07-13 11:52:38 +02:00
Helldez
fbab83cd06 docs(bench): commit the 2026-07-12 device matrix and align the benchmark docs
The 2026-07-12 base matrix (Qwen and Gemma, cache/lane/overlap sweep) that the
README and benchmarks.md cite was left untracked; commit it so the cited numbers
have their provenance in-tree, and align the three benchmark docs with it.
2026-07-13 11:41:21 +02:00
Helldez
d862300d59 docs: full Qwen+Gemma benchmark matrix with compute/IO decode split
256-token steady-state runs on the OnePlus 15R for two MoE families across the
cache/lane matrix. Adds benchmarks.md (the measured tables), the committed
drivers (bench-run.sh, bench-matrix.ps1, bench-analyze.py), and per-row flag +
model-file provenance so every number reproduces from the logs.

Records the decode compute+I/O split: at a 4 GiB cache decode is compute-bound
(I/O ~0.1 s/tok), so the streaming ceiling is the SoC's in-RAM speed (~7 tok/s
Qwen, ~5.3 Gemma). Ignores .bench/ and local *.log scratch.
2026-07-11 22:44:43 +02:00
Helldez
111fcdfaf8 docs: add measured desktop over-RAM benchmark (Qwen3-30B-A3B on 14.8 GiB PC) 2026-07-10 19:34:01 +02:00
Helldez
369ee5de96 docs: architecture, seam, MoE streaming, benchmarks, and agent guide 2026-07-10 18:18:24 +02:00