A session opened with --decide answers which of a list of choices the model would pick, read from
the next-token distribution after one prefill, with no decode. The state after a shared prefix is
kept and restored when the next prefix extends it. Android app: a Choose from options switch.
With --prefill-device, a decision is prefilled by the chat turn's placement rule and keeps no
prefix state: llama.cpp saves a sequence through KV views that do not follow the moved model state
(gate G18g). Also fixes the engine version, stuck at 0.23.0 since 0.24.0. App 0.27.0 (42).
Wide prefill graphs run on the Hexagon NPU through a two-layer arena streamed from flash; decode stays on the CPU. The app offers it in its own NPU section, off by default, Snapdragon only. A missing or unopenable device leaves the run on the CPU. release-apk builds the Hexagon backend and a skel per NPU generation in a separate, secret-free job. Bundles the llama.cpp bump to bmoe/expert-ready-hook-2609 (K-quants on the NPU). App 0.26.0 (41).
nemotron_h_moe is the third expert layout: gate-less. Each expert is
up, ReLU^2, down, so the registry row names ffn_up_exps and
ffn_down_exps and leaves the tail slot empty, as the fused gemma4 row
does. The Mamba2/attention blocks, the shared expert, the optional
latent projections and the MTP block all stay on the resident side of
the seam. No llama.cpp change and no submodule bump: the pinned tree
already builds nemotron_h_moe.
make-tiny-moe.py learns the whole shape in miniature (hybrid stack,
latent projections, biased sigmoid router, shared expert, a trailing
MTP block that is never loaded), and it runs as a third byte-identity
gate. Every identity gate passes on it.
The architecture never puts two MoE blocks next to each other, so the
forward predictors (predict-prefetch, route-ahead, the stale half of
predict-log) have no next layer to target. The gate reads that from the
file and reports those checks N/A instead of failing or passing them
vacuously; an unreadable file keeps them strict.
Ornith-1.5-35B-A3B is qwen35moe and needs no engine change. Both models
join the Android catalog at Q4_K_M; neither has device numbers yet.
Also releases 0.25.0: versionCode 40, versionName 0.25.0, dated changelog.
Experimental, off by default. Before a decode routing is committed, every
expert already in the LRU cache gets its score raised by L times the
token's score range and the top-k is taken again, so a near-tie goes to
the expert already in RAM (Skliar et al., arXiv:2412.00099). The same
number of experts runs; fewer are read from flash. Scores are read from
the tensor the graph itself sorted, exact for any gating function.
Desktop, Qwen3.6-35B Q4_K_M at L=0.15: 258 to 119 MiB of flash per
token, 2.37 to 3.84 tok/s, perplexity +1 to 4 %, tinyMMLU 88 to 84/100,
HumanEval-50 42 = 42. The on-device A/B is still owed, hence experimental.
Also: --ppl / --ppl-step / --ppl-list / --ppl-choices (teacher-forced
perplexity, one token per decode so cache-dependent policies are priced
where they act), scripts/tinymmlu-bench.py, scripts/humaneval-bench.py,
gates G8d/G8e, app switch "Prefer cached experts" under Experimental,
docs/cache-aware-substitution.md.
The benchmark call is pinned, and both pages were written for the maintainer
rather than for the people landing on them. community-benchmarks.md opened with
a list of hardware we want, which reads as an entry requirement, and neither
page ever answered the first question a contributor has: what do I set?
community-benchmarks.md:
- a "start here" for the three cases someone is actually in (PC or laptop,
Android phone, Apple hardware), each with the command and what to paste back
- the settings-override table, which used to be one buried sentence
- an explicit adb protocol for phones
- "hardware we want to see" moved to the end as open questions: any hardware is
a useful row, the list is what we cannot answer ourselves
benchmark-method.md:
- reference device out of the opening, named once at the end as the provenance
of the published numbers
- new "choosing the parameters": lossless, lossy and experimental knobs kept in
separate tables, each with its default and the telemetry field that says
whether moving it worked
- the hard-won rules kept as method rather than as the story of one session
The community protocol now pins --ubatch 512, which the app has always done and
bench-report.sh never did: prefill width costs resident memory the expert cache
would otherwise get, so a host row was running a different configuration from
the app it is compared against. UBATCH= overrides it.
Two platform limits documented for the first time: macOS has no O_DIRECT and
the engine does not call the F_NOCACHE equivalent, so expert reads there go
through the page cache while the metrics still report o_direct=1; and there is
no iOS target at all. Both were already true.
Qwen3.8-Flash-Next (qwen4exp): 125B total, ~6B active, 512 routed experts at
top-10 plus one shared, 48 hybrid gated-delta SSM / sparse attention layers,
and a 51B n-gram embedding table. One registry row streams the experts; a
dense-policy guard keeps the n-gram table (larger than any phone's RAM)
mmap'd under every mode so pinned and anonymous dense weights survive load.
Runs on the 12 GB test phone at ~2 tok/s with pinned dense weights, compute-
bound, and sits in the app catalog as a three-shard download. Submodule
pinned to upstream master b10666, the first with the architecture merged,
with the expert-ready hook on top. README hero clip, changelog and docs
updated. App 0.22.0 (versionCode 37).
Every published number comes from one phone and one laptop. This adds what a
contributor needs to add a row from hardware we do not own, without reading code:
- scripts/bench-report.sh runs the fixed README protocol on any Linux/macOS host
(256 greedy tokens, the reference prompt, auto cache, 4 lanes, overlap, dense
weights out of the page cache), records CPU / RAM / drive and the drive's
measured O_DIRECT rate at 512 KiB requests from the model file itself, and
prints the two markdown tables a report needs. Every figure is read from the
CSV `# summary` trailer by key name.
- .github/ISSUE_TEMPLATE/benchmark-report.yml collects hardware, model, engine
version and the pasted tables; tok/s alone is not accepted as a row.
- docs/community-benchmarks.md holds the protocol, the hardware wanted and why,
the reference models in the catalog quants, the meaning of each column, and
the results table seeded with the README rows.
- README, CONTRIBUTING and the docs index point at it.
* Mark the two doc-invoked host scripts executable
README, AGENTS, CONTRIBUTING and docs/seam.md run scripts/build-host.sh
directly, and docs/benchmarks.md runs scripts/bench-run.sh the same way,
but both are committed 100644, so a fresh POSIX clone answers permission
denied at the first documented step. Windows trees never see the mode
bit, which is likely how it went unnoticed.
* Mark the remaining host scripts executable too
Per the discussion in #163: leaving only the two doc-invoked scripts
executable hides the next instance of the same failure behind an
inconsistent set, so all five scripts/*.sh get the bit.
* feat(moe): stream split multi-shard ggufs natively + DeepSeek V4 Flash recipe
Hugging Face rejects single files above 50 GB, so every large model ships as
-00001-of-0000N.gguf shards; until now the streamer assumed one file, forcing
a merge with double the disk. gguf_offsets now fans the first shard out to the
whole set and resolves every tensor to (shard, offset); the expert streamer
and the dense loader open one positioned reader per shard and route each read
by the tensor's shard index. Pass the first shard, exactly as llama.cpp takes
it; a missing sibling fails the load with the shard named.
Add the deepseek4 recipe row: V3.2-style routing (256 routed experts, a
per-expert bias like lfm2moe, an always-on shared expert that stays resident)
over the standard split expert suffixes. The V4 compressed-attention machinery
is dense-side llama.cpp code, invisible to the streaming seam.
The byte-identity gates gain a 4-shard qwen3moe fixture (metadata-only first
shard, the layout large quants actually use); make-tiny-moe.py learns
--split-max-tensors. All gates pass, split included.
* fix(moe): cache auto must budget for the anon dense conversion
The auto budget read MemAvailable while the dense weights were still reclaimable
page cache, then dense-weights=anon converted them into buffers the kernel cannot
take back: the same bytes planned twice. Latent since the anon policy shipped
(dense sets were 2-3 GiB and explicit budgets were the benched path); DeepSeek V4
Flash's 6.5 GiB dense set turned it into a device-taking overcommit on first load.
The budget now deducts the pending conversion and says so in the log.
* fix(moe): review pass on the multi-shard path
Three defects the split rewrite introduced, none of which the gates could see:
- The shard index rode in an int8_t, so a model past 127 shards wrapped to a
negative index into the reader vector. The bounds check could never catch it:
it validated the untruncated value. Widened to int16_t, which covers the whole
-%05d-of-%05d filename space.
- DenseWeights::warm() reused one flag as both the inner loop condition and the
partial-warm report, so the first shard that failed to open silently skipped
the warm-up of every later shard. Per-shard condition, sticky report.
- The dense readers stayed allocated for the session after read_anonymous had
copied and rebound every tensor: fds and a per-lane bounce buffer per shard,
sitting next to a cache counting every MiB. Released at the end of init.
Also: the streaming banner read O_DIRECT off shard 0, which under the
small-first-shard layout is metadata only and too short to verify, so it could
claim a mode the shards carrying experts had not got. It now reports the weakest
of the readers.
* build: the engine version says 0.19.0, like the changelog does
The version is declared in CMakeLists.txt and reported by `--version` and by the
run-parameter preamble of every metrics CSV, so a committed benchmark file names
the engine that produced it. This release section landed while the number stayed
at 0.18.0, which would have stamped the wrong engine on every CSV this branch
produces, defeating the one purpose the string has.
* build(android): stage an explicit library list, and force GGML_OPENCL off
The staging step swept every libggml*.so it found anywhere in the build
tree into the app's jniLibs, and never cleaned jniLibs between builds.
The android build directory's cmake cache still carried GGML_OPENCL=ON
from a GPU experiment, so a libggml-opencl.so that no shipped
configuration ever loads was rebuilt on every run and rode along into
the v0.16.0 and v0.17.0 APKs.
The script now pins GGML_OPENCL=OFF at configure time (a cached ON from
an old experiment no longer survives), wipes jniLibs/*.so before
staging, and copies an explicit list of libraries. If a submodule bump
adds a library the CLI needs, the copy fails the build loudly instead of
a stray binary shipping silently.
* docs(agent): commits are grouped - one coherent change, no single-tweak commits
Compute buffers are reserved for the worst-case graph, and this engine had
always set n_ubatch = n_ctx so that any fitting prompt prefills in one pass.
That quietly ties RESIDENT MEMORY to the context rather than to the work:
320 MiB reserved at n_ctx 2048, falling to 80 MiB at 512, scaling exactly with
the context and nothing else.
On an engine whose entire problem is that the expert cache and the dense weights
compete for the same RAM, that is a real budget and it was invisible. Decode is
unaffected -- a decode graph is one token wide whatever this says -- so the cost
is prefill throughput, which processes a long prompt in more, smaller passes.
Default 0 keeps the previous behaviour exactly.
validate() rejects a negative value and one above n_ctx: reserving for a batch
that cannot arrive is the inverse of what the knob is for.
Found while chasing what looked like a catastrophic slowdown and turned out to
be this reservation pushing an already-tight system into reclaim at nearly
6 000 major faults per token, none of them the workload's fault. That happened
during GPU work, where the reservation was largest, but the coupling it exposed
is the engine's own and applies to every configuration. Salvaged here because
the GPU offload it was found alongside measured negative and was closed (#104).
Also fixed, all found alongside and unrelated to each other:
* A stray % in the --dense-weights help text was read by printf as a conversion
specifier, so that line printed garbage and read past the argument list.
* scripts/build-android.ps1 judged cmake by whether it wrote to stderr rather
than by its exit code, which broke it under Windows PowerShell 5.1 -- the NDK
toolchain prints progress to stderr and 5.1 turns that into a terminating
error under ErrorActionPreference = Stop.
* .gitignore now covers third_party/opencl-sdk/, 5.3 MB of Khronos headers and
an ICD loader that fetch-opencl-android.ps1 drops into the tree and nothing
was ignoring.
* feat(moe): --drop-cold-experts, spend quality only where it buys I/O
Turbo top-k drops the tail of a routing whether or not those experts were
already in RAM. A resident expert costs no flash read, so that trade pays
quality for nothing on the ~80% of decode routings that are cache hits.
This skips a routed expert only when it is a cache MISS and the router
weighted it below frac x (1/top-k). Replayed over the committed route
traces at frac 1.0, decode phase, that avoids 66% of flash reads for 9.5%
of the router's weight mass, where --n-expert-used 5 avoids 23% for a
comparable 10.6% -- about 3x the reads at the same quality cost.
Implementation. The decision needs the FINAL router weights, which arrive
several nodes after the topk where the streamer normally loads, so
load_layer() is deferred to the terminal node of the layer's weight chain.
Which node that is depends on the model's gating, so the hook learns it
from the graph rather than carrying an architecture table; if it fails to
arrive the hook forgets it and re-learns rather than re-betting. A dropped
slot has its weight zeroed and its expert id repointed at the routing's
top-weighted expert: an unread expert can sit in reserved-but-uncommitted
VM and mul_mat_id would touch it anyway, so the kernel is given memory
that is certainly resident and multiplies it by exactly zero.
Requires the LRU cache -- with --cache-mb 0 residency reads all-miss and
the policy would silently degenerate into an unconditional weight cut.
Prefill is excluded by default. The top expert is always pinned, so no
routing can be emptied at any threshold.
Gates: G8a/G8a' prove the deferral and the learned terminal node are
transparent (byte-identical output, zero drops, at a threshold below any
producible weight); G8b that full strength against a constantly-evicting
cache never reaches an unloaded slot; G8c that at top-k 1 dropping is a
no-op, pinning both the top-expert guarantee and the threshold tracking
the effective top-k.
Three existing metrics shift meaning under dropping and the docs now say
so: cache_hit_pct rises without the cache serving more (a dropped routing
is a miss that is never looked up), and token/layer_demand measure what
was staged rather than routed. prefetch.md's "cannot change output" is
scoped, limitations.md gains the non-reproducibility entry, and
benchmark-method.md warns that reversing the run order cannot distinguish
a moved drop rate from a contaminated cell.
Off by default in the CLI and in the app. The output is not reproducible
-- what gets dropped depends on what the cache held -- so it carries no
rows in the README tables, and switching it on by default waits on a
published on-device A/B rather than on the replay argument alone.
* feat(app): default cache-aware dropping to 75%, measured on device
Qwen3.6-35B-A3B (top-k 8 of 256), in-app, cache 3000, one variable
changed: 2.549 tok/s off, 3.938 at F=0.75 (+55%), 4.702 at F=1.0 (+84%),
with flash reads falling 248 -> 163 -> 48 GiB. Per-token bootstrap
intervals separate every pair except off vs 0.50, which overlaps -- at
half the uniform share the policy drops 2.7% of routings and buys
nothing, which doubles as a negative control that the machinery is free
when it does not fire.
Run order was 1.0, off, 0.5, 0.75, so the two fastest cells are the first
and the LAST; thermal drift would have made the last the worst. The
mechanism orders by threshold even though the run order does not.
The replay turned out conservative rather than optimistic. It is
documented as an upper bound because it cannot model the cache changing
in response to dropping: at F=0.75 it was accurate (37% predicted, 34%
measured), at F=1.0 it understated (66% predicted, 81% measured). Avoided
reads free cache capacity, which raises the hit rate, which leaves fewer
misses to drop.
75% rather than 100% is deliberate: it takes the larger part of the win
for half the discarded routings (14% against 28%). Quality is still
unquantified -- no perplexity number and no side-by-side exists -- so the
conservative end of a measured range is the defensible default. The CLI
stays off; the byte-identity gates need a deterministic default.
Also records cache_hit_pct rising 67.8 -> 90.7% as the documented
accounting artefact rather than the cache serving more, and majflt/token
as dominated by each run's starting memory state, not by the threshold.
This reverts 45a90a2.
The feature merged before the evidence for its shipping default did. The
replay numbers argue the shape of the trade is favourable, but no on-device
A/B is published in this repository, and the app default it landed with
(75%) changes model output for every user of the demo app -- and changes it
non-reproducibly, which no other setting in this engine does.
Nothing was found wrong with the code. This is a sequencing decision: the
work returns as a pull request, with the app default back to off, so the
measurement lands before the default does.
Reverted rather than force-pushed: main is public and this commit was
already pushed, so the history stays honest about what happened.
* feat(moe): --drop-cold-experts, spend quality only where it buys I/O
Turbo top-k drops the tail of a routing whether or not those experts were
already in RAM. A resident expert costs no flash read, so that trade pays
quality for nothing on the ~80% of decode routings that are cache hits.
This adds the cache-aware version: skip a routed expert only when it is a
cache MISS and the router weighted it below frac x (1/top-k). Replayed over
the committed route traces at frac 1.0, decode phase, that avoids 66% of
flash reads for 9.5% of the router's weight mass, against 59%/37% for
--n-expert-used 3 — about 3x the reads avoided at a comparable cost. The
threshold is a curve, not a switch: 0.75 trades 4.4% of the mass for 37% of
the reads, better than --n-expert-used 5 on both axes.
Implementation. The decision needs the FINAL router weights, which arrive
several nodes after the topk where the streamer normally loads, so with the
policy armed load_layer() is deferred to the terminal node of the layer's
weight chain. Which node that is depends on the model's gating, so the hook
learns it from the graph rather than carrying an architecture table; until
it is known a layer loads at its topk node undropped. A dropped slot has its
weight zeroed and its expert id repointed at the routing's top-weighted
expert: an expert we decline to read may sit in reserved-but-uncommitted VM
and mul_mat_id would still touch it, so the kernel is given memory that is
certainly resident and multiplies it by exactly zero. Survivors are rescaled
by default, since a systematically shrunk expert output perturbs the residual
stream more than the missing contribution does.
Prefill is excluded by default (cold cache, ~4x the weight mass discarded,
and compute-bound anyway). The largest weight in a routing is always at least
the uniform share, so frac <= 1 can never empty a layer; validate() enforces
the bound and the top expert is pinned regardless.
Gates: G8a proves the deferral and the learned terminal node are transparent
(a threshold below any producible weight leaves the output byte-identical),
G8b that full strength with the cache off never reaches an unloaded expert.
Unlike every other knob this one is state-dependent: what gets dropped
depends on what the cache held, so output is not reproducible across runs.
Off by default, not in the app's settings, and NOT yet measured on device —
the numbers above are a static replay and an upper bound. docs/expert-
dropping.md states what is owed before it is recommended anywhere.
* feat(app): expose cache-aware expert dropping in Settings
Speed / quality -> Drop cold experts, as a percentage of the uniform
share (off / 50 / 75 / 100). The engine takes a fraction; the app stores
integer rungs, so the setting divides by 100 on the way to the flag.
Disabled in mmap mode: the policy asks the expert source what is resident,
and there is no expert source without the streamer. Included in the session
signature, so changing it reopens the session rather than being ignored by
a process already loaded.
Off by default. This exists so the A/B can be run where the engine actually
ships -- through the app, not a pushed CLI binary.
* fix(moe): require the cache for dropping, and correct what it reports
Review of the first two commits found the policy could be armed in a
configuration where it is not cache-aware at all, and that two of the
numbers it reports were wrong.
- Require the LRU cache. With --cache-mb 0 query_residency answers
all-miss, so the policy silently degenerated into an unconditional
weight cut -- exactly what --n-expert-used already does, under a flag
claiming to consult residency. validate() now rejects it, as it already
did for --prefetch, and the app gates the setting on the same condition.
- Fix experts_routed. It was incremented inside apply_drop, so it counted
what the policy examined rather than what the router selected: layers
before the terminal weight node is learned, and every un-armed phase,
were missing from the denominator. The reported drop rate was a fraction
of the wrong thing.
- Re-learn instead of re-betting. If the node learned as terminal does not
arrive, the deferral now also forgets it, so the next graph loads at the
topk node while it re-learns. Deferring again on a stale guess would
repeat the fault every token against a graph that had moved.
- Point the gates at a real cache. G8a/G8b ran with the cache off, where
the shared-slot path has no reserved-but-uncommitted memory -- so the id
repointing, which is the design's whole safety argument, was never
exercised. They now run against a constantly-evicting budget. Adds G8a'
(asserts routings were examined and none dropped, so an inert-threshold
flake fails legibly) and G8c (at top-k 1 dropping is a proven no-op,
pinning both the top-expert guarantee and the threshold being taken
against the effective top-k).
Docs: three metrics change meaning under dropping and none of them said
so. A dropped routing is a miss that is never looked up, so cache_hit_pct
rises without the cache serving more, and token/layer_demand measure what
was staged rather than routed -- documented in telemetry.md, pressure.md
(size the cache with dropping off, then turn it on) and metrics.h.
prefetch.md's "cannot change output" is scoped: under dropping a correct
guess un-drops an expert. limitations.md gains the non-reproducibility
entry, benchmark-method.md the axis plus a warning that reversing the run
order cannot distinguish a moved drop rate from a contaminated cell, and
architecture.md/runtime.h no longer claim unconditional determinism.
Fixes two anchors the README rename broke, and a changelog sentence that
quoted the equal-I/O row while drawing the equal-quality conclusion.
App: Drop cold experts defaults to 75%. The default is a product decision
taken on the maintainer's device; no benchmark for it is published here,
and docs/expert-dropping.md says that plainly instead of implying a
measured figure. The CLI stays off by default -- the byte-identity gates
need a deterministic default.
Both of these tools produce figures quoted as evidence in docs/roadmap.md, and both had a path
that answered rather than errored when its input could not support an answer.
route-replay.py cost() defaulted a layer with no recorded expert_bytes to zero bytes. A free layer
is admitted without charge, never counts against the budget and is never evicted, so a trace whose
per-layer preamble is missing did not fail -- it printed a complete table in which every policy
scored the same near-perfect number. Reproduced on a gate-model trace with the preamble stripped:
the old code prints 96.9% across all six policies, the new code names the layer and exits. The
recorded traces behind the published curves all carry complete preambles, so no result in docs/
changes.
An unrecognised --policies name fell through every branch of victim() to the LRU default and was
tabulated under its own column header, claiming to compare a policy that never ran. argparse
choices= cannot express a comma-joined list, so the names are checked after the split.
bmoe-iobench hardcoded align = 4096 while alignment is the variable it exists to characterise; on
a 16 KiB-page device it measured the wrong one. It now calls pio::vm_page(), the engine's own
source. Its usage text also described --slice-kb as "bytes per read, default 4096" for a value
multiplied by 1024, leaving every printed figure open to being read off by 1024x.
Verified: replay accepted a real qwen3moe gate trace and reproduced the known LRU-cliff shape
(0% below one token cycle); rejected `--policies lur`; failed loudly on the stripped preamble.
bmoe-iobench rebuilt and swept, still negotiating O_DIRECT at the queried page size. Host gates
7/7, clang-format 18 clean.
No engine, CLI or app code is touched, so the Android version is unchanged.
Adds the two instruments the measurements were taken with, the data, and fixes to
the maintained docs that turned out to assert things that do not hold. No engine
code changes: the one candidate that was implemented is a measured regression and
stays on its branch.
tools/bmoe-iobench (new, BMOE_BUILD_TOOLS=OFF by default) sweeps flash read
bandwidth against lane count and read size. It drives bmoe::FileReader -- the
engine's own read path, so alignment, the bounce buffer and the O_DIRECT
verify/fallback are part of what is measured -- but links nothing else, since the
I/O layer has no llama.cpp dependency. It therefore cross-compiles in seconds and
cannot perturb the streamer. --compute-load adds CPU contention, because the
streamer reads while ggml's threads spin and an idle-CPU number is not the
condition it operates under.
scripts/route-replay.py (new, stdlib only) replays the committed route traces
through hypothetical cache policies at zero device cost. It reproduces the
recorded on-device hit rate to the decimal on all three captures and independently
predicts a historical budget-shrink measurement it was not calibrated against.
What they found, and what it invalidated:
- roadmap.md opened its read-bandwidth theme on the premise that effective
O_DIRECT bandwidth sits far below the drive's ceiling because routed slices are
scattered. Measured, reads saturate at 2 lanes and are flat above 256 KiB, so
scatter is cheap here. Runtime read coalescing is retired (0.6-4 % of a layer's
routed experts are id-adjacent); the expert-contiguous repack survives but needs
a different justification.
- prefetch.md opened on "MoE routing has strong temporal locality". It is 17.9 %
on gpt-oss, beaten there by a static hot list, and --prefetch 1 is a 2x
slowdown. The mechanism and its correctness argument stand; the bet is annotated
with what it returns and with the case still open (top-6 models).
- cache-sizing.md said a too-small budget makes the hit rate "collapse". It makes
it exactly 0.0 %, and the boundary is one token cycle -- not cache_min_mb, which
is a different quantity roughly n_layer smaller and protects today's models only
by coincidence. Reproduced on device at a budget the CLI accepts.
bench-data/2026-07-20-cache-replay/ carries the three notes, the curves and every
run CSV, indexed by a README stating the six verdicts.
The per-node compute trace pays ~3000 barriers per token, which serializes
the graph against the expert stream: on a model that streams heavily the
trace mostly measures its own serialization (Qwen3-30B: 9.4 s/token traced
vs 0.39 untraced), so its absolutes cannot be compared across models.
Layer granularity isolates only the first node of each layer (~n_layer
barriers per token). Operator coalescing and the async expert prefetch
survive, so the traced numbers stay close to an untraced run. Rows share
the per-node schema with op LAYER: name blk.<il> aggregates one layer''s
segment, pre the embedding lookup, post the last layer''s tail plus
final norm and LM head (closed by the session right after llama_decode,
since the tail has no successor boundary to observe it).
The granularity flows RunConfig -> SessionConfig -> RouterHook; the routing
nodes the streamer isolates anyway also close a segment, a barrier that
exists untraced too. decode-analyze.py detects the granularity and prints
the per-segment table.
Gates: all 6 pass (byte-identity qwen3moe + gemma4). Smoke-tested on the
tiny-moe models with streaming on.
The directory carried a name from an earlier project. It is a plain rename: the
constant, the three benchmark script defaults (all already env-overridable), and
the comments. Recorded measurements under docs/bench-data keep the old path —
they say where those runs actually read their model from.
Devices with models already pushed there:
adb shell mv /data/local/tmp/shardllm /data/local/tmp/bmoe
scripts/ had drifted into a pile of near-duplicate one-off drivers. Four PS1
files each carried a byte-identical Run-Cfg (cooldown, adb into bench-run.sh,
filter the perf lines, pull the artifacts); three Python tools each re-derived
the same readers, one of them saying so out loud ("Mirrors scripts/
route-analyze.py's reader").
bench-lib.ps1 now holds the shared driver plumbing, so a driver is its config
matrix and nothing else. trace_io.py holds the reading contract every artifact
shares: `key=value` tokens on `#` lines, rows below, unknown keys kept.
The device model paths were hardcoded in scripts that were otherwise
parameterised; they are now defaults behind -Qwen/-Gemma params (PS1) and
${VAR:-default} (sh), which is what they always were in spirit.
Retired framing, pruned from the tools that still run:
- bench-analyze.py listed `sg_ov` (speculative gating — removed from the
engine, PR #15) and labelled --cache-mb auto "adaptive cache" with a
`resizes` column. The governor that resized is gone, so that count is 0 for
every run bmoe-cli can produce, and a column that is structurally always 0
reads as a finding rather than a blank. Dropped; auto is "auto-sized".
- bench-pr23-c2000.ps1 grepped for `moe-spec-gate:`, a line the engine no
longer emits. Its --prefetch A/B is still valid, so it moves to the lib.
bench-matrix-rework.ps1 and bench-pr23-summary.py are NOT retrofitted: they only
re-derive tables already published under docs/bench-data, and their spec-gate
cells cannot run against a current build. They are marked ARCHIVED with why —
deleting them would leave published numbers with no visible derivation.
bench-lib.ps1 also documents what the copies silently carried: the cooldown is a
timer, and a timer does not guarantee a thermal baseline (see the contaminated
matrices where tok/s tracked run order). Fixing that needs device calibration;
saying so beats leaving the next reader to rediscover it.
Verified against committed data, not just by inspection:
- route-analyze --view hot/reuse/overlap/cache and decode-analyze: output
byte-identical before and after (docs/bench-data/2026-07-15-route-trace).
- bench-analyze on docs/bench-data/2026-07-13: every figure matches the
committed summary.md digit for digit; only labels changed and sg_ov is gone.
- All five PS1 files parse; the dot-sourced Invoke-BenchCfg builds the exact
same adb command string the copies did.
The per-token CSV reports compute as a residual (wall - io - mgmt), so everything the
engine does not itself clock is pooled into it: page faults, scheduler stalls, and the
matmuls. A residual cannot say which. That is the whole reason gpt-oss-120b reads as
"compute-bound" at 1.7 s/token while faulting 8.4k pages per token — the flash wait is
billed to compute because it happens under llama_decode.
--compute-trace measures it instead. Asking the eval callback to isolate a node makes
ggml compute exactly up to it and synchronize, so the wall delta between consecutive
boundaries is that node's real compute time; sampling major faults across the same
boundaries attributes the >RAM stall to the node that paid it. Still no llama.cpp patch:
this rides the public cb_eval ABI, whose ask/no-ask contract already specifies the
isolation. It costs a barrier per node and forbids operator coalescing, so it is a
diagnostic — a traced run is not a benchmark run, and only the shares are meaningful.
Unlike the other traces it does not need --moe-stream: it times the graph, so a dense
mmap baseline can be traced and compared against a streamed run.
--io-trace records one row per pread: latency, requested vs aligned bytes, lane, and the
(layer, expert, projection) it serves — values already computed at every enqueue site and
until now discarded. This is where the flash floor is: the aggregate 760 MiB/s sits far
below the drive's sequential ceiling because routed slices are scattered, and the trace
says whether that is per-read latency, request size, or lanes idling. It also measures
the adjacency the roadmap's read-coalescing and expert-contiguous-layout items assume.
Node classification stays out of the engine: which node is attention vs dense FFN vs
expert matmul is naming policy that varies by architecture, so the rows carry the raw op
and name and scripts/decode-analyze.py classifies. Verified on both gate models that the
generated text, cache hit rate and bytes read are identical with the traces on.
The six oldest bench scripts defaulted their output dir to one developer's
home. They all accept an override as arg 1 / -OutDir, so only the default was
unusable elsewhere: derive it from the script's own location instead, the way
build-host.sh already does. Same resolved path, no absolute literal.
Also ignore the root-level scratch a session leaves behind. Curated CSVs live
in docs/bench-data/, so a root-anchored rule cannot swallow them.
Packs a --route-trace session into one self-contained HTML page: the step x layer matrix
with the expert ids in its cells (miss/hit coloured, ids repeated from the previous step
marked), a miss-density heatmap of the whole run, hot experts per layer, the per-token
metrics with majflt highlighted, and the raw rows behind a filter.
Companion to route-analyze.py, which answers the same questions in a terminal. Stdlib
only, and the data is packed into the page, so it opens with no network and nothing
installed. Models are discovered from the session directory rather than hardcoded.
The page rounds weights to 4 decimals for size; the CSVs beside it keep full precision,
and it says so.
The trace is a long-format CSV; this reads it and answers the questions that shape
streaming speed: the step x layer matrix itself (with a '*' on each expert id the same
layer also routed on the previous step), routing concentration per layer (what a warm-up
should preload), reuse distance per (layer, expert) (LRU vs pinning, and how big a cache
buys what), overlap with earlier steps (whether temporal prefetch can predict), the
cumulative unique-expert curve against the cache budget, hit rate and prefetch
usefulness, routing entropy, and bytes per step and layer.
Stdlib only, like bench-analyze.py — nothing to install.
docs/telemetry.md gains the format (v1), the column semantics, and the two asymmetries
that would otherwise be misread: residency is per routing while expert_bytes is per read
(prefill dedups), and the last layer legitimately has a single prefill step because
llama.cpp gathers only the output token before its FFN.
Close the loop opened by the per-token warm-up analysis: add a "The fix" section
to warmup-analysis.md and fold the warm-up into adaptive-cache.md. Include the
before/after per-token CSVs (bench-data/2026-07-14-warmup/) and the reproducer
script (scripts/bench-warmonly.sh).
Also record the rejected alternative — reserving the dense bytes in the auto
budget — which lowered Gemma's hit rate (4000->2909 MiB, 83%->73%) without the
warm-up's benefit, and was dropped in its favour.
Document the first on-device run of a 58 GB / 120B MoE (gpt-oss-120b at
5.2x device RAM on the OnePlus 15R): a top-k x lanes x prefetch sweep and
an mmap baseline, streaming at up to 7.7x a plain mmap load (top-k 2).
- docs/benchmarks.md: full 12-cell matrix + 2 mmap rows, with honest
caveats (24-token probe, not steady state; the k=4 rows straddled USB
interruptions and are marked/excluded) and a quality note -- --no-think
drops gpt-oss's reasoning, so default top-4 answers 17x23 wrong while
k=2/3 answer right.
- docs/benchmark-method.md: how to benchmark harmony models (--no-think to
reach the answer, the /data path for O_DIRECT, the reasoning trade-off).
- README + CHANGELOG: headline (a 120B on a phone) and the gpt-oss section.
- scripts/gptoss-matrix.sh + gptoss-mmap.sh: the exact drivers.
- docs/bench-data/2026-07-14/: raw summary log + the two mmap per-token CSVs.
Speculative gating was the only feature that broke the ports-and-adapters
seam: it made router_hook reach into architecture-specific router math
(RouterPre), spawned a second thread inside the eval-callback bridge, and
inlined predictor logic into the streaming hot path. It was experimental and
default-off, and never paid its way in steady-state decode on device.
Removing it collapses MoeRecipe back to {arch, expert suffixes} — the header's
stated design intent — and router_hook back to capture -> gather -> load_layer
plus temporal prefetch. The shared speculative-prefetch queue in
expert_stream_source (used by --prefetch) is untouched.
- delete core/src/moe/spec_dot.{h,cpp} and docs/spec-gating.md
- strip RouterPre + router-node fields from recipe.h and every registry row
- drop spec_gate / spec_recall_* config, the run() wiring, and the
moe_spec_recall_pct / moe_spec_auto_off summary fields (+ CLI flags,
BMOE_SPEC_GATE env, BMOE_DONE + CSV columns, moe-spec-gate print)
- remove the Android "Speculative gating" toggle and specGate setting
- drop gates G6a-d; G1-G5 and S1-S3 still prove streamed == resident
- clean the reusable bench scripts of --spec-gate; keep docs/bench-data as an
archive of the historical measurements
Host byte-identity gates pass for qwen3moe and gemma4.
Pinning armv8.2-a+dotprod+i8mm+fp16 built i8mm (armv8.6, matmul-int8) instructions into the
CLI; on a device without them — e.g. the Snapdragon 865 (OnePlus 8 Pro), which has dotprod and
fp16 but not i8mm — the prefill GEMM hits an illegal instruction and the process dies with
SIGILL. dotprod + fp16 (armv8.2, ~2018+) is present on every SoC that can realistically run a
>RAM MoE model, so a single dotprod baseline is device-agnostic across the viable range and keeps
every kernel that matters for Q4_K experts. (ggml's per-device runtime dispatch is not usable
here: GGML_CPU_ALL_VARIANTS + GGML_BACKEND_DL split the CPU backend into dlopen'd variant .so's,
which drops the fork's statically-linked expert-ready overlap hook.)
Extend the per-run CSV summary trailer with the fields the rework added —
cache budget and resize count (--cache-mb auto), speculative-gating recall,
useful-hit rate, speculative bytes read, and the auto-off flag — so a single
CSV carries the full benchmark record. bench-analyze.py parses them (keys are
order-independent) and emits a third 'adaptive cache & speculative gating'
table alongside throughput and device-pressure. Add bench-matrix-rework.ps1:
the 2026-07-13 config matrix (fixed reference, adaptive ± ceiling, spec-gate)
for both models, reusing bench-run.sh and the tag naming the analyzer expects.
Measured on the OnePlus 15R on top of each model's best config. Both features are byte-correct
and run end-to-end (recall/useful metrics prove the path), but neither improves steady-state
256-token throughput: temporal prefetch is neutral (overlap already hides read latency, no idle
flash bandwidth), and speculative gating is net-harmful — it predicts well (88% recall on Qwen)
but reads 24-44 GB speculatively, saturating the flash and halving throughput. The real lever is
cache size, and Qwen is already compute-bound at ~5.1 tok/s. See docs/bench-data/2026-07-12-pr23.
Extends bench-run.sh to capture the moe-prefetch/moe-spec-gate lines; adds the A/B drivers and a
summary parser.
Add docs/prefetch.md (the temporal-locality bet, the correctness argument, and telemetry),
document the moe-prefetch summary line, record the feature in the changelog, extend
bench-analyze with prefetch config rows for a device A/B, and expose a prefetch-depth row in the
Android settings (session argv, so changing it reopens the session). Kotlin compiles.
vm-commit, eviction and LRU bookkeeping were hidden inside the compute residual, making the
first tokens after prefill read as pure matmul when the real cost is cache churn. Time that
staging work into mgmt_ns_ and report it per token (mgmt_ms), in the CSV, and in the
moe-stream summary (cache mgmt). compute_ms is now documented as a residual
(wall - io - mgmt, or wall - stall - mgmt under overlap), not a measured quantity. Bytes
served are unchanged; gates G1-G4 pass byte-identically on qwen3moe and gemma4.
seam.md gains a section on the hook, its call site and threading contract, and an
explicit sunset condition. README and architecture.md qualify the no-fork claim:
the streaming seam is fork-free, --overlap is the one sunsetted exception. Bench
scripts gain the overlap configs and the stall column.
256-token steady-state runs on the OnePlus 15R for two MoE families across the
cache/lane matrix. Adds benchmarks.md (the measured tables), the committed
drivers (bench-run.sh, bench-matrix.ps1, bench-analyze.py), and per-row flag +
model-file provenance so every number reproduces from the logs.
Records the decode compute+I/O split: at a 4 GiB cache decode is compute-bound
(I/O ~0.1 s/tok), so the streaming ceiling is the SoC's in-RAM speed (~7 tok/s
Qwen, ~5.3 Gemma). Ignores .bench/ and local *.log scratch.
The template-driven chat refactor linked llama.cpp's common (llama-common),
adding a runtime dependency on libllama-common.so. Both staging paths — the
local build-android.ps1 filter and the CI jniLibs step — matched only
libggml* and libllama.so exactly, so the new lib was never bundled and the
app crashed at first inference with 'library libllama-common.so not found'.
Widen both filters to libllama* to catch it.
Extend make-tiny-moe.py with --arch {qwen3moe,gemma4}. The gemma4 fixture
emits the fused ffn_gate_up_exps + ffn_down_exps expert tensors, a resident
shared expert, an interleaved dense layer and a mixed sliding-window / full
attention pattern, so the gates exercise the two-projection streaming path.
The qwen3moe output is byte-identical to before, so the existing gate is
unchanged. tests/CMakeLists.txt now wires one generate-fixture + gate per
architecture.
bmoe-cli links the c++_shared STL, so its runtime shared library must ship in
the APK next to it. Copy it from the NDK sysroot during staging so a clean
checkout produces a runnable app instead of one that fails at exec.