The combination produces wrong bytes with no error reported. Measured against a real model: same revision, same gguf, same flags, correct text on Android and garbage on Windows. Everything on this side is proven correct (requests, destination arithmetic, commit ranges, alignment, the partial-read retry), so the fault is below us; the difference is that in place the DMA target is freshly committed shared memory a compute thread is concurrently reading, where the bounce had a private buffer and a CPU copy ordered against the ready flag. Serial is correct on Windows, so only the overlapped path falls back, and it says so rather than degrading in silence. Tracked in issue #149.
G8 and G9 were already taken, by cache-aware expert dropping and by the
prediction probe, so this branch's two gates answered to labels that name a
different feature in the same binary's output. A failing "G9" could not be
read without opening the file. They become G11 and G12.
The index at the top of the file had also stopped describing the file below
it: it predates both the drop gates and these, so the map named neither. It
now lists them, and says that gate numbers are allocated once and never
reused, which is the rule this collision broke.
The +20% in this branch was measured with short single-turn CLI runs, which is
exactly the regime where a page leak cannot accumulate. In a long app session
the same build measured slower than the baseline until the boundary pages were
released. Both the win and the regression are now recorded, along with why the
earlier measurement could not have caught it, and neither number is presented as
settled until a long session re-measures it.
The zero-copy placement moves each layer buffer off the page boundary on
purpose, which makes the first and last page of EVERY expert slice shared with
the neighbouring expert. Eviction releases only the pages an entry fully
covers — correct, and previously exact, because the page-aligned placement had
no partial pages at all. With the shift it leaks two pages per eviction: the
cache accounts the bytes as freed while the kernel still holds them.
Measured in the app, where a session evicts thousands of entries and lives long
enough for it to accumulate: the process's anonymous footprint grew ~40 MiB a
turn and went to zram, and the decode paid it back as major faults — 1-7 per
token without the flag, 66 on the first turn with it and 332 on the second,
rising turn over turn, which is the shape of a leak rather than a cost. Net
effect in the app: 5.02-5.58 tok/s without, 4.41-4.57 with. Reversing the run
order changed nothing, so it was not warm-up or thermal drift.
A boundary page is now released as well, once the neighbour sharing it is gone —
`cvalid_` covers demand reads (an entry is valid from the moment it is staged)
and `spec_remaining_` covers speculative ones. Both are read without the lanes'
mutex, safe in the one direction that matters: staging happens on this thread so
nothing raises them concurrently, and a lane can only lower spec_remaining_
toward zero, so a stale read is stale-high and merely defers a release to the
next eviction. The first and last expert of a buffer border its own padding,
which nobody else owns.
Why the earlier device runs missed it: they were short, single-turn, in a fresh
process — 24 to 40 tokens, where the leak has no time to accumulate. The app's
long session is the only place it shows, which is the trap this project already
recorded once: a single-shot bench cannot see a reclaim problem.
Gates 7/7 on the host. The claim this branch carries still needs re-measuring in
the app, on a long session, before its headline number means anything.
Deciding the zero-copy path from the destination's alignment alone was wrong,
and it broke the dense loader: it reads through its own FileReader into
separately sized per-tensor buffers, and the aligned window is longer than the
tensor it was asked for. The most ordinary case of all triggers it — a
page-aligned tensor offset read into a page-aligned buffer matches the remainder
test with shift 0, then overruns the buffer by up to a page. On device with
--dense-weights ahwb (the app's default, which the earlier CLI runs did not use)
this surfaced as `pread failed` and a failed session open.
read() now takes allow_in_place, which is a promise: the pages either side of the
request, out to the enclosing alignment boundaries, are the caller's to clobber.
Only the expert cache can make it, and only for the buffers actually placed for
it — recorded per buffer rather than re-derived, because a remainder of zero is a
valid placement and must not be confused with a declined one. That distinction
was a second latent bug: with the remainder zero the extra page was not reserved,
so the last expert's window would have run past the end of the reservation.
G9 gates the combination that failed (zero-copy + the anon dense loader). It
bites on any model, unlike G8/G4e, whose expert-side placement goes inert on a
tiny test model. Verified on device with the exact failing configuration: 120/120
buffers placed, generation byte-identical with the flag on and off, 3.39 -> 4.00
tok/s in the same pair. Gates 7/7.
The engine flag is useless in the app without a way to turn it on, and the
device A/B this change still owes has to be run through the app as well as the
CLI. The switch sits under Direct I/O and greys out without it, since O_DIRECT
is what forces the bounce copy in the first place.
Worded as a pure speed knob, because it is one: unlike expert dropping or
route-ahead it is lossless by construction — the same bytes land in the same
places, and the generated text is byte-identical with it on and off.
It joins the session signature, so flipping it reopens the session rather than
silently carrying the old setting into the next generation.
O_DIRECT wants the file offset, the length and the buffer address aligned.
Expert slices are a whole number of pages long, but a gguf's tensor data does
not start on a page boundary: general.alignment is 32 by default, so on the
measured model family every expert offset sits a constant 1152 bytes past one.
The reader therefore pulled the enclosing aligned window into a per-lane bounce
buffer and memcpy'd the payload back by that remainder — over 200 MiB a token
of pure shift correction, on a device whose decode is already memory-bound.
That shift is only needed because the destination did not share the file's
remainder. It can: this engine reserves the per-layer buffers itself and rebinds
tensor->data onto them. With --odirect-zero-copy each buffer is placed at an
address carrying its own tensor's remainder; every expert inherits it because
the per-expert stride is a multiple of the page size, so the aligned window maps
onto the buffer in place and the read needs no copy anywhere.
The window overhangs its neighbours' slices. That is safe by construction, not
by luck: under this placement buffer and file differ by a constant offset, so
the overhung bytes receive their own correct file contents, and they land in
pages the entry's commit already covers exactly — eviction never releases a page
shared with a neighbour, which was already true before this change.
Measured on device (Qwen3.6-35B-A3B Q4_0, cache 2000 MiB, overlap, 4 lanes,
interleaved cells so thermal drift hits both): +20% tok/s, 3.65 -> 4.40 median,
compute residual 0.180 -> 0.143 s/token, identical bytes read (230.77 MiB/token
in every cell) and byte-identical generated text with the flag on and off. The
phone was warm and its baseline drifted 3.90 -> 3.40 across the run, so the
ratio is the claim, not the absolutes; a cool-device confirmation is owed, which
is why the flag ships off.
This is not specific to any decode feature — every expert read goes through this
path, and so does the dense loader's.
The host gates cannot prove the path: a tiny test model's per-expert stride is
not a multiple of the page size, so the placement declines and G8/G4e show only
that the flag is harmless. Rather than let that read as a pass, the engine
reports which happened (odirect-zero-copy ON|INERT — n/m layer buffers placed).
The device run above is the real proof. Gates 7/7.
* feat(app): sharded model downloads — DeepSeek V4 Flash in the catalog, gpt-oss one-tap
A catalog entry can now list shard files. They download sequentially through one
WorkManager chain (per-file HTTP Range resume, one aggregate progress bar, free
space checked once against the whole remaining set), the model picker offers only
the first shard — the file the engine opens — and deleting a sharded entry deletes
the whole set, so no 40 GB tail is ever orphaned.
DeepSeek V4 Flash UD-IQ2_M (~91 GB, three shards) joins the catalog, and
gpt-oss-120b turns from a "merge it on a PC" manual recipe into a one-tap
download; a merged single file from an earlier release still counts as on-device.
App version 0.19.0 (versionCode 34).
* docs(readme): DeepSeek V4 Flash becomes the flagship claim, video slot staged
* docs(readme): standardize benchmark tables (slowest to fastest, one label scheme), drop dashes
* fix(app): a split model's first shard is a MoE model too
The picker's MoE filter looked for an expert tensor inside the file it was
handed. A split gguf's first shard carries the metadata and, in the layout
large quants ship in, almost no tensors: DeepSeek V4 Flash was therefore
classified dense and never appeared in the model dropdown, with all 91 GB
sitting on the device. The header walk now also accepts the metadata key
<arch>.expert_count, which is the definitive MoE signal and always lives in
the first shard; the tensor-name check stays as the fallback.
* fix(app): bound the session's ubatch so compute buffers stop eating the model's RAM
A session opened at ctx 4096 with no --ubatch reserves compute buffers for the
whole width, though decode only ever computes one token. That reservation is
memory the expert cache and the dense weights do not get, and the CLI has
measured it as an 18% decode lever for a while; the app never passed the flag,
so every in-app run since gave it away.
On DeepSeek V4 it is not 18% but the whole result: the same configuration read
14.58 s/token in-app against 2.22 s over adb, with identical flash I/O (2.74 vs
2.89 s) and 3.6x the major faults. The 13.9 s of 'compute' were page faults, the
process swapping while it worked. Prefill pays instead, and barely: chunking it
costs ~7.7x the flash reads for ~6% of prefill wall time.
* feat(app): context is a setting, not a constant
The session opened at a fixed 4096 tokens. That is also memory — the KV cache is
sized for it once at open — so on a model that already fills RAM it competes with
the weights, and there was no way to trade conversation length for room without a
rebuild. It joins the other tunables (default unchanged), and the ubatch is
clamped to it so a graph is never reserved wider than the context. The service
reads the running session's context from its own argv for the 'ctx used/total'
readout, so the number describes the process rather than the current setting.
Measured on DeepSeek V4: the KV is 44 MiB at 512 and ~270 MiB at 4096, small
thanks to the compressed attention, so on that model the setting is not the lever
its size suggests. It is on models with ordinary attention.
* fix(app): review pass on sharded downloads
Four defects, all from the same blind spot: code that asked whether a filename
belongs to a catalog entry compared it against the entry's own name, which for a
sharded model is one file out of several.
- A sharded entry never reached ON_DEVICE. The present-files set was recomputed
only when the SELECTABLE model list changed, but shards 2..N are hidden from it
by design, so finishing a 41 GB shard left it byte-identical: the row offered
Download for a model already fully downloaded, and pressing it did nothing
until the app was restarted. Keyed on the in-flight names as well, which change
exactly when a shard starts or finishes.
- Shards also rendered as pasted-URL downloads, whose Cancel deleted the .part
the worker was still writing while cancelling nothing (the chain is registered
under the entry name). The transfer then ran on an unlinked file for tens of GB
before failing to finalize.
- A sharded gpt-oss also appeared under Imported models, with a Delete that
removed shard 1 and orphaned the rest — the exact failure the delete dialog
exists to prevent.
- That dialog replaced the entry name with the shard names instead of adding
them, so a gpt-oss merged by an earlier release became undeletable.
ModelCatalog.fileNamesOf/isCatalogFile is now the single answer to 'does this
file belong to an entry', and all four sites go through it. Also: a queued shard
reports zero bytes, so aggregate progress now falls back to its .part length
rather than appearing to lose ground on a resumed 50 GB transfer.
* feat(moe): stream split multi-shard ggufs natively + DeepSeek V4 Flash recipe
Hugging Face rejects single files above 50 GB, so every large model ships as
-00001-of-0000N.gguf shards; until now the streamer assumed one file, forcing
a merge with double the disk. gguf_offsets now fans the first shard out to the
whole set and resolves every tensor to (shard, offset); the expert streamer
and the dense loader open one positioned reader per shard and route each read
by the tensor's shard index. Pass the first shard, exactly as llama.cpp takes
it; a missing sibling fails the load with the shard named.
Add the deepseek4 recipe row: V3.2-style routing (256 routed experts, a
per-expert bias like lfm2moe, an always-on shared expert that stays resident)
over the standard split expert suffixes. The V4 compressed-attention machinery
is dense-side llama.cpp code, invisible to the streaming seam.
The byte-identity gates gain a 4-shard qwen3moe fixture (metadata-only first
shard, the layout large quants actually use); make-tiny-moe.py learns
--split-max-tensors. All gates pass, split included.
* fix(moe): cache auto must budget for the anon dense conversion
The auto budget read MemAvailable while the dense weights were still reclaimable
page cache, then dense-weights=anon converted them into buffers the kernel cannot
take back: the same bytes planned twice. Latent since the anon policy shipped
(dense sets were 2-3 GiB and explicit budgets were the benched path); DeepSeek V4
Flash's 6.5 GiB dense set turned it into a device-taking overcommit on first load.
The budget now deducts the pending conversion and says so in the log.
* fix(moe): review pass on the multi-shard path
Three defects the split rewrite introduced, none of which the gates could see:
- The shard index rode in an int8_t, so a model past 127 shards wrapped to a
negative index into the reader vector. The bounds check could never catch it:
it validated the untruncated value. Widened to int16_t, which covers the whole
-%05d-of-%05d filename space.
- DenseWeights::warm() reused one flag as both the inner loop condition and the
partial-warm report, so the first shard that failed to open silently skipped
the warm-up of every later shard. Per-shard condition, sticky report.
- The dense readers stayed allocated for the session after read_anonymous had
copied and rebound every tensor: fds and a per-lane bounce buffer per shard,
sitting next to a cache counting every MiB. Released at the end of init.
Also: the streaming banner read O_DIRECT off shard 0, which under the
small-first-shard layout is metadata only and too short to verify, so it could
claim a mode the shards carrying experts had not got. It now reports the weakest
of the readers.
* build: the engine version says 0.19.0, like the changelog does
The version is declared in CMakeLists.txt and reported by `--version` and by the
run-parameter preamble of every metrics CSV, so a committed benchmark file names
the engine that produced it. This release section landed while the number stayed
at 0.18.0, which would have stamped the wrong engine on every CSV this branch
produces, defeating the one purpose the string has.
* fix(app): show the whole run configuration a metrics CSV carries (#136)
The engine has written every resolved knob into the `# bmoe_metrics v2` preamble for
several releases; the app displayed hand-picked subsets of it. The header card of an
opened file listed eight fields, and the compare view rendered a 17-key whitelist that
had drifted behind the metrics sink. Expert dropping, predictive prefetch and its
speculation width, the sampling parameters and the engine build were all in the file and
none of them reachable — a saved run could not say whether it dropped experts, let alone
at what fraction, and an A/B differing only by one of those levers rendered two
configuration cards that looked identical.
Both views now go through one renderer over the whole preamble: the curated keys in
order, then every remaining key under its own name. The order list stops being a filter,
so a knob added to the sink becomes visible without an app change — which is the actual
fix, the missing keys were only the symptom. The single-file card keeps its glance line
and gains the full table behind a tap; `drop` and `predict` also join the short run label
the compare legends show.
Version bumped to 0.18.1 (versionCode 33).
* fix(app): state the whole run configuration on the main screen too (#136)
The reminder line under the prompt was the same hand-picked subset as the metrics
views: cache, lanes, overlap, threads, top-k. It left out the dense-weight policy,
both prefetches and cache-aware expert dropping - which defaults to 75%, so the
out-of-the-box configuration changed the answers and the screen said nothing.
The short line now carries what makes a run a different KIND of run, gated exactly
as sessionArgv gates the flags themselves so it cannot name a lever the CLI is
never told about. The whole configuration sits one tap below it, read back from
that same argv rather than from a second list kept by hand: a knob added to
sessionArgv shows up on its own, which is the property this display kept losing.
* fix(app): explain the configuration the metrics views now show (#136)
Surfacing the whole preamble is half an answer while the keys it names go
unexplained: several are unguessable from the key alone. The glossary behind ?
gains a second section describing every configuration key, worded from where each
knob is defined, and is now reachable from Compare too. It also finally describes
loop_overhead_ms, a column the engine has written since 0.17.0.
predict_spec_max renders as inert when predictive prefetch was off. The engine
records its own default (2) there - the one field of that block session.cpp does
not neutralise - so a file claimed two speculated misses per layer for a run that
speculated nothing. Only the record was wrong: with the feature off the value is
never read, so no run behaved differently than reported.
* docs(readme): reposition around large MoE models in general, add logo and TOC
The front page led with gpt-oss-120b and read as a single-stunt repo. It now
leads with what the engine is for: running MoE models past RAM on phones and
PCs, on llama.cpp's public API, with every llama.cpp quantization coming for
free. The 120B stays as the most extreme proof, not the pitch.
Structure follows the usual professional layout: logo (chip-buddy, light and
dark variants under docs/assets/logo/), badges, table of contents, a "Why this
exists" section that covers all three regimes (far past RAM, just past RAM,
and models that barely fit, which streaming keeps inside a chosen budget),
and Features grouped the way the app's Settings groups them. Benchmark tables
and the evidence sections are unchanged in substance; the test device is
phrased generically.
* docs(agents): make AGENTS.md the canonical agent guide, CLAUDE.md points to it
The agent guide lived in CLAUDE.md with AGENTS.md as a stub pointing at it,
which is backwards: AGENTS.md is the cross-tool convention (Codex, Copilot,
Cursor and others read it), CLAUDE.md is one tool's name for the same file.
The full guide now lives in AGENTS.md and CLAUDE.md is a one-line import.
While moving it, the guide gains the release rules that were only tribal
knowledge: release APKs come from the release-apk workflow, never a local
build; every released feature bumps versionCode/versionName in the same PR;
release titles are the bare version; on-device validation before a release.
The style section now names the CI clang-format version (18), since a
mismatched local formatter passes locally and fails the check.
* docs(readme): state per-platform host status honestly
Linux is CI-verified, Windows is where the desktop bench ran (but needs CMake
directly, not the bash script, and MSVC's Release output path), macOS builds
from the same sources but is unvalidated and has no O_DIRECT.
A cold layer's batch became visible to the I/O lanes only after every
miss took its page commits - up to three vm_commit syscalls per cold
expert of bookkeeping sitting in front of the first byte of I/O, which
is the latency-to-first-slice the sidecar refutation (PR #90) identified
as the binding constraint. (#118)
With the flag on, only the first present projection - the one
mul_mat_id blocks on first - is committed up front; its jobs publish
and wake the lanes immediately, and the remaining projections are
committed and appended while the lanes already read.
The drain protocol grew the one thing this needs:
- io_drain copies each job out under the lock, so jobs_ growing (and
possibly reallocating) mid-batch cannot leave a worker holding a
dangling reference;
- the worker wait predicate admits next_idx_ < batch_njobs_, so a
worker that drained wave one and left comes back for a batch that
grew in the SAME generation - the gen comparison alone never would;
- a wave-two commit failure goes fatal and wakes the ready waiters,
because wave one already published flags this batch will never flip.
Batch completion cannot fire between the waves: the only thread that
waits on done_cnt_ == batch_njobs_ is the eval thread, and it is the
one appending wave two.
Overlap + LRU cache only (validate() enforces both); recorded in the
CSV preamble as io_two_wave. Default off: the win is bounded by the
commit cost per cold layer, and the failure mode of a drain-protocol
bug is a hang the host cannot reproduce - so a new gate (G4d) holds
two-wave output byte-identical to serial streaming, and the flag stays
off until the on-device A/B (#120) says the win is real.
Every per-token line repeated the whole answer and reasoning so far, so
a generation of n tokens wrote, JSON-escaped and made the app parse
O(n^2) bytes - megabytes of pipe traffic to deliver a few kilobytes of
text on a reasoning model (#119).
The line now carries delta_reasoning/delta_text - the tail since the
previous line - and the reader appends. A pure append-only protocol
cannot express the one thing common_chat_parse does retroactively:
when a closing think tag arrives, text already reported as answer
becomes reasoning. That case falls back to a full snapshot with
"reset":1, and the reader replaces instead of appending.
Both emitters (one-shot --progress and --session) share the single
format string, so they changed together; the app's TelemetryParser
accumulates in StringBuilders (appending to a String re-copied the
whole answer per token) and resets them with the generation. The full
final text still travels in BMOE_DONE, untouched.
The engine-side re-parse per token remains (common_chat_parse cannot
resume); this removes the pipe, escape and app-parse cost, which is
what loop_overhead_ms can now see. On-device numbers are part of the
#120 A/B.
Closes#119.
Three findings from the 2026-07-28 audit's leftover list (#123), all on
paths that run for every routed expert of every layer of every token:
- on_expert_ready looked the expert tensor up in an unordered_map; the
map is static after init, so it is now a flat sorted array probed with
a binary search, and the pre-block spin of 2048 sched_yield syscalls
(up to a millisecond of scheduler churn per genuinely slow slice) is
now 256 single-instruction pauses (isb on arm64, pause on x86) before
the thread registers as a waiter.
- a cache hit in load_layer_async paid an LRU unlink+push in
touch_entry that the token-major promote loop overwrote
unconditionally two steps later; touch_entry now takes promote=false
from the overlap path. The serial path, which has no trailing loop
for first touches, keeps promoting. Final LRU order is unchanged.
- the gguf header was parsed once for the tensor offsets and once for
the model info — two full KV walks of a multi-GB file's header.
read_gguf_meta parses once and serves both; the session's lazy
accessor now hands the same parse to the top-k override, the route
trace, the run info and the streamer.
Items 4-6 of #123 stay parked: the no-consumer metrics guard only
matters to a library embedder, the O_DIRECT misreport only to macOS,
and the llama.cpp context defaults want a bench, not a code change.
Byte-identity gates pass; LRU semantics and readiness protocol are
untouched.
Release assets were built on a developer machine and uploaded by hand,
which is how a stale cmake cache shipped an OpenCL backend into two
releases. A release-apk workflow now runs when a release is published:
clean checkout of the tag, NDK build of the CLI with the same flags and
explicit staging list as scripts/build-android.ps1, APK build signed
with the stable release key from repository secrets, a content check
that fails on any stray library, and the assets attached to the
release. workflow_dispatch allows rebuilding assets for an old tag.
Both build types now sign with the stable key when available, so the
debug APK also updates in place instead of demanding an uninstall that
wipes downloaded models.
* build(android): stage an explicit library list, and force GGML_OPENCL off
The staging step swept every libggml*.so it found anywhere in the build
tree into the app's jniLibs, and never cleaned jniLibs between builds.
The android build directory's cmake cache still carried GGML_OPENCL=ON
from a GPU experiment, so a libggml-opencl.so that no shipped
configuration ever loads was rebuilt on every run and rode along into
the v0.16.0 and v0.17.0 APKs.
The script now pins GGML_OPENCL=OFF at configure time (a cached ON from
an old experiment no longer survives), wipes jniLibs/*.so before
staging, and copies an explicit list of libraries. If a submodule bump
adds a library the CLI needs, the copy fails the build loudly instead of
a stray binary shipping silently.
* docs(agent): commits are grouped - one coherent change, no single-tweak commits
build_moe_ffn emits a chain of router-weight nodes per layer — ffn_moe_weights,
then optionally _softmax or _norm, then optionally _scaled — each a refinement
of the last. The engine asked the eval callback to isolate all of them, because
last-wins is how it stays architecture-independent: it learns which node ends
the chain instead of tabulating the gating per model.
But that learning finishes on the first graph. From then on only the terminal
node is read: the drop policy decides there, and the route trace's last-wins
gather lands there. The other two or three asks per layer were still made, and
each one is a graph split plus a full compute-thread synchronization, on a
tensor of a few floats — per layer, per token, for the whole run.
Ask for the rest of the chain only while term_variant_[il] is unlearned. That
covers the first graph of a run and any graph after close_drop_layer forgets a
terminal that stopped appearing, so re-widening is automatic and the recovery
path for a graph that moved is unchanged: the deferral still fails to be
honoured, the layer still loads undropped, the terminal is still forgotten.
The drop policy is on by default in the app, so this cost sat inside every
--drop-cold-experts measurement taken so far.
Gates 7/7, run three times, including the two that exercise the deferral (G8a,
G8c). Route trace verified by hand on Qwen3-30B: 1920/1920 rows still carry a
non-zero router weight, range 0.0011-0.9862 — the terminal-only gather sees
exactly what the full-chain one did.
No throughput claim: still no device. Folded into the owed A/B (#120).
Carve the accumulated Unreleased section into a dated release, bump the app's
versionCode and versionName, and say in the README what the release changed
about how a run reports itself.
0.17.0 is the engine/core audit: three latent correctness bugs, four pieces of
per-token work that ran whether or not anything consumed them, and the two
measurement gaps that made the rest hard to judge — a metrics file that did not
record half the configuration it was produced under, and a tok/s that excluded
everything between the decodes.
The audit flagged the bounce buffer as the largest avoidable CPU cost on the
read path: every streamed byte is pulled into a per-lane staging buffer and
then memcpy'd to its cache slot, so the CPU touches 100-200 MB per token twice,
on the I/O lanes, alongside a decode that is already compute-bound.
Skipping the copy needs the read to be block-aligned at offset, length AND
destination. Measured against the three shipped models: bytes-per-expert IS
4096-aligned, but the tensor's file offset is not — 512, 1152 and 2272 bytes
past a block on Qwen3-30B, Qwen3.6-35B and gpt-oss-120b respectively, because
gguf aligns tensor data to general.alignment (32), not to a device block. So an
O_DIRECT window always rounds outward and always spills one block into the
neighbouring experts' bytes.
That spill writes the file's own bytes into buffer positions that hold exactly
those bytes, so it is content-identical. It is still not safe: the pages belong
to other cache entries, which may have been evicted, and touching them faults a
page back with nothing accounting for it — the residency budget stops matching
what is resident, and that budget is what keeps a >RAM model running. Offsetting
each buffer's base by file_off % align, the trick that makes the destination
inherit the file's misalignment, shrinks the spill but does not remove it.
An implementation of the safe subset — take the direct path only when all three
are already aligned — was written and thrown away: it is dead code on every
model in the catalog, and a branch on the hot read path that never fires, with a
startup line that would always report the fallback, is worse than the honest
note.
No code change; the finding is the deliverable.
Two gaps, both about a benchmark file being unable to explain itself.
The CSV preamble had eighteen keys and had fallen behind several releases.
Missing: n_ubatch, which sets the compute-buffer reservation and therefore
moves the very memory columns printed underneath it; predict_log,
predict_spec_max and prefetch_sync, so a probed run was indistinguishable from
a benchmark run the docs explicitly say it is not; drop_renorm and
drop_prefill, which change how much mass dropping discards; cache_floor_mb, the
input behind an auto-sized cache_mb; load_all, whose read_bytes mean something
else entirely; every sampling parameter, so a stochastic run read as a greedy
one; and compute_trace_layers. All are recorded now, under a bumped
"# bmoe_metrics v2" banner.
`think` stays out on purpose: it is a property of a request, not of a session,
so one value in a session-wide preamble would be wrong for every turn that
asked for the other.
The file also now names the build that wrote it. There was no version string
anywhere in the engine — the CMake project had none — so a committed CSV could
only be dated by the commit that copied it in. Added as a project VERSION, a
BMOE_VERSION define, bmoe::version(), `bmoe-cli --version`, and engine= in the
preamble. It sits on its own line so the model= line still BEGINS with model=,
which is how the app's CSV reader finds a run's name; the app's lookup is made
order-independent too, so the next key to be appended cannot break it again.
Second gap: wall_ms brackets llama_decode and nothing else. That is what makes
compute_ms a clean residual, and it also means sampling, detokenization,
rendering and the sink writes are outside wall_ms, outside gen_seconds and
outside the reported tok/s. Work moved into or out of that region was
unmeasurable by the number the project optimizes. loop_overhead_ms now reports
it per token, and loop_overhead_s/tok in the summary closes the accounting with
the tail after the last token that no row can carry.
docs/telemetry.md documents the preamble — it specified every other # block but
not this one — the new column, and the stall_ms divisor: stall is a per-thread
mean, so a stall that is not simultaneous across threads is under-stated and
compute_ms absorbs the difference.
Gates 7/7; preamble, column and --version verified against the tiny gate model.
Both readers checked: scripts/bench-analyze.py reads columns by name and skips
unknown # lines; the app's Csv.read keeps unknown keys and picks the new column
up automatically.
Producing TokenMetrics::text means parsing the entire generation so far —
common_chat_parse takes the whole string and cannot resume — so the cost is
O(n) per token and O(n-squared) across a turn, growing as the answer does. Off
the chat path it was not free either: shown_view returns the raw string, so
every token copied everything generated so far into the metrics struct.
It ran for every token of every run. The plain CLI path writes m.piece and
never looks at m.text; a benchmark run reads neither. Only the BMOE_PROGRESS
line protocol actually renders it.
GenerateRequest::render_text now carries that decision. The session loop sets
it (the protocol puts the parsed answer on every line); run() ties it to
--progress; it defaults to true so an embedder that has never seen the flag
behaves exactly as before. Note this cost sat OUTSIDE the wall bracket that
feeds gen_seconds, so it never showed up in the reported tok/s while being paid
on every benchmark run.
The end-of-turn parse was also done twice — once for RunResult::generated_text,
once to commit the assistant message to history — over the same string with the
same parameters. Parse once, use twice; the prefilled-turn and parse-failure
fallbacks are unchanged and now stated in one place.
Also: three back-to-back stats() snapshots to read three fields become one, and
--compute-trace / --io-trace are actually passed to the session loop. The CLI
parsed those flags, opened the files, wrote their headers and then dropped both
sinks when entering --session, so the user got an empty trace and no
explanation.
Gates 7/7. Both CLI output paths verified by hand on the tiny gate model:
--progress still accumulates the parsed text per line, the plain path still
streams pieces.
Two defects in the overlap/speculation paths, found auditing core/ with a
performance lens.
prefetch() pushed a speculative read job per projection as it built them, and
set the entry's pending count only after the last one succeeded. A commit that
failed part-way through an expert therefore returned with jobs already queued
against a count that was never set. A worker draining one of those decrements
spec_remaining_ from zero to negative: the entry can never reach zero, so it
never completes; quiesce_spec never sees it in spec_touched_, so its pages are
never released; and the "already queued" test at the top of prefetch skips that
expert forever after. One transient failure, permanent damage — in precisely
the low-memory situation this path exists to degrade gracefully in.
Stage an expert's jobs locally and publish them as a unit once all of its
projections are committed; on a part-way failure, hand back the pages it did
commit and stop. The staging vectors are members, so the path stays
allocation-free after the first call like the rest of the per-load scratch.
The commits also move OUT of io_mtx_: prefetch runs on the eval thread right
after a real batch was published, so a syscall per projection inside the lock
stalled the lanes trying to pull real read indices out of it — undercutting the
one invariant speculation has to keep.
Separately, every completed slice read took ready_mtx_ and called notify_all,
waking every compute thread blocked on any expert so each could re-check a
predicate that was almost never its own, and paying the mutex even when nobody
was blocked — the usual case, since a slice normally lands inside the spin.
Waiters now register in an atomic count that the publisher consults first.
Both the registration and the flag publication are seq_cst, so the two cannot
miss each other: either the waiter observes the flag and never sleeps, or the
publisher observes the registration and notifies.
Considered and left alone: folding the done_cnt_ increment into the same
critical section as the next-index fetch. It saves one uncontended mutex
acquisition per job against reads that cost ~100 us, which is not worth
perturbing the drain protocol for.
Gates 7/7, and the overlap set (G4a/b/c) run five times over to give a missed
wakeup a chance to hang. Format clean.
The eval callback sees every node of every graph. For each one the ask pass ran
sscanf("ffn_moe_topk-%d") — and, with the drop policy armed, match_weights on
top of it, and with the probe on, a third — to decide a question the first eight
characters already answer. sscanf drags a format-string parser and locale
machinery through a hot path that a fixed-length memcmp settles.
Gate all three matchers behind one memcmp against "ffn_moe_", the prefix every
routing node in this engine shares. The overwhelming majority of a graph's nodes
fail it and leave without touching anything else. Parse the trailing layer index
with a digit loop instead of "%d"; that is also stricter, rejecting a tail like
"12abc" that sscanf would have read as layer 12 — the same strictness node_layer
already applies next door.
Remember a layer's terminal weight-chain node as the INDEX of its name rather
than the name. The four chain names are mutually exclusive (the '-' check makes
them so), so the index identifies it exactly, and the old std::string assignment
was heap-allocating on every weight node of every layer of every token whenever
dropping was armed — the names run past the small-string buffer — with a string
compare next to it. term_node_ becomes term_variant_, a vector<int8_t>.
No routing decision changes. Byte-identity gates pass 7/7, including the two
that exercise the deferral this touches (G8a, G8c).
weight_at is the single source of truth for where token j's slot k sits inside
a weight node — the route trace reads through it and the drop policy writes
through it, so the two agreeing is what keeps the policy from zeroing another
token's slot.
It told the 3-D [1, nu, nt] node apart from the norm variant's pre-reshape 2-D
[nu, nt] by asking whether ne[0] == 1. That works only while the routing is
wider than one expert. At n_expert_used == 1 the 2-D node IS [1, nt]: it passes
the test, gets read with nb[2] as the token stride instead of nb[1], and every
token after the first is taken from the wrong row. With --drop-cold-experts
armed in prefill the same misreading sends the policy's zeroing and renorm
writes to the wrong slot.
Match the full extents against the routing's own nu and nt instead, which the
two callers already have. That separates the shapes exactly; where both fit —
a single token at top-1 — they name the same element, so either answer is
right. An unrecognised shape keeps the old test rather than inventing an offset.
Latent, not active: no model in the catalog routes top-1, which is also why the
top-1 gate (G8c) passes either way — it decodes one token at a time, and at
nt == 1 the wrong stride is multiplied by zero.
Three defects found auditing the I/O layer, all in the same seam between what
O_DIRECT requires and what the buffered fallback actually needs.
FileReader::read freed the lane's bounce before allocating its replacement but
left bounce_sz_ at the old capacity. After a failed realloc the lane advertised
a buffer it no longer had: the next smaller read passed the capacity test, took
the null pointer and preadd into it, and kept failing for the rest of the run.
Clear the size with the pointer.
The buffered path — taken when the platform refuses O_DIRECT, or when the
open-time verify catches storage that mis-serves it — shared the direct path's
mechanics for no reason. Buffered pread has no alignment constraint, so the
outward-aligned window, the staging buffer and the interior memcpy were pure
cost in the mode that is already the slow one. Read straight into the caller's
memory and report the bytes actually moved; the aligned window is a direct-mode
concept and only direct mode returns it now.
The per-lane buffered tail fd was opened without checking. When it failed, a
read reaching the file's sub-alignment EOF tail fell back to the O_DIRECT fd,
which must reject a length that short — so fd exhaustion surfaced as an
unexplained read error far from its cause. Report it at open, and name the
missing fd if a tail read is ever attempted without it.
Also: vm_reserve gains MAP_NORESERVE, so the address-only reservation is one by
contract and not merely by the default overcommit heuristic, and file_size uses
fstat rather than mutating the shared fd position to answer a question about the
file.
Host gates pass (7/7).
The pull_request trigger ignores markdown, docs/, LICENSE and .gitignore; the push
trigger does not — so merging a docs PR re-ran format and the full submodule build
plus ctest on a tree the PR had already gated.
paths-ignore cannot simply be copied onto the push trigger: that trigger also carries
the release tags, and a tag push has no commit range for a path filter to read, which
would silence the tag build that attaches the APK. So the question is answered once by
a ten-second guard job that diffs the pushed range, and the build jobs are gated on it.
Everything that is not a branch push — pull requests, manual runs, tags — answers yes
unconditionally, as does a push whose previous head is unknown (branch creation, force
push): the guard only ever saves work it can prove is redundant.
--ubatch is a memory-reservation knob, not a capability, and the list is what a
reader scans to learn what the engine does. The remaining entries now say what
each setting is for; the numbers stay where they can carry their caveats — the
benchmark tables and the bench-data write-ups they link to.
* chore(release): 0.16.0 (versionCode 30)
Carves the accumulated [Unreleased] entries into a dated section and advances the
example app to match: --ubatch N, and the expert-prediction work — --predict-log
plus --predict-prefetch, the latter shipping off by default because its own
matched-pair measurement refuted it.
* docs(changelog): restore the blank line before 0.15.1's Fixed heading
* feat(moe): --predict-log, measure how predictable expert routing is
Temporal prefetch was built on a predictor nobody had priced. The
previous-token bet turned out to be right ~38% of the time on Qwen and
~18% on gpt-oss, which cannot pay for the reads it speculates. The
lesson was not that prefetch is impossible but that a predictor should
be measured before it is wired into anything. This adds the instrument,
not a policy.
--predict-log ranks each layer's experts a layer early, by running the
NEXT layer's router matrix on the CURRENT layer's gate input. The
residual stream barely moves between layers, so the stale input ranks
nearly as the real one will -- the mechanism FATE (arXiv 2502.12224)
reports 78.8% for. It is training-free and changes no model: the router
matrix is a dense weight already resident, and the prediction is one
GEMV per layer. The matrix is learned from the graph (the gate matmul's
first source) rather than looked up by tensor name, so no architecture
is named anywhere in the path; being a weight leaf, its pointer stays
valid into layers the current token has not reached.
Three predictors are scored against the routing the router actually
produced, so they are comparable on one run: the stale gate, the
previous-token bet --prefetch already places, and a zero-staleness
control. The control is the load-bearing part. It shares every line of
code with the prediction under test and differs only in using the
layer's own matrix, so it must reproduce the selection llama.cpp
computes from those same two tensors. A transposed matrix, a mis-strided
row or the wrong token of the batch collapses it toward chance while the
stale figure would stay superficially plausible; an architecture that
selects by something other than raw-logit ranking (an additive bias,
group-limited routing) puts it below 100% and by that much the stale
figure understates the method. The CLI says so rather than letting the
gap be blamed on staleness.
Reported per layer as well as in aggregate, because an aggregate
flatters a prefetch: what a prefetch costs is set by the layers it gets
wrong, and a MoE model's first layers route far less predictably than
its last. Denominators are printed per predictor -- the stale gate
structurally cannot rank layer 0 (nothing precedes it) or the first
token of a run, and those routings are counted as unscored rather than
folded in, since a routing that was not ranked is not a wrong guess. A
predictor with no routings at a layer prints "-", never 0.0.
Diagnostics only: nothing it computes reaches load_layer, the cache or
the graph, so a probed run reads exactly the bytes an unprobed one does.
G9a gates that byte identity and G9b gates the control, which reads
100.0% on the tiny model. It is not free -- one isolated node and two
GEMVs per layer on the eval thread -- so a probed run is not a benchmark
run, and it requires --moe-stream since routing does not depend on how
the weights reached memory.
docs/expert-prediction.md also records the caveat the number will need:
a high score would say the routing is knowable earlier, not that knowing
it earlier makes decode faster. On a flash already saturated, starting a
read sooner adds no bandwidth -- which is why prefetch, layer-LFU and
the expert sidecar all lost despite improving the metric each was
designed around.
* feat(moe): --predict-prefetch, speculate on the stale-gate prediction
The probe said the routing is knowable a layer early (~89% of routed
slots on a 128-expert model, vs ~43% for the previous-token bet the
temporal prefetch acts on). This wires that prediction into the existing
speculative read path -- same cache buffers, same accounting, same
settle, same moe-prefetch summary line (tagged [stale-gate]) -- so the
only thing that changes is which guess rides the idle lanes.
Two decisions carry the design:
Speculation is issued AFTER the current layer's load, not at prediction
time. Every load path begins by quiescing speculation, and a layer's own
load sits a few graph nodes after its gate matmul -- reads queued at
prediction time would be cancelled before a lane picked them up. Each of
the three load sites (plain topk, deferred drop, drop fallback) issues
the pending next-layer prediction right after its load_layer, restoring
the same read-ahead window the temporal prefetch gets. On the tiny-model
gates this is the difference between 0% and 27% of speculated experts
proving useful -- the latter matching the probe's measured accuracy on
that model, which is the accounting agreeing with itself.
It is drop-aware. With --drop-cold-experts armed, a predicted expert
whose predicted routing weight (softmax over the predicted top-k)
falls below the drop threshold is not speculated: if it misses, the
policy discards it unread, so reading it ahead would spend the exact
I/O the policy exists to save. The top prediction is always kept,
mirroring the policy's own pin of the top-weighted expert. The known
interplay is inherited from the temporal prefetch and deliberate: a
correct guess un-drops an expert, buying quality at the same threshold
rather than speed.
The routing width the prefetch predicts at is learned from the topk
node, not read from config, so an --n-expert-used override stays honest
with no extra plumbing. Mutually exclusive with --prefetch (two
predictors would double-speculate the same future); requires the LRU
cache; decode only. The control GEMV remains probe-only, so the
production path costs one gate GEMV per MoE layer per token.
Gates: G10a proves byte-identity through the speculative path
(prefetch-sync, forced small cache, hits and evictions both occur);
G10b proves the run actually speculated and that useful-hit accounting
tracks the probe's accuracy. Off by default, pending an on-device A/B.
* feat(moe): cap predictive speculation at the top 2 predicted misses per layer
Speculating the whole predicted routing was measured on device at -38%
against its own baseline despite every intermediate metric improving
(hit 77.6->88.9%, stall 32->9 ms, 84% of speculations useful): the +33%
flash bytes and the vm commits fighting a full cache (major faults x3)
cost several times the stall removed. The stall a prefetch can remove is
head-of-line only -- overlap already hides the tail behind the expert
matmul -- so the cap keeps the part of the bet that can pay and drops
the part that provably cannot. Residency-aware: a predicted expert
already in cache does not burn a slot of the cap.
* perf(moe): rebuild the predictive prefetch around its measured costs
The observer-tax run priced the naive implementation: ~35-45 ms per
GEMV pass (ggml_fp16_to_fp32 is a function call per weight element --
21M calls/token) and ~20 ms/token for the extra isolated node, against
a speculation machinery that costs ~15-25. The GEMV and the barrier
were the feature; this commit removes both.
- gate_scores converts F16 natively on aarch64 (one instruction,
vectorizable) instead of a function call per element.
- The prefetch no longer isolates the gate matmul: the ask pass hands
over its source pointers for free, and the gate-input row is read at
the topk callback with no barrier of its own. A sampled watchdog
(the zero-staleness control, every 512 routings) validates the
barrier-less read and disarms the prefetch out loud if the memory
planner ever reuses that buffer -- without it, a future llama.cpp
bump could silently turn the predictor into a noise generator.
- The GEMV runs on a dedicated worker at a TWO-layer horizon: one
layer ahead has no landing spot (the callbacks between a layer's
gate and its own load are microseconds apart), while at l+2 the
worker has a whole layer for a ~0.5M-MAC job and the result inherits
the same post-load issue window as before. The probe now also scores
stale-2, so the extra layer of staleness is priced per model rather
than assumed.
- The prediction's residents are RETAINED (new IExpertSource::retain:
move-to-MRU, deliberately not a cache hit so the hit-rate metric
stays honest) -- protecting a predicted expert costs zero bytes,
unlike prefetching it. Only predicted misses are speculated, still
capped at 2.
The probe path keeps its barrier and both GEMVs: a probed run is
diagnostics, and its job is to be right, not fast.
* feat(moe): --predict-spec-max N — how much flash the prediction may spend (0 = retention only)
The cap was a constant; the retention-only point (0) is the config the
whole experiment now hinges on -- the prediction protecting predicted
residents from eviction while spending no flash at all -- and a
measured constant that cannot be varied is not a mechanism. Validated
[0, 8]; retention happens at every value because it is free.
* docs(predict): record the 2026-07-23 campaign — accuracy confirmed, throughput verdict open
Accuracy on device: stale-gate 88.6% (Qwen3-30B) / 80.7% (Qwen3.6),
control exactly 100.0% on every layer of both, prev-token 43/35% --
corroborating the offline route-trace estimates.
Throughput: the day's C-vs-B losses are recorded WITH their
invalidation. Re-running the reference on the by-then-hot device gave
3.93 tok/s against the cool morning's 6.54 with byte-identical I/O,
hit and drop counts -- the engine is deterministic, the -40% was
silent thermal capping, and every variant had been compared against
the cool number. The one thermally matched pair that was measured
(speculation on 8 io lanes, -28%, effective flash bandwidth 585->392
MiB/s) kills the more-lanes hypothesis specifically; spec-max 2 and
retention-only still owe a matched cool pair.
What did survive: the observer-tax decomposition (109 ms/token: the
per-element exported-function F16 conversion at 21M calls/token, plus
the barrier), and retention moving hit rate 0.1pp on a 3000 MiB cache
-- the offline replay bound confirmed from inside the engine.
* feat(app): expose the predictive prefetch as an experimental Streaming toggle
Off by default. Gated on streaming + a live cache like the temporal
prefetch, and the two settings disable each other in the UI -- the
engine refuses the pair, and a control the engine will reject is worse
than one that cannot be set. The spec-max rung selector (0/1/2/4)
surfaces the retention-only point, which is the configuration the open
throughput question most needs measured from the app. Session
signature includes both fields so flipping them reopens the process.
No versionCode bump: this is a PR-branch test build, not a release.
* docs(predict): record the 2026-07-24 matched pairs — read-ahead refuted, retention hit-neutral
A four-cell session (B, retention-only, spec-2, B sentinel) run at fixed 30 s
spacing re-proved the thermal-contamination mechanism (sentinel −17%, clusters
silently capped from cell 2) and yielded one genuinely matched pair: spec-2 vs
the B sentinel at the same caps and battery temperature, 3.14 vs 3.96 tok/s
(−21%) with hit rate up 4.3pp and 79% of speculations useful. Speculation
improves every metric it owns and still loses the wall clock — the flash has no
spare bandwidth to spend. Retention-only again moved the hit rate by nothing
(77.2% vs 77.6%), as the offline replay bound predicted.
* feat(app): contrast the two prefetch predictors in the UI, default spec-max to 0
The temporal and predictive toggles now say what actually differs — the bet
("repeats the previous token", ~40%) vs the question ("ask the next router a
layer early", ~85%) — instead of describing mechanisms side by side. Spec-max
defaults to 0 (retention only): the matched-pair A/B showed the read-ahead
losing −21% on a saturated flash, so 0 is the only rung the measurements did
not refute, and the helper text says so.
* chore(predict): file the changelog under Unreleased, drop a dead member, record the verdict
Three loose ends found reviewing the branch for merge:
- The changelog entries had been appended to the already-released 0.15.1 section;
they belong under [Unreleased], where the release commit carves them out.
- nu_hint_ was written on every routing and read nowhere: a leftover of the
synchronous first design, whose successor passes the routing width straight into
the prediction job.
- docs/roadmap.md still closed the routing-prediction question on the 2026-07-12
removal. It now records what reopening it with a training-free predictor found:
the accuracy is real and the throughput is not, for the same reason more lanes
and the sidecar lost.
Compute buffers are reserved for the worst-case graph, and this engine had
always set n_ubatch = n_ctx so that any fitting prompt prefills in one pass.
That quietly ties RESIDENT MEMORY to the context rather than to the work:
320 MiB reserved at n_ctx 2048, falling to 80 MiB at 512, scaling exactly with
the context and nothing else.
On an engine whose entire problem is that the expert cache and the dense weights
compete for the same RAM, that is a real budget and it was invisible. Decode is
unaffected -- a decode graph is one token wide whatever this says -- so the cost
is prefill throughput, which processes a long prompt in more, smaller passes.
Default 0 keeps the previous behaviour exactly.
validate() rejects a negative value and one above n_ctx: reserving for a batch
that cannot arrive is the inverse of what the knob is for.
Found while chasing what looked like a catastrophic slowdown and turned out to
be this reservation pushing an already-tight system into reclaim at nearly
6 000 major faults per token, none of them the workload's fault. That happened
during GPU work, where the reservation was largest, but the coupling it exposed
is the engine's own and applies to every configuration. Salvaged here because
the GPU offload it was found alongside measured negative and was closed (#104).
Also fixed, all found alongside and unrelated to each other:
* A stray % in the --dense-weights help text was read by printf as a conversion
specifier, so that line printed garbage and read past the argument list.
* scripts/build-android.ps1 judged cmake by whether it wrote to stderr rather
than by its exit code, which broke it under Windows PowerShell 5.1 -- the NDK
toolchain prints progress to stderr and 5.1 turns that into a terminating
error under ErrorActionPreference = Stop.
* .gitignore now covers third_party/opencl-sdk/, 5.3 MB of Khronos headers and
an ICD loader that fetch-opencl-android.ps1 drops into the tree and nothing
was ignoring.
Qwen3.6-35B (22.3 GB, ~1.5x RAM) on a Windows x86 laptop, engine
unmodified: baseline 4.78 tok/s, best recipe --overlap --cache-mb auto
--drop-cold-experts 0.75 = 7.33 (+55%), above the phone's 5.0-5.8 on
the same model.
Streamed decode is DRAM-bandwidth-bound there, the opposite of the
phone: compute sits at ~0.11 s/token in every cell (a ~9 tok/s ceiling
at zero I/O), doubling compute threads buys 2 ms, and io8 = io4 to the
millisecond once the cache is fixed — round 1's apparent +17% from
lanes was entirely a --cache-mb auto budget confound (5.3-7.4 GiB
across cells), recorded with its resolution. Aggregate flash read
stays ~900 MiB/s on a ~3 GB/s NVMe regardless of lanes, and with
overlap on it drops to ~330 MiB/s: async reads compete with FFN
compute for the same DRAM bandwidth.
README gets a dedicated Desktop section (table + the flip) replacing
the old one-line quick check; raw CSVs and findings.md land under
docs/bench-data/2026-07-24-desktop-qwen36/.
Ships the narrow-routing warning that landed after 0.15.0 was tagged.
0.15.0 shipped the app default at 75% with no caveat for models that
route few experts. On gpt-oss (4 of 128) that default puts the threshold
at 18.8% of the routing against the 9.4% every published number was
measured at, and nothing said so. The warning closes that gap in the
engine and in the app.
Also corrects a contradiction inside the 0.15.0 section itself: the
app-default entry still called the quality cost unquantified while the
Measured block below it reported the GSM8K result.
The threshold is a fraction of the uniform share 1/top-k, so what it
removes scales with how wide the routing is. At top-k 8 -- where every
number in docs/expert-dropping.md was collected -- 0.75 means "below 9.4%
of the routing", a tail trim. At top-k 4 it means "below 18.8%", and at
top-k 2 "below 37.5%", which on a miss discards the whole minority
expert: closer to halving the routing than trimming it, and unmeasured.
The engine now says so once at load when dropping is armed and the
effective top-k is 4 or fewer, quoting the actual share for the model in
hand. It warns rather than clamping or refusing: it cannot know whether
that trade is acceptable for a given model and task, and silently
adjusting a number the caller chose would be worse than a loud caveat.
MoeStreamConfig::drop_low_topk_warn is documented as an EVIDENCE
boundary, not a physical one -- nothing in the streaming path reads it.
The app shows the same caveat inline under the setting, in the error
colour, computed from the width the loaded model reports rather than
assumed -- and only once a model is loaded, since guessing would be worse
than staying quiet. gpt-oss is the case this exists for: it routes 4 of
128, and the app default is 75%.
To make that possible, BMOE_READY gains n_expert_used (the effective
width after any override, 0 on a non-MoE model) and Session exposes
n_expert_used() for embedders. Additive: older consumers ignore it.
Say explicitly that the in-app downloader takes any direct gguf URL, so
any model from the supported architecture families streams the same way
as the catalog entries. Also swaps an em-dash for a colon in that
paragraph.
Turbo top-k and cache-aware dropping made the same point in three places:
two long Features bullets and a top-k-only benchmark section that ended by
introducing the other knob. Now it is one three-line Features bullet and
one section, "Trading quality for speed", that states both knobs' numbers,
their determinism difference, and the GSM8K caveat once.
Also drops a near-verbatim duplicated paragraph in expert-dropping.md and
repoints its README anchor at the renamed section.
NaN compares false against every bound, so the plain min/max range check
accepted it — and it slipped past the cache-required check the same way.
Downstream the hook's `frac > 0.0f` clamp turned it into "off", so nothing
unsafe ran, but a config that should error out was validating clean.
The range check is now a negated inclusive range, which NaN fails.
Covered in config_test. Found by a pre-release audit.
Closes the open item the throughput measurement left behind. 15 questions
taken verbatim from the GSM8K test split, same model and config as the
throughput A/B (Qwen3.6-35B-A3B, cache 3000, 4 lanes, overlap, ahwb,
no-think), four cells differing only in the threshold:
off 12/15 0 routings dropped
0.50 13/15 ~3%
0.75 13/15 ~14%
1.00 13/15 ~28%
Twelve of the fifteen questions give an identical final answer in all four
cells. All variation sits on two questions, and it flips in BOTH
directions as the threshold rises rather than worsening with it -- the
signature of a perturbation on problems already at the edge of the model's
competence, not of accumulating damage. Reply length is flat and no cell
produced a missing #### marker or an empty reply.
Decoding is greedy, so there is no sampling noise: every difference
between cells is caused by the policy. That makes the twelve identical
answers a real statement rather than a coincidence. It does not make 13
against 12 an improvement -- perturbing a question the model already got
wrong can land either side of the right answer, and at 6.7 points per
question this sample cannot establish the sign of the effect. It rules out
a collapse, not a subtle cost, and the write-up says so.
The grading rule was fixed before any output was read (number after the
last ####, falling back to the last number in the reply, applied
identically to every cell) and answers.csv carries every reply in full so
it can be audited. The prompt set, driver and grader are committed with
the results.
README gains a section stating how the three kinds of evidence are kept
apart -- gates assert correctness, device runs measure speed, a public
benchmark checks quality -- and where each comes from, including the
GSM8K source and licence.
The description explained the mechanism twice and the choice not at all,
which is backwards: a user opening this setting has already decided to
try it and only needs to know which rung to pick.
The rung labels now carry the trade ("barely bites" / "recommended" /
"fastest, roughest") and the blurb loses the paragraph that restated the
mechanism, gaining instead what separates 75 from 100: the measured
speeds, and that the top rung discards twice as much of the routing to
get there. Both numbers are marked as one model, since that is all that
has been measured.
* feat(moe): --drop-cold-experts, spend quality only where it buys I/O
Turbo top-k drops the tail of a routing whether or not those experts were
already in RAM. A resident expert costs no flash read, so that trade pays
quality for nothing on the ~80% of decode routings that are cache hits.
This skips a routed expert only when it is a cache MISS and the router
weighted it below frac x (1/top-k). Replayed over the committed route
traces at frac 1.0, decode phase, that avoids 66% of flash reads for 9.5%
of the router's weight mass, where --n-expert-used 5 avoids 23% for a
comparable 10.6% -- about 3x the reads at the same quality cost.
Implementation. The decision needs the FINAL router weights, which arrive
several nodes after the topk where the streamer normally loads, so
load_layer() is deferred to the terminal node of the layer's weight chain.
Which node that is depends on the model's gating, so the hook learns it
from the graph rather than carrying an architecture table; if it fails to
arrive the hook forgets it and re-learns rather than re-betting. A dropped
slot has its weight zeroed and its expert id repointed at the routing's
top-weighted expert: an unread expert can sit in reserved-but-uncommitted
VM and mul_mat_id would touch it anyway, so the kernel is given memory
that is certainly resident and multiplies it by exactly zero.
Requires the LRU cache -- with --cache-mb 0 residency reads all-miss and
the policy would silently degenerate into an unconditional weight cut.
Prefill is excluded by default. The top expert is always pinned, so no
routing can be emptied at any threshold.
Gates: G8a/G8a' prove the deferral and the learned terminal node are
transparent (byte-identical output, zero drops, at a threshold below any
producible weight); G8b that full strength against a constantly-evicting
cache never reaches an unloaded slot; G8c that at top-k 1 dropping is a
no-op, pinning both the top-expert guarantee and the threshold tracking
the effective top-k.
Three existing metrics shift meaning under dropping and the docs now say
so: cache_hit_pct rises without the cache serving more (a dropped routing
is a miss that is never looked up), and token/layer_demand measure what
was staged rather than routed. prefetch.md's "cannot change output" is
scoped, limitations.md gains the non-reproducibility entry, and
benchmark-method.md warns that reversing the run order cannot distinguish
a moved drop rate from a contaminated cell.
Off by default in the CLI and in the app. The output is not reproducible
-- what gets dropped depends on what the cache held -- so it carries no
rows in the README tables, and switching it on by default waits on a
published on-device A/B rather than on the replay argument alone.
* feat(app): default cache-aware dropping to 75%, measured on device
Qwen3.6-35B-A3B (top-k 8 of 256), in-app, cache 3000, one variable
changed: 2.549 tok/s off, 3.938 at F=0.75 (+55%), 4.702 at F=1.0 (+84%),
with flash reads falling 248 -> 163 -> 48 GiB. Per-token bootstrap
intervals separate every pair except off vs 0.50, which overlaps -- at
half the uniform share the policy drops 2.7% of routings and buys
nothing, which doubles as a negative control that the machinery is free
when it does not fire.
Run order was 1.0, off, 0.5, 0.75, so the two fastest cells are the first
and the LAST; thermal drift would have made the last the worst. The
mechanism orders by threshold even though the run order does not.
The replay turned out conservative rather than optimistic. It is
documented as an upper bound because it cannot model the cache changing
in response to dropping: at F=0.75 it was accurate (37% predicted, 34%
measured), at F=1.0 it understated (66% predicted, 81% measured). Avoided
reads free cache capacity, which raises the hit rate, which leaves fewer
misses to drop.
75% rather than 100% is deliberate: it takes the larger part of the win
for half the discarded routings (14% against 28%). Quality is still
unquantified -- no perplexity number and no side-by-side exists -- so the
conservative end of a measured range is the defensible default. The CLI
stays off; the byte-identity gates need a deterministic default.
Also records cache_hit_pct rising 67.8 -> 90.7% as the documented
accounting artefact rather than the cache serving more, and majflt/token
as dominated by each run's starting memory state, not by the threshold.
This reverts 45a90a2.
The feature merged before the evidence for its shipping default did. The
replay numbers argue the shape of the trade is favourable, but no on-device
A/B is published in this repository, and the app default it landed with
(75%) changes model output for every user of the demo app -- and changes it
non-reproducibly, which no other setting in this engine does.
Nothing was found wrong with the code. This is a sequencing decision: the
work returns as a pull request, with the app default back to off, so the
measurement lands before the default does.
Reverted rather than force-pushed: main is public and this commit was
already pushed, so the history stays honest about what happened.
* feat(moe): --drop-cold-experts, spend quality only where it buys I/O
Turbo top-k drops the tail of a routing whether or not those experts were
already in RAM. A resident expert costs no flash read, so that trade pays
quality for nothing on the ~80% of decode routings that are cache hits.
This adds the cache-aware version: skip a routed expert only when it is a
cache MISS and the router weighted it below frac x (1/top-k). Replayed over
the committed route traces at frac 1.0, decode phase, that avoids 66% of
flash reads for 9.5% of the router's weight mass, against 59%/37% for
--n-expert-used 3 — about 3x the reads avoided at a comparable cost. The
threshold is a curve, not a switch: 0.75 trades 4.4% of the mass for 37% of
the reads, better than --n-expert-used 5 on both axes.
Implementation. The decision needs the FINAL router weights, which arrive
several nodes after the topk where the streamer normally loads, so with the
policy armed load_layer() is deferred to the terminal node of the layer's
weight chain. Which node that is depends on the model's gating, so the hook
learns it from the graph rather than carrying an architecture table; until
it is known a layer loads at its topk node undropped. A dropped slot has its
weight zeroed and its expert id repointed at the routing's top-weighted
expert: an expert we decline to read may sit in reserved-but-uncommitted VM
and mul_mat_id would still touch it, so the kernel is given memory that is
certainly resident and multiplies it by exactly zero. Survivors are rescaled
by default, since a systematically shrunk expert output perturbs the residual
stream more than the missing contribution does.
Prefill is excluded by default (cold cache, ~4x the weight mass discarded,
and compute-bound anyway). The largest weight in a routing is always at least
the uniform share, so frac <= 1 can never empty a layer; validate() enforces
the bound and the top expert is pinned regardless.
Gates: G8a proves the deferral and the learned terminal node are transparent
(a threshold below any producible weight leaves the output byte-identical),
G8b that full strength with the cache off never reaches an unloaded expert.
Unlike every other knob this one is state-dependent: what gets dropped
depends on what the cache held, so output is not reproducible across runs.
Off by default, not in the app's settings, and NOT yet measured on device —
the numbers above are a static replay and an upper bound. docs/expert-
dropping.md states what is owed before it is recommended anywhere.
* feat(app): expose cache-aware expert dropping in Settings
Speed / quality -> Drop cold experts, as a percentage of the uniform
share (off / 50 / 75 / 100). The engine takes a fraction; the app stores
integer rungs, so the setting divides by 100 on the way to the flag.
Disabled in mmap mode: the policy asks the expert source what is resident,
and there is no expert source without the streamer. Included in the session
signature, so changing it reopens the session rather than being ignored by
a process already loaded.
Off by default. This exists so the A/B can be run where the engine actually
ships -- through the app, not a pushed CLI binary.
* fix(moe): require the cache for dropping, and correct what it reports
Review of the first two commits found the policy could be armed in a
configuration where it is not cache-aware at all, and that two of the
numbers it reports were wrong.
- Require the LRU cache. With --cache-mb 0 query_residency answers
all-miss, so the policy silently degenerated into an unconditional
weight cut -- exactly what --n-expert-used already does, under a flag
claiming to consult residency. validate() now rejects it, as it already
did for --prefetch, and the app gates the setting on the same condition.
- Fix experts_routed. It was incremented inside apply_drop, so it counted
what the policy examined rather than what the router selected: layers
before the terminal weight node is learned, and every un-armed phase,
were missing from the denominator. The reported drop rate was a fraction
of the wrong thing.
- Re-learn instead of re-betting. If the node learned as terminal does not
arrive, the deferral now also forgets it, so the next graph loads at the
topk node while it re-learns. Deferring again on a stale guess would
repeat the fault every token against a graph that had moved.
- Point the gates at a real cache. G8a/G8b ran with the cache off, where
the shared-slot path has no reserved-but-uncommitted memory -- so the id
repointing, which is the design's whole safety argument, was never
exercised. They now run against a constantly-evicting budget. Adds G8a'
(asserts routings were examined and none dropped, so an inert-threshold
flake fails legibly) and G8c (at top-k 1 dropping is a proven no-op,
pinning both the top-expert guarantee and the threshold being taken
against the effective top-k).
Docs: three metrics change meaning under dropping and none of them said
so. A dropped routing is a miss that is never looked up, so cache_hit_pct
rises without the cache serving more, and token/layer_demand measure what
was staged rather than routed -- documented in telemetry.md, pressure.md
(size the cache with dropping off, then turn it on) and metrics.h.
prefetch.md's "cannot change output" is scoped: under dropping a correct
guess un-drops an expert. limitations.md gains the non-reproducibility
entry, benchmark-method.md the axis plus a warning that reversing the run
order cannot distinguish a moved drop rate from a contaminated cell, and
architecture.md/runtime.h no longer claim unconditional determinism.
Fixes two anchors the README rename broke, and a changelog sentence that
quoted the equal-I/O row while drawing the equal-quality conclusion.
App: Drop cold experts defaults to 75%. The default is a product decision
taken on the maintainer's device; no benchmark for it is published here,
and docs/expert-dropping.md says that plainly instead of implying a
measured figure. The CLI stays off by default -- the byte-identity gates
need a deterministic default.
* feat(dense): add --dense-weights ahwb, dense weights in reclaim-exempt memory
The dense weights must stay resident — every token touches them — and
docs/android-memory.md finds every lever for holding them there closed: mlock is
capped at 64 KiB by the vendor, the cgroup protections are v2-only, MGLRU is
disabled at runtime, and MADV_COLD only redirects reclaim. The exception measured
in 0.13.5 is dma-buf: its pages stay pinned for the buffer's lifetime because a
device may DMA from them, an unprivileged app can allocate one through
AHardwareBuffer, and it reads at exactly anonymous-memory speed.
This wires that allocation into the dense-weights policy. `ahwb` is `anon` with a
single substitution — pio::pinned_alloc instead of the heap — leaving the O_DIRECT
read, the tensor rebind and the mmap handback identical, so an A/B between the two
moves one variable rather than comparing two code paths. Allocation is per tensor,
which keeps it well under the 2047 MiB lock ceiling; a tensor that did exceed it
fails the run instead of quietly taking an anon buffer, since a silent mix would
corrupt the comparison in the direction that flatters the feature. For the same
reason the mode refuses to start on platforms without such an allocation rather
than falling back, which would let an A/B become a mode against itself.
dense_resident_frac keeps working under it — mincore does report on the dma-buf
mapping, which was not obvious — and there it doubles as the falsification test:
pinned pages that fall below 1.0 disprove reclaim-exemption directly.
Exposed as a Dense weights -> Pinned (experimental) setting in the example app,
default off.
What is NOT established is that any of this helps, and the mode should not be
turned on because the reasoning is good. Reclaim-exempt memory does not create
memory: under a >RAM model the RAM the dense weights stop yielding is taken from
the expert cache or from the page cache feeding the stream. That is the trade that
already refuted the bulk restore (#28) and the per-layer LFU cap, both of which
delivered exactly the local gain they predicted and lost throughput anyway. The
deciding A/B is owed and must be run in the app, not over adb: single-shot adb runs
never idle, and this class of bug lives in the reclaim the app's engine suffers
while it sits.
Verified: all 7 host gates pass; on device the three modes generate identical text
on the tiny MoE model, ahwb allocates its 39 pinned buffers, and dense_resident_frac
reaches 1.000 under it.
* test(dense): measure --dense-weights ahwb at +17.9%, and correct the mechanism
In-app on a 1354-token generation (Qwen3.6-35B-A3B, k=8, cache 3000, same session
and binary as its control): 2.588 -> 3.053 tok/s, bootstrap intervals disjoint.
dense_resident_frac reads exactly 1.000, minimum included, in every pinned run, so
reclaim-exemption is now measured rather than inferred.
The mechanism is not the predicted one, and that correction is worth more than the
number. Major faults are EQUAL between the modes (265 vs 257): anon already keeps
the dense weights off the flash. What it does not prevent is the kernel taking ~15%
of them into zram, where a later touch costs a minor fault plus a decompression —
a cost that appears in no I/O counter and no fault counter, so it lands in
compute_ms, which is a residual rather than a measurement of arithmetic. The entire
delta shows up there (298 -> 241 ms) while io_ms, stall_ms and cache hit rate stay
within 1%, and swap falls 562 -> 294 MiB. anon protects the dense weights from
flash; ahwb also protects them from zram.
The trade this was expected to lose on does not appear: the expert cache is
untouched, hit rate identical to the decimal, because the dense set (~1.6 GiB) is
small next to a 3000 MiB cache budget. That is also why it should NOT be extended to
the cache without sizing the prize first — only ~294 MiB of that budget sits in zram,
against a cost of 3 GiB of rigid LMK-accounted memory and the loss of the
reserve/commit/evict elasticity the cache is built on.
Default stays anon. In the decisive pair ahwb ran first and an order effect cannot be
excluded — the reversed pair is owed — and this is one device, one model, one config.
Two negative results are committed alongside so the reasoning is checkable: three
67-74 token pairs that are ALL inconclusive (per-token CV 33-71%, every interval
overlapping), because reclaim accumulates and short turns never build up enough of
it; and a cross-day pair reading +63.6% that is not usable, since anon alone moved
+38.8% between the two days.
Transferable: compute_ms has been absorbing zram decompression all along, so earlier
"this regime is compute-bound" conclusions deserve re-examination.
* feat(tools): add bmoe-membench, a read-bandwidth probe for pinned memory
The dense weights are the one part of the model that must stay resident, and no
lever documented in docs/android-memory.md can keep them there: RLIMIT_MEMLOCK is
capped at 64 KiB by the vendor, the cgroup knobs are v2-only, and MADV_COLD only
redirects reclaim rather than preventing it. One allocation an unprivileged app can
make is exempt by construction — a dma-buf, whose pages stay pinned because a device
may DMA from them at any time. Userspace reaches it through AHardwareBuffer.
Whether that is usable hinges on a property nobody publishes: gralloc decides per
allocation whether a buffer is CPU-cacheable, and uncached memory loses the cache
line, the prefetcher and most memory-level parallelism. A dense matmul streams
weights, so it would pay that in full. On a device whose flash serves 1.3-2.5 GB/s,
uncached pinned weights would read SLOWER than refaulting them off storage — so the
idea is gated on a bandwidth measurement, not on a design argument.
bmoe-membench measures exactly that: the same sequential read kernel over an
anonymous mapping (what --dense-weights anon already allocates) and over a locked
AHardwareBuffer BLOB, reporting the ratio. The two outcomes are an order of
magnitude apart, so the tool is built to be unambiguous rather than precise.
--probe-max additionally reports the largest BLOB the driver will hand over, since
no limit is documented and the dense working set is multiple GiB.
It links nothing beyond the stdlib, so it cross-compiles in seconds and also builds
on the host, where the anon row still runs. No engine, CLI or app code is touched.
* test(memory): measure reclaim-exempt memory — bandwidth gate passes, size capped at 2047 MiB
A locked AHardwareBuffer BLOB reads at exactly anonymous-memory speed: within 0.5% on
both CPU clusters single-threaded (30.4k and 45.2k MiB/s) and at 4 threads (59.0k),
and the CPU_READ_OFTEN hint makes no difference to it. The way this idea could have
died on arrival was an uncached mapping — gralloc chooses cacheability per allocation,
and uncached memory would read at or below the flash bandwidth it is meant to save,
making pinned dense weights slower than refaulting them. It does not. The gate passes.
The binding constraint turned out to be size, in a place the documentation does not
point at. Allocation succeeds up to the 4 GiB format cap implied by the 32-bit BLOB
width, but AHardwareBuffer_lock — which is what yields a usable CPU pointer — refuses
at exactly 2048 MiB with EINVAL while 2047 MiB succeeds. A ceiling landing precisely
on 2^31 is a signed 32-bit type in the lock path, not memory exhaustion. The usable
unit is therefore 2047 MiB and anything larger must span several buffers, which this
engine can do since dense_weights already allocates per tensor.
Two corrections to the tool, both cases of it producing a confident wrong number:
- --probe-max probed allocation only, and so reported roughly double the usable size.
It now locks every candidate, and refines to 1 MiB so a ceiling sitting on a power
of two is visible as evidence about its cause.
- --modes ran anon first whatever order was asked for, and a single pass turned out to
rank CPU clusters rather than allocators: the first, un-pinned run reported a clean
0.67x that inverted when the modes were swapped, because the scheduler's choice of
core moves this number 1.5x. Modes now run in the order given, --repeat interleaves
them, and the header says to pin with taskset.
What this does not show is that pinning helps. Reclaim-exempt memory does not create
memory: under a >RAM model the RAM the dense weights stop yielding has to come from
the expert cache or from the page cache feeding the stream, which is the same trade
that refuted the bulk restore and the per-layer LFU cap. Nor did any run here put the
device under the pressure a >RAM decode creates, so reclaim-exemption itself remains
an inference from how dma-buf works. Recorded as an open, unproven lever.
docs/android-memory.md gains the dma-buf row its lever table was missing, the roadmap
records the open question, and the raw runs are committed alongside the analysis.
The README claimed "everything runs on stock llama.cpp" and described the
submodule as plain, which has been inaccurate since the submodule was
repointed at the bmoe/expert-ready-hook branch of Helldez/llama.cpp. The
headline benchmark configuration uses --overlap, so the front-page numbers
come from a build that carries the hook.
docs/seam.md and docs/architecture.md already documented the hook, its
zero-cost-when-unregistered property and its sunset condition; only the
README was out of step. State the same thing where people actually read it:
public API for the whole streaming path, one ~25-line exception for
--overlap, dropped when upstream ships an equivalent.
Also drops the "on stock llama.cpp" contrast from the prior-art paragraph,
which credited flash-moe for running through a community fork while
implicitly claiming we carry none.