The README explains --decide as choosing from a list with one prefill, and its NPU prefill section describes the routed expert arena with its per-model numbers.
The NPU prefill's expert arena read every expert of every layer ahead of its
routing. It now reads, ahead of a layer's routing, the experts the previous
graph routed there, and at the routing node whatever the routing adds. The
matmul reads only routed experts, so the output is bit for bit the same. A layer
routing more than --prefill-routed-full (0.85) of its experts gets the next one
read whole; --no-prefill-routed restores whole layers everywhere.
Phone, Hexagon v81 NPU, top-4, same session, every answer identical:
Qwen3.6-35B-A3B Q4_0 7.68 -> 4.16 s, Q4_K_M 9.95 -> 5.37 s, Gemma 4 26B-A4B
Q4_K_M 6.69 -> 3.70 s, Nemotron 3.5 30B-A3B Q4_0 7.42 -> 6.81 s.
Also: --decide-probe (experimental per-decision expert usage and layer-exit
answers), BMOE_DECIDE prefill_dev_* counters, gates G17f/G17g, app 0.28.0.
A session opened with --decide answers which of a list of choices the model would pick, read from
the next-token distribution after one prefill, with no decode. The state after a shared prefix is
kept and restored when the next prefix extends it. Android app: a Choose from options switch.
With --prefill-device, a decision is prefilled by the chat turn's placement rule and keeps no
prefix state: llama.cpp saves a sequence through KV views that do not follow the moved model state
(gate G18g). Also fixes the engine version, stuck at 0.23.0 since 0.24.0. App 0.27.0 (42).
Nemotron-3.5-Lightning-30B-A3B now comes from ggml-org at Q4_0 (~18.9 GB) instead of a third-party Q4_K_M (~25.5 GB); Ornith-1.5-35B-A3B leaves the catalog (qwen35moe stays supported). Ships with the next release, under 0.27.0.
Wide prefill graphs run on the Hexagon NPU through a two-layer arena streamed from flash; decode stays on the CPU. The app offers it in its own NPU section, off by default, Snapdragon only. A missing or unopenable device leaves the run on the CPU. release-apk builds the Hexagon backend and a skel per NPU generation in a separate, secret-free job. Bundles the llama.cpp bump to bmoe/expert-ready-hook-2609 (K-quants on the NPU). App 0.26.0 (41).
nemotron_h_moe is the third expert layout: gate-less. Each expert is
up, ReLU^2, down, so the registry row names ffn_up_exps and
ffn_down_exps and leaves the tail slot empty, as the fused gemma4 row
does. The Mamba2/attention blocks, the shared expert, the optional
latent projections and the MTP block all stay on the resident side of
the seam. No llama.cpp change and no submodule bump: the pinned tree
already builds nemotron_h_moe.
make-tiny-moe.py learns the whole shape in miniature (hybrid stack,
latent projections, biased sigmoid router, shared expert, a trailing
MTP block that is never loaded), and it runs as a third byte-identity
gate. Every identity gate passes on it.
The architecture never puts two MoE blocks next to each other, so the
forward predictors (predict-prefetch, route-ahead, the stale half of
predict-log) have no next layer to target. The gate reads that from the
file and reports those checks N/A instead of failing or passing them
vacuously; an unreadable file keeps them strict.
Ornith-1.5-35B-A3B is qwen35moe and needs no engine change. Both models
join the Android catalog at Q4_K_M; neither has device numbers yet.
Also releases 0.25.0: versionCode 40, versionName 0.25.0, dated changelog.
Dates the section, and puts it back in the order the file's own header claims to
follow: Added, Changed, Fixed, with the two Fixed blocks the release accumulated
merged into one. No entry is added, removed or reworded.
llama.cpp maps every gguf it loads and keeps the mapping for the model's lifetime. On
Windows that is expensive in a way nothing had attributed: while a section of a file is
alive, NTFS serialises concurrent unbuffered reads on that file, and a lane opened while
the section existed keeps serialising against it after the section is gone. Four I/O lanes
therefore delivered exactly one lane's throughput, which is why lanes and threads have
always measured dead on the desktop host and why the engine read at about a third of what
the drive can serve.
--release-mmap hands the mapping back after load: unmap the file, close its section, reopen
the reader lanes. Both halves are needed. Whether it is safe is decided by looking rather
than by reasoning, the engine asks the OS whether any weight the capture pass observed still
points inside a mapping of the model files, and declines if any does. Off by default,
because the check answers for the pointers the capture saw and for no others.
Host A/B on Qwen3.6-35B-A3B Q4_K_M: 3.16 to 4.63 tok/s (+46%), flash stall per token 0.182
to 0.074, with bytes read, hit rate, evictions and re-reads identical to the digit and the
generated text byte-identical. On the phone the read path is flat (f2fs does not serialise)
but CPU per token falls about 9%; that cell is two short runs per variant and is recorded as
a direction, not a number.
Also fixes a bug this uncovered, independent of the flag: when a gguf carries no
output.weight, llama.cpp builds the output head from the token embedding table and the model
holds two identically named tensors over the same bytes. The capture pass keyed its map by
name, so --dense-weights anon and ahwb rebound one and left the twin reading the mmap for
the whole run, which on a model past RAM means the output projection served by page faults
from flash. The capture now records every distinct leaf object by address and the dense
policy rebinds every tensor over one file range onto the same buffer.
Adds bmoe-iobench --mmap / --reopen-lanes / --range-mb / --fresh, the cells that isolate the
mechanism, a mapping_release unit test on both platforms, the app switch "Release the model
mapping", and the bench findings. README, architecture, AGENTS and roadmap updated, the last
correcting a diagnosis this refutes.
The refreshed clip went in unlabelled, so the frame said nothing about which
model was running while the DeepSeek hero next to it names its own. Same
treatment now: a black band over the app title bar carrying the model and
quantization, "running on a phone, airplane mode" under it, and a second band
over the navigation bar with the real-time note and the run's tok/s.
Typography is measured off the DeepSeek asset rather than guessed - Arial 14 for
the subtitle (the same string lands on the same pixels, x 72 to 287), Arial Bold
16 for the model name, Arial 15 for the caption, and its purple. The palette is
built with stats_mode=full: on static overlay text the diff mode spends no
colours and quantizes white down to 249.
The caption reads "real time" rather than the DeepSeek clip's "from here on -
real time", which was there because the master it was cut from had sped-up
stretches ahead of that point. This clip is real time end to end.
The clip in the README was the UD-IQ3_XXS file at 2.03 tok/s. This is a new
in-app run on the 12 GB test phone: 82 tokens at 3.48 tok/s, real time, with
prefill 6.87 s (3.9 tok/s), 13487 MB streamed, a 1000/1000 MiB expert cache at
56 % hit, 70 major faults per token, pinned dense weights, overlap on, four
lanes, cold experts dropped at 100 %.
The file is the DevQuasar Q2_K build the catalog gained in #189, not the one the
old clip used, so the caption names the quantization and the size follows it:
six shards, 80,447,449,856 bytes. The two builds are the same size on disk to
within 2 %, because the 28.8 GB n-gram table dominates both and stays at IQ4_NL
in each. What differs is the dense side — 2.4 GB pinned against 4.3 — which is
the cache room the app README already describes.
The architecture table carried two stale claims: a single ~2 tok/s figure that
belonged to UD-IQ3_XXS alone, and a note that the model was pinned to an
unmerged upstream PR. Upstream support has been in since b10666.
On a 12 GB phone this model is capped by its dense side: 4.3 GB pinned
and walked every token in the UD-IQ3_XXS, which leaves the expert cache
1000-1500 MiB. Every published dynamic quant keeps that side at 5-8 bit
whatever its overall size. DevQuasar's plain Q2_K is the one file with
the dense side at 2-4 bit (2.4 GB pinned), at the price of coarser
experts (no importance matrix). The entry sits next to the UD file, its
blurb states the trade, and the docs that list the catalog follow.
The legacy-merged-file rule in ModelCatalog.statusOf (a single-file
gpt-oss from an earlier release still counts as on-device) also fired
for every entry whose fileName is its own first shard, which is the
shape DeepSeek V4 Flash and Qwen3.8-Flash-Next have. That shard is the
smallest of the set and lands seconds into the download, so the row
read as on-device while shards 2..N were still in flight: no progress
bar, a Run on the incomplete set failed at load, and if the chain died
there was no Download button left to resume it.
The rule now applies only when the entry's fileName is not one of its
shards. The app module gains its first JVM unit test (JUnit 4) covering
the rule, and CI runs the unit tests after the debug APK build.
A first command with nothing but -m, -p and -t ran plain llama.cpp on mmap: streaming
off, cache off, and a dense policy that only applies once streaming is on. Nothing in
the report said so, because the moe-stream: block only prints when streaming is enabled,
so a baseline run read as a measurement of this engine and got reported as one (#186).
A mode: line is now printed on every run. With streaming off on a MoE architecture the
build has a recipe for it says the run is a baseline and names the flag; on any other
model it says the architecture is not one this build streams. RunSummary carries the
model's arch so the CLI can tell those apart.
--cache-mb defaults to auto whenever --moe-stream is on. The previous default of 0 meant
the cache was off, which re-reads every routed expert from flash every token. On a 16 GB
host streaming Qwen3.6-35B-A3B Q4_K_M, 63 tokens, -t 8 --overlap: 1.164 to 2.351 tok/s
and 585 to 238 MiB per token, same output. The default is resolved in the CLI, not in
the library, so an embedder passing 0 still means no cache; an explicit --cache-mb or
BMOE_CACHE_MB still wins, including an explicit 0.
Reported by @eiffel31.
Three lines still described the state before #182 and #184.
- README said macOS is not exercised by CI. It is now compile-checked on every
pull request and release tag, alongside Windows.
- benchmark-method told a Mac reader that o_direct is a no-op there, which would
have them ignore the one field #182 made truthful.
- The platform-caveat paragraph still said macOS does not bypass the page cache.
limitations.md and community-benchmarks.md were already correct; these three were
missed in the same pass.
Adds a build-only macos-14 job so the __APPLE__ I/O paths are compiled on every PR and release tag, not just by release-host at release time. Mirrors the macos-arm64 release-host configuration and smoke-tests the resulting binary.
Follow-up to #182.
A direct request on Apple opens normally and applies fcntl(F_NOCACHE, 1) to the descriptor instead of silently returning a buffered fd. F_NOCACHE is a caching hint, not an I/O mode: no alignment contract, no DMA promise, so a direct reader on Apple keeps plain pread semantics (pio::direct_needs_alignment() splits "uncached descriptor" from "alignment-constrained reads") and skips the O_DIRECT bounce path.
Independently, the o_direct telemetry (CSV preamble, decode-trace header, streaming banner) now reports what the shard opens achieved, the AND across shards after every platform refusal and open-time downgrade, instead of the requested configuration, on every platform.
Measured on a 16 GB Apple-silicon Mac, model on an external volume, 256-token protocol, interleaved A B B A A B: decode 0.94 vs 0.61 tok/s (+56%, non-overlapping), byte stream identical between arms, cache-hit equal; the gain is the buffered arm's page-cache pollution doubling the compute residual while stall stays flat.
Addresses the macOS half of #179.
Experimental, off by default. Before a decode routing is committed, every
expert already in the LRU cache gets its score raised by L times the
token's score range and the top-k is taken again, so a near-tie goes to
the expert already in RAM (Skliar et al., arXiv:2412.00099). The same
number of experts runs; fewer are read from flash. Scores are read from
the tensor the graph itself sorted, exact for any gating function.
Desktop, Qwen3.6-35B Q4_K_M at L=0.15: 258 to 119 MiB of flash per
token, 2.37 to 3.84 tok/s, perplexity +1 to 4 %, tinyMMLU 88 to 84/100,
HumanEval-50 42 = 42. The on-device A/B is still owed, hence experimental.
Also: --ppl / --ppl-step / --ppl-list / --ppl-choices (teacher-forced
perplexity, one token per decode so cache-dependent policies are priced
where they act), scripts/tinymmlu-bench.py, scripts/humaneval-bench.py,
gates G8d/G8e, app switch "Prefer cached experts" under Experimental,
docs/cache-aware-substitution.md.
Dense tables the graph only gathers rows from (the token embedding, on
most models) are bound to reserved address space and fetched in 16 KiB
slabs inside a bounded LRU window, instead of being read whole and kept
resident. Which tables qualify is decided from the captured graph, not
from a name list. Byte-identical to the resident reference; -497 MiB
pinned on Qwen3.8-Flash-Next and -515 MiB on Qwen3.6-35B on the 12 GB
test phone, throughput neutral, off by default. Gates G15a/G15b.
App 0.24.0 (unreleased), new page docs/row-gathered-tables.md.
A console program started from Explorer gets a console of its own, prints
its usage because no model was given, exits, and the console vanishes with
it: a window that flashes and disappears, indistinguishable from a crash.
- cli: when the console was created for this process alone, the no-model
path says it is a command-line program and waits for Enter before
closing. From a terminal nothing changes.
- release-host: the archive README.txt opens with that fact and a complete
command, and no longer points a Windows user at a bash script as the
only instruction. The Windows build links the MSVC runtime statically so
the exe does not depend on the VC++ Redistributable.
- 0.23.0: version bump and changelog entry.
Split the prompt phase into the same wall-additive terms decode already reports:
prefill_cpu_s / prefill_read_mib / prefill_io_s / prefill_stall_s / prefill_mgmt_s,
as session-level deltas of the streamer's cumulative counters across the prefill
chunks. Session layer only; the streamer is untouched. The keys ride both
BMOE_DONE and the CSV `# summary` trailer, appended so existing readers ignore them.
Closes#173.
The benchmark call is pinned, and both pages were written for the maintainer
rather than for the people landing on them. community-benchmarks.md opened with
a list of hardware we want, which reads as an entry requirement, and neither
page ever answered the first question a contributor has: what do I set?
community-benchmarks.md:
- a "start here" for the three cases someone is actually in (PC or laptop,
Android phone, Apple hardware), each with the command and what to paste back
- the settings-override table, which used to be one buried sentence
- an explicit adb protocol for phones
- "hardware we want to see" moved to the end as open questions: any hardware is
a useful row, the list is what we cannot answer ourselves
benchmark-method.md:
- reference device out of the opening, named once at the end as the provenance
of the published numbers
- new "choosing the parameters": lossless, lossy and experimental knobs kept in
separate tables, each with its default and the telemetry field that says
whether moving it worked
- the hard-won rules kept as method rather than as the story of one session
The community protocol now pins --ubatch 512, which the app has always done and
bench-report.sh never did: prefill width costs resident memory the expert cache
would otherwise get, so a host row was running a different configuration from
the app it is compared against. UBATCH= overrides it.
Two platform limits documented for the first time: macOS has no O_DIRECT and
the engine does not call the F_NOCACHE equivalent, so expert reads there go
through the page cache while the metrics still report o_direct=1; and there is
no iOS target at all. Both were already true.
The dense policy assumed the dense set fits in RAM. Qwen3.8-Flash-Next breaks
that with its 51B n-gram table (per_layer_token_embd, ~28.8 GB at IQ4_NL),
which every mode failed on in its own way. A dense tensor larger than the
kernel's MemAvailable is now held back under every mode: it stays mmap'd
with MADV_RANDOM, leaves the warm sweep, the residency sensor and the auto
cache budget, and one stderr line names it. The bound is size, not access
shape: a row-gathered table that fits keeps its mode. Inert on every other
supported model (Qwen3.6-35B matches its baselines to the decimal under
anon, warm, mmap and auto); gates pass; the guard fires on device and pinned
dense weights survive load.
Qwen3.8-Flash-Next (qwen4exp): 125B total, ~6B active, 512 routed experts at
top-10 plus one shared, 48 hybrid gated-delta SSM / sparse attention layers,
and a 51B n-gram embedding table. One registry row streams the experts; a
dense-policy guard keeps the n-gram table (larger than any phone's RAM)
mmap'd under every mode so pinned and anonymous dense weights survive load.
Runs on the 12 GB test phone at ~2 tok/s with pinned dense weights, compute-
bound, and sits in the app catalog as a three-shard download. Submodule
pinned to upstream master b10666, the first with the architecture merged,
with the expert-ready hook on top. README hero clip, changelog and docs
updated. App 0.22.0 (versionCode 37).
A benchmark contributor should be able to download one file and run
scripts/bench-report.sh without a toolchain. release-host builds a static,
portable CLI (GGML_NATIVE=OFF; x86_64 assumes AVX2, aarch64 armv8.2-a+dotprod)
from a clean checkout of the tag for linux-x86_64, linux-aarch64, macos-arm64
and windows-x86_64, and attaches the archives to the release. workflow_dispatch
with upload=false keeps them as artifacts instead, for trying the workflow on a
branch without touching a release.
Every published number comes from one phone and one laptop. This adds what a
contributor needs to add a row from hardware we do not own, without reading code:
- scripts/bench-report.sh runs the fixed README protocol on any Linux/macOS host
(256 greedy tokens, the reference prompt, auto cache, 4 lanes, overlap, dense
weights out of the page cache), records CPU / RAM / drive and the drive's
measured O_DIRECT rate at 512 KiB requests from the model file itself, and
prints the two markdown tables a report needs. Every figure is read from the
CSV `# summary` trailer by key name.
- .github/ISSUE_TEMPLATE/benchmark-report.yml collects hardware, model, engine
version and the pasted tables; tok/s alone is not accepted as a row.
- docs/community-benchmarks.md holds the protocol, the hardware wanted and why,
the reference models in the catalog quants, the meaning of each column, and
the results table seeded with the README rows.
- README, CONTRIBUTING and the docs index point at it.
Overlap stall is the union of stalled intervals, not the per-thread mean; the app panel draws compute / flash wait / cache mgmt / unattributed from measured terms. Fixes#98.
The guard added in #167 writes cache_cycle_mb into the metrics preamble, but the
app never learned the key. MetricsScreen falls back to showing unknown keys under
their raw name, sorted, at the end of the list -- so the one number this release is
about arrived as "cache_cycle_mb" detached from the cache group it belongs to,
while every sibling (cache_mb, cache_floor_mb, cache_ceil_mb) reads as prose with
an explanation behind it.
Two rows, in the two places every other config field is declared: the label in
MetricsScreen's CONFIG_ORDER, right after cache_ceil_mb so it lands with the rest
of the cache block, and the description in MetricFields.
Also makes the CHANGELOG line true: it claimed the number was recorded, and now it
is also readable where people actually look at it.
Below one token's worst-case routed bytes, global LRU evicts precisely
what it is about to read: the hit rate is 0% while the run still pays
the cache's management time and its RAM. The only protection so far was
cache_min_mb, a fixed floor that happens to sit above the cycle for the
shipped models at their default top-k and stops holding the moment
--n-expert-used widens the routing without touching the budget.
The cycle is priced at init from the model's shape alone — every bound
layer's entry_bytes times min(top_k, n_expert), at the top-k the run
actually applies — and recorded as cache_cycle_mb in the metrics
preamble next to the budget it should be judged against, so a committed
CSV answers on its own whether its cache could ever have hit. The engine
prints one stderr line when the resolved budget falls under it. It warns
rather than refuses: the budget is legal and the output byte-identical,
and the engine states the fact and leaves the choice — the argument that
--cache-mb 0 is strictly better below the cliff stays in
docs/cache-sizing.md, not in the engine's output.
Docs updated in the same commit: cache-sizing.md (the guard section,
past tense), roadmap.md (the item moves from still-worth-doing to
shipped) and telemetry.md (the new preamble key).
* Mark the two doc-invoked host scripts executable
README, AGENTS, CONTRIBUTING and docs/seam.md run scripts/build-host.sh
directly, and docs/benchmarks.md runs scripts/bench-run.sh the same way,
but both are committed 100644, so a fresh POSIX clone answers permission
denied at the first documented step. Windows trees never see the mode
bit, which is likely how it went unnoticed.
* Mark the remaining host scripts executable too
Per the discussion in #163: leaving only the two doc-invoked scripts
executable hides the next instance of the same failure behind an
inconsistent set, so all five scripts/*.sh get the bit.
project(VERSION) in the top-level CMakeLists.txt still said 0.19.0 after 0.20.0
was tagged. It is the single source of BMOE_VERSION -- the comment above it says
so, and that is the point of declaring it in one place -- so every 0.20.0 build
self-reports as 0.19.0, both in `bmoe-cli --version` and in the engine= line of
every metrics CSV it writes. Any CSV committed since 2026-08-17 names the wrong
engine, which is the kind of error that quietly poisons a benchmark read months
later. Nothing else was affected. Reported by gjjkbssg (#163).
Bumped to 0.21.0, the version being developed, with the app's versionCode and
versionName moved in step.
Batched with it, since both are small and both ship in the same release:
- Finer expert-cache rungs below 2000 MiB in the app (#146). The ladder went
500 -> 1000 -> 2000, and that x2 is where the choice is sharp: on an 8 GB phone
1000 MiB runs and 2000 MiB gets the app killed by the OS, so the step handed
the user a cliff instead of a setting. 1250, 1500 and 1750 fill it; above 2000
the existing 1000 MiB step is already a small fraction of the budget and is
unchanged. Reported by eiffel31 (#146).
Ling-3.0-flash (127B total, ~5B active, 512 routed experts) becomes one
registry row: standard split expert suffixes, biased top-k (the lfm2moe
pattern), a resident shared expert, leading dense blocks that never bind,
and a NextN/MTP block llama.cpp does not load by default. The hybrid
KDA/MLA attention stack is dense-side llama.cpp machinery, invisible to
the streaming seam.
The submodule moves to fork branch bmoe/expert-ready-hook-ling3: the same
single expert-ready-hook commit, cherry-picked clean onto upstream 3733366
("model : BailingMoE3 Support"). The old fork branch stays, so the commit
the previous pin names remains reachable. The jump crosses 496 upstream
commits and forces three adaptations, all on our side of the seam:
- llama_model_params lost use_mmap for the load_mode enum; the streamer's
required layout is now pinned with LLAMA_LOAD_MODE_MMAP.
- The nextn/MTP tensors are now skipped at load unless load_mtp is set,
and it defaults to off, so --mtp would have built its second context
over a block whose tensors were never loaded. The block is now requested
exactly when speculation asks for it.
- gpt-oss/harmony now declares its thinking tags, so the think-control
probe decides by mechanism: a prefill that moves past the closed span
into further structure of the format itself is binding; one that just
ends at the closing tag is only a suggestion (LFM2.5 stays reported as
uncontrollable).
Byte-identity gates pass on the new base (qwen3moe, gemma4, 4-shard
split). PC smoke on Qwen3.6-35B-A3B matches the recorded baseline on
every applicable cell, --mtp included (17/19 drafts accepted, 3.43 tokens
per verify decode). On device, Ling-3.0-flash loads and streams over its
512 experts; --mtp on it stops at draft rollback because upstream has no
recurrent-state rollback for bailingmoe3 yet (llm_arch_supports_rs_rollback),
a llama.cpp limitation, not a seam one.
App version 0.20.0 (versionCode 35).
nttld/setup-ndk and softprops/action-gh-release ran from mutable major
tags inside the APK job, which holds a contents:write token. Whoever
controls those tags upstream could have pointed them at unreviewed code
with permission to rewrite this repo's release assets.
Both now resolve to a fixed commit, with the human-readable version in a
trailing comment. GitHub-owned actions stay on major tags: pinning them
would cost an SHA bump on every upstream release for a materially
smaller risk.
The signing keystore was never reachable this way. It lives only in
release-apk.yml, which runs no third-party action and triggers only on
events that already require write access.
The DeepSeek figure no longer drags a paragraph of caveats and a second model's numbers behind it; what follows the hero is one line introducing the video underneath, which is what that space is for.
gpt-oss-120b no longer 'still needs a manual copy to the device'. That stopped being true when sharded downloads landed: it is a catalog entry with two shards and a single progress bar. A README that tells people to do work the app already does for them is worse than one that says nothing.
The benchmark tables name the column Active experts rather than k, matching what the app and the docs call it.
Trading quality for speed is gone as a section. In its place the Speed and quality feature block gets an introduction that says what those settings have in common, how they differ from each other, and why they exist at all on a device far past its memory, without quoting a single figure. The figures were never the point there, and each one needs its method beside it to be worth reading, which is what the linked docs are for.
Desktop says plainly that it is not the primary target for now.
Both features shipped in this release with no runtime coverage at all: validation tests only. That is the same shape of hole that let a zero-copy corruption reach the edge of a merge with five green checks behind it.
G13 gates the speculative loop through the n-gram source. Abstention is forced by requiring a match longer than the whole generation, so the plain path is taken by construction and the identity is structural rather than lucky; the first draft of this gate left the matcher at its default, expected a synthetic model to repeat nothing, and it drafted anyway. A third cell then turns drafting on and asserts the machinery survives, since a batched verify may legitimately order a near-tie differently and there is no reference output for that.
G14 gates route-ahead in two halves. A horizon past the last layer can override nothing, so the run must be byte-identical: that covers the plumbing without touching the lossy part. A horizon of one must commit real routings and still generate, with the override count asserted so an inert policy cannot pass vacuously.
G11 and G12 are reserved for the zero-copy branch and left unused here, which is the collision this file's index now exists to prevent.
README: the features section is one row per setting, giving the name the app uses, the flag and its accepted values. The benchmark tables all take the same shape, so their widths stop depending on how long a configuration description happened to be.
* feat(app): settings grouped by purpose, and four defects an audit found
Settings now show the recommended configuration first and fold everything else
into a collapsed Experimental group per category. That is a statement about
evidence, not about how finished the code is: inside are the levers measured on
one device, measured once, or still owed a measurement. They stay in the
release build, because testing them on other hardware is what this app is for
and a lever nobody can reach is a lever nobody can refute. The caveat is stated
once in the group header instead of leaking into some descriptions and not
others.
Every description was rewritten to say what the setting does for the person
reading it. Out went the measured figures, which need the device, the model and
the day beside them to mean anything and have none of that room under a switch,
and out went the implementation names: O_DIRECT, top-k, dma-buf, mmap and KV
cache are not what someone deciding whether to turn something on needs to know.
The metrics screen keeps the flag names, deliberately: there the reader is
matching the UI against a CSV column and the technical name IS the vocabulary.
Four defects, all found by auditing rather than by anything failing:
The session signature is now derived from the argv instead of being a
hand-written list beside it. Those two had to be kept in step with nothing
enforcing it, and forgetting a field is a silent bug: the setting appears to
change while the engine keeps running the old configuration. Three of four
rebases this week collided on exactly that list.
A malformed end-of-turn summary no longer strands the UI. The whole handler sat
inside a catch with no failure branch, so a parse error left the state in
GENERATING with no turn committed and nothing said. The streamed answer is now
kept, the reason is shown, and the state returns to READY.
MainActivity drops from about 1050 lines to under 700: the model download and
import UI moves to ModelPickerUi.kt, which shares nothing with the chat screen.
No logic moved, only its address.
Dead code removed: a field whose own comment described a use it did not have,
two functions nobody called, and five string resources describing a UI two
rewrites ago.
* docs: record the settings regrouping and the signature fix
Rule 6: the changelog and the docs a change invalidates ship with it. The app README described the settings screen as it was before the regrouping, and explained one experimental lever in terms of predictor accuracy percentages that the UI no longer shows.
* docs: the improved DeepSeek hero recording
* feat(app): keep the flag vocabulary in Settings, and make Experimental read as a boundary
The first pass at rewriting the descriptions went too far: it renamed the controls into consumer phrasing and lost the vocabulary that lets a setting here be matched against the CLI, the CSV preamble and the docs. Labels are back to the flag's own names; the descriptions are shorter than the originals rather than longer, and still carry no measured figures. The Experimental group gets a divider and a tonal bar: collapsed, it is the only thing between the recommended configuration and the levers that can change the reply, so it has to look like a boundary rather than one more row.
The flagship demo is now the 284B model generating on a 12 GB phone at about
1 tok/s, with its own recording, instead of gpt-oss carrying that slot. The
quoted 0.94 tok/s is the app's own reading and the sentence next to it says
what produced it: cold-expert dropping at full strength, which trades quality.
gpt-oss keeps its lossless and knob-on figures one paragraph down, and the
three-model clip stays where it was.
The model is named by its exact release, 0731, everywhere it appears rather
than only in the opening line. Anyone reproducing this needs to know which
DeepSeek V4 Flash it was.
The list of what ctest gates claimed the LRU cache, evictions, overlap,
temporal prefetch, the dense rebind and multi-turn sessions. It omitted
predictive prefetch and the split multi-shard model, both of which are gated,
and it did not mention that a lossy knob only gets machinery gates. Fixed to
match what the suite actually runs today.
Also: docs/ngram.md was missing from the documentation index despite being a
250-line document for a shipped flag, and the methodology caveat carried an
editorialising aside about future storage in the one paragraph that has to
stay dry.
Four things an audit found, none of which changes engine behaviour.
The demo app declares android:appCategory="game". Vendor performance layers
read that attribute to pick a governor profile, and on the OxygenOS test device
it moved the app onto the boosted path: the foreground CPU ceiling went from
1.9/1.65 GHz to the hardware maximum of 3.32/3.80 GHz, measured before and
after. Decode is the most CPU-hungry thing a phone does outside a game. The
effect belongs to the vendor rather than to Android, and Samsung's game service
has historically throttled what it classifies this way, so the manifest, the
changelog and the app README all say to treat a per-device figure as a
measurement. In-app numbers from before this are not comparable with numbers
from after it.
CI now enforces the versions it claims. The format job installs clang-format-18
by name instead of whatever the runner image ships, which happened to be 18 and
would have started failing every PR against an unannounced version on the next
image bump. The APK job builds with NDK r27c, the release that produces
published APKs, so CI stops validating a build nobody installs. It also passes
-DGGML_OPENCL=OFF, the flag whose absence once shipped a stray backend into two
releases. checkout moves to v5, since v4 pins a deprecated Node runtime.
--help lists every one of the fifty flags the CLI accepts. Six were missing.
--io-trace is the one that mattered: a fully documented, guarded diagnostic
that the usage text never mentioned, so the only way to find it was to read
docs/telemetry.md.
.gitignore covers .claude/, which until now was excluded only by a
machine-local ignore file. A clone elsewhere would have shown a second checkout
with build output and .so binaries as untracked, which is precisely the
situation the never-"git add -A" rule exists to survive.
* feat(engine): --route-ahead N — commit decode routing to the N-layers-early prediction
Every prefetch lives under the same ceiling: layer L's routing needs layer L-1's
output, so any predictor working earlier is approximate and every speculated read
can miss. This inverts the bet. The expert selection of decode layer L is REPLACED
by the ranking layer L's own gate matrix produced on the hidden state N layers back
in the same forward pass, so the selection is known N layers early and cannot miss.
With the cache on, those reads are issued the moment the selection is fixed.
Lossy by construction: it changes the output, and roughly a fifth of slots route to
a different expert than the router chose at N=1. Off by default, mutually exclusive
with both prefetchers and with the prediction probe, since each would speculate on a
future this policy has already decided.
Quality is measured rather than assumed: the committed generations match the
baseline on a four-prompt objective battery, on a long essay, and on a second model
of a different generation and quantization; output stays deterministic across
repeated runs, which expert dropping does not. See docs/route-ahead.md.
Squashed from the seventeen commits of exp/route-ahead: the branch predated the
multi-shard and speculative-decoding work, and replaying it commit by commit meant
resolving the same two collisions seventeen times over. The history is preserved on
the pull request; what lands here is what a squash-merge would have produced anyway.
* fix(engine): refuse route-ahead alongside self-speculation, and say why
Running the two together on a real model committed NOTHING: 0 routings taken,
249 passed through. A verify decode is several positions wide and the policy
correctly declines each one. But it still charged for itself — the prediction
GEMVs ran (2.6 ms/token of worker CPU) and its early reads degraded into
ordinary speculation, falling from 100% useful to 81%.
Cost with no commitment is worse than either feature alone, and nothing told
the user. validate() now rejects the pair the way it already rejects
route-ahead beside the two prefetchers, the app stops emitting the flag and
greys the row out, and the config test covers both draft sources.
Making the combination work is a different change: commit the whole verify
batch to one selection. That is written and measured, and it costs draft
acceptance (70% to 53%) while route-ahead alone still won, so the exclusion
is the honest state today rather than a limitation to be worked around.
Found by the desktop smoke run, not by the gates.
* feat(engine): MTP self-speculative decoding for Qwen3.5/3.6 (proposal)
Qwen3.5/3.6 ship a trained multi-token-prediction block inside the gguf. With
--mtp that head drafts --mtp-draft continuation tokens and the target verifies
all of them in one wider decode, confirming the longest prefix whose argmax
equals what the target itself would have produced. Nothing is approximated and
no weight is skipped, so the quality is the full model's — but it is NOT
byte-identical the way --overlap and --prefetch are, and must not be used in a
byte-identity gate: a verify pass evaluates 1+N positions in one batch, and a
batched matmul is not bit-identical to N single-token ones, so a near-tie can
flip. Off by default.
The prize is that a decode's dominant cost, moving the dense weights and the
routed expert slices, is paid once per group instead of once per token. The
counterweight is that the verify positions route independently, so a layer's
read set widens toward N*k wherever adjacent tokens disagree, and the draft
pass routes through the MTP block's own expert layer on top. Measured on the
desktop host (DRAM-bandwidth-bound, model streamed at ~1.4x RAM): +15.1% at
draft 3 with the host's best recipe (7.12 -> 8.19 tok/s), +29% without the
lossy drop knob, acceptance falling from 71% at draft 2 to 52% at draft 4, and
flash bytes per token rising 19.7 -> 33.7 MiB as the widening predicts. Draft 3
is the optimum here; 4 is worse than 2. On a flash-I/O-bound phone that balance
can invert, so the flag ships off pending the device A/B.
The orchestration is llama.cpp's own (common/speculative.h, public headers
only): no fork, no patch, no submodule bump. Self-speculation is one model with
two contexts over it — the target, created with n_rs_seq so a rejected tail is
rewound from a bounded snapshot rather than replayed, and a draft context
created with ctx_type = LLAMA_CONTEXT_TYPE_MTP. The engine builds the draft
context itself rather than through common_speculative_init_from_params because
the eval callback is per-context: the streamer only sees the MTP block's expert
layer if the draft context carries the same cb_eval.
The MTP block is streamed like any other layer. It sits at layer index n_layer,
contiguous with the trunk and using the same tensor naming, so the hook and the
expert source are sized n_layer + n_layer_nextn; left at n_layer its experts
stay silently mmap-resident. Two consequences that are easy to get wrong: the
capture warm-up has to run on the draft context too (the MTP graph is built
nowhere else), and prefill is fed through the driver so the draft context's KV
reaches the last prompt position.
The loop accepts BEFORE catching the draft context up, so the catch-up runs on
the accepted prefix instead of the whole verify batch. Acceptance depends only
on the target's logits, which are already in hand once the decode returns, and
the rejected tail was being computed only to be deleted a few statements later.
The resulting state is identical — the driver seeds from row
min(n_accepted, n_rows-1), the same row under either batch, and the surviving
KV is exactly the range the rollback used to carve out — while skipping
n_draft - n_accepted positions through the MTP block per group. Since that
block carries its own MoE FFN, on a streamed device those are expert reads that
no longer happen. It also removes the draft context's rollback entirely: it is
never given a tail to drop.
Requires an MTP-converted gguf (most quantisations strip the nextn tensors) and
greedy decoding; both are rejected at load with a message rather than silently
ignored, as is a n_ubatch narrower than the verify batch, which would split the
graph back into single-token passes and spend the draft for nothing.
Telemetry: an "mtp:" summary line, an mtp_batch per-token CSV column (a verify
decode's whole cost is charged to its group's first row, the rest carry zeros),
mtp_drafted / mtp_accepted / mtp_decodes in the CSV trailer and in BMOE_DONE,
and mtp / mtp_draft_max in the CSV preamble. The Android app exposes the flag
and the draft width, off by default.
Host gates pass. Validated on Qwen3.6-35B-A3B-MXFP4 with the streamed recipe:
draft 1 and draft 3 produce identical text, which is the invariant a broken
accept/rollback path would violate. Device A/B still owed.
* perf(mtp): shrink the draft context, make its cost measurable, record the device verdict
The first on-device A/B says MTP loses at every draft width, and the counters
say why. Same gguf with the flag on and off, shipping recipe (overlap, 3000 MiB
cache, pinned dense, drop 0.75), Qwen3.6-35B-A3B-Q4_K_M streamed:
off 5.82 / 6.14 tok/s 69.3 MiB/tok 69-109 majflt/tok
--mtp-draft 2 5.59 93.8 230
--mtp-draft 3 4.38 106.6 633
Speculation is working - 2.35-2.52 tokens per verify decode, 52-69% acceptance
- and still losing, because the prize does not exist in this regime.
stall_s/tok is 0.025-0.027 in every one of those runs, MTP on or off: 11-16% of
the token. This configuration is compute-bound, and what MTP amortises is weight
movement. The costs meanwhile are real and monotonic in the draft width: the read
set widens (+35%, +54% flash bytes per token), CPU per token rises (+28%, +67%),
and the draft context's memory tips the device into a fault storm.
Two things follow, and both are engine bugs rather than facts of nature.
The draft context's graph width drops from 256 to 32. Compute buffers are
reserved for the widest ubatch and the dominant term scales with
ubatch x vocabulary; on device that reservation measured 493 MiB - for a context
that evaluates ONE token per draft step and is handed at most 1 + draft_max
positions by the catch-up, with no logits asked for. Only prefill ever feeds it a
wide batch, and that is one layer, so splitting it costs very little. On this
engine memory is never free: it is the expert cache's, and the cache is what
decides whether the widened verify read set is a hit or a flash read.
And the cost of speculation is now measured instead of inferred. Drafting happens
between decodes, so it never entered wall_ms and tok/s never included it - a
speculated run could report a rate the user was not experiencing. New
mtp_draft_ms per-token column (a slice of loop_overhead_ms, not an addition),
mtp_draft_s/tok in the CSV trailer, mtp_draft_s_tok and loop_overhead_s_tok in
BMOE_DONE, and a second "mtp:" summary line printing the effective rate next to
the reported one.
Adds --mtp-p-min F, which stops drafting once the head's confidence in what it is
proposing falls below F. The draft loop already had this floor and the engine was
passing 0, so it always drafted the full width however unsure the head was - with
roughly half the drafts rejected at draft 3, that is the cheapest waste available
to cut. On a streamed device it pays twice: a draft not made is a pass through the
MTP block (which carries its own MoE FFN, so its own expert reads) that never
happens, AND one fewer independently routed position in the verify batch. Default
0, the setting the host numbers were measured at; the useful value is a property
of a device's balance between drafting cost and acceptance, so it is a knob to
measure rather than a constant to guess.
The Android app now reads the mtp_* keys it was already being sent: acceptance,
tokens per pass, and the effective rate. Before this the UI could not tell whether
speculation had run at all - only the session CSV could - which made the A/B this
commit reports impossible to run from the phone.
Neither mitigation changes the regime. The honest expectation is nearer
break-even, not a win, and the flag stays off by default.
Host gates pass. Note the noise floor: the two off runs did byte-identical work
and still differ by 5.6% in tok/s, and the runs were back-to-back without thermal
gating - the mechanism counters are the trustworthy part, not the exact deltas.
* perf(mtp): split the drafting flash cost from the widened verify batch
A speculated run streams more bytes per token for two unrelated reasons: the
MTP block carries its own MoE FFN, so every draft pass routes experts of its
own, and the verify batch widens the trunk's read set wherever adjacent
positions disagree. They need opposite fixes -- a narrower draft attacks the
first, only better agreement attacks the second -- and the route trace can
separate neither, since its framing brackets the target decode while the head
only ever runs in the draft context.
Measure the head's share directly by bracketing both drafting passes with the
expert source's byte counter, and report it as a third mtp: summary line.
Also record the branch-deletion rule in AGENTS.md: a branch list should only
show work in flight, and a rejected PR loses nothing.
* feat(engine): n-gram prompt-lookup draft source, and the measurement that closes it
The flash split added last commit said where MTP's cost actually is: at draft 3 on
the host, the head's own routing was 2.9% of the extra bytes a speculated run
streams and the widened verify batch was the other 97.1%. So a cheaper draft
producer is worth almost nothing, and the only property that could matter is one
the head does not have -- the ability to decline to draft at zero cost.
--ngram is that source. It takes the last few tokens, finds where that run occurred
before in the prompt or in what has been generated, and proposes whatever followed.
No head, no draft context, no decode, no expert read, and it works on any gguf
including the ones --mtp refuses for want of a nextn block. Below --ngram-min-match
it proposes nothing and the step falls through to a plain single-token decode.
Measured on the host, Qwen3.6-35B-A3B-MXFP4 streamed, 256 greedy tokens, cells
back-to-back with off run twice:
prose off 5.80 / 6.59 mtp3 7.32 eff ngram3 6.51 (cov 7.4%)
copy-heavy off 5.45 / 5.65 mtp3 6.43 eff ngram3 5.24 (cov 15%)
The zero-cost claim holds exactly -- mtp_draft_s/tok reads 0.0000 in every n-gram
cell, against 0.020-0.023 for the head plus the ~500 MiB of expert cache its draft
context takes. But the floor turns out to be per STEP, not per run: the 15% of steps
that did draft widened the read set to 67.2 MiB/token against 48-58 at baseline and,
at 44% acceptance, bought 1.20 tokens per decode. That is not enough to earn the
widening back, and a modest fraction of such steps sinks the run.
A --ngram-min-match sweep settles it rather than leaving it open. Raising the gate
3 -> 5 -> 8 lifts acceptance 44% -> 75% while coverage collapses 15% -> 3.4%, and
narrowing to --draft 1 reaches 82.6% -- the head's own figure on this prompt. Every
cell climbs toward baseline from BELOW and none crosses it; the best configuration
found lands on the floor. A knob whose optimum is its own disablement is not a
tuning problem. Acceptance, not drafting cost, is what pays for a widened batch, and
what a trained head buys is being right often enough to justify a batch that has
already been widened.
--ngram ships off. It is kept because it is the only speculation available on a
model with no head, because the per-step floor is real, and because the counters it
adds make the next speculation claim falsifiable.
Wiring. MtpConfig became SpecConfig with DraftSource {none, mtp, ngram}, and
--mtp-draft became --draft: the width belongs to the verify batch, not to whoever
filled it. --mtp and --ngram are rejected together rather than resolved by flag
order. In the session the gate split in two -- spec_on (wide batch, acceptance,
rollback: both sources) against mtp_on (draft context, common/speculative.h, the
catch-up: the head only) -- which is what lets the n-gram source reuse the whole
verify half while allocating nothing.
A step that drafts nothing now takes the plain path: llama_batch_get_one with a
logits row of -1, byte for byte the unspeculated decode. It used to build the wide
batch anyway. Required for --ngram, and it tightens --mtp-p-min's zero-draft steps
for free.
The matcher is pure policy over token ids with no llama.cpp at all -- not even
llama.h, since llama_token is int32_t -- so it sits on the clean side of the seam,
adds no dependency on the common layer, and is unit-tested with no model
(tests/ngram_test.cpp covers tie-breaks, clipping, self-match exclusion and the gate
boundary). Telemetry: spec= / spec_draft_max= / ngram_min_match= in the CSV
preamble, a new drafted_steps key in the trailer and BMOE_DONE, and an ngram: line
reporting coverage -- without which a delta cannot be divided by the fraction of the
run it applies to. The per-token and trailer counters keep their mtp_ names: they
always described the loop rather than a source, spec_* already means the temporal
prefetch in that trailer, and renaming would break every CSV already holding a
measurement. The Android setting became a three-way picker, migrating the old
boolean preference.
The device A/B agrees and adds a cost the host could not show. Thermally gated cells
(a 120 s settle, then a battery-temperature gate, so all six start between 35.3 and
36.4 C): prose 4.90 inside a 4.59-5.17 band, copy-heavy 3.14 against 4.43 -- a 29%
loss, worse than MTP's 18%. Major faults per token go 126 -> 1427 for a source that
allocates no draft context at all, and that is the rollback snapshots: n_rs_seq =
draft_max is asked for by ANY speculation, since rejecting a draft means rewinding the
KV, and on a hybrid attention/SSM model that snapshot is a real allocation scaling with
the context. The n-gram source escapes MTP's draft context but not the loop's own
memory, and on device that memory is the expert cache's.
The same run re-measured MTP with the thermal confound removed -- 3.64 effective
against 4.43, so the earlier device verdict was not an artefact of benching without a
cooldown gate -- and reproduced the flash split at 3.7% head against 96.3% widened
verify batch, matching the host's 2.9-3.0%.
Byte-identity gates pass; speculation stays out of them for the reason docs/mtp.md
gives.
The app's CSV configuration surface follows: the three new preamble keys get their own
glossary entries rather than falling through to the unexplained-key renderer, and the
draft source joins the short run label. A speculated run is not the same KIND of run --
under speculation a decode confirms a whole group, so its per-token rows are not even
accounted the same way -- and two compare legends differing by it must not read alike.
* feat(app): sharded model downloads — DeepSeek V4 Flash in the catalog, gpt-oss one-tap
A catalog entry can now list shard files. They download sequentially through one
WorkManager chain (per-file HTTP Range resume, one aggregate progress bar, free
space checked once against the whole remaining set), the model picker offers only
the first shard — the file the engine opens — and deleting a sharded entry deletes
the whole set, so no 40 GB tail is ever orphaned.
DeepSeek V4 Flash UD-IQ2_M (~91 GB, three shards) joins the catalog, and
gpt-oss-120b turns from a "merge it on a PC" manual recipe into a one-tap
download; a merged single file from an earlier release still counts as on-device.
App version 0.19.0 (versionCode 34).
* docs(readme): DeepSeek V4 Flash becomes the flagship claim, video slot staged
* docs(readme): standardize benchmark tables (slowest to fastest, one label scheme), drop dashes
* fix(app): a split model's first shard is a MoE model too
The picker's MoE filter looked for an expert tensor inside the file it was
handed. A split gguf's first shard carries the metadata and, in the layout
large quants ship in, almost no tensors: DeepSeek V4 Flash was therefore
classified dense and never appeared in the model dropdown, with all 91 GB
sitting on the device. The header walk now also accepts the metadata key
<arch>.expert_count, which is the definitive MoE signal and always lives in
the first shard; the tensor-name check stays as the fallback.
* fix(app): bound the session's ubatch so compute buffers stop eating the model's RAM
A session opened at ctx 4096 with no --ubatch reserves compute buffers for the
whole width, though decode only ever computes one token. That reservation is
memory the expert cache and the dense weights do not get, and the CLI has
measured it as an 18% decode lever for a while; the app never passed the flag,
so every in-app run since gave it away.
On DeepSeek V4 it is not 18% but the whole result: the same configuration read
14.58 s/token in-app against 2.22 s over adb, with identical flash I/O (2.74 vs
2.89 s) and 3.6x the major faults. The 13.9 s of 'compute' were page faults, the
process swapping while it worked. Prefill pays instead, and barely: chunking it
costs ~7.7x the flash reads for ~6% of prefill wall time.
* feat(app): context is a setting, not a constant
The session opened at a fixed 4096 tokens. That is also memory — the KV cache is
sized for it once at open — so on a model that already fills RAM it competes with
the weights, and there was no way to trade conversation length for room without a
rebuild. It joins the other tunables (default unchanged), and the ubatch is
clamped to it so a graph is never reserved wider than the context. The service
reads the running session's context from its own argv for the 'ctx used/total'
readout, so the number describes the process rather than the current setting.
Measured on DeepSeek V4: the KV is 44 MiB at 512 and ~270 MiB at 4096, small
thanks to the compressed attention, so on that model the setting is not the lever
its size suggests. It is on models with ordinary attention.
* fix(app): review pass on sharded downloads
Four defects, all from the same blind spot: code that asked whether a filename
belongs to a catalog entry compared it against the entry's own name, which for a
sharded model is one file out of several.
- A sharded entry never reached ON_DEVICE. The present-files set was recomputed
only when the SELECTABLE model list changed, but shards 2..N are hidden from it
by design, so finishing a 41 GB shard left it byte-identical: the row offered
Download for a model already fully downloaded, and pressing it did nothing
until the app was restarted. Keyed on the in-flight names as well, which change
exactly when a shard starts or finishes.
- Shards also rendered as pasted-URL downloads, whose Cancel deleted the .part
the worker was still writing while cancelling nothing (the chain is registered
under the entry name). The transfer then ran on an unlinked file for tens of GB
before failing to finalize.
- A sharded gpt-oss also appeared under Imported models, with a Delete that
removed shard 1 and orphaned the rest — the exact failure the delete dialog
exists to prevent.
- That dialog replaced the entry name with the shard names instead of adding
them, so a gpt-oss merged by an earlier release became undeletable.
ModelCatalog.fileNamesOf/isCatalogFile is now the single answer to 'does this
file belong to an entry', and all four sites go through it. Also: a queued shard
reports zero bytes, so aggregate progress now falls back to its .part length
rather than appearing to lose ground on a resumed 50 GB transfer.
* feat(moe): stream split multi-shard ggufs natively + DeepSeek V4 Flash recipe
Hugging Face rejects single files above 50 GB, so every large model ships as
-00001-of-0000N.gguf shards; until now the streamer assumed one file, forcing
a merge with double the disk. gguf_offsets now fans the first shard out to the
whole set and resolves every tensor to (shard, offset); the expert streamer
and the dense loader open one positioned reader per shard and route each read
by the tensor's shard index. Pass the first shard, exactly as llama.cpp takes
it; a missing sibling fails the load with the shard named.
Add the deepseek4 recipe row: V3.2-style routing (256 routed experts, a
per-expert bias like lfm2moe, an always-on shared expert that stays resident)
over the standard split expert suffixes. The V4 compressed-attention machinery
is dense-side llama.cpp code, invisible to the streaming seam.
The byte-identity gates gain a 4-shard qwen3moe fixture (metadata-only first
shard, the layout large quants actually use); make-tiny-moe.py learns
--split-max-tensors. All gates pass, split included.
* fix(moe): cache auto must budget for the anon dense conversion
The auto budget read MemAvailable while the dense weights were still reclaimable
page cache, then dense-weights=anon converted them into buffers the kernel cannot
take back: the same bytes planned twice. Latent since the anon policy shipped
(dense sets were 2-3 GiB and explicit budgets were the benched path); DeepSeek V4
Flash's 6.5 GiB dense set turned it into a device-taking overcommit on first load.
The budget now deducts the pending conversion and says so in the log.
* fix(moe): review pass on the multi-shard path
Three defects the split rewrite introduced, none of which the gates could see:
- The shard index rode in an int8_t, so a model past 127 shards wrapped to a
negative index into the reader vector. The bounds check could never catch it:
it validated the untruncated value. Widened to int16_t, which covers the whole
-%05d-of-%05d filename space.
- DenseWeights::warm() reused one flag as both the inner loop condition and the
partial-warm report, so the first shard that failed to open silently skipped
the warm-up of every later shard. Per-shard condition, sticky report.
- The dense readers stayed allocated for the session after read_anonymous had
copied and rebound every tensor: fds and a per-lane bounce buffer per shard,
sitting next to a cache counting every MiB. Released at the end of init.
Also: the streaming banner read O_DIRECT off shard 0, which under the
small-first-shard layout is metadata only and too short to verify, so it could
claim a mode the shards carrying experts had not got. It now reports the weakest
of the readers.
* build: the engine version says 0.19.0, like the changelog does
The version is declared in CMakeLists.txt and reported by `--version` and by the
run-parameter preamble of every metrics CSV, so a committed benchmark file names
the engine that produced it. This release section landed while the number stayed
at 0.18.0, which would have stamped the wrong engine on every CSV this branch
produces, defeating the one purpose the string has.