A session opened with --decide answers which of a list of choices the model would pick, read from
the next-token distribution after one prefill, with no decode. The state after a shared prefix is
kept and restored when the next prefix extends it. Android app: a Choose from options switch.
With --prefill-device, a decision is prefilled by the chat turn's placement rule and keeps no
prefix state: llama.cpp saves a sequence through KV views that do not follow the moved model state
(gate G18g). Also fixes the engine version, stuck at 0.23.0 since 0.24.0. App 0.27.0 (42).
nemotron_h_moe is the third expert layout: gate-less. Each expert is
up, ReLU^2, down, so the registry row names ffn_up_exps and
ffn_down_exps and leaves the tail slot empty, as the fused gemma4 row
does. The Mamba2/attention blocks, the shared expert, the optional
latent projections and the MTP block all stay on the resident side of
the seam. No llama.cpp change and no submodule bump: the pinned tree
already builds nemotron_h_moe.
make-tiny-moe.py learns the whole shape in miniature (hybrid stack,
latent projections, biased sigmoid router, shared expert, a trailing
MTP block that is never loaded), and it runs as a third byte-identity
gate. Every identity gate passes on it.
The architecture never puts two MoE blocks next to each other, so the
forward predictors (predict-prefetch, route-ahead, the stale half of
predict-log) have no next layer to target. The gate reads that from the
file and reports those checks N/A instead of failing or passing them
vacuously; an unreadable file keeps them strict.
Ornith-1.5-35B-A3B is qwen35moe and needs no engine change. Both models
join the Android catalog at Q4_K_M; neither has device numbers yet.
Also releases 0.25.0: versionCode 40, versionName 0.25.0, dated changelog.
* feat(moe): stream split multi-shard ggufs natively + DeepSeek V4 Flash recipe
Hugging Face rejects single files above 50 GB, so every large model ships as
-00001-of-0000N.gguf shards; until now the streamer assumed one file, forcing
a merge with double the disk. gguf_offsets now fans the first shard out to the
whole set and resolves every tensor to (shard, offset); the expert streamer
and the dense loader open one positioned reader per shard and route each read
by the tensor's shard index. Pass the first shard, exactly as llama.cpp takes
it; a missing sibling fails the load with the shard named.
Add the deepseek4 recipe row: V3.2-style routing (256 routed experts, a
per-expert bias like lfm2moe, an always-on shared expert that stays resident)
over the standard split expert suffixes. The V4 compressed-attention machinery
is dense-side llama.cpp code, invisible to the streaming seam.
The byte-identity gates gain a 4-shard qwen3moe fixture (metadata-only first
shard, the layout large quants actually use); make-tiny-moe.py learns
--split-max-tensors. All gates pass, split included.
* fix(moe): cache auto must budget for the anon dense conversion
The auto budget read MemAvailable while the dense weights were still reclaimable
page cache, then dense-weights=anon converted them into buffers the kernel cannot
take back: the same bytes planned twice. Latent since the anon policy shipped
(dense sets were 2-3 GiB and explicit budgets were the benched path); DeepSeek V4
Flash's 6.5 GiB dense set turned it into a device-taking overcommit on first load.
The budget now deducts the pending conversion and says so in the log.
* fix(moe): review pass on the multi-shard path
Three defects the split rewrite introduced, none of which the gates could see:
- The shard index rode in an int8_t, so a model past 127 shards wrapped to a
negative index into the reader vector. The bounds check could never catch it:
it validated the untruncated value. Widened to int16_t, which covers the whole
-%05d-of-%05d filename space.
- DenseWeights::warm() reused one flag as both the inner loop condition and the
partial-warm report, so the first shard that failed to open silently skipped
the warm-up of every later shard. Per-shard condition, sticky report.
- The dense readers stayed allocated for the session after read_anonymous had
copied and rebound every tensor: fds and a per-lane bounce buffer per shard,
sitting next to a cache counting every MiB. Released at the end of init.
Also: the streaming banner read O_DIRECT off shard 0, which under the
small-first-shard layout is metadata only and too short to verify, so it could
claim a mode the shards carrying experts had not got. It now reports the weakest
of the readers.
* build: the engine version says 0.19.0, like the changelog does
The version is declared in CMakeLists.txt and reported by `--version` and by the
run-parameter preamble of every metrics CSV, so a committed benchmark file names
the engine that produced it. This release section landed while the number stayed
at 0.18.0, which would have stamped the wrong engine on every CSV this branch
produces, defeating the one purpose the string has.
Extend make-tiny-moe.py with --arch {qwen3moe,gemma4}. The gemma4 fixture
emits the fused ffn_gate_up_exps + ffn_down_exps expert tensors, a resident
shared expert, an interleaved dense layer and a mixed sliding-window / full
attention pattern, so the gates exercise the two-projection streaming path.
The qwen3moe output is byte-identical to before, so the existing gate is
unchanged. tests/CMakeLists.txt now wires one generate-fixture + gate per
architecture.