Commit graph

44 commits

Author SHA1 Message Date
Raffaele
3170385fad
feat(prefill): the NPU prefill reads only routed experts (+ --decide-probe); 0.28.0 (#208)
The NPU prefill's expert arena read every expert of every layer ahead of its
routing. It now reads, ahead of a layer's routing, the experts the previous
graph routed there, and at the routing node whatever the routing adds. The
matmul reads only routed experts, so the output is bit for bit the same. A layer
routing more than --prefill-routed-full (0.85) of its experts gets the next one
read whole; --no-prefill-routed restores whole layers everywhere.

Phone, Hexagon v81 NPU, top-4, same session, every answer identical:
Qwen3.6-35B-A3B Q4_0 7.68 -> 4.16 s, Q4_K_M 9.95 -> 5.37 s, Gemma 4 26B-A4B
Q4_K_M 6.69 -> 3.70 s, Nemotron 3.5 30B-A3B Q4_0 7.42 -> 6.81 s.

Also: --decide-probe (experimental per-decision expert usage and layer-exit
answers), BMOE_DECIDE prefill_dev_* counters, gates G17f/G17g, app 0.28.0.
2026-09-29 16:08:23 +02:00
Raffaele
1d37057fac
feat: --decide, choose from a list from one prefill, with no decode (#201)
Some checks failed
ci / changes (push) Has been cancelled
ci / host-windows (push) Has been cancelled
ci / host-macos (push) Has been cancelled
ci / android-apk (push) Has been cancelled
ci / format (push) Has been cancelled
ci / host-linux (push) Has been cancelled
A session opened with --decide answers which of a list of choices the model would pick, read from
the next-token distribution after one prefill, with no decode. The state after a shared prefix is
kept and restored when the next prefix extends it. Android app: a Choose from options switch.

With --prefill-device, a decision is prefilled by the chat turn's placement rule and keeps no
prefix state: llama.cpp saves a sequence through KV views that do not follow the moved model state
(gate G18g). Also fixes the engine version, stuck at 0.23.0 since 0.24.0. App 0.27.0 (42).
2026-09-28 12:01:25 +02:00
Raffaele
10af539a63
feat: prefill on the NPU, decode on the CPU (--prefill-device), in the release APK (#200)
Wide prefill graphs run on the Hexagon NPU through a two-layer arena streamed from flash; decode stays on the CPU. The app offers it in its own NPU section, off by default, Snapdragon only. A missing or unopenable device leaves the run on the CPU. release-apk builds the Hexagon backend and a skel per NPU generation in a separate, secret-free job. Bundles the llama.cpp bump to bmoe/expert-ready-hook-2609 (K-quants on the NPU). App 0.26.0 (41).
2026-09-28 09:52:46 +02:00
Raffaele
674e7dc136
feat(cli): say which mode a run used, and give --moe-stream a cache by default (#187)
A first command with nothing but -m, -p and -t ran plain llama.cpp on mmap: streaming
off, cache off, and a dense policy that only applies once streaming is on. Nothing in
the report said so, because the moe-stream: block only prints when streaming is enabled,
so a baseline run read as a measurement of this engine and got reported as one (#186).

A mode: line is now printed on every run. With streaming off on a MoE architecture the
build has a recipe for it says the run is a baseline and names the flag; on any other
model it says the architecture is not one this build streams. RunSummary carries the
model's arch so the CLI can tell those apart.

--cache-mb defaults to auto whenever --moe-stream is on. The previous default of 0 meant
the cache was off, which re-reads every routed expert from flash every token. On a 16 GB
host streaming Qwen3.6-35B-A3B Q4_K_M, 63 tokens, -t 8 --overlap: 1.164 to 2.351 tok/s
and 585 to 238 MiB per token, same output. The default is resolved in the CLI, not in
the library, so an embedder passing 0 still means no cache; an explicit --cache-mb or
BMOE_CACHE_MB still wins, including an explicit 0.

Reported by @eiffel31.
2026-08-29 20:42:17 +02:00
gjjkbssg
e02b477f91
feat(io): uncached reads on macOS via F_NOCACHE, and o_direct from the open's outcome (#182)
A direct request on Apple opens normally and applies fcntl(F_NOCACHE, 1) to the descriptor instead of silently returning a buffered fd. F_NOCACHE is a caching hint, not an I/O mode: no alignment contract, no DMA promise, so a direct reader on Apple keeps plain pread semantics (pio::direct_needs_alignment() splits "uncached descriptor" from "alignment-constrained reads") and skips the O_DIRECT bounce path.

Independently, the o_direct telemetry (CSV preamble, decode-trace header, streaming banner) now reports what the shard opens achieved, the AND across shards after every platform refusal and open-time downgrade, instead of the requested configuration, on every platform.

Measured on a 16 GB Apple-silicon Mac, model on an external volume, 256-token protocol, interleaved A B B A A B: decode 0.94 vs 0.61 tok/s (+56%, non-overlapping), byte stream identical between arms, cache-hit equal; the gain is the buffered arm's page-cache pollution doubling the compute residual while stall stays flat.

Addresses the macOS half of #179.
2026-08-29 16:05:11 +02:00
Raffaele
4334c89616
feat(moe): cache-aware expert substitution (--expert-substitute, experimental) and --ppl (#171)
Some checks are pending
ci / changes (push) Waiting to run
ci / format (push) Blocked by required conditions
ci / host-linux (push) Blocked by required conditions
ci / host-windows (push) Waiting to run
ci / android-apk (push) Waiting to run
Experimental, off by default. Before a decode routing is committed, every
expert already in the LRU cache gets its score raised by L times the
token's score range and the top-k is taken again, so a near-tie goes to
the expert already in RAM (Skliar et al., arXiv:2412.00099). The same
number of experts runs; fewer are read from flash. Scores are read from
the tensor the graph itself sorted, exact for any gating function.

Desktop, Qwen3.6-35B Q4_K_M at L=0.15: 258 to 119 MiB of flash per
token, 2.37 to 3.84 tok/s, perplexity +1 to 4 %, tinyMMLU 88 to 84/100,
HumanEval-50 42 = 42. The on-device A/B is still owed, hence experimental.

Also: --ppl / --ppl-step / --ppl-list / --ppl-choices (teacher-forced
perplexity, one token per decode so cache-dependent policies are priced
where they act), scripts/tinymmlu-bench.py, scripts/humaneval-bench.py,
gates G8d/G8e, app switch "Prefer cached experts" under Experimental,
docs/cache-aware-substitution.md.
2026-08-29 10:13:47 +02:00
Raffaele
9153dcc27a
feat(moe): serve row-gathered dense tables from flash (--row-stream) (#180)
Dense tables the graph only gathers rows from (the token embedding, on
most models) are bound to reserved address space and fetched in 16 KiB
slabs inside a bounded LRU window, instead of being read whole and kept
resident. Which tables qualify is decided from the captured graph, not
from a name list. Byte-identical to the resident reference; -497 MiB
pinned on Qwen3.8-Flash-Next and -515 MiB on Qwen3.6-35B on the 12 GB
test phone, throughput neutral, off by default. Gates G15a/G15b.
App 0.24.0 (unreleased), new page docs/row-gathered-tables.md.
2026-08-29 10:04:02 +02:00
gjjkbssg
039b8f1e2f
feat(telemetry): add prefill-phase attribution (#175)
Some checks are pending
ci / changes (push) Waiting to run
ci / format (push) Blocked by required conditions
ci / host-linux (push) Blocked by required conditions
ci / host-windows (push) Waiting to run
ci / android-apk (push) Waiting to run
Split the prompt phase into the same wall-additive terms decode already reports:
prefill_cpu_s / prefill_read_mib / prefill_io_s / prefill_stall_s / prefill_mgmt_s,
as session-level deltas of the streamer's cumulative counters across the prefill
chunks. Session layer only; the streamer is untouched. The keys ride both
BMOE_DONE and the CSV `# summary` trailer, appended so existing readers ignore them.

Closes #173.
2026-08-28 19:46:43 +02:00
gjjkbssg
7a8abcdda7
fix(telemetry): attribute critical-path stall explicitly (#169)
Some checks are pending
ci / changes (push) Waiting to run
ci / format (push) Blocked by required conditions
ci / host-linux (push) Blocked by required conditions
ci / host-windows (push) Waiting to run
ci / android-apk (push) Waiting to run
Overlap stall is the union of stalled intervals, not the per-thread mean; the app panel draws compute / flash wait / cache mgmt / unattributed from measured terms. Fixes #98.
2026-08-27 17:45:51 +02:00
gjjkbssg
0f193d3c05
feat(moe): warn when the expert cache budget sits below one token's cycle (#167)
Below one token's worst-case routed bytes, global LRU evicts precisely
what it is about to read: the hit rate is 0% while the run still pays
the cache's management time and its RAM. The only protection so far was
cache_min_mb, a fixed floor that happens to sit above the cycle for the
shipped models at their default top-k and stops holding the moment
--n-expert-used widens the routing without touching the budget.

The cycle is priced at init from the model's shape alone — every bound
layer's entry_bytes times min(top_k, n_expert), at the top-k the run
actually applies — and recorded as cache_cycle_mb in the metrics
preamble next to the budget it should be judged against, so a committed
CSV answers on its own whether its cache could ever have hit. The engine
prints one stderr line when the resolved budget falls under it. It warns
rather than refuses: the budget is legal and the output byte-identical,
and the engine states the fact and leaves the choice — the argument that
--cache-mb 0 is strictly better below the cliff stays in
docs/cache-sizing.md, not in the engine's output.

Docs updated in the same commit: cache-sizing.md (the guard section,
past tense), roadmap.md (the item moves from still-worth-doing to
shipped) and telemetry.md (the new preamble key).
2026-08-25 12:48:22 +02:00
Helldez
4645dc6b09
feat(engine): --route-ahead N — commit decode routing to the N-layers-early prediction (#142)
* feat(engine): --route-ahead N — commit decode routing to the N-layers-early prediction

Every prefetch lives under the same ceiling: layer L's routing needs layer L-1's
output, so any predictor working earlier is approximate and every speculated read
can miss. This inverts the bet. The expert selection of decode layer L is REPLACED
by the ranking layer L's own gate matrix produced on the hidden state N layers back
in the same forward pass, so the selection is known N layers early and cannot miss.
With the cache on, those reads are issued the moment the selection is fixed.

Lossy by construction: it changes the output, and roughly a fifth of slots route to
a different expert than the router chose at N=1. Off by default, mutually exclusive
with both prefetchers and with the prediction probe, since each would speculate on a
future this policy has already decided.

Quality is measured rather than assumed: the committed generations match the
baseline on a four-prompt objective battery, on a long essay, and on a second model
of a different generation and quantization; output stays deterministic across
repeated runs, which expert dropping does not. See docs/route-ahead.md.

Squashed from the seventeen commits of exp/route-ahead: the branch predated the
multi-shard and speculative-decoding work, and replaying it commit by commit meant
resolving the same two collisions seventeen times over. The history is preserved on
the pull request; what lands here is what a squash-merge would have produced anyway.

* fix(engine): refuse route-ahead alongside self-speculation, and say why

Running the two together on a real model committed NOTHING: 0 routings taken,
249 passed through. A verify decode is several positions wide and the policy
correctly declines each one. But it still charged for itself — the prediction
GEMVs ran (2.6 ms/token of worker CPU) and its early reads degraded into
ordinary speculation, falling from 100% useful to 81%.

Cost with no commitment is worse than either feature alone, and nothing told
the user. validate() now rejects the pair the way it already rejects
route-ahead beside the two prefetchers, the app stops emitting the flag and
greys the row out, and the config test covers both draft sources.

Making the combination work is a different change: commit the whole verify
batch to one selection. That is written and measured, and it costs draft
acceptance (70% to 53%) while route-ahead alone still won, so the exclusion
is the honest state today rather than a limitation to be worked around.

Found by the desktop smoke run, not by the gates.
2026-08-02 00:26:38 +02:00
Helldez
49c72e7ce5
feat(engine): self-speculative decoding — the model's own MTP head, or n-gram lookup (#134)
* feat(engine): MTP self-speculative decoding for Qwen3.5/3.6 (proposal)

Qwen3.5/3.6 ship a trained multi-token-prediction block inside the gguf. With
--mtp that head drafts --mtp-draft continuation tokens and the target verifies
all of them in one wider decode, confirming the longest prefix whose argmax
equals what the target itself would have produced. Nothing is approximated and
no weight is skipped, so the quality is the full model's — but it is NOT
byte-identical the way --overlap and --prefetch are, and must not be used in a
byte-identity gate: a verify pass evaluates 1+N positions in one batch, and a
batched matmul is not bit-identical to N single-token ones, so a near-tie can
flip. Off by default.

The prize is that a decode's dominant cost, moving the dense weights and the
routed expert slices, is paid once per group instead of once per token. The
counterweight is that the verify positions route independently, so a layer's
read set widens toward N*k wherever adjacent tokens disagree, and the draft
pass routes through the MTP block's own expert layer on top. Measured on the
desktop host (DRAM-bandwidth-bound, model streamed at ~1.4x RAM): +15.1% at
draft 3 with the host's best recipe (7.12 -> 8.19 tok/s), +29% without the
lossy drop knob, acceptance falling from 71% at draft 2 to 52% at draft 4, and
flash bytes per token rising 19.7 -> 33.7 MiB as the widening predicts. Draft 3
is the optimum here; 4 is worse than 2. On a flash-I/O-bound phone that balance
can invert, so the flag ships off pending the device A/B.

The orchestration is llama.cpp's own (common/speculative.h, public headers
only): no fork, no patch, no submodule bump. Self-speculation is one model with
two contexts over it — the target, created with n_rs_seq so a rejected tail is
rewound from a bounded snapshot rather than replayed, and a draft context
created with ctx_type = LLAMA_CONTEXT_TYPE_MTP. The engine builds the draft
context itself rather than through common_speculative_init_from_params because
the eval callback is per-context: the streamer only sees the MTP block's expert
layer if the draft context carries the same cb_eval.

The MTP block is streamed like any other layer. It sits at layer index n_layer,
contiguous with the trunk and using the same tensor naming, so the hook and the
expert source are sized n_layer + n_layer_nextn; left at n_layer its experts
stay silently mmap-resident. Two consequences that are easy to get wrong: the
capture warm-up has to run on the draft context too (the MTP graph is built
nowhere else), and prefill is fed through the driver so the draft context's KV
reaches the last prompt position.

The loop accepts BEFORE catching the draft context up, so the catch-up runs on
the accepted prefix instead of the whole verify batch. Acceptance depends only
on the target's logits, which are already in hand once the decode returns, and
the rejected tail was being computed only to be deleted a few statements later.
The resulting state is identical — the driver seeds from row
min(n_accepted, n_rows-1), the same row under either batch, and the surviving
KV is exactly the range the rollback used to carve out — while skipping
n_draft - n_accepted positions through the MTP block per group. Since that
block carries its own MoE FFN, on a streamed device those are expert reads that
no longer happen. It also removes the draft context's rollback entirely: it is
never given a tail to drop.

Requires an MTP-converted gguf (most quantisations strip the nextn tensors) and
greedy decoding; both are rejected at load with a message rather than silently
ignored, as is a n_ubatch narrower than the verify batch, which would split the
graph back into single-token passes and spend the draft for nothing.

Telemetry: an "mtp:" summary line, an mtp_batch per-token CSV column (a verify
decode's whole cost is charged to its group's first row, the rest carry zeros),
mtp_drafted / mtp_accepted / mtp_decodes in the CSV trailer and in BMOE_DONE,
and mtp / mtp_draft_max in the CSV preamble. The Android app exposes the flag
and the draft width, off by default.

Host gates pass. Validated on Qwen3.6-35B-A3B-MXFP4 with the streamed recipe:
draft 1 and draft 3 produce identical text, which is the invariant a broken
accept/rollback path would violate. Device A/B still owed.

* perf(mtp): shrink the draft context, make its cost measurable, record the device verdict

The first on-device A/B says MTP loses at every draft width, and the counters
say why. Same gguf with the flag on and off, shipping recipe (overlap, 3000 MiB
cache, pinned dense, drop 0.75), Qwen3.6-35B-A3B-Q4_K_M streamed:

    off             5.82 / 6.14 tok/s   69.3 MiB/tok    69-109 majflt/tok
    --mtp-draft 2   5.59                93.8            230
    --mtp-draft 3   4.38               106.6            633

Speculation is working - 2.35-2.52 tokens per verify decode, 52-69% acceptance
- and still losing, because the prize does not exist in this regime.
stall_s/tok is 0.025-0.027 in every one of those runs, MTP on or off: 11-16% of
the token. This configuration is compute-bound, and what MTP amortises is weight
movement. The costs meanwhile are real and monotonic in the draft width: the read
set widens (+35%, +54% flash bytes per token), CPU per token rises (+28%, +67%),
and the draft context's memory tips the device into a fault storm.

Two things follow, and both are engine bugs rather than facts of nature.

The draft context's graph width drops from 256 to 32. Compute buffers are
reserved for the widest ubatch and the dominant term scales with
ubatch x vocabulary; on device that reservation measured 493 MiB - for a context
that evaluates ONE token per draft step and is handed at most 1 + draft_max
positions by the catch-up, with no logits asked for. Only prefill ever feeds it a
wide batch, and that is one layer, so splitting it costs very little. On this
engine memory is never free: it is the expert cache's, and the cache is what
decides whether the widened verify read set is a hit or a flash read.

And the cost of speculation is now measured instead of inferred. Drafting happens
between decodes, so it never entered wall_ms and tok/s never included it - a
speculated run could report a rate the user was not experiencing. New
mtp_draft_ms per-token column (a slice of loop_overhead_ms, not an addition),
mtp_draft_s/tok in the CSV trailer, mtp_draft_s_tok and loop_overhead_s_tok in
BMOE_DONE, and a second "mtp:" summary line printing the effective rate next to
the reported one.

Adds --mtp-p-min F, which stops drafting once the head's confidence in what it is
proposing falls below F. The draft loop already had this floor and the engine was
passing 0, so it always drafted the full width however unsure the head was - with
roughly half the drafts rejected at draft 3, that is the cheapest waste available
to cut. On a streamed device it pays twice: a draft not made is a pass through the
MTP block (which carries its own MoE FFN, so its own expert reads) that never
happens, AND one fewer independently routed position in the verify batch. Default
0, the setting the host numbers were measured at; the useful value is a property
of a device's balance between drafting cost and acceptance, so it is a knob to
measure rather than a constant to guess.

The Android app now reads the mtp_* keys it was already being sent: acceptance,
tokens per pass, and the effective rate. Before this the UI could not tell whether
speculation had run at all - only the session CSV could - which made the A/B this
commit reports impossible to run from the phone.

Neither mitigation changes the regime. The honest expectation is nearer
break-even, not a win, and the flag stays off by default.

Host gates pass. Note the noise floor: the two off runs did byte-identical work
and still differ by 5.6% in tok/s, and the runs were back-to-back without thermal
gating - the mechanism counters are the trustworthy part, not the exact deltas.

* perf(mtp): split the drafting flash cost from the widened verify batch

A speculated run streams more bytes per token for two unrelated reasons: the
MTP block carries its own MoE FFN, so every draft pass routes experts of its
own, and the verify batch widens the trunk's read set wherever adjacent
positions disagree. They need opposite fixes -- a narrower draft attacks the
first, only better agreement attacks the second -- and the route trace can
separate neither, since its framing brackets the target decode while the head
only ever runs in the draft context.

Measure the head's share directly by bracketing both drafting passes with the
expert source's byte counter, and report it as a third mtp: summary line.

Also record the branch-deletion rule in AGENTS.md: a branch list should only
show work in flight, and a rejected PR loses nothing.

* feat(engine): n-gram prompt-lookup draft source, and the measurement that closes it

The flash split added last commit said where MTP's cost actually is: at draft 3 on
the host, the head's own routing was 2.9% of the extra bytes a speculated run
streams and the widened verify batch was the other 97.1%. So a cheaper draft
producer is worth almost nothing, and the only property that could matter is one
the head does not have -- the ability to decline to draft at zero cost.

--ngram is that source. It takes the last few tokens, finds where that run occurred
before in the prompt or in what has been generated, and proposes whatever followed.
No head, no draft context, no decode, no expert read, and it works on any gguf
including the ones --mtp refuses for want of a nextn block. Below --ngram-min-match
it proposes nothing and the step falls through to a plain single-token decode.

Measured on the host, Qwen3.6-35B-A3B-MXFP4 streamed, 256 greedy tokens, cells
back-to-back with off run twice:

    prose        off 5.80 / 6.59    mtp3 7.32 eff    ngram3 6.51  (cov 7.4%)
    copy-heavy   off 5.45 / 5.65    mtp3 6.43 eff    ngram3 5.24  (cov 15%)

The zero-cost claim holds exactly -- mtp_draft_s/tok reads 0.0000 in every n-gram
cell, against 0.020-0.023 for the head plus the ~500 MiB of expert cache its draft
context takes. But the floor turns out to be per STEP, not per run: the 15% of steps
that did draft widened the read set to 67.2 MiB/token against 48-58 at baseline and,
at 44% acceptance, bought 1.20 tokens per decode. That is not enough to earn the
widening back, and a modest fraction of such steps sinks the run.

A --ngram-min-match sweep settles it rather than leaving it open. Raising the gate
3 -> 5 -> 8 lifts acceptance 44% -> 75% while coverage collapses 15% -> 3.4%, and
narrowing to --draft 1 reaches 82.6% -- the head's own figure on this prompt. Every
cell climbs toward baseline from BELOW and none crosses it; the best configuration
found lands on the floor. A knob whose optimum is its own disablement is not a
tuning problem. Acceptance, not drafting cost, is what pays for a widened batch, and
what a trained head buys is being right often enough to justify a batch that has
already been widened.

--ngram ships off. It is kept because it is the only speculation available on a
model with no head, because the per-step floor is real, and because the counters it
adds make the next speculation claim falsifiable.

Wiring. MtpConfig became SpecConfig with DraftSource {none, mtp, ngram}, and
--mtp-draft became --draft: the width belongs to the verify batch, not to whoever
filled it. --mtp and --ngram are rejected together rather than resolved by flag
order. In the session the gate split in two -- spec_on (wide batch, acceptance,
rollback: both sources) against mtp_on (draft context, common/speculative.h, the
catch-up: the head only) -- which is what lets the n-gram source reuse the whole
verify half while allocating nothing.

A step that drafts nothing now takes the plain path: llama_batch_get_one with a
logits row of -1, byte for byte the unspeculated decode. It used to build the wide
batch anyway. Required for --ngram, and it tightens --mtp-p-min's zero-draft steps
for free.

The matcher is pure policy over token ids with no llama.cpp at all -- not even
llama.h, since llama_token is int32_t -- so it sits on the clean side of the seam,
adds no dependency on the common layer, and is unit-tested with no model
(tests/ngram_test.cpp covers tie-breaks, clipping, self-match exclusion and the gate
boundary). Telemetry: spec= / spec_draft_max= / ngram_min_match= in the CSV
preamble, a new drafted_steps key in the trailer and BMOE_DONE, and an ngram: line
reporting coverage -- without which a delta cannot be divided by the fraction of the
run it applies to. The per-token and trailer counters keep their mtp_ names: they
always described the loop rather than a source, spec_* already means the temporal
prefetch in that trailer, and renaming would break every CSV already holding a
measurement. The Android setting became a three-way picker, migrating the old
boolean preference.

The device A/B agrees and adds a cost the host could not show. Thermally gated cells
(a 120 s settle, then a battery-temperature gate, so all six start between 35.3 and
36.4 C): prose 4.90 inside a 4.59-5.17 band, copy-heavy 3.14 against 4.43 -- a 29%
loss, worse than MTP's 18%. Major faults per token go 126 -> 1427 for a source that
allocates no draft context at all, and that is the rollback snapshots: n_rs_seq =
draft_max is asked for by ANY speculation, since rejecting a draft means rewinding the
KV, and on a hybrid attention/SSM model that snapshot is a real allocation scaling with
the context. The n-gram source escapes MTP's draft context but not the loop's own
memory, and on device that memory is the expert cache's.

The same run re-measured MTP with the thermal confound removed -- 3.64 effective
against 4.43, so the earlier device verdict was not an artefact of benching without a
cooldown gate -- and reproduced the flash split at 3.7% head against 96.3% widened
verify batch, matching the host's 2.9-3.0%.

Byte-identity gates pass; speculation stays out of them for the reason docs/mtp.md
gives.

The app's CSV configuration surface follows: the three new preamble keys get their own
glossary entries rather than falling through to the unexplained-key renderer, and the
draft source joins the short run label. A speculated run is not the same KIND of run --
under speculation a decode confirms a whole group, so its per-token rows are not even
accounted the same way -- and two compare legends differing by it must not read alike.
2026-08-02 00:09:39 +02:00
Helldez
ad13038ce0
feat(moe): --io-two-wave, publish a layer's read batch in two waves (#128)
A cold layer's batch became visible to the I/O lanes only after every
miss took its page commits - up to three vm_commit syscalls per cold
expert of bookkeeping sitting in front of the first byte of I/O, which
is the latency-to-first-slice the sidecar refutation (PR #90) identified
as the binding constraint. (#118)

With the flag on, only the first present projection - the one
mul_mat_id blocks on first - is committed up front; its jobs publish
and wake the lanes immediately, and the remaining projections are
committed and appended while the lanes already read.

The drain protocol grew the one thing this needs:

- io_drain copies each job out under the lock, so jobs_ growing (and
  possibly reallocating) mid-batch cannot leave a worker holding a
  dangling reference;
- the worker wait predicate admits next_idx_ < batch_njobs_, so a
  worker that drained wave one and left comes back for a batch that
  grew in the SAME generation - the gen comparison alone never would;
- a wave-two commit failure goes fatal and wakes the ready waiters,
  because wave one already published flags this batch will never flip.

Batch completion cannot fire between the waves: the only thread that
waits on done_cnt_ == batch_njobs_ is the eval thread, and it is the
one appending wave two.

Overlap + LRU cache only (validate() enforces both); recorded in the
CSV preamble as io_two_wave. Default off: the win is bounded by the
commit cost per cold layer, and the failure mode of a drain-protocol
bug is a hang the host cannot reproduce - so a new gate (G4d) holds
two-wave output byte-identical to serial streaming, and the flag stays
off until the on-device A/B (#120) says the win is real.
2026-07-28 15:46:52 +02:00
Helldez
cc4aed1304
perf(cli): BMOE_PROGRESS carries the answer as a delta, not cumulatively (#127)
Every per-token line repeated the whole answer and reasoning so far, so
a generation of n tokens wrote, JSON-escaped and made the app parse
O(n^2) bytes - megabytes of pipe traffic to deliver a few kilobytes of
text on a reasoning model (#119).

The line now carries delta_reasoning/delta_text - the tail since the
previous line - and the reader appends. A pure append-only protocol
cannot express the one thing common_chat_parse does retroactively:
when a closing think tag arrives, text already reported as answer
becomes reasoning. That case falls back to a full snapshot with
"reset":1, and the reader replaces instead of appending.

Both emitters (one-shot --progress and --session) share the single
format string, so they changed together; the app's TelemetryParser
accumulates in StringBuilders (appending to a String re-copied the
whole answer per token) and resets them with the generation. The full
final text still travels in BMOE_DONE, untouched.

The engine-side re-parse per token remains (common_chat_parse cannot
resume); this removes the pipe, escape and app-parse cost, which is
what loop_overhead_ms can now see. On-device numbers are part of the
#120 A/B.

Closes #119.
2026-07-28 15:38:34 +02:00
Helldez
c9c5d37367
feat(metrics): record the whole run configuration, and what tok/s omits (#116)
Two gaps, both about a benchmark file being unable to explain itself.

The CSV preamble had eighteen keys and had fallen behind several releases.
Missing: n_ubatch, which sets the compute-buffer reservation and therefore
moves the very memory columns printed underneath it; predict_log,
predict_spec_max and prefetch_sync, so a probed run was indistinguishable from
a benchmark run the docs explicitly say it is not; drop_renorm and
drop_prefill, which change how much mass dropping discards; cache_floor_mb, the
input behind an auto-sized cache_mb; load_all, whose read_bytes mean something
else entirely; every sampling parameter, so a stochastic run read as a greedy
one; and compute_trace_layers. All are recorded now, under a bumped
"# bmoe_metrics v2" banner.

`think` stays out on purpose: it is a property of a request, not of a session,
so one value in a session-wide preamble would be wrong for every turn that
asked for the other.

The file also now names the build that wrote it. There was no version string
anywhere in the engine — the CMake project had none — so a committed CSV could
only be dated by the commit that copied it in. Added as a project VERSION, a
BMOE_VERSION define, bmoe::version(), `bmoe-cli --version`, and engine= in the
preamble. It sits on its own line so the model= line still BEGINS with model=,
which is how the app's CSV reader finds a run's name; the app's lookup is made
order-independent too, so the next key to be appended cannot break it again.

Second gap: wall_ms brackets llama_decode and nothing else. That is what makes
compute_ms a clean residual, and it also means sampling, detokenization,
rendering and the sink writes are outside wall_ms, outside gen_seconds and
outside the reported tok/s. Work moved into or out of that region was
unmeasurable by the number the project optimizes. loop_overhead_ms now reports
it per token, and loop_overhead_s/tok in the summary closes the accounting with
the tail after the last token that no row can carry.

docs/telemetry.md documents the preamble — it specified every other # block but
not this one — the new column, and the stall_ms divisor: stall is a per-thread
mean, so a stall that is not simultaneous across threads is under-stated and
compute_ms absorbs the difference.

Gates 7/7; preamble, column and --version verified against the tiny gate model.
Both readers checked: scripts/bench-analyze.py reads columns by name and skips
unknown # lines; the app's Csv.read keeps unknown keys and picks the new column
up automatically.
2026-07-28 11:49:34 +02:00
Helldez
523413e974
perf(engine): build a token's rendered view only when something reads it (#115)
Producing TokenMetrics::text means parsing the entire generation so far —
common_chat_parse takes the whole string and cannot resume — so the cost is
O(n) per token and O(n-squared) across a turn, growing as the answer does. Off
the chat path it was not free either: shown_view returns the raw string, so
every token copied everything generated so far into the metrics struct.

It ran for every token of every run. The plain CLI path writes m.piece and
never looks at m.text; a benchmark run reads neither. Only the BMOE_PROGRESS
line protocol actually renders it.

GenerateRequest::render_text now carries that decision. The session loop sets
it (the protocol puts the parsed answer on every line); run() ties it to
--progress; it defaults to true so an embedder that has never seen the flag
behaves exactly as before. Note this cost sat OUTSIDE the wall bracket that
feeds gen_seconds, so it never showed up in the reported tok/s while being paid
on every benchmark run.

The end-of-turn parse was also done twice — once for RunResult::generated_text,
once to commit the assistant message to history — over the same string with the
same parameters. Parse once, use twice; the prefilled-turn and parse-failure
fallbacks are unchanged and now stated in one place.

Also: three back-to-back stats() snapshots to read three fields become one, and
--compute-trace / --io-trace are actually passed to the session loop. The CLI
parsed those flags, opened the files, wrote their headers and then dropped both
sinks when entering --session, so the user got an empty trace and no
explanation.

Gates 7/7. Both CLI output paths verified by hand on the tiny gate model:
--progress still accumulates the parsed text per line, the plain path still
streams pieces.
2026-07-28 11:48:52 +02:00
Helldez
9e04e04d53
feat(moe): measure how predictable expert routing is, then act on it — and refute it (#101)
* feat(moe): --predict-log, measure how predictable expert routing is

Temporal prefetch was built on a predictor nobody had priced. The
previous-token bet turned out to be right ~38% of the time on Qwen and
~18% on gpt-oss, which cannot pay for the reads it speculates. The
lesson was not that prefetch is impossible but that a predictor should
be measured before it is wired into anything. This adds the instrument,
not a policy.

--predict-log ranks each layer's experts a layer early, by running the
NEXT layer's router matrix on the CURRENT layer's gate input. The
residual stream barely moves between layers, so the stale input ranks
nearly as the real one will -- the mechanism FATE (arXiv 2502.12224)
reports 78.8% for. It is training-free and changes no model: the router
matrix is a dense weight already resident, and the prediction is one
GEMV per layer. The matrix is learned from the graph (the gate matmul's
first source) rather than looked up by tensor name, so no architecture
is named anywhere in the path; being a weight leaf, its pointer stays
valid into layers the current token has not reached.

Three predictors are scored against the routing the router actually
produced, so they are comparable on one run: the stale gate, the
previous-token bet --prefetch already places, and a zero-staleness
control. The control is the load-bearing part. It shares every line of
code with the prediction under test and differs only in using the
layer's own matrix, so it must reproduce the selection llama.cpp
computes from those same two tensors. A transposed matrix, a mis-strided
row or the wrong token of the batch collapses it toward chance while the
stale figure would stay superficially plausible; an architecture that
selects by something other than raw-logit ranking (an additive bias,
group-limited routing) puts it below 100% and by that much the stale
figure understates the method. The CLI says so rather than letting the
gap be blamed on staleness.

Reported per layer as well as in aggregate, because an aggregate
flatters a prefetch: what a prefetch costs is set by the layers it gets
wrong, and a MoE model's first layers route far less predictably than
its last. Denominators are printed per predictor -- the stale gate
structurally cannot rank layer 0 (nothing precedes it) or the first
token of a run, and those routings are counted as unscored rather than
folded in, since a routing that was not ranked is not a wrong guess. A
predictor with no routings at a layer prints "-", never 0.0.

Diagnostics only: nothing it computes reaches load_layer, the cache or
the graph, so a probed run reads exactly the bytes an unprobed one does.
G9a gates that byte identity and G9b gates the control, which reads
100.0% on the tiny model. It is not free -- one isolated node and two
GEMVs per layer on the eval thread -- so a probed run is not a benchmark
run, and it requires --moe-stream since routing does not depend on how
the weights reached memory.

docs/expert-prediction.md also records the caveat the number will need:
a high score would say the routing is knowable earlier, not that knowing
it earlier makes decode faster. On a flash already saturated, starting a
read sooner adds no bandwidth -- which is why prefetch, layer-LFU and
the expert sidecar all lost despite improving the metric each was
designed around.

* feat(moe): --predict-prefetch, speculate on the stale-gate prediction

The probe said the routing is knowable a layer early (~89% of routed
slots on a 128-expert model, vs ~43% for the previous-token bet the
temporal prefetch acts on). This wires that prediction into the existing
speculative read path -- same cache buffers, same accounting, same
settle, same moe-prefetch summary line (tagged [stale-gate]) -- so the
only thing that changes is which guess rides the idle lanes.

Two decisions carry the design:

Speculation is issued AFTER the current layer's load, not at prediction
time. Every load path begins by quiescing speculation, and a layer's own
load sits a few graph nodes after its gate matmul -- reads queued at
prediction time would be cancelled before a lane picked them up. Each of
the three load sites (plain topk, deferred drop, drop fallback) issues
the pending next-layer prediction right after its load_layer, restoring
the same read-ahead window the temporal prefetch gets. On the tiny-model
gates this is the difference between 0% and 27% of speculated experts
proving useful -- the latter matching the probe's measured accuracy on
that model, which is the accounting agreeing with itself.

It is drop-aware. With --drop-cold-experts armed, a predicted expert
whose predicted routing weight (softmax over the predicted top-k)
falls below the drop threshold is not speculated: if it misses, the
policy discards it unread, so reading it ahead would spend the exact
I/O the policy exists to save. The top prediction is always kept,
mirroring the policy's own pin of the top-weighted expert. The known
interplay is inherited from the temporal prefetch and deliberate: a
correct guess un-drops an expert, buying quality at the same threshold
rather than speed.

The routing width the prefetch predicts at is learned from the topk
node, not read from config, so an --n-expert-used override stays honest
with no extra plumbing. Mutually exclusive with --prefetch (two
predictors would double-speculate the same future); requires the LRU
cache; decode only. The control GEMV remains probe-only, so the
production path costs one gate GEMV per MoE layer per token.

Gates: G10a proves byte-identity through the speculative path
(prefetch-sync, forced small cache, hits and evictions both occur);
G10b proves the run actually speculated and that useful-hit accounting
tracks the probe's accuracy. Off by default, pending an on-device A/B.

* feat(moe): cap predictive speculation at the top 2 predicted misses per layer

Speculating the whole predicted routing was measured on device at -38%
against its own baseline despite every intermediate metric improving
(hit 77.6->88.9%, stall 32->9 ms, 84% of speculations useful): the +33%
flash bytes and the vm commits fighting a full cache (major faults x3)
cost several times the stall removed. The stall a prefetch can remove is
head-of-line only -- overlap already hides the tail behind the expert
matmul -- so the cap keeps the part of the bet that can pay and drops
the part that provably cannot. Residency-aware: a predicted expert
already in cache does not burn a slot of the cap.

* perf(moe): rebuild the predictive prefetch around its measured costs

The observer-tax run priced the naive implementation: ~35-45 ms per
GEMV pass (ggml_fp16_to_fp32 is a function call per weight element --
21M calls/token) and ~20 ms/token for the extra isolated node, against
a speculation machinery that costs ~15-25. The GEMV and the barrier
were the feature; this commit removes both.

- gate_scores converts F16 natively on aarch64 (one instruction,
  vectorizable) instead of a function call per element.
- The prefetch no longer isolates the gate matmul: the ask pass hands
  over its source pointers for free, and the gate-input row is read at
  the topk callback with no barrier of its own. A sampled watchdog
  (the zero-staleness control, every 512 routings) validates the
  barrier-less read and disarms the prefetch out loud if the memory
  planner ever reuses that buffer -- without it, a future llama.cpp
  bump could silently turn the predictor into a noise generator.
- The GEMV runs on a dedicated worker at a TWO-layer horizon: one
  layer ahead has no landing spot (the callbacks between a layer's
  gate and its own load are microseconds apart), while at l+2 the
  worker has a whole layer for a ~0.5M-MAC job and the result inherits
  the same post-load issue window as before. The probe now also scores
  stale-2, so the extra layer of staleness is priced per model rather
  than assumed.
- The prediction's residents are RETAINED (new IExpertSource::retain:
  move-to-MRU, deliberately not a cache hit so the hit-rate metric
  stays honest) -- protecting a predicted expert costs zero bytes,
  unlike prefetching it. Only predicted misses are speculated, still
  capped at 2.

The probe path keeps its barrier and both GEMVs: a probed run is
diagnostics, and its job is to be right, not fast.

* feat(moe): --predict-spec-max N — how much flash the prediction may spend (0 = retention only)

The cap was a constant; the retention-only point (0) is the config the
whole experiment now hinges on -- the prediction protecting predicted
residents from eviction while spending no flash at all -- and a
measured constant that cannot be varied is not a mechanism. Validated
[0, 8]; retention happens at every value because it is free.

* docs(predict): record the 2026-07-23 campaign — accuracy confirmed, throughput verdict open

Accuracy on device: stale-gate 88.6% (Qwen3-30B) / 80.7% (Qwen3.6),
control exactly 100.0% on every layer of both, prev-token 43/35% --
corroborating the offline route-trace estimates.

Throughput: the day's C-vs-B losses are recorded WITH their
invalidation. Re-running the reference on the by-then-hot device gave
3.93 tok/s against the cool morning's 6.54 with byte-identical I/O,
hit and drop counts -- the engine is deterministic, the -40% was
silent thermal capping, and every variant had been compared against
the cool number. The one thermally matched pair that was measured
(speculation on 8 io lanes, -28%, effective flash bandwidth 585->392
MiB/s) kills the more-lanes hypothesis specifically; spec-max 2 and
retention-only still owe a matched cool pair.

What did survive: the observer-tax decomposition (109 ms/token: the
per-element exported-function F16 conversion at 21M calls/token, plus
the barrier), and retention moving hit rate 0.1pp on a 3000 MiB cache
-- the offline replay bound confirmed from inside the engine.

* feat(app): expose the predictive prefetch as an experimental Streaming toggle

Off by default. Gated on streaming + a live cache like the temporal
prefetch, and the two settings disable each other in the UI -- the
engine refuses the pair, and a control the engine will reject is worse
than one that cannot be set. The spec-max rung selector (0/1/2/4)
surfaces the retention-only point, which is the configuration the open
throughput question most needs measured from the app. Session
signature includes both fields so flipping them reopens the process.

No versionCode bump: this is a PR-branch test build, not a release.

* docs(predict): record the 2026-07-24 matched pairs — read-ahead refuted, retention hit-neutral

A four-cell session (B, retention-only, spec-2, B sentinel) run at fixed 30 s
spacing re-proved the thermal-contamination mechanism (sentinel −17%, clusters
silently capped from cell 2) and yielded one genuinely matched pair: spec-2 vs
the B sentinel at the same caps and battery temperature, 3.14 vs 3.96 tok/s
(−21%) with hit rate up 4.3pp and 79% of speculations useful. Speculation
improves every metric it owns and still loses the wall clock — the flash has no
spare bandwidth to spend. Retention-only again moved the hit rate by nothing
(77.2% vs 77.6%), as the offline replay bound predicted.

* feat(app): contrast the two prefetch predictors in the UI, default spec-max to 0

The temporal and predictive toggles now say what actually differs — the bet
("repeats the previous token", ~40%) vs the question ("ask the next router a
layer early", ~85%) — instead of describing mechanisms side by side. Spec-max
defaults to 0 (retention only): the matched-pair A/B showed the read-ahead
losing −21% on a saturated flash, so 0 is the only rung the measurements did
not refute, and the helper text says so.

* chore(predict): file the changelog under Unreleased, drop a dead member, record the verdict

Three loose ends found reviewing the branch for merge:

- The changelog entries had been appended to the already-released 0.15.1 section;
  they belong under [Unreleased], where the release commit carves them out.
- nu_hint_ was written on every routing and read nowhere: a leftover of the
  synchronous first design, whose successor passes the routing width straight into
  the prediction job.
- docs/roadmap.md still closed the routing-prediction question on the 2026-07-12
  removal. It now records what reopening it with a training-free predictor found:
  the accuracy is real and the throughput is not, for the same reason more lanes
  and the sidecar lost.
2026-07-28 09:21:41 +02:00
Helldez
6441494b76
feat(moe): warn when cache-aware dropping meets a narrow routing (#99)
The threshold is a fraction of the uniform share 1/top-k, so what it
removes scales with how wide the routing is. At top-k 8 -- where every
number in docs/expert-dropping.md was collected -- 0.75 means "below 9.4%
of the routing", a tail trim. At top-k 4 it means "below 18.8%", and at
top-k 2 "below 37.5%", which on a miss discards the whole minority
expert: closer to halving the routing than trimming it, and unmeasured.

The engine now says so once at load when dropping is armed and the
effective top-k is 4 or fewer, quoting the actual share for the model in
hand. It warns rather than clamping or refusing: it cannot know whether
that trade is acceptable for a given model and task, and silently
adjusting a number the caller chose would be worse than a loud caveat.
MoeStreamConfig::drop_low_topk_warn is documented as an EVIDENCE
boundary, not a physical one -- nothing in the streaming path reads it.

The app shows the same caveat inline under the setting, in the error
colour, computed from the width the loaded model reports rather than
assumed -- and only once a model is loaded, since guessing would be worse
than staying quiet. gpt-oss is the case this exists for: it routes 4 of
128, and the app default is 75%.

To make that possible, BMOE_READY gains n_expert_used (the effective
width after any override, 0 on a non-MoE model) and Session exposes
n_expert_used() for embedders. Additive: older consumers ignore it.
2026-07-23 10:08:25 +02:00
Helldez
bea5a0b99e
feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95)
* feat(moe): --drop-cold-experts, spend quality only where it buys I/O

Turbo top-k drops the tail of a routing whether or not those experts were
already in RAM. A resident expert costs no flash read, so that trade pays
quality for nothing on the ~80% of decode routings that are cache hits.

This skips a routed expert only when it is a cache MISS and the router
weighted it below frac x (1/top-k). Replayed over the committed route
traces at frac 1.0, decode phase, that avoids 66% of flash reads for 9.5%
of the router's weight mass, where --n-expert-used 5 avoids 23% for a
comparable 10.6% -- about 3x the reads at the same quality cost.

Implementation. The decision needs the FINAL router weights, which arrive
several nodes after the topk where the streamer normally loads, so
load_layer() is deferred to the terminal node of the layer's weight chain.
Which node that is depends on the model's gating, so the hook learns it
from the graph rather than carrying an architecture table; if it fails to
arrive the hook forgets it and re-learns rather than re-betting. A dropped
slot has its weight zeroed and its expert id repointed at the routing's
top-weighted expert: an unread expert can sit in reserved-but-uncommitted
VM and mul_mat_id would touch it anyway, so the kernel is given memory
that is certainly resident and multiplies it by exactly zero.

Requires the LRU cache -- with --cache-mb 0 residency reads all-miss and
the policy would silently degenerate into an unconditional weight cut.
Prefill is excluded by default. The top expert is always pinned, so no
routing can be emptied at any threshold.

Gates: G8a/G8a' prove the deferral and the learned terminal node are
transparent (byte-identical output, zero drops, at a threshold below any
producible weight); G8b that full strength against a constantly-evicting
cache never reaches an unloaded slot; G8c that at top-k 1 dropping is a
no-op, pinning both the top-expert guarantee and the threshold tracking
the effective top-k.

Three existing metrics shift meaning under dropping and the docs now say
so: cache_hit_pct rises without the cache serving more (a dropped routing
is a miss that is never looked up), and token/layer_demand measure what
was staged rather than routed. prefetch.md's "cannot change output" is
scoped, limitations.md gains the non-reproducibility entry, and
benchmark-method.md warns that reversing the run order cannot distinguish
a moved drop rate from a contaminated cell.

Off by default in the CLI and in the app. The output is not reproducible
-- what gets dropped depends on what the cache held -- so it carries no
rows in the README tables, and switching it on by default waits on a
published on-device A/B rather than on the replay argument alone.

* feat(app): default cache-aware dropping to 75%, measured on device

Qwen3.6-35B-A3B (top-k 8 of 256), in-app, cache 3000, one variable
changed: 2.549 tok/s off, 3.938 at F=0.75 (+55%), 4.702 at F=1.0 (+84%),
with flash reads falling 248 -> 163 -> 48 GiB. Per-token bootstrap
intervals separate every pair except off vs 0.50, which overlaps -- at
half the uniform share the policy drops 2.7% of routings and buys
nothing, which doubles as a negative control that the machinery is free
when it does not fire.

Run order was 1.0, off, 0.5, 0.75, so the two fastest cells are the first
and the LAST; thermal drift would have made the last the worst. The
mechanism orders by threshold even though the run order does not.

The replay turned out conservative rather than optimistic. It is
documented as an upper bound because it cannot model the cache changing
in response to dropping: at F=0.75 it was accurate (37% predicted, 34%
measured), at F=1.0 it understated (66% predicted, 81% measured). Avoided
reads free cache capacity, which raises the hit rate, which leaves fewer
misses to drop.

75% rather than 100% is deliberate: it takes the larger part of the win
for half the discarded routings (14% against 28%). Quality is still
unquantified -- no perplexity number and no side-by-side exists -- so the
conservative end of a measured range is the defensible default. The CLI
stays off; the byte-identity gates need a deterministic default.

Also records cache_hit_pct rising 67.8 -> 90.7% as the documented
accounting artefact rather than the cache serving more, and majflt/token
as dominated by each run's starting memory state, not by the threshold.
2026-07-22 17:21:55 +02:00
Helldez
79611654cd Revert "feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#94)"
This reverts 45a90a2.

The feature merged before the evidence for its shipping default did. The
replay numbers argue the shape of the trade is favourable, but no on-device
A/B is published in this repository, and the app default it landed with
(75%) changes model output for every user of the demo app -- and changes it
non-reproducibly, which no other setting in this engine does.

Nothing was found wrong with the code. This is a sequencing decision: the
work returns as a pull request, with the app default back to off, so the
measurement lands before the default does.

Reverted rather than force-pushed: main is public and this commit was
already pushed, so the history stays honest about what happened.
2026-07-22 15:47:14 +02:00
Helldez
45a90a2df5
feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#94)
* feat(moe): --drop-cold-experts, spend quality only where it buys I/O

Turbo top-k drops the tail of a routing whether or not those experts were
already in RAM. A resident expert costs no flash read, so that trade pays
quality for nothing on the ~80% of decode routings that are cache hits.

This adds the cache-aware version: skip a routed expert only when it is a
cache MISS and the router weighted it below frac x (1/top-k). Replayed over
the committed route traces at frac 1.0, decode phase, that avoids 66% of
flash reads for 9.5% of the router's weight mass, against 59%/37% for
--n-expert-used 3 — about 3x the reads avoided at a comparable cost. The
threshold is a curve, not a switch: 0.75 trades 4.4% of the mass for 37% of
the reads, better than --n-expert-used 5 on both axes.

Implementation. The decision needs the FINAL router weights, which arrive
several nodes after the topk where the streamer normally loads, so with the
policy armed load_layer() is deferred to the terminal node of the layer's
weight chain. Which node that is depends on the model's gating, so the hook
learns it from the graph rather than carrying an architecture table; until
it is known a layer loads at its topk node undropped. A dropped slot has its
weight zeroed and its expert id repointed at the routing's top-weighted
expert: an expert we decline to read may sit in reserved-but-uncommitted VM
and mul_mat_id would still touch it, so the kernel is given memory that is
certainly resident and multiplies it by exactly zero. Survivors are rescaled
by default, since a systematically shrunk expert output perturbs the residual
stream more than the missing contribution does.

Prefill is excluded by default (cold cache, ~4x the weight mass discarded,
and compute-bound anyway). The largest weight in a routing is always at least
the uniform share, so frac <= 1 can never empty a layer; validate() enforces
the bound and the top expert is pinned regardless.

Gates: G8a proves the deferral and the learned terminal node are transparent
(a threshold below any producible weight leaves the output byte-identical),
G8b that full strength with the cache off never reaches an unloaded expert.

Unlike every other knob this one is state-dependent: what gets dropped
depends on what the cache held, so output is not reproducible across runs.
Off by default, not in the app's settings, and NOT yet measured on device —
the numbers above are a static replay and an upper bound. docs/expert-
dropping.md states what is owed before it is recommended anywhere.

* feat(app): expose cache-aware expert dropping in Settings

Speed / quality -> Drop cold experts, as a percentage of the uniform
share (off / 50 / 75 / 100). The engine takes a fraction; the app stores
integer rungs, so the setting divides by 100 on the way to the flag.

Disabled in mmap mode: the policy asks the expert source what is resident,
and there is no expert source without the streamer. Included in the session
signature, so changing it reopens the session rather than being ignored by
a process already loaded.

Off by default. This exists so the A/B can be run where the engine actually
ships -- through the app, not a pushed CLI binary.

* fix(moe): require the cache for dropping, and correct what it reports

Review of the first two commits found the policy could be armed in a
configuration where it is not cache-aware at all, and that two of the
numbers it reports were wrong.

- Require the LRU cache. With --cache-mb 0 query_residency answers
  all-miss, so the policy silently degenerated into an unconditional
  weight cut -- exactly what --n-expert-used already does, under a flag
  claiming to consult residency. validate() now rejects it, as it already
  did for --prefetch, and the app gates the setting on the same condition.

- Fix experts_routed. It was incremented inside apply_drop, so it counted
  what the policy examined rather than what the router selected: layers
  before the terminal weight node is learned, and every un-armed phase,
  were missing from the denominator. The reported drop rate was a fraction
  of the wrong thing.

- Re-learn instead of re-betting. If the node learned as terminal does not
  arrive, the deferral now also forgets it, so the next graph loads at the
  topk node while it re-learns. Deferring again on a stale guess would
  repeat the fault every token against a graph that had moved.

- Point the gates at a real cache. G8a/G8b ran with the cache off, where
  the shared-slot path has no reserved-but-uncommitted memory -- so the id
  repointing, which is the design's whole safety argument, was never
  exercised. They now run against a constantly-evicting budget. Adds G8a'
  (asserts routings were examined and none dropped, so an inert-threshold
  flake fails legibly) and G8c (at top-k 1 dropping is a proven no-op,
  pinning both the top-expert guarantee and the threshold being taken
  against the effective top-k).

Docs: three metrics change meaning under dropping and none of them said
so. A dropped routing is a miss that is never looked up, so cache_hit_pct
rises without the cache serving more, and token/layer_demand measure what
was staged rather than routed -- documented in telemetry.md, pressure.md
(size the cache with dropping off, then turn it on) and metrics.h.
prefetch.md's "cannot change output" is scoped: under dropping a correct
guess un-drops an expert. limitations.md gains the non-reproducibility
entry, benchmark-method.md the axis plus a warning that reversing the run
order cannot distinguish a moved drop rate from a contaminated cell, and
architecture.md/runtime.h no longer claim unconditional determinism.
Fixes two anchors the README rename broke, and a changelog sentence that
quoted the equal-I/O row while drawing the equal-quality conclusion.

App: Drop cold experts defaults to 75%. The default is a product decision
taken on the maintainer's device; no benchmark for it is published here,
and docs/expert-dropping.md says that plainly instead of implying a
measured figure. The CLI stays off by default -- the byte-identity gates
need a deterministic default.
2026-07-22 15:08:49 +02:00
Helldez
3e6f657133 fix(chat): require evidence a model reasons before calling it uncontrollable
think_ctl reported none for any model whose template ignores enable_thinking while its handler
declares reasoning tags. But handlers publish that tag pair for a whole family, not per model: the
non-reasoning members advertise a <think> they never emit. LFM2-8B-A1B and LFM2.5-Instruct came out
uncontrollable, so the app disabled their Thinking switch and explained it with "this model always
reasons" — about models that never reason at all. A control that was merely moot became a control
that was taken away with a false reason.

none now requires positive evidence of both halves: the model declares a reasoning span AND its
template actually uses it. That second test is the one llama.cpp itself applies before wiring up
reasoning extraction for the family, so it costs nothing and stays model-supplied.

A second path had the same flaw: a template where the continuation hook does nothing also fell into
none, which caught any plain non-reasoning template. Reordered so the span question is asked first
and none is reachable only through it. Everything with nothing to suppress reports template — what
these models did before any of this existed.

Verified: LFM2.5-8B-A1B still none on-device. Host gates pin LFM2-8B-A1B, LFM2.5-Instruct, a plain
template and Gemma 4 (which reads the flag and never reaches the tag test) as template.
2026-07-19 21:37:48 +02:00
Helldez
4f5d4c2246 fix(chat): only prefill past reasoning the model cannot decline
The prefill landed on every model whose template ignores enable_thinking. On-device that made
LFM2.5 strictly worse: handed a pre-closed empty reasoning span it reasons straight past it and
emits the reasoning UNTAGGED into the answer, where before it at least parsed into a collapsible
block. 500 tokens without reaching an answer.

The two families were never the same case. gpt-oss declares no reasoning tags because reasoning is
a channel the format separates structurally, so starting the turn past it is not something the
model can decline. LFM2.5 declares <think>/</think>: the span is the model's own to open and close,
so handing it an empty closed one is a suggestion, and this model was not trained on that
convention. Liquid ships a separate non-reasoning checkpoint rather than an off switch.

So the probe now reads that distinction off the tags the model itself declares — no model names,
no template string matching. Tags declared and the flag inert means the request cannot be honoured:
report none, disable the control, say why. No tags means the prefill binds.

Measured on-device, three models:
  LFM2.5-8B-A1B   none      reasoning tagged into reasoning_content, answer clean ("4")
  gpt-oss-120b    prefill   no reasoning, direct answer (the retired hardcoded path's behaviour)
  Qwen3.6-35B-A3B template   unchanged, no reasoning, direct answer
2026-07-19 21:08:29 +02:00
Helldez
aa6a7fdafd fix(chat): honour "thinking off" on templates that ignore enable_thinking
Turning Thinking off set the template variable enable_thinking and stopped there. That
variable is only a request to the model's own chat template, and a template is free to
ignore it. LFM2.5's never reads it, so the rendered prompt was byte-identical with thinking
on and off, the model reasoned anyway, and nothing reported that the setting had been
dropped (#82).

Detection is measured, not assumed: at open() the template is rendered with the flag on and
off and the two prompts compared, then rendered once more with a continuation. That answers
the only question that matters — does this template react — for any model in any language.
common_chat_templates_support_enable_thinking cannot answer it: per handler it is a
hardcoded literal reporting "this model can reason", not "this template reads the variable".

Enforcement uses llama.cpp's continuation hook: a synthetic trailing assistant message with
continue_final_message makes upstream's per-template handler render that family's own
"reasoning is over" span into the prompt. The model resumes at the first token of its answer
with the reasoning already behind it. This is what a template that implements the toggle
natively does (Qwen3 renders <think></think> for enable_thinking=false), so it needs no
cooperation from the template and no sampler — it works on the greedy path the byte-identity
gates run on.

Because the span comes from upstream's handler, the engine names no markers of its own. That
retires the two hardcoded harmony strings in the decode path: priming gpt-oss to answer
without reasoning was a literal "<|start|>assistant" suffix test and a literal
"<|channel|>final<|message|>" appended to the prompt. gpt-oss now takes the same generic path
as every other family and resumes at the same point, so a submodule bump that changes those
markers needs no engine change.

Models where neither mechanism exists are reported rather than fought: BMOE_READY gains
think_ctl (template | prefill | none) and the app shows the Thinking switch disabled, with
the reason, instead of offering a control that does nothing.

Deliberately not done: forcing the reasoning block closed on logits. Measured on-device in
the closed PR #83, it made LFM2.5 strictly worse — the model reopened the block, then
abandoned the tags and reasoned in plain prose into the answer. Suppression belongs in the
prompt, before the model commits to reasoning, not mid-generation.

tests/think_control_test.cpp pins all three regimes against the vendored templates with no
model: Qwen3 as template, LFM2.5 and gpt-oss as prefill with the span asserted CLOSED, a
plain template as none, and fail-open when the probe cannot run.
2026-07-19 20:26:22 +02:00
Helldez
3db882d278 feat(trace): add layer-granularity compute trace (--compute-trace-layers)
The per-node compute trace pays ~3000 barriers per token, which serializes
the graph against the expert stream: on a model that streams heavily the
trace mostly measures its own serialization (Qwen3-30B: 9.4 s/token traced
vs 0.39 untraced), so its absolutes cannot be compared across models.

Layer granularity isolates only the first node of each layer (~n_layer
barriers per token). Operator coalescing and the async expert prefetch
survive, so the traced numbers stay close to an untraced run. Rows share
the per-node schema with op LAYER: name blk.<il> aggregates one layer''s
segment, pre the embedding lookup, post the last layer''s tail plus
final norm and LM head (closed by the session right after llama_decode,
since the tail has no successor boundary to observe it).

The granularity flows RunConfig -> SessionConfig -> RouterHook; the routing
nodes the streamer isolates anyway also close a segment, a barrier that
exists untraced too. decode-analyze.py detects the granularity and prints
the per-segment table.

Gates: all 6 pass (byte-identity qwen3moe + gemma4). Smoke-tested on the
tiny-moe models with streaming on.
2026-07-19 09:21:39 +02:00
Helldez
55b8579396 feat(core): surface the reasoning span alongside the answer (#70)
Since #49 wired the chat parser correctly, a reasoning model's thinking is
stripped from the shown answer unconditionally — including when the user asked
for thinking. With the in-app Thinking toggle ON, the answer area stays blank
while the model reasons and only the final answer ever appears; on a slow
streamed decode that reads as a hang for the whole thinking span.

The reasoning was being parsed and thrown away: shown_text() kept only
common_chat_parse(...).content and dropped .reasoning_content, and no field
downstream carried it. Surface it instead of discarding it:

  - TokenMetrics gains `reasoning`, RunResult gains `reasoning_text`; session's
    shown_view returns {content, reasoning} from a single parse (the partial
    parse already fills reasoning_content incrementally, so it streams).
  - The line protocol carries a `reasoning` field on BMOE_PROGRESS and
    BMOE_DONE, kept apart from `text` so a UI can render it as a distinct
    thinking block rather than inline. Documented in docs/telemetry.md.

The answer in `text`/generated_text is unchanged (reasoning still stripped),
so this is display-only: the byte-identity gates are untouched. The chat-parse
host gate now also asserts a partial parse exposes the reasoning span — the
contract the live thinking block depends on.
2026-07-18 22:43:23 +02:00
Helldez
1b00e2ffb9 docs(telemetry): note the compute_ms clamp breaks the wall-additive identity
compute_ms is clamped at 0, so wall = compute + flash + mgmt does not hold in
the pathological overlap case. A consumer inverting the residual to recover the
flash-wait term (wall − compute − mgmt) over-attributes to flash when the clamp
fires; read io_ms (serial) / stall_ms (overlap) directly instead. Documents the
sharp edge raised in #47 rather than adding a redundant flash_wait_ms column.
2026-07-18 11:01:27 +02:00
Helldez
75b5a634fc refactor(moe): retire the adaptive cache governor; fixed LRU + one-shot auto
Measured net loss on >RAM models (cache-off is the ceiling), so the runtime governor, --cache-dynamic/--cache-gov2, and the sense/resize loop go. Kept: the fixed --cache-mb N LRU, --cache-mb auto sizing once at load, cache-off default. Telemetry aligned (dense_resident_frac made live; resident_frac/cache_cuts removed). App: one 3-way dense-weights selector. Host+ctest green.
2026-07-17 08:55:18 +02:00
Helldez
9a1d1f8a1c feat(metrics): on-device memory telemetry, pressure sensing, and an adaptive cache governor
The measure-your-own-memory series: fault/CPU decomposition, the anon/file RSS split, mincore residency sensors, the Android metrics screens, and a --cache-dynamic governor that sizes the expert cache to what the device concedes. (The governor is retired further down this history; the telemetry stays.)
2026-07-17 08:54:06 +02:00
Helldez
f9e408f542 feat(metrics): --compute-trace and --io-trace decompose the decode
The per-token CSV reports compute as a residual (wall - io - mgmt), so everything the
engine does not itself clock is pooled into it: page faults, scheduler stalls, and the
matmuls. A residual cannot say which. That is the whole reason gpt-oss-120b reads as
"compute-bound" at 1.7 s/token while faulting 8.4k pages per token — the flash wait is
billed to compute because it happens under llama_decode.

--compute-trace measures it instead. Asking the eval callback to isolate a node makes
ggml compute exactly up to it and synchronize, so the wall delta between consecutive
boundaries is that node's real compute time; sampling major faults across the same
boundaries attributes the >RAM stall to the node that paid it. Still no llama.cpp patch:
this rides the public cb_eval ABI, whose ask/no-ask contract already specifies the
isolation. It costs a barrier per node and forbids operator coalescing, so it is a
diagnostic — a traced run is not a benchmark run, and only the shares are meaningful.
Unlike the other traces it does not need --moe-stream: it times the graph, so a dense
mmap baseline can be traced and compared against a streamed run.

--io-trace records one row per pread: latency, requested vs aligned bytes, lane, and the
(layer, expert, projection) it serves — values already computed at every enqueue site and
until now discarded. This is where the flash floor is: the aggregate 760 MiB/s sits far
below the drive's sequential ceiling because routed slices are scattered, and the trace
says whether that is per-read latency, request size, or lanes idling. It also measures
the adjacency the roadmap's read-coalescing and expert-contiguous-layout items assume.

Node classification stays out of the engine: which node is attention vs dense FFN vs
expert matmul is naming policy that varies by architecture, so the rows carry the raw op
and name and scripts/decode-analyze.py classifies. Verified on both gate models that the
generated text, cache hit rate and bytes read are identical with the traces on.
2026-07-15 20:38:56 +02:00
Helldez
71a2c4023a docs: point the telemetry contract and the data index at the route-trace session
Two pointers that belong with the capture: the bench-data index gains its row (stating up
front that the session's tok/s are trace-on and must not feed the benchmarks tables), and
the route-trace section of telemetry.md links the real traces that show the format in use.
2026-07-15 09:53:18 +02:00
Helldez
f6600993f0 feat(scripts): route-analyze.py + document the route trace format
The trace is a long-format CSV; this reads it and answers the questions that shape
streaming speed: the step x layer matrix itself (with a '*' on each expert id the same
layer also routed on the previous step), routing concentration per layer (what a warm-up
should preload), reuse distance per (layer, expert) (LRU vs pinning, and how big a cache
buys what), overlap with earlier steps (whether temporal prefetch can predict), the
cumulative unique-expert curve against the cache budget, hit rate and prefetch
usefulness, routing entropy, and bytes per step and layer.

Stdlib only, like bench-analyze.py — nothing to install.

docs/telemetry.md gains the format (v1), the column semantics, and the two asymmetries
that would otherwise be misread: residency is per routing while expert_bytes is per read
(prefill dedups), and the last layer legitimately has a single prefill step because
llama.cpp gathers only the output token before its FFN.
2026-07-15 08:44:28 +02:00
Helldez
584e96ca6a docs(telemetry): document the compute decomposition and stall floor
Add majflt / cpu_ms to the BMOE_PROGRESS, BMOE_DONE and CSV schemas and explain
how they attribute the compute_ms residual (fault stall vs. throttled core vs.
genuine matmul), plus the new compute: summary line. Document why stall_ms has a
structural floor above zero: the router picks a token's experts only just before
the FFN needs them, so a cache miss forces an on-demand read the overlap cannot
hide, and the residual stall tracks the miss rate (never zero below ~100% hit).
Extend the warmup-analysis CSV schema table with the two columns and record the
CHANGELOG entries.
2026-07-15 07:03:04 +02:00
Helldez
51933985d7 feat(android): richer telemetry — prefill, TTFT, streamed MB, cache, temperature
The live panel showed only decode tok/s, the compute/flash split and cache
hit rate. Surface the rest of what the engine already reports, plus a device
temperature reading:

- prefill rate (tok/s) and time-to-first-token (model load + prompt prefill)
- flash streamed this turn (MB) and expert-cache footprint (resident/budget MiB)
- live battery temperature (BatteryManager, no permission) as a thermal-headroom
  proxy — read on the Android side, it does not travel through the engine

The first four were already computed; only the streamed total needed a new
read_mib field on the BMOE_DONE session line (docs/telemetry.md updated). The
prefill rate and TTFT are also folded into the per-turn transcript line and the
summary. Bumps the app to 0.5.0 (versionCode 6).
2026-07-14 11:01:28 +02:00
Helldez
9eea9743b6 refactor(moe): remove speculative gating to restore the modular seam
Speculative gating was the only feature that broke the ports-and-adapters
seam: it made router_hook reach into architecture-specific router math
(RouterPre), spawned a second thread inside the eval-callback bridge, and
inlined predictor logic into the streaming hot path. It was experimental and
default-off, and never paid its way in steady-state decode on device.

Removing it collapses MoeRecipe back to {arch, expert suffixes} — the header's
stated design intent — and router_hook back to capture -> gather -> load_layer
plus temporal prefetch. The shared speculative-prefetch queue in
expert_stream_source (used by --prefetch) is untouched.

- delete core/src/moe/spec_dot.{h,cpp} and docs/spec-gating.md
- strip RouterPre + router-node fields from recipe.h and every registry row
- drop spec_gate / spec_recall_* config, the run() wiring, and the
  moe_spec_recall_pct / moe_spec_auto_off summary fields (+ CLI flags,
  BMOE_SPEC_GATE env, BMOE_DONE + CSV columns, moe-spec-gate print)
- remove the Android "Speculative gating" toggle and specGate setting
- drop gates G6a-d; G1-G5 and S1-S3 still prove streamed == resident
- clean the reusable bench scripts of --spec-gate; keep docs/bench-data as an
  archive of the historical measurements

Host byte-identity gates pass for qwen3moe and gemma4.
2026-07-14 10:41:27 +02:00
Helldez
a13fc99c9b fix(android): session-reload race + device-agnostic defaults and telemetry
Changing the model or any streaming setting restarts the engine session, but the torn-down
session's thread — unblocked the moment its process is destroyed — ran its finally/waitFor with
shuttingDown=false and reset the UI to IDLE (or ERROR "bmoe-cli exited") and nulled the process
handles, clobbering the fresh session that was already loading. Each session now carries an epoch;
a superseded thread no longer touches the shared process, UI state, or foreground service.

Also, per device-agnostic feedback:
- Default expert cache is a fixed 3000 MiB (was auto-capped); no benchmark- or device-specific
  tuning in the defaults.
- Settings help text is neutral — describes what each knob does, with no measured numbers or
  device/storage claims. Experimental knobs (prefetch, spec-gate) still marked experimental.
- The prefill phase after load is now signalled in the UI (a slow prefill no longer looks stuck).
- At the end of a run the compute and flash-I/O meters show the per-token AVERAGE, not the last
  token, alongside the average tok/s. BMOE_DONE gains prefill_tps, compute_s_tok, io_s_tok.
2026-07-13 13:56:50 +02:00
Helldez
c6128298ba docs: multi-turn chat, per-turn metrics, and the new telemetry fields
session.md documents the engine-held conversation, clear_kv new-chat/continue semantics, the
KV prefix-diff reuse, the SWA full-reprefill fallback and the thinking-on re-prefill cost.
telemetry.md adds the new BMOE_PROGRESS (mgmt/stall/read_mb) and BMOE_DONE (n_prompt/n_past,
cache resident/budget, spec recall, stall/mgmt per token) fields. README describes the
multi-turn app and the winning-recipe defaults.
2026-07-13 12:15:52 +02:00
Helldez
3c8d0fabca docs: speculative gating rework — off-thread prediction, cold inserts, auto-off
Describe the worker-thread prediction (only the hidden-state snapshot stays
on the eval thread), the NEON dot kernels, the cold-end LRU insertion for
speculative experts, and the recall self-governor (--spec-recall-min). Note
in telemetry.md that the prediction CPU no longer inflates the compute
residual, and document the auto-disabled suffix on the moe-spec-gate line.
Device A/B throughput numbers are refreshed separately after re-measuring.
2026-07-13 08:07:12 +02:00
Helldez
0de88f26df docs+android: speculative gating guide, telemetry, and settings toggle
Add docs/spec-gating.md (the residual-slowly-changes idea, the per-architecture recipe fields,
the byte-safety argument, and recall telemetry), document the moe-spec-gate summary line, record
the feature in the changelog, and expose a speculative-gating toggle in the Android settings
(session argv, so changing it reopens the session). Kotlin compiles.
2026-07-12 20:20:31 +02:00
Helldez
e1f8b977fc docs+android: temporal prefetch guide, telemetry, and settings row
Add docs/prefetch.md (the temporal-locality bet, the correctness argument, and telemetry),
document the moe-prefetch summary line, record the feature in the changelog, extend
bench-analyze with prefetch config rows for a device A/B, and expose a prefetch-depth row in the
Android settings (session argv, so changing it reopens the session). Kotlin compiles.
2026-07-12 20:08:14 +02:00
Helldez
b562ce04f4 docs: session mode protocol, architecture, and CHANGELOG
Document the --session stdin/stdout JSON line protocol (telemetry.md), add session.md covering
the open/generate/cancel model, the warm-cache guarantee (S1/S2), independent-prompts vs chat,
and fixed context, note Session as the composition root in architecture.md, and record the
feature in the changelog.
2026-07-12 19:48:49 +02:00
Helldez
7a744e5a59 feat(moe): surface cache-management cost as a mgmt telemetry term
vm-commit, eviction and LRU bookkeeping were hidden inside the compute residual, making the
first tokens after prefill read as pure matmul when the real cost is cache churn. Time that
staging work into mgmt_ns_ and report it per token (mgmt_ms), in the CSV, and in the
moe-stream summary (cache mgmt). compute_ms is now documented as a residual
(wall - io - mgmt, or wall - stall - mgmt under overlap), not a measured quantity. Bytes
served are unchanged; gates G1-G4 pass byte-identically on qwen3moe and gemma4.
2026-07-12 19:28:03 +02:00
Helldez
078703a5e6 docs: document the expert-ready fork hook and overlap telemetry
seam.md gains a section on the hook, its call site and threading contract, and an
explicit sunset condition. README and architecture.md qualify the no-fork claim:
the streaming seam is fork-free, --overlap is the one sunsetted exception. Bench
scripts gain the overlap configs and the stall column.
2026-07-12 08:54:19 +02:00
Helldez
369ee5de96 docs: architecture, seam, MoE streaming, benchmarks, and agent guide 2026-07-10 18:18:24 +02:00