Commit graph

40 commits

Author SHA1 Message Date
Helldez
45a90a2df5
feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#94)
* feat(moe): --drop-cold-experts, spend quality only where it buys I/O

Turbo top-k drops the tail of a routing whether or not those experts were
already in RAM. A resident expert costs no flash read, so that trade pays
quality for nothing on the ~80% of decode routings that are cache hits.

This adds the cache-aware version: skip a routed expert only when it is a
cache MISS and the router weighted it below frac x (1/top-k). Replayed over
the committed route traces at frac 1.0, decode phase, that avoids 66% of
flash reads for 9.5% of the router's weight mass, against 59%/37% for
--n-expert-used 3 — about 3x the reads avoided at a comparable cost. The
threshold is a curve, not a switch: 0.75 trades 4.4% of the mass for 37% of
the reads, better than --n-expert-used 5 on both axes.

Implementation. The decision needs the FINAL router weights, which arrive
several nodes after the topk where the streamer normally loads, so with the
policy armed load_layer() is deferred to the terminal node of the layer's
weight chain. Which node that is depends on the model's gating, so the hook
learns it from the graph rather than carrying an architecture table; until
it is known a layer loads at its topk node undropped. A dropped slot has its
weight zeroed and its expert id repointed at the routing's top-weighted
expert: an expert we decline to read may sit in reserved-but-uncommitted VM
and mul_mat_id would still touch it, so the kernel is given memory that is
certainly resident and multiplies it by exactly zero. Survivors are rescaled
by default, since a systematically shrunk expert output perturbs the residual
stream more than the missing contribution does.

Prefill is excluded by default (cold cache, ~4x the weight mass discarded,
and compute-bound anyway). The largest weight in a routing is always at least
the uniform share, so frac <= 1 can never empty a layer; validate() enforces
the bound and the top expert is pinned regardless.

Gates: G8a proves the deferral and the learned terminal node are transparent
(a threshold below any producible weight leaves the output byte-identical),
G8b that full strength with the cache off never reaches an unloaded expert.

Unlike every other knob this one is state-dependent: what gets dropped
depends on what the cache held, so output is not reproducible across runs.
Off by default, not in the app's settings, and NOT yet measured on device —
the numbers above are a static replay and an upper bound. docs/expert-
dropping.md states what is owed before it is recommended anywhere.

* feat(app): expose cache-aware expert dropping in Settings

Speed / quality -> Drop cold experts, as a percentage of the uniform
share (off / 50 / 75 / 100). The engine takes a fraction; the app stores
integer rungs, so the setting divides by 100 on the way to the flag.

Disabled in mmap mode: the policy asks the expert source what is resident,
and there is no expert source without the streamer. Included in the session
signature, so changing it reopens the session rather than being ignored by
a process already loaded.

Off by default. This exists so the A/B can be run where the engine actually
ships -- through the app, not a pushed CLI binary.

* fix(moe): require the cache for dropping, and correct what it reports

Review of the first two commits found the policy could be armed in a
configuration where it is not cache-aware at all, and that two of the
numbers it reports were wrong.

- Require the LRU cache. With --cache-mb 0 query_residency answers
  all-miss, so the policy silently degenerated into an unconditional
  weight cut -- exactly what --n-expert-used already does, under a flag
  claiming to consult residency. validate() now rejects it, as it already
  did for --prefetch, and the app gates the setting on the same condition.

- Fix experts_routed. It was incremented inside apply_drop, so it counted
  what the policy examined rather than what the router selected: layers
  before the terminal weight node is learned, and every un-armed phase,
  were missing from the denominator. The reported drop rate was a fraction
  of the wrong thing.

- Re-learn instead of re-betting. If the node learned as terminal does not
  arrive, the deferral now also forgets it, so the next graph loads at the
  topk node while it re-learns. Deferring again on a stale guess would
  repeat the fault every token against a graph that had moved.

- Point the gates at a real cache. G8a/G8b ran with the cache off, where
  the shared-slot path has no reserved-but-uncommitted memory -- so the id
  repointing, which is the design's whole safety argument, was never
  exercised. They now run against a constantly-evicting budget. Adds G8a'
  (asserts routings were examined and none dropped, so an inert-threshold
  flake fails legibly) and G8c (at top-k 1 dropping is a proven no-op,
  pinning both the top-expert guarantee and the threshold being taken
  against the effective top-k).

Docs: three metrics change meaning under dropping and none of them said
so. A dropped routing is a miss that is never looked up, so cache_hit_pct
rises without the cache serving more, and token/layer_demand measure what
was staged rather than routed -- documented in telemetry.md, pressure.md
(size the cache with dropping off, then turn it on) and metrics.h.
prefetch.md's "cannot change output" is scoped: under dropping a correct
guess un-drops an expert. limitations.md gains the non-reproducibility
entry, benchmark-method.md the axis plus a warning that reversing the run
order cannot distinguish a moved drop rate from a contaminated cell, and
architecture.md/runtime.h no longer claim unconditional determinism.
Fixes two anchors the README rename broke, and a changelog sentence that
quoted the equal-I/O row while drawing the equal-quality conclusion.

App: Drop cold experts defaults to 75%. The default is a product decision
taken on the maintainer's device; no benchmark for it is published here,
and docs/expert-dropping.md says that plainly instead of implying a
measured figure. The CLI stays off by default -- the byte-identity gates
need a deterministic default.
2026-07-22 15:08:49 +02:00
Helldez
719478908f
feat(dense): --dense-weights ahwb — dense weights in memory Android cannot reclaim (#93)
* feat(dense): add --dense-weights ahwb, dense weights in reclaim-exempt memory

The dense weights must stay resident — every token touches them — and
docs/android-memory.md finds every lever for holding them there closed: mlock is
capped at 64 KiB by the vendor, the cgroup protections are v2-only, MGLRU is
disabled at runtime, and MADV_COLD only redirects reclaim. The exception measured
in 0.13.5 is dma-buf: its pages stay pinned for the buffer's lifetime because a
device may DMA from them, an unprivileged app can allocate one through
AHardwareBuffer, and it reads at exactly anonymous-memory speed.

This wires that allocation into the dense-weights policy. `ahwb` is `anon` with a
single substitution — pio::pinned_alloc instead of the heap — leaving the O_DIRECT
read, the tensor rebind and the mmap handback identical, so an A/B between the two
moves one variable rather than comparing two code paths. Allocation is per tensor,
which keeps it well under the 2047 MiB lock ceiling; a tensor that did exceed it
fails the run instead of quietly taking an anon buffer, since a silent mix would
corrupt the comparison in the direction that flatters the feature. For the same
reason the mode refuses to start on platforms without such an allocation rather
than falling back, which would let an A/B become a mode against itself.

dense_resident_frac keeps working under it — mincore does report on the dma-buf
mapping, which was not obvious — and there it doubles as the falsification test:
pinned pages that fall below 1.0 disprove reclaim-exemption directly.

Exposed as a Dense weights -> Pinned (experimental) setting in the example app,
default off.

What is NOT established is that any of this helps, and the mode should not be
turned on because the reasoning is good. Reclaim-exempt memory does not create
memory: under a >RAM model the RAM the dense weights stop yielding is taken from
the expert cache or from the page cache feeding the stream. That is the trade that
already refuted the bulk restore (#28) and the per-layer LFU cap, both of which
delivered exactly the local gain they predicted and lost throughput anyway. The
deciding A/B is owed and must be run in the app, not over adb: single-shot adb runs
never idle, and this class of bug lives in the reclaim the app's engine suffers
while it sits.

Verified: all 7 host gates pass; on device the three modes generate identical text
on the tiny MoE model, ahwb allocates its 39 pinned buffers, and dense_resident_frac
reaches 1.000 under it.

* test(dense): measure --dense-weights ahwb at +17.9%, and correct the mechanism

In-app on a 1354-token generation (Qwen3.6-35B-A3B, k=8, cache 3000, same session
and binary as its control): 2.588 -> 3.053 tok/s, bootstrap intervals disjoint.
dense_resident_frac reads exactly 1.000, minimum included, in every pinned run, so
reclaim-exemption is now measured rather than inferred.

The mechanism is not the predicted one, and that correction is worth more than the
number. Major faults are EQUAL between the modes (265 vs 257): anon already keeps
the dense weights off the flash. What it does not prevent is the kernel taking ~15%
of them into zram, where a later touch costs a minor fault plus a decompression —
a cost that appears in no I/O counter and no fault counter, so it lands in
compute_ms, which is a residual rather than a measurement of arithmetic. The entire
delta shows up there (298 -> 241 ms) while io_ms, stall_ms and cache hit rate stay
within 1%, and swap falls 562 -> 294 MiB. anon protects the dense weights from
flash; ahwb also protects them from zram.

The trade this was expected to lose on does not appear: the expert cache is
untouched, hit rate identical to the decimal, because the dense set (~1.6 GiB) is
small next to a 3000 MiB cache budget. That is also why it should NOT be extended to
the cache without sizing the prize first — only ~294 MiB of that budget sits in zram,
against a cost of 3 GiB of rigid LMK-accounted memory and the loss of the
reserve/commit/evict elasticity the cache is built on.

Default stays anon. In the decisive pair ahwb ran first and an order effect cannot be
excluded — the reversed pair is owed — and this is one device, one model, one config.

Two negative results are committed alongside so the reasoning is checkable: three
67-74 token pairs that are ALL inconclusive (per-token CV 33-71%, every interval
overlapping), because reclaim accumulates and short turns never build up enough of
it; and a cross-day pair reading +63.6% that is not usable, since anon alone moved
+38.8% between the two days.

Transferable: compute_ms has been absorbing zram decompression all along, so earlier
"this regime is compute-bound" conclusions deserve re-examination.
2026-07-21 11:09:51 +02:00
Helldez
d61dd55d26 fix(cli): honour the documented flag-wins-over-env contract
The header comment and the usage text both promise that an explicit flag beats the matching
BMOE_* variable, but the overrides decided "was this flag passed?" by asking whether the field
still held its default. Those are different questions for any flag whose default is a value a
user might deliberately pass: --cache-mb 0 (cache off), --io-threads 4, --prefetch 0 and
--n-expert-used 0 all looked untouched and were overridden anyway. The Android app passes two
of them explicitly on every run.

The parse loop now records the flags it was given and the overrides consult that set, so
passing a flag its default value is still a choice. Values that do arrive from the environment
are validated exactly as before: resolution still happens between parsing and validate().

This also retires the special case that kept --cache-mb auto safe. auto and 0 share one
spelling, so asking whether that spelling was typed covers both, and the cache_auto clause
that used to stand in for the question goes away.

Verified against the tiny gate model, reading io_threads and cache_mb back out of the CSV
preamble: env alone still applies, flag alone still applies, flag plus env now resolves to the
flag, and --cache-mb auto still computes its budget with BMOE_CACHE_MB set.
2026-07-20 07:38:53 +02:00
Helldez
aa6a7fdafd fix(chat): honour "thinking off" on templates that ignore enable_thinking
Turning Thinking off set the template variable enable_thinking and stopped there. That
variable is only a request to the model's own chat template, and a template is free to
ignore it. LFM2.5's never reads it, so the rendered prompt was byte-identical with thinking
on and off, the model reasoned anyway, and nothing reported that the setting had been
dropped (#82).

Detection is measured, not assumed: at open() the template is rendered with the flag on and
off and the two prompts compared, then rendered once more with a continuation. That answers
the only question that matters — does this template react — for any model in any language.
common_chat_templates_support_enable_thinking cannot answer it: per handler it is a
hardcoded literal reporting "this model can reason", not "this template reads the variable".

Enforcement uses llama.cpp's continuation hook: a synthetic trailing assistant message with
continue_final_message makes upstream's per-template handler render that family's own
"reasoning is over" span into the prompt. The model resumes at the first token of its answer
with the reasoning already behind it. This is what a template that implements the toggle
natively does (Qwen3 renders <think></think> for enable_thinking=false), so it needs no
cooperation from the template and no sampler — it works on the greedy path the byte-identity
gates run on.

Because the span comes from upstream's handler, the engine names no markers of its own. That
retires the two hardcoded harmony strings in the decode path: priming gpt-oss to answer
without reasoning was a literal "<|start|>assistant" suffix test and a literal
"<|channel|>final<|message|>" appended to the prompt. gpt-oss now takes the same generic path
as every other family and resumes at the same point, so a submodule bump that changes those
markers needs no engine change.

Models where neither mechanism exists are reported rather than fought: BMOE_READY gains
think_ctl (template | prefill | none) and the app shows the Thinking switch disabled, with
the reason, instead of offering a control that does nothing.

Deliberately not done: forcing the reasoning block closed on logits. Measured on-device in
the closed PR #83, it made LFM2.5 strictly worse — the model reopened the block, then
abandoned the tags and reasoned in plain prose into the answer. Suppression belongs in the
prompt, before the model commits to reasoning, not mid-generation.

tests/think_control_test.cpp pins all three regimes against the vendored templates with no
model: Qwen3 as template, LFM2.5 and gpt-oss as prefill with the span asserted CLOSED, a
plain template as none, and fail-open when the probe cannot run.
2026-07-19 20:26:22 +02:00
Helldez
3058493523
Merge pull request #73 from Helldez/feat/71-default-params
feat(defaults): align n_predict at 128 and fix the app expert cache at 2000 MiB
2026-07-19 09:31:55 +02:00
Helldez
eb455bc4e6 feat(defaults): align n_predict at 128 and fix the app expert cache at 2000 MiB
The generation default diverged across surfaces (CLI 32, app 48) and both
budgets truncate most answers mid-sentence, which reads as broken rather
than slow on a first run. 128 is the smallest choice that lets a typical
answer finish; CLI and app now share it so there is one documented default.

The app's expert cache moves from Auto (ceil 3000) to a fixed 2000 MiB:
Auto sizes to whatever RAM happens to be free at load, so first
impressions and benchmarks varied with unrelated device state. 2000 sits
above the engine's 1500 MiB floor (no --force-cache) and Auto stays
selectable for devices where a fixed budget is wrong. Existing installs
keep their saved prefs; only fresh installs see the new defaults.

Closes #71.
2026-07-19 09:23:22 +02:00
Helldez
3db882d278 feat(trace): add layer-granularity compute trace (--compute-trace-layers)
The per-node compute trace pays ~3000 barriers per token, which serializes
the graph against the expert stream: on a model that streams heavily the
trace mostly measures its own serialization (Qwen3-30B: 9.4 s/token traced
vs 0.39 untraced), so its absolutes cannot be compared across models.

Layer granularity isolates only the first node of each layer (~n_layer
barriers per token). Operator coalescing and the async expert prefetch
survive, so the traced numbers stay close to an untraced run. Rows share
the per-node schema with op LAYER: name blk.<il> aggregates one layer''s
segment, pre the embedding lookup, post the last layer''s tail plus
final norm and LM head (closed by the session right after llama_decode,
since the tail has no successor boundary to observe it).

The granularity flows RunConfig -> SessionConfig -> RouterHook; the routing
nodes the streamer isolates anyway also close a segment, a barrier that
exists untraced too. decode-analyze.py detects the granularity and prints
the per-segment table.

Gates: all 6 pass (byte-identity qwen3moe + gemma4). Smoke-tested on the
tiny-moe models with streaming on.
2026-07-19 09:21:39 +02:00
Helldez
55b8579396 feat(core): surface the reasoning span alongside the answer (#70)
Since #49 wired the chat parser correctly, a reasoning model's thinking is
stripped from the shown answer unconditionally — including when the user asked
for thinking. With the in-app Thinking toggle ON, the answer area stays blank
while the model reasons and only the final answer ever appears; on a slow
streamed decode that reads as a hang for the whole thinking span.

The reasoning was being parsed and thrown away: shown_text() kept only
common_chat_parse(...).content and dropped .reasoning_content, and no field
downstream carried it. Surface it instead of discarding it:

  - TokenMetrics gains `reasoning`, RunResult gains `reasoning_text`; session's
    shown_view returns {content, reasoning} from a single parse (the partial
    parse already fills reasoning_content incrementally, so it streams).
  - The line protocol carries a `reasoning` field on BMOE_PROGRESS and
    BMOE_DONE, kept apart from `text` so a UI can render it as a distinct
    thinking block rather than inline. Documented in docs/telemetry.md.

The answer in `text`/generated_text is unchanged (reasoning still stripped),
so this is display-only: the byte-identity gates are untouched. The chat-parse
host gate now also asserts a partial parse exposes the reasoning span — the
contract the live thinking block depends on.
2026-07-18 22:43:23 +02:00
Helldez
b359ccdbd9 feat(core): opt-in sampling (temperature, top-p, top-k, seed) (#51)
Generation was greedy-only: argmax() always took the highest-logit token and
RunConfig carried no sampling knobs, so every caller was stuck with
deterministic decoding — the wrong default for anything conversational.

Add a SamplingConfig (temp, top_k, top_p, seed) to RunConfig/SessionConfig.
temp <= 0 (the default) keeps the exact argmax path, so a caller that passes
nothing sees today's behaviour and the byte-identity gates stay meaningful.
temp > 0 builds a per-session chain top_k -> top_p -> temp -> dist using only
the public llama_sampler_* API (common_sampler lives in the non-stable common
layer, so it is avoided per hard rule 1). The chain is freed with the session
and its RNG resets on clear_kv, so a fixed seed reproduces a fresh chat.

The seed default is pinned to LLAMA_DEFAULT_SEED via a static_assert, so the
llama-free config header cannot drift from the backend constant it mirrors.

CLI: --temp, --top-k, --top-p, --seed, resolved in main.cpp (the only place
env overrides are read). Surfacing the knobs in the Android app is left to a
separate change, as the issue asks.

Adds tests/config_test.cpp: pure validate() unit tests (no model), covering
the pre-existing rules and the new sampling ranges — and making good on the
long-standing "validate() is unit-tested" claim in config.h, which had no test.
2026-07-18 10:40:13 +02:00
Helldez
4be2e177c1
feat(cli): default --dense-weights to anon, matching the Android app (#55)
The library/CLI default was `warm` while the Android app has always shipped
`anon` ("the default — wins on >RAM models"). Align the CLI to the app: set
`DenseWeightsMode::Anonymous` as the config default so `bmoe-cli` without a
`--dense-weights` flag now reads the dense weights via O_DIRECT into anon
buffers, the policy that pays on the >RAM models the engine targets (3.2x on
gpt-oss). `warm` and `mmap` stay available for RAM-fitting models.

Byte-identity gates re-run green — anon rebind is already proven identical to
the mmap reference (G6/G7), so the default flip changes no output. Docs and the
CLI --help updated to reflect anon as the default.
2026-07-18 09:39:18 +02:00
Helldez
285c617e40 refactor(engine): dedupe the session plumbing
Three copies of the same knowledge, each free to drift from the others.

RunConfig -> SessionConfig was spelled out field by field in both run() and the
CLI's interactive loop. Adding a field to RunConfig then depended on remembering
to touch both; session_config_from() makes it one edit.

The BMOE_LOAD/BMOE_PROGRESS emission was copied between the one-shot --progress
path and the session loop. The Android app parses one parser's worth of protocol,
so the two lines must be byte-identical — a shared emit_progress_line() makes that
true rather than merely intended.

open() reopened and reparsed the gguf header for each question it asked of it (the
arch-prefixed override key, the route trace's effective top-k, the run info's
top-k/expert count). It now reads at most once, lazily: the callers are conditional
— a run with no override and no trace asks nothing — so an eager read would be work
the common path never needs.

Deliberately NOT done: splitting generate(). Its per-token metric block is not
self-contained — it mutates eight loop-carried accumulators, so extracting it means
inventing a struct to thread them through. That is more code than it removes, on a
hot loop, for a legibility gain alone.

Gates pass. CLI --progress and --cache-mb auto verified against the tiny model.
2026-07-17 10:19:32 +02:00
Helldez
228e162465 chore: sweep the retired governor's residue from core and CLI
Retiring the adaptive cache governor left state that nothing reads and comments
that describe a control loop the engine no longer has. The budget is sized once
at init and then fixed for the run; make the code say only that.

Dead state, all write-only since the governor went:
  - cache_auto_, cache_target_ and slot_sz_ are gone; cache_floor_ and
    cache_hard_cap_ demote to locals in the auto-sizing block, which is the only
    place they ever meant anything.
  - DenseWeights::rewarm() had no caller: it was the "reclaim happened, warm the
    dense set again" hook the governor drove.
  - ProcessMemory::rss_peak_bytes (VmHWM) was populated and never read. The bench
    script samples VmHWM straight from /proc, so nothing is lost.

Comments that had started to lie: config.h still described a budget "re-checked
during generation", session.h "complements --cache-mb auto's automatic tracking",
expert_source.h a budget that "moves under --cache-mb auto". And the CLI printed
"(auto, resized N×)" where N is now structurally always 0.

cache_resizes stays: it still counts an app's explicit set_cache_budget_mb calls
(Android onTrimMemory), which the shrink gate exercises.

Gates pass (byte-identity, qwen3moe + gemma4).
2026-07-17 10:04:27 +02:00
Helldez
75b5a634fc refactor(moe): retire the adaptive cache governor; fixed LRU + one-shot auto
Measured net loss on >RAM models (cache-off is the ceiling), so the runtime governor, --cache-dynamic/--cache-gov2, and the sense/resize loop go. Kept: the fixed --cache-mb N LRU, --cache-mb auto sizing once at load, cache-off default. Telemetry aligned (dense_resident_frac made live; resident_frac/cache_cuts removed). App: one 3-way dense-weights selector. Host+ctest green.
2026-07-17 08:55:18 +02:00
Helldez
847676d715 refactor(io): FileReader + DenseWeights modules — per-consumer O_DIRECT
Extract the pooled positioned reader (O_DIRECT chosen per reader, not one global flag) and the dense-weights policy+sensor into their own modules; ExpertStreamSource composes them. So the dense and expert O_DIRECT choices are independent.
2026-07-17 08:55:17 +02:00
Helldez
e78fb78194 feat(moe): --dense-weights — dense (non-expert) weight residency policy
Read each dense weight once via O_DIRECT into an anon buffer and rebind (anon), page-cache them at load (warm), or leave them mmap'd (mmap). RouterHook captures the non-expert weight leaves; the streamer reads+rebinds and drops the mmap copy. Byte-identity gates G6/G7.
2026-07-17 08:55:17 +02:00
Helldez
9a1d1f8a1c feat(metrics): on-device memory telemetry, pressure sensing, and an adaptive cache governor
The measure-your-own-memory series: fault/CPU decomposition, the anon/file RSS split, mincore residency sensors, the Android metrics screens, and a --cache-dynamic governor that sizes the expert cache to what the device concedes. (The governor is retired further down this history; the telemetry stays.)
2026-07-17 08:54:06 +02:00
Helldez
f9e408f542 feat(metrics): --compute-trace and --io-trace decompose the decode
The per-token CSV reports compute as a residual (wall - io - mgmt), so everything the
engine does not itself clock is pooled into it: page faults, scheduler stalls, and the
matmuls. A residual cannot say which. That is the whole reason gpt-oss-120b reads as
"compute-bound" at 1.7 s/token while faulting 8.4k pages per token — the flash wait is
billed to compute because it happens under llama_decode.

--compute-trace measures it instead. Asking the eval callback to isolate a node makes
ggml compute exactly up to it and synchronize, so the wall delta between consecutive
boundaries is that node's real compute time; sampling major faults across the same
boundaries attributes the >RAM stall to the node that paid it. Still no llama.cpp patch:
this rides the public cb_eval ABI, whose ask/no-ask contract already specifies the
isolation. It costs a barrier per node and forbids operator coalescing, so it is a
diagnostic — a traced run is not a benchmark run, and only the shares are meaningful.
Unlike the other traces it does not need --moe-stream: it times the graph, so a dense
mmap baseline can be traced and compared against a streamed run.

--io-trace records one row per pread: latency, requested vs aligned bytes, lane, and the
(layer, expert, projection) it serves — values already computed at every enqueue site and
until now discarded. This is where the flash floor is: the aggregate 760 MiB/s sits far
below the drive's sequential ceiling because routed slices are scattered, and the trace
says whether that is per-read latency, request size, or lanes idling. It also measures
the adjacency the roadmap's read-coalescing and expert-contiguous-layout items assume.

Node classification stays out of the engine: which node is attention vs dense FFN vs
expert matmul is naming policy that varies by architecture, so the rows carry the raw op
and name and scripts/decode-analyze.py classifies. Verified on both gate models that the
generated text, cache hit rate and bytes read are identical with the traces on.
2026-07-15 20:38:56 +02:00
Helldez
a0d54066f8 feat(telemetry): per-step per-layer MoE route trace (--route-trace)
The per-token metrics say how long a token took and how many expert bytes it pulled.
They cannot say WHICH experts each layer routed, how the router weighted them, or
whether a routed expert was already resident — which is what decides the streaming
cost, and what a warm-up heuristic has to be designed against.

--route-trace PATH writes one row per (step, layer, slot): the cells of a step x layer
matrix whose payload is the routed expert ids, plus the applied routing weight, the
cache state at that instant (miss / hit / prefetch-hit) and the flash bytes the routing
reads. Both prefill and decode are traced, on one absolute step axis.

The capture rides the existing eval-callback seam: the stream pass additionally asks for
each layer's router-weight node. llama.cpp builds a chain (ffn_moe_weights -> _softmax |
_norm -> _scaled), so taking the LAST one offered per layer yields the applied weight
without a per-architecture table. It observes only — the ids handed to load_layer are
the same traced or not, and the gates still pass — but each extra ask is another graph
barrier, so it stays off unless requested and a traced run is not a benchmark run.

Three defaulted virtuals on IExpertSource let the tracer describe what a routing cost
without changing it: settle_spec() (integrate landed prefetches, or a correct guess
reads as a miss), query_residency() and expert_bytes(). Dense bytes per layer come from
the gguf and are static — dense weights are mmap-resident and never streamed, so the
trace states them once rather than pretending to measure them per step.

Verified: byte-identity gates pass (qwen3moe + gemma4); generation is identical with and
without the trace; prefill's last layer correctly reports the final prompt position,
since llama.cpp gathers only the output token before that layer's FFN.
2026-07-15 08:44:28 +02:00
Helldez
aca596d6d4 feat(metrics): decompose per-token compute into CPU-time and major faults
The per-token compute_ms is a residual (wall - stall - mgmt), so it silently
absorbed anything not otherwise timed: dense-weight page faults on a >RAM model
and scheduler idle under a frequency cap both inflated it, indistinguishably from
genuine matmul. Two directly-measured counters, sampled around llama_decode with
no submodule patch, decompose it:

- major_faults() (getrusage ru_majflt): major page faults this token. Non-zero
  means a mmap-resident dense weight was re-faulted from flash inside the decode.
  Experts stream via O_DIRECT and never fault, so this isolates the dense
  residency stall.
- process_cpu_seconds() (CLOCK_PROCESS_CPUTIME_ID): CPU time across all threads.
  Divided by wall x threads it yields occupancy: near 1.0 is compute-bound, well
  below flags a throttled or preempted core.

Both live behind the platform_io seam and return 0 where unmeasurable (Windows
host build), treated as "unmeasured" rather than zero work. Surfaced as majflt /
cpu_ms per-token TokenMetrics and majflt_per_token / cpu_s_per_token run
averages, through the CSV sink (new trailing columns + summary keys), the
BMOE_PROGRESS / BMOE_DONE session lines, and a compute: line in the one-shot
summary reporting occupancy and faults/token. Measured for the mmap baseline too,
where dense faults also occur. Bytes served are unchanged; byte-identity gates
pass.
2026-07-15 07:02:55 +02:00
Helldez
188e321fa5 feat(engine): warm dense weights into the page cache at load
Only expert tensors are streamed; every other weight (embeddings, attention,
norms, lm_head) stays mmap-resident and is demand-paged by the kernel. On a
model far larger than RAM those dense pages fault in lazily, one scattered 4 KiB
page at a time, inside the first decodes — surfacing as a large compute_ms
residual that only settles after ~20 tokens (see docs/warmup-analysis.md).

Read them eagerly instead: after the streamer binds the expert offsets, sweep
the file's non-expert byte ranges (the complement of those offsets) once,
sequentially, into the page cache before the first token. The ranges are derived
at runtime from the discovered expert offsets — no per-model constants. It
touches neither the expert cache nor the budget, so the streaming path is
byte-identical (cache_hit% unchanged step-for-step); it only moves cold faults
out of the hot path and into load time.

Measured on gpt-oss-120b (62 GB, ~5.2x RAM): first-five-token wall average drops
~20x and the first token goes from ~18 s to ~1 s. Models whose dense set is ~1 GB
(Qwen, Gemma) are neutral. On by default; --no-warm-dense opts out for A/B runs.
2026-07-14 18:44:03 +02:00
Helldez
51933985d7 feat(android): richer telemetry — prefill, TTFT, streamed MB, cache, temperature
The live panel showed only decode tok/s, the compute/flash split and cache
hit rate. Surface the rest of what the engine already reports, plus a device
temperature reading:

- prefill rate (tok/s) and time-to-first-token (model load + prompt prefill)
- flash streamed this turn (MB) and expert-cache footprint (resident/budget MiB)
- live battery temperature (BatteryManager, no permission) as a thermal-headroom
  proxy — read on the Android side, it does not travel through the engine

The first four were already computed; only the streamed total needed a new
read_mib field on the BMOE_DONE session line (docs/telemetry.md updated). The
prefill rate and TTFT are also folded into the per-turn transcript line and the
summary. Bumps the app to 0.5.0 (versionCode 6).
2026-07-14 11:01:28 +02:00
Helldez
9eea9743b6 refactor(moe): remove speculative gating to restore the modular seam
Speculative gating was the only feature that broke the ports-and-adapters
seam: it made router_hook reach into architecture-specific router math
(RouterPre), spawned a second thread inside the eval-callback bridge, and
inlined predictor logic into the streaming hot path. It was experimental and
default-off, and never paid its way in steady-state decode on device.

Removing it collapses MoeRecipe back to {arch, expert suffixes} — the header's
stated design intent — and router_hook back to capture -> gather -> load_layer
plus temporal prefetch. The shared speculative-prefetch queue in
expert_stream_source (used by --prefetch) is untouched.

- delete core/src/moe/spec_dot.{h,cpp} and docs/spec-gating.md
- strip RouterPre + router-node fields from recipe.h and every registry row
- drop spec_gate / spec_recall_* config, the run() wiring, and the
  moe_spec_recall_pct / moe_spec_auto_off summary fields (+ CLI flags,
  BMOE_SPEC_GATE env, BMOE_DONE + CSV columns, moe-spec-gate print)
- remove the Android "Speculative gating" toggle and specGate setting
- drop gates G6a-d; G1-G5 and S1-S3 still prove streamed == resident
- clean the reusable bench scripts of --spec-gate; keep docs/bench-data as an
  archive of the historical measurements

Host byte-identity gates pass for qwen3moe and gemma4.
2026-07-14 10:41:27 +02:00
Helldez
a13fc99c9b fix(android): session-reload race + device-agnostic defaults and telemetry
Changing the model or any streaming setting restarts the engine session, but the torn-down
session's thread — unblocked the moment its process is destroyed — ran its finally/waitFor with
shuttingDown=false and reset the UI to IDLE (or ERROR "bmoe-cli exited") and nulled the process
handles, clobbering the fresh session that was already loading. Each session now carries an epoch;
a superseded thread no longer touches the shared process, UI state, or foreground service.

Also, per device-agnostic feedback:
- Default expert cache is a fixed 3000 MiB (was auto-capped); no benchmark- or device-specific
  tuning in the defaults.
- Settings help text is neutral — describes what each knob does, with no measured numbers or
  device/storage claims. Experimental knobs (prefetch, spec-gate) still marked experimental.
- The prefill phase after load is now signalled in the UI (a slow prefill no longer looks stuck).
- At the end of a run the compute and flash-I/O meters show the per-token AVERAGE, not the last
  token, alongside the average tok/s. BMOE_DONE gains prefill_tps, compute_s_tok, io_s_tok.
2026-07-13 13:56:50 +02:00
Helldez
d05f9d9b2e feat(session): multi-turn chat with KV prefix reuse
The Session now owns the conversation. In chat mode it re-renders the model's chat
template over the whole history each turn, then diffs the resulting tokens against the
tokens already decoded into the KV (kv_tokens) and prefills only the diverging suffix —
so a follow-up turn pays for its own tokens, not a full re-prefill of the conversation
(prefill is the slow phase on device). clear_kv=true still means "new chat": it drops
the KV and history, keeping the one-shot run() path (and the byte-identity gates) exactly
as before. A partial KV removal that SWA memory (Gemma) refuses falls back to a full
re-prefill. Cancel rolls the turn back to the reused prefix, leaving prior turns usable.

BMOE_DONE now reports the per-turn prefilled count (n_prompt) and total context (n_past),
plus cache resident/budget, spec recall and stall/mgmt per-token; BMOE_PROGRESS adds
mgmt_ms, stall_ms and read_mb. All additive JSON fields.
2026-07-13 12:03:53 +02:00
Helldez
10fbafa4eb feat(moe): --n-expert-used to override active experts per token
Add a knob to reduce the model's top-k expert routing (e.g. 8 -> 6 on
Qwen3-30B-A3B): fewer active experts cut per-token compute and, under
--moe-stream, flash I/O roughly in proportion. It is a speed/quality
trade-off — fewer experts changes the output — so it is opt-in and
defaults to the model's own count (0 = no override).

Implemented purely through llama.cpp's public kv_overrides on the
arch-prefixed expert_used_count metadata key, applied at model load:
the compute graph then emits a narrower top-k and the whole streaming
path adapts on its own (the router hook reads the top-k width from the
graph, the expert source works off a runtime id count). No llama.cpp
patch, no per-architecture constants — the arch is peeked from the gguf
before load via the public gguf API to build the override key.

- config: n_expert_used on RunConfig, pure lower-bound check in validate()
- runtime: read_gguf_model_info() peek + kv_override injection at load,
  with a clean out-of-range error instead of llama.cpp's GGML_ASSERT
- cli: --n-expert-used flag and BMOE_N_EXPERT_USED env override
- gates unchanged and green (default is a no-op, byte-identity holds)
2026-07-13 11:52:38 +02:00
Helldez
ea37f9ed6a feat(moe): --cache-ceil-mb upper bound on the auto cache budget
--cache-mb auto took all the headroom above the floor, which can leave the
system with only the 1536 MiB floor free when the marginal hit-rate gain no
longer justifies it. Add --cache-ceil-mb: the auto budget (init and runtime
grow-back) is capped at min(ceiling, full expert-set size). 0 keeps the
previous behaviour (cap only at the expert-set size).
2026-07-13 09:31:16 +02:00
Helldez
3768aa8985 feat(moe): runtime cache-budget adaptation under memory pressure
With --cache-mb auto the budget now tracks free RAM during generation, not
just at init. On the eval thread, inside load_layer's mgmt window, a throttled
re-probe (~every 128 layer loads) shrinks the budget by the shortfall when
available memory dips under the floor — the following eviction loop drains to
it — and grows it back toward the init target when memory recovers, with
hysteresis. Also adds set_cache_budget() (Session::set_cache_budget_mb): an
explicit resize from outside a decode, for an app's onTrimMemory callback.
The budget stays strictly positive (the LRU buffers exist for the session;
load_layer keys shared-slot vs LRU off cache_max_ == 0, so a runtime zero
must never happen). Budget and resize count are surfaced in Stats and the
moe-cache summary line under auto.
2026-07-13 08:20:30 +02:00
Helldez
e1d78d1855 feat(moe): auto cache budget from available memory (--cache-mb auto)
Instead of a hand-tuned --cache-mb, size the expert cache to the device.
With --cache-mb auto the engine, once the full expert-set size is known,
sets the budget to (available RAM - cache_floor_mb, default 1536) clamped to
[cache_min_mb, total expert bytes]; if memory is unknown it falls back to the
floor. cache_auto is mutually exclusive with an explicit cache_mb, and counts
as 'cache on' for the --prefetch/--spec-gate requirement. Adds --cache-floor-mb.
This is the knob that keeps the phone responsive without the user guessing a
number that OOMs on one model and underfills on another. Runtime shrink under
pressure follows in a later commit; this lands the init-time sizing.
2026-07-13 08:13:41 +02:00
Helldez
7d6a4c68c9 feat(moe): auto-disable speculative gating when router recall stays low
Speculative gating only pays for its extra I/O and per-layer graph barrier
when the router it forecasts is actually predictable. On architectures where
cross-layer routing correlates weakly the predictor wastes reads for no
throughput gain, and previously nothing turned it off. Add a self-governor:
once at least spec_recall_warmup predictions have been scored (past prompt-
transition noise), if cumulative recall is below spec_recall_min_pct the
engine latches spec-gating off for the rest of the run — which also drops
the router-input barrier. Defaults 75/512; 0 disables the check. Exposed as
--spec-recall-min and surfaced in the moe-spec-gate summary line. New gate
G6d proves the mid-run on-to-off transition is byte-identical (the governor
fires at 25 percent recall on the synthetic model).
2026-07-13 08:00:50 +02:00
Helldez
b82f8ca224 style: clang-format the session, prefetch and spec-gating sources 2026-07-12 20:23:12 +02:00
Helldez
466316ab0f feat(moe): speculative gating — predict the next layer's experts from its router
Where temporal prefetch guesses the next layers from the previous token's routing, speculative
gating (--spec-gate) predicts the NEXT MoE layer directly: it isolates the router-input node,
reads the current layer's hidden state (the residual stream changes slowly between layers), runs
the next layer's gate on it, and prefetches its top-k. The prediction only feeds the byte-safe
prefetch queue, so a mispredict wastes a read but never changes output.

Recipe-driven, no hardcoding: each row carries the router-input node name, the mmap-resident gate
weight (and per-channel scale), and a pre-transform. qwen3moe/qwen2moe route on the post-norm
hidden (kNone); gemma4 routes on attn_out with rms_norm * 1/sqrt(n_embd) * scale (kRmsScaled) and
has interleaved dense layers, handled by targeting the next bound MoE layer. n_expert_used and the
rms epsilon come from generic gguf metadata keys. F32/F16 gate weights supported; a quantized gate
disables the feature (logged once). Recall (predicted vs actual routing) is measured and reported.

Gates G6a/b/c pass byte-identically on qwen3moe and gemma4 (b forces synchronous completion to
exercise predict→integrate→hit). End-to-end on the tiny models: gemma4 reaches 91% router recall,
confirming the transform, while output stays byte-identical.
2026-07-12 20:18:35 +02:00
Helldez
4d57da5edb feat(moe): temporal prefetch of the next layers' experts on idle lanes
While a token computes layer l, the previous token's routing at layers l+1..l+K is a strong
predictor of this token's — so speculatively read those experts on the otherwise-idle I/O lanes,
turning the next layer's read into a cache hit. RouterHook records each layer's last-token
routing and calls the new IExpertSource::prefetch() hint; ExpertStreamSource services it with a
low-priority speculative queue that never delays real work: all LRU mutation stays on the eval
thread (prefetch commits + enqueues, quiesce integrates completed reads and discards the rest at
the next real load), while workers only read bytes and yield the moment a real batch arrives. A
speculative slice is the identical read (same file offset, same buffer) a real miss would issue,
so integration is byte-safe by construction. --prefetch K (env BMOE_PREFETCH), cache required;
telemetry reports speculative MiB and useful-hit rate. Gates G5a/b/c pass byte-identically on
qwen3moe and gemma4 — G5c forces synchronous completion to exercise integrate-then-hit
deterministically (observed 3 experts prefetched, 1 useful).
2026-07-12 20:05:28 +02:00
Helldez
2363471f79 feat(cli): --session keeps the process alive across prompts over stdin
Adds an interactive mode that opens one Session and serves many prompts, so the model load
and the expert-cache warm-up are paid once instead of per prompt. Requests are one JSON line
per prompt on stdin ({cmd:generate|cancel|close}); responses use the BMOE_* line protocol on
stdout (BMOE_READY / BEGIN / PROGRESS / LOAD / DONE / ERROR), a superset of --progress so the
existing per-token parser is reused. A stdin reader thread applies cancel immediately (via the
session abort callback) while the main thread drives generation. The one-shot argv path is
unchanged. Verified end-to-end on the tiny model: two warm generates produce byte-identical
output and the cache hit rate carries over between them.
2026-07-12 19:40:36 +02:00
Helldez
7a744e5a59 feat(moe): surface cache-management cost as a mgmt telemetry term
vm-commit, eviction and LRU bookkeeping were hidden inside the compute residual, making the
first tokens after prefill read as pure matmul when the real cost is cache churn. Time that
staging work into mgmt_ns_ and report it per token (mgmt_ms), in the CSV, and in the
moe-stream summary (cache mgmt). compute_ms is now documented as a residual
(wall - io - mgmt, or wall - stall - mgmt under overlap), not a measured quantity. Bytes
served are unchanged; gates G1-G4 pass byte-identically on qwen3moe and gemma4.
2026-07-12 19:28:03 +02:00
Helldez
fcc80ad0a5 feat(cli): add --overlap flag and stall telemetry
--overlap (env BMOE_OVERLAP, flag wins) enables the overlapped decode path. The
CSV sink appends stall_ms as the last column and stall_s/tok to the summary line
(positional append, backward compatible); the run summary prints per-token stall.
2026-07-12 08:54:19 +02:00
Helldez
481bf76b02 fix: honor the thinking toggle for all model families
The chat template was always rendered with enable_thinking=true, so a reasoning
model kept emitting its thinking channel; Gemma's leaked into the answer because
the display-time parser does not recognise its reasoning format. Thread a think
flag through RunConfig and set inputs.enable_thinking, and drive it from a
--no-think CLI flag. The app now passes --no-think for every architecture instead
of appending Qwen's /no_think only, so the toggle works for Gemma too.
2026-07-11 22:44:43 +02:00
Helldez
4fe7cfc393 feat: report prefill, model-load and TTFT in the run summary
Additive telemetry only. runtime times model load + streaming setup as one
'load' span and prefill as a separate span, so TTFT = load + prefill; the CLI
prints them and the CSV summary carries n_prompt/load_s/prefill_s/prefill_tps.
No change to the streaming or decode path.
2026-07-11 22:44:43 +02:00
Helldez
4b0ba10161 feat(engine): apply the model family's chat template, not just Qwen ChatML
--chatml wrapped every prompt in Qwen ChatML, so instruct models with a
different turn format (Gemma 4's <start_of_turn> turns) produced raw-completion
output. Make the wrapper arch-aware: read general.architecture and emit the
family's turn format (gemma -> <start_of_turn>, everything else -> ChatML).

The engine has no Jinja engine, so it can't render the gguf's embedded
tokenizer.chat_template (Gemma 4's is a 16 KB Jinja program with tool-call
macros that llama.cpp's non-Jinja llama_chat_apply_template cannot parse); the
small stable turn format per family is the pragmatic, dependency-free path. Add
a branch when supporting another instruct family.
2026-07-11 17:05:01 +02:00
Helldez
d2d5dc9ee7 style: apply clang-format across sources 2026-07-10 18:34:36 +02:00
Helldez
5273af6487 feat(cli): bmoe-cli with streaming flags, ChatML, and machine telemetry 2026-07-10 18:18:19 +02:00