Commit graph

16 commits

Author SHA1 Message Date
Helldez
10fbafa4eb feat(moe): --n-expert-used to override active experts per token
Add a knob to reduce the model's top-k expert routing (e.g. 8 -> 6 on
Qwen3-30B-A3B): fewer active experts cut per-token compute and, under
--moe-stream, flash I/O roughly in proportion. It is a speed/quality
trade-off — fewer experts changes the output — so it is opt-in and
defaults to the model's own count (0 = no override).

Implemented purely through llama.cpp's public kv_overrides on the
arch-prefixed expert_used_count metadata key, applied at model load:
the compute graph then emits a narrower top-k and the whole streaming
path adapts on its own (the router hook reads the top-k width from the
graph, the expert source works off a runtime id count). No llama.cpp
patch, no per-architecture constants — the arch is peeked from the gguf
before load via the public gguf API to build the override key.

- config: n_expert_used on RunConfig, pure lower-bound check in validate()
- runtime: read_gguf_model_info() peek + kv_override injection at load,
  with a clean out-of-range error instead of llama.cpp's GGML_ASSERT
- cli: --n-expert-used flag and BMOE_N_EXPERT_USED env override
- gates unchanged and green (default is a no-op, byte-identity holds)
2026-07-13 11:52:38 +02:00
Helldez
ea37f9ed6a feat(moe): --cache-ceil-mb upper bound on the auto cache budget
--cache-mb auto took all the headroom above the floor, which can leave the
system with only the 1536 MiB floor free when the marginal hit-rate gain no
longer justifies it. Add --cache-ceil-mb: the auto budget (init and runtime
grow-back) is capped at min(ceiling, full expert-set size). 0 keeps the
previous behaviour (cap only at the expert-set size).
2026-07-13 09:31:16 +02:00
Helldez
3768aa8985 feat(moe): runtime cache-budget adaptation under memory pressure
With --cache-mb auto the budget now tracks free RAM during generation, not
just at init. On the eval thread, inside load_layer's mgmt window, a throttled
re-probe (~every 128 layer loads) shrinks the budget by the shortfall when
available memory dips under the floor — the following eviction loop drains to
it — and grows it back toward the init target when memory recovers, with
hysteresis. Also adds set_cache_budget() (Session::set_cache_budget_mb): an
explicit resize from outside a decode, for an app's onTrimMemory callback.
The budget stays strictly positive (the LRU buffers exist for the session;
load_layer keys shared-slot vs LRU off cache_max_ == 0, so a runtime zero
must never happen). Budget and resize count are surfaced in Stats and the
moe-cache summary line under auto.
2026-07-13 08:20:30 +02:00
Helldez
e1d78d1855 feat(moe): auto cache budget from available memory (--cache-mb auto)
Instead of a hand-tuned --cache-mb, size the expert cache to the device.
With --cache-mb auto the engine, once the full expert-set size is known,
sets the budget to (available RAM - cache_floor_mb, default 1536) clamped to
[cache_min_mb, total expert bytes]; if memory is unknown it falls back to the
floor. cache_auto is mutually exclusive with an explicit cache_mb, and counts
as 'cache on' for the --prefetch/--spec-gate requirement. Adds --cache-floor-mb.
This is the knob that keeps the phone responsive without the user guessing a
number that OOMs on one model and underfills on another. Runtime shrink under
pressure follows in a later commit; this lands the init-time sizing.
2026-07-13 08:13:41 +02:00
Helldez
7d6a4c68c9 feat(moe): auto-disable speculative gating when router recall stays low
Speculative gating only pays for its extra I/O and per-layer graph barrier
when the router it forecasts is actually predictable. On architectures where
cross-layer routing correlates weakly the predictor wastes reads for no
throughput gain, and previously nothing turned it off. Add a self-governor:
once at least spec_recall_warmup predictions have been scored (past prompt-
transition noise), if cumulative recall is below spec_recall_min_pct the
engine latches spec-gating off for the rest of the run — which also drops
the router-input barrier. Defaults 75/512; 0 disables the check. Exposed as
--spec-recall-min and surfaced in the moe-spec-gate summary line. New gate
G6d proves the mid-run on-to-off transition is byte-identical (the governor
fires at 25 percent recall on the synthetic model).
2026-07-13 08:00:50 +02:00
Helldez
b82f8ca224 style: clang-format the session, prefetch and spec-gating sources 2026-07-12 20:23:12 +02:00
Helldez
466316ab0f feat(moe): speculative gating — predict the next layer's experts from its router
Where temporal prefetch guesses the next layers from the previous token's routing, speculative
gating (--spec-gate) predicts the NEXT MoE layer directly: it isolates the router-input node,
reads the current layer's hidden state (the residual stream changes slowly between layers), runs
the next layer's gate on it, and prefetches its top-k. The prediction only feeds the byte-safe
prefetch queue, so a mispredict wastes a read but never changes output.

Recipe-driven, no hardcoding: each row carries the router-input node name, the mmap-resident gate
weight (and per-channel scale), and a pre-transform. qwen3moe/qwen2moe route on the post-norm
hidden (kNone); gemma4 routes on attn_out with rms_norm * 1/sqrt(n_embd) * scale (kRmsScaled) and
has interleaved dense layers, handled by targeting the next bound MoE layer. n_expert_used and the
rms epsilon come from generic gguf metadata keys. F32/F16 gate weights supported; a quantized gate
disables the feature (logged once). Recall (predicted vs actual routing) is measured and reported.

Gates G6a/b/c pass byte-identically on qwen3moe and gemma4 (b forces synchronous completion to
exercise predict→integrate→hit). End-to-end on the tiny models: gemma4 reaches 91% router recall,
confirming the transform, while output stays byte-identical.
2026-07-12 20:18:35 +02:00
Helldez
4d57da5edb feat(moe): temporal prefetch of the next layers' experts on idle lanes
While a token computes layer l, the previous token's routing at layers l+1..l+K is a strong
predictor of this token's — so speculatively read those experts on the otherwise-idle I/O lanes,
turning the next layer's read into a cache hit. RouterHook records each layer's last-token
routing and calls the new IExpertSource::prefetch() hint; ExpertStreamSource services it with a
low-priority speculative queue that never delays real work: all LRU mutation stays on the eval
thread (prefetch commits + enqueues, quiesce integrates completed reads and discards the rest at
the next real load), while workers only read bytes and yield the moment a real batch arrives. A
speculative slice is the identical read (same file offset, same buffer) a real miss would issue,
so integration is byte-safe by construction. --prefetch K (env BMOE_PREFETCH), cache required;
telemetry reports speculative MiB and useful-hit rate. Gates G5a/b/c pass byte-identically on
qwen3moe and gemma4 — G5c forces synchronous completion to exercise integrate-then-hit
deterministically (observed 3 experts prefetched, 1 useful).
2026-07-12 20:05:28 +02:00
Helldez
2363471f79 feat(cli): --session keeps the process alive across prompts over stdin
Adds an interactive mode that opens one Session and serves many prompts, so the model load
and the expert-cache warm-up are paid once instead of per prompt. Requests are one JSON line
per prompt on stdin ({cmd:generate|cancel|close}); responses use the BMOE_* line protocol on
stdout (BMOE_READY / BEGIN / PROGRESS / LOAD / DONE / ERROR), a superset of --progress so the
existing per-token parser is reused. A stdin reader thread applies cancel immediately (via the
session abort callback) while the main thread drives generation. The one-shot argv path is
unchanged. Verified end-to-end on the tiny model: two warm generates produce byte-identical
output and the cache hit rate carries over between them.
2026-07-12 19:40:36 +02:00
Helldez
7a744e5a59 feat(moe): surface cache-management cost as a mgmt telemetry term
vm-commit, eviction and LRU bookkeeping were hidden inside the compute residual, making the
first tokens after prefill read as pure matmul when the real cost is cache churn. Time that
staging work into mgmt_ns_ and report it per token (mgmt_ms), in the CSV, and in the
moe-stream summary (cache mgmt). compute_ms is now documented as a residual
(wall - io - mgmt, or wall - stall - mgmt under overlap), not a measured quantity. Bytes
served are unchanged; gates G1-G4 pass byte-identically on qwen3moe and gemma4.
2026-07-12 19:28:03 +02:00
Helldez
fcc80ad0a5 feat(cli): add --overlap flag and stall telemetry
--overlap (env BMOE_OVERLAP, flag wins) enables the overlapped decode path. The
CSV sink appends stall_ms as the last column and stall_s/tok to the summary line
(positional append, backward compatible); the run summary prints per-token stall.
2026-07-12 08:54:19 +02:00
Helldez
481bf76b02 fix: honor the thinking toggle for all model families
The chat template was always rendered with enable_thinking=true, so a reasoning
model kept emitting its thinking channel; Gemma's leaked into the answer because
the display-time parser does not recognise its reasoning format. Thread a think
flag through RunConfig and set inputs.enable_thinking, and drive it from a
--no-think CLI flag. The app now passes --no-think for every architecture instead
of appending Qwen's /no_think only, so the toggle works for Gemma too.
2026-07-11 22:44:43 +02:00
Helldez
4fe7cfc393 feat: report prefill, model-load and TTFT in the run summary
Additive telemetry only. runtime times model load + streaming setup as one
'load' span and prefill as a separate span, so TTFT = load + prefill; the CLI
prints them and the CSV summary carries n_prompt/load_s/prefill_s/prefill_tps.
No change to the streaming or decode path.
2026-07-11 22:44:43 +02:00
Helldez
4b0ba10161 feat(engine): apply the model family's chat template, not just Qwen ChatML
--chatml wrapped every prompt in Qwen ChatML, so instruct models with a
different turn format (Gemma 4's <start_of_turn> turns) produced raw-completion
output. Make the wrapper arch-aware: read general.architecture and emit the
family's turn format (gemma -> <start_of_turn>, everything else -> ChatML).

The engine has no Jinja engine, so it can't render the gguf's embedded
tokenizer.chat_template (Gemma 4's is a 16 KB Jinja program with tool-call
macros that llama.cpp's non-Jinja llama_chat_apply_template cannot parse); the
small stable turn format per family is the pragmatic, dependency-free path. Add
a branch when supporting another instruct family.
2026-07-11 17:05:01 +02:00
Helldez
d2d5dc9ee7 style: apply clang-format across sources 2026-07-10 18:34:36 +02:00
Helldez
5273af6487 feat(cli): bmoe-cli with streaming flags, ChatML, and machine telemetry 2026-07-10 18:18:19 +02:00