feat(engine): self-speculative decoding — the model's own MTP head, or n-gram lookup (#134)

* feat(engine): MTP self-speculative decoding for Qwen3.5/3.6 (proposal)

Qwen3.5/3.6 ship a trained multi-token-prediction block inside the gguf. With
--mtp that head drafts --mtp-draft continuation tokens and the target verifies
all of them in one wider decode, confirming the longest prefix whose argmax
equals what the target itself would have produced. Nothing is approximated and
no weight is skipped, so the quality is the full model's — but it is NOT
byte-identical the way --overlap and --prefetch are, and must not be used in a
byte-identity gate: a verify pass evaluates 1+N positions in one batch, and a
batched matmul is not bit-identical to N single-token ones, so a near-tie can
flip. Off by default.

The prize is that a decode's dominant cost, moving the dense weights and the
routed expert slices, is paid once per group instead of once per token. The
counterweight is that the verify positions route independently, so a layer's
read set widens toward N*k wherever adjacent tokens disagree, and the draft
pass routes through the MTP block's own expert layer on top. Measured on the
desktop host (DRAM-bandwidth-bound, model streamed at ~1.4x RAM): +15.1% at
draft 3 with the host's best recipe (7.12 -> 8.19 tok/s), +29% without the
lossy drop knob, acceptance falling from 71% at draft 2 to 52% at draft 4, and
flash bytes per token rising 19.7 -> 33.7 MiB as the widening predicts. Draft 3
is the optimum here; 4 is worse than 2. On a flash-I/O-bound phone that balance
can invert, so the flag ships off pending the device A/B.

The orchestration is llama.cpp's own (common/speculative.h, public headers
only): no fork, no patch, no submodule bump. Self-speculation is one model with
two contexts over it — the target, created with n_rs_seq so a rejected tail is
rewound from a bounded snapshot rather than replayed, and a draft context
created with ctx_type = LLAMA_CONTEXT_TYPE_MTP. The engine builds the draft
context itself rather than through common_speculative_init_from_params because
the eval callback is per-context: the streamer only sees the MTP block's expert
layer if the draft context carries the same cb_eval.

The MTP block is streamed like any other layer. It sits at layer index n_layer,
contiguous with the trunk and using the same tensor naming, so the hook and the
expert source are sized n_layer + n_layer_nextn; left at n_layer its experts
stay silently mmap-resident. Two consequences that are easy to get wrong: the
capture warm-up has to run on the draft context too (the MTP graph is built
nowhere else), and prefill is fed through the driver so the draft context's KV
reaches the last prompt position.

The loop accepts BEFORE catching the draft context up, so the catch-up runs on
the accepted prefix instead of the whole verify batch. Acceptance depends only
on the target's logits, which are already in hand once the decode returns, and
the rejected tail was being computed only to be deleted a few statements later.
The resulting state is identical — the driver seeds from row
min(n_accepted, n_rows-1), the same row under either batch, and the surviving
KV is exactly the range the rollback used to carve out — while skipping
n_draft - n_accepted positions through the MTP block per group. Since that
block carries its own MoE FFN, on a streamed device those are expert reads that
no longer happen. It also removes the draft context's rollback entirely: it is
never given a tail to drop.

Requires an MTP-converted gguf (most quantisations strip the nextn tensors) and
greedy decoding; both are rejected at load with a message rather than silently
ignored, as is a n_ubatch narrower than the verify batch, which would split the
graph back into single-token passes and spend the draft for nothing.

Telemetry: an "mtp:" summary line, an mtp_batch per-token CSV column (a verify
decode's whole cost is charged to its group's first row, the rest carry zeros),
mtp_drafted / mtp_accepted / mtp_decodes in the CSV trailer and in BMOE_DONE,
and mtp / mtp_draft_max in the CSV preamble. The Android app exposes the flag
and the draft width, off by default.

Host gates pass. Validated on Qwen3.6-35B-A3B-MXFP4 with the streamed recipe:
draft 1 and draft 3 produce identical text, which is the invariant a broken
accept/rollback path would violate. Device A/B still owed.

* perf(mtp): shrink the draft context, make its cost measurable, record the device verdict

The first on-device A/B says MTP loses at every draft width, and the counters
say why. Same gguf with the flag on and off, shipping recipe (overlap, 3000 MiB
cache, pinned dense, drop 0.75), Qwen3.6-35B-A3B-Q4_K_M streamed:

    off             5.82 / 6.14 tok/s   69.3 MiB/tok    69-109 majflt/tok
    --mtp-draft 2   5.59                93.8            230
    --mtp-draft 3   4.38               106.6            633

Speculation is working - 2.35-2.52 tokens per verify decode, 52-69% acceptance
- and still losing, because the prize does not exist in this regime.
stall_s/tok is 0.025-0.027 in every one of those runs, MTP on or off: 11-16% of
the token. This configuration is compute-bound, and what MTP amortises is weight
movement. The costs meanwhile are real and monotonic in the draft width: the read
set widens (+35%, +54% flash bytes per token), CPU per token rises (+28%, +67%),
and the draft context's memory tips the device into a fault storm.

Two things follow, and both are engine bugs rather than facts of nature.

The draft context's graph width drops from 256 to 32. Compute buffers are
reserved for the widest ubatch and the dominant term scales with
ubatch x vocabulary; on device that reservation measured 493 MiB - for a context
that evaluates ONE token per draft step and is handed at most 1 + draft_max
positions by the catch-up, with no logits asked for. Only prefill ever feeds it a
wide batch, and that is one layer, so splitting it costs very little. On this
engine memory is never free: it is the expert cache's, and the cache is what
decides whether the widened verify read set is a hit or a flash read.

And the cost of speculation is now measured instead of inferred. Drafting happens
between decodes, so it never entered wall_ms and tok/s never included it - a
speculated run could report a rate the user was not experiencing. New
mtp_draft_ms per-token column (a slice of loop_overhead_ms, not an addition),
mtp_draft_s/tok in the CSV trailer, mtp_draft_s_tok and loop_overhead_s_tok in
BMOE_DONE, and a second "mtp:" summary line printing the effective rate next to
the reported one.

Adds --mtp-p-min F, which stops drafting once the head's confidence in what it is
proposing falls below F. The draft loop already had this floor and the engine was
passing 0, so it always drafted the full width however unsure the head was - with
roughly half the drafts rejected at draft 3, that is the cheapest waste available
to cut. On a streamed device it pays twice: a draft not made is a pass through the
MTP block (which carries its own MoE FFN, so its own expert reads) that never
happens, AND one fewer independently routed position in the verify batch. Default
0, the setting the host numbers were measured at; the useful value is a property
of a device's balance between drafting cost and acceptance, so it is a knob to
measure rather than a constant to guess.

The Android app now reads the mtp_* keys it was already being sent: acceptance,
tokens per pass, and the effective rate. Before this the UI could not tell whether
speculation had run at all - only the session CSV could - which made the A/B this
commit reports impossible to run from the phone.

Neither mitigation changes the regime. The honest expectation is nearer
break-even, not a win, and the flag stays off by default.

Host gates pass. Note the noise floor: the two off runs did byte-identical work
and still differ by 5.6% in tok/s, and the runs were back-to-back without thermal
gating - the mechanism counters are the trustworthy part, not the exact deltas.

* perf(mtp): split the drafting flash cost from the widened verify batch

A speculated run streams more bytes per token for two unrelated reasons: the
MTP block carries its own MoE FFN, so every draft pass routes experts of its
own, and the verify batch widens the trunk's read set wherever adjacent
positions disagree. They need opposite fixes -- a narrower draft attacks the
first, only better agreement attacks the second -- and the route trace can
separate neither, since its framing brackets the target decode while the head
only ever runs in the draft context.

Measure the head's share directly by bracketing both drafting passes with the
expert source's byte counter, and report it as a third mtp: summary line.

Also record the branch-deletion rule in AGENTS.md: a branch list should only
show work in flight, and a rejected PR loses nothing.

* feat(engine): n-gram prompt-lookup draft source, and the measurement that closes it

The flash split added last commit said where MTP's cost actually is: at draft 3 on
the host, the head's own routing was 2.9% of the extra bytes a speculated run
streams and the widened verify batch was the other 97.1%. So a cheaper draft
producer is worth almost nothing, and the only property that could matter is one
the head does not have -- the ability to decline to draft at zero cost.

--ngram is that source. It takes the last few tokens, finds where that run occurred
before in the prompt or in what has been generated, and proposes whatever followed.
No head, no draft context, no decode, no expert read, and it works on any gguf
including the ones --mtp refuses for want of a nextn block. Below --ngram-min-match
it proposes nothing and the step falls through to a plain single-token decode.

Measured on the host, Qwen3.6-35B-A3B-MXFP4 streamed, 256 greedy tokens, cells
back-to-back with off run twice:

    prose        off 5.80 / 6.59    mtp3 7.32 eff    ngram3 6.51  (cov 7.4%)
    copy-heavy   off 5.45 / 5.65    mtp3 6.43 eff    ngram3 5.24  (cov 15%)

The zero-cost claim holds exactly -- mtp_draft_s/tok reads 0.0000 in every n-gram
cell, against 0.020-0.023 for the head plus the ~500 MiB of expert cache its draft
context takes. But the floor turns out to be per STEP, not per run: the 15% of steps
that did draft widened the read set to 67.2 MiB/token against 48-58 at baseline and,
at 44% acceptance, bought 1.20 tokens per decode. That is not enough to earn the
widening back, and a modest fraction of such steps sinks the run.

A --ngram-min-match sweep settles it rather than leaving it open. Raising the gate
3 -> 5 -> 8 lifts acceptance 44% -> 75% while coverage collapses 15% -> 3.4%, and
narrowing to --draft 1 reaches 82.6% -- the head's own figure on this prompt. Every
cell climbs toward baseline from BELOW and none crosses it; the best configuration
found lands on the floor. A knob whose optimum is its own disablement is not a
tuning problem. Acceptance, not drafting cost, is what pays for a widened batch, and
what a trained head buys is being right often enough to justify a batch that has
already been widened.

--ngram ships off. It is kept because it is the only speculation available on a
model with no head, because the per-step floor is real, and because the counters it
adds make the next speculation claim falsifiable.

Wiring. MtpConfig became SpecConfig with DraftSource {none, mtp, ngram}, and
--mtp-draft became --draft: the width belongs to the verify batch, not to whoever
filled it. --mtp and --ngram are rejected together rather than resolved by flag
order. In the session the gate split in two -- spec_on (wide batch, acceptance,
rollback: both sources) against mtp_on (draft context, common/speculative.h, the
catch-up: the head only) -- which is what lets the n-gram source reuse the whole
verify half while allocating nothing.

A step that drafts nothing now takes the plain path: llama_batch_get_one with a
logits row of -1, byte for byte the unspeculated decode. It used to build the wide
batch anyway. Required for --ngram, and it tightens --mtp-p-min's zero-draft steps
for free.

The matcher is pure policy over token ids with no llama.cpp at all -- not even
llama.h, since llama_token is int32_t -- so it sits on the clean side of the seam,
adds no dependency on the common layer, and is unit-tested with no model
(tests/ngram_test.cpp covers tie-breaks, clipping, self-match exclusion and the gate
boundary). Telemetry: spec= / spec_draft_max= / ngram_min_match= in the CSV
preamble, a new drafted_steps key in the trailer and BMOE_DONE, and an ngram: line
reporting coverage -- without which a delta cannot be divided by the fraction of the
run it applies to. The per-token and trailer counters keep their mtp_ names: they
always described the loop rather than a source, spec_* already means the temporal
prefetch in that trailer, and renaming would break every CSV already holding a
measurement. The Android setting became a three-way picker, migrating the old
boolean preference.

The device A/B agrees and adds a cost the host could not show. Thermally gated cells
(a 120 s settle, then a battery-temperature gate, so all six start between 35.3 and
36.4 C): prose 4.90 inside a 4.59-5.17 band, copy-heavy 3.14 against 4.43 -- a 29%
loss, worse than MTP's 18%. Major faults per token go 126 -> 1427 for a source that
allocates no draft context at all, and that is the rollback snapshots: n_rs_seq =
draft_max is asked for by ANY speculation, since rejecting a draft means rewinding the
KV, and on a hybrid attention/SSM model that snapshot is a real allocation scaling with
the context. The n-gram source escapes MTP's draft context but not the loop's own
memory, and on device that memory is the expert cache's.

The same run re-measured MTP with the thermal confound removed -- 3.64 effective
against 4.43, so the earlier device verdict was not an artefact of benching without a
cooldown gate -- and reproduced the flash split at 3.7% head against 96.3% widened
verify batch, matching the host's 2.9-3.0%.

Byte-identity gates pass; speculation stays out of them for the reason docs/mtp.md
gives.

The app's CSV configuration surface follows: the three new preamble keys get their own
glossary entries rather than falling through to the unexplained-key renderer, and the
draft source joins the short run label. A speculated run is not the same KIND of run --
under speculation a decode confirms a whole group, so its per-token rows are not even
accounted the same way -- and two compare legends differing by it must not read alike.
This commit is contained in:
Helldez 2026-08-02 00:09:39 +02:00 • committed by GitHub
parent 03753187b1
commit 49c72e7ce5
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
29 changed files with 2075 additions and 101 deletions

View file

@ -58,6 +58,14 @@ check). Nothing below needs a storage permission except the last option.
2. **Any other model** — under **Other model**, paste a direct gguf URL (e.g. a Hugging Face
`…/resolve/main/model.gguf` link), or pick a `.gguf` already on the device to import it.
You do **not** need a special file for **Guess ahead → Model's own head (MTP)**: the catalog's
Qwen3.6 entry already carries the `nextn` block the MTP head lives in, as do Qwen3.6's ordinary
quantisations generally. A gguf named `-MTP-` is the same head at a different quantisation.
On a model with no head — anything that is not Qwen3.5/3.6 — the engine refuses to open rather
than silently decoding one token at a time, so a wrong file fails immediately and says why.
**Guess ahead → Repeated text (n-gram)** has no such requirement: it guesses from the text
rather than from the weights, so it works on every model in the catalog.
In-app downloads and picker imports both land in the app's internal storage (`filesDir`, a
real f2fs/ext4 volume), so the streamed expert reads use O_DIRECT at full speed. Only models
read from the emulated external dirs (adb-pushed to `/sdcard/Download`) fall back to buffered

View file

@ -8,10 +8,10 @@ import android.content.Context
* pair of switches, which could express the same policy two ways (or contradict each other).
*/
enum class DenseWeights(val flag: String, val label: String, val blurb: String) {
MMAP("mmap", "Mmap (baseline)", "Leave the dense weights mmap'd; the kernel faults them in. Slow first tokens on a >RAM model — the A/B baseline."),
WARM("warm", "Warm at load", "Page-cache the dense weights once at load, so the first tokens don't fault them in a page at a time. Best when the model fits in RAM."),
ANON("anon", "Anon (O_DIRECT)", "Read the dense weights via O_DIRECT into our own buffers so a reclaim swaps to zram (fast) instead of a slow flash refault. Costs a private copy. The default — it wins on >RAM models."),
AHWB("ahwb", "Pinned (dma-buf)", "As Anon, but into dma-buf memory the kernel cannot reclaim at all — not even to zram, which is what Anon still pays for. Measured +17.9% on a long generation; the gain needs a real conversation to appear, since short turns never build up enough reclaim. Off by default: measured on one device."),
MMAP("mmap", "Mmap (baseline)", "Leave the always-needed weights to the kernel, which faults them in as they are touched. The plain baseline to compare the others against."),
WARM("warm", "Warm at load", "Read them once at load so the first tokens do not fault them in a page at a time. Suits a model that fits in memory comfortably."),
ANON("anon", "Anon (O_DIRECT)", "Hold them in the app's own memory instead of the file cache. When the system reclaims, they are compressed rather than dropped and re-read from flash. Costs a private copy; the default."),
AHWB("ahwb", "Pinned (dma-buf)", "As Anon, but in memory the system cannot reclaim at all — not even by compressing it, which Anon still pays for. The gain shows up over a long conversation, once reclaim has had time to bite."),
}
/**
@ -51,6 +51,35 @@ data class AppSettings(
// share itself). Stored as an Int because the settings are integer rungs; the flag takes a
// fraction. LOSSY and cache-dependent — it changes the output, and not reproducibly.
val dropColdPct: Int = 75,
// Which source drafts for self-speculation: "off", "mtp" or "ngram". Both verify the same way —
// one wider decode, greedy acceptance — and differ only in what a draft costs.
//
// "mtp" uses the model's own head, so it needs a gguf carrying the nextn block, which Qwen3.5/3.6
// do in their ordinary quantisations — the catalog's Qwen3.6 included, so no special "-MTP-"
// download. On anything else the engine refuses to open, so the UI states the requirement rather
// than letting the session fail.
//
// "ngram" drafts by looking the recent tokens up in the prompt and in what has been generated:
// no head, no draft context, no expert read, and it works on every model. It also drafts nothing
// when it has no confident match, so a step without one costs exactly a plain decode — the
// property the head does not have.
//
// Off by default. The head is a clear win where decode is DRAM-bound (desktop, +15%) and lost on
// this phone at every draft width; the lookup exists because the measured reason for that loss
// was the widened verify batch, not the drafting — which is what this toggle now lets an A/B
// separate.
val spec: String = SPEC_OFF,
// Tokens drafted per verify pass, shared by both sources. 3 is the measured optimum for the head
// on desktop: acceptance falls as the draft widens, and past its horizon the extra drafts are
// paid for and thrown away. On this phone the device A/B measured 2 better than 3 and both below
// the baseline, which is why the confidence floor below exists.
val mtpDraft: Int = 3,
// Stop drafting when the head's own probability for what it is proposing drops below this
// percentage (0 = never stop, draft the full width however unsure it is). Makes the width
// adaptive per step, which on this device pays twice: a draft not made is one fewer pass
// through the MTP block AND one fewer independently routed position in the verify batch. 0 is
// the setting the desktop numbers were measured at, so it stays the default until the A/B says.
val mtpPMinPct: Int = 0,
val thinking: Boolean = false, // reasoning; off passes --no-think (enable_thinking=false)
val metricsCsv: Boolean = true, // write the engine's per-token CSV for this session (--csv)
) {
@ -121,6 +150,15 @@ data class AppSettings(
// of the uniform share; the setting is stored as a percentage.
if (dropColdPct > 0 && cacheOn) a += listOf("--drop-cold-experts", (dropColdPct / 100.0).toString())
}
// Outside the streaming block on purpose: speculation is a decode-loop change, not a
// residency policy, so it applies to the mmap baseline too — which is what makes an A/B of
// the two against each other meaningful.
if (spec == SPEC_MTP) {
a += listOf("--mtp", "--draft", mtpDraft.toString())
if (mtpPMinPct > 0) a += listOf("--mtp-p-min", (mtpPMinPct / 100.0).toString())
} else if (spec == SPEC_NGRAM) {
a += listOf("--ngram", "--draft", mtpDraft.toString())
}
return a
}
@ -132,7 +170,8 @@ data class AppSettings(
*/
fun sessionSignature(modelPath: String): String =
listOf(modelPath, mmap, cacheMb, cacheCeilMb, ioThreads, threads, nExpertUsed, sessionCtx, oDirect,
overlap, denseWeights, prefetchLayers, predictPrefetch, predictSpecMax, dropColdPct)
overlap, denseWeights, prefetchLayers, predictPrefetch, predictSpecMax, dropColdPct,
spec, mtpDraft, mtpPMinPct)
.joinToString("|")
fun save(ctx: Context) {
@ -149,6 +188,7 @@ data class AppSettings(
.putInt("predictSpecMax", predictSpecMax)
.putInt("dropColdPct", dropColdPct)
.putInt("sessionCtx", sessionCtx)
.putString("spec", spec).putInt("mtpDraft", mtpDraft).putInt("mtpPMinPct", mtpPMinPct)
.putBoolean("thinking", thinking)
.putBoolean("metricsCsv", metricsCsv)
.apply()
@ -180,6 +220,22 @@ data class AppSettings(
// truncate most answers mid-sentence, which reads as broken rather than slow.
const val DEFAULT_N_PREDICT = 128
// The draft sources, as stored. Strings rather than an enum ordinal: a preference that
// survives an app update must not depend on the order this list happens to be written in.
const val SPEC_OFF = "off"
const val SPEC_MTP = "mtp"
const val SPEC_NGRAM = "ngram"
val SPEC_CHOICES = arrayOf(SPEC_OFF, SPEC_MTP, SPEC_NGRAM)
// Draft widths worth offering. The useful range is small and not monotonic: acceptance
// falls as the draft widens while tokens-per-decode rises, and on desktop the two cross at
// 3 — 4 measured WORSE than 2. Stopping at 5 keeps the picker honest about that.
val MTP_DRAFT_CHOICES = intArrayOf(1, 2, 3, 4, 5)
// Confidence floors worth offering, as percentages. 0 keeps the current behaviour (draft
// the full width unconditionally); the rest trade speculative reach for wasted drafts.
val MTP_P_MIN_CHOICES = intArrayOf(0, 40, 60, 80)
/**
* A fresh CSV path for a session about to open, under the app's own external files dir —
* no permission needed to write, and `adb pull`-able without root:
@ -267,6 +323,17 @@ data class AppSettings(
predictSpecMax = p.getInt("predictSpecMax", d.predictSpecMax),
dropColdPct = p.getInt("dropColdPct", d.dropColdPct),
sessionCtx = p.getInt("sessionCtx", d.sessionCtx),
spec = run {
val saved = p.getString("spec", null)
when {
saved != null && saved in SPEC_CHOICES -> saved
// Migrate the old boolean pref from an install that only had the head.
p.getBoolean("mtp", false) -> SPEC_MTP
else -> d.spec
}
},
mtpDraft = p.getInt("mtpDraft", d.mtpDraft),
mtpPMinPct = p.getInt("mtpPMinPct", d.mtpPMinPct),
thinking = p.getBoolean("thinking", d.thinking),
metricsCsv = p.getBoolean("metricsCsv", d.metricsCsv),
)

View file

@ -102,6 +102,10 @@ object ConfigFields {
ConfigField("top_p", "sampling candidates kept by cumulative probability"),
ConfigField("seed", "sampling seed. Inert at temperature 0; otherwise the only thing that makes a run repeatable"),
ConfigField("compute_trace_layers", "per-layer compute tracing. It instruments the graph, so a traced run is a diagnostic rather than a benchmark"),
ConfigField("spec", "which source drafted for speculative decoding: off, mtp (the model's own trained head) or ngram (repeated text looked up in the prompt and the answer so far). Both verify identically, so this changes what a draft cost, not what was accepted"),
ConfigField("spec_draft_max", "tokens drafted per verify pass. The pass is one wider than this, and every position in it routes its own experts — which is what speculation trades for confirming several tokens at once"),
ConfigField("mtp_p_min", "the head stopped drafting below this confidence, making the width adaptive. 0 = always draft the full width. Applies to spec=mtp only"),
ConfigField("ngram_min_match", "shortest run of repeated tokens the lookup would draft from. Below it the step drafted nothing and cost exactly an unspeculated decode. Applies to spec=ngram only"),
)
private val byName = all.associateBy { it.name }

View file

@ -47,6 +47,10 @@ internal class Csv(
// legends that differ only by this used to read as identical (#136).
if (info["predict_prefetch"] == "1") "predict" else null,
info["drop_cold_frac"]?.takeIf { (it.toFloatOrNull() ?: 0f) > 0f }?.let { "drop $it" },
// Same argument, and stronger: under speculation a decode confirms a whole group, so
// the per-token rows are not even accounted the same way (mtp_batch). Two runs that
// differ by this must never read as the same kind of run.
info["spec"]?.takeIf { it != "off" }?.let { "$it ${info["spec_draft_max"]}" },
).joinToString(" · ")
}

View file

@ -257,6 +257,19 @@ class RunService : Service() {
val cacheBudgetMib = o.optDouble("cache_budget_mib", -1.0)
val majfltPerTok = o.optDouble("majflt_tok", -1.0)
val cpuSPerTok = o.optDouble("cpu_s_tok", -1.0)
// Self-speculation. All 0 with MTP off, which is what makes them safe to read
// unconditionally — and reading them is the only way the UI can say whether
// speculation ran and what it earned.
val mtpDrafted = o.optLong("mtp_drafted", 0)
val mtpAccepted = o.optLong("mtp_accepted", 0)
val mtpDecodes = o.optLong("mtp_decodes", 0)
val mtpDraftSTok = o.optDouble("mtp_draft_s_tok", 0.0)
val draftedSteps = o.optLong("drafted_steps", 0)
val loopOverheadSTok = o.optDouble("loop_overhead_s_tok", 0.0)
// What the user actually waits: tok_s counts decode time only, so drafting — which
// happens between decodes — is invisible to it. With MTP off the gap is ~0 and this
// equals tokS; with it on the two can differ by tens of percent.
val effTokS = if (tokS > 0) 1.0 / (1.0 / tokS + loopOverheadSTok) else -1.0
// Time-to-first-token: the model load plus this turn's prompt prefill.
val ttft = if (loadS >= 0 && prefill >= 0) loadS + prefill else -1.0
val cancelled = o.optBoolean("cancelled")
@ -272,6 +285,21 @@ class RunService : Service() {
if (ttft >= 0) append(String.format(loc, " | TTFT %.2fs", ttft))
if (hit >= 0) append(String.format(loc, " | cache %.0f%%", hit))
if (cancelled) append(" | cancelled")
if (mtpDecodes > 0 && mtpDrafted > 0) {
append(String.format(loc, "\nGuessing: %d/%d kept (%.0f%%), %.2f tok per pass",
mtpAccepted, mtpDrafted, 100.0 * mtpAccepted / mtpDrafted,
tokens.toDouble() / mtpDecodes))
if (mtpDraftSTok > 0) {
append(String.format(loc, " | guessing costs %.3fs/tok → %.2f tok/s real",
mtpDraftSTok, effTokS))
}
// Only the n-gram source ever abstains, so a coverage below every step is what
// says the rest of the turn ran at the plain, unspeculated cost.
if (draftedSteps in 1 until mtpDecodes) {
append(String.format(loc, " | guessed on %.0f%% of passes",
100.0 * draftedSteps / mtpDecodes))
}
}
}
// Compact one-line metrics shown under this turn's answer in the transcript.
val turnMetrics = buildString {
@ -282,6 +310,12 @@ class RunService : Service() {
}
if (nPast >= 0) append(String.format(loc, " · ctx %d/%d", nPast, sessionCtx))
if (hit >= 0) append(String.format(loc, " · cache %.0f%%", hit))
// With speculation on, the headline rate leaves out the drafting between decodes;
// show what the turn really ran at, and how often the guesses were right.
if (mtpDecodes > 0 && mtpDrafted > 0) {
if (effTokS > 0) append(String.format(loc, " · %.1f real", effTokS))
append(String.format(loc, " · %.0f%% kept", 100.0 * mtpAccepted / mtpDrafted))
}
if (cancelled) append(" · cancelled")
}
val tel = telemetry.current.copy(
@ -290,6 +324,9 @@ class RunService : Service() {
prefillTps = prefillTps, ttftS = ttft, readMib = readMib,
cacheResidentMib = cacheResidentMib, cacheBudgetMib = cacheBudgetMib,
avgMajfltPerTok = majfltPerTok, avgCpuSPerTok = cpuSPerTok,
mtpDrafted = mtpDrafted, mtpAccepted = mtpAccepted, mtpDecodes = mtpDecodes,
draftedSteps = draftedSteps,
mtpDraftSPerTok = mtpDraftSTok, loopOverheadSPerTok = loopOverheadSTok,
)
RunBus.update {
val answer = if (text.isNotEmpty()) text else it.answer

View file

@ -51,7 +51,7 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
// inert (the CLI omits --moe-stream and all sub-flags), so they are disabled.
SwitchRow(
"mmap baseline (no streaming)",
"Load the whole model via llama.cpp mmap — the baseline to compare against",
"Load the model the ordinary way, without streaming experts. The baseline to compare against",
current.mmap,
) { onChange(current.copy(mmap = it)) }
@ -69,14 +69,13 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
enabled = stream,
) { onChange(current.copy(cacheMb = it)) }
Text(
"Larger cache = fewer flash reads per token, but more RAM — and RAM the kernel takes " +
"back is paid for twice. Auto sizes to free RAM once at load, then holds. On a >RAM " +
"model, off is usually the ceiling. " +
"Experts already in memory cost no read at all, so a bigger cache means less waiting on " +
"flash — paid for in memory the rest of the system no longer has. Auto sizes it once " +
"at load from what is free, then holds that budget. " +
if (AppSettings.cacheNeedsForce(current.cacheMb))
"500 and 1000 are below the engine's floor (--force-cache): a cache under one " +
"token's routed experts can only thrash. Kept for measuring where the cache " +
"stops earning its memory."
else "500 and 1000 sit below the engine's floor and need --force-cache.",
"The smallest rungs sit below the engine's own floor: a cache too small to hold " +
"one token's experts evicts them before they are reused, and only churns."
else "The smallest rungs sit below the engine's floor and have to be forced.",
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
IntSetting(
@ -85,25 +84,25 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
enabled = stream && current.cacheMb == AppSettings.CACHE_AUTO,
) { onChange(current.copy(cacheCeilMb = it)) }
Text(
"Upper bound on the Auto budget at load, so it does not over-ask on devices with " +
"tight free RAM (MemAvailable counts the model's own weights as free).",
"Caps what Auto may claim. The system reports the model's own mapped weights as free " +
"memory, so left uncapped Auto can ask for more than really exists.",
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
IntSetting("Parallel I/O lanes", AppSettings.IO_CHOICES, current.ioThreads, enabled = stream) {
onChange(current.copy(ioThreads = it))
}
Text(
"Number of parallel flash-read threads for the expert stream.",
"How many expert reads are in flight at once. More lanes help only until the flash itself is saturated.",
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
SwitchRow(
"Direct I/O (O_DIRECT)",
"Bypass the page cache for expert reads. Falls back to buffered automatically if unsupported",
"Read experts straight from flash, so the system does not keep a second copy of what the cache above already holds. Falls back automatically where unsupported",
current.oDirect, enabled = stream,
) { onChange(current.copy(oDirect = it)) }
SwitchRow(
"I/O–compute overlap",
"Read the next experts while the current layer computes, hiding read latency",
"Start the next reads while the current layer is still computing, so waiting for flash happens behind the work instead of after it",
current.overlap, enabled = stream,
) { onChange(current.copy(overlap = it)) }
LabeledDropdown(
@ -124,15 +123,15 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
enabled = stream && cacheOn && !current.predictPrefetch,
) { onChange(current.copy(prefetchLayers = it)) }
Text(
"Experimental. Bets each layer will reuse the experts it picked for the previous " +
"token and reads them ahead on idle lanes — right ~40% of the time. Needs the cache on.",
"Experimental. Bets a layer will reuse the experts it picked for the previous token, and " +
"reads them ahead on lanes that would otherwise sit idle. Needs the cache.",
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
SwitchRow(
"Predictive prefetch (experimental)",
"Instead of betting on the previous token, asks the next layer's own router one " +
"layer early — right ~85% of the time. Reads ahead within the budget below and " +
"keeps what the prediction names. Needs the cache; replaces temporal prefetch.",
"Experimental. Rather than betting on the previous token, runs the next layer's own " +
"router early to ask which experts it will actually want. Reads ahead within the " +
"budget below and protects what it names. Needs the cache; replaces the above.",
current.predictPrefetch, enabled = stream && cacheOn && current.prefetchLayers == 0,
) { onChange(current.copy(predictPrefetch = it)) }
if (current.predictPrefetch) {
@ -147,15 +146,59 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
enabled = stream && cacheOn,
) { onChange(current.copy(predictSpecMax = it)) }
Text(
"How much flash the prediction may spend per layer. 0 reads nothing ahead and only " +
"protects predicted experts already in RAM from eviction. Measured: reading ahead " +
"lost its matched A/B — the flash has no spare bandwidth — so 0 is the honest default.",
"How much reading ahead the prediction may pay for. At zero it reads nothing and only " +
"keeps the experts it names from being evicted — the safe setting, since reading " +
"ahead competes for the same flash the current token is waiting on.",
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
Section("Speed / quality") {
// Speculation sits at the top of this section because it is the only entry here that
// does NOT trade quality: it changes how many tokens a decode confirms, not what the
// model computes. Not gated on the streamer — it is a decode-loop change, so it
// applies to the mmap baseline too.
LabeledDropdown(
"Guess ahead",
listOf("Off", "Model's own head (MTP)", "Repeated text (n-gram)"),
AppSettings.SPEC_CHOICES.indexOf(current.spec).coerceAtLeast(0),
) { onChange(current.copy(spec = AppSettings.SPEC_CHOICES[it])) }
Text(
"Guess the next few tokens, then check the whole group in one pass and keep only what " +
"the model itself would have produced. Nothing is skipped or approximated. The gain " +
"is reading the weights once for several tokens instead of once each; the cost is " +
"that checking several at a time makes every layer touch more experts.\n\n" +
"The head is a small extra part of the model trained to guess — accurate, but only " +
"some models carry it, and running it costs a pass of its own. The n-gram guess " +
"instead looks for the text repeating itself, which costs nothing at all and works " +
"on every model, but only has something to say when the model is quoting or " +
"editing. When it has nothing, that token runs exactly as if this were off.",
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
if (current.spec != AppSettings.SPEC_OFF) {
IntSetting(
"Tokens guessed per pass", AppSettings.MTP_DRAFT_CHOICES, current.mtpDraft,
) { onChange(current.copy(mtpDraft = it)) }
Text(
"How far ahead to guess. Further means more tokens confirmed per pass, but the " +
"guesses grow less reliable and a wrong one is paid for and thrown away. " +
"The best setting is rarely the largest.",
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
if (current.spec == AppSettings.SPEC_MTP) {
IntSetting(
"Guess only when confident", AppSettings.MTP_P_MIN_CHOICES, current.mtpPMinPct,
format = { if (it == 0) "Always guess" else "Above $it%" },
) { onChange(current.copy(mtpPMinPct = it)) }
Text(
"Stop guessing as soon as the model is unsure, instead of always filling the pass. " +
"A guess not made costs nothing and keeps the pass narrow, so fewer experts have " +
"to be read — but it also gives up the tokens that guess might have won.",
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
// Active-expert (top-k) override is a load-time kv_override, valid in both streaming
// and mmap mode, so it is not gated on the streamer.
IntSetting(
@ -168,8 +211,8 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
},
) { onChange(current.copy(nExpertUsed = it)) }
Text(
"Route fewer experts per token than the model's default. Faster and lighter on flash, " +
"but the output changes — a speed/quality trade-off.",
"Consult fewer experts per token than the model asks for. Cuts both the computing and " +
"the reading, and changes the reply — a deliberate trade of quality for speed.",
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
IntSetting(
@ -191,14 +234,12 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
(current.cacheMb == AppSettings.CACHE_AUTO || current.cacheMb > 0),
) { onChange(current.copy(dropColdPct = it)) }
Text(
"Reading an expert that is not already in RAM is what slows a token down. This skips " +
"one — but only if the router barely wanted it, below the chosen share of an even " +
"split. With 8 active experts an even split is 12.5%, so 75% skips anything under " +
"9.4%. Experts already in RAM always run; the strongest is never skipped.\n\n" +
"50% barely fires and changes little. 75% measured +55% faster and 100% +84% on one " +
"model, but the top rung discards twice as much of the routing for that extra speed." +
"Waiting for an expert that is not already in memory is what slows a token down. This " +
"skips one — but only when it is missing AND the router barely wanted it, below the " +
"chosen share of an even split. Experts already in memory always run, and the " +
"strongest is never skipped, so quality is spent only where it buys back a read." +
"\n\nThe reply changes, and unlike Active experts not the same way twice: what gets " +
"skipped depends on what the cache happened to hold.",
"skipped depends on what the cache happened to be holding.",
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
// The threshold is a share of 1/top-k, so a narrow routing changes what the same
@ -208,10 +249,9 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
val topk = ui.nExpertUsed
if (current.dropColdPct > 0 && topk != null && topk in 1..4) {
Text(
"⚠ This model routes only $topk experts per token, so ${current.dropColdPct}% here " +
"means \"skip below ${"%.1f".format(current.dropColdPct.toDouble() / topk)}%\" — a much " +
"bigger share of the reply than the 9.4% this setting was measured at (8 experts). " +
"Check the answers, or turn it off for this model.",
"⚠ This model routes only $topk experts per token, so the same share covers far more " +
"of the reply than it does on a model that routes many. Check the answers, or " +
"turn this off for this model.",
fontSize = 12.sp, color = MaterialTheme.colorScheme.error,
)
}
@ -258,9 +298,8 @@ fun SettingsScreen(current: AppSettings, onChange: (AppSettings) -> Unit, onBack
Section("Diagnostics") {
SwitchRow(
"Metrics CSV",
"Write one CSV per session with every token's timings, page faults (count and MiB), " +
"cache budget, and the memory split — anon (the expert cache), file (the model), " +
"swap. Takes effect on the next session; share it from the menu",
"Write one CSV per session: every token's timings, its page faults, the cache budget " +
"and where memory sat. Takes effect on the next session; share it from the menu",
current.metricsCsv,
) { onChange(current.copy(metricsCsv = it)) }
}

View file

@ -46,8 +46,40 @@ data class Telemetry(
// Run averages of the compute decomposition, from the final summary (BMOE_DONE); -1 until done.
var avgMajfltPerTok: Double = -1.0, // major page faults per token over the run
var avgCpuSPerTok: Double = -1.0, // CPU-seconds per token (summed across threads) over the run
// Self-speculation counters from BMOE_DONE; all 0 when speculation was off. Without these the UI
// cannot tell whether it ran at all, let alone whether it earned its keep. They describe the
// loop, not a source: the n-gram lookup and the MTP head report through the same fields.
var mtpDrafted: Long = 0,
var mtpAccepted: Long = 0,
var mtpDecodes: Long = 0,
// Passes that guessed anything. Below mtpDecodes it means the source abstained on the rest,
// which ran at exactly the unspeculated cost — only the n-gram source ever does that.
var draftedSteps: Long = 0,
// Seconds per token spent drafting, and the whole between-decode gap. tok/s counts decode time
// ONLY, so these are time the user waits that the headline rate does not include.
var mtpDraftSPerTok: Double = 0.0,
var loopOverheadSPerTok: Double = 0.0,
) {
val tokensPerSecond: Double get() = if (wallMs > 0) 1000.0 / wallMs else 0.0
/** Share of drafts the model itself confirmed, or -1 when nothing was drafted. */
val mtpAcceptancePct: Double get() =
if (mtpDrafted > 0) 100.0 * mtpAccepted / mtpDrafted else -1.0
/** Tokens confirmed per verify pass — what the speculation actually bought. */
val mtpTokensPerDecode: Double get() =
if (mtpDecodes > 0 && step > 0) step.toDouble() / mtpDecodes else -1.0
/**
* The rate the user actually experiences: decode time PLUS the gap between decodes, where
* drafting lives. Reporting only the decode rate flatters speculation, because the drafting it
* adds happens outside the measured window.
*/
val effectiveTokensPerSecond: Double get() {
if (avgTokensPerSecond <= 0) return -1.0
val perTok = 1.0 / avgTokensPerSecond + loopOverheadSPerTok
return if (perTok > 0) 1.0 / perTok else -1.0
}
}
/**