* feat(engine): MTP self-speculative decoding for Qwen3.5/3.6 (proposal)
Qwen3.5/3.6 ship a trained multi-token-prediction block inside the gguf. With
--mtp that head drafts --mtp-draft continuation tokens and the target verifies
all of them in one wider decode, confirming the longest prefix whose argmax
equals what the target itself would have produced. Nothing is approximated and
no weight is skipped, so the quality is the full model's — but it is NOT
byte-identical the way --overlap and --prefetch are, and must not be used in a
byte-identity gate: a verify pass evaluates 1+N positions in one batch, and a
batched matmul is not bit-identical to N single-token ones, so a near-tie can
flip. Off by default.
The prize is that a decode's dominant cost, moving the dense weights and the
routed expert slices, is paid once per group instead of once per token. The
counterweight is that the verify positions route independently, so a layer's
read set widens toward N*k wherever adjacent tokens disagree, and the draft
pass routes through the MTP block's own expert layer on top. Measured on the
desktop host (DRAM-bandwidth-bound, model streamed at ~1.4x RAM): +15.1% at
draft 3 with the host's best recipe (7.12 -> 8.19 tok/s), +29% without the
lossy drop knob, acceptance falling from 71% at draft 2 to 52% at draft 4, and
flash bytes per token rising 19.7 -> 33.7 MiB as the widening predicts. Draft 3
is the optimum here; 4 is worse than 2. On a flash-I/O-bound phone that balance
can invert, so the flag ships off pending the device A/B.
The orchestration is llama.cpp's own (common/speculative.h, public headers
only): no fork, no patch, no submodule bump. Self-speculation is one model with
two contexts over it — the target, created with n_rs_seq so a rejected tail is
rewound from a bounded snapshot rather than replayed, and a draft context
created with ctx_type = LLAMA_CONTEXT_TYPE_MTP. The engine builds the draft
context itself rather than through common_speculative_init_from_params because
the eval callback is per-context: the streamer only sees the MTP block's expert
layer if the draft context carries the same cb_eval.
The MTP block is streamed like any other layer. It sits at layer index n_layer,
contiguous with the trunk and using the same tensor naming, so the hook and the
expert source are sized n_layer + n_layer_nextn; left at n_layer its experts
stay silently mmap-resident. Two consequences that are easy to get wrong: the
capture warm-up has to run on the draft context too (the MTP graph is built
nowhere else), and prefill is fed through the driver so the draft context's KV
reaches the last prompt position.
The loop accepts BEFORE catching the draft context up, so the catch-up runs on
the accepted prefix instead of the whole verify batch. Acceptance depends only
on the target's logits, which are already in hand once the decode returns, and
the rejected tail was being computed only to be deleted a few statements later.
The resulting state is identical — the driver seeds from row
min(n_accepted, n_rows-1), the same row under either batch, and the surviving
KV is exactly the range the rollback used to carve out — while skipping
n_draft - n_accepted positions through the MTP block per group. Since that
block carries its own MoE FFN, on a streamed device those are expert reads that
no longer happen. It also removes the draft context's rollback entirely: it is
never given a tail to drop.
Requires an MTP-converted gguf (most quantisations strip the nextn tensors) and
greedy decoding; both are rejected at load with a message rather than silently
ignored, as is a n_ubatch narrower than the verify batch, which would split the
graph back into single-token passes and spend the draft for nothing.
Telemetry: an "mtp:" summary line, an mtp_batch per-token CSV column (a verify
decode's whole cost is charged to its group's first row, the rest carry zeros),
mtp_drafted / mtp_accepted / mtp_decodes in the CSV trailer and in BMOE_DONE,
and mtp / mtp_draft_max in the CSV preamble. The Android app exposes the flag
and the draft width, off by default.
Host gates pass. Validated on Qwen3.6-35B-A3B-MXFP4 with the streamed recipe:
draft 1 and draft 3 produce identical text, which is the invariant a broken
accept/rollback path would violate. Device A/B still owed.
* perf(mtp): shrink the draft context, make its cost measurable, record the device verdict
The first on-device A/B says MTP loses at every draft width, and the counters
say why. Same gguf with the flag on and off, shipping recipe (overlap, 3000 MiB
cache, pinned dense, drop 0.75), Qwen3.6-35B-A3B-Q4_K_M streamed:
off 5.82 / 6.14 tok/s 69.3 MiB/tok 69-109 majflt/tok
--mtp-draft 2 5.59 93.8 230
--mtp-draft 3 4.38 106.6 633
Speculation is working - 2.35-2.52 tokens per verify decode, 52-69% acceptance
- and still losing, because the prize does not exist in this regime.
stall_s/tok is 0.025-0.027 in every one of those runs, MTP on or off: 11-16% of
the token. This configuration is compute-bound, and what MTP amortises is weight
movement. The costs meanwhile are real and monotonic in the draft width: the read
set widens (+35%, +54% flash bytes per token), CPU per token rises (+28%, +67%),
and the draft context's memory tips the device into a fault storm.
Two things follow, and both are engine bugs rather than facts of nature.
The draft context's graph width drops from 256 to 32. Compute buffers are
reserved for the widest ubatch and the dominant term scales with
ubatch x vocabulary; on device that reservation measured 493 MiB - for a context
that evaluates ONE token per draft step and is handed at most 1 + draft_max
positions by the catch-up, with no logits asked for. Only prefill ever feeds it a
wide batch, and that is one layer, so splitting it costs very little. On this
engine memory is never free: it is the expert cache's, and the cache is what
decides whether the widened verify read set is a hit or a flash read.
And the cost of speculation is now measured instead of inferred. Drafting happens
between decodes, so it never entered wall_ms and tok/s never included it - a
speculated run could report a rate the user was not experiencing. New
mtp_draft_ms per-token column (a slice of loop_overhead_ms, not an addition),
mtp_draft_s/tok in the CSV trailer, mtp_draft_s_tok and loop_overhead_s_tok in
BMOE_DONE, and a second "mtp:" summary line printing the effective rate next to
the reported one.
Adds --mtp-p-min F, which stops drafting once the head's confidence in what it is
proposing falls below F. The draft loop already had this floor and the engine was
passing 0, so it always drafted the full width however unsure the head was - with
roughly half the drafts rejected at draft 3, that is the cheapest waste available
to cut. On a streamed device it pays twice: a draft not made is a pass through the
MTP block (which carries its own MoE FFN, so its own expert reads) that never
happens, AND one fewer independently routed position in the verify batch. Default
0, the setting the host numbers were measured at; the useful value is a property
of a device's balance between drafting cost and acceptance, so it is a knob to
measure rather than a constant to guess.
The Android app now reads the mtp_* keys it was already being sent: acceptance,
tokens per pass, and the effective rate. Before this the UI could not tell whether
speculation had run at all - only the session CSV could - which made the A/B this
commit reports impossible to run from the phone.
Neither mitigation changes the regime. The honest expectation is nearer
break-even, not a win, and the flag stays off by default.
Host gates pass. Note the noise floor: the two off runs did byte-identical work
and still differ by 5.6% in tok/s, and the runs were back-to-back without thermal
gating - the mechanism counters are the trustworthy part, not the exact deltas.
* perf(mtp): split the drafting flash cost from the widened verify batch
A speculated run streams more bytes per token for two unrelated reasons: the
MTP block carries its own MoE FFN, so every draft pass routes experts of its
own, and the verify batch widens the trunk's read set wherever adjacent
positions disagree. They need opposite fixes -- a narrower draft attacks the
first, only better agreement attacks the second -- and the route trace can
separate neither, since its framing brackets the target decode while the head
only ever runs in the draft context.
Measure the head's share directly by bracketing both drafting passes with the
expert source's byte counter, and report it as a third mtp: summary line.
Also record the branch-deletion rule in AGENTS.md: a branch list should only
show work in flight, and a rejected PR loses nothing.
* feat(engine): n-gram prompt-lookup draft source, and the measurement that closes it
The flash split added last commit said where MTP's cost actually is: at draft 3 on
the host, the head's own routing was 2.9% of the extra bytes a speculated run
streams and the widened verify batch was the other 97.1%. So a cheaper draft
producer is worth almost nothing, and the only property that could matter is one
the head does not have -- the ability to decline to draft at zero cost.
--ngram is that source. It takes the last few tokens, finds where that run occurred
before in the prompt or in what has been generated, and proposes whatever followed.
No head, no draft context, no decode, no expert read, and it works on any gguf
including the ones --mtp refuses for want of a nextn block. Below --ngram-min-match
it proposes nothing and the step falls through to a plain single-token decode.
Measured on the host, Qwen3.6-35B-A3B-MXFP4 streamed, 256 greedy tokens, cells
back-to-back with off run twice:
prose off 5.80 / 6.59 mtp3 7.32 eff ngram3 6.51 (cov 7.4%)
copy-heavy off 5.45 / 5.65 mtp3 6.43 eff ngram3 5.24 (cov 15%)
The zero-cost claim holds exactly -- mtp_draft_s/tok reads 0.0000 in every n-gram
cell, against 0.020-0.023 for the head plus the ~500 MiB of expert cache its draft
context takes. But the floor turns out to be per STEP, not per run: the 15% of steps
that did draft widened the read set to 67.2 MiB/token against 48-58 at baseline and,
at 44% acceptance, bought 1.20 tokens per decode. That is not enough to earn the
widening back, and a modest fraction of such steps sinks the run.
A --ngram-min-match sweep settles it rather than leaving it open. Raising the gate
3 -> 5 -> 8 lifts acceptance 44% -> 75% while coverage collapses 15% -> 3.4%, and
narrowing to --draft 1 reaches 82.6% -- the head's own figure on this prompt. Every
cell climbs toward baseline from BELOW and none crosses it; the best configuration
found lands on the floor. A knob whose optimum is its own disablement is not a
tuning problem. Acceptance, not drafting cost, is what pays for a widened batch, and
what a trained head buys is being right often enough to justify a batch that has
already been widened.
--ngram ships off. It is kept because it is the only speculation available on a
model with no head, because the per-step floor is real, and because the counters it
adds make the next speculation claim falsifiable.
Wiring. MtpConfig became SpecConfig with DraftSource {none, mtp, ngram}, and
--mtp-draft became --draft: the width belongs to the verify batch, not to whoever
filled it. --mtp and --ngram are rejected together rather than resolved by flag
order. In the session the gate split in two -- spec_on (wide batch, acceptance,
rollback: both sources) against mtp_on (draft context, common/speculative.h, the
catch-up: the head only) -- which is what lets the n-gram source reuse the whole
verify half while allocating nothing.
A step that drafts nothing now takes the plain path: llama_batch_get_one with a
logits row of -1, byte for byte the unspeculated decode. It used to build the wide
batch anyway. Required for --ngram, and it tightens --mtp-p-min's zero-draft steps
for free.
The matcher is pure policy over token ids with no llama.cpp at all -- not even
llama.h, since llama_token is int32_t -- so it sits on the clean side of the seam,
adds no dependency on the common layer, and is unit-tested with no model
(tests/ngram_test.cpp covers tie-breaks, clipping, self-match exclusion and the gate
boundary). Telemetry: spec= / spec_draft_max= / ngram_min_match= in the CSV
preamble, a new drafted_steps key in the trailer and BMOE_DONE, and an ngram: line
reporting coverage -- without which a delta cannot be divided by the fraction of the
run it applies to. The per-token and trailer counters keep their mtp_ names: they
always described the loop rather than a source, spec_* already means the temporal
prefetch in that trailer, and renaming would break every CSV already holding a
measurement. The Android setting became a three-way picker, migrating the old
boolean preference.
The device A/B agrees and adds a cost the host could not show. Thermally gated cells
(a 120 s settle, then a battery-temperature gate, so all six start between 35.3 and
36.4 C): prose 4.90 inside a 4.59-5.17 band, copy-heavy 3.14 against 4.43 -- a 29%
loss, worse than MTP's 18%. Major faults per token go 126 -> 1427 for a source that
allocates no draft context at all, and that is the rollback snapshots: n_rs_seq =
draft_max is asked for by ANY speculation, since rejecting a draft means rewinding the
KV, and on a hybrid attention/SSM model that snapshot is a real allocation scaling with
the context. The n-gram source escapes MTP's draft context but not the loop's own
memory, and on device that memory is the expert cache's.
The same run re-measured MTP with the thermal confound removed -- 3.64 effective
against 4.43, so the earlier device verdict was not an artefact of benching without a
cooldown gate -- and reproduced the flash split at 3.7% head against 96.3% widened
verify batch, matching the host's 2.9-3.0%.
Byte-identity gates pass; speculation stays out of them for the reason docs/mtp.md
gives.
The app's CSV configuration surface follows: the three new preamble keys get their own
glossary entries rather than falling through to the unexplained-key renderer, and the
draft source joins the short run label. A speculated run is not the same KIND of run --
under speculation a decode confirms a whole group, so its per-token rows are not even
accounted the same way -- and two compare legends differing by it must not read alike.
18 KiB
MTP decoding (--mtp)
--mtp lets the model draft its own continuation with the multi-token-prediction (MTP) head
trained into the gguf, then verifies every drafted token in a single wider decode. Off by default.
Unlike turbo top-k and expert dropping, this one is not a quality trade: a draft is accepted only when it equals the target model's own argmax at that position, so nothing is approximated and no weight is skipped. The quality is the full model's.
It is exact, but it is not bit-reproducible
Read that claim precisely, because it is weaker than the one the rest of this engine makes.
--overlap, --prefetch and the whole streaming path are byte-identical: they change how
weights reach the CPU and nothing else, and the gates prove the streamed
output equals the resident one exactly. --mtp is not in that class. Verification evaluates
1 + N positions in a single batch, and a batched matmul is not bit-identical to N separate
single-token matmuls — different blocking, different summation order, different last bits. On a
near-tie between two candidate tokens, that is enough to flip the argmax. The accepted sequence is
always exactly the argmax sequence of the batched forward pass; that pass is just not the same
arithmetic as the unbatched one.
Measured on Qwen3.6-35B-A3B-MXFP4, 128 greedy tokens, streamed with --overlap --cache-mb auto:
| comparison | result |
|---|---|
| plain vs plain (control) | identical, 608/608 chars |
plain vs plain at --ubatch 1, no MTP at all |
differs, 608/612 chars |
plain vs --mtp --draft 1 |
diverges at char 151 — one token |
plain vs --mtp --draft 3 |
diverges at char 151 — the same token |
--draft 1 vs --draft 3 |
identical |
The second row is the one that settles it, and it involves no speculation: --ubatch 1 only
forces the prompt to be prefilled one token at a time instead of in one wide pass. Same model, same
weights, same flags otherwise — and the greedy output already changes. Batch width moves the last
bits on this backend, full stop; speculation merely makes every decode a wide batch and inherits it.
The last row rules out the other candidate. Rolling back a rejected tail has to restore recurrent
state (Qwen3.5/3.6 is hybrid, which is why the target context is created with n_rs_seq), and a
broken rollback is a real failure mode — it is what SGLang diagnosed on Ascend NPU
(#25587). But a rollback bug scales with how
much gets rolled back: draft 1 rewinds at most one position, draft 3 up to three. Those two runs are
byte-identical to each other, so the rollback is not what is moving the output.
One token differed (Request → Query), the continuation stayed coherent, and every speculated run
agreed with every other regardless of draft width. A logic error in the accept/rollback path would
do neither of those things: it would vary with the draft width and compound over the generation.
This is expected, not a defect of this engine, and it is documented upstream:
- llama.cpp, in its own server README, on why prompt caching can change results: "Because (depending on the backend) the logits are not guaranteed to be bit-for-bit identical for different batch sizes (prompt processing vs. token generation) enabling this option can cause nondeterministic results." Speculation makes every decode a wide batch, so it lands in exactly that case.
- vLLM, whose speculative decoding is validated by a "Greedy Sampling Equality" test, still states that it is "theoretically lossless up to the precision limits of hardware numerics" and that "changes in batch size may cause variations in logprobs and output probabilities."
- The failure mode of a genuinely broken batched speculative implementation looks different:
"Batch Speculative Decoding Done Right" describes
desynchronised position ids, attention masks and KV state producing repetitive tokens or
gibberish and near-zero exact-match against standard decoding — not one flipped near-tie inside
an otherwise identical transcript. A duplicated token (
"renowned for for") is the classic accept/rollback off-by-one; a substituted one is not.
Whether it shows up at all is a property of the backend, not of the idea. SGLang reports MTP
output identical to non-speculative decoding on NVIDIA, where the kernels happen to be
batch-invariant, and divergent on Ascend NPU. On ggml's CPU path — GEMV for one token, blocked GEMM
plus mul_mat_id expert grouping for several — batch invariance does not hold, as the --ubatch 1
row above shows directly. Do not port the NVIDIA expectation here.
So: same model, same weights, no approximation — but do not expect a --mtp transcript to match an
unspeculated one character for character, and do not use it in a byte-identity gate.
Why a wider decode can be faster
A decode's cost is dominated by moving weights, not by arithmetic. One token reads the dense weights
plus its routed experts and does a handful of GEMVs against them; the hardware spends its time
waiting for bytes. Verifying N positions in one decode reads those same weights once and turns
the GEMVs into small GEMMs. If the draft is usually right, the engine confirms several tokens for
roughly the cost of one.
That is the same lever in both regimes this engine runs in, for different reasons:
- DRAM-bandwidth-bound (the desktop host): the dense weights are re-read from DRAM every token, and that traffic is the measured ceiling. Amortising it over a group attacks the bottleneck head-on.
- Flash-streamed (a >RAM model on a phone): what binds decode is latency-to-ready of the expert slices, not raw bandwidth (the sidecar refutation). A group pays that latency once.
Why it can lose
The N verify positions route independently. Each one picks its own top-k experts, so a layer's
read set widens from k toward N × k distinct experts wherever routing diverges across adjacent
tokens. In the streamed regime those extra experts are extra flash reads, and they can cost more
than the amortisation saves. The draft itself is not free either: the MTP block carries its own MoE
FFN, so every draft pass routes through one additional expert layer.
The engine therefore treats it as a measurement, not an improvement: default off, and the numbers below decide whether it ships enabled on a given device.
Requirements
The gguf must carry the nextn block — and on Qwen3.6 the ordinary quantisations already do.
Verified from the gguf headers: bartowski's Q4_0 and Q4_K_M and unsloth's MXFP4_MOE are all
qwen35moe with nextn_predict_layers = 1 and the same twenty blk.40.* tensors (full attention,
its own 256 routed experts, shared experts, plus nextn.eh_proj / enorm / hnorm /
shared_head_norm). The KV blocks differ only in provenance keys. A file named -MTP- buys a
quantisation, not a head — do not go looking for a special conversion.
There is no silent-fallback case to worry about either. llama_model_n_layer_nextn() reads the KV
key nextn_predict_layers, and llama.cpp loads the MTP tensors as required when it is greater than
zero, so a gguf that kept the key but dropped the tensors would not open at all. If a model opens
with --mtp, the head is genuinely there; if it has no head, the engine fails at load with a
message rather than quietly decoding one token at a time.
One consequence for measurement: an MTP A/B must be the same file with the flag on and off.
Comparing an -MTP--named file against a differently-quantised baseline confounds the feature with
the quantisation change.
On device it loses, and the counters say why
Measured on the test phone (12 GB), Qwen3.6-35B-A3B-Q4_K_M streamed with the shipping recipe
(--overlap, 3000 MiB cache, pinned dense, --drop-cold-experts 0.75), same file with the flag on
and off:
| tok/s | MiB/token | majflt/token | cpu-s/token | cache hit | acceptance | |
|---|---|---|---|---|---|---|
| off | 5.82 / 6.14 | 69.3 | 69–109 | 0.43–0.45 | 73.5% | — |
--draft 2 |
5.59 | 93.8 | 230 | 0.57 | 65.4% | 69% |
--draft 3 |
4.38 | 106.6 | 633 | 0.74 | 66.7% | 52% |
Speculation is working — 2.35–2.52 tokens per verify decode — and still losing, because the prize
does not exist in this regime. stall_s/tok is 0.025–0.027 in every one of those runs, MTP on or
off: 11–16% of the token. This configuration is compute-bound, and what MTP amortises is weight
movement. Meanwhile all three costs are real and monotonic in the draft width: the read set widens
(+35%, +54% flash bytes per token), the draft context's memory tips the device into a fault storm
(major faults per token 3–9×), and the drafting itself adds CPU (+28%, +67%).
Two mitigations came out of that measurement and are in the engine now. The draft context's graph
width was cut from 256 to 32, because compute buffers are reserved for the widest ubatch and the
output buffer scales with ubatch × vocabulary — 493 MiB reserved on device for a context that
evaluates one token at a time. And --mtp-p-min makes the draft width adaptive. Neither changes the
regime: the honest expectation is nearer break-even, not a win.
The two off runs did byte-identical work and still differ by 5.6% in tok/s, so treat that as the
noise floor: draft 2 is at its edge, draft 3 is well outside it. The runs were also back-to-back
without thermal gating, so the mechanism counters are the trustworthy part, not the exact deltas.
Stopping early when the head is unsure (--mtp-p-min)
The draft loop already knows how confident the head is in each token it proposes. --mtp-p-min F
stops drafting as soon as that confidence falls below F; at the default 0 it never stops and
always drafts the full width.
On a streamed device this pays twice, which is why it is the cheapest lever available: a draft not made is a pass through the MTP block — with its own MoE FFN, and therefore its own expert reads — that never happens, and one fewer position in the verify batch, so one fewer independent routing widening the layer's read set. A draft that is made and then rejected costs both and buys nothing, and at draft 3 above roughly half of them were rejected.
It is a knob rather than a default because the useful value is a property of the device's balance between drafting cost and acceptance, and 0 is what the host numbers were measured at.
Speculation requires greedy decoding: validate() rejects --mtp together with --temp > 0,
because acceptance under a sampling chain would depend on which draws happened to agree and the run
would no longer be the single-token run it claims to be.
How it is wired
The whole draft/verify orchestration is llama.cpp's, reached through common/speculative.h, which
depends only on public headers. No llama.cpp patch, no fork, no submodule pin of our own — the
same rule the rest of the engine follows (seam.md).
Self-speculation means one model and two contexts over it:
- the target context, created with
n_rs_seq = --draftso a rejected tail is rewound from a bounded snapshot instead of replayed; - the draft context, created with
ctx_type = LLAMA_CONTEXT_TYPE_MTPso llama.cpp builds the nextn graph, and carrying the same eval callback as the target — the router hook is per-context, not per-model.
That callback is what makes the MTP block visible to the streamer. The head lives at layer index
n_layer — contiguous with the trunk, same ffn_{gate,up,down}_exps tensor naming, same expert
count — so with --mtp on, the hook and the expert source are sized n_layer + n_layer_nextn
and the head's experts are streamed, cached and dropped like any other layer's. Two things follow
from that and are easy to get wrong:
- The capture warm-up has to run on the draft context too. The MTP graph is only ever built by a decode there, so a warm-up on the target alone never harvests the head's expert tensors, and they would stay mmap-resident — the fault storm streaming exists to avoid.
- Prefill goes through the driver as well. The draft context's KV must reach the last prompt position or the first draft is conditioned on a state that never saw the prompt.
Per step the loop is: draft --draft tokens from the head, decode
[confirmed token, drafts…] asking for logits at every position, accept the longest prefix whose
argmax matches, let the draft context catch up on that prefix, roll the target's KV back over the
rejected tail, and read the next token for free out of the logits at the first unverified position.
The catch-up runs on the accepted prefix, not the whole batch
That ordering is deliberate and it is not the obvious one. The catch-up re-runs the MTP block over the batch so the draft context ends up holding rows conditioned on the target's own hidden states rather than the head's guesses, and the next draft is seeded from the row at the accepted position. The natural place to put it is right after the decode, on the whole verify batch — which is what upstream's own loop does. But that computes the rejected tail too, and the rejected tail is deleted a few statements later: those drafts were wrong, and nothing downstream ever reads their rows.
Acceptance depends only on the target's logits, which are already in hand once the decode returns, so
the tail never has to be submitted at all. Accepting first and handing the catch-up only
[confirmed token, accepted drafts…] leaves identical state — the driver seeds from row
min(n_accepted, n_rows-1), which is the same row under either batch, and the surviving KV is
exactly the range the rollback used to carve out — while skipping n_draft − n_accepted positions
through the MTP block. It also removes the draft context's rollback entirely: it is never given a
tail to drop.
The saving is small on a host, where that block is one layer of arithmetic. It is not small on a streamed device: the MTP block carries its own MoE FFN, so every skipped position is a routing that no longer happens and a set of expert slices that no longer has to be read. At draft 3 with acceptance around 60%, roughly one position per group stops being computed and streamed.
Reading the numbers
mtp: 41/63 drafts accepted (65.1%), 2.31 tokens per verify decode (57 decodes for 132 tokens)
- Acceptance is a property of the model's trained head, not of this engine. It bounds everything else: at 0% the feature is pure overhead.
- Tokens per verify decode is what was actually bought — the factor the decode count fell by. It
is always below
1 + --draft. - The effective rate, on the second line, is the one to believe.
tok/sis computed from decode time alone and drafting happens between decodes, so the headline rate leaves out everything speculation adds.mtp_draft_s/tokin the CSV trailer andmtp_draft_msper row measure it directly, instead of leaving it to be inferred by differencing against an unspeculated run. - The flash split, on the third line, says how much of the run's streamed MiB the head's own routing pulled and how much was the widened verify batch. Speculation grows bytes per token for two unrelated reasons and they need opposite fixes: a narrower draft attacks the head's share, while the verify union only shrinks if adjacent positions agree more. The route trace cannot separate them — it brackets the target decode, and the head only ever runs in the draft context.
token_demand_MiB keeps its mechanical meaning but changes scope: it is the distinct expert bytes
one decode routes, and under speculation a decode is a group, so the figure is the group's union
— which is exactly the working set the cache has to stage. It is the same widening the "why it can
lose" section describes, measured. Read it against cache_budget_MiB as usual, but never compare it
across a speculated and an unspeculated run.
In a CSV, mtp_batch says how many tokens each decode confirmed. The group's entire cost is
charged to its first row and the rest carry zeros, so a group's per-token cost is that row's
wall_ms / mtp_batch. Never average rows from an mtp=1 file together with an unspeculated one.
See telemetry.md.
Measured
Desktop host (8 cores, 14.8 GB RAM, NVMe), Qwen3.6-35B-A3B-MXFP4_MOE (20.7 GiB, ~1.4× RAM), streamed, greedy, 256 tokens, one prompt, single session. This host is DRAM-bandwidth-bound, which is the regime the amortisation targets.
With the host's measured best recipe (--overlap --cache-mb auto --drop-cold-experts 0.75):
| draft width | tok/s | vs off | acceptance | tokens / verify decode | flash MiB/token |
|---|---|---|---|---|---|
| off | 7.12 | — | — | 1.00 | 19.7 |
--draft 2 |
7.66 | +7.6% | 71.4% | 2.42 | 30.3 |
--draft 3 |
8.19 | +15.1% | 60.1% | 2.78 | 33.7 |
--draft 4 |
7.64 | +7.3% | 52.4% | 3.08 | 31.1 |
Without the lossy drop knob (--overlap --cache-mb auto, 128 tokens), the gain is larger — 4.28 →
5.52 tok/s, +29% at draft 3 — because dropping had already removed some of the compute the
amortisation is buying back.
Two things to read off the table. First, acceptance falls as the draft widens while tokens per decode still rises, and the two effects cross: 3 is the optimum here, 4 is worse than 2. Second, the widening is real and visible — flash bytes per token rise from 19.7 to 33.7 MiB because the verify positions route independently. It still wins because compute per token falls further (0.115 → 0.088 s) than the extra I/O costs on this host. On a flash-I/O-bound device that balance can invert, which is why the flag ships off and the phone A/B is a separate measurement.