The NPU prefill's expert arena read every expert of every layer ahead of its routing. It now reads, ahead of a layer's routing, the experts the previous graph routed there, and at the routing node whatever the routing adds. The matmul reads only routed experts, so the output is bit for bit the same. A layer routing more than --prefill-routed-full (0.85) of its experts gets the next one read whole; --no-prefill-routed restores whole layers everywhere. Phone, Hexagon v81 NPU, top-4, same session, every answer identical: Qwen3.6-35B-A3B Q4_0 7.68 -> 4.16 s, Q4_K_M 9.95 -> 5.37 s, Gemma 4 26B-A4B Q4_K_M 6.69 -> 3.70 s, Nemotron 3.5 30B-A3B Q4_0 7.42 -> 6.81 s. Also: --decide-probe (experimental per-decision expert usage and layer-exit answers), BMOE_DECIDE prefill_dev_* counters, gates G17f/G17g, app 0.28.0.
14 KiB
Prefill on the NPU (--prefill-device)
Prefill and decode want opposite hardware. A prefill graph is hundreds or thousands of tokens wide, so a matrix engine runs it many times faster than the CPU cores. A decode graph is one token wide, and on a phone's unified memory the cost of crossing to an accelerator and back eats what it saves: measured on a Hexagon NPU, the decode token got slower (see "Why not decode" below).
--prefill-device D gives each phase its hardware. Wide prefill graphs run on ggml device D (the
Hexagon NPU is HTP0); decode stays on the CPU, exactly as without the flag. With --moe-stream it
works on a model larger than RAM: the NPU never holds the model, only two layers of it at a time.
Measured
Qwen3.6-35B-A3B, Q4_0 gguf (20.8 GB) on a 12 GB phone with a Hexagon v81 NPU and UFS 4 storage,
streamed (--moe-stream --overlap --dense-weights ahwb --cache-mb 1500 -t 4), ubatch 2048 on the
NPU, 512 on the CPU. The Q4_K_M row is the same model in the quantisation the app's catalog ships,
on the 2026-09-26 llama.cpp base, with 8 loaders:
| prompt | CPU prefill | NPU prefill | |
|---|---|---|---|
| 121 tokens | 9.95 s (12.2 tok/s) | 9.5 s (12.7 tok/s) | parity: flash bound |
| 1418 tokens, prose | 63.8 s (22.2 tok/s) | 8.2 s (172 tok/s) | 7.8x |
| 1921 tokens | 106.5 s (18.0 tok/s) | 11.2 s (171 tok/s) | 9.5x |
| 1418 tokens, prose, Q4_K_M | 80.6 s (17.6 tok/s) | 10.0 s (141 tok/s) | 8.1x |
Decode after the 1418-token prompt: 3.41 tok/s on the CPU-only run, 3.25 tok/s with the NPU prefill (both decode on the CPU; the difference is the memory the NPU's slots hold).
Gemma 4 26B-A4B-it Q4_K_M (17 GB), same phone and settings but --cache-mb 2000, a 238-token
prompt, a 2048-token context:
| prefill | decode | |
|---|---|---|
| CPU | 16.2 s (14.7 tok/s) | 3.60 tok/s (8192-token context) |
| NPU | 5.85 s (40.7 tok/s) | 3.30 tok/s |
Memory is the limit on a model with a large KV cache. The device path holds the two expert slots
(1.1 GB on this model), the device's compute buffers (0.8 GB at ubatch 2048) and the model state in
the device's host buffer. On a 12 GB phone Gemma 4 at an 8192-token context and a 2000 MiB expert
cache already leaves about 330 MB free on the CPU alone; with the device path on top it thrashes. At
2048 it fits. A smaller --cache-mb or context makes the room.
Short prompts are flash bound. The table above was measured reading the whole expert set from flash once per graph (17.4 GB here), whatever the prompt length. Past roughly a thousand tokens that read hides behind the NPU's compute; under a few hundred it is the whole cost. The arena now reads only the experts a graph routes to (below), which on short prompts is about half of them.
How it works
Moving the weights per graph, not per model
llama.cpp's scheduler runs each op on the backend that holds its weight, and it decides that again
for every graph it builds. Where a weight lives is three public fields of its ggml_tensor:
buffer, data, extra. So before a wide prefill graph the engine rebinds the layer weights onto
the NPU's buffers, and afterwards back onto the CPU's. Nothing in llama.cpp is patched. The device
joins the scheduler at load as the model's only non-CPU device, with no layer assigned to it, and
op_offload is off, so the device runs nothing it was not handed.
The one hazard is graph reuse: llama.cpp skips re-scheduling a graph shaped like the previous one.
The rule that makes the rebind safe is about widths. The prompt goes in one ubatch per decode;
pieces at least --prefill-min-tokens wide (default 32) run on the device and a shorter tail runs
on the CPU, so a device graph and a CPU graph never share a shape. Speculative decoding widens CPU
graphs and runs a second context over the same weights, so the two are refused together for now.
The arena: two layers of the model at a time
A model larger than RAM cannot give the NPU a copy of itself. The NPU gets two layer-sized slots instead, and every layer's weights are bound to one of them, alternating. While the NPU computes layer k out of one slot, loader threads fill the other with layer k+1:
- experts are read from the gguf with
O_DIRECT, one expert at a time, and handed to the backend withggml_backend_tensor_seton a per-expert view. That call is where the Hexagon backend repacks them into its matrix-engine tiles, so the repack runs in parallel across the loaders; - the other layer weights (attention, norms, shared experts) are already resident on the host, so filling them is a copy, not a read.
Pacing uses points the graph already offers. The experts wait at the layer's routing node, which the streamer knows how to isolate; there the arena reads the routed ids and loads what is missing (below). The other weights are needed before the routing, so they wait at the last node of the previous layer, which the capture pass learns per layer because no node name is common to every architecture. A graph that skips a pacing point fails the decode rather than compute on a slot that never filled. Measured: the slots cost about 900 MB for the model above, instead of the 21 GB the model is.
Reading only the routed experts (default; --no-prefill-routed reads whole layers)
Loading every expert assumes a wide graph routes to nearly all of them. A phone agent's prompt does
not: measured with --decide-probe on the same model with top-4 routing, a 130 to 480-token prompt
routes to about 128 of each layer's 256 experts, and a 55-token one to about 64.
A layer is read in two parts. Ahead of its routing, while the NPU computes
the layer before, the loaders read the experts the previous graph routed at that layer: consecutive
decisions of an agent route much alike. At the layer's routing node the arena reads the routed ids
(ggml_backend_tensor_get, wherever the tensor lives) and queues what the prediction missed ahead of
everything else, then waits for it. The expert matmul reads only routed experts, so what the slot
holds for the others never reaches the result: the output is the same bit for bit. A layer that
routes to more than --prefill-routed-full of its experts (0.85 by default) gets the next layer
read whole, as --no-prefill-routed does everywhere, so a long prompt, which routes to everything,
does not pay for a prediction it cannot use. The first graph of a session has no prediction and reads whole layers.
Measured on the phone above, Qwen3.6-35B-A3B Q4_0 streamed with the settings of an on-device
Android UI agent (--overlap --io-threads 4 --dense-weights ahwb -t 6 -c 2048 --n-expert-used 4, ubatch 2048), 35
decisions: 31 over Android screens (129 to 476 tokens) and 4 short questions. Same session, same
build, whole layers and routed:
| whole layers | routed | |
|---|---|---|
| prefill, median | 7.68 s | 4.16 s |
| arena reads, median | 17.0 GiB | 9.2 GiB |
| graph waiting on the arena, median | 6.4 s | 2.9 s |
choice_logp identical |
35 of 35 |
The prediction missed 11% of the routed experts. It misses more when the prompt changes domain (the short questions after the screens: about 40%), and those prompts still ran in 2.6 to 2.9 s.
Other models on the same phone, same settings, 18 of the screen decisions each, every answer identical between the two:
| model | experts routed per layer | whole layers | routed | |
|---|---|---|---|---|
| Qwen3.6-35B-A3B Q4_K_M | ~50% of 256 | 9.95 s | 5.37 s | 1.85x |
| Gemma 4 26B-A4B Q4_K_M | ~58% of 128 | 6.69 s | 3.70 s | 1.81x |
| Nemotron 3.5 30B-A3B Q4_0 | ~79% of 128 | 7.42 s | 6.81 s | 1.09x |
The gain is what each layer leaves unrouted: Nemotron's prompts route to most of its experts, many
of its layers cross the 0.85 fallback, and it barely gains. The figures are for top-4 routing; with
a model's own top-8 a prompt routes to more experts and the gain is smaller (not measured).
Per-decision numbers: bench-data/2026-09-29-prefill-routed/.
What else had to move
- The KV cache and recurrent state move once, at load, into the NPU's host buffer type: memory the
CPU reads directly and the NPU addresses too, so decode and prefill share one cache. Left in plain
CPU memory, every attention of a device graph would run on the CPU. The Hexagon backend exposes that
buffer type only with
GGML_HEXAGON_HOSTBUF=1, which the app sets. The buffers llama.cpp first allocated the state in stay allocated, since only llama.cpp can free them, but nothing reads them after the move, so their pages are handed back to the kernel (and again after eachllama_memory_clear, which rewrites them). Kept resident they would double the KV cache: 1760 MiB on Gemma 4 26B-A4B at an 8192-token context, where Qwen3.6, mostly linear attention, has 143 MiB. - Weights that are not a matmul's matrix. A backend may hold a
WEIGHTSbuffer in a form only its matmul kernels address: with DMA64 on (the default above Hexagon v79), the Hexagon backend maps such a buffer for DMA only, and most of its other kernels refuse it at run time, which aborts the graph. Gemma 4 met it first: it broadcasts a per-expert scale withREPEAT. The capture pass records every layer weight some op reads other than as the matrix of aMUL_MAT/MUL_MAT_ID(through views too), and those go to a second pair of slots in plain device memory, in their own type. The matrices keep theWEIGHTSslots. No op moves: the scheduler still runs each on the device, now on memory every kernel can read. - Weight types the device refuses. A "Q4_0" gguf is a mix: the one above keeps its shared experts in Q5_0 and four attention projections in Q6_K, which the Hexagon matmul does not take. Left alone they ran on the CPU inside every device graph (65 matmuls, 111 splits, 17.6 s instead of 11.2 s). Such a weight is carried to the device in the nearest type it takes (Q8_0 for a quantised grid, F16 or F32 for a float one), converted once at load. The file and the CPU decode are untouched. Which type is asked of the device with a probe matmul, not read off a list.
- Compute buffers. llama.cpp reserves compute memory at load for the widest graph with every weight on the CPU, a graph this session never runs there, and it reserves a logit row per token of the ubatch, each the width of the vocabulary: 2.2 GB at ubatch 2048, which pushed decode into thrashing (0.5 tok/s). The reservation is redone with the weights on the device, and the logit rows are capped at 128, which brought the CPU's buffer to 154 MB.
Requirements
- A model whose expert tensors the NPU's
MUL_MAT_IDtakes: Q4_0, Q4_1, Q8_0, IQ4_NL, MXFP4, and since the 2026-09-26 llama.cpp base the K-quants Q4_K, Q5_K and Q6_K, which is what a Q4_K_M is made of. A dense weight in any other type is converted for the device (see above); Q3_K and below are not taken for experts. gpt-oss is natively MXFP4. - The Hexagon backend in the build:
scripts/build-hexagon-android.shbuilds the CLI, the backend and one DSP-side skel per NPU generation (v73 to v81) inside upstream's Snapdragon toolchain container, andscripts/stage-hexagon-jnilibs.ps1stages them into the app. The release APK is built the same way by CI, so it carries all of them. --prefill-loaders N(default 8) sets the threads that fill the slots, apart from--io-threads, which stays the decode's read lanes. Each loader reads and repacks, and a K-quant repack is CPU-heavy: on a Q4_K_M, 4 loaders left the NPU waiting 10.3 s of a 15.1 s prefill, 8 left it 5.0 of 10.0.- On device:
ADSP_LIBRARY_PATHpointing at the directory with thelibggml-htp-v*.soskels (fastrpc resolves the one for the phone's NPU through it), andGGML_HEXAGON_HOSTBUF=1.GGML_HEXAGON_OPPOLL=1makes the host poll for the DSP instead of waiting on an interrupt, which halves the cost of each crossing.
When the device is not there the run does not fail: a name the registry does not know (no backend in
the build, or a phone without the fastrpc driver, where Hexagon registers nothing) and a device that
does not open (a Snapdragon older than v73, which registers and then refuses a session) both leave the
whole run on the CPU, with a bmoe: line on stderr saying which, and prefill_dev_tokens stays 0.
The device is opened once before the load to find out, and the context reuses that session.
Without --prefill-device a Hexagon build keeps the NPU out of the run altogether. llama.cpp, given
no devices, lists every GPU-type device and opens a backend on each, and Hexagon reports itself as
one; the engine drops any such device that can reach neither a host buffer type nor host pointers,
since with no layer assigned it could do nothing but open a DSP session. See docs/seam.md.
Correctness
Gates G16 and G17 run the whole path against a loopback rpc-server fronting the CPU: the same
kernels, so placement is the only difference, and output, perplexity and the bytes the arena reads
must all match an all-CPU run bit for bit, with and without a cache, with a slowed loader, and across
several generates in one session. Removing either wait in the arena fails them (checked). G17f runs
the routed arena the same way (fallback disabled, so every layer after the first graph is read from
a prediction plus the routing node) and requires fewer bytes read; G17g skips the routing-node reads
and requires the perplexity to change, so G17f cannot pass by luck. The RPC
backend is only that test fixture: it is built with the tests alone, on 127.0.0.1, and neither the
CLI nor the app accepts an RPC endpoint.
On the NPU itself the matrix engine computes in fp16, so the output is not bit-identical to the CPU's.
Price it with --ppl on the same Q4_0 model with and without the flag before relying on it.
Telemetry
BMOE_DONE and the CSV trailer carry prefill_dev_tokens, prefill_dev_nodes (nodes the device
actually computed), prefill_dev_read_mib and prefill_dev_stall_s; BMOE_DECIDE carries the same
device counters plus, in routed mode, prefill_dev_routed and prefill_dev_demand (experts routed,
and those read at the routing node). See telemetry.md.
Why not decode
Measured before this feature: the NPU runs the isolated q4_0 matmul 2.4x faster than the CPU at batch 1 and 20x at 512, and still made the decode token slower, 11% over adb and 31% in the app. With experts streamed on the CPU, a token crosses to the device and back about twice per layer, 91 times on a 40-layer model, at 0.3 to 0.5 ms each. Putting whole layers on the NPU removes the crossings, but on unified memory those layers take RAM from the expert cache one for one, and it lands at parity. A prefill graph crosses the same boundaries once per thousands of tokens instead.