BigMoeOnEdge/cli
Helldez 447a0f82f3 perf(io): read expert slices straight into the cache, with no bounce copy
O_DIRECT wants the file offset, the length and the buffer address aligned.
Expert slices are a whole number of pages long, but a gguf's tensor data does
not start on a page boundary: general.alignment is 32 by default, so on the
measured model family every expert offset sits a constant 1152 bytes past one.
The reader therefore pulled the enclosing aligned window into a per-lane bounce
buffer and memcpy'd the payload back by that remainder — over 200 MiB a token
of pure shift correction, on a device whose decode is already memory-bound.

That shift is only needed because the destination did not share the file's
remainder. It can: this engine reserves the per-layer buffers itself and rebinds
tensor->data onto them. With --odirect-zero-copy each buffer is placed at an
address carrying its own tensor's remainder; every expert inherits it because
the per-expert stride is a multiple of the page size, so the aligned window maps
onto the buffer in place and the read needs no copy anywhere.

The window overhangs its neighbours' slices. That is safe by construction, not
by luck: under this placement buffer and file differ by a constant offset, so
the overhung bytes receive their own correct file contents, and they land in
pages the entry's commit already covers exactly — eviction never releases a page
shared with a neighbour, which was already true before this change.

Measured on device (Qwen3.6-35B-A3B Q4_0, cache 2000 MiB, overlap, 4 lanes,
interleaved cells so thermal drift hits both): +20% tok/s, 3.65 -> 4.40 median,
compute residual 0.180 -> 0.143 s/token, identical bytes read (230.77 MiB/token
in every cell) and byte-identical generated text with the flag on and off. The
phone was warm and its baseline drifted 3.90 -> 3.40 across the run, so the
ratio is the claim, not the absolutes; a cool-device confirmation is owed, which
is why the flag ships off.

This is not specific to any decode feature — every expert read goes through this
path, and so does the dense loader's.

The host gates cannot prove the path: a tiny test model's per-expert stride is
not a multiple of the page size, so the placement declines and G8/G4e show only
that the flag is harmless. Rather than let that read as a pass, the engine
reports which happened (odirect-zero-copy ON|INERT — n/m layer buffers placed).
The device run above is the real proof. Gates 7/7.
2026-08-01 23:46:43 +02:00
..
CMakeLists.txt chore: sweep the retired governor's residue from core and CLI 2026-07-17 10:04:27 +02:00
main.cpp perf(io): read expert slices straight into the cache, with no bounce copy 2026-08-01 23:46:43 +02:00