mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
O_DIRECT wants the file offset, the length and the buffer address aligned. Expert slices are a whole number of pages long, but a gguf's tensor data does not start on a page boundary: general.alignment is 32 by default, so on the measured model family every expert offset sits a constant 1152 bytes past one. The reader therefore pulled the enclosing aligned window into a per-lane bounce buffer and memcpy'd the payload back by that remainder — over 200 MiB a token of pure shift correction, on a device whose decode is already memory-bound. That shift is only needed because the destination did not share the file's remainder. It can: this engine reserves the per-layer buffers itself and rebinds tensor->data onto them. With --odirect-zero-copy each buffer is placed at an address carrying its own tensor's remainder; every expert inherits it because the per-expert stride is a multiple of the page size, so the aligned window maps onto the buffer in place and the read needs no copy anywhere. The window overhangs its neighbours' slices. That is safe by construction, not by luck: under this placement buffer and file differ by a constant offset, so the overhung bytes receive their own correct file contents, and they land in pages the entry's commit already covers exactly — eviction never releases a page shared with a neighbour, which was already true before this change. Measured on device (Qwen3.6-35B-A3B Q4_0, cache 2000 MiB, overlap, 4 lanes, interleaved cells so thermal drift hits both): +20% tok/s, 3.65 -> 4.40 median, compute residual 0.180 -> 0.143 s/token, identical bytes read (230.77 MiB/token in every cell) and byte-identical generated text with the flag on and off. The phone was warm and its baseline drifted 3.90 -> 3.40 across the run, so the ratio is the claim, not the absolutes; a cool-device confirmation is owed, which is why the flag ships off. This is not specific to any decode feature — every expert read goes through this path, and so does the dense loader's. The host gates cannot prove the path: a tiny test model's per-expert stride is not a multiple of the page size, so the placement declines and G8/G4e show only that the flag is harmless. Rather than let that read as a pass, the engine reports which happened (odirect-zero-copy ON|INERT — n/m layer buffers placed). The device run above is the real proof. Gates 7/7. |
||
|---|---|---|
| .. | ||
| CMakeLists.txt | ||
| main.cpp | ||