koboldcpp/ggml
Eurekatic 4aa6ffba25
sycl: reduce redundant work in Q4_K multi-column MMVQ (#27062)
* sycl: Q4_K Weight unpack optimization and reuse between destination Columns

* sycl: Q4_K small N (N=2..4) + two output rows by subgroup reuse of activation between two rows.

* sycl: gate Q4_K two-row reuse for small N=2

* sycl: Fix on magic number now uses Q4_K_MMVQ_ROW_PAIR_MIN_NROWS=6272 for it, added tests for coverage around Q4_K_MMVQ_ROW_PAIR_MIN_NROWS with perf support to test Q4_K MUL_MAT, applied the same  reuse pattern to the activation as the weights.

Assisted-by: GPT-5.6 Sol

---------

Co-authored-by: RaulAbejonDelgado <raul.abejon.delgado@gmail.com>
2026-09-03 14:59:06 +08:00
..
cmake kleidiai: Rework KleidiAI Build System/Integration (#26077) 2026-08-25 14:07:29 -07:00
include CUDA + ggml: add sparse-fa for DSV4/GLM (#27970) 2026-09-02 17:27:37 +03:00
src sycl: reduce redundant work in Q4_K multi-column MMVQ (#27062) 2026-09-03 14:59:06 +08:00
.gitignore
CMakeLists.txt metal : add metallib build support for xcframework (#28163) 2026-09-02 07:45:56 +08:00