koboldcpp/ggml/src
Piotr Wilkin (ilintar) f4e276a206
ggml-cuda : convert contiguous tensors four elements at a time (#29155)
convert_unary handles the contiguous case through the general strided kernel,
one element per thread: each lane reads 4 bytes and writes 2. Converting the
activations for a bf16 matrix multiplication that way moves 126 MB in 1021 us
on gfx1151, about 65% of what the memory system can do.

Give the contiguous path its own kernel that takes four elements per thread
through a vector type, so a warp loads 512 bytes at a time instead of 128. It
is used only when the element count is a multiple of four and both pointers
carry the alignment the vector type needs, and falls back to the strided
kernel otherwise.

Model level, Qwen3.8-Next-Flash IQ3_XXS on gfx1151, llama-bench -ub 2048 -r 6,
mean of the last 3 reps, ABBA counterbalanced:

    pp2048   688.0 680.0  ->  694.3 691.1   +1.26%
    tg128     24.8  24.8  ->   24.8  24.8   +0.14%

Every conversion in a prefill takes the new kernel (kernel trace: 1146
convert_unary_cont_vec4, no convert_unary). Output is bit identical; MUL_MAT,
MUL_MAT_ID, CPY, CONT, GET_ROWS and SET_ROWS pass.

Assisted-by: Claude Opus 5
2026-09-21 18:00:51 +02:00
..
ggml-blas llama: add default load-mode auto, which avoids mmap on iGPUs (#26081) 2026-08-11 09:20:46 +03:00
ggml-cann ggml: add SWIGLU_CLAMP (#27930) 2026-08-30 23:00:02 +08:00
ggml-cpu ggml-cpu: ARM Repack kernels for Q1_0 (#23492) 2026-09-21 11:04:51 +03:00
ggml-cuda ggml-cuda : convert contiguous tensors four elements at a time (#29155) 2026-09-21 18:00:51 +02:00
ggml-et ggml: add SWIGLU_CLAMP (#27930) 2026-08-30 23:00:02 +08:00
ggml-hexagon hexagon: overhaul of buffer and DMA handling to support 64bit mappings + improvements (#29197) 2026-09-21 11:00:28 +03:00
ggml-hip CUDA: replace GGML_FA_ALL_QUANTS with GGML_FA_QUANTS, more control over what is compiled (#28079) 2026-09-09 12:50:08 +02:00
ggml-metal ggml-metal : simplify fusion pattern op list declaration (#29206) 2026-09-21 12:37:24 +03:00
ggml-musa CUDA: replace GGML_FA_ALL_QUANTS with GGML_FA_QUANTS, more control over what is compiled (#28079) 2026-09-09 12:50:08 +02:00
ggml-opencl opencl: add support for bin kernel flash_attn_f32_f16_bin (#29046) 2026-09-18 16:32:31 -07:00
ggml-openvino openvino : Update OpenVINO to 2026.4;fix clangd,MSVC warnings; (#29009) 2026-09-17 12:46:14 +02:00
ggml-rpc rpc : invalidate cached compute graph when a referenced buffer is freed (#24292) 2026-09-16 14:03:11 +03:00
ggml-sycl sycl : support gated DSV4_HC_PRE and optional HC_POST comb matrix (#29132) 2026-09-21 13:59:38 +03:00
ggml-virtgpu ggml: allow passing alloc dependencies in graph_optimize (#27301) 2026-08-30 11:34:20 +08:00
ggml-vulkan vulkan: add IQ3_S MMQ matmul kernels (#28822) 2026-09-18 11:46:02 +03:00
ggml-webgpu webgpu : add fused gdn + cpy (#28976) 2026-09-21 10:39:30 +03:00
ggml-zdnn llama: add default load-mode auto, which avoids mmap on iGPUs (#26081) 2026-08-11 09:20:46 +03:00
ggml-zendnn CI: Use LLVM's OpenMP over MSVC_DEBUG_non_redist on Windows (#26678) 2026-08-20 15:42:26 +02:00
CMakeLists.txt ggml : replace compile definitions with version.h.in (#28364) 2026-09-04 10:28:23 +02:00
ggml-alloc.c ggml : fix ggml_clamp (#27644) 2026-08-24 10:43:04 +03:00
ggml-backend-dl.cpp hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150) 2026-01-29 12:33:21 -08:00
ggml-backend-dl.h hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150) 2026-01-29 12:33:21 -08:00
ggml-backend-impl.h sync : ggml (#28379) 2026-09-04 14:39:19 +03:00
ggml-backend-meta.cpp ggml-meta: propagate buffer usage and call init on the new tensors (#27586) 2026-08-26 08:27:51 +03:00
ggml-backend-reg.cpp ggml : don't crash when backend search path can't be read (#28271) 2026-09-04 10:24:06 +03:00
ggml-backend.cpp ggml : handle graph buffer reservation failure (#26070) 2026-09-18 12:31:19 +03:00
ggml-common.h AVX2: Speed up large batch size prompt processing of IQ models (#27402) 2026-08-31 14:33:50 -04:00
ggml-feats.h ggml : fix arm builds, unused var (#26991) 2026-08-13 07:57:24 +03:00
ggml-impl.h ggml : update ggml_prec specification (#26675) 2026-09-08 09:06:24 +03:00
ggml-opt.cpp fix: free ctx_copy in ggml_opt_free to plug per-training-session leak (#21592) 2026-04-08 17:40:15 +02:00
ggml-quants.c Add Q2_0 quantization: type definition and CPU backend (#24448) 2026-07-07 12:05:47 -07:00
ggml-quants.h Add Q2_0 quantization: type definition and CPU backend (#24448) 2026-07-07 12:05:47 -07:00
ggml-threading.cpp
ggml-threading.h
ggml-version.h.in ggml : replace compile definitions with version.h.in (#28364) 2026-09-04 10:28:23 +02:00
ggml.c ggml : fix dimension and stride truncation in ggml_permute (#29227) 2026-09-21 17:25:44 +03:00
ggml.cpp ggml : Print backtrace on uncaught C++ exceptions (ggml/1232) 2025-06-01 13:43:57 +03:00
gguf.cpp gguf : align the data section relative to the GGUF start, not the file (#28993) 2026-09-17 09:19:44 +02:00