koboldcpp/ggml/src
2026-09-17 11:16:35 +03:00
..
ggml-blas llama: add default load-mode auto, which avoids mmap on iGPUs (#26081) 2026-08-11 09:20:46 +03:00
ggml-cann ggml: add SWIGLU_CLAMP (#27930) 2026-08-30 23:00:02 +08:00
ggml-cpu spacemit : fix wrong transpose function for int16 data (#25161) 2026-09-16 14:19:47 +03:00
ggml-cuda CUDA/HIP: improve access patterns in im2col (#28013) 2026-09-16 13:46:21 +02:00
ggml-et ggml: add SWIGLU_CLAMP (#27930) 2026-08-30 23:00:02 +08:00
ggml-hexagon hexagon: Support for K-Quants Q4_K and Q6_K (#28994) 2026-09-16 09:00:31 -07:00
ggml-hip CUDA: replace GGML_FA_ALL_QUANTS with GGML_FA_QUANTS, more control over what is compiled (#28079) 2026-09-09 12:50:08 +02:00
ggml-metal qwen4exp: add hc ops (#28901) 2026-09-16 16:00:01 +08:00
ggml-musa CUDA: replace GGML_FA_ALL_QUANTS with GGML_FA_QUANTS, more control over what is compiled (#28079) 2026-09-09 12:50:08 +02:00
ggml-opencl opencl: fix various warnings (#28984) 2026-09-17 09:48:05 +03:00
ggml-openvino OpenVINO: optimize stateful decode and GPU MoE inference (#28638) 2026-09-15 12:29:19 +03:00
ggml-rpc rpc : invalidate cached compute graph when a referenced buffer is freed (#24292) 2026-09-16 14:03:11 +03:00
ggml-sycl sycl : fix the B70 mem allocate error when >19.3GB (#28953) 2026-09-17 09:56:03 +03:00
ggml-virtgpu ggml: allow passing alloc dependencies in graph_optimize (#27301) 2026-08-30 11:34:20 +08:00
ggml-vulkan vulkan: split buffers and debug code into separate files, add shared headers (#28732) 2026-09-17 11:16:35 +03:00
ggml-webgpu webgpu: align tensor bindings to the type block size (#28382) 2026-09-12 06:40:21 +02:00
ggml-zdnn llama: add default load-mode auto, which avoids mmap on iGPUs (#26081) 2026-08-11 09:20:46 +03:00
ggml-zendnn CI: Use LLVM's OpenMP over MSVC_DEBUG_non_redist on Windows (#26678) 2026-08-20 15:42:26 +02:00
CMakeLists.txt ggml : replace compile definitions with version.h.in (#28364) 2026-09-04 10:28:23 +02:00
ggml-alloc.c ggml : fix ggml_clamp (#27644) 2026-08-24 10:43:04 +03:00
ggml-backend-dl.cpp hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150) 2026-01-29 12:33:21 -08:00
ggml-backend-dl.h hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150) 2026-01-29 12:33:21 -08:00
ggml-backend-impl.h sync : ggml (#28379) 2026-09-04 14:39:19 +03:00
ggml-backend-meta.cpp ggml-meta: propagate buffer usage and call init on the new tensors (#27586) 2026-08-26 08:27:51 +03:00
ggml-backend-reg.cpp ggml : don't crash when backend search path can't be read (#28271) 2026-09-04 10:24:06 +03:00
ggml-backend.cpp ggml: skip 0-sized ids tensor when offloading selected experts (#28739) 2026-09-11 15:17:08 +02:00
ggml-common.h AVX2: Speed up large batch size prompt processing of IQ models (#27402) 2026-08-31 14:33:50 -04:00
ggml-feats.h ggml : fix arm builds, unused var (#26991) 2026-08-13 07:57:24 +03:00
ggml-impl.h ggml : update ggml_prec specification (#26675) 2026-09-08 09:06:24 +03:00
ggml-opt.cpp fix: free ctx_copy in ggml_opt_free to plug per-training-session leak (#21592) 2026-04-08 17:40:15 +02:00
ggml-quants.c Add Q2_0 quantization: type definition and CPU backend (#24448) 2026-07-07 12:05:47 -07:00
ggml-quants.h Add Q2_0 quantization: type definition and CPU backend (#24448) 2026-07-07 12:05:47 -07:00
ggml-threading.cpp ggml : build backends as libraries (#10256) 2024-11-14 18:04:35 +01:00
ggml-threading.h remove CMAKE_WINDOWS_EXPORT_ALL_SYMBOLS (#10797) 2024-12-12 19:02:49 +01:00
ggml-version.h.in ggml : replace compile definitions with version.h.in (#28364) 2026-09-04 10:28:23 +02:00
ggml.c qwen4exp: add hc ops (#28901) 2026-09-16 16:00:01 +08:00
ggml.cpp ggml : Print backtrace on uncaught C++ exceptions (ggml/1232) 2025-06-01 13:43:57 +03:00
gguf.cpp gguf : align the data section relative to the GGUF start, not the file (#28993) 2026-09-17 09:19:44 +02:00