..
ggml-blas
llama: add default load-mode auto, which avoids mmap on iGPUs ( #26081 )
2026-08-11 09:20:46 +03:00
ggml-cann
ggml: add SWIGLU_CLAMP ( #27930 )
2026-08-30 23:00:02 +08:00
ggml-cpu
ggml: avoid KleidiAI buffer type init on dispatch ( #27891 )
2026-09-02 09:16:15 +03:00
ggml-cuda
ggml-cuda : remove unused vars ( #28235 )
2026-09-02 18:54:11 +02:00
ggml-et
ggml: add SWIGLU_CLAMP ( #27930 )
2026-08-30 23:00:02 +08:00
ggml-hexagon
ggml-hexagon: add F16 support for unary ops ( #28228 )
2026-09-02 12:59:36 -07:00
ggml-hip
ggml-hip : remove -funsafe-math-optimizations ( #26696 )
2026-08-13 08:38:02 +02:00
ggml-metal
metal : add fa-vec tunings for M3 ( #28236 )
2026-09-02 20:13:12 +02:00
ggml-musa
ggml-cuda: native bf16 flash attention for vec kernel ( #20525 )
2026-03-22 11:05:51 +01:00
ggml-opencl
opencl: fix out‐of‐bound reads in the Adreno image kernels ( #27632 )
2026-09-01 22:28:45 -07:00
ggml-openvino
ggml: add SWIGLU_CLAMP ( #27930 )
2026-08-30 23:00:02 +08:00
ggml-rpc
rpc: avoid serializing buffers from other servers ( #26500 )
2026-08-30 20:26:16 +03:00
ggml-sycl
sycl: reduce redundant work in Q4_K multi-column MMVQ ( #27062 )
2026-09-03 14:59:06 +08:00
ggml-virtgpu
ggml: allow passing alloc dependencies in graph_optimize ( #27301 )
2026-08-30 11:34:20 +08:00
ggml-vulkan
vulkan: handle larger batch sizes (>4) efficiently for IQ3_S mat-vec ( #27449 )
2026-09-02 09:14:52 +03:00
ggml-webgpu
webgpu : avoid crash when offset is not multiple of 4 in WebGPU ggml_backend_tensor_get() implementation ( #28045 )
2026-08-31 16:04:38 +02:00
ggml-zdnn
llama: add default load-mode auto, which avoids mmap on iGPUs ( #26081 )
2026-08-11 09:20:46 +03:00
ggml-zendnn
CI: Use LLVM's OpenMP over MSVC_DEBUG_non_redist on Windows ( #26678 )
2026-08-20 15:42:26 +02:00
CMakeLists.txt
CI: Use LLVM's OpenMP over MSVC_DEBUG_non_redist on Windows ( #26678 )
2026-08-20 15:42:26 +02:00
ggml-alloc.c
ggml : fix ggml_clamp ( #27644 )
2026-08-24 10:43:04 +03:00
ggml-backend-dl.cpp
hexagon: enable offloading to Hexagon on Windows on Snapdragon ( #19150 )
2026-01-29 12:33:21 -08:00
ggml-backend-dl.h
hexagon: enable offloading to Hexagon on Windows on Snapdragon ( #19150 )
2026-01-29 12:33:21 -08:00
ggml-backend-impl.h
ggml: allow passing alloc dependencies in graph_optimize ( #27301 )
2026-08-30 11:34:20 +08:00
ggml-backend-meta.cpp
ggml-meta: propagate buffer usage and call init on the new tensors ( #27586 )
2026-08-26 08:27:51 +03:00
ggml-backend-reg.cpp
ggml-et: Initial ET backend ( #24179 )
2026-07-10 12:38:34 +08:00
ggml-backend.cpp
ggml : add MUL_MAT to the list of ops that may need additional memory (for WebGPU) ( #28071 )
2026-08-31 10:17:23 +02:00
ggml-common.h
AVX2: Speed up large batch size prompt processing of IQ models ( #27402 )
2026-08-31 14:33:50 -04:00
ggml-feats.h
ggml : fix arm builds, unused var ( #26991 )
2026-08-13 07:57:24 +03:00
ggml-impl.h
ggml: add graph_reused ( #21764 )
2026-04-16 17:21:28 +08:00
ggml-opt.cpp
fix: free ctx_copy in ggml_opt_free to plug per-training-session leak ( #21592 )
2026-04-08 17:40:15 +02:00
ggml-quants.c
Add Q2_0 quantization: type definition and CPU backend ( #24448 )
2026-07-07 12:05:47 -07:00
ggml-quants.h
Add Q2_0 quantization: type definition and CPU backend ( #24448 )
2026-07-07 12:05:47 -07:00
ggml-threading.cpp
ggml-threading.h
ggml.c
finetune: fix no KV cache ( #27199 )
2026-09-02 23:53:32 +02:00
ggml.cpp
gguf.cpp
gguf : harden loader against malformed tensor dims and metadata types ( #25596 )
2026-08-12 15:07:48 +03:00