koboldcpp

mirror of https://github.com/LostRuins/koboldcpp.git synced 2026-05-31 05:03:44 +00:00

History

Oliver Simons 6ed481eea4 CUDA: Check PTX version on host side to guard PDL dispatch (#23530 ) * CUDA: Check PTX version on host side to guard PDL dispatch Checking on `__CUDA_ARCH_LIST__` alone is insufficient for JIT, as this variable doesn't differentiate between compiling for say sm_90, sm_90a or sm_90f (so forward-jittable PTX vs. arch/family-specific PTX). Thus, one can have a bug when compiling with `DCMAKE_CUDA_ARCHITECTURES="89;90a"`, where current code would wrongly dispatch to PDL on sm_90/sm_120 in forward-JIT mode. This PR fixes this issue by checking `cudaFuncAttributes::ptxVersion` of the incoming kernel at runtime. A check on ptxVersion alone is sufficient, as device-codes will always be >= ptxVersion (and any violation of this would be a severe bug in CUDA/nvcc), see: https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/#gpu-code-code-code * Implement MurmurHash3 mixer for better hash distribution Magic constants were taken from boost: `2698b43803/include/boost/container_hash/detail/hash_mix.hpp (L19-L65)` * Update ggml/src/ggml-cuda/common.cuh Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Address review comments, make seed non-zero * Apply code-formatting * Replace std::size_t -> size_t for consistency --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>		2026-05-29 12:28:18 +02:00
..
cmake	ggml : Parallelize quant LUT init (#23595 )	2026-05-25 10:15:46 +03:00
include	ggml.h: correct ggml_silu_back arg docstring (a=dy, b=x) (ggml/1500)	2026-05-25 12:38:01 +03:00
src	CUDA: Check PTX version on host side to guard PDL dispatch (#23530 )	2026-05-29 12:28:18 +02:00
.gitignore	vulkan : cmake integration (#8119 )	2024-07-13 18:12:39 +02:00
CMakeLists.txt	ggml : bump version to 0.13.1 (ggml/1523)	2026-05-29 09:56:08 +03:00