koboldcpp

mirror of https://github.com/LostRuins/koboldcpp.git synced 2026-04-28 03:30:20 +00:00

Author	SHA1	Message	Date
Concedo	fa3f86ee70	added simplepod cloud template to readme	2026-04-17 10:58:44 +08:00
Concedo	aed18cc901	swa padding default to 0	2026-04-17 10:54:14 +08:00
Sigbjørn Skjæret	30dce2cf29	cli : use get_media_marker (#22017 )	2026-04-17 00:12:31 +02:00
Xuan-Son Nguyen	089dd41fe3	cmake: use glob to collect src/models sources (#22005 )	2026-04-16 23:25:16 +02:00
nullname	85dde8dc4a	hexagon: optimize HMX matmul operations (#21071 ) * optimize hmx_mat_mul functions by calculating row and column tiles upfront * refactor core_dot_chunk_fp16 to use size_t for tile counts and improve readability * wip * set scale outside of loop * wip * refactor core_mma_chunk_fp16 and mat_mul_qk_0_d16a32 to use size_t for tile counts * wip * wip * refactor transfer_output_chunk_fp16_to_fp32 to use size_t for dimensions * refactor core_dot_chunk_fp16 to use size_t for tile row stride calculation * wip * refactor hmx_mat_mul functions to use hvx_vec_splat_f16 for column scales initialization * refactor hmx_mat_mul_permuted_w16a32_batched to streamline scale setting and locking * refactor core_dot_chunk_fp16 to improve tile stride calculations for output * refactor hmx_mat_mul functions to use Q6_V_vsplat_R for column scales initialization * fix compiling error * wip * optimize row and column tile indexing in core_mma_chunk_fp16 function * wip * Revert "wip" This reverts commit cde679eff79c4a28dd2d89d32f710015e09592b6. * Add size limit check for HAP_mmap in htp_iface_mmap and drop_mmap functions * wip	2026-04-16 13:48:34 -07:00
Xuan-Son Nguyen	4fbdabdc61	model: using single llm_build per arch (#21970 ) * model: using single llm_build per arch * fix merge * nits	2026-04-16 21:10:22 +02:00
shaofeiqi	e45dbdece8	opencl: add q5_K gemm and gemv kernels for Adreno (#21595 )	2026-04-16 12:08:33 -07:00
Pascal	4adac43f6f	server: tests: fetch random media marker via /apply-template (#21962 ) (#21980 ) * server: tests: fetch random media marker via /apply-template (#21962 fix) * server: allow pinning media marker via LLAMA_MEDIA_MARKER env var get_media_marker() checks LLAMA_MEDIA_MARKER at first call and uses it as-is if set, falling back to the random marker otherwise. Tests no longer need to fetch the marker dynamically via /apply-template: the fixture sets LLAMA_MEDIA_MARKER=<__media__> so the hardcoded prompts work as before. Address review feedback from ngxson * server: make get_media_marker() thread-safe via magic statics Use a C++11 static local with a lambda initializer instead of a global static with an empty-check. The runtime guarantees initialization exactly once without explicit locking. Address review feedback from ggerganov * nits * nits	2026-04-16 20:46:21 +03:00
Concedo	b5e317e015	SWA fix attempt 2	2026-04-17 00:33:45 +08:00
PikaPikachu	9db77a020c	model : refactor QKV into common build_qkv and create_tensor_qkv helpers (#21245 ) * model : refactor QKV into common build_qkv and create_tensor_qkv helpers * model : extend build_qkv to bert/mpt/dbrx/olmo/lfm2/nemotron-h/granite-hybrid/gemma3n-iswa/t5-dec and fix wqkv_s	2026-04-16 17:41:34 +02:00
Concedo	ab2c596718	updated lite	2026-04-16 23:21:57 +08:00
Sigbjørn Skjæret	f772f6e434	model : support NVFP4 tensors for Gemma4 (#21971 ) * support nvfp4 tensors for Gemma4 * add wo_s to build_attn * add wo_s to build_attn * fix glm4	2026-04-16 16:51:47 +02:00
Ruben Ortlam	b572d1ecd6	codeowners: add team member comments (#21714 )	2026-04-16 13:13:11 +03:00
Anav Prasad	03b3d07798	Convert: Fix NemotronH Config Parsing (#21664 ) * fix NemotronH vocab loading by using trust_remote_code for unsupported config patterns * fix NemotronH tokenizer loading by overriding set_vocab with trust_remote_code	2026-04-16 13:11:45 +03:00
Aman Gupta	3f7c29d318	ggml: add graph_reused (#21764 ) * ggml: add graph_reused * use versioning instead of reuse flag * increment version with atomic * use top bits for split numbering * add assert * move counter to ggml.c * set uid in split_graph only * fix windows * address further review comments * get next_uid rather than doing bit manipulation * rename + add comment about uid	2026-04-16 17:21:28 +08:00
Concedo	ae292c496e	handle SWA conflicting with rewind, increased default SWA padding.	2026-04-16 17:00:26 +08:00
Kusha Gharahi	ae2d34899e	metal: Implement ROLL op (#21946 ) * nix: support unified apple-sdk * Impl roll op for Metal * Revert "nix: support unified apple-sdk" This reverts commit abfa473360471532c547de8b202c780507924d4b. * update ops.md * update op docs	2026-04-16 11:54:37 +03:00
Concedo	0251c6dbde	added swa padding controls	2026-04-16 16:21:48 +08:00
rehan-10xengineer	1e796eb41f	ggml-cpu: add 128-bit RVV implementation for Quantization Vector Dot (#20633 ) * ggml-cpu: add 128-bit impls for i-quants, ternary quants * ggml-cpu: add 128-bit impls for iq2_xs, iq3_s, iq3_xxs, tq2_0 Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai> * ggml-cpu: refactor; add rvv checks --------- Co-authored-by: taimur-10x <taimur.ahmad@10xengineers.ai> Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>	2026-04-16 11:15:15 +03:00
rehan-10xengineer	5637536517	ggml : implemented simd_gemm kernel for riscv vector extension (#20627 ) Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>	2026-04-16 11:14:26 +03:00
Yuannan	90fb96a7b3	devops : added spirv-headers to nix (#21965 )	2026-04-16 11:12:52 +03:00
Reese Levine	82677a6ede	ggml-webgpu: compute pass batching and removing profiling overhead (#21873 ) * Update register tiling matmul to use f32 accumulation * fix profiling code * Fix register tiling matmul for chrome, i'm blaming dawn * Update batch tuning value for iOS * compile fix * Fix use of new load function * Move to a single query set for GPU profiling * Move to batching compute passes when not profiling * Refactor build_multi * remove iOS throttling now that we're batching compute passes	2026-04-16 11:12:19 +03:00
Ludovic Henry	8612ed18b7	ci : Use ggml-org/ccache-action on RISC-V as well (#21632 )	2026-04-16 11:11:25 +03:00
Concedo	a9e817fb4c	smartcache off when fastforward off	2026-04-16 15:29:23 +08:00
Concedo	535df844dd	touchup for min/max tokens ui	2026-04-16 14:56:22 +08:00
Katostrofik	b1be68e8ca	[SYCL] Fix Q8_0 reorder: garbage on 2nd prompt + crash on full VRAM (#21638 ) * [SYCL] Fix Q8_0 reorder: add missing dequantize path for GEMM The Q8_0 reorder optimization (#21527) was missing a reorder-aware dequantizer for the GEMM code path used during prompt processing. After token generation reordered Q8_0 weights (via DMMV/MMVQ), the next prompt processing pass would read them with the standard dequantizer, producing garbage output. Add dequantize_block_q8_0_reorder() and wire it into both ggml_get_to_fp16_sycl() and ggml_get_to_fp32_sycl(), matching the pattern already used by Q4_0, Q4_K, and Q6_K. Fixes #21589 AI (Claude) was used to assist with root cause investigation and writing the kernel code. All code was human-reviewed and tested on real hardware. * SYCL: fix reorder crash when device memory is full The reorder optimization allocates a temporary buffer the full size of the weight tensor on the device. When VRAM is nearly full (large models on a single GPU), this allocation fails and the subsequent memcpy crashes on a NULL pointer. Fix: try device allocation first, fall back to host memory if device memory is full. The reorder kernel still works correctly reading from host memory over PCIe. This is slower for the one-time reorder (~21 t/s vs ~38 t/s on Intel Arc Pro B70), but the optimization is preserved for all subsequent inference. If both device and host allocation fail, skip the reorder and fall back to the unoptimized kernel path. Also fixes a bug where opt_for_reorder() marked tensors as reordered even when the reorder was skipped due to allocation failure. This caused DMMV/MMVQ kernels to read the original AoS data as if it were SoA, producing garbage output or NaN results. Tested on Intel Arc Pro B70 (32GB) with Q8_0, Q4_K_M models. Coding was AI-assisted (Claude), reviewed and tested on hardware by a human. Fixes #20478 * SYCL: add RAII temp buffer class + macro guard for host fallback Replace sycl_ext_malloc_with_fallback/sycl_ext_free_fallback free functions with sycl_reorder_temp_buffer RAII class. The host_fallback bool is now a private member, and cleanup happens automatically at scope exit. Add GGML_SYCL_HOST_MEM_FALLBACK cmake option (default ON) to guard the host memory fallback code path. Device access to host memory requires Linux kernel 6.8+ (Ubuntu 26.04+); users on older kernels can set -DGGML_SYCL_HOST_MEM_FALLBACK=OFF to disable it. Addresses arthw's review on PR #21638. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * SYCL: document GGML_SYCL_HOST_MEM_FALLBACK build option in SYCL.md Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * SYCL: add reorder-aware DMMV dequantizers for Q4_K and Q6_K Q4_K and Q6_K had reorder support for MMVQ and GEMM paths but not DMMV. When the DMMV path encountered reordered data it would abort. Add DMMV kernels that read from the SOA reorder layout for both types. Same math as the non-reorder versions, different memory access pattern. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-16 08:34:05 +03:00
Llama	c592bd01da	Pass img_min_params and img_max_params to ctx_clip_params (#2133 ) * Pass img_min_params and img_max_params to ctx_clip_params These values determine the minimum and maximum size (in tokens) of vision embeddings. The default value of -1 uses a model-dependent default size, for example for Gemma 4 the default is a 280 token embedding. For higher quality results (at the cost of using more memory and slower speed) you can increase the size of the embedding to 1120 tokens. * Change dict to mydict to match change to method	2026-04-16 12:27:06 +08:00
Concedo	a9f9e9a38b	rename the filepaths for clarity (+1 squashed commits) Squashed commits: [fa8fc6914] rename the filepaths for clarity	2026-04-16 12:17:23 +08:00
Concedo	45737effd3	refactor for clarity	2026-04-16 10:53:35 +08:00
Xuan-Son Nguyen	408225bb1a	server: use random media marker (#21962 ) * server: use random media marker * nits * remove legacy <__image__> token * revert special char in random	2026-04-15 23:52:22 +02:00
Ruben Ortlam	b3d758750a	vulkan: optimize im2col (#21713 ) * vulkan: improve im2col memory write layout * cap workgroups * minimal device tuning * use vendor_id instead of subgroup size	2026-04-15 19:04:51 +02:00
Pasha Khosravi	7e72b38bc1	cuda: Q1_0 initial backend (#21629 ) * [cuda] initial Q1_0 backend * remove unused code, fix AMD MMA guard * attempt to support dp4a * Apply suggestions from code review Co-authored-by: Johannes Gäßler <johannesg@5d6.de> --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>	2026-04-15 18:38:38 +02:00
Reese Levine	20d3bc2cc8	ggml-webgpu: Fix dequantization helpers to not pass in pointers (#21872 ) * Fix dequantization helpers to not pass in pointers * Increase XIELU precision	2026-04-15 09:14:40 -07:00
Concedo	9042f3fec8	updated lite	2026-04-15 22:52:40 +08:00
Rose	2f67e9f096	new baseconfig setting that aworks in router mode (#2130 ) * new baseconfig setting that aworks in router mode * re-added fix that prevents unneccessary model reload * fixed the fix * swapped order of baseconfig <-> override * fix indent * simplify baseconfig, if specified AND restart_override_config_target is NOT, it simply replaces the field (+1 squashed commits) Squashed commits: [95e816b16] simplify baseconfig, if specified AND restart_override_config_target is NOT, it simply replaces the field --------- Co-authored-by: Concedo <39025047+LostRuins@users.noreply.github.com>	2026-04-15 22:50:47 +08:00
Johannes Gäßler	a6206958d2	CUDA: require explicit opt-in for P2P access (#21910 )	2026-04-15 16:01:46 +02:00
Johannes Gäßler	014dca49d6	CUDA: manage NCCL communicators in context (#21891 ) * CUDA: manage NCCL communicators in context * add check that all backends are CUDA * remove unused vector, limit init to > 1 GPUs * fix warnings * fix cuda device, cache allreduce	2026-04-15 15:58:40 +02:00
Valeriy Dubov	adb541a6ad	rpc : add native RDMA transport for RPC backend (RoCEv2) (#20590 )	2026-04-15 16:44:02 +03:00
Xuan-Son Nguyen	80d8770804	docs: more extensive RoPE documentation [no ci] (#21953 ) * more extensive ggml_rope documentation * add more docs * nits	2026-04-15 14:45:16 +02:00
Ruben Ortlam	8dc530b86d	ci: disable test-backend-ops on Vulkan llvmpipe run and resture default timeout (#21901 )	2026-04-15 10:55:21 +02:00
Piotr Wilkin (ilintar)	e1a9a6dcbe	autoparser: support case of JSON_NATIVE with per-call markers (test case: Reka-Edge) (#21892 )	2026-04-15 10:51:50 +02:00
Matt	e39eba26f3	read n_ctx back after making llama_context (#21939 )	2026-04-15 15:24:57 +08:00
Concedo	ac29e6f0c0	Merge branch 'upstream' into concedo_experimental # Conflicts: # .devops/vulkan.Dockerfile # .github/workflows/build-self-hosted.yml # .github/workflows/build.yml # .github/workflows/release.yml # .github/workflows/server-self-hosted.yml # docs/build.md # ggml/src/ggml-hexagon/htp/CMakeLists.txt # ggml/src/ggml-hexagon/htp/hex-utils.h # ggml/src/ggml-hexagon/htp/hmx-matmul-ops.c # ggml/src/ggml-hexagon/htp/hmx-utils.h # ggml/src/ggml-hexagon/htp/htp-ctx.h # ggml/src/ggml-hexagon/htp/htp-ops.h # ggml/src/ggml-hexagon/htp/hvx-base.h # ggml/src/ggml-hexagon/htp/main.c # ggml/src/ggml-webgpu/ggml-webgpu.cpp # tests/test-backend-ops.cpp # tests/test-mtmd-c-api.c	2026-04-15 15:15:19 +08:00
Yiwei Shao	5d14e5d19b	hexagon: optimization for HMX mat_mul (#21554 ) * hexagon: add async HMX worker Introduce hmx-worker (dedicated thread for HMX compute) to overlap HMX matmul with HVX dequant/DMA stages in the pipeline path, replacing the previous synchronous HMX calls that blocked the main thread. * hexagon: cost-based VTCM chunk search for out-stationary matmul * hexagon: fix futex race in hmx_worker_drain Store the boolean to local variable avoid atomic load twice * hex-mm: hmx optimize scatter/transpose and use HMX intrinsics * hex-vmem: drop vmem limit a touch under 3GB on v73 * hexagon: add fwd declaration of htp_context * hex-hmx: replace hmx-worker with hmx-queue that mimics dma-queue interface Simplifies the overall implemantion, reduces thread wakeup roundtrips. * hex-mm: add debug log to hmx work func called from hmx-queue * Update hmx-queue.h Co-authored-by: Max Krasnyansky <max.krasnyansky@gmail.com> --------- Co-authored-by: Kim-Chyan Gan <kgan@qti.qualcomm.com> Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com> Co-authored-by: Max Krasnyansky <max.krasnyansky@gmail.com>	2026-04-14 14:09:03 -07:00
Concedo	c6b59fc2c7	autoswap some edge conditions	2026-04-14 23:02:29 +08:00
Xuan-Son Nguyen	fae3a28070	ggml : remove ggml-ext.h (#21869 ) * ggml: correct placement of ggml-ext.h * ggml : remove ggml-ext.h --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2026-04-14 17:32:58 +03:00
Georgi Gerganov	c0de6eda72	metal : fix FA support logic (#21898 )	2026-04-14 17:32:29 +03:00
Xuan-Son Nguyen	707c0b7a6e	mtmd: add mtmd_image_tokens_get_decoder_pos() API (#21851 ) * mtmd: add mtmd_image_tokens_get_decoder_pos() API * consistent naming * fix build	2026-04-14 16:07:41 +02:00
Jeff Bolz	1f30ac0cea	vulkan: Programmatically add RoundingModeRTE to all shaders when the device supports it (#21572 ) * vulkan: Programmatically add RoundingModeRTE to all shaders when the device supports it * use FetchContent to get SPIRV-Headers * Fetch spirv-headers unconditionally * remove fetchcontent, rely on installed headers * fix ubuntu job * Update docs/build.md	2026-04-14 15:17:45 +02:00
Concedo	236ae27329	Merge branch 'upstream' into concedo_experimental # Conflicts: # .github/workflows/close-issue.yml # docs/multimodal.md # embd_res/templates/deepseek-ai-DeepSeek-V3.2.jinja # ggml/CMakeLists.txt # ggml/src/ggml-webgpu/ggml-webgpu.cpp # ggml/src/ggml-webgpu/wgsl-shaders/mul_mat_decls.tmpl # ggml/src/ggml-webgpu/wgsl-shaders/mul_mat_reg_tile.wgsl # ggml/src/ggml-webgpu/wgsl-shaders/mul_mat_subgroup_matrix.wgsl # tests/peg-parser/test-gbnf-generation.cpp # tests/test-chat.cpp	2026-04-14 21:01:41 +08:00

... 2 3 4 5 6 ...

13000 commits