koboldcpp

mirror of https://github.com/LostRuins/koboldcpp.git synced 2026-05-10 04:00:53 +00:00

Author	SHA1	Message	Date
Wagner Bruna	e971eaefe3	fix: qwenvl.hpp was renamed to llm.hpp	2025-12-01 19:15:38 -03:00
Wagner Bruna	53407c866d	fix mistral vocab filename and qwen2vl tensor name	2025-12-01 17:31:02 -03:00
Wagner Bruna	07f53a4d84	fix qwen2vl tensor detection	2025-12-01 17:04:56 -03:00
Wagner Bruna	3db48a1536	sd: sync to master-383-20eb674	2025-12-01 17:04:56 -03:00
Concedo	438eae7105	mistral2 vocab for sdcpp	2025-12-01 22:32:58 +08:00
Concedo	177e0d7515	strip_oaicontent_of_media placeholder (+2 squashed commit) Squashed commit: [7ccd52ef4] placeholder [71fd2d7bb] strip_oaicontent_of_media	2025-12-01 01:29:57 +08:00
Concedo	addea2b62a	Merge commit '`47a268ea50`' into concedo_experimental # Conflicts: # ggml/src/ggml-vulkan/vulkan-shaders/vulkan-shaders-gen.cpp # tests/test-backend-ops.cpp	2025-11-30 17:28:07 +08:00
Concedo	95be49ac19	sync write_output_files function	2025-11-30 17:21:38 +08:00
Concedo	ef992b4ab7	cleanup	2025-11-30 15:55:58 +08:00
Concedo	bf5efcf86d	Merge commit '`d82b7a7c1d`' into concedo_experimental # Conflicts: # ci/run.sh # ggml/CMakeLists.txt # ggml/src/CMakeLists.txt # ggml/src/ggml-cuda/common.cuh # tests/CMakeLists.txt	2025-11-30 15:43:11 +08:00
Concedo	65a3b75dac	rnn warning fix	2025-11-30 12:55:43 +08:00
Concedo	2985575be4	allow assistant prefills, fixed showgui issue	2025-11-30 12:52:28 +08:00
Concedo	925e7f8f6d	added a secondary terminal mirror for linux	2025-11-29 21:53:51 +08:00
Ruben Garcia	06d39dff73	Fix warnings (#1864 )	2025-11-29 20:18:38 +08:00
Concedo	9999b8950d	cleaner resizing	2025-11-29 18:01:49 +08:00
Ruben Ortlam	47a268ea50	Vulkan: MMVQ Integer Dot K-Quant and MUL_MAT_ID support (#16900 ) * vulkan: split mul_mmq_funcs for mul_mat_vecq use * add mxfp4 mmvq * add q2_k mmvq * add q3_k mmvq * add q4_k and q5_k mmvq * add q6_k mmvq * handle 4x4 quants per mmvq thread * enable MUL_MAT_ID mmvq support * enable subgroup optimizations for mul_mat_vec_id shaders * device tuning * request prealloc_y sync after quantization * fix indentation * fix llvmpipe test failures * fix mul_mat_id mmvq condition * fix unused variable warning	2025-11-29 09:37:22 +01:00
Jeff Bolz	59d8d4e963	vulkan: improve topk perf for large k, fix overflow in unit tests (#17582 )	2025-11-29 08:39:57 +01:00
Aleksei Nikiforov	d82b7a7c1d	gguf-py : fix passing non-native endian tensors (editor-gui and new-metadata) (#17553 ) gguf_new_metadata.py reads data from reader. Reader doesn't byteswap tensors to native endianness. But writer does expect tensors in native endianness to convert them into requested endianness. There are two ways to fix this: update reader and do conversion to native endianness and back, or skip converting endianness in writer in this particular USE-case. gguf_editor_gui.py doesn't allow editing or viewing tensor data. Let's go with skipping excessive byteswapping. If eventually capability to view or edit tensor data is added, tensor data should be instead byteswapped when reading it.	2025-11-28 20:53:01 +01:00
DAN™	03914c7ef8	common : move all common_chat_parse_* to chat-parser.cpp. (#17481 )	2025-11-28 19:29:36 +01:00
o7si	3ce7a65c2f	server: fix: /metrics endpoint returning JSON-escaped Prometheus format (#17386 ) * fix: /metrics endpoint returning JSON-escaped Prometheus format * mod: remove string overload from ok() method	2025-11-28 19:14:00 +01:00
Diego Devesa	e072b2052e	ggml : add GGML_SCHED_NO_REALLOC option to disable reallocations in ggml_backend_sched (#17276 ) * ggml : add GGML_SCHED_NO_REALLOC option to disable reallocations in ggml_backend_sched Enabled in ggml-ci for testing. * llama : update worst-case graph for unified cache * ci : disable op offload in some tests * fix spelling --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2025-11-28 17:33:23 +02:00
Concedo	0ccb298087	Merge commit '`ddf9f94389`' into concedo_experimental # Conflicts: # examples/model-conversion/scripts/causal/run-converted-model.sh # examples/model-conversion/scripts/causal/run-org-model.py # src/CMakeLists.txt # src/llama-quant.cpp # tools/server/README.md	2025-11-28 23:27:50 +08:00
R0CKSTAR	c6f7a423c8	[MUSA] enable fp16/fast_fp16/bf16_mma on PH1 (#17551 ) Some checks failed Python Type-Check / pyright type-check (push) Waiting to run Details Python check requirements.txt / check-requirements (push) Has been cancelled Details Check Pre-Tokenizer Hashes / pre-tokenizer-hashes (push) Has been cancelled Details * [MUSA] enable fp16/fast_fp16/bf16_mma on PH1 Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com> * Update ggml/src/ggml-cuda/fattn-vec.cuh Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Update ggml/src/ggml-cuda/fattn-vec.cuh Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Update ggml/src/ggml-cuda/fattn-tile.cuh Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Address review comments Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com> --------- Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com> Co-authored-by: Johannes Gäßler <johannesg@5d6.de>	2025-11-28 14:08:29 +01:00
Aman Gupta	2e7ef98f18	ggml-cuda: add stricter checking for fusion (#17568 ) * ggml-cuda: make conditions for fusion more explicit * ggml-cuda: remove size check as std::equal already does it	2025-11-28 20:34:51 +08:00
Fredrik Hultin	ddf9f94389	server : add Anthropic Messages API support (#17570 ) * server : add Anthropic Messages API support * remove -@pytest.mark.slow from tool calling/jinja tests * server : remove unused code and slow/skip on test_anthropic_vision_base64_with_multimodal_model in test_anthropic_api.py * server : removed redundant n field logic in anthropic_params_from_json * server : use single error object instead of error_array in streaming response handler for /v1/chat/completions and use unordered_set instead of set in to_json_anthropic_stream() * server : refactor Anthropic API to use OAI conversion * make sure basic test always go first * clean up * clean up api key check, add test --------- Co-authored-by: Xuan Son Nguyen <son@huggingface.co>	2025-11-28 12:57:04 +01:00
Piotr Wilkin (ilintar)	ff55414c42	model : Qwen3 Next (#16095 ) * Qwen3 Next - cleaned up version * Whitespaces and stuff * Correct minor errors * Update src/llama-model.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Misc. fixes. * Clean up code, add missing hybrid qualifier * Did someone transpose the SOLVE_TRI result matrix? Perhaps... * Whitespace * Proper tensors for cb calls * Use llama-graph.h vertical alignment * BROKEN: chunking * Set new tensors as inputs. * Proper chunk logic * It's the circle of life... * More shenanigans for n_seq > 1 * Nail in the coffin? * Fix Windows build * Eh, one fails on Windows, the other fails on Mac... just use general capture. * quant : cleanup * model : cleanup * qwen3 : cleanup * cont : cleanup * cont : cleanup * ggml : revert change * qwen3 : cleanup * cont : cleanup * Readd cmath * qwen3 : fix typo * Update convert_hf_to_gguf.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Usual suspects * fix my bad suggestion --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2025-11-28 12:02:56 +01:00
Concedo	d2d05bd365	Merge branch 'upstream' into concedo_experimental # Conflicts: # ggml/src/ggml-rpc/ggml-rpc.cpp	2025-11-28 18:45:43 +08:00
Concedo	b30f09db80	autoscroll fixes	2025-11-28 18:31:29 +08:00
Johannes Gäßler	73955f7d2a	CUDA: no FP16 arithmetic for vector FA kernel (#17558 )	2025-11-28 10:29:09 +01:00
Jeff Bolz	35cf8887e1	vulkan: Implement GGML_OP_TRI (#17503 ) * vulkan: Implement GGML_OP_TRI * check types match	2025-11-28 10:07:29 +01:00
Radoslav Gerganov	15d2b46b4d	rpc : cache and reuse compute graphs (#15405 ) Store the last computed graph and reuse it when possible. Also do not return response from GRAPH_COMPUTE and assume it always completes successfully. If this this is not the case, the server closes the connection. This saves us a network round trip to the server.	2025-11-28 08:33:51 +00:00
yulo	6bca76ff5e	HIP: enable mul_mat_f for RDNA4 (#17437 ) * enable mmf for rdna4 * move some mmvf to mmf * revert lds128 for wmma loading * Revert "revert lds128 for wmma loading" This reverts commit db9ae8b6b4738a5def5b393caa1611d52133e9b5. * Revert "enable mmf for rdna4" This reverts commit 698c9f24187b990e35c3b73a8067e5387e6ddbd4. * Revert "move some mmvf to mmf" This reverts commit 99b92bd6653cc8593607f641e44606391691792f. * enable mul_mat for rdna4 --------- Co-authored-by: zhang hui <you@example.com>	2025-11-28 08:24:30 +01:00
Concedo	6aa79513a9	Merge branch 'cuda-fa-vec-fix-overflow-2' into concedo_experimental	2025-11-28 13:27:16 +08:00
Concedo	eda4a312cb	Merge branch 'upstream' into concedo_experimental # Conflicts: # .devops/vulkan.Dockerfile # ggml/src/ggml-cpu/CMakeLists.txt # ggml/src/ggml-opencl/CMakeLists.txt # ggml/src/ggml-opencl/ggml-opencl.cpp # ggml/src/ggml-sycl/common.hpp # tests/test-backend-ops.cpp # tools/server/README.md	2025-11-28 13:22:02 +08:00
Concedo	e570478275	limit cuda arches + scale tweaks	2025-11-28 13:05:11 +08:00
Concedo	9a46faa1c3	fix for override tensors not passing correctly	2025-11-28 13:03:40 +08:00
Piotr Wilkin (ilintar)	cd0e3a7a3b	SOLVE_TRI CUDA kernel for small matrices (#17457 ) Some checks failed Python Type-Check / pyright type-check (push) Has been cancelled Details	2025-11-28 12:15:32 +08:00
Neo Zhang Jianyu	efaaccdd69	refactor pad_reflect_1d to make the UT case pass (#17204 ) Co-authored-by: Zhang Jianyu <zhang.jianyu@outlook.com>	2025-11-28 08:50:56 +08:00
Johannes Gäßler	b13fcf85c5	CUDA: no FP16 arithmetic for vector FA kernel	2025-11-27 21:13:46 +01:00
Jeff Bolz	4abef75f2c	vulkan: Implement SOLVE_TRI (#17486 ) * vulkan: Implement SOLVE_TRI * load B matrix through shared memory * use FLOAT_TYPE	2025-11-27 15:48:00 +01:00
Georgi Gerganov	c386114922	arch : add description about LLM_TENSOR_INFOS (#17550 )	2025-11-27 16:34:13 +02:00
Georgi Gerganov	6783b11fb0	models : fix LFM2 tensors (#17548 )	2025-11-27 16:04:29 +02:00
matt23654	909072abcf	cuda : fix UMA detection on discrete GPUs. (#17537 )	2025-11-27 13:35:35 +02:00
Alberto Cabrera Pérez	cd8370b408	ggml-cpu: aarm64: q4_K repack gemm and gemv implementations (dotprod only) (#17494 ) * Enabled q4_K_4x8 path * Fixed generic Q4_K 8x4 implementation * wip: dotprod gemm * Working arm q4_K dotprod gemm Signed-off-by: Alberto Cabrera <alberto.cabrera@liquid.ai> * Undo acc rename Signed-off-by: Alberto Cabrera <alberto.cabrera@liquid.ai> * Q4_K arm dotprod gemm Signed-off-by: Alberto Cabrera <alberto.cabrera@liquid.ai> * Fix: q4_qs reinterpret from uint to int Signed-off-by: Alberto Cabrera <alberto.cabrera@liquid.ai> * Removed comments * Fixed macro guards * Fixed unused vars in generic implementation * Fixed unused vars in 8x4 repack * Fixed unused vars in generic implementation, unneeded comment * Missing arch fallback for x86 * minor : style --------- Signed-off-by: Alberto Cabrera <alberto.cabrera@liquid.ai> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2025-11-27 13:25:14 +02:00
Eric Curtin	d21a76ac38	devops: Add build-essential to Ubuntu 26.04 image (#17531 ) This is no longer passing the build, needs more packages. Signed-off-by: Eric Curtin <eric.curtin@docker.com>	2025-11-27 18:35:47 +08:00
Aleksei Nikiforov	4fcd87cf7c	gguf-py : skip endian-conversion of MXFP4 data (#17523 ) * gguf_convert_endian.py: skip MXFP4 data * Use gguf.constants.GGML_QUANT_SIZES to determine block sizes	2025-11-27 11:35:38 +01:00
Acly	b78db3bd50	vulkan : move contiguous checks to device_supports_op (#17490 ) * vulkan : remove op_supports_incontiguous and add missing constraints in device_supports_op * im2col: remove contraints on src0 (kernel input)	2025-11-27 06:54:19 +01:00
Jeff Bolz	142df17c9c	vulkan: use a fixed 1KB buffer for the add_rms_fusion opt (#17514 )	2025-11-27 06:32:30 +01:00
Concedo	7527f1eff0	handle media for jinja path (+1 squashed commits) Squashed commits: [29d47d6b7] handle media for jinja path	2025-11-27 11:40:08 +08:00
Concedo	782ec5bffe	bad identifier name	2025-11-27 11:07:13 +08:00

1 2 3 4 5 ...

10544 commits