koboldcpp

mirror of https://github.com/LostRuins/koboldcpp.git synced 2025-09-11 09:34:37 +00:00

Author	SHA1	Message	Date
Djip007	0bb2919335	llama : change cpu_buft_list order: ACCEL -> GPU host -> CPU extra -> CPU (#12632 ) this allow to use GPU host when possible over CPU repack. this have the same effect to resolve this issues (#12498) without completely disable CPU extra buffer. Co-authored-by: philou <philou@framework>	2025-03-29 14:07:37 +01:00
Jay	a69f846351	cmake : fix ccache conflict (#12522 ) If users already set CMAKE_C_COMPILER_LAUNCHER globally, setting it in cmake again will lead to conflict and compile fail. Signed-off-by: Jay <BusyJay@users.noreply.github.com>	2025-03-29 11:04:58 +01:00
hipudding	d07a0d7a79	CANN : remove clang-format in ggml-cann (#12607 )	2025-03-29 11:03:28 +01:00
Concedo	396875e1c4	update api docs and lite	2025-03-29 15:39:25 +08:00
Sigbjørn Skjæret	3714c3ee1a	llama : fix incorrect Qwen2Moe ffn_moe_out graph callback (#12631 )	2025-03-28 22:13:02 +01:00
Georgi Gerganov	b4ae50810e	metal : improve FA + improve MoE (#12612 ) * ggml : FA with different K, V head sizes (CPU) ggml-ci * metal : add FA with HS=192 * metal : extend FA to support different K and V head sizes ggml-ci * metal : add FA vector kernels for heads K 192 and V 128 ggml-ci * ggml : restrict op on other backends to equal head sizes ggml-ci * metal : optimize FA-vec kernel ggml-ci * metal : FA remove mq registers * metal : improve MoE mul_mat_id condition ggml-ci * metal : fix comments + remove unnecessary addition ggml-ci * metal : avoid too much shared memory usage with mul_mat_id ggml-ci	2025-03-28 20:21:59 +02:00
Icenowy Zheng	b86f600723	vulkan: fix coopmat shader generation when cross-compiling (#12272 ) * vulkan: fix coopmat shader generation when cross-compiling Previously the status of coopmat{,2} support isn't passed to the vulkan-shaders-gen project building on the host, which leads to build failure because of the cross-compiling code expecting coopmat{,2} shaders that didn't get generated. Fix this by passing the coopmat{,2} support status to vulkan-shaders subproject. Signed-off-by: Icenowy Zheng <uwu@icenowy.me> * Only call coop-mat shaders once * Fix whitespace --------- Signed-off-by: Icenowy Zheng <uwu@icenowy.me> Co-authored-by: bandoti <141645996+bandoti@users.noreply.github.com>	2025-03-28 14:51:06 -03:00
Johannes Gäßler	dd373dd3bf	llama: fix error on bad grammar (#12628 )	2025-03-28 18:08:52 +01:00
Benson Wong	5d01670266	server : include speculative decoding stats when timings_per_token is enabled (#12603 ) * Include speculative decoding stats when timings_per_token is true New fields added to the `timings` object: - draft_n : number of draft tokens generated - draft_accepted_n : number of draft tokens accepted - draft_accept_ratio: ratio of accepted/generated * Remove redundant draft_accept_ratio var * add draft acceptance rate to server console output	2025-03-28 10:05:44 +02:00
Radoslav Gerganov	ef03229ff4	rpc : update README for cache usage (#12620 )	2025-03-28 09:44:13 +02:00
amritahs-ibm	13731766db	llamafile : ppc64le GEMV forwarding for FP32. (#12594 ) This patch enables usage of MMA when one of the dimensions of the matrix(ie either M or N) is 1. This is useful in case of token generation where N < 2. The concept of 'GEMV Forwarding' is used where when one of the matrix has a single row/column, the elements are broadcasted, instead of using packing routine to prepack the matrix elements. This change results in 5% - 15% improvement in total speed(ie all tokens/total time), across various batch sizes. This is in comparision with the corresponding dot product implementation. The patch is tested with FP32 models of Meta-Lllama-3-8B, Mistral-7B, Llama-2-7B-chat-hf on a IBM POWER10 machine. Signed-off-by: Amrita H S <amritahs@linux.vnet.ibm.com>	2025-03-28 09:43:22 +02:00
Radoslav Gerganov	ab6ab8f809	rpc : send hash when tensor data is above some fixed threshold (#12496 ) * rpc : send hash when tensor data is above some fixed threshold ref #10095 * rpc : put cache under $HOME/.cache/llama.cpp * try to fix win32 build * another try to fix win32 build * remove llama as dependency	2025-03-28 08:18:04 +02:00
Piotr	2099a9d5db	server : Support listening on a unix socket (#12613 ) * server : Bump cpp-httplib to include AF_UNIX windows support Signed-off-by: Piotr Stankiewicz <piotr.stankiewicz@docker.com> * server : Allow running the server example on a unix socket Signed-off-by: Piotr Stankiewicz <piotr.stankiewicz@docker.com> --------- Signed-off-by: Piotr Stankiewicz <piotr.stankiewicz@docker.com>	2025-03-27 23:41:04 +01:00
Georgi Gerganov	2969019837	media : add SVG logo [no ci] (#12616 )	2025-03-27 23:09:05 +02:00
lhez	5dec47dcd4	opencl: add multi and vision rope, `gelu_quick` and `im2col` (#12600 ) * opencl: add `im2col` * opencl: add `gelu_quick` * opencl: add mrope * opencl: add vision rope	2025-03-27 08:08:08 -07:00
Si1w	f125b8dccf	llama : add PLM GGUF Conversion & Inference Support (#12457 ) * add edgellm model arch[conversation feature doesn't work] * remove output.weight layer for edgellm arch * [Model] update the name of the model * update the name of model arch in convert gguf * [Model] Refarctor the model arch into llama-model * [Bug] Fix the bug in create attn kv * [Code] Fix editorconfig erros * [Code] Remove Trailing whitespace * [Code] Remove Trailing whitespace * [Code] Change the order of model arch in list * [Code] Fix flake8 Lint errors * Remove trailing white space * [Code] Remove call in model arch	2025-03-27 12:49:15 +02:00
HighDoping	953c2a62cf	model : restore support for T5Encoder (#12590 )	2025-03-27 11:43:33 +01:00
Csaba Kecskemeti	d5c6309d91	convert : Support Qwen2_5_VLForConditionalGeneration (#12595 )	2025-03-27 11:11:23 +01:00
Georgi Gerganov	029c693fdc	sync : ggml ggml-ci	2025-03-27 10:09:29 +02:00
Georgi Gerganov	771d84371c	scripts : update sync + fix cmake merge ggml-ci	2025-03-27 10:09:29 +02:00
Georgi Gerganov	df0665a483	sync : ggml ggml-ci	2025-03-27 09:04:38 +02:00
Georgi Gerganov	0306aad1ca	cmake : sync/merge PowerPC build commands (#0 )	2025-03-27 09:04:38 +02:00
amritahs-ibm	c7b43ab608	llamafile : ppc64le MMA implementation for Q4_0. (#12489 ) This change upstreams llamafile's cpu matrix multiplication kernels for ppc64le ISA using MMA builtins. This patch handles matrix multiplication between quantised datatypes, block_q4_0 and block_q8_0. This change results in 5% - 50% improvement in total speed(ie all tokens/total time), across various batch sizes. The patch is tested with Meta-Lllama-3-8B, Mistral-7B, Llama-2-7B-chat-hf models on a IBM POWER10 machine. Signed-off-by: Amrita H S <amritahs@linux.vnet.ibm.com>	2025-03-27 08:51:47 +02:00
xctan	24feaec057	ggml : riscv: add 128-bit RVV support (#12530 ) * ggml : add 128-bit RVV support * ggml : revert to old RVV 256+ q2_K, q3_K, q4_K, q6_K impl * remove trailing whitespaces * restructure vector length selection code	2025-03-27 08:38:34 +02:00
Georgi Gerganov	f28bc4c286	llama : make loras compatible with repacking (#12593 ) * llama : make loras compatible with repacking ggml-ci * cont : simplify ggml-ci * cont : add TODO [no ci]	2025-03-27 08:24:10 +02:00
Concedo	6a709be50a	replace deprecated	2025-03-27 10:27:20 +08:00
Akarshan Biswas	f17a3bb4e8	SYCL: implement memset ggml backend buffer interface (#12580 ) * SYCL: implement memset ggml backend buffer interface * use GGML_ABORT macro * Do not wait for all queues to finish for memset operation	2025-03-27 09:46:00 +08:00
Slobodan Josic	bd40678df7	HIP: Add support for RDNA4 targets (#12372 )	2025-03-26 23:46:30 +01:00
Georgi Gerganov	b3298fa47a	metal : refactor mat-vec code (#12569 ) * metal : refactor mat-vec code ggml-ci * metal : rename all_sum -> sum_all ggml-ci * metal : fix comments [no ci] * metal : fix nr constant [no ci] * metal : mv q6_K support nr0 > 1 ggml-ci * metal : reduce register pressure ggml-ci * metal : fix typo [no ci] * metal : reduce register pressure ggml-ci	2025-03-26 21:38:38 +02:00
Michał Moskal	2447ad8a98	upgrade to llguidance 0.7.10 (#12576 )	2025-03-26 11:06:09 -07:00
Ivy233	02082f1519	clip: Fix llama-llava-clip-quantize-cli quantization error under CUDA backend (#12566 ) * [Fix] Compiling clip-quantize-cli and running it in a CUDA environment will cause ggml_fp16_to_fp32 to report an error when trying to access video memory. You need to switch to the CPU backend to run quantize. After the fix, it will automatically run in the CPU backend and will no longer be bound to CUDA. * [Fix]Roll back the signature and implementation of clip_model_load, and change the call in clip_model_quantize to clip_init.	2025-03-26 15:06:04 +01:00
Concedo	4bf675a83d	Merge branch 'upstream' into concedo_experimental # Conflicts: # ci/README.md # docs/build.md # examples/run/run.cpp	2025-03-26 21:22:57 +08:00
Concedo	b4a8a5a278	Added CLI chat mode minor cli fixes (+1 squashed commits) Squashed commits: [60af39a9] Added CLI chat mode	2025-03-26 21:01:58 +08:00
Georgi Gerganov	df4d20cd53	convert : fix squeeze for ssm_conv tensors (#12573 ) * convert : fix squeeze for ssm_conv tensors * convert : match ssm_conv tensors by type --------- Co-authored-by: Francis Couture-Harpin <git@compilade.net>	2025-03-26 08:21:05 -04:00
Georgi Gerganov	5ed38b6852	ggml : fix MUL_MAT_ID repack with Q8_K (#12544 ) * ggml : fix MUL_MAT_ID repack with Q8_K ggml-ci * ggml : improve repack templates ggml-ci	2025-03-26 13:02:00 +02:00
Concedo	75e7902789	add localtunnel fallback (+1 squashed commits) Squashed commits: [ff0a63f6] add localtunnel fallback	2025-03-26 17:35:59 +08:00
R0CKSTAR	fd7855f8f5	doc: [MUSA] minor changes (#12583 ) Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>	2025-03-26 09:09:48 +02:00
Sigbjørn Skjæret	53af4dba42	convert: fix Mistral3/Gemma3 model hparams init (#12571 ) * Fix Mistral3/Gemma3 model hparams init * set positional args correctly * use existing hparams if passed	2025-03-25 23:03:10 +01:00
Eric Curtin	ef19c71769	run: de-duplicate fmt and format functions and optimize (#11596 )	2025-03-25 18:46:11 +01:00
Concedo	ea358369cc	Merge branch 'upstream' into concedo_experimental # Conflicts: # ci/README.md # ci/run.sh # docs/backend/CUDA-FEDORA.md # docs/build.md # docs/install.md # ggml/src/ggml-cpu/CMakeLists.txt # ggml/src/ggml-cuda/common.cuh # tests/test-backend-ops.cpp	2025-03-26 00:18:01 +08:00
Concedo	2bdf1dacff	embeddings done	2025-03-25 22:41:46 +08:00
Dan Johansson	053b3f9aae	ggml-cpu : update KleidiAI to v1.5.0 (#12568 ) ggml-cpu : bug fix related to KleidiAI LHS packing Signed-off-by: Dan Johansson <dan.johansson@arm.com>	2025-03-25 13:10:18 +02:00
Akarshan Biswas	e2f560175a	SYCL: disable Q4_0 reorder optimization (#12560 ) ggml-ci	2025-03-25 18:40:18 +08:00
Dan Johansson	36ee06dd2d	docs : add build instructions for KleidiAI (#12563 ) Signed-off-by: Dan Johansson <dan.johansson@arm.com>	2025-03-25 11:35:20 +02:00
R0CKSTAR	3cd3a39532	ci: [MUSA] add CI and update doc (#12562 ) Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>	2025-03-25 09:45:08 +02:00
Georgi Gerganov	2d77d88e70	context : fix worst-case reserve outputs (#12545 ) ggml-ci	2025-03-25 09:19:23 +02:00
Akarshan Biswas	c95fa362b3	ci: [SYCL] ggml-ci Use main GPU and enable sysman (#12547 )	2025-03-24 19:35:38 +02:00
lhez	2b65ae3029	opencl: simplify kernel embedding logic in cmakefile (#12503 ) Co-authored-by: Max Krasnyansky <quic_maxk@quicinc.com>	2025-03-24 09:20:47 -07:00
Concedo	82f2654049	wip embeddings model	2025-03-25 00:18:02 +08:00
Akarshan Biswas	48d7021c61	CI: fix SYCL build (#12546 )	2025-03-24 14:58:32 +02:00

... 4 5 6 7 8 ...

7620 commits