Commit graph

15728 commits

Author SHA1 Message Date
Concedo
635f05b657 wip kcpp agent 2026-09-21 01:28:40 +08:00
Concedo
a1b8a76cec added a simple kobold_agent.py 2026-09-21 00:11:24 +08:00
Concedo
0e49e296a6 updated lite (+1 squashed commits)
Squashed commits:

[5f32d56b4] updated lite
2026-09-20 23:18:41 +08:00
Concedo
e5e51240b5 limit max preview size 2026-09-20 20:57:30 +08:00
Concedo
ef51111b48 improved autoswap, added --autoswapthreshold which allows autoswap to trigger only if target swap exceeds threshold 2026-09-20 19:33:55 +08:00
Concedo
f44706a190 updated readme (+1 squashed commits)
Squashed commits:

[90fdd39c7] updated readme
2026-09-20 18:37:35 +08:00
Concedo
8a9833dad1 updated lite 2026-09-20 11:37:07 +08:00
Concedo
1859dbed49 preserved tokens 2026-09-20 10:35:37 +08:00
Concedo
1e300ce8b1 512 as the vae tiling threshold for extra safety 2026-09-19 23:12:19 +08:00
Concedo
9998be8dfb revert cuda changes to fix p40 vram usage 2026-09-19 19:01:56 +08:00
Concedo
cc9fa7d5e4 fix sd logger move out of makefiles 2026-09-19 18:04:08 +08:00
Wagner Bruna
469f456003
sd: sync with master-852-14eddb3 (#2457)
* bump GGML_MAX_NAME to 160

https://github.com/leejet/stable-diffusion.cpp/pull/1950

* sd: sync with master-852-14eddb3

* fix int8 and fp8 tensor size validation

* temporarily disable LoRA caching
2026-09-19 17:20:14 +08:00
Concedo
8c5df0a015 Merge branch 'upstream' into concedo_experimental
# Conflicts:
#	ggml/src/ggml-hexagon/ggml-hexagon.cpp
#	ggml/src/ggml-hexagon/htp/CMakeLists.txt
#	ggml/src/ggml-hexagon/htp/flash-attn-ops.c
#	ggml/src/ggml-hexagon/htp/hmx-fa-kernels.h
#	ggml/src/ggml-hexagon/htp/htp-ctx.h
#	ggml/src/ggml-hexagon/htp/htp-ops.h
#	ggml/src/ggml-hexagon/htp/im2col-ops.c
#	ggml/src/ggml-hexagon/htp/main.c
#	ggml/src/ggml-opencl/CMakeLists.txt
#	ggml/src/ggml-opencl/ggml-opencl.cpp
#	tests/test-backend-ops.cpp
#	tests/test-llama-archs.cpp
2026-09-19 17:00:05 +08:00
Georgi Gerganov
60b06ab9a9
metal : fix FA support checks (#29122) 2026-09-19 11:33:03 +03:00
Georgi Gerganov
efa28e950e
test-llama-archs : generate dummy test vocab (#29084)
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
2026-09-19 11:27:46 +03:00
Georgi Gerganov
59fc5a1ca3
metal : support qwen4exp hc ops (#29000)
Add support for the new DSV4 HC op variants used by qwen4exp:
- hc_pre with per-element sigmoid gate (gated variant)
- hc_post with identity mixing (comb == nullptr)

Assisted-by: pi:llama.cpp/Qwen3.8-27B
2026-09-19 11:27:30 +03:00
TheArchitectit
b23701f77d
cuda : fix CUB argsort corruption caused by in-place keys (#28389)
argsort_f32_i32_cuda_cub called the one-shot DeviceRadixSort::SortPairs
API with d_keys_in == d_keys_out (temp_keys, temp_keys). CUB's internal
double-buffer ping-pong requires distinct key buffers: with aliased
buffers the sort partially overwrites its own input mid-pass and emits a
corrupted permutation, surfacing as intermittent garbage indices (e.g.
backend top_k over a 248k-column vocab on Maxwell/CUDA 12.5/CCCL 2.x,
which then triggered out-of-bounds gathers in downstream get_rows).

Use a distinct keys-out buffer for all six call sites (plain and
segmented, ascending and descending, size-query and execute).

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Oliver Simons <osimons@nvidia.com>
2026-09-19 07:32:52 +02:00
Concedo
a9b4c044dc fixed: Generation failures return an error object: HTTP 500 before headers, or an SSE error after streaming starts.
Custom stop strings report "stop".
Fake-stream chunks retain a unique, consistent response ID.
Anthropic usage reports correct input/output counts.
2026-09-19 11:49:20 +08:00
Concedo
a5ca9da631 fix audio gen ui, fix finish reason 2026-09-19 11:18:57 +08:00
Concedo
a8b0069fac Merge branch 'upstream' into concedo_experimental
# Conflicts:
#	.devops/openvino.Dockerfile
#	.github/workflows/build-and-test-snapdragon.yml
#	.github/workflows/build-android.yml
#	.github/workflows/build-cache.yml
#	.github/workflows/build-openvino.yml
#	.github/workflows/build-self-hosted.yml
#	.github/workflows/copilot-setup-steps.yml
#	.github/workflows/gguf-publish.yml
#	.github/workflows/make-release.yml
#	.github/workflows/release.yml
#	.github/workflows/winget.yml
#	docs/backend/OPENVINO.md
#	examples/convert-llama2c-to-ggml/CMakeLists.txt
#	ggml/src/ggml-openvino/ggml-decoder.cpp
#	ggml/src/ggml-openvino/ggml-decoder.h
#	ggml/src/ggml-openvino/ggml-openvino-extra.cpp
#	ggml/src/ggml-openvino/ggml-openvino.cpp
#	ggml/src/ggml-openvino/ggml-quants.cpp
#	ggml/src/ggml-openvino/ggml-quants.h
#	ggml/src/ggml-openvino/model-cache.cpp
#	ggml/src/ggml-openvino/openvino/frontend.h
#	ggml/src/ggml-openvino/openvino/op/add_id.cpp
#	ggml/src/ggml-openvino/openvino/op/cont.cpp
#	ggml/src/ggml-openvino/openvino/op/flash_attn_ext.cpp
#	ggml/src/ggml-openvino/openvino/op/gated_delta_net.cpp
#	ggml/src/ggml-openvino/openvino/op/im2col.cpp
#	ggml/src/ggml-openvino/openvino/op/mul_mat_id.cpp
#	ggml/src/ggml-openvino/openvino/op/pad.cpp
#	ggml/src/ggml-openvino/openvino/op/repeat.cpp
#	ggml/src/ggml-openvino/openvino/op/rms_norm.cpp
#	ggml/src/ggml-openvino/openvino/op/view.cpp
#	ggml/src/ggml-openvino/openvino/pass/kv_state_seq_axis.cpp
#	ggml/src/ggml-openvino/openvino/translate_session.cpp
#	ggml/src/ggml-openvino/openvino/utils.cpp
#	ggml/src/ggml-openvino/openvino/utils.h
#	ggml/src/ggml-openvino/utils.cpp
#	ggml/src/ggml-openvino/utils.h
#	ggml/src/ggml-webgpu/ggml-webgpu-shader-lib.hpp
#	ggml/src/ggml-webgpu/ggml-webgpu.cpp
#	pocs/CMakeLists.txt
#	tests/CMakeLists.txt
#	tests/test-backend-ops.cpp
#	tools/ui/tests/stories/a11y/ChatScreenForm.a11y.stories.svelte
2026-09-19 10:52:50 +08:00
Concedo
dbf4704e7e Merge commit 'f172be756a' into concedo_experimental
# Conflicts:
#	.github/workflows/make-release.yml
#	ggml/src/ggml-vulkan/CMakeLists.txt
#	ggml/src/ggml-vulkan/ggml-vulkan.cpp
#	scripts/check-apiabi-compat.sh
#	scripts/make-release-checks.sh
#	tests/test-gguf.cpp
2026-09-19 10:48:29 +08:00
Concedo
d1c93e8e89 Merge commit 'c9a5eeeb34' into concedo_experimental
# Conflicts:
#	.github/workflows/build-self-hosted.yml
#	.github/workflows/check-vendor.yml
#	.github/workflows/pre-tokenizer-hashes.yml
#	.github/workflows/python-check-requirements.yml
#	.github/workflows/python-lint.yml
#	.github/workflows/python-type-check.yml
#	.github/workflows/update-ops-docs.yml
#	CODEOWNERS
#	ggml/include/ggml-sycl.h
#	ggml/src/ggml-hexagon/ggml-hexagon.cpp
#	ggml/src/ggml-hexagon/htp/hmx-mm-kernels-tiled.h
#	ggml/src/ggml-hexagon/htp/htp-ops.h
#	ggml/src/ggml-hexagon/htp/hvx-mm-kernels-flat.h
#	ggml/src/ggml-hexagon/htp/hvx-mm-kernels-tiled.h
#	ggml/src/ggml-hexagon/htp/matmul-ops.c
#	ggml/src/ggml-hexagon/htp/matmul-ops.h
#	ggml/src/ggml-opencl/ggml-opencl.cpp
#	ggml/src/ggml-sycl/fusion.cpp
#	ggml/src/ggml-sycl/ggml-sycl.cpp
#	ggml/src/ggml-sycl/ssm_conv.cpp
#	ggml/src/ggml-sycl/ssm_conv.hpp
#	tests/test-backend-ops.cpp
#	tests/test-llama-archs.cpp
#	tools/llama-bench/llama-bench.cpp
2026-09-19 10:10:02 +08:00
dsproule
60081bb2b5
opencl: add support for bin kernel flash_attn_f32_f16_bin (#29046)
* opencl: add `flash_attn_f32_f16_bin`

* opencl: guarded prefill fa
2026-09-18 16:32:31 -07:00
Todor Boinovski
2b1847030c
hexagon: add ROLL op support (#29105) 2026-09-18 15:05:10 -07:00
Todor Boinovski
50631b3d2c
hexagon: im2col update (#29103)
* ggml-hexagon: accept 1D and padded IM2COL ops

* ggml-hexagon: make pure-DDR IM2COL kernel is_2D-aware

* ggml-hexagon: extend IM2COL DMA patch-embed fast path to 1D

* ggml-hexagon: add blocked-staging general IM2COL DMA kernel
2026-09-18 14:20:48 -07:00
Todor Boinovski
18a04f09c2
hexagon: HMX flash-attention head_dim padding (support DK=DV=72) (#26539)
Allow HMX flash-attention to run with head_dim not a multiple of 64
(e.g. SigLIP head_dim=72), by operating on DK/DV rounded up to 64 with
zero-filled tail lanes.
2026-09-18 13:15:08 -07:00
shaofeiqi
ec92815050
opencl: add bin kernel kernel_gemm_noshuffle_q6_k_f32_32b_trans_ila_a8_bin (#28678)
* opencl: add A8 Q6_K non-MoE binary kernel

* opencl: fix layout compatibility
2026-09-18 10:50:15 -07:00
Concedo
da73eecc5a don't build unwanted targets (e.g. avx for arm) 2026-09-19 01:02:19 +08:00
bri-prism
4fea119de3
ggml-cpu: add F16 input to the FWHT (#27779)
Some checks failed
Python Type-Check / python type-check (push) Has been cancelled
Update Operations Documentation / update-ops-docs (push) Has been cancelled
Copilot Setup Steps / copilot-setup-steps (push) Has been cancelled
Check Pre-Tokenizer Hashes / pre-tokenizer-hashes (push) Has been cancelled
Python check requirements.txt / check-requirements (push) Has been cancelled
* ggml-cpu: add F16 input to the FWHT

The CPU FWHT accepts F32 input only. This change makes the source type a
template parameter. The CPU path now accepts F16 input and F32 input.

The CPU MUL_MAT reference now converts an F16 src1 to F32. It does this when
the caller sets the Hadamard hint.

No backend has an F16 FWHT kernel yet. The test cases come with the backend
changes that add one.

* ggml-cpu: assert the F16 FWHT input path, and use the bulk converter

Address review feedback.

The F16 branch writes plain floats into wdata, which is only correct when
vec_dot_type is F32. That invariant held because supports_op only accepts an
F16 src1 for the Hadamard hint with F32 src0 and dst, but nothing enforced it.
Assert it next to the existing src1 type check so widening supports_op cannot
silently break the write.

Replace the hand-rolled conversion loop with ggml_cpu_fp16_to_fp32.
2026-09-18 17:17:38 +03:00
Alexey Kopytko
5b335f413e
ggml : check for allocation failures to prevent crashes (#28149)
* ggml : check for allocation failures to prevent crashes

* wording
2026-09-18 16:58:25 +03:00
Pascal
542348a35c
Model-Saver: Write the SWA pattern, 15 more architectures roundtrip (#29042)
* llama: read the SWA pattern as a period or a per-layer array

Add llama_model_base::load_swa_pattern(), which reads
sliding_window_pattern either as one flag per layer or as a period
expanded by set_swa_pattern(), and use it in every loader that reads
the key as a period.

These loaders silently ignored an array and applied their default
period, although the converters of olmo2, gemma3n and exaone4 write
arrays. The published GGUFs match the defaults, so their outputs do
not change. The loaders that already accepted both forms lose their
duplicated scalar-then-array block, and use their declared default
period when the key is absent.

* model-saver: write the SWA pattern and the MLA SWA geometry

Write sliding_window_pattern as one flag per layer, nextn layers
included, for every model using SWA. The array is never collapsed to
a scalar, since the loaders read a scalar as a period.

Also write the MLA key/value lengths and KV LoRA rank of the SWA
layers, required by dots3note.

This enables the saver for plamo3, gemma3, cohere2, cohere2moe,
olmo2, exaone-moe, afmoe, mimo2, spark2_5, muse-glimmer, mellum,
laguna, granite_swa, dots3note and maple, all passing the bit-exact
roundtrip of test-llama-archs.
2026-09-18 15:20:03 +02:00
Aaron Teo
d663dd3f3a
ci: change ubuntu-latest to ubuntu-24.04 (#29079) 2026-09-18 21:17:19 +08:00
Masashi Yoshimura
44be98f057
ggml-webgpu: fix supports_op condition for GET_ROWS (#28978)
* fix get_rows vec4 handling

* Add src strides checking to vec4_aligned of get_rows and the new test case.
2026-09-18 20:47:07 +09:00
z
911f6cdc8a
ggml : handle graph buffer reservation failure (#26070) 2026-09-18 12:31:19 +03:00
Daniel Varga
bbd488c42a
vulkan: add IQ3_S MMQ matmul kernels (#28822)
* vulkan: add IQ3_S MMQ matmul kernels

* Make block_a_to_shmem do 2-byte loads (110 bytes is divisible by 2)

* Align the check, IQ3_S is also using K tile size
2026-09-18 11:46:02 +03:00
Sait Furkan Teke
dc85f89c7e
vocab : add ufakzeka pre-tokenizer (#29033)
* vocab : add ufakzeka pre-tokenizer

* vocab : move ufakzeka to the models list and regenerate the hash mapping
2026-09-18 11:45:11 +03:00
Nikita Gordeev
8ed1a55efc
cmake : fix build when GGML_CPU=OFF and GGML_CUDA=ON (#29026)
* fix: build fails when GGML_CPU=OFF and GGML_CUDA=ON

* fix: eol in examples/convert-llama2c-to-ggml/CMakeLists.txt file
2026-09-18 11:44:03 +03:00
yanghong
bb11ebb682
gguf-py: fix Q8_1 block size in GGML_QUANT_SIZES (2+2+32) (#29036)
* gguf-py: fix Q8_1 block size in GGML_QUANT_SIZES

* --whitespace

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
2026-09-18 11:43:06 +03:00
Sigbjørn Skjæret
f03cf3e9b8
ci : disable GHA cache for copilot (#29068) 2026-09-18 10:20:53 +02:00
Sigbjørn Skjæret
bdcbaaf6e7
ci : bump android-actions/setup-android to 4.0.4 (#29065) 2026-09-18 09:04:14 +02:00
drluoto
5c53396b89
vulkan: raise the hoisted row-id limit for mul_mat_id from 256 to 512 experts (#28501)
* vulkan: raise the hoisted row-id limit for mul_mat_id to 512 experts

The expert-count shader (count_experts.comp) sizes its shared arrays
with BLOCK_SIZE, which is 256. Because of that, row-id hoisting is
switched off for any model with more than 256 experts, and every
mul_mat_id workgroup has to rescan the whole ids tensor on its own.
Qwen3.8-Flash-Next has 512 experts and was quietly running on that
slow path.

This change sizes the arrays with a separate MAX_EXPERTS constant (512),
clears them in a loop instead of one entry per thread, and raises the
matching limit on the host side.

On Strix Halo at batch 2048 the expert matmuls drop from 12.5 to 9.5 ms
(iq3_s) and from 14.0 to 7.5 ms (iq4_nl) per op, and prompt processing
gets about 19 % faster at 8k tokens. test-backend-ops MUL_MAT_ID passes
(891/891) with new 512-expert test cases.

Assisted-by: Claude Fable 5.1

* vulkan: raise the hoisted row-id limit for mul_mat_id to 1024 experts

Follow-up to review feedback: 1024 matches LLAMA_MAX_EXPERTS instead of
stopping at 512. The three shared arrays in count_experts.comp grow to
3 * 1024 * 4 = 12 KiB, which fits the 16 KiB that Vulkan guarantees for
maxComputeSharedMemorySize.

Adds mul_mat_id test cases at 1024 experts alongside the existing 512
ones. test-backend-ops MUL_MAT_ID passes on Vulkan (RADV, Strix Halo,
Radeon 8060S): 889/889.
2026-09-18 09:00:15 +02:00
henk717
6c8df269ea
Optimize koboldcpp.sh for local users (#2444)
* koboldcpp.sh is now optimized for local

* Restore 12.1 behavior when build system has no GPU

* Fix PORTABLE_SO text replace fail

* Restore accidentally deleted file

* More koboldcpp.sh fixes

* Move vulkan noavx2 to portable section

* Fix syntax on portable env

* Fix rocm env variables in CI
2026-09-18 10:49:52 +08:00
Sigbjørn Skjæret
972d2313bc
ci : add missing evict-old-files (#29041) 2026-09-17 19:05:40 +02:00
Concedo
7f1f979185 ubatch slider, make sliders more compact 2026-09-18 00:59:17 +08:00
Concedo
fb14096627 allow music llm mode to gen audio codes 2026-09-18 00:15:44 +08:00
Concedo
173d340c43 wip on adding ubatch 2026-09-18 00:15:26 +08:00
Pedro Cuenca
c77ae695c9
rpc : skip ACCEL devices (#29020) 2026-09-17 18:36:53 +03:00
Concedo
fc1046959e added section for lite.koboldai.net 2026-09-17 23:20:02 +08:00
Concedo
1c52d02b2f updated sdui 2026-09-17 23:02:24 +08:00
Concedo
b71ac32806 smartcache spam off during quiet mode 2026-09-17 22:57:53 +08:00