Concedo
afc8de2c6b
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# .devops/cpu.Dockerfile
# .devops/cuda.Dockerfile
# .devops/intel.Dockerfile
# .devops/musa.Dockerfile
# .devops/openvino.Dockerfile
# .devops/rocm.Dockerfile
# .devops/vulkan.Dockerfile
# .devops/zendnn.Dockerfile
# .github/workflows/build-webgpu.yml
# .github/workflows/release.yml
# ggml/src/ggml-webgpu/ggml-webgpu-shader-lib.hpp
# ggml/src/ggml-webgpu/ggml-webgpu.cpp
# ggml/src/ggml-webgpu/wgsl-shaders/binary.wgsl
# ggml/src/ggml-webgpu/wgsl-shaders/concat.wgsl
# ggml/src/ggml-webgpu/wgsl-shaders/mul_mat_decls.tmpl
# ggml/src/ggml-webgpu/wgsl-shaders/scale.wgsl
# ggml/src/ggml-webgpu/wgsl-shaders/unary.wgsl
# tests/CMakeLists.txt
# tests/test-backend-ops.cpp
# tests/test-mtmd-c-api.c
# tools/cli/cli.cpp
# tools/mtmd/CMakeLists.txt
# tools/server/README.md
2026-06-10 17:21:05 +08:00
Concedo
6bae15da71
fix build, remove clip quantize (+1 squashed commits)
...
Squashed commits:
[09ffa906b] fix build, remove clip quantize
2026-06-10 00:42:34 +08:00
Concedo
fe13a2989d
migration to mtmd complete
2026-06-09 23:09:45 +08:00
Concedo
0740ce70b7
revert incorrect token count for mtmd
2026-06-09 22:46:17 +08:00
Concedo
ac8e77d82e
Revert "cleanup round 2"
...
This reverts commit c7e39f9c97 .
2026-06-09 22:33:00 +08:00
Concedo
da0ba5cefd
Revert "cleanup round 3"
...
This reverts commit 0ead9907f0 .
2026-06-09 22:32:42 +08:00
Concedo
0ead9907f0
cleanup round 3
2026-06-09 21:52:46 +08:00
Concedo
c7e39f9c97
cleanup round 2
2026-06-09 21:49:20 +08:00
Concedo
f5acccad63
cleanup round 1
2026-06-09 21:13:48 +08:00
Concedo
90a14cecf8
mtmd checkpoint 2
2026-06-09 17:55:17 +08:00
Concedo
794a271cfa
upgrade to mtmd checkpoint 1
2026-06-09 16:06:30 +08:00
Concedo
9755b05556
g4ua audio fix
2026-06-05 10:59:24 +08:00
Concedo
b8acf16e73
hacky fix that worked for gemma4 uv
2026-06-05 01:40:08 +08:00
Concedo
2cfcd40fd4
gemma4uv vision still a little buggy
2026-06-05 00:53:23 +08:00
Concedo
7fb55b1e32
does not work for e4b
2026-06-04 10:44:40 +08:00
Concedo
e434b3e61b
kcpp image handling fix
2026-05-26 14:32:53 +08:00
Concedo
ce3aa09b99
cache dir is null
2026-05-23 17:39:09 +08:00
Concedo
4bbbd55be6
rpc implementation is complete
2026-05-23 17:11:30 +08:00
Concedo
f85cc79526
make swa default on models that support it. removed --useswa, added --noswa
2026-05-22 23:38:33 +08:00
Concedo
694e8824c5
mmproj autofit reworked
2026-05-22 20:36:16 +08:00
Concedo
095bf63b58
prep for rpc
2026-05-20 23:29:49 +08:00
Concedo
9203b6a051
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# .github/labeler.yml
# .github/workflows/build-self-hosted.yml
# .github/workflows/release.yml
# .github/workflows/server-sanitize.yml
# .github/workflows/server-self-hosted.yml
# .github/workflows/server.yml
# .github/workflows/ui-build.yml
# .github/workflows/ui-ci.yml
# .github/workflows/ui-publish.yml
# .gitignore
# CMakeLists.txt
# CODEOWNERS
# scripts/ui-download.cmake
# scripts/xxd.cmake
# tests/test-backend-ops.cpp
# tests/test-reasoning-budget.cpp
# tools/CMakeLists.txt
# tools/server/CMakeLists.txt
# tools/server/README.md
2026-05-16 22:56:33 +08:00
Concedo
80ce8a50b3
allow token bans and eos handling in
2026-05-16 15:20:46 +08:00
askmyteapot
3174fccb83
Fix linker error for kcpp_permit_any_repack ( #2210 )
...
Fixes `gpttype_adapter.lib(gpttype_adapter.obj) : error LNK2019: unresolved external symbol "int kcpp_permit_any_repack"`
It's a `bool`, not an `int`
2026-05-16 08:55:56 +08:00
Concedo
66a7b5e5de
prevent repacking if mmap is used, reduces memory footprint
2026-05-15 16:40:53 +08:00
Concedo
286e62267e
adjust batching eligibility
2026-05-11 21:54:32 +08:00
Concedo
bfaddd7a3b
added support for added memory and gemma and glm prompt fixes for batching mode
2026-05-10 23:39:03 +08:00
Concedo
33ca75d56f
ci for tools upload, minor function reordering
2026-05-10 23:10:43 +08:00
AlpinDale
c03302b670
feat: add a primitive form of continuous batching ( #2167 )
...
* feat: add a primitive form of continuous batching
* fix: deadlock in batching fallback
* fix: windows build
* chore: suppress the contbatch arg from --help
* feat: batch-aware rep_pen_slope
* fix: automatically disable shifting when batching is enabled
* fix: mixed-path state corruption
* fix: attempt to fully separate the two pipelines
* added a semaphore to prevent non-batchable requests from starting while batched requests are running
---------
Co-authored-by: Concedo <39025047+LostRuins@users.noreply.github.com>
2026-05-10 17:50:31 +08:00
Concedo
eb30b29d69
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# .github/workflows/gguf-publish.yml
# CODEOWNERS
# examples/sycl/test.sh
# pyproject.toml
# tools/mtmd/CMakeLists.txt
# tools/mtmd/README.md
2026-05-08 14:48:57 +08:00
Concedo
950676fdb7
split utils.cpp into 2 files to support sd.cpp
2026-05-04 15:04:12 +08:00
Concedo
8b62e7b667
allow splitmode to be set independently, enable tensor parallelism
2026-05-02 16:41:28 +08:00
Concedo
0755f27372
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# .devops/openvino.Dockerfile
# .github/workflows/build-self-hosted.yml
# .github/workflows/build.yml
# common/chat.cpp
# docs/backend/OPENVINO.md
# examples/speculative-simple/speculative-simple.cpp
# ggml/src/ggml-hexagon/ggml-hexagon.cpp
# ggml/src/ggml-hexagon/htp/CMakeLists.txt
# ggml/src/ggml-hexagon/htp/htp-ctx.h
# ggml/src/ggml-hexagon/htp/htp-ops.h
# ggml/src/ggml-hexagon/htp/main.c
# ggml/src/ggml-hexagon/libggml-htp.inf
# ggml/src/ggml-openvino/ggml-decoder.cpp
# ggml/src/ggml-openvino/ggml-openvino-extra.cpp
# ggml/src/ggml-openvino/ggml-openvino.cpp
# ggml/src/ggml-openvino/ggml-quants.cpp
# ggml/src/ggml-openvino/openvino/op/rope.cpp
# ggml/src/ggml-openvino/openvino/op_table.cpp
# ggml/src/ggml-openvino/openvino/op_table.h
# ggml/src/ggml-openvino/openvino/translate_session.cpp
# ggml/src/ggml-openvino/openvino/utils.cpp
# ggml/src/ggml-openvino/openvino/utils.h
# ggml/src/ggml-openvino/utils.cpp
# ggml/src/ggml-openvino/utils.h
# ggml/src/ggml-sycl/common.hpp
# ggml/src/ggml-sycl/convert.cpp
# ggml/src/ggml-sycl/convert.hpp
# ggml/src/ggml-sycl/gemm.hpp
# ggml/src/ggml-sycl/ggml-sycl.cpp
# ggml/src/ggml-sycl/set_rows.cpp
# ggml/src/ggml-webgpu/ggml-webgpu-shader-lib.hpp
# ggml/src/ggml-webgpu/ggml-webgpu.cpp
# scripts/sync_vendor.py
# tests/CMakeLists.txt
# tests/test-chat.cpp
# tools/cli/cli.cpp
# tools/mtmd/CMakeLists.txt
# tools/server/CMakeLists.txt
2026-04-23 00:55:05 +08:00
Concedo
becf70d49b
fixed logspam for fit
2026-04-23 00:43:09 +08:00
Concedo
96ec87127a
updated colab, handle connection dropping during prompt processing
2026-04-21 21:46:13 +08:00
Concedo
4629b49afb
updated to handle changes for clip_is_mrope
2026-04-21 19:34:32 +08:00
Concedo
19a12bb080
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# CODEOWNERS
# common/CMakeLists.txt
# ggml/CMakeLists.txt
# ggml/src/ggml-webgpu/ggml-webgpu-shader-lib.hpp
# ggml/src/ggml-webgpu/ggml-webgpu.cpp
# ggml/src/ggml-webgpu/wgsl-shaders/common_decls.tmpl
# ggml/src/ggml-webgpu/wgsl-shaders/mul_mat_vec.wgsl
# scripts/sync-ggml.last
# tools/cli/cli.cpp
# tools/llama-bench/llama-bench.cpp
# tools/perplexity/perplexity.cpp
2026-04-21 18:53:03 +08:00
Concedo
1feba4e4ea
fixed koboldcpp.sh, fixed vision max/min when one param is missing, fixed processing count wrong, updated lite
2026-04-21 18:36:47 +08:00
Concedo
71b4107bb6
fixed terminal logs
2026-04-19 11:31:12 +08:00
Concedo
e5eab545f3
handle override jinja template
2026-04-19 00:30:28 +08:00
Concedo
17c754a5fc
improved reasoning budget
2026-04-18 17:19:09 +08:00
Concedo
0b37cb9a57
added preliminary support for reasoning budget
2026-04-18 11:56:33 +08:00
Concedo
9a38091207
support q5_1 kv
2026-04-17 17:06:15 +08:00
Concedo
cccb45a00a
summary outputs include processed amt
2026-04-17 14:22:51 +08:00
Concedo
64ce5fca15
better approach when SWA window exceeded, simply refill the window. this is not 100% correct but good enough for fastforward users. Disable FF or increase window if not good enough
2026-04-17 11:44:13 +08:00
Concedo
b5e317e015
SWA fix attempt 2
2026-04-17 00:33:45 +08:00
Concedo
ae292c496e
handle SWA conflicting with rewind, increased default SWA padding.
2026-04-16 17:00:26 +08:00
Concedo
0251c6dbde
added swa padding controls
2026-04-16 16:21:48 +08:00
Concedo
535df844dd
touchup for min/max tokens ui
2026-04-16 14:56:22 +08:00
Llama
c592bd01da
Pass img_min_params and img_max_params to ctx_clip_params ( #2133 )
...
* Pass img_min_params and img_max_params to ctx_clip_params
These values determine the minimum and maximum size (in
tokens) of vision embeddings. The default value of -1
uses a model-dependent default size, for example for
Gemma 4 the default is a 280 token embedding. For higher
quality results (at the cost of using more memory and
slower speed) you can increase the size of the embedding
to 1120 tokens.
* Change dict to mydict to match change to method
2026-04-16 12:27:06 +08:00