Commit graph

1652 commits

Author SHA1 Message Date
Concedo
bcc718d836 optimize keepalive handling 2026-09-13 21:50:46 +08:00
Concedo
ee68cf222c fix genkey image gen status leak 2026-09-08 18:15:18 +08:00
Concedo
e722ab2af2 keepalive interval 60s 2026-09-08 16:42:39 +08:00
Concedo
80d28478a1 hack to keep alive long image gen request connections 2026-09-08 16:33:12 +08:00
Wagner Bruna
78c1294245
fix --port argument when launching the gui (#2445) 2026-09-08 00:00:52 +08:00
Concedo
65512740ac improve sdui image recovery 2026-09-07 23:54:43 +08:00
Concedo
8e6ba2630b add ffn cpu flag 2026-09-06 18:08:03 +08:00
Concedo
4e4d88141c make help menu more accessible 2026-09-06 11:05:57 +08:00
Concedo
d65bed3752 hide cache slots by default if smartcache is off 2026-09-05 22:11:13 +08:00
Concedo
4685f6cea6 safetensors file analyze 2026-09-05 20:36:54 +08:00
yhz5613813
db5d5bfe5f
fix: handle invalid numeric API parameters (#2433) 2026-09-05 00:44:34 +08:00
Wagner Bruna
4d5779d1bb
sd: allow changing flow_shift through the api (#2429) 2026-09-04 10:17:13 +08:00
Concedo
62ce691117 bump version 2026-09-01 17:48:38 +08:00
Concedo
cec13b80a4 wip on true media references for h3 video gen 2026-08-30 14:49:24 +08:00
Concedo
f178512db3 bump default amt to gen 2026-08-28 18:36:26 +08:00
Concedo
cf113d8123 fix metal builds 2026-08-28 18:07:09 +08:00
Concedo
35de51e213 fixing ling template 2026-08-28 18:07:03 +08:00
Concedo
fe36aa1959 add directio support 2026-08-24 22:59:18 +08:00
Concedo
414fadaffc adjust cpu failsafe triggering issue 2026-08-21 21:48:18 +08:00
Concedo
da4540629a up version 2026-08-20 18:43:06 +08:00
Concedo
1fc6e11ab8 minor fix for erquint 2026-08-20 13:21:46 +08:00
Concedo
dfab7c1bf0 stop seq fix 2026-08-12 21:34:31 +08:00
Concedo
7fd4acc35c reasoning budget for muse glimmer 2026-08-11 15:43:23 +08:00
Concedo
b39ff27d6f muse glimmer jinja and tool calls working 2026-08-11 15:06:30 +08:00
Concedo
b3d0475aae muse glimmer templates 2026-08-10 21:44:56 +08:00
Concedo
59d92956f7 use python csv writer for benchmark 2026-08-09 09:28:18 +08:00
Concedo
8a16f96307 Merge branch 'upstream' into concedo_experimental
# Conflicts:
#	.github/workflows/build-apple.yml
#	.github/workflows/build-self-hosted.yml
#	.github/workflows/release.yml
#	SECURITY.md
#	build-xcframework.sh
#	ci/run.sh
#	docs/development/HOWTO-add-model.md
#	examples/model-conversion/scripts/causal/convert-model.sh
#	examples/model-conversion/scripts/embedding/convert-model.sh
#	scripts/sync_vendor.py
#	scripts/ui-assets.cmake
#	tests/test-arg-parser.cpp
#	tests/test-backend-sampler.cpp
#	tests/test-grammar-parser.cpp
#	tests/test-llama-archs.cpp
#	tests/test-sampling.cpp
#	tools/cli/README.md
#	tools/completion/README.md
#	tools/mtmd/CMakeLists.txt
#	tools/mtmd/mtmd.h
#	tools/mtmd/tests/test-deepseek-ocr.py
#	tools/server/README.md
#	tools/tts/CMakeLists.txt
#	tools/tts/convert_pt_to_hf.py
2026-08-07 20:46:56 +08:00
Concedo
0132829017 openai image edit endpoint 2026-08-07 14:40:54 +08:00
Concedo
927b997345 autoswap for oai images 2026-08-07 14:30:23 +08:00
Concedo
61bfce83de fix some issues with the preview image: Preview generation is disabled by default and only done when requested
Cleared stale generation state at job start/end.
Fixed the animated preview GIF buffer leak.
2026-08-07 00:19:39 +08:00
Wagner Bruna
e3cb5e9e44
sd: support for the /sdapi/v1/progress endpoint (#2316)
Co-authored-by: LostRuins Concedo <39025047+LostRuins@users.noreply.github.com>
2026-08-06 23:52:31 +08:00
Concedo
a8a8371229 increase max lora to 10 2026-08-06 22:31:52 +08:00
Concedo
348f7bf7f4 fix autoswap 2026-08-06 22:01:45 +08:00
Tai An
9fdd21de1b
fix(router): don't let an unmatched model field block autoswap (#2384) (#2387)
In autoswap mode, a POST to /v1/completions or /v1/chat/completions
carrying a `model` name that is not an entry in the admin dir set
`model_switch_pass = True` before checking the whitelist. No swap was
performed, but the flag suppressed the request-type dispatch below it,
so the text model was never loaded on demand.

The same requests without a `model` field, and every other model type
(stt/tts/embed/music/image), skip that branch entirely and load fine --
which is why only chat was affected, and why sending one model-less
request worked around it. It also recurs after --adminunloadtimeout
fires, since the "nomodel" state is recovered from by that same
dispatch.

Only set the flag on the path that actually issues the reload.
2026-08-06 21:58:44 +08:00
Julien BODIN
152e080b6a
Add Mistral [THINK]/[/THINK] thinking format (mistral3 arch) (#2380)
The reasoning budget derived from reasoning_effort never applied to Mistral
models. gpttype_adapter.cpp picks the think delimiters from a switch on the
model architecture, and mistral3 has no case, so it falls back to <think> /
</think>. Those are not vocabulary tokens for Ministral-3, so TokenizeString
returns more than one token each, the expected_start/end_tokens guard clears
all three vectors, and apply_reasoning_budget() returns at its first if.
The parameter is accepted, converted and passed down to the sampler, then
dropped on a size check, with nothing logged.

Adding the mistral3 case arms the budget. [THINK] and [/THINK] are single
vocabulary tokens (ids 34 and 35 on Ministral-3), so the size guard passes.

The thinkformats entry is a separate fix for a separate defect: without it the
thinking block was never split out, so it leaked into content with its [THINK]
marker still in it, instead of going to reasoning_content.

Measured on Ministral-3-14B-Reasoning-2512 (IQ4_XS, ctx 8192, --jinja), 5 real
prompts x 3 samples per cell, max_tokens 3000 (so a 750-token budget at "low"):

  reasoning_effort   thinking words before      thinking words after
  none               311 - 2314                 7 (the forced-close phrase)
  low                340 - 2255                 521 - 574

Forced closes: 0/15 before, 14/15 after at "low" and 15/15 at "none". Three
samples per cell because this model's variance at temperature 0.7 spans a
factor of 4 on an identical payload — a single sample per cell cannot tell an
effect from noise.

No regression on a non-reasoning mistral3 model: Ministral-3-8B-Instruct with
reasoning_effort "low" returns finish_reason "stop", a normal answer and zero
forced closes, since apply_reasoning_budget() bails out when the start marker
never appears.
2026-08-05 18:51:55 +08:00
Concedo
4dc9df05f6 increase max images 2026-08-05 14:30:32 +08:00
Concedo
bf9b9bd455 type check hardening 2026-08-02 10:48:23 +08:00
Concedo
9e64023c7b mcp media strip normalize 2026-08-02 10:46:26 +08:00
Tai An
4423b3af55
fix(api): strip MCP image base64 from tool results in the jinja path (#2374) (#2376)
When a tool/MCP result carries an image, the OpenAI-compatible chat
adapter's jinja code path left the base64 payload in the rendered
prompt as plain text (a single 1024x1024 jpeg bloated the context by
~120k tokens), while the legacy path already stripped it via
strip_mcpcontent_of_media.

- format_jinja now strips the base64 from tool-role string content
  before rendering, matching the legacy path; the image itself is
  still swept out and attached separately.
- sweep_media_from_messages now also recognizes MCP-style image
  content blocks (type == "image") inside a content list, so images
  delivered that way are attached instead of dropped.
2026-08-02 10:40:17 +08:00
Concedo
da8a5e4461 bump version 2026-08-02 10:33:54 +08:00
Concedo
57250b17dd untoggle no_host to try 2026-08-02 01:07:56 +08:00
Concedo
7a69646196 qwen3tts support languages 2026-07-27 22:45:04 +08:00
Concedo
304bd3d119 no host by default 2026-07-27 21:03:34 +08:00
Wagner Bruna
fd44ba2c61
sd: sync with master-795-87a0177 (#2338)
* sd: sync with master-782-b290693

* sd: sync with master-788-8a51eb9

* sd: sync with master-789-5114672

* sd: sync with master-795-87a0177

* sd: expose ref_image_args and make it trigger edit mode
2026-07-25 19:04:50 +08:00
Concedo
832ffa569e try to set the env vars first for cuda 2026-07-22 23:00:42 +08:00
Concedo
d097e369f5 fix for https://github.com/LostRuins/koboldcpp/discussions/2358
Revert "fix device enumeration issues"

This reverts commit 287f3cdb53d4c663d3108c6ad6e4e45624c45104. (+4 squashed commit)

Squashed commit:

[38bf57e59] Revert "try init cuda even earlier"

This reverts commit fa3d4dfd22284099d1484f3597ae60162d420e1b.

[fa3d4dfd2] try init cuda even earlier

[287f3cdb5] fix device enumeration issues

[02f9543f1] fix for https://github.com/LostRuins/koboldcpp/discussions/2358
2026-07-22 22:50:16 +08:00
Concedo
01f2fa51e4 patch from https://github.com/LostRuins/koboldcpp/pull/2352 2026-07-20 18:29:50 +08:00
DilanRG
a5560fc61d
fix(gui): preserve disabled VAE tiling in configs (#2346)
Treat the presence of sdtiledvae separately from its truthiness so the documented zero value survives a launcher config reload. Fixes #2065.

Co-authored-by: LK-Customs <lolikun@sunpraise-mc.com>
2026-07-19 17:37:37 +08:00
Concedo
768a2ca38c ipv4 and ipv6 total 40 threads for the http server 2026-07-18 18:55:46 +08:00
Tai An
8b85ba1a33
fix(api): don't iterate a string banned_tokens character by character (#2333) (#2336)
* fix(api): don't iterate a string banned_tokens character by character (#2333)

banned_tokens/banned_strings are consumed as a list of substrings. When a
caller supplies a bare string instead, `banned_tokens[:ban_token_max]`
slices it and `for tok in banned_tokens` walks it one character at a time,
so every letter in the value becomes its own banned substring. Banning
common letters like "e" makes generation collapse into garbage, with no
error to point at the cause.

Coerce a string value into a single-element list, matching how the OpenAI
`stop` parameter is already normalized in transform_genparams.

* fix(api): parse a JSON-array string in banned_tokens instead of inerting it

Per review: coercing the string form to a single-element list left the
reported gendefaults case a seemingly-effective no-op. Parse a JSON array
supplied as a string so the reported config bans "exclude" as intended,
and keep the bare-string form ("exclude") working as a single substring ban.

---------

Co-authored-by: Anai-Guo <antai12232931@anaiguo.com>
2026-07-18 10:46:02 +08:00