Commit graph

15773 commits

Author SHA1 Message Date
Concedo
26b9e94272 add ask_user tool 2026-09-22 20:58:26 +08:00
Concedo
d8975ee2dd agent interrupt 2026-09-22 20:38:45 +08:00
Concedo
d6bb510251 rename --sdt5xxl to --sdllm flag 2026-09-22 20:15:14 +08:00
Concedo
7dd09f714b try to improve vision prompting 2026-09-22 18:13:08 +08:00
Concedo
60618ebeb0 agent auto approval 2026-09-22 16:55:08 +08:00
Concedo
1cded239f9 set api key and model 2026-09-22 16:03:03 +08:00
Concedo
f694cd7083 view image tool 2026-09-22 15:51:30 +08:00
Concedo
8e6d4df92f bump agent requirements (+1 squashed commits)
Squashed commits:

[60291a266] bump agent requirements
2026-09-21 23:12:45 +08:00
Wagner Bruna
b3887208fc
sd: round info floating point parameters (#2484) 2026-09-21 21:31:49 +08:00
Concedo
018decfd69 fallback terminals for linux 2026-09-21 21:27:29 +08:00
Concedo
c49fbd7003 cleanup and remove sdcpp_logger_adapter 2026-09-21 21:00:24 +08:00
Wagner Bruna
3f490d1133
sd: fix: inhibit ggml log changes from sdcpp code (#2483)
* sd: apply logging changes from master-885-b8248a8

* sd: fix: inhibit ggml log changes from sdcpp code

* sd: do not set ggml logging callback, leave the default

* sd: do not set ggml logging callback

* sd: fix kcpp log level test
2026-09-21 20:56:15 +08:00
Concedo
a223943815 Merge branch 'upstream' into concedo_experimental
# Conflicts:
#	docs/ops.md
#	docs/ops/Hexagon.csv
#	examples/parallel/parallel.cpp
#	ggml/src/ggml-hexagon/ggml-hexagon.cpp
#	ggml/src/ggml-hexagon/htp/act-ops.c
#	ggml/src/ggml-hexagon/htp/argsort-ops.c
#	ggml/src/ggml-hexagon/htp/get-rows-ops.c
#	ggml/src/ggml-hexagon/htp/htp-ctx.h
#	ggml/src/ggml-hexagon/htp/htp-ops.h
#	ggml/src/ggml-hexagon/htp/htp-tensor.h
#	ggml/src/ggml-hexagon/htp/main.c
#	ggml/src/ggml-webgpu/ggml-webgpu-shader-lib.hpp
#	ggml/src/ggml-webgpu/ggml-webgpu.cpp
#	ggml/src/ggml-webgpu/wgsl-shaders/gated_delta_net.wgsl
#	tests/fusion/MTL.csv
#	tests/peg-parser/test-unicode.cpp
#	tests/test-backend-ops.cpp
#	tests/test-chat-peg-parser.cpp
#	tests/test-chat.cpp
#	tests/test-json-schema-to-grammar.cpp
#	tests/test-llama-archs.cpp
#	tools/ui/tests/stories/a11y/ChatScreenForm.a11y.stories.svelte
2026-09-21 20:54:47 +08:00
Concedo
b74d33d042 fix model swapping in routermode when loaded from a config 2026-09-21 19:13:18 +08:00
Concedo
f52f792937 fixed glob 2026-09-21 19:04:26 +08:00
Concedo
f3d8ff11d9 add session compaction 2026-09-21 19:04:18 +08:00
Yangyu Chen
68d9053afd
cuda : tune MMVQ to MMQ crossover for SM70 (Volta) (#28912)
Some checks failed
Python Type-Check / python type-check (push) Has been cancelled
Update Operations Documentation / update-ops-docs (push) Has been cancelled
* tune MMVQ to MMQ crossover for SM70 (Volta)

Signed-off-by: Yangyu Chen <cyy@cyyself.name>

* Apply suggestion from @JohannesGaessler

* Apply suggestion from @JohannesGaessler

* Apply suggestion from @JohannesGaessler

---------

Signed-off-by: Yangyu Chen <cyy@cyyself.name>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
2026-09-21 10:45:31 +03:00
Niklas Wenzel
8aa161b54a
metal : fix deprecation warnings from macOS 27 SDK (#29136) 2026-09-21 10:44:40 +03:00
Concedo
7f37b91331 improve agent shell detection 2026-09-21 15:41:02 +08:00
Masashi Yoshimura
932a68e068
webgpu : add fused gdn + cpy (#28976) 2026-09-21 10:39:30 +03:00
David Friehs
62668d6b26
convert: enable --fuse-qkv for muse-glimmer (#29203) 2026-09-21 10:38:28 +03:00
Concedo
bb6e2fb587 remove list_directory tool (+1 squashed commits)
Squashed commits:

[2c12cc940] remove list_directory tool
2026-09-21 15:16:51 +08:00
Concedo
71fafd7f61 bump agent minimum defaults (+1 squashed commits)
Squashed commits:

[bb5989996] bump agent minimum defaults
2026-09-21 14:58:47 +08:00
Concedo
b66d3eb717 allow wt if available 2026-09-21 14:16:54 +08:00
Concedo
003fa7e2f3 edit tool prompt, default temp 0.5 (+1 squashed commits)
Squashed commits:

[2c81e9a5d] edit prompt in agent tool
2026-09-21 13:53:23 +08:00
Concedo
9752998ee0 ensure some agent launch params 2026-09-21 11:13:15 +08:00
Johannes Gäßler
ce8caa6e60
CUDA: tune FA for Gemma 4 on Ampere or newer (#29152) 2026-09-20 22:20:12 +02:00
Concedo
f9880f4669 kcpp agent mcp tool calling works 2026-09-21 02:34:07 +08:00
Concedo
d2a2223e34 kcpp agent can set max tokens 2026-09-21 02:10:01 +08:00
Concedo
5c9f22972a refine kcpp agent (+1 squashed commits)
Squashed commits:

[01007173b] refine kcpp agent
2026-09-21 02:01:15 +08:00
Concedo
34b81e2ada bump default ctx to 16k 2026-09-21 01:32:59 +08:00
Concedo
635f05b657 wip kcpp agent 2026-09-21 01:28:40 +08:00
Concedo
a1b8a76cec added a simple kobold_agent.py 2026-09-21 00:11:24 +08:00
Concedo
0e49e296a6 updated lite (+1 squashed commits)
Squashed commits:

[5f32d56b4] updated lite
2026-09-20 23:18:41 +08:00
Georgi Gerganov
a894dae939
metal : support arbitrary hc in dsv4_hc_pre (#29169)
the dsv4_hc_pre kernels hardcoded hc = 4 via a constexpr used with
simd_shuffle, so the op was rejected by supports_op for any other hc
and fell back to CPU. Kimi-K3 uses dsv4_hc_pre with hc equal to the
number of banked checkpoints in the cross-layer residual stack, which
grows with the layer index.

pass n_hc as a function constant (FC_DSV4_HC) with per-n_hc pipeline
variants, and loop over it in both pre kernels with direct loads

add test-backend-ops cases for hc = 1, 2, 3, 5, 8 and 65, gated and
not gated

Assisted-by: pi:llama.cpp/Qwen3.8-27B
2026-09-20 17:52:20 +03:00
Concedo
e5e51240b5 limit max preview size 2026-09-20 20:57:30 +08:00
Aldehir Rojas
3d82ef62d4
common/peg : handle invalid utf-8 sequences in the AST (#29161)
* common/peg : handle invalid utf-8 sequences in the AST

* cont : return maximal subpart per Unicode recommendations

* cont : remove strict argument
2026-09-20 06:51:40 -05:00
Concedo
ef51111b48 improved autoswap, added --autoswapthreshold which allows autoswap to trigger only if target swap exceeds threshold 2026-09-20 19:33:55 +08:00
Concedo
f44706a190 updated readme (+1 squashed commits)
Squashed commits:

[90fdd39c7] updated readme
2026-09-20 18:37:35 +08:00
Aman Gupta
3cf03257f2
CUDA: enable sparse fa for qwen4 (#28770) 2026-09-20 16:08:11 +08:00
Aleksander Grygier
b23efaa2ef
ui: Fix mobile breakpoint + content overflow issues (#29108)
* ui : let the chat column shrink below its content width

The chat column is a flex item, so its automatic minimum size kept it as wide as the widest row inside it. Message rows cap at max-w-3xl plus padding, so a narrower window pushed a page-level horizontal scrollbar.

Set min-w-0 on the column so the inner scroll containers take over.

Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash

* ui : wrap markdown tables in a scroll container

Markdown tables render as a bare <table>, which keeps its content-driven minimum width and can stretch the chat column past the window. The table-wrapper CSS already existed, but nothing produced the wrapper.

Add a rehype plugin that wraps each table in div.table-wrapper, following the existing enhance-* plugins.

Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash

* ui : scroll long inline content inside markdown blocks

Long unbreakable content (inline code, paths, hashes) widened the message row and spilled over the neighbour elements. Give each markdown block a horizontal scroll container, and the content root one as well, since the trailing block renders with display: contents and has no box of its own.

Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash

* ui : use exact transition properties for markdown images

transition: all repainted every property and 300ms felt sluggish. Name transform and box-shadow at 200ms ease-out, and gate the hover scale behind (hover: hover) and (pointer: fine) so touch taps do not trigger it.

Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash

* ui : fit wide image attachments to the message width

Attachment thumbnails used a fixed height with w-auto, so a wide image kept its aspect-driven width and, being flex-shrink-0 in a right-aligned bubble, overflowed to the left of the message row.

Cap the thumbnail with max-height and max-width instead of a fixed height so it scales down proportionally, and let it shrink outside the single-row carousel.

Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash

* ui : keep long tool call titles inside the message row

A tool title could not shrink below its content, so a long path escaped the message row. Let the title span shrink and scroll, and for the file tools put the value on its own line only when it does not fit, with the value as the only scroll container.

Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash

* ui : render get info as a collapsible block with a table

get_info rendered its own always-open row with the values trailing the label. Use the shared ToolCallBlock chrome so it collapses like the other tools, and list os and cwd as table rows with the key as a row header.

The error and pending states now show inside the body, including the plain-string errors the server tools path produces.

Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash

* test : pin the server mode in the add menu a11y story

The story asserts the add menu's first enabled item is the reasoning submenu, which is mounted only outside router mode. The vitest dev server proxies /props to whichever server is running, so the assertion depended on the machine's server mode and failed whenever a router was up.

Pin the mode in the story, including props.role so a re-detection cannot flip it back.

Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash

* ui: wrap long markdown tokens instead of scrolling every block

Making each markdown block and the content root a horizontal scroll
container turns any hover transform into a scrollbar: the blockquote
translate and the image zoom overflow their block and flash a scrollbar
under it. Each block also becomes a block formatting context, so the
paragraph margins stop collapsing across blocks and the spacing doubles.

Drop both overflow-x rules and let long unbreakable tokens wrap with
overflow-wrap: break-word on the content root. break-word leaves the
min-content width untouched, so wide tables and code blocks keep
scrolling inside their own containers.

---------

Co-authored-by: Pascal <admin@serveurperso.com>
2026-09-20 08:59:43 +03:00
Andrei
4260903678
fix(mamba) : make time-step projection input contiguous (#28832)
* mamba : make time-step projection input contiguous

Assisted-by: ChatGPT

* mamba : skip contiguous copy after normalization

Assisted-by: ChatGPT
2026-09-20 07:58:14 +03:00
bri-prism
9a9f939b80
metal: add F16 input to the FWHT (#29094)
* metal: add F16 input to the FWHT

The Metal FWHT kernel accepts F32 input only. This change makes the source
type a template parameter, so the kernel reads an F16 source directly instead
of requiring a converted copy. The F32 instantiations are unchanged.

The pipeline name now carries the source type, and supports_op accepts an F16
src1 for the Hadamard hint at the four sizes the kernels cover. Every other
F16 src1 path still goes through ggml_metal_supports_mul_mat_op.

These are the test cases mentioned in #27779.

test-backend-ops on M5 Pro: MUL_MAT_HADAMARD 16/16, MUL_MAT 1265/1265.

* metal: ask the same FWHT question in supports_op and the dispatch

supports_op admitted an F16 src1 on the type, the hint and the width alone, but the
dispatch also requires src1 and dst to be contiguous and the same shape. A Hadamard
hinted MUL_MAT that passed the first and failed the second reached the generic path,
which has no F32 src0 by F16 src1 kernel, and aborted on a nil pipeline:

  kernel not found in any metal library: base = 'kernel_mul_mv_f32_f16_4'
  ggml_metal_encoder_set_pipeline: nil Metal pipeline

ggml_metal_use_fwht now holds the whole condition and both callers use it, so they
cannot drift apart again. The added test case has src1 and dst of different shapes,
which aborted before this change and is declined by the Metal backend after it.

* metal: branchless butterfly select in the FWHT simdgroup kernel

Review suggestion. Replaces the ternary in the shuffle stages with
val2 - val + 2*((lane & i) == 0)*val, which is the same value without the
select.

Measured on M5 Pro, interleaved A/B, five rounds, first discarded, on a
Hadamard matmul with block 512 and 65536 rows so the kernel rather than the
launch dominates: 1324.6 us before, 1285.0 us after, a 3.0% gain, and faster
in every round. At the shapes already in the perf suite the op runs 1.6 to
3.9 us against a 1.6 us launch floor, so the difference is not visible there.

FOR_UNROLL on the same loops was also measured and made no difference, the
delta changing sign between rounds, so it is not included.

* metal: move the FWHT dispatch predicates to ggml-metal-common

Review feedback. ggml_metal_use_fwht and ggml_metal_fwht_supported_size were
static inline in ggml-metal-device.h. They now follow the
ggml_metal_op_mul_mat_use_mm pattern: declared in ggml-metal-common.h and
implemented in ggml-metal-common.cpp, which is already the home for helpers
shared between supports_op and the op dispatch. The predicate is named
ggml_metal_op_mul_mat_use_fwht to sit alongside the _use_mm pair it parallels.

This also fixes the macos-latest-arm64 build. The header needed ggml-impl.h
for ggml_get_op_params_i32, but ggml-metal-device.h is reached from
tools/tuning through ggml-metal-tuning.h, and that target does not have
ggml/src on its include path. ggml-metal-common.cpp already includes
ggml-impl.h, so the accessor is used normally there and the header goes back
to needing nothing extra.

* metal: keep the FWHT size check internal and group the dispatch helpers

Applies the patch from the review. ggml_metal_fwht_supported_size becomes
static in ggml-metal-common.cpp since nothing outside it needs the size list,
which also drops stdint.h from the header again, and
ggml_metal_op_mul_mat_use_fwht joins the existing _use_mm declarations under
their shared comment instead of carrying its own block.

* tests: drop the mismatched-shape Hadamard case

I added a case with m != k to cover an abort, but the hint is a promise that
src0 is a Hadamard matrix, so src0 is square and dst has the same shape as
src1. Every other case in the suite holds to that. The case was not a valid
op, and on CPU it compared the FWHT against a real matmul of a non-square
src0, which cannot agree.

The supports_op and dispatch conditions still come from one predicate, which
is what keeps them from disagreeing on contiguity.
2026-09-20 07:57:30 +03:00
Concedo
8a9833dad1 updated lite 2026-09-20 11:37:07 +08:00
Concedo
1859dbed49 preserved tokens 2026-09-20 10:35:37 +08:00
Aldehir Rojas
f072b10371
chat : fix gemma4 required tool grammar (#29115) 2026-09-19 18:59:43 -05:00
Toby
59657a613a
chat : add dedicated Ling 3.0 (Bailing V3) parser (#28682)
* chat: add dedicated Ling 3.0 (Bailing V3) parser

Ling 3.0 Flash templates pre-open the think block in the generation
prompt, so the model never emits an opening <think>, and a tool call can
arrive before any </think>. The generated autoparser terminated reasoning
only at the close tag, which classified such tool calls entirely as
reasoning_content: clients received content="" with no tool_calls and
agent loops died as reasoning-only turns.

Adds a specialized parser that terminates reasoning at the think close
tag or at a <tool_call> start, mirroring the hand-written Qwen3-Coder and
Kimi K3 parsers and the reference vLLM/SGLang Ling3 parser (which treats
<tool_call> as an implicit reasoning terminator). Detection is gated on
the <role>...</role> section markers, unique to this family among the
tagged-argument templates.

Adds the Ling 3.0 Flash chat template and tests covering the
unclosed-think tool call (full parse and streaming), healthy closed-think
paths, trailing prose, parallel calls, marker-like strings in argument
values, string-union and non-string argument types, and
reasoning_format=none.

Assisted-by: Kimi Code

* tests : move Ling 3.0 test

---------

Co-authored-by: aetherbird <aetherbird@users.noreply.github.com>
Co-authored-by: Alde Rojas <hello@alde.dev>
2026-09-19 18:35:44 -05:00
Aparna M P
e613ef2c81
hexagon: enable I32 GET_ROWS (#29116) 2026-09-19 09:48:31 -07:00
Aparna M P
851cb34f21
hexagon: add support for GEGLU_QUICK (#29114) 2026-09-19 09:48:07 -07:00
Aparna M P
7d4b92bb9b
hexagon: enable support for TOP_K op (#29113)
* hexagon: enable support for TOP_K op

* hex-topk: thread single-row TOP_K, raise VTCM-based size cap

* hex-topk: fix TOP_K mdev row partitioning

* hex-topk: optimize TOP_K large-row selection

* hexagon: clean up comment formatting

* hex-docs: update TOP_K support listings
2026-09-19 09:16:03 -07:00