koboldcpp/tests
drluoto 5c53396b89
vulkan: raise the hoisted row-id limit for mul_mat_id from 256 to 512 experts (#28501)
* vulkan: raise the hoisted row-id limit for mul_mat_id to 512 experts

The expert-count shader (count_experts.comp) sizes its shared arrays
with BLOCK_SIZE, which is 256. Because of that, row-id hoisting is
switched off for any model with more than 256 experts, and every
mul_mat_id workgroup has to rescan the whole ids tensor on its own.
Qwen3.8-Flash-Next has 512 experts and was quietly running on that
slow path.

This change sizes the arrays with a separate MAX_EXPERTS constant (512),
clears them in a loop instead of one entry per thread, and raises the
matching limit on the host side.

On Strix Halo at batch 2048 the expert matmuls drop from 12.5 to 9.5 ms
(iq3_s) and from 14.0 to 7.5 ms (iq4_nl) per op, and prompt processing
gets about 19 % faster at 8k tokens. test-backend-ops MUL_MAT_ID passes
(891/891) with new 512-expert test cases.

Assisted-by: Claude Fable 5.1

* vulkan: raise the hoisted row-id limit for mul_mat_id to 1024 experts

Follow-up to review feedback: 1024 matches LLAMA_MAX_EXPERTS instead of
stopping at 512. The three shared arrays in count_experts.comp grow to
3 * 1024 * 4 = 12 KiB, which fits the 16 KiB that Vulkan guarantees for
maxComputeSharedMemorySize.

Adds mul_mat_id test cases at 1024 experts alongside the existing 512
ones. test-backend-ops MUL_MAT_ID passes on Vulkan (RADV, Strix Halo,
Radeon 8060S): 889/889.
2026-09-18 09:00:15 +02:00
..
fusion tests : add fusion baseline README and broaden fusion CI triggers (#28893) 2026-09-14 15:45:05 +03:00
peg-parser common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
snapshots tests : fix typo in test-quant-type-selection for nemotron 3 nano (#28835) 2026-09-13 18:50:46 +02:00
.gitignore metal : single-source fusion table + fusion debug rework (#28164) 2026-09-11 12:41:54 +03:00
CMakeLists.txt cmake : use PROJECT_SOURCE_DIR instead of CMAKE_SOURCE_DIR (#28771) 2026-09-15 05:26:09 +02:00
gguf-model-data.cpp ci : add [no release] keyword + fix sanitizer builds (#23728) 2026-05-26 19:05:48 +03:00
gguf-model-data.h tests : add unit test coverage for llama_tensor_get_type (#20112) 2026-04-02 22:53:58 +02:00
test-alloc.cpp ggml: allow passing alloc dependencies in graph_optimize (#27301) 2026-08-30 11:34:20 +08:00
test-arg-parser.cpp spec: Add benchmark-only synthetic speculative acceptance options (#27711) 2026-08-27 13:53:42 +03:00
test-autorelease.cpp tests : avoid building get-model.cpp many times (#26317) 2026-07-30 19:34:04 +03:00
test-backend-ops.cpp vulkan: raise the hoisted row-id limit for mul_mat_id from 256 to 512 experts (#28501) 2026-09-18 09:00:15 +02:00
test-backend-sampler.cpp tests : disable backend sampler hip multi output (#26878) 2026-08-11 07:21:32 +03:00
test-barrier.cpp Fix race conditions in threadpool when dealing with dynamic/frequent n_threads changes (#17748) 2025-12-10 12:32:23 -08:00
test-batch-alloc.cpp llama-batch: add unit test (#25471) 2026-07-10 11:04:31 +08:00
test-c.c ggml : remove kompute backend (#14501) 2025-07-03 07:48:32 +03:00
test-chat-analysis.cpp test: move tools/parser to tests (#27548) 2026-08-23 18:38:51 +02:00
test-chat-auto-parser.cpp test: move tools/parser to tests (#27548) 2026-08-23 18:38:51 +02:00
test-chat-peg-parser.cpp common : implement common_schema internal representation for JSON schemas (#28736) 2026-09-12 16:14:50 -05:00
test-chat-template.cpp test: move tools/parser to tests (#27548) 2026-08-23 18:38:51 +02:00
test-chat.cpp chat : improve parsing of complex types in qwen3-coder (#28742) 2026-09-12 19:08:52 -05:00
test-col2im-1d.cpp ggml : add GGML_OP_COL2IM_1D (#24206) 2026-06-09 12:01:37 +03:00
test-double-float.cpp ggml : minor naming changes (#8433) 2024-07-12 10:46:02 +03:00
test-export-graph-ops.cpp tests: export-graph-ops: exit gracefully when called w/o arguments (#25619) 2026-07-14 13:15:41 +03:00
test-fusion.cpp metal : single-source fusion table + fusion debug rework (#28164) 2026-09-11 12:41:54 +03:00
test-gbnf-validator.cpp cmake : do not include ./src as public for libllama (#13062) 2025-04-24 16:00:10 +03:00
test-gguf-model-data.cpp tests : add unit test coverage for llama_tensor_get_type (#20112) 2026-04-02 22:53:58 +02:00
test-gguf.cpp gguf : align the data section relative to the GGUF start, not the file (#28993) 2026-09-17 09:19:44 +02:00
test-grammar-integration.cpp common : implement common_schema internal representation for JSON schemas (#28736) 2026-09-12 16:14:50 -05:00
test-grammar-llguidance.cpp tool/ex/tests: consistently free ctx, then model (#18168) 2025-12-22 11:00:37 +01:00
test-grammar-parser.cpp grammar : degrade max repetition >= 2000 to unbounded (#26613) 2026-08-05 07:39:10 -05:00
test-jinja.cpp jinja : support dot property integer literals (#28817) 2026-09-12 23:49:53 +03:00
test-json-schema-to-grammar.cpp common : implement common_schema internal representation for JSON schemas (#28736) 2026-09-12 16:14:50 -05:00
test-json-schema.cpp common : implement common_schema internal representation for JSON schemas (#28736) 2026-09-12 16:14:50 -05:00
test-llama-archs.cpp model : add support for HrmTextForCausalLM (DFM Mimir 1B) (#27625) 2026-09-16 15:18:45 +02:00
test-llama-grammar.cpp common/grammar: fix grammar parsing issues to prevent stack overflow and hangs (#18604) 2026-03-21 18:43:35 +01:00
test-log.cpp common: Intentionally leak logger instance to fix hanging on Windows (#22273) 2026-04-29 10:58:43 +03:00
test-lora-conversion-inference.sh cli: new CLI experience (#17824) 2025-12-10 15:28:59 +01:00
test-model-load-cancel.cpp tests : avoid building get-model.cpp many times (#26317) 2026-07-30 19:34:04 +03:00
test-model-resolution.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
test-mtmd-c-api.c mtmd: add mtmd_tokenize_from_parts() (#28250) 2026-09-02 21:20:10 +02:00
test-mtmd-impl.cpp mtmd: add mtmd_tokenize_from_parts() (#28250) 2026-09-02 21:20:10 +02:00
test-opt.cpp tests : fix test-opt with GGML_BACKEND_DL (#15599) 2025-08-26 22:14:38 +02:00
test-peg-parser.cpp Autoparser - complete refactoring of parser architecture (#18675) 2026-03-06 21:01:00 +01:00
test-quant-type-selection.cpp tests : fix typo in test-quant-type-selection for nemotron 3 nano (#28835) 2026-09-13 18:50:46 +02:00
test-quantize-fns.cpp tests: extend test-quantize-fns to test nrc=2 (i8mm) kernels (#16234) 2026-09-12 02:19:37 +08:00
test-quantize-perf.cpp ci: run the x64 and arm ci on the github machines instead (#16183) 2025-09-25 08:06:06 +03:00
test-quantize-stats.cpp cmake : introduce semantic versioning (#26839) 2026-08-12 14:15:03 +02:00
test-reasoning-budget.cpp common : add support for multiple end sequences in the reasoning budget sampler (#25544) 2026-07-25 11:58:09 +02:00
test-recurrent-state-rollback.cpp model : support Kimi-K3 recurrent-state rollback (#28466) 2026-09-08 11:31:48 +08:00
test-rope.cpp ggml-cpu: templateify ggml_compute_forward_rope_f32 and _f16 (#16805) 2025-11-11 13:33:24 +02:00
test-rpc-multi-server.cpp rpc: avoid serializing buffers from other servers (#26500) 2026-08-30 20:26:16 +03:00
test-rpc-multi-server.sh rpc: avoid serializing buffers from other servers (#26500) 2026-08-30 20:26:16 +03:00
test-rset-release.cpp tests : avoid building get-model.cpp many times (#26317) 2026-07-30 19:34:04 +03:00
test-sampling.cpp llama : add missing headers (#28566) 2026-09-08 12:59:53 +02:00
test-save-load-state.cpp metal : single-source fusion table + fusion debug rework (#28164) 2026-09-11 12:41:54 +03:00
test-state-restore-fragmented.cpp common : only load backends when required (#22290) 2026-05-05 09:23:50 +02:00
test-thread-safety.cpp tests : synchronize contexts at end of test-thread-safety (#24935) 2026-06-25 09:22:51 +03:00
test-tokenizer-0.cpp tool/ex/tests: consistently free ctx, then model (#18168) 2025-12-22 11:00:37 +01:00
test-tokenizer-0.py requirements : update transformers to 5.5.1 (#21617) 2026-04-09 12:36:29 +02:00
test-tokenizer-0.sh model : add Jina Embeddings v5 Nano (partial EuroBERT) support (#19826) 2026-02-26 12:14:09 +01:00
test-tokenizer-1-bpe.cpp tool/ex/tests: consistently free ctx, then model (#18168) 2025-12-22 11:00:37 +01:00
test-tokenizer-1-spm.cpp tool/ex/tests: consistently free ctx, then model (#18168) 2025-12-22 11:00:37 +01:00
test-tokenizer-random.py requirements : update transformers to 5.5.1 (#21617) 2026-04-09 12:36:29 +02:00
test-tokenizers-repo.sh devops: add s390x & ppc64le CI (#15925) 2025-09-27 02:03:33 +08:00
test-unicode.cpp unicode : include '~' in collapsed symbol class (#26972) 2026-08-18 15:15:22 +02:00
testing.h chat: refactor handling supports_string_content / supports_typed_content (#27130) 2026-08-16 12:45:33 +02:00