koboldcpp/tests
anujj 3466812d1f
cuda: fuse MoE weighted expert reduction (#25952)
* cuda : fuse MoE weighted reduction (mul + view + add)

The MoE combine tail currently writes weighted expert outputs to
global memory before reducing them. That intermediate global-memory
traffic is the main cost. The production baseline generally runs two
physical fused kernels; this path runs one.

This change matches the full expert-weighting plus ordered-reduction
subgraph and replaces it with one weighted-reduction kernel.

Supported graphs:
- unscaled: experts * router_weights
- scaled:   (experts * expert_scale) * router_weights

k = 2..15 is handled by one runtime-k kernel.

Matching is structural: op sequence, shapes, strides, expert views,
and the left-to-right ADD chain. The fused kernel keeps that same
reduction order. Results are not claimed bit-identical; CUDA FP32
contraction can change rounding slightly.

Allocator integration uses add_alloc_dep from the graph-optimizer
API so experts, router weights, and optional expert scales stay live
until the fused destination is written. Memory ranges are rechecked
before the fused kernel runs.

Unrecognized or unsafe graphs are left alone and keep the existing
per-op path. Set GGML_CUDA_MOE_WEIGHTED_REDUCTION=0 to disable the
fusion.

test-backend-ops covers scaled/unscaled, aligned/unaligned, and
representative values across k=2..15, plus a k=16 case that must
stay on the per-op path.

* Pruned the test matrix from 15 to 6

* Addressed the aman and olivers review comments
2026-09-01 21:48:47 +02:00
..
peg-parser common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
snapshots tests : avoid building get-model.cpp many times (#26317) 2026-07-30 19:34:04 +03:00
.gitignore tests : add unit test coverage for llama_tensor_get_type (#20112) 2026-04-02 22:53:58 +02:00
CMakeLists.txt metal: enable Metal 4.0 tensor API on M5+/A19+ (#27461) 2026-09-01 12:02:42 +03:00
gguf-model-data.cpp ci : add [no release] keyword + fix sanitizer builds (#23728) 2026-05-26 19:05:48 +03:00
gguf-model-data.h tests : add unit test coverage for llama_tensor_get_type (#20112) 2026-04-02 22:53:58 +02:00
test-alloc.cpp ggml: allow passing alloc dependencies in graph_optimize (#27301) 2026-08-30 11:34:20 +08:00
test-arg-parser.cpp spec: Add benchmark-only synthetic speculative acceptance options (#27711) 2026-08-27 13:53:42 +03:00
test-autorelease.cpp tests : avoid building get-model.cpp many times (#26317) 2026-07-30 19:34:04 +03:00
test-backend-ops.cpp cuda: fuse MoE weighted expert reduction (#25952) 2026-09-01 21:48:47 +02:00
test-backend-sampler.cpp tests : disable backend sampler hip multi output (#26878) 2026-08-11 07:21:32 +03:00
test-barrier.cpp Fix race conditions in threadpool when dealing with dynamic/frequent n_threads changes (#17748) 2025-12-10 12:32:23 -08:00
test-batch-alloc.cpp llama-batch: add unit test (#25471) 2026-07-10 11:04:31 +08:00
test-c.c ggml : remove kompute backend (#14501) 2025-07-03 07:48:32 +03:00
test-chat-analysis.cpp test: move tools/parser to tests (#27548) 2026-08-23 18:38:51 +02:00
test-chat-auto-parser.cpp test: move tools/parser to tests (#27548) 2026-08-23 18:38:51 +02:00
test-chat-peg-parser.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
test-chat-template.cpp test: move tools/parser to tests (#27548) 2026-08-23 18:38:51 +02:00
test-chat.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
test-col2im-1d.cpp ggml : add GGML_OP_COL2IM_1D (#24206) 2026-06-09 12:01:37 +03:00
test-double-float.cpp ggml : minor naming changes (#8433) 2024-07-12 10:46:02 +03:00
test-export-graph-ops.cpp tests: export-graph-ops: exit gracefully when called w/o arguments (#25619) 2026-07-14 13:15:41 +03:00
test-gbnf-validator.cpp cmake : do not include ./src as public for libllama (#13062) 2025-04-24 16:00:10 +03:00
test-gguf-model-data.cpp tests : add unit test coverage for llama_tensor_get_type (#20112) 2026-04-02 22:53:58 +02:00
test-gguf.cpp gguf : harden loader against malformed tensor dims and metadata types (#25596) 2026-08-12 15:07:48 +03:00
test-grammar-integration.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
test-grammar-llguidance.cpp tool/ex/tests: consistently free ctx, then model (#18168) 2025-12-22 11:00:37 +01:00
test-grammar-parser.cpp grammar : degrade max repetition >= 2000 to unbounded (#26613) 2026-08-05 07:39:10 -05:00
test-jinja.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
test-json-schema-to-grammar.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
test-llama-archs.cpp qwen4exp: fix seq_cp, block position keying, mtmd input, cuda abort, add tests (#27941) 2026-09-01 13:22:04 +03:00
test-llama-grammar.cpp common/grammar: fix grammar parsing issues to prevent stack overflow and hangs (#18604) 2026-03-21 18:43:35 +01:00
test-log.cpp common: Intentionally leak logger instance to fix hanging on Windows (#22273) 2026-04-29 10:58:43 +03:00
test-lora-conversion-inference.sh cli: new CLI experience (#17824) 2025-12-10 15:28:59 +01:00
test-model-load-cancel.cpp tests : avoid building get-model.cpp many times (#26317) 2026-07-30 19:34:04 +03:00
test-model-resolution.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
test-mtmd-c-api.c mtmd: add chunk save/load function (#26645) 2026-08-06 19:46:40 +02:00
test-mtmd-impl.cpp mtmd: add mtmd_bitmap_set_mergeable (#27348) 2026-08-19 13:48:22 +02:00
test-opt.cpp tests : fix test-opt with GGML_BACKEND_DL (#15599) 2025-08-26 22:14:38 +02:00
test-peg-parser.cpp Autoparser - complete refactoring of parser architecture (#18675) 2026-03-06 21:01:00 +01:00
test-quant-type-selection.cpp tests : avoid building get-model.cpp many times (#26317) 2026-07-30 19:34:04 +03:00
test-quantize-fns.cpp Add Q2_0 quantization: type definition and CPU backend (#24448) 2026-07-07 12:05:47 -07:00
test-quantize-perf.cpp ci: run the x64 and arm ci on the github machines instead (#16183) 2025-09-25 08:06:06 +03:00
test-quantize-stats.cpp cmake : introduce semantic versioning (#26839) 2026-08-12 14:15:03 +02:00
test-reasoning-budget.cpp common : add support for multiple end sequences in the reasoning budget sampler (#25544) 2026-07-25 11:58:09 +02:00
test-recurrent-state-rollback.cpp DeepseekV4: fix rollback with multi-seq (#26756) 2026-08-23 13:57:49 +03:00
test-rope.cpp ggml-cpu: templateify ggml_compute_forward_rope_f32 and _f16 (#16805) 2025-11-11 13:33:24 +02:00
test-rpc-multi-server.cpp rpc: avoid serializing buffers from other servers (#26500) 2026-08-30 20:26:16 +03:00
test-rpc-multi-server.sh rpc: avoid serializing buffers from other servers (#26500) 2026-08-30 20:26:16 +03:00
test-rset-release.cpp tests : avoid building get-model.cpp many times (#26317) 2026-07-30 19:34:04 +03:00
test-sampling.cpp llama : support multi-output backend sampling (#25532) 2026-08-10 16:58:56 +03:00
test-save-load-state.cpp qwen4exp: fix seq_cp, block position keying, mtmd input, cuda abort, add tests (#27941) 2026-09-01 13:22:04 +03:00
test-state-restore-fragmented.cpp common : only load backends when required (#22290) 2026-05-05 09:23:50 +02:00
test-thread-safety.cpp tests : synchronize contexts at end of test-thread-safety (#24935) 2026-06-25 09:22:51 +03:00
test-tokenizer-0.cpp tool/ex/tests: consistently free ctx, then model (#18168) 2025-12-22 11:00:37 +01:00
test-tokenizer-0.py requirements : update transformers to 5.5.1 (#21617) 2026-04-09 12:36:29 +02:00
test-tokenizer-0.sh model : add Jina Embeddings v5 Nano (partial EuroBERT) support (#19826) 2026-02-26 12:14:09 +01:00
test-tokenizer-1-bpe.cpp tool/ex/tests: consistently free ctx, then model (#18168) 2025-12-22 11:00:37 +01:00
test-tokenizer-1-spm.cpp tool/ex/tests: consistently free ctx, then model (#18168) 2025-12-22 11:00:37 +01:00
test-tokenizer-random.py requirements : update transformers to 5.5.1 (#21617) 2026-04-09 12:36:29 +02:00
test-tokenizers-repo.sh devops: add s390x & ppc64le CI (#15925) 2025-09-27 02:03:33 +08:00
test-unicode.cpp unicode : include '~' in collapsed symbol class (#26972) 2026-08-18 15:15:22 +02:00
testing.h chat: refactor handling supports_string_content / supports_typed_content (#27130) 2026-08-16 12:45:33 +02:00