koboldcpp/tests
itsnotoger 2d8d612e4c
kv-cache : optimize restoring non-contiguous cells (#27991)
* kv cache : batch state restore scatter reads per contiguous run

When restoring state into non-contiguous destination cells (e.g. a
prompt-cache snapshot into a fragmented ring), state_read_data issued
one small copy per KV cell - ~1.4M copies of a few KiB each for a
40k+ token restore, taking 25-63 s on the CUDA backend.

The snapshot stores cell rows in cell order, so a maximal run of
consecutive destination indices maps to one contiguous block and can
be restored with a single copy. Precompute the runs once and use them
in all three scatter loops (K, V, transposed V). Byte-identical.

The on-device reader copies with a byte cursor when the read and
write chunking differs, so the batched reads are safe for it as well.
Batching makes equal tensor counts with a different split reachable
(save ranges [2,1] vs restore runs [1,2]); the next commit teaches the
reader's 1:1 path to fall back to the byte cursor in that case.

Verified in a production setup: 1,363,616 copies / 25-63 s -> 224
copies / 221-424 ms for the same restores (42,603 cells, 4 runs).

Assisted-by: Claude Code (unsloth/qwen3.8-27b)

* context : fall back to the byte cursor when read and write chunking differ

the on-device reader copies saved state back with a 1:1 copy by tensor
index whenever the write and read sides recorded the same number of
tensors, guarded by a per-tensor size assert.

equal tensor counts do not imply equal chunking: a state restore may
batch its reads per contiguous run of destination cells while the save
used per-range reads, so both sides can record two tensors that split
the same data differently, and the assert aborts in all builds.

compare the per-tensor sizes and only take the 1:1 path when the
chunking actually matches, otherwise fall through to the existing
byte-cursor copy. both sides enumerate the same logical data in the
same order, so the cursor copy is well-defined across tensor
boundaries.

Assisted-by: Claude Code (unsloth/qwen3.8-27b)

* tests : cover state restore scatter reads on host and on-device paths

decode the same prefix on two sequences, interleaving the seq 0 cells
between the seq 1 cells, so the seq 1 cells are isolated from each
other in the kv cache (three cells, two saved ranges). save the seq 1
state, free the interleaved seq 0 cells, and restore: the destination
is then non-contiguous (two runs), and the restore-side chunking has
the same tensor count as the save-side with a different split, so the
scatter path is batched per contiguous run and the on-device reader's
byte-cursor fallback is exercised.

the restored state is saved again on the host and compared byte for
byte with the first save: the blob is serialized in sequence cell
order, so the two saves are identical if and only if the scatter
restore wrote exactly the same KV content. this documents the
byte-identical guarantee of the run-batched scatter reads.

one test per io backend: the host (CPU) path and the on-device path.

Assisted-by: Claude Code (unsloth/qwen3.8-27b)
2026-08-31 19:49:58 +03:00
..
peg-parser common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
snapshots tests : avoid building get-model.cpp many times (#26317) 2026-07-30 19:34:04 +03:00
.gitignore tests : add unit test coverage for llama_tensor_get_type (#20112) 2026-04-02 22:53:58 +02:00
CMakeLists.txt tests : run test-save-load-state across all architectures (#27755) 2026-08-28 09:45:19 +03:00
gguf-model-data.cpp ci : add [no release] keyword + fix sanitizer builds (#23728) 2026-05-26 19:05:48 +03:00
gguf-model-data.h tests : add unit test coverage for llama_tensor_get_type (#20112) 2026-04-02 22:53:58 +02:00
test-alloc.cpp ggml: allow passing alloc dependencies in graph_optimize (#27301) 2026-08-30 11:34:20 +08:00
test-arg-parser.cpp spec: Add benchmark-only synthetic speculative acceptance options (#27711) 2026-08-27 13:53:42 +03:00
test-autorelease.cpp tests : avoid building get-model.cpp many times (#26317) 2026-07-30 19:34:04 +03:00
test-backend-ops.cpp CUDA: extend MOE fusion to specdec, earlier MOE glu fusion and topk-router fusion were restricted to 1 token (#27621) 2026-08-31 19:22:28 +08:00
test-backend-sampler.cpp tests : disable backend sampler hip multi output (#26878) 2026-08-11 07:21:32 +03:00
test-barrier.cpp Fix race conditions in threadpool when dealing with dynamic/frequent n_threads changes (#17748) 2025-12-10 12:32:23 -08:00
test-batch-alloc.cpp llama-batch: add unit test (#25471) 2026-07-10 11:04:31 +08:00
test-c.c ggml : remove kompute backend (#14501) 2025-07-03 07:48:32 +03:00
test-chat-analysis.cpp test: move tools/parser to tests (#27548) 2026-08-23 18:38:51 +02:00
test-chat-auto-parser.cpp test: move tools/parser to tests (#27548) 2026-08-23 18:38:51 +02:00
test-chat-peg-parser.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
test-chat-template.cpp test: move tools/parser to tests (#27548) 2026-08-23 18:38:51 +02:00
test-chat.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
test-col2im-1d.cpp ggml : add GGML_OP_COL2IM_1D (#24206) 2026-06-09 12:01:37 +03:00
test-double-float.cpp ggml : minor naming changes (#8433) 2024-07-12 10:46:02 +03:00
test-export-graph-ops.cpp tests: export-graph-ops: exit gracefully when called w/o arguments (#25619) 2026-07-14 13:15:41 +03:00
test-gbnf-validator.cpp cmake : do not include ./src as public for libllama (#13062) 2025-04-24 16:00:10 +03:00
test-gguf-model-data.cpp tests : add unit test coverage for llama_tensor_get_type (#20112) 2026-04-02 22:53:58 +02:00
test-gguf.cpp gguf : harden loader against malformed tensor dims and metadata types (#25596) 2026-08-12 15:07:48 +03:00
test-grammar-integration.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
test-grammar-llguidance.cpp tool/ex/tests: consistently free ctx, then model (#18168) 2025-12-22 11:00:37 +01:00
test-grammar-parser.cpp grammar : degrade max repetition >= 2000 to unbounded (#26613) 2026-08-05 07:39:10 -05:00
test-jinja.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
test-json-schema-to-grammar.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
test-llama-archs.cpp tests : run test-save-load-state across all architectures (#27755) 2026-08-28 09:45:19 +03:00
test-llama-grammar.cpp common/grammar: fix grammar parsing issues to prevent stack overflow and hangs (#18604) 2026-03-21 18:43:35 +01:00
test-log.cpp common: Intentionally leak logger instance to fix hanging on Windows (#22273) 2026-04-29 10:58:43 +03:00
test-lora-conversion-inference.sh cli: new CLI experience (#17824) 2025-12-10 15:28:59 +01:00
test-model-load-cancel.cpp tests : avoid building get-model.cpp many times (#26317) 2026-07-30 19:34:04 +03:00
test-model-resolution.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
test-mtmd-c-api.c mtmd: add chunk save/load function (#26645) 2026-08-06 19:46:40 +02:00
test-mtmd-impl.cpp mtmd: add mtmd_bitmap_set_mergeable (#27348) 2026-08-19 13:48:22 +02:00
test-opt.cpp tests : fix test-opt with GGML_BACKEND_DL (#15599) 2025-08-26 22:14:38 +02:00
test-peg-parser.cpp Autoparser - complete refactoring of parser architecture (#18675) 2026-03-06 21:01:00 +01:00
test-quant-type-selection.cpp tests : avoid building get-model.cpp many times (#26317) 2026-07-30 19:34:04 +03:00
test-quantize-fns.cpp Add Q2_0 quantization: type definition and CPU backend (#24448) 2026-07-07 12:05:47 -07:00
test-quantize-perf.cpp ci: run the x64 and arm ci on the github machines instead (#16183) 2025-09-25 08:06:06 +03:00
test-quantize-stats.cpp cmake : introduce semantic versioning (#26839) 2026-08-12 14:15:03 +02:00
test-reasoning-budget.cpp common : add support for multiple end sequences in the reasoning budget sampler (#25544) 2026-07-25 11:58:09 +02:00
test-recurrent-state-rollback.cpp DeepseekV4: fix rollback with multi-seq (#26756) 2026-08-23 13:57:49 +03:00
test-rope.cpp ggml-cpu: templateify ggml_compute_forward_rope_f32 and _f16 (#16805) 2025-11-11 13:33:24 +02:00
test-rpc-multi-server.cpp rpc: avoid serializing buffers from other servers (#26500) 2026-08-30 20:26:16 +03:00
test-rpc-multi-server.sh rpc: avoid serializing buffers from other servers (#26500) 2026-08-30 20:26:16 +03:00
test-rset-release.cpp tests : avoid building get-model.cpp many times (#26317) 2026-07-30 19:34:04 +03:00
test-sampling.cpp llama : support multi-output backend sampling (#25532) 2026-08-10 16:58:56 +03:00
test-save-load-state.cpp kv-cache : optimize restoring non-contiguous cells (#27991) 2026-08-31 19:49:58 +03:00
test-state-restore-fragmented.cpp common : only load backends when required (#22290) 2026-05-05 09:23:50 +02:00
test-thread-safety.cpp tests : synchronize contexts at end of test-thread-safety (#24935) 2026-06-25 09:22:51 +03:00
test-tokenizer-0.cpp tool/ex/tests: consistently free ctx, then model (#18168) 2025-12-22 11:00:37 +01:00
test-tokenizer-0.py requirements : update transformers to 5.5.1 (#21617) 2026-04-09 12:36:29 +02:00
test-tokenizer-0.sh model : add Jina Embeddings v5 Nano (partial EuroBERT) support (#19826) 2026-02-26 12:14:09 +01:00
test-tokenizer-1-bpe.cpp tool/ex/tests: consistently free ctx, then model (#18168) 2025-12-22 11:00:37 +01:00
test-tokenizer-1-spm.cpp tool/ex/tests: consistently free ctx, then model (#18168) 2025-12-22 11:00:37 +01:00
test-tokenizer-random.py requirements : update transformers to 5.5.1 (#21617) 2026-04-09 12:36:29 +02:00
test-tokenizers-repo.sh devops: add s390x & ppc64le CI (#15925) 2025-09-27 02:03:33 +08:00
test-unicode.cpp unicode : include '~' in collapsed symbol class (#26972) 2026-08-18 15:15:22 +02:00
testing.h chat: refactor handling supports_string_content / supports_typed_content (#27130) 2026-08-16 12:45:33 +02:00