koboldcpp/tools
Xuan-Son Nguyen 82d6bb284d
server: refactor subproc handling (#28555)
* server: refactor subproc handling

* fix Windows build

* download: keep concurrent downloads of one blob apart

Every process writes the same path + .downloadInProgress, so a second
download of the same blob finds that file, takes it for its own partial
transfer and asks for the bytes after it, which produces a corrupt
result. The in-progress file now carries the pid of the process writing
it.

std::rename also replaces an existing destination on POSIX but fails on
Windows, so a download whose blob appeared in the meantime is dropped
after every retry and an etag rewrite silently keeps the old value.
std::filesystem::rename has the POSIX behaviour everywhere, and the
error now carries the reason reported by the system.

* Revert "download: keep concurrent downloads of one blob apart"

This reverts commit 917b83f149c625527f872fb2cf41289358fa5371.

* tests: serialize the router tests that download the same model

Parallel workers share one cache, so the two tests fetch the same blob
into the same in-progress file and race to rename it. They now take a
file lock around the download, like the session fixture does for the
preset models.

* Revert "tests: serialize the router tests that download the same model"

This reverts commit c368a4a98c677938ca87002edb6186ff2c02fd83.

---------

Co-authored-by: Pascal <admin@serveurperso.com>
2026-09-12 00:53:07 +02:00
..
batched-bench
cli args: officially deprecate --mmap|mlock|dio (#28334) 2026-09-09 18:36:27 +08:00
completion args: officially deprecate --mmap|mlock|dio (#28334) 2026-09-09 18:36:27 +08:00
cvector-generator
export-lora
fit-params fit: also take into account n_streams (#27496) 2026-08-22 16:16:06 +02:00
gguf-split
imatrix
llama-bench args: officially deprecate --mmap|mlock|dio (#28334) 2026-09-09 18:36:27 +08:00
mtmd cmake : add PCH and unity build to improve build times (#28091) 2026-09-11 13:01:29 +02:00
perplexity quant : Optimise memory usage by evicting weights after processing each layer (#22877) 2026-08-18 16:22:32 +02:00
quantize quantize: cap working memory size to avoid loading big tensors onto RAM (#27795) 2026-08-27 18:31:13 +02:00
results
rpc rpc: avoid serializing buffers from other servers (#26500) 2026-08-30 20:26:16 +03:00
server server: refactor subproc handling (#28555) 2026-09-12 00:53:07 +02:00
tokenize
tts args: add --video-* CLI arguments (#24318) 2026-08-27 12:11:12 +02:00
tuning ggml : update ggml_prec specification (#26675) 2026-09-08 09:06:24 +03:00
ui ui: Improve Chat Messages rendering performance (#28460) 2026-09-06 10:52:40 +02:00
CMakeLists.txt metal : per-device tuned (Q, NE) for flash-attn vec (#26570) 2026-08-24 19:22:27 +03:00