koboldcpp/common
王金旭 662a0b0121
Some checks failed
Python Type-Check / python type-check (push) Has been cancelled
spec : fuse the DFlash encoder into the KV cache injection (#27310)
* dflash : fuse the encoder into the KV injection decode

The encoder is a single fc + norm, but running it as a separate
llama_encode forced a device-to-host round trip of its output before the
injection decode could re-upload it, plus a second graph build per
round. Fold the encoder into the decoder's embd branch and feed the
target features directly to one llama_decode.

Assisted-by: Claude Fable

* nit

* Apply batched suggestions from code review

Co-authored-by: Ruixiang Wang <wangruixiang07@outlook.com>

* Fix missing references from renaming

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
Co-authored-by: Ruixiang Wang <wangruixiang07@outlook.com>
2026-08-31 11:19:20 +02:00
..
jinja common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
arg.cpp common: rename --tensor-read-lazy to --lazy-mode, add -lzm shorthand (#27969) 2026-08-30 09:18:10 +03:00
arg.h server: add dedup-cache-models preset option (#27346) 2026-08-19 11:04:26 +02:00
base64.hpp
build-info.cpp.in cmake : introduce semantic versioning (#26839) 2026-08-12 14:15:03 +02:00
build-info.h cmake : introduce semantic versioning (#26839) 2026-08-12 14:15:03 +02:00
chat-auto-parser-generator.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
chat-auto-parser-helpers.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
chat-auto-parser-helpers.h
chat-auto-parser.h common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
chat-diff-analyzer.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
chat-peg-parser.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
chat-peg-parser.h common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
chat.cpp chat : scope qwen3-coder workarounds (#27679) 2026-08-25 09:33:00 -05:00
chat.h common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
CMakeLists.txt common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
common.cpp bench: add --tensor-read-lazy (#27881) 2026-08-28 20:51:05 +02:00
common.h bench: add --tensor-read-lazy (#27881) 2026-08-28 20:51:05 +02:00
console.cpp
console.h
debug.cpp
debug.h
download.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
download.h server: add dedup-cache-models preset option (#27346) 2026-08-19 11:04:26 +02:00
fit.cpp fit: also take into account n_streams (#27496) 2026-08-22 16:16:06 +02:00
fit.h fit: also take into account n_streams (#27496) 2026-08-22 16:16:06 +02:00
hf-cache.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
hf-cache.h
http.h cli : move to HTTP-based implementation (#24948) 2026-07-08 14:52:43 +02:00
imatrix-loader.cpp fix: check gguf array type before reading (#27075) 2026-08-15 11:45:30 +02:00
imatrix-loader.h
json-schema-to-grammar.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
json-schema-to-grammar.h common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
json.cpp common: json.h: fix clang lto (#27575) 2026-08-23 01:11:10 +02:00
json.h common: json.h: fix clang lto (#27575) 2026-08-23 01:11:10 +02:00
llguidance.cpp llama : support multi-output backend sampling (#25532) 2026-08-10 16:58:56 +03:00
log.cpp
log.h
ngram-cache.cpp
ngram-cache.h
ngram-map.cpp speculative : fix out-of-bounds read in ngram-map on prompt shrink (#23936) 2026-07-07 10:25:04 +03:00
ngram-map.h
ngram-mod.cpp
ngram-mod.h
peg-parser.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
peg-parser.h common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
preset.cpp common: support --models-dir loading MTP assistant models (#24431) 2026-08-15 13:17:35 +02:00
preset.h common: add system-level config file (#26118) 2026-08-13 00:02:27 +02:00
reasoning-budget.cpp llama : support multi-output backend sampling (#25532) 2026-08-10 16:58:56 +03:00
reasoning-budget.h common : add support for multiple end sequences in the reasoning budget sampler (#25544) 2026-07-25 11:58:09 +02:00
sampling.cpp llama : support multi-output backend sampling (#25532) 2026-08-10 16:58:56 +03:00
sampling.h llama : support multi-output backend sampling (#25532) 2026-08-10 16:58:56 +03:00
speculative.cpp spec : fuse the DFlash encoder into the KV cache injection (#27310) 2026-08-31 11:19:20 +02:00
speculative.h spec: Add benchmark-only synthetic speculative acceptance options (#27711) 2026-08-27 13:53:42 +03:00
subproc.cpp common: add subproc.h wrapper, disabled on android/ios (#26102) 2026-07-26 20:54:25 +02:00
subproc.h common: add subproc.h wrapper, disabled on android/ios (#26102) 2026-07-26 20:54:25 +02:00
trie.cpp common : add support for multiple end sequences in the reasoning budget sampler (#25544) 2026-07-25 11:58:09 +02:00
trie.h common : add support for multiple end sequences in the reasoning budget sampler (#25544) 2026-07-25 11:58:09 +02:00
unicode.cpp
unicode.h