koboldcpp/common
Bartosz Taudul ba8e0eddfb
common : skip device_info loop if it's not going to be printed (#26692)
The device_info loop iterates over the discovered devices and gets
the available and total memory counts. With the CUDA backend (and
possibly others too) this requires creating a GPU context, which,
in case of CUDA, results in a 550 MB VRAM allocation.

For this information to be used in any way, the log verbosity must
be set to LOG_LEVEL_TRACE. If it's not, including in the default
configuration, the contexts get created, memory sizes get queried,
then the log function quietly discards the data.

In certain cases the user may not want to use any GPU resources.
The device_loop iteration is the only place touching the GPU that
cannot be skipped.

Fix by checking the verbosity level and skipping the loop if there
would be no output.
2026-08-23 14:39:16 +02:00
..
jinja common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
arg.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
arg.h server: add dedup-cache-models preset option (#27346) 2026-08-19 11:04:26 +02:00
base64.hpp llava : expose as a shared library for downstream projects (#3613) 2023-11-07 00:36:23 +03:00
build-info.cpp.in cmake : introduce semantic versioning (#26839) 2026-08-12 14:15:03 +02:00
build-info.h cmake : introduce semantic versioning (#26839) 2026-08-12 14:15:03 +02:00
chat-auto-parser-generator.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
chat-auto-parser-helpers.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
chat-auto-parser-helpers.h chat : avoid including json in chat.h (#21306) 2026-04-03 09:07:59 +03:00
chat-auto-parser.h common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
chat-diff-analyzer.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
chat-peg-parser.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
chat-peg-parser.h common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
chat.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
chat.h common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
CMakeLists.txt common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
common.cpp common : skip device_info loop if it's not going to be printed (#26692) 2026-08-23 14:39:16 +02:00
common.h mtmd: add --mmproj-device argument (#23255) 2026-08-20 18:45:37 +02:00
console.cpp cli: fix stripping of \n in multiline input (#21485) 2026-04-06 20:54:06 +02:00
console.h cli : add command and file auto-completion (#19985) 2026-03-05 10:47:28 +01:00
debug.cpp common: fix missing exports in llama-common (#22340) 2026-04-27 08:06:39 +03:00
debug.h common: fix missing exports in llama-common (#22340) 2026-04-27 08:06:39 +03:00
download.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
download.h server: add dedup-cache-models preset option (#27346) 2026-08-19 11:04:26 +02:00
fit.cpp fit: also take into account n_streams (#27496) 2026-08-22 16:16:06 +02:00
fit.h fit: also take into account n_streams (#27496) 2026-08-22 16:16:06 +02:00
hf-cache.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
hf-cache.h server: (router) add model management API (#23976) 2026-06-17 18:04:58 +02:00
http.h cli : move to HTTP-based implementation (#24948) 2026-07-08 14:52:43 +02:00
imatrix-loader.cpp fix: check gguf array type before reading (#27075) 2026-08-15 11:45:30 +02:00
imatrix-loader.h Move duplicated imatrix code into single common imatrix-loader.cpp (#22445) 2026-06-04 17:45:40 +02:00
json-schema-to-grammar.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
json-schema-to-grammar.h common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
json.cpp common: json.h: fix clang lto (#27575) 2026-08-23 01:11:10 +02:00
json.h common: json.h: fix clang lto (#27575) 2026-08-23 01:11:10 +02:00
llguidance.cpp llama : support multi-output backend sampling (#25532) 2026-08-10 16:58:56 +03:00
log.cpp common: update logging to enforce max_capacity and optimize queue resizing (#24490) 2026-06-17 09:19:11 +03:00
log.h logs : reduce (#23021) 2026-05-14 13:05:52 +03:00
ngram-cache.cpp spec : add self‑speculative decoding (no draft model required) + refactor (#18471) 2026-01-28 19:42:42 +02:00
ngram-cache.h spec : add self‑speculative decoding (no draft model required) + refactor (#18471) 2026-01-28 19:42:42 +02:00
ngram-map.cpp speculative : fix out-of-bounds read in ngram-map on prompt shrink (#23936) 2026-07-07 10:25:04 +03:00
ngram-map.h fix: correct misspellings in code comments (#21217) 2026-03-31 13:50:51 +02:00
ngram-mod.cpp ngram-mod : Add missing include (#23857) 2026-05-29 09:21:37 +03:00
ngram-mod.h ngram-mod : fix build [no ci] (#19216) 2026-01-30 21:27:27 +02:00
peg-parser.cpp common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
peg-parser.h common: add json.h abstraction (#27511) 2026-08-22 16:28:28 +02:00
preset.cpp common: support --models-dir loading MTP assistant models (#24431) 2026-08-15 13:17:35 +02:00
preset.h common: add system-level config file (#26118) 2026-08-13 00:02:27 +02:00
reasoning-budget.cpp llama : support multi-output backend sampling (#25532) 2026-08-10 16:58:56 +03:00
reasoning-budget.h common : add support for multiple end sequences in the reasoning budget sampler (#25544) 2026-07-25 11:58:09 +02:00
sampling.cpp llama : support multi-output backend sampling (#25532) 2026-08-10 16:58:56 +03:00
sampling.h llama : support multi-output backend sampling (#25532) 2026-08-10 16:58:56 +03:00
speculative.cpp fit: also take into account n_streams (#27496) 2026-08-22 16:16:06 +02:00
speculative.h common : auto-detect spec type from draft GGUF metadata (#26814) 2026-08-13 12:34:27 +02:00
subproc.cpp common: add subproc.h wrapper, disabled on android/ios (#26102) 2026-07-26 20:54:25 +02:00
subproc.h common: add subproc.h wrapper, disabled on android/ios (#26102) 2026-07-26 20:54:25 +02:00
trie.cpp common : add support for multiple end sequences in the reasoning budget sampler (#25544) 2026-07-25 11:58:09 +02:00
trie.h common : add support for multiple end sequences in the reasoning budget sampler (#25544) 2026-07-25 11:58:09 +02:00
unicode.cpp common/parser: handle reasoning budget (#20297) 2026-03-11 10:26:12 +01:00
unicode.h common/parser: handle reasoning budget (#20297) 2026-03-11 10:26:12 +01:00