Jonas Jankaitis
0324696b8e
fit : count nextn (MTP) blocks in n_gpu_layers so front layers stay on GPU ( #26177 )
Check Pre-Tokenizer Hashes / pre-tokenizer-hashes (push) Has been cancelled
Python Type-Check / python type-check (push) Has been cancelled
2026-07-27 16:21:37 +03:00
Aaron Teo
e6dd0e29a6
args: refactor mlock/mmap/directio into load-mode ( #20834 )
...
* args: overhaul mmap/mlock/dio into single arg
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* docs: update docs with llama-gen-docs
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* chore: satisfy code quality
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* args: make the `+` sign an actual modifier now
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* chore: general code clean up + comments
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* arg: fix deprecated flags support
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* arg: quick sanity check
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* bench: sync llama-bench argument parsing
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* fix: bugfix variable behaviour + llama-bench lm column size
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* arg: inverse commands should do the opposite instead of doing nothing
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* bench: fix incorrect dash
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* bench: fix missing modifiers for deprecated flags
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* llama: switch back to thread_local
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* arg: switch back to single enum
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* docs: update arg docs
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* chore: fix missing `mlock` from llama_load_mode_from_str + cleanup llama-bench
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* llama: fix mlock not activating
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* arg: add deprecation warning when old and new flags are combined
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* arg: cont add comment for todo in the future
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* docs: sync with upstream
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* docs: re-sync with upstream again
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
---------
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
2026-07-23 20:32:56 +08:00
Georgi Gerganov
27c8bb4f63
logs : reduce v2 ( #25078 )
...
* server : reduce logs
* cont : common
* cont : spec
* cont : CMN_ -> COM_
2026-06-28 08:52:15 +03:00
Georgi Gerganov
d8a24ccee2
fit : wrap llama_device_memory_data ( #24522 )
2026-06-13 08:09:52 +03:00
Aman Gupta
83eebe9d08
server: add margin for draft model for fit ( #23485 )
2026-05-24 14:43:08 +08:00
Georgi Gerganov
67b2b7f2f2
logs : reduce ( #23021 )
...
Python Type-Check / python type-check (push) Waiting to run
Check Pre-Tokenizer Hashes / pre-tokenizer-hashes (push) Has been cancelled
Python check requirements.txt / check-requirements (push) Has been cancelled
Update Operations Documentation / update-ops-docs (push) Has been cancelled
* logs : reduce
* args : fix envs
* server : fix build
* common : print verbosity level at start
* server : clean-up logs
* server : print prompt processing timings + sampling params
* minor : whitespaces
2026-05-14 13:05:52 +03:00
fl0rianr
a0101225bc
common: do not fit to unknown device memory ( #22614 )
...
* common: do not fit to unknown device memory
Signed-off-by: Florian Reinle <f.reinle@otec.de>
* common: preserve host fallback for non-GPU fit devices
Signed-off-by: Florian Reinle <f.reinle@otec.de>
* common: keep unknown GPU fit memory at zero
Signed-off-by: Florian Reinle <f.reinle@otec.de>
---------
Signed-off-by: Florian Reinle <f.reinle@otec.de>
2026-05-06 17:03:45 +02:00
rankaiyx
42401c72b8
Fix type casting for unaccounted memory calculation ( #22424 )
Check Pre-Tokenizer Hashes / pre-tokenizer-hashes (push) Has been cancelled
Python check requirements.txt / check-requirements (push) Has been cancelled
Python Type-Check / python type-check (push) Has been cancelled
2026-04-27 14:31:13 +02:00
Georgi Gerganov
cfe9838d26
fit-params : refactor + add option to output estimated memory per device ( #22171 )
...
* fit-params : add option to output estimated memory per device
* cont : minor
* cont : refactor
* cont : move fit params implementation to libcommon
* cont : header
* cont : headers
* cont : codeowners
2026-04-21 09:54:36 +03:00