mirror of
https://github.com/LostRuins/koboldcpp.git
synced 2026-07-22 23:33:34 +00:00
* llama: save more VRAM by reserving n_outputs == n_seqs when possible * add n_outputs_per_seq * move n_outputs_max to server-context * change ubatch to batch everywhere |
||
|---|---|---|
| .. | ||
| llama-cpp.h | ||
| llama.h | ||