Find a file
Konrad Moren 96278e39fc
CUDA: Add backend sampler for penalties sampler (#25262)
* sampling: enhance penalty handling in common_sampler_init

- Set default value for penalty_last_n based on model context if not specified.
- Ensure penalty_last_n and n_prev are non-negative.
- Update llama_sampler_penalties structure to inherit from llama_sampler_backend and add backend input handling for penalties.
- Implement backend initialization and application logic for penalties, including frequency and presence adjustments.

* tests: add backend penalties sampling tests and utility functions

- Introduced `accept_prompt` and `unique_prompt_tokens` functions to handle prompt acceptance and token uniqueness.
- Implemented `compare_penalties_logits` to compare logits from backend and CPU samplers with penalties.
- Added `test_backend_penalties_sampling` to validate backend penalties with various configurations.
- Enhanced the test suite for better coverage of penalty handling in sampling.

* sampling: add support for top-k penalties in backend sampling

* sampling: add fix to ensure  stable numerical results. Preserve masked logits as -Inf and no longer generate NaN.

* sampling: enhance penalty comparison tests with masking penalties logic

* add comments on padding

* sampling: add comments on modifications

* add the unit test to cover masked-out token as -INF

* validate repeat penalty to ensure it is finite and greater than 0; add tests for invalid values

* refactor: test functions to share logic and be less verbose

* add test to cover case where previously penalized token is not part of candidates

* remove comments

* remove redundant penalty_last_n initialization and validation in common_sampler_init

* add support for penalties in sampler chain with configurable positions

* add validation for penalty parameters and enhance tests for non-finite values

* add context parameter to common_sampler_init and set default for penalty_last_n

* add llama_n_ctx parameter to common_sampler_init for improved sampler initialization

* replace penalty_last_n x n_candidates comparison matrix with a vocabulary-sized count tensor

* add tests for backend penalties sampling without filler entries , token_count.size() == n_active == n_max == 64

* add test for backend penalties sampling  after top-p with large history window

* remove as unused

* add is_disabled method, tensor logits reshape, add rest review suggestions

* clarify comment
2026-08-03 14:26:09 +02:00
.devops devops : add llama in all docker images (#25035) 2026-06-26 15:15:48 +02:00
.gemini contributing: tighten AI usage policy (#18388) 2025-12-29 16:01:32 +01:00
.github vulkan: update vulkan sdk to 1.4.357.0 (#26303) 2026-07-31 07:27:03 -05:00
.pi/gg pi : remove docs from system prompt (#24791) 2026-06-19 09:34:00 +03:00
app app : allow --version, --licenses & --help (#25054) 2026-06-26 23:18:11 +02:00
benches benches : add Nemotron 3 Nano on DGX Spark (#20652) 2026-03-16 21:50:43 +02:00
ci common : fix env names to all have LLAMA_ARG_ prefix (#23778) 2026-05-27 14:52:47 +03:00
cmake cmake : do not check for bin install dir (#23234) 2026-05-18 02:33:14 +02:00
common CUDA: Add backend sampler for penalties sampler (#25262) 2026-08-03 14:26:09 +02:00
conversion model: MTP support for Qwen3-Next (#25589) 2026-08-03 17:15:01 +09:00
docs SYCL: add oneMKL GEMM flash attention for XMX-accelerated prompt proc… (#25025) 2026-07-31 10:43:16 +03:00
examples args: refactor mlock/mmap/directio into load-mode (#20834) 2026-07-23 20:32:56 +08:00
ggml CUDA: Fix data-races when reusing SMEM in block_reduce (#26385) 2026-08-03 14:22:44 +02:00
gguf-py model: MTP support for Qwen3-Next (#25589) 2026-08-03 17:15:01 +09:00
grammars docs : fix typos in CUDA-FEDORA.md and grammars/README.md (#24459) 2026-06-15 01:33:38 +08:00
include CUDA: Add backend sampler for penalties sampler (#25262) 2026-08-03 14:26:09 +02:00
licenses refactor : remove libcurl, use OpenSSL when available (#18828) 2026-01-14 18:02:47 +01:00
media media : add transparent icon svg and png [no ci] (#15891) 2025-09-10 14:51:28 +03:00
models common/chat: add specialized minimax m3 parser (#26210) 2026-07-28 04:27:20 -05:00
pocs libs : rename libcommon -> libllama-common (#21936) 2026-04-17 11:11:46 +03:00
requirements model: add Mellum architecture (#23966) 2026-06-02 22:11:12 +03:00
scripts sync : ggml 2026-07-30 15:44:24 +03:00
skills agents: clarify comment style and jinja knowledge (#26405) 2026-08-01 18:45:46 +02:00
src CUDA: Add backend sampler for penalties sampler (#25262) 2026-08-03 14:26:09 +02:00
tests CUDA: Add backend sampler for penalties sampler (#25262) 2026-08-03 14:26:09 +02:00
tools CUDA: Add backend sampler for penalties sampler (#25262) 2026-08-03 14:26:09 +02:00
vendor vendor : update BoringSSL to 0.20260730.0 (#26353) 2026-08-01 20:53:00 +02:00
.clang-format fix: apply clang-format to CUDA macros (#16017) 2025-09-16 08:59:19 +02:00
.clang-tidy clang-tidy : disable warning about performance enum size (#16127) 2025-09-22 19:57:46 +02:00
.dockerignore docker : prebuild web UI for s390x build [no release] (#24829) 2026-06-20 05:54:42 -05:00
.ecrc
.editorconfig ui: Restructure repo to use tools/ui folder and ui / UI / llama-ui / LLAMA_UI naming (#23064) 2026-05-16 02:02:40 +02:00
.flake8 llama : move end-user examples to tools directory (#13249) 2025-05-02 20:27:13 +02:00
.gitignore ui: PWA support (#23871) 2026-06-12 15:53:26 +02:00
.gitmodules ggml : remove kompute backend (#14501) 2025-07-03 07:48:32 +03:00
.pre-commit-config.yaml
AGENTS.md agents: clarify comment style and jinja knowledge (#26405) 2026-08-01 18:45:46 +02:00
AUTHORS authors : update (#19263) 2026-02-02 08:51:25 +02:00
build-xcframework.sh xcframework : disable mtmd video on i/tv/visionos (#25018) 2026-06-26 00:13:59 +02:00
CLAUDE.md contributing: tighten AI usage policy (#18388) 2025-12-29 16:01:32 +01:00
CMakeLists.txt common: add subproc.h wrapper, disabled on android/ios (#26102) 2026-07-26 20:54:25 +02:00
CMakePresets.json cmake : Add CMake presets for Linux and GCC (#14656) 2025-07-13 08:12:36 +03:00
CODEOWNERS HIP: remove rocWMMA FlashAttention (#26046) 2026-07-24 17:53:54 +02:00
CONTRIBUTING.md contrib : add guideline about the "merge ready" label (#26178) 2026-07-28 08:41:04 +03:00
convert_hf_to_gguf.py convert: add option to create separate dspark GGUF (#26452) 2026-08-02 23:16:31 +08:00
convert_hf_to_gguf_update.py Add support for Laguna XS.2 & M.1 (#25165) 2026-07-22 09:54:08 +08:00
convert_llama_ggml_to_gguf.py ci : switch from pyright to ty (#20826) 2026-03-21 08:54:34 +01:00
convert_lora_to_gguf.py convert : fix lora base model arch retrieval (#24621) 2026-06-15 00:55:26 +02:00
flake.nix fix(nix): remove non-functional llama-cpp cachix cache from flake.nix (#15295) 2025-08-13 11:21:31 -07:00
LICENSE docs : Minor cleanups (#19252) 2026-02-02 08:38:55 +02:00
Makefile make : remove make in favor of CMake (#15449) 2025-08-20 13:31:16 +03:00
mypy.ini
pyproject.toml model: add Mellum architecture (#23966) 2026-06-02 22:11:12 +03:00
pyrightconfig.json ci : switch from pyright to ty (#20826) 2026-03-21 08:54:34 +01:00
README.md readme : refresh (#26280) 2026-07-30 16:14:37 +03:00
requirements.txt tool-call: fix Qwen 2.5 Coder support, add micro benchmarks, support trigger patterns for lazy grammars (#12034) 2025-03-05 13:05:13 +00:00
SECURITY.md binaries : Improve rpc-server and export-graph-ops names. (#25045) 2026-06-27 10:31:29 +03:00
ty.toml mtmd : DeepSeek-OCR image processing fixes, img_tool::resize padding refactor (#23345) 2026-05-20 17:37:10 +02:00

llama.cpp

llama

Quick start

A few options to get llama.cpp installed on your machine:

Once installed:

# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with `llama cli` VLM session with llama cli Built-in web UI against `llama serve` running Qwen 3.6 Built-in web UI against llama serve

Description

The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.

  • Plain C/C++ implementation without any dependencies
  • Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
  • AVX, AVX2, AVX512 and AMX support for x86 architectures
  • RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
  • 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
  • Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
  • Vulkan and SYCL backend support
  • CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity

The llama.cpp project is build on top of the ggml library.

Supported backends

Backend Target devices
BLAS All
BLIS All
CANN Ascend NPU
CUDA Nvidia GPU
HIP AMD GPU
Hexagon [In Progress] Snapdragon
IBM zDNN IBM Z & LinuxONE
MUSA Moore Threads GPU
Metal Apple Silicon
OpenCL Adreno GPU
OpenVINO [In Progress] Intel CPUs, GPUs, and NPUs
RPC All
SYCL Intel GPU
VirtGPU VirtGPU APIR
Vulkan GPU
WebGPU All
ZenDNN AMD CPU

Documentation

Tools

Development

Contributing

  • Contributors can open PRs
  • Collaborators will be invited based on contributions
  • Maintainers can push to branches in the llama.cpp repo and merge PRs into the master branch
  • Any help with managing issues, PRs and projects is very appreciated!
  • Read the CONTRIBUTING.md for more information

Acknowledgements

  • yhirose/cpp-httplib - Single-header HTTP server, used by llama-server - MIT license
  • stb-image - Single-header image format decoder, used by multimodal subsystem - Public domain
  • nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
  • miniaudio.h - Single-header audio format decoder, used by multimodal subsystem - Public domain
  • subprocess.h - Single-header process launching solution for C and C++ - Public domain