Find a file
Georgi Gerganov 7bb0fc18f6
metal : add sparse FA (#28098)
* metal : support n_kv_max sparse mask hint in flash attention vec kernel

- add kernel_flash_attn_ext_vec_idx: compacts finite mask entries into
  a per-row index list (Hillis-Steele scan, one threadgroup per row)
- extend vec FA kernel with optional sparse index gathering (FC slot 5)
- add host-side gate: sparse path when n_kv_max > 0, mask present,
  supported head sizes / KV types, n_kv_max <= 4096
- new buffer region extra_idx for the index list
- pipeline getter extended with has_sparse param
- add test cases: head sizes, quant types, nb>1, nr23 variants,
  sinks, ALiBi, softcap, permute, v_view_of_k, no-mask fallback

Note: multi-row (nb*nr23[1] > 1) cases still failing - rid mapping
in the store phase needs revisiting for the sparse path.

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* metal : fix sparse flash attention row addressing

- kernel_flash_attn_ext_vec_idx: mask param is half* but nb31 is a byte
  stride, so the per-row mask offset was scaled by 2x; cast to char*
  before applying the byte strides
- kernel_flash_attn_ext_vec: sparse pidx param is char* so the per-row
  element offset was under-scaled by sizeof(int); scale it by sizeof(int)
  to get the correct byte offset
- fixes the multi-row (nb*nr23[1] > 1) sparse flash attention failures

Assisted-by: pi:llama.cpp/DeepSeek-v4-0731

* cont : use sparse vec FA for prefill

* metal : single-pass flash attention sparse index compaction

The idx kernel previously read the mask row twice: once to count the finite
entries (for the prefix scan) and again to recover their positions. Since the
kernel is memory-bound, this doubled the mask traffic.

Keep the finite positions in a per-thread register array during the count
pass and write them out directly, avoiding the second mask read. A dense
mask with more than NLOCAL finite entries in a slice falls back to re-reading
the mask to write the remaining positions.

Assisted-by: pi:llama.cpp/DeepSeek-v4-0731

* tests : add perf cases for sparse flash attention prefill

Measure the sparse vec FA kernel across KV sizes, n_kv_max hints and batch
sizes. Run with:

    ./build/bin/test-backend-ops -b MTL0 -o FLASH_ATTN_EXT -p "n_kv_max=[1-9]" perf

Assisted-by: pi:llama.cpp/DeepSeek-v4-0731

* qwen4 : enable sparse attention

* cont : adjust nsg

* cont : sync test-backend-ops

* cont : disable Qwen4 for now

* cont : clean-up + tests
2026-09-03 13:51:13 +03:00
.devops OpenVINO: Update OV to 2026.3.1, whisper.cpp support, Qwen3.5 on NPU, and new ops (#27843) 2026-08-28 14:42:07 +03:00
.gemini contributing: tighten AI usage policy (#18388) 2025-12-29 16:01:32 +01:00
.github ci : check for missing autoreleasepools (#27884) 2026-09-02 20:54:48 +03:00
.pi/gg ci : add older, min and dry-run options to ccache-clear (#27504) 2026-08-22 11:31:30 +03:00
app cmake : introduce semantic versioning (#26839) 2026-08-12 14:15:03 +02:00
benches benches : add Nemotron 3 Nano on DGX Spark (#20652) 2026-03-16 21:50:43 +02:00
ci ci : add check for unzip (#28082) 2026-08-31 12:17:51 +02:00
cmake CI: Use LLVM's OpenMP over MSVC_DEBUG_non_redist on Windows (#26678) 2026-08-20 15:42:26 +02:00
common common, server : enable preserve_reasoning kwarg by default, log its effective state (#28174) 2026-09-02 19:19:54 +03:00
conversion convert : skip bias_vl tensor in DeepSeek-V4 DSpark conversion (#28294) 2026-09-03 10:37:23 +03:00
docs sycl : support limit max alloc memory within 2GB for host-pinned memory (#27559) 2026-09-01 13:35:47 +03:00
examples finetune: fix no KV cache (#27199) 2026-09-02 23:53:32 +02:00
ggml metal : add sparse FA (#28098) 2026-09-03 13:51:13 +03:00
gguf-py model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
grammars docs : fix typos in CUDA-FEDORA.md and grammars/README.md (#24459) 2026-06-15 01:33:38 +08:00
include bench: add --tensor-read-lazy (#27881) 2026-08-28 20:51:05 +02:00
licenses refactor : remove libcurl, use OpenSSL when available (#18828) 2026-01-14 18:02:47 +01:00
media media : add transparent icon svg and png [no ci] (#15891) 2025-09-10 14:51:28 +03:00
models model: add Kimi-K3 text model (#26185) 2026-08-15 17:11:05 +02:00
pocs libs : rename libcommon -> libllama-common (#21936) 2026-04-17 11:11:46 +03:00
requirements requirements: use stable torch packages on s390x (#26864) 2026-08-11 21:58:53 +08:00
scripts vendor : update cpp-httplib to 0.54.0 (#27919) 2026-08-30 09:01:51 +03:00
skills test: move tools/parser to tests (#27548) 2026-08-23 18:38:51 +02:00
src metal : add sparse FA (#28098) 2026-09-03 13:51:13 +03:00
tests metal : add sparse FA (#28098) 2026-09-03 13:51:13 +03:00
tools mtmd : add const in various places (#28307) 2026-09-03 12:12:49 +02:00
vendor vendor : update cpp-httplib to 0.54.0 (#27919) 2026-08-30 09:01:51 +03:00
.clang-format fix: apply clang-format to CUDA macros (#16017) 2025-09-16 08:59:19 +02:00
.clang-tidy clang-tidy : disable warning about performance enum size (#16127) 2025-09-22 19:57:46 +02:00
.dockerignore docker : prebuild web UI for s390x build [no release] (#24829) 2026-06-20 05:54:42 -05:00
.ecrc
.editorconfig ui: Restructure repo to use tools/ui folder and ui / UI / llama-ui / LLAMA_UI naming (#23064) 2026-05-16 02:02:40 +02:00
.flake8
.gitignore ui: PWA support (#23871) 2026-06-12 15:53:26 +02:00
.gitmodules
.pre-commit-config.yaml
AGENTS.md arg: remove -no-cnv from cli [no ci] (#27542) 2026-08-22 15:53:56 +02:00
AUTHORS readme : update status badges + regen AUTHORS (#27317) 2026-08-18 14:35:04 +03:00
build-xcframework.sh metal : add metallib build support for xcframework (#28163) 2026-09-02 07:45:56 +08:00
CLAUDE.md contributing: tighten AI usage policy (#18388) 2025-12-29 16:01:32 +01:00
CMakeLists.txt ci: Clean up UI builds from releases (#27706) 2026-08-26 14:12:09 +02:00
CMakePresets.json cmake : Add CMake presets for Linux and GCC (#14656) 2025-07-13 08:12:36 +03:00
CODEOWNERS AVX2: Speed up large batch size prompt processing of IQ models (#27402) 2026-08-31 14:33:50 -04:00
CONTRIBUTING.md contrib : recommend waiting for CI before merging (#27603) 2026-08-23 15:56:47 +03:00
convert_hf_to_gguf.py convert: add option to create separate dspark GGUF (#26452) 2026-08-02 23:16:31 +08:00
convert_hf_to_gguf_update.py Add support for Laguna XS.2 & M.1 (#25165) 2026-07-22 09:54:08 +08:00
convert_llama_ggml_to_gguf.py ci : switch from pyright to ty (#20826) 2026-03-21 08:54:34 +01:00
convert_lora_to_gguf.py convert : fix lora base model arch retrieval (#24621) 2026-06-15 00:55:26 +02:00
flake.nix fix(nix): remove non-functional llama-cpp cachix cache from flake.nix (#15295) 2025-08-13 11:21:31 -07:00
LICENSE docs : Minor cleanups (#19252) 2026-02-02 08:38:55 +02:00
Makefile make : remove make in favor of CMake (#15449) 2025-08-20 13:31:16 +03:00
mypy.ini
pyproject.toml model: add Mellum architecture (#23966) 2026-06-02 22:11:12 +03:00
pyrightconfig.json ci : switch from pyright to ty (#20826) 2026-03-21 08:54:34 +01:00
README.md readme : update links (#27617) 2026-08-23 20:55:56 +03:00
requirements.txt
SECURITY.md security : clarify about AI-generated reports (#26579) 2026-08-05 13:27:06 +02:00
ty.toml mtmd : DeepSeek-OCR image processing fixes, img_tool::resize padding refactor (#23345) 2026-05-20 17:37:10 +02:00

llama.cpp

llama

Quick start

A few options to get llama.cpp installed on your machine:

Once installed:

# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with `llama cli` VLM session with llama cli Built-in web UI against `llama serve` running Qwen 3.6 Built-in web UI against llama serve

Description

The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.

  • Plain C/C++ implementation without any dependencies
  • Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
  • AVX, AVX2, AVX512 and AMX support for x86 architectures
  • RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
  • 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
  • Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
  • Vulkan and SYCL backend support
  • CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity

The llama.cpp project is build on top of the ggml library.

Supported backends

Backend Target devices
BLAS All
BLIS All
CANN Ascend NPU
CUDA Nvidia GPU
HIP AMD GPU
Hexagon [In Progress] Snapdragon
IBM zDNN IBM Z & LinuxONE
MUSA Moore Threads GPU
Metal Apple Silicon
OpenCL Adreno GPU
OpenVINO [In Progress] Intel CPUs, GPUs, and NPUs
RPC All
SYCL Intel GPU
VirtGPU VirtGPU APIR
Vulkan GPU
WebGPU All
ZenDNN AMD CPU

Documentation

Tools

Development

Contributing

  • Contributors can open PRs
  • Collaborators will be invited based on contributions
  • Maintainers can push to branches in the llama.cpp repo and merge PRs into the master branch
  • Any help with managing issues, PRs and projects is very appreciated!
  • Read the CONTRIBUTING.md for more information

Acknowledgements

  • yhirose/cpp-httplib - Single-header HTTP server, used by llama-server - MIT license
  • nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
  • nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
  • mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
  • sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain