koboldcpp/src/models
Concedo 2d357d8359 Merge branch 'upstream' into concedo_experimental
# Conflicts:
#	.devops/nix/package.nix
#	.github/ISSUE_TEMPLATE/config.yml
#	.github/workflows/make-release.yml
#	docs/autoparser.md
#	flake.nix
#	ggml/src/ggml-opencl/ggml-opencl.cpp
#	ggml/src/ggml-openvino/ggml-openvino.cpp
#	ggml/src/ggml-sycl/ggml-sycl.cpp
#	ggml/src/ggml-sycl/norm.cpp
#	ggml/src/ggml-sycl/norm.hpp
#	ggml/src/ggml-webgpu/ggml-webgpu.cpp
#	models/templates/README.md
#	scripts/make-release-checks.sh
#	scripts/ui-assets.cmake
#	tests/test-backend-ops.cpp
#	tests/test-chat.cpp
#	tests/test-llama-archs.cpp
#	tools/cli/README.md
#	tools/completion/README.md
#	tools/server/CMakeLists.txt
#	tools/server/README.md
2026-09-08 11:59:02 +08:00
..
afmoe.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
apertus.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
arcee.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
arctic.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
arwkv7.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
baichuan.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
bailingmoe.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
bailingmoe2.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
bailingmoe3.cpp models : fix GDN normalization from max to rsqrt (#28068) 2026-09-06 18:46:21 +02:00
bert.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
bitnet.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
bloom.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
chameleon.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
chatglm.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
clip.cpp llama: Restore quantization of mmprojs (#26818) 2026-08-10 11:58:32 +02:00
codeshell.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
cogvlm.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
cohere2.cpp Add arch support for cohere2-MoE (#24260) 2026-06-13 19:49:00 +02:00
cohere2moe.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
command-r.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
dbrx.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
deci.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
deepseek.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
deepseek2.cpp convert : add --fuse-qkv flag to fuse Q/K/V into QKV during HF-to-GGUF conversion (#22780) 2026-09-07 00:47:05 +08:00
deepseek2ocr.cpp convert : add --fuse-qkv flag to fuse Q/K/V into QKV during HF-to-GGUF conversion (#22780) 2026-09-07 00:47:05 +08:00
deepseek4.cpp model, mtmd: fix gemma4 vision handling (#28335) 2026-09-04 12:23:27 +02:00
deepseek32.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
delta-net-base.cpp llama: refactor fused ops (#24646) 2026-07-08 18:18:09 +08:00
dflash.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
dots1.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
dots3note.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
dream.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
eagle3.cpp spec: add eagle3-v3 support for gpt-oss model (#25794) 2026-07-28 09:58:16 +03:00
ernie4-5-moe.cpp model : NvFP4 quantized LM head support (#23046) 2026-05-16 11:09:27 +02:00
ernie4-5.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
eurobert.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
exaone-moe.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
exaone.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
exaone4.cpp model : load hparams.n_layer_nextn before n_layer() calls (#28159) 2026-09-01 13:55:45 +02:00
falcon-h1.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
falcon.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
gemma-embedding.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
gemma.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
gemma2.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
gemma3.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
gemma3n.cpp convert : add --fuse-qkv flag to fuse Q/K/V into QKV during HF-to-GGUF conversion (#22780) 2026-09-07 00:47:05 +08:00
gemma4-assistant.cpp model : fix gemma4-assistant (#28183) 2026-09-01 19:58:44 +03:00
gemma4.cpp Merge branch 'upstream' into concedo_experimental 2026-09-08 11:59:02 +08:00
glm-dsa.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
glm4-moe.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
glm4.cpp model : load hparams.n_layer_nextn before n_layer() calls (#28159) 2026-09-01 13:55:45 +02:00
gpt2.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
gptneox.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
granite-hybrid.cpp model : GraniteSWAForCausalLM / GraniteMoeSWAForCausalLM (#25505) 2026-08-19 16:53:31 +02:00
granite-moe.cpp model : GraniteSWAForCausalLM / GraniteMoeSWAForCausalLM (#25505) 2026-08-19 16:53:31 +02:00
granite-swa.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
granite-switch.cpp model : GraniteSWAForCausalLM / GraniteMoeSWAForCausalLM (#25505) 2026-08-19 16:53:31 +02:00
granite.cpp model : GraniteSWAForCausalLM / GraniteMoeSWAForCausalLM (#25505) 2026-08-19 16:53:31 +02:00
grok.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
grovemoe.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
hunyuan-dense.cpp model: move load_hparams and load_tensors to per-model definition (#22004) 2026-05-04 12:36:59 +02:00
hunyuan-moe.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
hunyuan-vl.cpp model : NvFP4 quantized LM head support (#23046) 2026-05-16 11:09:27 +02:00
hy-v3.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
hy-v4.cpp model : add Tencent Hy 4 (hy_v4) preview architecture support (#28127) 2026-09-04 14:31:36 +02:00
internlm2.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
jais.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
jais2.cpp convert : add --fuse-qkv flag to fuse Q/K/V into QKV during HF-to-GGUF conversion (#22780) 2026-09-07 00:47:05 +08:00
jamba.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
jina-bert-v2.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
jina-bert-v3.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
kimi-k3.cpp models : fix GDN normalization from max to rsqrt (#28068) 2026-09-06 18:46:21 +02:00
kimi-linear.cpp convert : add --fuse-qkv flag to fuse Q/K/V into QKV during HF-to-GGUF conversion (#22780) 2026-09-07 00:47:05 +08:00
laguna.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
lfm2.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
lfm2moe.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
llada-moe.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
llada.cpp convert : add --fuse-qkv flag to fuse Q/K/V into QKV during HF-to-GGUF conversion (#22780) 2026-09-07 00:47:05 +08:00
llama-embed.cpp model: move load_hparams and load_tensors to per-model definition (#22004) 2026-05-04 12:36:59 +02:00
llama.cpp spec: add EAGLE3 speculative decoding support (#18039) 2026-06-12 10:21:06 +03:00
llama4.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
maincoder.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
mamba-base.cpp mamba2 : Flatten in/out projections to dispatch GEMM instead of GEMV (#27513) 2026-08-24 09:25:11 +03:00
mamba.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
mamba2.cpp mamba2: remove hardcoded 2x expansion factor and invalid d_inner % d_state check (#23082) 2026-06-26 08:50:54 +03:00
mellum.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
mimo2.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
minicpm.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
minicpm3.cpp model: use ggml_rope_set_offset() (#27382) 2026-08-21 18:54:29 +02:00
minimax-01.cpp tests : run test-save-load-state across all architectures (#27755) 2026-08-28 09:45:19 +03:00
minimax-m2.cpp convert : add --fuse-qkv flag to fuse Q/K/V into QKV during HF-to-GGUF conversion (#22780) 2026-09-07 00:47:05 +08:00
minimax-m3.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
mistral3.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
mistral4.cpp model: move load_hparams and load_tensors to per-model definition (#22004) 2026-05-04 12:36:59 +02:00
models.h models : fix GDN normalization from max to rsqrt (#28068) 2026-09-06 18:46:21 +02:00
modern-bert.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
mpt.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
muse-glimmer.cpp model: Muse Glimmer Support (#26841) 2026-08-10 13:07:27 +02:00
nanbeige.cpp models : support nanbeige4.2-3B (#27730) 2026-08-27 07:55:31 +03:00
nemotron-h-moe.cpp model: add MTP support for Nemotron model (#26725) 2026-08-10 11:25:24 +03:00
nemotron-h.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
nemotron.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
neo-bert.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
nomic-bert-moe.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
nomic-bert.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
olmo.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
olmo2.cpp convert : add --fuse-qkv flag to fuse Q/K/V into QKV during HF-to-GGUF conversion (#22780) 2026-09-07 00:47:05 +08:00
olmoe.cpp convert : add --fuse-qkv flag to fuse Q/K/V into QKV during HF-to-GGUF conversion (#22780) 2026-09-07 00:47:05 +08:00
openai-moe.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
openelm.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
orion.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
paddleocr.cpp model : NvFP4 quantized LM head support (#23046) 2026-05-16 11:09:27 +02:00
pangu-embed.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
phi2.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
phi3.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
phimoe.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
plamo.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
plamo2.cpp ggml : recurrent state rollback for ggml_ssm_scan (#26623) 2026-08-14 17:20:40 +03:00
plamo3.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
plm.cpp model: use ggml_rope_set_offset() (#27382) 2026-08-21 18:54:29 +02:00
pockettts.cpp mtmd: support pocket-tts (#26871) 2026-08-11 14:18:30 +02:00
qwen.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
qwen2.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
qwen2moe.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
qwen2vl.cpp model : NvFP4 quantized LM head support (#23046) 2026-05-16 11:09:27 +02:00
qwen3.cpp spec: add EAGLE3 speculative decoding support (#18039) 2026-06-12 10:21:06 +03:00
qwen3moe.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
qwen3next.cpp convert : add --fuse-qkv flag to fuse Q/K/V into QKV during HF-to-GGUF conversion (#22780) 2026-09-07 00:47:05 +08:00
qwen3tts.cpp mtmd: support Qwen3-TTS (note: breaking change to llama-tts binary) (#26254) 2026-08-04 17:26:15 +02:00
qwen3vl.cpp mtmd: support Qwen3-TTS (note: breaking change to llama-tts binary) (#26254) 2026-08-04 17:26:15 +02:00
qwen3vlmoe.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
qwen4exp.cpp models : fix GDN normalization from max to rsqrt (#28068) 2026-09-06 18:46:21 +02:00
qwen35.cpp convert : add --fuse-qkv flag to fuse Q/K/V into QKV during HF-to-GGUF conversion (#22780) 2026-09-07 00:47:05 +08:00
qwen35moe.cpp convert : add --fuse-qkv flag to fuse Q/K/V into QKV during HF-to-GGUF conversion (#22780) 2026-09-07 00:47:05 +08:00
refact.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
rnd1.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
rwkv6-base.cpp models : deduplicate delta-net graphs for Qwen family (#19597) 2026-02-16 14:35:04 +02:00
rwkv6.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
rwkv6qwen2.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
rwkv7-base.cpp models : deduplicate delta-net graphs for Qwen family (#19597) 2026-02-16 14:35:04 +02:00
rwkv7.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
seed-oss.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
smallthinker.cpp model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) 2026-09-03 08:53:08 +02:00
smollm3.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
spark2-5.cpp [Model] Support for Spark2_5ForCausalLM implementation (#27868) 2026-09-06 17:43:58 +02:00
stablelm.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
starcoder.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
starcoder2.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
step35.cpp convert : add --fuse-qkv flag to fuse Q/K/V into QKV during HF-to-GGUF conversion (#22780) 2026-09-07 00:47:05 +08:00
t5.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
t5encoder.cpp model: move load_hparams and load_tensors to per-model definition (#22004) 2026-05-04 12:36:59 +02:00
talkie.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00
wavtokenizer-dec.cpp model : NvFP4 quantized LM head support (#23046) 2026-05-16 11:09:27 +02:00
xverse.cpp hparams : refactor hparams.n_layer (#24060) 2026-06-05 11:09:36 +03:00