Find a file
2026-10-03 00:23:59 +08:00
.github ci : use hf-jobs-cpu-xl runner in server sanitize workflow (#29297) 2026-09-23 21:30:39 +03:00
common Merge branch 'upstream' into concedo_experimental 2026-09-24 16:44:08 +08:00
conversion model : add Ling 3.0 VL support (#29151) 2026-09-24 08:57:31 +02:00
embd_res bump default ctx to 16k 2026-09-21 01:32:59 +08:00
examples model-conversion : add causal-compare-logits recipe (#29305) 2026-09-23 12:51:35 +02:00
ggml revert repack.cpp changes 2026-09-24 17:24:20 +08:00
gguf-py Merge branch 'upstream' into concedo_experimental 2026-09-24 16:44:08 +08:00
include Merge branch 'upstream' into concedo_experimental 2026-09-19 17:00:05 +08:00
kcpp_adapters fixing ling template 2026-08-28 18:07:03 +08:00
lib remove clblast, part 2 2026-01-23 14:09:46 +08:00
media update previews 2025-09-11 19:00:26 +08:00
otherarch sd: sync with master-866-42d6c0a (#2458) 2026-09-24 16:21:20 +08:00
scripts scripts : make-release-desc - link previous release in changelog title (#29336) 2026-09-24 07:46:38 +03:00
src remove some logspam (+1 squashed commits) 2026-09-27 01:39:20 +08:00
tests cuda : add conv3d with implicit GEMM (#29137) 2026-09-24 10:24:57 +03:00
tools missing file 2026-09-24 16:45:02 +08:00
vendor Merge commit '4ceb171910' into concedo_experimental 2026-09-24 16:37:57 +08:00
.clang-format fix: apply clang-format to CUDA macros (#16017) 2025-09-16 08:59:19 +02:00
.editorconfig ui: Restructure repo to use tools/ui folder and ui / UI / llama-ui / LLAMA_UI naming (#23064) 2026-05-16 02:02:40 +02:00
.gitignore ui: PWA support (#23871) 2026-06-12 15:53:26 +02:00
android_install.sh updated android termux script (+1 squashed commits) 2025-08-15 00:16:33 +08:00
aria2c-win.exe embed aria2c for windows, add slowness check with highpriority recommendation (+1 squashed commits) 2025-05-06 18:56:02 +08:00
CMakeLists.txt fix cmakelists 2026-09-24 16:24:26 +08:00
colab.ipynb default 40k ctx colab 2026-06-30 23:14:47 +08:00
convert_hf_to_gguf.py Merge branch 'upstream' into concedo_experimental 2026-09-08 11:59:02 +08:00
convert_hf_to_gguf_update.py vocab : add ufakzeka pre-tokenizer (#29033) 2026-09-18 11:45:11 +03:00
convert_llama_ggml_to_gguf.py ci : switch from pyright to ty (#20826) 2026-03-21 08:54:34 +01:00
convert_lora_to_gguf.py convert : fix lora base model arch retrieval (#24621) 2026-06-15 00:55:26 +02:00
create_ver_file.bat cleanup, try to add version tagging 2024-11-23 12:59:06 +08:00
create_ver_file.sh KoboldCpp.sh updates (#1562) 2025-05-26 15:24:49 +08:00
cudart64_12.dll breaking change: due to cuda12 upgrade, release filenames will change. standardize them to windows naming for the future. (+1 squashed commits) 2025-06-06 14:02:34 +08:00
cudart64_110.dll updated runtimes to henky version 2023-07-18 18:48:54 +08:00
environment-nocuda.yaml remove CLBlast, part 1 2026-01-23 13:50:12 +08:00
environment.yaml remove CLBlast, part 1 2026-01-23 13:50:12 +08:00
expose.cpp stability and memory access fixes (codex generated/reviewed) 2026-09-13 22:27:35 +08:00
expose.h wip on adding ubatch 2026-09-18 00:15:26 +08:00
glslc-linux KoboldCpp.sh updates (#1562) 2025-05-26 15:24:49 +08:00
glslc.exe Merge branch 'upstream' into concedo_experimental 2025-04-01 20:16:07 +08:00
gpttype_adapter.cpp wip on adding ubatch 2026-09-18 00:15:26 +08:00
hash.cpp Merge commit '533b18257b' into concedo_experimental 2026-08-20 18:42:18 +08:00
json_to_gbnf.py Merge branch 'upstream' into concedo_experimental 2026-09-04 16:59:17 +08:00
kcpp_agent.py agent allows self signed ssl 2026-09-24 22:31:17 +08:00
kcpp_backend.cpp fix whisper device selection 2026-08-22 18:16:59 +08:00
kcpp_backend.h no need backend specific tensor split 2026-08-23 00:25:53 +08:00
koboldcpp.py remove some logspam (+1 squashed commits) 2026-09-27 01:39:20 +08:00
koboldcpp.sh wip kcpp agent 2026-09-21 01:28:40 +08:00
LICENSE.md update license, added backwards compatibility with both ggml model formats, fixed context length issues. 2023-03-20 23:43:35 +08:00
make_pyinstaller.bat wip kcpp agent 2026-09-21 01:28:40 +08:00
make_pyinstaller.sh wip kcpp agent 2026-09-21 01:28:40 +08:00
make_pyinstaller_cuda.bat wip kcpp agent 2026-09-21 01:28:40 +08:00
make_pyinstaller_cuda_oldpc.bat wip kcpp agent 2026-09-21 01:28:40 +08:00
make_pyinstaller_oldpc.bat wip kcpp agent 2026-09-21 01:28:40 +08:00
Makefile sd: sync with master-866-42d6c0a (#2458) 2026-09-24 16:21:20 +08:00
MIT_LICENSE_GGML_SDCPP_LLAMACPP_ONLY.md bundle AGPL license and llama.cpp's MIT license into binaries. clarified some licensing terms, updated readme (+1 squashed commits) 2025-05-18 02:21:27 +08:00
model_adapter.cpp prevent segfault on load fail 2026-09-13 22:01:18 +08:00
model_adapter.h fix streaming race condition 2026-08-30 18:04:49 +08:00
mypy.ini convert : partially revert PR #4818 (#5041) 2024-01-20 18:14:18 -05:00
niko.ico fixed some build errors on linux, changed icon resolution, added more error printing 2023-05-22 12:18:42 +08:00
nikogreen.ico wip on unified cublas integration, add all the small libraries but exclude the large ones 2023-06-29 18:35:31 +08:00
README.md Updated and modernized readme 2026-10-03 00:23:59 +08:00
Remote-Link.cmd improved remotelink cmd, fixed lib unload, updated class.py 2023-09-25 17:50:00 +08:00
requirements.txt mtp init -2 2026-06-21 10:21:48 +08:00
simplecpuinfo better way of checking for avx2 support 2025-06-22 22:56:50 +08:00
simplecpuinfo.cpp better way of checking for avx2 support 2025-06-22 22:56:50 +08:00
simplecpuinfo.exe better way of checking for avx2 support 2025-06-22 22:56:50 +08:00
version.txt add product version 2025-07-26 18:35:37 +08:00
version_template.txt add product version 2025-07-26 18:35:37 +08:00

KoboldCpp: Run local AI models with a built-in web UI

KoboldCpp is free and open-source software for running GGUF large language models (LLMs) on your own computer. Chat with an AI assistant, write stories, roleplay, or connect other apps to a local API. KoboldCpp runs on CPU or GPU and also includes text, image, video, speech and music generation with compatible models, an integrated agent, a bundled KoboldAI Lite WebUI, and many additional powerful features.

Inspired by KoboldAI and built on llama.cpp

One executable file, no installation required. Ready-to-run downloads are available for Windows, Linux, and macOS.

Download KoboldCpp | Documentation and FAQ | API reference | Discord community

Integrated Web UI Roleplay Chat Mode GUI Launcher Messenger Chat Mode Image Generation UI Assistant UI

Features

  • Local text generation: Run any GGUF language model to generate text in chat, adventure, instruct, or story writing modes. Compatible vision models also support image understanding.
  • Image generation and editing: Supports Stable Diffusion 1.5, SDXL, SD3, Flux, Qwen Image, Ideogram, Z-Image, Klein, Krea2 and more.
  • Video generation: Supports WAN 2.2, LTX2.3, MiniMax H3 and more.
  • Voice recognition: Speech-to-text with Whisper, and multimodal audio from Gemma4 E2B and E4B.
  • Speech generation: Text to speech with Qwen3TTS, Kokoro, OuteTTS, Parler, and Dia.
  • Music generation: ACE Step 1.5 and ACE Step XL.
  • Tools and agents: MCP server support, tool calling, web search, retrieval-augmented generation (RAG) through TextDB, and an integrated KoboldCpp Agent for writing and editing code, running programs, and scheduling tasks.
    • To start the agent, select Launch KoboldCpp Agent in the launcher's Admin tab, or add --agent to your launch command.
  • Writing and roleplay tools: Bundled KoboldAI Lite WebUI includes multiple UI themes, editing tools, memory, world info, author's notes, characters, scenarios, and persistent story saves. Import Tavern character cards and other supported formats from file or external sites. Also includes the classic llama.cpp WebUI.
  • App integrations: Compatible endpoints for KoboldAI, OpenAI, Ollama, A1111/Forge, ComfyUI, Whisper transcription, XTTS, and OpenAI speech clients. See APIs and integrations.
  • Fully Portable - Single standalone executable for Windows, Linux or macOS, with no installation required and no external dependencies. Runs on CPU or GPU, with full or partial offloading. Can also run on Colab, Docker, also supports other platforms if self-compiled (like Android via Termux and Raspberry PI).

Phishing Scam Alert ⚠️

Quick start

  1. Download KoboldCpp for your operating system from the latest KoboldCpp release. See the platform instructions below for help choosing a file.
  2. Download a GGUF text model. Models are separate from the software. If unsure, start with the example models, or open Get Help and pick from Newbie Templates in the launcher for an easy setup.
  3. Open KoboldCpp and select your model in the GGUF Text Model field. Choose hardware settings suited to your computer. Generally the defaults should work, see GPU and performance settings if needed.
    • Need help launching? See the platform instructions.
    • A dedicated GPU is optional; the model size and context length determine how much memory you need.
  4. Click Launch and wait for the model to load. Keep KoboldCpp running while you use it.
  5. Connect to the Web UI in your browser once ready at http://localhost:5001

Download and Run

You should choose the correct KoboldCpp executable from the Assets section of the latest KoboldCpp release. Here are direct links and a quick overview for each supported platform. Models must be obtained separately

Windows

  • Download koboldcpp.exe and double-click it to open the launcher.
  • If you do not need NVIDIA CUDA support, koboldcpp-nocuda.exe is a smaller download with CPU and Vulkan support.
  • If you're using a PC with an older CPU or GPU, try koboldcpp-oldpc.exe if you encounter compatibility issues.
  • KoboldCpp can also be run using the command line, for example koboldcpp.exe --model "C:\Models\model.gguf" For more info, please check koboldcpp.exe --help
  • One liner setup for windows:
cmd /c "curl -fLo koboldcpp.exe https://github.com/LostRuins/koboldcpp/releases/latest/download/koboldcpp.exe && koboldcpp.exe"

Linux

  • Download koboldcpp-linux-x64 for an x86-64 Linux system. Make it executable with chmod +x koboldcpp-linux-x64, then launch it from a terminal in the download folder with ./koboldcpp-linux-x64.
  • Use koboldcpp-linux-x64-nocuda if you do not need CUDA, or try koboldcpp-linux-x64-oldpc for older hardware. Substitute that filename in the commands above.
  • For command-line usage see ./koboldcpp-linux-x64 --help, models can be loaded directly with ./koboldcpp-linux-x64 --model /path/to/model.gguf
  • For other hardware or distributions that cannot run the binaries, see building from source.
  • One liner setup for linux:
curl -fLo koboldcpp-linux-x64 https://github.com/LostRuins/koboldcpp/releases/latest/download/koboldcpp-linux-x64 && chmod +x koboldcpp-linux-x64 && ./koboldcpp-linux-x64

macOS

  • Download koboldcpp-mac-arm64 for an Apple Silicon (M-series) ARM64 Mac. Make it executable with chmod +x koboldcpp-mac-arm64, then launch it from a terminal in the download folder with ./koboldcpp-mac-arm64.
  • If macOS blocks the app, follow Apple's instructions to whitelist it in security settings under System Settings and Privacy & Security. A macOS launch walkthrough is also available.
  • For command-line usage see ./koboldcpp-mac-arm64 --help, models can be loaded directly with ./koboldcpp-mac-arm64 --model /path/to/model.gguf
  • Intel Macs require a source build.

Android and other platforms

Android users can install through Termux. Source builds also support platforms such as OpenBSD and Raspberry Pi, see the build instructions and KoboldCpp wiki.

Other ways to run KoboldCpp

External providers without a local model

  • KoboldCpp allows connecting the web UI directly with a supported external AI provider instead of using a local model. Currently, AI Horde, OpenAI Compatible, Anthropic, OpenRouter, Gemini, Grok, Mistral are among the supported services.
  • To connect, start with --nomodel or select Allow Launch Without Models in the launcher's Loaded Files tab. Alternatively, you can also use the online KoboldAI Lite WebUI directly.

Cloud GPUs and public demo

Docker

  • Caution: The official KoboldCpp Docker image is intended for experts only, primarily for cloud GPU rentals. It uses an x86-64 Ubuntu environment internally and expects an NVIDIA or AMD GPU.
  • Docker may perform poorly on some Windows or macOS setups, and ARM systems may fail to run it. CPU feature detection can also incorrectly select slower fallback binaries on some systems.
  • For local use, you're recommended to start with the prebuilt binaries.

Obtaining a GGUF model

KoboldCpp does not include model files. For local text generation, huggingface.co hosts many GGUF models, including Bartowski's model collection. Image generation, music and audio features use their own model files and settings, and CivitAI has a good source of image models. Start with a smaller model if you are unsure what your computer can run. Alternatively, click 'Get Help' in the GUI launcher and browse the 'Newbie Templates' (recommended)

KoboldCpp also retains backward compatibility with legacy GGML .bin models, though some newer features may be unavailable.

To enable vision, load the matching MMProj file in Mmproj File under the launcher's Loaded Files tab, or add --mmproj /path/to/mmproj.gguf to your launch command alongside the text model.

To convert your own models, use the GGUF conversion and quantization tools: run convert_hf_to_gguf.py, then quantize_gguf.exe to quantize the result.

Troubleshooting and Improving Performance

KoboldCpp provides many hardware configurations that can affect performance. Generally the default configuration should work decently, however you can make some adjustments to optimize your experience.

  • System runs out of RAM: Try a smaller model or a shorter context. Reducing GPU layers can increase system RAM usage by moving more model weights off the GPU.
    • Adjusting Context Size: Use --contextsize N to set the maximum context length: how much text the model can work with at once, measured in tokens. Larger contexts need more memory.
    • Adjusting Batch Size: Use --batchsize N to adjust prompt-processing batch size. A smaller batch can use less memory but might be slower. Set -1 to disable batching.
  • GPU runs out of VRAM: Try fewer GPU layers, a smaller model, or a shorter context. Autofit is an estimate and may need manual adjustment.
  • Generation is slow: Check that the intended GPU backend is selected and that layers are offloaded. CPU-only generation works, but speed depends on your hardware and model.
    • GPU Acceleration: Windows and Linux users with GPUs can use --usecuda flag (Nvidia Only), or --usevulkan (AMD, Nvidia, Intel GPUs) for GPU acceleration, make sure you select the correct .exe with CUDA support. This is also selectable in the hardware preset in the GUI launcher.
    • GPU Layer Offloading: Add --gpulayers N to offload model layers to the GPU. The default, -1, enables autofit; 0 disables GPU offloading. Lower the layer count if you run out of GPU memory.
  • An older computer crashes at startup: Some devices lack newer CPU instruction support. Try an oldpc release or use --noavx2. See the release notes for hardware compatibility.
  • The browser cannot connect: Wait for model loading to finish and check the address printed in the terminal. The default is localhost:5001; a custom --port changes it.
  • A model will not load: Check the terminal error, your available memory, and whether your KoboldCpp version supports the model. You can trigger debug mode with the --debugmode flag or launch toggle. Try the latest release and check the wiki or create a Github issue to report a bug.
  • Model is incoherent: You might be using an incorrect chat template. Try relaunch with --jinjatools to use the included Jinja template, or enable the Jinja toggle in the GUI.
  • For more information, be sure to run the program with the --help flag.

APIs and integrations

KoboldCpp serves many APIs alongside multiple bundled web UIs. Simply connect your software With the default port:

Interface Base URL
KoboldCpp Default API base http://localhost:5001
OpenAI-compatible API http://localhost:5001/v1
Interactive API documentation http://localhost:5001/api
KoboldAI Lite web UI http://localhost:5001
llama.cpp web UI http://localhost:5001/lcpp
StableUI Image Gen UI http://localhost:5001/sdui
MusicUI Music Gen UI http://localhost:5001/musicui

Additional APIs supported: KoboldAI, OpenAI, Anthropic, Ollama, AUTOMATIC1111, ComfyUI, XTTS

Image, speech, embedding and music generation require the corresponding models to be loaded. Replace the host and port when connecting to a remote server or using a custom --port. For other apps, select a compatible API type and point the app at your running KoboldCpp server.

KoboldCpp and KoboldAI API Documentation

Compiling KoboldCpp From Source Code

Use a source build if a prebuilt binary does not suit your platform or you want to develop KoboldCpp. Manual builds require Git, Python 3, a C/C++ toolchain, and the development libraries for your chosen GPU backend.

Optional Python runtime dependencies include customtkinter and Tk support for the GUI launcher, jinja2 for chat templates, and psutil for system information. Install the Python packages you need in your Python environment with python -m pip install -r requirements.txt; Tk may require a separate package from your operating system. The automated Linux build script manages its own dependencies.

Start by cloning the repository and entering its directory:

git clone https://github.com/LostRuins/koboldcpp.git
cd koboldcpp

Run the following build commands from that directory. Use make -jN to compile with N parallel jobs. For manual builds intended for other machines, add LLAMA_PORTABLE=1; this avoids optimizing only for the build machine, but platform and runtime requirements still apply.

Compiling on Linux

Automated build: koboldcpp.sh uses a local micromamba/conda environment to obtain dependencies and build the libraries. Install curl and bzip2 first.

./koboldcpp.sh             # Build as needed and open the launcher (requires X11)
./koboldcpp.sh --help      # Show command-line options
./koboldcpp.sh rebuild     # Refresh the environment and rebuild after updates
./koboldcpp.sh dist        # Package a standalone binary in dist/

To build for other machines, use KCPP_PORTABLE=1 ./koboldcpp.sh dist. The packaged binary still depends on the target system's compatibility with the Linux environment used to build it.

Manual build: Run make for a CPU build, or choose a backend below and install its prerequisites.

Backend Build command Prerequisite
CPU make C/C++ toolchain
Vulkan make LLAMA_VULKAN=1 Vulkan SDK
NVIDIA CUDA make LLAMA_CUBLAS=1 CUDA Toolkit
AMD ROCm make LLAMA_HIPBLAS=1 ROCm development libraries
CUDA and Vulkan make LLAMA_CUBLAS=1 LLAMA_VULKAN=1 Both toolkits

After building, launch with python3 koboldcpp.py --model /path/to/model.gguf.

Compiling on Windows

  1. Install the standard x64 version of w64devkit, not the i686 variant.
  2. Open its integrated terminal in the repository directory.
  3. Run make for a CPU build, or make LLAMA_VULKAN=1 for Vulkan. This produces the DLLs used by koboldcpp.py.
  4. Launch with python koboldcpp.py --model "C:\Models\model.gguf".

CUDA builds require Visual Studio, CMake, and the CUDA Toolkit. Open the project's CMake configuration in Visual Studio, build it, and copy koboldcpp_cublas.dll beside koboldcpp.py. The Makefile's LLAMA_CUBLAS=1 option is for Linux. Portable CUDA executables must include matching cublas, cublasLt, and cudart libraries from the same CUDA Toolkit family used for the build.

Packaging an executable: Install PyInstaller and the Python modules collected by make_pyinstaller.bat. Build the CPU and Vulkan libraries with make LLAMA_VULKAN=1 LLAMA_PORTABLE=1 to provide the script's required DLLs, then run the batch file from a Windows command prompt. It produces dist/koboldcpp-nocuda.exe; see the Windows release workflow for CUDA packaging.

If replacing bundled Vulkan libraries, put the matching .lib files in kcpp_src/lib and their .dll files in the repository root, then rebuild. This is an advanced configuration.

Compiling on macOS

Run make for a CPU build. For Metal GPU support, install the Apple command-line developer tools and build with:

make LLAMA_METAL=1
python3 koboldcpp.py --model /path/to/model.gguf --gpulayers -1

Compiling on OpenBSD

Install GNU Make with pkg_add gmake, then use gmake for a CPU build. For Vulkan, install vulkan-loader (also included as a dependency of vulkan-tools) and shaderc:

pkg_add gmake vulkan-loader shaderc
ulimit -d 8388608
gmake LLAMA_VULKAN=1
python3 koboldcpp.py --model /path/to/model.gguf

Run package installation with the required system privileges. The ulimit setting raises the data-size limit for compilation. If the build reports that ggml-vulkan-shaders.hpp is missing, check that glslc from shaderc is installed.

Compiling on Android (Termux installation)

Install Termux from F-Droid, then choose an automated or manual setup.

Automated setup: Download and run the Android installer. Its interactive menu offers installation with a starter model or without a model.

curl -sSL https://raw.githubusercontent.com/LostRuins/koboldcpp/concedo/android_install.sh | sh

Manual setup: In Termux, install the dependencies, clone the repository if you have not already done so, and build:

pkg update
pkg upgrade
pkg install openssl wget git python clang make
git clone https://github.com/LostRuins/koboldcpp.git
cd koboldcpp
make
python koboldcpp.py --model /path/to/model.gguf

Download a small GGUF model before the final command, and replace the model path with its location. Open localhost:5001 in your mobile browser once loading completes. For portable ARM builds, LLAMA_PORTABLE=1 disables native ARM instruction optimizations.

If package installation fails, update your packages or use termux-change-repo to choose another mirror. These instructions cover CPU builds; GPU acceleration depends on the device and its drivers.

Help and community

Start with the KoboldCpp FAQ and knowledge base, then search existing issues and discussions. If you still need help, open an issue or join the KoboldAI Discord.

For troubleshooting, include your operating system, hardware, KoboldCpp version, model filename, launch settings, and relevant error output.

Third Party Resources

These community projects may be outdated or unmaintained. Contact their maintainers for support.

License

KoboldCpp and KoboldAI Lite are licensed under the GNU AGPL v3.0, unless a file states otherwise. Bundled components retain their respective licenses, including the MIT license for GGML, llama.cpp, and stable-diffusion.cpp.

KoboldCpp builds on the work of these projects:

For enquiries, contact @concedo on discord, message u/HadesThrowaway on reddit, or find LostRuins on github.