openclaw/docs/plugins/reference/llama-cpp.md
Patrick Erichsen 914530a957
fix(plugins): restore search icons and separate official results (#148518)
* fix(plugins): join official installations to their published listings

* fix(plugins): resolve published catalog icon paths

* feat(plugins): group mixed searches by official and community publishers

* refactor(plugins): reuse catalog field parsing and simplify test selection

* test(plugins): preserve install intent with published counterparts
2026-09-14 14:40:42 -07:00

1.4 KiB

summary read_when title
Managed and external llama.cpp servers for GGUF chat and embeddings.
You are installing, configuring, or auditing the llama-cpp plugin
Llama Cpp plugin reference

Managed and external llama.cpp servers for GGUF chat and embeddings.

Distribution

  • Package: @openclaw/llama-cpp-provider
  • Install route: npm or ClawHub: clawhub:@openclaw/llama-cpp-provider

Surface

  • Providers: llama-cpp
  • Contracts: embeddingProviders

Default text model

During interactive setup, OpenClaw installs a pinned, verified llama-server and offers Gemma 4 E4B IT Q4_K_M as an approximately 5.0 GB download. The model offer requires at least 16 GiB of total RAM. Existing cached models are still detected on smaller machines.

To use another model, set params.modelPath to any custom GGUF. Custom models are not subject to the bundled-download RAM requirement. On machines below the requirement, you can also run a smaller model through Ollama or LM Studio, or choose a cloud provider.