openclaw/extensions/llama-cpp
Peter Steinberger 6d9efe2f5e
fix(setup): validate LM Studio and llama.cpp server URLs inline (#160720)
* fix(lmstudio): validate setup URLs before discovery

Reject malformed interactive endpoints in the provider-owned prompt instead
of sending Control UI and CLI operators into an unreachable-server retry.
Retain the authored draft and the existing HTTP(S) host-shorthand support.
Validate before the permissive path/default normalization can hide a missing
host. Reuse the existing host normalizer; no config or SDK surface changes.

The registered setup-entry regression fails against the original validator.
Real Gateway/Chromium proof rejects the invalid draft inline, then configures
the corrected local stub endpoint and sends chat through the selected model.
Owner and sibling coverage: 115 tests, 13.804 seconds wall; new case 2 ms.
Production delta +19 lines, tests +28, docs +3. The added code owns recoverable
prompt validation; existing prompt types move rather than being duplicated.

* fix(llama-cpp): validate server URLs in setup

Run the existing endpoint parser inside the provider-owned URL validator so
Control UI and CLI users can correct an invalid URL without losing setup.
Previously the parser threw after submission, closing the editable prompt
with a raw Invalid URL error. Preserve host shorthand and the existing
HTTP(S) and embedded-credential rules; no config or SDK surface changes.

The setup-entry regression fails against the original nonempty validator.
Real Gateway/Chromium proof keeps the draft with actionable inline guidance,
then activates the corrected local stub endpoint and sends chat through it.
Owner and sibling coverage: 115 tests, 13.804 seconds wall; new case 2 ms.
Production delta +10 lines, tests +29, docs +4. Growth converts the existing
throwing parser into a recoverable prompt result without duplicating policy.

* fix(lmstudio): reject credential-bearing setup URLs inline

* test(setup): use credential-free embedded-user URLs in validation fixtures

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-28 21:36:22 -07:00
..
assets improve(plugins): give bundled logos consistent white icon tiles (#155259) 2026-09-23 19:09:26 -07:00
src fix(setup): validate LM Studio and llama.cpp server URLs inline (#160720) 2026-09-28 21:36:22 -07:00
index.test.ts test(core,plugins): remove low-value tests (batch d025) (#158960) 2026-09-26 15:22:53 +00:00
index.ts
openclaw.plugin.json feat(plugins): assign one purpose category to every bundled plugin (#142760) 2026-09-10 20:44:20 -07:00
package.json chore(deps): update fs-safe to 0.21.2 (#160617) 2026-09-28 20:12:01 -07:00
provider-policy-api.test.ts fix(doctor): preserve upgrade settings and explain local memory setup (#150927) 2026-09-17 13:16:48 -07:00
provider-policy-api.ts
README.md fix(llama-cpp): preserve configured model preset settings (#139646) 2026-09-05 22:20:51 -07:00

@openclaw/llama-cpp-provider

Official llama.cpp provider for managed and external OpenClaw model servers.

The llama-cpp provider either installs a pinned, integrity-verified llama-server under OpenClaw's localService supervisor or connects to a server that you already operate. Both choices use llama-cpp/<model> references and OpenClaw's normal OpenAI-compatible chat transport. Local embeddings require the managed choice.

Install

openclaw plugins install @openclaw/llama-cpp-provider

Restart the Gateway after installing or updating the plugin. Interactive setup shows Managed local server and Existing llama-server under one Local llama.cpp group.

Configure managed text inference

After explicit consent, OpenClaw installs the matching server build and a recommended chat model that fits the Gateway's memory, GPU, and free disk space. The download also includes the configured local embedding model, or EmbeddingGemma by default (approximately 0.3 GB). See the provider guide for current recommendations.

When local memory search is configured and chat setup is unavailable or declined, OpenClaw offers a separate embedding-only setup. After explicit consent, it installs only the server and EmbeddingGemma. It leaves the current chat model unchanged. Move any llama.cpp chat routes and remove its configured chat model entries first. Remove an existing external server config before retrying embedding-only setup.

Custom GGUF models remain supported through params.modelPath. Rerun llama.cpp setup after changing the model so OpenClaw can verify the file and regenerate the managed router preset.

See the llama.cpp provider guide for platform requirements, custom GGUF configuration, diagnostics, and repair.

Connect to an existing server

Choose Existing llama-server during setup and enter the endpoint and optional API key. OpenClaw passively discovers single-model and router catalogs. It never installs, starts, stops, or reconfigures the external process.

See the llama.cpp provider guide for authentication, router behavior, manual configuration, and troubleshooting.

Configure embeddings

Set memory.search.provider to local. The plugin preserves the historical local embedding provider and index identity while serving requests through the managed server's /v1/embeddings endpoint.

Package

  • Plugin id: llama-cpp
  • Provider id: llama-cpp
  • Package: @openclaw/llama-cpp-provider
  • Minimum OpenClaw host: 2026.6.2