diff --git a/docs/.i18n/glossary.zh-CN.json b/docs/.i18n/glossary.zh-CN.json index e48b0a0eb31e..596e06f076ee 100644 --- a/docs/.i18n/glossary.zh-CN.json +++ b/docs/.i18n/glossary.zh-CN.json @@ -2899,6 +2899,10 @@ "source": "llmman (local models)", "target": "llmman(本地模型)" }, + { + "source": "llmman (local + hybrid local/hosted models)", + "target": "llmman(本地 + 本地/托管混合模型)" + }, { "source": "llama.cpp (managed or existing server)", "target": "llama.cpp(托管或现有服务器)" diff --git a/docs/cli/infer.md b/docs/cli/infer.md index 535a4a33fe6e..3b30c9d4546e 100644 --- a/docs/cli/infer.md +++ b/docs/cli/infer.md @@ -134,6 +134,8 @@ openclaw infer model run --local --model anthropic/claude-sonnet-4-6 --prompt "R openclaw infer model run --local --model cerebras/zai-glm-4.7 --prompt "Reply with exactly: pong" --json openclaw infer model run --local --model google/gemini-2.5-flash --prompt "Reply with exactly: pong" --json openclaw infer model run --local --model groq/llama-3.1-8b-instant --prompt "Reply with exactly: pong" --json +openclaw infer model run --local --model llmman/qwen3.8 --prompt "Reply with exactly: pong" --json +openclaw infer model run --local --model llmman/gemma4:e4b --prompt "Describe this image." --file ./photo.jpg --json openclaw infer model run --local --model mistral/mistral-medium-3-5 --prompt "Reply with exactly: pong" --json openclaw infer model run --local --model mistral/mistral-small-latest --prompt "Reply with exactly: pong" --json openclaw infer model run --local --model openai/gpt-5.6-luna --prompt "Reply with exactly: pong" --json @@ -197,6 +199,7 @@ Notes: - Use `--timeout-ms` for slow local vision models or cold Ollama starts. - For `image describe`, an explicit `--model` (must be an image-capable ``) runs first, then tries configured `agents.defaults.imageModel.fallbacks` if that call fails. Input-preparation errors (missing file, unsupported URL) fail before any fallback attempt, and the model must be image-capable in the model catalog or provider config. - For local Ollama vision models, pull the model first and set `OLLAMA_API_KEY` to any placeholder value, for example `ollama-local`. See [Ollama](/providers/ollama#vision-and-image-description). +- For llmman vision models such as `llmman/gemma4:e4b`, configure the `llmman` provider with `input: ["text", "image"]` on the model entry and set `LLMMAN_API_KEY` to a placeholder such as `llmman-local`. See [llmman](/providers/llmman#vision-and-image-description). ## Audio diff --git a/docs/concepts/model-providers.md b/docs/concepts/model-providers.md index 15be7c48c82e..5246fd41f1e8 100644 --- a/docs/concepts/model-providers.md +++ b/docs/concepts/model-providers.md @@ -64,6 +64,7 @@ Each entry points at the page that now holds the content. - [Synthetic](/concepts/model-providers/custom-providers#synthetic) - [MiniMax](/concepts/model-providers/custom-providers#minimax) - [llama.cpp](/concepts/model-providers/custom-providers#llama-cpp) +- [llmman](/concepts/model-providers/custom-providers#llmman) - [LM Studio](/concepts/model-providers/custom-providers#lm-studio) - [Ollama](/concepts/model-providers/custom-providers#ollama) - [vLLM](/concepts/model-providers/custom-providers#vllm) diff --git a/docs/concepts/model-providers/custom-providers.md b/docs/concepts/model-providers/custom-providers.md index 7e9d17634511..2bb3c114d0f5 100644 --- a/docs/concepts/model-providers/custom-providers.md +++ b/docs/concepts/model-providers/custom-providers.md @@ -236,6 +236,43 @@ openclaw plugins install @openclaw/llama-cpp-provider Both use `llama-cpp/` references. See [llama.cpp](/plugins/llama-cpp) for setup, discovery, authentication, and managed local embeddings. +### llmman + +llmman is configured via `models.providers` as an OpenAI-compatible local server. It pulls models as OCI artifacts and serves them through upstream `llama-server`, `vllm`, or `mlx-lm`, and can pair a local model with a hosted one under a single model id: + +- Provider: `llmman` (custom; `api: "openai-completions"`) +- Auth: none enforced; set `LLMMAN_API_KEY=llmman-local` and use `apiKey: "${LLMMAN_API_KEY}"` +- Default base URL: `http://127.0.0.1:17434/v1` +- Example model: `llmman/qwen3.8` +- Hybrid example: `llmman/llmman.hybrid/qwen3.8,openai/gpt-5.6-luna` + +```bash +llmman pull qwen3.8 +llmman serve +``` + +```json5 +{ + agents: { + defaults: { model: { primary: "llmman/qwen3.8" } }, + }, + models: { + providers: { + llmman: { + baseUrl: "http://127.0.0.1:17434/v1", + apiKey: "${LLMMAN_API_KEY}", + api: "openai-completions", + models: [ + { id: "qwen3.8", name: "Qwen3.8 (llmman)", reasoning: true, input: ["text", "image"] }, + ], + }, + }, + }, +} +``` + +See [/providers/llmman](/providers/llmman) for setup, hybrid local + hosted routing, vision, and troubleshooting. + ### LM Studio LM Studio ships as a bundled provider plugin which uses the native API: diff --git a/docs/gateway/local-model-services.md b/docs/gateway/local-model-services.md index 8cef47e2eab1..4e94f5bad12a 100644 --- a/docs/gateway/local-model-services.md +++ b/docs/gateway/local-model-services.md @@ -92,13 +92,13 @@ During `memory_search`, managed embedding startup uses `readyTimeoutMs` instead ## llmman example -llmman is a custom OpenAI-compatible `/v1` backend, so the same `localService` API works with an `llmman` provider entry. It listens on `127.0.0.1:17434` by default; `LLMMAN_HOST` overrides the bind address, while `LLMMAN_LLM_LIBRARY` overrides GPU auto-detection. Its API has no authentication, so keep the default loopback bind unless a trusted network boundary restricts access. +llmman is a custom OpenAI-compatible `/v1` backend, so the same `localService` API works with an `llmman` provider entry. It listens on `127.0.0.1:17434` by default; `LLMMAN_HOST` overrides the bind address, while `LLMMAN_LLM_LIBRARY` overrides GPU auto-detection. Its API has no authentication, so keep the default loopback bind unless a trusted network boundary restricts access. llmman has no `/health` route; use `/v1/models` or `/api/version` as `healthUrl`. ```json5 { agents: { defaults: { - model: { primary: "llmman/gemma4" }, + model: { primary: "llmman/qwen3.8" }, }, }, models: { @@ -106,12 +106,12 @@ llmman is a custom OpenAI-compatible `/v1` backend, so the same `localService` A providers: { llmman: { baseUrl: "http://127.0.0.1:17434/v1", - apiKey: "llmman-local", + apiKey: "${LLMMAN_API_KEY}", api: "openai-completions", timeoutSeconds: 300, localService: { command: "/opt/homebrew/bin/llmman", - args: ["serve", "gemma4"], + args: ["serve"], env: { LLMMAN_CONTEXT_LENGTH: "65536" }, healthUrl: "http://127.0.0.1:17434/v1/models", readyTimeoutMs: 180000, @@ -119,13 +119,13 @@ llmman is a custom OpenAI-compatible `/v1` backend, so the same `localService` A }, models: [ { - id: "gemma4", - name: "Gemma 4 (llmman)", - reasoning: false, - input: ["text"], + id: "qwen3.8", + name: "Qwen3.8 (llmman)", + reasoning: true, + input: ["text", "image"], cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }, contextWindow: 65536, - maxTokens: 4096, + maxTokens: 8192, }, ], }, @@ -134,7 +134,7 @@ llmman is a custom OpenAI-compatible `/v1` backend, so the same `localService` A } ``` -Replace `command` with the result of `which llmman` on the machine running OpenClaw. Full llmman setup: [llmman](/providers/llmman). +Replace `command` with the result of `which llmman` on the machine running OpenClaw, and set `LLMMAN_API_KEY=llmman-local` in `~/.openclaw/.env`. Bare `llmman serve` requires no model argument and loads the model on the first request that names it; an optional model argument preloads it instead. Daemon settings such as `LLMMAN_CONTEXT_LENGTH` go in `env`. Full llmman setup, including hybrid local + hosted routing: [llmman](/providers/llmman). ## ds4 example @@ -182,6 +182,6 @@ Full setup, context sizing, and verification commands: [ds4](/providers/ds4). Local model setup, provider choices, and safety guidance. - Run OpenClaw through the llmman OpenAI-compatible local server. + Local models, hybrid local + hosted routing, and on-demand startup with llmman. diff --git a/docs/gateway/local-models.md b/docs/gateway/local-models.md index 221c438ce213..086fa7707897 100644 --- a/docs/gateway/local-models.md +++ b/docs/gateway/local-models.md @@ -22,14 +22,15 @@ For custom servers, leave room for the full OpenClaw prompt, tools, history, and ## Pick a backend -| Backend | Use when | -| ---------------------------------------------------- | ---------------------------------------------------------------------------------- | -| [ds4](/providers/ds4) | Local DeepSeek V4 Flash on macOS Metal with OpenAI-compatible tool calls | -| LiteLLM / OAI-proxy / custom OpenAI-compatible proxy | You front another model API and need OpenClaw to treat it as OpenAI | -| [llama.cpp](/plugins/llama-cpp) | Hardware-aware model selection, verified downloads, and an OpenClaw-managed server | -| [LM Studio](/providers/lmstudio) | First-time local setup, GUI loader, native Responses API | -| MLX / vLLM / SGLang | High-throughput self-hosted serving with an OpenAI-compatible HTTP endpoint | -| [Ollama](/providers/ollama) | CLI workflow, model library, hands-off systemd service | +| Backend | Use when | +| ---------------------------------------------------- | -------------------------------------------------------------------------------------------- | +| [ds4](/providers/ds4) | Local DeepSeek V4 Flash on macOS Metal with OpenAI-compatible tool calls | +| LiteLLM / OAI-proxy / custom OpenAI-compatible proxy | You front another model API and need OpenClaw to treat it as OpenAI | +| [llama.cpp](/plugins/llama-cpp) | Hardware-aware model selection, verified downloads, and an OpenClaw-managed server | +| [llmman](/providers/llmman) | OCI-registry model pulls, upstream llama.cpp/vLLM/MLX engines, hybrid local + hosted routing | +| [LM Studio](/providers/lmstudio) | First-time local setup, GUI loader, native Responses API | +| MLX / vLLM / SGLang | High-throughput self-hosted serving with an OpenAI-compatible HTTP endpoint | +| [Ollama](/providers/ollama) | CLI workflow, model library, hands-off systemd service | Use `api: "openai-responses"` when the backend supports it (LM Studio does). Otherwise use `api: "openai-completions"`. If `api` is omitted on a custom provider with a `baseUrl`, OpenClaw defaults to `openai-completions`. @@ -129,6 +130,8 @@ Setup checklist: For local-first with a hosted safety net, swap `primary`/`fallbacks` order and keep the same `providers` block and `models.mode: "merge"`. +OpenClaw fallbacks switch models per turn on provider errors. For per-request routing that keeps small prompts on the local model and sends only oversized ones to a hosted model, see [Hybrid inference](/providers/llmman#hybrid-inference) with llmman. + ### Regional hosting / data routing Hosted MiniMax/Kimi/GLM variants also exist on OpenRouter with region-pinned endpoints (for example, US-hosted). Pick the regional variant to keep traffic in your chosen jurisdiction while keeping `models.mode: "merge"` for Anthropic/OpenAI fallbacks. Local-only is still the strongest privacy path. Hosted regional routing is the middle ground when you need provider features but want control over data flow. diff --git a/docs/help/faq-models.md b/docs/help/faq-models.md index e2db7c92b7d3..d02a33b4bd01 100644 --- a/docs/help/faq-models.md +++ b/docs/help/faq-models.md @@ -80,11 +80,17 @@ troubleshooting, see the main [FAQ](/help/faq). cloud models such as `kimi-k2.5:cloud` need no local pull. To switch manually: `openclaw models list`, then `openclaw models set ollama/`. + [llmman](/providers/llmman) is the alternative when you want models pulled + from OCI registries or Hugging Face, unmodified upstream `llama-server`, + `vllm`, or `mlx-lm` engines, or hybrid routing that keeps small requests on + a local model such as `qwen3.8` and overflows large ones to a hosted model. + Smaller/heavily quantized models are more vulnerable to prompt injection. Use large models for any bot with tool access; if you use small models anyway, enable sandboxing and strict tool allowlists. - Docs: [Ollama](/providers/ollama), [Local models](/gateway/local-models), + Docs: [Ollama](/providers/ollama), [llmman](/providers/llmman), + [Local models](/gateway/local-models), [Model providers](/concepts/model-providers), [Security](/gateway/security), [Sandboxing](/gateway/sandboxing). diff --git a/docs/providers/index.md b/docs/providers/index.md index 6b3a4633af22..fa10ace36bb6 100644 --- a/docs/providers/index.md +++ b/docs/providers/index.md @@ -53,7 +53,7 @@ Looking for chat channel docs (WhatsApp/Telegram/Discord/Slack/Mattermost (plugi - [Kilocode](/providers/kilocode) - [LiteLLM (unified gateway)](/providers/litellm) - [llama.cpp (managed or existing server)](/plugins/llama-cpp) -- [llmman (local models)](/providers/llmman) +- [llmman (local + hybrid local/hosted models)](/providers/llmman) - [LM Studio (local models)](/providers/lmstudio) - [LongCat](/providers/longcat) - [MiniMax](/providers/minimax) diff --git a/docs/providers/llmman.md b/docs/providers/llmman.md index 654e25de554e..7d1121f03481 100644 --- a/docs/providers/llmman.md +++ b/docs/providers/llmman.md @@ -1,70 +1,130 @@ --- -summary: "Run OpenClaw through llmman (OpenAI-compatible local server)" +summary: "Run OpenClaw with llmman (local models, hosted providers, and hybrid local + hosted routing)" read_when: - - You want to run OpenClaw against a local llmman server - - You are serving Gemma or another model through llmman - - You need the exact OpenClaw compat flags for llmman + - You want to run OpenClaw against local GGUF or safetensors models through llmman + - You want hybrid inference that keeps small requests local and overflows large ones to a hosted model + - You need llmman setup, configuration, vision, or troubleshooting guidance title: "llmman" --- -[llmman](https://github.com/llmmanorg/llmman) pulls GGUF/safetensors models from OCI registries and serves them behind Ollama-, OpenAI-, and Anthropic-compatible APIs. It uses `llama-server` for GGUF models and `vllm` or `mlx_lm.server` for safetensors models. OpenClaw talks to it through the generic `openai-completions` adapter. +[llmman](https://github.com/llmmanorg/llmman) pulls models as OCI artifacts from +any registry or Hugging Face and serves them through unmodified upstream +engines: `llama-server` for GGUF, `vllm` or `mlx-lm` for safetensors. One +daemon exposes Ollama-, OpenAI-, and Anthropic-compatible APIs and can also +forward requests to hosted providers. OpenClaw talks to it through the generic +`openai-completions` adapter on `/v1`. -| Property | Value | -| ---------------- | ------------------------------------------------------------ | -| Provider id | `llmman` (custom; configure under `models.providers.llmman`) | -| Plugin | none — not a bundled OpenClaw provider plugin | -| Auth env var | none required; any value works, `llmman serve` has no auth | -| API | OpenAI-compatible (`openai-completions`) | -| Default base URL | `http://127.0.0.1:17434/v1` | +Three modes are supported: - - `llmman` is a custom self-hosted OpenAI-compatible backend, not a dedicated OpenClaw provider plugin: you configure it under `models.providers.llmman` instead of picking an onboarding auth choice. For a bundled plugin with auto-discovery, see [SGLang](/providers/sglang) or [vLLM](/providers/vllm). - +| Mode | What it uses | +| ----------------- | ---------------------------------------------------------------------------------------------------------------- | +| Local only | `llmman serve` on the Gateway host or LAN, serving pulled models such as `qwen3.8` | +| Hybrid | One `llmman.hybrid/,/` ref; llmman picks the local or hosted side per request | +| Hosted via llmman | `llmman.provider//` refs; llmman forwards to a hosted provider while keeping one local endpoint | + +| Property | Value | +| ---------------- | ---------------------------------------------------------------------------------------- | +| Provider id | `llmman` (custom; configure under `models.providers.llmman`) | +| Plugin | none; not a bundled OpenClaw provider plugin, so models are listed explicitly | +| API | OpenAI-compatible (`api: "openai-completions"`) | +| Default base URL | `http://127.0.0.1:17434/v1` | +| Auth | llmman has no authentication; OpenClaw sends whatever `apiKey` you configure as a bearer | +| Reference model | `qwen3.8` (Qwen3.8 27B, vision-capable, 262,144-token native context) | + + +`llmman serve` has no authentication and no TLS. Keep the default loopback bind unless a trusted network boundary restricts access, and never expose it on a public interface. + - Version scope: this page is verified against [llmman b315](https://github.com/llmmanorg/llmman/releases/tag/b315), commit [`0e7a3ed`](https://github.com/llmmanorg/llmman/commit/0e7a3ed815d49a74d7aad1b1c70b5eb6c3013b18). +Version scope: this page is verified against [llmman v0.1.334](https://github.com/llmmanorg/llmman/releases/tag/v0.1.334), commit [`22b6d73`](https://github.com/llmmanorg/llmman/commit/22b6d7330e1b0882401e83e35fe592817eb7c816). Model names, hybrid routing, vision, and tool calling on this page were exercised against that build with `qwen3.8`. +## Auth rules + + + + `llmman serve` never checks credentials. OpenClaw still needs a non-empty `apiKey` on the provider entry so the provider counts as configured. This page uses `LLMMAN_API_KEY=llmman-local` with `apiKey: "${LLMMAN_API_KEY}"`, mirroring the `OLLAMA_API_KEY=ollama-local` convention. A literal `apiKey: "llmman-local"` works too. + + + For `llmman.hybrid/...` and `llmman.provider/...` refs, llmman forwards the bearer OpenClaw presents to the hosted provider as that provider's API key. The one exception is the literal placeholder `llmman`, which tells the daemon to use its own key. See [Hybrid inference](#hybrid-inference) for both patterns. A marker such as `llmman-local` would be sent to the hosted provider and rejected there. + + + A daemon started with `LLMMAN_HOST=0.0.0.0` in its environment binds every interface with no auth. Point OpenClaw at a LAN host only inside a network you trust; there is no credential OpenClaw can send that llmman would enforce. A daemon reachable off loopback also refuses to spend its own hosted-provider key for callers that presented none. + + + `${LLMMAN_API_KEY}` resolves from the Gateway process environment or `~/.openclaw/.env`. If the variable is missing, OpenClaw logs a config warning and treats the provider as unavailable, so put the export where the Gateway can read it. See [Environment](/help/environment). + + + ## Getting started - + ```bash - LLMMAN_CONTEXT_LENGTH=65536 llmman serve gemma4 + curl -fsSL https://llmmanorg.github.io/install.sh | sh # Linux, macOS + brew install llmmanorg/tap/llmman # Homebrew ``` - `llmman serve` listens on `127.0.0.1:17434` by default. Set `LLMMAN_HOST` before startup to override the bind address; there are no `--host`/`--port` flags. GPU acceleration (CUDA, ROCm, Vulkan, or Metal) is auto-detected; set `LLMMAN_LLM_LIBRARY` to override it because there is no `--device` flag. The model argument is optional — omit it to start the server and load models on the first request that names them instead. - - The example fixes the server context at 65,536 tokens and uses the same value in OpenClaw below. If you change `LLMMAN_CONTEXT_LENGTH`, keep the OpenClaw model's `contextWindow` at or below that value. + Windows: `irm https://llmmanorg.github.io/install.ps1 | iex` or `winget install llmmanorg.llmman`. Other options: `cargo binstall llmman`. - + + ```bash + llmman pull qwen3.8 + llmman serve + ``` + + Bare names such as `qwen3.8` resolve to Docker Hub's curated `docker.io/ai/:latest`. `owner/repo` names resolve to `hf.co/owner/repo`; a full reference (`ghcr.io/...`, `hf.co/...`) is used as-is. + + `llmman serve` requires no arguments (an optional model argument preloads that model): it listens on `127.0.0.1:17434` and loads any pulled model on the first request that names it, then unloads it after five idle minutes. GPU acceleration (CUDA, ROCm, Vulkan, Metal) is auto-detected and the matching `llama-server` is downloaded if none is on `PATH`. Everything else is tuned through the daemon's environment: `LLMMAN_HOST` for the bind address, `LLMMAN_CONTEXT_LENGTH` for the server context, `LLMMAN_KEEP_ALIVE` for the idle unload timer (see [Advanced configuration](#advanced-configuration)). + + By default llmman uses up to 262,144 tokens, capped to the model's trained context (262,144 for `qwen3.8`) and, on out-of-memory, retries with the context halved down to a 16,384 floor. `llmman ps` shows the context a loaded model actually got; keep the OpenClaw model's `contextWindow` at or below that value. + + + ```bash - curl http://127.0.0.1:17434/v1/models curl http://127.0.0.1:17434/api/version + curl http://127.0.0.1:17434/v1/models + llmman ps ``` - `llmman serve` has no dedicated `/health` route at the top level; use `/v1/models` or `/api/version` for a readiness probe. + There is no `/health` route; use `/api/version` or `/v1/models` as the readiness probe. `/v1/models` lists fully qualified ids such as `docker.io/ai/qwen3.8:latest`; requests may use either that form or the short name. - - Add an explicit provider entry and point your default model at it. See the config example below. + + Add to `~/.openclaw/.env` (or export in the Gateway's shell): + + ```bash + LLMMAN_API_KEY=llmman-local + ``` + + + + Add the config below, then: + + ```bash + openclaw models list --provider llmman + openclaw models set llmman/qwen3.8 + ``` + + +`llmman launch openclaw --model qwen3.8` can bootstrap a first-run OpenClaw install by running non-interactive onboarding against the llmman endpoint. It only applies when no `openclaw.json` exists yet; for an existing install use the explicit config on this page. + + ## Full config example -Gemma 4 on a local `llmman` server: +Qwen3.8 on a local llmman server: ```json5 { agents: { defaults: { - model: { primary: "llmman/gemma4" }, + model: { primary: "llmman/qwen3.8" }, models: { - "llmman/gemma4": { - alias: "Gemma 4 (llmman)", - }, + "llmman/qwen3.8": { alias: "Qwen3.8 (llmman)" }, }, }, }, @@ -73,17 +133,18 @@ Gemma 4 on a local `llmman` server: providers: { llmman: { baseUrl: "http://127.0.0.1:17434/v1", - apiKey: "llmman-local", + apiKey: "${LLMMAN_API_KEY}", api: "openai-completions", + timeoutSeconds: 300, models: [ { - id: "gemma4", - name: "Gemma 4 (llmman)", - reasoning: false, - input: ["text"], + id: "qwen3.8", + name: "Qwen3.8 (llmman)", + reasoning: true, + input: ["text", "image"], cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }, contextWindow: 65536, - maxTokens: 4096, + maxTokens: 8192, }, ], }, @@ -92,9 +153,202 @@ Gemma 4 on a local `llmman` server: } ``` -## On-demand startup +`models.mode: "merge"` keeps hosted providers available as fallbacks. `timeoutSeconds` gives cold model loads and long generations room before the model request timeout fires. -OpenClaw can start `llmman` itself only when an `llmman/...` model is selected. Add `localService` to the same provider entry: +## Model discovery + +llmman is not a bundled OpenClaw plugin, so there is no implicit discovery. List every model you want under `models.providers.llmman.models` with a provider-local `id` (no `llmman/` prefix). + +| Behavior | Detail | +| ------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| Model ids | Short (`qwen3.8`, `qwen3.5:9b`), Hugging Face (`unsloth/Qwen3.5-0.8B-GGUF`), or fully qualified (`docker.io/ai/qwen3.8:latest`) all work | +| Default tag | `:latest` when omitted | +| Capabilities | `llmman show ` and `POST /api/show` report `completion` plus `vision` when a companion `mmproj` projector is present; mark such models `input: ["text", "image"]` | +| Reasoning | Set `reasoning: true` for thinking models such as Qwen3.8; llmman returns thinking as `reasoning_content` | +| Context | `llmman show ` prints the trained context length; `llmman ps` prints the context the running server was started with | +| Costs | All `0`; the hosted half of a hybrid ref is billed by that provider, not tracked by OpenClaw | + +```bash +llmman list +llmman show qwen3.8 +openclaw models list --provider llmman +``` + +To add a model, pull it and add a matching entry: + +```bash +llmman pull qwen3.5:9b +``` + +### Smoke tests + +A narrow text probe that skips the full agent tool surface: + +```bash +LLMMAN_API_KEY=llmman-local \ + openclaw infer model run \ + --local \ + --model llmman/qwen3.8 \ + --prompt "Reply with exactly: pong" \ + --json +``` + +Add `--file` with an image for a lean vision-model probe (PNG/JPEG/WebP; +non-image files are rejected before llmman is called; use +`openclaw infer audio transcribe` for audio): + +```bash +LLMMAN_API_KEY=llmman-local \ + openclaw infer model run \ + --local \ + --model llmman/gemma4:e4b \ + --prompt "Describe this image in one sentence." \ + --file ./photo.jpg \ + --json +``` + +Neither path loads chat tools, memory, or session context. If a probe succeeds +while normal agent replies fail, the issue is usually tool-schema handling or +context pressure in the backend, not the endpoint; see +[Troubleshooting](#troubleshooting). + +A full agent turn with tool calling is the real test: + +```bash +openclaw agent --local --session-id llmman-smoke \ + --message "Read the file ./README.md with a tool and summarize it in one sentence." +``` + +## Hybrid inference + +llmman can pair a local model with a hosted one under a single model name and +choose a side per request. OpenClaw configures the pair once as an ordinary +model id and gets local-first inference with hosted overflow, without an +agent-level fallback switch. + +The reference is `llmman.hybrid/,/`. With `qwen3.8` +as the local half and OpenAI's `gpt-5.6-luna` as the hosted half: + +```text +llmman.hybrid/qwen3.8,openai/gpt-5.6-luna +``` + +Which side serves a request: + +1. **`x-llmman-route: local` or `cloud`** request header wins. Any other value is a `400`. +2. **Otherwise, size.** A request body larger than the local context can hold goes to the hosted model. The budget is 4 bytes per token of `LLMMAN_CONTEXT_LENGTH`; `LLMMAN_HYBRID_LOCAL_BYTES` sets it directly and `0` disables the size rule. +3. **Otherwise, local.** + +If a request llmman kept local is then refused by the local backend as larger +than its context, llmman resends it to the hosted half before anything reaches +OpenClaw. A `local` pin is never overridden this way. Every routed request is +logged with the side and the reason. The raw completion response may report +the backend model or GGUF path; OpenClaw's result envelope retains the +configured pair ref. + +### Hosted-provider key + +The hosted half authenticates like any llmman `--provider` request. Pick one: + + + + Set the llmman provider's `apiKey` to the hosted provider's key. llmman forwards it per request and never persists it. Local-only models on the same provider entry ignore the value. + + ```bash + # ~/.openclaw/.env + LLMMAN_API_KEY=sk-... # your OpenAI API key + ``` + + Keep `apiKey: "${LLMMAN_API_KEY}"` in the config below. + + + + Give the daemon its own key and have OpenClaw send the literal placeholder `llmman`: + + ```bash + llmman config set providers.openai.api_key sk-... + llmman serve + ``` + + `OPENAI_API_KEY` in the daemon's environment works too and overrides the config file. + + ```bash + # ~/.openclaw/.env + LLMMAN_API_KEY=llmman + ``` + + The placeholder must be exactly `llmman`; any other bearer is treated as a real key and forwarded to the hosted provider. llmman only spends its own key for loopback callers, so this pattern requires the daemon and the Gateway on the same host. + + + + +### Hybrid config + +```json5 +{ + agents: { + defaults: { + model: { primary: "llmman/llmman.hybrid/qwen3.8,openai/gpt-5.6-luna" }, + models: { + "llmman/llmman.hybrid/qwen3.8,openai/gpt-5.6-luna": { alias: "Qwen3.8 + Luna" }, + "llmman/qwen3.8": { alias: "Qwen3.8 local" }, + }, + }, + }, + models: { + mode: "merge", + providers: { + llmman: { + baseUrl: "http://127.0.0.1:17434/v1", + apiKey: "${LLMMAN_API_KEY}", + api: "openai-completions", + timeoutSeconds: 300, + models: [ + { + id: "llmman.hybrid/qwen3.8,openai/gpt-5.6-luna", + name: "Qwen3.8 + GPT-5.6 Luna (hybrid)", + reasoning: true, + input: ["text", "image"], + cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }, + contextWindow: 65536, + maxTokens: 8192, + }, + { + id: "qwen3.8", + name: "Qwen3.8 (llmman)", + reasoning: true, + input: ["text", "image"], + cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }, + contextWindow: 65536, + maxTokens: 8192, + }, + ], + }, + }, + }, +} +``` + +The daemon computes one hybrid byte budget at startup: 4 bytes per token of +`LLMMAN_CONTEXT_LENGTH`, or `262144 × 4 = 1048576` bytes when unset. This budget +does not follow a model's trained-context cap or a later out-of-memory +reduction. Set `LLMMAN_CONTEXT_LENGTH=65536` in the daemon's environment to +match the example, or set `LLMMAN_HYBRID_LOCAL_BYTES` to pick the byte budget +directly. An explicit `LLMMAN_CONTEXT_LENGTH=0` disables size-based routing +unless a positive `LLMMAN_HYBRID_LOCAL_BYTES` supplies a budget; setting the +byte override to `0` also disables that rule. Local context-refusal fallback +still applies unless the request is pinned local. + +Set `contextWindow` on the pair to the local model's usable token context. OpenClaw then compacts +around the local model's limit, so most turns stay local; llmman still +overflows to `gpt-5.6-luna` when a request exceeds it. Set it to the hosted +model's window instead if you prefer fewer compactions and more hosted +traffic. + +### Pinning a side + +OpenClaw sends provider-level `headers` on every request, so a second provider +entry on the same base URL can force one side of the pair: ```json5 { @@ -102,26 +356,90 @@ OpenClaw can start `llmman` itself only when an `llmman/...` model is selected. providers: { llmman: { baseUrl: "http://127.0.0.1:17434/v1", - apiKey: "llmman-local", + apiKey: "${LLMMAN_API_KEY}", api: "openai-completions", - timeoutSeconds: 300, - localService: { - command: "/opt/homebrew/bin/llmman", - args: ["serve", "gemma4"], - env: { LLMMAN_CONTEXT_LENGTH: "65536" }, - healthUrl: "http://127.0.0.1:17434/v1/models", - readyTimeoutMs: 180000, - idleStopMs: 0, - }, models: [ { - id: "gemma4", - name: "Gemma 4 (llmman)", - reasoning: false, - input: ["text"], - cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }, - contextWindow: 65536, - maxTokens: 4096, + id: "llmman.hybrid/qwen3.8,openai/gpt-5.6-luna", + name: "Hybrid", + input: ["text", "image"], + }, + ], + }, + "llmman-cloud": { + baseUrl: "http://127.0.0.1:17434/v1", + apiKey: "${LLMMAN_API_KEY}", + api: "openai-completions", + headers: { "x-llmman-route": "cloud" }, + models: [ + { + id: "llmman.hybrid/qwen3.8,openai/gpt-5.6-luna", + name: "Hybrid (hosted)", + input: ["text", "image"], + }, + ], + }, + }, + }, + agents: { + defaults: { + model: { + primary: "llmman/llmman.hybrid/qwen3.8,openai/gpt-5.6-luna", + fallbacks: ["llmman-cloud/llmman.hybrid/qwen3.8,openai/gpt-5.6-luna"], + }, + }, + }, +} +``` + +Switch with `/model llmman-cloud/...` for a turn that should go hosted, or use +`headers: { "x-llmman-route": "local" }` for an entry that must never leave +the machine. + +### Hybrid versus OpenClaw fallbacks + +| Mechanism | Decides | Switches on | +| --------------------------------- | -------------------------- | -------------------------------------------------------------- | +| llmman hybrid pair | Per request, inside llmman | Request size versus local context, or an explicit route header | +| `agents.defaults.model.fallbacks` | Per turn, inside OpenClaw | Provider errors, timeouts, rate limits, auth failures | + +They compose. A common shape is a hybrid pair as `primary` with a direct hosted +model as a fallback for when llmman itself is down: + +```json5 +{ + agents: { + defaults: { + model: { + primary: "llmman/llmman.hybrid/qwen3.8,openai/gpt-5.6-luna", + fallbacks: ["openai/gpt-5.6-luna"], + }, + }, + }, +} +``` + +### Hosted models through llmman + +`llmman.provider//` forwards to a hosted provider with no +local half. Use it when you want every model, local or hosted, behind one +endpoint and one key-handling story: + +```json5 +{ + models: { + providers: { + llmman: { + baseUrl: "http://127.0.0.1:17434/v1", + apiKey: "${LLMMAN_API_KEY}", + api: "openai-completions", + models: [ + { id: "qwen3.8", name: "Qwen3.8 (llmman)", reasoning: true, input: ["text", "image"] }, + { + id: "llmman.provider/openai/gpt-5.6-luna", + name: "GPT-5.6 Luna via llmman", + reasoning: true, + input: ["text", "image"], }, ], }, @@ -130,61 +448,510 @@ OpenClaw can start `llmman` itself only when an `llmman/...` model is selected. } ``` -`command` must be an absolute path. Run `which llmman` on the Gateway host and use that path. Full field reference: [Local model services](/gateway/local-model-services). +The provider catalog comes from [models.dev](https://models.dev) and is cached +by the daemon. `llmman providers` shows which providers have a key; +`llmman list --provider openai` lists that provider's models and prices. For +direct hosted access without llmman in the path, configure the +[OpenAI](/providers/openai) provider instead. + +## Vision and image description + +Models that ship a companion `mmproj` projector are vision-capable; `qwen3.8` +and `gemma4:e4b` are. `llmman show ` logs `found companion mmproj file` +and `/api/show` reports a `vision` capability. Mark those models +`input: ["text", "image"]` so image attachments are injected into agent turns. + +```bash +llmman pull gemma4:e4b +openclaw infer image describe --file ./photo.jpg --model llmman/gemma4:e4b --json +``` + +`--model` must be a full `` ref. Use `infer image describe` +for OpenClaw's image-understanding flow and configured `imageModel`; use +`infer model run --file` for a raw multimodal probe with a custom prompt. + +To make a llmman model the default image-understanding provider for inbound +media: + +```json5 +{ + agents: { + defaults: { + imageModel: { + primary: "llmman/gemma4:e4b", + }, + }, + }, + models: { + providers: { + llmman: { + baseUrl: "http://127.0.0.1:17434/v1", + apiKey: "${LLMMAN_API_KEY}", + api: "openai-completions", + models: [{ id: "gemma4:e4b", name: "Gemma 4 E4B (llmman)", input: ["text", "image"] }], + }, + }, + }, + tools: { + media: { + image: { + timeoutSeconds: 180, + }, + }, + }, +} +``` + +OpenClaw rejects image-description requests for models not marked +image-capable. Slow local vision models can need a longer image-understanding +timeout than hosted models; `models.providers.llmman.timeoutSeconds` still +governs the underlying HTTP request for normal model calls. + +## Configuration + + + + ```json5 + { + models: { + providers: { + llmman: { + baseUrl: "http://127.0.0.1:17434/v1", + apiKey: "${LLMMAN_API_KEY}", + api: "openai-completions", + timeoutSeconds: 300, + models: [ + { + id: "qwen3.8", + name: "Qwen3.8 (llmman)", + reasoning: true, + input: ["text", "image"], + contextWindow: 65536, + maxTokens: 8192, + }, + ], + }, + }, + }, + } + ``` + + + + + Start the server on the GPU box with `LLMMAN_HOST=0.0.0.0` (and optionally `LLMMAN_CONTEXT_LENGTH=65536`) in its environment, then point OpenClaw at it: + + ```bash + llmman serve + ``` + + ```json5 + { + models: { + providers: { + llmman: { + baseUrl: "http://gpu-box.local:17434/v1", + apiKey: "${LLMMAN_API_KEY}", + api: "openai-completions", + timeoutSeconds: 420, + models: [ + { + id: "qwen3.8", + name: "Qwen3.8 (gpu-box)", + reasoning: true, + input: ["text", "image"], + contextWindow: 65536, + maxTokens: 8192, + }, + ], + }, + }, + }, + } + ``` + + + The remote daemon has no authentication. Only do this inside a trusted network, and do not use hybrid refs with "llmman holds the key" against a non-loopback daemon; it refuses to spend its own key for remote callers. + + + + + + OpenClaw starts llmman on demand when a `llmman/...` model is requested. This example keeps the daemon running until OpenClaw exits (`idleStopMs: 0`). Set a positive `idleStopMs` to stop an OpenClaw-started daemon after that many idle milliseconds; this is separate from llmman unloading idle models: + + ```json5 + { + models: { + providers: { + llmman: { + baseUrl: "http://127.0.0.1:17434/v1", + apiKey: "${LLMMAN_API_KEY}", + api: "openai-completions", + timeoutSeconds: 300, + localService: { + command: "/opt/homebrew/bin/llmman", + args: ["serve"], + env: { LLMMAN_CONTEXT_LENGTH: "65536" }, + healthUrl: "http://127.0.0.1:17434/v1/models", + readyTimeoutMs: 180000, + idleStopMs: 0, + }, + models: [ + { + id: "qwen3.8", + name: "Qwen3.8 (llmman)", + reasoning: true, + input: ["text", "image"], + contextWindow: 65536, + maxTokens: 8192, + }, + ], + }, + }, + }, + } + ``` + + `command` must be an absolute path; use the output of `which llmman` on the Gateway host. `env` is how the daemon's settings such as `LLMMAN_CONTEXT_LENGTH` are supplied here. `healthUrl` must be `/v1/models` or `/api/version`, since llmman has no `/health`. The first request after startup also loads the model, so keep `timeoutSeconds` generous. Full field reference: [Local model services](/gateway/local-model-services). + + + + +## Common recipes + +Replace model ids with names from `llmman list` or +`openclaw models list --provider llmman`. + + + + ```bash + llmman pull qwen3.8 + llmman serve + echo 'LLMMAN_API_KEY=llmman-local' >> ~/.openclaw/.env + openclaw models set llmman/qwen3.8 + ``` + + Use the [Full config example](#full-config-example) for the provider entry, and set `LLMMAN_CONTEXT_LENGTH=65536` in the daemon's environment if you want the server context to match it exactly. + + + + + The [Hybrid config](#hybrid-config) above: `llmman.hybrid/qwen3.8,openai/gpt-5.6-luna` as `primary`, `LLMMAN_API_KEY` set to the OpenAI key, `LLMMAN_CONTEXT_LENGTH` matching the pair's `contextWindow`. Add `openai/gpt-5.6-luna` to `fallbacks` so a stopped daemon does not block replies. + + + + Local models served through a custom `openai-completions` provider do not enable [Tool Search](/tools/tool-search) automatically. Turn it on to keep optional capabilities available while loading their schemas only when needed, and cap the context to what the host can run with `LLMMAN_CONTEXT_LENGTH=32768` in the daemon's environment: + + ```json5 + { + agents: { + defaults: { + model: { primary: "llmman/qwen3.8" }, + }, + }, + tools: { + toolSearch: { mode: "tools" }, + }, + models: { + providers: { + llmman: { + baseUrl: "http://127.0.0.1:17434/v1", + apiKey: "${LLMMAN_API_KEY}", + api: "openai-completions", + models: [ + { + id: "qwen3.8", + name: "Qwen3.8 (llmman)", + reasoning: true, + input: ["text", "image"], + contextWindow: 32768, + contextTokens: 32768, + maxTokens: 4096, + }, + ], + }, + }, + }, + } + ``` + + Use `compat.supportsTools: false` only when the model or server reliably fails on tool schemas; it disables tool use entirely. For a deliberately narrower agent, prefer `tools.profile` or a per-agent tool policy. + + + + + Custom provider ids when running more than one daemon; each gets its own host, models, and timeout: + + ```json5 + { + models: { + providers: { + "llmman-fast": { + baseUrl: "http://mini.local:17434/v1", + apiKey: "${LLMMAN_API_KEY}", + api: "openai-completions", + models: [{ id: "qwen3.5:9b", name: "qwen3.5:9b", input: ["text"], contextWindow: 32768 }], + }, + "llmman-large": { + baseUrl: "http://gpu-box.local:17434/v1", + apiKey: "${LLMMAN_API_KEY}", + api: "openai-completions", + timeoutSeconds: 420, + models: [{ id: "qwen3.8", name: "qwen3.8", reasoning: true, input: ["text", "image"], contextWindow: 131072 }], + }, + }, + }, + agents: { + defaults: { + model: { + primary: "llmman-fast/qwen3.5:9b", + fallbacks: ["llmman-large/qwen3.8"], + }, + }, + }, + } + ``` + + llmman can also pool several daemons itself: `llmman config set aggregation.peers ,` (or `LLMMAN_PEERS`) makes one daemon forward to peers, so OpenClaw sees a single endpoint whose model list spans the group. + + + + + Model ids are whatever llmman resolves. Pull with the full reference and use the same string as the OpenClaw model id: + + ```bash + llmman pull hf.co/unsloth/Qwen3.5-0.8B-GGUF + llmman pull ghcr.io/myorg/private-model:v1 + ``` + + ```json5 + { + models: { + providers: { + llmman: { + baseUrl: "http://127.0.0.1:17434/v1", + apiKey: "${LLMMAN_API_KEY}", + api: "openai-completions", + models: [ + { id: "hf.co/unsloth/Qwen3.5-0.8B-GGUF", name: "Qwen3.5 0.8B", input: ["text"] }, + { id: "ghcr.io/myorg/private-model:v1", name: "Private model", input: ["text"] }, + ], + }, + }, + }, + } + ``` + + The agent ref is then `llmman/hf.co/unsloth/Qwen3.5-0.8B-GGUF`. Run `llmman login ` first for private registries. + + + + +### Model selection + +```json5 +{ + agents: { + defaults: { + model: { + primary: "llmman/qwen3.8", + fallbacks: ["llmman/qwen3.5:9b", "openai/gpt-5.6-luna"], + }, + }, + }, +} +``` + +For slow local models, prefer provider-scoped tuning before raising the whole +agent runtime timeout: `models.providers.llmman.timeoutSeconds` covers +connection setup, headers, body streaming, and the total guarded-fetch abort +for that provider's model requests only. + +### Quick verification + +```bash +# llmman daemon visible to this machine +curl http://127.0.0.1:17434/api/version +llmman ps + +# OpenClaw catalog and selected model +openclaw models list --provider llmman +openclaw models status + +# Direct model smoke +LLMMAN_API_KEY=llmman-local openclaw infer model run \ + --local --model llmman/qwen3.8 --prompt "Reply with exactly: ok" --json +``` + +For remote hosts, replace `127.0.0.1` with the `baseUrl` host. If `curl` works +but OpenClaw does not, check whether the Gateway runs on a different machine, +container, or service account. ## Advanced configuration - - `llmman` resolves and loads the requested model, rewrites its id for the selected backend, and adds generation defaults such as `repeat_penalty`. It forwards message content and tool schemas without normalizing them, so compatibility for those fields depends on the selected backend and model. + + `LLMMAN_CONTEXT_LENGTH` is the server-side context (there is no flag). Semantics by backend: - - If OpenClaw runs fail with: + - **llama-server (GGUF):** set, it is passed as `--ctx-size` for generation models; `0` means the trained context. Unset, llmman uses `262144` or the model's trained context if smaller, and on out-of-memory retries with the context halved down to a 16,384 floor. + - **vLLM (safetensors):** a positive value becomes `--max-model-len`; unset uses the vLLM default. + - **mlx_lm.server:** not forwarded. - ```text - messages[1].content: invalid type: sequence, expected a string - ``` + `LLMMAN_NUM_PARALLEL` scales `--ctx-size` up by that factor so each slot keeps the full context. - set `compat.requiresStringContent: true` in the model entry. OpenClaw then flattens pure text content parts into plain strings before sending the request. - - - - - - If a model accepts small direct `/v1/chat/completions` requests but fails on full OpenClaw agent-runtime turns, try disabling the tool schema surface first: + On the OpenClaw side, `contextWindow` declares the model's window and `contextTokens` caps active input. Keep `contextWindow` at or below the server value; OpenClaw derives compaction and preflight thresholds from it. OpenClaw's `contextWindow` does not change llmman's hybrid byte budget; configure the daemon separately as described in [Hybrid config](#hybrid-config). ```json5 - compat: { - supportsTools: false + { + models: { + providers: { + llmman: { + models: [ + { + id: "qwen3.8", + name: "qwen3.8", + contextWindow: 65536, + contextTokens: 49152, + maxTokens: 8192, + }, + ], + }, + }, + }, } ``` - That reduces prompt pressure on stricter local backends. If tiny direct requests still work but normal OpenClaw agent turns keep crashing inside `llama-server`, treat it as an upstream model/server limitation rather than an OpenClaw transport issue. + + + + Qwen3.8 thinks by default; llmman returns the reasoning as `reasoning_content`, which OpenClaw's `openai-completions` adapter separates from the final text. Requests are proxied to `llama-server`, so `chat_template_kwargs` passes through. To turn thinking off for agent turns with a local Qwen model: + + ```json5 + { + agents: { + defaults: { + models: { + "llmman/qwen3.8": { + params: { + chat_template_kwargs: { enable_thinking: false }, + }, + }, + }, + }, + }, + } + ``` + + For per-run or session control, declare the local Qwen model's thinking format: + + ```json5 + { + models: { + providers: { + llmman: { + models: [ + { + id: "qwen3.8", + name: "Qwen3.8 (llmman)", + reasoning: true, + input: ["text", "image"], + compat: { thinkingFormat: "qwen-chat-template" }, + }, + ], + }, + }, + }, + } + ``` + + With this declaration, `openclaw agent --model llmman/qwen3.8 --thinking off`, `/think off`, and `openclaw infer model run --local --model llmman/qwen3.8 --thinking off --prompt "Reply with exactly: pong" --json` map the thinking setting to `chat_template_kwargs.enable_thinking`. Without it, the generic proxy defaults do not send this control or `reasoning_effort`. + + The lean `infer model run` path does not read the agent-level `params` recipe above; use the compatibility declaration and `--thinking off` for that probe. Do not combine a fixed `enable_thinking` agent param with per-run control, since the fixed param overrides the generated value. Apply Qwen-specific controls to a hybrid ref only if both its local and hosted backends accept them. - - Test both layers once configured: + + Models load on their first request, each in its own backend subprocess, and unload after five idle minutes. Tune with: + + | Variable | Meaning | + | -------------------------- | ---------------------------------------------------------------------------------------- | + | `LLMMAN_KEEP_ALIVE` | Idle unload timer (default `5m`); `0` unloads right after each request | + | `LLMMAN_MAX_LOADED_MODELS` | Cap on concurrently loaded models; idle ones are evicted LRU, busy ones return `503` | + | `LLMMAN_MAX_QUEUE` | Pending-request cap before `503` (default `512`) | + | `LLMMAN_LOAD_TIMEOUT` | Load stall timeout (default `10m`) | + + `llmman ps` shows loaded models with their context and expiry; `llmman stop ` unloads one now. A first request after startup or an idle unload pays the load cost, so set `timeoutSeconds` on the provider and raise `LLMMAN_KEEP_ALIVE` to keep the daemon warm for chat surfaces. + + + + + llmman probes CUDA, ROCm, Vulkan (Linux/Windows) or Metal (macOS) and downloads a matching `llama-server` release if none is on `PATH`. Override with `LLMMAN_LLM_LIBRARY`: `cpu`, `cuda`, `cuda13`, `rocm`, `vulkan`, or `metal`. Other knobs: `LLMMAN_FLASH_ATTENTION` (`on`/`off`/`auto`), `LLMMAN_KV_CACHE_TYPE` (`f16`, `q8_0`, `q4_0`), `LLMMAN_SCHED_SPREAD` for multi-GPU layer splitting, `LLMMAN_IGPU_ENABLE` to count integrated GPUs. On Linux, `llmman serve --ociman docker|podman` runs `llama-server` from the `ghcr.io/ggml-org/llama.cpp` images instead of a local binary. `LLMMAN_DEBUG=1` prints the probe result. + + + + llmman serves `/v1/embeddings` for GGUF embedding models, so [memory search](/concepts/memory) can use it through the generic `openai-compatible` embedding provider: ```bash - curl http://127.0.0.1:17434/v1/chat/completions \ - -H 'content-type: application/json' \ - -d '{"model":"gemma4","messages":[{"role":"user","content":"What is 2 + 2?"}],"stream":false}' + llmman pull embeddinggemma ``` - ```bash - openclaw infer model run \ - --model llmman/gemma4 \ - --prompt "What is 2 + 2? Reply with one short sentence." \ - --json + ```json5 + { + memory: { + search: { + provider: "openai-compatible", + model: "embeddinggemma", + remote: { + baseUrl: "http://127.0.0.1:17434/v1", + apiKey: "${LLMMAN_API_KEY}", + }, + }, + }, + } ``` - If the first command works but the second fails, see Troubleshooting below. + Embedding models are capped to their trained context regardless of `LLMMAN_CONTEXT_LENGTH`. See [Memory config](/reference/memory-config#remote-endpoint-config) for the remaining fields. + + + + + llmman also implements Ollama's native `/api/chat`, `/api/tags`, `/api/show`, and `/api/ps`, and `OLLAMA_HOST=127.0.0.1:17434 ollama run ` works against it. Prefer `api: "openai-completions"` from OpenClaw anyway: llmman's `/api/show` reports only `completion` and `vision` capabilities, so the bundled Ollama plugin's discovery would mark every llmman model `compat.supportsTools: false`. If you do point the Ollama plugin at `http://127.0.0.1:17434` (no `/v1`), list models explicitly instead of relying on discovery. + + + + llmman forwards message content and tool schemas to the backend without normalizing them, so compatibility depends on the selected engine and model. Structured content parts (text + image) and OpenClaw's full tool schema work with `qwen3.8` on `llama-server`. If a different backend or model rejects them: + + - `messages[].content: invalid type: sequence, expected a string` → set `compat.requiresStringContent: true` on the model entry. OpenClaw then flattens pure text content parts into plain strings. + - `400 JSON schema conversion failed` → `llama-server` could not compile a tool schema into its grammar subset. Update OpenClaw first; if a third-party tool or MCP server contributes the offending schema, disable it for that agent, and use `compat.supportsTools: false` only as a last resort. + + ```json5 + { + models: { + providers: { + llmman: { + models: [ + { + id: "qwen3.8", + name: "qwen3.8", + compat: { + requiresStringContent: true, + }, + }, + ], + }, + }, + }, + } + ``` - Because `llmman` uses the generic `openai-completions` adapter (not `openai-responses`), native-OpenAI-only request shaping never applies: no `service_tier`, no Responses `store`, no prompt-cache hints, and no OpenAI reasoning-compat payload shaping get sent. + Because llmman is a non-native `openai-completions` endpoint, OpenClaw treats it as a proxy route: no `service_tier`, no Responses `store`, no prompt-cache hints, no OpenAI reasoning-compat payload shaping, no hidden OpenClaw attribution headers, and `compat.supportsDeveloperRole` is forced to `false`. Vendor-specific fields can be merged into the request body with `agents.defaults.models["llmman/"].params.extra_body`. + + + + Local llmman models are free, so set all costs to `0`. The hosted half of a hybrid or `llmman.provider/...` ref is billed by that provider; `llmman list --provider openai` shows its per-million-token prices. @@ -193,28 +960,60 @@ OpenClaw can start `llmman` itself only when an `llmman/...` model is selected. `llmman serve` is not running or is not reachable at the configured address. The default is `127.0.0.1:17434`; if you set `LLMMAN_HOST`, update the OpenClaw `baseUrl` and `healthUrl` to match. + + ```bash + llmman serve + curl http://127.0.0.1:17434/api/version + ``` + - - Set `compat.requiresStringContent: true` in the model entry (see above). + + `models.providers.llmman.apiKey: *** env var "LLMMAN_API_KEY"` in the config warnings means the substitution found no value. Add `LLMMAN_API_KEY=llmman-local` to `~/.openclaw/.env`, or replace `"${LLMMAN_API_KEY}"` with the literal `"llmman-local"`. + + + + The model is not pulled, or the id in config does not match what llmman resolves. Compare against `llmman list`; a short name maps to `docker.io/ai/:latest`, and an `owner/repo` name maps to `hf.co/owner/repo`. + + ```bash + llmman pull qwen3.8 + llmman list + ``` + + + + + Large models can take minutes to load, especially on the first request after an idle unload. Raise `models.providers.llmman.timeoutSeconds`, warm the model with a first `openclaw infer model run`, and consider a longer `LLMMAN_KEEP_ALIVE` on the daemon. + + + + llmman routed the request to the hosted half and found no usable key. Either the bearer OpenClaw sent was a marker (`llmman-local`) rather than a real key, or you used the `llmman` placeholder without giving the daemon its own `OPENAI_API_KEY`, or the daemon is bound off loopback and refuses to spend its own key. See [Hosted-provider key](#hosted-provider-key). + + + + Check the daemon log; every routed request records the side and the reason. Routing is by request body size against `LLMMAN_CONTEXT_LENGTH` (4 bytes per token) unless `x-llmman-route` is set. Lower `LLMMAN_HYBRID_LOCAL_BYTES` to overflow sooner, raise it to stay local longer, or pin a side with a provider `headers` entry as in [Pinning a side](#pinning-a-side). - Both probes are tool-free, so `compat.supportsTools` cannot change this failure. Check the configured base URL and model id, inspect the `llmman`/backend logs, and compare the two request payloads and responses. + Both probes are tool-free, so `compat.supportsTools` cannot change this failure. Check the configured base URL, model id, and `LLMMAN_API_KEY`, inspect the daemon and backend logs, and compare the two request payloads. - The agent turn includes a larger prompt and may include tool schemas. Try `compat.supportsTools: false` to isolate tool-schema pressure (see the tool-schema caveat above). + The agent turn adds a larger prompt and tool schemas. A `400 JSON schema conversion failed` is `llama-server` rejecting a tool schema; update OpenClaw and check third-party tools or MCP servers. Otherwise enable [Tool Search](/tools/tool-search) to defer schemas, confirm the server's actual context allocation, and use `compat.supportsTools: false` only as a last resort. See [Smaller or stricter backends](/gateway/local-models#smaller-or-stricter-backends). - - If schema errors are gone but the spawned `llama-server` still crashes on larger agent turns, treat it as an upstream `llama.cpp` or model limitation. Reduce prompt pressure or switch backend/model. + + If schema errors are gone but the spawned `llama-server` still crashes on larger turns, treat it as an upstream `llama.cpp` or model limitation. Lower `LLMMAN_CONTEXT_LENGTH`, set `LLMMAN_KV_CACHE_TYPE=q8_0` to reduce memory, or switch the backend or model. + + + + Confirm the model's chat template supports tool calling and that the request reached `/v1/chat/completions` (not the Ollama or Anthropic surfaces via a proxy). If the model only calls tools when forced, set `params.extra_body.tool_choice: "required"` on that model ref as described in [Local models](/gateway/local-models#other-openai-compatible-local-proxies). - -For general help, see [Troubleshooting](/help/troubleshooting) and [FAQ](/help/faq). - + +More help: [Troubleshooting](/help/troubleshooting) and [FAQ](/help/faq). + ## Related @@ -225,10 +1024,16 @@ For general help, see [Troubleshooting](/help/troubleshooting) and [FAQ](/help/f Starting local model servers on demand for configured providers. + + Direct hosted access to the models used as the hybrid overflow half. + + + `openclaw infer model run` and the other one-shot probes used on this page. + + + Overview of all providers, model refs, and failover behavior. + Debugging local OpenAI-compatible backends that pass probes but fail agent runs. - - Overview of all providers, model refs, and failover behavior. -