docs(providers): expand llmman guidance and add hybrid inference (#139606)

* docs(providers): expand llmman guidance and add hybrid inference

Rewrite docs/providers/llmman.md into a full provider topic page: modes
table, auth rules, getting started, model discovery, smoke tests, vision,
configuration tabs, recipes, advanced accordions, troubleshooting.

Add a Hybrid inference section covering llmman.hybrid/<local>,<provider>/<model>
refs (qwen3.8 + openai/gpt-5.6-luna), routing rules, key handling,
x-llmman-route pinning via provider headers, and how it composes with
OpenClaw fallbacks. Document llmman.provider/... hosted refs.

Use qwen3.8 as the reference model throughout, call `llmman serve` with no
arguments everywhere (daemon settings go in its environment), and adopt the
LLMMAN_API_KEY=llmman-local / apiKey: "${LLMMAN_API_KEY}" convention.

Cross-link from local-models, local-model-services, model-providers,
infer CLI, provider index, and the models FAQ. Add the new provider index
label to the zh-CN glossary.

Verified against llmman v0.1.334 with qwen3.8: text and vision probes via
openclaw infer model run, a full tool-calling agent turn, hybrid routing
(local by default, cloud on pin or oversized body), and
chat_template_kwargs passthrough for thinking control.

Co-authored-by: sallyom <11166065+sallyom@users.noreply.github.com>

* docs(llmman): correct verified runtime guidance

Clarify Qwen thinking compatibility, daemon-wide hybrid budgets, and idle service lifetime. Preserve the existing integration and live proof.

Co-authored-by: sallyom <11166065+sallyom@users.noreply.github.com>

* docs(llmman): align service startup reference

Keep the shared local-service example consistent with llmman optional model preloading.

Co-authored-by: sallyom <11166065+sallyom@users.noreply.github.com>

---------

Co-authored-by: sallyom <11166065+sallyom@users.noreply.github.com>
This commit is contained in:
Eric Curtin 2026-09-19 18:07:03 +01:00 • committed by GitHub
parent 1f96db3d8c
commit f786111973
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
9 changed files with 978 additions and 119 deletions

View file

@ -2899,6 +2899,10 @@
"source": "llmman (local models)",
"target": "llmman(本地模型)"
},
{
"source": "llmman (local + hybrid local/hosted models)",
"target": "llmman(本地 + 本地/托管混合模型)"
},
{
"source": "llama.cpp (managed or existing server)",
"target": "llama.cpp(托管或现有服务器)"

View file

@ -134,6 +134,8 @@ openclaw infer model run --local --model anthropic/claude-sonnet-4-6 --prompt "R
openclaw infer model run --local --model cerebras/zai-glm-4.7 --prompt "Reply with exactly: pong" --json
openclaw infer model run --local --model google/gemini-2.5-flash --prompt "Reply with exactly: pong" --json
openclaw infer model run --local --model groq/llama-3.1-8b-instant --prompt "Reply with exactly: pong" --json
openclaw infer model run --local --model llmman/qwen3.8 --prompt "Reply with exactly: pong" --json
openclaw infer model run --local --model llmman/gemma4:e4b --prompt "Describe this image." --file ./photo.jpg --json
openclaw infer model run --local --model mistral/mistral-medium-3-5 --prompt "Reply with exactly: pong" --json
openclaw infer model run --local --model mistral/mistral-small-latest --prompt "Reply with exactly: pong" --json
openclaw infer model run --local --model openai/gpt-5.6-luna --prompt "Reply with exactly: pong" --json
@ -197,6 +199,7 @@ Notes:
- Use `--timeout-ms` for slow local vision models or cold Ollama starts.
- For `image describe`, an explicit `--model` (must be an image-capable `<provider/model>`) runs first, then tries configured `agents.defaults.imageModel.fallbacks` if that call fails. Input-preparation errors (missing file, unsupported URL) fail before any fallback attempt, and the model must be image-capable in the model catalog or provider config.
- For local Ollama vision models, pull the model first and set `OLLAMA_API_KEY` to any placeholder value, for example `ollama-local`. See [Ollama](/providers/ollama#vision-and-image-description).
- For llmman vision models such as `llmman/gemma4:e4b`, configure the `llmman` provider with `input: ["text", "image"]` on the model entry and set `LLMMAN_API_KEY` to a placeholder such as `llmman-local`. See [llmman](/providers/llmman#vision-and-image-description).
## Audio

View file

@ -64,6 +64,7 @@ Each entry points at the page that now holds the content.
- <a id="synthetic" />[Synthetic](/concepts/model-providers/custom-providers#synthetic)
- <a id="minimax" />[MiniMax](/concepts/model-providers/custom-providers#minimax)
- <a id="llama.cpp" /><a id="llama-cpp" />[llama.cpp](/concepts/model-providers/custom-providers#llama-cpp)
- <a id="llmman" />[llmman](/concepts/model-providers/custom-providers#llmman)
- <a id="lm-studio" />[LM Studio](/concepts/model-providers/custom-providers#lm-studio)
- <a id="ollama" />[Ollama](/concepts/model-providers/custom-providers#ollama)
- <a id="vllm" />[vLLM](/concepts/model-providers/custom-providers#vllm)

View file

@ -236,6 +236,43 @@ openclaw plugins install @openclaw/llama-cpp-provider
Both use `llama-cpp/<model>` references. See [llama.cpp](/plugins/llama-cpp) for setup,
discovery, authentication, and managed local embeddings.
### llmman
llmman is configured via `models.providers` as an OpenAI-compatible local server. It pulls models as OCI artifacts and serves them through upstream `llama-server`, `vllm`, or `mlx-lm`, and can pair a local model with a hosted one under a single model id:
- Provider: `llmman` (custom; `api: "openai-completions"`)
- Auth: none enforced; set `LLMMAN_API_KEY=llmman-local` and use `apiKey: "${LLMMAN_API_KEY}"`
- Default base URL: `http://127.0.0.1:17434/v1`
- Example model: `llmman/qwen3.8`
- Hybrid example: `llmman/llmman.hybrid/qwen3.8,openai/gpt-5.6-luna`
```bash
llmman pull qwen3.8
llmman serve
```
```json5
{
agents: {
defaults: { model: { primary: "llmman/qwen3.8" } },
},
models: {
providers: {
llmman: {
baseUrl: "http://127.0.0.1:17434/v1",
apiKey: "${LLMMAN_API_KEY}",
api: "openai-completions",
models: [
{ id: "qwen3.8", name: "Qwen3.8 (llmman)", reasoning: true, input: ["text", "image"] },
],
},
},
},
}
```
See [/providers/llmman](/providers/llmman) for setup, hybrid local + hosted routing, vision, and troubleshooting.
### LM Studio
LM Studio ships as a bundled provider plugin which uses the native API:

View file

@ -92,13 +92,13 @@ During `memory_search`, managed embedding startup uses `readyTimeoutMs` instead
## llmman example
llmman is a custom OpenAI-compatible `/v1` backend, so the same `localService` API works with an `llmman` provider entry. It listens on `127.0.0.1:17434` by default; `LLMMAN_HOST` overrides the bind address, while `LLMMAN_LLM_LIBRARY` overrides GPU auto-detection. Its API has no authentication, so keep the default loopback bind unless a trusted network boundary restricts access.
llmman is a custom OpenAI-compatible `/v1` backend, so the same `localService` API works with an `llmman` provider entry. It listens on `127.0.0.1:17434` by default; `LLMMAN_HOST` overrides the bind address, while `LLMMAN_LLM_LIBRARY` overrides GPU auto-detection. Its API has no authentication, so keep the default loopback bind unless a trusted network boundary restricts access. llmman has no `/health` route; use `/v1/models` or `/api/version` as `healthUrl`.
```json5
{
agents: {
defaults: {
model: { primary: "llmman/gemma4" },
model: { primary: "llmman/qwen3.8" },
},
},
models: {
@ -106,12 +106,12 @@ llmman is a custom OpenAI-compatible `/v1` backend, so the same `localService` A
providers: {
llmman: {
baseUrl: "http://127.0.0.1:17434/v1",
apiKey: "llmman-local",
apiKey: "${LLMMAN_API_KEY}",
api: "openai-completions",
timeoutSeconds: 300,
localService: {
command: "/opt/homebrew/bin/llmman",
args: ["serve", "gemma4"],
args: ["serve"],
env: { LLMMAN_CONTEXT_LENGTH: "65536" },
healthUrl: "http://127.0.0.1:17434/v1/models",
readyTimeoutMs: 180000,
@ -119,13 +119,13 @@ llmman is a custom OpenAI-compatible `/v1` backend, so the same `localService` A
},
models: [
{
id: "gemma4",
name: "Gemma 4 (llmman)",
reasoning: false,
input: ["text"],
id: "qwen3.8",
name: "Qwen3.8 (llmman)",
reasoning: true,
input: ["text", "image"],
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
contextWindow: 65536,
maxTokens: 4096,
maxTokens: 8192,
},
],
},
@ -134,7 +134,7 @@ llmman is a custom OpenAI-compatible `/v1` backend, so the same `localService` A
}
```
Replace `command` with the result of `which llmman` on the machine running OpenClaw. Full llmman setup: [llmman](/providers/llmman).
Replace `command` with the result of `which llmman` on the machine running OpenClaw, and set `LLMMAN_API_KEY=llmman-local` in `~/.openclaw/.env`. Bare `llmman serve` requires no model argument and loads the model on the first request that names it; an optional model argument preloads it instead. Daemon settings such as `LLMMAN_CONTEXT_LENGTH` go in `env`. Full llmman setup, including hybrid local + hosted routing: [llmman](/providers/llmman).
## ds4 example
@ -182,6 +182,6 @@ Full setup, context sizing, and verification commands: [ds4](/providers/ds4).
Local model setup, provider choices, and safety guidance.
</Card>
<Card title="llmman" href="/providers/llmman" icon="cpu">
Run OpenClaw through the llmman OpenAI-compatible local server.
Local models, hybrid local + hosted routing, and on-demand startup with llmman.
</Card>
</CardGroup>

View file

@ -22,14 +22,15 @@ For custom servers, leave room for the full OpenClaw prompt, tools, history, and
## Pick a backend
| Backend | Use when |
| ---------------------------------------------------- | ---------------------------------------------------------------------------------- |
| [ds4](/providers/ds4) | Local DeepSeek V4 Flash on macOS Metal with OpenAI-compatible tool calls |
| LiteLLM / OAI-proxy / custom OpenAI-compatible proxy | You front another model API and need OpenClaw to treat it as OpenAI |
| [llama.cpp](/plugins/llama-cpp) | Hardware-aware model selection, verified downloads, and an OpenClaw-managed server |
| [LM Studio](/providers/lmstudio) | First-time local setup, GUI loader, native Responses API |
| MLX / vLLM / SGLang | High-throughput self-hosted serving with an OpenAI-compatible HTTP endpoint |
| [Ollama](/providers/ollama) | CLI workflow, model library, hands-off systemd service |
| Backend | Use when |
| ---------------------------------------------------- | -------------------------------------------------------------------------------------------- |
| [ds4](/providers/ds4) | Local DeepSeek V4 Flash on macOS Metal with OpenAI-compatible tool calls |
| LiteLLM / OAI-proxy / custom OpenAI-compatible proxy | You front another model API and need OpenClaw to treat it as OpenAI |
| [llama.cpp](/plugins/llama-cpp) | Hardware-aware model selection, verified downloads, and an OpenClaw-managed server |
| [llmman](/providers/llmman) | OCI-registry model pulls, upstream llama.cpp/vLLM/MLX engines, hybrid local + hosted routing |
| [LM Studio](/providers/lmstudio) | First-time local setup, GUI loader, native Responses API |
| MLX / vLLM / SGLang | High-throughput self-hosted serving with an OpenAI-compatible HTTP endpoint |
| [Ollama](/providers/ollama) | CLI workflow, model library, hands-off systemd service |
Use `api: "openai-responses"` when the backend supports it (LM Studio does). Otherwise use `api: "openai-completions"`. If `api` is omitted on a custom provider with a `baseUrl`, OpenClaw defaults to `openai-completions`.
@ -129,6 +130,8 @@ Setup checklist:
For local-first with a hosted safety net, swap `primary`/`fallbacks` order and keep the same `providers` block and `models.mode: "merge"`.
OpenClaw fallbacks switch models per turn on provider errors. For per-request routing that keeps small prompts on the local model and sends only oversized ones to a hosted model, see [Hybrid inference](/providers/llmman#hybrid-inference) with llmman.
### Regional hosting / data routing
Hosted MiniMax/Kimi/GLM variants also exist on OpenRouter with region-pinned endpoints (for example, US-hosted). Pick the regional variant to keep traffic in your chosen jurisdiction while keeping `models.mode: "merge"` for Anthropic/OpenAI fallbacks. Local-only is still the strongest privacy path. Hosted regional routing is the middle ground when you need provider features but want control over data flow.

View file

@ -80,11 +80,17 @@ troubleshooting, see the main [FAQ](/help/faq).
cloud models such as `kimi-k2.5:cloud` need no local pull. To switch
manually: `openclaw models list`, then `openclaw models set ollama/<model>`.
[llmman](/providers/llmman) is the alternative when you want models pulled
from OCI registries or Hugging Face, unmodified upstream `llama-server`,
`vllm`, or `mlx-lm` engines, or hybrid routing that keeps small requests on
a local model such as `qwen3.8` and overflows large ones to a hosted model.
Smaller/heavily quantized models are more vulnerable to prompt injection.
Use large models for any bot with tool access; if you use small models
anyway, enable sandboxing and strict tool allowlists.
Docs: [Ollama](/providers/ollama), [Local models](/gateway/local-models),
Docs: [Ollama](/providers/ollama), [llmman](/providers/llmman),
[Local models](/gateway/local-models),
[Model providers](/concepts/model-providers), [Security](/gateway/security),
[Sandboxing](/gateway/sandboxing).

View file

@ -53,7 +53,7 @@ Looking for chat channel docs (WhatsApp/Telegram/Discord/Slack/Mattermost (plugi
- [Kilocode](/providers/kilocode)
- [LiteLLM (unified gateway)](/providers/litellm)
- [llama.cpp (managed or existing server)](/plugins/llama-cpp)
- [llmman (local models)](/providers/llmman)
- [llmman (local + hybrid local/hosted models)](/providers/llmman)
- [LM Studio (local models)](/providers/lmstudio)
- [LongCat](/providers/longcat)
- [MiniMax](/providers/minimax)

File diff suppressed because it is too large Load diff