openclaw/docs/providers/nvidia.md
Vincent Koc 71b21c8857
docs: fix 20 concrete gateway/provider defects from the ux audit (#143983)
* docs: fix concrete gateway/provider setup defects from the ux audit

Closes 20 `ux`-kind audit findings in docs/gateway/ and docs/providers/
that named an objectively checkable defect: a missing prerequisite, a step
with no command, an unexplained placeholder, or a silent flag divergence.

- gateway/1password: the token-file block reads and then unsets
  $OP_SERVICE_ACCOUNT_TOKEN, but nothing told the reader to set it, so a
  pasted block wrote an empty token file (r3-1654).
- gateway/prometheus: "Restart the Gateway" had no restart command (r3-1618).
- gateway/doctor/config-migrations: "fix those keys by hand against the
  current config reference" had no link (r3-1530).
- gateway/bonjour: dropped an undefined acronym ("OCM"), defined nowhere in
  docs/ (r3-1641).
- providers/litellm: `litellm --model claude-opus-4-6` needs the upstream
  provider key, named only in a collapsed accordion 110 lines below (r3-2201).
- providers/vllm: step 1 said to start a vLLM server but gave no command
  (r3-2132).
- providers/vercel-ai-gateway: install step omitted `openclaw gateway
  restart`, which the next step needs (r3-2190).
- providers/synthetic: "Verify the default model" had no verify command
  (r3-2205).
- providers/gmi: the non-interactive hint named one flag, no command
  (r5-0059).
- providers/stepfun: both steps titled "Non-interactive alternative" omitted
  `--non-interactive` and `--accept-risk` (r3-2167).
- providers/cerebras, fireworks, meta, tencent: two non-interactive commands
  per page differed only by `--mode local`, unexplained; `--mode` defaults to
  `local` (docs/cli/onboard.md) (r3-2153, r3-2151, r3-2155, r3-2188).
- providers/xiaomi: TTS example used `apiKey: "xiaomi_api_key"`, drifting from
  the canonical docs/tools/tts/configuration.md (r3-2136).
- providers/nvidia: the custom provider id read as a reserved value (r3-2145).
- providers/baseten: the daemon-env note stated the problem without the fix
  the other eight copies give (r3-2150).
- providers/groq: a Step code fence was unindented out of its container
  (r3-2194).
- providers/fal, zai: neither page said where the API key comes from
  (r3-2142, r3-2139).

* docs: keep provider onboarding on the Gateway host (ClawSweeper P2)

Remote-client onboarding writes gateway.remote connection settings only:
runNonInteractiveRemoteSetup returns before local provider setup and never
consumes the provider credential (src/commands/onboard-non-interactive/
remote.ts), and docs/start/wizard-cli-reference.md:434 already says so.

The --mode remote alternative I added to cerebras, fireworks, meta and
tencent would therefore have left the provider unconfigured. Replaced it
with the Gateway-host instruction on all four pages; the --mode local
equivalence statement, which was the point of the finding, stays.
2026-09-10 19:56:13 +08:00

9.8 KiB

summary read_when title
Use NVIDIA's OpenAI-compatible API in OpenClaw
You want to use open models in OpenClaw for free
You need NVIDIA_API_KEY setup
You want to use Nemotron 3 Ultra through NVIDIA
NVIDIA

NVIDIA serves open models for free through an OpenAI-compatible API at https://integrate.api.nvidia.com/v1, authenticated with an API key from build.nvidia.com. OpenClaw defaults the NVIDIA provider to Nemotron 3 Ultra, NVIDIA's 550B total / 55B active reasoning model for long-context agentic work.

Getting started

Create an API key at [build.nvidia.com](https://build.nvidia.com/settings/api-keys). ```bash export NVIDIA_API_KEY="nvapi-..." openclaw onboard --auth-choice nvidia-api-key ``` ```bash openclaw models set nvidia/nvidia/nemotron-3-ultra-550b-a55b ```

For non-interactive setup, pass the key directly:

openclaw onboard --auth-choice nvidia-api-key --nvidia-api-key "nvapi-..."
`--nvidia-api-key` lands the key in shell history and `ps` output. Prefer the `NVIDIA_API_KEY` environment variable when possible.

Config example

{
  env: { vars: { NVIDIA_API_KEY: "nvapi-..." } },
  models: {
    providers: {
      nvidia: {
        baseUrl: "https://integrate.api.nvidia.com/v1",
        api: "openai-completions",
      },
    },
  },
  agents: {
    defaults: {
      model: { primary: "nvidia/nvidia/nemotron-3-ultra-550b-a55b" },
    },
  },
}

Live model catalog

When an NVIDIA API key is configured, setup and model-selection paths check https://integrate.api.nvidia.com/v1/models for available model IDs, cached for 30 seconds. NVIDIA's public https://assets.ngc.nvidia.com/products/api-catalog/featured-models.json feed provides ranking and token limits, cached for 24 hours. Featured models appear first only while the inference inventory still lists them; other available bundled chat models follow. A fresh inventory can restore a previously hidden model that NVIDIA has republished.

The inventory also contains embeddings and other non-chat endpoints, without capability metadata. OpenClaw therefore offers only exact models with bundled chat metadata or valid featured-model metadata; it does not guess capabilities from model names. Unknown IDs can still be configured explicitly; listing alone does not prove chat compatibility. This is not a complete automatic catalog of every NVIDIA model.

Both public fetches use fixed HTTPS hosts and send no credentials. A failed inventory or featured request marks discovery unavailable and retains the last successful catalog for the same provider configuration and credentials. Failed featured metadata cannot silently remove previously discovered models. A successful empty inventory clears discovered models, even if the featured feed fails. Without NVIDIA auth, browsing uses the bundled catalog without fetching.

Nemotron 3.5 Lightning

nvidia/nemotron-3.5-lightning-30b-a3b is NVIDIA's smaller 30B total / 3B active reasoning model for agentic work. The bundled row records its 1M context and a 16,384-token output budget matching NVIDIA's hosted example. Select it with:

openclaw models set nvidia/nvidia/nemotron-3.5-lightning-30b-a3b

Lightning is selectable when the live inventory lists it even if it is absent from the featured feed. Nemotron 3 Ultra remains the default.

Nemotron 3 Ultra

Nemotron 3 Ultra is the default NVIDIA model in OpenClaw. NVIDIA's build page for nvidia/nemotron-3-ultra-550b-a55b lists it as an available free endpoint with a 1M-token context specification.

The bundled Ultra row sends chat_template_kwargs: { enable_thinking: false, force_nonempty_content: true } by default so normal chat output stays in the visible answer instead of exposing reasoning text.

Use Ultra for the highest-capability NVIDIA default. Select Nemotron 3.5 Lightning or Nemotron 3 Super when you want a smaller Nemotron option, or choose one of the third-party models hosted in NVIDIA's catalog when their context, latency, or behavior fits better.

Bundled fallback catalog

The bundled rows provide known chat metadata and an offline fallback. Deprecated compatibility rows keep existing exact model references recognizable but stay out of model pickers.

Model ref Name Context Max output
nvidia/nvidia/nemotron-3-ultra-550b-a55b Nemotron 3 Ultra 550B 1,048,576 8,192
nvidia/nvidia/nemotron-3.5-lightning-30b-a3b Nemotron 3.5 Lightning 30B 1,048,576 16,384
nvidia/nvidia/nemotron-3-super-120b-a12b Nemotron 3 Super 120B 1,000,000 8,192
nvidia/z-ai/glm-5.2 GLM 5.2 202,752 8,192
nvidia/moonshotai/kimi-k2.6 Kimi K2.6 262,144 65,536
nvidia/minimaxai/minimax-m3 Minimax M3 196,608 8,192
nvidia/deepseek-ai/deepseek-v4-pro DeepSeek V4 Pro 262,144 16,384

The full compatibility catalog also retains these shipped refs for existing configurations and migration: nvidia/qwen/qwen3.5-397b-a17b, nvidia/moonshotai/kimi-k2.5, nvidia/z-ai/glm-5.1, nvidia/z-ai/glm5, and nvidia/minimaxai/minimax-m2.7. These references stay hidden from bundled and offline model pickers unless NVIDIA republishes them in its inference inventory. NVIDIA has retired the Qwen endpoint, so requests using its model reference no longer work. Migrate existing Qwen configurations to an active model.

Advanced configuration

The provider auto-enables when the `NVIDIA_API_KEY` environment variable is set or a key was stored during onboarding. No explicit provider config is required beyond the key. OpenClaw uses NVIDIA's inference inventory for availability and its featured feed for ranking. Exact bundled metadata preserves reasoning and image capabilities omitted by the featured feed. Deprecated exact-reference compatibility rows stay hidden from the offline fallback; fresh inventory can restore models that NVIDIA has republished. Costs default to `0` in source since NVIDIA currently offers free API access for the listed models. OpenClaw talks to NVIDIA with the `openai-completions` adapter against the standard `/v1` chat completions route. Any OpenAI-compatible tooling should work out of the box with the NVIDIA base URL. NVIDIA's Ultra sample request uses `chat_template_kwargs.enable_thinking` and `reasoning_budget` for reasoning output. OpenClaw's bundled Ultra row disables template thinking by default for normal chat use. If you need to opt into NVIDIA reasoning output or force other NVIDIA-specific request fields, set per-model params and keep provider-specific overrides scoped to the NVIDIA model:
```json5
{
  agents: {
    defaults: {
      models: {
        "nvidia/nvidia/nemotron-3-ultra-550b-a55b": {
          params: {
            chat_template_kwargs: { enable_thinking: true },
            extra_body: { reasoning_budget: 16384 },
          },
        },
      },
    },
  },
}
```

`params.chat_template_kwargs` merges into any `chat_template_kwargs`
already on the request instead of replacing the whole object.
`params.extra_body` is the final OpenAI-compatible request-body override
and overwrites colliding payload keys, so use it only for fields NVIDIA
documents for the selected endpoint.
Some NVIDIA-hosted custom models can take longer than the default ~120s model idle watchdog before they emit a first response chunk. For custom NVIDIA provider entries, raise the provider timeout instead of the whole agent runtime timeout; `timeoutSeconds` covers provider HTTP requests and raises the idle/stream watchdog ceiling for that provider. The provider id below (`custom-integrate-api-nvidia-com`) is a name you choose, not a reserved value; any id works as long as your model refs use the same prefix:
```json5
{
  models: {
    providers: {
      "custom-integrate-api-nvidia-com": {
        baseUrl: "https://integrate.api.nvidia.com/v1",
        api: "openai-completions",
        apiKey: "NVIDIA_API_KEY",
        timeoutSeconds: 300,
      },
    },
  },
  agents: {
    defaults: {
      models: {
        "custom-integrate-api-nvidia-com/meta/llama-3.1-70b-instruct": {
          params: { thinking: "off" },
        },
      },
    },
  },
}
```
NVIDIA models are currently free to use. Check [build.nvidia.com](https://build.nvidia.com/) for the latest availability and rate-limit details. Choosing providers, model refs, and failover behavior. Full config reference for agents, models, and providers.