openclaw/docs/reference/token-use.md
Ayaan Zaidi 504d2905d6
fix(models): apply downloaded catalogs without a Gateway restart (#158000)
Related: #140086, #141885, #156535, #157838

## What Problem This Solves

A running Gateway doesn't see a newly downloaded hosted model catalog until it restarts. This PR publishes each accepted catalog through the existing prepared-runtime owner, without a restart. Model rows and their prices switch together as one generation, which keeps the invariant from #140086.

## User Impact

- Compatible downloads are adopted at the Gateway's background catalog check, or after an explicit `models.list` refresh. That refresh returns the currently accepted rows right away and runs adoption afterward.
- A turn admitted on catalog N keeps N's rows and prices until it finishes. New turns use N+1. Rows and prices are never mixed.
- A concurrent auth or config publication no longer postpones adoption to the next scheduled check (up to 6 h). Adoption waits for that publication to settle, then retries, up to 3 attempts.
- An owner whose build failed or timed out ends the adoption instead of waiting on unbounded work. Gateway shutdown cancels an adoption that is still preparing.
- Malformed, schema-invalid, too-new (`minVersion`) and older catalogs are rejected, and the previously accepted catalog stays in use.
- Changing `models.catalogRefresh.url` no longer needs a restart: the previous source's catalog stops applying, and the mirror's catalog is adopted at the next catalog check.

**Bad-catalog exposure:** with live apply, a *valid but wrong* published catalog reaches running Gateways at their next catalog check (at most every 6 h) or on the next explicit `models.list` refresh. It no longer waits for a restart. Recovery uses existing mechanisms only:
- Republish a corrected catalog with a newer `generatedAt`; Gateways adopt it the same way.
- Operators can set `models.catalogRefresh.enabled: false`, which withdraws remote rows and prices without a restart (covered by the Gateway integration test).

This PR adds no new kill switch, config option or env knob.

### Compatibility

No config keys, defaults, types, validation, stored rows, protocol or SDK contracts change. The only config-surface change is the `models.catalogRefresh.url` help text, which drops the stale "Changes apply after a Gateway restart" sentence, and its regenerated config-doc baseline hash. Existing configs validate unchanged and need no Doctor migration (maintainer confirmation: https://github.com/openclaw/openclaw/pull/158000#issuecomment-5913462783). Startup behavior is unchanged. Upgrade impact for existing installs: an accepted download activates at the next catalog check instead of the next restart.

## Why This Change Was Made

- `prepared-model-runtime.configured-refresh.ts` builds a complete candidate generation of the configured owners under the new catalog. One serialized commit then publishes rows, the accepted bundle, the pricing context and the reply-dispatch projection together.
- Adoption re-reads the stored catalog until the config it read under is still current, so a stale caller can't cancel a current adoption.
- Each preparation attempt has its own abort signal. A config advance restarts only the attempt; a newer catalog or shutdown ends the whole adoption, including pricing preparation.
- Between attempts, adoption waits only on publication gates: a pending replacement or an owner's pending publication.
- Adopted owners install the same plugin-retirement recovery as configured publication (#161267). After commit, a lost Gateway plugin loan republishes them through the normal recovery. Before commit, it restarts the adoption attempt, and the commit refuses any candidate whose plugin generation retired.

### Why downloads were restart-only, and what this keeps

Restart-only activation was a mechanism, not the goal. #140086 chose it to stop rows and prices from different catalog versions mixing, and #157838 was merged as "the prerequisite for applying new remote catalogs without a Gateway restart (rows and prices must switch together)". This PR is that follow-up. Every requirement those PRs set still holds:

| Original requirement | Source | How it holds here |
|---|---|---|
| Rows and prices from one catalog version; never mixed across reloads or new requests | #140086 | One serialized commit publishes owners, bundle, pricing context and dispatch; pricing contexts are keyed by the exact accepted catalog. Integration test: new rows appear only with new prices |
| Admitted work keeps its pair | #140086, #157838 | Runs carry their plugin generation's catalog; usage operations capture one pricing context. Integration test and live proof: the in-flight turn keeps the old price |
| Startup absence is a real state (no downloaded rows without prices) | #140086 | Absence → catalog goes through the same atomic commit; overlay absence tests unchanged |
| Worker replacement inherits the host's accepted pair, not a later download | #140086 | The commit updates the inherited pair; later workers and a worker-exit recovery keep it (overlay and integration tests) |
| Current enablement and source URL still gate eligibility | #140086, #156535 | Checked on every read and before adoption, including the default-install v1 fallback; disablement withdraws rows and prices together (integration test) |
| Bad or superseded downloads never replace the active pair | #140086, #141885 | Compatibility, `minVersion`, revision and `generatedAt` checks; stale reads can't cancel a current adoption (regression test) |
| Failed catalog checks retry at the remaining fresh interval, not a full TTL | #141885 | Unchanged scheduler behavior; the deleted notice test's retry case is restored for failed adoption (fails if the retry falls back to the full TTL) |
| Operators learn when a downloaded catalog is not yet active | #141885 | No longer needed: downloads activate at the next check. The restart notice and its tests are removed; `models refresh` says when a running Gateway applies the update |
| Billing-route prices switch with their rows | #156535 | `upstreamPricing` and `providerPricing` are part of the accepted catalog pair |

## Evidence

**Regressions.** Each fails with its fix reverted and passes with it:
- *Retries a scheduled adoption when its pending auth owner settles.* Runs through the real Gateway update scheduler. Reverted, it logs `remote model catalog check superseded; deferred to the next check`.
- *Does not let a read under a superseded config cancel the current adoption.* Reverted, both calls end `superseded`.
- *Ends adoption instead of joining a timed-out owner build.* Reverted, adoption never settles.
- *Does not hold Gateway shutdown on an adoption's pricing preparation.* Reverted, shutdown waits on the held preparation until the test times out.
- *Recovers adopted owners when their borrowed Gateway plugin retires after commit / before commit.* Without the recovery, both fail: `Prepared model runtime plugin generation retired` and `prepared reply dispatch runtime owner was not published`.
- *Uses the remaining stored TTL after a fresh startup check when adoption fails.* With the retry reverted to the full TTL, the second check doesn't run.

**Suites:**

| Suite | Result |
|---|---|
| `prepared-model-runtime.remote-publication.test.ts` | 10/10 |
| Gateway integration (`models-list.remote-catalog`) | v1 and v2 pass. Config and auth churn during preparation end `published` on the settled owners. Also covers retained admitted runs, rejected and stale bundles, worker replacement and disablement |
| `prepared-model-runtime*`, `server-plugin-reload*`, `update-startup`, and all PR-touched test files | pass |
| `tsgo:core`, all `tsgo:test:src` shards | pass |
| oxlint and oxfmt on changed files; `config:docs:check`, `config:schema:check`; max-lines, assertion-safety and test-timeout-race ratchets | pass |

Tests wait on owned completion signals (`withinTest`), not wall-clock deadlines.

**Live proof** on an isolated Gateway built from `cb8c9197fd` (no provider mocks). Later commits add plugin-retirement recovery for adopted owners, covered by the regression tests above, and rebases onto `main`. It used a real OpenAI key through `openai/gpt-4.1-mini`, and the build stamp was set before the real catalog's publication date. A client polled `models.list` back to back over one WebSocket for the whole run (1023 polls, no errors). One Gateway process (PID unchanged) and no restart:
1. The stored catalog was seeded with an older revision of the real `catalog.openclaw.ai` v2 catalog: generated 2 days earlier, `gpt-4.1-mini` priced ×10, plus one extra kimi row. `models.list` listed the extra row, and a turn priced **$4.00 / $16.00 per M** input/output.
2. `openclaw models refresh` downloaded the real catalog (`updated`, 1039 models). The listed rows didn't change for the next 7.1 s, and a turn in that window still priced **$4.00 / $16.00 per M**: a download stays inactive until the Gateway adopts it.
3. A long turn was admitted on the older catalog, then `models.list {refresh:true}` returned the older rows (extra kimi row still listed) and started adoption. The new catalog was visible 0.9 s later, while the long turn was still running: the extra kimi row was gone.
4. The in-flight turn finished at **$4.00 / $16.00 per M** (older rows and prices). The next turn priced **$0.40 / $1.60 per M** (real catalog).
5. `models.catalogRefresh.url` was moved to a local mirror of the real catalog through `config.patch`, and the mirror's catalog was adopted without a restart. The mirror then served malformed JSON: `models refresh` failed with `SyntaxError`, the model list was unchanged, and the next turn still priced $0.40 / $1.60 per M.

**Model picker during republication (also on `main`).** Right after the new generation commits, `models.list` shows the new generation's configured and static rows until its full catalog loads, then the full list. In the live run this lasted 109 ms. The same poller against a `main` build shows the same window after a `models.*` config reload (20 → 6 → 16 rows for about 350 ms), so this PR adds a new trigger for an existing behavior. It doesn't change it. The short list comes from the new generation, so rows and prices stay paired.

**Published-driver upgrade cells.** Candidate tarball built from a fresh clone at `6f682c7944` with the canonical Docker packaging script and no build-time overrides. sha256 `46d881a0…ef0ad`; embedded commit `6f682c7944`, version 2026.9.7. `6f682c7944` already includes the shutdown-cancellation and pricing-deadline commits. The current head the current head differs from it only by rebases onto `main`: `main` had moved the scheduled catalog check into `update-startup-catalog.ts`, and this PR's adoption call moved there unchanged (`git range-diff` shows no other production change; a `remoteCatalog: null` test-fixture field moved to main's relocated `cli-compaction.test-support.ts`).

| Driver → candidate | Scenario | Result |
|---|---|---|
| `openclaw@2026.9.6` | base | passed (930 s); updater outcome success, no recovery |
| `openclaw@2026.9.6` | plugin-deps-cleanup | passed (923 s); updater outcome success, no recovery |
| `openclaw@2026.9.7` (latest) | base | passed (813 s); updater outcome success, no recovery |

In every cell:
- Migration, post-Doctor config validation, survival, plugin-dependency cleanup and runtime-deps repair checks passed.
- The candidate Gateway logged ready, then its catalog check fetched and saved the hosted catalog about 0.2 s after starting. The hosted catalog is older than the candidate's build stamp, so adoption ends `unchanged`, which isn't logged. After a 300 s settlement window, `/readyz` (`ready:true`, nothing failing) and a Gateway `status` RPC passed, and the Gateway shut down cleanly.
- To show a logged terminal outcome, each cell was repeated with a loopback mirror serving a newer copy of the same catalog. Each check logged `remote model catalog applied` about 0.26 s after it started, followed by `/readyz` and `status` passing.

Not run: `openclaw@2026.9.7` plugin-deps-cleanup. At the previous head it failed inside the 2026.9.7 driver's retained-runtime verification (`Retained runtime entry does not reference its inventoried file: dist/a2ui-…mjs`), identically for a merge-base control package with none of this PR's commits.

Harness note for the 2026.9.7 cell: unchanged, the upgrade harness can't run against 2026.9.7. It seeds the retired `tools.toolSearch {mode:"code"}` setting, which 2026.9.7 rejects. The 2026.9.7 cell used the harness's existing Tool Search "absent" mode, a one-line local change that skips only that seed and its check. The 2026.9.6 cells used the unchanged harness.

**CI:** fully green on the final head ([run 36818367970](https://github.com/openclaw/openclaw/actions/runs/36818367970)). Earlier heads hit failures that reproduce on `main` in code this PR doesn't touch: `update-cli.target-schema` (main reproduction: https://github.com/openclaw/openclaw/pull/158000#issuecomment-5913464830), the type-suppression inventory, the Windows partition owner test and the Windows backup-rename test. Main has since fixed the last three. Config compatibility confirmation: https://github.com/openclaw/openclaw/pull/158000#issuecomment-5913462783.

No overlap with Pash/Sarah changes.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-10-01 13:27:55 +08:00

16 KiB

summary read_when title
How OpenClaw builds prompt context and reports token usage + costs
Explaining token usage, costs, or context windows
Debugging context growth or compaction behavior
Token use and costs

OpenClaw tracks tokens, not characters. Tokens are model-specific, but most OpenAI-style models average ~4 characters per token for English text.

How the system prompt is built

OpenClaw assembles its own system prompt on every run. It includes:

  • Tool list + short descriptions
  • Skills list (metadata only; instructions load on demand with read). Native Codex turns on the managed bundled app-server get the compact skills block in parent-local model request instructions. Connections without that relay use thread developer instructions; other harnesses get it in the normal prompt surface. Bounded by skills.limits.maxSkillsPromptChars, with optional per-agent override at agents.entries.*.skillsLimits.maxSkillsPromptChars.
  • Self-update instructions
  • Workspace + bootstrap files (AGENTS.md, SOUL.md, IDENTITY.md, USER.md, BOOTSTRAP.md when new, plus MEMORY.md when present). Large injected files are truncated by agents.defaults.bootstrapMaxChars (default: 20000); total bootstrap injection is capped by agents.defaults.bootstrapTotalMaxChars (default: 60000).
    • Native Codex turns do not paste raw MEMORY.md when memory tools are available for that workspace; they get a small memory pointer in parent-local request instructions instead and use memory tools on demand. If tools are disabled, memory search is unavailable, or the active workspace differs from the agent memory workspace, MEMORY.md falls back to the normal bounded turn-context path.
    • Lowercase root memory.md is never injected. It is legacy repair input for openclaw doctor --fix, which migrates it into MEMORY.md.
    • memory/*.md daily files are not part of the normal bootstrap prompt; they stay on-demand via memory tools on ordinary turns. Reset/startup model runs can prepend a one-shot startup-context block with recent daily memory for that first turn, controlled by agents.defaults.startupContext. Bare chat /new and /reset are acknowledged without invoking the model.
    • Post-compaction AGENTS.md excerpts require explicit agents.defaults.compaction.postCompactionSections opt-in; plugins can add other context through before_prompt_build.
  • Time (UTC + user timezone)
  • Reply tags + heartbeat behavior
  • Runtime metadata (host/OS/model/thinking)

See the full breakdown in System Prompt.

When documenting credentials or auth snippets, use the Secret Placeholder Conventions to avoid secret-scanner false positives in docs-only changes.

What counts in the context window

Everything the model receives counts toward the context limit:

  • System prompt (all sections above)
  • Conversation history (user + assistant messages)
  • Tool calls and tool results
  • Attachments/transcripts (images, audio, files)
  • Compaction summaries and pruning artifacts
  • Provider wrappers or safety headers (not visible, but still counted)

Runtime-heavy surfaces have their own explicit caps under agents.defaults.contextLimits (per-agent overrides under agents.entries.*.contextLimits):

Key Purpose
memoryGetMaxChars Max characters memory_get returns before truncation.
postCompactionMaxChars Max characters retained from AGENTS.md during post-compaction refresh.

These are bounded runtime excerpts and injected runtime-owned blocks, separate from bootstrap limits, startup-context limits, and skills prompt limits.

OpenClaw derives the live tool-result cap from the effective model context window: 16000 chars below 100K tokens, 32000 chars at 100K+ tokens, 64000 chars at 200K+ tokens. The runtime context-share guard also caps a single tool result at 30% of the context window.

Large provider windows are not enabled automatically when they materially change cost or latency. For example, direct OpenAI GPT-5.5 and GPT-5.6 models publish a 1050000 token total window, but OpenClaw defaults their active runtime budget to 272000 tokens. The opt-in 922000 input budget reserves the full 128000 output allowance, and OpenAI applies higher long-context pricing to the entire request once input exceeds 272000 tokens. See OpenAI context window defaults.

For images, OpenClaw downscales transcript/tool image payloads before provider calls. Tune with agents.defaults.imageMaxDimensionPx (default: 1200):

  • Lower values reduce vision-token usage and payload size.
  • Higher values preserve more visual detail for OCR/UI-heavy screenshots.

For a practical breakdown (per injected file, tools, skills, and system prompt size), use /context list or /context detail. See Context.

How to see current token usage

In chat:

  • /status -> emoji-rich status card with the session model, context usage, last response input/output tokens, and cost from recorded billing or local pricing for the active model.
  • /usage off|tokens|full -> appends a per-response usage footer to every reply. Persists per session (stored as responseUsage).
    • /usage reset (aliases: inherit, clear, default) clears the session override so it re-inherits the configured default.
    • /usage tokens shows turn token/cache details.
    • /usage full shows compact model/context/cost details. Cost comes from a recorded amount or usage metadata with local pricing for the active model. Custom messages.usageTemplate layouts can include token/cache fields.
  • /usage cost -> local cost summary from OpenClaw session logs.

Other surfaces:

  • Control UI: the working indicator and completed-run recap show cumulative output tokens for that run, including its model calls across tool use and retries. Counts update when the runtime reports completed-response usage, not on every streamed text fragment. Reloading an active run restores its latest count. This counter excludes input tokens and is separate from the composer context-window meter and persisted billing summaries.
  • TUI/Web TUI: /status and /usage are supported.
  • CLI: openclaw status --usage and openclaw channels list show normalized provider quota windows (X% left, not per-response costs). Usage-window providers, checked against 2026.9.3: Claude (Anthropic), ClawRouter, Copilot (GitHub), DeepSeek, MiniMax, OpenAI, OpenRouter, Venice, xAI, Xiaomi, Xiaomi Token Plan, and z.ai. Provider plugins supply these snapshots, so an installed plugin can add one.

Usage surfaces normalize common provider-native field aliases before display. For OpenAI-family Responses traffic, that includes both input_tokens/output_tokens and prompt_tokens/completion_tokens, so transport-specific field names do not change /status, /usage, or session summaries. Gemini CLI usage is normalized too: the default stream-json parser reads assistant message events, and stats.cached maps to cacheRead, with stats.input_tokens - stats.cached used when the CLI omits an explicit stats.input field. Legacy JSON overrides still read reply text from response.

For native OpenAI-family Responses traffic, WebSocket/SSE usage aliases normalize the same way, and totals fall back to normalized input + output when total_tokens is missing or 0.

When the current session snapshot is sparse, /status and session_status can recover token/cache counters and the active runtime model label from the most recent transcript usage log. Existing nonzero live values still take precedence over transcript fallback values, and larger prompt-oriented transcript totals can win when stored totals are missing or smaller.

Usage auth for provider quota windows comes from provider-specific hooks first; if a provider has no hook (or the hook does not resolve a token), OpenClaw falls back to matching OAuth/API-key credentials from auth profiles, env, or config.

Assistant transcript entries persist the same normalized usage shape, including usage.cost when the runtime calculates an estimate or the provider reports a billed amount. This gives /usage cost and transcript-backed session status a stable source even after the live runtime state is gone.

OpenClaw keeps provider usage accounting separate from the current context snapshot. Provider usage.total can include cached input, output, and multiple tool-loop model calls, so it is useful for cost and telemetry but can overstate the live context window. Context displays and diagnostics use the latest prompt snapshot (promptTokens, or the last model call when no prompt snapshot is available) for context.used.

Native Codex turn usage sums the reported counts from each unique completed model response, including responses before a retry or cancellation. Missing response counts stay unknown; they do not erase already observed usage. A missing final response snapshot leaves context usage unavailable.

Cost estimation (when shown)

Costs are estimated from your model pricing config:

models.providers.<provider>.models[].cost

These are USD per 1M tokens for input, output, cacheRead, and cacheWrite. If both pricing and a recorded amount are missing, /usage full omits cost; use /usage tokens or a custom messages.usageTemplate when you need token/cache details in every reply. Cost display is not limited to API-key auth: non-API-key providers such as aws-sdk can show estimated cost when their configured model entry includes local pricing and the provider returns usage metadata.

When a model publishes tieredPricing, each request selects one tier using its total prompt input: uncached input plus cache reads and cache writes. Output tokens do not select the tier. The selected rates apply to every token bucket in that request, rather than only to tokens above a threshold. Ranges are half-open [start, end); an open-ended final range uses [start].

Turn totals sum costs calculated for each model request, retaining tier and model boundaries across tool loops and retries. If an older or external runtime provides only aggregate tokens, flat-rate estimates remain available. A tiered aggregate without complete per-request costs omits the cost instead of treating the summed tokens as one large request. Provider-billed totals, including zero, take precedence over catalog estimates and remain visible even when token counts are unavailable. Unknown token counts are not inferred from a billed amount.

Transcript reports preserve valid recorded per-call totals and allocations, including priority/flex adjustments, and use current catalog pricing only for missing costs or unknown-price zero placeholders. Anthropic fast-mode estimates multiply base and tier rates alike, preserving tier thresholds and mixed 5-minute/1-hour cache-write pricing.

Omitting cost, or setting it to {}, inherits the catalog pricing schedule. Explicit flat or all-zero model prices do not inherit a catalog tier schedule. Omitted flat-rate fields can still inherit catalog defaults. An explicit tieredPricing schedule takes precedence over the catalog schedule.

Pricing updates ship in the hosted model catalog alongside model metadata. Its publisher reads public pricing sources, including OpenCode's official catalog and Venice's public model API when the provider declares the native source. Base rates and context tiers come from the same source; usage rendering makes no network requests. Hosted updates activate with the Gateway's next prepared catalog generation, without restarting. Each estimation operation captures one pricing context, so a publication cannot change its rates halfway through. Set models.catalogRefresh.enabled: false to disable hosted catalog traffic on offline or restricted networks; bundled pricing still works. Agent-local models.json prices take precedence over explicit models.providers.*.models[].cost entries, and both override catalog estimates, including explicit flat and zero rates.

When the Gateway writes updated agent-local models.json prices, subsequent local estimates use those rates without a restart. Recorded per-call costs keep their original amounts.

OpenRouter :nitro and :floor routing shortcuts use the base model's catalog estimate when the exact shortcut has no price. Recorded costs and explicit prices keep their precedence. Private endpoints and other model variants do not use this fallback. Priority and flex billing can differ from the base estimate.

Cache TTL and pruning impact

Provider prompt caching only applies within the cache TTL window. OpenClaw can optionally run cache-ttl pruning: it prunes the session once the cache TTL has expired, then resets the cache window so subsequent requests re-use the freshly cached context instead of re-caching the full history. This keeps cache write costs lower when a session goes idle past the TTL.

Configure it in Gateway configuration and see the behavior details in Session pruning.

Heartbeat can keep the cache warm across idle gaps. If your model cache TTL is 1h, setting the heartbeat interval just under that (e.g., 55m) can avoid re-caching the full prompt, reducing cache write costs.

In multi-agent setups, you can keep one shared model config and tune cache behavior per agent with agents.entries.*.params.cacheRetention.

For a full knob-by-knob guide, see Prompt Caching.

For Anthropic API pricing, cache reads are significantly cheaper than input tokens, while cache writes are billed at a higher multiplier. See Anthropic's prompt caching pricing for the latest rates and TTL multipliers: https://platform.claude.com/docs/en/build-with-claude/prompt-caching

Example: keep 1h cache warm with heartbeat

agents:
  defaults:
    model:
      primary: "anthropic/claude-opus-4-6"
    models:
      "anthropic/claude-opus-4-6":
        params:
          cacheRetention: "long"
    heartbeat:
      every: "55m"

Example: mixed traffic with per-agent cache strategy

agents:
  defaults:
    model:
      primary: "anthropic/claude-opus-4-6"
    models:
      "anthropic/claude-opus-4-6":
        params:
          cacheRetention: "long" # default baseline for most agents
  list:
    - id: "research"
      default: true
      heartbeat:
        every: "55m" # keep long cache warm for deep sessions
    - id: "alerts"
      params:
        cacheRetention: "none" # avoid cache writes for bursty notifications

agents.entries.*.params merges on top of the selected model's params, so you can override only cacheRetention and inherit other model defaults unchanged.

Anthropic 1M context

OpenClaw sizes GA-capable Claude 4.x models such as Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6 with Anthropic's 1M context window. You do not need params.context1m: true for those models.

agents:
  defaults:
    models:
      "anthropic/claude-opus-4-6":
        alias: opus

Older configs can keep context1m: true, but OpenClaw no longer sends Anthropic's retired context-1m-2025-08-07 beta header for this setting and does not expand unsupported older Claude models to 1M.

Requirement: the credential must be eligible for long-context usage. If not, Anthropic responds with a provider-side rate limit error for that request.

If you authenticate Anthropic with OAuth/subscription tokens (sk-ant-oat-*), OpenClaw preserves the OAuth-required Anthropic beta headers while stripping the retired context-1m-* beta if it remains in older config.

Tips for reducing token pressure

  • Use /compact to summarize long sessions.
  • Trim large tool outputs in your workflows.
  • Lower agents.defaults.imageMaxDimensionPx for screenshot-heavy sessions.
  • Keep skill descriptions short (skill list is injected into the prompt).
  • Prefer smaller models for verbose, exploratory work.

See Skills for the exact skill list overhead formula.