openclaw/docs/tools/thinking.md
RoboClaw 1caaef4b6f
feat(ui): select Ultrafast for supported accounts (#160352)
* feat(ui): select Ultrafast for supported accounts

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(gateway): preserve Ultrafast compatibility and account authority

Negotiate speed decoding per connection and keep canonical session state intact. Fence personal-account catalog requests at final guarded HTTP dispatch. Refresh the measured UI boot manifest without changing performance limits.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* refactor(openai): keep account catalog outcomes together

Move the existing account-scoped result projection into the already imported catalog helper without changing its behavior. Keep the provider owner below its existing line-count ratchet.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* test: refresh Ultrafast tool prompt fixtures

Regenerate the canonical Codex dynamic-tool fixtures for the authorized Ultrafast speed value. Update only that enum value and its derived size/hash metadata; keep snapshot checks enabled.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix: stop account catalog requests after default unlink

Bind automatic account selections to the canonical profile writer's committed link authority and carry the real request scope through final guarded dispatch. Keep explicit retained account selections usable after unlink and evict failed discovery custody. Preserve the existing active Auto Ultrafast opt-in while explicit Fast stays priority.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* refactor(codex): use narrowed Auto activation flag

Keep the reviewed Auto tier predicate while satisfying the typed boolean lint contract.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(ui): share speed applicability for optional Ultrafast

Respect the selected request mapping before offering an optional entitled tier. Preserve all three choices on eligible routes and clearable stored preferences on unsupported routes. Reproduced the contradictory catalog regression and passed 59 unit/Chromium cases; scoped independent review found no actionable P0/P1.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(ci): export manifest in same-revision preflight harness

Restore the trusted file omitted when the inline manifest moved out of ci.yml in 9d75a8fe87. Actual preflight job109720001696 failed before tests with MODULE_NOT_FOUND; real Git push and PR materialization fixtures reproduce the same omission. Keep the canonical index export owner, regenerate its workflow projection, and preserve all source/credential guards. Fifteen materialization variants, five import/size checks, root test types, and focused independent review pass.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(ci): satisfy extracted manifest static contracts

Repair inherited check-lint failures from the manifest extraction without changing CI routing: avoid namespace shadowing, retain error cause and the diagnostic callback string contract, preserve nonmutating shard copies, and apply required branch syntax. All1259 scripts lint clean;45 planner/import/size cases, formatting, UI i18n, styles, ratchet and independent review pass.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(ci): centralize dependency-free workflow flag parsing

Remove the extracted manifest local coercion helper and preserve its exact narrow Boolean grammar under the existing script argument owner. Register the canonical declaration rather than weakening the guard, and carry its runtime through trusted preflight materialization and fixtures.77 argument cases,11 declaration-guard cases,71 scoped integration cases, types, lint, export scans and remaining guard commands pass; independent review has no actionable P0/P1.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(gateway): separate model publication authority from selection scope

Keep the actual request lifetime in selected-account HTTP assertions without treating every anonymous unscoped catalog read as a personal account projection. Restore the established models.list response shape; no assertions weakened.177 model/catalog/session cases and10physical HTTP authority cases pass, together with types, lint, ratchet and independent review.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix: preserve session response argument tuples

Forward the original response tuple while projecting successful legacy payloads, without appending optional undefined arguments.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(codex): preserve existing Ultrafast opt-ins

Keep the v2026.9.7 Fast and active Auto opt-in semantics while adding explicit per-session Ultrafast. Standard still clears the tier. Cover cold and warm native turn requests without requiring a migration.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* test(ui): distinguish speed labels from model names

Match the complete Effort and Speed section labels rather than the Speed only fixture model. Retain the independent absence assertions for reasoning and speed controls. All four failures reproduced before the repair; the complete 13-case bundled browser file passes afterward.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* test(openai): intercept the shared transcription socket

Update the two stale socket mock registrations after the upstream transport consolidation. Keep the actual provider/session code, fake peers, assertions, timeouts, and Bun transport guard unchanged. All 86 OpenAI shard files pass: 1263 passed and one existing skip.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* test(gateway): construct complete session reset callers

Replace partial caller objects cast as never with the existing typed session mutation client fixture. Preserve provenance and required-sandbox assertions and the production client capability contract. Both CI failures reproduce before the repair; all 15 reset-model cases pass afterward.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix: confirm the selected Ultrafast command mode

Share the direct command, directive reply, and system-event confirmation formatter so the saved Ultrafast tier is named accurately. Include the accepted manual value in help and docs without advertising an unverified optional native-menu choice. Preserve boolean Fast, Auto, reset, authorization and persistence behavior. Three regressions fail before repair; 274 focused cases pass afterward.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

---------

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
2026-09-30 07:45:41 +00:00

23 KiB

summary read_when title
Directive syntax for /think, /fast, /verbose, /trace, and reasoning visibility
Adjusting thinking, fast-mode, or verbose directive parsing or defaults
Thinking levels

What it does

  • Inline directive in any inbound body: /t <level>, /think:<level>, or /thinking <level>.
  • Levels (aliases): off | minimal | low | medium | high | xhigh | adaptive | max | ultra, roughly mirroring Anthropic's classic "think" < "think hard" < "think harder" < "ultrathink" magic-word ladder:
    • minimal ~ "think"
    • low ~ "think hard"
    • medium ~ "think harder"
    • high ~ "ultrathink" (max budget)
    • xhigh ~ "ultrathink+" (GPT-5.2+ and Codex models, plus Anthropic Claude Opus 4.7+ effort)
    • adaptive → provider-managed adaptive thinking (supported for Claude 4.6 on Anthropic/Bedrock, Anthropic Claude Opus 4.7+, and Google Gemini dynamic thinking)
    • max → provider max reasoning (Anthropic Claude Opus 4.7+; Ollama maps this to its highest native think effort)
    • ultra → harness-level planning, execution, verification, and proactive sub-agent orchestration; available for every model on the OpenClaw and Claude Code runtimes, and supported native reasoning models on Codex
    • x-high, x_high, extra-high, extra high, and extra_high map to xhigh.
    • highest maps to high; maximum maps to max.
  • Provider notes:
    • Thinking menus and pickers are provider-profile driven. Provider plugins declare the exact level set for the selected model, including labels such as binary on.
    • adaptive, xhigh, and max are advertised only when the provider/model supports them. Ultra is a separate harness mode, not an additional provider API effort. Typed directives for unsupported native levels are rejected with that model's valid options.
    • Existing stored unsupported levels are remapped by provider profile rank. When adaptive is not selectable, it uses the provider's declared non-off default; otherwise its ranked fallback preserves enabled thinking, usually medium. xhigh and max fall back to the largest supported non-off level for the selected model.
    • Anthropic Claude 4.6 models default to adaptive when no explicit thinking level is set.
    • Anthropic Claude Opus 4.8 and Opus 4.7 keep thinking off unless you explicitly set a thinking level. Opus 4.8's provider-owned effort default is high after adaptive thinking is enabled.
    • Anthropic Claude Opus 4.7+ maps /think xhigh to adaptive thinking plus output_config.effort: "xhigh", because /think is a thinking directive and xhigh is the Opus effort setting.
    • Anthropic Claude Opus 4.7+ also exposes /think max; it maps to the same provider-owned max effort path.
    • Direct DeepSeek V4 models expose /think xhigh|max; both map to DeepSeek reasoning_effort: "max" while lower non-off levels map to high.
    • OpenRouter-routed DeepSeek V4 models expose /think xhigh and send OpenRouter-supported reasoning.effort values instead of DeepSeek-native top-level reasoning_effort. Lower non-off levels map to high, and stored max overrides fall back to xhigh.
    • Ollama thinking-capable models expose /think low|medium|high|max. Verified full-effort Ollama Cloud families such as GLM 5.2 and DeepSeek V4 send each matching native think effort, including max; other models and local Ollama keep the compatible high mapping for /think max.
    • OpenAI GPT models map /think through the selected model and auth route's effort support. On supported OpenAI Platform/API-key routes, /think off sends explicit reasoning.effort: "none", including when using the Codex runtime. Subscription routes that do not support none retain the provider or native runtime's default reasoning behavior; off does not guarantee zero reasoning there. Supervised native Codex threads keep their own thinking settings.
    • GPT-6 Astra and GPT-5.6 Sol and Terra expose native /think ultra through the Codex runtime with either Platform API-key or ChatGPT subscription auth. OpenClaw preserves Ultra for ordinary turns; /btw intentionally runs at off. Codex owns proactive delegation and the model-specific inference effort (Astra uses xhigh). Ultra is not sent as a raw Responses API reasoning effort. Other Codex models with a native reasoning effort can also use Ultra: Codex selects their supported inference effort while retaining its native delegation policy. Models with no native effort choices can use host-bootstrapped Ultra through the OpenClaw runtime.
    • The OpenClaw and Claude Code runtimes expose logical /think ultra for all models. They select the highest supported native effort and add run-scoped planning and verification guidance. Delegation guidance appears only when sessions_spawn is available; Ultra does not grant tools or bypass their policy. Nonreasoning models remain nonreasoning, and models without an effort control retain the provider default.
    • Changing effort in a cached OpenAI conversation can append a configuration_update while keeping the original request-level effort for prompt reuse. The latest update determines effective effort; an unchanged top-level field is not a downgrade.
    • Custom OpenAI-compatible catalog entries can opt into /think xhigh by setting models.providers.<provider>.models[].compat.supportedReasoningEfforts to include "xhigh". This uses the same compat metadata that maps outbound OpenAI reasoning effort payloads, so menus, session validation, agent CLI, and llm-task agree with transport behavior.
    • Since 2026.4.26, stale configured OpenRouter Hunter Alpha refs skip proxy reasoning injection because that retired route could return final answer text through reasoning fields.
    • Google Gemini maps /think adaptive to Gemini's provider-owned dynamic thinking. Gemini 3 requests omit a fixed thinkingLevel, while Gemini 2.5 requests send thinkingBudget: -1; fixed levels still map to the closest Gemini thinkingLevel or budget for that model family.
    • MiniMax M2.x (minimax/MiniMax-M2*) on the Anthropic-compatible streaming path defaults to thinking: { type: "disabled" } unless you explicitly set thinking in model params or request params. This avoids leaked reasoning_content deltas from M2.x's non-native Anthropic stream format. MiniMax-M3 (and M3.x) is exempt: M3 emits proper Anthropic thinking blocks and returns empty content when thinking is disabled, so OpenClaw keeps M3 on the provider's omitted/adaptive thinking path.
    • Z.AI (zai/*) is binary (on/off) for most GLM models. GLM-5.2 and GLM-5.3 are the exceptions. GLM-5.2 exposes /think off|low|high|max with an off default, maps low and high to Z.AI reasoning_effort: "high", and maps max to reasoning_effort: "max". GLM-5.3 exposes /think low|high|max with a max default, maps off, minimal, and low to reasoning_effort: "low", medium and high to "high", and xhigh, adaptive, and max to "max".
    • Moonshot API Kimi K3 (moonshot/kimi-k3) always thinks at max, sends reasoning_effort: "max", omits the K2 thinking field and fixed sampling overrides, and preserves K3-supported tool choices. Kimi Code K3 (kimi/k3 and kimi/k3-256k) exposes the full /think ladder with a high default: off sends thinking.type: "disabled", minimal/low map to low effort, medium/high/adaptive to high effort, and xhigh/max to max effort. Kimi Code refs also include kimi/kimi-for-coding and kimi/kimi-for-coding-highspeed. Kimi K2.7 Code (moonshot/kimi-k2.7-code and moonshot/kimi-k2.7-code-highspeed) always thinks, exposes only on, and omits both outbound thinking and reasoning_effort. Other moonshot/* models map /think off to thinking: { type: "disabled" } and any non-off level to thinking: { type: "enabled" }. When K2 thinking is enabled, Moonshot only accepts tool_choice auto|none; OpenClaw normalizes incompatible values to auto.

Resolution order

  1. Inline directive on the message (applies only to that message).
  2. Session override (set by sending a directive-only message).
  3. Per-agent default (agents.entries.*.thinkingDefault in config).
  4. Per-agent model default (agents.entries.*.models["<provider>/<model>"].params.thinking in config).
  5. Shared model default (agents.defaults.models["<provider>/<model>"].params.thinking in config).
  6. Global default (agents.defaults.thinkingDefault in config).
  7. Fallback: provider-declared default when available; otherwise reasoning-capable models resolve to medium or the nearest supported non-off level for that model, and non-reasoning models stay off.

Setting a model default

Use params.thinking to set the default for one configured model without changing the default for your other models. The key must match the provider and model you actually select, including any model path exposed by a custom provider.

Replace <provider>/<model> with a configured model's full ID, then merge this entry into your existing model configuration:

{
  agents: {
    defaults: {
      models: {
        "<provider>/<model>": {
          params: { thinking: "high" },
        },
      },
    },
  },
}

The provider must already be configured, and the model must support the selected thinking level.

To change that model's default for one agent, put the same entry under agents.entries.<agent>.models. This overrides the shared model setting. Model params.thinking accepts the same aliases as /think; false and "disabled" also select off.

An inline directive, a saved session override, or a per-agent thinkingDefault still takes precedence. Send /think default to clear a saved session override; check the per-agent setting if the model default still does not take effect.

Setting a session default

  • Send a message that is only the directive (whitespace allowed), e.g. /think:medium or /t high.
  • That sticks for the current session (per-sender by default). Use /think default to clear the session override and inherit the configured/provider default; aliases include inherit, clear, reset, and unpin.
  • /think off stores an explicit off override until you change or clear it. Whether the upstream model can disable thinking depends on the selected provider and auth route.
  • Confirmation reply is sent (Thinking level set to high. / Thinking disabled.). If the level is invalid (e.g. /thinking big), the command is rejected with a hint and the session state is left unchanged.
  • Send /think (or /think:) with no argument to see the current thinking level.

Application by agent

  • Embedded OpenClaw: the resolved level is passed to the in-process OpenClaw agent runtime.
  • Claude CLI backend: concrete levels are mapped to Claude Code --effort; models that allow fixed budgets also receive a matching MAX_THINKING_TOKENS launch value. adaptive removes both configured effort flags and fixed-budget overrides, delegating effective thinking to Claude Code's environment, settings, and model defaults. See CLI backends.

Fast mode (/fast)

  • Levels: auto|on|off|default.
  • Directive-only message toggles a session fast-mode override and replies Fast mode set to auto., Fast mode enabled., or Fast mode disabled.. Use /fast default to clear the session override and inherit the configured default; aliases include inherit, clear, reset, and unpin.
  • Send /fast (or /fast status) with no mode to see the current effective fast-mode state.
  • OpenClaw resolves fast mode in this order:
    1. Inline /fast auto|on|off override on the current message
    2. Stored session override from a directive-only message (/fast default clears this layer)
    3. Per-agent default (agents.entries.*.fastModeDefault)
    4. Global default (agents.defaults.fastModeDefault)
    5. Per-model config (agents.defaults.models["<provider>/<model>"].params.fastMode)
    6. Fallback: off
  • Valid model-scoped params.fastMode / params.fast_mode values and valid cutoff keys are typed agent-runtime controls. They do not count as authored provider request params and do not select OpenClaw or Codex by themselves. Pin agentRuntime.id: "openclaw" or agentRuntime.id: "codex" when a recipe depends on one runtime.
  • auto keeps the session/config mode as auto but resolves each new model call independently. Calls that start before the auto cutoff have fast mode enabled; later retry, fallback, tool-result, or continuation calls start with fast mode disabled. The cutoff defaults to 60 seconds; set agents.defaults.models["<provider>/<model>"].params.fastAutoOnSeconds on the active model to change it.
  • For openai/*, fast mode maps to OpenAI API Fast mode (formerly Priority processing). OpenClaw currently sends service_tier=priority on supported Responses requests.
  • The Control UI stores Standard as fastMode: false, Fast as true, and Ultrafast as "ultrafast" in the same session preference. Ultrafast is offered only when support is confirmed for the selected model and account. Codex rechecks that support at each turn request and falls back to Fast when support is unavailable; a saved preference never grants access.
  • On Codex harness turns, the shared runtime control supersedes a configured native app-server tier: Fast on sends priority, Fast off sends null to clear the OpenClaw-owned tier, and auto decides for each model call. A configured Codex tier is used only when no shared Fast-mode run control is supplied. The existing appServer.enableUltrafast: true opt-in can upgrade Fast or active Auto to supported Ultrafast; Standard still clears the tier. See Codex harness.
  • For direct API-key anthropic/* requests, Opus 5 and Opus 4.8 use native speed=fast. Other supported models use Priority Tier: on sets service_tier=auto, off sets service_tier=standard_only. Sonnet 5 supports neither mapping; OAuth requests receive neither field.
  • For minimax/* on the Anthropic-compatible path, /fast on (or params.fastMode: true) rewrites MiniMax-M2.7 to MiniMax-M2.7-highspeed.
  • Explicit Anthropic serviceTier / service_tier model params override the fast-mode default when both are set. OpenClaw still skips Anthropic service-tier injection for non-Anthropic proxy base URLs.
  • /status reports the resolved OpenClaw policy (on, off, or auto) and the selected runtime. It does not report the upstream service tier actually honored or returned for a completed request. See OpenAI Fast mode for provider details.
  • The Control UI disables Fast choices confirmed to have no effect on the selected request. Existing saved preferences remain visible and clearable. When applicability is unknown, controls retain their existing behavior; availability does not promise vendor entitlement or faster responses.

Verbose directives (/verbose or /v)

  • Levels: on (minimal) | full | off (default).
  • Directive-only message toggles session verbose and replies Verbose logging enabled. / Verbose logging disabled.; invalid levels return a hint without changing state.
  • /verbose off stores an explicit session override; clear it via the Sessions UI by choosing inherit.
  • Authorized external channel senders may persist the session verbose override. Internal gateway/webchat clients need operator.admin to persist it.
  • Inline directive affects only that message; session/global defaults apply otherwise.
  • Send /verbose (or /verbose:) with no argument to see the current verbose level.
  • When verbose is on, agents that emit structured tool results send each tool call back as its own safe metadata-only message. Shell tools show their label without command text. These tool summaries are sent as soon as each tool starts (separate bubbles), not as streaming deltas.
  • Tool failure summaries remain visible in normal mode, but raw error detail suffixes are hidden unless verbose is full.
  • When verbose is full, tool outputs are also forwarded after completion (separate bubble, truncated to a safe length). If you toggle /verbose on|full|off while a run is in-flight, subsequent tool bubbles honor the new setting.
  • agents.defaults.toolProgressDetail controls the shape of /verbose tool summaries and progress-draft tool lines. Use "explain" (default) for compact human labels and "raw" for unabridged non-shell detail. Standalone shell summaries require /verbose full for command text; progress drafts require the channel's explicit streaming.*.commandText: "raw" opt-in. Per-agent agents.entries.*.toolProgressDetail overrides the default.
    • /verbose on: 🛠️ Exec
    • /verbose full + explain: 🛠️ Exec: check JS syntax for /tmp/app.js
    • /verbose full + raw: 🛠️ Exec: check JS syntax for /tmp/app.js, node --check /tmp/app.js

Plugin trace directives (/trace)

  • Levels: on | off (default).
  • Directive-only message toggles session plugin trace output and replies Plugin trace enabled. / Plugin trace disabled..
  • Inline directive affects only that message; session/global defaults apply otherwise.
  • Send /trace (or /trace:) with no argument to see the current trace level.
  • /trace is narrower than /verbose: it only exposes plugin-owned trace/debug lines such as Active Memory debug summaries.
  • Trace lines can appear in /status and as a follow-up diagnostic message after the normal assistant reply.

Reasoning visibility (/reasoning)

  • Levels: on|off|stream.
  • Directive-only message toggles whether thinking blocks are shown in replies.
  • When enabled, reasoning is sent as a separate message prefixed with Thinking.
  • stream: streams reasoning while the reply is generating when the active channel supports reasoning previews, then sends the final answer without reasoning. Channel previews remove recognized internal runtime context before delivery; the original reasoning remains unchanged for model replay.
  • Control UI history shows saved reasoning only for on, with View → Reasoning enabled. off and stream keep it hidden, including after reload.
  • Visible Control UI reasoning preserves Markdown paragraphs and fenced code blocks, including blank lines inside code.
  • Alias: /reason.
  • Send /reasoning (or /reasoning:) with no argument to see the current reasoning level.
  • Resolution order: inline directive, then session override, then per-agent default (agents.entries.*.reasoningDefault), then global default (agents.defaults.reasoningDefault), then fallback (off).

Malformed local-model reasoning tags are handled conservatively. Closed <think>...</think> blocks stay hidden on normal replies, and unclosed reasoning after already visible text is also hidden. If a reply is fully wrapped in a single unclosed opening tag and would otherwise deliver as empty text, OpenClaw removes the malformed opening tag and delivers the remaining text.

Heartbeats

  • Heartbeat probe body is the configured heartbeat prompt (default: Follow the heartbeat monitor scratch context when provided. Recurring tasks are automations; create or change their schedules with the automations tool, not heartbeat scratch. Do not infer or repeat old tasks from prior chats. If nothing needs attention, reply NO_REPLY.). Inline directives in a heartbeat message apply as usual (but avoid changing session defaults from heartbeats).
  • Heartbeat delivery uses the last outbound-capable non-reasoning payload. Separate reasoning or Thinking payloads remain internal, and a reasoning-only heartbeat result produces no alert.

Web chat UI

  • Model, thinking-level, and fast-mode overrides can be changed in an existing session with operator.write; administrator access is not required for these three controls. Read-only clients cannot change them.
  • These are session preferences for subsequent turns, not a promise to change an already-running model call. The composer disables the controls while a reply is running and while a model change is being applied.
  • The web chat thinking selector shows the explicit session override, or the inherited configured/provider default when no override is stored.
  • Refreshing, reloading, or compacting a conversation keeps an inherited choice inherited; it does not store the resolved level as an override. While model metadata is loading, refreshes retain the known thinking profile for the same model and runtime.
  • Selecting a level on the effort slider writes an explicit session override immediately via sessions.patch; it does not wait for the next send and it is not a one-shot thinkingOnce override.
  • Sending while model, reasoning, or speed picker changes are still being applied waits for every pending picker patch; if a change fails, the message stays unsent for review.
  • The effort control displays the resolved level, such as Medium or Off. To clear an override and return to inheritance, send /think default.
  • Explicit picker choices use their direct level labels while preserving provider labels when present (for example Maximum for a provider-labeled max option).
  • The picker uses thinkingLevels returned by the gateway session row/defaults, with thinkingOptions kept as a legacy label list. The browser UI does not keep its own provider regex list; plugins own model-specific level sets.
  • /think:<level> still works and updates the same stored session level, so chat directives and the picker stay in sync.

Provider profiles

  • Provider plugins can expose resolveThinkingProfile(ctx) to define the model's supported levels and default.
  • Provider plugins that proxy Claude models should reuse resolveClaudeThinkingProfile(modelId) from openclaw/plugin-sdk/provider-model-shared so direct Anthropic and proxy catalogs stay aligned.
  • Each profile level has a stored canonical id (off, minimal, low, medium, high, xhigh, adaptive, max, or ultra) and may include a display label. Binary providers use { id: "low", label: "on" }.
  • Profile hooks receive merged catalog facts when available, including reasoning, thinkingLevelMap, compat.thinkingFormat, compat.supportsReasoningEffort, and compat.supportedReasoningEfforts. Use those facts to expose binary or custom profiles only when the configured request contract supports the matching payload. A null entry in thinkingLevelMap removes that level before choosing a default.
  • Tool plugins that need to validate an explicit thinking override should use api.runtime.agent.resolveThinkingPolicy({ provider, model, agentRuntime }) plus api.runtime.agent.normalizeThinkingLevel(...); they should not keep their own provider/model level lists. Pass agentRuntime when the tool owns the execution path, such as an always-embedded run.
  • Tool plugins with access to configured custom model metadata can pass catalog into resolveThinkingPolicy so compat.supportedReasoningEfforts opt-ins are reflected in plugin-side validation.
  • Published legacy hooks (supportsXHighThinking, isBinaryThinking, and resolveDefaultThinkingLevel) remain as compatibility adapters, but new custom level sets should use resolveThinkingProfile.
  • Gateway rows/defaults expose thinkingLevels, thinkingOptions, and thinkingDefault so ACP/chat clients render the same profile ids and labels that runtime validation uses.