Extend the existing model matrix with isolated Gateway tasks, fixed-workload build comparisons, and separate task and interview evidence. Preserve failed trials, exact source identities, task effects, logical cell completion, and observed preview coverage.
Validation: 187 focused tests, changed-file and targeted type/lint/docs checks, exact-candidate runtime build, actual OpenAI smoke, read-only replay of 24 original trials plus the final smoke, and fresh P2 review. Original scores and interview qualifications remain preserved.
Preserve actual Unix child termination signals through startup respawns so supervisors stop interrupted runtimes without entering watcher recovery. Keep explicit numeric exits and Windows termination behavior intact.
Fixes#144200.
* docs(tools): split the browser page by reader job
docs/tools/browser.md was 55,693 characters across 21 H2 sections mixing
explanation, how-to, reference, and troubleshooting. Move each section into
docs/tools/browser/ and keep /tools/browser as the index.
The split is content-preserving: the nine child pages plus the three blocks
the index still publishes (lede, What you get, Related) reconstruct the
original body byte-for-byte (sha256 a28d5ea4...). Only the four H2 lines that
became child page titles are not reproduced verbatim; each keeps its anchor
on the index.
Every one of the 43 pre-split anchor IDs -- headings, their percent-encoded
and cleaned variants, and the Accordion and Tab titles -- is authored as an
<a id> stub on the index, so existing /tools/browser#... links still resolve.
Per-anchor redirects are not possible: redirectSource() rejects sources
containing [?#].
* docs(tools): retarget cross-references orphaned by the browser split
Eleven link targets, no prose changes. Three were same-page or same-route
anchors inside the moved content: /tools/browser/setup's [Configuration] link
was genuinely broken (its target moved to the configuration page), and the
[Profiles] and [Custom Chrome MCP launch] links resolved only through the
index's anchor stub. The remaining eight are inbound deep links from other
pages, retargeted at the page that now holds the section, matching the
code-mode split precedent.
* docs(i18n): add zh-CN glossary entries for the browser child page titles
check-docs-i18n-glossary requires a source term for every changed doc label.
## What Problem This Solves
Fixes an issue where users running `openclaw agent` could not identify an accepted Gateway run from the human diagnostic after the connection closed or the CLI wait timed out. The run ID already appeared in JSON failure output, but the human hint omitted it.
## Why This Change Was Made
The hint reads the existing error-bound Gateway metadata and names the accepted run. Timeout failures also mention `--timeout <seconds>`. Both hints retain the instruction to check Gateway status and the session transcript before retrying.
The original automatic embedded fallback reported in #111407 was already removed from main. This PR completes the remaining diagnostic improvement; it does not introduce a reattachment command or guarantee that the run is still active.
## User Impact
Operators can correlate an ambiguous failure with the accepted run before risking duplicate execution. The original error, nonzero exit, canonical JSON envelope (`ok: false`, `error.type`, `error.message`, optional `runId` and `origin`), retry behavior, and cancellation behavior remain unchanged. Failures without an observed accepted ID keep the generic hint.
Fixes#111407.
## Evidence
Verified candidate `da178e1210818e8a8a1ec64145d871a30a4228f5` against pinned main `fc55b5b936` on one isolated Linux Testbox, using real `pnpm openclaw agent` processes, a real Gateway, an authenticated production WebSocket client, a loopback fault proxy, and a deterministic local provider.
- Before: accepted-run transport close and timeout retain the ID in JSON but omit it from the human hint. After: both human hints include the exact accepted ID; only timeout includes timeout guidance. Both exit unsuccessfully and retain the same sanitized transport error.
- Baseline and candidate controls cover human/JSON output, pre-acceptance failure, missing or unobserved acceptance IDs, cached final errors, transient and exhausted handshakes, normal completion, signals and `chat.abort`, local state-lock refusal, and standalone local execution/signals. Accepted transport-loss cases submit one CLI turn and one provider request. Cached-final proof uses one explicit fixture seed followed by the original CLI request, with one provider request total.
- Focused tests: 126 agent-command cases, 19 shared-failure cases, and the contributed real-Gateway end-to-end test passed. Changed-file formatting/lint and documentation MDX checks passed.
- Exact-head [CI run](https://github.com/openclaw/openclaw/actions/runs/34204549245) is green.
The timeout control uses the documented CLI timeout; its Gateway turn may already have ended. Process-level proof covers observable output and execution; original error-object identity and source-only invariants are covered by the focused tests and code review. No external inference or production channel was used.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
Each fix replaces a statement that contradicts the CLI implementation,
a TypeScript type, or the tool catalog. No prose or structure changes.
docs/cli/agent.md:11 — said "The explicit `--local` flag is the only
embedded execution path". `agent exec` is also embedded: the same page
at line 21 says it "runs one embedded agent turn without connecting to
a Gateway", and src/commands/agent-exec.ts:210 documents it as "Run one
isolated embedded agent turn and project its stable CLI result."
docs/cli/secrets.md:27 — the recommended operator loop ran `openclaw
secrets configure` with no `--plan-out`, then applied
`/tmp/openclaw-secrets-plan.json` on the next line. A plan file is only
written when the flag is passed: src/cli/secrets-cli.ts:255 guards the
write with `if (opts.planOut) { ... writeFileSync(opts.planOut, ...) }`,
and there is no default plan path. Added the flag to line 27.
docs/gateway/config-tools.md:26,46 — the `coding` profile row and the
`group:media` row listed `image` as a tool id. The catalog registers
`view_image`: src/agents/tool-catalog.ts:427 `id: "view_image"`, and
src/agents/tools/image-tool.ts:807 `name: "view_image"`. No tool with
id `image` exists. The same page already says so at line 149: "The image
inspection tool is `view_image`."
docs/cli/plugins.md:36,40 — the synopsis omitted flags the commands
accept. src/cli/plugins-cli.ts:136 gives `enable` `--accept-capabilities`;
src/cli/plugins-cli.ts:180,186 give `install` `--accept-capabilities` and
`--acknowledge-install-policy-warning`.
docs/cli/devices.md — the command reference had no section for
`openclaw devices join-code`, which src/cli/devices-cli.ts:43-47
registers with the description "Mint a single-use node onboarding URL".
Added a section matching its actual option surface.
docs/cli/index.md:165,200 — the command tree omitted `secrets store`
(registered at src/cli/secrets-store-cli.ts:164 and documented in the
`openclaw secrets` table at docs/cli/secrets.md:18) and `plugins pack`
(registered at src/cli/plugins-cli.ts:277 and listed in the plugins
synopsis at docs/cli/plugins.md:49).
docs/plugins/hooks.md:497 — the `BeforeToolCallResult` type block omitted
`scope`, while docs/plugins/plugin-permission-requests.md:93 tells plugin
authors to "Set `requireApproval.scope`". The real type has it:
src/plugins/hook-before-tool-call-result.ts:21 `scope?: ApprovalScope;`.
docs/diagnostics/flags.md:51 — the multi-flag example enabled
`"gateway.*"`. No diagnostic flag lives in a `gateway.` namespace: the
call sites are `timeline`, `diagnostics.timeline`, `plugin.load-profile`,
`ingress.timing` and `health`, and the page's own Known flags table
(lines 24-34) lists none. Replaced it with `health`, a real flag.
Closes audit findings: r3-0182, r3-0198, r3-0335, r3-0128, r3-0199, r3-0195, r5-0022, r3-0281
Extend the existing Code Mode evaluation harness with opt-in large-result reduction, independent-read composition and dependent-chain tasks. Preserve default matrix size and distinguish missing telemetry from observed zeros.
Focused tests, negative controls, dry-run evidence, changed checks, independent review and exact-head CI pass. No live model performance result is claimed.
Worked on by:
- @Takhoffman
Co-authored-by: Takhoffman <781889+Takhoffman@users.noreply.github.com>
* fix(agents): fit compacted context and prioritize foreground replies
Size automatic client-side compaction against the prepared foreground request,
including its actual queued context. Preserve pending input, complete exchanges,
and transient prompt ownership across compaction and initial steering. When
fixed context consumes the preferred reserve target, allow only a strictly
smaller history replacement instead of rejecting useful progress.
Run optional Gateway maintenance after delivery under fresh session admission;
preempt and drain it before the next foreground turn. Keep required checkpoints,
one-shot CLI boundaries, native compaction ownership, and explicit session fences.
Related: #121617, #53008
* test(agents): align maintenance fixtures with foreground ownership
Exercise required checkpoints through compaction, wait for optional maintenance settlement, and retain cancellation and delivery guards. Give retry-only sessions a valid context window and verify the actual remaining construction deadline. Qualify early-preflight headroom in the docs based on live measurements.
Validation: 51 SDK, construction, and direct-reply tests; 196 command and CLI owner tests; changed-scope test and docs gates; independent P0-P2 review. Production runtime files are unchanged from ae631c61.
* refactor(agents): remove overwritten model-switch result stores
Fallback provider and model are recorded from the completed attempt before the only return. Model-switch retries cannot expose the intermediate catch-block values. Removing those dead stores also keeps the merged module below the existing line limit after main adds error-transcript forwarding.
Validation: 131 model-switch tests, full changed-scope checks, and independent P0-P2 review pass. No inference, maintenance, or model-selection behavior changes.
* fix(agents): preserve pending turns through maintenance
Keep required memory checkpoints on a detached view of the processed conversation,
with private execution identity and the original sandbox, tool restrictions, and
prompt-cache ownership. Preserve admitted input independently of request sizing.
Carry intentional automatic compaction skips through their owner so foreground
preparation can continue, while preserving manual SDK results and failure handling.
Exclude hidden maintenance from displayed activity without changing cancellation
or operational liveness.
Regression coverage includes real SQLite admission, restricted memory writes,
compaction, and exact pending-input delivery. Focused suites, changed checks, and
independent review pass; fresh CLI and Tauri measurements follow on this revision.
* test(gateway): align activity mocks with session presentation
Preserve the raw operational progress mocks while supplying the separate
session-presentation query used by activity projections. Update shared fixture
consumers and reset paths without changing assertions or production behavior.
All six hosted CI failures reproduce before this fixture repair. The seven full
owner and sibling suites (183 tests), scoped changed checks, and independent
review pass afterward. Runtime files remain identical to bbbad8fd live proof.
* fix(agents): keep memory maintenance out of user transcripts
Run required and optional memory inference on private conversation copies.
Preserve source policy, prepared model facts, and cache identity while keeping
canonical compaction and memory outcomes with the original session owner.
Interrupted maintenance no longer becomes queued user context in the next
request. Owner cancellation does not count as a failed memory attempt.
* test(agents): use canonical history in delivery maintenance fixture
* fix(agents): prevent nested maintenance dependency cycles
* fix(runtime): retain cleanup ownership across agent teardown
Keep command results and cancellation separate from cleanup uncertainty.
MCP, LSP, Codex, CLI and Gateway owners retain cleanup evidence before
releasing resources or replacing a prior execution. Keep bounded agent
state when closure cannot be confirmed and fence retained host writes.
Stage 3 consumer extraction from #135868, following the process foundation
in #139517. Remove duplicate cleanup timers, one-use adapters, and the
unused asynchronous finalizer callback while preserving plugin disposal
compatibility.
* test(runtime): simplify cleanup ownership fixtures
* fix(runtime): fence retired handles and check transport exit evidence
* test(runtime): align cleanup proof with suite boundaries
* fix: retain local-model tools without changing sibling agents
Infer structured Tool Search from the final resolved provider route instead
of writing global lean mode during model setup. Preserve explicit user
choices and migrate only the lean flag owned by the retired setup marker.
Align CLI and native Ollama context caps and scale the built-in compaction
reserve to small context windows. Preserve ordinary forced message delivery
through the shared tool catalog while keeping private replies restricted.
Full build, focused owner tests, config baseline, and independent review
passed. Fresh CLI/native Ollama evaluation is tracked before landing.
Fixes#138753.
* test: cover private delivery and isolate local-model regressions
* fix: clarify deferred tool calls and trim repeated metadata
Keep full structured call details while presenting only exact tool identity
and the unchanged target result to the model. Explain tools-mode wrapping
and compact automation summaries without changing execution permissions.
Validation: 468 focused tests, failing-before compact-output regression,
and independent Codex review. Fresh real-model setup and recall proof follows.
* test: refresh automation prompt snapshots
Regenerate the canonical prompt fixtures for compact list summaries and
full get details. Snapshot drift check and all16 reconstruction tests pass.
Production remains byte-identical to8fa7d1488c3.
* fix: retain exact timing in compact automation lists
Include existing at/every/cron schedules after the normal visibility filter
so disabled jobs and recurring intervals can be understood in one tool call.
Keep event commands, working directories, payloads, and delivery definitions
behind the existing full-job read. Preserve scheduleKind for API compatibility.
The real local-model trial found the job but invented its schedule because
only the kind was exposed. Owner regressions reproduce that missing state.
* test: dispose assistant panel session capabilities
The panel fixture owns its application-level session capability. Removing
DOM alone leaves subscription retries alive and contaminates later fake
clocks. Register disposal at creation without changing runtime behavior or
weakening the Sessions typing timer assertion.
The original 11-file worker order reproduces the failure before cleanup and
passes all 233 tests afterward. Independent review is clean.
* fix: expose readable run times in compact automation results
Add exact ISO next/last run dates using the canonical timestamp formatter.
Keep the shipped millisecond fields for programmatic clients and preserve
null for absent run times. Models can now report dates without calendar
arithmetic on epoch values.
All 253 Gateway tests pass; five old-producer failures establish timestamp
and absence coverage. Independent review is clean.
* fix: preserve current Ollama context caps during doctor
Do not turn catalog windows or provider output budgets into stronger
num_ctx pins when a current contextTokens cap is already present. Avoid
provider-wide synthesis for mixed current/legacy models; migrate uncapped
native siblings individually and preserve existing explicit overrides.
Retain the shipped native legacy budget precedence and the compatibility
adapter's own API-specific behavior. Simplify the private API classifiers
and budget resolver while fixing the migration owner.
Live reproduction: Doctor tried to pin 262144 over the onboarded 32768 cap.
Eight pre-fix regressions fail; all 98 migration tests pass after the fix.
Independent review is clean. Existing oversized pins remain operator-owned.
* fix(ollama): preserve prepared tool text and continuation hints
Keep caller-budgeted explicit tool text intact when adding structured
fallback content. Bound structured data separately so large file pages
and Tool Search results retain their complete instructions.
Live native wire reproduction showed a 15,976-character tool result
truncated to 8,014 characters, deleting its continuation offset.
Regression coverage preserves long text, Unicode and whitespace while
retaining structured-data bounds and redaction.
* fix(agents): share the command budget with post-turn maintenance
Use the foreground run's remaining allowance for memory flushing and
compaction instead of restarting a full timeout for each phase. Cancel
and settle maintenance before returning the completed reply, while
preserving caller cancellation, restart fencing, and accepted compaction
session/count facts.
A live local-model turn completed in 462 seconds, but its CLI waiter
expired at 630 seconds while a separate memory flush was still running.
The existing model and Gateway deadlines remain unchanged.
Regression coverage includes exhausted and unlimited budgets, shared
phase expiry, committed successors, CLI siblings, and caller abort.
* fix(ollama): preserve native chat history on overflow
Default local native requests to truncate:false and shift:false so context
pressure reaches OpenClaw's compaction/recovery path instead of silently
removing conversation history. Preserve explicit model parameters and keep
hosted routes unchanged. Document the verified runtime and partial-output
behavior.
* fix(ollama): preserve hosted defaults through custom proxies
Exclude explicit Ollama Cloud provider routes from local history-preserving
request defaults even when their model and proxy URL lack cloud markers.
Reuse the existing cloud-origin check and cover the real native request.
* test(ollama): expect history preservation in timeout requests
Keep the complete request expectation aligned with the intentional native
truncate:false and shift:false defaults. Preserve the timeout and every
other existing assertion. No production behavior changes.
* fix(agents): use provider compaction effort and one summary format
Native Ollama summaries default to thinking off through prepared provider
metadata; explicit compaction settings and hosted routes retain precedence.
Use one primary summary format across history, prefix, stage and fallback
requests, preserving focus and existing output budgets.
Three live 4B compactions hit the existing 180-second watchdog with low
effort. A controlled thinking-off request completed in 39 seconds but
exposed conflicting formats. Owner regressions cover both failures.
No new operator configuration, schema, or timeout increases.
* fix: allow environment-only agent runs without a native harness
* chore(ui): reconcile measured main startup baseline
Unmodified main b6a4d0dad4 fails the startup gzip check at 350377 B in Linux CI; a pristine source build measures 350363 B. Record the measured main baseline while retaining the fixed cap and existing growth and variance allowances.
* fix(browser): avoid standalone routing errors and cropped screenshots
Route implicit standalone browser calls locally while preserving Gateway-owned routing. Capture through the session that owns viewport or device emulation, serialize mutations, and preserve successful viewport changes after later device errors.
* test(browser): complete control-server connection fixture
* fix(protocol): retain named schema types in declarations
* fix(browser): bound interrupted captures and release stale metrics
* fix(browser): distinguish embedded contexts from Gateway owners
Mark local embedded RPC contexts at their owner so standalone browser calls can use local control without Gateway discovery. Preserve caller and ambient Gateway bindings after retirement, and replace mocked ownership checks with real local-admission coverage.
* fix(cli): keep agent exec run errors when cleanup fails
After a failed run, attach cleanup as a secondary envelope field
instead of replacing the run error. Successful runs that only fail
cleanup still report the cleanup error.
Signed-off-by: Sebastien Tardif <sebtardif@ncf.ca>
* fix(cli): preserve agent exec failures during cleanup
Keep the primary run envelope and exit code when cleanup also fails.
Report secondary cleanup errors on stderr without extending the JSON result.
Co-authored-by: Sebastien Tardif <sebtardif@ncf.ca>
---------
Signed-off-by: Sebastien Tardif <sebtardif@ncf.ca>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: Sebastien Tardif <sebtardif@ncf.ca>
* fix(cli): reject a blank --session-id or --session-key instead of ignoring it
`openclaw agent --agent ops --session-id "$SESSION_ID" --message hi` with an
empty SESSION_ID does not fail. It runs the turn in the ops agent's default
session, so a caller that meant to target one session silently writes into
another.
Every downstream layer trims the value and treats blank as absent.
`normalizeSessionKeyOptsForDispatch` computes `hasExplicitSessionTarget` from
`Boolean(opts.sessionId?.trim())`, so a blank id reads as "no explicit target"
and the dispatch falls through to the agent default. `--session-key` takes the
same path, and `src/agents/command/session.ts` normalizes blank to undefined
as well, so nothing further down can tell the two cases apart.
`agentCliCommand` already rejects the sibling case three lines up:
if (opts.agent !== undefined && !opts.agent.trim()) {
throw new Error("--agent must not be blank");
}
This applies that same rule to the other two selector flags. An omitted flag is
still absent and still resolves normally; only a flag that is present and blank
now fails.
* fix(cli): reject explicit blank agent selectors
Validate all documented selectors at the shared CLI command boundary before
session normalization or dispatch. Preserve omitted and valid selectors,
cover local and Gateway paths, and document the explicit-blank contract.
Co-authored-by: MarMar Labs <marmar9615@icloud.com>
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: MarMar Labs <marmar9615@icloud.com>
* fix(usage): correct cached long-context cost estimates
Keep Venice extended schedules in the owning manifest, select per-call tiers from total prompt input, and preserve explicit pricing and provider-billed totals through run accounting. Fixes#133343.
* test(usage): assert worker monetary result metadata
Keep observed Gateway run provenance in agent --json failures so automation can correlate accepted and cached run errors. Preserve existing error fields, human output, cancellation behavior, and retry policy; local request keys are never reported as Gateway provenance.
The gateway-routed path already mapped terminal run status to an exit code,
but the local embedded path did not, so a failed turn exited 0 while its own
JSON envelope reported status "error". Route both through the canonical
agent-run terminal outcome, and fail closed on an unrecognized status since
the gateway response carries an open string.
* fix(skills): expand explicit references on agent turns
Route generic Gateway, CLI, webhook, and local agent turns through the same explicit skill-reference renderer as channel auto-replies. Keep original transcript text, preserve unknown slash behavior, and fail visibly for allowlist-hidden skills.
Maintainer review: scoped Option 1 — generic agent turns expand both $skill-name and leading /skill-name args through shared skill rendering; they do not run the channel command dispatcher, and all other slash commands retain their existing behavior.
* fix(skills): bound explicit reference prompts
* fix(skills): prefer allowed reference collisions
* fix(skills): preserve command invocation boundaries
* fix(skills): reject hidden channel slash commands
* perf(skills): skip literal dollar discovery
* feat(cli): run agent exec against the ambient config, composed in memory
Exec previously ignored the operator's config entirely, so a one-shot turn
could not reach configured providers, credentials, or agentRuntime harness
selection. It now layers config the way other folder-scoped coding CLIs do.
The composed config is published as this process's runtime snapshot rather
than serialized to a temp file and re-read through OPENCLAW_CONFIG_PATH. The
snapshot is the only in-process config cache, so the file only ever fed it --
while writing env-substituted provider keys to disk where the run's own exec
tool could read them.
* fix(cli): resolve exec stored credentials from the configured agent dir
* chore(scripts): allow agent exec the file-scoped config loader at its process boundary
* test(cli): cover the exec credential default and pinned-config flags
* feat(harness): report copilot code-mode engagement on the attempt result
* test(copilot): prove code-mode engagement through the production tool bridge
* docs: describe the normalized codeModeEngaged value for native harnesses
* feat(agents): add per-run stats to embedded agent run meta
Adds codeModeEngaged, assistantTurns, bridgeCalls, and costUsd to
EmbeddedAgentRunMeta.agentMeta and mirrors them on the agent exec --json
envelope. Code-mode engagement is stamped from the tool-surface truth,
round trips accumulate across attempts beside usage, bridge counts come
from the run's tool-search catalog counters, and cost reuses the shared
model pricing helpers (cache tiers included, omitted without cost data).
* fix(agents): accumulate bridge call counts across run attempts
Attempt cleanup clears the per-attempt tool-search catalog, so retries and
fallbacks discarded earlier bridge counts. Fold each attempt's bridgeCalls
into the run accumulator beside assistantTurns and stamp the cumulative
totals into agentMeta, matching the documented per-run contract.
A Gateway timeout or closed connection now fails the command with an
actionable stderr hint instead of silently re-running the whole turn
embedded under a fresh gateway-fallback-* session. The silent fallback
could double-execute side effects (the Gateway may still finish an
accepted turn), returned context-free answers to --session-key callers,
and ran with the CLI host's local config. --local remains the only
embedded execution path.
Also deletes the resultMetaOverrides plumbing (the fallback was its only
writer) and the fallback marker fields added in #111645.
Merges the Clownfish-repaired contributor branch for #93351. The latest repair preserves inline --message whitespace, adds --message-file coverage for gateway and local embedded runs, and the PR is clean/mergeable on head 4897f2fc20.
Clean up local Claude stdio one-shot runs before returning from embedded `openclaw agent --local`, including bundle MCP loopback teardown for local process resources.
Keeps gateway-owned MCP loopback cleanup internal to the Gateway, documents the local-vs-gateway behavior, and aligns the stale OpenAI provider-runtime fixture with the current unsupported Codex mini route.