Commit graph

50 commits

Author SHA1 Message Date
Josh Avant
3b0b12dea4
fix: show failed native tools in Codex transcripts (#164116) 2026-10-03 08:37:14 -05:00
Peter Steinberger
899a0ca681
fix(codex): preserve remote execution approval failures (#164237)
Treat expired, denied, and unavailable Codex launch approvals as execution
coordination failures instead of provider failures. Keep the original
bounded recovery message across live replies and saved history without
rotating healthy credentials or recording auth-profile failure state.

Preserve sign-in guidance and failure recording for genuine credential
errors. Approval enforcement and exec-server behavior are unchanged.
2026-10-03 03:37:03 -07:00
RoboClaw
6046d4fcb6
feat: use MCP plugin apps across conversations and workspace files (#161747)
* feat: use MCP plugin apps across conversations and workspace files

Extend the opt-in MCP Apps host with discovery and entrypoints, settings, rich forms, scoped file editing, multimodal context, deep links, and native Codex session preparation. Preserve existing requester, approval, runtime, and file authority owners.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix: keep app context compact and review large tool payloads

Reuse the canonical bounded plugin approval preview without rejecting otherwise valid large App arguments. Keep the same one-shot approval and live-authority gates. Bound context-strip icons and scrolling so attachments leave the composer usable.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix: reveal pending input above fullscreen MCP apps

Keep App fullscreen rendering in the existing browser top layer and return to inline for same-conversation questions or approvals without replacing the iframe. Consolidate host-context, resource, and input notifications under the existing bridge lifetime.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* test: preserve full attempt inputs in native assignment fixtures

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix: release session access when MCP app launch preparation fails

Preserve the upstream optional ifMatch contract and document unconditional-save semantics explicitly.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(ui): defer question controls and preserve canonical keyboard values

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(mcp): preserve app authority through startup and user turns

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(codex): share native app startup across concurrent discoveries

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* docs(mcp): explain native app prompting and request limits

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* feat: use MCP plugin apps across conversations and workspace files

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: f2f8d735-c1ac-4a0d-84f3-1576113c1baf

* fix(mcp): keep form contracts and test helpers with their owners

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* feat: use MCP plugin apps across conversations and workspace files

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: 2513c3ef-4af8-4722-b802-ffa426df486a

* fix(codex): revalidate MCP App authority before native retries

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* test(ui): provide chat identity fixture context dependencies

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* feat: use MCP plugin apps across conversations and workspace files

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: 85c132dc-d926-46ca-b57e-2d9e01f15487

* fix: integrate MCP app owners with current main

* fix(gateway): accept option labels from installed clients for rich forms

* chore(format): anchor the native apps formatter ignore

* perf(ui): load the MCP app link parser on demand

Keep startup click interception parser-free and capture the normalized href before lazy navigation. Preserve ordinary, modified and download links; drop malformed plugin links and cancel pending navigation on disposal. Router/parser proof: 17 tests passed in 4.50s wall. Startup gzip: 374086 -> 373613 B; 159 B still above the unchanged 373454 B limit.

* perf(ui): register MCP app English with lazy consumers

Move the MCP App catalog out of startup English and register it synchronously at every consumer. Keep the shared namespace anchor and host catalog composition so all text and source order remain byte-identical. i18n verification passes without baseline changes. Startup gzip: 373613 -> 373103 B, 351 B below the unchanged 373454 B limit.

* fix(ui): intercept only well-formed MCP app links

Require the plugin/app/tool shape before cancelling browser navigation so ordinary ChatGPT plugin pages keep their default behavior. Preserve lazy strict parsing and accepted native and web app links.

* fix(gateway): watch MCP app files directly so loaded macOS hosts still notify

Watch the bound file to avoid FSEvents directory event drops under load. Re-arm on atomic replacement, retry ENOENT once on the next immediate turn, and retain subscription authority and cleanup. Cover replacement followed by a plain write through the registered resource routes.

* fix(gateway): use the errno guard when rearming MCP app watches

Use the shared filesystem error guard instead of an unchecked type assertion. Keep the single immediate ENOENT retry and all subscription behavior unchanged.

* fix(codex): read the canonical MCP transport field

Main migrates MCP transport aliases before runtime (#162256) and retired the SDK re-export of the alias resolver, so the native app catalog reads the canonical server.transport field the same way the agentsapi plugin does.

* fix(gateway): keep MCP app file subscriptions alive across rename gaps

Keep resource subscriptions alive while editors move the old file aside before installing its replacement. Poll missing paths at 250 ms, return to the inode watcher when the file reappears, and release polling on subscription close. Recheck after polling registration to cover replacement before its first stat; share the rearm guard so late callbacks cannot leak watchers.

* test(codex): type the rooted thread policy support from attempt fixtures

The rooted policy support that main added types its lifecycle input from the raw thread signature, while this branch's binding fixtures carry full attempt params. Derive it from the shared attempt-thread fixture type, as the sibling policy-refresh support already does.

* test(ui): exercise the MCP app link probe through the click boundary

The eager link probe was exported only for its unit test, which the production dead-export scan rejects. Keep it module-private and assert the same accepted and rejected links through real click interception.

* style(lint): clear MCP app lint errors

CI run 36988865432 jobs check-lint-core-1 and check-lint-core-2 rejected shadowed stat callback variables, a returning Promise executor, and reassigned projection parameters. Rename the inner variables and use local projection bindings and an executor block without changing behavior.

* test(agents): align bundle MCP fixture ownership

CI run 36988865432 job checks-node-changed-compact-large-14 failed seven merge cases because the fixture omitted the loader-required pluginIdsByServer map. Type the fixture against the producer contract and verify ownership survives only for unshadowed enabled bundle servers, preserving the migrated transport behavior.

* fix(auto-reply): register turns before MCP context leasing

CI run 36988865432 job checks-node-changed-compact-large-33-2 exposed an asynchronous MCP lease before synchronous run registration. Prepare App context after ownership registration and image admission, revalidate requester authority around the lease, and retain commit and rollback settlement with the execution outcome owner.

* test(gateway): admit MCP shutdown requests through current policy

CI run 36988865432 job checks-node-changed-compact-large-9 timed out because the fixture ignored an admission rejection before its upstream call. Publish matching Gateway and runtime config, create the session through its RPC owner, and observe early request settlement while preserving all shutdown ordering and cleanup assertions.

* refactor(apple): drop the label-only question toggle

Periphery flagged toggleOption(questionID🏷️) as dead: production selects by canonical value since the rich-form change, and only tests still called the label overload. Tests now toggle by value; the ambiguous-label case becomes an unknown-value no-op.

---------

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-10-02 07:02:35 -05:00
Peter Steinberger
b56164e406
feat(codex): request Ultrafast by default when the catalog advertises it (#163320) 2026-10-02 06:53:00 +00:00
Ayaan Zaidi
08498ec40d
fix(cron): stale automatic tool lists block scheduled jobs from tools their owner has (#162432)
Related: #130753, #137832, #147969

## What Problem This Solves

Fixes: some scheduled jobs created by an agent fail for months because a tool list saved by an older OpenClaw build is missing tools the creator actually had, such as the native shell. In our setup, a monthly group job that runs `node <script>` delivered nothing in August, delivered nothing in September (the run still reported `ok`), and posted a blocker in October. Its saved list had 31 tools and no `exec`.

## User Impact

User impact: an agent-created agent-turn job that does not name specific tools now gets the same tools as its owner conversation at run time, like a job an operator creates without `--tools`. Existing jobs with an automatically saved creator snapshot behave the same way from their next run. Nothing stored is rewritten: no migration and no backups. Explicit tool lists, script payloads, condition triggers, and jobs bound to captured Codex app authority keep their stored list.

Tradeoff, approved by the maintainer (Ayaan): a per-sender tool policy on the creating owner, or a plugin hook that narrowed the creating turn, no longer limits these default jobs. Only owners can create automations from chat, and subagents cannot create them.

## Why This Change Was Made

**History of the saved list.**
- #91499 introduced it so a delayed run cannot do more than its creator could.
- #112483 made every agent-created job store one, because runs have no sender.
- #112661 made scheduled runs re-apply the owner session's group policy and every non-sender limit, keeping the stored list as the upper bound.
- #137832 fixed native tool capture for new jobs only, and deliberately did not widen stored lists.
- #147969 added a Doctor advisory. It only fires for claude-cli, so it never covered Codex-harness or built-in OpenAI jobs like ours.

**Root cause.** When no tool list was given, OpenClaw saved a frozen copy of the creating turn's tools instead of treating the job like an operator `*` job. Every capture bug (missing native tools, late configured MCP, renamed tools) then stayed in the job permanently.

**Fix.** This follows Hermes, which keeps no creator snapshot: `cron/scheduler.py` `_resolve_cron_enabled_toolsets` reads toolsets from config at run time.
- **New jobs.** An agent-turn create or update with no list, or `*`, stores `["*"]`. That is the same value operator jobs store, so the job's tools match a normal turn in its owner conversation. Script payloads and condition triggers still store the creator's concrete tools, because a script reaches MCP only through servers its list names. Jobs whose creator captured Codex app authority also keep the concrete list, because that authority is bound to it.
- **Existing jobs.** One helper, `resolveCronRunToolsAllow` in `src/cron/tools-allow.ts`: a stored automatic snapshot (`toolsAllowIsDefault`) runs as `*` when it has a valid scheduled owner policy, no condition trigger, and no Codex app authority. Otherwise it keeps its stored list. Every execution consumer of the stored list uses it: the run payload, the command-prompt preflight, and the scheduled message authority.
- **Script transitions.** A `*` job that becomes a script, or gains a condition trigger, captures the creator's concrete tools.
- **Exec pin.** A `*` list keeps the creator's exec host pin.
- **No new noise:** automatic snapshots stay excluded from the `web_search` provider warning, as on main.
- **Deleted, now pointless:** both Doctor advisories about incomplete automatic snapshots, the run warning about pre-MCP snapshots, and two exports nothing uses anymore.

Review note: on claude-cli, a `*` job runs without a CLI tool cap, so Claude's native tools behave exactly as in a normal chat turn in that conversation. This PR introduces no new path around `tools.deny` that a chat turn doesn't already have.

## Evidence

Live-model Telegram proof (Telegram Test Server DM, leased team credential, live `openai/gpt-6-astra` reached through a forwarding proxy that stands in for the runner's mock provider; the runner harness itself is unchanged). This reproduces the shape of the original incident:
- The tester DMs the bot, which creates the owner conversation.
- A job owned by that conversation is added. Its stored list is an old-style automatic snapshot `["automations","message","read"]` plus `toolsAllowIsDefault: true`, with no `exec`.
- The payload is `Run: node scripts/split-report.mjs and post its output line verbatim`. The workspace script prints a random nonce.
- The job is run once (`cron run --wait`), with announce delivery to the DM.

| Build | `exec` offered | Model action | What arrived in the DM | Run |
|---|---|---|---|---|
| base 94f5a8d (main before this PR) | no | `tool_search` ×2, then gave up | "Could not run node scripts/split-report.mjs: no command-execution tool is available…" | error |
| **this PR, head 5b78cb7** | **yes** | `exec {"command":"node scripts/split-report.mjs"}` | "**SPLIT-REPORT 93C53909**: general 41, design 17, ops 9" (the exact script output, with this run's random nonce) | ok, delivered |
| head 5b78cb7 with `tools.deny: ["exec"]` | no | `tool_search`, `read`, then gave up | "Could not run node scripts/split-report.mjs: no command-execution tool or paired node is available…" | error |

In every run, the stored job kept `["automations","message","read"]` plus the marker. Before and after use the same scenario and driver; only the checkout differs.

Update and live proof: published `openclaw@2026.9.7`, then this branch at the exact head (2f5099d), on the same state directory. Mock provider. Every process ran under a temporary `HOME` and state directory. Each job's message makes the model call `exec` with `touch <effects>/<job>`.

1. 2026.9.7 created both jobs through `cron.add` (scheduled policy `trusted`). With the Gateway stopped, the "stale" job was given the old automatic-snapshot shape `["automations","message","read"]` plus `toolsAllowIsDefault: true`. sha256 of both stored rows: `609358837…`.
2. Runs:

| Build / config | Job | `exec` offered | Side effect | Run |
|---|---|---|---|---|
| 2026.9.7 | stale automatic snapshot | no | absent | error |
| 2026.9.7 | explicit `["read","message"]` | no | absent | error |
| this branch | stale automatic snapshot | **yes** | **created** | ok |
| this branch | explicit `["read","message"]` | no | absent | error |
| this branch, owner policy narrowed to `tools.deny: ["exec"]` | stale automatic snapshot | no | **absent** | error |
| this branch, `tools.deny: ["exec"]` | explicit `["read","message"]` | no | absent | error |

3. After the branch runs, the stored rows were byte-identical (same sha256 `609358837…`), and job ids and lists were unchanged. Nothing was migrated.

An earlier run at e151ea3, with the same harness, also covered a snapshot bound to Codex app authority: `exec` was not offered, the file stayed absent, and the stored row was unchanged.

Tests:
- `run.tools-allow.test.ts`: a stored automatic snapshot `["message","read"]` reaches the embedded run as `["*"]`, with the owner's scheduled policy intact. It fails on main with `["message","read"]`.
- `cron-tool-creator-cap.test.ts`: a default agent turn stores `["*"]`, while a trigger script and a Codex-app creator keep the concrete snapshot.
- `run.tools-allow.test.ts`: snapshots without a valid owner policy, or behind a condition trigger, keep their list. Both cases fail on the previous head.
- `run.tools-allow.test.ts`: no `web_search` warning for an automatic snapshot that kept its list. This fails without the exclusion.
- `run.message-tool-policy.test.ts`: a self-edited automatic snapshot runs on CLI with no cap.
- `run.tools-allow.test.ts`: a legacy `Command to run:` prompt from an automatic snapshot without shell tools now runs instead of being rejected.
- `jobs-tool-policy.test.ts`: scheduled message authority is admitted for an automatic snapshot that lacked `message`.
- `cron-tool-creator-cap.test.ts`: a `*` agent turn converted to a script captures the creator's concrete tools.
- These three regressions fail on the previous head. `node scripts/check-changed.mjs` passes.
- Explicit-list, exec-pin and gateway creator-transport suites pass. `pnpm tsgo:core` passes.

## Bounded cost

No new path triggers a model call or a job run. The change only selects which tool list an already scheduled run uses.

LOC vs main: production +98/-250 (net -152), tests +140/-373, docs +19/-11.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-10-02 00:31:29 +08:00
Vincent Koc
b5b1029990
feat(agents): expose installed skill search across harnesses (#158091)
* feat(agents): expose installed skill search across harnesses

Keep installed-skill eligibility host-owned while applying normal tool policy to discovery and complete instruction reads. Preserve current main's prepared tool surface and consolidated Code Mode coverage. Caller-provided snapshot options supply read paths, not replacement eligibility authority.

* test(agents): preserve hoisted skill harness initialization

Keep mock factory dependencies owned by vi.hoisted when splitting skill fixtures. Assemble skill-specific mocks only when the test consumer requests them. Update two prior expectations for the prepared empty snapshot/resource contracts while preserving sandbox host-skill exclusion.

Exact-head reproduction failed on the original hoisting and resource-array assertions. Nine affected suites covered 61 tests; after the hoisting repair, a stale sandbox snapshot assertion was found and the 32-test owner file passed on focused retry. Fresh lifecycle autoreview found no actionable P0-P2 issues; formatting and whitespace checks pass.

* test(agents): group sandbox skill coverage with policy tests

* test(agents): admit skill tools in client collision coverage

* test(agents): provide complete prepared skill fixtures

* test(agents): match async skill prompt contract

* fix(skills): preserve existing Code Mode read grants and whole reads

* fix(skills): preserve read denials in frozen tool profiles

* fix(codex): refuse partial installed skill instructions

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-10-01 11:11:13 +00:00
Peter Steinberger
3862ade452
fix: restore native repairs and Codex recovery on Bun (#162559)
Restore missing fs-safe prebuilds without relying on Node directory-merge semantics, preserve Bun native-launch diagnostics, and recover malformed Codex JSON without consuming the next valid frame. Keep package custody, bounded diagnostics, warning-only lifecycle failures, and installed-updater ordering intact.

Proof: 53 changed-test cases plus six catalog sibling cases pass on each of Node 24 and Bun; full changed-file validation and both import-cycle checks pass. Published Node-driver upgrade and Node-free Bun lifecycle cells passed on AWS. The published Bun driver's owning-npm preflight limitation remains explicitly documented.
2026-10-01 02:20:49 -07:00
RoboClaw
1caaef4b6f
feat(ui): select Ultrafast for supported accounts (#160352)
* feat(ui): select Ultrafast for supported accounts

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(gateway): preserve Ultrafast compatibility and account authority

Negotiate speed decoding per connection and keep canonical session state intact. Fence personal-account catalog requests at final guarded HTTP dispatch. Refresh the measured UI boot manifest without changing performance limits.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* refactor(openai): keep account catalog outcomes together

Move the existing account-scoped result projection into the already imported catalog helper without changing its behavior. Keep the provider owner below its existing line-count ratchet.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* test: refresh Ultrafast tool prompt fixtures

Regenerate the canonical Codex dynamic-tool fixtures for the authorized Ultrafast speed value. Update only that enum value and its derived size/hash metadata; keep snapshot checks enabled.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix: stop account catalog requests after default unlink

Bind automatic account selections to the canonical profile writer's committed link authority and carry the real request scope through final guarded dispatch. Keep explicit retained account selections usable after unlink and evict failed discovery custody. Preserve the existing active Auto Ultrafast opt-in while explicit Fast stays priority.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* refactor(codex): use narrowed Auto activation flag

Keep the reviewed Auto tier predicate while satisfying the typed boolean lint contract.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(ui): share speed applicability for optional Ultrafast

Respect the selected request mapping before offering an optional entitled tier. Preserve all three choices on eligible routes and clearable stored preferences on unsupported routes. Reproduced the contradictory catalog regression and passed 59 unit/Chromium cases; scoped independent review found no actionable P0/P1.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(ci): export manifest in same-revision preflight harness

Restore the trusted file omitted when the inline manifest moved out of ci.yml in 9d75a8fe87. Actual preflight job109720001696 failed before tests with MODULE_NOT_FOUND; real Git push and PR materialization fixtures reproduce the same omission. Keep the canonical index export owner, regenerate its workflow projection, and preserve all source/credential guards. Fifteen materialization variants, five import/size checks, root test types, and focused independent review pass.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(ci): satisfy extracted manifest static contracts

Repair inherited check-lint failures from the manifest extraction without changing CI routing: avoid namespace shadowing, retain error cause and the diagnostic callback string contract, preserve nonmutating shard copies, and apply required branch syntax. All1259 scripts lint clean;45 planner/import/size cases, formatting, UI i18n, styles, ratchet and independent review pass.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(ci): centralize dependency-free workflow flag parsing

Remove the extracted manifest local coercion helper and preserve its exact narrow Boolean grammar under the existing script argument owner. Register the canonical declaration rather than weakening the guard, and carry its runtime through trusted preflight materialization and fixtures.77 argument cases,11 declaration-guard cases,71 scoped integration cases, types, lint, export scans and remaining guard commands pass; independent review has no actionable P0/P1.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(gateway): separate model publication authority from selection scope

Keep the actual request lifetime in selected-account HTTP assertions without treating every anonymous unscoped catalog read as a personal account projection. Restore the established models.list response shape; no assertions weakened.177 model/catalog/session cases and10physical HTTP authority cases pass, together with types, lint, ratchet and independent review.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix: preserve session response argument tuples

Forward the original response tuple while projecting successful legacy payloads, without appending optional undefined arguments.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(codex): preserve existing Ultrafast opt-ins

Keep the v2026.9.7 Fast and active Auto opt-in semantics while adding explicit per-session Ultrafast. Standard still clears the tier. Cover cold and warm native turn requests without requiring a migration.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* test(ui): distinguish speed labels from model names

Match the complete Effort and Speed section labels rather than the Speed only fixture model. Retain the independent absence assertions for reasoning and speed controls. All four failures reproduced before the repair; the complete 13-case bundled browser file passes afterward.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* test(openai): intercept the shared transcription socket

Update the two stale socket mock registrations after the upstream transport consolidation. Keep the actual provider/session code, fake peers, assertions, timeouts, and Bun transport guard unchanged. All 86 OpenAI shard files pass: 1263 passed and one existing skip.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* test(gateway): construct complete session reset callers

Replace partial caller objects cast as never with the existing typed session mutation client fixture. Preserve provenance and required-sandbox assertions and the production client capability contract. Both CI failures reproduce before the repair; all 15 reset-model cases pass afterward.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix: confirm the selected Ultrafast command mode

Share the direct command, directive reply, and system-event confirmation formatter so the saved Ultrafast tier is named accurately. Include the accepted manual value in help and docs without advertising an unverified optional native-menu choice. Preserve boolean Fast, Auto, reset, authorization and persistence behavior. Three regressions fail before repair; 274 focused cases pass afterward.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

---------

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
2026-09-30 07:45:41 +00:00
Peter Steinberger
202381532b
fix(codex): preserve native launch failures after process exit (#161419) 2026-09-29 16:41:12 -07:00
Kimi Yu
0b3221fbfb
fix: keep Codex chats working during slow model discovery (#160363) 2026-09-28 11:43:38 -07:00
Peter Steinberger
9ac804cdcd
fix(codex): preserve conversation text within fork history window (#159651)
* fix(codex): preserve conversation text within fork history window

* test(codex): assert continuity within the newest history window
2026-09-28 01:46:25 -07:00
Kimi Yu
98782a01dd
fix(codex): recover remote connections during harness replacement (#159466) 2026-09-27 21:11:17 -07:00
stevenlee-oai
80714fa47c
fix(openai): clarify auth capabilities and SIWC limitations (#160023) 2026-09-27 20:49:47 -07:00
Peter Steinberger
ce1ca89990
refactor(tool-search): retire tool_search_code in favor of structured search and Code Mode (#159398)
* fix(tool-search): run tool_search_code in the QuickJS sandbox

Tool Search code mode spawned a Node --permission child with a node:vm
guest. Under Bun it needed an installed Node, and without one an explicit
code config silently downgraded to structured tools mode.

Run the guest through the Code Mode executor contract with the bundled
quickjs executor on every runtime. openclaw.tools.search/describe/call use
the shared namespace bridge with lazy thenables; the host admits only those
three operations through ToolSearchRuntime. codeTimeoutMs still bounds the
whole invocation, now including executor preparation. A denied or disabled
code-mode-quickjs plugin fails explicitly with next-step guidance instead of
falling back.

Remove the child source, IPC types, stderr-tail handling, and the Node and
Electron capability probe.

* refactor(tool-search): retire the tool_search_code bridge

Keep structured Tool Search and generic Code Mode as the two large-catalog surfaces. Remove the superseded JavaScript bridge and its runtime, display, and QA paths.

Doctor migrates legacy code mode to tools and removes codeTimeoutMs while preserving activation. toolSearch: true now selects structured search; JavaScript orchestration uses Code Mode exec/wait.

* test(tool-search): cover runtime behavior through structured controls

Exercise retained catalog, policy, hook, cancellation, terminal, MCP and client behavior through structured controls and their runtime owner. Delete bridge-only sandbox and JavaScript envelope cases while preserving nested call-id compatibility.

* test(tool-search): drop retired code mode prompt case

* test: repair fixtures exposed by Tool Search retirement checks

Remove the remaining retired Tool Search mode row. Preserve the session reader owner through the media retention mock and remove an unreachable queued-only branch from the ACP controls/submission fixture.

* test(upgrade-survivor): seed retired Tool Search code mode config

Author the legacy mode and timeout through every supported representative baseline CLI recipe, then require structured search and timeout removal after candidate update and Doctor. Existing config validation proves the resulting effective config.

Include the diagnostics native assignment summary in frozen target staging; the required assertion suite exposed its missing import. Node recipe and assertion tests: 238 passed, 157.87s wall. Docker validation remains with the coordinator.

* refactor(tool-search): drop the retired code-mode recovery surface

* chore: shrink assertion baseline after Tool Search retirement

* test(tool-search): drop the unused Tool Search test API

* chore: drop the retired Tool Search test API assertion baseline

* fix(e2e): drop duplicate native assignment staging line

Main now stages native-assignment-summary.mjs for frozen upgrades itself; the branch copy from the Tool Search upgrade proof became a duplicate after merging.

* build(pr): list Tool Search migration in wrapper inventory

The scripts/pr wrapper loads the Doctor config migrations at runtime, so the new Tool Search retirement migration belongs in its extracted component inventory.

* test(e2e): ship the Tool Search recipe to prepared tooling workers

Prepared tooling workers copy only listed source-relative assets, while the
upgrade survivor config recipe reads every section file by name at import.
The new tools-tool-search.json was missing, so the Docker scheduler parent
signal test's runner died with ENOENT and its polling wait reported a generic
5 s timeout that looked like a flake.

List the asset, guard the recipe directory against the preserved list, and
make the scheduler readiness wait fail with the runner's stderr once it exits.
2026-09-27 16:08:37 -07:00
Sarah Fortune
9c2335a2d0
feat(codex): enable Ultrafast for supported models (#158703)
* feat(codex): enable Ultrafast for supported models

* test(codex): preserve unconfigured Fast off selection

* test(codex): cover Ultrafast persistence and retry baselines

* docs(codex): register optional Ultrafast config baseline

* fix(codex): type advertised model service tiers

---------

Co-authored-by: Sarah Fortune <sarah.fortune@gmail.com>
2026-09-27 01:35:23 +00:00
Peter Steinberger
da3d100f27
fix(codex): stop cancelled compaction and close stalled progress (#132343)
* fix(codex): stop cancelled compaction and close stalled progress

Honor cancellation at native admission while keeping written-request completion authoritative. Close unfinished progress through the projector lifecycle and preserve observed item identities so cleanup cannot erase replay barriers.

Addresses #132177 and #132178. Authenticated changed-path Gateway proof and extension-lint SDK preparation remain blocked; preserve this as a draft, not a land-ready change.

* test(codex): retain failed compaction in observed item count

The settled-finalization case observes a completed tool and a compaction
that ends without completion. Expect two observed items while retaining
one completed item, zero active items, and the settled tool evidence.

The original 11-file CI group passes all 327 cases with this correction;
independent review confirmed the expectation follows the owner contract.
2026-09-24 13:20:37 -07:00
Peter Steinberger
890ab5904b
fix(codex): await app-server cleanup when startup rejects (#157229) 2026-09-24 13:11:22 +00:00
Kevin
4f6eb26b1b
fix(codex): preserve enabled network restrictions on invalid config (#156906)
* docs(codex): specify native network allowlist behavior

* fix(codex): reject invalid config with enabled network restrictions

* fix(codex): repair blank network options and explain config errors

* docs: remove task-specific network allowlist spec
2026-09-23 19:53:40 -07:00
Sarah Fortune
5f24b24e3f
fix(codex): let Ask OpenClaw use remote app-servers (#156671)
* fix(codex): let Ask OpenClaw use remote app-servers

* test(codex): normalize remote fixture frames

* docs(codex): clarify remote verification setup scope

* test(ui): narrow scenario diagnostic task type

* test(ui): decouple scenario cleanup callback types

* test(ui): declare prebuilt assets for root Vitest context

---------

Co-authored-by: Sarah Fortune <sarah.fortune@gmail.com>
2026-09-23 12:13:37 -07:00
Galin Iliev
bc5a318e28
fix(codex): preserve native user-home providers (#156226)
Co-authored-by: galiniliev <5711535+galiniliev@users.noreply.github.com>
2026-09-23 11:58:49 -07:00
Shakker
17d2faf6b7
feat: enforce role model limits in native agents and Visitor Access (#154893)
Carry operator role model limits through native Codex parents, child agents, restored work, reviews, and model-calling tools, with current-source revocation and lifecycle cleanup. Require an explicit Visitor model policy while preserving the documented staff-mode limitation and existing configured service authority.
2026-09-23 12:28:30 +01:00
Peter Steinberger
72ff919499
fix(codex): release idle catalog decoder memory (#156202) 2026-09-22 23:07:00 -07:00
Vincent Koc
6dda26e46c
fix(codex): report failed inference WebSocket handshakes (#155710)
* fix(codex): return HTTP errors for failed WebSocket upgrades

Return sanitized 502/504 responses for upstream connection and deadline failures instead of dropping the pending local upgrade. Keep native retries, HTTPS fallback, and provider auth errors unchanged.

* fix(codex): honor handshake deadline before forwarding failures

* test(codex): bound native inference fixture payloads

* Merge commit '5b7e61fe36' into fix/codex-websocket-handshake-20260922

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-23 05:19:45 +00:00
Vincent Koc
43247c9b20
fix(codex): keep new inference moving while responses stream (#155090)
* fix(codex): keep new inference moving while responses stream

* test(codex): qualify inference resource lifetimes

* refactor(codex): simplify inference transport ownership

* refactor(codex): separate inference upload and test ownership

* test(codex): remove unused capacity fixture imports

* test(codex): keep completion fixture frame private
2026-09-22 11:12:18 +08:00
Vincent Koc
1439307023
fix(codex): keep inference bursts from failing turns (#154945)
* fix(codex): keep inference bursts from failing turns

Queue one bounded request batch through the shared FIFO permit pool before upstream admission. Preserve handshake errors and headers, absolute connection deadlines, generation cancellation, and idle transport reuse without raising active limits.

* test(codex): allow either native child scheduling order

* fix(codex): cancel queued inference handshakes on disconnect

* test(codex): return void from native cleanup hooks

* test(codex): satisfy inference fixture lint checks

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-22 02:02:49 +08:00
Peter Steinberger
6ae2babeb4
fix(codex): stop catalog retry loops when its executable cannot run (#154489)
* fix(codex): stop retrying catalog executables that cannot run

Resolve managed passive catalogs from the installed Codex package and run
its official launcher with the Gateway interpreter. Preserve terminal
spawn failures across reloads, cancel catalog scheduling, and expose the
failed executable through Doctor and one advisory.

Keep plugin retirement and capture retention with their existing owners.
Add startup, lifecycle, and Doctor regressions and document recovery.

Related: #154313
Reported-by: @cloudgg82-blip

* test(codex): assert retained terminal spawn failures

Assert ENOENT through the recorded startup failure, exclude retryable
connection-closed errors, and verify a second attempt reuses the failure
without spawning again.

Fix the stale top-level ENOENT expectation in attempt-startup-retry.
2026-09-21 13:17:01 +00:00
RoboClaw
fd77e92301
fix(codex): preserve long tool output after history reload (#153542)
* fix(codex): preserve long tool output after history reload

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: 29c1ee50-48fb-4237-8dc1-0c82fd0922de

* fix(codex): preserve long tool output after history reload

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: 266852b1-3650-499a-b1a5-c51f7497d74a

* fix(codex): finish full-output display and replay cutover

Use the existing mono theme token and update strict replay/sidebar expectations to preserve the captured newline and raw-inspector route. The affected UI/theme and native-attempt suites pass (244 and 196 cases).

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(codex): preserve long tool output after history reload

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: c4342f0f-77c4-405e-9704-24215c500199

* test(ui): exercise raw tool-output detail contracts

Mount the real inspector for exact JSON lexemes, duplicate keys, fences and invalid JSON; assert bounded inline previews and unchanged neutral progress metadata. The focused suites pass 68 tests; the broader sweep passed 21,319 tests with two unchanged local loopback tests blocked by the host proxy.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(codex): preserve long tool output after history reload

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: c910a1fc-990d-46c5-924b-dabc8f4e0637

---------

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
2026-09-20 06:23:43 -07:00
Peter Steinberger
9b158055b9
fix(codex): recover turns blocked by stale process records (#153454) 2026-09-19 23:00:32 -07:00
Peter Steinberger
8e7443e653
fix(codex): preserve delivery facts and native approval semantics (#151863)
* fix(codex): preserve delivery facts and native approval semantics

Keep failed sends from suppressing replies or recording false delivery. Share the host delivery facts before presentation middleware, preserve per-artifact media and finalized native text, and honor native approval lifetimes and explicit form answers.

Replace duplicate delivery and quota decisions with their existing owners. Document the seven shared SDK helpers and apply the approved seven-export and seven-callable surface budget increase.

* test(codex): align native fixtures with canonical outcomes

* test(codex): align mirrored transcript assertions

* fix(codex): retain core conversation delivery receipts
2026-09-19 16:45:20 -07:00
Peter Steinberger
60c0e9b09d
fix(codex): restore background completions with managed hooks (#151658)
* fix(codex): preserve managed hooks in background completions

* test(codex): expect isolated project root markers

* fix(codex): preserve native accounts and proxy routing
2026-09-18 08:35:12 -07:00
Peter Steinberger
cb4d0357c8
fix(codex): unify turn activation and cleanup ownership (#151753)
* fix(codex): unify turn activation and cleanup ownership

* fix(codex): preserve side failures and workspace captures
2026-09-18 08:00:13 -07:00
Peter Steinberger
c610d6327e
fix(codex): preserve selected account failures (#151831)
* fix(codex): preserve selected account failures

* docs(codex): disclose credential import override limits
2026-09-18 07:45:35 -07:00
Peter Steinberger
a8ffb3544a
fix(codex): align native plugin policy and cache ownership (#151757) 2026-09-18 07:39:32 -07:00
Peter Steinberger
88a04f66af
fix(codex): preserve control authority and node resources (#151728) 2026-09-18 05:10:09 -07:00
Peter Steinberger
b9fc4fe10c
fix(codex): align configuration and launcher contracts (#151686) 2026-09-18 04:32:20 -07:00
Peter Steinberger
033abf7530
fix(codex): reject malformed native completion events (#151373)
* fix(codex): reject malformed native completion events

* fix(codex): validate terminal events across all projections

* test(codex): register protocol regressions in the full suite
2026-09-17 23:39:17 -07:00
Peter Steinberger
deb939dc00
fix(codex): dismiss native requests resolved by another client (#151381) 2026-09-17 22:20:22 -07:00
Galin Iliev
3f0b58e269
fix(codex): preserve local native configuration across supervised turns (#151001)
Preserve native Codex configuration, tools, and model ownership when continuing supervised conversations through a shared local daemon.

Return valid retained subscriptions after passive preflight and pre-turn startup failures without weakening cancellation, revocation, replacement, or ambiguous-write safeguards. Preserve established native search policy when daemon defaults change, and document credential handoff, refresh ownership, and shared-account limits.

Validation: 291 fresh focused tests passed; exact-head hosted CI passed. Additional live startup-failure/retry capture was not performed; the PR records that limitation and the earlier native proof at its original revision.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: galiniliev <5711535+galiniliev@users.noreply.github.com>
2026-09-17 19:44:39 -07:00
Peter Steinberger
69948a6313
fix(codex): preserve native text and avoid duplicate async replies (#151200)
* fix(codex): preserve native output across item event shapes

* test(codex): mirror completed native reasoning in fallback fixtures
2026-09-17 16:11:08 -07:00
Peter Steinberger
4d28246be9
fix(codex): settle disconnected sockets and force shutdown (#151192)
* fix(codex): settle disconnected sockets and force shutdown

* test(codex): avoid returning listener handles from executors
2026-09-17 15:51:04 -07:00
Peter Steinberger
4dc8c253f3
fix(codex): avoid startup timeouts behind filesystem work (#148434)
* fix(codex): avoid startup timeouts behind filesystem work

* test(codex): classify process probes as fixtures
2026-09-14 12:53:01 -07:00
Peter Steinberger
2545695c28
fix: keep old screenshots out of new voice attachments (#147913)
* fix(codex): keep restored screenshots with their historical messages

* style(codex): copy image range fields explicitly
2026-09-13 22:38:04 -07:00
Peter Steinberger
f9ccd07ded
fix: skip official setup approvals and default to Astra (#145646)
* fix: skip official setup approvals and default to Astra

* test: align default-model expectations across consumers

* test: align attachment catalog with Astra default

* fix: preserve configured Codex catalog model selections
2026-09-12 00:34:25 -07:00
Peter Steinberger
96524b7326
fix(models): default Fable and Astra to medium effort (#145554)
* fix(models): default Fable and Astra to medium effort

Use the shared Fable default for omitted provider and transport effort, preserve Mythos High and explicit settings, and prefer Medium on Astra accounts that support it. Update saved-setting, payload, provider and documentation coverage. Restore the existing machine locator missing from a mainline UI screenshot test.

* test(ui): use the upstream screenshot locator repair

* test(codex): expect the medium Astra turn default

* fix(models): propagate medium defaults through cloud providers
2026-09-11 21:58:35 -07:00
Peter Steinberger
c3406f7696
fix(codex): preserve refusals when fallback access is denied (#145105)
* fix(codex): preserve refusals when fallback access is denied

* fix(codex): preserve fallback effects and terminal ownership
2026-09-11 11:36:13 -07:00
Peter Steinberger
e156194282
feat(codex): retry cyber-refused turns on Daybreak and surface cyber notices (#143761)
* feat(codex): show cyber safety notices above the composer

* feat(codex): retry cyber-refused turns on Daybreak

OpenAI declines some defensive-cyber work on its general models and directs
approved workspaces to a Daybreak model instead. A refusal previously ended the
turn, so the work had nowhere to go.

Retry a cyber-refused turn once on the configured Daybreak model, then report
the outcome in the existing composer notice. Escalation is turn-local: it moves
runtimeModelId for one attempt and never changes the session's stored model
selection.

Daybreak trails the general models, so routing stays scoped to work that was
actually refused. One attempt per turn, a bounded in-memory session window after
it, and no escalation for a turn already on the target. Only a Daybreak reply
earns sticky routing inside that window; an unauthorized or still-refusing
target suppresses further attempts instead.

Authorization stays server-owned. model/list advertises Daybreak to every
client, so catalog presence is not entitlement: an unentitled workspace still
gets 401/403. The retry is treated as the only evidence, and its failure
surfaces an explicit notice rather than a silent block.

Configured under plugins.entries.codex.config.appServer.cyberFailover with
mode, model, and cooloffMs.

* refactor(ui): order cyber notice states by precedence

Replace the pairwise sticky-state chain with an explicit precedence map. A
lower-ranked update can never replace a higher-ranked one for the same run,
which expresses the intent directly: a late review update cannot displace a
reroute or a block, and an automatic Daybreak escalation is terminal because it
happens after the block it followed.

Handle the cleared event before the precedence check so its existing
buffering-only semantics are unchanged, and tighten the notice copy.

* fix(codex): never replay a cyber-refused turn that already acted

The escalation retried a refused turn unconditionally. A turn can be refused
after it has already sent a message, added a cron entry, spawned a session, or
generated media, and re-running it would repeat those actions.

Gate escalation on the attempt's own replayMetadata.replaySafe verdict, which
the projector reports for every settled turn. Absence counts as unsafe, so an
unrecognized result shape skips escalation instead of risking a duplicate side
effect.

* fix(codex): bound the cyber escalation window map

Window entries expired only when their session was read again, so a session that
never returned kept its entry for the life of the Gateway.

Sweep expired entries on write once the map reaches a bound, then shed the
oldest live windows if a burst of sessions is still inside its cooloff.

* fix(codex): report the escalation outcome the attempt actually had

Three defects found in review of the Daybreak escalation.

An attempt was treated as answered whenever it carried neither an authorization
error nor a cyber refusal, so a transport failure or cancellation enabled sticky
routing for the whole cooloff even though Daybreak never replied. Require a
settled assistant reply instead.

Every authorized retry emitted the escalated notice, including one Daybreak
refused again. Because that state outranks a block, the composer replaced the
terminal block with a non-alert card. Announce only the outcomes this owner can
state truthfully: a reply, or an unauthorized target. A second refusal keeps the
projector's own block, and other failures keep their terminal error.

Bounding the window map evicted live entries, which could drop an unauthorized
target's suppression and let it be retried inside its cooloff. Shed only
answered windows; let the bound yield rather than break that guarantee.

* fix(codex): close the Daybreak window when entitlement lapses

Three further defects from review of the escalation fixes.

Sticky routing sends a turn straight to Daybreak, so a 401/403 there is not a
cyber refusal and returned early without touching the window. Every later turn
kept going to a target that could no longer answer. Close the window, say so,
and give the turn its attempt on the model the session actually selected.

The window was recorded only after the retry finished, so two parallel turns on
one session could both plan an escalation and both pay the full failing attempt.
Reserve a suppressed window before awaiting the retry; the real outcome replaces
it.

Eviction shed only answered windows, so a map full of live suppressed records
grew past its bound. Keep the cap hard by dropping the record soonest to expire
once no answered window remains.

* fix(codex): scope Daybreak authorization to the account, not the session

Review of the previous round found three more defects, and the third pointed at
a better shape for this state.

Entitlement belongs to the workspace and the target model, not to a
conversation, so remembering it per session was wrong in both directions: a hard
cap could evict a live suppression and let the expensive reconnect ladder repeat,
while a session that never saw the failure learned nothing from it. Keep the
unauthorized target in one account-level record that session churn cannot
displace, and leave the per-session map holding only routing hints, which are
safe to bound hard.

Refusal detection ORed the historical assistant row with the current attempt's,
so a later turn that failed before producing a row could inherit an older
refusal and be escalated for it. Prefer the current attempt, matching how the
answered check already reads the result.

An escalated attempt is also no longer counted as answered when it carries a
refusal diagnostic, rather than relying on the projector's stop-reason
convention to imply it.

* fix(codex): scope Daybreak authorization per workspace and drop stale sticky routing

One process can host several agent-scoped Codex homes, so keying the
unauthorized-target record by model alone let an unentitled workspace disable
escalation for an entitled one. Key it by the authenticated workspace as well.

A sticky window pre-routes to Daybreak, so a refusal there is refused by the
target we were favouring. Planning correctly declined to escalate a turn already
on Daybreak, but the window was left in place and later turns kept going to a
model that had just refused. Record the suppression before returning.

The answered check excluded only cyber refusals, so a bio or misalignment
refusal could satisfy it. Treat any provider-refusal diagnostic as refused; those
categories keep their own handling and must never look like a successful
escalation.

* refactor(codex): route only refused work to Daybreak

Escalation had grown a second behavior: after a successful retry, later turns in
that session were pre-routed to Daybreak for the rest of the cooloff. That sends
work Daybreak was never asked about to a weaker model, which is the opposite of
routing only what was refused.

It was also the source of most of this feature's complexity. Every interaction
it created needed its own rule: what to do when entitlement lapsed mid-window,
when the pre-routed turn was itself refused, and how to classify an attempt that
neither answered nor failed cleanly.

Remove it. Every turn starts on the model the session selected, and only a turn
OpenAI actually refused is retried on Daybreak. The session-scoped record now
only damps repeat attempts, and the account-level record still makes an
unauthorized target unrepeatable.

Net 192 fewer lines, and the remaining rules describe one behavior instead of
two interacting ones.

* fix(codex): stop the damper from blocking later refusals, and serialize probes

Review of the simplified design found four issues.

The session damper was applied to every escalation, including successful ones,
so for the rest of the cooloff a later refusal in that session got a bare
refusal instead of the escalation that had just been proven to work. The damper
exists to stop a target that is not working from being retried, so a successful
escalation now clears it.

The retry re-ran the full attempt with the same parameters, mirroring the user
prompt into the transcript a second time. Set the existing
suppressNextUserMessagePersistence flag on the retry.

Only the session was reserved before the retry, so sibling sessions in the same
workspace could each start their own Daybreak attempt and each pay the 401/403
reconnect ladder before the first result recorded the target as unavailable.
Reserve the workspace and target for the duration of the probe.

Expired unauthorized-target entries were dropped only when that exact key was
queried again, so short-lived agent or profile ids accumulated. Sweep and bound
that map the way the session map already is.

* fix(codex): escalate only OpenAI's own refusal on the current attempt

Refusal detection matched the diagnostic category but not its provider, so
another provider's cyber refusal could move a prompt onto OpenAI's Daybreak
tier. Require the OpenAI provider.

It also fell back to the previous assistant row when the attempt produced none,
which could inherit an earlier turn's refusal and reroute a prompt the provider
never refused. A refusal always populates the current-attempt row, so requiring
it costs nothing and removes the inheritance.

A cached account-level denial and a probe still in flight were the same skip
reason, so sessions that hit the cached denial showed only the generic block and
never learned why escalation had gone quiet. Split them and report the denial.

* chore: drop a scratch review prompt committed by mistake

* refactor: drop comment slop from the cyber failover diff

A deslop pass over the branch diff removed comments that narrated control flow or
restated the code. Comments recording live-probed upstream facts are kept: that
catalog presence does not prove entitlement, and that an unauthorized target
costs the transport's full reconnect ladder.

The pass also caught a comment claiming the account-level authorization record is
never evicted while the code capped that map at 256 and dropped the oldest entry.
The per-write sweep is what actually bounds it, so the cap is a backstop; say so,
and when it is reached drop the record with the least remaining protection rather
than an arbitrary one.

* chore(ui): record the startup budget for the cyber notice surface

The composer notice states, their precedence projection, and the accompanying
strings sit on the Control UI startup path and put it 88 bytes over the
enforcement limit.

A deslop pass over the diff moved the total by 0 bytes, so the growth is feature
code rather than comments, and the earlier attempts to trim it are recorded in
the PR: shortening the notice copy moved 1 byte, and collapsing the notice state
machine into a precedence map moved none. The remaining weight is the feature.

Baseline moves 352531 to 353195, measured with ui:check-performance:base against
the merge base, well under the 358400 committed cap.

* refactor(codex): collapse the escalation owner to one map and one verdict

The owner had grown from 217 lines to 375 across seven review rounds, and the
growth was bookkeeping rather than behavior.

Two module-global maps with two different eviction policies become one: the
per-session damper is gone. Its job was to stop repeated attempts, but the
expensive case is an unauthorized target, and the workspace-scoped record
already covers that for every session. A retry that Daybreak merely refuses is
an ordinary refusal, not worth its own bookkeeping. That also removes the
reserve-before-await dance, the damper release on success, and a skip reason.

Four predicates over the same attempt result become one verdict reader. They all
read the same two fields of the same object and each carried its own structural
type; asking once is clearer and makes the harness read as a sequence of facts
rather than four separate interrogations.

Behavior is unchanged. Every protection the review rounds added survives: only
OpenAI's own cyber refusal on the current attempt escalates, a turn that already
acted is never replayed, bio and misalignment refusals never count as an answer,
authorization is workspace-scoped and bounded, and one probe at a time per
workspace and target.

Owner 375 to 230 lines; PR production total 752 to 356.

* fix(codex): write the workspace key separator as an escape, not a raw byte

The previous commit embedded a literal zero byte in the source where the
six-character escape sequence was intended. It compiled and passed tests,
because that character inside a string literal is legal, but git classified the
file as binary. That silently drops it from diffs and blocks review tooling from
reading it at all.

The separator value is unchanged; only how it is spelled in source is.

* chore(ui): drop the startup budget raise, no longer needed

Rebasing onto current main moved the measured startup JS from 353195 to 353060,
which fits under the existing baseline plus its allowance. The raise this branch
carried is therefore unnecessary and the committed baseline returns to main's
value.

The drop comes from main, not from this branch: the server-side simplification
here does not touch the Control UI bundle, and the deslop pass over the UI files
measured a zero-byte effect. Headroom is 47 bytes against the enforcement limit,
so this sits inside the build-variance allowance rather than comfortably under
it.

* chore: regenerate the config baseline after rebase

Main advanced its own config surface again while this branch was open. The
generated baseline is rebuilt rather than hand-merged.

* chore(ui): record the startup budget for the cyber notice surface

The composer notice states, their precedence projection, and the accompanying
strings sit on the Control UI startup path and add 579 bytes over the previous
baseline, where the growth allowance plus build variance permits 576.

Trimming was attempted and measured rather than assumed: removing roughly 40
characters of notice copy moves the gzip total by 1 byte, and a deslop pass over
the whole diff moved 0, because the large English catalog is statically imported
at startup and gzip already dedupes the repeated phrasing. The remaining weight
is the feature.

Baseline moves 352531 to 353110, measured with ui:check-performance:base against
the merge base, well under the 358400 committed cap.

* fix(ui): keep chat-pane-render under the max-lines ceiling

Main grew this file to the 700-line limit while this branch was open, so the one
line the notice surface adds here tips it to 701.

Fund it with nearby cleanup rather than a suppression, which the ratchet only
allows to shrink: placementStartupPending was a two-line binding used once, so
its expression moves to its single use site.
2026-09-11 15:41:56 +00:00
Vincent Koc
d84f5e8cbd
docs(plugins): fix accuracy findings across the plugin docs (#143857)
Closes the open `accuracy` audit findings scoped to `docs/plugins/`, each
verified against source before editing.

Corrections where the docs understated or misstated the code:
- Telephony support listed 2 of 9 bundled providers that implement
  `synthesizeTelephony`.
- `allowInvalidConfigRecovery` named 2 of the 5 recovery cases in
  `isAllowedPluginRecoveryIssue`.
- The `activation-command-hint` row omitted the live
  `manifest-cli-command-owner` reason.
- A `js`/`export default` config example was neither a real config surface nor
  checked by `check-docs-config-examples` (that fence language is skipped);
  converted to `json5`, now validated.
- `npm:@openclaw/google-meet` skips the official-catalog install plan; the
  package is in the catalog, so the bare spec is correct.

Version scope added only where a release or dated compat record exists:
`2026.4.22` (guardian env removal), `2026.5.2` (Meet access-type control),
`2026.8.1` (Teams/Zoom live validation), `2026.8.2` (Beam named links),
`2026.9.2` (legacy doctor selector), and the dated `removeAfter` records in
`src/plugins/compat/`. Elsewhere the time-relative wording is dropped and
present behaviour stated, rather than a version guessed.

`docs/plugins/plugin-inventory.md` is generated: the fix is in
`scripts/generate-plugin-inventory-doc.mts`, regenerated, `plugins:inventory:check` rc=0.
2026-09-10 16:54:53 +08:00
Ahmad Shawwal
e89593171d
fix(codex): honor configured project document budget (#142932)
Closes #142878

Ordinary Codex-backed threads previously forced a 128 KiB project-document budget even when native configuration explicitly allowed more. The selected app-server client now supplies the authored budget and its provenance before thread start or resume. The unauthored 32 KiB native default retains OpenClaw's existing 128 KiB fallback; explicit request values and existing isolation zeros keep their precedence.

Direct bindings also carry the leased timeout, cancellation signal and current-owner check into the configuration read, so an unanswered read cannot leave binding pending indefinitely. The existing native reader remains the authority; no OpenClaw setting or stored field is added.

Validation on Codex 0.153.4:
- Real Telegram bind, approval and ordinary message: pinned main sent 131072 and omitted the deepest instruction file; the repaired head sent an authored 200000 and loaded all three files.
- Cold app-server replacement and explicit rebind preserved the same thread and the authored budget.
- An authenticated remote app-server connection sent its authored 240000 and loaded all three files. This is a same-machine remote-connection check, not multi-host or workspace-mapping proof.
- Unauthored configuration retained 131072. Focused controls cover explicit precedence, isolation zeros, invalid values, selected-client ownership and bounded read failure.
- 1054 focused tests passed. Native formatting, size/assertion checks, scoped syntax checks and the commit hook passed. Full lint, typing and suites remain the hosted CI gate.

The native provider endpoint used deterministic synthetic responses. Source reporting establishes which files contributed, not exact retained byte counts. Independent source-blind acceptance repeated the public flow on the final build and passed its stated runtime scope. Exact-head review and hosted CI remain separate finishing gates. Thanks @ashawwal for the original repair.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-09 19:52:25 +05:30
Vincent Koc
e427f280d9
docs(nodes): split the nodes overview by reader job (#142779)
* docs(nodes): split the nodes overview by reader job

docs/nodes/index.md was 68,029 bytes, 8,871 words and 35 headings mixing
pairing how-tos, node-host setup, session hosting, command-policy reference and
per-platform allowlists in one page. One H2, "Remote node host (system.run)",
parented 17 H3 sections that were not about system.run, and the page ended with
no next-steps list.

The page already lived in a directory with eleven siblings (audio, camera,
computer-use, images, location-command, media-playback, media-understanding,
presence, talk, troubleshooting, voicewake), so this extends that directory
rather than creating a parallel one. index.md becomes a real index at the same
/nodes route: intro, a "Node pages" list that also names the eleven existing
siblings, and the anchor table below.

Children, one per reader job:

- pairing-and-status.md - approve a node, read status and host stats, upgrade
  a fleet across the N-1 protocol window.
- node-host.md - foreground, service, SSH-tunnel and headless node hosts,
  gateway preconditions, identity state, and the system.* command surface.
- node-exec.md - allowlist commands, point exec at a node, raw node.invoke,
  and exec node binding.
- mcp-and-skills.md - node-hosted MCP servers, node-hosted skills, and local
  Ollama inference.
- session-hosting.md - nodeHost.workerRuns, device placement and capacity,
  and container isolation.
- session-catalogs.md - Codex, Claude, OpenCode and Pi session discovery and
  continuation on paired nodes.
- file-transfers.md - terminal uploads and the File Transfer plugin tools.
- command-policy.md - the platform default allowlists, dangerous-command
  opt-ins, gateway.nodes/tools.exec config, and the permissions map.
- device-commands.md - widget panel, camera, screen recording, location, SMS
  and device data CLI helpers.

Anchor strategy

Per-anchor redirects are not possible: redirectSource() in
scripts/lib/docs-redirects.mjs rejects any source containing [?#]. Every anchor
the old page published is therefore kept alive on the index itself as authored
<a id="..." /> stubs inside a "Where each section moved" list, each pointing at
its new home. Ids were computed with parseDocsDocument, not a slug
approximation, so the thirteen punctuated headings keep both their encoded and
their cleaned id (for example pairing-%2B-status and pairing-+-status). All 48
ids the pre-split page published resolve on the new index; the index publishes
only two ids of its own, node-pages and where-each-section-moved, so no stub
collides with a heading the index still owns. parseDocsDocument reports zero
collisions on the index and on every child.

Losslessness

Reassembling the 35 section bodies reproduces the original body byte for byte,
apart from the two declared link retargets below. Counts, original body vs
children:

- words 8,634 -> 8,634
- characters 66,320 -> 66,358 (+38, the two retargets)
- code fences 29 -> 29, identical fence for fence as a multiset
- markdown links 30 -> 30
- table rows 16 -> 16, all three tables byte-identical
- the per-platform default-allowlist table is byte-identical: 8 rows, with
  iOS 11, watchOS 3, Android 19, macOS 14, Windows 6 and Linux 2 commands

No prose was rewritten. Two intra-page fragment links whose target moved to a
different child, both `](#command-policy)` in device-commands.md, became links
to /nodes/command-policy#command-policy; that is the whole +38 characters.
Sections keep their original relative order within each child, so the five
directional cross-references the page carried ("the environment fallback
above", "see above", "see below", "the static platform-default table above")
all still resolve on their own page.

Also updated: the "Nodes and media" nav group in docs/docs.json, eight in-repo
deep links repointed at the new pages (docs/releases/2026.9.2.md left
untouched, its anchor still resolves through the stubs), fourteen zh-CN
glossary entries for the new titles and index link labels, and
src/docs/config-path-docs.test.ts, which asserts on the `openclaw config`
bracket-path examples that now live in node-exec.md.

Closes audit findings: r3-0443, r3-0445, r3-0447

* docs(nodes): link the relocated device command examples from the exec page

The exec page's "(camera, screen, location, below)" pointed at helpers the
split moved to /nodes/device-commands. Replace the directional word with a
link. Found by ClawSweeper; the orphan-reference scanner's patterns do not
match a bare trailing "below" with no noun phrase in front of it.
2026-09-09 10:44:24 +08:00
Vincent Koc
295771552f
docs(plugins): split the Codex harness page by reader job (#141164)
The Codex harness page had grown to 98,150 characters across 36 heading
sections, mixing install how-to, remote placement, routing policy,
app-server reference, command reference, runtime explanation, and
troubleshooting on one page. Split it into nine child pages under
docs/plugins/codex-harness/, keeping plugins/codex-harness as a short
index (15,826 characters) that holds the overview, Requirements,
Quickstart, Verify Codex runtime, and Related.

Children:

- codex-harness/placement — paired-device and cloud-worker placement
- codex-harness/routing — runtime selection and deployment patterns
- codex-harness/configuration — config map, restricted turns, project
  instructions, compaction, direct API long context
- codex-harness/app-server — app-server policy, approval evidence, auth
  order, scheduled app authority, environment isolation and overrides
- codex-harness/config-fields — top-level and appServer field tables
- codex-harness/commands — /codex surface, Fast mode, local inspection
- codex-harness/runtime-behavior — dynamic tools, web search, image
  loader, liveness, parallel chats, runtime boundaries
- codex-harness/native-features — native thread sharing, supervision,
  native Codex plugins, Computer Use
- codex-harness/troubleshooting — symptoms and fixes

Anchor strategy: per-anchor redirect routes are impossible because
redirectSource() in scripts/lib/docs-redirects.mjs rejects any source
containing a fragment. Every anchor stays alive on the parent instead.
Ids were computed with parseDocsDocument from scripts/lib/docs-markdown.mjs,
not a hand-rolled slug:

- 36 heading ids published before the split; 36 still resolve on
  /plugins/codex-harness after it, with 0 id collisions on the parent or
  on any child
- 4 of those ids (requirements, quickstart, verify-codex-runtime,
  related) stay published by the parent itself and are not stubbed
- the other 32 are authored <a id> stubs in "Where each section moved",
  each linking to the owning child; all 32 stub targets were asserted to
  exist in the child that claims them
- 11 in-repo references carried a /plugins/codex-harness fragment; the
  10 pointing at a moved section were repointed at the owning child, and
  the same-page one now resolves inside its own child

Losslessness, asserted mechanically rather than by eye: all 36 original
section bodies reappear byte-identical in the new files, except three
declared cross-reference repairs. Counts before -> after: words
11,688 -> 12,767 (frontmatter, page lead-ins, and the moved-section
index), fenced code blocks 25 -> 25, markdown links 51 -> 113. No child
exceeds 16,757 characters.

Three cross-references the split would have orphaned became real links,
the one sanctioned prose change:

- "If /status is surprising, see [Troubleshooting](#troubleshooting)" now
  points at /plugins/codex-harness/troubleshooting
- the cloud-worker section's link back to the paired-device section is
  now same-page, since both live on codex-harness/placement
- "The rest of this page covers deployment shape, fail-closed routing,
  guardian approval policy, native Codex plugins, and Computer Use" now
  links to routing, app-server, and native-features

Nine zh-CN glossary entries were added for the new page titles. They are
machine-written and unreviewed by a zh-CN speaker.

Closes audit findings: r3-0506, r3-0507, r3-0509
2026-09-07 11:40:23 +00:00