Commit graph

18123 commits

Author SHA1 Message Date
Peter Steinberger
6e987af9bc
fix(launcher): keep the Node compile-cache path short on Windows (#162841)
2026.9.7 lengthened the compile-cache marker from <mtime>-<size> to build-<buildId>, which pushed cache paths on Windows into the directory-length window where Node 24's module.enableCompileCache() never returns (nodejs/node#66438); every OpenClaw process start could then spin at full CPU. The shared cache owner now uses a 16-character hash of the full build id as the marker, refuses Windows cache paths over 200 characters with a recorded warning (cache off for that process), and no longer propagates an unsafe inherited cache path to children. Legacy namespace handling and POSIX behavior are unchanged; native Windows/Node 24.18 completes the packaged CLI with a 220-character TEMP in 2.4 s.

Closes #162821
2026-10-01 11:29:41 -07:00
RoboClaw
b0632d0eff
fix(agents): explain paused progress-card checklists (#162903)
Keep blocked work visible in a Markdown-only card when no authorized step can proceed, rather than repeatedly saving unfinished checklist steps. Preserve the bounded completion check unchanged.\n\nRefs #162878

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: RomneyDa <6581799+RomneyDa@users.noreply.github.com>
Co-authored-by: fuller-stack-dev <263060202+fuller-stack-dev@users.noreply.github.com>
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>
2026-10-01 11:27:00 -07:00
Peter Steinberger
d89da7595d
fix(cron): keep legacy repair in Doctor (#160448)
Runtime preserves authored cron definitions and consumes canonical ownership, delivery, and schedules. Doctor performs supported repairs with verified backups, while update rehearsals retain live legacy sources for the live import. Preserve durable receipts, explicit scope boundaries, and deletion fences.

Published 2026.9.7 updater acceptance passed on Testbox. The separately matched native published-driver CI timeout also fails on clean main; the PR records that inherited failure and its limits.
2026-10-01 18:24:12 +00:00
Peter Steinberger
335d435bde
fix(update): keep progress writes from invalidating snapshots (#162604)
* fix(update): retain progress writer through candidate snapshots

* fix(doctor): refresh source state after the open-file probe

* test(tui): supply the retained writer admission fixture

* fix(update): preserve canary receipt failures through cleanup

Keep initial admission and completion receipt refusals distinct from candidate check failures while preserving owned teardown.
2026-10-01 18:15:42 +00:00
Ayaan Zaidi
98e86147b1
fix(agents): stop repeating a progress message when the turn ends right after it (#162367)
Related: #110565, #115104, #119605

## What Problem This Solves

Fixes: a chat gets two nearly identical messages when a turn sends a progress update with `message(final=false)` and then ends without a final answer.

In our setup a group turn sent "started the run, I'll report back" as progress and stopped empty. Settled-turn finalization then made an extra model call, which restated the same progress, and both messages landed 12 s apart.

## User Impact

If the last thing a turn does is post progress to the current conversation and then it stops empty or with `NO_REPLY`, that progress message is the reply. No second model call, no duplicate. Turns that did more work after or alongside the progress, or that sent nothing, still get finalization.

## Why This Change Was Made

Provenance:

- Settled-turn finalization (#110565, closes #108738) exists so users are never left with no reply after tools ran but the model said nothing. Fallback text after an empty finalizer came later in #133520; #138645 keeps that fallback private in message-tool-only groups.
- `final:false` vs `final:true` (#105365, then the shared contract in #119605) keeps progress from ending a run early and keeps a later failure visible after progress.
- #115104 (fixing #111764) made marked progress *permit* finalization. #111764's caveat was progress sent early in a long tool run, followed by an empty final completion, which was swallowed silently. #127070 applied the same rule to queued follow-ups.

Root cause: none of these checks look at order. A progress message sent as the turn's **last** tool batch was written after the model had seen every other tool result. When the model then stops empty, the finalizer gets the same transcript and no new information, so all it can do is restate the progress. That is the duplicate. The #111764 case is different: there, tool work came after the progress, and that case is unchanged.

Fix: the decision happens when the runtime settles the attempt (`completeEmbeddedAttemptResult`), not per assistant message. The embedded subscriber only records, per provider turn (reset on `turn_start`), whether every tool that finished in that turn was a complete, non-error, non-partial `final:false` text send to the current source, and it keeps that fact for the latest turn that ran tools. At settlement, if that holds, the run ended normally (`terminal.kind === "ok"`), and the terminal assistant response stopped with empty text or `NO_REPLY`, the attempt result marks that latest progress send final and sets `sourceReplyDeliveryState: "delivered"`. Every downstream consumer (finalization gate, terminal resolution, follow-up delivery, run-entry terminal reply) then sees the same completed-reply fact an explicit `final:true` produces. No consumer special case, no prompt or tool-text change. Native-async fragments do not split a turn: an async send followed by a later tool turn, or async work and progress in the same response, still finalizes. Progress reactions, partial or errored sends, closing text, and work in the same batch keep today's behavior.

Options considered:
- Give the finalizer the sent text and accept `NO_REPLY`: this still costs a model call, depends on the model's judgment, and needs a new "empty finalizer is fine" path next to the #133520 fallback.
- Tighten the `final` tool description: models already end empty after progress, so this cannot guarantee anything.

Hermes (`agent/turn_empty_response.py`) handles this the same way. When an empty response follows visible content with only housekeeping tools after it, Hermes reuses that content as the final answer and marks it as already shown so it is not re-sent. It nudges once only when substantive tools ran after the content.

Scope: this changes only the embedded runner. Codex app-server turns can plausibly hit the same duplicate: Codex marks `final:false` in `dynamic-tools.ts` and can complete a turn with no agent message. But "nothing ran after the progress" there depends on item order across native Codex items (shell, apply_patch) that never pass through the OpenClaw dynamic-tool bridge. That makes the Codex event projector a separate owner, so it is a follow-up.

No overlap with Pash/Sarah changes: `a2bbcbf406` and #154217 touch unrelated hunks in these files, and intentional `NO_REPLY` after a delivered message stays honored.

## Bounded cost

- The new path only removes a model call: when the latest tool batch was only complete progress sends and the model then stops empty, settled finalization is skipped (0 extra calls instead of up to 2).
- Every other path is unchanged: settled finalization keeps its existing cap of 2 tool-free attempts (`MAX_EMPTY_SETTLED_FINALIZATION_ATTEMPTS`, covered by `settled-turn-finalization.test.ts`). The finalizer has no tools, so it cannot send progress and cannot re-enter this path.
- The settlement check runs once per attempt result and does not schedule retries, runs, or model calls.
- Tested: `attempt-result.test.ts` drives the real `runAgentLoop` with the real subscriber, then the real `completeEmbeddedAttemptResult` and `resolveSettledToolTerminalContinuationInstruction`. Progress-last followed by an empty terminal gives a delivered reply and no finalizer. Two cases still finalize: an async progress send, then an empty tail, then a read, then an empty terminal; and async read plus progress in one response, then an empty terminal. Each failed on the head before its fix with `expected 'delivered' to be 'missing'`. The QA scenario asserts 2 model requests (no finalization request) for progress-then-empty and 3 for the write-then-empty control.

## Evidence

QA lab, mock-openai, isolated gateway child, qa-channel group with `visibleReplies: message_tool`. New scenario `group-progress-then-empty-finalization`:

- Before (origin/main + scenario): fail. The progress room got `["Started the run, I will report back. QA-GROUP-PROGRESS-OK", "Still running, I will report back. QA-GROUP-PROGRESS-OK"]`.
- After: pass. The progress room got 1 post and made 2 model requests (no finalization request). The control room (write, then empty stop, nothing sent) still ran finalization, made 3 model requests, and posted `QA-GROUP-WRITE-FINALIZED-OK` once.

Settlement regressions in `attempt-result.test.ts` (real agent loop): on the previous per-message head, the async-continuation case fails with `expected 'delivered' to be 'missing'`. Subscriber batch-fact table: only-progress is true; same-batch work, a later batch, a reaction, and a partial/errored send are false. Focused suites pass: attempt-result (43), suppression (32), tools handler, settled-tool evidence, settled-turn finalization, attempt-execution-phase, and lifecycle.

Real gateway (qa-channel, round 3 head): the scenario passes with 1 progress post and 2 requests, and the write control finalizes with 3 requests. The scenario is the only CI-runnable check of the consumer chain after the attempt result: run-loop finalization admission, terminal resolution, payload building, and channel delivery. The loop tests stop at the delivery fact, so a consumer that ignored it would still pass them, but here it would post a second message.

### Live model, Telegram Test Server group

Setup: a real user in the QA group, `messages.groupChat.visibleReplies: message_tool`, live OpenAI `gpt-5.5` through a pass-through proxy, isolated gateway. The prompt asked the model to post a `final:false` progress update saying the report run started and then stop. In both runs below the model did that unsteered: one successful `message(final:false)`, then an empty stop. In each run the bot's transient status draft (`Working`) appears and is deleted after delivery; that is existing progress-draft behavior and the same before and after.

| Run | Bot posts that remain | Model requests (main agent) | Finalizer |
|---|---|---|---|
| Before, origin/main `7dd6ab7` | 2: `Monthly report run has started — I’ll report back here when it finishes.` (t=13.1 s), then `Done.` (t=17.7 s) | 3 | ran (`settled post-tool turn lacked a final answer`) |
| After, `e27810b` | 1: `Monthly report run started. I’ll report back here when it finishes.` (t=13.8 s) | 2 | not run |

Control on `e27810b` (finalizer must still run after real work): the model posted `Checking the clock now.` with `final:false`, then ran `exec date -u`. The proxy replaced only the model's next response with an empty turn, which is the incident's "ends empty" shape; the finalizer request itself went to the live model. The finalizer ran and the group received `The UTC time is Thu Oct 1 13:08:11 UTC 2026.`

An earlier variant that told the model to end with `NO_REPLY` produced one post on both refs: on main the live finalizer also answered `NO_REPLY` twice, and the private fallback stayed private per #138645. Main still spent 2 extra model calls there; the fix spends none.

LOC vs merge base: production +47/-3 (`src/agents` subscriber + attempt-result), docs +5/-1, tests and QA fixtures +441.

Test cost (`pnpm test <file> --maxWorkers=1`, wall, head e27810b, shared loaded host): `attempt-result.test.ts` 46.4 s for 44 tests, with the 3 new real-loop cases at 0.2 to 0.6 s each; `embedded-agent-subscribe.subscribe-embedded-agent-session.source-reply-suppression.test.ts` 53.4 s for 32 tests, with the 5 batch-fact cases at 4 to 160 ms each. Most of the wall time is import and transform. The new cases use no timers, sleeps, or gateway boots. No attributable CI timing exists yet for this head.

CI on 959c8b0 (run 36873121248, after one rerun of failed jobs):
- `gateway-timeout-recovery-subagent.e2e.test.ts`: fails 2 of 2 locally on clean origin/main. Known flake #161139, with fixture fix #162260 open. The scenario never uses the message tool.
- `src/gateway/server.chat-steer-direct-command.test.ts` (skills watchers left open): fails the same way on four unrelated PRs today: `steipete/session-store-agent-db-drain`, `chore/chrome-devtools-mcp-1.10.1`, `chore/npm-12-1-20261001`, `codex/repair-child-completion`.
- `node-worker-launch-wire.e2e.test.ts` and `telegram-model-picker-prepared-gateway.e2e.test.ts`: red on the first attempt and green on rerun; both pass locally on head and on main.
- `src/auto-reply/reply/dispatch-from-config.secrets.test.ts` "quiet group link": still red. Only the file's first case fails, at its default 1 s `vi.waitFor`; that case took 1437 ms in CI while the following cases took about 60 ms. The same three-file shard passes locally on the head, on head merged with current main, with `CI=true`, and twice on main under `taskpolicy -b`. The test drives the `secrets` tool, which never touches the message-tool progress path, the new `turn_start` reset, or attempt-result settlement. I could not reproduce it on main or find it on another PR's CI from the last two days (93 failed node jobs scanned), so this one is not proven unrelated.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-10-02 02:12:10 +08:00
Vincent Koc
c0c8fc9a39
feat(tooling): collect local compiler performance evidence (#156836)
* feat(tooling): add opt-in compiler performance evidence

* test(tooling): narrow compiler metrics artifact paths

* style(tooling): use braced compiler metrics guards

* chore(tooling): integrate main compiler ownership

* style(ci): format tsgo compiler version lookup

* feat(ci): integrate compiler metrics with current runner

Preserve the current prepared compiler owner, Kysely preparation, static diagnostics, signal callbacks and output joins while retaining opt-in metrics. Cover combined CI diagnostics and metrics for successful and diagnostic compiler exits.

The previous focused tests and profiles remain bound to the older composition. This integration is source-reviewed; exact-head hosted qualification is pending.

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-10-02 01:08:24 +07:00
Peter Steinberger
84c6eb8fc5
improve: sessions_send delivers a peer's reply once instead of running automatic agent ping-pong (#162227)
* refactor(agents): replace peer announcement loops with one reply delivery

Deliver delayed peer results once to the requester and return waited replies inline without a continuation. Preserve child completion custody and same-session generation-bound channel delivery. Use exact-key saved-route lookup, retire legacy peer control tokens from runtime decisions, and retain historical display suppression.

Simplify the announce owner, preserve required missing child output, and keep wake retries within the remaining operation budget. Update session guidance and the opt-in live peer scenario.

Candidate regression scenarios pass. Broad focused tests and changed-lane validation remain incomplete under host load; original-code regression proof is still pending. No live provider test was run.

* refactor(subagents): consolidate completion custody and lifecycle owners

Share terminal-effect persistence and requester-wake settlement adapters while preserving their distinct authority, frozen-wave, and publication guards. Remove unused preparation and accessor paths and reuse existing parameter contracts.

Propagate unknown SQLite write outcomes before scheduling interrupted completion recovery during restart drain. The focused candidate regression passes; original-code failure proof and the broader validation matrix remain pending.

* test(agents): align completion assertions with reply ownership

Assert that an ACP target already owned by its requester keeps task completion and starts no detached reply. Keep isolated cron fallback routing and participant-store checks while asserting no detached wait.

Remove assertions on the retired duplicate steerMessage field; retain complete source-reply content, ordering, exclusions, visibility and batch settlement checks. Remove the redundant inner same-session return while retaining the enclosing branch exit.

Validation: scoped Codex review at P2, formatting and diff checks passed. Two attempts on the coordinator Testbox lease failed during file sync before installation or tests. Baseline regression proof, the full matrix and the changed-lane gate remain pending.

* test(agents): inject the gateway caller into same-session delivery routing tests

* test(subagents): route direct requester-yield callers through prepared cron authority

* refactor(subagents): keep the requester wake currency assertion module-private

* fix(agents): fence failed and delayed subagent cleanup

Restore the previous resident row when same-ID registration persistence fails.
Carry collector cleanup authority into attachment removal, and bind lifecycle
grace timers to their captured registration and generation.

Regression coverage exercises failed writes, retirement during cleanup, and
same-ID successors without changing storage or timer durations.

* fix(agents): retain paused state in compact registry reads

Read pauseReason from existing stored JSON for ordinary and private-parent
records so cold projections preserve the resident paused status. Derive compact
SQL row types from the query and shrink the removed-assertion allowance.

No schema, serialized representation, or migration changes are required.

* fix(agents): preserve spawn cleanup and request scope

Route fork lookup failures through provisional-child cleanup. Reuse the spawn
admission owner for visible work with the resolved requester agent identity.
Pass trusted-creation transport timeouts through the argument the transport reads.

Consolidate in-process dispatch while preserving creation custody, source
fencing, per-interface deadlines, and signed fallback behavior.

* refactor(agents): simplify subagent registry and spawn owners

Remove unused registry APIs, callback declarations, presentation adapters,
registration inputs, and bootstrap scaffolding. Share registration ACK control
flow while retaining both durable writes and the original settlement error owner.

Complete keyed capability lookups, reuse heartbeat and timeout owners, carry
resolved model references, and remove obsolete ACP transcript work and relay
options. Preserve legacy stored attachment safety and protocol-v4 fallbacks.

* refactor(agents): simplify completion consumers and registry publication

Move full-registry fixture writes into test support and require named production mutations. Consolidate publication subscriptions while preserving projection ordering and persistence observer exception isolation. Remove duplicate completion retry scaffolding, runtime barrels, test-only exports, duplicate types, and single-caller adapters.

* chore(agents): shrink retired test API assertion allowance

* test(gateway): preserve real registry operations in agent fixtures

* test(subagents): preserve named publication fixture preconditions

Name all retired fixture rows when replacing describe/list facts. Warm the compact cache before its initial publication so the existing ordering assertion still detects an unintended SQL reload.

* fix(plugin-sdk): preserve capability store compatibility

Retain the public record-or-lookup store contract, including normalized first-match identity lookup and canonical depth fallback for partial records. Keep migrated internal callers on keyed lookups. Restore the base handoff declarations and requester query/caller linkage so the existing SDK declaration gate stays unchanged.

* test(gateway): publish fixture memory ownership

Publish the staged in-memory owner through commitOwnership before asserting describe and list projections. Named durable writes no longer trigger the whole-index rebuild that accidentally discovered the fixture row. Preserve the existing assertions.

* refactor(agents): tighten completion projections and ownership facts

Project only formatter-consumed child fields without mutating authoritative rows, and carry the session resolver's required requesterOwned boolean directly. Preserve formatting, nested identity references, authority checks, and public signatures while resolving the full core lint findings.

* test(agents): retire obsolete benchmark registry mock

Remove the full-registry writer no-op from the memory benchmark's typed module mock after the production writer cutover. Preserve the named-write mock and benchmark behavior.

* test(agents): await requester settlement in live fixtures

Child cleanup can publish before the requester delivery acknowledgement. Join persistence publications until the requester wake settles, preserving the existing delivery assertions and deadlines.

* refactor(agents): retire hidden sessions_send message aliases

Require the canonical message argument already declared by the public tool schema. Remove alias-only reasoning stripping and keep canonical message indentation unchanged.

* fix(agents): decode compact subagent metadata canonically

Replace the separate compact decoder with the canonical codec and session-list projector. Keep bounded SQL metadata projection, private envelopes, and duplicate-key last-value semantics without transferring retained results. Add real-reader duplicate-key and retained-content regression coverage. Schema, persisted bytes, and update behavior are unchanged.

* refactor(agents): remove redundant delivery comments

* perf(agents): restore compact reader after benchmark regression

Revert 2c0fd630c7 as required by the performance gate. Across five scoped 2000-row reads per revision, the base median was 227.31 ms and the candidate median was 429.25 ms (88.84% slower). Retain the sessions_send alias retirement and comment cleanup; leave the duplicate-key reader fix for separately approved work.

* test(agents): relocate sessions_send alias regression coverage

Keep all four rejection cases in the assembled-tool suite where retired preparation coverage freed space. The oversized sessions test file shrinks again; no assertions, mocks, timeouts, skips, or line-cap baselines change.

* fix(agents): retain sessions_send record input validation

Use the shared isRecord guard from the retired alias normalizer instead of a new type assertion. Preserve canonical missing-message errors for non-record input and satisfy the assertion safety gate without a suppression.

* refactor(agents): keep sessions_send replies on the requester key

Remove the legacy key-only DM-to-main reply remap and its routing exceptions. Use the resolved caller key for reply context, provenance, watches and followup preparation. Preserve exact-incarnation and authorization checks.

* test(agents): bind native followup custody in admission fixture

Keep the key-only DM requester authority test on the native child completion path. Settle the existing completion owner before mock acceptance and reset queued mock implementations between cases.

* fix(agents): reconcile subagent callers after main merge

Use the phase-aware publication API in newer main callers and explicit changed IDs in the prepared-read fixture. Keep retained regressions in their current sibling owners and place requester adoption in the lifecycle controller without weakening guards. Preserve heartbeat narrowing after shared-owner integration.

* test(agents): clean up merged fixture imports and names

* test(agents): refresh single-reply prompt snapshots

Regenerate prompt fixtures for the intentional sessions_send description change. The four-file delta contains only the description and derived sizes and hashes. Reproduced the hosted drift on Testbox, then passed prompt:snapshots:check and the snapshot-only changed checks.

* test(agents): await owned descendant settlement

* fix(agents): retain the leaf followup owner contract

Restore the standalone completion-owner interface and implements check from main. Deriving that contract from the implementation class creates a type-only cycle through the cohort projection. Runtime behavior and the exposed owner methods remain unchanged.

* fix(agents): disambiguate the sessions reply target resolver

Give the async sessions_send reply lookup a distinct exported name from the general synchronous outbound session resolver. Update its sole production caller and direct tests without an alias or behavior change.

* test(agents): remove the retired steer fixture argument

* test(agents): drop the retired sessions_send delivery mode from the follow-up yield live test
2026-10-01 17:28:00 +00:00
Peter Steinberger
778e1f37ae
fix(ci): fetch boundary bases for shallow PR checkouts
Reuse the extension-lint comparison-base action for PR boundary checks. Keep its bounded exact-SHA fetch and existing fallback policy, and leave checkout refs unchanged.

A depth-one merge stays an ancestry boundary even after its parent is fetched. Validate the pinned raw first parent, then compare the two trees; retain ordinary ancestry validation for other shapes. Shallow checkout regressions cover package, core and deleted public SDK changes, including blobless base inventory hydration.
2026-10-01 10:26:23 -07:00
Peter Steinberger
b0324be63c
test(ui): record renderer stall evidence when Control UI e2e waits fail (#162767)
* test(ui): record renderer stall evidence when Control UI e2e waits fail

Two scheduled main runs failed intermittently in comment flows without
evidence of what the renderer was doing: a pin click hung for 30 s with the
failure read missing its deadline, and a comment delete never showed up
during a 15 s poll. Neither reproduced locally.

Mocked-Gateway pages now arm a stall probe before navigation. A renderer
that misses the failure-read deadline reports main-thread busy time by kind
and its paused JavaScript stack, then resumes. A stall that ended appears as
script-attributed long animation frames. Public output keeps only bundle
paths, positions, function names, and listener tag and event names.

* test(ui): arm the stall probe before tests open CDP sessions

PR CI showed 11 mobile safe-area geometry failures: installMockGateway
attached the probe after those tests had set
Emulation.setSafeAreaInsetsOverride on their own CDP session, and Chromium
drops an earlier session's override at the next navigation once a later
session attaches. A session attached first leaves later overrides intact.

The shared suite's withPage now arms the probe right after newPage(),
before test code runs, and installation is best effort for test doubles,
closed pages, and non-Chromium contexts.
2026-10-01 17:17:23 +00:00
Peter Steinberger
4e6eb74beb
fix(ci): avoid overlay copies in published-driver updates (#162858)
The published-driver cell timed out on scheduled main runs: the Docker fixture forced tens of thousands of OverlayFS copies on the candidate tree during the managed update. The cell now stages the update on a fresh private native volume per invocation (cleaned up on exit), which brings retention from 39 s to 12 s and the whole cell to about four minutes; checksum failures print both digests and reject before the update starts. The earlier "digest-mismatch: error" log line was an action setting echo, not a mismatch.
2026-10-01 10:14:24 -07:00
Ayaan Zaidi
2f4d4751e5
fix(cron): delegated blocked runs report ok and silent blocked jobs get auto-disabled (#162686)
Related: #162391

## What Problem This Solves

Fixes two follow-ups to #162391:
- A scheduled run that hands its work to a subagent can deliver the child's `AUTOMATION_FAILED` reply verbatim and still record the run as `ok`. This happens on announce jobs, and on `delivery: none` jobs since #162474 started recording the child's answer.
- A silent (`delivery: none`) job whose agent keeps reporting a blocked task gets auto-disabled after 10 runs, and that posts an auto-disable notice even though the job has nowhere to notify.

## User Impact

- When a delegated run's settled child answer starts with `AUTOMATION_FAILED`, the run is recorded as an error with the child's explanation. Announce and current-session delivery send only the explanation, never the token. `delivery: none` runs record it without sending anything, and keep the run transcript as other failed executions do.
- A `delivery: none` job with no failure-alert route still records agent-reported failures in run history, job status, and error backoff (escalating to hourly, like any failing job). Those failures never auto-disable the job or post the auto-disable notice, so the job stays silent.
- Unchanged:
  - Runtime errors on silent jobs still auto-disable at 10 as before.
  - Announce and webhook jobs, and jobs with a configured failure route, still auto-disable on agent-reported failures.
  - `NO_REPLY` behavior.

## Why This Change Was Made

- **Settled child answers.** #162391 classified the token in `resolveCronPayloadOutcome`, which runs on the parent's reply before #117308's descendant settlement. Two places in `dispatchCronDelivery` then adopt the settled child's reply as the run's answer: the announce settlement (#117308) and the no-delivery settlement (#162474). Neither classified it. Both now go through one adoption helper that calls the shared parser `readAutomationFailedReport`, so the run's effective terminal answer is classified after settlement. `buildDeliveryState` derives one execution-failed flag from that result and uses it for both the completion status and the transcript-cleanup decision, so a failed no-delivery run is no longer cleaned up as a quiet success. `run-finalize` consumes the result the same way it already handles `pendingPresentationWarningError`: error status, permanent classification, explanation as the summary.
- **Auto-disable.** Maintainer decision: an agent-reported blocked outcome on a job with no notification owner must not lead to auto-disable. The permanent error classification now carries `reportedByAgent`. `applyJobResult` in `timer-outcomes.ts` still counts every error in `consecutiveErrors`, so status and error backoff are unchanged. It skips only the auto-disable decision, and only when all of these hold: the failure was agent-reported, `resolveFailureAlert` finds no route, and the delivery mode is `none` (webhook jobs still auto-disable).
- Hermes handles the same protocol the same way (`[CRON_FAILURE]`): its cron hint asks the agent to report a delegated child's failure on the first line. No new tool is needed.

## Bounded cost

- No new turns, runs, waits, or notifications. Parsing happens once per settled answer inside the existing finalization.
- Agent-reported failures stay permanent, so the scheduler never retries them early.
- A silent blocked job keeps the normal error backoff (30 s, 1 min, 5 min, 15 min, then hourly), so it can never run more often than an identical job on origin/main. The only difference is that it keeps running at that backed-off cadence after the 10th failure instead of being disabled.

## Evidence

### Live model, Telegram Test Server (`telegram-e2e-userbot`, DM)

Setup: the repo runner unchanged (`run-mock-sut-user-e2e.mjs --backend mock --source-gateway --dm`, Convex lease). `E2E_MOCK_SERVER_PATH` points to a throwaway proxy that stands in for the mock provider and forwards every request unchanged to the real OpenAI API. The proxy reads the key from a private file, so the Gateway only ever holds the runner's dummy key. `E2E_ROOT_CONFIG_PATCH` selects `openai/gpt-6-astra` (`tools.toolSearch: false`). A scenario `command` step creates the jobs with `openclaw automations add` and runs them. No model output was steered: the live parent chose `sessions_spawn`, and the live child and the live silent job wrote their own `AUTOMATION_FAILED` reports because they had no shell tool. The only steer is timing in case (b): before each run, `cron.update` sets `nextRunAtMs` to now and moves `lastRunAtMs` back 2 h, so the error backoff (up to 1 h) counts as served. The Gateway's own scheduler then runs the job as a normal scheduled (not forced) run.

Before = origin/main `d8a110b1a1`, plus this PR's exact base `5c4e9aba92` for case (a). After = this head `8b77715ba4f`, with all three cases in one Gateway run.

**(a) Announce job, live parent delegates to one child, the child reports `AUTOMATION_FAILED`**
- before (`d8a110b1a15`): `cron run --wait` took 99.4 s. Run `ok`, `completionStatus: succeeded`, `deliveryStatus: delivered`. The DM received bot message 43601: `AUTOMATION_FAILED` / `The command could not run because no shell execution tool is available in this session.` The token is visible.
- before (`5c4e9aba92`, the PR base): 80.7 s. Same result: run `ok`, `delivered`, and bot message 43705 reads `AUTOMATION_FAILED` / `The command could not run because no shell execution tool is available in this session.`
- after: 63.5 s. Run `error`, `completionStatus: failed`, `deliveryStatus: delivered`. Error and summary are both the explanation. The DM received bot message 43712 with only `The command could not run because no shell execution tool is available in this session.` No token.
- Pre-existing flake, not in this diff: on both sides, the announce settlement sometimes ends with `cron child-session handoff completed without a final assistant payload` even though the live child answered. That happened on 1 of 3 origin/main-side attempts (a second `d8a110b1a1` run) and on 2 of 4 attempts on this branch (both on `bfe1f0f2169`). In those cases nothing reaches the chat. The PR's change only runs after a reply has been read, so it is not involved. The no-delivery path (case c) settled on every attempt.

**(c) Same delegated job with `delivery: none` (the #162474 path)**
- before: 55.2 s. Run `ok`, `succeeded`, `deliveryStatus: not-requested`, summary `AUTOMATION_FAILED ⏎ The command could not run because no shell execution tool is available in this session.` The token is recorded as an ok answer.
- after: 36.3 s. Run `error`, `failed`, `not-requested`. Error and summary are both `The command could not run because no shell execution tool is available in this session.` The job stays enabled. No chat message.

**(b) `delivery: none` job, no failure-alert route, the live model reports `AUTOMATION_FAILED` on every run (11 scheduled runs)**
- before: runs 1–10 were all `error` (`Could not run \`node scripts/ledger-sync.mjs\`: this session only provides a file-reading tool…`), with `consecutiveErrors` 1…10. After run 10 the job was `enabled: false` with `autoDisabled: {reason: "consecutive-failures", consecutiveErrors: 10}`, so run 11 never happened. The DM received bot message 43644 (heartbeat relay): `⚠️ Your "Nightly ledger sync" automation was automatically disabled after 10 consecutive failures. … openclaw automations enable <id>`. 10 runs took 504 s.
- after: all 11 runs were `error` with the same explanation, and `consecutiveErrors` went 1…11, which keeps feeding backoff. After run 11 the job was still `enabled: true` with no `autoDisabled`. History shows 11 `error` entries. Across the whole run, the DM received exactly one bot message: 43712 from case (a). Cases (c) and (b) sent none, including during a 120 s wait after the loop. 11 runs took 526 s.

Provider evidence (proxy log): before had 1 parent `sessions_spawn` turn, 1 child turn, 10 silent-job turns, and 1 relay turn carrying the auto-disable notice. After had 2 parent `sessions_spawn` turns, 2 child turns, 11 silent-job turns, no auto-disable relay turn, and only one routine heartbeat poll, which answered `NO_REPLY`.

Environment note: in `--source-gateway` mode on both refs, Telegram inbound polling could not start (the ingress worker's plugin capture has no `dist/plugin-sdk`). So the opening DM turn was not processed, and the jobs are owned by `agent:main:main`, the DM's owner session under the default `dmScope`. Outbound delivery and relays reached the DM normally. This is the same for before and after.

### Tests

Each new test fails on the code it fixes and passes on this head:
- `delivery-dispatch.double-announce.test.ts`, real `dispatchCronDelivery`:
  - Announce settlement: a settled child reply `AUTOMATION_FAILED ⏎ …` is delivered as the explanation only and recorded as `agentReportedFailure`.
  - No-delivery settlement: uses the real `waitForDescendantSubagentResult`, with the registry mocked only at its edge. The child's report is recorded as `agentReportedFailure` with the explanation as output and summary. Nothing is sent, and the `deleteAfterRun` transcript is not deleted. Classification fails without the shared adoption helper (checked on the rebased predecessor commit), and retention fails on `bfe1f0f2169`.
- `timer-outcomes.regression.test.ts`, real `applyJobResult`, 11 agent-reported failures from a clean streak:
  - `delivery: none`: still enabled, streak 11, no auto-disable notice, backoff delays 5 min, 15 min, 1 h, 1 h. This fails on `ee6d9e07fdf` (no streak, no backoff).
  - Webhook and announce jobs: auto-disabled at 10 with one notice, same backoff.

Measured on this head, run separately (vitest `Duration`):
- `delivery-dispatch.double-announce`: 84 tests, 22.1 s.
- `timer-outcomes.regression`: 20 tests, 15.5 s.
- `run.meta-error-status`: 14 tests, 20.0 s.

All pass. `pnpm tsgo:core` passes, and oxlint and oxfmt on the changed files are clean.

Docs: `delivery.md` documents the silent agent-failure exception to the 10-failure auto-disable. `payloads.md` notes that a delegated run's settled answer is classified the same way.

LOC vs origin/main: production +54/−25, tests +100/−1, docs +2/−2.

### Review

- Adversarial re-review round 1 (on `0530dc9f`): fixed the one real finding. Webhook jobs had been exempted too; the exemption now requires delivery mode `none`, with a webhook test case.
- Round 2 (on `ee6d9e07fdf`): NOT LAND, one finding. Skipping the increment also froze error backoff, so a silent blocked job could keep running every minute. Fixed per maintainer decision with no second counter: the streak always increments, and the exemption moved to the auto-disable decision.
- Round 3 (on `bfe1f0f2169`, rebased): LAND, no findings. The rebase onto origin/main also routes the #162474 no-delivery settlement through the same adoption helper.
- ClawSweeper on `bfe1f0f2169`: fixed both findings. (P2) A failed no-delivery child run deleted its transcript; it is now kept. (P3) The auto-disable exception is now documented. Also added per-file test timings.
- Round 4 (on `8b77715ba4f`): LAND, no findings.
- Accepted tradeoffs: the exemption is not re-derived during finalized-run startup recovery. A silent job whose streak is already 10 or more from reported failures auto-disables on its next runtime error.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-10-02 01:00:18 +08:00
Peter Steinberger
d188233947
fix(talk): preserve inherited realtime settings on upgrade (#162644)
* fix(talk): persist inherited realtime settings in Doctor

* fix(talk): normalize inherited realtime input without assertions

* fix(tooling): include Talk normalizer in trusted wrapper sources
2026-10-01 16:58:58 +00:00
Peter Steinberger
a4a78b738e
fix(update): package activation recovery rejects Bun-hosted installs (#162586)
Package recovery treated checked Bun executables as Node and dropped configured macOS SQLite selection. Carry admitted runtime facts and the shared environment choice into durable standalone recovery, preserving version-1 journals, backup custody, and separately owned service restart.

Verified focused suites on Node and Bun, real macOS custom-library recovery from a clean shell, published-driver updates and Node-free interrupted recovery on AWS, static/import checks, and independent review. CI fixture and cell-lifetime repairs retain the existing assertions and product timeout policies.
2026-10-01 09:54:24 -07:00
Peter Steinberger
c66d7fbc08
chore(storage): correct worker-only SQLite inventory classifications (#162826)
* chore(storage): classify traced worker-only SQLite kernels

* docs(storage): refresh inventory location after main merge
2026-10-01 09:41:47 -07:00
Mert Başar
2f1f8beed7
feat: optionally omit tools on conversational turns (#153340)
* feat(core): gate decision tool prefilter behind decisionAssistance labs and harness capability (#155314, #155316)

- Connect prompt-build prefilter to canonical isDecisionAssistanceEligible Labs consent contract (#155314)
- Gate prefilter evaluation on harnessSupportsTurnScopedToolRestrictions capability declaration (#155316)
- Preserve tool policy when harness is unsupported (e.g. Codex app-server), pending action instructions are present, or attempt is cancelled
- Update test suite with schema-backed decisionAssistance configurations and per-agent decisionModel overrides

Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>

* feat: align conversational tool filtering with Decision assistance

Preserve the contributor core consumer while binding canonical consent, actual harness support, prepared config, and live run authority. Validate the runtime deadline and real ONNX prompt submission without claiming universal classification or latency gains.

Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>

* fix: disclose automatic Decision filtering in configuration help

Resolve the remaining configuration-help review finding. Explain explicit consent, inherited Decision models, built-in filtering, preserved required tools, request transfer, and hosted costs. Regenerate the config documentation baseline without changing settings or runtime behavior.

Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>

* feat: filter conversational follow-ups with bounded Decision context

Use one provider-neutral batch with separate missing-context and next-response tool-need judgments, following the documented Jev/Noul semantics. Preserve first-turn eligibility, safe bounded history, permissions, required tools and per-turn restoration. Measure actual foreground tool definitions with DEBUG-only scalar diagnostics.

Live model-quality testing is deferred to local testing; no classifier-specific thresholds or adapter framework are introduced.

Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>

* fix: restore prompt-phase CI and simplify Decision assistance copy

Bring the shared prompt-phase fixture and Tool Search assembly override up to the current dispatch diagnostics contract. Type the assembly mock and complete policy result so missing fields fail checking instead of silently blocking the primary stream. Preserve the original four replay/queued-context assertions and cover DEBUG-on/off diagnostics without adding cases.

Construct modified Decision answers on fixture-owned copies rather than mutating the read-only API result. Keep the Labs toggle and configuration help generic; supported uses and full behavior/privacy details remain in canonical docs. No provider, saved-consent, permission, or classifier-policy change.

Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>

* fix: preserve prompt hook context in tool prefilter

Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>

* fix: preserve tool and provider availability through Decision filtering

Revalidate Decision eligibility at the foreground dispatch boundary and withdraw only its optional cap when the published opt-in or model selection changes. Restore the current permitted schemas, catalog, callability and owned prompt guidance without discarding independent hook restrictions or required tools.

Keep caller-budget expiry out of shared provider outage accounting while preserving genuine provider failures and physical request settlement. Replace incomplete context-engine hook fixtures with the typed real runner, retain the original assertions, and keep the context limit private for the production dead-code gate.

Preserve the hook-context and rubric-9 change; document and test the separate 6000/8000 UTF-16 evidence-text bounds and the existing forward-looking saved-intent contract. Add failure-first composed final-wire and real loopback HTTP/provider-host regressions, including subsequent explicit evaluation on the same provider.

Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>

* fix: restore foreground result-budget fixture typecheck

Model the result-budget fixture's foreground-only session explicitly as not compacting. Keep the required production compaction contract and all existing result/persistence assertions intact.

Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>

* test: consolidate Decision assistance boundary coverage

Remove overlapping test layers and repeated fixtures while retaining the distinct context, policy, and dispatch boundaries. Runtime validation remains pending because the secretless runner broker is unavailable.

Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>

* fix: stop automatic Decision calls after opt-out

Address MertBasar0’s pre-dispatch opt-out report by reusing the automatic consumer eligibility fence after awaited operator preparation. Preserve explicit Decision calls and root/scoped authority and cancellation checks. Extend the existing composed regression at the operator boundary.

Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>

* fix: stop Decision evidence after admission is revoked

Carry automatic consumer admission through the existing provider host and TypeSafe adapter to the final synchronous guarded-fetch boundary. Retain observed revocation or authority failure across adapter sanitization and asynchronous cleanup without treating either as a provider outage. Preserve explicit Decision evaluation, caller cancellation, deadlines, and provider retirement.

Rebase the contributor commits onto pinned main and retain their previously published integration fixes and test organization. Prove the cold-load and transport-preparation windows with the registered TypeSafe plugin and a real loopback request counter, including same-host explicit calls after revocations.

Co-authored-by: MertBasar0 <mertbasar0@hotmail.com>
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>

* fix(agents): make decision assistance opt-out admission-only

Keep Labs eligibility at provider dispatch while preserving independent live guards for admitted evaluations. Let in-flight results survive opt-out and cover subsequent disabled turns.

* test: close skills watchers in Discord capture fixture

---------

Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
Co-authored-by: MertBasar0 <mertbasar0@hotmail.com>
2026-10-01 09:33:15 -07:00
Ayaan Zaidi
08498ec40d
fix(cron): stale automatic tool lists block scheduled jobs from tools their owner has (#162432)
Related: #130753, #137832, #147969

## What Problem This Solves

Fixes: some scheduled jobs created by an agent fail for months because a tool list saved by an older OpenClaw build is missing tools the creator actually had, such as the native shell. In our setup, a monthly group job that runs `node <script>` delivered nothing in August, delivered nothing in September (the run still reported `ok`), and posted a blocker in October. Its saved list had 31 tools and no `exec`.

## User Impact

User impact: an agent-created agent-turn job that does not name specific tools now gets the same tools as its owner conversation at run time, like a job an operator creates without `--tools`. Existing jobs with an automatically saved creator snapshot behave the same way from their next run. Nothing stored is rewritten: no migration and no backups. Explicit tool lists, script payloads, condition triggers, and jobs bound to captured Codex app authority keep their stored list.

Tradeoff, approved by the maintainer (Ayaan): a per-sender tool policy on the creating owner, or a plugin hook that narrowed the creating turn, no longer limits these default jobs. Only owners can create automations from chat, and subagents cannot create them.

## Why This Change Was Made

**History of the saved list.**
- #91499 introduced it so a delayed run cannot do more than its creator could.
- #112483 made every agent-created job store one, because runs have no sender.
- #112661 made scheduled runs re-apply the owner session's group policy and every non-sender limit, keeping the stored list as the upper bound.
- #137832 fixed native tool capture for new jobs only, and deliberately did not widen stored lists.
- #147969 added a Doctor advisory. It only fires for claude-cli, so it never covered Codex-harness or built-in OpenAI jobs like ours.

**Root cause.** When no tool list was given, OpenClaw saved a frozen copy of the creating turn's tools instead of treating the job like an operator `*` job. Every capture bug (missing native tools, late configured MCP, renamed tools) then stayed in the job permanently.

**Fix.** This follows Hermes, which keeps no creator snapshot: `cron/scheduler.py` `_resolve_cron_enabled_toolsets` reads toolsets from config at run time.
- **New jobs.** An agent-turn create or update with no list, or `*`, stores `["*"]`. That is the same value operator jobs store, so the job's tools match a normal turn in its owner conversation. Script payloads and condition triggers still store the creator's concrete tools, because a script reaches MCP only through servers its list names. Jobs whose creator captured Codex app authority also keep the concrete list, because that authority is bound to it.
- **Existing jobs.** One helper, `resolveCronRunToolsAllow` in `src/cron/tools-allow.ts`: a stored automatic snapshot (`toolsAllowIsDefault`) runs as `*` when it has a valid scheduled owner policy, no condition trigger, and no Codex app authority. Otherwise it keeps its stored list. Every execution consumer of the stored list uses it: the run payload, the command-prompt preflight, and the scheduled message authority.
- **Script transitions.** A `*` job that becomes a script, or gains a condition trigger, captures the creator's concrete tools.
- **Exec pin.** A `*` list keeps the creator's exec host pin.
- **No new noise:** automatic snapshots stay excluded from the `web_search` provider warning, as on main.
- **Deleted, now pointless:** both Doctor advisories about incomplete automatic snapshots, the run warning about pre-MCP snapshots, and two exports nothing uses anymore.

Review note: on claude-cli, a `*` job runs without a CLI tool cap, so Claude's native tools behave exactly as in a normal chat turn in that conversation. This PR introduces no new path around `tools.deny` that a chat turn doesn't already have.

## Evidence

Live-model Telegram proof (Telegram Test Server DM, leased team credential, live `openai/gpt-6-astra` reached through a forwarding proxy that stands in for the runner's mock provider; the runner harness itself is unchanged). This reproduces the shape of the original incident:
- The tester DMs the bot, which creates the owner conversation.
- A job owned by that conversation is added. Its stored list is an old-style automatic snapshot `["automations","message","read"]` plus `toolsAllowIsDefault: true`, with no `exec`.
- The payload is `Run: node scripts/split-report.mjs and post its output line verbatim`. The workspace script prints a random nonce.
- The job is run once (`cron run --wait`), with announce delivery to the DM.

| Build | `exec` offered | Model action | What arrived in the DM | Run |
|---|---|---|---|---|
| base 94f5a8d (main before this PR) | no | `tool_search` ×2, then gave up | "Could not run node scripts/split-report.mjs: no command-execution tool is available…" | error |
| **this PR, head 5b78cb7** | **yes** | `exec {"command":"node scripts/split-report.mjs"}` | "**SPLIT-REPORT 93C53909**: general 41, design 17, ops 9" (the exact script output, with this run's random nonce) | ok, delivered |
| head 5b78cb7 with `tools.deny: ["exec"]` | no | `tool_search`, `read`, then gave up | "Could not run node scripts/split-report.mjs: no command-execution tool or paired node is available…" | error |

In every run, the stored job kept `["automations","message","read"]` plus the marker. Before and after use the same scenario and driver; only the checkout differs.

Update and live proof: published `openclaw@2026.9.7`, then this branch at the exact head (2f5099d), on the same state directory. Mock provider. Every process ran under a temporary `HOME` and state directory. Each job's message makes the model call `exec` with `touch <effects>/<job>`.

1. 2026.9.7 created both jobs through `cron.add` (scheduled policy `trusted`). With the Gateway stopped, the "stale" job was given the old automatic-snapshot shape `["automations","message","read"]` plus `toolsAllowIsDefault: true`. sha256 of both stored rows: `609358837…`.
2. Runs:

| Build / config | Job | `exec` offered | Side effect | Run |
|---|---|---|---|---|
| 2026.9.7 | stale automatic snapshot | no | absent | error |
| 2026.9.7 | explicit `["read","message"]` | no | absent | error |
| this branch | stale automatic snapshot | **yes** | **created** | ok |
| this branch | explicit `["read","message"]` | no | absent | error |
| this branch, owner policy narrowed to `tools.deny: ["exec"]` | stale automatic snapshot | no | **absent** | error |
| this branch, `tools.deny: ["exec"]` | explicit `["read","message"]` | no | absent | error |

3. After the branch runs, the stored rows were byte-identical (same sha256 `609358837…`), and job ids and lists were unchanged. Nothing was migrated.

An earlier run at e151ea3, with the same harness, also covered a snapshot bound to Codex app authority: `exec` was not offered, the file stayed absent, and the stored row was unchanged.

Tests:
- `run.tools-allow.test.ts`: a stored automatic snapshot `["message","read"]` reaches the embedded run as `["*"]`, with the owner's scheduled policy intact. It fails on main with `["message","read"]`.
- `cron-tool-creator-cap.test.ts`: a default agent turn stores `["*"]`, while a trigger script and a Codex-app creator keep the concrete snapshot.
- `run.tools-allow.test.ts`: snapshots without a valid owner policy, or behind a condition trigger, keep their list. Both cases fail on the previous head.
- `run.tools-allow.test.ts`: no `web_search` warning for an automatic snapshot that kept its list. This fails without the exclusion.
- `run.message-tool-policy.test.ts`: a self-edited automatic snapshot runs on CLI with no cap.
- `run.tools-allow.test.ts`: a legacy `Command to run:` prompt from an automatic snapshot without shell tools now runs instead of being rejected.
- `jobs-tool-policy.test.ts`: scheduled message authority is admitted for an automatic snapshot that lacked `message`.
- `cron-tool-creator-cap.test.ts`: a `*` agent turn converted to a script captures the creator's concrete tools.
- These three regressions fail on the previous head. `node scripts/check-changed.mjs` passes.
- Explicit-list, exec-pin and gateway creator-transport suites pass. `pnpm tsgo:core` passes.

## Bounded cost

No new path triggers a model call or a job run. The change only selects which tool list an already scheduled run uses.

LOC vs main: production +98/-250 (net -152), tests +140/-373, docs +19/-11.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-10-02 00:31:29 +08:00
Peter Steinberger
f833bb4256
fix(update): await progress receipts before advancing commands (#162098)
* fix(update): await progress receipts before advancing commands

* fix(update): preserve Doctor refusal and progress worker boundaries

Keep the execution guard contract in the existing parameter types module
without an implementation import cycle. Join actual reporting failures with
settled Doctor errors while preserving later typed authority refusals.

Run the fresh-package staging regression in the worker-capable harness now
that its real progress callbacks use the shared-state writer. Preserve both
replay cases and their file and backup assertions.

* test(update): await candidate progress receipt settlement

* fix(update): retain recovery after progress receipt refusal
2026-10-01 09:26:08 -07:00
Peter Steinberger
44aaaa7c57
refactor(channels): move repeated contract boilerplate onto shared helpers (#162827)
Seven channel plugins hand-wrote the same secret-contract setup, and
Zalouser duplicated migration and prefix parsing that shared helpers
already own. A new Plugin SDK helper, createChannelSecretContract on
openclaw/plugin-sdk/channel-secret-basic-runtime, builds the contract
from each channel's declaration, and Zalouser uses the existing helpers.
Feishu and Matrix stay out because other work is changing them.

The public SDK surface budget grows by exactly this one export (public
exports and callable exports each +1), approved by Peter on 2026-10-01.
clickclack, googlechat, irc, slack, sms and zalo now require OpenClaw
2026.9.8 because they call it; published plugins already require a
matching host, and installs fall back to the newest compatible version.

Registration and secret-contract output are identical to main across
eight channels; 6,619 channel tests pass. Production -110 net.

Release note: The clickclack, Google Chat, IRC, Slack, SMS and Zalo
plugins require OpenClaw 2026.9.8 or newer.
2026-10-01 09:24:55 -07:00
Peter Steinberger
6dbca39af1
fix(text): keep grapheme lookups on cluster boundaries under Bun (#162749)
Fix Bun message chunking and terminal truncation selecting the preceding grapheme when JSC containing() probes an emoji high surrogate. Route all seven lookups through normalization-core, normalize numeric indexes before the offset, and retain the new eager dependency in the native PR wrapper inventory. Upstream engine fix: oven-sh/WebKit#753.

Node 24 and fork Bun each pass 586 tests (one skipped) across the requested 39 files. Node matches native containing() for all 1,236 corpus lookups. Changed-file checks, import-cycle validation, wrapper closure checks, and independent P2 reviews pass.

The remaining CI cron copy-fault and Discord skills-watcher teardown failures reproduce on clean main and use the authorized native pre-existing-failure exception. The original wrapper inventory regression was fixed.
2026-10-01 09:24:26 -07:00
Peter Steinberger
b5dd0b887c
fix(doctor): close Doctor-owned SQLite handles before the NOCOW store rewrite (#162809)
Preserve the existing maintenance shutdown owner and cover live inspection-child closure plus shared and agent handles through the real Doctor health flow. The suspected self-handle leak was not reproduced; distinguish fuser holder PIDs from inspection failures without relaxing refusal.

Includes the independently landed WAL inventory rescan for integration. Validation: 27 focused cases on Blacksmith Testbox, production and all 27 test type graphs, scoped lint and repository guards; isolated Codex review found no actionable P0-P2 findings. Commit hooks skipped after remote validation to honor the Mac host restriction.
2026-10-01 09:20:16 -07:00
Peter Steinberger
3771e26e81
fix(doctor): defer unrequested NOCOW rewrites during updates (#162786)
* fix(doctor): defer unrequested NOCOW rewrites during updates

* chore(ci): raise the environment-variable count budget for the Doctor NOCOW opt-in
2026-10-01 09:14:20 -07:00
Shakker
72fc2b8829
fix: invite GitHub visitors without a public email (#162666)
Target GitHub visitor invitations by immutable account ID without requiring a public email. Preserve explicit email invitations, verified identity bindings, grant lifetimes, revocation and restart recovery.
2026-10-01 17:12:53 +01:00
Peter Steinberger
01e328528a
feat(macos): add Gateway hosting controls and Node migration (#161779) 2026-10-01 09:02:40 -07:00
Peter Steinberger
27ee259fb8
perf(ui): resolve warm chat routes before Gateway connect (#162794)
Agent discovery became live-only in #161019, but short-route cache
resolution still required agentsList to provide routing defaults. A warm
reload could finish reading its transcript snapshot and still wait for
Gateway connect before returning the route.

Project only mainKey and scope through the existing session cache lifecycle.
Retire these hints on connect, credential/Gateway changes, and disposal.
Keep agent discovery live-only, preserve reserved main-key URL parsing,
and retain exact-session live revalidation and refusal on reconnect.

Bundled Chromium reloads on two isolated Gateways using the same compiled
runtime at 83a974025c, identical synthetic sessions, immutable assets, and
TCP delays: four interleaved rounds per RTT, alternating before/after order.
Median visible time (ms): 0 RTT 321 -> 314; 80 RTT 436.5 -> 316.5;
200 RTT 766 -> 404.5. Per-run before/after values:
0: 329,313,343,307 / 315,313,313,396
80: 429,439,441,434 / 327,323,310,290
200: 770,760,762,776 / 400,417,403,406
At 200 RTT, the original trace read the snapshot at 337 ms, received
connect at 671 ms, and rendered at 759 ms. The fixed trace read at 358 ms,
rendered at 436 ms, and received connect at 723 ms.

Regression proof fails on original production code and passes with the fix.
Relevant role-discovery, routing, cache, and reconnect coverage passes:
270 tests across 16 focused files. Repaired the cold-route fixture's missing
session capability without changing its assertions. Per-file test wall cost:
route-warm-boot 3.49 s; session-roster-warm-boot 10.05 s; route 3.84 s.

Validation: core/UI/core-test/root-test tsgo; both import-cycle checks (0);
all five typed core lint stripes; oxfmt; diff check; bundled UI build and
performance budgets; independent Codex review scoped-clean through P2.
2026-10-01 08:52:28 -07:00
Ayaan Zaidi
e35e8591b8
fix(cron): failing automations get a second repair turn when the error changes, or a repair instead of the final alert (#162702)
## What Problem This Solves

Fixes two gaps in the owner-conversation repair from #161007: a job that keeps failing can get a second repair turn when its error changes, and a one-shot job that has used up its retries gets a repair request instead of its alert.

## User Impact

- Each failure streak gets one repair request, even if its cause changes later. The next failure sends the normal alert once, naming the repair, as before. Later alerts in the streak follow the usual dedupe and cooldown. Only a successful run starts a new streak, which can be repaired again.
- A job that will not run again keeps the existing failure alert and gets no repair request. This covers a disabled job and an `at` job with no retry scheduled, for example one whose retries ran out or whose error is permanent.

## Why This Change Was Made

**Provenance.** #161007 (fb2052a33f) added `failureAlertIncident.repair`. The repair request takes the incident's alert and cooldown slot, and the next failure skips dedupe and cooldown once to send the fallback alert. That fallback goes through `startFailureAlertCycle` (#146582, 5c815031c3), which rebuilds the incident from the new signature and so dropped `repair`. After that, a later failure with a different cause, once the cooldown had passed, looked like a new streak and got a second repair turn, which #161007 meant to prevent ("the streak is never repaired twice"). Separately, the one-shot retry policy (#24355) disables an `at` job when its retries run out. With `failureAlert.after` equal to that attempt, the repair replaced the job's final alert, and no later run could send the fallback.

**Fix (`failure-alerts.ts`, `startup-run-repair.ts`):**
- `startFailureAlertCycle` keeps `repair` and marks it `alerted: true`. Any alert after a repair request is its fallback. The marker lasts until `resolveFailureIncident` (success) deletes the incident. A repair is only requested when the incident has no `repair`, and the fallback bypass only runs while `alerted` is unset. Everything else from #161007 is unchanged: the repair still takes the alert and cooldown slot, and the fallback still fires once. A persisted marker from #161007 without `alerted` gets its pending fallback.
- Repair also requires `isJobEnabled(job)`, checked after scheduling settles. `applyJobResult` already finalizes notifications after a one-shot is disabled or a retry is scheduled. On restart, `markInterruptedStartupRun` disabled a retired one-shot only after finalizing notifications. It now does that first (#131473 logic, unchanged otherwise), so an interrupted one-shot that will be replayed is still repaired and a retired one alerts. Recurring jobs whose next run is recomputed later in recovery also stay eligible.

Persisted state: `repair: { atMs }` gains an optional `alerted: true`, and its lifetime becomes the streak.

**Bounded cost.** At most one repair turn per failure streak. The number of messages does not increase: a repair turn replaces an alert, the fallback is the single alert it was before, and terminal one-shots send the same single alert as before #161007. No new config, warnings, or message types.

Hermes check: Hermes kanban's `failure_limit` auto-block also keeps streak-scoped state that only a success resets. This PR follows the same approach.

## Evidence

- `src/cron/service.failure-repair.test.ts` (real `CronService`):
  - error A ×2 → repair; A → alert; cooldown passes; error B → alert, not a second repair; success; A ×2 → repair again. On origin/main, B started a second repair.
  - an `at` job with `failureAlert.after: 4`, run on schedule (`due`) until its retries are exhausted → disabled, one alert, no repair. On origin/main: one repair, no alert.
  - restart-interrupted owned jobs at the threshold (real `markInterruptedStartupRun`): recurring → repair; replayed one-shot → repair; retired one-shot → alert. It fails without the reorder (the retired one-shot is repaired) and with a `nextRunAtMs` gate (the replayed one-shot alerts).
- Related suites pass: failure-repair, failure-alert, failure-alerts persistence/recovery, startup-run-repair, run-recovery, Gateway `server.cron-failure-repair`.
- **Live Telegram Test Server, live model (`openai/gpt-6-astra`), DM owner conversation.** Each run uses an isolated source Gateway with a fresh state dir. A proxy in front of the real OpenAI API injects provider rejections only into the automation's own model calls: cause A is a 400 `invalid_value` (classified `format`), and cause B is a 404 `model_not_found`. The owner chat and the repair turns go to the live model. Scenario: owned hourly job, `failureAlert {after: 2, cooldownMs: 60s}`. Runs A, A, A; wait 65s; B. Then an owned `at` job (`after: 1`) whose single scheduled run fails with A.

  | Step | origin/main cron files (unchanged since `b2fc927`) | PR head `f064aed` |
  |---|---|---|
  | A ×2 | repair turn → *"Inbox sync is failing because the provider rejects its request format… Please provide the provider's detailed invalid_value error…"* | same repair turn and reply |
  | A (3rd) | alert *"failed 3 times / An automatic repair was requested…"* | same alert |
  | B after cooldown | **second repair turn** (model saw a brief for failure 4) → *"Inbox sync now fails because the provider cannot find openai/gpt-6-astra. Which available provider/model should I use…?"* and no alert | alert *"Automation "Inbox sync" failed 4 times / An automatic repair was requested… / Cause: model_not_found…"* and no repair turn |
  | terminal one-shot (disabled after the run) | **repair turn** → *"Report export failed because the provider rejected the request format. Please provide…"* for a job that will never run again | alert *"Automation "Report export" failed 1 times / Cause: format"* and no repair turn |

  Repair briefs the live model received: origin/main `[1, 2, 4]` (failure counts), PR head `[2]`. The runner completed on the PR head (lease released, scratch removed, no listeners left). The origin/main run hit the scenario command's 540s timeout at the very end, after it had recorded every step.
- **Store reopen (throwaway, real SQLite cron store, not committed).** A seeded #161007-shape incident `{signature, repair: {atMs}}` reloads as-is. The next failure sends the pending fallback (alerts 1, repairs 0) and persists `repair: {atMs, alerted: true}`. After a second `CronService` reopens the store, the marker still has `alerted: true`. A same-cause failure and a new-cause failure within the cooldown send nothing. A new cause after the cooldown sends one alert and no repair.
- Changed suite runtime: `pnpm vitest run src/cron/service.failure-repair.test.ts --maxWorkers=1` → 15 tests, 30.7s total (about 16s of it test time, the rest transform and import warmup).
- Throwaway real-Gateway smoke over WS (`cron.add`/`cron.run`, real cron service and `agent` dispatch, mocked isolated run; not committed): A,A → repair (`agent` dispatch 1); A → alert; A → deduped; B after cooldown → alert `delivered`, dispatch still 1; ok, then A,A → repair (dispatch 2); terminal `at` job with a permanent error → `enabled:false`, alert `delivered`, dispatch still 2. On origin/main, B after cooldown dispatched a second repair.

LOC vs origin/main: prod +31/−23 (`failure-alerts.ts`, `startup-run-repair.ts`, `types.ts`), test +98/−1, docs +1/−1.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-10-01 23:45:07 +08:00
Dallin Romney
9f9dab3306
fix(codex): avoid browser auth in CUA readiness (#162467) 2026-10-01 08:32:00 -07:00
Peter Steinberger
54885d54a0
fix(state): keep agent database admission open across transient lease contention (#162750)
A stalled cron sub-agent tool call drove sustained SQLite write contention; a state.lease operation with busy_timeout 0 failed with lock contention, and the per-agent database execution admission then closed permanently until restart because the execution owner retired logical admission after recoverable native-open failures and the worker transport dropped SQLite contention codes. Admission is now preserved across recoverable native-open failures, contention retries share the 25 ms cadence introduced for the lease heartbeat in #160702 and run only before application work starts, and shutdown, FIFO ownership and unknown-outcome safeguards are unchanged. Doctor preserves nested lease-loss diagnostics instead of flattening them; a nine-transaction Doctor probe confirms #160702 closes the heartbeat loss reported in #162083.

Closes #162579
Refs #162083
2026-10-01 08:16:07 -07:00
Shakker
cfe8791ca9
fix: reuse prepared test runtimes after source refreshes (#162744)
Reuse compatible test runtimes across source refreshes and test-only edits while preserving input invalidation and strict UI/deployment checks.

Related: #57204.
2026-10-01 16:08:26 +01:00
Peter Steinberger
779880c635
fix(ci): reuse candidate artifacts for published-driver updates (#162765)
The published-driver update cell rebuilt the candidate inside its own job and overran its ten-minute bound on the 12:46Z main hourly, cancelling the whole run. The cell now consumes the checksummed candidate from the same run's build-artifacts job (explicit digest comparison, portable across GNU and BSD tools), runs the managed update in about five minutes, and reports a command timeout as a job failure naming the active phase instead of cancelling the workflow. Scheduled-main coverage and omission on unrelated PRs are verified by the plan tests.
2026-10-01 08:04:15 -07:00
Vincent Koc
7996775452
chore(deps): upgrade bundled npm to 12.1.0 (#162581)
* chore(deps): upgrade bundled npm to 12.1.0

* fix(deps): preserve npm source policies

* fix(deps): honor effective npm source policies

* fix(deps): satisfy npm config declaration types

* fix(deps): align npm 12 runtime contracts

* fix(deps): use Windows-safe npm config probes
2026-10-01 21:59:49 +07:00
Peter Steinberger
cd8907aa0c
ci: select affected extension packages for PR boundary checks (#162640)
PR boundary selection: check only extension packages the PR's own diff can affect. When core/SDK declaration inputs change, select directly touched extension packages, a fixed smoke set of at most three broad SDK consumers, and packages that import a changed public plugin-sdk entry directly; skip transitive declaration fan-out. Hourly/schedule and release keep the full boundary check and the negative canary. Kill switch: repository variable OPENCLAW_CI_BOUNDARY_SELECTION=full restores full PR selection (unset means aggressive).

Backtest: all 10 historical PR boundary failures remain selected. 20-PR replay: modeled boundary median 7:22 -> 2:55 (conservative 4:49).

Merged past one inherited red: published-driver-update / Published driver update was cancelled at its 10-minute job timeout. It is cancelled the same way on main in hourlies 36863974207 and 36869865743, and its owner is fixing it (reuse build artifacts; timeout becomes a failure with reason). All other 67 jobs, including tooling, passed.
2026-10-01 07:36:04 -07:00
Peter Steinberger
defb7cf5a6
fix: prevent slow Gateway shutdown during memory sync (#162616)
* fix(state): retain publication leases while Gateway drains

* perf(state): preserve normal publication close scheduling
2026-10-01 14:04:43 +00:00
RoboClaw
98f9d597f7
feat: retain required results without model polling (#162707)
* feat: retain required results without model polling

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

* feat: retain required results without model polling

Worked on by:
- @VACInc

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
OpenClaw-Publication: cec817cf-dea5-4829-859b-02ad1b724e19

---------

Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com>
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
2026-10-01 09:58:49 -04:00
Patrick Erichsen
98f7c8d8cb
feat: migrate existing agents to local Claws (#162329)
* feat: migrate existing agents to local Claws

* fix(claws): keep migration config and provenance clean

* fix(claws): avoid duplicate migration helper export

* test(config): cover runtime snapshot MCP results

* fix(claws): reject Basic authorization during migration

* fix(claws): preserve adopted authority across lifecycle operations

* fix(claws): classify adopted ownership writes as CLI operations

* test(claws): exercise adopted consent through tool execution
2026-10-01 06:52:35 -07:00
Peter Steinberger
777421df55
fix(test): leaked skills watchers flake later fake-timer tests instead of failing their file (#162454)
skills.status and skill snapshot preparation start real @openclaw/fs-safe
watchers. In the shared-worker (non-isolated) Vitest lanes a file that never
closed them left them running after the module reset, unreachable, and their
re-armed poll and change-hint timers landed on whichever fake clock a later
file installed. That file's vi.runAllTimersAsync() then aborted after 10000
timers, far from the cause (session-catalog.session-share.test.ts; leaker
fixed in #162261, an earlier occurrence patched at the victim in #114751).

The non-isolated runner now remembers every real refresh.ts generation from
Vite's evaluated module graph, pairing each closeSkillsWatchers export with
the registry instance that generation imported. It captures at task
boundaries and before every vi.resetModules(), so generations a test erases
or shadows mid-test are still seen. After each file it closes leftover
watchers and fails that file with "skills watchers failed", naming it. No
production code changes.

The composite runner test gains a producer/observer pair: the producer
leaves two reset and shadowed generations open, and the observer in the next
file asserts both were closed. Without the drain, the reset wrapper, or the
import-graph pairing, the producer passes and the observer sees a live owner.
2026-10-01 06:49:49 -07:00
Vito Cappello
54a6149549
fix(telegram): restore progress when reply hooks suppress previews (#161546)
* fix(telegram): deliver progress when reply hooks suppress previews

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

* fix(telegram): honor global block-streaming opt-out

Preserve the explicit global off preference when reply hooks suppress previews, and cover both hook contracts through Telegram HTTP delivery.

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

* docs(telegram): clarify hooked fallback for multi-agent turns

Clarify that configured GroupThread turns retain their existing block policy: their previews are disabled independently of reply-modifying hooks. The new forced fallback applies to ordinary single-agent turns. No delivery-policy or test changes.

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

---------

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
2026-10-01 06:41:47 -07:00
Peter Steinberger
fffa1d0ac1
fix(macos): native tests leak spinning MLX reads and reconnecting Gateway channels (#162590)
TalkMLXSpeechSynthesizer races each stream read against a stall timeout with
AsyncTimeout, which cancels and abandons the losing read. The test fake's
nextEvent() busy-polled with Task.yield() and ignored cancellation, and its stale
modes never close, so every full suite run kept an orphaned read spinning a
cooperative thread for the rest of the process. The fake now stops waiting when
its read is cancelled, as the production stream transport does, and the
stale-stream-timeout test asserts the abandoned read retires.

Two GatewayChannelRequestTests never shut down their channel. Its 30 s watchdog
kept reconnecting through the test's start gate, and each attempt replaced the
parked continuation, printing 10-24 "wait() leaked its continuation" warnings
per run and accumulating parked connect tasks. Both tests now shut the channel
down, and the gate stays released so a late reconnect passes through.

The native-test doc records the rule: work a test starts must end with it.
2026-10-01 06:39:04 -07:00
Peter Steinberger
03ba3c4c14
fix(gateway): bound request starts per connection during setup bursts (#162668) 2026-10-01 13:29:24 +00:00
Ayaan Zaidi
0636d7e772
fix(cron): no-delivery spawn-only runs drop the child's result (#162474)
## What Problem This Solves

Fixes: when a `delivery: "none"` scheduled run's turn only hands work to a subagent, the run ends `ok` right away and the child's result is dropped.

In our setup, a monthly no-delivery job spawned a worker for its report. The worker reported that it was blocked (no shell tool), but the answer was never recorded and the run showed `ok`.

## User Impact

A `delivery.mode: "none"` run whose turn only spawned children now waits for them, but no longer than the run's existing deadline, and then:
- records the child's final reply as the run output (summary), without sending it anywhere;
- stays a quiet `ok` run when the child deliberately answers `NO_REPLY`;
- fails with the existing #117308 errors when the child times out or finishes without a reply, and keeps that run's transcript.

Announce jobs, webhook jobs, runs whose parent wrote its own reply, and `sessionTarget: "current"` jobs are unchanged.

## Why This Change Was Made

Provenance:
- #117308 added spawn-only settlement (wait for the children, adopt the child's reply, fail an empty or timed-out handoff), but only in the announce path.
- #135318 rejects `sessions_yield` in cron turns so the scheduler handles child output under the job's delivery policy, and nothing resumes a cron parent after its children finish. It deliberately left no-delivery jobs without a child wait.
- So nothing owned a no-delivery spawn-only run's child output: the subagent completion is marked `intentional_non_delivery` for a cron requester, and cron has already finalized `ok`.
- #162460 (merged) fixed the reader that this settlement depends on; before it, the run recorded the projected cron prompt instead of the child's reply.

**Decision:** the maintainer approved that no-delivery spawn-only runs settle by their children. This reverses #135318's no-wait choice for that one shape, gated on `deliveryPlan.mode === "none"`. Webhook jobs keep their policy, so a child's text never reaches a webhook endpoint through this path.

**Fix:** settle directly on the children's terminal results, at the follow-up owner (`subagent-followup.ts`).
- `waitForDescendantSubagentResult` reuses the existing descendant wait (`agent.wait` plus registry settlement, under the run deadline and abort signal). It stops as soon as the children settle, because the parent-synthesis grace exists only for parents that get resumed.
- It then reads the child's reply through the existing fallback reader. A child whose terminal reply is explicitly `silent` with an `ok` outcome settles as `NO_REPLY` rather than as "no output".
- `dispatchCronDelivery` records that result for none-mode spawn-only runs. A failed handoff no longer triggers delete-after-run cleanup, so the failed one-shot job keeps its transcript, as failed executions already do.

Hermes Agent's cron runs delegated work inline and takes the run's own final response. OpenClaw children settle after the parent turn, so the scheduler reads the children's results instead. It is the same idea: the run's output is the work's final answer, and nothing waits for a reply that will never come.

## Bounded cost

- **The wait:** the existing descendant wait, under the run's existing `timeoutMs` deadline and abort signal. It ends when the children settle; there is no extra 5 s synthesis poll. A timeout records an error; it does not retry.
- **No nested waits:** cron turns cannot `sessions_yield` (#135318). The wait observes the registry and the descendants' `agent.wait`; it starts no runs and wakes no parent.
- **Error status:** a failed handoff is an ordinary execution error. Explicit child silence is not an error. One-shot jobs retry only through the existing transient-retry classifier; recurring jobs only back off.
- **Model calls:** none added.

## Evidence

Live model, Telegram Test Server, real-user recorder (`telegram-e2e-userbot`, DM), with a live OpenAI model behind the provider slot and `tools.toolSearch: false`. The scenario creates an isolated `--no-deliver` cron job, then runs `cron run --wait` and `cron runs`. The live parent chose `sessions_spawn` itself and the child answered live. Only the parent's turn after the accepted spawn was steered to end empty, because that is the spawn-only shape.
- before (main `18154cae312`): the run was `ok` after 6.5 s with an empty summary. The child answered live (`Ganymede is the largest moon of Jupiter.`) but its reply was dropped.
- after (this head): the run was `ok` after 9.0 s, `deliveryStatus: not-requested`, and its summary was the child's live answer, `Ganymede is the largest moon of Jupiter.`
- In both runs the recorder captured zero bot messages in the DM.

Real flow, isolated Gateway with qa-lab mock-openai, temporary scenario (not committed): `delivery: none` jobs with `toolsAllow: [sessions_spawn, read]`, run through `cron.run`.
- previous head of this PR (announce-path reuse): visible child recorded after 8.6 s; a child answering `NO_REPLY` became `error: cron child-session handoff completed without a final assistant payload`.
- this head: visible child recorded after 3.7 s (no synthesis grace); `NO_REPLY` child gives `ok` with an empty summary. Zero outbound messages in both cases.

`delivery-dispatch.double-announce.test.ts` runs the real follow-up module (registry and Gateway mocked only at their edges), with coherent job/plan fixtures and the production 5 s grace under fake time:
- a child settling 3 s before the run deadline is recorded (on the previous head, the watchdog fired during the grace and the result was lost);
- an explicitly silent child gives a quiet `ok`;
- a webhook spawn-only job is unchanged;
- timeout and empty-child failures keep the transcript.

The first three fail on the previous head. `delivery-dispatch.double-announce`, `subagent-followup`, `run.session-lifecycle` and `run.meta-error-status` pass (164 tests).

No overlap with Pash/Sarah changes.

LOC vs main: production +81/−2, tests +153/−4, docs +1/−1.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-10-01 21:25:05 +08:00
Vincent Koc
9bf17b16a8
fix(macos): stop push-to-talk capture when its hold ends (#143064)
* fix(macos): stop push-to-talk capture when its hold ends

* fix(macos): remove unused push-to-talk key codes

* fix(macos): bind voice lifecycle to app-owned runtime

* style(mac): format voice runtime test catalog arguments
2026-10-01 12:59:11 +00:00
kevin2966n
d06974d245
fix(openai): preserve structured output on Responses paths (#114204)
Fixes #114200.

OpenAI Responses paths now send structured output correctly: a raw JSON Schema `responseFormat` is wrapped as `{type: "json_schema", name, schema}` and set as `text.format` for the direct OpenAI and Azure Responses providers. `json_object` and `text` pass through unchanged, Chat Completions is unchanged, and SDK retries stay at 0 so failover keeps owning retries.

Proof: isolated Gateway captured the raw schema on main and the wrapped `json_schema` with this change. Live check on the exact head with a real OpenAI key (gpt-4.1-mini, Responses): the schema turn returned HTTP 200 with schema-conformant JSON, and a control turn without `responseFormat` also succeeded. Regression tests fail on main and pass here.

Co-authored-by: Codex QA <codex-qa@local.invalid>
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-10-01 20:57:32 +08:00
Ayaan Zaidi
4d755295bc
fix: steered message no longer replaces a question whose model request failed (#162439)
## What Problem This Solves

Fixes: with the default `steer` queue mode, a message sent while the first message's model request is failing replaces the first message. The chat shows `⚠️ <model> request failed.` and then only the answer to the second message. The first question is never retried or answered.

## User Impact

When a model request fails while a newer message is waiting to steer into that turn, OpenClaw retries or falls back for the first message as usual. The newer message then runs as its own turn. Both questions get answered in order. The failure notice only appears if the first message's retries and fallbacks are all exhausted, the same as when nothing was steered.

## Why This Change Was Made

Root cause: after a model request failed, `AgentSession.handlePostAgentRun` still continued the session with the queued steer. The embedded runner owns retries for a failed request (`agent-project-settings.ts` turns session auto-retry off, #73781), so the failed message was never retried. The steered message ran in its place, against a transcript where the failed message looked handled. Per-input answer segments (#149925) then correctly reported the first segment's failure, so the user got `⚠️ request failed` plus the second answer.

Provenance:
- `steer` became the default in `4a6e10ece8` ("feat: default queueing to steer") and #77023, to keep the active turn responsive without starting a second run.
- The post-run `hasQueuedMessages() ? "continue"` came from #109709 so that messages queued by `agent_end` handlers are not stranded. That reason still holds after successful turns. This PR only excludes failed requests, where continuing skips the run owner's retry and fallback.
- Steered channel input already waits for its transcript commit (`agent-runner-steer-adoption.ts`, `waitForTranscriptCommit: true`). When the session settles without committing it, the steer is withdrawn from the runtime queue and the parked reservation falls back to the ordinary follow-up queue. That path runs after the reply operation clears. No new queue or retry logic is added.

Fix: one condition. After a failed request, queued input no longer continues the session (`agent-session-prompting.ts`). Docs: `docs/concepts/queue-steering.md`.

Hermes comparison: Hermes injects a steer only at a role-safe boundary after a tool result. With no tool row yet, or when the request fails, it requeues the steer as the next turn's user message, and each queued message gets its own turn. This change gives OpenClaw the same outcome by reusing its existing follow-up fallback.

Relation to #161069 (stalled-turn recovery): that PR covers turns aborted by stuck-session recovery. This one covers a provider failure while a steer is pending. The two don't overlap.

## Bounded cost

- Fewer model calls: before, a pending steer triggered one extra model call inside the failed run. Now it doesn't.
- The first message keeps its existing retry budget (`MAX_EMPTY_ERROR_RETRIES = 3`, plus the configured fallback chain). Nothing new is retried.
- A handed-back steer is queued at most once. The fallback drops its `steerPending` marker, so it drains as an ordinary follow-up turn and is never steered again. If that turn also fails, it ends with the normal failure notice and is not requeued. The worst case is one follow-up turn per inbound message.
- Tested: the session regression test asserts exactly one model request after the failure, with no continuation. The requeue-once path is existing behavior covered by `src/auto-reply/reply/queue/enqueue.steering.test.ts` (6/6 on this head).

## Evidence

**Telegram Test Server** (`telegram-e2e-userbot`, Convex-leased credential, real QA user in a DM, default queue mode). A deterministic mock provider fails agent-turn requests whose newest question is `17*23` three times (`response.failed` after 8 s). Like the live model in our earlier proof, it answers only the newest user question. The QA user sends `What is 17*23?` at +0.2 s and `Also, what is 19*21?` at +8.2 s, so the second message arrives during the first failing request.

- Before (built head with only this line reverted): the SUT sends `⚠️ openai/gpt-5.5 request failed.` at +13.1 s, then `399` at +14.1 s. The request after the failure carries both questions (`developer,user,user,user`). The first request is never retried, and `391` is never sent.
- After: three failing first-question requests (`[empty-error-retry]` attempts 1/3 to 3/3), none carrying the second question. The fourth request answers `391` at +30.5 s. The second message then runs as its own turn (`developer,user,user,assistant,user,user`) and answers `399` at +31.5 s. No failure notice is sent.

**Unit test:** `agent-session-loop-next-turn.test.ts` "does not answer a steer in place of a failed request". On `main` the steer is committed into the failed turn (`promise resolved instead of rejecting`). With this change there is one request, the steer's commit wait rejects so its caller requeues it, and nothing is left queued. File: 24/24 pass. Related suites pass: `sdk`, `agent-session-loop-correctness`, `attempt-prompt-submit` (+ steering, retention), `attempt-stream-prepare`, `attempt-session-replay` and `provider-review-continuation`, 231 tests.

**Telegram with a live OpenAI model** (same Telegram Test Server DM; a small local proxy in front of the real OpenAI API fails only the first two agent requests that carry 17*23 without 19*21, and forwards everything else to the live model):
- Before (this line reverted in the build): one failing request, then the next request carries both questions and goes to the live model, which answers `399`. The bot sends `⚠️ openai/gpt-6-astra request failed.` (+20.4 s), then `399` (+21.0 s). 17*23 is never answered.
- After: two failing requests for the first question alone, then the live model answers the retry with `391` (+24.1 s). The second message's own turn gets `399` (+26.0 s). No notice.
- Both runs used the fix build, before the branch was rebased onto newer main; the production change is the same. For the before run, only the fix line was reverted.

**Telegram, when the first question fails on every retry and fallback** (same harness; the mock fails every agent request whose newest question is 17*23, and the agent has a fallback model configured): 8 requests for the first question, 4 on `gpt-5.5` and 4 on `gpt-5.5-fallback`. None of them carries the second question. The model-fallback log shows `gpt-5.5` failing over to `gpt-5.5-fallback`, then nothing left to try. The SUT sends exactly two messages: one `⚠️ openai/gpt-5.5-fallback request failed.` at +70.9 s, then `399` at +71.9 s from the second message's own turn. No other notices.

**Web UI `chat.send` (real Gateway, operator WebSocket client of the kind the TUI uses, same mock):** the client sends `What is 17*23?`, then `Also, what is 19*21?` 8 s later, during the first failing request.
- Before (this line reverted in the build): the client's only answer is `399`. The request after the failure carries both questions (`developer,user,user,user`). `chat.history` ends `user 17*23, user 19*21, assistant 399`, so the first question is unanswered.
- After: the steer is withdrawn when the failed attempt settles. Its `chat.send` run closes with an empty final (+17.1 s), and `chat-send-agent-dispatch.ts` dispatches it again; it runs as a follow-up and is not steered into the retry. The first run retries three times, with no retried request carrying the second question, and answers `391` (+33.7 s). The second message's own run answers `399` (+34.0 s). `chat.history`: `user 17*23, assistant 391, user 19*21, assistant 399`.

Why a withdrawn Web UI steer can't be re-steered into the retry: with explicit `queueMode: "steer"`, chat.send's fallback dispatch passes `messageInjectionDisposition: "rejected"` (`src/gateway/server-methods/chat-send-agent-dispatch.ts:388-390`). `runReplyAgent` steers only when the disposition is `"none"` (`src/auto-reply/reply/agent-runner-run.ts:341-346`), so the message is queued as a follow-up instead (`:389`). With the default mode, chat.send makes no injection attempt of its own (`chat-send-admission.ts:272-275`) and the message takes the channel path: its parked reservation falls back with its steer marker removed (`queue/enqueue.ts:307-318`), and the queue drains only after the reply operation clears (`agent-runner-steer-adoption.ts:110-122`).

**Wall times** (`pnpm test <file>`, shared, heavily loaded host):
- `agent-session-loop-next-turn.test.ts`: 24/24, 121 s wall (vitest 106 s). The new test signals request start and steer acceptance with deferreds; it no longer polls.
- `enqueue.steering`: 6/6, 78 s.
- `attempt.queue-message`: 16/16, 98 s.
- `agent-session-handoff-adoption.integration`: 1/1, 63 s.
- `agent-runner-steer-adoption.question-recovery`: 25/25, 97 s.
- `chat-send-steering-custody`: 17/17, 90 s.
- `agent-runner.runreplyagent.e2e`: 223/224 under load. The one failure, "keeps the replacement source when retired admission completes", passes alone (60 s wall). That file mocks the embedded agent, so it never reaches the changed session code.

qa-lab note: the qa-channel serializes inbound per account, so the same-sender steer couldn't be reproduced there. The proof uses the Telegram Test Server instead.

LOC vs `origin/main`: production +3/−1, test +46, docs +1 (changed line).

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-10-01 20:42:48 +08:00
Peter Steinberger
5c8a714982
refactor(providers): deslop provider adapters (#162488)
* refactor(providers): deslop provider adapters

* refactor(providers): preserve the auth profile SDK declaration

* test(providers): complete the prepared auth bootstrap fixture

* refactor(providers): make detached catalog method contracts explicit
2026-10-01 05:41:53 -07:00
Stronggenetics
72f5840c11
fix(claude-cli): subagent completion turns run tool-free and fabricate tool calls (#162374)
* fix(claude-cli): subagent completion turns run tool-free and fabricate tool calls

When a subagent finished and handed its result to a requester running on
the Claude CLI backend, the requester's completion turn ran with no tools.
Native CLI tools cannot enforce the requester's persisted tool cap, so the
turn was forced tool-free. The model then wrote tool calls as plain text
and reported invented output as verified.

Restore the requester's policy-filtered tools for that turn on Claude CLI.
Native CLI tools stay disabled; the tools reach the CLI only through the
mediated MCP loopback grant.

- The attempt keeps requester tools for a verified completion on Claude
  CLI, and passes the verified handoff to the CLI runner. Node-hosted
  sessions and settle batches are excluded.
- CLI preparation resolves the whole surface through the existing policy
  projection with native tools empty. It refuses a backend that cannot
  enforce the cap.
- The loopback grant carries the handoff and its provenance, and stays
  bound to the source reply under message-tool-only delivery.
- Loopback tool resolution rebuilds the requester policy from the handoff,
  so inherited denies and the persisted cap still apply. A grant whose
  requester lineage no longer verifies fails closed, and the lineage is
  rechecked before a cached tool list is served.

Forged, missing, and unverified handoffs stay tool-free, as do other CLI
backends. Sibling modules hold the extracted policy steps so the touched
files stay within the line-cap ratchet.

Related: #121661

* fix(agents): keep the Claude CLI fallback prelude builder in the attempt

The first commit moved the only production caller of
buildClaudeCliFallbackContextPrelude into the helpers file that defines it,
so the production dead-code scan reported the export as unused.

Build the prelude in the attempt again, as on main. The helpers file keeps
only the condition for seeding it, which holds attempt-execution.ts inside
its line cap.

Related: #121661

* fix(gateway): recheck completion lineage at the final tool-effect fence

A Claude CLI completion grant verified its requester lineage only when the
loopback tool list was resolved. The child entry could be deleted or
re-parented while tool preparation, a before-tool hook or an approval
awaited, and the resolved tool then ran: dispatch authorization and the
tools' source-effect guard checked the grant and the run, not the lineage.

Compose lineage currentness into both. The loopback server now asks one
predicate for dispatch authorization, for the invocation check handed to
tools, and for the caller identity's receipt and approval authority that
write, exec and the other coding tools assert before their I/O. Grants
without a completion handoff skip the check.

Tests drive the real loopback server with a real grant, a persisted child
entry and the real write tool: allowed after an awaited hook; rejected
when the child is removed or re-parented during the hook, removed during a
hook approval wait, or removed while the authorized write waits before its
I/O; and rejected for a lineage that belongs to another requester.

Related: #121661

* fix(gateway): check completion lineage when a cached tool list is served

The loopback tool cache verified a completion grant's lineage on entry,
before its lookups await. A child removed or re-parented during that wait
could still get the previously cached tool list.

Run the check at the point a cached list is returned. A list that is built
fresh is already verified by tool resolution. The new test re-parents the
child while a cached lookup is pending and fails without this change.

Related: #121661

* test(gateway): pin the completion owner at the tool-effect fence

Lineage for a completion grant is the child's persisted completion owner
when one is set, and its spawner otherwise; that rule lives in the
requester-policy resolver and is unchanged here. Cover it at the loopback
fence: the completion owner may write while another session controls the
child, and handing the child to another completion owner during an awaited
before-tool hook rejects the write.

Related: #121661

* fix(claude-cli): narrow completion tool restoration to automatic replies

---------

Co-authored-by: Stronggenetics <285854990+Stronggenetics@users.noreply.github.com>
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
2026-10-01 08:38:15 -04:00
Peter Steinberger
25972a6a28
ci: published-driver update cell for updater and identity paths (#162629)
Adds a CI cell that installs the latest published stable openclaw as the driver and runs a managed update to the candidate built from the current revision, asserting a finished run, the candidate version, Gateway readiness, and no canary, identity, or lease warnings. It is path-gated to the updater, lease/identity, state-database-open, plugin native admission, and startup-trace surfaces and counted by the aggregate gate; on the dispatch-fallback path (checkout revision differs from github.sha) it skips with a recorded reason. Motivation: the four 2026.9.7-only update regressions (#162131, #162130, #161746, #162047) were landed by PRs whose tests exercised encoder and decoder in-process; this cell fails on the pre-fix tree. The reusable workflow checks out github.sha with read-only permissions, no persisted credentials, no cache writes, and no secrets.
2026-10-01 05:38:00 -07:00
Peter Steinberger
1b896e6327
perf(sessions): bound list materialization after invalidation (#162280)
* perf(sessions): bound list materialization after invalidation

After a catalog or metadata invalidation, sessions.list materialized the
display rows of every session before selecting the page, so a burst of
concurrent lists spent seconds in the materialize phase (avg 4.7 s, max
19 s on Team with 334 slow lists per hour).

Select first, then materialize only the rows the page returns, and share
the post-invalidation work across concurrent lists. The reply contract is
unchanged.

Testbox fixture (8,000 sessions, 20 concurrent lists): average
materialization 534 -> 97 ms after catalog invalidation and 1,239 -> 841 ms
after metadata invalidation; selected display work 8,000 -> 200 rows.

* test(sessions): align list fixtures with bounded materialization

Align session-list fixtures with metadata readiness and selected-page
materialization. Fixtures that need every display row now drain their own
setup explicitly, while list wait tests observe prepareSelection and
preserve its arguments.

The synthetic plugin role-revocation test still waited for the retired
ensureMaterialized list call and deterministically timed out after 120 s.
Observe the actual reader readiness call, assert that each injected wait
was consumed, and retain the revocation-error expectation and describe's
bulk-read isolation guard.

Testbox: 435 tests across 40 files passed; the reproduced 120,058 ms
revocation timeout now completes in 670 ms. PR #162280.
2026-10-01 12:31:33 +00:00
Peter Steinberger
28ac1ad96f
refactor(voice-call): decode legacy call logs only in Doctor (#162538)
* refactor(voice-call): move legacy call decoding into Doctor

* chore(voice-call): shrink retired decoder assertion baseline
2026-10-01 05:30:36 -07:00
Vincent Koc
94f5a8d586
feat(skills): search bounded installed skill instructions (#160538)
* fix(build): keep pinned pnpm probes from dirtying the lockfile

* feat(skills): search bounded installed instruction bodies

Index bounded instruction text lazily through the admitted catalog reader, retaining metadata-only search for unavailable or over-budget bodies. Revalidate reader authority on cache hits and preserve whole-body skill reads. Report index coverage in search results.

Validation: 33 focused tests across four files passed on the native proof checkout in 117.62s; fresh scoped autoreview clean through P2; all eight changed files pass oxfmt. The coding worktree is deliberately code-only, so formatting ran through the qualified independent tooling checkout before this commit.

* fix(skills): require current native read authority for body search

* test(skills): prove disk-backed search authority through dispatch

* test(skills): tie authority gates to test cancellation

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-10-01 12:19:06 +00:00
Peter Steinberger
55bbbf9ddc
refactor(doctor): route plugin config repairs through one executor (#162641)
Run registered setup migrations, Doctor contracts, legacy channel previews and interactive channel compatibility hooks through one executor. Preserve migration phases, public registration and adapter receivers while isolating failed candidates and retaining warning-only results.

Regression failures were reproduced on the base; focused and plugin-contract tests, changed checks, both cycle checks, SDK API comparison and independent reviews completed successfully.
2026-10-01 12:09:44 +00:00
Peter Steinberger
084990ed4f
fix(process): release idle Bun spawn broker IPC references (#162655)
Bun-hosted published updates could hang after candidate Doctor completed because the spawn broker retained an idle IPC channel. Honor optional channel reference controls while retaining live requests, native resource claims, and shutdown work.

Published openclaw@2026.9.7 updates passed on both Bun builds and a Node 24 control, including service replacement, authenticated readiness, and preserved backup bytes. The 22 focused broker cases pass on Node 24 and the prescribed Bun fork after a patch-identical rebase. The installed updater, delegated Doctor markers, schemas, and rollback contract are unchanged.

CI exception: both prebuilt UI global setups failed on inherited stale pnpm lockfile metadata before tests; independent main reproduction confirmed the cause, and main fixed it in 4b32b40150. Fail-fast cancellations remain unrun coverage. Landed through the authorized native pre-existing-failure route.
2026-10-01 05:05:10 -07:00