The upgrade-survivor candidate identity check surfaced a raw node:fs read error instead of the identity messages when the candidate payload was missing or unreadable (test/scripts/upgrade-survivor-candidate-identity red on main since #160448 narrowed the contract). The generic baseline contract is restored (the captured baseline stays mandatory for cron-owner-doctor; generic scenarios use baselineCommit), and both filesystem error boundaries now produce the contextual identity error with the original cause retained.
Refs #160448
After a Doctor error the package rollback captured its snapshot while retained writers were still active, so the rollback refused with a false "changed" verdict and the npm error-path result carried an unexpected shape (update-cli.package-lifecycle red on main since #162604). Rollback now drains retained writers before capture on POSIX by awaiting the cache close and re-checking ownership, without blocking POSIX snapshot readers; Windows keeps its existing removal probe. Every existing assertion is unchanged.
Refs #162604
Let session creators with session-scoped write access use shared GitHub publication from chat while keeping personal publication restricted.
Refs #162916.
The published-driver update cell ran serially after build-artifacts on every PR it was selected for, adding about 13.5 minutes of critical path. PR selection is now limited to the updater, activation, managed-service handoff, and canary-emitter owners (scheduled main and release validation still run the cell unconditionally); published-install caches have main-only writers with explicit cache-mode gating and read-only PR consumers; PR runs skip initial Doctor seeding and use shorter readiness waits. PR-mode cell measured at 3m32s with a cold image-builder cache. The cleanup tests no longer touch host Docker.
snapshotCache in src/agents/shell-snapshot.ts was a process-wide Map that retained one entry per distinct cwd/environment key for the Gateway's lifetime, so long-running Gateways with many sessions grew without bound. The owner now reuses the existing LRU cache with a documented 128-entry limit; the active key stays cached and a miss recomputes. At 201k synthetic keys RSS fell from 265 MiB to 232 MiB and live heap from 74 MiB to 6 MiB.
Closes#162582
* fix(gateway): explain preserved scope limits after pairing approval
Distinguish successful pairing that preserves a narrowed token from rejection. Keep the outcome internal, return existing non-retryable RPC guidance, and never release an insufficiently scoped token. Document same-origin recovery through the existing one-time owner dashboard handoff without changing approval policy.
Add a real Gateway RPC regression that fails on the former rejected response and verifies preserved read-only access, no returned token, and actionable recovery guidance.
* test(gateway): check scope outcome retryability through the RPC matcher
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Allow verified session-scoped callers to read shared GitHub publication options and receipts while preserving current authority, visibility, and personal account boundaries.
Refs #162916.
* perf(gateway): prepare history workers before the first chat read
The history worker reached the full session lookup/runtime graph through
source candidate selection. Branch snapshots also imported their host list
coordinator and writer-side reexports. Extract the shared candidate owner
and isolate branch snapshots/cache from host coordination without changing
selection, admission, watermark reuse, restoration, or cleanup.
Prioritize complete history preparation, including the response encoder,
and allow only this worker operation during foreground browser loading.
Main-thread handler/optional preparation remains idle-only. Preserve the
post-ready barrier, root admission, shutdown joins, and worker custody;
ready never awaits this work. No schema, protocol, public SDK, configuration,
dependency, persistence, or update behavior changes.
Reproduction: three isolated seeded starts on Node 26, zero added RTT.
Cold chat.startup server ms: 2020,586,2041; visible ms: 4440,2246,3739.
Warm server ms: 12,12,12; visible ms: 544,12098,16714. The last two warm
outliers occurred before WebSocket connection, not in chat.startup.
Four interleaved rounds against a preserved build of main 83a974025c,
same worktree/runtime paths and cloned synthetic fixtures (1500 turns plus
30 x 60 turns), fresh browser per load, immediate navigation after ready:
- Main cold server/visible ms: 647/2356,751/2557,556/2015,1034/2893.
- Candidate cold server/visible ms: 59/2059,51/1644,55/1760,89/1877.
- Main warm server/visible ms: 11/456,14/472,13/546,12/484.
- Candidate warm server/visible ms: 11/503,13/493,13/596,13/505.
Cold medians: server 699 -> 57 ms (-91.8%); visible 2456.5 -> 1818.5 ms
(-26.0%). Warm medians: server 12.5 -> 13 ms; visible 478 -> 504 ms.
Branch RPC cold medians: 657.5 -> 572.5 ms.
CPU diagnostics aligned to the first chat.startup server interval:
baseline 1171.4 ms under history-worker module loaders in a 1278 ms request;
candidate 0 ms in a 71 ms request. Candidate loading instead sampled
2635.7 ms after ready, before the request. Samples are approximate wall-time
attribution, not exclusive CPU accounting. Branch worker bootstrap still
costs time: the final loaded-host profile sampled 1079 ms loading in a
1930 ms branch RPC. No claim that branch cold loading is eliminated.
Full process-to-ready medians were 14352 -> 16372 ms on the shared host;
listening-to-ready medians were 674.5 -> 646.5 ms. Warmup remains after
ready; total restart-time improvement is not established.
Validation: 224 tests across 12 focused files pass. Import-boundary and
foreground-prewarm regressions fail before their fixes. Changed-file test
wrapper wall costs: session-history-read.imports 10.86s (2 tests),
server-startup-handler-prewarm 5.94s (12 tests), session-utils 51.46s
(104 existing tests; only its type import moved). The architecture guard
checks static worker boundaries; no timed performance assertions were added.
Core, core-test, root-test typechecks pass; both import-cycle checks report
zero; all five typed core lint stripes pass, with the affected Gateway
stripe refreshed after the final change. Full runtime/postbuild/UI builds,
UI performance checks, oxfmt, diff checks, and Codex P2 review pass.
* fix(gateway): drop unused discovery-cache type re-export
Callers import GatewaySessionStoreDiscoveryCache from its new owner,
session-utils-store-candidates.ts; the session-utils barrel re-export had
no consumers and failed the production deadcode check.
Share the boundary selector's package policy: directly changed plugins,
Telegram/Codex/Slack smoke for shared sources, and direct consumers of changed
public SDK entries, including test imports. Defer transitive and ambient
consumer fan-out to hourly main and release validation. Keep full lint for
lint/type policy and directly changed loose extension sources so their own
canonical lint coverage remains intact. Log selected roots and reasons in
addition to the existing job summary.
Twenty complete recent PR path scenarios reduce the median selection from
164 to 3 packages and aggregate selections from 2461 to 237; full cases fall
from 15 to 1. The one confirmed own-diff historical lint failure retains all
three failing roots, with no observed miss. This is limited historical proof.
The existing OPENCLAW_CI_EXTENSION_LINT_FULL=true or 1 switch restores full PR
coverage; unset keeps the aggressive default. No variable or setting changed.
Full SDK preparation and typed lint programs remain. A scoped-preparation
experiment did not improve the measured cold smoke stripes, so it is not part
of this change. Existing preparation plus selected lint takes 133-144 seconds
on a four-CPU, 16-GiB Linux Testbox; adding observed median setup models about
3m12s per cold job, not the two-minute target or measured Actions wall.
Proof: whole tooling across 927 files, full test types, changed-file and
boundary lint, architecture, source-contract checks, real preflight in both
checkout shapes, and P2 review. Three unrelated upgrade-survivor fixture
failures reproduce on the unmodified parent. The peer's declaration-fixture
repair is preserved and passes the final owner replay. Selector tests cost
3.26 seconds in the local single-file diagnostic and also passed on Linux.
The published-driver Docker script's direct-caller default promised the CI cell's envelope but still used now+525 after #162586 raised CI to 1125 s, so the release Docker E2E published-driver-update lane (which passes no deadline) stopped work at 465 s while hosted cells after #162858 need p50 521 s and up to 659 s. Direct callers now get the same 1125 s envelope, the release-lane watchdog outlives it at 20 minutes, docs/ci.md says twenty minutes, and a behavioral test runs the real CI budget step and the real script with a recording fake docker so the two envelopes cannot drift again.
Fix plain openclaw startup on trusted POSIX Bun-only global installs by
pinning the installed Bun executable. Keep paths as inert launcher data so
released updaters can relocate them safely, with existing ownership, Doctor
consent, and rollback handling.
Validated 114 focused tests per runtime, package/tarball integrity, native
published-9.7 plain/apostrophe upgrades, and byte-exact rollback controls.
Document the inherited 9.7 readiness wait and the 9.6 upgrade limitation.
CI exception: the only failing test job reproduces an inherited browser
mock-cache defect already fixed on main by #162796; all other selected
checks passed. Exact-head ClawSweeper found no actionable code issue.
OpenAI's gpt-6.1 generation publishes no dated snapshots either, so the Reef
guard rejected the Team configuration that pinned gpt-6.1-sol and the channel
could not start. Admit the exact id beside the gpt-5.6 ids, with the same
documented residual risk; bare family aliases stay rejected.
A retained legacy session transcript truncated to 0 bytes (an earlier process left two .jsonl.bak-<pid>-<timestamp> backups beside it) put Doctor into an unresolvable retained_plugin_source_conflict loop: with the file present recovery reported the source incomplete and protected, with it moved recovery reported the source changed, and every Gateway start logged a degraded state. The retained-source verifier required a transcript header and ignored the backup siblings. Recovery now verifies those backups, imports missing history through the existing importer with identity and ownership checks, and reversibly archives the empty original; canonical edits and deletions remain authoritative, and unrecoverable sources stay protected with exact manual restoration instructions.
Closes#162817
Doctor's own inspection workers and cached shared-state readers kept the conservative holder census from reclaiming abandoned updater runtimes.
Settle those resources through their existing owners while retaining maintenance authority, then run the unchanged census. Preserve independent holders, runtime-only service boundaries, and the refusal to restart after resource settlement fails.
Validated 53 focused/regression cases on Node and fork Bun, published 2026.9.7 upgrades and positive/negative custody on both runtimes, scoped-clean P2 review, and green exact-head CI/security checks.
Await buffered progress receipts in the fresh candidate receiver before finalization, preserving captured database context and live executor/requester admission. Refused receipts stop replay; uncertain writes retain the existing cleanup failure. This candidate-side change preserves published drivers, schemas, stored state, and recovery policy.
* fix(ui): localize onboarding welcome in Chinese
Use the selected Control UI locale in the existing connection metadata and translate onboarding prose and question labels through the wizard catalogs. Preserve command reply payloads and English fallback. Covers simplified and traditional Chinese without a protocol or configuration change.
* fix(ui): preserve onboarding locale during catalog loading
Use the i18n owner’s requested locale for the existing Gateway connection metadata, including while the lazy translation chunk is pending. Preserve rendering fallback and offline retry behavior. Add Gateway locale forwarding coverage, a held-chunk browser regression, and operator documentation. Remote proof passes 173 owner tests and four browser cases; the new browser case takes 734ms. No protocol or config surface changes.
Co-authored-by: Peter Steinberger <steipete@gmail.com>
When the candidate migration rehearsal is terminated at its budget, the update reported "The target configuration could not provide a usable inference route", a message from secondary repair setup, instead of what actually ran out of time. validateUpdateCandidateCanary now records the rehearsal step, elapsed seconds, and the last sanitized output in the failure fact and the update report. The flat 300 s cap in the report belongs to the installed 2026.9.3 driver; main already derives the rehearsal budget from database size (a 420 MB fixture rehearses in 24 s within an 802 s budget) and uses the responsive integrity child from #162213.
Refs #162737
## What Problem This Solves
Native chat lowercased complete session keys when admitting Gateway events. Distinct Matrix rooms/threads, Signal groups, or catalog conversations whose opaque IDs differed only by case could therefore appear in the selected transcript.
## User Impact
Native chat keeps those conversations separate while continuing to accept structural routing aliases. Stored keys, Gateway wire formats, main/global routing, and composer identity are unchanged.
## Why This Change Was Made
The sidebar already implemented the comparison contract used by the Control UI. Move that implementation to the existing shared session-key owner and use it for sidebar, native event filtering, and the public default-main matcher. Catalog bodies remain fully opaque; the canonical Control UI contract normalizes only their agent prefix. Matrix/Signal routing words retain their existing normalization.
A five-case regression passes actual `session.message` frames through the payload codec and registered dispatcher. Each case rejects a different opaque ID and accepts its structural alias. The existing public-matcher table also covers its separate Talk-facing API contract.
## Evidence
- Remote `check-changed` passed on Blacksmith Testbox.
- Both import-cycle checks passed with **0 cycles**.
- Focused source review and isolated independent P2 review completed. The one review concern about lowercasing the catalog discriminator was rejected against the existing Control UI source contract, which deliberately preserves the complete catalog body.
- Swift formatting and `git diff --check` passed.
- Exact-head hosted OpenClawKit CI passed 2,000 tests in 168 suites (35.259s), plus 21 NativeState tests. The new five-case frame/dispatch regression passed in 5.206s including concurrent-suite scheduling. Isolated-file wall time was not measured. The regression has not been executed against the original implementation: Blacksmith supports Linux only, and the AWS existing-host Mac route returned no available Dedicated Host. The original whole-key lowercase path and its event-dispatch effect were traced directly; this is source evidence, not an observed baseline test failure.
- Tests add no sleeps, polling, process boots, or production seams. The full macOS app and iOS smoke jobs remain separate hosted evidence.
Found during the sibling duplication investigation. The production change removes one net line by sharing the existing comparison owner; tests add 52 lines and documentation adds three.
## Fixes found along the way
The macOS cloud-worker fixture signaled readiness through file existence while `printf` was still writing its argument capture. Hosted run 36887822355 observed output ending at `--ephemeral`, before the final display-name arguments. Publish the complete capture with a same-directory temporary file and atomic rename; every original assertion and timeout stays intact. This fixture-only repair has a clean independent P2 review. Linux Testbox proof passed 25 runs per enrollment mode (50 total), two gated atomic-publication controls, and two original-publication adverse controls. The proof verified the exact committed fixture bytes and preserved every argument, including a Unicode path. Both cycle checks again reported zero. This proves shell fixture publication; updated-head Mac ProcessIdentity/AppKit and full native CI remain separate evidence.
That run also encountered `AXError.attributeUnsupported (-25205)` while the test helper requested the application's accessibility windows, before GatewayInstallerView's text/action assertions. The same helper failure is documented in #158049 without a proven fix. The affected owners are unchanged by this PR, recent inspected main runs passed, and no matching current-main failure has been established. The assertion is retained; this earlier failure is not claimed fixed or bypassed.
Keep pending attachment claims with the submit guard until durable admission. Retire only unchanged files in the captured current composer scope while preserving newer drafts and overlapping submissions.
Related: #137427. Broader pre-admission handoff work remains separate.
* fix(test): load compiled-subprocess declarations at collection
The first load of a compiled-subprocess declaration in a Vitest invocation
prepares the whole compiled worker generation (tens of seconds warm, minutes
cold). 387 test files first reached a declaration through an await import()
inside a test body or hook, so their first test or hook absorbed that
preparation and could time out.
Move those loads to collection at their owner: static subject imports,
static imports in the shared helpers that owned the lazy load, and a named
side-effect preload (src/test-utils/prepare-compiled-subprocesses.ts) for
suites that re-import their subject per test or whose first load happens in
production code they call. plugin-test-runtime keeps the host-capability
fixture lazy so its other consumers do not start loading that graph.
Document the rule in docs/help/testing/writing-tests.md.
335 of the 387 files now load their first declaration before collection
ends; 52 remain (48 extension tests needing an SDK preload, one package
test, three files over the line cap).
* fix(test): retain npm install fixture reuse
Preserve the existing failedSpawn reuse after merging main so the collection-time preload stays within the line-cap ratchet. Assertions and import ordering are unchanged.
* fix(test): keep memory-core facade cold-import assertion meaningful
The static subject import ran before beforeEach reset the loader mock, so the cold-import assertion passed vacuously. Restore the in-test subject imports after mock setup and preload the compiled-subprocess declaration during collection instead.
* fix: prevent delegated work from stopping silently
* fix: preserve frozen bootstrap metadata and refresh task prompts
* test: retain a foreign manager in frozen hydration proof
* test: exercise Corepack bootstrap warnings during hydration
* test: align completion fixtures with required private replies
Verify meaningful private outcomes, retained completion authority after inline waits, and one real write with no recovery replay. Start the browser fixture completion deadline when its held model is released.
* fix(node): recheck upload cancellation after opening snapshots
* test: group subagent typechecks with session ownership
2026.9.7 lengthened the compile-cache marker from <mtime>-<size> to build-<buildId>, which pushed cache paths on Windows into the directory-length window where Node 24's module.enableCompileCache() never returns (nodejs/node#66438); every OpenClaw process start could then spin at full CPU. The shared cache owner now uses a 16-character hash of the full build id as the marker, refuses Windows cache paths over 200 characters with a recorded warning (cache off for that process), and no longer propagates an unsafe inherited cache path to children. Legacy namespace handling and POSIX behavior are unchanged; native Windows/Node 24.18 completes the packaged CLI with a 220-character TEMP in 2.4 s.
Closes#162821
Keep blocked work visible in a Markdown-only card when no authorized step can proceed, rather than repeatedly saving unfinished checklist steps. Preserve the bounded completion check unchanged.\n\nRefs #162878
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: RomneyDa <6581799+RomneyDa@users.noreply.github.com>
Co-authored-by: fuller-stack-dev <263060202+fuller-stack-dev@users.noreply.github.com>
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>
Runtime preserves authored cron definitions and consumes canonical ownership, delivery, and schedules. Doctor performs supported repairs with verified backups, while update rehearsals retain live legacy sources for the live import. Preserve durable receipts, explicit scope boundaries, and deletion fences.
Published 2026.9.7 updater acceptance passed on Testbox. The separately matched native published-driver CI timeout also fails on clean main; the PR records that inherited failure and its limits.
Related: #110565, #115104, #119605
## What Problem This Solves
Fixes: a chat gets two nearly identical messages when a turn sends a progress update with `message(final=false)` and then ends without a final answer.
In our setup a group turn sent "started the run, I'll report back" as progress and stopped empty. Settled-turn finalization then made an extra model call, which restated the same progress, and both messages landed 12 s apart.
## User Impact
If the last thing a turn does is post progress to the current conversation and then it stops empty or with `NO_REPLY`, that progress message is the reply. No second model call, no duplicate. Turns that did more work after or alongside the progress, or that sent nothing, still get finalization.
## Why This Change Was Made
Provenance:
- Settled-turn finalization (#110565, closes#108738) exists so users are never left with no reply after tools ran but the model said nothing. Fallback text after an empty finalizer came later in #133520; #138645 keeps that fallback private in message-tool-only groups.
- `final:false` vs `final:true` (#105365, then the shared contract in #119605) keeps progress from ending a run early and keeps a later failure visible after progress.
- #115104 (fixing #111764) made marked progress *permit* finalization. #111764's caveat was progress sent early in a long tool run, followed by an empty final completion, which was swallowed silently. #127070 applied the same rule to queued follow-ups.
Root cause: none of these checks look at order. A progress message sent as the turn's **last** tool batch was written after the model had seen every other tool result. When the model then stops empty, the finalizer gets the same transcript and no new information, so all it can do is restate the progress. That is the duplicate. The #111764 case is different: there, tool work came after the progress, and that case is unchanged.
Fix: the decision happens when the runtime settles the attempt (`completeEmbeddedAttemptResult`), not per assistant message. The embedded subscriber only records, per provider turn (reset on `turn_start`), whether every tool that finished in that turn was a complete, non-error, non-partial `final:false` text send to the current source, and it keeps that fact for the latest turn that ran tools. At settlement, if that holds, the run ended normally (`terminal.kind === "ok"`), and the terminal assistant response stopped with empty text or `NO_REPLY`, the attempt result marks that latest progress send final and sets `sourceReplyDeliveryState: "delivered"`. Every downstream consumer (finalization gate, terminal resolution, follow-up delivery, run-entry terminal reply) then sees the same completed-reply fact an explicit `final:true` produces. No consumer special case, no prompt or tool-text change. Native-async fragments do not split a turn: an async send followed by a later tool turn, or async work and progress in the same response, still finalizes. Progress reactions, partial or errored sends, closing text, and work in the same batch keep today's behavior.
Options considered:
- Give the finalizer the sent text and accept `NO_REPLY`: this still costs a model call, depends on the model's judgment, and needs a new "empty finalizer is fine" path next to the #133520 fallback.
- Tighten the `final` tool description: models already end empty after progress, so this cannot guarantee anything.
Hermes (`agent/turn_empty_response.py`) handles this the same way. When an empty response follows visible content with only housekeeping tools after it, Hermes reuses that content as the final answer and marks it as already shown so it is not re-sent. It nudges once only when substantive tools ran after the content.
Scope: this changes only the embedded runner. Codex app-server turns can plausibly hit the same duplicate: Codex marks `final:false` in `dynamic-tools.ts` and can complete a turn with no agent message. But "nothing ran after the progress" there depends on item order across native Codex items (shell, apply_patch) that never pass through the OpenClaw dynamic-tool bridge. That makes the Codex event projector a separate owner, so it is a follow-up.
No overlap with Pash/Sarah changes: `a2bbcbf406` and #154217 touch unrelated hunks in these files, and intentional `NO_REPLY` after a delivered message stays honored.
## Bounded cost
- The new path only removes a model call: when the latest tool batch was only complete progress sends and the model then stops empty, settled finalization is skipped (0 extra calls instead of up to 2).
- Every other path is unchanged: settled finalization keeps its existing cap of 2 tool-free attempts (`MAX_EMPTY_SETTLED_FINALIZATION_ATTEMPTS`, covered by `settled-turn-finalization.test.ts`). The finalizer has no tools, so it cannot send progress and cannot re-enter this path.
- The settlement check runs once per attempt result and does not schedule retries, runs, or model calls.
- Tested: `attempt-result.test.ts` drives the real `runAgentLoop` with the real subscriber, then the real `completeEmbeddedAttemptResult` and `resolveSettledToolTerminalContinuationInstruction`. Progress-last followed by an empty terminal gives a delivered reply and no finalizer. Two cases still finalize: an async progress send, then an empty tail, then a read, then an empty terminal; and async read plus progress in one response, then an empty terminal. Each failed on the head before its fix with `expected 'delivered' to be 'missing'`. The QA scenario asserts 2 model requests (no finalization request) for progress-then-empty and 3 for the write-then-empty control.
## Evidence
QA lab, mock-openai, isolated gateway child, qa-channel group with `visibleReplies: message_tool`. New scenario `group-progress-then-empty-finalization`:
- Before (origin/main + scenario): fail. The progress room got `["Started the run, I will report back. QA-GROUP-PROGRESS-OK", "Still running, I will report back. QA-GROUP-PROGRESS-OK"]`.
- After: pass. The progress room got 1 post and made 2 model requests (no finalization request). The control room (write, then empty stop, nothing sent) still ran finalization, made 3 model requests, and posted `QA-GROUP-WRITE-FINALIZED-OK` once.
Settlement regressions in `attempt-result.test.ts` (real agent loop): on the previous per-message head, the async-continuation case fails with `expected 'delivered' to be 'missing'`. Subscriber batch-fact table: only-progress is true; same-batch work, a later batch, a reaction, and a partial/errored send are false. Focused suites pass: attempt-result (43), suppression (32), tools handler, settled-tool evidence, settled-turn finalization, attempt-execution-phase, and lifecycle.
Real gateway (qa-channel, round 3 head): the scenario passes with 1 progress post and 2 requests, and the write control finalizes with 3 requests. The scenario is the only CI-runnable check of the consumer chain after the attempt result: run-loop finalization admission, terminal resolution, payload building, and channel delivery. The loop tests stop at the delivery fact, so a consumer that ignored it would still pass them, but here it would post a second message.
### Live model, Telegram Test Server group
Setup: a real user in the QA group, `messages.groupChat.visibleReplies: message_tool`, live OpenAI `gpt-5.5` through a pass-through proxy, isolated gateway. The prompt asked the model to post a `final:false` progress update saying the report run started and then stop. In both runs below the model did that unsteered: one successful `message(final:false)`, then an empty stop. In each run the bot's transient status draft (`Working`) appears and is deleted after delivery; that is existing progress-draft behavior and the same before and after.
| Run | Bot posts that remain | Model requests (main agent) | Finalizer |
|---|---|---|---|
| Before, origin/main `7dd6ab7` | 2: `Monthly report run has started — I’ll report back here when it finishes.` (t=13.1 s), then `Done.` (t=17.7 s) | 3 | ran (`settled post-tool turn lacked a final answer`) |
| After, `e27810b` | 1: `Monthly report run started. I’ll report back here when it finishes.` (t=13.8 s) | 2 | not run |
Control on `e27810b` (finalizer must still run after real work): the model posted `Checking the clock now.` with `final:false`, then ran `exec date -u`. The proxy replaced only the model's next response with an empty turn, which is the incident's "ends empty" shape; the finalizer request itself went to the live model. The finalizer ran and the group received `The UTC time is Thu Oct 1 13:08:11 UTC 2026.`
An earlier variant that told the model to end with `NO_REPLY` produced one post on both refs: on main the live finalizer also answered `NO_REPLY` twice, and the private fallback stayed private per #138645. Main still spent 2 extra model calls there; the fix spends none.
LOC vs merge base: production +47/-3 (`src/agents` subscriber + attempt-result), docs +5/-1, tests and QA fixtures +441.
Test cost (`pnpm test <file> --maxWorkers=1`, wall, head e27810b, shared loaded host): `attempt-result.test.ts` 46.4 s for 44 tests, with the 3 new real-loop cases at 0.2 to 0.6 s each; `embedded-agent-subscribe.subscribe-embedded-agent-session.source-reply-suppression.test.ts` 53.4 s for 32 tests, with the 5 batch-fact cases at 4 to 160 ms each. Most of the wall time is import and transform. The new cases use no timers, sleeps, or gateway boots. No attributable CI timing exists yet for this head.
CI on 959c8b0 (run 36873121248, after one rerun of failed jobs):
- `gateway-timeout-recovery-subagent.e2e.test.ts`: fails 2 of 2 locally on clean origin/main. Known flake #161139, with fixture fix#162260 open. The scenario never uses the message tool.
- `src/gateway/server.chat-steer-direct-command.test.ts` (skills watchers left open): fails the same way on four unrelated PRs today: `steipete/session-store-agent-db-drain`, `chore/chrome-devtools-mcp-1.10.1`, `chore/npm-12-1-20261001`, `codex/repair-child-completion`.
- `node-worker-launch-wire.e2e.test.ts` and `telegram-model-picker-prepared-gateway.e2e.test.ts`: red on the first attempt and green on rerun; both pass locally on head and on main.
- `src/auto-reply/reply/dispatch-from-config.secrets.test.ts` "quiet group link": still red. Only the file's first case fails, at its default 1 s `vi.waitFor`; that case took 1437 ms in CI while the following cases took about 60 ms. The same three-file shard passes locally on the head, on head merged with current main, with `CI=true`, and twice on main under `taskpolicy -b`. The test drives the `secrets` tool, which never touches the message-tool progress path, the new `turn_start` reset, or attempt-result settlement. I could not reproduce it on main or find it on another PR's CI from the last two days (93 failed node jobs scanned), so this one is not proven unrelated.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* feat(tooling): add opt-in compiler performance evidence
* test(tooling): narrow compiler metrics artifact paths
* style(tooling): use braced compiler metrics guards
* chore(tooling): integrate main compiler ownership
* style(ci): format tsgo compiler version lookup
* feat(ci): integrate compiler metrics with current runner
Preserve the current prepared compiler owner, Kysely preparation, static diagnostics, signal callbacks and output joins while retaining opt-in metrics. Cover combined CI diagnostics and metrics for successful and diagnostic compiler exits.
The previous focused tests and profiles remain bound to the older composition. This integration is source-reviewed; exact-head hosted qualification is pending.
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
* refactor(agents): replace peer announcement loops with one reply delivery
Deliver delayed peer results once to the requester and return waited replies inline without a continuation. Preserve child completion custody and same-session generation-bound channel delivery. Use exact-key saved-route lookup, retire legacy peer control tokens from runtime decisions, and retain historical display suppression.
Simplify the announce owner, preserve required missing child output, and keep wake retries within the remaining operation budget. Update session guidance and the opt-in live peer scenario.
Candidate regression scenarios pass. Broad focused tests and changed-lane validation remain incomplete under host load; original-code regression proof is still pending. No live provider test was run.
* refactor(subagents): consolidate completion custody and lifecycle owners
Share terminal-effect persistence and requester-wake settlement adapters while preserving their distinct authority, frozen-wave, and publication guards. Remove unused preparation and accessor paths and reuse existing parameter contracts.
Propagate unknown SQLite write outcomes before scheduling interrupted completion recovery during restart drain. The focused candidate regression passes; original-code failure proof and the broader validation matrix remain pending.
* test(agents): align completion assertions with reply ownership
Assert that an ACP target already owned by its requester keeps task completion and starts no detached reply. Keep isolated cron fallback routing and participant-store checks while asserting no detached wait.
Remove assertions on the retired duplicate steerMessage field; retain complete source-reply content, ordering, exclusions, visibility and batch settlement checks. Remove the redundant inner same-session return while retaining the enclosing branch exit.
Validation: scoped Codex review at P2, formatting and diff checks passed. Two attempts on the coordinator Testbox lease failed during file sync before installation or tests. Baseline regression proof, the full matrix and the changed-lane gate remain pending.
* test(agents): inject the gateway caller into same-session delivery routing tests
* test(subagents): route direct requester-yield callers through prepared cron authority
* refactor(subagents): keep the requester wake currency assertion module-private
* fix(agents): fence failed and delayed subagent cleanup
Restore the previous resident row when same-ID registration persistence fails.
Carry collector cleanup authority into attachment removal, and bind lifecycle
grace timers to their captured registration and generation.
Regression coverage exercises failed writes, retirement during cleanup, and
same-ID successors without changing storage or timer durations.
* fix(agents): retain paused state in compact registry reads
Read pauseReason from existing stored JSON for ordinary and private-parent
records so cold projections preserve the resident paused status. Derive compact
SQL row types from the query and shrink the removed-assertion allowance.
No schema, serialized representation, or migration changes are required.
* fix(agents): preserve spawn cleanup and request scope
Route fork lookup failures through provisional-child cleanup. Reuse the spawn
admission owner for visible work with the resolved requester agent identity.
Pass trusted-creation transport timeouts through the argument the transport reads.
Consolidate in-process dispatch while preserving creation custody, source
fencing, per-interface deadlines, and signed fallback behavior.
* refactor(agents): simplify subagent registry and spawn owners
Remove unused registry APIs, callback declarations, presentation adapters,
registration inputs, and bootstrap scaffolding. Share registration ACK control
flow while retaining both durable writes and the original settlement error owner.
Complete keyed capability lookups, reuse heartbeat and timeout owners, carry
resolved model references, and remove obsolete ACP transcript work and relay
options. Preserve legacy stored attachment safety and protocol-v4 fallbacks.
* refactor(agents): simplify completion consumers and registry publication
Move full-registry fixture writes into test support and require named production mutations. Consolidate publication subscriptions while preserving projection ordering and persistence observer exception isolation. Remove duplicate completion retry scaffolding, runtime barrels, test-only exports, duplicate types, and single-caller adapters.
* chore(agents): shrink retired test API assertion allowance
* test(gateway): preserve real registry operations in agent fixtures
* test(subagents): preserve named publication fixture preconditions
Name all retired fixture rows when replacing describe/list facts. Warm the compact cache before its initial publication so the existing ordering assertion still detects an unintended SQL reload.
* fix(plugin-sdk): preserve capability store compatibility
Retain the public record-or-lookup store contract, including normalized first-match identity lookup and canonical depth fallback for partial records. Keep migrated internal callers on keyed lookups. Restore the base handoff declarations and requester query/caller linkage so the existing SDK declaration gate stays unchanged.
* test(gateway): publish fixture memory ownership
Publish the staged in-memory owner through commitOwnership before asserting describe and list projections. Named durable writes no longer trigger the whole-index rebuild that accidentally discovered the fixture row. Preserve the existing assertions.
* refactor(agents): tighten completion projections and ownership facts
Project only formatter-consumed child fields without mutating authoritative rows, and carry the session resolver's required requesterOwned boolean directly. Preserve formatting, nested identity references, authority checks, and public signatures while resolving the full core lint findings.
* test(agents): retire obsolete benchmark registry mock
Remove the full-registry writer no-op from the memory benchmark's typed module mock after the production writer cutover. Preserve the named-write mock and benchmark behavior.
* test(agents): await requester settlement in live fixtures
Child cleanup can publish before the requester delivery acknowledgement. Join persistence publications until the requester wake settles, preserving the existing delivery assertions and deadlines.
* refactor(agents): retire hidden sessions_send message aliases
Require the canonical message argument already declared by the public tool schema. Remove alias-only reasoning stripping and keep canonical message indentation unchanged.
* fix(agents): decode compact subagent metadata canonically
Replace the separate compact decoder with the canonical codec and session-list projector. Keep bounded SQL metadata projection, private envelopes, and duplicate-key last-value semantics without transferring retained results. Add real-reader duplicate-key and retained-content regression coverage. Schema, persisted bytes, and update behavior are unchanged.
* refactor(agents): remove redundant delivery comments
* perf(agents): restore compact reader after benchmark regression
Revert 2c0fd630c7 as required by the performance gate. Across five scoped 2000-row reads per revision, the base median was 227.31 ms and the candidate median was 429.25 ms (88.84% slower). Retain the sessions_send alias retirement and comment cleanup; leave the duplicate-key reader fix for separately approved work.
* test(agents): relocate sessions_send alias regression coverage
Keep all four rejection cases in the assembled-tool suite where retired preparation coverage freed space. The oversized sessions test file shrinks again; no assertions, mocks, timeouts, skips, or line-cap baselines change.
* fix(agents): retain sessions_send record input validation
Use the shared isRecord guard from the retired alias normalizer instead of a new type assertion. Preserve canonical missing-message errors for non-record input and satisfy the assertion safety gate without a suppression.
* refactor(agents): keep sessions_send replies on the requester key
Remove the legacy key-only DM-to-main reply remap and its routing exceptions. Use the resolved caller key for reply context, provenance, watches and followup preparation. Preserve exact-incarnation and authorization checks.
* test(agents): bind native followup custody in admission fixture
Keep the key-only DM requester authority test on the native child completion path. Settle the existing completion owner before mock acceptance and reset queued mock implementations between cases.
* fix(agents): reconcile subagent callers after main merge
Use the phase-aware publication API in newer main callers and explicit changed IDs in the prepared-read fixture. Keep retained regressions in their current sibling owners and place requester adoption in the lifecycle controller without weakening guards. Preserve heartbeat narrowing after shared-owner integration.
* test(agents): clean up merged fixture imports and names
* test(agents): refresh single-reply prompt snapshots
Regenerate prompt fixtures for the intentional sessions_send description change. The four-file delta contains only the description and derived sizes and hashes. Reproduced the hosted drift on Testbox, then passed prompt:snapshots:check and the snapshot-only changed checks.
* test(agents): await owned descendant settlement
* fix(agents): retain the leaf followup owner contract
Restore the standalone completion-owner interface and implements check from main. Deriving that contract from the implementation class creates a type-only cycle through the cohort projection. Runtime behavior and the exposed owner methods remain unchanged.
* fix(agents): disambiguate the sessions reply target resolver
Give the async sessions_send reply lookup a distinct exported name from the general synchronous outbound session resolver. Update its sole production caller and direct tests without an alias or behavior change.
* test(agents): remove the retired steer fixture argument
* test(agents): drop the retired sessions_send delivery mode from the follow-up yield live test
Reuse the extension-lint comparison-base action for PR boundary checks. Keep its bounded exact-SHA fetch and existing fallback policy, and leave checkout refs unchanged.
A depth-one merge stays an ancestry boundary even after its parent is fetched. Validate the pinned raw first parent, then compare the two trees; retain ordinary ancestry validation for other shapes. Shallow checkout regressions cover package, core and deleted public SDK changes, including blobless base inventory hydration.
* test(ui): record renderer stall evidence when Control UI e2e waits fail
Two scheduled main runs failed intermittently in comment flows without
evidence of what the renderer was doing: a pin click hung for 30 s with the
failure read missing its deadline, and a comment delete never showed up
during a 15 s poll. Neither reproduced locally.
Mocked-Gateway pages now arm a stall probe before navigation. A renderer
that misses the failure-read deadline reports main-thread busy time by kind
and its paused JavaScript stack, then resumes. A stall that ended appears as
script-attributed long animation frames. Public output keeps only bundle
paths, positions, function names, and listener tag and event names.
* test(ui): arm the stall probe before tests open CDP sessions
PR CI showed 11 mobile safe-area geometry failures: installMockGateway
attached the probe after those tests had set
Emulation.setSafeAreaInsetsOverride on their own CDP session, and Chromium
drops an earlier session's override at the next navigation once a later
session attaches. A session attached first leaves later overrides intact.
The shared suite's withPage now arms the probe right after newPage(),
before test code runs, and installation is best effort for test doubles,
closed pages, and non-Chromium contexts.
The published-driver cell timed out on scheduled main runs: the Docker fixture forced tens of thousands of OverlayFS copies on the candidate tree during the managed update. The cell now stages the update on a fresh private native volume per invocation (cleaned up on exit), which brings retention from 39 s to 12 s and the whole cell to about four minutes; checksum failures print both digests and reject before the update starts. The earlier "digest-mismatch: error" log line was an action setting echo, not a mismatch.
Related: #162391
## What Problem This Solves
Fixes two follow-ups to #162391:
- A scheduled run that hands its work to a subagent can deliver the child's `AUTOMATION_FAILED` reply verbatim and still record the run as `ok`. This happens on announce jobs, and on `delivery: none` jobs since #162474 started recording the child's answer.
- A silent (`delivery: none`) job whose agent keeps reporting a blocked task gets auto-disabled after 10 runs, and that posts an auto-disable notice even though the job has nowhere to notify.
## User Impact
- When a delegated run's settled child answer starts with `AUTOMATION_FAILED`, the run is recorded as an error with the child's explanation. Announce and current-session delivery send only the explanation, never the token. `delivery: none` runs record it without sending anything, and keep the run transcript as other failed executions do.
- A `delivery: none` job with no failure-alert route still records agent-reported failures in run history, job status, and error backoff (escalating to hourly, like any failing job). Those failures never auto-disable the job or post the auto-disable notice, so the job stays silent.
- Unchanged:
- Runtime errors on silent jobs still auto-disable at 10 as before.
- Announce and webhook jobs, and jobs with a configured failure route, still auto-disable on agent-reported failures.
- `NO_REPLY` behavior.
## Why This Change Was Made
- **Settled child answers.** #162391 classified the token in `resolveCronPayloadOutcome`, which runs on the parent's reply before #117308's descendant settlement. Two places in `dispatchCronDelivery` then adopt the settled child's reply as the run's answer: the announce settlement (#117308) and the no-delivery settlement (#162474). Neither classified it. Both now go through one adoption helper that calls the shared parser `readAutomationFailedReport`, so the run's effective terminal answer is classified after settlement. `buildDeliveryState` derives one execution-failed flag from that result and uses it for both the completion status and the transcript-cleanup decision, so a failed no-delivery run is no longer cleaned up as a quiet success. `run-finalize` consumes the result the same way it already handles `pendingPresentationWarningError`: error status, permanent classification, explanation as the summary.
- **Auto-disable.** Maintainer decision: an agent-reported blocked outcome on a job with no notification owner must not lead to auto-disable. The permanent error classification now carries `reportedByAgent`. `applyJobResult` in `timer-outcomes.ts` still counts every error in `consecutiveErrors`, so status and error backoff are unchanged. It skips only the auto-disable decision, and only when all of these hold: the failure was agent-reported, `resolveFailureAlert` finds no route, and the delivery mode is `none` (webhook jobs still auto-disable).
- Hermes handles the same protocol the same way (`[CRON_FAILURE]`): its cron hint asks the agent to report a delegated child's failure on the first line. No new tool is needed.
## Bounded cost
- No new turns, runs, waits, or notifications. Parsing happens once per settled answer inside the existing finalization.
- Agent-reported failures stay permanent, so the scheduler never retries them early.
- A silent blocked job keeps the normal error backoff (30 s, 1 min, 5 min, 15 min, then hourly), so it can never run more often than an identical job on origin/main. The only difference is that it keeps running at that backed-off cadence after the 10th failure instead of being disabled.
## Evidence
### Live model, Telegram Test Server (`telegram-e2e-userbot`, DM)
Setup: the repo runner unchanged (`run-mock-sut-user-e2e.mjs --backend mock --source-gateway --dm`, Convex lease). `E2E_MOCK_SERVER_PATH` points to a throwaway proxy that stands in for the mock provider and forwards every request unchanged to the real OpenAI API. The proxy reads the key from a private file, so the Gateway only ever holds the runner's dummy key. `E2E_ROOT_CONFIG_PATCH` selects `openai/gpt-6-astra` (`tools.toolSearch: false`). A scenario `command` step creates the jobs with `openclaw automations add` and runs them. No model output was steered: the live parent chose `sessions_spawn`, and the live child and the live silent job wrote their own `AUTOMATION_FAILED` reports because they had no shell tool. The only steer is timing in case (b): before each run, `cron.update` sets `nextRunAtMs` to now and moves `lastRunAtMs` back 2 h, so the error backoff (up to 1 h) counts as served. The Gateway's own scheduler then runs the job as a normal scheduled (not forced) run.
Before = origin/main `d8a110b1a1`, plus this PR's exact base `5c4e9aba92` for case (a). After = this head `8b77715ba4f`, with all three cases in one Gateway run.
**(a) Announce job, live parent delegates to one child, the child reports `AUTOMATION_FAILED`**
- before (`d8a110b1a15`): `cron run --wait` took 99.4 s. Run `ok`, `completionStatus: succeeded`, `deliveryStatus: delivered`. The DM received bot message 43601: `AUTOMATION_FAILED` / `The command could not run because no shell execution tool is available in this session.` The token is visible.
- before (`5c4e9aba92`, the PR base): 80.7 s. Same result: run `ok`, `delivered`, and bot message 43705 reads `AUTOMATION_FAILED` / `The command could not run because no shell execution tool is available in this session.`
- after: 63.5 s. Run `error`, `completionStatus: failed`, `deliveryStatus: delivered`. Error and summary are both the explanation. The DM received bot message 43712 with only `The command could not run because no shell execution tool is available in this session.` No token.
- Pre-existing flake, not in this diff: on both sides, the announce settlement sometimes ends with `cron child-session handoff completed without a final assistant payload` even though the live child answered. That happened on 1 of 3 origin/main-side attempts (a second `d8a110b1a1` run) and on 2 of 4 attempts on this branch (both on `bfe1f0f2169`). In those cases nothing reaches the chat. The PR's change only runs after a reply has been read, so it is not involved. The no-delivery path (case c) settled on every attempt.
**(c) Same delegated job with `delivery: none` (the #162474 path)**
- before: 55.2 s. Run `ok`, `succeeded`, `deliveryStatus: not-requested`, summary `AUTOMATION_FAILED ⏎ The command could not run because no shell execution tool is available in this session.` The token is recorded as an ok answer.
- after: 36.3 s. Run `error`, `failed`, `not-requested`. Error and summary are both `The command could not run because no shell execution tool is available in this session.` The job stays enabled. No chat message.
**(b) `delivery: none` job, no failure-alert route, the live model reports `AUTOMATION_FAILED` on every run (11 scheduled runs)**
- before: runs 1–10 were all `error` (`Could not run \`node scripts/ledger-sync.mjs\`: this session only provides a file-reading tool…`), with `consecutiveErrors` 1…10. After run 10 the job was `enabled: false` with `autoDisabled: {reason: "consecutive-failures", consecutiveErrors: 10}`, so run 11 never happened. The DM received bot message 43644 (heartbeat relay): `⚠️ Your "Nightly ledger sync" automation was automatically disabled after 10 consecutive failures. … openclaw automations enable <id>`. 10 runs took 504 s.
- after: all 11 runs were `error` with the same explanation, and `consecutiveErrors` went 1…11, which keeps feeding backoff. After run 11 the job was still `enabled: true` with no `autoDisabled`. History shows 11 `error` entries. Across the whole run, the DM received exactly one bot message: 43712 from case (a). Cases (c) and (b) sent none, including during a 120 s wait after the loop. 11 runs took 526 s.
Provider evidence (proxy log): before had 1 parent `sessions_spawn` turn, 1 child turn, 10 silent-job turns, and 1 relay turn carrying the auto-disable notice. After had 2 parent `sessions_spawn` turns, 2 child turns, 11 silent-job turns, no auto-disable relay turn, and only one routine heartbeat poll, which answered `NO_REPLY`.
Environment note: in `--source-gateway` mode on both refs, Telegram inbound polling could not start (the ingress worker's plugin capture has no `dist/plugin-sdk`). So the opening DM turn was not processed, and the jobs are owned by `agent:main:main`, the DM's owner session under the default `dmScope`. Outbound delivery and relays reached the DM normally. This is the same for before and after.
### Tests
Each new test fails on the code it fixes and passes on this head:
- `delivery-dispatch.double-announce.test.ts`, real `dispatchCronDelivery`:
- Announce settlement: a settled child reply `AUTOMATION_FAILED ⏎ …` is delivered as the explanation only and recorded as `agentReportedFailure`.
- No-delivery settlement: uses the real `waitForDescendantSubagentResult`, with the registry mocked only at its edge. The child's report is recorded as `agentReportedFailure` with the explanation as output and summary. Nothing is sent, and the `deleteAfterRun` transcript is not deleted. Classification fails without the shared adoption helper (checked on the rebased predecessor commit), and retention fails on `bfe1f0f2169`.
- `timer-outcomes.regression.test.ts`, real `applyJobResult`, 11 agent-reported failures from a clean streak:
- `delivery: none`: still enabled, streak 11, no auto-disable notice, backoff delays 5 min, 15 min, 1 h, 1 h. This fails on `ee6d9e07fdf` (no streak, no backoff).
- Webhook and announce jobs: auto-disabled at 10 with one notice, same backoff.
Measured on this head, run separately (vitest `Duration`):
- `delivery-dispatch.double-announce`: 84 tests, 22.1 s.
- `timer-outcomes.regression`: 20 tests, 15.5 s.
- `run.meta-error-status`: 14 tests, 20.0 s.
All pass. `pnpm tsgo:core` passes, and oxlint and oxfmt on the changed files are clean.
Docs: `delivery.md` documents the silent agent-failure exception to the 10-failure auto-disable. `payloads.md` notes that a delegated run's settled answer is classified the same way.
LOC vs origin/main: production +54/−25, tests +100/−1, docs +2/−2.
### Review
- Adversarial re-review round 1 (on `0530dc9f`): fixed the one real finding. Webhook jobs had been exempted too; the exemption now requires delivery mode `none`, with a webhook test case.
- Round 2 (on `ee6d9e07fdf`): NOT LAND, one finding. Skipping the increment also froze error backoff, so a silent blocked job could keep running every minute. Fixed per maintainer decision with no second counter: the streak always increments, and the exemption moved to the auto-disable decision.
- Round 3 (on `bfe1f0f2169`, rebased): LAND, no findings. The rebase onto origin/main also routes the #162474 no-delivery settlement through the same adoption helper.
- ClawSweeper on `bfe1f0f2169`: fixed both findings. (P2) A failed no-delivery child run deleted its transcript; it is now kept. (P3) The auto-disable exception is now documented. Also added per-file test timings.
- Round 4 (on `8b77715ba4f`): LAND, no findings.
- Accepted tradeoffs: the exemption is not re-derived during finalized-run startup recovery. A silent job whose streak is already 10 or more from reported failures auto-disables on its next runtime error.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(talk): persist inherited realtime settings in Doctor
* fix(talk): normalize inherited realtime input without assertions
* fix(tooling): include Talk normalizer in trusted wrapper sources
Package recovery treated checked Bun executables as Node and dropped configured macOS SQLite selection. Carry admitted runtime facts and the shared environment choice into durable standalone recovery, preserving version-1 journals, backup custody, and separately owned service restart.
Verified focused suites on Node and Bun, real macOS custom-library recovery from a clean shell, published-driver updates and Node-free interrupted recovery on AWS, static/import checks, and independent review. CI fixture and cell-lifetime repairs retain the existing assertions and product timeout policies.
* feat(core): gate decision tool prefilter behind decisionAssistance labs and harness capability (#155314, #155316)
- Connect prompt-build prefilter to canonical isDecisionAssistanceEligible Labs consent contract (#155314)
- Gate prefilter evaluation on harnessSupportsTurnScopedToolRestrictions capability declaration (#155316)
- Preserve tool policy when harness is unsupported (e.g. Codex app-server), pending action instructions are present, or attempt is cancelled
- Update test suite with schema-backed decisionAssistance configurations and per-agent decisionModel overrides
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
* feat: align conversational tool filtering with Decision assistance
Preserve the contributor core consumer while binding canonical consent, actual harness support, prepared config, and live run authority. Validate the runtime deadline and real ONNX prompt submission without claiming universal classification or latency gains.
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
* fix: disclose automatic Decision filtering in configuration help
Resolve the remaining configuration-help review finding. Explain explicit consent, inherited Decision models, built-in filtering, preserved required tools, request transfer, and hosted costs. Regenerate the config documentation baseline without changing settings or runtime behavior.
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
* feat: filter conversational follow-ups with bounded Decision context
Use one provider-neutral batch with separate missing-context and next-response tool-need judgments, following the documented Jev/Noul semantics. Preserve first-turn eligibility, safe bounded history, permissions, required tools and per-turn restoration. Measure actual foreground tool definitions with DEBUG-only scalar diagnostics.
Live model-quality testing is deferred to local testing; no classifier-specific thresholds or adapter framework are introduced.
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
* fix: restore prompt-phase CI and simplify Decision assistance copy
Bring the shared prompt-phase fixture and Tool Search assembly override up to the current dispatch diagnostics contract. Type the assembly mock and complete policy result so missing fields fail checking instead of silently blocking the primary stream. Preserve the original four replay/queued-context assertions and cover DEBUG-on/off diagnostics without adding cases.
Construct modified Decision answers on fixture-owned copies rather than mutating the read-only API result. Keep the Labs toggle and configuration help generic; supported uses and full behavior/privacy details remain in canonical docs. No provider, saved-consent, permission, or classifier-policy change.
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
* fix: preserve prompt hook context in tool prefilter
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
* fix: preserve tool and provider availability through Decision filtering
Revalidate Decision eligibility at the foreground dispatch boundary and withdraw only its optional cap when the published opt-in or model selection changes. Restore the current permitted schemas, catalog, callability and owned prompt guidance without discarding independent hook restrictions or required tools.
Keep caller-budget expiry out of shared provider outage accounting while preserving genuine provider failures and physical request settlement. Replace incomplete context-engine hook fixtures with the typed real runner, retain the original assertions, and keep the context limit private for the production dead-code gate.
Preserve the hook-context and rubric-9 change; document and test the separate 6000/8000 UTF-16 evidence-text bounds and the existing forward-looking saved-intent contract. Add failure-first composed final-wire and real loopback HTTP/provider-host regressions, including subsequent explicit evaluation on the same provider.
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
* fix: restore foreground result-budget fixture typecheck
Model the result-budget fixture's foreground-only session explicitly as not compacting. Keep the required production compaction contract and all existing result/persistence assertions intact.
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
* test: consolidate Decision assistance boundary coverage
Remove overlapping test layers and repeated fixtures while retaining the distinct context, policy, and dispatch boundaries. Runtime validation remains pending because the secretless runner broker is unavailable.
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
* fix: stop automatic Decision calls after opt-out
Address MertBasar0’s pre-dispatch opt-out report by reusing the automatic consumer eligibility fence after awaited operator preparation. Preserve explicit Decision calls and root/scoped authority and cancellation checks. Extend the existing composed regression at the operator boundary.
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
* fix: stop Decision evidence after admission is revoked
Carry automatic consumer admission through the existing provider host and TypeSafe adapter to the final synchronous guarded-fetch boundary. Retain observed revocation or authority failure across adapter sanitization and asynchronous cleanup without treating either as a provider outage. Preserve explicit Decision evaluation, caller cancellation, deadlines, and provider retirement.
Rebase the contributor commits onto pinned main and retain their previously published integration fixes and test organization. Prove the cold-load and transport-preparation windows with the registered TypeSafe plugin and a real loopback request counter, including same-host explicit calls after revocations.
Co-authored-by: MertBasar0 <mertbasar0@hotmail.com>
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
* fix(agents): make decision assistance opt-out admission-only
Keep Labs eligibility at provider dispatch while preserving independent live guards for admitted evaluations. Let in-flight results survive opt-out and cover subsequent disabled turns.
* test: close skills watchers in Discord capture fixture
---------
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
Co-authored-by: MertBasar0 <mertbasar0@hotmail.com>
Related: #130753, #137832, #147969
## What Problem This Solves
Fixes: some scheduled jobs created by an agent fail for months because a tool list saved by an older OpenClaw build is missing tools the creator actually had, such as the native shell. In our setup, a monthly group job that runs `node <script>` delivered nothing in August, delivered nothing in September (the run still reported `ok`), and posted a blocker in October. Its saved list had 31 tools and no `exec`.
## User Impact
User impact: an agent-created agent-turn job that does not name specific tools now gets the same tools as its owner conversation at run time, like a job an operator creates without `--tools`. Existing jobs with an automatically saved creator snapshot behave the same way from their next run. Nothing stored is rewritten: no migration and no backups. Explicit tool lists, script payloads, condition triggers, and jobs bound to captured Codex app authority keep their stored list.
Tradeoff, approved by the maintainer (Ayaan): a per-sender tool policy on the creating owner, or a plugin hook that narrowed the creating turn, no longer limits these default jobs. Only owners can create automations from chat, and subagents cannot create them.
## Why This Change Was Made
**History of the saved list.**
- #91499 introduced it so a delayed run cannot do more than its creator could.
- #112483 made every agent-created job store one, because runs have no sender.
- #112661 made scheduled runs re-apply the owner session's group policy and every non-sender limit, keeping the stored list as the upper bound.
- #137832 fixed native tool capture for new jobs only, and deliberately did not widen stored lists.
- #147969 added a Doctor advisory. It only fires for claude-cli, so it never covered Codex-harness or built-in OpenAI jobs like ours.
**Root cause.** When no tool list was given, OpenClaw saved a frozen copy of the creating turn's tools instead of treating the job like an operator `*` job. Every capture bug (missing native tools, late configured MCP, renamed tools) then stayed in the job permanently.
**Fix.** This follows Hermes, which keeps no creator snapshot: `cron/scheduler.py` `_resolve_cron_enabled_toolsets` reads toolsets from config at run time.
- **New jobs.** An agent-turn create or update with no list, or `*`, stores `["*"]`. That is the same value operator jobs store, so the job's tools match a normal turn in its owner conversation. Script payloads and condition triggers still store the creator's concrete tools, because a script reaches MCP only through servers its list names. Jobs whose creator captured Codex app authority also keep the concrete list, because that authority is bound to it.
- **Existing jobs.** One helper, `resolveCronRunToolsAllow` in `src/cron/tools-allow.ts`: a stored automatic snapshot (`toolsAllowIsDefault`) runs as `*` when it has a valid scheduled owner policy, no condition trigger, and no Codex app authority. Otherwise it keeps its stored list. Every execution consumer of the stored list uses it: the run payload, the command-prompt preflight, and the scheduled message authority.
- **Script transitions.** A `*` job that becomes a script, or gains a condition trigger, captures the creator's concrete tools.
- **Exec pin.** A `*` list keeps the creator's exec host pin.
- **No new noise:** automatic snapshots stay excluded from the `web_search` provider warning, as on main.
- **Deleted, now pointless:** both Doctor advisories about incomplete automatic snapshots, the run warning about pre-MCP snapshots, and two exports nothing uses anymore.
Review note: on claude-cli, a `*` job runs without a CLI tool cap, so Claude's native tools behave exactly as in a normal chat turn in that conversation. This PR introduces no new path around `tools.deny` that a chat turn doesn't already have.
## Evidence
Live-model Telegram proof (Telegram Test Server DM, leased team credential, live `openai/gpt-6-astra` reached through a forwarding proxy that stands in for the runner's mock provider; the runner harness itself is unchanged). This reproduces the shape of the original incident:
- The tester DMs the bot, which creates the owner conversation.
- A job owned by that conversation is added. Its stored list is an old-style automatic snapshot `["automations","message","read"]` plus `toolsAllowIsDefault: true`, with no `exec`.
- The payload is `Run: node scripts/split-report.mjs and post its output line verbatim`. The workspace script prints a random nonce.
- The job is run once (`cron run --wait`), with announce delivery to the DM.
| Build | `exec` offered | Model action | What arrived in the DM | Run |
|---|---|---|---|---|
| base 94f5a8d (main before this PR) | no | `tool_search` ×2, then gave up | "Could not run node scripts/split-report.mjs: no command-execution tool is available…" | error |
| **this PR, head 5b78cb7** | **yes** | `exec {"command":"node scripts/split-report.mjs"}` | "**SPLIT-REPORT 93C53909**: general 41, design 17, ops 9" (the exact script output, with this run's random nonce) | ok, delivered |
| head 5b78cb7 with `tools.deny: ["exec"]` | no | `tool_search`, `read`, then gave up | "Could not run node scripts/split-report.mjs: no command-execution tool or paired node is available…" | error |
In every run, the stored job kept `["automations","message","read"]` plus the marker. Before and after use the same scenario and driver; only the checkout differs.
Update and live proof: published `openclaw@2026.9.7`, then this branch at the exact head (2f5099d), on the same state directory. Mock provider. Every process ran under a temporary `HOME` and state directory. Each job's message makes the model call `exec` with `touch <effects>/<job>`.
1. 2026.9.7 created both jobs through `cron.add` (scheduled policy `trusted`). With the Gateway stopped, the "stale" job was given the old automatic-snapshot shape `["automations","message","read"]` plus `toolsAllowIsDefault: true`. sha256 of both stored rows: `609358837…`.
2. Runs:
| Build / config | Job | `exec` offered | Side effect | Run |
|---|---|---|---|---|
| 2026.9.7 | stale automatic snapshot | no | absent | error |
| 2026.9.7 | explicit `["read","message"]` | no | absent | error |
| this branch | stale automatic snapshot | **yes** | **created** | ok |
| this branch | explicit `["read","message"]` | no | absent | error |
| this branch, owner policy narrowed to `tools.deny: ["exec"]` | stale automatic snapshot | no | **absent** | error |
| this branch, `tools.deny: ["exec"]` | explicit `["read","message"]` | no | absent | error |
3. After the branch runs, the stored rows were byte-identical (same sha256 `609358837…`), and job ids and lists were unchanged. Nothing was migrated.
An earlier run at e151ea3, with the same harness, also covered a snapshot bound to Codex app authority: `exec` was not offered, the file stayed absent, and the stored row was unchanged.
Tests:
- `run.tools-allow.test.ts`: a stored automatic snapshot `["message","read"]` reaches the embedded run as `["*"]`, with the owner's scheduled policy intact. It fails on main with `["message","read"]`.
- `cron-tool-creator-cap.test.ts`: a default agent turn stores `["*"]`, while a trigger script and a Codex-app creator keep the concrete snapshot.
- `run.tools-allow.test.ts`: snapshots without a valid owner policy, or behind a condition trigger, keep their list. Both cases fail on the previous head.
- `run.tools-allow.test.ts`: no `web_search` warning for an automatic snapshot that kept its list. This fails without the exclusion.
- `run.message-tool-policy.test.ts`: a self-edited automatic snapshot runs on CLI with no cap.
- `run.tools-allow.test.ts`: a legacy `Command to run:` prompt from an automatic snapshot without shell tools now runs instead of being rejected.
- `jobs-tool-policy.test.ts`: scheduled message authority is admitted for an automatic snapshot that lacked `message`.
- `cron-tool-creator-cap.test.ts`: a `*` agent turn converted to a script captures the creator's concrete tools.
- These three regressions fail on the previous head. `node scripts/check-changed.mjs` passes.
- Explicit-list, exec-pin and gateway creator-transport suites pass. `pnpm tsgo:core` passes.
## Bounded cost
No new path triggers a model call or a job run. The change only selects which tool list an already scheduled run uses.
LOC vs main: production +98/-250 (net -152), tests +140/-373, docs +19/-11.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(update): await progress receipts before advancing commands
* fix(update): preserve Doctor refusal and progress worker boundaries
Keep the execution guard contract in the existing parameter types module
without an implementation import cycle. Join actual reporting failures with
settled Doctor errors while preserving later typed authority refusals.
Run the fresh-package staging regression in the worker-capable harness now
that its real progress callbacks use the shared-state writer. Preserve both
replay cases and their file and backup assertions.
* test(update): await candidate progress receipt settlement
* fix(update): retain recovery after progress receipt refusal
Seven channel plugins hand-wrote the same secret-contract setup, and
Zalouser duplicated migration and prefix parsing that shared helpers
already own. A new Plugin SDK helper, createChannelSecretContract on
openclaw/plugin-sdk/channel-secret-basic-runtime, builds the contract
from each channel's declaration, and Zalouser uses the existing helpers.
Feishu and Matrix stay out because other work is changing them.
The public SDK surface budget grows by exactly this one export (public
exports and callable exports each +1), approved by Peter on 2026-10-01.
clickclack, googlechat, irc, slack, sms and zalo now require OpenClaw
2026.9.8 because they call it; published plugins already require a
matching host, and installs fall back to the newest compatible version.
Registration and secret-contract output are identical to main across
eight channels; 6,619 channel tests pass. Production -110 net.
Release note: The clickclack, Google Chat, IRC, Slack, SMS and Zalo
plugins require OpenClaw 2026.9.8 or newer.
Fix Bun message chunking and terminal truncation selecting the preceding grapheme when JSC containing() probes an emoji high surrogate. Route all seven lookups through normalization-core, normalize numeric indexes before the offset, and retain the new eager dependency in the native PR wrapper inventory. Upstream engine fix: oven-sh/WebKit#753.
Node 24 and fork Bun each pass 586 tests (one skipped) across the requested 39 files. Node matches native containing() for all 1,236 corpus lookups. Changed-file checks, import-cycle validation, wrapper closure checks, and independent P2 reviews pass.
The remaining CI cron copy-fault and Discord skills-watcher teardown failures reproduce on clean main and use the authorized native pre-existing-failure exception. The original wrapper inventory regression was fixed.
Preserve the existing maintenance shutdown owner and cover live inspection-child closure plus shared and agent handles through the real Doctor health flow. The suspected self-handle leak was not reproduced; distinguish fuser holder PIDs from inspection failures without relaxing refusal.
Includes the independently landed WAL inventory rescan for integration. Validation: 27 focused cases on Blacksmith Testbox, production and all 27 test type graphs, scoped lint and repository guards; isolated Codex review found no actionable P0-P2 findings. Commit hooks skipped after remote validation to honor the Mac host restriction.