* fix(build): keep pinned pnpm probes from dirtying the lockfile
* feat(skills): search bounded installed instruction bodies
Index bounded instruction text lazily through the admitted catalog reader, retaining metadata-only search for unavailable or over-budget bodies. Revalidate reader authority on cache hits and preserve whole-body skill reads. Report index coverage in search results.
Validation: 33 focused tests across four files passed on the native proof checkout in 117.62s; fresh scoped autoreview clean through P2; all eight changed files pass oxfmt. The coding worktree is deliberately code-only, so formatting ran through the qualified independent tooling checkout before this commit.
* fix(skills): require current native read authority for body search
* test(skills): prove disk-backed search authority through dispatch
* test(skills): tie authority gates to test cancellation
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
Run registered setup migrations, Doctor contracts, legacy channel previews and interactive channel compatibility hooks through one executor. Preserve migration phases, public registration and adapter receivers while isolating failed candidates and retaining warning-only results.
Regression failures were reproduced on the base; focused and plugin-contract tests, changed checks, both cycle checks, SDK API comparison and independent reviews completed successfully.
Bun-hosted published updates could hang after candidate Doctor completed because the spawn broker retained an idle IPC channel. Honor optional channel reference controls while retaining live requests, native resource claims, and shutdown work.
Published openclaw@2026.9.7 updates passed on both Bun builds and a Node 24 control, including service replacement, authenticated readiness, and preserved backup bytes. The 22 focused broker cases pass on Node 24 and the prescribed Bun fork after a patch-identical rebase. The installed updater, delegated Doctor markers, schemas, and rollback contract are unchanged.
CI exception: both prebuilt UI global setups failed on inherited stale pnpm lockfile metadata before tests; independent main reproduction confirmed the cause, and main fixed it in 4b32b40150. Fail-fast cancellations remain unrun coverage. Landed through the authorized native pre-existing-failure route.
* chore(deps): update fs-safe to 0.22.0
* test: preserve realpath encoding in Windows path fixture
* test: close the plugin lifecycle fixture database
* test(memory): hold writer admission without reopening resources
* refactor(state): resolve Gateway profiles in state workers
Move the remaining Gateway email acquisition and canonical profile selection
callers onto existing state worker operations. Retain request and profile
authority across waits and recheck it before writes, network effects, and
metadata publication. Remove the synchronous Gateway access helpers.
Keep schemas, stored data, permissions, and update behavior unchanged.
Remaining profile reads/setters and account stores follow in a second slice.
Isolate session-catalog fixture writes from background reclamation so request
cleanup assertions remain deterministic across the worker wait.
* test(gateway): assert snapshot rejection at handler boundary
Profile acquisition now rejects discovery snapshots before model catalog loading. Preserve zero credential and HTTP effects while testing the direct handler rejection; WebSocket dispatch owns the error response.
* test(gateway): await concurrent metadata reader admission
Synchronize the concurrency fixture on all reader callbacks entering after asynchronous profile preparation. Preserve the original unrelated-creation, response, and cleanup assertions without timers or polling.
* refactor(state): register worker operations once per domain
Infer shared-state worker command contracts from lazy per-domain handler tables. Migrate Web Push, APNs, worktree registry operations, and fleet registry while preserving the existing broker and transaction owners.
* fix(gateway): retain one catalog acquisition authority check
Keep the shared post-await authority assertion before catalog effects and the final publication check, without duplicate session path validation. Migrate authorization fixtures to the async profile owner contracts and remove their retired resolver mock. Existing real Gateway freshness budgets and permission assertions remain unchanged.
* refactor(state): run profile setters in state workers
Move profile writer kernels behind the typed worker registry and route display-name and avatar RPCs through admitted worker writes. Recheck requester authority before refreshing connected profiles, use current committed display facts, and preserve committed self-merge responses. Keep role-policy reads as follow-up migration work; schemas and update behavior are unchanged.
* refactor(state): let profile handlers own kernel dispatch
* test(gateway): assemble synthetic question URL credentials
* test(gateway): synchronize concurrent project discovery
Fixes#158894.
OpenAI image generation over the Codex OAuth route no longer fails outright when the account cannot use the default Responses chat model. On a confirmed model-unavailable 400 (structured HTTP 400 with the exact evidenced error detail), it retries with the configured OpenAI chat models. Other 400s do not retry, and a working default is unchanged. The match lives in the OpenAI plugin; no public SDK surface is added.
Proof: fake OAuth profile, built CLI and a localhost Responses endpoint. main produces no image; this PR retries the configured fallback model and produces a PNG. Regression tests (including request-id-suffixed errors through the real HTTP normalizer) fail before and pass after.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
## What Problem This Solves
Fixes: a scheduled agent run that could not do its task is recorded as `ok`, so run history, `--wait`, and the configured failure alerts treat it as a success.
In our setup, a monthly isolated job delivered nothing for two months and every receipt said `ok`. One month the turn knew it was blocked (no shell tool) and said so in plain text, but nothing structured recorded that.
## User Impact
- A run can now fail itself. A final reply whose first line is exactly `AUTOMATION_FAILED` records the run as `error`, with the lines that follow as the error text. The existing failure handling then applies: error backoff, plus failure alerts and owner-conversation repair for jobs that already have a failure route or `failureAlert` config.
- Announce jobs still deliver the agent's explanation, with the token removed.
- These stay ok: `NO_REPLY` runs, and replies that mention the token anywhere other than as the exact first line.
- **No change to notification contracts.** An unconfigured `delivery: none` job still sends nothing to chat (no alert, no repair request). It now records the error in run history and job status (`lastRunStatus`, `consecutiveErrors`).
**Not fixed here:**
- A delegated (spawn-only) `delivery: none` run still ends `ok` and drops its child's result. The parent never sees the child's reply, so it can't report that failure with the token. Since #135318, cron turns also can't `sessions_yield`. A separate PR fixes settlement at #117308's spawn-only path.
- Notifying owners of failing `delivery: none` jobs is a pending product decision and is not part of this PR.
## Why This Change Was Made
Provenance:
- #110949 added the unattended-run preamble. It tells the model "If something failed, state plainly what failed … the scheduler owns retries and failure alerts", but no structured fact carried that statement. The scheduler only fails runs on run-level errors, error payloads, or exec-denied tool errors (`failure-signal.ts`, `execution_denied`).
- #126164 and #131228 deliberately keep execution status separate from delivery outcome. This PR keeps that separation: the token is an execution outcome the agent reports. It is parsed in the existing cron outcome owner (`resolveCronPayloadOutcome`), next to `failureSignal`, and is never derived from delivery.
The token is a strict control line, like `NO_REPLY`. Only an exact first line counts, so ordinary prose is never classified. Hermes Agent uses the same mechanism for cron (`[CRON_FAILURE]` on the first line; `cron/scheduler_prompt.py`, `cron/scheduler.py`). The preamble sentence is reworded in place: similar length, still static for prompt caching.
## Bounded cost
New or changed paths that could lead to a model call or a job run:
- **The token parse.** No new turns, runs, or notifications. It only changes the recorded status of the run that produced the reply. It is parsed only in cron finalization for that run. The delivered explanation is outbound text, never inbound, so there is no self-trigger.
- **Error status → scheduler.**
- One-shot jobs: an agent-reported failure carries `errorClassification: { kind: "permanent" }`, so it never enters the transient-retry path. This also means wording like "network timeout" in the explanation is never text-classified into a retry or an error reason. **Tested** in `run.meta-error-status.test.ts`.
- Recurring jobs: the existing error backoff only delays the next run (natural schedule or later), and the existing auto-disable stops the job after 10 consecutive errors. Nothing runs sooner than scheduled.
- **Repair turns.** Unchanged. They still happen only for jobs with an existing chat failure route (#161007): at most one per failure incident, and the incident marker is persisted in job state. A repair turn can't start runs, because automation control is owner-only and the request carries no sender identity. It is an ordinary conversation turn, not a cron run, so its own failure can't open another repair.
- **The preamble text.** Static; it adds no turns.
## Evidence
Isolated Gateway, qa-lab mock-openai, new scenario `qa/scenarios/scheduling/cron-blocked-outcome.yaml`:
- **Before** (origin/main): a `delivery: none` job whose reply starts with `AUTOMATION_FAILED` → `status: "ok"`, `completionStatus: "succeeded"`, and the summary holds the blocker text.
- **After:** see the result comment below. The scenario checks:
- An owned `delivery: none` job, run 3 times → each run is `error` with the reported reason; no repair request and no chat message.
- An unowned announce job, run twice → `error`; the reason is delivered without the token, and the existing `failed 2 times` alert fires.
- A mid-text quote, and `NO_REPLY` (with and without announce) → `ok`.
Focused tests (fail before, pass after): `src/cron/isolated-agent/run.meta-error-status.test.ts` covers token → `error` + permanent classification, and a mid-text quote → `ok` (two cases, about 1.7 s each in the shared harness). Also passing: `failure-alerts.{account-routing,persistence,recovery}`, `service.failure-repair`, and `run.message-tool-policy` (116 tests). `pnpm tsgo:core` and oxlint on the changed files are clean. `run.message-tool-policy.test.ts` no longer pins the full preamble wording; it keeps the ordering assertion.
LOC vs origin/main: production +36/−7, unit tests +24/−2, QA scenario and mock about +180, docs +1/−1.
## Upgrade compatibility
No migration is required. The diff touches no schema, store, codec or persisted type file. Run history keeps the existing `CronRunStatus` values (`ok | error | skipped`) and the existing error/permanent classification fields, so rows written before and after this change go through the same unchanged readers. The QA scenario's "repair" wording refers to the scheduler's automation repair request, not a data repair. Existing-state compatibility is verified by the unchanged `failure-alerts.persistence` and `service.failure-repair` suites passing on this head.
Select changed extension packages and consumers of changed public source through the existing import graph, including type-only imports. Keep the complete typed programs, native lint chunks, resource limits, artifact preparation, and non-extension checks.
Retain full extension lint for scheduled/hourly main and release validation, shared lint/type policy, uncertain prior source types, and the OPENCLAW_CI_EXTENSION_LINT_FULL repository switch. Read the pinned diff base so removing a global augmentation cannot hide its consumers. Publish selected package reasons in the check-plan summary.
Twenty recent PR path scenarios emit 54 -> 40 hosted extension-lint rows (25.9% fewer): four 3 -> 0, one 3 -> 1, two already 0 -> 0, and thirteen unchanged. On Linux with four CPUs and a 16 GiB limit, warm maximum full-stripe compute was 64.44 s and the qa-lab-only stripe was 14.58 s. Shared artifact preparation took 114.24 s separately; these are not end-to-end Actions job timings. Full typed programs are unchanged. Shared loose extension consumers still retain full coverage.
The bounded historical search found one source-attributed extension-lint failure among 42 inspected runs; all three failing packages remain covered. Ten causal runs were unavailable, so this is limited historical evidence.
Proof on Linux Testboxes: whole tooling plus both owning fast configs; full core/extension/root test types; final check-changed, typed lint, boundary guards, architecture, source contracts and dead exports; eight same/different-SHA native preflight cells; native before/after manifests for twenty PR scenarios; final helper parity for all twenty plus two probes; and P2 review. Initial new-fixture errors were corrected and replayed. The sole inherited suppression-inventory failure reproduces on the unmodified parent and passes with already-landed d346309979. No workflow dispatches or CI reruns. New helper tests cost 5.25 s locally, with their Linux replay included in the focused proof.
## What Problem This Solves
Fixes: Web UI (`chat.send`) turns that stall before replying still end with "⚠️ This turn was interrupted because it stopped making progress. Please try again." instead of the automatic recovery that channel DMs get since #161069.
## User Impact
- When a Web UI turn stalls before replying, the user now gets the recovered answer instead of the retry notice. The stalled run ends as aborted, and the answer arrives as its own run in the same chat, the same way any queued Web UI follow-up arrives. If the recovery also stalls, the existing notice is sent once.
- The recovery keeps its normal tools, except that Skill Workshop can't write to the skill library during a recovery. It answers "send a fresh message" instead.
- Group-thread participant turns keep the notice. That is now the only exception, and it's stated in the docs.
## Why This Change Was Made
#161069 kept the notice for Web UI turns because the recovery run had no owner that would deliver its answer ([maintainer decision](https://github.com/openclaw/openclaw/pull/161069#issuecomment-5908163404)). Provenance of that owner:
- #116649 binds every queued `chat.send` reply to its originating Gateway admission, so a late answer never goes out through a later same-session dispatcher.
- #129001 added `createSourceRetry`, which gives a front-requeued run of the same source a fresh owner, and the `pending` state: a source that never handed its request to the queue keeps its own terminal, so a finished run gets no duplicate or late reply.
- The recovery run inherited the stalled turn's still-`pending` owner, so its answer was dropped (`terminal-not-recorded`).
The fix:
- **The stalled turn hands its request to the recovery through the source's existing `createSourceRetry`.** A `pending` source that hands off is settled, so the original can never deliver once the retry exists, and the recovery owns the only delivery. The relaxed `pending` case is reachable only through the stall continuation, which runs after the stale watchdog aborted the source before any output. `createSourceRetry` has two callers:
- the follow-up source retry in `followup-delivery.ts`, which only sees owners that already queued (`deliver`/`drop`);
- the stall continuation.
- **The #161069 early return is narrowed.** Only a source that explicitly declares no owner (`drop`) keeps the notice.
- **Library authoring is not carried into the recovery.** Every authenticated Web UI human turn mints a library-authoring grant (`prepareGatewaySkillAuthoring`), and every run binds it at admission. By design (#134068) the grant refuses a replacement run, and the recovery is a new run, so it doesn't inherit the grant. Before this change the recovery failed with "Personal authoring cannot move to a replacement run" instead of answering.
- Rule: a recovery run never inherits skill-library authoring. Whatever the original grant's target, the recovery gets an expired personal grant. The Workshop tool can't fall back to the workspace tool (`createSkillWorkshopTool`), and every call returns `AUTHORITY_EXPIRED` asking for a fresh message. All other tools stay available, and no access is widened.
- Alternatives considered:
- The existing per-run `conversationToolPolicy` replaces the configured group policy instead of adding to it, so using it as a deny list could widen access in group sessions.
- Minting a fresh grant would relax #134068's replacement-run guard.
Hermes (sibling framework) returns its timeout result through each turn's own reply path on every surface. This PR follows that idea of one reply owner per source. Hermes has no auto-continuation to borrow.
### Group threads: left for a follow-up
Group-thread participant turns keep the notice. Participant reply options set `onQueuedFollowupReplyBatch: undefined` (#141411), which #116649's `in` check turns into a `drop` owner, so every queued participant reply is dropped. Omitting that key makes the recovery reach the group: in a qa-channel run the group got `STALLED-TURN-RECOVERED-OK` instead of the notice. The same run showed a second group message, arriving while the recovery ran, ending `skipped:reply-operation-active` with no reply. Participant reply options also clear `turnAdoptionLifecycle`, so a participant turn on a busy session never reaches queue policy (`allowGatewayQueueResolution` in `dispatch-from-config.lifecycle.ts`). Fixing that needs its own change to participant admission, so this PR doesn't enable group recovery.
## Bounded cost
- **Recovery runs at most once per stalled turn.** `continueStalledTurn` is armed only for direct turns and called once from the aborted-dispatch finisher. The recovery runs on the follow-up runner, which never arms it, so if the recovery also stalls the existing notice is sent and nothing else starts. Tested: "sends the Web UI notice, not a second recovery, when the recovery also stalls" (2 model turns total, empty queue).
- **One delivery per handed-off source.** After `createSourceRetry` from `pending`, the original owner is settled and refuses both delivery and further retries. Tested: "gives a pending source's only delivery to the retry it hands off to".
- **No recovery where its answer would be dropped.** Tested: "leaves the notice with a group-thread participant whose source declares no reply owner".
- **No wider tool access.** Tested: "never lets a personal/workspace-target recovery author skills while keeping its other tools". No tools are disabled.
## Evidence
- **Real Gateway `chat.send`.** A QA Gateway child, mock-openai with the stalled-turn fixture, `QA_DIAGNOSTIC_STUCK_SESSION_ABORT_MS=30000`, and an operator `GatewayClient` (TUI pattern) recording `chat` broadcasts:
- Before (`f525f46`): `+56.4s state=error "⚠️ This turn was interrupted because it stopped making progress. Please try again."`, then `aborted`. 5 mock requests, none guided.
- After: `+58.6s aborted` for the stalled run, then `+59.5s state=final` under the recovery's own run id with the recovered answer. 6 mock requests: 5 stalled, 1 guided success. The answer is persisted in `chat.history`.
- An intermediate run without the authoring change: the recovery was delivered through the new owner but failed with the authoring error. That run is what found the issue.
- **Unit tests:**
- The two Web UI tests in `agent-runner.stalled-turn.test.ts` fail on `main` (`continueStalledTurn` returned `false`) and pass now.
- The owner, drop and authoring tests pass.
- Wall times: `agent-runner.stalled-turn.test.ts` + `followup-delivery.test.ts`: 51 s (75 tests; plus `skill-workshop-tool.test.ts` 19 tests). `chat-send-late-followup.test.ts`: 2 s (6 tests).
- `node scripts/check-changed.mjs` passes (typecheck, lint, guards).
- **Head `c1bca05` (delivery path unchanged since):** the stalled run ends `aborted` at +59.1s, then `state=final` with the recovered answer under the recovery's run id at +60.1s. 6 mock requests, 1 guided. The answer is persisted in `chat.history`.
LOC vs `origin/main`: production +36/−6, tests +130/−5, docs +1/−1.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
Retire the five SDK compatibility facades under the approved September 30 owner decision, and migrate in-repository callers onto their focused contracts. Keep implementations with their canonical owners and remove forwarding exports that become unused after the cutover.
BREAKING CHANGE: remove openclaw/plugin-sdk/channel-lifecycle, channel-message, channel-reply-pipeline, config-runtime, and infra-runtime. Use channel-outbound/channel-inbound, config-contracts and focused configuration/infra entrypoints. The new system-event-runtime entrypoint supplies public snapshot inspection and consumption. Update affected external plugins before upgrading the host; this retirement does not certify universal external migration.
Package exports, SDK entry inventories, compatibility tombstones, migration docs, and surface budgets move together. Canonical API comparison confirms exactly five removed entrypoints, 869 removed export paths, 35 focused additions, and no retained-export signature changes. Net production reduction: 1,320 lines.
* feat(agents): expose installed skill search across harnesses
Keep installed-skill eligibility host-owned while applying normal tool policy to discovery and complete instruction reads. Preserve current main's prepared tool surface and consolidated Code Mode coverage. Caller-provided snapshot options supply read paths, not replacement eligibility authority.
* test(agents): preserve hoisted skill harness initialization
Keep mock factory dependencies owned by vi.hoisted when splitting skill fixtures. Assemble skill-specific mocks only when the test consumer requests them. Update two prior expectations for the prepared empty snapshot/resource contracts while preserving sandbox host-skill exclusion.
Exact-head reproduction failed on the original hoisting and resource-array assertions. Nine affected suites covered 61 tests; after the hoisting repair, a stale sandbox snapshot assertion was found and the 32-test owner file passed on focused retry. Fresh lifecycle autoreview found no actionable P0-P2 issues; formatting and whitespace checks pass.
* test(agents): group sandbox skill coverage with policy tests
* test(agents): admit skill tools in client collision coverage
* test(agents): provide complete prepared skill fixtures
* test(agents): match async skill prompt contract
* fix(skills): preserve existing Code Mode read grants and whole reads
* fix(skills): preserve read denials in frozen tool profiles
* fix(codex): refuse partial installed skill instructions
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
* fix(agents): recognize registered message waits after yield
Carry registry acceptance of an explicit message wait through the yield callback and current attempt into continuation evidence. Avoid a synthetic missing-continuation reply for registered waits while preserving diagnostics for refused or absent registration.
* test(agents): align yield callback expectations
* test(gateway): include yield registration evidence in callback assertion
The update database generation check compared physical fingerprints, so a SQLite checkpoint that merely folded already-captured WAL frames into the main file (on Windows, triggered by the exclusive removal probe) was treated as a later write and rollback was refused with "Databases changed after snapshot capture"; the Windows lifecycle test failed in isolation. The check now compares committed pages plus retained WAL commit evidence: checkpoint-only transitions and exact reversals that leave byte-identical content are admitted (identical content cannot lose data on restore; recorded at the comparer with a regression), while any newer data is still refused. The fixture's unsafe in-process snapshot mock is removed.
Refs #158163
## What Problem This Solves
Fixes: when an agent runs an automation with `automations run`, the call returns as soon as the run is queued, so the model never learns whether the run worked. It then posts guesses, or schedules a separate "verify" job. That job runs with only its own scope, can't see the result, and reports a false "blocked".
## User Impact
User impact: when an agent runs an automation, it gets the finished run (status, error, delivery status and summary) in the same tool result if the run ends within the call's `timeoutMs` (default 60 s, capped at 10 min). A longer run returns its `runId` and a note to check it later with `runs jobId runId`, not with a scheduled check. The CLI is unchanged.
## Why This Change Was Made
`cron.run` became enqueue-only in #40204 (closing #40192). Back then a synchronous run held the RPC until the 60 s gateway timeout and returned a misleading timeout error while the job kept running. #81929 later added `openclaw cron run --wait`, polling `cron.runs` from the client, and deliberately kept `cron.run` from holding the request open. The agent tool never got a wait at all. A client-side poller would not work for the tool anyway: management-grant turns can run a job, but `cron.runs` is not one of their allowed methods.
This PR keeps that invariant (no unbounded request, no false failure) and gives the one owner of run completion a bounded, opt-in wait:
- The cron service records each accepted manual run's settlement. `waitForManualRun(runId, timeoutMs, signal)` resolves when the run has written its history row. It is event-driven, with no polling.
- `cron.run` accepts an optional `waitTimeoutMs`. When the run finishes in time, the ack also carries `run`, read through the existing `cron.runs` handler so the same visibility and authority rules apply. When the wait ends first, the ack is unchanged and the run continues.
- When an agent turn runs a main-session job, or a `session:<key>` job in its own session, the call returns at once. The caller's turn holds the main lane and its own session lane: heartbeats skip while `CommandLane.Main` is busy, and a main-session run waits up to 2 min for that busy turn. Waiting there would only burn the budget.
- The `automations` tool's `run` passes its `timeoutMs` as the wait. `runs` now forwards `runId`. The tool descriptions now steer models away from verify jobs. Against a shipped Gateway, which rejects the new param before enqueueing, the tool retries the plain call once, the same way it already does for `compact` on `cron.list`.
- Accepted runs stay accepted, and outcomes go only to callers that may still read them. After the history read, the handler rechecks read authority (caller authority, agent runtime and any management grant), with no await between that check and the response. If the read or the recheck is refused, for example because a delegated grant lapsed or was revoked during a long wait, the caller gets the plain ack and nothing about the outcome. If the run finished but its history is not visible (for example a one-shot that deleted itself after succeeding), the ack carries `finished: true`. Management-only turns are pointed to the Automations page instead of `runs`.
The CLI keeps its client-side `--wait` polling in this PR. Moving it onto the Gateway wait is a separate, compatibility-reviewed follow-up.
## Evidence
- Real isolated Gateway, real Telegram plugin on the Crabline synthetic Bot API, and a deterministic Responses provider. An owner DM turn calls the `automations` tool (throwaway harness, not committed).
- Before (origin/main build): every `run` returned only `{ ok, enqueued, runId, processInstanceId }` in under 1 s, with no outcome. The report still reached the chat, so the model couldn't tell the run had succeeded.
- After: the isolated job's `run` returned in 4.3 s with `run { status: "ok", completionStatus: "succeeded", deliveryStatus: "delivered", summary: "FAST-REPORT-DELIVERED" }`, and the report reached the chat. A slow job (8 s, called with `timeoutMs: 2000`) returned after 2.6 s with its `runId` and the `runs jobId runId` note. The run kept going and delivered, and a later `runs` call with that `runId` returned the single finished entry. A main-session job run from the owner's main-session DM returned in 0.6 s with the note instead of blocking the turn.
- Gateway regression (`src/gateway/server.cron.test.ts`): `waitTimeoutMs: 50` returns only the `runId` while the run is blocked. `waitTimeoutMs: 10000` returns `run { runId, status: "ok", completionStatus: "succeeded", summary }`. On origin/main this fails because the closed schema rejects `waitTimeoutMs`.
- Same-session skip (`cron.validation.test.ts`): from an agent turn, `main` and `session:<own key>` jobs return the plain ack without waiting. `isolated` jobs wait.
- Authority chain (`cron.validation.test.ts`, through the real `cron.run` and `cron.runs` handlers with a stored run record): a caller that is still authorized gets `run { runId, status: "ok" }`. A caller revoked during the history read gets only the accepted ack. Before this change the revoked case turned into a request error.
- Tool (`cron-tool.test.ts`): the wait is capped at 10 min, and the RPC timeout is the wait plus 60 s. An unfinished run gets the `runs jobId runId` note.
- Tool contract (`cron-tool.output-contract.test.ts`): a shipped Gateway that rejects `waitTimeoutMs` still enqueues through a single retry, and the result stays valid against the output schema. The `run` output type only declares the outcome fields, so the generated code-mode declarations stay within their size allowance.
- Measured local wall times on head `cf942df0601` (one `node scripts/run-vitest.mjs <file>` per file; wrapper wall / Vitest duration): `cron-tool.output-contract` 5 tests, 24s / 7.0s; `cron-tool` 72 tests, 19s / 16.2s; `server-cron-lazy` 25 tests, 21s / 17.6s; `cron.validation` 261 tests, 24s / 21.4s; `server.cron` 17 tests, 46s / 43.8s. There are no CI seconds yet: CI skips the test jobs while the PR is a draft, and the only job that ran, "PR context and evidence", passed in 38s.
- `pnpm tsgo:core` is clean. Targeted tests pass: `server.cron`, `cron.validation`, `cron-tool`, `server-cron-lazy`.
### Bounded cost
- The wait is bounded by `timeoutMs` (default 60 s, capped at 10 min in the tool) and ends early if the caller aborts. The `cron.run` RPC timeout is the wait plus a 60 s enqueue budget, so the response always arrives before the client gives up.
- No new model calls or job runs. The waiter only watches a run that was already accepted. It does not retry, re-run or alert.
- No nested waits. Automation-run sessions are limited to self status/list/get/runs/remove, so a job's run cannot call `run`. A job running itself gets `already-running`. Main-session and same-session runs, which can only start after the caller's turn, skip the wait.
### Hermes
Hermes `cronjob(action="run")` sends the job to its async-delegation pool and returns a handle with "Do not wait or poll". The outcome comes back into the conversation as a new turn, and Hermes runs the job inline when no session can receive it. This PR borrows "the tool owns the follow-up, so the model doesn't poll". Feeding completions back into the conversation is not done here: OpenClaw already delivers run output to the job's target. It's a possible follow-up for long runs.
LOC (commit): production about +125/-12, tests about +96, docs +3.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(subagents): keep raw-key child cancel from stopping another agent's work
Cancelling a watched sessions_send child registered under a raw session key
such as `global` re-derived its owner from config and cleared every agent's
queue under that key. With two agents this aborted the default agent's run,
marked it aborted, and dropped its queued follow-ups and lane commands; even
with the owner resolved correctly, another agent's queued work was dropped.
Registration now records the known owning agent as optional `childAgentId`
for raw child keys (payload_json only; no DDL or schema-version bump). One
owner resolver feeds direct cancel, kill-scope discovery, the durable-kill
sweep, provisional reconciliation, and completion reads, and queue cleanup
uses the agent-scoped clearSessionLifecycleQueues owner. The raw-key
clearSessionQueues helper and its re-exports are deleted. Sweep session reads
move off the main thread onto worker reads.
* fix(subagents): align registry fixture with session owner reads
Configure the seeded fixture store and keep synthetic generation facts on the same session owner. Preserve real-store reads, identity and revision checks, and release invalidation; extract session mock wiring to keep the registry suite within its line-count budget.
* fix(anthropic): recognize sdk-ts entrypoint in Claude session catalog
Claude Code's bidirectional stream-json subprocess mode (used by the
anthropic plugin to serve adopted web sessions) writes transcript
lines tagged entrypoint: "sdk-ts". CLI_ENTRYPOINTS only recognized
"cli" and "sdk-cli", so discovery treated the first sdk-ts line as an
unrecognized entrypoint and stopped scanning the file, making the
session permanently invisible in the catalog even though its
transcript was otherwise valid.
Add "sdk-ts" to CLI_ENTRYPOINTS alongside the existing values so these
sessions are discovered, listed, and readable like any other Claude
CLI session.
* docs(anthropic): document sdk-ts as a recognized catalog entrypoint
Keep the Claude session catalog discovery docs in sync with the
CLI_ENTRYPOINTS fix: bidirectional stream-json sessions are now listed
alongside interactive and headless CLI sessions.
* docs(anthropic): scope sdk-ts discovery to Gateway reader
---------
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
Only classify an unstarted tool call with structured steering-skip evidence as an intentional omission. Preserve earlier genuine failures and keep other skipped or executed failures reportable.
Co-authored-by: RileyJJY <0668000974@xydigit.com>
* feat(ios): snooze sessions from the session menu
Add shared calendar presets and identity-bound snooze/wake patches while preserving existing Swift transport call sites. iOS gains Active, Snoozed, and Archived scopes, wake labels, and deadline-driven updates for Sessions, the sidebar, and Overview. Snoozing keeps ongoing work and the open conversation intact.
* test(ios): capture session snooze proof states
Extend the synthetic Gateway with awake and snoozed sessions. Capture Active and Snoozed lists, snooze presets, and Wake through the real iOS session menu, using All Sessions navigation and a row selector that excludes the retained sidebar.
* fix(ios): show the weekday for the next-week snooze preset
* chore(ios): record snooze strings in the native i18n source inventory
* fix(ios): format snooze wake labels through the localized catalog
Apple i18n verification rejects runtime-interpolated localized strings
because they bypass the generated catalog, so the "Wake session" menu item
and the "tomorrow" wake description now use the repository's
String(format: String(localized: "… %@")) pattern, and the native string
inventory records the format sources.
* refactor(ios): drop the superseded mutation lease initializer
* fix(ios): keep snoozed sessions browsable offline
Separate cached browsing scopes from connected mutation controls. Keep Archived connected-only and reset an archived selection to Active on disconnect. Cover scope availability and the offline cached roster, including snoozed rows and wake labels.
* chore(doctor): trace the standalone original-state capture phase
* fix(doctor): narrow the install root inside the traced capture
* perf(doctor): snapshot the pre-repair capture under maintenance custody
Use one database backup worker and local metadata sizing for standalone
Doctor. Maintenance custody selects safe in-process discovery copies and
generation sealing; existing process-local handles transparently retain
isolated acquisition, preserving capture availability for embedded callers.
Keep updater acquisition unchanged when omitted, along with both inventory
passes, payload verification, generation sealing, and manifest schema v2.
Validation: backup 14/14, Doctor fleet 6/6, unchanged fresh-preview 40/40,
recovery status 29/29, read-only 42/42, snapshot admission 9/9; core and
core-test type lanes, lint, formatting, artifact build, and fresh review.
Built-CLI sampling retains one backup worker, two workshop read-only
workers, and no node --eval children; macOS ACL helpers remain unchanged.
* fix(doctor): retire standalone captures older than 30 days
Retire only sealed schema-v2 baseline Doctor UUID captures older than
30 days with no outcome or update-run record. Preserve the current run,
symlinks, incomplete captures, update captures, and privacy markers.
Report every removal and retain failed removals with a warning. Document
the retention window and recommend verified backups for long-term copies.
Validation: backup and retention tests 14/14 (47.041s wall), existing
Doctor fleet run 6/6 (55.933s wall), unchanged fresh-preview 40/40,
focused recovery/read-only suites, type lanes, lint, formatting, build,
built-CLI sampler, and independent review. No CI run was dispatched.
* fix(doctor): size the maintenance capture deadline from the discovered inventory
* fix(doctor): drop the unused retention export
* fix(doctor): bound in-process maintenance copies and keep inventory reads cancellable
* fix(build): keep the capture acquisition type out of the owner cycle
Keep app and XCTest compilation on every admitted iOS smoke job. Select the
voice/media/typography and Access/chat lifecycle simulator groups from their
runtime, test, fixture and build owners, and omit simulator preparation when
neither group is selected. Preserve full scheduled/manual/release coverage.
Add OPENCLAW_CI_IOS_SIMULATOR_FULL to restore both PR groups, emit selection
reasons in preflight and job summaries, and include the planner in the trusted
preflight/platform harnesses. Retain every existing test case and assertion.
Test cost: the complete iOS workflow file passed 67 cases in 70.65s locally
while native compilation overlapped; the three integrated workflow/checkout
files passed 653 cases in 182.09s on Linux. Native build-only and lifecycle-only
paths passed with an unchanged app executable hash and no simulator use for build-only.
* feat(lobster): support structured workflow input
* docs(lobster): use current agent configuration examples
---------
Co-authored-by: Patrick Erichsen <patrick.a.erichsen@gmail.com>
Keep direct runtime consumers, expanding to depth two only below 20 direct
importers. Select adjacent tests only in directories with at most 30 tests;
larger directories use name prefixes and explicit ownership. Keep package
consumers direct so mixed import styles cannot bypass the hub cutoff.
Use config-loading and plugin-registry smoke within two Node rows. Preserve
the source-inventory guards, full-selection kill switch, and full hourly plans.
A controlled 48-diff replay drops from 2,193 to 1,083 rows, with 44 of 48 at or
below 40. Historical failure replay has 29 hits, one process-fixture miss
retained by hourly main, and three unavailable historical test files.
* fix(agents): await sessions_send reply custody transfer
Hold the requesting tool invocation until the existing continuation owner has
retained reply authority. Keep child result observation detached so a completed
requester turn cannot race asynchronous announce admission.
Cover delayed authority capture through the sessions_send entry point using
the existing source-route fixtures, and document reply lifetime behavior.
* test(agents): await reply admission in continuation proofs
* fix(agents): retain reply authority through child settlement
Bind accepted in-process agent turns to retained source authority instead of
their requester's turn-scoped participant guard. Keep participant checks at
admission and captured operator checks on source and creation guards.
Exercise native and peer sessions_send replies when the requester closes
before dispatch or completion, after releasing its original source hold.
The post-session plugin repair re-ran every default-phase plugin detector under
repair authority to certify completion that only deferred-migration resolution
and retained-source settlement consume. A fresh install has neither, so the
second detector pass (and the codex task detector's readonly SQLite child) now
runs only when deferred migrations are pending or the session-transcripts owner
registered its settlement hook, which it does only when the import retained
sources. completedPluginIds is undefined when nothing could consume it.
* fix(media): allow gateway media in workspace-scoped tools
Use the media store owner to retain browser screenshots and staged inbound files alongside the active session root. Keep sandbox bridge reads and canonical containment checks unchanged, including outside-file and symlink rejection.
* test(media): expect the Gateway media root in tool fixtures
Keep successful children with authoritative empty terminal replies in parent completion batches with their task identity, status, and explicit no-output result. Preserve intentional silence, announce-skip tokens, and stale-fallback suppression.
The six-child requester-settle regression fails on the base with child 0 missing. All 119 focused announce and completion tests pass, along with the core typecheck, formatting, targeted lint, and diff checks. Codex pre-commit review is clean through P2. Production LOC: +4/-9 (net -5).
Restore missing fs-safe prebuilds without relying on Node directory-merge semantics, preserve Bun native-launch diagnostics, and recover malformed Codex JSON without consuming the next valid frame. Keep package custody, bounded diagnostics, warning-only lifecycle failures, and installed-updater ordering intact.
Proof: 53 changed-test cases plus six catalog sibling cases pass on each of Node 24 and Bun; full changed-file validation and both import-cycle checks pass. Published Node-driver upgrade and Node-free Bun lifecycle cells passed on AWS. The published Bun driver's owning-npm preflight limitation remains explicitly documented.
CI still used the separately maintained Bun artifact and kept native-compiler tests on Node. Pin the OpenClaw Bun fork prerelease at 57fadf566d with release metadata verification and independent archive/executable checksums. Admit the 17 compiler files and library through their existing runtime owners, retain Node siblings and dual-mode coverage, and require native PTY success in the Bun-only smoke.
Proof: all ten Linux AWS selections passed (44,534 case executions), Node focused tests passed 243/243, Bun routing tests passed 208/208, and changed-file, workflow, and import-cycle checks passed. The old-first full smoke was fail/pass/pass/pass, attributing Chrome's fresh-host first launch to a shared startup issue. Eight forced-Bun ledger global-stub failures reproduce unchanged on main; that tooling suite stays on Node.
Supersedes only the pin portion of #159988. The batch-2 pin bump remains separate.
Run Node-only maintainer tooling fixtures through the existing Node-selection helpers while keeping Vitest on its selected runtime. Preserve package-integrity, update interruption, isolation, and declaration assertions, and restore Bun-hosted compensation coverage.
Proof: all 21 consumers pass on Node 24 and Bun (1043 passed, four existing skips each); changed-file checks and import-cycle checks pass; independent P2 and exact-head ClawSweeper reviews found no actionable defects.
Land through the authorized native exception for CI run 36837194384 attempt 1: the untouched Windows session-creation path-alias failure independently reproduces on main; nine downstream cancelled jobs remain unrun coverage.
Live update progress can leave committed data in an orphan WAL. Windows
rollback closes its native exclusion probe before replacing files, which
checkpoints that WAL and makes the saved physical write generation look
changed. The updater then refuses recovery and leaves the failed candidate
package installed.
Run the same Windows exclusion probe before eligible snapshot capture,
under the existing maintenance owner. Keep snapshot-only manual recovery
when another SQLite connection prevents exclusion, and preserve strict
fingerprint, ownership, and cleanup checks. The existing Windows lifecycle
regression and all its restoration/autostart assertions remain unchanged.
Regressed by 6e6458a98f (#162345), confirmed
by reverting only its progress.ts changes. This repair belongs to the
installed updater; a candidate cannot repair an already-running driver.
Validation: 252 selected lifecycle, rollback, progress, and direct-caller
tests passed. The known admission-ledger "Phase: verifying" failure
remains unchanged and is owned by #162495. Core types, all 27 core test
type graphs, changed-file checks, formatting, diff checks, and managed
Codex review through P2 passed.
Adds named storage locations as a generic, pluggable capability, with backup as its first consumer.
- Core storage owner (src/storage): storage.locations config, a location marker that binds identity (runtime never creates it, so unplugged disks and different disks at the same path are refused), client-side streaming encryption (scrypt key from a SecretRef passphrase, per-object HKDF keys, AES-256-GCM segments), and a built-in filesystem provider for external disks and mounts.
- Plugin SDK: api.registerStorageProvider plus manifest contracts.storageProviders; providers move opaque bytes only.
- Bundled cloudflare plugin: an r2 provider over the S3 API with conditional writes and bounded multipart uploads; auto-enabled when a location uses provider "r2".
- Backups: backup create --to <location> with verified archives, UTC retention, list/verify/restore --from, Gateway-owned offsite schedules (installed Git schedules unchanged), per-installation namespace claims fenced at publication and deletion, backup record for external jobs, backup.status RPC, Doctor/status hints, and a Systems page Backups section.
No config or state migration; the storage section is new and optional. Proof: live R2 and mounted-disk round trips, namespace takeover trace, and a published 2026.9.7 upgrade cell with an existing Git backup schedule.
Move durable session reads and upstream marker settlement to the existing database workers. Preserve idle-owner admission, exact-link CAS, durable event-before-marker ordering, and joined shutdown; isolate settlement failures to each provider outcome.
* refactor(state): register worker operations once per domain
Infer shared-state worker command contracts from lazy per-domain handler tables. Migrate Web Push, APNs, worktree registry operations, and fleet registry while preserving the existing broker and transaction owners.
* refactor(memory): move retained index reads into workers
* test(memory): align fixtures with worker read ownership
Repair the ten fixture failures exposed by retained worker reads. Install the embedding generation and token budget, retain a file-backed startup owner, and preserve the explicit scheduler-yield proof through source-wide snapshots.
Follow-up to #161053 (shared-session emoji reactions).
Reaction writes now run through the SQLite worker admission the sibling
session stores use (runOpenClawAgentWorkerWrite); the native path stays
only for process-held incognito databases the worker cannot reopen by
path, and the handler revalidates live authority around the awaited
write.
The plugin action dispatch path awaits onPlatformSendDispatch right
before the synchronous handoff fence, exactly like the send path, so the
reaction mirror re-reads the conversation binding at the final handoff
and refuses a message whose conversation was rebound while the action
runner prepared delivery.
Channels with one bot reaction per message (Telegram bots, WhatsApp)
declare the new optional ChannelPlugin.capabilities.reactionSlots =
"single"; when a person removes one emoji while others remain, the
mirror re-sets the newest surviving emoji instead of clearing the slot.
Multi-slot channels are unchanged.
Also trims redundant scaffolding in the reaction handler, kernel, UI
component and worker.
Proof: 119 focused tests across store, handler, dispatch and UI; mocked
Gateway reactions e2e; typecheck lanes; database-worker inventory check;
live two-person Gateway proof with the qa-channel mirror reporting
delivered through the final-dispatch hook.
The native Android chat shows reaction chips under saved prompts and
replies (emoji, count, highlighted when the viewer reacted, TalkBack
lists who reacted) plus an add-reaction chip; tapping a chip toggles the
viewer's own reaction, and the add chip or message actions open a Quick
reactions sheet with the Control UI's palette plus "More…" for any single
emoji. Updates arrive live from the session.reaction event; nothing
notifies.
Reaction wire models are selected into the generated GatewayProtocol.kt,
the hello summary retains sessionCap, sessions carry sharingRole and
visibility, and ChatReactions.kt owns state, the permission projection
ported from the web's canReactToSession, the palette and the
single-grapheme rule. Writes run FIFO per message with per-emoji response
guards, list snapshots advance revisions so late set responses cannot
overwrite refreshed state, and message ids are canonical entry ids from
__openclaw.id. Generated locale artifacts are left to the refresh
workflow.
Proof: pnpm android:assemble, focused Robolectric/Compose tests (hello
summary, controller list/set/event/race regressions, chips/palette),
ktlint and Android lint, protocol and i18n checks, emulator before/after
screenshots in the PR, Codex branch review with its findings fixed.
* fix(gateway): keep model metadata available during plugin drains
Keep the active model publication readable while admitted plugin work settles,
then retire it before resource replacement. Preserve execution fencing, auth
revocation, decision cancellation, rollback, and cleanup ownership.
Retain an automatic drain failure only while its original plugin configuration
delta remains unresolved. Allow explicit wait recovery and reversion alongside
changes to a different plugin without retrying unrelated failed work.
Validate on Blacksmith Testbox with real Gateway reader latency proof, 378
focused tests, 145 follow-up tests, negative controls, types, lint, and guards.
* fix(plugins): fence admission while preserving owned cleanup
* test(gateway): preserve plugin record helpers in reload fixture
* test: repair reload mocks and supervised process joins
* test(models): preserve queue receiver in drain observer
* fix(doctor): retain migrations for supported config formats
* fix(doctor): retain the typed sandbox scope for reporting
* fix(doctor): include agent policy in wrapper extraction
Defer unmapped and indirect runtime browser compositions to hourly main and
release validation. PRs keep five documented smoke files, edited tests,
transitive fixture/test imports, and declared runtime owners plus their
direct imports. Narrow full fallback to core E2E harness, UI bundle config,
and explicit app-shell inputs. Whole-route watches now name route entries
rather than every leaf under the route; source-backed CSS owners remain.
Shared shell/store cycles previously selected hundreds of unrelated routes
for a single leaf change. Across 20 recent UI PRs, median selected files
fall from 654 to 6; 15 select fewer than 60. Pack 30 files per browser row,
retaining existing full-mode caps and worker limits. Median modeled wall
falls from 641.2 to 219.6 seconds including a conservative 200-second reserve.
This is a native timing/LPT projection, not observed CI wall.
The original ten causal failing runs produce four file misses across three
runs: sidebar-customization, plugin-bundled-view-recovery, debug-diagnostics,
and config-controls-visual. All remain in the full hourly inventory. The
next hourly is the accepted safety net; queue/runtime mean detection within
one hour is not guaranteed. The existing full-PR kill switch is unchanged.
No browser assertion, screenshot, test deadline or product behavior changes.
Add planner coverage for deferred unknown/transitive owners, direct fixture
imports, full periodic inventory, kill switch and bounded matrix rows.