Related: #130753, #137832, #147969
## What Problem This Solves
Fixes: some scheduled jobs created by an agent fail for months because a tool list saved by an older OpenClaw build is missing tools the creator actually had, such as the native shell. In our setup, a monthly group job that runs `node <script>` delivered nothing in August, delivered nothing in September (the run still reported `ok`), and posted a blocker in October. Its saved list had 31 tools and no `exec`.
## User Impact
User impact: an agent-created agent-turn job that does not name specific tools now gets the same tools as its owner conversation at run time, like a job an operator creates without `--tools`. Existing jobs with an automatically saved creator snapshot behave the same way from their next run. Nothing stored is rewritten: no migration and no backups. Explicit tool lists, script payloads, condition triggers, and jobs bound to captured Codex app authority keep their stored list.
Tradeoff, approved by the maintainer (Ayaan): a per-sender tool policy on the creating owner, or a plugin hook that narrowed the creating turn, no longer limits these default jobs. Only owners can create automations from chat, and subagents cannot create them.
## Why This Change Was Made
**History of the saved list.**
- #91499 introduced it so a delayed run cannot do more than its creator could.
- #112483 made every agent-created job store one, because runs have no sender.
- #112661 made scheduled runs re-apply the owner session's group policy and every non-sender limit, keeping the stored list as the upper bound.
- #137832 fixed native tool capture for new jobs only, and deliberately did not widen stored lists.
- #147969 added a Doctor advisory. It only fires for claude-cli, so it never covered Codex-harness or built-in OpenAI jobs like ours.
**Root cause.** When no tool list was given, OpenClaw saved a frozen copy of the creating turn's tools instead of treating the job like an operator `*` job. Every capture bug (missing native tools, late configured MCP, renamed tools) then stayed in the job permanently.
**Fix.** This follows Hermes, which keeps no creator snapshot: `cron/scheduler.py` `_resolve_cron_enabled_toolsets` reads toolsets from config at run time.
- **New jobs.** An agent-turn create or update with no list, or `*`, stores `["*"]`. That is the same value operator jobs store, so the job's tools match a normal turn in its owner conversation. Script payloads and condition triggers still store the creator's concrete tools, because a script reaches MCP only through servers its list names. Jobs whose creator captured Codex app authority also keep the concrete list, because that authority is bound to it.
- **Existing jobs.** One helper, `resolveCronRunToolsAllow` in `src/cron/tools-allow.ts`: a stored automatic snapshot (`toolsAllowIsDefault`) runs as `*` when it has a valid scheduled owner policy, no condition trigger, and no Codex app authority. Otherwise it keeps its stored list. Every execution consumer of the stored list uses it: the run payload, the command-prompt preflight, and the scheduled message authority.
- **Script transitions.** A `*` job that becomes a script, or gains a condition trigger, captures the creator's concrete tools.
- **Exec pin.** A `*` list keeps the creator's exec host pin.
- **No new noise:** automatic snapshots stay excluded from the `web_search` provider warning, as on main.
- **Deleted, now pointless:** both Doctor advisories about incomplete automatic snapshots, the run warning about pre-MCP snapshots, and two exports nothing uses anymore.
Review note: on claude-cli, a `*` job runs without a CLI tool cap, so Claude's native tools behave exactly as in a normal chat turn in that conversation. This PR introduces no new path around `tools.deny` that a chat turn doesn't already have.
## Evidence
Live-model Telegram proof (Telegram Test Server DM, leased team credential, live `openai/gpt-6-astra` reached through a forwarding proxy that stands in for the runner's mock provider; the runner harness itself is unchanged). This reproduces the shape of the original incident:
- The tester DMs the bot, which creates the owner conversation.
- A job owned by that conversation is added. Its stored list is an old-style automatic snapshot `["automations","message","read"]` plus `toolsAllowIsDefault: true`, with no `exec`.
- The payload is `Run: node scripts/split-report.mjs and post its output line verbatim`. The workspace script prints a random nonce.
- The job is run once (`cron run --wait`), with announce delivery to the DM.
| Build | `exec` offered | Model action | What arrived in the DM | Run |
|---|---|---|---|---|
| base 94f5a8d (main before this PR) | no | `tool_search` ×2, then gave up | "Could not run node scripts/split-report.mjs: no command-execution tool is available…" | error |
| **this PR, head 5b78cb7** | **yes** | `exec {"command":"node scripts/split-report.mjs"}` | "**SPLIT-REPORT 93C53909**: general 41, design 17, ops 9" (the exact script output, with this run's random nonce) | ok, delivered |
| head 5b78cb7 with `tools.deny: ["exec"]` | no | `tool_search`, `read`, then gave up | "Could not run node scripts/split-report.mjs: no command-execution tool or paired node is available…" | error |
In every run, the stored job kept `["automations","message","read"]` plus the marker. Before and after use the same scenario and driver; only the checkout differs.
Update and live proof: published `openclaw@2026.9.7`, then this branch at the exact head (2f5099d), on the same state directory. Mock provider. Every process ran under a temporary `HOME` and state directory. Each job's message makes the model call `exec` with `touch <effects>/<job>`.
1. 2026.9.7 created both jobs through `cron.add` (scheduled policy `trusted`). With the Gateway stopped, the "stale" job was given the old automatic-snapshot shape `["automations","message","read"]` plus `toolsAllowIsDefault: true`. sha256 of both stored rows: `609358837…`.
2. Runs:
| Build / config | Job | `exec` offered | Side effect | Run |
|---|---|---|---|---|
| 2026.9.7 | stale automatic snapshot | no | absent | error |
| 2026.9.7 | explicit `["read","message"]` | no | absent | error |
| this branch | stale automatic snapshot | **yes** | **created** | ok |
| this branch | explicit `["read","message"]` | no | absent | error |
| this branch, owner policy narrowed to `tools.deny: ["exec"]` | stale automatic snapshot | no | **absent** | error |
| this branch, `tools.deny: ["exec"]` | explicit `["read","message"]` | no | absent | error |
3. After the branch runs, the stored rows were byte-identical (same sha256 `609358837…`), and job ids and lists were unchanged. Nothing was migrated.
An earlier run at e151ea3, with the same harness, also covered a snapshot bound to Codex app authority: `exec` was not offered, the file stayed absent, and the stored row was unchanged.
Tests:
- `run.tools-allow.test.ts`: a stored automatic snapshot `["message","read"]` reaches the embedded run as `["*"]`, with the owner's scheduled policy intact. It fails on main with `["message","read"]`.
- `cron-tool-creator-cap.test.ts`: a default agent turn stores `["*"]`, while a trigger script and a Codex-app creator keep the concrete snapshot.
- `run.tools-allow.test.ts`: snapshots without a valid owner policy, or behind a condition trigger, keep their list. Both cases fail on the previous head.
- `run.tools-allow.test.ts`: no `web_search` warning for an automatic snapshot that kept its list. This fails without the exclusion.
- `run.message-tool-policy.test.ts`: a self-edited automatic snapshot runs on CLI with no cap.
- `run.tools-allow.test.ts`: a legacy `Command to run:` prompt from an automatic snapshot without shell tools now runs instead of being rejected.
- `jobs-tool-policy.test.ts`: scheduled message authority is admitted for an automatic snapshot that lacked `message`.
- `cron-tool-creator-cap.test.ts`: a `*` agent turn converted to a script captures the creator's concrete tools.
- These three regressions fail on the previous head. `node scripts/check-changed.mjs` passes.
- Explicit-list, exec-pin and gateway creator-transport suites pass. `pnpm tsgo:core` passes.
## Bounded cost
No new path triggers a model call or a job run. The change only selects which tool list an already scheduled run uses.
LOC vs main: production +98/-250 (net -152), tests +140/-373, docs +19/-11.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
The update database generation check compared physical fingerprints, so a SQLite checkpoint that merely folded already-captured WAL frames into the main file (on Windows, triggered by the exclusive removal probe) was treated as a later write and rollback was refused with "Databases changed after snapshot capture"; the Windows lifecycle test failed in isolation. The check now compares committed pages plus retained WAL commit evidence: checkpoint-only transitions and exact reversals that leave byte-identical content are admitted (identical content cannot lose data on restore; recorded at the comparer with a regression), while any newer data is still refused. The fixture's unsafe in-process snapshot mock is removed.
Refs #158163
* chore(doctor): trace the standalone original-state capture phase
* fix(doctor): narrow the install root inside the traced capture
* perf(doctor): snapshot the pre-repair capture under maintenance custody
Use one database backup worker and local metadata sizing for standalone
Doctor. Maintenance custody selects safe in-process discovery copies and
generation sealing; existing process-local handles transparently retain
isolated acquisition, preserving capture availability for embedded callers.
Keep updater acquisition unchanged when omitted, along with both inventory
passes, payload verification, generation sealing, and manifest schema v2.
Validation: backup 14/14, Doctor fleet 6/6, unchanged fresh-preview 40/40,
recovery status 29/29, read-only 42/42, snapshot admission 9/9; core and
core-test type lanes, lint, formatting, artifact build, and fresh review.
Built-CLI sampling retains one backup worker, two workshop read-only
workers, and no node --eval children; macOS ACL helpers remain unchanged.
* fix(doctor): retire standalone captures older than 30 days
Retire only sealed schema-v2 baseline Doctor UUID captures older than
30 days with no outcome or update-run record. Preserve the current run,
symlinks, incomplete captures, update captures, and privacy markers.
Report every removal and retain failed removals with a warning. Document
the retention window and recommend verified backups for long-term copies.
Validation: backup and retention tests 14/14 (47.041s wall), existing
Doctor fleet run 6/6 (55.933s wall), unchanged fresh-preview 40/40,
focused recovery/read-only suites, type lanes, lint, formatting, build,
built-CLI sampler, and independent review. No CI run was dispatched.
* fix(doctor): size the maintenance capture deadline from the discovered inventory
* fix(doctor): drop the unused retention export
* fix(doctor): bound in-process maintenance copies and keep inventory reads cancellable
* fix(build): keep the capture acquisition type out of the owner cycle
Adds named storage locations as a generic, pluggable capability, with backup as its first consumer.
- Core storage owner (src/storage): storage.locations config, a location marker that binds identity (runtime never creates it, so unplugged disks and different disks at the same path are refused), client-side streaming encryption (scrypt key from a SecretRef passphrase, per-object HKDF keys, AES-256-GCM segments), and a built-in filesystem provider for external disks and mounts.
- Plugin SDK: api.registerStorageProvider plus manifest contracts.storageProviders; providers move opaque bytes only.
- Bundled cloudflare plugin: an r2 provider over the S3 API with conditional writes and bounded multipart uploads; auto-enabled when a location uses provider "r2".
- Backups: backup create --to <location> with verified archives, UTC retention, list/verify/restore --from, Gateway-owned offsite schedules (installed Git schedules unchanged), per-installation namespace claims fenced at publication and deletion, backup record for external jobs, backup.status RPC, Doctor/status hints, and a Systems page Backups section.
No config or state migration; the storage section is new and optional. Proof: live R2 and mounted-disk round trips, namespace takeover trace, and a published 2026.9.7 upgrade cell with an existing Git backup schedule.
doctor --fix heap-OOMed on a 739-agent state: the runtime tool-schema check resolved a model per agent workspace, and every plugin-load cache context cloned the full runtime config and retained a registry per workspace (516 MB of config clones at 2.4 GB heap). The loader now defers config capture until a cache miss and shares prepared provider registrations within a captured config generation for agents with matching configuration and plugin sources; a reentrant load during capture joins the in-flight load instead of registering twice. 300 synthetic agents: config clones 1,802 -> 3, registries 300 -> 1, peak RSS 1.25 GB -> 0.88 GB.
Closes#161869
openclaw update --no-restart still waited up to 30 minutes for Gateway readiness: the failure-recovery owner inferred a Gateway startup from the package mutation and observed readiness regardless of restart: false. Recovery now disables only the startup readiness wait when restart is explicitly false, keeps one-shot verification and the recovery diagnostics, and records restart: skipped by operator after service preparation; the default restart path is unchanged.
Refs #161795
`openclaw skills install <source> --as <name>` reported "Installed" for a skill that discovery can never load, such as a `SKILL.md` with no parseable description, so the skill silently never appeared. Install now checks the source against the discovery requirements before copying. It fails with an error naming what's missing and copies nothing.
Fixes#122298. Thanks @Yun-0000 for the fix and @suninweb for the report.
Proof: the freshly built CLI in a secretless container, on isolated profiles.
- On main, a malformed source was "installed" but was missing from `skills list`. With this change, the install exits 1, names the missing description, and copies nothing.
- A valid skill installs and appears on both builds.
- The regressions in the three named test files fail on main, and 90 targeted tests pass.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
The CLI update progress display polled getUpdateRun every 250 ms and, on the shared-state path, each synchronous read staged a full SQLite snapshot in a worker and blocked the updater's event loop. Progress now uses one reusable live read per session (at most one snapshot launch across all polls, none when the display is disabled) with cancellation fenced during suspension and disposal; phase rendering is unchanged. A foreground-handoff wait in the same owner dropped from 156 s to about 67 s. The pnpm stdin contract (#161866) is covered separately by #161920.
Closes#161867
Refs #161866
Doctor's session-repair step overflowed the stack (RangeError: Maximum call stack size exceeded) on large but healthy agent databases during 2026.9.7 update activation: canonical-owner resolution followed canonicalOwnerSessionKey alias chains recursively, and row-count-sized spreads sat on the same path. Resolution is now iterative with the same terminal-owner, missing-owner, and cycle outcomes, and the spreads are linear reductions with unchanged empty-input behavior. A healthy 12,000-owner / 200,000-event fixture fails pre-fix and passes post-fix; the published-driver upgrade preserved transcript bytes.
Closes#161979
Related: #140086, #141885, #156535, #157838
## What Problem This Solves
A running Gateway doesn't see a newly downloaded hosted model catalog until it restarts. This PR publishes each accepted catalog through the existing prepared-runtime owner, without a restart. Model rows and their prices switch together as one generation, which keeps the invariant from #140086.
## User Impact
- Compatible downloads are adopted at the Gateway's background catalog check, or after an explicit `models.list` refresh. That refresh returns the currently accepted rows right away and runs adoption afterward.
- A turn admitted on catalog N keeps N's rows and prices until it finishes. New turns use N+1. Rows and prices are never mixed.
- A concurrent auth or config publication no longer postpones adoption to the next scheduled check (up to 6 h). Adoption waits for that publication to settle, then retries, up to 3 attempts.
- An owner whose build failed or timed out ends the adoption instead of waiting on unbounded work. Gateway shutdown cancels an adoption that is still preparing.
- Malformed, schema-invalid, too-new (`minVersion`) and older catalogs are rejected, and the previously accepted catalog stays in use.
- Changing `models.catalogRefresh.url` no longer needs a restart: the previous source's catalog stops applying, and the mirror's catalog is adopted at the next catalog check.
**Bad-catalog exposure:** with live apply, a *valid but wrong* published catalog reaches running Gateways at their next catalog check (at most every 6 h) or on the next explicit `models.list` refresh. It no longer waits for a restart. Recovery uses existing mechanisms only:
- Republish a corrected catalog with a newer `generatedAt`; Gateways adopt it the same way.
- Operators can set `models.catalogRefresh.enabled: false`, which withdraws remote rows and prices without a restart (covered by the Gateway integration test).
This PR adds no new kill switch, config option or env knob.
### Compatibility
No config keys, defaults, types, validation, stored rows, protocol or SDK contracts change. The only config-surface change is the `models.catalogRefresh.url` help text, which drops the stale "Changes apply after a Gateway restart" sentence, and its regenerated config-doc baseline hash. Existing configs validate unchanged and need no Doctor migration (maintainer confirmation: https://github.com/openclaw/openclaw/pull/158000#issuecomment-5913462783). Startup behavior is unchanged. Upgrade impact for existing installs: an accepted download activates at the next catalog check instead of the next restart.
## Why This Change Was Made
- `prepared-model-runtime.configured-refresh.ts` builds a complete candidate generation of the configured owners under the new catalog. One serialized commit then publishes rows, the accepted bundle, the pricing context and the reply-dispatch projection together.
- Adoption re-reads the stored catalog until the config it read under is still current, so a stale caller can't cancel a current adoption.
- Each preparation attempt has its own abort signal. A config advance restarts only the attempt; a newer catalog or shutdown ends the whole adoption, including pricing preparation.
- Between attempts, adoption waits only on publication gates: a pending replacement or an owner's pending publication.
- Adopted owners install the same plugin-retirement recovery as configured publication (#161267). After commit, a lost Gateway plugin loan republishes them through the normal recovery. Before commit, it restarts the adoption attempt, and the commit refuses any candidate whose plugin generation retired.
### Why downloads were restart-only, and what this keeps
Restart-only activation was a mechanism, not the goal. #140086 chose it to stop rows and prices from different catalog versions mixing, and #157838 was merged as "the prerequisite for applying new remote catalogs without a Gateway restart (rows and prices must switch together)". This PR is that follow-up. Every requirement those PRs set still holds:
| Original requirement | Source | How it holds here |
|---|---|---|
| Rows and prices from one catalog version; never mixed across reloads or new requests | #140086 | One serialized commit publishes owners, bundle, pricing context and dispatch; pricing contexts are keyed by the exact accepted catalog. Integration test: new rows appear only with new prices |
| Admitted work keeps its pair | #140086, #157838 | Runs carry their plugin generation's catalog; usage operations capture one pricing context. Integration test and live proof: the in-flight turn keeps the old price |
| Startup absence is a real state (no downloaded rows without prices) | #140086 | Absence → catalog goes through the same atomic commit; overlay absence tests unchanged |
| Worker replacement inherits the host's accepted pair, not a later download | #140086 | The commit updates the inherited pair; later workers and a worker-exit recovery keep it (overlay and integration tests) |
| Current enablement and source URL still gate eligibility | #140086, #156535 | Checked on every read and before adoption, including the default-install v1 fallback; disablement withdraws rows and prices together (integration test) |
| Bad or superseded downloads never replace the active pair | #140086, #141885 | Compatibility, `minVersion`, revision and `generatedAt` checks; stale reads can't cancel a current adoption (regression test) |
| Failed catalog checks retry at the remaining fresh interval, not a full TTL | #141885 | Unchanged scheduler behavior; the deleted notice test's retry case is restored for failed adoption (fails if the retry falls back to the full TTL) |
| Operators learn when a downloaded catalog is not yet active | #141885 | No longer needed: downloads activate at the next check. The restart notice and its tests are removed; `models refresh` says when a running Gateway applies the update |
| Billing-route prices switch with their rows | #156535 | `upstreamPricing` and `providerPricing` are part of the accepted catalog pair |
## Evidence
**Regressions.** Each fails with its fix reverted and passes with it:
- *Retries a scheduled adoption when its pending auth owner settles.* Runs through the real Gateway update scheduler. Reverted, it logs `remote model catalog check superseded; deferred to the next check`.
- *Does not let a read under a superseded config cancel the current adoption.* Reverted, both calls end `superseded`.
- *Ends adoption instead of joining a timed-out owner build.* Reverted, adoption never settles.
- *Does not hold Gateway shutdown on an adoption's pricing preparation.* Reverted, shutdown waits on the held preparation until the test times out.
- *Recovers adopted owners when their borrowed Gateway plugin retires after commit / before commit.* Without the recovery, both fail: `Prepared model runtime plugin generation retired` and `prepared reply dispatch runtime owner was not published`.
- *Uses the remaining stored TTL after a fresh startup check when adoption fails.* With the retry reverted to the full TTL, the second check doesn't run.
**Suites:**
| Suite | Result |
|---|---|
| `prepared-model-runtime.remote-publication.test.ts` | 10/10 |
| Gateway integration (`models-list.remote-catalog`) | v1 and v2 pass. Config and auth churn during preparation end `published` on the settled owners. Also covers retained admitted runs, rejected and stale bundles, worker replacement and disablement |
| `prepared-model-runtime*`, `server-plugin-reload*`, `update-startup`, and all PR-touched test files | pass |
| `tsgo:core`, all `tsgo:test:src` shards | pass |
| oxlint and oxfmt on changed files; `config:docs:check`, `config:schema:check`; max-lines, assertion-safety and test-timeout-race ratchets | pass |
Tests wait on owned completion signals (`withinTest`), not wall-clock deadlines.
**Live proof** on an isolated Gateway built from `cb8c9197fd` (no provider mocks). Later commits add plugin-retirement recovery for adopted owners, covered by the regression tests above, and rebases onto `main`. It used a real OpenAI key through `openai/gpt-4.1-mini`, and the build stamp was set before the real catalog's publication date. A client polled `models.list` back to back over one WebSocket for the whole run (1023 polls, no errors). One Gateway process (PID unchanged) and no restart:
1. The stored catalog was seeded with an older revision of the real `catalog.openclaw.ai` v2 catalog: generated 2 days earlier, `gpt-4.1-mini` priced ×10, plus one extra kimi row. `models.list` listed the extra row, and a turn priced **$4.00 / $16.00 per M** input/output.
2. `openclaw models refresh` downloaded the real catalog (`updated`, 1039 models). The listed rows didn't change for the next 7.1 s, and a turn in that window still priced **$4.00 / $16.00 per M**: a download stays inactive until the Gateway adopts it.
3. A long turn was admitted on the older catalog, then `models.list {refresh:true}` returned the older rows (extra kimi row still listed) and started adoption. The new catalog was visible 0.9 s later, while the long turn was still running: the extra kimi row was gone.
4. The in-flight turn finished at **$4.00 / $16.00 per M** (older rows and prices). The next turn priced **$0.40 / $1.60 per M** (real catalog).
5. `models.catalogRefresh.url` was moved to a local mirror of the real catalog through `config.patch`, and the mirror's catalog was adopted without a restart. The mirror then served malformed JSON: `models refresh` failed with `SyntaxError`, the model list was unchanged, and the next turn still priced $0.40 / $1.60 per M.
**Model picker during republication (also on `main`).** Right after the new generation commits, `models.list` shows the new generation's configured and static rows until its full catalog loads, then the full list. In the live run this lasted 109 ms. The same poller against a `main` build shows the same window after a `models.*` config reload (20 → 6 → 16 rows for about 350 ms), so this PR adds a new trigger for an existing behavior. It doesn't change it. The short list comes from the new generation, so rows and prices stay paired.
**Published-driver upgrade cells.** Candidate tarball built from a fresh clone at `6f682c7944` with the canonical Docker packaging script and no build-time overrides. sha256 `46d881a0…ef0ad`; embedded commit `6f682c7944`, version 2026.9.7. `6f682c7944` already includes the shutdown-cancellation and pricing-deadline commits. The current head the current head differs from it only by rebases onto `main`: `main` had moved the scheduled catalog check into `update-startup-catalog.ts`, and this PR's adoption call moved there unchanged (`git range-diff` shows no other production change; a `remoteCatalog: null` test-fixture field moved to main's relocated `cli-compaction.test-support.ts`).
| Driver → candidate | Scenario | Result |
|---|---|---|
| `openclaw@2026.9.6` | base | passed (930 s); updater outcome success, no recovery |
| `openclaw@2026.9.6` | plugin-deps-cleanup | passed (923 s); updater outcome success, no recovery |
| `openclaw@2026.9.7` (latest) | base | passed (813 s); updater outcome success, no recovery |
In every cell:
- Migration, post-Doctor config validation, survival, plugin-dependency cleanup and runtime-deps repair checks passed.
- The candidate Gateway logged ready, then its catalog check fetched and saved the hosted catalog about 0.2 s after starting. The hosted catalog is older than the candidate's build stamp, so adoption ends `unchanged`, which isn't logged. After a 300 s settlement window, `/readyz` (`ready:true`, nothing failing) and a Gateway `status` RPC passed, and the Gateway shut down cleanly.
- To show a logged terminal outcome, each cell was repeated with a loopback mirror serving a newer copy of the same catalog. Each check logged `remote model catalog applied` about 0.26 s after it started, followed by `/readyz` and `status` passing.
Not run: `openclaw@2026.9.7` plugin-deps-cleanup. At the previous head it failed inside the 2026.9.7 driver's retained-runtime verification (`Retained runtime entry does not reference its inventoried file: dist/a2ui-…mjs`), identically for a merge-base control package with none of this PR's commits.
Harness note for the 2026.9.7 cell: unchanged, the upgrade harness can't run against 2026.9.7. It seeds the retired `tools.toolSearch {mode:"code"}` setting, which 2026.9.7 rejects. The 2026.9.7 cell used the harness's existing Tool Search "absent" mode, a one-line local change that skips only that seed and its check. The 2026.9.6 cells used the unchanged harness.
**CI:** fully green on the final head ([run 36818367970](https://github.com/openclaw/openclaw/actions/runs/36818367970)). Earlier heads hit failures that reproduce on `main` in code this PR doesn't touch: `update-cli.target-schema` (main reproduction: https://github.com/openclaw/openclaw/pull/158000#issuecomment-5913464830), the type-suppression inventory, the Windows partition owner test and the Windows backup-rename test. Main has since fixed the last three. Config compatibility confirmation: https://github.com/openclaw/openclaw/pull/158000#issuecomment-5913462783.
No overlap with Pash/Sarah changes.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* feat(sessions): import Claude Code and Codex transcripts into OpenClaw
Add sessions.catalog.import, which preserves a native catalog transcript
(Claude Code, Codex, OpenCode, Pi, or shared OpenClaw sessions) as an
ordinary OpenClaw session so it survives the source tool's cleanup or the
loss of the source computer. It reads through the existing catalog read
path, so Gateway-local, headless node, macOS app, and Linux app sources
all work without app changes.
Re-importing the same source reuses a deterministic agent-scoped session
key and appends only items not yet imported, so repeated imports act as
an explicit sync. Imports keep up to 50,000 items or 64 MiB (newest first)
and report complete: false when older history is cut. Continuation keeps
its 200-item / 512 KiB seed; adoption, continue, and fork are unchanged.
Surfaces: the sessions.catalog.import Gateway RPC (operator.write, same
row visibility as sessions.catalog.read), `openclaw sessions import`
(single transcript or --all with paging, --dry-run, --json), and an
"Import to OpenClaw" item in the Control UI catalog row menu. The shared
catalog history reader now pages at 50 items, the Claude and Codex
transcript read cap.
* chore: merge main into transcript import branch
* fix(sessions): fence catalog import source access through commit
Retain the catalog visibility owner's matched source and require current
read authority for destination creation, transcript appends, and state events.
Compare access facts without rejecting ordinary transcript generations.
Cover forbidden readers, pre-write sharing revocation, and mid-append
revocation through real Gateway owners. Share the integration state fixture.
Use typed CLI options and local paging state to satisfy assertion and lint gates.
* fix(sessions): default catalog imports to drafts
Preserve published copies on re-import and honor the Gateway no-drafts policy. Split cheap import authorization checks from retained release-tier owner integration proof. Refresh the Workboard asset manifest required by the generation gate.
* Merge remote-tracking branch 'origin/main' into steipete/transcript-import-mode-a262c5
* test(ui): expect the import action in adopted catalog session menus
* Merge remote-tracking branch 'origin/main' into steipete/transcript-import-mode-a262c5
* Merge remote-tracking branch 'origin/main' into steipete/transcript-import-mode-a262c5
* Merge remote-tracking branch 'origin/main' into steipete/transcript-import-mode-a262c5
* Merge remote-tracking branch 'origin/main' into steipete/transcript-import-mode-a262c5
* Merge remote-tracking branch 'origin/main' into steipete/transcript-import-mode-a262c5
# Conflicts:
# docs/nodes/session-catalogs.md
Co-authored-by: Peter Steinberger <steipete@gmail.com>
The managed-service handoff lease owner retained an uncertain foreground handoff forever: a dead handoff owner's lease could never be reclaimed and no repair path could establish descendant absence, so updates ended in reconcile:abandoned and operators had to intervene by hand. Reclaim now uses process-census evidence (unverified or unresolved census vetoes), generation-bound metadata on the existing handoff table (nullable additive columns, same store version, lazy idempotent ensure), an atomic ownership claim, failed-repair retention, and a settlement receipt recorded before release; openclaw update repair has an explicit reclaim admission; reboot reclamation preserves generation-bound repair metadata; a state-override path that could bypass the fail-closed admission is closed. Legacy rows without metadata are reclaimed only after a 45-minute floor plus the census.
Closes#159897
Refs #159909, #161090, #161721, #161761
The pending-migration integrity gate ran a synchronous, unbounded full PRAGMA integrity_check on the main thread, so doctor --fix printed "agent database schema migration pending; verifying integrity first" and then hung for 25+ minutes ignoring SIGTERM/SIGINT on large agent databases. Doctor now routes that check through the existing child-process integrity owner, emits 10-second heartbeat lines, and wires cancellation to its existing signal cleanup; integrity checks, backups, and admitted-write settlement are preserved. Candidate exits within 100 ms of SIGTERM where the original blocked; the published 2026.9.5 updater completed against the candidate with a verified backup.
Closes#161888
A retained .openclaw.package-backup-<pid>-<timestamp> symlink in the global node_modules (left by an earlier source-to-package update) made the updater inventory walk into an unrelated checkout and refuse the update with the host-owned plugin-link error. The updater now excludes its own recovery artifacts (package backups and activation control files) from inventory through one producer-owned naming predicate, and retires historical backups once a later update has completed, never the current run's backup. The already-installed 2026.9.6 driver still refuses on that first hop; the issue carries the recovery note.
Refs #161922
Follow-up to #161832: when archiving a verified-empty retired Telegram bindings file fails (for example the .migrated target already exists), the migration records a recoverable warning instead of refusing the update; non-empty or uncertain sources keep the preservation refusal. Sanitized refusal messages keep their diagnostic text.
Refs #161795
Thanks @ericcaiwx-star for the base fix in #161832.
An in-place update replaces the package's hashed chunks while the old Gateway is still running. Shutdown then lazily imported cleanup modules that no longer existed: 2026.9.6 exited with ERR_MODULE_NOT_FOUND for the transcript capture-operations chunk, and other teardown paths (MCP runtime retirement, provider local services, browser relay and Chrome MCP, worker temp-artifact cleanup) had the same hazard.
Cleanup owners are now retained when their resources are created, so teardown never first-imports code. Transcript capture ownership moves into the eagerly loaded capture-startup module, and shutdown persists heuristic notes without starting optional model inference. Provider local services register their stop callback when they start. MCP cleanup has a lightweight owner module. Browser teardown peeks the lazy accessors that production ingress loads. The worker pool preloads artifact cleanup in its async preparation step before creating resources, which keeps chalk/tslog out of the native hook relay's static graph.
The provider-runtime-lifecycle stable tsdown entry stays, because published-update compatibility bridges map older hashed chunks onto it, and Knip treats it as an entry. Protection applies to updates from a release containing this change; older running Gateways keep the limitation, as docs/install/updating.md describes.
Proof: a built-runtime smoke removed 7,011 hashed chunks after startup and still shut down cleanly (exit 0, 6.3 s). Focused shutdown, transcript, worker-pool, browser and MCP suites pass, along with core types, import-cycle checks, Knip and the CLI bootstrap guard. Codex autoreview is scoped-clean. The final CI failures were main-side inventory breaks (wrapper components, suppression count), fixed on main by b3b34beebe and bd11387977.
On Windows the managed update swap renamed the live package tree to its backup location once; a transient EPERM/EBUSY/EACCES from an antivirus, indexer, or Scheduled Task handle failed the swap and rolled the update back. The filesystem owner of the swap now retries those codes with bounded backoff (16 attempts), records each retry, re-checks package identity and update authority between attempts, and reports a named failure after the bound. POSIX behavior is unchanged.
Refs #162027
* fix(update): reject unattributed database writes during Doctor
Keep the captured fingerprint baseline through Doctor settlement. Maintenance ownership does not exclude independent SQLite writers, so any observed change, including a new database, must refuse automatic restoration through the existing receipt contract. Preserve current databases and the existing actionable recovery guidance without adding schema or transaction receipt infrastructure.
Forward-port the conservative safety correction from c3c20e0241 onto the current capture owner.
Co-authored-by: RomneyDa <6581799+RomneyDa@users.noreply.github.com>
Co-authored-by: vyctorbrzezowski <51521767+vyctorbrzezowski@users.noreply.github.com>
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
* test(update): expect Doctor-time writes to disable automatic database restore
Rebasing onto main surfaced two receipts that relied on interval-wide
attribution: the NOCOW physical replacement (#160877) and Doctor's own
schema upgrade during legacy-root relocation. Under the conservative
policy neither is attributable, so both receipts are ineligible for
automatic restoration.
---------
Co-authored-by: RomneyDa <6581799+RomneyDa@users.noreply.github.com>
Co-authored-by: vyctorbrzezowski <51521767+vyctorbrzezowski@users.noreply.github.com>
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
Doctor requester-authority validation copied the entire SQLite database in a synchronous worker on every authority callback (195 snapshot launches, 31 s for the configured-owner case; candidate Doctor 74 s). Maintenance now owns one reusable live reader; admission, physical file identity, and revocation are re-checked on every reused read, and the reader joins maintenance and database cleanup. Configured-owner case 31.4 s -> 3.6-4.6 s; snapshot launches <= 8 during admission, zero during maintenance.
Refs #161949
Refs #161867
* perf(state): retire periodic runtime integrity scans
Remove delayed and daily full-database scans from the Gateway while preserving requested agent quick checks, live-owner confirmation, and quarantine. Keep full verification with admission, migrations, and Doctor maintenance. No config or schema migration is required.
* fix(update): restore Windows task autostart after cancellation
Carry the existing restoration phase through Windows task recovery so SIGINT fences forward work without rejecting compensation. Preserve executor and native task ownership checks before side effects.
The original 40-file CI shard reproduced 821 passes and one SIGINT failure; it passes all 822 tests with this change. Testbox focused tests and changed checks passed, and independent Codex review found no actionable findings.
* fix(plugin-sdk): keep updater path context private
Pin the legacy home-directory facade to its existing eight exports so new internal updater helpers do not become public SDK contracts. Reduce the wildcard ratchet by one without expanding export or callable budgets.
All 11 SDK surface tests passed against the exact failed CI merge plus this fix on Testbox, along with core/script types, targeted lint, export guards, and formatting. Independent Codex review found no actionable findings.
Resolve read-only worker launches through the running updater's retained
generation. Keep session reuse generation-aware and join child work before
retiring its runtime files. This protects future self-updates; an already
loaded 2026.9.6 driver retains its first-hop limitation.
Refs #161865
The async predecessor-receipt read introduced in #161813 can cold-start a
SQLite reader after the updater's launch directory becomes unavailable.
Resolve snapshot inputs before selecting a safe worker cwd, so captured
absolute paths keep working without rebasing unresolved relative paths.
Retain a valid physical cwd while replacing an overlapping installation:
Node worker bootstrap reads cwd before worker code can recover. Keep user
and session paths anchored to the captured invocation directory, forward
that directory for local package artifacts, and restore it after worker
settlement when it still exists. Keep backup and rollback ownership intact.
Concurrent config write during admission: warn and re-read through the
existing config owner; invocation-relative paths keep their original base.
Doctor sees its own finished update row: terminal ledger settlement remains
authoritative before post-update Doctor; original-state capture is preserved.
The installed updater owns cwd protection. A candidate cannot patch an
older driver before staging; this enables the next update performed by the
corrected driver and adds no new candidate marker or environment protocol.
Published driver x candidate: openclaw@2026.9.5 installs this candidate in
the base/manual-restart cell; Doctor, Gateway readiness, and preservation
of unavailable-plugin configuration pass.
Validation: 34 focused local tests and 69 focused Linux tests; real removed
cwd, relative-path, and package-artifact regressions fail without the fix.
The full architecture command, five source-contract checks, explicit
boundary lint, and independent review through P2 pass.
## What Problem This Solves
When a scheduled automation keeps failing, OpenClaw posts "Automation X failed N times" to the chat and waits for the user to ask for a fix, even when the fix is something the agent in the conversation that created the job could make itself.
## User Impact
When a job created from a conversation (it has an owner session) reaches its failure-alert threshold (default: 2 failures in a row, same cooldown and incident dedupe), OpenClaw sends one repair request to that owner conversation instead of the first chat alert. The conversation handles it as an ordinary agent turn, as if it had received a message: its own session and transcript, its workspace, its normal tool policy, and its reply goes to that conversation's own route (chat, thread or forum topic). The request tells it to:
- fix the problem in the workspace if it can (for example the helper script or instructions file the job follows) and say what it fixed in one line;
- otherwise ask the user for exactly what it needs;
- say nothing if the failure was transient.
Heartbeat settings do not apply to the repair: `heartbeat.target: "none"`, `isolatedSession`, `lightContext` and `activeHours` do not route, isolate, trim or hold it, and it carries no first-heartbeat notice.
The normal alert is still sent when the job has no owner conversation, for operator-only jobs (command payloads, on-exit and stream schedules), for webhook alert routes, and when the job fails again after the repair request (once, noting that a repair was requested). A successful run clears the streak silently.
**Upgrade behavior change (intentional):** repair is on by default. Opt-out is per job: `failureAlert: false` disables both the repair request and the alert. No config key is added. The chat "Automation … recovered" notice is removed; recovery stays in automation history. This reverses the notice added in #146582.
## Why This Change Was Made
The owner conversation is where these jobs get fixed in practice: in our setup a user had to reply "just fix it" to a failure notice, and the agent in that conversation then repaired the job. This change starts that step automatically.
The repair is dispatched through the Gateway's existing `agent` method on the instance's lifecycle principal (`dispatchGatewayLifecycleMethod`, the same owner restart-sentinel continuations and exec-approval follow-ups use), with the owner `sessionKey` and `deliver: true`, so the agent method resolves the reply route from that session's stored delivery context, thread included. The cron service reaches it through one injected dependency (`runCronFailureRepair`, wired in `server-cron.ts`), so `src/cron` does not import Gateway code. Nothing in the `sessions_send` tool needed extracting: its default path also calls `agent`, but with `deliver: false` plus its own announce flow, which is what this does not want.
An earlier revision carried the request on a heartbeat wake (system event plus `requestHeartbeat`) and patched the result with a `heartbeat: { target: "last" }` override. That made the repair inherit heartbeat behavior it should not have: an isolated `:heartbeat` side session without the conversation's transcript, light context, the internal-handling prompt, and heartbeat routing. This revision removes that path (system event, heartbeat wake, override, and their docs and tests); heartbeat behavior is unchanged for everything else.
**Tradeoffs (maintainer decisions):**
- The request is scheduler-authored, not a user message: it runs on the Gateway's system principal with `internal_system` provenance (`sourceTool: "cron_failure_repair"`), so it carries no sender-owner identity. Under the existing trust model the turn gets the conversation's normal non-owner tool policy (workspace and exec tools, not owner-only control tools such as `automations`); a change to the job itself goes through the user's reply turn. No new authority plumbing.
- The job's name, payload and last error are wrapped as untrusted data in the request.
- The request goes to the job's owner at dispatch time. If it is lost (the job was removed in between, or the `agent` dispatch rejects), nothing extra is sent at that moment; the job's next failure sends the normal alert naming the repair.
- A repair reply is posted at any hour, like the failure alert it replaces.
Persisted state: the incident gains an optional `repair: { atMs }` marker. It means "this incident's first alert became a repair request". It survives restarts, and the next alert of that incident clears it.
## Evidence
- **Size vs `origin/main`** (merge base `accde9e31a`, head `435e097ee7c`): production +233/−85 (net +148), docs +14/−4, tests/QA +649/−49; 21 files. To stay within the line-cap ratchet, the Gateway dispatch lives in `server-cron-notifications.ts` next to the failure-alert transport (2 lines of wiring in `server-cron.ts`), `reconcileCronExitWatchers` moved to `cron-exit-watchers.ts`, and two input helpers moved from the qa-lab mock server to `mock-openai-input.ts`.
- **Our production shape, Telegram Test Server forum topic, live OpenAI model** (`telegram-e2e-userbot`, Convex-leased credential, the lease's forum supergroup with a run-owned topic created before recording and deleted after; fresh isolated gateway from this checkout via `--source-gateway` on ports 19951/19952; provider slot = logging proxy to the OpenAI API). Config mirrors ours: `agents.defaults.heartbeat = { every: "1h", activeHours: 08:00–23:00 Asia/Kolkata, target: "none", directPolicy: "allow", to: <the tester's DM>, lightContext: true, isolatedSession: true }`, `cron.failureAlert` and `messages` unset. The QA user posts in the topic, which creates the owner conversation `agent:main:telegram:group:<forum>:topic:<n>`. A command action seeds `scripts/meeting-sync.md` (step 1: `sleep 150`), adds an isolated agentTurn job "Meeting sync" owned by that topic and announcing to it (`timeoutSeconds: 45`), and force-runs it twice: both time out (`consecutiveErrors: 2`, notification status `not-requested`).
- Command (secrets omitted): `E2E_MOCK_SERVER_PATH=<logging proxy to the OpenAI API> E2E_ROOT_CONFIG_PATCH=<live model + the heartbeat block above, messages: null> node <runner with forum-topic setup hooks> --backend mock --source-gateway --gateway-port 19951 --mock-port 19952 --chat <lease forum> --scenario <scenario> --timeout-ms 1000000 --record events.ndjson --output summary.json` → exit 0.
- Topic timeline (all messages):
| Elapsed | Topic message (SUT unless noted) |
|---|---|
| 0.8 s | user: `Hi! This topic owns our hourly sync automations. Reply with one short sentence confirming you are here.` |
| 61 s | `I’m here in this topic for your hourly sync automations.` |
| 272 s | `Fixed scripts/meeting-sync.md by removing the 150-second wait that exceeded the automation’s 45-second timeout.` (the repair turn) |
| 393 s | `Sync: 3 new meetings.` (third run, `ok`, `consecutiveErrors: 0`) |
| 514 s | `Automation "Calendar sync" failed 2 times` / `Cause: timeout` / `Run started: …` (the unowned control's alert) |
| 620 s | user: `What did you change in the meeting sync automation earlier? One sentence.` |
| 651 s | `I updated scripts/meeting-sync.md to read the CSV immediately instead of waiting 150 seconds, which exceeded the automation’s 45-second timeout; the schedule stayed unchanged.` |
- The repair ran in the topic session itself: `chat.history` of `agent:main:telegram:group:<forum>:topic:<n>` holds, in order, the owner-ready exchange, the repair request (user role, `internal_system`), `read` → `exec` → `exec` → `edit scripts/meeting-sync.md` → `exec`, the one-line reply, the next run's result, and the follow-up exchange. No `…:topic:<n>:heartbeat` session exists (sessions list). Its 6 main-model requests carried the brief with the topic's ordinary tool surface. The step file changed about 30 s after the second failure.
- No `Automation "Meeting sync" failed …` alert and no `First heartbeat alert` text anywhere (7 SUT messages). Control: an unowned job "Calendar sync" with the same broken step fails twice and alerts (delivery status `delivered`); its step file stays unchanged. The provider log also shows one `NO_REPLY` turn without the brief right after the repair's `exec` calls that sent nothing.
- **Telegram Test Server DM owner, live OpenAI model, `activeHours` excluding now** (same harness with `--dm`, window 06:00–07:00 Asia/Kolkata at 21:06 IST, otherwise the same heartbeat block). The DM creates the owner conversation `agent:main:main`; the same owned job fails twice (timeouts). DM timeline: 29.5 s `I’m here in the chat that owns your hourly meeting sync automation.`, 310.8 s `Fixed meeting sync’s timeout by replacing its 150-second blocking wait with a nonblocking export-readiness check.` (the repair turn; it chose to delegate the edit to a subagent via `sessions_spawn`/`sessions_yield`), 339.1 s `Sync: 3 new meetings.` (third run `ok`). The repair request and reply are in `agent:main:main`'s transcript; no alert, no onboarding notice, 3 SUT messages.
- **Regression test** (`src/gateway/server.cron-failure-repair.test.ts`, new file, "repairs an owned job with an ordinary owner-topic turn whatever the heartbeat config"; 4.3–5.8 s): a real Gateway with its real cron service and `agent` method (agent command mocked at the runtime boundary), heartbeat `target: "none"`, `isolatedSession`, `lightContext` and an `activeHours` window that excludes now; an owned job owned by a Telegram topic session fails twice; the agent command runs once in `agent:main:telegram:group:<g>:topic:42` with `deliver: true`, `channel: "telegram"`, `to: <g>`, `threadId: 42` and the repair request as its message, and no announce is sent. Fails on the previous head (no agent run).
- **qa-lab end to end** (`cron-failure-repair-owner-conversation`, real gateway child, qa-channel, mock-openai, isolated state): `OPENCLAW_BUILD_PRIVATE_QA=1 OPENCLAW_ENABLE_PRIVATE_QA_CLI=1 pnpm openclaw --profile <isolated> qa suite --provider-mode mock-openai --scenario cron-failure-repair-owner-conversation` → `passed=1 failed=0`. An owned job fails twice and no alert is sent; the owner conversation receives the repair request, its ordinary turn fixes the workspace step file (`write`), the reply lands in that conversation, and the request and reply are in the owner session's own `chat.history`. Control: an unowned job with the same failures alerts, and the repaired job never does.
- **Service tests** (`src/cron/service.failure-repair.test.ts`, 11 tests): threshold → one `runCronFailureRepair` request to the owner session (no system event, no heartbeat wake) and no alert; the job's name, payload and error sit only inside `<untrusted-text>`; the next failure sends one alert naming the repair, later failures nothing more and never a second repair; a rejected repair request alerts on the next failure; success clears the incident silently; no owner → the old alert; agentTurn, systemEvent and script jobs repair, command jobs and on-exit or stream schedules alert. The heartbeat-route case in `src/cron/service.wake-now-real-heartbeat.test.ts` is removed (the file is back to `main`).
- On `435e097ee7c`: `node scripts/check-changed.mjs` exit 0 (line-cap ratchet, all tsgo shards, lint, import cycles, dead exports); cron repair/alert/persistence/real-heartbeat, gateway cron, cron-exit-watchers, server-cron, notifications and qa-lab mock-openai suites pass (30 files, 538 tests); `pnpm config:docs:check` is clean.
- **Existing-state upgrade (recorded on an earlier head, harness not kept):** published `openclaw@2026.9.6` created an owner session, an owned and an unowned agentTurn job with an announce failure route and each job's first failure; its own `openclaw update --tag <candidate tarball>` succeeded and both streaks and owners survived; one more failure each reached the threshold: the owned job's incident gained `repair: {atMs}` with delivery status `not-requested` and its repair request reached the model; the unowned job alerted. The persisted state is unchanged by this rework.
- **Config compatibility:** no config key is added or changed (`config:docs:check` OK). The behavior change is the default for owned chat-alerting jobs, documented above; per-job `failureAlert: false` opts out.
- Downgrade: a persisted `repair` marker is an extra field older releases ignore; an incident that was mid-repair at downgrade keeps its signature there, so the older release dedupes that same-cause failure (no new alert until the cause changes or the job recovers), as it does for any already-alerted incident.
- **Owner boundary:** `owner` is fixed at creation (`CronJobPatch` omits it), so it cannot be reassigned; dispatch reads the live job and sends nothing if the job is gone.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
The v13 wide-row state migration conflated a missing installed_plugin_index row with a row whose JSON was unparseable or shape-drifted, and dropped the table in both cases. An invalid row is now preserved in full in the existing diagnostic_events quarantine with repair guidance recorded, valid install records are retained independently of damaged metadata, and a preservation failure rolls the step back to v12. No schema change.
Closes#161329.
Landed under the pre-existing-red rule: the remaining CI failures were current main reds at this base (Codex app-server settlement fixture drift, fixed on main by 52e60fb42b; the subagent kill-tombstone / descendant-cancellation intermittent first seen on main hourly 36619831492).
`openclaw memory status --deep` kept reporting `dirty=true` for agents that had a system-only cron-base session (`agent:<id>:cron:<job-id>`). Metadata catch-up counted that session as eligible, but the indexer correctly excludes it, so re-indexing never cleared the flag (eligible = indexed + 1). Status and catch-up eligibility now use the indexer's own admission gate. What gets indexed is unchanged.
Fixes#161823. Thanks @alfred429 for the report and the source-level reproduction.
Proof: the real CLI on an isolated profile, in FTS-only mode with no external providers.
- On main, the agent stayed dirty after repeated indexing (eligible 2, indexed 1). With this change it's clean (1/1).
- A new normal session still marks the agent dirty (2/1) until it's indexed, then it's clean (2/2).
- The regression test fails before the fix and passes after, and 40 focused tests pass.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(exec): avoid extra prompts for audit suppression commands
Remove the command-text gate and its dedicated approval/reviewer plumbing. Suppression configuration and audit filtering remain unchanged; normal exec policies still decide execution. Keep only the shipped deprecated SDK signature as a no-op.
Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>
* fix(exec): avoid extra prompts for audit suppression commands
Worked on by:
- @jesse-merhi
Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>
OpenClaw-Publication: f1fe3778-77d9-432e-b98f-01f680f2cdd3
* fix(plugins): preserve deprecated suppression approval predicate
Retain the shipped SDK result during infra-runtime retirement without restoring internal execution callers. Document the distinct runtime and plugin migration outcomes.
Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>
---------
Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>
* fix(sessions): reject compaction of missing sessions
Return an actionable INVALID_REQUEST instead of a successful no-op for missing targets in both compaction modes. Preserve explicit no-op outcomes for existing empty sessions.
* fix(sessions): reject malformed reset agent selectors
Use strict session agent input validation before lifecycle cleanup so an invalid explicit selector cannot reset the default agent. Extract reset-target resolution from the oversized lifecycle service while retaining selected-global targeting.
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Confirm installation-only pending records for disabled channels and
installed plugins without config entries. Carry retained plugin IDs
through package convergence while preserving real migration obligations.
Fixes#161288 and #160523. Thanks @akennedytog and @g62nkcx4q7-wq.
The Gmail watcher stayed down for good when a Gateway restart briefly found its port still in use (EADDRINUSE). It now retries the bind a few times with bounded backoff and recovers once the port frees. If the port stays taken, it reports a clear error and stops retrying.
Fixes#161467. Reported by @kazuyuki-eguchi.
Proof: a real isolated Gateway, with the Gmail watcher on local fakes.
- Before the fix, the watcher stayed down after a transient port conflict. After it, the watcher recovered and answered.
- With the port held for the whole run, the watcher stopped after the initial attempt plus three retries, with a clear error.
- The regression test fails before the fix and passes after, and 31 focused tests pass.
- Updating from the published 2026.9.6 build to this build succeeded in two fresh runs (102 s and 97 s), and the installed build passed both the transient and the persistent control. The first updater run failed at service activation and didn't recur. The watcher can't reach that step: rehearsal disables hooks, and this diff doesn't touch activation or lease code.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
`openclaw update` could hang after printing its result because the retained updater runtime waited on the SQLite broker with no bound, and a force-exit watchdog was unsafe while accepted operations were still settling. The broker close is now split: settlement of accepted operations, pending opens and live references must complete; native close and worker termination are bounded afterwards, with the bound expiry recorded as a warning and the exit still happening.
Closes#160690.
Landed under the pre-existing-red rule: the remaining CI failures were current main reds in the merge window (check:architecture import cycle from b36eb3e7b1, fixed by 2a0a65c4ae; update-candidate-canary and cron service tests, fixed by 5f76cc437d).
Build on main's separate caller/service projections from 1794d8b4ef.
Propagate freshly read configuration through admission and already-current
completion, keeping explicit channel/source requests authoritative.
Refresh matching legacy projections and candidate validation after root or
include saves, preserving invalid-config, schema, and live-authority checks.
Concurrent config write during admission: warn and reread current bytes;
a save alone never triggers the former configuration-changed refusal.
An implicit stored-channel change resolves the whole target again, discards
old staged/admission/confirmation facts, and rechecks its original executor
bindings before execution.
Cached candidate admission carries the source snapshot it inspected. A changed
snapshot revokes old checks and reinspects the retained candidate before any
activation shortcut; only the admission result is published to history.
Late Git admission revalidates against a retained runnable candidate before
checkout. Service preparation and activation notification occur once.
Doctor sees its own finished update row: real original-state capture and
terminal publication remain visible to subsequent Doctor as completed;
no history or terminal-state policy is weakened.
Retain the Git candidate and private repository through finalization. Retire
candidate, repository, then previous-runtime backup after verification;
cleanup warnings retain dependent resources and recovery backups.
Published-driver x candidate: the installed updater runs first, so this
admission repair applies when that updater is installed. The candidate
continues to support the run/driver markers shipped by openclaw@2026.9.5.
The published-driver survivor cell passed with manual restart. Original
captures remain manual evidence, distinct from verified rollback snapshots;
rollback preserves later operator writes and existing ownership checks.
Proof: representative attribution 55/60 -> 59/60 reversing only the original
capture change; Linux original cluster/siblings 271/271, invariants 205/205,
execution 113/113, and the published-driver survivor cell passed. Local broad
proof covered 85 files; five stale fixture contracts were corrected and the
12-file follow-up passed 168/168, including both new cache regressions.
Combined coverage: 1,496 unique cases, with no weakened assertions.
The new cache cases took 4.358s/1.983s; the real Doctor invariant took 22.186s.
These exercise real config, admission and ledger boundaries.
Final local check-changed, production/test types, typed and boundary lint,
all five source-contract checks, and full Knip passed. Independent P2 review
found no actionable issues. Full owning configs are deferred under the
approved incident exception; focused proof ran with at most three workers.
Reuse managed-service publication checks at each runtime publication boundary,
comparing the publisher's exact physical outputs. Live disjoint runtimes stay
available; observed shared consumers must be stopped before replacement.
Document that stopped sibling profiles retain their own state and may need the
updated CLI's existing Doctor migration before resuming. No schema version or
automatic sibling migration is introduced.
Validation: 67 focused tests pass; five overlap regressions take 202 ms.
Fresh managed review is scoped-clean through P2. Isolated Linux native proof of
the composed selector/runtime candidate covers disjoint, overlapping live,
and stopped-sibling publication with preserved conversation witnesses.
Preserve the signal-driven recovery path when startup triage declines. Use the existing supervisor identity to gate terminal failed recovery, retain one retry after confirmed cleanup, and print platform-specific manual recovery steps. Refs #159539, #160173.
Preserve the caller Doctor projection separately when an explicit channel request selects another managed profile. Keep service capture, target selection, runtime materialization and writes owned by the selected profile, and retain strict source-bound schema validation.
Retain both published legacy convergence regressions, correct their file-backed environment fixture, and cover refusal without an explicit channel request. Extract admitted update orchestration without changing its body so the command stays within the existing line limit.
Canonical missing-copy handling in the shared archive walker returned no warnings, so an operator whose session archive copies were deleted got no signal. Missing copies are now warning-level facts (count plus up to five sample paths); blobs stay retained and migrations complete.
Refs #160770 (the reported stack overflow remains unreproduced; reporter asked for the full stack).
Landed under the pre-existing-red rule: remaining CI failures were current main reds (github-publication-personal-pending, server.sessions.list-changed, server.heartbeat-store-lifecycle.product-proof).