Apply the existing candidate flag before session catalogs, worker recovery,
and background lifetimes are prepared. Keep configuration, ownership,
schema, migration admission, and plugin runtime validation intact.
Prove the startup boundary and document manual recovery from the published
updater's shared five-minute deadline. The published 2026.9.4 driver upgrades
the synthetic ten-agent state successfully, with candidate readiness in 6.134s.
Thanks to @ZaneMc-AO and the other timeout reporters.
Related: #144858, #154381
Finalization waits through the existing readiness owner and defers Doctor, config, and plugin writes only for the same verified Gateway holder. Preserve ordinary maintenance for dead or unverified owners and retain pending recovery history.
* fix: finish updates after foreground Gateway shutdown
Wait for the same foreground state owner to settle before delegated update Doctor enters maintenance. Preserve the updater fence, physical lifecycle coordinator, five-minute bound, and manual restart ownership.
* fix: retain the full foreground restart settlement budget
Derive Doctor admission wait from the installation tick, restart drain, and shared service-stop allowances. Protect a legitimate slow foreground shutdown beyond the initial five-minute lifecycle timeout.
* fix: join foreground owner cleanup before update maintenance
Give updater-owned Doctor the existing shutdown reserve when the Gateway removed its owner row before releasing the physical lock. Reject newly appearing owners, cap retry sleeps to the remaining reserve, and use the repository deferred helper under ES2023.
* docs: clarify published foreground update recovery
Document that published 2026.9.5 has no installation-replacement watcher and keeps the supported operator stop, update, and launch sequence. Preserve the failed active no-restart cell and the existing update repair recovery path without granting new stop authority.
* test(cron): use shared fixture ownership for pacing cases
Retire the legacy per-case state directory and manual finally cleanup. The existing harness owns unique SQLite store partitions and file teardown drains shared database resources. Keep every pacing assertion and the fixed clock while avoiding deletion of a database directory before its owner closes and preserving test-body failures for Vitest.
* fix(plugins): report incomplete installs before capability consent
* fix(plugins): repair install-health test and type boundaries
* fix(doctor): preserve recorded plugin recovery selectors
* fix: refresh stale Gateway stop policies before maintenance
Main recognized historical systemd timeouts, but maintenance could stop the
resident before repair, and Doctor and already-current or no-restart updates
could preserve stale computed policy. A published 2026.9.5 resident also retains
its startup shutdown budget after the unit changes.
Refresh owned policy through the existing definition-mutation and backup owners,
confirm daemon reload, preserve operator drop-ins, and retain restore/reload/input
retirement ordering. Share maintenance between update and Doctor, and warn when
a non-stopping refresh still has a short effective manager timeout.
Publish process-owned shutdown budgets and lifecycle write-custody facts. Reuse
the suspension owner and existing update deadline: stop when idle, warn and stop
at the deadline for ordinary work or unknown custody, and refuse only current
reported write custody with its exact owner phase. Reread native policy when the
new Gateway accepts shutdown without resetting its elapsed budget or watchdog.
Thanks @ezimerman for the installed-unit and shutdown evidence.
Fixes#153153. Refs #150898, #152879, #153017.
* fix: preserve the resident shutdown budget in Gateway status
The kernel request-context adapter copied host lifecycle control methods but
dropped the recorded shutdown-budget getter. Real Linux package proof observed
a 325-second startup budget while status omitted it, forcing maintenance onto
the legacy unknown-budget path.
Forward the live getter through the existing adapter without changing request
authority or adding another budget owner. Add a real-kernel registered-status
regression for short and adequate budgets; the corrected test fails on the
original adapter and passes with the forwarding line.
Refs #153153.
* test: control the Gateway shutdown clock consistently
Restore the explicit node:perf_hooks performance import for the run-loop regressions that retain their own fake-clock assertions. They must control the same monotonic clock as the production shutdown-budget owner. Keep the deadline assertions, timers, and production behavior unchanged.
* fix: preserve Doctor legacy reads during Gateway preflight
The stale-Gateway probe opened the canonical owner-lease database without
Doctor's existing legacy-catalog admission. With a built install and a busy
Gateway port, that read created WAL/SHM files and rejected a supported repair
before maintenance could stop the Gateway.
Carry the existing read admission through restart inspection to the lease owner.
Keep ordinary restart validation unchanged and preserve canonical artifacts.
Use synthetic port, build, and reachability facts in the regression fixture while
retaining real lease reads and byte-preservation assertions.
Refs #153153.
* refactor: keep Gateway maintenance owners within line limits
The L903 stop-policy repair exceeded the existing line-growth gate after
composition with current restart and service identity handling. Move request
upgrade policy into its existing request owner and service revalidation into
one sibling implementation without changing their behavior.
Preserve original request admission time and Doctor's statically primed
maintenance facade across package replacement. The bounded drain, recorded
custody-only refusal, warning policy, and native backup/reload ordering remain
unchanged. No options, schema changes, suppressions, or baseline growth.
Validation: 398 focused tests, core typecheck, line-growth ratchet, and fresh
Codex P1 review passed. Full changed checks continue on the frozen source.
* test: retain backup custody coverage through the archive walker
The rebase incorporated the maintained archive walker, but the L903 custody
regression still referenced the retired tar-create mock. Use the existing
walker mock bound to the real backup command without changing assertions or
production code.
Validation: 132 backup, migration, cron, coordinator, and suspension tests
passed. Fresh Codex P1 review is clean. Full changed checks continue.
* fix: preserve maintenance lifecycle and wrapper contracts
Doctor fixtures advertised a running native service without the matching
resident identity, effective policy, or lifecycle readiness response. Supply
those facts through the same RPC and native-query boundaries used by the real
maintenance owner, including fresh readiness without a resident budget.
Keep exact suspension assertions current with the additive custody category.
Install the model-acquisition fixture's manager after normal PATH setup so
startup and shutdown observe the same policy under the original deadlines.
Include the custody owner in the canonical PR-wrapper source inventory.
Move the unchanged Doctor inspection assertion and shared fixtures into their
existing policy/support owners to preserve the line-growth ratchet. Do not
change the bounded-deferral, warning, or reported-write-custody refusal policy.
* fix: preserve Stop ownership through shutdown budget refresh
An asynchronous systemd budget read could resume after Stop captured a
foreground updater and re-arm the hard-exit worker. Keep watchdog admission
with the run loop's current successor owner, and represent an absent cleanup
deadline with no process-cleanup budget.
Preserve foreground no-op service ownership and the full native allowance
after verified parking. Keep one fake clock for boundary tests, prepare
fixture operations before held scripts, and join cancelled fixture work
before the next case. Move unchanged budget cases into their support owner.
Retain the ruling that unknown custody cannot block an update and only
reported write custody may refuse maintenance. Preserve all existing test
assertions and limits. The full Linux changed gate, 790 focused Linux tests,
and a fresh P1 review passed; local host limitations are recorded in the PR.
* refactor(doctor): extract update-run admission from doctor-maintenance
* fix: preserve native policy and write custody during maintenance
* fix: keep shutdown integration within source and type gates
* fix(test): own survivor model endpoint before baseline setup
---------
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
## What Problem This Solves
Fixes slow `update.status` polling when the shared SQLite database is large. The active-run and latest-run reads each prepared a full private database copy before selecting one row, repeating the work on every poll and multiplying it across concurrent viewers. Follow-up to #153660.
## User Impact
Status reads the two current run records through the existing database owner without copying an already-open database. Cold reads still prepare an artifact-preserving snapshot. Slow requests now log phase durations when diagnostics are enabled.
## Why This Change Was Made
Measurement identified history reads as the expensive phase. Install identity caching and concurrent checkout-refresh sharing already exist; this change leaves their lifecycle and timeouts intact. The ledger reader now supplies the fixed two-row status projection through its existing read-only owner. General history list/get requests continue using the worker. No new cache, configuration, schema, or updater execution behavior; sentinel and reconciliation contracts are preserved. Net production growth: 51 lines, including diagnostics.
## Evidence
Same-process comparison on Blacksmith Testbox (`blacksmith-testbox`, profile `openclaw-check`), lease `tbx_01m30m5gtt27p3z5ggbha72nja`: registered status handler, real SQLite/ledger, 512 MiB of synthetic unrelated payload, three sequential calls per reader. The before arm uses the original two awaited `listUpdateRunsAsync` calls; the after arm uses the new projection. Install-channel and sentinel dependencies are stubbed to isolate history cost.
| Call | Before | After | Full database copies |
| --- | ---: | ---: | --- |
| First | 1,368.85 ms | 1.42 ms | 2 → 0 |
| Second | 1,001.43 ms | 0.74 ms | 2 → 0 |
| Third | 1,093.22 ms | 0.92 ms | 2 → 0 |
Median: 1,093.22 → 0.92 ms (99.9% reduction in this fixture). The first slow warning attributed 1,362 ms to history and 1 ms to reconciliation. An independent run of the original reader before the production change measured 1,494 / 1,037 / 919 ms on lease `tbx_01m30k0nc6afybgwzj0r3tcgcx`.
The concurrent/fresh-status regression failed before the fix with six full copies across three calls, then passed. Existing cached-identity and shared-refresh tests passed. Coverage also includes cold WAL/artifact preservation, inherited snapshots, committed reads during a writer transaction, phase timing, disabled diagnostics, and preservation of errors/responses when logging fails.
<details>
<summary>156 passing tests, validation commands, and test costs</summary>
All 156 retained focused tests passed (plus the temporary benchmark), using:
(focused vitest command for the four update-status/update-run test files; see the test list below)
All selected changed-file gates completed across the validation runs. On `tbx_01m30m5gtt27p3z5ggbha72nja` ([backing run](https://github.com/openclaw/openclaw/actions/runs/35546226270)), the base-pinned `node scripts/check-changed.mjs --base f0f31b4208 --head f0f31b4208 -- <seven changed paths>` passed core and all core-test typechecks, boundary checks, formatting, ratchets, and dead-export scans, then caught two missing brace blocks in the new test code.
After fixing those style errors, the finishing command on `tbx_01m30pejcb6ez8amcbpnje68k6` ([backing run](https://github.com/openclaw/openclaw/actions/runs/35548385015)) exited **0**: full six-file Oxlint (zero warnings/errors), native state schema guard, database-first/media/sidecar/import-cycle/webhook/pairing guards, the changed test's `gateway-methods` typecheck, and these separate single-worker test runs:
| File | Tests | Command wall time |
| --- | ---: | ---: |
| `src/gateway/server-methods/update-status.test.ts` | 10 passed | 18.09 s |
| `src/gateway/server-methods/update-effective-channel.test.ts` | 11 passed | 10.14 s |
| `src/infra/update-run-reader.test.ts` | 19 passed | 14.60 s |
Each was measured through the repository test wrapper with a single worker and dot reporter.
The baseline lease ([backing run](https://github.com/openclaw/openclaw/actions/runs/35545287134)) required dependency installation and was replaced after two file-sync timeouts; a later sync timeout required the finishing lease. No timed-out sync ran tests. The first default changed-file check selected unrelated files against the lease's moving `origin/main`; final validation pins the actual candidate base `f0f31b4208` and the seven changed paths. All three leases were stopped after proof. No build or full suite ran on the Mac.
Independent Codex autoreview: scoped-clean through P2; subsequent source edits only added the two test brace blocks required by lint. The temporary large benchmark is excluded from per-PR tests and the commit.
</details>
Limits: synthetic handler measurement, not a live deployment or confirmation of the reported 37–52-second production latency. No production restart/deployment or published-driver × candidate installation test was performed; this change only affects the read projection and diagnostics, with no state format or update execution changes.
<details>
<summary>Landing CI repair and verification (2026-09-21)</summary>
The original head failed `checks-node-compact-large-2`, job `106182679540`, in `update-effective-channel.test.ts` (`reports slow status phases`, diagnostics disabled): its process-wide `performance.now()` assertion counted two legitimate SQLite coordination deadline reads. A cold-database reproduction captured both calls at `tryAcquireSqliteCoordinator`; the original assertion failed with exactly two calls. This is a PR-added test defect, not an updater behavior defect.
The test now uses a call-through spy to verify that disabled diagnostics constructs no stage-timing tracker. Real ledger reads and the existing response and warning assertions remain intact. The repaired assertion passes the same cold-database scenario, and deliberately forcing tracker construction still makes it fail. No production changes were needed for this CI repair.
On Blacksmith Testbox lease `tbx_01m30x9ycx2h1wtsas2nrvfdbd` ([backing run](https://github.com/openclaw/openclaw/actions/runs/35554888144)):
- The original affected group passed on pure main `894fb140a2` (125 files; 212.84 seconds through the shard wrapper).
- All seven candidate file hashes were verified after correcting the remote shell's working directory and installing the synced checkout's own frozen dependencies. Earlier runs from the hydrated main checkout are not counted as candidate proof.
- All 40 focused tests passed across `update-effective-channel.test.ts`, `update-status.test.ts`, and `update-run-reader.test.ts` through `node scripts/run-vitest.mjs ... --maxWorkers=1`. The changed file's measured command wall time was 12.34 seconds. The original CI gateway-methods segment took 146.62 seconds for 118 files.
- `node scripts/check-changed.mjs --base HEAD --head HEAD -- src/gateway/server-methods/update-effective-channel.test.ts` passed, including the gateway-methods test typecheck, lint, formatting, boundaries, and dead-export scans.
- Independent Codex autoreview of the repair was scoped-clean through P2. The temporary reproduction and mutation probes are excluded from the commit. The lease was stopped after validation.
The repair was folded into the existing single commit with its subject unchanged. No GitHub job was rerun to clear the original failure. Exact-head [CI run 35556160577](https://github.com/openclaw/openclaw/actions/runs/35556160577) is green at `c74340fcb8ea999fde49db2f74bb0104757f8840`: all 160 jobs completed successfully, both required `openclaw/ci-gate` entries passed, and the previously failing compact-large-2 job passed in 424 seconds. Native preparation accepted this hosted evidence without rebasing or changing the head.
Maintainer direction: Peter explicitly directed landing this PR and authorized the native admin route if a review requirement blocks. Published-driver × candidate installation proof remains unrun; the change preserves schemas, config, and updater execution and only changes status reads and opt-in diagnostics.
</details>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Observe the installed Gateway after update finalization fails, once child work and service restoration have settled. Reuse the shared readiness owner, preserve the original failure and unsafe-restart constraints, and atomically record observed recovery through the existing update ledger before terminal output and public reporting.
Preserve installed service identity across rollback, first-terminal-writer receipts, and main's maintenance custody and publication owners. Include the previously approved measured startup-budget refresh without increasing its growth or variance allowances.
Validation: exact-head CI passed; the campaign Codex review is scoped-clean. Focused integration and ledger regressions cover restoration, observation, publication, recovery identity, and authority. Installed recovery evidence remains historical; no final-head live recovery run is claimed.
Thanks to @0-danielviktorovich-0, @wlassalle724, @neangneatos, @ZSmallX, and @hxy8241 for the reports.
Fixes#152811, fixes#152788, fixes#152742, fixes#153000, fixes#152974.
Doctor now recovers verified session imports when the original legacy index is missing, a restored source has a different inode, or the import database has been replaced. The receipt owner verifies hashes and ordered canonical transcript events before refreshing existing receipt evidence; conflicting sources remain protected, and current session settings and deletions remain authoritative.
Gateway readiness can admit usable SQLite stores while retained inputs await Doctor repair. Readiness does not grant archival authority. Recovery reports include remaining issue codes without claiming successful validation.
Published 2026.9.4 and 2026.9.5 updater cells cover missing indexes, copied sources, and replaced databases, preserving transcripts, hashes, schemas, and SQLite integrity. Reordered replacement history remains refused. No schema, dependency, configuration, or CLI option changes.
Thanks @cpsleepy, @mmm7053455-tech, @blackeyes-boy, and @Albertyn87 for the reports and recovery evidence.
Fixes#152884.
Fixes#153606.
Fixes#153619.
Follow-up to #153097.
* fix: avoid local health timeouts for wildcard Gateways
* test: consolidate port listener fixture literals
* fix: preserve family-scoped port confirmation
* refactor(code-mode): execute JavaScript with typed API discovery
Keep schema-derived tool declarations available to agents while removing TypeScript compilation, optional preflight, and language selection from Code Mode. Retire the saved languages setting through shared Doctor/startup migration, preserving activation and limits. Existing TypeScript cells must be rewritten as JavaScript.
Related: #153889
* fix(code-mode): complete JavaScript cutover checks
* docs(config): regenerate JavaScript-only Code Mode baseline
* test(code-mode): match public names in live validation errors
* test(code-mode): avoid shadowing the selected call
* test(code-mode): allow fixture rereads in live evidence
Updates could report doctor-failed after successful checks when slow plugin inspection cleanup kept candidate Doctor open until the rehearsal deadline.
Publish completed checks before joining disposal inside the private state snapshot, and supervise cleanup in an owned worker. Preserve completed results with a duration warning only after confirmed cleanup; incomplete checks, real findings, independent unsuccessful exits, and caller cancellation remain failures. Give each validation process its existing configured or state-derived allowance. The candidate-side containment also works with the published 2026.9.4 updater, whose aggregate five-minute limit remains unchanged.
Fixes#153541. Thanks to @sebastian-hoebarth for the report and diagnostic bundles.
Evidence: exact-head CI gate passed; campaign Codex review scoped-clean; recorded Linux published-driver and candidate-to-candidate managed updates succeeded with a slow disposer, while incomplete checks preserved the running installation. Windows taskkill-shaped results are regression-covered; native Windows execution remains unproven because the supported Testbox provider cannot run that target. No persisted schema, config, CLI option, or dependency changes.
Verify macOS Gateway shutdown before reporting stop success. Both stop modes now use the launchd stop owner to unload the LaunchAgent and verify the original and subsequently observed process IDs have exited; --disable also preserves the disabled policy across logins.
Revalidate delegated update authority before native writes and stale-process cleanup. Improve profile-aware update and Doctor recovery commands without changing admission predicates, stored data, backup, rollback, or restoration ownership. Correct the receipt-compensation fixture so its synthetic process exits with fake bootout.
Related: #153704. Thanks @britrik for reporting the macOS shutdown and repair-admission failure.
Evidence: 429 initial targeted tests; 290 follow-up tests after the receipt fixture correction; isolated native macOS before/after stop proof, including a six-second drain; scoped-clean Codex review and exact-head CI. The reporter's exact default-stop case was not reproduced; deterministic regressions cover a loaded label or surviving PID after bootout.
Closes#152540
Credit to @yetval for the report and CLI reproduction.
## What Problem This Solves
Fixes: node approval documentation tells operators to run `openclaw approvals --node`, which the CLI rejects because `--node` belongs to the `get` and `set` subcommands.
## User Impact
User impact: operators can copy working commands to inspect or replace a node's exec approval policy from the Gateway.
## Why This Change Was Made
Both affected passages now use the canonical `approvals get --node` and `approvals set --node` forms, include the required `--file` or `--stdin` input for replacement, and link to the approvals CLI reference. The CLI contract is unchanged.
## Evidence
Head SHA: `a565f41297f4eb917d6bd7b8b941406642042d14`
- Red before: `pnpm openclaw approvals --node abc123` exited 1 with `OpenClaw does not recognize option "--node"`.
- `pnpm openclaw approvals get --help`: passed and exposed `--node <node>`.
- `pnpm openclaw approvals set --help`: passed and exposed `--node <node>`, `--file <path>`, and `--stdin`.
- `pnpm docs:check-mdx`: passed for 1,313 files.
- `pnpm docs:check-links`: passed with 14,114 internal links checked and 0 broken.
- `node scripts/check-changed.mjs -- docs/cli/node.md docs/tools/exec-approvals.md`: passed the docs-only lane.
- Changed-file `oxfmt --check` and `git diff --check`: passed.
`pnpm format:docs:check` could not run repository-wide on Windows because the generated formatter command exceeded the operating system command-line length. The same formatter passed both changed files, and the changed-file gate passed. No live paired node was used because this docs-only repair changes command spelling, which was verified against the real CLI command tree.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(browser): preserve snapshot identity through canceled captures
Keep native DOM bindings tied to their captured controls, fence canceled
snapshot writes, and preserve visible descendants under ignored AX roots.
Consolidate snapshot capture under the existing role snapshot owner.
Keep MCP labels scoped to their documents, finish cleanup before publishing,
and discard ambiguous cross-document IDs. Retire refs before every snapshot
refresh, including hidden and failed captures.
Preserve healthy sibling connections during tab closure, ignore superseded
extension bootstrap timeouts, deliver node timeout diagnostics through the
CLI, and recheck node access before browser actions.
* test(browser): preserve request budgets in CLI timeout assertions
Keep failed update operations identifiable in reports and issue titles while preserving redaction of private command arguments, paths, and diagnostic text. Separate stable step identifiers from command text in package, Git, candidate-validation, and CLI producers, and clear stale diagnostics after a successful retry.
Use one released-key adapter for ledger writes, restart sentinels, and report joins so restored readers retain exact package/fetch keys and measured exit codes. Preserve authoritative recovery keys and existing update admission, execution, rollback, schemas, defaults, and operator options.
Related: #153230, #152193. Thanks to @marcio-absmartly for identifying the indistinguishable failure reports.
Validation: scoped-clean recorded Codex review; green exact-head CI; 196 Gateway run-loop and 100 package recovery tests; changed-file gates; published 2026.9.4 reader replay against unchanged production bytes. Historical installed-runtime evidence covers upgrade, rollback, retry, and rerun at its recorded source head.
Closes#153730
## What Problem This Solves
Fixes: directory commands infer a configured channel and can return its contacts when `--channel` is explicitly empty or whitespace-only, such as an unset shell variable.
## User Impact
Only omitting `--channel` requests automatic selection. Explicit blank values now fail with `--channel must not be blank` before command bootstrap and directory setup or lookup. Scripts that previously passed blank values must omit the flag instead. Valid channel names, aliases, padded names, account defaults, and documented inference remain unchanged.
## Why This Change Was Made
Directory option parsing now uses the same strict selector rule as message and channel-auth commands. Those existing consumers share one validator without changing their validation positions or precedence. General channel normalization and durable plugin setup behavior are unchanged.
This rejects invalid input before command bootstrap, not before every process side effect: the CLI can still read best-effort proxy configuration before parsing options.
## Evidence
- Supported `pnpm openclaw directory` before/after proof on isolated current-main builds: an explicit blank channel returned the configured IRC contacts before the fix and a CLI error afterward. The six changed files match the prepared PR.
- Fourteen real CLI cases per phase passed, covering all four directory leaves, omitted/explicit/padded selectors, whitespace, `--channel=`, text and JSON output, invalid configuration, blank accounts, and both selector orders.
- Native startup traces show candidate blank selectors reject before the command's config-ready/plugin-registry stages; omitted and valid controls still enter them. Config bytes and canonical plugin installation records/artifacts remained unchanged in the explicitly enabled synthetic fixture. No external network request occurred.
- 168 focused directory, message, and channel-auth tests passed. Eight blank-channel cases fail on the original directory registration; existing blank-account controls pass. Shared table coverage replaces duplicated mock assertions and repeated precedence cases.
- Scoped type-aware lint, formatting, whitespace, both source line-limit ratchets, and executable-runtime builds passed. Required hosted CI covers full static/matrix integration.
The IRC directory is config-backed; no live IRC transport, Gateway, or zero-config-read claim is made. Runtime proof uses current-main builds with matching changed owners, not a full replay of the PR's older base.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(update): hand foreground replacement to the CLI owner
Keep the Gateway available through validation, close its installed module
graph before publication, and let a fresh successor verify completion.
Route RPC updates through the existing detached updater and retire the
duplicate synchronous finalizer and unconsumed private activation receiver.
Require current foreground ownership throughout admission, preserve legacy
native handoffs, and resolve session aliases at the RPC boundary. Preserve
actual same-revision outcomes, join cancellation, and retain backups while
replacement verification is pending.
* fix(update): refresh verified Git target admission
* fix(update): preserve promoted Git runtime freshness
Preserve candidate runtime timestamps during Git activation so copy traversal
cannot invert the build and runtime postbuild stamps. Ordinary CLI commands
keep the validated runtime identity, while genuinely newer builds still trigger
postbuild.
Cover real Git inventory and filesystem promotion with fresh and stale runtime
stamps, plus the ordinary --version launcher path. Validate 218 owner and sibling
tests, the standalone two-case regression, changed-code gates, and P2 review.
* fix(protocol): accept existing update acknowledgment fields
Declare the acknowledgment fields already emitted by update.run, regenerate the Swift model, and validate captured handler replies against the canonical result schema. Clarify that acknowledgment is distinct from durable completion.
* fix(update): refuse foreground updates without respawn
Reject OPENCLAW_NO_RESPAWN foreground requests before handoff with actionable recovery guidance, and recheck the policy after acknowledgement and before parking. Preserve rejected foreground parking decisions in the existing helper activation fence, including refusals arriving after the notice timeout; managed-service notices retain best-effort behavior.
* fix(update): retain foreground stop before closure
* test(update): consolidate refusal and lifecycle fixtures
* refactor: remove canary launch argument mutation
* fix: hosted Stop bypasses foreground updater settlement
Successful update repair now reports the historical run IDs whose acknowledgments were newly recorded, while preserving failed outcomes. The existing ledger writer returns whether it inserted the acknowledgment, and both repair paths use that result for JSON reporting.
Additional captured historical candidates go through the existing full finalizer. The captured 100-row history, 30-minute lightweight boundary, and live-updater protections remain intact. No schema, configuration, CLI option, or dependency changes.
Thanks to @IWhatsskill for the original report and discussion.
Validation: 235 focused tests across seven affected files and 1,183 routing tests passed on an isolated Testbox; scoped-clean Codex review and exact-head CI passed. Published-driver/candidate maintenance recovery remains a documented validation gap.
Related: #153228. Complements the historical-selection fix in #153261.
Co-authored-by: IWhatsskill <284122573+IWhatsskill@users.noreply.github.com>
## What Problem This Solves
Connects supported installed OpenCode, Qwen Code, Pi, and Kilo Code agents to Models settings, setup, the shared model picker, and chat. Each CLI owns its sign-in; this does not discover every arbitrary ACP-compatible executable.
## User Impact
- Enabling an installed agent finishes at the configuration acknowledgement, without waiting for model discovery.
- Native model choices survive reload. Disabling an agent stops new discovery and turns without interrupting admitted work or deleting history.
- When optional chat restrictions cannot be enforced, an administrator can explicitly choose **Continue for this chat** to use the native app's permissions. Required sandboxing and other mandatory boundaries remain enforced; agent and global defaults do not change.
- Continue retries the original refused message and attachments once. A newer draft is preserved rather than sent under an earlier confirmation. Native assistant output streams before the final response.
- Explicit permission and sandbox preferences survive same-chat resets and expiry. Native consent remains incarnation-bound and is retired on reset, fork, runtime change, or stronger settings.
- Empty native discovery does not remove unrelated API models or trigger false model-fact changes, including repeated empty reads.
## Why This Change Was Made
OpenClaw owns model selection, consent, transcripts, session lifecycle, and run authority. ACPX owns transport, model inspection, session-scoped client permissions, and final prompt admission. This consumes upstream ACPX [#622](https://github.com/openclaw/acpx/pull/622), [#623](https://github.com/openclaw/acpx/pull/623), and [#624](https://github.com/openclaw/acpx/pull/624), rather than maintaining private enforcement wrappers in OpenClaw.
Both dependencies now use published **acpx 0.17.1**. The temporary Git-source pin and source-build approval are removed. Exact, dated dependency cooldown exclusions remain; installer integrity checks and script-disabled managed installation are unchanged.
The final reconciliation keeps catalog requests with the existing catalog schema owner, removes an unused type facade, and preserves native observation identity at the merge producer. Native policy calculation lives in the existing execution-policy module instead of importing the full agent runner during preflight. Runtime availability, implicit fallback, and explicit native pins stay with the execution selector; chat preflight checks only the known host-only runtime's permission restrictions. Tests exercise real catalog invalidation rather than retired UI wording or private metadata.
## Evidence
- **Published-driver upgrade:** the actual npm `openclaw@2026.9.4` updater installed the standard candidate package. Staging, migration rehearsal/continuation, canary, swap, and Doctor succeeded. The existing session retained its ID, model, runtime, workspace, label, and permissions. Restart was explicit after `update --no-restart`; automatic OS-service restart is not claimed.
- **Managed plugin and native continuity:** the unchanged candidate plugin archive installed through the repository's canonical local prepublish npm registry fixture with normal official-package provenance. Its ACPX dependency matched public npm 0.17.1, and its SDK resolved the installed candidate core—not the source checkout. A native conversation continued across a real Gateway process restart. Reset cleared consent; a stale lifecycle grant was rejected without peer effects; fresh explicit consent restored execution. The peer was a deterministic ACP subprocess, not paid-provider inference. The candidate OpenClaw plugin itself is not publicly released.
- **Telegram Test Server:** six real-user sends and two callback selections produced native-runtime confirmation edits. Four persisted snapshots verified native provider/model/runtime selection, `/reset` retention, restart retention, and unchanged agent defaults. Zero inference requests. The unchanged QA runner used Bun after a host Node/Undici error; the product Gateway used Node and the built candidate. Leases, credentials, processes, and listeners were cleaned up.
- **Real provider:** OpenCode through ACPX 0.17.1 returned the requested exact smoke response using `opencode/big-pickle`, an isolated HOME, and no API-key environment variables.
- **Real UI/Gateway:** a synthetic ACP subprocess refused the first message under Read Only, automatically retried it after confirmation, and streamed assistant chunks while its final response was still held. The recording contains one session creation and one automatic `chat.send`. [Inspected before/after screenshots](https://github.com/openclaw/openclaw/pull/150224#issuecomment-5748103787) are embedded in the PR discussion and originating chat; [earlier consent/retry captures](https://github.com/openclaw/openclaw/pull/150224#issuecomment-5743925304) remain available.
- **Regression proof:** repeated empty native discovery failed before the producer fix: captured API rows disappeared, and configured API rows emitted false invalidations. Both cases pass after repair. A successful empty provider refresh still removes its own obsolete rows; native cancellation, final-prompt authority, and API/native route separation remain covered.
- **Lightweight preflight:** the existing lazy-handler regression reproduced a speech-runtime import during basic `chat.send` validation. Moving the shared policy implementation fixed it without changing policy decisions: all nine moved function bodies and denial-prompt initializers have identical tokens. The registered consent-refusal test, 25 native-process cases, 18 model-ownership cases, 168 harness-selection cases, and seven relocated pure policy cases passed. The seven policy cases were moved, not added; the oversized selection module no longer needs its old size exception.
- **Single availability owner:** hosted creation/recovery regressions exposed a duplicate availability check in chat preflight. Removing it restores inherited and recovered native selections without inventing fixture runtimes or changing stored pins. Four title/initial-send selection paths, tombstone recovery, and refusal before message persistence passed unchanged. Temporary core-runtime fixture pins were removed; execution still rejects unavailable explicit native ownership rather than silently switching to an API provider.
- **Focused checks:** more than 700 cases passed across the repair rounds, including registered native-process execution, Gateway creation/recovery/admission/compaction/queued Stop, protocol exports, plugin boundaries, catalog ownership, Telegram callbacks, and 12 browser scenarios. The policy-cutover tree passed a full build and core production/UI/core-test typechecks; the subsequent availability-owner correction passed focused execution, production types, and lint. Protocol generation/Swift/Kotlin checks, assertion safety, and runtime import-cycle checks passed. Startup JS measured 364,119 bytes gzip against the 364,774-byte limit. The complete protocol schema's serialized bytes stayed identical through its module relocation. Size exceptions were pruned, not relaxed.
Upgrade, Telegram, and UI recordings bind the compiled `99ecb40e5d` artifact. Later ACP changes affect empty catalog observations, schema placement, metadata-only policy loading, and test fixtures. Consent decisions and denial prompts are unchanged; native execution and pre-admission refusal were rechecked after the policy move. The final main merge also carries upstream update-service planning changes; retained upgrade proof still exercises the unchanged published 2026.9.4 driver with explicit no-restart, not a new automatic-service-restart claim. Historical observations are not relabeled as reruns of a later commit. Final exact-head review and hosted checks remain separate gates.
### Disclosed limits
- Direct local `npm-pack` installation imported successfully but correctly lacked official-plugin storage trust for native execution. The normal npm installation through the canonical prepublish registry fixture closed that execution gap without trust overrides.
- An optional custom `--bundle-plugin acpx` distribution failed Darwin staging in `fsevents@2.3.3` (`binding.gyp` missing). No flags or packaging code were weakened; the standard candidate installation remained intact. That custom-distribution route is not claimed green.
- Matching-backup restoration and a post-upgrade external MCP tool call were not exercised.
## Compatibility and intended permission contract
The maintainer-requested behavior is explicit administrator-only, per-chat delegation to the native app when optional restrictions cannot be enforced—not silent global relaxation or a second OpenClaw enforcement system. Mandatory boundaries, live run authority, and consent retirement remain required. The accompanying SDK additions use the existing harness, transcript, and turn-stream owners; existing silent transcript-write defaults remain unchanged.
The candidate plugin declares `compat.pluginApi: >=2026.9.5`; supported installation uses that SDK floor rather than relying on the older general install hint. The package proof upgraded the host before installing the candidate plugin.
Existing API-key setup and explicit ACP commands remain available. No database schema migration is introduced. No package release, operator Gateway redeployment, or OpenClaw merge is included in this preparation.
## Contributor context
Sandbox recovery and upstream-first integration requested by @obviyus. [Discussion, implementation and verification](https://team.openclaw.ai/chat/roboclaw/dashboard/c522995c-e4fe-471a-848e-0d8fc44edd5c).
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
Keep optional Doctor inspections and intentional channel-policy advisories visible as warnings during updates while preserving required configuration, state, and startup checks. Preserve complete bounded findings in named diagnostic artifacts and compact receipts in update history; retain the published rollback reader's 8 KiB file, 4 KiB prompt, and 768-byte error-field limits.
Publish admitted-run terminal outcomes and report paths through the existing terminal owner after executor settlement. Standalone finalization uses the existing maintenance owner to park and restore its Gateway, while inherited finalizers retain the parent's ownership.
Validation: reviewed source has passing focused regression, released-reader compatibility, changed-file, and import-cycle proof. Published 2026.9.4 upgrades and standalone finalizers passed on Linux and macOS at the earlier source documented in the PR; later diagnostic-envelope and exceptional-publication corrections have focused regression coverage. Exact-head CI is required by native preparation and merge.
Thanks to @davidchyi-beep, @BodegaClaw, @CHE10X, @jammyclaw, and Discord reporter Patryk for the reports and diagnostic evidence.
Fixes#152759. Fixes#152918.
Retain structured plugin load failures through Doctor completion after temporary inspection registries are released. Report plugin IDs, error classes, and source paths through existing health findings while preserving lazy metadata inventory.
Standalone non-interactive Doctor exits 1 for plugin load failures. Updater-invoked Doctor retains these failures as warnings and continues otherwise safe updates, including published 2026.9.5 parent-marker behavior. No persistent data-model change; scratch accumulation and capacity are outside this fix.
Validation: 375 repository tests and two published-driver boundary cases passed; runtime build, changed checks, and scoped Codex review passed.
Related: #153607 (Doctor silent-failure reporting only). Thanks @th-m-vogel for the report.
Share the existing cron cancellation call between forced and expired
active-work drains before server close joins connection-owned cleanup.
Successful ordinary drains and existing shutdown budgets stay intact.
Cover real SIGUSR1 and SIGTERM lifecycle handling, including retained
cleanup writes. Validate the published 2026.9.5 restart driver against an
isolated full candidate Gateway with active cron work and unchanged
systemd shutdown budgets.
## What Problem This Solves
Fixes `openclaw message send --media first.png --media second.png` silently dropping the first attachment. The CLI previously overwrote each earlier `--media` value before the shared send action could see it.
## User Impact
Repeated media flags now reach the existing send pipeline in their original order. Single-file and text-only sends are unchanged. Telegram uses its existing photo-album support; existing duplicate removal and caption behavior remain unchanged. This does not add a channel capability, alter message receipts, or resolve the broader album feature requests #13620 and #135545.
## Why This Change Was Made
The command reuses the existing repeated-option collector and passes its ordered list to the shared `mediaUrls` contract. The contributor's production repair and commits are preserved. Regression coverage is consolidated in the existing actual-CLI process suite rather than a new test that only spies on argument forwarding.
## Evidence
- The actual CLI regression failed on pinned main: serialized dry-run output contained only `second.png` instead of both files. The repaired process case passes; all 64 focused CLI process/helper tests pass. Scoped lint, formatting, both line ratchets, and runtime builds pass.
- Five network-denied actual-CLI dry-run cases per phase covered two files, one file, text only, repeated duplicates, and three media flags mixed with the message option.
- Ran both phases through the actual CLI, an isolated Gateway, the registered Telegram adapter, Telegram Test Server, and a separately authenticated QA user under the maintained credential lease.
- Native recipient observations confirmed the baseline delivered only the last photo. The candidate delivered red then blue in one two-photo album, preserved the single red photo and text-only controls, kept red/blue order after duplicate removal, and delivered red/blue/green in one three-photo album.
- Inspected the actual recipient minithumbnail bytes and native album/message ordering, not reconstructed chat images. This is attachment-delivery evidence, not a Telegram client layout or image-quality claim.
- Both native recording runs completed and removed credential/Gateway scratch. Synthetic Test Server DM messages remain as test evidence; no deletion is claimed. Baseline had no model calls; the candidate had two synthetic-provider calls independently attributable to an autonomous heartbeat and its recap, not media sends. No paid inference was used.
All three final changed files match the tested current-main replay. The change is confined to CLI collection, its process regression, and CLI documentation; no schema, permission, dependency, provider, or UI-rendering changes are included.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
Fixes#145058
## What Problem This Solves
Fixes `openclaw nodes invoke --idempotency-key ""` contacting the Gateway and returning a generic RPC error instead of identifying the empty CLI option.
## User Impact
An explicitly empty key now fails locally with `--idempotency-key` guidance before node lookup or any Gateway connection. Nonempty keys retain their exact bytes, including whitespace-only and padded values. Omitting the option still generates a key. No Gateway schema, pending-action identity, permission, or configuration contract changes.
## Why This Change Was Made
The existing invocation command validates the empty string alongside its other input checks. The original repair was simplified to a three-line guard, preserving the contributor's commits and the existing key-generation path. The CLI reference documents the option's behavior.
## Evidence
- Ran the actual CLI through `scripts/run-node.mjs`, the repository owner of `pnpm openclaw`, against a task-owned recording WebSocket Gateway fixture. Each phase covered 14 plain/JSON cases with isolated state and a loopback-only network fence.
- On pinned main, the empty key opened two connections and reached `node.list` and `node.invoke`; the CLI displayed the fixture-modeled schema error. The rebuilt candidate returned local option guidance with zero connections or requests.
- Whitespace-only, padded, ordinary, and omitted keys reached the wire correctly. Both shell-command denylist controls remained local failures. The fixture and CLI processes settled and disposable state was removed.
- The empty-key regression failed on the original command. All 62 tests in the existing CLI coverage and plugin-registration suites passed on the candidate; focused lint, formatting, both line ratchets, and runtime builds passed.
This is CLI/transport boundary proof, not execution of the real Gateway validator, a native node, or pending-action deduplication. The production file and changed regression region match the tested current-main replay. The first isolated package-manager attempt and an intermediate dirty-checkout rebuild were blocked from downloading dependencies; the final clean, rebuilt checkout passed without widening network access.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix: preserve restart ownership and chat position during startup
Keep foreground and update-owned restarts within their existing capped cleanup deadline instead of assuming an enclosing service will replace them. Anchor the chat position rail after its first visible-message measurement. Stabilize Doctor and memory SQLite fixtures and assert the actual Codex process-notification ordering contract.
* test: exercise doctor state isolation on Windows
* fix(ci): wait for exiting Darwin process groups
* fix(tasks): keep superseded progress from failing readers
* fix(tasks): preserve publication ownership through supersession
* test: acknowledge delivery fixtures with their claim owner
* test(qa): join parent acknowledgement before terminal assertions
* test(tasks): arm scan barrier after starting mutation
* test: retire task event owners at fixture boundaries
* test: share port claims and await hover recovery
Keep the active packaged CLI as the update target when the managed Gateway service points at another installation. Reconcile eligible previously running services through the existing native installer, with live authority checks, verified backup restoration, and visible recovery warnings. Doctor preserves already-stopped services and reports the profile-aware installer command; unresolved cross-install drift after state maintenance keeps the older service stopped until explicit recovery.
Status and Doctor report both package paths and versions even when the Gateway cannot answer. Deployment-owned definitions and pending package recovery remain protected by their existing owners. No new configuration options, CLI options, dependencies, or SQLite schema are introduced.
Validation includes the campaign's scoped-clean P0/P1 Codex review, green exact-head CI, focused owner regressions, and historical native launchd replacement and rollback. Published 2026.9.4 refused before candidate handoff without changing the running Gateway; successful old-driver adoption and fresh Linux custody recovery remain unqualified, and Windows evidence remains fixture-backed.
Closes#150744. Thanks @baiwei0427 for reporting the two-prefix failure and its manual repair.
Full read-only Doctor reports now reuse private state bytes through the existing report-scoped snapshot lifecycle, avoiding repeated database copies on large installations. Each new report still takes a fresh snapshot, selected checks remain on demand, and independent verification and writable inspectors retain their own copies.
Preserve completed update findings and exit codes when private snapshot disposal fails; record that disposal-only failure as a warning. No configuration, CLI, dependency, persisted schema, or data-format changes are required.
The matched Linux Testbox fixture reduced Doctor staging writes from 21.76 GB to 7.03 GB across seven reports (67.7%), with complete cleanup. Sampled peak temporary occupancy rose from 184 MiB to 339 MiB. Real-WAL tests cover within-report reuse, next-report freshness, unchanged source artifacts, and cleanup; the published openclaw@2026.9.4 updater successfully installed the packaged candidate.
Addresses the read-only Doctor portion of #153067. Related: #153276 and #153512. Thanks to @MasterSwords1 for analyzing snapshot costs and @dmlau76 for reporting the issue. This is an independent implementation.
Plain triage could discard failed update history and report “already resolved” from healthy Doctor results. Centralize update-history selection and completion checks in the resolution owner, preserve pending plugin migrations, and keep explicit failure artifacts tied to their original targets.
Accept later successful upgrades and verified rollbacks after validating the current installation, Doctor results, and managed Gateway readiness. Abandoned-run acknowledgement still requires fresh Doctor checks. Incomplete recovery stays visible with update repair guidance; historical ledger records remain intact.
Fixes#153378. Thanks to @johnnyjrizzo for the report and recovery evidence.
Validation: 300 focused tests across seven files and full check:changed passed in the repair lane. Retained isolated macOS PTY proof with published 2026.9.5 ledger state confirms pending migrations and missing completion remain unrepaired; current command regressions cover the later recovery and rollback fixes. The review of record is scoped-clean and exact-head CI passed.
Backups of active workspaces could abort when a file disappeared between directory enumeration and lstat/open. Open regular files before writing archive headers so vanished descendants are omitted with per-path warnings while surviving data remains verifiable and restorable.
Keep required roots and captures, unsafe source substitutions, and I/O failures as publication blockers. Use the existing resource inventory and tar Header/Pax encoder; retain truncation retries and keep omission reports outside the size-limited restore manifest. Normalize Windows link targets once before manifest recording and encoding, preserving literal POSIX backslashes.
This repairs explicitly requested full backups before updates. The automatic updater's separate configuration snapshot is unchanged. No configuration options, schemas, or dependencies change.
Evidence: 285 backup tests passed with 2 existing platform skips across 11 files; check-changed passed; independent P1 review scoped-clean; exact-head CI green. Windows tar path-mode coverage passed; native Windows execution was not independently verified.
Thanks to @migelizaga-rgb for the report and reproduction.
Fixes#153250. Related: #153315.
Reuse completed candidate validation and its safely retained migrated private state when automatic repair starts. A later check failure no longer triggers another multi-gigabyte snapshot and Doctor pass before discovering that inference is unavailable.
Keep the existing repair owner responsible for consuming the cached result once, validating every subsequent repair turn, draining workers, and cleanup. Incomplete migrations or unconfirmed teardown still require a fresh copy, and activation always requires fresh-copy validation. Keep source configuration hashes bound and bound diagnostic redaction to the existing IPC limit.
Recorded Linux package proof covers published 2026.9.4 to candidate, injected recovery failure with one snapshot and no repeated Doctor, and the next-version fixture. All 13,600 cache rows and database sizes were preserved. Exact-head CI passed; scoped review found no actionable findings. Runtime package evidence belongs to the earlier proof revision with review-confirmed continuity, and does not establish macOS or managed-service package-update coverage.
Thanks to @BodegaClaw for the measurements and follow-up evidence.
Refs: #153049, #153077, #144447, #153099, #144858.
Keep verification records in the quarantine store and let the existing
lease owner consume clean proof and certify the last checkpointed close.
Share gate eligibility with readiness, queue quick verification after
listening, and invalidate proof across crashes, aliases, and maintenance.
The warmed 899.5 MB fixture gate falls from 1094 ms to 17.8 ms (98.4%).
Full checks remain for updates, migrations, unclean or replaced files,
Doctor, shared state, snapshot fallback, and daily verification.
Closes#151301
## What Problem This Solves
Fixes browser storage commands silently trimming keys, so reading or setting `" account "` targets the distinct `account` entry instead.
## User Impact
Quoted keys retain their exact contents in both local and session storage. Setting a padded key no longer overwrites its unpadded neighbor. Ordinary keys, primitive-key coercion, value whitespace, omitted-key reads, and rejection of blank SET keys remain unchanged.
## Why This Change Was Made
The CLI lookup and the shared storage GET/SET routes now reuse the existing whitespace-preserving string helpers. Browser authorization, URL guards, storage formats and the Playwright storage owner are unchanged. Existing regression tests assert returned values and resulting storage state instead of forwarded call shapes.
## Evidence
- Ran 24 independently reset before/after cells through the actual built browser CLI and direct authenticated Gateway requests, using a task-owned real Chromium page. Raw CDP independently inspected both storage buckets.
- Baseline padded reads selected the unpadded entry; padded writes overwrote it while leaving the requested entry unchanged. Candidate reads and writes use the exact padded key and preserve its neighbor.
- Both storage kinds passed ordinary-key, padded-value, omitted-key listing, rejected blank SET/no-mutation, and numeric/boolean key controls. No egress attempts occurred, and all task-owned Gateway/client/browser resources settled.
- 57 focused CLI/route tests passed. Reintroducing trimming at the two owners caused five distinct regression failures while the two clear-storage controls still passed. Runtime build, scoped lint, formatting, whitespace and line-cap checks passed.
Proof used main `d0192ed506` with byte-identical contributor production files. The isolated Gateway used its shipped canary startup profile to suppress unrelated autonomous sidecars while preserving real plugin loading and method admission; this is not a full normal-startup or upgrade claim. The disposable Chromium proof replaces the proposal's optional direct-route-only fixture; retained tests cover the regression at both parsing boundaries without requiring a browser installation.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix: preserve Node debugger attachment on SIGUSR1
Use SIGUSR2 for Gateway restarts and reserve SIGUSR1 for Node's inspector.
Update internal callers, diagnostics, QA harnesses, and operator docs.
Route current unmanaged Gateways through owner-targeted restart RPC while
retaining SIGUSR1 only for verified legacy Gateway locks during upgrades.
Refuse self-signaling embedded Unix hosts without a restart handler.
Manual restart scripts must switch to SIGUSR2; prefer openclaw gateway restart.
* test: follow prepared runtime replacement owner
Avoid false Gateway timeout and unreachable reports during slow cold starts. Doctor, health, and status diagnostics now reuse the existing readiness owner and one monotonic deadline, preserve observed startup progress, and leave a Gateway that is still starting alone. Remove Doctor's duplicate restart polling loop.
Provider usage consumes the remaining allowance, reports Timeout before dispatch when exhausted, and cancels active usage work on expiry. Preserve explicit timeouts, target/auth selection, and real process/version/build failures.
Health and channel-status JSON can return a documented non-failing starting payload; consumers that need a complete snapshot must distinguish it and rerun after startup. No new configuration, flags, dependencies, or stored-data migration.
Validation: 592 cases across 37 files on isolated Linux Testbox; installed-package cold-start/down-Gateway matrix and unmodified 2026.9.5 updater-to-candidate Doctor proof; focused command/transport regressions and changed-file gates. The existing independent review is scoped-clean, and ClawSweeper reports no actionable correctness finding.
Thanks to @Suidge for the report and startup measurements.
Fixes#152970.
Doctor could leave a healthy systemd Gateway stopped when repeated update-admission checks exhausted restoration inspection. Use lightweight custody for metadata reads, retain the native service identity before shutdown, and restore with a warning after inconclusive diagnostics while verifying readiness. Revalidate broker, manager, unit, account, and current authority at activation; preserve explicit ownership refusals without an unbound fallback.
Keep failed Doctor lint output compatible with published updaters by retaining the readiness envelope, redacted error metadata, and exit 2. No persistent schema, configuration, CLI, or dependency contract changes.
Thanks @xilopaint for isolating the LoadUnit delays and providing the timing trace.
Validated by green exact-head CI, focused regression coverage, real Linux recovery and native refusal observations, and a published 2026.9.5-to-candidate update that completed successfully without rollback. The frozen released reader accepts success, warning, and repaired failure output. Independent Codex review found no actionable P0/P1 findings.
Fixes#152879.
Manual cron runs now reserve their durable SQLite receipt before the Gateway acknowledges command-lane admission. The receipt preserves the returned runId so existing startup recovery accounts for an acknowledged request interrupted before dispatch; recovery does not replay it. Pre-dispatch interruptions remain receipt accounting and are not shown in task-backed cron.runs history.
The existing cron admission and receipt owners retain caller revalidation, exact-reservation cleanup, and future schedule ownership. No configuration, dependency, schema-version, or retention changes are required; updater and rollback behavior are unchanged.
Related: #146956. Follow-up to #147054. The acknowledgement-loss finding was reproduced during managed restart validation.
Validation: isolated production CronService/SQLite subprocess acknowledgement, SIGKILL, restart, and exact-runId interruption proof; 105 focused cron tests, 25 on-exit/heartbeat checks, and 89 reused receipt/recovery/scheduling checks. Planned changed-file checks passed. The recorded Codex review was scoped-clean and the exact-head ClawSweeper review found no actionable defects.
Repair official npm plugin version drift during doctor --fix after a manual core installation, through Doctor's configured-plugin repair and the existing plugin updater. Preserve explicit tags, newer pins, and floating selectors; unavailable packages remain warnings while other repairs complete. Ordinary diagnosis and third-party plugins stay unchanged.
Keep installation transactions, lifecycle recovery, and persisted-index writes with their existing owners. Cover selector precedence, correction releases, persisted records, partial failures, and already-current no-op updates. Document manual recovery for the published updater timeout.
Thanks to Prince for the Discord report.
Validation: independent Codex scoped review; 387 focused tests across seven suites and changed-file checks passed. Recorded published-driver, packaged Doctor, partial-registry-failure, and managed-Gateway proof retain their original head bindings. Exact-head CI and native preparation gates verify the landing candidate.
Publication and confirmation now distinguish held workspace leases from unavailable storage and caller-canceled acquisition. The shared SQLite lease owner records these outcomes, replacing the publication-specific retry inference while preserving caller authority, native settlement, and retired-owner fencing.
Doctor records pre-grant cancellation as a visible inspection warning without changing its signal budget or mutating source state. No new configuration, CLI options, schemas, dependencies, or wait limits are introduced.
Validation: scoped-clean Codex review; native SQLite/worker, registered publication RPC, Doctor, updater-authority, and wrapper-boundary proof; exact-head CI gate success. The current candidate's published-updater through healthy Gateway restart scenario remains an explicitly accepted, bounded validation gap recorded in the PR body.
* fix: restore Mac presence context and Canvas media playback
* refactor(gateway): separate the node session type contract
* test(codex): cover active presence in prompt fixtures
* test: include active computer in Codex context ordering
* test: supply snapshot auth-profile store
* test: align presence search and CLI prompt expectations
Detect manual package replacement from the Gateway maintenance tick or a missing runtime import, close new-work admission, retain active reply completion and cleanup dependencies, and drain before native-supervisor handoff. A foreground Gateway exits with an explicit relaunch instruction. Preserve the explicit updater's successor ownership and expose replacement history through status and Doctor using the existing boot record.
Thanks @mattkanwisher for reporting update-triggered message loss. Fixes#152997. Related: #152466 and #150153.
Validation: native hosted CI/Testbox gates passed at the unchanged reviewed head. 201 focused tests passed, changed checks passed with the documented canary-verified nested-checkout lint substitute, and Codex P1 review was scoped-clean. Real Linux npm/systemd replacement proof preserved published-created agents, sessions, transcript, config, plugin, and cron state. A running published 2026.9.5 process required one operator restart; a running candidate completed its held reply, drained, restarted exactly once, and retired its plugin, capture, worker, and process resources. No schema, retention, option, or dependency change.
Repair now acknowledges outstanding abandoned updates in Doctor's latest 100 history records, including aged runs behind newer updates. Doctor stops repeating resolved repair guidance while the original failed outcomes, finish times, and unknown target-build evidence remain intact.
The repair command selects the outstanding records and delegates acknowledgement to the existing successful-finalization owner. Aged records still require full convergence; active updater admission and failed-repair behavior remain protected.
Thanks to @IWhatsskill for the source-cited report in #153228.
Validation: 233 focused tests across 8 files; changed-file checks; isolated CLI repair against records from the published 2026.9.5 ledger reduced Doctor guidance from two instructions to zero while preserving historical failures. Campaign autoreview was scoped-clean and exact-head CI passed.
Doctor can reject verified session SQLite migrations while a plugin remains pending, reporting false transcript errors or leaving plugin configuration locked after the legacy index is retired.
Reuse verified receipt source identities, accept verified archive-only prerequisites when no live legacy inputs remain, and let the existing post-session plugin owner settle migration and unlock retained configuration. Warning-only session commands exit successfully; genuine refusals preserve their actionable failure facts across Doctor and update reporting. Schemas, retention, dependencies, and options are unchanged.
Fixes#153079. Fixes#152744. Thanks @berklingtools for the recovery report and @wiipud and @ozp for diagnosing the index-free prerequisite failure. Additional pinned-plugin validation followed @ericpearson's Docker report.
Validation: published 2026.9.5 and candidate CLI fixtures preserved session/event counts, archive hashes, and SQLite integrity; candidate Doctor completed the pending plugin row and unlocked retained config. The 78 focused owner tests, 24 shard tests, changed-file checks, and exact-head CI passed. The scoped review was clean and both rebased commits are patch-identical to the reviewed versions. Full published-updater-to-candidate finalization and Linux service-restoration proof remain unperformed; service restoration is tracked separately in #153017.
* docs(providers): expand llmman guidance and add hybrid inference
Rewrite docs/providers/llmman.md into a full provider topic page: modes
table, auth rules, getting started, model discovery, smoke tests, vision,
configuration tabs, recipes, advanced accordions, troubleshooting.
Add a Hybrid inference section covering llmman.hybrid/<local>,<provider>/<model>
refs (qwen3.8 + openai/gpt-5.6-luna), routing rules, key handling,
x-llmman-route pinning via provider headers, and how it composes with
OpenClaw fallbacks. Document llmman.provider/... hosted refs.
Use qwen3.8 as the reference model throughout, call `llmman serve` with no
arguments everywhere (daemon settings go in its environment), and adopt the
LLMMAN_API_KEY=llmman-local / apiKey: "${LLMMAN_API_KEY}" convention.
Cross-link from local-models, local-model-services, model-providers,
infer CLI, provider index, and the models FAQ. Add the new provider index
label to the zh-CN glossary.
Verified against llmman v0.1.334 with qwen3.8: text and vision probes via
openclaw infer model run, a full tool-calling agent turn, hybrid routing
(local by default, cloud on pin or oversized body), and
chat_template_kwargs passthrough for thinking control.
Co-authored-by: sallyom <11166065+sallyom@users.noreply.github.com>
* docs(llmman): correct verified runtime guidance
Clarify Qwen thinking compatibility, daemon-wide hybrid budgets, and idle service lifetime. Preserve the existing integration and live proof.
Co-authored-by: sallyom <11166065+sallyom@users.noreply.github.com>
* docs(llmman): align service startup reference
Keep the shared local-service example consistent with llmman optional model preloading.
Co-authored-by: sallyom <11166065+sallyom@users.noreply.github.com>
---------
Co-authored-by: sallyom <11166065+sallyom@users.noreply.github.com>