* fix: preserve plugin settings through update recovery (#160344)
Keep incompatible local plugin configuration intact and complete deferred migration confirmation after unchanged package convergence. Use completion receipts to clear obsolete retry warnings without rewriting update history.
Fixes#159477. Thanks @EndeavorPioneer for the recovery report.
(cherry picked from commit 8e66b0b6c8)
* fix: prevent duplicate Gateways during container onboarding (#160193)
* fix: prevent duplicate Gateways during container onboarding
* fix: keep QA cron controls scoped to the child
Drop inherited parent cron suppression while preserving explicit child runtime
settings. Cover the failing inheritance contract at the environment owner.
Await node-list RPC responses after acknowledged rename and reconnect events
without racing the existing RPC deadline against a separate polling timeout.
(cherry picked from commit 54b1b1b63e)
* fix: deeply nested MCP tool results crash result projection (#160825)
Handle deeply nested MCP structuredContent as an actionable tool error instead of letting a serialization RangeError escape result projection. Preserve server content and report the same failure to Code Mode guest callers, without retaining the unprojectable value for downstream serialization.
Preserves ordinary shallow results and adds focused regression coverage.
Co-authored-by: wangmiao0668000666 <wang.miao86@xydigit.com>
Co-authored-by: Tak Hoffman <781889+Takhoffman@users.noreply.github.com>
Co-authored-by: Takhoffman <781889+Takhoffman@users.noreply.github.com>
(cherry picked from commit 452c808d5b)
* feat(openai): support GPT-6.1 Sol (#161400)
* feat(openai): support GPT-6.1 Sol
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
* feat(openai): support GPT-6.1 Sol
Worked on by:
- @steipete
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: 970e2aab-1efa-4534-be78-7b6ec08717b5
* fix(openai): preserve GPT-6.1 Sol request capabilities
Preserve mandatory reasoning through subscription discovery and configured Responses requests. Extract existing catalog readers and discovery coverage to respect file-size ratchets.
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
* test(openai): narrow discovered model request fixture
Require the discovered row before converting catalog modalities and optional context metadata to the typed chat runtime model.
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
---------
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
(cherry picked from commit 858993c1a4)
* fix(logging): redaction stalls for seconds on long plus-joined text (#161089)
* fix(logging): stop redaction stalling on long plus-joined text
The seven vendor-token patterns built on BASE64_SAFE_TOKEN_BOUNDARY ran
their data-URL negative lookbehind at every `+`, `/`, or `=` boundary of a
base64-class run, rescanning the whole run each time. On two-byte subjects
(any character above U+00FF) V8 no longer rejects those boundaries early,
so 100 KB of `+`-joined text took seconds; JSC shows it on more shapes.
Check the token with a lookahead first so the lookbehind runs only where a
complete token starts. The matched language and capture groups are
unchanged.
* refactor(logging): inline single-use redaction helpers
Compact the base64-safe token helper (identical pattern strings) and fold the naming-only readEnvAssignmentKey wrapper into its only caller, keeping the fix production-LOC negative.
(cherry picked from commit ff21cc82ce)
* fix(gmail): recover watcher after transient restart bind conflicts (#161503)
The Gmail watcher stayed down for good when a Gateway restart briefly found its port still in use (EADDRINUSE). It now retries the bind a few times with bounded backoff and recovers once the port frees. If the port stays taken, it reports a clear error and stops retrying.
Fixes#161467. Reported by @kazuyuki-eguchi.
Proof: a real isolated Gateway, with the Gmail watcher on local fakes.
- Before the fix, the watcher stayed down after a transient port conflict. After it, the watcher recovered and answered.
- With the port held for the whole run, the watcher stopped after the initial attempt plus three retries, with a clear error.
- The regression test fails before the fix and passes after, and 31 focused tests pass.
- Updating from the published 2026.9.6 build to this build succeeded in two fresh runs (102 s and 97 s), and the installed build passed both the transient and the persistent control. The first updater run failed at service activation and didn't recur. The watcher can't reach that step: rehearsal disables hooks, and this diff doesn't touch activation or lease code.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
(cherry picked from commit 8095304914)
* fix(sessions): reject missing and malformed session targets (#161533)
* fix(sessions): reject compaction of missing sessions
Return an actionable INVALID_REQUEST instead of a successful no-op for missing targets in both compaction modes. Preserve explicit no-op outcomes for existing empty sessions.
* fix(sessions): reject malformed reset agent selectors
Use strict session agent input validation before lifecycle cleanup so an invalid explicit selector cannot reset the default agent. Extract reset-target resolution from the oversized lifecycle service while retaining selected-global targeting.
Co-authored-by: Peter Steinberger <steipete@gmail.com>
(cherry picked from commit 670c2d0d6a)
* fix(agents): honor captured subagent announce skips (#161542)
Authoritative terminal snapshots bypassed ANNOUNCE_SKIP filtering and woke
the requester unnecessarily. Apply the existing completion selector before
expanding retained history, then suppress the successful announce while
preserving child cleanup. Keep global terminal snapshots, legacy fallback
behavior, and failed or timed-out child outcomes unchanged.
The real Gateway reproduction made three parent model requests instead of
two; five fixed repetitions avoid the extra turn. Three pre-fix regression
failures now pass. Blacksmith Testbox proof: 634 related unit cases, 87
announce E2E cases, 15 stub-provider operator scenarios repeated five times,
and the full changed-code gate. Streaming/reconnect/timeout-compaction
baseline flows passed once; their five-repeat rerun and process-restart
proof were not completed within the campaign time box. No live credentials.
Changed-test wall costs (pnpm test <file> --maxWorkers=1):
T2 COST FILE src/agents/subagents/announce/subagent-announce.test.ts
T2 TEST WALL 47.63 seconds
T2 COST FILE src/agents/subagents/announce/subagent-announce-output.test.ts
T2 TEST WALL 43.91 seconds
T2 COST FILE src/agents/subagents/announce/subagent-announce-result.test.ts
T2 TEST WALL 29.28 seconds
T2 COST FILE src/agents/subagents/completion/subagent-completion-result.test.ts
T2 TEST WALL 1.33 seconds
Co-authored-by: Peter Steinberger <steipete@gmail.com>
(cherry picked from commit 806f961301)
* fix(auto-reply): release canceled thinking-catalog waits (#161319)
* fix(auto-reply): release canceled thinking-catalog waits
Race the canceled reply's model-level and status waits against its own
AbortSignal without canceling shared catalog discovery.
Refs #161310
* fix(ci): keep database-worker test routes unique
The duplicate pdf-tool.resources.test.ts entry introduced in #161270
causes preflight to fail while splitting core-runtime-infra-storage-state.
Retain its original route once so PR test plans can be built.
* fix(auto-reply): keep abort assertion outside preprocessing import cycle
Move the existing AbortError assertion into a small reply leaf module so
model-level resolution can use it without importing the preprocessing graph.
The reply assertion behavior remains unchanged.
* test(ui): wait for shell viewport sync before splash geometry check
* fix(auto-reply): cancel earlier directive catalog waits
* chore: keep catalog cancellation PR scoped to reply handling
* test(auto-reply): avoid returning timer handle from executor
---------
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
(cherry picked from commit f18202c76b)
* fix(tts): pass the responding agent to summary model acquisition (#161454)
* fix(tts): pass the responding agent to summary model acquisition
summarizeText acquired its model without an agent id, so it fell back to the default agent and threw the no-explicit-owner error in multi-agent configs. Forward the reply's agent id from maybeApplyTtsToPayloadCore. The injected-deps path is unchanged.
Fixes#161431
* test(tts): cover multi-agent summary through real model selection
Drives summarizeText for agent 'work' in a two-agent config through the
real acquisition and selection path, stubbing only the completion call.
Asserts the ownership error is gone and the summary keeps using the
documented agents.defaults.model.primary fallback.
* test(tts): retain behavioral owner regression and document auth scope
---------
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
(cherry picked from commit 5073ee41ab)
* fix(imap): bound backlog message retention during sweeps (#160413)
The IMAP watcher collected every unseen message into memory before processing any of them, so a large backlog held every message body at once. Sweeps now fetch and process unseen messages in bounded batches, in UID order, with the existing cursor, retry and sender-gate behavior unchanged.
Fixes#160353. Thanks @Ayushdevo for the fix and @addyCooks for the report.
Proof: a real ImapFlow TCP connection to an in-process IMAP server holding 2,000 messages, run in a container.
- On main, the sweep held all 2,000 bodies at once. With this change it held at most 20, and every message was admitted exactly once in UID order.
- A rejection injected at UID 25 resumed from cursor 24.
- A sender-gated message (UID 30) was skipped once.
- The regression test fails before the fix and passes after, and the watcher suite passes 21/21.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
(cherry picked from commit ecbe1861ef)
* fix(cron): isolated setup times out while the thinking catalog hydrates (#161132)
Isolated cron turns spent their whole 60-second setup budget waiting for the native thinking catalog to load, so the setup watchdog killed runs whose catalog was slow (measured at 13–65 s on 2026.9.6). The cron path now bounds that wait the same way `loadFullModelCatalog` already does and falls back to the published catalog facts. The shared `loadNativeModelCatalog` stays unbounded for the other callers that depend on it returning the complete catalog.
Fixes#161112. Thanks @jayzhou2309 for the fix and @Flakedict for the report and the measurements.
Proof: a real isolated Gateway in a container, with the mock provider's native catalog publication held open.
- On main, cron setup hit the watchdog at 60,012 ms. With this change, the cron turn replied in 11,756 ms while the publication was still held.
- When the catalog is fast, the turn replied in 1,683 ms using the full catalog.
- The regression test fails before the fix and passes after, and the scoped tests pass 7/7.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
(cherry picked from commit 786ed6b566)
* fix(matrix): preserve direct mappings when account-data reads fail (#161727)
Preserve existing Matrix peer mappings when an account-data read fails. Keep best-effort fallback on read-only inspection and propagate read errors before writes.
The regression failed on the original code and passes after the fix. Direct-room, sibling sender/FIFO, changed-file and zero-cycle checks passed; repair documentation explains the failure behavior.
(cherry picked from commit 78d75e9350)
* fix(agents): detached runs keep their transcript when a settled turn is finalized (#161091)
* fix(agents): preserve detached finalization history without internal prompts
Keep the host-owned session manager and persistence mode during settled-turn finalization. Refresh one-shot user persistence suppression when reusing the guard so empty-answer retries do not retain internal recovery prompts. Preserve completed tool receipts and subsequent genuine user turns.
Regression fails before the guard repair with two retained internal prompts and passes afterward. Focused proof: 126 tests, core tsgo, scoped lint, formatting, and whitespace checks pass. Independent Codex review found no actionable P0-P2 findings. Both native-live cases remain enabled on main.
* test(agents): reuse a caller-owned admission without growing the attempt runner
* test(agents): pass the caller-owned admission from the typed override
(cherry picked from commit 8f924bff5d)
* fix(worker): stop durable session retries from freezing node hosts (#160812)
* fix(worker): stop durable session retries from freezing node hosts
* test(worker): name closing window harness as repro
* test(worker): fix loopback harness checks
* test(worker): adapt closing-window regression to Gateway tools
---------
Co-authored-by: Altay <altay@hey.com>
(cherry picked from commit 1de9a42f11)
* fix: preserve plugin install records during legacy updates (#161485)
The v13 wide-row state migration conflated a missing installed_plugin_index row with a row whose JSON was unparseable or shape-drifted, and dropped the table in both cases. An invalid row is now preserved in full in the existing diagnostic_events quarantine with repair guidance recorded, valid install records are retained independently of damaged metadata, and a preservation failure rolls the step back to v12. No schema change.
Closes#161329.
Landed under the pre-existing-red rule: the remaining CI failures were current main reds at this base (Codex app-server settlement fixture drift, fixed on main by 52e60fb42b; the subagent kill-tombstone / descendant-cancellation intermittent first seen on main hourly 36619831492).
(cherry picked from commit e1ea50caf7)
* fix(gateway): restore CLI-backed exec with secret egress proxy (#160760)
* fix(gateway): pass admitted run instance to loopback-mediated exec
* test(gateway): cover egress-enabled MCP exec grants
Prove the real MCP HTTP grant, cached tool construction, egress-enabled process launch, retired-grant rejection, and a later run in the same session. Route the fixture through the existing native database-worker test group.
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* test(ui): wait for transcript resize before measuring centering
The real layout owner publishes viewport width asynchronously after resize. Wait for rendered width readiness, then sample all centers atomically while retaining visible-box and one-pixel assertions. The focused catalog E2E now passes in secretless AWS; the real-browser mechanism reproduces the prior 160px phase and still rejects a 2px displacement.
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* fix(ci): restore Doctor plugin repair lint budget (#160715)
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* fix(ci): restore config and node adapter lint budgets (#160843)
* fix(ci): keep config env and worker tests within lint budgets
* test(gateway): share node capacity and rejection fixtures
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* refactor(talk): combine transcript early-return guards
Preserve echo-first short-circuit evaluation while restoring the existing relay file-length budget. No runtime behavior or lint limit changes.
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* refactor(ui): share model setup activation payloads
Reuse the same discovery-to-activation projection for prepared and directly selected candidates, retaining optional-field omission and excluding discovery metadata. Restore the existing model setup page line budget without changing UI behavior or lint limits.
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* test(gateway): join chat execution before metadata assertions
The operator.write verbose-level fixture asserted agent.wait success after a one-second observation window without joining its detached run. Reuse the existing execution observer before the same scoped RPC, preserving the timeout and every metadata/permission assertion.
Verified the complete chat file on CI's Node 24.19.0 in network-isolated execution: 41 tests passed. Independent P0-P3 review and targeted lint passed.
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* test(gateway): join recap producer lifecycle events
Wait for actual model entry before inherited connection drain and for the relocated recap post-commit publication before inspecting persisted state. Preserve model cancellation, old-owner fencing, exact target/store checks and existing persistence assertions. New waits follow test cancellation rather than timing a cold worker through an observation poll.
Verified both edited files (25 tests), the original receiving-merge peer groups (377 and 164 tests), focused lint, both owning type graphs and independent P0-P3 review.
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* fix(infra): retain explicit SQLite launch paths when cwd is unavailable
Keep supplied launch and transport facts, resolve relative arguments only against a genuine captured cwd, and anchor native inspections to selected absolute operation paths when ambient cwd is gone. Refuse unresolved cwd-dependent inputs rather than redirecting them. Prove both native transports after real directory removal; cold Node Worker bootstrap remains a separate lifecycle limitation.
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* test(gateway): coordinate monitor fault injection with worker admission
Use the existing managed state write transaction for persistent trigger creation and removal. Controlled scheduleUnowned interleaving reproduces both raw-DDL lock failures and passes both managed-DDL variants without changing reload assertions or timeouts. Ensure final trigger cleanup cannot skip service cleanup.
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* test(gateway): keep monitor fault injection in fixture support
Preserve the exact managed DDL statements and cleanup lifecycle while shrinking the over-cap reload test. Line-cap, suppression, assertion and import-cycle guards pass; the 115-case file, focused lint, owning types and independent P0-P3 review pass.
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* test(update): select host runtime for Homebrew guidance
Keep the Homebrew-specific guidance fixture independent of the container used for isolated proof. Container cases and all guidance assertions remain unchanged.
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
---------
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
(cherry picked from commit 6daee6e4bd)
* fix(sandbox): sessions fail with EACCES after skills refresh on read-only installs (#162304)
* fix(sandbox): keep synced skill copies removable on read-only installs
Sandbox skill sync copied bundled skills with fs.cp, which keeps source
modes. From a read-only install every copied directory became 0555, so the
next full re-sync (skills version bump or Gateway restart) failed with
EACCES unlinking skills/<skill>/SKILL.md, and library-pinned sessions
failed every turn.
Copied directories are now made owner-writable (file modes unchanged), and
removal repairs legacy read-only trees inside the synced child before one
retry. The repair stays inside the skills root and never follows symlinks.
The Claude CLI skills-plugin copy fallback uses the same helper.
* fix(sandbox): classify skill removal errors without a type assertion
(cherry picked from commit be93e85955)
* fix(agents): cron fallback delivers a truncated reply when history caps a long message (#160520)
When a run's final message was longer than the chat history display cap, the run-wait reply reader fell back to the capped history copy and delivered it with an internal `...(truncated)...` marker, cutting off the user's reply. The reader now recovers the full message instead of the display-capped projection. The capped display history keeps its provenance, and text the model writes itself, including a literal marker, is left alone.
Also addresses #82121 at the producer, without global stripping in the sanitizer. Thanks @jayzhou2309.
Proof: real isolated Gateways in a secretless container, with a scripted mock model, over the Gateway WebSocket and the actual run-wait reply reader.
- On main (`f57a952ab0aa`), a 14,329-character report came back as an 8,018-character copy ending in the truncation marker. With this change, it comes back complete and exact.
- Display-cap provenance is kept in history on both builds.
- Normal prose and an intentional standalone marker are unchanged.
- Three recovery cases fail on main and pass here, and the regression suite passes 63 tests.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
(cherry picked from commit ebf4bd70ca)
* fix(update): avoid terminal errors during pnpm updates (#161920)
The updater spawned the package-manager install with inherited terminal stdin but captured stdout/stderr; pnpm 12.1.0's global build-approval prompt then failed with "IO error: not a terminal" and aborted the install. The pnpm package-install step now receives EOF on stdin (pnpm skips the prompt when stdin is not a terminal; no --yes, so consent is unchanged); npm and Bun installs and every other bounded command keep their stdin behavior. docs/install/updating.md gains the first-hop recovery for already-installed updaters.
Closes#161866
Thanks @alkor2000 for the diagnosis and fix.
Co-authored-by: alkor2000 <200923177@qq.com>
(cherry picked from commit bbe0d114e4)
* fix(agents): preserve spawned session cwd at ingress (#162308)
Visible child sessions persist lineage and a managed cwd without an inherited workspace. Keep that cwd for subsequent runs while preserving explicit inherited workspace precedence and sandbox ownership checks. Existing session rows require no migration.
(cherry picked from commit 43d29bbdb1)
* fix: finished subagents keep child slots occupied after delivery expires (#162368)
* fix: release child slots when descendant delivery is suspended
* test: wait for collector completion publication
* ci: install ripgrep for baseline ratchets
* Revert "ci: install ripgrep for baseline ratchets"
This reverts commit 616d5cc3e04e77cf97c5cc6b7a201fb02e966a84.
* Revert "test: wait for collector completion publication"
This reverts commit a9902f4633b12fddccfe4780b3f08ed5ca52cda7.
---------
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
(cherry picked from commit c03ee95e10)
* fix(skills): sandboxed turns fail with EACCES when refreshing read-only skill copies (#162295)
* fix(skills): refresh read-only sandbox skill copies
Repair owner access through confined directory descriptors before removing stale materialized skill trees. Preserve read-only sandbox mounts and source permissions. Cover prior read-only copies, nested cleanup, symlink confinement, and mount identity.
* fix(skills): unlink stale top-level skill symlinks
(cherry picked from commit 81d89fb43b)
* fix: steered message no longer replaces a question whose model request failed (#162439)
Fixes: with the default `steer` queue mode, a message sent while the first message's model request is failing replaces the first message. The chat shows `⚠️ <model> request failed.` and then only the answer to the second message. The first question is never retried or answered.
When a model request fails while a newer message is waiting to steer into that turn, OpenClaw retries or falls back for the first message as usual. The newer message then runs as its own turn. Both questions get answered in order. The failure notice only appears if the first message's retries and fallbacks are all exhausted, the same as when nothing was steered.
Root cause: after a model request failed, `AgentSession.handlePostAgentRun` still continued the session with the queued steer. The embedded runner owns retries for a failed request (`agent-project-settings.ts` turns session auto-retry off, #73781), so the failed message was never retried. The steered message ran in its place, against a transcript where the failed message looked handled. Per-input answer segments (#149925) then correctly reported the first segment's failure, so the user got `⚠️ request failed` plus the second answer.
Provenance:
- `steer` became the default in `4a6e10ece8` ("feat: default queueing to steer") and #77023, to keep the active turn responsive without starting a second run.
- The post-run `hasQueuedMessages() ? "continue"` came from #109709 so that messages queued by `agent_end` handlers are not stranded. That reason still holds after successful turns. This PR only excludes failed requests, where continuing skips the run owner's retry and fallback.
- Steered channel input already waits for its transcript commit (`agent-runner-steer-adoption.ts`, `waitForTranscriptCommit: true`). When the session settles without committing it, the steer is withdrawn from the runtime queue and the parked reservation falls back to the ordinary follow-up queue. That path runs after the reply operation clears. No new queue or retry logic is added.
Fix: one condition. After a failed request, queued input no longer continues the session (`agent-session-prompting.ts`). Docs: `docs/concepts/queue-steering.md`.
Hermes comparison: Hermes injects a steer only at a role-safe boundary after a tool result. With no tool row yet, or when the request fails, it requeues the steer as the next turn's user message, and each queued message gets its own turn. This change gives OpenClaw the same outcome by reusing its existing follow-up fallback.
Relation to #161069 (stalled-turn recovery): that PR covers turns aborted by stuck-session recovery. This one covers a provider failure while a steer is pending. The two don't overlap.
- Fewer model calls: before, a pending steer triggered one extra model call inside the failed run. Now it doesn't.
- The first message keeps its existing retry budget (`MAX_EMPTY_ERROR_RETRIES = 3`, plus the configured fallback chain). Nothing new is retried.
- A handed-back steer is queued at most once. The fallback drops its `steerPending` marker, so it drains as an ordinary follow-up turn and is never steered again. If that turn also fails, it ends with the normal failure notice and is not requeued. The worst case is one follow-up turn per inbound message.
- Tested: the session regression test asserts exactly one model request after the failure, with no continuation. The requeue-once path is existing behavior covered by `src/auto-reply/reply/queue/enqueue.steering.test.ts` (6/6 on this head).
**Telegram Test Server** (`telegram-e2e-userbot`, Convex-leased credential, real QA user in a DM, default queue mode). A deterministic mock provider fails agent-turn requests whose newest question is `17*23` three times (`response.failed` after 8 s). Like the live model in our earlier proof, it answers only the newest user question. The QA user sends `What is 17*23?` at +0.2 s and `Also, what is 19*21?` at +8.2 s, so the second message arrives during the first failing request.
- Before (built head with only this line reverted): the SUT sends `⚠️ openai/gpt-5.5 request failed.` at +13.1 s, then `399` at +14.1 s. The request after the failure carries both questions (`developer,user,user,user`). The first request is never retried, and `391` is never sent.
- After: three failing first-question requests (`[empty-error-retry]` attempts 1/3 to 3/3), none carrying the second question. The fourth request answers `391` at +30.5 s. The second message then runs as its own turn (`developer,user,user,assistant,user,user`) and answers `399` at +31.5 s. No failure notice is sent.
**Unit test:** `agent-session-loop-next-turn.test.ts` "does not answer a steer in place of a failed request". On `main` the steer is committed into the failed turn (`promise resolved instead of rejecting`). With this change there is one request, the steer's commit wait rejects so its caller requeues it, and nothing is left queued. File: 24/24 pass. Related suites pass: `sdk`, `agent-session-loop-correctness`, `attempt-prompt-submit` (+ steering, retention), `attempt-stream-prepare`, `attempt-session-replay` and `provider-review-continuation`, 231 tests.
**Telegram with a live OpenAI model** (same Telegram Test Server DM; a small local proxy in front of the real OpenAI API fails only the first two agent requests that carry 17*23 without 19*21, and forwards everything else to the live model):
- Before (this line reverted in the build): one failing request, then the next request carries both questions and goes to the live model, which answers `399`. The bot sends `⚠️ openai/gpt-6-astra request failed.` (+20.4 s), then `399` (+21.0 s). 17*23 is never answered.
- After: two failing requests for the first question alone, then the live model answers the retry with `391` (+24.1 s). The second message's own turn gets `399` (+26.0 s). No notice.
- Both runs used the fix build, before the branch was rebased onto newer main; the production change is the same. For the before run, only the fix line was reverted.
**Telegram, when the first question fails on every retry and fallback** (same harness; the mock fails every agent request whose newest question is 17*23, and the agent has a fallback model configured): 8 requests for the first question, 4 on `gpt-5.5` and 4 on `gpt-5.5-fallback`. None of them carries the second question. The model-fallback log shows `gpt-5.5` failing over to `gpt-5.5-fallback`, then nothing left to try. The SUT sends exactly two messages: one `⚠️ openai/gpt-5.5-fallback request failed.` at +70.9 s, then `399` at +71.9 s from the second message's own turn. No other notices.
**Web UI `chat.send` (real Gateway, operator WebSocket client of the kind the TUI uses, same mock):** the client sends `What is 17*23?`, then `Also, what is 19*21?` 8 s later, during the first failing request.
- Before (this line reverted in the build): the client's only answer is `399`. The request after the failure carries both questions (`developer,user,user,user`). `chat.history` ends `user 17*23, user 19*21, assistant 399`, so the first question is unanswered.
- After: the steer is withdrawn when the failed attempt settles. Its `chat.send` run closes with an empty final (+17.1 s), and `chat-send-agent-dispatch.ts` dispatches it again; it runs as a follow-up and is not steered into the retry. The first run retries three times, with no retried request carrying the second question, and answers `391` (+33.7 s). The second message's own run answers `399` (+34.0 s). `chat.history`: `user 17*23, assistant 391, user 19*21, assistant 399`.
Why a withdrawn Web UI steer can't be re-steered into the retry: with explicit `queueMode: "steer"`, chat.send's fallback dispatch passes `messageInjectionDisposition: "rejected"` (`src/gateway/server-methods/chat-send-agent-dispatch.ts:388-390`). `runReplyAgent` steers only when the disposition is `"none"` (`src/auto-reply/reply/agent-runner-run.ts:341-346`), so the message is queued as a follow-up instead (`:389`). With the default mode, chat.send makes no injection attempt of its own (`chat-send-admission.ts:272-275`) and the message takes the channel path: its parked reservation falls back with its steer marker removed (`queue/enqueue.ts:307-318`), and the queue drains only after the reply operation clears (`agent-runner-steer-adoption.ts:110-122`).
**Wall times** (`pnpm test <file>`, shared, heavily loaded host):
- `agent-session-loop-next-turn.test.ts`: 24/24, 121 s wall (vitest 106 s). The new test signals request start and steer acceptance with deferreds; it no longer polls.
- `enqueue.steering`: 6/6, 78 s.
- `attempt.queue-message`: 16/16, 98 s.
- `agent-session-handoff-adoption.integration`: 1/1, 63 s.
- `agent-runner-steer-adoption.question-recovery`: 25/25, 97 s.
- `chat-send-steering-custody`: 17/17, 90 s.
- `agent-runner.runreplyagent.e2e`: 223/224 under load. The one failure, "keeps the replacement source when retired admission completes", passes alone (60 s wall). That file mocks the embedded agent, so it never reaches the changed session code.
qa-lab note: the qa-channel serializes inbound per account, so the same-sender steer couldn't be reproduced there. The proof uses the Telegram Test Server instead.
LOC vs `origin/main`: production +3/−1, test +46, docs +1 (changed line).
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
(cherry picked from commit 4d755295bc)
* fix(telegram): restore progress when reply hooks suppress previews (#161546)
* fix(telegram): deliver progress when reply hooks suppress previews
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* fix(telegram): honor global block-streaming opt-out
Preserve the explicit global off preference when reply hooks suppress previews, and cover both hook contracts through Telegram HTTP delivery.
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* docs(telegram): clarify hooked fallback for multi-agent turns
Clarify that configured GroupThread turns retain their existing block policy: their previews are disabled independently of reply-modifying hooks. The new forced fallback applies to ordinary single-agent turns. No delivery-policy or test changes.
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
---------
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
(cherry picked from commit 54a6149549)
* fix(webchat): keep saved replies visible during history refresh (#162763)
* fix(webchat): keep saved replies visible during history refresh
Fence pre-commit history snapshots at transcript publication before run settlement. Reuse the existing queued refresh so both active and late background replies survive history races.
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* fix(webchat): keep saved replies visible during history refresh
Worked on by:
- @VACInc
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
OpenClaw-Publication: 18964aa1-fac1-42b9-8777-682b1d10ef09
---------
Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com>
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
(cherry picked from commit e42c52b7d8)
* fix(cron): stale automatic tool lists block scheduled jobs from tools their owner has (#162432)
Related: #130753, #137832, #147969
Fixes: some scheduled jobs created by an agent fail for months because a tool list saved by an older OpenClaw build is missing tools the creator actually had, such as the native shell. In our setup, a monthly group job that runs `node <script>` delivered nothing in August, delivered nothing in September (the run still reported `ok`), and posted a blocker in October. Its saved list had 31 tools and no `exec`.
User impact: an agent-created agent-turn job that does not name specific tools now gets the same tools as its owner conversation at run time, like a job an operator creates without `--tools`. Existing jobs with an automatically saved creator snapshot behave the same way from their next run. Nothing stored is rewritten: no migration and no backups. Explicit tool lists, script payloads, condition triggers, and jobs bound to captured Codex app authority keep their stored list.
Tradeoff, approved by the maintainer (Ayaan): a per-sender tool policy on the creating owner, or a plugin hook that narrowed the creating turn, no longer limits these default jobs. Only owners can create automations from chat, and subagents cannot create them.
**History of the saved list.**
- #91499 introduced it so a delayed run cannot do more than its creator could.
- #112483 made every agent-created job store one, because runs have no sender.
- #112661 made scheduled runs re-apply the owner session's group policy and every non-sender limit, keeping the stored list as the upper bound.
- #137832 fixed native tool capture for new jobs only, and deliberately did not widen stored lists.
- #147969 added a Doctor advisory. It only fires for claude-cli, so it never covered Codex-harness or built-in OpenAI jobs like ours.
**Root cause.** When no tool list was given, OpenClaw saved a frozen copy of the creating turn's tools instead of treating the job like an operator `*` job. Every capture bug (missing native tools, late configured MCP, renamed tools) then stayed in the job permanently.
**Fix.** This follows Hermes, which keeps no creator snapshot: `cron/scheduler.py` `_resolve_cron_enabled_toolsets` reads toolsets from config at run time.
- **New jobs.** An agent-turn create or update with no list, or `*`, stores `["*"]`. That is the same value operator jobs store, so the job's tools match a normal turn in its owner conversation. Script payloads and condition triggers still store the creator's concrete tools, because a script reaches MCP only through servers its list names. Jobs whose creator captured Codex app authority also keep the concrete list, because that authority is bound to it.
- **Existing jobs.** One helper, `resolveCronRunToolsAllow` in `src/cron/tools-allow.ts`: a stored automatic snapshot (`toolsAllowIsDefault`) runs as `*` when it has a valid scheduled owner policy, no condition trigger, and no Codex app authority. Otherwise it keeps its stored list. Every execution consumer of the stored list uses it: the run payload, the command-prompt preflight, and the scheduled message authority.
- **Script transitions.** A `*` job that becomes a script, or gains a condition trigger, captures the creator's concrete tools.
- **Exec pin.** A `*` list keeps the creator's exec host pin.
- **No new noise:** automatic snapshots stay excluded from the `web_search` provider warning, as on main.
- **Deleted, now pointless:** both Doctor advisories about incomplete automatic snapshots, the run warning about pre-MCP snapshots, and two exports nothing uses anymore.
Review note: on claude-cli, a `*` job runs without a CLI tool cap, so Claude's native tools behave exactly as in a normal chat turn in that conversation. This PR introduces no new path around `tools.deny` that a chat turn doesn't already have.
Live-model Telegram proof (Telegram Test Server DM, leased team credential, live `openai/gpt-6-astra` reached through a forwarding proxy that stands in for the runner's mock provider; the runner harness itself is unchanged). This reproduces the shape of the original incident:
- The tester DMs the bot, which creates the owner conversation.
- A job owned by that conversation is added. Its stored list is an old-style automatic snapshot `["automations","message","read"]` plus `toolsAllowIsDefault: true`, with no `exec`.
- The payload is `Run: node scripts/split-report.mjs and post its output line verbatim`. The workspace script prints a random nonce.
- The job is run once (`cron run --wait`), with announce delivery to the DM.
| Build | `exec` offered | Model action | What arrived in the DM | Run |
|---|---|---|---|---|
| base 94f5a8d (main before this PR) | no | `tool_search` ×2, then gave up | "Could not run node scripts/split-report.mjs: no command-execution tool is available…" | error |
| **this PR, head 5b78cb7** | **yes** | `exec {"command":"node scripts/split-report.mjs"}` | "**SPLIT-REPORT 93C53909**: general 41, design 17, ops 9" (the exact script output, with this run's random nonce) | ok, delivered |
| head 5b78cb7 with `tools.deny: ["exec"]` | no | `tool_search`, `read`, then gave up | "Could not run node scripts/split-report.mjs: no command-execution tool or paired node is available…" | error |
In every run, the stored job kept `["automations","message","read"]` plus the marker. Before and after use the same scenario and driver; only the checkout differs.
Update and live proof: published `openclaw@2026.9.7`, then this branch at the exact head (2f5099d), on the same state directory. Mock provider. Every process ran under a temporary `HOME` and state directory. Each job's message makes the model call `exec` with `touch <effects>/<job>`.
1. 2026.9.7 created both jobs through `cron.add` (scheduled policy `trusted`). With the Gateway stopped, the "stale" job was given the old automatic-snapshot shape `["automations","message","read"]` plus `toolsAllowIsDefault: true`. sha256 of both stored rows: `609358837…`.
2. Runs:
| Build / config | Job | `exec` offered | Side effect | Run |
|---|---|---|---|---|
| 2026.9.7 | stale automatic snapshot | no | absent | error |
| 2026.9.7 | explicit `["read","message"]` | no | absent | error |
| this branch | stale automatic snapshot | **yes** | **created** | ok |
| this branch | explicit `["read","message"]` | no | absent | error |
| this branch, owner policy narrowed to `tools.deny: ["exec"]` | stale automatic snapshot | no | **absent** | error |
| this branch, `tools.deny: ["exec"]` | explicit `["read","message"]` | no | absent | error |
3. After the branch runs, the stored rows were byte-identical (same sha256 `609358837…`), and job ids and lists were unchanged. Nothing was migrated.
An earlier run at e151ea3, with the same harness, also covered a snapshot bound to Codex app authority: `exec` was not offered, the file stayed absent, and the stored row was unchanged.
Tests:
- `run.tools-allow.test.ts`: a stored automatic snapshot `["message","read"]` reaches the embedded run as `["*"]`, with the owner's scheduled policy intact. It fails on main with `["message","read"]`.
- `cron-tool-creator-cap.test.ts`: a default agent turn stores `["*"]`, while a trigger script and a Codex-app creator keep the concrete snapshot.
- `run.tools-allow.test.ts`: snapshots without a valid owner policy, or behind a condition trigger, keep their list. Both cases fail on the previous head.
- `run.tools-allow.test.ts`: no `web_search` warning for an automatic snapshot that kept its list. This fails without the exclusion.
- `run.message-tool-policy.test.ts`: a self-edited automatic snapshot runs on CLI with no cap.
- `run.tools-allow.test.ts`: a legacy `Command to run:` prompt from an automatic snapshot without shell tools now runs instead of being rejected.
- `jobs-tool-policy.test.ts`: scheduled message authority is admitted for an automatic snapshot that lacked `message`.
- `cron-tool-creator-cap.test.ts`: a `*` agent turn converted to a script captures the creator's concrete tools.
- These three regressions fail on the previous head. `node scripts/check-changed.mjs` passes.
- Explicit-list, exec-pin and gateway creator-transport suites pass. `pnpm tsgo:core` passes.
No new path triggers a model call or a job run. The change only selects which tool list an already scheduled run uses.
LOC vs main: production +98/-250 (net -152), tests +140/-373, docs +19/-11.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
(cherry picked from commit 08498ec40d)
* fix(cron): deliver manual runs started from an agent turn after that turn ends (#162780)
Fixes: when an agent turn starts an automation with `automations` `run` (for example the owner says "run my report now"), the run fails with `attempt disposed before transcript write` once that turn ends. For a `sessionTarget: "current"` job the report is never delivered. For other jobs the report is still sent, but its copy in the conversation transcript is lost (a WARN).
User impact: "Run it now" from a chat delivers the report, the same as a scheduled run. No new settings, messages or warnings.
Root cause: the `automations` tool calls `cron.run` through the in-process gateway, so `enqueueRun` runs inside the calling turn's AsyncLocalStorage, including its owned transcript-write context (`withOwnedSessionTranscriptWrites`). The queued run inherits that context. When the run writes into the caller's session (the current-session report or the delivery mirror), `runWithOwnedSessionTranscriptWrite` matches the session and routes the write through the caller's attempt lifecycle. By then that turn has ended and its lifecycle is disposed, so the write is rejected.
Provenance:
- #104595 (`b6e95201f8`, "retain detached manual run admission") made the queued manual run its own gateway root (`runWithGatewayIndependentRootWorkContinuation(..., "cron:manual-run")`), so it no longer depends on the caller's request. That helper resets only gateway work admission, not other per-turn context.
- Later, #115404 (`d16e33e08e`) and #121113 (`ce53f7e82e`) carried the turn's transcript lifecycle and writer fence in an AsyncLocalStorage, so nested writes stay serialized and fenced to the running attempt. #115404 also added `runWithoutOwnedSessionTranscriptWrites` for work that outlives its turn. Subagent completion (`b3d3f860e4a`), `sessions_send` follow-ups, media-generation completions, and session event wakes all use it. The manual-run detach point from #104595 never got it.
Fix: `enqueueRun` starts the queued run outside the caller's owned transcript context (`runWithoutOwnedSessionTranscriptWrites` around the existing independent root continuation). It reuses the helper the sibling detached paths use; no new mechanism. The invariants of those PRs still hold: the run is still a tracked independent gateway root (#104595); the caller's activation and commit guards still run before acceptance, and they don't touch transcripts; writes made during the turn are still fenced to its attempt (#115404/#121113). The queued run gets its own lifecycle when its attempt starts, the same as a scheduled run.
Hermes: `trigger_job` doesn't run the job inside the caller. It marks the job for the scheduler's next tick, so a manual run always starts from scheduler-owned context. This change does the same for OpenClaw's queued manual run, at the point where it detaches.
- **Live, Telegram Test Server DM, live OpenAI model** (`telegram-e2e-userbot`, Convex-leased credential, fresh isolated gateway built from each ref, ports 19951/19952, provider slot = pass-through proxy to the OpenAI API; the tester is the configured owner, `session.dmScope: "main"`). The owner creates `current-check` (`sessionTarget: "current"`, `payload: agentTurn "Reply with exactly CURRENT-REPORT"`, `delivery: announce`, disabled), then sends: "Use the automations tool to run the automation named current-check now (action run, runMode force …). When the run call returns, reply with exactly RUN-DONE".
- Command (secrets omitted): `E2E_MOCK_SERVER_PATH=<live proxy> node .agents/skills/telegram-e2e-userbot/scripts/run-mock-sut-user-e2e.mjs --gateway-port 19951 --mock-port 19952 --dm --timeout-ms 420000 --scenario <scenario> --record events.ndjson --output summary.json` → exit 0 on both refs. The scenario waits 90 s after the last reply.
- **Before, `origin/main` `26bcc94353c`:** DM shows `SETUP-DONE` (20.8 s) and `RUN-DONE` (87.7 s), then nothing more. The run receipt is `error: attempt disposed before transcript write`. The gateway logs it at the cron run session (`outcome=error`).
- **After, this branch:** DM shows `SETUP-DONE` (24.2 s), `RUN-DONE` (94.4 s), **`CURRENT-REPORT` (96.4 s)**. The run receipt is `ok`, and the gateway log has no `attempt disposed` lines. The WARN/ERROR lines are the same harness startup lines as in the before run.
- **Regression test** (`src/cron/service/ops.regression.test.ts`, "runs a manual run queued from an agent turn outside that turn's transcript lifecycle"): `enqueueRun` is called inside a real attempt transcript lifecycle (`createEmbeddedAttemptTranscriptLifecycle` + `withOwnedSessionTranscriptWrites`). The calling turn disposes, then the run writes into the caller session through the real `runWithOwnedSessionTranscriptWrite`. On `origin/main` the run finishes `error`; with this change it finishes `ok` and the write lands.
- `pnpm exec vitest run src/cron/service/ops.regression.test.ts src/cron/service/ops.run-admission-cleanup.test.ts src/cron/service/ops.run-admission.test.ts src/cron/service/manual-ack-durability.test.ts` → 4 files, 46 tests passed.
- `node scripts/check-changed.mjs` on `997394f3a3a` → exit 0 (format, lint, core and test typecheck, line-cap and dead-export ratchets).
- Test cost: `pnpm test src/cron/service/ops.regression.test.ts --maxWorkers=1` → 15 tests passed, 33.9 s wall (vitest 30.7 s, mostly transform; the new test itself takes about 40 ms).
- Size vs `origin/main`: production +6/−2 ignoring whitespace (+75/−71 raw, because the existing callback is re-indented one level); test +58.
Not changed here: in both runs the `current` run waits behind the caller turn, and the tool's run wait returns "Not finished yet" after about 60 s before the turn replies. That wait is a separate issue.
One AsyncLocalStorage `exit` per queued manual run. No new state, timers, retries or queues. The run's own transcript writes go through its own attempt lifecycle, as for any scheduled run.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
(cherry picked from commit 510beb8d52)
* fix(codex): prevent Gateway heap exhaustion with large agent fleets (#162912)
Closes#162802
Fixes Gateway heap exhaustion during Codex session discovery on large agent fleets. Thanks to @609NFT for the report and allocation profiles.
Large fleets can finish startup and retain their native Codex session catalog without retaining a whole fleet configuration for every agent/home pair. No schema, stored-data, configuration, or catalog-output changes are required.
The existing config-identity cache now owns one captured configuration per generation. Agent/home entries keep their separate directories and cloned connection options, but share that captured configuration. Config reload isolation and source-specific backoff remain unchanged.
No overlap with Pash/Sarah changes.
Codex source inspection: `codex-rs/app-server/src/request_processors/thread_processor.rs:2524–2608` in the sibling Codex checkout confirms that `thread/list` consumes pagination and filter parameters; the fleet configuration capture being repaired belongs to OpenClaw, not the native listing protocol.
Compared baseline `7ef388bac8` with candidate `f0ae922f60922f2b357bc3d5af22a2142799fbe0` in isolated Linux arm64 Docker containers, Node 24.21.0. The synthetic profile contained 739 agents, eight explicit Codex homes, and three native sessions (380,439 bytes of configuration).
| Check | Baseline | Candidate |
| --- | --- | --- |
| Real Gateway, 2,560 MiB heap | Heap OOM at 100.3 s | Survived the six-minute window; clean shutdown |
| Sampled catalog allocations under `resolveRequestOptions` → `structuredClone` | 2.248 GB | 5.48 MB |
| Post-startup heap, 150–360 s | Process already terminated | Bounded at 377–566 MiB |
| Complete startup, separate 6,144 MiB control without profiling | 98.34 s | 93.83 s |
| HTTP listening, same startup control | 50.52 s | 49.56 s |
The larger-heap startup control lets the baseline complete startup rather than comparing a successful candidate against an OOM.
- Real `sessions.catalog.list` Gateway RPCs returned exactly the three expected native session IDs for agents `agent000`, `agent001`, and `agent738`.
- Exact-head live gate: a real OpenAI `gpt-5.4-mini` turn through the candidate Gateway's Codex harness returned `CATALOG_LIVE_OK` on the 739-agent profile (5.449 s). Its temporary read-only credential file and container were removed.
- Regression failed on baseline: four fleet config clones instead of one across multiple agents/homes. It passes on the candidate and also checks reload allocation isolation; the regression itself took 5 ms.
- All 38 focused tests passed across `session-catalog-listing-cache.test.ts`, `session-catalog-request-lifetime.test.ts`, `session-catalog-homes.test.ts`, and `session-catalog-backoff.test.ts` (38.67 s wall).
- Standalone `pnpm test extensions/codex/src/session-catalog-listing-cache.test.ts --maxWorkers=1` passed (32.61 s wall, including runner preparation).
- Built and exercised the real candidate runtime with `pnpm build`. Broader static and project-wide validation is left to CI.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
(cherry picked from commit fd0124688d)
* fix(agents): preserve complete CLI subagent answers (#162843)
Fixes#162777.
The claude-cli transcript writer now records `__openclaw.runId` on the terminal assistant row, matching the field other writers set and the completion reader expects. CLI-backed subagents now deliver their complete final answer instead of the truncated-by-retention fallback.
Proof: real isolated Gateway with a fake claude-cli backend. main delivered the fallback and dropped the tail of a 14,025-character child answer; with this change the parent received all 400 lines including the tail marker. A failed CLI child still reports its error with no fabricated final. Live control: one real OpenAI embedded subagent turn delivered its exact answer. Regression tests fail before and pass after.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
(cherry picked from commit d7cea26df2)
* fix(reef): admit gpt-6.1-sol as a documented immutable guard model (#162955)
OpenAI's gpt-6.1 generation publishes no dated snapshots either, so the Reef
guard rejected the Team configuration that pinned gpt-6.1-sol and the channel
could not start. Admit the exact id beside the gpt-5.6 ids, with the same
documented residual risk; bare family aliases stay rejected.
(cherry picked from commit d77cdd2af6)
* fix: remote MCP plugins fail to start in native agent sessions (#162376)
* fix: remote MCP plugins fail to start in native agent sessions
* test: align remote MCP assertions with normalized transports
(cherry picked from commit 4587ef903c)
* fix(llama-cpp): managed server fails on clean Windows installs (#163093)
* fix(llama-cpp): ship Windows VC runtime app-local
* fix(llama-cpp): make Windows runtime staging fallback-only
(cherry picked from commit 45762faaa8)
* fix(secrets): keep store entry kind when rotating a value (#158968) (#160926)
* fix(secrets): keep store entry kind when rotating a value
`secrets store set <NAME> --value-file <path>` resolved the entry kind
from the name heuristic whenever no host-policy flag was passed, so
rotating the value of an entry created with `--kind secret` under a name
the heuristic does not classify as sensitive silently converted it into a
readable `env` entry and deleted its `--allow-host` allowlist. `import`
classified every entry from its name for the same reason.
Both paths now inherit the stored kind, and that inheritance is resolved
inside the authoritative store write transaction rather than captured
before the CLI awaits its value or its confirmation. A protection change
committed while one of those waits is pending is therefore no longer
overwritten by the pending write, and the value it carries lands under
the newer policy instead of reaching agent subprocesses as plaintext.
Because a kind resolved at write time can invalidate a value the CLI
already accepted, `import` now hands its whole batch to a single
store-owned write. `writeSecretStoreEntries` resolves every live kind and
validates every resulting entry inside one transaction before it writes
any of them, so an entry that a concurrent protection change turns into
an empty secret still fails the command with nothing committed, matching
the existing no-write-on-validation-failure contract.
An explicit `--kind` still overrides both the stored kind and the name
heuristic in either direction.
(cherry picked from commit 66a8c01c8db76bf63ab79c506880f462856b2d1e)
* fix(secrets): satisfy CI pools, formatters, and assertion ratchet
- Run the kind-inheritance purge through the expiry kernel directly so it
does not require the host-broker state worker (the CLI project executes
test files off the main thread).
- Replace uncommented kind casts with runtime narrowing for the assertion
SAFETY ratchet; reformat to oxfmt.
Co-authored-by: yetval <yetvald@gmail.com>
* chore(ci): rerun checks
* ci: retrigger checks for flaky shard reruns
* ci: retrigger checks (flaky infra shards)
* test(ci): dedupe pdf-tool.resources in database worker core paths
670e3fb4ea (#161270) re-added src/agents/tools/pdf-tool.resources.test.ts
to databaseWorkerCoreTestFiles, which already listed it (16b17b22eb). The
changed-node shard planner now rejects repeated files in a split timing
generation, so preflight threw 'split timing generation repeats files for
core-runtime-infra-storage-state' for every PR touching any non-README path.
* fix(secrets): enforce literal input safety at commit
Carry argv provenance to the store transaction and refuse literal values when
kind inheritance resolves to a secret. Preserve explicit env reclassification,
cover concurrent protection, and clarify the set/import kind rules.
Strengthen rotation-value, committed-kind output, and redacted batch no-write
regressions. Existing Gateway writer paths remain unchanged.
Co-authored-by: zachisfine <131436334+zachisfine@users.noreply.github.com>
Co-authored-by: yetval <yetvald@gmail.com>
---------
Co-authored-by: yetval <yetvald@gmail.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: zachisfine <131436334+zachisfine@users.noreply.github.com>
(cherry picked from commit 0ae13a5620)
* chore(release): prepare 2026.8.35
* fix(release): adapt backports for 2026.8
* docs(changelog): refresh 2026.8.35 ledger
* fix: repair extended-stable backport integration
* fix: preserve cleanup failures on 2026.8
* test: complete worker cleanup before gateway close
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: wangmiao0668000666 <wang.miao86@xydigit.com>
Co-authored-by: Tak Hoffman <781889+Takhoffman@users.noreply.github.com>
Co-authored-by: RoboClaw <services+roboclaw@openclaw.org>
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
Co-authored-by: Baumus <126391633+Baumus@users.noreply.github.com>
Co-authored-by: Y.B. <paranoyouz@gmail.com>
Co-authored-by: Ayush Tiwari <147244828+Ayushdevo@users.noreply.github.com>
Co-authored-by: Jay Zhou <zhoujunbai123@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: RileyJJY <0668000974@xydigit.com>
Co-authored-by: Altay <altay@hey.com>
Co-authored-by: Vito Cappello <hixvac@gmail.com>
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
Co-authored-by: alkor2000 <131229172+alkor2000@users.noreply.github.com>
Co-authored-by: alkor2000 <200923177@qq.com>
Co-authored-by: Sam Armstrong <armstrongsam25@gmail.com>
Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com>
Co-authored-by: Jony <619963502@qq.com>
Co-authored-by: Patrick Erichsen <patrick.a.erichsen@gmail.com>
Co-authored-by: zachisfine <131436334+zachisfine@users.noreply.github.com>
Co-authored-by: yetval <yetvald@gmail.com>
* fix(models): reuse prepared catalogs when browsing inventory
Inventory callers forced full discovery on every browse while trying to
refresh auth-stale content. Give the prepared catalog owner a distinct stale
refresh intent, preserve explicit forced refresh, and leave ordinary turn
reads on their existing facts. Apply the policy to commands, CLI, and RPC.
Correct the Ollama fixture response contract and await full catalog readiness.
Verify discovered-only models survive callbacks and RPC while discovery is
unavailable, and verify explicit refresh still reaches the provider.
Proof: before the fix, the real Gateway made 8 discovery calls after a 3-call
warmup, and the owner policy regressions failed. Prepared catalog and auth
reload owner tests now pass (43 tests).
(cherry picked from commit 1dfddb6e8fa7f57fcc60c4f01384b679e88ad77a)
(cherry picked from commit 7e9563ccff)
* test(release): align extended-stable validation fixtures
* test(sessions): isolate heartbeat wake cleanup
* fix(deps): update joi for extended stable
* test(gateway): share title-retention disk fixtures across isolated readers (#134642)
* fix(deps): update brace-expansion to 5.0.12
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(plugins): recover orphan installs without weakening ownership
Closes#134321
## What Problem This Solves
Operators could not run `openclaw plugins update --all` or uninstall a path-source plugin after both tracked plugin paths disappeared. Registry refresh succeeded but left the durable install record behind, so both commands failed on the same ownership check.
## Why This Change Was Made
Plugin lifecycle ownership now distinguishes an active package from an exact orphan record. Update and uninstall may act on an orphan only after duplicate-path and conflicting-plugin checks pass. Inspection, capability consent, active packages, and restored packages keep strict package ownership.
Post-update reconciliation accepts an unchanged safe orphan while another package updates. If an orphan install record changes, the replacement must restore a strictly valid package before any update can commit.
Uninstall removes owner-keyed plugin and channel policy plus the durable install record. Child policy remains based only on authoritative discovered children.
## User Impact
Operators can update their remaining plugins and remove stale path-source install records without editing OpenClaw state by hand. Ambiguous, conflicting, and invalid replacement packages still fail closed.
## Evidence
Current `main` at `77fee95791` failed through the real command-line flow after a local plugin was installed and both tracked paths were moved away:
```text
$ openclaw plugins update --all --dry-run
Plugin "orphan-proof" package owner "orphan-proof" has no authoritative runtime child list.
Exit 1
$ openclaw plugins uninstall orphan-proof --force
Plugin "orphan-proof" package owner "orphan-proof" has no authoritative runtime child list.
Exit 1
$ openclaw plugins registry --refresh
Plugin registry refreshed: 0/0 enabled plugins indexed.
Exit 0
$ openclaw plugins update --all --dry-run
Plugin "orphan-proof" package owner "orphan-proof" has no authoritative runtime child list.
Exit 1
```
Exact head `06de2f66734fd338e457f50a3060e2d625288274` passed the same flow:
```text
$ openclaw plugins update --all --dry-run
Skipping "orphan-proof" (source: path).
Exit 0
$ openclaw plugins uninstall orphan-proof --force
Uninstalled plugin "orphan-proof". Removed: config entry, install record.
Exit 0
$ openclaw plugins update --all --dry-run
No tracked plugins or hook packs to update.
Exit 0
```
The contributor head initially accepted a changed orphan install record with no authoritative children. The new owner-boundary regression failed with `{ ok: true }` before the maintainer repair and passes after it.
Validation:
- `node scripts/run-vitest.mjs src/plugins/manifest-registry-installed.test.ts src/plugins/management-service.uninstall-ownership.test.ts src/plugins/plugin-package-update.test.ts src/cli/plugins-cli.update.test.ts src/cli/plugins-cli.uninstall.test.ts` — 129 passed.
- Targeted Oxfmt and Oxlint — passed.
- Max-lines and assertion-safety ratchets — passed.
- Coercion, deprecated-API, dead-export, native-schema, and import-cycle guards — passed.
- `git diff --check` — passed.
- Local type checking was not run because host policy requires CI for OpenClaw type checks.
The proof fixture, isolated states, removed plugin copies, and proof manifest remain in the durable maintainer worktree.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(doctor): repair keyed multi-agent rosters without ownership (#134706)
* fix: prevent device pairing migration crash on contaminated records
Closes#134340
## What Problem This Solves
Operators upgrading from the retired JSON device-pairing store can have pending requests or malformed values inside `devices/paired.json`. One bad record aborts the full JSON-to-SQLite import, keeps every valid pairing out of SQLite, and repeats the failed migration on each gateway start.
The pairing migration failure is caught and logged. It does not stop gateway startup by itself. The field report also contained a separate legacy session-store blocker.
## Why This Change Was Made
The migration now validates every field that SQLite binds directly. It skips records with invalid required pairing data. When only optional metadata is malformed, it omits that field and preserves the usable pairing. The original legacy store is archived only after the safe import completes.
Validation includes required timestamps, optional `lastSeenAtMs`, and optional text columns. Separate warnings report skipped records and omitted optional fields. The SQLite mapper remains strict, and no schema or fallback path changes.
## User Impact
Valid device pairings survive an upgrade even when the legacy file also contains pending requests or malformed records. The gateway completes the import once and does not retry it after restart.
## Evidence
- Current-main red at `77fee95791`: a normal isolated gateway logged `NOT NULL constraint failed: device_pairing_paired.created_at_ms` for a pending-shaped leaf and retained `paired.json`.
- Anti-cheat red: changing the malformed leaf to `lastSeenAtMs: 1e100` logged `cannot store REAL value in INTEGER column device_pairing_paired.last_seen_at_ms`.
- Regression red on contributor head `ee7ea2bfbc5bde3e1b6f90ea58e4d070b7e7ff61`: the expanded boundary test failed with `Provided value cannot be bound to SQLite parameter 47`.
- Fixed-head boundary proof: `node scripts/run-vitest.mjs src/infra/device-pairing-migration.test.ts` passed 7/7; targeted formatting and Oxlint passed.
- Optional-preservation red on `4f47648f1d6bcf279697d4e8ec5044960bb27cf5`: the corrected regression failed because only one row imported instead of preserving three usable pairings.
- Final real gateway proof: startup returned `status=started`, imported three usable pairings, stored malformed optional metadata as `NULL`, skipped one record with invalid required data, archived `paired.json`, and restarted with no migration log.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(doctor): preserve per-agent memory search during upgrades
Closes#134256
Reported by @vdruts.
## What Problem This Solves
Fixes an issue where operators running `openclaw doctor --fix` after an upgrade could silently lose every per-agent memory-search override when Doctor converted `agents.list` to keyed `agents.entries` first. This removed settings such as `extraPaths`, `sources`, `experimental.sessionMemory`, and `enabled`.
## Why This Change Was Made
Doctor now visits both keyed agent entries and legacy list entries at the migration owner. It moves `memorySearch` before unknown-key cleanup, preserves explicit canonical values, and applies the existing provider, nested-field, and retired-path migrations to the moved data.
The shared agent-scope traversal replaces two duplicate traversal implementations. The regression was introduced by #112678, which moved roster normalization before the older list-only migration from #111527.
## User Impact
Per-agent memory-search settings now survive `openclaw doctor --fix` and remain attached to the correct agent. Operators no longer lose per-agent retrieval paths silently during this upgrade path.
## Evidence
- Red on current-main baseline `8abb80073c`: [Testbox run 33466170600](https://github.com/openclaw/openclaw/actions/runs/33466170600) removed the legacy block and retained only the pre-existing canonical field.
- Green on exact head `983df6b6ddf33ffad80f32874966ffeeaf0a803f`: [Testbox run 33467597600](https://github.com/openclaw/openclaw/actions/runs/33467597600) ran the real CLI, preserved supported fields and canonical precedence, verified the authored input, and produced byte-identical config on a second Doctor run.
- Regression proof: the strengthened Doctor process test fails on the baseline and passes on the head.
- `node scripts/run-vitest.mjs src/commands/doctor-config-preflight.process.test.ts` — 10 passed.
- `node scripts/run-vitest.mjs src/commands/doctor/shared/legacy-config-migrate.test.ts src/commands/doctor/shared/legacy-config-migrations.runtime.retired.memory-qmd.test.ts src/commands/doctor/shared/legacy-config-migrations.runtime.retired.test.ts` — 294 passed.
- Focused format, Oxlint, Doctor registry, coercion helper, max-lines, assertion safety, import-cycle, temp-directory, conflict-marker, and dead-export checks passed.
- Autoreview on the exact local diff: clean, no actionable P0 findings.
Agent task: https://chatgpt.com/codex/tasks/01a040d7-aa3d-7433-83eb-de7f49979443
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(doctor): detect shared auth migration by provenance (#134808)
* fix(cron): refuse stale doctor store rewrites
Prevent `openclaw doctor` from replacing cron definitions or runtime-owned state captured before its interactive repair prompt.
Doctor now rechecks a bounded definition-and-order fingerprint inside the SQLite write transaction. Definition changes stop the repair with a retry warning. Runtime columns and current runtime authority remain owned by the Gateway, while repaired authorization-input changes fail closed into recovery.
Preserve old embedded-authority migration and current quarantine recovery.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(gateway): refresh model catalog after auth changes (#134361)
* fix(models): refresh catalog content after credential-identity changes
* refactor(agents): split auth-profile mutation lineage out of runtime-snapshots
Split the max-lines boundary by moving persisted mutation lineage out of runtime snapshots; behavior is unchanged.
* refactor(agents): annotate the moved test-api global assertion for the safety ratchet
* test(agents): type the catalog worker mock and shrink the assertion baseline
* perf(agents): bound auth publication provider facts
* perf(gateway): keep auth status off catalog discovery
* fix(agents): reject retired catalog owners
* test(gateway): resolve the auth-status handler result before racing it
* fix(models): gate catalog refresh by read intent
* test(models): prove deferred catalog reads stay static
* test(models): type the unscoped catalog-refresh row context
* fix(openai): repair stale doctor route pins
Declare canonical OpenAI session-route ownership at the OpenAI plugin manifest boundary.
Clear automatic fallback-origin provenance when Doctor removes the stale model override. Preserve explicit model choices and valid locked harness sessions.
Fixes#101672.
Credit @849261680 for the original report-driven repair.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(auth): verify shared credential migration receipts (#134952)
Attach expected credential fingerprints to both credential-store receipt
variants so interrupted shared-store imports use the existing verification
and archive recovery flow. State-only receipts retain their own verification.
Regression coverage interrupts archival and removes the destination before
retrying both legacy auth file formats against both SQLite destinations.
The shared-store cases fail before this change; all 90 focused tests pass.
Related to #134608, reported by @erameson. This preserves the existing
non-replay policy for completed receipts; recovery of already-completed
receipts requires a separate maintainer decision.
* fix(cron): keep model aliases scoped to the selected owner (#135069)
Reuse the published model owner metadata for default, agent, subagent, payload, session, and Gmail selections. Consolidate override arguments while retaining agent-specific policy and existing precedence. Six regression cases fail before and pass after; 83 related tests pass. Related: #130324, #130706.
* fix(memory): resolve profile auth for compatible embeddings (#134744)
* fix(memory): resolve profile auth for compatible embeddings
* test(memory): close profile fixture SQLite handles
* fix(memory): preserve credential ownership in embedding auth
Resolve saved profile bindings through the canonical terminal auth capability while preserving literal, blank, explicit-header and destination behavior. Retain general model-auth precedence and shared guard policy; exercise real HTTP, SQLite and public CLI indexing/search. Reuse the canonical doctor CI fixture and repair executable-preflight test isolation.
Co-authored-by: 0x-Parzival <0x-Parzival@users.noreply.github.com>
* fix(memory): preserve auth policy for configured embeddings
Carry the selected provider API into the canonical auth-mode check, so direct OpenAI routes reject saved token profiles before embedding HTTP. Check auth migration readiness before classifying provider-entry credentials, including warm snapshots after a legacy credentials file is restored. Keep explicit remote credentials, literal keys, and canonical SQLite profile ownership intact.
Regression coverage exercises provider/API/credential-mode combinations and cold-empty, warm-empty, and populated canonical stores with real embedding HTTP and SQLite.
Co-authored-by: 0xParzival <145645180+0x-Parzival@users.noreply.github.com>
* test(pr): reuse landed cross-checkout fixture
Adopt the exact cross-checkout fixture already landed in main by #134740, resolving the overlapping cleanup while preserving the isolated executable command stubs. The embedding and auth implementation is unchanged.
Co-authored-by: 0xParzival <145645180+0x-Parzival@users.noreply.github.com>
---------
Co-authored-by: 0x-Parzival <0x-Parzival@users.noreply.github.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: 0xParzival <145645180+0x-Parzival@users.noreply.github.com>
* fix(auth): recover credentials stranded by completed migrations (#135346)
* fix(auth): recover inconsistent completed credential migrations
* test(auth): brace completed receipt recovery branches
* fix(auth): relogin repairs stale profile order after upgrade
Closes#135140
Reported by @mattcbianco.
## What Problem This Solves
Fixes an issue where users who signed in to OpenAI again after shared auth-store relocation could still get `Explicit auth order for openai has no usable profiles.` before a model request started.
## Why This Change Was Made
Login credentials and per-agent profile order have separate SQLite owners after relocation. Profile promotion now reads the effective credential view before it updates the agent-local order, and it preserves the inherited profile ID when it saves that order.
This keeps explicit order fail-closed behavior unchanged. It does not add a dispatch fallback or copy shared OAuth credentials into the agent database.
## User Impact
A successful OpenAI relogin now places the new usable profile before stale entries, so the next `openai/*` model call can start with the login that just succeeded.
## Evidence
- Current-main baseline: `2e3f734017`.
- Before: Codex reported a healthy ChatGPT login and OpenClaw listed a healthy OAuth profile, but the selected Codex runtime had no effective profile. Real Luna and Sol turns stopped before provider dispatch with `Explicit auth order for openai has no usable profiles.`
- After: OpenClaw device login stored the fresh OAuth profile, agent-local order became `[fresh, stale]`, and status reported a usable Codex runtime route.
- Live exact-head result: a real `openclaw agent --local` Luna turn returned exactly `ISSUE_135140_EXACT_HEAD_GREEN_2` through provider `openai`.
- Shipped CLI regression: `models auth login --provider openai --agent main` failed on the baseline because local order stayed absent; passed after the fix with `[fresh, stale]`.
- Focused CLI, owner, and sibling tests: 131 passed across login promotion, auth upsert/rollback, shared-store relocation, and runtime auth planning.
- Oxfmt, Oxlint, and `git diff --check`: passed.
- Autoreview P0: clean.
Production delta: `+4/-5` (net `-1`). Test delta: `+108/-0`.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(auth): preserve custom provider SecretRefs in Doctor
Fixes#135079.
Thanks @dannevang for the detailed report and two-provider reproduction.
## Problem
After an upgrade, `openclaw doctor` could treat custom-provider environment references and generated catalog values as plaintext API keys. It persisted them as auth profiles, and those profiles could shadow the provider's direct environment credential.
## Root cause
Config substitution discarded the fact that a resolved string came from an authored environment SecretRef. Doctor then classified model catalog strings without the provider's authored resolution facts. Runtime planning and status also treated stored profiles as usable ahead of an unqualified provider SecretRef.
## Fix
- Record resolved environment SecretRef provenance during config substitution.
- Make Doctor keep an authored provider SecretRef authoritative over root and plugin catalogs.
- Remove only persisted profiles whose values exactly match that provider's reference markers.
- Make runtime planning and model status prefer the authored provider SecretRef.
- Reuse the already loaded authored model config during status readback.
The production delta is +122 lines. The added code carries provider-specific pending/resolved provenance across config and worker boundaries and performs exact-value cleanup at the credential-migration owner. A global marker heuristic or downstream guard would not preserve this ownership or safely repair existing rows.
## Proof
- Baseline main `9d10dcb5d3`: Doctor migrated two `${VAR}` references plus stale or bare generated values; model status selected stored profiles over fresh direct environment values.
- Exact head `fc74120ab55d071e7c876b2f00c9db47032b8bb8`: Doctor migrated no model credentials, preserved both authored references, left no `authProfiles.store` row, and status selected two distinct environment credentials with zero profiles.
- Main commits after the baseline did not touch any changed file.
- Regression tests were red before the fix and green at the exact head.
- Focused tests, the remote changed-scope gate, and the exact-head remote build passed. The first PR CI run exposed source-authority and max-lines defects; the exact failing tests now pass locally.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(docker): install libgomp1 in the runtime image for managed llama.cpp (#134532)
* fix(docker): install libgomp1 in the runtime image for managed llama.cpp
The official ghcr.io/openclaw/openclaw image (bookworm-slim) never
installed libgomp1. The verified llama.cpp release binaries downloaded
by managed local llama-server setup are linked against libgomp, so the
--version preflight probe fails with "libgomp.so.1: cannot open shared
object file", making local embeddings unreachable in the official
Docker image.
Fixes#134439.
* fix(docker): clarify managed llama-server runtime requirement
Keep the mandatory OpenMP runtime in the image and identify its preflight contract in the existing comment and regression test.
Co-authored-by: chelsealong <chelsealong@126.com>
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: chelsealong <chelsealong@126.com>
* fix(devices): allow authorized scoped node token management (#135617)
* fix(cli): allow authorized node token rotation and revocation
* test(cli): keep token scope fixture indices typed
* test(ui): scope catalog marquee assertions to session rows
* fix(msteams): thread context is dropped when a channel reply has no replyToId (#129345)
* fix(msteams): thread context is dropped when a channel reply has no replyToId
Teams channel activities frequently arrive without `replyToId`. Thread context
resolution keyed exclusively off that field, so parent-message and sibling-reply
context silently never loaded for those turns and the agent answered without any
of the thread it was replying in.
Fall back to the thread root already parsed from the conversation id. The Graph
message URL builder in `attachments/graph.ts` resolves the same value this way
via `threadRootMessageId`, and `resolveMSTeamsRouteSessionKey` likewise accepts
both, so this aligns thread context with the two adjacent call sites rather than
introducing a new source of truth.
The self-comparison guard keeps a top-level post from being treated as its own
parent, mirroring the `threadRootMessageId !== messageId` check in graph.ts.
Behaviour is unchanged whenever `replyToId` is present.
* fix(msteams): resolve thread context against the conversation thread root
Review follow-up: the fallback chose `replyToId` ahead of the conversation's
embedded thread root, which contradicts the canonical precedence used elsewhere
in the plugin.
`resolveMSTeamsRouteSessionKey` reads `conversationMessageId ?? replyToId`, and
`attachments/graph.ts` resolves the Graph message URL from `threadRootMessageId`
the same way. On a deep reply, where both identifiers are present and differ, the
previous ordering routed the session to the thread root while fetching parent and
sibling context beneath the nested reply — attaching context from one message to
the session of another.
Take the conversation root first and keep `replyToId` as the fallback for
activities whose conversation id carries no thread suffix. The self-comparison
guard is retained so a top-level post is not treated as its own parent, matching
the `threadRootMessageId !== messageId` check in graph.ts.
Adds a handler-level regression covering both identifiers present and different;
it fails under the previous ordering.
* refactor(msteams): tighten the thread-root invariant comment
Condenses the nine-line block to the three-line form used elsewhere in this file,
keeping the two load-bearing facts: why the conversation root wins over replyToId,
and that replyToId remains the fallback. No behavior change.
* test(msteams): prove channel thread retrieval at the Graph boundary
The existing thread-parent tests stub ../graph-thread.js, so they pin the
selection rule but never reach the retrieval boundary the repair actually moves.
This asserts on the HTTP requests the process issues instead: thread-context.ts,
graph-thread.ts and graph.ts all run unmocked, and the only stub is globalThis.fetch.
The stub has to be a vi.fn(). fetchWithSsrFGuard treats a fetch carrying a .mock
property as a test-installed mock (isMockedFetch) and lets it win; a plain function
is bypassed in favour of undici dispatcher-aware fetch, and records nothing.
For a channel activity with no replyToId the process now issues:
GET /v1.0/teams/{aad}/channels/{channel}/messages/{root}
GET /v1.0/teams/{aad}/channels/{channel}/messages/{root}/replies?$top=50
and the retrieved parent reaches the agent. With threadParentId reverted to
activity.replyToId that set is empty, so the first assertion fails on zero requests.
* test(msteams): prove channel thread retrieval through the registered handler
Replaces the direct-invocation variant with one shaped after
approval-control.mock-gateway.test.ts: the activity enters through
registerMSTeamsHandlers and is driven by the registered Bot Framework onMessage
callback, so the inbound path runs as registered rather than a handler factory
being called directly.
The message handler, thread-context.ts, graph-thread.ts and graph.ts all run
unmocked; Graph is observed at the fetch boundary. The stub has to be a vi.fn()
because fetchWithSsrFGuard only lets a fetch carrying a .mock property win over
undici dispatcher-aware fetch (isMockedFetch).
For a channel activity delivered without replyToId the process issues:
GET /v1.0/teams/{aad}/channels/{channel}/messages/{root}
GET /v1.0/teams/{aad}/channels/{channel}/messages/{root}/replies?$top=50
and the retrieved parent reaches the agent as a system event. With threadParentId
reverted to activity.replyToId neither request is issued.
* fix(test): restore ambient fetch and correct the mock-gateway proof signatures
Three defects in the proof added by the previous commit, all caught by CI:
- The fetch stub was never restored, so it leaked into sibling suites sharing the
worker and extensions/msteams/src/qa/bot-framework-server.test.ts began seeing 200
for requests it expects to reject. An afterEach now restores the ambient fetch.
The local full-suite run passed only because vitest happened to order the files
the other way round; running the two files together reproduces it.
- OpenClawConfig was imported from ./runtime-api.js; the module sits one level up.
- The registered onMessage callback takes (context, next, turnAdoptionLifecycle);
both call sites passed only the context.
* fix(test): drop redundant String() conversions in the mock-gateway proof
typescript(no-unnecessary-type-conversion) flagged both call sites: the optional
chain already yields a string once defaulted, so String() changes neither type nor
value. The rule is type-aware, so a bare oxlint run over the single file did not
surface it; only the project shard does.
---------
Co-authored-by: Alexander Malysh <malysh00@gmail.com>
Co-authored-by: Alexander Malysh <a.malysh@venista.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(ui): ignore tool-only model auth rows (#134218)
* fix(ui): ignore tool-only model auth rows
* fix(ui): ignore tool-only model auth rows
* fix(ui): ignore tool-only model auth rows
* fix(ui): ignore tool-only model auth rows
* fix(ui): preserve model accounts when filtering tool credentials
Use canonical monitored-auth evidence for Model Providers cards. Preserve
OAuth and non-expiring token profiles, alias logout targets, configured
models and usage while omitting tool-only API-key rows. Remove duplicate
usage-presence bookkeeping.
Refs #134218. Production LOC -3; tests +56.
Co-authored-by: wantosure <wantosure@126.com>
* test(ui): assert usage failures without provider cards
Replace the stale OpenAI-card readiness wait with the visible usage
failure outcome and assert that an unconfigured provider stays absent.
Preserve the RPC assertion and all configured-provider sibling cases.
Refs #134218. Production unchanged; test LOC net zero.
Co-authored-by: wantosure <wantosure@126.com>
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: wantosure <wantosure@126.com>
* fix(cron): kind change fails with "command env must be an object" when env is omitted (#134639)
* fix(cron): omit undefined payload fields
* fix(cron): preserve valid payload kind conversions
Omit absent optional payload fields at the replacement producer and remove
unreachable same-kind fallback branches. Keep strict validation and cover
actual service persistence, malformed env, and tool restrictions.
Co-authored-by: Solomon Neas <me@solomonneas.dev>
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: Solomon Neas <me@solomonneas.dev>
* fix(backup): exclude workspaces before traversal
Carry disabled workspace roots into the shared backup inventory before tar and SQLite traversal. Preserve nested agent state and keep absolute-link rejection for included content.
Fixes#135218. Supersedes #135243 and #133990.
Thanks @jarvismazz and @yubingjiaocn for the reports. Thanks @ericcaiwx-star and @ruel225 for the original repair work.
Co-authored-by: ericcaiwx-star <ericcaiwx@gmail.com>
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(infra): drop empty PATHEXT entries from Windows extension lookup (#135807)
* fix(infra): drop empty PATHEXT entries from Windows extension lookup
resolveWindowsExecutableExtensions split PATHEXT on ";" without dropping
empty components. A trailing ";", which is common in real PATHEXT values,
leaves "" in the list, so an extensionless file matches even when the
caller passed includeExtensionless: false. windows-command.ts and tui.ts
both pass false.
The sibling resolveWindowsExecutableExtSet in the same file already
filters, so this was just the one path missing it.
* fix(infra): reuse normalized Windows executable extensions
Keep extensionless command lookup explicit by sharing the existing PATHEXT parser. Remove duplicate parsing and cover empty suffixes, explicit filenames, and the command-resolution caller.
Co-authored-by: Teddy Tennant <teddytennant@icloud.com>
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: Teddy Tennant <teddytennant@icloud.com>
* fix(anthropic): context usage is lost when a proxy omits message_delta usage (#135467)
* fix(anthropic): keep context usage when message_delta omits usage
The Anthropic SDK lane guarded the shared usage accumulator against an omitted
message_delta usage object; the managed transport lane called it unconditionally,
where the payload ?? {} fallback overwrote the message_start prompt snapshot with
contextUsage: unavailable. Anthropic-compatible routes whose provider is not
"anthropic" are forced onto that transport, so the population the guard was written
for never had it.
Move the presence check into the accumulator that owns the payload contract and drop
the duplicated guard in the SDK lane, so both lanes follow one rule. A present but
empty usage object still resolves to unavailable, matching the previous SDK check.
* test(anthropic): assert delta context usage through the usage object
expect(usage.contextUsage) narrowed to the object literal returned by the local
emptyUsage helper, which does not declare contextUsage, so the core test-types
stripe rejected it. Read it through toMatchObject like the neighboring cases.
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: Yigtwxx <yigiterdogan023@gmail.com>
* fix(systemd): protect credentials in managed backups (#131786)
Sanitize and privatize managed unit backups, restore them atomically on publication failure, surface unsafe remnants, and remove them across uninstall paths.
Closes#131781. Supersedes #131784.
Co-authored-by: vyctorbrzezowski <51521767+vyctorbrzezowski@users.noreply.github.com>
Worked on by:
- @vyctorbrzezowski
* fix(agents): preserve requester MCP runtime ownership (#134819)
## What Problem This Solves
Requester-scoped MCP resolution can own serialized work for a runtime key while that runtime has no lease. Idle or cap cleanup could dispose the runtime and its transport before the work finished.
## Why This Change Was Made
The lifecycle owner now treats the existing requester work chain as active ownership. Idle cleanup skips that key, and cap cleanup does not queue an eviction behind that work. Required session disposal still drains the same chain.
## User Impact
Requester-scoped MCP transports and sessions now survive resolver work. Normal idle cleanup still removes the runtime after that work drains.
## Evidence
- Current-main loopback reproduction: the production manager returned one idle eviction while requester resolution was held; expected zero.
- Exact candidate loopback proof: the shipped harness runtime retained the same runtime object and MCP session identifier through the protected sweep, then the later idle sweep removed the runtime and closed the session.
- Cap regression: the original PR head disposed the refreshed requester runtime; the repaired head preserved it.
- [Exact-head agent shard](https://github.com/openclaw/openclaw/actions/runs/33675253407/job/100398989009): 89 files and 1,395 tests passed, including the real loopback boundary.
- Formatting and `git diff --check`: passed.
- Exact-head Autoreview P0: clean.
## Proof Boundary
The proof uses the production MCP runtime manager, the pinned MCP SDK 1.30.0 client and server, and a real loopback Streamable HTTP session. It does not claim external MCP service or provider authentication behavior.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(plugins): refresh persisted registry when source mounts change (#136517)
* fix(plugins): retain mounted source ownership through loading and restart
* ci: refresh checks against current main
* fix: gateway restart hangs while followup drain retries a draining error (#136713)
Closes#136684
## What Problem This Solves
Fixes an issue where operators restarting a Gateway with a queued follow-up could wait several minutes for shutdown. The queued turn retried after restart admission rejected it, so the old Gateway stayed active until its shutdown deadline.
## Why This Change Was Made
The follow-up queue now follows the existing one-way restart signal. When restart drain commits, it retires process-local queue work and its old callback before the Gateway inspects active work. A reversible restart-signal fence still waits for settlement and retries after rollback.
Durable accepted-input recovery remains the restart owner. The old in-memory queue and execution authority do not cross restart, as documented in `docs/gateway/restart-recovery.md`.
## User Impact
A Gateway with a queued follow-up can finish a supervised restart on the normal shutdown path instead of emitting a retry burst until forced termination. If restart signal admission rolls back, the queued follow-up resumes normally.
## Evidence
Current-main red at `3100ba535f`:
- Blacksmith child-Gateway run: https://github.com/openclaw/openclaw/actions/runs/33711422936
- One held model turn, one confirmed queued same-session follow-up, then `SIGTERM`.
- Queued case reached the real follow-up admission path and emitted `7` `followup queue drain failed` events.
- No-queue control emitted `0` drain failures.
Exact repaired tree at `2baff4f9ed3f84dc9493149dcc20fed441c9baea`:
- Blacksmith exact-tree run: https://github.com/openclaw/openclaw/actions/runs/33714572460
- The same child-Gateway scenario emitted no retry burst and exited in `3.465s`.
- `src/auto-reply/reply/queue.drain-restart.test.ts`: `29` passed.
- `test/gateway-queued-session-rotation.e2e.test.ts`: `2` passed.
- Materialized source hashes matched the local commit on https://github.com/openclaw/openclaw/actions/runs/33715159753.
- Type-aware Oxlint passed on all touched files.
- Oxfmt and `git diff --check` passed.
The focused tests also cover:
- a typed draining error while restart admission stays open;
- restart-signal rollback and queue resume;
- restart-signal promotion to one-way drain;
- synchronous retirement of queued turn ownership and stale callbacks.
Production LOC: `+33/-2` (net `+31`). The growth binds queue authority to the existing restart signal and closes the user-visible shutdown lifecycle gap. Tests: `+316/-0`.
## AI Assistance
AI-assisted repair. The contributor's original restart-fence distinction and credit are preserved. Maintainer follow-up added the real Gateway reproduction, lifecycle retirement, and exact-head proof.
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(auth): preserve OAuth metadata during session SDK refresh (#127988)
The shipped agent sessions SDK replaced a persisted OAuth credential with the provider refresh result. Providers such as Anthropic return only rotated token fields, so the write removed identity and plan metadata.
Merge the locked stored credential beneath the provider result at the session AuthStorage persistence boundary. Rotated tokens still win, while omitted persisted fields survive.
Tests:
- Published `openclaw/plugin-sdk/agent-sessions` path with real SQLite-backed AuthStorage refresh: red on current main, green on the repaired head.
- `node scripts/run-vitest.mjs src/agents/sessions/auth-storage.test.ts`: 23 passed.
- Focused formatting, lint, and diff checks passed.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(doctor): recognize configured memory provider secret references (#137782)
Recognize structured provider SecretRefs at Doctor's mapped-provider credential-presence boundary without resolving secret bytes or accepting empty environment markers.
Preserve missing-key warnings for absent credentials and preserve the live Gateway's owner-specific unresolved-secret diagnostics.
Verified with the real doctor command on baseline main and the exact candidate: provider-owned and remote-memory references, absent keys, bare markers, and two live unresolved-reference controls. All 102 focused tests pass. Synthetic credentials were used with container networking disconnected; no provider authentication is claimed.
Closes#137742.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(sessions): preserve logical shared-store ownership (#138380)
* fix: full-access tasks lose tools after Gateway restart (#138701)
* fix: preserve full access when continuing interrupted tasks
* test: exercise full-access continuation across a live restart
* fix: retain replay safety beyond the recent transcript window
* test: make live restart checkpoints suspend explicitly
* fix: keep unchanged files out of session diffs (#138877)
* fix(auth): recover billing-disabled profiles in minutes instead of hours (#136364)
Reduce the initial billing-disable window from five hours to ten minutes for stored auth profiles and inline API keys. Inline keys become eligible after expiry; another billing failure starts a new ten-minute window. Stored-profile recovery probes and already-active persisted deadlines remain unchanged.
Correct the recovery-policy documentation and source comment. Verified with 598 focused auth/fallback tests and docs checks on Cloudflare, real SQLite persistence and eligibility proof, independent review, and successful CI on c44f81187bd745c70b23a360fe15496c40251bf9. Live provider recharge was not exercised; the maintainer accepted that documented proof limitation.
Fixes#135835. Thanks @holny.
Co-authored-by: holny <12433219+holny@users.noreply.github.com>
Co-authored-by: Altay <altay@hey.com>
* fix(auth): keep an unset default model unchanged during login (#139218)
Preserve an absent default model during ordinary provider login, just as existing model values are preserved. Remove only the provider-injected default while keeping connection settings, aliases, and credentials.
Keep --set-default as the explicit model-selection action. Extend the existing command regression across supported model shapes.
Related: #136257
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(gateway): avoid crashes when an upgrading client disconnects (#139061)
* fix(gateway): avoid crashes during interrupted upgrades
* test(gateway): bound upgrade regression websocket payloads
* fix(shell-env): preserve CLI terminal ownership during login imports (#139308)
Start the shared login-shell probe in a separate session so interactive Bash cannot take the CLI's controlling terminal. Preserve interactive startup exports and existing filtering, framing, timeout, executable lookup, and cache behavior.
Verified the current-main Doctor SIGTTOU failure and candidate completion using the built CLI in a real job-control terminal. Added a real-Bash startup regression and verified JSON output, imported keys, terminal restoration, and timeout behavior.
Closes#139257.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(auto-reply): fold trailing transcript growth into memory-flush fresh totals (#138928)
* fix(auto-reply): fold trailing transcript growth into memory-flush fresh totals
runMemoryFlushIfNeeded persisted totalTokensFresh=true using only the last
provider usage anchor, ignoring messages appended after it. A later
preflight compaction check trusts totalTokensFresh and skips re-reading the
transcript, so real growth after the anchor could silently escape both the
early memory-flush trigger and the blocking compaction gate.
Project trailing messages the same way runSessionCompactionIfNeeded already
does before persisting the total as fresh.
* refactor(auto-reply): share anchor-plus-trailing prompt token calc
Memory flush duplicated the anchor-plus-trailing-projection calculation
that preflight compaction (estimatePromptTokensFromSessionTranscript)
already used. Extract resolveAnchoredPromptTokens() and have both call
sites use it, so the two calculations can no longer drift.
* fix(auto-reply): share fresh usage projection without duplicate accounting
Include transcript growth after the latest provider usage report before
memory maintenance saves a fresh total for preflight compaction. Reuse the
existing provider-visible estimator and canonical freshness validation,
removing the extra wrapper and redundant freshness check.
Related: #138871
Co-authored-by: chelsealong <chelsealong@126.com>
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: chelsealong <chelsealong@126.com>
* fix(cron): preserve pending paced checks across restart (#140083)
Let the shared missed-slot decision honor the accepted future paced
occurrence before startup admission or backoff repair can replace it.
Keep ordinary unpaced and overdue catch-up behavior unchanged.
* fix(daemon): systemd install no longer aborts on "ownership could not be verified" after a clock step (#137253)
Replace wall-clock deadline arithmetic with `performance.now()` for systemd ownership probes, definition mutation, effective-manager reads, and legacy metadata refresh.
This preserves the existing shared timeout budgets, fallback behavior, and fail-closed ownership checks while preventing NTP, resume, or manual clock steps from collapsing later calls to the 1 ms floor or inflating them beyond the configured budget.
Regression coverage exercises forward and backward wall-clock steps across all four deadline owners. Exact-head CI, Workflow Sanity, independent P0/P1 review, and the durable ClawSweeper review passed.
Co-authored-by: coderdailyone <codecookie@proton.me>
* fix(config): apply environment changes after in-process restart (#141296)
## What Problem This Solves
Fixes an issue where users could update or remove a configured environment value after an in-process Gateway restart, but the process kept using its old value. A full process restart applied the saved configuration correctly.
## Why This Change Was Made
Snapshot cleanup discarded environment ownership without removing the values it owned. The next startup could then mistake those surviving values for inherited environment settings. Keep that ownership through snapshot cleanup, fence old rollback transactions, and retain explicit full-reset behavior.
## User Impact
Accepted configuration changes can replace or remove their environment values after restart. Inherited environment values keep their precedence, and config-selected paths remain stable.
## Evidence
- A real Gateway using main dcc68718f5 accepts a replacement after a successful same-process restart, but public config.get still resolves the old value despite matching new applied/config revisions.
- On c7b393e10286b8b43c316cc7229f3bca94786493, the real Gateway completes restart, replacement, and removal. Public reads return the new value after replacement and the normal unresolved placeholder after removal. The inherited control stays unchanged throughout. Both config writes and final shutdown exit 0.
- Both runs use the real source CLI entry with synthetic config and normal token authentication. The candidate uses the Gateway's existing signed-device WebSocket client for public reads and the normal config set/unset CLI for writes. Complete private command, response, lifecycle, source, and resource receipts are retained. The interrupted baseline removal read is not claimed as proof.
- All 58 focused tests pass across runtime-snapshot.env.test.ts, config.env-vars.test.ts, and runtime-snapshot.test.ts. Controls cover late rollback, ambient precedence, aliases, casing, blocked variables, and explicit reset. Pinned formatting and focused lint pass.
- [Required CI passed](https://github.com/openclaw/openclaw/actions/runs/34136580481) on the exact candidate. No provider call or local build was needed for this proof.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix: pasted text is missing in model-locked Codex sessions (#142021)
* test(models): consolidate runtime policy matrix
* test(models): extract registry fixtures to satisfy lint
* fix: deliver pasted text to model-locked harness sessions
* fix(docker): restore source builds after dependency ownership guard (#142073)
* fix(docker): restore source builds after dependency ownership guard
* fix(docker): preserve existing lifecycle copy prefix
* fix(security): preserve existing WhatsApp group allowlists (#142589)
* fix(doctor): avoid installing unselected search providers (#142640)
Respect a normalized explicit web-search provider during plugin repair.
Only derive search plugin candidates from environment credentials when
no provider is selected, retaining independent plugin configuration.
Cover selection normalization, environment-only discovery, and existing
plugin disable and deny controls at the repair boundary.
* fix(doctor): migrate legacy Codex music model selectors (#143235)
Doctor left a legacy music selection such as `openai-codex/gpt-5.4` unchanged after `doctor --fix`. Its media selector scan and repair covered image and video but omitted music.
Include music in both existing loops. Doctor now reports the affected setting and saves eligible references under `openai`, using the same primary/fallback handling and provider blockers as the other media selectors.
Validation:
- Required CI passed on `51f0f308`: [run 34372705782](https://github.com/openclaw/openclaw/actions/runs/34372705782).
- Actual CLI proof reproduced the saved-state failure on baseline build `006aaacc` and verified the fix on candidate merge build `0b17e570`, which contains this PR head.
- Scalar and structured selectors passed repair, public config readback, and repeated repair. Suffixes, fallback order, explicit metadata, and canonical/custom references stayed intact.
- Non-interactive Doctor preview named the music setting without changing saved configuration. A blocked legacy provider retained its selection and authored settings while an eligible fallback migrated.
- All six regression cases passed in [the required CI job](https://github.com/openclaw/openclaw/actions/runs/34372705782/job/102537913364). They cover both legacy aliases, selector shapes, provider/model blockers, and agreement with image/video. The native pre-commit hook passed for all three files.
This completes the music-specific omission left after #142284. Proof uses synthetic configuration and covers stored references; it makes no live generation or authentication claim.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(doctor): preserve selected model metadata in diagnostics (#144332)
Keep active tool schema diagnostics on the primary model already selected
from manifest aliases. A second alias pass could use another model's
context window while labeling the tools with the selected model ID.
The regression reproduces middle selecting final's 4096-token context
instead of its own 32000-token context.
Related: #130706
* fix(discord): report parent channel denies for thread permissions (#121144)
* fix(discord): use parent permissions for thread diagnostics
Punchcard-Session: brisk-lantern-meadow-g5
* test(discord): preserve voice send diagnostics
* fix(discord): keep send diagnostics operation-specific
* fix(discord): share thread channel classification
* fix(config): avoid false model changes when saving settings (#144807)
* fix(doctor): preserve legacy registry files when quarantine fails (#144787)
Closes#144773
## What Problem This Solves
Doctor could delete malformed legacy sandbox registry data after its quarantine rename failed, then report a quarantine location that had not received the bytes. Both container and browser registry repair used the same failure path.
## Why This Change Was Made
The existing quarantine owner now reports success only after rename succeeds. It preserves the source and propagates other rename errors; a source that disappears maps to the existing missing result. Doctor retains its normal warning handling for this optional repair.
The regression cases now share the existing migration fixture and real SQLite writers. The separate mocked-writer fixture is removed, and the successful-quarantine test verifies the actual destination bytes.
## User Impact
Failed quarantine leaves the malformed bytes available for manual repair and reports the failure. Valid migration, existing SQLite rows, sharded precedence, read-only inspection, empty/missing inputs, and maintenance ownership retain their contracts. This changes no schema, configuration, retention policy, dependency, or overall Doctor exit policy.
## Evidence
Pinned main `1d4afe3966` was exercised through the built public `openclaw doctor --fix --non-interactive` command in isolated offline state. For each registry kind, a real nonempty destination directory caused `rename` to return `EISDIR`; the source was then unlinked and Doctor reported successful quarantine. Original synthetic input and destination bytes were preserved independently.
The corrected source was independently exercised through the same public repair for both registry kinds. The native filesystem trace still records `EISDIR`, but the source and existing destination remain byte-for-byte intact, and Doctor reports:
```text
doctor:sandbox run failed: EISDIR: illegal operation on a directory,
```
| Before | After |
| --- | --- |
| Failed rename was followed by source deletion and a successful-quarantine claim. | Failed rename preserves both inputs and reports failure without claiming quarantine. |
| The claimed destination had not received the malformed bytes. | After resolving the owned collision, repair moves the identical bytes to the reported file; a repeat leaves the result unchanged. |
Independent acceptance also verified real-clock quarantine, registered read-only lint, empty/missing inputs, and real SQLite migration with existing-row conflicts and mixed valid/invalid shards. Whole canonical rows, including JSON and update timestamps, remained unchanged; canonical rows take precedence over shards, and shards over monolithic input. No database writer was mocked in this public proof.
An initial fixture reused an old fixed timestamp and was correctly refused by maintenance before reaching quarantine. Its failed and prematurely batched dependent steps are retained. The corrected fixture uses a contemporaneous timestamp and stops dependent actions on nonzero repair exit. Product leases, workers, permission checks and deadlines were unchanged. The original baseline proof remains valid and was not replayed.
The consolidated regressions produced the intended four failures on baseline. All 34 focused tests across three files pass with the repair. Native formatting, syntax lint, line-size and assertion-safety checks, and the complete runtime/UI artifact build passed. Full type checks and broad suites remain hosted. The tested source tree matches published head `84dffef7a099a9918fd08ae63993dfef4df788d9`; signing changed no runtime input. The retained build keeps its original pre-signing stamp.
The source-disappearance test is a deterministic filesystem-spy primitive, not a public concurrent-process claim. No host-power-loss or elapsed lease-expiry behavior is claimed. Complete private proof and failure history are retained; publication omits account and infrastructure details.
Thanks @Alix-007 for the original repair and regression coverage.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix: retain route ownership while discovering conversations (#146927)
* fix(install): preserve runtime links when Node validation fails (#147221)
* fix(install): preserve runtime links when Node validation fails
* fix(install): preserve relative runtime targets and empty aliases
* test(install): preserve runtime fixture tuple types
* fix: doctor reports missing OAuth dir for channel config with no installed plugin (#147888)
Doctor no longer reports a missing OAuth directory or offers to create credentials storage solely because leftover channel configuration requests pairing for an unavailable plugin. Reuse the existing effective installed-plugin set in the state-integrity check, while preserving CRITICAL diagnostics and repair prompts for installed pairing plugins and explicit OAuth-directory overrides.
The missing-plugin regression fails with the production change reverted and passes after restoration. The complete state-integrity test file passes all 30 tests, including an explicit installed-Telegram CRITICAL assertion; exact-head CI is green.
Fixes#147772.
Thanks @zhangguiping-xydt for the implementation.
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(doctor): stopped private inputs replay after session repair (#148625)
* fix(doctor): retain private input receipts during canonical repair
* test(doctor): use lint-safe receipt rekey expectation
* fix(channels): settle task-scoped context leases once (#148922)
* fix(push): normalize malformed proxy auth errors (#149112)
* fix(cron): reject invalid stagger before silent schedule changes (#151740)
Closes#146446
## What Problem This Solves
Fixes raw Gateway cron requests reporting success after silently discarding an invalid explicit `staggerMs`, substituting a default or retaining the previous value while other requested changes commit.
## User Impact
Invalid explicit stagger values now return `INVALID_REQUEST` before creating or updating a job. This includes explicit `null`: the existing schedule schema permits an optional integer, not a nullable value. Supported integer strings, zero, numeric flooring/clamping, omitted values, and ordinary CLI duration parsing retain their behavior. No operator migration is required; Doctor's existing legacy schedule repair is unchanged.
## Why This Change Was Made
The shared schedule normalizer rejects an invalid authored value before create defaults or update retention erase it. The fix uses the existing parser and request-error boundary rather than broadening the parser or adding a second validation policy. Non-cron schedules still discard stale stagger fields.
## Evidence
The authenticated comparison used Linux arm64 and Node 24.19.0, with registered `openclaw gateway call` commands, a task-only token, actual Gateway handlers, and SQLite readback. Baseline: `bef4c20d45`. Candidate: that same main revision plus this PR's two-file delta from `f9be3cb73438f6f68c4ead73103d413564398210`—not an execution of the full original PR checkout.
| Request | Baseline | Candidate |
| --- | --- | --- |
| Hourly add with `"1e3"` | Success; stored `300000` | `INVALID_REQUEST`; no mutation |
| Non-hourly add with `"1e3"` | Success; stagger omitted | `INVALID_REQUEST`; no mutation |
| Rename/update with `"42.8"`, same expression | Success; rename committed and previous stagger retained | `INVALID_REQUEST`; no mutation |
| Rename/update with `"42.8"`, changed expression | Success; rename/expression committed with default stagger | `INVALID_REQUEST`; no mutation |
For each rejected candidate request, all cron-table rows and authenticated views remained unchanged. All 20 expected-observation checks passed on each side, including integer strings, numeric/string zero, omitted add/update behavior, numeric floor/clamp, non-cron stale fields, explicit `null`, and a wrong-token control. Physical candidate restart preserved authenticated `cron.get` responses and complete serialized cron-table values.
Five focused files passed **167 tests**, covering normalization, stagger behavior, persisted schedules, and Doctor legacy schedule repair. With the original normalizer and the same regression tests, **8 failed and 95 passed**, demonstrating the invalid-input regression. Doctor coverage is focused-test evidence, not a published-updater upgrade test.
The initial fixture attempts are not product failures: the first stopped before mutation checks because native Workshop monitors were enabled in the fixture. Baseline restart initially compared `cron.list` projections with `cron.get` and then compared SQLite null-prototype records with JSON objects. Retained same-API responses and exact serialized row values confirmed persistence without replay; the failed fixture results remain retained.
During the completed comparison, scheduling and all jobs stayed disabled; native Workshop monitor rows were retained disabled. No scheduling timers, jobs, provider requests, channel delivery, or execution-delay measurement were exercised.
Existing hosted CI is not all green: [run 35338846156](https://github.com/openclaw/openclaw/actions/runs/35338846156) reports websocket transcript-event timeouts in `session-message-events.test.ts` in `checks-node-compact-large-4`, with the aggregate CI gate failing. No CI replay or unrelated repair is included in this PR.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(cli): keep model validation out of session execution (#154763)
* fix(agents): treat empty credential environment variables as unresolved (#155558)
Closes#153484
## What Problem This Solves
The session credential resolver returned the environment variable name when the variable existed but was empty.
## Why This Change Was Made
Distinguish a present empty variable from an absent variable in cached and uncached resolution. Empty values remain unresolved; non-empty variables override literals; absent names retain literal fallback.
## User Impact
SDK consumers of `openclaw/plugin-sdk/agent-sessions` no longer receive the variable name as credential/header material for an explicitly empty variable. Required headers use the existing resolution error. This is verified SDK hardening, not a claim that normal CLI provider headers use this resolver.
## Evidence
- Exercised the built public `ModelRegistry` export with an authored models.json provider header. The export is present in release v2026.9.5. Before the fix, empty returned the variable name; after the fix, it returned the existing failed-to-resolve-header error.
- Non-empty resolved the synthetic value; unset preserved the literal name. Repeated after integrating current main.
- Focused resolver suite: six tests passed; command wall 14.43s including worker preparation, test cases 15ms total.
- Simplification kept the two existing resolver owners and corrected the absent-variable test to exercise the deliberately unset name. No new abstraction/config/schema.
- CLI models probes and local agent requests were also inspected but use a different header path; they are not evidence for this fix.
## Compatibility
Shell command resolution and unset-name literal fallback are unchanged.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(doctor): apply ambient-owner fallback to dream diary and all memory target handlers (#156437)
Apply the existing ambient-owner fallback in the shared Doctor memory target resolver, matching doctor.memory.status. Older clients that omit agentId can read the dream diary and run the five maintenance handlers when a configured owner is unambiguous. Explicit selection still wins, and ownerless multi-agent fleets still receive the selection error.
Thanks @Orionation for the fix and @Cyb3rb1ade for reporting the legacy-client failures. The configured-owner case is addressed; selecting an agent remains necessary for ownerless fleets.
Validation: scoped review found no accepted/actionable P0 or P1 findings, the author supplied a resolver comparison, and exact-head hosted CI passed. No fresh after-fix Gateway trace was captured for this landing. Production +6/-1; tests unchanged.
Closes#138369
Co-authored-by: Jonathan Sieling <jonathan@tailoredmonkey.com>
* fix(sandbox): unblock scoped recreation across runtime targets (#157808)
* fix(sandbox): unblock scoped recreation across runtime targets
* test(sandbox): validate mocked process arguments
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
* fix(security): parse gpt generation numbers in the audit tier check (#140012)
Fix security audit warnings that incorrectly classify numbered GPT-6 models as below GPT-5. Compare numeric generations in the shared model-hygiene classifier while preserving older-model warnings and Azure's legacy gpt-35-turbo naming contract.
The repair covers primary models, fallbacks, image models, and agent overrides without changing provider routing, authorization, or unknown-name advisory policy. Broader unknown-name behavior remains separate in #160580.
Fixes#139751
Thanks @holny for the fix, regression coverage, and isolated CLI evidence.
Co-authored-by: holny <holny@foxmail.com>
Co-authored-by: IWhatsskill <284122573+IWhatsskill@users.noreply.github.com>
* fix(doctor): report shared auth health without a default agent (#134902)
* fix(doctor): report shared auth health without a default agent
* fix(ci): stabilize tooling timing and fixture cleanup
Use fake time for the existing readiness-body polling test, preserve explicit
success and delegated return statuses in Bash 5.2 EXIT cleanup, and remove a
repeated seven-compiler declaration invalidation build already covered by
the shared SDK writer.
Validation: 146 readiness/fixture tests and two real declaration-build cases
pass, along with 15 shell controls across Bash 3.2, 5.2, and 5.3. The scoped
changed gate and Codex review pass.
* fix(doctor): converge redundant shared auth relocation subsets (#135096)
* fix(doctor): converge redundant shared auth subsets
Preserve richer canonical credential stores and original source receipts through cleanup. Keep differing credentials and state rows fail-closed with actionable diagnostics. Reported by @jodok and independently reproduced by @felixboenkost-droid in #132605.
* fix(doctor): satisfy shared auth relocation validation
* fix(doctor): report retained device-auth files (#133978)
Surface retained client device-auth JSON in ordinary Doctor output and opted-in
lint checks, including when identity, pairing state, or Gateway health is absent.
Resolve only retired-file diagnostics and profile-aware repair hints from the
original health context environment; keep pairing and token reads on the active
state view.
Describe remaining migration or cleanup debt without inferring import history.
Keep the legacy-file token guard and Doctor-owned migration unchanged, share
friendly and structured findings, and remove redundant string-formatting paths.
Refs: #127973, #133978.
Reported by @yetval.
Co-authored-by: yetval <yetvald@gmail.com>
* fix(update): stop detached updates after chat ownership is revoked (#137312)
* fix(update): recheck chat ownership before detached update launch
* test(update): keep helper owner cases in lifecycle support
* docs(update): describe the implemented session-less notice route
* test(update): match conditional Vitest API in owner cases
* fix(doctor): inspect source auth stores during full lint (#137610)
* fix(doctor): inspect source auth stores during full lint
* test(ci): freeze release candidate fixture clock
* refactor(doctor): simplify scoped environment restoration
* test(release): preserve elapsed time in fixture clocks
* test(auth): join migration fixture callbacks before teardown
* fix(gateway): recover queued RPC replies across restart (#138565)
Bind RPC and scheduled lifetime resolvers to their exact Gateway host. Reuse the canonical host for HTTP hooks and worker dispatch so shutdown retains pending reply recovery without weakening execution fences.
Fixes#138564
* fix(doctor): preserve legacy ownership during config repair (#138837)
Carry the retained default-agent owner through legacy migration, persist the
canonical roster through the config writer, and verify startup repair with
that same ownership preparation. This prevents a migratable legacy fleet
from being replaced by stale last-known-good configuration during update.
Verify the actual persisted roster and retained settings, keep ambiguous
owners rejected, and document migration before recovery.
Validation: 77 owner tests, changed checks, and independent P0-P2 review.
Release-note context: preserves update settings and legacy default-agent
ownership through automatic Doctor config migration.
* fix: prevent browser profile permission hangs in Doctor (#139612)
* fix: prevent browser profile permission hangs in Doctor
* docs: explain explicit browser registration repair
* fix(gateway): prevent systemd restarts from hanging (#140914)
With OPENCLAW_NO_RESPAWN enabled, a managed restart could reopen the Gateway inside the original process while systemd waited for it to exit. Use the existing external-restart action to drain, close, retain the recorded reason, and exit so the service manager creates the successor.
Verified through real Linux user-systemd restart, all five drain modes, ordinary stop/start, unmanaged SIGUSR1, independent runtime acceptance, and 384 focused lifecycle controls. Preserve existing managed-update ownership and service settings.
Thanks to @rboy1 for the report. Closes#140821.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(doctor): recover legacy node tokens with invalid scopes (#143599)
Doctor detects legacy node tokens that inherited operator scopes, but its suggested rotation preserves those scopes and is denied. Add `devices rotate --no-scopes` and use it in Doctor's node recovery guidance. Rotation removes only the matching stale local node credential in the same transaction, while preserving operator pairing, approved scopes, and existing authorization and token-delivery rules.
Reproduced through public pairing on v2026.3.11, then upgraded the same isolated state to baseline `ac5ba7d738`. The corrected command succeeds, removes the matching cache, and preserves operator access through restart and node reconnect. Omitted scopes retain their existing meaning; `--scope` and `--no-scopes` conflict.
Validation:
- 203 focused CLI, Doctor, pairing, credential-store, and Gateway authorization tests.
- Formatter, syntax lint, size/assertion guards, and a full scenario build.
- Independent public CLI and signed-protocol acceptance: recovery, rejected scope/role changes, distinct-device bearer withholding, self-token delivery, revocation, interrupted response, node restart, hosting-state reset, and operator-rotation continuity.
- Fresh security-aware source review. Hosted CI and the exact-head landing gate run on this PR.
The real node fixture used Linux. Native macOS executable discovery was not run; existing real-store tests cover the shared generation reset. No worker session was launched.
Fixes#143263. Thanks @matthewmoroz for the report.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(skills): refresh session snapshots on gateway restart (#122110)
Closes#122107
With skill watching disabled, a persisted session can keep its old skill catalog after `openclaw gateway restart`. The unmanaged Gateway restarts in the same process, so the skill snapshot generation does not advance. The next turn still receives the old skill description even though the file changed.
Advance the existing skill generation at Gateway startup, before serving turns. This refreshes both snapshot caches on the same session's next turn. Keep the current startup and shutdown owners, consolidate the contributor's two-start regression into the maintained boot suite, and document same-session refresh. Managed library selections keep their pinned revisions.
- Built public Gateway and agent CLI: main retained the old description after a completed same-process restart; the candidate delivered the revised description to the same session. Complete provider requests and public JSON confirm the catalog content. The provider was a deterministic local fixture, not a live model service.
- Default watcher, unchanged skill, configuration reload, session/history continuity, and deferred/coalesced safe restarts were exercised through public commands. Bare unauthorized restart signals remained rejected on main and candidate. The operator CLI's existing intent-based behavior with the flag disabled was compared separately and was unchanged.
- The candidate's coalesced restart used the existing forced HTTP connection-close fallback after its one-second grace period; all turns and history survived. The matched baseline closed cleanly. This proof does not establish equal close timing or warning frequency.
- The startup regression fails on main for the missing generation advance and passes with the fix. 252 focused tests cover startup, snapshots, watchers, remote-node refresh, config invalidation, workshop lifecycle, cache pruning helpers, and restart ownership/fallbacks. Targeted lint, formatting, assertion/line-count checks, and documentation links pass.
- The normal watcher already refreshes changed skills on main. This repair covers startup invalidation independently of watcher events. Older closed reports #22517, #22525, and #22568 remain overlap history; their broader symptoms are not claimed resolved by this proof.
Preserves @sappkevin's four commits and original repair intent.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(doctor): session SQLite import drops Codex assistant messages from legacy transcripts (#140296)
Fixes#140100.
Preserve legacy assistant events during Doctor SQLite import by deferring
provider-metadata rewrites until the staging cursor is exhausted. Keep
byte-exact replay checks against normalized prior rows and reuse the existing
per-repair scratch set; no persistent schema or archive policy changes.
The importer and Doctor regressions cover nonterminal assistant rows. Existing
exact-head CI and the contributor's Linux CLI proof show all six events
preserved in order instead of four. Fresh maintainer review found no actionable
issue. Recovery into an already-partial destination still does not restore its
original physical order or active-history projection; that limitation remains
explicit in the PR evidence.
Co-authored-by: Yalçın Doksanbir <yalcindoksanbir@gmail.com>
* fix(cron): forward claude-cli auth profile on scheduled runs to prevent OAuth expiration (#144266)
* fix(cron): forward claude-cli auth profile on scheduled runs
* fix(agents): resolve managed CLI auth for queued runs
Use the existing CLI credential owner for queued candidates, preserving explicit account boundaries and native login. Isolate persisted cron auth fixtures.
Completes the managed CLI credential handoff for cron and queued runs described in openclaw/openclaw#144266 and openclaw/openclaw#144047.
---------
Co-authored-by: Jason (Json) <263060202+fuller-stack-dev@users.noreply.github.com>
* fix(discord): preserve sender-owned line limits and reply scope (#137969)
Co-authored-by: OpenClaw Assistant <assistant@openclaw.local>
* fix(install): never replace or break an existing nvm during installation (#145317)
Preserve compatible active Node runtimes and existing nvm installations in install.sh, the website installer synced from this repository. The installer no longer replaces system Node, rewrites shell profiles, or writes an npm prefix that breaks an existing nvm. With nvm present, reuse a compatible installed Node or ask before installing one; unattended runs that need preparation stop with the exact commands.
The 220-test installer suite and real Debian before/after proof cover preserved system Node, npm configuration, profiles, nvm defaults, and working nvm use. Exact-head CI is green. Website Installer Sync propagation remains a post-merge release-ops follow-up.
Fixes#145292. Thanks to @jb-lopez for the report and diagnostics.
* fix(sessions): recover omitted historical transcripts in Doctor (#120781)
* fix(sessions): stop doctor archiving live transcripts as orphans (#73471)
`openclaw doctor` treats a transcript as an orphan unless its path is the
CURRENT `sessionId` of some entry in `sessions.json`. That is not what a
session store holds, so the scan reports live history as unreferenced and
offers to archive it. Renamed files are then hard-deleted 30 days later by
`cleanupArchivedSessionTranscripts`.
Two independent causes:
1. `resolveSessionFilePath` can only rebuild `<sessionId>.jsonl` unless the
entry carries `sessionFile`, and `sessionFile` is persisted for the live
session only. Transcripts generated as `<iso-stamp>_<sessionId>.jsonl`
therefore never resolve for any *previous* session in a compaction chain.
Falls back to a name-based lookup via the existing
`extractGeneratedTranscriptSessionId` helper, only when the canonical path
is absent, so a brand-new session still gets the plain name.
2. The orphan scan reads `entry.sessionId` alone, ignoring
`usageFamilySessionIds` (every prior id in a compaction chain),
`compactionCheckpoints[].sessionId` and `systemPromptReport.sessionId`.
Referenced ids are now collected from all four and matched against the
session id parsed out of each file name.
Measured on an install with 84 primary transcripts: 58 reported orphans
before, 0 after — every one was a false positive.
Both tests fail on the unpatched build.
* fix(sessions): recover omitted historical transcripts in Doctor
* style(sessions): satisfy historical recovery lint guards
* fix(sessions): retain archived shared aliases without reclassification
* test(ci): include historical discovery in Doctor ownership plan
* fix(doctor): admit transcript-only migration targets
---------
Co-authored-by: Jason (Json) <263060202+fuller-stack-dev@users.noreply.github.com>
* fix(update): recognize custom npm prefix installations (#146091)
Recognize npm ownership for custom prefix installations such as nvm with ~/.npm-global when global-root probes disagree and the prefix has no npm executable. Use npm's configured prefix and the installed OpenClaw launcher, share the resolver with update status, and include inspected paths in unknown-owner refusals. Windows npm shims resolve their package entrypoint instead of the earlier node.exe check.
Keep update target selection on the running package prefix. Older installed updaters still need the explicit NPM_CONFIG_PREFIX workaround for the first hop. Full published-driver-to-candidate upgrade preservation remains unverified; focused ownership, target-selection, status, and CLI admission tests pass, including the failing-before/passing-after complete Windows shim regression.
Fixes#146038
Reported by Discord user Rex Horizon in the support thread "As usual update fails" (#146038).
* fix(context): preserve provider-scoped prompt budgets before reply maintenance (#141052)
* fix(context): preserve provider-scoped prompt budgets before reply maintenance
* fix: preserve literal identities in reply catalog selection
Keep configured and prepared rows distinct when their display keys collide, and prefer an exact model ID before alias matching. Covers both selected IDs and row orders.
* test: align prepared budget coverage with automatic model policy
* test: avoid shadowing expected context window
---------
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
* fix(sessions): prevent Doctor and startup memory exhaustion (#149704)
* fix(sessions): bound legacy migration transcript reads
Prevent openclaw doctor --fix, openclaw update, and Gateway startup
from exhausting memory on large single-agent stores. Read only
legacy-targeted claims and share a streamed digest across inspection,
copy staging, and transactional source cleanup.
* fix(sessions): isolate the transcript digest contract
Keep lifecycle types independent of the SQLite reader and session scope
to remove the type-import cycle caught by architecture CI. Preserve the
streamed digest, copy validation, and transactional deletion behavior.
* fix(daemon): preserve inline auth through service upgrades (#149686)
Keep active env SecretRef values already present in legacy service environments and migrate them through the existing owner-only environment-file writers. Refs #111578.
* fix(update): verify paired trusted-proxy gateways with service credentials (#153540)
* fix(update): verify paired trusted-proxy gateways with service credentials
* refactor(gateway): keep call identity resolution with device auth
* test(update): isolate restart polling from native credential storage
* fix: background continuations lose session access and GitHub tools (#141865)
* fix: background continuations lose session access and GitHub tools
Worked on by:
- @Takhoffman
Co-authored-by: Takhoffman <781889+Takhoffman@users.noreply.github.com>
OpenClaw-Publication: 867311a3-8ac4-4412-9b5d-e45028eed55e
* fix: preserve Windows launcher line endings during publication
Remove the unrelated line-ending normalization introduced by the publication snapshot. The launcher is byte-identical to the PR base.
Worked on by:
- @Takhoffman
Co-authored-by: Takhoffman <781889+Takhoffman@users.noreply.github.com>
* refactor: isolate background GitHub availability preparation
Keep the shared attempt dispatcher within its existing line limit by extracting
the bounded availability stage. Preserve adopted session identity, explicit
capability values, lazy loading, and the canonical live-attempt guard.
Worked on by:
- @Takhoffman
Co-authored-by: Takhoffman <781889+Takhoffman@users.noreply.github.com>
---------
Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com>
Co-authored-by: Takhoffman <781889+Takhoffman@users.noreply.github.com>
Co-authored-by: brokemac79 <255583030+brokemac79@users.noreply.github.com>
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
* fix(doctor): decode session headers across read chunk boundaries (#154034)
Doctor could skip stale transcript-path repair when a valid session header extended past its fixed 8192-byte read window. Read the complete first line with bounded incremental UTF-8 decoding so long headers and split multibyte characters preserve the existing session identity and collision safeguards.
Release-note context: restore Doctor's repair of stale session transcript paths for headers that cross a read chunk boundary. No configuration change or operator action is required. Thanks to @xydt-juyaohui for the fix.
Evidence: the existing regression fails before the fix and passes afterward (8 tests passed), and the reported Windows migration-owner flow repairs both short and boundary-split headers. The approved head has green hosted CI and a scoped-clean Codex review. Published-updater-to-candidate proof and per-file timing measurements remain unrun in this landing lane.
Follow-up: header parse failures still skip silently; record an explicit Doctor non-outcome in follow-up work.
Co-authored-by: 琚耀辉0668001366 <ju.yaohui@xydigit.com>
* fix: faulty compaction probe wedges run steering and restart aborts (#154868)
Closes#154700
Fixes: a failing active-run compaction probe escapes message steering and interrupts restart cancellation of other compacting runs. Reply-owned runs also bypassed the guarded registry path.
User impact: an unreadable compaction state refuses steering safely, chat follow-ups remain queued until the active run finishes, and compacting-mode cancellation continues to other eligible runs. Healthy steering is unchanged. No configuration, permissions, storage, or public API changes.
One contained probe reader now serves embedded queue admission, reply-owned injection, and both restart-abort owners. An indeterminate state rejects steering and skips that handle during compacting-only cancellation. A weak set records the first probe failure per handle so repeated admission checks do not flood logs or retain finished runs.
The existing abortability and supersede readers remain beside the compaction reader in the contributor's extracted module. Reply-owned handles use the same containment instead of another exception boundary.
Isolated Gateway with a plugin-owned active run and a scripted local HTTP provider; the plugin's `isCompacting()` throws. No product source was altered to inject the fault.
- **Before, main `9665d17902cf`:** `sessions_send(mode=steer)` surfaced the raw `compaction probe unavailable` exception. A `chat.send(queueMode=steer)` follow-up was incorrectly delivered to the faulty active backend; the provider recorded the injection.
- **After:** the same tool call returns the existing structured `queue_message_failed reason=compacting ... gatewayHealth=live` rejection. Chat follow-ups are not injected into the faulty backend; the provider sees the next request only after the first finishes. Both turns complete successfully. Repeated checks across the owners produce one warning for that handle.
- **Control, before and after:** a healthy probe accepts the follow-up with `targetDisposition=steered`, and the provider records that exact message.
- **Lifecycle:** the isolated Gateway restarts and answers `health` successfully after the runs settle. The contributor's live queued-compaction regression also proves that a faulty preceding handle does not prevent cancellation, settlement, and registry eviction of the eligible compaction. The reply-registry regression covers the corresponding reply-owned abort sweep.
Focused checks after merging main:
```text
runs.test.ts + runs.steering.test.ts: 69 passed
reply-run-registry.test.ts: 105 passed
Combined wrapper wall time: 30.34 s
compact.hooks.test.ts, faulty-probe case: passed (660 ms test body)
reply-registry throwing-probe and compacting-abort cases: 3 passed
Combined focused wrapper wall time: 30.09 s
```
The CLI lifecycle import-boundary check exposed an eager diagnostic-facade import from the new shared reader. The reader now imports the same logger directly from its lightweight owner. The existing lifecycle check passes all 3 cases, including signal handling after chunk rotation; rerunning it with both steering/reply owner suites passed 153 tests in 29.23 s without changing test mocks.
The live proof uses synthetic local model responses, not an external model provider. It distinguishes a contained rejection from a raw exception; it does not claim that the baseline Gateway process itself crashed.
Co-authored-by: sxh <sunxianhong@ncti-gba.cn>
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(ui): publish rejected goal operations (#156363)
* fix(commands): deduplicate session lifecycle keys (#158487)
* fix(plugins): preserve runtime artifact preference (#158487)
* fix(agents): strip private skill authoring from public ingress (#159220)
* fix(plugins): preserve include ownership during uninstall (#159945)
* fix(channels): restore system-agent approval reactions (#159132)
* fix(moonshot): reject inherited thinking-model keys (#158437)
* fix(tencent): ignore inherited effort-map keys (#158437)
* fix(workers): observe already-aborted operation promises (#158228)
* fix(gateway): reject non-object Claude history rows (#158228)
* fix(copilot): recover pool after synchronous factory failures (#157819)
* fix(tlon): preserve IPv6 ship origins (#157825)
* fix(sms): clean revoked inbound media (#157825)
* fix(slack): preserve replacement webhook registrations (#158085)
* fix(matrix): accept system-agent approval reactions (#158164)
* fix(channels): retain typing tick exclusivity across restart (#158830)
* fix(acp): format unset optional tool arguments (#159417)
* fix(meeting): clean up partial audio startup (#157982)
* fix(exec): retain prepared plugin environment (#158378)
* fix(mxc): clean temp directory after bridge failure (#158532)
* fix(discord): preserve batched native reply identities (#158886)
* fix(skills): count omitted dangerous exec aliases once (#158906)
* fix(sandbox): preserve remote mutation path bytes (#159346)
* fix(browser): accept standard CDP header containers (#159430)
* fix(plugins): preserve sparse catalog preferences (#159945)
* fix(doctor): preserve complete plugin inventory (#153901)
* fix(fs): preserve bounded filesystem correctness (#157524)
* chore: prepare extended-stable 2026.8.34
* docs: add 2026.8.34 release notes
* fix(fs): use typed errno checks for JSON symlinks
* chore: shrink assertion safety baseline
* fix: integrate cumulative 2026.8.34 backports
* fix: integrate release candidate validation repairs
* fix: close release review findings
* fix: repair cumulative backport integration
* fix: preserve system-agent channel custody
* fix: complete stable backport integration
* fix: deliver system-agent approvals to native channels
* fix: align native approval architecture checks
* fix: fence delegated approvals at final effect
* fix: close delegated approval final effects
* fix: fence delegated config publication
* fix: fence delegated config backup effects
* test: prove delegated config final effects
* test: bound delegated approval proof wait
* fix: fence every delegated config effect
* fix: close delegated config I/O races
* fix: preserve delegated backup boundaries
* fix: keep delegated include writes root-bound
* test: await delegated approval publication
* fix: reject unfenced delegated include writes
* test: register delegated include todo suppression
* fix: reject delegated includes before preflight
---------
Co-authored-by: ToToKr <friendnt@g.skku.edu>
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: IvanShang <77005282+qdivan@users.noreply.github.com>
Co-authored-by: Jason (Json) <263060202+fuller-stack-dev@users.noreply.github.com>
Co-authored-by: Yuval Dinodia <102706514+yetval@users.noreply.github.com>
Co-authored-by: Yzx <53250620+849261680@users.noreply.github.com>
Co-authored-by: 0xParzival <145645180+0x-Parzival@users.noreply.github.com>
Co-authored-by: 0x-Parzival <0x-Parzival@users.noreply.github.com>
Co-authored-by: chelsealong <chelsealong@126.com>
Co-authored-by: Alexander Malysh <amalysh@kannel.org>
Co-authored-by: Alexander Malysh <malysh00@gmail.com>
Co-authored-by: Alexander Malysh <a.malysh@venista.com>
Co-authored-by: wantosure <wantosure@126.com>
Co-authored-by: Solomon Neas <41877493+solomonneas@users.noreply.github.com>
Co-authored-by: Solomon Neas <me@solomonneas.dev>
Co-authored-by: ericcaiwx-star <ericcaiwx@gmail.com>
Co-authored-by: Teddy Tennant <teddytennant@icloud.com>
Co-authored-by: Yiğit ERDOĞAN <yigiterdogan023@gmail.com>
Co-authored-by: Vyctor H. Brzezowski <krzyszchweski@gmail.com>
Co-authored-by: qingminlong <qing.minlong@xydigit.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: holny <holny@foxmail.com>
Co-authored-by: holny <12433219+holny@users.noreply.github.com>
Co-authored-by: Altay <altay@hey.com>
Co-authored-by: Doksanbir <ylcn91@users.noreply.github.com>
Co-authored-by: coderdailyone <128658232+coderdailyone@users.noreply.github.com>
Co-authored-by: coderdailyone <codecookie@proton.me>
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
Co-authored-by: Agustin Rivera <31522568+eleqtrizit@users.noreply.github.com>
Co-authored-by: Alix-007 <li.long15@xydigit.com>
Co-authored-by: xingzhou <zhang.guiping@xydigit.com>
Co-authored-by: Orion <orionsieling@gmail.com>
Co-authored-by: Jonathan Sieling <jonathan@tailoredmonkey.com>
Co-authored-by: IWhatsskill <284122573+IWhatsskill@users.noreply.github.com>
Co-authored-by: yetval <yetvald@gmail.com>
Co-authored-by: NianJiu <3235467914@qq.com>
Co-authored-by: sappkevin <109696421+sappkevin@users.noreply.github.com>
Co-authored-by: Yalçın Doksanbir <yalcindoksanbir@gmail.com>
Co-authored-by: MasterSwords1 <142840981+MasterSwords1@users.noreply.github.com>
Co-authored-by: JC <anyech@users.noreply.github.com>
Co-authored-by: OpenClaw Assistant <assistant@openclaw.local>
Co-authored-by: leochun92-sudo <295001307+leochun92-sudo@users.noreply.github.com>
Co-authored-by: Mike <5853861+snotty@users.noreply.github.com>
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
Co-authored-by: RoboClaw <services+roboclaw@openclaw.org>
Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com>
Co-authored-by: Takhoffman <781889+Takhoffman@users.noreply.github.com>
Co-authored-by: brokemac79 <255583030+brokemac79@users.noreply.github.com>
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
Co-authored-by: juyaohuidt <ju.yaohui@xydigit.com>
Co-authored-by: sxh <xxh517598@gmail.com>
Co-authored-by: sxh <sunxianhong@ncti-gba.cn>
* fix(agents): scope append-only runtime context to prefix-binding Claude models (#136990)
* fix(agents): scope append-only runtime context to prefix-binding Claude models
Append-only runtime-context replay (persisted carriers, retained inline
inbound metadata, no consecutive-user merge on the Messages API) exists
because Claude Fable 5.1 binds thinking to the exact request prefix. It was
enabled for every Anthropic-family route, so Opus 5, Sonnet 5, Opus 4.8,
Sonnet 4.6 and Haiku 4.5 sessions paid for persisted carriers on every later
request (cache-read billing plus context) without any prefix to protect.
- Add `bindsClaudeThinkingPrefix` to the llm-core Claude contracts (Fable 5.1
and Mythos 5.1 identities, including cloud ids and canonical deployment
metadata) and derive `appendOnlyRuntimeContext` from it in the core
fallback and the shared Anthropic replay helpers; GitHub Copilot stays
transient. Non-binding Claude models return to transient carriers and
normal user-turn merging.
- Next-turn carriers now carry only the delimited context body. The shared
instruction (runtime context for the user request it follows, not
user-authored, keep internal details private, continue without waiting)
moves once into the stable system prompt prefix for every provider, so it
is cached instead of re-sent per turn. Runtime-event prefaces are
unchanged; historical prefaced carriers are still stripped from
user-visible surfaces.
- Docs: anthropic provider page, transcript hygiene, system prompt.
Live proof on a recording proxy with drop_block binding controls: an Opus 5
session merged consecutive user turns and kept its model-switch carrier
transient; a Fable 5.1 session switched to Opus 5 and back, then ran three
more turns with tool loops while the persisted compact carrier stayed in
place, thinking blocks accumulated (1 -> 3) and every response reported
input_transformations: [].
* fix(agents): limit prefix-binding replay to Fable 5.1 and match compact carriers in e2e
`bindsClaudeThinkingPrefix` now covers only the Fable 5.1 identity. Mythos 5.1
is not registered in the bundled Anthropic catalog and has no live replay
proof, so it keeps transient carriers and normal user-turn merging like every
other non-Fable model; extending the predicate requires live proof for the
new model. Predicate, transcript-policy, replay-helper and provider plugin
tests assert the narrowed set.
The qa-lab gateway RPC chat e2e test still imported the removed next-turn
carrier preface; its carrier filter now matches the delimiter-only shape.
* test(agents): expect merged user turns for non-binding Claude replay
(cherry picked from commit de7518bd05)
* fix(anthropic): preserve cache reuse across transient runtime context (#140621)
Anthropic tool loops could repeatedly write growing conversation history to the
prompt cache because a moving runtime-context carrier owned the checkpoint.
Use the existing replay lifecycle contract to exclude transient carriers from
both request builders while preserving retained anchors. Also omit empty beta
headers that caused direct API requests with thinking disabled to fail.
Add deterministic regression coverage and a packaged Docker test that verifies
real cache reads and incremental writes across two tool continuations and a new
user turn in both builders. Require that lane in stable/full release validation,
with explicit API-key preflight, no request retries, and a declared runtime entry
for dependency analysis.
Validation includes focused package/model/session and workflow tests, build and
package integrity checks, static checks, independent review, mock Docker, and
eight successful live Anthropic Docker requests. Hosted CI run 34379052492 passed
on the final PR head at attempt 2, after one unchanged browser-animation rerun.
Closes#140607
Thanks to @LightningWareLLC for reporting the regression and contributing the
lifecycle-aware cache repair.
Co-authored-by: LightningWareLLC <271410939+LightningWareLLC@users.noreply.github.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
(cherry picked from commit faaf4b0fbd)
* fix(anthropic): bind cache carriers to replay policy
* fix(release): qualify prefix-bound anthropic replay
* fix(github-copilot): keep replay context transient
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: LightningWare LLC <admin@lightningware.com>
Co-authored-by: LightningWareLLC <271410939+LightningWareLLC@users.noreply.github.com>
* fix(update): allow updates without a Gateway service (#136798)
The installed 2026.8.2 updater refuses the release smoke before package
mutation or candidate handoff. Current main has the same defect: strict
systemd command inspection cannot distinguish a missing manager from an
uninspectable existing service.
Let the native service-state owner record affirmative manager/unit absence,
then require strict Gateway lock inspection and a free configured port at
update preflight. Preserve unknown-state refusal and report skipped restart
when nothing is running. Lock discovery retains its existing contract.
Use the shipped baseline's documented --no-restart path only after proving
the dedicated smoke container idle, then repeat the candidate update with
default restart policy. Join the heartbeat timer exposed by that idle check.
Document the one-time 2026.8.2 upgrade workaround.
Proof: original baseline and packed-main Docker runs refused with exit 1;
managerless and strict-lock regressions fail before the fix. Final focused
suites pass 532 tests; the full update CLI suite passes 370 with 1 skipped.
Final Docker baseline/manual plus candidate/default updates and doctor
steps pass; pnpm check:changed passes. Independent review is clean through P2.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit b026948c8f)
* test(infra): split gateway lock inspection coverage
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
CI setup could fail before lint or tests when Matrix crypto native-library downloads hit a transient release-CDN error. Hosted store-only readers used a different cache path from the publisher, and hybrid mode seeded only the Blacksmith Linux backend.
Align the existing setup action and warmer on the workspace-local pnpm store, publish completed native installation outputs before unrelated build work, and add one install-only hosted warmer row in hybrid mode. Bound installed-Node discovery to four levels to avoid traversing bundled npm dependencies. The existing warmer remains the sole publisher and PR consumers remain restore-only.
Current-head hosted warm setup completed in 22 seconds with 1,436 packages reused, zero downloads, and Matrix postinstall skipped, compared with the recorded ten-job 45-second baseline median. This is an observed sample, not a measured fleet-wide median change. Cold caches and changed native inputs still require upstream assets. No application runtime, configuration, dependency, or permission changes.
Validation: exact-head PR CI run 35197161825 returned GREEN; the accepted Codex review and exact-head ClawSweeper review found no actionable issues. The PR also records passing targeted Node/cache guards, workflow checks, and cold/warm native reuse proof.
Related: #150637, #150526.
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Backport the Homebrew updater ownership and stable opt-path repair from upstream PR #141011 / commit 8db38db911.
Co-authored-by: Max Xu <xuhuan@live.cn>
Backport the validated package-manager fallthrough from #134386 and the bounded stale Winget registration repair from #112055.
Co-authored-by: luyifan <al3060388206@gmail.com>