Await buffered progress receipts in the fresh candidate receiver before finalization, preserving captured database context and live executor/requester admission. Refused receipts stop replay; uncertain writes retain the existing cleanup failure. This candidate-side change preserves published drivers, schemas, stored state, and recovery policy.
* refactor(agents): replace peer announcement loops with one reply delivery
Deliver delayed peer results once to the requester and return waited replies inline without a continuation. Preserve child completion custody and same-session generation-bound channel delivery. Use exact-key saved-route lookup, retire legacy peer control tokens from runtime decisions, and retain historical display suppression.
Simplify the announce owner, preserve required missing child output, and keep wake retries within the remaining operation budget. Update session guidance and the opt-in live peer scenario.
Candidate regression scenarios pass. Broad focused tests and changed-lane validation remain incomplete under host load; original-code regression proof is still pending. No live provider test was run.
* refactor(subagents): consolidate completion custody and lifecycle owners
Share terminal-effect persistence and requester-wake settlement adapters while preserving their distinct authority, frozen-wave, and publication guards. Remove unused preparation and accessor paths and reuse existing parameter contracts.
Propagate unknown SQLite write outcomes before scheduling interrupted completion recovery during restart drain. The focused candidate regression passes; original-code failure proof and the broader validation matrix remain pending.
* test(agents): align completion assertions with reply ownership
Assert that an ACP target already owned by its requester keeps task completion and starts no detached reply. Keep isolated cron fallback routing and participant-store checks while asserting no detached wait.
Remove assertions on the retired duplicate steerMessage field; retain complete source-reply content, ordering, exclusions, visibility and batch settlement checks. Remove the redundant inner same-session return while retaining the enclosing branch exit.
Validation: scoped Codex review at P2, formatting and diff checks passed. Two attempts on the coordinator Testbox lease failed during file sync before installation or tests. Baseline regression proof, the full matrix and the changed-lane gate remain pending.
* test(agents): inject the gateway caller into same-session delivery routing tests
* test(subagents): route direct requester-yield callers through prepared cron authority
* refactor(subagents): keep the requester wake currency assertion module-private
* fix(agents): fence failed and delayed subagent cleanup
Restore the previous resident row when same-ID registration persistence fails.
Carry collector cleanup authority into attachment removal, and bind lifecycle
grace timers to their captured registration and generation.
Regression coverage exercises failed writes, retirement during cleanup, and
same-ID successors without changing storage or timer durations.
* fix(agents): retain paused state in compact registry reads
Read pauseReason from existing stored JSON for ordinary and private-parent
records so cold projections preserve the resident paused status. Derive compact
SQL row types from the query and shrink the removed-assertion allowance.
No schema, serialized representation, or migration changes are required.
* fix(agents): preserve spawn cleanup and request scope
Route fork lookup failures through provisional-child cleanup. Reuse the spawn
admission owner for visible work with the resolved requester agent identity.
Pass trusted-creation transport timeouts through the argument the transport reads.
Consolidate in-process dispatch while preserving creation custody, source
fencing, per-interface deadlines, and signed fallback behavior.
* refactor(agents): simplify subagent registry and spawn owners
Remove unused registry APIs, callback declarations, presentation adapters,
registration inputs, and bootstrap scaffolding. Share registration ACK control
flow while retaining both durable writes and the original settlement error owner.
Complete keyed capability lookups, reuse heartbeat and timeout owners, carry
resolved model references, and remove obsolete ACP transcript work and relay
options. Preserve legacy stored attachment safety and protocol-v4 fallbacks.
* refactor(agents): simplify completion consumers and registry publication
Move full-registry fixture writes into test support and require named production mutations. Consolidate publication subscriptions while preserving projection ordering and persistence observer exception isolation. Remove duplicate completion retry scaffolding, runtime barrels, test-only exports, duplicate types, and single-caller adapters.
* chore(agents): shrink retired test API assertion allowance
* test(gateway): preserve real registry operations in agent fixtures
* test(subagents): preserve named publication fixture preconditions
Name all retired fixture rows when replacing describe/list facts. Warm the compact cache before its initial publication so the existing ordering assertion still detects an unintended SQL reload.
* fix(plugin-sdk): preserve capability store compatibility
Retain the public record-or-lookup store contract, including normalized first-match identity lookup and canonical depth fallback for partial records. Keep migrated internal callers on keyed lookups. Restore the base handoff declarations and requester query/caller linkage so the existing SDK declaration gate stays unchanged.
* test(gateway): publish fixture memory ownership
Publish the staged in-memory owner through commitOwnership before asserting describe and list projections. Named durable writes no longer trigger the whole-index rebuild that accidentally discovered the fixture row. Preserve the existing assertions.
* refactor(agents): tighten completion projections and ownership facts
Project only formatter-consumed child fields without mutating authoritative rows, and carry the session resolver's required requesterOwned boolean directly. Preserve formatting, nested identity references, authority checks, and public signatures while resolving the full core lint findings.
* test(agents): retire obsolete benchmark registry mock
Remove the full-registry writer no-op from the memory benchmark's typed module mock after the production writer cutover. Preserve the named-write mock and benchmark behavior.
* test(agents): await requester settlement in live fixtures
Child cleanup can publish before the requester delivery acknowledgement. Join persistence publications until the requester wake settles, preserving the existing delivery assertions and deadlines.
* refactor(agents): retire hidden sessions_send message aliases
Require the canonical message argument already declared by the public tool schema. Remove alias-only reasoning stripping and keep canonical message indentation unchanged.
* fix(agents): decode compact subagent metadata canonically
Replace the separate compact decoder with the canonical codec and session-list projector. Keep bounded SQL metadata projection, private envelopes, and duplicate-key last-value semantics without transferring retained results. Add real-reader duplicate-key and retained-content regression coverage. Schema, persisted bytes, and update behavior are unchanged.
* refactor(agents): remove redundant delivery comments
* perf(agents): restore compact reader after benchmark regression
Revert 2c0fd630c7 as required by the performance gate. Across five scoped 2000-row reads per revision, the base median was 227.31 ms and the candidate median was 429.25 ms (88.84% slower). Retain the sessions_send alias retirement and comment cleanup; leave the duplicate-key reader fix for separately approved work.
* test(agents): relocate sessions_send alias regression coverage
Keep all four rejection cases in the assembled-tool suite where retired preparation coverage freed space. The oversized sessions test file shrinks again; no assertions, mocks, timeouts, skips, or line-cap baselines change.
* fix(agents): retain sessions_send record input validation
Use the shared isRecord guard from the retired alias normalizer instead of a new type assertion. Preserve canonical missing-message errors for non-record input and satisfy the assertion safety gate without a suppression.
* refactor(agents): keep sessions_send replies on the requester key
Remove the legacy key-only DM-to-main reply remap and its routing exceptions. Use the resolved caller key for reply context, provenance, watches and followup preparation. Preserve exact-incarnation and authorization checks.
* test(agents): bind native followup custody in admission fixture
Keep the key-only DM requester authority test on the native child completion path. Settle the existing completion owner before mock acceptance and reset queued mock implementations between cases.
* fix(agents): reconcile subagent callers after main merge
Use the phase-aware publication API in newer main callers and explicit changed IDs in the prepared-read fixture. Keep retained regressions in their current sibling owners and place requester adoption in the lifecycle controller without weakening guards. Preserve heartbeat narrowing after shared-owner integration.
* test(agents): clean up merged fixture imports and names
* test(agents): refresh single-reply prompt snapshots
Regenerate prompt fixtures for the intentional sessions_send description change. The four-file delta contains only the description and derived sizes and hashes. Reproduced the hosted drift on Testbox, then passed prompt:snapshots:check and the snapshot-only changed checks.
* test(agents): await owned descendant settlement
* fix(agents): retain the leaf followup owner contract
Restore the standalone completion-owner interface and implements check from main. Deriving that contract from the implementation class creates a type-only cycle through the cohort projection. Runtime behavior and the exposed owner methods remain unchanged.
* fix(agents): disambiguate the sessions reply target resolver
Give the async sessions_send reply lookup a distinct exported name from the general synchronous outbound session resolver. Update its sole production caller and direct tests without an alias or behavior change.
* test(agents): remove the retired steer fixture argument
* test(agents): drop the retired sessions_send delivery mode from the follow-up yield live test
* test(ui): record renderer stall evidence when Control UI e2e waits fail
Two scheduled main runs failed intermittently in comment flows without
evidence of what the renderer was doing: a pin click hung for 30 s with the
failure read missing its deadline, and a comment delete never showed up
during a 15 s poll. Neither reproduced locally.
Mocked-Gateway pages now arm a stall probe before navigation. A renderer
that misses the failure-read deadline reports main-thread busy time by kind
and its paused JavaScript stack, then resumes. A stall that ended appears as
script-attributed long animation frames. Public output keeps only bundle
paths, positions, function names, and listener tag and event names.
* test(ui): arm the stall probe before tests open CDP sessions
PR CI showed 11 mobile safe-area geometry failures: installMockGateway
attached the probe after those tests had set
Emulation.setSafeAreaInsetsOverride on their own CDP session, and Chromium
drops an earlier session's override at the next navigation once a later
session attaches. A session attached first leaves later overrides intact.
The shared suite's withPage now arms the probe right after newPage(),
before test code runs, and installation is best effort for test doubles,
closed pages, and non-Chromium contexts.
* fix(update): await progress receipts before advancing commands
* fix(update): preserve Doctor refusal and progress worker boundaries
Keep the execution guard contract in the existing parameter types module
without an implementation import cycle. Join actual reporting failures with
settled Doctor errors while preserving later typed authority refusals.
Run the fresh-package staging regression in the worker-capable harness now
that its real progress callbacks use the shared-state writer. Preserve both
replay cases and their file and backup assertions.
* test(update): await candidate progress receipt settlement
* fix(update): retain recovery after progress receipt refusal
Preserve the existing maintenance shutdown owner and cover live inspection-child closure plus shared and agent handles through the real Doctor health flow. The suspected self-handle leak was not reproduced; distinguish fuser holder PIDs from inspection failures without relaxing refusal.
Includes the independently landed WAL inventory rescan for integration. Validation: 27 focused cases on Blacksmith Testbox, production and all 27 test type graphs, scoped lint and repository guards; isolated Codex review found no actionable P0-P2 findings. Commit hooks skipped after remote validation to honor the Mac host restriction.
A stalled cron sub-agent tool call drove sustained SQLite write contention; a state.lease operation with busy_timeout 0 failed with lock contention, and the per-agent database execution admission then closed permanently until restart because the execution owner retired logical admission after recoverable native-open failures and the worker transport dropped SQLite contention codes. Admission is now preserved across recoverable native-open failures, contention retries share the 25 ms cadence introduced for the lease heartbeat in #160702 and run only before application work starts, and shutdown, FIFO ownership and unknown-outcome safeguards are unchanged. Doctor preserves nested lease-loss diagnostics instead of flattening them; a nine-transaction Doctor probe confirms #160702 closes the heartbeat loss reported in #162083.
Closes#162579
Refs #162083
Reuse compatible test runtimes across source refreshes and test-only edits while preserving input invalidation and strict UI/deployment checks.
Related: #57204.
* perf(sessions): bound list materialization after invalidation
After a catalog or metadata invalidation, sessions.list materialized the
display rows of every session before selecting the page, so a burst of
concurrent lists spent seconds in the materialize phase (avg 4.7 s, max
19 s on Team with 334 slow lists per hour).
Select first, then materialize only the rows the page returns, and share
the post-invalidation work across concurrent lists. The reply contract is
unchanged.
Testbox fixture (8,000 sessions, 20 concurrent lists): average
materialization 534 -> 97 ms after catalog invalidation and 1,239 -> 841 ms
after metadata invalidation; selected display work 8,000 -> 200 rows.
* test(sessions): align list fixtures with bounded materialization
Align session-list fixtures with metadata readiness and selected-page
materialization. Fixtures that need every display row now drain their own
setup explicitly, while list wait tests observe prepareSelection and
preserve its arguments.
The synthetic plugin role-revocation test still waited for the retired
ensureMaterialized list call and deterministically timed out after 120 s.
Observe the actual reader readiness call, assert that each injected wait
was consumed, and retain the revocation-error expectation and describe's
bulk-read isolation guard.
Testbox: 435 tests across 40 files passed; the reproduced 120,058 ms
revocation timeout now completes in 670 ms. PR #162280.
* refactor(state): resolve Gateway profiles in state workers
Move the remaining Gateway email acquisition and canonical profile selection
callers onto existing state worker operations. Retain request and profile
authority across waits and recheck it before writes, network effects, and
metadata publication. Remove the synchronous Gateway access helpers.
Keep schemas, stored data, permissions, and update behavior unchanged.
Remaining profile reads/setters and account stores follow in a second slice.
Isolate session-catalog fixture writes from background reclamation so request
cleanup assertions remain deterministic across the worker wait.
* test(gateway): assert snapshot rejection at handler boundary
Profile acquisition now rejects discovery snapshots before model catalog loading. Preserve zero credential and HTTP effects while testing the direct handler rejection; WebSocket dispatch owns the error response.
* test(gateway): await concurrent metadata reader admission
Synchronize the concurrency fixture on all reader callbacks entering after asynchronous profile preparation. Preserve the original unrelated-creation, response, and cleanup assertions without timers or polling.
* refactor(state): register worker operations once per domain
Infer shared-state worker command contracts from lazy per-domain handler tables. Migrate Web Push, APNs, worktree registry operations, and fleet registry while preserving the existing broker and transaction owners.
* fix(gateway): retain one catalog acquisition authority check
Keep the shared post-await authority assertion before catalog effects and the final publication check, without duplicate session path validation. Migrate authorization fixtures to the async profile owner contracts and remove their retired resolver mock. Existing real Gateway freshness budgets and permission assertions remain unchanged.
* refactor(state): run profile setters in state workers
Move profile writer kernels behind the typed worker registry and route display-name and avatar RPCs through admitted worker writes. Recheck requester authority before refreshing connected profiles, use current committed display facts, and preserve committed self-merge responses. Keep role-policy reads as follow-up migration work; schemas and update behavior are unchanged.
* refactor(state): let profile handlers own kernel dispatch
* test(gateway): assemble synthetic question URL credentials
* test(gateway): synchronize concurrent project discovery
* fix(subagents): keep raw-key child cancel from stopping another agent's work
Cancelling a watched sessions_send child registered under a raw session key
such as `global` re-derived its owner from config and cleared every agent's
queue under that key. With two agents this aborted the default agent's run,
marked it aborted, and dropped its queued follow-ups and lane commands; even
with the owner resolved correctly, another agent's queued work was dropped.
Registration now records the known owning agent as optional `childAgentId`
for raw child keys (payload_json only; no DDL or schema-version bump). One
owner resolver feeds direct cancel, kill-scope discovery, the durable-kill
sweep, provisional reconciliation, and completion reads, and queue cleanup
uses the agent-scoped clearSessionLifecycleQueues owner. The raw-key
clearSessionQueues helper and its re-exports are deleted. Sweep session reads
move off the main thread onto worker reads.
* fix(subagents): align registry fixture with session owner reads
Configure the seeded fixture store and keep synthetic generation facts on the same session owner. Preserve real-store reads, identity and revision checks, and release invalidation; extract session mock wiring to keep the registry suite within its line-count budget.
Run Node-only maintainer tooling fixtures through the existing Node-selection helpers while keeping Vitest on its selected runtime. Preserve package-integrity, update interruption, isolation, and declaration assertions, and restore Bun-hosted compensation coverage.
Proof: all 21 consumers pass on Node 24 and Bun (1043 passed, four existing skips each); changed-file checks and import-cycle checks pass; independent P2 and exact-head ClawSweeper reviews found no actionable defects.
Land through the authorized native exception for CI run 36837194384 attempt 1: the untouched Windows session-creation path-alias failure independently reproduces on main; nine downstream cancelled jobs remain unrun coverage.
Move durable session reads and upstream marker settlement to the existing database workers. Preserve idle-owner admission, exact-link CAS, durable event-before-marker ordering, and joined shutdown; isolate settlement failures to each provider outcome.
* refactor(state): register worker operations once per domain
Infer shared-state worker command contracts from lazy per-domain handler tables. Migrate Web Push, APNs, worktree registry operations, and fleet registry while preserving the existing broker and transaction owners.
* refactor(memory): move retained index reads into workers
* test(memory): align fixtures with worker read ownership
Repair the ten fixture failures exposed by retained worker reads. Install the embedding generation and token budget, retain a file-backed startup owner, and preserve the explicit scheduler-yield proof through source-wide snapshots.
Follow-up to #161053 (shared-session emoji reactions).
Reaction writes now run through the SQLite worker admission the sibling
session stores use (runOpenClawAgentWorkerWrite); the native path stays
only for process-held incognito databases the worker cannot reopen by
path, and the handler revalidates live authority around the awaited
write.
The plugin action dispatch path awaits onPlatformSendDispatch right
before the synchronous handoff fence, exactly like the send path, so the
reaction mirror re-reads the conversation binding at the final handoff
and refuses a message whose conversation was rebound while the action
runner prepared delivery.
Channels with one bot reaction per message (Telegram bots, WhatsApp)
declare the new optional ChannelPlugin.capabilities.reactionSlots =
"single"; when a person removes one emoji while others remain, the
mirror re-sets the newest surviving emoji instead of clearing the slot.
Multi-slot channels are unchanged.
Also trims redundant scaffolding in the reaction handler, kernel, UI
component and worker.
Proof: 119 focused tests across store, handler, dispatch and UI; mocked
Gateway reactions e2e; typecheck lanes; database-worker inventory check;
live two-person Gateway proof with the qa-channel mirror reporting
delivered through the final-dispatch hook.
* refactor(cron): await standalone quarantine registration
* docs(db): refresh quarantine worker inventory
* test(cron): colocate legacy crontab warning coverage
Move the existing warning cases to their owning suite while preserving their assertions. This leaves room for the quarantine SQL regression under the Doctor fixture line-count ratchet and avoids unnecessary SQLite fixture setup.
* docs(sqlite): refresh worker inventory after main merge
Regenerate the existing inventory from the merged source. Keep the quarantine classification and PR production delta unchanged.
Related: #140086, #141885, #156535, #157838
## What Problem This Solves
A running Gateway doesn't see a newly downloaded hosted model catalog until it restarts. This PR publishes each accepted catalog through the existing prepared-runtime owner, without a restart. Model rows and their prices switch together as one generation, which keeps the invariant from #140086.
## User Impact
- Compatible downloads are adopted at the Gateway's background catalog check, or after an explicit `models.list` refresh. That refresh returns the currently accepted rows right away and runs adoption afterward.
- A turn admitted on catalog N keeps N's rows and prices until it finishes. New turns use N+1. Rows and prices are never mixed.
- A concurrent auth or config publication no longer postpones adoption to the next scheduled check (up to 6 h). Adoption waits for that publication to settle, then retries, up to 3 attempts.
- An owner whose build failed or timed out ends the adoption instead of waiting on unbounded work. Gateway shutdown cancels an adoption that is still preparing.
- Malformed, schema-invalid, too-new (`minVersion`) and older catalogs are rejected, and the previously accepted catalog stays in use.
- Changing `models.catalogRefresh.url` no longer needs a restart: the previous source's catalog stops applying, and the mirror's catalog is adopted at the next catalog check.
**Bad-catalog exposure:** with live apply, a *valid but wrong* published catalog reaches running Gateways at their next catalog check (at most every 6 h) or on the next explicit `models.list` refresh. It no longer waits for a restart. Recovery uses existing mechanisms only:
- Republish a corrected catalog with a newer `generatedAt`; Gateways adopt it the same way.
- Operators can set `models.catalogRefresh.enabled: false`, which withdraws remote rows and prices without a restart (covered by the Gateway integration test).
This PR adds no new kill switch, config option or env knob.
### Compatibility
No config keys, defaults, types, validation, stored rows, protocol or SDK contracts change. The only config-surface change is the `models.catalogRefresh.url` help text, which drops the stale "Changes apply after a Gateway restart" sentence, and its regenerated config-doc baseline hash. Existing configs validate unchanged and need no Doctor migration (maintainer confirmation: https://github.com/openclaw/openclaw/pull/158000#issuecomment-5913462783). Startup behavior is unchanged. Upgrade impact for existing installs: an accepted download activates at the next catalog check instead of the next restart.
## Why This Change Was Made
- `prepared-model-runtime.configured-refresh.ts` builds a complete candidate generation of the configured owners under the new catalog. One serialized commit then publishes rows, the accepted bundle, the pricing context and the reply-dispatch projection together.
- Adoption re-reads the stored catalog until the config it read under is still current, so a stale caller can't cancel a current adoption.
- Each preparation attempt has its own abort signal. A config advance restarts only the attempt; a newer catalog or shutdown ends the whole adoption, including pricing preparation.
- Between attempts, adoption waits only on publication gates: a pending replacement or an owner's pending publication.
- Adopted owners install the same plugin-retirement recovery as configured publication (#161267). After commit, a lost Gateway plugin loan republishes them through the normal recovery. Before commit, it restarts the adoption attempt, and the commit refuses any candidate whose plugin generation retired.
### Why downloads were restart-only, and what this keeps
Restart-only activation was a mechanism, not the goal. #140086 chose it to stop rows and prices from different catalog versions mixing, and #157838 was merged as "the prerequisite for applying new remote catalogs without a Gateway restart (rows and prices must switch together)". This PR is that follow-up. Every requirement those PRs set still holds:
| Original requirement | Source | How it holds here |
|---|---|---|
| Rows and prices from one catalog version; never mixed across reloads or new requests | #140086 | One serialized commit publishes owners, bundle, pricing context and dispatch; pricing contexts are keyed by the exact accepted catalog. Integration test: new rows appear only with new prices |
| Admitted work keeps its pair | #140086, #157838 | Runs carry their plugin generation's catalog; usage operations capture one pricing context. Integration test and live proof: the in-flight turn keeps the old price |
| Startup absence is a real state (no downloaded rows without prices) | #140086 | Absence → catalog goes through the same atomic commit; overlay absence tests unchanged |
| Worker replacement inherits the host's accepted pair, not a later download | #140086 | The commit updates the inherited pair; later workers and a worker-exit recovery keep it (overlay and integration tests) |
| Current enablement and source URL still gate eligibility | #140086, #156535 | Checked on every read and before adoption, including the default-install v1 fallback; disablement withdraws rows and prices together (integration test) |
| Bad or superseded downloads never replace the active pair | #140086, #141885 | Compatibility, `minVersion`, revision and `generatedAt` checks; stale reads can't cancel a current adoption (regression test) |
| Failed catalog checks retry at the remaining fresh interval, not a full TTL | #141885 | Unchanged scheduler behavior; the deleted notice test's retry case is restored for failed adoption (fails if the retry falls back to the full TTL) |
| Operators learn when a downloaded catalog is not yet active | #141885 | No longer needed: downloads activate at the next check. The restart notice and its tests are removed; `models refresh` says when a running Gateway applies the update |
| Billing-route prices switch with their rows | #156535 | `upstreamPricing` and `providerPricing` are part of the accepted catalog pair |
## Evidence
**Regressions.** Each fails with its fix reverted and passes with it:
- *Retries a scheduled adoption when its pending auth owner settles.* Runs through the real Gateway update scheduler. Reverted, it logs `remote model catalog check superseded; deferred to the next check`.
- *Does not let a read under a superseded config cancel the current adoption.* Reverted, both calls end `superseded`.
- *Ends adoption instead of joining a timed-out owner build.* Reverted, adoption never settles.
- *Does not hold Gateway shutdown on an adoption's pricing preparation.* Reverted, shutdown waits on the held preparation until the test times out.
- *Recovers adopted owners when their borrowed Gateway plugin retires after commit / before commit.* Without the recovery, both fail: `Prepared model runtime plugin generation retired` and `prepared reply dispatch runtime owner was not published`.
- *Uses the remaining stored TTL after a fresh startup check when adoption fails.* With the retry reverted to the full TTL, the second check doesn't run.
**Suites:**
| Suite | Result |
|---|---|
| `prepared-model-runtime.remote-publication.test.ts` | 10/10 |
| Gateway integration (`models-list.remote-catalog`) | v1 and v2 pass. Config and auth churn during preparation end `published` on the settled owners. Also covers retained admitted runs, rejected and stale bundles, worker replacement and disablement |
| `prepared-model-runtime*`, `server-plugin-reload*`, `update-startup`, and all PR-touched test files | pass |
| `tsgo:core`, all `tsgo:test:src` shards | pass |
| oxlint and oxfmt on changed files; `config:docs:check`, `config:schema:check`; max-lines, assertion-safety and test-timeout-race ratchets | pass |
Tests wait on owned completion signals (`withinTest`), not wall-clock deadlines.
**Live proof** on an isolated Gateway built from `cb8c9197fd` (no provider mocks). Later commits add plugin-retirement recovery for adopted owners, covered by the regression tests above, and rebases onto `main`. It used a real OpenAI key through `openai/gpt-4.1-mini`, and the build stamp was set before the real catalog's publication date. A client polled `models.list` back to back over one WebSocket for the whole run (1023 polls, no errors). One Gateway process (PID unchanged) and no restart:
1. The stored catalog was seeded with an older revision of the real `catalog.openclaw.ai` v2 catalog: generated 2 days earlier, `gpt-4.1-mini` priced ×10, plus one extra kimi row. `models.list` listed the extra row, and a turn priced **$4.00 / $16.00 per M** input/output.
2. `openclaw models refresh` downloaded the real catalog (`updated`, 1039 models). The listed rows didn't change for the next 7.1 s, and a turn in that window still priced **$4.00 / $16.00 per M**: a download stays inactive until the Gateway adopts it.
3. A long turn was admitted on the older catalog, then `models.list {refresh:true}` returned the older rows (extra kimi row still listed) and started adoption. The new catalog was visible 0.9 s later, while the long turn was still running: the extra kimi row was gone.
4. The in-flight turn finished at **$4.00 / $16.00 per M** (older rows and prices). The next turn priced **$0.40 / $1.60 per M** (real catalog).
5. `models.catalogRefresh.url` was moved to a local mirror of the real catalog through `config.patch`, and the mirror's catalog was adopted without a restart. The mirror then served malformed JSON: `models refresh` failed with `SyntaxError`, the model list was unchanged, and the next turn still priced $0.40 / $1.60 per M.
**Model picker during republication (also on `main`).** Right after the new generation commits, `models.list` shows the new generation's configured and static rows until its full catalog loads, then the full list. In the live run this lasted 109 ms. The same poller against a `main` build shows the same window after a `models.*` config reload (20 → 6 → 16 rows for about 350 ms), so this PR adds a new trigger for an existing behavior. It doesn't change it. The short list comes from the new generation, so rows and prices stay paired.
**Published-driver upgrade cells.** Candidate tarball built from a fresh clone at `6f682c7944` with the canonical Docker packaging script and no build-time overrides. sha256 `46d881a0…ef0ad`; embedded commit `6f682c7944`, version 2026.9.7. `6f682c7944` already includes the shutdown-cancellation and pricing-deadline commits. The current head the current head differs from it only by rebases onto `main`: `main` had moved the scheduled catalog check into `update-startup-catalog.ts`, and this PR's adoption call moved there unchanged (`git range-diff` shows no other production change; a `remoteCatalog: null` test-fixture field moved to main's relocated `cli-compaction.test-support.ts`).
| Driver → candidate | Scenario | Result |
|---|---|---|
| `openclaw@2026.9.6` | base | passed (930 s); updater outcome success, no recovery |
| `openclaw@2026.9.6` | plugin-deps-cleanup | passed (923 s); updater outcome success, no recovery |
| `openclaw@2026.9.7` (latest) | base | passed (813 s); updater outcome success, no recovery |
In every cell:
- Migration, post-Doctor config validation, survival, plugin-dependency cleanup and runtime-deps repair checks passed.
- The candidate Gateway logged ready, then its catalog check fetched and saved the hosted catalog about 0.2 s after starting. The hosted catalog is older than the candidate's build stamp, so adoption ends `unchanged`, which isn't logged. After a 300 s settlement window, `/readyz` (`ready:true`, nothing failing) and a Gateway `status` RPC passed, and the Gateway shut down cleanly.
- To show a logged terminal outcome, each cell was repeated with a loopback mirror serving a newer copy of the same catalog. Each check logged `remote model catalog applied` about 0.26 s after it started, followed by `/readyz` and `status` passing.
Not run: `openclaw@2026.9.7` plugin-deps-cleanup. At the previous head it failed inside the 2026.9.7 driver's retained-runtime verification (`Retained runtime entry does not reference its inventoried file: dist/a2ui-…mjs`), identically for a merge-base control package with none of this PR's commits.
Harness note for the 2026.9.7 cell: unchanged, the upgrade harness can't run against 2026.9.7. It seeds the retired `tools.toolSearch {mode:"code"}` setting, which 2026.9.7 rejects. The 2026.9.7 cell used the harness's existing Tool Search "absent" mode, a one-line local change that skips only that seed and its check. The 2026.9.6 cells used the unchanged harness.
**CI:** fully green on the final head ([run 36818367970](https://github.com/openclaw/openclaw/actions/runs/36818367970)). Earlier heads hit failures that reproduce on `main` in code this PR doesn't touch: `update-cli.target-schema` (main reproduction: https://github.com/openclaw/openclaw/pull/158000#issuecomment-5913464830), the type-suppression inventory, the Windows partition owner test and the Windows backup-rename test. Main has since fixed the last three. Config compatibility confirmation: https://github.com/openclaw/openclaw/pull/158000#issuecomment-5913462783.
No overlap with Pash/Sarah changes.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(release): require signed publication tags
* fix(release): accept SSH-signed tag retries
* fix(release): reject lightweight signed-commit tags
* fix(release): re-sign local publication tags
* fix(release): pin signed tags across publication
* test(release): stage signed-tag finalization helper
* test(release): model signed finalization tag
* test(release): model signed Android tag resolution
* test(release): model signed finalization refs
#157413 added a birthtimeNs equality term to the shared-state database identity check and re-asserted it right after schema setup. On hosts where Node reports Linux ctime as birthtime (statx unavailable), every legitimate migration write failed the identity check, so the Gateway looped on "Existing shared-state database generation changed" and Doctor took a fresh backup each cycle. One canonical identity helper now applies a process-stable host policy (birthtime participates only where the host reports a real birthtime; device and inode checks remain) for shared state, agent databases, session projections, and the worker identity sibling. The migration backup owner reuses bounded candidate-scoped backup groups, verifies every member, and distinguishes interrupted capture from damaged completed groups.
Closes#161746
Crabbox 0.69.0 ships openclaw/crabbox#2613, which #161465 needs to recover from backends that cannot capture native snapshots; older CLIs still pause the profile. It also carries GCP boot-disk selection per machine family (#2608) and definite capacity rejections (#2616). The managed fallback now downloads v0.69.0.
Move health, repair reporting, and legacy repair preparation onto the existing
shared-state read worker. Keep the ordered quarantine query, payload decoder,
source and snapshot ownership, and schema errors unchanged.
Group the existing Cron read dispatch in its domain and separate quarantine
store tests while preserving their assertions and shared fixture.
Prove warm and cold Doctor observations without caller-thread quarantine SQL,
existing store and dispatch contracts, and the real Doctor-to-Gateway process
flow on isolated Linux. Preserve schemas, retention, and migration writes.
* feat(sessions): snooze sessions from the sidebar menu
Active session lists need a way to defer a conversation without archiving
it or interrupting its work.
Store and protocol retain optional wake and set timestamps in entry JSON
and row snapshots, reserving the core names against plugin slot collisions.
Scoped patches validate durable identity, future wake
times, and eligible roots. Pin and archive clear snooze; snooze keeps pin.
Gateway user interactions and completed activity clear snooze against the
current entry, including snoozes set during a run. System events and
preserved-state runs keep it. Existing row publications carry the change.
The Control UI sidebar adds calendar-aware presets, Snoozed filtering,
Wake and Undo actions, localized wake labels, and deadline invalidation
through the existing attention controller. Root rows reuse the existing
trail, and adopted catalog rows share the visibility projection.
Docs explain the active-session overlay and patch contract. Coherent
sibling modules and shared test fixtures keep the original owners within
their line budgets without removing coverage.
Release-note context: snooze a sidebar session until a preset time, send
it a message, or let its run finish to bring it back. Snoozing never stops
or blocks the session, and its Gateway-owned state follows every client.
* chore(sessions): restore the plugin-seam note on lifecycle timestamps
* fix(ui): keep snooze preset times compact and teach the mock fixture snooze
The Snooze submenu labels already name the day ("Tomorrow"), so the time
column now shows only the clock time except for "Next week", which keeps
its weekday. The shared Control UI session fixture now applies snoozedUntil
patches with the same rules as the Gateway (null wakes, archive and pin
clear, a new wake time restamps), so the mocked dev server and e2e harness
hide and restore snoozed rows like a real Gateway.
* fix(sessions): drop snooze metadata when an entry is archived
Automatic archival (active-session cap, age retention, stale dashboards)
writes archive facts without the sessions.patch path, so a snoozed session
kept its wake time through the archive and a later restore left it hidden
from the Active list. The canonical entry shape now removes snoozedUntil
and snoozedAt from any archived entry, which covers every persisted write,
every read projection, and the existing-entry projection the restore patch
uses.
* chore(protocol): regenerate Swift models for session snooze fields
* test(gateway): keep the snooze session-utils test in inventory order
* test(ui): walk the Snoozed status option in the sidebar filter keyboard test
* test(ui): expect Snoozed after one arrow from Active in the owner filter e2e
Use the existing memory retrieval pool for noncreating index selection, counts, and vector inspection. Capture caller placement and matching facts before dispatch, share the existing scrub rules, and return only selected row identities and counts. Keep final lineage checks and write settlement with their existing transaction owner.
Validation: causal host count regression, 150 owning tests, native fetched-row and byte bounds, captured-input and read-error controls, full changed-file gate, and independent review.
* perf(state): retire periodic runtime integrity scans
Remove delayed and daily full-database scans from the Gateway while preserving requested agent quick checks, live-owner confirmation, and quarantine. Keep full verification with admission, migrations, and Doctor maintenance. No config or schema migration is required.
* fix(update): restore Windows task autostart after cancellation
Carry the existing restoration phase through Windows task recovery so SIGINT fences forward work without rejecting compensation. Preserve executor and native task ownership checks before side effects.
The original 40-file CI shard reproduced 821 passes and one SIGINT failure; it passes all 822 tests with this change. Testbox focused tests and changed checks passed, and independent Codex review found no actionable findings.
* fix(plugin-sdk): keep updater path context private
Pin the legacy home-directory facade to its existing eight exports so new internal updater helpers do not become public SDK contracts. Reduce the wildcard ratchet by one without expanding export or callable budgets.
All 11 SDK surface tests passed against the exact failed CI merge plus this fix on Testbox, along with core/script types, targeted lint, export guards, and formatting. Independent Codex review found no actionable findings.
* fix(update): settle service receipts and interrupted rollback
Await the existing bound worker phase receipt before native service stop,
then revalidate the original executor/requester. Retain already admitted
compensation through signal settlement, including its later phase writes,
without retaining the unbounded forward operation.
Keep current-core and rollback authority policies distinct. Preserve the
existing ledger kernel, records, schema and recovery semantics.
Related: #161385
Canonical stale-baseline context: #161766
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
* fix(test): register post-update worker fixture once
Keep its existing worker registration instead of adding a duplicate that
prevents CI compact timing generation. Preserve the strict uniqueness
guard and the three newly required worker routes.
The actual CI merge failed manifest generation and its existing registry
regression; the baseline and one-entry correction passed the same selector.
* test(update): retain real owners in module failure fixtures
Load the mutable-signal and execution-guard owners in the existing VM
fixture so post-update failures reach their original backup-retention
assertions. Keep all19 inner cases and the synthetic service boundary.
The original three-file CI group failed on the unexpected dependency call;
the two real-module additions restore17 outer passes and19 inner passes.
---------
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
Apply the existing exclusive-plan barrier to full-suite and explicit-parallel routes without changing deadlines, worker limits, or failure policy.
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Reuse current change snapshots and placement projections for ordered disk-probe discovery. Preserve live pruning and post-probe ownership checks, probe policy, and startup settlement. Prove the actual monitor avoids host inventory SQL with real SQLite storage and retains existing disk-pressure behavior.
* refactor(cron): move scheduler state writes into the worker
Keep skipped-run and startup transactions in the existing Cron actor. Preserve exact reservation cleanup policies, one-shot unknown settlement, and committed history/notification attribution across routing refreshes.
* test(cron): assert captured routing in watchdog alerts
* fix(cron): retain startup notification routing through commit
* test(prometheus): isolate metrics startup from Control UI builds
* fix(cron): compose current webhook outcome fixtures
Receive canonical 884cdb12ad and provide the required captured routing in its direct failure-alert fixture. Preserve unknown timeout/redirect outcomes and proven no-send classification. Move the existing payload-format case into the presentation owner without changing its assertions or the file limit.
* perf(memory): move generated cache writes to publication worker
Retain the captured published database and its writer admission while staging bounded cache rows. Recheck revision and session tombstones in the native transaction, and invalidate conflicting generations before releasing admitted work or a refused publication attempt.
* test(memory): keep bound capacity fixture cleanup synchronous
Preserve the real backend close receiver and both cleanup errors. Capture the fixture before manager startup and verify the actual database source; retain the producer synchronous close type without changing emitted runtime code.
Reuse existing recovery discovery and current placement projections for auto-suspend planning. Preserve candidate order, durable recovery blocks, live per-candidate checks, and final reclaim authority. The existing real-reclaim regression now detects host SQL during the unexpired scan.
Shared-state transaction diagnostics reported an executing worker command
as the default "state.write" label, hiding the owner of multi-second
holds. Reuse the worker command context the SQLite broker already installs
while preserving explicit operation labels, and give synchronous
plugin-state mutations their operation name.
Recording one cron result also read and decoded every historical run of
that job while holding the shared writer lock. Filter the history rows by
run ID before decoding, keeping store partition checks, the released-row
fallback, ordering, and first-terminal-result protection.
Testbox fixture (1.13 GB state store, 2,000 history rows per job, 64 KiB
payloads): maximum cron writer hold 259.6 -> 3.17 ms; refused peer probes
981 -> 6 across the same five writes.
Bun builds that release SQLite native handles still paid for conservative worker retirement because lifecycle policy checked the runtime name. Capture the native-close capability after library selection and reuse it across broker placement, retirement, reader cleanup, and inherited worker admission. Bound probe termination to five seconds; timeout/error admission returns conservatively while unreferenced cleanup retains any unconfirmed live worker’s private directory. Stock Bun and Windows stay conservative; Node retains its existing policy. Register the disposable native probe with the SQLite guard and include the diagnostic parser, probe, and worker in the maintainer wrapper source inventory so isolated provisioning loads its complete dependency closure.
Proof: recovery median 11.4 s -> 0.1 s with 0/0 measured transcript worker creations/retirements. Gateway 50x25 load replies improved 128 -> 263 versus Node 273. Node 24 and the signed Bun test fork cover touched lifecycle and startup tests; changed checks, import-cycle checks, and full build verify the rebased tree.
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
People who can send to or suggest in a session can put emoji reactions on
any saved prompt or reply in the Control UI. Everyone reading the session
sees the chips update live, the agent sees each reaction as a System line
on its next turn (never woken for it), and reactions on prompts that came
from a channel are mirrored back to that channel as the bot's reaction.
Operators whose role only permits viewing see the chips without controls.
Storage is a same-version `session_reactions` companion table in the
per-agent database, installed on open like the session-sharing tables and
inert for older builds; reactions never touch transcript bytes. Two Gateway
methods (`session.reactions.set`/`list`) and one session event carry the
state, the hello frame projects the operator's session cap so the UI applies
the server's rule, and reaction reads run in the admitted history worker. Channel
mirroring keeps one bot reaction per emoji, dispatches in commit order, and
rechecks reactor authority and the source conversation before channel I/O.
The picker is a tapback-style pill below the message with six quick emoji,
pressed state for your own, keyboard navigation, and a custom entry that
applies a complete emoji the moment it lands (IME-safe, single emoji only).
Chips animate on change and list reactors with "You" first.
Live proof on a real Gateway with two identities and a QA channel; storage,
Gateway, UI, e2e and authority tests; published-2026.9.6 upgrade check.
* refactor(update): persist execution phases in the worker
Await validation, activation, and inspected Git target receipts through the existing shared-state mutation owner. Preserve phase ordering, terminal-row no-ops, and the synchronous ledger companion.
Keep native receiver preparation with candidate validation and recheck the captured invocation before its probe after schema revalidation.
Validation: original-caller causal controls; 93 final caller cases with 263 applicable surrounding-contract cases; full changed-source gate, including all 27 core-test type graphs; independent review. Installed package qualification follows in the PR.
Related: #161104, #160592.
* refactor(update): bind phase writes to execution custody
Keep the original run and write capture together in the existing execution guard. Callers still await each receipt and recheck live authority before continuing. Preserve the unchanged worker transaction and synchronous ledger contract.