A retained .openclaw.package-backup-<pid>-<timestamp> symlink in the global node_modules (left by an earlier source-to-package update) made the updater inventory walk into an unrelated checkout and refuse the update with the host-owned plugin-link error. The updater now excludes its own recovery artifacts (package backups and activation control files) from inventory through one producer-owned naming predicate, and retires historical backups once a later update has completed, never the current run's backup. The already-installed 2026.9.6 driver still refuses on that first hop; the issue carries the recovery note.
Refs #161922
Follow-up to #161832: when archiving a verified-empty retired Telegram bindings file fails (for example the .migrated target already exists), the migration records a recoverable warning instead of refusing the update; non-empty or uncertain sources keep the preservation refusal. Sanitized refusal messages keep their diagnostic text.
Refs #161795
Thanks @ericcaiwx-star for the base fix in #161832.
An in-place update replaces the package's hashed chunks while the old Gateway is still running. Shutdown then lazily imported cleanup modules that no longer existed: 2026.9.6 exited with ERR_MODULE_NOT_FOUND for the transcript capture-operations chunk, and other teardown paths (MCP runtime retirement, provider local services, browser relay and Chrome MCP, worker temp-artifact cleanup) had the same hazard.
Cleanup owners are now retained when their resources are created, so teardown never first-imports code. Transcript capture ownership moves into the eagerly loaded capture-startup module, and shutdown persists heuristic notes without starting optional model inference. Provider local services register their stop callback when they start. MCP cleanup has a lightweight owner module. Browser teardown peeks the lazy accessors that production ingress loads. The worker pool preloads artifact cleanup in its async preparation step before creating resources, which keeps chalk/tslog out of the native hook relay's static graph.
The provider-runtime-lifecycle stable tsdown entry stays, because published-update compatibility bridges map older hashed chunks onto it, and Knip treats it as an entry. Protection applies to updates from a release containing this change; older running Gateways keep the limitation, as docs/install/updating.md describes.
Proof: a built-runtime smoke removed 7,011 hashed chunks after startup and still shut down cleanly (exit 0, 6.3 s). Focused shutdown, transcript, worker-pool, browser and MCP suites pass, along with core types, import-cycle checks, Knip and the CLI bootstrap guard. Codex autoreview is scoped-clean. The final CI failures were main-side inventory breaks (wrapper components, suppression count), fixed on main by b3b34beebe and bd11387977.
On Windows the managed update swap renamed the live package tree to its backup location once; a transient EPERM/EBUSY/EACCES from an antivirus, indexer, or Scheduled Task handle failed the swap and rolled the update back. The filesystem owner of the swap now retries those codes with bounded backoff (16 attempts), records each retry, re-checks package identity and update authority between attempts, and reports a named failure after the bound. POSIX behavior is unchanged.
Refs #162027
* fix(update): reject unattributed database writes during Doctor
Keep the captured fingerprint baseline through Doctor settlement. Maintenance ownership does not exclude independent SQLite writers, so any observed change, including a new database, must refuse automatic restoration through the existing receipt contract. Preserve current databases and the existing actionable recovery guidance without adding schema or transaction receipt infrastructure.
Forward-port the conservative safety correction from c3c20e0241 onto the current capture owner.
Co-authored-by: RomneyDa <6581799+RomneyDa@users.noreply.github.com>
Co-authored-by: vyctorbrzezowski <51521767+vyctorbrzezowski@users.noreply.github.com>
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
* test(update): expect Doctor-time writes to disable automatic database restore
Rebasing onto main surfaced two receipts that relied on interval-wide
attribution: the NOCOW physical replacement (#160877) and Doctor's own
schema upgrade during legacy-root relocation. Under the conservative
policy neither is attributable, so both receipts are ineligible for
automatic restoration.
---------
Co-authored-by: RomneyDa <6581799+RomneyDa@users.noreply.github.com>
Co-authored-by: vyctorbrzezowski <51521767+vyctorbrzezowski@users.noreply.github.com>
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
Doctor requester-authority validation copied the entire SQLite database in a synchronous worker on every authority callback (195 snapshot launches, 31 s for the configured-owner case; candidate Doctor 74 s). Maintenance now owns one reusable live reader; admission, physical file identity, and revocation are re-checked on every reused read, and the reader joins maintenance and database cleanup. Configured-owner case 31.4 s -> 3.6-4.6 s; snapshot launches <= 8 during admission, zero during maintenance.
Refs #161949
Refs #161867
* perf(state): retire periodic runtime integrity scans
Remove delayed and daily full-database scans from the Gateway while preserving requested agent quick checks, live-owner confirmation, and quarantine. Keep full verification with admission, migrations, and Doctor maintenance. No config or schema migration is required.
* fix(update): restore Windows task autostart after cancellation
Carry the existing restoration phase through Windows task recovery so SIGINT fences forward work without rejecting compensation. Preserve executor and native task ownership checks before side effects.
The original 40-file CI shard reproduced 821 passes and one SIGINT failure; it passes all 822 tests with this change. Testbox focused tests and changed checks passed, and independent Codex review found no actionable findings.
* fix(plugin-sdk): keep updater path context private
Pin the legacy home-directory facade to its existing eight exports so new internal updater helpers do not become public SDK contracts. Reduce the wildcard ratchet by one without expanding export or callable budgets.
All 11 SDK surface tests passed against the exact failed CI merge plus this fix on Testbox, along with core/script types, targeted lint, export guards, and formatting. Independent Codex review found no actionable findings.
Resolve read-only worker launches through the running updater's retained
generation. Keep session reuse generation-aware and join child work before
retiring its runtime files. This protects future self-updates; an already
loaded 2026.9.6 driver retains its first-hop limitation.
Refs #161865
The async predecessor-receipt read introduced in #161813 can cold-start a
SQLite reader after the updater's launch directory becomes unavailable.
Resolve snapshot inputs before selecting a safe worker cwd, so captured
absolute paths keep working without rebasing unresolved relative paths.
Retain a valid physical cwd while replacing an overlapping installation:
Node worker bootstrap reads cwd before worker code can recover. Keep user
and session paths anchored to the captured invocation directory, forward
that directory for local package artifacts, and restore it after worker
settlement when it still exists. Keep backup and rollback ownership intact.
Concurrent config write during admission: warn and re-read through the
existing config owner; invocation-relative paths keep their original base.
Doctor sees its own finished update row: terminal ledger settlement remains
authoritative before post-update Doctor; original-state capture is preserved.
The installed updater owns cwd protection. A candidate cannot patch an
older driver before staging; this enables the next update performed by the
corrected driver and adds no new candidate marker or environment protocol.
Published driver x candidate: openclaw@2026.9.5 installs this candidate in
the base/manual-restart cell; Doctor, Gateway readiness, and preservation
of unavailable-plugin configuration pass.
Validation: 34 focused local tests and 69 focused Linux tests; real removed
cwd, relative-path, and package-artifact regressions fail without the fix.
The full architecture command, five source-contract checks, explicit
boundary lint, and independent review through P2 pass.
## What Problem This Solves
When a scheduled automation keeps failing, OpenClaw posts "Automation X failed N times" to the chat and waits for the user to ask for a fix, even when the fix is something the agent in the conversation that created the job could make itself.
## User Impact
When a job created from a conversation (it has an owner session) reaches its failure-alert threshold (default: 2 failures in a row, same cooldown and incident dedupe), OpenClaw sends one repair request to that owner conversation instead of the first chat alert. The conversation handles it as an ordinary agent turn, as if it had received a message: its own session and transcript, its workspace, its normal tool policy, and its reply goes to that conversation's own route (chat, thread or forum topic). The request tells it to:
- fix the problem in the workspace if it can (for example the helper script or instructions file the job follows) and say what it fixed in one line;
- otherwise ask the user for exactly what it needs;
- say nothing if the failure was transient.
Heartbeat settings do not apply to the repair: `heartbeat.target: "none"`, `isolatedSession`, `lightContext` and `activeHours` do not route, isolate, trim or hold it, and it carries no first-heartbeat notice.
The normal alert is still sent when the job has no owner conversation, for operator-only jobs (command payloads, on-exit and stream schedules), for webhook alert routes, and when the job fails again after the repair request (once, noting that a repair was requested). A successful run clears the streak silently.
**Upgrade behavior change (intentional):** repair is on by default. Opt-out is per job: `failureAlert: false` disables both the repair request and the alert. No config key is added. The chat "Automation … recovered" notice is removed; recovery stays in automation history. This reverses the notice added in #146582.
## Why This Change Was Made
The owner conversation is where these jobs get fixed in practice: in our setup a user had to reply "just fix it" to a failure notice, and the agent in that conversation then repaired the job. This change starts that step automatically.
The repair is dispatched through the Gateway's existing `agent` method on the instance's lifecycle principal (`dispatchGatewayLifecycleMethod`, the same owner restart-sentinel continuations and exec-approval follow-ups use), with the owner `sessionKey` and `deliver: true`, so the agent method resolves the reply route from that session's stored delivery context, thread included. The cron service reaches it through one injected dependency (`runCronFailureRepair`, wired in `server-cron.ts`), so `src/cron` does not import Gateway code. Nothing in the `sessions_send` tool needed extracting: its default path also calls `agent`, but with `deliver: false` plus its own announce flow, which is what this does not want.
An earlier revision carried the request on a heartbeat wake (system event plus `requestHeartbeat`) and patched the result with a `heartbeat: { target: "last" }` override. That made the repair inherit heartbeat behavior it should not have: an isolated `:heartbeat` side session without the conversation's transcript, light context, the internal-handling prompt, and heartbeat routing. This revision removes that path (system event, heartbeat wake, override, and their docs and tests); heartbeat behavior is unchanged for everything else.
**Tradeoffs (maintainer decisions):**
- The request is scheduler-authored, not a user message: it runs on the Gateway's system principal with `internal_system` provenance (`sourceTool: "cron_failure_repair"`), so it carries no sender-owner identity. Under the existing trust model the turn gets the conversation's normal non-owner tool policy (workspace and exec tools, not owner-only control tools such as `automations`); a change to the job itself goes through the user's reply turn. No new authority plumbing.
- The job's name, payload and last error are wrapped as untrusted data in the request.
- The request goes to the job's owner at dispatch time. If it is lost (the job was removed in between, or the `agent` dispatch rejects), nothing extra is sent at that moment; the job's next failure sends the normal alert naming the repair.
- A repair reply is posted at any hour, like the failure alert it replaces.
Persisted state: the incident gains an optional `repair: { atMs }` marker. It means "this incident's first alert became a repair request". It survives restarts, and the next alert of that incident clears it.
## Evidence
- **Size vs `origin/main`** (merge base `accde9e31a`, head `435e097ee7c`): production +233/−85 (net +148), docs +14/−4, tests/QA +649/−49; 21 files. To stay within the line-cap ratchet, the Gateway dispatch lives in `server-cron-notifications.ts` next to the failure-alert transport (2 lines of wiring in `server-cron.ts`), `reconcileCronExitWatchers` moved to `cron-exit-watchers.ts`, and two input helpers moved from the qa-lab mock server to `mock-openai-input.ts`.
- **Our production shape, Telegram Test Server forum topic, live OpenAI model** (`telegram-e2e-userbot`, Convex-leased credential, the lease's forum supergroup with a run-owned topic created before recording and deleted after; fresh isolated gateway from this checkout via `--source-gateway` on ports 19951/19952; provider slot = logging proxy to the OpenAI API). Config mirrors ours: `agents.defaults.heartbeat = { every: "1h", activeHours: 08:00–23:00 Asia/Kolkata, target: "none", directPolicy: "allow", to: <the tester's DM>, lightContext: true, isolatedSession: true }`, `cron.failureAlert` and `messages` unset. The QA user posts in the topic, which creates the owner conversation `agent:main:telegram:group:<forum>:topic:<n>`. A command action seeds `scripts/meeting-sync.md` (step 1: `sleep 150`), adds an isolated agentTurn job "Meeting sync" owned by that topic and announcing to it (`timeoutSeconds: 45`), and force-runs it twice: both time out (`consecutiveErrors: 2`, notification status `not-requested`).
- Command (secrets omitted): `E2E_MOCK_SERVER_PATH=<logging proxy to the OpenAI API> E2E_ROOT_CONFIG_PATCH=<live model + the heartbeat block above, messages: null> node <runner with forum-topic setup hooks> --backend mock --source-gateway --gateway-port 19951 --mock-port 19952 --chat <lease forum> --scenario <scenario> --timeout-ms 1000000 --record events.ndjson --output summary.json` → exit 0.
- Topic timeline (all messages):
| Elapsed | Topic message (SUT unless noted) |
|---|---|
| 0.8 s | user: `Hi! This topic owns our hourly sync automations. Reply with one short sentence confirming you are here.` |
| 61 s | `I’m here in this topic for your hourly sync automations.` |
| 272 s | `Fixed scripts/meeting-sync.md by removing the 150-second wait that exceeded the automation’s 45-second timeout.` (the repair turn) |
| 393 s | `Sync: 3 new meetings.` (third run, `ok`, `consecutiveErrors: 0`) |
| 514 s | `Automation "Calendar sync" failed 2 times` / `Cause: timeout` / `Run started: …` (the unowned control's alert) |
| 620 s | user: `What did you change in the meeting sync automation earlier? One sentence.` |
| 651 s | `I updated scripts/meeting-sync.md to read the CSV immediately instead of waiting 150 seconds, which exceeded the automation’s 45-second timeout; the schedule stayed unchanged.` |
- The repair ran in the topic session itself: `chat.history` of `agent:main:telegram:group:<forum>:topic:<n>` holds, in order, the owner-ready exchange, the repair request (user role, `internal_system`), `read` → `exec` → `exec` → `edit scripts/meeting-sync.md` → `exec`, the one-line reply, the next run's result, and the follow-up exchange. No `…:topic:<n>:heartbeat` session exists (sessions list). Its 6 main-model requests carried the brief with the topic's ordinary tool surface. The step file changed about 30 s after the second failure.
- No `Automation "Meeting sync" failed …` alert and no `First heartbeat alert` text anywhere (7 SUT messages). Control: an unowned job "Calendar sync" with the same broken step fails twice and alerts (delivery status `delivered`); its step file stays unchanged. The provider log also shows one `NO_REPLY` turn without the brief right after the repair's `exec` calls that sent nothing.
- **Telegram Test Server DM owner, live OpenAI model, `activeHours` excluding now** (same harness with `--dm`, window 06:00–07:00 Asia/Kolkata at 21:06 IST, otherwise the same heartbeat block). The DM creates the owner conversation `agent:main:main`; the same owned job fails twice (timeouts). DM timeline: 29.5 s `I’m here in the chat that owns your hourly meeting sync automation.`, 310.8 s `Fixed meeting sync’s timeout by replacing its 150-second blocking wait with a nonblocking export-readiness check.` (the repair turn; it chose to delegate the edit to a subagent via `sessions_spawn`/`sessions_yield`), 339.1 s `Sync: 3 new meetings.` (third run `ok`). The repair request and reply are in `agent:main:main`'s transcript; no alert, no onboarding notice, 3 SUT messages.
- **Regression test** (`src/gateway/server.cron-failure-repair.test.ts`, new file, "repairs an owned job with an ordinary owner-topic turn whatever the heartbeat config"; 4.3–5.8 s): a real Gateway with its real cron service and `agent` method (agent command mocked at the runtime boundary), heartbeat `target: "none"`, `isolatedSession`, `lightContext` and an `activeHours` window that excludes now; an owned job owned by a Telegram topic session fails twice; the agent command runs once in `agent:main:telegram:group:<g>:topic:42` with `deliver: true`, `channel: "telegram"`, `to: <g>`, `threadId: 42` and the repair request as its message, and no announce is sent. Fails on the previous head (no agent run).
- **qa-lab end to end** (`cron-failure-repair-owner-conversation`, real gateway child, qa-channel, mock-openai, isolated state): `OPENCLAW_BUILD_PRIVATE_QA=1 OPENCLAW_ENABLE_PRIVATE_QA_CLI=1 pnpm openclaw --profile <isolated> qa suite --provider-mode mock-openai --scenario cron-failure-repair-owner-conversation` → `passed=1 failed=0`. An owned job fails twice and no alert is sent; the owner conversation receives the repair request, its ordinary turn fixes the workspace step file (`write`), the reply lands in that conversation, and the request and reply are in the owner session's own `chat.history`. Control: an unowned job with the same failures alerts, and the repaired job never does.
- **Service tests** (`src/cron/service.failure-repair.test.ts`, 11 tests): threshold → one `runCronFailureRepair` request to the owner session (no system event, no heartbeat wake) and no alert; the job's name, payload and error sit only inside `<untrusted-text>`; the next failure sends one alert naming the repair, later failures nothing more and never a second repair; a rejected repair request alerts on the next failure; success clears the incident silently; no owner → the old alert; agentTurn, systemEvent and script jobs repair, command jobs and on-exit or stream schedules alert. The heartbeat-route case in `src/cron/service.wake-now-real-heartbeat.test.ts` is removed (the file is back to `main`).
- On `435e097ee7c`: `node scripts/check-changed.mjs` exit 0 (line-cap ratchet, all tsgo shards, lint, import cycles, dead exports); cron repair/alert/persistence/real-heartbeat, gateway cron, cron-exit-watchers, server-cron, notifications and qa-lab mock-openai suites pass (30 files, 538 tests); `pnpm config:docs:check` is clean.
- **Existing-state upgrade (recorded on an earlier head, harness not kept):** published `openclaw@2026.9.6` created an owner session, an owned and an unowned agentTurn job with an announce failure route and each job's first failure; its own `openclaw update --tag <candidate tarball>` succeeded and both streaks and owners survived; one more failure each reached the threshold: the owned job's incident gained `repair: {atMs}` with delivery status `not-requested` and its repair request reached the model; the unowned job alerted. The persisted state is unchanged by this rework.
- **Config compatibility:** no config key is added or changed (`config:docs:check` OK). The behavior change is the default for owned chat-alerting jobs, documented above; per-job `failureAlert: false` opts out.
- Downgrade: a persisted `repair` marker is an extra field older releases ignore; an incident that was mid-repair at downgrade keeps its signature there, so the older release dedupes that same-cause failure (no new alert until the cause changes or the job recovers), as it does for any already-alerted incident.
- **Owner boundary:** `owner` is fixed at creation (`CronJobPatch` omits it), so it cannot be reassigned; dispatch reads the live job and sends nothing if the job is gone.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
The v13 wide-row state migration conflated a missing installed_plugin_index row with a row whose JSON was unparseable or shape-drifted, and dropped the table in both cases. An invalid row is now preserved in full in the existing diagnostic_events quarantine with repair guidance recorded, valid install records are retained independently of damaged metadata, and a preservation failure rolls the step back to v12. No schema change.
Closes#161329.
Landed under the pre-existing-red rule: the remaining CI failures were current main reds at this base (Codex app-server settlement fixture drift, fixed on main by 52e60fb42b; the subagent kill-tombstone / descendant-cancellation intermittent first seen on main hourly 36619831492).
`openclaw memory status --deep` kept reporting `dirty=true` for agents that had a system-only cron-base session (`agent:<id>:cron:<job-id>`). Metadata catch-up counted that session as eligible, but the indexer correctly excludes it, so re-indexing never cleared the flag (eligible = indexed + 1). Status and catch-up eligibility now use the indexer's own admission gate. What gets indexed is unchanged.
Fixes#161823. Thanks @alfred429 for the report and the source-level reproduction.
Proof: the real CLI on an isolated profile, in FTS-only mode with no external providers.
- On main, the agent stayed dirty after repeated indexing (eligible 2, indexed 1). With this change it's clean (1/1).
- A new normal session still marks the agent dirty (2/1) until it's indexed, then it's clean (2/2).
- The regression test fails before the fix and passes after, and 40 focused tests pass.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(exec): avoid extra prompts for audit suppression commands
Remove the command-text gate and its dedicated approval/reviewer plumbing. Suppression configuration and audit filtering remain unchanged; normal exec policies still decide execution. Keep only the shipped deprecated SDK signature as a no-op.
Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>
* fix(exec): avoid extra prompts for audit suppression commands
Worked on by:
- @jesse-merhi
Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>
OpenClaw-Publication: f1fe3778-77d9-432e-b98f-01f680f2cdd3
* fix(plugins): preserve deprecated suppression approval predicate
Retain the shipped SDK result during infra-runtime retirement without restoring internal execution callers. Document the distinct runtime and plugin migration outcomes.
Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>
---------
Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>
* fix(sessions): reject compaction of missing sessions
Return an actionable INVALID_REQUEST instead of a successful no-op for missing targets in both compaction modes. Preserve explicit no-op outcomes for existing empty sessions.
* fix(sessions): reject malformed reset agent selectors
Use strict session agent input validation before lifecycle cleanup so an invalid explicit selector cannot reset the default agent. Extract reset-target resolution from the oversized lifecycle service while retaining selected-global targeting.
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Confirm installation-only pending records for disabled channels and
installed plugins without config entries. Carry retained plugin IDs
through package convergence while preserving real migration obligations.
Fixes#161288 and #160523. Thanks @akennedytog and @g62nkcx4q7-wq.
The Gmail watcher stayed down for good when a Gateway restart briefly found its port still in use (EADDRINUSE). It now retries the bind a few times with bounded backoff and recovers once the port frees. If the port stays taken, it reports a clear error and stops retrying.
Fixes#161467. Reported by @kazuyuki-eguchi.
Proof: a real isolated Gateway, with the Gmail watcher on local fakes.
- Before the fix, the watcher stayed down after a transient port conflict. After it, the watcher recovered and answered.
- With the port held for the whole run, the watcher stopped after the initial attempt plus three retries, with a clear error.
- The regression test fails before the fix and passes after, and 31 focused tests pass.
- Updating from the published 2026.9.6 build to this build succeeded in two fresh runs (102 s and 97 s), and the installed build passed both the transient and the persistent control. The first updater run failed at service activation and didn't recur. The watcher can't reach that step: rehearsal disables hooks, and this diff doesn't touch activation or lease code.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
`openclaw update` could hang after printing its result because the retained updater runtime waited on the SQLite broker with no bound, and a force-exit watchdog was unsafe while accepted operations were still settling. The broker close is now split: settlement of accepted operations, pending opens and live references must complete; native close and worker termination are bounded afterwards, with the bound expiry recorded as a warning and the exit still happening.
Closes#160690.
Landed under the pre-existing-red rule: the remaining CI failures were current main reds in the merge window (check:architecture import cycle from b36eb3e7b1, fixed by 2a0a65c4ae; update-candidate-canary and cron service tests, fixed by 5f76cc437d).
Build on main's separate caller/service projections from 1794d8b4ef.
Propagate freshly read configuration through admission and already-current
completion, keeping explicit channel/source requests authoritative.
Refresh matching legacy projections and candidate validation after root or
include saves, preserving invalid-config, schema, and live-authority checks.
Concurrent config write during admission: warn and reread current bytes;
a save alone never triggers the former configuration-changed refusal.
An implicit stored-channel change resolves the whole target again, discards
old staged/admission/confirmation facts, and rechecks its original executor
bindings before execution.
Cached candidate admission carries the source snapshot it inspected. A changed
snapshot revokes old checks and reinspects the retained candidate before any
activation shortcut; only the admission result is published to history.
Late Git admission revalidates against a retained runnable candidate before
checkout. Service preparation and activation notification occur once.
Doctor sees its own finished update row: real original-state capture and
terminal publication remain visible to subsequent Doctor as completed;
no history or terminal-state policy is weakened.
Retain the Git candidate and private repository through finalization. Retire
candidate, repository, then previous-runtime backup after verification;
cleanup warnings retain dependent resources and recovery backups.
Published-driver x candidate: the installed updater runs first, so this
admission repair applies when that updater is installed. The candidate
continues to support the run/driver markers shipped by openclaw@2026.9.5.
The published-driver survivor cell passed with manual restart. Original
captures remain manual evidence, distinct from verified rollback snapshots;
rollback preserves later operator writes and existing ownership checks.
Proof: representative attribution 55/60 -> 59/60 reversing only the original
capture change; Linux original cluster/siblings 271/271, invariants 205/205,
execution 113/113, and the published-driver survivor cell passed. Local broad
proof covered 85 files; five stale fixture contracts were corrected and the
12-file follow-up passed 168/168, including both new cache regressions.
Combined coverage: 1,496 unique cases, with no weakened assertions.
The new cache cases took 4.358s/1.983s; the real Doctor invariant took 22.186s.
These exercise real config, admission and ledger boundaries.
Final local check-changed, production/test types, typed and boundary lint,
all five source-contract checks, and full Knip passed. Independent P2 review
found no actionable issues. Full owning configs are deferred under the
approved incident exception; focused proof ran with at most three workers.
Reuse managed-service publication checks at each runtime publication boundary,
comparing the publisher's exact physical outputs. Live disjoint runtimes stay
available; observed shared consumers must be stopped before replacement.
Document that stopped sibling profiles retain their own state and may need the
updated CLI's existing Doctor migration before resuming. No schema version or
automatic sibling migration is introduced.
Validation: 67 focused tests pass; five overlap regressions take 202 ms.
Fresh managed review is scoped-clean through P2. Isolated Linux native proof of
the composed selector/runtime candidate covers disjoint, overlapping live,
and stopped-sibling publication with preserved conversation witnesses.
Preserve the signal-driven recovery path when startup triage declines. Use the existing supervisor identity to gate terminal failed recovery, retain one retry after confirmed cleanup, and print platform-specific manual recovery steps. Refs #159539, #160173.
Preserve the caller Doctor projection separately when an explicit channel request selects another managed profile. Keep service capture, target selection, runtime materialization and writes owned by the selected profile, and retain strict source-bound schema validation.
Retain both published legacy convergence regressions, correct their file-backed environment fixture, and cover refusal without an explicit channel request. Extract admitted update orchestration without changing its body so the command stays within the existing line limit.
Canonical missing-copy handling in the shared archive walker returned no warnings, so an operator whose session archive copies were deleted got no signal. Missing copies are now warning-level facts (count plus up to five sample paths); blobs stay retained and migrations complete.
Refs #160770 (the reported stack overflow remains unreproduced; reporter asked for the full stack).
Landed under the pre-existing-red rule: remaining CI failures were current main reds (github-publication-personal-pending, server.sessions.list-changed, server.heartbeat-store-lifecycle.product-proof).
* fix(update): preserve original state before direct updates
Retain manual recovery evidence before direct update initialization and
Doctor state relocation, and carry the original reference through delegated
Doctor calls. Reuse the existing capture format, maintenance custody and
configuration reader, while deferring debug persistence until admission.
Distinguish manual and incomplete captures, and report restoration only
from evidence bound to the same manifest. Published legacy and inherited
updaters retain their existing limits; this does not add automatic restore.
Co-authored-by: Jason (Json) <263060202+fuller-stack-dev@users.noreply.github.com>
* test(update): admit contention fixture as an npm install
Preserve the real SQLite contention and history assertions by supplying package-manager ownership in the existing fixture mock. Clarify why updater debug persistence remains deferred until Doctor has admitted the schema.
* fix(tooling): include debug deferral in PR wrappers
* fix(update): preserve early refusal history and diagnostics
Capture originals before admitting refused updates to history. Preserve managed installation context and structured pending diagnostics without finalizing an uncertain run. Keep JSON refusals and ledger-only repair consistent with the recorded outcome.
* refactor(update): isolate initialization admission types
* fix(update): report tracing deferral during dry runs
Warn on stderr when the selected update environment enables HTTP capture,
while preserving tracing deferral and machine-readable JSON output. Extend
the existing preview controls and align the capacity refusal assertion
with the recorded failed-run contract.
---------
Co-authored-by: Jason (Json) <263060202+fuller-stack-dev@users.noreply.github.com>
A config change applied by an in-process restart whose startup refused the new config left the process alive but serving nothing: triage removed the refused setting, but the restart loop discarded triage completion and waited indefinitely. The startup owner now retries once after confirmed triage cleanup and exits nonzero when recovery fails, so the supervisor restarts it.
Closes#159539.
Landed under the pre-existing-red rule (#161124 and session-store main reds).
The process census treated unfamiliar-but-readable argv as a failed inspection, so an unrelated `bun run --silent` process vetoed Doctor's legacy plugin-capture and retained-runtime cleanup. The census now returns a closed holder | non-holder | unresolved result: only positive evidence of foreignness permits cleanup, and every unresolved evidence source (unreadable argv/cwd/package identity, batch timeout) keeps the veto.
Closes#159950. Replaces #159951 (thanks @Marvinthebored for the finding).
Landed under the pre-existing-red rule: remaining CI failures are current main reds from #161124 (server-restart-sentinel, github-shared-publication-events, server.sessions.list-changed, github-publication-personal-pending).
Classify fresh inspection before service-config repair and again before activation. Initially stopped Gateways remain untouched on inconclusive inspection, while Doctor-owned stops retain best-effort restoration and verified running services retain health checks.
Refs #159577. Pre-fix controls reproduced unwanted starts/restarts. Final owner suite passed 3x locally; 32 importer files and the five-suite 3x matrix passed on Testbox tbx_01m3jx5am10xv49bp07664y9an. Final check-changed and direct-route P1 autoreview passed.
Keep incompatible local plugin configuration intact and complete deferred migration confirmation after unchanged package convergence. Use completion receipts to clear obsolete retry warnings without rewriting update history.
Fixes#159477. Thanks @EndeavorPioneer for the recovery report.
* fix: allow Doctor repair with canonical service home overrides
Compare process and runtime homes independently with the OS account home so a canonical OPENCLAW_HOME in service metadata does not block Doctor credential migration. Preserve relocated-home isolation, active-writer checks, and the existing migration owner.
* test: keep inherited Vitest preloads out of verifier children
Carry the preload-owner fix and regression coverage from b9cd492ac6
(#160024). The failing CI merge inherited a newer preload that tried to
resolve vitest/runtime relative to temporary JavaScript fixture entrypoints.
It crashed before those children could register their IPC handlers.
Apply the adapter only to recognized Vitest workers, preserving their
package-local runtime while leaving ordinary forks, threads, and evals
alone. Keep the accepted canonical-home repair unchanged.
Validated the four original failures and the preload regressions before
repair, three focused passes, the original CI shard, changed checks, and
fresh P1 review on a Linux Testbox.
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
* test(update): align canonical home recovery guidance
---------
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
* fix(cron): report active-run cancellation on remove
* fix(cron): document active-run removal for managers
* test(cron): expect cancellation when removal aborts a live run
Sibling tests still deep-equal the old removal result, so an in-flight run now fails them.
Co-authored-by: Cursor <cursoragent@cursor.com>
* docs(cron): clarify self-removal cancellation exception
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(update): recover Git Gateways after activation Doctor failures
Keep requester checks inside retained maintenance custody and preserve read intent through borrowed ownership. Register Git rollback before activation Doctor, snapshot databases through the existing owner, and verify source safety before database restoration.
Refs #159839. Reported by @hannesrudolph.
* test(update): cover maintenance authority and Git rollback
Reproduce cold maintenance admission and cached policy revocation, reject new work during scope closure, and exercise Git migration failure plus source-edit refusal through the existing database recovery fixture. Document candidate-side authority repair and installed-driver rollback behavior.
* fix(test): include upstream reliability assertion export
Carry the already-landed fix da2efd83b8 so package activation tests typecheck against the integrated recovery owner. The assertion implementation and production runtime are unchanged.
* fix: recover the Gateway after activation Doctor fails
Keep Doctor policy reads and nested migration cleanup inside held maintenance authority. Bind Doctor generation receipts to the pre-migration database backup so attributable writes can be restored and later writes are preserved. Restart and verify the appropriate Gateway after either recovery outcome.
* fix(nodes): explain and recover session-host setup problems
Report unsafe workspace ancestry before dispatch, resume pending node pairing
after approval, explain runtime command availability at its authority owner,
and log inventory publication failures only when they change. Share inventory
validation and remove superseded policy and projection helpers.
* test(nodes): assert specific dispatch remediation messages
* test(nodes): align pairing retry coverage with node recovery
* fix(nodes): preserve desktop access after hosting failures
Keep connected-node environment availability separate from session hosting readiness and retain diagnostic placement refusal. Include Codex plugin installation in missing-command remediation. Cover the inventory-to-environment boundary with real connected-node enumeration.
* refactor(nodes): share runner declaration validation
Keep the status-wait capability added on main while extending the existing strict record schema with bounded hosting diagnostics. Remove the superseded declaration parser and clone closed worker-host snapshots through one path. Production code remains net zero against the refreshed base.
* fix(nodes): retain exhaustive command diagnostic states
* refactor(nodes): preserve explicit diagnostic return flow
* fix(update): restart previous Git runtime when branch reflog is missing
* fix(update): surface retained Git rollback diagnostics to the operator
* fix(update): keep the detached rollback advisory within report limits
* fix(update): suggest a compare-and-swap branch restore after detached rollback
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Add an opt-in slack-huddles plugin that joins active Slack huddles as a
dedicated signed-in Slack user. Slack has no bot or app huddle API, so the
plugin drives the Slack web client in the OpenClaw Chrome profile on the
shared meeting runtime, with in-page audio capture and a virtual microphone
for agent, bidi, and transcribe modes, mirroring zoom-meetings and
teams-meetings.
Membership is proven only by Slack's in-huddle channel header on the
requested channel. Sessions are workspace-scoped, joins are always muted,
talk-back unmutes only after Slack reports the virtual input, and every
click rechecks live authority. The shared meeting status source gains
optional liveOwnershipSource and afterAudioRoutingSource hooks. They
recheck ownership after awaited device and sink work, and roll back only
this pass's effects. Absent hooks emit no code, so the generated Zoom and
Teams scripts are byte-identical.
The plugin is disabled by default and requires OpenClaw 2026.9.8 or newer,
the first release that can carry the ownership hook.
Release note: Add an opt-in Slack huddles plugin that joins active huddles
as a dedicated signed-in Slack user through the shared meeting runtime.
Requires OpenClaw 2026.9.8 or newer.
Closes#160689
## What Problem This Solves
Fixes: `openclaw doctor --session-sqlite inspect` keeps reporting `plugin_migration_source_retained` after deferred plugin migrations complete, and repeated `openclaw doctor --fix` never clears it, when the import receipt includes unindexed session history with `.trajectory-path.json` pointer sidecars.
## User Impact
User impact: after upgrading with deferred plugin migrations (for example Codex or Brave), `doctor --fix` archives each unindexed transcript together with its trajectory pointer, and installs already left with stranded pointers get them archived on the next `doctor --fix`, so the warning clears.
## Why This Change Was Made
A deferred import receipt captures each discovered transcript together with its trajectory and pointer sidecar. When the plugin later completes, settlement no longer rediscovers that history, so its transcripts are moved by the unreferenced-JSONL sweep, which moved only `.jsonl` files. The pointer (a `.json` file) stayed live. The retained check fires while any receipt source is live, so the warning never cleared.
Doctor's archive sweep now moves a transcript's pointer sidecar with it. For installs an earlier release already left in that state, the same sweep also archives a receipt-captured pointer whose receipt-verified transcript is no longer live. Both paths use the existing archive move: identity and byte checks, a migration-manifest entry, and `doctor --session-sqlite restore` recovery. Nothing is deleted. Pointers still referenced by a live index, retained for another owner, in conflict with the receipt, or (during settlement) outside the receipt are left alone.
Update behavior: no schema, receipt or stored-field changes. The installed updater runs first, unchanged. The candidate's Doctor keys on the deferred-plugin session receipt and migration manifests that 2026.9.5/2026.9.6 already write, so the first `doctor --fix` after updating repairs existing stranded installs. Agent/state DB `user_version` and schema are identical to `main` after the same migration.
Related: introduced when deferred-import settlement began archiving receipt-captured unindexed history (#150015, first shipped in 2026.9.5).
No overlap with Pash/Sarah changes.
Thanks @spikewillcocks for the detailed report and receipt analysis.
## Evidence
Real CLI, isolated `--profile p160689base` / `p160689cand` with task-owned `OPENCLAW_HOME`, `OPENCLAW_STATE_DIR` and `OPENCLAW_CONFIG_PATH`; every Doctor run first printed the resolved state dir, config path and session DB path from the same binary. State: a file-era `sessions.json` (1 indexed session) plus 3 unindexed transcripts, each with `.trajectory.jsonl` and `.trajectory-path.json` (pointer format from v2026.9.5), plus unrelated `notes.txt` and `custom-settings.json`. The **published 2026.9.5** package ran `doctor --fix` with `brave` configured but its package unreachable, which deferred the migration and wrote the import receipt. The source checkout then completed the `brave` migration, as in the report.
- **Base** (`origin/main` 9366fd94e1): settlement archived the transcripts and left `history-{1,2,3}.trajectory-path.json` live. `inspect` → `1 issue(s)` `[plugin_migration_source_retained] … await archival. Run openclaw doctor --fix to finish.` A second `doctor --fix` changed nothing and `inspect` still reported it.
- **Candidate, same 2026.9.5 state**: `doctor --fix` archived all 4 pointers with their transcripts into `agents/main/session-sqlite-import-archive/` (for example `archive-tier.history-1.trajectory-path.json.imported-…`, manifest kind `trajectory`, reason `unreferenced-history`). Each archive's SHA-256 matches the receipt. `inspect` → `0 issue(s)`. A second `doctor --fix` wrote a run manifest with `completedMoves=0`, and no other file changed.
- **Candidate on the stranded base state** (the reporter's situation): `doctor --fix` → `Archived 3 legacy transcript artifact(s)`, the three stranded pointers now in the archive with receipt-matching SHA-256. `inspect` → `0 issue(s)`. A second run → `completedMoves=0`.
- **Controls**: `notes.txt` and `custom-settings.json` in the same sessions folder were byte-identical in every run. The regression test covers the explicit settlement path (`settleRetainedDoctorSessionSources`). There, a transcript and pointer written after the receipt (live, not covered by it) stay byte-identical, and so does an unrelated file. When the plugin completes during config preflight instead, Doctor's existing sweep archives every unreferenced JSONL in the folder, including ones written after the receipt. `main` does the same with `late.jsonl`, but leaves `late.trajectory-path.json` orphaned; the candidate moves that pointer with its transcript. A pointer still referenced by a live index or retained for another owner is skipped by the same guards that protect transcripts.
- Schema: after the same migration, base and candidate agent DB `user_version=23`, state DB `user_version=19`, with identical `sqlite_master` hashes.
- Regression test `doctor-session-sqlite.shared-orphan.test.ts` covers both paths. First, settlement of a deferred receipt with unindexed history and a pointer. Second, repair of a pointer an earlier settlement left behind; that half requires a distinct second archive move with receipt-matching bytes. The test fails on `main` because the pointer is still live after settlement. With only the stranded-pointer pass disabled, it fails at the repair step because the pointer stays live. It passes with the fix. Related suites (`deferred-plugin`, `manifests`, `retained-source-verification`, `active-settlement`, `indexless`, `archive-safety`, `recovery-shared-owners`, `discovery`, `recovery`, `receipt-recovery`, `doctor-session-sqlite`) pass: 12 files, 139 tests.
- Test cost: the new test is 1.7s locally and 1.8s in CI (`checks-node-changed-compact-large-19-1`, shard `agentic-commands-doctor-sessions-cron-hosted-1`, run 36491519850). `pnpm test src/commands/doctor-session-sqlite.shared-orphan.test.ts --maxWorkers=1` takes 34.4s wall for the whole file, which is mostly transform; its three tests run in about 6s. It uses no timers, sleeps, polling or process boots, and reuses the file's existing fixture and imports.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* prototype(plugins): show per-plugin auth alerts from live MCP checks
Place the shared alert beneath the individual plugin hero on main 018639af3b. Read a real Notion HTTP OAuth challenge and use a Notion-specific read-only account probe for the local stdio server. Keep the app inventory isolated and the prototype out of production imports.
* prototype(plugins): simplify auth alerts to one line
* feat(plugins): connect detail auth alerts to MCP OAuth
* fix(plugins): align auth alert with the main content column
* feat(plugins): show installed accounts and credential sections
* fix(plugins): use the supported credential projection copy
* feat(plugins): simplify setup states and retain edit actions
* fix(plugins): reuse MCP ownership across inspections
* fix(plugins): keep MCP cache types in the metadata owner
* test(plugins): preserve real install failure classification in fixtures
* fix: migrate retained plugin settings before update activation
Complete selected plugin-owned config repairs through the existing update and replacement-install lifecycle before activation. Preserve pending obligations with actual config rollback, retain published package generations when rollback is unconfirmed, and refuse unresolved active inputs without suppressing committed partial updates.
* test: align replacement-install warning fixture snapshots
* fix(update): bound stalled runtime cleanup and reclaim legacy projections
Forward-port retained runtime review corrections from c9de04ac1a. Preserve the recorded watchdog outcome and retained files when workers cannot close, recognize legacy pnpm projections within the same store, and inspect service TMP/TEMP roots without guessing relative paths.
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
* test(doctor): type the bounded projection iterator fixture
Narrow the first directory entry and declare the generator's undefined return so the synthetic opendir iterator matches Dir's AsyncIterator contract under the core test typecheck.
* fix(update): keep failed recovery barriers nonzero after the output watchdog
A rejected mutation or recovery drain must turn a recorded successful command result into exit 1. Preserve recorded nonzero results and the original result for disposable retained-worker cleanup stalls. Document the distinction and extend the existing real-child exit cases.
PR #160164 follow-up. Before the fix, three new failure cases exited 0 instead of 1; after the fix, all 11 retained-runtime exit cases pass. The regression file took 13.50s wall with --maxWorkers=1. After merging main, the five requested test files pass 79/79 in 99.18s. Core, other, infra, and config-cli typechecks, targeted oxlint, formatting, diff checks, and independent review pass.
The published-driver test file passes 11/11 twice on the original branch and once on main. The reported case takes 7.33s and 6.82s on the branch versus 7.88s on main; its stable finalizer exits naturally without invoking the output watchdog. No PR-specific exit regression reproduced, so no CI-related production or test change was made.
* fix(update): keep nonzero exits and in-flight runtime removal ahead of the watchdog
Preserve a pending nonzero exit when the output watchdog records success. Once retained-runtime directory removal starts, keep it joined through success or error; worker and preparation stalls before removal retain the bounded watchdog behavior. Document the installed-updater behavior.
Real-child regressions failed before the fix for pending exit precedence, interrupted removal, and removal error settlement. Validation: 80 tests across retained-runtime exit/runtime, signal-exit-barrier, and one-shot-exit passed with maxWorkers=2 in 72.86s. The 14-case exit file passed with maxWorkers=1 in 15.73s wall. Core, infra-test, and config-cli-test typechecks, targeted oxlint, oxfmt, and git diff --check passed. Independent P2 autoreview was scoped-clean.
* refactor(update): drop the retained-runtime exit watchdog pending broker settlement split
Remove the watchdog finalizer escape, retirement abort/removal flags, immediate warning path, recorded exit-code state, and watchdog-only process tests. Retained SQLite work and runtime removal stay joined through the existing settlement owner. A mandatory broker settlement versus disposable termination split remains a separate follow-up.
Keep Doctor legacy pnpm projection reclamation and service TMP/TEMP/TMPDIR discovery, including Windows casing and drive-relative paths. Main already preserves failed barriers and pending nonzero exits without watchdog state; retain that implementation and cover those outcomes directly in signal-exit-barrier tests. Recovery docs describe the reduced scope.
Validation: 101 tests passed across six files with maxWorkers=2 (101.54 seconds); core typecheck and state-logging, config-cli, services, and other core test typecheck lanes passed; targeted oxlint and oxfmt passed; Madge found zero cycles; git diff --check passed. Signal barrier cost with maxWorkers=1: seven tests passed, 3.28 seconds command wall (1.74 seconds Vitest). Autoreview against the merged main base was scoped-clean at P1. Native Windows and published-driver upgrade proof were not run in this scoped reduction.
---------
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
* fix(update): preserve retained rollback ownership (WIP)
Forward-port rollback-ownership and runtime-restoration corrections from 008d0c4c05. Preserve operator edits and changed branch refs, and record completed runtime restoration before rechecking authority. Validation remains incomplete under host load; do not publish.
* refactor(update): separate rollback steps and transaction fixtures
Keep retained Git rollback ownership and runtime restore replay fixes within existing file limits. Move rollback commands to the Git step owner and transaction regressions to sibling test support; preserve the original regression coverage. Completes the implementation extraction from 7d8c28fa1e; validation and review are recorded in the PR.
* fix(update): preserve staged content during retained rollback
Replace retained reset --keep with a non-forcing detached checkout, conditional branch-ref update, and reattachment. The old reset discarded staged content in unchanged files when the working copy contained a different edit. A real-Git regression reproduced the index losing the staged bytes before this fix.
* fix(update): preserve branches claimed by linked worktrees
* fix(update): use Git's own branch custody check before rollback ref changes
Linked worktrees in a paused rebase or bisect detach HEAD, so worktree porcelain misses their branch reservation. Probe with a no-op branch reset, which Git refuses for every worktree owner state.
* fix(update): keep branch custody through retained rollback ref writes
Keep HEAD attached during retained source restoration and verify the reflog transition before restoring the runtime.
Restore claimed dev refs with create-only compensation after CAS deletion. Preserve operator edits and report explicit recovery steps.
* fix(update): keep the created dev branch on retained rollback
Leave the created dev ref intact during retained rollback and report a quoted manual cleanup hint while restoring the original checkout and runtime.
Remove post-delete custody scans and compensation, which cannot exclude in-flight claims or safely repair after authority revocation. Preserve the existing rewrite and non-retained cleanup paths.
* fix(update): keep newer branch writes out of rollback recovery guidance
Require an existing, readable branch reflog before retained rollback rewrites
source. If the ref changes after checkout, retain both runtimes and report the
expected, current, and reflog-previous commits without a forced-ref command.
Only offer conditional restoration of an overwritten earlier write while the
branch still points to the rollback commit.
Extend the real-Git transaction table with the post-checkout writer and
pre-mutation reflog refusal, and document the installed-updater behavior.
Both regressions fail against the previous implementation; the existing
pre-checkout writer recovery remains covered.
Validation: 101 candidate/runtime tests pass (176.9s wall); core and infra-test
typechecks, targeted oxlint, oxfmt, and git diff --check pass. The three focused
real-Git cases take 6.4s test time on one worker. Isolated P2 review is clean.
* fix(daemon): keep a supported recorded Bun on implicit service reinstall
Forced reinstall already retained a supported Node recorded in the managed
service, but an unpinned service recorded on Bun fell back to Node. On a
Bun-only host that made every implicit reinstall, including the updater's
own service refresh, fail; with Node present it silently moved the service
to Node.
Implicit `gateway install` and `node install` now retain a supported recorded
Bun or Node executable. Interactive runtime pickers (configure, Doctor's
reinstall branch, onboarding) default to the recorded runtime and use its
path only when it stays selected; flows without a picker retain it directly.
Explicit --runtime/--runtime-path, pins, wrappers, and the fallback for a
missing or unsupported recorded runtime are unchanged. No pin is created.
* fix(daemon): use the running Bun when a Bun-only host has no Node
Fall back only after implicit Node discovery finds no supported executable, and require the existing capability probe to approve the running Bun. Carry explicit runtime intent through shared service planning, retain pins/wrappers/recorded runtimes, and never persist the fallback as a pin.
Validation: 250 distinct focused tests passed across nine files (164 tests in 39.70s; 100 in 20.85s, with overlapping regression coverage). The registered install regression failed before the fix. Both tsgo:core and tsgo:test:root, changed-file oxfmt, git diff --check, and scoped P2 Codex autoreview passed. No CI reruns were requested.
* fix(daemon): keep an interactive runtime choice ahead of the Bun-only fallback
Compute the setup suggestion before configure, Doctor, or advanced onboarding
prompts, using the shared Bun-only discovery policy. Treat picker results and
pins as explicit runtime intent so selecting unavailable Node fails before
service replacement. Retain recorded runtime paths, preserve quickstart's
implicit fallback, and keep automatic choices unpinned.
Doctor uses the same suggestion for its non-interactive runtime fallback.
Service-install consent and update deferral remain unchanged. No state or
runtime-pin migration is needed.
Validation: the configure regression reached installation before the fix and
now fails with the existing unavailable-Node diagnostic before install.
164 tests passed across six focused files in 218.77s wall, including cold
compiler preparation; the commands shard passed 85 tests in 22.04s and the
onboarding shard passed 70 tests in 12.36s. tsgo:core (18.83s),
tsgo:test:root (67.31s), changed-file oxfmt, git diff --check, and scoped P2
Codex autoreview passed. No CI runs were dispatched or retried.
Addresses the interactive runtime selection finding on #160189.
* fix(gateway): report the effective QuickStart service runtime
* fix(wizard): translate the QuickStart Bun runtime note