Commit graph

10786 commits

Author SHA1 Message Date
Peter Steinberger
1470a9bb09 fix(release): record npm's bundled dependencies in the npm lock report
Since #160224 the root package depends on npm, whose v3 lock lists bundled packages under node_modules/npm/node_modules/* with inBundle and no resolved/integrity. The dependency release evidence job rejected them as unsupported lock entries. Bundled entries are now accepted only when their nearest non-bundled carrier has a verified registry tarball, and each report entry records them as bundledDependencies with that parent. Ordinary entries without resolved/integrity stay fatal.
2026-09-28 21:29:01 -07:00
Peter Steinberger
20d728a72c
perf(test): stripe changed-plan source scans across worker threads (#160773)
The changed-test planner reads the whole ~42k-file tooling inventory in
its native source-facts child whenever a wide consumer query or full
import graph is needed, and tokenized it on one thread. After main stopped
parsing unmatched files in term queries, the flaky PR-exempt ownership
case got cheaper, but the next integration case still pays the full
single-threaded graph parse.

Large scan requests now stripe files across up to eight worker threads
that reuse the scanner module, keep matchingOnly semantics, and return
rows to their request slots after every worker has joined. A regression
test also pins that a wide consumer query leaves unmatched files for the
later full-graph walk.
2026-09-28 20:30:08 -07:00
Peter Steinberger
fb5f1954a2
perf(sqlite): prepare NOCOW stores and add offline btrfs repair (#160877)
* perf(sqlite): prepare NOCOW stores and add offline btrfs repair

* fix(sqlite): complete NOCOW integrity and tooling integration

* test(gateway): leave terminal fixture cleanup with its suite owner
2026-09-29 03:25:36 +00:00
Peter Steinberger
59490dc752
fix(pr): settle mergeability after verified main advances
A complete prior-CI REST observation can report UNKNOWN while GitHub
recalculates mergeability after main moves. Permit a bounded three-read
settlement window only after proving that forward transition. Keep the
original known projections and all identity, policy, and check facts pinned;
unknown or changed known projections still cannot authorize dispatch.

Track the latest validated main separately from the retained intent anchor,
so repeated same-main UNKNOWN and rollback to another descendant of the old
anchor remain refusals. Preserve local-only final reads and live authority
validation after the complete REST work.

Cover the actual native admission path at later stability and final authority,
plus exhausted waits, identity/policy drift, conflicts, rewind/divergence,
missing final objects, revoked authority, and retained-outcome fencing.
2026-09-28 20:17:16 -07:00
Peter Steinberger
0eb4ad70c4 fix(e2e): size update lane budgets from hosted 4-vCPU measurements
Hosted run 36506342273 exhausted the first-hop container's 1800s budget.
Project 1558s + 560s final candidate hop + ~5s assertions ~= 2125s;
x ~1.5 gives 3200s inner, plus 300s host margin gives 3500s per lane.
Six sources still need two waves at npm weight limit 5; 2 x 3500s
plus 10m setup ~= 127m, rounded to a 130m self-upgrade chunk.

Restart-auth exceeded 1515s in run 36506342273: x ~1.5 gives 2280s
inner plus 300s host margin, or 2580s per lane. Run 36506210440
measured an 856s restart update; 4-CPU Crabbox measured 993s, so
set a lane-only 1500s command timeout and retain the global 900s default.

The OpenAI/recovery row's five weight-3 npm lanes serialize at limit 5:
30 + 30 + 20 + 25 + 43 = 148m, plus 10m setup => 160m rounded.
The chat lane overlaps. CI Docker seed excludes restart-auth and keeps
its 60m job budget. No product code or scheduler concurrency changes.

Validation: all 185 tests in the three focused files passed (71.13s wall,
one worker). Bash syntax, oxfmt, oxlint, and git diff --check passed.
Independent Codex autoreview found no actionable P0/P1 findings.
2026-09-28 20:03:19 -07:00
Peter Steinberger
58b1860232
test(core,codex,ui): remove low-value tests (batch d092) (#160879)
* test(doctor): deslop t0131 tests

* test(ui): deslop t0124 tests

* test(logging): deslop t0134 tests

* test(outbound): deslop t0094 tests

* test(config): deslop t0135 tests

* test(status): deslop t0111 tests

* test(config): deslop t0137 tests

* test(acp): deslop t0146 tests

* test(codex): deslop t0141 tests

* test(config): deslop t0136 tests

* test(ci): align cache warm seed after status test cleanup

* test(ci): declare cache warm fixture configs explicitly
2026-09-29 02:39:55 +00:00
Peter Steinberger
ae6a88438d
perf(sessions): keep cold snapshots out of metadata reads (#160358)
* perf(sessions): separate hot facts from cold snapshots

Move saved skill, prompt-report, and diff snapshots into keyed agent-schema-24 rows. Keep full-entry compatibility through same-statement hydration, preserve revision fencing and Doctor/lifecycle ownership, and return immutable borrowed metadata without redundant clones. Metadata fixture list allocation falls about 96%; generic patch cost remains essentially unchanged.

* test(sessions): adapt storage fixtures to split snapshots

* test(sessions): read split snapshots in upgrade assertions

* ci: include main UI type shard rebalance

Apply main commit 52ae7e314d unchanged to unblock the inherited ui-other root-budget failure. Session storage and the measured product code are unchanged.

* fix(sessions): complete cold snapshot reader cutover

Hydrate bounded full-entry summaries through the canonical snapshot query, use prepared Kysely migration DML, and avoid repeated JSON bindings in snapshot upserts. Keep allocation and consistency limits intact while moving existing fixtures to split storage.

Includes the unchanged Doctor e2e mock repair from main 162dea17b0. Testbox tbx_01m3m1gy44kbw82fjdeb0y98fh passed architecture, focused lint and types, 242 focused tests, and the allocation benchmark. Fresh review through P2 was clean.

* test(sessions): synchronize recovery and heartbeat fixtures

Hold deferred recovery repair at its scheduler boundary while the real successor claims and releases ownership. The scheduler has its own retained execution proof, so teardown no longer drains unrelated timers during native SQLite work.

Wait for heartbeat delivery or its actual settlement instead of imposing a separate arrival timer. Preserve awareness/reset assertions and join admitted delivery after releasing the hook.

Original constrained shard replays passed and do not establish the exact CI trigger. Candidate proof passed 572 shard tests plus both scheduler-owner cases, focused lint/types and independent review through P2.

* test(ui): await startup request before coalescing assertion
2026-09-29 02:39:21 +00:00
Peter Steinberger
de86805f1e
fix: retain upgrade-survivor repair diagnostics and deadline coverage (#160896) 2026-09-29 02:37:19 +00:00
Josh Avant
35263bd17a
fix(ci): stop obsolete security reviews after PR lifecycle changes (#160890)
* fix(ci): stop obsolete security reviews after PR lifecycle changes

* fix(ci): preserve security review evidence after merges
2026-09-28 21:34:29 -05:00
Peter Steinberger
039b32a68a
perf(crabbox): reuse source mirrors across commits
Keep the derived mirror cache keyed to its retained Git directory and ref,
while each command continues sealing the full current commit witness.
Same-ref commits now preserve unchanged source files and warm indexes.
Existing full-witness cache entries rebuild once under the new key.

Regression proof covers empty and changed-source commits, byte-identical
cold capsules, ref/Git-directory changes, and mid-freeze revision changes.
The capsule owner suite, context controls, changed checks, and independent
review passed. Existing byte, policy, custody, and cleanup checks remain.
2026-09-28 19:14:07 -07:00
Peter Steinberger
9148fbb46a
fix(nodes): explain and recover session-host setup problems (#160188)
* fix(nodes): explain and recover session-host setup problems

Report unsafe workspace ancestry before dispatch, resume pending node pairing
after approval, explain runtime command availability at its authority owner,
and log inventory publication failures only when they change. Share inventory
validation and remove superseded policy and projection helpers.

* test(nodes): assert specific dispatch remediation messages

* test(nodes): align pairing retry coverage with node recovery

* fix(nodes): preserve desktop access after hosting failures

Keep connected-node environment availability separate from session hosting readiness and retain diagnostic placement refusal. Include Codex plugin installation in missing-command remediation. Cover the inventory-to-environment boundary with real connected-node enumeration.

* refactor(nodes): share runner declaration validation

Keep the status-wait capability added on main while extending the existing strict record schema with bounded hosting diagnostics. Remove the superseded declaration parser and clone closed worker-host snapshots through one path. Production code remains net zero against the refreshed base.

* fix(nodes): retain exhaustive command diagnostic states

* refactor(nodes): preserve explicit diagnostic return flow
2026-09-28 19:13:09 -07:00
Peter Steinberger
e187219fa2
fix(node): headless nodes with default plugins never activate automatic updates (#160790)
* fix(file-transfer): declare node commands idle between invokes

A default headless node host could never pass the auto-update idle
barrier: file.stat, file.create, workspace.memory and workspace.skills
registered without hasActiveWork, and the node host treats a missing
idle hook as active work. Their work is invoke-scoped and already
counted by the runtime's in-flight invoke tracking.

* test(e2e): prove default-plugin nodes activate node auto-updates

* test(e2e): read the node-host subsystem console prefix

* docs(install): explain recovering 2026.9.6 nodes stuck before automatic updates
2026-09-28 19:12:09 -07:00
Peter Steinberger
73182e751e fix(ci): add config-env-values to the PR wrapper inventory
PR #160843 moved the config env allowlist into src/config/config-env-values.ts,
which src/config/config-env-vars.ts re-exports, but scripts/pr-lib/wrapper-components.txt
was not updated. The materialized origin/main trust anchor therefore fails every
scripts/pr review-init with ERR_MODULE_NOT_FOUND, and
test/scripts/eager-import-closure.test.ts is red on main (5 of 16 cases).
Adding the inventory line restores the closure; the test passes 16/16.
2026-09-28 19:11:07 -07:00
Peter Steinberger
d5447648d4
fix(pr): validate main movement within prior-CI REST reads
Route the beginning and end of an active prior-CI REST observation through
the existing forward-main ancestry and merge-composition owner. Recheck
branch policy without replacing the original check snapshot, and preserve
strict main stability for ordinary REST admission and inactive recovery.

Verify the original main anchor and both endpoints locally after final
authority validation, suppress lazy fetch throughout composition, and keep
transient read bounds out of retained outcomes. Report endpoint and timing
facts without claiming an unobserved GitHub transport route.

Prove the former pre-dispatch refusal and valid forward movement, with
negative controls for rewind, divergence, conflict, empty change, policy
and authority drift, missing local endpoints, and uncertain outcomes.
2026-09-28 19:04:03 -07:00
Peter Steinberger
2fa29e2099
fix(release): keep stable publication handoffs moving (#160859)
* fix(release): keep stable publication handoffs moving

The Full Release Validation helper gave up with "discovery exhausted" while
the exact parent run sat in the runner queue before uploading its input
witness (2026.9.7 FRV 36466521174 waited 52 minutes for its first job). It now
keeps polling while that run is still active, bounded at three hours; a
completed run without a witness still fails closed.

release:stable now flips the GitHub release to Latest before the beta
dist-tag sync, so GitHub goes public as soon as core npm is visible.

The release skills record the 2026.9.6 publication lessons: early GitHub
activation as standing stable policy (finalize_release_before_docker=true on
the direct route), detached ClawHub child monitoring and per-package recovery,
bootstrap approvals, the VCR mirror follow-up dispatch, macOS preflight
ordering against the update compatibility inventory, and recording the shipped
tarball first at closeout.

* fix(release): emit early GitHub activation in stable publish commands

The candidate and publish-preflight helpers now add
finalize_release_before_docker=true for final versions on npm latest, so the
printed direct dispatch activates GitHub right after npm verification. Beta,
extended-stable and prepared routes are unchanged. release:stable strips the
input because its flip-github phase activates the release itself and must not
leave Docker waiting on the extra activation gate.
2026-09-28 19:02:35 -07:00
Peter Steinberger
7d0c9758f1
ci: scope retention sweep to planner changes
Move the full-repository PR-exempt retention sweep out of protected planner
integration coverage and into the existing PR-exempt tier. Explicit policy
watches and direct edits opt it into relevant PRs; hourly main and full
release verification retain its canonical tooling owner.

Preserve the landed case body and every assertion, including the total-row
cap check. Keep the other planner regressions protected and extend the
existing source-owner policy table to cover the new opt-in paths.

The moved case passes in 26.90s on two CPUs, and the remaining integration
file passes all nine cases. No timeout, packer, routing, or shared-helper
changes. P2 review is clean.

Validation: the complete 884-file tooling inventory passed on Linux
Testbox: 26,693 passing cases and 475 existing skips. Two completed file
reports were retained after an interrupted lease; the remaining 882 files
passed with an exact inventory reconciliation and source verification.
Full test-types, changed-file checks, boundary lint, and formatting passed.
2026-09-28 18:50:53 -07:00
Dallin Romney
0edcf447a5
fix(release): extend npm readback timeout (#160796) 2026-09-28 18:29:38 -07:00
Peter Steinberger
91149044f4
test(agents,ui,tooling): remove low-value tests (batch d090) (#160776)
* test(ui): deslop t0107 tests

* test(qa-lab): deslop t0091 tests

* test(agents): deslop t0114 tests

* test(code-mode): deslop t0100 tests

* test(status): deslop t0110 tests

* test(agents): deslop t0115 tests

* test(cron): deslop t0117 tests

* test(scripts): deslop t0123 tests

Remove redundant Docker harness source inventories and repeated wrapper
matrices. Keep independently repaired shell/process behaviors, release
boundaries, cleanup ownership and private diagnostic publication checks.
No production or shared test-support changes.

Tests: 347 -> 150 (-197, 56.8%); both Testbox runs passed.
Test LOC: 9113 -> 5981 (-3132, 34.4%).
Vitest duration: 84.48s -> 61.17s on the same Blacksmith Testbox lease.
Coverage: statements 90 -> 90; branches 69 -> 69; functions 27 -> 27.
No per-file drops or excluded config groups.
Formatting, diff-check and plain oxlint pass. Changed-file check: passed on Blacksmith Testbox after replacing an expired
lease; all root tsgo shards, typed lint, dead-export and boundary checks pass.
The expired-lease attempt executed no checker payload.

The 50% test-count target is met. The LOC target is not: remaining fixtures
exercise distinct past regressions and security boundaries. The per-test
pass-3 retention exceptions below use R2 for demonstrated historical fixes
and R3 for security or destructive-ownership boundaries. Table descriptions
name retained distinct cases. Coordinator review owns acceptance of the LOC
shortfall.

isolates helper snippets from shell hooks at SHLVL=%s | R2 58dce76576: inherited startup/logout hooks executed in fixture shells; rows SHLVL=0 and 1 cover top-level and nested shell invocation.
treats Docker registry auth 5xx failures as transient build failures | R2 f29248fa62: Docker OAuth registry 5xx responses were not recognized as retryable.
detects compiler processes killed by the OOM killer | R2 2bb6b870aa: compiler killed-signal diagnostics were omitted from OOM recognition.
retries Corepack connect timeouts without misreading Dockerfile comments as OOM | R2 2bb6b870aa: Corepack connect errors were missed while generic Killed comments falsely matched OOM.
routes standalone Docker smoke runs through the timeout-aware helper | R2 f1ceed94db: cleanup/install smoke Docker runs bypassed the timeout owner and could hang.
runs the sandbox browser sidecar proof from the package-installed image | R2 2c393ecf55: static SDK import captured process-stable roots before fixture HOME/state/config initialization.
cleans all sidecar modes without touching another run on the same Gateway workspace | R2 2c393ecf55; R3: cleanup labels lacked per-run ownership and could delete another run's containers.
gives cleanup-smoke builds enough Node heap while preserving explicit callers | R2 e3bab80bda: cleanup builds lacked heap headroom; repair supplies a default while preserving caller heap settings.
rejects invalid cleanup-smoke log byte limits | R2 a7b52ecad9: malformed limits silently defaulted rather than failing before setup.
normalizes zero-padded cleanup-smoke log byte limits | R2 901f963f62: bounded-tail assertions retain the repair replacing unbounded cleanup failure logs; a7b52ecad9 retained decimal normalization.
prints Docker MCP client logs through the bounded helper | R2 5d7e0b73a7: MCP runner success/failure paths printed entire client logs instead of bounded tails.
prints in-container Docker client logs through bounded helpers | R2 cdbf6d95ac: media and chat-tools scenarios independently bypassed bounded log output.
runs cleanup smoke on the native ARM platform instead of pulling an amd64 tag | R2 d3ab7e92ef: cleanup smoke hardcoded amd64 on ARM hosts.
lets Testbox fall back to building when a reused Docker image is missing | R2 5c591a4e13: missing/pull-failed images stopped Testbox and the live CLI runner skipped the fallback owner.
resolves source and compiled candidate test-state entrypoints | R2 88fc335323: test-state invocation hardcoded a TypeScript path and failed for compiled frozen candidates.
runs current TypeScript and frozen JavaScript Docker harness entrypoints | R2 1fb715853c: package scenarios selected incompatible source/compiled entrypoint runtimes.
rejects malformed Docker E2E resource limits before a suite starts | R2 0df60ad306: malformed resource knobs were passed through Docker setup rather than rejected at the owner.
keeps Testbox image-build fallback before isolating live MCP code-mode runtime flags | R2 e84b719c99: OPENCLAW_TESTBOX was unset before image fallback consumed it.
wraps centralized Docker builds with the timeout helper | R2 6ef0cbb94f: centralized image builds invoked Docker without a timeout.
stops the tracked build command without retrying when interrupted | R2 d31f4e2d62: interrupted builds left tracked work running or entered retry handling; TERM and INT retain distinct exit semantics.
normalizes zero-padded centralized Docker build heartbeat intervals | R2 17795c6c4c: raw 08 reached Bash arithmetic rather than normalized decimal.
normalizes zero-padded centralized Docker build retry counts | R2 37eea55afa: raw zero-padded retry counts reached Bash arithmetic.
rejects invalid centralized Docker build %s before invoking docker | R2 37eea55afa: rows retry count/2x and heartbeat interval/soon cover separate owners that silently defaulted instead of rejecting before Docker invocation.
fails centralized Docker builds fast when timeout is unavailable | R2 6ef0cbb94f: required-bound build entrypoints otherwise executed Docker unbounded.
keeps reused Docker image probes behind the timeout-aware helper | R2 55af31e0c6: image inspection/pull probes bypassed command timeouts.
explains how to opt out when Docker rejects default resource limits | R2 ff9b291673: unsupported cgroup limits lacked actionable diagnostics; bounded FIFO capture and cleanup preserve the repair.
rejects invalid Docker run pids limits before invoking docker | R2 6af1b97b1d: malformed PIDs limits reached Docker rather than failing at admission.
removes functional Docker build package inputs after the build | R2 abc7b7b331: generated tarball and BuildKit context directories leaked after image builds.
keeps caller-provided functional Docker build packages | R3: generated-input cleanup must not delete a caller-owned tarball.
cleans generated package mounts after harness Docker runs | R2 b377618fae; R3: generated mounted packages leaked; failure inspection precedes container cleanup and caller-owned packages survive.
propagates shared E2E command timeouts into package-backed containers | R2 d0cb7ba55b: host command timeout overrides did not cross the Docker environment boundary.
cleans the heartbeat command when the wrapper is terminated | R2 6bfd47af38: terminating the logger wrapper left its active child running.
cleans harness containers when heartbeat-wrapped Docker runs are terminated | R2 6bfd47af38: composed logger/harness interruption leaked owned containers and cid directories.
normalizes zero-padded Docker E2E stats heartbeat intervals | R2 2102166f86: stats-loop arithmetic interpreted zero-padded intervals incorrectly.
derives the browser CDP image from the shared functional image | R2 cdb2fa35e2: shared-image overrides reused an image lacking Chromium instead of building the required derived image.
fails fast on invalid browser CDP snapshot byte limits | R2 1fb11ab306: runner ignored and failed to validate the byte-limit setting.
forwards browser CDP snapshot byte limits into the Docker runner | R2 1fb11ab306: valid host byte-limit settings were not forwarded into the container.
uses Playwright Chromium for the browser CDP snapshot image | R2 fd2e4da006: distro Chromium/manual startup was incompatible with the browser smoke's managed runtime path.
opens the browser CDP fixture before snapshotting | R2 fd2e4da006: doctor ran before browser open and raced readiness.
fails Docker commands fast when timeout is unavailable | R2 3736d7b60b: missing timeout implementations silently executed commands without a bound.
preserves $helper status $commandStatus in $mode after exit $entryStatus | R2 ec66133ba0: bare returns restored the EXIT trap status; rows docker_e2e_docker_cmd/normal, docker_e2e_docker_run_cmd/normal, docker_e2e_docker_cmd/no-diagnostics, docker_e2e_docker_cmd/node-watchdog each preserve status 43 after exit 0 across independent repaired returns; platform shell matrix retained.
uses a Node watchdog for Docker commands when timeout is unavailable | R2 10056c9346: Node-capable hosts without GNU timeout could not run Docker; fallback preserves bounded execution, stdin and child status.
adds default Docker run resource limits without overriding explicit limits | R2 a372429a96: Docker harness containers lacked resource bounds; explicit caller constraints must remain authoritative.
escalates Docker watchdog children that ignore parent SIG${shellSignal} | R2 32f98d7fe8: TERM/143 and HUP/129 rows cover escalation of signal-ignoring children and previously absent HUP forwarding.
uses gtimeout when timeout is unavailable | R2 3736d7b60b: gtimeout-capable hosts were treated as lacking a timeout implementation.
passes plugin lifecycle sampler timeout overrides into Docker | R2 71cb60706b: phase/grace timeout overrides did not cross the Docker boundary.
rejects invalid plugin lifecycle Docker %s overrides before package setup | R2 7207072436: phase timeout/150ms and CPU ratio/0 rows cover separate integer and positive-decimal validation failures before package setup.
wraps direct Docker E2E npm installs with the shared timeout helper | R2 9777526eaa: actual scenario npm installs bypassed phase timeouts.
keeps upgrade survivor mutable state off the host-mounted artifact tree | R2 9fef53c3b1: runtime state/cache/temp paths incorrectly used host artifact mounts rather than container-local roots.
starts the upgrade survivor plugin registry before updates with scenario-owned config | R2 b729806758 and e81d7e62e8: ClawHub setup followed restart-environment capture and synthetic plugin credentials were globally injected rather than scenario-scoped.
keeps upgrade survivor wrappers and the embedded payload valid bash | R2 33a5f37937: nested single-quoted assertion code corrupted embedded shell payload syntax.
wraps package-backed scenario OpenClaw CLI calls with the shared timeout helper | R2 d0cb7ba55b: package CLI scenario calls could hang without command bounds.
preserves actionable, secret-safe typed onboarding failure diagnostics | R2 60aac74672; R3: diagnostics exposed auth/config data and gateway tokens, while fd cleanup permanently redirected stderr.
propagates frozen typed-onboarding ERR traps through nested helpers | R2 2d25402ecb: frozen runner lacked bash -E and suppressed nested failure diagnostics.
prints channel-add failures through the shared E2E logger | R2 f25f7429df; R3: channel-add raw-file diagnostics bypassed the shared redacting logger.
keeps append-only mock E2E state under per-run scratch roots | R2 dcf21ac3ad: shared append-only request/state files contaminated subsequent runs.

kills timed Docker scenario runners after the grace period | R2 d5bf325126: forced termination after the timeout grace prevents timed-out scenario processes from surviving.
propagates HTTP probe failures through command substitution | R2 56ffe5bc2e: explicit return preserves probe failure because command substitution does not inherit errexit; success and failure exercise different outcomes of this caller.
records an interrupted upgrade survivor phase as failed | R2 9de3ca5fc9: signal handling and completion accounting prevent interrupted success; separate summary output avoids corrupting the interrupted command artifact.
keeps multi-node update Docker artifacts isolated by default | R2 d5df1a1cd6: per-run directories prevent separate multi-node runs from overwriting shared default artifacts.
reuses the shared bare image for multi-node update targeted runs | R2 68ead4dd80: production replaced the hardcoded image and corrected skip-build routing so a targeted run actually uses the requested shared image.
bounds upgrade survivor foreground OpenClaw CLI calls | R2 c965b3a1ae: timeout wrappers at actual foreground CLI callers prevent hung upgrade checks; shared timeout-helper tests cannot protect missing caller wrappers.
starts the %s auth probe under the manager that owns its restart and stop | R2 1c33687051: published and current rows cover separate preparation entry points; both formerly adopted foreground CLI processes whose respawn descendants escaped service restart and stop custody.
returns the gateway readiness failure when startup is called conditionally | R2 1c33687051: run.sh:start_gateway must explicitly return a failed readiness status under conditional invocation; the current-install preparation table calls a different owner.
scopes candidate setup Doctor markers without creating legacy device identities | R2 30aa2794d9 and f584661fac: Doctor-only migration markers preserve plugin convergence, and removing legacy identity seeding prevents replacement of the canonical device identity.
restores the canonical authored config after %s failure | R2 30aa2794d9: doctor=41, readiness=42, service-env=43, and install=44 cover distinct failure stages whose original status and authored config must survive preparation.
prefers restore failure and retains the authored config snapshot | R2 30aa2794d9: the preparation caller prioritizes restore failure over an earlier Doctor failure and retains the recovery snapshot; the sibling helper test does not cover caller precedence.
keeps upgrade survivor auto-auth success summary set -u safe | R2 0f67474251: default startup_summary prevents nounset failure when auto-auth intentionally skips the start_seconds assignment.
stops promptly when the systemctl target is a zombie with spaces and parentheses in comm | R2 1ecd53fe27: validated proc-stat tail parsing replaces incorrect whitespace-field parsing for zombie process names.
waits for a killable systemctl target when proc stat is unreadable or malformed | R2 1ecd53fe27, adapted by 99e7d3e7a5: unreadable and malformed proc-stat inputs cannot be treated as confirmed exit or release a still-killable process from stop custody.
records delegated post-core systemd callers only with the environment marker | R2 dfd4cc8a5d: recognizes delegated post-core workers while preserving ordinary update attribution and rejecting missing or unreadable delegation markers.
retains the original post-core %s result separately from exit %i via %s | R3: public projection removes secret and unapproved result fields; error/0/update proves result versus process-exit independence, and warning/0/--post-core also protects delegated-worker recognition fixed by dfd4cc8a5d.
leaves post-core capture unavailable for %s without changing the child outcome | R3: doctor and worker restrict eligible observation process/thread; missing-context and delegated-missing-context require the marker for both entry forms; wrong-file and outside-tmp enforce the named temporary result boundary; symlink and hardlink enforce file ownership; oversize enforces bounded private input; blocked-output rejects an unsafe observation destination.
scopes the passive preload to the original update invocation in %s | R3: run.sh and upgrade-survivor-docker.sh are distinct update callers; private instrumentation and observation roots stay scoped to the selected update and its child, without leaking into later CLI invocations.
requires an update-owned service replacement after consent recovery (%s) | R2 56ffe5bc2e: pid-only rejects replacement without the updater restart request, request-only rejects a restart record without replacement, and replaced accepts both; stale earlier restart evidence cannot claim recovery success.
observes or explicitly omits persisted plugin identity without changing source artifacts (%s) | R3: sqlite and historical exercise separate persisted-index reader routes while publishing only selected identity metadata, excluding private config/package fields, avoiding plugin-code execution, and preserving source artifacts.
refuses unsafe plugin identity input (%s) | R3: index-symlink and index-hardlink protect persisted-index ownership; index-oversize and index-cap bound private index bytes and records; root-symlink rejects an escaped plugin root; doctor-symlink and doctor-hardlink protect selected Doctor-artifact ownership; doctor-oversize bounds private artifact reads.
retains a failed service child and only sanitized diagnostics (candidate redactor: %s) | R3: true supplies hostile installed candidate redactor code that must not execute; frozen host redaction, private-file exclusion, allowlisted projection, and secret removal protect publication.
retains supervisor bootstrap stderr without inventing a child exit in %s | R2 b85a049da5: update-restart-auth.sh previously discarded bootstrap stderr and reported fabricated zero child-exit values without an exit observation.
collects before cleanup without masking failure or breaking success in %s | R2 b85a049da5: run.sh and upgrade-survivor-docker.sh have separate exit handlers; capture precedes cleanup so shutdown cannot replace initial failure evidence, and capture failure cannot replace the original result; rows cover failure with capture, failure without capture, successful completion without capture, and incomplete zero exit.
publishes only on the host and preserves Docker outcomes (published baseline: %s) | R3: false and true cover current and published Docker routes; public destinations never enter candidate mounts, private/stale snapshots are excluded, and published summaries and optional logs are projected and redacted without reusing prior-run evidence.
stops supervised gateway restarts after the systemd burst limit | R2 fb0812c857: systemd-like supervision bounds rapid crash restarts instead of allowing unbounded fixture restart behavior.
terminates supervised gateway descendants at the systemd stop timeout | R2 fb0812c857: process-group drain replaces leader-only stopping; the graceful parent may exit while a term-ignoring descendant still requires termination.
drains the previous gateway process group before restarting | R2 fb0812c857: the replacement starts only after the previous process group has settled, preventing descendant overlap across restart.
rejects invalid upgrade survivor Docker %s before Docker setup | R2 cc3d346c15 and 8c3185d55c: start budget=90s exercises malformed positive integer, probe timeout=soon exercises malformed nonnegative integer, and probe attempt timeout=0 exercises the positive-only zero rejection before Docker setup.
bounds upgrade survivor failure log diagnostics | R2 8cba5f7efd: bounded logger calls replace unbounded cat operations at actual failure callers; helper-level truncation tests cannot detect callers bypassing the bounded logger.

R2 a3d5e5bc72 | preserves caller-owned file descriptors around harness runs | Bash 3-compatible stdin descriptor allocation must preserve an already-open caller descriptor.
R2 b377618fae | cleans release package runner $label on $scenario [label=release-upgrade-user-journey, scenario=generated-success] | Generated package mounts previously leaked; exercise owned tarball cleanup through the real runner while retaining run evidence.
R3 | cleans release package runner $label on $scenario [label=release-user-journey, scenario=provided-run-failure] | Destructive cleanup must preserve a caller-provided, unmarked tarball even when the container fails.
R3 | cleans release package runner $label on $scenario [label=release-media-memory, scenario=marked-other-name] | A generated-package marker must not authorize deletion of a differently named caller artifact.
R2 113af2fcbd | preserves failing heredoc output and status through Docker E2E heartbeat logging | Background heartbeat execution lost stdin before explicit descriptor forwarding; retain payload and failing command status.
R2 45d173a3ee | copies the pnpm lockfile into the runtime image before normalizing its permissions | Runtime Docker stage attempted chmod before the lockfile was copied.
R2 1421c21640 | verifies fs-safe through a pnpm-style linked package root | Linked pnpm package roots require resolving the physical package location before loading fs-safe.
R2 1421c21640 | builds and cleans package-lane images without touching shared image tags | Package-install proof must build and retire its own bare/musl tags, preserving shared functional images.
R2 dfa1a51225 | cleans the direct binding runner after %s [scenario=invalid summary] | An unfocused Vitest run must fail the wrapper instead of being accepted as the intended three-test smoke.
R2 6bfd47af38 | cleans the actual cron runner and its harness on %s [signal=SIGINT, status=130] | Interrupted harness cleanup must terminate the Docker child and retire container/log resources through the INT trap.
R2 6bfd47af38 | cleans the actual cron runner and its harness on %s [signal=SIGTERM, status=143] | Interrupted harness cleanup must terminate the Docker child and retire container/log resources through the TERM trap.
R2 6bfd47af38 | cleans the actual cron runner and its harness on %s [signal=SIGHUP, status=129] | Interrupted harness cleanup must terminate the Docker child and retire container/log resources through the HUP trap.
R2 ffb4f0d5c3 | stages the committed $layout bundle-MCP client through the real runner: $scenario [layout=June, scenario=success] | Preserve the shipped June client's import depth, committed bytes and ESM scope instead of executing dirty checkout decoys.
R3 | stages the committed $layout bundle-MCP client through the real runner: $scenario [layout=June, scenario=empty extraction] | Missing staged regular files must prevent selected-client Docker execution.
R3 | stages the committed $layout bundle-MCP client through the real runner: $scenario [layout=June, scenario=altered extraction] | Extracted client bytes must match the authorized selected commit before Docker execution.
R2 ffb4f0d5c3 | stages the committed $layout bundle-MCP client through the real runner: $scenario [layout=July, scenario=success] | Preserve the independently shipped July client's deeper import layout and selected-source module scope.
R3 | passes source-qualified overrides without leaking frozen control-plane identity | Containers receive resolved compatibility facts without selected/tooling authorization identities; boundary established by 07fe26ecf8.
R2 fc2724d831 | copies the complete bun harness closure into the package-install lane | Protect the missing dependency closure that broke bun package proof after #129552.
R3 | proves gateway suspension across a same-container process restart | Capability output remains host-owned, container writes use the host UID, and unsupported suspension is omitted only through frozen-target authorization.
R2 91067e340a | runs root lifecycles from the dependency inputs copied by %s [file=Dockerfile] | Execute root lifecycle hooks from the root image's actual dependency-stage COPY closure; extends the missing-hook regression protection to its independent recipe.
R2 dbc08f64c1 | runs root lifecycles from the dependency inputs copied by %s [file=scripts/docker/cleanup-smoke/Dockerfile] | Missing prepare-git-hooks input broke cleanup-image dependency installation; execute lifecycle hooks using only copied inputs.
R2 cbf23a1eba | selects one release-owned Windows helper mount | Frozen upgrade scenario imports require the selected release's helper, without a competing trusted-harness mount at the same destination.
R2 3009338a9a | keeps a stalled multi-node health request inside the probe deadline | A stalled fetch previously outlived the health probe deadline; prove abort and failure with accelerated clock progression.
R2 5ce9b61ceb | reports the installed doctor switch unit through the systemd manager | Exercise managed-container service inspection through the production systemd consumer and the real fixture manager protocol.
R2 83dce46a7d | distinguishes a missing named doctor switch unit from failed or unsupported inspection | Failed, unsupported or unreadable service inspection must not become authoritative absence during sealed-definition update handling.
R2 7444f28d2d | mounts the %s Doctor contract and canonical-path shims from the same checkout [targetMode=selected] | Selected-release scenario and canonical-path service shims must come from the same selected checkout.
R2 8d909ed0da | passes installer tag env to bash, not curl | Installer tag variables assigned to curl did not reach the installer shell.
R2 dfa1a51225 | keeps the plugin binding command escape Docker smoke focused | An extra -- prevented Vitest filter arguments from reaching the runner; retain focused argv and summary enforcement.

* test(discord): deslop t0122 tests

* test(agents): deslop t0113 tests

* test: retain distinct safety and recovery coverage

* test: preserve native compiler runtime routing after consolidation
2026-09-29 01:07:27 +00:00
Mitchell Etzel
6eeb120d69
fix(gateway): derive the darwin stop budget from the launchd job (#157007)
* fix(gateway): derive the darwin stop budget from the launchd job

resolveGatewayShutdownBudget derives the stop deadline from restart
ownership, so a darwin host running with OPENCLAW_SUPERVISOR_MODE=external
resolves drain=315000ms whatever launchd actually enforces.

The linux-gated systemd probe becomes a platform dispatch, and a new
readLaunchdStopTimeout mirrors readSystemdStopTimeout: it reads the running
job's effective exit timeout and accepts it only when the printed pid is
this process. run-loop.ts is untouched because the reader detects launchd
from the environment itself.

Closes #156968

* fix(gateway): accept the launchd launcher parent and cap its deadline

The darwin reader accepted the printed job only when its pid was this
process. The installed service can keep a launcher parent while the serving
Gateway runs as its child, so the job prints the launcher's pid, the
enforcing job was rejected, and the Gateway fell back to 20 seconds in the
one layout where an operator's ExitTimeOut was meant to apply.

Accepting that job at face value would be worse than rejecting it.
node-runtime-recovery.mjs builds the launcher's own reap timer from the
compile-time LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS rather than from the job, so
on a forwarded stop it re-sends SIGTERM to its child at 18000ms and SIGKILLs
it at 19000ms. Adopting a 90 second job deadline in that layout would plan a
75000ms drain and then lose it to its own parent, which is the truncated
drain this change exists to prevent. The reader now accepts process.pid or
process.ppid and takes min(jobExitTimeout, 20000) in the launcher case,
naming the cap in the source string only when it binds. A shorter job
deadline still applies in both layouts, because launchd reaps the whole job
regardless of the launcher.

The docs claimed the deadline is read at startup and when accepting
shutdown. run-loop.ts gates the budget refresh on linux at both consumption
sites, so on darwin it is read once, at startup, and nativeStopBudget has no
shutdown-time consumer there. The page now says startup-only and drops the
sentence about a failed shutdown reread retaining the startup budget, which
described linux behaviour only.

* fix(gateway): scope the darwin stop budget to launchd-driven stops

Address the two review findings on the previous revision.

The launcher parent is now identified by an explicit
OPENCLAW_LAUNCHER_STOP_TIMEOUT_MS declaration rather than a bare ppid
match, so an unknown parent keeps the job's full deadline instead of
being capped on a guess.

The darwin reader inspects the launchd job only while a stop is under
way, and adopts that job's deadline only when the job reports the
SIGTERMed state. An external SIGTERM keeps the platform-neutral drain,
because launchd's ExitTimeOut bounds only a stop launchd itself is
running.

The job state is read with an anchored single-tab pattern. launchctl
print emits further "state = active" lines inside the resource and
jetsam coalition blocks, and the shared key/value parser keeps the last
occurrence, so the shared parser could never observe SIGTERMed. A test
fixture reproduces the nested coalition shape.

* fix(gateway): keep a failed launchd inspection off the native stop budget

`readLaunchdStopTimeout` answered a failed job inspection with the Gateway's own
330000ms stop policy wrapped in a non-null result. `resolveGatewayShutdownBudget`
classifies every non-null result as a native stop budget, so a probe that
established nothing still set `nativeStopBudget`. On an externally supervised
darwin Gateway that caps a longer requested restart drain and can arm a forced
exit against a deadline no supervisor was confirmed to be enforcing.

The reader now reports its two answers separately. `stop` carries a deadline only
while launchd is confirmed to be stopping the job; `warning` reaches the operator
either way. A failed inspection reports `stop: null` and leaves the caller on the
platform-neutral policy it had already resolved, which is the same number as
before, now classified honestly. The reader no longer imports
GATEWAY_SERVICE_STOP_TIMEOUT_MS, because choosing the Gateway's policy number was
never its job.

Tests assert the flag and the downstream restart drain in the failure case, paired
with a confirmed-deadline case so deleting the read outright cannot pass both.

Also fixes TS2532 on the job-state regex: the named group is optional under
`noUncheckedIndexedAccess`, so `.trim()` needed the optional chain.

* test(infra): isolate the hoisted spawn mock and ratchet the OPENCLAW_* count

`spawn` is hoisted once per file, so its call log survived across cases and
`toHaveBeenCalledExactlyOnceWith` could only hold for the first one. Clear it in
`beforeEach`.

CI reports `OPENCLAW_* count 488 exceeds budget 487; update
config/env-var-count-budget.txt` as a warning. The new name is the launcher's
OPENCLAW_LAUNCHER_STOP_TIMEOUT_MS declaration, so the count is correct and the
ratchet moves to 488 with the reason recorded in the file.

* style(infra): apply oxfmt to the new launchd reader assertions

Run by the repo's own `pnpm run format`; joins one over-wrapped assertion line.
`pnpm run format:check` then reports "All matched files use the correct format."

* test(cli): mock the launchd reader and split the darwin run-loop cases out

Activating the darwin stop-budget refresh made `run-loop.test.ts` exercise the real
`readLaunchdStopTimeout`, which spawns three `launchctl print` calls against live
subprocess I/O while the suite holds `vi.useFakeTimers()`. One case ran 120039ms
before failing and the abandoned loop leaked into later cases, for eight failures
in total. All eight clear by mocking the reader; no existing expectation was wrong.

The mock could not simply be added. `run-loop.test.ts` is 3813 lines against a 1000
line cap where `check-line-cap-ratchet` rejects any growth, and an earlier attempt
failed with `3501 -> 3510 counted lines`. Taking the remedy that ratchet names, the
darwin cases move to `run-loop.launchd.test.ts` and the original file shrinks by 82
lines. It keeps the mock too, because four of those cases are registered into it by
shared `register*` helpers whose `it.each` arrays interleave systemd and launchd
variants and cannot be split without touching those support files.

The new file adds one case the default mock would otherwise hide: a job reporting a
30000ms deadline bounds the stop at 25000ms and arms the force exit there.

`node-runtime-recovery.test.ts` asserts the replacement child's spawn env exactly,
so it gains the launcher's own OPENCLAW_LAUNCHER_STOP_TIMEOUT_MS declaration rather
than a loosened matcher. That argv is a non-foreground `doctor --fix`, so the value
is the 1s signal exit grace plus the 1s force-kill grace.

* fix(gateway): only retain a startup budget when the probe was inconclusive

Extending the stop-budget refresh to darwin gave the retained-budget safety net a
new and wrong trigger. On a launchd-owned Gateway the startup budget is already
native, and an in-process restart signals the process without launchd running the
stop, so the job still prints `running` and the reader reports `stop: null`. The
old condition treated any null as unconfirmed, so every in-process restart of a
default macOS install logged "Retaining the startup shutdown budget of 15000ms
because the current supervisor stop timeout could not be confirmed." The timeout
was confirmed; it was confirmed not to apply.

`readNativeStopTimeout` now reports `inconclusive` alongside the deadline, and only
an inconclusive probe retains. Linux keeps its exact previous condition, an absent
unit or a warned read. The budget number was already unaffected, so this is the
message and the wasted `launchctl print` calls, not the deadline.

Two cases pin it: a confirmed not-stopping read warns nothing, and a failed
inspection still retains and still says so.

Also fixes two wording defects found in review. `unresolved()` no longer takes a
label it only ever interpolated as "the launchd job the configured label"; each
failure already names its own target. And the `readJobState` comment said `state`
appears four times in `launchctl print` output when its own next clause counts
three, which is what a live LaunchDaemon prints.

* style(cli): apply oxfmt to the new retained-budget assertions

Run by the repo's own `pnpm run format`; joins one over-wrapped assertion line.
Whitespace only, in a test file, so it cannot change runtime behaviour.

* fix(gateway): a defaulted deadline is not an inconclusive probe

`inconclusive` keyed off the warning alone, so the defaulted-value case counted as a
failed probe. There launchd is confirmed to be stopping the job and only its
deadline had to be guessed, so a clock is genuinely running and there is nothing to
retain. Under launchd ownership that combination logged both "using 20000ms
default" and "could not be confirmed" for the same stop, the second being false.

Now only a warning with no deadline at all counts as inconclusive. A third case
pins it, asserting the defaulted read warns exactly once and about the deadline
rather than about confirmation.

Also corrects a test fixture comment that still said `launchctl print` emits `state`
four times; the arrangement it documents, and a live LaunchDaemon, both show three.

* test(cli): correct the regression guard's account of the reported drain

The comment said active work "had time to finish" during the reported 315 second
drain. The reporting host's log shows the opposite: that drain hit its own timeout
with five tasks still active, and the Sep 20 stop with six. The drain was being
used, which is a stronger reason for the guard, not a weaker one.

* fix(infra): stop exporting a type nothing outside the module reads

Knip's all-exports gate reads LaunchdStopTimeout as dead: the only
public surface is readLaunchdStopTimeout's LaunchdStopRead return, and
nothing imports the nested shape by name. Keep it module-local.

* fix(gateway): keep a proportional drain and cap an undeclared launcher

The fixed 10s reserve and 5s exit margin were sized against the 315s policy
drain. Subtracting them outright from a launchd job's own ExitTimeOut spent a
short deadline entirely on overhead: 5 seconds funded the margin alone and 15
funded margin plus reserve, so active work drained for zero milliseconds in
both. Each allowance now takes at most a share of what it is carved from, so a
deadline long enough to fund them is unchanged and a shorter one keeps a
proportional drain. Drain is now positive for every positive deadline and never
decreases as ExitTimeOut grows.

The shutdown log also subtracted the unscaled reserve constant when reporting
drain, which understated a short budget by the amount the scaled reserve gives
back, so it now reports the margin and drain actually spent.

An already-running launcher published before OPENCLAW_LAUNCHER_STOP_TIMEOUT_MS
existed arms the same reap timer and declares nothing, which is the upgrade
shape: replacing files cannot change a launcher that is already running. Reading
that silence as "no deadline" let a long custom ExitTimeOut be budgeted past a
force-kill the parent was already counting down. The launcher's arithmetic moves
into the shared budget module so the serving Gateway reconstructs the same timer,
used only where the printed job carries OpenClaw's own label and its pid is this
process's immediate parent. In that position the parent is the process launchd
started for OpenClaw's job and the recovery launcher is the only path that puts a
Gateway underneath it, so an unrelated process manager holding the parent slot
still caps nothing.

* fix(infra): gate the launcher reconstruction on the recovery respawn marker

Holding the launchd job's pid is not evidence of a reap timer. An operator
wrapper can keep that pid and start the Gateway itself while running none, so
reconstructing a cap from the parent relation alone would cut a drain nothing was
going to interrupt. The recovery launcher has stamped OPENCLAW_NODE_UPDATE_RESPAWNED
on every child it respawns since long before it declared a stop timer, so that
marker is present in exactly the upgrade case and absent for any other parent.

* fix(infra): declare the new shared budget helpers for the type checker

The root module is typed by a hand-written declaration file rather than being
compiled, so new exports are invisible to src without being declared there.

* docs(gateway): scope the positive-drain claim to the allocation

The elapsed cost of resolving the budget is debited after the allocation, so a
deadline shorter than that cost can still leave nothing to spend. Claiming every
positive deadline yields a positive drain overstated it.

* test(cli): accept the retained-budget attribution on a synthetic launchd label

The fixture declares a launchd label with no matching job, so the per-stop
re-inspection cannot read one and the startup budget is retained. The deadline is
unchanged at 15000ms; only the source attribution differs. The assertion still
fails if the budget comes from the platform-neutral policy instead.

* fix(gateway): derive the launcher's reap deadline instead of declaring it

Measurement showed the declared value and the derived value are the same number:
a published launcher and a candidate launcher both arm an 18000ms exit grace, and
a candidate Gateway resolved an identical 19000ms cap and 14231ms budget under
each. The declaration was therefore carrying no information the child could not
compute, while adding an OPENCLAW_* name and leaving the upgrade path conditional
on which build of the launcher happened to be running.

Both sides now read one expression in gateway-shutdown-budget.mjs, the launcher to
arm its escalation and the Gateway to bound its budget, so they cannot disagree and
there is no published-versus-candidate launcher distinction left to prove. The cap
stays gated on OPENCLAW_NODE_UPDATE_RESPAWNED so a parent that did not respawn this
process still caps nothing.

Drops the env-var count back to 486, so the name ratchet no longer needs an owner
waiver, and takes config/env-var-count-budget.txt out of the diff entirely.

Test deadlines move off 90 seconds onto 55, because launchd clamps ExitTimeOut at
60 and a 90-second case cannot occur on macOS 27.

* test(infra): restore the recovery spawn-env cases to their base form

The launcher no longer declares a deadline, so the assertions these cases grew have
nothing left to check and their explanatory comment described behavior that is gone.

* fix(gateway): cap every respawning launcher, not just the Node-recovery one

All three runRespawnedChild call sites reach the same launcher through the same
function and arm the same escalation, but only the Node-recovery one sets
OPENCLAW_NODE_UPDATE_RESPAWNED. Gating on that marker alone left the two
compile-cache respawns uncapped, and the packaged one can wrap a foreground
gateway run on an installed service: the job's deadline would then be budgeted past
a force-kill the parent had already armed, which is the hazard this cap exists to
prevent.

The marker set moves next to the deadline it authorises, so adding a respawn path
cannot silently escape the cap. Keeping the literals in the root module also keeps
them out of the env-count ratchet's scope: reading the marker from counted
production source promoted a previously test-only name and held the count at 487
even after the declared-timer variable was removed.

Also applies oxfmt's own wrap to the launcher-relation expression.

* fix(infra): re-export the respawn marker set through the infra barrel

The constant existed only on the root module, so the test importing it from the
barrel got undefined and it.each(undefined) threw during collection: the suite
reported zero tests rather than a failure, and typecheck flagged the missing member.
Re-exporting the binding adds no OPENCLAW_* literal under src, so the env-var count
stays at 486.

* fix(gateway): address the pre-review findings on the launchd stop budget

Docs were stale or wrong in four places. The systemd section still promised a fixed
10 second reserve and 5 second margin, but the share caps are platform-neutral and
do change a Linux unit whose TimeoutStopSec is under 25 seconds; that radius is now
stated with the 20 second and 15 second cases spelled out. The direct-SIGTERM
paragraph claimed the platform-neutral drain "is the deadline that actually governs
that stop", which the branch's own kill -TERM capture contradicts: no supervisor
deadline governs that stop at all and the fallback is the template value. The
updated-launcher requirement no longer applies to the stop budget and says so. The
marker sentence named one marker when three are honoured, and one measured figure
was stale.

Two claims in the code were too strong. The marker list asserted every setter arms
the shared escalation, which is false for the compile-cache name: entry.compile-cache
sets it for a runner that reaps on a fixed short grace. That runner refuses a
foreground Gateway run off Windows and this deadline is only read during a darwin
stop, so it cannot be the parent here, and the comment now says that instead of
implying exclusivity. The sync note next to the escalation now records that the other
runner keeps its own copies of the graces.

An absent job state was indistinguishable from a job launchd is not stopping, so a
macOS that printed the block differently would silently revert to the platform-neutral
policy. When a deadline parsed but the state did not, that now warns and marks the read
inconclusive rather than passing as a positive answer.

Tests: launchd cases move off 90 and 315 second deadlines, which launchd clamps to 60
and cannot produce; the marker list is restated locally and tied back to the
implementation so dropping a marker fails; and the real-process case now asserts the
darwin probe actually ran, which it could not before.

Retires five exports with no production consumer that the dead-export scan flagged.

* test(infra): move the empty job state onto the missing-state warning

An empty value yields no state to recognise, so it reaches the same branch as a
missing line and now carries the same warning. Keeping it in the unrecognised-state
table asserted the absence of a field the reader deliberately reports.

* fix(infra): treat a whitespace-only job state as missing, and sharpen two docs claims

The state value is trimmed after matching, so a line carrying only whitespace yielded
an empty string rather than undefined and skipped the warning while behaving exactly
like a missing line.

The systemd section now says the exit margin shrinks below a 20-second deadline, not
just the reserve. The direct-signal paragraph distinguished only the launchd-supervised
budget; external supervisor mode keeps the platform-neutral policy and arms no
force-exit timer, and a signal delivered to the job pid does reach a resident recovery
launcher, which arms its own reap timer even though launchd is not stopping the job.

* docs(gateway): correct the loaded launchd job deadline guidance

launchctl print reports the job launchd has loaded, not the plist on disk,
so editing a loaded job's ExitTimeOut changes nothing until the job is
reloaded. Measured on macOS 27 against a scratch job: the plist moves from
20 to 47 while the loaded job keeps reporting 20, and it still reports 20
after launchctl kickstart -k, which restarts the process without reloading
the job. Only bootout then bootstrap makes it report 47, and that restarts
the Gateway with it.

State what reading per stop actually buys instead: the Gateway never plans
against a deadline it cached at its own startup.

* fix(gateway): keep the full cleanup reserve wherever a deadline funds it

The reserve was capped at half the shutdown budget whenever the budget could
not fund it outright. That kept a drain on short deadlines, but it also
reallocated deadlines that already worked: a 20 second ExitTimeOut, which is
both the shipped LaunchAgent template and launchd's own default, moved from a
10 second reserve and a 5 second drain to 7.5 seconds of each. Cleanup costing
between 7.5 and 10 seconds finished before and would have been cut off.

Bound the reserve by what keeping a drain actually requires instead. The floor
is 5 seconds, which is the drain the 20 second template already yields once the
fixed margin and reserve are subtracted, so funding the full reserve alongside
it takes 15 seconds of budget: exactly what that template resolves to. Every
deadline from 20 seconds up therefore keeps the allocation it had, and only a
deadline the fixed subtraction had already driven under a 5 second drain gives
any reserve up, never below half the budget.

Both allocations stay monotone in the deadline, and the 5 second case is
unchanged at 1875ms of each.

* fix(gateway): correct shutdown-budget probe-debit overclaims

- gateway-shutdown-budget.mjs: replace the floor-invariant claim with the
  real bound (reserve unchanged only at a >=15000ms post-margin budget,
  i.e. a >=20s deadline once the launchd probe's own cost, up to three
  launchctl print calls at 2000ms each, is subtracted); name the 8-20s
  sub-band reserve loss the old comment denied, including 16-19s where a
  flat subtraction already left a positive drain while still funding the
  reserve in full
- docs/gateway/restart-recovery.md: drop the systemd section's "down to
  the millisecond" promise and the launchd section's "exact allocation
  ... before this change" claim; both now state the inspection cost that
  comes off instead
- run-loop-shutdown-budget.test.ts: add a real 13ms elapsed-debit case at
  a 20s exit timeout (asserts reserve=9987, drain=5000, contrasting with
  the existing acceptedAtMs=MAX_SAFE_INTEGER zero-debit fixtures) and
  mirror darwin's 19/16/15/10s sub-20s it.each against a systemd stub,
  which previously had no sub-20s coverage

* test(cli): wrap the flat-subtraction assertion to satisfy oxfmt

The sub-20-second systemd cases added in 60e5eb509b6d left the flat-subtraction
comparison on a single line past the formatter's width, which failed oxfmt --check
and took check-lint and check-docs down with it. Lint itself reported zero warnings
and zero errors. Formatting only, no assertion or value changed.

* docs(gateway): bound the 20-second allocation claim by the inspection debit

The previous wording paired 'gives up exactly what the probe cost' with a probe
bounded near 6 seconds, and those two cannot both hold. Which allowance pays the
debit turns on a 5 second threshold: at or under it the reserve absorbs the whole
cost and the 5 second drain floor is untouched, so a 20 second deadline resolves
10000 minus the debit. Past it the floor is share-bounded as well and the two
converge on half the remainder, so a 6 second debit splits 9000 into 4500/4500
rather than retaining the old reserve. Documents the threshold in the docs and the
root-module comment, and pins the over-threshold case with a test so the boundary
is measured rather than reasoned.

* docs(gateway): qualify the full-reserve claim by the probe debit

The drain-floor comment asserted that a whole 10 second reserve follows
for every ExitTimeOut of 20 seconds or more once the probe that reads the
deadline is subtracted. The probe cost is charged before the budget is
split, so a 20 second ExitTimeOut clears the 15 second budget that a whole
reserve needs only when that probe costs nothing. At the 13ms measured on
this rig the reserve is 9987, and a whole reserve at that cost needs 20013.

The same comment already states this correctly six lines down, where it
resolves a 20 second deadline to 10000 - cost. This removes the
contradiction by qualifying the claim at its first statement instead of
restating the arithmetic twice.

Comment-only: no non-comment line changes and the export list is unchanged.

* fix(gateway): honor unlimited launchd stops

Preserve the observed launcher cap only for launchd-driven stops, and shorten redundant budget documentation without changing the measured split.

Co-authored-by: Patrick-Erichsen <20157849+Patrick-Erichsen@users.noreply.github.com>

* test(gateway): use action-based drain budget

Co-authored-by: Patrick-Erichsen <20157849+Patrick-Erichsen@users.noreply.github.com>

* style(gateway): format launchd timeout assertions

Co-authored-by: Patrick-Erichsen <20157849+Patrick-Erichsen@users.noreply.github.com>

* fix(ci): complete frozen Docker planner import closures

---------

Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com>
Co-authored-by: Patrick-Erichsen <20157849+Patrick-Erichsen@users.noreply.github.com>
Co-authored-by: Patrick Erichsen <patrick.a.erichsen@gmail.com>
2026-09-28 18:04:14 -07:00
Peter Steinberger
5d605627c5
perf(sessions): yield during SQLite entry write contention (#160801)
* perf(sessions): yield during SQLite entry write contention

Keep the Gateway event loop responsive while session-entry patches wait for another SQLite writer. Reuse begin-only asynchronous admission with zero native busy waiting, retaining FIFO ordering, connection custody, commit revalidation, and non-replay of admitted writes. No schema or configuration changes.

* fix(sessions): enforce synchronous callbacks during yielding admission

Name yielding admission separately from retired async-transaction APIs and register its synchronous callback with the source guard. Make the Slack Stop-owner fixture await its competing committed session write instead of depending on updater microtask ordering.
2026-09-29 01:03:21 +00:00
Peter Steinberger
25df9ed0f5
fix(doctor): preserve the live plugin index during read-only checks (#160400)
* fix(plugins): keep Doctor native captures out of live state

* test(plugins): use complete metadata snapshot fixtures

* fix(doctor): retain native captures in the profile context

* fix(doctor): include capture resolver in trusted wrapper
2026-09-29 00:52:38 +00:00
Jason (Json)
4a448bdd26
fix(ui): release memory when closing terminals (#160748)
* fix(ui): release closed terminal memory

* test(channels): drop redundant Synology builder inventory

* test(ui): use browser global type in terminal fixture

* test(agents): observe outbox admission ordering

* test(agents): use typed deferred in outbox ordering proof

* fix(release): verify planner dependencies in frozen tooling
2026-09-28 18:46:45 -06:00
Peter Steinberger
e066d8f473
fix(scripts): local macOS packaging fails on the second run in the same checkout (#160809)
A second packaging run in the same checkout reused apps/macos/.build/<arch> and died in verify_snapshot_swift_lock with only "exit code 1": swift-cmark had moved from the committed 0.8.0 to 0.9.0.

Each run synthesizes the package root under a new work directory, so the scratch path's workspace state still records OpenClawKit under the previous (deleted) root. With the committed originHash stale since #159535, the first SwiftPM command (an unguarded `unedit`) ran SwiftPM's full resolver, which could not load OpenClawKit at the recorded path, dropped that subtree's pins and floated swift-cmark, rewriting Package.resolved before the lock guard took its snapshot.

Resolve in SwiftPM lock-file mode instead and drop the pre-build unedit: it checks out exactly the committed pins, rebinds local packages, returns a leftover edited Peekaboo to its pin, and heals an already drifted scratch path. verify_snapshot_swift_lock now names each pin that differs, and a harness test requires the locked lock-file resolve to be the first SwiftPM operation.
2026-09-28 17:39:48 -07:00
Peter Steinberger
fd17d340ad
perf(plugins): reduce repeated stream payload inspection (#160208)
Share data classification only within one synchronous iterator-result read. Preserve mutation rechecks and nested consumer admission while reducing repeated payload walks. Add an old-code-failing regression and a reproducible streaming allocation benchmark.
2026-09-29 00:34:08 +00:00
Peter Steinberger
c432db8665
feat(slack-huddles): join Slack huddles as a signed-in Slack user (#159879)
Add an opt-in slack-huddles plugin that joins active Slack huddles as a
dedicated signed-in Slack user. Slack has no bot or app huddle API, so the
plugin drives the Slack web client in the OpenClaw Chrome profile on the
shared meeting runtime, with in-page audio capture and a virtual microphone
for agent, bidi, and transcribe modes, mirroring zoom-meetings and
teams-meetings.

Membership is proven only by Slack's in-huddle channel header on the
requested channel. Sessions are workspace-scoped, joins are always muted,
talk-back unmutes only after Slack reports the virtual input, and every
click rechecks live authority. The shared meeting status source gains
optional liveOwnershipSource and afterAudioRoutingSource hooks. They
recheck ownership after awaited device and sink work, and roll back only
this pass's effects. Absent hooks emit no code, so the generated Zoom and
Teams scripts are byte-identical.

The plugin is disabled by default and requires OpenClaw 2026.9.8 or newer,
the first release that can carry the ownership hook.

Release note: Add an opt-in Slack huddles plugin that joins active huddles
as a dedicated signed-in Slack user through the shared meeting runtime.
Requires OpenClaw 2026.9.8 or newer.
2026-09-28 17:34:03 -07:00
Peter Steinberger
80fd8670e3
refactor(scripts): deslop tooling scripts (#160798) 2026-09-28 17:20:56 -07:00
Josh Avant
2a778ef887
chore(ios): simplify store release workflow (#160791) 2026-09-28 19:00:37 -05:00
Peter Steinberger
91a86e494e
refactor(plugins): deslop plugin runtime fifth pass (#159945)
* refactor(plugins): deslop plugin runtime fifth pass

Consolidate repeated provider, manifest, install, catalog, hook and registration projections under their existing owners. Preserve plugin loading, SDK, caching and source-security contracts.

Preserve provider wizard modelSelection preferences through registration so configured onboarding choices retain their documented model-picker behavior.

* fix(plugins): preserve build and type cleanup contracts

* fix(plugins): retain distinct public credential source types

* fix(plugins): preserve sparse catalog matching and cleanup contracts

* fix(plugins): preserve include ownership during uninstall

* fix(plugins): include request policy in PR wrapper inventory
2026-09-28 16:59:20 -07:00
Peter Steinberger
682ffe55a9 fix(release): admit the provider catalog read by legacy-operator survivor planning
ac59199ad3 made docker-e2e-plan.mts import scripts/lib/official-external-provider-catalog.json and scripts/lib/record-shared.mjs and read the candidate's provider catalog through the frozen target reader for legacy-operator-state lanes. Frozen admission hydrates only declared source paths in a no-lazy-fetch checkout, so targeted survivor dispatches failed with 'unable to read selected source (cat-file)', and the admission fixture lacked both tooling imports. Declare both files in the tooling closure and hydrate the candidate catalog for survivor lanes.
2026-09-28 16:56:52 -07:00
Peter Steinberger
ae1d3da7a6 fix(ci): restore frozen planner dependency closure
Include the provider catalog and record helper added by ac59199 in frozen
tooling admission and all affected copied harnesses. Verify their bytes
through the existing closure owner. Keep child evidence isolation covered
with the child-only Codex install helper now that record-shared is common.

The original package failure reproduces locally: 93 failed and 462 passed.
Both added integrity cases fail before the closure repair. Afterward, all
555 package tests and 445 closure sibling tests pass. Independent review
is clean through P2; the full selected static check passes in 1448.701s.

Measured one-worker wall times: package 218.673s and closure siblings
267.426s. The two added integrity cases take 208ms and 202ms. No timeout,
retry, baseline, or integrity check is relaxed.

Related CI: 36494713554 and 36494667447. Separate startup assertion and
Synology caller fixes landed in 8799587 and c4845a6 and are not duplicated.
The separate macOS native reconnect failure in main CI 36492763105 remains
causally unproven and is not claimed fixed here.
2026-09-28 16:46:48 -07:00
Peter Steinberger
7c1838a418
fix(ci): scope untouched line-limit diagnostics in changed checks (#158995) 2026-09-28 16:35:33 -07:00
Peter Steinberger
db5f2df56c fix(release): retry the shipped 2026.9.6 Windows updater after its exit-13 liveness defect
Windows packaged upgrade from openclaw@2026.9.6 to the 2026.9.7 candidate failed on Node 24 and 26.1.0 because the shipped 9.6 updater exits 13 (unsettled top-level await) before switching the install: it awaits an unreferenced 300 ms cleanup timer after cancelling the retained-runtime progress probe. 02720df7e6 fixes it for 2026.9.7+, but cannot repair the executing 9.6 parent, and it is intermittent on native Windows.

Classify exactly that signature (win32, exit 13, the warning, empty stdout, install still on a baseline older than 2026.9.7), retry the same update once as a user would, and only if the retry fails the same way use the existing verified direct candidate install. Both outcomes are recorded as updater fallback evidence; any other failure still fails the lane.
2026-09-28 16:23:05 -07:00
Patrick Erichsen
f9358275ab
fix(agentmail): allow official ClawHub channel to load (#155097)
* fix(agentmail): allow official ClawHub channel to load

* docs(agentmail): keep setup guide vendor-owned

* fix(agentmail): require API key for configured state

* fix(plugin): verify pinned digest before legacy trust
2026-09-28 16:10:33 -07:00
Peter Steinberger
e5d4585279
test(agents,channels,cli): remove low-value tests (batch d091) (#160766)
* test(heartbeat): deslop t0067 tests

* test(agents): deslop t0119 tests

* test(telegram): deslop t0118 tests

* test(discord): deslop t0121 tests

* test(agents): deslop t0120 tests

* test(state): deslop t0129 tests

* test(browser): deslop t0128 tests

* test(status): deslop t0112 tests

* test(matrix): deslop t0125 tests

* test(cli): deslop t0082 tests

* test(agents): remove unused compaction fixture exports
2026-09-28 22:52:38 +00:00
Peter Steinberger
1fd2ad9b56 fix(e2e): size upgrade-survivor Gateway readiness for saturated runners
Release Checks 36479006821 saw both the published 2026.9.6 baseline and the
2026.9.7 candidate bind HTTP after 56-58 s and miss the fixed 90 s readiness
window on saturated runners, while the same phases pass in 36-41 s on idle
ones. Raise the default startup budget to 300 s and make the readiness wait
honor the configured budget instead of a hardcoded 360 polls, so the
OPENCLAW_UPGRADE_SURVIVOR_START_BUDGET_SECONDS override actually extends it.
2026-09-28 15:46:43 -07:00
Peter Steinberger
181405f477
fix(ci): planner ownership census times out on loaded runners (#160586)
* perf(ci): scan wide import-graph frontiers in one pass

The PR-exempt ownership census plans a PR that edits all 906 PR-exempt
test files. Its consumer check sends ~2,700 terms through the wide native
reader path over ~42.5k tracked files. #160531 and #160517 each added a
match-before-parse filter, so that path spawned one reader child to match
every file and a second to re-read and parse the 2,136 matches.
readImportGraphEdges already requests matchingOnly, so drop the extra
pass: one child matches, parses only uncached matches, and still leaves
nonmatches uncached.

Build the Aho-Corasick automaton as a dense Int32Array over the UTF-16
code units the terms use, completing failure rows breadth-first, so the
scan takes one table load per code unit instead of Map lookups and
failure walks. Files with no match skip projecting every requested term.
Matching semantics, reference checks, and ordering are unchanged.

For the real 906-file request, reader output is byte-identical. Reader
CPU fell from ~12.1s over two passes to ~5.9s. Inside the Vitest worker
the census case's reader cost fell from 16.5s to 8.8s, and the case from
39.1s to 30.4s on a comparably loaded host. Timeout, assertions, and file
coverage are unchanged.

* perf(ci): bound term-matcher rows to the ASCII alphabet

Search terms come from tracked paths, which have no character-diversity
bound. One dense column per distinct UTF-16 unit could therefore grow the
transition table by nodes x alphabet on Unicode-heavy frontiers.

Keep dense Int32Array rows only for the ASCII units the terms use, at most
129 columns per node, and route other term units through sparse per-node
edges with the classic failure walk. Size the table once to its node upper
bound so it never reallocates; only rows of real nodes are written. Fold
isReferenceCharacter into an equivalent character-class test, which is
exhaustively equal for every UTF-16 unit and NaN.

The matcher contract fixture gains terms whose failure links cross dense
and sparse edges. A 24,000-case randomized differential against main's
matcher and a String.includes oracle shows no differences, and reader
output for the 906-file census request stays byte-identical.
2026-09-28 15:45:40 -07:00
Peter Steinberger
ac59199ad3 fix(e2e): stage candidate providers for legacy-operator survivor lanes
Legacy published baselines (2026.7.35, 2026.6.34) author plugins.allow from
their bundled provider inventory. The candidate's doctor restores those
now-external providers at the exact candidate version, which does not exist
on npm or ClawHub before publication, so the prepublish registry must carry
them. Stage every official provider from the selected candidate's external
provider catalog for legacy-operator-state lanes, loaded once per plan.
2026-09-28 15:37:59 -07:00
Peter Steinberger
d978a19298
perf(cron): narrow delivery-target session reads (#160216)
* perf(cron): narrow delivery-target session reads

* fix(ci): keep plugin retention proof on Node
2026-09-28 22:20:34 +00:00
Jason (Json)
28e16eedc3
fix: dashboards require approval after native app authentication (#159860)
* fix: dashboards require approval after native app authentication

* fix: register native auth test entries and remove the client type cycle

* fix: recover native dashboard auth on older WebViews and reconnects

* fix: preserve native dashboard auth across current app paths

* fix(ci): use idle hosted tooling workers during overflow

* fix(ci): preserve oversized tooling refusal

* test: load first-hop helpers in the survivor fixture

* test(macos): wire conversation auth fixture and tokenless route

---------

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-28 16:16:53 -06:00
Peter Steinberger
2e7f4f8fcd
fix(ci): pack hybrid hourly tooling within the row cap
Use the existing hosted retry file packing for hybrid tooling so an
indivisible file can share its spare worker without increasing its price.
The exact hourly hybrid plan drops from 80 total rows to 78, within the
unchanged 77 Node plus two dist limit. GitHub remains at 76 total rows.

Add the hybrid hourly cap regression and retain the GitHub ownership and
cap assertions. Preserve the complete file/config/environment inventory,
release descriptors, timing weights, runner policy, and worker limits.

Validation: 26,658 tooling tests and all test types passed on Linux
Testbox; script types and lint passed there. The connection ended after
those gates, so final root-test lint and format checks passed locally.
The integration file passed all 10 cases locally. P2 autoreview is clean.
2026-09-28 15:04:46 -07:00
Peter Steinberger
0a5ef7d9e0
fix(node-host): updates never install on Bun-only macOS and Linux hosts (#160575)
* fix(node-host): prepare updates on Bun-only POSIX hosts

Read exact registry manifests and stream SHA-512-verified archives in process, then use shared Bun staging inside the private node runtime generation. Preserve npm preparation on Node and Windows, including existing schema and candidate checks.

Cover Bun preparation, private bins, retained generations, invalid registry archives, loopback discovery, and Windows npm selection. Remove the Linux smoke blocker and document the remaining Windows limitation.

* refactor(update): share registry package document reads

Consolidate registry URL construction, HTTP fetching, JSON consumption, deadlines, and response cleanup for update discovery and node runtime preparation. Preserve caller metadata mapping, HTTP error strings, diagnostic labels, and body timeout policies.

Remove the redundant registry status wrapper and share the existing discovery error return. This removes 25 production lines without changing preparation or its regression tests.

* fix(node-host): skip npm freshness probing for Bun preparation

* fix(node-host): create the private Bun node_modules root before staging

* fix(node-host): seed the private Bun global project manifest

* fix(node-host): skip Node engine checks for candidates run on Bun

* docs(bun): link node update preparation history row

* refactor(node-host): move registry archive reads out of install-source-utils

install-source-utils sits under plugin install owners; importing the shared
registry document reader from it closed an import cycle through the provider
and plugin metadata graph. The in-process manifest and archive reads now live
in a leaf module that only the node host imports.
2026-09-28 14:59:44 -07:00
Peter Steinberger
87702484d1
fix: worker turns fail on older session-host nodes after the Gateway adds a worker tool (#160147)
* fix(gateway): keep worker turns working on session hosts that predate new worker tools

A Gateway built from main added the `presence` worker session tool, and
every worker-turn launch to a published openclaw@2026.9.6 session-host
node failed with `INVALID_REQUEST: invalid worker launch descriptor`
followed by `node worker cancellation timed out`. The installed node
supervisor validates `toolAuthority.allowedToolNames` against a closed
vocabulary, and bundle refresh does not replace the supervisor.

The Gateway now negotiates the launch vocabulary per node, following the
existing capability pattern: it advertises
`node-worker-launch-tool-names-v1`, updated nodes declare
`workerHost.launchToolNames` only to Gateways that advertise it, and the
Gateway treats an absent declaration as the frozen 2026.9.6 vocabulary.
Tool authority filters the projected tool set by the destination
vocabulary, so turn authorization and the launch descriptor stay
identical. Unknown declared names are ignored, so future worker tools
need no further capability.

Update behavior: Gateway-first updates keep older nodes hosting turns
without newer tools (one info log names them); node-first updates keep
the declaration unchanged for older Gateways; updated pairs get presence.

* fix(scripts): add worker tool authority to the PR wrapper inventory

src/infra/node-runner-inventory.ts is in the trusted-anchor PR wrapper's runtime import closure and now imports src/worker/tool-authority.ts for launch tool-name negotiation; the extracted wrapper could not resolve it.

* test(gateway): await the shared async node tunnel manager fixture

Main moved createManager into node-worker-tunnel.test-support.ts as an async helper (#160415); the launch-vocabulary lifecycle test still called it synchronously.
2026-09-28 21:58:24 +00:00
Peter Steinberger
0b44647a90
fix(scripts): accept the CI workflow's fail-fast expression in admin landing (#160721) 2026-09-28 14:54:22 -07:00
Josh Avant
a19ae3eda4
fix(qa): unblock isolated harness tool and evidence checks (#160527)
* fix(qa): recognize settled subagent batches

Share requester-settle wake recognition between mock input classification and completion handling. Accept the current batch wording while retaining installed-candidate session wording and the existing provenance check.

The real catalog-only handoff spawned and completed a child on the baseline, but QA ignored its completion and timed out. Two focused regressions failed before this fix. The real handoff now passes, all 70 owner/sibling tests pass, and check-changed passes. Standalone changed-test wall times: input 3.21s; handoff 33.76s.

* fix(qa): admit canonical repository checkpoint commands

* fix(ui): stabilize Run Inspector evidence collection

Bind rendered evidence to the public run, execution, and selected decision
receipt identities. Add a shared collector that uses the component-owned
route model and a separate page, preserving the caller's Chat surface and
unsent draft through collection and later reload.

Cover exact run/execution selection, receipt cursor reload, back navigation,
wrong requested identity, missing receipt, and page cleanup in the existing
mock-Gateway Chromium harness. Document the collector and per-tab auth setup.

The regression failed on baseline before product edits because the rendered
run identity attribute was absent. Terminal R's optional assistant transcript
idempotency key is a distinct producer gap; this does not manufacture that
key, relabel historical cells, or change product appearance.

* fix(telegram): confine QA runtime and preserve readiness evidence

Run standard-library drivers through existing Python without UV inline-script virtual environments. Confine private runtime state, drain readiness pipes, retain structural diagnostics, and require verified teardown receipts before releasing recovery state. Preserve doctor, recovery, group, and published-upgrade callers.

Validation: native sandbox regression failed before the fix; 140 harness tests pass, and 72 final focused owner/sibling tests pass in 5.73s. Build exits 0. Broad changed checks stop on two unchanged TS2459 package-update test errors. Focused lint matches all 135 baseline findings with zero additions. No live credentials or Telegram sends.

* fix(telegram): avoid native scenario barrier watchers

* fix(qa): share production publication guard contract

* fix(qa): require named message availability in mock provider

A catalog dispatcher does not advertise every delivery tool. Require the
exact message declaration and a usable invocation surface, including trusted
system/developer instruction carriers. Finish with the recovered child result
when message is absent.

Prove the absent-message regression through the mock HTTP provider and retain
catalog delivery, namespace, text-only fanout, and subagent handoff coverage.

* fix(qa): pass checkout roots to script scenarios

Expand the selected repoRoot separately from each scenario outputDir in the
maintained test-file runner. Prove the argument and CWD contract through a
real subprocess with external artifacts and fresh passing producer evidence.

* fix(qa): distinguish tool declarations from instruction prose

* fix(qa): bind forked-context evidence to native receipts

* fix(qa): bind repeated-ingress MCP scheduler

* fix(qa): add owned provider continuation checkpoints

Let maintained mock scenarios hold one session continuation before response
bytes and observe its request cursor and tool-call identity. Reuse the
provider request log and scenario signal/shutdown lifecycle; replacement
requests proceed normally without touching Gateway decision state.

* fix(qa): keep harness proofs within their owners

* test(qa): exclude all registered runtime consumers

* fix(qa): preserve runtime inputs and enforce checkpoint launches

* test(qa): skip runtime proofs when tools are absent

* fix(qa): complete checkpoint launcher test admission

* test(qa): compile repeated ingress child before execution
2026-09-28 16:48:47 -05:00
Kimi Yu
9190ad7c12
fix(codex): retain completed command output with Codex 0.158.0 (#160487) 2026-09-28 14:38:08 -07:00
Peter Steinberger
b779775249
fix(release): admit 2026.9.7 recovery helper and plugin scan inventory (#160717)
fix(release): admit 2026.9.7 recovery helper and plugin scan inventory

Full Release Validation for 2026.9.7 (parent 36466521174) failed two
release-tooling checks that main never exercises:

- The installed-package verifier applied the 6 MiB root-dist cap to
  dist/package-update-activation-recovery.mjs, the ~66 MB sealed
  self-contained recovery helper added by #158491. Treat it like the other
  self-contained bundles (80 MiB cap, no reparse), as #155877 did for the
  sqlite-store worker. Every main-based FRV since #158491 failed here.
- plugin-npm-security-scan had no reviewed inventory for release/2026.9.7.
  Freeze the 9.6 counts, record four post-9.6 test-file-only findings
  (codex auth-refresh harness spawn, feishu proxy env test, imessage
  client test #156334, signal socket probe count after #159648) in the
  current inventory, and map release/2026.9.7 to it, as #153372 did for 9.6.
2026-09-28 14:36:10 -07:00
Peter Steinberger
8857192115
fix(test): include all runners for package directories (#160653)
* fix(test): run package directory tests in their owning projects

* ci: refresh package directory checks after planner repair
2026-09-28 14:36:06 -07:00
Irish Joseph
68a9ce0324
fix(browser): read Windows browser version via environment data so spaced paths stop breaking the PowerShell probe (#160432)
* fix(browser): read Windows browser version without breaking spaced paths in PowerShell probe

Windows PowerShell appends any extra -Command argument to the script text,
so passing the executable path as an argument made the version probe fail
with a ParserError for any path containing a space (including the default
Chrome install path), and the child stderr leaked into doctor --json.
Pass the path via environment data and capture probe stdout only.

* test(browser): prove Windows version probe on spaced paths natively

* test(browser): keep failed Windows metadata probes quiet

---------

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-29 02:51:17 +05:30
Peter Steinberger
de73318f5d
test(agents,plugins,ui): remove low-value tests (batch d087) (#160605)
* test(discord): deslop t0071 tests

* test(plugins): deslop t0093 tests

* test(ui): deslop t0089 tests

* test(cron): deslop t0087 tests

* test(cli): deslop t0104 tests

* test(codex): deslop t0096 tests

* test(codex): deslop t0105 tests

* test(agents): deslop t0097 tests

* test(agents): deslop t0098 tests

* test: preserve independent contracts in batch d087

* test(agents): preserve independent tool-wrapper fixtures

* test: fix typed lint in batch d087 fixtures
2026-09-28 21:15:35 +00:00
Dallin Romney
cb8e8c7324
fix(release): scan trusted release-tool dependencies (#160694)
* fix(release): scan trusted release-tool locks

* test(release): refresh Vercel lock integrity

* test(e2e): include first-hop timing fixture
2026-09-28 14:09:00 -07:00